0.0.4-gptq-llama-triton
1b5d/llm-api0.0.4-gptq-llama-tritonJun 16, 2023by 1b5d
AI Summary
This release adds Triton GPU kernel support for improved performance and restructures model directories to prevent configuration conflicts when switching between models.
Key Highlights
- Separated models into their own subdirectories to prevent overriding configs
- Added Triton support
New Features
- Model subdirectory isolation
- Triton GPU kernel support
Full Release Notes
- Separate models in their own sub directories to prevent overriding configs when changing models - Add triton support