0.0.4-gptq-llama-triton

1b5d/llm-api0.0.4-gptq-llama-tritonJun 16, 2023by 1b5d

AI Summary

This release adds Triton GPU kernel support for improved performance and restructures model directories to prevent configuration conflicts when switching between models.

Key Highlights

  • Separated models into their own subdirectories to prevent overriding configs
  • Added Triton support

New Features

  • Model subdirectory isolation
  • Triton GPU kernel support

Full Release Notes

- Separate models in their own sub directories to prevent overriding configs when changing models
- Add triton support