0.0.1-gptq-llama-cuda
1b5d/llm-api0.0.1-gptq-llama-cudaApr 23, 2023by 1b5d
AI Summary
Initial release providing support for running Llama-based model inference on GPU using GPTQ-for-llama quantization.
Key Highlights
- Initial release
- Llama-based model support
- GPU inference capability
- GPTQ-for-llama integration
New Features
- GPU inference for Llama-based models
- GPTQ quantization support
Full Release Notes
Support Llama based models inference on GPU using GPTQ-for-llama