0.0.1-gptq-llama-cuda

1b5d/llm-api0.0.1-gptq-llama-cudaApr 23, 2023by 1b5d

AI Summary

Initial release providing support for running Llama-based model inference on GPU using GPTQ-for-llama quantization.

Key Highlights

  • Initial release
  • Llama-based model support
  • GPU inference capability
  • GPTQ-for-llama integration

New Features

  • GPU inference for Llama-based models
  • GPTQ quantization support

Full Release Notes

Support Llama based models inference on GPU using GPTQ-for-llama