v0.5.18

ictnlp/LLaMA-Omniv0.5.18Jul 3, 2026by decolua

AI Summary

This release focuses on improving usage tracking by adding cached token calculation and fixes for various providers including Claude, Kiro, and Kimchi. It also introduces new models and capabilities for NVIDIA and ClinePass.

Key Highlights

  • Cached token tracking and cost correction
  • New models and capabilities added for NVIDIA
  • ClinePass provider support added
  • Fixes for streaming request logs and SSE pipe crashes

New Features

  • Track cached tokens + cost
  • Show reset credit expiry details (Codex)
  • NVIDIA new models and capabilities
  • ClinePass provider support

Full Release Notes

Features:
- Usage: track cached tokens + correct input/output/cache cost (#2209)
- Codex: show reset credit expiry details (#2290)
- NVIDIA: add new models and capabilities
- ClinePass: add provider support

Fixes:
- Usage: dedupe streaming request-details log entries
- Claude: drop foreign thinking signatures in passthrough
- Prevent non-SSE stream pipe crash and cross-IdP account overwrites (#2244)
- Kiro: route IdC auth to regional CodeWhisperer surface (#2297)
- Kiro: add Claude Sonnet 5 model support (#2264)
- Xiaomi-tokenplan: region selector, key validation, multi-connection (#2251)
- Translator: strict Anthropic content block compliance (#2225)
- Kimchi: strip reasoning_content echo to bound multi-turn input tokens
- Kimchi: bump User-Agent to kimchi/0.1.40 (#2256)
- Codebuddy-cn: strip empty tool_calls arrays to preserve reasoning
- Antigravity: preserve Claude tool delta index (#2223)
- MITM: generate root CA on server startup (#2228)