v0.5.18

ed-donner/llm_engineeringv0.5.18Jul 3, 2026by decolua

AI Summary

Adds cached token tracking with cost calculation and introduces new provider support, while resolving issues with Claude passthrough and IdC routing.

Key Highlights

  • Cached token tracking and cost calculation
  • New provider support (ClinePass, NVIDIA models)
  • Fixes for Claude passthrough and IdC auth routing
  • Improvements to Kimi and Kimchi handling

New Features

  • Usage: track cached tokens + correct input/output/cache cost
  • ClinePass: add provider support
  • NVIDIA: add new models and capabilities

Full Release Notes

Features:
- Usage: track cached tokens + correct input/output/cache cost (#2209)
- Codex: show reset credit expiry details (#2290)
- NVIDIA: add new models and capabilities
- ClinePass: add provider support

Fixes:
- Usage: dedupe streaming request-details log entries
- Claude: drop foreign thinking signatures in passthrough
- Prevent non-SSE stream pipe crash and cross-IdP account overwrites (#2244)
- Kiro: route IdC auth to regional CodeWhisperer surface (#2297)
- Kiro: add Claude Sonnet 5 model support (#2264)
- Xiaomi-tokenplan: region selector, key validation, multi-connection (#2251)
- Translator: strict Anthropic content block compliance (#2225)
- Kimchi: strip reasoning_content echo to bound multi-turn input tokens
- Kimchi: bump User-Agent to kimchi/0.1.40 (#2256)
- Codebuddy-cn: strip empty tool_calls arrays to preserve reasoning
- Antigravity: preserve Claude tool delta index (#2223)
- MITM: generate root CA on server startup (#2228)