v0.5.20

ictnlp/LLaMA-Omniv0.5.20Jul 7, 2026by decolua

AI Summary

This release introduces a per-model thinking level picker that appends a suffix to model names for forced reasoning effort across providers. It includes various fixes for reasoning content handling, token counting, and new Farsi language support.

Key Highlights

  • Thinking level picker with suffix support for forced reasoning effort
  • Farsi (fa) language support added via i18n
  • RTK adds JS-native git-log filter
  • Fixes for Claude reasoning effort and Volcengine-ark max_tokens

New Features

  • Thinking level picker on provider page
  • Farsi language support
  • RTK git-log filter
  • Caveman targeted style rules

Full Release Notes

Features:
- Thinking: per-model thinking level picker on provider page, appends (level) suffix to copied model names for forced reasoning effort across all formats (openai, claude, gemini, deepseek, kimi, qwen, zai, minimax, hunyuan, step)
- RTK: add JS-native git-log filter (#2423)
- Caveman: add targeted upstream-aligned style rules (#2424)
- i18n: add Farsi (fa) language support (#2385)

Fixes:
- Thinking: strip (level) suffix from upstream body.model so providers no longer reject requests
- Translator: preserve developer instructions in openai-responses conversion (#2434)
- count_tokens: count structured Anthropic blocks (#2419)
- Volcengine-ark: clamp GLM-5 max_tokens to model output ceiling (#2428)
- Kimi: normalize reasoning_effort to backend enum (#2427)
- Claude: reconcile max_tokens vs thinking budget and lift per-model ceiling (#2381)
- Kiro: deliver system prompt natively, add Opus 4.5/4.7/4.8, tolerate dash version ids (#2366)
- Headroom: proxy dashboard through app (#2372)
- MITM: recover from stale lock file on server start

Docs:
- README: swap in Vietnamese tutorial video; add English and Urdu/Hindi tutorials (#2305)
- Add CLAUDE.md guidance for Claude Code (#2354)