v1.98.0
BerriAI/litellmv1.98.0Aug 23, 2026by yuneng-berri
AI Summary
This release introduces comprehensive Provisioned Throughput (PTU) configuration capabilities, adds vector store index listing, and enhances rate limiting with configurable estimated output tokens.
Key Highlights
- Full PTU configuration support including flat costs and daily rollups.
- New rate limiting feature with configurable estimated output tokens.
- Vector store index listing endpoint added to the proxy.
- Proxy SSE keepalive heartbeat to prevent load-balancer timeouts.
- Guardrail load isolation to prevent cascading failures.
New Features
- PTU flat cost configuration and attribution
- Rate limiting with estimated output tokens
- Vector store index listing endpoint
- SSE keepalive heartbeat
- Guardrail load isolation
- SAML SSO detection
Full Release Notes
## Verify Docker Image Signature
All LiteLLM Docker images are signed with [cosign](https://docs.sigstore.dev/cosign/overview/). Every release is signed with the same key introduced in [commit `0112e53`](https://github.com/BerriAI/litellm/commit/0112e53046018d726492c814b3644b7d376029d0).
**Verify using the pinned commit hash (recommended):**
A commit hash is cryptographically immutable, so this is the strongest way to ensure you are using the original signing key:
```bash
cosign verify \
--key https://raw.githubusercontent.com/BerriAI/litellm/0112e53046018d726492c814b3644b7d376029d0/cosign.pub \
ghcr.io/berriai/litellm:v1.98.0
```
**Verify using the release tag (convenience):**
Tags are protected in this repository and resolve to the same key. This option is easier to read but relies on tag protection rules:
```bash
cosign verify \
--key https://raw.githubusercontent.com/BerriAI/litellm/v1.98.0/cosign.pub \
ghcr.io/berriai/litellm:v1.98.0
```
Expected output:
```
The following checks were performed on each of these signatures:
- The cosign claims were validated
- The signatures were verified against the specified public key
```
---
## What's Changed
* fix(bedrock): drop toolSpec.strict for Claude Sonnet 5 on Converse by @kr0k in https://github.com/BerriAI/litellm/pull/33196
* fix(batches): attribute Vertex passthrough batch cost to key/team/tags by @yucheng-berri in https://github.com/BerriAI/litellm/pull/34456
* docs: rewrite the CLAUDE.md comment rule with explicit exceptions by @devin-ai-integration[bot] in https://github.com/BerriAI/litellm/pull/36301
* fix(proxy): scope file list pagination cursors to the caller by @devin-ai-integration[bot] in https://github.com/BerriAI/litellm/pull/36093
* fix(proxy): skip prisma-dependent hooks when no database is attached by @mateo-berri in https://github.com/BerriAI/litellm/pull/36273
* fix(proxy): report has_more false on caller-scoped file list pages by @mateo-berri in https://github.com/BerriAI/litellm/pull/36326
* fix(proxy): restore management_v1 query-param validation under fastapi>=0.140.7 by @HuanQian571 in https://github.com/BerriAI/litellm/pull/35773
* fix(proxy): stop /{provider}/v1/files from capturing /openai_passthrough by @devin-ai-integration[bot] in https://github.com/BerriAI/litellm/pull/36092
* chore(typing): remove 914 basedpyright Any errors across 16 hotspot files by @mateo-berri in https://github.com/BerriAI/litellm/pull/36386
* fix(router): keep batch fallbacks inside the model group that owns the file by @mateo-berri in https://github.com/BerriAI/litellm/pull/36181
* feat(ptu): configure provisioned-throughput flat cost on a model deployment by @yucheng-berri in https://github.com/BerriAI/litellm/pull/35341
* docs: clarify the CLAUDE.md comment exceptions are any-of by @devin-ai-integration[bot] in https://github.com/BerriAI/litellm/pull/36421
* docs: replace the Changes PR template section with Caveats by @devin-ai-integration[bot] in https://github.com/BerriAI/litellm/pull/36423
* fix(bedrock): enable native structured output for GLM 5 and DeepSeek V3.2 by @alexshtf in https://github.com/BerriAI/litellm/pull/35669
* feat(ptu): daily rollup writes per-model PTU flat cost by active hour by @yucheng-berri in https://github.com/BerriAI/litellm/pull/35343
* feat(logging): add opt-in session_id and trace_id correlation to JSON log records via contextvars by @deepanshululla in https://github.com/BerriAI/litellm/pull/34418
* feat(ptu): surface PTU flat cost on the daily activity read path by @yucheng-berri in https://github.com/BerriAI/litellm/pull/35391
* feat(router): add per-deployment allowed_fails_policy and cooldown_time override support by @deepanshululla in https://github.com/BerriAI/litellm/pull/34416
* feat(ptu): add PTU inputs to the model form and flat cost to the Usage page by @yucheng-berri in https://github.com/BerriAI/litellm/pull/35393
* fix(cost): price dict-shaped image input token details at the image rate by @vairodp in https://github.com/BerriAI/litellm/pull/33490
* fix(model_prices): refresh deprecation dates, correct xAI pricing and add missing provider models by @devin-ai-integration[bot] in https://github.com/BerriAI/litellm/pull/36403
* feat(ptu): gate PTU flat-cost attribution behind an opt-in env var by @yucheng-berri in https://github.com/BerriAI/litellm/pull/36138
* ci: cache Prisma CLI and engine binaries, split test timeout from setup by @mateo-berri in https://github.com/BerriAI/litellm/pull/36417
* feat(rate limiting): configurable estimated output tokens per key, team and model by @yassin-berriai in https://github.com/BerriAI/litellm/pull/36143
* fix(ui): hide admin-only Logs tabs from roles that cannot call their endpoints by @yuneng-berri in https://github.com/BerriAI/litellm/pull/36333
* test(proxy): guard management_v1 against fastapi names removed in supported releases by @yuneng-berri in https://github.com/BerriAI/litellm/pull/36336
* fix(ui): gate policy and prompt lookups on an admin capability by @yuneng-berri in https://github.com/BerriAI/litellm/pull/36335
* build(deps): bump pypdf to 6.15.0 to clear osv-scan by @devin-ai-integration[bot] in https://github.com/BerriAI/litellm/pull/36350
* fix(proxy): isolate guardrail load failures per row by @yucheng-berri in https://github.com/BerriAI/litellm/pull/36432
* fix(ui): gate organization and agent usage views behind capabilities by @yuneng-berri in https://github.com/BerriAI/litellm/pull/36334
* fix(reset_budget_job): atomic budget cascade with chunked reset scans by @ryan-crabbe-berri in https://github.com/BerriAI/litellm/pull/36287
* feat(proxy): add GET /v1/indexes to list vector store indexes by @ryan-crabbe-berri in https://github.com/BerriAI/litellm/pull/36289
* feat(ui): show vector store indexes on the Vector Stores page by @ryan-crabbe-berri in https://github.com/BerriAI/litellm/pull/36306
* fix(proxy): treat SAML as configured in UI SSO detection by @fancybear-dev in https://github.com/BerriAI/litellm/pull/36196
* fix(bedrock): reject Anthropic server-side web_search tool with actionable error by @ryan-crabbe-berri in https://github.com/BerriAI/litellm/pull/36473
* fix(ui): open the classifier prompt editor above the edit auto-router form by @tin-berri in https://github.com/BerriAI/litellm/pull/36438
* fix(arize): trace MCP tool calls instead of crashing on CallToolResult by @yucheng-berri in https://github.com/BerriAI/litellm/pull/36453
* refactor(ui): make illegal DataTable prop combinations unrepresentable by @yuneng-berri in https://github.com/BerriAI/litellm/pull/36470
* fix(ui): scope Virtual Keys and Logs team lists to the caller by @yuneng-berri in https://github.com/BerriAI/litellm/pull/36472
* fix(ui): gate the Old Usage page behind a proxy-admin capability by @yuneng-berri in https://github.com/BerriAI/litellm/pull/36469
* docs(terraform): describe the provider release as automatic by @yuneng-berri in https://github.com/BerriAI/litellm/pull/36467
* feat(proxy): add per-deployment keepalive_seconds SSE heartbeat to prevent load-balancer timeout on long streams by @deepanshululla in https://github.com/BerriAI/litellm/pull/34423
* fix(router): cool down failed fallback deployments and correct cooldown TTL after Redis backfill by @deepanshululla in https://github.com/BerriAI/litellm/pull/35104
* perf(spend): write each daily spend batch in one upsert statement by @yassin-berriai in https://github.com/BerriAI/litellm/pull/36448
* fix(ui): gate four sidebar pages on the roles their endpoints allow by @yuneng-berri in https://github.com/BerriAI/litellm/pull/36475
* fix(ui): restore the Logs Deleted Teams tab for organization admins by @yuneng-berri in https://github.com/BerriAI/litellm/pull/36478
* fix(websearch): stop leaking interception control fields to providers by @yassin-berriai in https://github.com/BerriAI/litellm/pull/36480
* test(e2e): cover the Anthropic web_search server tool on Bedrock by @yassin-berriai in https://github.com/BerriAI/litellm/pull/36443
* fix(router): warn when a deployment's credentials contradict its provider by @yassin-berriai in https://github.com/BerriAI/litellm/pull/36486
* fix: net prompt-caching savings against the cache-write premium by @tin-berri in https://github.com/BerriAI/litellm/pull/36452
* feat(ui): deployment affinity toggle for the auto-router by @tin-berri in https://github.com/BerriAI/litellm/pull/36302
* fix(bedrock): use deployment credentials for AWS requests by @daleselaji-dev in https://github.com/BerriAI/litellm/pull/36160
* fix(anthropic): preserve midturn system corrections by @eugene-yao-zocdoc in https://github.com/BerriAI/litellm/pull/34290
* fix(email): stop duplicate legacy invitation email and fix its onboarding link by @mubashir1osmani in https://github.com/BerriAI/litellm/pull/36455
* feat(ui): show models under each tier in routing benchmark chart by @tin-berri in https://github.com/BerriAI/litellm/pull/36291
* fix(proxy): inject streaming usage cost on openai passthrough streams by @mateo-berri in https://github.com/BerriAI/litellm/pull/36503
* docs: require a user flow and live-proxy proof in bug reports by @devin-ai-integration[bot] in https://github.com/BerriAI/litellm/pull/36498
* fix(proxy): add config_updated_at audit timestamp for virtual keys by @ryan-crabbe-berri in https://github.com/BerriAI/litellm/pull/36488
* docs: require a user flow and a stuck-at proof in feature requests by @devin-ai-integration[bot] in https://github.com/BerriAI/litellm/pull/36500
* feat(router): add required-AND (&) tag prefix and allow_fail_open flag by @deepanshululla in https://github.com/BerriAI/litellm/pull/36193
* feat(proxy): per-key prompt caching toggle via enable_prompt_caching by @ryan-crabbe-berri in https://github.com/BerriAI/litellm/pull/36466
* fix(bedrock): send tool-search beta header for Haiku 4.5 on Invoke /v1/messages by @mateo-berri in https://github.com/BerriAI/litellm/pull/36502
* fix(bedrock): preserve adaptive thinking effort through the /v1/messages bridge by @mateo-berri in https://github.com/BerriAI/litellm/pull/36507
* ci: retry transient network fetch failures in lint workflow by @mateo-berri in https://github.com/BerriAI/litellm/pull/36563
* fix(ui): stub useIsOrgAdmin in UsageTab tests so useCan needs no QueryClient by @ryan-crabbe-berri in https://github.com/BerriAI/litellm/pull/36565
* fix(alerting): dedupe scheduled Slack spend reports across pods by @ryan-crabbe-berri in https://github.com/BerriAI/litellm/pull/36489
* chore(typing): clear 1.6k basedpyright Any errors across 56 files by @mateo-berri in https://github.com/BerriAI/litellm/pull/36543
* fix(bedrock): add text block to converse user messages carrying documents by @mateo-berri in https://github.com/BerriAI/litellm/pull/36499
* fix(deps): ship boto3 with the base SDK so bedrock works out of the box by @mubashir1osmani in https://github.com/BerriAI/litellm/pull/36568
* fix(model_prices): add provider-announced deprecation dates for Bedrock, Mistral, Cohere and Gemini models by @devin-ai-integration[bot] in https://github.com/BerriAI/litellm/pull/36538
* chore: bump litellm-enterprise 0.1.54 -> 0.1.55, litellm-proxy-extras 0.4.84 -> 0.4.85, litellm 1.97.0 -> 1.98.0 by @yuneng-berri in https://github.com/BerriAI/litellm/pull/36577
* fix(bedrock_guardrails): skip ApplyGuardrail when there is no content to scan by @yucheng-berri in https://github.com/BerriAI/litellm/pull/36441
* fix(e2e): assert on the gen-AI span that served the stream, not the span count by @yassin-berriai in https://github.com/BerriAI/litellm/pull/36582
* test(e2e): harden vendor API coverage by @mubashir1osmani in https://github.com/BerriAI/litellm/pull/34557
* test(e2e): add reproducers for passthrough and model budget gaps by @mubashir1osmani in https://github.com/BerriAI/litellm/pull/34657
* test(e2e): cover google-native generateContent framing and prometheus queue time by @mubashir1osmani in https://github.com/BerriAI/litellm/pull/34650
* chore(ci): promote internal staging to main by @tin-berri in https://github.com/BerriAI/litellm/pull/36560
* feat(router): make routing groups callable as virtual models and list them in /v1/models by @tin-berri in https://github.com/BerriAI/litellm/pull/36519
* fix(xai): bill web_search from server_side_tool_usage_details by @geraint0923 in https://github.com/BerriAI/litellm/pull/30817
* fix(responses): init completed_response on bridge streaming iterator (#35411) by @devin-ai-integration[bot] in https://github.com/BerriAI/litellm/pull/35413
* fix(batches): attribute Anthropic passthrough batch cost to the creating key, team and tags by @yucheng-berri in https://github.com/BerriAI/litellm/pull/36468
* feat(dashscope): add latest Model Studio models to the cost map by @devin-ai-integration[bot] in https://github.com/BerriAI/litellm/pull/36496
* fix(proxy): track streamed passthrough Responses cost by @william-xue in https://github.com/BerriAI/litellm/pull/36529
* fix(model_prices): advertise native structured output on every Bedrock DeepSeek V3.2 and GLM 5 id by @yuneng-berri in https://github.com/BerriAI/litellm/pull/36597
* test(bedrock): repoint live Claude tests off the retired Claude 3 Sonnet by @yuneng-berri in https://github.com/BerriAI/litellm/pull/36600
* fix(anthropic): preserve speed=fast in usage for /v1/messages and pass-through by @devin-ai-integration[bot] in https://github.com/BerriAI/litellm/pull/36447
* fix(proxy): forward resolved provider and deployment pricing in /cost/estimate by @devin-ai-integration[bot] in https://github.com/BerriAI/litellm/pull/35880
* feat(proxy): global SSE keepalive ping interval for OpenAI-shaped streaming routes by @devin-ai-integration[bot] in https://github.com/BerriAI/litellm/pull/36154
* fix(responses): preserve Codex namespace tool calls by @dcadenas in https://github.com/BerriAI/litellm/pull/32536
* fix(nvidia_nim): preserve image passages and stop sending top_k to /v1/ranking by @atomic in https://github.com/BerriAI/litellm/pull/34177
* fix: refactor HTTP handler initialization with client support by @Praveen11558 in https://github.com/BerriAI/litellm/pull/30952
* feat(lint): gate writable TypedDict fields with LIT012 by @mateo-berri in https://github.com/BerriAI/litellm/pull/36590
* perf(proxy): stagger scheduled background jobs across jobs and pods by @yassin-berriai in https://github.com/BerriAI/litellm/pull/36589
* test: remove four mirror test files that exercise none of their module by @yuneng-berri in https://github.com/BerriAI/litellm/pull/34635
* fix(router): stop re-applying router-selecting request tags to the routed tier's deployments by @mateo-berri in https://github.com/BerriAI/litellm/pull/36628
* test: remove tests that never execute by @yuneng-berri in https://github.com/BerriAI/litellm/pull/36681
* fix(ui): align spend and budget columns by @daniel-meismer-zocdoc in https://github.com/BerriAI/litellm/pull/35176
* test: rename tests that a later definition shadowed by @yuneng-berri in https://github.com/BerriAI/litellm/pull/36685
* fix(passthrough): carry the budget reservation into request metadata by @yassin-berriai in https://github.com/BerriAI/litellm/pull/36592
* fix(mcp): bound MCP client requests with a session read timeout by @yassin-berriai in https://github.com/BerriAI/litellm/pull/36675
* fix(proxy): log requests rejected for an unparsable body in spend logs by @yassin-berriai in https://github.com/BerriAI/litellm/pull/36673
* refactor(ui): migrate cost-optimization to shadcn by @yuneng-berri in https://github.com/BerriAI/litellm/pull/36629
* refactor(ui): migrate cost-tracking to shadcn by @yuneng-berri in https://github.com/BerriAI/litellm/pull/36631
* refactor(ui): migrate admin-panel to shadcn by @yuneng-berri in https://github.com/BerriAI/litellm/pull/36635
* refactor(ui): migrate users dashboard to shadcn by @yuneng-berri in https://github.com/BerriAI/litellm/pull/36642
* refactor(ui): migrate prompts to shadcn by @yuneng-berri in https://github.com/BerriAI/litellm/pull/36643
* refactor(ui): migrate team settings to shadcn by @yuneng-berri in https://github.com/BerriAI/litellm/pull/36641
* refactor(ui): migrate models-and-endpoints to shadcn by @yuneng-berri in https://github.com/BerriAI/litellm/pull/36648
* refactor(ui): migrate policy impact popover to shadcn by @yuneng-berri in https://github.com/BerriAI/litellm/pull/36653
* fix(proxy): expand config-defined model access groups when resolving team models for /v2/model/info by @ryan-crabbe-berri in https://github.com/BerriAI/litellm/pull/34211
* fix(batches): strip NUL bytes from passthrough batch tags before the managed object write by @yucheng-berri in https://github.com/BerriAI/litellm/pull/36688
* test(e2e-ui): verify UI mutations against the API instead of trusting the toast by @yuneng-berri in https://github.com/BerriAI/litellm/pull/36632
* fix(proxy): serialize model reconciles so concurrent model writes stop evicting each other by @yuneng-berri in https://github.com/BerriAI/litellm/pull/36687
* chore(e2e): port the compat-matrix cron publisher to tests/e2e/claude_code by @mateo-berri in https://github.com/BerriAI/litellm/pull/36465
* fix(router): never price a strategy-router alias by @tin-berri in https://github.com/BerriAI/litellm/pull/36691
* feat(model_prices): add NVIDIA Nemotron 3.5 Lightning on OpenRouter and DeepInfra by @devin-ai-integration[bot] in https://github.com/BerriAI/litellm/pull/36696
* feat(terraform/aws): make VPC, Aurora, and Redis optional by @yassin-berriai in https://github.com/BerriAI/litellm/pull/36676
* feat(ui): warn in the Admin UI when no Redis is configured by @devin-ai-integration[bot] in https://github.com/BerriAI/litellm/pull/36495
* fix(ui): show and edit key-level router settings on a virtual key by @yassin-berriai in https://github.com/BerriAI/litellm/pull/36674
* fix(router): forward auto-router alias params from the marker entry, not the first same-name deployment by @mateo-berri in https://github.com/BerriAI/litellm/pull/36626
* fix(bedrock_mantle): 1M context window and long-context pricing for GPT-5.6 Sol/Terra/Luna by @devin-ai-integration[bot] in https://github.com/BerriAI/litellm/pull/36698
* fix(model_prices): sync the Groq registry with Groq's docs by @devin-ai-integration[bot] in https://github.com/BerriAI/litellm/pull/36664
* fix(router): let untagged requests bypass a tagged pre-routing strategy on shared model names by @mateo-berri in https://github.com/BerriAI/litellm/pull/36627
* fix(spend): stop losing spend log rows when a flush is cancelled by @devin-ai-integration[bot] in https://github.com/BerriAI/litellm/pull/34826
* docs(claude): drop the @ prefix from the PR template path by @devin-ai-integration[bot] in https://github.com/BerriAI/litellm/pull/36726
* fix(langfuse): emit otel trace version and release on the keys langfuse v4 reads by @yucheng-berri in https://github.com/BerriAI/litellm/pull/36702
* test(interactions): follow Google spec drift replacing Turn with typed steps by @mateo-berri in https://github.com/BerriAI/litellm/pull/36730
* refactor(ui): migrate team detail controls to shadcn by @yuneng-berri in https://github.com/BerriAI/litellm/pull/36695
* refactor(ui): migrate guardrail and duration controls to shadcn by @yuneng-berri in https://github.com/BerriAI/litellm/pull/36693
* refactor(ui): migrate guardrails-monitor, projects, logs to shadcn by @yuneng-berri in https://github.com/BerriAI/litellm/pull/34606
* refactor(ui): migrate search and user controls to shadcn by @yuneng-berri in https://github.com/BerriAI/litellm/pull/36694
* fix(guardrails): scan and re-emit raw Anthropic SSE streams in the bedrock post-call hook by @yucheng-berri in https://github.com/BerriAI/litellm/pull/36598
* fix(helm): render nodeSelector on the migrations job by @yuneng-berri in https://github.com/BerriAI/litellm/pull/36747
* fix(langfuse): coerce header-sourced mask and trace-update steering values by @yucheng-berri in https://github.com/BerriAI/litellm/pull/36740
* refactor(ui): migrate usage tables to shared DataTable by @yuneng-berri in https://github.com/BerriAI/litellm/pull/36707
* refactor(ui): migrate guardrails monitor table to shared DataTable by @yuneng-berri in https://github.com/BerriAI/litellm/pull/36709
* refactor(ui): migrate guardrails content tables to shared DataTable by @yuneng-berri in https://github.com/BerriAI/litellm/pull/36708
* feat(gemini): day-0 pricing for gemini-3.7-flash by @mateo-berri in https://github.com/BerriAI/litellm/pull/36792
* ci: promote staging to main by @mateo-berri in https://github.com/BerriAI/litellm/pull/36725
* build(deps): bump nanoid to 3.3.18 to clear osv-scan by @devin-ai-integration[bot] in https://github.com/BerriAI/litellm/pull/36787
* fix(router): stop scoring system prompt text for code/technical complexity by @tin-berri in https://github.com/BerriAI/litellm/pull/36721
* feat(complexity_router): calibrate the classifier rubric with worked examples, selectable per router by @tin-berri in https://github.com/BerriAI/litellm/pull/36578
* fix(interactions): map step and turn history to Responses API roles and content types by @mateo-berri in https://github.com/BerriAI/litellm/pull/36733
* fix(ui): restore playground model filtering by endpoint by @mubashir1osmani in https://github.com/BerriAI/litellm/pull/36130
* fix(proxy/batches): stop forwarding custom_llm_provider twice in list and cancel by @anxkhn in https://github.com/BerriAI/litellm/pull/32813
* refactor(ui): migrate TokenFlow and JsonViewer to shadcn by @yuneng-berri in https://github.com/BerriAI/litellm/pull/36735
* feat: pre-adoption shadow eval for the auto-router (blind pairwise judge, derived state) by @tin-berri in https://github.com/BerriAI/litellm/pull/36587
* refactor(ui): migrate SimpleMessageBlock and SimpleToolCallBlock to shadcn by @yuneng-berri in https://github.com/BerriAI/litellm/pull/36737
* refactor(ui): migrate HistoryTree and CollapsibleMessage to shadcn by @yuneng-berri in https://github.com/BerriAI/litellm/pull/36738
* refactor: replace Any with precise types across responses, proxy, and llms modules by @mateo-berri in https://github.com/BerriAI/litellm/pull/36763
* refactor(ui): migrate TruncatedValue and OutputCard to shadcn by @yuneng-berri in https://github.com/BerriAI/litellm/pull/36739
* refactor(ui): migrate SectionHeader and ToolsSection to shadcn by @yuneng-berri in https://github.com/BerriAI/litellm/pull/36793
* feat(ui): migrate playground chat controls to shadcn by @mubashir1osmani in https://github.com/BerriAI/litellm/pull/36129
* feat(xai): day-0 pricing for grok-4.6 by @mateo-berri in https://github.com/BerriAI/litellm/pull/36805
* feat(ui): highlight Auto Router in the navbar announcement by @devin-ai-integration[bot] in https://github.com/BerriAI/litellm/pull/36315
* test(e2e): assert the model allow-list permits, not only denies by @yuneng-berri in https://github.com/BerriAI/litellm/pull/36823
* fix(proxy): tolerate a concurrent creator when creating spend views by @yuneng-berri in https://github.com/BerriAI/litellm/pull/36824
* fix(proxy): honor explicit null budget_duration on team and key create + clearable UI dropdowns by @ryan-crabbe-berri in https://github.com/BerriAI/litellm/pull/36699
* feat(model_prices): add meta/muse-spark-1.2 and its contributor tier by @mateo-berri in https://github.com/BerriAI/litellm/pull/36717
* fix(auth): carry team grants in lite login session tokens by @yassin-berriai in https://github.com/BerriAI/litellm/pull/36826
* feat(ui): show provider prompt cache tokens in chat response metrics by @yassin-berriai in https://github.com/BerriAI/litellm/pull/36827
* fix(auth): stop the team fallback from widening model access by @yassin-berriai in https://github.com/BerriAI/litellm/pull/36837
* fix(proxy/team): resolve member_delete cleanup by user id, not the addressed email by @yassin-berriai in https://github.com/BerriAI/litellm/pull/36839
* fix(cli): launch agents as a child process on Windows by @yassin-berriai in https://github.com/BerriAI/litellm/pull/36822
* feat(ui): shadow evals tab beside auto-router usage by @tin-berri in https://github.com/BerriAI/litellm/pull/36588
* feat(cli): make the hidden `lite` command list configurable by @yassin-berriai in https://github.com/BerriAI/litellm/pull/36816
* feat(azure_ai): add Fireworks FW model pricing on Azure AI Foundry by @emerzon in https://github.com/BerriAI/litellm/pull/35613
* fix: enable xhigh reasoning support for gpt-5.4-mini models by @emerzon in https://github.com/BerriAI/litellm/pull/26909
* feat(azure-ai): add Grok 4.3 model metadata by @emerzon in https://github.com/BerriAI/litellm/pull/27932
* feat(ui): render request metrics on the /ui/chat surface by @yassin-berriai in https://github.com/BerriAI/litellm/pull/36845
* fix(ui): stop a deselected MCP server keeping its grant on a virtual key by @yassin-berriai in https://github.com/BerriAI/litellm/pull/36840
* fix(team): sweep dangling team references and cache on team delete by @yassin-berriai in https://github.com/BerriAI/litellm/pull/36819
* fix(mcp): resolve admin OAuth sessions from any worker via DB-backed drafts by @yassin-berriai in https://github.com/BerriAI/litellm/pull/36844
* refactor(ui): migrate usage to shadcn by @yuneng-berri in https://github.com/BerriAI/litellm/pull/36834
* refactor(ui): migrate guardrails-monitor to shadcn by @yuneng-berri in https://github.com/BerriAI/litellm/pull/36838
* refactor(ui): migrate playground to shadcn by @yuneng-berri in https://github.com/BerriAI/litellm/pull/36847
* refactor(ui): migrate guardrails to shadcn by @yuneng-berri in https://github.com/BerriAI/litellm/pull/36832
* fix(batches): stop uncostable batches from starving the cost poll page by @devin-ai-integration[bot] in https://github.com/BerriAI/litellm/pull/36714
* perf(spend-logs): bound retention cleanup so one run cannot saturate the database by @yassin-berriai in https://github.com/BerriAI/litellm/pull/36594
* fix(proxy): fail config load when a callbacks entry is not dispatchable by @yassin-berriai in https://github.com/BerriAI/litellm/pull/36858
* fix(bedrock): hoist custom.defer_loading before dropping custom on invoke tools by @yassin-berriai in https://github.com/BerriAI/litellm/pull/36855
* fix(access groups): sync assigned_key_ids from the key write paths by @yassin-berriai in https://github.com/BerriAI/litellm/pull/36843
* fix(mcp): expose client HTTP headers to logging callbacks and hooks by @devin-ai-integration[bot] in https://github.com/BerriAI/litellm/pull/36724
* fix(ptu): stop per-token billing on a PTU-configured deployment by @yucheng-berri in https://github.com/BerriAI/litellm/pull/36829
* fix(ui): add nvidia riva to the model provider list by @devin-ai-integration[bot] in https://github.com/BerriAI/litellm/pull/36769
* fix(scripts): end make check with a ran/skipped summary and verdict by @mateo-berri in https://github.com/BerriAI/litellm/pull/36864
* fix(proxy): track spend for OpenAI passthrough /v1/embeddings by @lostmartian in https://github.com/BerriAI/litellm/pull/36660
* test(proxy): stop monkeypatch.undo re-planting fixture-mocked prisma_client by @mateo-berri in https://github.com/BerriAI/litellm/pull/36872
* fix(access groups): sync assigned_team_ids from the team write paths by @yassin-berriai in https://github.com/BerriAI/litellm/pull/36825
* ci: drop the CircleCI ui_build and ui_unit_tests jobs by @yuneng-berri in https://github.com/BerriAI/litellm/pull/36893
* fix(langfuse)!: source the emitted metadata blob from StandardLoggingPayload by @yucheng-berri in https://github.com/BerriAI/litellm/pull/36744
* refactor(ui): migrate Navbar off antd to shadcn by @yuneng-berri in https://github.com/BerriAI/litellm/pull/36902
* refactor(ui): migrate log details drawer off antd to shadcn by @yuneng-berri in https://github.com/BerriAI/litellm/pull/36904
* refactor(ui): migrate AI Hub off antd and tremor to shadcn by @yuneng-berri in https://github.com/BerriAI/litellm/pull/36908
* refactor(ui): move the shared dropdowns and selectors onto shadcn primitives by @yuneng-berri in https://github.com/BerriAI/litellm/pull/36924
* refactor(ui): move the root-level dashboard components onto shadcn primitives by @yuneng-berri in https://github.com/BerriAI/litellm/pull/36927
* refactor(ui): move the settings page and bulk user invite onto shadcn primitives by @yuneng-berri in https://github.com/BerriAI/litellm/pull/36936
* refactor(ui): move the cost tracking components onto shadcn primitives by @yuneng-berri in https://github.com/BerriAI/litellm/pull/36955
* ci: drop the duplicate proxy_unit_tests letter-shard workflow by @yuneng-berri in https://github.com/BerriAI/litellm/pull/36866
* refactor(ui): migrate shared common_components off antd and tremor by @yuneng-berri in https://github.com/BerriAI/litellm/pull/36910
* refactor(ui): migrate key info and permissions views off antd and tremor by @yuneng-berri in https://github.com/BerriAI/litellm/pull/36913
* feat(proxy): serve Anthropic-native /v1/models for Claude Code gateway discovery by @Ar-maan05 in https://github.com/BerriAI/litellm/pull/35455
* refactor(ui): migrate router settings and shared badges off antd and tremor by @yuneng-berri in https://github.com/BerriAI/litellm/pull/36915
* refactor(ui): move the model hub and model select onto shadcn primitives by @yuneng-berri in https://github.com/BerriAI/litellm/pull/36918
* fix(ui): keep the cost tracking removal confirmation open until it settles by @yuneng-berri in https://github.com/BerriAI/litellm/pull/36960
* refactor(ui): declare DateRangePickerValue locally instead of importing it from tremor by @yuneng-berri in https://github.com/BerriAI/litellm/pull/36962
* fix(main): an explicit provider outranks a known OpenAI model name by @FahimaGold in https://github.com/BerriAI/litellm/pull/36800
* fix(exception_mapping): bare 429 in an error body no longer outranks the status code by @FahimaGold in https://github.com/BerriAI/litellm/pull/36705
* refactor(ui): move MCP permission panels onto shadcn primitives by @yuneng-berri in https://github.com/BerriAI/litellm/pull/36964
* refactor(ui): migrate ten small dashboard files off antd and tremor by @yuneng-berri in https://github.com/BerriAI/litellm/pull/36966
* fix(proxy): force prisma recreate on postgres cached-plan error by @devin-ai-integration[bot] in https://github.com/BerriAI/litellm/pull/36428
* fix(transcription): stop a zero output rate from zeroing transcription cost by @hMED22 in https://github.com/BerriAI/litellm/pull/36914
* fix(langfuse): restrict trace steering keys to real langfuse trace fields by @yucheng-berri in https://github.com/BerriAI/litellm/pull/36862
* Revert "fix(auth): stop the team fallback from widening model access" (#36837) by @yuneng-berri in https://github.com/BerriAI/litellm/pull/36982
* fix(ui): show zeroed auto-router usage stats when a window has no sessions by @tin-berri in https://github.com/BerriAI/litellm/pull/36868
* fix(mcp): keep admin-entered oauth endpoints in management reads by @devin-ai-integration[bot] in https://github.com/BerriAI/litellm/pull/36888
* fix(ui): distinguish hosted and local vLLM in the provider dropdown by @mateo-berri in https://github.com/BerriAI/litellm/pull/36974
* fix(openai,azure): return a length-truncated 200 when the output budget fits no token by @yassin-berriai in https://github.com/BerriAI/litellm/pull/36859
* fix(proxy): always emit the Anthropic /v1/models token limits, null when unknown by @yassin-berriai in https://github.com/BerriAI/litellm/pull/36961
* feat(helm): add startupProbe and hpa.behavior to the componentized chart by @Louis-Vauterin in https://github.com/BerriAI/litellm/pull/36382
* fix(proxy): serve aggregate MCP endpoint on bare /mcp instead of 307-redirecting by @tin-berri in https://github.com/BerriAI/litellm/pull/34845
* feat(shadow_eval): add reverse-direction shadow eval jobs by @tin-berri in https://github.com/BerriAI/litellm/pull/36865
* fix(proxy): requeue Redis spend buffer transactions when the DB commit fails by @devin-ai-integration[bot] in https://github.com/BerriAI/litellm/pull/33881
* feat(search): add Nimble as a search provider by @ilchemla in https://github.com/BerriAI/litellm/pull/36347
* fix(mcp): drop caller host and configured upstream headers from logged metadata by @yucheng-berri in https://github.com/BerriAI/litellm/pull/36901
* fix(azure_ai): recognize real Search doc endpoints so teams can read/write via passthrough by @devin-ai-integration[bot] in https://github.com/BerriAI/litellm/pull/36798
* fix(anthropic): aggregate 5m/1h cache-write split across iterations path by @devin-ai-integration[bot] in https://github.com/BerriAI/litellm/pull/34860
* fix(anthropic cost): apply regional geo uplift to cached tokens by @devin-ai-integration[bot] in https://github.com/BerriAI/litellm/pull/34850
* fix(ui): match the MCP servers count badge to its sibling permission badges by @yuneng-berri in https://github.com/BerriAI/litellm/pull/36984
* fix(batches): mark terminal batch with no output file as processed in CheckBatchCost by @devin-ai-integration[bot] in https://github.com/BerriAI/litellm/pull/35360
* fix(caching): cache anthropic /v1/messages responses, including streaming by @devin-ai-integration[bot] in https://github.com/BerriAI/litellm/pull/34581
* fix(anthropic_messages): make tool_result images visible to OpenAI-compatible providers by @hMED22 in https://github.com/BerriAI/litellm/pull/34462
* feat(fireworks_ai): translate NIM/vLLM extra params to Fireworks-native args by @milesadkins in https://github.com/BerriAI/litellm/pull/35969
* fix(ui): stop the models tab strip from scrolling vertically by @yuneng-berri in https://github.com/BerriAI/litellm/pull/36993
* fix(ui): anchor chips-combobox popups to the field instead of the inner input by @yuneng-berri in https://github.com/BerriAI/litellm/pull/36995
* feat(proxy): per-component response cost headers by @erensh27 in https://github.com/BerriAI/litellm/pull/36965
* fix(cost): track OpenAI/Azure web search tool cost per call by @devin-ai-integration[bot] in https://github.com/BerriAI/litellm/pull/35286
* fix(bedrock): resolve aliases in batch file records by @daleselaji-dev in https://github.com/BerriAI/litellm/pull/36159
* fix: report real token usage on guardrail-blocked /v1/responses replies by @guptaishaan in https://github.com/BerriAI/litellm/pull/36907
* fix(proxy): requeue spend logs when the DB write fails with a transport error by @devin-ai-integration[bot] in https://github.com/BerriAI/litellm/pull/36716
* fix(cost): tiered pricing supports cache creation cost and is all-or-nothing by @devin-ai-integration[bot] in https://github.com/BerriAI/litellm/pull/36720
* fix(vertex_ai): translate /v1/embeddings batch rows to the Gemini embedding shape by @devin-ai-integration[bot] in https://github.com/BerriAI/litellm/pull/35092
* docs(claude): require ReadOnly on every TypedDict field (LIT012) by @devin-ai-integration[bot] in https://github.com/BerriAI/litellm/pull/37005
* refactor(ui): migrate access group create modal to RHF + zod + shadcn by @ryan-crabbe-berri in https://github.com/BerriAI/litellm/pull/37033
* refactor(ui): re-sync badge and skeleton onto the base-vega shadcn style by @yuneng-berri in https://github.com/BerriAI/litellm/pull/36991
* feat(ui): link user detail team names to team pages by @ryan-crabbe-berri in https://github.com/BerriAI/litellm/pull/37022
* fix(model_prices): correct DeepSeek V4 max output tokens by @devin-ai-integration[bot] in https://github.com/BerriAI/litellm/pull/36925
* fix(ui): rename models table Status column to Source by @ryan-crabbe-berri in https://github.com/BerriAI/litellm/pull/37021
* chore: bump litellm-enterprise 0.1.55 -> 0.1.56, litellm-proxy-extras 0.4.85 -> 0.4.86 by @yuneng-berri in https://github.com/BerriAI/litellm/pull/37045
* feat(proxy): gate the Global Control Plane worker registry on an enterprise license by @yassin-berriai in https://github.com/BerriAI/litellm/pull/36996
* fix(model_prices): add gemini 3.1 flash tts preview and legacy OpenAI shutdown dates by @devin-ai-integration[bot] in https://github.com/BerriAI/litellm/pull/36788
* fix(panw_prisma_airs): surface scan_id on allowed requests by @devin-ai-integration[bot] in https://github.com/BerriAI/litellm/pull/37037
* fix(model_map): flag native structured outputs on Anthropic-direct claude-sonnet-5 and claude-haiku-4-5 by @anmolg1997 in https://github.com/BerriAI/litellm/pull/35930
* fix(router): stop get_router_model_info from wiping cached pricing by @devin-ai-integration[bot] in https://github.com/BerriAI/litellm/pull/36985
* fix(redis): unwrap decorated __init__s when deriving the from_url kwargs allowlist by @anmolg1997 in https://github.com/BerriAI/litellm/pull/36654
* fix(proxy): reserve the larger declared output budget for TPM limits by @yassin-berriai in https://github.com/BerriAI/litellm/pull/37001
* fix(databricks): surface provider usage, including prompt-cache counts, in streaming chunks by @pokepoke81 in https://github.com/BerriAI/litellm/pull/36943
* fix(spend): give a batch's cost row a primary key of its own by @marty-sullivan in https://github.com/BerriAI/litellm/pull/36876
* feat: shadow eval samples /v1/messages and /v1/responses traffic by @tin-berri in https://github.com/BerriAI/litellm/pull/36830
* fix(ptu): stop a PTU deployment billing for grounded search by @yucheng-berri in https://github.com/BerriAI/litellm/pull/37043
* fix(fireworks_ai): support router slugs via routers/ prefix by @heathriel in https://github.com/BerriAI/litellm/pull/34257
* fix(bedrock): register managed-batch litellm_params so they stop leaking to the provider (internal copy of #36633) by @mateo-berri in https://github.com/BerriAI/litellm/pull/37048
* fix(bedrock): resolve the managed-batch output bucket on every path that reads it by @mateo-berri in https://github.com/BerriAI/litellm/pull/37047
* fix(bedrock): resolve the managed-batch output bucket on every path that reads it by @marty-sullivan in https://github.com/BerriAI/litellm/pull/36634
* feat(scripts): queue heavy gates behind a machine-wide slot lock by @mateo-berri in https://github.com/BerriAI/litellm/pull/36988
* feat(mcp): scope gateway session bearers to the RFC 8707 resource by @tin-berri in https://github.com/BerriAI/litellm/pull/35045
* feat(ui): direction picker and reverse-mode display for shadow evals by @tin-berri in https://github.com/BerriAI/litellm/pull/36994
* fix(guardrails): return the full PANW AIRS scan response on blocked requests by @devin-ai-integration[bot] in https://github.com/BerriAI/litellm/pull/37036
* fix(passthrough): stop forwarding client Accept-Encoding upstream by @mateo-berri in https://github.com/BerriAI/litellm/pull/37058
* fix(batches): account a managed batch's cost exactly once by @mateo-berri in https://github.com/BerriAI/litellm/pull/37050
* fix(panw_prisma_airs): scan tool call args as plain text, not a tool_event by @devin-ai-integration[bot] in https://github.com/BerriAI/litellm/pull/37038
* feat(lint): exempt TypedDict-annotated dict literals from LIT002 by @mateo-berri in https://github.com/BerriAI/litellm/pull/36869
* docs(claude): tell agents to let heavy gates queue for machine-wide slots by @devin-ai-integration[bot] in https://github.com/BerriAI/litellm/pull/37057
* test: unstick the suites CircleCI is failing on by @yuneng-berri in https://github.com/BerriAI/litellm/pull/37059
* docs(github): proof-of-fix section shows only the latest run as Before/After with nested cases by @mateo-berri in https://github.com/BerriAI/litellm/pull/37063
* test(e2e): assert provider error shape instead of pinned prose by @yuneng-berri in https://github.com/BerriAI/litellm/pull/37065
* fix(ui): de-duplicate the reset budget option and polish shadcn surfaces by @yuneng-berri in https://github.com/BerriAI/litellm/pull/37010
* chore: rebuild Admin UI bundle from litellm_internal_staging by @yuneng-berri in https://github.com/BerriAI/litellm/pull/37066
* test(e2e/ui): assert the log drawer chevrons by their lucide classes by @yuneng-berri in https://github.com/BerriAI/litellm/pull/37069
* chore(ci): promote internal staging to main by @yuneng-berri in https://github.com/BerriAI/litellm/pull/37042
* fix(ui): keep completion-mode models in the playground chat dropdown (backport to rc/1.98.0) by @yuneng-berri in https://github.com/BerriAI/litellm/pull/37955
## New Contributors
* @kr0k made their first contribution in https://github.com/BerriAI/litellm/pull/33196
* @HuanQian571 made their first contribution in https://github.com/BerriAI/litellm/pull/35773
* @alexshtf made their first contribution in https://github.com/BerriAI/litellm/pull/35669
* @vairodp made their first contribution in https://github.com/BerriAI/litellm/pull/33490
* @fancybear-dev made their first contribution in https://github.com/BerriAI/litellm/pull/36196
* @daleselaji-dev made their first contribution in https://github.com/BerriAI/litellm/pull/36160
* @eugene-yao-zocdoc made their first contribution in https://github.com/BerriAI/litellm/pull/34290
* @geraint0923 made their first contribution in https://github.com/BerriAI/litellm/pull/30817
* @william-xue made their first contribution in https://github.com/BerriAI/litellm/pull/36529
* @dcadenas made their first contribution in https://github.com/BerriAI/litellm/pull/32536
* @atomic made their first contribution in https://github.com/BerriAI/litellm/pull/34177
* @Praveen11558 made their first contribution in https://github.com/BerriAI/litellm/pull/30952
* @anxkhn made their first contribution in https://github.com/BerriAI/litellm/pull/32813
* @lostmartian made their first contribution in https://github.com/BerriAI/litellm/pull/36660
* @FahimaGold made their first contribution in https://github.com/BerriAI/litellm/pull/36800
* @Louis-Vauterin made their first contribution in https://github.com/BerriAI/litellm/pull/36382
* @ilchemla made their first contribution in https://github.com/BerriAI/litellm/pull/36347
* @milesadkins made their first contribution in https://github.com/BerriAI/litellm/pull/35969
* @erensh27 made their first contribution in https://github.com/BerriAI/litellm/pull/36965
* @guptaishaan made their first contribution in https://github.com/BerriAI/litellm/pull/36907
* @pokepoke81 made their first contribution in https://github.com/BerriAI/litellm/pull/36943
* @heathriel made their first contribution in https://github.com/BerriAI/litellm/pull/34257
**Full Changelog**: https://github.com/BerriAI/litellm/compare/v1.97.0...v1.98.0