v1.99.0-rc.1

Pre-release

BerriAI/litellmv1.99.0-rc.1Aug 23, 2026by github-actions[bot]

AI Summary

This release focuses on spend tracking accuracy, adding router_model_name to auto-routed responses, and fixing Bedrock Mantle context overflow errors. It also introduces a per-key Savings tab in the UI and updates pricing for Cognition and Bedrock models.

Key Highlights

  • Improved spend tracking with per-request auto-router savings.
  • Added router_model_name to auto-routed response bodies.
  • Fixed Bedrock Mantle context overflow and response header forwarding.
  • New per-key Savings tab in the Admin UI.
  • Pricing updates for Cognition swe-1.7 and Bedrock gpt-5.6.

New Features

  • Per-key Savings tab in key detail page
  • Router model name in auto-routed responses
  • Spend tracking for auto-router savings
  • Bedrock Mantle validation fixes
  • Cognition swe-1.7 pricing

Full Release Notes

## Verify Docker Image Signature

All LiteLLM Docker images are signed with [cosign](https://docs.sigstore.dev/cosign/overview/). Every release is signed with the same key introduced in [commit `0112e53`](https://github.com/BerriAI/litellm/commit/0112e53046018d726492c814b3644b7d376029d0).

**Verify using the pinned commit hash (recommended):**

A commit hash is cryptographically immutable, so this is the strongest way to ensure you are using the original signing key:

```bash
cosign verify \
  --key https://raw.githubusercontent.com/BerriAI/litellm/0112e53046018d726492c814b3644b7d376029d0/cosign.pub \
  ghcr.io/berriai/litellm:v1.99.0-rc.1
```

**Verify using the release tag (convenience):**

Tags are protected in this repository and resolve to the same key. This option is easier to read but relies on tag protection rules:

```bash
cosign verify \
  --key https://raw.githubusercontent.com/BerriAI/litellm/v1.99.0-rc.1/cosign.pub \
  ghcr.io/berriai/litellm:v1.99.0-rc.1
```

Expected output:

```
The following checks were performed on each of these signatures:
  - The cosign claims were validated
  - The signatures were verified against the specified public key
```

---
## What's Changed
* fix(scim): match group members by SSO identity or email before creating a placeholder by @devin-ai-integration[bot] in https://github.com/BerriAI/litellm/pull/37686
* fix(cost): match streamed Messages usage cost to the recorded spend by @devin-ai-integration[bot] in https://github.com/BerriAI/litellm/pull/35114
* test: enforce PT012 so a pytest.raises block cannot hide dead assertions by @ryan-crabbe-berri in https://github.com/BerriAI/litellm/pull/37748
* fix(streaming): price partial-stream spend rows at the real model and keep prompt and cache fields by @mateo-berri in https://github.com/BerriAI/litellm/pull/37734
* fix(fal_ai): price gpt-image-2 per size and quality from request params by @mateo-berri in https://github.com/BerriAI/litellm/pull/37751
* test: enforce PT011 and PT014 so a broad pytest.raises cannot pass on the wrong error by @ryan-crabbe-berri in https://github.com/BerriAI/litellm/pull/37769
* fix: populate team member emails missing from the roster snapshot by @ryan-crabbe-berri in https://github.com/BerriAI/litellm/pull/37759
* fix(bedrock): forward provider response headers on chat completions by @devin-ai-integration[bot] in https://github.com/BerriAI/litellm/pull/37003
* fix(cognition): price swe-1.7 at the standard tier, add swe-1.7-lightning by @mateo-berri in https://github.com/BerriAI/litellm/pull/37763
* fix(sagemaker_chat): send the inference component header and honor hf_model_name by @mateo-berri in https://github.com/BerriAI/litellm/pull/37766
* fix(spend-tracking): hash raw api keys before persisting to spend logs by @nitishagar in https://github.com/BerriAI/litellm/pull/30736
* fix(proxy): make per-model budgets track spend, enforce, and report the same counter by @yassin-berriai in https://github.com/BerriAI/litellm/pull/37736
* fix(spend): bound each spend-log write statement by row count as well as bytes by @yassin-berriai in https://github.com/BerriAI/litellm/pull/37758
* feat(proxy): add router_model_name to auto-routed response bodies by @tin-berri in https://github.com/BerriAI/litellm/pull/37725
* fix(responses): preserve Bedrock Mantle validation errors by @uuzzrm in https://github.com/BerriAI/litellm/pull/36580
* fix: omit thinking.type=disabled for always-on thinking Claude models by @tin-berri in https://github.com/BerriAI/litellm/pull/37510
* test(lint): clear the two PT011/PT012 violations left on the test tree by @tin-berri in https://github.com/BerriAI/litellm/pull/37864
* fix(proxy): read batch records the same way the upload validation does by @yucheng-berri in https://github.com/BerriAI/litellm/pull/37776
* fix(proxy): scan batch records with the content hooks that are not guardrails by @yucheng-berri in https://github.com/BerriAI/litellm/pull/37786
* fix(ptu): require an operator-declared id on a config.yaml reservation by @yucheng-berri in https://github.com/BerriAI/litellm/pull/37794
* fix(ptu): never retract a flat charge for a deployment the run cannot see by @yucheng-berri in https://github.com/BerriAI/litellm/pull/37793
* chore(codeowners): own the proxy-extras migrations directory by @yuneng-berri in https://github.com/BerriAI/litellm/pull/37876
* revert(spend-logs): drop the endTime backfill migration for spend log timestamps by @yuneng-berri in https://github.com/BerriAI/litellm/pull/37875
* fix(mcp): deny the interactive dcr_bridge authorize for a user without server access by @tin-berri in https://github.com/BerriAI/litellm/pull/37865
* fix(responses): map Bedrock Mantle context overflow to ContextWindowExceededError by @yassin-berriai in https://github.com/BerriAI/litellm/pull/37862
* test: enforce F811 so a duplicate definition cannot silently replace the first by @ryan-crabbe-berri in https://github.com/BerriAI/litellm/pull/37878
* fix(guardrails): run policy pipelines when the caller sends its own metadata (/v1/messages, Claude Code) by @devin-ai-integration[bot] in https://github.com/BerriAI/litellm/pull/36889
* fix(vertex_ai): convert messages to contents in gemini count_tokens by @sailikhithk in https://github.com/BerriAI/litellm/pull/36981
* fix(router): don't log 'Could not identify azure model' when the deployment name resolves from the cost map by @tin-berri in https://github.com/BerriAI/litellm/pull/37869
* fix(ui): drive auto-router usage from the shared cost-optimization time picker by @tin-berri in https://github.com/BerriAI/litellm/pull/37871
* feat(ui): add per-key Savings tab to key detail page by @tin-berri in https://github.com/BerriAI/litellm/pull/37693
* fix(redis): reset only the failed node on a cluster client timeout, not the whole client by @yassin-berriai in https://github.com/BerriAI/litellm/pull/37863
* test: reject assertions on a caught error inside except (ruff PT017) by @ryan-crabbe-berri in https://github.com/BerriAI/litellm/pull/37887
* fix(mcp): strip root_path before matching the per-server MCP route spelling by @tin-berri in https://github.com/BerriAI/litellm/pull/35576
* feat(spend): surface per-request auto-router savings to logging callbacks by @tin-berri in https://github.com/BerriAI/litellm/pull/37894
* fix(proxy): group Codex turns under one session id by @mubashir1osmani in https://github.com/BerriAI/litellm/pull/37895
* fix(logging): preserve uvicorn color_message args during secret redaction by @mubashir1osmani in https://github.com/BerriAI/litellm/pull/37122
* feat(llm): add moonshot/kimi-k3 to model prices and context window map by @tin-berri in https://github.com/BerriAI/litellm/pull/37552
* feat(providers): add SCX.ai as a JSON-configured OpenAI-compatible provider by @bhuvan2134686 in https://github.com/BerriAI/litellm/pull/34752
* feat(bedrock): serve gpt-5.6 cross-region inference profiles on bedrock runtime by @longwind48 in https://github.com/BerriAI/litellm/pull/37821
* fix(batches): decode model-encoded output file id so completed batches book spend by @mateo-berri in https://github.com/BerriAI/litellm/pull/37573
* fix(model_prices): correct gemini-3.1-flash-lite-image capabilities and dedupe its entries by @emerzon in https://github.com/BerriAI/litellm/pull/36849
* test: say whether a match= pattern is a regex or a literal (ruff RUF043) by @ryan-crabbe-berri in https://github.com/BerriAI/litellm/pull/37901
* fix(mcp): resolve admin OAuth sessions to the same server set the connect page shows by @tin-berri in https://github.com/BerriAI/litellm/pull/37900
* fix(anthropic_messages): gate sampling params on /v1/messages like /chat/completions by @tin-berri in https://github.com/BerriAI/litellm/pull/37868
* test(e2e): pin the openai websocket passthrough prefixes by @mateo-berri in https://github.com/BerriAI/litellm/pull/37903
* fix(model-costs): apply GPT-5.6 Sol promotional pricing cut by @devin-ai-integration[bot] in https://github.com/BerriAI/litellm/pull/37880
* fix(pricing): add undated azure aliases for gpt-audio-mini and gpt-realtime-mini by @tin-berri in https://github.com/BerriAI/litellm/pull/37867
* fix(anthropic): resolve the provider exactly once on /v1/messages by @tin-berri in https://github.com/BerriAI/litellm/pull/37757
* test(e2e): record and replay the non-streaming provider flows by @mateo-berri in https://github.com/BerriAI/litellm/pull/37909
* test(mcp): retire the last file of the dead tests/litellm mirror by @yuneng-berri in https://github.com/BerriAI/litellm/pull/37824
* perf(ci): cache the Rust build the unit shards compile from scratch by @yuneng-berri in https://github.com/BerriAI/litellm/pull/37795
* test(proxy): pin the request-body rules `proxy/_types.py` enforces by @yuneng-berri in https://github.com/BerriAI/litellm/pull/37811
* test(main): pin what a streamed response costs, end to end by @yuneng-berri in https://github.com/BerriAI/litellm/pull/37812
* test(proxy): pin what a failed request records as usage and spend by @yuneng-berri in https://github.com/BerriAI/litellm/pull/37813
* test(llm_http_handler): pin the websocket and callback gates the request path branches on by @yuneng-berri in https://github.com/BerriAI/litellm/pull/37814
* fix(ci): stop the mutation report publishing a score it never measured by @yuneng-berri in https://github.com/BerriAI/litellm/pull/37825
* test(bedrock): let monkeypatch own bedrock_request_metadata_fields by @yuneng-berri in https://github.com/BerriAI/litellm/pull/37840
* test(audit-logs): let monkeypatch own the audit log and s3 callback globals by @yuneng-berri in https://github.com/BerriAI/litellm/pull/37842
* test: unwind environment writes in tests/test_litellm with monkeypatch by @yuneng-berri in https://github.com/BerriAI/litellm/pull/37806
* test(datadog): restore an empty DD_API_KEY instead of unsetting it by @yuneng-berri in https://github.com/BerriAI/litellm/pull/37832
* test(cost-calc): stop 182 global writes leaking out of the cost-calc suites by @yuneng-berri in https://github.com/BerriAI/litellm/pull/37815
* test(key-management): unwind the global writes the key tests scaffold around by @yuneng-berri in https://github.com/BerriAI/litellm/pull/37822
* test(callbacks): unwind the callbacks global the policy engine and realtime tests scaffold around by @yuneng-berri in https://github.com/BerriAI/litellm/pull/37826
* test(guardrails): stop the onyx and prompt security tests leaking env vars on failure by @yuneng-berri in https://github.com/BerriAI/litellm/pull/37828
* test(e2e): bypass the proxy response cache in the mid-conversation system and fallback tests by @yuneng-berri in https://github.com/BerriAI/litellm/pull/37915
* docs(terraform/provider): the provider now ships at the LiteLLM version by @yuneng-berri in https://github.com/BerriAI/litellm/pull/37912
* test: stop the zai tests from leaking env and litellm globals into the session by @yuneng-berri in https://github.com/BerriAI/litellm/pull/37834
* test(s3): stop the logger tests leaking s3_callback_params on failure by @yuneng-berri in https://github.com/BerriAI/litellm/pull/37831
* test(http-handler): drop the save/restore scaffolding around litellm globals by @yuneng-berri in https://github.com/BerriAI/litellm/pull/37839
* test(interactions): drop the save/restore scaffolding around the legacy flag by @yuneng-berri in https://github.com/BerriAI/litellm/pull/37841
* test(bedrock): drop the leftover set_verbose from the embedding tests by @yuneng-berri in https://github.com/BerriAI/litellm/pull/37844
* test: drop the leftover set_verbose from eleven test files by @yuneng-berri in https://github.com/BerriAI/litellm/pull/37845
* fix(ptu): warn when config.yaml declares PTU while attribution is off by @yucheng-berri in https://github.com/BerriAI/litellm/pull/37898
* test(anthropic): let monkeypatch own litellm.callbacks in the cache control tests by @yuneng-berri in https://github.com/BerriAI/litellm/pull/37847
* test(router): let monkeypatch own expose_router_debug_in_errors by @yuneng-berri in https://github.com/BerriAI/litellm/pull/37848
* fix(proxy): stop forwarding a client Anthropic OAuth token to Bedrock and Vertex by @mateo-berri in https://github.com/BerriAI/litellm/pull/37905
* chore: bump litellm-enterprise 0.1.58 -> 0.1.59, litellm-proxy-extras 0.4.88 -> 0.4.89 by @yuneng-berri in https://github.com/BerriAI/litellm/pull/37939
* test: unshadow the module handles the F811 sweep left behind by @yuneng-berri in https://github.com/BerriAI/litellm/pull/37914
* test: drop the cwd-relative sys.path.insert calls from the test suite by @yuneng-berri in https://github.com/BerriAI/litellm/pull/37802
* chore(codeowners): add yuneng-berri as owner of the CODEOWNERS file by @yuneng-berri in https://github.com/BerriAI/litellm/pull/37944
* fix(files): list and page unscoped managed files locally by @rimysore in https://github.com/BerriAI/litellm/pull/37855
* fix(otel): emit LLM Call spans for speech, image, moderation, ocr and transcription by @devin-ai-integration[bot] in https://github.com/BerriAI/litellm/pull/37752
* fix(ui): keep completion-mode models in the playground chat dropdown by @yuneng-berri in https://github.com/BerriAI/litellm/pull/37954
* fix(types): silence pydantic ReadOnly warning on StandardLoggingRoutingDecision by @mateo-berri in https://github.com/BerriAI/litellm/pull/37949
* fix(websearch_interception): end the turn when the agentic loop hits its ceiling by @mateo-berri in https://github.com/BerriAI/litellm/pull/37911
* fix(proxy): keep a failed prisma generate from failing the migration entrypoint by @yuneng-berri in https://github.com/BerriAI/litellm/pull/37947
* fix(docker): bump wolfi-base digest for busybox 1.38.0-r1 and openssl 3.6.3-r5 by @yuneng-berri in https://github.com/BerriAI/litellm/pull/37950
* fix: don't retire a completed batch from cost recovery while output_file_id is lagging by @syahra712 in https://github.com/BerriAI/litellm/pull/37715
* fix(guardrails): pii/pci guardrail masking gaps in SpendLogs, debug logs, and logging_only response by @yassin-berriai in https://github.com/BerriAI/litellm/pull/37965
* fix(proxy): make /team/member_delete's four cleanups atomic by @yassin-berriai in https://github.com/BerriAI/litellm/pull/37959
* fix(auth): resolve team object_permission independently in the unresolvable-team fallback by @yassin-berriai in https://github.com/BerriAI/litellm/pull/37960
* fix(auth): stop the team fallback from widening model access by @yassin-berriai in https://github.com/BerriAI/litellm/pull/37962
* fix(prometheus): fold auth/pre-call time into litellm_request_total_latency_metric by @yassin-berriai in https://github.com/BerriAI/litellm/pull/37958
* fix(responses): mint Responses API item IDs in the completion bridge by @mateo-berri in https://github.com/BerriAI/litellm/pull/37946
* test(e2e): harden the suite against response-cache cross-talk, slow providers and single upstream blips by @yuneng-berri in https://github.com/BerriAI/litellm/pull/37957
* fix(proxy): omit litellm_batch_guardrail when no guardrail acted by @yucheng-berri in https://github.com/BerriAI/litellm/pull/37964
* fix(responses-bridge): preserve reasoning input items and signed thinking blocks by @harryzhou2000 in https://github.com/BerriAI/litellm/pull/36355
* chore: rebuild Admin UI bundle by @yuneng-berri in https://github.com/BerriAI/litellm/pull/37976
* test: add regression coverage for twelve closed issues by @mateo-berri in https://github.com/BerriAI/litellm/pull/37974
* chore(ci): promote internal staging to main by @yuneng-berri in https://github.com/BerriAI/litellm/pull/37913

## New Contributors
* @uuzzrm made their first contribution in https://github.com/BerriAI/litellm/pull/36580
* @sailikhithk made their first contribution in https://github.com/BerriAI/litellm/pull/36981
* @bhuvan2134686 made their first contribution in https://github.com/BerriAI/litellm/pull/34752
* @longwind48 made their first contribution in https://github.com/BerriAI/litellm/pull/37821
* @syahra712 made their first contribution in https://github.com/BerriAI/litellm/pull/37715
* @harryzhou2000 made their first contribution in https://github.com/BerriAI/litellm/pull/36355

**Full Changelog**: https://github.com/BerriAI/litellm/compare/v1.99.0-dev.2...v1.99.0-rc.1