v1.97.0
BerriAI/litellmv1.97.0Aug 16, 2026by yuneng-berri
AI Summary
A feature-rich release adding Cursor thinking support, guardrails integration, and complexity router session affinity, alongside improvements to the playground, spend dashboard, and OpenTelemetry metrics.
Key Highlights
- Support for Cursor thinking/fast model-name suffixes.
- New guardrails integration for Rubrik.
- Complexity router now supports session affinity.
- Playground gains a non-streaming response toggle.
- Spend dashboard updates showing auto-router savings.
New Features
- Cursor thinking model support
- Rubrik guardrails integration
- Complexity router session affinity
- Playground non-streaming toggle
- PTU flat cost configuration
- OTel service tier attributes
Full Release Notes
## Verify Docker Image Signature All LiteLLM Docker images are signed with [cosign](https://docs.sigstore.dev/cosign/overview/). Every release is signed with the same key introduced in [commit `0112e53`](https://github.com/BerriAI/litellm/commit/0112e53046018d726492c814b3644b7d376029d0). **Verify using the pinned commit hash (recommended):** A commit hash is cryptographically immutable, so this is the strongest way to ensure you are using the original signing key: ```bash cosign verify \ --key https://raw.githubusercontent.com/BerriAI/litellm/0112e53046018d726492c814b3644b7d376029d0/cosign.pub \ ghcr.io/berriai/litellm:v1.97.0 ``` **Verify using the release tag (convenience):** Tags are protected in this repository and resolve to the same key. This option is easier to read but relies on tag protection rules: ```bash cosign verify \ --key https://raw.githubusercontent.com/BerriAI/litellm/v1.97.0/cosign.pub \ ghcr.io/berriai/litellm:v1.97.0 ``` Expected output: ``` The following checks were performed on each of these signatures: - The cosign claims were validated - The signatures were verified against the specified public key ``` --- ## What's Changed * feat(proxy): resolve Cursor thinking/fast model-name suffixes on /cursor/chat/completions by @mateo-berri in https://github.com/BerriAI/litellm/pull/35554 * fix(team-callbacks): actually stop logging when disable_logging is called by @yucheng-berri in https://github.com/BerriAI/litellm/pull/35520 * refactor(lint): drop redundant !s f-string conversion flags and fix displaced import-group comments by @mateo-berri in https://github.com/BerriAI/litellm/pull/35546 * fix(proxy): backfill null user_email on existing users during JWT auth by @devin-ai-integration[bot] in https://github.com/BerriAI/litellm/pull/34588 * feat(playground): add non-streaming response toggle by @ryan-crabbe-berri in https://github.com/BerriAI/litellm/pull/35560 * feat(teams): apply default organization to new teams from default team settings by @ryan-crabbe-berri in https://github.com/BerriAI/litellm/pull/35540 * fix(ui): block Playground page for viewer roles on direct URL access by @ryan-crabbe-berri in https://github.com/BerriAI/litellm/pull/35676 * fix(caching): close evicted LLM clients so their connections are reclaimed by @yassin-berriai in https://github.com/BerriAI/litellm/pull/35492 * chore(deps): update brace-expansion, postcss, and gitpython to current patch releases by @yuneng-berri in https://github.com/BerriAI/litellm/pull/35692 * refactor(ui): rename the create MCP server component to PascalCase by @yuneng-berri in https://github.com/BerriAI/litellm/pull/35686 * fix(openai): drop undefined Union from owns_wrapped_http_client annotation by @ryan-crabbe-berri in https://github.com/BerriAI/litellm/pull/35706 * fix(openai): drop the undefined Union from owns_wrapped_http_client by @devin-ai-integration[bot] in https://github.com/BerriAI/litellm/pull/35704 * chore(ui): note Google's Agent Platform rename in vector store setup by @ryan-crabbe-berri in https://github.com/BerriAI/litellm/pull/28076 * fix(proxy): apply key/team router_settings.model_group_alias by @yassin-berriai in https://github.com/BerriAI/litellm/pull/35486 * feat(complexity_router): default session affinity off and expose it in the UI by @tin-berri in https://github.com/BerriAI/litellm/pull/35714 * fix(datadog): read team callback dd_* params from kwargs instead of blocked dynamic params (#35115 port) by @ryan-crabbe-berri in https://github.com/BerriAI/litellm/pull/35687 * refactor(ui): extract the MCP create form's logic and field groups by @yuneng-berri in https://github.com/BerriAI/litellm/pull/35694 * test(ui): tier the MCP create tests into unit and integration by @yuneng-berri in https://github.com/BerriAI/litellm/pull/35697 * fix(proxy): redact credential headers from request logging copies by @yucheng-berri in https://github.com/BerriAI/litellm/pull/35678 * feat(guardrails/rubrik): prompt moderation, response-text blocking, streaming buffer, failure logging by @devin-ai-integration[bot] in https://github.com/BerriAI/litellm/pull/35722 * fix(ui): render Responses API request and response in the logs drawer by @yuneng-berri in https://github.com/BerriAI/litellm/pull/35718 * fix(ui): hide guardrail review buttons from non-admin users by @ryan-crabbe-berri in https://github.com/BerriAI/litellm/pull/27535 * feat(team): custom metadata validation hook for team create and update by @yuneng-berri in https://github.com/BerriAI/litellm/pull/33353 * ci(circleci): install a pinned Rust toolchain on the Linux jobs by @yuneng-berri in https://github.com/BerriAI/litellm/pull/35519 * fix(bedrock): stop forwarding no-op toolSpec.strict to Converse by @tin-berri in https://github.com/BerriAI/litellm/pull/35688 * fix(ui): reject an auto-router keyword rule left empty instead of dropping it by @tin-berri in https://github.com/BerriAI/litellm/pull/35705 * fix(guardrails/rubrik): attribute blocked requests to the caller that made them by @yucheng-berri in https://github.com/BerriAI/litellm/pull/35734 * fix(responses): forward client headers to the provider on /v1/responses by @devin-ai-integration[bot] in https://github.com/BerriAI/litellm/pull/34531 * feat(spend): add net auto-router savings to the cost-optimization dashboard by @tin-berri in https://github.com/BerriAI/litellm/pull/35521 * chore(typing): clear basedpyright Any errors in budget reset, access groups, and cache settings by @mateo-berri in https://github.com/BerriAI/litellm/pull/35719 * fix(spend): read what a request cost from the record instead of pricing it again by @tin-berri in https://github.com/BerriAI/litellm/pull/35736 * perf: install hiredis so redis-py parses replies with its C parser by @Classic298 in https://github.com/BerriAI/litellm/pull/35709 * feat(ui): show auto-router savings on the cost-optimization dashboard by @tin-berri in https://github.com/BerriAI/litellm/pull/35522 * perf: build log messages lazily so filtered-out log records cost nothing by @Classic298 in https://github.com/BerriAI/litellm/pull/35703 * fix(proxy): retry model cost map fetch with Retry-After-aware backoff and keep current map on reload failure by @ryan-crabbe-berri in https://github.com/BerriAI/litellm/pull/35739 * feat(otel): stamp service tier attributes on inference spans by @devin-ai-integration[bot] in https://github.com/BerriAI/litellm/pull/35679 * fix(proxy): log the model cost map reload failure lazily by @tin-berri in https://github.com/BerriAI/litellm/pull/35750 * fix(groq): translate web_search_options to the browser_search tool by @hMED22 in https://github.com/BerriAI/litellm/pull/34971 * feat(ui): add admin-configurable user banner by @yuneng-berri in https://github.com/BerriAI/litellm/pull/35729 * fix(e2e): make spend-counter redis connection env-driven for non-cluster deployments by @yuneng-berri in https://github.com/BerriAI/litellm/pull/35732 * fix(proxy): make /cursor/chat/completions work with Cursor agent mode by @tin-berri in https://github.com/BerriAI/litellm/pull/34029 * fix(proxy): propagate user_email and bind api_key on JWT auth attribution paths by @devin-ai-integration[bot] in https://github.com/BerriAI/litellm/pull/34331 * chore(build): move the Admin UI toolchain to Node 24 by @yuneng-berri in https://github.com/BerriAI/litellm/pull/35801 * test(e2e): vendor API strategy coverage across endpoints by @mubashir1osmani in https://github.com/BerriAI/litellm/pull/34649 * chore(deps): upgrade cryptography to 50.0.0 by @yuneng-berri in https://github.com/BerriAI/litellm/pull/35803 * test(e2e): cover legacy text /completions endpoint by @mubashir1osmani in https://github.com/BerriAI/litellm/pull/34431 * feat(gemini): add gemini-robotics-er-2-preview and gemini-robotics-er-1.6-preview by @devin-ai-integration[bot] in https://github.com/BerriAI/litellm/pull/35555 * test(e2e): move load/perf testing out of the main suite and drop the vllm passthrough test by @yuneng-berri in https://github.com/BerriAI/litellm/pull/35820 * feat(lint): enforce Final on locals and freeze function parameters (LIT010, LIT011) by @mateo-berri in https://github.com/BerriAI/litellm/pull/35807 * chore: bump litellm-proxy-extras 0.4.81 -> 0.4.82, litellm 1.96.0 -> 1.97.0 by @yuneng-berri in https://github.com/BerriAI/litellm/pull/35810 * fix(bedrock): drop conflicting tool_choice.type when toolConfig.toolChoice is set by @devin-ai-integration[bot] in https://github.com/BerriAI/litellm/pull/35738 * docs(CLAUDE.md): prefer commas over semicolons when replacing em dashes by @mateo-berri in https://github.com/BerriAI/litellm/pull/35825 * chore(lint): zero out basedpyright headroom for purely local rules by @mateo-berri in https://github.com/BerriAI/litellm/pull/35828 * test(e2e): retry provider-transient statuses at the transport with bounded backoff by @yuneng-berri in https://github.com/BerriAI/litellm/pull/35824 * chore(ci): promote internal staging to main by @yuneng-berri in https://github.com/BerriAI/litellm/pull/35836 * refactor(ui): route MCP session tokens through the shared storage helper by @yuneng-berri in https://github.com/BerriAI/litellm/pull/35835 * docs(helm): replace the classic chart's 128Mi resource example with the documented 4Gi sizing by @yassin-berriai in https://github.com/BerriAI/litellm/pull/35830 * fix(proxy): persist periodic reload schedule state so status survives restarts and fires without store_model_in_db by @ryan-crabbe-berri in https://github.com/BerriAI/litellm/pull/35165 * fix(router): eagerly fetch Vertex AI deferred stream to surface HTTP errors in _acompletion fallback path by @deepanshululla in https://github.com/BerriAI/litellm/pull/34627 * fix(azure_storage): honor AZURE_STORAGE_ENDPOINT_SUFFIX for sovereign clouds by @yucheng-berri in https://github.com/BerriAI/litellm/pull/35806 * fix(proxy): apply key_alias/key_hash filters to all /key/list visibility branches by @ryan-crabbe-berri in https://github.com/BerriAI/litellm/pull/35840 * fix(proxy): enforce per-model budgets against resolved cursor model variants by @mateo-berri in https://github.com/BerriAI/litellm/pull/35834 * feat(ui): reorder Add Auto Router into name + template, with a collapsible detailed config by @tin-berri in https://github.com/BerriAI/litellm/pull/35746 * test: repair three failing suites on litellm_internal_staging by @yuneng-berri in https://github.com/BerriAI/litellm/pull/35845 * fix(guardrails): scan model output on the /openai/v1/responses alias by @yucheng-berri in https://github.com/BerriAI/litellm/pull/35818 * ci: pin Node on the Playwright UI lanes so npm ci meets the engines floor by @yuneng-berri in https://github.com/BerriAI/litellm/pull/35848 * fix(pricing): apply OpenAI's gpt-5.6 terra/luna cut to Azure cost map by @mubashir1osmani in https://github.com/BerriAI/litellm/pull/35481 * feat(spend): add caller-scoped key/user/team/organization spend report endpoints by @yuneng-berri in https://github.com/BerriAI/litellm/pull/35725 * revert: "fix(caching): close evicted LLM clients so their connections are reclaimed (#35492)" by @mateo-berri in https://github.com/BerriAI/litellm/pull/35856 * refactor(repositories): add prisma protocol seams and a spend-reset unit of work by @mateo-berri in https://github.com/BerriAI/litellm/pull/35748 * perf(streaming): assemble streamed tool-call arguments in linear time by @mateo-berri in https://github.com/BerriAI/litellm/pull/35826 * fix(s3_v2): sign S3 object URLs with S3SigV4Auth so encoded paths verify by @devin-ai-integration[bot] in https://github.com/BerriAI/litellm/pull/35726 * test(e2e): self-seed the ui suite's password-login users in global setup by @yuneng-berri in https://github.com/BerriAI/litellm/pull/35863 * fix(claude-code): create-only skill registration with a PUT update route (LIT-4110) by @ryan-crabbe-berri in https://github.com/BerriAI/litellm/pull/31752 * fix(proxy): fix zguard httpcode when block input by @jwang-gif in https://github.com/BerriAI/litellm/pull/31948 * fix(lint): pick the merge-aware base so in-progress merges are not blamed for base drift by @mateo-berri in https://github.com/BerriAI/litellm/pull/35868 * chore: bump litellm-proxy-extras 0.4.82 -> 0.4.83 by @yuneng-berri in https://github.com/BerriAI/litellm/pull/35877 * feat(ui): add Test Routing to the auto router create form by @devin-ai-integration[bot] in https://github.com/BerriAI/litellm/pull/35859 * fix(ui): derive auto-router preset tests from the bundled preset JSON by @tin-berri in https://github.com/BerriAI/litellm/pull/35882 * revert: "test(e2e): vendor API strategy coverage across endpoints" (#34649) by @yuneng-berri in https://github.com/BerriAI/litellm/pull/35881 * chore(deps): bump grpc and golang.org/x modules in the terraform provider by @yuneng-berri in https://github.com/BerriAI/litellm/pull/35844 * test(e2e): skip view-backed global spend probes pending LIT-5211 by @yuneng-berri in https://github.com/BerriAI/litellm/pull/35875 * fix(lint): move the basedpyright heap flag into the type check gate by @mateo-berri in https://github.com/BerriAI/litellm/pull/35869 * chore(ci): promote internal staging to main by @yuneng-berri in https://github.com/BerriAI/litellm/pull/35876 * feat(ui): add role capability gating, migrate Tool Policies route by @yuneng-berri in https://github.com/BerriAI/litellm/pull/35812 * refactor(ui): inject the fetch client's base url instead of reading it at import by @yuneng-berri in https://github.com/BerriAI/litellm/pull/35802 * chore: remove unused .flake8 config and flake8 dev dependency by @mateo-berri in https://github.com/BerriAI/litellm/pull/35888 * chore: stop advising pre-commit and bootstrap by @mateo-berri in https://github.com/BerriAI/litellm/pull/35884 * fix(auth): name enable_jwt_auth when a JWT-shaped key is rejected by @yassin-berriai in https://github.com/BerriAI/litellm/pull/35831 * feat(auto-router): make reminder marker pair configurable by @akapur99 in https://github.com/BerriAI/litellm/pull/35874 * fix(UI): update anthropic model presets by @tin-berri in https://github.com/BerriAI/litellm/pull/35896 * fix(bootstrap): switch to the dashboard node floor via nvm or fnm by @mateo-berri in https://github.com/BerriAI/litellm/pull/35895 * perf(pre-commit): run python, dashboard, and gen-api checks concurrently by @mateo-berri in https://github.com/BerriAI/litellm/pull/35903 * feat(spend): derive a default auto-router savings baseline from the hardest tier by @tin-berri in https://github.com/BerriAI/litellm/pull/35907 * fix(http_handler): self-heal handler clients closed after cache eviction by @mateo-berri in https://github.com/BerriAI/litellm/pull/35862 * fix(cost_tracking): keep OpenAI prompt cache token details through usage reassembly by @devin-ai-integration[bot] in https://github.com/BerriAI/litellm/pull/34812 * fix(cost): bill gpt-5.6 prompt cache reads at the cache read rate by @devin-ai-integration[bot] in https://github.com/BerriAI/litellm/pull/34957 * fix(batches): account for Responses API usage by @rimysore in https://github.com/BerriAI/litellm/pull/35367 * ci: retry Codecov uploads and stop failing jobs on OIDC token flakes by @mateo-berri in https://github.com/BerriAI/litellm/pull/35251 * feat(complexity_router): let operators rename the four complexity tiers by @akapur99 in https://github.com/BerriAI/litellm/pull/35893 * chore(lint): zero stale ruff and LIT headroom and strip inert type: ignore comments by @mateo-berri in https://github.com/BerriAI/litellm/pull/35928 * chore(lint): zero out seven more purely local basedpyright rules by @mateo-berri in https://github.com/BerriAI/litellm/pull/35927 * chore(ui): zero stale headroom on local dashboard eslint budgets by @mateo-berri in https://github.com/BerriAI/litellm/pull/35929 * fix(managed-files): skip rows without file objects by @rimysore in https://github.com/BerriAI/litellm/pull/35365 * fix(router): redact fallback tracebacks at the call site and cover the sync deferred stream by @yassin-berriai in https://github.com/BerriAI/litellm/pull/35843 * fix(migrations): recover from an interrupted Prisma toolchain install by @yassin-berriai in https://github.com/BerriAI/litellm/pull/35832 * fix(lint): bring basedpyright rule counts back under their budget limits by @mateo-berri in https://github.com/BerriAI/litellm/pull/35962 * chore(ui): don't zero out stale headroom except no-console by @mateo-berri in https://github.com/BerriAI/litellm/pull/35964 * fix(proxy): give proxy_admin_viewer read parity with proxy_admin by @ryan-crabbe-berri in https://github.com/BerriAI/litellm/pull/35851 * refactor(ui): address UI lint budget issues by refactoring UI by @tin-berri in https://github.com/BerriAI/litellm/pull/35960 * fix(ci): make the env-key doc gate see get_secret_bool reads by @yassin-berriai in https://github.com/BerriAI/litellm/pull/35833 * fix(caching): re-land evicted LLM client closing (#35492) atop self-healing handlers by @mateo-berri in https://github.com/BerriAI/litellm/pull/35870 * fix(proxy): keep the connected DB client when a startup health check fails by @yassin-berriai in https://github.com/BerriAI/litellm/pull/35837 * chore(lint): remove litellm/types from the ruff lint exclusion by @mateo-berri in https://github.com/BerriAI/litellm/pull/35926 * feat(sgr): make the gateway middleware the source of truth for successful requests by @yassin-berriai in https://github.com/BerriAI/litellm/pull/35717 * feat(auto-router): let operators replace the LLM classifier's system prompt by @akapur99 in https://github.com/BerriAI/litellm/pull/35855 * fix(docker): bake the pip image's prisma engines at a world-readable path by @yassin-berriai in https://github.com/BerriAI/litellm/pull/35976 * fix(auth): return 403 from the OAuth2 enterprise gate by @yassin-berriai in https://github.com/BerriAI/litellm/pull/35838 * fix(router): keep custom model_info across a price data reload by @yassin-berriai in https://github.com/BerriAI/litellm/pull/35491 * fix(proxy): resolve pass-through credentials live from router deployments by @mateo-berri in https://github.com/BerriAI/litellm/pull/35916 * fix(ci): fetch only head and merge-base in lint jobs instead of every branch by @mateo-berri in https://github.com/BerriAI/litellm/pull/35982 * fix(autorouter): match CJK keyword_tier_rules that regex word boundaries miss by @akapur99 in https://github.com/BerriAI/litellm/pull/35984 * feat(spend): rebuild the auto-router benchmarks backend as a per-session rollup by @tin-berri in https://github.com/BerriAI/litellm/pull/35910 * refactor(ui): replace hand-rolled query-param routing with nuqs by @ryan-crabbe-berri in https://github.com/BerriAI/litellm/pull/35871 * fix(docker): bake the componentized prisma engines at /opt/prisma so any uid can start by @yassin-berriai in https://github.com/BerriAI/litellm/pull/35989 * fix(migrations): keep the toolchain heal from raising on an unreadable nodeenv cache by @yassin-berriai in https://github.com/BerriAI/litellm/pull/35986 * fix(bedrock): sign Bedrock managed-file S3 requests with S3SigV4Auth by @yassin-berriai in https://github.com/BerriAI/litellm/pull/35983 * chore(typing): replace Any seams with real types across responses, proxy, and provider adapters by @mateo-berri in https://github.com/BerriAI/litellm/pull/35809 * fix(ai21): resolve the documented AI21_API_KEY instead of a misspelled name by @yassin-berriai in https://github.com/BerriAI/litellm/pull/35985 * fix(docker): fail the image build when the generated prisma engine paths drift off /opt/prisma by @yassin-berriai in https://github.com/BerriAI/litellm/pull/35979 * fix(jina_ai): resolve the documented JINA_API_KEY as a fallback by @yassin-berriai in https://github.com/BerriAI/litellm/pull/35992 * fix(proxy): only treat a recoverable database outage as grounds to serve without one by @yassin-berriai in https://github.com/BerriAI/litellm/pull/35864 * fix(ci): make every remaining CI checkout shallow by @mateo-berri in https://github.com/BerriAI/litellm/pull/35997 * fix(auto-router): stop the embedding model's context window from failing long requests by @akapur99 in https://github.com/BerriAI/litellm/pull/35956 * fix(ci): make the env-key doc gate see bare get_secret and get_secret_str reads by @yassin-berriai in https://github.com/BerriAI/litellm/pull/35996 * fix(logging): extend secret redaction to records litellm does not emit directly by @yassin-berriai in https://github.com/BerriAI/litellm/pull/35977 * test(utils): pin the register_model replay test to the recorded half by @yassin-berriai in https://github.com/BerriAI/litellm/pull/35994 * fix(ci): run every helm test suite, not just the first one per file by @yassin-berriai in https://github.com/BerriAI/litellm/pull/35993 * ci: fail the build when a test file or Dockerfile is invoked by no job by @yassin-berriai in https://github.com/BerriAI/litellm/pull/35991 * fix(langfuse): stop a collected httpx handler from closing a shared client by @yassin-berriai in https://github.com/BerriAI/litellm/pull/35981 * fix(bedrock): grant bedrock:CountTokens in OIDC session policy by @devin-ai-integration[bot] in https://github.com/BerriAI/litellm/pull/33145 * feat(pre-commit): save full lint output to a per-worktree log file by @mateo-berri in https://github.com/BerriAI/litellm/pull/36004 * feat(ui): match auto-router preset models against deployments' underlying model IDs by @tin-berri in https://github.com/BerriAI/litellm/pull/35972 * fix(core_helpers): map generic 'error' finish_reason to 'stop' by @devin-ai-integration[bot] in https://github.com/BerriAI/litellm/pull/33972 * fix(proxy)!: apply request-parameter checks consistently across body, path and form inputs by @yuneng-berri in https://github.com/BerriAI/litellm/pull/36011 * fix: rebuild models_by_provider in add_known_models so cost map reloads reach wildcard expansion by @ryan-crabbe-berri in https://github.com/BerriAI/litellm/pull/36010 * feat(complexity_router): report LLM classifier cost per request via routing_decision and x-litellm-classifier-cost header by @tin-berri in https://github.com/BerriAI/litellm/pull/36015 * fix(model-prices): correct replicate model key typo by @AkashNaickar in https://github.com/BerriAI/litellm/pull/34800 * fix(proxy): register managed batch output files on terminal retrieve by @Souravrajvi0 in https://github.com/BerriAI/litellm/pull/34092 * perf(pre-commit): fetch basedpyright base counts from CI artifacts by @mateo-berri in https://github.com/BerriAI/litellm/pull/35970 * fix(ui): sync projects list page index to ?page= so back and reload keep the page by @ryan-crabbe-berri in https://github.com/BerriAI/litellm/pull/36003 * fix(ui): link project page keys to their virtual key detail by @ryan-crabbe-berri in https://github.com/BerriAI/litellm/pull/36002 * refactor(ui): drop unreferenced locals from dashboard route components by @yuneng-berri in https://github.com/BerriAI/litellm/pull/35819 * fix(ui): opening a project now pushes ?project= so back and deep links work by @ryan-crabbe-berri in https://github.com/BerriAI/litellm/pull/36001 * refactor(ui): drop unreferenced locals from shared dashboard components by @yuneng-berri in https://github.com/BerriAI/litellm/pull/35821 * refactor(ui): drop unreferenced locals from tests and narrow destructures by @yuneng-berri in https://github.com/BerriAI/litellm/pull/36025 * fix(guardrails): allow litellm_content_filter to run on post_mcp_call by @mateo-berri in https://github.com/BerriAI/litellm/pull/35980 * fix(guardrails): scan /v1/messages tool traffic by @mateo-berri in https://github.com/BerriAI/litellm/pull/35999 * refactor(ui): drop dead locals and unused React state across the dashboard by @yuneng-berri in https://github.com/BerriAI/litellm/pull/36026 * feat(ui): add the auto-router usage tab to cost optimization by @tin-berri in https://github.com/BerriAI/litellm/pull/35995 * fix(managed_files): derive unified output file ids deterministically so concurrent registrations converge by @mateo-berri in https://github.com/BerriAI/litellm/pull/36019 * fix(proxy): send keepalive pings on anthropic messages SSE streams during upstream silence by @mateo-berri in https://github.com/BerriAI/litellm/pull/36024 * fix(managed_files): return unified ids from unscoped file listing by @mateo-berri in https://github.com/BerriAI/litellm/pull/36031 * fix(arize_phoenix): lowercase OTLP/gRPC auth metadata key by @devin-ai-integration[bot] in https://github.com/BerriAI/litellm/pull/34883 * fix(auto-router): accept every reminder marker pair a harness emits by @tin-berri in https://github.com/BerriAI/litellm/pull/36029 * fix(pricing): sync flex/priority tier keys to dated OpenAI snapshot variants by @mateo-berri in https://github.com/BerriAI/litellm/pull/35923 * fix(cost): bill reasoning tokens at the service tier output rate by @mateo-berri in https://github.com/BerriAI/litellm/pull/35925 * fix(proxy): include today's UTC bucket when a daily activity range ends at the caller's current day by @tin-berri in https://github.com/BerriAI/litellm/pull/36051 * fix: expired-miss share over all measured turns + cost-optimization tab labels by @tin-berri in https://github.com/BerriAI/litellm/pull/36037 * fix(router): include Bedrock batch/S3 fields and model in deployment credentials by @mpcusack-altos in https://github.com/BerriAI/litellm/pull/24548 * fix(batch): track cost for managed batches with no attributable key/u… by @elinacse in https://github.com/BerriAI/litellm/pull/35468 * feat(guardrails): add scan_only_tool_results to scope unified guardrails to tool results by @mateo-berri in https://github.com/BerriAI/litellm/pull/36014 * fix(cost): stop token-pricing the placeholder input on file content calls by @devin-ai-integration[bot] in https://github.com/BerriAI/litellm/pull/35140 * fix(proxy): fetch background responses through the router in CheckResponsesCost by @devin-ai-integration[bot] in https://github.com/BerriAI/litellm/pull/35137 * fix(proxy): yaml store_prompts_in_spend_logs should take precedence over DB cached value by @Praveena-617 in https://github.com/BerriAI/litellm/pull/35769 * fix(lint): measure the basedpyright budget gate in a gate-owned venv by @mateo-berri in https://github.com/BerriAI/litellm/pull/36050 * docs: cap all GitHub comments at 15-25 words, curb semicolon splices by @mateo-berri in https://github.com/BerriAI/litellm/pull/36059 * chore(lint): name MappingProxyType in the mutable-collection fix messages by @mateo-berri in https://github.com/BerriAI/litellm/pull/36072 * test: roll back runtime model registrations between tests by @mateo-berri in https://github.com/BerriAI/litellm/pull/36039 * refactor(types): cut 653 implicit and explicit Any diagnostics across 11 modules by @mateo-berri in https://github.com/BerriAI/litellm/pull/36054 * fix(proxy): stop resolving the UI session sentinel team on /search_tools/list by @yuneng-berri in https://github.com/BerriAI/litellm/pull/36061 * fix(batches): persist managed file ids for cancelled/failed/expired batches by @mateo-berri in https://github.com/BerriAI/litellm/pull/36048 * fix(batches): register managed output files on batch cancel by @mateo-berri in https://github.com/BerriAI/litellm/pull/36034 * fix(proxy): allow non-admins to reach /user/daily/activity/aggregated by @yuneng-berri in https://github.com/BerriAI/litellm/pull/36062 * fix(anthropic): coerce explicit additionalProperties to false in output_format schema by @dkindlund in https://github.com/BerriAI/litellm/pull/35811 * fix(batches): prevent managed file fallbacks by @rimysore in https://github.com/BerriAI/litellm/pull/35371 * chore: ignore the mechanical lint and typing sweeps in git blame by @mateo-berri in https://github.com/BerriAI/litellm/pull/36076 * fix(proxy): warn at startup when max_budget is set but no database is connected by @devin-ai-integration[bot] in https://github.com/BerriAI/litellm/pull/36041 * fix(proxy): promote caller metadata trace fields into litellm_metadata by @yucheng-berri in https://github.com/BerriAI/litellm/pull/35866 * feat(terraform): sync provider 0.3.0 from the mirror and cut 0.4.0 by @yuneng-berri in https://github.com/BerriAI/litellm/pull/36098 * fix(guardrails): honor configured timeout in Zscaler AI Guard by @yucheng-berri in https://github.com/BerriAI/litellm/pull/36110 * fix(logging): fall back to litellm_metadata when metadata is empty by @yucheng-berri in https://github.com/BerriAI/litellm/pull/36105 * fix(proxy): re-assert the authenticated identity on passthrough requests by @yucheng-berri in https://github.com/BerriAI/litellm/pull/36121 * chore: bump litellm-enterprise 0.1.53 -> 0.1.54, litellm-proxy-extras 0.4.83 -> 0.4.84 by @yuneng-berri in https://github.com/BerriAI/litellm/pull/36139 * fix(ui): match auto-router preset models against wildcard-expanded model groups by @tin-berri in https://github.com/BerriAI/litellm/pull/36111 * test(router): assert the auto-router max_input_chars kwarg by @yuneng-berri in https://github.com/BerriAI/litellm/pull/36109 * fix(ui): allow clearing a key's budget reset from the Edit Key form by @ryan-crabbe-berri in https://github.com/BerriAI/litellm/pull/36140 * fix(managed_files): skip unparseable rows when listing managed files by @mateo-berri in https://github.com/BerriAI/litellm/pull/36021 * fix(a2a): stop writing per-caller headers onto the shared cached httpx client by @yassin-berriai in https://github.com/BerriAI/litellm/pull/35978 * build(deps): bump h2 to 4.4.1 and js-yaml to 4.3.1 by @yuneng-berri in https://github.com/BerriAI/litellm/pull/36147 * chore: promote staging to main by @mateo-berri in https://github.com/BerriAI/litellm/pull/36057 * fix(azure_sentinel): respect AZURE_AUTHORITY_HOST and derive the Azure Monitor audience per cloud by @yucheng-berri in https://github.com/BerriAI/litellm/pull/36137 * fix(bedrock): pass SSE-KMS key through to the batch input-file S3 upload by @devin-ai-integration[bot] in https://github.com/BerriAI/litellm/pull/35148 * fix(anthropic adapter): stop indexing choices[0] on choiceless streaming chunks by @devin-ai-integration[bot] in https://github.com/BerriAI/litellm/pull/35314 * fix(bedrock): normalize /v1/completions and /v1/responses batch records by @devin-ai-integration[bot] in https://github.com/BerriAI/litellm/pull/35675 * fix(proxy): return the real status code when a credential update is rejected by @yucheng-berri in https://github.com/BerriAI/litellm/pull/36166 * fix(proxy): improve Headroom /v1/compress HTTP 404 diagnostics by @aayush598 in https://github.com/BerriAI/litellm/pull/35952 * fix(proxy): invalidate cached project object on project update and delete by @ryan-crabbe-berri in https://github.com/BerriAI/litellm/pull/36028 * feat(proxy): add apply_user_budget_to_team_keys opt-in by @ryan-crabbe-berri in https://github.com/BerriAI/litellm/pull/36102 * fix(proxy): stop alerting on health probes that lose the planned engine-restart race by @yassin-berriai in https://github.com/BerriAI/litellm/pull/36141 * test(docker): gate the componentized gateway and backend images on an arbitrary-uid offline boot by @yassin-berriai in https://github.com/BerriAI/litellm/pull/36136 * fix(http): stop pooled clients persisting cookies on the aiohttp jar too by @yassin-berriai in https://github.com/BerriAI/litellm/pull/36149 * fix(router): bound fallback-walk work and error-log volume by @yassin-berriai in https://github.com/BerriAI/litellm/pull/36148 * ci: wire credential_endpoints tests into the proxy endpoints job by @cursor[bot] in https://github.com/BerriAI/litellm/pull/36187 * docs(keys): document /key/info fields and clarify budget_reset_at is the next reset by @ryan-crabbe-berri in https://github.com/BerriAI/litellm/pull/36127 * fix(azure_sentinel): add AZURE_SENTINEL_AUTHORITY_HOST as a Sentinel scoped override by @yucheng-berri in https://github.com/BerriAI/litellm/pull/36165 * docs(pr-template): add a User Flow section with authoring instructions by @mateo-berri in https://github.com/BerriAI/litellm/pull/36162 * fix(proxy): derive config agent ids from agent_name so grants survive secret rotation by @ryan-crabbe-berri in https://github.com/BerriAI/litellm/pull/36020 * chore(ui): regenerate schema.d.ts for the /key/info docstring update by @devin-ai-integration[bot] in https://github.com/BerriAI/litellm/pull/36210 * build(deps): bump gitpython to 3.1.58 to clear osv-scan on staging by @devin-ai-integration[bot] in https://github.com/BerriAI/litellm/pull/36212 * fix(proxy): deny agent access when key and team grants resolve to nothing by @ryan-crabbe-berri in https://github.com/BerriAI/litellm/pull/36221 * build(deps): defer the second pypdf advisory until the 6.15.0 bump by @devin-ai-integration[bot] in https://github.com/BerriAI/litellm/pull/36218 * fix(a2a): align agent list annotation and test with the tuple return type by @devin-ai-integration[bot] in https://github.com/BerriAI/litellm/pull/36217 * ci: always run the UI API types sync check so it can be required by @devin-ai-integration[bot] in https://github.com/BerriAI/litellm/pull/36213 * build(deps): bump nanoid to 3.3.17 in the dashboard lockfile by @devin-ai-integration[bot] in https://github.com/BerriAI/litellm/pull/36227 * feat(ui): show user email or alias in usage data export by @ryan-crabbe-berri in https://github.com/BerriAI/litellm/pull/36232 * feat(auto-router): track turns per complexity tier (LIT-5302) by @tin-berri in https://github.com/BerriAI/litellm/pull/36209 * fix(websearch): restore snippet text in native web_search_tool_result blocks (LIT-5315) by @tin-berri in https://github.com/BerriAI/litellm/pull/36228 * fix(proxy): resolve entity access groups in the model listing endpoints by @devin-ai-integration[bot] in https://github.com/BerriAI/litellm/pull/36230 * fix(ui): let access groups be a team's only model source, with hover provenance by @ryan-crabbe-berri in https://github.com/BerriAI/litellm/pull/36234 * fix(managed_files): return unified output file ids from GET /batches by @mateo-berri in https://github.com/BerriAI/litellm/pull/36049 * test(proxy): compare empty agent list to the tuple get_agent_list returns by @devin-ai-integration[bot] in https://github.com/BerriAI/litellm/pull/36225 * fix(otel): name the RPC system and upstream on MCP tool-call spans by @yucheng-berri in https://github.com/BerriAI/litellm/pull/35857 * fix(guardrails): chunk oversized Bedrock ApplyGuardrail requests instead of failing by @yucheng-berri in https://github.com/BerriAI/litellm/pull/36119 * test(e2e): settle control-plane writes across every replica, not just one by @yuneng-berri in https://github.com/BerriAI/litellm/pull/36247 * fix(responses): forward allowed_openai_params through the chat completions bridge by @devin-ai-integration[bot] in https://github.com/BerriAI/litellm/pull/35885 * test(proxy): assert the copy _add_team_member_budget_table returns by @devin-ai-integration[bot] in https://github.com/BerriAI/litellm/pull/36244 * chore(ui): regenerate dashboard api types for tier_turns by @devin-ai-integration[bot] in https://github.com/BerriAI/litellm/pull/36243 * refactor(types): declare mirrored pricing fields on ModelInfo by @devin-ai-integration[bot] in https://github.com/BerriAI/litellm/pull/36215 * fix(lint): make strict-gate noqas survive base ruff and flag stale ones by @mateo-berri in https://github.com/BerriAI/litellm/pull/36257 * fix(vertex_ai): surface real error/status on vertex batch create instead of IndexError 500 by @devin-ai-integration[bot] in https://github.com/BerriAI/litellm/pull/35141 * ci: give the remaining pull_request workflows a concurrency group by @devin-ai-integration[bot] in https://github.com/BerriAI/litellm/pull/36252 * refactor(lint): graduate zero-violation strict rules and guard the budget ratchet by @mateo-berri in https://github.com/BerriAI/litellm/pull/36161 * fix(proxy): enforce require_managed_files on every route that accepts a raw provider id by @devin-ai-integration[bot] in https://github.com/BerriAI/litellm/pull/35551 * chore(typing): clear 1.4k basedpyright Any errors across 21 hotspot files by @mateo-berri in https://github.com/BerriAI/litellm/pull/36282 * test: roll back live router replay membership between tests by @mateo-berri in https://github.com/BerriAI/litellm/pull/36278 * chore(ci): sync main into internal staging by @yuneng-berri in https://github.com/BerriAI/litellm/pull/36288 * build(lint): rename make pre-commit to make check with a working-tree fallback by @mateo-berri in https://github.com/BerriAI/litellm/pull/36277 * fix(ui): show team BYOK models in team fallback settings by @devin-ai-integration[bot] in https://github.com/BerriAI/litellm/pull/36241 * fix(otel): mark v2 server spans as failed for pre-call errors by @devin-ai-integration[bot] in https://github.com/BerriAI/litellm/pull/34546 * fix(websearch_interception): bill intercepted searches to the calling key by @devin-ai-integration[bot] in https://github.com/BerriAI/litellm/pull/35708 * chore: remove pre-commit rule by @mateo-berri in https://github.com/BerriAI/litellm/pull/36295 * docs: clarify guideline priority ordering in CLAUDE.md by @devin-ai-integration[bot] in https://github.com/BerriAI/litellm/pull/36296 * feat(router): independent, default-on deployment affinity for the auto-router by @tin-berri in https://github.com/BerriAI/litellm/pull/36146 * test: repair stale CircleCI contracts by @yuneng-berri in https://github.com/BerriAI/litellm/pull/36293 * chore(ci): promote internal staging to main by @yuneng-berri in https://github.com/BerriAI/litellm/pull/36286 * chore: rebuild Admin UI bundle for the 2026-08-08 release by @yuneng-berri in https://github.com/BerriAI/litellm/pull/36297 * chore(ci): promote internal staging to main by @yuneng-berri in https://github.com/BerriAI/litellm/pull/36304 ## New Contributors * @rimysore made their first contribution in https://github.com/BerriAI/litellm/pull/35367 * @AkashNaickar made their first contribution in https://github.com/BerriAI/litellm/pull/34800 * @Souravrajvi0 made their first contribution in https://github.com/BerriAI/litellm/pull/34092 * @elinacse made their first contribution in https://github.com/BerriAI/litellm/pull/35468 * @aayush598 made their first contribution in https://github.com/BerriAI/litellm/pull/35952 * @cursor[bot] made their first contribution in https://github.com/BerriAI/litellm/pull/36187 **Full Changelog**: https://github.com/BerriAI/litellm/compare/v1.96.0...v1.97.0