v0.2.14
ruvnet/RuVectorv0.2.14Sep 1, 2026by github-actions[bot]
AI Summary
This release introduces the PageIndex SDK with a 'Flash' engine for faster indexing, unified local and cloud clients, and enhanced agent integration. It adds support for protocol-specific chat lanes and refined process display options.
Key Highlights
- Flash engine speeds up indexing by using layout stats instead of LLMs for structure generation.
- Unified client API allows seamless switching between local and cloud modes.
- Enhanced agent integration with MCP tools and serialization methods.
- Flexible chat surfaces with protocol support (Responses/Messages) and streaming.
Breaking Changes
- Method names `responses()` and `messages()` are renamed to `chat(protocol="responses")` and `chat(protocol="messages")`.
- Dict keys for `show_process` changed from `tool_calls`/`tool_results` to `tool_call`/`tool_result`.
- Old method names raise errors, and `ChatProcessOptions` fields align with event types.
New Features
- Protocol-specific chat methods (`protocol="responses"`, `protocol="messages")
- Enhanced `chat()` method with `instructions`, `max_turns`, `backend`, `extra_headers`, `extra_body`, and `reasoning_effort`
- Typed config shapes (`IndexConfig`, `ChatConfig`, etc.)
- Strict validation and error reporting for unknown keys and empty values
Full Release Notes
- The **PageIndex SDK**, **local** or **cloud** — vectorless, reasoning-based RAG, end to end.
- **Much faster indexing** — the **PageIndex Flash** engine gets the tree from layout stats: no LLM involved for the structure generation itself, LLMs only write the node summaries, and tree expansion proposes a wave of nodes concurrently.
```python
client = PageIndexClient()
client.submit_document("report.pdf")
client.chat("What does the report conclude?")
```
Index, to chat, to agent integration, one client.
Local mode needs no server, no vector DB, no PageIndex API key.
## Highlights
- **Flash engine**: the local default (`mode="standard"` keeps the classic LLM pipeline). Embedded bookmarks are consumed when trustworthy, and tree optimization is on by default — `optimize="merge"` for the deterministic LLM-free pass, `"full"` (default) adds LLM expand, which runs a wave of nodes concurrently instead of one round-trip at a time.
- **One complete surface, local and cloud**: `PageIndexLocalClient(storage_path=...)` is the same client as cloud — submit, tree, page content, chat, and agent tools all present in both modes, so code moves between them unchanged.
- **Cloud documents, your own model**: `api_key` decides where your documents live; a configured chat model decides who answers — and the two combine. `PageIndexClient(api_key="pi-...", chat_model="openai/gpt-5.2")` runs the same in-process document-QA engine over the live cloud tool set. Page content flows through your process to your provider on your credentials; `doc_id` targets at the prompt level; `enable_citations` stays with the managed chat.
- **Agent integration**: the cloud MCP tool contract, in-process — `client.agent_tools()` (plain functions), `as_openai_tools()`, `as_anthropic_tools()`, `as_claude_mcp()`, plus one-call `openai_agent_config()` / `anthropic_runner_config()` / `claude_agent_config()` bundles and `agent_instructions()` for the system prompt. Cloud clients get the live server tool set over the MCP bridge (read-only endpoint by default); local clients get the in-process subset with the same schemas and envelopes — agent prompts port unchanged.
- **Chat surfaces**: `chat()` — question in, answer out, on any backend; `chat(stream=True)` shows the run as it happens, thinking and tool calls woven into the text, or as typed events via `.events`; `chat(protocol="responses" | "messages")` drives the OpenAI Responses or Anthropic Messages API natively with that protocol's own shapes, and `chat_completions()` keeps the OpenAI-compatible envelope — all with `doc_id` targeting, streaming, honest usage accounting, and prompt-cache continuity across turns. Transcripts append verbatim: a protocol lane's output goes back into the next request unchanged.
- **Model & connection knobs**: `index_model` / `chat_model`, `index_backend` / `chat_backend` (and per-call `backend`) passed verbatim to each lane — LiteLLM-routed providers, keyless OpenAI-compatible servers, Azure/Bedrock/Vertex included.
- **`index=` / `chat=` slots**: the grouped spelling of the flat arguments — a string shorthand or a mapping (`index={"model": ..., "storage_path": ...}`, `chat={"model": ..., "backend": ...}`). `"cloud"` / `"local"` name a side, an optional `mode=` cross-checks it, and `PageIndexLocalClient` / `PageIndexCloudClient` take the same slots. `PageIndexCloudClient()` reads `PAGEINDEX_API_KEY`; a bare `PageIndexClient()` stays local no matter what the environment holds.
- **Typed config shapes**: `IndexConfig` / `CloudIndexConfig` / `LocalIndexConfig` / `ChatConfig`, with `py.typed` shipped so your type checker sees them.
- **Nothing fails quietly**: unknown keys, mixed sides, empty values and mode/content conflicts refuse at construction with the legal vocabulary in the message; dead credentials or a missing model fail the indexing run instead of storing a document with blank summaries; every cloud error carries its HTTP status.
- **Dependencies**: Python >= 3.10; `openai-agents` in the base install (the chat engine); `[anthropic]` and `[claude]` extras for those SDKs.
## Also in 0.2.14
- **The protocol doors move behind the front door.** `responses()` and `messages()` are now `chat(protocol="responses")` and `chat(protocol="messages")` — same engines, that protocol's own input and output shapes (transcript items or content blocks in, the response envelope or native event stream out), and `show_process` does not apply there because the transcript is the process. **Breaking**: the old method names raise, and the message names the new spelling. `chat_completions()` is unchanged.
- **`chat()` gains the knobs that lost their public home**: `instructions` (appended after the managed prompt), `max_turns`, `backend`, `extra_headers`, and `extra_body` for the provider's own request fields (`thinking`, `top_k`, `max_output_tokens`, …), merged last so they win. `reasoning_effort` lands in each lane's native spelling — LiteLLM's `reasoning_effort`, Responses `reasoning.effort`, Messages `output_config.effort`. `protocol="messages"` needs `model=` naming a Claude model.
- **The process display speaks one vocabulary.** Woven lines are labeled by their event type names — `[thinking]`, `[tool_call]`, `[tool_result]` — flush left, no arrows, no indent; and the `show_process` dict keys are those same names: `thinking` / `tool_call` / `tool_result` (+ `max_chars`), matching `.events`. **Breaking against 0.2.13's one-day-old spelling**: `{"tool_calls": ...}` / `{"tool_results": ...}` now raise, with the valid keys named in the error; `ChatProcessOptions` fields follow.
- **Docs**: the chat page at [docs.pageindex.ai](https://docs.pageindex.ai/sdk/chat) documents `show_process`, `.events`, and the protocol lanes.
**Full Changelog**: https://github.com/VectifyAI/PageIndex/compare/v0.2.13...v0.2.14