v0.1.13

feyninc/pulpiev0.1.13Aug 10, 2026by warren618

AI Summary

Fixes critical 'grounding' bugs that caused query refusals, introduces a comprehensive finance-math library (src/quantlib), and adds institutional research tools and slash commands.

Key Highlights

  • Fix for the identity/grounding gate causing query refusals
  • Introduction of src/quantlib (265 tested finance-math functions)
  • Valuation engine (DCF, comps, three-statement)
  • Institutional research slash commands (/comps, /dcf, etc.)
  • New data tools: institutional holdings, ETF holdings, prediction markets, research papers

New Features

  • src/quantlib finance-math layer
  • Valuation engine (run_dcf, run_comps, threestatement)
  • Institutional research slash commands
  • Investor lenses skill
  • Five ready-to-schedule research playbooks
  • get_institutional_holdings tool
  • etf_holdings tool
  • prediction_market tool
  • research_papers tool
  • Run manifest and audit ledger for governance

Full Release Notes

# v0.1.13 — The grounding repair, a finance-math layer, and institutional research

`pip install -U vibe-trading-ai`

**408 commits and 162 merged pull requests since [v0.1.12](https://github.com/HKUDS/Vibe-Trading/releases/tag/v0.1.12) (2026-07-22)** — the largest release to date, from 35 contributors. Full detail in [CHANGELOG.md](https://github.com/HKUDS/Vibe-Trading/blob/main/CHANGELOG.md).

---

## 🛡 The headline is a fix, not a feature

The single most user-visible defect in the 0.1.12 line: a **well-formed question would spend minutes on real tool calls and then refuse to answer**, returning

> 当前无法安全确认标的身份或价格证据,因此没有生成交易结论。请确认候选证券代码和交易所后再继续。
> *(cannot safely confirm instrument identity or price evidence)*

The identity/grounding gate was rejecting answers it already had the evidence for. Seven independent root causes, each now fixed and covered by two-sided guard tests (healthy samples must pass; bad samples must still be blocked) plus mutation tests:

| # | Root cause | Effect |
|---|-----------|--------|
| 1 | `.SS` and `.SH` treated as different instruments | **Every Shanghai ticker was permanently `ambiguous`** |
| 2 | A+H dual listings and bare HK codes collapsed to `ambiguous` | Dead end instead of a shortlist |
| 3 | Yahoo returns HTTP 400 for *every* CJK query; recorded as a source **failure** | Escalated to blocking `invalidated` instead of "not listed here" |
| 4 | A failed side query could demote an already-locked identity | Sticky aggregate status with no recovery |
| 5 | A hardcoded per-tool whitelist decided which bare tickers could match | **11 of the 17 documented argument spellings were blocked** (3/14 → 14/14 after the fix) |
| 6 | Source and currency had to be written in ASCII | Chinese answers rejected for writing `雅虎`/`腾讯`, or `元` instead of `人民币` |
| 7 | A thousands separator split the clause mid-number | `¥1,309.22` compared as `1` against the observed range → false price conflict |

Conceptual questions with no instrument at all, and comparison reports, no longer dead-end either.

**A quote outside recorded OHLC evidence is still refused.** The gate got more precise, not weaker — that is the point of the two-sided tests.

Verified end to end against a real LLM, real tools and real network: a `600519.SH` query that previously took **6m14s and ended in the canned refusal** now answers in **7–9s with zero rejections**; `茅台` in 14s; a knowledge question in 14s.

This also closes the narrower cases from the 0.1.12 line — numbers that were never prices (confidence scores, indicator readings, moving-average windows, year-less dates like `8/5`, percentage ranges, and a trading plan's own trigger levels, where `close ≥ 6.45` is a condition rather than a quote) — and makes a many-candidate shortlist count as an answer rather than a stalled resolution (#1001, #983, #955).

---

## 🧮 `src/quantlib` — a tested finance-math layer

**265 tested functions across 19 modules** — every module the tool allowlists now exports. Skills now *import* these instead of carrying formulas inside markdown code blocks — if you find a pricing formula living in a `SKILL.md`, that is a bug, not a pattern.

| Module | Coverage |
|--------|----------|
| `options` | Black-Scholes price + greeks, implied-volatility inversion |
| `fixedincome` | Bond math, Nelson-Siegel / Svensson curve fitting |
| `credit` | Altman Z-score, Merton / KMV distance-to-default |
| `timeseries` | Stationarity, cointegration, GARCH, bootstrap |
| `risk` · `var_backtest` | VaR / CVaR / EVT and their backtests |
| `attribution` | Brinson-Fachler decomposition |
| `performance` · `fundmath` | TWR / MWR / Modified Dietz; XIRR / MOIC / DPI / TVPI |
| `factormodel` · `eventstudy` | Factor regressions, event studies |
| `multipletesting` · `crossvalidation` | Deflated significance, purged CV |
| `impact` | Market-impact models |

The read-only `quantlib_call` tool reaches all of it through one contract, so the finance math works on the CLI, the Web UI, the REST API **and MCP** — including deployments **where `bash` is gated off**. It is structurally not a shell — module allowlist, `__all__`-only dispatch, `export_*` refused. Econometrics needs the `stats` extra; those functions lazy-import and name the missing one.

---

## 💰 Valuation engine

A valuation engine that refuses to invent its own inputs. The one rule in `contracts.py`: **a missing input makes a model NOT RUNNABLE and is never silently defaulted** — every default in a valuation model is an opinion wearing a constant's clothes.

- **`run_dcf`** — FCFF bridge, WACC build, mid-year discounting, net-debt bridge, WACC×g sensitivity grid. Dual terminal value: each method is cross-checked against the other's implied multiple and implied g.
- **`run_comps`** — EV bridge, LTM + calendar-year calendarisation, multiple matrix. A peer with a non-positive denominator is **excluded and reported**, never averaged in as a negative multiple.
- **`threestatement`** — linked projection with a hard balance assertion, an explicit revolver plug, and an iterated interest↔debt circularity that must converge or raise.

Artifacts are input-hashed and versioned, with xlsx / pptx export.

---

## 🏛 Institutional research

Six slash commands — **`/comps` `/dcf` `/attrib` `/memo` `/earnings` `/screen`** — each carrying a step skeleton and an arithmetic-consistent worked example: the Brinson decomposition sums exactly to active return, and the earnings bridge sums exactly to the EPS delta.

**Investor lenses** become a standalone skill: named-investor reasoning frameworks as stackable analysis overlays, decoupled from the data layer. Each lens is an operating procedure — priority signals, disqualifying conditions, typical misuse — not a biography, and names no tool.

**Five ready-to-schedule research playbooks** (premarket brief, earnings-season tracker, portfolio checkup, A-share money flow, institutional-holdings diff), reachable three ways: auth-gated REST routes, a `vibe-trading playbook` CLI subcommand, and a `/playbook` slash command. Templates state their data needs in natural language rather than naming tools, so coverage can grow without editing them, and every one mandates naming a missing input instead of filling it from memory.

---

## 📊 Four new read-only data tools — all on free public sources

- **`get_institutional_holdings`** — SEC 13F-HR books in manager / ticker-holders / top-managers modes, with **quarter-over-quarter position diffs** for factor use. Cover-page totals carry `value_units` + basis, because pre-2023 filings report in thousands.
- **`etf_holdings`** — cross-market look-through. SEC N-PORT for US; for A-shares the **semi-annual/annual reports carry the full book**, not the quarterly top ten: `510300` returns **342 rows covering 98.66% of net assets** versus 10 rows / 22.74%. `coverage` separates `full_portfolio` from `top_n_disclosed`, and every response stamps the report period — a full book is always the older disclosure.
- **`prediction_market`** — event-contract search / event / market / history, with prices converted to **implied probability and the unit labelled**, so downstream never reads 0.63 as dollars. Read-only by construction: no order path.
- **`research_papers`** — arXiv + OpenAlex search/read with source-anchored claim extraction. Anything not literally present in the source is left empty and marked `not stated in source`, and a paper's claimed performance is never presented as our backtest result.

---

## 🧾 Governance wired into every run

- A **run manifest** hashes the prompt, the skill contents, the tool registry and the package versions — so "what methodology produced that number?" is answerable a month later.
- The **audit ledger** chains each record to its predecessor's hash and fsyncs. Editing or deleting a record is detectable, and an edit that recomputes its own hash is still caught one record later via `prev_hash_mismatch`. Timestamps are always caller-supplied; no module here calls `datetime.now()`.
- **Trace redaction is sink-aware**: `content` is released only in the tool-RESULT sink and stays redacted in the fail-closed ARGUMENTS sink used by tool-call arguments and the live audit ledger. `env` is never released. Result strings are pattern-scrubbed, since shell output arrives as a JSON envelope.
- All **30 swarm presets** were re-audited — a deliverable no granted tool can compute is now declared as such instead of invented.

---

## 🧱 Platform

- **Desktop shell** — a source-first Electron host owning the backend lifecycle: random loopback port, per-launch secret, five-locale startup recovery, owned-process cleanup (#923). Windows packaging assembles a **checksum-pinned embedded Python 3.12 runtime** with x64 NSIS review/signing paths, plus Electron `safeStorage` for an allowlisted credential set — the renderer can set or clear secrets but never read them, plaintext config migrates once, and both unsigned-review and signed builds fail closed on the wrong signature state (#1015). *No installer artifact was published from that PR.*
- **eToro** joins as the **13th broker connector** with path-separated demo/real profiles; demo keys structurally reach only `/demo` paths (#989).
- **Korea (KRX: KOSPI/KOSDAQ)** becomes the **9th backtest engine** — execution-time ±30% band on the unified tick grid, structurally long-only, config-driven 2026 0.20% securities transaction tax (#693).
- **Canadian equities** end to end — `.TO`/`.V` classified in CAD, Yahoo → yfinance → local, Canada-specific GlobalEquity rules, `XIC.TO` benchmark, mixed-currency aggregation refused (#1024, #1019, #1037, closes #952).
- **OpenBB Workspace bridge** (#817) and a read-only **Taiwan snapshot** tool (#848).
- **`src/entities`** — a typed entity + irregular dated cash-flow substrate for NAVs, capital calls and coupons, deliberately parallel to the bar engines so a `nav` column can never reach one and get priced as a close. Surfaced by `cashflow_performance`.
- **`orderbook_depth`** — crypto L2 ladder via ccxt: spread bps, depth imbalance, and impact cost of a stated notional walked through the real book both ways.
- **ModelScope** joins the built-in providers (#1011); **`vibe-trading update`** self-upgrades and distinguishes wheel installs from editable checkouts (#1020); **Settings gains live model discovery** and records the immutable provider/model/reasoning identity that actually served each reply (#924).
- **The MCP surface grows to 70 tools.** `alpha_zoo` + bounded `alpha_bench` (#979) and **QVeris** discovery/inspect/execute (#976, cost quote read from the marketplace rather than trusted from the caller) join, and so do six read-only analytics tools that had reached the agent but never MCP — `quantlib_call`, `cashflow_performance`, `orderbook_depth`, `sentiment`, `technical_indicators` and `get_fundamentals`. Order-placing tools remain structurally un-exposed: the mirrored registration path refuses any class whose `is_readonly` is not `True`.
- **Memory Tier 2** — hierarchical routing, BM25 semantic linking, TF-IDF compression and an FTS5 index with CJK bigram tokenization, behind a one-line `VT_MEMORY=off|on|full` preset. All off by default (#815, #733).
- New read-only tools: **`sentiment`** (#939), **`technical_indicators`** (#921), analytic **`options_payoff`** (#946).

---

## 🔒 Sandbox

A closed gap: generated strategy code could **import the broker layer**, and could reach `socket` / `subprocess` / `os.system` / `ctypes` **through a renamed binding**. Both were accepted before; both are now refused. `src.quantlib` still imports.

---

## ✅ Correctness — the expensive ones

- **SEC reporting periods are keyed on their `(start, end)` span.** A 10-Q files the true quarter *and* the year-to-date frame under the same end date and fiscal period, so `period="annual"` had been returning a single quarter for AAPL FY2018–2020 — **a 4.2× understatement** — and every fiscal-Q4 slot in a quarterly series carried the full-year figure.
- **Tushare A-share prices are corporate-action adjusted** in both the factor bench and backtests. A raw close-to-close return across an ex-date was off by **up to 47 percentage points** (300750.SZ, 2023-04-26). The CSI300 bench also masks each date to its point-in-time index membership.
- **`bar_returns` no longer erases the move across a trading halt** longer than the forward-fill window — the resumption move was silently recorded as 0, understating volatility and inflating Sharpe.
- **Annualisation now covers all 24 data sources** at every interval, with a coverage test that fails CI when a loader lands without entries.
- **Cross-market composite backtests refuse a mixed-currency code set** instead of summing CNY, USD and KRW into one equity curve.
- **Option legs are marked at the volatility they were opened at**, removing a fabricated day-zero P&L of up to +93% of premium.
- A 19-PR **interval-normalization sweep** — lowercase `1h/4h/1d/1w` accepted everywhere, unsupported intervals fail fast instead of silently returning daily bars.
- **Two quantlib modules were allowlisted but unreachable.** `attribution` (Brinson-Fachler) and `impact` (market-impact models) carried no `__all__`, and `quantlib_call` dispatches on `__all__` alone — so the tool listed **zero** functions for both while the package docstring advertised them. Both now export, and a guard test fails when any allowlisted module exposes nothing; the existing test only asserted that each module *imports*, which could never catch this.
- **The published MCP manifest under-reported the server.** `SKILL.md`'s tool count was derived by counting `@mcp.tool` decorators, which ignores every tool registered through the mirrored path — so four institutional-research tools were live over MCP and absent from the manifest, and the contract test asserted that absence was correct. Both tests now measure `mcp.list_tools()`.
- **Resource leaks closed**: HTTP throttle sweeps stale buckets interval-aware, the rate limiter no longer grows unboundedly with unique client IPs, the event bus notifies and removes subscribers on clear, and `_json_loads` is guarded against corrupted JSON in database columns.

---

## 🙏 Contributors

@santhreal · @shadowinlife · @Robin1987China · @he-yufeng · @QCYTSN · @Shizoqua · @honginp · @cgycorey · @wiliao · @ngoanpv · @x-lambda · @ofeksh-tr · @00EVA · @zwrong · @yrk111222 · @su322 · @hhj123123 · @dineeshd · @sambazhu · @ddy4633 · @tyj147454413-cmd · @y85998607 · @JungHoonGhae · @shugaoye · @TSENGCHIENFENG · @darkknight4563 · @MuggleJinx · @klmtseng · @ebujinovch · @g0rdonL · @AmirF194 · @Echoandelementwebsites · @yagnikpipaliya · @dvirarad · @1anter

---

**Full changelog:** https://github.com/HKUDS/Vibe-Trading/compare/v0.1.12...v0.1.13