• 349 commits on main in three weeks
  • Windows portability from an outside contributor, 5 merged PRs
  • Reads traces from TypeScript, Python, Go and Rust with no SDK

A local-first debugger for AI-agent runs. Point an OpenTelemetry exporter at it and every run becomes a timeline, a span tree, a conversation and the exact tool inputs and outputs behind a result. A captured failure becomes a versioned regression case in one click, and the same checks gate CI through JSON and JUnit reports. Capture, search and deterministic evaluations all run on your machine, with no hosted account and no provider key.

A local daemon ingests OTLP over HTTP on 127.0.0.1:5947 and stores every trace in SQLite on the machine that produced it. The interface reads each run as a timeline, a span tree, a conversation, the tool inputs and outputs, and the errors. It speaks the OpenTelemetry GenAI and OpenInference conventions, so the example agents in TypeScript, Python, Go and Rust connect with no Run Phantom SDK at all.

Comparison is the core move. Two runs line up across inputs, outputs, tools, model, usage and timing, and a repeated call that cannot be matched one-to-one is reported as ambiguous instead of guessed. Evaluate run turns a captured failure into a versioned regression case — an expected response, a JSON value, a tool call, an argument, an error condition or a resource budget — and an experiment freezes the dataset revision, the candidate evidence and the evaluator versions, so a result means the same thing next week.

CI gets a contract, not a dashboard. The check exits 0 when everything passes, 1 on a failure or an inconclusive result, and 2 on an operational error, and an inconclusive result lands in JUnit as an error, never a skip. Repeated trials run 2 to 20 experiments and report pass, fail and unknown counts with missing-evidence bounds. Replay sends a fresh request to a project-owned endpoint, and Ask Claude Code opens a trace-aware session straight from the run.

Under it: Bun, Express, Drizzle and SQLite behind a React 19 interface, an MCP server that exposes the traces to other agents, a browser verification workspace, and a team service with Viewer, Editor and Admin roles and write-only ingestion keys. Saved checks keep their frozen evidence after the live trace is deleted. 84 test files cover the daemon, the evaluations and the interface, with Playwright end to end. A recorded source audit read 25 external projects at pinned commits — Langfuse, Arize Phoenix, inspect_ai, promptfoo, OpenLLMetry and DeepEval among them — and copied none of their code.

349 commits landed on main in three weeks, and an outside contributor made it portable to Windows across five merged pull requests.