I build AI systems that prove they work before they ship.

Graph-grounded answers that cite their sources, gateways that give models tools without handing over keys, voice agents that let you interrupt, and routers that cut the bill. Each one measured, reproducible, and honest when it fails.

Six systems, one discipline: measure first.

Six systems for putting models into production.

Hover or focus a panel to open it. Each links to its own design.

Multi-hop answers with a citation path

Regulatory filings become a property graph of clauses, amendments and enforcements. Queries traverse the graph and the vectors together, and a reflection loop refuses to answer without source-node citations.

  • LangGraph
  • Neo4j
  • Qdrant
  • Cohere Rerank
Open design

Tools for models, never the credentials

An MCP resource server behind OAuth 2.1 with mandatory PKCE. Tool discovery is filtered by grants, a server-side broker holds every upstream secret, and mutating actions pause for human approval.

  • FastMCP
  • OAuth 2.1
  • Redis
  • PostgreSQL
Open design

Barge-in under 500 milliseconds

A full-duplex WebRTC agent with edge voice-activity detection. The playback buffer clears on interruption, tools run mid-sentence behind a filler, and packet loss is concealed rather than heard.

  • LiveKit
  • Silero VAD
  • Cartesia
  • Twilio SIP
Open design

Continuous evaluation with a calibrated judge

Golden datasets, Ragas and DeepEval metrics, regression detection against a baseline, and a CI gate that blocks the merge. It replays from recorded cassettes, so CI runs without an API key.

  • DeepEval
  • Ragas
  • Langfuse
  • GitHub Actions
Open design

A web agent that shows its working

Vision-grounded navigation over legacy portals: numbered marks, DOM pruning and coordinate clicks. Every screenshot, mark and decision is recorded, so a run can be replayed step by step.

  • browser-use
  • Playwright
  • Claude Computer Use
  • Pydantic
Open design

The small model answers until it is unsure

A QLoRA-tuned Llama 3.2 3B served by vLLM, behind a LiteLLM router that escalates low-confidence requests to a frontier model. Most traffic never leaves your own hardware.

  • Unsloth QLoRA
  • vLLM
  • LiteLLM
  • NVIDIA L4
Open design

Evaluation is a gate, not a report

Scores are shown against their threshold, and a regression blocks the merge.

baselinecandidate0.85 faithfulness bar

Frontier quality, small-model cost

direct frontier$15,000
SLM + router$2,510

Multi-hop, with receipts

34%91%

Vector-only versus graph-grounded accuracy, as claimed in the project brief.

A conversation, not a queue

2–4s<500ms

Cascaded HTTP voice bot versus streaming WebRTC turnaround.

A model in production is a hypothesis. I treat it like one: write down what should be true, build the instrument that measures it, record every run so it can be replayed, and let the measurement decide what ships. It's slower for a week and much faster for a year.

query 07 · 3 hops · 2 statutes The answer traces AMENDS and ENFORCES edges to the clause it cites. No path, no answer.
barge-in · buffer cleared The caller cuts in mid-sentence. The agent stops before the next word reaches the speaker.
1,000,000 requests · 85% local The modelled bill falls from $15,000 to $2,510 a month, and the routing rule is on the page.

Let's make your model measurable.

Compliance retrieval, agent tooling, voice, evaluation and cost-aware model routing.

Start a conversation