Production Agents 2026
MCP connects agents to tools. A2A coordinates independent agents, including through stateful tasks. Learn where each belongs and what neither protocol solves.

I design retrieval, agent, and evaluation systems to stay grounded in evidence, bounded by clear controls, and measured against reality when the happy path ends.
See how I workReliability loop
Evidence before confidence
What I work on
The work spans the full reliability stack: finding the right evidence, coordinating model behavior, measuring quality, and translating research into products people can trust.
Designing retrieval-augmented systems that ground model output in real sources: hybrid retrieval (lexical + dense), semantic chunking, reranking, and citation-grounded answers.
Building agent systems that plan, call tools, and manage state reliably, with deterministic guards, retries, and explicit boundaries instead of hopeful prompting.
Treating evaluation as engineering: golden datasets, faithfulness and groundedness metrics, regression gates, and adversarial testing so quality is measured, not assumed.
Bringing PhD-level research rigor to production: information retrieval, network analysis, and adversarial/applied machine learning, translated into systems that ship.
Latest writing
Field notes and long-form essays on building AI systems that are grounded, bounded, and measurable.
New research series
Production Agents 2026
Five source-led guides to protocols, autonomy, evaluation, security, and memory.
Production Agents 2026
MCP connects agents to tools. A2A coordinates independent agents, including through stateful tasks. Learn where each belongs and what neither protocol solves.

Production Agents 2026
Build agent memory as governed state: control writes, preserve provenance, supersede stale beliefs, enforce deletion, and verify actions against evidence.

Production Agents 2026
Treat prompt injection as a source-to-sink security problem. Contain the agent, limit authority, and gate consequential actions outside the model.

Production Agents 2026
Evaluate the whole agent system: verify final state, inspect traces, control the environment, repeat trials, and turn real failures into regression tests.

Production Agents 2026
Task horizon, uninterrupted runtime, and parallel agent-hours are different clocks. Learn how to read agent autonomy research without overstating it.

Start a conversation
I'm always interested in thoughtful conversations about retrieval, agents, evaluation, and AI for learning.