Runtime reliability of LLM agent systems · Est. 2026

I make production LLM-agent systems reliable — and cheap. Field notes from operating them: real traces, real invoices, real 2am incidents.

Operator, not researcher. If a language model could have written the post from public knowledge, it isn't here — every piece carries at least one number that could only come from running the thing.

Agent system unreliable or expensive?

I've probably seen the failure mode. Reliability & cost audits, embedded engineering, and advisory.

Work with me →