Runtime reliability of LLM agent systems · Est. 2026
I make production LLM-agent systems reliable — and cheap. Field notes from operating them: real traces, real invoices, real 2am incidents.
Operator, not researcher. If a language model could have written the post from public knowledge, it isn't here — every piece carries at least one number that could only come from running the thing.
Five agent harnesses, one DeepSeek Flash model: same answers, 4.46× the bill
Same LLM, five coding harnesses, 120 runs: 80% of the token-cost gap between Claude Code, Codex CLI, aider, opencode and pi is scaffolding, not work.
Two Claude Code accounts on one laptop
CLAUDE_CONFIG_DIR moves the whole config directory, not just the login. Two directories, two shell functions, done.
No posts with that tag yet.
Agent system unreliable or expensive?
I've probably seen the failure mode. Reliability & cost audits, embedded engineering, and advisory.