I build agent runtimes that run real operations — and the evaluation tooling that proves whether they actually work.
For eleven years I've built production software — distributed data infrastructure at Amazon Web Services, 10TB/day pipelines at a real-money gaming company, and a developer-tools startup I co-founded and led as CTO.
Today I build Kavi, a multi-tenant AI agent runtime that runs a company's content, ads and research operations — 230 typed capabilities and 22 durable workflows, with tenants isolated by a database-level wall rather than by convention.
The thread running through it: I care whether the measurement is real. Most of what gets reported as an AI improvement is noise that nobody put error bars on. I write tools that tell the difference.
The two best coding agents on SWE-bench Verified are tied at 79.2%. They disagree on 36 tasks and split them 18–18. That is not a ranking — that is a coin landing heads 18 times out of 36.
97% of adjacent leaderboard pairs are statistically indistinguishable
The benchmark cannot resolve anything under 5.4 points
Wilson intervals, exact McNemar, minimal detectable difference
Zero compute — reproducible from data already published
I'm open to work on AI evaluation, agent infrastructure, and hard measurement problems.
The fastest route is email: tapish303@gmail.com