AI Systems Engineer

I build agent runtimes that run real operations — and the evaluation tooling that proves whether they actually work.

TR

Hi, I'm Tapish.

For eleven years I've built production software — distributed data infrastructure at Amazon Web Services, 10TB/day pipelines at a real-money gaming company, and a developer-tools startup I co-founded and led as CTO.

Today I build Kavi, a multi-tenant AI agent runtime that runs a company's content, ads and research operations — 230 typed capabilities and 22 durable workflows, with tenants isolated by a database-level wall rather than by convention.

The thread running through it: I care whether the measurement is real. Most of what gets reported as an AI improvement is noise that nobody put error bars on. I write tools that tell the difference.

Where I've worked

Khiladi ProVibinexMegashots (GetMega) Amazon Web ServicesNetskope / Sift SecurityNebulaa Innovations
FEATURED WORK

evalstats

The #1 spot on SWE-bench is a coin flip

The two best coding agents on SWE-bench Verified are tied at 79.2%. They disagree on 36 tasks and split them 18–18. That is not a ranking — that is a coin landing heads 18 times out of 36.

97% of adjacent leaderboard pairs are statistically indistinguishable

The benchmark cannot resolve anything under 5.4 points

Wilson intervals, exact McNemar, minimal detectable difference

Zero compute — reproducible from data already published

Read the essay View the code

Education

M.S. EECS · University of California, Merced B.Tech CSE · IIT Jodhpur

Interested in Collaborating?

I'm open to work on AI evaluation, agent infrastructure, and hard measurement problems.

×

Get in Touch

The fastest route is email: tapish303@gmail.com