what's checkable, and what isn't yet.
this page exists because everyone in this category claims the same things. below is what you can go and verify yourself, and the numbers from the work so far.
a yc-backed us startup runs agent systems i built, against live customer traffic.
not a pilot, not a demo environment. real requests from real customers, with the failure handling that implies.
artifact.engineer →the cto of mastra, the ai agent framework, referred me personally.
after roughly ten merged pull requests to the framework itself. the contributions are public and dated; so is the referral.
the pull requests →candidate sourcing was entirely manual — hunting for candidates by hand, one role at a time.
- an existing product replaced it, at a fraction of their prior sourcing cost
- found in two days, as part of a simplification audit
- the time went straight into calling and booking client calls
every quote was assembled by hand from three systems, and the friday exposure report was rebuilt from scratch each week.
- custom build — nothing off the shelf fit
- six weeks to handover
- monitoring, runbook and training included
clients are anonymised because most of them would rather not advertise which parts of their operation were manual. the numbers are the part that matters.
the counters start at the first audit and are updated as engagements close. every “don't build” outcome gets a write-up on this page: the situation, what i recommended instead, the named product and its real price, and what i didn't earn. a number on its own is a claim like any other — the write-ups are the part worth reading.
want the awkward version?
ask me on the call what i've got wrong, or which engagement i'd do differently. it's a shorter answer than you'd expect, and it's a real one.