Measured
Baseline recorded before deployment, same metric after, same method.
No logo wall. Several of these clients can't be named, and our own deployments are the ones we can describe in full detail. So each entry says what was measured, what wasn't, and what we'd change — including two that didn't work.
Most agency case studies quote a percentage with no indication of how it was arrived at. We label each figure so you can weigh it: measured against a recorded baseline, estimated from sampling, or reported by the client. A reported figure isn't worthless — it's just not the same as a measured one.
Baseline recorded before deployment, same metric after, same method.
Derived from a statistically meaningful subset, not the whole population.
Stated by the client. We believe it; we didn't independently verify it.
B2B software company, 4,200 tickets a month, median first response over eleven hours. The team wasn't slow — every reply required checking four systems. We put the research in front of the reply.
Read the deploymentTwenty-two reps, a CRM updated the night before each pipeline review, and a quarterly forecast variance the CFO had stopped taking seriously. The data wasn't wrong so much as retrospective.
Read the deploymentOur own delivery organisation: meeting decisions vanishing into transcripts nobody read, status updates written by hand every Friday, onboarding chased over WhatsApp. Everything in the registry earned its place here before we sold it.
Read the deploymentClient can't be named and the numbers are theirs, not ours. What we can describe is the shape: exceptions explained in a sentence rather than dumped in a queue, and a hard rule that the agent never posts to the ledger.
Read the deploymentFour weeks in, the agent was abstaining on nearly half of all questions. It wasn't broken. The handbook contradicted itself across three versions, and the honest answer was that no system could resolve what the organisation hadn't decided.
Read what went wrongA Gulf-based operator whose entire automation stack stalled on email approvals nobody opened. The agent didn't change. The channel did — and the workflow started completing the same day instead of the same week.
Read the deploymentThe agent worked. Accuracy was good, the drafts were useful, and the reps stopped using it anyway — because leadership had announced it in the same meeting as a headcount review. A lesson about sequencing, not engineering.
Read what went wrongForty engineers, eleven repositories, and the same five review comments repeated forever. We made the correct structure the default rather than the requirement.
Read the deploymentSwipe →
Connecting the API is never the hard part. Making the agent behave sensibly on the twentieth edge case is where the timeline actually goes.
The single strongest predictor of success across all eight. Where documentation was current, agents worked. Where it was contradictory, no amount of engineering rescued it.
In none of these deployments did a client reduce staff. What changed was what those staff spent the day doing, and how long customers waited.
Teams told the agent would help them used it. Teams who suspected it was measuring them found reasons not to. CS-07 is the whole argument in one deployment.
Any vendor claiming eight successes out of eight is either new, selectively remembering, or counting delivery rather than adoption. Projects fail — usually for organisational reasons that were visible before anyone wrote code.
Both failures here were preventable, and both are now questions we ask during the Sprint. That's the actual value of documenting them: CS-05 became our rule about written knowledge, and CS-07 became our rule about how a deployment is announced internally. You benefit from two clients' bad weeks.
See what the Sprint checks for →Start with five days and one workflow. We'll tell you honestly which column you'd end up in — and if it's the wrong one, we'll say that instead of selling you a build.