Patterns
Pattern 6

Who owns the evals

The software is replaceable. Eighteen months of your best people's judgement, written down and handed over, is not.

Rohit Chikballapur · 22 August 2026 · 4 min read

The most valuable thing in your company rarely sits in a document. It's the reason an underwriter reads one file twice and signs off another in ninety seconds, the way a maintenance engineer knows from the sound that a bearing has about three weeks left in it, the four questions a good analyst asks about a set of numbers before deciding whether to believe them.

None of that is written down anywhere and some of it can't be. It's the residue of people being wrong in expensive ways and adjusting, over years, and it's the actual difference between two companies that look identical on paper.

Any serious AI project has to get at this knowledge, because a system that doesn't encode it will be confidently wrong in precisely the places where your people are careful. There's no route around the extraction. The only real question is where it ends up.

From the inside the process looks entirely benign. A vendor arrives and everyone is cooperative, because everyone wants the project to work. Your best people are made available, correctly, since they're the ones who know. They walk through the process, explain the exceptions, provide examples of hard cases, and explain why each one was decided the way it was. When the system gets something wrong in testing, they explain what a right answer would have looked like and why.

That last activity is the important one and it has a name in the trade. Those are evaluations. Test cases with correct answers attached, plus the reasoning that makes them correct. They're what the system gets tuned against, they're what tells you whether a new model version is an improvement or a regression, and they are the single most expensive artefact in any AI deployment, because the only way to produce them is to spend the time of the people who know. A serious system accumulates a few thousand over a couple of years, and every one of them is a fragment of your institution's judgement made explicit in a form a machine can check against.

Go and read your contract to find out who owns them. In most agreements I've looked at the answer is either the vendor or nobody. Sometimes there's language about customer data that clearly covers your documents and just as clearly wasn't drafted with test cases and labels in mind. Occasionally there's a clause granting the vendor rights to use "learnings from the engagement to improve the service", which is a sentence worth stopping on and reading twice.

This matters more than the software does, because the software is replaceable and the evals aren't. Switch vendors in year three and the application is a rebuild, which is annoying but tractable, maybe six months. The evals are eighteen months of your best people's time, and without them you're starting from zero with the same people, who are now older and busier and in one or two cases retired. That asymmetry, not the integrations or the switching costs in the ordinary procurement sense, is the real lock-in. You made your institutional judgement legible, once, at considerable expense, and it's sitting in a repository you don't control.

The vendor, meanwhile, has ended up with something considerably more valuable than your subscription. They have a working model of how a good operator in your industry thinks, generalised across you and everybody else who went through the same discovery. Which is why the second implementation takes half as long as yours did and the fifth takes a quarter of it. You funded the learning curve that gets sold to your competitor at a discount.

I don't think this is usually malicious, for what it's worth. Nobody sets out to strip mine a customer's expertise, and most of the vendors doing it would describe what they're doing as product improvement, which from where they sit it genuinely is. It's that the contract was written about software, and the thing of value turned out not to be software.

The fix is procurement, not strategy, and it's cheap if you do it before signing. Name the artefacts explicitly in the agreement, not as "customer data" but as the evaluation sets, the labelled exceptions, the failure taxonomy, the annotations your staff produced and the process logic derived from your operations. Assert ownership of all of it in machine readable form, delivered on request during the engagement rather than only at termination. Then require a periodic export and actually take it, because an ownership clause you've never exercised is one whose limits you'll discover at the worst possible moment. And restrict derivative use separately: a vendor getting generally better at its craft because it worked with you is unavoidable and fine, a vendor productising your specific exception handling is a different thing, and the line between those two belongs in the contract rather than in the relationship.

Six posts in, these are all the same argument, so let me say it plainly. Intelligence is becoming a commodity, sold by a handful of companies, improving on its own schedule regardless of what you spend on it. The scarce things are what surrounds it: knowing when the output is wrong, knowing which parts of the work should exist at all, and holding onto the encoded judgement of the people who've been doing it for twenty years. The industry is selling the commodity and quietly acquiring the scarce part as a side effect of the sale. That trade is available to renegotiate, and almost nobody is renegotiating it, mostly because it doesn't appear anywhere on the invoice.

That's the sixth pattern, and the last one for now.

The question this pattern answers

Who owns the evals when an AI vendor relationship ends?

In most contracts, the vendor does, or nobody — even though the evals (labelled test cases plus the reasoning that makes them correct) are the most expensive artefact in the deployment, built from a company's best people making their judgement explicit. The software is replaceable in months. Re-earning that judgement from zero is not. The fix is procurement, not strategy: name the evaluation sets and failure taxonomy explicitly in the contract, and require them exportable on request during the engagement, not only at termination.

The series

That is the last pattern, for now.

Read the series from the start

These posts come out of advisory work on AI initiatives that stalled. If one of them describes where you are, the first conversation is a straight read on whether it is recoverable.

Schedule a call