The structure of a rescue engagement.
Whether it is an enterprise AI project or an AI vendor's strategic account, the arc is often the same: diagnose where the system actually stands, build the checks that live outside the AI, calibrate who sees what, and prove the value in the owner's terms. What differs is on whose behalf we operate.
We start with a single business workflow and build from there.
Most AI reviews begin with either the technology: audit the model, benchmark the vendors; or the organisation: interview everyone, map everything. Both these approaches take months, and neither delivers the data the sponsor needs to decide what to do. We start by following one key business process end to end, and asking where correctness gets decided, by whom, and what happens when it is wrong. The answer usually explains the last six months of internal debate, and often takes days to find.
Follow one key business workflow
One real case, traced from input to consequence. Where does correctness get decided, who owns it, and what happens when it is wrong? Nearly every struggling initiative gives the same answer: nowhere, no one, and nothing.
Build a check that lives outside the AI
Verification that cannot fail the way the model fails: the totals a document prints itself without determined arithmetic, a ledger that already closed, what the technician actually found. Wrong answers stop looking identical to right ones, and the magnitude of the gap points to a root cause.
Calibrate, then demonstrate
The harness clears the routine cases; the doubtful ones reach your people with the problem located. Thresholds are set with the executives who own the risk, and the result is reported in the sponsor's terms: the share of work running unreviewed, the hours returned, the errors caught.
AI reads the documents no software could. That is the benefit as well as the risk.
The value of these systems is that they read what traditional software never could: emails, statements, work orders, contracts, spec sheets, call recordings. The risk lives in the same sentence. Unstructured inputs are exactly where a model misreads quietly, e.g. a column shifted, a sign flipped, a clause skipped and the output looks exactly like a correct one. Every source below is somewhere we have seen it happen, and somewhere a harness has something solid to anchor to.
Where these failures startThe harness decides what a person looks at.
Reviewing everything an AI produces saves nobody any time; reviewing nothing means every error ships and keeps shipping until a customer finds it. The working answer sits between the two. The harness clears the routine cases, and the ones that fail its check reach a person with the problem already located, so the review takes minutes rather than starting from scratch.
Where that threshold sits is a business call. An irreversible action deserves a human even at high confidence; a reversible one can run. Setting those thresholds with the executives who own the outcome is part of every engagement, and it is usually the first time anyone in the room has had to name them.
AI Investment Rescue
Your AI is live. The value still doesn't appear on the radar.
For the executive who owns an AI spend that is under scrutiny. The initiative was reasonable, the tools are running, and the question in the room has changed to what it returned. We find what is holding the system back and turn it into production-worthy AI your people can trust; measured in your terms instead of the vendor's.
One key process, end to end
We trace real cases through the system and find where the failure often appears, what is recoverable, and what is beyond saving. Within days, you decide what happens next with a map in hand.
The checks the system shipped without
Verification that lives outside the AI, with failure modes the model does not share. The doubtful cases presented to your users with the failure point marked.
Live cycles, then the number
Thresholds set with the executives who own the risk, adjusted on live results, and reported as what the CFO actually asked for: the share of work running without review, the hours returned, the errors caught before they shipped.
- Value diagnostic · days · fixed fee
- Rescue engagement · 6–12 weeks
- Embedded partner · 6–12 months
Strategic Account Rescue
Retain your largest customers.
For the founder or delivery leader watching a flagship enterprise deployment slip. The customer is escalating, engineering is firefighting, and the reference account the next raise depends on is starting to feel fragile. We embed as your senior delivery team until the account holds on its own.
The delivery assessment
Days inside the deployment: what is actually failing, whether it is recoverable, and what it will take. You get a concrete recovery plan instead of another status meeting.
Senior hands, on both fronts
Forward-deployed engineers working with the customer's teams and yours — and someone senior in the steering committee who can explain what happened, what the plan is, and why continuing to invest is right.
Your team keeps the playbook
Documentation, delivery playbooks, and implementation patterns, so the next deployment runs without us. We do not build a dependency on ourselves.
- Delivery assessment · days
- Account recovery · 6–12 weeks
- Delivery partner · ongoing
Built right from day one
The cheapest rescue is the one you never need.
For teams starting an AI initiative now, with the budget approved and the vendor shortlist open. Most of what we rescue was designed without verification, calibration, or a human handover — none of which is expensive at the start and all of which is painful to retrofit. We design them in from the first architecture session.
The harness in the architecture
Where correctness will be decided, what the checks anchor to, and where the human sits — settled before the first line of code, while all of it is still cheap.
Proven on live work
Built into the process your team actually runs and validated on real cases before it scales. What works is rolled out; what does not is cut early, while cutting is still easy.
Trust from the first week
Adoption is designed in as a constraint, and the sponsor has a value number they can defend from the first quarter.
- Architecture sprint · 2–3 weeks
- Build & deploy · 3–9 months
- Fractional AI lead · 6–12 months
How the method actually works, in practice.
Why start with one business workflow instead of auditing the whole AI programme?
Auditing the model or interviewing everyone both take months and neither delivers the data a sponsor actually needs. Following one real case — a statement, a quote, an alert — end to end, and asking where correctness gets decided and what happens when it's wrong, usually explains the last six months of internal debate, and it takes days rather than months to find.
What is 'the harness,' and why can't the AI check its own work?
The harness is whatever sits outside the model and verifies its output without sharing its failure modes: the totals a document prints on itself, a ledger that already closed, what a technician actually found. An AI checking its own output, or a second model grading the first, shares the same blind spots as the system it's checking — which produces a confident answer, not a correct one.
How do you decide what the AI can do on its own versus what needs a human?
The harness clears the routine cases and routes the doubtful ones to your people with the problem already located. The threshold — what runs unreviewed, what waits for a human — is set with the executives who own the risk and adjusted as live results come in, because that line is a business decision, not a technical one.
Can this be built in from day one instead of used only to rescue a stalled project?
Yes — that's the third door. For teams starting an initiative now, we design the harness, the calibration loop, and the human handover into the architecture from the first session, so the system earns trust in its first week in production instead of losing it slowly over months. It's the same discipline, applied before there's anything to rescue.
Wherever you start, the discipline is the same.
Diagnose, harness, calibrate, demonstrate — one method across all three doors, with a partner in the room rather than a deck in your inbox. The person who scopes your engagement runs it. We don't hand off.
Schedule a callWe generally first engage with the executives who own the outcome before procurement teams.