AI / Wanderings 2026
The AI Diligence Test: Follow One Customer Outcome
AI can make an 18-person company look like an 80-person machine. Before you invest, trace one accepted customer outcome through every model call, retry, human review and failure path.

An AI company can look excellent for six months.
Revenue jumps because every customer wants a pilot. Product ships weekly. A team of 18 appears to do the work of 80. Gross margins look attractive because inference, retries, implementation, human review, refunds and the founder answering support tickets at midnight sit in different parts of the accounts.
Then the first renewal cycle arrives.
That is when the demo story meets the operating system.
Investors still need to understand the market, product, team, customers and cash. But AI has moved the assumptions beneath those familiar questions. It changes which costs scale, how quickly competing products converge, where defensibility sits and how an operational failure can spread.
It also allows a fragile service business to look like software.
The most useful diligence test is therefore simple: choose one accepted customer outcome and trace it all the way through the company. Follow every model call, retry, human intervention, exception, contractual promise and failure path required to produce it.
If that workflow works economically and reliably, conviction can grow. If nobody can map it, the uncertainty belongs in the valuation.
Start with one accepted customer outcome
Gross margin is unusually easy to flatter in an AI business.
The model bill appears in cost of sales. Implementation sits in product. Human review sits in customer success. Failed outputs disappear inside a pilot, supported by people whose time has somehow become free.
The spreadsheet may be correct and still tell the wrong story.
Start with one real production unit: a claim processed, a ticket resolved, a document reviewed, a qualified lead produced or a transaction completed. Then attach every cost required to deliver an acceptable result.
Include model calls, retrieval, storage, retries, fallback models, human review, integration work, support, compliance and incident handling. Include the cost of mistakes.
A model call that costs three cents and triggers a €50 manual review is not a three-cent transaction. Finance departments tend to notice this later.
The right denominator is the accepted customer outcome. Cost per request tells us little if requests must be rerun, checked or repaired. Ask how contribution margin changes as volume rises and edge cases accumulate.
Pricing deserves the same scrutiny. Seat-based pricing becomes peculiar when the product reduces the number of people required to do the work. The customer may need fewer seats precisely when the product becomes more useful.
Outcome pricing can fit better, although it introduces attribution disputes, cash timing and contract risk. The diligence question is whether price tracks customer value and gross profit improves as usage grows.
Find out whether speed produces learning
AI compresses build time. It also compresses copy time.
A polished demo says less than it did a few years ago. A competent team can assemble one quickly with models and infrastructure available to every other competent team.
The more useful evidence is how quickly the organisation turns customer experience into a reliable product change.
Compare the last six months of releases with customer complaints, evaluation results, rollbacks and incidents. Ask which decisions were reversed. Measure how long it took to move from a customer problem to a tested fix.
I would also ask what the team stopped building. A company that only shows acceleration may have speed and no judgement.
Every model, prompt or workflow change can alter quality in places the team never touched directly. AI-native companies therefore need test cases drawn from real usage, clear acceptance thresholds, version tracking and a named owner for the result.
Fast releases with weak evaluation create invisible liabilities. They reach customers before they reach the board pack.
Identify what the company actually owns
Foundation models make new capabilities available to many companies at once. A product lead can disappear when a provider releases a better model or a competitor combines the same components with stronger distribution.
Map the full stack:
- Which models provide the underlying capability?
- Which orchestration, evaluation and workflow layers has the company built?
- What customer knowledge or proprietary data improves the product?
- Which integrations make the product part of daily work?
- What distribution, permission or trust would take a competitor time to reproduce?
Code still matters, but code volume is weak evidence of defensibility. A large codebase can also be an archaeological site.
Look for advantages that strengthen with use. Each customer interaction should improve routing, evaluation, workflow design or commercial understanding. Customer data sitting unused in a warehouse creates storage bills. It becomes an advantage only when the company can turn it into a better result within the limits of privacy, contracts and regulation.
Model dependency belongs in the financial model. What happens if the provider doubles its prices, changes its terms, limits capacity or releases a competing feature?
“Model agnostic” is easy to put on a slide. Proving it requires a live switching test, including the time, cost and quality loss involved.
Diligence the revenue twice: pilot and renewal
AI creates a peculiar commercial pattern. Curiosity can make the first meeting and pilot relatively easy. Security reviews, workflow changes, budget ownership and employee adoption make production much harder.
A paid pilot is positive evidence. Renewal is stronger.
Review customers by cohort. Track who signed, activated, reached production, expanded, renewed and left. Separate contracted value from real usage. Measure the support burden for each customer and whether it falls as the product matures.
Customer reference calls should answer five direct questions:
- What work changed after the product went live?
- What measurable result did the customer receive?
- Where does a person still need to step in?
- Who owns the budget and the renewal decision?
- What would happen during the first week without the product?
The happiest reference customer may be using a founder-supported version that no other customer receives. Compare the reference call with usage data, support history and contract terms.
Revenue quality also depends on how much service work is hidden inside the software price. Implementation can be perfectly sensible. Pretending it scales like software becomes expensive.
Inspect the work design, not the headcount
A small team can produce remarkable output with AI. Headcount reduction alone tells us little about quality, resilience or how many people are cleaning up behind the machine.
Look at the design of the work. Who sets the goal? Who checks the output? How do exceptions reach a person? Who can stop the system? What happens outside normal working hours?
Ten people may be enough. Ten exhausted people plus an unmonitored agent create a rather different investment.
AI can also increase key-person risk. If one engineer understands the prompts, evaluation set, vendor relationships and failure cases, the apparent efficiency depends on one person’s calendar.
Inspect documentation, access controls, incident ownership and whether another person can operate the system. Ask management which decisions remain with people and why.
A system drafting an internal note can have broad freedom. A system issuing refunds, changing medical records or moving money needs very different boundaries. Good operators assign autonomy according to the cost and reversibility of failure.
Price risk by consequence, autonomy and scale
Generic AI risk lists become useless when everything receives the same red label.
Evaluate each workflow on three factors:
- What happens when the output is wrong?
- How much authority does the system have?
- How widely and quickly can the error spread?
A support tool suggesting a reply and an agent authorised to send payments do not deserve the same controls.
For every material workflow, document the data entering the system, the actions it can take, the human checkpoints, the audit trail and the rollback process. Review model changes, evaluation reports, security tests and the incident log.
Zero recorded failures can mean excellent controls. It can also mean nobody is looking.
Agentic systems need particular attention because the risk moves from bad text to bad action. Once an agent can send emails, edit a CRM, approve refunds, access customer records or deploy code, permissions become part of the investment case.
Business continuity belongs here too. What happens during a provider outage? Can the company move to another model? Which customer promises depend on a capability controlled by somebody else?
Read the contracts beside the architecture. A company may have promised a level of accuracy, security or availability that its current system cannot consistently provide.
Test the live system
The data room captures the version of the company that management chose to document. An AI business reveals far more through its behaviour.
Run representative workflows, including boring edge cases. Follow one customer task from input to accepted result. Compare output quality across product versions. Reconstruct the cost and inspect the manual work around it.
Use AI during diligence too. It can search contracts for inconsistent obligations, group support failures, map technical dependencies and analyse customer cohorts quickly. The investment committee still owns the judgement.
Every AI-generated diligence finding should link back to its source. A confident paragraph without provenance is another risk factor dressed as productivity.
Keep a dated evidence pack. Record which model, product version and configuration were tested, because the system may behave differently by the time the investment committee meets.
Eight questions before the investment committee
Before approving an investment or acquisition, I would want clear answers to these questions:
- What is the production unit of customer value?
- What is the fully loaded contribution margin per accepted outcome?
- Which advantage becomes stronger with every customer?
- How long and how much would it cost to change the main model provider?
- What evidence connects pilots with production use, expansion and renewal?
- Where does human judgement remain, and what does it cost?
- Which machine actions could cause expensive or irreversible harm?
- What did the team stop, change or learn during the last six months?
Build the one-page workflow before you decide
Add one production-workflow page beside the financial model. Show every machine cost, human intervention, exception, contractual promise and failure path attached to one accepted customer outcome.
If the economics work and the company learns faster without accumulating uncontrolled risk, the investment deserves deeper conviction.
If management cannot build the page, price that uncertainty before you approve the deal.