Imagine a company starts using AI to prepare its weekly marketing reports. The first few look good. They arrive faster, the writing is clear, and the team spends less time moving numbers between spreadsheets.
A few months later, those reports are helping decide where the company spends its money.
Somewhere along the way, a tool that helped prepare information became part of the decision-making process. I think that transition deserves more attention than it usually gets. The company has become more dependent on the system. Has its understanding of the system improved at the same pace?
In my previous article, I wrote about how AI changes the way service businesses deliver their work. This is the next question I would ask: once AI has a role in the business, what evidence tells us that it is doing that job well?
Start with the responsibility you have given it
An AI system that suggests ten possible headlines has a different responsibility from one that decides which customers receive attention from your sales team. You can discard an unhelpful headline immediately. A promising customer filtered out of a sales queue may never become visible to anyone.
The cost of an error depends on where the output goes and what people do with it.
Before choosing a performance metric, I would follow the work all the way through. What information enters the system? What does it produce? Who checks it? What decision follows? Who experiences the consequences?
Take the marketing report. An incomplete source file could lead to a misleading comparison. That comparison could become a confident recommendation. A manager could then reduce spending on a campaign that was performing well.
Every step might look reasonable to the person handling it. The problem becomes visible when you examine the whole chain.

Give “good” a meaning you can test
We use words such as accurate, reliable, and useful very easily. They become harder to use when somebody asks what would count as evidence against them.
Suppose a sales tool promises to identify promising prospects. What does that mean? Companies that match your target profile? People likely to reply? Buyers likely to become profitable customers? Each definition would produce a different evaluation.
I would start with a narrow claim: given the information available, the system should identify companies that meet these specific criteria and show the evidence behind each recommendation.
That gives the team something to check. They can examine suitable and unsuitable companies, include cases with incomplete information, and compare the system’s decisions with a carefully reviewed reference set. Disagreements in that reference set deserve attention too. Human judgments can be inconsistent.
The team also needs to decide which mistakes matter most. Recommending an unsuitable prospect creates extra work. Missing an excellent one creates an opportunity cost that may be much harder to observe.
For a business owner, that is where evaluation becomes useful: it connects the system’s behaviour to the reason it was introduced.
A good average can hide a bad experience
Imagine a hypothetical test of 1,000 prospect classifications. The system makes the correct decision in 920 cases. An overall accuracy of 92% sounds encouraging.

Now separate the results. Among 900 companies with detailed websites, it gets 882 right: 98%. Among 100 companies with sparse websites, it gets only 38 right: 38%.
Both results are contained inside that reassuring 92%.
If smaller companies often have less information online, this would give us a reason to investigate whether they are being overlooked. We would still need to examine the errors: did the system reject suitable companies, recommend unsuitable ones, or do both?
These numbers are illustrative, but the question they raise is practical. Which situations does your overall performance figure combine?
Language, customer type, document quality, and unusual requests may all deserve separate attention. The relevant categories depend on the work.
I would also ask who is missing from the test. A dataset containing only completed applications tells you little about people who could not finish the process.
Treat an audit as a question with evidence
This is where AI auditing becomes relevant. A credible audit systematically examines a system against defined claims or expectations. It should make clear what was tested, how it was tested, what happened, and what remains uncertain.
As a reader, I would look past the verdict and examine its boundaries. Which version was evaluated? Was it tested in a setting similar to ours? Did the evaluation include difficult cases? Could the auditors report uncomfortable findings? What did they lack access to?
A report can be useful within a limited scope. Problems arise when that limited finding becomes a much broader promise in a sales presentation.
For example, evidence that an AI system summarises documents accurately would not establish that its business recommendations are sound. The recommendation requires additional judgments about relevance, priorities, and consequences.
I find it helpful to organise an evaluation around five questions:
- What exactly should it do? Define the task, the context, and which errors would be unacceptable.
- How could we find out if it fails? Design tests that reflect real use, including difficult cases and incomplete information.
- What do the results actually support? Examine the failures, the missing data, and the limits of the conclusions.
- What will we do with what we learn? Choose an action and assign someone responsibility for following it through.
- When will we check again? Set a review date and identify changes that should trigger another evaluation.

These questions give a team a practical starting point. A formal audit may require additional methods and specialist expertise, particularly when the system affects people’s access to important opportunities or services.
Human review needs working conditions
“A person checks it” can sound reassuring. I would want to understand what that person is actually able to check.
Do they have the source material? Do they understand the subject? Is there enough time to investigate a doubtful conclusion? Can they reject the output, and is there a clear way to escalate a problem?
In the reporting example, a reviewer who can see only the finished summary may miss the incomplete source file. Giving that person a checklist will have limited value unless they can trace the recommendation back to the underlying information.
Review creates value when the reviewer can challenge the work. Its effectiveness should be tested too: which mistakes are caught, which reach the client, and how much effort does correction require?
That effort belongs in the economics of the system. If preparation becomes faster but verification becomes substantially harder, the business needs to measure the combined result.

Decide what happens when a test fails
Finding a problem should lead to a decision somebody owns.
Sometimes the next step is to improve the source data and repeat the test. Sometimes it is to keep the system on a narrower task, require review before an output is used, or pause a particular use while the problem is investigated.
Those choices should reflect the consequences of getting it wrong. There is little value in an escalation process that raises an alert but leaves everyone unsure who can stop the workflow.
I would agree on that responsibility before expanding the system’s role.
For the prospecting tool, that might mean allowing it to recommend companies while requiring a person to examine the ones it would exclude. The team could then test whether this arrangement catches the missed opportunities without creating more work than it saves.
A correction is another claim about the system. It needs evidence too.
Confidence needs maintenance
An evaluation describes the system under particular conditions. A new model, different source material, revised instructions, or a new customer group can change those conditions.
I would keep a record of the version tested, the important failures, the corrections made, and the cases that need checking again. Changes to the workflow should trigger a judgment about retesting. Routine monitoring can help identify problems between deeper evaluations.
For a small company, the starting point can be modest: one important workflow, a documented set of test cases, clear criteria, and someone responsible for follow-up. As the consequences grow, the depth and independence of the evaluation should grow with them.
What interests me about this is the possibility of expanding AI’s role with a clearer understanding of what we are relying on. A team that knows where a system performs well, where it fails, and how to respond has a better basis for deciding what to automate next.
Before giving AI another responsibility, I would ask one question:
What evidence would make us comfortable relying on it here—and who will notice when that evidence no longer holds?


Leave a Reply