AI Engineering 5 min read

The Test Every AI Agent Has to Pass Before It Talks to a Customer

Everyone ships AI agents. Almost nobody will say how they know one works. Here is the eval process we run: the fixed test set, the independent scorer, the agreed pass bar.

Sophie Costello AI & Digital Strategist
6:28
1x

Who scores it

Not the person who built it. This is the single cheapest quality control available and almost nobody does it, because it feels like distrust. It is not. The builder knows what they intended, so they read the intention into the output. A second person reads what is actually there.

For volume we grade with a model as a first pass and have a human check every disagreement and a random sample of the agreements. Anthropic's guidance on developing tests and the OpenAI evals repo are both reasonable starting points if you want to build this yourself rather than buy it.

The case every test set should carry

In November 2022 a man named Jake Moffatt asked Air Canada's website chatbot about bereavement fares. The bot told him he could buy a ticket at full price and claim the discount back within 90 days. That was wrong. The airline's actual policy required the request before purchase, not after.

He bought the ticket. Air Canada refused the refund. In February 2024 the British Columbia Civil Resolution Tribunal found for him and ordered the airline to pay $812.02.

The part worth sitting with is the defence. Air Canada argued the chatbot was a separate legal entity, responsible for its own words. Tribunal member Christopher Rivers called that submission remarkable, and found the airline had not taken reasonable care to keep its chatbot accurate.

That is the argument for evals in one sentence. Not that the model was wrong, because models are wrong sometimes. That nobody had checked, and "the AI said it" is not a defence anyone accepts.

Run it against the four failure modes above and it is a textbook case of the first. The answer was fluent, well structured, specific about a number of days, and false. It would have survived a skim review. A test set of thirty real bereavement enquiries with their correct answers would have caught it in an afternoon, for a fraction of what the ruling cost, and orders of magnitude less than the coverage cost.

Two more for the must-refuse pile

A Chevrolet dealership's assistant was talked into agreeing to sell a Tahoe for one dollar, and confirmed it as "a legally binding offer, no takesies backsies". The screenshot passed twenty million views. DPD switched off part of its chatbot after a customer prompted it to swear at him and write poems criticising the company.

Do you know what your agent's pass bar was?

Most agents went live on one good run in front of the team.

Neither is a model failure in the interesting sense. Both are cases nobody wrote a must-refuse test for. That is the third failure mode, and it is the one that ends up in a screenshot rather than a ledger. What happens when an agent skips this kind of testing is the same story with your CRM in the middle of it.

The lesson across all three: a test set is not a gate you pass once. It is a ledger of every mistake anyone in your category has already made, so you cannot make it again. Ours carries Air Canada. Yours should too, and it costs you nothing to add it today.

Where this sits with the Australian guidance

Since 1 May 2026 the ASD’s Australian Cyber Security Centre has had published guidance on careful adoption of agentic AI services, which is the current Australian government word on this exact question: an agent acting on its own, near real customers and real data. Read it and you will notice the process above is what it asks for, in plainer language. Know what the agent is allowed to do. Test it against cases you wrote down in advance. Have a person who is accountable for the output. Keep a record of what you tested, so a decision can be reviewed later by someone who was not in the room.

We are not claiming the guidance blesses our method. It is the other way around. If you already run a fixed test set, an independent scorer and an agreed pass bar, you can answer a regulator, an insurer or a board with evidence instead of a demo. If you do not, the guidance is a reasonable place to find the words for why you should. Worth deciding whether you need an agent at all first, because a workflow with three branches needs none of this and does the job more often than vendors admit.

What this costs

Building a first test set is roughly two days of work for a competent person who knows the business. Running it is minutes. Re-running it on every change is the whole reason it is worth building, because the second time you ship a prompt tweak that breaks something nobody was watching, you will find out in five minutes instead of five weeks.

If that sounds like a lot of process for a chatbot, consider what you already do before a new hire answers a customer unsupervised. You would not skip that. An agent is a system with your name on it, deployed at a scale a person cannot reach, and the NIST AI Risk Management Framework is worth ten minutes if you need the formal language for a board.

The one thing to take away

Ask any vendor pitching you an agent for their test set and their pass bar. Not their demo. If they cannot produce both in writing, they have not tested it, they have tried it. Those are different words for a reason.

That question applies to platforms too, HubSpot included. What HubSpot’s own agent tooling can and can’t do is worth knowing before you scope anything on top of it, because the test set is still yours to write either way.

Need help with an agent?

We're Neighbourhood. We build the AI and the revenue system it runs on. AI and RevOps engineering for Australian teams. Diamond HubSpot Partner, 17 HubSpot Impact Awards. How we build and test AI engineering work like this is on our service page, test sets included.

Would your agent pass a test set you did not write?

The person who built it is usually the one who signs it off.

Give us a shout and tell us what's broken.

Neighbourhood

Neighbourhood is a HubSpot Diamond Partner in Brisbane. We build AI systems and the revenue operations they run on, for businesses across Australia and New Zealand.