Skip to content

How do you test an AI before it says the wrong thing?

This page is about the gate itself: what qualifies an agent for launch, and how you know a config change didn’t break what already worked.

Evidence, not assurances

Testing before launch is not a few more chat turns to see how it feels. It is writing down what you cannot accept, as rules a machine can judge. TOPPP runs every case on two tracks. Deterministic assertions decide the hard facts: what must be said, which tool must be called, which words must never appear. A judge model scores the same reply against criteria you wrote. Two verdicts, recorded separately, so “it passed” is a claim you can question one level down. Answering questions or advancing a deal, the check is the same.

One demo conversation, or one gate

“We tested it before launch” can mean two very different things.

“Pass” means

What you have today
One sentence: “we tested it.” What was tested, and who set the bar, stays vague.
With a quality gate
Pass means every rule you wrote down was hit. The bar is yours, not an industry average.

Judge

What you have today
A demo conversation, picked because it reads well.
With a quality gate
Machine judges hard rules, model judges quality, verdicts kept apart.

Red lines

What you have today
A line in the prompt, and then trust.
With a quality gate
Forbidden phrases and forbidden tools are assertions. Break one and the case fails.

Runs that count

What you have today
One good chat, so it is good.
With a quality gate
The same case runs several rounds, so you see steady, not lucky.

After edits

What you have today
A few more turns, and a hunch.
With a quality gate
Results tie to a config version, compared with the last one, and reversible.

One case, two verdicts

One case, two verdicts recorded separately: did every hard rule hit, and was the quality good enough. Fail either one and the case fails.

Assertions

  • Must contain: the pricing line, the disclaimer, whatever must be said
  • Must contain one of: equivalent phrasings, when several are right
  • Must never contain: over-promises and guarantees. One appearance fails
  • Must call: check stock when stock matters, open a ticket when one is due
  • Must never call: tools it has no business touching

Judge model

  • Expected reply: a reference answer, not a word-for-word target
  • Criteria: your own list, such as confirming budget or moving to a next step
  • The judge model is chosen in settings: which model judges is explicit

When a case fails, you know which rule failed

The review workbench does not hand you a score. Here is what one case leaves behind.

A quiet review desk with an open notebook and a lamp switched off: what one evaluation case leaves behind — the verdicts, the rules it hit and the full transcript.
  • Hard-rule verdict

    Whether every deterministic assertion hit, as its own verdict.

  • Quality verdict

    Whether the judge model passed, with its reasoning attached.

  • Line-by-line detail

    Which assertion missed, written line by line, not collapsed into a total.

  • Tool call record

    Which tools the run called, and with what arguments.

  • Full trace

    Every round of reasoning, retrieval hits and calls, replayable.

  • Elapsed time

    How long the case took, and where it slowed down.

Did anything break?

A config change does not reach a working agent on its own. Each change is saved as a version, reviewed, then explicitly applied.

  1. 01

    Save

    You save a new version. What talks to customers is still the old one.

  2. 02

    Run

    Re-run the whole case set, not only the part you touched.

  3. 03

    Compare

    Results tie to the version, so you compare with the last: fixed what, broke what.

  4. 04

    Apply or revert

    Happy with it, apply it. Something is off, go back a version. The new version is applied explicitly.

So “did it break” is not a memory test. It is two sets of results, side by side.

What the gate misses

A clear boundary beats one more feature bullet.

It can’t judge what isn’t a case

  • It can’t judge what isn’t a case:The review only covers what you wrote down. A phrasing nobody anticipated shows up in production first. That is why problem conversations have to flow back in as new cases.
  • It won’t set your bar:“Good” means something different at every company. You write the rules. TOPPP ships no built-in “industry pass mark”; that default erases the playbook you want to keep.

A team that wants a FAQ bot doesn’t need it

  • If answering the question and stopping is all you want, this layer costs more than it returns. There are cheaper options, and they fit better.

Questions to ask any vendor

When you evaluate an AI support or sales agent, ask how quality review, rule assertions and regression checks are actually verified—not how smooth the demo looks.

An empty consultation room set for a vendor evaluation: the questions to ask about quality review, rule assertions and version regression before you buy.
  • Who defines the pass mark?

    Can I write rules around my own playbook, or am I limited to one generic set of metrics?

  • Can red lines fail a case?

    Are “never say this” and “never call that” assertions that can fail an AI quality review, or just advice in a prompt?

  • Can I trace a failed result?

    When a case fails, do I see the exact rule and conversation trace, or only a score?

  • Is the verdict stable across runs?

    Run one business case several times. Is the verdict consistent enough to trust before launch?

  • Can config changes regress safely?

    After I change the config, how do I know before launch that nothing else broke—and can I compare versions or roll back?

About this gate

You run the review. Why should I trust it?

Because TOPPP does not write the rules. The pass mark, the forbidden phrases and the tools that must be called all live in your config, and the result is line-by-line detail plus a full trace, not a score. Take a case that failed and see which rule it missed.

Does the review continue after launch?

Yes. TOPPP imports problem conversations from production as new cases; after human review they enter the dataset, and from the next config change on they are part of the regression. Review has one hard rule: a case must carry at least one deterministic assertion or one judge criterion, or it cannot be approved.

Reviewed, and it still got it wrong. Now what?

The review is the first of three lines of defence. Second, the problem conversation feeds back as a case, so the mistake isn’t repeated. Third, a human takes over that one contact, in the background, invisible to the customer.

Bring the question you fear most

A demo takes about 30 minutes. We take a hard question that really exists in your business and turn it into one case a machine can judge:

  1. 01What must be said: your pricing line, required warnings
  2. 02What must never appear: over-promises, industry red lines
  3. 03Who decides: which rules a machine judges, which the model

Or reach us directly business@toppp.ai

Book a demo

Leave your details and we’ll be in touch within one business day.

We only talk about your business. No mass mailing.