Skip to content

Your AI got it wrong. Which step broke?

Open the replay for that turn. It shows what the agent reasoned, which tools it called, and how long each step took. Knowledge-base, stage, tool, and config-version problems stop looking alike.

Three things to check after a failure

The final reply alone leaves you guessing. These three make it reviewable.

  • The turn itself

    What it reasoned, which tools it called, how long each step took, replayed turn by turn.

  • The version it ran

    Config changes are saved as a version and then applied, so you can tell which one was live.

  • Verdicts, line by line

    Pull the conversation into the review workbench and score it against your own criteria, failures included.

Four common failures, four layers

One sentence, “it got it wrong”, can be four different things underneath. Each takes a different fix.

Knowledge base

What you see
It states what your material never said
Which layer
Knowledge base
In the replay
Whether the turn queried the knowledge base, and what it got
The fix
Fix the document and re-chunk it. One fix covers the whole team that uses it.

Stage and script

What you see
Facts right, but it went too far — over-promised, or skipped a required step
Which layer
Stage and script
In the replay
Which stage it reached, and that stage’s goal and instructions
The fix
Rewrite that stage’s instructions, and the condition for moving on.

Tools

What you see
A required tool never ran — no stock check, no slot held, no ticket
Which layer
Tools
In the replay
Which tools the turn called, and with what arguments
The fix
Change the available tools, and make “must call” and “must not call” hard rules.

Config version

What you see
Fine yesterday, wrong today
Which layer
Config version
In the replay
The config version it ran on
The fix
Roll back. Every change is saved as a new version before it goes live.

The model, or the knowledge base?

These two need opposite fixes: a rewrite, or a missing document. Work the three steps in order.

  1. 01

    Start with what it found

    Did the turn query the knowledge base, and what came back? If nothing did, the model is not your problem.

  2. 02

    Then check what is in there

    Knowledge is its own layer: documents, chunks and retrieval are each inspectable. If it was never there, no instruction will fix it.

  3. 03

    Instructions last

    It is in there, it was retrieved, and the answer is still off. That is a scripting problem.

Locating it is half the job. Confirming the fix held is the other half, and that is what how to verify before launch covers.

When this layer is not worth it

You only want a FAQ bot

  • If answering the question ends the job, and the questions barely vary, replay buys little. Cheaper options exist; here is how they line up.

Not every business needs to see this deep. Three honest cases.

  • You only need a record:If you only need what happened, a chat archive is enough. What the line should have said still comes from your top closer.
  • No one will act on it:Tracing pays off in the change that follows. If the playbook never moves, this is just logs no one opens.

If answering the question ends the job, and the questions barely vary, replay buys little. Cheaper options exist; here is how they line up.

Five questions to ask every vendor

An open notebook on a desk by the window of an empty office, ready for the five traceability questions to ask every AI sales vendor.
  • Can you open one real conversation and show the full process?

    If only the final reply is stored, failure still starts with guessing.

  • Did this turn query the knowledge base, and what came back?

    It separates a missing document from a bad answer.

  • Which config version ran here, what was the one before, and can we roll back?

    “It worked yesterday” is usually a version.

  • Who sets the criteria for a good answer, and can we see every verdict?

    Generic criteria flatten whatever made your team better.

  • Can this failed conversation become a test case the next release has to pass?

    If it cannot, the same mistake comes back.

Ask TOPPP too. The ones that go unanswered are where you will get stuck later. In insurance or aesthetics, where any sentence can be pulled back up months later, ask harder.

Three questions we get about replay

Can a non-technical manager read a replay?

Yes. TOPPP lays out the process of each turn: what the agent reasoned, which tools it called, how long each took. Reading it takes knowledge of your own sales playbook, not code.

Can you really tell a model problem from a knowledge-base problem?

Yes. TOPPP shows whether the turn queried the knowledge base and what came back. Nothing there means the fault is not the model; retrieved and still wrong means the instructions.

How do we stop it happening again?

TOPPP lets you pull that conversation into the review workbench as a test case, so every config change has to clear it before going live.

Talk to it live, with the process on screen

A demo takes about 30 minutes: you play a buyer in your industry and chat with it — the conversation on the left, what it is thinking and calling on the right.

  1. 01The turn itself
  2. 02The version it ran
  3. 03Verdicts, line by line

Or reach us directly business@toppp.ai

Book a demo

Leave your details and we’ll be in touch within one business day.

We only talk about your business. No mass mailing.