Jazyk / Language

EN
RegisterGet first meetings
Back to Leeeds
AI sales agent16 August 2026

The question is not whether an AI will ever get it wrong. It is whether you find out before the customer does.

Saved questions are replayed after every change to your material and show which answers moved. And for every sent reply it is traceable what it was built from, whether a human rewrote it, and who released it.

The AI agent test suite with the results of the last run
Replayed after a changeCompares figures, not wordingSends nothingExports to a spreadsheet

The short verdict

You upload a new price list. From that moment the agent answers differently - and you do not know to which questions until somebody complains. That is the least comfortable property of language models: they change quietly.

So Leeeds can store the questions you get most often and replay them after every change to your material. The result does not just say pass or fail; it says these three answers moved since last time.

Alongside it sits the audit export: one row per sent reply, with what it answered, from which sources, whether a human rewrote it before sending, and who approved it.

Why reading the replies is not enough

Reading every reply works for a week. At twenty conversations a day nobody does it, the check quietly switches off - and that is exactly when the price list changes.

Comparing texts is not enough either. The model rewords a sentence on every run, so a character-by-character comparison would report a change every time. Nobody opens that report after the second week.

So the substance is compared instead: the figures in the answer and the sources it leans on. “The price is 1,490 EUR” and “Our price comes to 1,490 EUR” are the same. 1,990, or the same price suddenly from a different document, is a change.

What a run looks like

Material changed → the suite queues itself → questions are replayed on the dry path → compared with the previous run.

A scenario is a question and an expectation

You save a question you actually get, and optionally what you expect back: that it cites a particular source and contains a particular text. A scenario is tested against one specific agent, because that agent decides what it may draw on.

It starts itself after a change

After material is uploaded or a fact edited the suite queues itself, with a five-minute settle and a pause between runs - so editing ten prices means one run, not ten. It can be started by hand at any time.

Nothing is sent

Every scenario runs on the same dry path as the test in the agent dialog. No email leaves, no slot is held, no credit is spent. A run is safe in the middle of live traffic.

The result says what moved

Next to the number of scenarios that passed sits the number that matters: how many answers changed since the last run. For each failure you see why - missing text, wrong source, a risk flag.

The export traces a single reply

The audit export downloads as a spreadsheet for a chosen period: what went out, what it answered, from which parts of which sources, whether a human rewrote it before sending, who approved it and with what risks.

What it is for

Before a price change. Run the suite, change the list, run again. The difference is a list of answers that moved - five minutes of reading instead of two hundred conversations.

When a new agent goes live. Twenty questions you get most often are the cheapest way to find out whether the base is ready, without testing it on customers.

For larger and regulated clients. The audit export answers “how did you arrive at that”, which a legal department asks sooner or later. The “was edited” column stands on its own, because it is the first thing an auditor looks for.

In a dispute. When the agent gets something wrong for the first time, the difference between a company that knows from what and by whom, and one that does not, is decisive.

What a suite typically catches

A new price list changed an answer you did not expect - a lead-time question that mentioned price in passing.
An expired fact: an answer that carried a figure for a year suddenly has none.
A replaced document has a different name and the answer now leans on a different source, even though it reads the same.
A scenario that stopped passing after facts were narrowed to a specific agent - a setting nobody remembered.
An answer that did not change even after you edited the material, because it is written in a way the agent cannot find.

Why it is built this way

Silence is not proof

01

A change announces itself

Models change quietly. The suite is the one place that change is visible before a customer sees it.

No noise

02

Figures and sources, not wording

A reworded sentence is not reported as a change; otherwise nobody would open the report.

A ticket upmarket

03

An export to a spreadsheet

Exactly what larger and regulated clients want to see before they sign.

What to take from this

An AI you can trust is not one that never errs. It is one where you can tell when, where and why - and have it in writing.

Save the twenty questions you get most often. It is half an hour of work, and from then on you know about every change.

Want to know what a change to your material did to the answers?

Save your questions and Leeeds replays them after every change.

I want the AI agent

Leeeds · acquisition automation with an AI sales agent