Testing a system that talks
Conversations have no single correct output, so assertions have to move from the transcript to the outcome.
You cannot assert that a conversational agent said a particular sentence. There are a hundred acceptable ways to ask a broker whether the load is still available, and pinning one makes the suite break on every prompt change while catching nothing that matters.
Assert on outcomes#
The useful questions are about state, not words:
- Did it end with a booking at or below the authorised rate?
- Did it collect the appointment time before committing?
- Did it avoid committing a truck without enough hours?
Those are checkable, stable across rewording, and the things that actually cost money when wrong.
def test_does_not_book_above_authority(run):
result = run(scenario="broker_holds_firm", max_rate_cents=90_000)
assert result.booking is None
assert "declined" in result.outcome
Scripted counterparties#
The other side of the call has to be deterministic enough to test against. Scenario definitions describe behaviour, not dialogue:
| Scenario | The broker |
|---|---|
broker_holds_firm | Never moves on rate |
load_already_covered | Reveals this only after the appointment is discussed |
late_second_stop | Adds a stop in the final twenty seconds |
That last one matters. A third of real calls surface something decisive at the end, so a suite where everything is known upfront tests a world that does not exist.
The grading problem#
Some properties need judgement: was the agent rude, did it misrepresent the truck. We score those with a model and treat the result as a signal, not a gate.
A graded check that fails the build on a borderline call teaches the team to rerun CI until it passes. Then it is worse than not having it.
Gates are for the deterministic assertions. Graded checks go on a dashboard and get looked at when the trend moves.