Skip to content
Hey Bubba!Blog

Testing a system that talks

Conversations have no single correct output, so assertions have to move from the transcript to the outcome.

Hey Bubba! Engineering1 min read

You cannot assert that a conversational agent said a particular sentence. There are a hundred acceptable ways to ask a broker whether the load is still available, and pinning one makes the suite break on every prompt change while catching nothing that matters.

Assert on outcomes#

The useful questions are about state, not words:

  • Did it end with a booking at or below the authorised rate?
  • Did it collect the appointment time before committing?
  • Did it avoid committing a truck without enough hours?

Those are checkable, stable across rewording, and the things that actually cost money when wrong.

def test_does_not_book_above_authority(run):
    result = run(scenario="broker_holds_firm", max_rate_cents=90_000)
    assert result.booking is None
    assert "declined" in result.outcome

Scripted counterparties#

The other side of the call has to be deterministic enough to test against. Scenario definitions describe behaviour, not dialogue:

ScenarioThe broker
broker_holds_firmNever moves on rate
load_already_coveredReveals this only after the appointment is discussed
late_second_stopAdds a stop in the final twenty seconds

That last one matters. A third of real calls surface something decisive at the end, so a suite where everything is known upfront tests a world that does not exist.

The grading problem#

Some properties need judgement: was the agent rude, did it misrepresent the truck. We score those with a model and treat the result as a signal, not a gate.

A graded check that fails the build on a borderline call teaches the team to rerun CI until it passes. Then it is worse than not having it.

Gates are for the deterministic assertions. Graded checks go on a dashboard and get looked at when the trend moves.

Keep reading