Insight Test automation

Deep dive: Evaluating AI

Output Against Acceptable Ranges

Deep Dive: Evaluating AI

Refleqt

October 2, 2026 • 7 min leestijd

Deel deze insight

The most fundamental technique

n the last post, I ran through six techniques for testing AI-generated output, and one of them keeps coming up in every conversation I have with QA leads and IT managers: evaluation against acceptable ranges, instead of exact matches.

It's the most foundational technique of the bunch. If your team only adopts one new habit when it starts testing AI, this should be it. So let's slow down and actually walk through what it looks like in practice.

Why "the old way" breaks down

Picture a normal software test: you call a function, you know exactly what it should return, and you compare the two. If they match, green check. If not, red X. This works because traditional software is deterministic: same input, same output.

AI-generated output doesn't play by that rule. Ask a model to reply to "Can I get a refund?" three times, and you might get three differently worded, equally correct answers. If your test is looking for one exact string, it will fail the second and third response even though there's nothing actually wrong with them.

Here's the shift in a nutshell:

diagram-1-comparison

Verschuiving binnen testing

Notice what changed on the right side. We stopped asking "does this match one specific answer?" and started asking "does this fall inside a set of acceptable answers?" That's the whole idea behind rubric-based, range-based evaluation.

So what does "acceptable" actually mean?

This is the part that trips people up, understandably, because "acceptable" sounds vague. However, it doesn't have to be. In practice, you translate it into a small set of concrete, checkable criteria, each with its own bar to clear. Something like:

  • Accuracy: Is the information factually correct and consistent with your source of truth?
  • Completeness: Does it actually answer what was asked, without leaving out anything essential?
  • Tone of voice: Does it match the voice you want (empathetic, formal, concise, on-brand)?
  • Policy compliance: Does it stay within legal, regulatory, or company guardrails? (Refund promises, medical claims, and financial advice are classic danger zones here.)
  • Format and length: Does it fit the constraints of where it'll be shown? A chat bubble, an email, a one-line summary, etc.
None of these have one "correct" phrasing. But every single one of them is measurable. That's the trick: you're not giving up on rigor, you're just moving the rigor from "matching text" to "meeting criteria."

What this looks like when you actually run it

Here's what that could look like inside an evaluation tool, for a batch of test cases run against a customer-support bot. One case has been flagged and expanded to show its rubric breakdown:

fabricated-tool-screenshot

AI testing in de realiteit

A few things worth pointing out here, because they're the details that make this genuinely useful rather than just a nicer-looking checklist:

Each criterion has its own pass bar. "Policy compliance" is held to a stricter standard (must score a perfect 5) than "tone" (a 3 or above is fine). Not everything deserves the same level of strictness, and pretending otherwise is how rubrics become either too rigid or too soft.

Criteria are weighted. Factual accuracy and policy compliance matter more here than response length, so they carry more weight in the final number. This weighting should reflect your actual business risk, not be split evenly out of convenience.

A single flagged criterion doesn't necessarily fail the whole response. In the example, the response overshoots the length guideline, but everything else is strong. The composite score is still comfortably above the release threshold. This is important: it lets a response move forward while still surfacing the issue for a human to glance at, rather than forcing an all-or-nothing gate on every imperfection.

There's still a bottom line. All of this flexibility doesn't mean "anything goes." You still end up with a composite score and a clear threshold for what's shippable. The nuance is in how you get there, not in whether there's a decision at the end.

Building your own rubric

If you're setting one of these up for the first time, it helps to work through it in a specific order rather than listing criteria off the top of your head:

diagram-3-steps

De juiste stappen

Start from the outcome you actually care about ("the customer walks away with the right answer and feels heard"), not from a generic checklist copied from somewhere else. Then decompose that outcome into criteria you can actually measure, set a bar for each one, and finally decide how much each one should count toward the final call. Skipping straight to "let's list some criteria" is the most common way these rubrics end up shallow or disconnected from what the business actually cares about.

How this gets automated

None of this is useful if a person has to read every single output and score it by hand. That defeats the purpose the moment your AI system is handling more than a handful of interactions a day. In practice, teams automate this rubric-based scoring in a few overlapping ways:

  • Rule-based checks: for the objective stuff: length limits, required disclaimers, forbidden words or phrases, formatting rules. These are cheap, fast, and deterministic, even though what they're checking isn't.
  • Semantic similarity scoring for the fuzzier criteria: comparing the AI's output to a reference answer not word-for-word, but in meaning, using embedding-based comparisons.
  • A second AI model acting as the evaluator, applying the rubric the same way a human reviewer would, at a scale no human team could match. This is often called "LLM-as-a-judge," and it's powerful enough that it deserves its own deep dive. That's coming next in this series.
Most mature setups combine all three: cheap rule-based checks catch the obvious problems first, and the more expensive semantic or AI-judged checks handle the nuance.

Best practices before you roll this out

A short list worth keeping close if you're introducing this at your organization:

  • Write the rubric with the people who understand the risk, not just the testers. Legal, compliance, and product teams often see failure modes that QA alone would miss.
  • Keep criteria few and meaningful. Five focused criteria you actually monitor beat fifteen you'll never look at again.
  • Revisit thresholds regularly. What counted as "good enough" at launch might not hold once the system is handling more edge cases or higher-stakes conversations.
  • Log everything, not just the failures. Passing responses that sit right at the edge of a threshold are often your earliest warning sign of drift.
  • Treat a flag as information, not automatically a blocker. The goal is visibility, not paralysis.

 

Where this leaves us

Range-based, rubric-driven evaluation is the foundation everything else in AI testing gets built on top of. It's what makes it possible to say something meaningful about an AI system's quality, even when that system is allowed to say things differently every time you ask.

It's also, frankly, the technique with the best ratio of effort to payoff. You don't need exotic infrastructure to start — a clear rubric, a few thresholds, and the discipline to apply them consistently will get most teams most of the way there.

Next up in this series: how a second AI model can act as the judge scoring these rubrics at scale, what "LLM-as-a-judge" actually looks like under the hood, and where it can go wrong.

Deel deze insight

milan

Milan Meuleman

Business development & sales

Contact Refleqt today

Would you like more control over software quality, test automation, or performance? We are happy to explore together how we can support your team with an approach that works in practice.