Skip to main content

AI & LLM Testing Best Practices

This page lists good habits for testing AI and LLM features. They help you measure real quality (not a single lucky answer), catch the failures that matter โ€” wrong facts, leaks, unsafe actions โ€” and keep confidence as models and prompts change. They apply to SDETs, ML engineers and product teams.

Each practice has:

  • Do โ€“ the good way.
  • Why โ€“ the reason in simple words.
How to use this page

Read it once after Part 2 of the AI/LLM testing cheat sheet, then use the checklist at the end before shipping an AI feature.


Contentsโ€‹

  1. Evaluate, Don't Assert Once
  2. Build the Dataset With Experts
  3. Test Meaning, Not Exact Text
  4. Make Injection Testing Standard
  5. Pin Versions and Run Regression
  6. Calibrate Your Judges
  7. Enforce Agent Limits in Code
  8. Watch Cost, Latency and Data
  9. Keep Humans Accountable
  10. Test the Non-AI Parts Too
  11. Checklist Before Release

1. Evaluate, Don't Assert Onceโ€‹

In short: one green run proves nothing about a system that varies.

  • Do run a dataset of cases and report rates per category, with thresholds. Why: "94% of golden answers correct, 100% injection resistance" is a quality signal; a single passing test isn't.
  • Do run each case several times. Why: the same prompt can pass once and fail the next time.
  • Do set a threshold per category, not one overall number. Why: injection must be 100%; a 90% "overall" score can hide total failure on safety.

2. Build the Dataset With Expertsโ€‹

In short: your eval is only as good as its cases.

  • Do write golden cases with people who know the domain (support, legal, medical). Why: only they know the right answer and the dangerous wrong ones.
  • Do store must-include and must-not-include facts, not just questions. Why: the must-not list is your hallucination guard.
  • Do grow the set from real usage: every complaint and odd answer becomes a case. Why: the dataset comes to represent what users actually ask.
  • Do cover edge cases: empty, very long, other languages, ambiguous, adversarial. Why: that's where models fail.

3. Test Meaning, Not Exact Textโ€‹

In short: correct answers can be worded many ways; wrong ones can be worded like right ones.

  • Do check for required facts, or compare by meaning (embeddings / a judge). Why: == "expected string" fails on a correct paraphrase and passes on a confident lie with the right words.
  • Do use must_not_include for the specific wrong facts you fear. Why: it catches hallucinations precisely.
  • Do validate structured output by parsing and schema, not string match. Why: models add stray prose around JSON; parse it.

4. Make Injection Testing Standardโ€‹

In short: every AI feature that takes user input (or reads documents) needs an attack set.

  • Do keep a red-team set of injection and jailbreak prompts; run it every build. Why: prompt injection is the #1 LLM risk and it regresses easily.
  • Do test indirect injection: hide instructions in a document/email/page the AI reads. Why: it's the attack teams forget, and RAG systems are exposed to it.
  • Do check two things per attack: the secret/instructions don't leak, and the model doesn't act on the injection. Why: a refusal that still performs the injected action is a fail.
  • Do treat model output as untrusted when it's rendered or executed (XSS, SQL). Why: LLM05 โ€” output handling is where injection becomes a normal web vuln.

5. Pin Versions and Run Regressionโ€‹

In short: AI behaviour changes when the model, prompt or data changes โ€” even with no code change.

  • Do pin the exact model version, not "latest". Why: a silent model update can change every answer.
  • Do re-run the full eval set on every model, prompt or knowledge-base change. Why: that's the only thing that changed, and it changed behaviour.
  • Do compare per-category scores against the last release. Why: a new model can improve overall while getting worse where it matters.

6. Calibrate Your Judgesโ€‹

In short: an LLM judge is useful but not automatically right.

  • Do give the judge a clear rubric, a small scale, and ask for a reason. Why: vague rubrics give inconsistent scores.
  • Do check the judge against human grades on a sample before trusting it. Why: an uncalibrated judge can be confidently wrong, like any model.
  • Do use a different or stronger model as the judge where possible. Why: a model judging itself is biased toward its own style.
  • Don't let a judge be the only gate on high-stakes answers. Why: medical/legal/financial answers need human review.

7. Enforce Agent Limits in Codeโ€‹

In short: a prompt can be talked around; code can't.

  • Do enforce hard limits in the tool implementation (refund โ‰ค order total, allowed actions). Why: "never refund more than the total" in the prompt is a suggestion, not a guarantee (LLM06).
  • Do require confirmation for destructive actions. Why: an agent shouldn't delete or email on a misread instruction.
  • Do give agents least privilege โ€” only the tools the task needs. Why: limits the damage of any single mistake or injection.
  • Do test tool-error handling and termination. Why: agents loop or crash when a tool fails unexpectedly.

8. Watch Cost, Latency and Dataโ€‹

In short: "works" for an AI feature includes fast enough, cheap enough, and private.

  • Do measure latency (p95/p99) and cost per request in tests. Why: a correct answer that takes 30 s or costs a fortune isn't shippable.
  • Do test caps and timeouts against huge prompts and loops (LLM10). Why: unbounded consumption is a real denial-of-wallet risk.
  • Do agree what data may be sent to which model/eval tool; mask personal data. Why: prompts and eval runs can leak customer data to third parties.

9. Keep Humans Accountableโ€‹

In short: AI assists; people decide.

  • Do have a named human sign off tests, releases and any AI-drafted work. Why: accountability can't be delegated to a model.
  • Do review AI-written test code and cases like any pull request. Why: AI produces plausible-but-wrong tests too.
  • Do measure whether AI assistance actually helps (time saved, escaped bugs). Why: keep what works, drop what doesn't โ€” don't adopt on faith.

10. Test the Non-AI Parts Tooโ€‹

In short: an AI feature is still software.

  • Do test the surrounding functional, integration, security and accessibility behaviour. Why: the bug is often in the plumbing, not the model.
  • Do test fallbacks: model down, slow, or returns garbage. Why: the feature must degrade gracefully, not crash.

11. Checklist Before Releaseโ€‹

  • Golden dataset (built with experts) passes its accuracy threshold
  • Hallucination guards (must-not-include) in place and passing
  • Injection/red-team set at 100%, including indirect injection
  • Structured output parsed and schema-validated
  • Each case run multiple times; results stable enough
  • Model version pinned; full eval re-run after the latest change
  • RAG (if any): retrieval, groundedness, citations, "unknown" handling, permissions
  • Agents (if any): tool correctness, hard limits in code, confirmation, termination
  • Latency and cost within budget; caps/timeouts tested
  • Data-handling agreed; personal data masked in prompts and eval runs
  • Judges calibrated against humans; high-stakes answers human-reviewed
  • Fallback behaviour tested; a human signed off

Need more detail? Cheat sheet ยท Quick reference ยท Full guide