Skip to main content

AI & LLM Testing Learning Path: Start Here

In short: you learn AI testing in 6 milestones by building one evaluation harness for a support bot and growing it — golden answers, attack prompts, structured-output checks, RAG and agent tests. It runs without an API key (the "model" is a stand-in), so you can swap in a real model whenever you like.

How to use this page

Read this page once to see the plan. Then, for each milestone, follow the same four steps: learn → build → check → commit. Come back here whenever you're unsure what to do next.

Before you start​

  • Python 3.10+ (for pytest); Node.js 20+ from Milestone 3 (for promptfoo).
  • Basic Python: functions, dicts, assert.
  • The lab — evals/bot.py and friends.
  • Optional, for real-model milestones: an API key for Anthropic or OpenAI.
  • Helpful: the API testing cheat sheet if you'll call a model's API.

The pages in this learning path​

PageWhat it's forWhen to open it
This roadmapThe plan and self-checksAt the start of each milestone
AI/LLM testing cheat sheetLearn each idea, with runnable codeThe "learn" step
Milestones & Mini-ProjectsTasks, expected results, solutionsThe "build" and "check" steps
Quick ReferenceAssertions, OWASP LLM Top 10, metrics, attacks, toolsAny time
Best PracticesHabits that make evals trustworthyAfter Milestone 2, then before shipping
AI product & LLM testing guideThe service and AI-augmented QAFor the deeper "why"

The milestones​

MilestoneYou learnYou buildRough time
1Why AI differs, golden datasets, the harnessAccuracy tests1 week
2Hallucination, prompt injection, OWASP LLM Top 10An attack set1–2 weeks
3Structured output, determinism, promptfooConfig-based evals1 week
4RAG: retrieval, groundedness, citationsRAG tests1–2 weeks
5Agents, tool limits, LLM-as-judgeAgent + judge tests1–2 weeks
6Regression, cost, an evaluation reportA CI eval and report1 week

Times assume about 5 hours a week.

For each milestone: learn (cheat-sheet sections) → build (the tasks) → check (run the tests / solution) → commit.

Milestone 1: The harness​

Learn: Why AI testing is different · What to test · Evaluation, not pass/fail · Golden datasets · Metrics · The harness

Build: Accuracy tests

Check yourself:

  • Why can't you assert output == expected for an LLM?
  • Why store facts instead of exact strings in the golden set?
  • What does the must_not_include list protect against?

Milestone 2: Attacks & safety​

Learn: Accuracy & hallucination · Prompt injection & the OWASP LLM Top 10 · Quick Reference: Attack prompts, OWASP LLM Top 10

Build: An attack set

Then read: Best Practices, sections 1–4.

Check yourself:

  • What two things must you check for each attack prompt?
  • What is indirect prompt injection?
  • Why is injection resistance a 100% threshold, not 95%?

Milestone 3: Output & promptfoo​

Learn: Structured output · Determinism & temperature · promptfoo · Quick Reference: promptfoo

Build: Config-based evals

Check yourself:

  • How do you check JSON output properly?
  • Why run each case several times with a real model?
  • When would you choose promptfoo over pytest?

Milestone 4: RAG​

Learn: RAG testing · Quick Reference: RAG

Build: RAG tests

Check yourself:

  • Why test retrieval and the answer separately?
  • What is groundedness?
  • What should a RAG bot do when no document matches?

Milestone 5: Agents & judges​

Learn: Agents & tool use · LLM-as-a-judge · Quick Reference: Agents, Judge rubric

Build: Agent + judge tests

Then read: Best Practices, sections 6–7.

Check yourself:

  • Why enforce agent limits in code, not the prompt?
  • How do you calibrate an LLM judge?
  • When must a human review, not a judge?

Milestone 6: Regression, cost & report​

Learn: Regression after changes · Cost & latency · Bias & safety · AI-augmented QA

Build: A CI eval and report

Then read: Best Practices, sections 5, 8–11.

Check yourself:

  • Why pin the model version?
  • How can a new model improve overall but still be a regression?
  • What belongs in an evaluation report?

When you get stuck​

ProblemWhat to do
Tests pass but feel too easyThe stand-in bot is deterministic and careful by design — swap in a real model to see failures
A real model's answers varyTest facts/meaning, not exact text; run each case several times
No API keyEverything here runs on the stand-in; add a key only for the real-model tasks
promptfoo can't find the providerUse python:provider.py with the path relative to the config file
Judge disagrees with youTighten the rubric; calibrate on more human-graded samples

What's next​

Good resources: OWASP Top 10 for LLM Applications, the promptfoo docs, and Anthropic's and OpenAI's own guides on evaluating models.