AI & LLM Testing Learning Path: Start Here
In short: you learn AI testing in 6 milestones by building one evaluation harness for a support bot and growing it — golden answers, attack prompts, structured-output checks, RAG and agent tests. It runs without an API key (the "model" is a stand-in), so you can swap in a real model whenever you like.
Read this page once to see the plan. Then, for each milestone, follow the same four steps: learn → build → check → commit. Come back here whenever you're unsure what to do next.
Before you start
- Python 3.10+ (for pytest); Node.js 20+ from Milestone 3 (for promptfoo).
- Basic Python: functions, dicts,
assert. - The lab —
evals/bot.pyand friends. - Optional, for real-model milestones: an API key for Anthropic or OpenAI.
- Helpful: the API testing cheat sheet if you'll call a model's API.
The pages in this learning path
| Page | What it's for | When to open it |
|---|---|---|
| This roadmap | The plan and self-checks | At the start of each milestone |
| AI/LLM testing cheat sheet | Learn each idea, with runnable code | The "learn" step |
| Milestones & Mini-Projects | Tasks, expected results, solutions | The "build" and "check" steps |
| Quick Reference | Assertions, OWASP LLM Top 10, metrics, attacks, tools | Any time |
| Best Practices | Habits that make evals trustworthy | After Milestone 2, then before shipping |
| AI product & LLM testing guide | The service and AI-augmented QA | For the deeper "why" |
The milestones
| Milestone | You learn | You build | Rough time |
|---|---|---|---|
| 1 | Why AI differs, golden datasets, the harness | Accuracy tests | 1 week |
| 2 | Hallucination, prompt injection, OWASP LLM Top 10 | An attack set | 1–2 weeks |
| 3 | Structured output, determinism, promptfoo | Config-based evals | 1 week |
| 4 | RAG: retrieval, groundedness, citations | RAG tests | 1–2 weeks |
| 5 | Agents, tool limits, LLM-as-judge | Agent + judge tests | 1–2 weeks |
| 6 | Regression, cost, an evaluation report | A CI eval and report | 1 week |
Times assume about 5 hours a week.
For each milestone: learn (cheat-sheet sections) → build (the tasks) → check (run the tests / solution) → commit.
Milestone 1: The harness
Learn: Why AI testing is different · What to test · Evaluation, not pass/fail · Golden datasets · Metrics · The harness
Build: Accuracy tests
Check yourself:
- Why can't you assert
output == expectedfor an LLM? - Why store facts instead of exact strings in the golden set?
- What does the
must_not_includelist protect against?
Milestone 2: Attacks & safety
Learn: Accuracy & hallucination · Prompt injection & the OWASP LLM Top 10 · Quick Reference: Attack prompts, OWASP LLM Top 10
Build: An attack set
Then read: Best Practices, sections 1–4.
Check yourself:
- What two things must you check for each attack prompt?
- What is indirect prompt injection?
- Why is injection resistance a 100% threshold, not 95%?
Milestone 3: Output & promptfoo
Learn: Structured output · Determinism & temperature · promptfoo · Quick Reference: promptfoo
Build: Config-based evals
Check yourself:
- How do you check JSON output properly?
- Why run each case several times with a real model?
- When would you choose promptfoo over pytest?
Milestone 4: RAG
Learn: RAG testing · Quick Reference: RAG
Build: RAG tests
Check yourself:
- Why test retrieval and the answer separately?
- What is groundedness?
- What should a RAG bot do when no document matches?
Milestone 5: Agents & judges
Learn: Agents & tool use · LLM-as-a-judge · Quick Reference: Agents, Judge rubric
Build: Agent + judge tests
Then read: Best Practices, sections 6–7.
Check yourself:
- Why enforce agent limits in code, not the prompt?
- How do you calibrate an LLM judge?
- When must a human review, not a judge?
Milestone 6: Regression, cost & report
Learn: Regression after changes · Cost & latency · Bias & safety · AI-augmented QA
Build: A CI eval and report
Then read: Best Practices, sections 5, 8–11.
Check yourself:
- Why pin the model version?
- How can a new model improve overall but still be a regression?
- What belongs in an evaluation report?
When you get stuck
| Problem | What to do |
|---|---|
| Tests pass but feel too easy | The stand-in bot is deterministic and careful by design — swap in a real model to see failures |
| A real model's answers vary | Test facts/meaning, not exact text; run each case several times |
| No API key | Everything here runs on the stand-in; add a key only for the real-model tasks |
| promptfoo can't find the provider | Use python:provider.py with the path relative to the config file |
| Judge disagrees with you | Tighten the rubric; calibrate on more human-graded samples |
What's next
- Security of AI features: security cheat sheet and OWASP LLM Top 10.
- Load-test AI endpoints: performance cheat sheet.
- Build with models: the Claude cheat sheet and MCP & AI agents guide.
- Certifications: ISTQB CT-AI (AI Testing) and CT-GenAI (Testing with Generative AI).
Good resources: OWASP Top 10 for LLM Applications, the promptfoo docs, and Anthropic's and OpenAI's own guides on evaluating models.