AI & LLM Testing Best Practices
Habits that make AI feature testing trustworthy — evaluate not assert, build golden and attack sets with experts, test meaning not text, pin versions and run regression, calibrate judges, enforce agent limits in code, watch cost and data handling, keep humans accountable, and a pre-release checklist.
AI & LLM Testing Learning Path: Start Here
How to learn AI and LLM testing in six milestones — build an evaluation harness, test accuracy and hallucination, prompt injection and safety, structured output and promptfoo, RAG and agents, and a full evaluation report — runnable offline.
AI & LLM Testing Milestones & Mini-Projects
Tasks with expected results for each of the six AI/LLM testing milestones — the harness and golden set, attacks and injection, structured output and promptfoo, RAG, agents and judges, and a CI eval with a report — runnable offline.
AI & LLM Testing Quick Reference
Copy-paste reference for testing AI and LLM features — what to test, assertion types, pytest and promptfoo snippets, OWASP LLM Top 10, RAG and agent checklists, metrics and thresholds, attack prompts, judge rubrics and tools.