Skip to main content

4 docs tagged with "ai-testing"

View all tags

AI & LLM Testing Best Practices

Habits that make AI feature testing trustworthy — evaluate not assert, build golden and attack sets with experts, test meaning not text, pin versions and run regression, calibrate judges, enforce agent limits in code, watch cost and data handling, keep humans accountable, and a pre-release checklist.

AI & LLM Testing Learning Path: Start Here

How to learn AI and LLM testing in six milestones — build an evaluation harness, test accuracy and hallucination, prompt injection and safety, structured output and promptfoo, RAG and agents, and a full evaluation report — runnable offline.

AI & LLM Testing Milestones & Mini-Projects

Tasks with expected results for each of the six AI/LLM testing milestones — the harness and golden set, attacks and injection, structured output and promptfoo, RAG, agents and judges, and a CI eval with a report — runnable offline.

AI & LLM Testing Quick Reference

Copy-paste reference for testing AI and LLM features — what to test, assertion types, pytest and promptfoo snippets, OWASP LLM Top 10, RAG and agent checklists, metrics and thresholds, attack prompts, judge rubrics and tools.