Built here
RAG pipeline evaluation
/ai--rag-eval
/ai--rag-eval is a Claude Code skill in the AI & Agents section. Builds a test harness that measures a RAG pipeline: a golden question set, retrieval metrics, groundedness checks and a regression gate.
Author's description
Evaluate a RAG pipeline - retrieval metrics, groundedness, golden set, regression
Use it when
- You already have a RAG system and do not know if it answers well
- You are about to change chunking, embeddings or the prompt and want to measure the effect
- You want to catch made-up answers and retrieval misses
Not for
- Building the RAG pipeline from scratch (that is what ai--embeddings is for)
- Replacing review by a human expert on your content
What you get
A versioned golden question set in your repo and one command that prints a scorecard with recall@k, MRR, groundedness, correctness, abstention, cost and latency.
How to ask for it
/ai--rag-eval measure whether my RAG pipeline retrieves the right documents/ai--rag-eval build a golden set to evaluate RAG answers/ai--rag-eval check whether a recent change regressed our RAG quality
Install
curl -fsSL https://raw.githubusercontent.com/sgomez-dev/claude-skills/main/install.sh | bashAfter installing with the script, type /ai--rag-eval. Using Cursor, Windsurf or Codex? Platform guides
Questions about this skill
- How many questions does the golden set need?
- Between 30 and 100, each with the expected answer and the chunks that hold the evidence. It includes questions the corpus cannot answer, to check that the system admits it does not know.
- Why does it measure retrieval first?
- Because retrieval is usually the bottleneck: if recall@k is low, no prompt work will fix the answers. The skill notes that most hallucinations turn out to be retrieval misses.
- Does it use an LLM as a judge?
- Yes, for groundedness: a model lists the claims not supported by the retrieved context. It applies judge-bias controls such as a fixed rubric and spot-checks against human labels.
Demo coming soon
Pairs well with
- Recipe: /pipeline--llm-app LLM App