# RAG pipeline evaluation (/ai--rag-eval)

/ai--rag-eval is a Claude Code skill in the AI & Agents section. Builds a test harness that measures a RAG pipeline: a golden question set, retrieval metrics, groundedness checks and a regression gate.

- Web version: https://skills.sgomez.dev/en/s/ai--rag-eval
- Section: [AI & Agents](https://skills.sgomez.dev/en/ai.md)
- Author: Santiago Gómez de la Torre
- License: MIT
- Source: https://github.com/sgomez-dev/claude-skills/blob/main/skills/ai/rag-eval.md
- Updated 10 Jul 2026

## Use it when

- You already have a RAG system and do not know if it answers well
- You are about to change chunking, embeddings or the prompt and want to measure the effect
- You want to catch made-up answers and retrieval misses

## Not for

- Building the RAG pipeline from scratch (that is what ai--embeddings is for)
- Replacing review by a human expert on your content

## What you get

A versioned golden question set in your repo and one command that prints a scorecard with recall@k, MRR, groundedness, correctness, abstention, cost and latency.

## How to ask for it

- `/ai--rag-eval measure whether my RAG pipeline retrieves the right documents`
- `/ai--rag-eval build a golden set to evaluate RAG answers`
- `/ai--rag-eval check whether a recent change regressed our RAG quality`

## Install

macOS · Linux:

```
curl -fsSL https://raw.githubusercontent.com/sgomez-dev/claude-skills/main/install.sh | bash
```

Windows:

```
irm https://raw.githubusercontent.com/sgomez-dev/claude-skills/main/install.ps1 | iex
```

Claude Code plugin:

```
/plugin marketplace add sgomez-dev/claude-skills
/plugin install ai-skills@claude-skills-collection
```

## Permissions

- Reads: `**/*`
- Writes: `**/*`
- Runs: `project test/eval runners`
- Network: No
- Destructive: No

## Author's description

Evaluate a RAG pipeline - retrieval metrics, groundedness, golden set, regression

## Pairs well with

- [Production chatbot scaffold (/ai--chatbot-scaffold)](https://skills.sgomez.dev/en/s/ai--chatbot-scaffold.md): Scaffolds a chatbot on your existing stack: streaming responses, persistent conversation history, optional RAG grounding and a feedback loop.
- [Embeddings and semantic search (/ai--embeddings)](https://skills.sgomez.dev/en/s/ai--embeddings.md): Implements vector embeddings in your app for semantic search, RAG or similarity matching: ingestion, chunking, querying and relevance tuning.
- [LLM safety guardrails (/ai--guardrails)](https://skills.sgomez.dev/en/s/ai--guardrails.md): Adds input and output filters, prompt-injection defense and personal-data handling around an AI feature, then proves them with adversarial tests.
- Recipe: [/pipeline--llm-app](https://github.com/sgomez-dev/claude-skills/blob/main/pipelines/llm-app.yaml): LLM App

## Questions about this skill

### How many questions does the golden set need?

Between 30 and 100, each with the expected answer and the chunks that hold the evidence. It includes questions the corpus cannot answer, to check that the system admits it does not know.

### Why does it measure retrieval first?

Because retrieval is usually the bottleneck: if recall@k is low, no prompt work will fix the answers. The skill notes that most hallucinations turn out to be retrieval misses.

### Does it use an LLM as a judge?

Yes, for groundedness: a model lists the claims not supported by the retrieved context. It applies judge-bias controls such as a fixed rubric and spot-checks against human labels.


- [How we review this](https://skills.sgomez.dev/en/methodology.md)
