Built here
LLM output evaluation
/ai--llm-eval
/ai--llm-eval is a Claude Code skill in the AI & Agents section. Sets up an evaluation suite for an AI feature: test cases, rubrics, code-based checks, LLM-as-judge with bias controls and a CI gate.
Author's description
LLM output evals - rubrics, LLM-as-judge with bias controls, CI integration
How to ask for it
/ai--llm-eval create a rubric to score my assistant's answers/ai--llm-eval set up an LLM-as-judge with bias controls/ai--llm-eval wire model evals into our CI pipeline
Install
curl -fsSL https://raw.githubusercontent.com/sgomez-dev/claude-skills/main/install.sh | bashAfter installing with the script, type /ai--llm-eval. Using Cursor, Windsurf or Codex? Platform guides
Demo coming soon