Skip to content
Built here

LLM output evaluation

/ai--llm-eval

/ai--llm-eval is a Claude Code skill in the AI & Agents section. Sets up an evaluation suite for an AI feature: test cases, rubrics, code-based checks, LLM-as-judge with bias controls and a CI gate.

Author's description

LLM output evals - rubrics, LLM-as-judge with bias controls, CI integration

How to ask for it

  • /ai--llm-eval create a rubric to score my assistant's answers
  • /ai--llm-eval set up an LLM-as-judge with bias controls
  • /ai--llm-eval wire model evals into our CI pipeline

Install

curl -fsSL https://raw.githubusercontent.com/sgomez-dev/claude-skills/main/install.sh | bash

After installing with the script, type /ai--llm-eval. Using Cursor, Windsurf or Codex? Platform guides

Demo coming soon

↑↓ move · Enter open · Esc close