# Claude Code skills for AI agents, RAG and MCP

Claude Code skills to build agents and chatbots, evaluate RAG and models, create MCP servers and cut the cost of LLM calls.

These Claude Code skills are meant for people building products with language models. To start, [design an AI agent](https://skills.sgomez.dev/en/s/ai--agent-builder.md) defines the loop, the tools and the memory, [scaffold a chatbot](https://skills.sgomez.dev/en/s/ai--chatbot-scaffold.md) sets up streaming responses and persistent conversations, and [build an MCP server](https://skills.sgomez.dev/en/s/ai--mcp-server.md) exposes your project's functions through the Model Context Protocol.

To make it work well, [improve your prompts](https://skills.sgomez.dev/en/s/ai--prompt-engineer.md) reviews and rewrites instructions and tool definitions, [evaluate an AI feature](https://skills.sgomez.dev/en/s/ai--llm-eval.md) prepares test cases and rubrics, and [measure a RAG pipeline](https://skills.sgomez.dev/en/s/ai--rag-eval.md) builds a golden question set and retrieval metrics. For safety, [add guardrails to an AI app](https://skills.sgomez.dev/en/s/ai--guardrails.md) covers input and output filters and prompt-injection defense.

If cost or machine learning is your concern, [cut token spending](https://skills.sgomez.dev/en/s/ai--llm-cost-optimizer.md) measures where tokens go and applies model routing and caching, [train a model](https://skills.sgomez.dev/en/s/ml--model-training.md) starts from a trivial baseline, and [forecast a time series](https://skills.sgomez.dev/en/s/ml--time-series-forecast.md) uses backtesting to validate the forecast. Each one starts from a plain-language request in Claude Code, and you review the result before using it.

Web version: https://skills.sgomez.dev/en/ai
Updated 15 Sept 2026

## Skills (32)

- [Web and social media research (/agent-reach)](https://skills.sgomez.dev/en/s/agent-reach.md): Searches and reads content from the web and 16 platforms (Reddit, YouTube, X, LinkedIn, GitHub, RSS…) to research a topic. Read-only: it never posts.
- [Building an AI agent (/ai--agent-builder)](https://skills.sgomez.dev/en/s/ai--agent-builder.md): Designs and implements an LLM agent in your codebase: the loop, tools, memory, stop limits and a task-level eval suite to run before shipping.
- [Integrating an AI API (/ai--ai-integration)](https://skills.sgomez.dev/en/s/ai--ai-integration.md): Adds an AI API integration (Claude, OpenAI) to your app: client setup, retries, error handling, streaming, typed requests and tests with mocked responses.
- [Production chatbot scaffold (/ai--chatbot-scaffold)](https://skills.sgomez.dev/en/s/ai--chatbot-scaffold.md): Scaffolds a chatbot on your existing stack: streaming responses, persistent conversation history, optional RAG grounding and a feedback loop.
- [LLM context window tuning (/ai--context-engineering)](https://skills.sgomez.dev/en/s/ai--context-engineering.md): Audits what goes into your AI app's context window and decides what to retrieve, summarize or compact, then checks the result with an eval.
- [Embeddings and semantic search (/ai--embeddings)](https://skills.sgomez.dev/en/s/ai--embeddings.md): Implements vector embeddings in your app for semantic search, RAG or similarity matching: ingestion, chunking, querying and relevance tuning.
- [Fine-tuning a language model (/ai--fine-tuning)](https://skills.sgomez.dev/en/s/ai--fine-tuning.md): Decides honestly whether an LLM needs fine-tuning or better prompts, and if so prepares data, trains and compares before and after on held-out data.
- [LLM safety guardrails (/ai--guardrails)](https://skills.sgomez.dev/en/s/ai--guardrails.md): Adds input and output filters, prompt-injection defense and personal-data handling around an AI feature, then proves them with adversarial tests.
- [Cutting LLM costs (/ai--llm-cost-optimizer)](https://skills.sgomez.dev/en/s/ai--llm-cost-optimizer.md): Measures where your app spends tokens, then applies model routing, caching and prompt compression, keeping each change only if an eval shows quality holds.
- [LLM output evaluation (/ai--llm-eval)](https://skills.sgomez.dev/en/s/ai--llm-eval.md): Sets up an evaluation suite for an AI feature: test cases, rubrics, code-based checks, LLM-as-judge with bias controls and a CI gate.
- [LLM call monitoring (/ai--llm-observability)](https://skills.sgomez.dev/en/s/ai--llm-observability.md): Instruments your AI app with traces, token and cost tracking, quality signals, dashboards and alerts for every model call.
- [Building an MCP server (/ai--mcp-server)](https://skills.sgomez.dev/en/s/ai--mcp-server.md): Builds a Model Context Protocol server that exposes your project's functions, APIs or data to clients like Claude Code, with transport, auth and tests.
- [Multi-agent system design (/ai--multi-agent)](https://skills.sgomez.dev/en/s/ai--multi-agent.md): Designs a system of several AI agents with an orchestration pattern, contracts between agents and failure handling, or advises a single agent instead.
- [Better AI prompts (/ai--prompt-engineer)](https://skills.sgomez.dev/en/s/ai--prompt-engineer.md): Reviews and rewrites your system prompts, user prompts and tool definitions, then gives an A/B comparison of the original and improved versions.
- [RAG pipeline evaluation (/ai--rag-eval)](https://skills.sgomez.dev/en/s/ai--rag-eval.md): Builds a test harness that measures a RAG pipeline: a golden question set, retrieval metrics, groundedness checks and a regression gate.
- [Semantic cache for LLM calls (/ai--semantic-cache)](https://skills.sgomez.dev/en/s/ai--semantic-cache.md): Adds a cache in front of LLM calls so repeated or near-duplicate requests get a stored answer, and measures whether it helps without serving wrong ones.
- [Reliable LLM structured output (/ai--structured-output)](https://skills.sgomez.dev/en/s/ai--structured-output.md): Makes a model return dependable JSON or typed objects: schema design, native structured modes, validation and bounded repair loops.
- [LLM tool calling (/ai--tool-calling)](https://skills.sgomez.dev/en/s/ai--tool-calling.md): Audits or implements the tools a model can call: clear schemas, parallel calls, error handling that lets the model recover, and an accuracy eval.
- [Real-time voice agent (/ai--voice-agent)](https://skills.sgomez.dev/en/s/ai--voice-agent.md): Designs a low-latency voice agent that handles interruptions: speech-to-text and text-to-speech choice, latency budget and conversation design.
- [In-depth web research (/deep-research)](https://skills.sgomez.dev/en/s/deep-research.md): Applies a multi-phase research method, with several angles and sources, before answering or creating content that needs up-to-date information.
- [Skills health check (/meta--health-check)](https://skills.sgomez.dev/en/s/meta--health-check.md): Checks every installed skill for structure, declared permissions and safety, then produces a health report of what passes and what fails.
- [Running a skill pipeline (/meta--pipeline-run)](https://skills.sgomez.dev/en/s/meta--pipeline-run.md): Runs a chain of skills defined in a pipeline file, with quality gates between steps and a final summary showing the result of each step.
- [Creating a new skill (/meta--skill-forge)](https://skills.sgomez.dev/en/s/meta--skill-forge.md): Generates a new skill from a plain-language description, together with its permission manifest and a test file.
- [Project-specific skills setup (/meta--skills-init)](https://skills.sgomez.dev/en/s/meta--skills-init.md): Scans your project, detects its tech stack and activates only the relevant skills, saving the choice as a project configuration file.
- [Preparing an ML dataset (/ml--dataset-prep)](https://skills.sgomez.dev/en/s/ml--dataset-prep.md): Turns a raw dataset into a clean, well-split, leakage-free base for modeling: profiling, cleaning, train/validation/test splits and balance checks.
- [Feature engineering for ML (/ml--feature-engineering)](https://skills.sgomez.dev/en/s/ml--feature-engineering.md): Designs features for a tabular ML problem inside a leakage-safe pipeline: encodings, scaling, interactions and date-based features.
- [MLOps setup (/ml--mlops-pipeline)](https://skills.sgomez.dev/en/s/ml--mlops-pipeline.md): Gives an ML project reproducible training, experiment tracking, a model registry and CI, sized to the team instead of a full platform.
- [Deploying an ML model (/ml--model-deployment)](https://skills.sgomez.dev/en/s/ml--model-deployment.md): Takes a trained model to production: batch or online serving pattern, prediction API, packaging, monitoring, drift detection and rollback plan.
- [ML model evaluation (/ml--model-evaluation)](https://skills.sgomez.dev/en/s/ml--model-evaluation.md): Evaluates a model beyond one headline number: metrics matched to error costs, calibration, per-slice performance and error-by-error analysis.
- [Training an ML model (/ml--model-training)](https://skills.sgomez.dev/en/s/ml--model-training.md): Trains a model methodically: a trivial baseline first, then the simplest model that could work, with cross-validation and tuning.
- [Recommender system (/ml--recommender-system)](https://skills.sgomez.dev/en/s/ml--recommender-system.md): Builds a recommender that beats a popularity baseline: picks the approach, handles cold start and evaluates it offline.
- [Time series forecasting (/ml--time-series-forecast)](https://skills.sgomez.dev/en/s/ml--time-series-forecast.md): Builds a forecast with naive baselines, seasonality, backtesting and prediction intervals, without leaking future data.
