# Preparing an ML dataset (/ml--dataset-prep)

/ml--dataset-prep is a Claude Code skill in the AI & Agents section. Turns a raw dataset into a clean, well-split, leakage-free base for modeling: profiling, cleaning, train/validation/test splits and balance checks.

- Web version: https://skills.sgomez.dev/en/s/ml--dataset-prep
- Section: [AI & Agents](https://skills.sgomez.dev/en/ai.md)
- Author: Santiago Gómez de la Torre
- License: MIT
- Source: https://github.com/sgomez-dev/claude-skills/blob/main/skills/ml/dataset-prep.md
- Updated 10 Jul 2026

## How to ask for it

- `/ml--dataset-prep prepare my customer dataset for training`
- `/ml--dataset-prep check for data leakage between train and test`
- `/ml--dataset-prep balance the classes in this imbalanced dataset`

## Install

macOS · Linux:

```
curl -fsSL https://raw.githubusercontent.com/sgomez-dev/claude-skills/main/install.sh | bash
```

Windows:

```
irm https://raw.githubusercontent.com/sgomez-dev/claude-skills/main/install.ps1 | iex
```

Claude Code plugin:

```
/plugin marketplace add sgomez-dev/claude-skills
/plugin install ml-skills@claude-skills-collection
```

## Permissions

- Reads: `**/*.py`, `**/*.ipynb`, `**/*.csv`, `**/*.parquet`, `**/*.json`, `**/*.yaml`, `requirements.txt`, `pyproject.toml`
- Writes: `**/*.py`, `data/**`, `**/*.md`, `**/*.yaml`
- Runs: `python`, `pip`
- Network: No
- Destructive: No

## Author's description

Prepare a dataset for ML — cleaning, splits, leakage checks, balance, versioning

- [How we review this](https://skills.sgomez.dev/en/methodology.md)
