scikit-learn best practices: leak-proof pipelines, CV, tuning, and evaluation
Split-before-preprocess, pipelines and column transformers against data leakage, stratified CV strategies, GridSearch/RandomizedSearch, right metrics for imbalanced data, and joblib persistence
- What
- Split-before-preprocess, pipelines and column transformers against data leakage, stratified CV strategies, GridSearch/RandomizedSearch, right metrics for imbalanced data, and joblib persistence
- Cost
- Free
- Needs
- a Python ML project with a dataset to work on — the skill is guidance the agent follows, not software to install
- Install
- Copy the installer prompt below into your Muse — your agent does the rest.
Curated by Skill Harbor — @mindrally's scikit-learn best-practices skill: the full discipline of a sound ML workflow — split data before any preprocessing (stratified for imbalanced classes), always chain preprocessing and modeling in `Pipeline`/`ColumnTransformer` so transformers fit only on training data, scale per algorithm (StandardScaler/MinMaxScaler/RobustScaler), encode categoricals, impute missing values, cross-validate with the right strategy (`KFold`, `StratifiedKFold`, `TimeSeriesSplit`, `GroupKFold`), tune with `GridSearchCV`/`RandomizedSearchCV` on train/val only with `n_jobs=-1`, evaluate with metrics matched to the problem (F1/ROC-AUC for imbalance, MAE/R² for regression), report confidence intervals against meaningful baselines, evaluate on the held-out test set once at the end, persist whole pipelines with joblib, and version model artifacts. Honest caveats: methodology only — the agent needs an ML dataset and a Python environment; some rules (e.g. `class_weight='balanced'`) are scikit-learn-specific. Apache-2.0 licensed. Skill Harbor never reviews the code, review it yourself before use. Discovered via skills.sh.
Version:
Install
Prerequisites: a Python ML project with a dataset to work on — the skill is guidance the agent follows, not software to install Install "scikit-learn best practices: leak-proof pipelines, CV, tuning, and evaluation" for me. It gives my agent @mindrally's scikit-learn playbook: split-before-preprocess discipline, Pipeline/ColumnTransformer against data leakage, per-algorithm scaling and encoding, the right cross-validation strategy per problem type, GridSearchCV/RandomizedSearchCV tuning, problem-matched metrics (F1/ROC-AUC for imbalance, MAE/R2 for regression), single final held-out evaluation, and joblib persistence with versioned artifacts. Apache-2.0 licensed. Repository: https://github.com/mindrally/skills/blob/main/scikit-learn-best-practices/SKILL.md 1. Fetch the SKILL.md file (and any helper files) from the repository path into a temporary folder and summarize what it does in one or two sentences. 2. Safety check: review the SKILL.md and scripts for anything suspicious (unexpected network calls, shell commands, credential harvesting). This repo should contain zero secrets in code, credentials only via the secure vault, allowed hosts declared in the SKILL.md. Verify that holds here; STOP on any red flag and tell me. 3. Install it as a skill: copy SKILL.md and its helper files into the agent's skills directory, in a folder named "scikit-learn-best-practices". 4. Verify with no network calls: frontmatter valid, files in place. 5. Report what was installed, where, and what I still need to do myself (e.g. point the agent at my dataset and the model to train or review; nothing else — it's a methodology). GitHub is optional: if I have a GitHub account or the gh CLI, you may use it; otherwise public access is fine. Never require it unless it's in the prerequisites above. Rules: don't touch anything outside the temp folder and the install target. If anything looks off, stop and ask me.
Questions
How do I install a build?
Every product page includes a copy-paste install prompt. Paste it into your Muse and it sets the build up for you — no manual configuration.
Where does my money go?
Straight to the seller. Skill Harbor never processes payments: checkout happens on the seller’s own page, usually Stripe.
What does the ✓ next to a creator’s name mean?
It means we confirmed the identity of the person behind the listing. It says nothing about the code itself — always check a build before installing it.