Credit risk data cleaning and variable screening pipeline
Pre-loan modeling data cleaning — abnormal period filtering, missing-rate/IV/PSI/Null Importance/correlation screening and Excel reporting
- What
- Pre-loan modeling data cleaning — abnormal period filtering, missing-rate/IV/PSI/Null Importance/correlation screening and Excel reporting
- Cost
- Free
- Needs
- Python 3 with the bundled scripts' dependencies; a credit-data file (parquet recommended) as input; nothing else — no accounts or keys required
- Install
- Copy the installer prompt below into your Muse — your agent does the rest.
Curated by Skill Harbor — a data-science skill from the `github/awesome-copilot` community collection: an 11-step pipeline that cleans raw credit data and screens variables before pre-loan modeling — load and format raw data, analyze organization sample counts and bad-sample rates, separate out-of-sample data, filter abnormal months, compute missing rates, drop high-missing-rate features, filter low-IV and unstable high-PSI features, apply Null Importance denoising (label-permutation noise removal), drop high-correlation features, and export a full Excel report (15 sheets: summary, per-organization statistics, missing-rate/IV/PSI detail and distribution sheets, each screening step's results). Every step runs independently without deleting the original data, supports interactive parameters with defaults, and accelerates IV/PSI calculations with multiprocessing. By @github, listed here with credit to its creator. Honest caveats: niche domain — pre-loan credit-risk modeling; report sheets are in Chinese; needs Python and the bundled scripts with a parquet (or similar) data file as input. MIT-licensed. Skill Harbor never reviews the code, review it yourself before use. Discovered via skills.sh.
Version:
Install
Prerequisites: Python 3 with the bundled scripts' dependencies; a credit-data file (parquet recommended) as input; nothing else — no accounts or keys required Install "Credit risk data cleaning and variable screening pipeline" for me. It gives the agent the github/awesome-copilot 11-step data-cleaning pipeline for pre-loan credit modeling: organization sample analysis, out-of-sample separation, abnormal-month filtering, missing-rate calculation, high-missing/low-IV/high-PSI screening, Null Importance denoising, high-correlation removal, and a full Excel report — each step independent, interactive parameters with defaults, multiprocessing for IV/PSI. MIT-licensed. Repository: https://github.com/github/awesome-copilot/blob/main/skills/datanalysis-credit-risk/SKILL.md 1. Fetch the SKILL.md file (and any helper files) from the repository path into a temporary folder and summarize what it does in one or two sentences. 2. Safety check: review the SKILL.md and scripts for anything suspicious (unexpected network calls, shell commands, credential harvesting). This repo should contain zero secrets in code, credentials only via the secure vault, allowed hosts declared in the SKILL.md. Verify that holds here; STOP on any red flag and tell me. 3. Install it as a skill: copy SKILL.md and its helper files into the agent's skills directory, in a folder named "datanalysis-credit-risk". 4. Verify with no network calls: frontmatter valid, files in place. 5. Report what was installed, where, and what I still need to do myself (e.g. prepare Python 3 with the dependencies; point the agent at the credit-data file; note the report sheets are in Chinese). GitHub is optional: if I have a GitHub account or the gh CLI, you may use it; otherwise public access is fine. Never require it unless it's in the prerequisites above. Rules: don't touch anything outside the temp folder and the install target. If anything looks off, stop and ask me.
Questions
How do I install a build?
Every product page includes a copy-paste install prompt. Paste it into your Muse and it sets the build up for you — no manual configuration.
Where does my money go?
Straight to the seller. Skill Harbor never processes payments: checkout happens on the seller’s own page, usually Stripe.
What does the ✓ next to a creator’s name mean?
It means we confirmed the identity of the person behind the listing. It says nothing about the code itself — always check a build before installing it.