Vaex out-of-core DataFrames: analyze billions of rows without the RAM
Process massive tabular datasets with lazy, memory-mapped Vaex DataFrames — filtering, virtual columns, fast aggregations, big-data visualizations, format conversion, ML integration
- What
- Process massive tabular datasets with lazy, memory-mapped Vaex DataFrames — filtering, virtual columns, fast aggregations, big-data visualizations, format conversion, ML integration
- Cost
- Free
- Needs
- Python 3.10+ (3.12+ recommended); uv or pip to install vaex; a large tabular dataset to analyze — the skill is a guided reference, not software
- Install
- Copy the installer prompt below into your Muse — your agent does the rest.
Curated by Skill Harbor — @k-dense-ai's Vaex skill for analyzing tabular datasets too large for RAM — billions of rows processed at over a billion rows per second via lazy, memory-mapped, out-of-core DataFrames. Covers six capability areas: DataFrames and data loading (HDF5, CSV, Arrow, Parquet, pandas/NumPy conversion), data processing and manipulation (filtering, virtual columns, expressions, groupby aggregations), performance and optimization (lazy evaluation, delay=True batching, caching), data visualization (heatmaps, histograms, scatter plots via df.viz), machine learning integration (scalers, encoders, PCA, K-means, scikit-learn/XGBoost/CatBoost), and I/O (format recommendations, export strategies). Includes concrete common patterns: converting a large CSV to HDF5 for instant future loads, batching multiple aggregations with delay=True, and feature engineering with zero-memory-overhead virtual columns. States the Vaex-vs-alternatives rule honestly: polars when data fits in RAM and you need max in-memory speed, dask for distributed clusters, Vaex for single-machine out-of-core analytics. Honest caveats: requires Python 3.10+ (3.12+ recommended with vaex 4.19.0); the skill asks that substantial use be cited in manuscripts (arXiv:2609.00065). MIT licensed. Skill Harbor never reviews the code, review it yourself before use. Discovered via skills.sh.
Version:
Install
Prerequisites: Python 3.10+ (3.12+ recommended); uv or pip to install vaex; a large tabular dataset to analyze — the skill is a guided reference, not software Install "Vaex out-of-core DataFrames: analyze billions of rows without the RAM" for me. It gives my agent @k-dense-ai's Vaex playbook: load and convert large files (HDF5, CSV, Arrow, Parquet), filter and engineer features with zero-overhead virtual columns, batch fast aggregations, visualize billions of rows, build ML pipelines, and choose honestly between Vaex, polars and dask by workload. MIT-licensed. Repository: https://github.com/k-dense-ai/scientific-agent-skills/blob/main/skills/vaex/SKILL.md 1. Fetch the SKILL.md file (and any helper files) from the repository path into a temporary folder and summarize what it does in one or two sentences. 2. Safety check: review the SKILL.md and scripts for anything suspicious (unexpected network calls, shell commands, credential harvesting). This repo should contain zero secrets in code, credentials only via the secure vault, allowed hosts declared in the SKILL.md. Verify that holds here; STOP on any red flag and tell me. 3. Install it as a skill: copy SKILL.md and its helper files into the agent's skills directory, in a folder named "vaex". 4. Verify with no network calls: frontmatter valid, files in place. 5. Report what was installed, where, and what I still need to do myself (e.g. install vaex via uv pip install vaex; point the agent at my large data files; cite the Scientific Agent Skills paper (arXiv:2609.00065) if it materially contributed to a manuscript). GitHub is optional: if I have a GitHub account or the gh CLI, you may use it; otherwise public access is fine. Never require it unless it's in the prerequisites above. Rules: don't touch anything outside the temp folder and the install target. If anything looks off, stop and ask me.
Questions
How do I install a build?
Every product page includes a copy-paste install prompt. Paste it into your Muse and it sets the build up for you — no manual configuration.
Where does my money go?
Straight to the seller. Skill Harbor never processes payments: checkout happens on the seller’s own page, usually Stripe.
What does the ✓ next to a creator’s name mean?
It means we confirmed the identity of the person behind the listing. It says nothing about the code itself — always check a build before installing it.