Shuffle Parquet Files Online
Randomly shuffle the row order of Parquet files directly in your browser. Useful for randomising data before sampling or ML train/test splits.
How to shuffle Parquet files
- Drop your file. The row count, column count and first 200 rows are shown.
- Click Shuffle rows. Every row gets a random position in one pass. There are no settings.
- Check the preview, which now shows the first 200 rows of the shuffled result.
- Click Shuffle rows again for a different order, or download the result in the same format you uploaded.
Your file is processed locally in your browser and is never uploaded. The free limit is 50 MB per file; larger files work if your device has the memory for them.
Worked example
A small image-classification dataset lists its files grouped by label, all cats first and then all dogs. Taking the first 80% of rows for training would leave almost no dogs in the training set.
Input (Parquet)
Parquet file (binary, columnar) — shown as a table with its schema
| image_id | label | width_px | height_px |
|---|---|---|---|
| img_0001.jpg | cat | 640 | 480 |
| img_0002.jpg | cat | 800 | 600 |
| img_0003.jpg | cat | 640 | 640 |
| img_0004.jpg | dog | 1024 | 768 |
| img_0005.jpg | dog | 640 | 480 |
Schema: image_id VARCHAR, label VARCHAR, width_px BIGINT, height_px BIGINT
Settings
- No options: click Shuffle rows
Result
| image_id | label | width_px | height_px |
|---|---|---|---|
| img_0004.jpg | dog | 1024 | 768 |
| img_0002.jpg | cat | 800 | 600 |
| img_0005.jpg | dog | 640 | 480 |
| img_0001.jpg | cat | 640 | 480 |
| img_0003.jpg | cat | 640 | 640 |
The same five rows are there with every value unchanged, but the labels are now mixed through the file. This is one possible order. Each click produces a new random order and there is no seed to repeat it, so keep the downloaded file if you need the same split later. The first four rows of the shuffled file now make a random 80% sample.
Working with Parquet files
Parquet files are often written sorted by time or by ID, and readers use the minimum and maximum stored for each row group to skip data they do not need. After a shuffle every row group spans the full range of values, so those statistics stop helping filtered queries. That is fine for a training file that is read in full. Keep the sorted original for analytics.
Compression can also get worse. Sorted columns compress well because neighbouring values repeat, and shuffling breaks up those runs. The shuffled file can be noticeably larger than the input for that reason alone. The schema is kept exactly: column names, types, nested structs and lists are unchanged, and timestamps keep full precision because the in-browser engine writes the Parquet file directly. Row groups are rebuilt from scratch, so their number and size can differ from the input file.
Frequently Asked Questions
Why is my shuffled Parquet file bigger than the original?
Sorted data compresses well because similar values sit next to each other. Shuffling spreads them out, so run-length and dictionary encoding save less. The data is the same.
Are Parquet column types kept after shuffling?
Yes. The file is rewritten with the same column names and types. Only the row order, the row group layout and the compression change.
Can I get the same shuffle again?
Not from this tool. Each run uses a new random order with no seed. Save the shuffled file and reuse it. For a repeatable order, use the SQL Query tool and sort by a hash of a key column, for example ORDER BY hash(id).
Does shuffling change any values?
No values are changed and every row is kept. The file is written again by the exporter, though, so formatting can change. For example, a CSV value of 1.50 in a number column is written as 1.5.
How do I make a train and test split after shuffling?
Download the shuffled file, then use the Split tool with the size of your training set as the row count. The first file is your training set and the second your test set. For one random subset without a split, the Sample tool's random mode does it in one step.
Related Tools
Split Parquet Files Online
Split Parquet files into multiple smaller files by row count, directly in your browser.
Extract Head of Parquet Files Online
Extract the first N rows from Parquet files directly in your browser. Choose how many rows to keep and download the result — no upload required.
Sample Parquet Files Online
Sample rows from Parquet files — first N, last N, or random — directly in your browser.
Shuffle CSV Files Online
Randomly shuffle the row order of CSV files directly in your browser. Useful for randomising data before sampling or ML train/test splits.
Shuffle Excel Files Online
Randomly shuffle the row order of Excel files directly in your browser. Useful for randomising data before sampling or ML train/test splits.
Shuffle JSON Files Online
Randomly shuffle the row order of JSON files directly in your browser. Useful for randomising data before sampling or ML train/test splits.
Parquet Viewer Online
View and inspect Parquet files directly in your browser. Browse rows, check column names and data types — no upload required, your data stays on your device.
Convert Parquet to CSV Online
Convert Parquet files to CSV format directly in your browser. No upload required — your data never leaves your device.