SmartQueryTools

Shuffle Parquet Files Online

Randomly shuffle the row order of Parquet files directly in your browser. Useful for randomising data before sampling or ML train/test splits.

How to shuffle Parquet files

  1. Drop your file. The row count, column count and first 200 rows are shown.
  2. Click Shuffle rows. Every row gets a random position in one pass. There are no settings.
  3. Check the preview, which now shows the first 200 rows of the shuffled result.
  4. Click Shuffle rows again for a different order, or download the result in the same format you uploaded.

Your file is processed locally in your browser and is never uploaded. The free limit is 50 MB per file; larger files work if your device has the memory for them.

Worked example

A small image-classification dataset lists its files grouped by label, all cats first and then all dogs. Taking the first 80% of rows for training would leave almost no dogs in the training set.

Input (Parquet)

Parquet file (binary, columnar) — shown as a table with its schema

image_idlabelwidth_pxheight_px
img_0001.jpgcat640480
img_0002.jpgcat800600
img_0003.jpgcat640640
img_0004.jpgdog1024768
img_0005.jpgdog640480

Schema: image_id VARCHAR, label VARCHAR, width_px BIGINT, height_px BIGINT

Settings

  • No options: click Shuffle rows

Result

image_idlabelwidth_pxheight_px
img_0004.jpgdog1024768
img_0002.jpgcat800600
img_0005.jpgdog640480
img_0001.jpgcat640480
img_0003.jpgcat640640

The same five rows are there with every value unchanged, but the labels are now mixed through the file. This is one possible order. Each click produces a new random order and there is no seed to repeat it, so keep the downloaded file if you need the same split later. The first four rows of the shuffled file now make a random 80% sample.

Working with Parquet files

Parquet files are often written sorted by time or by ID, and readers use the minimum and maximum stored for each row group to skip data they do not need. After a shuffle every row group spans the full range of values, so those statistics stop helping filtered queries. That is fine for a training file that is read in full. Keep the sorted original for analytics.

Compression can also get worse. Sorted columns compress well because neighbouring values repeat, and shuffling breaks up those runs. The shuffled file can be noticeably larger than the input for that reason alone. The schema is kept exactly: column names, types, nested structs and lists are unchanged, and timestamps keep full precision because the in-browser engine writes the Parquet file directly. Row groups are rebuilt from scratch, so their number and size can differ from the input file.

Frequently Asked Questions

Why is my shuffled Parquet file bigger than the original?

Sorted data compresses well because similar values sit next to each other. Shuffling spreads them out, so run-length and dictionary encoding save less. The data is the same.

Are Parquet column types kept after shuffling?

Yes. The file is rewritten with the same column names and types. Only the row order, the row group layout and the compression change.

Can I get the same shuffle again?

Not from this tool. Each run uses a new random order with no seed. Save the shuffled file and reuse it. For a repeatable order, use the SQL Query tool and sort by a hash of a key column, for example ORDER BY hash(id).

Does shuffling change any values?

No values are changed and every row is kept. The file is written again by the exporter, though, so formatting can change. For example, a CSV value of 1.50 in a number column is written as 1.5.

How do I make a train and test split after shuffling?

Download the shuffled file, then use the Split tool with the size of your training set as the row count. The first file is your training set and the second your test set. For one random subset without a split, the Sample tool's random mode does it in one step.

Related Tools