SmartQueryTools

Shuffle CSV Files Online

Randomly shuffle the row order of CSV files directly in your browser. Useful for randomising data before sampling or ML train/test splits.

How to shuffle CSV files

  1. Drop your file. The row count, column count and first 200 rows are shown.
  2. Click Shuffle rows. Every row gets a random position in one pass. There are no settings.
  3. Check the preview, which now shows the first 200 rows of the shuffled result.
  4. Click Shuffle rows again for a different order, or download the result in the same format you uploaded.

Your file is processed locally in your browser and is never uploaded. The free limit is 50 MB per file; larger files work if your device has the memory for them.

Worked example

A small image-classification dataset lists its files grouped by label, all cats first and then all dogs. Taking the first 80% of rows for training would leave almost no dogs in the training set.

Input (CSV)

image_id,label,width_px,height_px
img_0001.jpg,cat,640,480
img_0002.jpg,cat,800,600
img_0003.jpg,cat,640,640
img_0004.jpg,dog,1024,768
img_0005.jpg,dog,640,480

Settings

  • No options: click Shuffle rows

Result

image_idlabelwidth_pxheight_px
img_0004.jpgdog1024768
img_0002.jpgcat800600
img_0005.jpgdog640480
img_0001.jpgcat640480
img_0003.jpgcat640640

The same five rows are there with every value unchanged, but the labels are now mixed through the file. This is one possible order. Each click produces a new random order and there is no seed to repeat it, so keep the downloaded file if you need the same split later. The first four rows of the shuffled file now make a random 80% sample.

Working with CSV files

The header row stays at the top and only data rows move. A record that spans several lines inside quotes, such as a two-line address, is one row and moves as one unit. The CSV is parsed on load and written fresh on download, so the text of some values can change even though the data does not. Numbers lose trailing zeros, dates the reader recognised are written as YYYY-MM-DD, and fields are quoted only when they contain a comma, quote or line break.

Shuffling is a common first step before splitting a CSV into training and test files for Python or R. Pandas and R can shuffle too, but doing it here means the file you share already has the random order fixed in it. Everyone who uses that file gets the same split, which makes results easier to reproduce. The output uses commas as delimiters and LF line endings.

Frequently Asked Questions

Does the CSV header row get shuffled into the data?

No. The first line is read as column names and written back as the first line of the output.

Why do some numbers look different in the shuffled CSV?

The file is rewritten from parsed values. A number column that held 1.50 is written as 1.5, and recognised dates are written as YYYY-MM-DD. The values themselves are the same.

Can I get the same shuffle again?

Not from this tool. Each run uses a new random order with no seed. Save the shuffled file and reuse it. For a repeatable order, use the SQL Query tool and sort by a hash of a key column, for example ORDER BY hash(id).

Does shuffling change any values?

No values are changed and every row is kept. The file is written again by the exporter, though, so formatting can change. For example, a CSV value of 1.50 in a number column is written as 1.5.

How do I make a train and test split after shuffling?

Download the shuffled file, then use the Split tool with the size of your training set as the row count. The first file is your training set and the second your test set. For one random subset without a split, the Sample tool's random mode does it in one step.

Related Tools