SmartQueryTools

Shuffle YAML Files Online

Randomly shuffle the row order of YAML files directly in your browser. Useful for randomising data before sampling or ML train/test splits.

How to shuffle YAML files

  1. Drop your file. The row count, column count and first 200 rows are shown.
  2. Click Shuffle rows. Every row gets a random position in one pass. There are no settings.
  3. Check the preview, which now shows the first 200 rows of the shuffled result.
  4. Click Shuffle rows again for a different order, or download the result in the same format you uploaded.

Your file is processed locally in your browser and is never uploaded. The free limit is 50 MB per file; larger files work if your device has the memory for them.

Worked example

A small image-classification dataset lists its files grouped by label, all cats first and then all dogs. Taking the first 80% of rows for training would leave almost no dogs in the training set.

Input (YAML)

- image_id: img_0001.jpg
  label: cat
  width_px: 640
  height_px: 480
- image_id: img_0002.jpg
  label: cat
  width_px: 800
  height_px: 600
- image_id: img_0003.jpg
  label: cat
  width_px: 640
  height_px: 640
- image_id: img_0004.jpg
  label: dog
  width_px: 1024
  height_px: 768
- image_id: img_0005.jpg
  label: dog
  width_px: 640
  height_px: 480

Settings

  • No options: click Shuffle rows

Result

image_idlabelwidth_pxheight_px
img_0004.jpgdog1024768
img_0002.jpgcat800600
img_0005.jpgdog640480
img_0001.jpgcat640480
img_0003.jpgcat640640

The same five rows are there with every value unchanged, but the labels are now mixed through the file. This is one possible order. Each click produces a new random order and there is no seed to repeat it, so keep the downloaded file if you need the same split later. The first four rows of the shuffled file now make a random 80% sample.

Frequently Asked Questions

Can I get the same shuffle again?

Not from this tool. Each run uses a new random order with no seed. Save the shuffled file and reuse it. For a repeatable order, use the SQL Query tool and sort by a hash of a key column, for example ORDER BY hash(id).

Does shuffling change any values?

No values are changed and every row is kept. The file is written again by the exporter, though, so formatting can change. For example, a CSV value of 1.50 in a number column is written as 1.5.

How do I make a train and test split after shuffling?

Download the shuffled file, then use the Split tool with the size of your training set as the row count. The first file is your training set and the second your test set. For one random subset without a split, the Sample tool's random mode does it in one step.

Related Tools