SmartQueryTools

Sample Parquet Files Online

Sample rows from Parquet files — first N, last N, or random — directly in your browser.

How to sample Parquet files

  1. Drop your file onto the upload area. It is loaded into the in-browser engine and the first 200 rows are shown.
  2. Pick a Method: First takes rows from the top, Last takes rows from the bottom, and Random picks rows from anywhere in the file.
  3. Enter the Number of rows you want. The default is 100.
  4. Click Sample. The preview shows the sampled rows and how many were taken.
  5. Click Download to save the sample in the same format you uploaded.

Your file is processed locally in your browser and is never uploaded. The free limit is 50 MB per file; larger files work if your device has the memory for them.

Worked example

A facilities team logs a cold-room temperature every two hours and only needs the latest three readings to check a door alarm.

Input (Parquet)

Parquet file (binary, columnar) — shown as a table with its schema

reading_atsensortemp_c
2026-08-01 00:00:00cold-room-23.8
2026-08-01 02:00:00cold-room-23.9
2026-08-01 04:00:00cold-room-24.1
2026-08-01 06:00:00cold-room-26.7
2026-08-01 08:00:00cold-room-27.2
2026-08-01 10:00:00cold-room-24

Schema: reading_at TIMESTAMP, sensor VARCHAR, temp_c DOUBLE

Settings

  • Method: Last
  • Number of rows: 3

Result

reading_atsensortemp_c
2026-08-01 06:00:00cold-room-26.7
2026-08-01 08:00:00cold-room-27.2
2026-08-01 10:00:00cold-room-24

Last takes rows by their position in the file, not by the timestamp value. The file has six rows, so asking for three returns rows 4 to 6 in their original order. If the log were not already in time order, sort by reading_at first. Random with the same settings would return three rows from anywhere in the file.

Working with Parquet files

A Parquet sample keeps the source schema exactly, including decimal precision, timestamp units and nested struct or list columns. That makes it a good fixture for unit tests and for sharing a small, realistic slice of a production table without sending the whole file. The result is a new, much smaller Parquet file that any Parquet reader can open. Unlike a CSV sample, there is no risk of numbers or dates being re-read as text by the next tool.

Random sampling picks individual rows across all row groups, not whole row groups, so the sample is spread through the file. First and Last use the order rows are stored in. For time-partitioned data written in date order, Last therefore returns the most recent records. The file is fully loaded before sampling, so sampling does not reduce memory use for very large inputs. The row count in the preview matches the downloaded file exactly.

Frequently Asked Questions

Does a Parquet sample keep nested columns?

Yes. Struct, list and map columns are copied as they are, with the same types as the source file.

Is random sampling of Parquet done by row or by row group?

By row. Individual rows are chosen from anywhere in the file, so the sample is not biased toward particular row groups.

What happens if I ask for more rows than the file has?

You get every row. First and Last return the whole file, and Random returns all rows as well.

Is the random sample the same each time?

No. Every click of Sample draws a new random set, and the rows may not keep their file order. For a repeatable sample, use the SQL Query tool with USING SAMPLE reservoir(N ROWS) REPEATABLE (seed).

What is the difference between Sample and Head or Tail?

First and Last in this tool do the same as Head and Tail. Sample adds the Random method, which spreads the selection across the whole file.

Related Tools