SmartQueryTools

Split Parquet Files Online

Split Parquet files into multiple smaller files by row count, directly in your browser.

How to split Parquet files

  1. Drop your file onto the upload area. The tool loads it and shows the total row and column count.
  2. Enter the number of rows each part should hold in Rows per chunk. The default is 1000.
  3. Click Calculate chunks. A button appears for each part, labelled with its row range, for example Part 2 with rows 1,001–2,000.
  4. Click each part to download it. Every part is a separate file in the same format as the original, named with a _part1, _part2 suffix.

Your file is processed locally in your browser and is never uploaded. The free limit is 50 MB per file; larger files work if your device has the memory for them.

Worked example

An outdoor gear shop is uploading its catalogue to a marketplace that rejects files over a set row limit. Scaled down, the limit here is two products per file.

Input (Parquet)

Parquet file (binary, columnar) — shown as a table with its schema

skutitlepricein_stock
TR-100Trail runner shoe119true
TR-101Merino running sock18.5true
TR-102Rain shell jacket164false
TR-103Headlamp 400 lm42.99true
TR-104Hydration vest89true

Schema: sku VARCHAR, title VARCHAR, price DOUBLE, in_stock BOOLEAN

Settings

  • Rows per chunk: 2
  • Downloaded: Part 1 (rows 1–2)

Result

skutitlepricein_stock
TR-100Trail runner shoe119true
TR-101Merino running sock18.5true

Five rows at two per chunk gives three parts: rows 1–2, rows 3–4, and row 5 on its own. Part 1 is shown. Each part has the full header and the rows in their original file order. The last part is smaller because the row count does not divide evenly. Nothing is dropped or repeated between parts.

Working with Parquet files

Each part is written as a complete Parquet file with its own footer and the full original schema. Any part can be read alone, and all parts can be read together as one dataset by tools that accept a folder or a glob such as data_part*.parquet. Column types are unchanged: decimals stay exact, timestamps keep their stored precision, and nested structs and lists pass through as they are.

Parts follow the order rows were stored in the source file. Parquet compresses well, so a chunk of 100,000 rows can be much smaller than the same rows in CSV. Pick the chunk size by row count for downstream jobs, not by the byte size of the original. Splitting is also a simple way to break one large file into pieces for parallel processing. Each part is compressed on its own, so the parts together can be slightly larger than the source. Files are named like name_part1.parquet and name_part2.parquet.

Frequently Asked Questions

Can the split Parquet parts be read back as one dataset?

Yes. They share the same schema, so engines such as DuckDB, Spark or pandas can read them together with a wildcard path like data_part*.parquet.

Does splitting a Parquet file change its schema?

No. Every part has the same column names and types as the original.

Can I split by a column value, such as one file per region?

Not with this tool, which splits by row count only. Filter the file once per value instead, or use the SQL Query tool.

Do I have to download each part separately?

Yes. Each part has its own download button and saves as a separate file. There is no single zip download.

Can I split a very large file?

Files up to 50 MB can be loaded on the free tier. The whole file is read into browser memory before splitting, so the practical limit also depends on your device.

Related Tools