SmartQueryTools

Split Column in Parquet Files Online

Split a column into multiple columns by delimiter in Parquet files directly in your browser. Turn "First Last" into separate first and last name columns — no upload required.

How to split Column in Parquet files

  1. Drop your file onto the upload area. The first column is selected as the column to split, and the first 200 rows are shown.
  2. Pick the column to split and type the delimiter. The default is a single space. Any text works, including multi-character delimiters such as " | ".
  3. Set the number of parts, from 2 to 10. Each part gets a name box, pre-filled as column_1, column_2 and so on, which you can rename.
  4. Tick Remove original column if you only want the parts, then click Split Column.
  5. Check the preview and download the file. The new columns sit directly after the source column, or in its place if you removed it.

Your file is processed locally in your browser and is never uploaded. The free limit is 50 MB per file; larger files work if your device has the memory for them.

Worked example

A warehouse stock sheet packs category, colour and size into one SKU code. The buying team wants to filter by colour and size separately.

Input (Parquet)

Parquet file (binary, columnar) — shown as a table with its schema

skustockwarehouse
SHOE-RED-4214WLG
SHOE-BLK-393AKL
HAT-BLUE22WLG
SOCK-WHT-M-3PK40CHC

Schema: sku VARCHAR, stock BIGINT, warehouse VARCHAR

Settings

  • Column to split: sku
  • Delimiter: -
  • Number of parts: 3, named category, color, size
  • Remove original column: ticked

Result

categorycolorsizestockwarehouse
SHOERED4214WLG
SHOEBLK393AKL
HATBLUE22WLG
SOCKWHTM40CHC

Each part is taken by position. HAT-BLUE has only two parts, so its size is an empty string. SOCK-WHT-M-3PK has four, and anything after the third part is dropped, so 3PK is lost. Raise the number of parts to keep it. The size column is text even where it holds numbers like 42.

Working with Parquet files

The source column can be any type. It is cast to text before splitting, so an integer like 20260411 or a date column can be split too, although a date is split on its text form 2026-04-11. The new part columns are always strings in the output schema. If you want the pieces as integers or dates again, run Cast Columns afterwards.

Parquet readers distinguish an empty string from NULL, and this tool writes empty strings for missing parts. A NULL in the source column gives NULL in every part. If downstream SQL filters with IS NULL, rows with too few parts will not be caught, so filter on = '' as well or fill them. Removing the original column changes the schema, which can break a table definition that expects it, so leave the box unticked if the file feeds a fixed schema.

Frequently Asked Questions

What type are the new columns when I split a Parquet column?

Always strings (VARCHAR), whatever the source type. Use Cast Columns afterwards to turn numeric parts back into integers or decimals.

Does splitting a Parquet column write NULL for missing parts?

No. Missing parts are empty strings. Only a NULL source value gives NULL in the parts.

What happens when a value has more parts than I asked for?

The extra parts are dropped. Splitting "Mary Ann Lee" into 2 parts on a space gives "Mary" and "Ann", and "Lee" is lost. Set the number of parts to the largest count in your data.

What happens when a value has fewer parts?

The missing parts are empty strings, not NULL. A NULL in the source column gives NULL in every part.

Is the delimiter a regular expression?

No. It is matched as literal text and is case-sensitive. Use Regex Extract if you need a pattern.

Related Tools