SmartQueryTools

Trim Whitespace in Parquet Files Online

Trim leading and trailing whitespace from all text columns in Parquet files, directly in your browser. No upload required.

How to trim Whitespace in Parquet files

  1. Drop your file onto the upload area. The tool loads it and reports how many text columns will be trimmed.
  2. Check the column chips. Highlighted columns are text and will be trimmed. Number, date and boolean columns are shown unhighlighted and are left alone.
  3. Click Trim Whitespace. The preview updates to show the trimmed values.
  4. Click Download to save the cleaned file in the same format you uploaded.

Your file is processed locally in your browser and is never uploaded. The free limit is 50 MB per file; larger files work if your device has the memory for them.

Worked example

Wholesale orders were keyed in through a web form, and some product codes and customer names picked up stray spaces. A lookup on product code is failing for those rows.

Input (Parquet)

Parquet file (binary, columnar) — shown as a table with its schema

product_codecustomerqty
PX-204Blue Harbor Cafe 12
PX-310 Lindqvist AB4
PX-118Casa Moreno20
PX-204Blue Harbor Cafe6

Schema: product_code VARCHAR, customer VARCHAR, qty BIGINT

Settings

  • No options. Every text column is trimmed: product_code and customer.

Result

product_codecustomerqty
PX-204Blue Harbor Cafe12
PX-310Lindqvist AB4
PX-118Casa Moreno20
PX-204Blue Harbor Cafe6

Spaces at the start and end of each text value are removed, so both PX-204 rows now carry the same code and the same customer name. The double space inside "Blue Harbor Cafe" is kept, because only the ends of a value are trimmed. qty is numeric, so it is not touched.

Working with Parquet files

Parquet is often trimmed after it was built from a messy source. A pipeline may have copied padded CHAR fields from a database or legacy system straight into the file. Every top-level string column is trimmed. Integers, decimals, dates, timestamps and booleans are copied unchanged, with their exact stored types. The highlighted chips show exactly which columns count as text before you run it.

Struct and list columns are not trimmed. They are copied unchanged, even when they contain text, and so are binary columns. To trim a nested text field, run Flatten first so it becomes its own column. A value made only of spaces becomes an empty string, not NULL, and Parquet keeps the two apart. The output is a new Parquet file with the same schema. Row counts and column order do not change. Trimming before writing to a warehouse avoids padded keys that make joins silently miss.

Frequently Asked Questions

Are struct and list columns in my Parquet file trimmed?

No. Only plain string columns are trimmed. Struct, list and binary columns are copied unchanged, so the run still succeeds. Flatten the file first if you need to trim a nested field.

Does trimming turn blank Parquet strings into NULL?

No. A value of only spaces becomes an empty string. Parquet keeps empty strings and NULLs distinct, so use the SQL Query tool with NULLIF(col, '') if you want NULLs.

Which characters are removed?

Ordinary space characters at the start and end of each text value. Tabs, line breaks and non-breaking spaces are not removed, and spaces between words are kept.

Can I choose which columns to trim?

No. Every text column is trimmed automatically, and non-text columns are left alone. To trim only some columns, use the SQL Query tool with TRIM on those columns.

Why trim before deduplicating or joining?

Matching is exact, so "PX-204" and "PX-204 " count as different values. Trimming first lets duplicates, lookups and joins match as expected.

Related Tools