SmartQueryTools

Find & Replace in Parquet Files Online

Find and replace text values in Parquet files directly in your browser. Supports plain text and regex patterns across any column — no upload required.

How to find & Replace in Parquet files

  1. Drop your file onto the upload area. The tool shows the row count, how many text columns it found, and the first 200 rows.
  2. In Apply to column, keep All text columns or pick one column. Number, date and boolean columns are never changed.
  3. Type the text to find and the replacement. Leave Replace with empty to delete every match.
  4. Tick Use regular expression if the Find box holds a pattern rather than literal text, then click Find & Replace.
  5. Check the result preview and download the file in the same format.

Your file is processed locally in your browser and is never uploaded. The free limit is 50 MB per file; larger files work if your device has the memory for them.

Worked example

A contractor list was typed in by hand, so phone numbers mix brackets, spaces, dashes and a country prefix. The SMS gateway wants digits only.

Input (Parquet)

Parquet file (binary, columnar) — shown as a table with its schema

namephonecity
Aroha Ngata(021) 555-0143Auckland
Liam Park021 555 0198Hamilton
Mele Tupou+64 21 555 0112Auckland
Sam ReidNULLTauranga

Schema: name VARCHAR, phone VARCHAR, city VARCHAR

Settings

  • Apply to column: phone
  • Find: [^0-9]
  • Replace with: (empty)
  • Use regular expression: ticked

Result

namephonecity
Aroha Ngata0215550143Auckland
Liam Park0215550198Hamilton
Mele Tupou64215550112Auckland
Sam ReidNULLTauranga

The pattern [^0-9] matches any character that is not a digit, and every match in the cell is replaced, not just the first. Brackets, spaces, dashes and the plus sign are removed. The phone column stays text, so the leading zero survives. Sam's empty phone stays empty. Limiting the run to the phone column keeps the other columns untouched.

Working with Parquet files

Parquet columns keep their stored types, so only real string columns are offered. Integer, decimal, timestamp and boolean columns pass through unchanged, which makes an All text columns run safe on wide analytical tables. The replaced columns are written back as strings, and every other column keeps its exact Parquet type. Enum-like string columns such as status or region are the usual target, because one rule updates every row in the file.

Dictionary-encoded string columns, common for low-cardinality fields like country or status, are a good fit for bulk recoding such as replacing "UK" with "GB". The replacement is applied to every row, and the output file is rewritten with fresh row groups. Replacing inside nested struct or list columns is not supported. Those columns are not offered and are copied unchanged, so All text columns is safe on files with nested fields.

Frequently Asked Questions

Does find and replace change my Parquet schema?

No. Column names and types stay the same. Only the values inside the string columns you targeted change.

Does All text columns change struct or list columns in my Parquet file?

No. Only plain string columns are changed. Struct, list and binary columns are copied unchanged. Flatten the file first if you need to edit a nested field.

Is the search case-sensitive?

Yes, in both modes. With Use regular expression ticked, you can start the pattern with (?i) to ignore case, for example (?i)limited.

Can I use capture groups in the replacement?

Yes. In regex mode, write \1, \2 and so on in Replace with. For example, find (\d{4})-(\d{2}) and replace with \2/\1 to turn 2026-04 into 04/2026. The syntax is RE2, so lookaheads and lookbehinds are not supported.

Does it replace every match or only the first one in each cell?

Every match. Both plain text and regex mode replace all occurrences in each cell of the targeted columns.

Related Tools