SmartQueryTools

Deduplicate Parquet Files Online

Remove duplicate rows from Parquet files instantly in your browser. No upload, no server — 100% private.

How to deduplicate Parquet files

  1. Drop your file onto the upload area. It is loaded into the in-browser engine and the first 200 rows are shown.
  2. Choose the columns that define a duplicate. All columns are ticked by default, which removes only rows that are identical in every field.
  3. Untick columns to match on a subset. For example, keep only customer_id ticked to keep one row per customer.
  4. Click Deduplicate. The tool reports how many rows were removed and previews the result.
  5. Download the cleaned file in the same format you uploaded.

Your file is processed locally in your browser and is never uploaded. The free limit is 50 MB per file; larger files work if your device has the memory for them.

Worked example

A CRM export lists the same contact more than once because the sync job ran twice, and one person also signed up again later with a different date.

Input (Parquet)

Parquet file (binary, columnar) — shown as a table with its schema

emailnameplansignup_date
ana@example.comAna Ruizpro2026-01-04
ben@example.comBen Chofree2026-01-09
ana@example.comAna Ruizpro2026-01-04
cara@example.comCara Doylefree2026-02-11
ben@example.comBen Chopro2026-03-20

Schema: email VARCHAR, name VARCHAR, plan VARCHAR, signup_date DATE

Settings

  • Columns to match on: email only (name, plan and signup_date unticked)

Result

emailnameplansignup_date
ana@example.comAna Ruizpro2026-01-04
ben@example.comBen Chofree2026-01-09
cara@example.comCara Doylefree2026-02-11

Matching on email alone collapses five rows to three: one row per address. With all columns ticked, only the exact repeat of Ana's row would be removed. Ben's two rows would both stay because his plan and signup date differ. Which row survives for each email is not guaranteed, so sort first if you need the earliest or latest one.

Working with Parquet files

Parquet columns keep their stored types, so 1001 in an INT64 column and "1001" in a string column never match each other. Timestamps are compared at full precision. Two events one microsecond apart are distinct rows even if they look the same when displayed to the second. Round or truncate timestamps first if you want them treated as duplicates.

Struct and list columns are compared by value, element by element. When a file has large nested columns, it is faster and clearer to match on a few scalar key columns. The output is written as a new Parquet file with the same schema, and row groups and compression are rewritten.

Frequently Asked Questions

Can I deduplicate a Parquet file that has nested columns?

Yes. Struct, list and map columns are compared by value. To ignore a nested column when matching, untick it so it is carried through but not used as a key.

Will the deduplicated Parquet file keep my column types?

Yes. Column names and types are preserved exactly. Only the row count changes.

Which row is kept when duplicates are found?

One row per unique combination of the selected columns is kept. When you match on a subset of columns, which of the matching rows survives is not guaranteed. Sort the file first if you need the earliest or latest one, or use Top N per Group with N = 1.

What is the difference between Remove Duplicates and Find Duplicates?

Remove Duplicates writes a cleaned file with the extra rows dropped. Find Duplicates lists which rows are duplicated and how many times each one appears, so you can review them before deleting anything.

Can it catch near-duplicates such as typos in names?

No. Matching is exact. Use Fuzzy Deduplicate for names or addresses that differ slightly, such as "Acme Ltd" and "ACME Limited".

Related Tools