SmartQueryTools

Find Fuzzy Duplicates in Parquet Files Online

Find near-duplicate rows in Parquet files using Levenshtein edit distance or Jaro-Winkler similarity — all in your browser. Set your own threshold and download the matched pairs as CSV — no upload required.

How to find Fuzzy Duplicates in Parquet files

  1. Drop your file onto the upload area. Files with more than 5,000 rows show a notice, because only the first 5,000 rows are compared.
  2. Choose the column to check for near-duplicates, usually a name, company or address column. The first column is selected by default.
  3. Pick a similarity method. Levenshtein counts character edits, with a maximum distance of 1 to 4 (default 2). Jaro-Winkler gives a 0 to 1 score, with a minimum similarity slider from 0.70 to 0.99 (default 0.85).
  4. Click Find Fuzzy Duplicates. The tool lists matching pairs of rows with their values and score.
  5. Review the pairs and click Download Pairs CSV. The file is a list of pairs to check, not a cleaned copy of your data.

Your file is processed locally in your browser and is never uploaded. The free limit is 50 MB per file; larger files work if your device has the memory for them.

Worked example

An accounts payable team merged supplier lists from two systems and suspects some vendors were entered twice with small spelling differences.

Input (Parquet)

Parquet file (binary, columnar) — shown as a table with its schema

vendor_idvendor_namecity
V-001Northgate PlumbingLeeds
V-002Northgate Plumbing LtdLeeds
V-003Brightwater ElectricalYork
V-004Brightwater ElectricYork
V-005Harlow & SonsHull
V-006Northgate PlumbngLeeds

Schema: vendor_id VARCHAR, vendor_name VARCHAR, city VARCHAR

Settings

  • Column: vendor_name
  • Method: Jaro-Winkler
  • Min similarity: 0.85

Result

row_arow_bval_aval_bsimilarity
16Northgate PlumbingNorthgate Plumbng0.9889
34Brightwater ElectricalBrightwater Electric0.9818
12Northgate PlumbingNorthgate Plumbing Ltd0.9636
26Northgate Plumbing LtdNorthgate Plumbng0.9545

row_a and row_b are 1-based row positions in the file, and the most similar pairs come first. The three Northgate rows produce three pairs, because every pair is scored separately. Harlow & Sons matches nothing. With Levenshtein at distance 2 instead, only the typo (distance 1) and Electrical/Electric (distance 2) would be listed. Adding " Ltd" is 4 edits.

Working with Parquet files

Parquet files are often large, and fuzzy matching compares every row with every other row. The tool therefore uses only the first 5,000 rows, about 12.5 million comparisons. For a bigger file, reduce it first. Filter to one region or category, or use Unique Values on the name column so each distinct spelling is compared once rather than once per transaction. A sales table with two million rows may hold only a few thousand distinct customer names, which fits inside the limit.

Any column type can be chosen, and non-string values are converted to text before comparing. A DATE column compares strings like 2026-03-04. That can find transposed day and month typos, but it is rarely what you want. val_a and val_b keep the column's original type in the output. Results are capped at 1,000 pairs and downloaded as CSV, not Parquet, because the output is a review list rather than a new version of the table.

Frequently Asked Questions

Can I fuzzy-deduplicate a Parquet file with a million rows?

Only the first 5,000 rows are compared. Reduce the file first, for example with Unique Values on the target column, then run the fuzzy check on that list.

Why is the result a CSV and not a Parquet file?

The output is a list of matched row pairs to review, so it is always downloaded as CSV.

Does the tool remove the near-duplicates for me?

No. It lists pairs of similar rows with their row numbers and score. You decide which rows to keep, then remove the others yourself or with Filter.

Are exact duplicates included in the pairs?

No. Levenshtein pairs must be at least 1 edit apart, and Jaro-Winkler pairs must not be identical. Use Find Duplicates or Remove Duplicates for exact matches.

Which method should I choose?

Levenshtein suits short values and typos, such as codes or single words. Jaro-Winkler suits names and company names, because it rewards a shared start. For example "Robert" and "Robrt" score about 0.96.

Related Tools

Deduplicate Parquet Files Online

Remove duplicate rows from Parquet files instantly in your browser. No upload, no server — 100% private.

Filter Parquet Files Online

Filter rows in Parquet files by column value, directly in your browser. Your data stays on your device.

Convert Case in Parquet Files Online

Convert text columns to UPPERCASE, lowercase, or Title Case in Parquet files directly in your browser. Apply case conversion to any or all text columns at once — no upload required.

Find Fuzzy Duplicates in CSV Files Online

Find near-duplicate rows in CSV files using Levenshtein edit distance or Jaro-Winkler similarity — all in your browser. Set your own threshold and download the matched pairs as CSV — no upload required.

Find Fuzzy Duplicates in Excel Files Online

Find near-duplicate rows in Excel files using Levenshtein edit distance or Jaro-Winkler similarity — all in your browser. Set your own threshold and download the matched pairs as CSV — no upload required.

Find Fuzzy Duplicates in JSON Files Online

Find near-duplicate rows in JSON files using Levenshtein edit distance or Jaro-Winkler similarity — all in your browser. Set your own threshold and download the matched pairs as CSV — no upload required.

Parquet Viewer Online

View and inspect Parquet files directly in your browser. Browse rows, check column names and data types — no upload required, your data stays on your device.

Convert Parquet to CSV Online

Convert Parquet files to CSV format directly in your browser. No upload required — your data never leaves your device.