Find Fuzzy Duplicates in Parquet Files Online
Find near-duplicate rows in Parquet files using Levenshtein edit distance or Jaro-Winkler similarity — all in your browser. Set your own threshold and download the matched pairs as CSV — no upload required.
How to find Fuzzy Duplicates in Parquet files
- Drop your file onto the upload area. Files with more than 5,000 rows show a notice, because only the first 5,000 rows are compared.
- Choose the column to check for near-duplicates, usually a name, company or address column. The first column is selected by default.
- Pick a similarity method. Levenshtein counts character edits, with a maximum distance of 1 to 4 (default 2). Jaro-Winkler gives a 0 to 1 score, with a minimum similarity slider from 0.70 to 0.99 (default 0.85).
- Click Find Fuzzy Duplicates. The tool lists matching pairs of rows with their values and score.
- Review the pairs and click Download Pairs CSV. The file is a list of pairs to check, not a cleaned copy of your data.
Your file is processed locally in your browser and is never uploaded. The free limit is 50 MB per file; larger files work if your device has the memory for them.
Worked example
An accounts payable team merged supplier lists from two systems and suspects some vendors were entered twice with small spelling differences.
Input (Parquet)
Parquet file (binary, columnar) — shown as a table with its schema
| vendor_id | vendor_name | city |
|---|---|---|
| V-001 | Northgate Plumbing | Leeds |
| V-002 | Northgate Plumbing Ltd | Leeds |
| V-003 | Brightwater Electrical | York |
| V-004 | Brightwater Electric | York |
| V-005 | Harlow & Sons | Hull |
| V-006 | Northgate Plumbng | Leeds |
Schema: vendor_id VARCHAR, vendor_name VARCHAR, city VARCHAR
Settings
- Column: vendor_name
- Method: Jaro-Winkler
- Min similarity: 0.85
Result
| row_a | row_b | val_a | val_b | similarity |
|---|---|---|---|---|
| 1 | 6 | Northgate Plumbing | Northgate Plumbng | 0.9889 |
| 3 | 4 | Brightwater Electrical | Brightwater Electric | 0.9818 |
| 1 | 2 | Northgate Plumbing | Northgate Plumbing Ltd | 0.9636 |
| 2 | 6 | Northgate Plumbing Ltd | Northgate Plumbng | 0.9545 |
row_a and row_b are 1-based row positions in the file, and the most similar pairs come first. The three Northgate rows produce three pairs, because every pair is scored separately. Harlow & Sons matches nothing. With Levenshtein at distance 2 instead, only the typo (distance 1) and Electrical/Electric (distance 2) would be listed. Adding " Ltd" is 4 edits.
Working with Parquet files
Parquet files are often large, and fuzzy matching compares every row with every other row. The tool therefore uses only the first 5,000 rows, about 12.5 million comparisons. For a bigger file, reduce it first. Filter to one region or category, or use Unique Values on the name column so each distinct spelling is compared once rather than once per transaction. A sales table with two million rows may hold only a few thousand distinct customer names, which fits inside the limit.
Any column type can be chosen, and non-string values are converted to text before comparing. A DATE column compares strings like 2026-03-04. That can find transposed day and month typos, but it is rarely what you want. val_a and val_b keep the column's original type in the output. Results are capped at 1,000 pairs and downloaded as CSV, not Parquet, because the output is a review list rather than a new version of the table.
Frequently Asked Questions
Can I fuzzy-deduplicate a Parquet file with a million rows?
Only the first 5,000 rows are compared. Reduce the file first, for example with Unique Values on the target column, then run the fuzzy check on that list.
Why is the result a CSV and not a Parquet file?
The output is a list of matched row pairs to review, so it is always downloaded as CSV.
Does the tool remove the near-duplicates for me?
No. It lists pairs of similar rows with their row numbers and score. You decide which rows to keep, then remove the others yourself or with Filter.
Are exact duplicates included in the pairs?
No. Levenshtein pairs must be at least 1 edit apart, and Jaro-Winkler pairs must not be identical. Use Find Duplicates or Remove Duplicates for exact matches.
Which method should I choose?
Levenshtein suits short values and typos, such as codes or single words. Jaro-Winkler suits names and company names, because it rewards a shared start. For example "Robert" and "Robrt" score about 0.96.
Related Tools
Deduplicate Parquet Files Online
Remove duplicate rows from Parquet files instantly in your browser. No upload, no server — 100% private.
Filter Parquet Files Online
Filter rows in Parquet files by column value, directly in your browser. Your data stays on your device.
Convert Case in Parquet Files Online
Convert text columns to UPPERCASE, lowercase, or Title Case in Parquet files directly in your browser. Apply case conversion to any or all text columns at once — no upload required.
Find Fuzzy Duplicates in CSV Files Online
Find near-duplicate rows in CSV files using Levenshtein edit distance or Jaro-Winkler similarity — all in your browser. Set your own threshold and download the matched pairs as CSV — no upload required.
Find Fuzzy Duplicates in Excel Files Online
Find near-duplicate rows in Excel files using Levenshtein edit distance or Jaro-Winkler similarity — all in your browser. Set your own threshold and download the matched pairs as CSV — no upload required.
Find Fuzzy Duplicates in JSON Files Online
Find near-duplicate rows in JSON files using Levenshtein edit distance or Jaro-Winkler similarity — all in your browser. Set your own threshold and download the matched pairs as CSV — no upload required.
Parquet Viewer Online
View and inspect Parquet files directly in your browser. Browse rows, check column names and data types — no upload required, your data stays on your device.
Convert Parquet to CSV Online
Convert Parquet files to CSV format directly in your browser. No upload required — your data never leaves your device.