SmartQueryTools

Find Fuzzy Duplicates in CSV Files Online

Find near-duplicate rows in CSV files using Levenshtein edit distance or Jaro-Winkler similarity — all in your browser. Set your own threshold and download the matched pairs as CSV — no upload required.

How to find Fuzzy Duplicates in CSV files

  1. Drop your file onto the upload area. Files with more than 5,000 rows show a notice, because only the first 5,000 rows are compared.
  2. Choose the column to check for near-duplicates, usually a name, company or address column. The first column is selected by default.
  3. Pick a similarity method. Levenshtein counts character edits, with a maximum distance of 1 to 4 (default 2). Jaro-Winkler gives a 0 to 1 score, with a minimum similarity slider from 0.70 to 0.99 (default 0.85).
  4. Click Find Fuzzy Duplicates. The tool lists matching pairs of rows with their values and score.
  5. Review the pairs and click Download Pairs CSV. The file is a list of pairs to check, not a cleaned copy of your data.

Your file is processed locally in your browser and is never uploaded. The free limit is 50 MB per file; larger files work if your device has the memory for them.

Worked example

An accounts payable team merged supplier lists from two systems and suspects some vendors were entered twice with small spelling differences.

Input (CSV)

vendor_id,vendor_name,city
V-001,Northgate Plumbing,Leeds
V-002,Northgate Plumbing Ltd,Leeds
V-003,Brightwater Electrical,York
V-004,Brightwater Electric,York
V-005,Harlow & Sons,Hull
V-006,Northgate Plumbng,Leeds

Settings

  • Column: vendor_name
  • Method: Jaro-Winkler
  • Min similarity: 0.85

Result

row_arow_bval_aval_bsimilarity
16Northgate PlumbingNorthgate Plumbng0.9889
34Brightwater ElectricalBrightwater Electric0.9818
12Northgate PlumbingNorthgate Plumbing Ltd0.9636
26Northgate Plumbing LtdNorthgate Plumbng0.9545

row_a and row_b are 1-based row positions in the file, and the most similar pairs come first. The three Northgate rows produce three pairs, because every pair is scored separately. Harlow & Sons matches nothing. With Levenshtein at distance 2 instead, only the typo (distance 1) and Electrical/Electric (distance 2) would be listed. Adding " Ltd" is 4 edits.

Working with CSV files

CSV values reach the comparison exactly as they sit in the file, including leading and trailing spaces and letter case. Matching is case-sensitive, so "ACME LTD" and "Acme Ltd" are 5 edits apart and Jaro-Winkler scores them only 0.58. Run Trim Whitespace and Convert Case on the name column first, then run fuzzy matching on the cleaned file. This one step often finds more real duplicates than any threshold change.

Empty fields are skipped, so blank names never pair with each other. Numeric columns can be checked too. The values are compared as text, which makes Levenshtein distance 1 useful for spotting phone numbers or account codes with a single mistyped digit. Remember that a CSV column of codes like 00417 may have been read as the number 417, which changes what is compared. The pairs download is always CSV, whatever the input. Its val_a and val_b fields are quoted when a name contains a comma, such as "Smith, Jones & Co".

Frequently Asked Questions

Why does my CSV give no pairs for names that only differ in capitals?

Comparison is case-sensitive. Run Convert Case on the column first so capitals match, then check again.

Can I fuzzy-match a numeric ID column in a CSV?

Yes. Values are compared as text. Levenshtein at distance 1 finds IDs that differ by one digit. Leading zeros may have been dropped if the column was read as a number.

Does the tool remove the near-duplicates for me?

No. It lists pairs of similar rows with their row numbers and score. You decide which rows to keep, then remove the others yourself or with Filter.

Are exact duplicates included in the pairs?

No. Levenshtein pairs must be at least 1 edit apart, and Jaro-Winkler pairs must not be identical. Use Find Duplicates or Remove Duplicates for exact matches.

Which method should I choose?

Levenshtein suits short values and typos, such as codes or single words. Jaro-Winkler suits names and company names, because it rewards a shared start. For example "Robert" and "Robrt" score about 0.96.

Related Tools

Deduplicate CSV Files Online

Remove duplicate rows from CSV files instantly in your browser. No upload, no server — 100% private.

Filter CSV Files Online

Filter rows in CSV files by column value, directly in your browser. Your data stays on your device.

Convert Case in CSV Files Online

Convert text columns to UPPERCASE, lowercase, or Title Case in CSV files directly in your browser. Apply case conversion to any or all text columns at once — no upload required.

Find Fuzzy Duplicates in Excel Files Online

Find near-duplicate rows in Excel files using Levenshtein edit distance or Jaro-Winkler similarity — all in your browser. Set your own threshold and download the matched pairs as CSV — no upload required.

Find Fuzzy Duplicates in Parquet Files Online

Find near-duplicate rows in Parquet files using Levenshtein edit distance or Jaro-Winkler similarity — all in your browser. Set your own threshold and download the matched pairs as CSV — no upload required.

Find Fuzzy Duplicates in JSON Files Online

Find near-duplicate rows in JSON files using Levenshtein edit distance or Jaro-Winkler similarity — all in your browser. Set your own threshold and download the matched pairs as CSV — no upload required.

CSV Viewer Online

View and inspect CSV files directly in your browser. Browse rows, check column names and data types — no upload required, your data stays on your device.

Convert CSV to Parquet Online

Convert CSV files to Parquet format directly in your browser. No upload required — your data never leaves your device.