Find Fuzzy Duplicates in JSON Files Online
Find near-duplicate rows in JSON files using Levenshtein edit distance or Jaro-Winkler similarity — all in your browser. Set your own threshold and download the matched pairs as CSV — no upload required.
How to find Fuzzy Duplicates in JSON files
- Drop your file onto the upload area. Files with more than 5,000 rows show a notice, because only the first 5,000 rows are compared.
- Choose the column to check for near-duplicates, usually a name, company or address column. The first column is selected by default.
- Pick a similarity method. Levenshtein counts character edits, with a maximum distance of 1 to 4 (default 2). Jaro-Winkler gives a 0 to 1 score, with a minimum similarity slider from 0.70 to 0.99 (default 0.85).
- Click Find Fuzzy Duplicates. The tool lists matching pairs of rows with their values and score.
- Review the pairs and click Download Pairs CSV. The file is a list of pairs to check, not a cleaned copy of your data.
Your file is processed locally in your browser and is never uploaded. The free limit is 50 MB per file; larger files work if your device has the memory for them.
Worked example
An accounts payable team merged supplier lists from two systems and suspects some vendors were entered twice with small spelling differences.
Input (JSON)
[
{
"vendor_id": "V-001",
"vendor_name": "Northgate Plumbing",
"city": "Leeds"
},
{
"vendor_id": "V-002",
"vendor_name": "Northgate Plumbing Ltd",
"city": "Leeds"
},
{
"vendor_id": "V-003",
"vendor_name": "Brightwater Electrical",
"city": "York"
},
{
"vendor_id": "V-004",
"vendor_name": "Brightwater Electric",
"city": "York"
},
{
"vendor_id": "V-005",
"vendor_name": "Harlow & Sons",
"city": "Hull"
},
{
"vendor_id": "V-006",
"vendor_name": "Northgate Plumbng",
"city": "Leeds"
}
]Settings
- Column: vendor_name
- Method: Jaro-Winkler
- Min similarity: 0.85
Result
| row_a | row_b | val_a | val_b | similarity |
|---|---|---|---|---|
| 1 | 6 | Northgate Plumbing | Northgate Plumbng | 0.9889 |
| 3 | 4 | Brightwater Electrical | Brightwater Electric | 0.9818 |
| 1 | 2 | Northgate Plumbing | Northgate Plumbing Ltd | 0.9636 |
| 2 | 6 | Northgate Plumbing Ltd | Northgate Plumbng | 0.9545 |
row_a and row_b are 1-based row positions in the file, and the most similar pairs come first. The three Northgate rows produce three pairs, because every pair is scored separately. Harlow & Sons matches nothing. With Levenshtein at distance 2 instead, only the typo (distance 1) and Electrical/Electric (distance 2) would be listed. Adding " Ltd" is 4 edits.
Working with JSON files
The column picker works on top-level keys. A name stored inside a nested object, such as customer.name, cannot be selected directly. Flatten the JSON first. If you choose a key that holds an object or an array, it is converted to text before comparing, so two records can match on how their nested structure prints rather than on a real field. Number keys are compared as text too, so 1024 and 1042 are two edits apart under Levenshtein.
Objects missing the chosen key are skipped, as are empty strings. Row numbers follow the order of objects in the array, starting at 1, which makes it easy to find a record by index in code: array index = row number minus 1. The output is not written as JSON. It is a CSV of pairs with row_a, row_b, val_a, val_b and either edit_distance or similarity. Load that CSV in your script to look up both records by index and merge them.
Frequently Asked Questions
How do row_a and row_b map to my JSON array?
They are 1-based positions in the array. Subtract 1 to get the zero-based index used in most programming languages.
Can I match on a nested JSON field?
Flatten the file so the nested field becomes a top-level column, then select it.
Does the tool remove the near-duplicates for me?
No. It lists pairs of similar rows with their row numbers and score. You decide which rows to keep, then remove the others yourself or with Filter.
Are exact duplicates included in the pairs?
No. Levenshtein pairs must be at least 1 edit apart, and Jaro-Winkler pairs must not be identical. Use Find Duplicates or Remove Duplicates for exact matches.
Which method should I choose?
Levenshtein suits short values and typos, such as codes or single words. Jaro-Winkler suits names and company names, because it rewards a shared start. For example "Robert" and "Robrt" score about 0.96.
Related Tools
Deduplicate JSON Files Online
Remove duplicate rows from JSON files instantly in your browser. No upload, no server — 100% private.
Filter JSON Files Online
Filter rows in JSON files by column value, directly in your browser. Your data stays on your device.
Convert Case in JSON Files Online
Convert text columns to UPPERCASE, lowercase, or Title Case in JSON files directly in your browser. Apply case conversion to any or all text columns at once — no upload required.
Find Fuzzy Duplicates in CSV Files Online
Find near-duplicate rows in CSV files using Levenshtein edit distance or Jaro-Winkler similarity — all in your browser. Set your own threshold and download the matched pairs as CSV — no upload required.
Find Fuzzy Duplicates in Excel Files Online
Find near-duplicate rows in Excel files using Levenshtein edit distance or Jaro-Winkler similarity — all in your browser. Set your own threshold and download the matched pairs as CSV — no upload required.
Find Fuzzy Duplicates in Parquet Files Online
Find near-duplicate rows in Parquet files using Levenshtein edit distance or Jaro-Winkler similarity — all in your browser. Set your own threshold and download the matched pairs as CSV — no upload required.
JSON Viewer Online
View and inspect JSON files directly in your browser. Browse rows, check column names and data types — no upload required, your data stays on your device.
Convert JSON to Parquet Online
Convert JSON files to Parquet format directly in your browser. No upload required — your data never leaves your device.