SmartQueryTools

Deduplicate JSON Files Online

Remove duplicate rows from JSON files instantly in your browser. No upload, no server — 100% private.

How to deduplicate JSON files

  1. Drop your file onto the upload area. It is loaded into the in-browser engine and the first 200 rows are shown.
  2. Choose the columns that define a duplicate. All columns are ticked by default, which removes only rows that are identical in every field.
  3. Untick columns to match on a subset. For example, keep only customer_id ticked to keep one row per customer.
  4. Click Deduplicate. The tool reports how many rows were removed and previews the result.
  5. Download the cleaned file in the same format you uploaded.

Your file is processed locally in your browser and is never uploaded. The free limit is 50 MB per file; larger files work if your device has the memory for them.

Worked example

A CRM export lists the same contact more than once because the sync job ran twice, and one person also signed up again later with a different date.

Input (JSON)

[
  {
    "email": "ana@example.com",
    "name": "Ana Ruiz",
    "plan": "pro",
    "signup_date": "2026-01-04"
  },
  {
    "email": "ben@example.com",
    "name": "Ben Cho",
    "plan": "free",
    "signup_date": "2026-01-09"
  },
  {
    "email": "ana@example.com",
    "name": "Ana Ruiz",
    "plan": "pro",
    "signup_date": "2026-01-04"
  },
  {
    "email": "cara@example.com",
    "name": "Cara Doyle",
    "plan": "free",
    "signup_date": "2026-02-11"
  },
  {
    "email": "ben@example.com",
    "name": "Ben Cho",
    "plan": "pro",
    "signup_date": "2026-03-20"
  }
]

Settings

  • Columns to match on: email only (name, plan and signup_date unticked)

Result

emailnameplansignup_date
ana@example.comAna Ruizpro2026-01-04
ben@example.comBen Chofree2026-01-09
cara@example.comCara Doylefree2026-02-11

Matching on email alone collapses five rows to three: one row per address. With all columns ticked, only the exact repeat of Ana's row would be removed. Ben's two rows would both stay because his plan and signup date differ. Which row survives for each email is not guaranteed, so sort first if you need the earliest or latest one.

Working with JSON files

A JSON array of objects is loaded as a table. Each key becomes a column, and objects with missing keys get NULL in that column. Key order inside an object does not matter: {"a":1,"b":2} and {"b":2,"a":1} are the same row.

Nested objects become struct columns and are compared by their full contents. Matching on email when one record has a different address.city still treats the two as duplicates. The result is written back as a JSON array. Numbers stay numbers and booleans stay booleans.

Mixed value types in one key change the picture. If "id" is 1 in one record and "1" in another, the key loads as a raw JSON column and those two values are different, so both rows survive. Very large IDs above 9,007,199,254,740,991 are written back as strings to keep every digit. Records that were missing a key come back with that key set to null, so the output can be slightly larger than the input even after rows are removed.

Frequently Asked Questions

How are duplicate JSON objects with different key order handled?

Key order does not matter. Objects are compared field by field after loading, so {"id":1,"name":"A"} and {"name":"A","id":1} are duplicates.

Can I deduplicate JSON on a nested field such as user.id?

Not directly. The column picker works on top-level keys. Flatten the JSON first so user.id becomes its own column, then match on it.

Which row is kept when duplicates are found?

One row per unique combination of the selected columns is kept. When you match on a subset of columns, which of the matching rows survives is not guaranteed. Sort the file first if you need the earliest or latest one, or use Top N per Group with N = 1.

What is the difference between Remove Duplicates and Find Duplicates?

Remove Duplicates writes a cleaned file with the extra rows dropped. Find Duplicates lists which rows are duplicated and how many times each one appears, so you can review them before deleting anything.

Can it catch near-duplicates such as typos in names?

No. Matching is exact. Use Fuzzy Deduplicate for names or addresses that differ slightly, such as "Acme Ltd" and "ACME Limited".

Related Tools