SmartQueryTools

Deduplicate YAML Files Online

Remove duplicate rows from YAML files instantly in your browser. No upload, no server — 100% private.

How to deduplicate YAML files

  1. Drop your file onto the upload area. It is loaded into the in-browser engine and the first 200 rows are shown.
  2. Choose the columns that define a duplicate. All columns are ticked by default, which removes only rows that are identical in every field.
  3. Untick columns to match on a subset. For example, keep only customer_id ticked to keep one row per customer.
  4. Click Deduplicate. The tool reports how many rows were removed and previews the result.
  5. Download the cleaned file in the same format you uploaded.

Your file is processed locally in your browser and is never uploaded. The free limit is 50 MB per file; larger files work if your device has the memory for them.

Worked example

A CRM export lists the same contact more than once because the sync job ran twice, and one person also signed up again later with a different date.

Input (YAML)

- email: ana@example.com
  name: Ana Ruiz
  plan: pro
  signup_date: '2026-01-04'
- email: ben@example.com
  name: Ben Cho
  plan: free
  signup_date: '2026-01-09'
- email: ana@example.com
  name: Ana Ruiz
  plan: pro
  signup_date: '2026-01-04'
- email: cara@example.com
  name: Cara Doyle
  plan: free
  signup_date: '2026-02-11'
- email: ben@example.com
  name: Ben Cho
  plan: pro
  signup_date: '2026-03-20'

Settings

  • Columns to match on: email only (name, plan and signup_date unticked)

Result

emailnameplansignup_date
ana@example.comAna Ruizpro2026-01-04
ben@example.comBen Chofree2026-01-09
cara@example.comCara Doylefree2026-02-11

Matching on email alone collapses five rows to three: one row per address. With all columns ticked, only the exact repeat of Ana's row would be removed. Ben's two rows would both stay because his plan and signup date differ. Which row survives for each email is not guaranteed, so sort first if you need the earliest or latest one.

Frequently Asked Questions

Which row is kept when duplicates are found?

One row per unique combination of the selected columns is kept. When you match on a subset of columns, which of the matching rows survives is not guaranteed. Sort the file first if you need the earliest or latest one, or use Top N per Group with N = 1.

What is the difference between Remove Duplicates and Find Duplicates?

Remove Duplicates writes a cleaned file with the extra rows dropped. Find Duplicates lists which rows are duplicated and how many times each one appears, so you can review them before deleting anything.

Can it catch near-duplicates such as typos in names?

No. Matching is exact. Use Fuzzy Deduplicate for names or addresses that differ slightly, such as "Acme Ltd" and "ACME Limited".

Related Tools