SmartQueryTools

Repair CSV Files Online

Repair malformed CSV files directly in your browser. Fix unbalanced quotes, inconsistent column counts, BOM markers, and mixed line endings — no upload required.

How to repair CSV files

  1. Drop your CSV file. Its raw bytes are read and five checks run straight away: BOM marker, encoding, mixed line endings, inconsistent column count and unbalanced quotes.
  2. Read the results. Each check has a green, amber or red dot and a short detail, such as how many short and long rows were found in the first 100 rows.
  3. If any check fails, click Repair File. If every check passes, the tool reports that the file looks clean and there is nothing to repair.
  4. Compare the Before and After panels, which show the first 20 lines of each version.
  5. Click Download Repaired File. The file keeps its name with -repaired added before the extension.

Your file is processed locally in your browser and is never uploaded. The free limit is 50 MB per file; larger files work if your device has the memory for them.

Worked example

A water-testing lab exports sample results from an older instrument. When a test was not run, the export drops the trailing fields instead of leaving them empty. Some lines end in CRLF and others in LF, and the lab's database import rejects the file.

Input (CSV)

sample_id,site,ph,nitrate_mg_l
WS-0141,Kaituna Weir,7.2,1.8
WS-0142,Mill Creek,6.9
WS-0143,Kaituna Weir,7.4,2.1
WS-0144,Rowan Bridge

Settings

  • Mixed line endings: found CRLF and LF
  • Inconsistent column count: header has 4 columns, 2 short rows and 0 long rows in the first 5 rows
  • Unbalanced quotes: none found
  • Action: Repair File

Result

sample_idsitephnitrate_mg_l
WS-0141Kaituna Weir7.21.8
WS-0142Mill Creek6.9NULL
WS-0143Kaituna Weir7.42.1
WS-0144Rowan BridgeNULLNULL

Every line now ends in LF. The short rows are padded with commas up to the header's four fields, so the WS-0142 line becomes "WS-0142,Mill Creek,6.9," and WS-0144 gets two trailing commas. Any CSV reader loads those empty fields as missing values, as shown above. No row is removed and no existing value is changed.

Working with CSV files

The repair works on the raw text of the CSV rather than on a parsed table, which is why it can fix files that other tools refuse to open. All line endings become LF. A byte order mark at the start of the file is removed. A quoted field that is never closed is closed at the end of the line where it started. Every row with fewer fields than the header is padded with empty fields up to the header count. Rows with more fields than the header are only reported. They are not changed, because the tool cannot know which field is the extra one. Most long rows come from an unquoted comma inside a value, such as 1,200 in a price column.

Quotes are tracked across lines, as the CSV standard allows. A quoted field with a line break inside, such as a two-line address, is one valid record and is left alone. Only a quoted field that runs to the end of the file without closing counts as unbalanced. A quote in the middle of a value, such as 12" pipe, is treated as a plain character. The file is read as UTF-8. If the bytes are not valid UTF-8, the tool reads it as Windows-1252 (Latin-1) instead, says so, and saves the repaired file as UTF-8, so characters such as é survive. Detection samples the first 100 rows, but the repair runs on every row.

Frequently Asked Questions

Does Repair File fix rows with too many columns?

No. Long rows are counted and reported, before and after the repair, but they are not trimmed or merged. They usually come from an unquoted comma inside a value. Fix those lines at the source, or quote the affected field by hand.

Where does the tool add the missing quote?

At the end of the line where the unclosed quoted field starts. Valid quoted fields that span several lines are not touched, and quotes in the middle of a value, such as 12" pipe, are not treated as field quotes. Check the After panel to confirm the fix.

Does Repair File delete any rows?

No. Every line is kept, including blank lines. Only the BOM, the encoding, line endings, unclosed quotes and short rows are changed.

Why does the tool say my file is clean when another program cannot read it?

The column count check covers the first 100 rows, and the Repair File button only appears when a check fails. Problems further down the file and a semicolon or other delimiter are not detected. Try Change Delimiter for semicolon files, and Convert Encoding if the file uses an encoding other than UTF-8 or Windows-1252.

What does the repaired file look like on disk?

It is UTF-8 text with no byte order mark and LF line endings, the Unix convention. Excel, Python, R, databases and the other tools on this site all read LF files. Values and delimiters are unchanged apart from the padding and any closing quotes.

Related Tools