Compare Schema of Parquet Files Online
Compare the schemas of two Parquet files directly in your browser. See which columns exist in each file and spot type mismatches — no upload required.
How to compare Schema of Parquet files
- Drop the first file onto the File A area and the second onto File B. Both must be in the format this page is for.
- The comparison runs as soon as both files are loaded. There are no settings to choose.
- Read the four counts at the top: Matched, Type mismatches, Only in A and Only in B.
- Go through the column-by-column list. Each column has a coloured dot and shows its type in one or both files.
- Drop a new file on either side to compare again. Nothing is downloaded, because the result is an on-screen report.
Your file is processed locally in your browser and is never uploaded. The free limit is 50 MB per file; larger files work if your device has the memory for them.
Worked example
An inventory team loads each monthly stock count into the same database table. April's export came from a new warehouse system and the load failed. File A is March's export, shown below.
Input (Parquet)
Parquet file (binary, columnar) — shown as a table with its schema
| sku | product_name | qty_on_hand | unit_cost | last_counted |
|---|---|---|---|---|
| HB-M8-40 | Hex bolt M8x40 | 420 | 0.12 | 2026-03-31 |
| WS-M8 | Washer M8 | 1150 | 0.03 | 2026-03-31 |
| NT-M8 | Nyloc nut M8 | 610 | 0.05 | 2026-03-30 |
Schema: sku VARCHAR, product_name VARCHAR, qty_on_hand BIGINT, unit_cost DOUBLE, last_counted DATE
Settings
- File B (April) has the columns sku, product_name, qty_on_hand, unit_cost, warehouse, count_date
- One qty_on_hand value in File B is "12 (est)", so that column loads as text
- No other settings: the diff runs once both files are loaded
Result
| column | file_a_type | file_b_type | status |
|---|---|---|---|
| sku | VARCHAR | VARCHAR | Matched |
| product_name | VARCHAR | VARCHAR | Matched |
| qty_on_hand | BIGINT | VARCHAR | Type mismatch |
| unit_cost | DOUBLE | DOUBLE | Matched |
| last_counted | DATE | NULL | Only in A |
| warehouse | NULL | VARCHAR | Only in B |
| count_date | NULL | DATE | Only in B |
The summary shows 3 matched, 1 type mismatch, 1 only in A and 2 only in B. qty_on_hand is in both files, but one annotated value made the April column text. last_counted and count_date hold the same data under different names. Columns are matched by exact name, so the tool lists them as unrelated. Renaming count_date to last_counted would turn two differences into one match.
Working with Parquet files
Parquet files store their schema, so the diff shows what the writer chose rather than a guess. That makes it strict. An INT32 column from one job and an INT64 column from another appear as INTEGER and BIGINT and are flagged. Decimal precision and scale are part of the type, so DECIMAL(10,2) and DECIMAL(12,2) do not match. A timestamp with a time zone and one without are also different types.
Struct columns are compared by their full type text, including every field inside them. Adding one field to a nested struct flags the whole struct as a mismatch, and the tool does not point to the field that changed. Compare the two type strings shown for that row to find it. Run this check before appending Parquet files into one dataset, since many readers reject files whose schemas differ. Both files are loaded in full, so very large files need enough memory in the browser.
Frequently Asked Questions
Why does my Parquet schema diff show INTEGER against BIGINT?
One file stores the column as 32-bit integers and the other as 64-bit. The values fit in both, but some engines refuse to read such files as one dataset. Cast one side with Cast Column Types so both match.
Can I see which field changed inside a Parquet struct?
Not directly. The row for that column shows the full STRUCT(...) type from each file. Compare the two strings to find the added, removed or retyped field.
Does schema diff compare the data in the rows?
No. It compares column names and types only. The row count of each file is shown under its drop area, but values are never compared. Use Compare Files to find rows that differ.
Are column names matched case-sensitively?
Yes. Email and email count as two different columns, one only in A and one only in B. Column position does not matter: a column is matched by name wherever it sits.
Why is INTEGER against BIGINT reported as a mismatch?
Types are compared by their exact names. INTEGER and BIGINT, or DECIMAL(10,2) and DECIMAL(18,2), count as mismatches even though the values fit together. A wider type in the new file is usually safe to load. Text where a number used to be usually is not.
Related Tools
Rename Columns in Parquet Files Online
Rename columns in Parquet files instantly in your browser. No upload, no server — your data stays on your device.
Cast Column Types in Parquet Files Online
Change column data types in Parquet files directly in your browser. Cast text to numbers, dates to timestamps, or any supported type conversion — no upload required.
Merge Parquet Files Online
Merge and concatenate multiple Parquet files into one, directly in your browser.
Compare Schema of CSV Files Online
Compare the schemas of two CSV files directly in your browser. See which columns exist in each file and spot type mismatches — no upload required.
Compare Schema of Excel Files Online
Compare the schemas of two Excel files directly in your browser. See which columns exist in each file and spot type mismatches — no upload required.
Compare Schema of JSON Files Online
Compare the schemas of two JSON files directly in your browser. See which columns exist in each file and spot type mismatches — no upload required.
Parquet Viewer Online
View and inspect Parquet files directly in your browser. Browse rows, check column names and data types — no upload required, your data stays on your device.
Convert Parquet to CSV Online
Convert Parquet files to CSV format directly in your browser. No upload required — your data never leaves your device.