SmartQueryTools

Validate Parquet Files Online

Validate Parquet file structure in your browser. Check null counts, distinct values, and data types for every column — no upload required.

How to validate Parquet files

  1. Drop your file onto the upload area. The report starts as soon as the file has loaded. There are no settings.
  2. Read the four summary cards: total rows, columns, null-heavy columns and all-null columns.
  3. Scan the Column Report table for each column's detected type, null count, percentage of nulls and number of distinct values.
  4. Check the warnings under the table. Columns that are entirely null are flagged in red and columns that are more than half null in amber.

Your file is processed locally in your browser and is never uploaded. The free limit is 50 MB per file; larger files work if your device has the memory for them.

Worked example

A clinic exports a week of appointments before loading it into a reporting database. The analyst wants to know which fields are reliable before writing any queries.

Input (Parquet)

Parquet file (binary, columnar) — shown as a table with its schema

patient_refclinicappointment_dateno_shownotes
P-104Eastside2026-06-02falseNULL
P-221Eastside2026-06-02trueCalled to rebook
P-104Harbour2026-06-09falseNULL
P-318NULL2026-06-11falseNULL
P-450Harbour2026-06-11trueNULL

Schema: patient_ref VARCHAR, clinic VARCHAR, appointment_date DATE, no_show BOOLEAN, notes VARCHAR

Settings

  • No settings. The report is generated when the file loads.

Result

ColumnTypeNulls% NullDistinct
patient_refVARCHAR00.0%4
clinicVARCHAR120.0%2
appointment_dateDATE00.0%3
no_showBOOLEAN00.0%2
notesVARCHAR480.0%1

The summary cards read 5 rows, 5 columns, 1 null-heavy column and 0 all-null columns. notes is flagged in amber at 80% null. patient_ref has 4 distinct values in 5 rows, which shows P-104 appears twice, so it cannot be used as a unique key. Distinct counts ignore nulls, which is why clinic shows 2 rather than 3.

Working with Parquet files

Parquet stores its schema in the file, so the Type column shows the types the writer declared rather than a guess. You will see types such as DECIMAL(12,2), TIMESTAMP WITH TIME ZONE, or STRUCT(...) with the nested field list. Comparing this against the schema a downstream table expects is a quick way to catch a column written as DOUBLE where DECIMAL was intended.

Parquet footers already hold null counts per row group, but this report does not rely on them. It scans every value and counts nulls and distinct values exactly, so the numbers are correct even when a writer skipped statistics. Empty strings are values, not nulls, in Parquet. A column full of "" is 0% null here, which often points to an upstream job that wrote blanks instead of nulls. A column whose Distinct count equals Total rows, with 0% nulls, is a candidate primary key. A low distinct count on a column meant to be an ID usually means the writer repeated or truncated values.

Frequently Asked Questions

Does the Parquet validation read the footer statistics?

No. It scans the data and counts nulls and distinct values directly, so the report is accurate even if the file has no statistics or wrong ones.

How are struct columns shown in the Parquet report?

As one row with the full STRUCT type, including the nested field names. The distinct count compares whole struct values. To check nested fields one by one, select them as separate columns in the SQL Query tool.

Does Validate change or download my file?

No. It only produces the on-screen report. Nothing is modified and there is no download.

What counts as a null-heavy column?

Any column where more than 50% of values are null. The Null-heavy card counts all-null columns too, so a column that is 100% null is included in both cards.

Do distinct counts include nulls?

No. Distinct counts only non-null values. A column with values A, B and some nulls shows 2.

Related Tools