SmartQueryTools

Count Values in Parquet Files Online

Group and count rows by any column in Parquet files directly in your browser. Sort by frequency or value to find the most common entries — no upload required.

How to count Values in Parquet files

  1. Drop your file onto the upload area. The row count is shown and the first column is pre-selected.
  2. Choose a column in the Group by column list.
  3. Pick a sort order: Count (high → low), Count (low → high) or Value (A → Z).
  4. Click Count. The result has two columns, value and count, and the heading shows how many distinct values were found.
  5. Click Download to save the counts in the same format you uploaded, with _countby added to the file name.

Your file is processed locally in your browser and is never uploaded. The free limit is 50 MB per file; larger files work if your device has the memory for them.

Worked example

A city bike-share scheme wants to know which docking stations trips start from most often. A few trips from a faulty dock have no start station recorded.

Input (Parquet)

Parquet file (binary, columnar) — shown as a table with its schema

trip_idstart_stationduration_minrider_type
5001Quay St12member
5002Park Rd25casual
5003Quay St8member
5004NULL17casual
5005Quay St31casual
5006Park Rd9member

Schema: trip_id BIGINT, start_station VARCHAR, duration_min BIGINT, rider_type VARCHAR

Settings

  • Group by column: start_station
  • Sort by: Count (high → low)

Result

valuecount
Quay St3
Park Rd2
NULL1

Six trips collapse to three groups. Quay St has three trips, Park Rd two, and the trip with no station forms its own group with a null value. Missing values are counted, not dropped, so the counts always add up to the total row count. The output columns are always named value and count, whatever the source column was called.

Working with Parquet files

The value column keeps the type of the source column, and count is a 64-bit integer. Counting by an integer or date column in Parquet gives a correctly typed lookup table that can be joined back to the original file without casting.

Be careful with timestamp and float columns. A timestamp stored to the microsecond is almost unique per row, so counting by it returns nearly as many groups as rows. Truncate it to the day or hour with Date Truncate first. Float columns have the same problem with tiny rounding differences. Round Numbers helps there. Counting by a struct column groups on the whole nested value, which is rarely what you want. The result downloads as a two-column Parquet file. Counting a low-cardinality column, such as a status or country code, is also a quick check that a Parquet export contains every value you expect and no unexpected ones.

Frequently Asked Questions

How do I count Parquet rows per day from a timestamp column?

Run Date Truncate on the timestamp to the day level, download the result, then count by that column. Counting the raw timestamp gives almost one group per row.

What type is the count column in the Parquet output?

A 64-bit integer (BIGINT). The value column keeps the type of the column you grouped by.

Can I count by more than one column?

No. The tool groups by one column. For combinations of columns, use the Aggregate tool or the SQL Query tool with GROUP BY on several columns.

How are ties ordered when sorting by count?

Values with the same count have no fixed order between them. Choose Value (A → Z) if you need a stable, alphabetical order.

Does the on-screen table show every value?

The on-screen table shows the first 500 groups. The heading still gives the full number of distinct values, and the downloaded file contains every group.

Related Tools