SmartQueryTools

Filter Parquet Files Online

Filter rows in Parquet files by column value, directly in your browser. Your data stays on your device.

How to filter Parquet files

  1. Drop your file onto the upload area. It is loaded into the in-browser engine and the first 200 rows are shown.
  2. Pick the column to test from the Column list. It defaults to the first column in the file.
  3. Choose an operator: =, !=, >, <, >=, <=, LIKE, IS NULL or IS NOT NULL. Type the comparison value, unless you picked IS NULL or IS NOT NULL.
  4. Click Apply Filter. The tool shows how many rows match and previews the first 200 of them.
  5. Click Download to save the matching rows in the same format you uploaded.

Your file is processed locally in your browser and is never uploaded. The free limit is 50 MB per file; larger files work if your device has the memory for them.

Worked example

A bike-share operator wants to review only the longer trips from one morning's log, meaning rides of 30 minutes or more.

Input (Parquet)

Parquet file (binary, columnar) — shown as a table with its schema

trip_idstart_stationduration_minmember
501Harbour St12true
502Park Ave47false
503Harbour St31true
504Mill Rd8false
505Park Ave30true

Schema: trip_id BIGINT, start_station VARCHAR, duration_min BIGINT, member BOOLEAN

Settings

  • Column: duration_min
  • Operator: >=
  • Value: 30

Result

trip_idstart_stationduration_minmember
502Park Ave47false
503Harbour St31true
505Park Ave30true

duration_min is an integer column, so the value 30 is compared as a number, not as text. Three of five rows match. Trip 505 is kept because >= includes the boundary; with > it would be dropped. Matching rows keep their original order and every column is carried through unchanged.

Working with Parquet files

Parquet stores a type for every column, so filters behave predictably. Numbers compare as numbers, and a DATE or TIMESTAMP column accepts a value such as 2026-03-01, which is converted to a date before comparing. On a timestamp column that value means midnight. So > 2026-03-01 keeps events from 00:00:00.000001 onward and >= 2026-03-01 includes midnight itself. Decimal columns compare exactly, so = 19.99 matches a DECIMAL(10,2) value of 19.99 without floating-point surprises.

Parquet keeps NULL and the empty string apart. IS NULL finds only true nulls, and a string column holding "" needs = with an empty value instead. Struct and list columns appear in the column list but cannot be filtered by an inner field. Use the SQL Query tool for conditions on nested values. The downloaded Parquet file keeps the original schema, only with fewer rows. Filtering before sharing is a cheap way to cut a large Parquet file down to the rows a colleague actually needs.

Frequently Asked Questions

Can I filter a Parquet file on a nested field such as address.country?

Not directly. The picker lists top-level columns only. Use the SQL Query tool with a condition such as WHERE address.country = 'NZ', or pull the field into its own column first and then filter on it.

Does filtering change the Parquet column types?

No. The filtered file is written with the same column names and types. Only the row count changes.

Can I combine several conditions, such as region = "EU" and amount > 100?

Not in one pass. The tool applies a single condition. Download the result and filter it again for the next condition, which gives the same result as AND. For OR logic or several conditions at once, use the SQL Query tool.

How does the LIKE operator work?

LIKE matches a pattern. % stands for any run of characters and _ for exactly one, so Park% matches every value starting with "Park". Without a wildcard, LIKE behaves like =. The match is case-sensitive.

Does filtering change the original file?

No. Your file is read in the browser and never modified. The matching rows are written to a new file that you download.

Related Tools