SmartQueryTools

Manage Columns in Parquet Files Online

Drop or select specific columns from Parquet files directly in your browser. No upload required.

How to manage Columns in Parquet files

  1. Drop your file onto the upload area. Every column is listed as a chip with its detected type, and all are selected to start with.
  2. Click a chip to deselect a column you want to drop. It turns grey with a line through it. Click again to bring it back.
  3. Use Deselect all and then click the few columns you need when you only want to keep a handful.
  4. Click Apply to preview the result, then Download to save it in the same format you uploaded.

Your file is processed locally in your browser and is never uploaded. The free limit is 50 MB per file; larger files work if your device has the memory for them.

Worked example

An HR analyst needs to send a headcount list to a facilities team, without salary or national ID numbers.

Input (Parquet)

Parquet file (binary, columnar) — shown as a table with its schema

emp_idnamedeptsalarynational_id
2041Maria KeaneFinance68000KX 44 91 02 C
2042Dev PatelOperations54500LM 20 17 88 A
2043Sofia BrandtFinance71200PR 63 05 19 D

Schema: emp_id BIGINT, name VARCHAR, dept VARCHAR, salary BIGINT, national_id VARCHAR

Settings

  • Deselected: salary, national_id
  • Kept: emp_id, name, dept

Result

emp_idnamedept
2041Maria KeaneFinance
2042Dev PatelOperations
2043Sofia BrandtFinance

The two sensitive columns are removed entirely, not blanked, so they are absent from the downloaded file. The three kept columns stay in their original left-to-right order with their original types. Every row is kept, and no values change.

Working with Parquet files

Parquet is columnar, so each column is stored separately. Dropping unused columns shrinks the file roughly in proportion to the size of those columns, and every later query reads less. This is a common step before sharing a table externally, or before loading it into a tool that struggles with very wide schemas. It is also the simplest way to strip personal data before handing a file to a vendor, because dropped columns are absent from the output, not blanked.

Each chip shows the column's stored Parquet type, such as BIGINT, DECIMAL(18,2), TIMESTAMP or a STRUCT, so you can see what you are removing. A struct column is kept or dropped as a whole; you cannot drop just one field inside it here. The output is a new Parquet file whose schema contains only the kept columns, with their types unchanged. Row groups, compression and column statistics are regenerated as the new file is written.

Frequently Asked Questions

Will dropping columns make my Parquet file smaller?

Yes. Parquet stores each column separately, so removing a column removes its data from the file. Large text or nested columns give the biggest savings.

Can I remove one field from inside a Parquet struct column?

Not with this tool. Structs are kept or dropped as a whole. Use the SQL Query tool to rebuild the struct without that field.

Can I reorder columns with this tool?

No. Kept columns stay in their original left-to-right order. To change the order, use the SQL Query tool and list the columns in the order you want.

Does selecting columns change any values or types?

No. Kept columns are copied exactly, with the same types. Only the dropped columns are removed.

Can I drop every column?

No. At least one column must stay selected, and Apply is disabled until one is.

Related Tools