SmartQueryTools

Convert CSV to Parquet Online

Convert CSV files to Parquet format directly in your browser. No upload required — your data never leaves your device.

About converting CSV to Parquet

Converting CSV to Parquet is a common first step in a data pipeline. Parquet stores data by column, so a query that reads three columns out of forty only reads those three. Compression usually shrinks a CSV file to 20–30% of its original size. Column names and types are stored in the Parquet file footer, so DuckDB, Spark, Athena, BigQuery and pandas can read the types directly without guessing.

Typical uses include archiving large CSV exports from databases or SaaS tools, loading data into S3 or GCS for Athena or BigQuery, and preparing datasets for pandas, polars or DuckDB. If you export the same CSV every week and load it into an analytics system, converting it to Parquet once saves storage and makes every later query faster.

CSV has no types, so they are inferred from the values: integers, decimals, booleans, dates and timestamps are detected automatically. A single stray value, such as "n/a" in a numeric column, makes the whole column text. Clean or cast those columns first if the types matter downstream.

Not sure which format you need? Read the CSV vs Parquet comparison.

Worked example

A small orders export with an ID, a customer name, a date, an amount (one missing) and a true/false flag, converted with the default settings.

Input (CSV)

order_id,customer,order_date,amount,shipped
1001,Acme Ltd,2026-03-02,249.5,true
1002,Brightside Co,2026-03-02,1200,false
1003,Acme Ltd,2026-03-05,89.99,true
1004,Northwind,2026-03-07,,false

Output (Parquet)

Parquet file (binary, columnar) — shown as a table with its schema

order_idcustomerorder_dateamountshipped
1001Acme Ltd2026-03-02249.5true
1002Brightside Co2026-03-021200false
1003Acme Ltd2026-03-0589.99true
1004Northwind2026-03-07NULLfalse

Schema: order_id BIGINT, customer VARCHAR, order_date DATE, amount DOUBLE, shipped BOOLEAN

What changes when you convert CSV to Parquet

  • Column names come from the CSV header row and are stored in the Parquet schema.
  • Whole numbers become BIGINT, decimals become DOUBLE, and true/false becomes BOOLEAN.
  • ISO dates (2026-03-02) become DATE, and ISO date-times become TIMESTAMP.
  • Empty fields become NULL, not empty strings. In the example, order 1004 has a NULL amount.
  • A column with any value that does not parse as a number or date stays VARCHAR.

Your file is processed locally in your browser and is never uploaded. The free limit is 50 MB per file; larger files work if your device has the memory for them.

Frequently Asked Questions

How much smaller is Parquet than CSV?

Typically 3–5 times smaller, and more when columns have lots of repeated values such as country codes or status flags. Dictionary encoding stores each distinct value once per column chunk.

Can I control the column types when converting CSV to Parquet?

The converter infers types automatically. To force a type, for example keeping ZIP codes as text so leading zeros survive, run Cast Column Types on the CSV first or use the SQL Query tool with an explicit CAST.

Why did my numeric column come out as text in Parquet?

At least one value in that column is not a number, such as "N/A", "-" or a thousands separator like "1,200". Replace or clear those values, then convert again.

What is Parquet format?

Parquet is an open-source columnar storage format designed for efficient analytics. It compresses far better than CSV and is natively supported by Spark, Athena, BigQuery, Pandas, and DuckDB.

Related Tools