Convert Arrow to Parquet Online
Convert Arrow files to Parquet format directly in your browser. No upload required — your data never leaves your device.
About converting Arrow to Parquet
Arrow to Parquet moves a table from a format designed for speed in memory to one designed for storage. An Arrow IPC file is large because its buffers are saved exactly as they sit in RAM. Parquet encodes and compresses each column, so the same table usually shrinks several times over. That matters when a Feather cache or Polars scratch file turns out to be worth keeping, or needs to go to S3, a lakehouse table, BigQuery or Athena. None of those query Arrow IPC files directly.
The two formats are both columnar and share most of their type system, so structure carries across well. Nested structs and lists become Parquet nested groups, decimals keep their precision and scale, and timestamps stay timestamps. The output uses Snappy compression, which every Parquet reader supports.
What does not carry across is metadata. The pandas block in the Arrow schema, which records the index and any category dtypes, is not copied, so pandas.read_parquet gives back plain string columns where you had categories. Field-level metadata and Arrow extension type names are dropped too. If a downstream job depends on them, write Parquet from pyarrow directly with pyarrow.parquet.write_table.
Not sure which format you need? Read the Apache Arrow vs Parquet comparison.
Worked example
A small orders export with an ID, a customer name, a date, an amount (one missing) and a true/false flag, converted with the default settings.
Input (Arrow)
Arrow IPC file (binary, columnar) — shown as a table with its schema
| order_id | customer | order_date | amount | shipped |
|---|---|---|---|---|
| 1001 | Acme Ltd | 2026-03-02 | 249.5 | true |
| 1002 | Brightside Co | 2026-03-02 | 1200 | false |
| 1003 | Acme Ltd | 2026-03-05 | 89.99 | true |
| 1004 | Northwind | 2026-03-07 | NULL | false |
Schema: order_id BIGINT, customer VARCHAR, order_date DATE, amount DOUBLE, shipped BOOLEAN
Output (Parquet)
Parquet file (binary, columnar) — shown as a table with its schema
| order_id | customer | order_date | amount | shipped |
|---|---|---|---|---|
| 1001 | Acme Ltd | 2026-03-02 | 249.5 | true |
| 1002 | Brightside Co | 2026-03-02 | 1200 | false |
| 1003 | Acme Ltd | 2026-03-05 | 89.99 | true |
| 1004 | Northwind | 2026-03-07 | NULL | false |
Schema: order_id BIGINT, customer VARCHAR, order_date DATE, amount DOUBLE, shipped BOOLEAN
What changes when you convert Arrow to Parquet
- The Parquet file is Snappy-compressed and split into row groups of up to 122,880 rows, whatever record batch sizes the Arrow file used.
- int32 and int64 become INT32 and INT64, float64 becomes DOUBLE, bool becomes BOOLEAN, and utf8 or large_utf8 become UTF-8 strings.
- Dictionary-encoded columns are stored as plain strings in the schema. Parquet dictionary-encodes them again on disk, so the file stays small, but readers get strings rather than categories.
- Timezone-aware timestamps are stored as UTC instants. The zone name, such as Europe/Dublin, is not kept.
- Struct, list and map columns become nested Parquet columns and read back with the same shape in Spark, pandas, Polars and BigQuery.
Your file is processed locally in your browser and is never uploaded. The free limit is 50 MB per file; larger files work if your device has the memory for them.
Frequently Asked Questions
How much smaller will the Parquet file be than my Arrow file?
Often three to ten times smaller, depending on the data. Columns with many repeated values, sorted IDs or low-cardinality strings shrink the most. Random floats, such as embeddings, shrink the least.
Why are my pandas categories plain strings after reading the Parquet back?
The pandas metadata in the Arrow schema is not carried into the Parquet file. Restore them after reading with df["site"] = df["site"].astype("category"), or write the Parquet from pandas with df.to_parquet if you need them preserved.
Can I get Zstandard compression instead of Snappy?
Not in this converter. Write the file with pyarrow.parquet.write_table(table, "out.parquet", compression="zstd"), or with COPY tbl TO 'out.parquet' (FORMAT parquet, COMPRESSION zstd) in the DuckDB command-line tool.
What is Parquet format?
Parquet is an open-source columnar storage format designed for efficient analytics. It compresses far better than CSV and is natively supported by Spark, Athena, BigQuery, Pandas, and DuckDB.
Related Tools
Convert Parquet to Arrow Online
Convert Parquet files to Arrow format directly in your browser. No upload required — your data never leaves your device.
Parquet Viewer Online
View and inspect Parquet files directly in your browser. Browse rows, check column names and data types — no upload required, your data stays on your device.
Arrow Viewer Online
View and inspect Arrow files directly in your browser. Browse rows, check column names and data types — no upload required, your data stays on your device.
Filter Parquet Files Online
Filter rows in Parquet files by column value, directly in your browser. Your data stays on your device.
Convert Arrow to CSV Online
Convert Arrow files to CSV format directly in your browser. No upload required — your data never leaves your device.
Convert Arrow to Excel Online
Convert Arrow files to Excel format directly in your browser. No upload required — your data never leaves your device.
Convert Arrow to JSON Online
Convert Arrow files to JSON format directly in your browser. No upload required — your data never leaves your device.