SmartQueryTools

Convert CSV to Arrow Online

Convert CSV files to Arrow format directly in your browser. No upload required — your data never leaves your device.

About converting CSV to Arrow

Arrow IPC is the file form of the in-memory format used by pyarrow, Polars, Spark and many query engines. Converting CSV to Arrow does the parsing and type detection once. Every later load is close to a memory copy: no delimiter scanning, no number parsing and no guessing of types.

That makes the pair useful for anyone who reloads the same CSV many times in a notebook, for handing a typed dataset from a Python process to a Rust or JavaScript one, and for feeding Arrow-native browser tools such as Perspective or Arquero. The output is an Arrow IPC file (the Feather version 2 format), readable with pyarrow.ipc.open_file, pandas.read_feather, polars.read_ipc and the apache-arrow JavaScript library.

Arrow is not a storage format. The file is written uncompressed, so it is often about the size of the CSV and sometimes larger, and much larger than the same data in Parquet. Keep Arrow for fast local reloads and exchange between processes; use Parquet for archiving or cloud storage. Before downloading, check the column types listed after the file loads. A column you expected as BIGINT that shows VARCHAR has at least one value that is not a number.

Worked example

A small orders export with an ID, a customer name, a date, an amount (one missing) and a true/false flag, converted with the default settings.

Input (CSV)

order_id,customer,order_date,amount,shipped
1001,Acme Ltd,2026-03-02,249.5,true
1002,Brightside Co,2026-03-02,1200,false
1003,Acme Ltd,2026-03-05,89.99,true
1004,Northwind,2026-03-07,,false

Output (Arrow)

Arrow IPC file (binary, columnar) — shown as a table with its schema

order_idcustomerorder_dateamountshipped
1001Acme Ltd2026-03-02249.5true
1002Brightside Co2026-03-021200false
1003Acme Ltd2026-03-0589.99true
1004Northwind2026-03-07NULLfalse

Schema: order_id BIGINT, customer VARCHAR, order_date DATE, amount DOUBLE, shipped BOOLEAN

What changes when you convert CSV to Arrow

  • Each CSV column becomes a typed Arrow column: order_id is Int64, amount is Float64 and shipped is Bool.
  • order_date becomes an Arrow date column (date32 in pyarrow), not a string.
  • Text columns such as customer become UTF-8 string columns.
  • Order 1004's empty amount is recorded in the column's validity bitmap. It reads back as null in pyarrow and NaN in a pandas float column.
  • The file begins with the ARROW1 magic bytes and carries its own schema, so readers never re-infer types.

Your file is processed locally in your browser and is never uploaded. The free limit is 50 MB per file; larger files work if your device has the memory for them.

Frequently Asked Questions

How do I open the .arrow file in Python?

pyarrow.ipc.open_file("orders.arrow").read_all() returns a pyarrow Table. pandas.read_feather("orders.arrow") and polars.read_ipc("orders.arrow") also work, because an uncompressed Arrow IPC file is the same thing as a Feather version 2 file.

Is Arrow smaller than CSV?

Usually not by much, and sometimes it is larger. Every integer and decimal takes a fixed 8 bytes, and nothing is compressed. The benefit is load speed, not size. If size matters, convert CSV to Parquet instead.

Is the output the Arrow file format or the stream format?

The file format, with ARROW1 magic bytes at the start and end and a footer that allows random access to record batches. Readers that expect the stream format, such as pyarrow.ipc.open_stream, will reject it. Use open_file.

What is Arrow format?

Apache Arrow is a columnar in-memory format designed for zero-copy reads and high-speed exchange between data systems.

Related Tools