SmartQueryTools

Convert NDJSON to Arrow Online

Convert NDJSON files to Arrow format directly in your browser. No upload required — your data never leaves your device.

About converting NDJSON to Arrow

NDJSON to Arrow suits analysts who pull event data into Python or R again and again. Reading a large JSON Lines file with pandas.read_json(lines=True) is slow and infers types afresh on every run. An Arrow IPC file loads much faster with pyarrow or Polars, and its column types are settled once, at conversion time. Row order follows the file, so the first log line is row 0 of the table.

The schema is built from the whole file. Keys that appear on only some lines become nullable columns, integer and float fields get Int64 and Float64, and nested objects inside a line become Struct columns, so a request object keeps its method, url and headers fields. Conflicting types for the same key, which happen when several services write to one stream, produce a string column holding the raw JSON for each value.

Two things to plan for. ISO timestamps with milliseconds arrive as strings, so parse them in your code, or convert them first with the Date Parse tool if you want an Arrow timestamp column. And the IPC file is written uncompressed; for long-term storage of the same logs, Parquet is several times smaller.

Worked example

A small sample file, converted with the default settings.

Input (NDJSON)

{"ts":"2026-03-02T14:05:11.042Z","level":"info","service":"checkout","status":200,"latency_ms":41.7,"user_id":"u_1842"}
{"ts":"2026-03-02T14:05:12.310Z","level":"error","service":"checkout","status":502,"latency_ms":3012.4,"user_id":null}
{"ts":"2026-03-02T14:05:12.877Z","level":"info","service":"search","status":200,"latency_ms":12.9,"user_id":"u_0077"}
{"ts":"2026-03-02T14:05:14.105Z","level":"warn","service":"search","status":429,"latency_ms":0.8,"user_id":"u_1842"}

Output (Arrow)

Arrow IPC file (binary, columnar) — shown as a table with its schema

tslevelservicestatuslatency_msuser_id
2026-03-02T14:05:11.042Zinfocheckout20041.7u_1842
2026-03-02T14:05:12.310Zerrorcheckout5023012.4NULL
2026-03-02T14:05:12.877Zinfosearch20012.9u_0077
2026-03-02T14:05:14.105Zwarnsearch4290.8u_1842

Schema: ts VARCHAR, level VARCHAR, service VARCHAR, status BIGINT, latency_ms DOUBLE, user_id VARCHAR

What changes when you convert NDJSON to Arrow

  • ts is a Utf8 string column, because ISO timestamps with milliseconds are not parsed as dates.
  • status becomes Int64 and latency_ms becomes Float64.
  • user_id is a nullable Utf8 column, and the error line holds a null in it.
  • Nested objects become Struct columns and arrays become List columns.
  • The download is a single Arrow IPC file with all lines as record batches, not a line-based text file.

Your file is processed locally in your browser and is never uploaded. The free limit is 50 MB per file; larger files work if your device has the memory for them.

Frequently Asked Questions

How do I load the Arrow file in R?

Install the arrow package and call arrow::read_ipc_file("events.arrow"), or read_feather(), which reads the same format. You get a data frame, or an Arrow Table if you pass as_data_frame = FALSE.

Why is my timestamp column a string?

Timestamps like 2026-03-02T14:05:11.042Z are kept as text by the NDJSON reader. Convert them in pandas with pd.to_datetime, in Polars with .str.to_datetime(), or run the Date Parse tool on the NDJSON before converting.

Is Arrow better than NDJSON for pandas?

For repeated loads, yes. pandas.read_feather or pyarrow reads the Arrow file without parsing text, and integer columns with nulls keep their integer type instead of turning into floats when you use Arrow-backed dtypes.

What is Arrow format?

Apache Arrow is a columnar in-memory format designed for zero-copy reads and high-speed exchange between data systems.

Related Tools