Convert NDJSON to Parquet Online
Convert NDJSON files to Parquet format directly in your browser. No upload required — your data never leaves your device.
About converting NDJSON to Parquet
NDJSON is how logs and events land; Parquet is how they get queried. Converting a day of application logs, an elasticdump backup or a BigQuery JSON export to Parquet shrinks it, fixes the column types, and lets Athena, Spark or Polars read only the fields a query touches, such as status and latency_ms out of events that carry dozens of keys.
Records in a stream drift over time. A field added in a later release shows up only on newer lines, and one service may log user_id as a number while another logs it as a string. Every key found in the file becomes a column, older lines get NULL for fields they never had, and a key with conflicting types is stored as a JSON-typed text column rather than failing the load. Check the column list after loading; a JSON type is a sign of drift worth fixing at the source.
Nested payloads are kept as nested Parquet. elasticdump lines carry each document under _source, so the file gets a _source STRUCT with one field per document key, and a Pino err object becomes a STRUCT with type, message and stack. Most engines let you query _source.status directly. Log timestamps with milliseconds, as JavaScript loggers write them, are stored as text, so cast them before writing if you want a real TIMESTAMP column.
Not sure which format you need? Read the NDJSON vs CSV comparison.
Worked example
A small sample file, converted with the default settings.
Input (NDJSON)
{"ts":"2026-03-02T14:05:11.042Z","level":"info","service":"checkout","status":200,"latency_ms":41.7,"user_id":"u_1842"}
{"ts":"2026-03-02T14:05:12.310Z","level":"error","service":"checkout","status":502,"latency_ms":3012.4,"user_id":null}
{"ts":"2026-03-02T14:05:12.877Z","level":"info","service":"search","status":200,"latency_ms":12.9,"user_id":"u_0077"}
{"ts":"2026-03-02T14:05:14.105Z","level":"warn","service":"search","status":429,"latency_ms":0.8,"user_id":"u_1842"}Output (Parquet)
Parquet file (binary, columnar) — shown as a table with its schema
| ts | level | service | status | latency_ms | user_id |
|---|---|---|---|---|---|
| 2026-03-02T14:05:11.042Z | info | checkout | 200 | 41.7 | u_1842 |
| 2026-03-02T14:05:12.310Z | error | checkout | 502 | 3012.4 | NULL |
| 2026-03-02T14:05:12.877Z | info | search | 200 | 12.9 | u_0077 |
| 2026-03-02T14:05:14.105Z | warn | search | 429 | 0.8 | u_1842 |
Schema: ts VARCHAR, level VARCHAR, service VARCHAR, status BIGINT, latency_ms DOUBLE, user_id VARCHAR
What changes when you convert NDJSON to Parquet
- ts is stored as VARCHAR because the values carry milliseconds, a form the reader does not detect as a timestamp. The schema line shows it as text.
- status becomes BIGINT and latency_ms becomes DOUBLE. level, service and user_id become VARCHAR.
- The null user_id on the error line is stored as a Parquet NULL, not an empty string.
- Keys that appear on only some lines still become columns, and lines without them get NULL.
- A nested object on a line becomes a STRUCT column and an array becomes a LIST column.
Your file is processed locally in your browser and is never uploaded. The free limit is 50 MB per file; larger files work if your device has the memory for them.
Frequently Asked Questions
Some columns came out with type JSON. Why?
That key held different kinds of value on different lines, for example a number on some and a string on others, or null on every line. Rather than fail, the reader keeps the raw JSON text. Cast the column in the SQL Query tool, or fix the producer so it always logs the same type.
Can I query _source fields after converting an elasticdump file?
Yes. _source becomes a STRUCT column, so engines that read Parquet can select _source.status or _source.user.id directly. If a tool you use cannot read nested Parquet, run Flatten NDJSON first to pull _source fields into top-level columns.
How do I store ts as a real timestamp?
Before converting, open the NDJSON in the SQL Query tool and use CAST(ts AS TIMESTAMP), which reads the Z suffix and any +02:00 style offset and stores the value in UTC. The Date Parse tool can do the same without writing SQL.
What is Parquet format?
Parquet is an open-source columnar storage format designed for efficient analytics. It compresses far better than CSV and is natively supported by Spark, Athena, BigQuery, Pandas, and DuckDB.
Related Tools
Convert Parquet to NDJSON Online
Convert Parquet files to NDJSON format directly in your browser. No upload required — your data never leaves your device.
Parquet Viewer Online
View and inspect Parquet files directly in your browser. Browse rows, check column names and data types — no upload required, your data stays on your device.
NDJSON Viewer Online
View and inspect NDJSON files directly in your browser. Browse rows, check column names and data types — no upload required, your data stays on your device.
Filter Parquet Files Online
Filter rows in Parquet files by column value, directly in your browser. Your data stays on your device.
Convert NDJSON to CSV Online
Convert NDJSON files to CSV format directly in your browser. No upload required — your data never leaves your device.
Convert NDJSON to Excel Online
Convert NDJSON files to Excel format directly in your browser. No upload required — your data never leaves your device.
Convert NDJSON to JSON Online
Convert NDJSON files to JSON format directly in your browser. No upload required — your data never leaves your device.