SmartQueryTools

Convert Parquet to NDJSON Online

Convert Parquet files to NDJSON format directly in your browser. No upload required — your data never leaves your device.

About converting Parquet to NDJSON

Parquet to NDJSON, also called JSON Lines, writes one self-contained JSON object per line. It is the format that streaming and bulk-load systems prefer: BigQuery and Snowflake load jobs, the Elasticsearch and OpenSearch bulk API, Kafka producers, and command-line tools like jq, grep and split. A job can start on line one without reading the rest of the file.

A typical case is moving an event table or a feature store extract from a lakehouse into a search index or a message queue. Another is building a fine-tuning or evaluation dataset: Hugging Face and many ML tools read .jsonl directly, while the source data lives in Parquet. Unlike a JSON array, the file can be cut at any line break and each piece is still valid.

Event data usually carries 64-bit IDs, such as Twitter Snowflake-style IDs. Any integer larger than 9,007,199,254,740,991 is written as a quoted string, because a JavaScript number cannot hold it exactly. Smaller integers remain numbers. That keeps the digits correct, but a strict schema on the receiving side may reject a string where it expects an integer, so check the target mapping before a bulk load.

Worked example

A small sample file, converted with the default settings.

Input (Parquet)

Parquet file (binary, columnar) — shown as a table with its schema

event_iduser_idevent_typeevent_timerevenue
18392047112345678915521view2026-03-02 09:15:00.123NULL
18392047112345678925521add_to_cart2026-03-02 09:15:41.870NULL
18392047112345679075521purchase2026-03-02 09:17:05249.5
18392047112345680127310view2026-03-02 09:18:22.004NULL

Schema: event_id BIGINT, user_id BIGINT, event_type VARCHAR, event_time TIMESTAMP, revenue DOUBLE

Output (NDJSON)

{"event_id":"1839204711234567891","user_id":5521,"event_type":"view","event_time":"2026-03-02 09:15:00.123","revenue":null}
{"event_id":"1839204711234567892","user_id":5521,"event_type":"add_to_cart","event_time":"2026-03-02 09:15:41.870","revenue":null}
{"event_id":"1839204711234567907","user_id":5521,"event_type":"purchase","event_time":"2026-03-02 09:17:05","revenue":249.5}
{"event_id":"1839204711234568012","user_id":7310,"event_type":"view","event_time":"2026-03-02 09:18:22.004","revenue":null}

What changes when you convert Parquet to NDJSON

  • Each Parquet row becomes one line holding a compact JSON object. There is no enclosing array and no trailing comma.
  • BIGINT values within the safe range stay numbers, like user_id 5521. Larger ones, like event_id 1839204711234567891, become strings.
  • Timestamps become strings with millisecond precision, such as "2026-03-02 09:15:00.123". Microseconds and nanoseconds from Spark or pandas writers are cut to milliseconds.
  • Structs and lists stay nested inside each line, so a single line can be long. MAP columns keep their keys as object keys.
  • NULL becomes null. The key is still present on every line, so revenue appears as null on the three non-purchase events.

Your file is processed locally in your browser and is never uploaded. The free limit is 50 MB per file; larger files work if your device has the memory for them.

Frequently Asked Questions

Why did my event_id turn into a string in the NDJSON file?

The value is above 2^53, the largest integer a JavaScript or double-precision number stores exactly. Writing it as a number would silently change the last digits. BigQuery load jobs and many other loaders accept a quoted integer for an INT64 field.

Can I split the NDJSON output into smaller files?

Yes. Every line is a complete record, so split -l 100000 on Linux or macOS, or any line-based splitter, produces valid files. That is not true of a JSON array.

Is NDJSON the same as JSON Lines and .jsonl?

Yes, for practical purposes. The file uses one JSON object per line separated by LF. Rename it to .jsonl if a tool looks for that extension.

What is NDJSON format?

NDJSON (Newline-Delimited JSON) stores one JSON object per line, making it easy to stream, append, and process incrementally.

Related Tools