SmartQueryTools

Convert JSON to Parquet Online

Convert JSON files to Parquet format directly in your browser. No upload required — your data never leaves your device.

About converting JSON to Parquet

People convert JSON to Parquet to stop paying, in storage and in parse time, for key names repeated in every record. An array of 100,000 objects spells out every key 100,000 times. Parquet writes each column name once in the file footer and compresses the values with Snappy. API archives, mongoexport --jsonArray dumps and nightly snapshots pulled from a REST endpoint are the usual candidates before they go to S3, BigQuery or a pandas notebook.

Unlike the flat text formats, Parquet can keep the shape of the data. A nested object becomes a STRUCT column with named fields, an array of strings becomes a LIST, and an array of objects becomes a LIST of STRUCTs. Query engines can then read address.city or unnest line items without parsing JSON again. Types are fixed at conversion time, so check the column list shown after loading before you publish the file.

Three things commonly go wrong. Numbers that the API sends as strings, such as "12.50", stay VARCHAR. Timestamp detection is narrow: whole-second UTC values such as 2026-03-02T09:14:00Z become TIMESTAMP, but values with milliseconds or a +02:00 offset stay text. A key that holds a number in some records and a string in others becomes a JSON-typed text column. Fix these with a CAST in the SQL Query tool if downstream code expects real types.

Not sure which format you need? Read the JSON vs CSV comparison.

Worked example

A small sample file, converted with the default settings.

Input (JSON)

[
  {
    "id": 4101,
    "title": "Login fails on Safari 17",
    "state": "open",
    "comments": 3,
    "created": "2026-03-02",
    "assignee": "maria"
  },
  {
    "id": 4102,
    "title": "Export slow, then times out",
    "state": "closed",
    "comments": 8,
    "created": "2026-03-03",
    "assignee": null
  },
  {
    "id": 4103,
    "title": "Add dark mode",
    "state": "open",
    "comments": 12,
    "created": "2026-03-05",
    "assignee": "dev-team"
  },
  {
    "id": 4104,
    "title": "Crash on empty file",
    "state": "closed",
    "comments": 0,
    "created": "2026-03-06",
    "assignee": "li.wei"
  }
]

Output (Parquet)

Parquet file (binary, columnar) — shown as a table with its schema

idtitlestatecommentscreatedassignee
4101Login fails on Safari 17open32026-03-02maria
4102Export slow, then times outclosed82026-03-03NULL
4103Add dark modeopen122026-03-05dev-team
4104Crash on empty fileclosed02026-03-06li.wei

Schema: id BIGINT, title VARCHAR, state VARCHAR, comments BIGINT, created DATE, assignee VARCHAR

What changes when you convert JSON to Parquet

  • Keys become Parquet columns. In the example, id and comments become BIGINT, title, state and assignee become VARCHAR, and created becomes DATE.
  • Nested objects become STRUCT columns and arrays become LIST columns. Nothing is flattened or turned into text.
  • JSON numbers with a decimal point become DOUBLE, so 249.50 is stored as 249.5 and the original decimal scale is not kept.
  • null values and keys missing from a record are stored as NULL. Issue 4102 has a NULL assignee.
  • Timestamps with milliseconds or an offset, such as 2026-03-02T09:14:00.250+02:00, stay VARCHAR. Cast them to TIMESTAMP if you need time arithmetic.

Your file is processed locally in your browser and is never uploaded. The free limit is 50 MB per file; larger files work if your device has the memory for them.

Frequently Asked Questions

Does converting nested JSON to Parquet flatten it?

No. Objects become STRUCT columns and arrays become LIST columns, which Spark, Athena, BigQuery and Polars read natively. pandas with pyarrow returns them as Python dicts and lists. If the destination needs flat columns, for example an older BI tool, run Flatten JSON before converting.

Can I convert a GeoJSON file to Parquet with this tool?

Only in a limited way. A FeatureCollection is a single object, so the result is one row with a features LIST column. It is not GeoParquet: geometries are not encoded as WKB and no coordinate reference system metadata is written. Use GDAL (ogr2ogr) or a GIS tool for GeoParquet output.

Why is one of my numeric fields typed JSON or VARCHAR?

Either the values are quoted in the source ("price": "12.50"), which makes them strings, or the field is a number in some records and text in others, which makes it a JSON column. Use TRY_CAST in the SQL Query tool to turn it into a number, with unparseable values becoming NULL.

What is Parquet format?

Parquet is an open-source columnar storage format designed for efficient analytics. It compresses far better than CSV and is natively supported by Spark, Athena, BigQuery, Pandas, and DuckDB.

Related Tools