SmartQueryTools

Extract JSON Column from Parquet Files Online

Extract values from a JSON-encoded column in Parquet files into a new flat column using a JSON path expression. Runs in your browser.

How to extract JSON Column from Parquet files

  1. Drop your file onto the upload area. It is loaded into the in-browser engine and the first 200 rows are shown.
  2. Choose the JSON column. Only text columns are listed, and the first one is selected for you.
  3. Type a JSON path such as $.customer.country or $.tags[0]. The new column name fills in from the last part of the path, and you can change it.
  4. Pick Extract string value for plain text, or Extract raw JSON to keep quotes, objects and arrays as JSON text. Click Extract.
  5. Check the new column in the preview, then download the file in the same format you uploaded.

Your file is processed locally in your browser and is never uploaded. The free limit is 50 MB per file; larger files work if your device has the memory for them.

Worked example

A payments webhook log was exported with the whole event body in one payload column. The analyst needs the customer's country as its own column to count paid orders by country. One event is a refund, which has no customer object.

Input (Parquet)

Parquet file (binary, columnar) — shown as a table with its schema

event_idreceived_atpayload
12026-08-01 09:14:00{"type":"order.paid","customer":{"id":"c_81","country":"IE"},"total":42.5}
22026-08-01 09:20:31{"type":"order.paid","customer":{"id":"c_17","country":"DE"},"total":18}
32026-08-01 10:02:07{"type":"refund.created","order_id":"o_992"}
42026-08-01 10:45:50{"type":"order.paid","customer":{"id":"c_40","country":"US"},"total":99.9}

Schema: event_id BIGINT, received_at TIMESTAMP, payload VARCHAR

Settings

  • JSON column: payload
  • JSON path: $.customer.country
  • New column name: country (filled in from the path)
  • Extract mode: Extract string value

Result

event_idreceived_atpayloadcountry
12026-08-01 09:14:00{"type":"order.paid","customer":{"id":"c_81","country":"IE"},"total":42.5}IE
22026-08-01 09:20:31{"type":"order.paid","customer":{"id":"c_17","country":"DE"},"total":18}DE
32026-08-01 10:02:07{"type":"refund.created","order_id":"o_992"}NULL
42026-08-01 10:45:50{"type":"order.paid","customer":{"id":"c_40","country":"US"},"total":99.9}US

The path walks into the customer object and returns its country for each row. The refund event has no customer key, so the path finds nothing and the new column is NULL for that row. Nothing fails. The payload column is kept as it was. With Extract raw JSON, the values would keep their JSON quotes, as "IE" rather than IE.

Working with Parquet files

In Parquet, JSON usually lives in a string column: a raw event body, an API response, or a properties bag written by an ingestion tool. Those columns appear in the JSON column picker. Columns stored as real struct, list or map types do not appear, because they are not text. Their fields are already typed and can be read with dot notation in the SQL Query tool. Columns with the JSON type are listed too. Plain text columns are listed as well, so check you picked the right one. A column of ordinary text gives NULL on every row, and the not valid JSON count shows it.

The extracted column is added to the schema as a string column in both modes. A number pulled out of the JSON is stored as text, so cast it with Cast Column Types if you want to sum or average it later. The other columns keep their Parquet types exactly, including timestamps and decimals. The downloaded file is rewritten with fresh row groups and default compression.

Frequently Asked Questions

Why is my nested Parquet column not in the JSON column list?

It is stored as a struct, list or map, not as text. The picker only lists text columns. Nested Parquet types do not need JSON parsing. Query their fields directly in the SQL Query tool, for example SELECT address.city FROM your_table.

What type does the extracted column get in the Parquet output?

A string column. That holds even when the JSON value was a number or boolean. Use Cast Column Types afterwards to turn it into BIGINT, DOUBLE or BOOLEAN.

What JSON path syntax is supported?

Start with $ for the root, use dots for keys ($.user.email) and brackets for array positions ($.tags[0]). Negative positions count from the end, so $.tags[-1] is the last element. Quote keys that contain spaces: $."first name". Keys are case-sensitive.

What happens when the path does not exist or a row is not valid JSON?

Both give NULL for that row, and so does an empty cell. The run does not stop. After it finishes, the tool shows how many rows were not valid JSON, so you know whether to clean them.

Can I extract several fields in one run?

No, one path per run, and each run starts from the uploaded file. Download the result and load it again to add another field, or use the SQL Query tool with several json_extract_string calls in one SELECT.

Related Tools