SmartQueryTools

Extract JSON Column from Arrow Files Online

Extract values from a JSON-encoded column in Arrow files into a new flat column using a JSON path expression. Runs in your browser.

How to extract JSON Column from Arrow files

  1. Drop your file onto the upload area. It is loaded into the in-browser engine and the first 200 rows are shown.
  2. Choose the JSON column. Only text columns are listed, and the first one is selected for you.
  3. Type a JSON path such as $.customer.country or $.tags[0]. The new column name fills in from the last part of the path, and you can change it.
  4. Pick Extract string value for plain text, or Extract raw JSON to keep quotes, objects and arrays as JSON text. Click Extract.
  5. Check the new column in the preview, then download the file in the same format you uploaded.

Your file is processed locally in your browser and is never uploaded. The free limit is 50 MB per file; larger files work if your device has the memory for them.

Worked example

A payments webhook log was exported with the whole event body in one payload column. The analyst needs the customer's country as its own column to count paid orders by country. One event is a refund, which has no customer object.

Input (Arrow)

Arrow IPC file (binary, columnar) — shown as a table with its schema

event_idreceived_atpayload
12026-08-01 09:14:00{"type":"order.paid","customer":{"id":"c_81","country":"IE"},"total":42.5}
22026-08-01 09:20:31{"type":"order.paid","customer":{"id":"c_17","country":"DE"},"total":18}
32026-08-01 10:02:07{"type":"refund.created","order_id":"o_992"}
42026-08-01 10:45:50{"type":"order.paid","customer":{"id":"c_40","country":"US"},"total":99.9}

Schema: event_id BIGINT, received_at TIMESTAMP, payload VARCHAR

Settings

  • JSON column: payload
  • JSON path: $.customer.country
  • New column name: country (filled in from the path)
  • Extract mode: Extract string value

Result

event_idreceived_atpayloadcountry
12026-08-01 09:14:00{"type":"order.paid","customer":{"id":"c_81","country":"IE"},"total":42.5}IE
22026-08-01 09:20:31{"type":"order.paid","customer":{"id":"c_17","country":"DE"},"total":18}DE
32026-08-01 10:02:07{"type":"refund.created","order_id":"o_992"}NULL
42026-08-01 10:45:50{"type":"order.paid","customer":{"id":"c_40","country":"US"},"total":99.9}US

The path walks into the customer object and returns its country for each row. The refund event has no customer key, so the path finds nothing and the new column is NULL for that row. Nothing fails. The payload column is kept as it was. With Extract raw JSON, the values would keep their JSON quotes, as "IE" rather than IE.

Frequently Asked Questions

What JSON path syntax is supported?

Start with $ for the root, use dots for keys ($.user.email) and brackets for array positions ($.tags[0]). Negative positions count from the end, so $.tags[-1] is the last element. Quote keys that contain spaces: $."first name". Keys are case-sensitive.

What happens when the path does not exist or a row is not valid JSON?

Both give NULL for that row, and so does an empty cell. The run does not stop. After it finishes, the tool shows how many rows were not valid JSON, so you know whether to clean them.

Can I extract several fields in one run?

No, one path per run, and each run starts from the uploaded file. Download the result and load it again to add another field, or use the SQL Query tool with several json_extract_string calls in one SELECT.

Related Tools