SmartQueryTools

Convert YAML to Arrow Online

Convert YAML files to Arrow format directly in your browser. No upload required — your data never leaves your device.

About converting YAML to Arrow

YAML to Arrow is for Python, R and JavaScript users who want to load record-style YAML fast and with types. Parsing YAML is slow, even with a C-based loader. An Arrow IPC file opens almost instantly with pyarrow, polars, pandas.read_feather or the R arrow package, and it can be memory-mapped instead of read into RAM.

The output is an Arrow IPC file, also known as Feather version 2. Arrow has proper nested types, so a nested mapping becomes a Struct field and a YAML sequence becomes a List field. Nothing is flattened or turned into strings. This makes Arrow a good fit for YAML that mixes flat fields with small nested blocks, such as a list of services each with a ports list.

Column types are decided from all records together. If one record has count: 3 and another has count: "three", the whole column becomes a string column. Keys missing from a record are stored as nulls. Unquoted dates arrive as ISO timestamp text rather than a date type, and integers above 2^53 lose precision during parsing. Quote those values in the YAML if exact types matter. The file is not compressed. Convert to Parquet instead if you need a small file for storage.

Worked example

A small orders export with an ID, a customer name, a date, an amount (one missing) and a true/false flag, converted with the default settings.

Input (YAML)

- order_id: 1001
  customer: Acme Ltd
  order_date: '2026-03-02'
  amount: 249.5
  shipped: true
- order_id: 1002
  customer: Brightside Co
  order_date: '2026-03-02'
  amount: 1200
  shipped: false
- order_id: 1003
  customer: Acme Ltd
  order_date: '2026-03-05'
  amount: 89.99
  shipped: true
- order_id: 1004
  customer: Northwind
  order_date: '2026-03-07'
  amount: null
  shipped: false

Output (Arrow)

Arrow IPC file (binary, columnar) — shown as a table with its schema

order_idcustomerorder_dateamountshipped
1001Acme Ltd2026-03-02249.5true
1002Brightside Co2026-03-021200false
1003Acme Ltd2026-03-0589.99true
1004Northwind2026-03-07NULLfalse

Schema: order_id BIGINT, customer VARCHAR, order_date DATE, amount DOUBLE, shipped BOOLEAN

What changes when you convert YAML to Arrow

  • Each key becomes an Arrow field. order_id is Int64, amount Float64 and shipped Bool.
  • The quoted order_date values become a Date32 column.
  • Nested mappings become Struct fields, and sequences become List fields.
  • null values and missing keys are stored as nulls, as for order 1004's amount.
  • The file uses the Arrow IPC file layout with the ARROW1 header, without compression.

Your file is processed locally in your browser and is never uploaded. The free limit is 50 MB per file; larger files work if your device has the memory for them.

Frequently Asked Questions

How do I open the Arrow file in Python?

Use pyarrow.ipc.open_file(path).read_all(), polars.read_ipc(path) or pandas.read_feather(path). All three read the same file, including nested Struct and List columns.

Why is a numeric YAML field a string column in Arrow?

At least one record holds a non-numeric value for that key, such as "n/a" or a quoted number. Arrow columns have one type, so the whole column becomes text. Clean the YAML or cast the column afterwards.

Are YAML comments stored in the Arrow metadata?

No. Comments are discarded when the YAML is parsed, and the Arrow schema only holds field names and types.

What is Arrow format?

Apache Arrow is a columnar in-memory format designed for zero-copy reads and high-speed exchange between data systems.

Related Tools