SmartQueryTools

Convert YAML to Parquet Online

Convert YAML files to Parquet format directly in your browser. No upload required — your data never leaves your device.

About converting YAML to Parquet

YAML is pleasant to edit but slow to parse and has no compression, so it is a poor format for large datasets. Converting a YAML list of records to Parquet gives you a compact, typed file for pandas, polars, Spark, Athena or BigQuery, and a much faster load the next time the data is used.

Parquet keeps structure that flat formats lose. A nested mapping becomes a STRUCT column with named fields, and a YAML list becomes a LIST column. Tools that read Parquet can query nested fields directly, for example address.city in DuckDB or Spark SQL. There is no need to flatten first.

Types come from the parsed values. Integers become BIGINT, decimals DOUBLE, true and false BOOLEAN, and quoted ISO dates like '2026-03-02' become DATE. Unquoted dates are the exception. The YAML parser turns them into timestamps, and they arrive as text such as 2026-03-02T00:00:00.000Z. The parser follows YAML 1.2, so country codes like NO and words like yes stay strings. Very large integers above 9,007,199,254,740,991 lose precision during parsing, so quote long numeric IDs in the YAML.

Worked example

A small orders export with an ID, a customer name, a date, an amount (one missing) and a true/false flag, converted with the default settings.

Input (YAML)

- order_id: 1001
  customer: Acme Ltd
  order_date: '2026-03-02'
  amount: 249.5
  shipped: true
- order_id: 1002
  customer: Brightside Co
  order_date: '2026-03-02'
  amount: 1200
  shipped: false
- order_id: 1003
  customer: Acme Ltd
  order_date: '2026-03-05'
  amount: 89.99
  shipped: true
- order_id: 1004
  customer: Northwind
  order_date: '2026-03-07'
  amount: null
  shipped: false

Output (Parquet)

Parquet file (binary, columnar) — shown as a table with its schema

order_idcustomerorder_dateamountshipped
1001Acme Ltd2026-03-02249.5true
1002Brightside Co2026-03-021200false
1003Acme Ltd2026-03-0589.99true
1004Northwind2026-03-07NULLfalse

Schema: order_id BIGINT, customer VARCHAR, order_date DATE, amount DOUBLE, shipped BOOLEAN

What changes when you convert YAML to Parquet

  • Each list item becomes a row, and each key becomes a typed Parquet column.
  • order_id becomes BIGINT, amount DOUBLE, shipped BOOLEAN and the quoted order_date DATE.
  • Nested mappings become STRUCT columns, and YAML lists become LIST columns.
  • null, ~ and missing keys become NULL, as for order 1004's amount.
  • Comments, anchors and key order within nested mappings are not preserved as such. Anchors are expanded into full values.

Your file is processed locally in your browser and is never uploaded. The free limit is 50 MB per file; larger files work if your device has the memory for them.

Frequently Asked Questions

Does nested YAML survive conversion to Parquet?

Yes. Nested mappings become STRUCT columns and sequences become LIST columns, so the structure is kept. This is one of the few output formats where you do not need to flatten first.

Why is my date column VARCHAR in Parquet?

The dates are unquoted in the YAML, so the parser reads them as timestamps and they arrive as text like 2026-03-02T00:00:00.000Z. Quote the dates in the YAML source, or convert the column afterwards with the Date Parse or Cast Column Types tools.

What if records have different keys?

The Parquet schema includes every key found. Records that lack a key get NULL in that column. If the same key holds a number in one record and text in another, the column becomes VARCHAR.

What is Parquet format?

Parquet is an open-source columnar storage format designed for efficient analytics. It compresses far better than CSV and is natively supported by Spark, Athena, BigQuery, Pandas, and DuckDB.

Related Tools