SmartQueryTools

Unpivot Parquet Files Online

Reshape Parquet files from wide to long format directly in your browser. Melt multiple columns into variable/value pairs — no upload required.

How to unpivot Parquet files

  1. Drop your file. The first column is ticked as an ID column by default.
  2. Tick every column that identifies a row, such as an ID, a name or a region. ID columns are repeated on every output row.
  3. Leave the columns you want to melt unticked. Each unticked column becomes rows holding its column name and its value.
  4. Name the two new columns. They default to variable and value.
  5. Click Unpivot, check the long-format preview, and download the result.

Your file is processed locally in your browser and is never uploaded. The free limit is 50 MB per file; larger files work if your device has the memory for them.

Worked example

An HR team runs a two-question pulse survey. The export has one row per respondent and one column per question, but the dashboard tool wants one row per answer. One respondent skipped the second question.

Input (Parquet)

Parquet file (binary, columnar) — shown as a table with its schema

respondent_idteamq1_scoreq2_score
1041Support45
1042Sales2NULL
1043Support54

Schema: respondent_id BIGINT, team VARCHAR, q1_score BIGINT, q2_score BIGINT

Settings

  • ID columns (ticked): respondent_id, team
  • Value columns (unticked): q1_score, q2_score
  • Variable column name: question
  • Value column name: score

Result

respondent_idteamquestionscore
1041Supportq1_score4
1041Supportq2_score5
1042Salesq1_score2
1042Salesq2_scoreNULL
1043Supportq1_score5
1043Supportq2_score4

Each respondent now has one row per question, with the ID columns repeated. Respondent 1042 skipped question 2, and that row is kept with an empty score. So the result has 6 rows, 3 respondents times 2 questions. Filter out empty scores afterwards if you only want answered questions.

Working with Parquet files

Parquet columns are strictly typed, and the value columns you melt must combine into one type. INTEGER and DOUBLE combine into DOUBLE. DECIMAL(10,2) and BIGINT combine into a decimal wide enough for both, DECIMAL(21,2). DATE and TIMESTAMP combine into TIMESTAMP. A string column next to numeric ones, or any other mix that cannot combine, turns the whole value column into text. Unpivot the string columns in a separate pass if you want to keep numeric types.

The variable column is written as a Parquet string column holding the old column names. It has only a few distinct values, so dictionary encoding keeps it small even when the row count grows a lot. The ID columns keep their original types. Unpivoting a wide sensor or finance file multiplies the row count, but the Parquet output often grows much less than the row count does. Tick struct and list columns as ID columns unless you mean to melt them, since they cannot combine with plain numbers.

Frequently Asked Questions

What type is the value column in the Parquet output?

The common type of the columns you melted. All DOUBLE gives DOUBLE, and INTEGER mixed with DOUBLE also gives DOUBLE. Types that cannot combine, such as INTEGER and VARCHAR, give a string column. The variable column is always a string.

Why is the Parquet value column a string after unpivoting?

The unticked columns have types with no common type, such as INTEGER and VARCHAR, or a struct and a number, so all values were converted to text. Tick the odd column as an ID column to keep the numeric type.

Are empty values kept?

Yes. Each input row gives one output row for every value column, even when the value is empty (NULL). Filter the value column afterwards if you want only filled values.

Can I unpivot columns of different types?

Yes. Whole numbers and decimals combine into a number column, and dates and timestamps into timestamps. Any other mix, such as numbers and text, is converted to text so the run still works.

How do I go back from long to wide?

Use the Pivot tool. Put the variable column into the column headers and use the value column for the cells.

Related Tools