SmartQueryTools

Add UUID Column to Parquet Files Online

Add a UUID column to Parquet files directly in your browser. Generate random UUIDv4 or time-ordered UUIDv7 identifiers for every row. Choose the column name and position — no upload required.

How to add UUID Column to Parquet files

  1. Drop your file onto the upload area. The first 200 rows are shown.
  2. Choose the UUID version: UUIDv4 (random, the default) or UUIDv7 (time-ordered to the millisecond).
  3. Set the column name, which defaults to "id", and the position: first column (prepend) or last column (append).
  4. Click Add UUID Column. Every row gets its own identifier.
  5. Download the file in the same format you uploaded. Keep that downloaded copy, because running the tool again produces different UUIDs.

Your file is processed locally in your browser and is never uploaded. The free limit is 50 MB per file; larger files work if your device has the memory for them.

Worked example

A research team is about to share survey responses with an outside analyst. Rows have no identifier, and the team needs one so both sides can refer to the same response in follow-up questions.

Input (Parquet)

Parquet file (binary, columnar) — shown as a table with its schema

submitted_oncountrynps_score
2026-08-02Ireland9
2026-08-02Canada6
2026-08-03Ireland10
2026-08-04Japan7

Schema: submitted_on DATE, country VARCHAR, nps_score BIGINT

Settings

  • UUID version: UUIDv4 (random)
  • Column name: response_id
  • Column position: First column (prepend)

Result

response_idsubmitted_oncountrynps_score
3f9c2b1e-8a47-4d2c-9b6e-2c71d0a5e4f82026-08-02Ireland9
b07e41d9-5c3a-4f86-a1d2-9e8f6b3c7a152026-08-02Canada6
6a2d8e5f-1b94-47c0-8e3a-f5d1c9b2a0672026-08-03Ireland10
d4c81f36-e2a9-4b75-b608-1a7e3f9d52cb2026-08-04Japan7

A new response_id column is added as the first column and the other columns are unchanged. The UUIDs shown are examples. They are random, so you get different values every time you run the tool. The 4 at the start of the third group marks them as version 4. No row order or content is used to create them.

Working with Parquet files

Parquet has a native UUID logical type, and the new column is written with it: a 16-byte fixed-length binary value, not a 36-character string. That takes less than half the space of text and loads as a UUID in Spark, DuckDB, Polars and pyarrow. Readers that ignore logical types see the raw 16 bytes, so check your target tool before assuming it will display the usual hyphenated form.

UUIDv7 starts with a millisecond timestamp, so every id from one run shares nearly the same prefix, and ids from a later run sort after it. That helps when you append batches to a dataset over time: row-group statistics for each batch cover a narrow range and filters can skip them. Within a single run the ids are not in row order, because the bits after the timestamp are random. UUIDv4 values are spread across the whole range. The original schema is otherwise unchanged apart from the new column at the start or end.

Frequently Asked Questions

Is the UUID column stored as a string in Parquet?

No. It uses the Parquet UUID logical type, a 16-byte fixed-length binary. Tools that understand the type show it as a normal UUID.

Should I use UUIDv4 or UUIDv7 for a Parquet dataset?

UUIDv7 if you add batches over time and want later batches to sort after earlier ones. UUIDv4 if the id should reveal nothing about when a row was created. Neither version preserves row order inside one file, so add a row number too if you need that.

Are the UUIDs the same if I run the tool twice?

No. New values are generated on every run. Keep the downloaded file as the master copy. For ids that can be recreated from the data, use Hash & Anonymise Columns on a unique column instead.

What is the difference between UUIDv4 and UUIDv7?

UUIDv4 is random. UUIDv7 starts with a millisecond timestamp followed by random bits, so ids from a later run sort after ids from an earlier one. Rows generated in the same millisecond are in random order, so a whole file processed in one run is not sorted by its UUIDv7 column. Use Add Row Numbers if you need to keep the original order.

Can two rows get the same UUID?

In practice, no. A UUIDv4 has 122 random bits, so the chance of a repeat in a file of millions of rows is far too small to matter.

Related Tools

Manage Columns in Parquet Files Online

Drop or select specific columns from Parquet files directly in your browser. No upload required.

Hash & Anonymise Columns in Parquet Files Online

Anonymise or pseudonymise columns in Parquet files by replacing values with MD5, SHA-256, or DuckDB hashes — directly in your browser. Useful for GDPR compliance and sharing data without exposing PII — no upload required.

Split Parquet Files Online

Split Parquet files into multiple smaller files by row count, directly in your browser.

Add UUID Column to CSV Files Online

Add a UUID column to CSV files directly in your browser. Generate random UUIDv4 or time-ordered UUIDv7 identifiers for every row. Choose the column name and position — no upload required.

Add UUID Column to Excel Files Online

Add a UUID column to Excel files directly in your browser. Generate random UUIDv4 or time-ordered UUIDv7 identifiers for every row. Choose the column name and position — no upload required.

Add UUID Column to JSON Files Online

Add a UUID column to JSON files directly in your browser. Generate random UUIDv4 or time-ordered UUIDv7 identifiers for every row. Choose the column name and position — no upload required.

Parquet Viewer Online

View and inspect Parquet files directly in your browser. Browse rows, check column names and data types — no upload required, your data stays on your device.

Convert Parquet to CSV Online

Convert Parquet files to CSV format directly in your browser. No upload required — your data never leaves your device.