SmartQueryTools

Add Row Numbers to Parquet Files Online

Add a row number index column to Parquet files directly in your browser. Set the column name and starting number — no upload required.

How to add Row Numbers to Parquet files

  1. Drop your file onto the upload area. The row count is shown and the first 200 rows are previewed.
  2. Type a name for the new column. It defaults to row_number.
  3. Set Start at to the first number you want. It defaults to 1, and 0 or a negative number also works.
  4. Click Add Row Numbers. The new column appears as the first column in the preview.
  5. Click Download to save the numbered file in the same format you uploaded.

Your file is processed locally in your browser and is never uploaded. The free limit is 50 MB per file; larger files work if your device has the memory for them.

Worked example

A research team has a batch of survey responses and a second batch already numbered up to 100. They want the new batch to carry on from 101 so the two can be combined later.

Input (Parquet)

Parquet file (binary, columnar) — shown as a table with its schema

respondentcountryscorecompleted
r_8f2aIreland7true
r_03cdPortugal9true
r_b771IrelandNULLfalse
r_5e10Canada6true

Schema: respondent VARCHAR, country VARCHAR, score BIGINT, completed BOOLEAN

Settings

  • Column name: response_no
  • Start at: 101

Result

response_norespondentcountryscorecompleted
101r_8f2aIreland7true
102r_03cdPortugal9true
103r_b771IrelandNULLfalse
104r_5e10Canada6true

A new integer column called response_no is placed in front of the existing columns and counts up by one per row from 101. Numbers follow the order of the rows in the file. The other columns are unchanged, including the missing score for r_b771.

Working with Parquet files

The new column is added to the schema as a BIGINT, a 64-bit integer, whatever the size of the file. It is the first field in the schema, before the original columns. All existing column types, including decimals, timestamps and nested structs, are kept as they were.

Parquet readers often read row groups in parallel and may not return rows in file order. Adding an explicit number column fixes the original order as data. After that, any tool can restore it with ORDER BY row_number, even after the file has been filtered, split or merged. That makes this a useful step before handing a Parquet file to a distributed job. The numbers follow the order rows are read from the file. If you want them to follow a date or ID instead, sort first and then number the sorted output. Because the column is a plain integer, it also works as a simple partition or bucket key, for example row_number % 10.

Frequently Asked Questions

What type is the row number column in the Parquet output?

BIGINT, a signed 64-bit integer. It is large enough for any row count a browser can load.

Can I use the row number as a stable key in a data lake table?

Yes, as long as you number the file once and keep the column. Re-running the tool on a changed file will renumber every row from the start value.

Can I start numbering at 0?

Yes. Set Start at to 0 for zero-based indexes. Any whole number works, including negatives. An empty box or a value that is not a number uses the default of 1, and decimals are cut to a whole number.

In what order are rows numbered?

In the order they appear in the file. The tool does not sort. Sort the file first if you want numbers to follow a date or ID column.

What if I leave the column name blank?

The column is named row_number. Choose a name that is not already used by a column in your file.

Related Tools