SmartQueryTools

Normalize Columns in Parquet Files Online

Normalize numeric columns in Parquet files using min-max scaling (0–1) or z-score standardisation (mean=0, std=1). Adds new columns alongside the originals — no upload required.

How to normalize Columns in Parquet files

  1. Drop your file onto the upload area. It is loaded into the in-browser engine and the first 200 rows are shown.
  2. Choose a method: Min-Max (0–1), the default, or Z-Score (μ=0, σ=1).
  3. Tick the numeric columns to scale. Every numeric column is ticked to start with, and the table shows the output name for each one (_normalized or _zscore).
  4. Click Normalize. Each new column is placed directly after the column it was computed from, and the originals are kept.
  5. Download the result in the same format you uploaded.

Your file is processed locally in your browser and is never uploaded. The free limit is 50 MB per file; larger files work if your device has the memory for them.

Worked example

A coach compares five players on a reaction-time test in milliseconds and a vertical jump in centimetres. The units are different, so the raw numbers cannot be compared or combined directly.

Input (Parquet)

Parquet file (binary, columnar) — shown as a table with its schema

playerreaction_msjump_cm
Aoife18040
Bram22052
Chidi20046
Dana16060
Eli26035

Schema: player VARCHAR, reaction_ms BIGINT, jump_cm BIGINT

Settings

  • Method: Min-Max (0–1)
  • Columns: reaction_ms and jump_cm (both ticked)

Result

playerreaction_msreaction_ms_normalizedjump_cmjump_cm_normalized
Aoife1800.2400.2
Bram2200.6520.68
Chidi2000.4460.44
Dana1600601
Eli2601350

Reaction times run from 160 to 260 ms, a range of 100, so Aoife's 180 becomes (180 − 160) / 100 = 0.2. Jumps run from 35 to 60 cm, so Bram's 52 becomes 17 / 25 = 0.68. The fastest reaction scores 0 because min-max does not know that lower is better here. Use Calculate Column with 1 − reaction_ms_normalized to flip it.

Working with Parquet files

Parquet files used as model input often have dozens of numeric features. They are all listed with their stored type: INT32, INT64, FLOAT, DOUBLE or DECIMAL. Use None, then tick only the features you want scaled, so labels and ID columns are left alone. Timestamp, boolean and string columns are never offered. Nested struct fields cannot be picked either, so flatten them first if a feature lives inside one.

Every new column is written as DOUBLE, whatever the source type, and the rest of the schema is unchanged. The minimum, maximum, mean and standard deviation are worked out over the whole file, not per row group. A file written in many row groups gives the same result as one written in a single group. Keep a note of those statistics if you need to apply the same scaling to a test set later. This tool does not save them.

Frequently Asked Questions

What Parquet type do normalized columns get?

DOUBLE. That applies even when the source column is INT32 or DECIMAL, because the scaled values are fractions.

Can I reuse the same scaling on another Parquet file?

Not automatically. Each run computes min, max, mean and standard deviation from the file you load. To scale a test set with training statistics, write the formula with fixed numbers in the SQL Query tool.

Does z-score use the sample or population standard deviation?

The sample standard deviation, which divides by n − 1. Results therefore differ slightly from tools that use the population figure, such as scikit-learn's StandardScaler. The gap shrinks as the row count grows.

What happens if every value in a column is the same?

The range (for min-max) or standard deviation (for z-score) is zero. The tool returns NULL for that column instead of dividing by zero. Null input values also stay null.

Can I scale values within each group, such as per region?

No. Statistics are computed over the whole file. For per-group scaling, filter the file to one group at a time, or use the SQL Query tool with PARTITION BY in the window functions.

Related Tools

Add Calculated Column to Parquet Files Online

Add a new column to Parquet files computed from an arithmetic expression over existing columns. No formulas, no code — just point and click.

Compute Correlation Matrix for Parquet Files Online

Compute a Pearson correlation matrix for numeric columns in Parquet files directly in your browser. Instantly spot which variables move together — colour-coded heatmap, no upload required.

Round Numbers in Parquet Files Online

Round numeric columns in Parquet files to any number of decimal places directly in your browser. Set precision per column with a simple slider — no upload required.

Normalize Columns in CSV Files Online

Normalize numeric columns in CSV files using min-max scaling (0–1) or z-score standardisation (mean=0, std=1). Adds new columns alongside the originals — no upload required.

Normalize Columns in Excel Files Online

Normalize numeric columns in Excel files using min-max scaling (0–1) or z-score standardisation (mean=0, std=1). Adds new columns alongside the originals — no upload required.

Normalize Columns in JSON Files Online

Normalize numeric columns in JSON files using min-max scaling (0–1) or z-score standardisation (mean=0, std=1). Adds new columns alongside the originals — no upload required.

Parquet Viewer Online

View and inspect Parquet files directly in your browser. Browse rows, check column names and data types — no upload required, your data stays on your device.

Convert Parquet to CSV Online

Convert Parquet files to CSV format directly in your browser. No upload required — your data never leaves your device.