Normalize Columns in Parquet Files Online
Normalize numeric columns in Parquet files using min-max scaling (0–1) or z-score standardisation (mean=0, std=1). Adds new columns alongside the originals — no upload required.
How to normalize Columns in Parquet files
- Drop your file onto the upload area. It is loaded into the in-browser engine and the first 200 rows are shown.
- Choose a method: Min-Max (0–1), the default, or Z-Score (μ=0, σ=1).
- Tick the numeric columns to scale. Every numeric column is ticked to start with, and the table shows the output name for each one (_normalized or _zscore).
- Click Normalize. Each new column is placed directly after the column it was computed from, and the originals are kept.
- Download the result in the same format you uploaded.
Your file is processed locally in your browser and is never uploaded. The free limit is 50 MB per file; larger files work if your device has the memory for them.
Worked example
A coach compares five players on a reaction-time test in milliseconds and a vertical jump in centimetres. The units are different, so the raw numbers cannot be compared or combined directly.
Input (Parquet)
Parquet file (binary, columnar) — shown as a table with its schema
| player | reaction_ms | jump_cm |
|---|---|---|
| Aoife | 180 | 40 |
| Bram | 220 | 52 |
| Chidi | 200 | 46 |
| Dana | 160 | 60 |
| Eli | 260 | 35 |
Schema: player VARCHAR, reaction_ms BIGINT, jump_cm BIGINT
Settings
- Method: Min-Max (0–1)
- Columns: reaction_ms and jump_cm (both ticked)
Result
| player | reaction_ms | reaction_ms_normalized | jump_cm | jump_cm_normalized |
|---|---|---|---|---|
| Aoife | 180 | 0.2 | 40 | 0.2 |
| Bram | 220 | 0.6 | 52 | 0.68 |
| Chidi | 200 | 0.4 | 46 | 0.44 |
| Dana | 160 | 0 | 60 | 1 |
| Eli | 260 | 1 | 35 | 0 |
Reaction times run from 160 to 260 ms, a range of 100, so Aoife's 180 becomes (180 − 160) / 100 = 0.2. Jumps run from 35 to 60 cm, so Bram's 52 becomes 17 / 25 = 0.68. The fastest reaction scores 0 because min-max does not know that lower is better here. Use Calculate Column with 1 − reaction_ms_normalized to flip it.
Working with Parquet files
Parquet files used as model input often have dozens of numeric features. They are all listed with their stored type: INT32, INT64, FLOAT, DOUBLE or DECIMAL. Use None, then tick only the features you want scaled, so labels and ID columns are left alone. Timestamp, boolean and string columns are never offered. Nested struct fields cannot be picked either, so flatten them first if a feature lives inside one.
Every new column is written as DOUBLE, whatever the source type, and the rest of the schema is unchanged. The minimum, maximum, mean and standard deviation are worked out over the whole file, not per row group. A file written in many row groups gives the same result as one written in a single group. Keep a note of those statistics if you need to apply the same scaling to a test set later. This tool does not save them.
Frequently Asked Questions
What Parquet type do normalized columns get?
DOUBLE. That applies even when the source column is INT32 or DECIMAL, because the scaled values are fractions.
Can I reuse the same scaling on another Parquet file?
Not automatically. Each run computes min, max, mean and standard deviation from the file you load. To scale a test set with training statistics, write the formula with fixed numbers in the SQL Query tool.
Does z-score use the sample or population standard deviation?
The sample standard deviation, which divides by n − 1. Results therefore differ slightly from tools that use the population figure, such as scikit-learn's StandardScaler. The gap shrinks as the row count grows.
What happens if every value in a column is the same?
The range (for min-max) or standard deviation (for z-score) is zero. The tool returns NULL for that column instead of dividing by zero. Null input values also stay null.
Can I scale values within each group, such as per region?
No. Statistics are computed over the whole file. For per-group scaling, filter the file to one group at a time, or use the SQL Query tool with PARTITION BY in the window functions.
Related Tools
Add Calculated Column to Parquet Files Online
Add a new column to Parquet files computed from an arithmetic expression over existing columns. No formulas, no code — just point and click.
Compute Correlation Matrix for Parquet Files Online
Compute a Pearson correlation matrix for numeric columns in Parquet files directly in your browser. Instantly spot which variables move together — colour-coded heatmap, no upload required.
Round Numbers in Parquet Files Online
Round numeric columns in Parquet files to any number of decimal places directly in your browser. Set precision per column with a simple slider — no upload required.
Normalize Columns in CSV Files Online
Normalize numeric columns in CSV files using min-max scaling (0–1) or z-score standardisation (mean=0, std=1). Adds new columns alongside the originals — no upload required.
Normalize Columns in Excel Files Online
Normalize numeric columns in Excel files using min-max scaling (0–1) or z-score standardisation (mean=0, std=1). Adds new columns alongside the originals — no upload required.
Normalize Columns in JSON Files Online
Normalize numeric columns in JSON files using min-max scaling (0–1) or z-score standardisation (mean=0, std=1). Adds new columns alongside the originals — no upload required.
Parquet Viewer Online
View and inspect Parquet files directly in your browser. Browse rows, check column names and data types — no upload required, your data stays on your device.
Convert Parquet to CSV Online
Convert Parquet files to CSV format directly in your browser. No upload required — your data never leaves your device.