SmartQueryTools

Compute Correlation Matrix for CSV Files Online

Compute a Pearson correlation matrix for numeric columns in CSV files directly in your browser. Instantly spot which variables move together — colour-coded heatmap, no upload required.

How to compute Correlation Matrix for CSV files

  1. Drop your file onto the upload area. Every numeric column is found and selected.
  2. Click column names to leave out any you do not want, such as ID or ZIP code columns. At least two must stay selected.
  3. Click Compute Correlations. A colour-coded matrix appears, blue for positive and red for negative, with values to 3 decimals. Hover a cell for 6 decimals.
  4. Click Export Matrix CSV to download the matrix with values to 4 decimals.

Your file is processed locally in your browser and is never uploaded. The free limit is 50 MB per file; larger files work if your device has the memory for them.

Worked example

An estate agent has recent house sales and wants to see which property features move with the sale price before building a pricing model.

Input (CSV)

floor_area_sqft,bedrooms,age_years,sale_price
1450,3,32,612000
2100,4,8,845000
980,2,55,455000
1720,3,20,701000
2600,5,3,990000
1200,2,41,540000

Settings

  • Numeric columns: floor_area_sqft, bedrooms, age_years, sale_price (all selected)
  • Export: Export Matrix CSV

Result

floor_area_sqftbedroomsage_yearssale_price
floor_area_sqft1.00000.9830-0.96690.9991
bedrooms0.98301.0000-0.93060.9798
age_years-0.9669-0.93061.0000-0.9721
sale_price0.99910.9798-0.97211.0000

The output is always a CSV, whatever the input format, with an empty top-left header cell. Floor area tracks price almost perfectly at 0.9991. Age is strongly negative at -0.9721: older homes sold for less. The diagonal is 1 because each column matches itself, and the matrix is symmetric. With only six rows these values are fragile, and one unusual sale could move them a lot.

Working with CSV files

A column is offered only if the engine reads it as a number. In a CSV, one stray value such as "n/a", "-" or "1,250" turns the whole column into text and it drops out of the matrix without warning. If a column you expected is missing from the list, open the preview, find the odd values, and clean them with Find & Replace or Cast Columns before computing.

Empty cells load as NULL. Each pair of columns is computed from the rows where both values are present, so different cells in the matrix can be based on different row counts. A column that is mostly empty can show a strong correlation from only a handful of rows. Columns of 0 and 1 flags are read as integers and are included. Columns of true and false are read as booleans and are not. The heatmap is only as good as the load, so glance at the column list before you trust a missing relationship.

Frequently Asked Questions

Why is one of my CSV columns missing from the correlation list?

It was read as text because at least one value is not a plain number. Clean those values and cast the column, then drop the file in again.

How are blank CSV cells handled?

Each pair uses only the rows where both columns have a value. Blanks are skipped, not treated as zero.

Which correlation method is used?

Pearson correlation, which measures straight-line relationships. A strong curved relationship can still show a value near 0. Spearman and Kendall are not available.

Why does a cell show NaN or a dash?

NaN means one of the two columns has the same value in every row it shares with the other, or only one shared row exists, so there is no variation to correlate. A dash means there were no rows with values in both columns. In the CSV export a dash becomes an empty cell.

Is there a limit on the number of columns?

There is no fixed limit, but the matrix grows with the square of the column count. Deselect IDs and other columns that are not real measures to keep it readable.

Related Tools

Detect Outliers in CSV Files Online

Detect statistical outliers in CSV files directly in your browser. Flag or remove rows where numeric values exceed a chosen number of standard deviations from the mean — no upload required.

Normalize Columns in CSV Files Online

Normalize numeric columns in CSV files using min-max scaling (0–1) or z-score standardisation (mean=0, std=1). Adds new columns alongside the originals — no upload required.

Manage Columns in CSV Files Online

Drop or select specific columns from CSV files directly in your browser. No upload required.

Compute Correlation Matrix for Excel Files Online

Compute a Pearson correlation matrix for numeric columns in Excel files directly in your browser. Instantly spot which variables move together — colour-coded heatmap, no upload required.

Compute Correlation Matrix for Parquet Files Online

Compute a Pearson correlation matrix for numeric columns in Parquet files directly in your browser. Instantly spot which variables move together — colour-coded heatmap, no upload required.

Compute Correlation Matrix for JSON Files Online

Compute a Pearson correlation matrix for numeric columns in JSON files directly in your browser. Instantly spot which variables move together — colour-coded heatmap, no upload required.

CSV Viewer Online

View and inspect CSV files directly in your browser. Browse rows, check column names and data types — no upload required, your data stays on your device.

Convert CSV to Parquet Online

Convert CSV files to Parquet format directly in your browser. No upload required — your data never leaves your device.