SmartQueryTools

Data Profiler

Load any supported file to get an instant profile: row count, column types, null rates, distinct value counts, and numeric statistics. No upload — runs entirely in your browser.

Loading DuckDB engine…

Drop your data file here

CSV, TSV, Parquet, JSON, NDJSON, Arrow, Excel or YAML

or click to browse · max 50 MB

How to use the Data Profiler

  1. Drag your file onto the drop area, or click it to browse. CSV, TSV, Parquet, JSON, NDJSON, Arrow, Excel (.xlsx or .xls, first sheet) and YAML files work, up to 50 MB.
  2. The profile runs as soon as the file loads. There are no settings to choose.
  3. Read the summary cards: total rows, number of columns, how many columns are numeric, and how many are null-heavy (more than half empty).
  4. Scan the column profile table. Each column shows its detected type, null count, null percentage and number of distinct values. Numeric columns also show min, max, mean and standard deviation.
  5. Check the warnings under the table. Columns that are entirely empty are flagged in red, and columns more than 50% empty are flagged in amber.

Worked example

A sensor export has four rows. Before loading it into a database you want to know the column types and where the gaps are.

Input (sensors.csv)

id,city,temp_c,note
1,Oslo,4.5,
2,Oslo,5.0,
3,Lima,19.5,check
4,Lima,,

Output (profile)

Total rows: 4   Columns: 4   Numeric cols: 2   Null-heavy cols: 1

Column  Type     Nulls  % Null  Distinct  Min  Max   Mean    Std Dev
id      BIGINT   0      0.0%    4         1    4     2.5     1.291
city    VARCHAR  0      0.0%    2
temp_c  DOUBLE   1      25.0%   3         4.5  19.5  9.6667  8.5196
note    VARCHAR  3      75.0%   1

Warning: note: 75.0% of values are null (3 rows)

temp_c is detected as a decimal number and has one missing reading, so its mean is taken over the three values present. note is empty in three of four rows, which triggers the null-heavy warning. Distinct counts ignore empty values.

Frequently Asked Questions

Is my file uploaded to be profiled?

No. The profiler runs entirely in your browser using a SQL engine compiled to WebAssembly. Your file is read locally, the statistics are computed on your device, and nothing is sent to a server.

What does the profile include?

Row count, column names and detected data types, null counts and null rates per column, distinct value counts, and min, max, mean and standard deviation for numeric columns. These are the checks you would normally script in pandas with df.info() and df.describe(), without writing code.

How are column types detected?

For CSV, Excel and JSON files, the types are inferred by sampling the values in each column. A column of whole numbers becomes BIGINT, decimals become DOUBLE, and anything mixed stays as text (VARCHAR). Parquet and Arrow files keep the types stored in the file.

When should I profile a file?

Before loading data into a database, warehouse or machine learning pipeline, and whenever you receive a file from someone else. A quick profile catches unexpected nulls, wrongly typed columns and suspicious distinct counts before they cause failures further down the line.

Can I export the profile?

Not directly. The profile is shown on the page. For a custom summary you can save, load the file into the SQL Query Tool and run SUMMARIZE your_table, then download the result as CSV.

Related tools