SmartQueryTools

Convert TSV to Parquet Online

Convert TSV files to Parquet format directly in your browser. No upload required — your data never leaves your device.

About converting TSV to Parquet

Tab-delimited files are the default output of much of genomics and scientific computing, and they get big: expression matrices with thousands of sample columns, variant annotation tables, per-read QC reports. Converting TSV to Parquet makes those files small and fast to query. Parquet compresses each column separately, and a column of repeated values such as chromosome names or sample IDs shrinks to almost nothing.

Polars, pandas with pyarrow, R's arrow package, Spark and cloud query services all read Parquet, and they can read only the columns an analysis needs instead of parsing every tab on every line. That matters most for wide files. Reading 5 columns from a 2,000-column table skips the other 1,995 entirely.

Types are inferred from the TSV, and scientific missing-value markers are the main trap. R writes NA, VCF uses a dot, and database dumps use \N. None of these counts as missing, so a pvalue or QUAL column containing them is stored as VARCHAR instead of DOUBLE. Replace them with empty fields first; empty fields become real nulls. Comment lines are handled: ## metadata lines at the top of a VCF-style file are skipped, and a #CHROM header becomes a column named #CHROM.

Worked example

A small sample file, converted with the default settings.

Input (TSV)

symbol	chrom	description	tpm	detected
TP53	chr17	tumor protein p53	24.7	true
BRCA1	chr17	BRCA1 DNA repair associated	8.13	true
NT5C	chr17	5', 3'-nucleotidase, cytosolic	3.4	true
SREBF2	chr22	sterol regulatory element binding transcription factor 2		false

Output (Parquet)

Parquet file (binary, columnar) — shown as a table with its schema

symbolchromdescriptiontpmdetected
TP53chr17tumor protein p5324.7true
BRCA1chr17BRCA1 DNA repair associated8.13true
NT5Cchr175', 3'-nucleotidase, cytosolic3.4true
SREBF2chr22sterol regulatory element binding transcription factor 2NULLfalse

Schema: symbol VARCHAR, chrom VARCHAR, description VARCHAR, tpm DOUBLE, detected BOOLEAN

What changes when you convert TSV to Parquet

  • Column names come from the TSV header line and are stored in the Parquet schema.
  • tpm becomes DOUBLE and detected becomes BOOLEAN. symbol, chrom and description become VARCHAR.
  • SREBF2's empty tpm is stored as null, not as zero or NaN.
  • chrom repeats chr17 three times, and dictionary encoding stores each distinct value once per column chunk.
  • The file is compressed with Snappy, the default codec.
  • A column holding NA, . or \N stays VARCHAR, because those markers are not read as missing.

Your file is processed locally in your browser and is never uploaded. The free limit is 50 MB per file; larger files work if your device has the memory for them.

Frequently Asked Questions

Will R and Python read the Parquet with the right types?

Yes. arrow::read_parquet() in R, and pandas.read_parquet() or polars.read_parquet() in Python, read the stored types directly. DOUBLE columns come back as numeric or float64, BOOLEAN as logical or bool, and nulls as NA in R.

How do I stop NA values from turning numeric columns into text?

Re-export the table from R with write.table(df, sep = "\t", na = "") so missing values are written as empty fields, or replace NA with nothing before converting. Empty fields load as nulls and the column keeps its numeric type.

Is Parquet a good fit for a genes-by-samples count matrix?

Yes, if you read subsets of samples. Each sample column is stored and compressed separately, so loading 10 samples out of 500 reads only those 10. If you always load the whole matrix at once, the gain is mostly file size.

What is Parquet format?

Parquet is an open-source columnar storage format designed for efficient analytics. It compresses far better than CSV and is natively supported by Spark, Athena, BigQuery, Pandas, and DuckDB.

Related Tools