SmartQueryTools

Convert TSV to Arrow Online

Convert TSV files to Arrow format directly in your browser. No upload required — your data never leaves your device.

About converting TSV to Arrow

Arrow suits TSV data that gets reloaded again and again during an analysis. Parsing a large tab-delimited file in R or Python means scanning every byte, splitting on tabs and guessing types on each load. An Arrow IPC file stores already-typed columns in the layout they have in memory, so reloading is nearly instant and types never drift between runs.

It also fits workflows that mix languages. A table written as Arrow opens in R's arrow package, in pandas and Polars, in Julia's Arrow.jl and in the apache-arrow JavaScript library, with the same column types in each. That avoids R and Python disagreeing about how to parse the same TSV, for example over NA handling or whether a column of 0 and 1 is numeric.

Types are decided once, at conversion time, so fix the input first. Missing-value markers other than an empty field (NA, ., \N) make a column text, and a header line is needed for meaningful column names. The column list shown after the file loads gives each inferred type, so check it before downloading. The output is uncompressed, so for long-term storage or sending over a network, Parquet is the better target.

Worked example

A small sample file, converted with the default settings.

Input (TSV)

symbol	chrom	description	tpm	detected
TP53	chr17	tumor protein p53	24.7	true
BRCA1	chr17	BRCA1 DNA repair associated	8.13	true
NT5C	chr17	5', 3'-nucleotidase, cytosolic	3.4	true
SREBF2	chr22	sterol regulatory element binding transcription factor 2		false

Output (Arrow)

Arrow IPC file (binary, columnar) — shown as a table with its schema

symbolchromdescriptiontpmdetected
TP53chr17tumor protein p5324.7true
BRCA1chr17BRCA1 DNA repair associated8.13true
NT5Cchr175', 3'-nucleotidase, cytosolic3.4true
SREBF2chr22sterol regulatory element binding transcription factor 2NULLfalse

Schema: symbol VARCHAR, chrom VARCHAR, description VARCHAR, tpm DOUBLE, detected BOOLEAN

What changes when you convert TSV to Arrow

  • symbol, chrom and description become Arrow UTF-8 string columns.
  • tpm becomes a 64-bit float column and detected a boolean column.
  • SREBF2's empty tpm is marked null in the validity bitmap. R reads it as NA and pandas as NaN.
  • The file uses the Arrow IPC file format (Feather version 2), uncompressed, with the schema stored inside.
  • Column names match the TSV header exactly, including a leading # as in #CHROM.

Your file is processed locally in your browser and is never uploaded. The free limit is 50 MB per file; larger files work if your device has the memory for them.

Frequently Asked Questions

How do I open the Arrow file in R?

Install the arrow package and call arrow::read_feather("genes.arrow") or arrow::read_ipc_file("genes.arrow"). Both return a data frame. Pass as_data_frame = FALSE to keep it as an Arrow Table.

Why pick Arrow over Parquet for a TSV?

Pick Arrow when load speed and memory mapping matter more than size, such as a scratch file reloaded many times in one analysis. Pick Parquet for anything stored, shared or uploaded, because it is compressed and far smaller.

Can I memory-map the file?

Yes. Because it is uncompressed, pyarrow.memory_map combined with pyarrow.ipc.open_file reads columns straight from disk without copying them into memory first. That lets you open a file larger than your RAM and pull out only the columns you need.

What is Arrow format?

Apache Arrow is a columnar in-memory format designed for zero-copy reads and high-speed exchange between data systems.

Related Tools