Parse URL Column in Parquet Files Online
Parse URL columns in Parquet files and extract host, path, query string, and fragment into separate columns — directly in your browser, no upload required.
How to parse URL Column in Parquet files
- Drop your file onto the upload area. The tool picks the first text column whose name contains url, link, href or uri as the URL column.
- Check the URL column. Only text columns are listed.
- Choose which parts to extract: Host, Path, Query and Fragment. Host and Path are selected by default.
- Optionally type an Output column prefix. By default new columns are named after the URL column, such as landing_url_host.
- Click Parse URLs, check the preview, then download the file with the new columns added at the end.
Your file is processed locally in your browser and is never uploaded. The free limit is 50 MB per file; larger files work if your device has the memory for them.
Worked example
A marketing analyst has a sessions export with full landing page URLs and wants to report by site and page, and see which campaign tags were used.
Input (Parquet)
Parquet file (binary, columnar) — shown as a table with its schema
| session_id | landing_url | pageviews |
|---|---|---|
| s-8841 | https://www.example.com/pricing?utm_source=newsletter&utm_campaign=spring | 4 |
| s-8842 | https://blog.example.com/2026/05/launch-notes?ref=twitter | 1 |
| s-8843 | https://www.example.com/signup?ref=partner-17 | 6 |
| s-8844 | https://shop.example.co.uk/cart?step=2 | 3 |
Schema: session_id VARCHAR, landing_url VARCHAR, pageviews BIGINT
Settings
- URL column: landing_url
- Extract parts: Host, Path, Query
- Output column prefix: (blank, uses landing_url_)
Result
| session_id | landing_url | pageviews | landing_url_host | landing_url_path | landing_url_query |
|---|---|---|---|---|---|
| s-8841 | https://www.example.com/pricing?utm_source=newsletter&utm_campaign=spring | 4 | www.example.com | /pricing | utm_source=newsletter&utm_campaign=spring |
| s-8842 | https://blog.example.com/2026/05/launch-notes?ref=twitter | 1 | blog.example.com | /2026/05/launch-notes | ref=twitter |
| s-8843 | https://www.example.com/signup?ref=partner-17 | 6 | www.example.com | /signup | ref=partner-17 |
| s-8844 | https://shop.example.co.uk/cart?step=2 | 3 | shop.example.co.uk | /cart | step=2 |
The original columns are kept and one column per selected part is appended, named with the prefix plus the part. The host keeps subdomains, so www.example.com and blog.example.com stay separate. The query is returned as one string. Individual parameters such as utm_campaign are not split into their own columns, so use Split Column on & and then on = if you need them.
Working with Parquet files
Web analytics and log pipelines often write URLs into Parquet as plain string columns, which appear in the URL column list straight away. Columns stored as binary without a string annotation load as BLOB and are not listed. Cast them to VARCHAR first if your writer did that. A single request table can hold millions of rows, and parsing runs as one pass in your browser tab, so memory rather than time is the usual limit.
Each new column is written as a string column in the output Parquet file, alongside the original columns and their types. Host values repeat heavily, and Parquet dictionary-encodes repeated strings, so adding a host column costs far less space than its row count suggests. If you only need the host for grouping, parse it and then use Columns to drop the full URL before you share the file. Full URLs with query strings can carry session tokens or email addresses, so dropping them also reduces what you expose.
Frequently Asked Questions
Why is my URL column not listed for this Parquet file?
It is probably stored as binary (BLOB) rather than a string. Cast it to VARCHAR with Cast Columns, then parse it.
What type are the extracted columns in the Parquet output?
They are string (VARCHAR) columns. The original columns keep their types.
Which URL parts can the tool extract?
Host, path, query string and fragment. It does not split out the scheme, port or individual query parameters. Use Split Column or Regex Extract on the query column for single parameters such as utm_source.
How are the new columns named?
Prefix plus part name, for example landing_url_host. The default prefix is the URL column name followed by an underscore. Characters other than letters, digits and underscores in a custom prefix are replaced with underscores.
Is the original URL column changed?
No. The original column stays as it was, and the extracted parts are appended as new columns at the end.
Related Tools
Count Values in Parquet Files Online
Group and count rows by any column in Parquet files directly in your browser. Sort by frequency or value to find the most common entries — no upload required.
Split Column in Parquet Files Online
Split a column into multiple columns by delimiter in Parquet files directly in your browser. Turn "First Last" into separate first and last name columns — no upload required.
Aggregate Parquet Files Online
Group and aggregate Parquet files by any column directly in your browser. Calculate sum, average, min, max, and count for any numeric column — no upload required.
Parse URL Column in CSV Files Online
Parse URL columns in CSV files and extract host, path, query string, and fragment into separate columns — directly in your browser, no upload required.
Parse URL Column in Excel Files Online
Parse URL columns in Excel files and extract host, path, query string, and fragment into separate columns — directly in your browser, no upload required.
Parse URL Column in JSON Files Online
Parse URL columns in JSON files and extract host, path, query string, and fragment into separate columns — directly in your browser, no upload required.
Parquet Viewer Online
View and inspect Parquet files directly in your browser. Browse rows, check column names and data types — no upload required, your data stays on your device.
Convert Parquet to CSV Online
Convert Parquet files to CSV format directly in your browser. No upload required — your data never leaves your device.