SmartQueryTools

Hash & Anonymise Columns in Parquet Files Online

Anonymise or pseudonymise columns in Parquet files by replacing values with MD5, SHA-256, or DuckDB hashes — directly in your browser. Useful for GDPR compliance and sharing data without exposing PII — no upload required.

How to hash & Anonymise Columns in Parquet files

  1. Drop your file onto the upload area. The first 200 rows are shown with the row and column count.
  2. Choose a Hash algorithm: MD5 (32-character hex), SHA-256 (64-character hex) or DuckDB hash (a 64-bit integer written as text).
  3. Choose an Output mode. Replace swaps each selected column for its hash. Append keeps the original and adds a new column such as email_md5 at the end.
  4. Tick the columns to hash. Nothing is ticked by default.
  5. Click Hash Columns, check the preview, then download the file in the same format.

Your file is processed locally in your browser and is never uploaded. The free limit is 50 MB per file; larger files work if your device has the memory for them.

Worked example

A research team wants to share survey scores with an outside analyst without revealing respondent emails, while still showing which answers came from the same person.

Input (Parquet)

Parquet file (binary, columnar) — shown as a table with its schema

respondent_idemailage_bandnps_score
R-1001maria.lopez@example.org35-448
R-1002j.okafor@example.net25-346
R-1003maria.lopez@example.org35-449
R-1004NULL45-547

Schema: respondent_id VARCHAR, email VARCHAR, age_band VARCHAR, nps_score BIGINT

Settings

  • Hash algorithm: MD5
  • Output mode: Replace original column with hash value
  • Columns to hash: email

Result

respondent_idemailage_bandnps_score
R-10010774b6f01f630684b8643132792f1c4835-448
R-100204010e734bdab7aa07d0f05cd693ef8125-346
R-10030774b6f01f630684b8643132792f1c4835-449
R-1004NULL45-547

These are the real MD5 values of each address. R-1001 and R-1003 get the same hash because they share an email, so the analyst can still group by person without seeing the address. The missing email stays empty, because MD5 and SHA-256 of NULL are NULL. The column keeps its name and position in Replace mode.

Working with Parquet files

In Replace mode the hashed column changes type. An INT64 customer_id becomes a string column of hex digests in the output Parquet file, so downstream code that expects an integer key must be updated. In Append mode the original column keeps its type and a new string column is added. The file is rewritten with the same compression settings the engine uses by default, not the original writer's.

The text that gets hashed follows the stored type. A DECIMAL(10,2) value 1.5 is hashed as "1.50" and a timestamp as "2026-01-04 10:15:00.123456" with its fractional seconds. The DuckDB hash option works on the typed value directly, so the integer 1001 and the string "1001" give different results. Use MD5 or SHA-256 when the hashes must match between a Parquet file and a CSV or database export. A warehouse query such as md5(CAST(customer_id AS VARCHAR)) produces the same digest for the same text.

Frequently Asked Questions

Does hashing change my Parquet column types?

In Replace mode, yes. Each hashed column becomes a string column. In Append mode the originals keep their types and new string columns are added.

Will the DuckDB hash of an ID match between Parquet and CSV files?

Not always. That hash depends on the column type, and an ID stored as an integer in one file and text in the other gives different values. Use MD5 or SHA-256 for cross-file matching.

Is hashing the same as anonymisation?

No. It is pseudonymisation. The same input always gives the same hash, and no salt is added. Anyone with a list of likely values, such as known email addresses, can hash them and look for matches. Treat hashed personal data as still personal.

Which algorithm should I choose?

SHA-256 for anything shared outside your team. MD5 is shorter and fine for internal join keys, but it is no longer considered secure. The DuckDB hash option is fast but not a standard algorithm, so other systems cannot reproduce it.

Can I add a salt before hashing?

Not in this tool. Use the SQL Query tool instead, with an expression such as sha256('my-secret-salt' || email). Keep the salt private and use the same one for every file you want to join.

Related Tools

Manage Columns in Parquet Files Online

Drop or select specific columns from Parquet files directly in your browser. No upload required.

Deduplicate Parquet Files Online

Remove duplicate rows from Parquet files instantly in your browser. No upload, no server — 100% private.

Count Values in Parquet Files Online

Group and count rows by any column in Parquet files directly in your browser. Sort by frequency or value to find the most common entries — no upload required.

Hash & Anonymise Columns in CSV Files Online

Anonymise or pseudonymise columns in CSV files by replacing values with MD5, SHA-256, or DuckDB hashes — directly in your browser. Useful for GDPR compliance and sharing data without exposing PII — no upload required.

Hash & Anonymise Columns in Excel Files Online

Anonymise or pseudonymise columns in Excel files by replacing values with MD5, SHA-256, or DuckDB hashes — directly in your browser. Useful for GDPR compliance and sharing data without exposing PII — no upload required.

Hash & Anonymise Columns in JSON Files Online

Anonymise or pseudonymise columns in JSON files by replacing values with MD5, SHA-256, or DuckDB hashes — directly in your browser. Useful for GDPR compliance and sharing data without exposing PII — no upload required.

Parquet Viewer Online

View and inspect Parquet files directly in your browser. Browse rows, check column names and data types — no upload required, your data stays on your device.

Convert Parquet to CSV Online

Convert Parquet files to CSV format directly in your browser. No upload required — your data never leaves your device.