SmartQueryTools

Convert Case in Parquet Files Online

Convert text columns to UPPERCASE, lowercase, or Title Case in Parquet files directly in your browser. Apply case conversion to any or all text columns at once — no upload required.

How to convert Case in Parquet files

  1. Drop your file onto the upload area. Every text column is ticked by default and the first 200 rows are shown.
  2. Choose the case type: UPPERCASE, lowercase or Title Case. One case type applies to every ticked column in a run.
  3. Adjust the columns under Apply to columns. Only text columns are listed. The All and None links tick or clear them all.
  4. Click Convert Case, check the preview, and download the file in the same format.

Your file is processed locally in your browser and is never uploaded. The free limit is 50 MB per file; larger files work if your device has the memory for them.

Worked example

Newsletter sign-ups were merged from two web forms, and email addresses arrive in mixed case. The mailing tool treats Jo.Bloggs@Example.COM and jo.bloggs@example.com as different contacts.

Input (Parquet)

Parquet file (binary, columnar) — shown as a table with its schema

emailcountryvisits
Jo.Bloggs@Example.COMNZ3
mika.k@example.fifi1
ANNA.W@EXAMPLE.DEDE7
r.osei@example.ghGH2

Schema: email VARCHAR, country VARCHAR, visits BIGINT

Settings

  • Case type: lowercase
  • Apply to columns: email only (country unticked; visits is a number and is not listed)

Result

emailcountryvisits
jo.bloggs@example.comNZ3
mika.k@example.fifi1
anna.w@example.deDE7
r.osei@example.ghGH2

Only the email column changes. Addresses that were already lowercase are left as they were. The country column was unticked, so "fi" is still lowercase. A second run with UPPERCASE and only country ticked would fix it. The visits column is a number and passes through unchanged.

Working with Parquet files

Parquet columns keep their declared types, so the column list shows exactly the string columns. Integer, decimal, timestamp and boolean columns are not listed. The converted columns are written back as strings with the same names, and the rest of the schema is untouched. That makes it safe to normalise a single key column, such as a country or status code, in a large file before joining it to another table. It is also a quick way to normalise a region or category column before grouping the file.

Low-cardinality string columns in Parquet are often dictionary-encoded, and after conversion values like "Active", "ACTIVE" and "active" collapse to one value. A later Count By or Unique Values run on that column then reports one value instead of three. Struct and list columns cannot be case converted even when they contain strings, so they are not listed. Flatten the data first to convert a nested field.

Frequently Asked Questions

Does converting case change my Parquet column types?

No. Text columns stay strings and every other column keeps its original type. Only the letters inside the ticked columns change.

Can I convert strings inside a Parquet struct column?

No. Only top-level text columns can be converted. Flatten the struct into separate columns first, or use the SQL workspace with upper() or lower() on the nested field.

Can I apply different cases to different columns?

Not in one run. Each run applies one case type to all ticked columns. Run the tool again on the downloaded file with the other columns ticked.

Which columns can I convert?

Text columns only, and those are the only ones listed. Number, date and boolean columns have no letters to convert, so they are left unchanged.

Does it trim spaces or fix accents?

No. It only changes letter case. Accented letters are converted correctly, for example é to É, but spaces and punctuation are left alone. Use Trim Whitespace for spaces.

Related Tools