How to open a Parquet file on a Mac

Nothing on macOS opens a .parquet file by default. It is a compressed, columnar, binary format, so a text editor shows you noise and Preview shows you nothing.

What a Parquet file actually is

Parquet stores data by column rather than by row, compresses each column separately, and keeps a footer describing the whole file. That is why it is small and fast to query, and also why you cannot just look at it.

The upside is that a good reader never has to load the whole thing. It can read the footer, find the columns you asked for, and skip everything else. A five gigabyte Parquet can answer a question in under a second.

A Parquet file open in Datapuddle on macOS, each column header drawing its own distribution above the rows.
A 630 MB Parquet, opened without an import step. Click to enlarge.

Ways to open one

DuckDB's command line (free)

The quickest no-install-ceremony route, if you are comfortable in a terminal:

brew install duckdb
duckdb -c "SELECT * FROM 'data.parquet' LIMIT 20"
duckdb -c "DESCRIBE SELECT * FROM 'data.parquet'"

No import, no conversion. It reads the file where it sits.

Python

pandas.read_parquet or pyarrow.parquet.read_table both work and are the right answer if the next step is code. They load into memory, so a file larger than your RAM needs row-group iteration and some care.

Convert it to CSV

Possible, and usually a mistake. You lose the types, you lose the compression, and a Parquet that was 600 MB can land as several gigabytes of text that Excel still will not open.

Doing it in Datapuddle

Datapuddle is a Mac app with DuckDB compiled in, so Parquet is a first-class file type rather than something you import.

  1. Double-click the file in Finder, or drag it onto the app icon.
  2. It opens at once — the footer is read, not the file.
  3. The grid holds the whole result, so you can scroll to the end of a forty-million-row file rather than paging through it.

Globs work too, so SELECT * FROM 'logs/*.parquet' treats a directory of files as one table. Every column header draws its own distribution, null share and distinct count, which saves writing the GROUP BY you were about to write.

Which to use

Datapuddle on the Mac App Store — 7 day trial, then $39.99 once. Apple silicon, macOS 14 or later. · Docs Guides · Home