Skip to content

No account required

Start a project

The .parquet file.

The format to ask for when a CSV export is too big, too slow or too lossy. It is all three of those problems solved by storing columns instead of rows.
Layout
Columnar, with row groups
Types
Declared in the file, including decimals and timestamps
Compression
Snappy, gzip or zstd, per column
Typical size
Roughly a fifth of the equivalent CSV
Human readable
No
Read here as
Direct, with the file's own types kept
Status
Ready
Fidelity
Read directly
Largest file
25 MB free, 250 MB once a project is bought

What it actually is

Read columnwise, with the file's own types kept.

Parquet stores a table one column at a time rather than one row at a time. Every value in a column has the same type and usually similar values, which makes it compress extremely well and makes reading three columns out of two hundred cost almost nothing.

It also carries a schema. A date is a date, a decimal is a decimal with a stated precision, and a string that looks like a number stays a string. Everything a CSV loses at the boundary is preserved, which removes an entire class of analytical error before any analysis starts.

The trade is that it is not human readable and not editable in a spreadsheet. It is an interchange format between systems, which is exactly what an export to an analysis tool is.

Upload one

The analysis runs before there is anything to pay for.

You see what it found, and the evidence behind it, first.

No account needed to start. You only pay when you like what you see.

What goes wrong, and what is done about it

2 named failures, each with the repair.

Timestamps without a timezone

Parquet distinguishes a timestamp that is a point in time from one that is a wall-clock reading, and many writers emit the second while the reader assumes the first. A day boundary then moves by hours, which is enough to shift a daily total into the wrong day.

What is done: The declared logical type is used rather than inferred, and where a file stores local timestamps with no zone that is recorded as an assumption you can see and override.

A folder of files that is really one table

Warehouses write partitioned datasets: a directory of many .parquet files, often with the partition key encoded in the directory name rather than in the data. Uploading one part gives you a slice with a missing column and no indication that either is true.

What is done: Export the dataset as a single file, with the partition keys written into it as columns. A zipped directory arrives as one table per part with the folder path discarded, and nothing concatenates the parts.

Questions people ask

About this format, not about the product.

How do I get a Parquet file out of my database?

Most warehouses export it directly: BigQuery, Snowflake, Redshift and DuckDB all do. From a local CSV, DuckDB will convert one in a single statement. It is worth the step for anything above a few hundred thousand rows.

Is Parquet better than CSV for this?

Meaningfully, yes. Types survive, so leading zeros, long identifiers and dates arrive intact, and the file is small enough that a dataset which exceeds the upload limit as CSV usually fits comfortably.

Can I mix Parquet with spreadsheets in one project?

Yes, and every file is read whatever format it arrived in. The analysis runs on one table, the largest, so a Parquet transaction export and an Excel cost sheet are not joined together. Where the two look related, that is reported as a candidate relationship with a confidence.

Keep reading

The jobs people do with this file, and the formats beside it.