The .parquet file.
- Layout
- Columnar, with row groups
- Types
- Declared in the file, including decimals and timestamps
- Compression
- Snappy, gzip or zstd, per column
- Typical size
- Roughly a fifth of the equivalent CSV
- Human readable
- No
- Read here as
- Direct, with the file's own types kept
- Status
- Ready
- Fidelity
- Read directly
- Largest file
- 25 MB free, 250 MB once a project is bought
What it actually is
Read columnwise, with the file's own types kept.
Parquet stores a table one column at a time rather than one row at a time. Every value in a column has the same type and usually similar values, which makes it compress extremely well and makes reading three columns out of two hundred cost almost nothing.
It also carries a schema. A date is a date, a decimal is a decimal with a stated precision, and a string that looks like a number stays a string. Everything a CSV loses at the boundary is preserved, which removes an entire class of analytical error before any analysis starts.
The trade is that it is not human readable and not editable in a spreadsheet. It is an interchange format between systems, which is exactly what an export to an analysis tool is.
Upload one
The analysis runs before there is anything to pay for.
You see what it found, and the evidence behind it, first.
No account needed to start. You only pay when you like what you see.
What goes wrong, and what is done about it
2 named failures, each with the repair.
Timestamps without a timezone
Parquet distinguishes a timestamp that is a point in time from one that is a wall-clock reading, and many writers emit the second while the reader assumes the first. A day boundary then moves by hours, which is enough to shift a daily total into the wrong day.
What is done: The declared logical type is used rather than inferred, and where a file stores local timestamps with no zone that is recorded as an assumption you can see and override.
A folder of files that is really one table
Warehouses write partitioned datasets: a directory of many .parquet files, often with the partition key encoded in the directory name rather than in the data. Uploading one part gives you a slice with a missing column and no indication that either is true.
What is done: Export the dataset as a single file, with the partition keys written into it as columns. A zipped directory arrives as one table per part with the folder path discarded, and nothing concatenates the parts.
Questions people ask
About this format, not about the product.
How do I get a Parquet file out of my database?
Most warehouses export it directly: BigQuery, Snowflake, Redshift and DuckDB all do. From a local CSV, DuckDB will convert one in a single statement. It is worth the step for anything above a few hundred thousand rows.
Is Parquet better than CSV for this?
Meaningfully, yes. Types survive, so leading zeros, long identifiers and dates arrive intact, and the file is small enough that a dataset which exceeds the upload limit as CSV usually fits comfortably.
Can I mix Parquet with spreadsheets in one project?
Yes, and every file is read whatever format it arrived in. The analysis runs on one table, the largest, so a Parquet transaction export and an Excel cost sheet are not joined together. Where the two look related, that is reported as a candidate relationship with a confidence.
Keep reading
The jobs people do with this file, and the formats beside it.