The .dta file.
- Carries
- Variable labels, value label sets and dataset notes
- Missing values
- 27 distinct kinds, plain and .a through .z
- Dates
- Days, months or quarters since 1 January 1960, per variable format
- Versions
- Format 113 through 121 in common circulation
- Read here as
- Direct, with labels kept so the codebook does not need re-typing
- Status
- Ready
- Fidelity
- Read directly
- Largest file
- 25 MB free, 250 MB once a project is bought
What it actually is
Variable and value labels are kept, so the codebook does not have to be re-typed.
A .dta file holds the data alongside variable labels, value label sets, and a notes field that researchers frequently use to record how a variable was constructed. For reproducing someone else's analysis, that notes field is often the most valuable thing in the file.
Stata distinguishes 27 missing values: a plain one and .a through .z, conventionally used to record *why* something is missing. That distinction is real information and is destroyed by any export that maps them all to blank.
The format has several versions, and older files use a different string encoding. Both are read here, but a file round-tripped through an old version may have lost long strings on the way.
Upload one
The analysis runs before there is anything to pay for.
You see what it found, and the evidence behind it, first.
No account needed to start. You only pay when you like what you see.
What goes wrong, and what is done about it
2 named failures, each with the repair.
Extended missing values collapsing into one
A dataset that carefully records .a for "refused" and .b for "not applicable" loses the distinction the moment it is exported to a format with one kind of blank, and with it the ability to tell a non-response from an inapplicable question.
What is done: The same thing happens here: .a through .z arrive as an ordinary blank, and the reason a value is missing does not survive the read. Where that distinction is part of the analysis, recode it into a variable of its own before exporting.
Dates that are days since 1960
Stata's date origin is 1 January 1960, and the unit varies by variable format: days, weeks, months, quarters or years. Read as plain numbers, a 2024 date becomes a small integer and every time-based analysis is silently nonsense.
What is done: The variable's declared format supplies the unit, so a monthly date variable is read as months rather than days.
Questions people ask
About this format, not about the product.
Do I need Stata to use this?
No. The .dta file is read directly, including its value labels, and nothing in the analysis depends on having the software. Exporting to CSV first is what loses those labels, so upload the original rather than a conversion of it.
Which Stata versions are readable?
All the formats in common circulation. Where a file is from an old enough version that long strings were truncated on the way in, that is reported rather than presented as complete.
Are the dataset notes used?
No. The notes field is not read at all, so a note recording how a variable was derived does not arrive with the data. Paste anything that matters into a .txt and upload it alongside, which is read as context.
Keep reading
The jobs people do with this file, and the formats beside it.