Skip to content

No account required

Start a project

7 min · 4 sections

How to tell a real finding from a coincidence.

Most findings that fall apart do so for one of four reasons, and all four can be checked in minutes.

By Data Analysis App team

How do I tell a real finding from a coincidence?

Ask four questions of any number that surprised you. Who is missing from the data, since whatever excluded them is often the thing being measured. How many comparisons were made, because one striking result out of twenty is what noise looks like. What happens with the largest record removed. And whether it survives a different but equally defensible definition.

One evidence card separated from a field of possible signals by three verification layers.
One evidence card separated from a field of possible signals by three verification layers.

Section 01

Who is missing from the data?

The most dangerous bias is in the rows that are not there. Survey respondents who abandoned part-way, customers who churned before the export window, transactions that failed and were never recorded.

The question is not whether your sample is large. It is whether the people missing from it differ systematically from the people in it, and they usually do, because whatever caused them to be missing is often the thing you are measuring.

If you cannot say who is absent, you cannot say what the number means. That is not a counsel of despair; it is a prompt to go and find out, which is usually possible.

A survey where unhappy new customers abandoned at twice the rate of everyone else will show that new customers are happier than they are.

Section 02

How many things did you compare?

If you split your data twenty ways and one split shows a striking difference, that is roughly what you would expect from noise alone. The more comparisons you make, the more likely at least one looks remarkable.

This is not an argument against exploring. It is an argument for remembering how many stones you turned over, and for treating the one that looked interesting as a hypothesis to test rather than a conclusion to report.

The honest version of this finding is: "of the twenty segments we looked at, this one stood out: worth checking on fresh data."

Section 03

What happens if you remove the largest record?

A single large customer, order or respondent can carry an entire trend. Removing the biggest one and re-running is the fastest robustness check that exists, and it takes seconds.

If the finding survives, you have learned something real. If it vanishes, you have learned something more useful: your finding is about one record, which is a different and often more actionable story.

Section 04

Would it survive a different definition?

Revenue net or gross of refunds. Active users in the last 7 days or the last 30. A month by calendar or by billing period. Each choice is defensible, and a finding that only holds under one of them is a finding about the definition rather than about the business.

Write down the definition you used before you look at the result, not after. The temptation to adjust it once you have seen the number is very strong and completely invisible in the final report.

Every finding in a Data Analysis App project carries its definition, its filters, and what it does not establish, because a conclusion shipped without its limits is the failure this all guards against.

See it on a real project

The apparent 0.42-point gap fell to 0.07 once partial responses were included.

Is the difference between customer groups real?

Keep reading

Every check in this guide runs on every upload.

Drop the file in and the problems described here are tested before anything is reported: with the rows behind each repair kept, so you can see exactly what changed.

  • Subtotal rows detected and excluded
  • Duplicates matched on identity, not on whole rows
  • Text-typed numbers found, coerced and counted

No account needed to start. You only pay when you like what you see.