Skip to content

No account required

Start a project

6 min · 4 sections

Analysis for people running surveys.

Most survey findings that fall apart do so because of who answered, not because of what they said.

By Data Analysis App team

Why do my survey results change when I include partial responses?

Because the people who abandon a survey are not a random sample of the people who start one. Dropout usually correlates with the thing being measured, so excluding partial responses biases every comparison in the same direction. Check completion by segment before reporting any difference between groups; if dropout tracks the answer, the headline is measuring dropout.

Section 01

Completion bias beats sample size

A survey with two thousand responses and a 40% completion rate among unhappy respondents is less trustworthy than one with four hundred responses and even completion. Size does not fix a systematic absence.

The check is straightforward: compare who finished against who started, split by the thing you are measuring. If dropout correlates with the answer, your headline comparison is measuring dropout.

In the worked example, the tenure gap the survey was run to test disappeared entirely once partial responses were included.

Section 02

Reverse-coded questions quietly break the average

Negatively worded items have to be recoded before anything is averaged. That much is standard, but it is worth checking afterwards whether they still behave, because respondents frequently answer them as if they were worded positively.

An item that correlates oddly with the rest of the scale after recoding is not measuring what its neighbours measure, and including it drags the overall score toward noise.

Section 03

Small segments should be counts, not percentages

Below about thirty responses, one person moves a rate by more than the effect being measured. Reporting "Enterprise satisfaction fell 8 points" when Enterprise is twenty-four people is reporting one person changing their mind.

Counts are less satisfying and more honest, and they make the sample size impossible to overlook.

Section 04

Free text is evidence, not decoration

Written answers are usually summarised separately from the scores, which wastes both. A theme appearing in forty comments means something different depending on whether those forty people rated you two or five.

Grouping the text by theme and showing the mean score for each theme is the best thing you can do with it.

See it on a real project

The apparent 0.42-point gap fell to 0.07 once partial responses were included.

Is the difference between customer groups real?

Keep reading

Every check in this guide runs on every upload.

Drop the file in and the problems described here are tested before anything is reported: with the rows behind each repair kept, so you can see exactly what changed.

  • Subtotal rows detected and excluded
  • Duplicates matched on identity, not on whole rows
  • Text-typed numbers found, coerced and counted

No account needed to start. You only pay when you like what you see.