Lesson 13: Surveys
Lesson 13: Surveys
Lesson 12 gave you a way to go deep with a small number of people. This lesson gives you the complementary tool: a way to check, across a much larger population, whether what you learned from those few conversations actually generalizes. A survey is a structured research instrument administered to many respondents at once, designed to answer quantitative questions — how many, how much, how often — that a handful of interviews cannot reliably answer on their own, directly extending Lesson 11's qualitative/quantitative distinction into a specific, buildable method.
Surveys are also, unfortunately, one of the easiest research instruments to build badly while feeling confident you've built them well. A leading question (Lesson 11), a biased sample, or a poorly worded scale can produce a clean-looking chart of numbers that feels authoritative precisely because it's quantitative — while actually encoding all the same distortions a bad interview would, just with more decimal places. This lesson exists to make sure the numbers you get out of a survey are actually measuring what you think they're measuring.
Learning Objectives
- 1
Identify when a survey is the appropriate research method, versus when an interview or behavioral data is better suited to the question.
- 2
Distinguish well-constructed survey questions from leading, double-barreled, or ambiguous ones.
- 3
Explain the purpose and correct interpretation of common quantitative scales (Likert scales, NPS) and their known limitations.
- 4
Identify sampling bias in survey distribution and explain its effect on the validity of results.
- 5
Apply a basic checklist for evaluating whether a completed survey's findings are trustworthy enough to inform a decision.
Lesson 11 (User Research) and Lesson 12 (Customer Interviews). This lesson assumes fluency with the stated-versus-revealed-preference distinction and the qualitative/quantitative complementary-methods framework, and extends both into the specific mechanics of survey design.
When a Survey Is the Right Tool
When a Survey Is the Right Tool
Recall Lesson 11's framework: quantitative methods are well suited to how-many/how-much questions, where a large, representative sample is needed to support a claim about prevalence or scale. Surveys are the most common quantitative instrument for this purpose. A survey is the right tool when:
You already have a well-formed hypothesis (often generated through qualitative interviews, per Lesson 12) and need to know how widespread it is across the broader user base.
You need to measure something efficiently across a large population that would be prohibitively expensive to interview individually.
You need a baseline measurement to track over time (e.g., satisfaction tracked quarterly).
A survey is the wrong tool when you don't yet know what you're looking for — a survey can only ask about things the designer already thought to ask about, unlike an open-ended interview, which can follow an unexpected thread wherever it leads. This is why, per Lesson 11's complementary-methods framework, surveys typically work best after qualitative research has generated a specific hypothesis to test, not as a substitute for that earlier exploratory work.
Constructing Good Survey Questions
Constructing Good Survey Questions
Several specific, well-documented question-design failures distort survey results:
Leading questions (from Lesson 11): "How much do you love our new dashboard?" presupposes a positive reaction and biases responses toward it, compared to a neutral "How would you describe your experience with the new dashboard?"
Double-barreled questions: asking about two things at once, such as "How satisfied are you with our product's speed and ease of use?" — a respondent might feel very differently about speed than about ease of use, and a single combined answer obscures which one actually drove their response.
Ambiguous or vague terms: asking "How often do you use this feature regularly?" without defining "regularly" leaves each respondent to apply their own inconsistent standard, making aggregated results difficult to interpret meaningfully.
Unbalanced or missing answer options: a satisfaction scale that includes three positive options and only one negative option biases the aggregate distribution toward apparent satisfaction regardless of true sentiment.
Common Scales: Likert and NPS, and Their Limitations
Common Scales: Likert and NPS, and Their Limitations
Two widely used quantitative scales deserve specific attention, both for their usefulness and their well-documented limitations:
Likert scales (typically a 5- or 7-point agreement or satisfaction scale, e.g., "Strongly Disagree" to "Strongly Agree") allow respondents to express degree, not just direction, of sentiment. Their limitation: individual respondents interpret the middle and endpoint options differently (some respondents rarely select extreme options regardless of true sentiment, a pattern sometimes called central tendency bias), meaning cross-respondent comparisons carry more noise than the clean-looking numeric output might suggest.
Net Promoter Score (NPS), based on the single question "How likely are you to recommend this to a friend or colleague?" on a 0–10 scale, is widely used as a simple, trackable proxy for overall sentiment. Its well-documented limitation, directly connected to Lesson 11's stated-preference warning: NPS asks about a hypothetical future action (a recommendation that may never actually happen), not a real, revealed behavior, and a single number provides no insight into why a respondent scored as they did — precisely the kind of why/how gap that quantitative methods, per Lesson 11's framework, are not well suited to answer alone.
Both scales are useful for tracking directional change over time and for flagging where deeper qualitative investigation (returning to Lesson 12's interview techniques) is warranted — but neither should be treated as a complete, self-sufficient answer to "how are we doing and why."
Sampling Bias in Survey Distribution
Sampling Bias in Survey Distribution
Directly extending Lesson 11's representativeness concern, how a survey is distributed determines who is even eligible to respond, and this can silently and severely distort results before a single question is even answered. Common patterns:
In-app pop-up surveys reach only currently active users, systematically excluding churned users, users who found the product too frustrating to keep using, and non-users entirely — precisely the populations most likely to reveal the most serious, unaddressed problems.
Email surveys sent to an existing customer list exclude prospects who never converted and churned customers who may have been removed from the active list.
Surveys promoted through the product's own social media or community channels oversample the most engaged, enthusiastic segment of the user base, who are demonstrably not representative of the broader population.
A survey with a large number of responses can still produce badly misleading findings if the distribution method silently excluded the population segment most relevant to the question being asked — a large, biased sample is not more trustworthy than a small, biased one; it simply produces a false sense of confidence at a larger scale.
A Trustworthiness Checklist for Completed Surveys
A Trustworthiness Checklist for Completed Surveys
Extending Lesson 11's general trustworthiness checklist to the specific case of a completed survey:
Who was excluded by the distribution method, and does that exclusion matter for the specific question being asked?
Were questions neutral, single-barreled, and clearly defined, or did any suffer from the failures described above?
Does the survey measure stated preference (e.g., NPS, hypothetical satisfaction) or something closer to revealed preference (e.g., a question about actual, specific past behavior, similar in spirit to Lesson 12's past-behavior interview questions)?
Is the sample size and response rate sufficient to support the specific claim being made, and has non-response bias (are the people who didn't respond systematically different from those who did) been considered?
Was the survey used to test an existing hypothesis, generated through prior qualitative work, or was it used to explore blindly without a clear prior hypothesis — a pattern more likely to produce noisy, hard-to-interpret results?
Common Mistakes to Avoid
Using a survey to explore an open-ended question with no prior hypothesis
Surveys can only ask about what the designer already thought to ask, making them poorly suited to genuinely open-ended exploration — that work belongs to qualitative interviews (Lesson 12), with the survey following afterward to test a specific, now-formed hypothesis at scale.
Writing double-barreled questions
Combining two distinct dimensions ("speed and ease of use") into a single question makes the resulting number impossible to interpret cleanly, since a respondent's answer conflates two potentially very different underlying sentiments.
Distributing a survey only through channels that reach already-engaged users
In-app pop-ups, existing customer email lists, and community channels all systematically exclude churned users, frustrated non-adopters, and prospects — often the exact population whose perspective would be most revealing.
Treating NPS or a single satisfaction number as a complete answer
A single quantitative score can flag that something is wrong, or track directional change over time, but provides no insight into why — treating it as sufficient on its own skips the necessary qualitative follow-up.
Over-interpreting small differences in Likert scale averages
Given individual variation in how respondents use rating scales (central tendency bias and similar effects), small differences between, say, a 3.8 and a 4.1 average often reflect noise rather than a meaningful, decision-worthy difference in underlying sentiment.
The Survey Validity Chain
This lesson's mental model is the Survey Validity Chain — a sequence of checkpoints, any one of which can silently invalidate an otherwise clean-looking result.
A break at any single link invalidates the chain regardless of how clean the final numbers look — a perfectly worded, neutral question distributed only to the most engaged users is still compromised at the distribution link, and a representatively distributed survey full of leading questions is compromised at the question-design link. Checking the whole chain, not just the parts that are easiest to verify (the final numbers), is the core discipline this lesson asks for.
Key Takeaway: How will you apply "The Survey Validity Chain" when evaluating trade-offs in your product decisions?
Ready to test your product judgment?
Take the interactive practice quiz for Lesson 13 and build your skill radar dashboard.