Skip to content
Autofiller
Log inStart for free
Back to blog
Cover image reading How to Pilot Test a Survey, Before you launch it

How to Pilot Test a Survey Before You Launch It

A survey can fail in two quite different ways. People can misread the questions, so the answers you collect measure something other than what you intended. Or the survey itself can malfunction, with a branch that sends respondents to the wrong page, a required question nobody can answer, or a results sheet that silently drops a column. Pilot testing is how you find both kinds of problem while fixing them is still cheap, before the first real respondent opens the link.

Survey researchers usually treat these as two separate activities. The American Association for Public Opinion Research describes pretesting as understanding respondents' thought processes, including how they interpret each question and how they arrive at an answer, typically through cognitive interviews with people similar to those who will take the survey. It describes a pilot test as a check that all of the survey procedures, from recruiting respondents and administering the questionnaire to cleaning the data, work as intended. Pew Research Center calls pretesting an essential step in questionnaire design, especially when questions are being asked for the first time.

The two answer different questions, and neither substitutes for the other. A questionnaire can be perfectly understood and still break on a phone, and it can run flawlessly while asking something people interpret in three different ways. What follows covers how to pretest the wording, how to test the mechanics, and how to run a small pilot that ties everything together.

Pretest the questions with real people

A cognitive interview is a one-to-one session in which someone from your target audience works through the questionnaire while you learn how they are reading it. Some interviewers ask participants to think aloud as they answer. Others let them answer normally and then probe afterwards with questions such as what a particular phrase meant to them, what period of time they were thinking about, or how they chose between two nearby options. Both approaches aim to surface the gap between what you meant and what a respondent understood.

The people you interview matter more than how many of them there are. If the survey is for nurses, pretest it with nurses, because a colleague on your own team already knows what every question is supposed to mean and will fill the gaps without noticing. Run the interviews in small rounds, revise the wording between rounds, and then check whether the revision fixed the problem or simply moved it somewhere else.

Listen for hesitation, requests to reread a question, and answers that are sensible in themselves but respond to a different question from the one you asked. Pew's guidance gives useful examples of what to look for. Double-barreled questions, which ask about two things at once, are hard to answer and produce responses that are hard to interpret. Agree-or-disagree statements invite acquiescence, the tendency of some respondents to agree regardless of content. The order of answer options can also nudge people, since respondents to self-administered surveys tend to favor options near the top of a list. A pretest is where these problems show up as puzzled faces rather than as a skewed chart after the survey has closed.

Test the mechanics before anyone else sees the form

Once the wording has settled, the survey needs to work as software. Start by taking it yourself on the devices your respondents will use, including a small phone screen, and try to break it. Leave required questions blank, enter text where a number is expected, go back a page and change an earlier answer, and follow each branch of any skip logic. Keep a simple table of every route through the survey and tick each one off as you walk it, because it is easy to test the common path five times and the unusual one never.

Manual walkthroughs have limits. They are slow, they rarely cover every combination of answers, and they produce only a handful of rows in the results, which is not enough to show whether the export, the analysis script, or any connected automation copes with a realistic volume of responses. This is where generated test submissions earn their place. Autofiller was built for this step. You point it at a copy of a form you own or are authorized to test, choose how many responses to submit, and set answers to vary randomly or according to weights you pick, so that even rarely chosen options and branches get exercised. When the batch finishes, you can check that every branch received responses, that nobody who took one route has answers recorded for questions on another, and that your downstream spreadsheet and scripts handle the load.

Treat those submissions strictly as test data. They are generated answers, not respondents, and they tell you nothing about what real people think. Run them against a copy of the form, or clear them out completely before launch, and never mix them into the responses you analyze. Autofiller's Terms limit it to testing forms you own or are authorized to test and prohibit using it to skew or interfere with real surveys and research. Used this way, generated responses complement a pilot with real people rather than replacing it.

Run a small pilot and decide what to change

A pilot is a dress rehearsal with a small number of real respondents drawn from the same population and recruited the same way as the main sample. It checks whether the full procedure holds together, from the invitation through to a finished table of results. A workable sequence looks like this.

  1. Freeze a version of the questionnaire and record exactly what the pilot group will see, so that any change you make afterwards can be traced.
  2. Recruit a small group through the same channel you plan to use for the launch, because a pilot sent to colleagues will not reveal problems with the invitation or the audience.
  3. Watch completion times and the points where people abandon the survey, and read every open-ended answer rather than skimming a sample.
  4. Clean the pilot data and run your planned analysis on it in full, including any weighting, recoding, and charts, even though the numbers themselves are too few to interpret.
  5. List every problem you found, fix the questionnaire, and repeat the mechanical checks on the revised version before you launch.

Two judgment calls come up at the end of almost every pilot. The first is whether pilot responses can be kept. If you changed a question's wording, answer options, or position, the pilot answers to that question were collected under different conditions and generally should not be pooled with the main data. The second is when to stop iterating. There will always be one more phrasing to try, so aim for a questionnaire that people understand the same way, that works on every device and path, and whose data flows cleanly into your analysis, and then launch it.

Pilot testing takes some time up front, but it is far cheaper than discovering after a thousand responses that a key question was misread or that a branch skipped half of your sample. If your survey relies heavily on branching, the companion guide on designing and testing skip logic goes further into mapping and verifying every path, and the guide to testing a form's data pipeline covers what happens to responses after they are submitted.

Mira Lang
Growth & Experimentation Lead, Autofiller

Growth & experimentation lead focused on sustainable, user-centric growth through clean UX, sharp messaging, and experiments that teach you something.