Skip to content
Autofiller
Log inStart for free
Back to blog
Cover image reading Survey Response Quality, Catching careless and bogus responses

Survey Response Quality: How to Catch Careless and Bogus Responses

Every survey collects some responses that do not reflect what the respondent actually thinks. Some come from people who are rushing, bored, or distracted and click through without reading, which researchers often call careless or insufficient-effort responding. Others come from people who should not be in the sample at all, such as someone who misrepresents who they are to collect an incentive. Both kinds of response look like ordinary rows in a spreadsheet, and both can change your conclusions if you do not look for them.

It is tempting to assume that bad responses are random and simply add noise that washes out in a large sample. Research suggests otherwise. In a study of more than 60,000 online interviews, Pew Research Center found that widely used opt-in sources contained small but measurable shares of bogus respondents, about 4% to 7% depending on the source. Those respondents did not answer at random. They tended to choose positive answers, which introduced a systematic bias into estimates such as approval ratings.

The good news is that response quality is something you can plan for. This guide covers what low-quality responses tend to look like, how to design a questionnaire that makes them easier to detect, and how to test your screening rules before the real data arrives.

What low-quality responses look like

Careless responding tends to leave patterns in the data. Some respondents finish far faster than anyone reading the questions could. Some choose the same answer down an entire grid of rating questions, a pattern usually called straightlining. Others give contradictory answers to two questions that ask about the same thing, fail an instruction embedded in the survey, or type the same meaningless string into every open-ended box. None of these signals is conclusive on its own, since a fast reader can be careful and a consistent answer can be genuine, but together they build a picture.

Researchers have studied which of these indicators work and when. Adam W. Meade and S. Bartholomew Craig examined several methods in a 2012 paper in Psychological Methods, including special items designed to catch inattention, consistency indices, multivariate outlier analysis, and response time. They found two distinct patterns of careless responding, one random and one not, and concluded that different indices are needed to catch each. Paul G. Curran's 2016 review in the Journal of Experimental Social Psychology brings these techniques together and offers recommendations for using them in practice.

Relying on a single check is risky. In the Pew study, 84% of bogus respondents passed an attention-check question and 87% passed a check for answering too quickly, so two of the most common screens missed most of the problem cases. Several signals checked together, each catching a different pattern, give you a much better chance than any one rule.

Design the survey so quality can be measured

Most quality checks depend on decisions you make before launch. If you want to flag speeders, you need a record of how long each response took, and you need a sense of how long a careful respondent takes, which a pilot test can give you. If you want to check consistency, you need at least one pair of questions that should agree. An instructed-response item, such as a question that asks respondents to select a specific option, gives you a direct measure of attention, and Meade and Craig recommend including items like this before data collection rather than trying to infer attention afterwards. An open-ended question is also useful, because nonsensical or off-topic text is one of the easier signals to read.

Question design affects how much careless responding you get in the first place. Long grids of agree-or-disagree statements are tiring and invite straightlining, and Pew notes that this format is prone to acquiescence bias, the tendency of some respondents to agree regardless of what a statement says. Shorter blocks, varied question formats, and a questionnaire no longer than it needs to be all make careful answering easier.

Write your screening rules down before you see any data. Decide which indicators you will use, what threshold counts as a flag, and how many flags lead to exclusion. Rules chosen after looking at the results are easy to bend, even unintentionally, towards the answer you were hoping for. When you report the findings, say how many responses you removed and why, so that readers can judge the effect for themselves.

Test your screening rules before the data is real

A screening rule is a piece of logic, and like any logic it can be wrong. A speed threshold might be set in the wrong units, a straightlining check might look at the wrong columns after a question was added, or an attention item might be coded so that the correct answer counts as a failure. It is much better to find these mistakes before launch than to discover them while deciding which real respondents to exclude.

The research literature offers a useful model here. In the second study of their paper, Meade and Craig simulated data with known random response patterns to measure how well each indicator detected them. The same idea works at a smaller scale for your own survey. If you submit test responses whose properties you already know, you can check that your rules flag what they are supposed to flag and pass what they are supposed to pass.

Autofiller can produce this kind of test batch for a copy of a form you own or are authorized to test. Its random mode fills in answers without any underlying preference, so a batch of those responses should trip consistency checks, and weighted answers let you push a batch heavily towards particular options to see how your straightlining rule behaves. A handful of careful test responses that you enter yourself gives you the clean comparison. Because Autofiller's CSV export records every generated answer with its timestamp, you can run your cleaning script on the test responses and compare the rows it flags with the batch you know you sent.

These responses are test data and nothing else. They are generated, they do not represent any person, and they must never be mixed into the dataset you analyze or used to increase a sample. Run them on a copy of the form, or delete them before launch, and keep them out of anything you report. Autofiller's Terms prohibit using it to skew or interfere with genuine surveys and research, and using generated answers as real data would be research misconduct.

Response quality is ultimately a design problem more than a cleanup problem. A questionnaire that is clear and reasonably short, a recruitment source you trust, screening rules agreed in advance, and checks that have been tested before launch will do more for your data than any amount of filtering after the fact. For the mechanics of getting clean responses from the form into your analysis, see the guide to testing a form's data pipeline.

Mira Lang
Growth & Experimentation Lead, Autofiller

Growth & experimentation lead focused on sustainable, user-centric growth through clean UX, sharp messaging, and experiments that teach you something.