---
title: Voice Surveys Get Twice the Words. Naive Ones Lose Half the Room.
description: The research on voice surveys cuts both ways: spoken answers are dramatically richer, and record-into-a-box voice surveys bleed respondents. The design choice between those outcomes is yours.
canonical: https://chatwisp.ai/blog/voice-surveys-spoken-vs-typed-responses
date: 2026-08-05
---

# Voice Surveys Get Twice the Words. Naive Ones Lose Half the Room.

The best study I've found on voice versus typed survey answers is Höhne, Gavras, and Claassen's 2024 experiment in Social Science Computer Review. They randomly assigned 1,001 German smartphone respondents to answer the same open-ended questions by typing or by voice. The voice answers were more than twice as long across every question; on one topic they averaged 56 words spoken versus 20 typed. Spoken answers also covered more ground, averaging 3.42 distinct topics against 2.06 for text on one question.\n\nSame study, other hand: 51% of the voice group broke off before answering, versus 24% for text, and item non-response in the voice condition ran around 35% against roughly 2% for typing.\n\nMost people cite half of this study, whichever half they're selling. Read together, the two findings say something more useful: speaking produces richer answers than typing, and being told to record yourself into a silent box drives people away. Those aren't contradictory results. They're the same result. Voice is a [conversation](/blog/what-is-a-conversational-survey)al medium, and the surveys in that experiment asked people to use it non-conversationally, leaving a voice memo for a machine that gave nothing back. The respondents who stayed gave more than typists ever do; the format made too few of them want to stay.\n\n## Why speech runs richer\n\nNothing mysterious here. People speak at 130 to 150 words a minute and thumb-type on a phone at a fraction of that, so the effort cost per word collapses. Typing forces summarizing; speech lets people think out loud, qualify, digress into the anecdote that turns out to be the finding. That's also why the researchers found spoken answers ranged over more topics: the marginal cost of one more thought is nearly zero when you're talking.\n\nThere's a flip side worth being honest about. The same study found typed answers had more varied vocabulary; speech is looser and more repetitive. If you need polished prose, voice won't give it to you. But survey research doesn't need polish. It needs the third topic the respondent wouldn't have bothered to type, and that's exactly what voice buys.\n\n## The dead-microphone problem\n\nSo why did half the voice group walk? Put yourself in the respondent's position. A form asks you to tap record and talk at a static question. No reaction, no sign anyone heard, no idea if two sentences is too little or too much. It's the survey equivalent of leaving a voicemail for a stranger, and most people hate leaving voicemails. The richness of speech comes from conversation, and there was no conversation on offer, just a microphone where a text box used to be.\n\nThis is my actual argument, and it cuts against a chunk of the voice-survey category: bolting voice recording onto a static survey takes the worst of both formats. You keep the form's deadness and add speech's awkwardness. The break-off numbers are what that combination deserves.\n\n## What voice needs to work\n\nA voice survey earns its richer answers when the respondent is talking *with* something. An [interviewer](/blog/ai-moderated-interviews-guide) that speaks the question, reacts to the answer, and asks the follow-up a person would ask turns the recording task into a phone call, which is a behavior everyone already knows. That interviewer also solves the how-much-do-I-say problem, because the conversation signals when to keep going. This is why we built ChatWisp's voice mode as a full [speaking interviewer with selectable voices and speech models](/knowledge-base/ai-interviewer/voices-and-models) rather than a record button, and why we spend so much effort on [what respondents actually experience](/knowledge-base/account-team/respondent-experience) during one.\n\nThe second requirement is an escape hatch. Some people are on a train. Some are in an open-plan office; some simply prefer typing, and the Höhne study's non-[response rates](/blog/survey-fatigue-why-response-rates-keep-falling) show how expensive it is to force the issue. Every voice survey should let the respondent switch to text mid-stream without losing their place. Choice of channel is also an [accessibility issue](/knowledge-base/integrations/accessibility) before it's a preference issue: mandatory voice excludes people with speech differences exactly the way mandatory typing excludes people with motor or vision impairments.\n\nRun it this way and the economics are hard to argue with. You're collecting interview-depth material at survey cost, and the transcripts flow into [analysis](/knowledge-base/results/voice-and-workspace-insights) the same as typed answers. The 2x word count is sitting there in the literature. The design question is whether your survey gives people a reason to say those words to it.
