The Acoustic Return: A Research Pass on McLuhan, Voice Interfaces, and the Fifty-Year Visual Detour
The Acoustic Return: A Research Pass on McLuhan, Voice Interfaces, and the Fifty-Year Visual Detour
Introduction and Framing
The paper's central thesis — that Marshall McLuhan was directionally correct about electronic media restoring "acoustic space" but premature by roughly half a century because the internet was first built on a typographic substrate, and that the prediction is now being fulfilled as voice interfaces cross a quality threshold — sits at the intersection of well-developed theoretical literatures and a strikingly under-theorized empirical present. This research pass maps each of the six requested threads, identifies the strongest anchors, and flags the genuine gaps where the paper's contribution is most original.
The headline finding is this: McLuhan's acoustic-space thesis, Ong's secondary orality, and the cognitive-neuroscientific literatures on speech, inner speech, and embodied prosody are all richly developed individually. The history of voice interface technology is well documented in technical and market sources. What is almost entirely absent is sustained academic work that binds these literatures together to theorize the post-2022 transformer-driven voice transition as a McLuhanesque epoch shift — particularly any rigorous update of Ong's secondary orality framework for AI voice agents that have ingested the entire typographic archive. That gap is precisely the territory the paper proposes to occupy.
Thread 1: The History of Voice Interface Quality Thresholds
The long ramp (1952–2010)
The technical history is well established and uncontroversial. Bell Labs' "Audrey" (1952) recognized spoken digits zero through nine from a single trained speaker. DARPA's Speech Understanding Research program in the 1970s produced Carnegie Mellon's "Harpy" (1011-word vocabulary). The Hidden Markov Model, introduced into speech recognition in the 1970s–80s, was the dominant statistical paradigm for roughly three decades. IBM's ViaVoice and Dragon Dictate (1990, $9,000) brought consumer ASR online but required discrete dictation with pauses between words. Dragon NaturallySpeaking (1997, $150) achieved continuous recognition at roughly 100 words per minute after a per-user training session.The smartphone/cloud era (2007–2017)
Google Voice Search (2007–2008) marked the inflection from local statistical models to cloud-based recognition trained on enormous datasets — Google's English speech model was reported to contain on the order of 230 billion words by the early 2010s. Apple's Siri launched in 2011 (using technology licensed from Nuance/Dragon). Amazon Echo/Alexa launched in 2014, Google Home in 2016, Microsoft Cortana in 2014. The deep-learning revolution in ASR (RNN/LSTM models, then attention-based architectures) drove word-error rates down sharply between 2012 and 2017.The transformer/multimodal threshold (2020–2024)
The most consequential recent shift is end-to-end multimodal transformer models. OpenAI's GPT-4o ("o" for "omni"), released in May 2024, was trained end-to-end across text, vision, and audio in a single neural network rather than as a cascaded pipeline. Before GPT-4o, voice mode in ChatGPT used a three-model pipeline (Whisper ASR → GPT-4 → TTS) with average latency of 2.8 seconds for GPT-3.5 and 5.4 seconds for GPT-4. GPT-4o responds to audio inputs in as little as 232 ms with an average of 320 ms — comparable to the 210 ms typical of human conversational turn-taking. Crucially, because the same network processes audio directly, it can perceive and produce tone, multiple speakers, background noises, laughter, singing, and emotional expression rather than only transcribed words. The arXiv evaluations of GPT-4o voice mode and competing end-to-end large audio-language models (Moshi, Gemini, Qwen2-audio) confirm a qualitative jump in prosodic and paralinguistic capability over cascaded systems.This is the technical anchor for the paper's "threshold crossing" claim: roughly 2022–2024 is when conversational latency dropped beneath the human turn-taking floor and when the system began carrying paralinguistic information rather than reducing speech to text and back. The early-internet barriers — accuracy, latency, vocabulary, the need for stilted speaking, and inability to handle noise or interruption — were progressively removed; the last and most consequential ones fell with end-to-end multimodal transformers.
Adoption data
Industry-survey figures (Statista, eMarketer, Edison Research, Backlinko, Voicebot.ai) converge on the following picture. By 2024, there were roughly 8.4 billion voice-assistant-capable devices globally and approximately 145–153 million US voice-assistant users (eMarketer projects growth to 170 million by 2028 at ~3.3% CAGR). NPR/Edison Research found 62% of US adults use a voice assistant on any device; 36% own smart speakers. Roughly 33.6% of US internet users aged 16–64 use voice assistants weekly. Reported recognition accuracy is now ~93–95% for major assistants. Voice messaging (Thread 6 below) accounts for ~7 billion daily messages on WhatsApp alone (Meta, March 2022). Note that these are vendor-influenced statistics; the more conservative academic survey work (e.g., Pew, Edison Research's Infinite Dial) shows a slower S-curve and the pre-2024 plateau that eMarketer also documents. Important caveat: the quality shift from cascaded to end-to-end voice is so recent (mid-2024 onward) that adoption data lagging this technical jump are not yet available; the paper should treat post-GPT-4o usage data as emerging and partial.The "friction threshold" concept
There is no established academic literature using the specific term "friction threshold" for voice. The closest empirical anchors are: (a) Bourdin and Fayol's classic work on the working-memory cost of writing versus speaking (writing taxes working memory substantially more than speech because mechanics — spelling, motor control — compete with composition); (b) the StepWrite study (arXiv 2508.04011, 2025) showing reduced cognitive load when speech-to-text is scaffolded by adaptive LLM prompts compared with conventional dictation or ChatGPT voice; (c) standard typing-vs-speaking speed comparisons (~40–60 wpm typing for average users versus 150–200 wpm natural speech). These together provide the empirical scaffolding for the paper's threshold-crossing argument, but the precise theoretical formulation — that voice becomes lower-cognitive-effort than typing once recognition accuracy, latency, and interruption-handling all clear specific thresholds — appears to be original to the paper.Ambient computing
AirPods (2016), always-on smart-speaker microphones (Echo 2014; Google Home 2016), automotive voice (Ford SYNC, GM OnStar, embedded systems now in ~50 million new vehicles per year), and the integration of voice across smart-TV, watch, and IoT devices have moved voice from a discrete "session" interface (open the app, press the button) to an ambient interface available anywhere. The market data on cumulative Echo sales (reported around 600 million units by 2025) and the ~85–90% of smartphone users who have tried voice search are the strongest quantitative anchors. The qualitative argument — that ambient availability is itself part of what crosses the threshold, by removing the activation cost of "opening" a voice channel — is well supported by the industry literature but again under-theorized in academic media studies.Gaps in Thread 1: Rigorous academic measurement of post-2024 voice usage as a share of total daily human-computer linguistic exchange is missing. No peer-reviewed cognitive-load study yet compares end-to-end voice agents (GPT-4o, Gemini Live) against typing across realistic everyday tasks. The "friction threshold" as a theoretical construct is genuinely untheorized.
Thread 2: Neuroscience of Speech versus Writing as Cognitive Modes
Different neural substrates
Speaking and writing share core language regions (Broca's, Wernicke's, left posterior inferior temporal cortex) but have demonstrably dissociable circuits. Rapp, Fischer-Baum, and Miozzo (2015, Psychological Science) showed dissociations in aphasic patients where written morphology was preserved while spoken morphology was impaired (and vice versa), establishing that written language is not simply spoken language passed through a motor filter. Writing-specific activations consistently appear in left superior parietal lobule (Exner's area neighborhood), posterior middle/superior frontal gyri, and right cerebellum. A large stroke study (Nature Scientific Reports, 2019, n=740) found left prefrontal, left parietal, and left temporal regions as the consistent writing network, with right hemisphere and cerebellar contributions. Speech production engages bilateral motor cortex (oral articulators), supplementary motor area, insula, and the basal-ganglia–cerebellum loops governing prosodic timing and emotional modulation.The inner-speech/Vygotskian link
Vygotsky's developmental schema — social speech → private (out-loud) speech → inner speech — is one of the most robust theoretical frameworks in developmental psychology, and it is highly relevant to the paper. The literature (Alderson-Day & Fernyhough's 2015 Psychological Bulletin review "Inner Speech: Development, Cognitive Functions, Phenomenology, and Neurobiology" is the standard anchor) establishes that inner speech retains the structure of dialogue, serves self-regulatory functions, and is neurally instantiated in a network including left inferior frontal gyrus, superior temporal gyrus, and supplementary motor area — overlapping substantially with overt speech production but in a condensed/predicted form.The strong, defensible claim the paper can make: speaking aloud reactivates the same regulatory scaffold that inner speech recapitulates, because inner speech is developmentally derived from external dialogue. Typing, by contrast, recruits a substantially different mechanical/visuospatial circuit and does not engage prosodic and breath-regulatory loops. This is well supported by the existing neural-substrates literature. The specific further claim — that speaking is therefore closer to inner speech than typing is — is plausible and consistent with the literature but is not, as far as this research pass could find, the subject of a direct empirical test.
Embodied cognition, prosody, breath
The recent special issue of Language and Cognition on "Multimodal prosody" (Cambridge, 2024) is a strong anchor: prosody is not only an acoustic property of the speech signal but is embodied — co-produced with gesture, breathing, and limb rhythm, with mutual adaptation between speech and pointing/beat gestures. Work by Pilar Prieto and colleagues on embodied music/speech training shows that gesture and kinaesthetic movement aid pronunciation and pragmatic competence. Studies on speech breathing (Hoit & Hixon series; Serré et al. 2020) show speaker-specific breathing profiles that are remarkably stable within individuals. The PMC paper "Influence of bodily resonances on emotional prosody perception" (2022) and the 2025 Social Cognitive and Affective Neuroscience paper on neural correlates of vocal-cord vibratory feedback in emotional prosody production both support an embodied-interoceptive account: speaking engages vocal-cord vibration as interoceptive feedback that modulates emotion processing. None of this exists for typing, which is a purely mechanical fingertip-keyboard loop without breath, vibratory, or prosodic engagement.Therapeutic differential: expressive writing vs. expressive speaking
Pennebaker's expressive-writing protocol (15 minutes/day for 4 days about a trauma) has decades of meta-analytic support for physical and psychological health effects (Smyth 1998; Frisina, Borod & Lepore 2004; Pavlacic et al. 2019; Reinhold et al.). Critically, *Pennebaker and Seagal (1999) found that speaking produces comparable benefits to writing**. The therapeutic mechanism is articulation/disclosure, not the pen. This is a key empirical anchor for the paper: if speaking equals writing in therapeutic effect, then the historical privileging of expressive writing in clinical literature was partly an accident of the medium most easily archived and assigned — not evidence that writing is uniquely curative. The emerging voice-journaling literature (e.g., commercial products like Claire, plus reviews citing Bourdin–Fayol on cognitive load) is building on this exact insight.Dictation and the cognitive process of writing
The dictation/voice-to-text literature establishes that (a) people with dysgraphia, dyslexia, ADHD, and motor impairments produce longer and higher-quality texts via dictation than via typing/handwriting; (b) speaking bypasses the working-memory tax of grapheme retrieval, motor planning, and orthographic monitoring; (c) the form of dictated text differs systematically from typed text — longer, more parallel, less revised, closer to spoken syntax. The StepWrite paper (arXiv 2508.04011, 2025) is the most recent empirical anchor, showing that LLM-scaffolded speech-driven writing reduces cognitive load relative to both standard dictation and conversational voice assistants. This is directly relevant to the paper's claim that voice interfaces collapse the distinction between speaking and writing in a way that returns composition to a more oral substrate.Gaps in Thread 2: There is no neuroimaging study, as far as this pass could find, directly comparing composition-by-typing vs composition-by-speaking-to-an-AI (as distinct from dictation). The interoceptive/regulatory differences between the two modes are theoretically well-supported but empirically under-tested.
Thread 3: Developmental Effects of Literacy Training on Oral Capacity
The "great divide" debate
The strong version of the great-divide thesis — Goody & Watt's "The Consequences of Literacy" (Comparative Studies in Society and History, 1963), Havelock's Preface to Plato (1963), McLuhan's Gutenberg Galaxy (1962), and Ong's Orality and Literacy (1982) — held that alphabetic literacy radically restructures cognition, enabling abstraction, decontextualized reasoning, history, individualism, and the categorical separation of knower from known.The decisive empirical pushback came from Scribner and Cole's
The Psychology of Literacy (1981), which studied the Vai people of Liberia. The Vai have three literacies (their indigenous Vai script, learned at home; Arabic, learned for Quranic study; English, learned in school) that allowed Scribner and Cole to dissociate literacy itself from schooling. Their finding: there is no global cognitive "literacy effect"; specific literacies confer specific, narrow cognitive skills tied to the social practices in which they are embedded. This dissolved the strong great-divide thesis and reframed literacy as a set of socially organized practices rather than a unitary cognitive technology. Subsequent work by Brian Street, Ruth Finnegan, and David Olson modulated the picture: Olson's "metalinguistic/metarepresentational" account (literacy sponsors attention to linguistic form because writing turns language into an inspectable object) is now the most defensible refined version of a literacy-effect thesis.Print exposure and verbal skills
Stanovich and others (e.g., the Memory & Cognition line of work on print exposure) have shown within literate populations that variation in print exposure correlates with vocabulary, cultural knowledge, spelling, and verbal fluency, independent of general ability. This is real but modest, and it concerns verbal skills broadly, not a categorical divide.Suppression of the oral channel
The specific claim the paper wants to make — that heavy literacy training cognitively (not just culturally) suppresses or devalues the oral channel — is harder to ground empirically. The relevant evidence is indirect:- Ong himself argued that "chirographic" and typographic culture restructures consciousness and that highly literate adults have difficulty recovering primary-oral mental habits (formulaic thought, additive structure, aggregative epithets, situational rather than abstract categorization).
- Daniel Chandler's
Strongest anchors: Scribner & Cole (1981) for the empirical demolition of the strong great divide; Olson for the refined metalinguistic account; Ong (1982) for the cultural/cognitive consequences of typography. Gap: A rigorous account of what happens to oral fluency, narrative competence, and vocal expressiveness
within highly literate populations, and of adult re-acquisition of primary-oral mental habits, does not exist.Thread 4: Ong's Secondary Orality — Full Account and Limitations
What Ong said
In Orality and Literacy: The Technologizing of the Word (1982; revised edition 2002), Walter Ong distinguishes:- Primary orality: the thought and expression of cultures wholly untouched by writing. Properties (Ong's famous nine): additive rather than subordinative; aggregative rather than analytic; redundant or "copious"; conservative or traditionalist; close to the human lifeworld; agonistically toned; empathetic and participatory rather than objectively distanced; homeostatic (continuously updated to present needs); situational rather than abstract.
- Residual orality: oral habits persisting within literate cultures (e.g., dialogue in Plato, classical rhetoric).
- Secondary orality: a "new orality" arising from electronic media (telephone, radio, television) — "essentially a more deliberate and self-conscious orality, based permanently on the use of writing and print." Like primary orality, it fosters a group sense, participatory mystique, and concentration on the present moment; unlike primary orality, it is post-literate, planned, broadcast rather than face-to-face, and presupposes the typographic substrate.
Subsequent scholarship
Ong's framework has been productively extended:- Tom Pettitt and Lars Ole Sauerberg's "Gutenberg Parenthesis" thesis treats the 500-year period of print dominance as a historical interlude bracketed by oral cultures before and "post-Gutenberg" digital orality after. Internet knowledge, on this view, is increasingly formed through secondary orality.
Limitations of Ong's framework for the AI voice era
Ong wrote when the paradigmatic secondary-orality technologies were radio and television — broadcast, one-to-many, scripted by literate producers, and consumed (not produced) by audiences. The AI voice agent is categorically different in at least four respects:What Ong's taxonomy would predict about a voice interface that has absorbed the typographic archive and translates it back into responsive speech: properties from both columns simultaneously. From primary orality — participatory, situational, present-tense, empathetic, additive in conversational rhythm. From the typographic substrate — analytic, abstract, decontextualized, structurally complete sentences. From neither — the dyadic personalization, the asymmetric stake (the human is mortal; the agent is not), and the recursive self-reference of an interlocutor that can read itself as text and re-speak. The paper can productively argue that this is a
fourth category not anticipated by Ong.Thread 5: McLuhan's Specific Acoustic-Space Formulations
Origins
The term "acoustic space" originated in the mid-1950s at the University of Toronto. The behavioral psychologist E. A. Bott had described "auditory space" as having "no centre and no margins since we hear from all directions simultaneously." Bott's idea reached McLuhan via Carl Williams (Bott's former student and a participant in the Ford Foundation Culture and Communication seminars). McLuhan and Edmund "Ted" Carpenter, in the journal Explorations (eight issues, 1953–1959), reworked "auditory" into "acoustic" to abstract it from the physical ear. The seminal published statement is the article "Acoustic Space" in Carpenter and McLuhan, eds., Explorations in Communication: An Anthology (Beacon Press, 1960).The canonical formulations
The mantra "centre everywhere and margin nowhere" recurs throughout McLuhan's work. Findlay-White and Logan (Philosophies, 2016) trace it to medieval and Renaissance mystical sources — The Book of 24 Philosophers (12th century), Alain de Lille, Nicholas of Cusa, Giordano Bruno, Pascal — all of whom used variants to describe God or an infinite universe. McLuhan absorbed it as a Medieval/Renaissance scholar, partly through Joyce's Finnegans Wake (Bruno and Cusa are the two philosophers Joyce most frequently invokes).Key textual loci:
"Audile-tactile"
McLuhan's binary is not strictly auditory-vs-visual but audile-tactile (the integrated, synesthetic sensorium of oral cultures and of pre-print manuscript scribes) versus visual (the abstracted, hierarchical, linear-perspective sensorium installed by alphabetic literacy and intensified by print). In Gutenberg Galaxy he insists that "even the number is audile-tactile and infinitely repeatable" and that the visual abstraction "anesthetizes the other senses." A medium operates by altering the ratio among the senses; "any sense when stepped up to high intensity can act as an anesthetic for other senses." The restoration of acoustic space therefore means not merely "more sound" but a re-balancing toward the embodied, synesthetic, simultaneous mode.McLuhan on what restored acoustic culture would feel like
McLuhan repeatedly described the restored condition phenomenologically: simultaneous (not sequential), participatory (not detached), iconic (not perspectival), tribal/communal (the global village), present-tense, mythic, and emotionally engaged. He described the electric age as a "retribalization" and warned that the visual mind would find acoustic experience disorienting, vertiginous, even terrifying ("the dark of the mind, the world of emotion, primordial intuition, terror"). Importantly for the paper: McLuhan saw the direction of the shift as inevitable but the experience as potentially traumatic for typographic minds — a useful framing for what late-literate adults may now be experiencing as they shift to voice-first communication.Critical reception
Recent scholarship has both extended and contested McLuhan. Richard Cavell's McLuhan in Space: A Cultural Geography (2002) reconstructs the acoustic-space concept's intellectual genealogy. Donald Theall, R. Murray Schafer (with his "soundscape" project), and the Toronto School inheritors have variously refined or critiqued it. The Senses and Society (Tandfonline, 2016) defends "acoustic space" against its mystical-sounding excesses by emphasizing the unconscious-related, non-emplaced epistemological function the concept performs.Strongest anchors: McLuhan & Powers,
The Global Village (1989), Chapter 1; Carpenter & McLuhan, "Acoustic Space" in Explorations in Communication (1960); Gutenberg Galaxy (1962); Findlay-White & Logan (2016) for the medieval-philosophical genealogy; Cavell (2002) for the most authoritative scholarly reconstruction.Thread 6: Existing Work on Voice Interfaces as Cultural Shift
Voice notes as a communication form
The most empirically rich strand is on voice notes in messaging apps. Meta reported approximately 7 billion voice messages per day on WhatsApp as of March 2022; WhatsApp now reports more than 3 billion monthly active users. NPR coverage (April 2023) documented the rise of voice messaging among Gen Z, with reports of users sending 10–50 voice notes per day to close contacts. Reportedly 84% of Gen Z use voice notes regularly. HKS Misinformation Review has published academic work on WhatsApp audio in Lebanon, demonstrating that voice notes have become a primary vector for political communication and misinformation. Afrikan Insights and other regional sources document Northern Nigeria/Sahel populations preferring voice over text because typing standardized English is a friction barrier and because oral discourse is the deep cultural default — voice notes thus align digital communication with pre-existing communicative norms.Podcasting and oral narrative
Royston's "Podcasts and new orality in the African mediascape" (new media & society, 2023) is the strongest peer-reviewed anchor. Frontiers in Education (2024) treats podcasting as pedagogical re-oralization in the GenAI era. Podcasting Culture: Re-inventing Audio Storytelling (H-Net, 2024) frames podcasting as a "sonic renaissance" enabled by the 2007 smartphone revolution. Earlier work (e.g., "IText Reconfigured: The Rise of the Podcast") connected podcasting explicitly to Ong's secondary orality.McLuhan + voice technology
There is some popular and semi-academic writing connecting McLuhan's acoustic space directly to voice technology — Mark Logan's "Revisiting Our First Technology" (Medium/Moonshot Lab) on conversational interfaces and McLuhan; the Journal of Media Literacy "Being an AI in a Marshall McLuhan World" (2025); blog and Substack pieces such as the Lindy Newsletter's "The Return of Oral Culture." However, a sustained, peer-reviewed academic article explicitly arguing that the post-2022 voice-AI transition fulfills McLuhan's acoustic-space thesis after a typographic detour does not yet appear in the literature this pass surveyed. This is the paper's clearest contribution opportunity.Qualities of AI voice conversation that differ from human voice conversation
The HCI and AI-ethics literature is developing rapidly. Key threads:- Anthropomorphism and parasocial relations: Maeda & Quan-Haase (FAccT 2024) "When Human-AI Interactions Become Parasocial" theorizes voice quality, responsive timing, and emotional expressivity as parasocial amplifiers. The CASA paradigm (Clifford Nass and successors) shows that humans respond to voice interfaces as if they were social actors even when explicitly told they are machines.
- Asymmetric stake: the AI has no body, breath, memory across sessions (typically), or mortality; the human supplies all the embodied/interoceptive infrastructure of conversation.
- End-to-end paralinguistic capability: GPT-4o and successors can perceive tone, laughter, breathing, multiple speakers, and emotional inflection, and can produce these in response. The arXiv literature on Large Audio-Language Models (Qwen2-audio, Moshi, Gemini Live) documents this rapidly expanding capacity space.
- Latency and turn-taking: at ~320 ms response latency, AI voice now meets the conversational floor that distinguishes "talking" from "messaging."
- Persistent availability with no social cost: the AI never gets tired, bored, or judgmental — a relational structure with no precise human analogue.
The cognitive/psychological experience of switching modes
This is the most under-theorized area. Some practitioner writing (voice journaling apps, "Return of Oral Culture" essays) and a small phenomenological literature in podcasting studies exists, but no rigorous academic account of the subjective experience of adults transitioning from text-dominant to voice-dominant communication — the disinhibition, the changed pacing of thought, the re-engagement of breath and prosody, the felt sense of being "in conversation with the typographic archive" — has yet been written. The paper is positioned to be among the first.Summary: What Is Established, What Is Emerging, What Is Untheorized
Well-established (strong empirical/theoretical anchors)
- The technical history of ASR/voice interfaces from Audrey (1952) through Dragon (1990/1997) through Siri (2011) through GPT-4o (2024).
- The end-to-end multimodal transformer step-change in 2024 (GPT-4o latency 232–320 ms, paralinguistic capability, single-network audio-text-vision).
- McLuhan's acoustic-space formulations and their medieval-philosophical genealogy.
- Ong's
Emerging (active but not yet consolidated)
- Academic work on podcasting as Ong-style secondary or "new" orality (Royston 2023;
Genuinely untheorized (the paper's territory)
- The fifty-year visual detour itself as a periodization: that the internet from roughly 1990 to 2022 was an anomalous typographic intensification
Two methodological cautions for the paper
First, much of the voice-adoption statistical literature is vendor-influenced (Statista, eMarketer, Voicebot.ai, industry blogs). The paper should cite peer-reviewed and government/academic sources (Pew, Edison Research's Infinite Dial, NPR) where possible and treat industry projections (especially forecasts like "8.4 billion devices by 2024" or "75% of US households by 2025") as marketing-inflected. Second, the qualitative* transition driven by end-to-end multimodal voice (GPT-4o and successors) is so recent that the empirical literature on its cognitive and cultural effects is still mostly absent. The paper is writing about a transition in progress, which is both its opportunity and its limit — the strongest move is to acknowledge that the prediction is being fulfilled "in real time" and to ground the argument in (a) the technical threshold-crossing evidence, (b) the theoretical anchors (McLuhan, Ong, Vygotsky, Pennebaker), and (c) the demonstrable mode-shifts already visible in voice notes, podcasting, and dictation, while flagging the empirical work on AI-voice adoption and its cognitive effects as a research agenda the paper is opening rather than closing.47 facts · 25 assertions → Vygotsky · Alderson-Day · Exner's area · Nature Scientific Reports · Psychological Bulletin · Amazon Echo · Google Home · Ford SYNC. Every one is a verbatim span; nothing was paraphrased into the graph.
This is a signed piece; its findings carry their sources inline, in the text. The piece argues; the sources carry the proof.