The Premise That AI Has Made Reading Optional
A common claim about large language models is that they have made deep reading optional. If a machine can summarise any document, the argument runs, the human capacity to read slowly, parse a complex argument, and detect a buried premise has become a luxury skill rather than a workplace necessity. The empirical evidence is going the other way. The roles that have absorbed LLMs most thoroughly (law, consulting, research, technology, finance) are also the roles where the verbal reasoning bar has visibly risen in the past two years, because the quality of work now depends on the human being able to do something the model still cannot.
The shift is straightforward to describe. Before late 2022 a junior analyst's verbal reasoning workload was dominated by reading: parsing a hundred pages of source material, extracting the relevant claims, and synthesising them into a memo. After ChatGPT's public release in November 2022 and the rapid adoption of GPT-4 from March 2023 onward, the reading workload has shifted. The model handles the first-pass extraction. The human's verbal reasoning is now spent on something different and arguably harder: detecting where the model is confidently wrong, identifying claims that need verification, and constructing arguments the model could not have written.
What Large Language Models Do Well, What They Do Badly
The strengths of current models are real and consequential. They summarise dense source material faster than any human. They generate clean, grammatical first drafts in seconds. They can paraphrase, translate, and restructure text reliably enough that most knowledge workers now use them as a routine drafting tool. None of this is in dispute, and the productivity gains are documented in studies like the MIT-affiliated work by Noy and Zhang (2023) showing significant time savings on writing tasks for professionals using ChatGPT.
The weaknesses are equally well documented. Models hallucinate citations, attributing claims to studies that do not exist, and inventing authoritative-sounding sources that fall apart on basic verification. Law schools and legal practice journals have documented LLM hallucination rates in legal research that make the unverified use of model output professionally negligent. Models conflate similar but distinct concepts when summarising long material, particularly when the source contains nuance the training data underweighted. Models follow leading questions, agreeing with whichever premise the user appears to want confirmed, which the literature on sycophancy in language models (Perez et al. 2023, Sharma et al. 2023) has characterised in detail.
The implication for verbal reasoning is direct. A worker who uses an LLM without strong verbal reasoning produces text that looks polished but contains undetected errors, citations that do not check out, and conclusions that follow from premises the model invented. A worker with strong verbal reasoning uses the same model to draft faster, but reads the output the way a senior editor reads a junior's first draft: looking for the specific failure modes the model exhibits.
Why Verbal Reasoning Matters More, Not Less
The skills that have appreciated in value since 2022 are the skills the model lacks. The first is critical reading of AI output. A junior who hands a partner an AI-drafted memo containing a hallucinated case citation has damaged the partner's trust in everything else that junior produces. The juniors who survive this risk are the ones whose verbal reasoning is strong enough to recognise the model's confident-sounding fabrications and remove them before delivery.
The second is prompt construction. A prompt is, structurally, a verbal reasoning exercise. The user is decomposing a complex task into a sequence of instructions that the model will follow literally, anticipating where the model will misinterpret, and constructing the language to close those gaps. Wei and colleagues' 2022 work on chain-of-thought prompting demonstrated that the right verbal framing of a reasoning task could substantially improve model performance on tasks the model could not solve when prompted naively. The skill that produces good prompts is verbal reasoning applied to model behaviour.
The third is constructing arguments the model could not have written. LLMs produce the median response to any prompt within their training distribution. They cannot easily produce the original argument, the unexpected synthesis, or the response that depends on tacit knowledge the model does not have. Workers whose verbal reasoning is strong enough to construct these original arguments occupy the position that LLMs have not displaced.
The Hiring Implications
Top firms have already adjusted. The verbal reasoning bar in consulting and legal hiring has risen because the lower-bar work has been automated away. A 2023 candidate who could only deliver clean first-pass summarisation was hireable. The same candidate in 2025 is not, because the firm gets that output from the model. The candidate who is hireable now is the one who can read the model's output critically, design prompts that produce work the model would not generate naively, and contribute original argumentative work that the model cannot.
This shift is visible in interview structure. Magic Circle and US biglaw firms have added AI-assisted research simulations to assessment centres, where the candidate is given an AI-generated research output containing planted hallucinations and inconsistencies and asked to deliver a memo on the underlying question. The pass criterion is not just spotting the errors, it is using the salvageable parts of the output efficiently and producing a clean argumentative product.
Consulting firms have updated their case interview practice in parallel. Candidates are increasingly asked to react to AI-generated analyses of the case material, identifying which parts of the analysis are reliable, which conclusions overreach the underlying evidence, and what additional verbal reasoning is needed to deliver a defensible client recommendation.
How to Maintain and Sharpen Verbal Reasoning in an AI-Saturated Workflow
The risk for individual workers is that LLMs make the cognitive work of reading less necessary, which causes the underlying skill to atrophy. The literature on cognitive offloading (Sparrow, Liu, and Wegner's 2011 work on the so-called Google effect, and the subsequent work by Risko and Gilbert on cognitive offloading) suggests that when a tool reliably handles a cognitive task, human performance on that task degrades over time without it. The implication for verbal reasoning is that the skill needs deliberate exercise, not automatic preservation.
The discipline that maintains verbal reasoning in an AI-saturated workflow is straightforward but uncomfortable. Read primary sources directly at least once a week, without model assistance. Brief one document a week by hand, in the way you would have done in 2021. When the model produces a summary, read the original of at least one section to verify, and recalibrate your trust in the model based on the discrepancies you find. When you draft text, write the opening paragraph yourself, then let the model continue, then rewrite the model's continuation in your own voice. These practices keep the verbal reasoning muscles engaged.
The senior people who have most successfully integrated LLMs into their work are typically the ones who treat the model as a junior associate rather than as an oracle. They give the model defined tasks, they read its output critically, they reject or rewrite the parts they would have rejected from a human junior, and they preserve their own first-pass reasoning by doing the high-stakes work themselves.
What the Long Game Looks Like
The trajectory of model capability is unsettled and the next two to three years may shift the landscape again. What is unlikely to change is the value of human verbal reasoning to the work that depends on it. The lawyer who can read a statute with the model's analysis in one screen and an authoritative legislative history in another, integrate the two, and produce a brief that neither could have produced alone, is performing the augmented work that defines this period of professional history. The same is true for consultants reading models of an industry, researchers reading AI-generated literature reviews, journalists reading model-summarised primary sources, and policy analysts reading model-drafted regulatory submissions.
The candidates and workers who will absorb the LLM transition successfully are not the ones with the strongest model fluency. They are the ones whose underlying verbal reasoning is strong enough to use the model productively without being misled by it.
If you want to know where your verbal reasoning baseline sits before the next round of model improvements raises the bar again, take the Verbal Reasoning test to see your current score against items designed to measure the underlying skill, with diagnostic feedback on which specific sub-skills (inference, deduction, evaluation of arguments) would most benefit from deliberate practice.