Before There Were Tests
Numerical assessment predates psychology by millennia, but it was assessment of training rather than of ability. Scribal schools in Mesopotamia examined arithmetic competence; guild admission required demonstrated calculation; navigation and gunnery both had practical numerical entrance requirements.
What none of these had was the apparatus that turns an examination into a test: standardised items, a norm group, and evidence that the result predicted anything beyond itself.
Francis Galton's 1880s measurements and James McKeen Cattell's 1890s "mental tests" attempted the psychological version and picked the wrong variables — reaction time and sensory discrimination, which correlated with almost nothing of interest.
Binet: Numerical Items Enter the Scale
The 1905 scale of Alfred Binet and Théodore Simon, built to identify Parisian schoolchildren needing additional support, abandoned sensory measures for tasks resembling actual thinking.
Several were numerical: counting, making change, comparing quantities, repeating digit strings. The digit-span task in particular has outlived nearly everything else in the original scale, and survives in modern batteries essentially unchanged.
Its persistence is informative. Digit span turned out to be measuring working memory rather than number skill, and untangling those two threads occupied the field for the following seventy years.
1917: Numerical Testing at Scale
The US Army's group testing programme under Robert Yerkes produced Army Alpha for literate English-reading recruits, containing arithmetical reasoning items alongside verbal ones — word problems delivered on paper to a hall of candidates working against a clock.
The nonverbal Army Beta, built for recruits who were illiterate or did not read English, carried number-series and counting tasks in pictorial form. The existence of two instruments encoded an insight the field then partly forgot: a numerical word problem measures reading as well as arithmetic.
The programme's interpretive conclusions were badly wrong and were pressed into service supporting restrictive immigration policy in the 1920s. The format outlived the conclusions, and the episode remains the reason group-difference claims from ability testing warrant scrutiny about exposure rather than capacity.
Thurstone: Number as a Distinct Factor
Charles Spearman's 1904 general factor implied that a numerical score was mostly a window onto g. Louis Thurstone's 1938 Primary Mental Abilities disagreed, isolating Number as a factor distinguishable from Verbal Comprehension, Space, Reasoning, and the rest.
The practical descendant of that argument is the profile. If numerical ability is separable, then a candidate strong in verbal and weak in numerical is telling an employer something a single aggregate score would erase — and every multi-scale battery on the market exists because of it.
Wechsler and the Arithmetic Problem
David Wechsler's adult scale, published in 1939 and reissued as the WAIS in 1955, included an Arithmetic subtest: word problems solved mentally, without paper, against a time limit.
It sat inside the Verbal half, which was the structurally revealing mistake. Arithmetic is delivered verbally, but what it loads on is the capacity to hold and manipulate numbers in the moment — working memory, not stored knowledge.
The 2008 correction
The fourth edition abolished the verbal/performance halves in favour of four indexes. Arithmetic moved to the Working Memory Index alongside Digit Span, where its behaviour had always suggested it belonged.
This is the cleanest illustration available of why numerical reasoning is a hybrid: the same subtest was defensibly filed under two different constructs for fifty years.
Cattell, Horn, and the Hybrid
Raymond Cattell's fluid–crystallised distinction, developed from the 1940s and elaborated with John Horn in the 1960s, gave the hybrid a vocabulary.
Arithmetic procedures are crystallised — taught, stored, retrieved. The manipulation of intermediate results under load is fluid. Numerical scores consequently age in an intermediate pattern: better retained than abstract reasoning, less durable than vocabulary.
SHL and the Employer Era
Occupational testing became an industry in the 1970s. Saville and Holdsworth — SHL — was founded in the UK in 1977 and built the format that still dominates graduate recruitment.
Its defining move was replacing the abstract word problem with realistic business data: a table of revenues by region, a chart of quarterly output, followed by questions requiring extraction plus calculation.
Why that change mattered
Face validity improved — candidates and hiring managers both accept a test that looks like the job. But the change also imported a reading-comprehension component into a numerical measure, which is why candidates so often report running out of time on a paper whose arithmetic they found trivial.
The calculator question was settled in the same period, and settled differently by different publishers. Some permit one on the reasoning that real analysts have one; others forbid it to keep arithmetic fluency in the measurement. Both conventions are still in use, which is why reading the instructions is not a formality.
Online Delivery and Verification
Unproctored online administration in the 2000s solved logistics and created an obvious integrity gap — a numerical test taken at home can be taken with help.
The standard response was randomised item banking, drawing each candidate's paper from a large pool, paired with a shorter supervised verification test at a later stage whose score is checked for consistency against the unsupervised one.
Item-response-theory scoring made that comparison workable across non-identical papers, by estimating ability from the difficulty of the items answered correctly rather than from a raw total. Most high-volume graduate numerical testing now runs on some version of this two-stage design.
Are These Tests Still Trusted?
They remain in wide use, and the evidence base has been actively revised rather than overturned.
Schmidt and Hunter's 1998 meta-analysis placed general mental ability at the top of the predictor hierarchy, a figure repeated in textbooks for two decades. Sackett, Zhang, Berry and Lievens' 2022 re-examination argued the standard corrections had inflated those estimates and produced markedly lower corrected validities — lower, but still competitive with the alternatives.
The defensible position sits between the 1998 enthusiasm and outright dismissal: numerical tests measure something real and job-relevant, they work better as one input among several than as a gate, and the 1917 lesson that a word problem is partly a reading test has never stopped applying.