A construct becomes a research programme when someone builds a usable measure for it, and that is what Rotter did in 1966. The I-E Scale is a modest-looking questionnaire with two design decisions worth understanding — forced-choice pairs and deliberate filler items — both of which were attempts to stop respondents from simply giving the flattering answer. Sixty years of criticism have not made it obsolete so much as revealed exactly what it does and does not measure.
How the Scale Is Built
The instrument presents twenty-nine pairs of statements. In each pair, one option locates causes inside the person and the other locates them outside — for example, an item contrasting the idea that people earn the respect they get with the idea that respect depends on luck and circumstance. The respondent picks the statement they believe more strongly.
Six of the twenty-nine pairs are fillers that are not scored at all. They exist to make the questionnaire's purpose less obvious, on the reasonable assumption that a respondent who has worked out what is being measured will start answering accordingly. The remaining twenty-three produce a single score along the internal-external dimension.
Why Forced Choice
The format is the interesting part. On an ordinary agree-disagree scale, statements like "hard work usually pays off" attract broad agreement, because agreeing is both easy and flattering. The resulting scores cluster near the top and discriminate poorly between people.
Forcing a choice between two plausible beliefs removes that escape route. You cannot endorse both, so the answer reveals which way you lean when the two genuinely compete. It is a small piece of measurement design that does most of the work of keeping the scale honest.
The Unidimensionality Problem
The most substantial criticism is that the scale treats externality as one thing. Factor analyses of the I-E Scale repeatedly suggested it was not unidimensional, and Hanna Levenson made the argument explicit: attributing outcomes to chance and attributing them to powerful other people are different beliefs with different behavioural consequences.
Her IPC scales separate Internality, Powerful Others and Chance into three measures. The refinement matters because a single external score cannot distinguish a fatalist from a political operator — two people who behave nothing alike but who can land in the same place on Rotter's original instrument.
General Versus Domain-Specific
The I-E Scale measures a generalised expectancy, and generality has a price. Many people are internal about one domain and external about another — confident that career outcomes follow from their choices while regarding their health as largely a matter of genes and luck. Averaged into one number, those two beliefs disappear.
This is why domain-specific instruments were developed for health, work and academic settings, and why they usually predict domain outcomes better than the general scale. If you want to know how someone will behave about their health, ask about their health.
What Any Version Still Cannot Do
All of these instruments are self-reports about beliefs, which means they inherit the standard limits. They measure what you say you expect, not what you actually do; they shift with recent events, so a score taken shortly after a success or a loss will lean accordingly; and they carry cultural assumptions about individual agency that do not travel identically across societies.
None of that makes the measure useless. It makes it a description of your current default rather than a fixed reading of your character — which is the right way to hold the result of the locus of control test as well. For what the orientation means in practice, see internal vs external locus of control.