MBTI fails at compatibility for a reason that has nothing to do with whether MBTI is any good. Even a perfectly accurate personality instrument would fail at this job, because matching two people's traits is the wrong operation. MBTI just fails harder, for four specific and fixable-sounding reasons that turn out not to be fixable.
This is not an argument that type is worthless. It is an argument about what happens when you take an instrument built to describe one person and use it to score a pair.
Failure One: The Cutoff Is Arbitrary
Type is presented as a category. You are an Introvert or an Extravert, the way you are left-handed or right-handed. That is the whole premise of a compatibility chart: two categories meet, and the chart looks up the result.
But the underlying scores are not categorical. They are continuous, and they pile up in the middle, the way height does. Most people land near the dividing line on at least one letter, and often on two. The instrument then rounds them to a side.
Consider what that means in practice. Two people can differ by a single point on the thinking-feeling scale and be sorted into opposite categories. The chart treats them as opposites. Two other people can differ by forty points and share a letter, and the chart treats them as identical.
A compatibility chart built on rounded scores inherits every rounding error, and then multiplies it by four letters and two people. For a couple where both partners sit near the middle on two letters each, the chart is reporting a result that a slightly different day would have reversed.
Failure Two: Your Type Does Not Hold Still
If a category is going to predict anything about a relationship, it has to be stable. Relationships run for years. A label that changes between one Tuesday and the next cannot describe something that lasts.
Retest studies on the MBTI have consistently found that a substantial share of people receive a different four-letter code when they retake it weeks later, usually because one near-the-middle letter flipped. That is exactly what you would expect given failure one: the letters that flip are the ones that were close to the cutoff to begin with.
Now run that through a compatibility chart. If one of your letters flips, your entire row changes, and so does the list of types you are supposedly suited to. Nothing about you changed. Nothing about your partner changed. The verdict changed because a score moved two points across a line somebody drew.
This is the failure that people notice on their own, usually when a friend retests and comes back a different type while behaving identically. The instrument is not measuring nothing. It is measuring something real and then reporting it with a precision it does not have.
Failure Three: The Trait That Matters Most Is Not In There
This one is structural and it is the most serious.
When McCrae and Costa compared the MBTI against the five-factor model in 1989, the four MBTI scales lined up recognisably with four of the five broad dimensions: extraversion, openness, agreeableness and conscientiousness.
The fifth dimension, emotional stability, had no counterpart. The MBTI does not measure it, by design, because the instrument was built to describe preferences rather than adjustment.
The problem is which dimension got left out. In Kelly and Conley's long-term study of married couples, emotional stability was among the strongest personality predictors of both marital dissatisfaction and eventual divorce. Later research has kept pointing at the same dimension: how reactive someone is to stress, how quickly they recover, how much a small slight expands.
So the instrument omits precisely the trait with the clearest link to relationship outcomes, and then a chart is built on top of the remaining four and sold as a compatibility prediction. It is a weather forecast that does not include the weather.
Why the omission cannot be patched
You could imagine adding a fifth letter. The reason nobody does is that it breaks the premise.
The four MBTI letters are framed as preferences, where neither side is better: introversion is not worse than extraversion. Emotional stability is not like that. One end genuinely goes better for the person, and a type system built on "all types are equally good" cannot absorb a dimension where they are not.
If you want that dimension in the picture, it has to come from a different instrument. A Big Five assessment includes it, and the five traits explained covers what each one does.
Failure Four: Matching Is the Wrong Question
Suppose all three problems above were solved. Perfect measurement, perfectly stable, all five dimensions included. Compatibility charts would still not work, and this is the part that most people find genuinely surprising.
When researchers separate out how much of your relationship satisfaction comes from your own traits, how much comes from your partner's traits, and how much comes from the combination of the two, the pattern is consistent. Your own traits matter. Your partner's traits matter. The combination contributes very little on top.
Large-scale work using machine learning across many longitudinal datasets has reached the same place: relationship-specific perceptions and individual characteristics carry the predictive weight, while the specific pairing of two people's traits adds close to nothing once those are accounted for.
That is a devastating result for the entire premise of a compatibility chart, because a chart is nothing but a claim about the combination. It says: these two traits, together, produce this outcome. The combination is the one part that keeps coming out near zero.
What the combination is not
It is worth being precise about what this does and does not mean:
- It does not mean partners are interchangeable. Who you are with matters enormously. It means that what matters about them is largely what they are like, not how they slot against you.
- It does not mean similarity is bad. Similarity in values, goals and life stage matters. Similarity in personality traits is the thing that keeps testing near zero.
- It does not mean nothing predicts anything. Plenty predicts. It is just behaviour rather than disposition.
The Fifth Problem, Which Is About You Rather Than the Test
There is one more reason compatibility charts feel accurate when they are not, and it is the oldest finding in this whole area.
In 1949 Bertram Forer gave his students a personality description, told each of them it was individually written, and asked how accurate it was. They rated it highly. It was the same description for everyone, assembled from a newsstand astrology column.
Type descriptions are written the same way, and they have to be. A description that fits one sixteenth of the population would be commercially useless. So the language generalises, and the generalisation gets read as precision.
This matters for compatibility specifically, because the moment a chart tells you that your pairing struggles with follow-through, you will remember every time follow-through was a problem and forget the times it was not. The chart supplies the category and your memory supplies the evidence.
What Actually Predicts How a Relationship Goes
The predictive stuff is unglamorous and behavioural, which is exactly why it does not make good charts.
How the argument is conducted
Gottman's research on couples identified a small set of conflict behaviours that reliably precede deterioration: criticism and contempt, defensiveness, and stonewalling. Contempt is the standout. None of these belongs to a type. Any type can do all four, and most people do at least one under enough pressure.
The useful question is therefore not what type your partner is, but what the two of you do in the ninety seconds after something goes wrong. The conflict styles assessment is aimed at that, and how conflict styles play out in marriage covers what changes when the stakes are shared.
The pursue-and-withdraw loop
One partner escalates to get engagement, the other shuts down to reduce heat, and each behaviour makes the other worse. This pattern predicts distress far better than any trait pairing, and it is a loop rather than a personality. It can be interrupted by either person unilaterally, which is more than can be said for a type.
Where it comes from is usually attachment rather than type. The complete guide to attachment theory lays out the four styles, avoidant attachment in relationships covers the withdrawing side, and how to develop secure attachment is the part that says the pattern is changeable. An attachment styles assessment gets you a starting read.
Whether you can tell your partner something inconvenient
Responsiveness, meaning the sense that when you bring something to your partner they engage with it rather than deflect it, does more work than any trait combination. It is also observable within a week of knowing someone, and it does not require an instrument.
So What Is Type Actually For
Type is a decent vocabulary and a poor predictor. That is not a small thing. Having a neutral word for a recurring friction is genuinely useful, because it moves an argument from accusation to description.
Used that way, the full MBTI compatibility chart becomes a list of conversations worth having rather than a verdict. What it should not do is decide anything. The same instrument fails to predict work outcomes for the same reasons, which is laid out in whether MBTI predicts job performance.
The Bottom Line
MBTI compatibility fails four times over: it rounds continuous scores into categories, the categories move on retest, the dimension most tied to relationship outcomes is not in the instrument, and even a perfect instrument would fail because the combination of two people's traits explains very little on its own.
The version worth keeping is the vocabulary. Take the MBTI assessment to get the shared language, then spend the energy on the things that move: how you fight, how you repair, and whether you can say the difficult thing out loud. For the evidence behind all of this, what the research actually says about MBTI compatibility goes through the studies one by one.
