The honest summary of the research on MBTI compatibility is that there is barely any, and what exists mostly comes from the type community's own journal. That sounds like an attack. It is not. It is the single most useful fact about the subject, because it tells you what the confident charts are actually built on.
This page walks through what has been studied, what the publisher says about its own instrument, what the much larger neighbouring literature found, and the arithmetic that explains why most of the claims you have read were never tested by anyone.
The Short Version
- Direct research on MBTI type pairings and relationship outcomes is thin, and much of it appears in the specialist journal associated with the type community rather than in mainstream personality journals.
- The publisher does not make compatibility claims. The ethical guidelines point the other way.
- The bigger, better-powered literature next door, on trait similarity in couples, consistently finds that matching contributes very little once each person's own traits are accounted for.
- Reliability comes first, and it is a problem. A category that changes on retest cannot support a claim about a multi-year relationship.
- Most of the 136 possible pairings have never been studied at all. Not studied and found weak. Never studied.
How Much Direct Research Actually Exists
If you go looking for studies that took MBTI types, paired them up, and followed the couples, you find a small number of papers with modest samples and results that range from null to weak.
Several of them were published in the journal produced by the type community itself, which is not disqualifying but does matter: a field that mostly reviews its own work has a harder time finding out that it is wrong.
What you do not find is the thing the charts imply exists. There is no large longitudinal study following couples by type combination. There is no replicated finding that any particular pairing outlasts any other. The confident second column of every compatibility chart is not a summary of evidence, because there is no body of evidence to summarise.
Independent reviews of the instrument have been blunt about the wider evidence base. Pittenger's assessments in the 1990s and again in the 2000s concluded that the popular applications ran far ahead of what the data supported, and a US National Research Council committee that examined the instrument reached a similar conclusion about the research base behind its everyday uses.
What the Publisher Says
This is the part that surprises people who have only met MBTI through charts. The Myers & Briggs Foundation's ethical guidelines are explicit that type should not be used to make selection decisions or to restrict what someone is offered, and the published materials decline to rank type combinations.
There is no official compatibility chart because the organisation behind the instrument does not endorse the idea.
Every chart in circulation is therefore an unofficial product. That does not automatically make it wrong. It does mean that when a chart says a pairing is "ideal", nobody with access to the instrument's own data stands behind the claim.
The Evidence That Does Exist, Borrowed From Next Door
Compatibility charts implicitly rely on a much larger literature that they never cite: research on whether personality similarity predicts relationship outcomes at all. That literature is well-powered, uses continuous traits rather than categories, and has produced a remarkably consistent answer.
Your traits, their traits, and the match
The standard way to take this apart is to separate three things: how much your satisfaction depends on your own personality, how much on your partner's, and how much on the fit between the two. Dyrenforth and colleagues ran this on large national samples from three countries in 2010.
Your own traits mattered. Your partner's traits mattered. The similarity between the two added almost nothing on top. The match, which is the entire content of a compatibility chart, was the smallest of the three effects by a wide margin.
The machine-learning attempt
In 2020 a large collaboration led by Joel and Eastwick pooled dozens of longitudinal datasets on couples and asked what predicted relationship quality. The strongest predictors were relationship-specific: how appreciated someone felt, how satisfied they were with the sex, whether they perceived commitment.
Individual traits followed. What the partners' traits looked like in combination added essentially nothing beyond that.
This is the closest thing to a decisive test that the field has produced, and it went the same way as everything before it. If pairing mattered much, that study had every opportunity to find it.
Perceived similarity is the one that works
There is a real similarity effect, and it is worth understanding because it explains why type feels so useful. Montoya and colleagues' meta-analytic work separated actual similarity from perceived similarity.
Actual measured similarity predicted attraction in studies where people never interacted. Once people actually met and talked, it faded. Perceived similarity kept predicting throughout.
In other words, believing you are alike does work. Being alike, measurably, does much less. A type system that gives two people a shared vocabulary and a sense of being understood is operating on the variable that matters. It is just not doing it through the mechanism the chart claims.
Reliability Comes First, and It Is a Problem
Before asking whether a type predicts anything, there is a prior question: does the type hold still? Retest studies have repeatedly found that a meaningful share of people receive a different four-letter code weeks later, typically because one letter that sat near the midpoint flipped.
The reason is structural rather than sloppy. The underlying scores are continuous and cluster in the middle, so a large number of people are close to at least one dividing line. The instrument reports a category, and the category is more confident than the score underneath it.
For compatibility this is fatal in a specific way. A chart row is keyed to the whole four-letter code. If one letter flips, the entire row changes, and with it the list of types you are supposedly suited to. Nothing about you moved.
The Arithmetic Nobody Mentions
There are sixteen types, which gives 136 possible pairings once you include same-type couples. To make an evidence-based claim about a single pairing you would need a decent sample of couples in that pairing, followed over time, with outcomes measured.
Now consider how the types are distributed. Some are common and some are rare, and the rare-by-rare combinations are very rare indeed. Recruiting an adequately powered sample of, say, INFJ-with-ENTP couples and following them for a decade is a study nobody has run and nobody is going to run.
So the chart fills 136 cells from a theory. That is not fraud, and the people who built the charts were not pretending otherwise. But it does mean the correct reading of any specific cell is "this is what the model predicts", not "this is what was found".
Where the Predictive Signal Actually Is
The same research programmes that keep failing to find pairing effects keep finding other things, and they are consistent enough to act on.
- Emotional stability. Kelly and Conley's long-running study of married couples found it among the strongest personality predictors of dissatisfaction and divorce. It is the one Big Five dimension the MBTI does not measure, which is covered in why MBTI fails at compatibility and in the five traits explained.
- Conflict behaviour. Gottman's work on criticism and contempt and defensiveness identifies behaviours, not dispositions, and behaviours are the part that changes.
- Attachment security. The pursue-and-withdraw loop tracks distress better than any trait pairing. The complete guide to attachment theory is the map, and how to develop secure attachment is the part that says it is not fixed.
Notice the pattern. Everything with a decent evidence base is either a single-person characteristic or a behaviour the couple performs. Nothing on the list is a combination of two categories.
How to Read a Compatibility Claim
Four questions will sort most of what you encounter:
- Is the claim about a pairing, or about a trait? Trait claims sometimes have evidence. Pairing claims almost never do.
- Does it name an outcome that was measured? Satisfaction and stability are measurable. "Deep connection" is not, and cannot be wrong.
- Would the claim survive one letter flipping? If a two-point score change reverses the verdict, the verdict was never that solid.
- Who published it? A finding that has only ever appeared inside the community that sells the instrument is a finding waiting for an outside test.
What This Does Not Mean
None of this makes type useless, and overcorrecting is its own error. Two things remain true.
Type is a good vocabulary. Naming a recurring friction in neutral language genuinely helps, and the perceived-similarity finding suggests that feeling understood does real work in a relationship. That is a mechanism, not a placebo.
Absence of evidence is not proof of absence. Nobody has shown that pairings do nothing. What has been shown is that when similar questions were asked with better instruments and larger samples, the answer kept coming back near zero. That is a reason to be sceptical, not a proof.
The Bottom Line
The research on MBTI compatibility is limited to the point of being nearly absent, the publisher makes no compatibility claim, and the larger literature on personality matching in couples has repeatedly found the match itself to be the least important of the three effects it can measure.
Use type for what it demonstrably does: give two people words for a difference they already noticed. Take the MBTI assessment for the vocabulary, add a Big Five assessment if you want the dimension the four letters leave out, and read the full compatibility chart as a list of likely arguments rather than a list of verdicts.
If you want the comparison that puts all of this in perspective, compatibility tests versus astrology is the uncomfortable one.
