Retest research on the Myers-Briggs Type Indicator has produced one of the most quietly uncomfortable numbers in workplace psychology. Depending on the study, between 39% and 76% of people receive a different four-letter personality type when they retake the assessment just five weeks later. For teams that build communication strategies around those letters, that instability is not a trivia problem. It is an operating problem.
The data that got my attention
Psychologist David Pittenger reviewed the retest literature and reported the 39% to 76% type-change range on a five-week interval. A 1983 replication by McCarley and Carskadon found that 50% of participants changed classification on at least one of the four scales within five weeks. Even the MBTI manual’s own research reported 35% of individuals received a different four-letter type after four weeks.
One detail makes those numbers worse. The changes are not random. They cluster among people whose scores sit near the midpoint of a scale, exactly where the tool draws its binary line. Two respondents three points apart can receive opposite labels from nearly identical answers.
Why this matters now
Assessments are now standard workplace infrastructure. SHRM’s 2023 talent assessment survey found that 76% of organizations with 100 or more employees use pre-employment assessments. Mercer survey data puts personality tests among the most common types, at roughly 32% of employer usage.
That scale gives type instability real consequences. A manager who builds a coaching plan around an employee’s reported type, then watches the same employee test differently six weeks later, faces an awkward conversation. Neither the employee nor the manager is wrong. The labeling system is what wobbles.
There is also a hiring dimension. The Myers-Briggs Company itself states the instrument should not be used for selection. Sackett and colleagues’ 2022 re-analysis found structured interviews are the strongest single predictor of job performance, near .42 operational validity, while the best single personality trait predictor reached about .25. Personality data adds signal for coaching and team design. It was never built to be a gatekeeper.
What the research actually shows
The most rigorous recent summary of MBTI reliability is a 2017 systematic review and meta-analysis by Randall, Isaacson, and Ciro. Their team screened 221 studies, kept seven, and pooled three test-retest studies covering 314 participants. Results split sharply by subscale.
| MBTI subscale | Pooled test-retest reliability (Randall et al., 2017) | Meets the .70 benchmark |
|---|---|---|
| Extraversion-Introversion | .764 | Yes |
| Sensing-Intuition | .753 | Yes |
| Thinking-Feeling | .612 | No |
| Judging-Perceiving | .775 | Yes |
Three of four subscales cleared the .70 reliability threshold most psychometricians consider acceptable. The Thinking-Feeling subscale, at .612, did not. The authors cautioned that most included studies used college-age samples.
The Myers-Briggs Company disputes the harshest readings of this literature. The company reports dimension-level test-retest correlations of 0.81 to 0.86 over six to fifteen weeks on the current Global Step I instrument, and notes that older studies used outdated forms. Both findings can coexist: dimension scores are more stable than type labels, and the four-letter label is where instability concentrates.
That distinction is the real lesson. The weakness is not the underlying questions. It is the practice of sorting continuous scores into 16 boxes and building identity around the box. Type language still works for private self-reflection. Continuum-based tools work better when the output guides daily team interaction.
Everything DiSC, published by Wiley, is the clearest example of the continuum approach. Its computer-adaptive assessment places each person as a dot within the DiSC circle rather than assigning a fixed type. Published validation research reports median test-retest reliability of .86, built on roughly 50 years of research and more than 10 million learners.
A practical framework for leaders
None of this argues for abandoning personality tools. It argues for matching the tool to the job. Four rules keep a program on solid ground.
- Use assessments for conversation, not classification. A profile should open a discussion about working style, not become a label colleagues invoke in every disagreement.
- Prefer continua over boxes when output drives daily behavior. Tools that report a position on a spectrum tolerate borderline scores. Type systems convert those same scores into opposite identities.
- Keep assessments out of hiring decisions. The publisher says so directly for MBTI, and four-quadrant tools broadly lack the predictive validity for selection. Hire with structured interviews and work samples. Develop with personality data.
- Reassess without alarm. Scores drift, especially near scale midpoints. Treat movement as information about context and growth, not inconsistency.
One caution deserves emphasis. Birkeland and colleagues’ 2006 meta-analysis found job applicants score roughly half a standard deviation higher than non-applicants on desirable traits such as conscientiousness. Any personality data collected during hiring is a candidate’s best case, not a baseline.
The bottom line
The 76% type-change figure sounds like an indictment of one instrument. It is really an indictment of binary labels applied to continuous data. When leaders treat a profile as a fixed identity, they inherit the stability problem. When they treat it as a starting point for adaptation, the problem largely disappears. Continuum-based tools, with reliability above .85 on every scale, give teams a sturdier foundation.
Where to go from here
Before rolling any assessment across a team, check the validation research behind it: sample sizes, retest intervals, and scale-level reliability coefficients. Teams that want a behavioral tool designed for workplace communication can start with a DiSC workshop → built around adaptive assessment, style reading, and communication practice.
