The History of the Big Five Personality Test
From a 1936 word list to the gold-standard personality model- the real story of how psychologists spent 80 years arriving at the Big Five.

For 40 years, psychologists couldn't agree on how many personality traits existed. Estimates ranged from 3 to 4,500. What the Big Five represents isn't a breakthrough. It's an armistice - one reached only after decades of competing theories, failed experiments, and statistical arguments that forced researchers toward the same answer.
The Problem: 4,504 Adjectives
In 1936, Gordon Allport and Henry Odbert did something that sounds absurd in retrospect. They went through Webster's Dictionary and pulled every word in the English language that could describe a personality trait. They found 4,504 of them.
This wasn't pedantry. Allport and Odbert were testing a hypothesis that still guides personality research today - the lexical hypothesis. It holds that if a personality difference matters, languages will encode it. Important traits become words. Unimportant distinctions fade from speech.
The question they posed was harder: could that list of 4,504 adjectives collapse into a smaller set of fundamental dimensions? Could personality be described not through thousands of unique characteristics, but through a few core factors that explained how traits cluster and correlate?
Their work, published as "Trait-names: A psycho-lexical study" in Psychological Monographs, set the agenda for the next four decades. Everything that followed - the disagreements, the false starts, the eventual consensus - happened because of this single question.
Earlier Attempts Nobody Quite Remembers
William McDougall published a personality model with five factors in 1932 - four years before Allport and Odbert. It arrived, captured almost no attention, and vanished.
Franziska Baumgarten did lexical analysis in German in 1933, extracting 1,093 personality-related terms from the German language. Her work paralleled Allport and Odbert's findings, suggesting that five factors might be universal. She too was largely forgotten.
Allport and Odbert themselves never reduced their own list. They compiled it and stopped, leaving the reduction problem for others. This hesitation mattered because the next researcher to take it on would make a choice that nearly derailed the entire field.
Cattell's 16 Factors and Their Problem
Raymond Cattell took that 4,504-word list and spent most of the 1940s running factor analyses on personality ratings. Factor analysis is a statistical technique that looks for hidden patterns - if someone who is "talkative" tends to also be "outgoing," those traits are bundled into a single factor rather than treated separately.
Cattell landed on 16 factors and built the 16 Personality Factors test (16PF) around them. The test became influential in clinical and organizational psychology. But there was a problem researchers couldn't ignore: when others ran the same analyses on the same data, they didn't get 16 factors. They kept getting 5 or 6.
The culprit was methodological. Cattell used oblique rotations - a mathematical approach that allows factors to remain correlated with each other, keeping them separate. Other researchers used orthogonal rotations, which force factors to be independent, causing correlated traits to collapse together. Same data, same question, different answer based on how you turned the mathematical dial.
This revealed something crucial about factor analysis itself: the technique doesn't discover factors the way a geologist discovers fossils. It finds patterns according to rules you set before you start. Different rules yield different patterns.
The Air Force Study Nobody Expected
In 1961, Ernest Tupes and Raymond Christal at the U.S. Air Force Personnel Research Laboratory analyzed officer peer-ratings. They used Cattell's own rating system and ran the same factor analyses. Across 8 different samples, they kept finding the same 5 factors.
Their technical report, "Recurrent Personality Factors Based on Trait Ratings," sat buried in military literature for years. It wasn't published in mainstream personality research journals until 1992, after the Big Five had already gained traction elsewhere.
Warren Norman replicated the finding in 1963 with civilian samples, lending the 5-factor model academic credibility outside the military context. By the early 1960s, the model existed - five stable factors kept appearing in the data. But nobody called it "the Big Five" yet. The researchers simply noted that five factors kept showing up. That consistency was the data's real message.
The Name: 1981
Lewis Goldberg coined "Big Five" in a 1981 paper published in the Review of Personality and Social Psychology. His framing was deliberate: he called them "big" not because there were five - but because these five kept showing up across decades of studies by different researchers using different methods. "Big" meant robust. "Big" meant they survived every re-analysis.
Goldberg's label gave the model a brand. Research started organizing around it instead of debating whether 5 was the right number. The terminology shifted from "we keep finding five factors" to "the Big Five is the framework." The word "big" implied not discovery but consensus.
The Losers
Hans Eysenck argued for 3 factors - extraversion, neuroticism, and psychoticism - and dismissed Big Five openness and agreeableness as secondary or unstable. Cattell defended his 16-factor model for decades, never conceding that orthogonal rotation had any advantage over oblique.
Raymond Cloninger's 7-factor Temperament and Character Inventory gained traction in clinical psychiatry. Lee and Ashton published HEXACO in 2004, arguing for 6 factors by splitting off Honesty-Humility as a distinct dimension. HEXACO remains active in research.
The point isn't that these models are wrong. It's that the Big Five isn't "the answer" to personality structure. It's the model that won on replication. When researchers from different labs, different countries, and different rating methods kept landing on the same five factors, that consistency became the strongest argument the field had. Consensus through repetition.
Costa, McCrae, and the Test That Made It Real
In 1978, Paul Costa and Robert McCrae published the NEO model covering Neuroticism, Extraversion, and Openness. In 1985, they released the NEO Personality Inventory - the first instrument properly designed around the five-factor framework. In 1992, they expanded it to the NEO PI-R, adding Agreeableness and Conscientiousness to land on all five factors.
This was the inflection point. Until Costa and McCrae, the Big Five was a pattern researchers kept finding. With the NEO PI-R, it became something you could administer, score, and report to someone in five minutes.
Peter Saville had already commercialized a five-factor approach in HR with his Occupational Personality Questionnaires in 1984. But the NEO PI-R became the research standard - the instrument that clinicians, organizational psychologists, and academic researchers would use, cite, and validate across hundreds of studies.
When we built our test, we built it on the NEO-PI-R framework. Our test is based on the same methodology Costa and McCrae refined. Five minutes, no signup required.
Why Factor Analysis Keeps Finding Five
This deserves a plain explanation. Factor analysis finds patterns of co-variation. When traits like "talkative," "outgoing," and "energetic" appear together consistently in how people rate each other, they collapse into one factor - extraversion.
Across languages, cultures, age groups, and different rating methods, five factors emerged as the stable ones. They held up. The pattern wasn't unique to English or American samples. German lexicons, Chinese personality assessments, Japanese organizational ratings - the same five factors kept appearing.
This doesn't prove there are exactly five personality traits existing in nature. It proves that at one level of description - the level of how humans talk about and perceive differences in personality - five dimensions capture most of the variation. Below that level, traits subdivide further. Above it, you could compress to fewer factors. Five is where the data stabilized.
Learn more about how the Big Five measures personality.
The Current State
Since the 1990s, the Big Five has been the dominant framework in personality research. It's used in longitudinal studies tracking personality across decades, cross-cultural research validating the structure in 50+ languages, clinical screening for personality pathology, and organizational psychology for hiring and team building.
Critiques continue. HEXACO offers a sixth factor. Personality-process theories argue that contexts change behavior too much to reduce it to traits. Network models suggest personality operates as interconnected systems rather than independent dimensions.
But in most large analyses and most research contexts, the five-factor structure holds. This is why we chose it. When building a personality test, you choose the model with four decades of replication behind it. That's why TheBig5 is built on it.
Key Takeaways
- The Big Five emerged from a 40-year process of hypothesis-testing and debate, not a single discovery.
- Allport and Odbert's 1936 word list kicked off the entire research agenda.
- Cattell's 16 factors lost primarily because the statistical method (oblique rotation) created an artifact that other researchers couldn't replicate.
- The 1961 Air Force study by Tupes and Christal first demonstrated that five factors held stable across multiple samples.
- Lewis Goldberg coined "Big Five" in 1981, framing five as the robust factors that survived decades of re-analysis.
- Costa and McCrae's NEO-PI-R in 1992 created the first practical instrument based on all five factors, making the model testable at scale.
- Alternative models like Eysenck's 3-factor model, Cattell's 16PF, and the HEXACO still exist. The Big Five won through replication, not through proof that five is the "true" number.
FAQ
H3: Who invented the Big Five personality test?
No single person invented it. Allport and Odbert started the lexical work in 1936. Tupes and Christal demonstrated the five-factor pattern in 1961. Lewis Goldberg named it "Big Five" in 1981. Costa and McCrae built the NEO-PI-R test in 1992. Each contributed a piece. The Big Five is a collective discovery.
H3: When was the Big Five personality test created?
The research spanned decades. The foundational word list came in 1936. The five-factor pattern emerged in 1961. The name "Big Five" was coined in 1981. The first practical test (Costa and McCrae's NEO-PI) was published in 1985, with the full five-factor version (NEO-PI-R) in 1992. Different milestones mark different stages of development.
H3: What did researchers use before the Big Five?
Before the Big Five consolidated, personality psychology used competing frameworks. Eysenck's 3-factor model, Cattell's 16-factor model, and various one-off trait systems dominated. Researchers also used projective tests like the Rorschach and narrative approaches that didn't assume personality could be dimensionalized at all. The lexical approach was one of many competing methods.
H3: Why is it called the Big Five?
Lewis Goldberg chose the name to emphasize that five factors kept appearing across studies - not because five was proven correct, but because five was big enough to be robust and small enough to be useful. "Big" meant these factors survived decades of attempts to reduce, expand, or reorganize them. It was the factors' replication, not their metaphysical status, that earned the name.
H3: Is the Big Five the same as the Five Factor Model?
Yes, they're the same thing. "Big Five" is the common name; "Five Factor Model" is the formal name. Some researchers use the terms interchangeably. The Big Five / Five Factor Model describes the same five dimensions - Openness, Conscientiousness, Extraversion, Agreeableness, and Neuroticism - regardless of which name you use.
H3: What's the difference between Big Five and HEXACO?
HEXACO adds a sixth factor - Honesty-Humility - by splitting off traits related to moral character and integrity that the Big Five bundles into Agreeableness. Lee and Ashton published HEXACO in 2004 based on lexical analyses of multiple languages. Both models are valid; HEXACO argues that honesty and humility deserve their own dimension. The five-factor model has more research behind it; HEXACO is more precise for certain uses.
H3: Did the U.S. military help create the Big Five?
The military played a key role. Tupes and Christal's 1961 study at the U.S. Air Force Personnel Research Laboratory demonstrated that five factors held stable across officer samples. That study was central evidence that five factors were robust. However, the Air Force didn't invent the Big Five - they provided the data and analysis that helped establish it during a critical moment when the field was still unsettled.
H3: Why did Cattell's 16 factors lose out to the Big Five?
Cattell's 16 factors used an oblique rotation method that kept factors separate even when they correlated. Other researchers used orthogonal rotation, which forced factors to be independent. When orthogonal methods were applied, the 16 factors collapsed to 5 or 6. The five-factor pattern proved more stable across different methods and different researchers, so it eventually won on replicability grounds.
H3: Has the Big Five changed since the 1990s?
The core five-factor structure has remained stable since Costa and McCrae finalized the NEO-PI-R in 1992. Minor refinements have been made - facets (sub-dimensions within each factor) have been better defined, and alternative measures have been developed. But the fundamental structure hasn't changed. The Big Five has become more entrenched, not less, as it has replicated across cultures and languages.