What this test measures, and what it cannot

A temperament test is worth exactly as much as the method behind it. So here is ours in full: where the four types came from, what modern research supports and what it does not, how your answers turn into eight scales, two axes and four percentages, and where our instrument stops.

  • 01166 statements. 160 of them are scored across eight scales, six are control items and never touch your result.
  • 02Eight scales, two axes, four percentages. Every step of that path is written on this page, including the arithmetic.
  • 03No hidden coefficients, no personality database, no claim that a questionnaire can diagnose anything.
Two layers of the model. Below: eight measured scales. In the middle: two axes computed from six of them. Above: the four classical temperaments as regions of one field, not as boxes.

Where the four temperaments came from

The four types are old, and their original explanation was wrong. Both halves of that sentence matter, so here is the short history without the usual reverence.

  1. 01

    Ancient medicine

    In the Hippocratic tradition and later in Galen, health and character were explained by four bodily fluids, and a person's character came from their mixture. The word temperament itself means a mixture, and that is the only part of the idea that stayed.

  2. 02

    The Enlightenment

    Kant, in his lectures on anthropology published in 1798, kept the four names but described them as patterns of feeling and activity rather than as fluids. The four types quietly became a way of describing behaviour instead of physiology.

  3. 03

    Early laboratory psychology

    Wundt made the decisive move: he arranged the same four types along two dimensions, the strength of emotional response and the speed with which states change. That is the structure we still use, and it is why our model has two axes.

  4. 04

    Nervous system properties

    Pavlov's work on the properties of the nervous system, strength, balance and mobility, gave the four types a physiological reading again, and in several countries this is still how temperament is taught at school. We name that tradition, and we do not borrow its instruments.

As physiology, the theory of four fluids is dead and has been for a long time. What survived is the observation underneath it: people differ in the intensity and speed of their reactions, stably and from an early age. We build on the observation, not on the explanation. Stelmack and Stalikas, 1991

What survived, and what the evidence says

Four findings from research on temperament that our instrument leans on. Each of them is a statement about temperament in general, not a claim that our test is validated. That difference is the subject of section ten.

Differences in reactivity appear early

Infants differ in how strongly and how fast they react to anything unfamiliar, and those differences are measurable in the first years of life and predict behaviour years later. This is where the idea of temperament as inborn dynamics comes from.

Kagan et al., 1988

Roughly forty percent of the differences are inherited

A meta-analysis of twin and family studies puts the heritability of personality traits at about forty percent. Inherited does not mean fixed: the same number says that the majority of the variance comes from elsewhere.

Vukasović and Bratko, 2015

It is stable, and it is not frozen

Rank-order consistency, meaning whether the same person stays higher or lower than others, rises steadily with age and is already substantial in childhood. Average levels still shift across the lifespan, so a result is a snapshot of a stable tendency, not a sentence.

Roberts and DelVecchio, 2000

Temperament describes how, not what

In the formal and dynamic tradition, temperament covers energy, tempo, reactivity, endurance and similar properties: the manner of behaviour rather than its content, values or goals. Our eight scales are built in that language.

Strelau, 1996

What none of this proves: that humanity comes in exactly four kinds. The evidence supports measurable, partly inherited, fairly stable differences in the dynamics of behaviour. The four temperaments are a readable map drawn over those differences, and we use them as a map, not as a discovery.

The eight scales we actually measure

Everything on your result is computed from these eight numbers. Each scale is 20 statements, roughly half of them worded in reverse so that agreeing with everything gets you nowhere. Both ends of every scale are normal: high is not better.

  • Energy

    ENE

    How much activity a day asks of you and how long you can hold that level.

    MeasuredActive
  • Sociability

    SOC

    How much you seek contact with people and how much effort contact costs you.

    Self-containedOutgoing
  • Tempo

    TEM

    The speed of speech, movement and switching from one thing to the next.

    UnhurriedFast-moving
  • Flexibility

    FLE

    How easily you change plans, routines and the way something is usually done.

    ConsistentAdaptable
  • Reactivity

    REA

    How quickly and how strongly a feeling starts when something happens.

    UnruffledResponsive
  • Recovery

    REC

    How fast you come back to your usual state after a hard episode.

    Slow to resetResilient
  • Sensitivity

    SEN

    The level of noise, light, tone or detail at which you begin to react at all.

    Steady in noiseFinely tuned
  • Persistence

    PER

    How long you keep effort on one task after it stops being interesting.

    Quick to switchEnduring

Two of the eight, Flexibility and Persistence, stay out of the axes on purpose. They describe how a profile is coloured rather than where it sits on the map, and mixing them into the axes would blur the very thing the axes are for.

From answers to a formula

Four steps, no others. If you want to check the arithmetic on paper, everything you need is here.

  • 01Answers
  • 02Eight scales
  • 03Two axes
  • 04Four shares
  1. 01

    You answer

    166 statements on a five point agreement scale. Nothing is skipped, and no answer is weighted more than another.

  2. 02

    Scales are summed

    Each scale gives a raw sum from 20 to 100, and reversed statements are flipped first. The raw sum becomes a position on the scale: pct = (sum - 20) / 80 * 100.

  3. 03

    Two axes are computed

    Activation is the mean of Energy, Sociability and Tempo. Stability is the mean of Recovery and the inverted Reactivity and Sensitivity. Both end up on the same 0 to 100 field.

  4. 04

    Four shares are derived

    Each temperament gets the product of its two conditions: sanguine is activation times stability, choleric is activation times the opposite of stability, phlegmatic and melancholic take the mirrored pair. The four products are then scaled to add up to 100.

In symbols, with A for activation and S for stability, both from 0 to 100: w(sanguine) = A x S, w(choleric) = A x (100 - S), w(phlegmatic) = (100 - A) x S, w(melancholic) = (100 - A) x (100 - S). The four weights are normalised to 100 and rounded so that the percentages always add up exactly.

There is no step in between where a human decides anything. The same answers always give the same result, your share link carries the raw sums rather than the conclusions, and an old link keeps working even if we later adjust a threshold.

Why thirteen profiles and not sixteen

A two axis model cannot produce every combination people expect, and pretending otherwise would be the first place to cheat. So here is the consequence, stated plainly.

  • Sanguine
  • Choleric
  • Phlegmatic
  • Melancholic

Horizontal: Activation. Vertical: Stability.

If a book or a quiz told you that you are a sanguine melancholic, the combination is not nonsense, it just is not a blend in this model. It usually means a profile close to the centre.

Sanguine and melancholic sit in opposite corners of the field: one is high activation with high stability, the other is low activation with low stability. To be strongly both, a person would have to be on both sides of both axes at once. The arithmetic says the same thing: sanguine outranks choleric exactly when stability is above the middle, and melancholic outranks phlegmatic exactly when it is below. Those conditions cannot hold together, so if sanguine leads your formula, melancholic is never second.

What remains is four pure types, eight adjacent pairs in a fixed order, and one balanced profile for people close to the centre. Thirteen identities, all reachable, none invented to make the list longer. A profile counts as pure when the leading temperament is at least 14 points ahead of the second, and as balanced when the distance from the centre is smaller than 7 points.

The thirteen profiles

What a high score means here

Levels are the place where tests quietly start lying, so this is what our labels do and do not mean.

  • Low, under 30
  • Middle, from 30 to 70
  • High, above 70
  • Low, under 30

    The lower end of the scale as this test defines it. It describes a pole, not a deficiency, and both poles have their own costs and advantages.

  • Middle, from 30 to 70

    The widest band, and the most common place to land. A middle score means the property does not stand out in either direction, which is information, not a missing answer.

  • High, above 70

    The upper end of the same scale. It says the property is pronounced, and it says nothing at all about whether that is good.

These positions are measured against the scale of this test, not against a population. We do not show percentiles, and we will not show them until we have enough completed and validated protocols to build norms honestly. When norms appear, the wording on your result will change to say exactly who you are being compared with: people of the same age and sex who took this test, not humanity.

Control items and the reliability flag

Six of the 166 statements are not scored. They exist to tell us whether a set of answers can be read at all.

  1. 01

    Attention items

    A few statements have an obvious intended answer. Missing one is a normal slip, missing several means the questionnaire was being clicked through rather than read.

  2. 02

    Consistency pairs

    Some statements come in pairs that say the same thing in different words, or the exact opposite. Answers that contradict each other by a wide margin are counted as a failure of that pair.

  3. 03

    The flag

    Two or more failures raise a reliability flag. Your result still opens in full, the flag is shown next to it in plain words, and the paid analysis is told not to build confident conclusions on shaky data.

We do not block anyone, and we do not hide the flag from the person it concerns. What we also do not do is store your individual answers as a public record: what leaves your browser in a share link is the eight raw sums, the flag and a checksum.

What we store and for how long

What this test does not do

Five things, written down so that nobody has to guess where our claims end.

  • It does not diagnose anything

    This is not a clinical instrument, and none of our scales is a symptom. If something in your life feels beyond your control, a questionnaire is not the tool for it, and neither is a paid report from us.

  • It does not choose your profession

    Temperament describes the environment that costs you less energy, not the job title you should hold. Any test that lists occupations from four types is selling certainty it does not have.

  • It does not sort people into four kinds

    A quantitative review of taxometric research found that personality differences are almost always continuous rather than categorical. Our four names are regions of a continuous field, which is why your result is a formula with percentages and not a label.

    Haslam et al., 2012
  • A description that feels accurate proves nothing

    People accept vague, flattering descriptions as personally true, a demonstration first run in a classroom in 1949 and repeated ever since. That is why our text is tied to your numbers and why we tell you which scales produced each statement.

    Forer, 1949
  • A questionnaire sees only what you can see

    Self-report is accurate for internal and easily observed traits and weaker for the ones your friends judge better than you do. That limit applies to every self-report instrument, ours included.

    Vazire, 2010

How we check our own numbers

The honest position for a new instrument is not a claim of validation. It is a list of what has been done, what has not, and what will be published.

What has been done

The item bank was written to explicit rules, with reversed statements on every scale and no double barrelled wording. The formula was tested on 100000 simulated response sets: the profile survives small answer changes in about 96 percent of cases, and no change of two answers ever moves a profile by more than one neighbour.

What is not there yet

We have no population norms, no published internal consistency figures and no external validity study. Internal consistency will be reported per scale as the classical coefficient once enough valid protocols have accumulated, and it will be reported whatever the numbers look like.

What we will publish and what we never will

Derived statistics only: norm tables, reliability coefficients, correlations between scales. Never individual protocols, and never data that could identify a person. Only sessions that pass the validity checks go into the numbers.

If a competing test states its accuracy as a single percentage, that number is either about something else or about nothing. There is no such figure for personality instruments, and we will not invent one for ours. Cronbach, 1951

Sources

Every reference here is a real publication, and every link is a DOI that resolves to it. Each one supports a specific statement on this page, and none of them says anything about our test.

  1. 01

    Stelmack, R. M., Stalikas, A. (1991). Galen and the humour theory of temperament. Personality and Individual Differences.

    The history of the four types and of the physiology that was abandoned.

    10.1016/0191-8869(91)90111-NOpen the paper
  2. 02

    Kagan, J., Reznick, J. S., Snidman, N. (1988). Biological bases of childhood shyness. Science.

    Differences in reactivity to the unfamiliar are measurable early and persist.

    10.1126/science.3353713Open the paper
  3. 03

    Rothbart, M. K. (2007). Temperament, development, and personality. Current Directions in Psychological Science.

    How temperament is defined in current research and how it relates to personality.

    10.1111/j.1467-8721.2007.00505.xOpen the paper
  4. 04

    Zentner, M., Bates, J. E. (2008). Child temperament: an integrative review of concepts, research programs, and measures. International Journal of Developmental Science.

    A review of what different research programmes actually mean by temperament.

    10.3233/DEV-2008-21203Open the paper
  5. 05

    Strelau, J. (1996). The regulative theory of temperament: current status. Personality and Individual Differences.

    The formal and dynamic language our eight scales are written in.

    10.1016/0191-8869(95)00159-XOpen the paper
  6. 06

    Trofimova, I., Robbins, T. W. (2016). Temperament and arousal systems: a new synthesis of differential psychology and functional neurochemistry. Neuroscience & Biobehavioral Reviews.

    Why energy, tempo and endurance are treated as separate properties rather than one.

    10.1016/j.neubiorev.2016.03.008Open the paper
  7. 07

    Vukasović, T., Bratko, D. (2015). Heritability of personality: a meta-analysis of behavior genetic studies. Psychological Bulletin.

    The heritability figure of about forty percent quoted in section three.

    10.1037/bul0000017Open the paper
  8. 08

    Roberts, B. W., DelVecchio, W. F. (2000). The rank-order consistency of personality traits from childhood to old age: a quantitative review of longitudinal studies. Psychological Bulletin.

    Stability of relative position across the lifespan, and its limits.

    10.1037/0033-2909.126.1.3Open the paper
  9. 09

    Caspi, A., Roberts, B. W., Shiner, R. L. (2005). Personality development: stability and change. Annual Review of Psychology.

    Average levels of traits keep shifting with age, so a result is a snapshot.

    10.1146/annurev.psych.55.090902.141913Open the paper
  10. 10

    Digman, J. M. (1997). Higher-order factors of the Big Five. Journal of Personality and Social Psychology.

    Evidence that broad traits group into two higher order factors, as our axes do.

    10.1037/0022-3514.73.6.1246Open the paper
  11. 11

    DeYoung, C. G. (2006). Higher-order factors of the Big Five in a multi-informant sample. Journal of Personality and Social Psychology.

    The same two factor structure confirmed with ratings from other people.

    10.1037/0022-3514.91.6.1138Open the paper
  12. 12

    Haslam, N., Holland, E., Kuppens, P. (2012). Categories versus dimensions in personality and psychopathology: a quantitative review of taxometric research. Psychological Medicine.

    Personality differences are continuous, which is why we report percentages.

    10.1017/S0033291711001966Open the paper
  13. 13

    Gerlach, M., Farb, B., Revelle, W., Nunes Amaral, L. A. (2018). A robust data-driven approach identifies four personality types across four large data sets. Nature Human Behaviour.

    Data driven types can be found, and they are dense regions rather than boxes.

    10.1038/s41562-018-0419-zOpen the paper
  14. 14

    Forer, B. R. (1949). The fallacy of personal validation: a classroom demonstration of gullibility. The Journal of Abnormal and Social Psychology.

    Why a description can feel accurate and still mean nothing.

    10.1037/h0059240Open the paper
  15. 15

    Vazire, S. (2010). Who knows what about a person? The self-other knowledge asymmetry (SOKA) model. Journal of Personality and Social Psychology.

    Which traits self-report reads well and which ones other people read better.

    10.1037/a0017908Open the paper
  16. 16

    Cronbach, L. J. (1951). Coefficient alpha and the internal structure of tests. Psychometrika.

    The classical internal consistency coefficient we will report per scale.

    10.1007/BF02310555Open the paper

Two things this list is not. It is not a claim of endorsement: none of these authors knows about this project. It is not a bibliography for decoration: if a source stopped supporting a statement on the page, it would be removed.

Text last updated:

Frequent questions

01

Is the four temperaments theory scientific?

As a theory of bodily fluids, no, and it has not been for centuries. As a way of describing differences in the energy, speed and intensity of behaviour, the underlying observations are studied seriously and hold up. We use the four names as a readable map over eight measured scales, and we do not claim more than that.

02

Is temperament inborn, and can it change?

Both, in a specific sense. Twin and family studies put heritability at roughly forty percent, and relative position stays fairly stable over years. Average levels still shift with age and circumstances. A result describes a stable tendency, not a fixed trait, and it is a snapshot of the person you are now.

03

Could I get a different result by answering differently?

Yes, and that is not a flaw of this test alone: a self-report questionnaire measures what you report. The control items catch inattentive or contradictory answering, and the simulation shows that small honest differences rarely move the profile further than one neighbouring option.

04

Why do you not show percentiles or a percentage of accuracy?

Percentiles need norms, and norms need a large, filtered sample. Until we have one, every level is stated against the scale of this test, and the wording says so. A single accuracy percentage does not exist for personality instruments at all, so anyone showing one is describing something else.

05

Did you take items from an existing questionnaire?

No. All 166 statements were written for this test to our own rules, and no protected instrument is reproduced, translated or adapted anywhere in it. The classical tradition is named on this page as history and as context, which is what it is.

Now the part that only your own answers can fill in

The method above is the same for everybody. What it produces is not: eight scales, two axes and a formula that belongs to one person. It takes about twenty minutes, and the result costs nothing.

Take the test