Are Online IQ Tests Accurate? What Reliability and Norming Actually Mean
Why most online IQ tests cannot give a trustworthy score, what a proper test needs, and how to judge any test you find online.
Search for "IQ test" and you will find hundreds of free sites that promise your exact IQ in a few minutes. Most of them cannot deliver that, and the reason is not that the questions are bad. It is that a trustworthy IQ score needs things a website cannot easily provide. Knowing what those things are lets you judge any test, including ours.
What makes an IQ score valid
Psychometricians look at three things before they trust a test.
1. Norming. An IQ is not a count of correct answers. It is your performance compared with a large, representative sample of people your age, taken under the same conditions. Professional tests such as the WAIS-IV and the Stanford-Binet 5 were normed on thousands of people chosen to match national census figures for age, sex, region, ethnicity and education. A website that has never collected such a sample cannot say where you rank. Its IQ figure is produced by a formula the author invented.
2. Reliability. A reliable test gives similar scores when repeated and has items that consistently measure the same thing. This is usually reported as a coefficient (Cronbach's alpha, or test-retest correlation) and for professional IQ tests it is around 0.90 to 0.98 for full-scale scores. Most online tests publish nothing at all, so you cannot know.
3. Validity. A valid test measures what it claims to. For intelligence that means scores should correlate with other established tests and with outcomes such as school performance. Again, free online tests rarely show evidence of this.
Where online tests fall short
- No supervision. A professional administers the test one-on-one, controls timing, watches for fatigue and distraction, and scores open-ended answers by rule. A browser test cannot know whether you had the radio on, were interrupted, or got help.
- Short length. More items mean less error. A 20-question quiz cannot be as precise as a test with dozens of subtests.
- Selection bias. People who take online tests are not a random sample. Comparing yourself with them tells you little about the general population.
- Leaked answers and fixed item banks. A fixed set of questions gets shared, so scores inflate over time.
- Commercial incentive. Many sites flatter you with a high score because the result page leads to a paid certificate. If a test gives nearly everyone a score above 120, that is a sign the scale is wrong.
How big is the error on a short matrix test?
Take a test with 36 matrix-style items like ours. Its reliability might plausibly be around 0.85 (this is an assumption, since we have not yet measured it, and we say so in our methodology). With a standard deviation of 15, a reliability of 0.85 gives a standard error of measurement of 15 × √(1 − 0.85) ≈ 5.8 points. A 95% confidence interval is then roughly ±11 points. So someone who gets "115" should read it as "most likely between about 104 and 126". The difference between a test that reports one number and one that reports an interval is the difference between false precision and an honest estimate.
Is a matrix test even a good idea?
Matrix reasoning, the type used in Raven's Progressive Matrices, is one of the best single indicators of general reasoning ability. It correlates strongly with overall IQ and does not depend on language, which is why it suits an international site. But it covers only part of what a full test measures. The Wechsler batteries combine verbal comprehension, visual-spatial, fluid reasoning, working memory and processing speed. A matrix-only test will rank some people differently from a full battery.
How to judge any online IQ test
Ask these questions before you trust a result:
- Does the site say how the scale was normed, with a real sample size and population?
- Does it report reliability or a confidence interval?
- Does it tell you what the test does not measure?
- Is the score free, or does the site hide it behind a payment?
- Does it admit that it is not a clinical tool?
If the answer to most of these is no, treat the number as entertainment. If the answers are yes, treat it as an informed estimate, with the range in mind.
What our own test can and cannot claim
Our IQ test has not been standardised on a representative sample. It converts a guess-corrected score to the IQ scale using a stated statistical assumption, caps results at 135, reports a confidence interval and refuses to print a number when the performance is no better than chance. We are collecting anonymous item statistics so that, in time, item difficulty and reliability can be measured rather than assumed. Until then, the honest description is a screening tool for curiosity.
If you need a score you can rely on, for a school placement, a diagnostic question, a gifted programme or a legal reason, ask a qualified psychologist for an individually administered test. See the what is a good IQ score guide for how to read the result once you have it.
Sources
- American Educational Research Association, American Psychological Association & National Council on Measurement in Education (2014). Standards for Educational and Psychological Testing.
- Raven, J. (2000). The Raven's Progressive Matrices: Change and stability over culture and time. Cognitive Psychology, 41(1), 1–48.
- Carpenter, P. A., Just, M. A., & Shell, P. (1990). What one intelligence test measures: A theoretical account of the processing in the Raven Progressive Matrices Test. Psychological Review, 97(3), 404–431.
- Wechsler, D. (2008). WAIS-IV Technical and Interpretive Manual. Pearson.