Methodology: How CoreSkillAI Tests Are Built and Scored
How each CoreSkillAI test is designed, how scores and percentiles are calculated, what the limits are, and how we handle corrections.
This page explains how the CoreSkillAI tests are designed and scored, what each score can and cannot tell you, how the site is funded, and how to report an error. If a statement on any test page conflicts with this page, tell us and we will fix it.
Our principles
- Adapt established instruments, and say so. Each test is built on a published task or model: matrix reasoning (Raven, Carpenter, Just and Shell), digit span (Wechsler, Miller, Cowan), simple reaction time (Donders), hue arrangement (Farnsworth), the Stroop task, mental rotation (Shepard and Metzler), the four-branch model of emotional intelligence (Mayer and Salovey), and IPIP Big-Five markers (Goldberg).
- Be clear when something is an assumption. Most of our percentile figures come from a stated statistical model, not from a norm table we collected. Each test page says so.
- Prefer the honest answer to the impressive one. The IQ test caps scores at 135, shows a confidence interval, and refuses to print a number when performance is at chance level.
- Not clinical. These are educational and self-exploration tools. They cannot diagnose anything or replace professional assessment.
The tests at a glance
| Test | Format | Score | Normed on a sample? |
|---|---|---|---|
| IQ | 36 matrix puzzles, six options, 30 minutes | Guessing-corrected score mapped to the IQ scale (mean 100, SD 15); 95% interval; capped at 135 | No. Stated statistical assumption |
| Pattern recognition | 16 matrix puzzles, six options | Guessing-corrected score and approximate percentile | No |
| Mental rotation | 20 items, six options (one rotation, five mirror images), 6 minutes | Guessing-corrected score, approximate percentile | No |
| Working memory | Forward digit span, two tries per length | Longest sequence recalled | No. Approximate percentile based on typical adult span |
| Reaction time | 12 rounds, random delay | Median of valid rounds (120–1,000 ms) | No. Model centred on 261 ms, SD 51 ms |
| Focus (Stroop) | 32 trials, half incongruent | Interference: median incongruent minus median congruent time, correct trials only | No. Model centred on 120 ms |
| Colour vision | 4 rows of 8 movable hue chips | Total error score, converted to bands | No. Screens are not calibrated |
| Typing | 60 seconds, real text | Net WPM (or characters per minute) and accuracy | No. Reference curve near 52–55 WPM for Latin scripts |
| Big Five | 50 IPIP items, 10 per trait | Raw 10–50 per trait, shown out of 100 | No. Not a percentile |
| Emotional intelligence | 28 self-report statements, 7 per branch | Branch scores 60–140; overall is their average | No. Self-rating, not a percentile |
How scores are calculated
Correcting for guessing
For multiple-choice reasoning items we remove the credit that guessing alone would earn. With six options, a person answering at random gets about one item in six right. We convert your proportion correct as (proportion − 1/6) ÷ (1 − 1/6), floored at zero. If the result is at the level chance produces, the test says it cannot estimate a score.
Converting to a scale
The corrected proportion is placed on a normal scale under a stated assumption. For the IQ test, that gives a figure with mean 100 and standard deviation 15, which is then limited to the range our short test can support. These conversions are assumptions, not norm tables, and we label them as such.
Confidence intervals
Every test score has error. For the IQ test we assume a reliability of 0.85 (typical for matrix tests of this length), which gives a standard error of measurement of 15 × √(1 − 0.85) ≈ 5.8 points and a 95% interval of about ±11. This reliability is an assumption until we can estimate it from data.
Medians, filters and caps
Reaction times and Stroop times are skewed, with a few very slow trials, so we use medians. We drop responses that are too fast to be real reactions (anticipations) and responses so slow that they are lapses. A run is scored only if enough valid trials remain.
Reference figures we rely on
Where we cite a number, it comes from a source named on the page. Examples: about 52 WPM average typing speed (Dhakal et al., 2018, 168,000 participants); about 8% of men and 0.5% of women of Northern European descent with a red-green deficiency (Birch, 2012); a working memory capacity of about four chunks (Cowan, 2001). Our own articles list their sources at the end.
What data we collect
Results are computed in your browser. The IQ, pattern recognition and mental rotation tests can also send anonymous item-level data: which puzzle rules appeared, whether each was answered correctly, timings, the total score, the language and the date. We store no name, email, IP address, cookie or device identifier. Our aim is to use this to estimate item difficulty and reliability, so that the figures in the table above can eventually be measured instead of assumed. See the privacy policy for the full statement.
How we write and check content
- Articles and reference pages are written from published research and cite their sources.
- Numbers are checked against the source before publication. Where we cannot verify a figure, we leave it out.
- When a reader or a source shows us an error, we correct the page and, for substantive changes, note the update.
- Translations of the test pages are produced for each language and checked by automated tests for untranslated text, mixed scripts and missing sections. Languages that do not pass are set to "noindex" and left out of the sitemap until they do.
Funding and independence
CoreSkillAI is free to use and does not require an account. The site is funded by advertising (Google AdSense). Ads are placed by Google and are not related to your test results, and we do not sell personal data. Advertising does not influence what we write or how tests are scored. Page-level analytics are collected with Google Analytics. Details are in the privacy policy.
Who is responsible
The site is built and maintained by Eduard-Dragos Condria in Romania. He is a web developer, not a psychologist, which is one reason we cite sources for everything and direct readers to professionals for anything that needs one.
Known limitations
- No test here has been standardised on a representative sample of any population.
- Online conditions (screen, device, distractions) add noise, especially for reaction time, colour vision and typing.
- Short tests have wider error than long, professionally administered ones.
- Percentiles are approximate positions on stated models.
Corrections and contact
If you find a factual error, a scoring problem, a translation mistake or an accessibility barrier, please tell us on the contact page. Include the page address and what you think is wrong, and we will look into it.
Sources
- American Educational Research Association, American Psychological Association & National Council on Measurement in Education (2014). Standards for Educational and Psychological Testing.
- Carpenter, P. A., Just, M. A., & Shell, P. (1990). What one intelligence test measures. Psychological Review, 97(3), 404–431.
- Goldberg, L. R. (1992). The development of markers for the Big-Five factor structure. Psychological Assessment, 4(1), 26–42.
- Mayer, J. D., & Salovey, P. (1997). What is emotional intelligence? In P. Salovey & D. Sluyter (Eds.), Emotional Development and Emotional Intelligence.
- Dhakal, V., Feit, A. M., Kristensson, P. O., & Oulasvirta, A. (2018). Observations on typing from 136 million keystrokes. CHI 2018.
- Birch, J. (2012). Worldwide prevalence of red-green color deficiency. Journal of the Optical Society of America A, 29(3), 313–320.
- Cowan, N. (2001). The magical number 4 in short-term memory. Behavioral and Brain Sciences, 24(1), 87–114.