77,284 item-level answers to 5 for-fun quizzes, in 24 languages, taken by real visitors on JobCannon between March 2026 and August 2026. Free to download, free to reuse.
These 5 quizzes are not psychometric instruments. Nobody validated them. There is no norm sample behind them and no reliability coefficient, and they were never built to measure any particular construct. “Spirit animal” is not a trait. The result a taker sees is the largest of a handful of hand-authored category counters, and those categories were picked because they make a fun result page.
So please don’t read a mean here as a population estimate, and there is no published norm for any of the 5 to compare these numbers against. What the file is genuinely good for is the behaviour around the answers: how people respond to identical forced-choice items in 24 languages, how long they take, and how answer patterns and result shares move between language audiences.
If you want scored instruments with published norms behind them, that is a separate release — the JobCannon Psychometric Response Dataset, nine instruments, no overlapping rows. Both come out of the same generator under the same privacy rules, and we kept the two test lists apart on purpose, so any given response appears in exactly one of them. The reliability coefficients for those instruments are on the reliability page; you will not find one for anything on this page, and that is deliberate.
10 files, 6.2 MB: one CSV per quiz, three aggregate tables, a manifest and the dataset card. One row is one completed quiz — the raw per-item answers, the category totals the site computed from them, and the result the taker was shown. Item wording is not distributed; the answer columns hold values only.
Jungian archetypejungian_archetype
Spirit animalspirit_animal
Mental agemental_age
Aura colouraura_color
Past life erapast_life
| Quiz | Items | Results | Median time | Responses |
|---|---|---|---|---|
| Jungian archetypejungian_archetype | 24 | 12 | 5m 17s | 31,081 |
| Spirit animalspirit_animal | 12 | 8 | 2m 09s | 17,520 |
| Mental agemental_age | 12 | 5 | 2m 02s | 10,985 |
| Aura colouraura_color | 10 | 7 | 1m 52s | 10,929 |
| Past life erapast_life | 12 | 8 | 2m 27s | 6,769 |
| Total | 70 | 77,284 |
Three aggregate tables ship alongside, already summarised by quiz and language: coverage.csv (126 rows, rows per quiz × language, with the suppression flag), result_distribution_by_locale.csv (233 rows, result shares per quiz × language), dimension_means_by_locale.csv (290 rows, mean category scores per quiz × language). Every one of them carries an ALL pseudo-language row, so the overall shape is in the table rather than something you have to recompute.
All 4 mirrors serve the same 10 files as v1, checked file by file against Zenodo’s own checksums rather than by file count.
English is 46.0% of the corpus, which makes this one of the less English-heavy public files of its kind. Japanese alone is 26,769 rows and Arabic is 5,885 — both of them answering the same items, under the same scoring, as the English takers.
| Language | Responses | Share |
|---|---|---|
| Englishen | 35,541 | 46.0% |
| Japaneseja | 26,769 | 34.6% |
| Arabicar | 5,885 | 7.6% |
| Spanishes | 2,541 | 3.3% |
| Germande | 927 | 1.2% |
The remaining 19 languages are in the file too, down to a few dozen rows each, and a further 1,507 rows never had a language recorded at all. Those are kept as their own bucket rather than folded into a total: 24 languages and 25 buckets are both true, and only one of them is a language count.
The build reads exactly six columns out of the results table — language, answers, scores, result, duration, timestamp — and nothing else. That is the privacy boundary, and it is a boundary rather than a cleanup step: there is no user id, no name, no email, no referrer and no campaign field to remove afterwards, because none of them are ever read. The response id is a sequential integer counted inside each file, not the database id, and the timestamp is reduced to year and month.
79,933 responses finished in the collection window and 77,284 are published. The 2,649 difference — 3.3% — is rows that answered an older version of a quiz, which are not comparable to the current item set. Duration is left blank where it came out at zero or over two hours, roughly 2% of rows, since a tab left open overnight is not a completion time. Nothing is weighted and nothing is down-sampled.
The aggregate tables apply two thresholds: a quiz-and-language cell has to hold at least 100 rows to be reported at all, and inside a reported cell any result category under 30 rows folds into a single “other” bucket. 37 of 121 cells clear that floor; the rest are present but flagged as not publishable, so you can see what exists without us reporting a share computed from eleven people.
Published under CC BY 4.0. Cite the DOI rather than this page — it resolves to whichever version is current, so a citation written today still points at something real after the next release.
JobCannon (2026). JobCannon Entertainment Quiz Response Dataset (v1) [Data set]. Zenodo. https://doi.org/10.5281/zenodo.21950814
The biggest single file is Jungian archetype at 31,081 responses. If you use the data and publish something, we would like to hear about it — and if you find something wrong in it, we would like to hear about that faster.
First-party score result distributions for the validated assessments, their internal-consistency reliability, and the methodology behind each instrument — none of which apply to the quizzes on this page, which is the point of keeping the two releases apart.
No, and we would rather say so on the download page than in a footnote. These 5 quizzes were written to be fun. Nobody validated them, there is no norm sample behind any of them and no reliability coefficient, so a mean in this file is not a population estimate of anything. What the file is good for is behaviour around the answers: how people respond to the same forced-choice items in 24 languages, how long they take, and how result shares move between language audiences. If you need scored instruments with published norms, that is our separate psychometric release and the two share no rows.
One row is one completed quiz. It carries a per-file response id, the language, the year and month it was taken, how long it took in seconds, the result the taker was shown, the category totals the site computed, and the raw answer to every item. Item wording is not distributed — the answer columns hold values only. Four of the quizzes carry one score column per possible result; Mental age instead carries a single percentage and its result is a band.
No. The build selects six columns and nothing else, so there is no user id, no name, no email, no referrer and no campaign data to leak — the columns are not stripped afterwards, they are never read. The response id is a sequential integer counted inside each file rather than the database id, and the timestamp is cut down to year and month, so a row cannot be matched back to a visit.
Because we only kept rows that answered the quiz's current item count. Quizzes get re-cut over time, and a row from an older item set is not comparable to a row from the current one, so 2,649 rows — 3.3% of what finished — were left out. Nothing else was filtered: no weighting, no down-sampling, no removing rows that looked odd.
A quiz-and-language cell is only published once it holds at least 100 rows, and inside a published cell any result category under 30 rows is folded into a single "other" bucket. That leaves 37 of 121 cells reported and the rest flagged as not publishable. The item-level files are unaffected — the thresholds apply to the summary tables, where a thin cell would name a handful of people's results.
CC BY 4.0. Use it commercially, redistribute it, build on it — the one condition is attribution, and the citation block on this page is the form we would like. Cite the DOI rather than this page: the DOI resolves to whichever version is current, so a citation written today still points at something real after the next release.