Model performance
O.W.L. & N.E.W.T. Intelligence Index
Per-paper accuracy and equal-weight overall · higher is better
✦ The independent wizarding-lore index
The famous facts are merely the entrance examination. We measure what happens when the questions leave the Great Hall and descend into the Forbidden Section.
✦ The honours board
Consulting the records…
Model performance
Per-paper accuracy and equal-weight overall · higher is better
Current champion
— No complete resultBest deep-lore score
— N.E.W.T. examinationLargest knowledge drop
— O.W.L.s to N.E.W.T.sComplete ledger
| Rank | Model | O.W.L.s | N.E.W.T.s | Overall | Coverage | Details |
|---|
Run both papers and rebuild the site to inscribe a result.
✦ A page from the examination
Real questions from the public papers surface in the ink. First comes the prompt. Then, after a pause, the canonical answer writes itself into the record.
✦ The examination
Models sit both papers under identical closed-book settings. Broad familiarity cannot hide a collapse on obscure canon.
Ordinary Wizarding Lore
Recurring book canon beyond the famous entrance facts: aliases, dates, objects, creatures, and details a capable model should retain.
Nastily Exhausting Wizarding Trivia
Minor names, exact objects, prices, prose details, and ordered recall drawn from the books' deepest shelves.
✦ Transparent methodology
The presentation may be enchanted. The scoring emphatically is not. Every run is reproducible and every grade is inspectable.
No retrieval, browsing, tools, or source context.
One question and one concise answer in a fresh session.
Local normalization and required answer parts, never a second model as judge.
Dataset fingerprints keep unlike runs off the board.
✦ Enter the examination hall
Run locally through Ollama, or bring any OpenRouter model with a token.
$ newt-bench run --provider ollama \
--model YOUR_MODEL
$ newt-bench site build