O.W.L. & N.E.W.T. Wizarding model intelligence Read the experiment

The independent wizarding-lore index

Is your model a Muggle?

The famous facts are merely the entrance examination. We measure what happens when the questions leave the Great Hall and descend into the Forbidden Section.

0models examined
75book-canon questions
2independent exams
0LLM judges

The honours board

Wizarding knowledge, measured.

Recorded local results

Consulting the records…

Model performance

O.W.L. & N.E.W.T. Intelligence Index

Per-paper accuracy and equal-weight overall · higher is better

O.W.L.s N.E.W.T.s Equal-weight overall
0 models
0 models shown · † trained on the public exam Benchmark v0.4.0 · deterministic alias grading

Current champion

No complete result

Best deep-lore score

N.E.W.T. examination

Largest knowledge drop

O.W.L.s to N.E.W.T.s

Complete ledger

Model standings

Rank Model O.W.L.s N.E.W.T.s Overall Coverage Details

A page from the examination

The book
answers back.

Real questions from the public papers surface in the ink. First comes the prompt. Then, after a pause, the canonical answer writes itself into the record.

O.W.L. canon N.E.W.T. deep lore
N.E.W.T. · CHARACTERS NEWT-001
The question

The record replies

The ink is listening

The examination

Two papers.
Two depths.

Models sit both papers under identical closed-book settings. Broad familiarity cannot hide a collapse on obscure canon.

Paper IChallenging canon

Ordinary Wizarding Lore

O.W.L.s

Recurring book canon beyond the famous entrance facts: aliases, dates, objects, creatures, and details a capable model should retain.

30 questionsChallenging paper
Paper IIDeep book lore

Nastily Exhausting Wizarding Trivia

N.E.W.T.s

Minor names, exact objects, prices, prose details, and ordered recall drawn from the books' deepest shelves.

45 questionsHard paper

Transparent methodology

No divination.
Just evidence.

The presentation may be enchanted. The scoring emphatically is not. Every run is reproducible and every grade is inspectable.

  1. 01
    Closed book

    No retrieval, browsing, tools, or source context.

  2. 02
    Isolated turns

    One question and one concise answer in a fresh session.

  3. 03
    Exact aliases

    Local normalization and required answer parts, never a second model as judge.

  4. 04
    Sealed comparison

    Dataset fingerprints keep unlike runs off the board.

Enter the examination hall

Your model is
expected.

Run locally through Ollama, or bring any OpenRouter model with a token.

Examination incantation
$ newt-bench run --provider ollama \
    --model YOUR_MODEL

$ newt-bench site build