GuroSuite
Sign up
Menu
No model, no network — it is counting

Find the question that was wrong, not just hard

Two numbers for every item on your test, from the marks you have already entered. Difficulty says how many got it right. Discrimination says whether the learners who understood the topic did better on it than the ones who did not — and when the answer is no, the item is usually miskeyed.

The sample includes an item whose answer key turned out to be wrong — no account needed.

What the two numbers mean

They answer different questions, and only one of them tells you whether to throw an item away.

Difficulty (p) — and the name is backwards

The share of learners who got it right, so a high p is an easy item. On its own it says little: an item everyone passes and an item everyone fails are equally uninformative.

Discrimination (D) — the one that matters

Split the class by total score into the strongest and weakest papers, then ask whether the strong group got this item right more often. If they did not, the item is not measuring what the rest of the test measures.

A negative D is usually a wrong key

If the weak group did better than the strong group, something is wrong with the item — and by far the most common something is that the answer key is wrong. That one check is worth the whole exercise. No amount of staring at the paper finds it.

Which wrong answer they chose

A tally per option, so you can see whether a distractor pulled everybody or nobody. An option no learner picks is not testing anything and is worth rewriting before the next paper.

How the groups are split, and why it changes for a small class

The statistic is only as trustworthy as the two groups it compares.

The upper and lower 27%

Taking the extreme 27% at each end maximises the gap between the groups while keeping each one big enough to mean something. That is the whole reason for the odd-looking fraction — it is Kelley’s, and it assumes a reasonably large group.

Under 30 papers, halves instead

For a class of twelve, 27% is three papers, and three papers is noise — one absent learner would swing every D on the sheet by a third. Below 30 the class is split in half, which is the usual classroom practice and keeps each group worth reading. The sheet prints which basis it used.

See it on a real paper

A 20-item Grade 7 Science test, marked, with every item scored on both measures.

Run it on a test you have already marked

The marks are the only input. If you have entered a test, the analysis is already available.