Assessment Item Analysis From Your Marks Data

Assessment item analysis turns your marks into two numbers per question, a facility and a discrimination, and reads what a small cohort can tell you.

Item analysis tells you which questions on an exam did their job and which confused the room. Using the marks you already hold, you can calculate two numbers per question: the facility and the discrimination. Use them to clean up the assessment before it appears again. This guide explains both measures and how to compute them. It also covers the real limits of the numbers on a small cohort.

Acuity is Lectimax's cohort analytics module. It reads your marks data and surfaces the same patterns automatically. You read the analysis and decide what to act on; Acuity does the sums and the record-keeping.

What does facility tell you?

Facility is the share of students who answered a question correctly. A question that ninety percent of the class got right is easy. A question that thirty percent got right is hard. The number is just the percentage correct. Facility is the easier of the two measures to read, and it matters because a paper needs a spread. If every question is easy, the top of the class cannot separate from the middle. If every question is hard, the paper punishes the whole room. It tells you little about who knows what.

What discrimination tells you

Discrimination asks whether the students who did well overall also did well on this specific question. Split the class into a top group and a bottom group by total exam score, then compare each group's performance on the question. A question that discriminates well is one the top group answers and the bottom group misses. Such a question is pulling its weight. It separates the students who know the material from those who do not. So does a question where the strong students fall for the wrong option. That question is broken, no matter how it reads. Rewrite it or retire it. The two measures work together, and neither replaces the other. A high facility question tells you the item was easy. It does not tell you whether the strong students still answered it. That is exactly what discrimination measures. A question can be universally easy and still be fine. It can also be moderately hard and still be useless if the weakest students happen to answer it. Reading the pair together, rather than either number alone, is where the analysis starts to earn its keep.

Do it by hand on a small cohort

On a class of twenty to thirty, you can run item analysis in an afternoon with a spreadsheet. Rank the students by total score, take the top third and bottom third, and for each question record how many in each group answered correctly. Facility is the overall correct rate. Discrimination is the gap between the top and bottom group. A positive gap is a healthy question. A flat one is a question that tested nothing. A negative gap, where the weaker students outscored the stronger ones, is a flag. Read the question and its keying again. Keep the analysis simple on purpose: one row per question and one column for each group score is enough to rank the whole paper. The goal is to sort the items into keep, rewrite and retire, not to build a small analytics product. Where the numbers leave you unsure, they have still narrowed the list. Your reading of the flagged questions makes the final call.

Why are small cohorts noisy?

The real constraint is the cohort size. With twenty students, a single student moving between the top and bottom group changes every number on the page. So treat the measures as flags, not verdicts. A discrimination figure on a class of fifteen is a hint for you to reread the question, not a statistic you quote in a report. The same analysis on a cohort of a hundred or more becomes stable enough to trust for retiring or keeping items. Until the group is big enough, let the numbers point you toward questions that deserve a second look, then use your reading of the question to decide.

Batch the analysis once the data grows

A spreadsheet handles one paper. Its limit is the same as every manual method. It only sees what you enter, and it stops being updated when the term fills up. Acuity runs the per item scan across submissions and cohorts. It keeps the facility and discrimination figures current as marks land. That matters when you teach more than one group. The same question can behave differently in two sections. The score you trust is the one drawn from the combined data rather than a single exam. See the full Lectimax feature set.

FAQ

What is a good facility value? Aim for a spread across the paper rather than a single ideal number. Roughly between thirty and ninety percent correct is a workable range, with the middle of the class landing around the middle of the paper.

What does a negative discrimination mean? It usually means the question is mis-keyed, the wording is ambiguous, or the strong students read it differently from the weak ones. Read it and either rewrite or retire it before it returns in another exam.

Can I trust item analysis on a small class? Treat it as a pointer, not a proof. On a cohort of under thirty the numbers swing with one student, so use them to decide which questions to reread, and reserve hard edits for what you find when you read them. On larger cohorts the measures become reliable enough to retire questions on.

Acuity surfaces the numbers; you read them and say what they mean. Pair item analysis with a fair grading flow by reading what AI grading in higher education really is, or review Lectimax pricing and the full feature set.

Back to all posts