Skip to content
IdeaScore

BlogRubricsEvaluation

How to design a rubric committees trust

Criteria, descriptors, weights and anchoring, and why a 0 to 10 scale with written descriptors produces a ranking a committee can defend.

IdeaScore team7 min read

In short. A rubric earns a committee's trust when every number on it can be explained without the person who gave it in the room. That means few criteria, written descriptors at each anchor point, weights that state what the programme is buying, a scoring range wide enough to separate a good application from a very good one, and a calibration step that corrects for strict and generous evaluators before anything is ranked.

Most programmes inherit a rubric. Someone drafted it for a previous call, it was pasted into a slide, and it has been used ever since because rewriting it in the week before a deadline is nobody's idea of a good time. Then the committee meets, two members disagree about why an application sits at rank 9 rather than rank 3, and the rubric turns out to say nothing that helps.

This is a practical account of how to build one that does.

Criteria are not descriptors

A criterion is the thing you are judging: market opportunity, technical feasibility, team. A descriptor is the sentence that tells an evaluator what a particular score on that criterion looks like.

Almost every weak rubric has criteria and no descriptors. It lists five headings, gives each a maximum, and leaves the evaluator to invent the meaning of a 7. Two faculty members will invent two different meanings, and neither will be wrong, because nothing was written down.

Descriptors are the whole job. They are also the part that takes an afternoon rather than ten minutes, which is why they are usually skipped. Write them at three anchor points per criterion — low, middle, high — and let evaluators interpolate between them. A descriptor should describe observable evidence in the application, not a feeling about it. "No named competitor, market size asserted without a source" is a descriptor. "Weak market understanding" is a restatement of the score.

Keep the criteria count low. Five is a good number for a seed-fund call; more than seven and evaluators stop reading the descriptors and start scoring from impression, which is exactly the failure the rubric exists to prevent.

Why 0 to 10 with descriptors beats 1 to 5 without

Short scales are popular because they feel decisive. In practice a 1 to 5 scale with no descriptors collapses. Evaluators avoid 1 and 5 because both feel like a statement about a person, so the real range is 2 to 4, and a 200-application field is being sorted into three buckets.

A 0 to 10 range with descriptors behaves differently. Zero becomes usable, because it means something specific — the criterion was not addressed at all — rather than an insult. Half points let an evaluator say "better than the middle descriptor, not yet the high one" without agonising. And because descriptors anchor 2, 5 and 8, the scores in between are still tethered to written evidence.

The width matters at the top of the table. In a field of 181 applications the interesting decisions happen among the top thirty, where the gap between candidates is genuinely small. A scale that cannot express a small difference forces the committee to break ties with argument instead of evidence.

Weights say what the programme is buying

Weights are not a tuning knob for making the numbers come out right. They are a public statement of what the programme values, and they should be decided before a single application is read.

A deep-tech translational fund weights technical feasibility and intellectual property heavily and discounts early revenue. A student pre-incubation programme does close to the opposite: the team and the clarity of the problem carry the call, because at that stage there is very little else to look at. Neither is wrong. What is wrong is a rubric whose weights do not match the call document that applicants read.

Set weights as maximum points rather than percentages. Evaluators find "out of 12" easier to hold in their head than "24 per cent of the total", and a total that lands on a round number — 50 is a good one — makes every downstream conversation easier. A score of 46.5 out of 50 reads as 93 per cent without anyone reaching for a calculator.

A worked rubric out of 50

Here is a five-criterion rubric for a seed-fund call, with weights summing to 50. Each criterion is scored 0 to 10 in half points; the points awarded are that score divided by ten, multiplied by the criterion's maximum points.

CriterionScore rangeWeight (max points)Anchor at the high end
Problem and market0 to 1012Named customer segment, sourced market estimate, two named competitors
Technical feasibility and TRL0 to 1012Working prototype at TRL 4 or above, test data attached
Team0 to 1010Relevant domain experience, named roles, evidence of working together
Commercial plan0 to 1010Costed path to first revenue, pricing stated, pilot partner identified
Institutional and policy fit0 to 106Uses institute facilities, aligns with a stated policy priority

A worked application scores 9.5 on problem and market, 9.5 on technical feasibility, 9 on team, 9 on commercial plan and 9.5 on institutional fit. That converts to 11.4, 11.4, 9, 9 and 5.7 points, a total of 46.5 out of 50, or 93 per cent. The committee can see exactly where the missing three and a half points went, which is the only thing they will ask.

One detail worth copying: publish the maximum points beside each criterion in the call document. Many public funders publish their rubrics with their calls, and applicants write better applications when they know what the programme is measuring. It also removes the most common complaint after a decision, which is not "I scored badly" but "I did not know that was being marked".

Anchoring and the first ten applications

Evaluators anchor on whatever they read first. If the first three applications in a batch are strong, the fourth looks weaker than it is; if the first three are poor, an average application looks like a finalist.

Two habits reduce this. First, give every evaluator a short calibration set — three applications, chosen by the programme team to sit low, middle and high — and ask them to score those before their real batch. Discuss the results in a thirty-minute briefing. This is the single highest-return half hour in the whole call.

Second, randomise assignment order rather than sending batches in case-number order. Case numbers usually track submission time, and applications submitted in the last six hours before a deadline are systematically different from those submitted in the first week.

Calibration across strict and generous evaluators

Even with descriptors, evaluators differ. One reads the high anchor as a standard almost nobody meets; another treats it as the expected level for a serious application. Across a jury of fourteen, that spread is worth several ranks.

The standard correction is z-score normalisation: for each evaluator, convert their raw scores into how far each score sits from their own mean, measured in their own standard deviations. An evaluator whose scores average 32 out of 50 with a narrow spread and an evaluator averaging 41 with a wide spread then become comparable, because each application is being judged against the rest of that evaluator's batch rather than against an absolute.

Three practical cautions. Normalisation needs enough reads per evaluator to be meaningful — below roughly eight, the evaluator's own mean is too noisy to correct with. It assumes batches are of comparable quality, which is why random assignment matters. And it should be presented alongside the raw scores, never instead of them; a committee that cannot see both will not trust either.

If a criterion shows systematic disagreement — two evaluators consistently four points apart on feasibility while agreeing everywhere else — that is a descriptor problem, not an evaluator problem. Fix the descriptor before the next round.

What to publish, and when

Publish the criteria, the weights and the scale with the call. Publish the descriptors too, unless doing so would let applicants write to the rubric in a way that destroys its signal, which is rarer than people expect.

After the decision, give each applicant their own scores by criterion and the anonymised range for the round. Not the evaluator names, not the comments written for internal use — the numbers and the descriptors they were measured against. It costs nothing once the rubric is structured, and it converts the most difficult conversation a programme manager has into a short one.

In IdeaScore a rubric is a first-class object: criteria, descriptors, weights and the scale live on the programme, evaluators score against the descriptors in the scoring view, and the leaderboard can be read raw or calibrated. But the hard part is not the tool. It is sitting down with three colleagues for an afternoon and writing fifteen descriptors that mean the same thing to all of you.

See IdeaScore run a call with your own rubric

A 30-minute walkthrough with a founder, using your programme's form and criteria.