GuideEvaluationRounds
Running evaluation rounds
Assignment, conflict of interest, deadlines, blind review, calibration, promotion and the three leaderboard views a committee will ask for.
A round is a batch of reads against one rubric, ending in a promotion decision. This guide runs one from opening to export, using a screening round of 181 applications and a jury round of 40 as the worked example.
Before you start
- The rubric for this round, with descriptors written at each anchor. A round with an unfinished rubric should not open.
- The evaluator list, invitations accepted, conflicts declared.
- The reads per application and the evaluation deadline, both published to evaluators in advance.
- The promotion rule, written down before any score exists.
- A clean submissions table: eligibility checked, duplicates resolved, superseded records removed from the queue.
That last point is worth insisting on. Assigning a round over an unresolved duplicate means two evaluators score the same venture and somebody has to unpick it afterwards.
Step 1: open the round
Create the round on the programme, attach its rubric, and set the pool of submissions it covers — all approved submissions for a screening round, the promoted set for a jury round.
Set the reads per application here. Two per application with a tie-break is right for screening; three is the defensible standard for a jury round. The arithmetic follows directly: 181 applications at three reads each is 543 reads, which across 14 evaluators is about 39 reads each. If that number looks wrong, fix it now by screening harder or reducing reads, not in week three by chasing people.
Step 2: assign evaluators
Auto-assign by expertise where the programme carries sector tags, then review what it produced before releasing it.
- Run the assignment and check the distribution: every application has its full complement of reads, and no evaluator is far above their stated capacity.
- Check that the batches are randomised within expertise. Assigning in case-number order gives one evaluator the early, considered applications and another the last-six-hours rush, and calibration cannot fix that.
- Reassign by hand where a conflict, a withdrawal or an obvious mismatch appears.
- Release the assignments and notify evaluators in one action, so that everybody starts on the same day.
Step 3: collect conflict declarations
Conflicts are declared before assignment and rechecked after it. An evaluator who is a co-inventor, a supervisor, a departmental colleague or an investor in an applicant venture should not read it, and the declaration should be on record rather than in an email.
Where a conflict emerges mid-round, remove the assignment, discard any draft score, and reassign. Keep the record of what happened; a discarded score with a reason is defensible, a quietly deleted one is not.
Step 4: set deadlines and reminders
Give evaluators a deadline with a week of slack before the committee date, and schedule three reminders: at the halfway point, three days out, and on the deadline morning.
Watch progress by evaluator rather than in aggregate. An 80 per cent completion rate usually means most evaluators have finished and two have not started, and those two need a phone call rather than a fourth email. Expect to chase two or three people personally in any round; it is normal and it is faster than redistributing their batch.
If a batch must be redistributed, do it at least four days before the deadline. A reassigned batch given to someone on the deadline morning produces rushed scores that will distort the calibration.
Step 5: decide on blind review
Blind review hides applicant and team identity from evaluators: names, institution, contact details and, where the programme wants it, the deck cover.
It is worth using where the field includes both well-known and unknown teams from the same institution, and where the criteria can be judged from the substance of the application. It is not workable where the rubric scores the team explicitly — you cannot assess relevant experience against a hidden CV — so most programmes run screening blind and the jury round open, or keep a "team" criterion outside the blind portion.
Decide once per round, state it in the evaluator briefing, and make sure lineage markers on resubmissions do not leak the identity the round is hiding.
Step 6: calibrate before you rank
Evaluators differ in strictness, and across a jury of fourteen that difference is worth several ranks. Calibration corrects for it.
The correction is z-score normalisation: each evaluator's scores are expressed as distance from their own mean in their own standard deviations, so an evaluator averaging 32 out of 50 and one averaging 41 become comparable.
Three cautions:
- Enough reads per evaluator. Below about eight, an evaluator's own mean is too noisy to correct with. At 39 reads each this is comfortable.
- Comparable batches. Normalisation assumes each evaluator saw a representative slice of the field, which is why randomised assignment in step 2 matters.
- Show both. Present calibrated and raw scores side by side. A committee shown only the adjusted numbers will not trust them.
Before calibrating, look for systematic disagreement on a single criterion. Two evaluators four points apart on feasibility while agreeing everywhere else is a descriptor problem to fix for the next round, not an evaluator to discount.
Step 7: read the leaderboard
Three views answer three different questions.
| View | Shows | Use it to |
|---|---|---|
| Human | Evaluator scores only, raw or calibrated | Produce the ranking the committee signs off |
| AI | The research-based assessment on its own | See what the evidence says without the room |
| Blended | Both, weighted as the programme decided | Sort quickly and spot divergence |
Keep the human view as the default and the one of record. The AI view exists so that a large divergence between judgement and evidence is visible before the committee meets, not so that it decides anything. Where the two disagree sharply, read the application again; it usually means the pitch was persuasive and the evidence thin, or the reverse.
Rank 1 is worth a deliberate second read regardless. It will be the most scrutinised decision of the call.
Step 8: promote and export
- Apply the promotion rule exactly as written — top N by calibrated score, or a threshold with a cap, or whatever the call document says. Do not adjust it now that you can see the scores.
- Resolve ties by the published tie-break, and record how.
- Take the promotion list to the committee for sign-off. The programme team produces the list; the committee approves it. Keeping those roles separate is what makes an appeal survivable.
- Export for the meeting: the committee report as a spreadsheet with scores by criterion, and the funnel and throughput figures from the reports view.
- Create the next round over the promoted set, with its own rubric.
- Message every applicant, promoted or not, on the date you published. Send each applicant their own scores by criterion and the anonymised range for the round.
For a 181 to 40 to 12 funnel, the jury round opens over the 40 and the same eight steps run again with a heavier rubric and fewer reads in total.
Checklist before you close the round
- Rubric descriptors complete and unchanged since the round opened
- Duplicates resolved and superseded records out of the queue
- Every application has its full complement of reads
- Conflicts declared, conflicted assignments removed, replacements scored
- No evaluator materially over their stated capacity
- Blind setting applied consistently, lineage markers checked for leaks
- Calibration run, raw and calibrated scores both available
- Criterion-level disagreement reviewed and noted for the next round
- Promotion rule applied as published, ties resolved by the published tie-break
- Committee sign-off recorded, export produced, applicants messaged on the published date
See IdeaScore run a call with your own rubric
A 30-minute walkthrough with a founder, using your programme's form and criteria.