Skip to content
IdeaScore

BlogIdeaScore AIEvaluation

What grounded AI evaluation means

The difference between an assistant that writes a plausible paragraph and one that researches and cites, and where an AI co-evaluator belongs in scoring.

IdeaScore team7 min read

In short. Two systems can produce the same paragraph about an application and mean entirely different things by it. One has read the application and written what an assessment usually sounds like; the other has gone and looked, and can show you where each claim came from. Only the second is useful in an evaluation process, and only if human scores remain the ones that decide, the citations are checked, and the tool is used where evidence is scarce rather than where judgement is required.

The fair question any programme manager asks is this: if the summary reads as confident either way, how do I tell which kind I am looking at? The answer is that you look at what sits underneath it. Everything in this article is about that distinction.

Two different things with one name

An ungrounded assessment is produced from the application text and from general patterns. It knows what a market-opportunity paragraph sounds like. Given a submission that says "the Indian cold-chain logistics market is large and growing", it will produce a fluent paragraph about cold-chain logistics that is plausible, internally consistent, and entirely unverifiable. It may be right. You cannot tell.

A grounded assessment runs searches, reads what it finds, and writes claims that point back at specific pages. Given the same submission, it returns something closer to: the applicant cites no source for market size; a named industry body's 2025 report gives a figure of a different order; three companies operating in the same segment are named, two of them funded within the last eighteen months; the lead applicant's publication record is listed; no granted patent matches the described technology.

Those are checkable statements. An evaluator can click a source, disagree with it, and say so. That is the only property that matters. A grounded assessment is not more intelligent than an ungrounded one; it is accountable, and an evaluation process runs on accountability.

Where grounded research earns its place in a seed-fund call is the set of questions applicants are worst at answering about themselves: who else is doing this, how large the market actually is, whether the intellectual property is real, whether the named mentor exists and works where the application says, and whether the claimed policy alignment matches the policy.

What to check in a citation

A citation is not proof. It is an address. Some are load-bearing and some are decoration, and the difference is visible in about fifteen seconds per source.

CheckWhat you are looking forWhat a bad citation looks like
ResolvesThe link opens and the page existsA dead link, or a redirect to a homepage
SpecificThe page contains the claim, not the topicA sector landing page cited for a market figure
PrimaryA filing, a registry, a paper, an official pageA blog summarising a press release about a report
DatedThe page carries a date, and it is recent enoughAn undated page cited for a current figure
IndependentNot the applicant's own websiteThe company's own about page cited for team credentials

The last row is the one that matters most and is checked least. An applicant's own site is a fine source for what the applicant claims and a worthless source for whether it is true. A grounded tool should mark self-sourced claims as such, and a reviewer should treat them as the applicant's assertion rather than as external evidence.

Two more habits. Spot-check rather than verify everything: pull three citations at random from each assessment in the first twenty applications of a round, and if they hold up, trust the rest at the level of "useful context". And read the negative findings closely — "no granted patent found" is a claim about a search, not a fact about the world, and it should say which register was searched.

Human scores stay primary

The rule we would defend to any committee: the AI does not hold a score that counts.

A grounded assessment sits beside the rubric, not inside it. It can inform an evaluator's judgement on market opportunity; it cannot supply the number. The leaderboard should be readable in three views — human only, AI only, and blended — with the human view as the default and the one the committee signs off. The AI view exists so that a large divergence is visible, and the blended view exists as a sorting aid, not as a verdict.

This is partly a governance point and partly a practical one. Committees will not approve a ranking they cannot attribute to named people. Applicants who appeal will ask who decided. An evaluation process where a human evaluator can say "I read this, here is my score, and here is why I disagreed with the research note" is defensible. One where the number came from a system nobody can question is not.

The useful signal from divergence is diagnostic. When an application scores far higher from human evaluators than from the research, it often means the pitch was persuasive and the evidence thin. When it scores far lower, it often means the team undersold something real. Both are worth a second look before the committee meets; neither should change a score automatically.

Quotas, staleness and fair use

Grounded research costs real time and real money per application, which is why it comes with a daily quota rather than an unlimited tap. Plan for it like any other resource in the call.

Three things to set up before a round:

  • A quota that fits the funnel. If 181 applications need a research pass in week 5, a limit of 20 a day means nine days. Either raise the limit for that week or start earlier. Running the pass in batches across the completeness week is usually enough.
  • A request log. Which applications were researched, when, and by whom. This is what you produce when someone asks why one application had a research note and another did not.
  • A staleness rule. Research done in week 2 against an application resubmitted in week 4 describes the wrong document. Flag assessments older than the version they refer to, and re-run before the jury round rather than before screening.

Set the expectation with evaluators that a research note is a starting point with a timestamp, not a permanent fact about the application.

Where it helps most

Pre-screening is the clearest win. Completeness checks, eligibility rules, duplicate and resubmission detection, and a short plausibility read across 181 applications take a programme team days and take a machine an afternoon — and the output is a list to verify, not a decision to accept.

Benchmarking is the second. One evaluator reading application 37 of 39 has no reliable sense of where it sits in the field. A note that places a claimed market size, a TRL assertion or a team profile against the rest of the cohort restores that context cheaply.

Deck extraction is the quiet third. Pulling the stated problem, the ask, the traction numbers and the team slide out of a 24-slide deck into a structured summary saves each of three evaluators ten minutes on every application. At 543 reads, that is real time.

Where it does not belong

Final judgement, and the criteria that are irreducibly about people. Whether a founding team will hold together through a difficult eighteen months is not a researchable question, and a confident paragraph claiming otherwise is worse than silence.

It also does not belong in the rejection letter. An applicant told they were declined because of an automated finding will, correctly, ask to see the finding and challenge it. If the programme cannot defend it line by line, it should not have been the stated reason.

And it does not belong anywhere the source cannot be shown. An assessment that cannot be traced to a page should be treated as an opinion from an anonymous reviewer — which is to say, not used.

A short checklist before you switch it on

  1. Decide which criteria the research note may inform, and write it in the evaluator briefing.
  2. Set the leaderboard default to the human view and show the committee all three.
  3. Set a daily quota that clears the cohort within the completeness week.
  4. Turn on staleness flags and re-run research after the resubmission deadline.
  5. Spot-check three citations per assessment for the first twenty applications.
  6. Tell applicants, in the call document, that applications are researched and that humans score them.
  7. Keep the request log with the round's audit trail.

IdeaScore AI is built to this shape: it researches each application on the web, cites what it used, and reports alongside the human scores rather than in place of them. The checklist above is what a programme should ask of any such tool, ours included.

See IdeaScore run a call with your own rubric

A 30-minute walkthrough with a founder, using your programme's form and criteria.