Why your reviewers aren’t all scoring the same way — and what to do about it

by | Jul 28, 2026 | Article

Two reviewers. The same scoring rubric. The same grant application. And yet two completely different scores. Sound familiar?

This is called scoring variance, and it’s one of the most common problems in grant review, yet one of the least talked about. It cuts right to the heart of what grantmaking is meant to deliver: fairness and trust. When the outcome of an application depends more on who reads it than on how strong it is, your whole program loses credibility.

Luckily, scoring variance isn’t unsolvable. It has clear causes and the right structures can reduce it significantly.

How big is the problem, really?

Many people assume a good rubric will automatically lead to consistent results. Research tells a different story, and has done so for some time.

Back in 2012, an analysis of over 8,300 grant applications and more than 23,000 individual reviews at the Austrian Science Fund (FWF) found only weak agreement between reviewers assessing the same applications. A much more recent study from 2025, covering close to 135,000 reviews across four Norwegian funding bodies, confirms the same pattern. Even after one funder deliberately tried to improve agreement between reviewers, it still sat at only around 29 per cent afterwards.

None of this means reviewers are doing a poor job. It means a rubric on its own rarely guarantees consistent outcomes. That’s exactly where structured process design comes in. For more on how to measure fairness across your program, see our article on KPIs that make grant programs better.

The three most common causes of inconsistent scoring

1. Unclear criteria

A 1 to 5 scale sounds simple enough. But what does a 3 actually mean? Without a concrete description for each point on the scale, everyone interprets it differently. One reviewer treats a 3 as solidly average, another sees it as a real weakness.

2. Lack of calibration

Calibration means reviewers practise scoring together beforehand and compare notes on where they land. Skip this step and everyone brings their own personal yardstick to the table. Experienced reviewers are often the ones who most overestimate how well they actually understand the scale. One study of grant funding processes found that a short calibration session improved agreement noticeably for both new and experienced reviewers alike.

3. Peer influence

Review panels can develop a dynamic where a handful of opinions carry too much weight. Whoever speaks first, or loudest, tends to shape the discussion. Other reviewers then adjust their own scores to avoid conflict, and the outcome ends up reflecting group dynamics rather than the actual quality of the application.

What inconsistent scoring costs your program

Scoring variance has real, practical consequences. A genuinely strong application can be turned down simply because it landed with a stricter reviewer. A weaker one might get funded because a more generous reviewer happened to assess it. To applicants, that feels arbitrary, even when your process looks fair on paper.

It affects your team too. Once reviewers notice that outcomes depend on who was assigned what, trust in the whole process starts to slip. And you end up spending more time explaining outliers after the fact instead of focusing on the actual program work.

Structural fixes that actually work

The key is not to leave fairness to individual reviewers’ good intentions, but to build it into the process itself. Many organisations already lean on multi-stage review to help with this:

At Good Grants, 77 per cent of clients use multi-stage judging, where applications move through several rounds, such as an initial shortlist followed by a more detailed assessment. That structure creates natural room for the checks and calibration described below.

Scoring controls. Decide how many applications each reviewer handles and make sure assignments overlap. That way you can spot, later on, who’s consistently scoring stricter or more generously than their peers.

Minimum and maximum thresholds. Caps on individual criteria stop a single weak point from sinking an otherwise strong application, and stop one standout strength from carrying a weak one.

Blind review. When reviewers don’t know an applicant’s name, organisation or funding history, the assessment focuses more squarely on the substance of the application itself. This cuts down on unconscious bias considerably. We go into more detail on anonymised scoring and other scoring controls in our guide to fair grantmaking.

Calibration sessions. Before the real scoring round begins, have reviewers score two or three sample applications together and talk through the results. Sessions like these align everyone’s standards before actual decisions are riding on them. We walk through what one of these sessions can look like in practice in Is your grant review process really as fair as you think?

A modern grant management platform like Good Grants supports all of this directly. You can set up rubrics with clear criteria, control assignment overlap, automatically hide names for blind review and see scoring discrepancies between reviewers at a glance. Fairness becomes part of the process by design, rather than an extra task on top of it.

Calibration isn’t a one-off task

A common mistake is running calibration once at the start of a program and then forgetting about it. Scoring standards shift over the course of a review round, particularly when reviewers are working through a lot of applications in sequence. Fatigue, comparisons with previous applications and time pressure all quietly nudge scores in one direction or another.

Build in short check-ins during the scoring window itself. A quick comparison halfway through is often enough to catch bigger discrepancies early and talk them through before they shape the final decisions.

A step towards fairer scoring

Scoring variance won’t disappear with a single fix. It takes structure: clear criteria, calibration that happens more than once, and transparency about how individual reviewers’ scores compare with one another. Each of these steps makes your program a little fairer and your decisions a little easier to stand behind.

A good next step is to look at how consistently your own reviewers are currently scoring. Often, just glancing at the spread of scores per reviewer is enough to show you where calibration would help most. Good Grants is there to help you turn that insight into fair, effective decisions, one step at a time.

Categories

Follow our blog

This field is for validation purposes and should be left unchanged.
Name(Required)
Katia Ernst

Katia Ernst

Katia is a content specialist for the DACH market at Good Grants. She localises content for German-speaking audiences and writes about awards, grants, and program management. When she’s not working, you’ll probably find her performing improv theatre, practising yoga, or reading a good book at the beach.