Why ratings drift apart
Managers bring habits to rating. Lenient raters avoid difficult conversations by rating everyone high. Strict raters believe nobody deserves the top score. Central raters give almost everyone a 3. Recency bias lets the last two months outweigh the first ten, and the halo effect lets one strong trait colour every goal. None of this is bad faith, but together it means two employees with the same results can end up with ratings a full point apart.
Spot the pattern before the meeting
Compare the average and spread of proposed ratings for each manager. In a worked example, two regional teams in a Hyderabad distribution company do the same work and hit similar sales and collection targets. Ravi rates his ten people 5, 5, 4, 4, 4, 4, 4, 4, 4 and 3, an average of 4.1. Farah rates her ten people 4, 3, 3, 3, 3, 3, 3, 3, 2 and 2, an average of 2.9. With similar results, a gap of 1.2 points on average reflects rating habits rather than performance, and it is exactly what calibration must resolve.
Running a calibration meeting
The department head chairs, HR facilitates, and every manager in the department attends with evidence for each proposed rating. Discuss the extremes first, the 5s and the 1s and 2s, because they drive increments and exits. Then compare people in similar roles side by side on results against goals. Change a rating only when the evidence supports it, and write down the reason. Managers leave owning the final ratings, since they will explain them to their teams.
- Share proposed ratings and goal results a day before the meeting
- Review top and bottom ratings first
- Compare similar roles across teams, not across the whole company
- Record every change with its reason
- Agree how managers will explain changed ratings
Guideline distributions, not forced curves
A company might publish a guideline such as: we expect a small group at 5, most people at 3 and 4, and a few at 1 and 2. Used as a prompt for discussion, that helps lenient and strict raters meet in the middle. Used as a quota, it becomes a forced curve that punishes strong teams and rewards weak ones, and it is especially unfair in teams of fewer than fifteen people, where the numbers are too small for any curve to be meaningful.
Step by step
- Define the scale with written anchors. Describe what each rating means in behaviour and results. In ZeniaHR, set the rating scale on the review cycle so every team rates on the same scale.
- Share the guideline before rating starts. Tell managers the expected shape of ratings and that calibration will follow, so they rate with evidence from the start.
- Collect proposed ratings by a deadline. Ask for ratings with one line of evidence per goal. ZeniaHR's cycle summary shows reviews by status, so HR can see which teams are still pending.
- Compare averages and spread by manager. Calculate the average rating for each manager's team and look at how ratings spread. Flag teams that sit far above or below the rest without matching results.
- Hold calibration by department. Discuss extremes first, compare similar roles side by side, and change ratings only on evidence.
- Record every change with a reason. Keep a log of original rating, final rating and the reason, so any rating can be explained later to the employee or to senior management.
- Brief managers before they communicate. Make sure each manager understands and can explain the final ratings, especially any that changed in calibration.
See it on your own data
A 30-minute demo on a video call. We set up your departments, shifts and leave rules and show attendance, leave and payroll running for your team. Free for your first 50 employees.
Book a free demoSee pricingFrequently asked questions
What is rating normalization in performance appraisal?
Rating normalization is the process of adjusting performance ratings so they mean the same thing across managers and teams. It usually involves comparing rating patterns, holding calibration meetings where managers justify ratings with evidence, and changing ratings that reflect a manager's habits rather than the employee's results.
What is the bell curve method in appraisal?
The bell curve method asks managers to spread ratings in a fixed pattern, with a few people at the top, most in the middle and a few at the bottom. It can correct lenient rating, but when used as a strict quota, especially in small teams, it forces good performers into lower ratings and damages trust.
Is forced ranking fair?
Forced ranking can be unfair when it is applied as a quota, because it assumes every team has the same mix of performers. It works better as a guideline that prompts calibration discussion. Ratings should finally rest on each person's results against their goals, not on a required percentage per band.
Who should attend a calibration meeting?
The department head, who chairs, every manager whose team is being calibrated, and an HR partner who facilitates and records decisions. Keep the group to one department or function so participants know the roles being discussed. Senior leadership can review the combined results afterwards.