Every decade or so, American higher education rediscovers grading. The conversation tends to arrive with a familiar sense of urgency, usually built around the claim that something has gone terribly wrong. The most recent version, sharpened by Harvard’s debate over capping A grades, returns to a worry that never fully goes away: grade inflation has squeezed out the distinctions that make transcripts readable. The numbers at Harvard aren’t really in dispute. Median GPAs climbed from 3.64 for the Class of 2015 to 3.83 for the Class of 2025. The median grade is now solidly in the A range, which means honors designations and internal distinctions now rest on thinner and thinner margins.
Reform proposals are circulating — caps on top grades, percentile-based honors, transcript recalibration — all trying to put some distance back between students. What’s surprising is how fast the conversation has jumped to fixes without settling a more basic question first. Before anyone redesigns the grading system, it’s worth asking what grading is supposed to do. That sounds obvious. It isn’t. Grades don’t do just one thing, and the different things they do don’t always point in the same direction.
At the most immediate level, a grade evaluates work; it registers how well a student understood the material, built an argument and engaged with evidence. In that sense, grades are part of learning itself — a signal embedded in the process. But they don’t stay there. They travel outward to graduate programs, fellowship committees and employers who use transcripts as shortcuts for information they don’t otherwise have. In that role, differentiation isn’t just nice to have. A transcript that can’t tell one level of achievement from another has failed at one of the main jobs it’s expected to do.
Grades also shape behavior in ways that get little airtime in reform debates. Students adjust to incentive structures whether or not anyone intends them to. Which courses you take, how much intellectual risk you’re willing to assume, what kind of work feels worth doing — all of it bends, at least a little, around how evaluation is structured. A compressed grading environment produces one set of habits, a more spread-out one produces another. Neither is neutral, and both leave marks on how students actually experience their education.
There’s a third piece that tends to stay in the background even though it shapes the conversation. Grading practices send a signal about academic seriousness to people outside the institution, whether anyone intends that or not. A transcript is never just a record of what one student did. It also says something about how seriously the school takes evaluation itself.
Look at grading that way — not as a single tool but as something doing several jobs at once — and the debate starts to look different. What first seems like an argument about standards turns out to be about institutions being pulled in different directions simultaneously, each for reasons that make sense on their own terms.
Selective schools make this harder still. When admissions already filter for high achievement, you start with a classroom full of students who were near the top of every prior distribution. Finding meaningful differences within that group requires finer distinctions than in a more mixed setting. But variation is still there: anyone who has spent time evaluating student work knows this. The classroom measures what students produce, not their admissions file, and compression at the top end hides differences that still matter educationally.
At a liberal arts college like Macalester, this isn’t abstract. Classes are small enough that professors know the range of student work in detail, and students usually do too. The gap between solid work and genuinely exceptional work is visible — in class, in feedback, in how a conversation goes — even when both end up in the same grading band. That closeness is part of what makes these institutions worth attending. It builds a collaborative intellectual culture that’s hard to replicate at scale. But it also makes grade compression harder to ignore. When distinctions that are obvious in the room vanish on paper, grading starts to feel less like evaluation and more like translation, a rough rendering of judgments everyone present already understood but the transcript can’t quite carry.
None of this leads to a clean fix. Caps might widen the spread, but probably at the cost of a more competitive classroom atmosphere. Leaving distributions where they are avoids that pressure, but evaluative signals get harder to read the more students cluster near the top. Every adjustment moves costs around rather than removing them, which is part of why reaching for technical solutions first feels like getting ahead of the problem.
To start, universities would do better by clarifying what they think academic evaluation is for. If the primary purpose is developmental feedback, the system should be built around that. If it’s labor-market signaling, the design logic is different. Right now, grading tries to do all of it at once. This may be unavoidable, but it means the tensions are baked in, and no procedural fix is going to dissolve them.
Debates over grade inflation are really arguments about institutional purpose dressed up as arguments about mechanics. Until the underlying question gets asked directly — what is grading for, and which purpose comes first when they conflict? — the technical reforms will feel incomplete. You can’t calibrate a measurement system until you’ve decided what you’re trying to measure.
