Equity Research

Two ESG Raters Can Score the Same Company in Opposite Directions

Credit ratings from different agencies agree almost perfectly. ESG ratings from different providers agree about as often as a coin flip that has been weighted slightly.

↩ Looking BackPart of the 2020 to 2026 retrospective, written in July 2026. The date below marks the 2023 events this piece revisits, not when it was published, so it draws on everything known through mid 2026.
Nathan Xiang·June 14, 2023

The Finding

Researchers at MIT compared ESG ratings for the same companies across six major providers and found the average pairwise correlation was roughly 0.5.

For comparison, the long term ratings that Moody and Standard and Poor assign to the same issuers correlate at approximately 0.99. Two credit agencies almost never disagree about direction. Two ESG raters frequently do.

A correlation near 0.5 means a company in the top quartile at one provider can plausibly sit in the bottom half at another.

If your independent variable changes rank depending on which vendor you bought it from, it is not measuring one thing. It is measuring several different things that share a name.

Where the Disagreement Comes From

The useful part of the research was decomposing the divergence into three sources.

SourceMeaningShare of divergence
MeasurementSame attribute, different indicators and dataLargest
ScopeDifferent sets of attributes included at allSubstantial
WeightSame attributes, different importanceSmallest

The intuition most people start with is that raters simply weight things differently, and that turns out to be the least important cause. The bigger problem is that two raters trying to assess the same concept, say labour practices, choose different observable proxies and get different answers.

Why Measurement Diverges So Much

Credit rating has an outcome to anchor on. Default either happens or it does not, and models can be calibrated against decades of realised defaults.

ESG has no equivalent single outcome. There is no event that settles whether a governance score was correct. Without a target variable, there is nothing to calibrate against, and each provider methodology stays a defensible internal opinion.

Two further mechanics widen the gap. Most raters score relative to industry peers, so the peer group definition alone moves a score. And missing data is imputed differently, which systematically favours large companies that disclose more, regardless of how they behave.

The Rater Effect

The research also found a rater specific bias: a provider that scores a company highly in one category tends to score it highly across other categories too.

That is consistent with analysts forming an overall impression and letting it colour the components, which is the same halo problem that shows up in credit committees and equity research. It means the subscores are less independent than their presentation implies.

What This Breaks

A large literature tests whether ESG performance predicts returns, cost of capital, or volatility. Many of those studies use one provider ratings as the measure of ESG quality.

If the measure changes rank across providers, results can flip based on a vendor choice made before the analysis began. Several attempted replications using a different provider have produced weaker or reversed findings, which is what you would expect from a noisy independent variable.

The same problem hits index construction and any product whose holdings are determined by a threshold score.

Using Them Anyway

The ratings are not worthless. They are efficient aggregators of disclosure that would take an analyst weeks to gather, and the underlying indicator data is often more useful than the headline score.

The practical approach is to go down a level. Use the raw indicators rather than the composite, decide explicitly which attributes matter for your thesis, and check whether two providers disagree about a specific company before relying on either. Persistent disagreement on one name is itself a signal that something about that company is genuinely contested.

The Bottom Line

ESG ratings from different providers correlate around 0.5 against roughly 0.99 for credit ratings, and most of the divergence comes from measurement rather than weighting. The root cause is that there is no realised outcome to calibrate against, so no methodology can be shown to be wrong. Treat a composite ESG score as one vendor opinion, work from the underlying indicators, and never treat a single provider score as ground truth in an analysis.

Explore Teen Biz News →