Before MSH was released I posted a review of preseason grades for the previous four sets. This post brings that project up to date for MSH.
I looked at 9 sets of grades (Both Limited Level-Ups hosts, both contributors to drafttierlists, Limited Resources, casual_spike, Cardgamebase, and Draftsim (Set Review and Tierlist ). As ever we are only looking at commons and uncommons.
Chart here.
Some grades get updated after preview season ends. This is not useful for us, and so where I couldn't find them unadjusted the grades are simply omitted. I'm particularly interested in finding pre-release drafttierlist grades (their change tracking is somewhat tricky to parse, and I haven't tried to use it to recreate older set grades).
As you can see drafttierlist and Limited Level-Ups (as ever) did quite well; for everyone else this was a challenging set to grade (I didn't include Draftsim's Tierlist in any averages; not sure what's going on with that). To an unusual degree the graders' ability to assess how well a card would perform was tied up with how highly they assessed one color (White) and dismissed a second (Red). 5 of the 10 most underrated cards had White, 7 of the 10 most overrated had Red.
One thing I struggled with when putting together the previous post was whether or not to adjust GIH WR by average pick to get closer to some notion of true card strength and, if so, how. I couldn't quite crack that to my own satisfaction. But, since then, /u/oelarnes posted his own project (found here) in which he did just that (and more). That's a useful resource and it seemed to me possible (perhaps likely) that preview grades might track GIH WR more closely if you adjust for what he calls pick equity. I set aside his other adjustments for now.
Chart here.
As you can see there was a small but meaningful increase in correlation across the board. This is a neat result, and I think it points to some combination of two things. One, insofar as adjusting for pick rate more accurately captures 'true' card strength, then graders are doing better than they appear at judging card strength. Two, insofar as graders are judging card strength in relative isolation, then the adjustment proposed by oelarnes is at least roughly correct and an improvement over naive GIH WR.
The other thing that's clear from the above is that naively normalizing and averaging grades is on average slightly better than following the best graders.