We Retired Our Own Hit Rate. Here's How Diamond Radar Is Graded Now.

Diamond Radar's first season of grading went through three scoreboards: a bare hit rate, a hit rate with the chance rate beside it, and grading by score band. Here's why each one was replaced, what the page says today, and how to read it.

Shaun, the Headghoul of RGL11 min readMethodology

TL;DR. Diamond Radar is graded by score band now. Every card it had the stats to score before a roster update lands in a 10-point band, and each band's upgrade rate sits next to chance: the share of those cards that upgraded anyway. There's no single hit rate any more. I retired it, because the one I published undersold the model and was easy to bend.


The best way to judge Diamond Radar is a table with no headline number in it. It took me a full season and three versions of the accuracy page to get there.

This is that season. The last graded update was October 2, which San Diego Studio called its "Final Regular Season Roster Update," so these are the final regular-season numbers. What I published, why each version got replaced, and how to read the page.

Version one: a hit rate on its own

The accuracy page launched in June leading with one number: the share of Radar's calls that came true.

The update it launched on, June 12, had come in at 65.6%, 145 of 221 calls. I wrote a whole post about the misses instead of burying them. I'd do that again.

What I got wrong was the number itself. 65% sounds like a coin flip plus a bit. Nobody reading it knew what it should be compared to, and neither did the page.

Here's what it should be compared to. Most cards don't upgrade at a roster update. On June 12, only about 29% of the cards Radar graded went up, whatever Radar thought of them. That's chance: the rate the model has to beat.

So 65.6% wasn't a coin flip plus a bit. It was 2.2 times what you'd get pointing at cards blind. The bare number undersold the model, and I'm the one who published it that way.

Version two: the hit rate, with chance beside it

At the end of July I pulled the bare figure off the public page. In August it came back with a partner. The headline hit rate, and the one on every update, was printed next to the chance rate it had to beat, and right under it, the gap between them as "× better than chance".

That fixed the undersell. It didn't fix the other problem, which took me longer to see.

A hit rate only counts the cards Radar flagged, and "flagged" means scored above a line I chose. In June that line moved down from 50 to 40, because at 50 it caught only 145 of the 354 cards that really upgraded. By the September update it was back at 50, because a new version of the model scored cards higher across the board, and at 40 it would have flagged far more cards.

Both moves had a reason. But every time the line moved, the hit rate moved with it, whether the model had changed or not. The May 8 update alone grades at 77% with the line at 50 and 89% with it at 70. Same cards, same model, same scores.

A number I can change by moving a line isn't measuring the model. It's measuring my line.

Version three: every card, by score band

So since mid-September, the grade doesn't depend on a line.

Every card Radar had the stats to score before an update counts, flagged or not. (A card with no stats this season shows a 0 on the board and sits out of the grade.) Each one goes into a 10-point band by the score it held before the update. Then each band gets one honest question: how often did San Diego Studio actually upgrade these?

The page today. No overall hit rate. Every band next to the same bar to beat.The page today. No overall hit rate. Every band next to the same bar to beat.

This can't be bent the way the old number could. There's no line to move. (The page's "70 and above" figures are a fixed cut, the same for every update, so they can't drift the way the old line did.) A model that knows nothing gives you ten bands all sitting on chance. A model that knows something gives you a staircase.

A staircase, and the bottom of it matters as much as the top.A staircase, and the bottom of it matters as much as the top.

Over the 2026 regular season, it's a staircase. Cards Radar scored 90–100 upgraded 85% of the time against 24% chance. Cards it scored 0–9 upgraded 2% of the time. The bottom three bands sit under chance. That's the part I'd look at first, because it means a low score is information too. A model that's only right when it shouts isn't much of a model.

Every step climbs, all the way to the top: 74% in the 80s, 85% in the 90s. After the September update that wasn't true. The 90s had slipped just under the 80s, and the page showed it that way rather than pretend otherwise. October 2 put them back in order.

The glossary

The page uses a small set of words. Here's what each one means and what it doesn't.

Call. A card Radar flagged as an upgrade candidate before San Diego Studio published the update. Every score, flagged or not, is timestamped and frozen while the update is still unannounced, and the band table grades every one Radar had the stats to score. Anyone can explain a rating change afterwards. These were on record first.

Score band. A 10-point slice of Diamond Radar scores: 0–9 up to 90–100. Every graded card lands in exactly one, by the score it held before the update.

Hit rate. The share of a band that upgraded. It's still on the page, one per band, and in the band table it's never alone. The per-update archive further down still shows the old "N of M calls landed" counts from the line. Those are the numbers this post is about retiring.

Chance. The share of all graded cards that upgraded, whatever Radar scored them. It's measured on the cards Radar graded, not on every card in the game, and it's the bar every band has to clear. This season, 1 in 4.

× better than chance. A band's hit rate divided by chance. "3.5×" means that band upgraded about three and a half times as often as a card picked blind.

Sample size. How many cards are behind a number. There are two floors. Under 20 cards, the page won't print a percentage at all, because a rate on that few cards could honestly be anything from bad to excellent. From 20 to 49, it prints the number but won't lead with it. At 50, it can lead.

Top 5 / 10 / 20. The table counts every card Radar graded. These count only the highest-ranked cards on each update's board, which is how the page says the board is actually read.

Model version. Which build of Radar made the call. More on this below, because it's the one most likely to catch you out.

Three terms are retired from the grade: precision, recall and the flag line, or threshold. They were how a call got graded at a line. You'll still see them in older posts, and the per-update archive still counts calls at the old line. The band table doesn't use them.

A call is graded by the model that made it

Radar changed a lot this season. The May, June and July updates were scored by the first version. August was scored by v5. September was the first update scored by v11, the version running today, and October 2 was its second.

The page grades each update by whichever version actually made its calls. It doesn't re-badge old calls with today's version number. Those scores genuinely came from an older model, and relabelling them would hand the current build a track record it hasn't earned.

So the band table is a season's record, not one model's. The page says so, and it adds a separate line for the newest model on its own. Through the regular season: of the 342 cards v11 scored 70 or above in its two graded updates, 71% upgraded, against 22% chance.

Two updates is still two updates. That line will mean more after the next graded update.

The per-update list on the page also answers why there are only six. Because most roster updates are small. Of the 26 on record this season, 6 were big enough to grade. The other 20 are hidden, not missing.

What a score is, and what it isn't

The lesson I'd most like you to take from this season has a name. Nick Kurtz.

Going into the September update, Radar scored him in the 90s while he was on the injured list. When a player stops playing, his recent stretches go empty, and the model sets them aside and leans on what's left. Under the older models he did fade, from 47 on the day of his last game to 24 by late August. The new model at the end of August lifted him to 40. Two days later his last game fell out of the 30-day stretch, the model leaned on his season line alone, and he jumped to 96. He moved once after that, to 84 the day after the September update, and sat there every day through October 2. His card didn't move at the update. SDS left him at 86, and I wrote the whole story up in the September report card. We've kept him on the board on purpose since then, to see how the model handles a player who stops playing. The October 2 report card has what that showed.

Two things came out of it. The model has no injury data, and that's a known gap with work open on it. And the bigger one: a score is a chance, not a promise. The top band upgraded 85% of the time this season. That's very good, and it still means about one card in seven didn't go up.

That's also where Radar differs from picking a few players off a hot week of stats. A hot week is a hunch about the players you happened to look at. Radar scores the Live Series cards it can match to a real player before the update, puts the score on record, and then gets graded on every card it had the stats to score, including the ones it said no to. The method isn't that it's always right. It's that you can see exactly how often it is.

How to read the page in ninety seconds

  1. Check the bar to beat. Every band is read against it. Above it, the score told you something.
  2. Look at the staircase. It should climb, and the bottom should sit under chance.
  3. Check the card counts. A great rate on 30 cards is a lead, not a fact. The page marks those for you.
  4. Read the top of the board. The Top 5/10/20 rows are how you'd actually use it.
  5. Check which model. The season table pools versions. The newest model has its own line.

Then, before you spend stubs on a high score, read what a top Diamond Radar score is actually worth. A call being right and a trade making money aren't the same thing.

Why publish any of this

Because the alternative is asking you to trust a number I won't explain.

A published, checkable record is the one thing I can offer that's hard to fake. That only works if the record includes the update where we looked bad, the score that made no sense, and the number I retired.

I don't trade on Radar. I built it, I watch it all season, and I sit on my hands.

Figures are the final regular-season record, as the accuracy page showed them after the last regular-season graded update on October 2, 2026. Screenshots were taken logged out on the same day.

I'm Shaun. If something is wrong, tell me.