DiamondOps
55%of the cards we flag get upgraded

against a 26% background rate2.1× better than chance

1,061 calls across 4 scored roster updates, measured against 4,889 scored cards. Radar flags candidates before an update is announced; it does not guarantee upgrades.

Signal 70 and above: 76% upgraded — 3.0× better than chance, on n=165.

What actually happened, by signal strength

Every scored card sorted into 10 signal bands, and the share of each band that San Diego Studio actually upgraded. The dashed line is the background rate — the share of all scored cards that upgraded. Bars above it did better than chance.

Share of the band that upgradedBackground rate 26%

Faded bars have fewer than 50 cards behind them — signal 80–89 (n=39), signal 90–99 (n=28). That is not enough to separate a real edge from luck, so we do not lead with them. signal 90–99 posts the highest rate on the chart, on one of the smallest samples.

Show the numbers
Observed upgrade rate by Diamond Radar signal band, against a background rate of 26%
Signal bandUpgradedCardsvs chance
0–92%1,0400.1×
10–199%9170.3×
20–2923%9080.9×
30–3932%7651.3×
40–4945%5731.8×
50–5958%3502.2×
60–6968%1712.7×
70–7973%982.9×
80–8979%39(thin)3.1×
90–9982%28(thin)3.2×

Accuracy detail

Mixed methods

46%

Upgrade recall

Of real upgrades, how many DR caught

4

Windows scored

May 8Aug 14

1,061

Total predictions

584 correct

60%

Avg per-window

Across 4 scored windows

Methodology improvement

v13 windows · 619 predictions

Precision

65%

Hand-crafted attribute-gap heuristic + frequency-learned weights + hand-tuned boosts.

v51 window · 442 predictions

Precision

41% 23.4

Served valueGap baseline re-anchored (diamondops-7f2g2): the downside is measured from the first half of the CURRENT approved-major window instead of a never-refreshed window pinned to the card's earliest-ever signal crossing. Restores a real downside to 65% of degraded rows (incomplete estimates 39% -> ~16%) and abstains where no trustworthy baseline exists.

Each version ran on different roster windows — an observed trend, not a controlled comparison.

Diamond Radar flags players who may be upgraded at the next roster update — it doesn't guarantee upgrades. Recall measures how many actual upgrades DR identified; precision measures how often flagged players were actually upgraded.

Accuracy by tier

Red Diamond
not enough data — 2 calls (need 20)

recall not reported — only 0 real upgrades at this tier

Diamond
31%

18 of 59 calls correct · recall 56% of the 32 real upgrades at this tier

Gold
36%

27 of 75 calls correct · recall 34% of the 80 real upgrades at this tier

Silver
53%

89 of 167 calls correct · recall 41% of the 215 real upgrades at this tier

Bronze
62%

282 of 457 calls correct · recall 45% of the 626 real upgrades at this tier

Common
56%

168 of 301 calls correct · recall 55% of the 307 real upgrades at this tier

Precision — how often a flagged card at this tier was actually upgraded. Recall — how many real upgrades at this tier DR caught.

Tier-jump recall

54% — caught 193 of 357 real tier jumps

Gold→Diamond
9/21
Silver→Gold
36/68
Bronze→Silver
68/135
Common→Bronze
80/133

Of cards that crossed a rarity boundary, how many DR flagged for an upgrade.

Precision by confidence

How often a flagged card was actually upgraded, split by the confidence attached to each call. Bands are independent — a higher band does not always score better.

High
58%425 of 730
Medium
48%49 of 103
Low
48%110 of 228

Miss profile

The two ways a prediction can be wrong: calling an upgrade that did not happen, and missing one that did.

477

False alarms — we called it, it did not happen

93 of those were strong-signal calls — the model called them loudly and was wrong. Signal strength is a separate measure from the confidence bands above — the two are not the same thing.

676

Missed movers — it happened, we did not call it

Near miss we came close to calling it
264
Weak signal we saw something, but did not call it
311
Blind we did not see it coming
101

Rating accuracy

When a card was upgraded, how close the predicted overall was. Separate from whether the upgrade was called at all.

21%

Exact OVR

50%

Within 1

1.93

Avg error (OVR)

Across 584 scored upgrades. We predicted +1.2 on average against an actual +3.0 — the model under-calls how far a card moves.

By roster update window

15 windows with no major SDS changes hidden.

← Back to Diamond Radar · Methodology

What a “call” is

A call is a prediction Diamond Radar makes before San Diego Studio publishes a roster update — not an explanation written afterwards. Radar flags a card as an upgrade candidate while the update is still unannounced, and that flag is timestamped and frozen. When the update lands, the card either went up or it didn’t.

That ordering is the whole point, and it is what makes this page checkable. Anyone can explain a rating change after the fact. The calls above were on record first, and every scored update below shows them next to what San Diego Studio actually did.

How Radar decides a card is a candidate

The core idea is a gap. Every rated attribute on a card maps to something the real player actually does on a baseball field — contact and power map to how he hits, a pitcher’s ratings map to what he gets out of hitters. Radar reads the player’s real MLB production and compares it against what his card currently claims. A player producing far above his card’s ratings has an upward gap, and a persistent gap is what a roster update tends to correct.

A single hot week is noise, so the gap is measured over several time horizons at once — the last few days, the current update window, the past month, the season — and blended, weighted toward recent form. Horizons with no data are dropped and the rest re-weighted between them, so a player who missed a month is judged on what exists rather than penalised for the gap in the record.

Each horizon carries a confidence from sample size, and the bar adapts to the horizon rather than being fixed: a typical hitter accumulates a few dozen plate appearances in a week and several hundred across a season, and starters and relievers are held to their own workloads. A part-time player’s hot stretch counts for less than an everyday player’s, because it should.

The threshold is also tier-aware. High-overall cards have less headroom — a 95 has fewer attributes that can plausibly move — so their natural gaps are smaller, and a signal that means nothing on a Bronze card can be significant on a Diamond. A single fixed cutoff would flag low-rated players constantly and elite ones almost never. Radar predicts a direction and a size, and it can flag cards trending the other way too.

How to read this page

The headline is a pair, and it is meant to be read as one: the share of the cards Radar flagged that were upgraded, next to the share of all scored cards that were upgraded. The second number is what makes the first one mean anything. A hit rate on its own tells you nothing about whether the signal helped, because it does not say what would have happened without it.

The curve sorts every scored card by how strongly Radar flagged it, and shows what actually happened to each group. It is a count of outcomes, not a claim about them.

Nothing here is held back for a paid tier. A number can still be missing — either there is not yet enough behind it to report honestly, or it was not recorded for that stretch — and the page says which, in its place, rather than filling the gap with a figure we would not stand behind.

Roster updates are a human decision at San Diego Studio, and real-life production is one input into that decision, not the whole of it. A player can rake for a month and get nothing; another can be adjusted for reasons no model can see from a box score.

There is also a ceiling, and it is worth stating plainly: the strongest signal Radar produces does not reach certainty. No honest reading of this page should suggest a card is guaranteed, and anyone promising that is guessing.

Why most roster updates aren’t on this page

A roster update is San Diego Studio re-rating players in MLB The Show based on real-world performance. They arrive regularly, but most carry no meaningful attribute changes — nothing to predict and nothing to score, so no scorecard exists for them and none is invented here. Of the 19 update windows on record, 4 produced changes worth scoring and 15 carried none.

That is also why this record grows in months rather than weeks, and why a scored update is worth reading in full: each one is a batch of predictions resolved at a single moment, against a set of ratings that had not moved in weeks.

Every scored roster update

One page per update — the cards Diamond Radar called before San Diego Studio published it, and what they actually got.