StormEval
Storm-watch bulletin · 2026-10-03 08:53 UTC

Same name, same tasks, every day

Is your AI model getting worse?

StormEval runs the same private tasks against each model every day and compares it with its own past, so a model that gets worse under the same name shows up here with a date.

01 Station reports

ModelStatusPass rateOutput tokensDays of data
Claude Opus 5.5Claude Code · effort medium Waiting for first run – – 0
Claude Sonnet 5.5Claude Code · effort high Waiting for first run – – 0

02 How it works

  1. The same tasks every day.Each model gets 30 exact-answer reasoning problems, tuned to the difficulty it solves about half the time. The tasks are private and never change.
  2. Graded by a program.An answer is right or wrong. No model judges another model.
  3. Compared with itself.The last 7 days are tested against the model's first 14 days, task by task.
  4. Two signals."Degraded" means the pass rate fell. "Changed" means the model started spending a different number of tokens on the same problems while the pass rate held.