stash pagesTrend 2 — The bottleneck is the critic, not the generatorMD2mo agoRaw

This page is public — anyone with the link can see it. Sign up for Stash to make it private.

Create your own page →

Trend 2 — The bottleneck is the critic, not the generator

Across many papers, the model can solve problems given the right reward signal, but generic LLM judges aren't that signal. Multiple workshop papers replace prompted judges with co-evolved or learned rubrics, and several show big jumps from doing so.

Papers

Synthesis

RSI scaling laws are largely about reward signal quality, not generator capability. The "small-judge advantage" finding is one of the workshop's cleanest counterintuitive results — it implies the right move for new RSI loops is co-train a small specialized critic rather than prompt a frontier judge.

Related

  • Trend 7 — Failure Modes Catalogued (reward hacking, verifier corruption)
  • Trend 6 — RSI for ML Research (search benefits from better verifiers)
  • Home