Rubrics as Rewards: Reinforcement Learning Beyond Verifiable Domains
Unknown (OpenReview: 21UFlJrmS2) · paper · 2025
01
Summary
Proposes rubrics as a reward source for reinforcement learning in domains where a crisp verifiable outcome does not exist. A deliberate extension of RLVR-style methods past the easy cases.
02
Why it matters
The library's schema intentionally separates reward richness from repairability and input legibility. This entry is a technical illustration: a rubric can score lower than a verified outcome on richness while scoring *higher* on repairability — because the rubric names which dimension failed.
03
Source
04
Notes
An important entry for preserving the library's most-subtle disagreement: verifiable outcome ≠ diagnostic feedback. Rubrics as Rewards operationalises that distinction by trading some of the "cleanness" of a verified outcome (one bit: passed / failed) for the structured richness of a rubric (multi-dimensional, failure-mode-named, repairable). Read alongside • Expanding RLVR Across Diverse Domains — the verifiable-outcome pole. • Royzen: Standard Signal — the domain where the outcome is unusually verifiable. • Wallach/Jacobs et al — the measurement critique that applies to both rubrics and verifiable rewards.
05
Note on sourcing
OpenReview URL and title verified. Authors and exact date not confirmed; confirm from the forum page before citing.