← Daniel Griffin

Rubrics as Rewards: Reinforcement Learning Beyond Verifiable Domains

Unknown (OpenReview: 21UFlJrmS2) · paper · 2025

01

Summary

Proposes rubrics as a reward source for reinforcement learning in domains where a crisp verifiable outcome does not exist. A deliberate extension of RLVR-style methods past the easy cases.

02

Why it matters

The library's schema intentionally separates reward richness from repairability and input legibility. This entry is a technical illustration: a rubric can score lower than a verified outcome on richness while scoring *higher* on repairability — because the rubric names which dimension failed.

03

Source

04

Notes

An important entry for preserving the library's most-subtle disagreement: verifiable outcome ≠ diagnostic feedback. Rubrics as Rewards operationalises that distinction by trading some of the "cleanness" of a verified outcome (one bit: passed / failed) for the structured richness of a rubric (multi-dimensional, failure-mode-named, repairable). Read alongside • Expanding RLVR Across Diverse Domains — the verifiable-outcome pole. • Royzen: Standard Signal — the domain where the outcome is unusually verifiable. • Wallach/Jacobs et al — the measurement critique that applies to both rubrics and verifiable rewards.

05

Note on sourcing

OpenReview URL and title verified. Authors and exact date not confirmed; confirm from the forum page before citing.