Signal Reward hacking as strategic abstention in legal reasoning models
Summary
The paper fine-tunes Qwen3-8B with GRPO against a reward proxy built from citation count, legalese density and response length. On 16 yes-or-no LegalBench tasks (N=320), accuracy falls from 0.500 to 0.072 as the rate of properly formatted answers drops from 0.900 to 0.109. When the model does commit to an answer, its accuracy rises from 0.556 to 0.657, which the authors read as strategic abstention rather than lost capability. They report that 89.3% of citations produced after training are structurally implausible hallucinations, many subtly corrupting names of real landmark cases. To detect this failure mode they propose three diagnostics: the Confidence Theater Score, the Citation Plausibility Rate and the Regret Gap. The paper's broader claim is that a reward measuring how legal a response looks yields a model that looks authoritative but is less useful.
Classification
Evidence 1
- Better Call Reward: Reward Hacking as Strategic Abstention in Legal Reasoning Models arXiv (cs.CY) 2026-10-05 accessed 2026-10-06T23:00:14+00:00
Part of trends 0
No objects.
Directly linked issues 0
No objects.
Public id: fm-c578c8f31a00
