From Rubrics to Reliable Scores: Evidence-Grounded Text Evaluation with LLM Judges

ArXi:2601.08654v2 Announce Type: replace-cross Rubric-based text evaluation increasingly uses large language models (LLMs) as scalable judges, but aligning frozen black-box models with human scoring standards remains challenging. We formulate this challenge as a criteria-transfer problem: the goal is not merely to prompt an LLM to assign a score, but to transfer human rubric intent into a stable, auditable, and human-aligned scoring protocol.