Reassessing Extractive QA Datasets at Scale: LLM-as-a-Judge and In-Depth Analyses

ArXi:2504.11972v3 Announce Type: replace Extractive QA tasks are commonly evaluated using Exact Match (EM) and F1-score, but these metrics often fail to reflect true model performance. Recent studies have proposed using large language models (LLMs) as judges (LLM-as-a-judge), yet they often lack comprehensive evaluation across datasets and overlook key factors such as sensitivity to answer types, prompt variations, and self-preference bias.