REHEARSE: Experiential Rehearsal for Verbal Confidence Calibration in Large Language Models
arXiv CS.AI
•
Generative AI
Large language models (LLMs) often express verbal confidence that is poorly aligned with actual correctness, limiting their reliability in safety-critical applications. Existing prompt-based methods treat calibration largely as a one-shot inference problem, relying on either instance-level reasoning or post-hoc self-assessment. We