AI RESEARCH
Retrieval and competition: how a protein foundation model starts a protein
arXiv CS.AI
•
ArXi:2605.16331v2 Announce Type: replace-cross Protein language models are increasingly used to guide experimental and clinical decisions, yet it is often unclear whether a confident prediction reflects recognition of biological evidence or retrieval of a statistical default. We examine this distinction for a near-universal biological rule, that proteins begin with methionine, by tracing the computational pathway through which ESM2-8M produces this prediction. The model does not detect methionine at the masked position.