AI RESEARCH

Global PIQA: Evaluating Commonsense Reasoning Across 100+ Languages and Cultures

arXiv CS.CL

ArXi:2510.24081v2 Announce Type: replace To date, there exist almost no culturally-specific evaluation benchmarks for large language models (LLMs) that cover a large number of languages and cultures. In this paper, we present Global PIQA, a participatory commonsense reasoning benchmark for over 100 languages, constructed by hand by over 350 researchers from over 65 countries around the world. The 141 language varieties in Global PIQA cover five continents, 19 language families, and 24 writing systems.