Language Gaps in Biodiversity Knowledge
Multilingual Wikipedia and LLMs as instruments for measuring where conservation knowledge is missing
Research Question
Why is biodiversity knowledge unevenly distributed across languages and cultures — and how does that unevenness propagate into large language models and data infrastructure that conservation increasingly relies on?
Background
Conservation knowledge skews heavily toward English-language information, leaving knowledge held in other languages underused. Two published foundations set up this project: BioWikiNet, a dataset linking 1.27 million Wikipedia articles in 11 languages to the GBIF backbone taxonomy, and a benchmark of large language models on 21,955 IUCN Red List species showing that LLMs classify taxonomy well (94.9%) but systematically fail at conservation status (27.2%).
Approach
Quantify knowledge disparities across language editions of Wikipedia for roughly 2,000 threatened species, and test how those disparities surface in LLM behavior — including the risk that Data Deficient species with thin multilingual coverage stay invisible to both humans and models. Supported by JSPS KAKENHI Grant-in-Aid for Scientific Research (C) 26K09459 (FY2026–2028); see Funding.
Representative Findings
Findings from the KAKENHI project will appear here as they are published.