Language Gaps in Biodiversity Knowledge

Multilingual Wikipedia and LLMs as instruments for measuring where conservation knowledge is missing

Line-art illustration: a tree whose foliage is characters from different writing systems.

Research Question

Why is biodiversity knowledge unevenly distributed across languages and cultures — and how does that unevenness propagate into large language models and data infrastructure that conservation increasingly relies on?

Background

Conservation knowledge skews heavily toward English-language information, leaving knowledge held in other languages underused. Two published foundations set up this project: BioWikiNet, a dataset linking 1.27 million Wikipedia articles in 11 languages to the GBIF backbone taxonomy, and a benchmark of large language models on 21,955 IUCN Red List species showing that LLMs classify taxonomy well (94.9%) but systematically fail at conservation status (27.2%).

Approach

Quantify knowledge disparities across language editions of Wikipedia for roughly 2,000 threatened species, and test how those disparities surface in LLM behavior — including the risk that Data Deficient species with thin multilingual coverage stay invisible to both humans and models. Supported by JSPS KAKENHI Grant-in-Aid for Scientific Research (C) 26K09459 (FY2026–2028); see Funding.

Representative Findings

Findings from the KAKENHI project will appear here as they are published.

Data & Code