Benchmarking large language models for biodiversity assessment: Performance and biases across IUCN Red List species

Author

Shinya Uryu

Doi

Published in Ecological Informatics, vol. 96, article 103839. doi:10.1016/j.ecoinf.2026.103839 · arXiv:2510.02830

Evaluates five large language models on 21,955 IUCN Red List species across four tasks. LLMs handle taxonomic classification well (94.9%) but systematically fail at conservation status assessment (27.2%) — documenting how closed-book LLMs fail is itself the contribution.