Benchmarking large language models for biodiversity assessment: Performance and biases across IUCN Red List species
Published in Ecological Informatics, vol. 96, article 103839. doi:10.1016/j.ecoinf.2026.103839 · arXiv:2510.02830
Evaluates five large language models on 21,955 IUCN Red List species across four tasks. LLMs handle taxonomic classification well (94.9%) but systematically fail at conservation status assessment (27.2%) — documenting how closed-book LLMs fail is itself the contribution.