Cactus: Predicting the Conservation Status of Species
The Cactus exploratory action investigates a new approach to more rapidly estimate the conservation status of a large number of species. By combining artificial intelligence, statistical learning, ecological modeling, and large biodiversity databases, the project aims to better identify potentially threatened species. Observations gathered from Pl@ntNet provide a particularly valuable source of information for this research.
In the face of accelerating climate change, habitat transformation, and numerous pressures exerted by human activities, assessing the conservation status of species has become a priority. These assessments notably help determine which species and environments require protection, guide conservation policies, and anticipate the impacts of development projects.
The Red List of the International Union for Conservation of Nature (IUCN) currently serves as the global benchmark in this field. Each assessed species can be classified into a category reflecting its level of threat: “Least Concern”, “Vulnerable”, “Endangered”, “Critically Endangered”, or “Extinct”.
Scaling Up with Artificial Intelligence
Cactus explores whether it is possible to assess thousands of species simultaneously. The goal is not to replace IUCN expertise or its official assessments, but rather to develop complementary methods capable of:
- Detecting species whose status appears concerning;
- Flagging those that should be assessed or reassessed as a priority;
- Generating large-scale indicators;
- Helping specialists process volumes of data that have become too large for manual analysis.
Interconnecting Vast Data Sources
Digital infrastructures dedicated to biodiversity currently aggregate billions of data points: observation dates and locations, digitized herbarium specimens, biological traits, climate data, soil properties, land use, and satellite imagery. Cactus studies how to cross-reference these sources to gain a more comprehensive understanding of each species and its environment. In particular, the project draws upon:
- Occurrence data indicating where and when a species was observed;
- Observations generated by citizen science initiatives such as Pl@ntNet;
- Natural history collections and digitized herbarium specimens;
- Data on species traits;
- Environmental information such as climate, soil, vegetation, and land cover;
- Projections modeling potential changes in climate and habitats.
Based on this information, the models learn the connections between species, their habitats, and environmental conditions. They can then be used to estimate species distribution ranges, anticipate potential range shifts, and identify conditions likely to increase extinction risk.
Three Major Scientific Challenges
Having vast amounts of data does not mean it is complete or immediately usable. Cactus must address three particularly significant challenges:
- Absence is difficult to prove: Most biodiversity databases contain presence-only data, indicating that a species was observed at a specific time and location. In contrast, they provide very little reliable absence data.
- Observations are unevenly distributed: Available data reflects the habits and access limitations of observers. Some regions are heavily surveyed, while others are barely covered. Areas close to cities, roads, or trails are generally far better documented than remote or hard-to-reach areas.
- The majority of species remain under-documented: While a few species have an extensive number of observations, most are represented by very few data points. This imbalance, known as a “long-tail” distribution, represents a major challenge for artificial intelligence.
The Role of Pl@ntNet
Citizen science has substantially increased the volume and frequency of available plant observations. Every identification shared on Pl@ntNet can provide valuable data on the presence of a species in a specific place and time.
The project builds upon research conducted around Pl@ntNet on species distribution modeling using citizen-contributed observations. Over time, the methods investigated in Cactus could help maximize the value of the millions of observations collected by the platform. Pl@ntNet would thus evolve beyond an identification and sharing tool into an observatory capable of contributing to the detection of shifts in species distribution and identifying early warning signals for conservation.
Toward Decision-Support for Conservation
The project outlines a multi-stage process. Models first learn suitable habitats for species along with certain inter-species relationships. They can then estimate changes in populations or geographical ranges across different scenarios. Finally, this information is used to estimate the probability that a species falls into a specific threat category.
This approach remains exploratory. The available data is fragmentary, heterogeneous, and sometimes uncertain, and there is no guarantee that it contains all the necessary information on its own. Automated predictions must therefore be accompanied by confidence metrics and interpreted with caution.
Research at the Crossroads of Multiple Disciplines
Cactus brings together complementary expertise in artificial intelligence, statistical learning, ecology, modeling, and botany. By connecting these disciplines, the project paves a new way forward for large-scale biodiversity research. Beyond conservation status, the models developed could help provide a better understanding of where species live, how they interact, how they respond to environmental changes, and how the ecosystems they depend upon evolve.
- Program: Inria Exploratory Action 2020
- Fields: Biodiversity Conservation, Artificial Intelligence, Statistical Learning, Ecological Modeling, Citizen Science
- Link: https://www.inria.fr/fr/cactus-prediction-du-statut-de-conservation-des-especes