Transforming your observations into scientific knowledge: an interview with Raphaël Benerradi on the use of Pl@ntNet data for research

Could you introduce yourself and tell us about your background?

My name is Raphaël Benerradi, and I am a second-year PhD student within the IROKO team at Inria, based at the LIRMM laboratory in Montpellier. My PhD focuses on a topic at the interface between computer science, statistics, machine learning, and ecology.

I first studied computer science and mathematics at CentraleSupélec in Paris-Saclay, before completing my training with a dual degree at AgroParisTech, specialising in agronomy. This interdisciplinary background now allows me to work at the intersection of these two fields.

What is your connection with Pl@ntNet?

Pl@ntNet brings together several teams, and I am part of the one working on the computer science aspects of the project. We collaborate with the engineers and project leaders to exploit the data collected by the application from a research perspective. These discussions also provide opportunities to suggest new directions for the development of the application, ensuring that it meets both the needs of users and the challenges faced by science, particularly in the field of biodiversity conservation.

I am supported by a doctoral contract funded through a national doctoral school grant. My connection with Pl@ntNet mainly comes through the research team and my supervisors, Alexis Joly, Christophe Botella, and Maximilien Servajean.

Did you know about Pl@ntNet before joining the team?

Yes, but only as a user. I used the application to identify plants, without knowing that it was a French project, let alone one based in Montpellier.

When looking for a PhD topic, I knew I wanted to work on something combining computer science and statistics with ecology. It was only when I joined this PhD project associated with the Pl@ntNet team that I discovered the full scientific dimension behind the application.

Looking back, it actually makes a lot of sense: millions of people producing millions of observations naturally represent an exceptional resource for answering many research questions, particularly those related to biodiversity monitoring.

Could you tell us more specifically about your PhD topic?

I work on estimating spatial and temporal changes in the distribution of plant species. The objective is to produce maps showing where plants are present and to study how their distribution changes over time.

To do this, I use opportunistic data from Pl@ntNet, meaning observations freely collected by users. By combining statistics and machine learning, I develop models capable of estimating the probability of species presence across space and time.

What are the main challenges of this work?

The observations collected through Pl@ntNet are extremely numerous, but they are also biased because they do not follow any predefined sampling protocol. Users photograph plants wherever and whenever they choose, which naturally leads to an overrepresentation of certain locations, such as cities or tourist sites, and certain periods, especially spring. Conversely, some locations and periods of the year are underrepresented.

My work therefore consists of correcting these biases in order to distinguish genuine ecological changes from effects related to observation behaviour.

I develop statistical models that separate two different processes. The first describes the ecological factors explaining the presence of a species, such as climate or land use. The second models the observation process itself, meaning the probability that a species is actually detected by users depending on location, season, or other influencing factors.

I then enrich these models using machine learning approaches, particularly neural networks, to improve their accuracy and their ability to process the very large volumes of data available through Pl@ntNet.

A graphical representation of the number of observations collected through Pl@ntNet over the years, highlighting a clear example of observation bias: users submit far more observations during the spring season.

Are these methods entirely new?

In practice, no. They build upon existing research. However, in our case, we adapt them to the specific context of Pl@ntNet and its data. Occupancy models already existed in the scientific literature, but they were mainly used with data collected according to strict scientific sampling protocols.

These protocols produce higher-quality and less biased data. However, such surveys require careful organisation, trained citizen participants, and significant financial and time investments. Moreover, these approaches often focus on a specific area or a particular group of species of interest.

Our goal is not to replace botanical surveys or vegetation monitoring efforts, but rather to provide an alternative approach capable of producing results at a much larger scale. The objective is to use these massive datasets to obtain results that are close or comparable to those generated by studies based on strict sampling protocols. However, the aim is not to replace these established scientific monitoring approaches, which remain essential, but rather to complement them.

The massive datasets generated through Pl@ntNet make it possible to extend analyses to much larger spatial scales and to a greater number of species than would be feasible using traditional protocols alone.

My work therefore involves dealing with these massive opportunistic datasets from Pl@ntNet, developing computational approaches capable of processing these very large data volumes, and enriching them through deep learning methods developed within the team. The goal is to extract scientifically robust and usable results, even though the underlying data may contain biases.

What is the main contribution of Pl@ntNet to your research?

The essential contribution comes from the observations produced by users. My work relies on these data, but also on the quality of their identification. Changes in the identification algorithm can influence the observations available; this is why my models can also help detect potential anomalies or errors introduced during algorithm updates.

I carry out this work in particular in collaboration with Vanessa Hequet, a botanist at Pl@ntNet, who is specifically interested in detecting errors, anomalies, and biases.

For example, if a sudden change is observed in the data concerning a species, it may reflect a real ecological change, but it could also result from changes in user behaviour or from modifications to the identification algorithm. By identifying these different situations more clearly, it becomes possible to assess the quality of identifications, improve the algorithm, and, if necessary, strengthen the validation processes carried out by botanists.

This work is still exploratory, but it demonstrates how research can directly contribute to improving the quality of the data produced by Pl@ntNet.

In your opinion, what is the scientific value of this research, and how does it benefit Pl@ntNet?

Pl@ntNet data provide a unique opportunity to develop new statistical and machine learning methods capable of correcting the biases inherent in citizen science datasets.

Beyond the methodological interest, this research can generate valuable knowledge for biodiversity conservation, for example by detecting declines in plant species and helping guide conservation measures.

For Pl@ntNet, these studies make it possible to better exploit the data generated by users and to gain a deeper understanding of the identification algorithm, thereby improving its reliability and strengthening the confidence of both experts and users.

Are there other data sources comparable to Pl@ntNet?

There are several similar platforms, such as iNaturalist, which also collects opportunistic observations but covers a broader range of living organisms. Pl@ntNet nevertheless stands out because it focuses specifically on plants and because it is designed for a very broad audience, allowing it to collect observations from users with highly diverse backgrounds.

This diversity provides information that complements data from more structured citizen science programmes, such as Vigie-Nature, for example, whose datasets are generally more precise but much smaller in volume.

Pl@ntNet data therefore make it possible to cover much larger territories, including areas that are poorly studied through traditional protocols, such as urban environments. The two approaches are thus complementary: structured monitoring programmes provide highly accurate observations, while Pl@ntNet provides an unparalleled quantity of data.

Indeed, citizen science now plays an essential role, particularly in biodiversity monitoring. How do you see its place in current research?

Researchers alone cannot produce the volumes of data required to monitor changes in species distributions at large scales. The contributions of users, often made without them even realising that they are participating in research, therefore represent a valuable resource.

In the case of Pl@ntNet, the very large number of observations provides an exceptional opportunity to address scientific questions that would otherwise be completely inaccessible. Research now also relies on this collective contribution from citizens.

Would this type of research have been possible twenty years ago?

Probably not at this scale. Although natural history collections and biodiversity records already existed, digital platforms have profoundly transformed the possibilities available to researchers by making it possible to gather millions of observations.

Today, I am working with a dataset containing around 2.6 million high-quality observations covering nearly 10,000 plant species. It is still impossible to monitor all of these species individually, but these volumes already make it possible to analyse several hundred, or even a thousand species, which would have been extremely difficult to achieve twenty years ago using traditional tools and methodologies.

Platforms such as Pl@ntNet have therefore opened up new perspectives for biodiversity monitoring, while still complementing more traditional scientific approaches.

Looking ahead five years, what developments would you like to see in Pl@ntNet?

One possible application emerging from my PhD research would be for the application, at the moment of plant identification, to indicate whether a species is increasing or declining in a given region, or to highlight that an observation is particularly rare and scientifically valuable.

From the user perspective, I would also like the application to become more explainable. Beyond simply providing an identification result, it could explain why a species has been suggested, highlighting the plant characteristics that led to this prediction and, if necessary, recommending that users photograph specific details or take another angle to improve the identification.

In my view, this ability to explain the decisions made by the algorithm represents an important future development.