Skip to content

Understanding identification errors in Pl@ntNet

In 2026, Pl@ntNet recognizes approximately 85,000 species. In some cases, an identification error can reinforce itself over time: a common species is regularly confused with a similar-looking rare species, and these new observations then contribute to maintaining the confusion.

For several years, experienced members of the community, known as veilleurs (watchdogs), have been spotting and documenting these situations.

Crossing their work with a statistical analysis allows us to better understand the phenomenon and provide an initial order of magnitude: approximately 1,500 species could be affected, out of the roughly 85,000 species recognized by Pl@ntNet. This estimate remains uncertain and depends on several hypotheses.

The Pl@ntNet identification system relies notably on observations validated by the community. The more data available and correctly identified, the more it can contribute to improving the model.

But when confusion sets in between two closely related species, this mechanism can also amplify an error.

Take a common species, represented by thousands of observations, and a rare species that resembles it. If the model starts suggesting the rare species too often instead of the common one:

  1. some observations of the common species are recorded under the name of the rare species;

  2. these observations increase the amount of data associated with the rare species;

  3. the model progressively learns this confusion and may suggest the rare species even more frequently.

Since rare species generally have fewer observations, a relatively small amount of erroneous data can have a significant effect on their representation.

The watchdogs call this identification drift: an error that does not remain isolated but can become self-sustaining.

Watchdogs are experienced users (confirmed botanists or very knowledgeable amateurs) who spot drifts and participate in their correction through the collaborative validation system.

Over the years, their work has made it possible to build a catalog of documented cases: confused species, the evolution of the problem, and corrections made. Approximately 190 species have been studied in detail.

This expertise is particularly important because an unusual statistical trend alone is not enough to prove that a drift exists. Interpreting the data requires knowledge of the species and their context.

Watchdogs use several terms to describe the different cases encountered.

A source is generally a common species that becomes less frequently correctly recognized. Its observations are progressively attributed to another species.

A sink is the species that incorrectly receives these observations. It is often a rare species, whose observation database can be quickly affected by these errors.

An underdog is also a rare and under-recognized species, but without a drift necessarily being involved. It may simply lack observations or be difficult to identify.

Watchdogs often favor an approach called asymmetric curation: they prioritize cleaning the sink’s data. Correcting a small, heavily contaminated database can be more effective than directly correcting a large database of observations.

An independent analysis studied 795 species with more than 500 observations, representing about 2.5 million observations between 2017 and 2024.

For each species, the analysis estimates whether its probability of being detected by users increases or decreases over time. This evolution is represented by a slope:

  • a negative slope indicates that the species is being detected less and less;

  • a positive slope indicates that it is being detected more and more.

The hypothesis is that a source should generally show a negative slope, while a sink might show a positive slope.

The cross-referencing with cases already known by the watchdogs yields a mixed result.

For sources, negative slopes seem to be a good detection tool: about 80% of known sources appear in this list, with an estimated purity of about 74%.

For sinks, the result is much less useful. Only about 13% of known sinks are found. The main reason is that sinks are often rare and have fewer than 500 observations: they are therefore excluded from the analysis from the start.

Slopes can therefore help prioritize species to be examined, but they do not replace the expertise of the watchdogs.

Why corrections sometimes complicate the analysis

Section titled “Why corrections sometimes complicate the analysis”

Curation itself can modify the statistics used to detect drifts.

When a sink is cleaned, a portion of the removed observations is often re-validated as belonging to the source. The source then gains observations, while the sink loses them.

Over time, this can change their respective slopes:

  • a source that has already been corrected may no longer appear as a species in decline;

  • a cleaned sink may, conversely, appear among species with a negative slope.

It is therefore important to keep a history of corrections. A species already identified as a source may remain vulnerable to new drifts, even after being corrected.

What could be the scale of the phenomenon?

Section titled “What could be the scale of the phenomenon?”

By combining the cases documented by the watchdogs and the results of the statistical analysis, the following estimates are obtained:

  • approximately 450 sources;

  • approximately 1,000 sinks, with a more uncertain estimate between 600 and 1,300 depending on the method;

  • approximately 1,500 species affected in total by identification drift.

For comparison, Pl@ntNet recognizes about 85,000 species.

These figures should be interpreted with caution. The watchdogs’ catalog is not a random sample and focuses primarily on European flora. The statistical analysis only covers species with at least 500 observations and is based on a period from 2017 to 2024.

These results therefore provide an order of magnitude rather than a definitive measurement of the phenomenon.

Can a species be a source in one region and a sink in another? Can we identify the precise moment a drift begins? Is it possible to find sinks starting from already known sources?

In the longer term, the goal could be to develop tools capable of automatically flagging species that are beginning to drift. These tools would not replace human expertise but could help watchdogs spot problems earlier.

Identification drifts are a real phenomenon affecting a limited portion of the species recognized by Pl@ntNet.

The work of the watchdogs is essential for spotting, understanding, and correcting them. Statistical analysis, for its part, provides an additional means of identifying species that deserve special attention.

By combining botanical expertise and data analysis, Pl@ntNet has a better foundation for monitoring these drifts and intervening before they become more significant.