Biodiversity databases such as the Global Biodiversity Information Facility (GBIF) or Atlas of Living Australia (ALA) contain large amounts of open data, but also face persistent challenges in detecting 'wrong' points; observations of plants and animals that appear to be in the wrong place, be allocated to the wrong species, or both. Detecting these records is notoriously difficult for those who lack expert knowledge of species biogeography.
Currently, the error-detection tools available to these institutions are conceptually very simple, relying on detecting points that are outside of expert-provided polygons, or that are outliers in climate space relative to other members of their species. In contrast, there are no tools that use:
- More advanced statistical methods, such as latent variable models or machine learning;
- Information on data providers that may be indicative of shared 'types' of error (spatial, temporal or taxonomic); or
- Data on species relatedness to suggest alternative taxa that might resemble those in mislabelled observations
We propose bringing some example datasets to workshop alternative statistical methods that may improve detection of errors, relative to traditional methods.
Biodiversity databases such as the Global Biodiversity Information Facility (GBIF) or Atlas of Living Australia (ALA) contain large amounts of open data, but also face persistent challenges in detecting 'wrong' points; observations of plants and animals that appear to be in the wrong place, be allocated to the wrong species, or both. Detecting these records is notoriously difficult for those who lack expert knowledge of species biogeography.
Currently, the error-detection tools available to these institutions are conceptually very simple, relying on detecting points that are outside of expert-provided polygons, or that are outliers in climate space relative to other members of their species. In contrast, there are no tools that use:
We propose bringing some example datasets to workshop alternative statistical methods that may improve detection of errors, relative to traditional methods.