serhii.net

In the middle of the desert you can say anything you want

UNLISTED

07 Aug 2023

Better dataset description draft thing

\section{iNaturalist-300k} \subsection{Dataset collection} We compiled iNaturalist-300k from images of the iNaturalist Research-grade observation collection, with the goal of creating an additional dataset matching our criteria, complementary to Pl@ntNet-300k. Since new observations get added to the iNaturalist dataset every day, the research-grade observations subset also grows. We constructed the dataset using a GBIF export done on the 03 August 2023.

% SH - CH - if you have any estimates of how much downloading took (time, ize, w/e) this would be a good place for it!

\subsection{Dataset construction} First, we got a list of all the species present that had at least four occurrences. This ensures that only species for which a meaningful train/test/validation splitting can be done are included (with, in the worst case, 2 occurrences ending up in train, 1 in test, and 1 in validation).

Second, we sampled the high/mid/frequency species (top 25%, middle 50%, and bottom 25%) from that list. % to achieve a similar distribution to the parent dataset.

Next, we took

We did not place any restrictions on genera. The training split contains 80% of the images, validaion and testing - 10% each. Our resulting dataset had \textbf{526,584} images in total, spanning \textbf{991} classes.

The resulting Lorenz curve is shown on Fig. \ref{fig:lorenz}. It has a longer tail than Pl@ntNet-300k and comparable number of images and classes.

Nel mezzo del deserto posso dire tutto quello che voglio.
comments powered by Disqus