Better dataset description draft thing
\section{iNaturalist-300k} \subsection{Dataset collection} We compiled iNaturalist-300k from images of the iNaturalist Research-grade observation collection, with the goal of creating an additional dataset matching our criteria, complementary to Pl@ntNet-300k. Since new observations get added to the iNaturalist dataset every day, the research-grade observations subset also grows. We constructed the dataset using a GBIF export done on the 03 August 2023.
% SH - CH - if you have any estimates of how much downloading took (time, ize, w/e) this would be a good place for it!
\subsection{Dataset construction} First, we got a list of all the species present that had at least four occurrences. This ensures that only species for which a meaningful train/test/validation splitting can be done are included (with, in the worst case, 2 occurrences ending up in train, 1 in test, and 1 in validation).
Second, we sampled the high/mid/frequency species (top 25%, middle 50%, and bottom 25%) from that list. % to achieve a similar distribution to the parent dataset.
Next, we took
We did not place any restrictions on genera. The training split contains 80% of the images, validaion and testing - 10% each. Our resulting dataset had \textbf{526,584} images in total, spanning \textbf{991} classes.
The resulting Lorenz curve is shown on Fig. \ref{fig:lorenz}. It has a longer tail than Pl@ntNet-300k and comparable number of images and classes.