{
  "id": 328637,
  "title": "2nd place solution description",
  "url": "/competitions/geolifeclef-2022-lifeclef-2022-fgvc9/discussion/328637",
  "author_name": "Matsushita",
  "post_date": "2022-06-02T08:10:19.985000",
  "votes": 4,
  "comment_count": 0,
  "views": 0,
  "content": "<p>Well, I guess it's my time to explain the solution that clocked in second on the private leaderboard.<br>\nSimilar to the post by the winners (congratulations!) I will not go into every single detail; this will be made available through the technical report.</p>\n<p><strong>General overview</strong><br>\nMy solution is based around deep convolutional neural networks (CNNs) that process the satellite remote sensing (RS) products only. In detail, the base model consists of two CNN feature extractors (standard image classiﬁcation architectures that had their ﬁnal, fully-connected classiﬁcation layer dropped). They run in parallel and ingest different parts of the RS imagery: the ﬁrst receives the RGB portion of the dataset, the second a stack of altitude, near-infrared (NIR) and NDVI (normalised difference vegetation index: (NIR-red)/(NIR+red)). I previously used the land cover map as input instead of the NDVI, but providing continuous models like CNNs with a discrete map of ordinal class indices is non-straightforward and not imminently sensible, hence the NDVI.</p>\n<p>The two feature extractors process their assigned three-channel imagery in parallel, but do not share parameters. They each output a latent feature vector of the same size, which gets concatenated, subjected to heavy dropout (probability of 0.45 worked best), and converted into per-class activations by a single fully-connected layer that maps to the 17,034 species. A standard softmax-cross-entropy loss is used to train the model (more on that below).</p>\n<p><strong>What helped improve accuracy</strong><br>\nI ran extensive experiments and tried many different ideas; many without success (as is usually the case). However, the following four provided sufﬁcient individual boosts in accuracy:</p>\n<ol>\n<li><em>Correct model architectures:</em> I commenced with a ResNet-50 for each of the two feature extractor branches. This worked reasonably well, but I got an improvement of almost 2% simply by switching to Inception-v4. I also tried DenseNet-201, which performed similarly to ResNet-50. More complex and recent architectures were on the list as well, including ConvNext and a Vision Transformer (ViT B/16). Those two took forever to train, despite decently powerful hardware (see below), and just resulted in massive overﬁtting (35% top-30 on train, 5% on val).</li>\n<li><em>Pretraining:</em> deep learning models oftentimes require some form of pre-training to be able to adapt well, due to the enormous search space with millions of free parameters. Different means of pre-training were on my list, but what eventually worked best was to simply use the weights from ImageNet pre-training. See below for other ideas.</li>\n<li><em>Spatial block-label swap:</em> while the two points above are important to make the model work in the ﬁrst place, this here is the most relevant for the performance of my model by a long shot. This is a training concept that boils down to the presence-only label problem we are facing: a non-observation of a species at a particular location does not imply that the species is not there. This circumstance, together with the fact that each data point is assigned just one single species, technically violates the assumption made with by-the-book heuristics like cross-entropy loss: this one assumes that each data point has exactly one correct label and that all other classes are to be zero (it minimises the entropy in the predicted probability space). While the GeoLifeCLEF dataset is set up this way, this causes the prediction problem to be ill-posed. It is thus prudent to ﬁnd a training strategy that takes this into account. A popular idea is to use a temperature softmax, but this only softens the blows for the model to a certain extent. My solution, the spatial block-label swap, attempts to relax the strict one-class requirement as follows: it ﬁrst creates a spatial grid of square cells (0.01x0.01° lat/lon) and assigns each training point to its encompassing cell. During training, it then performs a look-up for each data point and randomly replaces its label with one of the neighbours in the same grid cell with a probability of 10%. This is a simple but sort of intuitive means of regularisation, and it helped the model gain another 2% of accuracy. I tried different replacement chance probabilities, as well as more sophisticated strategies (e.g., swapping according to encounter probabilities of a species within the grid cell), but the most simple method worked best.</li>\n<li><em>Ensembling:</em> as with many contests (I believe), an ensemble of different models worked best. For the submission I ensembled ten different models trained; some ResNet-50, some Inception-v4, one DenseNet-201. Each was trained with different slight alterations (e.g., different learning rates, different random seeds), but the general concept always stayed the same. I ensembled the output probabilities, but in a pseudo-conﬁdence-based manner: I performed test-time augmentation (also important) and recorded the variance of conﬁdence across the augmentations for each test sample and each model. I then normalised and inverted the variances per model, so that they were bound to [0,1], and applied a softmax across all ten model runs. This provided me with a crude notion of model conﬁdence. Multiplying the softmax scores with the mean model conﬁdences across augmentations and summing them together gave the ﬁnal prediction score, from which I drew the top-30 species. Ensembling gave yet another ~2% boost over the single models.</li>\n</ol>\n<p><strong>What I tried without success</strong><br>\nA setup and challenge like GeoLifeCLEF absolutely invites to try out many different ideas. Providing an exhaustive list of tricks I attempted is beyond the scope of this thread, but here's some of the more relevant ones that simply did not want to work:</p>\n<ul>\n<li><em>Other covariates:</em> I tried plugging in the environmental rasters (one value extracted per data point), as well as the GPS coordinates, into the model. None of this worked well. I attribute reasons to the difﬁculty of normalising features against the RS-emerging ones and the difﬁculty in extracting features in the ﬁrst place (going from environmental covariates to latent features requires usage of an MLP, which is always a bit messy). The winners have done the better move here by using a non-DL model for the covariates, such as a Random Forest. It shows that DL isn't always the answer (which sounds ironic, given that this is only what I used). Otherwise, encoding geospatial coordinates is not straightforward either. Since we only have the contiguous U.S. and France as study areas, we don't have the periodicity problem, but it still is nontrivial to use those. An MLP trained on the coordinates only, sometimes raw, sometimes sine-cosine-encoded, and once with a more advanced encoding (Theory-guided spatial encoding), just led to severe underﬁtting. I believe more research on this multimodal covariate integration would be an interesting avenue to explore further.</li>\n<li><em>Auxiliary prediction tasks:</em> at some point I tried predicting not just the species class, but the other layers of the taxonomy tree (genus, family, kingdom) with individual fully-connected layers on top. That worked ok, but did not improve species classiﬁcation.</li>\n<li><em>Predicting histogram densities:</em> with the spatial block-label swap strategy explained above, I tried predicting the occurrence histogram per grid cell per species. This provided too faint a learning signal, though.</li>\n<li><em>Advanced pre-training:</em> besides standard ImageNet weights, I tried self-supervised pre-training (in particular MoCo-v2, which is what the winner of last year's challenge used), and an own form of Model-Agnostic Meta-Learning (MAML); resp. the more lightweight alternative Almost No Inner Loop (ANIL). I also tried training models from scratch. In the end, ImageNet is what performed best.</li>\n<li><em>Addressing the long-tail:</em> the species classes are severely long-tailed (i.e., thousands of species have less than ten images and a few make up the vast majority of the dataset). That strongly confuses machine learning models by default, with the result that the rare species never get predicted. I tried many ideas to cope with this, from loss weights over special losses (e.g., Balanced Softmax) to ANIL pre-training (see above). In the end, doing nothing about it worked best—for a very simply reason: the test set could be assumed to be as unbalanced in species classes as the training and validation sets. Hence, giving more weight to the rare species is actually the opposite of what one wants to maximise performance. I could probably even have dropped many of the rare species with possibly performance gains, as the model had a less complicated task to solve (I didn't do that, though).</li>\n</ul>\n<p><strong>Setup</strong><br>\nI distributed training to three machines: two workstations (16-core CPU, 128 GB RAM, NVIDIA GeForce RTX 3090 each; Ubuntu 20.04 LTS) and an HPC (NVIDIA Tesla V100; RHEL 7). I used Python 3.8.10 and implemented my solution in PyTorch 1.9.0. Models came from either Torchvision (ResNet-50) or the PyTorch Image Models (TIMM) library (Inception-v4, DenseNet-201, Inception-v4, ViT B/16).</p>\n<p><strong>Lessons learnt</strong><br>\nThis would be a big one to cover. It was difﬁcult to get to the grounds of the performance of models, due to the sheer number of species classes. However, gradually getting a feel for how the models perform over different experiments was possible to an extent and certainly helped in settling in on standard parameter sets and avoiding pitfalls. There is a slew of ideas to be tested still, such as advanced data augmentation (I only used random flips, addition of Gaussian noise and normalisation), grid-searching hyperparameters, and other architectures I didn't try (EfﬁcientNet, for example). Otherwise, adequately using all available covariates on the one hand, and further exploring tweaks about the label situation on the other, are probably worth studying.</p>\n<p>Apologies for the long post; I hope it was at least somewhat insightful. Full details on the solution will be given in the technical report.</p>\n<p>Many thanks for everyone involved: the dataset creators, contest organisers, data collectors, and of course the competitors! It was quite exciting to see developments on the leaderboard, especially towards the end. Huge congratulations to the winners; your solution reads fantastically and that win is more than well-deserved!</p>\n<p>Thank you.</p>",
  "messages": [
    {
      "id": 1808874,
      "postDate": "2022-06-02T08:10:19.987Z",
      "content": "<p>Well, I guess it's my time to explain the solution that clocked in second on the private leaderboard.<br>\nSimilar to the post by the winners (congratulations!) I will not go into every single detail; this will be made available through the technical report.</p>\n<p><strong>General overview</strong><br>\nMy solution is based around deep convolutional neural networks (CNNs) that process the satellite remote sensing (RS) products only. In detail, the base model consists of two CNN feature extractors (standard image classiﬁcation architectures that had their ﬁnal, fully-connected classiﬁcation layer dropped). They run in parallel and ingest different parts of the RS imagery: the ﬁrst receives the RGB portion of the dataset, the second a stack of altitude, near-infrared (NIR) and NDVI (normalised difference vegetation index: (NIR-red)/(NIR+red)). I previously used the land cover map as input instead of the NDVI, but providing continuous models like CNNs with a discrete map of ordinal class indices is non-straightforward and not imminently sensible, hence the NDVI.</p>\n<p>The two feature extractors process their assigned three-channel imagery in parallel, but do not share parameters. They each output a latent feature vector of the same size, which gets concatenated, subjected to heavy dropout (probability of 0.45 worked best), and converted into per-class activations by a single fully-connected layer that maps to the 17,034 species. A standard softmax-cross-entropy loss is used to train the model (more on that below).</p>\n<p><strong>What helped improve accuracy</strong><br>\nI ran extensive experiments and tried many different ideas; many without success (as is usually the case). However, the following four provided sufﬁcient individual boosts in accuracy:</p>\n<ol>\n<li><em>Correct model architectures:</em> I commenced with a ResNet-50 for each of the two feature extractor branches. This worked reasonably well, but I got an improvement of almost 2% simply by switching to Inception-v4. I also tried DenseNet-201, which performed similarly to ResNet-50. More complex and recent architectures were on the list as well, including ConvNext and a Vision Transformer (ViT B/16). Those two took forever to train, despite decently powerful hardware (see below), and just resulted in massive overﬁtting (35% top-30 on train, 5% on val).</li>\n<li><em>Pretraining:</em> deep learning models oftentimes require some form of pre-training to be able to adapt well, due to the enormous search space with millions of free parameters. Different means of pre-training were on my list, but what eventually worked best was to simply use the weights from ImageNet pre-training. See below for other ideas.</li>\n<li><em>Spatial block-label swap:</em> while the two points above are important to make the model work in the ﬁrst place, this here is the most relevant for the performance of my model by a long shot. This is a training concept that boils down to the presence-only label problem we are facing: a non-observation of a species at a particular location does not imply that the species is not there. This circumstance, together with the fact that each data point is assigned just one single species, technically violates the assumption made with by-the-book heuristics like cross-entropy loss: this one assumes that each data point has exactly one correct label and that all other classes are to be zero (it minimises the entropy in the predicted probability space). While the GeoLifeCLEF dataset is set up this way, this causes the prediction problem to be ill-posed. It is thus prudent to ﬁnd a training strategy that takes this into account. A popular idea is to use a temperature softmax, but this only softens the blows for the model to a certain extent. My solution, the spatial block-label swap, attempts to relax the strict one-class requirement as follows: it ﬁrst creates a spatial grid of square cells (0.01x0.01° lat/lon) and assigns each training point to its encompassing cell. During training, it then performs a look-up for each data point and randomly replaces its label with one of the neighbours in the same grid cell with a probability of 10%. This is a simple but sort of intuitive means of regularisation, and it helped the model gain another 2% of accuracy. I tried different replacement chance probabilities, as well as more sophisticated strategies (e.g., swapping according to encounter probabilities of a species within the grid cell), but the most simple method worked best.</li>\n<li><em>Ensembling:</em> as with many contests (I believe), an ensemble of different models worked best. For the submission I ensembled ten different models trained; some ResNet-50, some Inception-v4, one DenseNet-201. Each was trained with different slight alterations (e.g., different learning rates, different random seeds), but the general concept always stayed the same. I ensembled the output probabilities, but in a pseudo-conﬁdence-based manner: I performed test-time augmentation (also important) and recorded the variance of conﬁdence across the augmentations for each test sample and each model. I then normalised and inverted the variances per model, so that they were bound to [0,1], and applied a softmax across all ten model runs. This provided me with a crude notion of model conﬁdence. Multiplying the softmax scores with the mean model conﬁdences across augmentations and summing them together gave the ﬁnal prediction score, from which I drew the top-30 species. Ensembling gave yet another ~2% boost over the single models.</li>\n</ol>\n<p><strong>What I tried without success</strong><br>\nA setup and challenge like GeoLifeCLEF absolutely invites to try out many different ideas. Providing an exhaustive list of tricks I attempted is beyond the scope of this thread, but here's some of the more relevant ones that simply did not want to work:</p>\n<ul>\n<li><em>Other covariates:</em> I tried plugging in the environmental rasters (one value extracted per data point), as well as the GPS coordinates, into the model. None of this worked well. I attribute reasons to the difﬁculty of normalising features against the RS-emerging ones and the difﬁculty in extracting features in the ﬁrst place (going from environmental covariates to latent features requires usage of an MLP, which is always a bit messy). The winners have done the better move here by using a non-DL model for the covariates, such as a Random Forest. It shows that DL isn't always the answer (which sounds ironic, given that this is only what I used). Otherwise, encoding geospatial coordinates is not straightforward either. Since we only have the contiguous U.S. and France as study areas, we don't have the periodicity problem, but it still is nontrivial to use those. An MLP trained on the coordinates only, sometimes raw, sometimes sine-cosine-encoded, and once with a more advanced encoding (Theory-guided spatial encoding), just led to severe underﬁtting. I believe more research on this multimodal covariate integration would be an interesting avenue to explore further.</li>\n<li><em>Auxiliary prediction tasks:</em> at some point I tried predicting not just the species class, but the other layers of the taxonomy tree (genus, family, kingdom) with individual fully-connected layers on top. That worked ok, but did not improve species classiﬁcation.</li>\n<li><em>Predicting histogram densities:</em> with the spatial block-label swap strategy explained above, I tried predicting the occurrence histogram per grid cell per species. This provided too faint a learning signal, though.</li>\n<li><em>Advanced pre-training:</em> besides standard ImageNet weights, I tried self-supervised pre-training (in particular MoCo-v2, which is what the winner of last year's challenge used), and an own form of Model-Agnostic Meta-Learning (MAML); resp. the more lightweight alternative Almost No Inner Loop (ANIL). I also tried training models from scratch. In the end, ImageNet is what performed best.</li>\n<li><em>Addressing the long-tail:</em> the species classes are severely long-tailed (i.e., thousands of species have less than ten images and a few make up the vast majority of the dataset). That strongly confuses machine learning models by default, with the result that the rare species never get predicted. I tried many ideas to cope with this, from loss weights over special losses (e.g., Balanced Softmax) to ANIL pre-training (see above). In the end, doing nothing about it worked best—for a very simply reason: the test set could be assumed to be as unbalanced in species classes as the training and validation sets. Hence, giving more weight to the rare species is actually the opposite of what one wants to maximise performance. I could probably even have dropped many of the rare species with possibly performance gains, as the model had a less complicated task to solve (I didn't do that, though).</li>\n</ul>\n<p><strong>Setup</strong><br>\nI distributed training to three machines: two workstations (16-core CPU, 128 GB RAM, NVIDIA GeForce RTX 3090 each; Ubuntu 20.04 LTS) and an HPC (NVIDIA Tesla V100; RHEL 7). I used Python 3.8.10 and implemented my solution in PyTorch 1.9.0. Models came from either Torchvision (ResNet-50) or the PyTorch Image Models (TIMM) library (Inception-v4, DenseNet-201, Inception-v4, ViT B/16).</p>\n<p><strong>Lessons learnt</strong><br>\nThis would be a big one to cover. It was difﬁcult to get to the grounds of the performance of models, due to the sheer number of species classes. However, gradually getting a feel for how the models perform over different experiments was possible to an extent and certainly helped in settling in on standard parameter sets and avoiding pitfalls. There is a slew of ideas to be tested still, such as advanced data augmentation (I only used random flips, addition of Gaussian noise and normalisation), grid-searching hyperparameters, and other architectures I didn't try (EfﬁcientNet, for example). Otherwise, adequately using all available covariates on the one hand, and further exploring tweaks about the label situation on the other, are probably worth studying.</p>\n<p>Apologies for the long post; I hope it was at least somewhat insightful. Full details on the solution will be given in the technical report.</p>\n<p>Many thanks for everyone involved: the dataset creators, contest organisers, data collectors, and of course the competitors! It was quite exciting to see developments on the leaderboard, especially towards the end. Huge congratulations to the winners; your solution reads fantastically and that win is more than well-deserved!</p>\n<p>Thank you.</p>",
      "rawMarkdown": "Well, I guess it's my time to explain the solution that clocked in second on the private leaderboard.\nSimilar to the post by the winners (congratulations!) I will not go into every single detail; this will be made available through the technical report.\n\n**General overview**\nMy solution is based around deep convolutional neural networks (CNNs) that process the satellite remote sensing (RS) products only. In detail, the base model consists of two CNN feature extractors (standard image classiﬁcation architectures that had their ﬁnal, fully-connected classiﬁcation layer dropped). They run in parallel and ingest different parts of the RS imagery: the ﬁrst receives the RGB portion of the dataset, the second a stack of altitude, near-infrared (NIR) and NDVI (normalised difference vegetation index: (NIR-red)/(NIR+red)). I previously used the land cover map as input instead of the NDVI, but providing continuous models like CNNs with a discrete map of ordinal class indices is non-straightforward and not imminently sensible, hence the NDVI.\n\nThe two feature extractors process their assigned three-channel imagery in parallel, but do not share parameters. They each output a latent feature vector of the same size, which gets concatenated, subjected to heavy dropout (probability of 0.45 worked best), and converted into per-class activations by a single fully-connected layer that maps to the 17,034 species. A standard softmax-cross-entropy loss is used to train the model (more on that below).\n\n\n**What helped improve accuracy**\nI ran extensive experiments and tried many different ideas; many without success (as is usually the case). However, the following four provided sufﬁcient individual boosts in accuracy:\n\n1. *Correct model architectures:* I commenced with a ResNet-50 for each of the two feature extractor branches. This worked reasonably well, but I got an improvement of almost 2% simply by switching to Inception-v4. I also tried DenseNet-201, which performed similarly to ResNet-50. More complex and recent architectures were on the list as well, including ConvNext and a Vision Transformer (ViT B/16). Those two took forever to train, despite decently powerful hardware (see below), and just resulted in massive overﬁtting (35% top-30 on train, 5% on val).\n2. *Pretraining:* deep learning models oftentimes require some form of pre-training to be able to adapt well, due to the enormous search space with millions of free parameters. Different means of pre-training were on my list, but what eventually worked best was to simply use the weights from ImageNet pre-training. See below for other ideas.\n3. *Spatial block-label swap:* while the two points above are important to make the model work in the ﬁrst place, this here is the most relevant for the performance of my model by a long shot. This is a training concept that boils down to the presence-only label problem we are facing: a non-observation of a species at a particular location does not imply that the species is not there. This circumstance, together with the fact that each data point is assigned just one single species, technically violates the assumption made with by-the-book heuristics like cross-entropy loss: this one assumes that each data point has exactly one correct label and that all other classes are to be zero (it minimises the entropy in the predicted probability space). While the GeoLifeCLEF dataset is set up this way, this causes the prediction problem to be ill-posed. It is thus prudent to ﬁnd a training strategy that takes this into account. A popular idea is to use a temperature softmax, but this only softens the blows for the model to a certain extent. My solution, the spatial block-label swap, attempts to relax the strict one-class requirement as follows: it ﬁrst creates a spatial grid of square cells (0.01x0.01° lat/lon) and assigns each training point to its encompassing cell. During training, it then performs a look-up for each data point and randomly replaces its label with one of the neighbours in the same grid cell with a probability of 10%. This is a simple but sort of intuitive means of regularisation, and it helped the model gain another 2% of accuracy. I tried different replacement chance probabilities, as well as more sophisticated strategies (e.g., swapping according to encounter probabilities of a species within the grid cell), but the most simple method worked best.\n3. *Ensembling:* as with many contests (I believe), an ensemble of different models worked best. For the submission I ensembled ten different models trained; some ResNet-50, some Inception-v4, one DenseNet-201. Each was trained with different slight alterations (e.g., different learning rates, different random seeds), but the general concept always stayed the same. I ensembled the output probabilities, but in a pseudo-conﬁdence-based manner: I performed test-time augmentation (also important) and recorded the variance of conﬁdence across the augmentations for each test sample and each model. I then normalised and inverted the variances per model, so that they were bound to [0,1], and applied a softmax across all ten model runs. This provided me with a crude notion of model conﬁdence. Multiplying the softmax scores with the mean model conﬁdences across augmentations and summing them together gave the ﬁnal prediction score, from which I drew the top-30 species. Ensembling gave yet another ~2% boost over the single models.\n\n\n**What I tried without success**\nA setup and challenge like GeoLifeCLEF absolutely invites to try out many different ideas. Providing an exhaustive list of tricks I attempted is beyond the scope of this thread, but here's some of the more relevant ones that simply did not want to work:\n- *Other covariates:* I tried plugging in the environmental rasters (one value extracted per data point), as well as the GPS coordinates, into the model. None of this worked well. I attribute reasons to the difﬁculty of normalising features against the RS-emerging ones and the difﬁculty in extracting features in the ﬁrst place (going from environmental covariates to latent features requires usage of an MLP, which is always a bit messy). The winners have done the better move here by using a non-DL model for the covariates, such as a Random Forest. It shows that DL isn't always the answer (which sounds ironic, given that this is only what I used). Otherwise, encoding geospatial coordinates is not straightforward either. Since we only have the contiguous U.S. and France as study areas, we don't have the periodicity problem, but it still is nontrivial to use those. An MLP trained on the coordinates only, sometimes raw, sometimes sine-cosine-encoded, and once with a more advanced encoding (Theory-guided spatial encoding), just led to severe underﬁtting. I believe more research on this multimodal covariate integration would be an interesting avenue to explore further.\n- *Auxiliary prediction tasks:* at some point I tried predicting not just the species class, but the other layers of the taxonomy tree (genus, family, kingdom) with individual fully-connected layers on top. That worked ok, but did not improve species classiﬁcation.\n- *Predicting histogram densities:* with the spatial block-label swap strategy explained above, I tried predicting the occurrence histogram per grid cell per species. This provided too faint a learning signal, though.\n- *Advanced pre-training:* besides standard ImageNet weights, I tried self-supervised pre-training (in particular MoCo-v2, which is what the winner of last year's challenge used), and an own form of Model-Agnostic Meta-Learning (MAML); resp. the more lightweight alternative Almost No Inner Loop (ANIL). I also tried training models from scratch. In the end, ImageNet is what performed best.\n- *Addressing the long-tail:* the species classes are severely long-tailed (i.e., thousands of species have less than ten images and a few make up the vast majority of the dataset). That strongly confuses machine learning models by default, with the result that the rare species never get predicted. I tried many ideas to cope with this, from loss weights over special losses (e.g., Balanced Softmax) to ANIL pre-training (see above). In the end, doing nothing about it worked best—for a very simply reason: the test set could be assumed to be as unbalanced in species classes as the training and validation sets. Hence, giving more weight to the rare species is actually the opposite of what one wants to maximise performance. I could probably even have dropped many of the rare species with possibly performance gains, as the model had a less complicated task to solve (I didn't do that, though).\n\n\n**Setup**\nI distributed training to three machines: two workstations (16-core CPU, 128 GB RAM, NVIDIA GeForce RTX 3090 each; Ubuntu 20.04 LTS) and an HPC (NVIDIA Tesla V100; RHEL 7). I used Python 3.8.10 and implemented my solution in PyTorch 1.9.0. Models came from either Torchvision (ResNet-50) or the PyTorch Image Models (TIMM) library (Inception-v4, DenseNet-201, Inception-v4, ViT B/16).\n\n\n**Lessons learnt**\nThis would be a big one to cover. It was difﬁcult to get to the grounds of the performance of models, due to the sheer number of species classes. However, gradually getting a feel for how the models perform over different experiments was possible to an extent and certainly helped in settling in on standard parameter sets and avoiding pitfalls. There is a slew of ideas to be tested still, such as advanced data augmentation (I only used random flips, addition of Gaussian noise and normalisation), grid-searching hyperparameters, and other architectures I didn't try (EfﬁcientNet, for example). Otherwise, adequately using all available covariates on the one hand, and further exploring tweaks about the label situation on the other, are probably worth studying.\n\n\n\nApologies for the long post; I hope it was at least somewhat insightful. Full details on the solution will be given in the technical report.\n\nMany thanks for everyone involved: the dataset creators, contest organisers, data collectors, and of course the competitors! It was quite exciting to see developments on the leaderboard, especially towards the end. Huge congratulations to the winners; your solution reads fantastically and that win is more than well-deserved!\n\nThank you.\n\n",
      "votes": 4
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "1808874": "Well, I guess it's my time to explain the solution that clocked in second on the private leaderboard.\nSimilar to the post by the winners (congratulations!) I will not go into every single detail; this will be made available through the technical report.\n\n**General overview**\nMy solution is based around deep convolutional neural networks (CNNs) that process the satellite remote sensing (RS) products only. In detail, the base model consists of two CNN feature extractors (standard image classiﬁcation architectures that had their ﬁnal, fully-connected classiﬁcation layer dropped). They run in parallel and ingest different parts of the RS imagery: the ﬁrst receives the RGB portion of the dataset, the second a stack of altitude, near-infrared (NIR) and NDVI (normalised difference vegetation index: (NIR-red)/(NIR+red)). I previously used the land cover map as input instead of the NDVI, but providing continuous models like CNNs with a discrete map of ordinal class indices is non-straightforward and not imminently sensible, hence the NDVI.\n\nThe two feature extractors process their assigned three-channel imagery in parallel, but do not share parameters. They each output a latent feature vector of the same size, which gets concatenated, subjected to heavy dropout (probability of 0.45 worked best), and converted into per-class activations by a single fully-connected layer that maps to the 17,034 species. A standard softmax-cross-entropy loss is used to train the model (more on that below).\n\n\n**What helped improve accuracy**\nI ran extensive experiments and tried many different ideas; many without success (as is usually the case). However, the following four provided sufﬁcient individual boosts in accuracy:\n\n1. *Correct model architectures:* I commenced with a ResNet-50 for each of the two feature extractor branches. This worked reasonably well, but I got an improvement of almost 2% simply by switching to Inception-v4. I also tried DenseNet-201, which performed similarly to ResNet-50. More complex and recent architectures were on the list as well, including ConvNext and a Vision Transformer (ViT B/16). Those two took forever to train, despite decently powerful hardware (see below), and just resulted in massive overﬁtting (35% top-30 on train, 5% on val).\n2. *Pretraining:* deep learning models oftentimes require some form of pre-training to be able to adapt well, due to the enormous search space with millions of free parameters. Different means of pre-training were on my list, but what eventually worked best was to simply use the weights from ImageNet pre-training. See below for other ideas.\n3. *Spatial block-label swap:* while the two points above are important to make the model work in the ﬁrst place, this here is the most relevant for the performance of my model by a long shot. This is a training concept that boils down to the presence-only label problem we are facing: a non-observation of a species at a particular location does not imply that the species is not there. This circumstance, together with the fact that each data point is assigned just one single species, technically violates the assumption made with by-the-book heuristics like cross-entropy loss: this one assumes that each data point has exactly one correct label and that all other classes are to be zero (it minimises the entropy in the predicted probability space). While the GeoLifeCLEF dataset is set up this way, this causes the prediction problem to be ill-posed. It is thus prudent to ﬁnd a training strategy that takes this into account. A popular idea is to use a temperature softmax, but this only softens the blows for the model to a certain extent. My solution, the spatial block-label swap, attempts to relax the strict one-class requirement as follows: it ﬁrst creates a spatial grid of square cells (0.01x0.01° lat/lon) and assigns each training point to its encompassing cell. During training, it then performs a look-up for each data point and randomly replaces its label with one of the neighbours in the same grid cell with a probability of 10%. This is a simple but sort of intuitive means of regularisation, and it helped the model gain another 2% of accuracy. I tried different replacement chance probabilities, as well as more sophisticated strategies (e.g., swapping according to encounter probabilities of a species within the grid cell), but the most simple method worked best.\n3. *Ensembling:* as with many contests (I believe), an ensemble of different models worked best. For the submission I ensembled ten different models trained; some ResNet-50, some Inception-v4, one DenseNet-201. Each was trained with different slight alterations (e.g., different learning rates, different random seeds), but the general concept always stayed the same. I ensembled the output probabilities, but in a pseudo-conﬁdence-based manner: I performed test-time augmentation (also important) and recorded the variance of conﬁdence across the augmentations for each test sample and each model. I then normalised and inverted the variances per model, so that they were bound to [0,1], and applied a softmax across all ten model runs. This provided me with a crude notion of model conﬁdence. Multiplying the softmax scores with the mean model conﬁdences across augmentations and summing them together gave the ﬁnal prediction score, from which I drew the top-30 species. Ensembling gave yet another ~2% boost over the single models.\n\n\n**What I tried without success**\nA setup and challenge like GeoLifeCLEF absolutely invites to try out many different ideas. Providing an exhaustive list of tricks I attempted is beyond the scope of this thread, but here's some of the more relevant ones that simply did not want to work:\n- *Other covariates:* I tried plugging in the environmental rasters (one value extracted per data point), as well as the GPS coordinates, into the model. None of this worked well. I attribute reasons to the difﬁculty of normalising features against the RS-emerging ones and the difﬁculty in extracting features in the ﬁrst place (going from environmental covariates to latent features requires usage of an MLP, which is always a bit messy). The winners have done the better move here by using a non-DL model for the covariates, such as a Random Forest. It shows that DL isn't always the answer (which sounds ironic, given that this is only what I used). Otherwise, encoding geospatial coordinates is not straightforward either. Since we only have the contiguous U.S. and France as study areas, we don't have the periodicity problem, but it still is nontrivial to use those. An MLP trained on the coordinates only, sometimes raw, sometimes sine-cosine-encoded, and once with a more advanced encoding (Theory-guided spatial encoding), just led to severe underﬁtting. I believe more research on this multimodal covariate integration would be an interesting avenue to explore further.\n- *Auxiliary prediction tasks:* at some point I tried predicting not just the species class, but the other layers of the taxonomy tree (genus, family, kingdom) with individual fully-connected layers on top. That worked ok, but did not improve species classiﬁcation.\n- *Predicting histogram densities:* with the spatial block-label swap strategy explained above, I tried predicting the occurrence histogram per grid cell per species. This provided too faint a learning signal, though.\n- *Advanced pre-training:* besides standard ImageNet weights, I tried self-supervised pre-training (in particular MoCo-v2, which is what the winner of last year's challenge used), and an own form of Model-Agnostic Meta-Learning (MAML); resp. the more lightweight alternative Almost No Inner Loop (ANIL). I also tried training models from scratch. In the end, ImageNet is what performed best.\n- *Addressing the long-tail:* the species classes are severely long-tailed (i.e., thousands of species have less than ten images and a few make up the vast majority of the dataset). That strongly confuses machine learning models by default, with the result that the rare species never get predicted. I tried many ideas to cope with this, from loss weights over special losses (e.g., Balanced Softmax) to ANIL pre-training (see above). In the end, doing nothing about it worked best—for a very simply reason: the test set could be assumed to be as unbalanced in species classes as the training and validation sets. Hence, giving more weight to the rare species is actually the opposite of what one wants to maximise performance. I could probably even have dropped many of the rare species with possibly performance gains, as the model had a less complicated task to solve (I didn't do that, though).\n\n\n**Setup**\nI distributed training to three machines: two workstations (16-core CPU, 128 GB RAM, NVIDIA GeForce RTX 3090 each; Ubuntu 20.04 LTS) and an HPC (NVIDIA Tesla V100; RHEL 7). I used Python 3.8.10 and implemented my solution in PyTorch 1.9.0. Models came from either Torchvision (ResNet-50) or the PyTorch Image Models (TIMM) library (Inception-v4, DenseNet-201, Inception-v4, ViT B/16).\n\n\n**Lessons learnt**\nThis would be a big one to cover. It was difﬁcult to get to the grounds of the performance of models, due to the sheer number of species classes. However, gradually getting a feel for how the models perform over different experiments was possible to an extent and certainly helped in settling in on standard parameter sets and avoiding pitfalls. There is a slew of ideas to be tested still, such as advanced data augmentation (I only used random flips, addition of Gaussian noise and normalisation), grid-searching hyperparameters, and other architectures I didn't try (EfﬁcientNet, for example). Otherwise, adequately using all available covariates on the one hand, and further exploring tweaks about the label situation on the other, are probably worth studying.\n\n\n\nApologies for the long post; I hope it was at least somewhat insightful. Full details on the solution will be given in the technical report.\n\nMany thanks for everyone involved: the dataset creators, contest organisers, data collectors, and of course the competitors! It was quite exciting to see developments on the leaderboard, especially towards the end. Huge congratulations to the winners; your solution reads fantastically and that win is more than well-deserved!\n\nThank you.\n\n"
  }
}