{
  "id": 327573,
  "title": "571st place solution",
  "url": "/competitions/birdclef-2022/writeups/thacrobatheskis-571st-place-solution",
  "author_name": "",
  "post_date": "2022-05-28T12:38:04.030Z",
  "votes": 30,
  "comment_count": 4,
  "views": 0,
  "content": "<p>It's probably a bit silly for me to make this note given how low my ranking was. But this was my first real kaggle competition, and I personally learned a lot, so I thought I'd write down some reflections—if for nothing else than for my present and future self. :)</p>\n<p>First: thanks to everyone who competed and shared their thoughts and questions, and especially to the hosts of this competition for putting so much effort in fostering community focused on using AI/ML for scientific research. It was a lot of fun to participate, and I'm looking forward to doing more kaggle competitions in the future.</p>\n<h2>Dataset</h2>\n<p>I first preprocessed the dataset into log magnitude spectrograms, with linear frequency bins from 1-10 kHz . The final resolution was a 128x256 spectrogram per 5 seconds of audio. For augmentation during training, I randomly masked out portions of the frequency and time axis, inspired by <a href=\"https://arxiv.org/abs/1904.08779\" target=\"_blank\">SpecAugment</a>.</p>\n<h2>Model</h2>\n<p>I saw that <a href=\"https://www.kaggle.com/kaerururu\" target=\"_blank\">@kaerururu</a> shared a <a href=\"https://www.kaggle.com/competitions/birdclef-2022/discussion/318081\" target=\"_blank\">well performing solution</a>, and that there was a ton of work that might be reusable from previous competitions. But instead of reusing their approaches, I decided to first try doing my own thing, originally planning to go back and try their ideas if I had time (but of course I ran out of time!).</p>\n<p>My first baseline model (which ended up being the best) was a simple 1D convolution over the time axis of the spectrograms as a feature extractor, followed by max pooling over time, then flattening the channels and frequency axis into a linear layer that projects onto the # of classes. </p>\n<pre><code>class DumbBirdClassifier(nn.Module):\n    def __init__(self):\n        super(DumbBirdClassifier, self).__init__()\n        self.conv = nn.Sequential(\n            nn.Conv1d(kernel_size=4, stride=2, padding=1, in_channels=128, out_channels=128),\n            nn.ReLU(),\n            nn.Conv1d(kernel_size=4, stride=2, padding=1, in_channels=128, out_channels=128),\n            nn.ReLU(),\n            nn.Conv1d(kernel_size=4, stride=2, padding=1, in_channels=128, out_channels=128),\n            nn.ReLU(),\n            nn.Conv1d(kernel_size=4, stride=2, padding=1, in_channels=128, out_channels=128),\n            nn.ReLU(),\n        )\n        self.pool = lambda x: x.max(-1).values\n        self.linear = nn.Sequential(\n            nn.Linear(in_features=128, out_features=128),\n            nn.ReLU(),\n            nn.Linear(in_features=128, out_features=NUM_CLASSES),\n            nn.Sigmoid(),\n        )\n\n    def forward(self, x):\n        return self.linear(self.pool(self.conv(x)))\n</code></pre>\n<p>I reasoned that collapsing the time axis with max pooling would make the model robust to different lengths of audio. The model must make evaluations on 5 sec-long clips, but it could be useful to be able to give it longer clips of audio during training, since only a 5 sec clip from a training recording would likely not include all or any of the labeled birds, and therefore add noise to the learning signal.</p>\n<p>That was my theory at least. In practice, I found that training this model on 30 sec clips did not perform better on the leaderboard than training it directly on 5 sec clips. </p>\n<p>It seems many other solution reused <a href=\"https://www.kaggle.com/code/hidehisaarai1213/introduction-to-sound-event-detection/notebook\" target=\"_blank\">SED models</a>, which sounds somewhat similar to this idea, except that it actually works--so I will be reading a lot more about that before BirdCLEF 2023 :)</p>\n<p>I also tried replacing my spectrograms and 1D convolutional feature extractor that I trained from scratch with the feature extractor of <a href=\"https://pytorch.org/audio/0.10.0/models.html#wav2vec2model\" target=\"_blank\">a pretrained Wav2Vec 2.0 model</a>. But finetuning didn't seem to perform better than my original baseline model.</p>\n<h2>Training</h2>\n<p>One question I struggled with was how to use the 90% of training data that didn't include any of the scored birds. I tried training a model on the full 152 classes, then finetuning on the just the 21 scored birds. But that didn't perform better on the leaderboard than just training directly on the 21 scored birds.</p>\n<p>The other huge issue I had was dealing with class imbalance. I read quite a bit about this problem, and tried several things: reweighing loss based on class frequency, <a href=\"https://www.kaggle.com/code/residentmario/undersampling-and-oversampling-imbalanced-data/notebook\" target=\"_blank\">oversampling</a> based on average or max class frequency, using <a href=\"https://arxiv.org/abs/1708.02002\" target=\"_blank\">focal loss</a> instead of BCE -- and combinations of those things. Oversampling seemed to perform the best, but in all cases the models generally performed well on frequent birds, and mostly ignored the less frequent ones.</p>\n<p>All models I trained overfit. I split the data randomly 80/20 and took the model that had the highest validation score (computed by replicating <a href=\"https://www.kaggle.com/dschettler8845\" target=\"_blank\">@dschettler8845</a>'s <a href=\"https://www.kaggle.com/competitions/birdclef-2022/discussion/314999\" target=\"_blank\">fantastic description of the score</a>). </p>\n<h2>Evaluation</h2>\n<p>I'm still a bit confused by how the evaluation data is formatted. Is each 1 minute clip just a concatenation of random 5 second clips, or are they continuous recordings? Are there birds in the clips that aren't present in one of the 21 scored birds? Is it possible that a 5 sec clip has no bird vocalizing at all? Is the distribution of birds at all similar to training, or are they somehow evenly distributed, or forced to be more even by just ignoring labels of frequent birds? So many questions that I probably should have asked in the discussion forum--though I'm not sure if the hosts could answer.</p>\n<h2>Surprises</h2>\n<p>Many solutions operated on Mel Spectrograms. I find this odd, since <a href=\"https://en.wikipedia.org/wiki/Mel_scale\" target=\"_blank\">the melscale</a> is (roughly) based on what humans can hear, not birds. I do regret not using a log frequency axis, such as <a href=\"https://en.wikipedia.org/wiki/Constant-Q_transform\" target=\"_blank\">Constant-Q</a> (<a href=\"https://librosa.org/doc/latest/generated/librosa.cqt.html\" target=\"_blank\">librosa</a>).</p>\n<p>Also, most solutions I've read about so far used 2D convolutions for feature extraction, often taking backbones pretrained on image tasks. I'm surprised by this for two reasons: (1) I'm sure there's probably <em>some</em> out-of-tune birds out there, but I don't think there's that much of a need for translational invariance on the frequency axis and (2) I wouldn't expect a model pretrained on images of real things to transfer well to spectrograms. It seems my intuition is off on both of these things.</p>\n<p>I was baffled to read <a href=\"https://www.kaggle.com/shinmurashinmura\" target=\"_blank\">@shinmurashinmura</a>'s <a href=\"https://www.kaggle.com/competitions/birdclef-2022/discussion/318999\" target=\"_blank\">note on how changing the threshold</a> from 0.5 -&gt; 0.01 increased their leaderboard score from 0.56 -&gt; 0.71. For my models, I found that there was a moderate increase on my score when a threshold change from 0.5 -&gt; 0.3 (0.51 -&gt; 0.57), but none beyond that. I think my confusion here might be due to how little I understood the evaluation data. </p>\n<h2>Conclusions</h2>\n<p>This competition was overwhelming, but quite fun for me, and a huge learning experience! I look forward to learning even more from reading carefully through the other solutions, and applying these learnings next time.</p>",
  "messages": [
    {
      "id": "1803537",
      "postDate": "05/27/2022 23:28:48",
      "content": "<p>It's probably a bit silly for me to make this note given how low my ranking was. But this was my first real kaggle competition, and I personally learned a lot, so I thought I'd write down some reflections—if for nothing else than for my present and future self. :)</p>\n<p>First: thanks to everyone who competed and shared their thoughts and questions, and especially to the hosts of this competition for putting so much effort in fostering community focused on using AI/ML for scientific research. It was a lot of fun to participate, and I'm looking forward to doing more kaggle competitions in the future.</p>\n<h2>Dataset</h2>\n<p>I first preprocessed the dataset into log magnitude spectrograms, with linear frequency bins from 1-10 kHz . The final resolution was a 128x256 spectrogram per 5 seconds of audio. For augmentation during training, I randomly masked out portions of the frequency and time axis, inspired by <a href=\"https://arxiv.org/abs/1904.08779\" target=\"_blank\">SpecAugment</a>.</p>\n<h2>Model</h2>\n<p>I saw that <a href=\"https://www.kaggle.com/kaerururu\" target=\"_blank\">@kaerururu</a> shared a <a href=\"https://www.kaggle.com/competitions/birdclef-2022/discussion/318081\" target=\"_blank\">well performing solution</a>, and that there was a ton of work that might be reusable from previous competitions. But instead of reusing their approaches, I decided to first try doing my own thing, originally planning to go back and try their ideas if I had time (but of course I ran out of time!).</p>\n<p>My first baseline model (which ended up being the best) was a simple 1D convolution over the time axis of the spectrograms as a feature extractor, followed by max pooling over time, then flattening the channels and frequency axis into a linear layer that projects onto the # of classes. </p>\n<pre><code>class DumbBirdClassifier(nn.Module):\n    def __init__(self):\n        super(DumbBirdClassifier, self).__init__()\n        self.conv = nn.Sequential(\n            nn.Conv1d(kernel_size=4, stride=2, padding=1, in_channels=128, out_channels=128),\n            nn.ReLU(),\n            nn.Conv1d(kernel_size=4, stride=2, padding=1, in_channels=128, out_channels=128),\n            nn.ReLU(),\n            nn.Conv1d(kernel_size=4, stride=2, padding=1, in_channels=128, out_channels=128),\n            nn.ReLU(),\n            nn.Conv1d(kernel_size=4, stride=2, padding=1, in_channels=128, out_channels=128),\n            nn.ReLU(),\n        )\n        self.pool = lambda x: x.max(-1).values\n        self.linear = nn.Sequential(\n            nn.Linear(in_features=128, out_features=128),\n            nn.ReLU(),\n            nn.Linear(in_features=128, out_features=NUM_CLASSES),\n            nn.Sigmoid(),\n        )\n\n    def forward(self, x):\n        return self.linear(self.pool(self.conv(x)))\n</code></pre>\n<p>I reasoned that collapsing the time axis with max pooling would make the model robust to different lengths of audio. The model must make evaluations on 5 sec-long clips, but it could be useful to be able to give it longer clips of audio during training, since only a 5 sec clip from a training recording would likely not include all or any of the labeled birds, and therefore add noise to the learning signal.</p>\n<p>That was my theory at least. In practice, I found that training this model on 30 sec clips did not perform better on the leaderboard than training it directly on 5 sec clips. </p>\n<p>It seems many other solution reused <a href=\"https://www.kaggle.com/code/hidehisaarai1213/introduction-to-sound-event-detection/notebook\" target=\"_blank\">SED models</a>, which sounds somewhat similar to this idea, except that it actually works--so I will be reading a lot more about that before BirdCLEF 2023 :)</p>\n<p>I also tried replacing my spectrograms and 1D convolutional feature extractor that I trained from scratch with the feature extractor of <a href=\"https://pytorch.org/audio/0.10.0/models.html#wav2vec2model\" target=\"_blank\">a pretrained Wav2Vec 2.0 model</a>. But finetuning didn't seem to perform better than my original baseline model.</p>\n<h2>Training</h2>\n<p>One question I struggled with was how to use the 90% of training data that didn't include any of the scored birds. I tried training a model on the full 152 classes, then finetuning on the just the 21 scored birds. But that didn't perform better on the leaderboard than just training directly on the 21 scored birds.</p>\n<p>The other huge issue I had was dealing with class imbalance. I read quite a bit about this problem, and tried several things: reweighing loss based on class frequency, <a href=\"https://www.kaggle.com/code/residentmario/undersampling-and-oversampling-imbalanced-data/notebook\" target=\"_blank\">oversampling</a> based on average or max class frequency, using <a href=\"https://arxiv.org/abs/1708.02002\" target=\"_blank\">focal loss</a> instead of BCE -- and combinations of those things. Oversampling seemed to perform the best, but in all cases the models generally performed well on frequent birds, and mostly ignored the less frequent ones.</p>\n<p>All models I trained overfit. I split the data randomly 80/20 and took the model that had the highest validation score (computed by replicating <a href=\"https://www.kaggle.com/dschettler8845\" target=\"_blank\">@dschettler8845</a>'s <a href=\"https://www.kaggle.com/competitions/birdclef-2022/discussion/314999\" target=\"_blank\">fantastic description of the score</a>). </p>\n<h2>Evaluation</h2>\n<p>I'm still a bit confused by how the evaluation data is formatted. Is each 1 minute clip just a concatenation of random 5 second clips, or are they continuous recordings? Are there birds in the clips that aren't present in one of the 21 scored birds? Is it possible that a 5 sec clip has no bird vocalizing at all? Is the distribution of birds at all similar to training, or are they somehow evenly distributed, or forced to be more even by just ignoring labels of frequent birds? So many questions that I probably should have asked in the discussion forum--though I'm not sure if the hosts could answer.</p>\n<h2>Surprises</h2>\n<p>Many solutions operated on Mel Spectrograms. I find this odd, since <a href=\"https://en.wikipedia.org/wiki/Mel_scale\" target=\"_blank\">the melscale</a> is (roughly) based on what humans can hear, not birds. I do regret not using a log frequency axis, such as <a href=\"https://en.wikipedia.org/wiki/Constant-Q_transform\" target=\"_blank\">Constant-Q</a> (<a href=\"https://librosa.org/doc/latest/generated/librosa.cqt.html\" target=\"_blank\">librosa</a>).</p>\n<p>Also, most solutions I've read about so far used 2D convolutions for feature extraction, often taking backbones pretrained on image tasks. I'm surprised by this for two reasons: (1) I'm sure there's probably <em>some</em> out-of-tune birds out there, but I don't think there's that much of a need for translational invariance on the frequency axis and (2) I wouldn't expect a model pretrained on images of real things to transfer well to spectrograms. It seems my intuition is off on both of these things.</p>\n<p>I was baffled to read <a href=\"https://www.kaggle.com/shinmurashinmura\" target=\"_blank\">@shinmurashinmura</a>'s <a href=\"https://www.kaggle.com/competitions/birdclef-2022/discussion/318999\" target=\"_blank\">note on how changing the threshold</a> from 0.5 -&gt; 0.01 increased their leaderboard score from 0.56 -&gt; 0.71. For my models, I found that there was a moderate increase on my score when a threshold change from 0.5 -&gt; 0.3 (0.51 -&gt; 0.57), but none beyond that. I think my confusion here might be due to how little I understood the evaluation data. </p>\n<h2>Conclusions</h2>\n<p>This competition was overwhelming, but quite fun for me, and a huge learning experience! I look forward to learning even more from reading carefully through the other solutions, and applying these learnings next time.</p>",
      "rawMarkdown": "It's probably a bit silly for me to make this note given how low my ranking was. But this was my first real kaggle competition, and I personally learned a lot, so I thought I'd write down some reflections—if for nothing else than for my present and future self. :)\n\nFirst: thanks to everyone who competed and shared their thoughts and questions, and especially to the hosts of this competition for putting so much effort in fostering community focused on using AI/ML for scientific research. It was a lot of fun to participate, and I'm looking forward to doing more kaggle competitions in the future.\n\n## Dataset\n\nI first preprocessed the dataset into log magnitude spectrograms, with linear frequency bins from 1-10 kHz . The final resolution was a 128x256 spectrogram per 5 seconds of audio. For augmentation during training, I randomly masked out portions of the frequency and time axis, inspired by [SpecAugment](https://arxiv.org/abs/1904.08779).\n\n## Model\n\nI saw that @kaerururu shared a [well performing solution](https://www.kaggle.com/competitions/birdclef-2022/discussion/318081), and that there was a ton of work that might be reusable from previous competitions. But instead of reusing their approaches, I decided to first try doing my own thing, originally planning to go back and try their ideas if I had time (but of course I ran out of time!).\n\nMy first baseline model (which ended up being the best) was a simple 1D convolution over the time axis of the spectrograms as a feature extractor, followed by max pooling over time, then flattening the channels and frequency axis into a linear layer that projects onto the # of classes. \n\n```\nclass DumbBirdClassifier(nn.Module):\n    def __init__(self):\n        super(DumbBirdClassifier, self).__init__()\n        self.conv = nn.Sequential(\n            nn.Conv1d(kernel_size=4, stride=2, padding=1, in_channels=128, out_channels=128),\n            nn.ReLU(),\n            nn.Conv1d(kernel_size=4, stride=2, padding=1, in_channels=128, out_channels=128),\n            nn.ReLU(),\n            nn.Conv1d(kernel_size=4, stride=2, padding=1, in_channels=128, out_channels=128),\n            nn.ReLU(),\n            nn.Conv1d(kernel_size=4, stride=2, padding=1, in_channels=128, out_channels=128),\n            nn.ReLU(),\n        )\n        self.pool = lambda x: x.max(-1).values\n        self.linear = nn.Sequential(\n            nn.Linear(in_features=128, out_features=128),\n            nn.ReLU(),\n            nn.Linear(in_features=128, out_features=NUM_CLASSES),\n            nn.Sigmoid(),\n        )\n    \n    def forward(self, x):\n        return self.linear(self.pool(self.conv(x)))\n```\n\nI reasoned that collapsing the time axis with max pooling would make the model robust to different lengths of audio. The model must make evaluations on 5 sec-long clips, but it could be useful to be able to give it longer clips of audio during training, since only a 5 sec clip from a training recording would likely not include all or any of the labeled birds, and therefore add noise to the learning signal.\n\nThat was my theory at least. In practice, I found that training this model on 30 sec clips did not perform better on the leaderboard than training it directly on 5 sec clips. \n\nIt seems many other solution reused [SED models](https://www.kaggle.com/code/hidehisaarai1213/introduction-to-sound-event-detection/notebook), which sounds somewhat similar to this idea, except that it actually works--so I will be reading a lot more about that before BirdCLEF 2023 :)\n\nI also tried replacing my spectrograms and 1D convolutional feature extractor that I trained from scratch with the feature extractor of [a pretrained Wav2Vec 2.0 model](https://pytorch.org/audio/0.10.0/models.html#wav2vec2model). But finetuning didn't seem to perform better than my original baseline model.\n\n## Training\n\nOne question I struggled with was how to use the 90% of training data that didn't include any of the scored birds. I tried training a model on the full 152 classes, then finetuning on the just the 21 scored birds. But that didn't perform better on the leaderboard than just training directly on the 21 scored birds.\n\nThe other huge issue I had was dealing with class imbalance. I read quite a bit about this problem, and tried several things: reweighing loss based on class frequency, [oversampling](https://www.kaggle.com/code/residentmario/undersampling-and-oversampling-imbalanced-data/notebook) based on average or max class frequency, using [focal loss](https://arxiv.org/abs/1708.02002) instead of BCE -- and combinations of those things. Oversampling seemed to perform the best, but in all cases the models generally performed well on frequent birds, and mostly ignored the less frequent ones.\n\nAll models I trained overfit. I split the data randomly 80/20 and took the model that had the highest validation score (computed by replicating @dschettler8845's [fantastic description of the score](https://www.kaggle.com/competitions/birdclef-2022/discussion/314999)). \n\n## Evaluation\n\nI'm still a bit confused by how the evaluation data is formatted. Is each 1 minute clip just a concatenation of random 5 second clips, or are they continuous recordings? Are there birds in the clips that aren't present in one of the 21 scored birds? Is it possible that a 5 sec clip has no bird vocalizing at all? Is the distribution of birds at all similar to training, or are they somehow evenly distributed, or forced to be more even by just ignoring labels of frequent birds? So many questions that I probably should have asked in the discussion forum--though I'm not sure if the hosts could answer.\n\n## Surprises\n\nMany solutions operated on Mel Spectrograms. I find this odd, since [the melscale](https://en.wikipedia.org/wiki/Mel_scale) is (roughly) based on what humans can hear, not birds. I do regret not using a log frequency axis, such as [Constant-Q](https://en.wikipedia.org/wiki/Constant-Q_transform) ([librosa](https://librosa.org/doc/latest/generated/librosa.cqt.html)).\n\nAlso, most solutions I've read about so far used 2D convolutions for feature extraction, often taking backbones pretrained on image tasks. I'm surprised by this for two reasons: (1) I'm sure there's probably *some* out-of-tune birds out there, but I don't think there's that much of a need for translational invariance on the frequency axis and (2) I wouldn't expect a model pretrained on images of real things to transfer well to spectrograms. It seems my intuition is off on both of these things.\n\nI was baffled to read @shinmurashinmura's [note on how changing the threshold](https://www.kaggle.com/competitions/birdclef-2022/discussion/318999) from 0.5 -> 0.01 increased their leaderboard score from 0.56 -> 0.71. For my models, I found that there was a moderate increase on my score when a threshold change from 0.5 -> 0.3 (0.51 -> 0.57), but none beyond that. I think my confusion here might be due to how little I understood the evaluation data. \n\n## Conclusions\n\nThis competition was overwhelming, but quite fun for me, and a huge learning experience! I look forward to learning even more from reading carefully through the other solutions, and applying these learnings next time.",
      "votes": null
    },
    {
      "id": "1803596",
      "postDate": "05/28/2022 01:11:29",
      "content": "<p>It is not silly to share anything. Kaggle competitions are hard and we are here to learn and share.</p>\n<p>From my (little) experience in kaggle I can tell you that I learn a lot more when I don't use public notebooks. I prefer to avoid forking them but I do levarage the ideas with my own implementation.</p>\n<p>Keep reading, learning and coding! Thanks for sharing.</p>",
      "rawMarkdown": "It is not silly to share anything. Kaggle competitions are hard and we are here to learn and share.\n\nFrom my (little) experience in kaggle I can tell you that I learn a lot more when I don't use public notebooks. I prefer to avoid forking them but I do levarage the ideas with my own implementation.\n\nKeep reading, learning and coding! Thanks for sharing.",
      "votes": null
    },
    {
      "id": "1804382",
      "postDate": "05/28/2022 20:53:14",
      "content": "<p>Hey, thanks for making an effort!</p>\n<p>The eval data is continuous 1-minute audio samples (this is typically what is meant by 'soundscapes'), with labels applied to each 5-second window.</p>",
      "rawMarkdown": "Hey, thanks for making an effort!\n\nThe eval data is continuous 1-minute audio samples (this is typically what is meant by 'soundscapes'), with labels applied to each 5-second window.",
      "votes": null
    },
    {
      "id": "1804452",
      "postDate": "05/29/2022 02:18:36",
      "content": "<p>I really appreciate your work as a newer kaggle user. I've done some work with neural networks and acoustic spectrogram data and I agree that it make way more sense to use time-axis convolutions. Maybe the frequency-dimension convolution doesn't matter very much or the network learns not to worry about it? Very interesting.</p>",
      "rawMarkdown": "I really appreciate your work as a newer kaggle user. I've done some work with neural networks and acoustic spectrogram data and I agree that it make way more sense to use time-axis convolutions. Maybe the frequency-dimension convolution doesn't matter very much or the network learns not to worry about it? Very interesting.",
      "votes": null
    },
    {
      "id": "1805090",
      "postDate": "05/29/2022 19:16:10",
      "content": "<p>Thanks! Ah ok, that makes sense—I had heard that term thrown around but wasn’t sure what it meant exactly.</p>",
      "rawMarkdown": "Thanks! Ah ok, that makes sense—I had heard that term thrown around but wasn’t sure what it meant exactly.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1803596,
      "author_name": "hinepo",
      "author_url": "",
      "post_date": "05/28/2022 01:11:29",
      "content": "<p>It is not silly to share anything. Kaggle competitions are hard and we are here to learn and share.</p>\n<p>From my (little) experience in kaggle I can tell you that I learn a lot more when I don't use public notebooks. I prefer to avoid forking them but I do levarage the ideas with my own implementation.</p>\n<p>Keep reading, learning and coding! Thanks for sharing.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1804382,
      "author_name": "tomdenton",
      "author_url": "",
      "post_date": "05/28/2022 20:53:14",
      "content": "<p>Hey, thanks for making an effort!</p>\n<p>The eval data is continuous 1-minute audio samples (this is typically what is meant by 'soundscapes'), with labels applied to each 5-second window.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1805090,
          "author_name": "robbynevels",
          "author_url": "",
          "post_date": "05/29/2022 19:16:10",
          "content": "<p>Thanks! Ah ok, that makes sense—I had heard that term thrown around but wasn’t sure what it meant exactly.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1804452,
      "author_name": "michaelmortenson",
      "author_url": "",
      "post_date": "05/29/2022 02:18:36",
      "content": "<p>I really appreciate your work as a newer kaggle user. I've done some work with neural networks and acoustic spectrogram data and I agree that it make way more sense to use time-axis convolutions. Maybe the frequency-dimension convolution doesn't matter very much or the network learns not to worry about it? Very interesting.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1803537": "It's probably a bit silly for me to make this note given how low my ranking was. But this was my first real kaggle competition, and I personally learned a lot, so I thought I'd write down some reflections—if for nothing else than for my present and future self. :)\n\nFirst: thanks to everyone who competed and shared their thoughts and questions, and especially to the hosts of this competition for putting so much effort in fostering community focused on using AI/ML for scientific research. It was a lot of fun to participate, and I'm looking forward to doing more kaggle competitions in the future.\n\n## Dataset\n\nI first preprocessed the dataset into log magnitude spectrograms, with linear frequency bins from 1-10 kHz . The final resolution was a 128x256 spectrogram per 5 seconds of audio. For augmentation during training, I randomly masked out portions of the frequency and time axis, inspired by [SpecAugment](https://arxiv.org/abs/1904.08779).\n\n## Model\n\nI saw that @kaerururu shared a [well performing solution](https://www.kaggle.com/competitions/birdclef-2022/discussion/318081), and that there was a ton of work that might be reusable from previous competitions. But instead of reusing their approaches, I decided to first try doing my own thing, originally planning to go back and try their ideas if I had time (but of course I ran out of time!).\n\nMy first baseline model (which ended up being the best) was a simple 1D convolution over the time axis of the spectrograms as a feature extractor, followed by max pooling over time, then flattening the channels and frequency axis into a linear layer that projects onto the # of classes. \n\n```\nclass DumbBirdClassifier(nn.Module):\n    def __init__(self):\n        super(DumbBirdClassifier, self).__init__()\n        self.conv = nn.Sequential(\n            nn.Conv1d(kernel_size=4, stride=2, padding=1, in_channels=128, out_channels=128),\n            nn.ReLU(),\n            nn.Conv1d(kernel_size=4, stride=2, padding=1, in_channels=128, out_channels=128),\n            nn.ReLU(),\n            nn.Conv1d(kernel_size=4, stride=2, padding=1, in_channels=128, out_channels=128),\n            nn.ReLU(),\n            nn.Conv1d(kernel_size=4, stride=2, padding=1, in_channels=128, out_channels=128),\n            nn.ReLU(),\n        )\n        self.pool = lambda x: x.max(-1).values\n        self.linear = nn.Sequential(\n            nn.Linear(in_features=128, out_features=128),\n            nn.ReLU(),\n            nn.Linear(in_features=128, out_features=NUM_CLASSES),\n            nn.Sigmoid(),\n        )\n    \n    def forward(self, x):\n        return self.linear(self.pool(self.conv(x)))\n```\n\nI reasoned that collapsing the time axis with max pooling would make the model robust to different lengths of audio. The model must make evaluations on 5 sec-long clips, but it could be useful to be able to give it longer clips of audio during training, since only a 5 sec clip from a training recording would likely not include all or any of the labeled birds, and therefore add noise to the learning signal.\n\nThat was my theory at least. In practice, I found that training this model on 30 sec clips did not perform better on the leaderboard than training it directly on 5 sec clips. \n\nIt seems many other solution reused [SED models](https://www.kaggle.com/code/hidehisaarai1213/introduction-to-sound-event-detection/notebook), which sounds somewhat similar to this idea, except that it actually works--so I will be reading a lot more about that before BirdCLEF 2023 :)\n\nI also tried replacing my spectrograms and 1D convolutional feature extractor that I trained from scratch with the feature extractor of [a pretrained Wav2Vec 2.0 model](https://pytorch.org/audio/0.10.0/models.html#wav2vec2model). But finetuning didn't seem to perform better than my original baseline model.\n\n## Training\n\nOne question I struggled with was how to use the 90% of training data that didn't include any of the scored birds. I tried training a model on the full 152 classes, then finetuning on the just the 21 scored birds. But that didn't perform better on the leaderboard than just training directly on the 21 scored birds.\n\nThe other huge issue I had was dealing with class imbalance. I read quite a bit about this problem, and tried several things: reweighing loss based on class frequency, [oversampling](https://www.kaggle.com/code/residentmario/undersampling-and-oversampling-imbalanced-data/notebook) based on average or max class frequency, using [focal loss](https://arxiv.org/abs/1708.02002) instead of BCE -- and combinations of those things. Oversampling seemed to perform the best, but in all cases the models generally performed well on frequent birds, and mostly ignored the less frequent ones.\n\nAll models I trained overfit. I split the data randomly 80/20 and took the model that had the highest validation score (computed by replicating @dschettler8845's [fantastic description of the score](https://www.kaggle.com/competitions/birdclef-2022/discussion/314999)). \n\n## Evaluation\n\nI'm still a bit confused by how the evaluation data is formatted. Is each 1 minute clip just a concatenation of random 5 second clips, or are they continuous recordings? Are there birds in the clips that aren't present in one of the 21 scored birds? Is it possible that a 5 sec clip has no bird vocalizing at all? Is the distribution of birds at all similar to training, or are they somehow evenly distributed, or forced to be more even by just ignoring labels of frequent birds? So many questions that I probably should have asked in the discussion forum--though I'm not sure if the hosts could answer.\n\n## Surprises\n\nMany solutions operated on Mel Spectrograms. I find this odd, since [the melscale](https://en.wikipedia.org/wiki/Mel_scale) is (roughly) based on what humans can hear, not birds. I do regret not using a log frequency axis, such as [Constant-Q](https://en.wikipedia.org/wiki/Constant-Q_transform) ([librosa](https://librosa.org/doc/latest/generated/librosa.cqt.html)).\n\nAlso, most solutions I've read about so far used 2D convolutions for feature extraction, often taking backbones pretrained on image tasks. I'm surprised by this for two reasons: (1) I'm sure there's probably *some* out-of-tune birds out there, but I don't think there's that much of a need for translational invariance on the frequency axis and (2) I wouldn't expect a model pretrained on images of real things to transfer well to spectrograms. It seems my intuition is off on both of these things.\n\nI was baffled to read @shinmurashinmura's [note on how changing the threshold](https://www.kaggle.com/competitions/birdclef-2022/discussion/318999) from 0.5 -> 0.01 increased their leaderboard score from 0.56 -> 0.71. For my models, I found that there was a moderate increase on my score when a threshold change from 0.5 -> 0.3 (0.51 -> 0.57), but none beyond that. I think my confusion here might be due to how little I understood the evaluation data. \n\n## Conclusions\n\nThis competition was overwhelming, but quite fun for me, and a huge learning experience! I look forward to learning even more from reading carefully through the other solutions, and applying these learnings next time.",
    "1803596": "It is not silly to share anything. Kaggle competitions are hard and we are here to learn and share.\n\nFrom my (little) experience in kaggle I can tell you that I learn a lot more when I don't use public notebooks. I prefer to avoid forking them but I do levarage the ideas with my own implementation.\n\nKeep reading, learning and coding! Thanks for sharing.",
    "1804382": "Hey, thanks for making an effort!\n\nThe eval data is continuous 1-minute audio samples (this is typically what is meant by 'soundscapes'), with labels applied to each 5-second window.",
    "1804452": "I really appreciate your work as a newer kaggle user. I've done some work with neural networks and acoustic spectrogram data and I agree that it make way more sense to use time-axis convolutions. Maybe the frequency-dimension convolution doesn't matter very much or the network learns not to worry about it? Very interesting.",
    "1805090": "Thanks! Ah ok, that makes sense—I had heard that term thrown around but wasn’t sure what it meant exactly."
  },
  "source": "meta"
}