{
  "id": 168657,
  "title": "what works in this competition?!",
  "url": "/competitions/birdsong-recognition/discussion/168657",
  "author_name": "",
  "post_date": "2020-07-21T11:59:05.590859900Z",
  "votes": 33,
  "comment_count": 16,
  "views": 0,
  "content": "<p>In this <a href=\"https://www.kaggle.com/ttahara/training-birdsong-baseline-resnest50-fast\">great kernel</a>, based off @hidehisaarai1213's work, @ttahara ahieves an LB score of 0.568 through training on a single fold.</p>\n\n<p>I have tried to achieve a comparable result using various building blocks that to my understanding should perform at least as well as the kernel, but my scores on a single fold ranges from 0.554 to maybe 0.562 (the best I got was 0.564 but that was just a single model, probably an outlier).</p>\n\n<p>Where lies the magic? Is it the training with cosine annealing? Is it using the smallest of resnest50 models? (through not paying attention I was using the default resnest50 which is slightly larger I believe).</p>\n\n<p>At this point I am a little bit out of ideas. Based on results, I don't think the method of generating spectrograms is the defining factor here.  The way I generate spectrograms in this <a href=\"https://www.kaggle.com/radek1/esp-starter-pack-from-training-to-submission\">kernel</a> seems to be at least as good (might even be better).</p>\n\n<p>The one thing I haven't tried (trying now) is training on resized spectrograms. But why should this work? Maybe, maybe somehow the shapes are easier for the model to work with when they are resized to 224 across the shorter dimension... But resizing doesn't introduce any new information??? Why should training a model on upscaled spectrograms work better? Especially that it requires reducing the batch size,  which I am not sure is that great for training with that many classes and data as noisy as this... </p>\n\n<p>Well, at this point, with a good dose of frustration so familiar to everyone who has done any machine learning at any point in their lives 😉, I would like to report I have no clue. I would like to find the reason - maybe indeed it is the example size 🤔, but at the same time I am hoping this is not the case. If training on bigger, upscaled spectrograms would work better, it would give advantage to people with better access to hardware, and would work like that probably in other CV competitions as well. I fully accept that if you start with an image that is 512x512 and you resize it to 128x128 you lose more information than if you resize it to 224x224. In this case, I am happy to agree that training on bigger examples should work better. But why should this mechanism hold for upscaling that introduces no new information and might actually (maybe?) produce some artifacts in the process or some other loss of fidelity?!</p>",
  "messages": [
    {
      "id": "938207",
      "postDate": "07/21/2020 11:59:05",
      "content": "<p>In this <a href=\"https://www.kaggle.com/ttahara/training-birdsong-baseline-resnest50-fast\">great kernel</a>, based off @hidehisaarai1213's work, @ttahara ahieves an LB score of 0.568 through training on a single fold.</p>\n\n<p>I have tried to achieve a comparable result using various building blocks that to my understanding should perform at least as well as the kernel, but my scores on a single fold ranges from 0.554 to maybe 0.562 (the best I got was 0.564 but that was just a single model, probably an outlier).</p>\n\n<p>Where lies the magic? Is it the training with cosine annealing? Is it using the smallest of resnest50 models? (through not paying attention I was using the default resnest50 which is slightly larger I believe).</p>\n\n<p>At this point I am a little bit out of ideas. Based on results, I don't think the method of generating spectrograms is the defining factor here.  The way I generate spectrograms in this <a href=\"https://www.kaggle.com/radek1/esp-starter-pack-from-training-to-submission\">kernel</a> seems to be at least as good (might even be better).</p>\n\n<p>The one thing I haven't tried (trying now) is training on resized spectrograms. But why should this work? Maybe, maybe somehow the shapes are easier for the model to work with when they are resized to 224 across the shorter dimension... But resizing doesn't introduce any new information??? Why should training a model on upscaled spectrograms work better? Especially that it requires reducing the batch size,  which I am not sure is that great for training with that many classes and data as noisy as this... </p>\n\n<p>Well, at this point, with a good dose of frustration so familiar to everyone who has done any machine learning at any point in their lives 😉, I would like to report I have no clue. I would like to find the reason - maybe indeed it is the example size 🤔, but at the same time I am hoping this is not the case. If training on bigger, upscaled spectrograms would work better, it would give advantage to people with better access to hardware, and would work like that probably in other CV competitions as well. I fully accept that if you start with an image that is 512x512 and you resize it to 128x128 you lose more information than if you resize it to 224x224. In this case, I am happy to agree that training on bigger examples should work better. But why should this mechanism hold for upscaling that introduces no new information and might actually (maybe?) produce some artifacts in the process or some other loss of fidelity?!</p>",
      "rawMarkdown": "In this [great kernel](https://www.kaggle.com/ttahara/training-birdsong-baseline-resnest50-fast), based off @hidehisaarai1213's work, @ttahara ahieves an LB score of 0.568 through training on a single fold.\n\nI have tried to achieve a comparable result using various building blocks that to my understanding should perform at least as well as the kernel, but my scores on a single fold ranges from 0.554 to maybe 0.562 (the best I got was 0.564 but that was just a single model, probably an outlier).\n\nWhere lies the magic? Is it the training with cosine annealing? Is it using the smallest of resnest50 models? (through not paying attention I was using the default resnest50 which is slightly larger I believe).\n\nAt this point I am a little bit out of ideas. Based on results, I don't think the method of generating spectrograms is the defining factor here.  The way I generate spectrograms in this [kernel](https://www.kaggle.com/radek1/esp-starter-pack-from-training-to-submission) seems to be at least as good (might even be better).\n\nThe one thing I haven't tried (trying now) is training on resized spectrograms. But why should this work? Maybe, maybe somehow the shapes are easier for the model to work with when they are resized to 224 across the shorter dimension... But resizing doesn't introduce any new information??? Why should training a model on upscaled spectrograms work better? Especially that it requires reducing the batch size, \n\nWell, at this point, with a good dose of frustration so familiar to everyone who has done any machine learning at any point in their lives 😉, I would like to report I have no clue. I would like to find the reason - maybe indeed it is the example size 🤔, but at the same time I am hoping this is not the case. If training on bigger, upscaled spectrograms would work better, it would give advantage to people with better access to hardware, and would work like that probably in other CV competitions as well. I fully accept that if you start with an image that is 512x512 and you resize it to 128x128 you lose more information than if you resize it to 224x224. In this case, I am happy to agree that training on bigger examples should work better. But why should this mechanism hold for upscaling that introduces no new information and might actually (maybe?) produce some artifacts in the process or some other loss of fidelity?!",
      "votes": null
    },
    {
      "id": "938370",
      "postDate": "07/21/2020 13:38:58",
      "content": "<p>A couple of ideas that I have:</p>\n<ol>\n<li>mels is a scale for humans but birds sing for birds and do not care about humans. and even for humans there are different scales. For instance, rnnoise lib uses some Bark scale to suppress noise. It is still beyond me to comprehend the implications though</li>\n<li>a \"song\" label from the same birds can have completely different sounds and thus spectograms i.e. either train a model on \"bird + type of song\" or time of the year or locaiton (i.e. pump up classes) + augmentations + noise or raw spectograms are a wrong go-to tool</li>\n<li>spectograms lose information by doing FFT by construction. Maybe they lose the wrong kind of information, i.e. time dimension can be more important than frequency. Or blending 2 models: 1 time heavy and 1 frequency heavy, might be the trick</li>\n<li>some more but I forgot and listing these out of my head :)</li>\n</ol>",
      "rawMarkdown": "A couple of ideas that I have:\n1. mels is a scale for humans but birds sing for birds and do not care about humans. and even for humans there are different scales. For instance, rnnoise lib uses some Bark scale to suppress noise. It is still beyond me to comprehend the implications though\n2. a \"song\" label from the same birds can have completely different sounds and thus spectograms i.e. either train a model on \"bird + type of song\" or time of the year or locaiton (i.e. pump up classes) + augmentations + noise or raw spectograms are a wrong go-to tool\n3. spectograms lose information by doing FFT by construction. Maybe they lose the wrong kind of information, i.e. time dimension can be more important than frequency. Or blending 2 models: 1 time heavy and 1 frequency heavy, might be the trick\n4. some more but I forgot and listing these out of my head :)",
      "votes": null
    },
    {
      "id": "938475",
      "postDate": "07/21/2020 15:01:04",
      "content": "<p>The other factor is the sensitivity of the f1_score to your choice of threshold. Reducing the threshold from 0.8 to 0.6 results in a significant (0.562 -&gt; 0.494) LB score drop using the same model:</p>\n<p><img src=\"https://www.googleapis.com/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F36054%2Ff0a98ee24655fbfca1b9ac2b5932443b%2Fscores.png\" alt=\"\"></p>",
      "rawMarkdown": "The other factor is the sensitivity of the f1_score to your choice of threshold. Reducing the threshold from 0.8 to 0.6 results in a significant (0.562 -&gt; 0.494) LB score drop using the same model:\n\n![](https://www.googleapis.com/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F36054%2Ff0a98ee24655fbfca1b9ac2b5932443b%2Fscores.png)",
      "votes": null
    },
    {
      "id": "939035",
      "postDate": "07/22/2020 01:48:17",
      "content": "<p>As for 1. , I think mel scale is ok, because annotation comes out of humans.</p>",
      "rawMarkdown": "As for 1. , I think mel scale is ok, because annotation comes out of humans.",
      "votes": null
    },
    {
      "id": "939043",
      "postDate": "07/22/2020 01:59:46",
      "content": "<p>Personally I decided to use variable threshold per class based on the result in train set. This cannot remove the sensitivity to threshold, but at least it can automatically decide what threshold to use; and as far as I have observed so far it's working fine (not very good but just fine).</p>",
      "rawMarkdown": "Personally I decided to use variable threshold per class based on the result in train set. This cannot remove the sensitivity to threshold, but at least it can automatically decide what threshold to use; and as far as I have observed so far it's working fine (not very good but just fine).",
      "votes": null
    },
    {
      "id": "939201",
      "postDate": "07/22/2020 05:00:42",
      "content": "<p>I published my baseline as <strong>just a baseline</strong>. It is unexpectedly strong and I don't know why🤔</p>\n<blockquote>\n  <p>I was using the default resnest50 which is slightly larger I believe.</p>\n</blockquote>\n<p>I think so. I suspect lighter models with special training techniques such as MIL(you've suggested) suit this competition.</p>\n<blockquote>\n  <p>The one thing I haven't tried (trying now) is training on resized spectrograms. But why should this work?</p>\n</blockquote>\n<p>I have the same doubt about model input sizes.  <br>\nOne hypothesis is that difference in resolution between pretraining task(ImageNet) and downstream task(this competition) influences models' performance. But I'm not certain about this idea.</p>\n<p>I'd like to get your opinion, <a href=\"https://www.kaggle.com/hidehisaarai1213\" target=\"_blank\">@hidehisaarai1213</a></p>",
      "rawMarkdown": "I published my baseline as **just a baseline**. It is unexpectedly strong and I don't know why🤔\n\n&gt;  I was using the default resnest50 which is slightly larger I believe.\n\nI think so. I suspect lighter models with special training techniques such as MIL(you've suggested) suit this competition.\n\n&gt; The one thing I haven't tried (trying now) is training on resized spectrograms. But why should this work?\n\nI have the same doubt about model input sizes.  \nOne hypothesis is that difference in resolution between pretraining task(ImageNet) and downstream task(this competition) influences models' performance. But I'm not certain about this idea.\n\nI'd like to get your opinion, @hidehisaarai1213",
      "votes": null
    },
    {
      "id": "939367",
      "postDate": "07/22/2020 07:34:06",
      "content": "<blockquote>\n  <p>lighter models with special training techniques such as MIL(you've suggested) suit this competition.</p>\n</blockquote>\n\n<p>I agree with this. I think that backbone structure does not make significant difference as long as they are expressive enough, however, the use of pretrained weight may affects. I tried pretrained (with AudioSet) PANNs / untrained PANNs &amp; pretrained (with ImageNet) ResNet / untrained ResNet and found that pretrained model does better, so if we want to try custom architecture as backbone it may be good to train it with ImageNet / AudioSet at first and then train it with the data of this competition.</p>\n\n<blockquote>\n  <p>resolution between pretraining task(ImageNet) and downstream task(this competition) influences models' performance</p>\n</blockquote>\n\n<p>Pretrained ResNet / ResNest are trained with 224 x 224 images thus kernels of CNN are trained to extract patterns in the image in that size. Maybe that is why resizing the input to 224 works better.</p>",
      "rawMarkdown": "&gt; lighter models with special training techniques such as MIL(you've suggested) suit this competition.\n\nI agree with this. I think that backbone structure does not make significant difference as long as they are expressive enough, however, the use of pretrained weight may affects. I tried pretrained (with AudioSet) PANNs / untrained PANNs &amp; pretrained (with ImageNet) ResNet / untrained ResNet and found that pretrained model does better, so if we want to try custom architecture as backbone it may be good to train it with ImageNet / AudioSet at first and then train it with the data of this competition.\n\n&gt; resolution between pretraining task(ImageNet) and downstream task(this competition) influences models' performance\n\nPretrained ResNet / ResNest are trained with 224 x 224 images thus kernels of CNN are trained to extract patterns in the image in that size. Maybe that is why resizing the input to 224 works better.",
      "votes": null
    },
    {
      "id": "939432",
      "postDate": "07/22/2020 08:37:04",
      "content": "<p>On one hand I am relieved to finally find the reason for the difference in performance… on the other hand the finding is a little bit disheartening…</p>\n\n<p>Training with spectrograms upscaled to 224x my model achieves 0.567, where training in the same fashion but on 128x achieved a score of 0.556.</p>\n\n<p>It is interesting, because the relative improvement is enormous. I am not sure if the change in performance is due to the model being pretrained on 224x224 images. The shapes the model would see on imagenet are very different to what it sees here. I am thinking it might rather be due to the model architecture being fitted to 224x224 examples.</p>\n\n<p>I don't really know yet what this result means - have to ponder on it a little bit.</p>\n\n<p>I think it would be useful to run training on 256x or maybe even bigger sizes? One could use the best weights from 224x as a starting point. I am not sure if Kaggle hw would permit, but would be interesting to see how the starter packs would perform with such an increase in size, or even larger.</p>\n\n<p>Right now I am training with the smaller resnest and will also redo the training with res34. If there were no difference in performance being able to train with smaller models would be useful, could help to iterate faster, can also be useful for applications as inference with smaller models is cheaper / faster.</p>",
      "rawMarkdown": "On one hand I am relieved to finally find the reason for the difference in performance… on the other hand the finding is a little bit disheartening…\n\nTraining with spectrograms upscaled to 224x my model achieves 0.567, where training in the same fashion but on 128x achieved a score of 0.556.\n\nIt is interesting, because the relative improvement is enormous. I am not sure if the change in performance is due to the model being pretrained on 224x224 images. The shapes the model would see on imagenet are very different to what it sees here. I am thinking it might rather be due to the model architecture being fitted to 224x224 examples.\n\nI don't really know yet what this result means - have to ponder on it a little bit.\n\nI think it would be useful to run training on 256x or maybe even bigger sizes? One could use the best weights from 224x as a starting point. I am not sure if Kaggle hw would permit, but would be interesting to see how the starter packs would perform with such an increase in size, or even larger.\n\nRight now I am training with the smaller resnest and will also redo the training with res34. If there were no difference in performance being able to train with smaller models would be useful, could help to iterate faster, can also be useful for applications as inference with smaller models is cheaper / faster.",
      "votes": null
    },
    {
      "id": "939531",
      "postDate": "07/22/2020 09:45:56",
      "content": "<p>The point of resizing is somewhat misplaced. The dimensions of spectogram is the result of your direct choice coming out of the FFT. You can directly produce any pixel dimensions you like by choosing appropriate parameters FFT, i.e. your sound clip length + window length + your sampling rate will give you the exact image dimension on the x side.</p>",
      "rawMarkdown": "The point of resizing is somewhat misplaced. The dimensions of spectogram is the result of your direct choice coming out of the FFT. You can directly produce any pixel dimensions you like by choosing appropriate parameters FFT, i.e. your sound clip length + window length + your sampling rate will give you the exact image dimension on the x side.",
      "votes": null
    },
    {
      "id": "939543",
      "postDate": "07/22/2020 09:52:12",
      "content": "<p>In Audio Tagging 2019 competition I also found that upscaled images worked great (128x128 crop upscale to 256x256). Interestingly it worked better than cropping 128x256 and upscaling to 256x256 and the models were not pretrained. </p>",
      "rawMarkdown": "In Audio Tagging 2019 competition I also found that upscaled images worked great (128x128 crop upscale to 256x256). Interestingly it worked better than cropping 128x256 and upscaling to 256x256 and the models were not pretrained.",
      "votes": null
    },
    {
      "id": "939974",
      "postDate": "07/22/2020 15:50:01",
      "content": "<p>Thanks for answering!</p>\n<p>The Relation between pretrained models on images and audio tasks is slightly strange and interesting.</p>",
      "rawMarkdown": "Thanks for answering!\n\nThe Relation between pretrained models on images and audio tasks is slightly strange and interesting.",
      "votes": null
    },
    {
      "id": "961533",
      "postDate": "08/07/2020 09:09:41",
      "content": "<p>I quit using variable threshold. Working on threshold optimization is muddy and has no significant difference. Maybe good to use 0.5: I checked some technical reports of DCASE challenges and they usually use threshold of 0.5 when they are not certain about test data.</p>",
      "rawMarkdown": "I quit using variable threshold. Working on threshold optimization is muddy and has no significant difference. Maybe good to use 0.5: I checked some technical reports of DCASE challenges and they usually use threshold of 0.5 when they are not certain about test data.",
      "votes": null
    },
    {
      "id": "962579",
      "postDate": "08/08/2020 08:46:45",
      "content": "<p>Great work Upvoted for you!</p>",
      "rawMarkdown": "Great work Upvoted for you!",
      "votes": null
    },
    {
      "id": "965734",
      "postDate": "08/10/2020 19:56:00",
      "content": "<p>Yes, but in many cases the bird was seen, so the annotation is not <em>completely</em> based off the audio</p>",
      "rawMarkdown": "Yes, but in many cases the bird was seen, so the annotation is not _completely_ based off the audio",
      "votes": null
    },
    {
      "id": "991343",
      "postDate": "08/30/2020 10:34:46",
      "content": "<blockquote>\n  <p>Where lies the magic?</p>\n</blockquote>\n<p>Maybe this notebook is just overfitting to public LB.  </p>",
      "rawMarkdown": "> Where lies the magic?\n\nMaybe this notebook is just overfitting to public LB.",
      "votes": null
    },
    {
      "id": "994929",
      "postDate": "09/02/2020 04:31:19",
      "content": "<p><a href=\"https://www.kaggle.com/hidehisaarai1213\" target=\"_blank\">@hidehisaarai1213</a> , thanks for sharing. Are you using a threshold higher than 0.5? I find when using additional labels, such as adding secondary labels, a higher threshold has a better score. This could be due to the nature of public LB, but challenging as it is hard to tell if this threshold is optimal for private LB</p>",
      "rawMarkdown": "hidehisaarai1213 , thanks for sharing. Are you using a threshold higher than 0.5? I find when using additional labels, such as adding secondary labels, a higher threshold has a better score. This could be due to the nature of public LB, but challenging as it is hard to tell if this threshold is optimal for private LB",
      "votes": null
    },
    {
      "id": "995116",
      "postDate": "09/02/2020 07:15:26",
      "content": "<blockquote>\n  <p>Are you using a threshold higher than 0.5?</p>\n</blockquote>\n<p>I did, but not now. I've fixed the threshold to 0.5.</p>\n<blockquote>\n  <p>I find when using additional labels, such as adding secondary labels, a higher threshold has a better score.</p>\n</blockquote>\n<p>Interesting observation. I think this comes from the nature of train data. You may label nocall / some call events but not those of species in secondary labels as those of species in secondary labels, because the label is weak. This will make the model produce higher probability on nuisance signals and lead us to get lots of false positives.</p>",
      "rawMarkdown": "> Are you using a threshold higher than 0.5?\n\nI did, but not now. I've fixed the threshold to 0.5.\n\n> I find when using additional labels, such as adding secondary labels, a higher threshold has a better score.\n\nInteresting observation. I think this comes from the nature of train data. You may label nocall / some call events but not those of species in secondary labels as those of species in secondary labels, because the label is weak. This will make the model produce higher probability on nuisance signals and lead us to get lots of false positives.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 938370,
      "author_name": "snovik1975",
      "author_url": "",
      "post_date": "07/21/2020 13:38:58",
      "content": "<p>A couple of ideas that I have:</p>\n<ol>\n<li>mels is a scale for humans but birds sing for birds and do not care about humans. and even for humans there are different scales. For instance, rnnoise lib uses some Bark scale to suppress noise. It is still beyond me to comprehend the implications though</li>\n<li>a \"song\" label from the same birds can have completely different sounds and thus spectograms i.e. either train a model on \"bird + type of song\" or time of the year or locaiton (i.e. pump up classes) + augmentations + noise or raw spectograms are a wrong go-to tool</li>\n<li>spectograms lose information by doing FFT by construction. Maybe they lose the wrong kind of information, i.e. time dimension can be more important than frequency. Or blending 2 models: 1 time heavy and 1 frequency heavy, might be the trick</li>\n<li>some more but I forgot and listing these out of my head :)</li>\n</ol>",
      "votes": null,
      "replies": [
        {
          "id": 939035,
          "author_name": "hidehisaarai1213",
          "author_url": "",
          "post_date": "07/22/2020 01:48:17",
          "content": "<p>As for 1. , I think mel scale is ok, because annotation comes out of humans.</p>",
          "votes": null,
          "replies": [
            {
              "id": 965734,
              "author_name": "marcogorelli",
              "author_url": "",
              "post_date": "08/10/2020 19:56:00",
              "content": "<p>Yes, but in many cases the bird was seen, so the annotation is not <em>completely</em> based off the audio</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 938475,
      "author_name": "ramarlina",
      "author_url": "",
      "post_date": "07/21/2020 15:01:04",
      "content": "<p>The other factor is the sensitivity of the f1_score to your choice of threshold. Reducing the threshold from 0.8 to 0.6 results in a significant (0.562 -&gt; 0.494) LB score drop using the same model:</p>\n<p><img src=\"https://www.googleapis.com/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F36054%2Ff0a98ee24655fbfca1b9ac2b5932443b%2Fscores.png\" alt=\"\"></p>",
      "votes": null,
      "replies": [
        {
          "id": 939043,
          "author_name": "hidehisaarai1213",
          "author_url": "",
          "post_date": "07/22/2020 01:59:46",
          "content": "<p>Personally I decided to use variable threshold per class based on the result in train set. This cannot remove the sensitivity to threshold, but at least it can automatically decide what threshold to use; and as far as I have observed so far it's working fine (not very good but just fine).</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 961533,
          "author_name": "hidehisaarai1213",
          "author_url": "",
          "post_date": "08/07/2020 09:09:41",
          "content": "<p>I quit using variable threshold. Working on threshold optimization is muddy and has no significant difference. Maybe good to use 0.5: I checked some technical reports of DCASE challenges and they usually use threshold of 0.5 when they are not certain about test data.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 994929,
          "author_name": "alanchn31",
          "author_url": "",
          "post_date": "09/02/2020 04:31:19",
          "content": "<p><a href=\"https://www.kaggle.com/hidehisaarai1213\" target=\"_blank\">@hidehisaarai1213</a> , thanks for sharing. Are you using a threshold higher than 0.5? I find when using additional labels, such as adding secondary labels, a higher threshold has a better score. This could be due to the nature of public LB, but challenging as it is hard to tell if this threshold is optimal for private LB</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 995116,
          "author_name": "hidehisaarai1213",
          "author_url": "",
          "post_date": "09/02/2020 07:15:26",
          "content": "<blockquote>\n  <p>Are you using a threshold higher than 0.5?</p>\n</blockquote>\n<p>I did, but not now. I've fixed the threshold to 0.5.</p>\n<blockquote>\n  <p>I find when using additional labels, such as adding secondary labels, a higher threshold has a better score.</p>\n</blockquote>\n<p>Interesting observation. I think this comes from the nature of train data. You may label nocall / some call events but not those of species in secondary labels as those of species in secondary labels, because the label is weak. This will make the model produce higher probability on nuisance signals and lead us to get lots of false positives.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 939201,
      "author_name": "ttahara",
      "author_url": "",
      "post_date": "07/22/2020 05:00:42",
      "content": "<p>I published my baseline as <strong>just a baseline</strong>. It is unexpectedly strong and I don't know why🤔</p>\n<blockquote>\n  <p>I was using the default resnest50 which is slightly larger I believe.</p>\n</blockquote>\n<p>I think so. I suspect lighter models with special training techniques such as MIL(you've suggested) suit this competition.</p>\n<blockquote>\n  <p>The one thing I haven't tried (trying now) is training on resized spectrograms. But why should this work?</p>\n</blockquote>\n<p>I have the same doubt about model input sizes.  <br>\nOne hypothesis is that difference in resolution between pretraining task(ImageNet) and downstream task(this competition) influences models' performance. But I'm not certain about this idea.</p>\n<p>I'd like to get your opinion, <a href=\"https://www.kaggle.com/hidehisaarai1213\" target=\"_blank\">@hidehisaarai1213</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 939367,
          "author_name": "hidehisaarai1213",
          "author_url": "",
          "post_date": "07/22/2020 07:34:06",
          "content": "<blockquote>\n  <p>lighter models with special training techniques such as MIL(you've suggested) suit this competition.</p>\n</blockquote>\n\n<p>I agree with this. I think that backbone structure does not make significant difference as long as they are expressive enough, however, the use of pretrained weight may affects. I tried pretrained (with AudioSet) PANNs / untrained PANNs &amp; pretrained (with ImageNet) ResNet / untrained ResNet and found that pretrained model does better, so if we want to try custom architecture as backbone it may be good to train it with ImageNet / AudioSet at first and then train it with the data of this competition.</p>\n\n<blockquote>\n  <p>resolution between pretraining task(ImageNet) and downstream task(this competition) influences models' performance</p>\n</blockquote>\n\n<p>Pretrained ResNet / ResNest are trained with 224 x 224 images thus kernels of CNN are trained to extract patterns in the image in that size. Maybe that is why resizing the input to 224 works better.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 939974,
          "author_name": "ttahara",
          "author_url": "",
          "post_date": "07/22/2020 15:50:01",
          "content": "<p>Thanks for answering!</p>\n<p>The Relation between pretrained models on images and audio tasks is slightly strange and interesting.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 939531,
      "author_name": "snovik1975",
      "author_url": "",
      "post_date": "07/22/2020 09:45:56",
      "content": "<p>The point of resizing is somewhat misplaced. The dimensions of spectogram is the result of your direct choice coming out of the FFT. You can directly produce any pixel dimensions you like by choosing appropriate parameters FFT, i.e. your sound clip length + window length + your sampling rate will give you the exact image dimension on the x side.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 991343,
      "author_name": "cpmpml",
      "author_url": "",
      "post_date": "08/30/2020 10:34:46",
      "content": "<blockquote>\n  <p>Where lies the magic?</p>\n</blockquote>\n<p>Maybe this notebook is just overfitting to public LB.  </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 939432,
      "author_name": "radek1",
      "author_url": "",
      "post_date": "07/22/2020 08:37:04",
      "content": "<p>On one hand I am relieved to finally find the reason for the difference in performance… on the other hand the finding is a little bit disheartening…</p>\n\n<p>Training with spectrograms upscaled to 224x my model achieves 0.567, where training in the same fashion but on 128x achieved a score of 0.556.</p>\n\n<p>It is interesting, because the relative improvement is enormous. I am not sure if the change in performance is due to the model being pretrained on 224x224 images. The shapes the model would see on imagenet are very different to what it sees here. I am thinking it might rather be due to the model architecture being fitted to 224x224 examples.</p>\n\n<p>I don't really know yet what this result means - have to ponder on it a little bit.</p>\n\n<p>I think it would be useful to run training on 256x or maybe even bigger sizes? One could use the best weights from 224x as a starting point. I am not sure if Kaggle hw would permit, but would be interesting to see how the starter packs would perform with such an increase in size, or even larger.</p>\n\n<p>Right now I am training with the smaller resnest and will also redo the training with res34. If there were no difference in performance being able to train with smaller models would be useful, could help to iterate faster, can also be useful for applications as inference with smaller models is cheaper / faster.</p>",
      "votes": null,
      "replies": [
        {
          "id": 939543,
          "author_name": "mnpinto",
          "author_url": "",
          "post_date": "07/22/2020 09:52:12",
          "content": "<p>In Audio Tagging 2019 competition I also found that upscaled images worked great (128x128 crop upscale to 256x256). Interestingly it worked better than cropping 128x256 and upscaling to 256x256 and the models were not pretrained. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 962579,
      "author_name": "",
      "author_url": "",
      "post_date": "08/08/2020 08:46:45",
      "content": "<p>Great work Upvoted for you!</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "938207": "In this [great kernel](https://www.kaggle.com/ttahara/training-birdsong-baseline-resnest50-fast), based off @hidehisaarai1213's work, @ttahara ahieves an LB score of 0.568 through training on a single fold.\n\nI have tried to achieve a comparable result using various building blocks that to my understanding should perform at least as well as the kernel, but my scores on a single fold ranges from 0.554 to maybe 0.562 (the best I got was 0.564 but that was just a single model, probably an outlier).\n\nWhere lies the magic? Is it the training with cosine annealing? Is it using the smallest of resnest50 models? (through not paying attention I was using the default resnest50 which is slightly larger I believe).\n\nAt this point I am a little bit out of ideas. Based on results, I don't think the method of generating spectrograms is the defining factor here.  The way I generate spectrograms in this [kernel](https://www.kaggle.com/radek1/esp-starter-pack-from-training-to-submission) seems to be at least as good (might even be better).\n\nThe one thing I haven't tried (trying now) is training on resized spectrograms. But why should this work? Maybe, maybe somehow the shapes are easier for the model to work with when they are resized to 224 across the shorter dimension... But resizing doesn't introduce any new information??? Why should training a model on upscaled spectrograms work better? Especially that it requires reducing the batch size, \n\nWell, at this point, with a good dose of frustration so familiar to everyone who has done any machine learning at any point in their lives 😉, I would like to report I have no clue. I would like to find the reason - maybe indeed it is the example size 🤔, but at the same time I am hoping this is not the case. If training on bigger, upscaled spectrograms would work better, it would give advantage to people with better access to hardware, and would work like that probably in other CV competitions as well. I fully accept that if you start with an image that is 512x512 and you resize it to 128x128 you lose more information than if you resize it to 224x224. In this case, I am happy to agree that training on bigger examples should work better. But why should this mechanism hold for upscaling that introduces no new information and might actually (maybe?) produce some artifacts in the process or some other loss of fidelity?!",
    "938370": "A couple of ideas that I have:\n1. mels is a scale for humans but birds sing for birds and do not care about humans. and even for humans there are different scales. For instance, rnnoise lib uses some Bark scale to suppress noise. It is still beyond me to comprehend the implications though\n2. a \"song\" label from the same birds can have completely different sounds and thus spectograms i.e. either train a model on \"bird + type of song\" or time of the year or locaiton (i.e. pump up classes) + augmentations + noise or raw spectograms are a wrong go-to tool\n3. spectograms lose information by doing FFT by construction. Maybe they lose the wrong kind of information, i.e. time dimension can be more important than frequency. Or blending 2 models: 1 time heavy and 1 frequency heavy, might be the trick\n4. some more but I forgot and listing these out of my head :)",
    "938475": "The other factor is the sensitivity of the f1_score to your choice of threshold. Reducing the threshold from 0.8 to 0.6 results in a significant (0.562 -&gt; 0.494) LB score drop using the same model:\n\n![](https://www.googleapis.com/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F36054%2Ff0a98ee24655fbfca1b9ac2b5932443b%2Fscores.png)",
    "939035": "As for 1. , I think mel scale is ok, because annotation comes out of humans.",
    "939043": "Personally I decided to use variable threshold per class based on the result in train set. This cannot remove the sensitivity to threshold, but at least it can automatically decide what threshold to use; and as far as I have observed so far it's working fine (not very good but just fine).",
    "939201": "I published my baseline as **just a baseline**. It is unexpectedly strong and I don't know why🤔\n\n&gt;  I was using the default resnest50 which is slightly larger I believe.\n\nI think so. I suspect lighter models with special training techniques such as MIL(you've suggested) suit this competition.\n\n&gt; The one thing I haven't tried (trying now) is training on resized spectrograms. But why should this work?\n\nI have the same doubt about model input sizes.  \nOne hypothesis is that difference in resolution between pretraining task(ImageNet) and downstream task(this competition) influences models' performance. But I'm not certain about this idea.\n\nI'd like to get your opinion, @hidehisaarai1213",
    "939367": "&gt; lighter models with special training techniques such as MIL(you've suggested) suit this competition.\n\nI agree with this. I think that backbone structure does not make significant difference as long as they are expressive enough, however, the use of pretrained weight may affects. I tried pretrained (with AudioSet) PANNs / untrained PANNs &amp; pretrained (with ImageNet) ResNet / untrained ResNet and found that pretrained model does better, so if we want to try custom architecture as backbone it may be good to train it with ImageNet / AudioSet at first and then train it with the data of this competition.\n\n&gt; resolution between pretraining task(ImageNet) and downstream task(this competition) influences models' performance\n\nPretrained ResNet / ResNest are trained with 224 x 224 images thus kernels of CNN are trained to extract patterns in the image in that size. Maybe that is why resizing the input to 224 works better.",
    "939432": "On one hand I am relieved to finally find the reason for the difference in performance… on the other hand the finding is a little bit disheartening…\n\nTraining with spectrograms upscaled to 224x my model achieves 0.567, where training in the same fashion but on 128x achieved a score of 0.556.\n\nIt is interesting, because the relative improvement is enormous. I am not sure if the change in performance is due to the model being pretrained on 224x224 images. The shapes the model would see on imagenet are very different to what it sees here. I am thinking it might rather be due to the model architecture being fitted to 224x224 examples.\n\nI don't really know yet what this result means - have to ponder on it a little bit.\n\nI think it would be useful to run training on 256x or maybe even bigger sizes? One could use the best weights from 224x as a starting point. I am not sure if Kaggle hw would permit, but would be interesting to see how the starter packs would perform with such an increase in size, or even larger.\n\nRight now I am training with the smaller resnest and will also redo the training with res34. If there were no difference in performance being able to train with smaller models would be useful, could help to iterate faster, can also be useful for applications as inference with smaller models is cheaper / faster.",
    "939531": "The point of resizing is somewhat misplaced. The dimensions of spectogram is the result of your direct choice coming out of the FFT. You can directly produce any pixel dimensions you like by choosing appropriate parameters FFT, i.e. your sound clip length + window length + your sampling rate will give you the exact image dimension on the x side.",
    "939543": "In Audio Tagging 2019 competition I also found that upscaled images worked great (128x128 crop upscale to 256x256). Interestingly it worked better than cropping 128x256 and upscaling to 256x256 and the models were not pretrained.",
    "939974": "Thanks for answering!\n\nThe Relation between pretrained models on images and audio tasks is slightly strange and interesting.",
    "961533": "I quit using variable threshold. Working on threshold optimization is muddy and has no significant difference. Maybe good to use 0.5: I checked some technical reports of DCASE challenges and they usually use threshold of 0.5 when they are not certain about test data.",
    "962579": "Great work Upvoted for you!",
    "965734": "Yes, but in many cases the bird was seen, so the annotation is not _completely_ based off the audio",
    "991343": "> Where lies the magic?\n\nMaybe this notebook is just overfitting to public LB.",
    "994929": "hidehisaarai1213 , thanks for sharing. Are you using a threshold higher than 0.5? I find when using additional labels, such as adding secondary labels, a higher threshold has a better score. This could be due to the nature of public LB, but challenging as it is hard to tell if this threshold is optimal for private LB",
    "995116": "> Are you using a threshold higher than 0.5?\n\nI did, but not now. I've fixed the threshold to 0.5.\n\n> I find when using additional labels, such as adding secondary labels, a higher threshold has a better score.\n\nInteresting observation. I think this comes from the nature of train data. You may label nocall / some call events but not those of species in secondary labels as those of species in secondary labels, because the label is weak. This will make the model produce higher probability on nuisance signals and lead us to get lots of false positives."
  },
  "source": "meta"
}