{
  "id": 220304,
  "title": "11th place, The 0.931 Magic Explained: Image Classification",
  "url": "/competitions/rfcx-species-audio-detection/writeups/cpmp-11th-place-the-0-931-magic-explained-image-cl",
  "author_name": "",
  "post_date": "2021-02-22T12:49:16.677Z",
  "votes": 98,
  "comment_count": 63,
  "views": 0,
  "content": "<p>First of all, thanks to the host, for providing yet another very interesting audio challenge.  The fact that it was only partly labelled was a significant difference with previous bird song competition.</p>\n<p>Second, congrats to all those who managed to pass the 0.95 bar on public LB.  I couldn't, and, as I write this before competition deadline, I don't know why.  I tried a lot of things, and it looks like my approach has some fundamental limit.</p>\n<p>Yet it was good enough to produce a first submission at 0.931, placing directly at 3rd spot while the competition had started two months earlier.</p>\n<p>This looked great to me.  In hindsight, if my first sub had been weaker, then I would not have stick to its model and would have explored other models probably, like SED or Transformers.</p>\n<p>Anyway, no need to complain, I learned a lot of stuff along the way, like how to efficiently implement teacher student training of all sorts.  I hope this knowledge will be useful in the future.</p>\n<p>Back to the topic, my approach was extremely simple:  each row of train data, TP or FP, gives us a label for a crop in the (log mel) spectrogram of the corresponding recording.  If time is x axis and frequency the y axis, as is generally the case, then t_min, t_max gives bounds on x axis, and f_min, f_max gives bounds on the y axis.</p>\n<p>We then have a 26 multi label classification problem (24 species but two species have 2 song types.  I treated each species/song type as a different class).  This is easily handled with BCE Loss.</p>\n<p>The only little caveat is that we are given 26 classes (species + song type) but we get only one class label, 0 or 1 per image.  We only have to mask the loss for other classes and that's it!</p>\n<p>I didn't know it when I did it, but a similar way has been used by some of the host of the competition, in this paper (not the one shared in the forum): <a href=\"https://www.sciencedirect.com/science/article/abs/pii/S0003682X20304795\" target=\"_blank\">https://www.sciencedirect.com/science/article/abs/pii/S0003682X20304795</a></p>\n<p>The other caveat is that the competition metric works with a label for every class.  Which we don't have in train data.  However the competition metric is very similar to a roc auc score per recording: when a pair of predictions is in the wrong order, i.e. a positive label has a prediction lower than another negative label prediction, then the metric is lowered.  As a proxy I decided to use roc-auc on my multi label classification problem.  Correlation with public LB is noisy, but it was good enough to let me make progress without submitting for a while.</p>\n<p>What worked best for me was to not resize the crops.  It means my model had to learn from sometimes tiny images.  To make it work by batch I pad all images to 4 seconds on the x axis.  Crops longer than that were resized on the x axis.  Shorter ones were padded with 0. One thing that helped was to add a positional encoding on the frequency axis.  Indeed, CNNs are good at learning translation independent representations, and here we don't want the model to be frequency independent.  I simply added a linear gradient on the frequency axis to all my crops.</p>\n<p>For the rest my model is exactly what I used and shared in the previous bird song competition: <a href=\"https://www.kaggle.com/c/birdsong-recognition/discussion/183219\" target=\"_blank\">https://www.kaggle.com/c/birdsong-recognition/discussion/183219</a>  Just using the code I shared there was almost good enough to get 0.931.  The only differences are that I add noise as in the first solution in that competition.  I also did not use a no call class here, nor secondary labels.</p>\n<p>For prediction I predict on slidding crops of  each test recording and take the overall max prediction.  This is slow as I need to do each of the 26 classes separately.  This is also maybe where I lost against others: my model cannot learn long range temporal patterns, nor class interactions.</p>\n<p>With the above I entered high with an Efficient B0 model, and moved to 0.945 in few submissions with Efficientnet B3 . Then I got stuck for the remainder of the competition.  </p>\n<p>I was convinced that semi supervised learning was the key, and I implemented all sorts of methods, from Google (noisy student), Facebook, others (mean student).  They all improved weaker models but could not improve my best models.</p>\n<p>In the last days I looked for external data with the hope that it would make a difference. Curating all this and identifying which species correspond to the species_id we have took some time and I only submitted models trained with it today.  They are in same range as previous ones unfortunately.  with a bit more time I am sure it could improve score, but I doubt it would be significant..</p>\n<p>For matching species to species id I used my best model and predicted the external data. It would be interesting to see if I got this mapping right.  Here is what I converged to  :</p>\n<p>0                        E. gryllus<br>\n1         Leutherodactylus brittoni<br>\n2           Leptodactylu albilabris<br>\n3                          E. coqui<br>\n4                       E. hedricki<br>\n5                 Setophaga angelae<br>\n6          Melanerpes portoricensis<br>\n7                  Coereba flaveola<br>\n8                       E. locustus<br>\n9                Margarops fuscatus<br>\n10          Loxigilla portoricensis<br>\n11                 Vireo altiloquus<br>\n12                 E. portoricensis<br>\n13                Megascops nudipes<br>\n14                     E. richmondi<br>\n15             Patagioenas squamosa<br>\n16    Eleutherodactylus antillensis<br>\n17                  Turdus plumbeus<br>\n18                      E. unicolor<br>\n19               Coccyzus vieilloti<br>\n20                  Todus mexicanus<br>\n21                    E  wightmanae<br>\n22         Nesospingus speculiferus<br>\n23          Spindalis portoricensis</p>\n<p>The picture in the paper shared in the forum  helped to disambiguate few cases: <a href=\"https://reader.elsevier.com/reader/sd/pii/S1574954120300637\" target=\"_blank\">https://reader.elsevier.com/reader/sd/pii/S1574954120300637</a>  The paper also gives the list of species.  My final selected subs did not include models trained on external data, given they were not improving.</p>\n<p>This concludes my experience in this competition. I am looking forward to see how so many teams passed me during the competition.  There is certainly a lot to be learned.</p>\n<p>Edit.  I am very pleased to get a solo gold in a deep learning competition.,  This is a first for me, and it was my goal here.</p>\n<p>Edit 2:  The models I trained last day with external data are actually better than the ones without. The best one has a private LB of 0.950 (5 folds). However, they are way better on private LB but not on public LB.  Selecting them would have been an act of faith.  And late submission show they are not good enough to change my rank.  No regrets then.</p>\n<p>Edit 3  Using Chris Deotte post processing. my best selected sub gets 0.7390 on private LB.  It means that PP was what I missed and that my modeling approach was good enough probably.  I'll definitely look at test prediction distribution from now on!</p>",
  "messages": [
    {
      "id": "1207589",
      "postDate": "02/18/2021 00:00:47",
      "content": "<p>First of all, thanks to the host, for providing yet another very interesting audio challenge.  The fact that it was only partly labelled was a significant difference with previous bird song competition.</p>\n<p>Second, congrats to all those who managed to pass the 0.95 bar on public LB.  I couldn't, and, as I write this before competition deadline, I don't know why.  I tried a lot of things, and it looks like my approach has some fundamental limit.</p>\n<p>Yet it was good enough to produce a first submission at 0.931, placing directly at 3rd spot while the competition had started two months earlier.</p>\n<p>This looked great to me.  In hindsight, if my first sub had been weaker, then I would not have stick to its model and would have explored other models probably, like SED or Transformers.</p>\n<p>Anyway, no need to complain, I learned a lot of stuff along the way, like how to efficiently implement teacher student training of all sorts.  I hope this knowledge will be useful in the future.</p>\n<p>Back to the topic, my approach was extremely simple:  each row of train data, TP or FP, gives us a label for a crop in the (log mel) spectrogram of the corresponding recording.  If time is x axis and frequency the y axis, as is generally the case, then t_min, t_max gives bounds on x axis, and f_min, f_max gives bounds on the y axis.</p>\n<p>We then have a 26 multi label classification problem (24 species but two species have 2 song types.  I treated each species/song type as a different class).  This is easily handled with BCE Loss.</p>\n<p>The only little caveat is that we are given 26 classes (species + song type) but we get only one class label, 0 or 1 per image.  We only have to mask the loss for other classes and that's it!</p>\n<p>I didn't know it when I did it, but a similar way has been used by some of the host of the competition, in this paper (not the one shared in the forum): <a href=\"https://www.sciencedirect.com/science/article/abs/pii/S0003682X20304795\" target=\"_blank\">https://www.sciencedirect.com/science/article/abs/pii/S0003682X20304795</a></p>\n<p>The other caveat is that the competition metric works with a label for every class.  Which we don't have in train data.  However the competition metric is very similar to a roc auc score per recording: when a pair of predictions is in the wrong order, i.e. a positive label has a prediction lower than another negative label prediction, then the metric is lowered.  As a proxy I decided to use roc-auc on my multi label classification problem.  Correlation with public LB is noisy, but it was good enough to let me make progress without submitting for a while.</p>\n<p>What worked best for me was to not resize the crops.  It means my model had to learn from sometimes tiny images.  To make it work by batch I pad all images to 4 seconds on the x axis.  Crops longer than that were resized on the x axis.  Shorter ones were padded with 0. One thing that helped was to add a positional encoding on the frequency axis.  Indeed, CNNs are good at learning translation independent representations, and here we don't want the model to be frequency independent.  I simply added a linear gradient on the frequency axis to all my crops.</p>\n<p>For the rest my model is exactly what I used and shared in the previous bird song competition: <a href=\"https://www.kaggle.com/c/birdsong-recognition/discussion/183219\" target=\"_blank\">https://www.kaggle.com/c/birdsong-recognition/discussion/183219</a>  Just using the code I shared there was almost good enough to get 0.931.  The only differences are that I add noise as in the first solution in that competition.  I also did not use a no call class here, nor secondary labels.</p>\n<p>For prediction I predict on slidding crops of  each test recording and take the overall max prediction.  This is slow as I need to do each of the 26 classes separately.  This is also maybe where I lost against others: my model cannot learn long range temporal patterns, nor class interactions.</p>\n<p>With the above I entered high with an Efficient B0 model, and moved to 0.945 in few submissions with Efficientnet B3 . Then I got stuck for the remainder of the competition.  </p>\n<p>I was convinced that semi supervised learning was the key, and I implemented all sorts of methods, from Google (noisy student), Facebook, others (mean student).  They all improved weaker models but could not improve my best models.</p>\n<p>In the last days I looked for external data with the hope that it would make a difference. Curating all this and identifying which species correspond to the species_id we have took some time and I only submitted models trained with it today.  They are in same range as previous ones unfortunately.  with a bit more time I am sure it could improve score, but I doubt it would be significant..</p>\n<p>For matching species to species id I used my best model and predicted the external data. It would be interesting to see if I got this mapping right.  Here is what I converged to  :</p>\n<p>0                        E. gryllus<br>\n1         Leutherodactylus brittoni<br>\n2           Leptodactylu albilabris<br>\n3                          E. coqui<br>\n4                       E. hedricki<br>\n5                 Setophaga angelae<br>\n6          Melanerpes portoricensis<br>\n7                  Coereba flaveola<br>\n8                       E. locustus<br>\n9                Margarops fuscatus<br>\n10          Loxigilla portoricensis<br>\n11                 Vireo altiloquus<br>\n12                 E. portoricensis<br>\n13                Megascops nudipes<br>\n14                     E. richmondi<br>\n15             Patagioenas squamosa<br>\n16    Eleutherodactylus antillensis<br>\n17                  Turdus plumbeus<br>\n18                      E. unicolor<br>\n19               Coccyzus vieilloti<br>\n20                  Todus mexicanus<br>\n21                    E  wightmanae<br>\n22         Nesospingus speculiferus<br>\n23          Spindalis portoricensis</p>\n<p>The picture in the paper shared in the forum  helped to disambiguate few cases: <a href=\"https://reader.elsevier.com/reader/sd/pii/S1574954120300637\" target=\"_blank\">https://reader.elsevier.com/reader/sd/pii/S1574954120300637</a>  The paper also gives the list of species.  My final selected subs did not include models trained on external data, given they were not improving.</p>\n<p>This concludes my experience in this competition. I am looking forward to see how so many teams passed me during the competition.  There is certainly a lot to be learned.</p>\n<p>Edit.  I am very pleased to get a solo gold in a deep learning competition.,  This is a first for me, and it was my goal here.</p>\n<p>Edit 2:  The models I trained last day with external data are actually better than the ones without. The best one has a private LB of 0.950 (5 folds). However, they are way better on private LB but not on public LB.  Selecting them would have been an act of faith.  And late submission show they are not good enough to change my rank.  No regrets then.</p>\n<p>Edit 3  Using Chris Deotte post processing. my best selected sub gets 0.7390 on private LB.  It means that PP was what I missed and that my modeling approach was good enough probably.  I'll definitely look at test prediction distribution from now on!</p>",
      "rawMarkdown": "First of all, thanks to the host, for providing yet another very interesting audio challenge.  The fact that it was only partly labelled was a significant difference with previous bird song competition.\n\nSecond, congrats to all those who managed to pass the 0.95 bar on public LB.  I couldn't, and, as I write this before competition deadline, I don't know why.  I tried a lot of things, and it looks like my approach has some fundamental limit.\n\nYet it was good enough to produce a first submission at 0.931, placing directly at 3rd spot while the competition had started two months earlier.\n\nThis looked great to me.  In hindsight, if my first sub had been weaker, then I would not have stick to its model and would have explored other models probably, like SED or Transformers.\n\nAnyway, no need to complain, I learned a lot of stuff along the way, like how to efficiently implement teacher student training of all sorts.  I hope this knowledge will be useful in the future.\n\nBack to the topic, my approach was extremely simple:  each row of train data, TP or FP, gives us a label for a crop in the (log mel) spectrogram of the corresponding recording.  If time is x axis and frequency the y axis, as is generally the case, then t_min, t_max gives bounds on x axis, and f_min, f_max gives bounds on the y axis.\n\nWe then have a 26 multi label classification problem (24 species but two species have 2 song types.  I treated each species/song type as a different class).  This is easily handled with BCE Loss.\n\nThe only little caveat is that we are given 26 classes (species + song type) but we get only one class label, 0 or 1 per image.  We only have to mask the loss for other classes and that's it!\n\nI didn't know it when I did it, but a similar way has been used by some of the host of the competition, in this paper (not the one shared in the forum): https://www.sciencedirect.com/science/article/abs/pii/S0003682X20304795\n\nThe other caveat is that the competition metric works with a label for every class.  Which we don't have in train data.  However the competition metric is very similar to a roc auc score per recording: when a pair of predictions is in the wrong order, i.e. a positive label has a prediction lower than another negative label prediction, then the metric is lowered.  As a proxy I decided to use roc-auc on my multi label classification problem.  Correlation with public LB is noisy, but it was good enough to let me make progress without submitting for a while.\n\nWhat worked best for me was to not resize the crops.  It means my model had to learn from sometimes tiny images.  To make it work by batch I pad all images to 4 seconds on the x axis.  Crops longer than that were resized on the x axis.  Shorter ones were padded with 0. One thing that helped was to add a positional encoding on the frequency axis.  Indeed, CNNs are good at learning translation independent representations, and here we don't want the model to be frequency independent.  I simply added a linear gradient on the frequency axis to all my crops.\n\nFor the rest my model is exactly what I used and shared in the previous bird song competition: https://www.kaggle.com/c/birdsong-recognition/discussion/183219  Just using the code I shared there was almost good enough to get 0.931.  The only differences are that I add noise as in the first solution in that competition.  I also did not use a no call class here, nor secondary labels.\n\nFor prediction I predict on slidding crops of  each test recording and take the overall max prediction.  This is slow as I need to do each of the 26 classes separately.  This is also maybe where I lost against others: my model cannot learn long range temporal patterns, nor class interactions.\n\nWith the above I entered high with an Efficient B0 model, and moved to 0.945 in few submissions with Efficientnet B3 . Then I got stuck for the remainder of the competition.  \n\nI was convinced that semi supervised learning was the key, and I implemented all sorts of methods, from Google (noisy student), Facebook, others (mean student).  They all improved weaker models but could not improve my best models.\n\nIn the last days I looked for external data with the hope that it would make a difference. Curating all this and identifying which species correspond to the species_id we have took some time and I only submitted models trained with it today.  They are in same range as previous ones unfortunately.  with a bit more time I am sure it could improve score, but I doubt it would be significant..\n\nFor matching species to species id I used my best model and predicted the external data. It would be interesting to see if I got this mapping right.  Here is what I converged to  :\n\n0                        E. gryllus\n1         Leutherodactylus brittoni\n2           Leptodactylu albilabris\n3                          E. coqui\n4                       E. hedricki\n5                 Setophaga angelae\n6          Melanerpes portoricensis\n7                  Coereba flaveola\n8                       E. locustus\n9                Margarops fuscatus\n10          Loxigilla portoricensis\n11                 Vireo altiloquus\n12                 E. portoricensis\n13                Megascops nudipes\n14                     E. richmondi\n15             Patagioenas squamosa\n16    Eleutherodactylus antillensis\n17                  Turdus plumbeus\n18                      E. unicolor\n19               Coccyzus vieilloti\n20                  Todus mexicanus\n21                    E  wightmanae\n22         Nesospingus speculiferus\n23          Spindalis portoricensis\n\nThe picture in the paper shared in the forum  helped to disambiguate few cases: https://reader.elsevier.com/reader/sd/pii/S1574954120300637  The paper also gives the list of species.  My final selected subs did not include models trained on external data, given they were not improving.\n\nThis concludes my experience in this competition. I am looking forward to see how so many teams passed me during the competition.  There is certainly a lot to be learned.\n\nEdit.  I am very pleased to get a solo gold in a deep learning competition.,  This is a first for me, and it was my goal here.\n\nEdit 2:  The models I trained last day with external data are actually better than the ones without. The best one has a private LB of 0.950 (5 folds). However, they are way better on private LB but not on public LB.  Selecting them would have been an act of faith.  And late submission show they are not good enough to change my rank.  No regrets then.\n\nEdit 3  Using Chris Deotte post processing. my best selected sub gets 0.7390 on private LB.  It means that PP was what I missed and that my modeling approach was good enough probably.  I'll definitely look at test prediction distribution from now on!",
      "votes": null
    },
    {
      "id": "1207601",
      "postDate": "02/18/2021 00:10:03",
      "content": "<p>Very nice writeup. We used a very similar approach, which we really kicked into action at the last 8 days; will share details tomorrow… congrats on the solo medal!</p>",
      "rawMarkdown": "Very nice writeup. We used a very similar approach, which we really kicked into action at the last 8 days; will share details tomorrow... congrats on the solo medal!",
      "votes": null
    },
    {
      "id": "1207603",
      "postDate": "02/18/2021 00:10:29",
      "content": "<p>Congrats <a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a> . Great solo finish! I'm glad you won Gold. You helped many teams by showing them what was possible!</p>",
      "rawMarkdown": "Congrats @cpmpml . Great solo finish! I'm glad you won Gold. You helped many teams by showing them what was possible!",
      "votes": null
    },
    {
      "id": "1207608",
      "postDate": "02/18/2021 00:14:45",
      "content": "<p>Great work. So many techniques that are close to what I tried, but just never put them together in the right way. </p>\n<p>It seemed like the public pannloss that was floating around was trying to implement that masked loss like was mentioned in the paper, but was kind of implemented wrong. </p>",
      "rawMarkdown": "Great work. So many techniques that are close to what I tried, but just never put them together in the right way. \n\nIt seemed like the public pannloss that was floating around was trying to implement that masked loss like was mentioned in the paper, but was kind of implemented wrong.",
      "votes": null
    },
    {
      "id": "1207612",
      "postDate": "02/18/2021 00:16:37",
      "content": "<p>Congrats on the solo gold! I have one question: did you crop y-axis (the frequency axis) too when doing training and inference?</p>",
      "rawMarkdown": "Congrats on the solo gold! I have one question: did you crop y-axis (the frequency axis) too when doing training and inference?",
      "votes": null
    },
    {
      "id": "1207613",
      "postDate": "02/18/2021 00:16:52",
      "content": "<p>tx.  Not only did I show it, but I shared I was using an image classification model.  Next time I'll share less maybe ;)</p>\n<p>Congrats on your solo gold too.</p>",
      "rawMarkdown": "tx.  Not only did I show it, but I shared I was using an image classification model.  Next time I'll share less maybe ;)\n\nCongrats on your solo gold too.",
      "votes": null
    },
    {
      "id": "1207614",
      "postDate": "02/18/2021 00:17:04",
      "content": "<p>Congratulations! Could you share how exactly did you do the loss masking? I have tried this early in competition (and actually specifically inspired by your approach to secondary labels in BirdCall), but it backfired quite heavily to me (I only did time based crops, frequency cropping may be even more important, but intuitively I thought even without frequency cropping loss masking should make more sense…)</p>",
      "rawMarkdown": "Congratulations! Could you share how exactly did you do the loss masking? I have tried this early in competition (and actually specifically inspired by your approach to secondary labels in BirdCall), but it backfired quite heavily to me (I only did time based crops, frequency cropping may be even more important, but intuitively I thought even without frequency cropping loss masking should make more sense...)",
      "votes": null
    },
    {
      "id": "1207615",
      "postDate": "02/18/2021 00:18:07",
      "content": "<p>Tx.  My limited skills in computer vision may be what prevented me from getting a higher score.  Looking forward to your writeup.  I am sure I'll learn stuff.  And congrats as well on your result.</p>",
      "rawMarkdown": "Tx.  My limited skills in computer vision may be what prevented me from getting a higher score.  Looking forward to your writeup.  I am sure I'll learn stuff.  And congrats as well on your result.",
      "votes": null
    },
    {
      "id": "1207616",
      "postDate": "02/18/2021 00:19:09",
      "content": "<p>Tx.  I didn't look at public notebooks at all.  Maybe I should have…  It also looks like others used the same overall method, but got better results than me.  </p>",
      "rawMarkdown": "Tx.  I didn't look at public notebooks at all.  Maybe I should have...  It also looks like others used the same overall method, but got better results than me.",
      "votes": null
    },
    {
      "id": "1207617",
      "postDate": "02/18/2021 00:19:27",
      "content": "<p>Yes.  I cropped by f_min and f_max.</p>",
      "rawMarkdown": "Yes.  I cropped by f_min and f_max.",
      "votes": null
    },
    {
      "id": "1207619",
      "postDate": "02/18/2021 00:20:15",
      "content": "<p>Congrats, amazing job. We were indeed very impressed by the 931 start. I dont think you revealed anything with the image model though, so dont worry about that. Everyone uses image models in these types of problems.</p>",
      "rawMarkdown": "Congrats, amazing job. We were indeed very impressed by the 931 start. I dont think you revealed anything with the image model though, so dont worry about that. Everyone uses image models in these types of problems.",
      "votes": null
    },
    {
      "id": "1207622",
      "postDate": "02/18/2021 00:21:58",
      "content": "<p><a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a>   <br>\n\"is slow as I need to do each of the 26 classes separately\"<br>\n\"Yes. I cropped by f_min and f_max.\"<br>\n…</p>\n<p>see my solution to solve that<br>\n<a href=\"https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/220309\" target=\"_blank\">https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/220309</a></p>",
      "rawMarkdown": "cpmpml   \n\"is slow as I need to do each of the 26 classes separately\"\n\"Yes. I cropped by f_min and f_max.\"\n...\n\nsee my solution to solve that\nhttps://www.kaggle.com/c/rfcx-species-audio-detection/discussion/220309",
      "votes": null
    },
    {
      "id": "1207623",
      "postDate": "02/18/2021 00:22:46",
      "content": "<p>Did the extra class for the songs help anything? I had it on the list since day 1 but never got around to try it. I think it was for low-frequency classes but would need to re-check.</p>",
      "rawMarkdown": "Did the extra class for the songs help anything? I had it on the list since day 1 but never got around to try it. I think it was for low-frequency classes but would need to re-check.",
      "votes": null
    },
    {
      "id": "1207627",
      "postDate": "02/18/2021 00:26:25",
      "content": "<p>I am not sure about what is not clear.  You compute the BCE loss for all targets (no reduction), then multiply by the mask (1 for the song class to be predicted, 0 for all other classes), then take the mean for the batch.</p>",
      "rawMarkdown": "I am not sure about what is not clear.  You compute the BCE loss for all targets (no reduction), then multiply by the mask (1 for the song class to be predicted, 0 for all other classes), then take the mean for the batch.",
      "votes": null
    },
    {
      "id": "1207628",
      "postDate": "02/18/2021 00:26:45",
      "content": "<p>Congrats on your solo gold <a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a> . Your entrance to the competition was inspirational 💪</p>",
      "rawMarkdown": "Congrats on your solo gold @cpmpml . Your entrance to the competition was inspirational 💪",
      "votes": null
    },
    {
      "id": "1207630",
      "postDate": "02/18/2021 00:27:15",
      "content": "<p>Sorry, I did not use it here.  My bad.  Let me fix the writeup.</p>",
      "rawMarkdown": "Sorry, I did not use it here.  My bad.  Let me fix the writeup.",
      "votes": null
    },
    {
      "id": "1207637",
      "postDate": "02/18/2021 00:34:05",
      "content": "<p>Congrats! The gradient idea is nice. I have also tried applying fixed grid lines on the images. That helped.</p>",
      "rawMarkdown": "Congrats! The gradient idea is nice. I have also tried applying fixed grid lines on the images. That helped.",
      "votes": null
    },
    {
      "id": "1207650",
      "postDate": "02/18/2021 00:41:35",
      "content": "<p>Congratulations <a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a> for strong solo finish! 0.931 for the first sub was quite impressive! We learned a lot from you, thank you!</p>",
      "rawMarkdown": "Congratulations @cpmpml for strong solo finish! 0.931 for the first sub was quite impressive! We learned a lot from you, thank you!",
      "votes": null
    },
    {
      "id": "1207651",
      "postDate": "02/18/2021 00:41:39",
      "content": "<p>Exactly, sounds quite straightforward to me, so just checking whether I am not missing anything. So something like below, right?</p>\n<pre><code>class RainforestLossMasked(nn.Module):\n    def __init__(self):\n        super().__init__()    \n        self.bce_prob  = nn.BCELoss(reduction='none')\n\n    def forward(self, out, target):                \n        loss = self.bce_prob(out, target)     \n        with torch.no_grad():\n            loss[target&lt;1] = 0   \n        loss = loss.mean()\n        return loss\n</code></pre>",
      "rawMarkdown": "Exactly, sounds quite straightforward to me, so just checking whether I am not missing anything. So something like below, right?\n\n```\nclass RainforestLossMasked(nn.Module):\n    def __init__(self):\n        super().__init__()    \n        self.bce_prob  = nn.BCELoss(reduction='none')\n\n    def forward(self, out, target):                \n        loss = self.bce_prob(out, target)     \n        with torch.no_grad():\n            loss[target<1] = 0   \n        loss = loss.mean()\n        return loss\n```",
      "votes": null
    },
    {
      "id": "1207662",
      "postDate": "02/18/2021 00:49:35",
      "content": "<p>Thanks.  You finished strongly too.  I am sure that if your studies had left you more time then you would have passed me.  Also, I guess that most teams reused your SED model from last competition.  This is also something to be proud of.</p>",
      "rawMarkdown": "Thanks.  You finished strongly too.  I am sure that if your studies had left you more time then you would have passed me.  Also, I guess that most teams reused your SED model from last competition.  This is also something to be proud of.",
      "votes": null
    },
    {
      "id": "1207664",
      "postDate": "02/18/2021 00:50:25",
      "content": "<p>Grid may be more effective actually, I'll try it next time.  Congrats on your strong silver.</p>",
      "rawMarkdown": "Grid may be more effective actually, I'll try it next time.  Congrats on your strong silver.",
      "votes": null
    },
    {
      "id": "1207672",
      "postDate": "02/18/2021 00:55:37",
      "content": "<p>Congrats on 12th place and gold medal <a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a> </p>",
      "rawMarkdown": "Congrats on 12th place and gold medal @cpmpml",
      "votes": null
    },
    {
      "id": "1207683",
      "postDate": "02/18/2021 01:05:33",
      "content": "<p>Congrats! thank you for making this competition so much more competitive :) also, I wanted to ask how much did masking the loss and using a 26 classes model instead of 24 classes help independently?</p>",
      "rawMarkdown": "Congrats! thank you for making this competition so much more competitive :) also, I wanted to ask how much did masking the loss and using a 26 classes model instead of 24 classes help independently?",
      "votes": null
    },
    {
      "id": "1207686",
      "postDate": "02/18/2021 01:06:32",
      "content": "<p>How much Position Information Do Convolutional Neural Networks Encode?<br>\n<a href=\"https://openreview.net/pdf?id=rJeB36NKvB\" target=\"_blank\">https://openreview.net/pdf?id=rJeB36NKvB</a></p>\n<p>It turns out that CNN knows position of image. Because of effects of padding, CNN knows where the  input window is in the image, if the effective receptive field is large (e.g. same size of larger than the size of the input image). Hence adding position gradient is sometimes not necessary</p>\n<p>quote:<br>\n\"Experiments reveal that positional information is available to a strong degree\"</p>\n<p>\"Results point to zero padding and borders as an anchor from which spatial information is derived and eventually propagated over the whole image as spatial abstraction occurs\"</p>\n<hr>\n<p>I think there are also some interesting papers that experiment positional encoding used in transformer for CNN</p>",
      "rawMarkdown": "How much Position Information Do Convolutional Neural Networks Encode?\nhttps://openreview.net/pdf?id=rJeB36NKvB\n\nIt turns out that CNN knows position of image. Because of effects of padding, CNN knows where the  input window is in the image, if the effective receptive field is large (e.g. same size of larger than the size of the input image). Hence adding position gradient is sometimes not necessary\n\nquote:\n\"Experiments reveal that positional information is available to a strong degree\"\n\n\"Results point to zero padding and borders as an anchor from which spatial information is derived and eventually propagated over the whole image as spatial abstraction occurs\"\n\n---\n\nI think there are also some interesting papers that experiment positional encoding used in transformer for CNN",
      "votes": null
    },
    {
      "id": "1207696",
      "postDate": "02/18/2021 01:20:03",
      "content": "<p>I only tried 24 classes omce, and it was a bit worse.</p>\n<p>Congrats on your strong finish.  Looking forward to your writeup.  </p>\n<p>I wish I had waited two more week before submitting ;)</p>\n<p>Just kidding, I am sure others would have submitted high scores soon.</p>",
      "rawMarkdown": "I only tried 24 classes omce, and it was a bit worse.\n\nCongrats on your strong finish.  Looking forward to your writeup.  \n\nI wish I had waited two more week before submitting ;)\n\nJust kidding, I am sure others would have submitted high scores soon.",
      "votes": null
    },
    {
      "id": "1207697",
      "postDate": "02/18/2021 01:22:38",
      "content": "<p>something like this:</p>\n<pre><code>                def use_tp_fp_binary_cross_entropy(logit, label, is_tp):\n                    batch_size, num_label = logit.shape\n                    batch_size = label.shape # e.g. [ 2,2,6,8, ... 23]\n                    batch_size = is_tp.shape # e.g. [ 1,0,0,1, ... 1]\n\n                    l = logit.gather(1, label.unsqueeze(1))\n                    l = l.reshape(-1)\n                    t = is_tp.float()\n\n                    p = torch.sigmoid(l)\n                    logp = - torch.log(torch.clamp(p, 1e-4, 1-1e-4))\n                    logn = - torch.log(torch.clamp(1-p, 1e-4, 1-1e-4))\n                    loss = t*logp +(1-t)*logn\n\n                    loss  = loss.mean()\n                    return loss\n</code></pre>",
      "rawMarkdown": "something like this:\n\n```\n                def use_tp_fp_binary_cross_entropy(logit, label, is_tp):\n                    batch_size, num_label = logit.shape\n                    batch_size = label.shape # e.g. [ 2,2,6,8, ... 23]\n                    batch_size = is_tp.shape # e.g. [ 1,0,0,1, ... 1]\n                 \n                    l = logit.gather(1, label.unsqueeze(1))\n                    l = l.reshape(-1)\n                    t = is_tp.float()\n\n                    p = torch.sigmoid(l)\n                    logp = - torch.log(torch.clamp(p, 1e-4, 1-1e-4))\n                    logn = - torch.log(torch.clamp(1-p, 1e-4, 1-1e-4))\n                    loss = t*logp +(1-t)*logn\n\n                    loss  = loss.mean()\n                    return loss\n\n\n\n```",
      "votes": null
    },
    {
      "id": "1207700",
      "postDate": "02/18/2021 01:26:49",
      "content": "<p>You mentioned focal loss earlier in the competition. Did you ever try that in place of bce in your final solution? In my experience that has only ever been useful for extreme class imbalance problems like segmentation where there is a small number of positives to a large number of negatives. </p>\n<p>I feel like this is another one of the competitions where I had most of the ingredients for the top solutions out but couldn't quite find the right recipe to combine them all in. I tried the masked loss, frequency banded model that I implemented a bit differently than yours, but similar concept. I mentioned to our team at one point the linear gradient to show position more explicitly to the cnn. </p>",
      "rawMarkdown": "You mentioned focal loss earlier in the competition. Did you ever try that in place of bce in your final solution? In my experience that has only ever been useful for extreme class imbalance problems like segmentation where there is a small number of positives to a large number of negatives. \n\nI feel like this is another one of the competitions where I had most of the ingredients for the top solutions out but couldn't quite find the right recipe to combine them all in. I tried the masked loss, frequency banded model that I implemented a bit differently than yours, but similar concept. I mentioned to our team at one point the linear gradient to show position more explicitly to the cnn.",
      "votes": null
    },
    {
      "id": "1207701",
      "postDate": "02/18/2021 01:27:45",
      "content": "<p>Final question: did you validate against the crops with your image model or did you apply the same test time rolling validation? I'd assume you validated against the crops since the rolling window inference was slow. </p>",
      "rawMarkdown": "Final question: did you validate against the crops with your image model or did you apply the same test time rolling validation? I'd assume you validated against the crops since the rolling window inference was slow.",
      "votes": null
    },
    {
      "id": "1207702",
      "postDate": "02/18/2021 01:29:00",
      "content": "<p>Good question!</p>\n<p>Yes I tried it at some point and it was a bit worse than vanilla bce.</p>\n<p>Focal loss only worked for me in segmentation tasks.  This is the original use case of focal loss unless mistaken.</p>",
      "rawMarkdown": "Good question!\n\nYes I tried it at some point and it was a bit worse than vanilla bce.\n\nFocal loss only worked for me in segmentation tasks.  This is the original use case of focal loss unless mistaken.",
      "votes": null
    },
    {
      "id": "1207704",
      "postDate": "02/18/2021 01:31:06",
      "content": "<p>I used 5 fold CV on my crops.  And I used roc-auc.</p>",
      "rawMarkdown": "I used 5 fold CV on my crops.  And I used roc-auc.",
      "votes": null
    },
    {
      "id": "1207708",
      "postDate": "02/18/2021 01:34:45",
      "content": "<p>Thank you! will post a writeup soon. I am sure quite a few entered after looking at your 0.931 post and also many that already entered started to look for alternative solutions too haha!</p>\n<p>how about masking the loss, how much did it help?</p>",
      "rawMarkdown": "Thank you! will post a writeup soon. I am sure quite a few entered after looking at your 0.931 post and also many that already entered started to look for alternative solutions too haha!\n\nhow about masking the loss, how much did it help?",
      "votes": null
    },
    {
      "id": "1207715",
      "postDate": "02/18/2021 01:44:36",
      "content": "<p>Congrats on your solo gold. Your discussion comments helped me many times.</p>",
      "rawMarkdown": "Congrats on your solo gold. Your discussion comments helped me many times.",
      "votes": null
    },
    {
      "id": "1207728",
      "postDate": "02/18/2021 02:12:12",
      "content": "<p>Kind of kicking myself for not exploring this further. The first thing I sent to the team when I joined them was showing them a model I had that was kind of similar to yours. Instead of cropping though I simply masked out the regions outside of the frequency band and time that were irrelevant. So I would crop around the region of the time and then I would mask out the frequencies that were irrelevant to that prediction. </p>\n<p><img src=\"https://i.imgur.com/tJPa7Sv.png\" alt=\"\"></p>\n<p>Would look something like this. I made it so they were long enough that I never had to do any cropping, only ever padding near the beginning or end of the audio. Would use similar procedure to you at test time, but I could stack all of them together fairly easily so it was 16(num_time_steps)*24(num_frequency_ranges), 128(num_mel_bins), 500 (time) and it was fairly fast to do inference. On the crops themselves I got validation performance of .985 at times, but when I applied the rolling window validation I would get poor results like .6-.7. </p>\n<p>I tried the masked loss, but I did not apply it to this specific model and I never made a submission with it because the windowed validation looked poor. I considered the linear gradient to give position but ended up not using it in this setup because I figured the model already had its relative position based on the amount of 0 padding above and below the unmasked signal. Might have to fiddle with that and see if it was actually good. </p>",
      "rawMarkdown": "Kind of kicking myself for not exploring this further. The first thing I sent to the team when I joined them was showing them a model I had that was kind of similar to yours. Instead of cropping though I simply masked out the regions outside of the frequency band and time that were irrelevant. So I would crop around the region of the time and then I would mask out the frequencies that were irrelevant to that prediction. \n\n![](https://i.imgur.com/tJPa7Sv.png)\n\nWould look something like this. I made it so they were long enough that I never had to do any cropping, only ever padding near the beginning or end of the audio. Would use similar procedure to you at test time, but I could stack all of them together fairly easily so it was 16(num_time_steps)*24(num_frequency_ranges), 128(num_mel_bins), 500 (time) and it was fairly fast to do inference. On the crops themselves I got validation performance of .985 at times, but when I applied the rolling window validation I would get poor results like .6-.7. \n\nI tried the masked loss, but I did not apply it to this specific model and I never made a submission with it because the windowed validation looked poor. I considered the linear gradient to give position but ended up not using it in this setup because I figured the model already had its relative position based on the amount of 0 padding above and below the unmasked signal. Might have to fiddle with that and see if it was actually good.",
      "votes": null
    },
    {
      "id": "1207730",
      "postDate": "02/18/2021 02:14:10",
      "content": "<p>Thank you for the great write-up and congratz on solo gold!<br>\nI thought <code>Coereba flaveola</code> was <code>species_id=1</code> but I was not sure.</p>",
      "rawMarkdown": "Thank you for the great write-up and congratz on solo gold!\nI thought `Coereba flaveola` was `species_id=1` but I was not sure.",
      "votes": null
    },
    {
      "id": "1207774",
      "postDate": "02/18/2021 03:13:23",
      "content": "<p>Congrats on your medal. As always, thanks for sharing  <br>\nSorry for my lack of understanding, when you mention</p>\n<blockquote>\n  <p>For prediction I predict on slidding crops of each test recording and take the overall max prediction. This is slow as I need to do each of the 26 classes separately. This is also maybe where I lost against others: my model cannot learn long range temporal patterns, nor class interactions.</p>\n</blockquote>\n<p>Does this mean u train 26 models for each class? From what i am reading, if you are masking your bce, you are able to output multi label predictions for each sample so where do the 26 classes separately come from? Will be good to answer! </p>",
      "rawMarkdown": "Congrats on your medal. As always, thanks for sharing  \nSorry for my lack of understanding, when you mention\n\n> For prediction I predict on slidding crops of each test recording and take the overall max prediction. This is slow as I need to do each of the 26 classes separately. This is also maybe where I lost against others: my model cannot learn long range temporal patterns, nor class interactions.\n\nDoes this mean u train 26 models for each class? From what i am reading, if you are masking your bce, you are able to output multi label predictions for each sample so where do the 26 classes separately come from? Will be good to answer!",
      "votes": null
    },
    {
      "id": "1207786",
      "postDate": "02/18/2021 03:26:11",
      "content": "<p><a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a>  If we crop by the f_min and f_max, doesn't that mean the image generated per species is not comparable anymore (since the y axis are representing different frequency region).  Would you explain how can that work ?<br>\nAnd big congratulations on the solo gold !</p>",
      "rawMarkdown": "cpmpml  If we crop by the f_min and f_max, doesn't that mean the image generated per species is not comparable anymore (since the y axis are representing different frequency region).  Would you explain how can that work ?\nAnd big congratulations on the solo gold !",
      "votes": null
    },
    {
      "id": "1208111",
      "postDate": "02/18/2021 07:00:30",
      "content": "<p>I train one model, but crops are different for each class.  I need to apply the model to each crop separately.</p>",
      "rawMarkdown": "I train one model, but crops are different for each class.  I need to apply the model to each crop separately.",
      "votes": null
    },
    {
      "id": "1208118",
      "postDate": "02/18/2021 07:03:04",
      "content": "<blockquote>\n  <p>Instead of cropping though I simply masked out the regions outside of the frequency band and time that were irrelevant. So I would crop around the region of the time and then I would mask out the frequencies that were irrelevant to that prediction. </p>\n</blockquote>\n<p>That's what I did.  Crop then pad.  Sorry if this is not clear enough in my post.</p>\n<p>It looks like you had the same idea as me ;)</p>",
      "rawMarkdown": ">  Instead of cropping though I simply masked out the regions outside of the frequency band and time that were irrelevant. So I would crop around the region of the time and then I would mask out the frequencies that were irrelevant to that prediction. \n\nThat's what I did.  Crop then pad.  Sorry if this is not clear enough in my post.\n\nIt looks like you had the same idea as me ;)",
      "votes": null
    },
    {
      "id": "1208120",
      "postDate": "02/18/2021 07:03:44",
      "content": "<blockquote>\n  <p>I thought Coereba flaveola was species_id=1 but I was not sure.</p>\n</blockquote>\n<p>Based on what data?</p>",
      "rawMarkdown": "> I thought Coereba flaveola was species_id=1 but I was not sure.\n\nBased on what data?",
      "votes": null
    },
    {
      "id": "1208132",
      "postDate": "02/18/2021 07:12:55",
      "content": "<p>I just checked, assuming all other targets are 0 instead of masking costs 0.01 on LB.  I tried both BCE and softmax in that case.</p>",
      "rawMarkdown": "I just checked, assuming all other targets are 0 instead of masking costs 0.01 on LB.  I tried both BCE and softmax in that case.",
      "votes": null
    },
    {
      "id": "1208138",
      "postDate": "02/18/2021 07:15:43",
      "content": "<blockquote>\n  <p>It turns out that CNN knows position of image. Because of effects of padding</p>\n</blockquote>\n<p>This is what i thought as well.  But I explored, for instance I read the paper you quote and others on positional encoding. It turns out that adding some positional encoding helps.</p>",
      "rawMarkdown": "> It turns out that CNN knows position of image. Because of effects of padding\n\nThis is what i thought as well.  But I explored, for instance I read the paper you quote and others on positional encoding. It turns out that adding some positional encoding helps.",
      "votes": null
    },
    {
      "id": "1208140",
      "postDate": "02/18/2021 07:19:37",
      "content": "<p>I did ask you for how your method differs from mine.  I hope you will explain.</p>\n<p>Tiny crops of code is not the same as a written explanation.  </p>\n<p>From what I see, you crop by f_min and f_max, hence I don't see how it is different from what I did.</p>",
      "rawMarkdown": "I did ask you for how your method differs from mine.  I hope you will explain.\n\nTiny crops of code is not the same as a written explanation.  \n\nFrom what I see, you crop by f_min and f_max, hence I don't see how it is different from what I did.",
      "votes": null
    },
    {
      "id": "1208158",
      "postDate": "02/18/2021 07:30:31",
      "content": "<p>Thanks for sharing your codes.  Mine is different.</p>\n<p>Given that each crop has a single class target, I just use a single target overall.  I know what class I am predicting when I crop, hence I don't need to give it to the model.  My code is then very simple:</p>\n<pre><code>        loss_fct = nn.BCEWithLogitsLoss()\n        logits = self.head(x)\n        mask = input_dict['mask']\n        preds = (logits * mask).sum(-1, keepdim=True)\n        mono_targets = input_dict['mono_target']\n        loss = loss_fct(preds, mono_targets)\n</code></pre>\n<p>This does no work:</p>\n<pre><code>        with torch.no_grad():\n            loss[target&lt;1] = 0   \n</code></pre>\n<p>because FP crops have  a target of 0 and you want the model to learn that.  You must mask only the loss for other classes.</p>\n<p>How is this</p>\n<pre><code>                    p = torch.sigmoid(l)\n                    logp = - torch.log(torch.clamp(p, 1e-4, 1-1e-4))\n                    logn = - torch.log(torch.clamp(1-p, 1e-4, 1-1e-4))\n                    loss = t*logp +(1-t)*logn\n\n                    loss  = loss.mean()\n</code></pre>\n<p>different from </p>\n<pre><code>    loss_fct = nn.BCEWithLogitsLoss()\n    loss = loss_fct(l, t)\n</code></pre>",
      "rawMarkdown": "Thanks for sharing your codes.  Mine is different.\n\nGiven that each crop has a single class target, I just use a single target overall.  I know what class I am predicting when I crop, hence I don't need to give it to the model.  My code is then very simple:\n\n```\n        loss_fct = nn.BCEWithLogitsLoss()\n        logits = self.head(x)\n        mask = input_dict['mask']\n        preds = (logits * mask).sum(-1, keepdim=True)\n        mono_targets = input_dict['mono_target']\n        loss = loss_fct(preds, mono_targets)\n```\n\nThis does no work:\n\n```\n        with torch.no_grad():\n            loss[target<1] = 0   \n\n```\nbecause FP crops have  a target of 0 and you want the model to learn that.  You must mask only the loss for other classes.\n\nHow is this\n\n```\n                    p = torch.sigmoid(l)\n                    logp = - torch.log(torch.clamp(p, 1e-4, 1-1e-4))\n                    logn = - torch.log(torch.clamp(1-p, 1e-4, 1-1e-4))\n                    loss = t*logp +(1-t)*logn\n\n                    loss  = loss.mean()\n\n```\n\ndifferent from \n\n```\n    loss_fct = nn.BCEWithLogitsLoss()\n    loss = loss_fct(l, t)\n```",
      "votes": null
    },
    {
      "id": "1208163",
      "postDate": "02/18/2021 07:32:21",
      "content": "<blockquote>\n  <p>doesn't that mean the image generated per species is not comparable anymore </p>\n</blockquote>\n<p>Yes, this is why I need to perform inference separately for each class.  </p>",
      "rawMarkdown": ">  doesn't that mean the image generated per species is not comparable anymore \n\nYes, this is why I need to perform inference separately for each class.",
      "votes": null
    },
    {
      "id": "1208375",
      "postDate": "02/18/2021 09:22:08",
      "content": "<p>I spent a few days trying almost exactly the same idea as yours, although I couldn't make it work, it looks like it wasn't that bad after all :)<br>\nAnd congratz on the impressive performance !</p>",
      "rawMarkdown": "I spent a few days trying almost exactly the same idea as yours, although I couldn't make it work, it looks like it wasn't that bad after all :)\nAnd congratz on the impressive performance !",
      "votes": null
    },
    {
      "id": "1208618",
      "postDate": "02/18/2021 11:24:38",
      "content": "<p>Thanks.  We all tried stuff that failed for us but worked for others apparently.  Mine was pseudo labeling.  </p>",
      "rawMarkdown": "Thanks.  We all tried stuff that failed for us but worked for others apparently.  Mine was pseudo labeling.",
      "votes": null
    },
    {
      "id": "1208631",
      "postDate": "02/18/2021 11:38:09",
      "content": "<p>S7 is correct for that one</p>",
      "rawMarkdown": "S7 is correct for that one",
      "votes": null
    },
    {
      "id": "1208639",
      "postDate": "02/18/2021 11:50:31",
      "content": "<p><a href=\"https://www.kaggle.com/CPMP\" target=\"_blank\">@CPMP</a>. I have completed the writeup and outlined the similarities and differences between your and my approaches. under the discussion \"Devil is in detail but it seems we have a similar method. I welcome comments about what we did differently …. \" Please check it.</p>\n<p>you mentioned here that \"… models probably, like SED or Transformers.\". I am curious how will you apply transformer here. Can you give some hints?</p>",
      "rawMarkdown": "CPMP. I have completed the writeup and outlined the similarities and differences between your and my approaches. under the discussion \"Devil is in detail but it seems we have a similar method. I welcome comments about what we did differently .... \" Please check it.\n\nyou mentioned here that \"... models probably, like SED or Transformers.\". I am curious how will you apply transformer here. Can you give some hints?",
      "votes": null
    },
    {
      "id": "1208677",
      "postDate": "02/18/2021 12:31:14",
      "content": "<p>Do you mean for each class slide for x axis(time) and fixed crop for y axis (as we know frequency)? So we move along x and take max?</p>\n<p>Thank you and congratulations!  </p>",
      "rawMarkdown": "Do you mean for each class slide for x axis(time) and fixed crop for y axis (as we know frequency)? So we move along x and take max?\n\nThank you and congratulations!",
      "votes": null
    },
    {
      "id": "1208710",
      "postDate": "02/18/2021 12:57:07",
      "content": "<p>Congrats on the solo gold! I have two quick questions: </p>\n<ol>\n<li>How do you crop the frequencies, do you specify the <code>fmin</code> and <code>fmax</code> in <code>librosa.feature.melspectrogram</code>, and the frequency sizes of generated features are still controlled by <code>n_mels</code>; or you use an overall initial range and crop the generated features on frequency axis later? If it's the latter, how to do it, could you provide a few lines of code? And how to make them in batch, also padding like time axis?</li>\n<li>Did you use the FP exactly as you use TP? Is there any trick? Cause the labels in FP are wrong and I have tried to use them, it really harms the performance.</li>\n</ol>",
      "rawMarkdown": "Congrats on the solo gold! I have two quick questions: \n1. How do you crop the frequencies, do you specify the `fmin` and `fmax` in `librosa.feature.melspectrogram`, and the frequency sizes of generated features are still controlled by `n_mels`; or you use an overall initial range and crop the generated features on frequency axis later? If it's the latter, how to do it, could you provide a few lines of code? And how to make them in batch, also padding like time axis?\n2. Did you use the FP exactly as you use TP? Is there any trick? Cause the labels in FP are wrong and I have tried to use them, it really harms the performance.",
      "votes": null
    },
    {
      "id": "1208857",
      "postDate": "02/18/2021 14:16:42",
      "content": "<p>Thanks for answering. </p>\n<p>My understanding is that yes, the sliding must be across the time axis for each frequency axis associated with the classes. </p>",
      "rawMarkdown": "Thanks for answering. \n\nMy understanding is that yes, the sliding must be across the time axis for each frequency axis associated with the classes.",
      "votes": null
    },
    {
      "id": "1208880",
      "postDate": "02/18/2021 14:34:44",
      "content": "<p>why sum? not just max?</p>",
      "rawMarkdown": "why sum? not just max?",
      "votes": null
    },
    {
      "id": "1208916",
      "postDate": "02/18/2021 15:03:20",
      "content": "<blockquote>\n  <p>I guess that most teams reused your SED model from last competition. This is also something to be proud of.</p>\n</blockquote>\n<p>Happy to know many teams used SED model. Unfortunately we couldn't make it work better than other kind of models, but it turned out some top performing teams successfully used it. I feel like I found tons of things to learn from this competition.</p>",
      "rawMarkdown": ">  I guess that most teams reused your SED model from last competition. This is also something to be proud of.\n\nHappy to know many teams used SED model. Unfortunately we couldn't make it work better than other kind of models, but it turned out some top performing teams successfully used it. I feel like I found tons of things to learn from this competition.",
      "votes": null
    },
    {
      "id": "1208977",
      "postDate": "02/18/2021 15:47:21",
      "content": "<ol>\n<li><p>I cropped the spectrogram after generating it.  I used fmin=90 and fmax=14000 for generating spectrogram.  To find which pixels to crop, librosa has utilities functions that map frequencies to the number of mel.</p></li>\n<li><p>I used FP and TP the same way.  They give a 0 or a 1 label for an image.</p></li>\n</ol>\n<p>What is wrong with FP labels?  They were good enough to get me a gold medal ;)</p>",
      "rawMarkdown": "1. I cropped the spectrogram after generating it.  I used fmin=90 and fmax=14000 for generating spectrogram.  To find which pixels to crop, librosa has utilities functions that map frequencies to the number of mel.\n\n2. I used FP and TP the same way.  They give a 0 or a 1 label for an image.\n\nWhat is wrong with FP labels?  They were good enough to get me a gold medal ;)",
      "votes": null
    },
    {
      "id": "1209018",
      "postDate": "02/18/2021 16:20:45",
      "content": "<blockquote>\n  <p>Do you mean for each class slide for x axis(time) and fixed crop for y axis (as we know frequency)? So we move along x and take max?</p>\n</blockquote>\n<p>yes.</p>",
      "rawMarkdown": "> Do you mean for each class slide for x axis(time) and fixed crop for y axis (as we know frequency)? So we move along x and take max?\n\nyes.",
      "votes": null
    },
    {
      "id": "1209052",
      "postDate": "02/18/2021 16:49:35",
      "content": "<p>I feel like one thing that people are not understanding with the FPs is that they are actually our best 0's. Some people have been making the mistake of using them as 1's which is explicitly what we dont want. </p>",
      "rawMarkdown": "I feel like one thing that people are not understanding with the FPs is that they are actually our best 0's. Some people have been making the mistake of using them as 1's which is explicitly what we dont want.",
      "votes": null
    },
    {
      "id": "1209074",
      "postDate": "02/18/2021 17:03:41",
      "content": "<blockquote>\n  <p>they are actually our best 0's.</p>\n</blockquote>\n<p><a href=\"https://www.kaggle.com/ryches\" target=\"_blank\">@ryches</a> Oh! So I should train my model to predict the species listed in FP do not appear instead appearing, is it how you did <a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a> ? <br>\nIt's so bad I figure this out so late…</p>",
      "rawMarkdown": "> they are actually our best 0's.\n\n@ryches Oh! So I should train my model to predict the species listed in FP do not appear instead appearing, is it how you did @cpmpml ? \nIt's so bad I figure this out so late...",
      "votes": null
    },
    {
      "id": "1209166",
      "postDate": "02/18/2021 18:19:40",
      "content": "<p>Yes, the false positives are regions that showed up using the template matching but an expert reviewed and said specifically the species was not present. So training your model to detect false positives the same as true positives is not what we want to do</p>",
      "rawMarkdown": "Yes, the false positives are regions that showed up using the template matching but an expert reviewed and said specifically the species was not present. So training your model to detect false positives the same as true positives is not what we want to do",
      "votes": null
    },
    {
      "id": "1209170",
      "postDate": "02/18/2021 18:22:16",
      "content": "<p>wow that's huge for me thanks for checking and sharing!</p>",
      "rawMarkdown": "wow that's huge for me thanks for checking and sharing!",
      "votes": null
    },
    {
      "id": "1209304",
      "postDate": "02/18/2021 20:08:45",
      "content": "<blockquote>\n  <p>they are actually our best 0's. </p>\n</blockquote>\n<p>Indeed.  I'd even say they are our ONLY 0's.</p>",
      "rawMarkdown": "> they are actually our best 0's. \n\nIndeed.  I'd even say they are our ONLY 0's.",
      "votes": null
    },
    {
      "id": "1209311",
      "postDate": "02/18/2021 20:12:16",
      "content": "<p>Tx, will check your writeup next.</p>\n<p>For transformers I refer to a speech to text model that was released recently by facebook I think.  I haven't looked in detail, maybe it is not applicable because ground truth is too sparse here.</p>",
      "rawMarkdown": "Tx, will check your writeup next.\n\nFor transformers I refer to a speech to text model that was released recently by facebook I think.  I haven't looked in detail, maybe it is not applicable because ground truth is too sparse here.",
      "votes": null
    },
    {
      "id": "1209314",
      "postDate": "02/18/2021 20:15:04",
      "content": "<p><a href=\"https://www.kaggle.com/philippsinger\" target=\"_blank\">@philippsinger</a> Thanks, and thanks for addressing my concerns about oversharing ;)</p>\n<p>It was a bit stressful to wake up every morning and see someone pass me while I was stuck.  But the story has a happy ending, everything is good.</p>\n<p>Congrats on your nth win in a row, this is amazing.</p>",
      "rawMarkdown": "philippsinger Thanks, and thanks for addressing my concerns about oversharing ;)\n\nIt was a bit stressful to wake up every morning and see someone pass me while I was stuck.  But the story has a happy ending, everything is good.\n\nCongrats on your nth win in a row, this is amazing.",
      "votes": null
    },
    {
      "id": "1212813",
      "postDate": "02/21/2021 16:21:58",
      "content": "<p>I didn't think of post processing but I did think of pseudo labeling during the competition, yet could not make it work.  After reading all the writeups where teams who passed me successfully used PL I revisited what i did.  The issue was that instead of randomly sample crops for PL, I imposed a distribution that matches training samples, i.e. same number of positive pseudo labels per class.  When I remove this bias then pseudo labeling works.  In my first experiment, a single model gets a 0.01 boost on public and private LB.  With tuning and iterations I now see how I could have moved higher.  </p>\n<p>I am not sure why I imposed this sampling bias.  I'll try to be more careful next time.</p>",
      "rawMarkdown": "I didn't think of post processing but I did think of pseudo labeling during the competition, yet could not make it work.  After reading all the writeups where teams who passed me successfully used PL I revisited what i did.  The issue was that instead of randomly sample crops for PL, I imposed a distribution that matches training samples, i.e. same number of positive pseudo labels per class.  When I remove this bias then pseudo labeling works.  In my first experiment, a single model gets a 0.01 boost on public and private LB.  With tuning and iterations I now see how I could have moved higher.  \n\nI am not sure why I imposed this sampling bias.  I'll try to be more careful next time.",
      "votes": null
    },
    {
      "id": "1330750",
      "postDate": "06/01/2021 04:41:18",
      "content": "<p>Thanks for the excellent and detailed write-up! I hope I wont be too late to ask a few questions wrt ur approach:</p>\n<ol>\n<li>u mentioned u made the time-axis of each crop the same (by cropping/ padding), how did handle the differences in frequency-axis? (as in different species has different height)</li>\n<li>may I learn more abt how u applied positional encoding along frequency-axis?  </li>\n</ol>\n<p>Thanks!</p>",
      "rawMarkdown": "Thanks for the excellent and detailed write-up! I hope I wont be too late to ask a few questions wrt ur approach:\n1. u mentioned u made the time-axis of each crop the same (by cropping/ padding), how did handle the differences in frequency-axis? (as in different species has different height)\n2. may I learn more abt how u applied positional encoding along frequency-axis?  \n\nThanks!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1207601,
      "author_name": "antorsae",
      "author_url": "",
      "post_date": "02/18/2021 00:10:03",
      "content": "<p>Very nice writeup. We used a very similar approach, which we really kicked into action at the last 8 days; will share details tomorrow… congrats on the solo medal!</p>",
      "votes": null,
      "replies": [
        {
          "id": 1207615,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "02/18/2021 00:18:07",
          "content": "<p>Tx.  My limited skills in computer vision may be what prevented me from getting a higher score.  Looking forward to your writeup.  I am sure I'll learn stuff.  And congrats as well on your result.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1207603,
      "author_name": "cdeotte",
      "author_url": "",
      "post_date": "02/18/2021 00:10:29",
      "content": "<p>Congrats <a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a> . Great solo finish! I'm glad you won Gold. You helped many teams by showing them what was possible!</p>",
      "votes": null,
      "replies": [
        {
          "id": 1207613,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "02/18/2021 00:16:52",
          "content": "<p>tx.  Not only did I show it, but I shared I was using an image classification model.  Next time I'll share less maybe ;)</p>\n<p>Congrats on your solo gold too.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1207619,
          "author_name": "philippsinger",
          "author_url": "",
          "post_date": "02/18/2021 00:20:15",
          "content": "<p>Congrats, amazing job. We were indeed very impressed by the 931 start. I dont think you revealed anything with the image model though, so dont worry about that. Everyone uses image models in these types of problems.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1209314,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "02/18/2021 20:15:04",
          "content": "<p><a href=\"https://www.kaggle.com/philippsinger\" target=\"_blank\">@philippsinger</a> Thanks, and thanks for addressing my concerns about oversharing ;)</p>\n<p>It was a bit stressful to wake up every morning and see someone pass me while I was stuck.  But the story has a happy ending, everything is good.</p>\n<p>Congrats on your nth win in a row, this is amazing.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1207608,
      "author_name": "ryches",
      "author_url": "",
      "post_date": "02/18/2021 00:14:45",
      "content": "<p>Great work. So many techniques that are close to what I tried, but just never put them together in the right way. </p>\n<p>It seemed like the public pannloss that was floating around was trying to implement that masked loss like was mentioned in the paper, but was kind of implemented wrong. </p>",
      "votes": null,
      "replies": [
        {
          "id": 1207616,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "02/18/2021 00:19:09",
          "content": "<p>Tx.  I didn't look at public notebooks at all.  Maybe I should have…  It also looks like others used the same overall method, but got better results than me.  </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1207612,
      "author_name": "jihangz",
      "author_url": "",
      "post_date": "02/18/2021 00:16:37",
      "content": "<p>Congrats on the solo gold! I have one question: did you crop y-axis (the frequency axis) too when doing training and inference?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1207617,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "02/18/2021 00:19:27",
          "content": "<p>Yes.  I cropped by f_min and f_max.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1207786,
          "author_name": "nyleve",
          "author_url": "",
          "post_date": "02/18/2021 03:26:11",
          "content": "<p><a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a>  If we crop by the f_min and f_max, doesn't that mean the image generated per species is not comparable anymore (since the y axis are representing different frequency region).  Would you explain how can that work ?<br>\nAnd big congratulations on the solo gold !</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1208163,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "02/18/2021 07:32:21",
          "content": "<blockquote>\n  <p>doesn't that mean the image generated per species is not comparable anymore </p>\n</blockquote>\n<p>Yes, this is why I need to perform inference separately for each class.  </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1207614,
      "author_name": "rytisva88",
      "author_url": "",
      "post_date": "02/18/2021 00:17:04",
      "content": "<p>Congratulations! Could you share how exactly did you do the loss masking? I have tried this early in competition (and actually specifically inspired by your approach to secondary labels in BirdCall), but it backfired quite heavily to me (I only did time based crops, frequency cropping may be even more important, but intuitively I thought even without frequency cropping loss masking should make more sense…)</p>",
      "votes": null,
      "replies": [
        {
          "id": 1207627,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "02/18/2021 00:26:25",
          "content": "<p>I am not sure about what is not clear.  You compute the BCE loss for all targets (no reduction), then multiply by the mask (1 for the song class to be predicted, 0 for all other classes), then take the mean for the batch.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1207651,
          "author_name": "rytisva88",
          "author_url": "",
          "post_date": "02/18/2021 00:41:39",
          "content": "<p>Exactly, sounds quite straightforward to me, so just checking whether I am not missing anything. So something like below, right?</p>\n<pre><code>class RainforestLossMasked(nn.Module):\n    def __init__(self):\n        super().__init__()    \n        self.bce_prob  = nn.BCELoss(reduction='none')\n\n    def forward(self, out, target):                \n        loss = self.bce_prob(out, target)     \n        with torch.no_grad():\n            loss[target&lt;1] = 0   \n        loss = loss.mean()\n        return loss\n</code></pre>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1207697,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "02/18/2021 01:22:38",
          "content": "<p>something like this:</p>\n<pre><code>                def use_tp_fp_binary_cross_entropy(logit, label, is_tp):\n                    batch_size, num_label = logit.shape\n                    batch_size = label.shape # e.g. [ 2,2,6,8, ... 23]\n                    batch_size = is_tp.shape # e.g. [ 1,0,0,1, ... 1]\n\n                    l = logit.gather(1, label.unsqueeze(1))\n                    l = l.reshape(-1)\n                    t = is_tp.float()\n\n                    p = torch.sigmoid(l)\n                    logp = - torch.log(torch.clamp(p, 1e-4, 1-1e-4))\n                    logn = - torch.log(torch.clamp(1-p, 1e-4, 1-1e-4))\n                    loss = t*logp +(1-t)*logn\n\n                    loss  = loss.mean()\n                    return loss\n</code></pre>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1208158,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "02/18/2021 07:30:31",
          "content": "<p>Thanks for sharing your codes.  Mine is different.</p>\n<p>Given that each crop has a single class target, I just use a single target overall.  I know what class I am predicting when I crop, hence I don't need to give it to the model.  My code is then very simple:</p>\n<pre><code>        loss_fct = nn.BCEWithLogitsLoss()\n        logits = self.head(x)\n        mask = input_dict['mask']\n        preds = (logits * mask).sum(-1, keepdim=True)\n        mono_targets = input_dict['mono_target']\n        loss = loss_fct(preds, mono_targets)\n</code></pre>\n<p>This does no work:</p>\n<pre><code>        with torch.no_grad():\n            loss[target&lt;1] = 0   \n</code></pre>\n<p>because FP crops have  a target of 0 and you want the model to learn that.  You must mask only the loss for other classes.</p>\n<p>How is this</p>\n<pre><code>                    p = torch.sigmoid(l)\n                    logp = - torch.log(torch.clamp(p, 1e-4, 1-1e-4))\n                    logn = - torch.log(torch.clamp(1-p, 1e-4, 1-1e-4))\n                    loss = t*logp +(1-t)*logn\n\n                    loss  = loss.mean()\n</code></pre>\n<p>different from </p>\n<pre><code>    loss_fct = nn.BCEWithLogitsLoss()\n    loss = loss_fct(l, t)\n</code></pre>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1207622,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "02/18/2021 00:21:58",
      "content": "<p><a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a>   <br>\n\"is slow as I need to do each of the 26 classes separately\"<br>\n\"Yes. I cropped by f_min and f_max.\"<br>\n…</p>\n<p>see my solution to solve that<br>\n<a href=\"https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/220309\" target=\"_blank\">https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/220309</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 1208140,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "02/18/2021 07:19:37",
          "content": "<p>I did ask you for how your method differs from mine.  I hope you will explain.</p>\n<p>Tiny crops of code is not the same as a written explanation.  </p>\n<p>From what I see, you crop by f_min and f_max, hence I don't see how it is different from what I did.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1208639,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "02/18/2021 11:50:31",
          "content": "<p><a href=\"https://www.kaggle.com/CPMP\" target=\"_blank\">@CPMP</a>. I have completed the writeup and outlined the similarities and differences between your and my approaches. under the discussion \"Devil is in detail but it seems we have a similar method. I welcome comments about what we did differently …. \" Please check it.</p>\n<p>you mentioned here that \"… models probably, like SED or Transformers.\". I am curious how will you apply transformer here. Can you give some hints?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1209311,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "02/18/2021 20:12:16",
          "content": "<p>Tx, will check your writeup next.</p>\n<p>For transformers I refer to a speech to text model that was released recently by facebook I think.  I haven't looked in detail, maybe it is not applicable because ground truth is too sparse here.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1207623,
      "author_name": "philippsinger",
      "author_url": "",
      "post_date": "02/18/2021 00:22:46",
      "content": "<p>Did the extra class for the songs help anything? I had it on the list since day 1 but never got around to try it. I think it was for low-frequency classes but would need to re-check.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1207630,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "02/18/2021 00:27:15",
          "content": "<p>Sorry, I did not use it here.  My bad.  Let me fix the writeup.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1207628,
      "author_name": "snnclsr",
      "author_url": "",
      "post_date": "02/18/2021 00:26:45",
      "content": "<p>Congrats on your solo gold <a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a> . Your entrance to the competition was inspirational 💪</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1207637,
      "author_name": "sgalib",
      "author_url": "",
      "post_date": "02/18/2021 00:34:05",
      "content": "<p>Congrats! The gradient idea is nice. I have also tried applying fixed grid lines on the images. That helped.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1207664,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "02/18/2021 00:50:25",
          "content": "<p>Grid may be more effective actually, I'll try it next time.  Congrats on your strong silver.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1207686,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "02/18/2021 01:06:32",
          "content": "<p>How much Position Information Do Convolutional Neural Networks Encode?<br>\n<a href=\"https://openreview.net/pdf?id=rJeB36NKvB\" target=\"_blank\">https://openreview.net/pdf?id=rJeB36NKvB</a></p>\n<p>It turns out that CNN knows position of image. Because of effects of padding, CNN knows where the  input window is in the image, if the effective receptive field is large (e.g. same size of larger than the size of the input image). Hence adding position gradient is sometimes not necessary</p>\n<p>quote:<br>\n\"Experiments reveal that positional information is available to a strong degree\"</p>\n<p>\"Results point to zero padding and borders as an anchor from which spatial information is derived and eventually propagated over the whole image as spatial abstraction occurs\"</p>\n<hr>\n<p>I think there are also some interesting papers that experiment positional encoding used in transformer for CNN</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1208138,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "02/18/2021 07:15:43",
          "content": "<blockquote>\n  <p>It turns out that CNN knows position of image. Because of effects of padding</p>\n</blockquote>\n<p>This is what i thought as well.  But I explored, for instance I read the paper you quote and others on positional encoding. It turns out that adding some positional encoding helps.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1207650,
      "author_name": "hidehisaarai1213",
      "author_url": "",
      "post_date": "02/18/2021 00:41:35",
      "content": "<p>Congratulations <a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a> for strong solo finish! 0.931 for the first sub was quite impressive! We learned a lot from you, thank you!</p>",
      "votes": null,
      "replies": [
        {
          "id": 1207662,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "02/18/2021 00:49:35",
          "content": "<p>Thanks.  You finished strongly too.  I am sure that if your studies had left you more time then you would have passed me.  Also, I guess that most teams reused your SED model from last competition.  This is also something to be proud of.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1208916,
          "author_name": "hidehisaarai1213",
          "author_url": "",
          "post_date": "02/18/2021 15:03:20",
          "content": "<blockquote>\n  <p>I guess that most teams reused your SED model from last competition. This is also something to be proud of.</p>\n</blockquote>\n<p>Happy to know many teams used SED model. Unfortunately we couldn't make it work better than other kind of models, but it turned out some top performing teams successfully used it. I feel like I found tons of things to learn from this competition.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1207672,
      "author_name": "duykhanh99",
      "author_url": "",
      "post_date": "02/18/2021 00:55:37",
      "content": "<p>Congrats on 12th place and gold medal <a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a> </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1207683,
      "author_name": "dicksonchin93",
      "author_url": "",
      "post_date": "02/18/2021 01:05:33",
      "content": "<p>Congrats! thank you for making this competition so much more competitive :) also, I wanted to ask how much did masking the loss and using a 26 classes model instead of 24 classes help independently?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1207696,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "02/18/2021 01:20:03",
          "content": "<p>I only tried 24 classes omce, and it was a bit worse.</p>\n<p>Congrats on your strong finish.  Looking forward to your writeup.  </p>\n<p>I wish I had waited two more week before submitting ;)</p>\n<p>Just kidding, I am sure others would have submitted high scores soon.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1207708,
          "author_name": "dicksonchin93",
          "author_url": "",
          "post_date": "02/18/2021 01:34:45",
          "content": "<p>Thank you! will post a writeup soon. I am sure quite a few entered after looking at your 0.931 post and also many that already entered started to look for alternative solutions too haha!</p>\n<p>how about masking the loss, how much did it help?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1208132,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "02/18/2021 07:12:55",
          "content": "<p>I just checked, assuming all other targets are 0 instead of masking costs 0.01 on LB.  I tried both BCE and softmax in that case.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1209170,
          "author_name": "dicksonchin93",
          "author_url": "",
          "post_date": "02/18/2021 18:22:16",
          "content": "<p>wow that's huge for me thanks for checking and sharing!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1207700,
      "author_name": "ryches",
      "author_url": "",
      "post_date": "02/18/2021 01:26:49",
      "content": "<p>You mentioned focal loss earlier in the competition. Did you ever try that in place of bce in your final solution? In my experience that has only ever been useful for extreme class imbalance problems like segmentation where there is a small number of positives to a large number of negatives. </p>\n<p>I feel like this is another one of the competitions where I had most of the ingredients for the top solutions out but couldn't quite find the right recipe to combine them all in. I tried the masked loss, frequency banded model that I implemented a bit differently than yours, but similar concept. I mentioned to our team at one point the linear gradient to show position more explicitly to the cnn. </p>",
      "votes": null,
      "replies": [
        {
          "id": 1207702,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "02/18/2021 01:29:00",
          "content": "<p>Good question!</p>\n<p>Yes I tried it at some point and it was a bit worse than vanilla bce.</p>\n<p>Focal loss only worked for me in segmentation tasks.  This is the original use case of focal loss unless mistaken.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1207701,
      "author_name": "ryches",
      "author_url": "",
      "post_date": "02/18/2021 01:27:45",
      "content": "<p>Final question: did you validate against the crops with your image model or did you apply the same test time rolling validation? I'd assume you validated against the crops since the rolling window inference was slow. </p>",
      "votes": null,
      "replies": [
        {
          "id": 1207704,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "02/18/2021 01:31:06",
          "content": "<p>I used 5 fold CV on my crops.  And I used roc-auc.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1207715,
      "author_name": "hayaocelot",
      "author_url": "",
      "post_date": "02/18/2021 01:44:36",
      "content": "<p>Congrats on your solo gold. Your discussion comments helped me many times.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1207728,
      "author_name": "ryches",
      "author_url": "",
      "post_date": "02/18/2021 02:12:12",
      "content": "<p>Kind of kicking myself for not exploring this further. The first thing I sent to the team when I joined them was showing them a model I had that was kind of similar to yours. Instead of cropping though I simply masked out the regions outside of the frequency band and time that were irrelevant. So I would crop around the region of the time and then I would mask out the frequencies that were irrelevant to that prediction. </p>\n<p><img src=\"https://i.imgur.com/tJPa7Sv.png\" alt=\"\"></p>\n<p>Would look something like this. I made it so they were long enough that I never had to do any cropping, only ever padding near the beginning or end of the audio. Would use similar procedure to you at test time, but I could stack all of them together fairly easily so it was 16(num_time_steps)*24(num_frequency_ranges), 128(num_mel_bins), 500 (time) and it was fairly fast to do inference. On the crops themselves I got validation performance of .985 at times, but when I applied the rolling window validation I would get poor results like .6-.7. </p>\n<p>I tried the masked loss, but I did not apply it to this specific model and I never made a submission with it because the windowed validation looked poor. I considered the linear gradient to give position but ended up not using it in this setup because I figured the model already had its relative position based on the amount of 0 padding above and below the unmasked signal. Might have to fiddle with that and see if it was actually good. </p>",
      "votes": null,
      "replies": [
        {
          "id": 1208118,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "02/18/2021 07:03:04",
          "content": "<blockquote>\n  <p>Instead of cropping though I simply masked out the regions outside of the frequency band and time that were irrelevant. So I would crop around the region of the time and then I would mask out the frequencies that were irrelevant to that prediction. </p>\n</blockquote>\n<p>That's what I did.  Crop then pad.  Sorry if this is not clear enough in my post.</p>\n<p>It looks like you had the same idea as me ;)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1207730,
      "author_name": "bayartsogtya",
      "author_url": "",
      "post_date": "02/18/2021 02:14:10",
      "content": "<p>Thank you for the great write-up and congratz on solo gold!<br>\nI thought <code>Coereba flaveola</code> was <code>species_id=1</code> but I was not sure.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1208120,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "02/18/2021 07:03:44",
          "content": "<blockquote>\n  <p>I thought Coereba flaveola was species_id=1 but I was not sure.</p>\n</blockquote>\n<p>Based on what data?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1208631,
          "author_name": "philippsinger",
          "author_url": "",
          "post_date": "02/18/2021 11:38:09",
          "content": "<p>S7 is correct for that one</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1207774,
      "author_name": "germmie",
      "author_url": "",
      "post_date": "02/18/2021 03:13:23",
      "content": "<p>Congrats on your medal. As always, thanks for sharing  <br>\nSorry for my lack of understanding, when you mention</p>\n<blockquote>\n  <p>For prediction I predict on slidding crops of each test recording and take the overall max prediction. This is slow as I need to do each of the 26 classes separately. This is also maybe where I lost against others: my model cannot learn long range temporal patterns, nor class interactions.</p>\n</blockquote>\n<p>Does this mean u train 26 models for each class? From what i am reading, if you are masking your bce, you are able to output multi label predictions for each sample so where do the 26 classes separately come from? Will be good to answer! </p>",
      "votes": null,
      "replies": [
        {
          "id": 1208111,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "02/18/2021 07:00:30",
          "content": "<p>I train one model, but crops are different for each class.  I need to apply the model to each crop separately.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1208677,
          "author_name": "egm108",
          "author_url": "",
          "post_date": "02/18/2021 12:31:14",
          "content": "<p>Do you mean for each class slide for x axis(time) and fixed crop for y axis (as we know frequency)? So we move along x and take max?</p>\n<p>Thank you and congratulations!  </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1208857,
          "author_name": "germmie",
          "author_url": "",
          "post_date": "02/18/2021 14:16:42",
          "content": "<p>Thanks for answering. </p>\n<p>My understanding is that yes, the sliding must be across the time axis for each frequency axis associated with the classes. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1208880,
          "author_name": "egm108",
          "author_url": "",
          "post_date": "02/18/2021 14:34:44",
          "content": "<p>why sum? not just max?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1209018,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "02/18/2021 16:20:45",
          "content": "<blockquote>\n  <p>Do you mean for each class slide for x axis(time) and fixed crop for y axis (as we know frequency)? So we move along x and take max?</p>\n</blockquote>\n<p>yes.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1208375,
      "author_name": "theoviel",
      "author_url": "",
      "post_date": "02/18/2021 09:22:08",
      "content": "<p>I spent a few days trying almost exactly the same idea as yours, although I couldn't make it work, it looks like it wasn't that bad after all :)<br>\nAnd congratz on the impressive performance !</p>",
      "votes": null,
      "replies": [
        {
          "id": 1208618,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "02/18/2021 11:24:38",
          "content": "<p>Thanks.  We all tried stuff that failed for us but worked for others apparently.  Mine was pseudo labeling.  </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1208710,
      "author_name": "karlyukang",
      "author_url": "",
      "post_date": "02/18/2021 12:57:07",
      "content": "<p>Congrats on the solo gold! I have two quick questions: </p>\n<ol>\n<li>How do you crop the frequencies, do you specify the <code>fmin</code> and <code>fmax</code> in <code>librosa.feature.melspectrogram</code>, and the frequency sizes of generated features are still controlled by <code>n_mels</code>; or you use an overall initial range and crop the generated features on frequency axis later? If it's the latter, how to do it, could you provide a few lines of code? And how to make them in batch, also padding like time axis?</li>\n<li>Did you use the FP exactly as you use TP? Is there any trick? Cause the labels in FP are wrong and I have tried to use them, it really harms the performance.</li>\n</ol>",
      "votes": null,
      "replies": [
        {
          "id": 1208977,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "02/18/2021 15:47:21",
          "content": "<ol>\n<li><p>I cropped the spectrogram after generating it.  I used fmin=90 and fmax=14000 for generating spectrogram.  To find which pixels to crop, librosa has utilities functions that map frequencies to the number of mel.</p></li>\n<li><p>I used FP and TP the same way.  They give a 0 or a 1 label for an image.</p></li>\n</ol>\n<p>What is wrong with FP labels?  They were good enough to get me a gold medal ;)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1209052,
          "author_name": "ryches",
          "author_url": "",
          "post_date": "02/18/2021 16:49:35",
          "content": "<p>I feel like one thing that people are not understanding with the FPs is that they are actually our best 0's. Some people have been making the mistake of using them as 1's which is explicitly what we dont want. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1209074,
          "author_name": "karlyukang",
          "author_url": "",
          "post_date": "02/18/2021 17:03:41",
          "content": "<blockquote>\n  <p>they are actually our best 0's.</p>\n</blockquote>\n<p><a href=\"https://www.kaggle.com/ryches\" target=\"_blank\">@ryches</a> Oh! So I should train my model to predict the species listed in FP do not appear instead appearing, is it how you did <a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a> ? <br>\nIt's so bad I figure this out so late…</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1209166,
          "author_name": "ryches",
          "author_url": "",
          "post_date": "02/18/2021 18:19:40",
          "content": "<p>Yes, the false positives are regions that showed up using the template matching but an expert reviewed and said specifically the species was not present. So training your model to detect false positives the same as true positives is not what we want to do</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1209304,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "02/18/2021 20:08:45",
          "content": "<blockquote>\n  <p>they are actually our best 0's. </p>\n</blockquote>\n<p>Indeed.  I'd even say they are our ONLY 0's.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1212813,
      "author_name": "cpmpml",
      "author_url": "",
      "post_date": "02/21/2021 16:21:58",
      "content": "<p>I didn't think of post processing but I did think of pseudo labeling during the competition, yet could not make it work.  After reading all the writeups where teams who passed me successfully used PL I revisited what i did.  The issue was that instead of randomly sample crops for PL, I imposed a distribution that matches training samples, i.e. same number of positive pseudo labels per class.  When I remove this bias then pseudo labeling works.  In my first experiment, a single model gets a 0.01 boost on public and private LB.  With tuning and iterations I now see how I could have moved higher.  </p>\n<p>I am not sure why I imposed this sampling bias.  I'll try to be more careful next time.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1330750,
      "author_name": "alexlwh",
      "author_url": "",
      "post_date": "06/01/2021 04:41:18",
      "content": "<p>Thanks for the excellent and detailed write-up! I hope I wont be too late to ask a few questions wrt ur approach:</p>\n<ol>\n<li>u mentioned u made the time-axis of each crop the same (by cropping/ padding), how did handle the differences in frequency-axis? (as in different species has different height)</li>\n<li>may I learn more abt how u applied positional encoding along frequency-axis?  </li>\n</ol>\n<p>Thanks!</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1207589": "First of all, thanks to the host, for providing yet another very interesting audio challenge.  The fact that it was only partly labelled was a significant difference with previous bird song competition.\n\nSecond, congrats to all those who managed to pass the 0.95 bar on public LB.  I couldn't, and, as I write this before competition deadline, I don't know why.  I tried a lot of things, and it looks like my approach has some fundamental limit.\n\nYet it was good enough to produce a first submission at 0.931, placing directly at 3rd spot while the competition had started two months earlier.\n\nThis looked great to me.  In hindsight, if my first sub had been weaker, then I would not have stick to its model and would have explored other models probably, like SED or Transformers.\n\nAnyway, no need to complain, I learned a lot of stuff along the way, like how to efficiently implement teacher student training of all sorts.  I hope this knowledge will be useful in the future.\n\nBack to the topic, my approach was extremely simple:  each row of train data, TP or FP, gives us a label for a crop in the (log mel) spectrogram of the corresponding recording.  If time is x axis and frequency the y axis, as is generally the case, then t_min, t_max gives bounds on x axis, and f_min, f_max gives bounds on the y axis.\n\nWe then have a 26 multi label classification problem (24 species but two species have 2 song types.  I treated each species/song type as a different class).  This is easily handled with BCE Loss.\n\nThe only little caveat is that we are given 26 classes (species + song type) but we get only one class label, 0 or 1 per image.  We only have to mask the loss for other classes and that's it!\n\nI didn't know it when I did it, but a similar way has been used by some of the host of the competition, in this paper (not the one shared in the forum): https://www.sciencedirect.com/science/article/abs/pii/S0003682X20304795\n\nThe other caveat is that the competition metric works with a label for every class.  Which we don't have in train data.  However the competition metric is very similar to a roc auc score per recording: when a pair of predictions is in the wrong order, i.e. a positive label has a prediction lower than another negative label prediction, then the metric is lowered.  As a proxy I decided to use roc-auc on my multi label classification problem.  Correlation with public LB is noisy, but it was good enough to let me make progress without submitting for a while.\n\nWhat worked best for me was to not resize the crops.  It means my model had to learn from sometimes tiny images.  To make it work by batch I pad all images to 4 seconds on the x axis.  Crops longer than that were resized on the x axis.  Shorter ones were padded with 0. One thing that helped was to add a positional encoding on the frequency axis.  Indeed, CNNs are good at learning translation independent representations, and here we don't want the model to be frequency independent.  I simply added a linear gradient on the frequency axis to all my crops.\n\nFor the rest my model is exactly what I used and shared in the previous bird song competition: https://www.kaggle.com/c/birdsong-recognition/discussion/183219  Just using the code I shared there was almost good enough to get 0.931.  The only differences are that I add noise as in the first solution in that competition.  I also did not use a no call class here, nor secondary labels.\n\nFor prediction I predict on slidding crops of  each test recording and take the overall max prediction.  This is slow as I need to do each of the 26 classes separately.  This is also maybe where I lost against others: my model cannot learn long range temporal patterns, nor class interactions.\n\nWith the above I entered high with an Efficient B0 model, and moved to 0.945 in few submissions with Efficientnet B3 . Then I got stuck for the remainder of the competition.  \n\nI was convinced that semi supervised learning was the key, and I implemented all sorts of methods, from Google (noisy student), Facebook, others (mean student).  They all improved weaker models but could not improve my best models.\n\nIn the last days I looked for external data with the hope that it would make a difference. Curating all this and identifying which species correspond to the species_id we have took some time and I only submitted models trained with it today.  They are in same range as previous ones unfortunately.  with a bit more time I am sure it could improve score, but I doubt it would be significant..\n\nFor matching species to species id I used my best model and predicted the external data. It would be interesting to see if I got this mapping right.  Here is what I converged to  :\n\n0                        E. gryllus\n1         Leutherodactylus brittoni\n2           Leptodactylu albilabris\n3                          E. coqui\n4                       E. hedricki\n5                 Setophaga angelae\n6          Melanerpes portoricensis\n7                  Coereba flaveola\n8                       E. locustus\n9                Margarops fuscatus\n10          Loxigilla portoricensis\n11                 Vireo altiloquus\n12                 E. portoricensis\n13                Megascops nudipes\n14                     E. richmondi\n15             Patagioenas squamosa\n16    Eleutherodactylus antillensis\n17                  Turdus plumbeus\n18                      E. unicolor\n19               Coccyzus vieilloti\n20                  Todus mexicanus\n21                    E  wightmanae\n22         Nesospingus speculiferus\n23          Spindalis portoricensis\n\nThe picture in the paper shared in the forum  helped to disambiguate few cases: https://reader.elsevier.com/reader/sd/pii/S1574954120300637  The paper also gives the list of species.  My final selected subs did not include models trained on external data, given they were not improving.\n\nThis concludes my experience in this competition. I am looking forward to see how so many teams passed me during the competition.  There is certainly a lot to be learned.\n\nEdit.  I am very pleased to get a solo gold in a deep learning competition.,  This is a first for me, and it was my goal here.\n\nEdit 2:  The models I trained last day with external data are actually better than the ones without. The best one has a private LB of 0.950 (5 folds). However, they are way better on private LB but not on public LB.  Selecting them would have been an act of faith.  And late submission show they are not good enough to change my rank.  No regrets then.\n\nEdit 3  Using Chris Deotte post processing. my best selected sub gets 0.7390 on private LB.  It means that PP was what I missed and that my modeling approach was good enough probably.  I'll definitely look at test prediction distribution from now on!",
    "1207601": "Very nice writeup. We used a very similar approach, which we really kicked into action at the last 8 days; will share details tomorrow... congrats on the solo medal!",
    "1207603": "Congrats @cpmpml . Great solo finish! I'm glad you won Gold. You helped many teams by showing them what was possible!",
    "1207608": "Great work. So many techniques that are close to what I tried, but just never put them together in the right way. \n\nIt seemed like the public pannloss that was floating around was trying to implement that masked loss like was mentioned in the paper, but was kind of implemented wrong.",
    "1207612": "Congrats on the solo gold! I have one question: did you crop y-axis (the frequency axis) too when doing training and inference?",
    "1207613": "tx.  Not only did I show it, but I shared I was using an image classification model.  Next time I'll share less maybe ;)\n\nCongrats on your solo gold too.",
    "1207614": "Congratulations! Could you share how exactly did you do the loss masking? I have tried this early in competition (and actually specifically inspired by your approach to secondary labels in BirdCall), but it backfired quite heavily to me (I only did time based crops, frequency cropping may be even more important, but intuitively I thought even without frequency cropping loss masking should make more sense...)",
    "1207615": "Tx.  My limited skills in computer vision may be what prevented me from getting a higher score.  Looking forward to your writeup.  I am sure I'll learn stuff.  And congrats as well on your result.",
    "1207616": "Tx.  I didn't look at public notebooks at all.  Maybe I should have...  It also looks like others used the same overall method, but got better results than me.",
    "1207617": "Yes.  I cropped by f_min and f_max.",
    "1207619": "Congrats, amazing job. We were indeed very impressed by the 931 start. I dont think you revealed anything with the image model though, so dont worry about that. Everyone uses image models in these types of problems.",
    "1207622": "cpmpml   \n\"is slow as I need to do each of the 26 classes separately\"\n\"Yes. I cropped by f_min and f_max.\"\n...\n\nsee my solution to solve that\nhttps://www.kaggle.com/c/rfcx-species-audio-detection/discussion/220309",
    "1207623": "Did the extra class for the songs help anything? I had it on the list since day 1 but never got around to try it. I think it was for low-frequency classes but would need to re-check.",
    "1207627": "I am not sure about what is not clear.  You compute the BCE loss for all targets (no reduction), then multiply by the mask (1 for the song class to be predicted, 0 for all other classes), then take the mean for the batch.",
    "1207628": "Congrats on your solo gold @cpmpml . Your entrance to the competition was inspirational 💪",
    "1207630": "Sorry, I did not use it here.  My bad.  Let me fix the writeup.",
    "1207637": "Congrats! The gradient idea is nice. I have also tried applying fixed grid lines on the images. That helped.",
    "1207650": "Congratulations @cpmpml for strong solo finish! 0.931 for the first sub was quite impressive! We learned a lot from you, thank you!",
    "1207651": "Exactly, sounds quite straightforward to me, so just checking whether I am not missing anything. So something like below, right?\n\n```\nclass RainforestLossMasked(nn.Module):\n    def __init__(self):\n        super().__init__()    \n        self.bce_prob  = nn.BCELoss(reduction='none')\n\n    def forward(self, out, target):                \n        loss = self.bce_prob(out, target)     \n        with torch.no_grad():\n            loss[target<1] = 0   \n        loss = loss.mean()\n        return loss\n```",
    "1207662": "Thanks.  You finished strongly too.  I am sure that if your studies had left you more time then you would have passed me.  Also, I guess that most teams reused your SED model from last competition.  This is also something to be proud of.",
    "1207664": "Grid may be more effective actually, I'll try it next time.  Congrats on your strong silver.",
    "1207672": "Congrats on 12th place and gold medal @cpmpml",
    "1207683": "Congrats! thank you for making this competition so much more competitive :) also, I wanted to ask how much did masking the loss and using a 26 classes model instead of 24 classes help independently?",
    "1207686": "How much Position Information Do Convolutional Neural Networks Encode?\nhttps://openreview.net/pdf?id=rJeB36NKvB\n\nIt turns out that CNN knows position of image. Because of effects of padding, CNN knows where the  input window is in the image, if the effective receptive field is large (e.g. same size of larger than the size of the input image). Hence adding position gradient is sometimes not necessary\n\nquote:\n\"Experiments reveal that positional information is available to a strong degree\"\n\n\"Results point to zero padding and borders as an anchor from which spatial information is derived and eventually propagated over the whole image as spatial abstraction occurs\"\n\n---\n\nI think there are also some interesting papers that experiment positional encoding used in transformer for CNN",
    "1207696": "I only tried 24 classes omce, and it was a bit worse.\n\nCongrats on your strong finish.  Looking forward to your writeup.  \n\nI wish I had waited two more week before submitting ;)\n\nJust kidding, I am sure others would have submitted high scores soon.",
    "1207697": "something like this:\n\n```\n                def use_tp_fp_binary_cross_entropy(logit, label, is_tp):\n                    batch_size, num_label = logit.shape\n                    batch_size = label.shape # e.g. [ 2,2,6,8, ... 23]\n                    batch_size = is_tp.shape # e.g. [ 1,0,0,1, ... 1]\n                 \n                    l = logit.gather(1, label.unsqueeze(1))\n                    l = l.reshape(-1)\n                    t = is_tp.float()\n\n                    p = torch.sigmoid(l)\n                    logp = - torch.log(torch.clamp(p, 1e-4, 1-1e-4))\n                    logn = - torch.log(torch.clamp(1-p, 1e-4, 1-1e-4))\n                    loss = t*logp +(1-t)*logn\n\n                    loss  = loss.mean()\n                    return loss\n\n\n\n```",
    "1207700": "You mentioned focal loss earlier in the competition. Did you ever try that in place of bce in your final solution? In my experience that has only ever been useful for extreme class imbalance problems like segmentation where there is a small number of positives to a large number of negatives. \n\nI feel like this is another one of the competitions where I had most of the ingredients for the top solutions out but couldn't quite find the right recipe to combine them all in. I tried the masked loss, frequency banded model that I implemented a bit differently than yours, but similar concept. I mentioned to our team at one point the linear gradient to show position more explicitly to the cnn.",
    "1207701": "Final question: did you validate against the crops with your image model or did you apply the same test time rolling validation? I'd assume you validated against the crops since the rolling window inference was slow.",
    "1207702": "Good question!\n\nYes I tried it at some point and it was a bit worse than vanilla bce.\n\nFocal loss only worked for me in segmentation tasks.  This is the original use case of focal loss unless mistaken.",
    "1207704": "I used 5 fold CV on my crops.  And I used roc-auc.",
    "1207708": "Thank you! will post a writeup soon. I am sure quite a few entered after looking at your 0.931 post and also many that already entered started to look for alternative solutions too haha!\n\nhow about masking the loss, how much did it help?",
    "1207715": "Congrats on your solo gold. Your discussion comments helped me many times.",
    "1207728": "Kind of kicking myself for not exploring this further. The first thing I sent to the team when I joined them was showing them a model I had that was kind of similar to yours. Instead of cropping though I simply masked out the regions outside of the frequency band and time that were irrelevant. So I would crop around the region of the time and then I would mask out the frequencies that were irrelevant to that prediction. \n\n![](https://i.imgur.com/tJPa7Sv.png)\n\nWould look something like this. I made it so they were long enough that I never had to do any cropping, only ever padding near the beginning or end of the audio. Would use similar procedure to you at test time, but I could stack all of them together fairly easily so it was 16(num_time_steps)*24(num_frequency_ranges), 128(num_mel_bins), 500 (time) and it was fairly fast to do inference. On the crops themselves I got validation performance of .985 at times, but when I applied the rolling window validation I would get poor results like .6-.7. \n\nI tried the masked loss, but I did not apply it to this specific model and I never made a submission with it because the windowed validation looked poor. I considered the linear gradient to give position but ended up not using it in this setup because I figured the model already had its relative position based on the amount of 0 padding above and below the unmasked signal. Might have to fiddle with that and see if it was actually good.",
    "1207730": "Thank you for the great write-up and congratz on solo gold!\nI thought `Coereba flaveola` was `species_id=1` but I was not sure.",
    "1207774": "Congrats on your medal. As always, thanks for sharing  \nSorry for my lack of understanding, when you mention\n\n> For prediction I predict on slidding crops of each test recording and take the overall max prediction. This is slow as I need to do each of the 26 classes separately. This is also maybe where I lost against others: my model cannot learn long range temporal patterns, nor class interactions.\n\nDoes this mean u train 26 models for each class? From what i am reading, if you are masking your bce, you are able to output multi label predictions for each sample so where do the 26 classes separately come from? Will be good to answer!",
    "1207786": "cpmpml  If we crop by the f_min and f_max, doesn't that mean the image generated per species is not comparable anymore (since the y axis are representing different frequency region).  Would you explain how can that work ?\nAnd big congratulations on the solo gold !",
    "1208111": "I train one model, but crops are different for each class.  I need to apply the model to each crop separately.",
    "1208118": ">  Instead of cropping though I simply masked out the regions outside of the frequency band and time that were irrelevant. So I would crop around the region of the time and then I would mask out the frequencies that were irrelevant to that prediction. \n\nThat's what I did.  Crop then pad.  Sorry if this is not clear enough in my post.\n\nIt looks like you had the same idea as me ;)",
    "1208120": "> I thought Coereba flaveola was species_id=1 but I was not sure.\n\nBased on what data?",
    "1208132": "I just checked, assuming all other targets are 0 instead of masking costs 0.01 on LB.  I tried both BCE and softmax in that case.",
    "1208138": "> It turns out that CNN knows position of image. Because of effects of padding\n\nThis is what i thought as well.  But I explored, for instance I read the paper you quote and others on positional encoding. It turns out that adding some positional encoding helps.",
    "1208140": "I did ask you for how your method differs from mine.  I hope you will explain.\n\nTiny crops of code is not the same as a written explanation.  \n\nFrom what I see, you crop by f_min and f_max, hence I don't see how it is different from what I did.",
    "1208158": "Thanks for sharing your codes.  Mine is different.\n\nGiven that each crop has a single class target, I just use a single target overall.  I know what class I am predicting when I crop, hence I don't need to give it to the model.  My code is then very simple:\n\n```\n        loss_fct = nn.BCEWithLogitsLoss()\n        logits = self.head(x)\n        mask = input_dict['mask']\n        preds = (logits * mask).sum(-1, keepdim=True)\n        mono_targets = input_dict['mono_target']\n        loss = loss_fct(preds, mono_targets)\n```\n\nThis does no work:\n\n```\n        with torch.no_grad():\n            loss[target<1] = 0   \n\n```\nbecause FP crops have  a target of 0 and you want the model to learn that.  You must mask only the loss for other classes.\n\nHow is this\n\n```\n                    p = torch.sigmoid(l)\n                    logp = - torch.log(torch.clamp(p, 1e-4, 1-1e-4))\n                    logn = - torch.log(torch.clamp(1-p, 1e-4, 1-1e-4))\n                    loss = t*logp +(1-t)*logn\n\n                    loss  = loss.mean()\n\n```\n\ndifferent from \n\n```\n    loss_fct = nn.BCEWithLogitsLoss()\n    loss = loss_fct(l, t)\n```",
    "1208163": ">  doesn't that mean the image generated per species is not comparable anymore \n\nYes, this is why I need to perform inference separately for each class.",
    "1208375": "I spent a few days trying almost exactly the same idea as yours, although I couldn't make it work, it looks like it wasn't that bad after all :)\nAnd congratz on the impressive performance !",
    "1208618": "Thanks.  We all tried stuff that failed for us but worked for others apparently.  Mine was pseudo labeling.",
    "1208631": "S7 is correct for that one",
    "1208639": "CPMP. I have completed the writeup and outlined the similarities and differences between your and my approaches. under the discussion \"Devil is in detail but it seems we have a similar method. I welcome comments about what we did differently .... \" Please check it.\n\nyou mentioned here that \"... models probably, like SED or Transformers.\". I am curious how will you apply transformer here. Can you give some hints?",
    "1208677": "Do you mean for each class slide for x axis(time) and fixed crop for y axis (as we know frequency)? So we move along x and take max?\n\nThank you and congratulations!",
    "1208710": "Congrats on the solo gold! I have two quick questions: \n1. How do you crop the frequencies, do you specify the `fmin` and `fmax` in `librosa.feature.melspectrogram`, and the frequency sizes of generated features are still controlled by `n_mels`; or you use an overall initial range and crop the generated features on frequency axis later? If it's the latter, how to do it, could you provide a few lines of code? And how to make them in batch, also padding like time axis?\n2. Did you use the FP exactly as you use TP? Is there any trick? Cause the labels in FP are wrong and I have tried to use them, it really harms the performance.",
    "1208857": "Thanks for answering. \n\nMy understanding is that yes, the sliding must be across the time axis for each frequency axis associated with the classes.",
    "1208880": "why sum? not just max?",
    "1208916": ">  I guess that most teams reused your SED model from last competition. This is also something to be proud of.\n\nHappy to know many teams used SED model. Unfortunately we couldn't make it work better than other kind of models, but it turned out some top performing teams successfully used it. I feel like I found tons of things to learn from this competition.",
    "1208977": "1. I cropped the spectrogram after generating it.  I used fmin=90 and fmax=14000 for generating spectrogram.  To find which pixels to crop, librosa has utilities functions that map frequencies to the number of mel.\n\n2. I used FP and TP the same way.  They give a 0 or a 1 label for an image.\n\nWhat is wrong with FP labels?  They were good enough to get me a gold medal ;)",
    "1209018": "> Do you mean for each class slide for x axis(time) and fixed crop for y axis (as we know frequency)? So we move along x and take max?\n\nyes.",
    "1209052": "I feel like one thing that people are not understanding with the FPs is that they are actually our best 0's. Some people have been making the mistake of using them as 1's which is explicitly what we dont want.",
    "1209074": "> they are actually our best 0's.\n\n@ryches Oh! So I should train my model to predict the species listed in FP do not appear instead appearing, is it how you did @cpmpml ? \nIt's so bad I figure this out so late...",
    "1209166": "Yes, the false positives are regions that showed up using the template matching but an expert reviewed and said specifically the species was not present. So training your model to detect false positives the same as true positives is not what we want to do",
    "1209170": "wow that's huge for me thanks for checking and sharing!",
    "1209304": "> they are actually our best 0's. \n\nIndeed.  I'd even say they are our ONLY 0's.",
    "1209311": "Tx, will check your writeup next.\n\nFor transformers I refer to a speech to text model that was released recently by facebook I think.  I haven't looked in detail, maybe it is not applicable because ground truth is too sparse here.",
    "1209314": "philippsinger Thanks, and thanks for addressing my concerns about oversharing ;)\n\nIt was a bit stressful to wake up every morning and see someone pass me while I was stuck.  But the story has a happy ending, everything is good.\n\nCongrats on your nth win in a row, this is amazing.",
    "1212813": "I didn't think of post processing but I did think of pseudo labeling during the competition, yet could not make it work.  After reading all the writeups where teams who passed me successfully used PL I revisited what i did.  The issue was that instead of randomly sample crops for PL, I imposed a distribution that matches training samples, i.e. same number of positive pseudo labels per class.  When I remove this bias then pseudo labeling works.  In my first experiment, a single model gets a 0.01 boost on public and private LB.  With tuning and iterations I now see how I could have moved higher.  \n\nI am not sure why I imposed this sampling bias.  I'll try to be more careful next time.",
    "1330750": "Thanks for the excellent and detailed write-up! I hope I wont be too late to ask a few questions wrt ur approach:\n1. u mentioned u made the time-axis of each crop the same (by cropping/ padding), how did handle the differences in frequency-axis? (as in different species has different height)\n2. may I learn more abt how u applied positional encoding along frequency-axis?  \n\nThanks!"
  },
  "source": "meta"
}