{
  "id": 183219,
  "title": "18th place solution: efficientnet b3",
  "url": "/competitions/birdsong-recognition/writeups/birdcall-of-duty-18th-place-solution-efficientnet-",
  "author_name": "",
  "post_date": "2021-04-04T04:14:35.247Z",
  "votes": 68,
  "comment_count": 18,
  "views": 0,
  "content": "<p>First of all I want to thank the host and Kaggle for this very challenging competition, with a truly hidden test set.  A bit too hidden maybe, but thanks to the community latecomers like us could get a submission template up and running in a couple of days.  I also want to thank my team mate Kazuki, without whom I would probably have given up after many failed attempts to beat the public notebook baseline…</p>\n<p><strong>Overview</strong></p>\n<p>Our best subs are single efficientnet models trained on log mel spectrograms.  For our baseline I started from scratch rather than reusing the excellent baselines that were available.  Reason is that I enter Kaggle competition to learn, and I learn more when I try from scratch than when I  modify someone else' code.  We then evolved that baseline as we could in the two weeks we had before the end of competition.</p>\n<p>Given the overall approach is well known and probably used by most participants here I will only discuss the items that may be a bit different from what others did.</p>\n<p><strong>Training data clips</strong></p>\n<p>It was clear from host that training on random 5 second clips had a drawback: some of the clips may not contain the target bird.  We then used a simple hypothesis: clips were stripped, i.e. periods without a song at the beginning or at the end were removed for the sake of reducing storage needs.  We therefore trained on first 5 seconds or last 5 seconds of clips, assuming these would contain the target bird.  We preprocessed all data to be sampled at 32 kHz.</p>\n<p><strong>Noise</strong><br>\nWe added noise extracted from the two test sequences made available, a bit like what Theo Viel did.  But we used the meta data to extract sub sequences without bird call, then we merged these sequence with a smooth transition between them.  We then added a random clip of the merges sequences to our training clips</p>\n<p><strong>No Call</strong><br>\nWe added the freefield1010 clips that were labelled as nocall to our training data.  We added a 265th class to represent the no call.  As a result our model could predict both one or more birds, and a nocall. Adding this data and the nocall class was probably the most important single improvement we saw in CvV and LB scores.  It is what led us to pass the public notebook baseline.</p>\n<p><strong>Multi Bird Clips</strong><br>\nThe main documented difference between train and test data is that train is a multi class data while test is a multi label data.  Therefore we implemented a mixup variant were up to 3 clips could be merged.  This is not really mixup as the target for the merged clip is the maximum of the targets of each merged clip.</p>\n<p><strong>Secondary labels</strong><br>\nPrimary labels were noisy, but secondary labels were even noisier.  As a result we masked the loss for secondary labels as we didn't want to force the model to learn a presence or an absence when we don't know.  We therefore defined a secondary mask that nullifies the BCE loss for secondary labels.  For instance, assuming only 3 ebird_code b0, b1, and b2, and a clip with primary label b0 and secondary label b1, then these two target values are possible:</p>\n<p>[1, 0, 0]</p>\n<p>[1, 1, 0]</p>\n<p>The secondary mask is therefore:</p>\n<p>[1, 0, 1]</p>\n<p>For merged clips, a target is masked if it it not one of the primary labels and if it is one of the secondary labels.</p>\n<p><strong>Loss Function</strong></p>\n<p>We use binary cross entropy on one hot encoding of ebird codes.  Using bce rather than softmax makes sense as bce extends to multi label seamlessly.  We tried dice loss to directly optimize F! score, but for some reason this led to very strong overfitting.</p>\n<p><strong>Class Weights</strong></p>\n<p>The number of record per species is not always 100.  In order to not penalize the less frequent one we use class weights inversely proportional to class frequencies.  And for the nocall class we set it to 1 even though it was way more frequent than each bird classes to make sure the model learns about nocall correctly.</p>\n<p><strong>Model</strong></p>\n<p>Our best model was efficientnet on log mel spectrograms.  We resized images to be twice the size of effnet images: 240x480 for effnet b1, 260x520 for effnet b2, and 300x600 for effnet b3.  We started from efficientnet_pytorch pretrained models.  We tried the attention head from PANNs models but it led to severe overfitting.  I am not sure why to be honest, maybe we did something wrong.</p>\n<p><strong>Training</strong><br>\nNothings fancy, adam optimizer and cosine scheduler with 60 epochs.  In general last epoch was the one with best score and we used last epoch weights for scoring.</p>\n<p><strong>Log Mel Spectrogram</strong></p>\n<p>Nothing fancy, except that we saw a lot of power in low frequencies of the first spectrograms we created. As a result we clipped frequency to be at least 300 Hz. We also clipped them to be below 16 kHz given test data was sampled at twice that frequency.  </p>\n<p><strong>Augmentations</strong></p>\n<p>Time and pitch variations were implements in a very simple way: modify the length of the clipped sequence, and modify the sampling rate, without modifying the data itself.  For instance, we would read 5.2 seconds of a clip instead of 5 seconds, and we could tell librosa that the sampling rate was 0.9*32 kHz.  We then compute the hop so that the number of stft is equal to the image width for the effnet model we are training.  We also computed the number of mel bins to be equal to the height of the image.  As a result we never had to resample data nor resize images, which speed up training and inferencing quite a bit.  There was an issue with that still: this led us to use a high nftt value of 2048 which lead to poor time resolution.  We ended up with nfft of 1048 and a resize of the images in the height dimension.</p>\n<p><strong>Cross Validation</strong></p>\n<p>We started with cross validating on our training data (first or last 5 seconds of clips) but it was rapidly clear that it was not representative of the LB score.  And given we had very few submissions, we could not perform LB probing.  We therefore spent some time during last week to create a CV score that was correlated with the public LB.  Our CV score is computed using multi bird clips for 46% of the score, and nocall clips for 54% of the score.  We added warblrb and birdvox nocall clips to the freefield1010 for valuating no call performance.  We tuned the proportion of each possible number of clips in multi bird clips, and the amount of noise until we could find a relatively good CV LB relationship, see the picture below: x is cv, y is public lb.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F75976%2F39f916bec9be955cfa2a8fd5e8384b3b%2Fcv_lb.png?generation=1600261413279842&amp;alt=media\" alt=\"\"></p>\n<p>The correlation with private LB is also quite good:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F75976%2F02c89402f0dfd6e9069f36e67fc32e85%2Fcv_private.png?generation=1600261852596913&amp;alt=media\" alt=\"\"></p>\n<p>We then used it to guide our last few submissions training and thresholding.  Our CV said 0.7 was best, and late submissions proved it was right.  We ended up selecting our 2 best CV submissions, which were also the 2 best public LB submissions, and also the two best private Lb submissions.</p>\n<p><strong>Conclusion</strong></p>\n<p>These were probably the most intensive two weeks I had on Kaggle since a long time.  My first two subs 14 days ago were a failure, then a LB score of 0. Kazuki had started a bit earlier, but not much. I am very happy about where we landed, and I am not sure we would have done much better with few more days.  Maybe using effnet b4 or b5 wold have moved us higher but I am not sure.  I am looking for gold medalist solutions to see what we missed.  I'm sure I'll learn quite a bit.</p>",
  "messages": [
    {
      "id": "1012236",
      "postDate": "09/16/2020 01:33:32",
      "content": "<p>First of all I want to thank the host and Kaggle for this very challenging competition, with a truly hidden test set.  A bit too hidden maybe, but thanks to the community latecomers like us could get a submission template up and running in a couple of days.  I also want to thank my team mate Kazuki, without whom I would probably have given up after many failed attempts to beat the public notebook baseline…</p>\n<p><strong>Overview</strong></p>\n<p>Our best subs are single efficientnet models trained on log mel spectrograms.  For our baseline I started from scratch rather than reusing the excellent baselines that were available.  Reason is that I enter Kaggle competition to learn, and I learn more when I try from scratch than when I  modify someone else' code.  We then evolved that baseline as we could in the two weeks we had before the end of competition.</p>\n<p>Given the overall approach is well known and probably used by most participants here I will only discuss the items that may be a bit different from what others did.</p>\n<p><strong>Training data clips</strong></p>\n<p>It was clear from host that training on random 5 second clips had a drawback: some of the clips may not contain the target bird.  We then used a simple hypothesis: clips were stripped, i.e. periods without a song at the beginning or at the end were removed for the sake of reducing storage needs.  We therefore trained on first 5 seconds or last 5 seconds of clips, assuming these would contain the target bird.  We preprocessed all data to be sampled at 32 kHz.</p>\n<p><strong>Noise</strong><br>\nWe added noise extracted from the two test sequences made available, a bit like what Theo Viel did.  But we used the meta data to extract sub sequences without bird call, then we merged these sequence with a smooth transition between them.  We then added a random clip of the merges sequences to our training clips</p>\n<p><strong>No Call</strong><br>\nWe added the freefield1010 clips that were labelled as nocall to our training data.  We added a 265th class to represent the no call.  As a result our model could predict both one or more birds, and a nocall. Adding this data and the nocall class was probably the most important single improvement we saw in CvV and LB scores.  It is what led us to pass the public notebook baseline.</p>\n<p><strong>Multi Bird Clips</strong><br>\nThe main documented difference between train and test data is that train is a multi class data while test is a multi label data.  Therefore we implemented a mixup variant were up to 3 clips could be merged.  This is not really mixup as the target for the merged clip is the maximum of the targets of each merged clip.</p>\n<p><strong>Secondary labels</strong><br>\nPrimary labels were noisy, but secondary labels were even noisier.  As a result we masked the loss for secondary labels as we didn't want to force the model to learn a presence or an absence when we don't know.  We therefore defined a secondary mask that nullifies the BCE loss for secondary labels.  For instance, assuming only 3 ebird_code b0, b1, and b2, and a clip with primary label b0 and secondary label b1, then these two target values are possible:</p>\n<p>[1, 0, 0]</p>\n<p>[1, 1, 0]</p>\n<p>The secondary mask is therefore:</p>\n<p>[1, 0, 1]</p>\n<p>For merged clips, a target is masked if it it not one of the primary labels and if it is one of the secondary labels.</p>\n<p><strong>Loss Function</strong></p>\n<p>We use binary cross entropy on one hot encoding of ebird codes.  Using bce rather than softmax makes sense as bce extends to multi label seamlessly.  We tried dice loss to directly optimize F! score, but for some reason this led to very strong overfitting.</p>\n<p><strong>Class Weights</strong></p>\n<p>The number of record per species is not always 100.  In order to not penalize the less frequent one we use class weights inversely proportional to class frequencies.  And for the nocall class we set it to 1 even though it was way more frequent than each bird classes to make sure the model learns about nocall correctly.</p>\n<p><strong>Model</strong></p>\n<p>Our best model was efficientnet on log mel spectrograms.  We resized images to be twice the size of effnet images: 240x480 for effnet b1, 260x520 for effnet b2, and 300x600 for effnet b3.  We started from efficientnet_pytorch pretrained models.  We tried the attention head from PANNs models but it led to severe overfitting.  I am not sure why to be honest, maybe we did something wrong.</p>\n<p><strong>Training</strong><br>\nNothings fancy, adam optimizer and cosine scheduler with 60 epochs.  In general last epoch was the one with best score and we used last epoch weights for scoring.</p>\n<p><strong>Log Mel Spectrogram</strong></p>\n<p>Nothing fancy, except that we saw a lot of power in low frequencies of the first spectrograms we created. As a result we clipped frequency to be at least 300 Hz. We also clipped them to be below 16 kHz given test data was sampled at twice that frequency.  </p>\n<p><strong>Augmentations</strong></p>\n<p>Time and pitch variations were implements in a very simple way: modify the length of the clipped sequence, and modify the sampling rate, without modifying the data itself.  For instance, we would read 5.2 seconds of a clip instead of 5 seconds, and we could tell librosa that the sampling rate was 0.9*32 kHz.  We then compute the hop so that the number of stft is equal to the image width for the effnet model we are training.  We also computed the number of mel bins to be equal to the height of the image.  As a result we never had to resample data nor resize images, which speed up training and inferencing quite a bit.  There was an issue with that still: this led us to use a high nftt value of 2048 which lead to poor time resolution.  We ended up with nfft of 1048 and a resize of the images in the height dimension.</p>\n<p><strong>Cross Validation</strong></p>\n<p>We started with cross validating on our training data (first or last 5 seconds of clips) but it was rapidly clear that it was not representative of the LB score.  And given we had very few submissions, we could not perform LB probing.  We therefore spent some time during last week to create a CV score that was correlated with the public LB.  Our CV score is computed using multi bird clips for 46% of the score, and nocall clips for 54% of the score.  We added warblrb and birdvox nocall clips to the freefield1010 for valuating no call performance.  We tuned the proportion of each possible number of clips in multi bird clips, and the amount of noise until we could find a relatively good CV LB relationship, see the picture below: x is cv, y is public lb.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F75976%2F39f916bec9be955cfa2a8fd5e8384b3b%2Fcv_lb.png?generation=1600261413279842&amp;alt=media\" alt=\"\"></p>\n<p>The correlation with private LB is also quite good:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F75976%2F02c89402f0dfd6e9069f36e67fc32e85%2Fcv_private.png?generation=1600261852596913&amp;alt=media\" alt=\"\"></p>\n<p>We then used it to guide our last few submissions training and thresholding.  Our CV said 0.7 was best, and late submissions proved it was right.  We ended up selecting our 2 best CV submissions, which were also the 2 best public LB submissions, and also the two best private Lb submissions.</p>\n<p><strong>Conclusion</strong></p>\n<p>These were probably the most intensive two weeks I had on Kaggle since a long time.  My first two subs 14 days ago were a failure, then a LB score of 0. Kazuki had started a bit earlier, but not much. I am very happy about where we landed, and I am not sure we would have done much better with few more days.  Maybe using effnet b4 or b5 wold have moved us higher but I am not sure.  I am looking for gold medalist solutions to see what we missed.  I'm sure I'll learn quite a bit.</p>",
      "rawMarkdown": "First of all I want to thank the host and Kaggle for this very challenging competition, with a truly hidden test set.  A bit too hidden maybe, but thanks to the community latecomers like us could get a submission template up and running in a couple of days.  I also want to thank my team mate Kazuki, without whom I would probably have given up after many failed attempts to beat the public notebook baseline...\n\n**Overview**\n\nOur best subs are single efficientnet models trained on log mel spectrograms.  For our baseline I started from scratch rather than reusing the excellent baselines that were available.  Reason is that I enter Kaggle competition to learn, and I learn more when I try from scratch than when I  modify someone else' code.  We then evolved that baseline as we could in the two weeks we had before the end of competition.\n\nGiven the overall approach is well known and probably used by most participants here I will only discuss the items that may be a bit different from what others did.\n\n**Training data clips**\n\nIt was clear from host that training on random 5 second clips had a drawback: some of the clips may not contain the target bird.  We then used a simple hypothesis: clips were stripped, i.e. periods without a song at the beginning or at the end were removed for the sake of reducing storage needs.  We therefore trained on first 5 seconds or last 5 seconds of clips, assuming these would contain the target bird.  We preprocessed all data to be sampled at 32 kHz.\n\n**Noise**\nWe added noise extracted from the two test sequences made available, a bit like what Theo Viel did.  But we used the meta data to extract sub sequences without bird call, then we merged these sequence with a smooth transition between them.  We then added a random clip of the merges sequences to our training clips\n\n**No Call**\nWe added the freefield1010 clips that were labelled as nocall to our training data.  We added a 265th class to represent the no call.  As a result our model could predict both one or more birds, and a nocall. Adding this data and the nocall class was probably the most important single improvement we saw in CvV and LB scores.  It is what led us to pass the public notebook baseline.\n\n**Multi Bird Clips**\nThe main documented difference between train and test data is that train is a multi class data while test is a multi label data.  Therefore we implemented a mixup variant were up to 3 clips could be merged.  This is not really mixup as the target for the merged clip is the maximum of the targets of each merged clip.\n\n**Secondary labels**\nPrimary labels were noisy, but secondary labels were even noisier.  As a result we masked the loss for secondary labels as we didn't want to force the model to learn a presence or an absence when we don't know.  We therefore defined a secondary mask that nullifies the BCE loss for secondary labels.  For instance, assuming only 3 ebird_code b0, b1, and b2, and a clip with primary label b0 and secondary label b1, then these two target values are possible:\n\n[1, 0, 0]\n\n[1, 1, 0]\n\nThe secondary mask is therefore:\n\n[1, 0, 1]\n\nFor merged clips, a target is masked if it it not one of the primary labels and if it is one of the secondary labels.\n\n**Loss Function**\n\nWe use binary cross entropy on one hot encoding of ebird codes.  Using bce rather than softmax makes sense as bce extends to multi label seamlessly.  We tried dice loss to directly optimize F! score, but for some reason this led to very strong overfitting.\n\n**Class Weights**\n\nThe number of record per species is not always 100.  In order to not penalize the less frequent one we use class weights inversely proportional to class frequencies.  And for the nocall class we set it to 1 even though it was way more frequent than each bird classes to make sure the model learns about nocall correctly.\n\n**Model**\n\nOur best model was efficientnet on log mel spectrograms.  We resized images to be twice the size of effnet images: 240x480 for effnet b1, 260x520 for effnet b2, and 300x600 for effnet b3.  We started from efficientnet_pytorch pretrained models.  We tried the attention head from PANNs models but it led to severe overfitting.  I am not sure why to be honest, maybe we did something wrong.\n\n**Training**\nNothings fancy, adam optimizer and cosine scheduler with 60 epochs.  In general last epoch was the one with best score and we used last epoch weights for scoring.\n\n\n**Log Mel Spectrogram**\n\nNothing fancy, except that we saw a lot of power in low frequencies of the first spectrograms we created. As a result we clipped frequency to be at least 300 Hz. We also clipped them to be below 16 kHz given test data was sampled at twice that frequency.  \n\n**Augmentations**\n\nTime and pitch variations were implements in a very simple way: modify the length of the clipped sequence, and modify the sampling rate, without modifying the data itself.  For instance, we would read 5.2 seconds of a clip instead of 5 seconds, and we could tell librosa that the sampling rate was 0.9*32 kHz.  We then compute the hop so that the number of stft is equal to the image width for the effnet model we are training.  We also computed the number of mel bins to be equal to the height of the image.  As a result we never had to resample data nor resize images, which speed up training and inferencing quite a bit.  There was an issue with that still: this led us to use a high nftt value of 2048 which lead to poor time resolution.  We ended up with nfft of 1048 and a resize of the images in the height dimension.\n\n**Cross Validation**\n\nWe started with cross validating on our training data (first or last 5 seconds of clips) but it was rapidly clear that it was not representative of the LB score.  And given we had very few submissions, we could not perform LB probing.  We therefore spent some time during last week to create a CV score that was correlated with the public LB.  Our CV score is computed using multi bird clips for 46% of the score, and nocall clips for 54% of the score.  We added warblrb and birdvox nocall clips to the freefield1010 for valuating no call performance.  We tuned the proportion of each possible number of clips in multi bird clips, and the amount of noise until we could find a relatively good CV LB relationship, see the picture below: x is cv, y is public lb.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F75976%2F39f916bec9be955cfa2a8fd5e8384b3b%2Fcv_lb.png?generation=1600261413279842&alt=media)\n\nThe correlation with private LB is also quite good:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F75976%2F02c89402f0dfd6e9069f36e67fc32e85%2Fcv_private.png?generation=1600261852596913&alt=media)\n\nWe then used it to guide our last few submissions training and thresholding.  Our CV said 0.7 was best, and late submissions proved it was right.  We ended up selecting our 2 best CV submissions, which were also the 2 best public LB submissions, and also the two best private Lb submissions.\n\n**Conclusion**\n\nThese were probably the most intensive two weeks I had on Kaggle since a long time.  My first two subs 14 days ago were a failure, then a LB score of 0. Kazuki had started a bit earlier, but not much. I am very happy about where we landed, and I am not sure we would have done much better with few more days.  Maybe using effnet b4 or b5 wold have moved us higher but I am not sure.  I am looking for gold medalist solutions to see what we missed.  I'm sure I'll learn quite a bit.",
      "votes": null
    },
    {
      "id": "1012259",
      "postDate": "09/16/2020 02:05:51",
      "content": "<p>Congratulations <a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a> garu</p>",
      "rawMarkdown": "Congratulations @cpmpml garu",
      "votes": null
    },
    {
      "id": "1012438",
      "postDate": "09/16/2020 05:14:38",
      "content": "<blockquote>\n  <p>These were probably the most intensive two weeks I had on Kaggle since a long time.</p>\n</blockquote>\n<p>I quite agree. Congratulations, and thanks for sharing your approach.</p>",
      "rawMarkdown": "> These were probably the most intensive two weeks I had on Kaggle since a long time.\n\nI quite agree. Congratulations, and thanks for sharing your approach.",
      "votes": null
    },
    {
      "id": "1012440",
      "postDate": "09/16/2020 05:17:04",
      "content": "<p>Congratulations <a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a> and thanks for the lovely writeup  , time and again you always inspire me and one of the key things that I have learned from you is to write your own code and enter any competition for learning ,and one day medals will start to follow . Although writing my own code failed me during this competition but the learning was immense <br>\nThanks</p>",
      "rawMarkdown": "Congratulations @cpmpml and thanks for the lovely writeup  , time and again you always inspire me and one of the key things that I have learned from you is to write your own code and enter any competition for learning ,and one day medals will start to follow . Although writing my own code failed me during this competition but the learning was immense \nThanks",
      "votes": null
    },
    {
      "id": "1012462",
      "postDate": "09/16/2020 05:34:19",
      "content": "<p>Masking secondary is really an interesting approach to reduce the label noise!<br>\nCool stuff. I wonder why nocall as additional class didnt work for me</p>",
      "rawMarkdown": "Masking secondary is really an interesting approach to reduce the label noise!\nCool stuff. I wonder why nocall as additional class didnt work for me",
      "votes": null
    },
    {
      "id": "1012723",
      "postDate": "09/16/2020 08:49:16",
      "content": "<p>nocall worked for us maybe because of class weights.  I updated the post to describe what we did.</p>",
      "rawMarkdown": "nocall worked for us maybe because of class weights.  I updated the post to describe what we did.",
      "votes": null
    },
    {
      "id": "1012725",
      "postDate": "09/16/2020 08:49:47",
      "content": "<p>Thank you.  yes, progress is slower if you write your own code but you'll go farther.</p>",
      "rawMarkdown": "Thank you.  yes, progress is slower if you write your own code but you'll go farther.",
      "votes": null
    },
    {
      "id": "1012726",
      "postDate": "09/16/2020 08:50:36",
      "content": "<p>Thanks.  Indeed, a number of late comers did well here, some in gold unless mistaken.  I didn't track precisely hence I prefer not to name people in case I forget some ;)  I hope you'll share your approach as well!  Congrats on the result even if you dropped a bit in private LB.  I hope you're not too disappointed.</p>",
      "rawMarkdown": "Thanks.  Indeed, a number of late comers did well here, some in gold unless mistaken.  I didn't track precisely hence I prefer not to name people in case I forget some ;)  I hope you'll share your approach as well!  Congrats on the result even if you dropped a bit in private LB.  I hope you're not too disappointed.",
      "votes": null
    },
    {
      "id": "1012728",
      "postDate": "09/16/2020 08:52:03",
      "content": "<p>Thank you.</p>",
      "rawMarkdown": "Thank you.",
      "votes": null
    },
    {
      "id": "1013021",
      "postDate": "09/16/2020 13:13:29",
      "content": "<p>I was never worried, you ended in the top as usual ☝️👍</p>",
      "rawMarkdown": "I was never worried, you ended in the top as usual ☝️👍",
      "votes": null
    },
    {
      "id": "1013030",
      "postDate": "09/16/2020 13:19:39",
      "content": "<p>Indeed, I was way more worried than you were ;)</p>",
      "rawMarkdown": "Indeed, I was way more worried than you were ;)",
      "votes": null
    },
    {
      "id": "1013180",
      "postDate": "09/16/2020 14:51:50",
      "content": "<p>Well done</p>",
      "rawMarkdown": "Well done",
      "votes": null
    },
    {
      "id": "1013218",
      "postDate": "09/16/2020 15:10:58",
      "content": "<p>Great to see your effort..</p>",
      "rawMarkdown": "Great to see your effort..",
      "votes": null
    },
    {
      "id": "1013349",
      "postDate": "09/16/2020 16:50:19",
      "content": "<p>I don't think that I did as much as some people ended up in the gold range, so nothing really regret about, probably just a little bit.</p>",
      "rawMarkdown": "I don't think that I did as much as some people ended up in the gold range, so nothing really regret about, probably just a little bit.",
      "votes": null
    },
    {
      "id": "1013409",
      "postDate": "09/16/2020 17:25:00",
      "content": "<p>Congrats <a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a> on a strong finish.<br>\nNice idea with regard to training on either first 5 or last 5 seconds part of the clip.</p>\n<p>Quick question on noise - What was the noise level you used ?</p>",
      "rawMarkdown": "Congrats @cpmpml on a strong finish.\nNice idea with regard to training on either first 5 or last 5 seconds part of the clip.\n\nQuick question on noise - What was the noise level you used ?",
      "votes": null
    },
    {
      "id": "1013479",
      "postDate": "09/16/2020 17:57:16",
      "content": "<p>Thanks.  In the writeup:</p>\n<blockquote>\n  <p>We added noise extracted from the two test sequences made available, a bit like what Theo Viel did. But we used the meta data to extract sub sequences without bird call, then we merged these sequence with a smooth transition between them. We then added a random clip of the merges sequences to our training clips</p>\n</blockquote>\n<p>We multiplied the noise by a random number between 0.2 and 1.</p>",
      "rawMarkdown": "Thanks.  In the writeup:\n\n> We added noise extracted from the two test sequences made available, a bit like what Theo Viel did. But we used the meta data to extract sub sequences without bird call, then we merged these sequence with a smooth transition between them. We then added a random clip of the merges sequences to our training clips\n\nWe multiplied the noise by a random number between 0.2 and 1.",
      "votes": null
    },
    {
      "id": "1013906",
      "postDate": "09/17/2020 03:54:45",
      "content": "<p>It's nice to see Efficient Models nowadays. </p>",
      "rawMarkdown": "It's nice to see Efficient Models nowadays.",
      "votes": null
    },
    {
      "id": "1019409",
      "postDate": "09/20/2020 12:10:42",
      "content": "<p>Congrats .. Thanks for sharing </p>",
      "rawMarkdown": "Congrats .. Thanks for sharing",
      "votes": null
    },
    {
      "id": "1020499",
      "postDate": "09/21/2020 07:51:02",
      "content": "<p>Nice writeup. Impressive. </p>",
      "rawMarkdown": "Nice writeup. Impressive.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1012259,
      "author_name": "gopidurgaprasad",
      "author_url": "",
      "post_date": "09/16/2020 02:05:51",
      "content": "<p>Congratulations <a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a> garu</p>",
      "votes": null,
      "replies": [
        {
          "id": 1012728,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "09/16/2020 08:52:03",
          "content": "<p>Thank you.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1012438,
      "author_name": "iafoss",
      "author_url": "",
      "post_date": "09/16/2020 05:14:38",
      "content": "<blockquote>\n  <p>These were probably the most intensive two weeks I had on Kaggle since a long time.</p>\n</blockquote>\n<p>I quite agree. Congratulations, and thanks for sharing your approach.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1012726,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "09/16/2020 08:50:36",
          "content": "<p>Thanks.  Indeed, a number of late comers did well here, some in gold unless mistaken.  I didn't track precisely hence I prefer not to name people in case I forget some ;)  I hope you'll share your approach as well!  Congrats on the result even if you dropped a bit in private LB.  I hope you're not too disappointed.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1013349,
          "author_name": "iafoss",
          "author_url": "",
          "post_date": "09/16/2020 16:50:19",
          "content": "<p>I don't think that I did as much as some people ended up in the gold range, so nothing really regret about, probably just a little bit.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1012440,
      "author_name": "tanulsingh077",
      "author_url": "",
      "post_date": "09/16/2020 05:17:04",
      "content": "<p>Congratulations <a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a> and thanks for the lovely writeup  , time and again you always inspire me and one of the key things that I have learned from you is to write your own code and enter any competition for learning ,and one day medals will start to follow . Although writing my own code failed me during this competition but the learning was immense <br>\nThanks</p>",
      "votes": null,
      "replies": [
        {
          "id": 1012725,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "09/16/2020 08:49:47",
          "content": "<p>Thank you.  yes, progress is slower if you write your own code but you'll go farther.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1012462,
      "author_name": "yaroshevskiy",
      "author_url": "",
      "post_date": "09/16/2020 05:34:19",
      "content": "<p>Masking secondary is really an interesting approach to reduce the label noise!<br>\nCool stuff. I wonder why nocall as additional class didnt work for me</p>",
      "votes": null,
      "replies": [
        {
          "id": 1012723,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "09/16/2020 08:49:16",
          "content": "<p>nocall worked for us maybe because of class weights.  I updated the post to describe what we did.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1013021,
      "author_name": "kirderf",
      "author_url": "",
      "post_date": "09/16/2020 13:13:29",
      "content": "<p>I was never worried, you ended in the top as usual ☝️👍</p>",
      "votes": null,
      "replies": [
        {
          "id": 1013030,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "09/16/2020 13:19:39",
          "content": "<p>Indeed, I was way more worried than you were ;)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1013180,
      "author_name": "bhanvimenghani",
      "author_url": "",
      "post_date": "09/16/2020 14:51:50",
      "content": "<p>Well done</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1013409,
      "author_name": "watzisname",
      "author_url": "",
      "post_date": "09/16/2020 17:25:00",
      "content": "<p>Congrats <a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a> on a strong finish.<br>\nNice idea with regard to training on either first 5 or last 5 seconds part of the clip.</p>\n<p>Quick question on noise - What was the noise level you used ?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1013479,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "09/16/2020 17:57:16",
          "content": "<p>Thanks.  In the writeup:</p>\n<blockquote>\n  <p>We added noise extracted from the two test sequences made available, a bit like what Theo Viel did. But we used the meta data to extract sub sequences without bird call, then we merged these sequence with a smooth transition between them. We then added a random clip of the merges sequences to our training clips</p>\n</blockquote>\n<p>We multiplied the noise by a random number between 0.2 and 1.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1013906,
      "author_name": "feronial",
      "author_url": "",
      "post_date": "09/17/2020 03:54:45",
      "content": "<p>It's nice to see Efficient Models nowadays. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1020499,
      "author_name": "feronial",
      "author_url": "",
      "post_date": "09/21/2020 07:51:02",
      "content": "<p>Nice writeup. Impressive. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1013218,
      "author_name": "rahuljain63",
      "author_url": "",
      "post_date": "09/16/2020 15:10:58",
      "content": "<p>Great to see your effort..</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1019409,
      "author_name": "pinakimishrads",
      "author_url": "",
      "post_date": "09/20/2020 12:10:42",
      "content": "<p>Congrats .. Thanks for sharing </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1012236": "First of all I want to thank the host and Kaggle for this very challenging competition, with a truly hidden test set.  A bit too hidden maybe, but thanks to the community latecomers like us could get a submission template up and running in a couple of days.  I also want to thank my team mate Kazuki, without whom I would probably have given up after many failed attempts to beat the public notebook baseline...\n\n**Overview**\n\nOur best subs are single efficientnet models trained on log mel spectrograms.  For our baseline I started from scratch rather than reusing the excellent baselines that were available.  Reason is that I enter Kaggle competition to learn, and I learn more when I try from scratch than when I  modify someone else' code.  We then evolved that baseline as we could in the two weeks we had before the end of competition.\n\nGiven the overall approach is well known and probably used by most participants here I will only discuss the items that may be a bit different from what others did.\n\n**Training data clips**\n\nIt was clear from host that training on random 5 second clips had a drawback: some of the clips may not contain the target bird.  We then used a simple hypothesis: clips were stripped, i.e. periods without a song at the beginning or at the end were removed for the sake of reducing storage needs.  We therefore trained on first 5 seconds or last 5 seconds of clips, assuming these would contain the target bird.  We preprocessed all data to be sampled at 32 kHz.\n\n**Noise**\nWe added noise extracted from the two test sequences made available, a bit like what Theo Viel did.  But we used the meta data to extract sub sequences without bird call, then we merged these sequence with a smooth transition between them.  We then added a random clip of the merges sequences to our training clips\n\n**No Call**\nWe added the freefield1010 clips that were labelled as nocall to our training data.  We added a 265th class to represent the no call.  As a result our model could predict both one or more birds, and a nocall. Adding this data and the nocall class was probably the most important single improvement we saw in CvV and LB scores.  It is what led us to pass the public notebook baseline.\n\n**Multi Bird Clips**\nThe main documented difference between train and test data is that train is a multi class data while test is a multi label data.  Therefore we implemented a mixup variant were up to 3 clips could be merged.  This is not really mixup as the target for the merged clip is the maximum of the targets of each merged clip.\n\n**Secondary labels**\nPrimary labels were noisy, but secondary labels were even noisier.  As a result we masked the loss for secondary labels as we didn't want to force the model to learn a presence or an absence when we don't know.  We therefore defined a secondary mask that nullifies the BCE loss for secondary labels.  For instance, assuming only 3 ebird_code b0, b1, and b2, and a clip with primary label b0 and secondary label b1, then these two target values are possible:\n\n[1, 0, 0]\n\n[1, 1, 0]\n\nThe secondary mask is therefore:\n\n[1, 0, 1]\n\nFor merged clips, a target is masked if it it not one of the primary labels and if it is one of the secondary labels.\n\n**Loss Function**\n\nWe use binary cross entropy on one hot encoding of ebird codes.  Using bce rather than softmax makes sense as bce extends to multi label seamlessly.  We tried dice loss to directly optimize F! score, but for some reason this led to very strong overfitting.\n\n**Class Weights**\n\nThe number of record per species is not always 100.  In order to not penalize the less frequent one we use class weights inversely proportional to class frequencies.  And for the nocall class we set it to 1 even though it was way more frequent than each bird classes to make sure the model learns about nocall correctly.\n\n**Model**\n\nOur best model was efficientnet on log mel spectrograms.  We resized images to be twice the size of effnet images: 240x480 for effnet b1, 260x520 for effnet b2, and 300x600 for effnet b3.  We started from efficientnet_pytorch pretrained models.  We tried the attention head from PANNs models but it led to severe overfitting.  I am not sure why to be honest, maybe we did something wrong.\n\n**Training**\nNothings fancy, adam optimizer and cosine scheduler with 60 epochs.  In general last epoch was the one with best score and we used last epoch weights for scoring.\n\n\n**Log Mel Spectrogram**\n\nNothing fancy, except that we saw a lot of power in low frequencies of the first spectrograms we created. As a result we clipped frequency to be at least 300 Hz. We also clipped them to be below 16 kHz given test data was sampled at twice that frequency.  \n\n**Augmentations**\n\nTime and pitch variations were implements in a very simple way: modify the length of the clipped sequence, and modify the sampling rate, without modifying the data itself.  For instance, we would read 5.2 seconds of a clip instead of 5 seconds, and we could tell librosa that the sampling rate was 0.9*32 kHz.  We then compute the hop so that the number of stft is equal to the image width for the effnet model we are training.  We also computed the number of mel bins to be equal to the height of the image.  As a result we never had to resample data nor resize images, which speed up training and inferencing quite a bit.  There was an issue with that still: this led us to use a high nftt value of 2048 which lead to poor time resolution.  We ended up with nfft of 1048 and a resize of the images in the height dimension.\n\n**Cross Validation**\n\nWe started with cross validating on our training data (first or last 5 seconds of clips) but it was rapidly clear that it was not representative of the LB score.  And given we had very few submissions, we could not perform LB probing.  We therefore spent some time during last week to create a CV score that was correlated with the public LB.  Our CV score is computed using multi bird clips for 46% of the score, and nocall clips for 54% of the score.  We added warblrb and birdvox nocall clips to the freefield1010 for valuating no call performance.  We tuned the proportion of each possible number of clips in multi bird clips, and the amount of noise until we could find a relatively good CV LB relationship, see the picture below: x is cv, y is public lb.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F75976%2F39f916bec9be955cfa2a8fd5e8384b3b%2Fcv_lb.png?generation=1600261413279842&alt=media)\n\nThe correlation with private LB is also quite good:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F75976%2F02c89402f0dfd6e9069f36e67fc32e85%2Fcv_private.png?generation=1600261852596913&alt=media)\n\nWe then used it to guide our last few submissions training and thresholding.  Our CV said 0.7 was best, and late submissions proved it was right.  We ended up selecting our 2 best CV submissions, which were also the 2 best public LB submissions, and also the two best private Lb submissions.\n\n**Conclusion**\n\nThese were probably the most intensive two weeks I had on Kaggle since a long time.  My first two subs 14 days ago were a failure, then a LB score of 0. Kazuki had started a bit earlier, but not much. I am very happy about where we landed, and I am not sure we would have done much better with few more days.  Maybe using effnet b4 or b5 wold have moved us higher but I am not sure.  I am looking for gold medalist solutions to see what we missed.  I'm sure I'll learn quite a bit.",
    "1012259": "Congratulations @cpmpml garu",
    "1012438": "> These were probably the most intensive two weeks I had on Kaggle since a long time.\n\nI quite agree. Congratulations, and thanks for sharing your approach.",
    "1012440": "Congratulations @cpmpml and thanks for the lovely writeup  , time and again you always inspire me and one of the key things that I have learned from you is to write your own code and enter any competition for learning ,and one day medals will start to follow . Although writing my own code failed me during this competition but the learning was immense \nThanks",
    "1012462": "Masking secondary is really an interesting approach to reduce the label noise!\nCool stuff. I wonder why nocall as additional class didnt work for me",
    "1012723": "nocall worked for us maybe because of class weights.  I updated the post to describe what we did.",
    "1012725": "Thank you.  yes, progress is slower if you write your own code but you'll go farther.",
    "1012726": "Thanks.  Indeed, a number of late comers did well here, some in gold unless mistaken.  I didn't track precisely hence I prefer not to name people in case I forget some ;)  I hope you'll share your approach as well!  Congrats on the result even if you dropped a bit in private LB.  I hope you're not too disappointed.",
    "1012728": "Thank you.",
    "1013021": "I was never worried, you ended in the top as usual ☝️👍",
    "1013030": "Indeed, I was way more worried than you were ;)",
    "1013180": "Well done",
    "1013218": "Great to see your effort..",
    "1013349": "I don't think that I did as much as some people ended up in the gold range, so nothing really regret about, probably just a little bit.",
    "1013409": "Congrats @cpmpml on a strong finish.\nNice idea with regard to training on either first 5 or last 5 seconds part of the clip.\n\nQuick question on noise - What was the noise level you used ?",
    "1013479": "Thanks.  In the writeup:\n\n> We added noise extracted from the two test sequences made available, a bit like what Theo Viel did. But we used the meta data to extract sub sequences without bird call, then we merged these sequence with a smooth transition between them. We then added a random clip of the merges sequences to our training clips\n\nWe multiplied the noise by a random number between 0.2 and 1.",
    "1013906": "It's nice to see Efficient Models nowadays.",
    "1019409": "Congrats .. Thanks for sharing",
    "1020499": "Nice writeup. Impressive."
  },
  "source": "meta"
}