{
  "id": 412707,
  "title": "2nd place solution: SED + CNN with 7 models ensemble",
  "url": "/competitions/birdclef-2023/writeups/griffith-2nd-place-solution-sed-cnn-with-7-models-",
  "author_name": "",
  "post_date": "2024-04-04T10:27:54.573Z",
  "votes": 67,
  "comment_count": 37,
  "views": 0,
  "content": "<p>Congratulations to all the winners! Thanks to Kaggle and Cornell Lab of Ornithology for hosting this interesting competition.</p>\n<p>This is my first solo gold medal and I am glad to have this result.</p>\n<p>This competition shared a lot of similarity to the past BirdClef competitions(2020/2021/2022). Thus I spent a lot of time gathering the solution shared by the top teams in the past competitions. Special thanks to all of you for sharing such important information!</p>\n<p>Let me briefly introduce my solution. I will update the solution for more details in a couple of days.</p>\n<h1>Most important (7 models ensemble!)</h1>\n<p>Please see the notebook below.<br>\n<a href=\"https://www.kaggle.com/code/honglihang/openvino-is-all-you-need\" target=\"_blank\">openvino is all you need!!</a></p>\n<h1>Training data</h1>\n<p>Here is my training data.</p>\n<ul>\n<li>2023/2022/2021/2020 competition data</li>\n<li><a href=\"https://www.kaggle.com/competitions/birdclef-2023/discussion/398318\" target=\"_blank\">2020 additional competition data</a></li>\n<li>additional training data from xeno-canto, including 2023 comp species in both foreground and background(records with 2023 comp species only in background which is less than 60 seconds are included). </li>\n</ul>\n<p>I intended to collect more records from ebird site, but I realized that ebird data is not public and cannot be used. I <a href=\"https://www.kaggle.com/competitions/birdclef-2023/discussion/393023\" target=\"_blank\">asked the host</a> and confirmed that. Thanks <a href=\"https://www.kaggle.com/tomdenton\" target=\"_blank\">@tomdenton</a> answering my questions.</p>\n<p>Thus, my training pipeline does not contain records from ebird site.</p>\n<h1>Model Architecture</h1>\n<p>First, I used SED architecture. The same as yours.</p>\n<p>backbones are:</p>\n<ul>\n<li>tf_efficientnetv2_s_in21k</li>\n<li>seresnext26t_32x4d</li>\n<li>tf_efficientnet_b3_ns</li>\n</ul>\n<p>All of them are trained on 10sec clip.</p>\n<p>Second, I used CNN proposed by <a href=\"https://www.kaggle.com/competitions/birdclef-2021/discussion/243463\" target=\"_blank\">2nd place of 2021 competition</a></p>\n<p>backbones are:</p>\n<ul>\n<li>tf_efficientnetv2_s_in21k</li>\n<li>resnet34d</li>\n<li>tf_efficientnet_b3_ns</li>\n<li>tf_efficientnet_b0_ns</li>\n</ul>\n<p>All except b0 are trained on 15sec clip. b0 is trained on 20sec clip.</p>\n<h1>Pseudo Labeling and Hand Labeling</h1>\n<p>I have used SED model to generate pseudo label and extracted the potential nocall using quantile threshold. Then, I hand labeled the potential nocall by hearing the record. I hand labeled about 1800 records but did not see improvement. Maybe the pseudo label contains more FP rather than FN. I did not have time to further investigate the prediction.</p>\n<h1>Model Training</h1>\n<p>augmentations:</p>\n<ul>\n<li>GaussianNoise</li>\n<li>PinkNoise</li>\n<li>Gain</li>\n<li>NoiseInjection</li>\n<li>Background Noise(nocall in 2020, 2021 comp + rainforest + environment sound + nocall in freefield1010, warblrb, birdvox)</li>\n<li>PitchShift</li>\n<li>TimeShift</li>\n<li>FrequencyMasking</li>\n<li>TimeMasking</li>\n<li>OR Mixup on waveforms</li>\n<li>Mixup on spectrograms.</li>\n<li><a href=\"https://www.kaggle.com/competitions/birdsong-recognition/discussion/183269\" target=\"_blank\">With a probability of 0.5 lowered the upper frequencies</a></li>\n<li>self mixup for records with 2023 species only in background.(60sec waveform -&gt; split to 6 * 10sec -&gt; np.sum(audios,axis=0) to get a 10sec clip)</li>\n</ul>\n<p>I have used weights (computed by primary_label and secondary_labels) for Dataloader in order to cope with unbalanced dataset.</p>\n<h1>Training stages</h1>\n<p>For training I have used 2 stage training:</p>\n<ol>\n<li>Pretrain on all data(834 species).</li>\n<li>Finetune on 2023 species(264 species).</li>\n</ol>\n<p>In both stages, I first train model with CrossEntropyLoss, and then train on with BCEWithLogitsLoss(reduction='sum'). Model converges faster with CrossEntropyLoss than BCEWithLogitsLoss, but BCEWithLogitsLoss gives better score.</p>\n<p>To give more diversity, models are trained on different windows and different mixup rate, and some of them only trained on CrossEntropyLoss. And also 3 of the models are fintuned on 30s clip.</p>\n<h1>CV strategy</h1>\n<ul>\n<li>For each validation sample - slice the first 60 seconds to pieces -&gt; predict each piece -&gt; max(sample_predictions, dim=pieces).</li>\n</ul>\n<p>CV does not show correlation with LB, but it seems that the right ways to improve the LB are those which do not significantly decrease CV. So I monitored the CV when tuning the pipeline.</p>\n<h1>Inference</h1>\n<p>For SED model, feed model 10 sec chunk BUT apply head only on centered 5 sec reduced CNN image and use max(framewise, dim=time).</p>\n<p>Also, tta(2s) is used for SED model.</p>\n<p>Important: convert pytorch model to openvino model significantly reduce inference time(about 40%). (eca_nfnet_l0 backbone ONXX cannot be converted to openvino because the stdconv layer in timm use train mode of F.batch_norm in forward method). That is the magic of ensembling 7 models.</p>\n<h1>Ensemble</h1>\n<p>I spent quite a lot of time understanding the metrics. The ensemble is as followed:</p>\n<ol>\n<li>(weighted average, 0.84 on LB, 0.76 on private) Apply weighted average on raw logit. This ensemble does not make sense for me because the output logit of models differ and should not be simply added, otherwise the result is biased. But considering that the absolute value of logit may also contribute to the score and it does give the best LB, so I choose it for final submission to have a gamble.</li>\n<li>(rank average, 0.83 on LB, 0.75 on private) Convert the logit to rank and apply weighted average on rank. I think this is the reasonable way to ensemble, considering that I don't have a reliable CV. <strong>Be careful that in this comp submission is padded with 5 rows of 1, thus the ranking should start from 0 to prevent the largest ranking to be 1 after convert rankings to percentile form.(For example, 0, 0.333, 0.666 is good while 0.333, 0.666, 1.0 is bad. otherwise the score will decrease due to the potential false positive of those ranked 1.0) </strong></li>\n</ol>\n<p>Basically all of my single model's performance reaches about 0.81-0.82 on LB, so the weight of each model is similar to give diversity although higher public lb requires large weight of sed v2s and sed b3_ns.</p>\n<p>Looking at the private LB, weighted average seems to be better. emmm……Why?</p>\n<h1>What did not work</h1>\n<ul>\n<li><a href=\"https://www.kaggle.com/competitions/birdsong-recognition/discussion/183269\" target=\"_blank\">contrast</a></li>\n<li>EMA</li>\n<li>BirdNet</li>\n<li>adjusting logit according to previous and next 5s clip</li>\n<li>q transform</li>\n</ul>\n<p>Inference Notebook: <a href=\"https://www.kaggle.com/code/honglihang/2nd-place-solution-inference-kernel\" target=\"_blank\">https://www.kaggle.com/code/honglihang/2nd-place-solution-inference-kernel</a><br>\ngithub: <a href=\"https://github.com/LIHANG-HONG/birdclef2023-2nd-place-solution\" target=\"_blank\">https://github.com/LIHANG-HONG/birdclef2023-2nd-place-solution</a></p>",
  "messages": [
    {
      "id": "2272995",
      "postDate": "05/25/2023 00:05:40",
      "content": "<p>Congratulations to all the winners! Thanks to Kaggle and Cornell Lab of Ornithology for hosting this interesting competition.</p>\n<p>This is my first solo gold medal and I am glad to have this result.</p>\n<p>This competition shared a lot of similarity to the past BirdClef competitions(2020/2021/2022). Thus I spent a lot of time gathering the solution shared by the top teams in the past competitions. Special thanks to all of you for sharing such important information!</p>\n<p>Let me briefly introduce my solution. I will update the solution for more details in a couple of days.</p>\n<h1>Most important (7 models ensemble!)</h1>\n<p>Please see the notebook below.<br>\n<a href=\"https://www.kaggle.com/code/honglihang/openvino-is-all-you-need\" target=\"_blank\">openvino is all you need!!</a></p>\n<h1>Training data</h1>\n<p>Here is my training data.</p>\n<ul>\n<li>2023/2022/2021/2020 competition data</li>\n<li><a href=\"https://www.kaggle.com/competitions/birdclef-2023/discussion/398318\" target=\"_blank\">2020 additional competition data</a></li>\n<li>additional training data from xeno-canto, including 2023 comp species in both foreground and background(records with 2023 comp species only in background which is less than 60 seconds are included). </li>\n</ul>\n<p>I intended to collect more records from ebird site, but I realized that ebird data is not public and cannot be used. I <a href=\"https://www.kaggle.com/competitions/birdclef-2023/discussion/393023\" target=\"_blank\">asked the host</a> and confirmed that. Thanks <a href=\"https://www.kaggle.com/tomdenton\" target=\"_blank\">@tomdenton</a> answering my questions.</p>\n<p>Thus, my training pipeline does not contain records from ebird site.</p>\n<h1>Model Architecture</h1>\n<p>First, I used SED architecture. The same as yours.</p>\n<p>backbones are:</p>\n<ul>\n<li>tf_efficientnetv2_s_in21k</li>\n<li>seresnext26t_32x4d</li>\n<li>tf_efficientnet_b3_ns</li>\n</ul>\n<p>All of them are trained on 10sec clip.</p>\n<p>Second, I used CNN proposed by <a href=\"https://www.kaggle.com/competitions/birdclef-2021/discussion/243463\" target=\"_blank\">2nd place of 2021 competition</a></p>\n<p>backbones are:</p>\n<ul>\n<li>tf_efficientnetv2_s_in21k</li>\n<li>resnet34d</li>\n<li>tf_efficientnet_b3_ns</li>\n<li>tf_efficientnet_b0_ns</li>\n</ul>\n<p>All except b0 are trained on 15sec clip. b0 is trained on 20sec clip.</p>\n<h1>Pseudo Labeling and Hand Labeling</h1>\n<p>I have used SED model to generate pseudo label and extracted the potential nocall using quantile threshold. Then, I hand labeled the potential nocall by hearing the record. I hand labeled about 1800 records but did not see improvement. Maybe the pseudo label contains more FP rather than FN. I did not have time to further investigate the prediction.</p>\n<h1>Model Training</h1>\n<p>augmentations:</p>\n<ul>\n<li>GaussianNoise</li>\n<li>PinkNoise</li>\n<li>Gain</li>\n<li>NoiseInjection</li>\n<li>Background Noise(nocall in 2020, 2021 comp + rainforest + environment sound + nocall in freefield1010, warblrb, birdvox)</li>\n<li>PitchShift</li>\n<li>TimeShift</li>\n<li>FrequencyMasking</li>\n<li>TimeMasking</li>\n<li>OR Mixup on waveforms</li>\n<li>Mixup on spectrograms.</li>\n<li><a href=\"https://www.kaggle.com/competitions/birdsong-recognition/discussion/183269\" target=\"_blank\">With a probability of 0.5 lowered the upper frequencies</a></li>\n<li>self mixup for records with 2023 species only in background.(60sec waveform -&gt; split to 6 * 10sec -&gt; np.sum(audios,axis=0) to get a 10sec clip)</li>\n</ul>\n<p>I have used weights (computed by primary_label and secondary_labels) for Dataloader in order to cope with unbalanced dataset.</p>\n<h1>Training stages</h1>\n<p>For training I have used 2 stage training:</p>\n<ol>\n<li>Pretrain on all data(834 species).</li>\n<li>Finetune on 2023 species(264 species).</li>\n</ol>\n<p>In both stages, I first train model with CrossEntropyLoss, and then train on with BCEWithLogitsLoss(reduction='sum'). Model converges faster with CrossEntropyLoss than BCEWithLogitsLoss, but BCEWithLogitsLoss gives better score.</p>\n<p>To give more diversity, models are trained on different windows and different mixup rate, and some of them only trained on CrossEntropyLoss. And also 3 of the models are fintuned on 30s clip.</p>\n<h1>CV strategy</h1>\n<ul>\n<li>For each validation sample - slice the first 60 seconds to pieces -&gt; predict each piece -&gt; max(sample_predictions, dim=pieces).</li>\n</ul>\n<p>CV does not show correlation with LB, but it seems that the right ways to improve the LB are those which do not significantly decrease CV. So I monitored the CV when tuning the pipeline.</p>\n<h1>Inference</h1>\n<p>For SED model, feed model 10 sec chunk BUT apply head only on centered 5 sec reduced CNN image and use max(framewise, dim=time).</p>\n<p>Also, tta(2s) is used for SED model.</p>\n<p>Important: convert pytorch model to openvino model significantly reduce inference time(about 40%). (eca_nfnet_l0 backbone ONXX cannot be converted to openvino because the stdconv layer in timm use train mode of F.batch_norm in forward method). That is the magic of ensembling 7 models.</p>\n<h1>Ensemble</h1>\n<p>I spent quite a lot of time understanding the metrics. The ensemble is as followed:</p>\n<ol>\n<li>(weighted average, 0.84 on LB, 0.76 on private) Apply weighted average on raw logit. This ensemble does not make sense for me because the output logit of models differ and should not be simply added, otherwise the result is biased. But considering that the absolute value of logit may also contribute to the score and it does give the best LB, so I choose it for final submission to have a gamble.</li>\n<li>(rank average, 0.83 on LB, 0.75 on private) Convert the logit to rank and apply weighted average on rank. I think this is the reasonable way to ensemble, considering that I don't have a reliable CV. <strong>Be careful that in this comp submission is padded with 5 rows of 1, thus the ranking should start from 0 to prevent the largest ranking to be 1 after convert rankings to percentile form.(For example, 0, 0.333, 0.666 is good while 0.333, 0.666, 1.0 is bad. otherwise the score will decrease due to the potential false positive of those ranked 1.0) </strong></li>\n</ol>\n<p>Basically all of my single model's performance reaches about 0.81-0.82 on LB, so the weight of each model is similar to give diversity although higher public lb requires large weight of sed v2s and sed b3_ns.</p>\n<p>Looking at the private LB, weighted average seems to be better. emmm……Why?</p>\n<h1>What did not work</h1>\n<ul>\n<li><a href=\"https://www.kaggle.com/competitions/birdsong-recognition/discussion/183269\" target=\"_blank\">contrast</a></li>\n<li>EMA</li>\n<li>BirdNet</li>\n<li>adjusting logit according to previous and next 5s clip</li>\n<li>q transform</li>\n</ul>\n<p>Inference Notebook: <a href=\"https://www.kaggle.com/code/honglihang/2nd-place-solution-inference-kernel\" target=\"_blank\">https://www.kaggle.com/code/honglihang/2nd-place-solution-inference-kernel</a><br>\ngithub: <a href=\"https://github.com/LIHANG-HONG/birdclef2023-2nd-place-solution\" target=\"_blank\">https://github.com/LIHANG-HONG/birdclef2023-2nd-place-solution</a></p>",
      "rawMarkdown": "Congratulations to all the winners! Thanks to Kaggle and Cornell Lab of Ornithology for hosting this interesting competition.\n\nThis is my first solo gold medal and I am glad to have this result.\n\nThis competition shared a lot of similarity to the past BirdClef competitions(2020/2021/2022). Thus I spent a lot of time gathering the solution shared by the top teams in the past competitions. Special thanks to all of you for sharing such important information!\n\nLet me briefly introduce my solution. I will update the solution for more details in a couple of days.\n\n# Most important (7 models ensemble!)\nPlease see the notebook below.\n[openvino is all you need!!](https://www.kaggle.com/code/honglihang/openvino-is-all-you-need)\n\n# Training data\n\nHere is my training data.\n\n- 2023/2022/2021/2020 competition data\n- [2020 additional competition data](https://www.kaggle.com/competitions/birdclef-2023/discussion/398318)\n- additional training data from xeno-canto, including 2023 comp species in both foreground and background(records with 2023 comp species only in background which is less than 60 seconds are included). \n\nI intended to collect more records from ebird site, but I realized that ebird data is not public and cannot be used. I [asked the host](https://www.kaggle.com/competitions/birdclef-2023/discussion/393023) and confirmed that. Thanks [@tomdenton](https://www.kaggle.com/tomdenton) answering my questions.\n\nThus, my training pipeline does not contain records from ebird site.\n\n# Model Architecture\n\nFirst, I used SED architecture. The same as yours.\n\nbackbones are:\n\n- tf_efficientnetv2_s_in21k\n- seresnext26t_32x4d\n- tf_efficientnet_b3_ns\n\nAll of them are trained on 10sec clip.\n\nSecond, I used CNN proposed by [2nd place of 2021 competition](https://www.kaggle.com/competitions/birdclef-2021/discussion/243463)\n\nbackbones are:\n\n- tf_efficientnetv2_s_in21k\n- resnet34d\n- tf_efficientnet_b3_ns\n- tf_efficientnet_b0_ns\n\nAll except b0 are trained on 15sec clip. b0 is trained on 20sec clip.\n\n# Pseudo Labeling and Hand Labeling\n\nI have used SED model to generate pseudo label and extracted the potential nocall using quantile threshold. Then, I hand labeled the potential nocall by hearing the record. I hand labeled about 1800 records but did not see improvement. Maybe the pseudo label contains more FP rather than FN. I did not have time to further investigate the prediction.\n\n# Model Training\n\naugmentations:\n\n- GaussianNoise\n- PinkNoise\n- Gain\n- NoiseInjection\n- Background Noise(nocall in 2020, 2021 comp + rainforest + environment sound + nocall in freefield1010, warblrb, birdvox)\n- PitchShift\n- TimeShift\n- FrequencyMasking\n- TimeMasking\n- OR Mixup on waveforms\n- Mixup on spectrograms.\n- [With a probability of 0.5 lowered the upper frequencies](https://www.kaggle.com/competitions/birdsong-recognition/discussion/183269)\n- self mixup for records with 2023 species only in background.(60sec waveform -> split to 6 \\* 10sec -> np.sum(audios,axis=0) to get a 10sec clip)\n\nI have used weights (computed by primary_label and secondary_labels) for Dataloader in order to cope with unbalanced dataset.\n\n# Training stages\n\nFor training I have used 2 stage training:\n\n1. Pretrain on all data(834 species).\n2. Finetune on 2023 species(264 species).\n\nIn both stages, I first train model with CrossEntropyLoss, and then train on with BCEWithLogitsLoss(reduction='sum'). Model converges faster with CrossEntropyLoss than BCEWithLogitsLoss, but BCEWithLogitsLoss gives better score.\n\nTo give more diversity, models are trained on different windows and different mixup rate, and some of them only trained on CrossEntropyLoss. And also 3 of the models are fintuned on 30s clip.\n\n# CV strategy\n\n- For each validation sample - slice the first 60 seconds to pieces -> predict each piece -> max(sample_predictions, dim=pieces).\n\nCV does not show correlation with LB, but it seems that the right ways to improve the LB are those which do not significantly decrease CV. So I monitored the CV when tuning the pipeline.\n\n# Inference\n\nFor SED model, feed model 10 sec chunk BUT apply head only on centered 5 sec reduced CNN image and use max(framewise, dim=time).\n\nAlso, tta(2s) is used for SED model.\n\nImportant: convert pytorch model to openvino model significantly reduce inference time(about 40%). (eca_nfnet_l0 backbone ONXX cannot be converted to openvino because the stdconv layer in timm use train mode of F.batch_norm in forward method). That is the magic of ensembling 7 models.\n\n# Ensemble\n\nI spent quite a lot of time understanding the metrics. The ensemble is as followed:\n\n1. (weighted average, 0.84 on LB, 0.76 on private) Apply weighted average on raw logit. This ensemble does not make sense for me because the output logit of models differ and should not be simply added, otherwise the result is biased. But considering that the absolute value of logit may also contribute to the score and it does give the best LB, so I choose it for final submission to have a gamble.\n2. (rank average, 0.83 on LB, 0.75 on private) Convert the logit to rank and apply weighted average on rank. I think this is the reasonable way to ensemble, considering that I don't have a reliable CV. <strong>Be careful that in this comp submission is padded with 5 rows of 1, thus the ranking should start from 0 to prevent the largest ranking to be 1 after convert rankings to percentile form.(For example, 0, 0.333, 0.666 is good while 0.333, 0.666, 1.0 is bad. otherwise the score will decrease due to the potential false positive of those ranked 1.0) </strong>\n\nBasically all of my single model's performance reaches about 0.81-0.82 on LB, so the weight of each model is similar to give diversity although higher public lb requires large weight of sed v2s and sed b3_ns.\n\nLooking at the private LB, weighted average seems to be better. emmm......Why?\n\n# What did not work\n\n- [contrast](https://www.kaggle.com/competitions/birdsong-recognition/discussion/183269)\n- EMA\n- BirdNet\n- adjusting logit according to previous and next 5s clip\n- q transform\n\nInference Notebook: https://www.kaggle.com/code/honglihang/2nd-place-solution-inference-kernel\ngithub: https://github.com/LIHANG-HONG/birdclef2023-2nd-place-solution",
      "votes": null
    },
    {
      "id": "2273003",
      "postDate": "05/25/2023 00:23:15",
      "content": "<p>Huge congrats!!! 🎉🎉 Thanks for sharing, I was really curious about the top solutions. I'll read this in detail after having my (missing) night sleep.</p>",
      "rawMarkdown": "Huge congrats!!! 🎉🎉 Thanks for sharing, I was really curious about the top solutions. I'll read this in detail after having my (missing) night sleep.",
      "votes": null
    },
    {
      "id": "2273006",
      "postDate": "05/25/2023 00:27:20",
      "content": "<p>Congratulations! It must have been a fulfilling journey :)</p>",
      "rawMarkdown": "Congratulations! It must have been a fulfilling journey :)",
      "votes": null
    },
    {
      "id": "2273013",
      "postDate": "05/25/2023 00:36:22",
      "content": "<p>Congratulations, thanks for sharing</p>",
      "rawMarkdown": "Congratulations, thanks for sharing",
      "votes": null
    },
    {
      "id": "2273018",
      "postDate": "05/25/2023 00:51:53",
      "content": "<p>Congratulations, and thanks for sharing! We tried onnx but not openvino. Onnx worked very well individually, but when we tried to ensemble the models, the results were very bad, almost all the time we ended up with a timeout or some kind of error, probably because we made some mistake during the ensemble. For that reason, we skip onnx, and our ensemble was with only 4 Pytorch models.</p>",
      "rawMarkdown": "Congratulations, and thanks for sharing! We tried onnx but not openvino. Onnx worked very well individually, but when we tried to ensemble the models, the results were very bad, almost all the time we ended up with a timeout or some kind of error, probably because we made some mistake during the ensemble. For that reason, we skip onnx, and our ensemble was with only 4 Pytorch models.",
      "votes": null
    },
    {
      "id": "2273051",
      "postDate": "05/25/2023 01:34:15",
      "content": "<p>Congratulations on achieving second place as a solo participant! Incorporating \"heavy\" augmentations like pitch shifting seems like it would significantly extend the training time. Have you taken any measures to address this?</p>",
      "rawMarkdown": "Congratulations on achieving second place as a solo participant! Incorporating \"heavy\" augmentations like pitch shifting seems like it would significantly extend the training time. Have you taken any measures to address this?",
      "votes": null
    },
    {
      "id": "2273056",
      "postDate": "05/25/2023 01:43:23",
      "content": "<p>Congratulations and thanks for sharing</p>",
      "rawMarkdown": "Congratulations and thanks for sharing",
      "votes": null
    },
    {
      "id": "2273190",
      "postDate": "05/25/2023 04:17:10",
      "content": "<p>Congratulations ✨ and thanks for sharing, I am just a beginner getting started with Deep Learning this year only, for me it is both fascinating and intimidating how you can figure out so many things while approaching a problem. Do you have any tips about how I can become more efficient in problem solving like this. Thank you for your time.</p>",
      "rawMarkdown": "Congratulations ✨ and thanks for sharing, I am just a beginner getting started with Deep Learning this year only, for me it is both fascinating and intimidating how you can figure out so many things while approaching a problem. Do you have any tips about how I can become more efficient in problem solving like this. Thank you for your time.",
      "votes": null
    },
    {
      "id": "2273253",
      "postDate": "05/25/2023 05:01:57",
      "content": "<p>Amazing 💯 Congratulations and thanks for sharing 😊</p>",
      "rawMarkdown": "Amazing 💯 Congratulations and thanks for sharing 😊",
      "votes": null
    },
    {
      "id": "2273274",
      "postDate": "05/25/2023 05:13:01",
      "content": "<p>Congratulations and thanks for sharing! I would like to learn Openvino from your notebook.</p>",
      "rawMarkdown": "Congratulations and thanks for sharing! I would like to learn Openvino from your notebook.",
      "votes": null
    },
    {
      "id": "2273425",
      "postDate": "05/25/2023 06:42:18",
      "content": "<p>Incredible !!! Congratulations 🎉. </p>",
      "rawMarkdown": "Incredible !!! Congratulations 🎉.",
      "votes": null
    },
    {
      "id": "2273477",
      "postDate": "05/25/2023 07:20:54",
      "content": "<p>Congratulations…..</p>\n<p>Ensembling trick of openvino is indeed something to add to my knowledge arsenal. Great work!!!</p>",
      "rawMarkdown": "Congratulations.....\n\nEnsembling trick of openvino is indeed something to add to my knowledge arsenal. Great work!!!",
      "votes": null
    },
    {
      "id": "2273496",
      "postDate": "05/25/2023 07:33:56",
      "content": "<p>Congratulations!</p>\n<p>Amazing stuff!</p>",
      "rawMarkdown": "Congratulations!\n\nAmazing stuff!",
      "votes": null
    },
    {
      "id": "2273602",
      "postDate": "05/25/2023 09:16:02",
      "content": "<p>Very nice, congratulations and thanks for the quick writeup! Will look into openvino as well, nice find.</p>\n<blockquote>\n  <p>I hand labeled about 1800 records but did not see improvement.</p>\n</blockquote>\n<p>So in the end, you neither used manual nor pseudo-labels? I experimented with pseudo-labels on the test set (with much weaker models than the top solutions), but hardly got an improvement either, hence I'm curious.</p>\n<blockquote>\n  <p>self mixup for records with 2023 species only in background</p>\n</blockquote>\n<p>Interesting idea! What was the intuition behind that? Make it more likely that the bird in question is actually part of the 10s clip? How much did it help, compared to other tricks?</p>\n<blockquote>\n  <p>In both stages, I first train model with CrossEntropyLoss, and then train on with BCEWithLogitsLoss(reduction='sum').</p>\n</blockquote>\n<p>Again, interesting idea, will keep this in my head!</p>\n<blockquote>\n  <p>For SED model, feed model 10 sec chunk BUT apply head only on centered 5 sec reduced CNN image and use max(framewise, dim=time).</p>\n</blockquote>\n<p>This sounds expensive. Does this double the inference time or did you feed longer chunks than 10s and then extracted multiple 5sec chunks from that?</p>\n<blockquote>\n  <p>Also, tta(2s) is used for SED model.</p>\n</blockquote>\n<p>tta = test-time augmentation? Was there even time for that? :-O</p>\n<p>Again, congratulations for your work!</p>",
      "rawMarkdown": "Very nice, congratulations and thanks for the quick writeup! Will look into openvino as well, nice find.\n\n> I hand labeled about 1800 records but did not see improvement.\n\nSo in the end, you neither used manual nor pseudo-labels? I experimented with pseudo-labels on the test set (with much weaker models than the top solutions), but hardly got an improvement either, hence I'm curious.\n\n> self mixup for records with 2023 species only in background\n\nInteresting idea! What was the intuition behind that? Make it more likely that the bird in question is actually part of the 10s clip? How much did it help, compared to other tricks?\n\n> In both stages, I first train model with CrossEntropyLoss, and then train on with BCEWithLogitsLoss(reduction='sum').\n\nAgain, interesting idea, will keep this in my head!\n\n> For SED model, feed model 10 sec chunk BUT apply head only on centered 5 sec reduced CNN image and use max(framewise, dim=time).\n\nThis sounds expensive. Does this double the inference time or did you feed longer chunks than 10s and then extracted multiple 5sec chunks from that?\n\n> Also, tta(2s) is used for SED model.\n\ntta = test-time augmentation? Was there even time for that? :-O\n\nAgain, congratulations for your work!",
      "votes": null
    },
    {
      "id": "2274088",
      "postDate": "05/25/2023 15:57:42",
      "content": "<p><a href=\"https://www.kaggle.com/honglihang\" target=\"_blank\">@honglihang</a> I wanted to reach out and express my heartfelt gratitude for sharing your incredible 2nd place solution in the recent competition. Your approach, combining Sound Event Detection (SED) with Convolutional Neural Networks (CNN) and utilizing an ensemble of seven models, is truly remarkable.</p>",
      "rawMarkdown": "honglihang I wanted to reach out and express my heartfelt gratitude for sharing your incredible 2nd place solution in the recent competition. Your approach, combining Sound Event Detection (SED) with Convolutional Neural Networks (CNN) and utilizing an ensemble of seven models, is truly remarkable.",
      "votes": null
    },
    {
      "id": "2276700",
      "postDate": "05/27/2023 06:23:07",
      "content": "<p>Thanks! It is really a long journey for me and I slept for two days to recover from lack of sleep </p>",
      "rawMarkdown": "Thanks! It is really a long journey for me and I slept for two days to recover from lack of sleep",
      "votes": null
    },
    {
      "id": "2276704",
      "postDate": "05/27/2023 06:24:25",
      "content": "<p>Thanks you! We should all thanks to those who shared these amazing stuffs in the past few years.</p>",
      "rawMarkdown": "Thanks you! We should all thanks to those who shared these amazing stuffs in the past few years.",
      "votes": null
    },
    {
      "id": "2276706",
      "postDate": "05/27/2023 06:24:37",
      "content": "<p>Thank you!</p>",
      "rawMarkdown": "Thank you!",
      "votes": null
    },
    {
      "id": "2276721",
      "postDate": "05/27/2023 06:37:25",
      "content": "<p>Hi, thanks for your question.</p>\n<blockquote>\n  <p>I hand labeled about 1800 records but did not see improvement.</p>\n</blockquote>\n<p>I used both pseudo label and hand label, although there is no improvement seen in public LB, I have not confirm whether it has improvement in private LB.</p>\n<blockquote>\n  <p>self mixup for records with 2023 species only in background.</p>\n</blockquote>\n<p>The idea comes from that I found mixup in waveform is effective. So I think I can take use of the records with 2023 species only in background by applying the same idea because I no longer need to care about where the bird call appears in background by doing this mixup to fully use the record. I have not confirmed whether it is useful, but there was a marginal improvement in the score.</p>\n<blockquote>\n  <p>For SED model, feed model 10 sec chunk BUT apply head only on centered 5 sec reduced CNN image and use max(framewise, dim=time).</p>\n</blockquote>\n<p>Yes, it is expensive which double the inference time. So I only use 10s, not any longer…</p>\n<blockquote>\n  <p>Also, tta(2s) is used for SED model.</p>\n</blockquote>\n<p>TTA is only applied on attention head so it is super fast. The encoder extract the 10s feature and the target 5s is in the center of the matrix (2.5s to 7.5s). For tta, I take the feature from 1.5s to 6.5s and 3.5s to 8.5s.</p>",
      "rawMarkdown": "Hi, thanks for your question.\n>I hand labeled about 1800 records but did not see improvement.\n\nI used both pseudo label and hand label, although there is no improvement seen in public LB, I have not confirm whether it has improvement in private LB.\n>self mixup for records with 2023 species only in background.\n\nThe idea comes from that I found mixup in waveform is effective. So I think I can take use of the records with 2023 species only in background by applying the same idea because I no longer need to care about where the bird call appears in background by doing this mixup to fully use the record. I have not confirmed whether it is useful, but there was a marginal improvement in the score.\n\n>For SED model, feed model 10 sec chunk BUT apply head only on centered 5 sec reduced CNN image and use max(framewise, dim=time).\n\nYes, it is expensive which double the inference time. So I only use 10s, not any longer...\n\n>Also, tta(2s) is used for SED model.\n\nTTA is only applied on attention head so it is super fast. The encoder extract the 10s feature and the target 5s is in the center of the matrix (2.5s to 7.5s). For tta, I take the feature from 1.5s to 6.5s and 3.5s to 8.5s.",
      "votes": null
    },
    {
      "id": "2276724",
      "postDate": "05/27/2023 06:38:39",
      "content": "<p>Thank you! It is also my first time to try openvino.</p>",
      "rawMarkdown": "Thank you! It is also my first time to try openvino.",
      "votes": null
    },
    {
      "id": "2276725",
      "postDate": "05/27/2023 06:39:00",
      "content": "<p>Thanks! Happy kaggling!</p>",
      "rawMarkdown": "Thanks! Happy kaggling!",
      "votes": null
    },
    {
      "id": "2276728",
      "postDate": "05/27/2023 06:42:39",
      "content": "<p>Thank you! The same to you. Your solution is super incredible and intersting. I will take a deep look into your solution after I get enougth sleep~</p>\n<p>By the way, have you already got the winner email from the kaggle team? I haven't…</p>",
      "rawMarkdown": "Thank you! The same to you. Your solution is super incredible and intersting. I will take a deep look into your solution after I get enougth sleep~\n\nBy the way, have you already got the winner email from the kaggle team? I haven't...",
      "votes": null
    },
    {
      "id": "2276730",
      "postDate": "05/27/2023 06:44:36",
      "content": "<p>Thank you. I appreciate you sharing the data in 2020 comp and I took full use of it. I used <a href=\"https://github.com/asteroid-team/torch-audiomentations\" target=\"_blank\">torch-audiomentations</a> to perform pitch shift and time shift on GPU</p>",
      "rawMarkdown": "Thank you. I appreciate you sharing the data in 2020 comp and I took full use of it. I used [torch-audiomentations](https://github.com/asteroid-team/torch-audiomentations) to perform pitch shift and time shift on GPU",
      "votes": null
    },
    {
      "id": "2276732",
      "postDate": "05/27/2023 06:45:21",
      "content": "<p>Thank you! happy kaggle</p>",
      "rawMarkdown": "Thank you! happy kaggle",
      "votes": null
    },
    {
      "id": "2276733",
      "postDate": "05/27/2023 06:45:51",
      "content": "<p>Thank you and congratulations to you too!</p>",
      "rawMarkdown": "Thank you and congratulations to you too!",
      "votes": null
    },
    {
      "id": "2276736",
      "postDate": "05/27/2023 06:47:57",
      "content": "<p>Thank you! Also congrats to you!<br>\nThe other top solutions are very incredible to me. I will also read in detial after getting back to sleep..</p>",
      "rawMarkdown": "Thank you! Also congrats to you!\nThe other top solutions are very incredible to me. I will also read in detial after getting back to sleep..",
      "votes": null
    },
    {
      "id": "2276738",
      "postDate": "05/27/2023 06:51:45",
      "content": "<p>The same to your team! Congratulations! It has been a really a fulfilling(and short sleep) journey for me since the RSNA.<br>\nLooks that you have made birdnet work, I will read your solution in detail after getting enough sleep.</p>",
      "rawMarkdown": "The same to your team! Congratulations! It has been a really a fulfilling(and short sleep) journey for me since the RSNA.\nLooks that you have made birdnet work, I will read your solution in detail after getting enough sleep.",
      "votes": null
    },
    {
      "id": "2276739",
      "postDate": "05/27/2023 06:55:23",
      "content": "<p>Also congratulations to you! Thanks for sharing the information about onnx. I am also new to onnx and openvino and maybe it is lucky for me to try openvino at first because I learnt about openvino in a project at work</p>",
      "rawMarkdown": "Also congratulations to you! Thanks for sharing the information about onnx. I am also new to onnx and openvino and maybe it is lucky for me to try openvino at first because I learnt about openvino in a project at work",
      "votes": null
    },
    {
      "id": "2276743",
      "postDate": "05/27/2023 07:00:31",
      "content": "<p>Thank you! I started kaggle about 2 years ago and also had a hard time to understand the solutions. I think it is a process of try and error based on the knowledge shared by others. It is better to get new ideas through reading about the papers and the past top solutions.</p>",
      "rawMarkdown": "Thank you! I started kaggle about 2 years ago and also had a hard time to understand the solutions. I think it is a process of try and error based on the knowledge shared by others. It is better to get new ideas through reading about the papers and the past top solutions.",
      "votes": null
    },
    {
      "id": "2276834",
      "postDate": "05/27/2023 08:35:37",
      "content": "<p>Thanks!<br>\nI also haven't got it yet.</p>",
      "rawMarkdown": "Thanks!\nI also haven't got it yet.",
      "votes": null
    },
    {
      "id": "2276977",
      "postDate": "05/27/2023 10:58:58",
      "content": "<p>Awesome approach to use openvino for boosting inference times over ensembling <a href=\"https://www.kaggle.com/honglihang\" target=\"_blank\">@honglihang</a> ! Congrats !</p>",
      "rawMarkdown": "Awesome approach to use openvino for boosting inference times over ensembling @honglihang ! Congrats !",
      "votes": null
    },
    {
      "id": "2277433",
      "postDate": "05/27/2023 18:12:15",
      "content": "<p>That is one amazing solution to the competition!</p>",
      "rawMarkdown": "That is one amazing solution to the competition!",
      "votes": null
    },
    {
      "id": "2279058",
      "postDate": "05/29/2023 06:30:52",
      "content": "<p>Thanks for your reply.<br>\nI still haven't received yet, maybe it cost sometime….</p>",
      "rawMarkdown": "Thanks for your reply.\nI still haven't received yet, maybe it cost sometime....",
      "votes": null
    },
    {
      "id": "2280136",
      "postDate": "05/29/2023 22:08:59",
      "content": "<p>Awesome job!!</p>",
      "rawMarkdown": "Awesome job!!",
      "votes": null
    },
    {
      "id": "2285781",
      "postDate": "06/03/2023 02:04:06",
      "content": "<p>Great stuff <a href=\"https://www.kaggle.com/honglihang\" target=\"_blank\">@honglihang</a> congrats on your first solo gold and thanks for sharing this writeup + code 🎉 🔥</p>",
      "rawMarkdown": "Great stuff @honglihang congrats on your first solo gold and thanks for sharing this writeup + code 🎉 🔥",
      "votes": null
    },
    {
      "id": "2638854",
      "postDate": "02/06/2024 14:18:58",
      "content": "<p>Here a little late, but…<br>\nYour solution is incredibly inventive and probably the easiest to follow compared to all the ones I've viewed from the last 3 years. I hope to use it to finetune on just 2 bird species. I was wondering if you still have your .ckpt files that i could use, so that I don't have to train all over again? Thanks :)</p>",
      "rawMarkdown": "Here a little late, but...\nYour solution is incredibly inventive and probably the easiest to follow compared to all the ones I've viewed from the last 3 years. I hope to use it to finetune on just 2 bird species. I was wondering if you still have your .ckpt files that i could use, so that I don't have to train all over again? Thanks :)",
      "votes": null
    },
    {
      "id": "2782262",
      "postDate": "04/29/2024 08:23:10",
      "content": "<p>Hello!<br>\nThanks for the very nice written code!</p>\n<p>When i download the .csv file, and rename it to train.csv and put it in ./input, i get this error when trying to train:<br>\nKeyError: 'duration'<br>\nIt comes from <br>\nmodules/preprocess.py\", line 104, in preprocess<br>\n    (df[\"duration\"] &lt;= cfg.background_duration_thre)<br>\nAny idea on how to fix this?</p>",
      "rawMarkdown": "Hello!\nThanks for the very nice written code!\n\nWhen i download the .csv file, and rename it to train.csv and put it in ./input, i get this error when trying to train:\nKeyError: 'duration'\nIt comes from \nmodules/preprocess.py\", line 104, in preprocess\n    (df[\"duration\"] <= cfg.background_duration_thre)\nAny idea on how to fix this?",
      "votes": null
    },
    {
      "id": "3208332",
      "postDate": "05/23/2025 22:48:03",
      "content": "<p>Thanks for your contrubition. I have a question about re-producing the pipeline. Can I produce the same results via github or via kaggle notebook. As I understood, I need to construct the directory in github on my own hardware or maybe on colab. But, I can get the same result from the kaggle notebook by running it on kaggle kernel. <a href=\"https://www.kaggle.com/honglihang\" target=\"_blank\">@honglihang</a> </p>",
      "rawMarkdown": "Thanks for your contrubition. I have a question about re-producing the pipeline. Can I produce the same results via github or via kaggle notebook. As I understood, I need to construct the directory in github on my own hardware or maybe on colab. But, I can get the same result from the kaggle notebook by running it on kaggle kernel. @honglihang",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2273003,
      "author_name": "snnclsr",
      "author_url": "",
      "post_date": "05/25/2023 00:23:15",
      "content": "<p>Huge congrats!!! 🎉🎉 Thanks for sharing, I was really curious about the top solutions. I'll read this in detail after having my (missing) night sleep.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2276736,
          "author_name": "honglihang",
          "author_url": "",
          "post_date": "05/27/2023 06:47:57",
          "content": "<p>Thank you! Also congrats to you!<br>\nThe other top solutions are very incredible to me. I will also read in detial after getting back to sleep..</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2273006,
      "author_name": "harshitsheoran",
      "author_url": "",
      "post_date": "05/25/2023 00:27:20",
      "content": "<p>Congratulations! It must have been a fulfilling journey :)</p>",
      "votes": null,
      "replies": [
        {
          "id": 2276738,
          "author_name": "honglihang",
          "author_url": "",
          "post_date": "05/27/2023 06:51:45",
          "content": "<p>The same to your team! Congratulations! It has been a really a fulfilling(and short sleep) journey for me since the RSNA.<br>\nLooks that you have made birdnet work, I will read your solution in detail after getting enough sleep.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2273013,
      "author_name": "liuzhangzhen",
      "author_url": "",
      "post_date": "05/25/2023 00:36:22",
      "content": "<p>Congratulations, thanks for sharing</p>",
      "votes": null,
      "replies": [
        {
          "id": 2276733,
          "author_name": "honglihang",
          "author_url": "",
          "post_date": "05/27/2023 06:45:51",
          "content": "<p>Thank you and congratulations to you too!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2273018,
      "author_name": "maxdiazbattan",
      "author_url": "",
      "post_date": "05/25/2023 00:51:53",
      "content": "<p>Congratulations, and thanks for sharing! We tried onnx but not openvino. Onnx worked very well individually, but when we tried to ensemble the models, the results were very bad, almost all the time we ended up with a timeout or some kind of error, probably because we made some mistake during the ensemble. For that reason, we skip onnx, and our ensemble was with only 4 Pytorch models.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2276739,
          "author_name": "honglihang",
          "author_url": "",
          "post_date": "05/27/2023 06:55:23",
          "content": "<p>Also congratulations to you! Thanks for sharing the information about onnx. I am also new to onnx and openvino and maybe it is lucky for me to try openvino at first because I learnt about openvino in a project at work</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2273051,
      "author_name": "nomorevotch",
      "author_url": "",
      "post_date": "05/25/2023 01:34:15",
      "content": "<p>Congratulations on achieving second place as a solo participant! Incorporating \"heavy\" augmentations like pitch shifting seems like it would significantly extend the training time. Have you taken any measures to address this?</p>",
      "votes": null,
      "replies": [
        {
          "id": 2276730,
          "author_name": "honglihang",
          "author_url": "",
          "post_date": "05/27/2023 06:44:36",
          "content": "<p>Thank you. I appreciate you sharing the data in 2020 comp and I took full use of it. I used <a href=\"https://github.com/asteroid-team/torch-audiomentations\" target=\"_blank\">torch-audiomentations</a> to perform pitch shift and time shift on GPU</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2273056,
      "author_name": "m7mdalbaddawi",
      "author_url": "",
      "post_date": "05/25/2023 01:43:23",
      "content": "<p>Congratulations and thanks for sharing</p>",
      "votes": null,
      "replies": [
        {
          "id": 2276732,
          "author_name": "honglihang",
          "author_url": "",
          "post_date": "05/27/2023 06:45:21",
          "content": "<p>Thank you! happy kaggle</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2273190,
      "author_name": "aadityabansalcodes",
      "author_url": "",
      "post_date": "05/25/2023 04:17:10",
      "content": "<p>Congratulations ✨ and thanks for sharing, I am just a beginner getting started with Deep Learning this year only, for me it is both fascinating and intimidating how you can figure out so many things while approaching a problem. Do you have any tips about how I can become more efficient in problem solving like this. Thank you for your time.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2276743,
          "author_name": "honglihang",
          "author_url": "",
          "post_date": "05/27/2023 07:00:31",
          "content": "<p>Thank you! I started kaggle about 2 years ago and also had a hard time to understand the solutions. I think it is a process of try and error based on the knowledge shared by others. It is better to get new ideas through reading about the papers and the past top solutions.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2273253,
      "author_name": "scipygaurav",
      "author_url": "",
      "post_date": "05/25/2023 05:01:57",
      "content": "<p>Amazing 💯 Congratulations and thanks for sharing 😊</p>",
      "votes": null,
      "replies": [
        {
          "id": 2276700,
          "author_name": "honglihang",
          "author_url": "",
          "post_date": "05/27/2023 06:23:07",
          "content": "<p>Thanks! It is really a long journey for me and I slept for two days to recover from lack of sleep </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2273274,
      "author_name": "atsunorifujita",
      "author_url": "",
      "post_date": "05/25/2023 05:13:01",
      "content": "<p>Congratulations and thanks for sharing! I would like to learn Openvino from your notebook.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2276728,
          "author_name": "honglihang",
          "author_url": "",
          "post_date": "05/27/2023 06:42:39",
          "content": "<p>Thank you! The same to you. Your solution is super incredible and intersting. I will take a deep look into your solution after I get enougth sleep~</p>\n<p>By the way, have you already got the winner email from the kaggle team? I haven't…</p>",
          "votes": null,
          "replies": [
            {
              "id": 2276834,
              "author_name": "atsunorifujita",
              "author_url": "",
              "post_date": "05/27/2023 08:35:37",
              "content": "<p>Thanks!<br>\nI also haven't got it yet.</p>",
              "votes": null,
              "replies": [
                {
                  "id": 2279058,
                  "author_name": "honglihang",
                  "author_url": "",
                  "post_date": "05/29/2023 06:30:52",
                  "content": "<p>Thanks for your reply.<br>\nI still haven't received yet, maybe it cost sometime….</p>",
                  "votes": null,
                  "replies": []
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 2273425,
      "author_name": "adegladius",
      "author_url": "",
      "post_date": "05/25/2023 06:42:18",
      "content": "<p>Incredible !!! Congratulations 🎉. </p>",
      "votes": null,
      "replies": [
        {
          "id": 2276725,
          "author_name": "honglihang",
          "author_url": "",
          "post_date": "05/27/2023 06:39:00",
          "content": "<p>Thanks! Happy kaggling!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2273477,
      "author_name": "alabibojesomo",
      "author_url": "",
      "post_date": "05/25/2023 07:20:54",
      "content": "<p>Congratulations…..</p>\n<p>Ensembling trick of openvino is indeed something to add to my knowledge arsenal. Great work!!!</p>",
      "votes": null,
      "replies": [
        {
          "id": 2276724,
          "author_name": "honglihang",
          "author_url": "",
          "post_date": "05/27/2023 06:38:39",
          "content": "<p>Thank you! It is also my first time to try openvino.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2273496,
      "author_name": "muhinyiawandegwa",
      "author_url": "",
      "post_date": "05/25/2023 07:33:56",
      "content": "<p>Congratulations!</p>\n<p>Amazing stuff!</p>",
      "votes": null,
      "replies": [
        {
          "id": 2276706,
          "author_name": "honglihang",
          "author_url": "",
          "post_date": "05/27/2023 06:24:37",
          "content": "<p>Thank you!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2273602,
      "author_name": "janschl",
      "author_url": "",
      "post_date": "05/25/2023 09:16:02",
      "content": "<p>Very nice, congratulations and thanks for the quick writeup! Will look into openvino as well, nice find.</p>\n<blockquote>\n  <p>I hand labeled about 1800 records but did not see improvement.</p>\n</blockquote>\n<p>So in the end, you neither used manual nor pseudo-labels? I experimented with pseudo-labels on the test set (with much weaker models than the top solutions), but hardly got an improvement either, hence I'm curious.</p>\n<blockquote>\n  <p>self mixup for records with 2023 species only in background</p>\n</blockquote>\n<p>Interesting idea! What was the intuition behind that? Make it more likely that the bird in question is actually part of the 10s clip? How much did it help, compared to other tricks?</p>\n<blockquote>\n  <p>In both stages, I first train model with CrossEntropyLoss, and then train on with BCEWithLogitsLoss(reduction='sum').</p>\n</blockquote>\n<p>Again, interesting idea, will keep this in my head!</p>\n<blockquote>\n  <p>For SED model, feed model 10 sec chunk BUT apply head only on centered 5 sec reduced CNN image and use max(framewise, dim=time).</p>\n</blockquote>\n<p>This sounds expensive. Does this double the inference time or did you feed longer chunks than 10s and then extracted multiple 5sec chunks from that?</p>\n<blockquote>\n  <p>Also, tta(2s) is used for SED model.</p>\n</blockquote>\n<p>tta = test-time augmentation? Was there even time for that? :-O</p>\n<p>Again, congratulations for your work!</p>",
      "votes": null,
      "replies": [
        {
          "id": 2276721,
          "author_name": "honglihang",
          "author_url": "",
          "post_date": "05/27/2023 06:37:25",
          "content": "<p>Hi, thanks for your question.</p>\n<blockquote>\n  <p>I hand labeled about 1800 records but did not see improvement.</p>\n</blockquote>\n<p>I used both pseudo label and hand label, although there is no improvement seen in public LB, I have not confirm whether it has improvement in private LB.</p>\n<blockquote>\n  <p>self mixup for records with 2023 species only in background.</p>\n</blockquote>\n<p>The idea comes from that I found mixup in waveform is effective. So I think I can take use of the records with 2023 species only in background by applying the same idea because I no longer need to care about where the bird call appears in background by doing this mixup to fully use the record. I have not confirmed whether it is useful, but there was a marginal improvement in the score.</p>\n<blockquote>\n  <p>For SED model, feed model 10 sec chunk BUT apply head only on centered 5 sec reduced CNN image and use max(framewise, dim=time).</p>\n</blockquote>\n<p>Yes, it is expensive which double the inference time. So I only use 10s, not any longer…</p>\n<blockquote>\n  <p>Also, tta(2s) is used for SED model.</p>\n</blockquote>\n<p>TTA is only applied on attention head so it is super fast. The encoder extract the 10s feature and the target 5s is in the center of the matrix (2.5s to 7.5s). For tta, I take the feature from 1.5s to 6.5s and 3.5s to 8.5s.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2274088,
      "author_name": "bilalwaseer",
      "author_url": "",
      "post_date": "05/25/2023 15:57:42",
      "content": "<p><a href=\"https://www.kaggle.com/honglihang\" target=\"_blank\">@honglihang</a> I wanted to reach out and express my heartfelt gratitude for sharing your incredible 2nd place solution in the recent competition. Your approach, combining Sound Event Detection (SED) with Convolutional Neural Networks (CNN) and utilizing an ensemble of seven models, is truly remarkable.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2276704,
          "author_name": "honglihang",
          "author_url": "",
          "post_date": "05/27/2023 06:24:25",
          "content": "<p>Thanks you! We should all thanks to those who shared these amazing stuffs in the past few years.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2276977,
      "author_name": "suraj520",
      "author_url": "",
      "post_date": "05/27/2023 10:58:58",
      "content": "<p>Awesome approach to use openvino for boosting inference times over ensembling <a href=\"https://www.kaggle.com/honglihang\" target=\"_blank\">@honglihang</a> ! Congrats !</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2277433,
      "author_name": "salehahmad1",
      "author_url": "",
      "post_date": "05/27/2023 18:12:15",
      "content": "<p>That is one amazing solution to the competition!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2280136,
      "author_name": "joaopedrorebeschini",
      "author_url": "",
      "post_date": "05/29/2023 22:08:59",
      "content": "<p>Awesome job!!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2285781,
      "author_name": "pardeep19singh",
      "author_url": "",
      "post_date": "06/03/2023 02:04:06",
      "content": "<p>Great stuff <a href=\"https://www.kaggle.com/honglihang\" target=\"_blank\">@honglihang</a> congrats on your first solo gold and thanks for sharing this writeup + code 🎉 🔥</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2638854,
      "author_name": "ameeassad",
      "author_url": "",
      "post_date": "02/06/2024 14:18:58",
      "content": "<p>Here a little late, but…<br>\nYour solution is incredibly inventive and probably the easiest to follow compared to all the ones I've viewed from the last 3 years. I hope to use it to finetune on just 2 bird species. I was wondering if you still have your .ckpt files that i could use, so that I don't have to train all over again? Thanks :)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2782262,
      "author_name": "jonathancarlson2",
      "author_url": "",
      "post_date": "04/29/2024 08:23:10",
      "content": "<p>Hello!<br>\nThanks for the very nice written code!</p>\n<p>When i download the .csv file, and rename it to train.csv and put it in ./input, i get this error when trying to train:<br>\nKeyError: 'duration'<br>\nIt comes from <br>\nmodules/preprocess.py\", line 104, in preprocess<br>\n    (df[\"duration\"] &lt;= cfg.background_duration_thre)<br>\nAny idea on how to fix this?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3208332,
      "author_name": "ayagler",
      "author_url": "",
      "post_date": "05/23/2025 22:48:03",
      "content": "<p>Thanks for your contrubition. I have a question about re-producing the pipeline. Can I produce the same results via github or via kaggle notebook. As I understood, I need to construct the directory in github on my own hardware or maybe on colab. But, I can get the same result from the kaggle notebook by running it on kaggle kernel. <a href=\"https://www.kaggle.com/honglihang\" target=\"_blank\">@honglihang</a> </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2272995": "Congratulations to all the winners! Thanks to Kaggle and Cornell Lab of Ornithology for hosting this interesting competition.\n\nThis is my first solo gold medal and I am glad to have this result.\n\nThis competition shared a lot of similarity to the past BirdClef competitions(2020/2021/2022). Thus I spent a lot of time gathering the solution shared by the top teams in the past competitions. Special thanks to all of you for sharing such important information!\n\nLet me briefly introduce my solution. I will update the solution for more details in a couple of days.\n\n# Most important (7 models ensemble!)\nPlease see the notebook below.\n[openvino is all you need!!](https://www.kaggle.com/code/honglihang/openvino-is-all-you-need)\n\n# Training data\n\nHere is my training data.\n\n- 2023/2022/2021/2020 competition data\n- [2020 additional competition data](https://www.kaggle.com/competitions/birdclef-2023/discussion/398318)\n- additional training data from xeno-canto, including 2023 comp species in both foreground and background(records with 2023 comp species only in background which is less than 60 seconds are included). \n\nI intended to collect more records from ebird site, but I realized that ebird data is not public and cannot be used. I [asked the host](https://www.kaggle.com/competitions/birdclef-2023/discussion/393023) and confirmed that. Thanks [@tomdenton](https://www.kaggle.com/tomdenton) answering my questions.\n\nThus, my training pipeline does not contain records from ebird site.\n\n# Model Architecture\n\nFirst, I used SED architecture. The same as yours.\n\nbackbones are:\n\n- tf_efficientnetv2_s_in21k\n- seresnext26t_32x4d\n- tf_efficientnet_b3_ns\n\nAll of them are trained on 10sec clip.\n\nSecond, I used CNN proposed by [2nd place of 2021 competition](https://www.kaggle.com/competitions/birdclef-2021/discussion/243463)\n\nbackbones are:\n\n- tf_efficientnetv2_s_in21k\n- resnet34d\n- tf_efficientnet_b3_ns\n- tf_efficientnet_b0_ns\n\nAll except b0 are trained on 15sec clip. b0 is trained on 20sec clip.\n\n# Pseudo Labeling and Hand Labeling\n\nI have used SED model to generate pseudo label and extracted the potential nocall using quantile threshold. Then, I hand labeled the potential nocall by hearing the record. I hand labeled about 1800 records but did not see improvement. Maybe the pseudo label contains more FP rather than FN. I did not have time to further investigate the prediction.\n\n# Model Training\n\naugmentations:\n\n- GaussianNoise\n- PinkNoise\n- Gain\n- NoiseInjection\n- Background Noise(nocall in 2020, 2021 comp + rainforest + environment sound + nocall in freefield1010, warblrb, birdvox)\n- PitchShift\n- TimeShift\n- FrequencyMasking\n- TimeMasking\n- OR Mixup on waveforms\n- Mixup on spectrograms.\n- [With a probability of 0.5 lowered the upper frequencies](https://www.kaggle.com/competitions/birdsong-recognition/discussion/183269)\n- self mixup for records with 2023 species only in background.(60sec waveform -> split to 6 \\* 10sec -> np.sum(audios,axis=0) to get a 10sec clip)\n\nI have used weights (computed by primary_label and secondary_labels) for Dataloader in order to cope with unbalanced dataset.\n\n# Training stages\n\nFor training I have used 2 stage training:\n\n1. Pretrain on all data(834 species).\n2. Finetune on 2023 species(264 species).\n\nIn both stages, I first train model with CrossEntropyLoss, and then train on with BCEWithLogitsLoss(reduction='sum'). Model converges faster with CrossEntropyLoss than BCEWithLogitsLoss, but BCEWithLogitsLoss gives better score.\n\nTo give more diversity, models are trained on different windows and different mixup rate, and some of them only trained on CrossEntropyLoss. And also 3 of the models are fintuned on 30s clip.\n\n# CV strategy\n\n- For each validation sample - slice the first 60 seconds to pieces -> predict each piece -> max(sample_predictions, dim=pieces).\n\nCV does not show correlation with LB, but it seems that the right ways to improve the LB are those which do not significantly decrease CV. So I monitored the CV when tuning the pipeline.\n\n# Inference\n\nFor SED model, feed model 10 sec chunk BUT apply head only on centered 5 sec reduced CNN image and use max(framewise, dim=time).\n\nAlso, tta(2s) is used for SED model.\n\nImportant: convert pytorch model to openvino model significantly reduce inference time(about 40%). (eca_nfnet_l0 backbone ONXX cannot be converted to openvino because the stdconv layer in timm use train mode of F.batch_norm in forward method). That is the magic of ensembling 7 models.\n\n# Ensemble\n\nI spent quite a lot of time understanding the metrics. The ensemble is as followed:\n\n1. (weighted average, 0.84 on LB, 0.76 on private) Apply weighted average on raw logit. This ensemble does not make sense for me because the output logit of models differ and should not be simply added, otherwise the result is biased. But considering that the absolute value of logit may also contribute to the score and it does give the best LB, so I choose it for final submission to have a gamble.\n2. (rank average, 0.83 on LB, 0.75 on private) Convert the logit to rank and apply weighted average on rank. I think this is the reasonable way to ensemble, considering that I don't have a reliable CV. <strong>Be careful that in this comp submission is padded with 5 rows of 1, thus the ranking should start from 0 to prevent the largest ranking to be 1 after convert rankings to percentile form.(For example, 0, 0.333, 0.666 is good while 0.333, 0.666, 1.0 is bad. otherwise the score will decrease due to the potential false positive of those ranked 1.0) </strong>\n\nBasically all of my single model's performance reaches about 0.81-0.82 on LB, so the weight of each model is similar to give diversity although higher public lb requires large weight of sed v2s and sed b3_ns.\n\nLooking at the private LB, weighted average seems to be better. emmm......Why?\n\n# What did not work\n\n- [contrast](https://www.kaggle.com/competitions/birdsong-recognition/discussion/183269)\n- EMA\n- BirdNet\n- adjusting logit according to previous and next 5s clip\n- q transform\n\nInference Notebook: https://www.kaggle.com/code/honglihang/2nd-place-solution-inference-kernel\ngithub: https://github.com/LIHANG-HONG/birdclef2023-2nd-place-solution",
    "2273003": "Huge congrats!!! 🎉🎉 Thanks for sharing, I was really curious about the top solutions. I'll read this in detail after having my (missing) night sleep.",
    "2273006": "Congratulations! It must have been a fulfilling journey :)",
    "2273013": "Congratulations, thanks for sharing",
    "2273018": "Congratulations, and thanks for sharing! We tried onnx but not openvino. Onnx worked very well individually, but when we tried to ensemble the models, the results were very bad, almost all the time we ended up with a timeout or some kind of error, probably because we made some mistake during the ensemble. For that reason, we skip onnx, and our ensemble was with only 4 Pytorch models.",
    "2273051": "Congratulations on achieving second place as a solo participant! Incorporating \"heavy\" augmentations like pitch shifting seems like it would significantly extend the training time. Have you taken any measures to address this?",
    "2273056": "Congratulations and thanks for sharing",
    "2273190": "Congratulations ✨ and thanks for sharing, I am just a beginner getting started with Deep Learning this year only, for me it is both fascinating and intimidating how you can figure out so many things while approaching a problem. Do you have any tips about how I can become more efficient in problem solving like this. Thank you for your time.",
    "2273253": "Amazing 💯 Congratulations and thanks for sharing 😊",
    "2273274": "Congratulations and thanks for sharing! I would like to learn Openvino from your notebook.",
    "2273425": "Incredible !!! Congratulations 🎉.",
    "2273477": "Congratulations.....\n\nEnsembling trick of openvino is indeed something to add to my knowledge arsenal. Great work!!!",
    "2273496": "Congratulations!\n\nAmazing stuff!",
    "2273602": "Very nice, congratulations and thanks for the quick writeup! Will look into openvino as well, nice find.\n\n> I hand labeled about 1800 records but did not see improvement.\n\nSo in the end, you neither used manual nor pseudo-labels? I experimented with pseudo-labels on the test set (with much weaker models than the top solutions), but hardly got an improvement either, hence I'm curious.\n\n> self mixup for records with 2023 species only in background\n\nInteresting idea! What was the intuition behind that? Make it more likely that the bird in question is actually part of the 10s clip? How much did it help, compared to other tricks?\n\n> In both stages, I first train model with CrossEntropyLoss, and then train on with BCEWithLogitsLoss(reduction='sum').\n\nAgain, interesting idea, will keep this in my head!\n\n> For SED model, feed model 10 sec chunk BUT apply head only on centered 5 sec reduced CNN image and use max(framewise, dim=time).\n\nThis sounds expensive. Does this double the inference time or did you feed longer chunks than 10s and then extracted multiple 5sec chunks from that?\n\n> Also, tta(2s) is used for SED model.\n\ntta = test-time augmentation? Was there even time for that? :-O\n\nAgain, congratulations for your work!",
    "2274088": "honglihang I wanted to reach out and express my heartfelt gratitude for sharing your incredible 2nd place solution in the recent competition. Your approach, combining Sound Event Detection (SED) with Convolutional Neural Networks (CNN) and utilizing an ensemble of seven models, is truly remarkable.",
    "2276700": "Thanks! It is really a long journey for me and I slept for two days to recover from lack of sleep",
    "2276704": "Thanks you! We should all thanks to those who shared these amazing stuffs in the past few years.",
    "2276706": "Thank you!",
    "2276721": "Hi, thanks for your question.\n>I hand labeled about 1800 records but did not see improvement.\n\nI used both pseudo label and hand label, although there is no improvement seen in public LB, I have not confirm whether it has improvement in private LB.\n>self mixup for records with 2023 species only in background.\n\nThe idea comes from that I found mixup in waveform is effective. So I think I can take use of the records with 2023 species only in background by applying the same idea because I no longer need to care about where the bird call appears in background by doing this mixup to fully use the record. I have not confirmed whether it is useful, but there was a marginal improvement in the score.\n\n>For SED model, feed model 10 sec chunk BUT apply head only on centered 5 sec reduced CNN image and use max(framewise, dim=time).\n\nYes, it is expensive which double the inference time. So I only use 10s, not any longer...\n\n>Also, tta(2s) is used for SED model.\n\nTTA is only applied on attention head so it is super fast. The encoder extract the 10s feature and the target 5s is in the center of the matrix (2.5s to 7.5s). For tta, I take the feature from 1.5s to 6.5s and 3.5s to 8.5s.",
    "2276724": "Thank you! It is also my first time to try openvino.",
    "2276725": "Thanks! Happy kaggling!",
    "2276728": "Thank you! The same to you. Your solution is super incredible and intersting. I will take a deep look into your solution after I get enougth sleep~\n\nBy the way, have you already got the winner email from the kaggle team? I haven't...",
    "2276730": "Thank you. I appreciate you sharing the data in 2020 comp and I took full use of it. I used [torch-audiomentations](https://github.com/asteroid-team/torch-audiomentations) to perform pitch shift and time shift on GPU",
    "2276732": "Thank you! happy kaggle",
    "2276733": "Thank you and congratulations to you too!",
    "2276736": "Thank you! Also congrats to you!\nThe other top solutions are very incredible to me. I will also read in detial after getting back to sleep..",
    "2276738": "The same to your team! Congratulations! It has been a really a fulfilling(and short sleep) journey for me since the RSNA.\nLooks that you have made birdnet work, I will read your solution in detail after getting enough sleep.",
    "2276739": "Also congratulations to you! Thanks for sharing the information about onnx. I am also new to onnx and openvino and maybe it is lucky for me to try openvino at first because I learnt about openvino in a project at work",
    "2276743": "Thank you! I started kaggle about 2 years ago and also had a hard time to understand the solutions. I think it is a process of try and error based on the knowledge shared by others. It is better to get new ideas through reading about the papers and the past top solutions.",
    "2276834": "Thanks!\nI also haven't got it yet.",
    "2276977": "Awesome approach to use openvino for boosting inference times over ensembling @honglihang ! Congrats !",
    "2277433": "That is one amazing solution to the competition!",
    "2279058": "Thanks for your reply.\nI still haven't received yet, maybe it cost sometime....",
    "2280136": "Awesome job!!",
    "2285781": "Great stuff @honglihang congrats on your first solo gold and thanks for sharing this writeup + code 🎉 🔥",
    "2638854": "Here a little late, but...\nYour solution is incredibly inventive and probably the easiest to follow compared to all the ones I've viewed from the last 3 years. I hope to use it to finetune on just 2 bird species. I was wondering if you still have your .ckpt files that i could use, so that I don't have to train all over again? Thanks :)",
    "2782262": "Hello!\nThanks for the very nice written code!\n\nWhen i download the .csv file, and rename it to train.csv and put it in ./input, i get this error when trying to train:\nKeyError: 'duration'\nIt comes from \nmodules/preprocess.py\", line 104, in preprocess\n    (df[\"duration\"] <= cfg.background_duration_thre)\nAny idea on how to fix this?",
    "3208332": "Thanks for your contrubition. I have a question about re-producing the pipeline. Can I produce the same results via github or via kaggle notebook. As I understood, I need to construct the directory in github on my own hardware or maybe on colab. But, I can get the same result from the kaggle notebook by running it on kaggle kernel. @honglihang"
  },
  "source": "meta"
}