{
  "id": 220450,
  "title": "17th place solution",
  "url": "/competitions/rfcx-species-audio-detection/writeups/hiraiwa-17th-place-solution",
  "author_name": "",
  "post_date": "2021-02-18T10:06:51.443614700Z",
  "votes": 34,
  "comment_count": 15,
  "views": 0,
  "content": "<p>Congratulations to all the winners and gold getters, I guess those teams that broke the 0.950 wall have found the essence of this competition, which we couldn't.  <br>\n<br><br>\nFirstly, thanks to the host for holding quite an interesting competition. Partly labeled classification is a challenging task, which made this competition more interesting than a simple bioacoustics audio tagging competition.<br>\n<br><br>\nOur solution is a ranking average of image classification models and SED models. <a href=\"https://www.kaggle.com/yasufuminakama\" target=\"_blank\">Y.Nakama</a>, <a href=\"https://www.kaggle.com/kaerunantoka\" target=\"_blank\">kaerururu</a>, and I worked a lot on SED models but couldn't break 0.90 until we merge with <a href=\"https://www.kaggle.com/thiraiwa\" target=\"_blank\">Taku Hiraiwa</a>. His model was based on image classification similar to <a href=\"https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/220304\" target=\"_blank\">@cpmpml's model</a>. We found that quite good , so we focused on improving image classification models in the remained days.</p>\n<h2>Image classification models</h2>\n<p>It was <a href=\"https://www.kaggle.com/thiraiwa\" target=\"_blank\">Taku Hiraiwa</a>'s idea to only use the <em>annotated</em> part of the train data. To do so, we crop the image patches from log-melspectrogram of train data based on the <code>t_min</code>, <code>t_max</code>, <code>f_min</code>, <code>f_max</code> information of train_tp.csv and train_fp.csv, and resized the patch to fixed shape (say, 320 x 320 or so). The cropping is performed on the fly through training, so the part we crop out is randomized along with time axis.<br>\nWith these image patches we trained EfficientNet models, and monitored F1 score with threshold 0.5. <br>\nHere's the other details of image classification models.</p>\n<ul>\n<li>image size: varies from (244, 244) to (456, 456) between models</li>\n<li>backbone: EfficientNetB0 - B5 (used <a href=\"https://rwightman.github.io/pytorch-image-models/\" target=\"_blank\"><code>timm</code></a> and used <code>tf_efficientnet_b&lt;0-5&gt;_ns</code> weights).</li>\n<li>augmentation: GaussianNoise, Gain, PitchShift of <a href=\"https://github.com/iver56/audiomentations\" target=\"_blank\">audiomentations</a> on raw waveform. Also HorizontalFlip also had positive impact on LB slightly, so we used (but don't know why it worked).</li>\n<li>AdamW optimizer with linear warmup scheduler</li>\n<li>BCEFocalLoss<br>\n<br><br>\nIn the end, we trained a stacking model that takes the output of models below which achieve public 0.942:</li>\n</ul>\n<ol>\n<li><code>tf_efficientnet_b0_ns</code> image size 244</li>\n<li><code>tf_efficientnet_b0_ns</code> image size 320</li>\n<li><code>tf_efficientnet_b0_ns</code> image size 456</li>\n<li><code>tf_efficientnet_b1_ns</code> image size 456</li>\n<li><code>tf_efficientnet_b2_ns</code> image size 456</li>\n<li><code>tf_efficientnet_b3_ns</code> image size 456</li>\n<li><code>tf_efficientnet_b4_ns</code> image size 456</li>\n<li><code>tf_efficientnet_b5_ns</code> image size 456</li>\n</ol>\n<h2>SED models</h2>\n<p>All of our SED models use the head architecture introduced in <a href=\"https://github.com/qiuqiangkong/audioset_tagging_cnn\" target=\"_blank\">PANNs repository</a>. The CNN encoder is either EfficientNet or ResNeSt, and they are trained with weak/strong supervision. We tried a lot of things on this, but couldn't find factors that consistently work well - which means the LB scores varied quite randomly w.r.t CV score.<br>\nOur best SED model is rank average of 11 models(public: 0.901, privte: 0.911) below - each of them differs slightly so we describe the difference briefly.</p>\n<h3>2 x kaerururu's model (public: 0.882, 0.873)</h3>\n<ul>\n<li>Based on public starter SED (EffnB0) notebook (<a href=\"https://www.kaggle.com/gopidurgaprasad/rfcx-sed-model-stater\" target=\"_blank\">https://www.kaggle.com/gopidurgaprasad/rfcx-sed-model-stater</a>)</li>\n<li>3ch input</li>\n<li>10sec clip</li>\n<li>waveform mixup</li>\n<li>some augmentation (audiomentations)</li>\n<li>pseudo-labeled datasets (add labels on tp data)</li>\n<li>trained with tp and fp dataset (1st training)</li>\n<li>trained with pseudo-labeled tp (2nd training)</li>\n<li>tta=10</li>\n</ul>\n<h3>5 x arai's model (public: 0.879,0.880, 0.868, 0.874, 0.870)</h3>\n<ul>\n<li>Based on Birdcall's challenge 6th place (<a href=\"https://github.com/koukyo1994/kaggle-birdcall-6th-place\" target=\"_blank\">https://github.com/koukyo1994/kaggle-birdcall-6th-place</a>)</li>\n<li>ResNeSt50 encoder or EfficientNetB3 encoder</li>\n<li>AddPinkNoiseSNR / VolumeControl / PitchShift from <a href=\"https://www.kaggle.com/hidehisaarai1213/rfcx-audio-data-augmentation-japanese-english\" target=\"_blank\">My Notebook</a></li>\n<li>tp only</li>\n</ul>\n<h3>4 x Y.Nakama's model (public: 0.871, 0.863, 0.866, 0.870)</h3>\n<ul>\n<li>Based on Birdcall's challenge 6th place (<a href=\"https://github.com/koukyo1994/kaggle-birdcall-6th-place\" target=\"_blank\">https://github.com/koukyo1994/kaggle-birdcall-6th-place</a>)</li>\n<li>ResNeSt50 encoder or ResNet200D encoder</li>\n<li>mixup &amp; some augmentations</li>\n<li>2nd stage training<ul>\n<li>1st stage: weighted loss for framewise_logit &amp; logit</li>\n<li>2nd stage: loss for logit</li></ul></li>\n<li>tp only</li>\n</ul>",
  "messages": [
    {
      "id": "1208512",
      "postDate": "02/18/2021 10:06:51",
      "content": "<p>Congratulations to all the winners and gold getters, I guess those teams that broke the 0.950 wall have found the essence of this competition, which we couldn't.  <br>\n<br><br>\nFirstly, thanks to the host for holding quite an interesting competition. Partly labeled classification is a challenging task, which made this competition more interesting than a simple bioacoustics audio tagging competition.<br>\n<br><br>\nOur solution is a ranking average of image classification models and SED models. <a href=\"https://www.kaggle.com/yasufuminakama\" target=\"_blank\">Y.Nakama</a>, <a href=\"https://www.kaggle.com/kaerunantoka\" target=\"_blank\">kaerururu</a>, and I worked a lot on SED models but couldn't break 0.90 until we merge with <a href=\"https://www.kaggle.com/thiraiwa\" target=\"_blank\">Taku Hiraiwa</a>. His model was based on image classification similar to <a href=\"https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/220304\" target=\"_blank\">@cpmpml's model</a>. We found that quite good , so we focused on improving image classification models in the remained days.</p>\n<h2>Image classification models</h2>\n<p>It was <a href=\"https://www.kaggle.com/thiraiwa\" target=\"_blank\">Taku Hiraiwa</a>'s idea to only use the <em>annotated</em> part of the train data. To do so, we crop the image patches from log-melspectrogram of train data based on the <code>t_min</code>, <code>t_max</code>, <code>f_min</code>, <code>f_max</code> information of train_tp.csv and train_fp.csv, and resized the patch to fixed shape (say, 320 x 320 or so). The cropping is performed on the fly through training, so the part we crop out is randomized along with time axis.<br>\nWith these image patches we trained EfficientNet models, and monitored F1 score with threshold 0.5. <br>\nHere's the other details of image classification models.</p>\n<ul>\n<li>image size: varies from (244, 244) to (456, 456) between models</li>\n<li>backbone: EfficientNetB0 - B5 (used <a href=\"https://rwightman.github.io/pytorch-image-models/\" target=\"_blank\"><code>timm</code></a> and used <code>tf_efficientnet_b&lt;0-5&gt;_ns</code> weights).</li>\n<li>augmentation: GaussianNoise, Gain, PitchShift of <a href=\"https://github.com/iver56/audiomentations\" target=\"_blank\">audiomentations</a> on raw waveform. Also HorizontalFlip also had positive impact on LB slightly, so we used (but don't know why it worked).</li>\n<li>AdamW optimizer with linear warmup scheduler</li>\n<li>BCEFocalLoss<br>\n<br><br>\nIn the end, we trained a stacking model that takes the output of models below which achieve public 0.942:</li>\n</ul>\n<ol>\n<li><code>tf_efficientnet_b0_ns</code> image size 244</li>\n<li><code>tf_efficientnet_b0_ns</code> image size 320</li>\n<li><code>tf_efficientnet_b0_ns</code> image size 456</li>\n<li><code>tf_efficientnet_b1_ns</code> image size 456</li>\n<li><code>tf_efficientnet_b2_ns</code> image size 456</li>\n<li><code>tf_efficientnet_b3_ns</code> image size 456</li>\n<li><code>tf_efficientnet_b4_ns</code> image size 456</li>\n<li><code>tf_efficientnet_b5_ns</code> image size 456</li>\n</ol>\n<h2>SED models</h2>\n<p>All of our SED models use the head architecture introduced in <a href=\"https://github.com/qiuqiangkong/audioset_tagging_cnn\" target=\"_blank\">PANNs repository</a>. The CNN encoder is either EfficientNet or ResNeSt, and they are trained with weak/strong supervision. We tried a lot of things on this, but couldn't find factors that consistently work well - which means the LB scores varied quite randomly w.r.t CV score.<br>\nOur best SED model is rank average of 11 models(public: 0.901, privte: 0.911) below - each of them differs slightly so we describe the difference briefly.</p>\n<h3>2 x kaerururu's model (public: 0.882, 0.873)</h3>\n<ul>\n<li>Based on public starter SED (EffnB0) notebook (<a href=\"https://www.kaggle.com/gopidurgaprasad/rfcx-sed-model-stater\" target=\"_blank\">https://www.kaggle.com/gopidurgaprasad/rfcx-sed-model-stater</a>)</li>\n<li>3ch input</li>\n<li>10sec clip</li>\n<li>waveform mixup</li>\n<li>some augmentation (audiomentations)</li>\n<li>pseudo-labeled datasets (add labels on tp data)</li>\n<li>trained with tp and fp dataset (1st training)</li>\n<li>trained with pseudo-labeled tp (2nd training)</li>\n<li>tta=10</li>\n</ul>\n<h3>5 x arai's model (public: 0.879,0.880, 0.868, 0.874, 0.870)</h3>\n<ul>\n<li>Based on Birdcall's challenge 6th place (<a href=\"https://github.com/koukyo1994/kaggle-birdcall-6th-place\" target=\"_blank\">https://github.com/koukyo1994/kaggle-birdcall-6th-place</a>)</li>\n<li>ResNeSt50 encoder or EfficientNetB3 encoder</li>\n<li>AddPinkNoiseSNR / VolumeControl / PitchShift from <a href=\"https://www.kaggle.com/hidehisaarai1213/rfcx-audio-data-augmentation-japanese-english\" target=\"_blank\">My Notebook</a></li>\n<li>tp only</li>\n</ul>\n<h3>4 x Y.Nakama's model (public: 0.871, 0.863, 0.866, 0.870)</h3>\n<ul>\n<li>Based on Birdcall's challenge 6th place (<a href=\"https://github.com/koukyo1994/kaggle-birdcall-6th-place\" target=\"_blank\">https://github.com/koukyo1994/kaggle-birdcall-6th-place</a>)</li>\n<li>ResNeSt50 encoder or ResNet200D encoder</li>\n<li>mixup &amp; some augmentations</li>\n<li>2nd stage training<ul>\n<li>1st stage: weighted loss for framewise_logit &amp; logit</li>\n<li>2nd stage: loss for logit</li></ul></li>\n<li>tp only</li>\n</ul>",
      "rawMarkdown": "Congratulations to all the winners and gold getters, I guess those teams that broke the 0.950 wall have found the essence of this competition, which we couldn't.  \n<br/>\nFirstly, thanks to the host for holding quite an interesting competition. Partly labeled classification is a challenging task, which made this competition more interesting than a simple bioacoustics audio tagging competition.\n<br/>\n\nOur solution is a ranking average of image classification models and SED models. [Y.Nakama](https://www.kaggle.com/yasufuminakama), [kaerururu](https://www.kaggle.com/kaerunantoka), and I worked a lot on SED models but couldn't break 0.90 until we merge with [Taku Hiraiwa](https://www.kaggle.com/thiraiwa). His model was based on image classification similar to [@cpmpml's model](https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/220304). We found that quite good , so we focused on improving image classification models in the remained days.\n## Image classification models\nIt was [Taku Hiraiwa](https://www.kaggle.com/thiraiwa)'s idea to only use the *annotated* part of the train data. To do so, we crop the image patches from log-melspectrogram of train data based on the `t_min`, `t_max`, `f_min`, `f_max` information of train\\_tp.csv and train\\_fp.csv, and resized the patch to fixed shape (say, 320 x 320 or so). The cropping is performed on the fly through training, so the part we crop out is randomized along with time axis.\nWith these image patches we trained EfficientNet models, and monitored F1 score with threshold 0.5. \nHere's the other details of image classification models.\n\n* image size: varies from (244, 244) to (456, 456) between models\n* backbone: EfficientNetB0 - B5 (used [`timm`](https://rwightman.github.io/pytorch-image-models/) and used `tf_efficientnet_b<0-5>_ns` weights).\n* augmentation: GaussianNoise, Gain, PitchShift of [audiomentations](https://github.com/iver56/audiomentations) on raw waveform. Also HorizontalFlip also had positive impact on LB slightly, so we used (but don't know why it worked).\n* AdamW optimizer with linear warmup scheduler\n* BCEFocalLoss\n<br/>\nIn the end, we trained a stacking model that takes the output of models below which achieve public 0.942:\n\n1. `tf_efficientnet_b0_ns` image size 244\n2. `tf_efficientnet_b0_ns` image size 320\n3. `tf_efficientnet_b0_ns` image size 456\n4. `tf_efficientnet_b1_ns` image size 456\n5. `tf_efficientnet_b2_ns` image size 456\n6. `tf_efficientnet_b3_ns` image size 456\n7. `tf_efficientnet_b4_ns` image size 456\n8. `tf_efficientnet_b5_ns` image size 456\n\n## SED models\nAll of our SED models use the head architecture introduced in [PANNs repository](https://github.com/qiuqiangkong/audioset_tagging_cnn). The CNN encoder is either EfficientNet or ResNeSt, and they are trained with weak/strong supervision. We tried a lot of things on this, but couldn't find factors that consistently work well - which means the LB scores varied quite randomly w.r.t CV score.\nOur best SED model is rank average of 11 models(public: 0.901, privte: 0.911) below - each of them differs slightly so we describe the difference briefly.\n\n### 2 x kaerururu's model (public: 0.882, 0.873)\n\n* Based on public starter SED (EffnB0) notebook (https://www.kaggle.com/gopidurgaprasad/rfcx-sed-model-stater)\n* 3ch input\n* 10sec clip\n* waveform mixup\n* some augmentation (audiomentations)\n* pseudo-labeled datasets (add labels on tp data)\n* trained with tp and fp dataset (1st training)\n* trained with pseudo-labeled tp (2nd training)\n* tta=10\n\n### 5 x arai's model (public: 0.879,0.880, 0.868, 0.874, 0.870)\n\n* Based on Birdcall's challenge 6th place (https://github.com/koukyo1994/kaggle-birdcall-6th-place)\n* ResNeSt50 encoder or EfficientNetB3 encoder\n* AddPinkNoiseSNR / VolumeControl / PitchShift from [My Notebook](https://www.kaggle.com/hidehisaarai1213/rfcx-audio-data-augmentation-japanese-english)\n* tp only\n### 4 x Y.Nakama's model (public: 0.871, 0.863, 0.866, 0.870)\n\n* Based on Birdcall's challenge 6th place (https://github.com/koukyo1994/kaggle-birdcall-6th-place)\n* ResNeSt50 encoder or ResNet200D encoder\n* mixup & some augmentations\n* 2nd stage training\n  * 1st stage: weighted loss for framewise_logit & logit\n  * 2nd stage: loss for logit\n* tp only",
      "votes": null
    },
    {
      "id": "1208562",
      "postDate": "02/18/2021 10:41:36",
      "content": "<p>Congrats on strong finish <a href=\"https://www.kaggle.com/hidehisaarai1213\" target=\"_blank\">@hidehisaarai1213</a> and team and thanks for sharing details solution </p>",
      "rawMarkdown": "Congrats on strong finish @hidehisaarai1213 and team and thanks for sharing details solution",
      "votes": null
    },
    {
      "id": "1208635",
      "postDate": "02/18/2021 11:45:20",
      "content": "<p>Congratz ! Good to see you at the top of another audio competition :) </p>",
      "rawMarkdown": "Congratz ! Good to see you at the top of another audio competition :)",
      "votes": null
    },
    {
      "id": "1208655",
      "postDate": "02/18/2021 12:06:57",
      "content": "<p>Congratulations🎉🎉</p>",
      "rawMarkdown": "Congratulations🎉🎉",
      "votes": null
    },
    {
      "id": "1208690",
      "postDate": "02/18/2021 12:44:04",
      "content": "<p>Your solution (Mixing Image classification models and SED) is interesting!</p>\n<p>I have a question. <br>\nIn Image classification models, can models detect multi label?</p>",
      "rawMarkdown": "Your solution (Mixing Image classification models and SED) is interesting!\n\nI have a question. \nIn Image classification models, can models detect multi label?",
      "votes": null
    },
    {
      "id": "1209458",
      "postDate": "02/18/2021 22:59:16",
      "content": "<p>Wow, you pushed SED quite far!  I am glad you mentioned my Cornell model :)</p>\n<p>Congrats on the strong ending,  I was afraid you'd pass me.</p>",
      "rawMarkdown": "Wow, you pushed SED quite far!  I am glad you mentioned my Cornell model :)\n\nCongrats on the strong ending,  I was afraid you'd pass me.",
      "votes": null
    },
    {
      "id": "1209645",
      "postDate": "02/19/2021 01:49:16",
      "content": "<p>is there any experimental results on COLA from \"Contrastive learning of general purpose audio representations\"?</p>",
      "rawMarkdown": "is there any experimental results on COLA from \"Contrastive learning of general purpose audio representations\"?",
      "votes": null
    },
    {
      "id": "1209689",
      "postDate": "02/19/2021 02:30:25",
      "content": "<p>Thank you! You've also finished in good place!</p>",
      "rawMarkdown": "Thank you! You've also finished in good place!",
      "votes": null
    },
    {
      "id": "1209693",
      "postDate": "02/19/2021 02:33:26",
      "content": "<p>Same to you! Also good to see many competitors of Birdcall comes to this competition. At first it seemed similar to birdcall but it turned out it wasn't - both are interesting in different ways</p>",
      "rawMarkdown": "Same to you! Also good to see many competitors of Birdcall comes to this competition. At first it seemed similar to birdcall but it turned out it wasn't - both are interesting in different ways",
      "votes": null
    },
    {
      "id": "1209696",
      "postDate": "02/19/2021 02:33:51",
      "content": "<p>Thank you😊😊😊</p>",
      "rawMarkdown": "Thank you😊😊😊",
      "votes": null
    },
    {
      "id": "1209712",
      "postDate": "02/19/2021 02:43:49",
      "content": "<p>Thanks!</p>\n<blockquote>\n  <p>In Image classification models, can models detect multi label?</p>\n</blockquote>\n<p>No, it cannot. To be precise, image classification model only takes the region around where a specific species exist. We created dicts that tell the frequency range of species, which was calculated based on train data. When we crop image patch, we only take the region inside that frequency range. Of course, the frequency range is different between species, so the result of cropping have different shape each other, so we resize that to a fixed shape. <br>\nWhen on inference time, we crop 24 patches (each of them corresponds to each species) from one time window, and put that 24 patches into CNN, so it takes 24x longer than usual CNN forward path.</p>",
      "rawMarkdown": "Thanks!\n> In Image classification models, can models detect multi label?\n\nNo, it cannot. To be precise, image classification model only takes the region around where a specific species exist. We created dicts that tell the frequency range of species, which was calculated based on train data. When we crop image patch, we only take the region inside that frequency range. Of course, the frequency range is different between species, so the result of cropping have different shape each other, so we resize that to a fixed shape. \nWhen on inference time, we crop 24 patches (each of them corresponds to each species) from one time window, and put that 24 patches into CNN, so it takes 24x longer than usual CNN forward path.",
      "votes": null
    },
    {
      "id": "1209716",
      "postDate": "02/19/2021 02:49:32",
      "content": "<p>Thank you! </p>\n<blockquote>\n  <p>I am glad you mentioned my Cornell model :)</p>\n</blockquote>\n<p>Oh, I meant it turned out that our image classification model was similar to your model in this competition (I found that after we read you solution.), but yes we were also helped a lot from your solution in Cornell's competition. <br>\nI saw many top performing teams used loss masking technique, which was also used in your Cornell's solution. It's not sure whether they came up with that idea alone or inspired from yours, but I'm quite sure your solution had big impact on this competition.</p>",
      "rawMarkdown": "Thank you! \n> I am glad you mentioned my Cornell model :)\n\nOh, I meant it turned out that our image classification model was similar to your model in this competition (I found that after we read you solution.), but yes we were also helped a lot from your solution in Cornell's competition. \nI saw many top performing teams used loss masking technique, which was also used in your Cornell's solution. It's not sure whether they came up with that idea alone or inspired from yours, but I'm quite sure your solution had big impact on this competition.",
      "votes": null
    },
    {
      "id": "1209729",
      "postDate": "02/19/2021 02:53:18",
      "content": "<p>Not unfortunately.<br>\nAt first I thought it is worth trying because I thought the key of this competition is Few-shot learning, but after a while I found this competition is more like a PU learning competition, so I changed the priority of COLA or other audio representation learning methods.</p>",
      "rawMarkdown": "Not unfortunately.\nAt first I thought it is worth trying because I thought the key of this competition is Few-shot learning, but after a while I found this competition is more like a PU learning competition, so I changed the priority of COLA or other audio representation learning methods.",
      "votes": null
    },
    {
      "id": "1209806",
      "postDate": "02/19/2021 03:48:12",
      "content": "<p>Thank you for your reply. If that is so, how do you detect multi label( secodary label)?<br>\nIs it SED? If it is SED, It's a amazing ensemble result!</p>\n<p>I also tried Image classification with ArcFace. But score was too low.<br>\nBut mixing SED and ArcFace (simple prediction average ) improved a littel.<br>\nThis might be the same as your team.</p>",
      "rawMarkdown": "Thank you for your reply. If that is so, how do you detect multi label( secodary label)?\nIs it SED? If it is SED, It's a amazing ensemble result!\n\nI also tried Image classification with ArcFace. But score was too low.\nBut mixing SED and ArcFace (simple prediction average ) improved a littel.\nThis might be the same as your team.",
      "votes": null
    },
    {
      "id": "1211126",
      "postDate": "02/20/2021 01:15:17",
      "content": "<p>That is addressed by this</p>\n<blockquote>\n  <p>When on inference time, we crop 24 patches (each of them corresponds to each species) from one time window, and put that 24 patches into CNN, so it takes 24x longer than usual CNN forward path.</p>\n</blockquote>",
      "rawMarkdown": "That is addressed by this\n> When on inference time, we crop 24 patches (each of them corresponds to each species) from one time window, and put that 24 patches into CNN, so it takes 24x longer than usual CNN forward path.",
      "votes": null
    },
    {
      "id": "1211583",
      "postDate": "02/20/2021 10:40:00",
      "content": "<blockquote>\n  <p>. It's not sure whether they came up with that idea alone or inspired from yours, but I'm quite sure your solution had big impact on this competition.</p>\n</blockquote>\n<p>I am not sure it was influential given it was almost never cited in writeups.  I was mentioned only a couple of times unless mistaken, and you are the only one to mention explicitly my Cornell solution.</p>",
      "rawMarkdown": "> . It's not sure whether they came up with that idea alone or inspired from yours, but I'm quite sure your solution had big impact on this competition.\n\nI am not sure it was influential given it was almost never cited in writeups.  I was mentioned only a couple of times unless mistaken, and you are the only one to mention explicitly my Cornell solution.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1208562,
      "author_name": "duykhanh99",
      "author_url": "",
      "post_date": "02/18/2021 10:41:36",
      "content": "<p>Congrats on strong finish <a href=\"https://www.kaggle.com/hidehisaarai1213\" target=\"_blank\">@hidehisaarai1213</a> and team and thanks for sharing details solution </p>",
      "votes": null,
      "replies": [
        {
          "id": 1209689,
          "author_name": "hidehisaarai1213",
          "author_url": "",
          "post_date": "02/19/2021 02:30:25",
          "content": "<p>Thank you! You've also finished in good place!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1208635,
      "author_name": "theoviel",
      "author_url": "",
      "post_date": "02/18/2021 11:45:20",
      "content": "<p>Congratz ! Good to see you at the top of another audio competition :) </p>",
      "votes": null,
      "replies": [
        {
          "id": 1209693,
          "author_name": "hidehisaarai1213",
          "author_url": "",
          "post_date": "02/19/2021 02:33:26",
          "content": "<p>Same to you! Also good to see many competitors of Birdcall comes to this competition. At first it seemed similar to birdcall but it turned out it wasn't - both are interesting in different ways</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1208655,
      "author_name": "riadalmadani",
      "author_url": "",
      "post_date": "02/18/2021 12:06:57",
      "content": "<p>Congratulations🎉🎉</p>",
      "votes": null,
      "replies": [
        {
          "id": 1209696,
          "author_name": "hidehisaarai1213",
          "author_url": "",
          "post_date": "02/19/2021 02:33:51",
          "content": "<p>Thank you😊😊😊</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1208690,
      "author_name": "shinmurashinmura",
      "author_url": "",
      "post_date": "02/18/2021 12:44:04",
      "content": "<p>Your solution (Mixing Image classification models and SED) is interesting!</p>\n<p>I have a question. <br>\nIn Image classification models, can models detect multi label?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1209712,
          "author_name": "hidehisaarai1213",
          "author_url": "",
          "post_date": "02/19/2021 02:43:49",
          "content": "<p>Thanks!</p>\n<blockquote>\n  <p>In Image classification models, can models detect multi label?</p>\n</blockquote>\n<p>No, it cannot. To be precise, image classification model only takes the region around where a specific species exist. We created dicts that tell the frequency range of species, which was calculated based on train data. When we crop image patch, we only take the region inside that frequency range. Of course, the frequency range is different between species, so the result of cropping have different shape each other, so we resize that to a fixed shape. <br>\nWhen on inference time, we crop 24 patches (each of them corresponds to each species) from one time window, and put that 24 patches into CNN, so it takes 24x longer than usual CNN forward path.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1209806,
          "author_name": "shinmurashinmura",
          "author_url": "",
          "post_date": "02/19/2021 03:48:12",
          "content": "<p>Thank you for your reply. If that is so, how do you detect multi label( secodary label)?<br>\nIs it SED? If it is SED, It's a amazing ensemble result!</p>\n<p>I also tried Image classification with ArcFace. But score was too low.<br>\nBut mixing SED and ArcFace (simple prediction average ) improved a littel.<br>\nThis might be the same as your team.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1211126,
          "author_name": "hidehisaarai1213",
          "author_url": "",
          "post_date": "02/20/2021 01:15:17",
          "content": "<p>That is addressed by this</p>\n<blockquote>\n  <p>When on inference time, we crop 24 patches (each of them corresponds to each species) from one time window, and put that 24 patches into CNN, so it takes 24x longer than usual CNN forward path.</p>\n</blockquote>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1209458,
      "author_name": "cpmpml",
      "author_url": "",
      "post_date": "02/18/2021 22:59:16",
      "content": "<p>Wow, you pushed SED quite far!  I am glad you mentioned my Cornell model :)</p>\n<p>Congrats on the strong ending,  I was afraid you'd pass me.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1209716,
          "author_name": "hidehisaarai1213",
          "author_url": "",
          "post_date": "02/19/2021 02:49:32",
          "content": "<p>Thank you! </p>\n<blockquote>\n  <p>I am glad you mentioned my Cornell model :)</p>\n</blockquote>\n<p>Oh, I meant it turned out that our image classification model was similar to your model in this competition (I found that after we read you solution.), but yes we were also helped a lot from your solution in Cornell's competition. <br>\nI saw many top performing teams used loss masking technique, which was also used in your Cornell's solution. It's not sure whether they came up with that idea alone or inspired from yours, but I'm quite sure your solution had big impact on this competition.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1211583,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "02/20/2021 10:40:00",
          "content": "<blockquote>\n  <p>. It's not sure whether they came up with that idea alone or inspired from yours, but I'm quite sure your solution had big impact on this competition.</p>\n</blockquote>\n<p>I am not sure it was influential given it was almost never cited in writeups.  I was mentioned only a couple of times unless mistaken, and you are the only one to mention explicitly my Cornell solution.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1209645,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "02/19/2021 01:49:16",
      "content": "<p>is there any experimental results on COLA from \"Contrastive learning of general purpose audio representations\"?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1209729,
          "author_name": "hidehisaarai1213",
          "author_url": "",
          "post_date": "02/19/2021 02:53:18",
          "content": "<p>Not unfortunately.<br>\nAt first I thought it is worth trying because I thought the key of this competition is Few-shot learning, but after a while I found this competition is more like a PU learning competition, so I changed the priority of COLA or other audio representation learning methods.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1208512": "Congratulations to all the winners and gold getters, I guess those teams that broke the 0.950 wall have found the essence of this competition, which we couldn't.  \n<br/>\nFirstly, thanks to the host for holding quite an interesting competition. Partly labeled classification is a challenging task, which made this competition more interesting than a simple bioacoustics audio tagging competition.\n<br/>\n\nOur solution is a ranking average of image classification models and SED models. [Y.Nakama](https://www.kaggle.com/yasufuminakama), [kaerururu](https://www.kaggle.com/kaerunantoka), and I worked a lot on SED models but couldn't break 0.90 until we merge with [Taku Hiraiwa](https://www.kaggle.com/thiraiwa). His model was based on image classification similar to [@cpmpml's model](https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/220304). We found that quite good , so we focused on improving image classification models in the remained days.\n## Image classification models\nIt was [Taku Hiraiwa](https://www.kaggle.com/thiraiwa)'s idea to only use the *annotated* part of the train data. To do so, we crop the image patches from log-melspectrogram of train data based on the `t_min`, `t_max`, `f_min`, `f_max` information of train\\_tp.csv and train\\_fp.csv, and resized the patch to fixed shape (say, 320 x 320 or so). The cropping is performed on the fly through training, so the part we crop out is randomized along with time axis.\nWith these image patches we trained EfficientNet models, and monitored F1 score with threshold 0.5. \nHere's the other details of image classification models.\n\n* image size: varies from (244, 244) to (456, 456) between models\n* backbone: EfficientNetB0 - B5 (used [`timm`](https://rwightman.github.io/pytorch-image-models/) and used `tf_efficientnet_b<0-5>_ns` weights).\n* augmentation: GaussianNoise, Gain, PitchShift of [audiomentations](https://github.com/iver56/audiomentations) on raw waveform. Also HorizontalFlip also had positive impact on LB slightly, so we used (but don't know why it worked).\n* AdamW optimizer with linear warmup scheduler\n* BCEFocalLoss\n<br/>\nIn the end, we trained a stacking model that takes the output of models below which achieve public 0.942:\n\n1. `tf_efficientnet_b0_ns` image size 244\n2. `tf_efficientnet_b0_ns` image size 320\n3. `tf_efficientnet_b0_ns` image size 456\n4. `tf_efficientnet_b1_ns` image size 456\n5. `tf_efficientnet_b2_ns` image size 456\n6. `tf_efficientnet_b3_ns` image size 456\n7. `tf_efficientnet_b4_ns` image size 456\n8. `tf_efficientnet_b5_ns` image size 456\n\n## SED models\nAll of our SED models use the head architecture introduced in [PANNs repository](https://github.com/qiuqiangkong/audioset_tagging_cnn). The CNN encoder is either EfficientNet or ResNeSt, and they are trained with weak/strong supervision. We tried a lot of things on this, but couldn't find factors that consistently work well - which means the LB scores varied quite randomly w.r.t CV score.\nOur best SED model is rank average of 11 models(public: 0.901, privte: 0.911) below - each of them differs slightly so we describe the difference briefly.\n\n### 2 x kaerururu's model (public: 0.882, 0.873)\n\n* Based on public starter SED (EffnB0) notebook (https://www.kaggle.com/gopidurgaprasad/rfcx-sed-model-stater)\n* 3ch input\n* 10sec clip\n* waveform mixup\n* some augmentation (audiomentations)\n* pseudo-labeled datasets (add labels on tp data)\n* trained with tp and fp dataset (1st training)\n* trained with pseudo-labeled tp (2nd training)\n* tta=10\n\n### 5 x arai's model (public: 0.879,0.880, 0.868, 0.874, 0.870)\n\n* Based on Birdcall's challenge 6th place (https://github.com/koukyo1994/kaggle-birdcall-6th-place)\n* ResNeSt50 encoder or EfficientNetB3 encoder\n* AddPinkNoiseSNR / VolumeControl / PitchShift from [My Notebook](https://www.kaggle.com/hidehisaarai1213/rfcx-audio-data-augmentation-japanese-english)\n* tp only\n### 4 x Y.Nakama's model (public: 0.871, 0.863, 0.866, 0.870)\n\n* Based on Birdcall's challenge 6th place (https://github.com/koukyo1994/kaggle-birdcall-6th-place)\n* ResNeSt50 encoder or ResNet200D encoder\n* mixup & some augmentations\n* 2nd stage training\n  * 1st stage: weighted loss for framewise_logit & logit\n  * 2nd stage: loss for logit\n* tp only",
    "1208562": "Congrats on strong finish @hidehisaarai1213 and team and thanks for sharing details solution",
    "1208635": "Congratz ! Good to see you at the top of another audio competition :)",
    "1208655": "Congratulations🎉🎉",
    "1208690": "Your solution (Mixing Image classification models and SED) is interesting!\n\nI have a question. \nIn Image classification models, can models detect multi label?",
    "1209458": "Wow, you pushed SED quite far!  I am glad you mentioned my Cornell model :)\n\nCongrats on the strong ending,  I was afraid you'd pass me.",
    "1209645": "is there any experimental results on COLA from \"Contrastive learning of general purpose audio representations\"?",
    "1209689": "Thank you! You've also finished in good place!",
    "1209693": "Same to you! Also good to see many competitors of Birdcall comes to this competition. At first it seemed similar to birdcall but it turned out it wasn't - both are interesting in different ways",
    "1209696": "Thank you😊😊😊",
    "1209712": "Thanks!\n> In Image classification models, can models detect multi label?\n\nNo, it cannot. To be precise, image classification model only takes the region around where a specific species exist. We created dicts that tell the frequency range of species, which was calculated based on train data. When we crop image patch, we only take the region inside that frequency range. Of course, the frequency range is different between species, so the result of cropping have different shape each other, so we resize that to a fixed shape. \nWhen on inference time, we crop 24 patches (each of them corresponds to each species) from one time window, and put that 24 patches into CNN, so it takes 24x longer than usual CNN forward path.",
    "1209716": "Thank you! \n> I am glad you mentioned my Cornell model :)\n\nOh, I meant it turned out that our image classification model was similar to your model in this competition (I found that after we read you solution.), but yes we were also helped a lot from your solution in Cornell's competition. \nI saw many top performing teams used loss masking technique, which was also used in your Cornell's solution. It's not sure whether they came up with that idea alone or inspired from yours, but I'm quite sure your solution had big impact on this competition.",
    "1209729": "Not unfortunately.\nAt first I thought it is worth trying because I thought the key of this competition is Few-shot learning, but after a while I found this competition is more like a PU learning competition, so I changed the priority of COLA or other audio representation learning methods.",
    "1209806": "Thank you for your reply. If that is so, how do you detect multi label( secodary label)?\nIs it SED? If it is SED, It's a amazing ensemble result!\n\nI also tried Image classification with ArcFace. But score was too low.\nBut mixing SED and ArcFace (simple prediction average ) improved a littel.\nThis might be the same as your team.",
    "1211126": "That is addressed by this\n> When on inference time, we crop 24 patches (each of them corresponds to each species) from one time window, and put that 24 patches into CNN, so it takes 24x longer than usual CNN forward path.",
    "1211583": "> . It's not sure whether they came up with that idea alone or inspired from yours, but I'm quite sure your solution had big impact on this competition.\n\nI am not sure it was influential given it was almost never cited in writeups.  I was mentioned only a couple of times unless mistaken, and you are the only one to mention explicitly my Cornell solution."
  },
  "source": "meta"
}