{
  "id": 183300,
  "title": "5th place solution",
  "url": "/competitions/birdsong-recognition/discussion/183300",
  "author_name": "Oleg Yaroshevskiy",
  "post_date": "2020-09-16T06:53:15.200000",
  "votes": 74,
  "comment_count": 19,
  "views": 0,
  "content": "<p>A great journey, thanks for organizing it! Now every time I walk around I hear much more birds :) <br>\nI've started my machine learning journey with speech processing few years ago and that was a joy to play again with it. One more time to thank and mention <a href=\"https://www.kaggle.com/hidehisaarai1213\" target=\"_blank\">@hidehisaarai1213</a> who led us through the darkness here.</p>\n<p>From the very scratch I understood that the problem can be decomposed into two main problems:</p>\n<ul>\n<li>a domain shift (which is pretty tough us we don't have target test set)</li>\n<li>a clipwise to framewise classification transit (which also introduced us additional \"nocall\" class with huge influence on results, and labels noise)</li>\n</ul>\n<p>Tackling those two needed a validation/test set as close to target distribution as possible in terms of snr and nocall distributions. </p>\n<p><strong>Validation</strong></p>\n<p>Believe me or not but with 6 records from <a href=\"https://www.kaggle.com/c/birdsong-recognition/discussion/158877#911336\" target=\"_blank\">here</a> and 2 example audios I was able to get some correlation with LB even though it contained only 500-800 clips. But more important I was able to control % of \"nocall\"s and pick the right threshold based on them. Last two days I splitted those clips into 3 test sets with 50%, 60%, 80% \"nocalls\" to see possible scenarios on private leaderboard (thanks organizers private is almost equal to public) and to secure the score with my second submission. </p>\n<p>As a target metric I've used F1 by <code>average=\"samples\"</code> pointed by <a href=\"https://www.kaggle.com/cpmp\" target=\"_blank\">@cpmp</a>. By the end I've also monitored validation (default CV split) primary+secondary F1 score as my another decision-making metric. Primary F1 or mAP was not enough.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F805012%2Fa290c42c9808e69c800abb79f2889f22%2FScreenshot%20from%202020-09-16%2009-53-50.png?generation=1600239312223536&amp;alt=media\" alt=\"\"></p>\n<p>Adding separate nocall class didn't work for me for whatever reason.<br>\nAnd yeah, I trusted leaderboard too.</p>\n<p><strong>Domain shift</strong><br>\nAssuming we don't have the transit problem, this might be solved simply with augmentations. I've spent too much time on this and right now I understand that was not productive. My final models contain different augmentations configurations: </p>\n<ul>\n<li>backgrounds (all from external data thread)</li>\n<li>pink and brown noise</li>\n<li>pitch shift</li>\n<li>low pass filtering</li>\n<li>spec augments (time and frequency masking)</li>\n</ul>\n<p>As organizers informed that we can train on two test examples, I tried to collect batch norm statistics from those as a domain adaptation technique but it didn't work great. </p>\n<p><strong>Clipwise &gt; framewise transit p.1</strong></p>\n<p>That's my favorite part!<br>\nSo what we have here - every time we crop 5 seconds clip we have a chance to crop a nocall clip. So labels become really noisy. Even more it's hard to crop secondary classes. Easy way to tackle this? Label smoothing 0.2. I don't remember when I became a fan of label smoothing but it works well with noisy labels on practice. But that's not serious.</p>\n<p>People here tackled this task having an energy based cuts. And it really worked. So I've tried both approaches: soft and hard. By soft I mean random sampling based on energy, by hard - removing everything below normalized energy threshold.</p>\n<p>But what if we can use our model to extract these labels? Having avg/map pooling head gives a chance to get a free segmentation:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F805012%2Fb9938f7758cd413a47006ac98c5f62e0%2Fphoto5377374165436313567.jpg?generation=1600241278088906&amp;alt=media\" alt=\"\"><br>\nwhich you can use as soft labels or hard binarized labels. These methods showed really cool performance for primary labels. For secondary I kept label smoothing :)</p>\n<p><strong>Clipwise &gt; framewise transit p.2 (The most important part)</strong></p>\n<p>I found SED models comparably weak trained on 5 sec clips but as I trained them on longer clips (10-30 seconds) I noticed that due to labels noise reduction they show much better training loss/mAP. Moreover, it was possible to run inference on 5 sec clips with nice performance. So I took B4 EffNet, added the same Attention decision making head and … failed!</p>\n<p>I've spent a lot on this one. So why did Cnn14_DecisionLevelAtt trained on 30 sec clips work well on 5 sec clips? Playing with EffNet I found that problem was in receptive field - it's too big (630 compared to 200 or so). With wide receptive field attention head was not able to build a meaningful framewise feature maps. As I've changed EffNet kernels from 5x5 to 3x3 or reduced number of blocks - it solved problem but price was too big - pretrained weights. </p>\n<p>So what we have: </p>\n<ul>\n<li>longer input</li>\n<li>clipping based on some weakly labeled probabilities</li>\n<li>label smoothing for secondary</li>\n</ul>\n<p>Having these I've decided to simply run 5sec classification inference with no sophisticated postprocessing that might fail on private LB. As it always happens, I've found this too late so trained only two PANNs models (resnet38 and cnn14 mentioned above) on 12 and 15 second clips (128 mels) with mixup and didn't have time to configure them properly.</p>\n<p><strong>Finally</strong></p>\n<p>I've optimized inference to run all my models (cnn14, resnet38 and few effnets) in 20 mins with <code>kaiser_fast</code> resampling and around one hour with <code>kaiser_best</code>. I've picked thresholds based on my test set (0.3) and some skepticism about it (0.4). I found I have 670+ private score submissions from two weeks ago, which I can't explain. Probably, Kaggle leaderboard is another stochastic process.</p>\n<p>This has been a long journey in which I not only solved many competitions and learned just hundreds of tools and tricks, but also met wonderful minds from all over the world and found job of my dreams. I wish to meet you in person at Kaggle Days once covid is over. See you!</p>",
  "messages": [
    {
      "id": 1012563,
      "postDate": "2020-09-16T06:53:15.200Z",
      "content": "<p>A great journey, thanks for organizing it! Now every time I walk around I hear much more birds :) <br>\nI've started my machine learning journey with speech processing few years ago and that was a joy to play again with it. One more time to thank and mention <a href=\"https://www.kaggle.com/hidehisaarai1213\" target=\"_blank\">@hidehisaarai1213</a> who led us through the darkness here.</p>\n<p>From the very scratch I understood that the problem can be decomposed into two main problems:</p>\n<ul>\n<li>a domain shift (which is pretty tough us we don't have target test set)</li>\n<li>a clipwise to framewise classification transit (which also introduced us additional \"nocall\" class with huge influence on results, and labels noise)</li>\n</ul>\n<p>Tackling those two needed a validation/test set as close to target distribution as possible in terms of snr and nocall distributions. </p>\n<p><strong>Validation</strong></p>\n<p>Believe me or not but with 6 records from <a href=\"https://www.kaggle.com/c/birdsong-recognition/discussion/158877#911336\" target=\"_blank\">here</a> and 2 example audios I was able to get some correlation with LB even though it contained only 500-800 clips. But more important I was able to control % of \"nocall\"s and pick the right threshold based on them. Last two days I splitted those clips into 3 test sets with 50%, 60%, 80% \"nocalls\" to see possible scenarios on private leaderboard (thanks organizers private is almost equal to public) and to secure the score with my second submission. </p>\n<p>As a target metric I've used F1 by <code>average=\"samples\"</code> pointed by <a href=\"https://www.kaggle.com/cpmp\" target=\"_blank\">@cpmp</a>. By the end I've also monitored validation (default CV split) primary+secondary F1 score as my another decision-making metric. Primary F1 or mAP was not enough.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F805012%2Fa290c42c9808e69c800abb79f2889f22%2FScreenshot%20from%202020-09-16%2009-53-50.png?generation=1600239312223536&amp;alt=media\" alt=\"\"></p>\n<p>Adding separate nocall class didn't work for me for whatever reason.<br>\nAnd yeah, I trusted leaderboard too.</p>\n<p><strong>Domain shift</strong><br>\nAssuming we don't have the transit problem, this might be solved simply with augmentations. I've spent too much time on this and right now I understand that was not productive. My final models contain different augmentations configurations: </p>\n<ul>\n<li>backgrounds (all from external data thread)</li>\n<li>pink and brown noise</li>\n<li>pitch shift</li>\n<li>low pass filtering</li>\n<li>spec augments (time and frequency masking)</li>\n</ul>\n<p>As organizers informed that we can train on two test examples, I tried to collect batch norm statistics from those as a domain adaptation technique but it didn't work great. </p>\n<p><strong>Clipwise &gt; framewise transit p.1</strong></p>\n<p>That's my favorite part!<br>\nSo what we have here - every time we crop 5 seconds clip we have a chance to crop a nocall clip. So labels become really noisy. Even more it's hard to crop secondary classes. Easy way to tackle this? Label smoothing 0.2. I don't remember when I became a fan of label smoothing but it works well with noisy labels on practice. But that's not serious.</p>\n<p>People here tackled this task having an energy based cuts. And it really worked. So I've tried both approaches: soft and hard. By soft I mean random sampling based on energy, by hard - removing everything below normalized energy threshold.</p>\n<p>But what if we can use our model to extract these labels? Having avg/map pooling head gives a chance to get a free segmentation:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F805012%2Fb9938f7758cd413a47006ac98c5f62e0%2Fphoto5377374165436313567.jpg?generation=1600241278088906&amp;alt=media\" alt=\"\"><br>\nwhich you can use as soft labels or hard binarized labels. These methods showed really cool performance for primary labels. For secondary I kept label smoothing :)</p>\n<p><strong>Clipwise &gt; framewise transit p.2 (The most important part)</strong></p>\n<p>I found SED models comparably weak trained on 5 sec clips but as I trained them on longer clips (10-30 seconds) I noticed that due to labels noise reduction they show much better training loss/mAP. Moreover, it was possible to run inference on 5 sec clips with nice performance. So I took B4 EffNet, added the same Attention decision making head and … failed!</p>\n<p>I've spent a lot on this one. So why did Cnn14_DecisionLevelAtt trained on 30 sec clips work well on 5 sec clips? Playing with EffNet I found that problem was in receptive field - it's too big (630 compared to 200 or so). With wide receptive field attention head was not able to build a meaningful framewise feature maps. As I've changed EffNet kernels from 5x5 to 3x3 or reduced number of blocks - it solved problem but price was too big - pretrained weights. </p>\n<p>So what we have: </p>\n<ul>\n<li>longer input</li>\n<li>clipping based on some weakly labeled probabilities</li>\n<li>label smoothing for secondary</li>\n</ul>\n<p>Having these I've decided to simply run 5sec classification inference with no sophisticated postprocessing that might fail on private LB. As it always happens, I've found this too late so trained only two PANNs models (resnet38 and cnn14 mentioned above) on 12 and 15 second clips (128 mels) with mixup and didn't have time to configure them properly.</p>\n<p><strong>Finally</strong></p>\n<p>I've optimized inference to run all my models (cnn14, resnet38 and few effnets) in 20 mins with <code>kaiser_fast</code> resampling and around one hour with <code>kaiser_best</code>. I've picked thresholds based on my test set (0.3) and some skepticism about it (0.4). I found I have 670+ private score submissions from two weeks ago, which I can't explain. Probably, Kaggle leaderboard is another stochastic process.</p>\n<p>This has been a long journey in which I not only solved many competitions and learned just hundreds of tools and tricks, but also met wonderful minds from all over the world and found job of my dreams. I wish to meet you in person at Kaggle Days once covid is over. See you!</p>",
      "rawMarkdown": "A great journey, thanks for organizing it! Now every time I walk around I hear much more birds :) \nI've started my machine learning journey with speech processing few years ago and that was a joy to play again with it. One more time to thank and mention @hidehisaarai1213 who led us through the darkness here.\n\nFrom the very scratch I understood that the problem can be decomposed into two main problems:\n\n- a domain shift (which is pretty tough us we don't have target test set)\n- a clipwise to framewise classification transit (which also introduced us additional \"nocall\" class with huge influence on results, and labels noise)\n\nTackling those two needed a validation/test set as close to target distribution as possible in terms of snr and nocall distributions. \n\n**Validation**\n\nBelieve me or not but with 6 records from [here](https://www.kaggle.com/c/birdsong-recognition/discussion/158877#911336) and 2 example audios I was able to get some correlation with LB even though it contained only 500-800 clips. But more important I was able to control % of \"nocall\"s and pick the right threshold based on them. Last two days I splitted those clips into 3 test sets with 50%, 60%, 80% \"nocalls\" to see possible scenarios on private leaderboard (thanks organizers private is almost equal to public) and to secure the score with my second submission. \n\nAs a target metric I've used F1 by `average=\"samples\"` pointed by @cpmp. By the end I've also monitored validation (default CV split) primary+secondary F1 score as my another decision-making metric. Primary F1 or mAP was not enough.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F805012%2Fa290c42c9808e69c800abb79f2889f22%2FScreenshot%20from%202020-09-16%2009-53-50.png?generation=1600239312223536&alt=media)\n\nAdding separate nocall class didn't work for me for whatever reason.\nAnd yeah, I trusted leaderboard too.\n\n**Domain shift**\nAssuming we don't have the transit problem, this might be solved simply with augmentations. I've spent too much time on this and right now I understand that was not productive. My final models contain different augmentations configurations: \n- backgrounds (all from external data thread)\n- pink and brown noise\n- pitch shift\n- low pass filtering\n- spec augments (time and frequency masking)\n\nAs organizers informed that we can train on two test examples, I tried to collect batch norm statistics from those as a domain adaptation technique but it didn't work great. \n\n**Clipwise > framewise transit p.1**\n\nThat's my favorite part!\nSo what we have here - every time we crop 5 seconds clip we have a chance to crop a nocall clip. So labels become really noisy. Even more it's hard to crop secondary classes. Easy way to tackle this? Label smoothing 0.2. I don't remember when I became a fan of label smoothing but it works well with noisy labels on practice. But that's not serious.\n\nPeople here tackled this task having an energy based cuts. And it really worked. So I've tried both approaches: soft and hard. By soft I mean random sampling based on energy, by hard - removing everything below normalized energy threshold.\n\nBut what if we can use our model to extract these labels? Having avg/map pooling head gives a chance to get a free segmentation:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F805012%2Fb9938f7758cd413a47006ac98c5f62e0%2Fphoto5377374165436313567.jpg?generation=1600241278088906&alt=media)\nwhich you can use as soft labels or hard binarized labels. These methods showed really cool performance for primary labels. For secondary I kept label smoothing :)\n\n**Clipwise > framewise transit p.2 (The most important part)**\n\nI found SED models comparably weak trained on 5 sec clips but as I trained them on longer clips (10-30 seconds) I noticed that due to labels noise reduction they show much better training loss/mAP. Moreover, it was possible to run inference on 5 sec clips with nice performance. So I took B4 EffNet, added the same Attention decision making head and ... failed!\n\nI've spent a lot on this one. So why did Cnn14_DecisionLevelAtt trained on 30 sec clips work well on 5 sec clips? Playing with EffNet I found that problem was in receptive field - it's too big (630 compared to 200 or so). With wide receptive field attention head was not able to build a meaningful framewise feature maps. As I've changed EffNet kernels from 5x5 to 3x3 or reduced number of blocks - it solved problem but price was too big - pretrained weights. \n\nSo what we have: \n- longer input\n- clipping based on some weakly labeled probabilities\n- label smoothing for secondary\n\nHaving these I've decided to simply run 5sec classification inference with no sophisticated postprocessing that might fail on private LB. As it always happens, I've found this too late so trained only two PANNs models (resnet38 and cnn14 mentioned above) on 12 and 15 second clips (128 mels) with mixup and didn't have time to configure them properly.\n\n**Finally**\n\nI've optimized inference to run all my models (cnn14, resnet38 and few effnets) in 20 mins with `kaiser_fast` resampling and around one hour with `kaiser_best`. I've picked thresholds based on my test set (0.3) and some skepticism about it (0.4). I found I have 670+ private score submissions from two weeks ago, which I can't explain. Probably, Kaggle leaderboard is another stochastic process.\n\nThis has been a long journey in which I not only solved many competitions and learned just hundreds of tools and tricks, but also met wonderful minds from all over the world and found job of my dreams. I wish to meet you in person at Kaggle Days once covid is over. See you!",
      "votes": 74
    },
    {
      "id": 1012616,
      "postDate": "2020-09-16T07:25:29.870Z",
      "content": "<p>Congrats on solo gold medal and becoming GM <a href=\"https://www.kaggle.com/yaroshevskiy\" target=\"_blank\">@yaroshevskiy</a> </p>",
      "rawMarkdown": "Congrats on solo gold medal and becoming GM @yaroshevskiy ",
      "votes": 10,
      "replies": [
        {
          "id": 1012685,
          "postDate": "2020-09-16T08:14:06.787Z",
          "content": "<p>Thank you!</p>",
          "rawMarkdown": "Thank you!",
          "votes": 3
        }
      ]
    },
    {
      "id": 1012743,
      "postDate": "2020-09-16T09:10:17.563Z",
      "content": "<p>Congrats on the solo gold result!  I am glad you investigated why attention head on top of effnet wasn't working, I saw the same but didn't know why.  How did you ensemble your models?  We failed he few times we tried.</p>\n<p>It looks like you're the only top finisher to have set a CV that mimics test.  Another good move from you IMHO.</p>",
      "rawMarkdown": "Congrats on the solo gold result!  I am glad you investigated why attention head on top of effnet wasn't working, I saw the same but didn't know why.  How did you ensemble your models?  We failed he few times we tried.\n\nIt looks like you're the only top finisher to have set a CV that mimics test.  Another good move from you IMHO.",
      "votes": 6,
      "replies": [
        {
          "id": 1012931,
          "postDate": "2020-09-16T11:56:43.293Z",
          "content": "<p>Thanks!<br>\nI've tried mean averaging, voting and rank renormalization (with <code>scipy.interpolate.interp1d(probs, ranks, kind=\"linear\")</code>) but didn't get any local improvement for all of thm. So I just submitted mean averaged and leaderboard showed a great boost. Selecting the threshold based on test examples worked too. If I had more submissions I'd try also voting as top-1 performer. </p>",
          "rawMarkdown": "Thanks!\nI've tried mean averaging, voting and rank renormalization (with `scipy.interpolate.interp1d(probs, ranks, kind=\"linear\")`) but didn't get any local improvement for all of thm. So I just submitted mean averaged and leaderboard showed a great boost. Selecting the threshold based on test examples worked too. If I had more submissions I'd try also voting as top-1 performer. ",
          "votes": 4
        }
      ]
    },
    {
      "id": 1012724,
      "postDate": "2020-09-16T08:49:20.257Z",
      "content": "<p>Congrats on the solo gold Oleg! 🎉</p>",
      "rawMarkdown": "Congrats on the solo gold Oleg! 🎉",
      "votes": 4
    },
    {
      "id": 1025427,
      "postDate": "2020-09-24T14:44:05.037Z",
      "content": "<p>Congrats on GM</p>",
      "rawMarkdown": "Congrats on GM",
      "votes": 1
    },
    {
      "id": 1016763,
      "postDate": "2020-09-19T08:09:47.757Z",
      "content": "<p>Your story is really inspired me. I am a beginner in speech recognition. </p>",
      "rawMarkdown": "Your story is really inspired me. I am a beginner in speech recognition. ",
      "votes": 1
    },
    {
      "id": 1015242,
      "postDate": "2020-09-18T04:19:22.250Z",
      "content": "<p>Congrats on solo gold and becoming GM !  And thanks for your good write up.<br>\nI have some questions:</p>\n<blockquote>\n  <p>Having avg/map pooling head gives a chance to get a free segmentation.</p>\n</blockquote>\n<p>How did you choose between max and avg, I was very torn on which one to use in this comp, they were good and bad at times.</p>\n<blockquote>\n  <p>I found that problem was in receptive field - it's too big.</p>\n</blockquote>\n<p>What's an amazing finding! Could you talk more about how did you find it? </p>",
      "rawMarkdown": "Congrats on solo gold and becoming GM !  And thanks for your good write up.\nI have some questions:\n\n> Having avg/map pooling head gives a chance to get a free segmentation.\n\nHow did you choose between max and avg, I was very torn on which one to use in this comp, they were good and bad at times.\n\n> I found that problem was in receptive field - it's too big.\n\nWhat's an amazing finding! Could you talk more about how did you find it? ",
      "votes": 1,
      "replies": [
        {
          "id": 1019221,
          "postDate": "2020-09-20T09:29:30.403Z",
          "content": "<p>Truly to say, as I remember max and avg+max poolings worked better. </p>\n<p>Yes, I tried to train on different input lengths - 5/10/15/30 seconds and monitored my test set. I found that PANNs CNN14 generalized well to 5 sec (!) inference but EffNet didn't. I've tried different configurations: padding types, kernel sizes, number of blocks repeats to find that the problem was in receptive field. </p>",
          "rawMarkdown": "Truly to say, as I remember max and avg+max poolings worked better. \n\nYes, I tried to train on different input lengths - 5/10/15/30 seconds and monitored my test set. I found that PANNs CNN14 generalized well to 5 sec (!) inference but EffNet didn't. I've tried different configurations: padding types, kernel sizes, number of blocks repeats to find that the problem was in receptive field. ",
          "votes": 2
        },
        {
          "id": 1019282,
          "postDate": "2020-09-20T10:27:29.043Z",
          "content": "<p>Thanks for answering.</p>",
          "rawMarkdown": "Thanks for answering."
        }
      ]
    },
    {
      "id": 1013465,
      "postDate": "2020-09-16T17:51:33.833Z",
      "content": "<p>Nice solution! Our sanity-check dataset was not far from yours! But we had some variance on LB correlation.</p>",
      "rawMarkdown": "Nice solution! Our sanity-check dataset was not far from yours! But we had some variance on LB correlation.",
      "votes": 1
    },
    {
      "id": 1013229,
      "postDate": "2020-09-16T15:19:33.800Z",
      "content": "<p>Congratulations! Keep posting so we (the newbies) will learn about new methods and tricks in competition.</p>",
      "rawMarkdown": "Congratulations! Keep posting so we (the newbies) will learn about new methods and tricks in competition.",
      "votes": 1
    },
    {
      "id": 1012735,
      "postDate": "2020-09-16T09:00:43.543Z",
      "content": "<p>Thank you for sharing your solution, congratulation👍</p>",
      "rawMarkdown": "Thank you for sharing your solution, congratulation👍",
      "votes": 1
    },
    {
      "id": 1018591,
      "postDate": "2020-09-19T18:58:15.840Z",
      "content": "<p>congratulation :)</p>",
      "rawMarkdown": "congratulation :)",
      "votes": 2
    },
    {
      "id": 1013065,
      "postDate": "2020-09-16T13:39:11.233Z",
      "content": "<p>Wow, that's great 😍<br>\nThanks a lot for sharing.<br>\nCongratulations for solo GOLD and becoming GM 😍</p>",
      "rawMarkdown": "Wow, that's great 😍\nThanks a lot for sharing.\nCongratulations for solo GOLD and becoming GM 😍",
      "votes": 2
    },
    {
      "id": 1012954,
      "postDate": "2020-09-16T12:18:43.503Z",
      "content": "<p>congratulation to become GrandMaster :) </p>",
      "rawMarkdown": " congratulation to become GrandMaster :) ",
      "votes": 2
    },
    {
      "id": 1012639,
      "postDate": "2020-09-16T07:44:01.343Z",
      "content": "<p>Congratulations and looking forward for more details !<br>\nHope you stay safe where you are (the island of the gods must be very blissful without the crowd of tourists)</p>",
      "rawMarkdown": "Congratulations and looking forward for more details !\nHope you stay safe where you are (the island of the gods must be very blissful without the crowd of tourists)",
      "votes": 2,
      "replies": [
        {
          "id": 1012684,
          "postDate": "2020-09-16T08:13:59.387Z",
          "content": "<p>Yeah, thanks Bali for having me during pandemic ^^</p>",
          "rawMarkdown": "Yeah, thanks Bali for having me during pandemic ^^",
          "votes": 4
        }
      ]
    },
    {
      "id": 1022572,
      "postDate": "2020-09-22T16:08:50.457Z",
      "rawMarkdown": "",
      "votes": 1,
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 1012616,
      "author_name": "KhanhVD",
      "author_url": "",
      "post_date": "2020-09-16T07:25:29.870000",
      "content": "<p>Congrats on solo gold medal and becoming GM <a href=\"https://www.kaggle.com/yaroshevskiy\" target=\"_blank\">@yaroshevskiy</a> </p>",
      "votes": 10,
      "replies": [
        {
          "id": 1012685,
          "author_name": "Oleg Yaroshevskiy",
          "author_url": "",
          "post_date": "2020-09-16T08:14:06.787000",
          "content": "<p>Thank you!</p>",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 1012743,
      "author_name": "CPMP",
      "author_url": "",
      "post_date": "2020-09-16T09:10:17.563000",
      "content": "<p>Congrats on the solo gold result!  I am glad you investigated why attention head on top of effnet wasn't working, I saw the same but didn't know why.  How did you ensemble your models?  We failed he few times we tried.</p>\n<p>It looks like you're the only top finisher to have set a CV that mimics test.  Another good move from you IMHO.</p>",
      "votes": 6,
      "replies": [
        {
          "id": 1012931,
          "author_name": "Oleg Yaroshevskiy",
          "author_url": "",
          "post_date": "2020-09-16T11:56:43.293000",
          "content": "<p>Thanks!<br>\nI've tried mean averaging, voting and rank renormalization (with <code>scipy.interpolate.interp1d(probs, ranks, kind=\"linear\")</code>) but didn't get any local improvement for all of thm. So I just submitted mean averaged and leaderboard showed a great boost. Selecting the threshold based on test examples worked too. If I had more submissions I'd try also voting as top-1 performer. </p>",
          "votes": 4,
          "replies": []
        }
      ]
    },
    {
      "id": 1012724,
      "author_name": "Kostya Kravchenko",
      "author_url": "",
      "post_date": "2020-09-16T08:49:20.257000",
      "content": "<p>Congrats on the solo gold Oleg! 🎉</p>",
      "votes": 4,
      "replies": []
    },
    {
      "id": 1025427,
      "author_name": "Bo Peng",
      "author_url": "",
      "post_date": "2020-09-24T14:44:05.037000",
      "content": "<p>Congrats on GM</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1016763,
      "author_name": "Md. Abdullah Al Mamun",
      "author_url": "",
      "post_date": "2020-09-19T08:09:47.757000",
      "content": "<p>Your story is really inspired me. I am a beginner in speech recognition. </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1015242,
      "author_name": "Yu Kang",
      "author_url": "",
      "post_date": "2020-09-18T04:19:22.250000",
      "content": "<p>Congrats on solo gold and becoming GM !  And thanks for your good write up.<br>\nI have some questions:</p>\n<blockquote>\n  <p>Having avg/map pooling head gives a chance to get a free segmentation.</p>\n</blockquote>\n<p>How did you choose between max and avg, I was very torn on which one to use in this comp, they were good and bad at times.</p>\n<blockquote>\n  <p>I found that problem was in receptive field - it's too big.</p>\n</blockquote>\n<p>What's an amazing finding! Could you talk more about how did you find it? </p>",
      "votes": 1,
      "replies": [
        {
          "id": 1019221,
          "author_name": "Oleg Yaroshevskiy",
          "author_url": "",
          "post_date": "2020-09-20T09:29:30.403000",
          "content": "<p>Truly to say, as I remember max and avg+max poolings worked better. </p>\n<p>Yes, I tried to train on different input lengths - 5/10/15/30 seconds and monitored my test set. I found that PANNs CNN14 generalized well to 5 sec (!) inference but EffNet didn't. I've tried different configurations: padding types, kernel sizes, number of blocks repeats to find that the problem was in receptive field. </p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1019282,
          "author_name": "Yu Kang",
          "author_url": "",
          "post_date": "2020-09-20T10:27:29.043000",
          "content": "<p>Thanks for answering.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1013465,
      "author_name": "MPWARE",
      "author_url": "",
      "post_date": "2020-09-16T17:51:33.833000",
      "content": "<p>Nice solution! Our sanity-check dataset was not far from yours! But we had some variance on LB correlation.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1013229,
      "author_name": "Varonos",
      "author_url": "",
      "post_date": "2020-09-16T15:19:33.800000",
      "content": "<p>Congratulations! Keep posting so we (the newbies) will learn about new methods and tricks in competition.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1012735,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-09-16T09:00:43.543000",
      "content": "<p>Thank you for sharing your solution, congratulation👍</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1018591,
      "author_name": " Akash Gupta",
      "author_url": "",
      "post_date": "2020-09-19T18:58:15.840000",
      "content": "<p>congratulation :)</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 1013065,
      "author_name": "Abdur Rahim",
      "author_url": "",
      "post_date": "2020-09-16T13:39:11.233000",
      "content": "<p>Wow, that's great 😍<br>\nThanks a lot for sharing.<br>\nCongratulations for solo GOLD and becoming GM 😍</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 1012954,
      "author_name": "SeshuRaju 🧘‍♂️",
      "author_url": "",
      "post_date": "2020-09-16T12:18:43.503000",
      "content": "<p>congratulation to become GrandMaster :) </p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 1012639,
      "author_name": "yukiya",
      "author_url": "",
      "post_date": "2020-09-16T07:44:01.343000",
      "content": "<p>Congratulations and looking forward for more details !<br>\nHope you stay safe where you are (the island of the gods must be very blissful without the crowd of tourists)</p>",
      "votes": 2,
      "replies": [
        {
          "id": 1012684,
          "author_name": "Oleg Yaroshevskiy",
          "author_url": "",
          "post_date": "2020-09-16T08:13:59.387000",
          "content": "<p>Yeah, thanks Bali for having me during pandemic ^^</p>",
          "votes": 4,
          "replies": []
        }
      ]
    },
    {
      "id": 1022572,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-09-22T16:08:50.457000",
      "content": "",
      "votes": 1,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1012563": "A great journey, thanks for organizing it! Now every time I walk around I hear much more birds :) \nI've started my machine learning journey with speech processing few years ago and that was a joy to play again with it. One more time to thank and mention @hidehisaarai1213 who led us through the darkness here.\n\nFrom the very scratch I understood that the problem can be decomposed into two main problems:\n\n- a domain shift (which is pretty tough us we don't have target test set)\n- a clipwise to framewise classification transit (which also introduced us additional \"nocall\" class with huge influence on results, and labels noise)\n\nTackling those two needed a validation/test set as close to target distribution as possible in terms of snr and nocall distributions. \n\n**Validation**\n\nBelieve me or not but with 6 records from [here](https://www.kaggle.com/c/birdsong-recognition/discussion/158877#911336) and 2 example audios I was able to get some correlation with LB even though it contained only 500-800 clips. But more important I was able to control % of \"nocall\"s and pick the right threshold based on them. Last two days I splitted those clips into 3 test sets with 50%, 60%, 80% \"nocalls\" to see possible scenarios on private leaderboard (thanks organizers private is almost equal to public) and to secure the score with my second submission. \n\nAs a target metric I've used F1 by `average=\"samples\"` pointed by @cpmp. By the end I've also monitored validation (default CV split) primary+secondary F1 score as my another decision-making metric. Primary F1 or mAP was not enough.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F805012%2Fa290c42c9808e69c800abb79f2889f22%2FScreenshot%20from%202020-09-16%2009-53-50.png?generation=1600239312223536&alt=media)\n\nAdding separate nocall class didn't work for me for whatever reason.\nAnd yeah, I trusted leaderboard too.\n\n**Domain shift**\nAssuming we don't have the transit problem, this might be solved simply with augmentations. I've spent too much time on this and right now I understand that was not productive. My final models contain different augmentations configurations: \n- backgrounds (all from external data thread)\n- pink and brown noise\n- pitch shift\n- low pass filtering\n- spec augments (time and frequency masking)\n\nAs organizers informed that we can train on two test examples, I tried to collect batch norm statistics from those as a domain adaptation technique but it didn't work great. \n\n**Clipwise > framewise transit p.1**\n\nThat's my favorite part!\nSo what we have here - every time we crop 5 seconds clip we have a chance to crop a nocall clip. So labels become really noisy. Even more it's hard to crop secondary classes. Easy way to tackle this? Label smoothing 0.2. I don't remember when I became a fan of label smoothing but it works well with noisy labels on practice. But that's not serious.\n\nPeople here tackled this task having an energy based cuts. And it really worked. So I've tried both approaches: soft and hard. By soft I mean random sampling based on energy, by hard - removing everything below normalized energy threshold.\n\nBut what if we can use our model to extract these labels? Having avg/map pooling head gives a chance to get a free segmentation:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F805012%2Fb9938f7758cd413a47006ac98c5f62e0%2Fphoto5377374165436313567.jpg?generation=1600241278088906&alt=media)\nwhich you can use as soft labels or hard binarized labels. These methods showed really cool performance for primary labels. For secondary I kept label smoothing :)\n\n**Clipwise > framewise transit p.2 (The most important part)**\n\nI found SED models comparably weak trained on 5 sec clips but as I trained them on longer clips (10-30 seconds) I noticed that due to labels noise reduction they show much better training loss/mAP. Moreover, it was possible to run inference on 5 sec clips with nice performance. So I took B4 EffNet, added the same Attention decision making head and ... failed!\n\nI've spent a lot on this one. So why did Cnn14_DecisionLevelAtt trained on 30 sec clips work well on 5 sec clips? Playing with EffNet I found that problem was in receptive field - it's too big (630 compared to 200 or so). With wide receptive field attention head was not able to build a meaningful framewise feature maps. As I've changed EffNet kernels from 5x5 to 3x3 or reduced number of blocks - it solved problem but price was too big - pretrained weights. \n\nSo what we have: \n- longer input\n- clipping based on some weakly labeled probabilities\n- label smoothing for secondary\n\nHaving these I've decided to simply run 5sec classification inference with no sophisticated postprocessing that might fail on private LB. As it always happens, I've found this too late so trained only two PANNs models (resnet38 and cnn14 mentioned above) on 12 and 15 second clips (128 mels) with mixup and didn't have time to configure them properly.\n\n**Finally**\n\nI've optimized inference to run all my models (cnn14, resnet38 and few effnets) in 20 mins with `kaiser_fast` resampling and around one hour with `kaiser_best`. I've picked thresholds based on my test set (0.3) and some skepticism about it (0.4). I found I have 670+ private score submissions from two weeks ago, which I can't explain. Probably, Kaggle leaderboard is another stochastic process.\n\nThis has been a long journey in which I not only solved many competitions and learned just hundreds of tools and tricks, but also met wonderful minds from all over the world and found job of my dreams. I wish to meet you in person at Kaggle Days once covid is over. See you!",
    "1012616": "Congrats on solo gold medal and becoming GM @yaroshevskiy ",
    "1012743": "Congrats on the solo gold result!  I am glad you investigated why attention head on top of effnet wasn't working, I saw the same but didn't know why.  How did you ensemble your models?  We failed he few times we tried.\n\nIt looks like you're the only top finisher to have set a CV that mimics test.  Another good move from you IMHO.",
    "1012724": "Congrats on the solo gold Oleg! 🎉",
    "1025427": "Congrats on GM",
    "1016763": "Your story is really inspired me. I am a beginner in speech recognition. ",
    "1015242": "Congrats on solo gold and becoming GM !  And thanks for your good write up.\nI have some questions:\n\n> Having avg/map pooling head gives a chance to get a free segmentation.\n\nHow did you choose between max and avg, I was very torn on which one to use in this comp, they were good and bad at times.\n\n> I found that problem was in receptive field - it's too big.\n\nWhat's an amazing finding! Could you talk more about how did you find it? ",
    "1013465": "Nice solution! Our sanity-check dataset was not far from yours! But we had some variance on LB correlation.",
    "1013229": "Congratulations! Keep posting so we (the newbies) will learn about new methods and tricks in competition.",
    "1012735": "Thank you for sharing your solution, congratulation👍",
    "1018591": "congratulation :)",
    "1013065": "Wow, that's great 😍\nThanks a lot for sharing.\nCongratulations for solo GOLD and becoming GM 😍",
    "1012954": " congratulation to become GrandMaster :) ",
    "1012639": "Congratulations and looking forward for more details !\nHope you stay safe where you are (the island of the gods must be very blissful without the crowd of tourists)",
    "1022572": ""
  }
}