{
  "id": 220309,
  "title": "[completed] rank 19th solution in 13 days - LB 0.937/0.939 (public/private) 5 fold model",
  "url": "/competitions/rfcx-species-audio-detection/writeups/completed-rank-19th-solution-in-13-days-lb-0-937-0",
  "author_name": "",
  "post_date": "2021-02-18T12:00:41.400Z",
  "votes": 51,
  "comment_count": 9,
  "views": 0,
  "content": "<p>[summary]</p>\n<ul>\n<li><p>Design a sliding window conv classifier net. Trained with tp and fp annotations (no pesudo label) on resnet34, this architechiture achieved LB 0.937/0.939 (public/private) in 5 fold model. A single fold model gives 0.881/0.888 (public/private). The first max pooling resnet34 is removed to make the stride of the backbone 16 (instead of 32)</p></li>\n<li><p>Final submission is LB 0.945/0.943 (public rank 15/private rank 19) is an ensemble of this network with different CNN backbone.</p></li>\n</ul>\n<hr>\n<p>I am grateful for Kaggle and the host Rainforest Connection (RFCx) for organizing the competition. </p>\n<p>As a Z by HP &amp; NVidia global datascience ambassador under <a href=\"https://datascience.hp.com/us/en.html.HP\" target=\"_blank\">https://datascience.hp.com/us/en.html.HP</a> &amp; Nvidia has kindly provided a Z8 datascience workstation for my use in this competition. Without this powerful workstation, it would not be possible for me to develop my idea from scratch within 13 days. </p>\n<p>Because I can finish experiments at a very great speed (z8 has ax quadro rtx 8000, NVlink 96GB), I gain a lot of insights into model training and model design when the experimental feedback is almost instantaneous. I want to share these insights with the kaggle community. </p>\n<p>Hence I started another thread to guide and supervise kagglers who wants to improve their training skills. This is targeted to bring them from sliver to gold levels. You can refer to:<br>\n<a href=\"https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/220379\" target=\"_blank\">https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/220379</a><br>\n<a href=\"https://www.kaggle.com/c/hpa-single-cell-image-classification/discussion/217238\" target=\"_blank\">https://www.kaggle.com/c/hpa-single-cell-image-classification/discussion/217238</a></p>\n<hr>\n<p>data preprocessing</p>\n<ul>\n<li>log mel spectrogram with modified PCEN denoising (per- channel energy normalization). A 10sec clip has melspec of size 128x938<br>\nn_fft           = 2048<br>\nwin_length = 2048<br>\nhop_length = 512  <br>\nnum_freq   = 128</li>\n</ul>\n<hr>\n<p>augmentation</p>\n<ul>\n<li>I haven't tried much augmentation yet. For TP annotation, random shift by 0.02 sec. For FP annotation, random shift width of the annotation and then random flip, mixup, cut and paste of the original annotation and its shift version.</li>\n<li>Augmentation is done in the temporal time domain because of my code. (this is not the best solution)</li>\n<li>It can be observed that if I increased the FP augmentation, the bce loss of the validation FP decreases and the public LB score improved.</li>\n<li>heavy dropout in model</li>\n</ul>\n<hr>\n<p>model and loss</p>\n<ul>\n<li>please see the below images</li>\n<li>during training, I monitor the separate log loss of validation TP and FP. I also use the LRAP assumpting one label (i.e. top-1 correctness). This is to decide when to early stop or decrease learning rate. These are not the best approach.</li>\n</ul>\n<p><img src=\"https://i.ibb.co/txpjSWk/Selection-041.png\" alt=\"\"><br>\n<img src=\"https://i.ibb.co/C9fQh7P/Selection-042.png\" alt=\"\"><br>\n<img src=\"https://i.ibb.co/wymW9fB/Selection-040.png\" alt=\"\"><br>\n<img src=\"https://i.ibb.co/9hhCfrw/Selection-039.png\" alt=\"\"></p>\n<p>daily progress:<br>\n<img src=\"https://i.ibb.co/7kGcfpj/Selection-050.png\" alt=\"\"><br>\nI enter the competition are reading CPMP post, saying that he has 0.930 LB in the first submission. WOW! Reading between the lines, it means:</p>\n<ol>\n<li>the public kernels are doing it \"wrong\". It could be the problem setting, the network architecture, the data, or some magic features. there is something very fundamental that can do better than the public kernel.</li>\n<li>0.930 is quite high score (in the gold region). With 2 more week to the deadline, I decided to give it a try. if I solved this unobvious puzzle, i may end up in the gold region as well.</li>\n<li>the first step is to go through the public kernel (some PANN SED) and the baseline image classifier. I want to see what most people are doing (and could be improved)</li>\n<li>after working for a week, I realize that there are two major \"potential flaws\":<ul>\n<li>treating it as a multi-instance learning SED. Well, MIL could be a solution, but the problem is that we don't have bag label (clip label). Most MIL required bag level label, but lack instance level label (segment label).Here we have the opposite.</li>\n<li>not using FP in train. Most public kernel use only TP as train only. </li></ul></li>\n<li>hence I start to design my own network architecture. The first step is an efficient way to do crop classification based on mean annotation box. Hence I design the slide window network.</li>\n<li>The next step is to use TP+FP in training</li>\n<li>The last step is to use pesudo labels, but I don't have time to complete this. But I do have some initial experimental results on this. Top-1 (max over time) pesudo labels is about 95% accurate for LB score of 0.94. This is good enough for distillation.</li>\n<li>Because pesudo labeling requires many probing to prevent error propagation, I could not do it because I don't have sufficient slots in the last days. Worst still, there is no train data at clip level at all. This makes local testing impossible.</li>\n</ol>\n<p>I was able to train 5-fold efficientB0 and resnet50 in the last day. Because of the large GPU card of the HP workstation I am using, I can train large models with large batch size. When I compared training the same model with smaller batch size on my old machines, I find that the results are different and inferior, even if I use gradient accumulation.</p>\n<p>I strongly feel that we have reached a new era. I think this is also why the latest RTX3090 has larger GPU memory than the previous cards. The future is the transformers …. meaning more GPU memory. That's is how fast deep learning are moving!</p>",
  "messages": [
    {
      "id": "1207607",
      "postDate": "02/18/2021 00:13:24",
      "content": "<p>[summary]</p>\n<ul>\n<li><p>Design a sliding window conv classifier net. Trained with tp and fp annotations (no pesudo label) on resnet34, this architechiture achieved LB 0.937/0.939 (public/private) in 5 fold model. A single fold model gives 0.881/0.888 (public/private). The first max pooling resnet34 is removed to make the stride of the backbone 16 (instead of 32)</p></li>\n<li><p>Final submission is LB 0.945/0.943 (public rank 15/private rank 19) is an ensemble of this network with different CNN backbone.</p></li>\n</ul>\n<hr>\n<p>I am grateful for Kaggle and the host Rainforest Connection (RFCx) for organizing the competition. </p>\n<p>As a Z by HP &amp; NVidia global datascience ambassador under <a href=\"https://datascience.hp.com/us/en.html.HP\" target=\"_blank\">https://datascience.hp.com/us/en.html.HP</a> &amp; Nvidia has kindly provided a Z8 datascience workstation for my use in this competition. Without this powerful workstation, it would not be possible for me to develop my idea from scratch within 13 days. </p>\n<p>Because I can finish experiments at a very great speed (z8 has ax quadro rtx 8000, NVlink 96GB), I gain a lot of insights into model training and model design when the experimental feedback is almost instantaneous. I want to share these insights with the kaggle community. </p>\n<p>Hence I started another thread to guide and supervise kagglers who wants to improve their training skills. This is targeted to bring them from sliver to gold levels. You can refer to:<br>\n<a href=\"https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/220379\" target=\"_blank\">https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/220379</a><br>\n<a href=\"https://www.kaggle.com/c/hpa-single-cell-image-classification/discussion/217238\" target=\"_blank\">https://www.kaggle.com/c/hpa-single-cell-image-classification/discussion/217238</a></p>\n<hr>\n<p>data preprocessing</p>\n<ul>\n<li>log mel spectrogram with modified PCEN denoising (per- channel energy normalization). A 10sec clip has melspec of size 128x938<br>\nn_fft           = 2048<br>\nwin_length = 2048<br>\nhop_length = 512  <br>\nnum_freq   = 128</li>\n</ul>\n<hr>\n<p>augmentation</p>\n<ul>\n<li>I haven't tried much augmentation yet. For TP annotation, random shift by 0.02 sec. For FP annotation, random shift width of the annotation and then random flip, mixup, cut and paste of the original annotation and its shift version.</li>\n<li>Augmentation is done in the temporal time domain because of my code. (this is not the best solution)</li>\n<li>It can be observed that if I increased the FP augmentation, the bce loss of the validation FP decreases and the public LB score improved.</li>\n<li>heavy dropout in model</li>\n</ul>\n<hr>\n<p>model and loss</p>\n<ul>\n<li>please see the below images</li>\n<li>during training, I monitor the separate log loss of validation TP and FP. I also use the LRAP assumpting one label (i.e. top-1 correctness). This is to decide when to early stop or decrease learning rate. These are not the best approach.</li>\n</ul>\n<p><img src=\"https://i.ibb.co/txpjSWk/Selection-041.png\" alt=\"\"><br>\n<img src=\"https://i.ibb.co/C9fQh7P/Selection-042.png\" alt=\"\"><br>\n<img src=\"https://i.ibb.co/wymW9fB/Selection-040.png\" alt=\"\"><br>\n<img src=\"https://i.ibb.co/9hhCfrw/Selection-039.png\" alt=\"\"></p>\n<p>daily progress:<br>\n<img src=\"https://i.ibb.co/7kGcfpj/Selection-050.png\" alt=\"\"><br>\nI enter the competition are reading CPMP post, saying that he has 0.930 LB in the first submission. WOW! Reading between the lines, it means:</p>\n<ol>\n<li>the public kernels are doing it \"wrong\". It could be the problem setting, the network architecture, the data, or some magic features. there is something very fundamental that can do better than the public kernel.</li>\n<li>0.930 is quite high score (in the gold region). With 2 more week to the deadline, I decided to give it a try. if I solved this unobvious puzzle, i may end up in the gold region as well.</li>\n<li>the first step is to go through the public kernel (some PANN SED) and the baseline image classifier. I want to see what most people are doing (and could be improved)</li>\n<li>after working for a week, I realize that there are two major \"potential flaws\":<ul>\n<li>treating it as a multi-instance learning SED. Well, MIL could be a solution, but the problem is that we don't have bag label (clip label). Most MIL required bag level label, but lack instance level label (segment label).Here we have the opposite.</li>\n<li>not using FP in train. Most public kernel use only TP as train only. </li></ul></li>\n<li>hence I start to design my own network architecture. The first step is an efficient way to do crop classification based on mean annotation box. Hence I design the slide window network.</li>\n<li>The next step is to use TP+FP in training</li>\n<li>The last step is to use pesudo labels, but I don't have time to complete this. But I do have some initial experimental results on this. Top-1 (max over time) pesudo labels is about 95% accurate for LB score of 0.94. This is good enough for distillation.</li>\n<li>Because pesudo labeling requires many probing to prevent error propagation, I could not do it because I don't have sufficient slots in the last days. Worst still, there is no train data at clip level at all. This makes local testing impossible.</li>\n</ol>\n<p>I was able to train 5-fold efficientB0 and resnet50 in the last day. Because of the large GPU card of the HP workstation I am using, I can train large models with large batch size. When I compared training the same model with smaller batch size on my old machines, I find that the results are different and inferior, even if I use gradient accumulation.</p>\n<p>I strongly feel that we have reached a new era. I think this is also why the latest RTX3090 has larger GPU memory than the previous cards. The future is the transformers …. meaning more GPU memory. That's is how fast deep learning are moving!</p>",
      "rawMarkdown": "[summary]\n\n- Design a sliding window conv classifier net. Trained with tp and fp annotations (no pesudo label) on resnet34, this architechiture achieved LB 0.937/0.939 (public/private) in 5 fold model. A single fold model gives 0.881/0.888 (public/private). The first max pooling resnet34 is removed to make the stride of the backbone 16 (instead of 32)\n\n- Final submission is LB 0.945/0.943 (public rank 15/private rank 19) is an ensemble of this network with different CNN backbone.\n\n---\n\nI am grateful for Kaggle and the host Rainforest Connection (RFCx) for organizing the competition. \n\nAs a Z by HP & NVidia global datascience ambassador under https://datascience.hp.com/us/en.html.HP & Nvidia has kindly provided a Z8 datascience workstation for my use in this competition. Without this powerful workstation, it would not be possible for me to develop my idea from scratch within 13 days. \n\nBecause I can finish experiments at a very great speed (z8 has ax quadro rtx 8000, NVlink 96GB), I gain a lot of insights into model training and model design when the experimental feedback is almost instantaneous. I want to share these insights with the kaggle community. \n\nHence I started another thread to guide and supervise kagglers who wants to improve their training skills. This is targeted to bring them from sliver to gold levels. You can refer to:\nhttps://www.kaggle.com/c/rfcx-species-audio-detection/discussion/220379\nhttps://www.kaggle.com/c/hpa-single-cell-image-classification/discussion/217238\n\n---\ndata preprocessing\n- log mel spectrogram with modified PCEN denoising (per- channel energy normalization). A 10sec clip has melspec of size 128x938\nn_fft           = 2048\nwin_length = 2048\nhop_length = 512  \nnum_freq   = 128\n\n\n---\naugmentation\n- I haven't tried much augmentation yet. For TP annotation, random shift by 0.02 sec. For FP annotation, random shift width of the annotation and then random flip, mixup, cut and paste of the original annotation and its shift version.\n- Augmentation is done in the temporal time domain because of my code. (this is not the best solution)\n- It can be observed that if I increased the FP augmentation, the bce loss of the validation FP decreases and the public LB score improved.\n- heavy dropout in model\n\n---\nmodel and loss\n- please see the below images\n- during training, I monitor the separate log loss of validation TP and FP. I also use the LRAP assumpting one label (i.e. top-1 correctness). This is to decide when to early stop or decrease learning rate. These are not the best approach.\n\n\n![](https://i.ibb.co/txpjSWk/Selection-041.png)\n![](https://i.ibb.co/C9fQh7P/Selection-042.png)\n![](https://i.ibb.co/wymW9fB/Selection-040.png)\n![](https://i.ibb.co/9hhCfrw/Selection-039.png)\n\ndaily progress:\n![](https://i.ibb.co/7kGcfpj/Selection-050.png)\nI enter the competition are reading CPMP post, saying that he has 0.930 LB in the first submission. WOW! Reading between the lines, it means:\n1. the public kernels are doing it \"wrong\". It could be the problem setting, the network architecture, the data, or some magic features. there is something very fundamental that can do better than the public kernel.\n2. 0.930 is quite high score (in the gold region). With 2 more week to the deadline, I decided to give it a try. if I solved this unobvious puzzle, i may end up in the gold region as well.\n3. the first step is to go through the public kernel (some PANN SED) and the baseline image classifier. I want to see what most people are doing (and could be improved)\n4. after working for a week, I realize that there are two major \"potential flaws\":\n  - treating it as a multi-instance learning SED. Well, MIL could be a solution, but the problem is that we don't have bag label (clip label). Most MIL required bag level label, but lack instance level label (segment label).Here we have the opposite.\n  - not using FP in train. Most public kernel use only TP as train only. \n5. hence I start to design my own network architecture. The first step is an efficient way to do crop classification based on mean annotation box. Hence I design the slide window network.\n6. The next step is to use TP+FP in training\n7. The last step is to use pesudo labels, but I don't have time to complete this. But I do have some initial experimental results on this. Top-1 (max over time) pesudo labels is about 95% accurate for LB score of 0.94. This is good enough for distillation.\n8. Because pesudo labeling requires many probing to prevent error propagation, I could not do it because I don't have sufficient slots in the last days. Worst still, there is no train data at clip level at all. This makes local testing impossible.\n\nI was able to train 5-fold efficientB0 and resnet50 in the last day. Because of the large GPU card of the HP workstation I am using, I can train large models with large batch size. When I compared training the same model with smaller batch size on my old machines, I find that the results are different and inferior, even if I use gradient accumulation.\n\nI strongly feel that we have reached a new era. I think this is also why the latest RTX3090 has larger GPU memory than the previous cards. The future is the transformers .... meaning more GPU memory. That's is how fast deep learning are moving!",
      "votes": null
    },
    {
      "id": "1207639",
      "postDate": "02/18/2021 00:35:19",
      "content": "<p>anyone can suggest a good free image hosting website. The one I am using may delete the images after some time.</p>",
      "rawMarkdown": "anyone can suggest a good free image hosting website. The one I am using may delete the images after some time.",
      "votes": null
    },
    {
      "id": "1207642",
      "postDate": "02/18/2021 00:37:12",
      "content": "<p>I think github gist is good </p>",
      "rawMarkdown": "I think github gist is good",
      "votes": null
    },
    {
      "id": "1207646",
      "postDate": "02/18/2021 00:38:15",
      "content": "<p>Devil is in detail but it seems we have a similar method.  I welcome comments about what we did differently.  And congrats on the good result in few days.</p>",
      "rawMarkdown": "Devil is in detail but it seems we have a similar method.  I welcome comments about what we did differently.  And congrats on the good result in few days.",
      "votes": null
    },
    {
      "id": "1207754",
      "postDate": "02/18/2021 02:52:44",
      "content": "<p>Congratulations for strong solo finish👍 Looking forward to hear about 0.881(1fold)-&gt;0.937(5fold) boost! I was able to get 0.92x with single fold single model but my score worsened when doing 5fold ensembling..<br>\nEdit) Never mind, I checked and my other fold lb scores were bad, that was the reason</p>",
      "rawMarkdown": "Congratulations for strong solo finish👍 Looking forward to hear about 0.881(1fold)->0.937(5fold) boost! I was able to get 0.92x with single fold single model but my score worsened when doing 5fold ensembling..\nEdit) Never mind, I checked and my other fold lb scores were bad, that was the reason",
      "votes": null
    },
    {
      "id": "1207868",
      "postDate": "02/18/2021 04:17:59",
      "content": "<p>I've been using imgur</p>",
      "rawMarkdown": "I've been using imgur",
      "votes": null
    },
    {
      "id": "1207870",
      "postDate": "02/18/2021 04:21:02",
      "content": "<p><a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> Congratulations :D, can you share ur make_freq_mask function ? </p>",
      "rawMarkdown": "hengck23 Congratulations :D, can you share ur make_freq_mask function ?",
      "votes": null
    },
    {
      "id": "1208355",
      "postDate": "02/18/2021 09:00:01",
      "content": "<p><a href=\"https://www.kaggle.com/CPMP\" target=\"_blank\">@CPMP</a> Thanks for the comment, Your first 0.930 submission inspires me to enter the competition and race against time. Thanks a lot and congrats on your solo gold!</p>\n<p>here are what I think are the same and differences of our approaches:</p>\n<p>the same: high level objectives</p>\n<ul>\n<li>to use limited frequency fmin and fmax per class</li>\n<li>basically, train a classifier on the \"crop spectrograms\"</li>\n<li>make the CNN <strong>variant</strong> to the frequency information</li>\n</ul>\n<p>the differences: implementation details</p>\n<ul>\n<li><p>you crop the sepctrograms but I didn't. I have a larger receptive field (could be more signal or more noise). I use sliding roi window instead. There is no resizing of the roi to fixed size like 224x224. hence I preserved the same scale across all classes(with may or may not be a good thing??)</p></li>\n<li><p>most CNN do average pooling at the last layer before classification. Here I use a conv kernel to do aggregation instead. In this way I make the CNN <strong>variant</strong> to the frequency information.</p></li>\n</ul>\n<p>Apart from these, I think the rest are the same. I can see that the use of TP and FP and loss function are the same. There will be some differences in augmentation, etc.</p>",
      "rawMarkdown": "CPMP Thanks for the comment, Your first 0.930 submission inspires me to enter the competition and race against time. Thanks a lot and congrats on your solo gold!\n\nhere are what I think are the same and differences of our approaches:\n\nthe same: high level objectives\n  - to use limited frequency fmin and fmax per class\n  - basically, train a classifier on the \"crop spectrograms\"\n  - make the CNN **variant** to the frequency information\n\nthe differences: implementation details\n  - you crop the sepctrograms but I didn't. I have a larger receptive field (could be more signal or more noise). I use sliding roi window instead. There is no resizing of the roi to fixed size like 224x224. hence I preserved the same scale across all classes(with may or may not be a good thing??)\n\n  - most CNN do average pooling at the last layer before classification. Here I use a conv kernel to do aggregation instead. In this way I make the CNN **variant** to the frequency information.\n\nApart from these, I think the rest are the same. I can see that the use of TP and FP and loss function are the same. There will be some differences in augmentation, etc.",
      "votes": null
    },
    {
      "id": "1208501",
      "postDate": "02/18/2021 09:57:21",
      "content": "<p>Congrats <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> on your solo achievement. Waiting eagerly to see full writeup. I can learn a lot here!</p>",
      "rawMarkdown": "Congrats @hengck23 on your solo achievement. Waiting eagerly to see full writeup. I can learn a lot here!",
      "votes": null
    },
    {
      "id": "1209446",
      "postDate": "02/18/2021 22:45:48",
      "content": "<p>Thanks for adding great explanations. Our methods are very close, but your implementation is probably more efficient as you avoid my overlapping crops along the time dimension for a given f_min, f_max pair.  It works becaus eyou use max pooling where I used mean pooling.</p>\n<p>And thank you for mentioning my first sub as motivation.  Unless mistaken you're the only one acknowledging it helped.  Little fix: I got 0.931 at first sub, not 0.930 ;)</p>\n<p>Good work in few days, congrats!</p>",
      "rawMarkdown": "Thanks for adding great explanations. Our methods are very close, but your implementation is probably more efficient as you avoid my overlapping crops along the time dimension for a given f_min, f_max pair.  It works becaus eyou use max pooling where I used mean pooling.\n\nAnd thank you for mentioning my first sub as motivation.  Unless mistaken you're the only one acknowledging it helped.  Little fix: I got 0.931 at first sub, not 0.930 ;)\n\nGood work in few days, congrats!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1207639,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "02/18/2021 00:35:19",
      "content": "<p>anyone can suggest a good free image hosting website. The one I am using may delete the images after some time.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1207642,
          "author_name": "hidehisaarai1213",
          "author_url": "",
          "post_date": "02/18/2021 00:37:12",
          "content": "<p>I think github gist is good </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1207868,
          "author_name": "ryches",
          "author_url": "",
          "post_date": "02/18/2021 04:17:59",
          "content": "<p>I've been using imgur</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1207646,
      "author_name": "cpmpml",
      "author_url": "",
      "post_date": "02/18/2021 00:38:15",
      "content": "<p>Devil is in detail but it seems we have a similar method.  I welcome comments about what we did differently.  And congrats on the good result in few days.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1208355,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "02/18/2021 09:00:01",
          "content": "<p><a href=\"https://www.kaggle.com/CPMP\" target=\"_blank\">@CPMP</a> Thanks for the comment, Your first 0.930 submission inspires me to enter the competition and race against time. Thanks a lot and congrats on your solo gold!</p>\n<p>here are what I think are the same and differences of our approaches:</p>\n<p>the same: high level objectives</p>\n<ul>\n<li>to use limited frequency fmin and fmax per class</li>\n<li>basically, train a classifier on the \"crop spectrograms\"</li>\n<li>make the CNN <strong>variant</strong> to the frequency information</li>\n</ul>\n<p>the differences: implementation details</p>\n<ul>\n<li><p>you crop the sepctrograms but I didn't. I have a larger receptive field (could be more signal or more noise). I use sliding roi window instead. There is no resizing of the roi to fixed size like 224x224. hence I preserved the same scale across all classes(with may or may not be a good thing??)</p></li>\n<li><p>most CNN do average pooling at the last layer before classification. Here I use a conv kernel to do aggregation instead. In this way I make the CNN <strong>variant</strong> to the frequency information.</p></li>\n</ul>\n<p>Apart from these, I think the rest are the same. I can see that the use of TP and FP and loss function are the same. There will be some differences in augmentation, etc.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1207754,
      "author_name": "harangdev",
      "author_url": "",
      "post_date": "02/18/2021 02:52:44",
      "content": "<p>Congratulations for strong solo finish👍 Looking forward to hear about 0.881(1fold)-&gt;0.937(5fold) boost! I was able to get 0.92x with single fold single model but my score worsened when doing 5fold ensembling..<br>\nEdit) Never mind, I checked and my other fold lb scores were bad, that was the reason</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1207870,
      "author_name": "dathudeptrai",
      "author_url": "",
      "post_date": "02/18/2021 04:21:02",
      "content": "<p><a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> Congratulations :D, can you share ur make_freq_mask function ? </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1208501,
      "author_name": "rajkumarl",
      "author_url": "",
      "post_date": "02/18/2021 09:57:21",
      "content": "<p>Congrats <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> on your solo achievement. Waiting eagerly to see full writeup. I can learn a lot here!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1209446,
      "author_name": "cpmpml",
      "author_url": "",
      "post_date": "02/18/2021 22:45:48",
      "content": "<p>Thanks for adding great explanations. Our methods are very close, but your implementation is probably more efficient as you avoid my overlapping crops along the time dimension for a given f_min, f_max pair.  It works becaus eyou use max pooling where I used mean pooling.</p>\n<p>And thank you for mentioning my first sub as motivation.  Unless mistaken you're the only one acknowledging it helped.  Little fix: I got 0.931 at first sub, not 0.930 ;)</p>\n<p>Good work in few days, congrats!</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1207607": "[summary]\n\n- Design a sliding window conv classifier net. Trained with tp and fp annotations (no pesudo label) on resnet34, this architechiture achieved LB 0.937/0.939 (public/private) in 5 fold model. A single fold model gives 0.881/0.888 (public/private). The first max pooling resnet34 is removed to make the stride of the backbone 16 (instead of 32)\n\n- Final submission is LB 0.945/0.943 (public rank 15/private rank 19) is an ensemble of this network with different CNN backbone.\n\n---\n\nI am grateful for Kaggle and the host Rainforest Connection (RFCx) for organizing the competition. \n\nAs a Z by HP & NVidia global datascience ambassador under https://datascience.hp.com/us/en.html.HP & Nvidia has kindly provided a Z8 datascience workstation for my use in this competition. Without this powerful workstation, it would not be possible for me to develop my idea from scratch within 13 days. \n\nBecause I can finish experiments at a very great speed (z8 has ax quadro rtx 8000, NVlink 96GB), I gain a lot of insights into model training and model design when the experimental feedback is almost instantaneous. I want to share these insights with the kaggle community. \n\nHence I started another thread to guide and supervise kagglers who wants to improve their training skills. This is targeted to bring them from sliver to gold levels. You can refer to:\nhttps://www.kaggle.com/c/rfcx-species-audio-detection/discussion/220379\nhttps://www.kaggle.com/c/hpa-single-cell-image-classification/discussion/217238\n\n---\ndata preprocessing\n- log mel spectrogram with modified PCEN denoising (per- channel energy normalization). A 10sec clip has melspec of size 128x938\nn_fft           = 2048\nwin_length = 2048\nhop_length = 512  \nnum_freq   = 128\n\n\n---\naugmentation\n- I haven't tried much augmentation yet. For TP annotation, random shift by 0.02 sec. For FP annotation, random shift width of the annotation and then random flip, mixup, cut and paste of the original annotation and its shift version.\n- Augmentation is done in the temporal time domain because of my code. (this is not the best solution)\n- It can be observed that if I increased the FP augmentation, the bce loss of the validation FP decreases and the public LB score improved.\n- heavy dropout in model\n\n---\nmodel and loss\n- please see the below images\n- during training, I monitor the separate log loss of validation TP and FP. I also use the LRAP assumpting one label (i.e. top-1 correctness). This is to decide when to early stop or decrease learning rate. These are not the best approach.\n\n\n![](https://i.ibb.co/txpjSWk/Selection-041.png)\n![](https://i.ibb.co/C9fQh7P/Selection-042.png)\n![](https://i.ibb.co/wymW9fB/Selection-040.png)\n![](https://i.ibb.co/9hhCfrw/Selection-039.png)\n\ndaily progress:\n![](https://i.ibb.co/7kGcfpj/Selection-050.png)\nI enter the competition are reading CPMP post, saying that he has 0.930 LB in the first submission. WOW! Reading between the lines, it means:\n1. the public kernels are doing it \"wrong\". It could be the problem setting, the network architecture, the data, or some magic features. there is something very fundamental that can do better than the public kernel.\n2. 0.930 is quite high score (in the gold region). With 2 more week to the deadline, I decided to give it a try. if I solved this unobvious puzzle, i may end up in the gold region as well.\n3. the first step is to go through the public kernel (some PANN SED) and the baseline image classifier. I want to see what most people are doing (and could be improved)\n4. after working for a week, I realize that there are two major \"potential flaws\":\n  - treating it as a multi-instance learning SED. Well, MIL could be a solution, but the problem is that we don't have bag label (clip label). Most MIL required bag level label, but lack instance level label (segment label).Here we have the opposite.\n  - not using FP in train. Most public kernel use only TP as train only. \n5. hence I start to design my own network architecture. The first step is an efficient way to do crop classification based on mean annotation box. Hence I design the slide window network.\n6. The next step is to use TP+FP in training\n7. The last step is to use pesudo labels, but I don't have time to complete this. But I do have some initial experimental results on this. Top-1 (max over time) pesudo labels is about 95% accurate for LB score of 0.94. This is good enough for distillation.\n8. Because pesudo labeling requires many probing to prevent error propagation, I could not do it because I don't have sufficient slots in the last days. Worst still, there is no train data at clip level at all. This makes local testing impossible.\n\nI was able to train 5-fold efficientB0 and resnet50 in the last day. Because of the large GPU card of the HP workstation I am using, I can train large models with large batch size. When I compared training the same model with smaller batch size on my old machines, I find that the results are different and inferior, even if I use gradient accumulation.\n\nI strongly feel that we have reached a new era. I think this is also why the latest RTX3090 has larger GPU memory than the previous cards. The future is the transformers .... meaning more GPU memory. That's is how fast deep learning are moving!",
    "1207639": "anyone can suggest a good free image hosting website. The one I am using may delete the images after some time.",
    "1207642": "I think github gist is good",
    "1207646": "Devil is in detail but it seems we have a similar method.  I welcome comments about what we did differently.  And congrats on the good result in few days.",
    "1207754": "Congratulations for strong solo finish👍 Looking forward to hear about 0.881(1fold)->0.937(5fold) boost! I was able to get 0.92x with single fold single model but my score worsened when doing 5fold ensembling..\nEdit) Never mind, I checked and my other fold lb scores were bad, that was the reason",
    "1207868": "I've been using imgur",
    "1207870": "hengck23 Congratulations :D, can you share ur make_freq_mask function ?",
    "1208355": "CPMP Thanks for the comment, Your first 0.930 submission inspires me to enter the competition and race against time. Thanks a lot and congrats on your solo gold!\n\nhere are what I think are the same and differences of our approaches:\n\nthe same: high level objectives\n  - to use limited frequency fmin and fmax per class\n  - basically, train a classifier on the \"crop spectrograms\"\n  - make the CNN **variant** to the frequency information\n\nthe differences: implementation details\n  - you crop the sepctrograms but I didn't. I have a larger receptive field (could be more signal or more noise). I use sliding roi window instead. There is no resizing of the roi to fixed size like 224x224. hence I preserved the same scale across all classes(with may or may not be a good thing??)\n\n  - most CNN do average pooling at the last layer before classification. Here I use a conv kernel to do aggregation instead. In this way I make the CNN **variant** to the frequency information.\n\nApart from these, I think the rest are the same. I can see that the use of TP and FP and loss function are the same. There will be some differences in augmentation, etc.",
    "1208501": "Congrats @hengck23 on your solo achievement. Waiting eagerly to see full writeup. I can learn a lot here!",
    "1209446": "Thanks for adding great explanations. Our methods are very close, but your implementation is probably more efficient as you avoid my overlapping crops along the time dimension for a given f_min, f_max pair.  It works becaus eyou use max pooling where I used mean pooling.\n\nAnd thank you for mentioning my first sub as motivation.  Unless mistaken you're the only one acknowledging it helped.  Little fix: I got 0.931 at first sub, not 0.930 ;)\n\nGood work in few days, congrats!"
  },
  "source": "meta"
}