{
  "id": 220335,
  "title": "43rd place solution",
  "url": "/competitions/rfcx-species-audio-detection/writeups/43rd-place-solution",
  "author_name": "",
  "post_date": "2021-02-19T06:56:28.127Z",
  "votes": 33,
  "comment_count": 6,
  "views": 0,
  "content": "<p>First of all, many thanks to Kaggler.<br>\n I got a lot of ideas from Kaggler in the discussion. And it was fun.</p>\n<p>My solution summary is below.</p>\n<ul>\n<li>The high resolution of spectrogram</li>\n<li>Post-processing with moving average</li>\n<li>Teacher-student model for missing labels</li>\n</ul>\n<h1>1st stage: SED</h1>\n<p>I started from basic experiment with <a href=\"https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/211007\" target=\"_blank\">SED</a>. Why SED? Because I think SED is strong in the multi label task. </p>\n<p>I used log mel-spectrogram as the input of SED. Basic experiment involves data augmentation (Gaussian noise, SpecAugment and MixUP), backbone model choice and adjusting resolution of log mel-spectrogram. As a result, below condition was the best for me.</p>\n<ul>\n<li>No noise injection</li>\n<li>MixUp</li>\n<li>The best model architecture is EfficientNet</li>\n<li><strong>The higher the resolution</strong> of log mel-spectrogram, the better the result.</li>\n</ul>\n<h3>The resolution</h3>\n<p>The most important one is the resolution. Recently, in Kaggle computer vision solution, the higher the resolution of the image, the better the result. In spectrogram, the same phenomenon may happen. In mel-spectrogram, the resolution can be changed by adjusting \"hop_size\" and \"mel_bins\". Following result is changing the resolution with ResNest50(single model).</p>\n<table>\n<thead>\n<tr>\n<th>Resolution(Width-Height)</th>\n<th>public LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>1501-64(PANNs default)</td>\n<td>0.692</td>\n</tr>\n<tr>\n<td>3001-128</td>\n<td>0.805</td>\n</tr>\n<tr>\n<td>6001-64</td>\n<td>0.725</td>\n</tr>\n<tr>\n<td>1501-256</td>\n<td>0.761</td>\n</tr>\n<tr>\n<td>751-512</td>\n<td><strong>0.823</strong></td>\n</tr>\n<tr>\n<td>1501-512</td>\n<td>0.821</td>\n</tr>\n</tbody>\n</table>\n<p>The resolution was critical! According to experimental result, good resolution was \"high\" and similar to the square. 751-512 looks good. As a result, I chose 858-850. This configuration is as follows.</p>\n<pre><code>model_config = {\n    \"sample_rate\": 48000,\n    \"window_size\": 1024,\n    \"hop_size\": 560,\n    \"mel_bins\": 850,\n    \"fmin\": 50,\n    \"fmax\": 14000,\n    \"classes_num\": 24\n}\n</code></pre>\n<h3>Post-processing</h3>\n<p>I used <a href=\"https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/211007\" target=\"_blank\">Framewise output</a> for submission. It contains time and classes information. But there is a lot of false positive information in framewise output. Because they are not processing by a long time information. Therefore <strong>a short event of framewise output should be deleted.</strong> I prepared post-processing for framewise output. It is a moving average.</p>\n<p><img src=\"https://qiita-image-store.s3.ap-northeast-1.amazonaws.com/0/264781/2d0ad5b8-f046-64da-a98a-96f0ae79ba2e.png\" alt=\"image.png\"></p>\n<p>By taking a moving average in the time direction for each class, we can delete short events. This idea is based on <a href=\"http://dcase.community/documents/challenge2020/technical_reports/DCASE2020_Chan_6.pdf\" target=\"_blank\">the paper</a>[1]. The sample code is as follows.</p>\n<pre><code>def post_processing(data): # data.shape = (24, 600) # (classes, time)\n    result = []\n    for i in range(len(data)):\n        result.append(cv2.blur(data[i],(1,31)))\n    return result\n</code></pre>\n<p>I improved LB by using moving average. The following result is comparing post-processing with EfficientNetB3(single model).</p>\n<table>\n<thead>\n<tr>\n<th></th>\n<th>public LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>w/o post-processing</td>\n<td>0.785</td>\n</tr>\n<tr>\n<td>w/ post-processing</td>\n<td><strong>0.840</strong></td>\n</tr>\n</tbody>\n</table>\n<h3>Summary</h3>\n<ul>\n<li>MixUp(alpha=0.1)</li>\n<li>Epoch 30</li>\n<li>Adam(lr=0.001) + CosineAnnealing(T=10)</li>\n<li>Batchsize 6</li>\n<li>Use only tp label</li>\n<li>Get random tp clip 10 sec</li>\n<li>The resolution of log mel-spectrogram: 858-850</li>\n<li>Loss function: BCE</li>\n<li><a href=\"https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/211007\" target=\"_blank\">Weak label training</a></li>\n<li>Post-processing: moving average</li>\n</ul>\n<p>Then I got <strong>0.916 public LB</strong> with EfficientNetB0(5-folds average ensemble).</p>\n<h1>2nd stage: missing labels and the ensemble</h1>\n<p>I reported <a href=\"https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/208830#1139171\" target=\"_blank\">discovering missing labels and re-train</a>. And It didn't work. After that, I thought about missing labels again. My answer is that the model is not correct for discovering missing labels. There are a lot of missing labels around tp. Therefore the model is not correct. </p>\n<p><img src=\"https://qiita-image-store.s3.ap-northeast-1.amazonaws.com/0/264781/0e6009db-c1cf-7297-3ac4-df92dd3420c5.png\" alt=\"train.png\"></p>\n<p>To solve this issue, I used teacher-student model. </p>\n<p><img src=\"https://qiita-image-store.s3.ap-northeast-1.amazonaws.com/0/264781/79d8e1c8-6814-68c6-0a9d-d41d830fcdfb.png\" alt=\"geretation.png\"></p>\n<p>1st generation is similar to 1st stage. I gradually increased model prediction ratio. By using teacher-student model, I could discover missing labels. Specially, in strong label training, teacher-student model was effective. Following result is teacher-student model score with EfficientNetB0.</p>\n<p><img src=\"https://qiita-image-store.s3.ap-northeast-1.amazonaws.com/0/264781/ddded273-c2b1-92c1-3272-1e19377a5b77.png\" alt=\"image.png\"></p>\n<p>\"MixUp rate\" is probabilistic MixUp. This method is based on <a href=\"https://arxiv.org/abs/2102.01243\" target=\"_blank\">the paper</a>[2]. </p>\n<p>Finally, I made the ensemble of 1st stage model and 2nd stage model. Ensemble procedure is simple average. Then I got <strong>0.924 public LB.</strong></p>\n<h1>References</h1>\n<p>[1] Teck Kai Chan, Cheng Siong Chin1 and Ye Li, \"SEMI-SUPERVISED NMF-CNN FOR SOUND EVENT DETECTION\".<br>\n[2] Yuan Gong, Yu-An Chung, and James Glass, \"PSLA: Improving Audio Event Classification with<br>\nPretraining, Sampling, Labeling, and Aggregation\".</p>\n<h1>Appendix: the resolution and EfficientNet</h1>\n<p>Finally, I show interesting result. It is relationship between EfficientNet and the resolution. The following result is public LB(5-folds average ensemble).</p>\n<table>\n<thead>\n<tr>\n<th>Resolution(W-H)</th>\n<th>751-512</th>\n<th>751-751</th>\n<th>858-850</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>EfficientNetB0</td>\n<td>0.893</td>\n<td>0.904</td>\n<td><strong>0.916</strong></td>\n</tr>\n<tr>\n<td>EfficientNetB3</td>\n<td><strong>0.913</strong></td>\n<td>0.912</td>\n<td>0.900</td>\n</tr>\n</tbody>\n</table>\n<p>In B0, the higher resolution, the better result. But B3 was vice versa. Usually, the larger EfficientNet, the better at <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/154683\" target=\"_blank\">high resolution</a> it is. But the above is reverse. Why?</p>\n<p>Maybe <strong>domain shift</strong>(train: noisy sound -&gt; test: clean sound) is concerned. B3 has learned about <strong>train domain features</strong>(noisy sound). On the other hand, B0 has less representational ability than B3. Therefore B0 has learned the <strong>common features</strong> of the train and test domain with high resolution. Without domain shift, B3 would have also shown good results with high resolution.</p>",
  "messages": [
    {
      "id": "1207769",
      "postDate": "02/18/2021 03:05:33",
      "content": "<p>First of all, many thanks to Kaggler.<br>\n I got a lot of ideas from Kaggler in the discussion. And it was fun.</p>\n<p>My solution summary is below.</p>\n<ul>\n<li>The high resolution of spectrogram</li>\n<li>Post-processing with moving average</li>\n<li>Teacher-student model for missing labels</li>\n</ul>\n<h1>1st stage: SED</h1>\n<p>I started from basic experiment with <a href=\"https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/211007\" target=\"_blank\">SED</a>. Why SED? Because I think SED is strong in the multi label task. </p>\n<p>I used log mel-spectrogram as the input of SED. Basic experiment involves data augmentation (Gaussian noise, SpecAugment and MixUP), backbone model choice and adjusting resolution of log mel-spectrogram. As a result, below condition was the best for me.</p>\n<ul>\n<li>No noise injection</li>\n<li>MixUp</li>\n<li>The best model architecture is EfficientNet</li>\n<li><strong>The higher the resolution</strong> of log mel-spectrogram, the better the result.</li>\n</ul>\n<h3>The resolution</h3>\n<p>The most important one is the resolution. Recently, in Kaggle computer vision solution, the higher the resolution of the image, the better the result. In spectrogram, the same phenomenon may happen. In mel-spectrogram, the resolution can be changed by adjusting \"hop_size\" and \"mel_bins\". Following result is changing the resolution with ResNest50(single model).</p>\n<table>\n<thead>\n<tr>\n<th>Resolution(Width-Height)</th>\n<th>public LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>1501-64(PANNs default)</td>\n<td>0.692</td>\n</tr>\n<tr>\n<td>3001-128</td>\n<td>0.805</td>\n</tr>\n<tr>\n<td>6001-64</td>\n<td>0.725</td>\n</tr>\n<tr>\n<td>1501-256</td>\n<td>0.761</td>\n</tr>\n<tr>\n<td>751-512</td>\n<td><strong>0.823</strong></td>\n</tr>\n<tr>\n<td>1501-512</td>\n<td>0.821</td>\n</tr>\n</tbody>\n</table>\n<p>The resolution was critical! According to experimental result, good resolution was \"high\" and similar to the square. 751-512 looks good. As a result, I chose 858-850. This configuration is as follows.</p>\n<pre><code>model_config = {\n    \"sample_rate\": 48000,\n    \"window_size\": 1024,\n    \"hop_size\": 560,\n    \"mel_bins\": 850,\n    \"fmin\": 50,\n    \"fmax\": 14000,\n    \"classes_num\": 24\n}\n</code></pre>\n<h3>Post-processing</h3>\n<p>I used <a href=\"https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/211007\" target=\"_blank\">Framewise output</a> for submission. It contains time and classes information. But there is a lot of false positive information in framewise output. Because they are not processing by a long time information. Therefore <strong>a short event of framewise output should be deleted.</strong> I prepared post-processing for framewise output. It is a moving average.</p>\n<p><img src=\"https://qiita-image-store.s3.ap-northeast-1.amazonaws.com/0/264781/2d0ad5b8-f046-64da-a98a-96f0ae79ba2e.png\" alt=\"image.png\"></p>\n<p>By taking a moving average in the time direction for each class, we can delete short events. This idea is based on <a href=\"http://dcase.community/documents/challenge2020/technical_reports/DCASE2020_Chan_6.pdf\" target=\"_blank\">the paper</a>[1]. The sample code is as follows.</p>\n<pre><code>def post_processing(data): # data.shape = (24, 600) # (classes, time)\n    result = []\n    for i in range(len(data)):\n        result.append(cv2.blur(data[i],(1,31)))\n    return result\n</code></pre>\n<p>I improved LB by using moving average. The following result is comparing post-processing with EfficientNetB3(single model).</p>\n<table>\n<thead>\n<tr>\n<th></th>\n<th>public LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>w/o post-processing</td>\n<td>0.785</td>\n</tr>\n<tr>\n<td>w/ post-processing</td>\n<td><strong>0.840</strong></td>\n</tr>\n</tbody>\n</table>\n<h3>Summary</h3>\n<ul>\n<li>MixUp(alpha=0.1)</li>\n<li>Epoch 30</li>\n<li>Adam(lr=0.001) + CosineAnnealing(T=10)</li>\n<li>Batchsize 6</li>\n<li>Use only tp label</li>\n<li>Get random tp clip 10 sec</li>\n<li>The resolution of log mel-spectrogram: 858-850</li>\n<li>Loss function: BCE</li>\n<li><a href=\"https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/211007\" target=\"_blank\">Weak label training</a></li>\n<li>Post-processing: moving average</li>\n</ul>\n<p>Then I got <strong>0.916 public LB</strong> with EfficientNetB0(5-folds average ensemble).</p>\n<h1>2nd stage: missing labels and the ensemble</h1>\n<p>I reported <a href=\"https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/208830#1139171\" target=\"_blank\">discovering missing labels and re-train</a>. And It didn't work. After that, I thought about missing labels again. My answer is that the model is not correct for discovering missing labels. There are a lot of missing labels around tp. Therefore the model is not correct. </p>\n<p><img src=\"https://qiita-image-store.s3.ap-northeast-1.amazonaws.com/0/264781/0e6009db-c1cf-7297-3ac4-df92dd3420c5.png\" alt=\"train.png\"></p>\n<p>To solve this issue, I used teacher-student model. </p>\n<p><img src=\"https://qiita-image-store.s3.ap-northeast-1.amazonaws.com/0/264781/79d8e1c8-6814-68c6-0a9d-d41d830fcdfb.png\" alt=\"geretation.png\"></p>\n<p>1st generation is similar to 1st stage. I gradually increased model prediction ratio. By using teacher-student model, I could discover missing labels. Specially, in strong label training, teacher-student model was effective. Following result is teacher-student model score with EfficientNetB0.</p>\n<p><img src=\"https://qiita-image-store.s3.ap-northeast-1.amazonaws.com/0/264781/ddded273-c2b1-92c1-3272-1e19377a5b77.png\" alt=\"image.png\"></p>\n<p>\"MixUp rate\" is probabilistic MixUp. This method is based on <a href=\"https://arxiv.org/abs/2102.01243\" target=\"_blank\">the paper</a>[2]. </p>\n<p>Finally, I made the ensemble of 1st stage model and 2nd stage model. Ensemble procedure is simple average. Then I got <strong>0.924 public LB.</strong></p>\n<h1>References</h1>\n<p>[1] Teck Kai Chan, Cheng Siong Chin1 and Ye Li, \"SEMI-SUPERVISED NMF-CNN FOR SOUND EVENT DETECTION\".<br>\n[2] Yuan Gong, Yu-An Chung, and James Glass, \"PSLA: Improving Audio Event Classification with<br>\nPretraining, Sampling, Labeling, and Aggregation\".</p>\n<h1>Appendix: the resolution and EfficientNet</h1>\n<p>Finally, I show interesting result. It is relationship between EfficientNet and the resolution. The following result is public LB(5-folds average ensemble).</p>\n<table>\n<thead>\n<tr>\n<th>Resolution(W-H)</th>\n<th>751-512</th>\n<th>751-751</th>\n<th>858-850</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>EfficientNetB0</td>\n<td>0.893</td>\n<td>0.904</td>\n<td><strong>0.916</strong></td>\n</tr>\n<tr>\n<td>EfficientNetB3</td>\n<td><strong>0.913</strong></td>\n<td>0.912</td>\n<td>0.900</td>\n</tr>\n</tbody>\n</table>\n<p>In B0, the higher resolution, the better result. But B3 was vice versa. Usually, the larger EfficientNet, the better at <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/154683\" target=\"_blank\">high resolution</a> it is. But the above is reverse. Why?</p>\n<p>Maybe <strong>domain shift</strong>(train: noisy sound -&gt; test: clean sound) is concerned. B3 has learned about <strong>train domain features</strong>(noisy sound). On the other hand, B0 has less representational ability than B3. Therefore B0 has learned the <strong>common features</strong> of the train and test domain with high resolution. Without domain shift, B3 would have also shown good results with high resolution.</p>",
      "rawMarkdown": "First of all, many thanks to Kaggler.\n I got a lot of ideas from Kaggler in the discussion. And it was fun.\n\nMy solution summary is below.\n+ The high resolution of spectrogram\n+ Post-processing with moving average\n+ Teacher-student model for missing labels\n\n#1st stage: SED\nI started from basic experiment with [SED](https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/211007). Why SED? Because I think SED is strong in the multi label task. \n\nI used log mel-spectrogram as the input of SED. Basic experiment involves data augmentation (Gaussian noise, SpecAugment and MixUP), backbone model choice and adjusting resolution of log mel-spectrogram. As a result, below condition was the best for me.\n+ No noise injection\n+ MixUp\n+ The best model architecture is EfficientNet\n+ **The higher the resolution** of log mel-spectrogram, the better the result.\n\n### The resolution\nThe most important one is the resolution. Recently, in Kaggle computer vision solution, the higher the resolution of the image, the better the result. In spectrogram, the same phenomenon may happen. In mel-spectrogram, the resolution can be changed by adjusting \"hop_size\" and \"mel_bins\". Following result is changing the resolution with ResNest50(single model).\n\n|Resolution(Width-Height)|public LB|\n|---|---|\n|1501-64(PANNs default)|0.692\n|3001-128|0.805\n|6001-64|0.725\n|1501-256|0.761\n|751-512|**0.823**\n|1501-512|0.821\n\nThe resolution was critical! According to experimental result, good resolution was \"high\" and similar to the square. 751-512 looks good. As a result, I chose 858-850. This configuration is as follows.\n\n```\nmodel_config = {\n    \"sample_rate\": 48000,\n    \"window_size\": 1024,\n    \"hop_size\": 560,\n    \"mel_bins\": 850,\n    \"fmin\": 50,\n    \"fmax\": 14000,\n    \"classes_num\": 24\n}\n```\n\n### Post-processing\nI used [Framewise output](https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/211007) for submission. It contains time and classes information. But there is a lot of false positive information in framewise output. Because they are not processing by a long time information. Therefore **a short event of framewise output should be deleted.** I prepared post-processing for framewise output. It is a moving average.\n\n![image.png](https://qiita-image-store.s3.ap-northeast-1.amazonaws.com/0/264781/2d0ad5b8-f046-64da-a98a-96f0ae79ba2e.png)\n\nBy taking a moving average in the time direction for each class, we can delete short events. This idea is based on [the paper](http://dcase.community/documents/challenge2020/technical_reports/DCASE2020_Chan_6.pdf)[1]. The sample code is as follows.\n```\ndef post_processing(data): # data.shape = (24, 600) # (classes, time)\n    result = []\n    for i in range(len(data)):\n        result.append(cv2.blur(data[i],(1,31)))\n    return result\n```\n\nI improved LB by using moving average. The following result is comparing post-processing with EfficientNetB3(single model).\n\n||public LB|\n|---|---|\n|w/o post-processing|0.785|\n|w/ post-processing|**0.840**|\n \n###Summary\n+ MixUp(alpha=0.1)\n+ Epoch 30\n+ Adam(lr=0.001) + CosineAnnealing(T=10)\n+ Batchsize 6\n+ Use only tp label\n+ Get random tp clip 10 sec\n+ The resolution of log mel-spectrogram: 858-850\n+ Loss function: BCE\n+ [Weak label training](https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/211007)\n+ Post-processing: moving average\n\nThen I got **0.916 public LB** with EfficientNetB0(5-folds average ensemble).\n\n#2nd stage: missing labels and the ensemble\nI reported [discovering missing labels and re-train](https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/208830#1139171). And It didn't work. After that, I thought about missing labels again. My answer is that the model is not correct for discovering missing labels. There are a lot of missing labels around tp. Therefore the model is not correct. \n\n![train.png](https://qiita-image-store.s3.ap-northeast-1.amazonaws.com/0/264781/0e6009db-c1cf-7297-3ac4-df92dd3420c5.png)\n\nTo solve this issue, I used teacher-student model. \n\n![geretation.png](https://qiita-image-store.s3.ap-northeast-1.amazonaws.com/0/264781/79d8e1c8-6814-68c6-0a9d-d41d830fcdfb.png)\n\n1st generation is similar to 1st stage. I gradually increased model prediction ratio. By using teacher-student model, I could discover missing labels. Specially, in strong label training, teacher-student model was effective. Following result is teacher-student model score with EfficientNetB0.\n\n![image.png](https://qiita-image-store.s3.ap-northeast-1.amazonaws.com/0/264781/ddded273-c2b1-92c1-3272-1e19377a5b77.png)\n\n\"MixUp rate\" is probabilistic MixUp. This method is based on [the paper](https://arxiv.org/abs/2102.01243)[2]. \n\nFinally, I made the ensemble of 1st stage model and 2nd stage model. Ensemble procedure is simple average. Then I got **0.924 public LB.**\n\n#References\n[1] Teck Kai Chan, Cheng Siong Chin1 and Ye Li, \"SEMI-SUPERVISED NMF-CNN FOR SOUND EVENT DETECTION\".\n[2] Yuan Gong, Yu-An Chung, and James Glass, \"PSLA: Improving Audio Event Classification with\nPretraining, Sampling, Labeling, and Aggregation\".\n\n#Appendix: the resolution and EfficientNet\nFinally, I show interesting result. It is relationship between EfficientNet and the resolution. The following result is public LB(5-folds average ensemble).\n\n|Resolution(W-H)|751-512|751-751|858-850|\n|---|---|---|---|\n|EfficientNetB0|0.893|0.904|**0.916**|\n|EfficientNetB3|**0.913**|0.912|0.900|\n\nIn B0, the higher resolution, the better result. But B3 was vice versa. Usually, the larger EfficientNet, the better at [high resolution](https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/154683) it is. But the above is reverse. Why?\n\nMaybe **domain shift**(train: noisy sound -> test: clean sound) is concerned. B3 has learned about **train domain features**(noisy sound). On the other hand, B0 has less representational ability than B3. Therefore B0 has learned the **common features** of the train and test domain with high resolution. Without domain shift, B3 would have also shown good results with high resolution.",
      "votes": null
    },
    {
      "id": "1207777",
      "postDate": "02/18/2021 03:18:49",
      "content": "<p>Congrats on results and thanks for the writeup <a href=\"https://www.kaggle.com/shinmurashinmura\" target=\"_blank\">@shinmurashinmura</a> </p>",
      "rawMarkdown": "Congrats on results and thanks for the writeup @shinmurashinmura",
      "votes": null
    },
    {
      "id": "1208018",
      "postDate": "02/18/2021 05:52:26",
      "content": "<p>Interesting that the resolution had that large of an impact. Do you think that it was actually a more detail representation to learn from or do you think maybe the imagenet weights are more tuned to have square images and features and not good when the features are quite small and spread out far across the time axis?</p>",
      "rawMarkdown": "Interesting that the resolution had that large of an impact. Do you think that it was actually a more detail representation to learn from or do you think maybe the imagenet weights are more tuned to have square images and features and not good when the features are quite small and spread out far across the time axis?",
      "votes": null
    },
    {
      "id": "1208257",
      "postDate": "02/18/2021 08:18:20",
      "content": "<p>Maybe small feature is more critical. In the paper[2](state-of-the-art in audioset), EfficientNet (pretrained by ImageNet) is used. But resolution is not square.</p>",
      "rawMarkdown": "Maybe small feature is more critical. In the paper[2](state-of-the-art in audioset), EfficientNet (pretrained by ImageNet) is used. But resolution is not square.",
      "votes": null
    },
    {
      "id": "1208394",
      "postDate": "02/18/2021 09:29:50",
      "content": "<p>Congratz ! </p>\n<p>Your sharings durign the competition helped a lot of people, so thanks a lot for that. </p>",
      "rawMarkdown": "Congratz ! \n\nYour sharings durign the competition helped a lot of people, so thanks a lot for that.",
      "votes": null
    },
    {
      "id": "1215006",
      "postDate": "02/23/2021 09:03:42",
      "content": "<p>Thanks <a href=\"https://www.kaggle.com/shinmurashinmura\" target=\"_blank\">@shinmurashinmura</a> for the write up. I would like to ask when you use  a large number of mel bands (850), but n_fft =1024, would you end up with some empty bands?</p>",
      "rawMarkdown": "Thanks @shinmurashinmura for the write up. I would like to ask when you use  a large number of mel bands (850), but n_fft =1024, would you end up with some empty bands?",
      "votes": null
    },
    {
      "id": "1216353",
      "postDate": "02/24/2021 08:16:43",
      "content": "<p>Thank you for your comment. </p>\n<blockquote>\n  <p>would you end up with some empty bands?</p>\n</blockquote>\n<p>Yes. In 850 mel bands, empty bands is exist. But 850 is better than 512 or 751.<br>\nIt is strange. I cannot understand this reason.</p>",
      "rawMarkdown": "Thank you for your comment. \n\n> would you end up with some empty bands?\n\nYes. In 850 mel bands, empty bands is exist. But 850 is better than 512 or 751.\nIt is strange. I cannot understand this reason.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1207777,
      "author_name": "duykhanh99",
      "author_url": "",
      "post_date": "02/18/2021 03:18:49",
      "content": "<p>Congrats on results and thanks for the writeup <a href=\"https://www.kaggle.com/shinmurashinmura\" target=\"_blank\">@shinmurashinmura</a> </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1208018,
      "author_name": "ryches",
      "author_url": "",
      "post_date": "02/18/2021 05:52:26",
      "content": "<p>Interesting that the resolution had that large of an impact. Do you think that it was actually a more detail representation to learn from or do you think maybe the imagenet weights are more tuned to have square images and features and not good when the features are quite small and spread out far across the time axis?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1208257,
          "author_name": "shinmurashinmura",
          "author_url": "",
          "post_date": "02/18/2021 08:18:20",
          "content": "<p>Maybe small feature is more critical. In the paper[2](state-of-the-art in audioset), EfficientNet (pretrained by ImageNet) is used. But resolution is not square.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1208394,
      "author_name": "theoviel",
      "author_url": "",
      "post_date": "02/18/2021 09:29:50",
      "content": "<p>Congratz ! </p>\n<p>Your sharings durign the competition helped a lot of people, so thanks a lot for that. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1215006,
      "author_name": "thomeou",
      "author_url": "",
      "post_date": "02/23/2021 09:03:42",
      "content": "<p>Thanks <a href=\"https://www.kaggle.com/shinmurashinmura\" target=\"_blank\">@shinmurashinmura</a> for the write up. I would like to ask when you use  a large number of mel bands (850), but n_fft =1024, would you end up with some empty bands?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1216353,
          "author_name": "shinmurashinmura",
          "author_url": "",
          "post_date": "02/24/2021 08:16:43",
          "content": "<p>Thank you for your comment. </p>\n<blockquote>\n  <p>would you end up with some empty bands?</p>\n</blockquote>\n<p>Yes. In 850 mel bands, empty bands is exist. But 850 is better than 512 or 751.<br>\nIt is strange. I cannot understand this reason.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1207769": "First of all, many thanks to Kaggler.\n I got a lot of ideas from Kaggler in the discussion. And it was fun.\n\nMy solution summary is below.\n+ The high resolution of spectrogram\n+ Post-processing with moving average\n+ Teacher-student model for missing labels\n\n#1st stage: SED\nI started from basic experiment with [SED](https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/211007). Why SED? Because I think SED is strong in the multi label task. \n\nI used log mel-spectrogram as the input of SED. Basic experiment involves data augmentation (Gaussian noise, SpecAugment and MixUP), backbone model choice and adjusting resolution of log mel-spectrogram. As a result, below condition was the best for me.\n+ No noise injection\n+ MixUp\n+ The best model architecture is EfficientNet\n+ **The higher the resolution** of log mel-spectrogram, the better the result.\n\n### The resolution\nThe most important one is the resolution. Recently, in Kaggle computer vision solution, the higher the resolution of the image, the better the result. In spectrogram, the same phenomenon may happen. In mel-spectrogram, the resolution can be changed by adjusting \"hop_size\" and \"mel_bins\". Following result is changing the resolution with ResNest50(single model).\n\n|Resolution(Width-Height)|public LB|\n|---|---|\n|1501-64(PANNs default)|0.692\n|3001-128|0.805\n|6001-64|0.725\n|1501-256|0.761\n|751-512|**0.823**\n|1501-512|0.821\n\nThe resolution was critical! According to experimental result, good resolution was \"high\" and similar to the square. 751-512 looks good. As a result, I chose 858-850. This configuration is as follows.\n\n```\nmodel_config = {\n    \"sample_rate\": 48000,\n    \"window_size\": 1024,\n    \"hop_size\": 560,\n    \"mel_bins\": 850,\n    \"fmin\": 50,\n    \"fmax\": 14000,\n    \"classes_num\": 24\n}\n```\n\n### Post-processing\nI used [Framewise output](https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/211007) for submission. It contains time and classes information. But there is a lot of false positive information in framewise output. Because they are not processing by a long time information. Therefore **a short event of framewise output should be deleted.** I prepared post-processing for framewise output. It is a moving average.\n\n![image.png](https://qiita-image-store.s3.ap-northeast-1.amazonaws.com/0/264781/2d0ad5b8-f046-64da-a98a-96f0ae79ba2e.png)\n\nBy taking a moving average in the time direction for each class, we can delete short events. This idea is based on [the paper](http://dcase.community/documents/challenge2020/technical_reports/DCASE2020_Chan_6.pdf)[1]. The sample code is as follows.\n```\ndef post_processing(data): # data.shape = (24, 600) # (classes, time)\n    result = []\n    for i in range(len(data)):\n        result.append(cv2.blur(data[i],(1,31)))\n    return result\n```\n\nI improved LB by using moving average. The following result is comparing post-processing with EfficientNetB3(single model).\n\n||public LB|\n|---|---|\n|w/o post-processing|0.785|\n|w/ post-processing|**0.840**|\n \n###Summary\n+ MixUp(alpha=0.1)\n+ Epoch 30\n+ Adam(lr=0.001) + CosineAnnealing(T=10)\n+ Batchsize 6\n+ Use only tp label\n+ Get random tp clip 10 sec\n+ The resolution of log mel-spectrogram: 858-850\n+ Loss function: BCE\n+ [Weak label training](https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/211007)\n+ Post-processing: moving average\n\nThen I got **0.916 public LB** with EfficientNetB0(5-folds average ensemble).\n\n#2nd stage: missing labels and the ensemble\nI reported [discovering missing labels and re-train](https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/208830#1139171). And It didn't work. After that, I thought about missing labels again. My answer is that the model is not correct for discovering missing labels. There are a lot of missing labels around tp. Therefore the model is not correct. \n\n![train.png](https://qiita-image-store.s3.ap-northeast-1.amazonaws.com/0/264781/0e6009db-c1cf-7297-3ac4-df92dd3420c5.png)\n\nTo solve this issue, I used teacher-student model. \n\n![geretation.png](https://qiita-image-store.s3.ap-northeast-1.amazonaws.com/0/264781/79d8e1c8-6814-68c6-0a9d-d41d830fcdfb.png)\n\n1st generation is similar to 1st stage. I gradually increased model prediction ratio. By using teacher-student model, I could discover missing labels. Specially, in strong label training, teacher-student model was effective. Following result is teacher-student model score with EfficientNetB0.\n\n![image.png](https://qiita-image-store.s3.ap-northeast-1.amazonaws.com/0/264781/ddded273-c2b1-92c1-3272-1e19377a5b77.png)\n\n\"MixUp rate\" is probabilistic MixUp. This method is based on [the paper](https://arxiv.org/abs/2102.01243)[2]. \n\nFinally, I made the ensemble of 1st stage model and 2nd stage model. Ensemble procedure is simple average. Then I got **0.924 public LB.**\n\n#References\n[1] Teck Kai Chan, Cheng Siong Chin1 and Ye Li, \"SEMI-SUPERVISED NMF-CNN FOR SOUND EVENT DETECTION\".\n[2] Yuan Gong, Yu-An Chung, and James Glass, \"PSLA: Improving Audio Event Classification with\nPretraining, Sampling, Labeling, and Aggregation\".\n\n#Appendix: the resolution and EfficientNet\nFinally, I show interesting result. It is relationship between EfficientNet and the resolution. The following result is public LB(5-folds average ensemble).\n\n|Resolution(W-H)|751-512|751-751|858-850|\n|---|---|---|---|\n|EfficientNetB0|0.893|0.904|**0.916**|\n|EfficientNetB3|**0.913**|0.912|0.900|\n\nIn B0, the higher resolution, the better result. But B3 was vice versa. Usually, the larger EfficientNet, the better at [high resolution](https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/154683) it is. But the above is reverse. Why?\n\nMaybe **domain shift**(train: noisy sound -> test: clean sound) is concerned. B3 has learned about **train domain features**(noisy sound). On the other hand, B0 has less representational ability than B3. Therefore B0 has learned the **common features** of the train and test domain with high resolution. Without domain shift, B3 would have also shown good results with high resolution.",
    "1207777": "Congrats on results and thanks for the writeup @shinmurashinmura",
    "1208018": "Interesting that the resolution had that large of an impact. Do you think that it was actually a more detail representation to learn from or do you think maybe the imagenet weights are more tuned to have square images and features and not good when the features are quite small and spread out far across the time axis?",
    "1208257": "Maybe small feature is more critical. In the paper[2](state-of-the-art in audioset), EfficientNet (pretrained by ImageNet) is used. But resolution is not square.",
    "1208394": "Congratz ! \n\nYour sharings durign the competition helped a lot of people, so thanks a lot for that.",
    "1215006": "Thanks @shinmurashinmura for the write up. I would like to ask when you use  a large number of mel bands (850), but n_fft =1024, would you end up with some empty bands?",
    "1216353": "Thank you for your comment. \n\n> would you end up with some empty bands?\n\nYes. In 850 mel bands, empty bands is exist. But 850 is better than 512 or 751.\nIt is strange. I cannot understand this reason."
  },
  "source": "meta"
}