{
  "id": 97812,
  "title": "7th place solution with commentary kernel",
  "url": "/competitions/freesound-audio-tagging-2019/writeups/kaggler-ja-shirogane-7th-place-solution-with-comme",
  "author_name": "",
  "post_date": "2019-07-02T02:04:51.740Z",
  "votes": 11,
  "comment_count": 2,
  "views": 0,
  "content": "<h1>Freesound 7th place solution</h1>\n\n<p>Thank you for all the competitors, kaggle teams, and the host of this competition!\nWe enjoyed a lot during the competition and learned many things.  </p>\n\n<p>We especially want to thank <a href=\"/daisukelab\">@daisukelab</a> for his clear instruction with great kernels and datasets, <a href=\"/mhiro2\">@mhiro2</a> for sharing excellent training framework, and <a href=\"/sailorwei\">@sailorwei</a> for showing his Inception v3 model in his public kernel.</p>\n\n<p>The points where the solution of our team seems to be different from other teams are as follows.</p>\n\n<p>keypoint : Data Augmentation, Strength Adaptive Crop,Custom CNN, RandomResizedCrop</p>\n\n<p>The detailed explanation is in the following kernel, so please read it.  </p>\n\n<p><a href=\"https://www.kaggle.com/hidehisaarai1213/freesound-7th-place-solution\">https://www.kaggle.com/hidehisaarai1213/freesound-7th-place-solution</a>  </p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/564097/13702/Shirogane_solution.png\" alt=\"pipeline\">  </p>\n\n<h2>Data Augmentation</h2>\n\n<p>We created 7 augmented training dataset with <a href=\"http://sox.sourceforge.net/\">sox</a>.</p>\n\n<ul>\n<li>fade</li>\n<li>pitch * 2</li>\n<li>reverb</li>\n<li>treble &amp; bass * 2</li>\n<li>equalize </li>\n</ul>\n\n<p>We trained a togal of 4970 * 8 samples without leaks.</p>\n\n<h2>Crop policy</h2>\n\n<p>We use random crop, because we use fixed image size(128 * 128).\nRandom crop got a little better cv than the first 2 seconds cropped.\nAt first, it wa cropped uniformly as mhiro's kernel.</p>\n\n<h3>Strength Adaptive Crop</h3>\n\n<p>Many sound crip has important information at the first few seconds.\nSome sample has it in the middle of crip. However, due to the nature of recording, it is rare to have important information at th end of sounds.\nThe score drops about 0.03~0.04 when learning only the last few seconds.</p>\n\n<p>Then, We introduce <strong>Strength Adaptive Crop</strong>. \nWe tried to crop the place where the total of db is high preferentially. </p>\n\n<p>This method is very effective because most samples contain important information in places where the sound is loud.　　</p>\n\n<p>CV 0.01 up <br>\nLB 0.004~0.005 up  </p>\n\n<p>Detailed code is in <a href=\"https://www.kaggle.com/hidehisaarai1213/freesound-7th-place-solution\">this kernel</a>.</p>\n\n<h2>model structure</h2>\n\n<ul>\n<li>InceptionV3 3ch  </li>\n<li>InceptionV3 1ch  </li>\n<li>CustomCNN  </li>\n</ul>\n\n<p>CustomCNN is carefully designed to the characteristics of the sound. The details are in  <a href=\"https://www.kaggle.com/hidehisaarai1213/freesound-7th-place-solution\">this kernel</a></p>\n\n<h2>Augmentation in batch</h2>\n\n<ul>\n<li>Random erasing or Coarse Dropout</li>\n<li>Horizontal Flip</li>\n<li>mixup</li>\n<li>Random Resized Crop (only InceptionV3) </li>\n</ul>\n\n<h2>Training strategy</h2>\n\n<h3>TTA for validation</h3>\n\n<p>When RandomResizedCrop is used, val score fluctuate,so if val tta is not used, an appropriate epoch can not be selected. So, we used tta for validation to ensure that validation can be properly evaluated.  </p>\n\n<h3>Stage 1 : pretrain with noisy data(warm-up)</h3>\n\n<p>We used noisy data to 'pre-train' our model.\nLB : about 0.01 up</p>\n\n<h3>Stage 2 : Train with curated data 1</h3>\n\n<p>We used curated data to 'finetune' the models, which were 'pre-trained' with noisy data.</p>\n\n<h3>Stage 3 : Train with curated data 2(Inception only)</h3>\n\n<p>We used stage 2 best weight to stage 3 training without random resized crop. We don't know why, but lwlrap goes up without random resized crop in Inception model.</p>\n\n<h3>score</h3>\n\n<p>Accurate single model public score not measured.\n|  | public|private |\n|-----------|------------|------------|\n|Inception 3ch|0.724 over|0.73865|\n|Inception 1ch|??|0.73917|\n|CustomCNN|0.720 over|0.73103|</p>\n\n<h2>Ensemble</h2>\n\n<p>(Inception 3ch + Inception 1ch + CustomCNN) / 3  </p>\n\n<p>private score:0.75302</p>",
  "messages": [
    {
      "id": "564097",
      "postDate": "06/29/2019 01:23:16",
      "content": "<h1>Freesound 7th place solution</h1>\n\n<p>Thank you for all the competitors, kaggle teams, and the host of this competition!\nWe enjoyed a lot during the competition and learned many things.  </p>\n\n<p>We especially want to thank <a href=\"/daisukelab\">@daisukelab</a> for his clear instruction with great kernels and datasets, <a href=\"/mhiro2\">@mhiro2</a> for sharing excellent training framework, and <a href=\"/sailorwei\">@sailorwei</a> for showing his Inception v3 model in his public kernel.</p>\n\n<p>The points where the solution of our team seems to be different from other teams are as follows.</p>\n\n<p>keypoint : Data Augmentation, Strength Adaptive Crop,Custom CNN, RandomResizedCrop</p>\n\n<p>The detailed explanation is in the following kernel, so please read it.  </p>\n\n<p><a href=\"https://www.kaggle.com/hidehisaarai1213/freesound-7th-place-solution\">https://www.kaggle.com/hidehisaarai1213/freesound-7th-place-solution</a>  </p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/564097/13702/Shirogane_solution.png\" alt=\"pipeline\">  </p>\n\n<h2>Data Augmentation</h2>\n\n<p>We created 7 augmented training dataset with <a href=\"http://sox.sourceforge.net/\">sox</a>.</p>\n\n<ul>\n<li>fade</li>\n<li>pitch * 2</li>\n<li>reverb</li>\n<li>treble &amp; bass * 2</li>\n<li>equalize </li>\n</ul>\n\n<p>We trained a togal of 4970 * 8 samples without leaks.</p>\n\n<h2>Crop policy</h2>\n\n<p>We use random crop, because we use fixed image size(128 * 128).\nRandom crop got a little better cv than the first 2 seconds cropped.\nAt first, it wa cropped uniformly as mhiro's kernel.</p>\n\n<h3>Strength Adaptive Crop</h3>\n\n<p>Many sound crip has important information at the first few seconds.\nSome sample has it in the middle of crip. However, due to the nature of recording, it is rare to have important information at th end of sounds.\nThe score drops about 0.03~0.04 when learning only the last few seconds.</p>\n\n<p>Then, We introduce <strong>Strength Adaptive Crop</strong>. \nWe tried to crop the place where the total of db is high preferentially. </p>\n\n<p>This method is very effective because most samples contain important information in places where the sound is loud.　　</p>\n\n<p>CV 0.01 up <br>\nLB 0.004~0.005 up  </p>\n\n<p>Detailed code is in <a href=\"https://www.kaggle.com/hidehisaarai1213/freesound-7th-place-solution\">this kernel</a>.</p>\n\n<h2>model structure</h2>\n\n<ul>\n<li>InceptionV3 3ch  </li>\n<li>InceptionV3 1ch  </li>\n<li>CustomCNN  </li>\n</ul>\n\n<p>CustomCNN is carefully designed to the characteristics of the sound. The details are in  <a href=\"https://www.kaggle.com/hidehisaarai1213/freesound-7th-place-solution\">this kernel</a></p>\n\n<h2>Augmentation in batch</h2>\n\n<ul>\n<li>Random erasing or Coarse Dropout</li>\n<li>Horizontal Flip</li>\n<li>mixup</li>\n<li>Random Resized Crop (only InceptionV3) </li>\n</ul>\n\n<h2>Training strategy</h2>\n\n<h3>TTA for validation</h3>\n\n<p>When RandomResizedCrop is used, val score fluctuate,so if val tta is not used, an appropriate epoch can not be selected. So, we used tta for validation to ensure that validation can be properly evaluated.  </p>\n\n<h3>Stage 1 : pretrain with noisy data(warm-up)</h3>\n\n<p>We used noisy data to 'pre-train' our model.\nLB : about 0.01 up</p>\n\n<h3>Stage 2 : Train with curated data 1</h3>\n\n<p>We used curated data to 'finetune' the models, which were 'pre-trained' with noisy data.</p>\n\n<h3>Stage 3 : Train with curated data 2(Inception only)</h3>\n\n<p>We used stage 2 best weight to stage 3 training without random resized crop. We don't know why, but lwlrap goes up without random resized crop in Inception model.</p>\n\n<h3>score</h3>\n\n<p>Accurate single model public score not measured.\n|  | public|private |\n|-----------|------------|------------|\n|Inception 3ch|0.724 over|0.73865|\n|Inception 1ch|??|0.73917|\n|CustomCNN|0.720 over|0.73103|</p>\n\n<h2>Ensemble</h2>\n\n<p>(Inception 3ch + Inception 1ch + CustomCNN) / 3  </p>\n\n<p>private score:0.75302</p>",
      "rawMarkdown": "# Freesound 7th place solution\n\nThank you for all the competitors, kaggle teams, and the host of this competition!\nWe enjoyed a lot during the competition and learned many things.  \n\nWe especially want to thank @daisukelab for his clear instruction with great kernels and datasets, @mhiro2 for sharing excellent training framework, and @sailorwei for showing his Inception v3 model in his public kernel.\n\nThe points where the solution of our team seems to be different from other teams are as follows.\n\nkeypoint : Data Augmentation, Strength Adaptive Crop,Custom CNN, RandomResizedCrop\n\nThe detailed explanation is in the following kernel, so please read it.  \n  \nhttps://www.kaggle.com/hidehisaarai1213/freesound-7th-place-solution  \n  \n  \n![pipeline](https://storage.googleapis.com/kaggle-forum-message-attachments/564097/13702/Shirogane_solution.png)  \n  \n  \n  \n\n## Data Augmentation\nWe created 7 augmented training dataset with [sox](http://sox.sourceforge.net/).\n\n- fade\n- pitch * 2\n- reverb\n- treble &amp; bass * 2\n- equalize \n\nWe trained a togal of 4970 * 8 samples without leaks.\n\n## Crop policy\nWe use random crop, because we use fixed image size(128 * 128).\nRandom crop got a little better cv than the first 2 seconds cropped.\nAt first, it wa cropped uniformly as mhiro's kernel.\n\n### Strength Adaptive Crop\nMany sound crip has important information at the first few seconds.\nSome sample has it in the middle of crip. However, due to the nature of recording, it is rare to have important information at th end of sounds.\nThe score drops about 0.03~0.04 when learning only the last few seconds.\n\nThen, We introduce **Strength Adaptive Crop**. \nWe tried to crop the place where the total of db is high preferentially. \n\nThis method is very effective because most samples contain important information in places where the sound is loud.　　\n\nCV 0.01 up  \nLB 0.004~0.005 up  \n\nDetailed code is in [this kernel](https://www.kaggle.com/hidehisaarai1213/freesound-7th-place-solution).\n\n## model structure\n- InceptionV3 3ch  \n- InceptionV3 1ch  \n- CustomCNN  \n\n\nCustomCNN is carefully designed to the characteristics of the sound. The details are in  [this kernel](https://www.kaggle.com/hidehisaarai1213/freesound-7th-place-solution)\n\n## Augmentation in batch\n- Random erasing or Coarse Dropout\n- Horizontal Flip\n- mixup\n- Random Resized Crop (only InceptionV3) \n\n## Training strategy\n### TTA for validation\nWhen RandomResizedCrop is used, val score fluctuate,so if val tta is not used, an appropriate epoch can not be selected. So, we used tta for validation to ensure that validation can be properly evaluated.  \n\n### Stage 1 : pretrain with noisy data(warm-up)\nWe used noisy data to 'pre-train' our model.\nLB : about 0.01 up\n\n### Stage 2 : Train with curated data 1\nWe used curated data to 'finetune' the models, which were 'pre-trained' with noisy data.\n\n### Stage 3 : Train with curated data 2(Inception only)  \nWe used stage 2 best weight to stage 3 training without random resized crop. We don't know why, but lwlrap goes up without random resized crop in Inception model.\n\n### score\n\nAccurate single model public score not measured.\n|  | public|private |\n|-----------|------------|------------|\n|Inception 3ch|0.724 over|0.73865|\n|Inception 1ch|??|0.73917|\n|CustomCNN|0.720 over|0.73103|\n\n## Ensemble\n(Inception 3ch + Inception 1ch + CustomCNN) / 3  \n  \nprivate score:0.75302",
      "votes": null
    },
    {
      "id": "564102",
      "postDate": "06/29/2019 01:30:52",
      "content": "<p>Resizing makes performance worse in our case. Resizing sometimes makes a female voice like a male voice and vice versa. It is plausible for me that resizing only works as warm-up.</p>",
      "rawMarkdown": "Resizing makes performance worse in our case. Resizing sometimes makes a female voice like a male voice and vice versa. It is plausible for me that resizing only works as warm-up.",
      "votes": null
    },
    {
      "id": "564533",
      "postDate": "06/29/2019 15:07:55",
      "content": "<p>We did not notice that effect. If we were carefully modeling with the meaning of resizing in mind, LB score might have been a bit better. Thank you.</p>",
      "rawMarkdown": "We did not notice that effect. If we were carefully modeling with the meaning of resizing in mind, LB score might have been a bit better. Thank you.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 564102,
      "author_name": "osciiart",
      "author_url": "",
      "post_date": "06/29/2019 01:30:52",
      "content": "<p>Resizing makes performance worse in our case. Resizing sometimes makes a female voice like a male voice and vice versa. It is plausible for me that resizing only works as warm-up.</p>",
      "votes": null,
      "replies": [
        {
          "id": 564533,
          "author_name": "d1348k",
          "author_url": "",
          "post_date": "06/29/2019 15:07:55",
          "content": "<p>We did not notice that effect. If we were carefully modeling with the meaning of resizing in mind, LB score might have been a bit better. Thank you.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "564097": "# Freesound 7th place solution\n\nThank you for all the competitors, kaggle teams, and the host of this competition!\nWe enjoyed a lot during the competition and learned many things.  \n\nWe especially want to thank @daisukelab for his clear instruction with great kernels and datasets, @mhiro2 for sharing excellent training framework, and @sailorwei for showing his Inception v3 model in his public kernel.\n\nThe points where the solution of our team seems to be different from other teams are as follows.\n\nkeypoint : Data Augmentation, Strength Adaptive Crop,Custom CNN, RandomResizedCrop\n\nThe detailed explanation is in the following kernel, so please read it.  \n  \nhttps://www.kaggle.com/hidehisaarai1213/freesound-7th-place-solution  \n  \n  \n![pipeline](https://storage.googleapis.com/kaggle-forum-message-attachments/564097/13702/Shirogane_solution.png)  \n  \n  \n  \n\n## Data Augmentation\nWe created 7 augmented training dataset with [sox](http://sox.sourceforge.net/).\n\n- fade\n- pitch * 2\n- reverb\n- treble &amp; bass * 2\n- equalize \n\nWe trained a togal of 4970 * 8 samples without leaks.\n\n## Crop policy\nWe use random crop, because we use fixed image size(128 * 128).\nRandom crop got a little better cv than the first 2 seconds cropped.\nAt first, it wa cropped uniformly as mhiro's kernel.\n\n### Strength Adaptive Crop\nMany sound crip has important information at the first few seconds.\nSome sample has it in the middle of crip. However, due to the nature of recording, it is rare to have important information at th end of sounds.\nThe score drops about 0.03~0.04 when learning only the last few seconds.\n\nThen, We introduce **Strength Adaptive Crop**. \nWe tried to crop the place where the total of db is high preferentially. \n\nThis method is very effective because most samples contain important information in places where the sound is loud.　　\n\nCV 0.01 up  \nLB 0.004~0.005 up  \n\nDetailed code is in [this kernel](https://www.kaggle.com/hidehisaarai1213/freesound-7th-place-solution).\n\n## model structure\n- InceptionV3 3ch  \n- InceptionV3 1ch  \n- CustomCNN  \n\n\nCustomCNN is carefully designed to the characteristics of the sound. The details are in  [this kernel](https://www.kaggle.com/hidehisaarai1213/freesound-7th-place-solution)\n\n## Augmentation in batch\n- Random erasing or Coarse Dropout\n- Horizontal Flip\n- mixup\n- Random Resized Crop (only InceptionV3) \n\n## Training strategy\n### TTA for validation\nWhen RandomResizedCrop is used, val score fluctuate,so if val tta is not used, an appropriate epoch can not be selected. So, we used tta for validation to ensure that validation can be properly evaluated.  \n\n### Stage 1 : pretrain with noisy data(warm-up)\nWe used noisy data to 'pre-train' our model.\nLB : about 0.01 up\n\n### Stage 2 : Train with curated data 1\nWe used curated data to 'finetune' the models, which were 'pre-trained' with noisy data.\n\n### Stage 3 : Train with curated data 2(Inception only)  \nWe used stage 2 best weight to stage 3 training without random resized crop. We don't know why, but lwlrap goes up without random resized crop in Inception model.\n\n### score\n\nAccurate single model public score not measured.\n|  | public|private |\n|-----------|------------|------------|\n|Inception 3ch|0.724 over|0.73865|\n|Inception 1ch|??|0.73917|\n|CustomCNN|0.720 over|0.73103|\n\n## Ensemble\n(Inception 3ch + Inception 1ch + CustomCNN) / 3  \n  \nprivate score:0.75302",
    "564102": "Resizing makes performance worse in our case. Resizing sometimes makes a female voice like a male voice and vice versa. It is plausible for me that resizing only works as warm-up.",
    "564533": "We did not notice that effect. If we were carefully modeling with the meaning of resizing in mind, LB score might have been a bit better. Thank you."
  },
  "source": "meta"
}