{
  "id": 243514,
  "title": "87th solution",
  "url": "/competitions/birdclef-2021/writeups/omastar-87th-solution",
  "author_name": "",
  "post_date": "2021-06-02T23:29:28.178654800Z",
  "votes": 5,
  "comment_count": 3,
  "views": 0,
  "content": "<p>I appreciate the organizers and participants!<br>\nI am looking forward to seeing other solutions!</p>\n<p>I am not good at English, but I hope this helps anyone.</p>\n<p>This is my 87th(LB 90th) solution.</p>\n<h1>Solution</h1>\n<p>I referred <a href=\"https://www.kaggle.com/kkiller\" target=\"_blank\">@kkiller</a> 's nice notebook <a href=\"https://www.kaggle.com/kneroma/clean-fast-simple-bird-identifier-inference\" target=\"_blank\">https://www.kaggle.com/kneroma/clean-fast-simple-bird-identifier-inference</a></p>\n<p>My final solution is</p>\n<ol>\n<li>Ensemble 65models (13 types x 5 folds)</li>\n<li>Audio clip length 6 sec in inference time</li>\n<li>DataAugmentation in training time</li>\n</ol>\n<h2>1. Ensemble 65models</h2>\n<p>I noticed increasing model type is effective local score and LB for first a few submissions, so I trained various CNN types and selected nice models seeing the result of my local score in inference time.</p>\n<p>Model list is below</p>\n<ul>\n<li>EfficientNet-v2 large</li>\n<li>EfficientNet-v2 medium</li>\n<li>EfficientNet-v2 b3</li>\n<li>EfficientNet-b6</li>\n<li>EfficientNet-b4</li>\n<li>EfficientNet-b2</li>\n<li>EfficientNet-b0</li>\n<li>ResNeSt50d x3 (2 random seeds and train methods)</li>\n<li>ResNet200d</li>\n<li>DenseNet201</li>\n<li>DenseNet121</li>\n</ul>\n<p>Other than these models, I tried NFNet and other ResNet types, but I have no time for training convergence.</p>\n<p>Final prediction labels are selected by threshold=0.2.</p>\n<h2>2. Audio clip , 3.Data Augmentation</h2>\n<p>I decided that audio clip length is 6sec in inference for maximizing local score. ( I trained some models in 7sec crop and the others in holizontal random crops(5~7sec), but the score is maximized in 6sec crops. )</p>\n<p>Moreover, I used not only one hot label but also second_labels for teaching labels. </p>\n<ul>\n<li>primary_label = 0.9</li>\n<li>secondary_labels = 0.5</li>\n<li>other = 0.01</li>\n</ul>\n<p>And, I used BCEwithLogits loss.</p>\n<h2>Training settings</h2>\n<ul>\n<li>StratifiedKFold 5fold</li>\n<li>RandomSeed = 42, 46</li>\n<li>LearningRate = 0.0008</li>\n<li>CosineAnnealingLR</li>\n<li>Mixup</li>\n</ul>\n<h1>not improving</h1>\n<ol>\n<li>clip length unsemble in inference time</li>\n<li>typical data augmentation (parital dropout, random cutout, gaussian noise, random changing brightness and contrast)</li>\n<li>threshold based on every bird from appearance frequencies in training data</li>\n</ol>\n<p>At the end, I want to thank everyone in this competition and kaggle again!</p>",
  "messages": [
    {
      "id": "1333633",
      "postDate": "06/02/2021 23:29:28",
      "content": "<p>I appreciate the organizers and participants!<br>\nI am looking forward to seeing other solutions!</p>\n<p>I am not good at English, but I hope this helps anyone.</p>\n<p>This is my 87th(LB 90th) solution.</p>\n<h1>Solution</h1>\n<p>I referred <a href=\"https://www.kaggle.com/kkiller\" target=\"_blank\">@kkiller</a> 's nice notebook <a href=\"https://www.kaggle.com/kneroma/clean-fast-simple-bird-identifier-inference\" target=\"_blank\">https://www.kaggle.com/kneroma/clean-fast-simple-bird-identifier-inference</a></p>\n<p>My final solution is</p>\n<ol>\n<li>Ensemble 65models (13 types x 5 folds)</li>\n<li>Audio clip length 6 sec in inference time</li>\n<li>DataAugmentation in training time</li>\n</ol>\n<h2>1. Ensemble 65models</h2>\n<p>I noticed increasing model type is effective local score and LB for first a few submissions, so I trained various CNN types and selected nice models seeing the result of my local score in inference time.</p>\n<p>Model list is below</p>\n<ul>\n<li>EfficientNet-v2 large</li>\n<li>EfficientNet-v2 medium</li>\n<li>EfficientNet-v2 b3</li>\n<li>EfficientNet-b6</li>\n<li>EfficientNet-b4</li>\n<li>EfficientNet-b2</li>\n<li>EfficientNet-b0</li>\n<li>ResNeSt50d x3 (2 random seeds and train methods)</li>\n<li>ResNet200d</li>\n<li>DenseNet201</li>\n<li>DenseNet121</li>\n</ul>\n<p>Other than these models, I tried NFNet and other ResNet types, but I have no time for training convergence.</p>\n<p>Final prediction labels are selected by threshold=0.2.</p>\n<h2>2. Audio clip , 3.Data Augmentation</h2>\n<p>I decided that audio clip length is 6sec in inference for maximizing local score. ( I trained some models in 7sec crop and the others in holizontal random crops(5~7sec), but the score is maximized in 6sec crops. )</p>\n<p>Moreover, I used not only one hot label but also second_labels for teaching labels. </p>\n<ul>\n<li>primary_label = 0.9</li>\n<li>secondary_labels = 0.5</li>\n<li>other = 0.01</li>\n</ul>\n<p>And, I used BCEwithLogits loss.</p>\n<h2>Training settings</h2>\n<ul>\n<li>StratifiedKFold 5fold</li>\n<li>RandomSeed = 42, 46</li>\n<li>LearningRate = 0.0008</li>\n<li>CosineAnnealingLR</li>\n<li>Mixup</li>\n</ul>\n<h1>not improving</h1>\n<ol>\n<li>clip length unsemble in inference time</li>\n<li>typical data augmentation (parital dropout, random cutout, gaussian noise, random changing brightness and contrast)</li>\n<li>threshold based on every bird from appearance frequencies in training data</li>\n</ol>\n<p>At the end, I want to thank everyone in this competition and kaggle again!</p>",
      "rawMarkdown": "I appreciate the organizers and participants!\nI am looking forward to seeing other solutions!\n\nI am not good at English, but I hope this helps anyone.\n\nThis is my 87th(LB 90th) solution.\n\n# Solution\n\nI referred @kkiller 's nice notebook https://www.kaggle.com/kneroma/clean-fast-simple-bird-identifier-inference\n\nMy final solution is\n\n1. Ensemble 65models (13 types x 5 folds)\n2. Audio clip length 6 sec in inference time\n3. DataAugmentation in training time\n\n## 1. Ensemble 65models\n\nI noticed increasing model type is effective local score and LB for first a few submissions, so I trained various CNN types and selected nice models seeing the result of my local score in inference time.\n\nModel list is below\n\n- EfficientNet-v2 large\n- EfficientNet-v2 medium\n- EfficientNet-v2 b3\n- EfficientNet-b6\n- EfficientNet-b4\n- EfficientNet-b2\n- EfficientNet-b0\n- ResNeSt50d x3 (2 random seeds and train methods)\n- ResNet200d\n- DenseNet201\n- DenseNet121\n\nOther than these models, I tried NFNet and other ResNet types, but I have no time for training convergence.\n\nFinal prediction labels are selected by threshold=0.2.\n\n\n## 2. Audio clip , 3.Data Augmentation\n\nI decided that audio clip length is 6sec in inference for maximizing local score. ( I trained some models in 7sec crop and the others in holizontal random crops(5~7sec), but the score is maximized in 6sec crops. )\n\nMoreover, I used not only one hot label but also second_labels for teaching labels. \n\n- primary_label = 0.9\n- secondary_labels = 0.5\n- other = 0.01\n\nAnd, I used BCEwithLogits loss.\n\n## Training settings\n\n- StratifiedKFold 5fold\n- RandomSeed = 42, 46\n- LearningRate = 0.0008\n- CosineAnnealingLR\n- Mixup\n\n\n# not improving\n\n1. clip length unsemble in inference time\n2. typical data augmentation (parital dropout, random cutout, gaussian noise, random changing brightness and contrast)\n3. threshold based on every bird from appearance frequencies in training data\n\nAt the end, I want to thank everyone in this competition and kaggle again!",
      "votes": null
    },
    {
      "id": "1336039",
      "postDate": "06/04/2021 15:43:55",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/yoshito\" target=\"_blank\">@yoshito</a> thanks for sharing this. I'm new here and want to understand what does these numbers refer to in this statement:</p>\n<p>Moreover, I used not only one hot label but also second_labels for teaching labels.</p>\n<p>primary_label = 0.9<br>\nsecondary_labels = 0.5<br>\nother = 0.01</p>\n<p>In addition to above, could you also please tell me if by \"partial dropout\"  you mean partial drop of time and frequency sections in the melspecs?</p>",
      "rawMarkdown": "Hi @yoshito thanks for sharing this. I'm new here and want to understand what does these numbers refer to in this statement:\n\nMoreover, I used not only one hot label but also second_labels for teaching labels.\n\nprimary_label = 0.9\nsecondary_labels = 0.5\nother = 0.01\n\nIn addition to above, could you also please tell me if by \"partial dropout\"  you mean partial drop of time and frequency sections in the melspecs?",
      "votes": null
    },
    {
      "id": "1336445",
      "postDate": "06/04/2021 23:05:13",
      "content": "<p>Thank you for comment!</p>\n<p>The lables I used had no theoretical meaning.<br>\nI wanted to try other label strategies, but I have no time to try them.</p>\n<p>Partical dropout means CoarseDropout in pytorch's albumentation library.</p>",
      "rawMarkdown": "Thank you for comment!\n\nThe lables I used had no theoretical meaning.\nI wanted to try other label strategies, but I have no time to try them.\n\nPartical dropout means CoarseDropout in pytorch's albumentation library.",
      "votes": null
    },
    {
      "id": "1336620",
      "postDate": "06/05/2021 05:17:20",
      "content": "<p>Okay. Thanks!</p>",
      "rawMarkdown": "Okay. Thanks!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1336039,
      "author_name": "yashraizada",
      "author_url": "",
      "post_date": "06/04/2021 15:43:55",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/yoshito\" target=\"_blank\">@yoshito</a> thanks for sharing this. I'm new here and want to understand what does these numbers refer to in this statement:</p>\n<p>Moreover, I used not only one hot label but also second_labels for teaching labels.</p>\n<p>primary_label = 0.9<br>\nsecondary_labels = 0.5<br>\nother = 0.01</p>\n<p>In addition to above, could you also please tell me if by \"partial dropout\"  you mean partial drop of time and frequency sections in the melspecs?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1336445,
          "author_name": "yoshito",
          "author_url": "",
          "post_date": "06/04/2021 23:05:13",
          "content": "<p>Thank you for comment!</p>\n<p>The lables I used had no theoretical meaning.<br>\nI wanted to try other label strategies, but I have no time to try them.</p>\n<p>Partical dropout means CoarseDropout in pytorch's albumentation library.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1336620,
          "author_name": "yashraizada",
          "author_url": "",
          "post_date": "06/05/2021 05:17:20",
          "content": "<p>Okay. Thanks!</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1333633": "I appreciate the organizers and participants!\nI am looking forward to seeing other solutions!\n\nI am not good at English, but I hope this helps anyone.\n\nThis is my 87th(LB 90th) solution.\n\n# Solution\n\nI referred @kkiller 's nice notebook https://www.kaggle.com/kneroma/clean-fast-simple-bird-identifier-inference\n\nMy final solution is\n\n1. Ensemble 65models (13 types x 5 folds)\n2. Audio clip length 6 sec in inference time\n3. DataAugmentation in training time\n\n## 1. Ensemble 65models\n\nI noticed increasing model type is effective local score and LB for first a few submissions, so I trained various CNN types and selected nice models seeing the result of my local score in inference time.\n\nModel list is below\n\n- EfficientNet-v2 large\n- EfficientNet-v2 medium\n- EfficientNet-v2 b3\n- EfficientNet-b6\n- EfficientNet-b4\n- EfficientNet-b2\n- EfficientNet-b0\n- ResNeSt50d x3 (2 random seeds and train methods)\n- ResNet200d\n- DenseNet201\n- DenseNet121\n\nOther than these models, I tried NFNet and other ResNet types, but I have no time for training convergence.\n\nFinal prediction labels are selected by threshold=0.2.\n\n\n## 2. Audio clip , 3.Data Augmentation\n\nI decided that audio clip length is 6sec in inference for maximizing local score. ( I trained some models in 7sec crop and the others in holizontal random crops(5~7sec), but the score is maximized in 6sec crops. )\n\nMoreover, I used not only one hot label but also second_labels for teaching labels. \n\n- primary_label = 0.9\n- secondary_labels = 0.5\n- other = 0.01\n\nAnd, I used BCEwithLogits loss.\n\n## Training settings\n\n- StratifiedKFold 5fold\n- RandomSeed = 42, 46\n- LearningRate = 0.0008\n- CosineAnnealingLR\n- Mixup\n\n\n# not improving\n\n1. clip length unsemble in inference time\n2. typical data augmentation (parital dropout, random cutout, gaussian noise, random changing brightness and contrast)\n3. threshold based on every bird from appearance frequencies in training data\n\nAt the end, I want to thank everyone in this competition and kaggle again!",
    "1336039": "Hi @yoshito thanks for sharing this. I'm new here and want to understand what does these numbers refer to in this statement:\n\nMoreover, I used not only one hot label but also second_labels for teaching labels.\n\nprimary_label = 0.9\nsecondary_labels = 0.5\nother = 0.01\n\nIn addition to above, could you also please tell me if by \"partial dropout\"  you mean partial drop of time and frequency sections in the melspecs?",
    "1336445": "Thank you for comment!\n\nThe lables I used had no theoretical meaning.\nI wanted to try other label strategies, but I have no time to try them.\n\nPartical dropout means CoarseDropout in pytorch's albumentation library.",
    "1336620": "Okay. Thanks!"
  },
  "source": "meta"
}