{
  "id": 183266,
  "title": "33rd Place Solution",
  "url": "/competitions/birdsong-recognition/writeups/adityasinha-33rd-place-solution",
  "author_name": "",
  "post_date": "2020-09-16T04:58:51.270246500Z",
  "votes": 3,
  "comment_count": 1,
  "views": 0,
  "content": "<p>First of all thanks to competition organizer and Kaggle to host such interesting competition. For the first time I learned about audio features and how ML is applied on them. Quite a learning.</p>\n<p>Whatever improvement I got is during last 4 days.<br>\nPrevious to that I struggled to match the best public kernel.<br>\nAnyways the score is of single fold Resnest50 as shared in public kernel.</p>\n<h1>Training process: key points</h1>\n<p>Extracted nocall (duration for which no labels are present) from example_test_audio and BIRDCLEF 2020 data and used as noise as well as extra nocall class.<br>\nAs ResNet50 expect images with 3 channel and MelSpectrogram gives single channel I stacked three different Melspectrogram with window size of 1024,1536 ans 2048. Not sure is it a good technique as I have not evaluated the same model and training process with same channel duplicated.</p>\n<h1>Inference:</h1>\n<p>Same as available in public kernel with threshold of 0.5</p>\n<h1>Lots of areas for improvement</h1>\n<p>One idea was to select sample from competition training data based on available metadata like rating, duration, background birds.<br>\nTraining more models and ensemble.</p>\n<p>At the end local CV (CV of random clips from validation data) matches the private leaderboard.<br>\nAlso  I think I am bit lucky as I see quite a lot pvt leaderboard shakeup.</p>",
  "messages": [
    {
      "id": "1012426",
      "postDate": "09/16/2020 04:58:51",
      "content": "<p>First of all thanks to competition organizer and Kaggle to host such interesting competition. For the first time I learned about audio features and how ML is applied on them. Quite a learning.</p>\n<p>Whatever improvement I got is during last 4 days.<br>\nPrevious to that I struggled to match the best public kernel.<br>\nAnyways the score is of single fold Resnest50 as shared in public kernel.</p>\n<h1>Training process: key points</h1>\n<p>Extracted nocall (duration for which no labels are present) from example_test_audio and BIRDCLEF 2020 data and used as noise as well as extra nocall class.<br>\nAs ResNet50 expect images with 3 channel and MelSpectrogram gives single channel I stacked three different Melspectrogram with window size of 1024,1536 ans 2048. Not sure is it a good technique as I have not evaluated the same model and training process with same channel duplicated.</p>\n<h1>Inference:</h1>\n<p>Same as available in public kernel with threshold of 0.5</p>\n<h1>Lots of areas for improvement</h1>\n<p>One idea was to select sample from competition training data based on available metadata like rating, duration, background birds.<br>\nTraining more models and ensemble.</p>\n<p>At the end local CV (CV of random clips from validation data) matches the private leaderboard.<br>\nAlso  I think I am bit lucky as I see quite a lot pvt leaderboard shakeup.</p>",
      "rawMarkdown": "First of all thanks to competition organizer and Kaggle to host such interesting competition. For the first time I learned about audio features and how ML is applied on them. Quite a learning.\n\nWhatever improvement I got is during last 4 days.\nPrevious to that I struggled to match the best public kernel.\nAnyways the score is of single fold Resnest50 as shared in public kernel.\nTraining process: key points\n====================\nExtracted nocall (duration for which no labels are present) from example_test_audio and BIRDCLEF 2020 data and used as noise as well as extra nocall class.\nAs ResNet50 expect images with 3 channel and MelSpectrogram gives single channel I stacked three different Melspectrogram with window size of 1024,1536 ans 2048. Not sure is it a good technique as I have not evaluated the same model and training process with same channel duplicated.\n\nInference:\n===========\nSame as available in public kernel with threshold of 0.5\n\nLots of areas for improvement\n=======================\nOne idea was to select sample from competition training data based on available metadata like rating, duration, background birds.\nTraining more models and ensemble.\n\n\nAt the end local CV (CV of random clips from validation data) matches the private leaderboard.\nAlso  I think I am bit lucky as I see quite a lot pvt leaderboard shakeup.",
      "votes": null
    },
    {
      "id": "1012737",
      "postDate": "09/16/2020 09:01:01",
      "content": "<p>Thank you for sharing your solution, congratulation👍</p>",
      "rawMarkdown": "Thank you for sharing your solution, congratulation👍",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1012737,
      "author_name": "",
      "author_url": "",
      "post_date": "09/16/2020 09:01:01",
      "content": "<p>Thank you for sharing your solution, congratulation👍</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1012426": "First of all thanks to competition organizer and Kaggle to host such interesting competition. For the first time I learned about audio features and how ML is applied on them. Quite a learning.\n\nWhatever improvement I got is during last 4 days.\nPrevious to that I struggled to match the best public kernel.\nAnyways the score is of single fold Resnest50 as shared in public kernel.\nTraining process: key points\n====================\nExtracted nocall (duration for which no labels are present) from example_test_audio and BIRDCLEF 2020 data and used as noise as well as extra nocall class.\nAs ResNet50 expect images with 3 channel and MelSpectrogram gives single channel I stacked three different Melspectrogram with window size of 1024,1536 ans 2048. Not sure is it a good technique as I have not evaluated the same model and training process with same channel duplicated.\n\nInference:\n===========\nSame as available in public kernel with threshold of 0.5\n\nLots of areas for improvement\n=======================\nOne idea was to select sample from competition training data based on available metadata like rating, duration, background birds.\nTraining more models and ensemble.\n\n\nAt the end local CV (CV of random clips from validation data) matches the private leaderboard.\nAlso  I think I am bit lucky as I see quite a lot pvt leaderboard shakeup.",
    "1012737": "Thank you for sharing your solution, congratulation👍"
  },
  "source": "meta"
}