{
  "id": 183339,
  "title": "4-th place solution",
  "url": "/competitions/birdsong-recognition/writeups/dimabert-ususani-4-th-place-solution",
  "author_name": "",
  "post_date": "2020-09-17T07:10:35.870Z",
  "votes": 41,
  "comment_count": 10,
  "views": 0,
  "content": "<p>Hi ,Kagglers!</p>\n<p>It was an amazing competition! And I am happy to share training tips that helped us reach 4-th LB position</p>\n<h1>Data Preparation</h1>\n<p>We used resampled into 32 000 sample rate and normalized audio</p>\n<p>As most competitors, we used spectral features - Logmels</p>\n<p>For most of my models we have added external xeno-canto datasets:<br>\n<a href=\"https://www.kaggle.com/rohanrao/xeno-canto-bird-recordings-extended-a-m\" target=\"_blank\">https://www.kaggle.com/rohanrao/xeno-canto-bird-recordings-extended-a-m</a> <br>\n<a href=\"https://www.kaggle.com/rohanrao/xeno-canto-bird-recordings-extended-n-z\" target=\"_blank\">https://www.kaggle.com/rohanrao/xeno-canto-bird-recordings-extended-n-z</a><br>\nGreat thanks to <a href=\"https://www.kaggle.com/rohanrao\" target=\"_blank\">@rohanrao</a> </p>\n<p>Also we used RMS trimming, in order to get clips with bird calls. More details you can find in <a href=\"https://www.kaggle.com/vladimirsydor/4-th-place-solution-inference-and-training-tips?scriptVersionId=42796948\" target=\"_blank\">our notebook</a></p>\n<h1>Feature Extraction</h1>\n<p>As we were using pretrained CNNs, we have to support it with 3-channel input (or inplace first Conv)</p>\n<p>Worked:</p>\n<ul>\n<li>Repeat of Logmel 3 times - good baseline option</li>\n<li>Use <a href=\"https://pytorch.org/audio/functional.html#compute-deltas\" target=\"_blank\">deltas</a> - We have used concatenation of Logmel, 1-st order delta and 2-nd order delta. It worked the best</li>\n<li>Adding <code>secondary_labels</code> for training</li>\n</ul>\n<p>Not Worked great:</p>\n<ul>\n<li>Time and Frequency encoding - originally used by <a href=\"https://www.kaggle.com/ddanevskyi\" target=\"_blank\">@ddanevskyi</a> in <a href=\"https://www.kaggle.com/c/freesound-audio-tagging-2019/discussion/97926\" target=\"_blank\">Freesound</a>. But it does not work well here</li>\n<li>Adding some more features, like Loudness and Spectral Centroid</li>\n</ul>\n<h1>Validation Scheme</h1>\n<p>All our models were trained in cross-validation mode. So we had one fold for validation. Also we used <code>example_audio</code> as one more validation set. All in all we tracked such metrics:</p>\n<ul>\n<li>loss</li>\n<li>MaP score by one validation fold - <code>map_score</code></li>\n<li>Original competition F1 metric with threshold 0.5 on test set - <code>f1_test_median</code></li>\n<li>Original competition F1 metric with best threshold on test set</li>\n<li>Original competition F1 metric  with threshold 0.5 on validation fold (if we use <code>secondary_labels</code>)</li>\n<li>Original competition F1 metric  with threshold best threshold on validation fold (if we use <code>secondary_labels</code>)</li>\n</ul>\n<p>We made early stopping  and scheduling by MaP score, as it has converged the last one from all metrics</p>\n<p>Then we took 3 best checkpoints by <code>f1_test_median</code> and averaged weights matrices for each fold - some kind of SWA. Then blend 5 models (5 folds) and evaluate on test_set (example audio). This score correlates well with LB till 0.607 point. After this point test_set was nearly useless :)</p>\n<h1>Model</h1>\n<p>We used different EfficientNets (B3, B4, B5) pretrained on <a href=\"https://github.com/rwightman/gen-efficientnet-pytorch\" target=\"_blank\">noisy student</a>.</p>\n<p>We tried some classifier heads. But 2 Layer Dropout-&gt;Liner-&gt;Relu works the best. Also <br>\n <a href=\"https://arxiv.org/pdf/1905.09788.pdf\" target=\"_blank\">Multi-Sample Dropout</a> slightly boosts the performance and give some more stability</p>\n<p>We tried SeResnexts but they did not work at all for us. Also we tried model proposed by <a href=\"https://www.kaggle.com/ddanevskyi\" target=\"_blank\">@ddanevskyi</a> <a href=\"https://github.com/ex4sperans/freesound-classification\" target=\"_blank\">here</a> but it worked worse.</p>\n<h1>Training process</h1>\n<p>We used <code>Adam</code> optimizer and   <code>ReduceLROnPlateau</code> scheduler and <code>BCEwithLogits</code> loss</p>\n<p>Augmentations really boosted performance (~2%). We listened to example audio and tried to choose such augmentations, that can shift our train set to example audio:</p>\n<ul>\n<li>Gain (to make bird call less loud)</li>\n<li>Background noise - very and less loud. We have taken some background from <a href=\"https://www.kaggle.com/mmoreaux/environmental-sound-classification-50\" target=\"_blank\">here</a> and some 5 second clips directly from example audio. Finally we created such <a href=\"https://www.kaggle.com/vladimirsydor/cornelli-background-noises\" target=\"_blank\">background dataset</a></li>\n<li>LowFrequancy CutOff - we found out, that example audio has no lower frequency</li>\n</ul>\n<p>Also we used MixUp - we add audios and take max from two one-hot targets. As it was done by  <a href=\"https://www.kaggle.com/ddanevskyi\" target=\"_blank\">@ddanevskyi</a> in <a href=\"https://www.kaggle.com/c/freesound-audio-tagging-2019/discussion/97926\" target=\"_blank\">Freesound</a></p>\n<h1>Choose Final Blend</h1>\n<p>First we tried simply to Blend all our good models - 14 experiments (70 models) and it gave us 0.623 Public score and 0.669 Private score. But then we have taken 4 best experiments with external data and 3 best without external data, which gave us 0.624 Public and 0.67 Private scores</p>\n<p>All Training details and inference of best Blend you can find in our <a href=\"https://www.kaggle.com/vladimirsydor/4-th-place-solution-inference-and-training-tips?scriptVersionId=42796948\" target=\"_blank\">inference notebook</a> </p>\n<h1>Framework</h1>\n<p>For training and experiment monitoring was used <a href=\"https://pytorch.org/\" target=\"_blank\">Pytorch</a> and <a href=\"https://catalyst-team.github.io/catalyst/\" target=\"_blank\">Catalyst</a> frameworks. Great thanks to <a href=\"https://www.kaggle.com/scitator\" target=\"_blank\">@scitator</a> and catalyst team! </p>\n<h1>P.S</h1>\n<p>Great thanks to my teammate - <a href=\"https://www.kaggle.com/khapilins\" target=\"_blank\">@khapilins</a><br>\nAlso thanks to all DS community, especially to <a href=\"https://www.kaggle.com/ddanevskyi\" target=\"_blank\">@ddanevskyi</a>, <a href=\"https://www.kaggle.com/yaroshevskiy\" target=\"_blank\">@yaroshevskiy</a> and <a href=\"https://www.kaggle.com/frednavruzov\" target=\"_blank\">@frednavruzov</a>. They have taught me a lot and give inspiration to take part in Kaggle competitions.<br>\nAlso thanks to all Kaggle team and community.<br>\nAnd happy Kaggling! </p>",
  "messages": [
    {
      "id": "1012836",
      "postDate": "09/16/2020 10:32:49",
      "content": "<p>Hi ,Kagglers!</p>\n<p>It was an amazing competition! And I am happy to share training tips that helped us reach 4-th LB position</p>\n<h1>Data Preparation</h1>\n<p>We used resampled into 32 000 sample rate and normalized audio</p>\n<p>As most competitors, we used spectral features - Logmels</p>\n<p>For most of my models we have added external xeno-canto datasets:<br>\n<a href=\"https://www.kaggle.com/rohanrao/xeno-canto-bird-recordings-extended-a-m\" target=\"_blank\">https://www.kaggle.com/rohanrao/xeno-canto-bird-recordings-extended-a-m</a> <br>\n<a href=\"https://www.kaggle.com/rohanrao/xeno-canto-bird-recordings-extended-n-z\" target=\"_blank\">https://www.kaggle.com/rohanrao/xeno-canto-bird-recordings-extended-n-z</a><br>\nGreat thanks to <a href=\"https://www.kaggle.com/rohanrao\" target=\"_blank\">@rohanrao</a> </p>\n<p>Also we used RMS trimming, in order to get clips with bird calls. More details you can find in <a href=\"https://www.kaggle.com/vladimirsydor/4-th-place-solution-inference-and-training-tips?scriptVersionId=42796948\" target=\"_blank\">our notebook</a></p>\n<h1>Feature Extraction</h1>\n<p>As we were using pretrained CNNs, we have to support it with 3-channel input (or inplace first Conv)</p>\n<p>Worked:</p>\n<ul>\n<li>Repeat of Logmel 3 times - good baseline option</li>\n<li>Use <a href=\"https://pytorch.org/audio/functional.html#compute-deltas\" target=\"_blank\">deltas</a> - We have used concatenation of Logmel, 1-st order delta and 2-nd order delta. It worked the best</li>\n<li>Adding <code>secondary_labels</code> for training</li>\n</ul>\n<p>Not Worked great:</p>\n<ul>\n<li>Time and Frequency encoding - originally used by <a href=\"https://www.kaggle.com/ddanevskyi\" target=\"_blank\">@ddanevskyi</a> in <a href=\"https://www.kaggle.com/c/freesound-audio-tagging-2019/discussion/97926\" target=\"_blank\">Freesound</a>. But it does not work well here</li>\n<li>Adding some more features, like Loudness and Spectral Centroid</li>\n</ul>\n<h1>Validation Scheme</h1>\n<p>All our models were trained in cross-validation mode. So we had one fold for validation. Also we used <code>example_audio</code> as one more validation set. All in all we tracked such metrics:</p>\n<ul>\n<li>loss</li>\n<li>MaP score by one validation fold - <code>map_score</code></li>\n<li>Original competition F1 metric with threshold 0.5 on test set - <code>f1_test_median</code></li>\n<li>Original competition F1 metric with best threshold on test set</li>\n<li>Original competition F1 metric  with threshold 0.5 on validation fold (if we use <code>secondary_labels</code>)</li>\n<li>Original competition F1 metric  with threshold best threshold on validation fold (if we use <code>secondary_labels</code>)</li>\n</ul>\n<p>We made early stopping  and scheduling by MaP score, as it has converged the last one from all metrics</p>\n<p>Then we took 3 best checkpoints by <code>f1_test_median</code> and averaged weights matrices for each fold - some kind of SWA. Then blend 5 models (5 folds) and evaluate on test_set (example audio). This score correlates well with LB till 0.607 point. After this point test_set was nearly useless :)</p>\n<h1>Model</h1>\n<p>We used different EfficientNets (B3, B4, B5) pretrained on <a href=\"https://github.com/rwightman/gen-efficientnet-pytorch\" target=\"_blank\">noisy student</a>.</p>\n<p>We tried some classifier heads. But 2 Layer Dropout-&gt;Liner-&gt;Relu works the best. Also <br>\n <a href=\"https://arxiv.org/pdf/1905.09788.pdf\" target=\"_blank\">Multi-Sample Dropout</a> slightly boosts the performance and give some more stability</p>\n<p>We tried SeResnexts but they did not work at all for us. Also we tried model proposed by <a href=\"https://www.kaggle.com/ddanevskyi\" target=\"_blank\">@ddanevskyi</a> <a href=\"https://github.com/ex4sperans/freesound-classification\" target=\"_blank\">here</a> but it worked worse.</p>\n<h1>Training process</h1>\n<p>We used <code>Adam</code> optimizer and   <code>ReduceLROnPlateau</code> scheduler and <code>BCEwithLogits</code> loss</p>\n<p>Augmentations really boosted performance (~2%). We listened to example audio and tried to choose such augmentations, that can shift our train set to example audio:</p>\n<ul>\n<li>Gain (to make bird call less loud)</li>\n<li>Background noise - very and less loud. We have taken some background from <a href=\"https://www.kaggle.com/mmoreaux/environmental-sound-classification-50\" target=\"_blank\">here</a> and some 5 second clips directly from example audio. Finally we created such <a href=\"https://www.kaggle.com/vladimirsydor/cornelli-background-noises\" target=\"_blank\">background dataset</a></li>\n<li>LowFrequancy CutOff - we found out, that example audio has no lower frequency</li>\n</ul>\n<p>Also we used MixUp - we add audios and take max from two one-hot targets. As it was done by  <a href=\"https://www.kaggle.com/ddanevskyi\" target=\"_blank\">@ddanevskyi</a> in <a href=\"https://www.kaggle.com/c/freesound-audio-tagging-2019/discussion/97926\" target=\"_blank\">Freesound</a></p>\n<h1>Choose Final Blend</h1>\n<p>First we tried simply to Blend all our good models - 14 experiments (70 models) and it gave us 0.623 Public score and 0.669 Private score. But then we have taken 4 best experiments with external data and 3 best without external data, which gave us 0.624 Public and 0.67 Private scores</p>\n<p>All Training details and inference of best Blend you can find in our <a href=\"https://www.kaggle.com/vladimirsydor/4-th-place-solution-inference-and-training-tips?scriptVersionId=42796948\" target=\"_blank\">inference notebook</a> </p>\n<h1>Framework</h1>\n<p>For training and experiment monitoring was used <a href=\"https://pytorch.org/\" target=\"_blank\">Pytorch</a> and <a href=\"https://catalyst-team.github.io/catalyst/\" target=\"_blank\">Catalyst</a> frameworks. Great thanks to <a href=\"https://www.kaggle.com/scitator\" target=\"_blank\">@scitator</a> and catalyst team! </p>\n<h1>P.S</h1>\n<p>Great thanks to my teammate - <a href=\"https://www.kaggle.com/khapilins\" target=\"_blank\">@khapilins</a><br>\nAlso thanks to all DS community, especially to <a href=\"https://www.kaggle.com/ddanevskyi\" target=\"_blank\">@ddanevskyi</a>, <a href=\"https://www.kaggle.com/yaroshevskiy\" target=\"_blank\">@yaroshevskiy</a> and <a href=\"https://www.kaggle.com/frednavruzov\" target=\"_blank\">@frednavruzov</a>. They have taught me a lot and give inspiration to take part in Kaggle competitions.<br>\nAlso thanks to all Kaggle team and community.<br>\nAnd happy Kaggling! </p>",
      "rawMarkdown": "Hi ,Kagglers!\n\nIt was an amazing competition! And I am happy to share training tips that helped us reach 4-th LB position\n\n# Data Preparation\n\nWe used resampled into 32 000 sample rate and normalized audio\n\nAs most competitors, we used spectral features - Logmels\n\nFor most of my models we have added external xeno-canto datasets:\nhttps://www.kaggle.com/rohanrao/xeno-canto-bird-recordings-extended-a-m \nhttps://www.kaggle.com/rohanrao/xeno-canto-bird-recordings-extended-n-z\nGreat thanks to @rohanrao \n\nAlso we used RMS trimming, in order to get clips with bird calls. More details you can find in [our notebook](https://www.kaggle.com/vladimirsydor/4-th-place-solution-inference-and-training-tips?scriptVersionId=42796948)\n\n# Feature Extraction \n\nAs we were using pretrained CNNs, we have to support it with 3-channel input (or inplace first Conv)\n\nWorked:\n- Repeat of Logmel 3 times - good baseline option\n- Use [deltas](https://pytorch.org/audio/functional.html#compute-deltas) - We have used concatenation of Logmel, 1-st order delta and 2-nd order delta. It worked the best\n- Adding `secondary_labels` for training\n\nNot Worked great:\n- Time and Frequency encoding - originally used by @ddanevskyi in [Freesound](https://www.kaggle.com/c/freesound-audio-tagging-2019/discussion/97926). But it does not work well here\n- Adding some more features, like Loudness and Spectral Centroid\n\n# Validation Scheme\n\nAll our models were trained in cross-validation mode. So we had one fold for validation. Also we used `example_audio` as one more validation set. All in all we tracked such metrics:\n- loss\n- MaP score by one validation fold - `map_score`\n- Original competition F1 metric with threshold 0.5 on test set - `f1_test_median`\n-  Original competition F1 metric with best threshold on test set\n- Original competition F1 metric  with threshold 0.5 on validation fold (if we use `secondary_labels`)\n- Original competition F1 metric  with threshold best threshold on validation fold (if we use `secondary_labels`)\n\nWe made early stopping  and scheduling by MaP score, as it has converged the last one from all metrics\n\nThen we took 3 best checkpoints by `f1_test_median` and averaged weights matrices for each fold - some kind of SWA. Then blend 5 models (5 folds) and evaluate on test_set (example audio). This score correlates well with LB till 0.607 point. After this point test_set was nearly useless :)\n\n# Model\n\nWe used different EfficientNets (B3, B4, B5) pretrained on [noisy student](https://github.com/rwightman/gen-efficientnet-pytorch).\n\nWe tried some classifier heads. But 2 Layer Dropout->Liner->Relu works the best. Also \n [Multi-Sample Dropout](https://arxiv.org/pdf/1905.09788.pdf) slightly boosts the performance and give some more stability\n\nWe tried SeResnexts but they did not work at all for us. Also we tried model proposed by @ddanevskyi [here](https://github.com/ex4sperans/freesound-classification) but it worked worse.\n\n# Training process \n\nWe used `Adam` optimizer and   `ReduceLROnPlateau` scheduler and `BCEwithLogits` loss\n\nAugmentations really boosted performance (~2%). We listened to example audio and tried to choose such augmentations, that can shift our train set to example audio:\n- Gain (to make bird call less loud)\n- Background noise - very and less loud. We have taken some background from [here](https://www.kaggle.com/mmoreaux/environmental-sound-classification-50) and some 5 second clips directly from example audio. Finally we created such [background dataset](https://www.kaggle.com/vladimirsydor/cornelli-background-noises)\n- LowFrequancy CutOff - we found out, that example audio has no lower frequency\n\nAlso we used MixUp - we add audios and take max from two one-hot targets. As it was done by  @ddanevskyi in [Freesound](https://www.kaggle.com/c/freesound-audio-tagging-2019/discussion/97926)\n\n# Choose Final Blend\n\nFirst we tried simply to Blend all our good models - 14 experiments (70 models) and it gave us 0.623 Public score and 0.669 Private score. But then we have taken 4 best experiments with external data and 3 best without external data, which gave us 0.624 Public and 0.67 Private scores\n\nAll Training details and inference of best Blend you can find in our [inference notebook](https://www.kaggle.com/vladimirsydor/4-th-place-solution-inference-and-training-tips?scriptVersionId=42796948) \n\n# Framework \n\nFor training and experiment monitoring was used [Pytorch](https://pytorch.org/) and [Catalyst](https://catalyst-team.github.io/catalyst/) frameworks. Great thanks to @scitator and catalyst team! \n\n# P.S\n\nGreat thanks to my teammate - @khapilins\nAlso thanks to all DS community, especially to @ddanevskyi, @yaroshevskiy and @frednavruzov. They have taught me a lot and give inspiration to take part in Kaggle competitions.\nAlso thanks to all Kaggle team and community.\nAnd happy Kaggling!",
      "votes": null
    },
    {
      "id": "1012904",
      "postDate": "09/16/2020 11:33:39",
      "content": "<p>Congratulations</p>",
      "rawMarkdown": "Congratulations",
      "votes": null
    },
    {
      "id": "1012941",
      "postDate": "09/16/2020 12:06:48",
      "content": "<p>Congrats on the solution and result!  I didn't know about deltas, looks very interesting.</p>",
      "rawMarkdown": "Congrats on the solution and result!  I didn't know about deltas, looks very interesting.",
      "votes": null
    },
    {
      "id": "1012948",
      "postDate": "09/16/2020 12:12:15",
      "content": "<p>Congrats on 4th place and thanks for the writeup <a href=\"https://www.kaggle.com/vladimirsydor\" target=\"_blank\">@vladimirsydor</a> </p>",
      "rawMarkdown": "Congrats on 4th place and thanks for the writeup @vladimirsydor",
      "votes": null
    },
    {
      "id": "1012964",
      "postDate": "09/16/2020 12:28:16",
      "content": "<p><a href=\"https://www.kaggle.com/vladimirsydor\" target=\"_blank\">@vladimirsydor</a> Congratulations! and thanks for sharing solution, I have one simple question - what is your duration for training samples? Some top scorers show that they use 30 sec to gain chance to capture bird-singing part. Then I'm wondering if it is also your case. Anyway your notebook looks great.</p>",
      "rawMarkdown": "vladimirsydor Congratulations! and thanks for sharing solution, I have one simple question - what is your duration for training samples? Some top scorers show that they use 30 sec to gain chance to capture bird-singing part. Then I'm wondering if it is also your case. Anyway your notebook looks great.",
      "votes": null
    },
    {
      "id": "1013045",
      "postDate": "09/16/2020 13:26:22",
      "content": "<p>5 seconds for all models</p>",
      "rawMarkdown": "5 seconds for all models",
      "votes": null
    },
    {
      "id": "1013058",
      "postDate": "09/16/2020 13:34:54",
      "content": "<p>Thanks for quick answer.</p>",
      "rawMarkdown": "Thanks for quick answer.",
      "votes": null
    },
    {
      "id": "1013072",
      "postDate": "09/16/2020 13:41:37",
      "content": "<p>Great 😍 <br>\nAnd Congratulations on your first GOLD and becoming Master 👍</p>",
      "rawMarkdown": "Great 😍 \nAnd Congratulations on your first GOLD and becoming Master 👍",
      "votes": null
    },
    {
      "id": "1013550",
      "postDate": "09/16/2020 18:57:09",
      "content": "<p>congrats for that</p>",
      "rawMarkdown": "congrats for that",
      "votes": null
    },
    {
      "id": "1013631",
      "postDate": "09/16/2020 20:39:10",
      "content": "<p>hi . tankx for sharring solution.🙏 </p>",
      "rawMarkdown": "hi . tankx for sharring solution.🙏",
      "votes": null
    },
    {
      "id": "1018912",
      "postDate": "09/20/2020 04:55:47",
      "content": "<p>Great insights!</p>",
      "rawMarkdown": "Great insights!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1012904,
      "author_name": "fiyeroleung",
      "author_url": "",
      "post_date": "09/16/2020 11:33:39",
      "content": "<p>Congratulations</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1012941,
      "author_name": "cpmpml",
      "author_url": "",
      "post_date": "09/16/2020 12:06:48",
      "content": "<p>Congrats on the solution and result!  I didn't know about deltas, looks very interesting.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1012948,
      "author_name": "duykhanh99",
      "author_url": "",
      "post_date": "09/16/2020 12:12:15",
      "content": "<p>Congrats on 4th place and thanks for the writeup <a href=\"https://www.kaggle.com/vladimirsydor\" target=\"_blank\">@vladimirsydor</a> </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1012964,
      "author_name": "daisukelab",
      "author_url": "",
      "post_date": "09/16/2020 12:28:16",
      "content": "<p><a href=\"https://www.kaggle.com/vladimirsydor\" target=\"_blank\">@vladimirsydor</a> Congratulations! and thanks for sharing solution, I have one simple question - what is your duration for training samples? Some top scorers show that they use 30 sec to gain chance to capture bird-singing part. Then I'm wondering if it is also your case. Anyway your notebook looks great.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1013045,
          "author_name": "vladimirsydor",
          "author_url": "",
          "post_date": "09/16/2020 13:26:22",
          "content": "<p>5 seconds for all models</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1013058,
          "author_name": "daisukelab",
          "author_url": "",
          "post_date": "09/16/2020 13:34:54",
          "content": "<p>Thanks for quick answer.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1013072,
      "author_name": "rahim3",
      "author_url": "",
      "post_date": "09/16/2020 13:41:37",
      "content": "<p>Great 😍 <br>\nAnd Congratulations on your first GOLD and becoming Master 👍</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1013550,
      "author_name": "elvinagammed",
      "author_url": "",
      "post_date": "09/16/2020 18:57:09",
      "content": "<p>congrats for that</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1013631,
      "author_name": "zeintiz",
      "author_url": "",
      "post_date": "09/16/2020 20:39:10",
      "content": "<p>hi . tankx for sharring solution.🙏 </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1018912,
      "author_name": "varunyadav17",
      "author_url": "",
      "post_date": "09/20/2020 04:55:47",
      "content": "<p>Great insights!</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1012836": "Hi ,Kagglers!\n\nIt was an amazing competition! And I am happy to share training tips that helped us reach 4-th LB position\n\n# Data Preparation\n\nWe used resampled into 32 000 sample rate and normalized audio\n\nAs most competitors, we used spectral features - Logmels\n\nFor most of my models we have added external xeno-canto datasets:\nhttps://www.kaggle.com/rohanrao/xeno-canto-bird-recordings-extended-a-m \nhttps://www.kaggle.com/rohanrao/xeno-canto-bird-recordings-extended-n-z\nGreat thanks to @rohanrao \n\nAlso we used RMS trimming, in order to get clips with bird calls. More details you can find in [our notebook](https://www.kaggle.com/vladimirsydor/4-th-place-solution-inference-and-training-tips?scriptVersionId=42796948)\n\n# Feature Extraction \n\nAs we were using pretrained CNNs, we have to support it with 3-channel input (or inplace first Conv)\n\nWorked:\n- Repeat of Logmel 3 times - good baseline option\n- Use [deltas](https://pytorch.org/audio/functional.html#compute-deltas) - We have used concatenation of Logmel, 1-st order delta and 2-nd order delta. It worked the best\n- Adding `secondary_labels` for training\n\nNot Worked great:\n- Time and Frequency encoding - originally used by @ddanevskyi in [Freesound](https://www.kaggle.com/c/freesound-audio-tagging-2019/discussion/97926). But it does not work well here\n- Adding some more features, like Loudness and Spectral Centroid\n\n# Validation Scheme\n\nAll our models were trained in cross-validation mode. So we had one fold for validation. Also we used `example_audio` as one more validation set. All in all we tracked such metrics:\n- loss\n- MaP score by one validation fold - `map_score`\n- Original competition F1 metric with threshold 0.5 on test set - `f1_test_median`\n-  Original competition F1 metric with best threshold on test set\n- Original competition F1 metric  with threshold 0.5 on validation fold (if we use `secondary_labels`)\n- Original competition F1 metric  with threshold best threshold on validation fold (if we use `secondary_labels`)\n\nWe made early stopping  and scheduling by MaP score, as it has converged the last one from all metrics\n\nThen we took 3 best checkpoints by `f1_test_median` and averaged weights matrices for each fold - some kind of SWA. Then blend 5 models (5 folds) and evaluate on test_set (example audio). This score correlates well with LB till 0.607 point. After this point test_set was nearly useless :)\n\n# Model\n\nWe used different EfficientNets (B3, B4, B5) pretrained on [noisy student](https://github.com/rwightman/gen-efficientnet-pytorch).\n\nWe tried some classifier heads. But 2 Layer Dropout->Liner->Relu works the best. Also \n [Multi-Sample Dropout](https://arxiv.org/pdf/1905.09788.pdf) slightly boosts the performance and give some more stability\n\nWe tried SeResnexts but they did not work at all for us. Also we tried model proposed by @ddanevskyi [here](https://github.com/ex4sperans/freesound-classification) but it worked worse.\n\n# Training process \n\nWe used `Adam` optimizer and   `ReduceLROnPlateau` scheduler and `BCEwithLogits` loss\n\nAugmentations really boosted performance (~2%). We listened to example audio and tried to choose such augmentations, that can shift our train set to example audio:\n- Gain (to make bird call less loud)\n- Background noise - very and less loud. We have taken some background from [here](https://www.kaggle.com/mmoreaux/environmental-sound-classification-50) and some 5 second clips directly from example audio. Finally we created such [background dataset](https://www.kaggle.com/vladimirsydor/cornelli-background-noises)\n- LowFrequancy CutOff - we found out, that example audio has no lower frequency\n\nAlso we used MixUp - we add audios and take max from two one-hot targets. As it was done by  @ddanevskyi in [Freesound](https://www.kaggle.com/c/freesound-audio-tagging-2019/discussion/97926)\n\n# Choose Final Blend\n\nFirst we tried simply to Blend all our good models - 14 experiments (70 models) and it gave us 0.623 Public score and 0.669 Private score. But then we have taken 4 best experiments with external data and 3 best without external data, which gave us 0.624 Public and 0.67 Private scores\n\nAll Training details and inference of best Blend you can find in our [inference notebook](https://www.kaggle.com/vladimirsydor/4-th-place-solution-inference-and-training-tips?scriptVersionId=42796948) \n\n# Framework \n\nFor training and experiment monitoring was used [Pytorch](https://pytorch.org/) and [Catalyst](https://catalyst-team.github.io/catalyst/) frameworks. Great thanks to @scitator and catalyst team! \n\n# P.S\n\nGreat thanks to my teammate - @khapilins\nAlso thanks to all DS community, especially to @ddanevskyi, @yaroshevskiy and @frednavruzov. They have taught me a lot and give inspiration to take part in Kaggle competitions.\nAlso thanks to all Kaggle team and community.\nAnd happy Kaggling!",
    "1012904": "Congratulations",
    "1012941": "Congrats on the solution and result!  I didn't know about deltas, looks very interesting.",
    "1012948": "Congrats on 4th place and thanks for the writeup @vladimirsydor",
    "1012964": "vladimirsydor Congratulations! and thanks for sharing solution, I have one simple question - what is your duration for training samples? Some top scorers show that they use 30 sec to gain chance to capture bird-singing part. Then I'm wondering if it is also your case. Anyway your notebook looks great.",
    "1013045": "5 seconds for all models",
    "1013058": "Thanks for quick answer.",
    "1013072": "Great 😍 \nAnd Congratulations on your first GOLD and becoming Master 👍",
    "1013550": "congrats for that",
    "1013631": "hi . tankx for sharring solution.🙏",
    "1018912": "Great insights!"
  },
  "source": "meta"
}