{
  "id": 322478,
  "title": "Summary - 3rd Place solution",
  "url": "/competitions/kaggle-pog-series-s01e02/writeups/phaedrus-summary-3rd-place-solution",
  "author_name": "",
  "post_date": "2022-05-02T11:34:30.852826300Z",
  "votes": 16,
  "comment_count": 7,
  "views": 0,
  "content": "<p>Hello everyone.</p>\n<p>Thanks for <a href=\"https://www.kaggle.com/robikscube\" target=\"_blank\">@robikscube</a> for hosting this awesome competition. It was fun for to me tune my signal processing skills. </p>\n<p><strong>My approach to this competition:</strong></p>\n<p><strong>Datasets:</strong> </p>\n<p>I resampled to original audio files to 16k frequency and then created the following datasets for CNNs.</p>\n<p><strong>Mel Spec</strong>: I created 3 differently sized spec datasets for this competition. Appropriate min-max normalisation was done for each spectrogram </p>\n<table>\n<thead>\n<tr>\n<th>Hop Len</th>\n<th>N Mels</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>512</td>\n<td>128</td>\n</tr>\n<tr>\n<td>512</td>\n<td>64</td>\n</tr>\n<tr>\n<td>448</td>\n<td>160</td>\n</tr>\n</tbody>\n</table>\n<p><strong>CQT</strong>: I also experimented with CQT (Constant Q transform) datasets. It have a good boost in local ensemble CV.</p>\n<table>\n<thead>\n<tr>\n<th>Hop Len</th>\n<th>N Bins</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>512</td>\n<td>64</td>\n</tr>\n<tr>\n<td>448</td>\n<td>84</td>\n</tr>\n</tbody>\n</table>\n<p><strong>Validation</strong></p>\n<p>Validation was based on 5 fold mode ensemble. Mean logit/prob based ensemble wasn't corresponding as well with the LB.</p>\n<p><strong>Framework</strong><br>\nI primarily used fastai for this competition as I had a good fastai pipeline set from a previous signal processing competition.</p>\n<p><strong>Things that worked - Training</strong></p>\n<ul>\n<li>Reduced stride for 1st conv layers</li>\n<li>Mixup (0.25)</li>\n<li>Labelsmoothing (0.025)</li>\n<li>Noisy student (re) training</li>\n<li>Ranger, and Adam optimiser depending on model</li>\n<li>fit_flat_cos for normal training and fit_one_cycle for noisy training</li>\n</ul>\n<p><strong>Things that did not work - Training</strong></p>\n<ul>\n<li>Larger models</li>\n<li>Pseudo Labels</li>\n<li>Transformers (ViT, Deit)</li>\n</ul>\n<p><strong>Models</strong></p>\n<p>My final submit was based on 12 models which are trained on either CQT or Spec datasets.</p>\n<p>CQT - 4 ecaresnext50t_32x4d models trained different sized CQT datasets</p>\n<pre><code>l_1_CQT = inferenceCQT(_model = 'ecaresnext50t_32x4d',modelpth='pog2-ecaresnext50t-32x4d-cqt-512-64')\nl_2_CQT = inferenceCQT(_model = 'ecaresnext50t_32x4d',modelpth='pog2-ecaresnext50t-32x4d-cqt-512-64-mixup')\nl_3_CQT = inferenceCQT(_model = 'ecaresnext50t_32x4d',modelpth='pog2-ecaresnext50t-32x4d-cqt-448-84-tune')\nl_4_CQT = inferenceCQT(_model = 'ecaresnext50t_32x4d',modelpth='pog2-ecaresnext50t-32x4d-cqt-512-64-noisy')\n</code></pre>\n<p>Spec - 8 models on different sized spec datasets</p>\n<pre><code>l_1 = inference(_model='resnest50d',modelpth='pog2-resnest50d-monospec-hop448-mels160-ftune')\nl_2 = inference(_model='resnest50d',modelpth='pog2-resnest50d-monospec-hop448-mels160-noisytrain')\nl_3 = inference(_model = 'ecaresnext50t_32x4d',modelpth='pog2-ecaresnext50t-32x4d-monospec-hop448-mels160')\nl_4 = inference(_model = 'resnext50d_32x4d',modelpth='pog2-resnext50d-32x4d-monospec-hop448-mels160')\nl_5 = inference(_model = 'ecaresnext50t_32x4d',modelpth='pog2-ecaresnext50t-32x4d-monospec-hop448-mels64',sz=[64])\nl_6 = inference(_model = 'ecaresnext50t_32x4d',modelpth='pog2-ecaresnext50t-32x4d-spec-hop448-mels64-tune',sz=[64])\nl_7 = inference(_model = 'ecaresnext50t_32x4d',modelpth='pog2-ecaresnext50t-32x4d-spec-hop448-mels64-noisy',sz=[64])\nl_8 = inference(_model = 'resnext50d_32x4d',modelpth='pog2-resnext50t-32x4d-spec-hop448-mels64-noisy',sz=[64])\n</code></pre>\n<p>Inference pipeline is available at: <a href=\"https://www.kaggle.com/code/pheadrus/inference-pipeline-v0?scriptVersionId=93716687\" target=\"_blank\">https://www.kaggle.com/code/pheadrus/inference-pipeline-v0?scriptVersionId=93716687</a></p>\n<p>Thanks</p>",
  "messages": [
    {
      "id": "1774684",
      "postDate": "05/02/2022 11:34:30",
      "content": "<p>Hello everyone.</p>\n<p>Thanks for <a href=\"https://www.kaggle.com/robikscube\" target=\"_blank\">@robikscube</a> for hosting this awesome competition. It was fun for to me tune my signal processing skills. </p>\n<p><strong>My approach to this competition:</strong></p>\n<p><strong>Datasets:</strong> </p>\n<p>I resampled to original audio files to 16k frequency and then created the following datasets for CNNs.</p>\n<p><strong>Mel Spec</strong>: I created 3 differently sized spec datasets for this competition. Appropriate min-max normalisation was done for each spectrogram </p>\n<table>\n<thead>\n<tr>\n<th>Hop Len</th>\n<th>N Mels</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>512</td>\n<td>128</td>\n</tr>\n<tr>\n<td>512</td>\n<td>64</td>\n</tr>\n<tr>\n<td>448</td>\n<td>160</td>\n</tr>\n</tbody>\n</table>\n<p><strong>CQT</strong>: I also experimented with CQT (Constant Q transform) datasets. It have a good boost in local ensemble CV.</p>\n<table>\n<thead>\n<tr>\n<th>Hop Len</th>\n<th>N Bins</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>512</td>\n<td>64</td>\n</tr>\n<tr>\n<td>448</td>\n<td>84</td>\n</tr>\n</tbody>\n</table>\n<p><strong>Validation</strong></p>\n<p>Validation was based on 5 fold mode ensemble. Mean logit/prob based ensemble wasn't corresponding as well with the LB.</p>\n<p><strong>Framework</strong><br>\nI primarily used fastai for this competition as I had a good fastai pipeline set from a previous signal processing competition.</p>\n<p><strong>Things that worked - Training</strong></p>\n<ul>\n<li>Reduced stride for 1st conv layers</li>\n<li>Mixup (0.25)</li>\n<li>Labelsmoothing (0.025)</li>\n<li>Noisy student (re) training</li>\n<li>Ranger, and Adam optimiser depending on model</li>\n<li>fit_flat_cos for normal training and fit_one_cycle for noisy training</li>\n</ul>\n<p><strong>Things that did not work - Training</strong></p>\n<ul>\n<li>Larger models</li>\n<li>Pseudo Labels</li>\n<li>Transformers (ViT, Deit)</li>\n</ul>\n<p><strong>Models</strong></p>\n<p>My final submit was based on 12 models which are trained on either CQT or Spec datasets.</p>\n<p>CQT - 4 ecaresnext50t_32x4d models trained different sized CQT datasets</p>\n<pre><code>l_1_CQT = inferenceCQT(_model = 'ecaresnext50t_32x4d',modelpth='pog2-ecaresnext50t-32x4d-cqt-512-64')\nl_2_CQT = inferenceCQT(_model = 'ecaresnext50t_32x4d',modelpth='pog2-ecaresnext50t-32x4d-cqt-512-64-mixup')\nl_3_CQT = inferenceCQT(_model = 'ecaresnext50t_32x4d',modelpth='pog2-ecaresnext50t-32x4d-cqt-448-84-tune')\nl_4_CQT = inferenceCQT(_model = 'ecaresnext50t_32x4d',modelpth='pog2-ecaresnext50t-32x4d-cqt-512-64-noisy')\n</code></pre>\n<p>Spec - 8 models on different sized spec datasets</p>\n<pre><code>l_1 = inference(_model='resnest50d',modelpth='pog2-resnest50d-monospec-hop448-mels160-ftune')\nl_2 = inference(_model='resnest50d',modelpth='pog2-resnest50d-monospec-hop448-mels160-noisytrain')\nl_3 = inference(_model = 'ecaresnext50t_32x4d',modelpth='pog2-ecaresnext50t-32x4d-monospec-hop448-mels160')\nl_4 = inference(_model = 'resnext50d_32x4d',modelpth='pog2-resnext50d-32x4d-monospec-hop448-mels160')\nl_5 = inference(_model = 'ecaresnext50t_32x4d',modelpth='pog2-ecaresnext50t-32x4d-monospec-hop448-mels64',sz=[64])\nl_6 = inference(_model = 'ecaresnext50t_32x4d',modelpth='pog2-ecaresnext50t-32x4d-spec-hop448-mels64-tune',sz=[64])\nl_7 = inference(_model = 'ecaresnext50t_32x4d',modelpth='pog2-ecaresnext50t-32x4d-spec-hop448-mels64-noisy',sz=[64])\nl_8 = inference(_model = 'resnext50d_32x4d',modelpth='pog2-resnext50t-32x4d-spec-hop448-mels64-noisy',sz=[64])\n</code></pre>\n<p>Inference pipeline is available at: <a href=\"https://www.kaggle.com/code/pheadrus/inference-pipeline-v0?scriptVersionId=93716687\" target=\"_blank\">https://www.kaggle.com/code/pheadrus/inference-pipeline-v0?scriptVersionId=93716687</a></p>\n<p>Thanks</p>",
      "rawMarkdown": "Hello everyone.\n\nThanks for @robikscube for hosting this awesome competition. It was fun for to me tune my signal processing skills. \n\n**My approach to this competition:**\n\n**Datasets:** \n\nI resampled to original audio files to 16k frequency and then created the following datasets for CNNs.\n\n**Mel Spec**: I created 3 differently sized spec datasets for this competition. Appropriate min-max normalisation was done for each spectrogram \n\n| Hop Len |N Mels  |\n| --- | --- |\n| 512 | 128 |\n| 512 | 64 |\n| 448 | 160 |\n\n**CQT**: I also experimented with CQT (Constant Q transform) datasets. It have a good boost in local ensemble CV.\n\n| Hop Len |N Bins  |\n| --- | --- |\n| 512 | 64 |\n| 448 | 84 |\n\n**Validation**\n\nValidation was based on 5 fold mode ensemble. Mean logit/prob based ensemble wasn't corresponding as well with the LB.\n\n**Framework**\nI primarily used fastai for this competition as I had a good fastai pipeline set from a previous signal processing competition.\n\n \n**Things that worked - Training**\n\n- Reduced stride for 1st conv layers\n- Mixup (0.25)\n- Labelsmoothing (0.025)\n- Noisy student (re) training\n- Ranger, and Adam optimiser depending on model\n- fit_flat_cos for normal training and fit_one_cycle for noisy training\n\n**Things that did not work - Training**\n- Larger models\n- Pseudo Labels\n- Transformers (ViT, Deit)\n\n**Models**\n\nMy final submit was based on 12 models which are trained on either CQT or Spec datasets.\n\nCQT - 4 ecaresnext50t_32x4d models trained different sized CQT datasets\n\n```\nl_1_CQT = inferenceCQT(_model = 'ecaresnext50t_32x4d',modelpth='pog2-ecaresnext50t-32x4d-cqt-512-64')\nl_2_CQT = inferenceCQT(_model = 'ecaresnext50t_32x4d',modelpth='pog2-ecaresnext50t-32x4d-cqt-512-64-mixup')\nl_3_CQT = inferenceCQT(_model = 'ecaresnext50t_32x4d',modelpth='pog2-ecaresnext50t-32x4d-cqt-448-84-tune')\nl_4_CQT = inferenceCQT(_model = 'ecaresnext50t_32x4d',modelpth='pog2-ecaresnext50t-32x4d-cqt-512-64-noisy')\n\n```\n\nSpec - 8 models on different sized spec datasets\n\n```\nl_1 = inference(_model='resnest50d',modelpth='pog2-resnest50d-monospec-hop448-mels160-ftune')\nl_2 = inference(_model='resnest50d',modelpth='pog2-resnest50d-monospec-hop448-mels160-noisytrain')\nl_3 = inference(_model = 'ecaresnext50t_32x4d',modelpth='pog2-ecaresnext50t-32x4d-monospec-hop448-mels160')\nl_4 = inference(_model = 'resnext50d_32x4d',modelpth='pog2-resnext50d-32x4d-monospec-hop448-mels160')\nl_5 = inference(_model = 'ecaresnext50t_32x4d',modelpth='pog2-ecaresnext50t-32x4d-monospec-hop448-mels64',sz=[64])\nl_6 = inference(_model = 'ecaresnext50t_32x4d',modelpth='pog2-ecaresnext50t-32x4d-spec-hop448-mels64-tune',sz=[64])\nl_7 = inference(_model = 'ecaresnext50t_32x4d',modelpth='pog2-ecaresnext50t-32x4d-spec-hop448-mels64-noisy',sz=[64])\nl_8 = inference(_model = 'resnext50d_32x4d',modelpth='pog2-resnext50t-32x4d-spec-hop448-mels64-noisy',sz=[64])\n```\n\nInference pipeline is available at: https://www.kaggle.com/code/pheadrus/inference-pipeline-v0?scriptVersionId=93716687\n\n Thanks",
      "votes": null
    },
    {
      "id": "1775016",
      "postDate": "05/02/2022 16:41:13",
      "content": "<p>Thanks for sharing your solution</p>",
      "rawMarkdown": "Thanks for sharing your solution",
      "votes": null
    },
    {
      "id": "1775057",
      "postDate": "05/02/2022 17:12:51",
      "content": "<p>Hi Phaedra's, your dataset is great, interesting also. I will want you to tell me more about dataset because I have been listening  to video and audio about data set but I have not created even one dataset. so tell me more about dataset for me to understand it very well so I can create one for myself. thank you and remain blessed</p>",
      "rawMarkdown": "Hi Phaedra's, your dataset is great, interesting also. I will want you to tell me more about dataset because I have been listening  to video and audio about data set but I have not created even one dataset. so tell me more about dataset for me to understand it very well so I can create one for myself. thank you and remain blessed",
      "votes": null
    },
    {
      "id": "1775241",
      "postDate": "05/02/2022 20:19:42",
      "content": "<p>Thanks for sharing <a href=\"https://www.kaggle.com/pheadrus\" target=\"_blank\">@pheadrus</a>. Wow, we end up trying many common techniques. I need to take a look at what CQT is</p>",
      "rawMarkdown": "Thanks for sharing @pheadrus. Wow, we end up trying many common techniques. I need to take a look at what CQT is",
      "votes": null
    },
    {
      "id": "1775331",
      "postDate": "05/02/2022 22:37:50",
      "content": "<p>This is a really nicely written and informative post. Well done!</p>",
      "rawMarkdown": "This is a really nicely written and informative post. Well done!",
      "votes": null
    },
    {
      "id": "1775346",
      "postDate": "05/02/2022 23:01:39",
      "content": "<p><a href=\"https://www.kaggle.com/pheadrus\" target=\"_blank\">@pheadrus</a> Thank you for sharing this, it is very interesting. To be honest I have no clue how to do this, I didn't think it was possible to create a machine learning algorithm capable of predicting music genre. </p>",
      "rawMarkdown": "pheadrus Thank you for sharing this, it is very interesting. To be honest I have no clue how to do this, I didn't think it was possible to create a machine learning algorithm capable of predicting music genre.",
      "votes": null
    },
    {
      "id": "1777008",
      "postDate": "05/04/2022 10:53:26",
      "content": "<p>Its quite possible. :) </p>",
      "rawMarkdown": "Its quite possible. :)",
      "votes": null
    },
    {
      "id": "1777192",
      "postDate": "05/04/2022 13:22:47",
      "content": "<p>Awesome work <a href=\"https://www.kaggle.com/pheadrus\" target=\"_blank\">@pheadrus</a>! It was fun watching how strong your submissions were throughout the competition. This is a great writeup and there is a lot to learn from it.</p>",
      "rawMarkdown": "Awesome work @pheadrus! It was fun watching how strong your submissions were throughout the competition. This is a great writeup and there is a lot to learn from it.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1775016,
      "author_name": "kurianbenoy",
      "author_url": "",
      "post_date": "05/02/2022 16:41:13",
      "content": "<p>Thanks for sharing your solution</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1775057,
      "author_name": "nedolisachinyere",
      "author_url": "",
      "post_date": "05/02/2022 17:12:51",
      "content": "<p>Hi Phaedra's, your dataset is great, interesting also. I will want you to tell me more about dataset because I have been listening  to video and audio about data set but I have not created even one dataset. so tell me more about dataset for me to understand it very well so I can create one for myself. thank you and remain blessed</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1775241,
      "author_name": "dienhoa",
      "author_url": "",
      "post_date": "05/02/2022 20:19:42",
      "content": "<p>Thanks for sharing <a href=\"https://www.kaggle.com/pheadrus\" target=\"_blank\">@pheadrus</a>. Wow, we end up trying many common techniques. I need to take a look at what CQT is</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1775331,
      "author_name": "",
      "author_url": "",
      "post_date": "05/02/2022 22:37:50",
      "content": "<p>This is a really nicely written and informative post. Well done!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1775346,
      "author_name": "alanjo",
      "author_url": "",
      "post_date": "05/02/2022 23:01:39",
      "content": "<p><a href=\"https://www.kaggle.com/pheadrus\" target=\"_blank\">@pheadrus</a> Thank you for sharing this, it is very interesting. To be honest I have no clue how to do this, I didn't think it was possible to create a machine learning algorithm capable of predicting music genre. </p>",
      "votes": null,
      "replies": [
        {
          "id": 1777008,
          "author_name": "pheadrus",
          "author_url": "",
          "post_date": "05/04/2022 10:53:26",
          "content": "<p>Its quite possible. :) </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1777192,
      "author_name": "robikscube",
      "author_url": "",
      "post_date": "05/04/2022 13:22:47",
      "content": "<p>Awesome work <a href=\"https://www.kaggle.com/pheadrus\" target=\"_blank\">@pheadrus</a>! It was fun watching how strong your submissions were throughout the competition. This is a great writeup and there is a lot to learn from it.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1774684": "Hello everyone.\n\nThanks for @robikscube for hosting this awesome competition. It was fun for to me tune my signal processing skills. \n\n**My approach to this competition:**\n\n**Datasets:** \n\nI resampled to original audio files to 16k frequency and then created the following datasets for CNNs.\n\n**Mel Spec**: I created 3 differently sized spec datasets for this competition. Appropriate min-max normalisation was done for each spectrogram \n\n| Hop Len |N Mels  |\n| --- | --- |\n| 512 | 128 |\n| 512 | 64 |\n| 448 | 160 |\n\n**CQT**: I also experimented with CQT (Constant Q transform) datasets. It have a good boost in local ensemble CV.\n\n| Hop Len |N Bins  |\n| --- | --- |\n| 512 | 64 |\n| 448 | 84 |\n\n**Validation**\n\nValidation was based on 5 fold mode ensemble. Mean logit/prob based ensemble wasn't corresponding as well with the LB.\n\n**Framework**\nI primarily used fastai for this competition as I had a good fastai pipeline set from a previous signal processing competition.\n\n \n**Things that worked - Training**\n\n- Reduced stride for 1st conv layers\n- Mixup (0.25)\n- Labelsmoothing (0.025)\n- Noisy student (re) training\n- Ranger, and Adam optimiser depending on model\n- fit_flat_cos for normal training and fit_one_cycle for noisy training\n\n**Things that did not work - Training**\n- Larger models\n- Pseudo Labels\n- Transformers (ViT, Deit)\n\n**Models**\n\nMy final submit was based on 12 models which are trained on either CQT or Spec datasets.\n\nCQT - 4 ecaresnext50t_32x4d models trained different sized CQT datasets\n\n```\nl_1_CQT = inferenceCQT(_model = 'ecaresnext50t_32x4d',modelpth='pog2-ecaresnext50t-32x4d-cqt-512-64')\nl_2_CQT = inferenceCQT(_model = 'ecaresnext50t_32x4d',modelpth='pog2-ecaresnext50t-32x4d-cqt-512-64-mixup')\nl_3_CQT = inferenceCQT(_model = 'ecaresnext50t_32x4d',modelpth='pog2-ecaresnext50t-32x4d-cqt-448-84-tune')\nl_4_CQT = inferenceCQT(_model = 'ecaresnext50t_32x4d',modelpth='pog2-ecaresnext50t-32x4d-cqt-512-64-noisy')\n\n```\n\nSpec - 8 models on different sized spec datasets\n\n```\nl_1 = inference(_model='resnest50d',modelpth='pog2-resnest50d-monospec-hop448-mels160-ftune')\nl_2 = inference(_model='resnest50d',modelpth='pog2-resnest50d-monospec-hop448-mels160-noisytrain')\nl_3 = inference(_model = 'ecaresnext50t_32x4d',modelpth='pog2-ecaresnext50t-32x4d-monospec-hop448-mels160')\nl_4 = inference(_model = 'resnext50d_32x4d',modelpth='pog2-resnext50d-32x4d-monospec-hop448-mels160')\nl_5 = inference(_model = 'ecaresnext50t_32x4d',modelpth='pog2-ecaresnext50t-32x4d-monospec-hop448-mels64',sz=[64])\nl_6 = inference(_model = 'ecaresnext50t_32x4d',modelpth='pog2-ecaresnext50t-32x4d-spec-hop448-mels64-tune',sz=[64])\nl_7 = inference(_model = 'ecaresnext50t_32x4d',modelpth='pog2-ecaresnext50t-32x4d-spec-hop448-mels64-noisy',sz=[64])\nl_8 = inference(_model = 'resnext50d_32x4d',modelpth='pog2-resnext50t-32x4d-spec-hop448-mels64-noisy',sz=[64])\n```\n\nInference pipeline is available at: https://www.kaggle.com/code/pheadrus/inference-pipeline-v0?scriptVersionId=93716687\n\n Thanks",
    "1775016": "Thanks for sharing your solution",
    "1775057": "Hi Phaedra's, your dataset is great, interesting also. I will want you to tell me more about dataset because I have been listening  to video and audio about data set but I have not created even one dataset. so tell me more about dataset for me to understand it very well so I can create one for myself. thank you and remain blessed",
    "1775241": "Thanks for sharing @pheadrus. Wow, we end up trying many common techniques. I need to take a look at what CQT is",
    "1775331": "This is a really nicely written and informative post. Well done!",
    "1775346": "pheadrus Thank you for sharing this, it is very interesting. To be honest I have no clue how to do this, I didn't think it was possible to create a machine learning algorithm capable of predicting music genre.",
    "1777008": "Its quite possible. :)",
    "1777192": "Awesome work @pheadrus! It was fun watching how strong your submissions were throughout the competition. This is a great writeup and there is a lot to learn from it."
  },
  "source": "meta"
}