{
  "id": 321281,
  "title": "Quick Summary for #1 solution",
  "url": "/competitions/kaggle-pog-series-s01e02/writeups/dhoa-quick-summary-for-1-solution",
  "author_name": "",
  "post_date": "2022-04-26T18:05:10.287Z",
  "votes": 23,
  "comment_count": 12,
  "views": 0,
  "content": "<p>Hi everyone. </p>\n<p>Again thank you all for this opportunity, to learn and share in our data journey.</p>\n<p>I write here a quick summary of what I did for this competition.</p>\n<p>The inference notebook: <a href=\"https://www.kaggle.com/code/dienhoa/inference-submission-music-genre\" target=\"_blank\">https://www.kaggle.com/code/dienhoa/inference-submission-music-genre</a></p>\n<p>For training part, I trained locally and just uploaded here: <a href=\"https://www.kaggle.com/code/dienhoa/music-genre-resnext-mixup-kfold-mish-spectro\" target=\"_blank\">https://www.kaggle.com/code/dienhoa/music-genre-resnext-mixup-kfold-mish-spectro</a><br>\n(If you want to play with, please make sure you attach the data I mentioned below, and change the path in the notebook)</p>\n<p>The dataset, it is just the spectrogram computing from the original one: <a href=\"https://www.kaggle.com/datasets/dienhoa/music-genre-spectrogram-pogchamps\" target=\"_blank\">https://www.kaggle.com/datasets/dienhoa/music-genre-spectrogram-pogchamps</a></p>\n<p>If you see something confused, please tell me, maybe I forgot something after trying a lot of stuffs :D</p>\n<p>I used fast.ai and used a lot of tricks I got from Imagenette <a href=\"https://github.com/fastai/imagenette\" target=\"_blank\">https://github.com/fastai/imagenette</a> ( a simplified version of Imagenet - a competition to train model from scratch ).</p>\n<ul>\n<li>Crop randomly ~ 5s of audio (224 px in the spectrogram)</li>\n<li>7 folds xse_resnext50 </li>\n<li>simple computer vision augmentation: Reflection, Brightness, …</li>\n<li>MixUp</li>\n<li>LabelSmoothing</li>\n<li>Replace every MaxPool to BlurMaxPool (<a href=\"https://arxiv.org/pdf/1904.11486.pdf\" target=\"_blank\">https://arxiv.org/pdf/1904.11486.pdf</a>)</li>\n<li>fast.ai fit_flat_cos Learning Rate Scheduler</li>\n<li>Mish as activation function</li>\n<li>Self Attention</li>\n<li>Ranger Optimizer</li>\n<li>Stratified Kfolds - stats.mode the results</li>\n<li>Test Time Augmentation</li>\n</ul>\n<p>All of these tricks I got from the LeaderBoard of Imagenette: <a href=\"https://github.com/fastai/imagenette\" target=\"_blank\">https://github.com/fastai/imagenette</a></p>\n<p>If it is not clear I will try to write a details version.</p>\n<p>Thanks again everyone </p>",
  "messages": [
    {
      "id": "1768193",
      "postDate": "04/26/2022 04:06:22",
      "content": "<p>Hi everyone. </p>\n<p>Again thank you all for this opportunity, to learn and share in our data journey.</p>\n<p>I write here a quick summary of what I did for this competition.</p>\n<p>The inference notebook: <a href=\"https://www.kaggle.com/code/dienhoa/inference-submission-music-genre\" target=\"_blank\">https://www.kaggle.com/code/dienhoa/inference-submission-music-genre</a></p>\n<p>For training part, I trained locally and just uploaded here: <a href=\"https://www.kaggle.com/code/dienhoa/music-genre-resnext-mixup-kfold-mish-spectro\" target=\"_blank\">https://www.kaggle.com/code/dienhoa/music-genre-resnext-mixup-kfold-mish-spectro</a><br>\n(If you want to play with, please make sure you attach the data I mentioned below, and change the path in the notebook)</p>\n<p>The dataset, it is just the spectrogram computing from the original one: <a href=\"https://www.kaggle.com/datasets/dienhoa/music-genre-spectrogram-pogchamps\" target=\"_blank\">https://www.kaggle.com/datasets/dienhoa/music-genre-spectrogram-pogchamps</a></p>\n<p>If you see something confused, please tell me, maybe I forgot something after trying a lot of stuffs :D</p>\n<p>I used fast.ai and used a lot of tricks I got from Imagenette <a href=\"https://github.com/fastai/imagenette\" target=\"_blank\">https://github.com/fastai/imagenette</a> ( a simplified version of Imagenet - a competition to train model from scratch ).</p>\n<ul>\n<li>Crop randomly ~ 5s of audio (224 px in the spectrogram)</li>\n<li>7 folds xse_resnext50 </li>\n<li>simple computer vision augmentation: Reflection, Brightness, …</li>\n<li>MixUp</li>\n<li>LabelSmoothing</li>\n<li>Replace every MaxPool to BlurMaxPool (<a href=\"https://arxiv.org/pdf/1904.11486.pdf\" target=\"_blank\">https://arxiv.org/pdf/1904.11486.pdf</a>)</li>\n<li>fast.ai fit_flat_cos Learning Rate Scheduler</li>\n<li>Mish as activation function</li>\n<li>Self Attention</li>\n<li>Ranger Optimizer</li>\n<li>Stratified Kfolds - stats.mode the results</li>\n<li>Test Time Augmentation</li>\n</ul>\n<p>All of these tricks I got from the LeaderBoard of Imagenette: <a href=\"https://github.com/fastai/imagenette\" target=\"_blank\">https://github.com/fastai/imagenette</a></p>\n<p>If it is not clear I will try to write a details version.</p>\n<p>Thanks again everyone </p>",
      "rawMarkdown": "Hi everyone. \n\nAgain thank you all for this opportunity, to learn and share in our data journey.\n\nI write here a quick summary of what I did for this competition.\n\nThe inference notebook: https://www.kaggle.com/code/dienhoa/inference-submission-music-genre\n\nFor training part, I trained locally and just uploaded here: https://www.kaggle.com/code/dienhoa/music-genre-resnext-mixup-kfold-mish-spectro\n(If you want to play with, please make sure you attach the data I mentioned below, and change the path in the notebook)\n\nThe dataset, it is just the spectrogram computing from the original one: https://www.kaggle.com/datasets/dienhoa/music-genre-spectrogram-pogchamps\n\nIf you see something confused, please tell me, maybe I forgot something after trying a lot of stuffs :D\n\nI used fast.ai and used a lot of tricks I got from Imagenette https://github.com/fastai/imagenette ( a simplified version of Imagenet - a competition to train model from scratch ).\n- Crop randomly ~ 5s of audio (224 px in the spectrogram)\n- 7 folds xse_resnext50 \n- simple computer vision augmentation: Reflection, Brightness, ...\n- MixUp\n- LabelSmoothing\n- Replace every MaxPool to BlurMaxPool (https://arxiv.org/pdf/1904.11486.pdf)\n- fast.ai fit_flat_cos Learning Rate Scheduler\n- Mish as activation function\n- Self Attention\n- Ranger Optimizer\n- Stratified Kfolds - stats.mode the results\n- Test Time Augmentation\n\nAll of these tricks I got from the LeaderBoard of Imagenette: https://github.com/fastai/imagenette\n\nIf it is not clear I will try to write a details version.\n\nThanks again everyone",
      "votes": null
    },
    {
      "id": "1768197",
      "postDate": "04/26/2022 04:09:51",
      "content": "<p>Congratulations <a href=\"https://www.kaggle.com/dienhoa\" target=\"_blank\">@dienhoa</a> ! Well done on the 1st place finish. </p>",
      "rawMarkdown": "Congratulations @dienhoa ! Well done on the 1st place finish.",
      "votes": null
    },
    {
      "id": "1768205",
      "postDate": "04/26/2022 04:17:50",
      "content": "<p>Thanks <a href=\"https://www.kaggle.com/pheadrus\" target=\"_blank\">@pheadrus</a> I learned a lot from you during this competition  </p>",
      "rawMarkdown": "Thanks @pheadrus I learned a lot from you during this competition",
      "votes": null
    },
    {
      "id": "1768700",
      "postDate": "04/26/2022 15:17:32",
      "content": "<p>Nice work and huge feat to finish 1st place. Thanks for sharing your learnings on how you did that too!</p>",
      "rawMarkdown": "Nice work and huge feat to finish 1st place. Thanks for sharing your learnings on how you did that too!",
      "votes": null
    },
    {
      "id": "1768710",
      "postDate": "04/26/2022 15:32:10",
      "content": "<p>Congrats on 1st place. May I ask what your best single model score was?</p>\n<p>I see most of the tricks in the code, but I am not that familiar with self attention. Do the xse_resnext models have self attention already? </p>\n<p>And am I assuming correctly that the test-time augmentation runs the test dataset through the same augmentations used during training - 50 random crops with random transforms - or are the test crops/transforms handled differently?</p>\n<p>Thanks for the writeup and code :)</p>",
      "rawMarkdown": "Congrats on 1st place. May I ask what your best single model score was?\n\nI see most of the tricks in the code, but I am not that familiar with self attention. Do the xse_resnext models have self attention already? \n\nAnd am I assuming correctly that the test-time augmentation runs the test dataset through the same augmentations used during training - 50 random crops with random transforms - or are the test crops/transforms handled differently?\n\nThanks for the writeup and code :)",
      "votes": null
    },
    {
      "id": "1768869",
      "postDate": "04/26/2022 18:19:07",
      "content": "<p>Thanks <a href=\"https://www.kaggle.com/pcmiller\" target=\"_blank\">@pcmiller</a> . If you mean best single model = (no fold, no TTA) so I don't know because after some very early models, I've seen a big boost with TTA and always used it later. My model predicts with a crop ~ 5 seconds so I think it's more logic using TTA so I can take into account the whole music file.</p>\n<p>Actually, after verifying my best model, it is xse_resnext50 with 7 folds. I think many kagglers here count it as a single model :D</p>\n<p>I also don't know if the original xse_resnext (in the paper) have self attention or not or fastai add it. But in fastai I can activate it easily with <code>sa=1</code> in <code>xse_resnext101(n_out=19, act_cls=Mish, sa=1, pool=MaxPool, pretrained=False)</code>. More details you can find here: <a href=\"https://docs.fast.ai/layers.html#SelfAttention\" target=\"_blank\">https://docs.fast.ai/layers.html#SelfAttention</a></p>\n<p>For TTA, the same augmentations for both training set and test set. I think I just did cropping, reflection, changing brightness so it doesn't hurt too much the test set</p>\n<p>Hope it clears :)</p>",
      "rawMarkdown": "Thanks @pcmiller . If you mean best single model = (no fold, no TTA) so I don't know because after some very early models, I've seen a big boost with TTA and always used it later. My model predicts with a crop ~ 5 seconds so I think it's more logic using TTA so I can take into account the whole music file.\n\nActually, after verifying my best model, it is xse_resnext50 with 7 folds. I think many kagglers here count it as a single model :D\n\nI also don't know if the original xse_resnext (in the paper) have self attention or not or fastai add it. But in fastai I can activate it easily with `sa=1` in `xse_resnext101(n_out=19, act_cls=Mish, sa=1, pool=MaxPool, pretrained=False)`. More details you can find here: https://docs.fast.ai/layers.html#SelfAttention\n\nFor TTA, the same augmentations for both training set and test set. I think I just did cropping, reflection, changing brightness so it doesn't hurt too much the test set\n\nHope it clears :)",
      "votes": null
    },
    {
      "id": "1768978",
      "postDate": "04/26/2022 20:16:21",
      "content": "<p>Thanks for the response :)</p>",
      "rawMarkdown": "Thanks for the response :)",
      "votes": null
    },
    {
      "id": "1773263",
      "postDate": "05/01/2022 01:50:01",
      "content": "<p>Congratulations <a href=\"https://www.kaggle.com/dienhoa\" target=\"_blank\">@dienhoa</a> on fully deserved 1st place! A couple of questions if I may: <br>\n(1) What was your score LB for your pipeline without that maxBlurPool? <br>\n(2) What was your LB without smoothing labels (but with mixup)? I noticed, in my cases, it should be one or the other but not both (either smoothing or mixup).<br>\n(3) You used folds voting. What was just the mean of folds prediction probs/logits? In my case mean was always better (per model)</p>",
      "rawMarkdown": "Congratulations @dienhoa on fully deserved 1st place! A couple of questions if I may: \n(1) What was your score LB for your pipeline without that maxBlurPool? \n(2) What was your LB without smoothing labels (but with mixup)? I noticed, in my cases, it should be one or the other but not both (either smoothing or mixup).\n(3) You used folds voting. What was just the mean of folds prediction probs/logits? In my case mean was always better (per model)",
      "votes": null
    },
    {
      "id": "1773387",
      "postDate": "05/01/2022 05:42:00",
      "content": "<p>Congratulations <a href=\"https://www.kaggle.com/dienhoa\" target=\"_blank\">@dienhoa</a> <br>\nThanks for sharing</p>",
      "rawMarkdown": "Congratulations @dienhoa \nThanks for sharing",
      "votes": null
    },
    {
      "id": "1774036",
      "postDate": "05/01/2022 18:14:38",
      "content": "<p>Thanks <a href=\"https://www.kaggle.com/dmitrykonovalov\" target=\"_blank\">@dmitrykonovalov</a> . It's hard to answer because I didn't carefully log down everything during different experiments, I definitely need to journal for the next project (maybe with WandB or just a text file). And also, I exclusively tried a lot of TTA in the 2 final submissions to get a better score so the score from the previous submission is not very good.  </p>\n<ol>\n<li>If I remember correctly, with maxBlurPool my CV improved from 0.59x to 0.60x and the model converge faster (40 epochs instead of 50 epochs w/o maxBlurPool)</li>\n<li>These 2 (MixUp and LabelSmoothing) I always use together so I don't know if trying just one can improve the score or not.</li>\n<li>It's not actually folds voting. I have 7 folds and each folds predicts 50 different 5s-slices, so 350 predictions for voting. Because the model runs on a slice of just 5 seconds audio so I was thinking that maybe there are some slices that not relevant at all in the whole music, and can hurt the prediction if I use mean average. I did not test the submission's score with mean-avg</li>\n</ol>",
      "rawMarkdown": "Thanks @dmitrykonovalov . It's hard to answer because I didn't carefully log down everything during different experiments, I definitely need to journal for the next project (maybe with WandB or just a text file). And also, I exclusively tried a lot of TTA in the 2 final submissions to get a better score so the score from the previous submission is not very good.  \n\n1. If I remember correctly, with maxBlurPool my CV improved from 0.59x to 0.60x and the model converge faster (40 epochs instead of 50 epochs w/o maxBlurPool)\n2. These 2 (MixUp and LabelSmoothing) I always use together so I don't know if trying just one can improve the score or not.\n3. It's not actually folds voting. I have 7 folds and each folds predicts 50 different 5s-slices, so 350 predictions for voting. Because the model runs on a slice of just 5 seconds audio so I was thinking that maybe there are some slices that not relevant at all in the whole music, and can hurt the prediction if I use mean average. I did not test the submission's score with mean-avg",
      "votes": null
    },
    {
      "id": "1774132",
      "postDate": "05/01/2022 21:19:06",
      "content": "<p>Thank you and see you in a next sound competition :)</p>",
      "rawMarkdown": "Thank you and see you in a next sound competition :)",
      "votes": null
    },
    {
      "id": "1781628",
      "postDate": "05/08/2022 18:23:44",
      "content": "<p>Congratulations <a href=\"https://www.kaggle.com/dienhoa\" target=\"_blank\">@dienhoa</a>.</p>\n<p>I was trying to load your models standalone by installing fastai, korina and downloading data from here:</p>\n<p><a href=\"https://www.kaggle.com/datasets/dienhoa/inference-music-genre\" target=\"_blank\">https://www.kaggle.com/datasets/dienhoa/inference-music-genre</a></p>\n<p>The objective was to create an inference pipeline which takes a audio as input, and outputs the predictions of your models. What are steps required to create such a pipeline?</p>\n<p>When I am trying in standalone loading models, I am getting the below error. </p>\n<pre><code>File /opt/conda/lib/python3.8/site-packages/torch/serialization.py:1039, in _load.&lt;locals&gt;.UnpicklerWrapper.find_class(self, mod_name, name)\n   1037         pass\n   1038 mod_name = load_module_mapping.get(mod_name, mod_name)\n-&gt; 1039 return super().find_class(mod_name, name)\n\nModuleNotFoundError: No module named 'kornia.contrib.max_blur_pool'\n</code></pre>\n<p>Any clues on debugging this issue, and your thoughts on creating a inference pipeline will be highly appreciated.</p>\n<p>~ Kurian</p>",
      "rawMarkdown": "Congratulations @dienhoa.\n\nI was trying to load your models standalone by installing fastai, korina and downloading data from here:\n\nhttps://www.kaggle.com/datasets/dienhoa/inference-music-genre\n\nThe objective was to create an inference pipeline which takes a audio as input, and outputs the predictions of your models. What are steps required to create such a pipeline?\n\nWhen I am trying in standalone loading models, I am getting the below error. \n```\nFile /opt/conda/lib/python3.8/site-packages/torch/serialization.py:1039, in _load.<locals>.UnpicklerWrapper.find_class(self, mod_name, name)\n   1037         pass\n   1038 mod_name = load_module_mapping.get(mod_name, mod_name)\n-> 1039 return super().find_class(mod_name, name)\n\nModuleNotFoundError: No module named 'kornia.contrib.max_blur_pool'\n```\n\nAny clues on debugging this issue, and your thoughts on creating a inference pipeline will be highly appreciated.\n\n~ Kurian",
      "votes": null
    },
    {
      "id": "1781677",
      "postDate": "05/08/2022 20:16:26",
      "content": "<p>I think you get this problem when you install the latest version of Kornia (0.6.4). The code is ok with Kornia Kaggle version (0.5.8).</p>\n<p>I think <code>kornia.contrib.max_blur_pool</code> is replaced by <code>kornia.filters.max_blur_pool2d</code> in the latter version, however when I tried to use it, I got a type error, I don't remember exactly but something like Kornia doesn't accept fastai TensorImage as Input.</p>\n<p>Hope it helps and let me know if you have something news ;)</p>",
      "rawMarkdown": "I think you get this problem when you install the latest version of Kornia (0.6.4). The code is ok with Kornia Kaggle version (0.5.8).\n\nI think `kornia.contrib.max_blur_pool` is replaced by `kornia.filters.max_blur_pool2d` in the latter version, however when I tried to use it, I got a type error, I don't remember exactly but something like Kornia doesn't accept fastai TensorImage as Input.\n\nHope it helps and let me know if you have something news ;)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1768197,
      "author_name": "pheadrus",
      "author_url": "",
      "post_date": "04/26/2022 04:09:51",
      "content": "<p>Congratulations <a href=\"https://www.kaggle.com/dienhoa\" target=\"_blank\">@dienhoa</a> ! Well done on the 1st place finish. </p>",
      "votes": null,
      "replies": [
        {
          "id": 1768205,
          "author_name": "dienhoa",
          "author_url": "",
          "post_date": "04/26/2022 04:17:50",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/pheadrus\" target=\"_blank\">@pheadrus</a> I learned a lot from you during this competition  </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1768700,
      "author_name": "kevinkwan",
      "author_url": "",
      "post_date": "04/26/2022 15:17:32",
      "content": "<p>Nice work and huge feat to finish 1st place. Thanks for sharing your learnings on how you did that too!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1768710,
      "author_name": "pcmiller",
      "author_url": "",
      "post_date": "04/26/2022 15:32:10",
      "content": "<p>Congrats on 1st place. May I ask what your best single model score was?</p>\n<p>I see most of the tricks in the code, but I am not that familiar with self attention. Do the xse_resnext models have self attention already? </p>\n<p>And am I assuming correctly that the test-time augmentation runs the test dataset through the same augmentations used during training - 50 random crops with random transforms - or are the test crops/transforms handled differently?</p>\n<p>Thanks for the writeup and code :)</p>",
      "votes": null,
      "replies": [
        {
          "id": 1768869,
          "author_name": "dienhoa",
          "author_url": "",
          "post_date": "04/26/2022 18:19:07",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/pcmiller\" target=\"_blank\">@pcmiller</a> . If you mean best single model = (no fold, no TTA) so I don't know because after some very early models, I've seen a big boost with TTA and always used it later. My model predicts with a crop ~ 5 seconds so I think it's more logic using TTA so I can take into account the whole music file.</p>\n<p>Actually, after verifying my best model, it is xse_resnext50 with 7 folds. I think many kagglers here count it as a single model :D</p>\n<p>I also don't know if the original xse_resnext (in the paper) have self attention or not or fastai add it. But in fastai I can activate it easily with <code>sa=1</code> in <code>xse_resnext101(n_out=19, act_cls=Mish, sa=1, pool=MaxPool, pretrained=False)</code>. More details you can find here: <a href=\"https://docs.fast.ai/layers.html#SelfAttention\" target=\"_blank\">https://docs.fast.ai/layers.html#SelfAttention</a></p>\n<p>For TTA, the same augmentations for both training set and test set. I think I just did cropping, reflection, changing brightness so it doesn't hurt too much the test set</p>\n<p>Hope it clears :)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1768978,
          "author_name": "pcmiller",
          "author_url": "",
          "post_date": "04/26/2022 20:16:21",
          "content": "<p>Thanks for the response :)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1773263,
      "author_name": "dmitrykonovalov",
      "author_url": "",
      "post_date": "05/01/2022 01:50:01",
      "content": "<p>Congratulations <a href=\"https://www.kaggle.com/dienhoa\" target=\"_blank\">@dienhoa</a> on fully deserved 1st place! A couple of questions if I may: <br>\n(1) What was your score LB for your pipeline without that maxBlurPool? <br>\n(2) What was your LB without smoothing labels (but with mixup)? I noticed, in my cases, it should be one or the other but not both (either smoothing or mixup).<br>\n(3) You used folds voting. What was just the mean of folds prediction probs/logits? In my case mean was always better (per model)</p>",
      "votes": null,
      "replies": [
        {
          "id": 1774036,
          "author_name": "dienhoa",
          "author_url": "",
          "post_date": "05/01/2022 18:14:38",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/dmitrykonovalov\" target=\"_blank\">@dmitrykonovalov</a> . It's hard to answer because I didn't carefully log down everything during different experiments, I definitely need to journal for the next project (maybe with WandB or just a text file). And also, I exclusively tried a lot of TTA in the 2 final submissions to get a better score so the score from the previous submission is not very good.  </p>\n<ol>\n<li>If I remember correctly, with maxBlurPool my CV improved from 0.59x to 0.60x and the model converge faster (40 epochs instead of 50 epochs w/o maxBlurPool)</li>\n<li>These 2 (MixUp and LabelSmoothing) I always use together so I don't know if trying just one can improve the score or not.</li>\n<li>It's not actually folds voting. I have 7 folds and each folds predicts 50 different 5s-slices, so 350 predictions for voting. Because the model runs on a slice of just 5 seconds audio so I was thinking that maybe there are some slices that not relevant at all in the whole music, and can hurt the prediction if I use mean average. I did not test the submission's score with mean-avg</li>\n</ol>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1774132,
          "author_name": "dmitrykonovalov",
          "author_url": "",
          "post_date": "05/01/2022 21:19:06",
          "content": "<p>Thank you and see you in a next sound competition :)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1773387,
      "author_name": "amritharj",
      "author_url": "",
      "post_date": "05/01/2022 05:42:00",
      "content": "<p>Congratulations <a href=\"https://www.kaggle.com/dienhoa\" target=\"_blank\">@dienhoa</a> <br>\nThanks for sharing</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1781628,
      "author_name": "kurianbenoy",
      "author_url": "",
      "post_date": "05/08/2022 18:23:44",
      "content": "<p>Congratulations <a href=\"https://www.kaggle.com/dienhoa\" target=\"_blank\">@dienhoa</a>.</p>\n<p>I was trying to load your models standalone by installing fastai, korina and downloading data from here:</p>\n<p><a href=\"https://www.kaggle.com/datasets/dienhoa/inference-music-genre\" target=\"_blank\">https://www.kaggle.com/datasets/dienhoa/inference-music-genre</a></p>\n<p>The objective was to create an inference pipeline which takes a audio as input, and outputs the predictions of your models. What are steps required to create such a pipeline?</p>\n<p>When I am trying in standalone loading models, I am getting the below error. </p>\n<pre><code>File /opt/conda/lib/python3.8/site-packages/torch/serialization.py:1039, in _load.&lt;locals&gt;.UnpicklerWrapper.find_class(self, mod_name, name)\n   1037         pass\n   1038 mod_name = load_module_mapping.get(mod_name, mod_name)\n-&gt; 1039 return super().find_class(mod_name, name)\n\nModuleNotFoundError: No module named 'kornia.contrib.max_blur_pool'\n</code></pre>\n<p>Any clues on debugging this issue, and your thoughts on creating a inference pipeline will be highly appreciated.</p>\n<p>~ Kurian</p>",
      "votes": null,
      "replies": [
        {
          "id": 1781677,
          "author_name": "dienhoa",
          "author_url": "",
          "post_date": "05/08/2022 20:16:26",
          "content": "<p>I think you get this problem when you install the latest version of Kornia (0.6.4). The code is ok with Kornia Kaggle version (0.5.8).</p>\n<p>I think <code>kornia.contrib.max_blur_pool</code> is replaced by <code>kornia.filters.max_blur_pool2d</code> in the latter version, however when I tried to use it, I got a type error, I don't remember exactly but something like Kornia doesn't accept fastai TensorImage as Input.</p>\n<p>Hope it helps and let me know if you have something news ;)</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1768193": "Hi everyone. \n\nAgain thank you all for this opportunity, to learn and share in our data journey.\n\nI write here a quick summary of what I did for this competition.\n\nThe inference notebook: https://www.kaggle.com/code/dienhoa/inference-submission-music-genre\n\nFor training part, I trained locally and just uploaded here: https://www.kaggle.com/code/dienhoa/music-genre-resnext-mixup-kfold-mish-spectro\n(If you want to play with, please make sure you attach the data I mentioned below, and change the path in the notebook)\n\nThe dataset, it is just the spectrogram computing from the original one: https://www.kaggle.com/datasets/dienhoa/music-genre-spectrogram-pogchamps\n\nIf you see something confused, please tell me, maybe I forgot something after trying a lot of stuffs :D\n\nI used fast.ai and used a lot of tricks I got from Imagenette https://github.com/fastai/imagenette ( a simplified version of Imagenet - a competition to train model from scratch ).\n- Crop randomly ~ 5s of audio (224 px in the spectrogram)\n- 7 folds xse_resnext50 \n- simple computer vision augmentation: Reflection, Brightness, ...\n- MixUp\n- LabelSmoothing\n- Replace every MaxPool to BlurMaxPool (https://arxiv.org/pdf/1904.11486.pdf)\n- fast.ai fit_flat_cos Learning Rate Scheduler\n- Mish as activation function\n- Self Attention\n- Ranger Optimizer\n- Stratified Kfolds - stats.mode the results\n- Test Time Augmentation\n\nAll of these tricks I got from the LeaderBoard of Imagenette: https://github.com/fastai/imagenette\n\nIf it is not clear I will try to write a details version.\n\nThanks again everyone",
    "1768197": "Congratulations @dienhoa ! Well done on the 1st place finish.",
    "1768205": "Thanks @pheadrus I learned a lot from you during this competition",
    "1768700": "Nice work and huge feat to finish 1st place. Thanks for sharing your learnings on how you did that too!",
    "1768710": "Congrats on 1st place. May I ask what your best single model score was?\n\nI see most of the tricks in the code, but I am not that familiar with self attention. Do the xse_resnext models have self attention already? \n\nAnd am I assuming correctly that the test-time augmentation runs the test dataset through the same augmentations used during training - 50 random crops with random transforms - or are the test crops/transforms handled differently?\n\nThanks for the writeup and code :)",
    "1768869": "Thanks @pcmiller . If you mean best single model = (no fold, no TTA) so I don't know because after some very early models, I've seen a big boost with TTA and always used it later. My model predicts with a crop ~ 5 seconds so I think it's more logic using TTA so I can take into account the whole music file.\n\nActually, after verifying my best model, it is xse_resnext50 with 7 folds. I think many kagglers here count it as a single model :D\n\nI also don't know if the original xse_resnext (in the paper) have self attention or not or fastai add it. But in fastai I can activate it easily with `sa=1` in `xse_resnext101(n_out=19, act_cls=Mish, sa=1, pool=MaxPool, pretrained=False)`. More details you can find here: https://docs.fast.ai/layers.html#SelfAttention\n\nFor TTA, the same augmentations for both training set and test set. I think I just did cropping, reflection, changing brightness so it doesn't hurt too much the test set\n\nHope it clears :)",
    "1768978": "Thanks for the response :)",
    "1773263": "Congratulations @dienhoa on fully deserved 1st place! A couple of questions if I may: \n(1) What was your score LB for your pipeline without that maxBlurPool? \n(2) What was your LB without smoothing labels (but with mixup)? I noticed, in my cases, it should be one or the other but not both (either smoothing or mixup).\n(3) You used folds voting. What was just the mean of folds prediction probs/logits? In my case mean was always better (per model)",
    "1773387": "Congratulations @dienhoa \nThanks for sharing",
    "1774036": "Thanks @dmitrykonovalov . It's hard to answer because I didn't carefully log down everything during different experiments, I definitely need to journal for the next project (maybe with WandB or just a text file). And also, I exclusively tried a lot of TTA in the 2 final submissions to get a better score so the score from the previous submission is not very good.  \n\n1. If I remember correctly, with maxBlurPool my CV improved from 0.59x to 0.60x and the model converge faster (40 epochs instead of 50 epochs w/o maxBlurPool)\n2. These 2 (MixUp and LabelSmoothing) I always use together so I don't know if trying just one can improve the score or not.\n3. It's not actually folds voting. I have 7 folds and each folds predicts 50 different 5s-slices, so 350 predictions for voting. Because the model runs on a slice of just 5 seconds audio so I was thinking that maybe there are some slices that not relevant at all in the whole music, and can hurt the prediction if I use mean average. I did not test the submission's score with mean-avg",
    "1774132": "Thank you and see you in a next sound competition :)",
    "1781628": "Congratulations @dienhoa.\n\nI was trying to load your models standalone by installing fastai, korina and downloading data from here:\n\nhttps://www.kaggle.com/datasets/dienhoa/inference-music-genre\n\nThe objective was to create an inference pipeline which takes a audio as input, and outputs the predictions of your models. What are steps required to create such a pipeline?\n\nWhen I am trying in standalone loading models, I am getting the below error. \n```\nFile /opt/conda/lib/python3.8/site-packages/torch/serialization.py:1039, in _load.<locals>.UnpicklerWrapper.find_class(self, mod_name, name)\n   1037         pass\n   1038 mod_name = load_module_mapping.get(mod_name, mod_name)\n-> 1039 return super().find_class(mod_name, name)\n\nModuleNotFoundError: No module named 'kornia.contrib.max_blur_pool'\n```\n\nAny clues on debugging this issue, and your thoughts on creating a inference pipeline will be highly appreciated.\n\n~ Kurian",
    "1781677": "I think you get this problem when you install the latest version of Kornia (0.6.4). The code is ok with Kornia Kaggle version (0.5.8).\n\nI think `kornia.contrib.max_blur_pool` is replaced by `kornia.filters.max_blur_pool2d` in the latter version, however when I tried to use it, I got a type error, I don't remember exactly but something like Kornia doesn't accept fastai TensorImage as Input.\n\nHope it helps and let me know if you have something news ;)"
  },
  "source": "meta"
}