{
  "id": 321438,
  "title": "What is the trick that boosts your CV ?",
  "url": "/competitions/kaggle-pog-series-s01e02/discussion/321438",
  "author_name": "",
  "post_date": "2022-04-26T20:17:13.475331600Z",
  "votes": 7,
  "comment_count": 4,
  "views": 0,
  "content": "<p>Hi everyone,</p>\n<p>During this competition, there are some moments when I felt really stuck, and then a trick boosts my CV score by 0.3. 0.4% or reduce a lot of my training time. This is a very great feeling!</p>\n<ul>\n<li><p>Using precompute Spectrogram instead of the raw audio, this accelerates a lot of my training time (thanks to <a href=\"https://www.kaggle.com/theoviel\" target=\"_blank\">@theoviel</a> )</p></li>\n<li><p>The ImagNette tips: <br>\nxse_resnext50<br>\nMixUp<br>\nLabelSmoothing<br>\nfast.ai fit_flat_cos Learning Rate Scheduler<br>\nMish as activation function<br>\nSelf Attention<br>\nRanger Optimizer<br>\nMixup, LabelSmoothing</p></li>\n</ul>\n<p>They boost my CV score from 0.54x to 0.58x </p>\n<ul>\n<li><p>Training with 224 px Spectrogram instead of 128: 0.58x -&gt; 0.61x. I trained firstly with 128 to accelerate things but changing to 224 really boost my CV score (the happiest moment for me in this competition)</p></li>\n<li><p>To make the solution more robust in Private Leaderboard: TTA (simple but efficient because my model predict on a slice of 5s so randomly cropping then I can take into account the whole file)<br>\nMaxPoolBlur : this trick helps to improve the robustness too I think, it helps to overcome violating the Nyquist theorem of MaxPool (also reduces training time too)</p></li>\n</ul>\n<p>Can you share your Ahah moment with us ? </p>\n<p>Thanks</p>",
  "messages": [
    {
      "id": "1768979",
      "postDate": "04/26/2022 20:17:13",
      "content": "<p>Hi everyone,</p>\n<p>During this competition, there are some moments when I felt really stuck, and then a trick boosts my CV score by 0.3. 0.4% or reduce a lot of my training time. This is a very great feeling!</p>\n<ul>\n<li><p>Using precompute Spectrogram instead of the raw audio, this accelerates a lot of my training time (thanks to <a href=\"https://www.kaggle.com/theoviel\" target=\"_blank\">@theoviel</a> )</p></li>\n<li><p>The ImagNette tips: <br>\nxse_resnext50<br>\nMixUp<br>\nLabelSmoothing<br>\nfast.ai fit_flat_cos Learning Rate Scheduler<br>\nMish as activation function<br>\nSelf Attention<br>\nRanger Optimizer<br>\nMixup, LabelSmoothing</p></li>\n</ul>\n<p>They boost my CV score from 0.54x to 0.58x </p>\n<ul>\n<li><p>Training with 224 px Spectrogram instead of 128: 0.58x -&gt; 0.61x. I trained firstly with 128 to accelerate things but changing to 224 really boost my CV score (the happiest moment for me in this competition)</p></li>\n<li><p>To make the solution more robust in Private Leaderboard: TTA (simple but efficient because my model predict on a slice of 5s so randomly cropping then I can take into account the whole file)<br>\nMaxPoolBlur : this trick helps to improve the robustness too I think, it helps to overcome violating the Nyquist theorem of MaxPool (also reduces training time too)</p></li>\n</ul>\n<p>Can you share your Ahah moment with us ? </p>\n<p>Thanks</p>",
      "rawMarkdown": "Hi everyone,\n\nDuring this competition, there are some moments when I felt really stuck, and then a trick boosts my CV score by 0.3. 0.4% or reduce a lot of my training time. This is a very great feeling!\n\n- Using precompute Spectrogram instead of the raw audio, this accelerates a lot of my training time (thanks to @theoviel )\n\n- The ImagNette tips: \nxse_resnext50\nMixUp\nLabelSmoothing\nfast.ai fit_flat_cos Learning Rate Scheduler\nMish as activation function\nSelf Attention\nRanger Optimizer\nMixup, LabelSmoothing\n\nThey boost my CV score from 0.54x to 0.58x \n\n- Training with 224 px Spectrogram instead of 128: 0.58x -> 0.61x. I trained firstly with 128 to accelerate things but changing to 224 really boost my CV score (the happiest moment for me in this competition)\n\n- To make the solution more robust in Private Leaderboard: TTA (simple but efficient because my model predict on a slice of 5s so randomly cropping then I can take into account the whole file)\nMaxPoolBlur : this trick helps to improve the robustness too I think, it helps to overcome violating the Nyquist theorem of MaxPool (also reduces training time too)\n\nCan you share your Ahah moment with us ? \n\nThanks",
      "votes": null
    },
    {
      "id": "1769296",
      "postDate": "04/27/2022 06:14:45",
      "content": "<p>A lot of mine are similar to yours. My best models were eca-resnext50d.</p>\n<p>Aha moments:</p>\n<ol>\n<li>Reduced stride for 1st conv layer [(1,1) instead of typical (2,2)]. </li>\n<li>Noisy label training</li>\n</ol>",
      "rawMarkdown": "A lot of mine are similar to yours. My best models were eca-resnext50d.\n\nAha moments:\n1. Reduced stride for 1st conv layer [(1,1) instead of typical (2,2)]. \n2. Noisy label training",
      "votes": null
    },
    {
      "id": "1769782",
      "postDate": "04/27/2022 15:15:17",
      "content": "<p>Thanks <a href=\"https://www.kaggle.com/pheadrus\" target=\"_blank\">@pheadrus</a> . Do you have some articles that I can read about <code>Noisy label training</code> ? By the way, can I see your training notebook pls? I'm curious that you said in one discussion that, your model learns in 7,8 epochs, mine is about 40 epochs (&gt;2 hours) :D</p>",
      "rawMarkdown": "Thanks @pheadrus . Do you have some articles that I can read about `Noisy label training` ? By the way, can I see your training notebook pls? I'm curious that you said in one discussion that, your model learns in 7,8 epochs, mine is about 40 epochs (>2 hours) :D",
      "votes": null
    },
    {
      "id": "1769854",
      "postDate": "04/27/2022 17:15:12",
      "content": "<p>Umn, I am not sure about papers, but think of it more like student-teacher models, wherein, I use teacher preds as an aux pred task for student. Contrary to popular wisdom, I often use student models that are larger than teacher models. I will have to refine my notebook, but I can send you a rough draft over email in a bit.   </p>",
      "rawMarkdown": "Umn, I am not sure about papers, but think of it more like student-teacher models, wherein, I use teacher preds as an aux pred task for student. Contrary to popular wisdom, I often use student models that are larger than teacher models. I will have to refine my notebook, but I can send you a rough draft over email in a bit.",
      "votes": null
    },
    {
      "id": "1769918",
      "postDate": "04/27/2022 18:48:12",
      "content": "<p>Thanks a lot <a href=\"https://www.kaggle.com/pheadrus\" target=\"_blank\">@pheadrus</a> :)</p>",
      "rawMarkdown": "Thanks a lot @pheadrus :)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1769296,
      "author_name": "pheadrus",
      "author_url": "",
      "post_date": "04/27/2022 06:14:45",
      "content": "<p>A lot of mine are similar to yours. My best models were eca-resnext50d.</p>\n<p>Aha moments:</p>\n<ol>\n<li>Reduced stride for 1st conv layer [(1,1) instead of typical (2,2)]. </li>\n<li>Noisy label training</li>\n</ol>",
      "votes": null,
      "replies": [
        {
          "id": 1769782,
          "author_name": "dienhoa",
          "author_url": "",
          "post_date": "04/27/2022 15:15:17",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/pheadrus\" target=\"_blank\">@pheadrus</a> . Do you have some articles that I can read about <code>Noisy label training</code> ? By the way, can I see your training notebook pls? I'm curious that you said in one discussion that, your model learns in 7,8 epochs, mine is about 40 epochs (&gt;2 hours) :D</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1769854,
          "author_name": "pheadrus",
          "author_url": "",
          "post_date": "04/27/2022 17:15:12",
          "content": "<p>Umn, I am not sure about papers, but think of it more like student-teacher models, wherein, I use teacher preds as an aux pred task for student. Contrary to popular wisdom, I often use student models that are larger than teacher models. I will have to refine my notebook, but I can send you a rough draft over email in a bit.   </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1769918,
          "author_name": "dienhoa",
          "author_url": "",
          "post_date": "04/27/2022 18:48:12",
          "content": "<p>Thanks a lot <a href=\"https://www.kaggle.com/pheadrus\" target=\"_blank\">@pheadrus</a> :)</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1768979": "Hi everyone,\n\nDuring this competition, there are some moments when I felt really stuck, and then a trick boosts my CV score by 0.3. 0.4% or reduce a lot of my training time. This is a very great feeling!\n\n- Using precompute Spectrogram instead of the raw audio, this accelerates a lot of my training time (thanks to @theoviel )\n\n- The ImagNette tips: \nxse_resnext50\nMixUp\nLabelSmoothing\nfast.ai fit_flat_cos Learning Rate Scheduler\nMish as activation function\nSelf Attention\nRanger Optimizer\nMixup, LabelSmoothing\n\nThey boost my CV score from 0.54x to 0.58x \n\n- Training with 224 px Spectrogram instead of 128: 0.58x -> 0.61x. I trained firstly with 128 to accelerate things but changing to 224 really boost my CV score (the happiest moment for me in this competition)\n\n- To make the solution more robust in Private Leaderboard: TTA (simple but efficient because my model predict on a slice of 5s so randomly cropping then I can take into account the whole file)\nMaxPoolBlur : this trick helps to improve the robustness too I think, it helps to overcome violating the Nyquist theorem of MaxPool (also reduces training time too)\n\nCan you share your Ahah moment with us ? \n\nThanks",
    "1769296": "A lot of mine are similar to yours. My best models were eca-resnext50d.\n\nAha moments:\n1. Reduced stride for 1st conv layer [(1,1) instead of typical (2,2)]. \n2. Noisy label training",
    "1769782": "Thanks @pheadrus . Do you have some articles that I can read about `Noisy label training` ? By the way, can I see your training notebook pls? I'm curious that you said in one discussion that, your model learns in 7,8 epochs, mine is about 40 epochs (>2 hours) :D",
    "1769854": "Umn, I am not sure about papers, but think of it more like student-teacher models, wherein, I use teacher preds as an aux pred task for student. Contrary to popular wisdom, I often use student models that are larger than teacher models. I will have to refine my notebook, but I can send you a rough draft over email in a bit.",
    "1769918": "Thanks a lot @pheadrus :)"
  },
  "source": "meta"
}