{
  "id": 168608,
  "title": "[9th place] Short summary",
  "url": "/competitions/alaska2-image-steganalysis/discussion/168608",
  "author_name": "Psi",
  "post_date": "2020-07-21T08:15:49.002000",
  "votes": 57,
  "comment_count": 20,
  "views": 0,
  "content": "<p>Thanks to the hosts for this interesting competition and <a href=\"https://www.kaggle.com/christofhenkel\" target=\"_blank\">@christofhenkel</a> for the fun team!</p>\n<p>We are happy to make the jump on private LB which is as often is the case based on a strong blend that is not overfitting the public LB. That said, our best sub is both our best pub LB and private LB and we made the best selection.</p>\n<h3>Image norm</h3>\n<p>We observed slight differences in image channel distributions between train and test and doing local image normalization brought CV and LB closer together for us, which is why we sticked to it throughout our final models. That means we standardize each channel locally for each image.</p>\n<h3>Augmentations</h3>\n<p>Most models only use standard flips and transposes. Some models also add cutout and some add tiny random noise. For TTA we do either TTA4 or TTA8.</p>\n<h3>Models</h3>\n<p>We only use EfficientNets. We started fitting models only on fold 0, but then switched to fitting models \"blindly\" on full data as CV convergence always was stable for us. If we were unsure, we re-fitted on a fold on a new model and checked if the models overfit or not, and then trusted the full fits. This also meant that we had to blend a bit blindly though.</p>\n<p>We both fitted vanilla EfficientNets, but also got huge boosts on CV when changing the stride in the first layer to (1,1), as also other contestants did. This means that the models are fit for longer on the full resolution, and this is specifically helpful as a lot of information about the manipulation lies between neighboring pixels. We observed this facet also when randomly shuffling neighboring pixels in TTA and as soon as you shuffle further than one pixel away, the score deteriorates heavily.</p>\n<p>Our final sub consists of the following models:</p>\n<ul>\n<li>Vanilla EfficientNet B6 (one fold)</li>\n<li>Vanilla EfficientNet B8 (full data)</li>\n<li>Stride 1 EfficientNet B1 (full data)</li>\n<li>Stride 1 EfficientNet B2 (full data)</li>\n<li>Stride 1 EfficientNet B3 (full data)</li>\n</ul>\n<p>Unfortunately, we dont have the full logs plotted as some models finished on the last day, but this is how the training scores across models roughly look like (2x refers to stride 1 models):</p>\n<p><img src=\"https://i.imgur.com/ooXn6U9.png\" alt=\"\"></p>\n<p>You can see that B3 Stride 1 model was the best, also on a separate fit on CV. However, vanilla B6 model was the best on public LB, so things are a bit misleading there. On private LB actually B3 Stride 1 alone is our best one (925), followed by B8 (924) and so on. So actually matching the ranking on the plot above.</p>\n<h3>Fitting</h3>\n<p>All models are fitted over 40 epochs with cosine decay, AdamW, mixed precision, and varying batch sizes. We did not see changes in BS and LR to have much impact, which is why we basically left it untouched across models as we also did not have the time to tune it further.</p>\n<h3>Blending</h3>\n<p>Our final blend is composed of a 50-50 blend between (1) a blend between vanilla B6 and B8 (LB Score 936) and (2) a blend between epochs 30,35,39 of the B1,B2,B3 Stride 1 models (LB Score 935). The blend scored 937 on public LB and 928 on private LB.</p>",
  "messages": [
    {
      "id": 937880,
      "postDate": "2020-07-21T08:15:49.003Z",
      "content": "<p>Thanks to the hosts for this interesting competition and <a href=\"https://www.kaggle.com/christofhenkel\" target=\"_blank\">@christofhenkel</a> for the fun team!</p>\n<p>We are happy to make the jump on private LB which is as often is the case based on a strong blend that is not overfitting the public LB. That said, our best sub is both our best pub LB and private LB and we made the best selection.</p>\n<h3>Image norm</h3>\n<p>We observed slight differences in image channel distributions between train and test and doing local image normalization brought CV and LB closer together for us, which is why we sticked to it throughout our final models. That means we standardize each channel locally for each image.</p>\n<h3>Augmentations</h3>\n<p>Most models only use standard flips and transposes. Some models also add cutout and some add tiny random noise. For TTA we do either TTA4 or TTA8.</p>\n<h3>Models</h3>\n<p>We only use EfficientNets. We started fitting models only on fold 0, but then switched to fitting models \"blindly\" on full data as CV convergence always was stable for us. If we were unsure, we re-fitted on a fold on a new model and checked if the models overfit or not, and then trusted the full fits. This also meant that we had to blend a bit blindly though.</p>\n<p>We both fitted vanilla EfficientNets, but also got huge boosts on CV when changing the stride in the first layer to (1,1), as also other contestants did. This means that the models are fit for longer on the full resolution, and this is specifically helpful as a lot of information about the manipulation lies between neighboring pixels. We observed this facet also when randomly shuffling neighboring pixels in TTA and as soon as you shuffle further than one pixel away, the score deteriorates heavily.</p>\n<p>Our final sub consists of the following models:</p>\n<ul>\n<li>Vanilla EfficientNet B6 (one fold)</li>\n<li>Vanilla EfficientNet B8 (full data)</li>\n<li>Stride 1 EfficientNet B1 (full data)</li>\n<li>Stride 1 EfficientNet B2 (full data)</li>\n<li>Stride 1 EfficientNet B3 (full data)</li>\n</ul>\n<p>Unfortunately, we dont have the full logs plotted as some models finished on the last day, but this is how the training scores across models roughly look like (2x refers to stride 1 models):</p>\n<p><img src=\"https://i.imgur.com/ooXn6U9.png\" alt=\"\"></p>\n<p>You can see that B3 Stride 1 model was the best, also on a separate fit on CV. However, vanilla B6 model was the best on public LB, so things are a bit misleading there. On private LB actually B3 Stride 1 alone is our best one (925), followed by B8 (924) and so on. So actually matching the ranking on the plot above.</p>\n<h3>Fitting</h3>\n<p>All models are fitted over 40 epochs with cosine decay, AdamW, mixed precision, and varying batch sizes. We did not see changes in BS and LR to have much impact, which is why we basically left it untouched across models as we also did not have the time to tune it further.</p>\n<h3>Blending</h3>\n<p>Our final blend is composed of a 50-50 blend between (1) a blend between vanilla B6 and B8 (LB Score 936) and (2) a blend between epochs 30,35,39 of the B1,B2,B3 Stride 1 models (LB Score 935). The blend scored 937 on public LB and 928 on private LB.</p>",
      "rawMarkdown": "Thanks to the hosts for this interesting competition and @christofhenkel for the fun team!\n\nWe are happy to make the jump on private LB which is as often is the case based on a strong blend that is not overfitting the public LB. That said, our best sub is both our best pub LB and private LB and we made the best selection.\n\n### Image norm\n\nWe observed slight differences in image channel distributions between train and test and doing local image normalization brought CV and LB closer together for us, which is why we sticked to it throughout our final models. That means we standardize each channel locally for each image.\n\n### Augmentations\n\nMost models only use standard flips and transposes. Some models also add cutout and some add tiny random noise. For TTA we do either TTA4 or TTA8.\n\n### Models\n\nWe only use EfficientNets. We started fitting models only on fold 0, but then switched to fitting models \"blindly\" on full data as CV convergence always was stable for us. If we were unsure, we re-fitted on a fold on a new model and checked if the models overfit or not, and then trusted the full fits. This also meant that we had to blend a bit blindly though.\n\nWe both fitted vanilla EfficientNets, but also got huge boosts on CV when changing the stride in the first layer to (1,1), as also other contestants did. This means that the models are fit for longer on the full resolution, and this is specifically helpful as a lot of information about the manipulation lies between neighboring pixels. We observed this facet also when randomly shuffling neighboring pixels in TTA and as soon as you shuffle further than one pixel away, the score deteriorates heavily.\n\nOur final sub consists of the following models:\n\n- Vanilla EfficientNet B6 (one fold)\n- Vanilla EfficientNet B8 (full data)\n- Stride 1 EfficientNet B1 (full data)\n- Stride 1 EfficientNet B2 (full data)\n- Stride 1 EfficientNet B3 (full data)\n\nUnfortunately, we dont have the full logs plotted as some models finished on the last day, but this is how the training scores across models roughly look like (2x refers to stride 1 models):\n\n![](https://i.imgur.com/ooXn6U9.png)\n\nYou can see that B3 Stride 1 model was the best, also on a separate fit on CV. However, vanilla B6 model was the best on public LB, so things are a bit misleading there. On private LB actually B3 Stride 1 alone is our best one (925), followed by B8 (924) and so on. So actually matching the ranking on the plot above.\n\n### Fitting\n\nAll models are fitted over 40 epochs with cosine decay, AdamW, mixed precision, and varying batch sizes. We did not see changes in BS and LR to have much impact, which is why we basically left it untouched across models as we also did not have the time to tune it further.\n\n### Blending\n\nOur final blend is composed of a 50-50 blend between (1) a blend between vanilla B6 and B8 (LB Score 936) and (2) a blend between epochs 30,35,39 of the B1,B2,B3 Stride 1 models (LB Score 935). The blend scored 937 on public LB and 928 on private LB.\n\n",
      "votes": 57
    },
    {
      "id": 938690,
      "postDate": "2020-07-21T17:24:20.653Z",
      "content": "<p>Thanks for sharing awesome summary and congrats <a href=\"/philippsinger\">@philippsinger</a> for the gold medal!\nI watched you take part in the competition late and was curious about the result.\nAnd I feel a lot in this article.</p>\n\n<p>There is a part I didn't understand:\n<code>(2) a blend between epochs 30,35,39 of the B1,B2,B3 Stride 1 models (LB Score 935).</code>\nIt means that a blend between B1 of 30 epochs + B2 of 35 epochs + B3 of 39 epochs?\nIn simple terms, is it a result using best-checkpoints for B1, B2, B3?</p>\n\n<p>Thanks again. I really learn a lot.</p>",
      "rawMarkdown": "Thanks for sharing awesome summary and congrats @philippsinger for the gold medal!\nI watched you take part in the competition late and was curious about the result.\nAnd I feel a lot in this article.\n\nThere is a part I didn't understand:\n`(2) a blend between epochs 30,35,39 of the B1,B2,B3 Stride 1 models (LB Score 935).`\nIt means that a blend between B1 of 30 epochs + B2 of 35 epochs + B3 of 39 epochs?\nIn simple terms, is it a result using best-checkpoints for B1, B2, B3?\n\nThanks again. I really learn a lot.\n",
      "votes": 1,
      "replies": [
        {
          "id": 938694,
          "postDate": "2020-07-21T17:26:19.063Z",
          "content": "<p>Yes exactly, it is basically a checkpoint ensemble. I think last epoch would have been enough though.</p>",
          "rawMarkdown": "Yes exactly, it is basically a checkpoint ensemble. I think last epoch would have been enough though.",
          "votes": 1
        },
        {
          "id": 938702,
          "postDate": "2020-07-21T17:31:37.733Z",
          "content": "<p>Thanks! \nI hope your next competition is successful. 👍 </p>",
          "rawMarkdown": "Thanks! \nI hope your next competition is successful. 👍 ",
          "votes": 1
        },
        {
          "id": 938996,
          "postDate": "2020-07-22T00:10:01.723Z",
          "content": "<blockquote>\n  <p>I think last epoch would have been enough though.</p>\n</blockquote>\n<p>no way :P</p>",
          "rawMarkdown": "&gt;  I think last epoch would have been enough though.\n\nno way :P",
          "votes": 2
        },
        {
          "id": 939344,
          "postDate": "2020-07-22T07:16:55.367Z",
          "content": "<p>Oh,  congrats <a href=\"/christofhenkel\">@christofhenkel</a>!\n👍 </p>",
          "rawMarkdown": "Oh,  congrats @christofhenkel!\n👍 ",
          "votes": 1
        }
      ]
    },
    {
      "id": 938347,
      "postDate": "2020-07-21T13:22:32.210Z",
      "content": "<p>any chance to share a bit on what HW you use in this competition?<br>\nfor me,  this is one of the most computing resource intensive competitions I took part. </p>\n<p>spent a few hundred $ on vastai myself to use their V100 &amp; RTX Titan instance in the final week of the comp to run larger models, however even with that I won't have resource (or rather, time)  to run B7/B8</p>\n<p>in our best single model - a B5, I actually have to reset the cos annealing scheme and the number of the epoch so that it could finish before the competition deadline. </p>",
      "rawMarkdown": "any chance to share a bit on what HW you use in this competition?\nfor me,  this is one of the most computing resource intensive competitions I took part. \n\nspent a few hundred $ on vastai myself to use their V100 &amp; RTX Titan instance in the final week of the comp to run larger models, however even with that I won't have resource (or rather, time)  to run B7/B8\n\nin our best single model - a B5, I actually have to reset the cos annealing scheme and the number of the epoch so that it could finish before the competition deadline. ",
      "votes": 1,
      "replies": [
        {
          "id": 938399,
          "postDate": "2020-07-21T13:58:46.760Z",
          "content": "<p>This was indeed a compute heavy competition, we used V100s ourselves. Still, you can do a lot here with architectural changes, then things like hard sample mining also work well to reduce runtime. Compute can actually really bring you quite far here, just taking public kernel and fitting B6-B8 already brings you far. To reach top spots you still need to think out of the box though.</p>",
          "rawMarkdown": "This was indeed a compute heavy competition, we used V100s ourselves. Still, you can do a lot here with architectural changes, then things like hard sample mining also work well to reduce runtime. Compute can actually really bring you quite far here, just taking public kernel and fitting B6-B8 already brings you far. To reach top spots you still need to think out of the box though.",
          "votes": 2
        }
      ]
    },
    {
      "id": 938306,
      "postDate": "2020-07-21T12:53:52.977Z",
      "content": "<p>congrats @psi and <a href=\"/christofhenkel\">@christofhenkel</a> for the gold. There are some similarities with our solution. </p>",
      "rawMarkdown": "congrats @psi and @christofhenkel for the gold. There are some similarities with our solution. ",
      "votes": 1,
      "replies": [
        {
          "id": 938322,
          "postDate": "2020-07-21T13:05:19.990Z",
          "content": "<p>Thanks - how did you guys climb so high on public LB?</p>",
          "rawMarkdown": "Thanks - how did you guys climb so high on public LB?"
        }
      ]
    },
    {
      "id": 940077,
      "postDate": "2020-07-22T17:02:17.543Z",
      "content": "<p>Thanks for sharing and congratulations on your medal! How long did it take to train everything?</p>",
      "rawMarkdown": "Thanks for sharing and congratulations on your medal! How long did it take to train everything?"
    },
    {
      "id": 939534,
      "postDate": "2020-07-22T09:47:16.453Z",
      "content": "<p>Congratulations on top 10! \nYour method is clever, simple and informative!!\nCould you tell me the parameters of cosine annealing? I didn't use it because it's harder to adjust lr and epoch than ReduceOnplateau, but a lot of top teams use cosine annealing.</p>",
      "rawMarkdown": "Congratulations on top 10! \nYour method is clever, simple and informative!!\nCould you tell me the parameters of cosine annealing? I didn't use it because it's harder to adjust lr and epoch than ReduceOnplateau, but a lot of top teams use cosine annealing.",
      "replies": [
        {
          "id": 939736,
          "postDate": "2020-07-22T12:28:03.087Z",
          "content": "<p>40 epochs, 2 epochs warmup, cosine decay. Starting LR and BS did not matter much in this comp.</p>",
          "rawMarkdown": "40 epochs, 2 epochs warmup, cosine decay. Starting LR and BS did not matter much in this comp.",
          "votes": 1
        },
        {
          "id": 942092,
          "postDate": "2020-07-23T15:21:01.207Z",
          "content": "<p>Thanks!</p>",
          "rawMarkdown": "Thanks!"
        }
      ]
    },
    {
      "id": 939046,
      "postDate": "2020-07-22T02:05:31.210Z",
      "content": "<blockquote>\n  <p>What I mean is that we fitted the models on the full 100% training data without any validation fold and just looked at training score.</p>\n</blockquote>\n<p><a href=\"https://www.kaggle.com/philippsinger\" target=\"_blank\">@philippsinger</a>  Would you mind telling me what is the best way to tell when to stop when doing full 100% training data without validation ? Should we keep the number of epoch the same as in fold 0, or should we see the loss score or accuracy score until it reach the same number as fold 0 ? </p>\n<blockquote>\n  <p>If we were unsure, we re-fitted on a fold on a new model and checked if the models overfit or not, and then trusted the full fits.</p>\n</blockquote>\n<p>Does this mean taking the fulldata checkpoint and refit it on a fold ? (in this case for how many epoch and at which learning rate)</p>\n<p>Thank you very much for the write-ups and congratulations on the strong finish !</p>",
      "rawMarkdown": "&gt; What I mean is that we fitted the models on the full 100% training data without any validation fold and just looked at training score.\n\n@philippsinger  Would you mind telling me what is the best way to tell when to stop when doing full 100% training data without validation ? Should we keep the number of epoch the same as in fold 0, or should we see the loss score or accuracy score until it reach the same number as fold 0 ? \n\n&gt;  If we were unsure, we re-fitted on a fold on a new model and checked if the models overfit or not, and then trusted the full fits.\n\nDoes this mean taking the fulldata checkpoint and refit it on a fold ? (in this case for how many epoch and at which learning rate)\n\nThank you very much for the write-ups and congratulations on the strong finish !\n",
      "replies": [
        {
          "id": 939310,
          "postDate": "2020-07-22T06:52:12.130Z",
          "content": "<p>You start with using a validation fold, or better k-fold and analyze your training/validation curves. If your last epoch is always the best one, you can be pretty confident that fitting on full data and picking last epoch should work well. The only issue is that you have more training data, so potentially an earlier epoch would be better. However, in this case we saw on val that even if there was some overfit, val was staying stable, and did not deteriorate, so picking last epoch of full fit was fine.</p>",
          "rawMarkdown": "You start with using a validation fold, or better k-fold and analyze your training/validation curves. If your last epoch is always the best one, you can be pretty confident that fitting on full data and picking last epoch should work well. The only issue is that you have more training data, so potentially an earlier epoch would be better. However, in this case we saw on val that even if there was some overfit, val was staying stable, and did not deteriorate, so picking last epoch of full fit was fine.",
          "votes": 2
        }
      ]
    },
    {
      "id": 938222,
      "postDate": "2020-07-21T12:10:53.493Z",
      "content": "<p><a href=\"https://www.kaggle.com/philippsinger\" target=\"_blank\">@philippsinger</a> <br>\nhi, thank you for the awesome and straightforward write-up. A query, I didn't catch one of your words; <code>fitting models only on fold 0, but then switched to fitting models \"blindly\" on full data...</code> I guess, you had N fold, among them you've fitted the model in fold no. (N-1) and validate on Nth fold! Can u please clarify?</p>\n<p>Reducing stride looks like an almost common approach like others did. As you said, <code>this is specifically helpful as a lot of information about the manipulation lies between neighboring pixels.</code>..reducing stride makes lots sense.</p>",
      "rawMarkdown": "@philippsinger \nhi, thank you for the awesome and straightforward write-up. A query, I didn't catch one of your words; `fitting models only on fold 0, but then switched to fitting models \"blindly\" on full data...` I guess, you had N fold, among them you've fitted the model in fold no. (N-1) and validate on Nth fold! Can u please clarify?\n\nReducing stride looks like an almost common approach like others did. As you said, ` this is specifically helpful as a lot of information about the manipulation lies between neighboring pixels.`..reducing stride makes lots sense.",
      "replies": [
        {
          "id": 938238,
          "postDate": "2020-07-21T12:17:51.467Z",
          "content": "<p><a href=\"https://www.kaggle.com/ipythonx\" target=\"_blank\">@ipythonx</a> What I mean is that we fitted the models on the full 100% training data without any validation fold and just looked at training score.</p>",
          "rawMarkdown": "@ipythonx What I mean is that we fitted the models on the full 100% training data without any validation fold and just looked at training score.",
          "votes": 1
        },
        {
          "id": 938271,
          "postDate": "2020-07-21T12:32:55.140Z",
          "content": "<p>hmm, I see. Thank u. -)</p>",
          "rawMarkdown": "hmm, I see. Thank u. -)"
        }
      ]
    },
    {
      "id": 943674,
      "postDate": "2020-07-24T14:09:45.727Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 942403,
      "postDate": "2020-07-23T18:16:33.407Z",
      "content": "<p>thanks for sharing your experience</p>",
      "rawMarkdown": "thanks for sharing your experience"
    }
  ],
  "comments": [
    {
      "id": 938690,
      "author_name": "Heroseo",
      "author_url": "",
      "post_date": "2020-07-21T17:24:20.653000",
      "content": "<p>Thanks for sharing awesome summary and congrats <a href=\"/philippsinger\">@philippsinger</a> for the gold medal!\nI watched you take part in the competition late and was curious about the result.\nAnd I feel a lot in this article.</p>\n\n<p>There is a part I didn't understand:\n<code>(2) a blend between epochs 30,35,39 of the B1,B2,B3 Stride 1 models (LB Score 935).</code>\nIt means that a blend between B1 of 30 epochs + B2 of 35 epochs + B3 of 39 epochs?\nIn simple terms, is it a result using best-checkpoints for B1, B2, B3?</p>\n\n<p>Thanks again. I really learn a lot.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 938694,
          "author_name": "Psi",
          "author_url": "",
          "post_date": "2020-07-21T17:26:19.063000",
          "content": "<p>Yes exactly, it is basically a checkpoint ensemble. I think last epoch would have been enough though.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 938702,
          "author_name": "Heroseo",
          "author_url": "",
          "post_date": "2020-07-21T17:31:37.733000",
          "content": "<p>Thanks! \nI hope your next competition is successful. 👍 </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 938996,
          "author_name": "Dieter",
          "author_url": "",
          "post_date": "2020-07-22T00:10:01.723000",
          "content": "<blockquote>\n  <p>I think last epoch would have been enough though.</p>\n</blockquote>\n<p>no way :P</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 939344,
          "author_name": "Heroseo",
          "author_url": "",
          "post_date": "2020-07-22T07:16:55.367000",
          "content": "<p>Oh,  congrats <a href=\"/christofhenkel\">@christofhenkel</a>!\n👍 </p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 938347,
      "author_name": "Yifan Xie",
      "author_url": "",
      "post_date": "2020-07-21T13:22:32.210000",
      "content": "<p>any chance to share a bit on what HW you use in this competition?<br>\nfor me,  this is one of the most computing resource intensive competitions I took part. </p>\n<p>spent a few hundred $ on vastai myself to use their V100 &amp; RTX Titan instance in the final week of the comp to run larger models, however even with that I won't have resource (or rather, time)  to run B7/B8</p>\n<p>in our best single model - a B5, I actually have to reset the cos annealing scheme and the number of the epoch so that it could finish before the competition deadline. </p>",
      "votes": 1,
      "replies": [
        {
          "id": 938399,
          "author_name": "Psi",
          "author_url": "",
          "post_date": "2020-07-21T13:58:46.760000",
          "content": "<p>This was indeed a compute heavy competition, we used V100s ourselves. Still, you can do a lot here with architectural changes, then things like hard sample mining also work well to reduce runtime. Compute can actually really bring you quite far here, just taking public kernel and fitting B6-B8 already brings you far. To reach top spots you still need to think out of the box though.</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 938306,
      "author_name": "Giba",
      "author_url": "",
      "post_date": "2020-07-21T12:53:52.977000",
      "content": "<p>congrats @psi and <a href=\"/christofhenkel\">@christofhenkel</a> for the gold. There are some similarities with our solution. </p>",
      "votes": 1,
      "replies": [
        {
          "id": 938322,
          "author_name": "Psi",
          "author_url": "",
          "post_date": "2020-07-21T13:05:19.990000",
          "content": "<p>Thanks - how did you guys climb so high on public LB?</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 940077,
      "author_name": "Yassine Alouini",
      "author_url": "",
      "post_date": "2020-07-22T17:02:17.543000",
      "content": "<p>Thanks for sharing and congratulations on your medal! How long did it take to train everything?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 939534,
      "author_name": "cp_t2",
      "author_url": "",
      "post_date": "2020-07-22T09:47:16.453000",
      "content": "<p>Congratulations on top 10! \nYour method is clever, simple and informative!!\nCould you tell me the parameters of cosine annealing? I didn't use it because it's harder to adjust lr and epoch than ReduceOnplateau, but a lot of top teams use cosine annealing.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 939736,
          "author_name": "Psi",
          "author_url": "",
          "post_date": "2020-07-22T12:28:03.087000",
          "content": "<p>40 epochs, 2 epochs warmup, cosine decay. Starting LR and BS did not matter much in this comp.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 942092,
          "author_name": "cp_t2",
          "author_url": "",
          "post_date": "2020-07-23T15:21:01.207000",
          "content": "<p>Thanks!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 939046,
      "author_name": "yukiya",
      "author_url": "",
      "post_date": "2020-07-22T02:05:31.210000",
      "content": "<blockquote>\n  <p>What I mean is that we fitted the models on the full 100% training data without any validation fold and just looked at training score.</p>\n</blockquote>\n<p><a href=\"https://www.kaggle.com/philippsinger\" target=\"_blank\">@philippsinger</a>  Would you mind telling me what is the best way to tell when to stop when doing full 100% training data without validation ? Should we keep the number of epoch the same as in fold 0, or should we see the loss score or accuracy score until it reach the same number as fold 0 ? </p>\n<blockquote>\n  <p>If we were unsure, we re-fitted on a fold on a new model and checked if the models overfit or not, and then trusted the full fits.</p>\n</blockquote>\n<p>Does this mean taking the fulldata checkpoint and refit it on a fold ? (in this case for how many epoch and at which learning rate)</p>\n<p>Thank you very much for the write-ups and congratulations on the strong finish !</p>",
      "votes": 0,
      "replies": [
        {
          "id": 939310,
          "author_name": "Psi",
          "author_url": "",
          "post_date": "2020-07-22T06:52:12.130000",
          "content": "<p>You start with using a validation fold, or better k-fold and analyze your training/validation curves. If your last epoch is always the best one, you can be pretty confident that fitting on full data and picking last epoch should work well. The only issue is that you have more training data, so potentially an earlier epoch would be better. However, in this case we saw on val that even if there was some overfit, val was staying stable, and did not deteriorate, so picking last epoch of full fit was fine.</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 938222,
      "author_name": "Innat",
      "author_url": "",
      "post_date": "2020-07-21T12:10:53.493000",
      "content": "<p><a href=\"https://www.kaggle.com/philippsinger\" target=\"_blank\">@philippsinger</a> <br>\nhi, thank you for the awesome and straightforward write-up. A query, I didn't catch one of your words; <code>fitting models only on fold 0, but then switched to fitting models \"blindly\" on full data...</code> I guess, you had N fold, among them you've fitted the model in fold no. (N-1) and validate on Nth fold! Can u please clarify?</p>\n<p>Reducing stride looks like an almost common approach like others did. As you said, <code>this is specifically helpful as a lot of information about the manipulation lies between neighboring pixels.</code>..reducing stride makes lots sense.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 938238,
          "author_name": "Psi",
          "author_url": "",
          "post_date": "2020-07-21T12:17:51.467000",
          "content": "<p><a href=\"https://www.kaggle.com/ipythonx\" target=\"_blank\">@ipythonx</a> What I mean is that we fitted the models on the full 100% training data without any validation fold and just looked at training score.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 938271,
          "author_name": "Innat",
          "author_url": "",
          "post_date": "2020-07-21T12:32:55.140000",
          "content": "<p>hmm, I see. Thank u. -)</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 943674,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-07-24T14:09:45.727000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 942403,
      "author_name": "Atul Singh",
      "author_url": "",
      "post_date": "2020-07-23T18:16:33.407000",
      "content": "<p>thanks for sharing your experience</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "937880": "Thanks to the hosts for this interesting competition and @christofhenkel for the fun team!\n\nWe are happy to make the jump on private LB which is as often is the case based on a strong blend that is not overfitting the public LB. That said, our best sub is both our best pub LB and private LB and we made the best selection.\n\n### Image norm\n\nWe observed slight differences in image channel distributions between train and test and doing local image normalization brought CV and LB closer together for us, which is why we sticked to it throughout our final models. That means we standardize each channel locally for each image.\n\n### Augmentations\n\nMost models only use standard flips and transposes. Some models also add cutout and some add tiny random noise. For TTA we do either TTA4 or TTA8.\n\n### Models\n\nWe only use EfficientNets. We started fitting models only on fold 0, but then switched to fitting models \"blindly\" on full data as CV convergence always was stable for us. If we were unsure, we re-fitted on a fold on a new model and checked if the models overfit or not, and then trusted the full fits. This also meant that we had to blend a bit blindly though.\n\nWe both fitted vanilla EfficientNets, but also got huge boosts on CV when changing the stride in the first layer to (1,1), as also other contestants did. This means that the models are fit for longer on the full resolution, and this is specifically helpful as a lot of information about the manipulation lies between neighboring pixels. We observed this facet also when randomly shuffling neighboring pixels in TTA and as soon as you shuffle further than one pixel away, the score deteriorates heavily.\n\nOur final sub consists of the following models:\n\n- Vanilla EfficientNet B6 (one fold)\n- Vanilla EfficientNet B8 (full data)\n- Stride 1 EfficientNet B1 (full data)\n- Stride 1 EfficientNet B2 (full data)\n- Stride 1 EfficientNet B3 (full data)\n\nUnfortunately, we dont have the full logs plotted as some models finished on the last day, but this is how the training scores across models roughly look like (2x refers to stride 1 models):\n\n![](https://i.imgur.com/ooXn6U9.png)\n\nYou can see that B3 Stride 1 model was the best, also on a separate fit on CV. However, vanilla B6 model was the best on public LB, so things are a bit misleading there. On private LB actually B3 Stride 1 alone is our best one (925), followed by B8 (924) and so on. So actually matching the ranking on the plot above.\n\n### Fitting\n\nAll models are fitted over 40 epochs with cosine decay, AdamW, mixed precision, and varying batch sizes. We did not see changes in BS and LR to have much impact, which is why we basically left it untouched across models as we also did not have the time to tune it further.\n\n### Blending\n\nOur final blend is composed of a 50-50 blend between (1) a blend between vanilla B6 and B8 (LB Score 936) and (2) a blend between epochs 30,35,39 of the B1,B2,B3 Stride 1 models (LB Score 935). The blend scored 937 on public LB and 928 on private LB.\n\n",
    "938690": "Thanks for sharing awesome summary and congrats @philippsinger for the gold medal!\nI watched you take part in the competition late and was curious about the result.\nAnd I feel a lot in this article.\n\nThere is a part I didn't understand:\n`(2) a blend between epochs 30,35,39 of the B1,B2,B3 Stride 1 models (LB Score 935).`\nIt means that a blend between B1 of 30 epochs + B2 of 35 epochs + B3 of 39 epochs?\nIn simple terms, is it a result using best-checkpoints for B1, B2, B3?\n\nThanks again. I really learn a lot.\n",
    "938347": "any chance to share a bit on what HW you use in this competition?\nfor me,  this is one of the most computing resource intensive competitions I took part. \n\nspent a few hundred $ on vastai myself to use their V100 &amp; RTX Titan instance in the final week of the comp to run larger models, however even with that I won't have resource (or rather, time)  to run B7/B8\n\nin our best single model - a B5, I actually have to reset the cos annealing scheme and the number of the epoch so that it could finish before the competition deadline. ",
    "938306": "congrats @psi and @christofhenkel for the gold. There are some similarities with our solution. ",
    "940077": "Thanks for sharing and congratulations on your medal! How long did it take to train everything?",
    "939534": "Congratulations on top 10! \nYour method is clever, simple and informative!!\nCould you tell me the parameters of cosine annealing? I didn't use it because it's harder to adjust lr and epoch than ReduceOnplateau, but a lot of top teams use cosine annealing.",
    "939046": "&gt; What I mean is that we fitted the models on the full 100% training data without any validation fold and just looked at training score.\n\n@philippsinger  Would you mind telling me what is the best way to tell when to stop when doing full 100% training data without validation ? Should we keep the number of epoch the same as in fold 0, or should we see the loss score or accuracy score until it reach the same number as fold 0 ? \n\n&gt;  If we were unsure, we re-fitted on a fold on a new model and checked if the models overfit or not, and then trusted the full fits.\n\nDoes this mean taking the fulldata checkpoint and refit it on a fold ? (in this case for how many epoch and at which learning rate)\n\nThank you very much for the write-ups and congratulations on the strong finish !\n",
    "938222": "@philippsinger \nhi, thank you for the awesome and straightforward write-up. A query, I didn't catch one of your words; `fitting models only on fold 0, but then switched to fitting models \"blindly\" on full data...` I guess, you had N fold, among them you've fitted the model in fold no. (N-1) and validate on Nth fold! Can u please clarify?\n\nReducing stride looks like an almost common approach like others did. As you said, ` this is specifically helpful as a lot of information about the manipulation lies between neighboring pixels.`..reducing stride makes lots sense.",
    "943674": "",
    "942403": "thanks for sharing your experience"
  }
}