{
  "id": 168611,
  "title": " congrats to Guanshuo Xu! and [14th place] solution ",
  "url": "/competitions/alaska2-image-steganalysis/discussion/168611",
  "author_name": "Μαριος Μιχαηλιδης KazAnova",
  "post_date": "2020-07-21T08:27:05.949000",
  "votes": 47,
  "comment_count": 13,
  "views": 0,
  "content": "<p>This was a nice competition with lots of ideas and sharing on the forums - thank you to kaggle and the organisers for  hosting it. </p>\n<p>I would like to congratulate my colleague (at <a href=\"https://www.h2o.ai/\" target=\"_blank\">h2o.ai</a>) ,<a href=\"https://www.kaggle.com/wowfattie\" target=\"_blank\">Guanshuo Xu</a> <strong>for winning the competition and for becoming the new kaggle #1</strong>. It takes a lot of hard work and dedication to achieve this - well done!</p>\n<p>My solution is fairly simple. Where I think I did really well is I never became emotional and never tried things that I did not see them working in validation first. </p>\n<p>I used a single holdout (20%)  for validation (stratified by type) . I trained 2 models - an EfficientNet b3 and b4 on the training (80%) part of the data AND 2 more models on 100% of the data. So 4 models in total. Using a cosine learning schedule, all models (small and full ones), had very similar behaviour. </p>\n<p>I generated multiple checkpoints predictions (over 10) from each one of these models. I progressively added more augmetations, swapped optimizers , (lowered) learning rates and batch sizes. I would make changes every-time I would see the validation performance getting halted. Every single time, I did get improvements from these swaps. In total, I trained the b3 model for 150 epochs and the b4 120 epochs. </p>\n<p>Augmentations in stages:</p>\n<ul>\n<li>Vertical and horizontal flips</li>\n<li>Vertical and horizontal + transpose + rotate</li>\n<li>Vertical and horizontal + transpose  + rotate + cutout (1 hole, 80)</li>\n<li>Vertical and horizontal + transpose + rotate + cutout (2 holes, 64)</li>\n<li>Vertical and horizontal + transpose + rotate  + cutout (4 holes, 64) </li>\n</ul>\n<p>For TTA I used  Vertical , horizontal and  Vertical + horizontal </p>\n<p>my models are all in pytorch . At some point, I tried to run EfficientNetb3 in keras (TF) with same batch size, optimmizers and augmentations and performance was significantly worse (not sure why) .</p>\n<p>I got a lot of gain from stacking - around +0.003-4 in LB. For test predictions in stacking, I had 25% the predictions generated from the small(80%) model and 75% of the model trained with 100% of the data. </p>\n<p>NN with 2 hidden layers, leakyrelu and a bit of l2 regularization had the best cv for me (0.936). <a href=\"https://lightgbm.readthedocs.io/en/latest/Parameters.html\" target=\"_blank\">Lightgbm </a>with dart was a close second. <a href=\"http://scikit-learn.org/stable/modules/generated/sklearn.ensemble.ExtraTreesClassifier.html\" target=\"_blank\">EextraTreesClassificer  </a> added a little bit too. </p>",
  "messages": [
    {
      "id": 937902,
      "postDate": "2020-07-21T08:27:05.950Z",
      "content": "<p>This was a nice competition with lots of ideas and sharing on the forums - thank you to kaggle and the organisers for  hosting it. </p>\n<p>I would like to congratulate my colleague (at <a href=\"https://www.h2o.ai/\" target=\"_blank\">h2o.ai</a>) ,<a href=\"https://www.kaggle.com/wowfattie\" target=\"_blank\">Guanshuo Xu</a> <strong>for winning the competition and for becoming the new kaggle #1</strong>. It takes a lot of hard work and dedication to achieve this - well done!</p>\n<p>My solution is fairly simple. Where I think I did really well is I never became emotional and never tried things that I did not see them working in validation first. </p>\n<p>I used a single holdout (20%)  for validation (stratified by type) . I trained 2 models - an EfficientNet b3 and b4 on the training (80%) part of the data AND 2 more models on 100% of the data. So 4 models in total. Using a cosine learning schedule, all models (small and full ones), had very similar behaviour. </p>\n<p>I generated multiple checkpoints predictions (over 10) from each one of these models. I progressively added more augmetations, swapped optimizers , (lowered) learning rates and batch sizes. I would make changes every-time I would see the validation performance getting halted. Every single time, I did get improvements from these swaps. In total, I trained the b3 model for 150 epochs and the b4 120 epochs. </p>\n<p>Augmentations in stages:</p>\n<ul>\n<li>Vertical and horizontal flips</li>\n<li>Vertical and horizontal + transpose + rotate</li>\n<li>Vertical and horizontal + transpose  + rotate + cutout (1 hole, 80)</li>\n<li>Vertical and horizontal + transpose + rotate + cutout (2 holes, 64)</li>\n<li>Vertical and horizontal + transpose + rotate  + cutout (4 holes, 64) </li>\n</ul>\n<p>For TTA I used  Vertical , horizontal and  Vertical + horizontal </p>\n<p>my models are all in pytorch . At some point, I tried to run EfficientNetb3 in keras (TF) with same batch size, optimmizers and augmentations and performance was significantly worse (not sure why) .</p>\n<p>I got a lot of gain from stacking - around +0.003-4 in LB. For test predictions in stacking, I had 25% the predictions generated from the small(80%) model and 75% of the model trained with 100% of the data. </p>\n<p>NN with 2 hidden layers, leakyrelu and a bit of l2 regularization had the best cv for me (0.936). <a href=\"https://lightgbm.readthedocs.io/en/latest/Parameters.html\" target=\"_blank\">Lightgbm </a>with dart was a close second. <a href=\"http://scikit-learn.org/stable/modules/generated/sklearn.ensemble.ExtraTreesClassifier.html\" target=\"_blank\">EextraTreesClassificer  </a> added a little bit too. </p>",
      "rawMarkdown": "This was a nice competition with lots of ideas and sharing on the forums - thank you to kaggle and the organisers for  hosting it. \n\nI would like to congratulate my colleague (at [h2o.ai](https://www.h2o.ai/)) ,[Guanshuo Xu](https://www.kaggle.com/wowfattie) **for winning the competition and for becoming the new kaggle #1**. It takes a lot of hard work and dedication to achieve this - well done!\n\nMy solution is fairly simple. Where I think I did really well is I never became emotional and never tried things that I did not see them working in validation first. \n\nI used a single holdout (20%)  for validation (stratified by type) . I trained 2 models - an EfficientNet b3 and b4 on the training (80%) part of the data AND 2 more models on 100% of the data. So 4 models in total. Using a cosine learning schedule, all models (small and full ones), had very similar behaviour. \n\nI generated multiple checkpoints predictions (over 10) from each one of these models. I progressively added more augmetations, swapped optimizers , (lowered) learning rates and batch sizes. I would make changes every-time I would see the validation performance getting halted. Every single time, I did get improvements from these swaps. In total, I trained the b3 model for 150 epochs and the b4 120 epochs. \n\nAugmentations in stages:\n\n- Vertical and horizontal flips\n- Vertical and horizontal + transpose + rotate\n- Vertical and horizontal + transpose  + rotate + cutout (1 hole, 80)\n- Vertical and horizontal + transpose + rotate + cutout (2 holes, 64)\n- Vertical and horizontal + transpose + rotate  + cutout (4 holes, 64) \n\nFor TTA I used  Vertical , horizontal and  Vertical + horizontal \n\nmy models are all in pytorch . At some point, I tried to run EfficientNetb3 in keras (TF) with same batch size, optimmizers and augmentations and performance was significantly worse (not sure why) .\n\nI got a lot of gain from stacking - around +0.003-4 in LB. For test predictions in stacking, I had 25% the predictions generated from the small(80%) model and 75% of the model trained with 100% of the data. \n\nNN with 2 hidden layers, leakyrelu and a bit of l2 regularization had the best cv for me (0.936). [Lightgbm ](https://lightgbm.readthedocs.io/en/latest/Parameters.html)with dart was a close second. [EextraTreesClassificer  ](http://scikit-learn.org/stable/modules/generated/sklearn.ensemble.ExtraTreesClassifier.html) added a little bit too. \n\n\n\n\n\n\n\n\n\n\n\n\n",
      "votes": 47
    },
    {
      "id": 938936,
      "postDate": "2020-07-21T21:52:08.020Z",
      "content": "<p>Congrats !! I am wondering about stacking part, can you provide a snippet how to do it?</p>",
      "rawMarkdown": "Congrats !! I am wondering about stacking part, can you provide a snippet how to do it?",
      "votes": 1,
      "replies": [
        {
          "id": 938955,
          "postDate": "2020-07-21T22:25:23.980Z",
          "content": "<p>Not sure if you have access to <a href=\"https://www.coursera.org/lecture/competitive-data-science/stacking-Qdtt6\" target=\"_blank\">the following video</a>. It explains stacking (with a code example) around minute 7. </p>\n<p>Disclaimer: I am not good at explaining things in videos (or any) , but with some effort, you can figure it out!</p>",
          "rawMarkdown": "Not sure if you have access to [the following video](https://www.coursera.org/lecture/competitive-data-science/stacking-Qdtt6). It explains stacking (with a code example) around minute 7. \n\nDisclaimer: I am not good at explaining things in videos (or any) , but with some effort, you can figure it out!",
          "votes": 1
        },
        {
          "id": 938964,
          "postDate": "2020-07-21T22:41:46.660Z",
          "content": "<p>yes I can see the video. Thanks a lot 👍</p>",
          "rawMarkdown": "yes I can see the video. Thanks a lot 👍",
          "votes": 1
        }
      ]
    },
    {
      "id": 938843,
      "postDate": "2020-07-21T19:51:02.510Z",
      "content": "<p>nicely done <a href=\"https://www.kaggle.com/kazanova\" target=\"_blank\">@kazanova</a>! great to see you around the same LB   :)</p>",
      "rawMarkdown": "nicely done @kazanova! great to see you around the same LB   :)",
      "votes": 1
    },
    {
      "id": 938311,
      "postDate": "2020-07-21T12:57:22.217Z",
      "content": "<p>congrats <a href=\"/kazanova\">@kazanova</a> for the solo #14. Nice solution. </p>",
      "rawMarkdown": "congrats @kazanova for the solo #14. Nice solution. ",
      "votes": 1
    },
    {
      "id": 938291,
      "postDate": "2020-07-21T12:40:58.610Z",
      "content": "<p><a href=\"https://www.kaggle.com/kazanova\" target=\"_blank\">@kazanova</a><br>\nhello, thank u for ur write-up. You didn't do any model surgery like make stride 2 to 1, as other mention in the solution. Is there any particular reason? And what do you think the behind cause of different results of two different frameworks with same setup?</p>",
      "rawMarkdown": "@kazanova\nhello, thank u for ur write-up. You didn't do any model surgery like make stride 2 to 1, as other mention in the solution. Is there any particular reason? And what do you think the behind cause of different results of two different frameworks with same setup?",
      "votes": 1,
      "replies": [
        {
          "id": 938381,
          "postDate": "2020-07-21T13:44:33.643Z",
          "content": "<p>i did no model surgery - because I joined pretty late - around the 2 week mark and did not have much time to investigate these . My models are plain simple EfficientNets with nothing special in  them. </p>",
          "rawMarkdown": "i did no model surgery - because I joined pretty late - around the 2 week mark and did not have much time to investigate these . My models are plain simple EfficientNets with nothing special in  them. ",
          "votes": 2
        },
        {
          "id": 938382,
          "postDate": "2020-07-21T13:46:26.903Z",
          "content": "<blockquote>\n  <p>And what do you think the behind cause of different results of two different frameworks with same setup?</p>\n</blockquote>\n<p>Maybe something different in the weight initialisation? Otherwise, I have no idea why TF was that much worse…</p>",
          "rawMarkdown": "&gt; And what do you think the behind cause of different results of two different frameworks with same setup?\n\nMaybe something different in the weight initialisation? Otherwise, I have no idea why TF was that much worse...",
          "votes": 1
        },
        {
          "id": 938512,
          "postDate": "2020-07-21T15:23:30.833Z",
          "content": "<p>do note that the padding of tf and pytorch is different. \nalso, batch norm parameters are different (eps and momentum)\nfinally, the drop connect rate may be different in some pytorch implementation.</p>\n\n<p>not sure if that can be the cause?</p>",
          "rawMarkdown": "do note that the padding of tf and pytorch is different. \nalso, batch norm parameters are different (eps and momentum)\nfinally, the drop connect rate may be different in some pytorch implementation.\n\nnot sure if that can be the cause?",
          "votes": 2
        }
      ]
    },
    {
      "id": 937920,
      "postDate": "2020-07-21T08:37:52.043Z",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/kazanova\" target=\"_blank\">@kazanova</a> </p>\n<p>Congrats on the Victory 👏👏👏</p>\n<p>Sorry for asking a naive question but how did you ensure that the cutout didn't remove the parts of the image that had steg information. </p>",
      "rawMarkdown": "Hi @kazanova \n\nCongrats on the Victory 👏👏👏\n\nSorry for asking a naive question but how did you ensure that the cutout didn't remove the parts of the image that had steg information. ",
      "votes": 1,
      "replies": [
        {
          "id": 937923,
          "postDate": "2020-07-21T08:39:02.910Z",
          "content": "<p>I didn't !!! just normal cutout from albumentations!</p>",
          "rawMarkdown": "I didn't !!! just normal cutout from albumentations!",
          "votes": 2
        }
      ]
    },
    {
      "id": 969906,
      "postDate": "2020-08-14T03:49:54.503Z",
      "content": "<p>congrats <a href=\"https://www.kaggle.com/kazanova\" target=\"_blank\">@kazanova</a> 🙌🏻🙌🏻🙌🏻</p>",
      "rawMarkdown": "congrats @kazanova 🙌🏻🙌🏻🙌🏻"
    },
    {
      "id": 943680,
      "postDate": "2020-07-24T14:11:30.483Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 938936,
      "author_name": "Phi",
      "author_url": "",
      "post_date": "2020-07-21T21:52:08.020000",
      "content": "<p>Congrats !! I am wondering about stacking part, can you provide a snippet how to do it?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 938955,
          "author_name": "Μαριος Μιχαηλιδης KazAnova",
          "author_url": "",
          "post_date": "2020-07-21T22:25:23.980000",
          "content": "<p>Not sure if you have access to <a href=\"https://www.coursera.org/lecture/competitive-data-science/stacking-Qdtt6\" target=\"_blank\">the following video</a>. It explains stacking (with a code example) around minute 7. </p>\n<p>Disclaimer: I am not good at explaining things in videos (or any) , but with some effort, you can figure it out!</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 938964,
          "author_name": "Phi",
          "author_url": "",
          "post_date": "2020-07-21T22:41:46.660000",
          "content": "<p>yes I can see the video. Thanks a lot 👍</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 938843,
      "author_name": "Yifan Xie",
      "author_url": "",
      "post_date": "2020-07-21T19:51:02.510000",
      "content": "<p>nicely done <a href=\"https://www.kaggle.com/kazanova\" target=\"_blank\">@kazanova</a>! great to see you around the same LB   :)</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 938311,
      "author_name": "Giba",
      "author_url": "",
      "post_date": "2020-07-21T12:57:22.217000",
      "content": "<p>congrats <a href=\"/kazanova\">@kazanova</a> for the solo #14. Nice solution. </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 938291,
      "author_name": "Innat",
      "author_url": "",
      "post_date": "2020-07-21T12:40:58.610000",
      "content": "<p><a href=\"https://www.kaggle.com/kazanova\" target=\"_blank\">@kazanova</a><br>\nhello, thank u for ur write-up. You didn't do any model surgery like make stride 2 to 1, as other mention in the solution. Is there any particular reason? And what do you think the behind cause of different results of two different frameworks with same setup?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 938381,
          "author_name": "Μαριος Μιχαηλιδης KazAnova",
          "author_url": "",
          "post_date": "2020-07-21T13:44:33.643000",
          "content": "<p>i did no model surgery - because I joined pretty late - around the 2 week mark and did not have much time to investigate these . My models are plain simple EfficientNets with nothing special in  them. </p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 938382,
          "author_name": "Μαριος Μιχαηλιδης KazAnova",
          "author_url": "",
          "post_date": "2020-07-21T13:46:26.903000",
          "content": "<blockquote>\n  <p>And what do you think the behind cause of different results of two different frameworks with same setup?</p>\n</blockquote>\n<p>Maybe something different in the weight initialisation? Otherwise, I have no idea why TF was that much worse…</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 938512,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2020-07-21T15:23:30.833000",
          "content": "<p>do note that the padding of tf and pytorch is different. \nalso, batch norm parameters are different (eps and momentum)\nfinally, the drop connect rate may be different in some pytorch implementation.</p>\n\n<p>not sure if that can be the cause?</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 937920,
      "author_name": "Vishnu R",
      "author_url": "",
      "post_date": "2020-07-21T08:37:52.043000",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/kazanova\" target=\"_blank\">@kazanova</a> </p>\n<p>Congrats on the Victory 👏👏👏</p>\n<p>Sorry for asking a naive question but how did you ensure that the cutout didn't remove the parts of the image that had steg information. </p>",
      "votes": 1,
      "replies": [
        {
          "id": 937923,
          "author_name": "Μαριος Μιχαηλιδης KazAnova",
          "author_url": "",
          "post_date": "2020-07-21T08:39:02.910000",
          "content": "<p>I didn't !!! just normal cutout from albumentations!</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 969906,
      "author_name": "Santhoshkumar",
      "author_url": "",
      "post_date": "2020-08-14T03:49:54.503000",
      "content": "<p>congrats <a href=\"https://www.kaggle.com/kazanova\" target=\"_blank\">@kazanova</a> 🙌🏻🙌🏻🙌🏻</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 943680,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-07-24T14:11:30.483000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "937902": "This was a nice competition with lots of ideas and sharing on the forums - thank you to kaggle and the organisers for  hosting it. \n\nI would like to congratulate my colleague (at [h2o.ai](https://www.h2o.ai/)) ,[Guanshuo Xu](https://www.kaggle.com/wowfattie) **for winning the competition and for becoming the new kaggle #1**. It takes a lot of hard work and dedication to achieve this - well done!\n\nMy solution is fairly simple. Where I think I did really well is I never became emotional and never tried things that I did not see them working in validation first. \n\nI used a single holdout (20%)  for validation (stratified by type) . I trained 2 models - an EfficientNet b3 and b4 on the training (80%) part of the data AND 2 more models on 100% of the data. So 4 models in total. Using a cosine learning schedule, all models (small and full ones), had very similar behaviour. \n\nI generated multiple checkpoints predictions (over 10) from each one of these models. I progressively added more augmetations, swapped optimizers , (lowered) learning rates and batch sizes. I would make changes every-time I would see the validation performance getting halted. Every single time, I did get improvements from these swaps. In total, I trained the b3 model for 150 epochs and the b4 120 epochs. \n\nAugmentations in stages:\n\n- Vertical and horizontal flips\n- Vertical and horizontal + transpose + rotate\n- Vertical and horizontal + transpose  + rotate + cutout (1 hole, 80)\n- Vertical and horizontal + transpose + rotate + cutout (2 holes, 64)\n- Vertical and horizontal + transpose + rotate  + cutout (4 holes, 64) \n\nFor TTA I used  Vertical , horizontal and  Vertical + horizontal \n\nmy models are all in pytorch . At some point, I tried to run EfficientNetb3 in keras (TF) with same batch size, optimmizers and augmentations and performance was significantly worse (not sure why) .\n\nI got a lot of gain from stacking - around +0.003-4 in LB. For test predictions in stacking, I had 25% the predictions generated from the small(80%) model and 75% of the model trained with 100% of the data. \n\nNN with 2 hidden layers, leakyrelu and a bit of l2 regularization had the best cv for me (0.936). [Lightgbm ](https://lightgbm.readthedocs.io/en/latest/Parameters.html)with dart was a close second. [EextraTreesClassificer  ](http://scikit-learn.org/stable/modules/generated/sklearn.ensemble.ExtraTreesClassifier.html) added a little bit too. \n\n\n\n\n\n\n\n\n\n\n\n\n",
    "938936": "Congrats !! I am wondering about stacking part, can you provide a snippet how to do it?",
    "938843": "nicely done @kazanova! great to see you around the same LB   :)",
    "938311": "congrats @kazanova for the solo #14. Nice solution. ",
    "938291": "@kazanova\nhello, thank u for ur write-up. You didn't do any model surgery like make stride 2 to 1, as other mention in the solution. Is there any particular reason? And what do you think the behind cause of different results of two different frameworks with same setup?",
    "937920": "Hi @kazanova \n\nCongrats on the Victory 👏👏👏\n\nSorry for asking a naive question but how did you ensure that the cutout didn't remove the parts of the image that had steg information. ",
    "969906": "congrats @kazanova 🙌🏻🙌🏻🙌🏻",
    "943680": ""
  }
}