{
  "id": 168510,
  "title": "Private LB 13th Solution - so near yet so far :)",
  "url": "/competitions/alaska2-image-steganalysis/writeups/all-data-are-ext-private-lb-13th-solution-so-near-",
  "author_name": "",
  "post_date": "2020-07-21T01:07:33.023Z",
  "votes": 41,
  "comment_count": 7,
  "views": 0,
  "content": "<p>First of all, congratulation to all, and especially to the prize winner and gold medalists.<br>\nSecondly, ouch, narrowly missing out on gold.</p>\n<p>So here are some essential points of our solutions:</p>\n<ul>\n<li><p>DL framework: Pytorch</p></li>\n<li><p>Used Architecture: Effnet(timm/geffnet) B0/B2/B3/B4/B5, <a href=\"https://github.com/clovaai/rexnet\" target=\"_blank\">Rexnet</a> 1.3/1.5/2.0 </p></li>\n<li><p>Model customization: <a href=\"https://www.kaggle.com/haqishen\" target=\"_blank\">@haqishen</a>  and <a href=\"https://www.kaggle.com/garybios\" target=\"_blank\">@garybios</a> did some amazing work in adding convolution head and SE head to our chosen architecture, and dropping the last two blocks from before the final FC output also helps - I will let them expand on these. </p></li>\n<li><p>Upsampling Cover: We find that upsampling cover image to match amount to Stegos helps to improve both CV and LB</p></li>\n<li><p>Method of ensemble: we have trained about 80+ models, we have taken the top models with top 6CV in each fold, and did unweighted gmean </p></li>\n</ul>\n<p>We were particularly careful about overfitting the LB so throughout the competition we didn't use it for feedback. But in general, we observed that for the same fold of data, better CV usually (not always) yields better LB</p>\n<p>Best Single Model - B5 with upsampling covers, added convolution heads, drop blocks -&gt; Private LB 0.927, Public LB 0.928</p>\n<p>Congratulation to my deal teammates <a href=\"https://www.kaggle.com/garybios\" target=\"_blank\">@garybios</a> <a href=\"https://www.kaggle.com/haqishen\" target=\"_blank\">@haqishen</a>  <a href=\"https://www.kaggle.com/strideradu\" target=\"_blank\">@strideradu</a> and <a href=\"https://www.kaggle.com/yl1202\" target=\"_blank\">@yl1202</a> for the effort, I really had a great time in this short and yet intensive competition! :)</p>",
  "messages": [
    {
      "id": "937359",
      "postDate": "07/21/2020 00:55:50",
      "content": "<p>First of all, congratulation to all, and especially to the prize winner and gold medalists.<br>\nSecondly, ouch, narrowly missing out on gold.</p>\n<p>So here are some essential points of our solutions:</p>\n<ul>\n<li><p>DL framework: Pytorch</p></li>\n<li><p>Used Architecture: Effnet(timm/geffnet) B0/B2/B3/B4/B5, <a href=\"https://github.com/clovaai/rexnet\" target=\"_blank\">Rexnet</a> 1.3/1.5/2.0 </p></li>\n<li><p>Model customization: <a href=\"https://www.kaggle.com/haqishen\" target=\"_blank\">@haqishen</a>  and <a href=\"https://www.kaggle.com/garybios\" target=\"_blank\">@garybios</a> did some amazing work in adding convolution head and SE head to our chosen architecture, and dropping the last two blocks from before the final FC output also helps - I will let them expand on these. </p></li>\n<li><p>Upsampling Cover: We find that upsampling cover image to match amount to Stegos helps to improve both CV and LB</p></li>\n<li><p>Method of ensemble: we have trained about 80+ models, we have taken the top models with top 6CV in each fold, and did unweighted gmean </p></li>\n</ul>\n<p>We were particularly careful about overfitting the LB so throughout the competition we didn't use it for feedback. But in general, we observed that for the same fold of data, better CV usually (not always) yields better LB</p>\n<p>Best Single Model - B5 with upsampling covers, added convolution heads, drop blocks -&gt; Private LB 0.927, Public LB 0.928</p>\n<p>Congratulation to my deal teammates <a href=\"https://www.kaggle.com/garybios\" target=\"_blank\">@garybios</a> <a href=\"https://www.kaggle.com/haqishen\" target=\"_blank\">@haqishen</a>  <a href=\"https://www.kaggle.com/strideradu\" target=\"_blank\">@strideradu</a> and <a href=\"https://www.kaggle.com/yl1202\" target=\"_blank\">@yl1202</a> for the effort, I really had a great time in this short and yet intensive competition! :)</p>",
      "rawMarkdown": "First of all, congratulation to all, and especially to the prize winner and gold medalists.\nSecondly, ouch, narrowly missing out on gold.\n\nSo here are some essential points of our solutions:\n\n- DL framework: Pytorch\n\n- Used Architecture: Effnet(timm/geffnet) B0/B2/B3/B4/B5, [Rexnet](https://github.com/clovaai/rexnet) 1.3/1.5/2.0 \n\n- Model customization: @haqishen  and @garybios did some amazing work in adding convolution head and SE head to our chosen architecture, and dropping the last two blocks from before the final FC output also helps - I will let them expand on these. \n\n- Upsampling Cover: We find that upsampling cover image to match amount to Stegos helps to improve both CV and LB\n\n- Method of ensemble: we have trained about 80+ models, we have taken the top models with top 6CV in each fold, and did unweighted gmean \n\nWe were particularly careful about overfitting the LB so throughout the competition we didn't use it for feedback. But in general, we observed that for the same fold of data, better CV usually (not always) yields better LB\n\nBest Single Model - B5 with upsampling covers, added convolution heads, drop blocks -&gt; Private LB 0.927, Public LB 0.928\n\nCongratulation to my deal teammates @garybios @haqishen  @strideradu and @yl1202 for the effort, I really had a great time in this short and yet intensive competition! :)",
      "votes": null
    },
    {
      "id": "937428",
      "postDate": "07/21/2020 02:41:09",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/yifanxie\" target=\"_blank\">@yifanxie</a> ,</p>\n<p>Congratulations for the brilliant work. Training 80+ model says a lot. Where exactly did you train the model (Colab/Kaggle GPU/TPU) ? How long did it take for each model to train ?</p>",
      "rawMarkdown": "Hi @yifanxie ,\n\nCongratulations for the brilliant work. Training 80+ model says a lot. Where exactly did you train the model (Colab/Kaggle GPU/TPU) ? How long did it take for each model to train ?",
      "votes": null
    },
    {
      "id": "937687",
      "postDate": "07/21/2020 05:58:47",
      "content": "<p>Congrats <a href=\"https://www.kaggle.com/yifanxie\" target=\"_blank\">@yifanxie</a> and team</p>",
      "rawMarkdown": "Congrats @yifanxie and team",
      "votes": null
    },
    {
      "id": "938213",
      "postDate": "07/21/2020 12:05:44",
      "content": "<blockquote>\n  <p>Hi <a href=\"https://www.kaggle.com/yifanxie\" target=\"_blank\">@yifanxie</a> ,</p>\n  <p>Congratulations for the brilliant work. Training 80+ model says a lot. Where exactly did you train the model (Colab/Kaggle GPU/TPU) ? How long did it take for each model to train ?</p>\n</blockquote>\n<p>Just to clarify - when I say 80+ models, I am talking about during the whole competition that we trained about 80 models, this includes the very first baseline type of models that is shared on the kernel.<br>\nAs I said our 1st solution has 30 models - 6 models for each fold. </p>\n<p>We used a combination of personal own GPUs and cloud GPUs via Vast.ai. While I agree hardware is an import factor in this competition, I would say it is not the major factor - as shown in both 1st and 2nd ranked solution there are domain-specific information that we could use </p>\n<p>As mentioned by <a href=\"https://www.kaggle.com/haqishen\" target=\"_blank\">@haqishen</a> <a href=\"https://www.kaggle.com/c/alaska2-image-steganalysis/discussion/168542\" target=\"_blank\">here</a> even with B0 we can achieve quite good CV/LB scores, and there is quite a lot of skilful work in there by my teammates  :)</p>",
      "rawMarkdown": "&gt; Hi @yifanxie ,\n&gt; \n&gt; Congratulations for the brilliant work. Training 80+ model says a lot. Where exactly did you train the model (Colab/Kaggle GPU/TPU) ? How long did it take for each model to train ?\n\nJust to clarify - when I say 80+ models, I am talking about during the whole competition that we trained about 80 models, this includes the very first baseline type of models that is shared on the kernel.\nAs I said our 1st solution has 30 models - 6 models for each fold. \n\nWe used a combination of personal own GPUs and cloud GPUs via Vast.ai. While I agree hardware is an import factor in this competition, I would say it is not the major factor - as shown in both 1st and 2nd ranked solution there are domain-specific information that we could use \n\nAs mentioned by @haqishen [here](https://www.kaggle.com/c/alaska2-image-steganalysis/discussion/168542) even with B0 we can achieve quite good CV/LB scores, and there is quite a lot of skilful work in there by my teammates  :)",
      "votes": null
    },
    {
      "id": "938301",
      "postDate": "07/21/2020 12:49:29",
      "content": "<p>congrats <a href=\"/yifanxie\">@yifanxie</a> and team! You guys have a solid solution. </p>",
      "rawMarkdown": "congrats @yifanxie and team! You guys have a solid solution.",
      "votes": null
    },
    {
      "id": "938315",
      "postDate": "07/21/2020 12:59:21",
      "content": "<p><a href=\"https://www.kaggle.com/yifanxie\" target=\"_blank\">@yifanxie</a> congratulations =)</p>",
      "rawMarkdown": "yifanxie congratulations =)",
      "votes": null
    },
    {
      "id": "938544",
      "postDate": "07/21/2020 15:45:02",
      "content": "<p>Congrats <a href=\"/yifanxie\">@yifanxie</a> . Great work. Could you elaborate on 'B5 with upsampling covers, added covolution head, drop blocks'.</p>",
      "rawMarkdown": "Congrats @yifanxie . Great work. Could you elaborate on 'B5 with upsampling covers, added covolution head, drop blocks'.",
      "votes": null
    },
    {
      "id": "945483",
      "postDate": "07/25/2020 21:15:50",
      "content": "<blockquote>\n  <p><strong>Kalyan Kumar Pichuka wrote:</strong></p>\n  \n  <p>Congrats <a href=\"/yifanxie\">@yifanxie</a> . Great work. Could you elaborate on 'B5 with upsampling covers, added convolution head, drop blocks'.</p>\n</blockquote>\n\n<p>thanks. For the convolution head and drop blocks stuff, refer to <a href=\"/haqishen\">@haqishen</a> 's post <a href=\"https://www.kaggle.com/c/alaska2-image-steganalysis/discussion/168542\">here</a> </p>\n\n<p>For upsampling, we were training on 12 classes, but we kept the original binary class, and use that to make sure at each epoch the amount of Cover and Stego images involved are balanced.  We found this bring out better CV although with longer training cycle.</p>",
      "rawMarkdown": "&gt; **Kalyan Kumar Pichuka wrote:**\n&gt; \n&gt; Congrats @yifanxie . Great work. Could you elaborate on 'B5 with upsampling covers, added convolution head, drop blocks'.\n\nthanks. For the convolution head and drop blocks stuff, refer to @haqishen 's post [here](https://www.kaggle.com/c/alaska2-image-steganalysis/discussion/168542) \n\nFor upsampling, we were training on 12 classes, but we kept the original binary class, and use that to make sure at each epoch the amount of Cover and Stego images involved are balanced.  We found this bring out better CV although with longer training cycle.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 937428,
      "author_name": "shwetank3",
      "author_url": "",
      "post_date": "07/21/2020 02:41:09",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/yifanxie\" target=\"_blank\">@yifanxie</a> ,</p>\n<p>Congratulations for the brilliant work. Training 80+ model says a lot. Where exactly did you train the model (Colab/Kaggle GPU/TPU) ? How long did it take for each model to train ?</p>",
      "votes": null,
      "replies": [
        {
          "id": 938213,
          "author_name": "yifanxie",
          "author_url": "",
          "post_date": "07/21/2020 12:05:44",
          "content": "<blockquote>\n  <p>Hi <a href=\"https://www.kaggle.com/yifanxie\" target=\"_blank\">@yifanxie</a> ,</p>\n  <p>Congratulations for the brilliant work. Training 80+ model says a lot. Where exactly did you train the model (Colab/Kaggle GPU/TPU) ? How long did it take for each model to train ?</p>\n</blockquote>\n<p>Just to clarify - when I say 80+ models, I am talking about during the whole competition that we trained about 80 models, this includes the very first baseline type of models that is shared on the kernel.<br>\nAs I said our 1st solution has 30 models - 6 models for each fold. </p>\n<p>We used a combination of personal own GPUs and cloud GPUs via Vast.ai. While I agree hardware is an import factor in this competition, I would say it is not the major factor - as shown in both 1st and 2nd ranked solution there are domain-specific information that we could use </p>\n<p>As mentioned by <a href=\"https://www.kaggle.com/haqishen\" target=\"_blank\">@haqishen</a> <a href=\"https://www.kaggle.com/c/alaska2-image-steganalysis/discussion/168542\" target=\"_blank\">here</a> even with B0 we can achieve quite good CV/LB scores, and there is quite a lot of skilful work in there by my teammates  :)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 937687,
      "author_name": "vishnurapps",
      "author_url": "",
      "post_date": "07/21/2020 05:58:47",
      "content": "<p>Congrats <a href=\"https://www.kaggle.com/yifanxie\" target=\"_blank\">@yifanxie</a> and team</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 938315,
      "author_name": "ipythonx",
      "author_url": "",
      "post_date": "07/21/2020 12:59:21",
      "content": "<p><a href=\"https://www.kaggle.com/yifanxie\" target=\"_blank\">@yifanxie</a> congratulations =)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 938301,
      "author_name": "titericz",
      "author_url": "",
      "post_date": "07/21/2020 12:49:29",
      "content": "<p>congrats <a href=\"/yifanxie\">@yifanxie</a> and team! You guys have a solid solution. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 938544,
      "author_name": "kalyanpichuka",
      "author_url": "",
      "post_date": "07/21/2020 15:45:02",
      "content": "<p>Congrats <a href=\"/yifanxie\">@yifanxie</a> . Great work. Could you elaborate on 'B5 with upsampling covers, added covolution head, drop blocks'.</p>",
      "votes": null,
      "replies": [
        {
          "id": 945483,
          "author_name": "yifanxie",
          "author_url": "",
          "post_date": "07/25/2020 21:15:50",
          "content": "<blockquote>\n  <p><strong>Kalyan Kumar Pichuka wrote:</strong></p>\n  \n  <p>Congrats <a href=\"/yifanxie\">@yifanxie</a> . Great work. Could you elaborate on 'B5 with upsampling covers, added convolution head, drop blocks'.</p>\n</blockquote>\n\n<p>thanks. For the convolution head and drop blocks stuff, refer to <a href=\"/haqishen\">@haqishen</a> 's post <a href=\"https://www.kaggle.com/c/alaska2-image-steganalysis/discussion/168542\">here</a> </p>\n\n<p>For upsampling, we were training on 12 classes, but we kept the original binary class, and use that to make sure at each epoch the amount of Cover and Stego images involved are balanced.  We found this bring out better CV although with longer training cycle.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "937359": "First of all, congratulation to all, and especially to the prize winner and gold medalists.\nSecondly, ouch, narrowly missing out on gold.\n\nSo here are some essential points of our solutions:\n\n- DL framework: Pytorch\n\n- Used Architecture: Effnet(timm/geffnet) B0/B2/B3/B4/B5, [Rexnet](https://github.com/clovaai/rexnet) 1.3/1.5/2.0 \n\n- Model customization: @haqishen  and @garybios did some amazing work in adding convolution head and SE head to our chosen architecture, and dropping the last two blocks from before the final FC output also helps - I will let them expand on these. \n\n- Upsampling Cover: We find that upsampling cover image to match amount to Stegos helps to improve both CV and LB\n\n- Method of ensemble: we have trained about 80+ models, we have taken the top models with top 6CV in each fold, and did unweighted gmean \n\nWe were particularly careful about overfitting the LB so throughout the competition we didn't use it for feedback. But in general, we observed that for the same fold of data, better CV usually (not always) yields better LB\n\nBest Single Model - B5 with upsampling covers, added convolution heads, drop blocks -&gt; Private LB 0.927, Public LB 0.928\n\nCongratulation to my deal teammates @garybios @haqishen  @strideradu and @yl1202 for the effort, I really had a great time in this short and yet intensive competition! :)",
    "937428": "Hi @yifanxie ,\n\nCongratulations for the brilliant work. Training 80+ model says a lot. Where exactly did you train the model (Colab/Kaggle GPU/TPU) ? How long did it take for each model to train ?",
    "937687": "Congrats @yifanxie and team",
    "938213": "&gt; Hi @yifanxie ,\n&gt; \n&gt; Congratulations for the brilliant work. Training 80+ model says a lot. Where exactly did you train the model (Colab/Kaggle GPU/TPU) ? How long did it take for each model to train ?\n\nJust to clarify - when I say 80+ models, I am talking about during the whole competition that we trained about 80 models, this includes the very first baseline type of models that is shared on the kernel.\nAs I said our 1st solution has 30 models - 6 models for each fold. \n\nWe used a combination of personal own GPUs and cloud GPUs via Vast.ai. While I agree hardware is an import factor in this competition, I would say it is not the major factor - as shown in both 1st and 2nd ranked solution there are domain-specific information that we could use \n\nAs mentioned by @haqishen [here](https://www.kaggle.com/c/alaska2-image-steganalysis/discussion/168542) even with B0 we can achieve quite good CV/LB scores, and there is quite a lot of skilful work in there by my teammates  :)",
    "938301": "congrats @yifanxie and team! You guys have a solid solution.",
    "938315": "yifanxie congratulations =)",
    "938544": "Congrats @yifanxie . Great work. Could you elaborate on 'B5 with upsampling covers, added covolution head, drop blocks'.",
    "945483": "&gt; **Kalyan Kumar Pichuka wrote:**\n&gt; \n&gt; Congrats @yifanxie . Great work. Could you elaborate on 'B5 with upsampling covers, added convolution head, drop blocks'.\n\nthanks. For the convolution head and drop blocks stuff, refer to @haqishen 's post [here](https://www.kaggle.com/c/alaska2-image-steganalysis/discussion/168542) \n\nFor upsampling, we were training on 12 classes, but we kept the original binary class, and use that to make sure at each epoch the amount of Cover and Stego images involved are balanced.  We found this bring out better CV although with longer training cycle."
  },
  "source": "meta"
}