{
  "id": 168771,
  "title": "18th Place Brief Solution Overview",
  "url": "/competitions/alaska2-image-steganalysis/discussion/168771",
  "author_name": "Bojan Tunguz",
  "post_date": "2020-07-21T19:36:16.941000",
  "votes": 50,
  "comment_count": 10,
  "views": 0,
  "content": "<p>I would like to thank Kaggle for organizing such an interesting competition. Unfortunately, I had not paid much attention to it until just a couple of weeks ago, and am regretting not spending more time on it. Especially since models for the competition took <strong>really</strong> long time to train.</p>\n\n<p>Like many others, I decided to seriously join the competition after discovering the wonderful <a href=\"https://www.kaggle.com/shonenkov\">Alex Shonenkov</a> starter <a href=\"https://www.kaggle.com/shonenkov/train-inference-gpu-baseline\">kernel</a>. Most of my work was building upon his work, and experimenting with different network backbones, training schedules, ensembles, and to a lesser degree augmentations.</p>\n\n<p>My final ensemble consists of the following three EfficientNet model architectures:</p>\n\n<ul>\n<li><p>B1 - I have only trained one of these networks successfully. The best result can be found in <a href=\"https://www.kaggle.com/tunguz/best-b1-inference/\">this kernel</a>, and all the weights are in <a href=\"https://www.kaggle.com/tunguz/eb1-weights\">this dataset</a>. The best single model achieves 0.918 locally, 0.925 on public LB and 0.911 on private LB.</p></li>\n<li><p>B2 - I have trained several of these networks, and managed to get quite a bit of an improvement over Alex's original model. However, these networks have only made a minor contribution to my final blend. The best result can be found in <a href=\"https://www.kaggle.com/tunguz/best-b2-inference\">this kernel</a>, and all the weights are in <a href=\"https://www.kaggle.com/tunguz/eb2-weights/\">this dataset</a>. The best single model achieves 0.923 locally, 0.929 on public LB and 0.916 on private LB.</p></li>\n<li><p>B4 - These networks have been the main workhorses behind my solution. Unfortunately, they are really slow to train, and I pretty much pushed all of my compute resources to get the best solution with them. Did not have any additional capacity (compute + time) to train bigger models, which I suspect would have performed even better. The best result can be found in <a href=\"https://www.kaggle.com/tunguz/best-b4-inference\">this kernel</a>, and all the weights are in <a href=\"https://www.kaggle.com/tunguz/alaska2-eb4-model-weights\">this dataset</a>. The best single model achieves 0.930 locally, 0.940 on public LB and 0.925 on private LB.</p></li>\n</ul>\n\n<p><strong>Training Schedule</strong></p>\n\n<p>This is probably the biggest source of improvement for my networks over the base one that was found in the public kernel. I started with the same training schedule as in the original kernel, but then I <strong>re-trained</strong> all the weights - twice. The second time I decreased the LR to 0.0005, and increased patience to 2. I added JpegCompression to the augmentations, but otherwise did not change anything. I also increased training to 50 epochs, which took a really long time to go through, since each epoch would take about 2.5 hours on an V100 GPU.</p>\n\n<p><strong>Train/Val Split</strong></p>\n\n<p>For the most of the models I stuck to the original 80/20 split, but have used different seeds. The choice of split seems to have been very important in this competition - the final models varied between 0.933 and 0.940 on LB. I have also trained a couple of models with the 95/5 split, but since they were B2 networks, they did not contribute significantly to the final ensemble. </p>\n\n<p><strong>Final Blend</strong></p>\n\n<p>Since I've used different splits for my training, it was not possible to do a consistent local validation for the ensemble, so I was forced to rely on LB. Again, since I was only able to work on this competition for less than two weeks, that did not leave me a lot of submissions on which to base my validation strategy, so I had to rely on my intuition a lot. For instance, I did not use an equally weighted average for the blend of the same model on different splits, but based the weights on the individual model's performance on LB. It is highly likely that the weights I had finally used are suboptimal, but I feel fortunate that my two submissions were the best ones on private LB. </p>",
  "messages": [
    {
      "id": 938834,
      "postDate": "2020-07-21T19:36:16.940Z",
      "content": "<p>I would like to thank Kaggle for organizing such an interesting competition. Unfortunately, I had not paid much attention to it until just a couple of weeks ago, and am regretting not spending more time on it. Especially since models for the competition took <strong>really</strong> long time to train.</p>\n\n<p>Like many others, I decided to seriously join the competition after discovering the wonderful <a href=\"https://www.kaggle.com/shonenkov\">Alex Shonenkov</a> starter <a href=\"https://www.kaggle.com/shonenkov/train-inference-gpu-baseline\">kernel</a>. Most of my work was building upon his work, and experimenting with different network backbones, training schedules, ensembles, and to a lesser degree augmentations.</p>\n\n<p>My final ensemble consists of the following three EfficientNet model architectures:</p>\n\n<ul>\n<li><p>B1 - I have only trained one of these networks successfully. The best result can be found in <a href=\"https://www.kaggle.com/tunguz/best-b1-inference/\">this kernel</a>, and all the weights are in <a href=\"https://www.kaggle.com/tunguz/eb1-weights\">this dataset</a>. The best single model achieves 0.918 locally, 0.925 on public LB and 0.911 on private LB.</p></li>\n<li><p>B2 - I have trained several of these networks, and managed to get quite a bit of an improvement over Alex's original model. However, these networks have only made a minor contribution to my final blend. The best result can be found in <a href=\"https://www.kaggle.com/tunguz/best-b2-inference\">this kernel</a>, and all the weights are in <a href=\"https://www.kaggle.com/tunguz/eb2-weights/\">this dataset</a>. The best single model achieves 0.923 locally, 0.929 on public LB and 0.916 on private LB.</p></li>\n<li><p>B4 - These networks have been the main workhorses behind my solution. Unfortunately, they are really slow to train, and I pretty much pushed all of my compute resources to get the best solution with them. Did not have any additional capacity (compute + time) to train bigger models, which I suspect would have performed even better. The best result can be found in <a href=\"https://www.kaggle.com/tunguz/best-b4-inference\">this kernel</a>, and all the weights are in <a href=\"https://www.kaggle.com/tunguz/alaska2-eb4-model-weights\">this dataset</a>. The best single model achieves 0.930 locally, 0.940 on public LB and 0.925 on private LB.</p></li>\n</ul>\n\n<p><strong>Training Schedule</strong></p>\n\n<p>This is probably the biggest source of improvement for my networks over the base one that was found in the public kernel. I started with the same training schedule as in the original kernel, but then I <strong>re-trained</strong> all the weights - twice. The second time I decreased the LR to 0.0005, and increased patience to 2. I added JpegCompression to the augmentations, but otherwise did not change anything. I also increased training to 50 epochs, which took a really long time to go through, since each epoch would take about 2.5 hours on an V100 GPU.</p>\n\n<p><strong>Train/Val Split</strong></p>\n\n<p>For the most of the models I stuck to the original 80/20 split, but have used different seeds. The choice of split seems to have been very important in this competition - the final models varied between 0.933 and 0.940 on LB. I have also trained a couple of models with the 95/5 split, but since they were B2 networks, they did not contribute significantly to the final ensemble. </p>\n\n<p><strong>Final Blend</strong></p>\n\n<p>Since I've used different splits for my training, it was not possible to do a consistent local validation for the ensemble, so I was forced to rely on LB. Again, since I was only able to work on this competition for less than two weeks, that did not leave me a lot of submissions on which to base my validation strategy, so I had to rely on my intuition a lot. For instance, I did not use an equally weighted average for the blend of the same model on different splits, but based the weights on the individual model's performance on LB. It is highly likely that the weights I had finally used are suboptimal, but I feel fortunate that my two submissions were the best ones on private LB. </p>",
      "rawMarkdown": "I would like to thank Kaggle for organizing such an interesting competition. Unfortunately, I had not paid much attention to it until just a couple of weeks ago, and am regretting not spending more time on it. Especially since models for the competition took **really** long time to train.\n\nLike many others, I decided to seriously join the competition after discovering the wonderful [Alex Shonenkov](https://www.kaggle.com/shonenkov) starter [kernel](https://www.kaggle.com/shonenkov/train-inference-gpu-baseline). Most of my work was building upon his work, and experimenting with different network backbones, training schedules, ensembles, and to a lesser degree augmentations.\n\nMy final ensemble consists of the following three EfficientNet model architectures:\n\n* B1 - I have only trained one of these networks successfully. The best result can be found in [this kernel](https://www.kaggle.com/tunguz/best-b1-inference/), and all the weights are in [this dataset](https://www.kaggle.com/tunguz/eb1-weights). The best single model achieves 0.918 locally, 0.925 on public LB and 0.911 on private LB.\n\n* B2 - I have trained several of these networks, and managed to get quite a bit of an improvement over Alex's original model. However, these networks have only made a minor contribution to my final blend. The best result can be found in [this kernel](https://www.kaggle.com/tunguz/best-b2-inference), and all the weights are in [this dataset](https://www.kaggle.com/tunguz/eb2-weights/). The best single model achieves 0.923 locally, 0.929 on public LB and 0.916 on private LB.\n\n* B4 - These networks have been the main workhorses behind my solution. Unfortunately, they are really slow to train, and I pretty much pushed all of my compute resources to get the best solution with them. Did not have any additional capacity (compute + time) to train bigger models, which I suspect would have performed even better. The best result can be found in [this kernel](https://www.kaggle.com/tunguz/best-b4-inference), and all the weights are in [this dataset](https://www.kaggle.com/tunguz/alaska2-eb4-model-weights). The best single model achieves 0.930 locally, 0.940 on public LB and 0.925 on private LB.\n\n**Training Schedule**\n\nThis is probably the biggest source of improvement for my networks over the base one that was found in the public kernel. I started with the same training schedule as in the original kernel, but then I **re-trained** all the weights - twice. The second time I decreased the LR to 0.0005, and increased patience to 2. I added JpegCompression to the augmentations, but otherwise did not change anything. I also increased training to 50 epochs, which took a really long time to go through, since each epoch would take about 2.5 hours on an V100 GPU.\n\n**Train/Val Split**\n\nFor the most of the models I stuck to the original 80/20 split, but have used different seeds. The choice of split seems to have been very important in this competition - the final models varied between 0.933 and 0.940 on LB. I have also trained a couple of models with the 95/5 split, but since they were B2 networks, they did not contribute significantly to the final ensemble. \n\n**Final Blend**\n\nSince I've used different splits for my training, it was not possible to do a consistent local validation for the ensemble, so I was forced to rely on LB. Again, since I was only able to work on this competition for less than two weeks, that did not leave me a lot of submissions on which to base my validation strategy, so I had to rely on my intuition a lot. For instance, I did not use an equally weighted average for the blend of the same model on different splits, but based the weights on the individual model's performance on LB. It is highly likely that the weights I had finally used are suboptimal, but I feel fortunate that my two submissions were the best ones on private LB. \n\n\n",
      "votes": 50
    },
    {
      "id": 943289,
      "postDate": "2020-07-24T09:01:15.263Z",
      "content": "<p>Nice work. Always great to hear about approaches from the masters themselves. Learned from such discussions more than anything. Keep Sharing.😊 </p>",
      "rawMarkdown": "Nice work. Always great to hear about approaches from the masters themselves. Learned from such discussions more than anything. Keep Sharing.😊 ",
      "votes": 1
    },
    {
      "id": 939305,
      "postDate": "2020-07-22T06:49:18.603Z",
      "content": "<p>Thank you for the write ups, and congratulations on the result.</p>\n<blockquote>\n  <p>but then I re-trained all the weights - twice.</p>\n</blockquote>\n<p>Would you mind elaborating more on this ? Does it mean after first run (40 epoch ?), you restarted the training with different learning rate (for 50epoch?) and one more time restart adding jpeg compression ?<br>\nI have tried continuing the starter kernel for a long time, but it got trapped in a very small lr and never improve further.</p>",
      "rawMarkdown": "Thank you for the write ups, and congratulations on the result.\n\n&gt;  but then I re-trained all the weights - twice.\n\nWould you mind elaborating more on this ? Does it mean after first run (40 epoch ?), you restarted the training with different learning rate (for 50epoch?) and one more time restart adding jpeg compression ?\nI have tried continuing the starter kernel for a long time, but it got trapped in a very small lr and never improve further.\n",
      "votes": 1,
      "replies": [
        {
          "id": 939794,
          "postDate": "2020-07-22T13:15:59.293Z",
          "content": "<p>Yes, that's basically correct, except that the first time I restarted training with the <strong>same</strong> learning rate as the first time.</p>",
          "rawMarkdown": "Yes, that's basically correct, except that the first time I restarted training with the **same** learning rate as the first time.",
          "votes": 1
        }
      ]
    },
    {
      "id": 938874,
      "postDate": "2020-07-21T20:22:42.397Z",
      "content": "<p><a href=\"https://www.kaggle.com/tunguz\" target=\"_blank\">@tunguz</a> - excellent result :) Just curious did you also play around with the stride on the efficient net models? Also, curious to know what batch size you used?</p>",
      "rawMarkdown": "@tunguz - excellent result :) Just curious did you also play around with the stride on the efficient net models? Also, curious to know what batch size you used?",
      "votes": 1,
      "replies": [
        {
          "id": 938886,
          "postDate": "2020-07-21T20:37:23.050Z",
          "content": "<p>No, I did not find out about the stride trick until after the competition. </p>\n\n<p>Thanks for asking - the increased batch size certainly seemed to give me a boost. I was not able to go beyond batch size of 26 on a single GPU. Still figuring out how to effectively harness multi-GPU training, especially with PyTorch, which is still a foreign \"language\" to me. </p>",
          "rawMarkdown": "No, I did not find out about the stride trick until after the competition. \n\nThanks for asking - the increased batch size certainly seemed to give me a boost. I was not able to go beyond batch size of 26 on a single GPU. Still figuring out how to effectively harness multi-GPU training, especially with PyTorch, which is still a foreign \"language\" to me. ",
          "votes": 1
        }
      ]
    },
    {
      "id": 938841,
      "postDate": "2020-07-21T19:46:17.873Z",
      "content": "<p><a href=\"https://www.kaggle.com/shonenkov\">https://www.kaggle.com/shonenkov</a> work is awesome indeed! One question: haven't you tried other EfficientNet backbones (B5 and more)? Maybe it wasn't possible due to time and/or resources constraints? Congratulations on your medal anyway! </p>",
      "rawMarkdown": "https://www.kaggle.com/shonenkov work is awesome indeed! One question: haven't you tried other EfficientNet backbones (B5 and more)? Maybe it wasn't possible due to time and/or resources constraints? Congratulations on your medal anyway! ",
      "votes": 1
    },
    {
      "id": 942674,
      "postDate": "2020-07-23T23:34:52.513Z",
      "content": "<p>Hey, Nice work there! Do check and comment, would love to know your thoughts on it!</p>\n\n<p><a href=\"https://www.kaggle.com/lokeshrth4617/house-price-prediction-boosting-method?scriptVersionId=39394471\">Link</a></p>",
      "rawMarkdown": "Hey, Nice work there! Do check and comment, would love to know your thoughts on it!\n\n[Link](https://www.kaggle.com/lokeshrth4617/house-price-prediction-boosting-method?scriptVersionId=39394471)",
      "votes": -2
    },
    {
      "id": 942196,
      "postDate": "2020-07-23T16:23:57.717Z",
      "content": "<p>Thanks <a href=\"/tunguz\">@tunguz</a>   for your outstanding work and share with us</p>",
      "rawMarkdown": "Thanks @tunguz   for your outstanding work and share with us"
    },
    {
      "id": 939788,
      "postDate": "2020-07-22T13:13:40.673Z",
      "content": "<p>Thanks for sharing this! I am a beginner to computer vision, can you please suggest me sources to begin my journey?</p>",
      "rawMarkdown": "Thanks for sharing this! I am a beginner to computer vision, can you please suggest me sources to begin my journey?"
    },
    {
      "id": 943673,
      "postDate": "2020-07-24T14:09:24.853Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 943289,
      "author_name": "Jaseem C K",
      "author_url": "",
      "post_date": "2020-07-24T09:01:15.263000",
      "content": "<p>Nice work. Always great to hear about approaches from the masters themselves. Learned from such discussions more than anything. Keep Sharing.😊 </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 939305,
      "author_name": "yukiya",
      "author_url": "",
      "post_date": "2020-07-22T06:49:18.603000",
      "content": "<p>Thank you for the write ups, and congratulations on the result.</p>\n<blockquote>\n  <p>but then I re-trained all the weights - twice.</p>\n</blockquote>\n<p>Would you mind elaborating more on this ? Does it mean after first run (40 epoch ?), you restarted the training with different learning rate (for 50epoch?) and one more time restart adding jpeg compression ?<br>\nI have tried continuing the starter kernel for a long time, but it got trapped in a very small lr and never improve further.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 939794,
          "author_name": "Bojan Tunguz",
          "author_url": "",
          "post_date": "2020-07-22T13:15:59.293000",
          "content": "<p>Yes, that's basically correct, except that the first time I restarted training with the <strong>same</strong> learning rate as the first time.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 938874,
      "author_name": "RDizzl3",
      "author_url": "",
      "post_date": "2020-07-21T20:22:42.397000",
      "content": "<p><a href=\"https://www.kaggle.com/tunguz\" target=\"_blank\">@tunguz</a> - excellent result :) Just curious did you also play around with the stride on the efficient net models? Also, curious to know what batch size you used?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 938886,
          "author_name": "Bojan Tunguz",
          "author_url": "",
          "post_date": "2020-07-21T20:37:23.050000",
          "content": "<p>No, I did not find out about the stride trick until after the competition. </p>\n\n<p>Thanks for asking - the increased batch size certainly seemed to give me a boost. I was not able to go beyond batch size of 26 on a single GPU. Still figuring out how to effectively harness multi-GPU training, especially with PyTorch, which is still a foreign \"language\" to me. </p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 938841,
      "author_name": "Yassine Alouini",
      "author_url": "",
      "post_date": "2020-07-21T19:46:17.873000",
      "content": "<p><a href=\"https://www.kaggle.com/shonenkov\">https://www.kaggle.com/shonenkov</a> work is awesome indeed! One question: haven't you tried other EfficientNet backbones (B5 and more)? Maybe it wasn't possible due to time and/or resources constraints? Congratulations on your medal anyway! </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 942674,
      "author_name": "Lokesh Rathi",
      "author_url": "",
      "post_date": "2020-07-23T23:34:52.513000",
      "content": "<p>Hey, Nice work there! Do check and comment, would love to know your thoughts on it!</p>\n\n<p><a href=\"https://www.kaggle.com/lokeshrth4617/house-price-prediction-boosting-method?scriptVersionId=39394471\">Link</a></p>",
      "votes": -2,
      "replies": []
    },
    {
      "id": 942196,
      "author_name": "Mahmud Hasan",
      "author_url": "",
      "post_date": "2020-07-23T16:23:57.717000",
      "content": "<p>Thanks <a href=\"/tunguz\">@tunguz</a>   for your outstanding work and share with us</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 939788,
      "author_name": "srtpandata",
      "author_url": "",
      "post_date": "2020-07-22T13:13:40.673000",
      "content": "<p>Thanks for sharing this! I am a beginner to computer vision, can you please suggest me sources to begin my journey?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 943673,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-07-24T14:09:24.853000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "938834": "I would like to thank Kaggle for organizing such an interesting competition. Unfortunately, I had not paid much attention to it until just a couple of weeks ago, and am regretting not spending more time on it. Especially since models for the competition took **really** long time to train.\n\nLike many others, I decided to seriously join the competition after discovering the wonderful [Alex Shonenkov](https://www.kaggle.com/shonenkov) starter [kernel](https://www.kaggle.com/shonenkov/train-inference-gpu-baseline). Most of my work was building upon his work, and experimenting with different network backbones, training schedules, ensembles, and to a lesser degree augmentations.\n\nMy final ensemble consists of the following three EfficientNet model architectures:\n\n* B1 - I have only trained one of these networks successfully. The best result can be found in [this kernel](https://www.kaggle.com/tunguz/best-b1-inference/), and all the weights are in [this dataset](https://www.kaggle.com/tunguz/eb1-weights). The best single model achieves 0.918 locally, 0.925 on public LB and 0.911 on private LB.\n\n* B2 - I have trained several of these networks, and managed to get quite a bit of an improvement over Alex's original model. However, these networks have only made a minor contribution to my final blend. The best result can be found in [this kernel](https://www.kaggle.com/tunguz/best-b2-inference), and all the weights are in [this dataset](https://www.kaggle.com/tunguz/eb2-weights/). The best single model achieves 0.923 locally, 0.929 on public LB and 0.916 on private LB.\n\n* B4 - These networks have been the main workhorses behind my solution. Unfortunately, they are really slow to train, and I pretty much pushed all of my compute resources to get the best solution with them. Did not have any additional capacity (compute + time) to train bigger models, which I suspect would have performed even better. The best result can be found in [this kernel](https://www.kaggle.com/tunguz/best-b4-inference), and all the weights are in [this dataset](https://www.kaggle.com/tunguz/alaska2-eb4-model-weights). The best single model achieves 0.930 locally, 0.940 on public LB and 0.925 on private LB.\n\n**Training Schedule**\n\nThis is probably the biggest source of improvement for my networks over the base one that was found in the public kernel. I started with the same training schedule as in the original kernel, but then I **re-trained** all the weights - twice. The second time I decreased the LR to 0.0005, and increased patience to 2. I added JpegCompression to the augmentations, but otherwise did not change anything. I also increased training to 50 epochs, which took a really long time to go through, since each epoch would take about 2.5 hours on an V100 GPU.\n\n**Train/Val Split**\n\nFor the most of the models I stuck to the original 80/20 split, but have used different seeds. The choice of split seems to have been very important in this competition - the final models varied between 0.933 and 0.940 on LB. I have also trained a couple of models with the 95/5 split, but since they were B2 networks, they did not contribute significantly to the final ensemble. \n\n**Final Blend**\n\nSince I've used different splits for my training, it was not possible to do a consistent local validation for the ensemble, so I was forced to rely on LB. Again, since I was only able to work on this competition for less than two weeks, that did not leave me a lot of submissions on which to base my validation strategy, so I had to rely on my intuition a lot. For instance, I did not use an equally weighted average for the blend of the same model on different splits, but based the weights on the individual model's performance on LB. It is highly likely that the weights I had finally used are suboptimal, but I feel fortunate that my two submissions were the best ones on private LB. \n\n\n",
    "943289": "Nice work. Always great to hear about approaches from the masters themselves. Learned from such discussions more than anything. Keep Sharing.😊 ",
    "939305": "Thank you for the write ups, and congratulations on the result.\n\n&gt;  but then I re-trained all the weights - twice.\n\nWould you mind elaborating more on this ? Does it mean after first run (40 epoch ?), you restarted the training with different learning rate (for 50epoch?) and one more time restart adding jpeg compression ?\nI have tried continuing the starter kernel for a long time, but it got trapped in a very small lr and never improve further.\n",
    "938874": "@tunguz - excellent result :) Just curious did you also play around with the stride on the efficient net models? Also, curious to know what batch size you used?",
    "938841": "https://www.kaggle.com/shonenkov work is awesome indeed! One question: haven't you tried other EfficientNet backbones (B5 and more)? Maybe it wasn't possible due to time and/or resources constraints? Congratulations on your medal anyway! ",
    "942674": "Hey, Nice work there! Do check and comment, would love to know your thoughts on it!\n\n[Link](https://www.kaggle.com/lokeshrth4617/house-price-prediction-boosting-method?scriptVersionId=39394471)",
    "942196": "Thanks @tunguz   for your outstanding work and share with us",
    "939788": "Thanks for sharing this! I am a beginner to computer vision, can you please suggest me sources to begin my journey?",
    "943673": ""
  }
}