{
  "id": 71667,
  "title": "Few lessons learned (4th place)",
  "url": "/competitions/airbus-ship-detection/discussion/71667",
  "author_name": "Oleg Yaroshevskiy",
  "post_date": "2018-11-15T12:54:11.849000",
  "votes": 54,
  "comment_count": 8,
  "views": 0,
  "content": "<p>Thanks to all who has been a part of this competition - our second segmentation story in a row and finally gold for me and <a href=\"/ddanevskyi\">@ddanevskyi</a> , <a href=\"/vshmyhlo\">@vshmyhlo</a> . We hadn't stopped training our model till the last hour and that's a nice story about how we stuck around top-30 on public but kept challenging and improving our models till the end.</p>\n\n<p>Data science is more about understanding the task and therefore making proper validation. Keep calm watching public leaderboard - 12% of highly unbalanced dataset makes no real view. The story was not only about detecting ships but more to fight false positives (wave glare etc). We found having a good ship/no-ship classifier to be much more important for this competition than having a very good performing segmentation model. That's not hard part to obtain.</p>\n\n<p>About validation not much to say - folds based on this shinny <a href=\"https://www.kaggle.com/manuscrits/create-a-validation-dataset-correcting-the-leak/notebook#x\">kernel</a> (cheers <a href=\"/manuscrits\">@manuscrits</a> !). Clear,  no leaks, evaluated on ship images only with a small amount of no-ship images added (around 9k images totally). We got all kinds of U-Nets tuned from TGS Competition so we trained separately shallow/deep encoders to blend 4 of them as a final model. We made a really nice grid search over U-Net based architectures. </p>\n\n<p>Few words about lessons learned:</p>\n\n<ul>\n<li><p>Check your U-Net concatenations - if you want to predict small objects then try to avoid poolings/strides&gt;1 where they can be avoided. Also size matters - we found that missing central unet layer (&lt; x / 32) leaded to worst predictions of huge ships (up to 300px length).</p></li>\n<li><p>Deep networks not always lead to better results unless you tune and train them forever. Pick the best encoder to fit your particular resources, unet architecture, finetuning etc. Our best validation model was resnet34 (oh yes) trained 300 +epochs on cropped images. </p></li>\n<li><p>Finetuning on the full size images might boost you performance significantly if tuned properly (from 0.490 to 0.520 local-ships only validation as example).</p></li>\n<li><p>Part of success was an accurate intelligent 256 cropping.</p></li>\n<li><p>Some of losses lead to better shapes of boats but force ships to be more homogeneous which makes them hard to split. BCE-based will gave you worse shapes but more chances to split em.</p></li>\n</ul>\n\n<p>Also there are more to say about what additional channels might help you to fit instance segmentation, to split ships but we didn't achieve our best - unfortunately we didn't finalize our solutions for boundaries predictions (due to dice based loss probably). Looking forward to read about that from other teams.</p>\n\n<p>Thanks my talented team. Thanks the best community <a href=\"http://ods.ai\">ods.ai</a> for a wonderful sense of humor!</p>",
  "messages": [
    {
      "id": 421807,
      "postDate": "2018-11-15T12:54:11.850Z",
      "content": "<p>Thanks to all who has been a part of this competition - our second segmentation story in a row and finally gold for me and <a href=\"/ddanevskyi\">@ddanevskyi</a> , <a href=\"/vshmyhlo\">@vshmyhlo</a> . We hadn't stopped training our model till the last hour and that's a nice story about how we stuck around top-30 on public but kept challenging and improving our models till the end.</p>\n\n<p>Data science is more about understanding the task and therefore making proper validation. Keep calm watching public leaderboard - 12% of highly unbalanced dataset makes no real view. The story was not only about detecting ships but more to fight false positives (wave glare etc). We found having a good ship/no-ship classifier to be much more important for this competition than having a very good performing segmentation model. That's not hard part to obtain.</p>\n\n<p>About validation not much to say - folds based on this shinny <a href=\"https://www.kaggle.com/manuscrits/create-a-validation-dataset-correcting-the-leak/notebook#x\">kernel</a> (cheers <a href=\"/manuscrits\">@manuscrits</a> !). Clear,  no leaks, evaluated on ship images only with a small amount of no-ship images added (around 9k images totally). We got all kinds of U-Nets tuned from TGS Competition so we trained separately shallow/deep encoders to blend 4 of them as a final model. We made a really nice grid search over U-Net based architectures. </p>\n\n<p>Few words about lessons learned:</p>\n\n<ul>\n<li><p>Check your U-Net concatenations - if you want to predict small objects then try to avoid poolings/strides&gt;1 where they can be avoided. Also size matters - we found that missing central unet layer (&lt; x / 32) leaded to worst predictions of huge ships (up to 300px length).</p></li>\n<li><p>Deep networks not always lead to better results unless you tune and train them forever. Pick the best encoder to fit your particular resources, unet architecture, finetuning etc. Our best validation model was resnet34 (oh yes) trained 300 +epochs on cropped images. </p></li>\n<li><p>Finetuning on the full size images might boost you performance significantly if tuned properly (from 0.490 to 0.520 local-ships only validation as example).</p></li>\n<li><p>Part of success was an accurate intelligent 256 cropping.</p></li>\n<li><p>Some of losses lead to better shapes of boats but force ships to be more homogeneous which makes them hard to split. BCE-based will gave you worse shapes but more chances to split em.</p></li>\n</ul>\n\n<p>Also there are more to say about what additional channels might help you to fit instance segmentation, to split ships but we didn't achieve our best - unfortunately we didn't finalize our solutions for boundaries predictions (due to dice based loss probably). Looking forward to read about that from other teams.</p>\n\n<p>Thanks my talented team. Thanks the best community <a href=\"http://ods.ai\">ods.ai</a> for a wonderful sense of humor!</p>",
      "rawMarkdown": "Thanks to all who has been a part of this competition - our second segmentation story in a row and finally gold for me and @ddanevskyi , @vshmyhlo . We hadn't stopped training our model till the last hour and that's a nice story about how we stuck around top-30 on public but kept challenging and improving our models till the end.\n\nData science is more about understanding the task and therefore making proper validation. Keep calm watching public leaderboard - 12% of highly unbalanced dataset makes no real view. The story was not only about detecting ships but more to fight false positives (wave glare etc). We found having a good ship/no-ship classifier to be much more important for this competition than having a very good performing segmentation model. That's not hard part to obtain.\n\nAbout validation not much to say - folds based on this shinny [kernel][1] (cheers @manuscrits !). Clear,  no leaks, evaluated on ship images only with a small amount of no-ship images added (around 9k images totally). We got all kinds of U-Nets tuned from TGS Competition so we trained separately shallow/deep encoders to blend 4 of them as a final model. We made a really nice grid search over U-Net based architectures. \n\nFew words about lessons learned:\n\n - Check your U-Net concatenations - if you want to predict small objects then try to avoid poolings/strides&gt;1 where they can be avoided. Also size matters - we found that missing central unet layer (&lt; x / 32) leaded to worst predictions of huge ships (up to 300px length).\n\n - Deep networks not always lead to better results unless you tune and train them forever. Pick the best encoder to fit your particular resources, unet architecture, finetuning etc. Our best validation model was resnet34 (oh yes) trained 300 +epochs on cropped images. \n\n - Finetuning on the full size images might boost you performance significantly if tuned properly (from 0.490 to 0.520 local-ships only validation as example).\n\n - Part of success was an accurate intelligent 256 cropping.\n\n - Some of losses lead to better shapes of boats but force ships to be more homogeneous which makes them hard to split. BCE-based will gave you worse shapes but more chances to split em.\n\nAlso there are more to say about what additional channels might help you to fit instance segmentation, to split ships but we didn't achieve our best - unfortunately we didn't finalize our solutions for boundaries predictions (due to dice based loss probably). Looking forward to read about that from other teams.\n\nThanks my talented team. Thanks the best community [ods.ai][2] for a wonderful sense of humor!\n\n\n  [1]: https://www.kaggle.com/manuscrits/create-a-validation-dataset-correcting-the-leak/notebook#x\n  [2]: http://ods.ai",
      "votes": 54
    },
    {
      "id": 421825,
      "postDate": "2018-11-15T13:14:28.957Z",
      "content": "<p>I also want to make emphasis on how important smart random crop turned out to be. The main idea is to randomly crop the image in such a way that at least some part of the mask in present in the resulted crop, this little modification greatly boosted local score and with ship/no-ship classifier was second major improvement which took us to the 4th place.</p>",
      "rawMarkdown": "I also want to make emphasis on how important smart random crop turned out to be. The main idea is to randomly crop the image in such a way that at least some part of the mask in present in the resulted crop, this little modification greatly boosted local score and with ship/no-ship classifier was second major improvement which took us to the 4th place.",
      "votes": 12,
      "replies": [
        {
          "id": 558876,
          "postDate": "2019-06-23T05:03:50.160Z",
          "content": "<p>Can you explain the cropping method in detail? How do you ensure that the cropped image contains at least part of the mask? What do you do when the model predicts? Thank you very much!</p>",
          "rawMarkdown": "Can you explain the cropping method in detail? How do you ensure that the cropped image contains at least part of the mask? What do you do when the model predicts? Thank you very much!",
          "votes": -1
        }
      ]
    },
    {
      "id": 791931,
      "postDate": "2020-03-30T19:24:15.313Z",
      "content": "<p>Such a brilliant solution! It is always interesting to revisit old competitions to check if my skills have improved and if I can come up with such a solution. Should check again in few months. ;)</p>",
      "rawMarkdown": "Such a brilliant solution! It is always interesting to revisit old competitions to check if my skills have improved and if I can come up with such a solution. Should check again in few months. ;)",
      "votes": 1
    },
    {
      "id": 427690,
      "postDate": "2018-11-26T00:37:16.380Z",
      "content": "<p>Congrats <a href=\"/yaroshevskiy\">@yaroshevskiy</a>, and thanks for sharing.</p>",
      "rawMarkdown": "Congrats @yaroshevskiy, and thanks for sharing."
    },
    {
      "id": 422023,
      "postDate": "2018-11-15T17:24:20.517Z",
      "content": "<p>Congratulations!</p>",
      "rawMarkdown": "Congratulations!"
    },
    {
      "id": 421976,
      "postDate": "2018-11-15T16:17:32.077Z",
      "content": "<p>Congratulation! \nGood to hear that my kernel was used by some of the top competitors :)</p>",
      "rawMarkdown": "Congratulation! \nGood to hear that my kernel was used by some of the top competitors :)"
    },
    {
      "id": 421965,
      "postDate": "2018-11-15T16:00:24.360Z",
      "content": "<p>Congratulations! Really deserved :)</p>\n\n<p>Are you guys going to open source the code?</p>",
      "rawMarkdown": "Congratulations! Really deserved :)\n\nAre you guys going to open source the code?"
    },
    {
      "id": 423079,
      "postDate": "2018-11-17T12:23:22.310Z",
      "rawMarkdown": "",
      "votes": 7,
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 421825,
      "author_name": "Vlad Shmyhlo ",
      "author_url": "",
      "post_date": "2018-11-15T13:14:28.957000",
      "content": "<p>I also want to make emphasis on how important smart random crop turned out to be. The main idea is to randomly crop the image in such a way that at least some part of the mask in present in the resulted crop, this little modification greatly boosted local score and with ship/no-ship classifier was second major improvement which took us to the 4th place.</p>",
      "votes": 12,
      "replies": [
        {
          "id": 558876,
          "author_name": "Jone",
          "author_url": "",
          "post_date": "2019-06-23T05:03:50.160000",
          "content": "<p>Can you explain the cropping method in detail? How do you ensure that the cropped image contains at least part of the mask? What do you do when the model predicts? Thank you very much!</p>",
          "votes": -1,
          "replies": []
        }
      ]
    },
    {
      "id": 791931,
      "author_name": "Yassine Alouini",
      "author_url": "",
      "post_date": "2020-03-30T19:24:15.313000",
      "content": "<p>Such a brilliant solution! It is always interesting to revisit old competitions to check if my skills have improved and if I can come up with such a solution. Should check again in few months. ;)</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 427690,
      "author_name": "YaGana Sheriff-Hussaini",
      "author_url": "",
      "post_date": "2018-11-26T00:37:16.380000",
      "content": "<p>Congrats <a href=\"/yaroshevskiy\">@yaroshevskiy</a>, and thanks for sharing.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 422023,
      "author_name": "Aleksandr Zolotarev",
      "author_url": "",
      "post_date": "2018-11-15T17:24:20.517000",
      "content": "<p>Congratulations!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 421976,
      "author_name": "Maxime Riché",
      "author_url": "",
      "post_date": "2018-11-15T16:17:32.077000",
      "content": "<p>Congratulation! \nGood to hear that my kernel was used by some of the top competitors :)</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 421965,
      "author_name": "Eduardo Rocha de Andrade",
      "author_url": "",
      "post_date": "2018-11-15T16:00:24.360000",
      "content": "<p>Congratulations! Really deserved :)</p>\n\n<p>Are you guys going to open source the code?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 423079,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-11-17T12:23:22.310000",
      "content": "",
      "votes": 7,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "421807": "Thanks to all who has been a part of this competition - our second segmentation story in a row and finally gold for me and @ddanevskyi , @vshmyhlo . We hadn't stopped training our model till the last hour and that's a nice story about how we stuck around top-30 on public but kept challenging and improving our models till the end.\n\nData science is more about understanding the task and therefore making proper validation. Keep calm watching public leaderboard - 12% of highly unbalanced dataset makes no real view. The story was not only about detecting ships but more to fight false positives (wave glare etc). We found having a good ship/no-ship classifier to be much more important for this competition than having a very good performing segmentation model. That's not hard part to obtain.\n\nAbout validation not much to say - folds based on this shinny [kernel][1] (cheers @manuscrits !). Clear,  no leaks, evaluated on ship images only with a small amount of no-ship images added (around 9k images totally). We got all kinds of U-Nets tuned from TGS Competition so we trained separately shallow/deep encoders to blend 4 of them as a final model. We made a really nice grid search over U-Net based architectures. \n\nFew words about lessons learned:\n\n - Check your U-Net concatenations - if you want to predict small objects then try to avoid poolings/strides&gt;1 where they can be avoided. Also size matters - we found that missing central unet layer (&lt; x / 32) leaded to worst predictions of huge ships (up to 300px length).\n\n - Deep networks not always lead to better results unless you tune and train them forever. Pick the best encoder to fit your particular resources, unet architecture, finetuning etc. Our best validation model was resnet34 (oh yes) trained 300 +epochs on cropped images. \n\n - Finetuning on the full size images might boost you performance significantly if tuned properly (from 0.490 to 0.520 local-ships only validation as example).\n\n - Part of success was an accurate intelligent 256 cropping.\n\n - Some of losses lead to better shapes of boats but force ships to be more homogeneous which makes them hard to split. BCE-based will gave you worse shapes but more chances to split em.\n\nAlso there are more to say about what additional channels might help you to fit instance segmentation, to split ships but we didn't achieve our best - unfortunately we didn't finalize our solutions for boundaries predictions (due to dice based loss probably). Looking forward to read about that from other teams.\n\nThanks my talented team. Thanks the best community [ods.ai][2] for a wonderful sense of humor!\n\n\n  [1]: https://www.kaggle.com/manuscrits/create-a-validation-dataset-correcting-the-leak/notebook#x\n  [2]: http://ods.ai",
    "421825": "I also want to make emphasis on how important smart random crop turned out to be. The main idea is to randomly crop the image in such a way that at least some part of the mask in present in the resulted crop, this little modification greatly boosted local score and with ship/no-ship classifier was second major improvement which took us to the 4th place.",
    "791931": "Such a brilliant solution! It is always interesting to revisit old competitions to check if my skills have improved and if I can come up with such a solution. Should check again in few months. ;)",
    "427690": "Congrats @yaroshevskiy, and thanks for sharing.",
    "422023": "Congratulations!",
    "421976": "Congratulation! \nGood to hear that my kernel was used by some of the top competitors :)",
    "421965": "Congratulations! Really deserved :)\n\nAre you guys going to open source the code?",
    "423079": ""
  }
}