{
  "id": 64730,
  "title": "prediction is very good, but submission sore is very low",
  "url": "/competitions/airbus-ship-detection/discussion/64730",
  "author_name": "",
  "post_date": "2018-09-01T03:06:46.290965300Z",
  "votes": 6,
  "comment_count": 11,
  "views": 0,
  "content": "",
  "messages": [
    {
      "id": "379837",
      "postDate": "09/01/2018 03:06:46",
      "content": "",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "379838",
      "postDate": "09/01/2018 03:07:53",
      "content": "<p>final sore is 0.825.But I think the model predict very well</p>",
      "rawMarkdown": "final sore is 0.825.But I think the model predict very well",
      "votes": null
    },
    {
      "id": "379874",
      "postDate": "09/01/2018 05:49:53",
      "content": "<p>It is because you identify the ship hull and not the ships bounding box. The score is calculated according to the given bounding box mask and not the ships hull.</p>\n\n<p>See also my question:</p>\n\n<p>\"How do I draw the ship slanted bounding box?\"</p>\n\n<p><a href=\"https://www.kaggle.com/c/airbus-ship-detection/discussion/64648\">https://www.kaggle.com/c/airbus-ship-detection/discussion/64648</a></p>",
      "rawMarkdown": "It is because you identify the ship hull and not the ships bounding box. The score is calculated according to the given bounding box mask and not the ships hull.\n\nSee also my question:\n\n\"How do I draw the ship slanted bounding box?\"\n\nhttps://www.kaggle.com/c/airbus-ship-detection/discussion/64648",
      "votes": null
    },
    {
      "id": "380086",
      "postDate": "09/01/2018 16:38:41",
      "content": "<p>There are several things. First, the mask for each ship must be separated. So if you have 5 ships in an image, in your submission for the image you must have 5 lines, one for each predicted ship. Second, your model struggles from prediction ships for an empty image. If you predict something for an empty image, you automatically get zero score for it. If you submitted a solution with no ships at all, you would get ~0.85 score. You may consider stacking your model with one that predicts ship/nothing (like <a href=\"https://www.kaggle.com/iafoss/fine-tuning-resnet34-on-ship-detection\">https://www.kaggle.com/iafoss/fine-tuning-resnet34-on-ship-detection</a> that gives ~98% accuracy). Finally, you should keep in mind that the data was labeled by machine (likely SSD with rotating bounding boxes), and the test labeling is not ideal and has boxes rather than masks. Look at the following images. The first one is a predicted mask of the image that gives 0.82 dice. The second one is the original image with predicted (red) and \"ground truth\" (green) bounding boxes that gives 0.84 dice. You see that a model may perform better than one used for labeling the dataset, but the score is low since is calculated based on the machine labeling. So to get high score instead of building a model that is as accurate as possible one should create a model that mimics one used for labeling. </p>\n\n<p>BTW I'm curious, what the dice of your final image segmentation model?</p>\n\n<p><img src=\"https://image.ibb.co/cgLofK/ship_detection_mask.png\" alt=\"enter image description here\"> <img src=\"https://image.ibb.co/gJtKnz/ship_detection_box.png\" alt=\"enter image description here\"></p>",
      "rawMarkdown": "There are several things. First, the mask for each ship must be separated. So if you have 5 ships in an image, in your submission for the image you must have 5 lines, one for each predicted ship. Second, your model struggles from prediction ships for an empty image. If you predict something for an empty image, you automatically get zero score for it. If you submitted a solution with no ships at all, you would get ~0.85 score. You may consider stacking your model with one that predicts ship/nothing (like https://www.kaggle.com/iafoss/fine-tuning-resnet34-on-ship-detection that gives ~98% accuracy). Finally, you should keep in mind that the data was labeled by machine (likely SSD with rotating bounding boxes), and the test labeling is not ideal and has boxes rather than masks. Look at the following images. The first one is a predicted mask of the image that gives 0.82 dice. The second one is the original image with predicted (red) and \"ground truth\" (green) bounding boxes that gives 0.84 dice. You see that a model may perform better than one used for labeling the dataset, but the score is low since is calculated based on the machine labeling. So to get high score instead of building a model that is as accurate as possible one should create a model that mimics one used for labeling. \n\nBTW I'm curious, what the dice of your final image segmentation model?\n\n![enter image description here][1] ![enter image description here][2]\n\n\n  [1]: https://image.ibb.co/cgLofK/ship_detection_mask.png\n  [2]: https://image.ibb.co/gJtKnz/ship_detection_box.png",
      "votes": null
    },
    {
      "id": "390745",
      "postDate": "09/20/2018 18:58:14",
      "content": "<p>Very helpful discussion.</p>",
      "rawMarkdown": "Very helpful discussion.",
      "votes": null
    },
    {
      "id": "398642",
      "postDate": "10/04/2018 11:41:41",
      "content": "<p>At least part of the problem is the poor segmentation quality of the test images.\nFirst column is my prediction (and it's score), 2nd column is the same prediction but approximated to a rectangle, 3rd column is the ground truth and 4th column is the full image</p>\n\n<p><img src=\"https://i.imgur.com/dSE4lbm.png\" alt=\"Sample 1\">\n<img src=\"https://i.imgur.com/QSbBLJW.png\" alt=\"Sample 2\">\n<img src=\"https://i.imgur.com/28XED72.png\" alt=\"Sample 3\">\n<img src=\"https://i.imgur.com/Ic9g9RX.png\" alt=\"Sample 4\">\n<img src=\"https://i.imgur.com/WyETBls.png\" alt=\"Sample 5\">\n<img src=\"https://i.imgur.com/pLRp7v7.png\" alt=\"Sample 6\">\n<img src=\"https://i.imgur.com/P1blX3S.png\" alt=\"Sample 7\">\n<img src=\"https://i.imgur.com/sg7OkRc.png\" alt=\"Sample 8\"></p>\n\n<p>I am not claiming that my predictions are perfect, but I think that they are in this (and many more) cases as good as the ground truth (which has no consistency of what it includes), but scored very badly.</p>\n\n<p>Note: All of my samples are from single ship images to make the metric only depending on this one ship.</p>",
      "rawMarkdown": "At least part of the problem is the poor segmentation quality of the test images.\nFirst column is my prediction (and it's score), 2nd column is the same prediction but approximated to a rectangle, 3rd column is the ground truth and 4th column is the full image\n\n![Sample 1][1]\n![Sample 2][2]\n![Sample 3][3]\n![Sample 4][4]\n![Sample 5][5]\n![Sample 6][6]\n![Sample 7][7]\n![Sample 8][8]\n\nI am not claiming that my predictions are perfect, but I think that they are in this (and many more) cases as good as the ground truth (which has no consistency of what it includes), but scored very badly.\n\nNote: All of my samples are from single ship images to make the metric only depending on this one ship.\n\n  [1]: https://i.imgur.com/dSE4lbm.png\n  [2]: https://i.imgur.com/QSbBLJW.png\n  [3]: https://i.imgur.com/28XED72.png\n  [4]: https://i.imgur.com/Ic9g9RX.png\n  [5]: https://i.imgur.com/WyETBls.png\n  [6]: https://i.imgur.com/pLRp7v7.png\n  [7]: https://i.imgur.com/P1blX3S.png\n  [8]: https://i.imgur.com/sg7OkRc.png",
      "votes": null
    },
    {
      "id": "398651",
      "postDate": "10/04/2018 12:10:16",
      "content": "<p>Your predictions seem great. The data leakage make your rank lower (about 200 teams submit 1.000 kernel), but you may be winning potentially. </p>",
      "rawMarkdown": "Your predictions seem great. The data leakage make your rank lower (about 200 teams submit 1.000 kernel), but you may be winning potentially.",
      "votes": null
    },
    {
      "id": "399696",
      "postDate": "10/06/2018 14:10:10",
      "content": "<p>I don't think.\nThe score of an empty submission (no ships at all detected is 0.847). The 2nd column gets a score of only 0.850.</p>\n\n<p>Analysis shows, that I correctly identify 97% of the empty images and  only miss 0.2% of the images with ships. But I  get a score of 0.24 on the images with ships.</p>\n\n<p>Further analysis show, that I even find most of the ships in these images (IoU &gt;0) , but get an extremely low score, because the IoU (shown as second number after the score) is lower than 0.5 in many cases (so the score for that ship becomes 0). </p>\n\n<p>And the reason for the low IoU seems not to be the prediction but the bad ground truth.</p>\n\n<p>Hera are some samples</p>\n\n<p><img src=\"https://i.imgur.com/syWu6vV.png\" alt=\"Sample 1\">\n<img src=\"https://i.imgur.com/gvNHBA5.png\" alt=\"Sample 2\">\n<img src=\"https://i.imgur.com/JvZzTd1.png\" alt=\"Sample 3\">\n<img src=\"https://i.imgur.com/RDNz1Fx.png\" alt=\"Sample 4\">\n<img src=\"https://i.imgur.com/gIPBJ5l.png\" alt=\"Sample 5\">\n<img src=\"https://i.imgur.com/Tvgw5jF.png\" alt=\"Sample 6\">\n<img src=\"https://i.imgur.com/woOJtLJ.png\" alt=\"Sample 7\">\n<img src=\"https://i.imgur.com/hyP4gJw.png\" alt=\"Sample 8\"></p>\n\n<p><strong>Note:</strong> This are not the worst or only samples with wrong ground truth. This are average samples. I just selected some where my prediction are clearly better than the ground truth. I can find many many more.</p>",
      "rawMarkdown": "I don't think.\nThe score of an empty submission (no ships at all detected is 0.847). The 2nd column gets a score of only 0.850.\n\nAnalysis shows, that I correctly identify 97% of the empty images and  only miss 0.2% of the images with ships. But I  get a score of 0.24 on the images with ships.\n\nFurther analysis show, that I even find most of the ships in these images (IoU &gt;0) , but get an extremely low score, because the IoU (shown as second number after the score) is lower than 0.5 in many cases (so the score for that ship becomes 0). \n\nAnd the reason for the low IoU seems not to be the prediction but the bad ground truth.\n\nHera are some samples\n\n![Sample 1][1]\n![Sample 2][2]\n![Sample 3][3]\n![Sample 4][4]\n![Sample 5][5]\n![Sample 6][6]\n![Sample 7][7]\n![Sample 8][8]\n\n**Note:** This are not the worst or only samples with wrong ground truth. This are average samples. I just selected some where my prediction are clearly better than the ground truth. I can find many many more.\n\n\n  [1]: https://i.imgur.com/syWu6vV.png\n  [2]: https://i.imgur.com/gvNHBA5.png\n  [3]: https://i.imgur.com/JvZzTd1.png\n  [4]: https://i.imgur.com/RDNz1Fx.png\n  [5]: https://i.imgur.com/gIPBJ5l.png\n  [6]: https://i.imgur.com/Tvgw5jF.png\n  [7]: https://i.imgur.com/woOJtLJ.png\n  [8]: https://i.imgur.com/hyP4gJw.png",
      "votes": null
    },
    {
      "id": "399745",
      "postDate": "10/06/2018 16:31:29",
      "content": "<p>Very nice indeed. Only missing 0.2 % of the ships is extremely impressive. Especially given the very low number of pixels representing some of the boats.</p>\n\n<p>97% identification of empty picture is fine. On images with several ships the scoring punishes FP 4 times <strong>less</strong> than FN. But the same is not true on an image to image level. Here FP are punished just as hard as FN.</p>\n\n<p>The minimum 0.5 IoU limit has a significant influence where to ships are close together. If one draws only one box around two equal sized ships the score will be zero for both ships. If it is a small and a larger ship only the larger ship will score. I think that this could be one of the reasons for the relative high 0.5 limit.   </p>",
      "rawMarkdown": "Very nice indeed. Only missing 0.2 % of the ships is extremely impressive. Especially given the very low number of pixels representing some of the boats.\n\n97% identification of empty picture is fine. On images with several ships the scoring punishes FP 4 times **less** than FN. But the same is not true on an image to image level. Here FP are punished just as hard as FN.\n\nThe minimum 0.5 IoU limit has a significant influence where to ships are close together. If one draws only one box around two equal sized ships the score will be zero for both ships. If it is a small and a larger ship only the larger ship will score. I think that this could be one of the reasons for the relative high 0.5 limit.",
      "votes": null
    },
    {
      "id": "400037",
      "postDate": "10/07/2018 12:37:36",
      "content": "<p>I did not say, that I get 99.8% of the ships. I said that I miss 0.2% of the images with ships.</p>\n\n<p>The scoring metric suggests that the challenge is to separate ships (like it was (from what I have read) important in the Data Science Bowl 2018 to separate overlapping nuclei). But it is not: 84.78% of the images contain no ship, 9.60% contain just 1 ship. So not separating ships at all (but doing everything else correct) would score at least 0.9445.</p>\n\n<p>And even separating a small number of ships,  which are far apart in most cases, is easy in post processing with traditional CV algorithms.</p>\n\n<p>So the score is mostly (at least 94.45%) depending on the ship / background decision. And this part of the score is determined by the  ground truth quality. We can replicate the original mechanism by predicting a single rectangle for each ship (traditional CV). But the problem of wrong ground truth (missing parts of the ship or non minimal rectangle size) remains. This can lead (as my samples have shown) to a 0 score for a correctly identified ship.</p>\n\n<p>So maybe there is more usable information in the in the ground truth itself (maybe they only used certain rectangle sizes, angels or had a raster bigger than 1 px) </p>",
      "rawMarkdown": "I did not say, that I get 99.8% of the ships. I said that I miss 0.2% of the images with ships.\n\nThe scoring metric suggests that the challenge is to separate ships (like it was (from what I have read) important in the Data Science Bowl 2018 to separate overlapping nuclei). But it is not: 84.78% of the images contain no ship, 9.60% contain just 1 ship. So not separating ships at all (but doing everything else correct) would score at least 0.9445.\n\nAnd even separating a small number of ships,  which are far apart in most cases, is easy in post processing with traditional CV algorithms.\n\nSo the score is mostly (at least 94.45%) depending on the ship / background decision. And this part of the score is determined by the  ground truth quality. We can replicate the original mechanism by predicting a single rectangle for each ship (traditional CV). But the problem of wrong ground truth (missing parts of the ship or non minimal rectangle size) remains. This can lead (as my samples have shown) to a 0 score for a correctly identified ship.\n\nSo maybe there is more usable information in the in the ground truth itself (maybe they only used certain rectangle sizes, angels or had a raster bigger than 1 px)",
      "votes": null
    },
    {
      "id": "400213",
      "postDate": "10/07/2018 22:20:54",
      "content": "<p>If you want to do an analysis of the rectangle sizes and angles, you can use a dataset that I have created <a href=\"https://www.kaggle.com/iafoss/rotating-bounding-boxes-for-ship-localization\">https://www.kaggle.com/iafoss/rotating-bounding-boxes-for-ship-localization</a> and a kernel that demonstrates it <a href=\"https://www.kaggle.com/iafoss/rotating-bounding-boxes-ship-localization\">https://www.kaggle.com/iafoss/rotating-bounding-boxes-ship-localization</a> . However, I couldn't boost my score much when I tried to train a second level model that accounts for this information.</p>\n\n<p>Regarding overlapping ships, based on my testing, they do not contribute much to the error. That I saw the main issue are displaced bounding boxes for small and intermediate ships, which essentially you also were referring to.</p>",
      "rawMarkdown": "If you want to do an analysis of the rectangle sizes and angles, you can use a dataset that I have created https://www.kaggle.com/iafoss/rotating-bounding-boxes-for-ship-localization and a kernel that demonstrates it https://www.kaggle.com/iafoss/rotating-bounding-boxes-ship-localization . However, I couldn't boost my score much when I tried to train a second level model that accounts for this information.\n\nRegarding overlapping ships, based on my testing, they do not contribute much to the error. That I saw the main issue are displaced bounding boxes for small and intermediate ships, which essentially you also were referring to.",
      "votes": null
    },
    {
      "id": "412649",
      "postDate": "10/30/2018 14:56:56",
      "content": "<p>Could I please ask you how you went from heatmaps to rectangles? </p>",
      "rawMarkdown": "Could I please ask you how you went from heatmaps to rectangles?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 379838,
      "author_name": "liuyicheng5418",
      "author_url": "",
      "post_date": "09/01/2018 03:07:53",
      "content": "<p>final sore is 0.825.But I think the model predict very well</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 379874,
      "author_name": "petersorensen360",
      "author_url": "",
      "post_date": "09/01/2018 05:49:53",
      "content": "<p>It is because you identify the ship hull and not the ships bounding box. The score is calculated according to the given bounding box mask and not the ships hull.</p>\n\n<p>See also my question:</p>\n\n<p>\"How do I draw the ship slanted bounding box?\"</p>\n\n<p><a href=\"https://www.kaggle.com/c/airbus-ship-detection/discussion/64648\">https://www.kaggle.com/c/airbus-ship-detection/discussion/64648</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 380086,
      "author_name": "iafoss",
      "author_url": "",
      "post_date": "09/01/2018 16:38:41",
      "content": "<p>There are several things. First, the mask for each ship must be separated. So if you have 5 ships in an image, in your submission for the image you must have 5 lines, one for each predicted ship. Second, your model struggles from prediction ships for an empty image. If you predict something for an empty image, you automatically get zero score for it. If you submitted a solution with no ships at all, you would get ~0.85 score. You may consider stacking your model with one that predicts ship/nothing (like <a href=\"https://www.kaggle.com/iafoss/fine-tuning-resnet34-on-ship-detection\">https://www.kaggle.com/iafoss/fine-tuning-resnet34-on-ship-detection</a> that gives ~98% accuracy). Finally, you should keep in mind that the data was labeled by machine (likely SSD with rotating bounding boxes), and the test labeling is not ideal and has boxes rather than masks. Look at the following images. The first one is a predicted mask of the image that gives 0.82 dice. The second one is the original image with predicted (red) and \"ground truth\" (green) bounding boxes that gives 0.84 dice. You see that a model may perform better than one used for labeling the dataset, but the score is low since is calculated based on the machine labeling. So to get high score instead of building a model that is as accurate as possible one should create a model that mimics one used for labeling. </p>\n\n<p>BTW I'm curious, what the dice of your final image segmentation model?</p>\n\n<p><img src=\"https://image.ibb.co/cgLofK/ship_detection_mask.png\" alt=\"enter image description here\"> <img src=\"https://image.ibb.co/gJtKnz/ship_detection_box.png\" alt=\"enter image description here\"></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 390745,
      "author_name": "kamranash",
      "author_url": "",
      "post_date": "09/20/2018 18:58:14",
      "content": "<p>Very helpful discussion.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 398642,
      "author_name": "rue1401",
      "author_url": "",
      "post_date": "10/04/2018 11:41:41",
      "content": "<p>At least part of the problem is the poor segmentation quality of the test images.\nFirst column is my prediction (and it's score), 2nd column is the same prediction but approximated to a rectangle, 3rd column is the ground truth and 4th column is the full image</p>\n\n<p><img src=\"https://i.imgur.com/dSE4lbm.png\" alt=\"Sample 1\">\n<img src=\"https://i.imgur.com/QSbBLJW.png\" alt=\"Sample 2\">\n<img src=\"https://i.imgur.com/28XED72.png\" alt=\"Sample 3\">\n<img src=\"https://i.imgur.com/Ic9g9RX.png\" alt=\"Sample 4\">\n<img src=\"https://i.imgur.com/WyETBls.png\" alt=\"Sample 5\">\n<img src=\"https://i.imgur.com/pLRp7v7.png\" alt=\"Sample 6\">\n<img src=\"https://i.imgur.com/P1blX3S.png\" alt=\"Sample 7\">\n<img src=\"https://i.imgur.com/sg7OkRc.png\" alt=\"Sample 8\"></p>\n\n<p>I am not claiming that my predictions are perfect, but I think that they are in this (and many more) cases as good as the ground truth (which has no consistency of what it includes), but scored very badly.</p>\n\n<p>Note: All of my samples are from single ship images to make the metric only depending on this one ship.</p>",
      "votes": null,
      "replies": [
        {
          "id": 398651,
          "author_name": "toshik",
          "author_url": "",
          "post_date": "10/04/2018 12:10:16",
          "content": "<p>Your predictions seem great. The data leakage make your rank lower (about 200 teams submit 1.000 kernel), but you may be winning potentially. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 399696,
          "author_name": "rue1401",
          "author_url": "",
          "post_date": "10/06/2018 14:10:10",
          "content": "<p>I don't think.\nThe score of an empty submission (no ships at all detected is 0.847). The 2nd column gets a score of only 0.850.</p>\n\n<p>Analysis shows, that I correctly identify 97% of the empty images and  only miss 0.2% of the images with ships. But I  get a score of 0.24 on the images with ships.</p>\n\n<p>Further analysis show, that I even find most of the ships in these images (IoU &gt;0) , but get an extremely low score, because the IoU (shown as second number after the score) is lower than 0.5 in many cases (so the score for that ship becomes 0). </p>\n\n<p>And the reason for the low IoU seems not to be the prediction but the bad ground truth.</p>\n\n<p>Hera are some samples</p>\n\n<p><img src=\"https://i.imgur.com/syWu6vV.png\" alt=\"Sample 1\">\n<img src=\"https://i.imgur.com/gvNHBA5.png\" alt=\"Sample 2\">\n<img src=\"https://i.imgur.com/JvZzTd1.png\" alt=\"Sample 3\">\n<img src=\"https://i.imgur.com/RDNz1Fx.png\" alt=\"Sample 4\">\n<img src=\"https://i.imgur.com/gIPBJ5l.png\" alt=\"Sample 5\">\n<img src=\"https://i.imgur.com/Tvgw5jF.png\" alt=\"Sample 6\">\n<img src=\"https://i.imgur.com/woOJtLJ.png\" alt=\"Sample 7\">\n<img src=\"https://i.imgur.com/hyP4gJw.png\" alt=\"Sample 8\"></p>\n\n<p><strong>Note:</strong> This are not the worst or only samples with wrong ground truth. This are average samples. I just selected some where my prediction are clearly better than the ground truth. I can find many many more.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 399745,
          "author_name": "petersorensen360",
          "author_url": "",
          "post_date": "10/06/2018 16:31:29",
          "content": "<p>Very nice indeed. Only missing 0.2 % of the ships is extremely impressive. Especially given the very low number of pixels representing some of the boats.</p>\n\n<p>97% identification of empty picture is fine. On images with several ships the scoring punishes FP 4 times <strong>less</strong> than FN. But the same is not true on an image to image level. Here FP are punished just as hard as FN.</p>\n\n<p>The minimum 0.5 IoU limit has a significant influence where to ships are close together. If one draws only one box around two equal sized ships the score will be zero for both ships. If it is a small and a larger ship only the larger ship will score. I think that this could be one of the reasons for the relative high 0.5 limit.   </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 400037,
          "author_name": "rue1401",
          "author_url": "",
          "post_date": "10/07/2018 12:37:36",
          "content": "<p>I did not say, that I get 99.8% of the ships. I said that I miss 0.2% of the images with ships.</p>\n\n<p>The scoring metric suggests that the challenge is to separate ships (like it was (from what I have read) important in the Data Science Bowl 2018 to separate overlapping nuclei). But it is not: 84.78% of the images contain no ship, 9.60% contain just 1 ship. So not separating ships at all (but doing everything else correct) would score at least 0.9445.</p>\n\n<p>And even separating a small number of ships,  which are far apart in most cases, is easy in post processing with traditional CV algorithms.</p>\n\n<p>So the score is mostly (at least 94.45%) depending on the ship / background decision. And this part of the score is determined by the  ground truth quality. We can replicate the original mechanism by predicting a single rectangle for each ship (traditional CV). But the problem of wrong ground truth (missing parts of the ship or non minimal rectangle size) remains. This can lead (as my samples have shown) to a 0 score for a correctly identified ship.</p>\n\n<p>So maybe there is more usable information in the in the ground truth itself (maybe they only used certain rectangle sizes, angels or had a raster bigger than 1 px) </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 400213,
          "author_name": "iafoss",
          "author_url": "",
          "post_date": "10/07/2018 22:20:54",
          "content": "<p>If you want to do an analysis of the rectangle sizes and angles, you can use a dataset that I have created <a href=\"https://www.kaggle.com/iafoss/rotating-bounding-boxes-for-ship-localization\">https://www.kaggle.com/iafoss/rotating-bounding-boxes-for-ship-localization</a> and a kernel that demonstrates it <a href=\"https://www.kaggle.com/iafoss/rotating-bounding-boxes-ship-localization\">https://www.kaggle.com/iafoss/rotating-bounding-boxes-ship-localization</a> . However, I couldn't boost my score much when I tried to train a second level model that accounts for this information.</p>\n\n<p>Regarding overlapping ships, based on my testing, they do not contribute much to the error. That I saw the main issue are displaced bounding boxes for small and intermediate ships, which essentially you also were referring to.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 412649,
          "author_name": "radek1",
          "author_url": "",
          "post_date": "10/30/2018 14:56:56",
          "content": "<p>Could I please ask you how you went from heatmaps to rectangles? </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "379837": "",
    "379838": "final sore is 0.825.But I think the model predict very well",
    "379874": "It is because you identify the ship hull and not the ships bounding box. The score is calculated according to the given bounding box mask and not the ships hull.\n\nSee also my question:\n\n\"How do I draw the ship slanted bounding box?\"\n\nhttps://www.kaggle.com/c/airbus-ship-detection/discussion/64648",
    "380086": "There are several things. First, the mask for each ship must be separated. So if you have 5 ships in an image, in your submission for the image you must have 5 lines, one for each predicted ship. Second, your model struggles from prediction ships for an empty image. If you predict something for an empty image, you automatically get zero score for it. If you submitted a solution with no ships at all, you would get ~0.85 score. You may consider stacking your model with one that predicts ship/nothing (like https://www.kaggle.com/iafoss/fine-tuning-resnet34-on-ship-detection that gives ~98% accuracy). Finally, you should keep in mind that the data was labeled by machine (likely SSD with rotating bounding boxes), and the test labeling is not ideal and has boxes rather than masks. Look at the following images. The first one is a predicted mask of the image that gives 0.82 dice. The second one is the original image with predicted (red) and \"ground truth\" (green) bounding boxes that gives 0.84 dice. You see that a model may perform better than one used for labeling the dataset, but the score is low since is calculated based on the machine labeling. So to get high score instead of building a model that is as accurate as possible one should create a model that mimics one used for labeling. \n\nBTW I'm curious, what the dice of your final image segmentation model?\n\n![enter image description here][1] ![enter image description here][2]\n\n\n  [1]: https://image.ibb.co/cgLofK/ship_detection_mask.png\n  [2]: https://image.ibb.co/gJtKnz/ship_detection_box.png",
    "390745": "Very helpful discussion.",
    "398642": "At least part of the problem is the poor segmentation quality of the test images.\nFirst column is my prediction (and it's score), 2nd column is the same prediction but approximated to a rectangle, 3rd column is the ground truth and 4th column is the full image\n\n![Sample 1][1]\n![Sample 2][2]\n![Sample 3][3]\n![Sample 4][4]\n![Sample 5][5]\n![Sample 6][6]\n![Sample 7][7]\n![Sample 8][8]\n\nI am not claiming that my predictions are perfect, but I think that they are in this (and many more) cases as good as the ground truth (which has no consistency of what it includes), but scored very badly.\n\nNote: All of my samples are from single ship images to make the metric only depending on this one ship.\n\n  [1]: https://i.imgur.com/dSE4lbm.png\n  [2]: https://i.imgur.com/QSbBLJW.png\n  [3]: https://i.imgur.com/28XED72.png\n  [4]: https://i.imgur.com/Ic9g9RX.png\n  [5]: https://i.imgur.com/WyETBls.png\n  [6]: https://i.imgur.com/pLRp7v7.png\n  [7]: https://i.imgur.com/P1blX3S.png\n  [8]: https://i.imgur.com/sg7OkRc.png",
    "398651": "Your predictions seem great. The data leakage make your rank lower (about 200 teams submit 1.000 kernel), but you may be winning potentially.",
    "399696": "I don't think.\nThe score of an empty submission (no ships at all detected is 0.847). The 2nd column gets a score of only 0.850.\n\nAnalysis shows, that I correctly identify 97% of the empty images and  only miss 0.2% of the images with ships. But I  get a score of 0.24 on the images with ships.\n\nFurther analysis show, that I even find most of the ships in these images (IoU &gt;0) , but get an extremely low score, because the IoU (shown as second number after the score) is lower than 0.5 in many cases (so the score for that ship becomes 0). \n\nAnd the reason for the low IoU seems not to be the prediction but the bad ground truth.\n\nHera are some samples\n\n![Sample 1][1]\n![Sample 2][2]\n![Sample 3][3]\n![Sample 4][4]\n![Sample 5][5]\n![Sample 6][6]\n![Sample 7][7]\n![Sample 8][8]\n\n**Note:** This are not the worst or only samples with wrong ground truth. This are average samples. I just selected some where my prediction are clearly better than the ground truth. I can find many many more.\n\n\n  [1]: https://i.imgur.com/syWu6vV.png\n  [2]: https://i.imgur.com/gvNHBA5.png\n  [3]: https://i.imgur.com/JvZzTd1.png\n  [4]: https://i.imgur.com/RDNz1Fx.png\n  [5]: https://i.imgur.com/gIPBJ5l.png\n  [6]: https://i.imgur.com/Tvgw5jF.png\n  [7]: https://i.imgur.com/woOJtLJ.png\n  [8]: https://i.imgur.com/hyP4gJw.png",
    "399745": "Very nice indeed. Only missing 0.2 % of the ships is extremely impressive. Especially given the very low number of pixels representing some of the boats.\n\n97% identification of empty picture is fine. On images with several ships the scoring punishes FP 4 times **less** than FN. But the same is not true on an image to image level. Here FP are punished just as hard as FN.\n\nThe minimum 0.5 IoU limit has a significant influence where to ships are close together. If one draws only one box around two equal sized ships the score will be zero for both ships. If it is a small and a larger ship only the larger ship will score. I think that this could be one of the reasons for the relative high 0.5 limit.",
    "400037": "I did not say, that I get 99.8% of the ships. I said that I miss 0.2% of the images with ships.\n\nThe scoring metric suggests that the challenge is to separate ships (like it was (from what I have read) important in the Data Science Bowl 2018 to separate overlapping nuclei). But it is not: 84.78% of the images contain no ship, 9.60% contain just 1 ship. So not separating ships at all (but doing everything else correct) would score at least 0.9445.\n\nAnd even separating a small number of ships,  which are far apart in most cases, is easy in post processing with traditional CV algorithms.\n\nSo the score is mostly (at least 94.45%) depending on the ship / background decision. And this part of the score is determined by the  ground truth quality. We can replicate the original mechanism by predicting a single rectangle for each ship (traditional CV). But the problem of wrong ground truth (missing parts of the ship or non minimal rectangle size) remains. This can lead (as my samples have shown) to a 0 score for a correctly identified ship.\n\nSo maybe there is more usable information in the in the ground truth itself (maybe they only used certain rectangle sizes, angels or had a raster bigger than 1 px)",
    "400213": "If you want to do an analysis of the rectangle sizes and angles, you can use a dataset that I have created https://www.kaggle.com/iafoss/rotating-bounding-boxes-for-ship-localization and a kernel that demonstrates it https://www.kaggle.com/iafoss/rotating-bounding-boxes-ship-localization . However, I couldn't boost my score much when I tried to train a second level model that accounts for this information.\n\nRegarding overlapping ships, based on my testing, they do not contribute much to the error. That I saw the main issue are displaced bounding boxes for small and intermediate ships, which essentially you also were referring to.",
    "412649": "Could I please ask you how you went from heatmaps to rectangles?"
  },
  "source": "meta"
}