{
  "id": 201816,
  "title": "why your cv is not achieving 0.9+",
  "url": "/competitions/hubmap-kidney-segmentation/discussion/201816",
  "author_name": "",
  "post_date": "2020-12-07T00:27:23.411541900Z",
  "votes": 59,
  "comment_count": 15,
  "views": 0,
  "content": "<p><a href=\"https://www.kaggle.com/c/hubmap-kidney-segmentation/discussion/198116\" target=\"_blank\">https://www.kaggle.com/c/hubmap-kidney-segmentation/discussion/198116</a><br>\n<a href=\"https://www.kaggle.com/c/hubmap-kidney-segmentation/discussion/197664\" target=\"_blank\">https://www.kaggle.com/c/hubmap-kidney-segmentation/discussion/197664</a></p>\n<p>despite the data update, i think some annotations are still shifted </p>\n<p>here is 54f2eec69 <br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fa87bd651ba3b8f00a615429c23789cf0%2FSelection_115.png?generation=1607300478590964&amp;alt=media\" alt=\"\"><br>\nground truth is from the polygon of the json file</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fb29e2cb6497be3142ee610dbc323037a%2FSelection_114.png?generation=1607300502024509&amp;alt=media\" alt=\"\"><br>\nground truth is from the rle decoding of the csv file</p>\n<hr>\n<p>i train with 7 images and validate with the remaining one.  for images without shifted, validation is often 0.920 or greater.</p>\n<p>because we trained with flip augmentation, our model is not shifted-bias, even on the train set. (without augmentation, the model would have predicted the same shift as the ground truth)</p>\n<p>i now suspect the ground truth of public test set could also be shifted (I need more proofs on this). Because submitting single image gives an unusually low scope for images c68fe75ea<br>\nand afa5e8098</p>\n<p>how about the private test set?</p>\n<p>if you made the same observation, please post it here and we can then raise the issue to the organizers.</p>\n<p>(sometimes there might be bugs in my code, etc)</p>",
  "messages": [
    {
      "id": "1104440",
      "postDate": "12/07/2020 00:27:23",
      "content": "<p><a href=\"https://www.kaggle.com/c/hubmap-kidney-segmentation/discussion/198116\" target=\"_blank\">https://www.kaggle.com/c/hubmap-kidney-segmentation/discussion/198116</a><br>\n<a href=\"https://www.kaggle.com/c/hubmap-kidney-segmentation/discussion/197664\" target=\"_blank\">https://www.kaggle.com/c/hubmap-kidney-segmentation/discussion/197664</a></p>\n<p>despite the data update, i think some annotations are still shifted </p>\n<p>here is 54f2eec69 <br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fa87bd651ba3b8f00a615429c23789cf0%2FSelection_115.png?generation=1607300478590964&amp;alt=media\" alt=\"\"><br>\nground truth is from the polygon of the json file</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fb29e2cb6497be3142ee610dbc323037a%2FSelection_114.png?generation=1607300502024509&amp;alt=media\" alt=\"\"><br>\nground truth is from the rle decoding of the csv file</p>\n<hr>\n<p>i train with 7 images and validate with the remaining one.  for images without shifted, validation is often 0.920 or greater.</p>\n<p>because we trained with flip augmentation, our model is not shifted-bias, even on the train set. (without augmentation, the model would have predicted the same shift as the ground truth)</p>\n<p>i now suspect the ground truth of public test set could also be shifted (I need more proofs on this). Because submitting single image gives an unusually low scope for images c68fe75ea<br>\nand afa5e8098</p>\n<p>how about the private test set?</p>\n<p>if you made the same observation, please post it here and we can then raise the issue to the organizers.</p>\n<p>(sometimes there might be bugs in my code, etc)</p>",
      "rawMarkdown": "https://www.kaggle.com/c/hubmap-kidney-segmentation/discussion/198116\nhttps://www.kaggle.com/c/hubmap-kidney-segmentation/discussion/197664\n\ndespite the data update, i think some annotations are still shifted \n\nhere is 54f2eec69 \n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fa87bd651ba3b8f00a615429c23789cf0%2FSelection_115.png?generation=1607300478590964&alt=media)\nground truth is from the polygon of the json file\n\n\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fb29e2cb6497be3142ee610dbc323037a%2FSelection_114.png?generation=1607300502024509&alt=media)\nground truth is from the rle decoding of the csv file\n\n---\n\ni train with 7 images and validate with the remaining one.  for images without shifted, validation is often 0.920 or greater.\n\nbecause we trained with flip augmentation, our model is not shifted-bias, even on the train set. (without augmentation, the model would have predicted the same shift as the ground truth)\n\ni now suspect the ground truth of public test set could also be shifted (I need more proofs on this). Because submitting single image gives an unusually low scope for images c68fe75ea\nand afa5e8098\n\nhow about the private test set?\n\nif you made the same observation, please post it here and we can then raise the issue to the organizers.\n\n(sometimes there might be bugs in my code, etc)",
      "votes": null
    },
    {
      "id": "1104464",
      "postDate": "12/07/2020 01:13:30",
      "content": "<p>095bf7a1f is also shifted</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F6a933d9611d79f59f695e5199a6eeb35%2FSelection_119.png?generation=1607303579694469&amp;alt=media\" alt=\"\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F52807039d2fac3973e19b34e2ad4c676%2FSelection_118.png?generation=1607303608765779&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "095bf7a1f is also shifted\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F6a933d9611d79f59f695e5199a6eeb35%2FSelection_119.png?generation=1607303579694469&alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F52807039d2fac3973e19b34e2ad4c676%2FSelection_118.png?generation=1607303608765779&alt=media)",
      "votes": null
    },
    {
      "id": "1104467",
      "postDate": "12/07/2020 01:18:57",
      "content": "<p>Thanks for mentioning this. Aside from the issue with test set not being hidden, these potential annotation mistakes are starting to make me question whether if i should continue to work on this competition. Hopefully the organizers can work something out</p>",
      "rawMarkdown": "Thanks for mentioning this. Aside from the issue with test set not being hidden, these potential annotation mistakes are starting to make me question whether if i should continue to work on this competition. Hopefully the organizers can work something out",
      "votes": null
    },
    {
      "id": "1104473",
      "postDate": "12/07/2020 01:26:12",
      "content": "<p><a href=\"https://www.kaggle.com/shujun717\" target=\"_blank\">@shujun717</a> </p>\n<p>the fact that we find so many issues shows that we can understand/analyze the data/results pretty well. once the issues are solved, it would be a good competition with little shakeup.</p>",
      "rawMarkdown": "shujun717 \n\nthe fact that we find so many issues shows that we can understand/analyze the data/results pretty well. once the issues are solved, it would be a good competition with little shakeup.",
      "votes": null
    },
    {
      "id": "1104476",
      "postDate": "12/07/2020 01:32:01",
      "content": "<p>i guess its a good thing we've found them early the competition. it can be a really good competition if the issues can be properly resolved</p>",
      "rawMarkdown": "i guess its a good thing we've found them early the competition. it can be a really good competition if the issues can be properly resolved",
      "votes": null
    },
    {
      "id": "1104480",
      "postDate": "12/07/2020 01:37:03",
      "content": "<p>here is e79de561c (the updated data). small shift is still present</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F950257705328a38027df68e2de614d7e%2FSelection_123.png?generation=1607305162863558&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "here is e79de561c (the updated data). small shift is still present\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F950257705328a38027df68e2de614d7e%2FSelection_123.png?generation=1607305162863558&alt=media)",
      "votes": null
    },
    {
      "id": "1104499",
      "postDate": "12/07/2020 02:06:07",
      "content": "<p>I am also suspecting this, on some folds, using data augmentation actually hurts the CV.</p>",
      "rawMarkdown": "I am also suspecting this, on some folds, using data augmentation actually hurts the CV.",
      "votes": null
    },
    {
      "id": "1108813",
      "postDate": "12/11/2020 03:13:44",
      "content": "<p>This is the score that we trained on each id and predicted for each id. There are two groups.</p>\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/1108812/17659/hubmap_cm1.png\" alt=\"hubmup_cm1\"></p>\n<p>GroupA = ['cb2d976f4', '0486052bb', '2f6ecfcdf', 'aaa6a05cc']<br>\nGroupB = ['095bf7a1f', 'e79de561c', '1e2425f28', '54f2eec69']</p>\n<p>It predicts well within each group, but not between GroupA and GroupB.</p>\n<p>As <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> already pointed out, '54f2eec69', '095bf7a1f,' 'e79de561c' may have shifted labels, but in addition, '1e2425f28' may have the same problem, since it is not predicted well by GroupA.</p>",
      "rawMarkdown": "This is the score that we trained on each id and predicted for each id. There are two groups.\n\n![hubmup_cm1](https://storage.googleapis.com/kaggle-forum-message-attachments/1108812/17659/hubmap_cm1.png)\n\nGroupA = ['cb2d976f4', '0486052bb', '2f6ecfcdf', 'aaa6a05cc']\nGroupB = ['095bf7a1f', 'e79de561c', '1e2425f28', '54f2eec69']\n\nIt predicts well within each group, but not between GroupA and GroupB.\n\nAs @hengck23 already pointed out, '54f2eec69', '095bf7a1f,' 'e79de561c' may have shifted labels, but in addition, '1e2425f28' may have the same problem, since it is not predicted well by GroupA.",
      "votes": null
    },
    {
      "id": "1108873",
      "postDate": "12/11/2020 05:18:18",
      "content": "<p>i think '1e2425f28' has no annotation shift.</p>\n<p>it is correct to say that prediction is very much dependent on the input train image (I think there are too few training images)</p>\n<p>here are the predictions for using different training folds, see attachment.<br>\n<img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/1108873/17660/afa5e8098.gif\" alt=\"\"></p>",
      "rawMarkdown": "i think '1e2425f28' has no annotation shift.\n\nit is correct to say that prediction is very much dependent on the input train image (I think there are too few training images)\n\nhere are the predictions for using different training folds, see attachment.\n![](https://storage.googleapis.com/kaggle-forum-message-attachments/1108873/17660/afa5e8098.gif)",
      "votes": null
    },
    {
      "id": "1108916",
      "postDate": "12/11/2020 06:34:31",
      "content": "<p>The official answer was that.<br>\n<a href=\"https://www.kaggle.com/c/hubmap-kidney-segmentation/discussion/197664\" target=\"_blank\">https://www.kaggle.com/c/hubmap-kidney-segmentation/discussion/197664</a></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6004385%2Fdf123760cd5312d0f9fabad1f78ee0eb%2FFFA55CFE-03DF-43ae-ACDA-92F9BDE1FAAD.png?generation=1607668428357714&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "The official answer was that.\nhttps://www.kaggle.com/c/hubmap-kidney-segmentation/discussion/197664\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6004385%2Fdf123760cd5312d0f9fabad1f78ee0eb%2FFFA55CFE-03DF-43ae-ACDA-92F9BDE1FAAD.png?generation=1607668428357714&alt=media)",
      "votes": null
    },
    {
      "id": "1108918",
      "postDate": "12/11/2020 06:39:20",
      "content": "<p>i don't think official answer is correct. <br>\nhe checked only \"by eye\".</p>\n<p>but the data and analysis shows otherwise.</p>\n<p>if the annotation is correct, the difference between prediction and ground truth should be random. if the difference is a shift, then the ground truth itself is shifted</p>",
      "rawMarkdown": "i don't think official answer is correct. \nhe checked only \"by eye\".\n\nbut the data and analysis shows otherwise.\n\nif the annotation is correct, the difference between prediction and ground truth should be random. if the difference is a shift, then the ground truth itself is shifted",
      "votes": null
    },
    {
      "id": "1111215",
      "postDate": "12/13/2020 14:27:46",
      "content": "<p>This is a very rich thread. Based on the discussions on why we have difficulty getting beyond ~0.9, and considering the annotation task was performed by humans, it seems the small (but evident) shifts we see are related to Bayes optimum error (human-level performance). It should be nice to know if any measures were taken (or not) to avoid it in the test set.</p>\n<p>For one thing, in the RSNA Pneumonia Detection Challenge, the bounding boxes of the test set were annotated by 3 people and the intersection of those 3 bboxes was considered the ground truth, while in the training set each image was annotated by only one person. I wonder if any similar method was done in this HuBMAP challenge. <a href=\"https://www.kaggle.com/juyingnan\" target=\"_blank\">@juyingnan</a> </p>",
      "rawMarkdown": "This is a very rich thread. Based on the discussions on why we have difficulty getting beyond ~0.9, and considering the annotation task was performed by humans, it seems the small (but evident) shifts we see are related to Bayes optimum error (human-level performance). It should be nice to know if any measures were taken (or not) to avoid it in the test set.\n\nFor one thing, in the RSNA Pneumonia Detection Challenge, the bounding boxes of the test set were annotated by 3 people and the intersection of those 3 bboxes was considered the ground truth, while in the training set each image was annotated by only one person. I wonder if any similar method was done in this HuBMAP challenge. @juyingnan",
      "votes": null
    },
    {
      "id": "1116263",
      "postDate": "12/17/2020 02:21:39",
      "content": "<p>how to verify if the test images have wrong shifted ground truth annotation.<br>\nthe basic idea is just to shift prediction and test. the question is how to do it efficiently?</p>\n<ol>\n<li>select some very high confident instances in an image</li>\n<li>dilate your prediction and submit</li>\n</ol>\n<p>you can try on your validation set. the chnage in the score is different for annotation that is shifted and not shfited<br>\nif you have a lot of submissions to spare, you can make one score matrix:</p>\n<pre><code>                        xshift\n               -10,   -5, 0, 5, 10\nyshift -10           ....\n            -5   ...    s_ij ....\n             0          ....\n             5\n            10\n\ns_ij = LB score for the one image after shifted prediction\n</code></pre>\n<p>since the score matrix is symmetrical, you can reduce the probe from M^2 to just 2M by</p>\n<pre><code>xshift  =  -10,   -5, 0, 5, 10,   and yshift=0\nyshift  =  -10,   -5, 0, 5, 10,   and xshift=0\n</code></pre>\n<p>if you want to do it in just M probes:</p>\n<pre><code>start from xshift=0.\ntry xshift=+8, -8\nchoose the better score, e.g. +8. then you can throw away all candidates xshift&lt;0 (since this score matrix is convex and has only one global max)\nrepeat\n\nsame for yshift\n</code></pre>\n<p><br>\nexample of local validation image 'id = '095bf7a1f', scanning from shift=np.arange(-32,32,4), i.e. the tick marks should have been labelled -32 to 32</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fff41137c5fa9b7120fdc8cdac943a3d4%2FSelection_054.png?generation=1608172660248970&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "how to verify if the test images have wrong shifted ground truth annotation.\nthe basic idea is just to shift prediction and test. the question is how to do it efficiently?\n1. select some very high confident instances in an image\n2. dilate your prediction and submit\n\nyou can try on your validation set. the chnage in the score is different for annotation that is shifted and not shfited\nif you have a lot of submissions to spare, you can make one score matrix:\n```\n                        xshift\n               -10,   -5, 0, 5, 10\nyshift -10           ....\n            -5   ...    s_ij ....\n             0          ....\n             5\n            10\n\ns_ij = LB score for the one image after shifted prediction\n\n```\n\nsince the score matrix is symmetrical, you can reduce the probe from M^2 to just 2M by\n\n```   \nxshift  =  -10,   -5, 0, 5, 10,   and yshift=0\nyshift  =  -10,   -5, 0, 5, 10,   and xshift=0\n```\n\nif you want to do it in just M probes:\n\n```   \nstart from xshift=0.\ntry xshift=+8, -8\nchoose the better score, e.g. +8. then you can throw away all candidates xshift<0 (since this score matrix is convex and has only one global max)\nrepeat\n\nsame for yshift\n```   \nexample of local validation image 'id = '095bf7a1f', scanning from shift=np.arange(-32,32,4), i.e. the tick marks should have been labelled -32 to 32\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fff41137c5fa9b7120fdc8cdac943a3d4%2FSelection_054.png?generation=1608172660248970&alt=media)",
      "votes": null
    },
    {
      "id": "1118896",
      "postDate": "12/19/2020 14:00:56",
      "content": "<p>[deleted due to error on my part] </p>",
      "rawMarkdown": "[deleted due to error on my part]",
      "votes": null
    },
    {
      "id": "1121585",
      "postDate": "12/21/2020 18:49:29",
      "content": "<p>Hi there, <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> <br>\nDo you generate masks from polygons yourself or get pre-generated masks from Kaggle?<br>\nI just noticed that my generated from polygons masks differ 1 pixel or so from the masks provided by Kaggle in RLE form in train.csv. Below is a crop of the difference between the two masks.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F349155%2Fdd8565ae8aa18b4e811b09ed3f86648e%2FScreenshot%20from%202020-12-21%2020-48-53.png?generation=1608576591989739&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "Hi there, @hengck23 \nDo you generate masks from polygons yourself or get pre-generated masks from Kaggle?\nI just noticed that my generated from polygons masks differ 1 pixel or so from the masks provided by Kaggle in RLE form in train.csv. Below is a crop of the difference between the two masks.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F349155%2Fdd8565ae8aa18b4e811b09ed3f86648e%2FScreenshot%20from%202020-12-21%2020-48-53.png?generation=1608576591989739&alt=media)",
      "votes": null
    },
    {
      "id": "1121974",
      "postDate": "12/22/2020 04:39:25",
      "content": "<p>in training and cross valuation, i use the RLE from cvs file.</p>\n<p>when viewing results, i use cv2 fillpoly/polyline for polygon annotation from json file to draw the outline</p>",
      "rawMarkdown": "in training and cross valuation, i use the RLE from cvs file.\n\nwhen viewing results, i use cv2 fillpoly/polyline for polygon annotation from json file to draw the outline",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1104464,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "12/07/2020 01:13:30",
      "content": "<p>095bf7a1f is also shifted</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F6a933d9611d79f59f695e5199a6eeb35%2FSelection_119.png?generation=1607303579694469&amp;alt=media\" alt=\"\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F52807039d2fac3973e19b34e2ad4c676%2FSelection_118.png?generation=1607303608765779&amp;alt=media\" alt=\"\"></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1104467,
      "author_name": "shujun717",
      "author_url": "",
      "post_date": "12/07/2020 01:18:57",
      "content": "<p>Thanks for mentioning this. Aside from the issue with test set not being hidden, these potential annotation mistakes are starting to make me question whether if i should continue to work on this competition. Hopefully the organizers can work something out</p>",
      "votes": null,
      "replies": [
        {
          "id": 1104473,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "12/07/2020 01:26:12",
          "content": "<p><a href=\"https://www.kaggle.com/shujun717\" target=\"_blank\">@shujun717</a> </p>\n<p>the fact that we find so many issues shows that we can understand/analyze the data/results pretty well. once the issues are solved, it would be a good competition with little shakeup.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1104476,
          "author_name": "shujun717",
          "author_url": "",
          "post_date": "12/07/2020 01:32:01",
          "content": "<p>i guess its a good thing we've found them early the competition. it can be a really good competition if the issues can be properly resolved</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1104480,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "12/07/2020 01:37:03",
      "content": "<p>here is e79de561c (the updated data). small shift is still present</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F950257705328a38027df68e2de614d7e%2FSelection_123.png?generation=1607305162863558&amp;alt=media\" alt=\"\"></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1104499,
      "author_name": "tugstugi",
      "author_url": "",
      "post_date": "12/07/2020 02:06:07",
      "content": "<p>I am also suspecting this, on some folds, using data augmentation actually hurts the CV.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1108813,
      "author_name": "d1348k",
      "author_url": "",
      "post_date": "12/11/2020 03:13:44",
      "content": "<p>This is the score that we trained on each id and predicted for each id. There are two groups.</p>\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/1108812/17659/hubmap_cm1.png\" alt=\"hubmup_cm1\"></p>\n<p>GroupA = ['cb2d976f4', '0486052bb', '2f6ecfcdf', 'aaa6a05cc']<br>\nGroupB = ['095bf7a1f', 'e79de561c', '1e2425f28', '54f2eec69']</p>\n<p>It predicts well within each group, but not between GroupA and GroupB.</p>\n<p>As <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> already pointed out, '54f2eec69', '095bf7a1f,' 'e79de561c' may have shifted labels, but in addition, '1e2425f28' may have the same problem, since it is not predicted well by GroupA.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1108873,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "12/11/2020 05:18:18",
          "content": "<p>i think '1e2425f28' has no annotation shift.</p>\n<p>it is correct to say that prediction is very much dependent on the input train image (I think there are too few training images)</p>\n<p>here are the predictions for using different training folds, see attachment.<br>\n<img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/1108873/17660/afa5e8098.gif\" alt=\"\"></p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1108916,
      "author_name": "zzemily",
      "author_url": "",
      "post_date": "12/11/2020 06:34:31",
      "content": "<p>The official answer was that.<br>\n<a href=\"https://www.kaggle.com/c/hubmap-kidney-segmentation/discussion/197664\" target=\"_blank\">https://www.kaggle.com/c/hubmap-kidney-segmentation/discussion/197664</a></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6004385%2Fdf123760cd5312d0f9fabad1f78ee0eb%2FFFA55CFE-03DF-43ae-ACDA-92F9BDE1FAAD.png?generation=1607668428357714&amp;alt=media\" alt=\"\"></p>",
      "votes": null,
      "replies": [
        {
          "id": 1108918,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "12/11/2020 06:39:20",
          "content": "<p>i don't think official answer is correct. <br>\nhe checked only \"by eye\".</p>\n<p>but the data and analysis shows otherwise.</p>\n<p>if the annotation is correct, the difference between prediction and ground truth should be random. if the difference is a shift, then the ground truth itself is shifted</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1111215,
      "author_name": "felipekitamura",
      "author_url": "",
      "post_date": "12/13/2020 14:27:46",
      "content": "<p>This is a very rich thread. Based on the discussions on why we have difficulty getting beyond ~0.9, and considering the annotation task was performed by humans, it seems the small (but evident) shifts we see are related to Bayes optimum error (human-level performance). It should be nice to know if any measures were taken (or not) to avoid it in the test set.</p>\n<p>For one thing, in the RSNA Pneumonia Detection Challenge, the bounding boxes of the test set were annotated by 3 people and the intersection of those 3 bboxes was considered the ground truth, while in the training set each image was annotated by only one person. I wonder if any similar method was done in this HuBMAP challenge. <a href=\"https://www.kaggle.com/juyingnan\" target=\"_blank\">@juyingnan</a> </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1116263,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "12/17/2020 02:21:39",
      "content": "<p>how to verify if the test images have wrong shifted ground truth annotation.<br>\nthe basic idea is just to shift prediction and test. the question is how to do it efficiently?</p>\n<ol>\n<li>select some very high confident instances in an image</li>\n<li>dilate your prediction and submit</li>\n</ol>\n<p>you can try on your validation set. the chnage in the score is different for annotation that is shifted and not shfited<br>\nif you have a lot of submissions to spare, you can make one score matrix:</p>\n<pre><code>                        xshift\n               -10,   -5, 0, 5, 10\nyshift -10           ....\n            -5   ...    s_ij ....\n             0          ....\n             5\n            10\n\ns_ij = LB score for the one image after shifted prediction\n</code></pre>\n<p>since the score matrix is symmetrical, you can reduce the probe from M^2 to just 2M by</p>\n<pre><code>xshift  =  -10,   -5, 0, 5, 10,   and yshift=0\nyshift  =  -10,   -5, 0, 5, 10,   and xshift=0\n</code></pre>\n<p>if you want to do it in just M probes:</p>\n<pre><code>start from xshift=0.\ntry xshift=+8, -8\nchoose the better score, e.g. +8. then you can throw away all candidates xshift&lt;0 (since this score matrix is convex and has only one global max)\nrepeat\n\nsame for yshift\n</code></pre>\n<p><br>\nexample of local validation image 'id = '095bf7a1f', scanning from shift=np.arange(-32,32,4), i.e. the tick marks should have been labelled -32 to 32</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fff41137c5fa9b7120fdc8cdac943a3d4%2FSelection_054.png?generation=1608172660248970&amp;alt=media\" alt=\"\"></p>",
      "votes": null,
      "replies": [
        {
          "id": 1118896,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "12/19/2020 14:00:56",
          "content": "<p>[deleted due to error on my part] </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1121585,
      "author_name": "sakvaua",
      "author_url": "",
      "post_date": "12/21/2020 18:49:29",
      "content": "<p>Hi there, <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> <br>\nDo you generate masks from polygons yourself or get pre-generated masks from Kaggle?<br>\nI just noticed that my generated from polygons masks differ 1 pixel or so from the masks provided by Kaggle in RLE form in train.csv. Below is a crop of the difference between the two masks.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F349155%2Fdd8565ae8aa18b4e811b09ed3f86648e%2FScreenshot%20from%202020-12-21%2020-48-53.png?generation=1608576591989739&amp;alt=media\" alt=\"\"></p>",
      "votes": null,
      "replies": [
        {
          "id": 1121974,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "12/22/2020 04:39:25",
          "content": "<p>in training and cross valuation, i use the RLE from cvs file.</p>\n<p>when viewing results, i use cv2 fillpoly/polyline for polygon annotation from json file to draw the outline</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1104440": "https://www.kaggle.com/c/hubmap-kidney-segmentation/discussion/198116\nhttps://www.kaggle.com/c/hubmap-kidney-segmentation/discussion/197664\n\ndespite the data update, i think some annotations are still shifted \n\nhere is 54f2eec69 \n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fa87bd651ba3b8f00a615429c23789cf0%2FSelection_115.png?generation=1607300478590964&alt=media)\nground truth is from the polygon of the json file\n\n\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fb29e2cb6497be3142ee610dbc323037a%2FSelection_114.png?generation=1607300502024509&alt=media)\nground truth is from the rle decoding of the csv file\n\n---\n\ni train with 7 images and validate with the remaining one.  for images without shifted, validation is often 0.920 or greater.\n\nbecause we trained with flip augmentation, our model is not shifted-bias, even on the train set. (without augmentation, the model would have predicted the same shift as the ground truth)\n\ni now suspect the ground truth of public test set could also be shifted (I need more proofs on this). Because submitting single image gives an unusually low scope for images c68fe75ea\nand afa5e8098\n\nhow about the private test set?\n\nif you made the same observation, please post it here and we can then raise the issue to the organizers.\n\n(sometimes there might be bugs in my code, etc)",
    "1104464": "095bf7a1f is also shifted\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F6a933d9611d79f59f695e5199a6eeb35%2FSelection_119.png?generation=1607303579694469&alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F52807039d2fac3973e19b34e2ad4c676%2FSelection_118.png?generation=1607303608765779&alt=media)",
    "1104467": "Thanks for mentioning this. Aside from the issue with test set not being hidden, these potential annotation mistakes are starting to make me question whether if i should continue to work on this competition. Hopefully the organizers can work something out",
    "1104473": "shujun717 \n\nthe fact that we find so many issues shows that we can understand/analyze the data/results pretty well. once the issues are solved, it would be a good competition with little shakeup.",
    "1104476": "i guess its a good thing we've found them early the competition. it can be a really good competition if the issues can be properly resolved",
    "1104480": "here is e79de561c (the updated data). small shift is still present\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F950257705328a38027df68e2de614d7e%2FSelection_123.png?generation=1607305162863558&alt=media)",
    "1104499": "I am also suspecting this, on some folds, using data augmentation actually hurts the CV.",
    "1108813": "This is the score that we trained on each id and predicted for each id. There are two groups.\n\n![hubmup_cm1](https://storage.googleapis.com/kaggle-forum-message-attachments/1108812/17659/hubmap_cm1.png)\n\nGroupA = ['cb2d976f4', '0486052bb', '2f6ecfcdf', 'aaa6a05cc']\nGroupB = ['095bf7a1f', 'e79de561c', '1e2425f28', '54f2eec69']\n\nIt predicts well within each group, but not between GroupA and GroupB.\n\nAs @hengck23 already pointed out, '54f2eec69', '095bf7a1f,' 'e79de561c' may have shifted labels, but in addition, '1e2425f28' may have the same problem, since it is not predicted well by GroupA.",
    "1108873": "i think '1e2425f28' has no annotation shift.\n\nit is correct to say that prediction is very much dependent on the input train image (I think there are too few training images)\n\nhere are the predictions for using different training folds, see attachment.\n![](https://storage.googleapis.com/kaggle-forum-message-attachments/1108873/17660/afa5e8098.gif)",
    "1108916": "The official answer was that.\nhttps://www.kaggle.com/c/hubmap-kidney-segmentation/discussion/197664\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6004385%2Fdf123760cd5312d0f9fabad1f78ee0eb%2FFFA55CFE-03DF-43ae-ACDA-92F9BDE1FAAD.png?generation=1607668428357714&alt=media)",
    "1108918": "i don't think official answer is correct. \nhe checked only \"by eye\".\n\nbut the data and analysis shows otherwise.\n\nif the annotation is correct, the difference between prediction and ground truth should be random. if the difference is a shift, then the ground truth itself is shifted",
    "1111215": "This is a very rich thread. Based on the discussions on why we have difficulty getting beyond ~0.9, and considering the annotation task was performed by humans, it seems the small (but evident) shifts we see are related to Bayes optimum error (human-level performance). It should be nice to know if any measures were taken (or not) to avoid it in the test set.\n\nFor one thing, in the RSNA Pneumonia Detection Challenge, the bounding boxes of the test set were annotated by 3 people and the intersection of those 3 bboxes was considered the ground truth, while in the training set each image was annotated by only one person. I wonder if any similar method was done in this HuBMAP challenge. @juyingnan",
    "1116263": "how to verify if the test images have wrong shifted ground truth annotation.\nthe basic idea is just to shift prediction and test. the question is how to do it efficiently?\n1. select some very high confident instances in an image\n2. dilate your prediction and submit\n\nyou can try on your validation set. the chnage in the score is different for annotation that is shifted and not shfited\nif you have a lot of submissions to spare, you can make one score matrix:\n```\n                        xshift\n               -10,   -5, 0, 5, 10\nyshift -10           ....\n            -5   ...    s_ij ....\n             0          ....\n             5\n            10\n\ns_ij = LB score for the one image after shifted prediction\n\n```\n\nsince the score matrix is symmetrical, you can reduce the probe from M^2 to just 2M by\n\n```   \nxshift  =  -10,   -5, 0, 5, 10,   and yshift=0\nyshift  =  -10,   -5, 0, 5, 10,   and xshift=0\n```\n\nif you want to do it in just M probes:\n\n```   \nstart from xshift=0.\ntry xshift=+8, -8\nchoose the better score, e.g. +8. then you can throw away all candidates xshift<0 (since this score matrix is convex and has only one global max)\nrepeat\n\nsame for yshift\n```   \nexample of local validation image 'id = '095bf7a1f', scanning from shift=np.arange(-32,32,4), i.e. the tick marks should have been labelled -32 to 32\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fff41137c5fa9b7120fdc8cdac943a3d4%2FSelection_054.png?generation=1608172660248970&alt=media)",
    "1118896": "[deleted due to error on my part]",
    "1121585": "Hi there, @hengck23 \nDo you generate masks from polygons yourself or get pre-generated masks from Kaggle?\nI just noticed that my generated from polygons masks differ 1 pixel or so from the masks provided by Kaggle in RLE form in train.csv. Below is a crop of the difference between the two masks.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F349155%2Fdd8565ae8aa18b4e811b09ed3f86648e%2FScreenshot%20from%202020-12-21%2020-48-53.png?generation=1608576591989739&alt=media)",
    "1121974": "in training and cross valuation, i use the RLE from cvs file.\n\nwhen viewing results, i use cv2 fillpoly/polyline for polygon annotation from json file to draw the outline"
  },
  "source": "meta"
}