{
  "id": 234751,
  "title": "Segmentation boundary noise problem",
  "url": "/competitions/hubmap-kidney-segmentation/discussion/234751",
  "author_name": "lafe",
  "post_date": "2021-04-26T03:01:51.753000",
  "votes": 27,
  "comment_count": 7,
  "views": 0,
  "content": "<p>Update</p>\n<p>By analyzing the verification set pictures, I think that if the dice is lower than 0.95, most of the verification set pictures are not caused by edge problems. <br>\nThe dice value is too low and is caused by abnormal cropping of pictures, or missed and misdetected</p>\n<p>Each row of pictures is the label and predicted concat<br>\n(Original image + label, original image + prediction)<br>\n(label, mask)<br>\n(Intersection, Difference)<br>\n<img src=\"https://i.loli.net/2021/05/01/adjfTnNzbEQemX9.png\" alt=\"截屏2021-05-01 下午10.44.04.png\"></p>\n<p>--- 分割线<br>\n通过分析验证集图片，我认为如果dice值低于0.9的图片绝大部分不是由于边缘问题导致的, 更多的是分割图的预测产生漏检和误检, 或者是裁剪图片时产生了过小的mask. </p>\n<p>每一行图片为标签和预测的concat<br>\n(原图+label, 原图+预测)<br>\n(label, mask)<br>\n(交集, 差集)</p>\n<hr>\n<p>I divided the data set into a training set and a validation set. My best model can reach 0.937 in lb, and the validation set can reach 0.95+, but I found that my dice on the training set is only about 0.948, and the c68fe75ea In the figure, only 0.90 dice.<br>\nThe following figure is the data visualization effect: <br>\n(Yellow represents the intersection of the prediction and the mask, green represents the false detection of the prediction, and red represents the missed detection)</p>\n<p><img src=\"https://i.loli.net/2021/04/26/OudCTlXRHpU15Sm.jpg\" alt=\"WechatIMG405.jpeg\"><br>\n<img src=\"https://i.loli.net/2021/04/26/I3Nj4b8FW2srSkd.png\" alt=\"WechatIMG411.png\"></p>\n<p>I have a few guesses about this:</p>\n<ol>\n<li>The boundary labeling of cells contains noise. Although segmentation has a strong ability to fit data, such task networks with less diversity tend to learn the commonality of the target, which leads to confrontation, so in some cases, the network becomes more and more trained Poor, and the training set dice cannot be improved.</li>\n<li>Since the official said that the training set and test set labeling methods are the same, in order to improve the score, do you want the model to learn this artificial labeling noise?</li>\n</ol>\n<p>Program: <br>\nFor adjusting edges:<br>\n1.In the training phase, the mask is randomly shrunk by 1 to 3 pixels inward.</p>\n<ol>\n<li>Use SegFix post-processing</li>\n<li>Use multiple model results to refine the mask</li>\n</ol>\n<p>I haven't thought of a good way to learn noise.</p>\n<p>华丽的分割线--------<br>\n我将数据集分为训练集和验证集, 我最好的模型在lb可以达到0.937, 验证集可以到0.95+, 但是我发现, 我在训练集上的dice也只有0.948左右, 而且 c68fe75ea 这张图, 只有0.90的dice.<br>\n图片在上方 (黄色代表预测和mask相交,  绿色代表预测误检, 红色代表漏检)</p>\n<p>对此我有几个猜想:</p>\n<ol>\n<li>细胞的边界标注包含噪声, 虽然分割具有强拟合数据的能力, 但是多样性较少的此类任务网络往往会学习到目标的共性, 这就产生了对抗, 所以有些情况网络越训越差, 并且训练集dice不能提升. <br>\n2.由于官方表示训练集和测试集标注方式是一致的, 那么为了提升分数, 是否要让模型学习这种人工标注噪声?</li>\n</ol>\n<p>方案: <br>\n对于调整边缘:</p>\n<ol>\n<li>训练阶段随机使mask 向内收缩1到3个像素.</li>\n<li>使用 SegFix 后处理 </li>\n<li>使用多个模型结果对mask进行refine</li>\n</ol>\n<p>对于学习噪声我还没有想到好的方法.  哎…</p>",
  "messages": [
    {
      "id": 1284507,
      "postDate": "2021-04-26T03:01:51.753Z",
      "content": "<p>Update</p>\n<p>By analyzing the verification set pictures, I think that if the dice is lower than 0.95, most of the verification set pictures are not caused by edge problems. <br>\nThe dice value is too low and is caused by abnormal cropping of pictures, or missed and misdetected</p>\n<p>Each row of pictures is the label and predicted concat<br>\n(Original image + label, original image + prediction)<br>\n(label, mask)<br>\n(Intersection, Difference)<br>\n<img src=\"https://i.loli.net/2021/05/01/adjfTnNzbEQemX9.png\" alt=\"截屏2021-05-01 下午10.44.04.png\"></p>\n<p>--- 分割线<br>\n通过分析验证集图片，我认为如果dice值低于0.9的图片绝大部分不是由于边缘问题导致的, 更多的是分割图的预测产生漏检和误检, 或者是裁剪图片时产生了过小的mask. </p>\n<p>每一行图片为标签和预测的concat<br>\n(原图+label, 原图+预测)<br>\n(label, mask)<br>\n(交集, 差集)</p>\n<hr>\n<p>I divided the data set into a training set and a validation set. My best model can reach 0.937 in lb, and the validation set can reach 0.95+, but I found that my dice on the training set is only about 0.948, and the c68fe75ea In the figure, only 0.90 dice.<br>\nThe following figure is the data visualization effect: <br>\n(Yellow represents the intersection of the prediction and the mask, green represents the false detection of the prediction, and red represents the missed detection)</p>\n<p><img src=\"https://i.loli.net/2021/04/26/OudCTlXRHpU15Sm.jpg\" alt=\"WechatIMG405.jpeg\"><br>\n<img src=\"https://i.loli.net/2021/04/26/I3Nj4b8FW2srSkd.png\" alt=\"WechatIMG411.png\"></p>\n<p>I have a few guesses about this:</p>\n<ol>\n<li>The boundary labeling of cells contains noise. Although segmentation has a strong ability to fit data, such task networks with less diversity tend to learn the commonality of the target, which leads to confrontation, so in some cases, the network becomes more and more trained Poor, and the training set dice cannot be improved.</li>\n<li>Since the official said that the training set and test set labeling methods are the same, in order to improve the score, do you want the model to learn this artificial labeling noise?</li>\n</ol>\n<p>Program: <br>\nFor adjusting edges:<br>\n1.In the training phase, the mask is randomly shrunk by 1 to 3 pixels inward.</p>\n<ol>\n<li>Use SegFix post-processing</li>\n<li>Use multiple model results to refine the mask</li>\n</ol>\n<p>I haven't thought of a good way to learn noise.</p>\n<p>华丽的分割线--------<br>\n我将数据集分为训练集和验证集, 我最好的模型在lb可以达到0.937, 验证集可以到0.95+, 但是我发现, 我在训练集上的dice也只有0.948左右, 而且 c68fe75ea 这张图, 只有0.90的dice.<br>\n图片在上方 (黄色代表预测和mask相交,  绿色代表预测误检, 红色代表漏检)</p>\n<p>对此我有几个猜想:</p>\n<ol>\n<li>细胞的边界标注包含噪声, 虽然分割具有强拟合数据的能力, 但是多样性较少的此类任务网络往往会学习到目标的共性, 这就产生了对抗, 所以有些情况网络越训越差, 并且训练集dice不能提升. <br>\n2.由于官方表示训练集和测试集标注方式是一致的, 那么为了提升分数, 是否要让模型学习这种人工标注噪声?</li>\n</ol>\n<p>方案: <br>\n对于调整边缘:</p>\n<ol>\n<li>训练阶段随机使mask 向内收缩1到3个像素.</li>\n<li>使用 SegFix 后处理 </li>\n<li>使用多个模型结果对mask进行refine</li>\n</ol>\n<p>对于学习噪声我还没有想到好的方法.  哎…</p>",
      "rawMarkdown": "Update\n\n\nBy analyzing the verification set pictures, I think that if the dice is lower than 0.95, most of the verification set pictures are not caused by edge problems. \nThe dice value is too low and is caused by abnormal cropping of pictures, or missed and misdetected\n\n\nEach row of pictures is the label and predicted concat\n(Original image + label, original image + prediction)\n(label, mask)\n(Intersection, Difference)\n![截屏2021-05-01 下午10.44.04.png](https://i.loli.net/2021/05/01/adjfTnNzbEQemX9.png)\n\n--- 分割线\n通过分析验证集图片，我认为如果dice值低于0.9的图片绝大部分不是由于边缘问题导致的, 更多的是分割图的预测产生漏检和误检, 或者是裁剪图片时产生了过小的mask. \n\n每一行图片为标签和预测的concat\n(原图+label, 原图+预测)\n(label, mask)\n(交集, 差集)\n\n----------------------------------------\nI divided the data set into a training set and a validation set. My best model can reach 0.937 in lb, and the validation set can reach 0.95+, but I found that my dice on the training set is only about 0.948, and the c68fe75ea In the figure, only 0.90 dice.\nThe following figure is the data visualization effect: \n(Yellow represents the intersection of the prediction and the mask, green represents the false detection of the prediction, and red represents the missed detection)\n\n![WechatIMG405.jpeg](https://i.loli.net/2021/04/26/OudCTlXRHpU15Sm.jpg)\n![WechatIMG411.png](https://i.loli.net/2021/04/26/I3Nj4b8FW2srSkd.png)\n\nI have a few guesses about this:\n1. The boundary labeling of cells contains noise. Although segmentation has a strong ability to fit data, such task networks with less diversity tend to learn the commonality of the target, which leads to confrontation, so in some cases, the network becomes more and more trained Poor, and the training set dice cannot be improved.\n2. Since the official said that the training set and test set labeling methods are the same, in order to improve the score, do you want the model to learn this artificial labeling noise?\n\nProgram: \nFor adjusting edges:\n1.In the training phase, the mask is randomly shrunk by 1 to 3 pixels inward.\n2. Use SegFix post-processing\n3. Use multiple model results to refine the mask\n\nI haven't thought of a good way to learn noise.\n\n\n华丽的分割线--------\n我将数据集分为训练集和验证集, 我最好的模型在lb可以达到0.937, 验证集可以到0.95+, 但是我发现, 我在训练集上的dice也只有0.948左右, 而且 c68fe75ea 这张图, 只有0.90的dice.\n图片在上方 (黄色代表预测和mask相交,  绿色代表预测误检, 红色代表漏检)\n\n对此我有几个猜想:\n1. 细胞的边界标注包含噪声, 虽然分割具有强拟合数据的能力, 但是多样性较少的此类任务网络往往会学习到目标的共性, 这就产生了对抗, 所以有些情况网络越训越差, 并且训练集dice不能提升. \n2.由于官方表示训练集和测试集标注方式是一致的, 那么为了提升分数, 是否要让模型学习这种人工标注噪声?\n\n方案: \n对于调整边缘:\n1. 训练阶段随机使mask 向内收缩1到3个像素.\n2. 使用 SegFix 后处理 \n3. 使用多个模型结果对mask进行refine\n\n对于学习噪声我还没有想到好的方法.  哎...\n\n\n\n\n",
      "votes": 27
    },
    {
      "id": 1284752,
      "postDate": "2021-04-26T08:41:35.060Z",
      "content": "<p>I don't believe it makes sense to learn noise, it may be helpful if we can train more robust model. BTW, your score is already very high and impressive.</p>",
      "rawMarkdown": "I don't believe it makes sense to learn noise, it may be helpful if we can train more robust model. BTW, your score is already very high and impressive.",
      "votes": 3
    },
    {
      "id": 1286523,
      "postDate": "2021-04-28T05:09:37.347Z",
      "content": "<p>the gaps are not regular<br>\n<img src=\"https://i.loli.net/2021/04/28/8AYnCvKu314HhGL.png\" alt=\"1619586244_1_.png\"></p>\n<p><img src=\"https://i.loli.net/2021/04/28/E9irTfM3dVnohqe.png\" alt=\"1619586364_1_.png\"></p>",
      "rawMarkdown": "the gaps are not regular\n![1619586244_1_.png](https://i.loli.net/2021/04/28/8AYnCvKu314HhGL.png)\n\n![1619586364_1_.png](https://i.loli.net/2021/04/28/E9irTfM3dVnohqe.png)",
      "votes": 1
    },
    {
      "id": 1284541,
      "postDate": "2021-04-26T04:22:09.787Z",
      "content": "<p>There is a post <a href=\"https://www.kaggle.com/c/hubmap-kidney-segmentation/discussion/230089\" target=\"_blank\">here</a> on the evaluation metric.  Perhaps the idea could be adapted to dealing with edges in prediction or post process. </p>\n<p>Also this post <a href=\"https://www.kaggle.com/c/hubmap-kidney-segmentation/discussion/232124#1275721\" target=\"_blank\">here</a> on label errors. </p>",
      "rawMarkdown": "There is a post [here](https://www.kaggle.com/c/hubmap-kidney-segmentation/discussion/230089) on the evaluation metric.  Perhaps the idea could be adapted to dealing with edges in prediction or post process. \n\nAlso this post [here](https://www.kaggle.com/c/hubmap-kidney-segmentation/discussion/232124#1275721) on label errors. ",
      "votes": 1
    },
    {
      "id": 1284686,
      "postDate": "2021-04-26T07:45:03.347Z",
      "content": "<p>弹性形变可能可以有用。</p>",
      "rawMarkdown": "弹性形变可能可以有用。"
    },
    {
      "id": 1290005,
      "postDate": "2021-05-01T14:48:58.083Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 1288436,
      "postDate": "2021-04-30T02:23:31.023Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 1287980,
      "postDate": "2021-04-29T15:11:26Z",
      "content": "<p>Thank you for your sharing</p>",
      "rawMarkdown": "Thank you for your sharing"
    }
  ],
  "comments": [
    {
      "id": 1284752,
      "author_name": "Yi Wu",
      "author_url": "",
      "post_date": "2021-04-26T08:41:35.060000",
      "content": "<p>I don't believe it makes sense to learn noise, it may be helpful if we can train more robust model. BTW, your score is already very high and impressive.</p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 1286523,
      "author_name": "朴大福",
      "author_url": "",
      "post_date": "2021-04-28T05:09:37.347000",
      "content": "<p>the gaps are not regular<br>\n<img src=\"https://i.loli.net/2021/04/28/8AYnCvKu314HhGL.png\" alt=\"1619586244_1_.png\"></p>\n<p><img src=\"https://i.loli.net/2021/04/28/E9irTfM3dVnohqe.png\" alt=\"1619586364_1_.png\"></p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1284541,
      "author_name": "something4kag",
      "author_url": "",
      "post_date": "2021-04-26T04:22:09.787000",
      "content": "<p>There is a post <a href=\"https://www.kaggle.com/c/hubmap-kidney-segmentation/discussion/230089\" target=\"_blank\">here</a> on the evaluation metric.  Perhaps the idea could be adapted to dealing with edges in prediction or post process. </p>\n<p>Also this post <a href=\"https://www.kaggle.com/c/hubmap-kidney-segmentation/discussion/232124#1275721\" target=\"_blank\">here</a> on label errors. </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1284686,
      "author_name": "朴大福",
      "author_url": "",
      "post_date": "2021-04-26T07:45:03.347000",
      "content": "<p>弹性形变可能可以有用。</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1290005,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-05-01T14:48:58.083000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1288436,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-04-30T02:23:31.023000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1287980,
      "author_name": "Fengxiao",
      "author_url": "",
      "post_date": "2021-04-29T15:11:26",
      "content": "<p>Thank you for your sharing</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1284507": "Update\n\n\nBy analyzing the verification set pictures, I think that if the dice is lower than 0.95, most of the verification set pictures are not caused by edge problems. \nThe dice value is too low and is caused by abnormal cropping of pictures, or missed and misdetected\n\n\nEach row of pictures is the label and predicted concat\n(Original image + label, original image + prediction)\n(label, mask)\n(Intersection, Difference)\n![截屏2021-05-01 下午10.44.04.png](https://i.loli.net/2021/05/01/adjfTnNzbEQemX9.png)\n\n--- 分割线\n通过分析验证集图片，我认为如果dice值低于0.9的图片绝大部分不是由于边缘问题导致的, 更多的是分割图的预测产生漏检和误检, 或者是裁剪图片时产生了过小的mask. \n\n每一行图片为标签和预测的concat\n(原图+label, 原图+预测)\n(label, mask)\n(交集, 差集)\n\n----------------------------------------\nI divided the data set into a training set and a validation set. My best model can reach 0.937 in lb, and the validation set can reach 0.95+, but I found that my dice on the training set is only about 0.948, and the c68fe75ea In the figure, only 0.90 dice.\nThe following figure is the data visualization effect: \n(Yellow represents the intersection of the prediction and the mask, green represents the false detection of the prediction, and red represents the missed detection)\n\n![WechatIMG405.jpeg](https://i.loli.net/2021/04/26/OudCTlXRHpU15Sm.jpg)\n![WechatIMG411.png](https://i.loli.net/2021/04/26/I3Nj4b8FW2srSkd.png)\n\nI have a few guesses about this:\n1. The boundary labeling of cells contains noise. Although segmentation has a strong ability to fit data, such task networks with less diversity tend to learn the commonality of the target, which leads to confrontation, so in some cases, the network becomes more and more trained Poor, and the training set dice cannot be improved.\n2. Since the official said that the training set and test set labeling methods are the same, in order to improve the score, do you want the model to learn this artificial labeling noise?\n\nProgram: \nFor adjusting edges:\n1.In the training phase, the mask is randomly shrunk by 1 to 3 pixels inward.\n2. Use SegFix post-processing\n3. Use multiple model results to refine the mask\n\nI haven't thought of a good way to learn noise.\n\n\n华丽的分割线--------\n我将数据集分为训练集和验证集, 我最好的模型在lb可以达到0.937, 验证集可以到0.95+, 但是我发现, 我在训练集上的dice也只有0.948左右, 而且 c68fe75ea 这张图, 只有0.90的dice.\n图片在上方 (黄色代表预测和mask相交,  绿色代表预测误检, 红色代表漏检)\n\n对此我有几个猜想:\n1. 细胞的边界标注包含噪声, 虽然分割具有强拟合数据的能力, 但是多样性较少的此类任务网络往往会学习到目标的共性, 这就产生了对抗, 所以有些情况网络越训越差, 并且训练集dice不能提升. \n2.由于官方表示训练集和测试集标注方式是一致的, 那么为了提升分数, 是否要让模型学习这种人工标注噪声?\n\n方案: \n对于调整边缘:\n1. 训练阶段随机使mask 向内收缩1到3个像素.\n2. 使用 SegFix 后处理 \n3. 使用多个模型结果对mask进行refine\n\n对于学习噪声我还没有想到好的方法.  哎...\n\n\n\n\n",
    "1284752": "I don't believe it makes sense to learn noise, it may be helpful if we can train more robust model. BTW, your score is already very high and impressive.",
    "1286523": "the gaps are not regular\n![1619586244_1_.png](https://i.loli.net/2021/04/28/8AYnCvKu314HhGL.png)\n\n![1619586364_1_.png](https://i.loli.net/2021/04/28/E9irTfM3dVnohqe.png)",
    "1284541": "There is a post [here](https://www.kaggle.com/c/hubmap-kidney-segmentation/discussion/230089) on the evaluation metric.  Perhaps the idea could be adapted to dealing with edges in prediction or post process. \n\nAlso this post [here](https://www.kaggle.com/c/hubmap-kidney-segmentation/discussion/232124#1275721) on label errors. ",
    "1284686": "弹性形变可能可以有用。",
    "1290005": "",
    "1288436": "",
    "1287980": "Thank you for your sharing"
  }
}