{
  "id": 347967,
  "title": "[Experimental record] Will noisy label impede the performance of models?",
  "url": "/competitions/hubmap-organ-segmentation/discussion/347967",
  "author_name": "gray98",
  "post_date": "2022-08-26T06:44:54.491000",
  "votes": 7,
  "comment_count": 26,
  "views": 0,
  "content": "<p>As we all know, noisy label degrades the generalization performance in all machine learning tasks. However, I am not a pathologist, so I can not make the right judgement on whether this dataset contains noisy label or not. </p>\n<p>Many previous discussions have suspected that some prostate images do have omitted annotations. Therefore, I want to conduct some experiments to find out whether this dataset contains omissions or not and whether the performance of models will be improved when we find these omissions?</p>\n<p>LB-score for prostate:<br>\nTrain with all images: 0.19<br>\nTrain with only prostate images: 0.19</p>",
  "messages": [
    {
      "id": 1914519,
      "postDate": "2022-08-26T06:44:54.490Z",
      "content": "<p>As we all know, noisy label degrades the generalization performance in all machine learning tasks. However, I am not a pathologist, so I can not make the right judgement on whether this dataset contains noisy label or not. </p>\n<p>Many previous discussions have suspected that some prostate images do have omitted annotations. Therefore, I want to conduct some experiments to find out whether this dataset contains omissions or not and whether the performance of models will be improved when we find these omissions?</p>\n<p>LB-score for prostate:<br>\nTrain with all images: 0.19<br>\nTrain with only prostate images: 0.19</p>",
      "rawMarkdown": "As we all know, noisy label degrades the generalization performance in all machine learning tasks. However, I am not a pathologist, so I can not make the right judgement on whether this dataset contains noisy label or not. \n\nMany previous discussions have suspected that some prostate images do have omitted annotations. Therefore, I want to conduct some experiments to find out whether this dataset contains omissions or not and whether the performance of models will be improved when we find these omissions?\n\nLB-score for prostate:\nTrain with all images: 0.19\nTrain with only prostate images: 0.19",
      "votes": 7
    },
    {
      "id": 1918905,
      "postDate": "2022-08-30T01:29:35.930Z",
      "content": "<p>if you are interested in the topics of label noise, you can refer to<br>\n<a href=\"https://github.com/subeeshvasu/Awesome-Learning-with-Label-Noise\" target=\"_blank\">https://github.com/subeeshvasu/Awesome-Learning-with-Label-Noise</a></p>\n<p>in particular,<br>\n\"We observe a phenomenon that has been previously reported in the context of classification: the networks tend to first fit the clean pixel-level labels during an “early-learning” phase, before eventually memorizing the false annotations. However, in contrast to classification,<br>\nmemorization in segmentation does not arise simultaneously for all semantic categories\"</p>\n<p><a href=\"https://arxiv.org/pdf/2110.03740.pdf\" target=\"_blank\">https://arxiv.org/pdf/2110.03740.pdf</a><br>\n<a href=\"https://www.youtube.com/watch?v=joDAzTrNNnI\" target=\"_blank\">https://www.youtube.com/watch?v=joDAzTrNNnI</a><br>\npaper: Adaptive Early-Learning Correction for Segmentation from Noisy Annotations<br>\n<a href=\"https://ibb.co/xgG35hg\"><img src=\"https://i.ibb.co/PW4tCxW/Selection-101.png\" alt=\"Selection-101\"></a></p>\n<p>a quick way to test this is to over train a model, and submit at different iterations to see your LB.  </p>\n<p><a href=\"https://ibb.co/4Vqyh0S\"><img src=\"https://i.ibb.co/jrNKnjw/Selection-102.png\" alt=\"Selection-102\"></a></p>\n<hr>\n<p>warning:</p>\n<ul>\n<li>in most paper, the training set has corrupted labels, but the test/validation set may have clean labels.</li>\n<li>in kaggle labels are corrupted for all train and test set</li>\n</ul>",
      "rawMarkdown": "if you are interested in the topics of label noise, you can refer to\nhttps://github.com/subeeshvasu/Awesome-Learning-with-Label-Noise\n\nin particular,\n\"We observe a phenomenon that has been previously reported in the context of classification: the networks tend to first fit the clean pixel-level labels during an “early-learning” phase, before eventually memorizing the false annotations. However, in contrast to classification,\nmemorization in segmentation does not arise simultaneously for all semantic categories\"\n\nhttps://arxiv.org/pdf/2110.03740.pdf\nhttps://www.youtube.com/watch?v=joDAzTrNNnI\npaper: Adaptive Early-Learning Correction for Segmentation from Noisy Annotations\n<a href=\"https://ibb.co/xgG35hg\"><img src=\"https://i.ibb.co/PW4tCxW/Selection-101.png\" alt=\"Selection-101\" border=\"0\"></a>\n\n\na quick way to test this is to over train a model, and submit at different iterations to see your LB.  \n\n\n<a href=\"https://ibb.co/4Vqyh0S\"><img src=\"https://i.ibb.co/jrNKnjw/Selection-102.png\" alt=\"Selection-102\" border=\"0\"></a>\n\n---\n\nwarning:\n- in most paper, the training set has corrupted labels, but the test/validation set may have clean labels.\n- in kaggle labels are corrupted for all train and test set",
      "votes": 3,
      "replies": [
        {
          "id": 1918956,
          "postDate": "2022-08-30T02:22:51.767Z",
          "content": "<p>Thanks for sharing. It helps me a lot. My dissertation is about weakly-supervised semantic segmentation, so I also want to solve the same problem in your recommended paper. </p>\n<p>In local experiments, I found that relabel prostate annotations did improve performance. Although all data from this dataset has corrupted labels, a good model will show higher dice or more true positives. It is hard to make the prediction of prostate better, since it's already good.</p>\n<p>I want to deal with spleen images, which have ambiguous label in the edge of white pulp.</p>",
          "rawMarkdown": "Thanks for sharing. It helps me a lot. My dissertation is about weakly-supervised semantic segmentation, so I also want to solve the same problem in your recommended paper. \n\nIn local experiments, I found that relabel prostate annotations did improve performance. Although all data from this dataset has corrupted labels, a good model will show higher dice or more true positives. It is hard to make the prediction of prostate better, since it's already good.\n\nI want to deal with spleen images, which have ambiguous label in the edge of white pulp."
        }
      ]
    },
    {
      "id": 1931792,
      "postDate": "2022-09-09T03:59:57.770Z",
      "content": "<p>I generate some FTUs of kidney. The first row is the segmentation model trained with corrupted label (delete 50% FTUs). The second row is the segmentation model trained with complete label.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2325440%2Fde19ac1dcd88157b9418878c13e4429b%2Fgeneration.png?generation=1662704269058981&amp;alt=media\" alt=\"invert\"></p>",
      "rawMarkdown": "I generate some FTUs of kidney. The first row is the segmentation model trained with corrupted label (delete 50% FTUs). The second row is the segmentation model trained with complete label.\n![invert](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2325440%2Fde19ac1dcd88157b9418878c13e4429b%2Fgeneration.png?generation=1662704269058981&alt=media)",
      "votes": 1,
      "replies": [
        {
          "id": 1931838,
          "postDate": "2022-09-09T04:50:02.107Z",
          "content": "<p>how about results for missed detection or false positives? can you show also the original rgb input images?</p>\n<p>it seems that we are learning texture detector for the corrupted first case and   texture+shape detector for the sceond clean case</p>",
          "rawMarkdown": "how about results for missed detection or false positives? can you show also the original rgb input images?\n\nit seems that we are learning texture detector for the corrupted first case and   texture+shape detector for the sceond clean case",
          "votes": 1
        },
        {
          "id": 1931916,
          "postDate": "2022-09-09T05:46:14.827Z",
          "content": "<p>Thanks for your reply!</p>\n<p>The original rgb input images are torch.rand(1, 3, 768, 768). All results is from PSPNet (Resnet50 backbone) which is not as good as transformer. <br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2325440%2F141011ca0649f6b495e91d3ff1001335%2Fprocess.png?generation=1662703339727485&amp;alt=media\" alt=\"\"><br>\n(If you feel confused about the generation process, please tell me. I told this process to my friends, but most of them did not understand what I am doing)<br>\nThe generation process just likes \"Finding an input that maximizes a specific class\" in <a href=\"https://blog.keras.io/how-convolutional-neural-networks-see-the-world.html\" target=\"_blank\">https://blog.keras.io/how-convolutional-neural-networks-see-the-world.html</a></p>\n<p>\"how about results for missed detection or false positives?\"<br>\nFew False Positives but a lot of missed detections. These missed detections show inconspicuously higher activation. Maybe this is the reason why a low threshold of lung have a better result in LB. I need more time to redesign experiments and record detail numerical value.</p>",
          "rawMarkdown": "Thanks for your reply!\n\nThe original rgb input images are torch.rand(1, 3, 768, 768). All results is from PSPNet (Resnet50 backbone) which is not as good as transformer. \n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2325440%2F141011ca0649f6b495e91d3ff1001335%2Fprocess.png?generation=1662703339727485&alt=media)\n(If you feel confused about the generation process, please tell me. I told this process to my friends, but most of them did not understand what I am doing)\nThe generation process just likes \"Finding an input that maximizes a specific class\" in [https://blog.keras.io/how-convolutional-neural-networks-see-the-world.html](https://blog.keras.io/how-convolutional-neural-networks-see-the-world.html)\n\n\"how about results for missed detection or false positives?\"\nFew False Positives but a lot of missed detections. These missed detections show inconspicuously higher activation. Maybe this is the reason why a low threshold of lung have a better result in LB. I need more time to redesign experiments and record detail numerical value."
        },
        {
          "id": 1932188,
          "postDate": "2022-09-09T12:23:21.210Z",
          "content": "<p>i am not sure if my interpretation is correct or not.</p>\n<p>\" These missed detections show inconspicuously higher activation. \"</p>\n<p>this means that the miss FTU is found in the training input (that is why activation is high), but the part found are labelled as negative. This is only possible if the miss FTU looks like some background (or missed label) in  the training set.</p>\n<p>see if you can find \"these background (or missed label)\"</p>\n<p>if the hypothesis is correct, relabeling the missing label in ground truth should have improve performance (as the increase in FP for this would be less than the increase in additional TP)</p>",
          "rawMarkdown": "i am not sure if my interpretation is correct or not.\n\n\" These missed detections show inconspicuously higher activation. \"\n\nthis means that the miss FTU is found in the training input (that is why activation is high), but the part found are labelled as negative. This is only possible if the miss FTU looks like some background (or missed label) in  the training set.\n\nsee if you can find \"these background (or missed label)\"\n\nif the hypothesis is correct, relabeling the missing label in ground truth should have improve performance (as the increase in FP for this would be less than the increase in additional TP)"
        },
        {
          "id": 1946963,
          "postDate": "2022-09-20T07:29:08.930Z",
          "content": "<p>Thank you sir. After hearing your advise, I got some new results. To alleviate the negative influence of corrupted label, I add a regularization method to dynamically handle missed label. I observe that “the increase in FP for this would be less than the increase in additional TP” as you said. Miraculously, in early-learning stage, the training result under corrupted supervision (50%) is close to the counterpart under full supervision. Early-learning may only be observed when we adopt some methods to slow down overfitting.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2325440%2Fbb05f0bf0dc38fd44c07cf88394b2ebb%2F100epoch.png?generation=1663658928289987&amp;alt=media\" alt=\"\"><br>\nHowever, after early-learning stage, large oscillation starts, as shown in this figure.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2325440%2Fe8785346601ad1d1576a1934f5a6bbef%2F300epoch.png?generation=1663658891635505&amp;alt=media\" alt=\"\"></p>",
          "rawMarkdown": "Thank you sir. After hearing your advise, I got some new results. To alleviate the negative influence of corrupted label, I add a regularization method to dynamically handle missed label. I observe that “the increase in FP for this would be less than the increase in additional TP” as you said. Miraculously, in early-learning stage, the training result under corrupted supervision (50%) is close to the counterpart under full supervision. Early-learning may only be observed when we adopt some methods to slow down overfitting.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2325440%2Fbb05f0bf0dc38fd44c07cf88394b2ebb%2F100epoch.png?generation=1663658928289987&alt=media)\nHowever, after early-learning stage, large oscillation starts, as shown in this figure.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2325440%2Fe8785346601ad1d1576a1934f5a6bbef%2F300epoch.png?generation=1663658891635505&alt=media)\n"
        }
      ]
    },
    {
      "id": 1926592,
      "postDate": "2022-09-05T01:20:45.803Z",
      "content": "<p>i have been trying different seeds to identify bad labels (and good labels).<br>\nI think it would be overlap of the sets train images in the worst results </p>\n<pre><code>            #'../input/hubmap-submit-06-weight1/pvt_v2_b4-aug8b-768-different-seed-fold-0-swa.pth', #lb=0.78\n            #'../input/hubmap-submit-06-weight1/pvt_v2_b4-aug8b-768-different-seed-fold-1-swa.pth', #0.80\n            #'../input/hubmap-submit-06-weight1/pvt_v2_b4-aug8b-768-different-seed-fold-2-swa.pth', #0.79 \n            '../input/hubmap-submit-06-weight1/pvt_v2_b4-aug8b-768-different-seed-fold-3-swa.pth',  #0.75 !!!\n</code></pre>",
      "rawMarkdown": "i have been trying different seeds to identify bad labels (and good labels).\nI think it would be overlap of the sets train images in the worst results \n\n```\n\n            #'../input/hubmap-submit-06-weight1/pvt_v2_b4-aug8b-768-different-seed-fold-0-swa.pth', #lb=0.78\n            #'../input/hubmap-submit-06-weight1/pvt_v2_b4-aug8b-768-different-seed-fold-1-swa.pth', #0.80\n            #'../input/hubmap-submit-06-weight1/pvt_v2_b4-aug8b-768-different-seed-fold-2-swa.pth', #0.79 \n            '../input/hubmap-submit-06-weight1/pvt_v2_b4-aug8b-768-different-seed-fold-3-swa.pth',  #0.75 !!!\n\n```",
      "votes": 1,
      "replies": [
        {
          "id": 1926655,
          "postDate": "2022-09-05T03:02:16.407Z",
          "content": "<p>This looks like a lot of work. Finding bad label also is the thing I wanna explore. However, I just an engineering student without any histopathology expertise. So I try to ignore large loss pixels of certain percentage in training process, like a reversed version of OHEM or a simple version of co-teaching. But this is a bad method for this dataset, maybe because of a fixed certain percentage or the bad way of punishing hard pixels.</p>",
          "rawMarkdown": "This looks like a lot of work. Finding bad label also is the thing I wanna explore. However, I just an engineering student without any histopathology expertise. So I try to ignore large loss pixels of certain percentage in training process, like a reversed version of OHEM or a simple version of co-teaching. But this is a bad method for this dataset, maybe because of a fixed certain percentage or the bad way of punishing hard pixels."
        }
      ]
    },
    {
      "id": 1925791,
      "postDate": "2022-09-04T10:05:50.520Z",
      "content": "<p>\"In test and train process, regarding they as different classes can improve my model a little.\" do you mean add the cls head? or train with only one class images?</p>",
      "rawMarkdown": "\"In test and train process, regarding they as different classes can improve my model a little.\" do you mean add the cls head? or train with only one class images?",
      "votes": 1,
      "replies": [
        {
          "id": 1925897,
          "postDate": "2022-09-04T12:26:36.737Z",
          "content": "<p>I mean train with only one class images. In offline experiments, I find that if I train a binary segmentation model with all-class images, in validation set I observe some False positives are very … ridiculous. But when I train with only one class images, these ridiculous results are gone. This observation is from lung images. Training with only lung images improve my model in LB. I haven't had time to apply this to other organ images. If this trick improves your LB, please tell me:)</p>\n<p>I haven't tried multi-task learning yet, maybe multi-task learning can utilize the information of classification too.</p>",
          "rawMarkdown": "I mean train with only one class images. In offline experiments, I find that if I train a binary segmentation model with all-class images, in validation set I observe some False positives are very ... ridiculous. But when I train with only one class images, these ridiculous results are gone. This observation is from lung images. Training with only lung images improve my model in LB. I haven't had time to apply this to other organ images. If this trick improves your LB, please tell me:)\n\nI haven't tried multi-task learning yet, maybe multi-task learning can utilize the information of classification too."
        },
        {
          "id": 1926083,
          "postDate": "2022-09-04T14:48:15.237Z",
          "content": "<p>thanks a lot，i will try and feedback soon</p>",
          "rawMarkdown": "thanks a lot，i will try and feedback soon"
        },
        {
          "id": 1926090,
          "postDate": "2022-09-04T14:54:20.907Z",
          "content": "<p>Can you tell me your LB score of hubmap lung？</p>",
          "rawMarkdown": "Can you tell me your LB score of hubmap lung？"
        },
        {
          "id": 1926647,
          "postDate": "2022-09-05T02:37:58.803Z",
          "content": "<p>total hubmap 0.59, total lung 0.11 (train with all images / train with lung images),  hubmap lung 0.10</p>",
          "rawMarkdown": "total hubmap 0.59, total lung 0.11 (train with all images / train with lung images),  hubmap lung 0.10"
        },
        {
          "id": 1926664,
          "postDate": "2022-09-05T03:10:45.257Z",
          "content": "<p>Plus, I train lung images with many epochs. I set 1500 epochs.</p>",
          "rawMarkdown": "Plus, I train lung images with many epochs. I set 1500 epochs."
        }
      ]
    },
    {
      "id": 1925484,
      "postDate": "2022-09-04T02:25:16.033Z",
      "content": "<p>how about using deepdream to visualize what the CNN sees for the case trained with original corrupted train samples and another case with your cleaned train samples? <br>\n<a href=\"https://blog.keras.io/how-convolutional-neural-networks-see-the-world.html\" target=\"_blank\">https://blog.keras.io/how-convolutional-neural-networks-see-the-world.html</a><br>\n<img src=\"https://blog.keras.io/img/vgg16_filters_overview.jpg\" alt=\"https://blog.keras.io/img/vgg16_filters_overview.jpg\"></p>\n<p>maybe it is a good idea for judge prizes</p>\n<hr>",
      "rawMarkdown": "how about using deepdream to visualize what the CNN sees for the case trained with original corrupted train samples and another case with your cleaned train samples? \nhttps://blog.keras.io/how-convolutional-neural-networks-see-the-world.html\n![https://blog.keras.io/img/vgg16_filters_overview.jpg](https://blog.keras.io/img/vgg16_filters_overview.jpg)\n\nmaybe it is a good idea for judge prizes\n\n---\n ",
      "votes": 1,
      "replies": [
        {
          "id": 1925878,
          "postDate": "2022-09-04T12:11:25.567Z",
          "content": "<p>Leaderboard always disappoints me, but your advise always gives me confidence again! </p>\n<p>I read something about visualization, like this one: <a href=\"https://ai.googleblog.com/2015/06/inceptionism-going-deeper-into-neural.html\" target=\"_blank\">https://ai.googleblog.com/2015/06/inceptionism-going-deeper-into-neural.html</a>. The example of dumbbells proves that networks may learn co-occurrence in classification task, like dumbbells and arms regarded as dumbbells. Will networks learn wrong co-occurrence in segmentation task? When we invert a segmentation network, can it generate correct images which match their masks?</p>\n<p>I think I should stop thinking too many things and focus on what would happen when annotations are incomplete or contains mistake. Maybe before this Friday, I can complete my toy experiments and share the results.</p>",
          "rawMarkdown": "Leaderboard always disappoints me, but your advise always gives me confidence again! \n\nI read something about visualization, like this one: [https://ai.googleblog.com/2015/06/inceptionism-going-deeper-into-neural.html](https://ai.googleblog.com/2015/06/inceptionism-going-deeper-into-neural.html). The example of dumbbells proves that networks may learn co-occurrence in classification task, like dumbbells and arms regarded as dumbbells. Will networks learn wrong co-occurrence in segmentation task? When we invert a segmentation network, can it generate correct images which match their masks?\n\nI think I should stop thinking too many things and focus on what would happen when annotations are incomplete or contains mistake. Maybe before this Friday, I can complete my toy experiments and share the results."
        },
        {
          "id": 1925888,
          "postDate": "2022-09-04T12:21:04.200Z",
          "content": "<p>style transfer is also related to deepdream.</p>\n<p>it is also interesting to do inversion for HPA models on images from another domain (e.g. H&amp;E hubmap)</p>",
          "rawMarkdown": "style transfer is also related to deepdream.\n\nit is also interesting to do inversion for HPA models on images from another domain (e.g. H&E hubmap)"
        },
        {
          "id": 2004041,
          "postDate": "2022-10-26T03:18:32.393Z",
          "content": "<p>Really thanks for all your help during this competition. And I really obtain scientific prize thank to your help.</p>",
          "rawMarkdown": "Really thanks for all your help during this competition. And I really obtain scientific prize thank to your help."
        }
      ]
    },
    {
      "id": 1914564,
      "postDate": "2022-08-26T07:37:17.347Z",
      "content": "<p>Noisy labels would not be a problem if test set is all from HPA, but that's not the case. As you said, noisy labels degrade the generalization of your models. HuBMAP annotators are probably different than HPA annotators so we should somehow deal with the noisy labels.</p>",
      "rawMarkdown": "Noisy labels would not be a problem if test set is all from HPA, but that's not the case. As you said, noisy labels degrade the generalization of your models. HuBMAP annotators are probably different than HPA annotators so we should somehow deal with the noisy labels.",
      "votes": 1
    },
    {
      "id": 1914551,
      "postDate": "2022-08-26T07:28:48.233Z",
      "content": "<p>from the ranking at your submission page, you can tell which is better (i.e. an estimate of the third decimal place)</p>\n<hr>\n<p>the correct way to do is probably \"not to use your LB\". if you want to understand the problem and effects, you should do this:</p>\n<ol>\n<li>create dummy data. (i would use random noise background and use cv2.circle to draw objects and ground truth mask)</li>\n<li>now i proceed as normal, create a good dateset set with perfect ground truth and split into train validation set. train a model and record loss, etc…</li>\n<li>now if i purposely miss some labels  and created corrupted datasets, what will happen? i train model as usually using corrupted dataset.<br>\ni make observations to the training/validation metric loss, etc. more importantly, i should draw the prediction results and make visualization.<br>\nwill the learned model also made miss predictions?</li>\n<li>when i conduct the experiments using different % of missed labels, the model behavior changes</li>\n</ol>\n<p>such controlled experiments are useful for understands how things work. you can also work out what is the \"best validation metric/loss\" under given \"missing labels\" using controlled experiments and some statistics.</p>\n<hr>\n<p>actually the network don't learned individual annotations (unless you are overfitting). The network learned \"averaged annotation\".</p>",
      "rawMarkdown": "from the ranking at your submission page, you can tell which is better (i.e. an estimate of the third decimal place)\n\n---\n\nthe correct way to do is probably \"not to use your LB\". if you want to understand the problem and effects, you should do this:\n1. create dummy data. (i would use random noise background and use cv2.circle to draw objects and ground truth mask)\n2. now i proceed as normal, create a good dateset set with perfect ground truth and split into train validation set. train a model and record loss, etc...\n3. now if i purposely miss some labels  and created corrupted datasets, what will happen? i train model as usually using corrupted dataset.\ni make observations to the training/validation metric loss, etc. more importantly, i should draw the prediction results and make visualization.\nwill the learned model also made miss predictions?\n4. when i conduct the experiments using different % of missed labels, the model behavior changes\n\nsuch controlled experiments are useful for understands how things work. you can also work out what is the \"best validation metric/loss\" under given \"missing labels\" using controlled experiments and some statistics.\n\n---\n\nactually the network don't learned individual annotations (unless you are overfitting). The network learned \"averaged annotation\".",
      "votes": 1,
      "replies": [
        {
          "id": 1914622,
          "postDate": "2022-08-26T08:28:38.270Z",
          "content": "<p>Thanks for your inspiration and guidance. I will try it as your mentioned. </p>",
          "rawMarkdown": "Thanks for your inspiration and guidance. I will try it as your mentioned. "
        }
      ]
    },
    {
      "id": 1923225,
      "postDate": "2022-09-02T04:30:49.557Z",
      "content": "<p>It is my prediction of 11064 which probably has missing label. In my perspective, this result is satisfactory. I would appreciate if someone could provide a professional review.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2325440%2Fbe727ae57828477a3bd0234c3e4675bc%2F11064_label.png?generation=1662092699648610&amp;alt=media\" alt=\"11064\"></p>",
      "rawMarkdown": "It is my prediction of 11064 which probably has missing label. In my perspective, this result is satisfactory. I would appreciate if someone could provide a professional review.\n![11064](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2325440%2Fbe727ae57828477a3bd0234c3e4675bc%2F11064_label.png?generation=1662092699648610&alt=media)"
    },
    {
      "id": 1921967,
      "postDate": "2022-09-01T07:19:09.620Z",
      "content": "<p>In the past two days, I rethink the our task. We should segment FTUs in several organs, but these organs are independent. For example, the FTUs of spleens and the FTUs of lungs have no connection. In test and train process, regarding they as different classes can improve my model a little.</p>\n<p>I assume our annotations contain noisy. Plus, I add some noise (model prediction) into original annotations, so now it actually becomes a noisy label problem. Some noisy label researches like early-learning, focusing on easy samples, co-teaching did not show their power in my experiments of lung, maybe because the task itself is challenging.</p>\n<p>In <a href=\"https://www.kaggle.com/competitions/hubmap-organ-segmentation/discussion/344276\" target=\"_blank\">https://www.kaggle.com/competitions/hubmap-organ-segmentation/discussion/344276</a>, auriml reported an example with missing label. I check the prediction of this images, my model actually can detect it.</p>",
      "rawMarkdown": "In the past two days, I rethink the our task. We should segment FTUs in several organs, but these organs are independent. For example, the FTUs of spleens and the FTUs of lungs have no connection. In test and train process, regarding they as different classes can improve my model a little.\n\nI assume our annotations contain noisy. Plus, I add some noise (model prediction) into original annotations, so now it actually becomes a noisy label problem. Some noisy label researches like early-learning, focusing on easy samples, co-teaching did not show their power in my experiments of lung, maybe because the task itself is challenging.\n\nIn https://www.kaggle.com/competitions/hubmap-organ-segmentation/discussion/344276, auriml reported an example with missing label. I check the prediction of this images, my model actually can detect it."
    },
    {
      "id": 1917884,
      "postDate": "2022-08-29T06:18:47.550Z",
      "content": "<p>I try to automatically relabel prostate images and use the new annotation in the training process. However, the LB-score for prostate remains 0.19. My total score changed from 0.80 (which is very close to 0.81 according to my ranking) to 0.81 (which is the lowest 0.81). It can not prove that considering noisy label will improve the performance of my models.</p>",
      "rawMarkdown": "I try to automatically relabel prostate images and use the new annotation in the training process. However, the LB-score for prostate remains 0.19. My total score changed from 0.80 (which is very close to 0.81 according to my ranking) to 0.81 (which is the lowest 0.81). It can not prove that considering noisy label will improve the performance of my models."
    },
    {
      "id": 1955814,
      "postDate": "2022-09-26T06:37:24.400Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 1918905,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2022-08-30T01:29:35.930000",
      "content": "<p>if you are interested in the topics of label noise, you can refer to<br>\n<a href=\"https://github.com/subeeshvasu/Awesome-Learning-with-Label-Noise\" target=\"_blank\">https://github.com/subeeshvasu/Awesome-Learning-with-Label-Noise</a></p>\n<p>in particular,<br>\n\"We observe a phenomenon that has been previously reported in the context of classification: the networks tend to first fit the clean pixel-level labels during an “early-learning” phase, before eventually memorizing the false annotations. However, in contrast to classification,<br>\nmemorization in segmentation does not arise simultaneously for all semantic categories\"</p>\n<p><a href=\"https://arxiv.org/pdf/2110.03740.pdf\" target=\"_blank\">https://arxiv.org/pdf/2110.03740.pdf</a><br>\n<a href=\"https://www.youtube.com/watch?v=joDAzTrNNnI\" target=\"_blank\">https://www.youtube.com/watch?v=joDAzTrNNnI</a><br>\npaper: Adaptive Early-Learning Correction for Segmentation from Noisy Annotations<br>\n<a href=\"https://ibb.co/xgG35hg\"><img src=\"https://i.ibb.co/PW4tCxW/Selection-101.png\" alt=\"Selection-101\"></a></p>\n<p>a quick way to test this is to over train a model, and submit at different iterations to see your LB.  </p>\n<p><a href=\"https://ibb.co/4Vqyh0S\"><img src=\"https://i.ibb.co/jrNKnjw/Selection-102.png\" alt=\"Selection-102\"></a></p>\n<hr>\n<p>warning:</p>\n<ul>\n<li>in most paper, the training set has corrupted labels, but the test/validation set may have clean labels.</li>\n<li>in kaggle labels are corrupted for all train and test set</li>\n</ul>",
      "votes": 3,
      "replies": [
        {
          "id": 1918956,
          "author_name": "gray98",
          "author_url": "",
          "post_date": "2022-08-30T02:22:51.767000",
          "content": "<p>Thanks for sharing. It helps me a lot. My dissertation is about weakly-supervised semantic segmentation, so I also want to solve the same problem in your recommended paper. </p>\n<p>In local experiments, I found that relabel prostate annotations did improve performance. Although all data from this dataset has corrupted labels, a good model will show higher dice or more true positives. It is hard to make the prediction of prostate better, since it's already good.</p>\n<p>I want to deal with spleen images, which have ambiguous label in the edge of white pulp.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1931792,
      "author_name": "gray98",
      "author_url": "",
      "post_date": "2022-09-09T03:59:57.770000",
      "content": "<p>I generate some FTUs of kidney. The first row is the segmentation model trained with corrupted label (delete 50% FTUs). The second row is the segmentation model trained with complete label.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2325440%2Fde19ac1dcd88157b9418878c13e4429b%2Fgeneration.png?generation=1662704269058981&amp;alt=media\" alt=\"invert\"></p>",
      "votes": 1,
      "replies": [
        {
          "id": 1931838,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2022-09-09T04:50:02.107000",
          "content": "<p>how about results for missed detection or false positives? can you show also the original rgb input images?</p>\n<p>it seems that we are learning texture detector for the corrupted first case and   texture+shape detector for the sceond clean case</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1931916,
          "author_name": "gray98",
          "author_url": "",
          "post_date": "2022-09-09T05:46:14.827000",
          "content": "<p>Thanks for your reply!</p>\n<p>The original rgb input images are torch.rand(1, 3, 768, 768). All results is from PSPNet (Resnet50 backbone) which is not as good as transformer. <br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2325440%2F141011ca0649f6b495e91d3ff1001335%2Fprocess.png?generation=1662703339727485&amp;alt=media\" alt=\"\"><br>\n(If you feel confused about the generation process, please tell me. I told this process to my friends, but most of them did not understand what I am doing)<br>\nThe generation process just likes \"Finding an input that maximizes a specific class\" in <a href=\"https://blog.keras.io/how-convolutional-neural-networks-see-the-world.html\" target=\"_blank\">https://blog.keras.io/how-convolutional-neural-networks-see-the-world.html</a></p>\n<p>\"how about results for missed detection or false positives?\"<br>\nFew False Positives but a lot of missed detections. These missed detections show inconspicuously higher activation. Maybe this is the reason why a low threshold of lung have a better result in LB. I need more time to redesign experiments and record detail numerical value.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1932188,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2022-09-09T12:23:21.210000",
          "content": "<p>i am not sure if my interpretation is correct or not.</p>\n<p>\" These missed detections show inconspicuously higher activation. \"</p>\n<p>this means that the miss FTU is found in the training input (that is why activation is high), but the part found are labelled as negative. This is only possible if the miss FTU looks like some background (or missed label) in  the training set.</p>\n<p>see if you can find \"these background (or missed label)\"</p>\n<p>if the hypothesis is correct, relabeling the missing label in ground truth should have improve performance (as the increase in FP for this would be less than the increase in additional TP)</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1946963,
          "author_name": "gray98",
          "author_url": "",
          "post_date": "2022-09-20T07:29:08.930000",
          "content": "<p>Thank you sir. After hearing your advise, I got some new results. To alleviate the negative influence of corrupted label, I add a regularization method to dynamically handle missed label. I observe that “the increase in FP for this would be less than the increase in additional TP” as you said. Miraculously, in early-learning stage, the training result under corrupted supervision (50%) is close to the counterpart under full supervision. Early-learning may only be observed when we adopt some methods to slow down overfitting.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2325440%2Fbb05f0bf0dc38fd44c07cf88394b2ebb%2F100epoch.png?generation=1663658928289987&amp;alt=media\" alt=\"\"><br>\nHowever, after early-learning stage, large oscillation starts, as shown in this figure.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2325440%2Fe8785346601ad1d1576a1934f5a6bbef%2F300epoch.png?generation=1663658891635505&amp;alt=media\" alt=\"\"></p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1926592,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2022-09-05T01:20:45.803000",
      "content": "<p>i have been trying different seeds to identify bad labels (and good labels).<br>\nI think it would be overlap of the sets train images in the worst results </p>\n<pre><code>            #'../input/hubmap-submit-06-weight1/pvt_v2_b4-aug8b-768-different-seed-fold-0-swa.pth', #lb=0.78\n            #'../input/hubmap-submit-06-weight1/pvt_v2_b4-aug8b-768-different-seed-fold-1-swa.pth', #0.80\n            #'../input/hubmap-submit-06-weight1/pvt_v2_b4-aug8b-768-different-seed-fold-2-swa.pth', #0.79 \n            '../input/hubmap-submit-06-weight1/pvt_v2_b4-aug8b-768-different-seed-fold-3-swa.pth',  #0.75 !!!\n</code></pre>",
      "votes": 1,
      "replies": [
        {
          "id": 1926655,
          "author_name": "gray98",
          "author_url": "",
          "post_date": "2022-09-05T03:02:16.407000",
          "content": "<p>This looks like a lot of work. Finding bad label also is the thing I wanna explore. However, I just an engineering student without any histopathology expertise. So I try to ignore large loss pixels of certain percentage in training process, like a reversed version of OHEM or a simple version of co-teaching. But this is a bad method for this dataset, maybe because of a fixed certain percentage or the bad way of punishing hard pixels.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1925791,
      "author_name": "pky",
      "author_url": "",
      "post_date": "2022-09-04T10:05:50.520000",
      "content": "<p>\"In test and train process, regarding they as different classes can improve my model a little.\" do you mean add the cls head? or train with only one class images?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1925897,
          "author_name": "gray98",
          "author_url": "",
          "post_date": "2022-09-04T12:26:36.737000",
          "content": "<p>I mean train with only one class images. In offline experiments, I find that if I train a binary segmentation model with all-class images, in validation set I observe some False positives are very … ridiculous. But when I train with only one class images, these ridiculous results are gone. This observation is from lung images. Training with only lung images improve my model in LB. I haven't had time to apply this to other organ images. If this trick improves your LB, please tell me:)</p>\n<p>I haven't tried multi-task learning yet, maybe multi-task learning can utilize the information of classification too.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1926083,
          "author_name": "pky",
          "author_url": "",
          "post_date": "2022-09-04T14:48:15.237000",
          "content": "<p>thanks a lot，i will try and feedback soon</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1926090,
          "author_name": "pky",
          "author_url": "",
          "post_date": "2022-09-04T14:54:20.907000",
          "content": "<p>Can you tell me your LB score of hubmap lung？</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1926647,
          "author_name": "gray98",
          "author_url": "",
          "post_date": "2022-09-05T02:37:58.803000",
          "content": "<p>total hubmap 0.59, total lung 0.11 (train with all images / train with lung images),  hubmap lung 0.10</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1926664,
          "author_name": "gray98",
          "author_url": "",
          "post_date": "2022-09-05T03:10:45.257000",
          "content": "<p>Plus, I train lung images with many epochs. I set 1500 epochs.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1925484,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2022-09-04T02:25:16.033000",
      "content": "<p>how about using deepdream to visualize what the CNN sees for the case trained with original corrupted train samples and another case with your cleaned train samples? <br>\n<a href=\"https://blog.keras.io/how-convolutional-neural-networks-see-the-world.html\" target=\"_blank\">https://blog.keras.io/how-convolutional-neural-networks-see-the-world.html</a><br>\n<img src=\"https://blog.keras.io/img/vgg16_filters_overview.jpg\" alt=\"https://blog.keras.io/img/vgg16_filters_overview.jpg\"></p>\n<p>maybe it is a good idea for judge prizes</p>\n<hr>",
      "votes": 1,
      "replies": [
        {
          "id": 1925878,
          "author_name": "gray98",
          "author_url": "",
          "post_date": "2022-09-04T12:11:25.567000",
          "content": "<p>Leaderboard always disappoints me, but your advise always gives me confidence again! </p>\n<p>I read something about visualization, like this one: <a href=\"https://ai.googleblog.com/2015/06/inceptionism-going-deeper-into-neural.html\" target=\"_blank\">https://ai.googleblog.com/2015/06/inceptionism-going-deeper-into-neural.html</a>. The example of dumbbells proves that networks may learn co-occurrence in classification task, like dumbbells and arms regarded as dumbbells. Will networks learn wrong co-occurrence in segmentation task? When we invert a segmentation network, can it generate correct images which match their masks?</p>\n<p>I think I should stop thinking too many things and focus on what would happen when annotations are incomplete or contains mistake. Maybe before this Friday, I can complete my toy experiments and share the results.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1925888,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2022-09-04T12:21:04.200000",
          "content": "<p>style transfer is also related to deepdream.</p>\n<p>it is also interesting to do inversion for HPA models on images from another domain (e.g. H&amp;E hubmap)</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 2004041,
          "author_name": "gray98",
          "author_url": "",
          "post_date": "2022-10-26T03:18:32.393000",
          "content": "<p>Really thanks for all your help during this competition. And I really obtain scientific prize thank to your help.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1914564,
      "author_name": "Gunes Evitan",
      "author_url": "",
      "post_date": "2022-08-26T07:37:17.347000",
      "content": "<p>Noisy labels would not be a problem if test set is all from HPA, but that's not the case. As you said, noisy labels degrade the generalization of your models. HuBMAP annotators are probably different than HPA annotators so we should somehow deal with the noisy labels.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1914551,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2022-08-26T07:28:48.233000",
      "content": "<p>from the ranking at your submission page, you can tell which is better (i.e. an estimate of the third decimal place)</p>\n<hr>\n<p>the correct way to do is probably \"not to use your LB\". if you want to understand the problem and effects, you should do this:</p>\n<ol>\n<li>create dummy data. (i would use random noise background and use cv2.circle to draw objects and ground truth mask)</li>\n<li>now i proceed as normal, create a good dateset set with perfect ground truth and split into train validation set. train a model and record loss, etc…</li>\n<li>now if i purposely miss some labels  and created corrupted datasets, what will happen? i train model as usually using corrupted dataset.<br>\ni make observations to the training/validation metric loss, etc. more importantly, i should draw the prediction results and make visualization.<br>\nwill the learned model also made miss predictions?</li>\n<li>when i conduct the experiments using different % of missed labels, the model behavior changes</li>\n</ol>\n<p>such controlled experiments are useful for understands how things work. you can also work out what is the \"best validation metric/loss\" under given \"missing labels\" using controlled experiments and some statistics.</p>\n<hr>\n<p>actually the network don't learned individual annotations (unless you are overfitting). The network learned \"averaged annotation\".</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1914622,
          "author_name": "gray98",
          "author_url": "",
          "post_date": "2022-08-26T08:28:38.270000",
          "content": "<p>Thanks for your inspiration and guidance. I will try it as your mentioned. </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1923225,
      "author_name": "gray98",
      "author_url": "",
      "post_date": "2022-09-02T04:30:49.557000",
      "content": "<p>It is my prediction of 11064 which probably has missing label. In my perspective, this result is satisfactory. I would appreciate if someone could provide a professional review.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2325440%2Fbe727ae57828477a3bd0234c3e4675bc%2F11064_label.png?generation=1662092699648610&amp;alt=media\" alt=\"11064\"></p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1921967,
      "author_name": "gray98",
      "author_url": "",
      "post_date": "2022-09-01T07:19:09.620000",
      "content": "<p>In the past two days, I rethink the our task. We should segment FTUs in several organs, but these organs are independent. For example, the FTUs of spleens and the FTUs of lungs have no connection. In test and train process, regarding they as different classes can improve my model a little.</p>\n<p>I assume our annotations contain noisy. Plus, I add some noise (model prediction) into original annotations, so now it actually becomes a noisy label problem. Some noisy label researches like early-learning, focusing on easy samples, co-teaching did not show their power in my experiments of lung, maybe because the task itself is challenging.</p>\n<p>In <a href=\"https://www.kaggle.com/competitions/hubmap-organ-segmentation/discussion/344276\" target=\"_blank\">https://www.kaggle.com/competitions/hubmap-organ-segmentation/discussion/344276</a>, auriml reported an example with missing label. I check the prediction of this images, my model actually can detect it.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1917884,
      "author_name": "gray98",
      "author_url": "",
      "post_date": "2022-08-29T06:18:47.550000",
      "content": "<p>I try to automatically relabel prostate images and use the new annotation in the training process. However, the LB-score for prostate remains 0.19. My total score changed from 0.80 (which is very close to 0.81 according to my ranking) to 0.81 (which is the lowest 0.81). It can not prove that considering noisy label will improve the performance of my models.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1955814,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-09-26T06:37:24.400000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1914519": "As we all know, noisy label degrades the generalization performance in all machine learning tasks. However, I am not a pathologist, so I can not make the right judgement on whether this dataset contains noisy label or not. \n\nMany previous discussions have suspected that some prostate images do have omitted annotations. Therefore, I want to conduct some experiments to find out whether this dataset contains omissions or not and whether the performance of models will be improved when we find these omissions?\n\nLB-score for prostate:\nTrain with all images: 0.19\nTrain with only prostate images: 0.19",
    "1918905": "if you are interested in the topics of label noise, you can refer to\nhttps://github.com/subeeshvasu/Awesome-Learning-with-Label-Noise\n\nin particular,\n\"We observe a phenomenon that has been previously reported in the context of classification: the networks tend to first fit the clean pixel-level labels during an “early-learning” phase, before eventually memorizing the false annotations. However, in contrast to classification,\nmemorization in segmentation does not arise simultaneously for all semantic categories\"\n\nhttps://arxiv.org/pdf/2110.03740.pdf\nhttps://www.youtube.com/watch?v=joDAzTrNNnI\npaper: Adaptive Early-Learning Correction for Segmentation from Noisy Annotations\n<a href=\"https://ibb.co/xgG35hg\"><img src=\"https://i.ibb.co/PW4tCxW/Selection-101.png\" alt=\"Selection-101\" border=\"0\"></a>\n\n\na quick way to test this is to over train a model, and submit at different iterations to see your LB.  \n\n\n<a href=\"https://ibb.co/4Vqyh0S\"><img src=\"https://i.ibb.co/jrNKnjw/Selection-102.png\" alt=\"Selection-102\" border=\"0\"></a>\n\n---\n\nwarning:\n- in most paper, the training set has corrupted labels, but the test/validation set may have clean labels.\n- in kaggle labels are corrupted for all train and test set",
    "1931792": "I generate some FTUs of kidney. The first row is the segmentation model trained with corrupted label (delete 50% FTUs). The second row is the segmentation model trained with complete label.\n![invert](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2325440%2Fde19ac1dcd88157b9418878c13e4429b%2Fgeneration.png?generation=1662704269058981&alt=media)",
    "1926592": "i have been trying different seeds to identify bad labels (and good labels).\nI think it would be overlap of the sets train images in the worst results \n\n```\n\n            #'../input/hubmap-submit-06-weight1/pvt_v2_b4-aug8b-768-different-seed-fold-0-swa.pth', #lb=0.78\n            #'../input/hubmap-submit-06-weight1/pvt_v2_b4-aug8b-768-different-seed-fold-1-swa.pth', #0.80\n            #'../input/hubmap-submit-06-weight1/pvt_v2_b4-aug8b-768-different-seed-fold-2-swa.pth', #0.79 \n            '../input/hubmap-submit-06-weight1/pvt_v2_b4-aug8b-768-different-seed-fold-3-swa.pth',  #0.75 !!!\n\n```",
    "1925791": "\"In test and train process, regarding they as different classes can improve my model a little.\" do you mean add the cls head? or train with only one class images?",
    "1925484": "how about using deepdream to visualize what the CNN sees for the case trained with original corrupted train samples and another case with your cleaned train samples? \nhttps://blog.keras.io/how-convolutional-neural-networks-see-the-world.html\n![https://blog.keras.io/img/vgg16_filters_overview.jpg](https://blog.keras.io/img/vgg16_filters_overview.jpg)\n\nmaybe it is a good idea for judge prizes\n\n---\n ",
    "1914564": "Noisy labels would not be a problem if test set is all from HPA, but that's not the case. As you said, noisy labels degrade the generalization of your models. HuBMAP annotators are probably different than HPA annotators so we should somehow deal with the noisy labels.",
    "1914551": "from the ranking at your submission page, you can tell which is better (i.e. an estimate of the third decimal place)\n\n---\n\nthe correct way to do is probably \"not to use your LB\". if you want to understand the problem and effects, you should do this:\n1. create dummy data. (i would use random noise background and use cv2.circle to draw objects and ground truth mask)\n2. now i proceed as normal, create a good dateset set with perfect ground truth and split into train validation set. train a model and record loss, etc...\n3. now if i purposely miss some labels  and created corrupted datasets, what will happen? i train model as usually using corrupted dataset.\ni make observations to the training/validation metric loss, etc. more importantly, i should draw the prediction results and make visualization.\nwill the learned model also made miss predictions?\n4. when i conduct the experiments using different % of missed labels, the model behavior changes\n\nsuch controlled experiments are useful for understands how things work. you can also work out what is the \"best validation metric/loss\" under given \"missing labels\" using controlled experiments and some statistics.\n\n---\n\nactually the network don't learned individual annotations (unless you are overfitting). The network learned \"averaged annotation\".",
    "1923225": "It is my prediction of 11064 which probably has missing label. In my perspective, this result is satisfactory. I would appreciate if someone could provide a professional review.\n![11064](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2325440%2Fbe727ae57828477a3bd0234c3e4675bc%2F11064_label.png?generation=1662092699648610&alt=media)",
    "1921967": "In the past two days, I rethink the our task. We should segment FTUs in several organs, but these organs are independent. For example, the FTUs of spleens and the FTUs of lungs have no connection. In test and train process, regarding they as different classes can improve my model a little.\n\nI assume our annotations contain noisy. Plus, I add some noise (model prediction) into original annotations, so now it actually becomes a noisy label problem. Some noisy label researches like early-learning, focusing on easy samples, co-teaching did not show their power in my experiments of lung, maybe because the task itself is challenging.\n\nIn https://www.kaggle.com/competitions/hubmap-organ-segmentation/discussion/344276, auriml reported an example with missing label. I check the prediction of this images, my model actually can detect it.",
    "1917884": "I try to automatically relabel prostate images and use the new annotation in the training process. However, the LB-score for prostate remains 0.19. My total score changed from 0.80 (which is very close to 0.81 according to my ranking) to 0.81 (which is the lowest 0.81). It can not prove that considering noisy label will improve the performance of my models.",
    "1955814": ""
  }
}