{
  "id": 344276,
  "title": "Ground Truth: Are all functional units found in an image labeled? ",
  "url": "/competitions/hubmap-organ-segmentation/discussion/344276",
  "author_name": "",
  "post_date": "2022-08-14T17:05:44.142739200Z",
  "votes": 17,
  "comment_count": 16,
  "views": 0,
  "content": "<p><a href=\"https://www.kaggle.com/yashvrdnjain\" target=\"_blank\">@yashvrdnjain</a> <br>\nAt least in the images of lung tissue, it seems that not all aveoli (functional units) are being labeled with masks. I have asked an experienced pathologist and she has confirmed that for example on the image below the region highlighted with an arrow is also another alveolus but has not been masked.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F865377%2Fb339f98c78c778fbbb320631403e16e3%2FScreen%20Shot%202022-08-14%20at%2003.34.51.png?generation=1660496680735529&amp;alt=media\" alt=\"\"></p>\n<p>Thanks in advance for your answer,<br>\nBR,</p>",
  "messages": [
    {
      "id": "1898600",
      "postDate": "08/14/2022 17:05:44",
      "content": "<p><a href=\"https://www.kaggle.com/yashvrdnjain\" target=\"_blank\">@yashvrdnjain</a> <br>\nAt least in the images of lung tissue, it seems that not all aveoli (functional units) are being labeled with masks. I have asked an experienced pathologist and she has confirmed that for example on the image below the region highlighted with an arrow is also another alveolus but has not been masked.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F865377%2Fb339f98c78c778fbbb320631403e16e3%2FScreen%20Shot%202022-08-14%20at%2003.34.51.png?generation=1660496680735529&amp;alt=media\" alt=\"\"></p>\n<p>Thanks in advance for your answer,<br>\nBR,</p>",
      "rawMarkdown": "yashvrdnjain \nAt least in the images of lung tissue, it seems that not all aveoli (functional units) are being labeled with masks. I have asked an experienced pathologist and she has confirmed that for example on the image below the region highlighted with an arrow is also another alveolus but has not been masked.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F865377%2Fb339f98c78c778fbbb320631403e16e3%2FScreen%20Shot%202022-08-14%20at%2003.34.51.png?generation=1660496680735529&alt=media)\n\nThanks in advance for your answer,\nBR,",
      "votes": null
    },
    {
      "id": "1898665",
      "postDate": "08/14/2022 18:04:37",
      "content": "<p>There is unmasked labels all over dataset, the dataset is most definitely masked with their algorithm. Some most notable examples…</p>\n<ul>\n<li>RED: Mask</li>\n<li>BLUE: Predictions</li>\n<li>VIOLET: Mask &amp;&amp; Predictions</li>\n</ul>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2386017%2F012fa733e2aa324f672ab44505d0f9f4%2Ff7c24a6a-bee1-40b5-9c6b-db6928f2feee.png?generation=1660499929307393&amp;alt=media\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2386017%2F889e8d9f8e2b0624d22833fcfed704d0%2Fmedia_images_on_epoch_end_%20validate_761_15af07af5e2cbca49b9a.png?generation=1660499941166791&amp;alt=media\"></p>",
      "rawMarkdown": "There is unmasked labels all over dataset, the dataset is most definitely masked with their algorithm. Some most notable examples...\n\n* RED: Mask\n* BLUE: Predictions\n* VIOLET: Mask && Predictions\n\n<img width=\"300\"  src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2386017%2F012fa733e2aa324f672ab44505d0f9f4%2Ff7c24a6a-bee1-40b5-9c6b-db6928f2feee.png?generation=1660499929307393&alt=media\"></img>\n\n<img width=\"300\"  src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2386017%2F889e8d9f8e2b0624d22833fcfed704d0%2Fmedia_images_on_epoch_end_%20validate_761_15af07af5e2cbca49b9a.png?generation=1660499941166791&alt=media\"></img>",
      "votes": null
    },
    {
      "id": "1898686",
      "postDate": "08/14/2022 18:23:59",
      "content": "<p>Here are some more examples….<br>\nIf you train your model, there should be a generalization of model detection even if the label is missing. My model overcomes this problem quite nicely and I don't think there should be a problem if the dataset is the missing label. However, there will be problems with lung detection, my model decided to joust skip mask detection in the lung section. 😄</p>\n<p>I'm really curious how grandmasters decided to tackle this problem with lungs…</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2386017%2F4db024bf1463a420e5ca76ee486e3f16%2Fc57d55c4-6779-41b4-97fb-ef6769f71b79.png?generation=1660500985178077&amp;alt=media\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2386017%2F6bb17657625e85232355eb2fe6fa6aa7%2F20be363b-c67d-4a58-8130-9fdb1f2aec58.png?generation=1660500974523051&amp;alt=media\"></p>",
      "rawMarkdown": "Here are some more examples....\nIf you train your model, there should be a generalization of model detection even if the label is missing. My model overcomes this problem quite nicely and I don't think there should be a problem if the dataset is the missing label. However, there will be problems with lung detection, my model decided to joust skip mask detection in the lung section. 😄\n\nI'm really curious how grandmasters decided to tackle this problem with lungs...\n\n<img width=\"300\" src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2386017%2F4db024bf1463a420e5ca76ee486e3f16%2Fc57d55c4-6779-41b4-97fb-ef6769f71b79.png?generation=1660500985178077&alt=media\"></img>\n\n<img width=\"300\" src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2386017%2F6bb17657625e85232355eb2fe6fa6aa7%2F20be363b-c67d-4a58-8130-9fdb1f2aec58.png?generation=1660500974523051&alt=media\"></img>",
      "votes": null
    },
    {
      "id": "1898694",
      "postDate": "08/14/2022 18:29:42",
      "content": "<p>Labeling is very inconsistent. For large intestine and prostate some images have FPUs on the border segmented, while in others they are not. Oh, and don't get me started on lungs. It's a mess. Probably a result of some automatic segmenting. Kidneys are ok.</p>",
      "rawMarkdown": "Labeling is very inconsistent. For large intestine and prostate some images have FPUs on the border segmented, while in others they are not. Oh, and don't get me started on lungs. It's a mess. Probably a result of some automatic segmenting. Kidneys are ok.",
      "votes": null
    },
    {
      "id": "1898704",
      "postDate": "08/14/2022 18:39:47",
      "content": "<p>I'm wondering what special measures to take to tackle the lung problem?</p>",
      "rawMarkdown": "I'm wondering what special measures to take to tackle the lung problem?",
      "votes": null
    },
    {
      "id": "1898708",
      "postDate": "08/14/2022 18:44:47",
      "content": "<p>if you want, you can have 3 class : background, FTU-lablled, FTU-not-labelled</p>",
      "rawMarkdown": "if you want, you can have 3 class : background, FTU-lablled, FTU-not-labelled",
      "votes": null
    },
    {
      "id": "1898915",
      "postDate": "08/14/2022 23:58:00",
      "content": "<p>if you take a look at edges, its most likely manually anotated, but without a strict rule set for labels.<br>\nI believe, that for the alveoli only nice examples have been annotated. Which leads to many not annotated examples.<br>\nThe big question is: do we want to learn the task or the annotator?</p>",
      "rawMarkdown": "if you take a look at edges, its most likely manually anotated, but without a strict rule set for labels.\nI believe, that for the alveoli only nice examples have been annotated. Which leads to many not annotated examples.\nThe big question is: do we want to learn the task or the annotator?",
      "votes": null
    },
    {
      "id": "1899041",
      "postDate": "08/15/2022 03:08:57",
      "content": "<p>\" do we want to learn the task or the annotator?\"</p>\n<p>from application  point of view (judges prize) : learn the task</p>\n<p>from performance ranking competition point of view (supervised ML prize) : learn the annotator</p>\n<hr>\n<p>kaggle needs to develop their own code for data cleaning or organize a data cleaning up competition for us to help</p>",
      "rawMarkdown": "\" do we want to learn the task or the annotator?\"\n\nfrom application  point of view (judges prize) : learn the task\n\nfrom performance ranking competition point of view (supervised ML prize) : learn the annotator\n\n---\n\nkaggle needs to develop their own code for data cleaning or organize a data cleaning up competition for us to help",
      "votes": null
    },
    {
      "id": "1899601",
      "postDate": "08/15/2022 11:41:05",
      "content": "<p>Labeling can be an iterative process as well because there will always be errors in annotation tasks.<br>\nBy iteratively training models on k folds and predicting on the validation data, the worst errors could be found and corrected easily.</p>\n<p>Maybe it could also be a possibility, to actually let the users suggest label changes and let the hosts review them.</p>",
      "rawMarkdown": "Labeling can be an iterative process as well because there will always be errors in annotation tasks.\nBy iteratively training models on k folds and predicting on the validation data, the worst errors could be found and corrected easily.\n\nMaybe it could also be a possibility, to actually let the users suggest label changes and let the hosts review them.",
      "votes": null
    },
    {
      "id": "1899684",
      "postDate": "08/15/2022 12:54:00",
      "content": "<p>Hello, how dos you get the second set of annotation? I was following the RLE in <code>train.csv</code></p>",
      "rawMarkdown": "Hello, how dos you get the second set of annotation? I was following the RLE in `train.csv`",
      "votes": null
    },
    {
      "id": "1899753",
      "postDate": "08/15/2022 13:35:52",
      "content": "<p>The second set is just predictions from your model.</p>",
      "rawMarkdown": "The second set is just predictions from your model.",
      "votes": null
    },
    {
      "id": "1901484",
      "postDate": "08/16/2022 17:18:58",
      "content": "<p>A few quick thoughts:</p>\n<ul>\n<li>It's quite difficult to get labels for this kind of data. Manual work with segmentation masks is time consuming and in this case requires medical expertise. I can assure you that a great deal of effort went into labeling this dataset.</li>\n<li>It's also quite common for medical challenges to have limits on the quality of the labels. For example, I've worked on multiple medical classification problems where we knew the experts only agreed on the label about 80% of the time. Those competitions were still quite successful for both the hosts and competitors. </li>\n</ul>\n<p>To me, there are two questions to focus on when discussing label quality in the context of a competition:</p>\n<ul>\n<li>Is there a systematic error? For example, if the lower left quarter of each image was never labeled I would be concerned. That doesn't appear to be the case here.</li>\n<li>Is the label quality the limiting factor for model performance? We have run competitions in the past where there wasn't much signal and it was <em>all</em> extracted. In those cases the public and private leaderboard scores cluster closely around some less-than-perfect value. Scores in this competition have a wide range and continue to improve, which indicates to me that the models aren't currently limited by the label quality.</li>\n</ul>\n<p>If you're still concerned about the label quality I would recommend making an estimate of how much any label shortcomings would reduce the score of a truly perfect set of masks. If we're still a long way from that, then you don't have to worry that label quality is defining the competition.</p>",
      "rawMarkdown": "A few quick thoughts:\n\n- It's quite difficult to get labels for this kind of data. Manual work with segmentation masks is time consuming and in this case requires medical expertise. I can assure you that a great deal of effort went into labeling this dataset.\n- It's also quite common for medical challenges to have limits on the quality of the labels. For example, I've worked on multiple medical classification problems where we knew the experts only agreed on the label about 80% of the time. Those competitions were still quite successful for both the hosts and competitors. \n\nTo me, there are two questions to focus on when discussing label quality in the context of a competition:\n- Is there a systematic error? For example, if the lower left quarter of each image was never labeled I would be concerned. That doesn't appear to be the case here.\n- Is the label quality the limiting factor for model performance? We have run competitions in the past where there wasn't much signal and it was _all_ extracted. In those cases the public and private leaderboard scores cluster closely around some less-than-perfect value. Scores in this competition have a wide range and continue to improve, which indicates to me that the models aren't currently limited by the label quality.\n\nIf you're still concerned about the label quality I would recommend making an estimate of how much any label shortcomings would reduce the score of a truly perfect set of masks. If we're still a long way from that, then you don't have to worry that label quality is defining the competition.",
      "votes": null
    },
    {
      "id": "1901576",
      "postDate": "08/16/2022 19:01:08",
      "content": "<p>I have no problem with the dataset being incomplete, but I have problems when it comes to dataset transparency from the host. I get the point that this is the competition and not every detail around the private dataset should be revealed. But I would expect some level of transparency about…</p>\n<ul>\n<li>Who or what created labels?</li>\n<li>What is the expected accuracy of labels in the private and public datasets?</li>\n<li>Distribution of image sizes in the private dataset.</li>\n<li>Label class distribution on the private dataset. (Is the dataset balanced)</li>\n<li>Why are there no images in the public dataset 160x160 size but there are such images in the private dataset?</li>\n<li>etc…</li>\n</ul>\n<p>I would expect from you Kaggle Staff to ask the host before he creates a competition to provide such information so that we can really focus on providing the best-generalized models and not to create hacks around datasets. With such information, we could provide better models without revealing to much information about private datasets.</p>\n<p>I love Kaggle community and the fact that there is a sincere wish in the kaggle community to create the best of the best results and to provide good service to scientific world. But, hiding information, and not being transparent is something that makes me question the whole competition and their motives…</p>",
      "rawMarkdown": "I have no problem with the dataset being incomplete, but I have problems when it comes to dataset transparency from the host. I get the point that this is the competition and not every detail around the private dataset should be revealed. But I would expect some level of transparency about...\n\n- Who or what created labels?\n- What is the expected accuracy of labels in the private and public datasets?\n- Distribution of image sizes in the private dataset.\n- Label class distribution on the private dataset. (Is the dataset balanced)\n- Why are there no images in the public dataset 160x160 size but there are such images in the private dataset?\n- etc...\n\nI would expect from you Kaggle Staff to ask the host before he creates a competition to provide such information so that we can really focus on providing the best-generalized models and not to create hacks around datasets. With such information, we could provide better models without revealing to much information about private datasets.\n\nI love Kaggle community and the fact that there is a sincere wish in the kaggle community to create the best of the best results and to provide good service to scientific world. But, hiding information, and not being transparent is something that makes me question the whole competition and their motives...",
      "votes": null
    },
    {
      "id": "1901609",
      "postDate": "08/16/2022 19:23:52",
      "content": "<p><a href=\"https://www.kaggle.com/urosjarc\" target=\"_blank\">@urosjarc</a> our standard practice is always to disclose as little as possible about the private dataset, for obvious reasons. Frankly, the data description for this competition goes into an unusual level of detail about the hidden test. Deciding what to disclose is a joint decision between Kaggle and the host and I'm very comfortable with the status quo.</p>",
      "rawMarkdown": "urosjarc our standard practice is always to disclose as little as possible about the private dataset, for obvious reasons. Frankly, the data description for this competition goes into an unusual level of detail about the hidden test. Deciding what to disclose is a joint decision between Kaggle and the host and I'm very comfortable with the status quo.",
      "votes": null
    },
    {
      "id": "1902914",
      "postDate": "08/17/2022 02:03:28",
      "content": "<p><a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a> I don't think that answering such questions should compromise the competition in any way… This is my first competition so I don't know how things work here, but isn't it in the host's best interest to provide the community with as much information as possible he can to get the best models and to still make the competition valid?</p>\n<blockquote>\n  <p>our standard practice is always to disclose as little as possible about the private dataset</p>\n</blockquote>\n<p>Yea I get it it's competition something should be hidden if this is compeitition,  but still, I failed to realize to see how answering such questions will degrade competition validity.</p>",
      "rawMarkdown": "sohier I don't think that answering such questions should compromise the competition in any way... This is my first competition so I don't know how things work here, but isn't it in the host's best interest to provide the community with as much information as possible he can to get the best models and to still make the competition valid?\n\n> our standard practice is always to disclose as little as possible about the private dataset\n\nYea I get it it's competition something should be hidden if this is compeitition,  but still, I failed to realize to see how answering such questions will degrade competition validity.",
      "votes": null
    },
    {
      "id": "1903359",
      "postDate": "08/17/2022 10:59:30",
      "content": "<p><a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a> From the clinical point of view, could you please define/list which of the following structures in the lung are the target?</p>\n<ul>\n<li>All observable alveoli (including alveolar septa and enclosed air)</li>\n<li>All observable alveolar ducts</li>\n<li>All observable terminal bronchiole </li>\n</ul>\n<p>The FTU of the lung comprises all of the above within a lung lobule  (please refer to Pubmed source below and image from <a href=\"url\" target=\"_blank\">DOI:10.1136/jclinpath-2014-202685</a> ). <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F865377%2F31fd895f35c25c62fe6dfeba693de510%2Flung_FTU.png?generation=1660733618800459&amp;alt=media\" alt=\"\"></p>\n<p>As it can be observed, the anatomy of the lungs, where each lung lobule has a complex reticular structure with interconnected alveoli (as those have pores between them), makes it unfeasible to  attempt to separate alveoli into isolated units with separate masks. <br>\nI would appreciate clarification from the clinical staff involved on the labeling of this competition.</p>\n<p>Thanks in advance, </p>\n<p>Aurelia Bustos, MD, PhD<br>\nMedical Oncologist</p>",
      "rawMarkdown": "sohier From the clinical point of view, could you please define/list which of the following structures in the lung are the target?\n- All observable alveoli (including alveolar septa and enclosed air)\n- All observable alveolar ducts\n- All observable terminal bronchiole \n\nThe FTU of the lung comprises all of the above within a lung lobule  (please refer to Pubmed source below and image from [DOI:10.1136/jclinpath-2014-202685](url) ). ![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F865377%2F31fd895f35c25c62fe6dfeba693de510%2Flung_FTU.png?generation=1660733618800459&alt=media)\n\n\nAs it can be observed, the anatomy of the lungs, where each lung lobule has a complex reticular structure with interconnected alveoli (as those have pores between them), makes it unfeasible to  attempt to separate alveoli into isolated units with separate masks. \nI would appreciate clarification from the clinical staff involved on the labeling of this competition.\n\nThanks in advance, \n\nAurelia Bustos, MD, PhD\nMedical Oncologist",
      "votes": null
    },
    {
      "id": "1903901",
      "postDate": "08/17/2022 18:28:28",
      "content": "<p>The experts annotated whole, uncollapsed alveolar spaces, focusing on every individual instance that can be clearly recognized as an alveolus. </p>",
      "rawMarkdown": "The experts annotated whole, uncollapsed alveolar spaces, focusing on every individual instance that can be clearly recognized as an alveolus.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1898665,
      "author_name": "urosjarc",
      "author_url": "",
      "post_date": "08/14/2022 18:04:37",
      "content": "<p>There is unmasked labels all over dataset, the dataset is most definitely masked with their algorithm. Some most notable examples…</p>\n<ul>\n<li>RED: Mask</li>\n<li>BLUE: Predictions</li>\n<li>VIOLET: Mask &amp;&amp; Predictions</li>\n</ul>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2386017%2F012fa733e2aa324f672ab44505d0f9f4%2Ff7c24a6a-bee1-40b5-9c6b-db6928f2feee.png?generation=1660499929307393&amp;alt=media\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2386017%2F889e8d9f8e2b0624d22833fcfed704d0%2Fmedia_images_on_epoch_end_%20validate_761_15af07af5e2cbca49b9a.png?generation=1660499941166791&amp;alt=media\"></p>",
      "votes": null,
      "replies": [
        {
          "id": 1898686,
          "author_name": "urosjarc",
          "author_url": "",
          "post_date": "08/14/2022 18:23:59",
          "content": "<p>Here are some more examples….<br>\nIf you train your model, there should be a generalization of model detection even if the label is missing. My model overcomes this problem quite nicely and I don't think there should be a problem if the dataset is the missing label. However, there will be problems with lung detection, my model decided to joust skip mask detection in the lung section. 😄</p>\n<p>I'm really curious how grandmasters decided to tackle this problem with lungs…</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2386017%2F4db024bf1463a420e5ca76ee486e3f16%2Fc57d55c4-6779-41b4-97fb-ef6769f71b79.png?generation=1660500985178077&amp;alt=media\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2386017%2F6bb17657625e85232355eb2fe6fa6aa7%2F20be363b-c67d-4a58-8130-9fdb1f2aec58.png?generation=1660500974523051&amp;alt=media\"></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1898915,
          "author_name": "theudas",
          "author_url": "",
          "post_date": "08/14/2022 23:58:00",
          "content": "<p>if you take a look at edges, its most likely manually anotated, but without a strict rule set for labels.<br>\nI believe, that for the alveoli only nice examples have been annotated. Which leads to many not annotated examples.<br>\nThe big question is: do we want to learn the task or the annotator?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1899041,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "08/15/2022 03:08:57",
          "content": "<p>\" do we want to learn the task or the annotator?\"</p>\n<p>from application  point of view (judges prize) : learn the task</p>\n<p>from performance ranking competition point of view (supervised ML prize) : learn the annotator</p>\n<hr>\n<p>kaggle needs to develop their own code for data cleaning or organize a data cleaning up competition for us to help</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1899601,
          "author_name": "theudas",
          "author_url": "",
          "post_date": "08/15/2022 11:41:05",
          "content": "<p>Labeling can be an iterative process as well because there will always be errors in annotation tasks.<br>\nBy iteratively training models on k folds and predicting on the validation data, the worst errors could be found and corrected easily.</p>\n<p>Maybe it could also be a possibility, to actually let the users suggest label changes and let the hosts review them.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1899684,
          "author_name": "jirkaborovec",
          "author_url": "",
          "post_date": "08/15/2022 12:54:00",
          "content": "<p>Hello, how dos you get the second set of annotation? I was following the RLE in <code>train.csv</code></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1899753,
          "author_name": "sakvaua",
          "author_url": "",
          "post_date": "08/15/2022 13:35:52",
          "content": "<p>The second set is just predictions from your model.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1898694,
      "author_name": "sakvaua",
      "author_url": "",
      "post_date": "08/14/2022 18:29:42",
      "content": "<p>Labeling is very inconsistent. For large intestine and prostate some images have FPUs on the border segmented, while in others they are not. Oh, and don't get me started on lungs. It's a mess. Probably a result of some automatic segmenting. Kidneys are ok.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1898704,
          "author_name": "urosjarc",
          "author_url": "",
          "post_date": "08/14/2022 18:39:47",
          "content": "<p>I'm wondering what special measures to take to tackle the lung problem?</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1898708,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "08/14/2022 18:44:47",
      "content": "<p>if you want, you can have 3 class : background, FTU-lablled, FTU-not-labelled</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1901484,
      "author_name": "sohier",
      "author_url": "",
      "post_date": "08/16/2022 17:18:58",
      "content": "<p>A few quick thoughts:</p>\n<ul>\n<li>It's quite difficult to get labels for this kind of data. Manual work with segmentation masks is time consuming and in this case requires medical expertise. I can assure you that a great deal of effort went into labeling this dataset.</li>\n<li>It's also quite common for medical challenges to have limits on the quality of the labels. For example, I've worked on multiple medical classification problems where we knew the experts only agreed on the label about 80% of the time. Those competitions were still quite successful for both the hosts and competitors. </li>\n</ul>\n<p>To me, there are two questions to focus on when discussing label quality in the context of a competition:</p>\n<ul>\n<li>Is there a systematic error? For example, if the lower left quarter of each image was never labeled I would be concerned. That doesn't appear to be the case here.</li>\n<li>Is the label quality the limiting factor for model performance? We have run competitions in the past where there wasn't much signal and it was <em>all</em> extracted. In those cases the public and private leaderboard scores cluster closely around some less-than-perfect value. Scores in this competition have a wide range and continue to improve, which indicates to me that the models aren't currently limited by the label quality.</li>\n</ul>\n<p>If you're still concerned about the label quality I would recommend making an estimate of how much any label shortcomings would reduce the score of a truly perfect set of masks. If we're still a long way from that, then you don't have to worry that label quality is defining the competition.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1901576,
          "author_name": "urosjarc",
          "author_url": "",
          "post_date": "08/16/2022 19:01:08",
          "content": "<p>I have no problem with the dataset being incomplete, but I have problems when it comes to dataset transparency from the host. I get the point that this is the competition and not every detail around the private dataset should be revealed. But I would expect some level of transparency about…</p>\n<ul>\n<li>Who or what created labels?</li>\n<li>What is the expected accuracy of labels in the private and public datasets?</li>\n<li>Distribution of image sizes in the private dataset.</li>\n<li>Label class distribution on the private dataset. (Is the dataset balanced)</li>\n<li>Why are there no images in the public dataset 160x160 size but there are such images in the private dataset?</li>\n<li>etc…</li>\n</ul>\n<p>I would expect from you Kaggle Staff to ask the host before he creates a competition to provide such information so that we can really focus on providing the best-generalized models and not to create hacks around datasets. With such information, we could provide better models without revealing to much information about private datasets.</p>\n<p>I love Kaggle community and the fact that there is a sincere wish in the kaggle community to create the best of the best results and to provide good service to scientific world. But, hiding information, and not being transparent is something that makes me question the whole competition and their motives…</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1901609,
          "author_name": "sohier",
          "author_url": "",
          "post_date": "08/16/2022 19:23:52",
          "content": "<p><a href=\"https://www.kaggle.com/urosjarc\" target=\"_blank\">@urosjarc</a> our standard practice is always to disclose as little as possible about the private dataset, for obvious reasons. Frankly, the data description for this competition goes into an unusual level of detail about the hidden test. Deciding what to disclose is a joint decision between Kaggle and the host and I'm very comfortable with the status quo.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1902914,
          "author_name": "urosjarc",
          "author_url": "",
          "post_date": "08/17/2022 02:03:28",
          "content": "<p><a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a> I don't think that answering such questions should compromise the competition in any way… This is my first competition so I don't know how things work here, but isn't it in the host's best interest to provide the community with as much information as possible he can to get the best models and to still make the competition valid?</p>\n<blockquote>\n  <p>our standard practice is always to disclose as little as possible about the private dataset</p>\n</blockquote>\n<p>Yea I get it it's competition something should be hidden if this is compeitition,  but still, I failed to realize to see how answering such questions will degrade competition validity.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1903359,
      "author_name": "auriml",
      "author_url": "",
      "post_date": "08/17/2022 10:59:30",
      "content": "<p><a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a> From the clinical point of view, could you please define/list which of the following structures in the lung are the target?</p>\n<ul>\n<li>All observable alveoli (including alveolar septa and enclosed air)</li>\n<li>All observable alveolar ducts</li>\n<li>All observable terminal bronchiole </li>\n</ul>\n<p>The FTU of the lung comprises all of the above within a lung lobule  (please refer to Pubmed source below and image from <a href=\"url\" target=\"_blank\">DOI:10.1136/jclinpath-2014-202685</a> ). <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F865377%2F31fd895f35c25c62fe6dfeba693de510%2Flung_FTU.png?generation=1660733618800459&amp;alt=media\" alt=\"\"></p>\n<p>As it can be observed, the anatomy of the lungs, where each lung lobule has a complex reticular structure with interconnected alveoli (as those have pores between them), makes it unfeasible to  attempt to separate alveoli into isolated units with separate masks. <br>\nI would appreciate clarification from the clinical staff involved on the labeling of this competition.</p>\n<p>Thanks in advance, </p>\n<p>Aurelia Bustos, MD, PhD<br>\nMedical Oncologist</p>",
      "votes": null,
      "replies": [
        {
          "id": 1903901,
          "author_name": "yashvrdnjain",
          "author_url": "",
          "post_date": "08/17/2022 18:28:28",
          "content": "<p>The experts annotated whole, uncollapsed alveolar spaces, focusing on every individual instance that can be clearly recognized as an alveolus. </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1898600": "yashvrdnjain \nAt least in the images of lung tissue, it seems that not all aveoli (functional units) are being labeled with masks. I have asked an experienced pathologist and she has confirmed that for example on the image below the region highlighted with an arrow is also another alveolus but has not been masked.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F865377%2Fb339f98c78c778fbbb320631403e16e3%2FScreen%20Shot%202022-08-14%20at%2003.34.51.png?generation=1660496680735529&alt=media)\n\nThanks in advance for your answer,\nBR,",
    "1898665": "There is unmasked labels all over dataset, the dataset is most definitely masked with their algorithm. Some most notable examples...\n\n* RED: Mask\n* BLUE: Predictions\n* VIOLET: Mask && Predictions\n\n<img width=\"300\"  src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2386017%2F012fa733e2aa324f672ab44505d0f9f4%2Ff7c24a6a-bee1-40b5-9c6b-db6928f2feee.png?generation=1660499929307393&alt=media\"></img>\n\n<img width=\"300\"  src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2386017%2F889e8d9f8e2b0624d22833fcfed704d0%2Fmedia_images_on_epoch_end_%20validate_761_15af07af5e2cbca49b9a.png?generation=1660499941166791&alt=media\"></img>",
    "1898686": "Here are some more examples....\nIf you train your model, there should be a generalization of model detection even if the label is missing. My model overcomes this problem quite nicely and I don't think there should be a problem if the dataset is the missing label. However, there will be problems with lung detection, my model decided to joust skip mask detection in the lung section. 😄\n\nI'm really curious how grandmasters decided to tackle this problem with lungs...\n\n<img width=\"300\" src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2386017%2F4db024bf1463a420e5ca76ee486e3f16%2Fc57d55c4-6779-41b4-97fb-ef6769f71b79.png?generation=1660500985178077&alt=media\"></img>\n\n<img width=\"300\" src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2386017%2F6bb17657625e85232355eb2fe6fa6aa7%2F20be363b-c67d-4a58-8130-9fdb1f2aec58.png?generation=1660500974523051&alt=media\"></img>",
    "1898694": "Labeling is very inconsistent. For large intestine and prostate some images have FPUs on the border segmented, while in others they are not. Oh, and don't get me started on lungs. It's a mess. Probably a result of some automatic segmenting. Kidneys are ok.",
    "1898704": "I'm wondering what special measures to take to tackle the lung problem?",
    "1898708": "if you want, you can have 3 class : background, FTU-lablled, FTU-not-labelled",
    "1898915": "if you take a look at edges, its most likely manually anotated, but without a strict rule set for labels.\nI believe, that for the alveoli only nice examples have been annotated. Which leads to many not annotated examples.\nThe big question is: do we want to learn the task or the annotator?",
    "1899041": "\" do we want to learn the task or the annotator?\"\n\nfrom application  point of view (judges prize) : learn the task\n\nfrom performance ranking competition point of view (supervised ML prize) : learn the annotator\n\n---\n\nkaggle needs to develop their own code for data cleaning or organize a data cleaning up competition for us to help",
    "1899601": "Labeling can be an iterative process as well because there will always be errors in annotation tasks.\nBy iteratively training models on k folds and predicting on the validation data, the worst errors could be found and corrected easily.\n\nMaybe it could also be a possibility, to actually let the users suggest label changes and let the hosts review them.",
    "1899684": "Hello, how dos you get the second set of annotation? I was following the RLE in `train.csv`",
    "1899753": "The second set is just predictions from your model.",
    "1901484": "A few quick thoughts:\n\n- It's quite difficult to get labels for this kind of data. Manual work with segmentation masks is time consuming and in this case requires medical expertise. I can assure you that a great deal of effort went into labeling this dataset.\n- It's also quite common for medical challenges to have limits on the quality of the labels. For example, I've worked on multiple medical classification problems where we knew the experts only agreed on the label about 80% of the time. Those competitions were still quite successful for both the hosts and competitors. \n\nTo me, there are two questions to focus on when discussing label quality in the context of a competition:\n- Is there a systematic error? For example, if the lower left quarter of each image was never labeled I would be concerned. That doesn't appear to be the case here.\n- Is the label quality the limiting factor for model performance? We have run competitions in the past where there wasn't much signal and it was _all_ extracted. In those cases the public and private leaderboard scores cluster closely around some less-than-perfect value. Scores in this competition have a wide range and continue to improve, which indicates to me that the models aren't currently limited by the label quality.\n\nIf you're still concerned about the label quality I would recommend making an estimate of how much any label shortcomings would reduce the score of a truly perfect set of masks. If we're still a long way from that, then you don't have to worry that label quality is defining the competition.",
    "1901576": "I have no problem with the dataset being incomplete, but I have problems when it comes to dataset transparency from the host. I get the point that this is the competition and not every detail around the private dataset should be revealed. But I would expect some level of transparency about...\n\n- Who or what created labels?\n- What is the expected accuracy of labels in the private and public datasets?\n- Distribution of image sizes in the private dataset.\n- Label class distribution on the private dataset. (Is the dataset balanced)\n- Why are there no images in the public dataset 160x160 size but there are such images in the private dataset?\n- etc...\n\nI would expect from you Kaggle Staff to ask the host before he creates a competition to provide such information so that we can really focus on providing the best-generalized models and not to create hacks around datasets. With such information, we could provide better models without revealing to much information about private datasets.\n\nI love Kaggle community and the fact that there is a sincere wish in the kaggle community to create the best of the best results and to provide good service to scientific world. But, hiding information, and not being transparent is something that makes me question the whole competition and their motives...",
    "1901609": "urosjarc our standard practice is always to disclose as little as possible about the private dataset, for obvious reasons. Frankly, the data description for this competition goes into an unusual level of detail about the hidden test. Deciding what to disclose is a joint decision between Kaggle and the host and I'm very comfortable with the status quo.",
    "1902914": "sohier I don't think that answering such questions should compromise the competition in any way... This is my first competition so I don't know how things work here, but isn't it in the host's best interest to provide the community with as much information as possible he can to get the best models and to still make the competition valid?\n\n> our standard practice is always to disclose as little as possible about the private dataset\n\nYea I get it it's competition something should be hidden if this is compeitition,  but still, I failed to realize to see how answering such questions will degrade competition validity.",
    "1903359": "sohier From the clinical point of view, could you please define/list which of the following structures in the lung are the target?\n- All observable alveoli (including alveolar septa and enclosed air)\n- All observable alveolar ducts\n- All observable terminal bronchiole \n\nThe FTU of the lung comprises all of the above within a lung lobule  (please refer to Pubmed source below and image from [DOI:10.1136/jclinpath-2014-202685](url) ). ![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F865377%2F31fd895f35c25c62fe6dfeba693de510%2Flung_FTU.png?generation=1660733618800459&alt=media)\n\n\nAs it can be observed, the anatomy of the lungs, where each lung lobule has a complex reticular structure with interconnected alveoli (as those have pores between them), makes it unfeasible to  attempt to separate alveoli into isolated units with separate masks. \nI would appreciate clarification from the clinical staff involved on the labeling of this competition.\n\nThanks in advance, \n\nAurelia Bustos, MD, PhD\nMedical Oncologist",
    "1903901": "The experts annotated whole, uncollapsed alveolar spaces, focusing on every individual instance that can be clearly recognized as an alveolus."
  },
  "source": "meta"
}