{
  "id": 415521,
  "title": "What would be the best idea to handle the \"unsure\" category?",
  "url": "/competitions/hubmap-hacking-the-human-vasculature/discussion/415521",
  "author_name": "",
  "post_date": "2023-06-06T19:57:34.317638900Z",
  "votes": 3,
  "comment_count": 4,
  "views": 0,
  "content": "<p>As the question states above, what do you guys think is the best way to handle the \"unsure\" class of the annotations? Would it be too wasteful to not include it all during training or would it be too risky to include it since the experts are not sure if they are blood vessels or not? </p>\n<p>After filtering for Dataset 1 only (expert reviewed annotations), if I did not do it incorrectly, I have a total of 4406 individual annotations (note that this is different from # of tiles, since there are multiple # of annotations per tile), of which 3498 are blood vessels, only 36 are glomerulus and 872 are unsure. So we have a quite a bit of the unsure annotations. </p>\n<p>Any input would be appreciated!</p>",
  "messages": [
    {
      "id": "2290452",
      "postDate": "06/06/2023 19:57:34",
      "content": "<p>As the question states above, what do you guys think is the best way to handle the \"unsure\" class of the annotations? Would it be too wasteful to not include it all during training or would it be too risky to include it since the experts are not sure if they are blood vessels or not? </p>\n<p>After filtering for Dataset 1 only (expert reviewed annotations), if I did not do it incorrectly, I have a total of 4406 individual annotations (note that this is different from # of tiles, since there are multiple # of annotations per tile), of which 3498 are blood vessels, only 36 are glomerulus and 872 are unsure. So we have a quite a bit of the unsure annotations. </p>\n<p>Any input would be appreciated!</p>",
      "rawMarkdown": "As the question states above, what do you guys think is the best way to handle the \"unsure\" class of the annotations? Would it be too wasteful to not include it all during training or would it be too risky to include it since the experts are not sure if they are blood vessels or not? \n\nAfter filtering for Dataset 1 only (expert reviewed annotations), if I did not do it incorrectly, I have a total of 4406 individual annotations (note that this is different from # of tiles, since there are multiple # of annotations per tile), of which 3498 are blood vessels, only 36 are glomerulus and 872 are unsure. So we have a quite a bit of the unsure annotations. \n\nAny input would be appreciated!",
      "votes": null
    },
    {
      "id": "2290465",
      "postDate": "06/06/2023 20:13:10",
      "content": "<p>I'm not sure if this is the best approach, but for me, I think it would be wise to just multiply the pixel-wise categorical cross-entropy loss by the inverse 'unsure' mask to zero out the loss resulting from these areas. This essentially means that the loss function won't care what class the model will treat them as (glomerulus vs. blood vessel vs. background).</p>",
      "rawMarkdown": "I'm not sure if this is the best approach, but for me, I think it would be wise to just multiply the pixel-wise categorical cross-entropy loss by the inverse 'unsure' mask to zero out the loss resulting from these areas. This essentially means that the loss function won't care what class the model will treat them as (glomerulus vs. blood vessel vs. background).",
      "votes": null
    },
    {
      "id": "2300733",
      "postDate": "06/13/2023 11:55:44",
      "content": "<p>I went through the same process as well. Not sure exactly how I am going to deal with this type of class. Mohamed approach seems reasonable though!</p>",
      "rawMarkdown": "I went through the same process as well. Not sure exactly how I am going to deal with this type of class. Mohamed approach seems reasonable though!",
      "votes": null
    },
    {
      "id": "2345068",
      "postDate": "07/15/2023 03:54:08",
      "content": "<p>I was thinking about making class weights zero for unsure class during training.</p>",
      "rawMarkdown": "I was thinking about making class weights zero for unsure class during training.",
      "votes": null
    },
    {
      "id": "2357147",
      "postDate": "07/24/2023 15:11:54",
      "content": "<p>I treat unsure and glomerulus as not contributing to loss for now, with using torch.nn.CrossEntropyLoss(ignore_index=#x#).<br>\nI think it might also be good to treat it as a soft-label (assigning 0.5 to it maybe).</p>",
      "rawMarkdown": "I treat unsure and glomerulus as not contributing to loss for now, with using torch.nn.CrossEntropyLoss(ignore_index=#x#).\nI think it might also be good to treat it as a soft-label (assigning 0.5 to it maybe).",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2290465,
      "author_name": "momakmd",
      "author_url": "",
      "post_date": "06/06/2023 20:13:10",
      "content": "<p>I'm not sure if this is the best approach, but for me, I think it would be wise to just multiply the pixel-wise categorical cross-entropy loss by the inverse 'unsure' mask to zero out the loss resulting from these areas. This essentially means that the loss function won't care what class the model will treat them as (glomerulus vs. blood vessel vs. background).</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2300733,
      "author_name": "ignasialemany",
      "author_url": "",
      "post_date": "06/13/2023 11:55:44",
      "content": "<p>I went through the same process as well. Not sure exactly how I am going to deal with this type of class. Mohamed approach seems reasonable though!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2345068,
      "author_name": "aniarya",
      "author_url": "",
      "post_date": "07/15/2023 03:54:08",
      "content": "<p>I was thinking about making class weights zero for unsure class during training.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2357147,
      "author_name": "raki21",
      "author_url": "",
      "post_date": "07/24/2023 15:11:54",
      "content": "<p>I treat unsure and glomerulus as not contributing to loss for now, with using torch.nn.CrossEntropyLoss(ignore_index=#x#).<br>\nI think it might also be good to treat it as a soft-label (assigning 0.5 to it maybe).</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2290452": "As the question states above, what do you guys think is the best way to handle the \"unsure\" class of the annotations? Would it be too wasteful to not include it all during training or would it be too risky to include it since the experts are not sure if they are blood vessels or not? \n\nAfter filtering for Dataset 1 only (expert reviewed annotations), if I did not do it incorrectly, I have a total of 4406 individual annotations (note that this is different from # of tiles, since there are multiple # of annotations per tile), of which 3498 are blood vessels, only 36 are glomerulus and 872 are unsure. So we have a quite a bit of the unsure annotations. \n\nAny input would be appreciated!",
    "2290465": "I'm not sure if this is the best approach, but for me, I think it would be wise to just multiply the pixel-wise categorical cross-entropy loss by the inverse 'unsure' mask to zero out the loss resulting from these areas. This essentially means that the loss function won't care what class the model will treat them as (glomerulus vs. blood vessel vs. background).",
    "2300733": "I went through the same process as well. Not sure exactly how I am going to deal with this type of class. Mohamed approach seems reasonable though!",
    "2345068": "I was thinking about making class weights zero for unsure class during training.",
    "2357147": "I treat unsure and glomerulus as not contributing to loss for now, with using torch.nn.CrossEntropyLoss(ignore_index=#x#).\nI think it might also be good to treat it as a soft-label (assigning 0.5 to it maybe)."
  },
  "source": "meta"
}