{
  "id": 529364,
  "title": "Handling Missing Values for Subarticular Stenosis : Optimal Strategy?",
  "url": "/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/529364",
  "author_name": "",
  "post_date": "2024-08-20T11:05:53.015484Z",
  "votes": 3,
  "comment_count": 2,
  "views": 0,
  "content": "<p>I noticed that several participants have likely plotted the class distribution for the given dataset and observed that the labels for subarticular stenosis are frequently missing for the L1/L2 and L2/L3 regions. I'm curious about the strategies being employed to address this issue. Specifically:</p>\n<ul>\n<li>Are you excluding the affected study IDs from your analysis entirely?</li>\n<li>Alternatively, are you applying any imputation techniques to handle these missing values?<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5048844%2F044d9d4ab872cbbf2cd22cdd01cf24d4%2FScreenshot%202024-08-20%20163418.png?generation=1724151878359226&amp;alt=media\" alt=\"\"></li>\n</ul>\n<p>I came across an <a href=\"https://www.kaggle.com/code/hugowjd/rsna2024-lsdc-training-densenet#RSNA2024-LSDC-Training-DenseNet\" target=\"_blank\">approach</a> where the authors <strong>replace missing values with -100</strong>, stating that this allows their functions to ignore these entries during the calculation of loss and evaluation metrics. <br>\nMy question is: Is this method theoretically sound, and does setting the missing values to -100 effectively handle the problem without introducing bias or affecting model performance?</p>\n<p>I’d appreciate any insights or experiences on the best practices for dealing with these missing values in the context of this specific problem.</p>",
  "messages": [
    {
      "id": "2964914",
      "postDate": "08/20/2024 11:05:53",
      "content": "<p>I noticed that several participants have likely plotted the class distribution for the given dataset and observed that the labels for subarticular stenosis are frequently missing for the L1/L2 and L2/L3 regions. I'm curious about the strategies being employed to address this issue. Specifically:</p>\n<ul>\n<li>Are you excluding the affected study IDs from your analysis entirely?</li>\n<li>Alternatively, are you applying any imputation techniques to handle these missing values?<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5048844%2F044d9d4ab872cbbf2cd22cdd01cf24d4%2FScreenshot%202024-08-20%20163418.png?generation=1724151878359226&amp;alt=media\" alt=\"\"></li>\n</ul>\n<p>I came across an <a href=\"https://www.kaggle.com/code/hugowjd/rsna2024-lsdc-training-densenet#RSNA2024-LSDC-Training-DenseNet\" target=\"_blank\">approach</a> where the authors <strong>replace missing values with -100</strong>, stating that this allows their functions to ignore these entries during the calculation of loss and evaluation metrics. <br>\nMy question is: Is this method theoretically sound, and does setting the missing values to -100 effectively handle the problem without introducing bias or affecting model performance?</p>\n<p>I’d appreciate any insights or experiences on the best practices for dealing with these missing values in the context of this specific problem.</p>",
      "rawMarkdown": "I noticed that several participants have likely plotted the class distribution for the given dataset and observed that the labels for subarticular stenosis are frequently missing for the L1/L2 and L2/L3 regions. I'm curious about the strategies being employed to address this issue. Specifically:\n\n- Are you excluding the affected study IDs from your analysis entirely?\n- Alternatively, are you applying any imputation techniques to handle these missing values?\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5048844%2F044d9d4ab872cbbf2cd22cdd01cf24d4%2FScreenshot%202024-08-20%20163418.png?generation=1724151878359226&alt=media)\n\nI came across an [approach](https://www.kaggle.com/code/hugowjd/rsna2024-lsdc-training-densenet#RSNA2024-LSDC-Training-DenseNet) where the authors **replace missing values with -100**, stating that this allows their functions to ignore these entries during the calculation of loss and evaluation metrics. \nMy question is: Is this method theoretically sound, and does setting the missing values to -100 effectively handle the problem without introducing bias or affecting model performance?\n\nI’d appreciate any insights or experiences on the best practices for dealing with these missing values in the context of this specific problem.",
      "votes": null
    },
    {
      "id": "2964928",
      "postDate": "08/20/2024 11:27:36",
      "content": "<p>I think I remember I've excluded them, but may be would been better solution mask them at loss calculation.</p>\n<p>I mean, imput -100 when you can simply Loss(y_pred[mask],y_true[mask]). It just takes a bit of extra memmory.</p>\n<p>EDIT: Well though, you don't need to store the mask at memmory, just generate it with nans in y_true.</p>",
      "rawMarkdown": "I think I remember I've excluded them, but may be would been better solution mask them at loss calculation.\n\nI mean, imput -100 when you can simply Loss(y_pred[mask],y_true[mask]). It just takes a bit of extra memmory.\n\nEDIT: Well though, you don't need to store the mask at memmory, just generate it with nans in y_true.",
      "votes": null
    },
    {
      "id": "2964984",
      "postDate": "08/20/2024 12:58:24",
      "content": "<p><a href=\"https://www.kaggle.com/rahulnakka\" target=\"_blank\">@rahulnakka</a>, -100 solves the problem, consider it as a magic number for the torch cross entropy loss. By default -100 \"specifies a target value that is ignored and does not contribute to the input gradient\", so you are safe to use -100.  <a href=\"https://pytorch.org/docs/stable/generated/torch.nn.CrossEntropyLoss.html#torch.nn.CrossEntropyLoss:~:text=%3DNone%2C-,ignore_index%3D%2D100,-%2C%20reduce%3D\" target=\"_blank\">link</a></p>",
      "rawMarkdown": "rahulnakka, -100 solves the problem, consider it as a magic number for the torch cross entropy loss. By default -100 \"specifies a target value that is ignored and does not contribute to the input gradient\", so you are safe to use -100.  [link](https://pytorch.org/docs/stable/generated/torch.nn.CrossEntropyLoss.html#torch.nn.CrossEntropyLoss:~:text=%3DNone%2C-,ignore_index%3D%2D100,-%2C%20reduce%3D)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2964928,
      "author_name": "sacuscreed",
      "author_url": "",
      "post_date": "08/20/2024 11:27:36",
      "content": "<p>I think I remember I've excluded them, but may be would been better solution mask them at loss calculation.</p>\n<p>I mean, imput -100 when you can simply Loss(y_pred[mask],y_true[mask]). It just takes a bit of extra memmory.</p>\n<p>EDIT: Well though, you don't need to store the mask at memmory, just generate it with nans in y_true.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2964984,
      "author_name": "sergiosaharovskiy",
      "author_url": "",
      "post_date": "08/20/2024 12:58:24",
      "content": "<p><a href=\"https://www.kaggle.com/rahulnakka\" target=\"_blank\">@rahulnakka</a>, -100 solves the problem, consider it as a magic number for the torch cross entropy loss. By default -100 \"specifies a target value that is ignored and does not contribute to the input gradient\", so you are safe to use -100.  <a href=\"https://pytorch.org/docs/stable/generated/torch.nn.CrossEntropyLoss.html#torch.nn.CrossEntropyLoss:~:text=%3DNone%2C-,ignore_index%3D%2D100,-%2C%20reduce%3D\" target=\"_blank\">link</a></p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2964914": "I noticed that several participants have likely plotted the class distribution for the given dataset and observed that the labels for subarticular stenosis are frequently missing for the L1/L2 and L2/L3 regions. I'm curious about the strategies being employed to address this issue. Specifically:\n\n- Are you excluding the affected study IDs from your analysis entirely?\n- Alternatively, are you applying any imputation techniques to handle these missing values?\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5048844%2F044d9d4ab872cbbf2cd22cdd01cf24d4%2FScreenshot%202024-08-20%20163418.png?generation=1724151878359226&alt=media)\n\nI came across an [approach](https://www.kaggle.com/code/hugowjd/rsna2024-lsdc-training-densenet#RSNA2024-LSDC-Training-DenseNet) where the authors **replace missing values with -100**, stating that this allows their functions to ignore these entries during the calculation of loss and evaluation metrics. \nMy question is: Is this method theoretically sound, and does setting the missing values to -100 effectively handle the problem without introducing bias or affecting model performance?\n\nI’d appreciate any insights or experiences on the best practices for dealing with these missing values in the context of this specific problem.",
    "2964928": "I think I remember I've excluded them, but may be would been better solution mask them at loss calculation.\n\nI mean, imput -100 when you can simply Loss(y_pred[mask],y_true[mask]). It just takes a bit of extra memmory.\n\nEDIT: Well though, you don't need to store the mask at memmory, just generate it with nans in y_true.",
    "2964984": "rahulnakka, -100 solves the problem, consider it as a magic number for the torch cross entropy loss. By default -100 \"specifies a target value that is ignored and does not contribute to the input gradient\", so you are safe to use -100.  [link](https://pytorch.org/docs/stable/generated/torch.nn.CrossEntropyLoss.html#torch.nn.CrossEntropyLoss:~:text=%3DNone%2C-,ignore_index%3D%2D100,-%2C%20reduce%3D)"
  },
  "source": "meta"
}