{
  "id": 456763,
  "title": "Assistance Needed: Unraveling the Mystery Behind a 0.0 Score",
  "url": "/competitions/blood-vessel-segmentation/discussion/456763",
  "author_name": "",
  "post_date": "2023-11-21T15:24:01.119331500Z",
  "votes": 4,
  "comment_count": 1,
  "views": 0,
  "content": "<p>Hi everyone,</p>\n<p>I'm having a problem with my model for the segmentation task, and it looks like I'm not the only one. My model keeps getting a score of 0.0, and I can't figure out why. I've been following the usual steps, but something's not working.</p>\n<p>Here's what I've tried so far:</p>\n<p>I removed the tiny parts from my model's results.<br>\nI changed the threshold value, which is how I decide if a part of the image is something or not.<br>\nBut even after doing these, my score is still 0.0. I'm looking for advice from anyone who got their model to work:</p>\n<p>Are there any common mistakes that could make my score so low?<br>\nIs there something important I might be missing about how to process the results or set the threshold?<br>\nIf you had this problem and fixed it, what did you do?<br>\nI'd really appreciate any tips or stories about how you fixed a similar problem. Thanks a lot for your help!</p>",
  "messages": [
    {
      "id": "2533088",
      "postDate": "11/21/2023 15:24:01",
      "content": "<p>Hi everyone,</p>\n<p>I'm having a problem with my model for the segmentation task, and it looks like I'm not the only one. My model keeps getting a score of 0.0, and I can't figure out why. I've been following the usual steps, but something's not working.</p>\n<p>Here's what I've tried so far:</p>\n<p>I removed the tiny parts from my model's results.<br>\nI changed the threshold value, which is how I decide if a part of the image is something or not.<br>\nBut even after doing these, my score is still 0.0. I'm looking for advice from anyone who got their model to work:</p>\n<p>Are there any common mistakes that could make my score so low?<br>\nIs there something important I might be missing about how to process the results or set the threshold?<br>\nIf you had this problem and fixed it, what did you do?<br>\nI'd really appreciate any tips or stories about how you fixed a similar problem. Thanks a lot for your help!</p>",
      "rawMarkdown": "Hi everyone,\n\nI'm having a problem with my model for the segmentation task, and it looks like I'm not the only one. My model keeps getting a score of 0.0, and I can't figure out why. I've been following the usual steps, but something's not working.\n\nHere's what I've tried so far:\n\nI removed the tiny parts from my model's results.\nI changed the threshold value, which is how I decide if a part of the image is something or not.\nBut even after doing these, my score is still 0.0. I'm looking for advice from anyone who got their model to work:\n\nAre there any common mistakes that could make my score so low?\nIs there something important I might be missing about how to process the results or set the threshold?\nIf you had this problem and fixed it, what did you do?\nI'd really appreciate any tips or stories about how you fixed a similar problem. Thanks a lot for your help!",
      "votes": null
    },
    {
      "id": "2533313",
      "postDate": "11/21/2023 19:03:35",
      "content": "<p>Some general debugging advice that helps me at work and has already helped in this competition: always visualize your inputs and outputs.</p>\n<ol>\n<li>Try visualizing the input data right before model ingestion. Does it look like what you'd expect given your preprocessing chain? Does it match with what the model was trained on (i.e., norm stats didn't get changed by mistake)? Any deltas in preprocessing between train and inference could lead to 0 scores.</li>\n<li>Try visualizing the predicted masks on the training dataset, then your validation dataset. Odds are, your model converged to something, implying that it learned the training data. If you push the training data back through the model, you should see really good performance. If you don't, then it could be related to point 1 (preprocessing differences) or the another bug. If the training data looks good, visualize the predicted masks on your validation dataset. They shouldn't look perfect, but they should be logical.</li>\n<li>Make sure your RLE encode method accounts for empty masks.</li>\n<li>Make sure you're loading your trained model when inferencing. Load the checkpoint once and print the first layer weights. Repeat. If the weights are different, your model is not getting loaded correctly.</li>\n</ol>\n<p>If you verify all of the above and still have issues, it could be related to the competition scoring metric. The hosts have been debugging it since launch and there are dozens of threads on the topic.</p>\n<p>Some of the points above are trivial but not meant to be condescending in any way- I've made these mistakes dozens of times and I've seen industry veterans do the same. Just try to debug your workflow sequentially, one step at a time. Good luck! o7</p>",
      "rawMarkdown": "Some general debugging advice that helps me at work and has already helped in this competition: always visualize your inputs and outputs.\n\n1. Try visualizing the input data right before model ingestion. Does it look like what you'd expect given your preprocessing chain? Does it match with what the model was trained on (i.e., norm stats didn't get changed by mistake)? Any deltas in preprocessing between train and inference could lead to 0 scores.\n2. Try visualizing the predicted masks on the training dataset, then your validation dataset. Odds are, your model converged to something, implying that it learned the training data. If you push the training data back through the model, you should see really good performance. If you don't, then it could be related to point 1 (preprocessing differences) or the another bug. If the training data looks good, visualize the predicted masks on your validation dataset. They shouldn't look perfect, but they should be logical.\n3. Make sure your RLE encode method accounts for empty masks.\n4. Make sure you're loading your trained model when inferencing. Load the checkpoint once and print the first layer weights. Repeat. If the weights are different, your model is not getting loaded correctly.\n\nIf you verify all of the above and still have issues, it could be related to the competition scoring metric. The hosts have been debugging it since launch and there are dozens of threads on the topic.\n\nSome of the points above are trivial but not meant to be condescending in any way- I've made these mistakes dozens of times and I've seen industry veterans do the same. Just try to debug your workflow sequentially, one step at a time. Good luck! o7",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2533313,
      "author_name": "squidinator",
      "author_url": "",
      "post_date": "11/21/2023 19:03:35",
      "content": "<p>Some general debugging advice that helps me at work and has already helped in this competition: always visualize your inputs and outputs.</p>\n<ol>\n<li>Try visualizing the input data right before model ingestion. Does it look like what you'd expect given your preprocessing chain? Does it match with what the model was trained on (i.e., norm stats didn't get changed by mistake)? Any deltas in preprocessing between train and inference could lead to 0 scores.</li>\n<li>Try visualizing the predicted masks on the training dataset, then your validation dataset. Odds are, your model converged to something, implying that it learned the training data. If you push the training data back through the model, you should see really good performance. If you don't, then it could be related to point 1 (preprocessing differences) or the another bug. If the training data looks good, visualize the predicted masks on your validation dataset. They shouldn't look perfect, but they should be logical.</li>\n<li>Make sure your RLE encode method accounts for empty masks.</li>\n<li>Make sure you're loading your trained model when inferencing. Load the checkpoint once and print the first layer weights. Repeat. If the weights are different, your model is not getting loaded correctly.</li>\n</ol>\n<p>If you verify all of the above and still have issues, it could be related to the competition scoring metric. The hosts have been debugging it since launch and there are dozens of threads on the topic.</p>\n<p>Some of the points above are trivial but not meant to be condescending in any way- I've made these mistakes dozens of times and I've seen industry veterans do the same. Just try to debug your workflow sequentially, one step at a time. Good luck! o7</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2533088": "Hi everyone,\n\nI'm having a problem with my model for the segmentation task, and it looks like I'm not the only one. My model keeps getting a score of 0.0, and I can't figure out why. I've been following the usual steps, but something's not working.\n\nHere's what I've tried so far:\n\nI removed the tiny parts from my model's results.\nI changed the threshold value, which is how I decide if a part of the image is something or not.\nBut even after doing these, my score is still 0.0. I'm looking for advice from anyone who got their model to work:\n\nAre there any common mistakes that could make my score so low?\nIs there something important I might be missing about how to process the results or set the threshold?\nIf you had this problem and fixed it, what did you do?\nI'd really appreciate any tips or stories about how you fixed a similar problem. Thanks a lot for your help!",
    "2533313": "Some general debugging advice that helps me at work and has already helped in this competition: always visualize your inputs and outputs.\n\n1. Try visualizing the input data right before model ingestion. Does it look like what you'd expect given your preprocessing chain? Does it match with what the model was trained on (i.e., norm stats didn't get changed by mistake)? Any deltas in preprocessing between train and inference could lead to 0 scores.\n2. Try visualizing the predicted masks on the training dataset, then your validation dataset. Odds are, your model converged to something, implying that it learned the training data. If you push the training data back through the model, you should see really good performance. If you don't, then it could be related to point 1 (preprocessing differences) or the another bug. If the training data looks good, visualize the predicted masks on your validation dataset. They shouldn't look perfect, but they should be logical.\n3. Make sure your RLE encode method accounts for empty masks.\n4. Make sure you're loading your trained model when inferencing. Load the checkpoint once and print the first layer weights. Repeat. If the weights are different, your model is not getting loaded correctly.\n\nIf you verify all of the above and still have issues, it could be related to the competition scoring metric. The hosts have been debugging it since launch and there are dozens of threads on the topic.\n\nSome of the points above are trivial but not meant to be condescending in any way- I've made these mistakes dozens of times and I've seen industry veterans do the same. Just try to debug your workflow sequentially, one step at a time. Good luck! o7"
  },
  "source": "meta"
}