{
  "id": 552631,
  "title": "[Solved] 2D UNet having CV=0, Is this expected? How should I debug?",
  "url": "/competitions/czii-cryo-et-object-identification/discussion/552631",
  "author_name": "",
  "post_date": "2024-12-20T16:52:36.005833500Z",
  "votes": 1,
  "comment_count": 2,
  "views": 0,
  "content": "<p>Hello everyone!</p>\n<p>I have been trying to make 2D UNet work for a while but it seems like I am failing miserably in this(as I am getting a score of 0). My idea is the following:</p>\n<ol>\n<li>create segmentation mask using copick</li>\n<li>create 2D dataset of the tomograph along the depth dimension.</li>\n<li>use UNet model and cross-entropy loss to train the model</li>\n<li>During inference we will slice the tomograph and then predict it's corresponding heatmap and then stack all the heatmaps together to find the centeroids.</li>\n<li>Also while creating the 2D slices I include neighbouring depth images i.e +- 1 depth</li>\n<li>While training I try to maximize the <code>dice_score</code> which is defined as follows (I am not sure whether this is the correct metric that I should maximize for):</li>\n</ol>\n<pre><code> ():\n    \n\n    \n    \n    bg_preds = ~((ytrue == ) &amp; (ypred == ))\n\n    ytrue = ytrue[bg_preds]\n    ypred = ypred[bg_preds]\n\n     torch.count_nonzero(ytrue == ypred) / (ytrue)\n</code></pre>\n<p>With <code>efficientnet-b7</code> as the backbone I am able to achieve a \"dice_score\" of <code>0.2597</code> on the validation set. I interpret this score as \"Approximately 1 in 4 pixels are correctly predicted by the model\". Though when plotting the ce loss and the dice score of the validation set I get the following graph (which I thought was quite interesting, though I am not sure how to interpret it):</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4362752%2F2bdb10bcc9005972dd24375bedd4fd73%2FScreenshot%20from%202024-12-20%2022-12-37.png?generation=1734712969168320&amp;alt=media\" alt=\"\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4362752%2Fa491de5efddb60fcf7f29fa06c300bf0%2FScreenshot%20from%202024-12-20%2022-13-55.png?generation=1734712977863025&amp;alt=media\" alt=\"\"></p>\n<p>And here are the couple of predictions of the mask from the model:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4362752%2Fb05de67ae436b27e3ae5810b759df819%2FScreenshot%20from%202024-12-20%2022-19-24.png?generation=1734713250624707&amp;alt=media\" alt=\"\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4362752%2F25bbb600c38bcce78b4130cae040b1b5%2FScreenshot%20from%202024-12-20%2022-19-10.png?generation=1734713262702589&amp;alt=media\" alt=\"\"></p>\n<p>According to me the predictions looks ok or not as bad as to give me a score of 0 i.e. miss all the particles.</p>\n<p>So, I have the following questions:</p>\n<ol>\n<li>Is my idea sane? i.e. using 2D UNet to get segmentation masks and then stack them to predict centroids?</li>\n<li>If so, is the metric \"dice_score_fn\" that I have defined makes sense?</li>\n<li>How should I go about debugging this?</li>\n<li>what are other methods that I can try using 2D unet?</li>\n</ol>\n<p>I have published all the training code in the following notebook: <a href=\"https://www.kaggle.com/code/iamparadox/efficientnet-b7-2d-unet-training?scriptVersionId=214045197\" target=\"_blank\">https://www.kaggle.com/code/iamparadox/efficientnet-b7-2d-unet-training?scriptVersionId=214045197</a></p>\n<p>Any help regarding this is very much appreciated, Thanks!</p>",
  "messages": [
    {
      "id": "3077195",
      "postDate": "12/20/2024 16:52:36",
      "content": "<p>Hello everyone!</p>\n<p>I have been trying to make 2D UNet work for a while but it seems like I am failing miserably in this(as I am getting a score of 0). My idea is the following:</p>\n<ol>\n<li>create segmentation mask using copick</li>\n<li>create 2D dataset of the tomograph along the depth dimension.</li>\n<li>use UNet model and cross-entropy loss to train the model</li>\n<li>During inference we will slice the tomograph and then predict it's corresponding heatmap and then stack all the heatmaps together to find the centeroids.</li>\n<li>Also while creating the 2D slices I include neighbouring depth images i.e +- 1 depth</li>\n<li>While training I try to maximize the <code>dice_score</code> which is defined as follows (I am not sure whether this is the correct metric that I should maximize for):</li>\n</ol>\n<pre><code> ():\n    \n\n    \n    \n    bg_preds = ~((ytrue == ) &amp; (ypred == ))\n\n    ytrue = ytrue[bg_preds]\n    ypred = ypred[bg_preds]\n\n     torch.count_nonzero(ytrue == ypred) / (ytrue)\n</code></pre>\n<p>With <code>efficientnet-b7</code> as the backbone I am able to achieve a \"dice_score\" of <code>0.2597</code> on the validation set. I interpret this score as \"Approximately 1 in 4 pixels are correctly predicted by the model\". Though when plotting the ce loss and the dice score of the validation set I get the following graph (which I thought was quite interesting, though I am not sure how to interpret it):</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4362752%2F2bdb10bcc9005972dd24375bedd4fd73%2FScreenshot%20from%202024-12-20%2022-12-37.png?generation=1734712969168320&amp;alt=media\" alt=\"\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4362752%2Fa491de5efddb60fcf7f29fa06c300bf0%2FScreenshot%20from%202024-12-20%2022-13-55.png?generation=1734712977863025&amp;alt=media\" alt=\"\"></p>\n<p>And here are the couple of predictions of the mask from the model:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4362752%2Fb05de67ae436b27e3ae5810b759df819%2FScreenshot%20from%202024-12-20%2022-19-24.png?generation=1734713250624707&amp;alt=media\" alt=\"\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4362752%2F25bbb600c38bcce78b4130cae040b1b5%2FScreenshot%20from%202024-12-20%2022-19-10.png?generation=1734713262702589&amp;alt=media\" alt=\"\"></p>\n<p>According to me the predictions looks ok or not as bad as to give me a score of 0 i.e. miss all the particles.</p>\n<p>So, I have the following questions:</p>\n<ol>\n<li>Is my idea sane? i.e. using 2D UNet to get segmentation masks and then stack them to predict centroids?</li>\n<li>If so, is the metric \"dice_score_fn\" that I have defined makes sense?</li>\n<li>How should I go about debugging this?</li>\n<li>what are other methods that I can try using 2D unet?</li>\n</ol>\n<p>I have published all the training code in the following notebook: <a href=\"https://www.kaggle.com/code/iamparadox/efficientnet-b7-2d-unet-training?scriptVersionId=214045197\" target=\"_blank\">https://www.kaggle.com/code/iamparadox/efficientnet-b7-2d-unet-training?scriptVersionId=214045197</a></p>\n<p>Any help regarding this is very much appreciated, Thanks!</p>",
      "rawMarkdown": "Hello everyone!\n\nI have been trying to make 2D UNet work for a while but it seems like I am failing miserably in this(as I am getting a score of 0). My idea is the following:\n\n\n1.  create segmentation mask using copick\n2. create 2D dataset of the tomograph along the depth dimension.\n3. use UNet model and cross-entropy loss to train the model\n4. During inference we will slice the tomograph and then predict it's corresponding heatmap and then stack all the heatmaps together to find the centeroids.\n5. Also while creating the 2D slices I include neighbouring depth images i.e +- 1 depth\n6. While training I try to maximize the `dice_score` which is defined as follows (I am not sure whether this is the correct metric that I should maximize for):\n```python\n\ndef dice_score_fn(ytrue, ypred):\n    \"\"\"\n        count number of correctly predicted pixels that are not background pixels.\n    \"\"\"\n\n    # We are not intereseted in how many background pixels the model predicts correctly\n    # The idea behind this is generally the image will have 95%+ background pixels\n    bg_preds = ~((ytrue == 0) & (ypred == 0))\n\n    ytrue = ytrue[bg_preds]\n    ypred = ypred[bg_preds]\n   \n    return torch.count_nonzero(ytrue == ypred) / len(ytrue)\n``` \n\nWith `efficientnet-b7` as the backbone I am able to achieve a \"dice_score\" of `0.2597` on the validation set. I interpret this score as \"Approximately 1 in 4 pixels are correctly predicted by the model\". Though when plotting the ce loss and the dice score of the validation set I get the following graph (which I thought was quite interesting, though I am not sure how to interpret it):\n\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4362752%2F2bdb10bcc9005972dd24375bedd4fd73%2FScreenshot%20from%202024-12-20%2022-12-37.png?generation=1734712969168320&alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4362752%2Fa491de5efddb60fcf7f29fa06c300bf0%2FScreenshot%20from%202024-12-20%2022-13-55.png?generation=1734712977863025&alt=media)\n\nAnd here are the couple of predictions of the mask from the model:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4362752%2Fb05de67ae436b27e3ae5810b759df819%2FScreenshot%20from%202024-12-20%2022-19-24.png?generation=1734713250624707&alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4362752%2F25bbb600c38bcce78b4130cae040b1b5%2FScreenshot%20from%202024-12-20%2022-19-10.png?generation=1734713262702589&alt=media)\n\nAccording to me the predictions looks ok or not as bad as to give me a score of 0 i.e. miss all the particles.\n\n\nSo, I have the following questions:\n\n1. Is my idea sane? i.e. using 2D UNet to get segmentation masks and then stack them to predict centroids?\n2. If so, is the metric \"dice_score_fn\" that I have defined makes sense?\n3. How should I go about debugging this?\n4. what are other methods that I can try using 2D unet?\n\nI have published all the training code in the following notebook: https://www.kaggle.com/code/iamparadox/efficientnet-b7-2d-unet-training?scriptVersionId=214045197\n\nAny help regarding this is very much appreciated, Thanks!",
      "votes": null
    },
    {
      "id": "3077293",
      "postDate": "12/20/2024 18:47:51",
      "content": "<p>Hi. Those predictions corresponds to 1.2 CE loss? Are you sure?</p>\n<p>EDIT: Now I've seen weights 1.2 loss makes sense. I would say the problem is in inference code. Metric doesn't affect training. Make sure you predict global 3D centroids not 184 2D centroids.</p>",
      "rawMarkdown": "Hi. Those predictions corresponds to 1.2 CE loss? Are you sure?\n\nEDIT: Now I've seen weights 1.2 loss makes sense. I would say the problem is in inference code. Metric doesn't affect training. Make sure you predict global 3D centroids not 184 2D centroids.",
      "votes": null
    },
    {
      "id": "3077541",
      "postDate": "12/21/2024 05:17:59",
      "content": "<p>Hey <a href=\"https://www.kaggle.com/sacuscreed\" target=\"_blank\">@sacuscreed</a>, thanks for the pointer!, Yup the issue was in inference stage.</p>",
      "rawMarkdown": "Hey @sacuscreed, thanks for the pointer!, Yup the issue was in inference stage.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3077293,
      "author_name": "sacuscreed",
      "author_url": "",
      "post_date": "12/20/2024 18:47:51",
      "content": "<p>Hi. Those predictions corresponds to 1.2 CE loss? Are you sure?</p>\n<p>EDIT: Now I've seen weights 1.2 loss makes sense. I would say the problem is in inference code. Metric doesn't affect training. Make sure you predict global 3D centroids not 184 2D centroids.</p>",
      "votes": null,
      "replies": [
        {
          "id": 3077541,
          "author_name": "iamparadox",
          "author_url": "",
          "post_date": "12/21/2024 05:17:59",
          "content": "<p>Hey <a href=\"https://www.kaggle.com/sacuscreed\" target=\"_blank\">@sacuscreed</a>, thanks for the pointer!, Yup the issue was in inference stage.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3077195": "Hello everyone!\n\nI have been trying to make 2D UNet work for a while but it seems like I am failing miserably in this(as I am getting a score of 0). My idea is the following:\n\n\n1.  create segmentation mask using copick\n2. create 2D dataset of the tomograph along the depth dimension.\n3. use UNet model and cross-entropy loss to train the model\n4. During inference we will slice the tomograph and then predict it's corresponding heatmap and then stack all the heatmaps together to find the centeroids.\n5. Also while creating the 2D slices I include neighbouring depth images i.e +- 1 depth\n6. While training I try to maximize the `dice_score` which is defined as follows (I am not sure whether this is the correct metric that I should maximize for):\n```python\n\ndef dice_score_fn(ytrue, ypred):\n    \"\"\"\n        count number of correctly predicted pixels that are not background pixels.\n    \"\"\"\n\n    # We are not intereseted in how many background pixels the model predicts correctly\n    # The idea behind this is generally the image will have 95%+ background pixels\n    bg_preds = ~((ytrue == 0) & (ypred == 0))\n\n    ytrue = ytrue[bg_preds]\n    ypred = ypred[bg_preds]\n   \n    return torch.count_nonzero(ytrue == ypred) / len(ytrue)\n``` \n\nWith `efficientnet-b7` as the backbone I am able to achieve a \"dice_score\" of `0.2597` on the validation set. I interpret this score as \"Approximately 1 in 4 pixels are correctly predicted by the model\". Though when plotting the ce loss and the dice score of the validation set I get the following graph (which I thought was quite interesting, though I am not sure how to interpret it):\n\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4362752%2F2bdb10bcc9005972dd24375bedd4fd73%2FScreenshot%20from%202024-12-20%2022-12-37.png?generation=1734712969168320&alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4362752%2Fa491de5efddb60fcf7f29fa06c300bf0%2FScreenshot%20from%202024-12-20%2022-13-55.png?generation=1734712977863025&alt=media)\n\nAnd here are the couple of predictions of the mask from the model:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4362752%2Fb05de67ae436b27e3ae5810b759df819%2FScreenshot%20from%202024-12-20%2022-19-24.png?generation=1734713250624707&alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4362752%2F25bbb600c38bcce78b4130cae040b1b5%2FScreenshot%20from%202024-12-20%2022-19-10.png?generation=1734713262702589&alt=media)\n\nAccording to me the predictions looks ok or not as bad as to give me a score of 0 i.e. miss all the particles.\n\n\nSo, I have the following questions:\n\n1. Is my idea sane? i.e. using 2D UNet to get segmentation masks and then stack them to predict centroids?\n2. If so, is the metric \"dice_score_fn\" that I have defined makes sense?\n3. How should I go about debugging this?\n4. what are other methods that I can try using 2D unet?\n\nI have published all the training code in the following notebook: https://www.kaggle.com/code/iamparadox/efficientnet-b7-2d-unet-training?scriptVersionId=214045197\n\nAny help regarding this is very much appreciated, Thanks!",
    "3077293": "Hi. Those predictions corresponds to 1.2 CE loss? Are you sure?\n\nEDIT: Now I've seen weights 1.2 loss makes sense. I would say the problem is in inference code. Metric doesn't affect training. Make sure you predict global 3D centroids not 184 2D centroids.",
    "3077541": "Hey @sacuscreed, thanks for the pointer!, Yup the issue was in inference stage."
  },
  "source": "meta"
}