{
  "id": 556249,
  "title": " \"Differential cell counts using center-point networks\" - anybody had any luck implementing loss function?",
  "url": "/competitions/czii-cryo-et-object-identification/discussion/556249",
  "author_name": "",
  "post_date": "2025-01-12T07:27:51.414445800Z",
  "votes": null,
  "comment_count": 10,
  "views": 0,
  "content": "<p>I've tried to follow the pipeline demonstrated in this paper <a href=\"https://www.nature.com/articles/s41598-021-96067-3#Sec8\" target=\"_blank\">Differential cell counts using center-point networks</a> and code <a href=\"https://github.com/slee5777/DCNet/blob/main/DCNet_Simple_Code.ipynb\" target=\"_blank\">here</a>. </p>\n<p>The basic pipeline, create gaussian heatmaps from centre point, then use the loss below. Positive pixels are only the centre coordinate, and all other pixels are considered background (the loss is weighted by the gaussian, thus the loss for negative examples that are close to the centre are down-weighted).</p>\n<p>The model architecture is a 3d UNET (I also tested with 2d example), with 6 output channels (num class exc. background) with sigmoid activation.</p>\n<pre><code> ():\n    \n    pos_inds = gt.eq().()\n    neg_inds = gt.lt().()\n\n    neg_weights = torch.( - gt, )\n\n    loss = \n\n    pos_loss = torch.log(pred + eps) * torch.( - pred, ) * pos_inds\n    neg_loss = torch.log( - pred + eps) * torch.(pred, ) * neg_weights * neg_inds\n\n    num_pos  = pos_inds.().()\n    pos_loss = pos_loss.()\n    neg_loss = neg_loss.()\n\n     num_pos == : \n        loss = loss - neg_loss\n    :\n        loss = loss - (pos_loss + neg_loss) / num_pos\n     loss\n\n (nn.Module):\n    \n     ():\n        ().__init__()\n        .neg_loss = loss\n\n     ():\n\n        bs,ch = target.shape[:]\n        sout = torch.sigmoid(out)\n        loss1 = .neg_loss(sout, target.data)\n        loss1 = loss1.() / bs\n         loss1\n\n     ():  torch.sigmoid(x)\n</code></pre>\n<p>In a nutshell the model does not learn at all it generally just predicted everything with ~0.1 prob.</p>\n<p>To test things out I gave it a simple 2d example, with 1 centre coord (so 1 gaussian heatmap) and basically got it to overtrain for hundreds of epochs on that image. Still to no avail, it just reaches ~0.1 prob.</p>\n<p>I was skeptical of this loss in the first place because I simply can't see it working, where your ratio of positive samples to negative is tiny. I have tried scaling the loss by number of neg/pos pixels - but it still doesn't work, due to the same fact of pos/neg ratio of pixels impacting the gradient contributions. The thing is, a few papers have shown it working…. some famous ones, i.e. CornerNet</p>\n<p>Has anyone had any go at implementing this? Or anyone has any feedback, I am keen to learn.</p>",
  "messages": [
    {
      "id": "3094516",
      "postDate": "01/12/2025 07:27:51",
      "content": "<p>I've tried to follow the pipeline demonstrated in this paper <a href=\"https://www.nature.com/articles/s41598-021-96067-3#Sec8\" target=\"_blank\">Differential cell counts using center-point networks</a> and code <a href=\"https://github.com/slee5777/DCNet/blob/main/DCNet_Simple_Code.ipynb\" target=\"_blank\">here</a>. </p>\n<p>The basic pipeline, create gaussian heatmaps from centre point, then use the loss below. Positive pixels are only the centre coordinate, and all other pixels are considered background (the loss is weighted by the gaussian, thus the loss for negative examples that are close to the centre are down-weighted).</p>\n<p>The model architecture is a 3d UNET (I also tested with 2d example), with 6 output channels (num class exc. background) with sigmoid activation.</p>\n<pre><code> ():\n    \n    pos_inds = gt.eq().()\n    neg_inds = gt.lt().()\n\n    neg_weights = torch.( - gt, )\n\n    loss = \n\n    pos_loss = torch.log(pred + eps) * torch.( - pred, ) * pos_inds\n    neg_loss = torch.log( - pred + eps) * torch.(pred, ) * neg_weights * neg_inds\n\n    num_pos  = pos_inds.().()\n    pos_loss = pos_loss.()\n    neg_loss = neg_loss.()\n\n     num_pos == : \n        loss = loss - neg_loss\n    :\n        loss = loss - (pos_loss + neg_loss) / num_pos\n     loss\n\n (nn.Module):\n    \n     ():\n        ().__init__()\n        .neg_loss = loss\n\n     ():\n\n        bs,ch = target.shape[:]\n        sout = torch.sigmoid(out)\n        loss1 = .neg_loss(sout, target.data)\n        loss1 = loss1.() / bs\n         loss1\n\n     ():  torch.sigmoid(x)\n</code></pre>\n<p>In a nutshell the model does not learn at all it generally just predicted everything with ~0.1 prob.</p>\n<p>To test things out I gave it a simple 2d example, with 1 centre coord (so 1 gaussian heatmap) and basically got it to overtrain for hundreds of epochs on that image. Still to no avail, it just reaches ~0.1 prob.</p>\n<p>I was skeptical of this loss in the first place because I simply can't see it working, where your ratio of positive samples to negative is tiny. I have tried scaling the loss by number of neg/pos pixels - but it still doesn't work, due to the same fact of pos/neg ratio of pixels impacting the gradient contributions. The thing is, a few papers have shown it working…. some famous ones, i.e. CornerNet</p>\n<p>Has anyone had any go at implementing this? Or anyone has any feedback, I am keen to learn.</p>",
      "rawMarkdown": "I've tried to follow the pipeline demonstrated in this paper [Differential cell counts using center-point networks]( https://www.nature.com/articles/s41598-021-96067-3#Sec8) and code [here](https://github.com/slee5777/DCNet/blob/main/DCNet_Simple_Code.ipynb). \n\nThe basic pipeline, create gaussian heatmaps from centre point, then use the loss below. Positive pixels are only the centre coordinate, and all other pixels are considered background (the loss is weighted by the gaussian, thus the loss for negative examples that are close to the centre are down-weighted).\n\nThe model architecture is a 3d UNET (I also tested with 2d example), with 6 output channels (num class exc. background) with sigmoid activation.\n\n```python\ndef _neg_loss(pred, gt, eps=1e-12):\n    ''' Modified focal loss. Exactly the same as CornerNet.\n      Runs faster and costs a little bit more memory\n    Arguments:\n      pred (batch x c x h x w)\n      gt_regr (batch x c x h x w)\n    '''\n    pos_inds = gt.eq(1).float()\n    neg_inds = gt.lt(1).float()\n\n    neg_weights = torch.pow(1 - gt, 4)\n\n    loss = 0\n\n    pos_loss = torch.log(pred + eps) * torch.pow(1 - pred, 2) * pos_inds\n    neg_loss = torch.log(1 - pred + eps) * torch.pow(pred, 2) * neg_weights * neg_inds\n\n    num_pos  = pos_inds.float().sum()\n    pos_loss = pos_loss.sum()\n    neg_loss = neg_loss.sum()\n\n    if num_pos == 0: \n        loss = loss - neg_loss\n    else:\n        loss = loss - (pos_loss + neg_loss) / num_pos\n    return loss\n\nclass FocalLoss(nn.Module):\n    '''nn.Module warpper for focal loss'''\n    def __init__(self, loss=_neg_loss):\n        super().__init__()\n        self.neg_loss = loss\n\n    def forward(self, out, target):\n        \n        bs,ch = target.shape[:2]\n        sout = torch.sigmoid(out)\n        loss1 = self.neg_loss(sout, target.data)\n        loss1 = loss1.sum() / bs\n        return loss1\n    \n    def activation(self, x): return torch.sigmoid(x)\n```\n\nIn a nutshell the model does not learn at all it generally just predicted everything with ~0.1 prob.\n\nTo test things out I gave it a simple 2d example, with 1 centre coord (so 1 gaussian heatmap) and basically got it to overtrain for hundreds of epochs on that image. Still to no avail, it just reaches ~0.1 prob.\n\nI was skeptical of this loss in the first place because I simply can't see it working, where your ratio of positive samples to negative is tiny. I have tried scaling the loss by number of neg/pos pixels - but it still doesn't work, due to the same fact of pos/neg ratio of pixels impacting the gradient contributions. The thing is, a few papers have shown it working.... some famous ones, i.e. CornerNet\n\nHas anyone had any go at implementing this? Or anyone has any feedback, I am keen to learn.",
      "votes": null
    },
    {
      "id": "3094530",
      "postDate": "01/12/2025 07:50:47",
      "content": "<p>You might try using Tversky loss and setting the beta really high like 0.95 (alpha 0.5.)  That seems to help with what may be a similar problem with segmentation.  See discussion here:<br>\n<a href=\"https://www.kaggle.com/competitions/czii-cryo-et-object-identification/discussion/551740\" target=\"_blank\">https://www.kaggle.com/competitions/czii-cryo-et-object-identification/discussion/551740</a></p>",
      "rawMarkdown": "You might try using Tversky loss and setting the beta really high like 0.95 (alpha 0.5.)  That seems to help with what may be a similar problem with segmentation.  See discussion here:\n[https://www.kaggle.com/competitions/czii-cryo-et-object-identification/discussion/551740](https://www.kaggle.com/competitions/czii-cryo-et-object-identification/discussion/551740)",
      "votes": null
    },
    {
      "id": "3094584",
      "postDate": "01/12/2025 09:33:32",
      "content": "<p>Thanks again <a href=\"https://www.kaggle.com/davidlist\" target=\"_blank\">@davidlist</a> - just had a read, interesting - I moved away from Tversky to try this, but I'll go back and play around.</p>\n<p>I am keen to see how <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> weighted cross-entropy goes on heatmaps because I would imagine that suffers from the same problem - I tried to implement that but again the model failed to learn anything.</p>\n<p>An additional question, with multi-class Tversky Loss, do we include background if setting beta = 0.95, I  would guess not for the following reason: if we do include background then we just weight the false positive for background (for background class we predict a particle class) really high as well? Please correct me if I'm wrong.</p>",
      "rawMarkdown": "Thanks again @davidlist - just had a read, interesting - I moved away from Tversky to try this, but I'll go back and play around.\n\nI am keen to see how @hengck23 weighted cross-entropy goes on heatmaps because I would imagine that suffers from the same problem - I tried to implement that but again the model failed to learn anything.\n\nAn additional question, with multi-class Tversky Loss, do we include background if setting beta = 0.95, I  would guess not for the following reason: if we do include background then we just weight the false positive for background (for background class we predict a particle class) really high as well? Please correct me if I'm wrong.",
      "votes": null
    },
    {
      "id": "3094598",
      "postDate": "01/12/2025 09:59:20",
      "content": "<p>I'm just learning as I go here.  I'd try both ways and see if one works.  😀</p>",
      "rawMarkdown": "I'm just learning as I go here.  I'd try both ways and see if one works.  😀",
      "votes": null
    },
    {
      "id": "3094755",
      "postDate": "01/12/2025 13:53:12",
      "content": "<p>Interestingly enough, I have found cross entropy losses much easier to use than dice ones, though it's certainly possible I have been getting stuck in a local minimum with them. I wanted to see if different post-processing techniques would work better with different losses, but no dice (pun intended).</p>",
      "rawMarkdown": "Interestingly enough, I have found cross entropy losses much easier to use than dice ones, though it's certainly possible I have been getting stuck in a local minimum with them. I wanted to see if different post-processing techniques would work better with different losses, but no dice (pun intended).",
      "votes": null
    },
    {
      "id": "3094806",
      "postDate": "01/12/2025 14:48:38",
      "content": "<p>I have also been experimenting with various losses, but contrarily my models are performing way worse on CE loss than compared to Tversky. I know that <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> has trained models using CE loss from this post: <a href=\"https://www.kaggle.com/competitions/czii-cryo-et-object-identification/discussion/545221#3047384\" target=\"_blank\">https://www.kaggle.com/competitions/czii-cryo-et-object-identification/discussion/545221#3047384</a>, but I am not really able to replicate that result. I think selecting weights for CE loss is a very important hyperparameter, I tried various weights based on frequency but to no success. It might be the case that the model might be getting stuck in local optimum as pointed out by hengck, I have tried his recommend method i.e. optimizer reset, it does indeed increases the dice score in both losses, but the overall score f4 decreases. </p>\n<p>But Tversky loss is working greatly out of the box for me.</p>",
      "rawMarkdown": "I have also been experimenting with various losses, but contrarily my models are performing way worse on CE loss than compared to Tversky. I know that @hengck23 has trained models using CE loss from this post: https://www.kaggle.com/competitions/czii-cryo-et-object-identification/discussion/545221#3047384, but I am not really able to replicate that result. I think selecting weights for CE loss is a very important hyperparameter, I tried various weights based on frequency but to no success. It might be the case that the model might be getting stuck in local optimum as pointed out by hengck, I have tried his recommend method i.e. optimizer reset, it does indeed increases the dice score in both losses, but the overall score f4 decreases. \n\nBut Tversky loss is working greatly out of the box for me.",
      "votes": null
    },
    {
      "id": "3094906",
      "postDate": "01/12/2025 17:23:54",
      "content": "<p>Are you using binary masks for the labels? For CE I have been using gaussian heatmaps mostly.</p>",
      "rawMarkdown": "Are you using binary masks for the labels? For CE I have been using gaussian heatmaps mostly.",
      "votes": null
    },
    {
      "id": "3094937",
      "postDate": "01/12/2025 17:51:19",
      "content": "<p>I imagine 1 heatmap per particle, correct?</p>",
      "rawMarkdown": "I imagine 1 heatmap per particle, correct?",
      "votes": null
    },
    {
      "id": "3094945",
      "postDate": "01/12/2025 18:05:29",
      "content": "<p>Yup, I am using guassian heatmaps. The public LB score I got on that model was: 0.635, same as what hengck23 in his notebook: <a href=\"https://www.kaggle.com/code/hengck23/3d-unet-using-2d-image-encoder\" target=\"_blank\">https://www.kaggle.com/code/hengck23/3d-unet-using-2d-image-encoder</a>, which I thought was way to low for a single model. Hence I decided to experiment with other methods and losses like: TVersky which were giving me good baseline scores to improve upon. </p>\n<p>I will try to revisit this approach again, and see if I can make it work!</p>",
      "rawMarkdown": "Yup, I am using guassian heatmaps. The public LB score I got on that model was: 0.635, same as what hengck23 in his notebook: https://www.kaggle.com/code/hengck23/3d-unet-using-2d-image-encoder, which I thought was way to low for a single model. Hence I decided to experiment with other methods and losses like: TVersky which were giving me good baseline scores to improve upon. \n\nI will try to revisit this approach again, and see if I can make it work!",
      "votes": null
    },
    {
      "id": "3095066",
      "postDate": "01/12/2025 22:19:11",
      "content": "<p>Hey <a href=\"https://www.kaggle.com/andreizamfir\" target=\"_blank\">@andreizamfir</a> I mean one heatmap per particle type (5 channels), but each channel has all of the particles.</p>\n<p><a href=\"https://www.kaggle.com/iamparadox\" target=\"_blank\">@iamparadox</a> Ah, so interesting. I can confirm my best single model scores 0.736 without TTA or ensembling. However, it doesn't seem to benefit much from these techniques (yet) either.</p>",
      "rawMarkdown": "Hey @andreizamfir I mean one heatmap per particle type (5 channels), but each channel has all of the particles.\n\n@iamparadox Ah, so interesting. I can confirm my best single model scores 0.736 without TTA or ensembling. However, it doesn't seem to benefit much from these techniques (yet) either.",
      "votes": null
    },
    {
      "id": "3095131",
      "postDate": "01/13/2025 02:30:19",
      "content": "<p>That's a nice score!, with TVersky loss my best model scores 0.713 on public lb</p>",
      "rawMarkdown": "That's a nice score!, with TVersky loss my best model scores 0.713 on public lb",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3094530,
      "author_name": "davidlist",
      "author_url": "",
      "post_date": "01/12/2025 07:50:47",
      "content": "<p>You might try using Tversky loss and setting the beta really high like 0.95 (alpha 0.5.)  That seems to help with what may be a similar problem with segmentation.  See discussion here:<br>\n<a href=\"https://www.kaggle.com/competitions/czii-cryo-et-object-identification/discussion/551740\" target=\"_blank\">https://www.kaggle.com/competitions/czii-cryo-et-object-identification/discussion/551740</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 3094584,
          "author_name": "homiecal",
          "author_url": "",
          "post_date": "01/12/2025 09:33:32",
          "content": "<p>Thanks again <a href=\"https://www.kaggle.com/davidlist\" target=\"_blank\">@davidlist</a> - just had a read, interesting - I moved away from Tversky to try this, but I'll go back and play around.</p>\n<p>I am keen to see how <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> weighted cross-entropy goes on heatmaps because I would imagine that suffers from the same problem - I tried to implement that but again the model failed to learn anything.</p>\n<p>An additional question, with multi-class Tversky Loss, do we include background if setting beta = 0.95, I  would guess not for the following reason: if we do include background then we just weight the false positive for background (for background class we predict a particle class) really high as well? Please correct me if I'm wrong.</p>",
          "votes": null,
          "replies": [
            {
              "id": 3094598,
              "author_name": "davidlist",
              "author_url": "",
              "post_date": "01/12/2025 09:59:20",
              "content": "<p>I'm just learning as I go here.  I'd try both ways and see if one works.  😀</p>",
              "votes": null,
              "replies": [
                {
                  "id": 3094755,
                  "author_name": "chemdatafarmer",
                  "author_url": "",
                  "post_date": "01/12/2025 13:53:12",
                  "content": "<p>Interestingly enough, I have found cross entropy losses much easier to use than dice ones, though it's certainly possible I have been getting stuck in a local minimum with them. I wanted to see if different post-processing techniques would work better with different losses, but no dice (pun intended).</p>",
                  "votes": null,
                  "replies": [
                    {
                      "id": 3094806,
                      "author_name": "iamparadox",
                      "author_url": "",
                      "post_date": "01/12/2025 14:48:38",
                      "content": "<p>I have also been experimenting with various losses, but contrarily my models are performing way worse on CE loss than compared to Tversky. I know that <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> has trained models using CE loss from this post: <a href=\"https://www.kaggle.com/competitions/czii-cryo-et-object-identification/discussion/545221#3047384\" target=\"_blank\">https://www.kaggle.com/competitions/czii-cryo-et-object-identification/discussion/545221#3047384</a>, but I am not really able to replicate that result. I think selecting weights for CE loss is a very important hyperparameter, I tried various weights based on frequency but to no success. It might be the case that the model might be getting stuck in local optimum as pointed out by hengck, I have tried his recommend method i.e. optimizer reset, it does indeed increases the dice score in both losses, but the overall score f4 decreases. </p>\n<p>But Tversky loss is working greatly out of the box for me.</p>",
                      "votes": null,
                      "replies": [
                        {
                          "id": 3094906,
                          "author_name": "chemdatafarmer",
                          "author_url": "",
                          "post_date": "01/12/2025 17:23:54",
                          "content": "<p>Are you using binary masks for the labels? For CE I have been using gaussian heatmaps mostly.</p>",
                          "votes": null,
                          "replies": [
                            {
                              "id": 3094937,
                              "author_name": "andreizamfir",
                              "author_url": "",
                              "post_date": "01/12/2025 17:51:19",
                              "content": "<p>I imagine 1 heatmap per particle, correct?</p>",
                              "votes": null,
                              "replies": []
                            },
                            {
                              "id": 3094945,
                              "author_name": "iamparadox",
                              "author_url": "",
                              "post_date": "01/12/2025 18:05:29",
                              "content": "<p>Yup, I am using guassian heatmaps. The public LB score I got on that model was: 0.635, same as what hengck23 in his notebook: <a href=\"https://www.kaggle.com/code/hengck23/3d-unet-using-2d-image-encoder\" target=\"_blank\">https://www.kaggle.com/code/hengck23/3d-unet-using-2d-image-encoder</a>, which I thought was way to low for a single model. Hence I decided to experiment with other methods and losses like: TVersky which were giving me good baseline scores to improve upon. </p>\n<p>I will try to revisit this approach again, and see if I can make it work!</p>",
                              "votes": null,
                              "replies": [
                                {
                                  "id": 3095066,
                                  "author_name": "chemdatafarmer",
                                  "author_url": "",
                                  "post_date": "01/12/2025 22:19:11",
                                  "content": "<p>Hey <a href=\"https://www.kaggle.com/andreizamfir\" target=\"_blank\">@andreizamfir</a> I mean one heatmap per particle type (5 channels), but each channel has all of the particles.</p>\n<p><a href=\"https://www.kaggle.com/iamparadox\" target=\"_blank\">@iamparadox</a> Ah, so interesting. I can confirm my best single model scores 0.736 without TTA or ensembling. However, it doesn't seem to benefit much from these techniques (yet) either.</p>",
                                  "votes": null,
                                  "replies": [
                                    {
                                      "id": 3095131,
                                      "author_name": "iamparadox",
                                      "author_url": "",
                                      "post_date": "01/13/2025 02:30:19",
                                      "content": "<p>That's a nice score!, with TVersky loss my best model scores 0.713 on public lb</p>",
                                      "votes": null,
                                      "replies": []
                                    }
                                  ]
                                }
                              ]
                            }
                          ]
                        }
                      ]
                    }
                  ]
                }
              ]
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3094516": "I've tried to follow the pipeline demonstrated in this paper [Differential cell counts using center-point networks]( https://www.nature.com/articles/s41598-021-96067-3#Sec8) and code [here](https://github.com/slee5777/DCNet/blob/main/DCNet_Simple_Code.ipynb). \n\nThe basic pipeline, create gaussian heatmaps from centre point, then use the loss below. Positive pixels are only the centre coordinate, and all other pixels are considered background (the loss is weighted by the gaussian, thus the loss for negative examples that are close to the centre are down-weighted).\n\nThe model architecture is a 3d UNET (I also tested with 2d example), with 6 output channels (num class exc. background) with sigmoid activation.\n\n```python\ndef _neg_loss(pred, gt, eps=1e-12):\n    ''' Modified focal loss. Exactly the same as CornerNet.\n      Runs faster and costs a little bit more memory\n    Arguments:\n      pred (batch x c x h x w)\n      gt_regr (batch x c x h x w)\n    '''\n    pos_inds = gt.eq(1).float()\n    neg_inds = gt.lt(1).float()\n\n    neg_weights = torch.pow(1 - gt, 4)\n\n    loss = 0\n\n    pos_loss = torch.log(pred + eps) * torch.pow(1 - pred, 2) * pos_inds\n    neg_loss = torch.log(1 - pred + eps) * torch.pow(pred, 2) * neg_weights * neg_inds\n\n    num_pos  = pos_inds.float().sum()\n    pos_loss = pos_loss.sum()\n    neg_loss = neg_loss.sum()\n\n    if num_pos == 0: \n        loss = loss - neg_loss\n    else:\n        loss = loss - (pos_loss + neg_loss) / num_pos\n    return loss\n\nclass FocalLoss(nn.Module):\n    '''nn.Module warpper for focal loss'''\n    def __init__(self, loss=_neg_loss):\n        super().__init__()\n        self.neg_loss = loss\n\n    def forward(self, out, target):\n        \n        bs,ch = target.shape[:2]\n        sout = torch.sigmoid(out)\n        loss1 = self.neg_loss(sout, target.data)\n        loss1 = loss1.sum() / bs\n        return loss1\n    \n    def activation(self, x): return torch.sigmoid(x)\n```\n\nIn a nutshell the model does not learn at all it generally just predicted everything with ~0.1 prob.\n\nTo test things out I gave it a simple 2d example, with 1 centre coord (so 1 gaussian heatmap) and basically got it to overtrain for hundreds of epochs on that image. Still to no avail, it just reaches ~0.1 prob.\n\nI was skeptical of this loss in the first place because I simply can't see it working, where your ratio of positive samples to negative is tiny. I have tried scaling the loss by number of neg/pos pixels - but it still doesn't work, due to the same fact of pos/neg ratio of pixels impacting the gradient contributions. The thing is, a few papers have shown it working.... some famous ones, i.e. CornerNet\n\nHas anyone had any go at implementing this? Or anyone has any feedback, I am keen to learn.",
    "3094530": "You might try using Tversky loss and setting the beta really high like 0.95 (alpha 0.5.)  That seems to help with what may be a similar problem with segmentation.  See discussion here:\n[https://www.kaggle.com/competitions/czii-cryo-et-object-identification/discussion/551740](https://www.kaggle.com/competitions/czii-cryo-et-object-identification/discussion/551740)",
    "3094584": "Thanks again @davidlist - just had a read, interesting - I moved away from Tversky to try this, but I'll go back and play around.\n\nI am keen to see how @hengck23 weighted cross-entropy goes on heatmaps because I would imagine that suffers from the same problem - I tried to implement that but again the model failed to learn anything.\n\nAn additional question, with multi-class Tversky Loss, do we include background if setting beta = 0.95, I  would guess not for the following reason: if we do include background then we just weight the false positive for background (for background class we predict a particle class) really high as well? Please correct me if I'm wrong.",
    "3094598": "I'm just learning as I go here.  I'd try both ways and see if one works.  😀",
    "3094755": "Interestingly enough, I have found cross entropy losses much easier to use than dice ones, though it's certainly possible I have been getting stuck in a local minimum with them. I wanted to see if different post-processing techniques would work better with different losses, but no dice (pun intended).",
    "3094806": "I have also been experimenting with various losses, but contrarily my models are performing way worse on CE loss than compared to Tversky. I know that @hengck23 has trained models using CE loss from this post: https://www.kaggle.com/competitions/czii-cryo-et-object-identification/discussion/545221#3047384, but I am not really able to replicate that result. I think selecting weights for CE loss is a very important hyperparameter, I tried various weights based on frequency but to no success. It might be the case that the model might be getting stuck in local optimum as pointed out by hengck, I have tried his recommend method i.e. optimizer reset, it does indeed increases the dice score in both losses, but the overall score f4 decreases. \n\nBut Tversky loss is working greatly out of the box for me.",
    "3094906": "Are you using binary masks for the labels? For CE I have been using gaussian heatmaps mostly.",
    "3094937": "I imagine 1 heatmap per particle, correct?",
    "3094945": "Yup, I am using guassian heatmaps. The public LB score I got on that model was: 0.635, same as what hengck23 in his notebook: https://www.kaggle.com/code/hengck23/3d-unet-using-2d-image-encoder, which I thought was way to low for a single model. Hence I decided to experiment with other methods and losses like: TVersky which were giving me good baseline scores to improve upon. \n\nI will try to revisit this approach again, and see if I can make it work!",
    "3095066": "Hey @andreizamfir I mean one heatmap per particle type (5 channels), but each channel has all of the particles.\n\n@iamparadox Ah, so interesting. I can confirm my best single model scores 0.736 without TTA or ensembling. However, it doesn't seem to benefit much from these techniques (yet) either.",
    "3095131": "That's a nice score!, with TVersky loss my best model scores 0.713 on public lb"
  },
  "source": "meta"
}