{
  "id": 561483,
  "title": "26th Solution - 3D UNet + CCL with Morphological Post Processing + Custom Loss Function",
  "url": "/competitions/czii-cryo-et-object-identification/writeups/andrei-zamfir-26th-solution-3d-unet-ccl-with-morph",
  "author_name": "",
  "post_date": "2025-02-06T14:00:19.667Z",
  "votes": 15,
  "comment_count": 3,
  "views": 0,
  "content": "<p>Hello,</p>\n<p>First of all, I'd like to thank the competition hosts for the very interesting problem presented. I am in awe at the homogeneousity of the data, where our training 7 tomograms represent the data sufficiently well. The private-public split was also great, with minimal leaderboard changes.</p>\n<p>Second, I'd like to thank the community, as always, for all the insightful discussions on the forum and the Kaggle spirit of sharing ideas. None of this would have been possible for me without the collective effort of the other participants. Special thanks to <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a>, <a href=\"https://www.kaggle.com/fnands\" target=\"_blank\">@fnands</a>, <a href=\"https://www.kaggle.com/davidlist\" target=\"_blank\">@davidlist</a> and <a href=\"https://www.kaggle.com/sacuscreed\" target=\"_blank\">@sacuscreed</a>. I learned the basics of solving a computer vision problem from Hengck, and I learned a lot from dissecting his code.</p>\n<h1>Solution Summary</h1>\n<p>My solution used MONAi's full 3D UNet implementation, with a window size of 184x184x184. I generated multi-class hard labels, using spheres of 0.5 the estimated particle radius. </p>\n<p>I ended up using a 50% overlap and generated gaussian weights for volume reconstruction. For the data augmentation pipeline, I used random rotations in the X,Y plane, random flips on all spatial dimensions, random gaussian noise and random zoom (0.95-1.05). <br>\nI did also spend some time looking up ways of replicating the missing wedge artifact using affine transformations but ended up ditching the idea.</p>\n<p>I attempted training using multiple loss functions, settling on Tversky loss with multiple alpha/beta combinations, obtaining decent results, with best individual models scoring 0.70x-0.73x on the public LB. </p>\n<p>For test-time augmentations, I used 90-180-270 rotations in the X,Y plane and full spatial dimensions flips. I used two particular ways of ensembling the TTA and models together:</p>\n<ol>\n<li>Raw Logit averaging, where I take the mean of the TTAs, add them with the other model predicted mean of TTAs and perform another average, reconstruct the volume and then apply a final softmax.</li>\n<li>Softmax per TTA stack, then add max of probabilities together with the other model predictions for the subvolume, perform an average between models and reconstruct the volume.</li>\n</ol>\n<p>Both methods ended up scoring very similarly, although obviously required different thresholds for CCL, with the max of TTA being overconfident in predictions. Looking back, I imagine I could get more ensemble variety if I combined the two methods, which I have not done.</p>\n<p>Following thresholding to transform the output probabilities into hard binary masks, I played around with the following morphological operations: dilation, erosion, opening and closing. Fully dilating all my binary masks before performing CCL caused a jump on the LB score of around 0.010-0.015.</p>\n<p>I then played around with Optuna HPO studies to search for best threshold/morphological operation per particle in order to maximize the pixel FBeta (beta of 1 or 1.25 performed best, going above would sacrifice too much precision) score for said particle (not to be confused with competition FBeta).</p>\n<p>I have only validated on TS_99_9 and trained on the other 6 samples throughout the whole competition, and I tried finding a correlation between any local metric and the LB. I ended up using weighted pixel FBeta as I found the LB FBeta to be very unstable and not correlate too well on the leaderboard.</p>\n<h1>Additional thoughts</h1>\n<p>A good part of January I spent on analyzing this problem, how to find this correlation. On the one hand, we are not working on a segmentation problem per se, so comparing the dice metric between two models does not say the whole story, on the other hand, there is some correlation between dice/IoU or whatever other metric you want to use and the way it will translate over to LB.</p>\n<p>Say you already have a cluster of pixel predictions in the right place, and whatever peak detection method you use already 'scores' you that particle. Further improving those probabilities (getting them closer to 1, or covering more of the ground truth) improves dice without improving competition FBeta.</p>\n<p>This has led me to think about a loss function that could prioritize identifying as many objects as possible (to encourage object level recall, directly translating to competition FBeta). </p>\n<pre><code> CustomLoss(self, predictions, labels, metadata):\n\n     = torch.softmax(predictions, dim=)\n     = .\n     = \n     = e-\n     = \n     = \n     = \n     = \n     class_label, obj_id, obj_mask in metadata:\n\n         = pred_mask[class_label]\n             class_label == :\n                \n\n         = torch.sum(pred_probs * obj_mask)\n         = torch.sum(( - pred_probs) * obj_mask)  \n         = class_weights[class_label]\n\n         = tp/(tp+fn+eps)\n         = torch.sigmoid(sharpness_sigmoid * (recall - threshold_sigmoid))\n\n         recall&gt;threshold_sigmoid:\n             += \n\n         += (-torch.log(sigmoid_recall + eps)) * weight\n\n         += \n         += weight\n\n     = total_object_loss / (n_objects + eps)\n    (f'Custom loss is equal to {loss}!')\n    (f'Objects found with over {threshold_sigmoid} in recall: {objects_found} out of {object_counter}!')\n\n     loss\n</code></pre>\n<p>The above function runs through the metadata for a particular batch (which is generated outside for each batch, to maintain differentiability) and iterates through classes and individual objects and calculates recall (assume you have 3 ribosomes and 5 apo-ferritins in this batch, metadata would create 8 ground truth masks, one for each particle sphere).</p>\n<p>I then used sigmoid for soft thresholding and calculate loss as negative log of sigmoid of recall. The idea behind was to encourage the network to identify as many objects as possible (with the way metadata splits objects, only recall can be calculated, as the precision of the predictions would be way off considering it would take into account all predictions for that said class, dramatically decreasing precision).</p>\n<p>This loss was combined with Tversky loss with a higher alpha to prioritize precision, in the attempt of antagonizing all the positive predictions that the custom loss function would generate to reach 5% recall per object.</p>\n<p>This combination required a lot of tuning for the right recall threshold, the right class weights as well as the weights of the custom loss and Tversky loss blend (-log of epsilon would dominate Tversky which is bounded to [0,1]).</p>\n<p>Unfortunately, I only got decent results with it, not really improving my LB score and I had to give up on the idea as I lacked the GPU required to further delve into this, plus the competition deadline was approaching.</p>\n<p>In the end, I used a model trained with this combination, which at least helped the diversity of the ensemble blend.</p>\n<h1>Lessons learned</h1>\n<p>This was my first time dedicating 3 months to a competition and it was interesting to see how the leaderboard evolved, the discussions on the forum and the shared code snippets. I now have better expectations of how to manage my time, my GPU, and I think the greatest take away would be to always revisit earlier parts of my code pipeline and be more mindful of how and why they work together.</p>\n<p>I tend to tunnel vision on the task at hand and take the big picture less into consideration.</p>\n<p>Sometimes, less is more, and taking a break from a problem you can't solve makes you come back with a fresher perspective.</p>",
  "messages": [
    {
      "id": "3116872",
      "postDate": "02/06/2025 11:53:58",
      "content": "<p>Hello,</p>\n<p>First of all, I'd like to thank the competition hosts for the very interesting problem presented. I am in awe at the homogeneousity of the data, where our training 7 tomograms represent the data sufficiently well. The private-public split was also great, with minimal leaderboard changes.</p>\n<p>Second, I'd like to thank the community, as always, for all the insightful discussions on the forum and the Kaggle spirit of sharing ideas. None of this would have been possible for me without the collective effort of the other participants. Special thanks to <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a>, <a href=\"https://www.kaggle.com/fnands\" target=\"_blank\">@fnands</a>, <a href=\"https://www.kaggle.com/davidlist\" target=\"_blank\">@davidlist</a> and <a href=\"https://www.kaggle.com/sacuscreed\" target=\"_blank\">@sacuscreed</a>. I learned the basics of solving a computer vision problem from Hengck, and I learned a lot from dissecting his code.</p>\n<h1>Solution Summary</h1>\n<p>My solution used MONAi's full 3D UNet implementation, with a window size of 184x184x184. I generated multi-class hard labels, using spheres of 0.5 the estimated particle radius. </p>\n<p>I ended up using a 50% overlap and generated gaussian weights for volume reconstruction. For the data augmentation pipeline, I used random rotations in the X,Y plane, random flips on all spatial dimensions, random gaussian noise and random zoom (0.95-1.05). <br>\nI did also spend some time looking up ways of replicating the missing wedge artifact using affine transformations but ended up ditching the idea.</p>\n<p>I attempted training using multiple loss functions, settling on Tversky loss with multiple alpha/beta combinations, obtaining decent results, with best individual models scoring 0.70x-0.73x on the public LB. </p>\n<p>For test-time augmentations, I used 90-180-270 rotations in the X,Y plane and full spatial dimensions flips. I used two particular ways of ensembling the TTA and models together:</p>\n<ol>\n<li>Raw Logit averaging, where I take the mean of the TTAs, add them with the other model predicted mean of TTAs and perform another average, reconstruct the volume and then apply a final softmax.</li>\n<li>Softmax per TTA stack, then add max of probabilities together with the other model predictions for the subvolume, perform an average between models and reconstruct the volume.</li>\n</ol>\n<p>Both methods ended up scoring very similarly, although obviously required different thresholds for CCL, with the max of TTA being overconfident in predictions. Looking back, I imagine I could get more ensemble variety if I combined the two methods, which I have not done.</p>\n<p>Following thresholding to transform the output probabilities into hard binary masks, I played around with the following morphological operations: dilation, erosion, opening and closing. Fully dilating all my binary masks before performing CCL caused a jump on the LB score of around 0.010-0.015.</p>\n<p>I then played around with Optuna HPO studies to search for best threshold/morphological operation per particle in order to maximize the pixel FBeta (beta of 1 or 1.25 performed best, going above would sacrifice too much precision) score for said particle (not to be confused with competition FBeta).</p>\n<p>I have only validated on TS_99_9 and trained on the other 6 samples throughout the whole competition, and I tried finding a correlation between any local metric and the LB. I ended up using weighted pixel FBeta as I found the LB FBeta to be very unstable and not correlate too well on the leaderboard.</p>\n<h1>Additional thoughts</h1>\n<p>A good part of January I spent on analyzing this problem, how to find this correlation. On the one hand, we are not working on a segmentation problem per se, so comparing the dice metric between two models does not say the whole story, on the other hand, there is some correlation between dice/IoU or whatever other metric you want to use and the way it will translate over to LB.</p>\n<p>Say you already have a cluster of pixel predictions in the right place, and whatever peak detection method you use already 'scores' you that particle. Further improving those probabilities (getting them closer to 1, or covering more of the ground truth) improves dice without improving competition FBeta.</p>\n<p>This has led me to think about a loss function that could prioritize identifying as many objects as possible (to encourage object level recall, directly translating to competition FBeta). </p>\n<pre><code> CustomLoss(self, predictions, labels, metadata):\n\n     = torch.softmax(predictions, dim=)\n     = .\n     = \n     = e-\n     = \n     = \n     = \n     = \n     class_label, obj_id, obj_mask in metadata:\n\n         = pred_mask[class_label]\n             class_label == :\n                \n\n         = torch.sum(pred_probs * obj_mask)\n         = torch.sum(( - pred_probs) * obj_mask)  \n         = class_weights[class_label]\n\n         = tp/(tp+fn+eps)\n         = torch.sigmoid(sharpness_sigmoid * (recall - threshold_sigmoid))\n\n         recall&gt;threshold_sigmoid:\n             += \n\n         += (-torch.log(sigmoid_recall + eps)) * weight\n\n         += \n         += weight\n\n     = total_object_loss / (n_objects + eps)\n    (f'Custom loss is equal to {loss}!')\n    (f'Objects found with over {threshold_sigmoid} in recall: {objects_found} out of {object_counter}!')\n\n     loss\n</code></pre>\n<p>The above function runs through the metadata for a particular batch (which is generated outside for each batch, to maintain differentiability) and iterates through classes and individual objects and calculates recall (assume you have 3 ribosomes and 5 apo-ferritins in this batch, metadata would create 8 ground truth masks, one for each particle sphere).</p>\n<p>I then used sigmoid for soft thresholding and calculate loss as negative log of sigmoid of recall. The idea behind was to encourage the network to identify as many objects as possible (with the way metadata splits objects, only recall can be calculated, as the precision of the predictions would be way off considering it would take into account all predictions for that said class, dramatically decreasing precision).</p>\n<p>This loss was combined with Tversky loss with a higher alpha to prioritize precision, in the attempt of antagonizing all the positive predictions that the custom loss function would generate to reach 5% recall per object.</p>\n<p>This combination required a lot of tuning for the right recall threshold, the right class weights as well as the weights of the custom loss and Tversky loss blend (-log of epsilon would dominate Tversky which is bounded to [0,1]).</p>\n<p>Unfortunately, I only got decent results with it, not really improving my LB score and I had to give up on the idea as I lacked the GPU required to further delve into this, plus the competition deadline was approaching.</p>\n<p>In the end, I used a model trained with this combination, which at least helped the diversity of the ensemble blend.</p>\n<h1>Lessons learned</h1>\n<p>This was my first time dedicating 3 months to a competition and it was interesting to see how the leaderboard evolved, the discussions on the forum and the shared code snippets. I now have better expectations of how to manage my time, my GPU, and I think the greatest take away would be to always revisit earlier parts of my code pipeline and be more mindful of how and why they work together.</p>\n<p>I tend to tunnel vision on the task at hand and take the big picture less into consideration.</p>\n<p>Sometimes, less is more, and taking a break from a problem you can't solve makes you come back with a fresher perspective.</p>",
      "rawMarkdown": "Hello,\n\nFirst of all, I'd like to thank the competition hosts for the very interesting problem presented. I am in awe at the homogeneousity of the data, where our training 7 tomograms represent the data sufficiently well. The private-public split was also great, with minimal leaderboard changes.\n\nSecond, I'd like to thank the community, as always, for all the insightful discussions on the forum and the Kaggle spirit of sharing ideas. None of this would have been possible for me without the collective effort of the other participants. Special thanks to @hengck23, @fnands, @davidlist and @sacuscreed. I learned the basics of solving a computer vision problem from Hengck, and I learned a lot from dissecting his code.\n\n# Solution Summary\n\nMy solution used MONAi's full 3D UNet implementation, with a window size of 184x184x184. I generated multi-class hard labels, using spheres of 0.5 the estimated particle radius. \n\nI ended up using a 50% overlap and generated gaussian weights for volume reconstruction. For the data augmentation pipeline, I used random rotations in the X,Y plane, random flips on all spatial dimensions, random gaussian noise and random zoom (0.95-1.05). \nI did also spend some time looking up ways of replicating the missing wedge artifact using affine transformations but ended up ditching the idea.\n\nI attempted training using multiple loss functions, settling on Tversky loss with multiple alpha/beta combinations, obtaining decent results, with best individual models scoring 0.70x-0.73x on the public LB. \n\nFor test-time augmentations, I used 90-180-270 rotations in the X,Y plane and full spatial dimensions flips. I used two particular ways of ensembling the TTA and models together:\n\n1.  Raw Logit averaging, where I take the mean of the TTAs, add them with the other model predicted mean of TTAs and perform another average, reconstruct the volume and then apply a final softmax.\n2. Softmax per TTA stack, then add max of probabilities together with the other model predictions for the subvolume, perform an average between models and reconstruct the volume.\n\nBoth methods ended up scoring very similarly, although obviously required different thresholds for CCL, with the max of TTA being overconfident in predictions. Looking back, I imagine I could get more ensemble variety if I combined the two methods, which I have not done.\n\nFollowing thresholding to transform the output probabilities into hard binary masks, I played around with the following morphological operations: dilation, erosion, opening and closing. Fully dilating all my binary masks before performing CCL caused a jump on the LB score of around 0.010-0.015.\n\nI then played around with Optuna HPO studies to search for best threshold/morphological operation per particle in order to maximize the pixel FBeta (beta of 1 or 1.25 performed best, going above would sacrifice too much precision) score for said particle (not to be confused with competition FBeta).\n\nI have only validated on TS_99_9 and trained on the other 6 samples throughout the whole competition, and I tried finding a correlation between any local metric and the LB. I ended up using weighted pixel FBeta as I found the LB FBeta to be very unstable and not correlate too well on the leaderboard.\n\n# Additional thoughts\n\nA good part of January I spent on analyzing this problem, how to find this correlation. On the one hand, we are not working on a segmentation problem per se, so comparing the dice metric between two models does not say the whole story, on the other hand, there is some correlation between dice/IoU or whatever other metric you want to use and the way it will translate over to LB.\n\nSay you already have a cluster of pixel predictions in the right place, and whatever peak detection method you use already 'scores' you that particle. Further improving those probabilities (getting them closer to 1, or covering more of the ground truth) improves dice without improving competition FBeta.\n\nThis has led me to think about a loss function that could prioritize identifying as many objects as possible (to encourage object level recall, directly translating to competition FBeta). \n\n```\n\ndef CustomLoss(self, predictions, labels, metadata):\n    \n    pred_mask = torch.softmax(predictions, dim=0)\n    threshold_sigmoid = 0.05\n    sharpness_sigmoid = 10\n    eps = 1e-6\n    n_objects = 0\n    total_object_loss = 0\n    objects_found = 0\n    object_counter = 0\n    for class_label, obj_id, obj_mask in metadata:\n            \n        pred_probs = pred_mask[class_label]\n            if class_label == 2:\n                continue\n            \n        tp = torch.sum(pred_probs * obj_mask)\n        fn = torch.sum((1 - pred_probs) * obj_mask)  \n        weight = class_weights[class_label]\n            \n        recall = tp/(tp+fn+eps)\n        sigmoid_recall = torch.sigmoid(sharpness_sigmoid * (recall - threshold_sigmoid))\n            \n        if recall>threshold_sigmoid:\n            objects_found += 1\n            \n        total_object_loss += (-torch.log(sigmoid_recall + eps)) * weight\n                \n        object_counter += 1\n        n_objects += weight\n            \n    loss = total_object_loss / (n_objects + eps)\n    print(f'Custom loss is equal to {loss}!')\n    print(f'Objects found with over {threshold_sigmoid} in recall: {objects_found} out of {object_counter}!')\n        \n    return loss\n\n```\n\nThe above function runs through the metadata for a particular batch (which is generated outside for each batch, to maintain differentiability) and iterates through classes and individual objects and calculates recall (assume you have 3 ribosomes and 5 apo-ferritins in this batch, metadata would create 8 ground truth masks, one for each particle sphere).\n\nI then used sigmoid for soft thresholding and calculate loss as negative log of sigmoid of recall. The idea behind was to encourage the network to identify as many objects as possible (with the way metadata splits objects, only recall can be calculated, as the precision of the predictions would be way off considering it would take into account all predictions for that said class, dramatically decreasing precision).\n\nThis loss was combined with Tversky loss with a higher alpha to prioritize precision, in the attempt of antagonizing all the positive predictions that the custom loss function would generate to reach 5% recall per object.\n\nThis combination required a lot of tuning for the right recall threshold, the right class weights as well as the weights of the custom loss and Tversky loss blend (-log of epsilon would dominate Tversky which is bounded to [0,1]).\n\nUnfortunately, I only got decent results with it, not really improving my LB score and I had to give up on the idea as I lacked the GPU required to further delve into this, plus the competition deadline was approaching.\n\nIn the end, I used a model trained with this combination, which at least helped the diversity of the ensemble blend.\n\n# Lessons learned\n\nThis was my first time dedicating 3 months to a competition and it was interesting to see how the leaderboard evolved, the discussions on the forum and the shared code snippets. I now have better expectations of how to manage my time, my GPU, and I think the greatest take away would be to always revisit earlier parts of my code pipeline and be more mindful of how and why they work together.\n\nI tend to tunnel vision on the task at hand and take the big picture less into consideration.\n\nSometimes, less is more, and taking a break from a problem you can't solve makes you come back with a fresher perspective.",
      "votes": null
    },
    {
      "id": "3117084",
      "postDate": "02/06/2025 16:11:50",
      "content": "<p>If you use, validation data 'TS5_4' or 'TS_6_4' it might be higher. In my case, different validation data makes different score. Maximum 0.02 point.</p>",
      "rawMarkdown": "If you use, validation data 'TS5_4' or 'TS_6_4' it might be higher. In my case, different validation data makes different score. Maximum 0.02 point.",
      "votes": null
    },
    {
      "id": "3117155",
      "postDate": "02/06/2025 17:20:21",
      "content": "<p>Yes, definitely an area of improvement would have been to train using K-Fold or do model soups of different validation tomograms.<br>\nIt was an oversight to be this static in the training process but a great lesson nonetheless!</p>",
      "rawMarkdown": "Yes, definitely an area of improvement would have been to train using K-Fold or do model soups of different validation tomograms.\nIt was an oversight to be this static in the training process but a great lesson nonetheless!",
      "votes": null
    },
    {
      "id": "3117230",
      "postDate": "02/06/2025 18:35:52",
      "content": "<p>Yeah, those, TS_5_4 in particular, had fewer particles which meant you trained with more if you used it for validation.</p>",
      "rawMarkdown": "Yeah, those, TS_5_4 in particular, had fewer particles which meant you trained with more if you used it for validation.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3117084,
      "author_name": "junhanzangai",
      "author_url": "",
      "post_date": "02/06/2025 16:11:50",
      "content": "<p>If you use, validation data 'TS5_4' or 'TS_6_4' it might be higher. In my case, different validation data makes different score. Maximum 0.02 point.</p>",
      "votes": null,
      "replies": [
        {
          "id": 3117155,
          "author_name": "andreizamfir",
          "author_url": "",
          "post_date": "02/06/2025 17:20:21",
          "content": "<p>Yes, definitely an area of improvement would have been to train using K-Fold or do model soups of different validation tomograms.<br>\nIt was an oversight to be this static in the training process but a great lesson nonetheless!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 3117230,
          "author_name": "davidlist",
          "author_url": "",
          "post_date": "02/06/2025 18:35:52",
          "content": "<p>Yeah, those, TS_5_4 in particular, had fewer particles which meant you trained with more if you used it for validation.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3116872": "Hello,\n\nFirst of all, I'd like to thank the competition hosts for the very interesting problem presented. I am in awe at the homogeneousity of the data, where our training 7 tomograms represent the data sufficiently well. The private-public split was also great, with minimal leaderboard changes.\n\nSecond, I'd like to thank the community, as always, for all the insightful discussions on the forum and the Kaggle spirit of sharing ideas. None of this would have been possible for me without the collective effort of the other participants. Special thanks to @hengck23, @fnands, @davidlist and @sacuscreed. I learned the basics of solving a computer vision problem from Hengck, and I learned a lot from dissecting his code.\n\n# Solution Summary\n\nMy solution used MONAi's full 3D UNet implementation, with a window size of 184x184x184. I generated multi-class hard labels, using spheres of 0.5 the estimated particle radius. \n\nI ended up using a 50% overlap and generated gaussian weights for volume reconstruction. For the data augmentation pipeline, I used random rotations in the X,Y plane, random flips on all spatial dimensions, random gaussian noise and random zoom (0.95-1.05). \nI did also spend some time looking up ways of replicating the missing wedge artifact using affine transformations but ended up ditching the idea.\n\nI attempted training using multiple loss functions, settling on Tversky loss with multiple alpha/beta combinations, obtaining decent results, with best individual models scoring 0.70x-0.73x on the public LB. \n\nFor test-time augmentations, I used 90-180-270 rotations in the X,Y plane and full spatial dimensions flips. I used two particular ways of ensembling the TTA and models together:\n\n1.  Raw Logit averaging, where I take the mean of the TTAs, add them with the other model predicted mean of TTAs and perform another average, reconstruct the volume and then apply a final softmax.\n2. Softmax per TTA stack, then add max of probabilities together with the other model predictions for the subvolume, perform an average between models and reconstruct the volume.\n\nBoth methods ended up scoring very similarly, although obviously required different thresholds for CCL, with the max of TTA being overconfident in predictions. Looking back, I imagine I could get more ensemble variety if I combined the two methods, which I have not done.\n\nFollowing thresholding to transform the output probabilities into hard binary masks, I played around with the following morphological operations: dilation, erosion, opening and closing. Fully dilating all my binary masks before performing CCL caused a jump on the LB score of around 0.010-0.015.\n\nI then played around with Optuna HPO studies to search for best threshold/morphological operation per particle in order to maximize the pixel FBeta (beta of 1 or 1.25 performed best, going above would sacrifice too much precision) score for said particle (not to be confused with competition FBeta).\n\nI have only validated on TS_99_9 and trained on the other 6 samples throughout the whole competition, and I tried finding a correlation between any local metric and the LB. I ended up using weighted pixel FBeta as I found the LB FBeta to be very unstable and not correlate too well on the leaderboard.\n\n# Additional thoughts\n\nA good part of January I spent on analyzing this problem, how to find this correlation. On the one hand, we are not working on a segmentation problem per se, so comparing the dice metric between two models does not say the whole story, on the other hand, there is some correlation between dice/IoU or whatever other metric you want to use and the way it will translate over to LB.\n\nSay you already have a cluster of pixel predictions in the right place, and whatever peak detection method you use already 'scores' you that particle. Further improving those probabilities (getting them closer to 1, or covering more of the ground truth) improves dice without improving competition FBeta.\n\nThis has led me to think about a loss function that could prioritize identifying as many objects as possible (to encourage object level recall, directly translating to competition FBeta). \n\n```\n\ndef CustomLoss(self, predictions, labels, metadata):\n    \n    pred_mask = torch.softmax(predictions, dim=0)\n    threshold_sigmoid = 0.05\n    sharpness_sigmoid = 10\n    eps = 1e-6\n    n_objects = 0\n    total_object_loss = 0\n    objects_found = 0\n    object_counter = 0\n    for class_label, obj_id, obj_mask in metadata:\n            \n        pred_probs = pred_mask[class_label]\n            if class_label == 2:\n                continue\n            \n        tp = torch.sum(pred_probs * obj_mask)\n        fn = torch.sum((1 - pred_probs) * obj_mask)  \n        weight = class_weights[class_label]\n            \n        recall = tp/(tp+fn+eps)\n        sigmoid_recall = torch.sigmoid(sharpness_sigmoid * (recall - threshold_sigmoid))\n            \n        if recall>threshold_sigmoid:\n            objects_found += 1\n            \n        total_object_loss += (-torch.log(sigmoid_recall + eps)) * weight\n                \n        object_counter += 1\n        n_objects += weight\n            \n    loss = total_object_loss / (n_objects + eps)\n    print(f'Custom loss is equal to {loss}!')\n    print(f'Objects found with over {threshold_sigmoid} in recall: {objects_found} out of {object_counter}!')\n        \n    return loss\n\n```\n\nThe above function runs through the metadata for a particular batch (which is generated outside for each batch, to maintain differentiability) and iterates through classes and individual objects and calculates recall (assume you have 3 ribosomes and 5 apo-ferritins in this batch, metadata would create 8 ground truth masks, one for each particle sphere).\n\nI then used sigmoid for soft thresholding and calculate loss as negative log of sigmoid of recall. The idea behind was to encourage the network to identify as many objects as possible (with the way metadata splits objects, only recall can be calculated, as the precision of the predictions would be way off considering it would take into account all predictions for that said class, dramatically decreasing precision).\n\nThis loss was combined with Tversky loss with a higher alpha to prioritize precision, in the attempt of antagonizing all the positive predictions that the custom loss function would generate to reach 5% recall per object.\n\nThis combination required a lot of tuning for the right recall threshold, the right class weights as well as the weights of the custom loss and Tversky loss blend (-log of epsilon would dominate Tversky which is bounded to [0,1]).\n\nUnfortunately, I only got decent results with it, not really improving my LB score and I had to give up on the idea as I lacked the GPU required to further delve into this, plus the competition deadline was approaching.\n\nIn the end, I used a model trained with this combination, which at least helped the diversity of the ensemble blend.\n\n# Lessons learned\n\nThis was my first time dedicating 3 months to a competition and it was interesting to see how the leaderboard evolved, the discussions on the forum and the shared code snippets. I now have better expectations of how to manage my time, my GPU, and I think the greatest take away would be to always revisit earlier parts of my code pipeline and be more mindful of how and why they work together.\n\nI tend to tunnel vision on the task at hand and take the big picture less into consideration.\n\nSometimes, less is more, and taking a break from a problem you can't solve makes you come back with a fresher perspective.",
    "3117084": "If you use, validation data 'TS5_4' or 'TS_6_4' it might be higher. In my case, different validation data makes different score. Maximum 0.02 point.",
    "3117155": "Yes, definitely an area of improvement would have been to train using K-Fold or do model soups of different validation tomograms.\nIt was an oversight to be this static in the training process but a great lesson nonetheless!",
    "3117230": "Yeah, those, TS_5_4 in particular, had fewer particles which meant you trained with more if you used it for validation."
  },
  "source": "meta"
}