{
  "id": 561608,
  "title": "79th placed solution [2 Stage Process]",
  "url": "/competitions/czii-cryo-et-object-identification/discussion/561608",
  "author_name": "Olly Powell",
  "post_date": "2025-02-06T22:37:57.214000",
  "votes": 10,
  "comment_count": 0,
  "views": 0,
  "content": "<p>Congrats to all &amp; thanks to the organisers.  This was an interesting competition, I learned a lot from it.  </p>\n<p>I went down the path of trying to pour my computational resources only at the small crops of the larger volume that looked like they had a blob.   Stage 1 was the highest scoring public model, but with relaxed thresholds.  Stage 2 was to crop out 32x32x32 regions of interest for reclassification.</p>\n<p>It didn't work as well as I hoped, but I stuck with it anyway.   My local CV score using F4 metric on pre-processed cropps was in the mid-90's.  But this is irreverent if I was cropping the wrong locations come evaluation time.   </p>\n<p>With hind sight  once I had a couple of decent single-models, for classifying 3D crops,  I should have diverted my effort to the first stage of the process instead of relying on public models.  Since my method boosted the underlying first stage by 0.025, if this carried through to a higher performing first stage, it could have been quite useful.</p>\n<p>Things that worked for me, just not as well as the top solutions:</p>\n<ul>\n<li><p>Flattening the 3D 32x32x32 ROI into a 3-channel image 3x64x64.  Some but not all TIMM image backbones were able to work with such small image sizes.  The models were very fast.  My solution was an ensemble of 21  (7x backbones, 3xTTA)</p></li>\n<li><p>ViT models.  Seemed to generally out-perform CNN architecturs..  Especially MobileViT_xs</p></li>\n<li><p>Pretraining on synthetic data.  Starting with pretrained backbones but randomising the top 8 or so layers.   This really helped stabilise the subsequent fine-tuning on our much smaller quantity of real data.</p></li>\n<li><p>Considering my ensemble to be a single weak model, in addition to UNET &amp; YOLO predictions, all with low thresholds, I predicted a positive if all three agreed.  This was the most effective of many ensembling strategies I came up with.</p></li>\n<li><p>I also predicted a positive if only two model-types were in agreement, if one was a detector and one was my own but with a higher threshold.</p></li>\n</ul>\n<p>I have shared all my code on <a href=\"https://github.com/Wologman/Kaggle_CZII_CryoET\" target=\"_blank\">GitHub</a></p>",
  "messages": [
    {
      "id": 3117356,
      "postDate": "2025-02-06T22:37:57.213Z",
      "content": "<p>Congrats to all &amp; thanks to the organisers.  This was an interesting competition, I learned a lot from it.  </p>\n<p>I went down the path of trying to pour my computational resources only at the small crops of the larger volume that looked like they had a blob.   Stage 1 was the highest scoring public model, but with relaxed thresholds.  Stage 2 was to crop out 32x32x32 regions of interest for reclassification.</p>\n<p>It didn't work as well as I hoped, but I stuck with it anyway.   My local CV score using F4 metric on pre-processed cropps was in the mid-90's.  But this is irreverent if I was cropping the wrong locations come evaluation time.   </p>\n<p>With hind sight  once I had a couple of decent single-models, for classifying 3D crops,  I should have diverted my effort to the first stage of the process instead of relying on public models.  Since my method boosted the underlying first stage by 0.025, if this carried through to a higher performing first stage, it could have been quite useful.</p>\n<p>Things that worked for me, just not as well as the top solutions:</p>\n<ul>\n<li><p>Flattening the 3D 32x32x32 ROI into a 3-channel image 3x64x64.  Some but not all TIMM image backbones were able to work with such small image sizes.  The models were very fast.  My solution was an ensemble of 21  (7x backbones, 3xTTA)</p></li>\n<li><p>ViT models.  Seemed to generally out-perform CNN architecturs..  Especially MobileViT_xs</p></li>\n<li><p>Pretraining on synthetic data.  Starting with pretrained backbones but randomising the top 8 or so layers.   This really helped stabilise the subsequent fine-tuning on our much smaller quantity of real data.</p></li>\n<li><p>Considering my ensemble to be a single weak model, in addition to UNET &amp; YOLO predictions, all with low thresholds, I predicted a positive if all three agreed.  This was the most effective of many ensembling strategies I came up with.</p></li>\n<li><p>I also predicted a positive if only two model-types were in agreement, if one was a detector and one was my own but with a higher threshold.</p></li>\n</ul>\n<p>I have shared all my code on <a href=\"https://github.com/Wologman/Kaggle_CZII_CryoET\" target=\"_blank\">GitHub</a></p>",
      "rawMarkdown": "Congrats to all & thanks to the organisers.  This was an interesting competition, I learned a lot from it.  \n\nI went down the path of trying to pour my computational resources only at the small crops of the larger volume that looked like they had a blob.   Stage 1 was the highest scoring public model, but with relaxed thresholds.  Stage 2 was to crop out 32x32x32 regions of interest for reclassification.\n\nIt didn't work as well as I hoped, but I stuck with it anyway.   My local CV score using F4 metric on pre-processed cropps was in the mid-90's.  But this is irreverent if I was cropping the wrong locations come evaluation time.   \n\nWith hind sight  once I had a couple of decent single-models, for classifying 3D crops,  I should have diverted my effort to the first stage of the process instead of relying on public models.  Since my method boosted the underlying first stage by 0.025, if this carried through to a higher performing first stage, it could have been quite useful.\n\nThings that worked for me, just not as well as the top solutions:\n\n- Flattening the 3D 32x32x32 ROI into a 3-channel image 3x64x64.  Some but not all TIMM image backbones were able to work with such small image sizes.  The models were very fast.  My solution was an ensemble of 21  (7x backbones, 3xTTA)\n\n- ViT models.  Seemed to generally out-perform CNN architecturs..  Especially MobileViT_xs\n\n- Pretraining on synthetic data.  Starting with pretrained backbones but randomising the top 8 or so layers.   This really helped stabilise the subsequent fine-tuning on our much smaller quantity of real data.\n\n- Considering my ensemble to be a single weak model, in addition to UNET & YOLO predictions, all with low thresholds, I predicted a positive if all three agreed.  This was the most effective of many ensembling strategies I came up with.\n\n- I also predicted a positive if only two model-types were in agreement, if one was a detector and one was my own but with a higher threshold.\n\nI have shared all my code on [GitHub](https://github.com/Wologman/Kaggle_CZII_CryoET)",
      "votes": 10
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "3117356": "Congrats to all & thanks to the organisers.  This was an interesting competition, I learned a lot from it.  \n\nI went down the path of trying to pour my computational resources only at the small crops of the larger volume that looked like they had a blob.   Stage 1 was the highest scoring public model, but with relaxed thresholds.  Stage 2 was to crop out 32x32x32 regions of interest for reclassification.\n\nIt didn't work as well as I hoped, but I stuck with it anyway.   My local CV score using F4 metric on pre-processed cropps was in the mid-90's.  But this is irreverent if I was cropping the wrong locations come evaluation time.   \n\nWith hind sight  once I had a couple of decent single-models, for classifying 3D crops,  I should have diverted my effort to the first stage of the process instead of relying on public models.  Since my method boosted the underlying first stage by 0.025, if this carried through to a higher performing first stage, it could have been quite useful.\n\nThings that worked for me, just not as well as the top solutions:\n\n- Flattening the 3D 32x32x32 ROI into a 3-channel image 3x64x64.  Some but not all TIMM image backbones were able to work with such small image sizes.  The models were very fast.  My solution was an ensemble of 21  (7x backbones, 3xTTA)\n\n- ViT models.  Seemed to generally out-perform CNN architecturs..  Especially MobileViT_xs\n\n- Pretraining on synthetic data.  Starting with pretrained backbones but randomising the top 8 or so layers.   This really helped stabilise the subsequent fine-tuning on our much smaller quantity of real data.\n\n- Considering my ensemble to be a single weak model, in addition to UNET & YOLO predictions, all with low thresholds, I predicted a positive if all three agreed.  This was the most effective of many ensembling strategies I came up with.\n\n- I also predicted a positive if only two model-types were in agreement, if one was a detector and one was my own but with a higher threshold.\n\nI have shared all my code on [GitHub](https://github.com/Wologman/Kaggle_CZII_CryoET)"
  }
}