{
  "id": 561440,
  "title": "1st place solution [Object Detection Part] +Train Code +Inference +Models released",
  "url": "/competitions/czii-cryo-et-object-identification/discussion/561440",
  "author_name": "Eugene Khvedchenya",
  "post_date": "2025-02-06T06:09:24.375000",
  "votes": 117,
  "comment_count": 23,
  "views": 0,
  "content": "<pre><code>Throughout the competition there were numerous missile strikes, bombings, and other acts of war that have taken the lives of many innocent people in Ukraine. \nRockets from russia hit within few kilometers from my home in Odesa. Each day Kaggle users from Ukraine facing the chance of not waking up. Just keep in this mind while you read this solution writeup.\n\nI would like to thank the Armed Forces of Ukraine, the Security Service of Ukraine, Defence Intelligence of Ukraine, and the State Emergency Service of Ukraine for providing safety and security to participate in this great competition, complete this work, and help science, technology, and business not to stop but to move forward.\n</code></pre>\n<p><strong>TLDR</strong>: The solution in an ensemble of segmentation (3D Unets with ResNet &amp; B3 encoders) and object detection models (SegResNet and DynUnet backbones) from Monai. <br>\nWe accelerate model inference with NVidia TensorRT to achieve 200% speedup compared to eager PyTorch runtime and leverage two T4 GPUs to run predictions in parallel.</p>\n<p>Kudos to my teammate <a href=\"https://www.kaggle.com/christofhenkel\" target=\"_blank\">@christofhenkel</a> for his great work on ensembling our models into a final submission. Alone, we were able to get into Top-5, but when merged jumped immediately to the Top-1.<br>\nThis writeup covers my part of the solution. Please check out Christof's <a href=\"https://www.kaggle.com/competitions/czii-cryo-et-object-identification/discussion/561510\" target=\"_blank\">writeup</a> for segmentation approach and ensembling technique.</p>\n<p><strong>Update</strong> Training code, models and inference code for OD part released. Check it out here:  </p>\n<ul>\n<li>Training code <a href=\"https://github.com/BloodAxe/Kaggle-2024-CryoET\" target=\"_blank\">https://github.com/BloodAxe/Kaggle-2024-CryoET</a></li>\n<li>Pre-trained models dataset <a href=\"https://www.kaggle.com/datasets/bloodaxe/cryoet-detection-models\" target=\"_blank\">https://www.kaggle.com/datasets/bloodaxe/cryoet-detection-models</a> (Mirror on GitHub: <a href=\"https://github.com/BloodAxe/Kaggle-2024-CryoET/releases/tag/1.0\" target=\"_blank\">https://github.com/BloodAxe/Kaggle-2024-CryoET/releases/tag/1.0</a>)</li>\n<li>[Notebook] Conversion from ONNX to TensorRT engine <a href=\"https://www.kaggle.com/code/bloodaxe/convert-onnx-ensemble-to-tensorrt\" target=\"_blank\">https://www.kaggle.com/code/bloodaxe/convert-onnx-ensemble-to-tensorrt</a></li>\n<li>[Notebook] Inference notebook (OD only) <a href=\"https://www.kaggle.com/code/bloodaxe/v4-inference-notebook/notebook\" target=\"_blank\">https://www.kaggle.com/code/bloodaxe/v4-inference-notebook/notebook</a></li>\n<li>[Dataset] ONNX ensemble <a href=\"https://www.kaggle.com/datasets/bloodaxe/cryoet-onnx-models/\" target=\"_blank\">https://www.kaggle.com/datasets/bloodaxe/cryoet-onnx-models/</a></li>\n</ul>\n<h2>Introduction</h2>\n<p>My baseline was a segmentation (heatmap) 3D Unet built from the SegResNet backbone. As a supervised objective target, I built a five-class heatmap with Gaussian peaks for present ground-truth targets in volume. To train a model I used reduced focal loss from CenterNet paper and NMS postprocessing approach from the same paper. <br>\nAlthough this approach immediately gave ~0.740 LB score (silver zone at the moment of submitting), I ditched it as I anticipated many participants would be using the segmentation approach. It is the easiest to implement, however it doesn't necessarily suit the best for the point detection task (A lot of model's capacity dedicated to predict the curvature of Gaussian at the peak and not on the particles discrimination). Second reason is model diversity. So if I decide to team up in the future, there will be less diversity in the final model ensemble.</p>\n<h2>Modeling approach</h2>\n<p>My final approach uses SegResNet and DynUnet backbones from Monai in combination with a custom point detection head. Based on my experience creating YOLO-based object detection models, I implemented an anchor-free point detection model. Likewise to box detection, model predict class probabilities map and offsets to the center object. Unlike box detection, however there is no concept of bounding box and hence to use IoU between two points I used well-known object keypoint similarity measure <code>exp(-mse(x,y)/ (2 * radius ^ 2))</code> as proxy for IoU metric.</p>\n<p>To train such model I implemented a custom loss function that mimics PP-Yolo loss function with a few modifications:</p>\n<ul>\n<li>Assignment of ground-truth labels to predicted centers based on point-point IoU. Point-Point IoU metric.</li>\n<li>Taking Top-K predictions for each GT label for training</li>\n<li>Computing varifocal loss for class map prediction and IoU-based loss for distance regression</li>\n</ul>\n<p>My point detection approach worked well in terms of accuracy and training speed. On 4x3090 it took a mere two hours to train a single fold. The first submission of a single-fold detection model scored 0.752 on the LB. Once the modeling approach became clear, the next effort was to make it fast.</p>\n<p>In this competition, we need to process 500 scans within a 12-hour time limit.  This translates into ~1m30s per single study. In my models, I predict class-map <code>[B,C,D/2,H/2,W/2]</code> and offsets map <code>[B,3,D/2,H/2,W/2]</code> in stride 2 along depth, height, and width. </p>\n<p>For the object detection approach, there is no need to predict a full-resolution feature maps. Predicting smaller feature maps has a massive impact on model throughput. I've got almost 50% speed increase when changing to stride 2 output from stride 1 (full resolution).</p>\n<p>I ablate on stride 1, stride 2, stride 4, and stride 2 &amp; 4 outputs and found that:</p>\n<ul>\n<li>Stride 1 gives a slightly, slightly better (0.002) LB  score than stride 2</li>\n<li>Stride 4 also worked ok, but stride 2 was better in the F-beta score.</li>\n<li>Using stride 2 and 4 didn't bring any improvements over using only stride 2. An interesting observation when using the two-heads method: Large particles tend to migrate to stride 4 while small classes were present on the class-map with stride 2.</li>\n</ul>\n<p>In all my models I use stride 2 prediction maps.</p>\n<h2>Training</h2>\n<p>I used 5-fold cross-validation scheme for all models. Each fold used 2 studies for validation and 5 for training.</p>\n<p>Training epoch used fixed number of random crops per study and fixed number of random crops around each particle instance. For data augmentations I used random rotations along Z-axis (+- 180 degrees), slight rotations along X and Y axis (+-10), subtle scale jitter (+-5%), and random flips along X, Y, Z axes.<br>\nI experimented with erasing particles from the input volume, random copy-pasting particles from one study to another, doing mixup on the particle instances, but these augmentation techniques didn't bring any improvements in the final model. </p>\n<p>Validation epoch used sliding window approach with the same window size and overlap as in the inference kernel. During validation individual tiles accumulated to final classmap and offsets map and F-beta score was computed on the final maps. After each epoch, I computed per-class thresholds that maximizes F-beta score on the validation set. I saved top-5 models for training experiment which I later averaged which almost always increased the F-beta score.</p>\n<p>Training used 96x128x128px while validation was performed on 192x128x128 volumes.</p>\n<h2>Tiling</h2>\n<p>I used a sliding window approach to tile the input volume. The window size was 192x128x128px with 1x9x9 tiles configuration for input volume of 184x630x630. Individual tiles were aggregated to final classmap and offsets map using a weighted averaging approach where pixels on the border of the tile had lower weights than pixels in the center of the tile. For sliding average aggregation I used my own implementation.</p>\n<h2>Postprocessing</h2>\n<p>Having an aggregated classmap of [C, D, H, W] shape and offsets map of [3, D, H, W] shape, I used the following postprocessing steps:</p>\n<ul>\n<li>Centernet-like NMS to remove duplicate detections: <code>x = x * (x == maxpool3d(x, kernel_size=3, stride=1, padding=1))</code></li>\n<li>Take Top-K predictions for each class (I used 16K but lower numbers should also work fine I guess)</li>\n<li>Filter out predictions by confidence threshold (Each class has it's own confidence threshold)</li>\n<li>Apply greedy NMS using pairwise IoU distance (Similar to how Box NMS works) for detections of a single class.</li>\n<li>Scale pixels to Angstroms</li>\n</ul>\n<p>Since my models were directly predicting offset to the center of the particle, this approach is not sensitive to systematic shift in point coordinates (Some argue it's 0.5px, some claims it to be 1px).</p>\n<h2>Object Detection Ensemble</h2>\n<p>My final object ensemble is 5 models (folds) of OD SegResNet and 5 models of OD DynUnet (OD stands for Object Detection to differentiate them from segmentation-based models). <br>\nA 10 models in total. I used local F-beta metric to select final checkpoints that would make into ensemble.</p>\n<h3>SegResNet</h3>\n<table>\n<thead>\n<tr>\n<th>fold</th>\n<th>score</th>\n<th>AFRT</th>\n<th>BGT</th>\n<th>RBSM</th>\n<th>TRGLB</th>\n<th>VLP</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>0</td>\n<td>0.8457</td>\n<td>0.265</td>\n<td>0.29</td>\n<td>0.195</td>\n<td>0.15</td>\n<td>0.55</td>\n</tr>\n<tr>\n<td>1</td>\n<td>0.8366</td>\n<td>0.345</td>\n<td>0.11</td>\n<td>0.185</td>\n<td>0.135</td>\n<td>0.305</td>\n</tr>\n<tr>\n<td>2</td>\n<td>0.8046</td>\n<td>0.355</td>\n<td>0.28</td>\n<td>0.395</td>\n<td>0.34</td>\n<td>0.185</td>\n</tr>\n<tr>\n<td>3</td>\n<td>0.8398</td>\n<td>0.23</td>\n<td>0.345</td>\n<td>0.405</td>\n<td>0.36</td>\n<td>0.255</td>\n</tr>\n<tr>\n<td>4</td>\n<td>0.8437</td>\n<td>0.165</td>\n<td>0.35</td>\n<td>0.245</td>\n<td>0.235</td>\n<td>0.27</td>\n</tr>\n</tbody>\n</table>\n<h3>DynUnet</h3>\n<table>\n<thead>\n<tr>\n<th>fold</th>\n<th>score</th>\n<th>AFRT</th>\n<th>BGT</th>\n<th>RBSM</th>\n<th>TRGLB</th>\n<th>VLP</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>0</td>\n<td>0.8272</td>\n<td>0.145</td>\n<td>0.33</td>\n<td>0.215</td>\n<td>0.235</td>\n<td>0.405</td>\n</tr>\n<tr>\n<td>1</td>\n<td>0.8337</td>\n<td>0.39</td>\n<td>0.185</td>\n<td>0.215</td>\n<td>0.21</td>\n<td>0.235</td>\n</tr>\n<tr>\n<td>2</td>\n<td>0.7971</td>\n<td>0.425</td>\n<td>0.16</td>\n<td>0.55</td>\n<td>0.275</td>\n<td>0.65</td>\n</tr>\n<tr>\n<td>3</td>\n<td>0.8418</td>\n<td>0.385</td>\n<td>0.345</td>\n<td>0.345</td>\n<td>0.305</td>\n<td>0.21</td>\n</tr>\n<tr>\n<td>4</td>\n<td>0.842</td>\n<td>0.23</td>\n<td>0.34</td>\n<td>0.255</td>\n<td>0.375</td>\n<td>0.14</td>\n</tr>\n</tbody>\n</table>\n<p>As you can see on F-4 scores from individual folds, there is a serious gap between CV and LB scores. Me explanation of this phenomenon is big difference in number of studies in validation set that we use locally and LB. I found that extending validation set by using flipped and 90-degree rotated studies lowers the CV score and also makes F-beta curve mode smooth, which allow finding a more accurate and less noisy threshold.</p>\n<p>To figure out thresholds of the final ensemble I compute OOF F-beta curves for each class and select thresholds that maximize average F-beta score:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1504864%2Fd845bb6bd6eb49c923c2e75c8916f5f5%2Fplot_rotTrue_zFalse_yFalse_xFalse_96x3_96x9_slpaFalse.png?generation=1738822738382933&amp;alt=media\" alt=\"\"></p>\n<h2>Accelerating inference</h2>\n<p>Final submission uses TensorRT for inference.</p>\n<p>Initially, I used <code>torch.jit</code> which was enough for start, but as we teamed up with <a href=\"https://www.kaggle.com/christofhenkel\" target=\"_blank\">@christofhenkel</a> a more efficient approach was needed.<br>\nI tried using <code>onnxruntime</code> with <code>CUDAExecutionProvider</code> which gave me nearly the same throughput as <code>torch.jit</code>.<br>\nFinally, by enabling <code>TensorRTExecutionProvider</code> that leverages all the power of TensorRT I was able to achieve desired speedup of 200% compared to <code>torch.jit</code></p>\n<p>The proces of going from individual checkpoints to TensorRT engine is multi-stage:</p>\n<ol>\n<li>[Offline] Convert individual checkpoints into a single ONNX model containing all models and averaging their predictions</li>\n<li>[Kaggle] Convert ONNX model into TensorRT engine (And save it to disk). This step happens on Kaggle, using T4 GPU (a target GPU we use for inference) and it saved us 10 minutes of submission time for creating TensorRT engine.</li>\n<li>[Kaggle] Actual inference notebook. I split all test data in two chunks and use 2xT4 GPUs to process them in parallel. Then two submission shards were concatenated to form the final submission.</li>\n</ol>\n<p>To combine results of my ensemble with ensemble of <a href=\"https://www.kaggle.com/christofhenkel\" target=\"_blank\">@christofhenkel</a> we simply used weighted average of the predictions on the classmap level followed<br>\nby postprocessing described above. Please check second part of the solution <a href=\"https://www.kaggle.com/competitions/czii-cryo-et-object-identification/discussion/561510\" target=\"_blank\">here</a>.</p>\n<h2>Things that didn't quite work (for me)</h2>\n<ul>\n<li>Mixup, Copy-Paste, Random-Erasing (CV higher, LB lower)</li>\n<li>2.5-D models</li>\n<li>3D version of HRNet and ConvNext</li>\n<li>Gaussian noise, anisotropic scale jitter</li>\n<li>Knowledge distillation</li>\n</ul>\n<h2>References</h2>\n<ol>\n<li>Training code <a href=\"https://github.com/BloodAxe/Kaggle-2024-CryoET\" target=\"_blank\">https://github.com/BloodAxe/Kaggle-2024-CryoET</a></li>\n<li>Pre-trained models dataset <a href=\"https://www.kaggle.com/datasets/bloodaxe/cryoet-detection-models\" target=\"_blank\">https://www.kaggle.com/datasets/bloodaxe/cryoet-detection-models</a> (Mirror on GitHub: <a href=\"https://github.com/BloodAxe/Kaggle-2024-CryoET/releases/tag/1.0\" target=\"_blank\">https://github.com/BloodAxe/Kaggle-2024-CryoET/releases/tag/1.0</a>)</li>\n<li>[Notebook] Conversion from ONNX to TensorRT engine <a href=\"https://www.kaggle.com/code/bloodaxe/convert-onnx-ensemble-to-tensorrt\" target=\"_blank\">https://www.kaggle.com/code/bloodaxe/convert-onnx-ensemble-to-tensorrt</a></li>\n<li>[Notebook] Inference notebook (OD only) <a href=\"https://www.kaggle.com/code/bloodaxe/v4-inference-notebook/notebook\" target=\"_blank\">https://www.kaggle.com/code/bloodaxe/v4-inference-notebook/notebook</a></li>\n<li>[Dataset] ONNX ensemble  <a href=\"https://www.kaggle.com/datasets/bloodaxe/cryoet-onnx-models/\" target=\"_blank\">https://www.kaggle.com/datasets/bloodaxe/cryoet-onnx-models/</a></li>\n</ol>",
  "messages": [
    {
      "id": 3116611,
      "postDate": "2025-02-06T06:09:24.377Z",
      "content": "<pre><code>Throughout the competition there were numerous missile strikes, bombings, and other acts of war that have taken the lives of many innocent people in Ukraine. \nRockets from russia hit within few kilometers from my home in Odesa. Each day Kaggle users from Ukraine facing the chance of not waking up. Just keep in this mind while you read this solution writeup.\n\nI would like to thank the Armed Forces of Ukraine, the Security Service of Ukraine, Defence Intelligence of Ukraine, and the State Emergency Service of Ukraine for providing safety and security to participate in this great competition, complete this work, and help science, technology, and business not to stop but to move forward.\n</code></pre>\n<p><strong>TLDR</strong>: The solution in an ensemble of segmentation (3D Unets with ResNet &amp; B3 encoders) and object detection models (SegResNet and DynUnet backbones) from Monai. <br>\nWe accelerate model inference with NVidia TensorRT to achieve 200% speedup compared to eager PyTorch runtime and leverage two T4 GPUs to run predictions in parallel.</p>\n<p>Kudos to my teammate <a href=\"https://www.kaggle.com/christofhenkel\" target=\"_blank\">@christofhenkel</a> for his great work on ensembling our models into a final submission. Alone, we were able to get into Top-5, but when merged jumped immediately to the Top-1.<br>\nThis writeup covers my part of the solution. Please check out Christof's <a href=\"https://www.kaggle.com/competitions/czii-cryo-et-object-identification/discussion/561510\" target=\"_blank\">writeup</a> for segmentation approach and ensembling technique.</p>\n<p><strong>Update</strong> Training code, models and inference code for OD part released. Check it out here:  </p>\n<ul>\n<li>Training code <a href=\"https://github.com/BloodAxe/Kaggle-2024-CryoET\" target=\"_blank\">https://github.com/BloodAxe/Kaggle-2024-CryoET</a></li>\n<li>Pre-trained models dataset <a href=\"https://www.kaggle.com/datasets/bloodaxe/cryoet-detection-models\" target=\"_blank\">https://www.kaggle.com/datasets/bloodaxe/cryoet-detection-models</a> (Mirror on GitHub: <a href=\"https://github.com/BloodAxe/Kaggle-2024-CryoET/releases/tag/1.0\" target=\"_blank\">https://github.com/BloodAxe/Kaggle-2024-CryoET/releases/tag/1.0</a>)</li>\n<li>[Notebook] Conversion from ONNX to TensorRT engine <a href=\"https://www.kaggle.com/code/bloodaxe/convert-onnx-ensemble-to-tensorrt\" target=\"_blank\">https://www.kaggle.com/code/bloodaxe/convert-onnx-ensemble-to-tensorrt</a></li>\n<li>[Notebook] Inference notebook (OD only) <a href=\"https://www.kaggle.com/code/bloodaxe/v4-inference-notebook/notebook\" target=\"_blank\">https://www.kaggle.com/code/bloodaxe/v4-inference-notebook/notebook</a></li>\n<li>[Dataset] ONNX ensemble <a href=\"https://www.kaggle.com/datasets/bloodaxe/cryoet-onnx-models/\" target=\"_blank\">https://www.kaggle.com/datasets/bloodaxe/cryoet-onnx-models/</a></li>\n</ul>\n<h2>Introduction</h2>\n<p>My baseline was a segmentation (heatmap) 3D Unet built from the SegResNet backbone. As a supervised objective target, I built a five-class heatmap with Gaussian peaks for present ground-truth targets in volume. To train a model I used reduced focal loss from CenterNet paper and NMS postprocessing approach from the same paper. <br>\nAlthough this approach immediately gave ~0.740 LB score (silver zone at the moment of submitting), I ditched it as I anticipated many participants would be using the segmentation approach. It is the easiest to implement, however it doesn't necessarily suit the best for the point detection task (A lot of model's capacity dedicated to predict the curvature of Gaussian at the peak and not on the particles discrimination). Second reason is model diversity. So if I decide to team up in the future, there will be less diversity in the final model ensemble.</p>\n<h2>Modeling approach</h2>\n<p>My final approach uses SegResNet and DynUnet backbones from Monai in combination with a custom point detection head. Based on my experience creating YOLO-based object detection models, I implemented an anchor-free point detection model. Likewise to box detection, model predict class probabilities map and offsets to the center object. Unlike box detection, however there is no concept of bounding box and hence to use IoU between two points I used well-known object keypoint similarity measure <code>exp(-mse(x,y)/ (2 * radius ^ 2))</code> as proxy for IoU metric.</p>\n<p>To train such model I implemented a custom loss function that mimics PP-Yolo loss function with a few modifications:</p>\n<ul>\n<li>Assignment of ground-truth labels to predicted centers based on point-point IoU. Point-Point IoU metric.</li>\n<li>Taking Top-K predictions for each GT label for training</li>\n<li>Computing varifocal loss for class map prediction and IoU-based loss for distance regression</li>\n</ul>\n<p>My point detection approach worked well in terms of accuracy and training speed. On 4x3090 it took a mere two hours to train a single fold. The first submission of a single-fold detection model scored 0.752 on the LB. Once the modeling approach became clear, the next effort was to make it fast.</p>\n<p>In this competition, we need to process 500 scans within a 12-hour time limit.  This translates into ~1m30s per single study. In my models, I predict class-map <code>[B,C,D/2,H/2,W/2]</code> and offsets map <code>[B,3,D/2,H/2,W/2]</code> in stride 2 along depth, height, and width. </p>\n<p>For the object detection approach, there is no need to predict a full-resolution feature maps. Predicting smaller feature maps has a massive impact on model throughput. I've got almost 50% speed increase when changing to stride 2 output from stride 1 (full resolution).</p>\n<p>I ablate on stride 1, stride 2, stride 4, and stride 2 &amp; 4 outputs and found that:</p>\n<ul>\n<li>Stride 1 gives a slightly, slightly better (0.002) LB  score than stride 2</li>\n<li>Stride 4 also worked ok, but stride 2 was better in the F-beta score.</li>\n<li>Using stride 2 and 4 didn't bring any improvements over using only stride 2. An interesting observation when using the two-heads method: Large particles tend to migrate to stride 4 while small classes were present on the class-map with stride 2.</li>\n</ul>\n<p>In all my models I use stride 2 prediction maps.</p>\n<h2>Training</h2>\n<p>I used 5-fold cross-validation scheme for all models. Each fold used 2 studies for validation and 5 for training.</p>\n<p>Training epoch used fixed number of random crops per study and fixed number of random crops around each particle instance. For data augmentations I used random rotations along Z-axis (+- 180 degrees), slight rotations along X and Y axis (+-10), subtle scale jitter (+-5%), and random flips along X, Y, Z axes.<br>\nI experimented with erasing particles from the input volume, random copy-pasting particles from one study to another, doing mixup on the particle instances, but these augmentation techniques didn't bring any improvements in the final model. </p>\n<p>Validation epoch used sliding window approach with the same window size and overlap as in the inference kernel. During validation individual tiles accumulated to final classmap and offsets map and F-beta score was computed on the final maps. After each epoch, I computed per-class thresholds that maximizes F-beta score on the validation set. I saved top-5 models for training experiment which I later averaged which almost always increased the F-beta score.</p>\n<p>Training used 96x128x128px while validation was performed on 192x128x128 volumes.</p>\n<h2>Tiling</h2>\n<p>I used a sliding window approach to tile the input volume. The window size was 192x128x128px with 1x9x9 tiles configuration for input volume of 184x630x630. Individual tiles were aggregated to final classmap and offsets map using a weighted averaging approach where pixels on the border of the tile had lower weights than pixels in the center of the tile. For sliding average aggregation I used my own implementation.</p>\n<h2>Postprocessing</h2>\n<p>Having an aggregated classmap of [C, D, H, W] shape and offsets map of [3, D, H, W] shape, I used the following postprocessing steps:</p>\n<ul>\n<li>Centernet-like NMS to remove duplicate detections: <code>x = x * (x == maxpool3d(x, kernel_size=3, stride=1, padding=1))</code></li>\n<li>Take Top-K predictions for each class (I used 16K but lower numbers should also work fine I guess)</li>\n<li>Filter out predictions by confidence threshold (Each class has it's own confidence threshold)</li>\n<li>Apply greedy NMS using pairwise IoU distance (Similar to how Box NMS works) for detections of a single class.</li>\n<li>Scale pixels to Angstroms</li>\n</ul>\n<p>Since my models were directly predicting offset to the center of the particle, this approach is not sensitive to systematic shift in point coordinates (Some argue it's 0.5px, some claims it to be 1px).</p>\n<h2>Object Detection Ensemble</h2>\n<p>My final object ensemble is 5 models (folds) of OD SegResNet and 5 models of OD DynUnet (OD stands for Object Detection to differentiate them from segmentation-based models). <br>\nA 10 models in total. I used local F-beta metric to select final checkpoints that would make into ensemble.</p>\n<h3>SegResNet</h3>\n<table>\n<thead>\n<tr>\n<th>fold</th>\n<th>score</th>\n<th>AFRT</th>\n<th>BGT</th>\n<th>RBSM</th>\n<th>TRGLB</th>\n<th>VLP</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>0</td>\n<td>0.8457</td>\n<td>0.265</td>\n<td>0.29</td>\n<td>0.195</td>\n<td>0.15</td>\n<td>0.55</td>\n</tr>\n<tr>\n<td>1</td>\n<td>0.8366</td>\n<td>0.345</td>\n<td>0.11</td>\n<td>0.185</td>\n<td>0.135</td>\n<td>0.305</td>\n</tr>\n<tr>\n<td>2</td>\n<td>0.8046</td>\n<td>0.355</td>\n<td>0.28</td>\n<td>0.395</td>\n<td>0.34</td>\n<td>0.185</td>\n</tr>\n<tr>\n<td>3</td>\n<td>0.8398</td>\n<td>0.23</td>\n<td>0.345</td>\n<td>0.405</td>\n<td>0.36</td>\n<td>0.255</td>\n</tr>\n<tr>\n<td>4</td>\n<td>0.8437</td>\n<td>0.165</td>\n<td>0.35</td>\n<td>0.245</td>\n<td>0.235</td>\n<td>0.27</td>\n</tr>\n</tbody>\n</table>\n<h3>DynUnet</h3>\n<table>\n<thead>\n<tr>\n<th>fold</th>\n<th>score</th>\n<th>AFRT</th>\n<th>BGT</th>\n<th>RBSM</th>\n<th>TRGLB</th>\n<th>VLP</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>0</td>\n<td>0.8272</td>\n<td>0.145</td>\n<td>0.33</td>\n<td>0.215</td>\n<td>0.235</td>\n<td>0.405</td>\n</tr>\n<tr>\n<td>1</td>\n<td>0.8337</td>\n<td>0.39</td>\n<td>0.185</td>\n<td>0.215</td>\n<td>0.21</td>\n<td>0.235</td>\n</tr>\n<tr>\n<td>2</td>\n<td>0.7971</td>\n<td>0.425</td>\n<td>0.16</td>\n<td>0.55</td>\n<td>0.275</td>\n<td>0.65</td>\n</tr>\n<tr>\n<td>3</td>\n<td>0.8418</td>\n<td>0.385</td>\n<td>0.345</td>\n<td>0.345</td>\n<td>0.305</td>\n<td>0.21</td>\n</tr>\n<tr>\n<td>4</td>\n<td>0.842</td>\n<td>0.23</td>\n<td>0.34</td>\n<td>0.255</td>\n<td>0.375</td>\n<td>0.14</td>\n</tr>\n</tbody>\n</table>\n<p>As you can see on F-4 scores from individual folds, there is a serious gap between CV and LB scores. Me explanation of this phenomenon is big difference in number of studies in validation set that we use locally and LB. I found that extending validation set by using flipped and 90-degree rotated studies lowers the CV score and also makes F-beta curve mode smooth, which allow finding a more accurate and less noisy threshold.</p>\n<p>To figure out thresholds of the final ensemble I compute OOF F-beta curves for each class and select thresholds that maximize average F-beta score:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1504864%2Fd845bb6bd6eb49c923c2e75c8916f5f5%2Fplot_rotTrue_zFalse_yFalse_xFalse_96x3_96x9_slpaFalse.png?generation=1738822738382933&amp;alt=media\" alt=\"\"></p>\n<h2>Accelerating inference</h2>\n<p>Final submission uses TensorRT for inference.</p>\n<p>Initially, I used <code>torch.jit</code> which was enough for start, but as we teamed up with <a href=\"https://www.kaggle.com/christofhenkel\" target=\"_blank\">@christofhenkel</a> a more efficient approach was needed.<br>\nI tried using <code>onnxruntime</code> with <code>CUDAExecutionProvider</code> which gave me nearly the same throughput as <code>torch.jit</code>.<br>\nFinally, by enabling <code>TensorRTExecutionProvider</code> that leverages all the power of TensorRT I was able to achieve desired speedup of 200% compared to <code>torch.jit</code></p>\n<p>The proces of going from individual checkpoints to TensorRT engine is multi-stage:</p>\n<ol>\n<li>[Offline] Convert individual checkpoints into a single ONNX model containing all models and averaging their predictions</li>\n<li>[Kaggle] Convert ONNX model into TensorRT engine (And save it to disk). This step happens on Kaggle, using T4 GPU (a target GPU we use for inference) and it saved us 10 minutes of submission time for creating TensorRT engine.</li>\n<li>[Kaggle] Actual inference notebook. I split all test data in two chunks and use 2xT4 GPUs to process them in parallel. Then two submission shards were concatenated to form the final submission.</li>\n</ol>\n<p>To combine results of my ensemble with ensemble of <a href=\"https://www.kaggle.com/christofhenkel\" target=\"_blank\">@christofhenkel</a> we simply used weighted average of the predictions on the classmap level followed<br>\nby postprocessing described above. Please check second part of the solution <a href=\"https://www.kaggle.com/competitions/czii-cryo-et-object-identification/discussion/561510\" target=\"_blank\">here</a>.</p>\n<h2>Things that didn't quite work (for me)</h2>\n<ul>\n<li>Mixup, Copy-Paste, Random-Erasing (CV higher, LB lower)</li>\n<li>2.5-D models</li>\n<li>3D version of HRNet and ConvNext</li>\n<li>Gaussian noise, anisotropic scale jitter</li>\n<li>Knowledge distillation</li>\n</ul>\n<h2>References</h2>\n<ol>\n<li>Training code <a href=\"https://github.com/BloodAxe/Kaggle-2024-CryoET\" target=\"_blank\">https://github.com/BloodAxe/Kaggle-2024-CryoET</a></li>\n<li>Pre-trained models dataset <a href=\"https://www.kaggle.com/datasets/bloodaxe/cryoet-detection-models\" target=\"_blank\">https://www.kaggle.com/datasets/bloodaxe/cryoet-detection-models</a> (Mirror on GitHub: <a href=\"https://github.com/BloodAxe/Kaggle-2024-CryoET/releases/tag/1.0\" target=\"_blank\">https://github.com/BloodAxe/Kaggle-2024-CryoET/releases/tag/1.0</a>)</li>\n<li>[Notebook] Conversion from ONNX to TensorRT engine <a href=\"https://www.kaggle.com/code/bloodaxe/convert-onnx-ensemble-to-tensorrt\" target=\"_blank\">https://www.kaggle.com/code/bloodaxe/convert-onnx-ensemble-to-tensorrt</a></li>\n<li>[Notebook] Inference notebook (OD only) <a href=\"https://www.kaggle.com/code/bloodaxe/v4-inference-notebook/notebook\" target=\"_blank\">https://www.kaggle.com/code/bloodaxe/v4-inference-notebook/notebook</a></li>\n<li>[Dataset] ONNX ensemble  <a href=\"https://www.kaggle.com/datasets/bloodaxe/cryoet-onnx-models/\" target=\"_blank\">https://www.kaggle.com/datasets/bloodaxe/cryoet-onnx-models/</a></li>\n</ol>",
      "rawMarkdown": "```markdown\nThroughout the competition there were numerous missile strikes, bombings, and other acts of war that have taken the lives of many innocent people in Ukraine. \nRockets from russia hit within few kilometers from my home in Odesa. Each day Kaggle users from Ukraine facing the chance of not waking up. Just keep in this mind while you read this solution writeup.\n\nI would like to thank the Armed Forces of Ukraine, the Security Service of Ukraine, Defence Intelligence of Ukraine, and the State Emergency Service of Ukraine for providing safety and security to participate in this great competition, complete this work, and help science, technology, and business not to stop but to move forward.\n```\n\n**TLDR**: The solution in an ensemble of segmentation (3D Unets with ResNet & B3 encoders) and object detection models (SegResNet and DynUnet backbones) from Monai. \nWe accelerate model inference with NVidia TensorRT to achieve 200% speedup compared to eager PyTorch runtime and leverage two T4 GPUs to run predictions in parallel.\n\nKudos to my teammate @christofhenkel for his great work on ensembling our models into a final submission. Alone, we were able to get into Top-5, but when merged jumped immediately to the Top-1.\nThis writeup covers my part of the solution. Please check out Christof's [writeup](https://www.kaggle.com/competitions/czii-cryo-et-object-identification/discussion/561510) for segmentation approach and ensembling technique.\n\n**Update** Training code, models and inference code for OD part released. Check it out here:  \n\n- Training code https://github.com/BloodAxe/Kaggle-2024-CryoET\n- Pre-trained models dataset https://www.kaggle.com/datasets/bloodaxe/cryoet-detection-models (Mirror on GitHub: https://github.com/BloodAxe/Kaggle-2024-CryoET/releases/tag/1.0)\n- [Notebook] Conversion from ONNX to TensorRT engine https://www.kaggle.com/code/bloodaxe/convert-onnx-ensemble-to-tensorrt\n- [Notebook] Inference notebook (OD only) https://www.kaggle.com/code/bloodaxe/v4-inference-notebook/notebook\n- [Dataset] ONNX ensemble https://www.kaggle.com/datasets/bloodaxe/cryoet-onnx-models/\n\n## Introduction\n\nMy baseline was a segmentation (heatmap) 3D Unet built from the SegResNet backbone. As a supervised objective target, I built a five-class heatmap with Gaussian peaks for present ground-truth targets in volume. To train a model I used reduced focal loss from CenterNet paper and NMS postprocessing approach from the same paper. \nAlthough this approach immediately gave ~0.740 LB score (silver zone at the moment of submitting), I ditched it as I anticipated many participants would be using the segmentation approach. It is the easiest to implement, however it doesn't necessarily suit the best for the point detection task (A lot of model's capacity dedicated to predict the curvature of Gaussian at the peak and not on the particles discrimination). Second reason is model diversity. So if I decide to team up in the future, there will be less diversity in the final model ensemble.\n\n## Modeling approach \n\nMy final approach uses SegResNet and DynUnet backbones from Monai in combination with a custom point detection head. Based on my experience creating YOLO-based object detection models, I implemented an anchor-free point detection model. Likewise to box detection, model predict class probabilities map and offsets to the center object. Unlike box detection, however there is no concept of bounding box and hence to use IoU between two points I used well-known object keypoint similarity measure `exp(-mse(x,y)/ (2 * radius ^ 2))` as proxy for IoU metric.\n\nTo train such model I implemented a custom loss function that mimics PP-Yolo loss function with a few modifications:\n* Assignment of ground-truth labels to predicted centers based on point-point IoU. Point-Point IoU metric.\n* Taking Top-K predictions for each GT label for training\n* Computing varifocal loss for class map prediction and IoU-based loss for distance regression\n\nMy point detection approach worked well in terms of accuracy and training speed. On 4x3090 it took a mere two hours to train a single fold. The first submission of a single-fold detection model scored 0.752 on the LB. Once the modeling approach became clear, the next effort was to make it fast.\n\nIn this competition, we need to process 500 scans within a 12-hour time limit.  This translates into ~1m30s per single study. In my models, I predict class-map `[B,C,D/2,H/2,W/2]` and offsets map `[B,3,D/2,H/2,W/2]` in stride 2 along depth, height, and width. \n\nFor the object detection approach, there is no need to predict a full-resolution feature maps. Predicting smaller feature maps has a massive impact on model throughput. I've got almost 50% speed increase when changing to stride 2 output from stride 1 (full resolution).\n\nI ablate on stride 1, stride 2, stride 4, and stride 2 & 4 outputs and found that:\n* Stride 1 gives a slightly, slightly better (0.002) LB  score than stride 2\n* Stride 4 also worked ok, but stride 2 was better in the F-beta score.\n* Using stride 2 and 4 didn't bring any improvements over using only stride 2. An interesting observation when using the two-heads method: Large particles tend to migrate to stride 4 while small classes were present on the class-map with stride 2.\n\nIn all my models I use stride 2 prediction maps.\n\n## Training\n\nI used 5-fold cross-validation scheme for all models. Each fold used 2 studies for validation and 5 for training.\n\nTraining epoch used fixed number of random crops per study and fixed number of random crops around each particle instance. For data augmentations I used random rotations along Z-axis (+- 180 degrees), slight rotations along X and Y axis (+-10), subtle scale jitter (+-5%), and random flips along X, Y, Z axes.\nI experimented with erasing particles from the input volume, random copy-pasting particles from one study to another, doing mixup on the particle instances, but these augmentation techniques didn't bring any improvements in the final model. \n\nValidation epoch used sliding window approach with the same window size and overlap as in the inference kernel. During validation individual tiles accumulated to final classmap and offsets map and F-beta score was computed on the final maps. After each epoch, I computed per-class thresholds that maximizes F-beta score on the validation set. I saved top-5 models for training experiment which I later averaged which almost always increased the F-beta score.\n\nTraining used 96x128x128px while validation was performed on 192x128x128 volumes.\n\n## Tiling\n\nI used a sliding window approach to tile the input volume. The window size was 192x128x128px with 1x9x9 tiles configuration for input volume of 184x630x630. Individual tiles were aggregated to final classmap and offsets map using a weighted averaging approach where pixels on the border of the tile had lower weights than pixels in the center of the tile. For sliding average aggregation I used my own implementation.\n\n## Postprocessing\n\nHaving an aggregated classmap of [C, D, H, W] shape and offsets map of [3, D, H, W] shape, I used the following postprocessing steps:\n\n- Centernet-like NMS to remove duplicate detections: `x = x * (x == maxpool3d(x, kernel_size=3, stride=1, padding=1))`\n- Take Top-K predictions for each class (I used 16K but lower numbers should also work fine I guess)\n- Filter out predictions by confidence threshold (Each class has it's own confidence threshold)\n- Apply greedy NMS using pairwise IoU distance (Similar to how Box NMS works) for detections of a single class.\n- Scale pixels to Angstroms\n\nSince my models were directly predicting offset to the center of the particle, this approach is not sensitive to systematic shift in point coordinates (Some argue it's 0.5px, some claims it to be 1px).\n\n## Object Detection Ensemble\n\nMy final object ensemble is 5 models (folds) of OD SegResNet and 5 models of OD DynUnet (OD stands for Object Detection to differentiate them from segmentation-based models). \nA 10 models in total. I used local F-beta metric to select final checkpoints that would make into ensemble.\n\n### SegResNet\n\n| fold         |    score |   AFRT |   BGT |   RBSM |   TRGLB |   VLP |\n|:-------------|---------:|-------:|------:|-------:|--------:|------:|\n| 0            | 0.8457   |  0.265 | 0.29  |  0.195 |   0.15  | 0.55  |\n| 1            | 0.8366   |  0.345 | 0.11  |  0.185 |   0.135 | 0.305 |\n| 2            | 0.8046   |  0.355 | 0.28  |  0.395 |   0.34  | 0.185 |\n| 3            | 0.8398   |  0.23  | 0.345 |  0.405 |   0.36  | 0.255 |\n| 4            | 0.8437   |  0.165 | 0.35  |  0.245 |   0.235 | 0.27  |\n\n### DynUnet\n\n| fold         |    score |   AFRT |   BGT |   RBSM |   TRGLB |   VLP |\n|:-------------|---------:|-------:|------:|-------:|--------:|------:|\n| 0            | 0.8272   |  0.145 | 0.33  |  0.215 |   0.235 | 0.405 |\n| 1            | 0.8337   |  0.39  | 0.185 |  0.215 |   0.21  | 0.235 |\n| 2            | 0.7971   |  0.425 | 0.16  |  0.55  |   0.275 | 0.65  |\n| 3            | 0.8418   |  0.385 | 0.345 |  0.345 |   0.305 | 0.21  |\n| 4            | 0.842    |  0.23  | 0.34  |  0.255 |   0.375 | 0.14  |\n\nAs you can see on F-4 scores from individual folds, there is a serious gap between CV and LB scores. Me explanation of this phenomenon is big difference in number of studies in validation set that we use locally and LB. I found that extending validation set by using flipped and 90-degree rotated studies lowers the CV score and also makes F-beta curve mode smooth, which allow finding a more accurate and less noisy threshold.\n\nTo figure out thresholds of the final ensemble I compute OOF F-beta curves for each class and select thresholds that maximize average F-beta score:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1504864%2Fd845bb6bd6eb49c923c2e75c8916f5f5%2Fplot_rotTrue_zFalse_yFalse_xFalse_96x3_96x9_slpaFalse.png?generation=1738822738382933&alt=media)\n\n## Accelerating inference\n\nFinal submission uses TensorRT for inference.\n\nInitially, I used `torch.jit` which was enough for start, but as we teamed up with @christofhenkel a more efficient approach was needed.\nI tried using `onnxruntime` with `CUDAExecutionProvider` which gave me nearly the same throughput as `torch.jit`.\nFinally, by enabling `TensorRTExecutionProvider` that leverages all the power of TensorRT I was able to achieve desired speedup of 200% compared to `torch.jit`\n\nThe proces of going from individual checkpoints to TensorRT engine is multi-stage:\n\n1. [Offline] Convert individual checkpoints into a single ONNX model containing all models and averaging their predictions\n2. [Kaggle] Convert ONNX model into TensorRT engine (And save it to disk). This step happens on Kaggle, using T4 GPU (a target GPU we use for inference) and it saved us 10 minutes of submission time for creating TensorRT engine.\n3. [Kaggle] Actual inference notebook. I split all test data in two chunks and use 2xT4 GPUs to process them in parallel. Then two submission shards were concatenated to form the final submission.\n\nTo combine results of my ensemble with ensemble of @christofhenkel we simply used weighted average of the predictions on the classmap level followed\nby postprocessing described above. Please check second part of the solution [here](https://www.kaggle.com/competitions/czii-cryo-et-object-identification/discussion/561510).\n\n## Things that didn't quite work (for me)\n\n* Mixup, Copy-Paste, Random-Erasing (CV higher, LB lower)\n* 2.5-D models\n* 3D version of HRNet and ConvNext\n* Gaussian noise, anisotropic scale jitter\n* Knowledge distillation\n\n## References\n\n1. Training code https://github.com/BloodAxe/Kaggle-2024-CryoET\n2. Pre-trained models dataset https://www.kaggle.com/datasets/bloodaxe/cryoet-detection-models (Mirror on GitHub: https://github.com/BloodAxe/Kaggle-2024-CryoET/releases/tag/1.0)\n3. [Notebook] Conversion from ONNX to TensorRT engine https://www.kaggle.com/code/bloodaxe/convert-onnx-ensemble-to-tensorrt\n4. [Notebook] Inference notebook (OD only) https://www.kaggle.com/code/bloodaxe/v4-inference-notebook/notebook\n5. [Dataset] ONNX ensemble  https://www.kaggle.com/datasets/bloodaxe/cryoet-onnx-models/\n",
      "votes": 117
    },
    {
      "id": 3117046,
      "postDate": "2025-02-06T15:27:15.530Z",
      "content": "<p>Training code for OD models released! Check it here <a href=\"https://github.com/BloodAxe/Kaggle-2024-CryoET\" target=\"_blank\">https://github.com/BloodAxe/Kaggle-2024-CryoET</a></p>",
      "rawMarkdown": "Training code for OD models released! Check it here https://github.com/BloodAxe/Kaggle-2024-CryoET",
      "votes": 5,
      "replies": [
        {
          "id": 3122139,
          "postDate": "2025-02-12T10:15:37.800Z",
          "content": "<p>okiiee got it </p>",
          "rawMarkdown": "okiiee got it \n",
          "isDeleted": true
        }
      ]
    },
    {
      "id": 3117879,
      "postDate": "2025-02-07T10:58:27.523Z",
      "content": "<p>Congrats, and thank you for sharing your solution - so much for me to learn from here. </p>\n<p>My sincere support for the people of Ukraine, coming from Australia</p>",
      "rawMarkdown": "Congrats, and thank you for sharing your solution - so much for me to learn from here. \n\nMy sincere support for the people of Ukraine, coming from Australia",
      "votes": 3,
      "replies": [
        {
          "id": 3117882,
          "postDate": "2025-02-07T11:00:57.660Z",
          "content": "<p>Thanks for your kind words! I learned myself from winning solutions of past years and I guess it’s my turn to share the knowledge now :) </p>",
          "rawMarkdown": "Thanks for your kind words! I learned myself from winning solutions of past years and I guess it’s my turn to share the knowledge now :) ",
          "votes": 3
        }
      ]
    },
    {
      "id": 3117546,
      "postDate": "2025-02-07T03:53:05.467Z",
      "content": "<p>Congrats. Your insights are helpful</p>",
      "rawMarkdown": "Congrats. Your insights are helpful",
      "votes": 2
    },
    {
      "id": 3122380,
      "postDate": "2025-02-12T15:15:48.567Z",
      "content": "<p>congrats nicely done!</p>",
      "rawMarkdown": "congrats nicely done!"
    },
    {
      "id": 3122180,
      "postDate": "2025-02-12T10:59:53.777Z",
      "content": "<p>Congrats! Thank you for sharing a nice write-up and solution code !!!</p>",
      "rawMarkdown": "Congrats! Thank you for sharing a nice write-up and solution code !!!"
    },
    {
      "id": 3121687,
      "postDate": "2025-02-11T20:30:57.507Z",
      "content": "<p>yoyo is the best game</p>",
      "rawMarkdown": "yoyo is the best game"
    },
    {
      "id": 3121343,
      "postDate": "2025-02-11T13:52:43.420Z",
      "content": "<p>Congrats!  Thank you for sharing a well structured write up and your code.</p>\n<p>I found your approach quite interesting! Could you point me to the custom loss function used in your anchor-free point detection model? I'm curious to understand how it contributes to the overall performance.</p>\n<p>Looking forward to your insights! </p>",
      "rawMarkdown": "Congrats!  Thank you for sharing a well structured write up and your code.\n\nI found your approach quite interesting! Could you point me to the custom loss function used in your anchor-free point detection model? I'm curious to understand how it contributes to the overall performance.\n\nLooking forward to your insights! ",
      "replies": [
        {
          "id": 3123029,
          "postDate": "2025-02-13T09:32:50.977Z",
          "content": "<p>Sure, it can be found here <a href=\"https://github.com/BloodAxe/Kaggle-2024-CryoET/blob/master/cryoet/modelling/detection/functional.py#L122\" target=\"_blank\">https://github.com/BloodAxe/Kaggle-2024-CryoET/blob/master/cryoet/modelling/detection/functional.py#L122</a></p>",
          "rawMarkdown": "Sure, it can be found here https://github.com/BloodAxe/Kaggle-2024-CryoET/blob/master/cryoet/modelling/detection/functional.py#L122\n\n"
        }
      ]
    },
    {
      "id": 3121309,
      "postDate": "2025-02-11T13:16:47.510Z",
      "content": "<p>woww greattttttttttttttttttt</p>",
      "rawMarkdown": "woww greattttttttttttttttttt\n"
    },
    {
      "id": 3121217,
      "postDate": "2025-02-11T11:26:32.770Z",
      "content": "<p>Thanks for the write-up.<br>\nYou mention that MixUp didn't work for the object detection part but it seems that it has worked for the segmentation part.<br>\nDo you have any ideas why? </p>",
      "rawMarkdown": "Thanks for the write-up.\nYou mention that MixUp didn't work for the object detection part but it seems that it has worked for the segmentation part.\nDo you have any ideas why? ",
      "replies": [
        {
          "id": 3123030,
          "postDate": "2025-02-13T09:37:26.290Z",
          "content": "<p>It's a good question to ask. Frankly - I don't know. I imagine if we have at least 100 studies in train, the CV estimate would be less noisy and we could obtain more reliable signals on what work and what doesn't.</p>\n<p>The metric is very sensitive to confidence thresholds you choose, so LB validation was the only viable approach to tune thresholds, given the size of studies in public part of the LB. I don't quite like overfitting to public, but in this comp it was the only viable choice. Under these conditions, it is really hard to say confidently that one approach works or doesn't since you may be just in poor thresholds area and moving one threshold to +0.05 and another -0.05 would give you SOTA. I tried playing with thresholds a bit for mixup models but didn't spent more than 5 subs on finding best combination of thresholds. </p>",
          "rawMarkdown": "It's a good question to ask. Frankly - I don't know. I imagine if we have at least 100 studies in train, the CV estimate would be less noisy and we could obtain more reliable signals on what work and what doesn't.\n\nThe metric is very sensitive to confidence thresholds you choose, so LB validation was the only viable approach to tune thresholds, given the size of studies in public part of the LB. I don't quite like overfitting to public, but in this comp it was the only viable choice. Under these conditions, it is really hard to say confidently that one approach works or doesn't since you may be just in poor thresholds area and moving one threshold to +0.05 and another -0.05 would give you SOTA. I tried playing with thresholds a bit for mixup models but didn't spent more than 5 subs on finding best combination of thresholds. ",
          "votes": 1
        }
      ]
    },
    {
      "id": 3120466,
      "postDate": "2025-02-10T14:59:30.170Z",
      "content": "<p>I went over your code to study and it's just amazing! There are things that I missed because of lack of my knowledge about object detection but it was easy to follow and very well written. Thank you!</p>",
      "rawMarkdown": "I went over your code to study and it's just amazing! There are things that I missed because of lack of my knowledge about object detection but it was easy to follow and very well written. Thank you!"
    },
    {
      "id": 3119828,
      "postDate": "2025-02-09T18:03:17.720Z",
      "content": "<p>Huge congratulations, <a href=\"https://www.kaggle.com/bloodaxe\" target=\"_blank\">@bloodaxe</a> and <a href=\"https://www.kaggle.com/christofhenkel\" target=\"_blank\">@christofhenkel</a>! Your write-up is an absolute masterclass in efficient deep learning inference and ensembling techniques. The combination of SegResNet and DynUnet backbones, the anchor-free point detection approach, and your custom loss function inspired by PP-YOLO really showcases the depth of optimization that goes into top-tier solutions.</p>",
      "rawMarkdown": "Huge congratulations, @bloodaxe and @christofhenkel! Your write-up is an absolute masterclass in efficient deep learning inference and ensembling techniques. The combination of SegResNet and DynUnet backbones, the anchor-free point detection approach, and your custom loss function inspired by PP-YOLO really showcases the depth of optimization that goes into top-tier solutions."
    },
    {
      "id": 3119575,
      "postDate": "2025-02-09T12:10:31.190Z",
      "content": "<p>Awesome work with parallel computing the processing time and running both T4's! Thanks for sharing this.</p>",
      "rawMarkdown": "Awesome work with parallel computing the processing time and running both T4's! Thanks for sharing this.\n"
    },
    {
      "id": 3117293,
      "postDate": "2025-02-06T20:14:33.693Z",
      "content": "<p>Congrats! Thank you for sharing a nice write-up and your code.</p>\n<p>Could you point me to the custom loss function used in your \"anchor-free point detection model\"?</p>",
      "rawMarkdown": "Congrats! Thank you for sharing a nice write-up and your code.\n\nCould you point me to the custom loss function used in your \"anchor-free point detection model\"?",
      "replies": [
        {
          "id": 3117295,
          "postDate": "2025-02-06T20:19:15.120Z",
          "content": "<p>You can start from main loss function and dig as deep as you want from there: <a href=\"https://github.com/BloodAxe/Kaggle-2024-CryoET/blob/master/cryoet/modelling/detection/functional.py#L122\" target=\"_blank\">https://github.com/BloodAxe/Kaggle-2024-CryoET/blob/master/cryoet/modelling/detection/functional.py#L122</a></p>",
          "rawMarkdown": "You can start from main loss function and dig as deep as you want from there: https://github.com/BloodAxe/Kaggle-2024-CryoET/blob/master/cryoet/modelling/detection/functional.py#L122",
          "votes": 3
        }
      ]
    },
    {
      "id": 3116826,
      "postDate": "2025-02-06T10:53:50.597Z",
      "content": "<p>Congrats! Is there any chance you can show us your notebooks?</p>",
      "rawMarkdown": "Congrats! Is there any chance you can show us your notebooks?",
      "replies": [
        {
          "id": 3116830,
          "postDate": "2025-02-06T10:59:45.353Z",
          "content": "<p>We are still a bit overwhelmed by the result, and also need some sleep after intense last week. Will share more things in a bit.</p>",
          "rawMarkdown": "We are still a bit overwhelmed by the result, and also need some sleep after intense last week. Will share more things in a bit.",
          "votes": 9
        }
      ]
    },
    {
      "id": 3116617,
      "postDate": "2025-02-06T06:17:21.190Z",
      "content": "<p>Your GPU accelerating is amazing!</p>",
      "rawMarkdown": "Your GPU accelerating is amazing!"
    },
    {
      "id": 3122340,
      "postDate": "2025-02-12T14:03:47.947Z",
      "content": "<p>Congrats, thank you for sharing.</p>",
      "rawMarkdown": "Congrats, thank you for sharing."
    },
    {
      "id": 3121775,
      "postDate": "2025-02-11T23:43:40.147Z",
      "content": "<p>Thanks for sharing your solution. </p>",
      "rawMarkdown": "Thanks for sharing your solution. "
    }
  ],
  "comments": [
    {
      "id": 3117046,
      "author_name": "Eugene Khvedchenya",
      "author_url": "",
      "post_date": "2025-02-06T15:27:15.530000",
      "content": "<p>Training code for OD models released! Check it here <a href=\"https://github.com/BloodAxe/Kaggle-2024-CryoET\" target=\"_blank\">https://github.com/BloodAxe/Kaggle-2024-CryoET</a></p>",
      "votes": 5,
      "replies": [
        {
          "id": 3122139,
          "author_name": "",
          "author_url": "",
          "post_date": "2025-02-12T10:15:37.800000",
          "content": "<p>okiiee got it </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 3117879,
      "author_name": "homiecal",
      "author_url": "",
      "post_date": "2025-02-07T10:58:27.523000",
      "content": "<p>Congrats, and thank you for sharing your solution - so much for me to learn from here. </p>\n<p>My sincere support for the people of Ukraine, coming from Australia</p>",
      "votes": 3,
      "replies": [
        {
          "id": 3117882,
          "author_name": "Eugene Khvedchenya",
          "author_url": "",
          "post_date": "2025-02-07T11:00:57.660000",
          "content": "<p>Thanks for your kind words! I learned myself from winning solutions of past years and I guess it’s my turn to share the knowledge now :) </p>",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 3117546,
      "author_name": "Teena Sapra",
      "author_url": "",
      "post_date": "2025-02-07T03:53:05.467000",
      "content": "<p>Congrats. Your insights are helpful</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 3122380,
      "author_name": "Rohan Dsouza",
      "author_url": "",
      "post_date": "2025-02-12T15:15:48.567000",
      "content": "<p>congrats nicely done!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3122180,
      "author_name": "YanjieSong",
      "author_url": "",
      "post_date": "2025-02-12T10:59:53.777000",
      "content": "<p>Congrats! Thank you for sharing a nice write-up and solution code !!!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3121687,
      "author_name": "Rachit Pandey",
      "author_url": "",
      "post_date": "2025-02-11T20:30:57.507000",
      "content": "<p>yoyo is the best game</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3121343,
      "author_name": "Mueen Khan Khattak",
      "author_url": "",
      "post_date": "2025-02-11T13:52:43.420000",
      "content": "<p>Congrats!  Thank you for sharing a well structured write up and your code.</p>\n<p>I found your approach quite interesting! Could you point me to the custom loss function used in your anchor-free point detection model? I'm curious to understand how it contributes to the overall performance.</p>\n<p>Looking forward to your insights! </p>",
      "votes": 0,
      "replies": [
        {
          "id": 3123029,
          "author_name": "Eugene Khvedchenya",
          "author_url": "",
          "post_date": "2025-02-13T09:32:50.977000",
          "content": "<p>Sure, it can be found here <a href=\"https://github.com/BloodAxe/Kaggle-2024-CryoET/blob/master/cryoet/modelling/detection/functional.py#L122\" target=\"_blank\">https://github.com/BloodAxe/Kaggle-2024-CryoET/blob/master/cryoet/modelling/detection/functional.py#L122</a></p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 3121309,
      "author_name": "Aastha Garg",
      "author_url": "",
      "post_date": "2025-02-11T13:16:47.510000",
      "content": "<p>woww greattttttttttttttttttt</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3121217,
      "author_name": "Yassine Alouini",
      "author_url": "",
      "post_date": "2025-02-11T11:26:32.770000",
      "content": "<p>Thanks for the write-up.<br>\nYou mention that MixUp didn't work for the object detection part but it seems that it has worked for the segmentation part.<br>\nDo you have any ideas why? </p>",
      "votes": 0,
      "replies": [
        {
          "id": 3123030,
          "author_name": "Eugene Khvedchenya",
          "author_url": "",
          "post_date": "2025-02-13T09:37:26.290000",
          "content": "<p>It's a good question to ask. Frankly - I don't know. I imagine if we have at least 100 studies in train, the CV estimate would be less noisy and we could obtain more reliable signals on what work and what doesn't.</p>\n<p>The metric is very sensitive to confidence thresholds you choose, so LB validation was the only viable approach to tune thresholds, given the size of studies in public part of the LB. I don't quite like overfitting to public, but in this comp it was the only viable choice. Under these conditions, it is really hard to say confidently that one approach works or doesn't since you may be just in poor thresholds area and moving one threshold to +0.05 and another -0.05 would give you SOTA. I tried playing with thresholds a bit for mixup models but didn't spent more than 5 subs on finding best combination of thresholds. </p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 3120466,
      "author_name": "Sinan Calisir",
      "author_url": "",
      "post_date": "2025-02-10T14:59:30.170000",
      "content": "<p>I went over your code to study and it's just amazing! There are things that I missed because of lack of my knowledge about object detection but it was easy to follow and very well written. Thank you!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3119828,
      "author_name": "Lennon Shikhman",
      "author_url": "",
      "post_date": "2025-02-09T18:03:17.720000",
      "content": "<p>Huge congratulations, <a href=\"https://www.kaggle.com/bloodaxe\" target=\"_blank\">@bloodaxe</a> and <a href=\"https://www.kaggle.com/christofhenkel\" target=\"_blank\">@christofhenkel</a>! Your write-up is an absolute masterclass in efficient deep learning inference and ensembling techniques. The combination of SegResNet and DynUnet backbones, the anchor-free point detection approach, and your custom loss function inspired by PP-YOLO really showcases the depth of optimization that goes into top-tier solutions.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3119575,
      "author_name": "MK245",
      "author_url": "",
      "post_date": "2025-02-09T12:10:31.190000",
      "content": "<p>Awesome work with parallel computing the processing time and running both T4's! Thanks for sharing this.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3117293,
      "author_name": "Bartley",
      "author_url": "",
      "post_date": "2025-02-06T20:14:33.693000",
      "content": "<p>Congrats! Thank you for sharing a nice write-up and your code.</p>\n<p>Could you point me to the custom loss function used in your \"anchor-free point detection model\"?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 3117295,
          "author_name": "Eugene Khvedchenya",
          "author_url": "",
          "post_date": "2025-02-06T20:19:15.120000",
          "content": "<p>You can start from main loss function and dig as deep as you want from there: <a href=\"https://github.com/BloodAxe/Kaggle-2024-CryoET/blob/master/cryoet/modelling/detection/functional.py#L122\" target=\"_blank\">https://github.com/BloodAxe/Kaggle-2024-CryoET/blob/master/cryoet/modelling/detection/functional.py#L122</a></p>",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 3116826,
      "author_name": "madmax0404",
      "author_url": "",
      "post_date": "2025-02-06T10:53:50.597000",
      "content": "<p>Congrats! Is there any chance you can show us your notebooks?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 3116830,
          "author_name": "Dieter",
          "author_url": "",
          "post_date": "2025-02-06T10:59:45.353000",
          "content": "<p>We are still a bit overwhelmed by the result, and also need some sleep after intense last week. Will share more things in a bit.</p>",
          "votes": 9,
          "replies": []
        }
      ]
    },
    {
      "id": 3116617,
      "author_name": "Timmy Juicehouse",
      "author_url": "",
      "post_date": "2025-02-06T06:17:21.190000",
      "content": "<p>Your GPU accelerating is amazing!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3122340,
      "author_name": "Tommaso Facchin",
      "author_url": "",
      "post_date": "2025-02-12T14:03:47.947000",
      "content": "<p>Congrats, thank you for sharing.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3121775,
      "author_name": "sv3799",
      "author_url": "",
      "post_date": "2025-02-11T23:43:40.147000",
      "content": "<p>Thanks for sharing your solution. </p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3116611": "```markdown\nThroughout the competition there were numerous missile strikes, bombings, and other acts of war that have taken the lives of many innocent people in Ukraine. \nRockets from russia hit within few kilometers from my home in Odesa. Each day Kaggle users from Ukraine facing the chance of not waking up. Just keep in this mind while you read this solution writeup.\n\nI would like to thank the Armed Forces of Ukraine, the Security Service of Ukraine, Defence Intelligence of Ukraine, and the State Emergency Service of Ukraine for providing safety and security to participate in this great competition, complete this work, and help science, technology, and business not to stop but to move forward.\n```\n\n**TLDR**: The solution in an ensemble of segmentation (3D Unets with ResNet & B3 encoders) and object detection models (SegResNet and DynUnet backbones) from Monai. \nWe accelerate model inference with NVidia TensorRT to achieve 200% speedup compared to eager PyTorch runtime and leverage two T4 GPUs to run predictions in parallel.\n\nKudos to my teammate @christofhenkel for his great work on ensembling our models into a final submission. Alone, we were able to get into Top-5, but when merged jumped immediately to the Top-1.\nThis writeup covers my part of the solution. Please check out Christof's [writeup](https://www.kaggle.com/competitions/czii-cryo-et-object-identification/discussion/561510) for segmentation approach and ensembling technique.\n\n**Update** Training code, models and inference code for OD part released. Check it out here:  \n\n- Training code https://github.com/BloodAxe/Kaggle-2024-CryoET\n- Pre-trained models dataset https://www.kaggle.com/datasets/bloodaxe/cryoet-detection-models (Mirror on GitHub: https://github.com/BloodAxe/Kaggle-2024-CryoET/releases/tag/1.0)\n- [Notebook] Conversion from ONNX to TensorRT engine https://www.kaggle.com/code/bloodaxe/convert-onnx-ensemble-to-tensorrt\n- [Notebook] Inference notebook (OD only) https://www.kaggle.com/code/bloodaxe/v4-inference-notebook/notebook\n- [Dataset] ONNX ensemble https://www.kaggle.com/datasets/bloodaxe/cryoet-onnx-models/\n\n## Introduction\n\nMy baseline was a segmentation (heatmap) 3D Unet built from the SegResNet backbone. As a supervised objective target, I built a five-class heatmap with Gaussian peaks for present ground-truth targets in volume. To train a model I used reduced focal loss from CenterNet paper and NMS postprocessing approach from the same paper. \nAlthough this approach immediately gave ~0.740 LB score (silver zone at the moment of submitting), I ditched it as I anticipated many participants would be using the segmentation approach. It is the easiest to implement, however it doesn't necessarily suit the best for the point detection task (A lot of model's capacity dedicated to predict the curvature of Gaussian at the peak and not on the particles discrimination). Second reason is model diversity. So if I decide to team up in the future, there will be less diversity in the final model ensemble.\n\n## Modeling approach \n\nMy final approach uses SegResNet and DynUnet backbones from Monai in combination with a custom point detection head. Based on my experience creating YOLO-based object detection models, I implemented an anchor-free point detection model. Likewise to box detection, model predict class probabilities map and offsets to the center object. Unlike box detection, however there is no concept of bounding box and hence to use IoU between two points I used well-known object keypoint similarity measure `exp(-mse(x,y)/ (2 * radius ^ 2))` as proxy for IoU metric.\n\nTo train such model I implemented a custom loss function that mimics PP-Yolo loss function with a few modifications:\n* Assignment of ground-truth labels to predicted centers based on point-point IoU. Point-Point IoU metric.\n* Taking Top-K predictions for each GT label for training\n* Computing varifocal loss for class map prediction and IoU-based loss for distance regression\n\nMy point detection approach worked well in terms of accuracy and training speed. On 4x3090 it took a mere two hours to train a single fold. The first submission of a single-fold detection model scored 0.752 on the LB. Once the modeling approach became clear, the next effort was to make it fast.\n\nIn this competition, we need to process 500 scans within a 12-hour time limit.  This translates into ~1m30s per single study. In my models, I predict class-map `[B,C,D/2,H/2,W/2]` and offsets map `[B,3,D/2,H/2,W/2]` in stride 2 along depth, height, and width. \n\nFor the object detection approach, there is no need to predict a full-resolution feature maps. Predicting smaller feature maps has a massive impact on model throughput. I've got almost 50% speed increase when changing to stride 2 output from stride 1 (full resolution).\n\nI ablate on stride 1, stride 2, stride 4, and stride 2 & 4 outputs and found that:\n* Stride 1 gives a slightly, slightly better (0.002) LB  score than stride 2\n* Stride 4 also worked ok, but stride 2 was better in the F-beta score.\n* Using stride 2 and 4 didn't bring any improvements over using only stride 2. An interesting observation when using the two-heads method: Large particles tend to migrate to stride 4 while small classes were present on the class-map with stride 2.\n\nIn all my models I use stride 2 prediction maps.\n\n## Training\n\nI used 5-fold cross-validation scheme for all models. Each fold used 2 studies for validation and 5 for training.\n\nTraining epoch used fixed number of random crops per study and fixed number of random crops around each particle instance. For data augmentations I used random rotations along Z-axis (+- 180 degrees), slight rotations along X and Y axis (+-10), subtle scale jitter (+-5%), and random flips along X, Y, Z axes.\nI experimented with erasing particles from the input volume, random copy-pasting particles from one study to another, doing mixup on the particle instances, but these augmentation techniques didn't bring any improvements in the final model. \n\nValidation epoch used sliding window approach with the same window size and overlap as in the inference kernel. During validation individual tiles accumulated to final classmap and offsets map and F-beta score was computed on the final maps. After each epoch, I computed per-class thresholds that maximizes F-beta score on the validation set. I saved top-5 models for training experiment which I later averaged which almost always increased the F-beta score.\n\nTraining used 96x128x128px while validation was performed on 192x128x128 volumes.\n\n## Tiling\n\nI used a sliding window approach to tile the input volume. The window size was 192x128x128px with 1x9x9 tiles configuration for input volume of 184x630x630. Individual tiles were aggregated to final classmap and offsets map using a weighted averaging approach where pixels on the border of the tile had lower weights than pixels in the center of the tile. For sliding average aggregation I used my own implementation.\n\n## Postprocessing\n\nHaving an aggregated classmap of [C, D, H, W] shape and offsets map of [3, D, H, W] shape, I used the following postprocessing steps:\n\n- Centernet-like NMS to remove duplicate detections: `x = x * (x == maxpool3d(x, kernel_size=3, stride=1, padding=1))`\n- Take Top-K predictions for each class (I used 16K but lower numbers should also work fine I guess)\n- Filter out predictions by confidence threshold (Each class has it's own confidence threshold)\n- Apply greedy NMS using pairwise IoU distance (Similar to how Box NMS works) for detections of a single class.\n- Scale pixels to Angstroms\n\nSince my models were directly predicting offset to the center of the particle, this approach is not sensitive to systematic shift in point coordinates (Some argue it's 0.5px, some claims it to be 1px).\n\n## Object Detection Ensemble\n\nMy final object ensemble is 5 models (folds) of OD SegResNet and 5 models of OD DynUnet (OD stands for Object Detection to differentiate them from segmentation-based models). \nA 10 models in total. I used local F-beta metric to select final checkpoints that would make into ensemble.\n\n### SegResNet\n\n| fold         |    score |   AFRT |   BGT |   RBSM |   TRGLB |   VLP |\n|:-------------|---------:|-------:|------:|-------:|--------:|------:|\n| 0            | 0.8457   |  0.265 | 0.29  |  0.195 |   0.15  | 0.55  |\n| 1            | 0.8366   |  0.345 | 0.11  |  0.185 |   0.135 | 0.305 |\n| 2            | 0.8046   |  0.355 | 0.28  |  0.395 |   0.34  | 0.185 |\n| 3            | 0.8398   |  0.23  | 0.345 |  0.405 |   0.36  | 0.255 |\n| 4            | 0.8437   |  0.165 | 0.35  |  0.245 |   0.235 | 0.27  |\n\n### DynUnet\n\n| fold         |    score |   AFRT |   BGT |   RBSM |   TRGLB |   VLP |\n|:-------------|---------:|-------:|------:|-------:|--------:|------:|\n| 0            | 0.8272   |  0.145 | 0.33  |  0.215 |   0.235 | 0.405 |\n| 1            | 0.8337   |  0.39  | 0.185 |  0.215 |   0.21  | 0.235 |\n| 2            | 0.7971   |  0.425 | 0.16  |  0.55  |   0.275 | 0.65  |\n| 3            | 0.8418   |  0.385 | 0.345 |  0.345 |   0.305 | 0.21  |\n| 4            | 0.842    |  0.23  | 0.34  |  0.255 |   0.375 | 0.14  |\n\nAs you can see on F-4 scores from individual folds, there is a serious gap between CV and LB scores. Me explanation of this phenomenon is big difference in number of studies in validation set that we use locally and LB. I found that extending validation set by using flipped and 90-degree rotated studies lowers the CV score and also makes F-beta curve mode smooth, which allow finding a more accurate and less noisy threshold.\n\nTo figure out thresholds of the final ensemble I compute OOF F-beta curves for each class and select thresholds that maximize average F-beta score:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1504864%2Fd845bb6bd6eb49c923c2e75c8916f5f5%2Fplot_rotTrue_zFalse_yFalse_xFalse_96x3_96x9_slpaFalse.png?generation=1738822738382933&alt=media)\n\n## Accelerating inference\n\nFinal submission uses TensorRT for inference.\n\nInitially, I used `torch.jit` which was enough for start, but as we teamed up with @christofhenkel a more efficient approach was needed.\nI tried using `onnxruntime` with `CUDAExecutionProvider` which gave me nearly the same throughput as `torch.jit`.\nFinally, by enabling `TensorRTExecutionProvider` that leverages all the power of TensorRT I was able to achieve desired speedup of 200% compared to `torch.jit`\n\nThe proces of going from individual checkpoints to TensorRT engine is multi-stage:\n\n1. [Offline] Convert individual checkpoints into a single ONNX model containing all models and averaging their predictions\n2. [Kaggle] Convert ONNX model into TensorRT engine (And save it to disk). This step happens on Kaggle, using T4 GPU (a target GPU we use for inference) and it saved us 10 minutes of submission time for creating TensorRT engine.\n3. [Kaggle] Actual inference notebook. I split all test data in two chunks and use 2xT4 GPUs to process them in parallel. Then two submission shards were concatenated to form the final submission.\n\nTo combine results of my ensemble with ensemble of @christofhenkel we simply used weighted average of the predictions on the classmap level followed\nby postprocessing described above. Please check second part of the solution [here](https://www.kaggle.com/competitions/czii-cryo-et-object-identification/discussion/561510).\n\n## Things that didn't quite work (for me)\n\n* Mixup, Copy-Paste, Random-Erasing (CV higher, LB lower)\n* 2.5-D models\n* 3D version of HRNet and ConvNext\n* Gaussian noise, anisotropic scale jitter\n* Knowledge distillation\n\n## References\n\n1. Training code https://github.com/BloodAxe/Kaggle-2024-CryoET\n2. Pre-trained models dataset https://www.kaggle.com/datasets/bloodaxe/cryoet-detection-models (Mirror on GitHub: https://github.com/BloodAxe/Kaggle-2024-CryoET/releases/tag/1.0)\n3. [Notebook] Conversion from ONNX to TensorRT engine https://www.kaggle.com/code/bloodaxe/convert-onnx-ensemble-to-tensorrt\n4. [Notebook] Inference notebook (OD only) https://www.kaggle.com/code/bloodaxe/v4-inference-notebook/notebook\n5. [Dataset] ONNX ensemble  https://www.kaggle.com/datasets/bloodaxe/cryoet-onnx-models/\n",
    "3117046": "Training code for OD models released! Check it here https://github.com/BloodAxe/Kaggle-2024-CryoET",
    "3117879": "Congrats, and thank you for sharing your solution - so much for me to learn from here. \n\nMy sincere support for the people of Ukraine, coming from Australia",
    "3117546": "Congrats. Your insights are helpful",
    "3122380": "congrats nicely done!",
    "3122180": "Congrats! Thank you for sharing a nice write-up and solution code !!!",
    "3121687": "yoyo is the best game",
    "3121343": "Congrats!  Thank you for sharing a well structured write up and your code.\n\nI found your approach quite interesting! Could you point me to the custom loss function used in your anchor-free point detection model? I'm curious to understand how it contributes to the overall performance.\n\nLooking forward to your insights! ",
    "3121309": "woww greattttttttttttttttttt\n",
    "3121217": "Thanks for the write-up.\nYou mention that MixUp didn't work for the object detection part but it seems that it has worked for the segmentation part.\nDo you have any ideas why? ",
    "3120466": "I went over your code to study and it's just amazing! There are things that I missed because of lack of my knowledge about object detection but it was easy to follow and very well written. Thank you!",
    "3119828": "Huge congratulations, @bloodaxe and @christofhenkel! Your write-up is an absolute masterclass in efficient deep learning inference and ensembling techniques. The combination of SegResNet and DynUnet backbones, the anchor-free point detection approach, and your custom loss function inspired by PP-YOLO really showcases the depth of optimization that goes into top-tier solutions.",
    "3119575": "Awesome work with parallel computing the processing time and running both T4's! Thanks for sharing this.\n",
    "3117293": "Congrats! Thank you for sharing a nice write-up and your code.\n\nCould you point me to the custom loss function used in your \"anchor-free point detection model\"?",
    "3116826": "Congrats! Is there any chance you can show us your notebooks?",
    "3116617": "Your GPU accelerating is amazing!",
    "3122340": "Congrats, thank you for sharing.",
    "3121775": "Thanks for sharing your solution. "
  }
}