{
  "id": 561580,
  "title": "5th Place Solution",
  "url": "/competitions/czii-cryo-et-object-identification/writeups/youssef-ouertani-5th-place-solution",
  "author_name": "",
  "post_date": "2025-02-13T16:36:42.320Z",
  "votes": 21,
  "comment_count": 2,
  "views": 0,
  "content": "<p><strong>Introduction</strong></p>\n<p>First and foremost, I would like to express my gratitude to Allah, the competition hosts, Kaggle, and <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> for making this event possible. This competition provided an exciting opportunity to work with 3D volumetric data and develop an efficient solution for particle detection.</p>\n<p>I present my straightforward approach to solving the problem. Below, I outline the key steps of my solution, including data preparation, network architecture, training strategy, and inference techniques. I also discuss what worked and what didn’t, along with the final results achieved on the public and private leaderboards.</p>\n<p><strong>Data Preparation &amp; Loading</strong></p>\n<p><strong>Volume Normalization</strong></p>\n<ul>\n<li>The volumes were normalized by calculating the (5, 99) percentiles of the 7 volume datasets and averaging them to perform min-max scaling.</li>\n</ul>\n<p><strong>Label Preparation</strong></p>\n<ul>\n<li>The labels were created as spheres with a radius of log2(given_radius) * 0.8.</li>\n</ul>\n<p><strong>Training Data</strong></p>\n<ul>\n<li>The model was trained on batches of 4 patches, each of size 128x128x128.</li>\n<li>Patches were randomly sampled from the volumes during training.</li>\n</ul>\n<p><strong>Data Augmentation</strong></p>\n<ul>\n<li>Flipping along all 3 axes.</li>\n<li>Rotations of 90°, 180°, and 270° along the z-axis.</li>\n<li>Mean and standard deviation shifting using the following function</li>\n</ul>\n<pre><code> ():\n    factor =  / (shift * )\n    std = image.std()\n    mean = image.mean()\n    shift_mean = (torch.rand() / factor - shift).item()\n    shift_std = (torch.rand() / factor - shift).item()\n    new_mean = mean + mean * shift_mean\n    new_std = std + std * shift_std\n    new_image = (image - mean) / std * new_std + new_mean\n     new_image\n</code></pre>\n<p><strong>Network Architecture</strong></p>\n<p>The network architecture is inspired by DeepFinder, with the following modifications:</p>\n<ul>\n<li>Added a BatchNorm3d layer as the first input layer.</li>\n<li>Reduced the number of channels to 28, 32, and 36, resulting in a compact model size of 1.44 MB.</li>\n<li>Used trilinear interpolation for downsampling and upsampling, except for the final upsampling layer, which uses a transposed convolution.</li>\n</ul>\n<p>Here is a visualization of the architecture:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6498106%2F0b0f957d7842139008b83908a1670737%2FScreenshot%20from%202025-02-06%2019-33-03.png?generation=1738873845345750&amp;alt=media\" alt=\"\"></p>\n<p><strong>Training Strategy</strong></p>\n<ul>\n<li>Optimizer: Adam with a learning rate of 0.0001, beta1 of 0.9, and beta2 of 0.999.</li>\n<li>Loss Function: Label smoothing cross-entropy with a smoothing factor of 0.01.</li>\n<li>Precision: Training was conducted in float16 precision with gradient clipping applied.</li>\n<li>Model Ensembling: The final model consists of 4 seeds of the above architecture, trained on all 7 volumes.</li>\n</ul>\n<p><strong>Inference</strong></p>\n<ul>\n<li>Patch Splitting:</li>\n</ul>\n<p>For inference, the volumes were split into patches of size 128x128x128 with minimal overlap along the z-axis and overlap + 1 along the x and y axes.</p>\n<ul>\n<li>Test-Time Augmentation (TTA):</li>\n</ul>\n<p>Applied 3 flips and 3 rotations.</p>\n<ul>\n<li>Post-Processing:</li>\n</ul>\n<p>Connected components were applied to binary masks generated using a probability threshold for each particle.<br>\nComponents with an area less than 1/7th of the trained masks were removed.</p>\n<p><strong>Results</strong></p>\n<ul>\n<li>Public Leaderboard: 0.7798</li>\n<li>Private Leaderboard: 0.7825</li>\n</ul>\n<p><strong>What Didn’t Work</strong></p>\n<ul>\n<li>Multicascade Network: This approach did not yield improvements.</li>\n<li>Larger Models: These models tended to overfit quickly and performed worse than the compact architecture.</li>\n</ul>",
  "messages": [
    {
      "id": "3117266",
      "postDate": "02/06/2025 19:30:30",
      "content": "<p><strong>Introduction</strong></p>\n<p>First and foremost, I would like to express my gratitude to Allah, the competition hosts, Kaggle, and <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> for making this event possible. This competition provided an exciting opportunity to work with 3D volumetric data and develop an efficient solution for particle detection.</p>\n<p>I present my straightforward approach to solving the problem. Below, I outline the key steps of my solution, including data preparation, network architecture, training strategy, and inference techniques. I also discuss what worked and what didn’t, along with the final results achieved on the public and private leaderboards.</p>\n<p><strong>Data Preparation &amp; Loading</strong></p>\n<p><strong>Volume Normalization</strong></p>\n<ul>\n<li>The volumes were normalized by calculating the (5, 99) percentiles of the 7 volume datasets and averaging them to perform min-max scaling.</li>\n</ul>\n<p><strong>Label Preparation</strong></p>\n<ul>\n<li>The labels were created as spheres with a radius of log2(given_radius) * 0.8.</li>\n</ul>\n<p><strong>Training Data</strong></p>\n<ul>\n<li>The model was trained on batches of 4 patches, each of size 128x128x128.</li>\n<li>Patches were randomly sampled from the volumes during training.</li>\n</ul>\n<p><strong>Data Augmentation</strong></p>\n<ul>\n<li>Flipping along all 3 axes.</li>\n<li>Rotations of 90°, 180°, and 270° along the z-axis.</li>\n<li>Mean and standard deviation shifting using the following function</li>\n</ul>\n<pre><code> ():\n    factor =  / (shift * )\n    std = image.std()\n    mean = image.mean()\n    shift_mean = (torch.rand() / factor - shift).item()\n    shift_std = (torch.rand() / factor - shift).item()\n    new_mean = mean + mean * shift_mean\n    new_std = std + std * shift_std\n    new_image = (image - mean) / std * new_std + new_mean\n     new_image\n</code></pre>\n<p><strong>Network Architecture</strong></p>\n<p>The network architecture is inspired by DeepFinder, with the following modifications:</p>\n<ul>\n<li>Added a BatchNorm3d layer as the first input layer.</li>\n<li>Reduced the number of channels to 28, 32, and 36, resulting in a compact model size of 1.44 MB.</li>\n<li>Used trilinear interpolation for downsampling and upsampling, except for the final upsampling layer, which uses a transposed convolution.</li>\n</ul>\n<p>Here is a visualization of the architecture:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6498106%2F0b0f957d7842139008b83908a1670737%2FScreenshot%20from%202025-02-06%2019-33-03.png?generation=1738873845345750&amp;alt=media\" alt=\"\"></p>\n<p><strong>Training Strategy</strong></p>\n<ul>\n<li>Optimizer: Adam with a learning rate of 0.0001, beta1 of 0.9, and beta2 of 0.999.</li>\n<li>Loss Function: Label smoothing cross-entropy with a smoothing factor of 0.01.</li>\n<li>Precision: Training was conducted in float16 precision with gradient clipping applied.</li>\n<li>Model Ensembling: The final model consists of 4 seeds of the above architecture, trained on all 7 volumes.</li>\n</ul>\n<p><strong>Inference</strong></p>\n<ul>\n<li>Patch Splitting:</li>\n</ul>\n<p>For inference, the volumes were split into patches of size 128x128x128 with minimal overlap along the z-axis and overlap + 1 along the x and y axes.</p>\n<ul>\n<li>Test-Time Augmentation (TTA):</li>\n</ul>\n<p>Applied 3 flips and 3 rotations.</p>\n<ul>\n<li>Post-Processing:</li>\n</ul>\n<p>Connected components were applied to binary masks generated using a probability threshold for each particle.<br>\nComponents with an area less than 1/7th of the trained masks were removed.</p>\n<p><strong>Results</strong></p>\n<ul>\n<li>Public Leaderboard: 0.7798</li>\n<li>Private Leaderboard: 0.7825</li>\n</ul>\n<p><strong>What Didn’t Work</strong></p>\n<ul>\n<li>Multicascade Network: This approach did not yield improvements.</li>\n<li>Larger Models: These models tended to overfit quickly and performed worse than the compact architecture.</li>\n</ul>",
      "rawMarkdown": "**Introduction**\n\nFirst and foremost, I would like to express my gratitude to Allah, the competition hosts, Kaggle, and @hengck23 for making this event possible. This competition provided an exciting opportunity to work with 3D volumetric data and develop an efficient solution for particle detection.\n\nI present my straightforward approach to solving the problem. Below, I outline the key steps of my solution, including data preparation, network architecture, training strategy, and inference techniques. I also discuss what worked and what didn’t, along with the final results achieved on the public and private leaderboards.\n\n**Data Preparation & Loading**\n\n**Volume Normalization**\n- The volumes were normalized by calculating the (5, 99) percentiles of the 7 volume datasets and averaging them to perform min-max scaling.\n\n**Label Preparation**\n- The labels were created as spheres with a radius of log2(given_radius) * 0.8.\n\n**Training Data**\n- The model was trained on batches of 4 patches, each of size 128x128x128.\n- Patches were randomly sampled from the volumes during training.\n\n**Data Augmentation**\n\n- Flipping along all 3 axes.\n- Rotations of 90°, 180°, and 270° along the z-axis.\n- Mean and standard deviation shifting using the following function\n\n```python\ndef mean_std_shift(image, shift=0.03):\n    factor = 1 / (shift * 2)\n    std = image.std()\n    mean = image.mean()\n    shift_mean = (torch.rand(1) / factor - shift).item()\n    shift_std = (torch.rand(1) / factor - shift).item()\n    new_mean = mean + mean * shift_mean\n    new_std = std + std * shift_std\n    new_image = (image - mean) / std * new_std + new_mean\n    return new_image\n```\n**Network Architecture**\n\nThe network architecture is inspired by DeepFinder, with the following modifications:\n\n - Added a BatchNorm3d layer as the first input layer.\n - Reduced the number of channels to 28, 32, and 36, resulting in a compact model size of 1.44 MB.\n - Used trilinear interpolation for downsampling and upsampling, except for the final upsampling layer, which uses a transposed convolution.\n\nHere is a visualization of the architecture:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6498106%2F0b0f957d7842139008b83908a1670737%2FScreenshot%20from%202025-02-06%2019-33-03.png?generation=1738873845345750&alt=media)\n\n**Training Strategy**\n\n- Optimizer: Adam with a learning rate of 0.0001, beta1 of 0.9, and beta2 of 0.999.\n- Loss Function: Label smoothing cross-entropy with a smoothing factor of 0.01.\n- Precision: Training was conducted in float16 precision with gradient clipping applied.\n- Model Ensembling: The final model consists of 4 seeds of the above architecture, trained on all 7 volumes.\n\n**Inference**\n- Patch Splitting:\n\nFor inference, the volumes were split into patches of size 128x128x128 with minimal overlap along the z-axis and overlap + 1 along the x and y axes.\n\n- Test-Time Augmentation (TTA):\n\nApplied 3 flips and 3 rotations.\n\n- Post-Processing:\n\nConnected components were applied to binary masks generated using a probability threshold for each particle.\nComponents with an area less than 1/7th of the trained masks were removed.\n\n**Results**\n- Public Leaderboard: 0.7798\n- Private Leaderboard: 0.7825\n\n**What Didn’t Work**\n- Multicascade Network: This approach did not yield improvements.\n- Larger Models: These models tended to overfit quickly and performed worse than the compact architecture.",
      "votes": null
    },
    {
      "id": "3117305",
      "postDate": "02/06/2025 20:33:05",
      "content": "<p>Congratulations on your gold medal and the prize!</p>",
      "rawMarkdown": "Congratulations on your gold medal and the prize!",
      "votes": null
    },
    {
      "id": "3118068",
      "postDate": "02/07/2025 14:47:24",
      "content": "<p>Congratulations! Always proud to see fellow Tunisians achieve gold medals!</p>",
      "rawMarkdown": "Congratulations! Always proud to see fellow Tunisians achieve gold medals!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3117305,
      "author_name": "snnclsr",
      "author_url": "",
      "post_date": "02/06/2025 20:33:05",
      "content": "<p>Congratulations on your gold medal and the prize!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3118068,
      "author_name": "fahmiayari",
      "author_url": "",
      "post_date": "02/07/2025 14:47:24",
      "content": "<p>Congratulations! Always proud to see fellow Tunisians achieve gold medals!</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3117266": "**Introduction**\n\nFirst and foremost, I would like to express my gratitude to Allah, the competition hosts, Kaggle, and @hengck23 for making this event possible. This competition provided an exciting opportunity to work with 3D volumetric data and develop an efficient solution for particle detection.\n\nI present my straightforward approach to solving the problem. Below, I outline the key steps of my solution, including data preparation, network architecture, training strategy, and inference techniques. I also discuss what worked and what didn’t, along with the final results achieved on the public and private leaderboards.\n\n**Data Preparation & Loading**\n\n**Volume Normalization**\n- The volumes were normalized by calculating the (5, 99) percentiles of the 7 volume datasets and averaging them to perform min-max scaling.\n\n**Label Preparation**\n- The labels were created as spheres with a radius of log2(given_radius) * 0.8.\n\n**Training Data**\n- The model was trained on batches of 4 patches, each of size 128x128x128.\n- Patches were randomly sampled from the volumes during training.\n\n**Data Augmentation**\n\n- Flipping along all 3 axes.\n- Rotations of 90°, 180°, and 270° along the z-axis.\n- Mean and standard deviation shifting using the following function\n\n```python\ndef mean_std_shift(image, shift=0.03):\n    factor = 1 / (shift * 2)\n    std = image.std()\n    mean = image.mean()\n    shift_mean = (torch.rand(1) / factor - shift).item()\n    shift_std = (torch.rand(1) / factor - shift).item()\n    new_mean = mean + mean * shift_mean\n    new_std = std + std * shift_std\n    new_image = (image - mean) / std * new_std + new_mean\n    return new_image\n```\n**Network Architecture**\n\nThe network architecture is inspired by DeepFinder, with the following modifications:\n\n - Added a BatchNorm3d layer as the first input layer.\n - Reduced the number of channels to 28, 32, and 36, resulting in a compact model size of 1.44 MB.\n - Used trilinear interpolation for downsampling and upsampling, except for the final upsampling layer, which uses a transposed convolution.\n\nHere is a visualization of the architecture:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6498106%2F0b0f957d7842139008b83908a1670737%2FScreenshot%20from%202025-02-06%2019-33-03.png?generation=1738873845345750&alt=media)\n\n**Training Strategy**\n\n- Optimizer: Adam with a learning rate of 0.0001, beta1 of 0.9, and beta2 of 0.999.\n- Loss Function: Label smoothing cross-entropy with a smoothing factor of 0.01.\n- Precision: Training was conducted in float16 precision with gradient clipping applied.\n- Model Ensembling: The final model consists of 4 seeds of the above architecture, trained on all 7 volumes.\n\n**Inference**\n- Patch Splitting:\n\nFor inference, the volumes were split into patches of size 128x128x128 with minimal overlap along the z-axis and overlap + 1 along the x and y axes.\n\n- Test-Time Augmentation (TTA):\n\nApplied 3 flips and 3 rotations.\n\n- Post-Processing:\n\nConnected components were applied to binary masks generated using a probability threshold for each particle.\nComponents with an area less than 1/7th of the trained masks were removed.\n\n**Results**\n- Public Leaderboard: 0.7798\n- Private Leaderboard: 0.7825\n\n**What Didn’t Work**\n- Multicascade Network: This approach did not yield improvements.\n- Larger Models: These models tended to overfit quickly and performed worse than the compact architecture.",
    "3117305": "Congratulations on your gold medal and the prize!",
    "3118068": "Congratulations! Always proud to see fellow Tunisians achieve gold medals!"
  },
  "source": "meta"
}