{
  "id": 417281,
  "title": "11th place solution",
  "url": "/competitions/vesuvius-challenge-ink-detection/discussion/417281",
  "author_name": "",
  "post_date": "2023-06-15T03:37:22.580739500Z",
  "votes": 33,
  "comment_count": 12,
  "views": 0,
  "content": "<p>Congrats to all the winners, and thanks to hosts for interesting competition!<br>\nI’m very happy to become kaggle competition GM.</p>\n<h1>overview</h1>\n<ul>\n<li>2.5d + 1d pool encoder</li>\n<li>ensemble of 6 Unet models</li>\n</ul>\n<h1>model architecture</h1>\n<ul>\n<li>Grouping the images by threes.<ul>\n<li>z_len = img_num / 3</li>\n<li>images (batch_size*z_len, 3, height, width) →2dcnn encoder →reshape to (batch_size, z_len, output_channel, height, width) →avg pooling(z_axis) → decoder</li></ul></li>\n<li>decoder: Unet</li>\n</ul>\n<table>\n<thead>\n<tr>\n<th>backbone</th>\n<th>img_num</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>se_resnext50_32x4d</td>\n<td>30</td>\n</tr>\n<tr>\n<td>eca_nfnet_l1</td>\n<td>30</td>\n</tr>\n<tr>\n<td>tf_efficientnetv2_m</td>\n<td>30</td>\n</tr>\n<tr>\n<td>convnext_small_in22ft1k (best)</td>\n<td>30</td>\n</tr>\n<tr>\n<td>se_resnext101_32x4d</td>\n<td>21</td>\n</tr>\n<tr>\n<td>convnext_base_in22ft1k</td>\n<td>21</td>\n</tr>\n</tbody>\n</table>\n<h1>training</h1>\n<p>training pipeline is almost same to my public notebook.</p>\n<p><a href=\"https://www.kaggle.com/code/tanakar/2-5d-segmentaion-baseline-training\" target=\"_blank\">https://www.kaggle.com/code/tanakar/2-5d-segmentaion-baseline-training</a></p>\n<ul>\n<li>img_size = 224</li>\n<li>stride = 112</li>\n<li>epoch = 30</li>\n<li>fp16</li>\n<li>5fold<ul>\n<li>ink_id=2 is divided into 3 parts</li></ul></li>\n<li>bce + dice loss</li>\n<li>exclude areas with a mask value of 0</li>\n<li>augmentation<ul>\n<li>almost same to my public training notebook</li>\n<li>I added A.RandomRotate90 because rotating the image during inference improved public lb</li></ul></li>\n</ul>\n<h1>inference</h1>\n<ul>\n<li>stride = 56</li>\n<li>ensemble<ul>\n<li>simple average</li></ul></li>\n<li>threshold = 0.5</li>\n<li>exclude areas with img.sum() == 0</li>\n</ul>\n<h1>not worked</h1>\n<ul>\n<li>3d encoder</li>\n<li>segformer</li>\n<li>other backbone</li>\n<li>mixup</li>\n<li>more slices</li>\n</ul>\n<h1>inference code</h1>\n<p><a href=\"https://www.kaggle.com/code/tanakar/2-5d-segmentaion-ensemble-final-stride-4\" target=\"_blank\">https://www.kaggle.com/code/tanakar/2-5d-segmentaion-ensemble-final-stride-4</a></p>",
  "messages": [
    {
      "id": "2303034",
      "postDate": "06/15/2023 03:37:22",
      "content": "<p>Congrats to all the winners, and thanks to hosts for interesting competition!<br>\nI’m very happy to become kaggle competition GM.</p>\n<h1>overview</h1>\n<ul>\n<li>2.5d + 1d pool encoder</li>\n<li>ensemble of 6 Unet models</li>\n</ul>\n<h1>model architecture</h1>\n<ul>\n<li>Grouping the images by threes.<ul>\n<li>z_len = img_num / 3</li>\n<li>images (batch_size*z_len, 3, height, width) →2dcnn encoder →reshape to (batch_size, z_len, output_channel, height, width) →avg pooling(z_axis) → decoder</li></ul></li>\n<li>decoder: Unet</li>\n</ul>\n<table>\n<thead>\n<tr>\n<th>backbone</th>\n<th>img_num</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>se_resnext50_32x4d</td>\n<td>30</td>\n</tr>\n<tr>\n<td>eca_nfnet_l1</td>\n<td>30</td>\n</tr>\n<tr>\n<td>tf_efficientnetv2_m</td>\n<td>30</td>\n</tr>\n<tr>\n<td>convnext_small_in22ft1k (best)</td>\n<td>30</td>\n</tr>\n<tr>\n<td>se_resnext101_32x4d</td>\n<td>21</td>\n</tr>\n<tr>\n<td>convnext_base_in22ft1k</td>\n<td>21</td>\n</tr>\n</tbody>\n</table>\n<h1>training</h1>\n<p>training pipeline is almost same to my public notebook.</p>\n<p><a href=\"https://www.kaggle.com/code/tanakar/2-5d-segmentaion-baseline-training\" target=\"_blank\">https://www.kaggle.com/code/tanakar/2-5d-segmentaion-baseline-training</a></p>\n<ul>\n<li>img_size = 224</li>\n<li>stride = 112</li>\n<li>epoch = 30</li>\n<li>fp16</li>\n<li>5fold<ul>\n<li>ink_id=2 is divided into 3 parts</li></ul></li>\n<li>bce + dice loss</li>\n<li>exclude areas with a mask value of 0</li>\n<li>augmentation<ul>\n<li>almost same to my public training notebook</li>\n<li>I added A.RandomRotate90 because rotating the image during inference improved public lb</li></ul></li>\n</ul>\n<h1>inference</h1>\n<ul>\n<li>stride = 56</li>\n<li>ensemble<ul>\n<li>simple average</li></ul></li>\n<li>threshold = 0.5</li>\n<li>exclude areas with img.sum() == 0</li>\n</ul>\n<h1>not worked</h1>\n<ul>\n<li>3d encoder</li>\n<li>segformer</li>\n<li>other backbone</li>\n<li>mixup</li>\n<li>more slices</li>\n</ul>\n<h1>inference code</h1>\n<p><a href=\"https://www.kaggle.com/code/tanakar/2-5d-segmentaion-ensemble-final-stride-4\" target=\"_blank\">https://www.kaggle.com/code/tanakar/2-5d-segmentaion-ensemble-final-stride-4</a></p>",
      "rawMarkdown": "Congrats to all the winners, and thanks to hosts for interesting competition!\nI’m very happy to become kaggle competition GM.\n\n# overview\n- 2.5d + 1d pool encoder\n- ensemble of 6 Unet models\n\n# model architecture\n\n- Grouping the images by threes.\n    - z_len = img_num / 3\n    - images (batch_size*z_len, 3, height, width) →2dcnn encoder →reshape to (batch_size, z_len, output_channel, height, width) →avg pooling(z_axis) → decoder\n- decoder: Unet\n\n| backbone | img_num |\n| ---- | ---- |\nse_resnext50_32x4d | 30\neca_nfnet_l1 | 30\ntf_efficientnetv2_m | 30\nconvnext_small_in22ft1k (best) | 30\nse_resnext101_32x4d | 21\nconvnext_base_in22ft1k | 21\n\n# training\n\ntraining pipeline is almost same to my public notebook.\n\nhttps://www.kaggle.com/code/tanakar/2-5d-segmentaion-baseline-training\n\n- img_size = 224\n- stride = 112\n- epoch = 30\n- fp16\n- 5fold\n    - ink_id=2 is divided into 3 parts\n- bce + dice loss\n- exclude areas with a mask value of 0\n- augmentation\n    - almost same to my public training notebook\n    - I added A.RandomRotate90 because rotating the image during inference improved public lb\n\n\n# inference\n\n- stride = 56\n- ensemble\n    - simple average\n- threshold = 0.5\n- exclude areas with img.sum() == 0\n\n# not worked\n\n- 3d encoder\n- segformer\n- other backbone\n- mixup\n- more slices\n\n# inference code\n\nhttps://www.kaggle.com/code/tanakar/2-5d-segmentaion-ensemble-final-stride-4",
      "votes": null
    },
    {
      "id": "2303043",
      "postDate": "06/15/2023 03:57:39",
      "content": "<p>Congratulations….</p>",
      "rawMarkdown": "Congratulations….",
      "votes": null
    },
    {
      "id": "2303365",
      "postDate": "06/15/2023 09:09:03",
      "content": "<p>Congratulations <a href=\"https://www.kaggle.com/tanakar\" target=\"_blank\">@tanakar</a> on becoming GM &amp; gold position finish 🎉🎉🎉</p>",
      "rawMarkdown": "Congratulations @tanakar on becoming GM & gold position finish 🎉🎉🎉",
      "votes": null
    },
    {
      "id": "2303823",
      "postDate": "06/15/2023 13:49:34",
      "content": "<p>Congratulations on your third solo gold! I'm thrilled to see you achieve Grandmaster status. In this competition, many people used your notebook, and furthermore, you have accomplished the remarkable feat of securing a solo gold medal! I respect you !</p>",
      "rawMarkdown": "Congratulations on your third solo gold! I'm thrilled to see you achieve Grandmaster status. In this competition, many people used your notebook, and furthermore, you have accomplished the remarkable feat of securing a solo gold medal! I respect you !",
      "votes": null
    },
    {
      "id": "2303835",
      "postDate": "06/15/2023 13:58:18",
      "content": "<p>\"not worked :segformer \"</p>\n<p>i design a segmentation model and it doesn't work. what should i do next?</p>\n<p>How to debug?</p>\n<ol>\n<li><p>use 3 channels + adam at small learning rate 1e-4,1e-5 <br>\nif it works, it means that  you cannot randomly initialise you first convolution channel to n input channel.<br>\nyou should initalise the values correctly.</p></li>\n<li><p>remove decoder and just train encoder (output 1/32 scale)<br>\nif it works, it means your decoder parameters or design are wrong</p></li>\n<li><p>try MIT-b0,b1,b2 …..<br>\nall needs different parameters/hyperparamters. Usually the shallowest model is easier to train (but may not lead to best results)</p></li>\n</ol>\n<hr>\n<p>performance is limited by data. since it works for other model, it means that your segformer hyperparamters are not correct.</p>\n<hr>\n<p>usually ensemble of VIT and CNN usually improve results. so it is important to get VIT working</p>",
      "rawMarkdown": "\"not worked :segformer \"\n\ni design a segmentation model and it doesn't work. what should i do next?\n\nHow to debug?\n1. use 3 channels + adam at small learning rate 1e-4,1e-5 \nif it works, it means that  you cannot randomly initialise you first convolution channel to n input channel.\nyou should initalise the values correctly.\n\n2. remove decoder and just train encoder (output 1/32 scale)\nif it works, it means your decoder parameters or design are wrong\n\n3. try MIT-b0,b1,b2 .....\nall needs different parameters/hyperparamters. Usually the shallowest model is easier to train (but may not lead to best results)\n\n---\n\nperformance is limited by data. since it works for other model, it means that your segformer hyperparamters are not correct.\n\n---\n\nusually ensemble of VIT and CNN usually improve results. so it is important to get VIT working",
      "votes": null
    },
    {
      "id": "2304035",
      "postDate": "06/15/2023 16:22:03",
      "content": "<p>Congrats on the solo gold and your new GM status! And many thanks for your public notebook that was an important starting point for many of us!</p>",
      "rawMarkdown": "Congrats on the solo gold and your new GM status! And many thanks for your public notebook that was an important starting point for many of us!",
      "votes": null
    },
    {
      "id": "2304812",
      "postDate": "06/16/2023 08:54:52",
      "content": "<p>Congrats on your gold medal! And I can't thank you enough for sharing your great train and inference code in the beginning of the competition. This was my first Kaggle competition, and the first time I have used PyTorch. Your code was so clean, and so well structured that it really helped me a lot in getting up to speed with PyTorch 🙏</p>",
      "rawMarkdown": "Congrats on your gold medal! And I can't thank you enough for sharing your great train and inference code in the beginning of the competition. This was my first Kaggle competition, and the first time I have used PyTorch. Your code was so clean, and so well structured that it really helped me a lot in getting up to speed with PyTorch 🙏",
      "votes": null
    },
    {
      "id": "2304885",
      "postDate": "06/16/2023 09:41:46",
      "content": "<p>Congrats on solo gold! The grouping trick is simple and awesome!</p>",
      "rawMarkdown": "Congrats on solo gold! The grouping trick is simple and awesome!",
      "votes": null
    },
    {
      "id": "2305056",
      "postDate": "06/16/2023 12:16:51",
      "content": "<p>Congrats 🧨</p>",
      "rawMarkdown": "Congrats 🧨",
      "votes": null
    },
    {
      "id": "2305075",
      "postDate": "06/16/2023 12:25:34",
      "content": "<blockquote>\n  <p>Grouping the images by threes.<br>\n  z_len = img_num / 3<br>\n  images (batch_size*z_len, 3, height, width) →2dcnn encoder →reshape to (batch_size, z_len, output_channel, height, width) →avg &gt; pooling(z_axis) → decoder</p>\n</blockquote>\n<p>I have been looking at this for a while now, but I can't figure out either what's the purpose or how it's working. Would you mind sharing some more information about this? <a href=\"https://www.kaggle.com/traptinblur\" target=\"_blank\">@traptinblur</a> <a href=\"https://www.kaggle.com/tanakar\" target=\"_blank\">@tanakar</a> 🙏</p>",
      "rawMarkdown": "> Grouping the images by threes.\n> z_len = img_num / 3\n> images (batch_size*z_len, 3, height, width) →2dcnn encoder →reshape to (batch_size, z_len, output_channel, height, width) →avg > pooling(z_axis) → decoder\n\nI have been looking at this for a while now, but I can't figure out either what's the purpose or how it's working. Would you mind sharing some more information about this? @traptinblur @tanakar 🙏",
      "votes": null
    },
    {
      "id": "2305085",
      "postDate": "06/16/2023 12:35:21",
      "content": "<p>it would be easier to understand the effect of \"mean pooling\" by considering the mean pooling used in 2d imagenet classifier model.</p>\n<p>for example in resnet,efficentnet, the structure is:<br>\nconv feature map (e.g. 2048x7x7)--&gt; mean pool (2048) --&gt; linear classsifier</p>\n<p>you can then use class activation heatmap  to check which of the 7x7 feature map (xy location) is activated for high logit score of the class.</p>\n<p>hence \"mean pool \" can select where is the signal (here, it is localising the signal in 2d for image classification).</p>\n<hr>\n<p>now for our case, we do not know where is the ink in z direction (from 0 to 65)<br>\nwe can let the model to select the signal automatically by mean pooling in the z direction.</p>\n<p>e.g. <br>\n3d conv feature map : 1x64x32x32 --&gt;2048x8x1x1<br>\nmean pool over 8 : 2048x8x1x1 --&gt; 2048x1x1x1<br>\nclassify : 2048 --&gt; 1</p>",
      "rawMarkdown": "it would be easier to understand the effect of \"mean pooling\" by considering the mean pooling used in 2d imagenet classifier model.\n\nfor example in resnet,efficentnet, the structure is:\nconv feature map (e.g. 2048x7x7)--> mean pool (2048) --> linear classsifier\n\nyou can then use class activation heatmap  to check which of the 7x7 feature map (xy location) is activated for high logit score of the class.\n\nhence \"mean pool \" can select where is the signal (here, it is localising the signal in 2d for image classification).\n\n---\n\nnow for our case, we do not know where is the ink in z direction (from 0 to 65)\nwe can let the model to select the signal automatically by mean pooling in the z direction.\n\ne.g. \n3d conv feature map : 1x64x32x32 -->2048x8x1x1\nmean pool over 8 : 2048x8x1x1 --> 2048x1x1x1\nclassify : 2048 --> 1",
      "votes": null
    },
    {
      "id": "2305225",
      "postDate": "06/16/2023 14:17:10",
      "content": "<p>yeah, <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> had explained very clear, and I would like to add some observations from my experiments.<br>\nWhen we simply add more slices at the channel dimension, the 2d models performed worse.<br>\nThere is a fact that when you stack voxel at batch dimension, it’s like increasing the batch size, all the voxel you stacked are processed in the same way.<br>\n2d encoders generally have 3 channels input for the first conv pretrained weight. So stacking 3-slice voxel for 10 and 7 neighboring groups at batch dimension can treat input voxel like normal 3 channels images but process the 30 and 21 slices at the same time.</p>",
      "rawMarkdown": "yeah, @hengck23 had explained very clear, and I would like to add some observations from my experiments.\nWhen we simply add more slices at the channel dimension, the 2d models performed worse.\nThere is a fact that when you stack voxel at batch dimension, it’s like increasing the batch size, all the voxel you stacked are processed in the same way.\n2d encoders generally have 3 channels input for the first conv pretrained weight. So stacking 3-slice voxel for 10 and 7 neighboring groups at batch dimension can treat input voxel like normal 3 channels images but process the 30 and 21 slices at the same time.",
      "votes": null
    },
    {
      "id": "2314169",
      "postDate": "06/23/2023 08:30:47",
      "content": "<p>Congrats! And thank you for sharing <a href=\"https://www.kaggle.com/code/tanakar/2-5d-segmentaion-baseline-training\" target=\"_blank\">https://www.kaggle.com/code/tanakar/2-5d-segmentaion-baseline-training</a></p>",
      "rawMarkdown": "Congrats! And thank you for sharing https://www.kaggle.com/code/tanakar/2-5d-segmentaion-baseline-training",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2303043,
      "author_name": "abualabed",
      "author_url": "",
      "post_date": "06/15/2023 03:57:39",
      "content": "<p>Congratulations….</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2303365,
      "author_name": "conjuring92",
      "author_url": "",
      "post_date": "06/15/2023 09:09:03",
      "content": "<p>Congratulations <a href=\"https://www.kaggle.com/tanakar\" target=\"_blank\">@tanakar</a> on becoming GM &amp; gold position finish 🎉🎉🎉</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2303823,
      "author_name": "chumajin",
      "author_url": "",
      "post_date": "06/15/2023 13:49:34",
      "content": "<p>Congratulations on your third solo gold! I'm thrilled to see you achieve Grandmaster status. In this competition, many people used your notebook, and furthermore, you have accomplished the remarkable feat of securing a solo gold medal! I respect you !</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2303835,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "06/15/2023 13:58:18",
      "content": "<p>\"not worked :segformer \"</p>\n<p>i design a segmentation model and it doesn't work. what should i do next?</p>\n<p>How to debug?</p>\n<ol>\n<li><p>use 3 channels + adam at small learning rate 1e-4,1e-5 <br>\nif it works, it means that  you cannot randomly initialise you first convolution channel to n input channel.<br>\nyou should initalise the values correctly.</p></li>\n<li><p>remove decoder and just train encoder (output 1/32 scale)<br>\nif it works, it means your decoder parameters or design are wrong</p></li>\n<li><p>try MIT-b0,b1,b2 …..<br>\nall needs different parameters/hyperparamters. Usually the shallowest model is easier to train (but may not lead to best results)</p></li>\n</ol>\n<hr>\n<p>performance is limited by data. since it works for other model, it means that your segformer hyperparamters are not correct.</p>\n<hr>\n<p>usually ensemble of VIT and CNN usually improve results. so it is important to get VIT working</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2304035,
      "author_name": "headsortails",
      "author_url": "",
      "post_date": "06/15/2023 16:22:03",
      "content": "<p>Congrats on the solo gold and your new GM status! And many thanks for your public notebook that was an important starting point for many of us!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2304812,
      "author_name": "lucasvw",
      "author_url": "",
      "post_date": "06/16/2023 08:54:52",
      "content": "<p>Congrats on your gold medal! And I can't thank you enough for sharing your great train and inference code in the beginning of the competition. This was my first Kaggle competition, and the first time I have used PyTorch. Your code was so clean, and so well structured that it really helped me a lot in getting up to speed with PyTorch 🙏</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2304885,
      "author_name": "traptinblur",
      "author_url": "",
      "post_date": "06/16/2023 09:41:46",
      "content": "<p>Congrats on solo gold! The grouping trick is simple and awesome!</p>",
      "votes": null,
      "replies": [
        {
          "id": 2305075,
          "author_name": "lucasvw",
          "author_url": "",
          "post_date": "06/16/2023 12:25:34",
          "content": "<blockquote>\n  <p>Grouping the images by threes.<br>\n  z_len = img_num / 3<br>\n  images (batch_size*z_len, 3, height, width) →2dcnn encoder →reshape to (batch_size, z_len, output_channel, height, width) →avg &gt; pooling(z_axis) → decoder</p>\n</blockquote>\n<p>I have been looking at this for a while now, but I can't figure out either what's the purpose or how it's working. Would you mind sharing some more information about this? <a href=\"https://www.kaggle.com/traptinblur\" target=\"_blank\">@traptinblur</a> <a href=\"https://www.kaggle.com/tanakar\" target=\"_blank\">@tanakar</a> 🙏</p>",
          "votes": null,
          "replies": [
            {
              "id": 2305085,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "06/16/2023 12:35:21",
              "content": "<p>it would be easier to understand the effect of \"mean pooling\" by considering the mean pooling used in 2d imagenet classifier model.</p>\n<p>for example in resnet,efficentnet, the structure is:<br>\nconv feature map (e.g. 2048x7x7)--&gt; mean pool (2048) --&gt; linear classsifier</p>\n<p>you can then use class activation heatmap  to check which of the 7x7 feature map (xy location) is activated for high logit score of the class.</p>\n<p>hence \"mean pool \" can select where is the signal (here, it is localising the signal in 2d for image classification).</p>\n<hr>\n<p>now for our case, we do not know where is the ink in z direction (from 0 to 65)<br>\nwe can let the model to select the signal automatically by mean pooling in the z direction.</p>\n<p>e.g. <br>\n3d conv feature map : 1x64x32x32 --&gt;2048x8x1x1<br>\nmean pool over 8 : 2048x8x1x1 --&gt; 2048x1x1x1<br>\nclassify : 2048 --&gt; 1</p>",
              "votes": null,
              "replies": []
            },
            {
              "id": 2305225,
              "author_name": "traptinblur",
              "author_url": "",
              "post_date": "06/16/2023 14:17:10",
              "content": "<p>yeah, <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> had explained very clear, and I would like to add some observations from my experiments.<br>\nWhen we simply add more slices at the channel dimension, the 2d models performed worse.<br>\nThere is a fact that when you stack voxel at batch dimension, it’s like increasing the batch size, all the voxel you stacked are processed in the same way.<br>\n2d encoders generally have 3 channels input for the first conv pretrained weight. So stacking 3-slice voxel for 10 and 7 neighboring groups at batch dimension can treat input voxel like normal 3 channels images but process the 30 and 21 slices at the same time.</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2305056,
      "author_name": "sinothmabasa",
      "author_url": "",
      "post_date": "06/16/2023 12:16:51",
      "content": "<p>Congrats 🧨</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2314169,
      "author_name": "riow1983",
      "author_url": "",
      "post_date": "06/23/2023 08:30:47",
      "content": "<p>Congrats! And thank you for sharing <a href=\"https://www.kaggle.com/code/tanakar/2-5d-segmentaion-baseline-training\" target=\"_blank\">https://www.kaggle.com/code/tanakar/2-5d-segmentaion-baseline-training</a></p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2303034": "Congrats to all the winners, and thanks to hosts for interesting competition!\nI’m very happy to become kaggle competition GM.\n\n# overview\n- 2.5d + 1d pool encoder\n- ensemble of 6 Unet models\n\n# model architecture\n\n- Grouping the images by threes.\n    - z_len = img_num / 3\n    - images (batch_size*z_len, 3, height, width) →2dcnn encoder →reshape to (batch_size, z_len, output_channel, height, width) →avg pooling(z_axis) → decoder\n- decoder: Unet\n\n| backbone | img_num |\n| ---- | ---- |\nse_resnext50_32x4d | 30\neca_nfnet_l1 | 30\ntf_efficientnetv2_m | 30\nconvnext_small_in22ft1k (best) | 30\nse_resnext101_32x4d | 21\nconvnext_base_in22ft1k | 21\n\n# training\n\ntraining pipeline is almost same to my public notebook.\n\nhttps://www.kaggle.com/code/tanakar/2-5d-segmentaion-baseline-training\n\n- img_size = 224\n- stride = 112\n- epoch = 30\n- fp16\n- 5fold\n    - ink_id=2 is divided into 3 parts\n- bce + dice loss\n- exclude areas with a mask value of 0\n- augmentation\n    - almost same to my public training notebook\n    - I added A.RandomRotate90 because rotating the image during inference improved public lb\n\n\n# inference\n\n- stride = 56\n- ensemble\n    - simple average\n- threshold = 0.5\n- exclude areas with img.sum() == 0\n\n# not worked\n\n- 3d encoder\n- segformer\n- other backbone\n- mixup\n- more slices\n\n# inference code\n\nhttps://www.kaggle.com/code/tanakar/2-5d-segmentaion-ensemble-final-stride-4",
    "2303043": "Congratulations….",
    "2303365": "Congratulations @tanakar on becoming GM & gold position finish 🎉🎉🎉",
    "2303823": "Congratulations on your third solo gold! I'm thrilled to see you achieve Grandmaster status. In this competition, many people used your notebook, and furthermore, you have accomplished the remarkable feat of securing a solo gold medal! I respect you !",
    "2303835": "\"not worked :segformer \"\n\ni design a segmentation model and it doesn't work. what should i do next?\n\nHow to debug?\n1. use 3 channels + adam at small learning rate 1e-4,1e-5 \nif it works, it means that  you cannot randomly initialise you first convolution channel to n input channel.\nyou should initalise the values correctly.\n\n2. remove decoder and just train encoder (output 1/32 scale)\nif it works, it means your decoder parameters or design are wrong\n\n3. try MIT-b0,b1,b2 .....\nall needs different parameters/hyperparamters. Usually the shallowest model is easier to train (but may not lead to best results)\n\n---\n\nperformance is limited by data. since it works for other model, it means that your segformer hyperparamters are not correct.\n\n---\n\nusually ensemble of VIT and CNN usually improve results. so it is important to get VIT working",
    "2304035": "Congrats on the solo gold and your new GM status! And many thanks for your public notebook that was an important starting point for many of us!",
    "2304812": "Congrats on your gold medal! And I can't thank you enough for sharing your great train and inference code in the beginning of the competition. This was my first Kaggle competition, and the first time I have used PyTorch. Your code was so clean, and so well structured that it really helped me a lot in getting up to speed with PyTorch 🙏",
    "2304885": "Congrats on solo gold! The grouping trick is simple and awesome!",
    "2305056": "Congrats 🧨",
    "2305075": "> Grouping the images by threes.\n> z_len = img_num / 3\n> images (batch_size*z_len, 3, height, width) →2dcnn encoder →reshape to (batch_size, z_len, output_channel, height, width) →avg > pooling(z_axis) → decoder\n\nI have been looking at this for a while now, but I can't figure out either what's the purpose or how it's working. Would you mind sharing some more information about this? @traptinblur @tanakar 🙏",
    "2305085": "it would be easier to understand the effect of \"mean pooling\" by considering the mean pooling used in 2d imagenet classifier model.\n\nfor example in resnet,efficentnet, the structure is:\nconv feature map (e.g. 2048x7x7)--> mean pool (2048) --> linear classsifier\n\nyou can then use class activation heatmap  to check which of the 7x7 feature map (xy location) is activated for high logit score of the class.\n\nhence \"mean pool \" can select where is the signal (here, it is localising the signal in 2d for image classification).\n\n---\n\nnow for our case, we do not know where is the ink in z direction (from 0 to 65)\nwe can let the model to select the signal automatically by mean pooling in the z direction.\n\ne.g. \n3d conv feature map : 1x64x32x32 -->2048x8x1x1\nmean pool over 8 : 2048x8x1x1 --> 2048x1x1x1\nclassify : 2048 --> 1",
    "2305225": "yeah, @hengck23 had explained very clear, and I would like to add some observations from my experiments.\nWhen we simply add more slices at the channel dimension, the 2d models performed worse.\nThere is a fact that when you stack voxel at batch dimension, it’s like increasing the batch size, all the voxel you stacked are processed in the same way.\n2d encoders generally have 3 channels input for the first conv pretrained weight. So stacking 3-slice voxel for 10 and 7 neighboring groups at batch dimension can treat input voxel like normal 3 channels images but process the 30 and 21 slices at the same time.",
    "2314169": "Congrats! And thank you for sharing https://www.kaggle.com/code/tanakar/2-5d-segmentaion-baseline-training"
  },
  "source": "meta"
}