{
  "id": 417255,
  "title": "2nd place solution",
  "url": "/competitions/vesuvius-challenge-ink-detection/writeups/rtx23090-2nd-place-solution",
  "author_name": "",
  "post_date": "2023-06-26T01:12:28.293Z",
  "votes": 76,
  "comment_count": 21,
  "views": 0,
  "content": "<h2>tattaka &amp; mipypf part</h2>\n<h3>Summary</h3>\n<ul>\n<li>2.5D and 3D backbone and decoder without upsampling<ul>\n<li>1/32 Resolution is sufficient for this task</li>\n<li>By focusing resources on encoder, more layers can be used for training</li></ul></li>\n<li>Strong regularization</li>\n</ul>\n<h3>Data preprocessing</h3>\n<ul>\n<li>Group K-fold is used for the validation, but is divided into 3 parts because of the large ink_id=2</li>\n<li>Input resolution is 256x256 for 2.5D model and 192x192 for 3D model<ul>\n<li>Cut out and store in npy in 32x32 (64x64 when inferring) to speed up data loading</li></ul></li>\n<li>The label is downsampled to 1/32 resolution using bilinear interpolation</li>\n<li>Input images were normalized for each sample　</li>\n</ul>\n<pre><code>  \n  mean = img.mean(dim=(, , ), keepdim=)\n  std = img.std(dim=(, , ), keepdim=) + \n  img = (img - mean) / std\n</code></pre>\n<h3>Model</h3>\n<ul>\n<li>images(batch_size x channel x group x height x width) -&gt; 2dcnn backbone -&gt; pointwise conv2d neck -&gt; 3dcnn(ResBlockCSN like, 3 or 6 blocks) -&gt; avg + max pooling(z axis) -&gt; pointwise conv2d<ul>\n<li>2dcnn backbone: <ul>\n<li>resnetrs50</li>\n<li>convnext_tiny</li>\n<li>swinv2_tiny</li>\n<li>resnext50</li></ul></li>\n<li>The channel/group combinations used were 5x7 and 3x9. In other words, the middle 35 or 27 layers of the 65 layers are used.</li>\n<li>Referring to <a href=\"https://www.kaggle.com/tereka\" target=\"_blank\">@tereka</a>'s <a href=\"https://www.kaggle.com/competitions/nfl-player-contact-detection/discussion/392402#2170010\" target=\"_blank\">model architecture</a></li></ul></li>\n<li>images(batch_size x 1 x layers x height x width) -&gt; 3dcnn backbone -&gt; max pooling(z axis) -&gt; pointwise conv2d<ul>\n<li>3DCNN backbone: <ul>\n<li>resnet50-irCSN(layers: 32)</li>\n<li>resnet152-irCSN(layers: 24)</li></ul></li></ul></li>\n<li>loss: bce + global fbeta loss(calculate fbeta score in batch)</li>\n</ul>\n<h3>Training</h3>\n<ul>\n<li>amp</li>\n<li>EMA(decay=0.99)</li>\n<li>label smoothing(smooth = 0.1)</li>\n<li>drop_path_rate=0.2</li>\n<li>cutmix + mixup + manifold mixup</li>\n<li>heavy augmentation<ul>\n<li>cutout<ul>\n<li>modified to match the output resolution (1/32)</li></ul></li>\n<li>channel shuffle<ul>\n<li>Shuffle within the group after splitting</li></ul></li>\n<li>Random +-2 shift in z-direction of volume with 0.5 probability  </li>\n<li>other transforms</li></ul></li>\n</ul>\n<pre><code>albu.Compose(\n          [\n              albu.Flip(=0.5),\n              albu.RandomRotate90(=0.9),\n              albu.ShiftScaleRotate(\n                  =0.0625,\n                  =0.2,\n                  =15,\n                  =0.9,\n              ),\n              albu.OneOf(\n                  [\n                      albu.ElasticTransform(=0.3),\n                      albu.GaussianBlur(=0.3),\n                      albu.GaussNoise(=0.3),\n                      albu.OpticalDistortion(=0.3),\n                      albu.GridDistortion(=0.1),\n                      albu.PiecewiseAffine(=0.3),  # IAAPiecewiseAffine\n                  ],\n                  =0.9,\n              ),\n              albu.RandomBrightnessContrast(\n                  =0.3, =0.3, =0.3\n              ),\n              ToTensorV2(),\n          ]\n      )\n</code></pre>\n<h3>Inference</h3>\n<ul>\n<li>fp16 inference</li>\n<li>stride=32</li>\n<li>Use weights learned on val_ink_id=(1, 2a)</li>\n<li>TTA<ul>\n<li>h/v flip</li>\n<li>Switch tta each time stride </li></ul></li>\n<li><a href=\"https://www.kaggle.com/philippsinger\" target=\"_blank\">@philippsinger</a> proposed <a href=\"https://www.kaggle.com/competitions/birdclef-2021/discussion/243463\" target=\"_blank\">percentile threshold</a><ul>\n<li>Always predicts the same amount of positives, so it is independent of model performance and depends on the pos/neg ratio of the GT</li>\n<li>Calculated for all pixels except the area outside the mask</li>\n<li>0.9 and 0.93 were used</li>\n<li>We expected 0.90 to be a better score in private, but 0.93 was better in the end!</li></ul></li>\n</ul>\n<h2>yukke42 part</h2>\n<h3>Summary</h3>\n<ul>\n<li>3D encoder and 2D/1D encoder<ul>\n<li>1/2 or 1/4 resolution prediction</li>\n<li>very simple decoder<br>\n​</li></ul></li>\n</ul>\n<h3>Data preprocessing</h3>\n<ul>\n<li>split the 2nd fragment into two image vertically: 4 folds<br>\n​</li>\n</ul>\n<h3>Model</h3>\n<ul>\n<li>3D CNN encoder and 2D Encoder<ul>\n<li>encoderbased on <a href=\"https://www.kaggle.com/samfc10\" target=\"_blank\">@samfc10</a> 's Notebook: <a href=\"https://www.kaggle.com/code/samfc10/vesuvius-challenge-3d-resnet-training\" target=\"_blank\">Vesuvius Challenge - 3D ResNet Training</a><ul>\n<li>remove maxpooling after the 1st CNN</li>\n<li>use attention before reduce D-dim</li>\n<li>use resnet18 and resnet34</li></ul></li>\n<li>decoder<ul>\n<li>a single 2D CNN layer</li>\n<li>upsample with a nearest interpolation</li></ul></li>\n<li>output resolution is downsampled to 1/2. then upsample with a bilinear interpolation</li></ul></li>\n<li>3D transformer encoder and linear decoder<ul>\n<li>encoder: use <a href=\"https://pytorch.org/vision/main/models/video_mvit.html\" target=\"_blank\">MViTv2-s</a> of the PyTorch official implementation and a pre-trained model<ul>\n<li>modify forward function to get each scale output</li>\n<li>replace MaxPool3d into MaxPool2d for the reproducibility</li></ul></li>\n<li>decoder: a single linear and patch expanding to upscale low resolutions 3D images<ul>\n<li>patch expanding is from <a href=\"https://arxiv.org/abs/2105.05537\" target=\"_blank\">Swin-Unet</a></li></ul></li>\n<li>output resolution is downsampled to 1/2 or 1/4, then upsample with a bilinear interpolation<br>\n​</li></ul></li>\n</ul>\n<h3>Training</h3>\n<ul>\n<li>amp</li>\n<li>torch.compile</li>\n<li>label smoothing</li>\n<li>cutout</li>\n<li>cutmix</li>\n<li>mixup</li>\n<li>data augmentation<ul>\n<li>referred <a href=\"https://www.kaggle.com/tanakar\" target=\"_blank\">@tanakar</a> 's notebook: <a href=\"https://www.kaggle.com/code/tanakar/2-5d-segmentaion-baseline-training\" target=\"_blank\">2.5d segmentaion baseline [training]</a></li>\n<li>referred <a href=\"https://www.kaggle.com/tattaka\" target=\"_blank\">@tattaka</a> and <a href=\"https://www.kaggle.com/mipypf\" target=\"_blank\">@mipypf</a> 's</li></ul></li>\n<li>patch_size=224 and stride=112<ul>\n<li>stride=75 or stride=56 didn't work<br>\n​</li></ul></li>\n</ul>\n<h3>Inference</h3>\n<ul>\n<li>fp16 inference</li>\n<li>stride=75<ul>\n<li>better than stride=112</li></ul></li>\n<li>ignore edge of output prediction<ul>\n<li>use only the red area of prediction (figure below)<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1745801%2Fc309c518581d183e62897d4249eb7e7e%2Fimage.png?generation=1686801349236365&amp;alt=media\" alt=\"\"></li></ul></li>\n</ul>\n<h2>Code &amp; inference notebook</h2>\n<ul>\n<li>tattaka &amp; ron part<ul>\n<li>training code: <a href=\"https://github.com/mipypf/Vesuvius-Challenge/tree/winner_call/tattaka_ron\" target=\"_blank\">https://github.com/mipypf/Vesuvius-Challenge/tree/winner_call/tattaka_ron</a></li>\n<li>inference code: <a href=\"https://www.kaggle.com/code/mipypf/ink-segmentation-2-5d-3dcnn-resnet3dcsn-fp16fold01?scriptVersionId=132226669\" target=\"_blank\">https://www.kaggle.com/code/mipypf/ink-segmentation-2-5d-3dcnn-resnet3dcsn-fp16fold01?scriptVersionId=132226669</a></li></ul></li>\n<li>yukke42 part<ul>\n<li>training code: <a href=\"https://github.com/yukke42/kaggle-vesuvius-challenge-ink-detection\" target=\"_blank\">https://github.com/yukke42/kaggle-vesuvius-challenge-ink-detection</a></li></ul></li>\n</ul>",
  "messages": [
    {
      "id": "2302890",
      "postDate": "06/15/2023 00:11:38",
      "content": "<h2>tattaka &amp; mipypf part</h2>\n<h3>Summary</h3>\n<ul>\n<li>2.5D and 3D backbone and decoder without upsampling<ul>\n<li>1/32 Resolution is sufficient for this task</li>\n<li>By focusing resources on encoder, more layers can be used for training</li></ul></li>\n<li>Strong regularization</li>\n</ul>\n<h3>Data preprocessing</h3>\n<ul>\n<li>Group K-fold is used for the validation, but is divided into 3 parts because of the large ink_id=2</li>\n<li>Input resolution is 256x256 for 2.5D model and 192x192 for 3D model<ul>\n<li>Cut out and store in npy in 32x32 (64x64 when inferring) to speed up data loading</li></ul></li>\n<li>The label is downsampled to 1/32 resolution using bilinear interpolation</li>\n<li>Input images were normalized for each sample　</li>\n</ul>\n<pre><code>  \n  mean = img.mean(dim=(, , ), keepdim=)\n  std = img.std(dim=(, , ), keepdim=) + \n  img = (img - mean) / std\n</code></pre>\n<h3>Model</h3>\n<ul>\n<li>images(batch_size x channel x group x height x width) -&gt; 2dcnn backbone -&gt; pointwise conv2d neck -&gt; 3dcnn(ResBlockCSN like, 3 or 6 blocks) -&gt; avg + max pooling(z axis) -&gt; pointwise conv2d<ul>\n<li>2dcnn backbone: <ul>\n<li>resnetrs50</li>\n<li>convnext_tiny</li>\n<li>swinv2_tiny</li>\n<li>resnext50</li></ul></li>\n<li>The channel/group combinations used were 5x7 and 3x9. In other words, the middle 35 or 27 layers of the 65 layers are used.</li>\n<li>Referring to <a href=\"https://www.kaggle.com/tereka\" target=\"_blank\">@tereka</a>'s <a href=\"https://www.kaggle.com/competitions/nfl-player-contact-detection/discussion/392402#2170010\" target=\"_blank\">model architecture</a></li></ul></li>\n<li>images(batch_size x 1 x layers x height x width) -&gt; 3dcnn backbone -&gt; max pooling(z axis) -&gt; pointwise conv2d<ul>\n<li>3DCNN backbone: <ul>\n<li>resnet50-irCSN(layers: 32)</li>\n<li>resnet152-irCSN(layers: 24)</li></ul></li></ul></li>\n<li>loss: bce + global fbeta loss(calculate fbeta score in batch)</li>\n</ul>\n<h3>Training</h3>\n<ul>\n<li>amp</li>\n<li>EMA(decay=0.99)</li>\n<li>label smoothing(smooth = 0.1)</li>\n<li>drop_path_rate=0.2</li>\n<li>cutmix + mixup + manifold mixup</li>\n<li>heavy augmentation<ul>\n<li>cutout<ul>\n<li>modified to match the output resolution (1/32)</li></ul></li>\n<li>channel shuffle<ul>\n<li>Shuffle within the group after splitting</li></ul></li>\n<li>Random +-2 shift in z-direction of volume with 0.5 probability  </li>\n<li>other transforms</li></ul></li>\n</ul>\n<pre><code>albu.Compose(\n          [\n              albu.Flip(=0.5),\n              albu.RandomRotate90(=0.9),\n              albu.ShiftScaleRotate(\n                  =0.0625,\n                  =0.2,\n                  =15,\n                  =0.9,\n              ),\n              albu.OneOf(\n                  [\n                      albu.ElasticTransform(=0.3),\n                      albu.GaussianBlur(=0.3),\n                      albu.GaussNoise(=0.3),\n                      albu.OpticalDistortion(=0.3),\n                      albu.GridDistortion(=0.1),\n                      albu.PiecewiseAffine(=0.3),  # IAAPiecewiseAffine\n                  ],\n                  =0.9,\n              ),\n              albu.RandomBrightnessContrast(\n                  =0.3, =0.3, =0.3\n              ),\n              ToTensorV2(),\n          ]\n      )\n</code></pre>\n<h3>Inference</h3>\n<ul>\n<li>fp16 inference</li>\n<li>stride=32</li>\n<li>Use weights learned on val_ink_id=(1, 2a)</li>\n<li>TTA<ul>\n<li>h/v flip</li>\n<li>Switch tta each time stride </li></ul></li>\n<li><a href=\"https://www.kaggle.com/philippsinger\" target=\"_blank\">@philippsinger</a> proposed <a href=\"https://www.kaggle.com/competitions/birdclef-2021/discussion/243463\" target=\"_blank\">percentile threshold</a><ul>\n<li>Always predicts the same amount of positives, so it is independent of model performance and depends on the pos/neg ratio of the GT</li>\n<li>Calculated for all pixels except the area outside the mask</li>\n<li>0.9 and 0.93 were used</li>\n<li>We expected 0.90 to be a better score in private, but 0.93 was better in the end!</li></ul></li>\n</ul>\n<h2>yukke42 part</h2>\n<h3>Summary</h3>\n<ul>\n<li>3D encoder and 2D/1D encoder<ul>\n<li>1/2 or 1/4 resolution prediction</li>\n<li>very simple decoder<br>\n​</li></ul></li>\n</ul>\n<h3>Data preprocessing</h3>\n<ul>\n<li>split the 2nd fragment into two image vertically: 4 folds<br>\n​</li>\n</ul>\n<h3>Model</h3>\n<ul>\n<li>3D CNN encoder and 2D Encoder<ul>\n<li>encoderbased on <a href=\"https://www.kaggle.com/samfc10\" target=\"_blank\">@samfc10</a> 's Notebook: <a href=\"https://www.kaggle.com/code/samfc10/vesuvius-challenge-3d-resnet-training\" target=\"_blank\">Vesuvius Challenge - 3D ResNet Training</a><ul>\n<li>remove maxpooling after the 1st CNN</li>\n<li>use attention before reduce D-dim</li>\n<li>use resnet18 and resnet34</li></ul></li>\n<li>decoder<ul>\n<li>a single 2D CNN layer</li>\n<li>upsample with a nearest interpolation</li></ul></li>\n<li>output resolution is downsampled to 1/2. then upsample with a bilinear interpolation</li></ul></li>\n<li>3D transformer encoder and linear decoder<ul>\n<li>encoder: use <a href=\"https://pytorch.org/vision/main/models/video_mvit.html\" target=\"_blank\">MViTv2-s</a> of the PyTorch official implementation and a pre-trained model<ul>\n<li>modify forward function to get each scale output</li>\n<li>replace MaxPool3d into MaxPool2d for the reproducibility</li></ul></li>\n<li>decoder: a single linear and patch expanding to upscale low resolutions 3D images<ul>\n<li>patch expanding is from <a href=\"https://arxiv.org/abs/2105.05537\" target=\"_blank\">Swin-Unet</a></li></ul></li>\n<li>output resolution is downsampled to 1/2 or 1/4, then upsample with a bilinear interpolation<br>\n​</li></ul></li>\n</ul>\n<h3>Training</h3>\n<ul>\n<li>amp</li>\n<li>torch.compile</li>\n<li>label smoothing</li>\n<li>cutout</li>\n<li>cutmix</li>\n<li>mixup</li>\n<li>data augmentation<ul>\n<li>referred <a href=\"https://www.kaggle.com/tanakar\" target=\"_blank\">@tanakar</a> 's notebook: <a href=\"https://www.kaggle.com/code/tanakar/2-5d-segmentaion-baseline-training\" target=\"_blank\">2.5d segmentaion baseline [training]</a></li>\n<li>referred <a href=\"https://www.kaggle.com/tattaka\" target=\"_blank\">@tattaka</a> and <a href=\"https://www.kaggle.com/mipypf\" target=\"_blank\">@mipypf</a> 's</li></ul></li>\n<li>patch_size=224 and stride=112<ul>\n<li>stride=75 or stride=56 didn't work<br>\n​</li></ul></li>\n</ul>\n<h3>Inference</h3>\n<ul>\n<li>fp16 inference</li>\n<li>stride=75<ul>\n<li>better than stride=112</li></ul></li>\n<li>ignore edge of output prediction<ul>\n<li>use only the red area of prediction (figure below)<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1745801%2Fc309c518581d183e62897d4249eb7e7e%2Fimage.png?generation=1686801349236365&amp;alt=media\" alt=\"\"></li></ul></li>\n</ul>\n<h2>Code &amp; inference notebook</h2>\n<ul>\n<li>tattaka &amp; ron part<ul>\n<li>training code: <a href=\"https://github.com/mipypf/Vesuvius-Challenge/tree/winner_call/tattaka_ron\" target=\"_blank\">https://github.com/mipypf/Vesuvius-Challenge/tree/winner_call/tattaka_ron</a></li>\n<li>inference code: <a href=\"https://www.kaggle.com/code/mipypf/ink-segmentation-2-5d-3dcnn-resnet3dcsn-fp16fold01?scriptVersionId=132226669\" target=\"_blank\">https://www.kaggle.com/code/mipypf/ink-segmentation-2-5d-3dcnn-resnet3dcsn-fp16fold01?scriptVersionId=132226669</a></li></ul></li>\n<li>yukke42 part<ul>\n<li>training code: <a href=\"https://github.com/yukke42/kaggle-vesuvius-challenge-ink-detection\" target=\"_blank\">https://github.com/yukke42/kaggle-vesuvius-challenge-ink-detection</a></li></ul></li>\n</ul>",
      "rawMarkdown": "## tattaka & mipypf part\n### Summary\n* 2.5D and 3D backbone and decoder without upsampling\n  * 1/32 Resolution is sufficient for this task\n  * By focusing resources on encoder, more layers can be used for training\n* Strong regularization\n### Data preprocessing\n* Group K-fold is used for the validation, but is divided into 3 parts because of the large ink_id=2\n* Input resolution is 256x256 for 2.5D model and 192x192 for 3D model\n  * Cut out and store in npy in 32x32 (64x64 when inferring) to speed up data loading\n* The label is downsampled to 1/32 resolution using bilinear interpolation\n* Input images were normalized for each sample　\n  ```python\n  # img: (bs, layers, h, w)\n  mean = img.mean(dim=(1, 2, 3), keepdim=True)\n  std = img.std(dim=(1, 2, 3), keepdim=True) + 1e-6\n  img = (img - mean) / std\n  ```\n### Model\n* images(batch_size x channel x group x height x width) -> 2dcnn backbone -> pointwise conv2d neck -> 3dcnn(ResBlockCSN like, 3 or 6 blocks) -> avg + max pooling(z axis) -> pointwise conv2d\n  * 2dcnn backbone: \n      * resnetrs50\n      * convnext_tiny\n      * swinv2_tiny\n      * resnext50\n  * The channel/group combinations used were 5x7 and 3x9. In other words, the middle 35 or 27 layers of the 65 layers are used.\n  * Referring to @tereka's [model architecture](https://www.kaggle.com/competitions/nfl-player-contact-detection/discussion/392402#2170010)\n* images(batch_size x 1 x layers x height x width) -> 3dcnn backbone -> max pooling(z axis) -> pointwise conv2d\n  * 3DCNN backbone: \n      * resnet50-irCSN(layers: 32)\n      * resnet152-irCSN(layers: 24)\n* loss: bce + global fbeta loss(calculate fbeta score in batch)\n### Training\n* amp\n* EMA(decay=0.99)\n* label smoothing(smooth = 0.1)\n* drop_path_rate=0.2\n* cutmix + mixup + manifold mixup\n* heavy augmentation\n  * cutout\n      * modified to match the output resolution (1/32)\n  * channel shuffle\n      * Shuffle within the group after splitting\n  * Random +-2 shift in z-direction of volume with 0.5 probability  \n  * other transforms\n```\nalbu.Compose(\n          [\n              albu.Flip(p=0.5),\n              albu.RandomRotate90(p=0.9),\n              albu.ShiftScaleRotate(\n                  shift_limit=0.0625,\n                  scale_limit=0.2,\n                  rotate_limit=15,\n                  p=0.9,\n              ),\n              albu.OneOf(\n                  [\n                      albu.ElasticTransform(p=0.3),\n                      albu.GaussianBlur(p=0.3),\n                      albu.GaussNoise(p=0.3),\n                      albu.OpticalDistortion(p=0.3),\n                      albu.GridDistortion(p=0.1),\n                      albu.PiecewiseAffine(p=0.3),  # IAAPiecewiseAffine\n                  ],\n                  p=0.9,\n              ),\n              albu.RandomBrightnessContrast(\n                  brightness_limit=0.3, contrast_limit=0.3, p=0.3\n              ),\n              ToTensorV2(),\n          ]\n      )\n```\n### Inference\n* fp16 inference\n* stride=32\n* Use weights learned on val_ink_id=(1, 2a)\n* TTA\n  * h/v flip\n  * Switch tta each time stride \n* @philippsinger proposed [percentile threshold](https://www.kaggle.com/competitions/birdclef-2021/discussion/243463)\n  * Always predicts the same amount of positives, so it is independent of model performance and depends on the pos/neg ratio of the GT\n  * Calculated for all pixels except the area outside the mask\n  * 0.9 and 0.93 were used\n  * We expected 0.90 to be a better score in private, but 0.93 was better in the end!\n\n## yukke42 part\n### Summary\n- 3D encoder and 2D/1D encoder\n  - 1/2 or 1/4 resolution prediction\n  - very simple decoder\n​\n### Data preprocessing\n- split the 2nd fragment into two image vertically: 4 folds\n​\n### Model\n- 3D CNN encoder and 2D Encoder\n  - encoderbased on @samfc10 's Notebook: [Vesuvius Challenge - 3D ResNet Training](https://www.kaggle.com/code/samfc10/vesuvius-challenge-3d-resnet-training)\n      - remove maxpooling after the 1st CNN\n      - use attention before reduce D-dim\n      - use resnet18 and resnet34\n  - decoder\n      - a single 2D CNN layer\n      - upsample with a nearest interpolation\n  - output resolution is downsampled to 1/2. then upsample with a bilinear interpolation\n- 3D transformer encoder and linear decoder\n  - encoder: use [MViTv2-s](https://pytorch.org/vision/main/models/video_mvit.html) of the PyTorch official implementation and a pre-trained model\n      - modify forward function to get each scale output\n      - replace MaxPool3d into MaxPool2d for the reproducibility\n  - decoder: a single linear and patch expanding to upscale low resolutions 3D images\n      - patch expanding is from [Swin-Unet](https://arxiv.org/abs/2105.05537)\n  - output resolution is downsampled to 1/2 or 1/4, then upsample with a bilinear interpolation\n​\n### Training\n- amp\n- torch.compile\n- label smoothing\n- cutout\n- cutmix\n- mixup\n- data augmentation\n  - referred @tanakar 's notebook: [2.5d segmentaion baseline [training]](https://www.kaggle.com/code/tanakar/2-5d-segmentaion-baseline-training)\n  - referred @tattaka and @mipypf 's\n- patch_size=224 and stride=112\n  - stride=75 or stride=56 didn't work\n​\n### Inference\n- fp16 inference\n- stride=75\n  - better than stride=112\n- ignore edge of output prediction\n  - use only the red area of prediction (figure below)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1745801%2Fc309c518581d183e62897d4249eb7e7e%2Fimage.png?generation=1686801349236365&alt=media)\n\n\n## Code & inference notebook\n* tattaka & ron part\n    * training code: https://github.com/mipypf/Vesuvius-Challenge/tree/winner_call/tattaka_ron\n    * inference code: https://www.kaggle.com/code/mipypf/ink-segmentation-2-5d-3dcnn-resnet3dcsn-fp16fold01?scriptVersionId=132226669\n* yukke42 part\n    * training code: https://github.com/yukke42/kaggle-vesuvius-challenge-ink-detection",
      "votes": null
    },
    {
      "id": "2302910",
      "postDate": "06/15/2023 00:45:06",
      "content": "<p>congrats for the good results and high ranking!</p>\n<p>if possible, i would like to know your split and local CV score.<br>\nI was wondering if anyone can actually get 0.80 in CV ?</p>",
      "rawMarkdown": "congrats for the good results and high ranking!\n\nif possible, i would like to know your split and local CV score.\nI was wondering if anyone can actually get 0.80 in CV ?",
      "votes": null
    },
    {
      "id": "2302912",
      "postDate": "06/15/2023 00:47:18",
      "content": "<p>The local score for val_ink_id=1 was about 0.69 in ensemble (also 0.7, but failed to submit)</p>",
      "rawMarkdown": "The local score for val_ink_id=1 was about 0.69 in ensemble (also 0.7, but failed to submit)",
      "votes": null
    },
    {
      "id": "2302927",
      "postDate": "06/15/2023 00:59:20",
      "content": "<p>Congratulations for the 2nd rank.</p>\n<blockquote>\n  <p>We expected 0.90 to be a better score in private, but 0.93 was better in the end!</p>\n</blockquote>\n<p>Were best percentiles different between test fragment a and b?<br>\nIf so, how can you estimate the best percentile of fragment b?</p>",
      "rawMarkdown": "Congratulations for the 2nd rank.\n\n> We expected 0.90 to be a better score in private, but 0.93 was better in the end!\n\nWere best percentiles different between test fragment a and b?\nIf so, how can you estimate the best percentile of fragment b?",
      "votes": null
    },
    {
      "id": "2302930",
      "postDate": "06/15/2023 01:10:07",
      "content": "<p>wow 0.69 for ink_id = 1, super impressive, congrats!</p>",
      "rawMarkdown": "wow 0.69 for ink_id = 1, super impressive, congrats!",
      "votes": null
    },
    {
      "id": "2302931",
      "postDate": "06/15/2023 01:16:12",
      "content": "<p>Looking at our results, it seems that the optimal thresholds for fragment_a and fragment_b are not that different.<br>\nHowever, we were unable to find any evidence to support it during the competition.</p>",
      "rawMarkdown": "Looking at our results, it seems that the optimal thresholds for fragment_a and fragment_b are not that different.\nHowever, we were unable to find any evidence to support it during the competition.",
      "votes": null
    },
    {
      "id": "2302962",
      "postDate": "06/15/2023 01:51:57",
      "content": "<p>Congratulations for the great results</p>",
      "rawMarkdown": "Congratulations for the great results",
      "votes": null
    },
    {
      "id": "2303017",
      "postDate": "06/15/2023 03:08:00",
      "content": "<p>what is resnet50-irCSN?</p>",
      "rawMarkdown": "what is resnet50-irCSN?",
      "votes": null
    },
    {
      "id": "2303068",
      "postDate": "06/15/2023 04:21:57",
      "content": "<p><a href=\"https://github.com/open-mmlab/mmaction2/blob/main/configs/recognition/csn/README.md\" target=\"_blank\">https://github.com/open-mmlab/mmaction2/blob/main/configs/recognition/csn/README.md</a></p>",
      "rawMarkdown": "https://github.com/open-mmlab/mmaction2/blob/main/configs/recognition/csn/README.md",
      "votes": null
    },
    {
      "id": "2303163",
      "postDate": "06/15/2023 05:57:55",
      "content": "<p><a href=\"https://www.kaggle.com/tattaka\" target=\"_blank\">@tattaka</a> Congratulations on the shake-up to second place! That's an impressive achievement. It seems that using many layers resulted in a more robust model. I also learned the importance of applying strong augmentations.</p>\n<pre><code>/ Resolution  sufficient   task\n</code></pre>\n<p>I also had a similar idea, but I didn't try it out. I should have done it from the beginning.</p>\n<p>Perhaps, by reducing the resolution, the original text, which was recognized as 1000 pixels x 1000 pixels, became smaller and It might be that it contributed to improved recognition of the characters. This is just my speculation.</p>",
      "rawMarkdown": "tattaka Congratulations on the shake-up to second place! That's an impressive achievement. It seems that using many layers resulted in a more robust model. I also learned the importance of applying strong augmentations.\n\n\n~~~\n1/32 Resolution is sufficient for this task\n~~~\n\nI also had a similar idea, but I didn't try it out. I should have done it from the beginning.\n\nPerhaps, by reducing the resolution, the original text, which was recognized as 1000 pixels x 1000 pixels, became smaller and It might be that it contributed to improved recognition of the characters. This is just my speculation.",
      "votes": null
    },
    {
      "id": "2303221",
      "postDate": "06/15/2023 07:06:33",
      "content": "<p>Before starting the competition, when I read the host's <a href=\"https://github.com/educelab/ink-id/blob/develop/inkid/scripts/train_and_predict.py#L134\" target=\"_blank\">inkid repo</a>, I found that they didn't use upsampling.<br>\nSo when I actually calculated the maximum score that could be achieved by downsampling the label and undoing it, I found that I could get 0.98, so I knew that 1/32 resolution was enough. <br>\nAlso, by lowering the resolution, the difficulty of the task is also lowered, and we believe that we are able to reduce the burden on the machine learning model.</p>",
      "rawMarkdown": "Before starting the competition, when I read the host's [inkid repo](https://github.com/educelab/ink-id/blob/develop/inkid/scripts/train_and_predict.py#L134), I found that they didn't use upsampling.\nSo when I actually calculated the maximum score that could be achieved by downsampling the label and undoing it, I found that I could get 0.98, so I knew that 1/32 resolution was enough. \nAlso, by lowering the resolution, the difficulty of the task is also lowered, and we believe that we are able to reduce the burden on the machine learning model.",
      "votes": null
    },
    {
      "id": "2303779",
      "postDate": "06/15/2023 13:22:01",
      "content": "<p>Great method of confirmation! I think the benefits of this are incredible. Thank you for sharing.</p>",
      "rawMarkdown": "Great method of confirmation! I think the benefits of this are incredible. Thank you for sharing.",
      "votes": null
    },
    {
      "id": "2303791",
      "postDate": "06/15/2023 13:27:40",
      "content": "<p>in my experiments, just train a pvtv2-b3 transformer and use its feature map (1/32 scale) gives about 0.65 lb without tunning. validation is fragment one.</p>\n<p>but not all backbone work.</p>",
      "rawMarkdown": "in my experiments, just train a pvtv2-b3 transformer and use its feature map (1/32 scale) gives about 0.65 lb without tunning. validation is fragment one.\n\nbut not all backbone work.",
      "votes": null
    },
    {
      "id": "2303914",
      "postDate": "06/15/2023 14:49:28",
      "content": "<p>Congratulations!</p>",
      "rawMarkdown": "Congratulations!",
      "votes": null
    },
    {
      "id": "2304587",
      "postDate": "06/16/2023 05:22:46",
      "content": "<p><a href=\"https://www.kaggle.com/tattaka\" target=\"_blank\">@tattaka</a> </p>\n<p>I have another question:</p>\n<blockquote>\n  <p>1/32 Resolution is sufficient for this task</p>\n</blockquote>\n<p>Does this mean you are feeding 2d encoders final layer's activations (typically output_stride is 32) directly to the 3d encoder (i.e. without decoder)?</p>",
      "rawMarkdown": "tattaka \n\nI have another question:\n\n> 1/32 Resolution is sufficient for this task\n\nDoes this mean you are feeding 2d encoders final layer's activations (typically output_stride is 32) directly to the 3d encoder (i.e. without decoder)?",
      "votes": null
    },
    {
      "id": "2304721",
      "postDate": "06/16/2023 07:37:08",
      "content": "<p>Yes, let me explain in a little more detail.<br>\nThe input of the 2D encoder is (bs * groups, ch, h, w) and the output after going through the neck is (bs * groups, 512, h//32, w//32). Then permute it to (bs, ch, groups, h//32, w//32) and feed it to 3D CNN.</p>",
      "rawMarkdown": "Yes, let me explain in a little more detail.\nThe input of the 2D encoder is (bs * groups, ch, h, w) and the output after going through the neck is (bs * groups, 512, h//32, w//32). Then permute it to (bs, ch, groups, h//32, w//32) and feed it to 3D CNN.",
      "votes": null
    },
    {
      "id": "2304894",
      "postDate": "06/16/2023 09:54:05",
      "content": "<p>congrats for better results</p>",
      "rawMarkdown": "congrats for better results",
      "votes": null
    },
    {
      "id": "2305926",
      "postDate": "06/17/2023 02:25:18",
      "content": "<p>Congratulations!!!</p>",
      "rawMarkdown": "Congratulations!!!",
      "votes": null
    },
    {
      "id": "2310568",
      "postDate": "06/20/2023 13:22:05",
      "content": "<p>Congratulations on the success! The end was a photo finish! I hope you will post more details of your solution here or on Github or another similar platform.</p>",
      "rawMarkdown": "Congratulations on the success! The end was a photo finish! I hope you will post more details of your solution here or on Github or another similar platform.",
      "votes": null
    },
    {
      "id": "2315412",
      "postDate": "06/24/2023 05:54:07",
      "content": "<p>This is impressive! <br>\nI tried feeding (bs, groups, h, w) to 2D encoder (changing to -&gt; (bs, output channel, h, w)) and subsequently to 3D encoder but this approach has problem to lose temporal information. But using your method, it can be possible to retain temporal information and segmentation information at the same time.<br>\nRespect for your idea!</p>",
      "rawMarkdown": "This is impressive! \nI tried feeding (bs, groups, h, w) to 2D encoder (changing to -> (bs, output channel, h, w)) and subsequently to 3D encoder but this approach has problem to lose temporal information. But using your method, it can be possible to retain temporal information and segmentation information at the same time.\nRespect for your idea!",
      "votes": null
    },
    {
      "id": "2317125",
      "postDate": "06/25/2023 13:11:54",
      "content": "<p><strong><em>Update</em></strong><br>\nSubmission of tattaka and ron's part was set to public!</p>\n<ul>\n<li>training code: <a href=\"https://github.com/mipypf/Vesuvius-Challenge/tree/winner_call/tattaka_ron\" target=\"_blank\">https://github.com/mipypf/Vesuvius-Challenge/tree/winner_call/tattaka_ron</a></li>\n<li>inference code: <a href=\"https://www.kaggle.com/code/mipypf/ink-segmentation-2-5d-3dcnn-resnet3dcsn-fp16fold01?scriptVersionId=132226669\" target=\"_blank\">https://www.kaggle.com/code/mipypf/ink-segmentation-2-5d-3dcnn-resnet3dcsn-fp16fold01?scriptVersionId=132226669</a></li>\n</ul>",
      "rawMarkdown": "***Update***\nSubmission of tattaka and ron's part was set to public!\n* training code: https://github.com/mipypf/Vesuvius-Challenge/tree/winner_call/tattaka_ron\n* inference code: https://www.kaggle.com/code/mipypf/ink-segmentation-2-5d-3dcnn-resnet3dcsn-fp16fold01?scriptVersionId=132226669",
      "votes": null
    },
    {
      "id": "2317273",
      "postDate": "06/25/2023 15:02:11",
      "content": "<p>Update<br>\nyukke42's training code is also public! <br>\n<a href=\"https://github.com/yukke42/kaggle-vesuvius-challenge-ink-detection\" target=\"_blank\">https://github.com/yukke42/kaggle-vesuvius-challenge-ink-detection</a></p>",
      "rawMarkdown": "Update\nyukke42's training code is also public! \nhttps://github.com/yukke42/kaggle-vesuvius-challenge-ink-detection",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2302910,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "06/15/2023 00:45:06",
      "content": "<p>congrats for the good results and high ranking!</p>\n<p>if possible, i would like to know your split and local CV score.<br>\nI was wondering if anyone can actually get 0.80 in CV ?</p>",
      "votes": null,
      "replies": [
        {
          "id": 2302912,
          "author_name": "tattaka",
          "author_url": "",
          "post_date": "06/15/2023 00:47:18",
          "content": "<p>The local score for val_ink_id=1 was about 0.69 in ensemble (also 0.7, but failed to submit)</p>",
          "votes": null,
          "replies": [
            {
              "id": 2302930,
              "author_name": "fengqilong",
              "author_url": "",
              "post_date": "06/15/2023 01:10:07",
              "content": "<p>wow 0.69 for ink_id = 1, super impressive, congrats!</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2302927,
      "author_name": "tatamikenn",
      "author_url": "",
      "post_date": "06/15/2023 00:59:20",
      "content": "<p>Congratulations for the 2nd rank.</p>\n<blockquote>\n  <p>We expected 0.90 to be a better score in private, but 0.93 was better in the end!</p>\n</blockquote>\n<p>Were best percentiles different between test fragment a and b?<br>\nIf so, how can you estimate the best percentile of fragment b?</p>",
      "votes": null,
      "replies": [
        {
          "id": 2302931,
          "author_name": "tattaka",
          "author_url": "",
          "post_date": "06/15/2023 01:16:12",
          "content": "<p>Looking at our results, it seems that the optimal thresholds for fragment_a and fragment_b are not that different.<br>\nHowever, we were unable to find any evidence to support it during the competition.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2302962,
      "author_name": "robsonsan",
      "author_url": "",
      "post_date": "06/15/2023 01:51:57",
      "content": "<p>Congratulations for the great results</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2303017,
      "author_name": "cutebomb",
      "author_url": "",
      "post_date": "06/15/2023 03:08:00",
      "content": "<p>what is resnet50-irCSN?</p>",
      "votes": null,
      "replies": [
        {
          "id": 2303068,
          "author_name": "tattaka",
          "author_url": "",
          "post_date": "06/15/2023 04:21:57",
          "content": "<p><a href=\"https://github.com/open-mmlab/mmaction2/blob/main/configs/recognition/csn/README.md\" target=\"_blank\">https://github.com/open-mmlab/mmaction2/blob/main/configs/recognition/csn/README.md</a></p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2303163,
      "author_name": "chumajin",
      "author_url": "",
      "post_date": "06/15/2023 05:57:55",
      "content": "<p><a href=\"https://www.kaggle.com/tattaka\" target=\"_blank\">@tattaka</a> Congratulations on the shake-up to second place! That's an impressive achievement. It seems that using many layers resulted in a more robust model. I also learned the importance of applying strong augmentations.</p>\n<pre><code>/ Resolution  sufficient   task\n</code></pre>\n<p>I also had a similar idea, but I didn't try it out. I should have done it from the beginning.</p>\n<p>Perhaps, by reducing the resolution, the original text, which was recognized as 1000 pixels x 1000 pixels, became smaller and It might be that it contributed to improved recognition of the characters. This is just my speculation.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2303221,
          "author_name": "tattaka",
          "author_url": "",
          "post_date": "06/15/2023 07:06:33",
          "content": "<p>Before starting the competition, when I read the host's <a href=\"https://github.com/educelab/ink-id/blob/develop/inkid/scripts/train_and_predict.py#L134\" target=\"_blank\">inkid repo</a>, I found that they didn't use upsampling.<br>\nSo when I actually calculated the maximum score that could be achieved by downsampling the label and undoing it, I found that I could get 0.98, so I knew that 1/32 resolution was enough. <br>\nAlso, by lowering the resolution, the difficulty of the task is also lowered, and we believe that we are able to reduce the burden on the machine learning model.</p>",
          "votes": null,
          "replies": [
            {
              "id": 2303779,
              "author_name": "chumajin",
              "author_url": "",
              "post_date": "06/15/2023 13:22:01",
              "content": "<p>Great method of confirmation! I think the benefits of this are incredible. Thank you for sharing.</p>",
              "votes": null,
              "replies": [
                {
                  "id": 2303791,
                  "author_name": "hengck23",
                  "author_url": "",
                  "post_date": "06/15/2023 13:27:40",
                  "content": "<p>in my experiments, just train a pvtv2-b3 transformer and use its feature map (1/32 scale) gives about 0.65 lb without tunning. validation is fragment one.</p>\n<p>but not all backbone work.</p>",
                  "votes": null,
                  "replies": []
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 2303914,
      "author_name": "ahnheeyoung1",
      "author_url": "",
      "post_date": "06/15/2023 14:49:28",
      "content": "<p>Congratulations!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2304587,
      "author_name": "tatamikenn",
      "author_url": "",
      "post_date": "06/16/2023 05:22:46",
      "content": "<p><a href=\"https://www.kaggle.com/tattaka\" target=\"_blank\">@tattaka</a> </p>\n<p>I have another question:</p>\n<blockquote>\n  <p>1/32 Resolution is sufficient for this task</p>\n</blockquote>\n<p>Does this mean you are feeding 2d encoders final layer's activations (typically output_stride is 32) directly to the 3d encoder (i.e. without decoder)?</p>",
      "votes": null,
      "replies": [
        {
          "id": 2304721,
          "author_name": "tattaka",
          "author_url": "",
          "post_date": "06/16/2023 07:37:08",
          "content": "<p>Yes, let me explain in a little more detail.<br>\nThe input of the 2D encoder is (bs * groups, ch, h, w) and the output after going through the neck is (bs * groups, 512, h//32, w//32). Then permute it to (bs, ch, groups, h//32, w//32) and feed it to 3D CNN.</p>",
          "votes": null,
          "replies": [
            {
              "id": 2315412,
              "author_name": "ahnheeyoung1",
              "author_url": "",
              "post_date": "06/24/2023 05:54:07",
              "content": "<p>This is impressive! <br>\nI tried feeding (bs, groups, h, w) to 2D encoder (changing to -&gt; (bs, output channel, h, w)) and subsequently to 3D encoder but this approach has problem to lose temporal information. But using your method, it can be possible to retain temporal information and segmentation information at the same time.<br>\nRespect for your idea!</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2304894,
      "author_name": "shubhamrai2104",
      "author_url": "",
      "post_date": "06/16/2023 09:54:05",
      "content": "<p>congrats for better results</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2305926,
      "author_name": "lthoon",
      "author_url": "",
      "post_date": "06/17/2023 02:25:18",
      "content": "<p>Congratulations!!!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2310568,
      "author_name": "guillermoperezg",
      "author_url": "",
      "post_date": "06/20/2023 13:22:05",
      "content": "<p>Congratulations on the success! The end was a photo finish! I hope you will post more details of your solution here or on Github or another similar platform.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2317125,
      "author_name": "tattaka",
      "author_url": "",
      "post_date": "06/25/2023 13:11:54",
      "content": "<p><strong><em>Update</em></strong><br>\nSubmission of tattaka and ron's part was set to public!</p>\n<ul>\n<li>training code: <a href=\"https://github.com/mipypf/Vesuvius-Challenge/tree/winner_call/tattaka_ron\" target=\"_blank\">https://github.com/mipypf/Vesuvius-Challenge/tree/winner_call/tattaka_ron</a></li>\n<li>inference code: <a href=\"https://www.kaggle.com/code/mipypf/ink-segmentation-2-5d-3dcnn-resnet3dcsn-fp16fold01?scriptVersionId=132226669\" target=\"_blank\">https://www.kaggle.com/code/mipypf/ink-segmentation-2-5d-3dcnn-resnet3dcsn-fp16fold01?scriptVersionId=132226669</a></li>\n</ul>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2317273,
      "author_name": "yukke42",
      "author_url": "",
      "post_date": "06/25/2023 15:02:11",
      "content": "<p>Update<br>\nyukke42's training code is also public! <br>\n<a href=\"https://github.com/yukke42/kaggle-vesuvius-challenge-ink-detection\" target=\"_blank\">https://github.com/yukke42/kaggle-vesuvius-challenge-ink-detection</a></p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2302890": "## tattaka & mipypf part\n### Summary\n* 2.5D and 3D backbone and decoder without upsampling\n  * 1/32 Resolution is sufficient for this task\n  * By focusing resources on encoder, more layers can be used for training\n* Strong regularization\n### Data preprocessing\n* Group K-fold is used for the validation, but is divided into 3 parts because of the large ink_id=2\n* Input resolution is 256x256 for 2.5D model and 192x192 for 3D model\n  * Cut out and store in npy in 32x32 (64x64 when inferring) to speed up data loading\n* The label is downsampled to 1/32 resolution using bilinear interpolation\n* Input images were normalized for each sample　\n  ```python\n  # img: (bs, layers, h, w)\n  mean = img.mean(dim=(1, 2, 3), keepdim=True)\n  std = img.std(dim=(1, 2, 3), keepdim=True) + 1e-6\n  img = (img - mean) / std\n  ```\n### Model\n* images(batch_size x channel x group x height x width) -> 2dcnn backbone -> pointwise conv2d neck -> 3dcnn(ResBlockCSN like, 3 or 6 blocks) -> avg + max pooling(z axis) -> pointwise conv2d\n  * 2dcnn backbone: \n      * resnetrs50\n      * convnext_tiny\n      * swinv2_tiny\n      * resnext50\n  * The channel/group combinations used were 5x7 and 3x9. In other words, the middle 35 or 27 layers of the 65 layers are used.\n  * Referring to @tereka's [model architecture](https://www.kaggle.com/competitions/nfl-player-contact-detection/discussion/392402#2170010)\n* images(batch_size x 1 x layers x height x width) -> 3dcnn backbone -> max pooling(z axis) -> pointwise conv2d\n  * 3DCNN backbone: \n      * resnet50-irCSN(layers: 32)\n      * resnet152-irCSN(layers: 24)\n* loss: bce + global fbeta loss(calculate fbeta score in batch)\n### Training\n* amp\n* EMA(decay=0.99)\n* label smoothing(smooth = 0.1)\n* drop_path_rate=0.2\n* cutmix + mixup + manifold mixup\n* heavy augmentation\n  * cutout\n      * modified to match the output resolution (1/32)\n  * channel shuffle\n      * Shuffle within the group after splitting\n  * Random +-2 shift in z-direction of volume with 0.5 probability  \n  * other transforms\n```\nalbu.Compose(\n          [\n              albu.Flip(p=0.5),\n              albu.RandomRotate90(p=0.9),\n              albu.ShiftScaleRotate(\n                  shift_limit=0.0625,\n                  scale_limit=0.2,\n                  rotate_limit=15,\n                  p=0.9,\n              ),\n              albu.OneOf(\n                  [\n                      albu.ElasticTransform(p=0.3),\n                      albu.GaussianBlur(p=0.3),\n                      albu.GaussNoise(p=0.3),\n                      albu.OpticalDistortion(p=0.3),\n                      albu.GridDistortion(p=0.1),\n                      albu.PiecewiseAffine(p=0.3),  # IAAPiecewiseAffine\n                  ],\n                  p=0.9,\n              ),\n              albu.RandomBrightnessContrast(\n                  brightness_limit=0.3, contrast_limit=0.3, p=0.3\n              ),\n              ToTensorV2(),\n          ]\n      )\n```\n### Inference\n* fp16 inference\n* stride=32\n* Use weights learned on val_ink_id=(1, 2a)\n* TTA\n  * h/v flip\n  * Switch tta each time stride \n* @philippsinger proposed [percentile threshold](https://www.kaggle.com/competitions/birdclef-2021/discussion/243463)\n  * Always predicts the same amount of positives, so it is independent of model performance and depends on the pos/neg ratio of the GT\n  * Calculated for all pixels except the area outside the mask\n  * 0.9 and 0.93 were used\n  * We expected 0.90 to be a better score in private, but 0.93 was better in the end!\n\n## yukke42 part\n### Summary\n- 3D encoder and 2D/1D encoder\n  - 1/2 or 1/4 resolution prediction\n  - very simple decoder\n​\n### Data preprocessing\n- split the 2nd fragment into two image vertically: 4 folds\n​\n### Model\n- 3D CNN encoder and 2D Encoder\n  - encoderbased on @samfc10 's Notebook: [Vesuvius Challenge - 3D ResNet Training](https://www.kaggle.com/code/samfc10/vesuvius-challenge-3d-resnet-training)\n      - remove maxpooling after the 1st CNN\n      - use attention before reduce D-dim\n      - use resnet18 and resnet34\n  - decoder\n      - a single 2D CNN layer\n      - upsample with a nearest interpolation\n  - output resolution is downsampled to 1/2. then upsample with a bilinear interpolation\n- 3D transformer encoder and linear decoder\n  - encoder: use [MViTv2-s](https://pytorch.org/vision/main/models/video_mvit.html) of the PyTorch official implementation and a pre-trained model\n      - modify forward function to get each scale output\n      - replace MaxPool3d into MaxPool2d for the reproducibility\n  - decoder: a single linear and patch expanding to upscale low resolutions 3D images\n      - patch expanding is from [Swin-Unet](https://arxiv.org/abs/2105.05537)\n  - output resolution is downsampled to 1/2 or 1/4, then upsample with a bilinear interpolation\n​\n### Training\n- amp\n- torch.compile\n- label smoothing\n- cutout\n- cutmix\n- mixup\n- data augmentation\n  - referred @tanakar 's notebook: [2.5d segmentaion baseline [training]](https://www.kaggle.com/code/tanakar/2-5d-segmentaion-baseline-training)\n  - referred @tattaka and @mipypf 's\n- patch_size=224 and stride=112\n  - stride=75 or stride=56 didn't work\n​\n### Inference\n- fp16 inference\n- stride=75\n  - better than stride=112\n- ignore edge of output prediction\n  - use only the red area of prediction (figure below)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1745801%2Fc309c518581d183e62897d4249eb7e7e%2Fimage.png?generation=1686801349236365&alt=media)\n\n\n## Code & inference notebook\n* tattaka & ron part\n    * training code: https://github.com/mipypf/Vesuvius-Challenge/tree/winner_call/tattaka_ron\n    * inference code: https://www.kaggle.com/code/mipypf/ink-segmentation-2-5d-3dcnn-resnet3dcsn-fp16fold01?scriptVersionId=132226669\n* yukke42 part\n    * training code: https://github.com/yukke42/kaggle-vesuvius-challenge-ink-detection",
    "2302910": "congrats for the good results and high ranking!\n\nif possible, i would like to know your split and local CV score.\nI was wondering if anyone can actually get 0.80 in CV ?",
    "2302912": "The local score for val_ink_id=1 was about 0.69 in ensemble (also 0.7, but failed to submit)",
    "2302927": "Congratulations for the 2nd rank.\n\n> We expected 0.90 to be a better score in private, but 0.93 was better in the end!\n\nWere best percentiles different between test fragment a and b?\nIf so, how can you estimate the best percentile of fragment b?",
    "2302930": "wow 0.69 for ink_id = 1, super impressive, congrats!",
    "2302931": "Looking at our results, it seems that the optimal thresholds for fragment_a and fragment_b are not that different.\nHowever, we were unable to find any evidence to support it during the competition.",
    "2302962": "Congratulations for the great results",
    "2303017": "what is resnet50-irCSN?",
    "2303068": "https://github.com/open-mmlab/mmaction2/blob/main/configs/recognition/csn/README.md",
    "2303163": "tattaka Congratulations on the shake-up to second place! That's an impressive achievement. It seems that using many layers resulted in a more robust model. I also learned the importance of applying strong augmentations.\n\n\n~~~\n1/32 Resolution is sufficient for this task\n~~~\n\nI also had a similar idea, but I didn't try it out. I should have done it from the beginning.\n\nPerhaps, by reducing the resolution, the original text, which was recognized as 1000 pixels x 1000 pixels, became smaller and It might be that it contributed to improved recognition of the characters. This is just my speculation.",
    "2303221": "Before starting the competition, when I read the host's [inkid repo](https://github.com/educelab/ink-id/blob/develop/inkid/scripts/train_and_predict.py#L134), I found that they didn't use upsampling.\nSo when I actually calculated the maximum score that could be achieved by downsampling the label and undoing it, I found that I could get 0.98, so I knew that 1/32 resolution was enough. \nAlso, by lowering the resolution, the difficulty of the task is also lowered, and we believe that we are able to reduce the burden on the machine learning model.",
    "2303779": "Great method of confirmation! I think the benefits of this are incredible. Thank you for sharing.",
    "2303791": "in my experiments, just train a pvtv2-b3 transformer and use its feature map (1/32 scale) gives about 0.65 lb without tunning. validation is fragment one.\n\nbut not all backbone work.",
    "2303914": "Congratulations!",
    "2304587": "tattaka \n\nI have another question:\n\n> 1/32 Resolution is sufficient for this task\n\nDoes this mean you are feeding 2d encoders final layer's activations (typically output_stride is 32) directly to the 3d encoder (i.e. without decoder)?",
    "2304721": "Yes, let me explain in a little more detail.\nThe input of the 2D encoder is (bs * groups, ch, h, w) and the output after going through the neck is (bs * groups, 512, h//32, w//32). Then permute it to (bs, ch, groups, h//32, w//32) and feed it to 3D CNN.",
    "2304894": "congrats for better results",
    "2305926": "Congratulations!!!",
    "2310568": "Congratulations on the success! The end was a photo finish! I hope you will post more details of your solution here or on Github or another similar platform.",
    "2315412": "This is impressive! \nI tried feeding (bs, groups, h, w) to 2D encoder (changing to -> (bs, output channel, h, w)) and subsequently to 3D encoder but this approach has problem to lose temporal information. But using your method, it can be possible to retain temporal information and segmentation information at the same time.\nRespect for your idea!",
    "2317125": "***Update***\nSubmission of tattaka and ron's part was set to public!\n* training code: https://github.com/mipypf/Vesuvius-Challenge/tree/winner_call/tattaka_ron\n* inference code: https://www.kaggle.com/code/mipypf/ink-segmentation-2-5d-3dcnn-resnet3dcsn-fp16fold01?scriptVersionId=132226669",
    "2317273": "Update\nyukke42's training code is also public! \nhttps://github.com/yukke42/kaggle-vesuvius-challenge-ink-detection"
  },
  "source": "meta"
}