{
  "id": 417363,
  "title": "10th place solution",
  "url": "/competitions/vesuvius-challenge-ink-detection/discussion/417363",
  "author_name": "Feng Qilong",
  "post_date": "2023-06-15T12:16:31.555000",
  "votes": 22,
  "comment_count": 0,
  "views": 0,
  "content": "<h3>Summary</h3>\n<ul>\n<li>pretrain 3D encoder using the <a href=\"https://github.com/educelab/ink-id/blob/develop/inkid/model/model.py\" target=\"_blank\">ink-id InkClassifier3DCNN</a> sampled size = 64, stride = 32 (improved CV by 0.03-0.06)</li>\n<li>3D encoder -&gt; pool along z-axis -&gt; 2D FPN Decoder</li>\n<li>Simple self-implemented attention pooling</li>\n<li>Data Sampling -&gt; 1 region containing positive ink : 8 region has no ink at all</li>\n<li>Denoiser</li>\n</ul>\n<h3>Data preparation</h3>\n<ul>\n<li>5 folds, split ink_id = 2 into 3 equal regions along height</li>\n<li>Resolution of 224 , stride of 56, slice 20-36</li>\n<li>Sampled regions with ink and regions without ink using a ratio of 1:8 to generate more data, this also makes training more stable as compared to sampling with moving window</li>\n<li>Augmentation:</li>\n</ul>\n<pre><code>[\n        A.HorizontalFlip(=0.5),\n        A.VerticalFlip(=0.5),\n        A.RandomRotate90(=0.5),\n        A.Affine(=0, =0.1, scale=[0.9,1.5], =0, =0.5),\n        A.OneOf([\n            A.RandomToneCurve(=0.3, =0.2),\n            A.RandomBrightnessContrast(brightness_limit=(-0.1, 0.2), contrast_limit=(-0.4, 0.5), =, =, =0.8)\n        ], =0.5),\n        A.OneOf([\n            A.ShiftScaleRotate(=None, scale_limit=[-0.15, 0.15], rotate_limit=[-30, 30], =cv2.INTER_LINEAR,border_mode=cv2.BORDER_CONSTANT, =0, =None, shift_limit_x=[-0.1, 0.1],shift_limit_y=[-0.2, 0.2], =, =0.5),\n            A.ElasticTransform(=1, =20, =10, =cv2.INTER_LINEAR, =cv2.BORDER_CONSTANT, =0, =None, =, =, =0.5),\n            A.GridDistortion(=5, =0.3, =cv2.INTER_LINEAR, =cv2.BORDER_CONSTANT, =0, =None, =, =0.5),\n        ], =0.5),\n        A.OneOf([\n            A.GaussNoise(var_limit=[10, 50], =0.5),\n            A.GaussianBlur(=0.5),\n            A.MotionBlur(=0.5),\n        ], =0.5),\n        A.CoarseDropout(=3, =0.15, =0.25, =0, =0.5),\n        A.Normalize(\n            mean=[0]*in_chans, \n            std=[1]*in_chans, \n        ),\n        ToTensorV2(=),\n ]\n</code></pre>\n<h3>Model</h3>\n<ul>\n<li><a href=\"https://github.com/kenshohara/3D-ResNets-PyTorch\" target=\"_blank\">resnet3D</a> as backbone</li>\n<li>Attention pooling</li>\n</ul>\n<pre><code> (torch.nn.Module):\n     ():\n        ().__init__()\n        .attention_weights = nn.Parameter(torch.ones(, , depth, height, width))\n        .softmax = nn.Softmax(dim=)\n\n     ():\n        \n        attention_weights = .softmax(.attention_weights)\n        \n        pooled_output = torch.mul(attention_weights, x)\n        \n        pooled_output = torch.sum(pooled_output, dim=)\n         pooled_output\n</code></pre>\n<ul>\n<li>FPN as decoder</li>\n<li>Denoiser, this was inspired by the diffusion model, whereas the model predicts the noise</li>\n</ul>\n<pre><code> cfg.use_denoiser:\n            self.denoiser = smp.Unet(\n                =, #  \n                =,\n                =1,\n                =1,\n                =None,\n            )\n\n self.cfg.use_denoiser:\n            noise = self.denoiser(masks)\n            masks = masks - noise\n</code></pre>\n<ul>\n<li>loss: bce</li>\n</ul>\n<h3>Training</h3>\n<ul>\n<li>amp</li>\n<li>adamw</li>\n</ul>\n<h3>Inference</h3>\n<ul>\n<li>tta: rot90 * 3 + original</li>\n<li>threshold mainly determined by cv, tried to use a percentile of 93 but did not use as I was unsure about the ink composition in private dataset.</li>\n</ul>\n<h3>General pipeline</h3>\n<ul>\n<li>pretrain with ink-id InkClassifier3DCNN -&gt; save encoder -&gt; load into segementation model -&gt; train segmentation model -&gt; inference</li>\n</ul>\n<h3>Result</h3>\n<ul>\n<li>CV: 0.685 </li>\n<li>public LB: 0.69 </li>\n<li>private LB: 0.65</li>\n</ul>\n<h2>Code</h2>\n<ul>\n<li><a href=\"https://www.kaggle.com/code/fengqilong/vesuvius-inference\" target=\"_blank\">Inference</a></li>\n<li><a href=\"https://github.com/fengql123/kaggle-vesuvius-10th-place-solution/tree/main\" target=\"_blank\">Training Code</a></li>\n</ul>",
  "messages": [
    {
      "id": 2303690,
      "postDate": "2023-06-15T12:16:31.557Z",
      "content": "<h3>Summary</h3>\n<ul>\n<li>pretrain 3D encoder using the <a href=\"https://github.com/educelab/ink-id/blob/develop/inkid/model/model.py\" target=\"_blank\">ink-id InkClassifier3DCNN</a> sampled size = 64, stride = 32 (improved CV by 0.03-0.06)</li>\n<li>3D encoder -&gt; pool along z-axis -&gt; 2D FPN Decoder</li>\n<li>Simple self-implemented attention pooling</li>\n<li>Data Sampling -&gt; 1 region containing positive ink : 8 region has no ink at all</li>\n<li>Denoiser</li>\n</ul>\n<h3>Data preparation</h3>\n<ul>\n<li>5 folds, split ink_id = 2 into 3 equal regions along height</li>\n<li>Resolution of 224 , stride of 56, slice 20-36</li>\n<li>Sampled regions with ink and regions without ink using a ratio of 1:8 to generate more data, this also makes training more stable as compared to sampling with moving window</li>\n<li>Augmentation:</li>\n</ul>\n<pre><code>[\n        A.HorizontalFlip(=0.5),\n        A.VerticalFlip(=0.5),\n        A.RandomRotate90(=0.5),\n        A.Affine(=0, =0.1, scale=[0.9,1.5], =0, =0.5),\n        A.OneOf([\n            A.RandomToneCurve(=0.3, =0.2),\n            A.RandomBrightnessContrast(brightness_limit=(-0.1, 0.2), contrast_limit=(-0.4, 0.5), =, =, =0.8)\n        ], =0.5),\n        A.OneOf([\n            A.ShiftScaleRotate(=None, scale_limit=[-0.15, 0.15], rotate_limit=[-30, 30], =cv2.INTER_LINEAR,border_mode=cv2.BORDER_CONSTANT, =0, =None, shift_limit_x=[-0.1, 0.1],shift_limit_y=[-0.2, 0.2], =, =0.5),\n            A.ElasticTransform(=1, =20, =10, =cv2.INTER_LINEAR, =cv2.BORDER_CONSTANT, =0, =None, =, =, =0.5),\n            A.GridDistortion(=5, =0.3, =cv2.INTER_LINEAR, =cv2.BORDER_CONSTANT, =0, =None, =, =0.5),\n        ], =0.5),\n        A.OneOf([\n            A.GaussNoise(var_limit=[10, 50], =0.5),\n            A.GaussianBlur(=0.5),\n            A.MotionBlur(=0.5),\n        ], =0.5),\n        A.CoarseDropout(=3, =0.15, =0.25, =0, =0.5),\n        A.Normalize(\n            mean=[0]*in_chans, \n            std=[1]*in_chans, \n        ),\n        ToTensorV2(=),\n ]\n</code></pre>\n<h3>Model</h3>\n<ul>\n<li><a href=\"https://github.com/kenshohara/3D-ResNets-PyTorch\" target=\"_blank\">resnet3D</a> as backbone</li>\n<li>Attention pooling</li>\n</ul>\n<pre><code> (torch.nn.Module):\n     ():\n        ().__init__()\n        .attention_weights = nn.Parameter(torch.ones(, , depth, height, width))\n        .softmax = nn.Softmax(dim=)\n\n     ():\n        \n        attention_weights = .softmax(.attention_weights)\n        \n        pooled_output = torch.mul(attention_weights, x)\n        \n        pooled_output = torch.sum(pooled_output, dim=)\n         pooled_output\n</code></pre>\n<ul>\n<li>FPN as decoder</li>\n<li>Denoiser, this was inspired by the diffusion model, whereas the model predicts the noise</li>\n</ul>\n<pre><code> cfg.use_denoiser:\n            self.denoiser = smp.Unet(\n                =, #  \n                =,\n                =1,\n                =1,\n                =None,\n            )\n\n self.cfg.use_denoiser:\n            noise = self.denoiser(masks)\n            masks = masks - noise\n</code></pre>\n<ul>\n<li>loss: bce</li>\n</ul>\n<h3>Training</h3>\n<ul>\n<li>amp</li>\n<li>adamw</li>\n</ul>\n<h3>Inference</h3>\n<ul>\n<li>tta: rot90 * 3 + original</li>\n<li>threshold mainly determined by cv, tried to use a percentile of 93 but did not use as I was unsure about the ink composition in private dataset.</li>\n</ul>\n<h3>General pipeline</h3>\n<ul>\n<li>pretrain with ink-id InkClassifier3DCNN -&gt; save encoder -&gt; load into segementation model -&gt; train segmentation model -&gt; inference</li>\n</ul>\n<h3>Result</h3>\n<ul>\n<li>CV: 0.685 </li>\n<li>public LB: 0.69 </li>\n<li>private LB: 0.65</li>\n</ul>\n<h2>Code</h2>\n<ul>\n<li><a href=\"https://www.kaggle.com/code/fengqilong/vesuvius-inference\" target=\"_blank\">Inference</a></li>\n<li><a href=\"https://github.com/fengql123/kaggle-vesuvius-10th-place-solution/tree/main\" target=\"_blank\">Training Code</a></li>\n</ul>",
      "rawMarkdown": "### Summary\n- pretrain 3D encoder using the [ink-id InkClassifier3DCNN](https://github.com/educelab/ink-id/blob/develop/inkid/model/model.py) sampled size = 64, stride = 32 (improved CV by 0.03-0.06)\n- 3D encoder -> pool along z-axis -> 2D FPN Decoder\n- Simple self-implemented attention pooling\n- Data Sampling -> 1 region containing positive ink : 8 region has no ink at all\n- Denoiser\n\n### Data preparation\n- 5 folds, split ink_id = 2 into 3 equal regions along height\n- Resolution of 224 , stride of 56, slice 20-36\n- Sampled regions with ink and regions without ink using a ratio of 1:8 to generate more data, this also makes training more stable as compared to sampling with moving window\n- Augmentation:\n```\n[\n        A.HorizontalFlip(p=0.5),\n        A.VerticalFlip(p=0.5),\n        A.RandomRotate90(p=0.5),\n        A.Affine(rotate=0, translate_percent=0.1, scale=[0.9,1.5], shear=0, p=0.5),\n        A.OneOf([\n            A.RandomToneCurve(scale=0.3, p=0.2),\n            A.RandomBrightnessContrast(brightness_limit=(-0.1, 0.2), contrast_limit=(-0.4, 0.5), brightness_by_max=True, always_apply=False, p=0.8)\n        ], p=0.5),\n        A.OneOf([\n            A.ShiftScaleRotate(shift_limit=None, scale_limit=[-0.15, 0.15], rotate_limit=[-30, 30], interpolation=cv2.INTER_LINEAR,border_mode=cv2.BORDER_CONSTANT, value=0, mask_value=None, shift_limit_x=[-0.1, 0.1],shift_limit_y=[-0.2, 0.2], rotate_method='largest_box', p=0.5),\n            A.ElasticTransform(alpha=1, sigma=20, alpha_affine=10, interpolation=cv2.INTER_LINEAR, border_mode=cv2.BORDER_CONSTANT, value=0, mask_value=None, approximate=False, same_dxdy=False, p=0.5),\n            A.GridDistortion(num_steps=5, distort_limit=0.3, interpolation=cv2.INTER_LINEAR, border_mode=cv2.BORDER_CONSTANT, value=0, mask_value=None, normalized=True, p=0.5),\n        ], p=0.5),\n        A.OneOf([\n            A.GaussNoise(var_limit=[10, 50], p=0.5),\n            A.GaussianBlur(p=0.5),\n            A.MotionBlur(p=0.5),\n        ], p=0.5),\n        A.CoarseDropout(max_holes=3, max_width=0.15, max_height=0.25, mask_fill_value=0, p=0.5),\n        A.Normalize(\n            mean=[0]*in_chans, \n            std=[1]*in_chans, \n        ),\n        ToTensorV2(transpose_mask=True),\n ]\n```\n### Model\n- [resnet3D](https://github.com/kenshohara/3D-ResNets-PyTorch) as backbone\n- Attention pooling\n\n```\nclass AttentionPool(torch.nn.Module):\n    def __init__(self, depth, height, width):\n        super().__init__()\n        self.attention_weights = nn.Parameter(torch.ones(1, 1, depth, height, width))\n        self.softmax = nn.Softmax(dim=2)\n\n    def forward(self, x):\n        # Apply softmax along the depth dimension to obtain attention weights\n        attention_weights = self.softmax(self.attention_weights)\n        # Perform attention pooling by multiplying the attention weights with the input tensor\n        pooled_output = torch.mul(attention_weights, x)\n        # Sum the pooled output along the depth dimension\n        pooled_output = torch.sum(pooled_output, dim=2)\n        return pooled_output\n```\n- FPN as decoder\n- Denoiser, this was inspired by the diffusion model, whereas the model predicts the noise\n```\nif cfg.use_denoiser:\n            self.denoiser = smp.Unet(\n                encoder_name=\"tu-resnet10t\", # \"tu-resnet10t\" \"resnet18\"\n                encoder_weights=\"imagenet\",\n                in_channels=1,\n                classes=1,\n                activation=None,\n            )\n# inference\nif self.cfg.use_denoiser:\n            noise = self.denoiser(masks)\n            masks = masks - noise\n```\n- loss: bce\n\n### Training\n- amp\n- adamw\n\n### Inference\n- tta: rot90 * 3 + original\n- threshold mainly determined by cv, tried to use a percentile of 93 but did not use as I was unsure about the ink composition in private dataset.\n\n### General pipeline\n- pretrain with ink-id InkClassifier3DCNN -> save encoder -> load into segementation model -> train segmentation model -> inference\n\n### Result\n- CV: 0.685 \n- public LB: 0.69 \n- private LB: 0.65\n\n\n## Code\n- [Inference](https://www.kaggle.com/code/fengqilong/vesuvius-inference)\n- [Training Code](https://github.com/fengql123/kaggle-vesuvius-10th-place-solution/tree/main)",
      "votes": 22
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "2303690": "### Summary\n- pretrain 3D encoder using the [ink-id InkClassifier3DCNN](https://github.com/educelab/ink-id/blob/develop/inkid/model/model.py) sampled size = 64, stride = 32 (improved CV by 0.03-0.06)\n- 3D encoder -> pool along z-axis -> 2D FPN Decoder\n- Simple self-implemented attention pooling\n- Data Sampling -> 1 region containing positive ink : 8 region has no ink at all\n- Denoiser\n\n### Data preparation\n- 5 folds, split ink_id = 2 into 3 equal regions along height\n- Resolution of 224 , stride of 56, slice 20-36\n- Sampled regions with ink and regions without ink using a ratio of 1:8 to generate more data, this also makes training more stable as compared to sampling with moving window\n- Augmentation:\n```\n[\n        A.HorizontalFlip(p=0.5),\n        A.VerticalFlip(p=0.5),\n        A.RandomRotate90(p=0.5),\n        A.Affine(rotate=0, translate_percent=0.1, scale=[0.9,1.5], shear=0, p=0.5),\n        A.OneOf([\n            A.RandomToneCurve(scale=0.3, p=0.2),\n            A.RandomBrightnessContrast(brightness_limit=(-0.1, 0.2), contrast_limit=(-0.4, 0.5), brightness_by_max=True, always_apply=False, p=0.8)\n        ], p=0.5),\n        A.OneOf([\n            A.ShiftScaleRotate(shift_limit=None, scale_limit=[-0.15, 0.15], rotate_limit=[-30, 30], interpolation=cv2.INTER_LINEAR,border_mode=cv2.BORDER_CONSTANT, value=0, mask_value=None, shift_limit_x=[-0.1, 0.1],shift_limit_y=[-0.2, 0.2], rotate_method='largest_box', p=0.5),\n            A.ElasticTransform(alpha=1, sigma=20, alpha_affine=10, interpolation=cv2.INTER_LINEAR, border_mode=cv2.BORDER_CONSTANT, value=0, mask_value=None, approximate=False, same_dxdy=False, p=0.5),\n            A.GridDistortion(num_steps=5, distort_limit=0.3, interpolation=cv2.INTER_LINEAR, border_mode=cv2.BORDER_CONSTANT, value=0, mask_value=None, normalized=True, p=0.5),\n        ], p=0.5),\n        A.OneOf([\n            A.GaussNoise(var_limit=[10, 50], p=0.5),\n            A.GaussianBlur(p=0.5),\n            A.MotionBlur(p=0.5),\n        ], p=0.5),\n        A.CoarseDropout(max_holes=3, max_width=0.15, max_height=0.25, mask_fill_value=0, p=0.5),\n        A.Normalize(\n            mean=[0]*in_chans, \n            std=[1]*in_chans, \n        ),\n        ToTensorV2(transpose_mask=True),\n ]\n```\n### Model\n- [resnet3D](https://github.com/kenshohara/3D-ResNets-PyTorch) as backbone\n- Attention pooling\n\n```\nclass AttentionPool(torch.nn.Module):\n    def __init__(self, depth, height, width):\n        super().__init__()\n        self.attention_weights = nn.Parameter(torch.ones(1, 1, depth, height, width))\n        self.softmax = nn.Softmax(dim=2)\n\n    def forward(self, x):\n        # Apply softmax along the depth dimension to obtain attention weights\n        attention_weights = self.softmax(self.attention_weights)\n        # Perform attention pooling by multiplying the attention weights with the input tensor\n        pooled_output = torch.mul(attention_weights, x)\n        # Sum the pooled output along the depth dimension\n        pooled_output = torch.sum(pooled_output, dim=2)\n        return pooled_output\n```\n- FPN as decoder\n- Denoiser, this was inspired by the diffusion model, whereas the model predicts the noise\n```\nif cfg.use_denoiser:\n            self.denoiser = smp.Unet(\n                encoder_name=\"tu-resnet10t\", # \"tu-resnet10t\" \"resnet18\"\n                encoder_weights=\"imagenet\",\n                in_channels=1,\n                classes=1,\n                activation=None,\n            )\n# inference\nif self.cfg.use_denoiser:\n            noise = self.denoiser(masks)\n            masks = masks - noise\n```\n- loss: bce\n\n### Training\n- amp\n- adamw\n\n### Inference\n- tta: rot90 * 3 + original\n- threshold mainly determined by cv, tried to use a percentile of 93 but did not use as I was unsure about the ink composition in private dataset.\n\n### General pipeline\n- pretrain with ink-id InkClassifier3DCNN -> save encoder -> load into segementation model -> train segmentation model -> inference\n\n### Result\n- CV: 0.685 \n- public LB: 0.69 \n- private LB: 0.65\n\n\n## Code\n- [Inference](https://www.kaggle.com/code/fengqilong/vesuvius-inference)\n- [Training Code](https://github.com/fengql123/kaggle-vesuvius-10th-place-solution/tree/main)"
  }
}