{
  "id": 418921,
  "title": "12th place solution",
  "url": "/competitions/vesuvius-challenge-ink-detection/writeups/one-piece-12th-place-solution",
  "author_name": "",
  "post_date": "2023-06-23T08:20:07.487Z",
  "votes": 7,
  "comment_count": 2,
  "views": 0,
  "content": "<h1><strong>12th place solution</strong></h1>\n<p>Thank you to the organizers for hosting this exciting competition and congratulations to all the winners! This is my first time participating in a competition like this, and below is my solution and approach.</p>\n<h2><strong>overview</strong></h2>\n<p>we use 3D Resnet(<a href=\"https://github.com/kenshohara/3D-ResNets-PyTorch/blob/master/models/resnet.py\" target=\"_blank\">https://github.com/kenshohara/3D-ResNets-PyTorch/blob/master/models/resnet.py</a>) for our encoder and CNN decoder.</p>\n<p>We tried different decoders, such as Unet, Uperhead, CNN, and ResCNN. In the end, the CNN decoder, while being the simplest, achieved the best results.</p>\n<h2><strong>model architecture</strong></h2>\n<table>\n<thead>\n<tr>\n<th>backbone</th>\n<th>decoder</th>\n<th>Z-DIM</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>3DResnet-18</td>\n<td>CNN</td>\n<td>22</td>\n</tr>\n<tr>\n<td>3DResnet-34</td>\n<td>CNN</td>\n<td>22</td>\n</tr>\n</tbody>\n</table>\n<h2><strong>training</strong></h2>\n<p>Z_DIM:21-&gt;43<br>\nimg_size:224<br>\nstride_size:56</p>\n<h3>augmentation:</h3>\n<p>A.Rotate(limit=90, p=0.5),<br>\nA.HorizontalFlip(p=0.5),<br>\nA.VerticalFlip(p=0.5),<br>\nA.RandomBrightnessContrast(p=0.75),<br>\nA.ShiftScaleRotate(p=0.75),<br>\nA.OneOf([<br>\n        A.GaussNoise(),<br>\n        A.GaussianBlur(),<br>\n        A.MotionBlur(),<br>\n        ], p=0.5),<br>\nA.GridDistortion(num_steps=5, distort_limit=0.3, p=0.5),<br>\nA.CoarseDropout(max_holes=1, max_width=int(size * 0.3), max_height=int(size * 0.3), <br>\n                  mask_fill_value=0, p=0.5),<br>\nA.Normalize(<br>\n    mean= [0] * in_chans,<br>\n    std= [1] * in_chans<br>\n),<br>\nToTensorV2(transpose_mask=True)</p>\n<h3>training details</h3>\n<p>We used two different training methods. The first method involved selecting one folder from 1, 2, and 3 as the validation set and training three models. However, due to poor performance on validation set 2, we only used models trained with validation sets 1 and 3 for our 3DRsnet-18 model. The second method involved randomly selecting image blocks from 1, 2, and 3 as the validation set. We used this method to train 4 models: two based on 3DResnet-18 with a stride of 56 and 37, and another two based on 3DResnet-34 with a stride of 56 and 37.</p>\n<h2>others</h2>\n<p>BCELoss<br>\nAdamw<br>\nGradualWarmupSchedulerV2(<a href=\"https://www.kaggle.com/code/underwearfitting/single-fold-training-of-resnet200d-lb0-965\" target=\"_blank\">https://www.kaggle.com/code/underwearfitting/single-fold-training-of-resnet200d-lb0-965</a>)</p>\n<h2><strong>inference</strong></h2>\n<p>img_size:224<br>\nstride_size:224<br>\nth:0.5<br>\nTTA: 4 rotates, h/v flips<br>\nlast model:<br>\n(3DResnet-18 + CNN) * 4 + (3DResnet-34 ＋CNN) * 2<br>\npublic:0.786897<br>\nprivate:0.652892</p>\n<p>training code:<a href=\"https://github.com/tkz24589/Deep_Learning_work_space/tree/main/convolution/segmentation/vesuvius-challenge-ink-detection-tutorial\" target=\"_blank\">https://github.com/tkz24589/Deep_Learning_work_space/tree/main/convolution/segmentation/vesuvius-challenge-ink-detection-tutorial</a><br>\nnotebook example: <a href=\"https://www.kaggle.com/code/kongzhangtang/resnet3d-cnn\" target=\"_blank\">https://www.kaggle.com/code/kongzhangtang/resnet3d-cnn</a></p>",
  "messages": [
    {
      "id": "2314145",
      "postDate": "06/23/2023 07:53:02",
      "content": "<h1><strong>12th place solution</strong></h1>\n<p>Thank you to the organizers for hosting this exciting competition and congratulations to all the winners! This is my first time participating in a competition like this, and below is my solution and approach.</p>\n<h2><strong>overview</strong></h2>\n<p>we use 3D Resnet(<a href=\"https://github.com/kenshohara/3D-ResNets-PyTorch/blob/master/models/resnet.py\" target=\"_blank\">https://github.com/kenshohara/3D-ResNets-PyTorch/blob/master/models/resnet.py</a>) for our encoder and CNN decoder.</p>\n<p>We tried different decoders, such as Unet, Uperhead, CNN, and ResCNN. In the end, the CNN decoder, while being the simplest, achieved the best results.</p>\n<h2><strong>model architecture</strong></h2>\n<table>\n<thead>\n<tr>\n<th>backbone</th>\n<th>decoder</th>\n<th>Z-DIM</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>3DResnet-18</td>\n<td>CNN</td>\n<td>22</td>\n</tr>\n<tr>\n<td>3DResnet-34</td>\n<td>CNN</td>\n<td>22</td>\n</tr>\n</tbody>\n</table>\n<h2><strong>training</strong></h2>\n<p>Z_DIM:21-&gt;43<br>\nimg_size:224<br>\nstride_size:56</p>\n<h3>augmentation:</h3>\n<p>A.Rotate(limit=90, p=0.5),<br>\nA.HorizontalFlip(p=0.5),<br>\nA.VerticalFlip(p=0.5),<br>\nA.RandomBrightnessContrast(p=0.75),<br>\nA.ShiftScaleRotate(p=0.75),<br>\nA.OneOf([<br>\n        A.GaussNoise(),<br>\n        A.GaussianBlur(),<br>\n        A.MotionBlur(),<br>\n        ], p=0.5),<br>\nA.GridDistortion(num_steps=5, distort_limit=0.3, p=0.5),<br>\nA.CoarseDropout(max_holes=1, max_width=int(size * 0.3), max_height=int(size * 0.3), <br>\n                  mask_fill_value=0, p=0.5),<br>\nA.Normalize(<br>\n    mean= [0] * in_chans,<br>\n    std= [1] * in_chans<br>\n),<br>\nToTensorV2(transpose_mask=True)</p>\n<h3>training details</h3>\n<p>We used two different training methods. The first method involved selecting one folder from 1, 2, and 3 as the validation set and training three models. However, due to poor performance on validation set 2, we only used models trained with validation sets 1 and 3 for our 3DRsnet-18 model. The second method involved randomly selecting image blocks from 1, 2, and 3 as the validation set. We used this method to train 4 models: two based on 3DResnet-18 with a stride of 56 and 37, and another two based on 3DResnet-34 with a stride of 56 and 37.</p>\n<h2>others</h2>\n<p>BCELoss<br>\nAdamw<br>\nGradualWarmupSchedulerV2(<a href=\"https://www.kaggle.com/code/underwearfitting/single-fold-training-of-resnet200d-lb0-965\" target=\"_blank\">https://www.kaggle.com/code/underwearfitting/single-fold-training-of-resnet200d-lb0-965</a>)</p>\n<h2><strong>inference</strong></h2>\n<p>img_size:224<br>\nstride_size:224<br>\nth:0.5<br>\nTTA: 4 rotates, h/v flips<br>\nlast model:<br>\n(3DResnet-18 + CNN) * 4 + (3DResnet-34 ＋CNN) * 2<br>\npublic:0.786897<br>\nprivate:0.652892</p>\n<p>training code:<a href=\"https://github.com/tkz24589/Deep_Learning_work_space/tree/main/convolution/segmentation/vesuvius-challenge-ink-detection-tutorial\" target=\"_blank\">https://github.com/tkz24589/Deep_Learning_work_space/tree/main/convolution/segmentation/vesuvius-challenge-ink-detection-tutorial</a><br>\nnotebook example: <a href=\"https://www.kaggle.com/code/kongzhangtang/resnet3d-cnn\" target=\"_blank\">https://www.kaggle.com/code/kongzhangtang/resnet3d-cnn</a></p>",
      "rawMarkdown": "# **12th place solution**\nThank you to the organizers for hosting this exciting competition and congratulations to all the winners! This is my first time participating in a competition like this, and below is my solution and approach.\n## **overview**\nwe use 3D Resnet(https://github.com/kenshohara/3D-ResNets-PyTorch/blob/master/models/resnet.py) for our encoder and CNN decoder.\n\nWe tried different decoders, such as Unet, Uperhead, CNN, and ResCNN. In the end, the CNN decoder, while being the simplest, achieved the best results.\n##**model architecture**\n| backbone |decoder| Z-DIM |\n| --- | --- | --- |\n| 3DResnet-18 |CNN| 22 |\n| 3DResnet-34 |CNN| 22 |\n\n## **training**\n\nZ_DIM:21->43\nimg_size:224\nstride_size:56\n\n###augmentation:\nA.Rotate(limit=90, p=0.5),\nA.HorizontalFlip(p=0.5),\nA.VerticalFlip(p=0.5),\nA.RandomBrightnessContrast(p=0.75),\nA.ShiftScaleRotate(p=0.75),\nA.OneOf([\n        A.GaussNoise(),\n        A.GaussianBlur(),\n        A.MotionBlur(),\n        ], p=0.5),\nA.GridDistortion(num_steps=5, distort_limit=0.3, p=0.5),\nA.CoarseDropout(max_holes=1, max_width=int(size * 0.3), max_height=int(size * 0.3), \n                  mask_fill_value=0, p=0.5),\nA.Normalize(\n    mean= [0] * in_chans,\n    std= [1] * in_chans\n),\nToTensorV2(transpose_mask=True)\n\n###training details\nWe used two different training methods. The first method involved selecting one folder from 1, 2, and 3 as the validation set and training three models. However, due to poor performance on validation set 2, we only used models trained with validation sets 1 and 3 for our 3DRsnet-18 model. The second method involved randomly selecting image blocks from 1, 2, and 3 as the validation set. We used this method to train 4 models: two based on 3DResnet-18 with a stride of 56 and 37, and another two based on 3DResnet-34 with a stride of 56 and 37.\n\n##others\nBCELoss\nAdamw\nGradualWarmupSchedulerV2(https://www.kaggle.com/code/underwearfitting/single-fold-training-of-resnet200d-lb0-965)\n\n## **inference**\nimg_size:224\nstride_size:224\nth:0.5\nTTA: 4 rotates, h/v flips\nlast model:\n(3DResnet-18 + CNN) * 4 + (3DResnet-34 ＋CNN) * 2\npublic:0.786897\nprivate:0.652892\n\ntraining code:https://github.com/tkz24589/Deep_Learning_work_space/tree/main/convolution/segmentation/vesuvius-challenge-ink-detection-tutorial\nnotebook example: https://www.kaggle.com/code/kongzhangtang/resnet3d-cnn",
      "votes": null
    },
    {
      "id": "2319678",
      "postDate": "06/27/2023 08:50:39",
      "content": "<p>Congratulations on your gold medal on your first competition!  It looks like you used a very sound approach to trying a number of options and figuring out what worked best.  Thank you for sharing your summary and your code.</p>",
      "rawMarkdown": "Congratulations on your gold medal on your first competition!  It looks like you used a very sound approach to trying a number of options and figuring out what worked best.  Thank you for sharing your summary and your code.",
      "votes": null
    },
    {
      "id": "2319847",
      "postDate": "06/27/2023 10:36:05",
      "content": "<p>Thank you, and congratulations on winning first place!</p>",
      "rawMarkdown": "Thank you, and congratulations on winning first place!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2319678,
      "author_name": "socated",
      "author_url": "",
      "post_date": "06/27/2023 08:50:39",
      "content": "<p>Congratulations on your gold medal on your first competition!  It looks like you used a very sound approach to trying a number of options and figuring out what worked best.  Thank you for sharing your summary and your code.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2319847,
          "author_name": "kongzhangtang",
          "author_url": "",
          "post_date": "06/27/2023 10:36:05",
          "content": "<p>Thank you, and congratulations on winning first place!</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2314145": "# **12th place solution**\nThank you to the organizers for hosting this exciting competition and congratulations to all the winners! This is my first time participating in a competition like this, and below is my solution and approach.\n## **overview**\nwe use 3D Resnet(https://github.com/kenshohara/3D-ResNets-PyTorch/blob/master/models/resnet.py) for our encoder and CNN decoder.\n\nWe tried different decoders, such as Unet, Uperhead, CNN, and ResCNN. In the end, the CNN decoder, while being the simplest, achieved the best results.\n##**model architecture**\n| backbone |decoder| Z-DIM |\n| --- | --- | --- |\n| 3DResnet-18 |CNN| 22 |\n| 3DResnet-34 |CNN| 22 |\n\n## **training**\n\nZ_DIM:21->43\nimg_size:224\nstride_size:56\n\n###augmentation:\nA.Rotate(limit=90, p=0.5),\nA.HorizontalFlip(p=0.5),\nA.VerticalFlip(p=0.5),\nA.RandomBrightnessContrast(p=0.75),\nA.ShiftScaleRotate(p=0.75),\nA.OneOf([\n        A.GaussNoise(),\n        A.GaussianBlur(),\n        A.MotionBlur(),\n        ], p=0.5),\nA.GridDistortion(num_steps=5, distort_limit=0.3, p=0.5),\nA.CoarseDropout(max_holes=1, max_width=int(size * 0.3), max_height=int(size * 0.3), \n                  mask_fill_value=0, p=0.5),\nA.Normalize(\n    mean= [0] * in_chans,\n    std= [1] * in_chans\n),\nToTensorV2(transpose_mask=True)\n\n###training details\nWe used two different training methods. The first method involved selecting one folder from 1, 2, and 3 as the validation set and training three models. However, due to poor performance on validation set 2, we only used models trained with validation sets 1 and 3 for our 3DRsnet-18 model. The second method involved randomly selecting image blocks from 1, 2, and 3 as the validation set. We used this method to train 4 models: two based on 3DResnet-18 with a stride of 56 and 37, and another two based on 3DResnet-34 with a stride of 56 and 37.\n\n##others\nBCELoss\nAdamw\nGradualWarmupSchedulerV2(https://www.kaggle.com/code/underwearfitting/single-fold-training-of-resnet200d-lb0-965)\n\n## **inference**\nimg_size:224\nstride_size:224\nth:0.5\nTTA: 4 rotates, h/v flips\nlast model:\n(3DResnet-18 + CNN) * 4 + (3DResnet-34 ＋CNN) * 2\npublic:0.786897\nprivate:0.652892\n\ntraining code:https://github.com/tkz24589/Deep_Learning_work_space/tree/main/convolution/segmentation/vesuvius-challenge-ink-detection-tutorial\nnotebook example: https://www.kaggle.com/code/kongzhangtang/resnet3d-cnn",
    "2319678": "Congratulations on your gold medal on your first competition!  It looks like you used a very sound approach to trying a number of options and figuring out what worked best.  Thank you for sharing your summary and your code.",
    "2319847": "Thank you, and congratulations on winning first place!"
  },
  "source": "meta"
}