{
  "id": 417779,
  "title": "4th place solution",
  "url": "/competitions/vesuvius-challenge-ink-detection/writeups/posco-dx-heeyoung-ahn-4th-place-solution",
  "author_name": "",
  "post_date": "2023-07-13T14:37:18.687Z",
  "votes": 22,
  "comment_count": 2,
  "views": 0,
  "content": "<h1>Acknowledgement</h1>\n<p>I'm very grateful for organizing very challenging and fantastic challenge.<br>\nI feel like there are a lot of really great people and I respect them all.<br>\nI would like to express my gratitude to both the organizers and participants of the competition.</p>\n<p><br><br>\n<br></p>\n<h1>Solution</h1>\n<p>I would like to describe my solution in the order of contribution of performance improvement.<br>\n<strong>The number of ★ in the description below means the degree of contribution to performance improvement.</strong></p>\n<p><br><br>\n<br></p>\n<h3>1. Temporal random crop &amp; random paste &amp; random cutout (★★★★★)</h3>\n<h4>1.1. Thinking about the data itself</h4>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5725749%2Fb9819315244312a47802b241517dc57e%2F4.PNG?generation=1686971322910074&amp;alt=media\" alt=\"\"></p>\n<p>As you can see in the figure above, i thought that particular layers of fragment would not correspond to the same layers of other fragment. Using these points, i came up with an augmentation method that can give strong regularization to the model by reflecting the characteristics of the data.</p>\n<h4>1.2. Application</h4>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5725749%2F8601353d50ccdc56721bfc1341395bff%2F8.PNG?generation=1686981811287731&amp;alt=media\" alt=\"\"></p>\n<p>As you can see in the figure above, there are 3 steps.<br>\n1) temporal random crop</p>\n<ul>\n<li>First, I decided to use a total of 22 layers (21-42) out of 65 layers. Of these 22 layers, layers with a range of cropping_min to cropping_max are cropped randomly. After several ablation study, I set cropping_min = 12, cropping_max = 22.</li>\n</ul>\n<p>2) random paste</p>\n<ul>\n<li>The randomly cropped layers are attached to a random area of 22 while maintaining sequential information.<br>\n<strong>This crop &amp; paste method allows the model to learn wide, various range of layers rather than a fixed area of fragments, allowing generalization to learning various fragments.</strong></li>\n</ul>\n<p>3) random cutout</p>\n<ul>\n<li>Similar to spatial cutout augmentation, I applied temporal cutout augmentation. Among the pasted layers, 0 to 2 random layers are filled with 0 values.</li>\n</ul>\n<p>The above three processes can be seen simply by looking at the Pytorch Dataset code part below.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5725749%2F2cb3e87617e91bbff8df5dbdf213838e%2F9.PNG?generation=1686982042582252&amp;alt=media\" alt=\"\"></p>\n<p><br><br>\n<br></p>\n<h3>2. Weight choice with low false positive (★★★★)</h3>\n<p>As shown in the figure below, I saved the predicted mask for each epoch. The file name has the epoch, score, tp, fp, and fn, of which the fp value was important. <strong>Even if the score was similar, if the fp was large, it tended to be very bad in the public score.</strong> Therefore, I wanted to select a weight with a small fp value and a high score value, and I selected an appropriate fp value for each fold through several submissions. In addition, I tried to increase the generalization performance by ensemble models with a range of fp values.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5725749%2Faa46536cf349929607790d1c3a32a21c%2F55.PNG?generation=1687062720342906&amp;alt=media\" alt=\"\"></p>\n<p><br><br>\n<br></p>\n<h3>3. Used models &amp; ensemble (★★★★)</h3>\n<p>I ensembles 3 models.</p>\n<ul>\n<li><p>3D resnet152, 3D resnet200, 3D resnext101 with unet-like decoder.</p></li>\n<li><p>Refered models was firstly inspired by JEBASTIN NADAR (<a href=\"url\" target=\"_blank\">https://www.kaggle.com/code/samfc10/vesuvius-challenge-3d-resnet-training</a>) and applied. Thanks for your effort and respect you. <br>\n3D resnet152, 3D resnet200 is from <a href=\"https://github.com/kenshohara/3D-ResNets-PyTorch\" target=\"_blank\">https://github.com/kenshohara/3D-ResNets-PyTorch</a><br>\n3D resnext101 is from <a href=\"https://github.com/okankop/Efficient-3DCNNs\" target=\"_blank\">https://github.com/okankop/Efficient-3DCNNs</a><br>\nAll models have mit licenses and are not against the rules.</p></li>\n<li><p>3D Resnet was pretrained on Kinetic 710, and 3D resnext101 was pretrained on Kinetic 600.<br>\n<strong>There was a huge difference between being trained in kinetics and not being able to.</strong> Therefore, during the competition, I tried to find a model trained in the kinetic 600 or 700.</p></li>\n<li><p>All models are combined unet-like simple decoder. Features from 3D encoder are upsampled and concatenated using decoder.<br>\ndetails are in code.</p></li>\n</ul>\n<p><br><br>\n<br></p>\n<h3>4. Label smoothing(0.3), cutmix augmentation, data clipping (★★★)</h3>\n<ul>\n<li>Cross entropy with label smoothing improved performance. After doing ablation study, i set the parameter of label smoothing to 0.3</li>\n<li>Cutmix augmentation improved performance in cross validation &amp; public score</li>\n<li>Inspired by AJLAND(<a href=\"url\" target=\"_blank\">https://www.kaggle.com/code/ajland/eda-a-slice-by-slice-analysis</a>), I clipped the image below 50 and more than 200. Also thanks for your effort and respect you.<br>\n(image = np.clip(image, 50, 200))</li>\n</ul>\n<p><br><br>\n<br></p>\n<h3>5. Belief for local cross validation (★★)</h3>\n<p>As shown in the table below, the two models are the ones I submitted finally.<br>\n<strong>Ensemble 2 had a lower public score than ensemble1, but performed better in cross validation.</strong> There were other models that could be adopted, but considering that it came out well in the cross validation, it was adopted, and <strong>the performance was improved more in the private, resulting in a better first place.</strong></p>\n<table>\n<thead>\n<tr>\n<th>models</th>\n<th>cv score</th>\n<th>threshold</th>\n<th>public score</th>\n<th>private score</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Ensemble 1</td>\n<td>0.66 / 0.762/ 0.723 / 0.68</td>\n<td>0.5</td>\n<td>0.795835</td>\n<td>0.663542</td>\n</tr>\n<tr>\n<td>Ensemble 2</td>\n<td>0.67 / 0.764 / 0.724 / 0.69</td>\n<td>0.47</td>\n<td>0.789024</td>\n<td>0.674544</td>\n</tr>\n</tbody>\n</table>\n<p><br><br>\n<br></p>\n<h3>6. Other training details (★)</h3>\n<p><br></p>\n<h4>6.1. loss, optimizer, scheduler</h4>\n<ul>\n<li>loss : CrossEntropyLoss with label smoothing 0.3</li>\n<li>optimizer = AdamW(lr=1e-4)</li>\n<li>scheduler = cosine annealing with warmup</li>\n</ul>\n<p><br></p>\n<h4>6.2. cross validation strategy &amp; result</h4>\n<p>I trained the model with 4fold</p>\n<ul>\n<li>fold1 : fragment 1</li>\n<li>fold3 : fragment 3</li>\n<li>fold2, 4 : fragment 2 divided into 2 sub-frag(9506 x 14830 -&gt; (4300 x 14830) &amp; (5206 x 14830))</li>\n</ul>\n<p>The results are shown in the table below.<br>\nEnsemble means ensemble of 3 models with same weight.</p>\n<table>\n<thead>\n<tr>\n<th>models</th>\n<th>cv score</th>\n<th>threshold</th>\n<th>public score</th>\n<th>private score</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>3D Resnet152</td>\n<td>0.64 / 0.71 / 0.71 / 0.69</td>\n<td>0.5</td>\n<td>0.78</td>\n<td>-</td>\n</tr>\n<tr>\n<td>3D Resnet200</td>\n<td>0.66 / 0.71 / 0.69 / 0.66</td>\n<td>0.5</td>\n<td>0.77</td>\n<td>-</td>\n</tr>\n<tr>\n<td>3D Resnext101</td>\n<td>0.61 / 0.72 / 0.71 / 0.64</td>\n<td>0.5</td>\n<td>0.77</td>\n<td>-</td>\n</tr>\n<tr>\n<td>Ensemble</td>\n<td>0.66 / 0.762/ 0.723 / 0.68</td>\n<td>0.5</td>\n<td>0.795835</td>\n<td>0.663542</td>\n</tr>\n<tr>\n<td>Ensemble</td>\n<td>0.67 / 0.764 / 0.724 / 0.69</td>\n<td>0.47</td>\n<td>0.789024</td>\n<td>0.674544</td>\n</tr>\n</tbody>\n</table>\n<p><br></p>\n<h4>6.3. image augmentation</h4>\n<p><code>all probability set to 0.6</code></p>\n<ul>\n<li>Horizontal, Vertical Flip</li>\n<li>RandomGamma(limit=(50, 150))</li>\n<li>RandomBrightnessContrast(brightness=0.2, contrast=0.2)</li>\n<li>Gaussian Noise(10, 30), Gaussian blur</li>\n<li>shift 0.1, scale 0.1, rotate 360</li>\n<li>coarsedropout(holes=4, size=0.2 * image_size)</li>\n</ul>\n<h4>6.4 Data Preparation</h4>\n<ul>\n<li>stride rate 3</li>\n<li>image size 256</li>\n</ul>\n<h4>6.5 etc</h4>\n<ul>\n<li>To achieve consistent and stable results, I selected the best approach based on threshold 0.5.</li>\n<li>inference stride is (256 / 5)</li>\n</ul>\n<p><br><br>\n<br></p>\n<h3>7. Tried but not worked</h3>\n<ul>\n<li><p>mixup augmentation</p></li>\n<li><p>temporal channel shuffle augmentation</p></li>\n<li><p>3D transformer encoder(uniformerv2, video swin transformer) - cannot training… i don't know why loss didn't decrease.</p></li>\n<li><p>spatial TTA</p></li>\n<li><p>temporal TTA - because of my crop&amp;paste method mentioned above, i tried temporal TTA. In other words, I tried tta by cropping 22, 20, and 18 layers in various ways and then averaging 3 predictions, sometimes performance increased and sometimes decreased. I didn't apply it because it took too much time.</p></li>\n<li><p>other losses(tversky focal loss, fbeta loss)</p></li>\n</ul>\n<p><br><br>\n<br></p>\n<h3>8. Code</h3>\n<p>8.1. Github version(based on Docker)</p>\n<ul>\n<li><a href=\"https://github.com/AhnHeeYoung/Competition/tree/master/kaggle\" target=\"_blank\">https://github.com/AhnHeeYoung/Competition/tree/master/kaggle</a></li>\n</ul>\n<p>8.2. Kaggle Notebook version</p>\n<ul>\n<li>Training : <a href=\"https://www.kaggle.com/code/ahnheeyoung1/ink-detection-training\" target=\"_blank\">https://www.kaggle.com/code/ahnheeyoung1/ink-detection-training</a></li>\n<li>Inference : <a href=\"https://www.kaggle.com/code/ahnheeyoung1/ink-detection-inference/notebook?scriptVersionId=136610637\" target=\"_blank\">https://www.kaggle.com/code/ahnheeyoung1/ink-detection-inference/notebook?scriptVersionId=136610637</a></li>\n</ul>",
  "messages": [
    {
      "id": "2306168",
      "postDate": "06/17/2023 06:08:54",
      "content": "<h1>Acknowledgement</h1>\n<p>I'm very grateful for organizing very challenging and fantastic challenge.<br>\nI feel like there are a lot of really great people and I respect them all.<br>\nI would like to express my gratitude to both the organizers and participants of the competition.</p>\n<p><br><br>\n<br></p>\n<h1>Solution</h1>\n<p>I would like to describe my solution in the order of contribution of performance improvement.<br>\n<strong>The number of ★ in the description below means the degree of contribution to performance improvement.</strong></p>\n<p><br><br>\n<br></p>\n<h3>1. Temporal random crop &amp; random paste &amp; random cutout (★★★★★)</h3>\n<h4>1.1. Thinking about the data itself</h4>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5725749%2Fb9819315244312a47802b241517dc57e%2F4.PNG?generation=1686971322910074&amp;alt=media\" alt=\"\"></p>\n<p>As you can see in the figure above, i thought that particular layers of fragment would not correspond to the same layers of other fragment. Using these points, i came up with an augmentation method that can give strong regularization to the model by reflecting the characteristics of the data.</p>\n<h4>1.2. Application</h4>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5725749%2F8601353d50ccdc56721bfc1341395bff%2F8.PNG?generation=1686981811287731&amp;alt=media\" alt=\"\"></p>\n<p>As you can see in the figure above, there are 3 steps.<br>\n1) temporal random crop</p>\n<ul>\n<li>First, I decided to use a total of 22 layers (21-42) out of 65 layers. Of these 22 layers, layers with a range of cropping_min to cropping_max are cropped randomly. After several ablation study, I set cropping_min = 12, cropping_max = 22.</li>\n</ul>\n<p>2) random paste</p>\n<ul>\n<li>The randomly cropped layers are attached to a random area of 22 while maintaining sequential information.<br>\n<strong>This crop &amp; paste method allows the model to learn wide, various range of layers rather than a fixed area of fragments, allowing generalization to learning various fragments.</strong></li>\n</ul>\n<p>3) random cutout</p>\n<ul>\n<li>Similar to spatial cutout augmentation, I applied temporal cutout augmentation. Among the pasted layers, 0 to 2 random layers are filled with 0 values.</li>\n</ul>\n<p>The above three processes can be seen simply by looking at the Pytorch Dataset code part below.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5725749%2F2cb3e87617e91bbff8df5dbdf213838e%2F9.PNG?generation=1686982042582252&amp;alt=media\" alt=\"\"></p>\n<p><br><br>\n<br></p>\n<h3>2. Weight choice with low false positive (★★★★)</h3>\n<p>As shown in the figure below, I saved the predicted mask for each epoch. The file name has the epoch, score, tp, fp, and fn, of which the fp value was important. <strong>Even if the score was similar, if the fp was large, it tended to be very bad in the public score.</strong> Therefore, I wanted to select a weight with a small fp value and a high score value, and I selected an appropriate fp value for each fold through several submissions. In addition, I tried to increase the generalization performance by ensemble models with a range of fp values.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5725749%2Faa46536cf349929607790d1c3a32a21c%2F55.PNG?generation=1687062720342906&amp;alt=media\" alt=\"\"></p>\n<p><br><br>\n<br></p>\n<h3>3. Used models &amp; ensemble (★★★★)</h3>\n<p>I ensembles 3 models.</p>\n<ul>\n<li><p>3D resnet152, 3D resnet200, 3D resnext101 with unet-like decoder.</p></li>\n<li><p>Refered models was firstly inspired by JEBASTIN NADAR (<a href=\"url\" target=\"_blank\">https://www.kaggle.com/code/samfc10/vesuvius-challenge-3d-resnet-training</a>) and applied. Thanks for your effort and respect you. <br>\n3D resnet152, 3D resnet200 is from <a href=\"https://github.com/kenshohara/3D-ResNets-PyTorch\" target=\"_blank\">https://github.com/kenshohara/3D-ResNets-PyTorch</a><br>\n3D resnext101 is from <a href=\"https://github.com/okankop/Efficient-3DCNNs\" target=\"_blank\">https://github.com/okankop/Efficient-3DCNNs</a><br>\nAll models have mit licenses and are not against the rules.</p></li>\n<li><p>3D Resnet was pretrained on Kinetic 710, and 3D resnext101 was pretrained on Kinetic 600.<br>\n<strong>There was a huge difference between being trained in kinetics and not being able to.</strong> Therefore, during the competition, I tried to find a model trained in the kinetic 600 or 700.</p></li>\n<li><p>All models are combined unet-like simple decoder. Features from 3D encoder are upsampled and concatenated using decoder.<br>\ndetails are in code.</p></li>\n</ul>\n<p><br><br>\n<br></p>\n<h3>4. Label smoothing(0.3), cutmix augmentation, data clipping (★★★)</h3>\n<ul>\n<li>Cross entropy with label smoothing improved performance. After doing ablation study, i set the parameter of label smoothing to 0.3</li>\n<li>Cutmix augmentation improved performance in cross validation &amp; public score</li>\n<li>Inspired by AJLAND(<a href=\"url\" target=\"_blank\">https://www.kaggle.com/code/ajland/eda-a-slice-by-slice-analysis</a>), I clipped the image below 50 and more than 200. Also thanks for your effort and respect you.<br>\n(image = np.clip(image, 50, 200))</li>\n</ul>\n<p><br><br>\n<br></p>\n<h3>5. Belief for local cross validation (★★)</h3>\n<p>As shown in the table below, the two models are the ones I submitted finally.<br>\n<strong>Ensemble 2 had a lower public score than ensemble1, but performed better in cross validation.</strong> There were other models that could be adopted, but considering that it came out well in the cross validation, it was adopted, and <strong>the performance was improved more in the private, resulting in a better first place.</strong></p>\n<table>\n<thead>\n<tr>\n<th>models</th>\n<th>cv score</th>\n<th>threshold</th>\n<th>public score</th>\n<th>private score</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Ensemble 1</td>\n<td>0.66 / 0.762/ 0.723 / 0.68</td>\n<td>0.5</td>\n<td>0.795835</td>\n<td>0.663542</td>\n</tr>\n<tr>\n<td>Ensemble 2</td>\n<td>0.67 / 0.764 / 0.724 / 0.69</td>\n<td>0.47</td>\n<td>0.789024</td>\n<td>0.674544</td>\n</tr>\n</tbody>\n</table>\n<p><br><br>\n<br></p>\n<h3>6. Other training details (★)</h3>\n<p><br></p>\n<h4>6.1. loss, optimizer, scheduler</h4>\n<ul>\n<li>loss : CrossEntropyLoss with label smoothing 0.3</li>\n<li>optimizer = AdamW(lr=1e-4)</li>\n<li>scheduler = cosine annealing with warmup</li>\n</ul>\n<p><br></p>\n<h4>6.2. cross validation strategy &amp; result</h4>\n<p>I trained the model with 4fold</p>\n<ul>\n<li>fold1 : fragment 1</li>\n<li>fold3 : fragment 3</li>\n<li>fold2, 4 : fragment 2 divided into 2 sub-frag(9506 x 14830 -&gt; (4300 x 14830) &amp; (5206 x 14830))</li>\n</ul>\n<p>The results are shown in the table below.<br>\nEnsemble means ensemble of 3 models with same weight.</p>\n<table>\n<thead>\n<tr>\n<th>models</th>\n<th>cv score</th>\n<th>threshold</th>\n<th>public score</th>\n<th>private score</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>3D Resnet152</td>\n<td>0.64 / 0.71 / 0.71 / 0.69</td>\n<td>0.5</td>\n<td>0.78</td>\n<td>-</td>\n</tr>\n<tr>\n<td>3D Resnet200</td>\n<td>0.66 / 0.71 / 0.69 / 0.66</td>\n<td>0.5</td>\n<td>0.77</td>\n<td>-</td>\n</tr>\n<tr>\n<td>3D Resnext101</td>\n<td>0.61 / 0.72 / 0.71 / 0.64</td>\n<td>0.5</td>\n<td>0.77</td>\n<td>-</td>\n</tr>\n<tr>\n<td>Ensemble</td>\n<td>0.66 / 0.762/ 0.723 / 0.68</td>\n<td>0.5</td>\n<td>0.795835</td>\n<td>0.663542</td>\n</tr>\n<tr>\n<td>Ensemble</td>\n<td>0.67 / 0.764 / 0.724 / 0.69</td>\n<td>0.47</td>\n<td>0.789024</td>\n<td>0.674544</td>\n</tr>\n</tbody>\n</table>\n<p><br></p>\n<h4>6.3. image augmentation</h4>\n<p><code>all probability set to 0.6</code></p>\n<ul>\n<li>Horizontal, Vertical Flip</li>\n<li>RandomGamma(limit=(50, 150))</li>\n<li>RandomBrightnessContrast(brightness=0.2, contrast=0.2)</li>\n<li>Gaussian Noise(10, 30), Gaussian blur</li>\n<li>shift 0.1, scale 0.1, rotate 360</li>\n<li>coarsedropout(holes=4, size=0.2 * image_size)</li>\n</ul>\n<h4>6.4 Data Preparation</h4>\n<ul>\n<li>stride rate 3</li>\n<li>image size 256</li>\n</ul>\n<h4>6.5 etc</h4>\n<ul>\n<li>To achieve consistent and stable results, I selected the best approach based on threshold 0.5.</li>\n<li>inference stride is (256 / 5)</li>\n</ul>\n<p><br><br>\n<br></p>\n<h3>7. Tried but not worked</h3>\n<ul>\n<li><p>mixup augmentation</p></li>\n<li><p>temporal channel shuffle augmentation</p></li>\n<li><p>3D transformer encoder(uniformerv2, video swin transformer) - cannot training… i don't know why loss didn't decrease.</p></li>\n<li><p>spatial TTA</p></li>\n<li><p>temporal TTA - because of my crop&amp;paste method mentioned above, i tried temporal TTA. In other words, I tried tta by cropping 22, 20, and 18 layers in various ways and then averaging 3 predictions, sometimes performance increased and sometimes decreased. I didn't apply it because it took too much time.</p></li>\n<li><p>other losses(tversky focal loss, fbeta loss)</p></li>\n</ul>\n<p><br><br>\n<br></p>\n<h3>8. Code</h3>\n<p>8.1. Github version(based on Docker)</p>\n<ul>\n<li><a href=\"https://github.com/AhnHeeYoung/Competition/tree/master/kaggle\" target=\"_blank\">https://github.com/AhnHeeYoung/Competition/tree/master/kaggle</a></li>\n</ul>\n<p>8.2. Kaggle Notebook version</p>\n<ul>\n<li>Training : <a href=\"https://www.kaggle.com/code/ahnheeyoung1/ink-detection-training\" target=\"_blank\">https://www.kaggle.com/code/ahnheeyoung1/ink-detection-training</a></li>\n<li>Inference : <a href=\"https://www.kaggle.com/code/ahnheeyoung1/ink-detection-inference/notebook?scriptVersionId=136610637\" target=\"_blank\">https://www.kaggle.com/code/ahnheeyoung1/ink-detection-inference/notebook?scriptVersionId=136610637</a></li>\n</ul>",
      "rawMarkdown": "# Acknowledgement\nI'm very grateful for organizing very challenging and fantastic challenge.\nI feel like there are a lot of really great people and I respect them all.\nI would like to express my gratitude to both the organizers and participants of the competition.\n\n<br />\n<br />\n\n# Solution\nI would like to describe my solution in the order of contribution of performance improvement.\n**The number of ★ in the description below means the degree of contribution to performance improvement.**\n\n\n<br />\n<br />\n\n### 1. Temporal random crop & random paste & random cutout (★★★★★)\n#### 1.1. Thinking about the data itself\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5725749%2Fb9819315244312a47802b241517dc57e%2F4.PNG?generation=1686971322910074&alt=media)\n\n\nAs you can see in the figure above, i thought that particular layers of fragment would not correspond to the same layers of other fragment. Using these points, i came up with an augmentation method that can give strong regularization to the model by reflecting the characteristics of the data.\n\n\n\n#### 1.2. Application\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5725749%2F8601353d50ccdc56721bfc1341395bff%2F8.PNG?generation=1686981811287731&alt=media)\n\nAs you can see in the figure above, there are 3 steps.\n1) temporal random crop\n- First, I decided to use a total of 22 layers (21-42) out of 65 layers. Of these 22 layers, layers with a range of cropping_min to cropping_max are cropped randomly. After several ablation study, I set cropping_min = 12, cropping_max = 22.\n\n\n2) random paste\n- The randomly cropped layers are attached to a random area of 22 while maintaining sequential information.\n**This crop & paste method allows the model to learn wide, various range of layers rather than a fixed area of fragments, allowing generalization to learning various fragments.**\n\n3) random cutout\n- Similar to spatial cutout augmentation, I applied temporal cutout augmentation. Among the pasted layers, 0 to 2 random layers are filled with 0 values.\n\n\nThe above three processes can be seen simply by looking at the Pytorch Dataset code part below.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5725749%2F2cb3e87617e91bbff8df5dbdf213838e%2F9.PNG?generation=1686982042582252&alt=media)\n\n\n\n<br />\n<br />\n### 2. Weight choice with low false positive (★★★★)\n\nAs shown in the figure below, I saved the predicted mask for each epoch. The file name has the epoch, score, tp, fp, and fn, of which the fp value was important. **Even if the score was similar, if the fp was large, it tended to be very bad in the public score.** Therefore, I wanted to select a weight with a small fp value and a high score value, and I selected an appropriate fp value for each fold through several submissions. In addition, I tried to increase the generalization performance by ensemble models with a range of fp values.\n\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5725749%2Faa46536cf349929607790d1c3a32a21c%2F55.PNG?generation=1687062720342906&alt=media)\n\n\n<br />\n<br />\n### 3. Used models & ensemble (★★★★)\nI ensembles 3 models.\n- 3D resnet152, 3D resnet200, 3D resnext101 with unet-like decoder.\n\n- Refered models was firstly inspired by JEBASTIN NADAR ([https://www.kaggle.com/code/samfc10/vesuvius-challenge-3d-resnet-training](url)) and applied. Thanks for your effort and respect you. \n3D resnet152, 3D resnet200 is from https://github.com/kenshohara/3D-ResNets-PyTorch\n3D resnext101 is from https://github.com/okankop/Efficient-3DCNNs\nAll models have mit licenses and are not against the rules.\n\n- 3D Resnet was pretrained on Kinetic 710, and 3D resnext101 was pretrained on Kinetic 600.\n**There was a huge difference between being trained in kinetics and not being able to.** Therefore, during the competition, I tried to find a model trained in the kinetic 600 or 700.\n\n- All models are combined unet-like simple decoder. Features from 3D encoder are upsampled and concatenated using decoder.\ndetails are in code.\n\n\n\n<br />\n<br />\n### 4. Label smoothing(0.3), cutmix augmentation, data clipping (★★★)\n- Cross entropy with label smoothing improved performance. After doing ablation study, i set the parameter of label smoothing to 0.3\n- Cutmix augmentation improved performance in cross validation & public score\n- Inspired by AJLAND([https://www.kaggle.com/code/ajland/eda-a-slice-by-slice-analysis](url)), I clipped the image below 50 and more than 200. Also thanks for your effort and respect you.\n(image = np.clip(image, 50, 200))\n\n\n<br />\n<br />\n### 5. Belief for local cross validation (★★)\n\n\nAs shown in the table below, the two models are the ones I submitted finally.\n**Ensemble 2 had a lower public score than ensemble1, but performed better in cross validation.** There were other models that could be adopted, but considering that it came out well in the cross validation, it was adopted, and **the performance was improved more in the private, resulting in a better first place.**\n\n| models | cv score | threshold | public score | private score\n| --- | --- | --- | --- | --- |\n| Ensemble 1 | 0.66 / 0.762/ 0.723 / 0.68 | 0.5 | 0.795835 | 0.663542 |\n| Ensemble 2 | 0.67 / 0.764 / 0.724 / 0.69 | 0.47 | 0.789024 | 0.674544 |\n\n\n\n\n\n<br />\n<br />\n### 6. Other training details (★)\n\n<br />\n#### 6.1. loss, optimizer, scheduler\n- loss : CrossEntropyLoss with label smoothing 0.3\n- optimizer = AdamW(lr=1e-4)\n- scheduler = cosine annealing with warmup\n\n<br />\n#### 6.2. cross validation strategy & result\nI trained the model with 4fold\n- fold1 : fragment 1\n- fold3 : fragment 3\n- fold2, 4 : fragment 2 divided into 2 sub-frag(9506 x 14830 -> (4300 x 14830) & (5206 x 14830))\n\nThe results are shown in the table below.\nEnsemble means ensemble of 3 models with same weight.\n\n| models | cv score | threshold | public score | private score\n| --- | --- | --- | --- | --- |\n| 3D Resnet152 | 0.64 / 0.71 / 0.71 / 0.69 | 0.5 | 0.78 | - |\n| 3D Resnet200 | 0.66 / 0.71 / 0.69 / 0.66 | 0.5 | 0.77 | - |\n| 3D Resnext101 | 0.61 / 0.72 / 0.71 / 0.64 | 0.5 | 0.77 | - |\n| Ensemble | 0.66 / 0.762/ 0.723 / 0.68 | 0.5 | 0.795835 | 0.663542 |\n| Ensemble | 0.67 / 0.764 / 0.724 / 0.69 | 0.47 | 0.789024 | 0.674544 |\n\n\n\n\n\n<br />\n#### 6.3. image augmentation\n``` all probability set to 0.6 ```\n- Horizontal, Vertical Flip\n- RandomGamma(limit=(50, 150))\n- RandomBrightnessContrast(brightness=0.2, contrast=0.2)\n- Gaussian Noise(10, 30), Gaussian blur\n- shift 0.1, scale 0.1, rotate 360\n- coarsedropout(holes=4, size=0.2 * image_size)\n\n\n#### 6.4 Data Preparation\n- stride rate 3\n- image size 256\n\n#### 6.5 etc\n- To achieve consistent and stable results, I selected the best approach based on threshold 0.5.\n- inference stride is (256 / 5)\n\n\n\n<br />\n<br />\n### 7. Tried but not worked\n- mixup augmentation\n- temporal channel shuffle augmentation\n- 3D transformer encoder(uniformerv2, video swin transformer) - cannot training... i don't know why loss didn't decrease.\n- spatial TTA\n- temporal TTA - because of my crop&paste method mentioned above, i tried temporal TTA. In other words, I tried tta by cropping 22, 20, and 18 layers in various ways and then averaging 3 predictions, sometimes performance increased and sometimes decreased. I didn't apply it because it took too much time.\n\n- other losses(tversky focal loss, fbeta loss)\n\n\n\n<br />\n<br />\n### 8. Code\n8.1. Github version(based on Docker)\n- https://github.com/AhnHeeYoung/Competition/tree/master/kaggle\n\n8.2. Kaggle Notebook version\n- Training : https://www.kaggle.com/code/ahnheeyoung1/ink-detection-training\n- Inference : https://www.kaggle.com/code/ahnheeyoung1/ink-detection-inference/notebook?scriptVersionId=136610637",
      "votes": null
    },
    {
      "id": "2317415",
      "postDate": "06/25/2023 17:31:46",
      "content": "<p>Thanks for sharing and explaining your project! it makes the difference.</p>",
      "rawMarkdown": "Thanks for sharing and explaining your project! it makes the difference.",
      "votes": null
    },
    {
      "id": "2319408",
      "postDate": "06/27/2023 05:37:19",
      "content": "<p>I publised the (docker based) code on my github.<br>\n(<a href=\"https://github.com/AhnHeeYoung/Competition/tree/master/kaggle\" target=\"_blank\">https://github.com/AhnHeeYoung/Competition/tree/master/kaggle</a>)</p>\n<p>Thank you!</p>",
      "rawMarkdown": "I publised the (docker based) code on my github.\n(https://github.com/AhnHeeYoung/Competition/tree/master/kaggle)\n\nThank you!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2317415,
      "author_name": "guillermoperezg",
      "author_url": "",
      "post_date": "06/25/2023 17:31:46",
      "content": "<p>Thanks for sharing and explaining your project! it makes the difference.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2319408,
      "author_name": "ahnheeyoung1",
      "author_url": "",
      "post_date": "06/27/2023 05:37:19",
      "content": "<p>I publised the (docker based) code on my github.<br>\n(<a href=\"https://github.com/AhnHeeYoung/Competition/tree/master/kaggle\" target=\"_blank\">https://github.com/AhnHeeYoung/Competition/tree/master/kaggle</a>)</p>\n<p>Thank you!</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2306168": "# Acknowledgement\nI'm very grateful for organizing very challenging and fantastic challenge.\nI feel like there are a lot of really great people and I respect them all.\nI would like to express my gratitude to both the organizers and participants of the competition.\n\n<br />\n<br />\n\n# Solution\nI would like to describe my solution in the order of contribution of performance improvement.\n**The number of ★ in the description below means the degree of contribution to performance improvement.**\n\n\n<br />\n<br />\n\n### 1. Temporal random crop & random paste & random cutout (★★★★★)\n#### 1.1. Thinking about the data itself\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5725749%2Fb9819315244312a47802b241517dc57e%2F4.PNG?generation=1686971322910074&alt=media)\n\n\nAs you can see in the figure above, i thought that particular layers of fragment would not correspond to the same layers of other fragment. Using these points, i came up with an augmentation method that can give strong regularization to the model by reflecting the characteristics of the data.\n\n\n\n#### 1.2. Application\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5725749%2F8601353d50ccdc56721bfc1341395bff%2F8.PNG?generation=1686981811287731&alt=media)\n\nAs you can see in the figure above, there are 3 steps.\n1) temporal random crop\n- First, I decided to use a total of 22 layers (21-42) out of 65 layers. Of these 22 layers, layers with a range of cropping_min to cropping_max are cropped randomly. After several ablation study, I set cropping_min = 12, cropping_max = 22.\n\n\n2) random paste\n- The randomly cropped layers are attached to a random area of 22 while maintaining sequential information.\n**This crop & paste method allows the model to learn wide, various range of layers rather than a fixed area of fragments, allowing generalization to learning various fragments.**\n\n3) random cutout\n- Similar to spatial cutout augmentation, I applied temporal cutout augmentation. Among the pasted layers, 0 to 2 random layers are filled with 0 values.\n\n\nThe above three processes can be seen simply by looking at the Pytorch Dataset code part below.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5725749%2F2cb3e87617e91bbff8df5dbdf213838e%2F9.PNG?generation=1686982042582252&alt=media)\n\n\n\n<br />\n<br />\n### 2. Weight choice with low false positive (★★★★)\n\nAs shown in the figure below, I saved the predicted mask for each epoch. The file name has the epoch, score, tp, fp, and fn, of which the fp value was important. **Even if the score was similar, if the fp was large, it tended to be very bad in the public score.** Therefore, I wanted to select a weight with a small fp value and a high score value, and I selected an appropriate fp value for each fold through several submissions. In addition, I tried to increase the generalization performance by ensemble models with a range of fp values.\n\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5725749%2Faa46536cf349929607790d1c3a32a21c%2F55.PNG?generation=1687062720342906&alt=media)\n\n\n<br />\n<br />\n### 3. Used models & ensemble (★★★★)\nI ensembles 3 models.\n- 3D resnet152, 3D resnet200, 3D resnext101 with unet-like decoder.\n\n- Refered models was firstly inspired by JEBASTIN NADAR ([https://www.kaggle.com/code/samfc10/vesuvius-challenge-3d-resnet-training](url)) and applied. Thanks for your effort and respect you. \n3D resnet152, 3D resnet200 is from https://github.com/kenshohara/3D-ResNets-PyTorch\n3D resnext101 is from https://github.com/okankop/Efficient-3DCNNs\nAll models have mit licenses and are not against the rules.\n\n- 3D Resnet was pretrained on Kinetic 710, and 3D resnext101 was pretrained on Kinetic 600.\n**There was a huge difference between being trained in kinetics and not being able to.** Therefore, during the competition, I tried to find a model trained in the kinetic 600 or 700.\n\n- All models are combined unet-like simple decoder. Features from 3D encoder are upsampled and concatenated using decoder.\ndetails are in code.\n\n\n\n<br />\n<br />\n### 4. Label smoothing(0.3), cutmix augmentation, data clipping (★★★)\n- Cross entropy with label smoothing improved performance. After doing ablation study, i set the parameter of label smoothing to 0.3\n- Cutmix augmentation improved performance in cross validation & public score\n- Inspired by AJLAND([https://www.kaggle.com/code/ajland/eda-a-slice-by-slice-analysis](url)), I clipped the image below 50 and more than 200. Also thanks for your effort and respect you.\n(image = np.clip(image, 50, 200))\n\n\n<br />\n<br />\n### 5. Belief for local cross validation (★★)\n\n\nAs shown in the table below, the two models are the ones I submitted finally.\n**Ensemble 2 had a lower public score than ensemble1, but performed better in cross validation.** There were other models that could be adopted, but considering that it came out well in the cross validation, it was adopted, and **the performance was improved more in the private, resulting in a better first place.**\n\n| models | cv score | threshold | public score | private score\n| --- | --- | --- | --- | --- |\n| Ensemble 1 | 0.66 / 0.762/ 0.723 / 0.68 | 0.5 | 0.795835 | 0.663542 |\n| Ensemble 2 | 0.67 / 0.764 / 0.724 / 0.69 | 0.47 | 0.789024 | 0.674544 |\n\n\n\n\n\n<br />\n<br />\n### 6. Other training details (★)\n\n<br />\n#### 6.1. loss, optimizer, scheduler\n- loss : CrossEntropyLoss with label smoothing 0.3\n- optimizer = AdamW(lr=1e-4)\n- scheduler = cosine annealing with warmup\n\n<br />\n#### 6.2. cross validation strategy & result\nI trained the model with 4fold\n- fold1 : fragment 1\n- fold3 : fragment 3\n- fold2, 4 : fragment 2 divided into 2 sub-frag(9506 x 14830 -> (4300 x 14830) & (5206 x 14830))\n\nThe results are shown in the table below.\nEnsemble means ensemble of 3 models with same weight.\n\n| models | cv score | threshold | public score | private score\n| --- | --- | --- | --- | --- |\n| 3D Resnet152 | 0.64 / 0.71 / 0.71 / 0.69 | 0.5 | 0.78 | - |\n| 3D Resnet200 | 0.66 / 0.71 / 0.69 / 0.66 | 0.5 | 0.77 | - |\n| 3D Resnext101 | 0.61 / 0.72 / 0.71 / 0.64 | 0.5 | 0.77 | - |\n| Ensemble | 0.66 / 0.762/ 0.723 / 0.68 | 0.5 | 0.795835 | 0.663542 |\n| Ensemble | 0.67 / 0.764 / 0.724 / 0.69 | 0.47 | 0.789024 | 0.674544 |\n\n\n\n\n\n<br />\n#### 6.3. image augmentation\n``` all probability set to 0.6 ```\n- Horizontal, Vertical Flip\n- RandomGamma(limit=(50, 150))\n- RandomBrightnessContrast(brightness=0.2, contrast=0.2)\n- Gaussian Noise(10, 30), Gaussian blur\n- shift 0.1, scale 0.1, rotate 360\n- coarsedropout(holes=4, size=0.2 * image_size)\n\n\n#### 6.4 Data Preparation\n- stride rate 3\n- image size 256\n\n#### 6.5 etc\n- To achieve consistent and stable results, I selected the best approach based on threshold 0.5.\n- inference stride is (256 / 5)\n\n\n\n<br />\n<br />\n### 7. Tried but not worked\n- mixup augmentation\n- temporal channel shuffle augmentation\n- 3D transformer encoder(uniformerv2, video swin transformer) - cannot training... i don't know why loss didn't decrease.\n- spatial TTA\n- temporal TTA - because of my crop&paste method mentioned above, i tried temporal TTA. In other words, I tried tta by cropping 22, 20, and 18 layers in various ways and then averaging 3 predictions, sometimes performance increased and sometimes decreased. I didn't apply it because it took too much time.\n\n- other losses(tversky focal loss, fbeta loss)\n\n\n\n<br />\n<br />\n### 8. Code\n8.1. Github version(based on Docker)\n- https://github.com/AhnHeeYoung/Competition/tree/master/kaggle\n\n8.2. Kaggle Notebook version\n- Training : https://www.kaggle.com/code/ahnheeyoung1/ink-detection-training\n- Inference : https://www.kaggle.com/code/ahnheeyoung1/ink-detection-inference/notebook?scriptVersionId=136610637",
    "2317415": "Thanks for sharing and explaining your project! it makes the difference.",
    "2319408": "I publised the (docker based) code on my github.\n(https://github.com/AhnHeeYoung/Competition/tree/master/kaggle)\n\nThank you!"
  },
  "source": "meta"
}