{
  "id": 417274,
  "title": "6th place solution",
  "url": "/competitions/vesuvius-challenge-ink-detection/writeups/chumajin-6th-place-solution",
  "author_name": "",
  "post_date": "2023-07-03T05:07:54.860Z",
  "votes": 84,
  "comment_count": 42,
  "views": 0,
  "content": "<p>Thank you very much for organizing such an interesting competition. I am greatly thankful to the hosts and the Kaggle staff.</p>\n<p>Continuing from the previous competition, I am delighted to have won solo gold medal again, with a total of four medals (Table 1, NLP × 2, CV × 1). Additionally, it was my first time attempting the segmentation task, and I began with <a href=\"https://www.kaggle.com/tanakar\" target=\"_blank\">@tanakar</a> excellent notebook <a href=\"https://www.kaggle.com/code/tanakar/2-5d-segmentaion-baseline-training/comments\" target=\"_blank\">here</a>. I am truly grateful for that.</p>\n<h1>1. Summary</h1>\n<p>My approach involved an ensemble of EfficientNet and SegFormer models. I believed that the test data was rotated shown in other discussions, so I rotated the images during inference, which resulted in a significant boost at the beginning (Public LB 0.58 → 0.74). Additionally, my originality came from incorporating IR images into the training data, which gave me a CV score increase of 0.01 and an LB score increase of 0.01. I will now explain the details below.</p>\n<h1>2. Inference</h1>\n<h2>2.1 About test data and inference flow</h2>\n<p>Based on the brief LB probing and the information provided on the competition page, I had an idea of what the test data might look like. Here is an image that represents my understanding:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4001300%2Fb9c4e129fe3285876db0a7135f704edb%2Fsolution1.jpg?generation=1686795321876414&amp;alt=media\" alt=\"\"></p>\n<p>Therefore, I structured my inference code in the Kaggle notebook as follows:</p>\n<ol>\n<li>Concatenate(axis=1) fragments A and B.</li>\n<li>Rotate the image clockwise.</li>\n<li>Perform inference (original + h flip TTA).</li>\n<li>Rotate the prediction countor-clockwise back to its original position.</li>\n<li>Cut and encode each fragment A and B respectively.</li>\n</ol>\n<p>Of course, to reduce inference time, I skipped the inference for areas where the mask value was 0. Furthermore, instead of inferring fragment A and B separately, I concatenated them. This not only eliminated the 0 padding at the boundary between A and B but also allowed for continuous inference of the initial part of fragment B as a contiguous sequence. These led to a significant boost in my LB score (EfficientNet B4: 0.58 → 0.74)</p>\n<h2>2.2 Threshold</h2>\n<p>I believe many of you experienced the instability of the signal values. Therefore, I used the following function to rank the entire image and calculate percentiles. Then, by applying a threshold, I obtained a stable threshold value. This approach proved helpful not only during inference but also during ensemble processes. 2nd place solution also used the same way <a href=\"https://www.kaggle.com/competitions/vesuvius-challenge-ink-detection/discussion/417255\" target=\"_blank\">here</a></p>\n<pre><code>def get_percentile(array):\n     org_shape = array.shape\n    = array.reshape(-1)\n    = np.arange(len(array))[array.argsort().argsort()]\n    =/ array.max()\n    = array.reshape(org_shape)\n    array\n</code></pre>\n<p>For the Public LB, the optimal threshold was found to be 0.96. However, using fragment 3, I conducted a simulation to observe the correlation between the partially optimal threshold (around 10%) and the threshold for the remaining 90%. As a result, I noticed that the threshold was overfitting for the 10% portion (likely reducing noise), while for the remaining 90%, it was better to slightly lower the threshold below the optimal value (aiming for clearer extraction of text). In fact, when comparing the same model, a threshold of 0.95 performed slightly better for the private LB(but less than 0.01). For the final submission, I used different models: sub1 with a threshold of 0.96 and sub2 with a threshold of 0.95.</p>\n<h1>3. Training</h1>\n<p>The following is an overview of the training process. Similar to inference, I created three sets of data and took their averages. It should be noted that SegFormer differs from CNN as it can only utilize 3 channels. As mentioned earlier, incorporating IR images resulted in improvements in both CV and LB scores.</p>\n<h2>3.1 CNN + Unet</h2>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4001300%2Fa32d067916dee939f3301b72d8daf9f8%2Fsolution2.jpg?generation=1686795342236188&amp;alt=media\" alt=\"\"></p>\n<h2>3.2 SegFormer</h2>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4001300%2Fba6b7cdc7012775d09930b49179e0317%2Fsolution3.jpg?generation=1686795355254555&amp;alt=media\" alt=\"\"></p>\n<h2>3.3 Fine-tuned Parameters</h2>\n<ul>\n<li>Stride: image size // 4</li>\n<li>Optimizer: Adam</li>\n<li>Epochs: 20</li>\n<li>Early stopping: 4</li>\n<li>Scheduler: get_cosine_schedule_with_warmup (from transformers)</li>\n<li>Warm-up: 0.1</li>\n<li>Gradient norm: 10</li>\n<li>Loss function: SoftBCEWithLogitsLoss (segmentation model in PyTorch)</li>\n<li>Training excludes areas with a mask value of 0.</li>\n<li>TTA: Horizontal flip</li>\n</ul>\n<h2>3.4 Cross Validation</h2>\n<p>For submission1, I used a 7kfold cross-validation, and for submission2, I used a 10kfold cross-validation. Increasing the value of k-fold resulted in improvements in both CV and LB scores. I recall that increasing from 5-fold to 7-fold led to an improvement of approximately 0.1 in the Public LB score.</p>\n<h1>4 Final result</h1>\n<p>Ensemble was all mean value of predictions.</p>\n<p>sub1 : th 0.96, cv 0.740, public LB 0.811570, private LB 0.661339</p>\n<table>\n<thead>\n<tr>\n<th>model</th>\n<th>image size</th>\n<th>kfold</th>\n<th>cv</th>\n<th>public LB</th>\n<th>private LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>efficientnet_b7_ns</td>\n<td>608</td>\n<td>7 + fulltrain</td>\n<td>0.712</td>\n<td>0.80</td>\n<td>0.64</td>\n</tr>\n<tr>\n<td>efficientnet_b6_ns</td>\n<td>544</td>\n<td>7 + fulltrain</td>\n<td>0.702</td>\n<td>0.79</td>\n<td>0.64</td>\n</tr>\n<tr>\n<td>efficientnetv2_l_in21ft1k</td>\n<td>480</td>\n<td>7</td>\n<td>0.707</td>\n<td>0.79</td>\n<td>0.65</td>\n</tr>\n<tr>\n<td>tf_efficientnet_b8</td>\n<td>672</td>\n<td>7</td>\n<td>0.716</td>\n<td>0.79</td>\n<td>0.64</td>\n</tr>\n<tr>\n<td>segformer b3</td>\n<td>1024</td>\n<td>7</td>\n<td>0.738</td>\n<td>0.78</td>\n<td>0.66</td>\n</tr>\n</tbody>\n</table>\n<p>sub2 : th 0.95,cv 0.746 , public LB 0.799563, private LB 0.654812</p>\n<table>\n<thead>\n<tr>\n<th>model</th>\n<th>image size</th>\n<th>kfold</th>\n<th>cv</th>\n<th>public LB</th>\n<th>private LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>efficientnet_b7_ns</td>\n<td>608</td>\n<td>10</td>\n<td>0.722</td>\n<td>0.80</td>\n<td>0.65</td>\n</tr>\n<tr>\n<td>efficientnet_b6_ns</td>\n<td>544</td>\n<td>10</td>\n<td>0.720</td>\n<td>0.79</td>\n<td>0.63</td>\n</tr>\n<tr>\n<td>efficientnetv2_l_in21ft1k</td>\n<td>480</td>\n<td>10</td>\n<td>0.717</td>\n<td>0.79</td>\n<td>0.65</td>\n</tr>\n<tr>\n<td>segformer b3</td>\n<td>1024</td>\n<td>7</td>\n<td>0.738</td>\n<td>0.78</td>\n<td>0.66</td>\n</tr>\n</tbody>\n</table>\n<h2>4.1 Visualization of predictions</h2>\n<p>The following images visualize the predictions of submission1.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4001300%2F985af9f03d54b031cc820f29b63bc891%2Fpredictions.jpg?generation=1686889411081867&amp;alt=media\" alt=\"\"></p>\n<h1>5. My understanding</h1>\n<h2>5.1 Not working well</h2>\n<p>I tried several models such as ConvNext, Mask2Former, Swin Transformer + PSPNet, BeiT, and many others, but their effectiveness was not satisfactory for cv and lb. EfficientNet and mobilevit performed well and stable in this competition for me. I also experimented with SegFormer using various versions from b1 to b5. Although it showed good performance in cross-validation (CV), the leaderboard (LB) scores were poor and unstable. Five days before the end of the competition, when I plotted the relationship between CV and LB scores again, I noticed that larger models tended to overfit. They achieved good CV scores but had poor LB scores. After adjusting the layers used, I found that only b3 showed high LB scores, although I suspected it might be overfitting. When I used it in the ensemble, it significantly improved the LB scores, so I decided to include it. This discrepancy may be due to the limited amount of training data. I should have also tried regularization techniques such as dropout, freezing, and other strategies, but I ran out of time. Considering these options might have potentially improved the performance. <br>\n※ These are just my guesses.</p>\n<h2>5.2 Potential Successes That Were Not Implemented</h2>\n<p>Pre-training using IR images: Although it improved the CV performance, it resulted in a decline in LB scores, so it was not implemented.<br>\nIncluding EMNIST (external data) in the training dataset: While it improved the CV performance, it led to a deterioration in LB scores, so it was not implemented.</p>\n<h1>6. Acknowledgments</h1>\n<p>I could not have achieved these results on my own. I was greatly influenced by those who I have collaborated with in the past, and I am grateful for their contributions. I would also like to express my sincere gratitude to those who have shared their knowledge and insights through previous competitions. Thank you very much.</p>\n<p>training code : <a href=\"https://github.com/chumajin/kaggle-VCID\" target=\"_blank\">https://github.com/chumajin/kaggle-VCID</a><br>\ninference code : <a href=\"https://www.kaggle.com/code/chumajin/vcid-6th-place-inference\" target=\"_blank\">https://www.kaggle.com/code/chumajin/vcid-6th-place-inference</a></p>",
  "messages": [
    {
      "id": "2302987",
      "postDate": "06/15/2023 02:29:10",
      "content": "<p>Thank you very much for organizing such an interesting competition. I am greatly thankful to the hosts and the Kaggle staff.</p>\n<p>Continuing from the previous competition, I am delighted to have won solo gold medal again, with a total of four medals (Table 1, NLP × 2, CV × 1). Additionally, it was my first time attempting the segmentation task, and I began with <a href=\"https://www.kaggle.com/tanakar\" target=\"_blank\">@tanakar</a> excellent notebook <a href=\"https://www.kaggle.com/code/tanakar/2-5d-segmentaion-baseline-training/comments\" target=\"_blank\">here</a>. I am truly grateful for that.</p>\n<h1>1. Summary</h1>\n<p>My approach involved an ensemble of EfficientNet and SegFormer models. I believed that the test data was rotated shown in other discussions, so I rotated the images during inference, which resulted in a significant boost at the beginning (Public LB 0.58 → 0.74). Additionally, my originality came from incorporating IR images into the training data, which gave me a CV score increase of 0.01 and an LB score increase of 0.01. I will now explain the details below.</p>\n<h1>2. Inference</h1>\n<h2>2.1 About test data and inference flow</h2>\n<p>Based on the brief LB probing and the information provided on the competition page, I had an idea of what the test data might look like. Here is an image that represents my understanding:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4001300%2Fb9c4e129fe3285876db0a7135f704edb%2Fsolution1.jpg?generation=1686795321876414&amp;alt=media\" alt=\"\"></p>\n<p>Therefore, I structured my inference code in the Kaggle notebook as follows:</p>\n<ol>\n<li>Concatenate(axis=1) fragments A and B.</li>\n<li>Rotate the image clockwise.</li>\n<li>Perform inference (original + h flip TTA).</li>\n<li>Rotate the prediction countor-clockwise back to its original position.</li>\n<li>Cut and encode each fragment A and B respectively.</li>\n</ol>\n<p>Of course, to reduce inference time, I skipped the inference for areas where the mask value was 0. Furthermore, instead of inferring fragment A and B separately, I concatenated them. This not only eliminated the 0 padding at the boundary between A and B but also allowed for continuous inference of the initial part of fragment B as a contiguous sequence. These led to a significant boost in my LB score (EfficientNet B4: 0.58 → 0.74)</p>\n<h2>2.2 Threshold</h2>\n<p>I believe many of you experienced the instability of the signal values. Therefore, I used the following function to rank the entire image and calculate percentiles. Then, by applying a threshold, I obtained a stable threshold value. This approach proved helpful not only during inference but also during ensemble processes. 2nd place solution also used the same way <a href=\"https://www.kaggle.com/competitions/vesuvius-challenge-ink-detection/discussion/417255\" target=\"_blank\">here</a></p>\n<pre><code>def get_percentile(array):\n     org_shape = array.shape\n    = array.reshape(-1)\n    = np.arange(len(array))[array.argsort().argsort()]\n    =/ array.max()\n    = array.reshape(org_shape)\n    array\n</code></pre>\n<p>For the Public LB, the optimal threshold was found to be 0.96. However, using fragment 3, I conducted a simulation to observe the correlation between the partially optimal threshold (around 10%) and the threshold for the remaining 90%. As a result, I noticed that the threshold was overfitting for the 10% portion (likely reducing noise), while for the remaining 90%, it was better to slightly lower the threshold below the optimal value (aiming for clearer extraction of text). In fact, when comparing the same model, a threshold of 0.95 performed slightly better for the private LB(but less than 0.01). For the final submission, I used different models: sub1 with a threshold of 0.96 and sub2 with a threshold of 0.95.</p>\n<h1>3. Training</h1>\n<p>The following is an overview of the training process. Similar to inference, I created three sets of data and took their averages. It should be noted that SegFormer differs from CNN as it can only utilize 3 channels. As mentioned earlier, incorporating IR images resulted in improvements in both CV and LB scores.</p>\n<h2>3.1 CNN + Unet</h2>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4001300%2Fa32d067916dee939f3301b72d8daf9f8%2Fsolution2.jpg?generation=1686795342236188&amp;alt=media\" alt=\"\"></p>\n<h2>3.2 SegFormer</h2>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4001300%2Fba6b7cdc7012775d09930b49179e0317%2Fsolution3.jpg?generation=1686795355254555&amp;alt=media\" alt=\"\"></p>\n<h2>3.3 Fine-tuned Parameters</h2>\n<ul>\n<li>Stride: image size // 4</li>\n<li>Optimizer: Adam</li>\n<li>Epochs: 20</li>\n<li>Early stopping: 4</li>\n<li>Scheduler: get_cosine_schedule_with_warmup (from transformers)</li>\n<li>Warm-up: 0.1</li>\n<li>Gradient norm: 10</li>\n<li>Loss function: SoftBCEWithLogitsLoss (segmentation model in PyTorch)</li>\n<li>Training excludes areas with a mask value of 0.</li>\n<li>TTA: Horizontal flip</li>\n</ul>\n<h2>3.4 Cross Validation</h2>\n<p>For submission1, I used a 7kfold cross-validation, and for submission2, I used a 10kfold cross-validation. Increasing the value of k-fold resulted in improvements in both CV and LB scores. I recall that increasing from 5-fold to 7-fold led to an improvement of approximately 0.1 in the Public LB score.</p>\n<h1>4 Final result</h1>\n<p>Ensemble was all mean value of predictions.</p>\n<p>sub1 : th 0.96, cv 0.740, public LB 0.811570, private LB 0.661339</p>\n<table>\n<thead>\n<tr>\n<th>model</th>\n<th>image size</th>\n<th>kfold</th>\n<th>cv</th>\n<th>public LB</th>\n<th>private LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>efficientnet_b7_ns</td>\n<td>608</td>\n<td>7 + fulltrain</td>\n<td>0.712</td>\n<td>0.80</td>\n<td>0.64</td>\n</tr>\n<tr>\n<td>efficientnet_b6_ns</td>\n<td>544</td>\n<td>7 + fulltrain</td>\n<td>0.702</td>\n<td>0.79</td>\n<td>0.64</td>\n</tr>\n<tr>\n<td>efficientnetv2_l_in21ft1k</td>\n<td>480</td>\n<td>7</td>\n<td>0.707</td>\n<td>0.79</td>\n<td>0.65</td>\n</tr>\n<tr>\n<td>tf_efficientnet_b8</td>\n<td>672</td>\n<td>7</td>\n<td>0.716</td>\n<td>0.79</td>\n<td>0.64</td>\n</tr>\n<tr>\n<td>segformer b3</td>\n<td>1024</td>\n<td>7</td>\n<td>0.738</td>\n<td>0.78</td>\n<td>0.66</td>\n</tr>\n</tbody>\n</table>\n<p>sub2 : th 0.95,cv 0.746 , public LB 0.799563, private LB 0.654812</p>\n<table>\n<thead>\n<tr>\n<th>model</th>\n<th>image size</th>\n<th>kfold</th>\n<th>cv</th>\n<th>public LB</th>\n<th>private LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>efficientnet_b7_ns</td>\n<td>608</td>\n<td>10</td>\n<td>0.722</td>\n<td>0.80</td>\n<td>0.65</td>\n</tr>\n<tr>\n<td>efficientnet_b6_ns</td>\n<td>544</td>\n<td>10</td>\n<td>0.720</td>\n<td>0.79</td>\n<td>0.63</td>\n</tr>\n<tr>\n<td>efficientnetv2_l_in21ft1k</td>\n<td>480</td>\n<td>10</td>\n<td>0.717</td>\n<td>0.79</td>\n<td>0.65</td>\n</tr>\n<tr>\n<td>segformer b3</td>\n<td>1024</td>\n<td>7</td>\n<td>0.738</td>\n<td>0.78</td>\n<td>0.66</td>\n</tr>\n</tbody>\n</table>\n<h2>4.1 Visualization of predictions</h2>\n<p>The following images visualize the predictions of submission1.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4001300%2F985af9f03d54b031cc820f29b63bc891%2Fpredictions.jpg?generation=1686889411081867&amp;alt=media\" alt=\"\"></p>\n<h1>5. My understanding</h1>\n<h2>5.1 Not working well</h2>\n<p>I tried several models such as ConvNext, Mask2Former, Swin Transformer + PSPNet, BeiT, and many others, but their effectiveness was not satisfactory for cv and lb. EfficientNet and mobilevit performed well and stable in this competition for me. I also experimented with SegFormer using various versions from b1 to b5. Although it showed good performance in cross-validation (CV), the leaderboard (LB) scores were poor and unstable. Five days before the end of the competition, when I plotted the relationship between CV and LB scores again, I noticed that larger models tended to overfit. They achieved good CV scores but had poor LB scores. After adjusting the layers used, I found that only b3 showed high LB scores, although I suspected it might be overfitting. When I used it in the ensemble, it significantly improved the LB scores, so I decided to include it. This discrepancy may be due to the limited amount of training data. I should have also tried regularization techniques such as dropout, freezing, and other strategies, but I ran out of time. Considering these options might have potentially improved the performance. <br>\n※ These are just my guesses.</p>\n<h2>5.2 Potential Successes That Were Not Implemented</h2>\n<p>Pre-training using IR images: Although it improved the CV performance, it resulted in a decline in LB scores, so it was not implemented.<br>\nIncluding EMNIST (external data) in the training dataset: While it improved the CV performance, it led to a deterioration in LB scores, so it was not implemented.</p>\n<h1>6. Acknowledgments</h1>\n<p>I could not have achieved these results on my own. I was greatly influenced by those who I have collaborated with in the past, and I am grateful for their contributions. I would also like to express my sincere gratitude to those who have shared their knowledge and insights through previous competitions. Thank you very much.</p>\n<p>training code : <a href=\"https://github.com/chumajin/kaggle-VCID\" target=\"_blank\">https://github.com/chumajin/kaggle-VCID</a><br>\ninference code : <a href=\"https://www.kaggle.com/code/chumajin/vcid-6th-place-inference\" target=\"_blank\">https://www.kaggle.com/code/chumajin/vcid-6th-place-inference</a></p>",
      "rawMarkdown": "Thank you very much for organizing such an interesting competition. I am greatly thankful to the hosts and the Kaggle staff.\n\nContinuing from the previous competition, I am delighted to have won solo gold medal again, with a total of four medals (Table 1, NLP × 2, CV × 1). Additionally, it was my first time attempting the segmentation task, and I began with @tanakar excellent notebook [here](https://www.kaggle.com/code/tanakar/2-5d-segmentaion-baseline-training/comments). I am truly grateful for that.\n\n\n\n# 1. Summary\nMy approach involved an ensemble of EfficientNet and SegFormer models. I believed that the test data was rotated shown in other discussions, so I rotated the images during inference, which resulted in a significant boost at the beginning (Public LB 0.58 → 0.74). Additionally, my originality came from incorporating IR images into the training data, which gave me a CV score increase of 0.01 and an LB score increase of 0.01. I will now explain the details below.\n\n# 2. Inference\n## 2.1 About test data and inference flow\nBased on the brief LB probing and the information provided on the competition page, I had an idea of what the test data might look like. Here is an image that represents my understanding:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4001300%2Fb9c4e129fe3285876db0a7135f704edb%2Fsolution1.jpg?generation=1686795321876414&alt=media)\n\nTherefore, I structured my inference code in the Kaggle notebook as follows:\n\n1. Concatenate(axis=1) fragments A and B.\n2. Rotate the image clockwise.\n3. Perform inference (original + h flip TTA).\n4. Rotate the prediction countor-clockwise back to its original position.\n5. Cut and encode each fragment A and B respectively.\n\n\n\nOf course, to reduce inference time, I skipped the inference for areas where the mask value was 0. Furthermore, instead of inferring fragment A and B separately, I concatenated them. This not only eliminated the 0 padding at the boundary between A and B but also allowed for continuous inference of the initial part of fragment B as a contiguous sequence. These led to a significant boost in my LB score (EfficientNet B4: 0.58 → 0.74)\n\n\n## 2.2 Threshold\nI believe many of you experienced the instability of the signal values. Therefore, I used the following function to rank the entire image and calculate percentiles. Then, by applying a threshold, I obtained a stable threshold value. This approach proved helpful not only during inference but also during ensemble processes. 2nd place solution also used the same way [here](https://www.kaggle.com/competitions/vesuvius-challenge-ink-detection/discussion/417255)\n\n\n~~~\ndef get_percentile(array):\n     org_shape = array.shape\n     array = array.reshape(-1)\n     array = np.arange(len(array))[array.argsort().argsort()]\n     array = array / array.max()\n     array = array.reshape(org_shape)\n     return array\n~~~\n\nFor the Public LB, the optimal threshold was found to be 0.96. However, using fragment 3, I conducted a simulation to observe the correlation between the partially optimal threshold (around 10%) and the threshold for the remaining 90%. As a result, I noticed that the threshold was overfitting for the 10% portion (likely reducing noise), while for the remaining 90%, it was better to slightly lower the threshold below the optimal value (aiming for clearer extraction of text). In fact, when comparing the same model, a threshold of 0.95 performed slightly better for the private LB(but less than 0.01). For the final submission, I used different models: sub1 with a threshold of 0.96 and sub2 with a threshold of 0.95.\n\n\n\n# 3. Training\nThe following is an overview of the training process. Similar to inference, I created three sets of data and took their averages. It should be noted that SegFormer differs from CNN as it can only utilize 3 channels. As mentioned earlier, incorporating IR images resulted in improvements in both CV and LB scores.\n\n\n\n## 3.1 CNN + Unet\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4001300%2Fa32d067916dee939f3301b72d8daf9f8%2Fsolution2.jpg?generation=1686795342236188&alt=media)\n\n## 3.2 SegFormer\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4001300%2Fba6b7cdc7012775d09930b49179e0317%2Fsolution3.jpg?generation=1686795355254555&alt=media)\n\n## 3.3 Fine-tuned Parameters\n* Stride: image size // 4\n* Optimizer: Adam\n* Epochs: 20\n* Early stopping: 4\n* Scheduler: get_cosine_schedule_with_warmup (from transformers)\n* Warm-up: 0.1\n* Gradient norm: 10\n* Loss function: SoftBCEWithLogitsLoss (segmentation model in PyTorch)\n* Training excludes areas with a mask value of 0.\n* TTA: Horizontal flip\n\n## 3.4 Cross Validation\nFor submission1, I used a 7kfold cross-validation, and for submission2, I used a 10kfold cross-validation. Increasing the value of k-fold resulted in improvements in both CV and LB scores. I recall that increasing from 5-fold to 7-fold led to an improvement of approximately 0.1 in the Public LB score.\n\n# 4 Final result\n\nEnsemble was all mean value of predictions.\n\nsub1 : th 0.96, cv 0.740, public LB 0.811570, private LB 0.661339\n\n| model                     | image size | kfold         | cv     | public LB | private LB |\n|---------------------------|------------|---------------|--------|-----------|------------|\n| efficientnet_b7_ns        | 608        | 7 + fulltrain | 0.712  | 0.80      | 0.64       |\n| efficientnet_b6_ns        | 544        | 7 + fulltrain | 0.702  | 0.79      | 0.64       |\n| efficientnetv2_l_in21ft1k | 480        | 7             | 0.707  | 0.79      | 0.65       |\n| tf_efficientnet_b8        | 672        | 7             | 0.716  | 0.79      | 0.64       |\n| segformer b3              | 1024       | 7             | 0.738  | 0.78      | 0.66       |\n\n\nsub2 : th 0.95,cv 0.746 , public LB 0.799563, private LB 0.654812\n\n| model                     | image size | kfold | cv     | public LB | private LB |\n|---------------------------|------------|-------|--------|-----------|------------|\n| efficientnet_b7_ns        | 608        | 10    | 0.722  | 0.80      | 0.65       |\n| efficientnet_b6_ns        | 544        | 10    | 0.720  | 0.79      | 0.63       |\n| efficientnetv2_l_in21ft1k | 480        | 10    | 0.717  | 0.79      | 0.65       |\n| segformer b3              | 1024       | 7     | 0.738  | 0.78      | 0.66       |\n\n## 4.1 Visualization of predictions\n\nThe following images visualize the predictions of submission1.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4001300%2F985af9f03d54b031cc820f29b63bc891%2Fpredictions.jpg?generation=1686889411081867&alt=media)\n\n# 5. My understanding\n## 5.1 Not working well\n\nI tried several models such as ConvNext, Mask2Former, Swin Transformer + PSPNet, BeiT, and many others, but their effectiveness was not satisfactory for cv and lb. EfficientNet and mobilevit performed well and stable in this competition for me. I also experimented with SegFormer using various versions from b1 to b5. Although it showed good performance in cross-validation (CV), the leaderboard (LB) scores were poor and unstable. Five days before the end of the competition, when I plotted the relationship between CV and LB scores again, I noticed that larger models tended to overfit. They achieved good CV scores but had poor LB scores. After adjusting the layers used, I found that only b3 showed high LB scores, although I suspected it might be overfitting. When I used it in the ensemble, it significantly improved the LB scores, so I decided to include it. This discrepancy may be due to the limited amount of training data. I should have also tried regularization techniques such as dropout, freezing, and other strategies, but I ran out of time. Considering these options might have potentially improved the performance. \n※ These are just my guesses.\n\n## 5.2 Potential Successes That Were Not Implemented\nPre-training using IR images: Although it improved the CV performance, it resulted in a decline in LB scores, so it was not implemented.\nIncluding EMNIST (external data) in the training dataset: While it improved the CV performance, it led to a deterioration in LB scores, so it was not implemented.\n\n\n# 6. Acknowledgments\nI could not have achieved these results on my own. I was greatly influenced by those who I have collaborated with in the past, and I am grateful for their contributions. I would also like to express my sincere gratitude to those who have shared their knowledge and insights through previous competitions. Thank you very much.\n\ntraining code : https://github.com/chumajin/kaggle-VCID\ninference code : https://www.kaggle.com/code/chumajin/vcid-6th-place-inference",
      "votes": null
    },
    {
      "id": "2303004",
      "postDate": "06/15/2023 02:45:08",
      "content": "<p>Congratulations on 6th place!<br>\nWe weren't able to use IR images well, so I thought it was an impressive solution.<br>\nDoes that mean that the IR image is put into the Training Dataset as it is and treated in the same way as the original image?</p>",
      "rawMarkdown": "Congratulations on 6th place!\nWe weren't able to use IR images well, so I thought it was an impressive solution.\nDoes that mean that the IR image is put into the Training Dataset as it is and treated in the same way as the original image?",
      "votes": null
    },
    {
      "id": "2303012",
      "postDate": "06/15/2023 02:59:46",
      "content": "<p><a href=\"https://www.kaggle.com/tattaka\" target=\"_blank\">@tattaka</a> Thank you very much! Yes, that's correct. I stacked the same image multiple times as per the number of in_channels, and I used the original labels without any modifications.</p>",
      "rawMarkdown": "tattaka Thank you very much! Yes, that's correct. I stacked the same image multiple times as per the number of in_channels, and I used the original labels without any modifications.",
      "votes": null
    },
    {
      "id": "2303042",
      "postDate": "06/15/2023 03:49:57",
      "content": "<p>Congratulations on your high place! I wonder why such sensitivity to image rotation? I thought CNN was invariant to image rotations. After all, we do not recognize characters (A, B, C) and strings of characters, but only pixels. Can someone clarify?</p>",
      "rawMarkdown": "Congratulations on your high place! I wonder why such sensitivity to image rotation? I thought CNN was invariant to image rotations. After all, we do not recognize characters (A, B, C) and strings of characters, but only pixels. Can someone clarify?",
      "votes": null
    },
    {
      "id": "2303063",
      "postDate": "06/15/2023 04:15:38",
      "content": "<p>Congratulations on the gold and 6th place!  And thank you for the nice writeup.</p>\n<p>We were worried about threshold calibration for different models.  We briefly considered thresholding at a specific percentile, like you did, but felt that was too risky if the amount of ink on the test set happened to differ.  Curious if you worried about that at all?</p>",
      "rawMarkdown": "Congratulations on the gold and 6th place!  And thank you for the nice writeup.\n\nWe were worried about threshold calibration for different models.  We briefly considered thresholding at a specific percentile, like you did, but felt that was too risky if the amount of ink on the test set happened to differ.  Curious if you worried about that at all?",
      "votes": null
    },
    {
      "id": "2303094",
      "postDate": "06/15/2023 04:56:04",
      "content": "<p>\"I found that only b3 showed high LB scores\"  Was this also the case for the private lb as well?  If so, do you think this shows that the private and public lb scores were very highly correlated such that overfitting to the public lb was the optimal strategy in this competition?  </p>",
      "rawMarkdown": "\"I found that only b3 showed high LB scores\"  Was this also the case for the private lb as well?  If so, do you think this shows that the private and public lb scores were very highly correlated such that overfitting to the public lb was the optimal strategy in this competition?",
      "votes": null
    },
    {
      "id": "2303134",
      "postDate": "06/15/2023 05:23:18",
      "content": "<p><a href=\"https://www.kaggle.com/sapr3s\" target=\"_blank\">@sapr3s</a> Thank you very much! This is just my opinion, but I believe that it is not so much about the rotation itself, but rather aligning the images in the same orientation as during training. The angle of the text during training may have an influence on the results.</p>",
      "rawMarkdown": "sapr3s Thank you very much! This is just my opinion, but I believe that it is not so much about the rotation itself, but rather aligning the images in the same orientation as during training. The angle of the text during training may have an influence on the results.",
      "votes": null
    },
    {
      "id": "2303139",
      "postDate": "06/15/2023 05:31:04",
      "content": "<p><a href=\"https://www.kaggle.com/socated\" target=\"_blank\">@socated</a> Congratulations on the 1st place! I'm looking forward to hearing about your solution. Personally, I didn't have such concerns because I didn't observe such tendencies in frag1 to frag3. Instead, I thought that it might be influenced by the density of the text. The graph shows the label density, where the x-axis indicates the ratio of pixels with label 1 to the total number of pixels, and the y-axis represents the optimal threshold when using percentiles. </p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4001300%2Faa985dc98f0e0d70f4cf73acdbb567ae%2Fsolution4.jpg?generation=1686807056839803&amp;alt=media\" alt=\"\"></p>\n<p>From the graph, I confirmed that the optimal threshold changes when the text density varies significantly, but I believed that it would not fluctuate significantly within the same fragment. (I anticipated that the threshold for the private dataset would be slightly lower as the text density tends to be higher compared to the public dataset.)</p>",
      "rawMarkdown": "socated Congratulations on the 1st place! I'm looking forward to hearing about your solution. Personally, I didn't have such concerns because I didn't observe such tendencies in frag1 to frag3. Instead, I thought that it might be influenced by the density of the text. The graph shows the label density, where the x-axis indicates the ratio of pixels with label 1 to the total number of pixels, and the y-axis represents the optimal threshold when using percentiles. \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4001300%2Faa985dc98f0e0d70f4cf73acdbb567ae%2Fsolution4.jpg?generation=1686807056839803&alt=media)\n\nFrom the graph, I confirmed that the optimal threshold changes when the text density varies significantly, but I believed that it would not fluctuate significantly within the same fragment. (I anticipated that the threshold for the private dataset would be slightly lower as the text density tends to be higher compared to the public dataset.)",
      "votes": null
    },
    {
      "id": "2303147",
      "postDate": "06/15/2023 05:43:20",
      "content": "<p><a href=\"https://www.kaggle.com/petersk20\" target=\"_blank\">@petersk20</a> I didn't anticipate it, but surprisingly, the single model of SegFormer b3 achieved the highest private LB score(0.66) among my submissions. I was not very confident in the single model due to overfitting concerns. Since single models tend to have larger variances due to the limited training data, aligning them with the public LB may not be ideal. However, I knew from conducting simulations by dividing fragment3 that ensembling would reduce the variance and establish a correlation between the public LB and private LB.</p>",
      "rawMarkdown": "petersk20 I didn't anticipate it, but surprisingly, the single model of SegFormer b3 achieved the highest private LB score(0.66) among my submissions. I was not very confident in the single model due to overfitting concerns. Since single models tend to have larger variances due to the limited training data, aligning them with the public LB may not be ideal. However, I knew from conducting simulations by dividing fragment3 that ensembling would reduce the variance and establish a correlation between the public LB and private LB.",
      "votes": null
    },
    {
      "id": "2303171",
      "postDate": "06/15/2023 06:07:52",
      "content": "<p>Congratulations  6th place<br>\nthanks for share detail explanation <a href=\"https://www.kaggle.com/chumajin\" target=\"_blank\">@chumajin</a> </p>",
      "rawMarkdown": "Congratulations  6th place\nthanks for share detail explanation @chumajin",
      "votes": null
    },
    {
      "id": "2303228",
      "postDate": "06/15/2023 07:17:29",
      "content": "<p>Congratulations! The concatenation of hidden test fragments A and B is a fantastic approach! <br>\nBut how do you know the bonding surface of each test fragment? Those are different from the public test fragments, which can be estimated through the shape of train fragment 1. </p>",
      "rawMarkdown": "Congratulations! The concatenation of hidden test fragments A and B is a fantastic approach! \nBut how do you know the bonding surface of each test fragment? Those are different from the public test fragments, which can be estimated through the shape of train fragment 1.",
      "votes": null
    },
    {
      "id": "2303232",
      "postDate": "06/15/2023 07:19:57",
      "content": "<p>Congratulations, that's great</p>",
      "rawMarkdown": "Congratulations, that's great",
      "votes": null
    },
    {
      "id": "2303274",
      "postDate": "06/15/2023 07:47:32",
      "content": "<p>May I ask if I use TTA to rotate the images by 90, 180, and 270 and concatenate them with the original image, which is the same as the clockwise rotation effect you mentioned? I think these two should be equivalent.</p>",
      "rawMarkdown": "May I ask if I use TTA to rotate the images by 90, 180, and 270 and concatenate them with the original image, which is the same as the clockwise rotation effect you mentioned? I think these two should be equivalent.",
      "votes": null
    },
    {
      "id": "2303367",
      "postDate": "06/15/2023 09:10:37",
      "content": "<p>Congratulations <a href=\"https://www.kaggle.com/chumajin\" target=\"_blank\">@chumajin</a> 🎉🎉🎉</p>",
      "rawMarkdown": "Congratulations @chumajin 🎉🎉🎉",
      "votes": null
    },
    {
      "id": "2303635",
      "postDate": "06/15/2023 11:48:59",
      "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F72b2a382e1e9bdc9ad24d2d747566842%2FSelection_999(2229).png?generation=1686829690948484&amp;alt=media\" alt=\"\"></p>\n<p>\"2.1 About test data and inference flow<br>\nBased on the brief LB probing and the information provided on the competition page, I had an idea of what the test data might look like. Here is an image that represents my understanding:\"</p>\n<p><a href=\"https://www.kaggle.com/chumajin\" target=\"_blank\">@chumajin</a> </p>\n<p>based on your idea, i make a submission using the image below. you are right!!!!</p>",
      "rawMarkdown": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F72b2a382e1e9bdc9ad24d2d747566842%2FSelection_999(2229).png?generation=1686829690948484&alt=media)\n\n\n\"2.1 About test data and inference flow\nBased on the brief LB probing and the information provided on the competition page, I had an idea of what the test data might look like. Here is an image that represents my understanding:\"\n\n@chumajin \n\nbased on your idea, i make a submission using the image below. you are right!!!!",
      "votes": null
    },
    {
      "id": "2303640",
      "postDate": "06/15/2023 11:51:43",
      "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fd03ecdfb13db02524c85065717d4d18c%2FSelection_999(2230).png?generation=1686829844102390&amp;alt=media\" alt=\"\"></p>\n<p>image from my post<br>\n<a href=\"https://www.kaggle.com/competitions/vesuvius-challenge-ink-detection/discussion/407972\" target=\"_blank\">https://www.kaggle.com/competitions/vesuvius-challenge-ink-detection/discussion/407972</a></p>\n<p>(you also get the number of lines and number of charcters correct)</p>",
      "rawMarkdown": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fd03ecdfb13db02524c85065717d4d18c%2FSelection_999(2230).png?generation=1686829844102390&alt=media)\n\nimage from my post\nhttps://www.kaggle.com/competitions/vesuvius-challenge-ink-detection/discussion/407972\n\n(you also get the number of lines and number of charcters correct)",
      "votes": null
    },
    {
      "id": "2303673",
      "postDate": "06/15/2023 12:08:38",
      "content": "<p>you have just submitted an image or you made rotation during predcition?</p>",
      "rawMarkdown": "you have just submitted an image or you made rotation during predcition?",
      "votes": null
    },
    {
      "id": "2303682",
      "postDate": "06/15/2023 12:12:10",
      "content": "<p>it depends on your precision/recall of your model.</p>\n<p>if you use TTA, the erros may also increase by 4x</p>",
      "rawMarkdown": "it depends on your precision/recall of your model.\n\nif you use TTA, the erros may also increase by 4x",
      "votes": null
    },
    {
      "id": "2303713",
      "postDate": "06/15/2023 12:29:37",
      "content": "<p>just submitted an image </p>",
      "rawMarkdown": "just submitted an image",
      "votes": null
    },
    {
      "id": "2303735",
      "postDate": "06/15/2023 12:46:02",
      "content": "<p>You make a really good point. Thank you!</p>",
      "rawMarkdown": "You make a really good point. Thank you!",
      "votes": null
    },
    {
      "id": "2303761",
      "postDate": "06/15/2023 13:09:02",
      "content": "<p>Congratulations on 6th place and your second solo gold medal!</p>\n<blockquote>\n  <p>the test data was rotated shown in other discussions</p>\n</blockquote>\n<p>It seems I missed this information.😨</p>\n<p>Moreover, I wasn't able to tackle tasks such as threshold adjustments.<br>\nThank you for sharing the nice solution!</p>",
      "rawMarkdown": "Congratulations on 6th place and your second solo gold medal!\n\n> the test data was rotated shown in other discussions\n\nIt seems I missed this information.😨\n\nMoreover, I wasn't able to tackle tasks such as threshold adjustments.\nThank you for sharing the nice solution!",
      "votes": null
    },
    {
      "id": "2303785",
      "postDate": "06/15/2023 13:25:09",
      "content": "<p><a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> Great confirmation method and result ! Thank you very much. When I saw this image in your post for the first time, I also imagined that this image could be test data ! </p>",
      "rawMarkdown": "hengck23 Great confirmation method and result ! Thank you very much. When I saw this image in your post for the first time, I also imagined that this image could be test data !",
      "votes": null
    },
    {
      "id": "2303795",
      "postDate": "06/15/2023 13:30:33",
      "content": "<p><a href=\"https://www.kaggle.com/wangxuc\" target=\"_blank\">@wangxuc</a> Yes, perhaps adding rotation in aug could be effective. However, if rotation is not included, I think adding noise in TTA would be necessary.</p>",
      "rawMarkdown": "wangxuc Yes, perhaps adding rotation in aug could be effective. However, if rotation is not included, I think adding noise in TTA would be necessary.",
      "votes": null
    },
    {
      "id": "2303801",
      "postDate": "06/15/2023 13:34:03",
      "content": "<p>oh, my phd friend just tell me that you can download orginal images used in arvix pdf paper from:</p>\n<p>\" gzipped TeX, DVI, PostScript or HTML (.gz, .dvi.gz, .ps.gz or .html.gz) file depending on submission format. [ Download source ]\"</p>\n<p>… and i find a metric results csv file … and i think kaggler results exceeded the paper results</p>\n<p><a href=\"https://arxiv.org/format/2304.02084\" target=\"_blank\">https://arxiv.org/format/2304.02084</a></p>",
      "rawMarkdown": "oh, my phd friend just tell me that you can download orginal images used in arvix pdf paper from:\n\n\" gzipped TeX, DVI, PostScript or HTML (.gz, .dvi.gz, .ps.gz or .html.gz) file depending on submission format. [ Download source ]\"\n\n... and i find a metric results csv file ... and i think kaggler results exceeded the paper results\n\nhttps://arxiv.org/format/2304.02084",
      "votes": null
    },
    {
      "id": "2303804",
      "postDate": "06/15/2023 13:34:25",
      "content": "<p><a href=\"https://www.kaggle.com/riow1983\" target=\"_blank\">@riow1983</a> Thank you for comment !! Yes. Considering that the dataset before submitting is divided into two flags and the fragments are rotated, I can imagine the image mentioned above. Additionally, using the Kaggle technique called LB proving, for example, if the edges of masks for flag A and B have the same pixels, you can set the score to zero to confirm.</p>",
      "rawMarkdown": "riow1983 Thank you for comment !! Yes. Considering that the dataset before submitting is divided into two flags and the fragments are rotated, I can imagine the image mentioned above. Additionally, using the Kaggle technique called LB proving, for example, if the edges of masks for flag A and B have the same pixels, you can set the score to zero to confirm.",
      "votes": null
    },
    {
      "id": "2303831",
      "postDate": "06/15/2023 13:55:38",
      "content": "<p>wow <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> What a find! I should have taken more note of the discussion about test fragments being rotated.</p>",
      "rawMarkdown": "wow @hengck23 What a find! I should have taken more note of the discussion about test fragments being rotated.",
      "votes": null
    },
    {
      "id": "2304024",
      "postDate": "06/15/2023 16:15:45",
      "content": "<p>congrats on your 6th place</p>",
      "rawMarkdown": "congrats on your 6th place",
      "votes": null
    },
    {
      "id": "2304200",
      "postDate": "06/15/2023 19:13:35",
      "content": "<p>Hi! Congratulations on 6th place and thank you for nice write-up! Out of curiosity, could you post  the predictions on your validation folds? That would be nice to see the visual results for different models</p>",
      "rawMarkdown": "Hi! Congratulations on 6th place and thank you for nice write-up! Out of curiosity, could you post  the predictions on your validation folds? That would be nice to see the visual results for different models",
      "votes": null
    },
    {
      "id": "2304374",
      "postDate": "06/16/2023 00:31:38",
      "content": "<p>Thank you!, I understand.👍</p>",
      "rawMarkdown": "Thank you!, I understand.👍",
      "votes": null
    },
    {
      "id": "2304532",
      "postDate": "06/16/2023 04:25:18",
      "content": "<p><a href=\"https://www.kaggle.com/mkotyushev\" target=\"_blank\">@mkotyushev</a> Thank you for comment! I added the visualization of predictions (see ## 4.1).</p>",
      "rawMarkdown": "mkotyushev Thank you for comment! I added the visualization of predictions (see ## 4.1).",
      "votes": null
    },
    {
      "id": "2304759",
      "postDate": "06/16/2023 08:05:20",
      "content": "<p>Thanks so much for the detailed write-up and congrats on your 6th place. </p>\n<p>On the rotation of the test-set and your inference rotation back and forth trick: did you add 90degree rotations to the train augmentations? I wonder whether -even with adding the 90degree rotations- the inference rotation trick you mention is still improving the score?</p>\n<p>And a second question: how did you incorporate IR images into the training? I incorporated it by having two heads on the same model: one head predicted the mask (loss: Dice + BCE) and the other head predicted the IR image (loss: MSE). Then I added the two losses and take the gradient. Is this similar to how you did it as well? </p>\n<p>Last but not least, you mention the pre-training on IR images: how does this work? I'm not familiar with pre-training and would love to understand it better.</p>",
      "rawMarkdown": "Thanks so much for the detailed write-up and congrats on your 6th place. \n\nOn the rotation of the test-set and your inference rotation back and forth trick: did you add 90degree rotations to the train augmentations? I wonder whether -even with adding the 90degree rotations- the inference rotation trick you mention is still improving the score?\n\nAnd a second question: how did you incorporate IR images into the training? I incorporated it by having two heads on the same model: one head predicted the mask (loss: Dice + BCE) and the other head predicted the IR image (loss: MSE). Then I added the two losses and take the gradient. Is this similar to how you did it as well? \n\nLast but not least, you mention the pre-training on IR images: how does this work? I'm not familiar with pre-training and would love to understand it better.",
      "votes": null
    },
    {
      "id": "2305222",
      "postDate": "06/16/2023 14:15:06",
      "content": "<p><a href=\"https://www.kaggle.com/lucasvw\" target=\"_blank\">@lucasvw</a> Thank you for comment ! </p>\n<p>1) No, I did not intentionally include 90-degree rotations in the augmentation. I believed that it could introduce noise in this approach, as the direction of the characters was already known.</p>\n<p>2) The usage of IR images was not that complex; I simply stacked the same images and used them as training images. I also used the labels in the same positions as before.</p>\n<p>3) Regarding pretraining, the pretrained model I'm using is not specialized in characters. Therefore, I thought that pretraining in the domain of characters and then further fine-tuning could enhance its sensitivity.</p>",
      "rawMarkdown": "lucasvw Thank you for comment ! \n\n1) No, I did not intentionally include 90-degree rotations in the augmentation. I believed that it could introduce noise in this approach, as the direction of the characters was already known.\n\n2) The usage of IR images was not that complex; I simply stacked the same images and used them as training images. I also used the labels in the same positions as before.\n\n3) Regarding pretraining, the pretrained model I'm using is not specialized in characters. Therefore, I thought that pretraining in the domain of characters and then further fine-tuning could enhance its sensitivity.",
      "votes": null
    },
    {
      "id": "2305563",
      "postDate": "06/16/2023 19:11:09",
      "content": "<p>Congrats brother </p>",
      "rawMarkdown": "Congrats brother",
      "votes": null
    },
    {
      "id": "2305583",
      "postDate": "06/16/2023 19:27:32",
      "content": "<p>Congratulations <a href=\"https://www.kaggle.com/chumajin\" target=\"_blank\">@chumajin</a> !</p>\n<p>Interesting how yo have concatenated fragments A and B on the test set, in order to make predictions as one continuous sequence.</p>\n<p>Keep up the great work🔥💪</p>",
      "rawMarkdown": "Congratulations @chumajin !\n\nInteresting how yo have concatenated fragments A and B on the test set, in order to make predictions as one continuous sequence.\n\nKeep up the great work🔥💪",
      "votes": null
    },
    {
      "id": "2306699",
      "postDate": "06/17/2023 13:53:13",
      "content": "<p><a href=\"https://www.youtube.com/watch?v=tXGcs02y9r4\" target=\"_blank\">Percentiles - How to calculate Percentiles, Quartiles, … - YouTube</a></p>",
      "rawMarkdown": "[Percentiles - How to calculate Percentiles, Quartiles, ... - YouTube](https://www.youtube.com/watch?v=tXGcs02y9r4)",
      "votes": null
    },
    {
      "id": "2306843",
      "postDate": "06/17/2023 16:26:36",
      "content": "<p>Congrats on another solo gold!</p>",
      "rawMarkdown": "Congrats on another solo gold!",
      "votes": null
    },
    {
      "id": "2307050",
      "postDate": "06/17/2023 19:08:52",
      "content": "<p>Congrats mate! you are on fire! :)</p>",
      "rawMarkdown": "Congrats mate! you are on fire! :)",
      "votes": null
    },
    {
      "id": "2307119",
      "postDate": "06/17/2023 20:41:10",
      "content": "<p><a href=\"https://www.kaggle.com/imeintanis\" target=\"_blank\">@imeintanis</a> Thank you my friend!! </p>",
      "rawMarkdown": "imeintanis Thank you my friend!!",
      "votes": null
    },
    {
      "id": "2307120",
      "postDate": "06/17/2023 20:41:38",
      "content": "<p><a href=\"https://www.kaggle.com/ohanegby\" target=\"_blank\">@ohanegby</a> Thank you very much!!</p>",
      "rawMarkdown": "ohanegby Thank you very much!!",
      "votes": null
    },
    {
      "id": "2307850",
      "postDate": "06/18/2023 13:09:39",
      "content": "<p>Your approach and detailed explanation provide valuable insights into your solution. It's impressive how you utilized a combination of models and techniques to achieve your results, such as ensemble methods, rotating images during inference, and incorporating IR images into the training data. It's also interesting to see how you analyzed the relationship between CV and LB scores and made adjustments accordingly. Overall, your strategy and experimentation demonstrate a strong understanding of the problem and effective problem-solving skills. Congratulations on your success!</p>",
      "rawMarkdown": "Your approach and detailed explanation provide valuable insights into your solution. It's impressive how you utilized a combination of models and techniques to achieve your results, such as ensemble methods, rotating images during inference, and incorporating IR images into the training data. It's also interesting to see how you analyzed the relationship between CV and LB scores and made adjustments accordingly. Overall, your strategy and experimentation demonstrate a strong understanding of the problem and effective problem-solving skills. Congratulations on your success!",
      "votes": null
    },
    {
      "id": "2308047",
      "postDate": "06/18/2023 16:15:39",
      "content": "<p>Congrats man you are superb</p>",
      "rawMarkdown": "Congrats man you are superb",
      "votes": null
    },
    {
      "id": "2319606",
      "postDate": "06/27/2023 08:32:07",
      "content": "<p>I have released the training code and inference code publicly!<br>\ntraining code : <a href=\"https://github.com/chumajin/kaggle-VCID\" target=\"_blank\">https://github.com/chumajin/kaggle-VCID</a><br>\ninference code(kaggle) : <a href=\"https://www.kaggle.com/code/chumajin/vcid-6th-place-inference\" target=\"_blank\">https://www.kaggle.com/code/chumajin/vcid-6th-place-inference</a></p>",
      "rawMarkdown": "I have released the training code and inference code publicly!\ntraining code : https://github.com/chumajin/kaggle-VCID\ninference code(kaggle) : https://www.kaggle.com/code/chumajin/vcid-6th-place-inference",
      "votes": null
    },
    {
      "id": "2341363",
      "postDate": "07/12/2023 04:24:24",
      "content": "<p>Thanks for the information :)  would do me a favor.</p>",
      "rawMarkdown": "Thanks for the information :)  would do me a favor.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2303004,
      "author_name": "tattaka",
      "author_url": "",
      "post_date": "06/15/2023 02:45:08",
      "content": "<p>Congratulations on 6th place!<br>\nWe weren't able to use IR images well, so I thought it was an impressive solution.<br>\nDoes that mean that the IR image is put into the Training Dataset as it is and treated in the same way as the original image?</p>",
      "votes": null,
      "replies": [
        {
          "id": 2303012,
          "author_name": "chumajin",
          "author_url": "",
          "post_date": "06/15/2023 02:59:46",
          "content": "<p><a href=\"https://www.kaggle.com/tattaka\" target=\"_blank\">@tattaka</a> Thank you very much! Yes, that's correct. I stacked the same image multiple times as per the number of in_channels, and I used the original labels without any modifications.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2303042,
      "author_name": "sapr3s",
      "author_url": "",
      "post_date": "06/15/2023 03:49:57",
      "content": "<p>Congratulations on your high place! I wonder why such sensitivity to image rotation? I thought CNN was invariant to image rotations. After all, we do not recognize characters (A, B, C) and strings of characters, but only pixels. Can someone clarify?</p>",
      "votes": null,
      "replies": [
        {
          "id": 2303134,
          "author_name": "chumajin",
          "author_url": "",
          "post_date": "06/15/2023 05:23:18",
          "content": "<p><a href=\"https://www.kaggle.com/sapr3s\" target=\"_blank\">@sapr3s</a> Thank you very much! This is just my opinion, but I believe that it is not so much about the rotation itself, but rather aligning the images in the same orientation as during training. The angle of the text during training may have an influence on the results.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2303063,
      "author_name": "socated",
      "author_url": "",
      "post_date": "06/15/2023 04:15:38",
      "content": "<p>Congratulations on the gold and 6th place!  And thank you for the nice writeup.</p>\n<p>We were worried about threshold calibration for different models.  We briefly considered thresholding at a specific percentile, like you did, but felt that was too risky if the amount of ink on the test set happened to differ.  Curious if you worried about that at all?</p>",
      "votes": null,
      "replies": [
        {
          "id": 2303139,
          "author_name": "chumajin",
          "author_url": "",
          "post_date": "06/15/2023 05:31:04",
          "content": "<p><a href=\"https://www.kaggle.com/socated\" target=\"_blank\">@socated</a> Congratulations on the 1st place! I'm looking forward to hearing about your solution. Personally, I didn't have such concerns because I didn't observe such tendencies in frag1 to frag3. Instead, I thought that it might be influenced by the density of the text. The graph shows the label density, where the x-axis indicates the ratio of pixels with label 1 to the total number of pixels, and the y-axis represents the optimal threshold when using percentiles. </p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4001300%2Faa985dc98f0e0d70f4cf73acdbb567ae%2Fsolution4.jpg?generation=1686807056839803&amp;alt=media\" alt=\"\"></p>\n<p>From the graph, I confirmed that the optimal threshold changes when the text density varies significantly, but I believed that it would not fluctuate significantly within the same fragment. (I anticipated that the threshold for the private dataset would be slightly lower as the text density tends to be higher compared to the public dataset.)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2303094,
      "author_name": "petersk20",
      "author_url": "",
      "post_date": "06/15/2023 04:56:04",
      "content": "<p>\"I found that only b3 showed high LB scores\"  Was this also the case for the private lb as well?  If so, do you think this shows that the private and public lb scores were very highly correlated such that overfitting to the public lb was the optimal strategy in this competition?  </p>",
      "votes": null,
      "replies": [
        {
          "id": 2303147,
          "author_name": "chumajin",
          "author_url": "",
          "post_date": "06/15/2023 05:43:20",
          "content": "<p><a href=\"https://www.kaggle.com/petersk20\" target=\"_blank\">@petersk20</a> I didn't anticipate it, but surprisingly, the single model of SegFormer b3 achieved the highest private LB score(0.66) among my submissions. I was not very confident in the single model due to overfitting concerns. Since single models tend to have larger variances due to the limited training data, aligning them with the public LB may not be ideal. However, I knew from conducting simulations by dividing fragment3 that ensembling would reduce the variance and establish a correlation between the public LB and private LB.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2303171,
      "author_name": "pritomsh",
      "author_url": "",
      "post_date": "06/15/2023 06:07:52",
      "content": "<p>Congratulations  6th place<br>\nthanks for share detail explanation <a href=\"https://www.kaggle.com/chumajin\" target=\"_blank\">@chumajin</a> </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2303228,
      "author_name": "riow1983",
      "author_url": "",
      "post_date": "06/15/2023 07:17:29",
      "content": "<p>Congratulations! The concatenation of hidden test fragments A and B is a fantastic approach! <br>\nBut how do you know the bonding surface of each test fragment? Those are different from the public test fragments, which can be estimated through the shape of train fragment 1. </p>",
      "votes": null,
      "replies": [
        {
          "id": 2303804,
          "author_name": "chumajin",
          "author_url": "",
          "post_date": "06/15/2023 13:34:25",
          "content": "<p><a href=\"https://www.kaggle.com/riow1983\" target=\"_blank\">@riow1983</a> Thank you for comment !! Yes. Considering that the dataset before submitting is divided into two flags and the fragments are rotated, I can imagine the image mentioned above. Additionally, using the Kaggle technique called LB proving, for example, if the edges of masks for flag A and B have the same pixels, you can set the score to zero to confirm.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2303232,
      "author_name": "afarfrgvgrtb",
      "author_url": "",
      "post_date": "06/15/2023 07:19:57",
      "content": "<p>Congratulations, that's great</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2303274,
      "author_name": "wangxuc",
      "author_url": "",
      "post_date": "06/15/2023 07:47:32",
      "content": "<p>May I ask if I use TTA to rotate the images by 90, 180, and 270 and concatenate them with the original image, which is the same as the clockwise rotation effect you mentioned? I think these two should be equivalent.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2303682,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "06/15/2023 12:12:10",
          "content": "<p>it depends on your precision/recall of your model.</p>\n<p>if you use TTA, the erros may also increase by 4x</p>",
          "votes": null,
          "replies": [
            {
              "id": 2303735,
              "author_name": "wangxuc",
              "author_url": "",
              "post_date": "06/15/2023 12:46:02",
              "content": "<p>You make a really good point. Thank you!</p>",
              "votes": null,
              "replies": [
                {
                  "id": 2303795,
                  "author_name": "chumajin",
                  "author_url": "",
                  "post_date": "06/15/2023 13:30:33",
                  "content": "<p><a href=\"https://www.kaggle.com/wangxuc\" target=\"_blank\">@wangxuc</a> Yes, perhaps adding rotation in aug could be effective. However, if rotation is not included, I think adding noise in TTA would be necessary.</p>",
                  "votes": null,
                  "replies": [
                    {
                      "id": 2304374,
                      "author_name": "wangxuc",
                      "author_url": "",
                      "post_date": "06/16/2023 00:31:38",
                      "content": "<p>Thank you!, I understand.👍</p>",
                      "votes": null,
                      "replies": []
                    }
                  ]
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 2303367,
      "author_name": "conjuring92",
      "author_url": "",
      "post_date": "06/15/2023 09:10:37",
      "content": "<p>Congratulations <a href=\"https://www.kaggle.com/chumajin\" target=\"_blank\">@chumajin</a> 🎉🎉🎉</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2303635,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "06/15/2023 11:48:59",
      "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F72b2a382e1e9bdc9ad24d2d747566842%2FSelection_999(2229).png?generation=1686829690948484&amp;alt=media\" alt=\"\"></p>\n<p>\"2.1 About test data and inference flow<br>\nBased on the brief LB probing and the information provided on the competition page, I had an idea of what the test data might look like. Here is an image that represents my understanding:\"</p>\n<p><a href=\"https://www.kaggle.com/chumajin\" target=\"_blank\">@chumajin</a> </p>\n<p>based on your idea, i make a submission using the image below. you are right!!!!</p>",
      "votes": null,
      "replies": [
        {
          "id": 2303640,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "06/15/2023 11:51:43",
          "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fd03ecdfb13db02524c85065717d4d18c%2FSelection_999(2230).png?generation=1686829844102390&amp;alt=media\" alt=\"\"></p>\n<p>image from my post<br>\n<a href=\"https://www.kaggle.com/competitions/vesuvius-challenge-ink-detection/discussion/407972\" target=\"_blank\">https://www.kaggle.com/competitions/vesuvius-challenge-ink-detection/discussion/407972</a></p>\n<p>(you also get the number of lines and number of charcters correct)</p>",
          "votes": null,
          "replies": [
            {
              "id": 2303673,
              "author_name": "maksimovka",
              "author_url": "",
              "post_date": "06/15/2023 12:08:38",
              "content": "<p>you have just submitted an image or you made rotation during predcition?</p>",
              "votes": null,
              "replies": [
                {
                  "id": 2303713,
                  "author_name": "hengck23",
                  "author_url": "",
                  "post_date": "06/15/2023 12:29:37",
                  "content": "<p>just submitted an image </p>",
                  "votes": null,
                  "replies": [
                    {
                      "id": 2303785,
                      "author_name": "chumajin",
                      "author_url": "",
                      "post_date": "06/15/2023 13:25:09",
                      "content": "<p><a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> Great confirmation method and result ! Thank you very much. When I saw this image in your post for the first time, I also imagined that this image could be test data ! </p>",
                      "votes": null,
                      "replies": [
                        {
                          "id": 2303801,
                          "author_name": "hengck23",
                          "author_url": "",
                          "post_date": "06/15/2023 13:34:03",
                          "content": "<p>oh, my phd friend just tell me that you can download orginal images used in arvix pdf paper from:</p>\n<p>\" gzipped TeX, DVI, PostScript or HTML (.gz, .dvi.gz, .ps.gz or .html.gz) file depending on submission format. [ Download source ]\"</p>\n<p>… and i find a metric results csv file … and i think kaggler results exceeded the paper results</p>\n<p><a href=\"https://arxiv.org/format/2304.02084\" target=\"_blank\">https://arxiv.org/format/2304.02084</a></p>",
                          "votes": null,
                          "replies": []
                        }
                      ]
                    }
                  ]
                }
              ]
            },
            {
              "id": 2303831,
              "author_name": "adrianshedley",
              "author_url": "",
              "post_date": "06/15/2023 13:55:38",
              "content": "<p>wow <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> What a find! I should have taken more note of the discussion about test fragments being rotated.</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2303761,
      "author_name": "takamichitoda",
      "author_url": "",
      "post_date": "06/15/2023 13:09:02",
      "content": "<p>Congratulations on 6th place and your second solo gold medal!</p>\n<blockquote>\n  <p>the test data was rotated shown in other discussions</p>\n</blockquote>\n<p>It seems I missed this information.😨</p>\n<p>Moreover, I wasn't able to tackle tasks such as threshold adjustments.<br>\nThank you for sharing the nice solution!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2304024,
      "author_name": "arunodhayan",
      "author_url": "",
      "post_date": "06/15/2023 16:15:45",
      "content": "<p>congrats on your 6th place</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2304200,
      "author_name": "mkotyushev",
      "author_url": "",
      "post_date": "06/15/2023 19:13:35",
      "content": "<p>Hi! Congratulations on 6th place and thank you for nice write-up! Out of curiosity, could you post  the predictions on your validation folds? That would be nice to see the visual results for different models</p>",
      "votes": null,
      "replies": [
        {
          "id": 2304532,
          "author_name": "chumajin",
          "author_url": "",
          "post_date": "06/16/2023 04:25:18",
          "content": "<p><a href=\"https://www.kaggle.com/mkotyushev\" target=\"_blank\">@mkotyushev</a> Thank you for comment! I added the visualization of predictions (see ## 4.1).</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2304759,
      "author_name": "lucasvw",
      "author_url": "",
      "post_date": "06/16/2023 08:05:20",
      "content": "<p>Thanks so much for the detailed write-up and congrats on your 6th place. </p>\n<p>On the rotation of the test-set and your inference rotation back and forth trick: did you add 90degree rotations to the train augmentations? I wonder whether -even with adding the 90degree rotations- the inference rotation trick you mention is still improving the score?</p>\n<p>And a second question: how did you incorporate IR images into the training? I incorporated it by having two heads on the same model: one head predicted the mask (loss: Dice + BCE) and the other head predicted the IR image (loss: MSE). Then I added the two losses and take the gradient. Is this similar to how you did it as well? </p>\n<p>Last but not least, you mention the pre-training on IR images: how does this work? I'm not familiar with pre-training and would love to understand it better.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2305222,
          "author_name": "chumajin",
          "author_url": "",
          "post_date": "06/16/2023 14:15:06",
          "content": "<p><a href=\"https://www.kaggle.com/lucasvw\" target=\"_blank\">@lucasvw</a> Thank you for comment ! </p>\n<p>1) No, I did not intentionally include 90-degree rotations in the augmentation. I believed that it could introduce noise in this approach, as the direction of the characters was already known.</p>\n<p>2) The usage of IR images was not that complex; I simply stacked the same images and used them as training images. I also used the labels in the same positions as before.</p>\n<p>3) Regarding pretraining, the pretrained model I'm using is not specialized in characters. Therefore, I thought that pretraining in the domain of characters and then further fine-tuning could enhance its sensitivity.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2305563,
      "author_name": "manjeetsinghrawat",
      "author_url": "",
      "post_date": "06/16/2023 19:11:09",
      "content": "<p>Congrats brother </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2305583,
      "author_name": "vladiluzjr",
      "author_url": "",
      "post_date": "06/16/2023 19:27:32",
      "content": "<p>Congratulations <a href=\"https://www.kaggle.com/chumajin\" target=\"_blank\">@chumajin</a> !</p>\n<p>Interesting how yo have concatenated fragments A and B on the test set, in order to make predictions as one continuous sequence.</p>\n<p>Keep up the great work🔥💪</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2306699,
      "author_name": "dgi1995",
      "author_url": "",
      "post_date": "06/17/2023 13:53:13",
      "content": "<p><a href=\"https://www.youtube.com/watch?v=tXGcs02y9r4\" target=\"_blank\">Percentiles - How to calculate Percentiles, Quartiles, … - YouTube</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2306843,
      "author_name": "ohanegby",
      "author_url": "",
      "post_date": "06/17/2023 16:26:36",
      "content": "<p>Congrats on another solo gold!</p>",
      "votes": null,
      "replies": [
        {
          "id": 2307120,
          "author_name": "chumajin",
          "author_url": "",
          "post_date": "06/17/2023 20:41:38",
          "content": "<p><a href=\"https://www.kaggle.com/ohanegby\" target=\"_blank\">@ohanegby</a> Thank you very much!!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2307050,
      "author_name": "imeintanis",
      "author_url": "",
      "post_date": "06/17/2023 19:08:52",
      "content": "<p>Congrats mate! you are on fire! :)</p>",
      "votes": null,
      "replies": [
        {
          "id": 2307119,
          "author_name": "chumajin",
          "author_url": "",
          "post_date": "06/17/2023 20:41:10",
          "content": "<p><a href=\"https://www.kaggle.com/imeintanis\" target=\"_blank\">@imeintanis</a> Thank you my friend!! </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2307850,
      "author_name": "poojach7611",
      "author_url": "",
      "post_date": "06/18/2023 13:09:39",
      "content": "<p>Your approach and detailed explanation provide valuable insights into your solution. It's impressive how you utilized a combination of models and techniques to achieve your results, such as ensemble methods, rotating images during inference, and incorporating IR images into the training data. It's also interesting to see how you analyzed the relationship between CV and LB scores and made adjustments accordingly. Overall, your strategy and experimentation demonstrate a strong understanding of the problem and effective problem-solving skills. Congratulations on your success!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2308047,
      "author_name": "shareefmhs",
      "author_url": "",
      "post_date": "06/18/2023 16:15:39",
      "content": "<p>Congrats man you are superb</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2319606,
      "author_name": "chumajin",
      "author_url": "",
      "post_date": "06/27/2023 08:32:07",
      "content": "<p>I have released the training code and inference code publicly!<br>\ntraining code : <a href=\"https://github.com/chumajin/kaggle-VCID\" target=\"_blank\">https://github.com/chumajin/kaggle-VCID</a><br>\ninference code(kaggle) : <a href=\"https://www.kaggle.com/code/chumajin/vcid-6th-place-inference\" target=\"_blank\">https://www.kaggle.com/code/chumajin/vcid-6th-place-inference</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2341363,
      "author_name": "ouyangqianhe",
      "author_url": "",
      "post_date": "07/12/2023 04:24:24",
      "content": "<p>Thanks for the information :)  would do me a favor.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2302987": "Thank you very much for organizing such an interesting competition. I am greatly thankful to the hosts and the Kaggle staff.\n\nContinuing from the previous competition, I am delighted to have won solo gold medal again, with a total of four medals (Table 1, NLP × 2, CV × 1). Additionally, it was my first time attempting the segmentation task, and I began with @tanakar excellent notebook [here](https://www.kaggle.com/code/tanakar/2-5d-segmentaion-baseline-training/comments). I am truly grateful for that.\n\n\n\n# 1. Summary\nMy approach involved an ensemble of EfficientNet and SegFormer models. I believed that the test data was rotated shown in other discussions, so I rotated the images during inference, which resulted in a significant boost at the beginning (Public LB 0.58 → 0.74). Additionally, my originality came from incorporating IR images into the training data, which gave me a CV score increase of 0.01 and an LB score increase of 0.01. I will now explain the details below.\n\n# 2. Inference\n## 2.1 About test data and inference flow\nBased on the brief LB probing and the information provided on the competition page, I had an idea of what the test data might look like. Here is an image that represents my understanding:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4001300%2Fb9c4e129fe3285876db0a7135f704edb%2Fsolution1.jpg?generation=1686795321876414&alt=media)\n\nTherefore, I structured my inference code in the Kaggle notebook as follows:\n\n1. Concatenate(axis=1) fragments A and B.\n2. Rotate the image clockwise.\n3. Perform inference (original + h flip TTA).\n4. Rotate the prediction countor-clockwise back to its original position.\n5. Cut and encode each fragment A and B respectively.\n\n\n\nOf course, to reduce inference time, I skipped the inference for areas where the mask value was 0. Furthermore, instead of inferring fragment A and B separately, I concatenated them. This not only eliminated the 0 padding at the boundary between A and B but also allowed for continuous inference of the initial part of fragment B as a contiguous sequence. These led to a significant boost in my LB score (EfficientNet B4: 0.58 → 0.74)\n\n\n## 2.2 Threshold\nI believe many of you experienced the instability of the signal values. Therefore, I used the following function to rank the entire image and calculate percentiles. Then, by applying a threshold, I obtained a stable threshold value. This approach proved helpful not only during inference but also during ensemble processes. 2nd place solution also used the same way [here](https://www.kaggle.com/competitions/vesuvius-challenge-ink-detection/discussion/417255)\n\n\n~~~\ndef get_percentile(array):\n     org_shape = array.shape\n     array = array.reshape(-1)\n     array = np.arange(len(array))[array.argsort().argsort()]\n     array = array / array.max()\n     array = array.reshape(org_shape)\n     return array\n~~~\n\nFor the Public LB, the optimal threshold was found to be 0.96. However, using fragment 3, I conducted a simulation to observe the correlation between the partially optimal threshold (around 10%) and the threshold for the remaining 90%. As a result, I noticed that the threshold was overfitting for the 10% portion (likely reducing noise), while for the remaining 90%, it was better to slightly lower the threshold below the optimal value (aiming for clearer extraction of text). In fact, when comparing the same model, a threshold of 0.95 performed slightly better for the private LB(but less than 0.01). For the final submission, I used different models: sub1 with a threshold of 0.96 and sub2 with a threshold of 0.95.\n\n\n\n# 3. Training\nThe following is an overview of the training process. Similar to inference, I created three sets of data and took their averages. It should be noted that SegFormer differs from CNN as it can only utilize 3 channels. As mentioned earlier, incorporating IR images resulted in improvements in both CV and LB scores.\n\n\n\n## 3.1 CNN + Unet\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4001300%2Fa32d067916dee939f3301b72d8daf9f8%2Fsolution2.jpg?generation=1686795342236188&alt=media)\n\n## 3.2 SegFormer\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4001300%2Fba6b7cdc7012775d09930b49179e0317%2Fsolution3.jpg?generation=1686795355254555&alt=media)\n\n## 3.3 Fine-tuned Parameters\n* Stride: image size // 4\n* Optimizer: Adam\n* Epochs: 20\n* Early stopping: 4\n* Scheduler: get_cosine_schedule_with_warmup (from transformers)\n* Warm-up: 0.1\n* Gradient norm: 10\n* Loss function: SoftBCEWithLogitsLoss (segmentation model in PyTorch)\n* Training excludes areas with a mask value of 0.\n* TTA: Horizontal flip\n\n## 3.4 Cross Validation\nFor submission1, I used a 7kfold cross-validation, and for submission2, I used a 10kfold cross-validation. Increasing the value of k-fold resulted in improvements in both CV and LB scores. I recall that increasing from 5-fold to 7-fold led to an improvement of approximately 0.1 in the Public LB score.\n\n# 4 Final result\n\nEnsemble was all mean value of predictions.\n\nsub1 : th 0.96, cv 0.740, public LB 0.811570, private LB 0.661339\n\n| model                     | image size | kfold         | cv     | public LB | private LB |\n|---------------------------|------------|---------------|--------|-----------|------------|\n| efficientnet_b7_ns        | 608        | 7 + fulltrain | 0.712  | 0.80      | 0.64       |\n| efficientnet_b6_ns        | 544        | 7 + fulltrain | 0.702  | 0.79      | 0.64       |\n| efficientnetv2_l_in21ft1k | 480        | 7             | 0.707  | 0.79      | 0.65       |\n| tf_efficientnet_b8        | 672        | 7             | 0.716  | 0.79      | 0.64       |\n| segformer b3              | 1024       | 7             | 0.738  | 0.78      | 0.66       |\n\n\nsub2 : th 0.95,cv 0.746 , public LB 0.799563, private LB 0.654812\n\n| model                     | image size | kfold | cv     | public LB | private LB |\n|---------------------------|------------|-------|--------|-----------|------------|\n| efficientnet_b7_ns        | 608        | 10    | 0.722  | 0.80      | 0.65       |\n| efficientnet_b6_ns        | 544        | 10    | 0.720  | 0.79      | 0.63       |\n| efficientnetv2_l_in21ft1k | 480        | 10    | 0.717  | 0.79      | 0.65       |\n| segformer b3              | 1024       | 7     | 0.738  | 0.78      | 0.66       |\n\n## 4.1 Visualization of predictions\n\nThe following images visualize the predictions of submission1.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4001300%2F985af9f03d54b031cc820f29b63bc891%2Fpredictions.jpg?generation=1686889411081867&alt=media)\n\n# 5. My understanding\n## 5.1 Not working well\n\nI tried several models such as ConvNext, Mask2Former, Swin Transformer + PSPNet, BeiT, and many others, but their effectiveness was not satisfactory for cv and lb. EfficientNet and mobilevit performed well and stable in this competition for me. I also experimented with SegFormer using various versions from b1 to b5. Although it showed good performance in cross-validation (CV), the leaderboard (LB) scores were poor and unstable. Five days before the end of the competition, when I plotted the relationship between CV and LB scores again, I noticed that larger models tended to overfit. They achieved good CV scores but had poor LB scores. After adjusting the layers used, I found that only b3 showed high LB scores, although I suspected it might be overfitting. When I used it in the ensemble, it significantly improved the LB scores, so I decided to include it. This discrepancy may be due to the limited amount of training data. I should have also tried regularization techniques such as dropout, freezing, and other strategies, but I ran out of time. Considering these options might have potentially improved the performance. \n※ These are just my guesses.\n\n## 5.2 Potential Successes That Were Not Implemented\nPre-training using IR images: Although it improved the CV performance, it resulted in a decline in LB scores, so it was not implemented.\nIncluding EMNIST (external data) in the training dataset: While it improved the CV performance, it led to a deterioration in LB scores, so it was not implemented.\n\n\n# 6. Acknowledgments\nI could not have achieved these results on my own. I was greatly influenced by those who I have collaborated with in the past, and I am grateful for their contributions. I would also like to express my sincere gratitude to those who have shared their knowledge and insights through previous competitions. Thank you very much.\n\ntraining code : https://github.com/chumajin/kaggle-VCID\ninference code : https://www.kaggle.com/code/chumajin/vcid-6th-place-inference",
    "2303004": "Congratulations on 6th place!\nWe weren't able to use IR images well, so I thought it was an impressive solution.\nDoes that mean that the IR image is put into the Training Dataset as it is and treated in the same way as the original image?",
    "2303012": "tattaka Thank you very much! Yes, that's correct. I stacked the same image multiple times as per the number of in_channels, and I used the original labels without any modifications.",
    "2303042": "Congratulations on your high place! I wonder why such sensitivity to image rotation? I thought CNN was invariant to image rotations. After all, we do not recognize characters (A, B, C) and strings of characters, but only pixels. Can someone clarify?",
    "2303063": "Congratulations on the gold and 6th place!  And thank you for the nice writeup.\n\nWe were worried about threshold calibration for different models.  We briefly considered thresholding at a specific percentile, like you did, but felt that was too risky if the amount of ink on the test set happened to differ.  Curious if you worried about that at all?",
    "2303094": "\"I found that only b3 showed high LB scores\"  Was this also the case for the private lb as well?  If so, do you think this shows that the private and public lb scores were very highly correlated such that overfitting to the public lb was the optimal strategy in this competition?",
    "2303134": "sapr3s Thank you very much! This is just my opinion, but I believe that it is not so much about the rotation itself, but rather aligning the images in the same orientation as during training. The angle of the text during training may have an influence on the results.",
    "2303139": "socated Congratulations on the 1st place! I'm looking forward to hearing about your solution. Personally, I didn't have such concerns because I didn't observe such tendencies in frag1 to frag3. Instead, I thought that it might be influenced by the density of the text. The graph shows the label density, where the x-axis indicates the ratio of pixels with label 1 to the total number of pixels, and the y-axis represents the optimal threshold when using percentiles. \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4001300%2Faa985dc98f0e0d70f4cf73acdbb567ae%2Fsolution4.jpg?generation=1686807056839803&alt=media)\n\nFrom the graph, I confirmed that the optimal threshold changes when the text density varies significantly, but I believed that it would not fluctuate significantly within the same fragment. (I anticipated that the threshold for the private dataset would be slightly lower as the text density tends to be higher compared to the public dataset.)",
    "2303147": "petersk20 I didn't anticipate it, but surprisingly, the single model of SegFormer b3 achieved the highest private LB score(0.66) among my submissions. I was not very confident in the single model due to overfitting concerns. Since single models tend to have larger variances due to the limited training data, aligning them with the public LB may not be ideal. However, I knew from conducting simulations by dividing fragment3 that ensembling would reduce the variance and establish a correlation between the public LB and private LB.",
    "2303171": "Congratulations  6th place\nthanks for share detail explanation @chumajin",
    "2303228": "Congratulations! The concatenation of hidden test fragments A and B is a fantastic approach! \nBut how do you know the bonding surface of each test fragment? Those are different from the public test fragments, which can be estimated through the shape of train fragment 1.",
    "2303232": "Congratulations, that's great",
    "2303274": "May I ask if I use TTA to rotate the images by 90, 180, and 270 and concatenate them with the original image, which is the same as the clockwise rotation effect you mentioned? I think these two should be equivalent.",
    "2303367": "Congratulations @chumajin 🎉🎉🎉",
    "2303635": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F72b2a382e1e9bdc9ad24d2d747566842%2FSelection_999(2229).png?generation=1686829690948484&alt=media)\n\n\n\"2.1 About test data and inference flow\nBased on the brief LB probing and the information provided on the competition page, I had an idea of what the test data might look like. Here is an image that represents my understanding:\"\n\n@chumajin \n\nbased on your idea, i make a submission using the image below. you are right!!!!",
    "2303640": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fd03ecdfb13db02524c85065717d4d18c%2FSelection_999(2230).png?generation=1686829844102390&alt=media)\n\nimage from my post\nhttps://www.kaggle.com/competitions/vesuvius-challenge-ink-detection/discussion/407972\n\n(you also get the number of lines and number of charcters correct)",
    "2303673": "you have just submitted an image or you made rotation during predcition?",
    "2303682": "it depends on your precision/recall of your model.\n\nif you use TTA, the erros may also increase by 4x",
    "2303713": "just submitted an image",
    "2303735": "You make a really good point. Thank you!",
    "2303761": "Congratulations on 6th place and your second solo gold medal!\n\n> the test data was rotated shown in other discussions\n\nIt seems I missed this information.😨\n\nMoreover, I wasn't able to tackle tasks such as threshold adjustments.\nThank you for sharing the nice solution!",
    "2303785": "hengck23 Great confirmation method and result ! Thank you very much. When I saw this image in your post for the first time, I also imagined that this image could be test data !",
    "2303795": "wangxuc Yes, perhaps adding rotation in aug could be effective. However, if rotation is not included, I think adding noise in TTA would be necessary.",
    "2303801": "oh, my phd friend just tell me that you can download orginal images used in arvix pdf paper from:\n\n\" gzipped TeX, DVI, PostScript or HTML (.gz, .dvi.gz, .ps.gz or .html.gz) file depending on submission format. [ Download source ]\"\n\n... and i find a metric results csv file ... and i think kaggler results exceeded the paper results\n\nhttps://arxiv.org/format/2304.02084",
    "2303804": "riow1983 Thank you for comment !! Yes. Considering that the dataset before submitting is divided into two flags and the fragments are rotated, I can imagine the image mentioned above. Additionally, using the Kaggle technique called LB proving, for example, if the edges of masks for flag A and B have the same pixels, you can set the score to zero to confirm.",
    "2303831": "wow @hengck23 What a find! I should have taken more note of the discussion about test fragments being rotated.",
    "2304024": "congrats on your 6th place",
    "2304200": "Hi! Congratulations on 6th place and thank you for nice write-up! Out of curiosity, could you post  the predictions on your validation folds? That would be nice to see the visual results for different models",
    "2304374": "Thank you!, I understand.👍",
    "2304532": "mkotyushev Thank you for comment! I added the visualization of predictions (see ## 4.1).",
    "2304759": "Thanks so much for the detailed write-up and congrats on your 6th place. \n\nOn the rotation of the test-set and your inference rotation back and forth trick: did you add 90degree rotations to the train augmentations? I wonder whether -even with adding the 90degree rotations- the inference rotation trick you mention is still improving the score?\n\nAnd a second question: how did you incorporate IR images into the training? I incorporated it by having two heads on the same model: one head predicted the mask (loss: Dice + BCE) and the other head predicted the IR image (loss: MSE). Then I added the two losses and take the gradient. Is this similar to how you did it as well? \n\nLast but not least, you mention the pre-training on IR images: how does this work? I'm not familiar with pre-training and would love to understand it better.",
    "2305222": "lucasvw Thank you for comment ! \n\n1) No, I did not intentionally include 90-degree rotations in the augmentation. I believed that it could introduce noise in this approach, as the direction of the characters was already known.\n\n2) The usage of IR images was not that complex; I simply stacked the same images and used them as training images. I also used the labels in the same positions as before.\n\n3) Regarding pretraining, the pretrained model I'm using is not specialized in characters. Therefore, I thought that pretraining in the domain of characters and then further fine-tuning could enhance its sensitivity.",
    "2305563": "Congrats brother",
    "2305583": "Congratulations @chumajin !\n\nInteresting how yo have concatenated fragments A and B on the test set, in order to make predictions as one continuous sequence.\n\nKeep up the great work🔥💪",
    "2306699": "[Percentiles - How to calculate Percentiles, Quartiles, ... - YouTube](https://www.youtube.com/watch?v=tXGcs02y9r4)",
    "2306843": "Congrats on another solo gold!",
    "2307050": "Congrats mate! you are on fire! :)",
    "2307119": "imeintanis Thank you my friend!!",
    "2307120": "ohanegby Thank you very much!!",
    "2307850": "Your approach and detailed explanation provide valuable insights into your solution. It's impressive how you utilized a combination of models and techniques to achieve your results, such as ensemble methods, rotating images during inference, and incorporating IR images into the training data. It's also interesting to see how you analyzed the relationship between CV and LB scores and made adjustments accordingly. Overall, your strategy and experimentation demonstrate a strong understanding of the problem and effective problem-solving skills. Congratulations on your success!",
    "2308047": "Congrats man you are superb",
    "2319606": "I have released the training code and inference code publicly!\ntraining code : https://github.com/chumajin/kaggle-VCID\ninference code(kaggle) : https://www.kaggle.com/code/chumajin/vcid-6th-place-inference",
    "2341363": "Thanks for the information :)  would do me a favor."
  },
  "source": "meta"
}