{
  "id": 136815,
  "title": "10th Place: Single Model Using Snapshot Ensemble w/ code",
  "url": "/competitions/bengaliai-cv19/discussion/136815",
  "author_name": "Tawara",
  "post_date": "2020-03-18T00:14:01.561000",
  "votes": 83,
  "comment_count": 31,
  "views": 0,
  "content": "<h2>UPDATED</h2>\n\n<p>I added the additional study at the end of this post. <br>\n* link to discussion about DataAugmentation\n* links to discussion and notebook about visualization of sSE Block \n* tables of my late submission results (w/o RandomErasing, w/o sSE Block, w/ common sSE Block, and w/ <strong><em>Chris's Magic</em></strong>)\n* link to the source code repository on GitHub</p>\n\n<hr>\n\n<p>I was more surprised than pleased when the private leaderboard uncovered. I didn't imagine such a big shake happens.\nI still don't understand why got this place (sorry, since I had only 4 days, I was not able to do enough experiments) , and want to investigate it by late submission.</p>\n\n<p>Although there remain some mysteries, this is my first solo gold medal. I'm so glad to share my solution as a gold place one and become Competitions Master!</p>\n\n<p>Lastly, congrats to all the teams got the medal and participants finished this competition! And thanks to Kaggle and Bengali.AI for hosting this competition!</p>\n\n<p><br>\nHere, I'd like to share my solution's summary.</p>\n\n<h3>Resources</h3>\n\n<p>Kaggle notebooks and a local machine (GTX1080ti x 1)</p>\n\n<h2>Approach</h2>\n\n<p>At first, for tackling the unseen  graphemes problem, I split train data by multi-label stratified group K-fold (regarding a grapheme as a group) and tried training models. However, I was not able to train models successfully. I suspect this is because some components are contained only in one or few grapheme(s).</p>\n\n<p>My approach is very simple － training a model using all train data (because I want to include all components in training). This way, However, has a risk of overfitting. For preventing worsening of generalization performance, I train a model using cosine annealing scheduling and do snapshot ensemble.</p>\n\n<p>There is no special pre-processing nor post-processing.</p>\n\n<h2>Model</h2>\n\n<p>The model is based on SE-ResNext50 and not so special, but has something unique <strong>in global pooling step</strong>. The figure below is the model overview.</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F473234%2F30dfe0e6ec593a3824841290c0b5df98%2F10th_place_model_overview_v2.png?generation=1584668448561081&amp;alt=media\" alt=\"model overview\"></p>\n\n<p>Instead of simply applying Global Average Pooling to feature map and using its result as <strong>common input</strong> of each component's head, I used <strong>sSE Block</strong> and Global Average Pooling <strong>for each component</strong>.</p>\n\n<p>If you want to know more details about sSE Block, read the original paper:</p>\n\n<p><a href=\"https://arxiv.org/abs/1803.02579\">Concurrent Spatial and Channel Squeeze &amp; Excitation in Fully Convolutional Networks(Abhijit Guha Roy, et al., MICCAI 2018)</a></p>\n\n<p><strong>NOTE</strong>: For passing 1-channel images to the pre-trained model(3-channel), I use a simple way. See <a href=\"https://www.kaggle.com/c/bengaliai-cv19/discussion/136815#781819\">the comment bellow</a> for more details.</p>\n\n<h2>Training</h2>\n\n<h3>Data Augmentation</h3>\n\n<p>applying augmentations  by the following order, implemented by <a href=\"https://albumentations.readthedocs.io/en/latest/\">albumentations</a></p>\n\n<p>[Original Image: 137x236]\n-&gt; Padding(to 140x245) -&gt; Rotate(rotate limit=5, p=0.8)\n-&gt; Resize(to 128x224)    -&gt; RandomScale(scale limit=0.1, p=1.0)\n-&gt; Padding(to 146x256) -&gt; RandomCrop(to 128x224, p=1.0)\n-&gt; <strong>RandomErasing(mask-ratio-min=0.02, mask-ratio-max=0.4, p=0.5)</strong>\n-&gt; Normalize(by channel-wise mean and std of train data)\n-&gt; [Model Input: 128x224]</p>\n\n<p>As you see, <strong>I did not use CutMix/MixUp.</strong></p>\n\n<h3>Loss</h3>\n\n<p>calculating softmax cross entropy for each component, and averaging them by weights (<code>grapheme_root</code>:<code>vowel_diacritic</code>:<code>consonant_diacritic</code> = 2 : 1 : 1) </p>\n\n<h3>Optimization</h3>\n\n<ul>\n<li>Optimizer : SGD + NesterovAG(momentum: 0.9) with weight decay(1e-04)</li>\n<li>batch size: 64</li>\n<li>epoch: 105</li>\n<li>learning schedule: cosine annealing (3 cycle)\n<ul><li>35 epoch per cycle</li>\n<li>learning rate: max=1.5e-02, min=0.0</li></ul></li>\n</ul>\n\n<h2>Inference</h2>\n\n<h3>transform</h3>\n\n<p>applying the following transforms to images before feeding them into models</p>\n\n<p>[Original Image: 137x236]\n-&gt; Padding(to 140x245) -&gt; Resize(to 128x224)\n-&gt; Normalize(by channel-wise mean and std of train data)\n-&gt; [Model Input: 128x224]</p>\n\n<h3>ensemble</h3>\n\n<p>Obtaining models(35, 75,  and 105 epoch)' outputs using softmax by component, simply averaging them and applying argmax by component.</p>\n\n<h2>Why got this place ?</h2>\n\n<p>I don't still understand, but have some hypotheses.</p>\n\n<h3>RandomErasing is Good?</h3>\n\n<p>Many participants used CutMix or MixUp. While these techniques seem effective in making combinations of components, not effective in disentangling interdependence of components in original combinations, especially about MixUp.</p>\n\n<p>Because of this, may be, simple masking method like RandomErasing is good.</p>\n\n<h3>sSE Block is Good?</h3>\n\n<p>I'm not sure whether or not sSE Block is the best. But simply applying global average pooling looks not good for me because I suspect each component has spacial dependence.</p>\n\n<p>I suppose a kind of weighted global average pooling <strong>for each component</strong> is good.</p>\n\n<h3>Training a model by all data and Snapshot Ensemble is Good?</h3>\n\n<p>In previous competitions, training K-fold and K-fold averaging was the very effective way for me. But in this competition, it was very difficult for me to split K-fold because of unseen graphemes problem and rare components, I trained a model using all train data.</p>\n\n<p>Here are each cycle's and snapshot ensemble's scores.</p>\n\n<p>|       model         |  Public Score (Rank) | Private Score (Rank) |\n|:-------------------:|:-------:|:--------|\n| cycle 1 (35epoch)   | 0.9754 (288th)  |  0.9438 (23rd) |\n| cycle 2 (70epoch)   | 0.9821 (160th) |  0.9499 (14th) |\n| cycle 3 (105epoch)  | <strong><em>0.9843 (124th)</em></strong> |  0.9497 (14th) |\n| snapshot ensemble   | 0.9840 (131st) |  <strong><em>0.9536 (10th)</em></strong> |</p>\n\n<p>The model cycle 3 achieved the best public score, but this slightly overfitted.\nSnapshot ensemble achieved the best private score, which is <strong>0.0037</strong> higher than the best single model(cycle 2).</p>\n\n<p>I think snapshot ensemble is good when you want to train models <strong>using all train data</strong>.</p>\n\n<p><br>\nThat's all. Thank you for reading!</p>\n\n<p><br></p>\n\n<hr>\n\n<h2>Additional Study</h2>\n\n<h3>RandomErasing is Good?</h3>\n\n<ul>\n<li>Discussion: <a href=\"https://www.kaggle.com/c/bengaliai-cv19/discussion/137029\">CutMix/MixUp is <strong>Not</strong> All You Need?</a></li>\n</ul>\n\n<h3>sSE Block is Good?</h3>\n\n<ul>\n<li>Discussion: <a href=\"https://www.kaggle.com/c/bengaliai-cv19/discussion/137552\">Key of the 10th solution? : Where sSE Block looks</a></li>\n<li>Notebook: <a href=\"https://www.kaggle.com/ttahara/visualize-10th-place-model-where-sse-block-looks\">Visualize 10th place model: Where sSE Block looks?</a></li>\n</ul>\n\n<h3>Late Submission</h3>\n\nAll the results\n\n<ul>\n<li>w/o RandomErasing : not using RandomErasing</li>\n<li>w/o sSE Module : simply apply GAP to feature map extracted from SE-ResNeXt50 and feed it into each component's head (Dense -&gt; ReLU -&gt; Dropout -&gt; Dense)</li>\n<li>w/  <strong><em>Common</em></strong> sSE Module : apply one common sSE-Pooling to feature map and feed it into each component's head</li>\n</ul>\n\n<p>| model |       cycle       | Public Score | Private Score  |\n|:-------:|:-------------------:|:-------:|:--------|\n| w/o RandomErasing |  1 (35epoch)   | 0.9425  |  0.9187  |\n| 〃 | 2 (70epoch)   | 0.9479 |  0.9201  |\n| 〃 | 3 (105epoch)  | 0.9648 |  0.9345  |\n| 〃 | Snapshot Ensemble | 0.9643 |  0.9373 |\n| w/o sSE Module |  1 (35epoch)   | 0.9746 | 0.9431 |\n| 〃 | 2 (70epoch)   | 0.9824 |  0.9490  |\n| 〃 | 3 (105epoch)  | 0.9836 |  0.9485  |\n| 〃 | Snapshot Ensemble   | <strong>0.9843</strong> | 0.9517 |\n| w/ <strong><em>Common</em></strong> sSE Module | 1 (35epoch) | 0.9739 | 0.9433 |\n| 〃 | 2 (70epoch) | 0.9809 | 0.9474 |\n| 〃 | 3 (105epoch) | 0.9827 | 0.9496 |\n| 〃 | Snapshot Ensemble | 0.9832 | <strong>0.9527</strong> |</p>\n\nCompare results using Snapshot Ensemble (&amp; <strong><em>Magical</em></strong> Post-Processing)\n\n<p>| model | Public Score (rank) | Private Score (rank)  |\n|:--------:|:----------- -:|:---------------|\n| w/o RandomErasing | 0.9643 (1210th) | 0.9373 (66th) |\n| w/o sSE Module | 0.9843 (124th) | 0.9517 (12th) |\n| w/ <strong><em>Common</em></strong> sSE Module  | 0.9832 (149th) | 0.9527 (12th) |\n| final sub model (<strong><em>component-wise</em></strong> sSE Module) | 0.9840 (131st) |  0.9536 (10th) |\n| w/ <strong>Chris's Magic</strong>(-0.6,-0.6,-0.4) |  <strong><em>0.9848 (112th)</em></strong> | 0.9623 (4th) |\n| w/ <strong>Chris's Magic</strong>(-0.8,-0.8,-0.7) |  0.9828 (153rd) | <strong><em>0.9653 (3rd)</em></strong> |</p>\n\n<p>OMG! It's a really magic!</p>\n\n<p>If you want to use this magic (it is very easy to use but effective!), check this discussion: <a href=\"https://www.kaggle.com/c/bengaliai-cv19/discussion/136021\">https://www.kaggle.com/c/bengaliai-cv19/discussion/136021</a></p>\n\n<hr>\n\n<h2>code</h2>\n\n<p>I've published the source code on GitHub:\n<a href=\"https://github.com/tawatawara/kaggle-bengaliai-cv19\">https://github.com/tawatawara/kaggle-bengaliai-cv19</a></p>",
  "messages": [
    {
      "id": 777797,
      "postDate": "2020-03-18T00:14:01.563Z",
      "content": "<h2>UPDATED</h2>\n\n<p>I added the additional study at the end of this post. <br>\n* link to discussion about DataAugmentation\n* links to discussion and notebook about visualization of sSE Block \n* tables of my late submission results (w/o RandomErasing, w/o sSE Block, w/ common sSE Block, and w/ <strong><em>Chris's Magic</em></strong>)\n* link to the source code repository on GitHub</p>\n\n<hr>\n\n<p>I was more surprised than pleased when the private leaderboard uncovered. I didn't imagine such a big shake happens.\nI still don't understand why got this place (sorry, since I had only 4 days, I was not able to do enough experiments) , and want to investigate it by late submission.</p>\n\n<p>Although there remain some mysteries, this is my first solo gold medal. I'm so glad to share my solution as a gold place one and become Competitions Master!</p>\n\n<p>Lastly, congrats to all the teams got the medal and participants finished this competition! And thanks to Kaggle and Bengali.AI for hosting this competition!</p>\n\n<p><br>\nHere, I'd like to share my solution's summary.</p>\n\n<h3>Resources</h3>\n\n<p>Kaggle notebooks and a local machine (GTX1080ti x 1)</p>\n\n<h2>Approach</h2>\n\n<p>At first, for tackling the unseen  graphemes problem, I split train data by multi-label stratified group K-fold (regarding a grapheme as a group) and tried training models. However, I was not able to train models successfully. I suspect this is because some components are contained only in one or few grapheme(s).</p>\n\n<p>My approach is very simple － training a model using all train data (because I want to include all components in training). This way, However, has a risk of overfitting. For preventing worsening of generalization performance, I train a model using cosine annealing scheduling and do snapshot ensemble.</p>\n\n<p>There is no special pre-processing nor post-processing.</p>\n\n<h2>Model</h2>\n\n<p>The model is based on SE-ResNext50 and not so special, but has something unique <strong>in global pooling step</strong>. The figure below is the model overview.</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F473234%2F30dfe0e6ec593a3824841290c0b5df98%2F10th_place_model_overview_v2.png?generation=1584668448561081&amp;alt=media\" alt=\"model overview\"></p>\n\n<p>Instead of simply applying Global Average Pooling to feature map and using its result as <strong>common input</strong> of each component's head, I used <strong>sSE Block</strong> and Global Average Pooling <strong>for each component</strong>.</p>\n\n<p>If you want to know more details about sSE Block, read the original paper:</p>\n\n<p><a href=\"https://arxiv.org/abs/1803.02579\">Concurrent Spatial and Channel Squeeze &amp; Excitation in Fully Convolutional Networks(Abhijit Guha Roy, et al., MICCAI 2018)</a></p>\n\n<p><strong>NOTE</strong>: For passing 1-channel images to the pre-trained model(3-channel), I use a simple way. See <a href=\"https://www.kaggle.com/c/bengaliai-cv19/discussion/136815#781819\">the comment bellow</a> for more details.</p>\n\n<h2>Training</h2>\n\n<h3>Data Augmentation</h3>\n\n<p>applying augmentations  by the following order, implemented by <a href=\"https://albumentations.readthedocs.io/en/latest/\">albumentations</a></p>\n\n<p>[Original Image: 137x236]\n-&gt; Padding(to 140x245) -&gt; Rotate(rotate limit=5, p=0.8)\n-&gt; Resize(to 128x224)    -&gt; RandomScale(scale limit=0.1, p=1.0)\n-&gt; Padding(to 146x256) -&gt; RandomCrop(to 128x224, p=1.0)\n-&gt; <strong>RandomErasing(mask-ratio-min=0.02, mask-ratio-max=0.4, p=0.5)</strong>\n-&gt; Normalize(by channel-wise mean and std of train data)\n-&gt; [Model Input: 128x224]</p>\n\n<p>As you see, <strong>I did not use CutMix/MixUp.</strong></p>\n\n<h3>Loss</h3>\n\n<p>calculating softmax cross entropy for each component, and averaging them by weights (<code>grapheme_root</code>:<code>vowel_diacritic</code>:<code>consonant_diacritic</code> = 2 : 1 : 1) </p>\n\n<h3>Optimization</h3>\n\n<ul>\n<li>Optimizer : SGD + NesterovAG(momentum: 0.9) with weight decay(1e-04)</li>\n<li>batch size: 64</li>\n<li>epoch: 105</li>\n<li>learning schedule: cosine annealing (3 cycle)\n<ul><li>35 epoch per cycle</li>\n<li>learning rate: max=1.5e-02, min=0.0</li></ul></li>\n</ul>\n\n<h2>Inference</h2>\n\n<h3>transform</h3>\n\n<p>applying the following transforms to images before feeding them into models</p>\n\n<p>[Original Image: 137x236]\n-&gt; Padding(to 140x245) -&gt; Resize(to 128x224)\n-&gt; Normalize(by channel-wise mean and std of train data)\n-&gt; [Model Input: 128x224]</p>\n\n<h3>ensemble</h3>\n\n<p>Obtaining models(35, 75,  and 105 epoch)' outputs using softmax by component, simply averaging them and applying argmax by component.</p>\n\n<h2>Why got this place ?</h2>\n\n<p>I don't still understand, but have some hypotheses.</p>\n\n<h3>RandomErasing is Good?</h3>\n\n<p>Many participants used CutMix or MixUp. While these techniques seem effective in making combinations of components, not effective in disentangling interdependence of components in original combinations, especially about MixUp.</p>\n\n<p>Because of this, may be, simple masking method like RandomErasing is good.</p>\n\n<h3>sSE Block is Good?</h3>\n\n<p>I'm not sure whether or not sSE Block is the best. But simply applying global average pooling looks not good for me because I suspect each component has spacial dependence.</p>\n\n<p>I suppose a kind of weighted global average pooling <strong>for each component</strong> is good.</p>\n\n<h3>Training a model by all data and Snapshot Ensemble is Good?</h3>\n\n<p>In previous competitions, training K-fold and K-fold averaging was the very effective way for me. But in this competition, it was very difficult for me to split K-fold because of unseen graphemes problem and rare components, I trained a model using all train data.</p>\n\n<p>Here are each cycle's and snapshot ensemble's scores.</p>\n\n<p>|       model         |  Public Score (Rank) | Private Score (Rank) |\n|:-------------------:|:-------:|:--------|\n| cycle 1 (35epoch)   | 0.9754 (288th)  |  0.9438 (23rd) |\n| cycle 2 (70epoch)   | 0.9821 (160th) |  0.9499 (14th) |\n| cycle 3 (105epoch)  | <strong><em>0.9843 (124th)</em></strong> |  0.9497 (14th) |\n| snapshot ensemble   | 0.9840 (131st) |  <strong><em>0.9536 (10th)</em></strong> |</p>\n\n<p>The model cycle 3 achieved the best public score, but this slightly overfitted.\nSnapshot ensemble achieved the best private score, which is <strong>0.0037</strong> higher than the best single model(cycle 2).</p>\n\n<p>I think snapshot ensemble is good when you want to train models <strong>using all train data</strong>.</p>\n\n<p><br>\nThat's all. Thank you for reading!</p>\n\n<p><br></p>\n\n<hr>\n\n<h2>Additional Study</h2>\n\n<h3>RandomErasing is Good?</h3>\n\n<ul>\n<li>Discussion: <a href=\"https://www.kaggle.com/c/bengaliai-cv19/discussion/137029\">CutMix/MixUp is <strong>Not</strong> All You Need?</a></li>\n</ul>\n\n<h3>sSE Block is Good?</h3>\n\n<ul>\n<li>Discussion: <a href=\"https://www.kaggle.com/c/bengaliai-cv19/discussion/137552\">Key of the 10th solution? : Where sSE Block looks</a></li>\n<li>Notebook: <a href=\"https://www.kaggle.com/ttahara/visualize-10th-place-model-where-sse-block-looks\">Visualize 10th place model: Where sSE Block looks?</a></li>\n</ul>\n\n<h3>Late Submission</h3>\n\nAll the results\n\n<ul>\n<li>w/o RandomErasing : not using RandomErasing</li>\n<li>w/o sSE Module : simply apply GAP to feature map extracted from SE-ResNeXt50 and feed it into each component's head (Dense -&gt; ReLU -&gt; Dropout -&gt; Dense)</li>\n<li>w/  <strong><em>Common</em></strong> sSE Module : apply one common sSE-Pooling to feature map and feed it into each component's head</li>\n</ul>\n\n<p>| model |       cycle       | Public Score | Private Score  |\n|:-------:|:-------------------:|:-------:|:--------|\n| w/o RandomErasing |  1 (35epoch)   | 0.9425  |  0.9187  |\n| 〃 | 2 (70epoch)   | 0.9479 |  0.9201  |\n| 〃 | 3 (105epoch)  | 0.9648 |  0.9345  |\n| 〃 | Snapshot Ensemble | 0.9643 |  0.9373 |\n| w/o sSE Module |  1 (35epoch)   | 0.9746 | 0.9431 |\n| 〃 | 2 (70epoch)   | 0.9824 |  0.9490  |\n| 〃 | 3 (105epoch)  | 0.9836 |  0.9485  |\n| 〃 | Snapshot Ensemble   | <strong>0.9843</strong> | 0.9517 |\n| w/ <strong><em>Common</em></strong> sSE Module | 1 (35epoch) | 0.9739 | 0.9433 |\n| 〃 | 2 (70epoch) | 0.9809 | 0.9474 |\n| 〃 | 3 (105epoch) | 0.9827 | 0.9496 |\n| 〃 | Snapshot Ensemble | 0.9832 | <strong>0.9527</strong> |</p>\n\nCompare results using Snapshot Ensemble (&amp; <strong><em>Magical</em></strong> Post-Processing)\n\n<p>| model | Public Score (rank) | Private Score (rank)  |\n|:--------:|:----------- -:|:---------------|\n| w/o RandomErasing | 0.9643 (1210th) | 0.9373 (66th) |\n| w/o sSE Module | 0.9843 (124th) | 0.9517 (12th) |\n| w/ <strong><em>Common</em></strong> sSE Module  | 0.9832 (149th) | 0.9527 (12th) |\n| final sub model (<strong><em>component-wise</em></strong> sSE Module) | 0.9840 (131st) |  0.9536 (10th) |\n| w/ <strong>Chris's Magic</strong>(-0.6,-0.6,-0.4) |  <strong><em>0.9848 (112th)</em></strong> | 0.9623 (4th) |\n| w/ <strong>Chris's Magic</strong>(-0.8,-0.8,-0.7) |  0.9828 (153rd) | <strong><em>0.9653 (3rd)</em></strong> |</p>\n\n<p>OMG! It's a really magic!</p>\n\n<p>If you want to use this magic (it is very easy to use but effective!), check this discussion: <a href=\"https://www.kaggle.com/c/bengaliai-cv19/discussion/136021\">https://www.kaggle.com/c/bengaliai-cv19/discussion/136021</a></p>\n\n<hr>\n\n<h2>code</h2>\n\n<p>I've published the source code on GitHub:\n<a href=\"https://github.com/tawatawara/kaggle-bengaliai-cv19\">https://github.com/tawatawara/kaggle-bengaliai-cv19</a></p>",
      "rawMarkdown": "## UPDATED\nI added the additional study at the end of this post.  \n* link to discussion about DataAugmentation\n* links to discussion and notebook about visualization of sSE Block \n* tables of my late submission results (w/o RandomErasing, w/o sSE Block, w/ common sSE Block, and w/ **_Chris's Magic_**)\n* link to the source code repository on GitHub\n\n----\n\nI was more surprised than pleased when the private leaderboard uncovered. I didn't imagine such a big shake happens.\nI still don't understand why got this place (sorry, since I had only 4 days, I was not able to do enough experiments) , and want to investigate it by late submission.\n\nAlthough there remain some mysteries, this is my first solo gold medal. I'm so glad to share my solution as a gold place one and become Competitions Master!\n\nLastly, congrats to all the teams got the medal and participants finished this competition! And thanks to Kaggle and Bengali.AI for hosting this competition!\n\n<br>\nHere, I'd like to share my solution's summary.\n\n### Resources\nKaggle notebooks and a local machine (GTX1080ti x 1)\n\n## Approach\nAt first, for tackling the unseen  graphemes problem, I split train data by multi-label stratified group K-fold (regarding a grapheme as a group) and tried training models. However, I was not able to train models successfully. I suspect this is because some components are contained only in one or few grapheme(s).\n\nMy approach is very simple － training a model using all train data (because I want to include all components in training). This way, However, has a risk of overfitting. For preventing worsening of generalization performance, I train a model using cosine annealing scheduling and do snapshot ensemble.\n\nThere is no special pre-processing nor post-processing.\n\n## Model\n\nThe model is based on SE-ResNext50 and not so special, but has something unique **in global pooling step**. The figure below is the model overview.\n\n![model overview](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F473234%2F30dfe0e6ec593a3824841290c0b5df98%2F10th_place_model_overview_v2.png?generation=1584668448561081&amp;alt=media)\n\nInstead of simply applying Global Average Pooling to feature map and using its result as **common input** of each component's head, I used **sSE Block** and Global Average Pooling **for each component**.\n\nIf you want to know more details about sSE Block, read the original paper:\n\n[Concurrent Spatial and Channel Squeeze &amp; Excitation in Fully Convolutional Networks(Abhijit Guha Roy, et al., MICCAI 2018)](https://arxiv.org/abs/1803.02579)\n\n**NOTE**: For passing 1-channel images to the pre-trained model(3-channel), I use a simple way. See [the comment bellow](https://www.kaggle.com/c/bengaliai-cv19/discussion/136815#781819) for more details.\n\n## Training\n### Data Augmentation\napplying augmentations  by the following order, implemented by [albumentations](https://albumentations.readthedocs.io/en/latest/)\n\n[Original Image: 137x236]\n-&gt; Padding(to 140x245) -&gt; Rotate(rotate limit=5, p=0.8)\n-&gt; Resize(to 128x224)    -&gt; RandomScale(scale limit=0.1, p=1.0)\n-&gt; Padding(to 146x256) -&gt; RandomCrop(to 128x224, p=1.0)\n-&gt; **RandomErasing(mask-ratio-min=0.02, mask-ratio-max=0.4, p=0.5)**\n-&gt; Normalize(by channel-wise mean and std of train data)\n-&gt; [Model Input: 128x224]\n\nAs you see, **I did not use CutMix/MixUp.**\n\n### Loss\ncalculating softmax cross entropy for each component, and averaging them by weights (`grapheme_root`:`vowel_diacritic`:`consonant_diacritic` = 2 : 1 : 1) \n\n### Optimization\n- Optimizer : SGD + NesterovAG(momentum: 0.9) with weight decay(1e-04)\n- batch size: 64\n- epoch: 105\n- learning schedule: cosine annealing (3 cycle)\n  - 35 epoch per cycle\n  - learning rate: max=1.5e-02, min=0.0\n\n## Inference\n### transform\napplying the following transforms to images before feeding them into models\n\n[Original Image: 137x236]\n-&gt; Padding(to 140x245) -&gt; Resize(to 128x224)\n-&gt; Normalize(by channel-wise mean and std of train data)\n-&gt; [Model Input: 128x224]\n \n### ensemble\nObtaining models(35, 75,  and 105 epoch)' outputs using softmax by component, simply averaging them and applying argmax by component.\n\n## Why got this place ?\nI don't still understand, but have some hypotheses.\n\n### RandomErasing is Good?\nMany participants used CutMix or MixUp. While these techniques seem effective in making combinations of components, not effective in disentangling interdependence of components in original combinations, especially about MixUp.\n\nBecause of this, may be, simple masking method like RandomErasing is good.\n\n### sSE Block is Good?\nI'm not sure whether or not sSE Block is the best. But simply applying global average pooling looks not good for me because I suspect each component has spacial dependence.\n\nI suppose a kind of weighted global average pooling **for each component** is good.\n\n### Training a model by all data and Snapshot Ensemble is Good?\nIn previous competitions, training K-fold and K-fold averaging was the very effective way for me. But in this competition, it was very difficult for me to split K-fold because of unseen graphemes problem and rare components, I trained a model using all train data.\n\nHere are each cycle's and snapshot ensemble's scores.\n\n|       model         |  Public Score (Rank) | Private Score (Rank) |\n|:-------------------:|:-------:|:--------|\n| cycle 1 (35epoch)   | 0.9754 (288th)  |  0.9438 (23rd) |\n| cycle 2 (70epoch)   | 0.9821 (160th) |  0.9499 (14th) |\n| cycle 3 (105epoch)  | **_0.9843 (124th)_** |  0.9497 (14th) |\n| snapshot ensemble   | 0.9840 (131st) |  **_0.9536 (10th)_** |\n\nThe model cycle 3 achieved the best public score, but this slightly overfitted.\nSnapshot ensemble achieved the best private score, which is **0.0037** higher than the best single model(cycle 2).\n\nI think snapshot ensemble is good when you want to train models **using all train data**.\n\n<br>\nThat's all. Thank you for reading!\n\n<br>\n\n----\n\n## Additional Study\n\n### RandomErasing is Good?\n- Discussion: [CutMix/MixUp is **Not** All You Need?](https://www.kaggle.com/c/bengaliai-cv19/discussion/137029)\n\n### sSE Block is Good?\n- Discussion: [Key of the 10th solution? : Where sSE Block looks](https://www.kaggle.com/c/bengaliai-cv19/discussion/137552)\n- Notebook: [Visualize 10th place model: Where sSE Block looks?](https://www.kaggle.com/ttahara/visualize-10th-place-model-where-sse-block-looks)\n\n### Late Submission\n\n#### All the results\n- w/o RandomErasing : not using RandomErasing\n- w/o sSE Module : simply apply GAP to feature map extracted from SE-ResNeXt50 and feed it into each component's head (Dense -&gt; ReLU -&gt; Dropout -&gt; Dense)\n- w/  **_Common_** sSE Module : apply one common sSE-Pooling to feature map and feed it into each component's head\n\n| model |       cycle       | Public Score | Private Score  |\n|:-------:|:-------------------:|:-------:|:--------|\n| w/o RandomErasing |  1 (35epoch)   | 0.9425  |  0.9187  |\n| 〃 | 2 (70epoch)   | 0.9479 |  0.9201  |\n| 〃 | 3 (105epoch)  | 0.9648 |  0.9345  |\n| 〃 | Snapshot Ensemble | 0.9643 |  0.9373 |\n| w/o sSE Module |  1 (35epoch)   | 0.9746 | 0.9431 |\n| 〃 | 2 (70epoch)   | 0.9824 |  0.9490  |\n| 〃 | 3 (105epoch)  | 0.9836 |  0.9485  |\n| 〃 | Snapshot Ensemble   | **0.9843** | 0.9517 |\n| w/ **_Common_** sSE Module | 1 (35epoch) | 0.9739 | 0.9433 |\n| 〃 | 2 (70epoch) | 0.9809 | 0.9474 |\n| 〃 | 3 (105epoch) | 0.9827 | 0.9496 |\n| 〃 | Snapshot Ensemble | 0.9832 | **0.9527** |\n\n#### Compare results using Snapshot Ensemble (&amp; **_Magical_** Post-Processing)\n\n| model | Public Score (rank) | Private Score (rank)  |\n|:--------:|:----------- -:|:---------------|\n| w/o RandomErasing | 0.9643 (1210th) | 0.9373 (66th) |\n| w/o sSE Module | 0.9843 (124th) | 0.9517 (12th) |\n| w/ **_Common_** sSE Module  | 0.9832 (149th) | 0.9527 (12th) |\n| final sub model (**_component-wise_** sSE Module) | 0.9840 (131st) |  0.9536 (10th) |\n| w/ **Chris's Magic**(-0.6,-0.6,-0.4) |  **_0.9848 (112th)_** | 0.9623 (4th) |\n| w/ **Chris's Magic**(-0.8,-0.8,-0.7) |  0.9828 (153rd) | **_0.9653 (3rd)_** |\n\nOMG! It's a really magic!\n\nIf you want to use this magic (it is very easy to use but effective!), check this discussion: https://www.kaggle.com/c/bengaliai-cv19/discussion/136021\n\n----\n\n## code\n\nI've published the source code on GitHub:\nhttps://github.com/tawatawara/kaggle-bengaliai-cv19",
      "votes": 83
    },
    {
      "id": 781819,
      "postDate": "2020-03-21T17:05:19.840Z",
      "content": "<p>Thanks for sharing and congratulations on your solo gold!! Are you passing in the 3-channel image to your model? In your diagram, it's 1-channel; if it's 1-channel then there has to be a conv-layer to convert into 3-channel before passing in to pretrained serex50</p>",
      "rawMarkdown": "Thanks for sharing and congratulations on your solo gold!! Are you passing in the 3-channel image to your model? In your diagram, it's 1-channel; if it's 1-channel then there has to be a conv-layer to convert into 3-channel before passing in to pretrained serex50",
      "votes": 1,
      "replies": [
        {
          "id": 781853,
          "postDate": "2020-03-21T17:46:26.883Z",
          "content": "<p>Thanks!</p>\n\n<p>No, I'm passing <strong><em>1-channel images</em></strong> to SE-ResNeXt50, but this is the pre-trained model(3-channel) on ImageNet.\nI do this in a naive way.</p>\n\n<p>The first conv-layer of SE-ResNeXt50 is <code>Conv2D(in_channels=3, out_channels=64, kernel_size=(7,7), stride=(2,2), pad=(3,3))</code>, of which weight has the shape of <code>(64, 3, 7, 7)</code> as bellow. <br>\n(<strong>NOTE</strong>: I think this shape depends on your framework. I'm using Chainer and ChainerCV.)</p>\n\n<p><code>python\n&amp;gt;&amp;gt;&amp;gt; import numpy as np\n&amp;gt;&amp;gt;&amp;gt; import chainercv\n&amp;gt;&amp;gt;&amp;gt; model = chainercv.links.SEResNet50(pretrained_model=\"imagenet\")\n&amp;gt;&amp;gt;&amp;gt; model.conv1.conv.W.data.shape\n(64, 3, 7, 7)\n</code></p>\n\n<p>I simply average it along <code>axis=1</code> which is corresponded to <code>in_channels</code>, then add a new-axis to this average weight.</p>\n\n<p><code>python\n&amp;gt;&amp;gt;&amp;gt; w_mean = model.conv1.conv.W.data.mean(axis=1)\n&amp;gt;&amp;gt;&amp;gt; w_mean.shape\n(64, 7, 7)\n&amp;gt;&amp;gt;&amp;gt; w_mean = w_mean[:, np.newaxis, :, :]\n&amp;gt;&amp;gt;&amp;gt; w_mean.shape\n(64, 1, 7, 7)\n</code></p>\n\n<p>Finally, I use it as a weight of the first conv-layer of SE-ResNeXt50.</p>\n\n<p><code>\n&amp;gt;&amp;gt;&amp;gt; model.conv1.conv.W.data = w_mean\n&amp;gt;&amp;gt;&amp;gt; model.conv1.conv.W.data.shape\n(64, 1, 7, 7)\n&amp;gt;&amp;gt;&amp;gt; # # passing examples # #\n&amp;gt;&amp;gt;&amp;gt; x = np.random.rand(4, 1, 128, 224).astype(\"f\")\n&amp;gt;&amp;gt;&amp;gt; x.shape  # shape=(batch_size, channel, height, width)\n(4, 1, 128, 224)\n&amp;gt;&amp;gt;&amp;gt; y = model(x)\n&amp;gt;&amp;gt;&amp;gt; y.shape  # shape=(bs, output_class_num)\n(4, 1000)\n</code></p>\n\n<p>I once pushed <code>publish topic</code> by mistake, sorry😅 .</p>",
          "rawMarkdown": "Thanks!\n\nNo, I'm passing **_1-channel images_** to SE-ResNeXt50, but this is the pre-trained model(3-channel) on ImageNet.\nI do this in a naive way.\n\nThe first conv-layer of SE-ResNeXt50 is `Conv2D(in_channels=3, out_channels=64, kernel_size=(7,7), stride=(2,2), pad=(3,3))`, of which weight has the shape of `(64, 3, 7, 7)` as bellow.  \n(**NOTE**: I think this shape depends on your framework. I'm using Chainer and ChainerCV.)\n\n```python\n&gt;&gt;&gt; import numpy as np\n&gt;&gt;&gt; import chainercv\n&gt;&gt;&gt; model = chainercv.links.SEResNet50(pretrained_model=\"imagenet\")\n&gt;&gt;&gt; model.conv1.conv.W.data.shape\n(64, 3, 7, 7)\n```\n\nI simply average it along `axis=1` which is corresponded to `in_channels`, then add a new-axis to this average weight.\n\n```python\n&gt;&gt;&gt; w_mean = model.conv1.conv.W.data.mean(axis=1)\n&gt;&gt;&gt; w_mean.shape\n(64, 7, 7)\n&gt;&gt;&gt; w_mean = w_mean[:, np.newaxis, :, :]\n&gt;&gt;&gt; w_mean.shape\n(64, 1, 7, 7)\n```\n\nFinally, I use it as a weight of the first conv-layer of SE-ResNeXt50.\n\n```\n&gt;&gt;&gt; model.conv1.conv.W.data = w_mean\n&gt;&gt;&gt; model.conv1.conv.W.data.shape\n(64, 1, 7, 7)\n&gt;&gt;&gt; # # passing examples # #\n&gt;&gt;&gt; x = np.random.rand(4, 1, 128, 224).astype(\"f\")\n&gt;&gt;&gt; x.shape  # shape=(batch_size, channel, height, width)\n(4, 1, 128, 224)\n&gt;&gt;&gt; y = model(x)\n&gt;&gt;&gt; y.shape  # shape=(bs, output_class_num)\n(4, 1000)\n```\n\nI once pushed `publish topic` by mistake, sorry😅 .",
          "votes": 2
        },
        {
          "id": 781958,
          "postDate": "2020-03-21T19:38:26.817Z",
          "content": "<p>Thanks!! I'm currently trying to replicate your results with pytorch and with 3-channel image. I'll update my results once I'm done training. Just to be clear: you save model checkpoints at epoch 35, 70 and 105 epochs, right?</p>",
          "rawMarkdown": "Thanks!! I'm currently trying to replicate your results with pytorch and with 3-channel image. I'll update my results once I'm done training. Just to be clear: you save model checkpoints at epoch 35, 70 and 105 epochs, right?",
          "votes": 1
        },
        {
          "id": 781976,
          "postDate": "2020-03-21T20:05:52.110Z",
          "content": "<p>Thank you for trying to reproduce! Yes, I saved the model at epoch 35, 70, and 105.</p>\n\n<p>For your reference, I show you my final model's training curve.</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F473234%2Fc1a76e0dbf058bc81471e746027c8359%2FBengaliAI_CV2019_20200322045207.png?generation=1584820441866694&amp;alt=media\" alt=\"training curve\"></p>\n\n<ul>\n<li>blue line (left axis) : training loss\n<ul><li>It is jaggy because represents the loss of last iteration at each epoch.</li></ul></li>\n<li>pale red line (right axis): learning rate\n<ul><li>cosine anealing scheduling, lr-max=1.5e-02, lr-min=0.0</li></ul></li>\n<li>There is no validation loss <strong>because of training by all train data</strong>.</li>\n</ul>\n\n<p>Regards.</p>",
          "rawMarkdown": "Thank you for trying to reproduce! Yes, I saved the model at epoch 35, 70, and 105.\n\nFor your reference, I show you my final model's training curve.\n\n![training curve](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F473234%2Fc1a76e0dbf058bc81471e746027c8359%2FBengaliAI_CV2019_20200322045207.png?generation=1584820441866694&amp;alt=media)\n\n* blue line (left axis) : training loss\n  * It is jaggy because represents the loss of last iteration at each epoch.\n* pale red line (right axis): learning rate\n  * cosine anealing scheduling, lr-max=1.5e-02, lr-min=0.0\n* There is no validation loss **because of training by all train data**.\n\nRegards.",
          "votes": 2
        }
      ]
    },
    {
      "id": 778407,
      "postDate": "2020-03-18T12:16:08.607Z",
      "content": "<p>Nice approach </p>",
      "rawMarkdown": "Nice approach ",
      "votes": 1,
      "replies": [
        {
          "id": 778459,
          "postDate": "2020-03-18T13:11:56.327Z",
          "content": "<p>Thanks!</p>",
          "rawMarkdown": "Thanks!"
        }
      ]
    },
    {
      "id": 777954,
      "postDate": "2020-03-18T03:23:20.093Z",
      "content": "<p>I guess the biggest factor is training on the whole dataset? I tried training on different folds and ensemble but not sure if it does the same thing as training on the whole dataset. I encountered a big shake-down.</p>",
      "rawMarkdown": "I guess the biggest factor is training on the whole dataset? I tried training on different folds and ensemble but not sure if it does the same thing as training on the whole dataset. I encountered a big shake-down.",
      "votes": 1,
      "replies": [
        {
          "id": 778442,
          "postDate": "2020-03-18T12:56:29.510Z",
          "content": "<p>We tried a snapshot ensemble by training on the entire dataset, using the number of validation epochs as a guide for early stopping. We still shook down. So I think it has more to do with sSEBlock and RandomErasing. <a href=\"/ttahara\">@ttahara</a> , congratulations on your gold!</p>",
          "rawMarkdown": "We tried a snapshot ensemble by training on the entire dataset, using the number of validation epochs as a guide for early stopping. We still shook down. So I think it has more to do with sSEBlock and RandomErasing. @ttahara , congratulations on your gold!",
          "votes": 2
        },
        {
          "id": 778456,
          "postDate": "2020-03-18T13:10:49.823Z",
          "content": "<p>Thanks <a href=\"/returnofsputnik\">@returnofsputnik</a> !</p>\n\n<p>I have a similar thought as yours because my each cycle model got silver or gold place.\nIn my intuition, using RandomErasing <strong>but not using CutMix/MixUp</strong> is more important than sSE Block. </p>",
          "rawMarkdown": "Thanks @returnofsputnik !\n\nI have a similar thought as yours because my each cycle model got silver or gold place.\nIn my intuition, using RandomErasing **but not using CutMix/MixUp** is more important than sSE Block. \n",
          "votes": 1
        },
        {
          "id": 778464,
          "postDate": "2020-03-18T13:18:00.850Z",
          "content": "<p><a href=\"/ttahara\">@ttahara</a> Is there any differnce between <code>RandomErasing</code> and <code>Cutout</code> augmentations? Thank you</p>",
          "rawMarkdown": "@ttahara Is there any differnce between `RandomErasing` and `Cutout` augmentations? Thank you"
        },
        {
          "id": 778507,
          "postDate": "2020-03-18T14:16:05.627Z",
          "content": "<p>I have the opposite intuition. wouldn't erase/cutout force the model to remember combinations / associate different components?</p>",
          "rawMarkdown": "I have the opposite intuition. wouldn't erase/cutout force the model to remember combinations / associate different components?",
          "votes": 1
        },
        {
          "id": 778925,
          "postDate": "2020-03-18T21:25:14.580Z",
          "content": "<p><a href=\"/returnofsputnik\">@returnofsputnik</a> </p>\n\n<p>As far as I know, these two have the following differnces.( sorry if there is any mistake )</p>\n\n<p>|   property ＼ Method     |  Cutout | RandomErasing |\n|:------------|:--------------------------|:---------------|\n|  mask size  |  fixed (hyper param)  | <strong>random</strong> (decided from ranges of hyper param)|\n| mask aspect ratio | fixed | <strong>random</strong> (decided from ranges of hyper param) |\n| mask pixel value | fixed(mean value) | <strong>random</strong> (or fixed value) |\n| mask location | random (if mask crosses edges, the  overflown is cut ) | random (decided to not cross edges) |\n| mask num| 1 or <strong>more than 1</strong> (hyperparam) | 1 |</p>\n\n<p>if you want to know exact information, I recommend you to read the original papers and codes:</p>\n\n<ul>\n<li><code>Cutout</code>\n<ul><li><a href=\"https://arxiv.org/abs/1708.04552\">Improved Regularization of Convolutional Neural Networks with Cutout (Terrance DeVries, et al., arXiv:1708.04552) </a></li>\n<li><a href=\"https://github.com/uoguelph-mlrg/Cutout/blob/master/util/cutout.py\">Original Implementation(GitHub) </a></li>\n<li><a href=\"https://albumentations.readthedocs.io/en/latest/api/augmentations.html#albumentations.augmentations.transforms.Cutout\">Albumentations</a></li></ul></li>\n<li><code>RandomErasing</code>\n<ul><li><a href=\"https://arxiv.org/abs/1708.04896\">Random Erasing Data Augmentation(Zhun Zhong, et al., arXiv:1708.04896) </a></li>\n<li><a href=\"https://github.com/zhunzhong07/Random-Erasing/blob/master/utils/transforms.py\">Original Implementation(GitHub)</a></li></ul></li>\n</ul>",
          "rawMarkdown": "@returnofsputnik \n\nAs far as I know, these two have the following differnces.( sorry if there is any mistake )\n\n|   property ＼ Method     |  Cutout | RandomErasing |\n|:------------|:--------------------------|:---------------|\n|  mask size  |  fixed (hyper param)  | **random** (decided from ranges of hyper param)|\n| mask aspect ratio | fixed | **random** (decided from ranges of hyper param) |\n| mask pixel value | fixed(mean value) | **random** (or fixed value) |\n| mask location | random (if mask crosses edges, the  overflown is cut ) | random (decided to not cross edges) |\n| mask num| 1 or **more than 1** (hyperparam) | 1 |\n\nif you want to know exact information, I recommend you to read the original papers and codes:\n\n* `Cutout`\n  * [Improved Regularization of Convolutional Neural Networks with Cutout (Terrance DeVries, et al., arXiv:1708.04552) ](https://arxiv.org/abs/1708.04552)\n  * [Original Implementation(GitHub) ](https://github.com/uoguelph-mlrg/Cutout/blob/master/util/cutout.py)\n  * [Albumentations](https://albumentations.readthedocs.io/en/latest/api/augmentations.html#albumentations.augmentations.transforms.Cutout)\n* `RandomErasing`\n  * [Random Erasing Data Augmentation(Zhun Zhong, et al., arXiv:1708.04896) ](https://arxiv.org/abs/1708.04896)\n  * [Original Implementation(GitHub)](https://github.com/zhunzhong07/Random-Erasing/blob/master/utils/transforms.py)",
          "votes": 3
        },
        {
          "id": 778933,
          "postDate": "2020-03-18T21:35:09.297Z",
          "content": "<p><a href=\"/yl1202\">@yl1202</a> Thanks! that's a very useful comment for me.</p>\n\n<p>I think your intuition is right if RandomErasing/CutOut  are <strong>always</strong> applied, but practically, these are  <strong>stochastically</strong> applied (e.g. p=0.5).</p>",
          "rawMarkdown": "@yl1202 Thanks! that's a very useful comment for me.\n\nI think your intuition is right if RandomErasing/CutOut  are **always** applied, but practically, these are  **stochastically** applied (e.g. p=0.5).",
          "votes": 1
        }
      ]
    },
    {
      "id": 777911,
      "postDate": "2020-03-18T02:46:20.870Z",
      "content": "<p>👍 👍 👍 </p>",
      "rawMarkdown": "👍 👍 👍 ",
      "votes": 1,
      "replies": [
        {
          "id": 778458,
          "postDate": "2020-03-18T13:11:40.240Z",
          "content": "<p>Thanks!</p>",
          "rawMarkdown": "Thanks!"
        }
      ]
    },
    {
      "id": 777891,
      "postDate": "2020-03-18T02:29:23.500Z",
      "content": "<p>Hi, thank you for sharing and congrats on 10th place.\nI'm amazed you achieved that score without postprocessing such as <a href=\"https://www.kaggle.com/c/bengaliai-cv19/discussion/136021\">https://www.kaggle.com/c/bengaliai-cv19/discussion/136021</a></p>",
      "rawMarkdown": "Hi, thank you for sharing and congrats on 10th place.\nI'm amazed you achieved that score without postprocessing such as https://www.kaggle.com/c/bengaliai-cv19/discussion/136021",
      "votes": 1,
      "replies": [
        {
          "id": 777988,
          "postDate": "2020-03-18T04:03:55.437Z",
          "content": "<p>Thanks! I didn't think of that post-processing. I'll try it.</p>",
          "rawMarkdown": "Thanks! I didn't think of that post-processing. I'll try it.",
          "votes": 1
        }
      ]
    },
    {
      "id": 777828,
      "postDate": "2020-03-18T01:26:43.177Z",
      "content": "<p>Thank you for sharing and congratulations!!</p>",
      "rawMarkdown": "Thank you for sharing and congratulations!!",
      "votes": 1,
      "replies": [
        {
          "id": 777920,
          "postDate": "2020-03-18T02:51:09.950Z",
          "content": "<p>Thanks!</p>",
          "rawMarkdown": "Thanks!"
        }
      ]
    },
    {
      "id": 779039,
      "postDate": "2020-03-19T01:06:49.547Z",
      "content": "<p>Congrats on your 10th place and awesome solo gold! \nOne of the reasons why your model performs so well in Private may be your validation scheme. You used only unseen graphemes for validation until the middle stage of this competition (according to your tweet). I think it made your base model and approach robust to unseen graphemes, even though it might not be so apparent in your final model architecture and settings.</p>",
      "rawMarkdown": "Congrats on your 10th place and awesome solo gold! \nOne of the reasons why your model performs so well in Private may be your validation scheme. You used only unseen graphemes for validation until the middle stage of this competition (according to your tweet). I think it made your base model and approach robust to unseen graphemes, even though it might not be so apparent in your final model architecture and settings.",
      "votes": 2,
      "replies": [
        {
          "id": 779531,
          "postDate": "2020-03-19T12:54:32.750Z",
          "content": "<p>Thanks and congrats on your 6th place!</p>\n\n<p>Yes, you are right. For the first two days, I validated models by validation set split in that manner and decided several settings such as parameters of data augmentation.</p>\n\n<p>But validation loss would not decrease more at an early epoch (5 ~ 7). I was very troubled😅 </p>",
          "rawMarkdown": "Thanks and congrats on your 6th place!\n\nYes, you are right. For the first two days, I validated models by validation set split in that manner and decided several settings such as parameters of data augmentation.\n\nBut validation loss would not decrease more at an early epoch (5 ~ 7). I was very troubled😅 ",
          "votes": 2
        }
      ]
    },
    {
      "id": 777888,
      "postDate": "2020-03-18T02:27:29.290Z",
      "content": "<p>Really interesting, you didn't do any optimization on unseen graphemes but your models did it for you!\nThat's impressive!</p>",
      "rawMarkdown": "Really interesting, you didn't do any optimization on unseen graphemes but your models did it for you!\nThat's impressive!",
      "votes": 2,
      "replies": [
        {
          "id": 777932,
          "postDate": "2020-03-18T03:00:39.387Z",
          "content": "<p>Thanks and congrats on 8th place! I'm really surprised at this result, too.</p>\n\n<p>Although I still don't understand, I'm happy to get this place by very simple solution.</p>",
          "rawMarkdown": "Thanks and congrats on 8th place! I'm really surprised at this result, too.\n\nAlthough I still don't understand, I'm happy to get this place by very simple solution.",
          "votes": 1
        }
      ]
    },
    {
      "id": 795959,
      "postDate": "2020-04-03T06:41:11.283Z",
      "content": "<p>Hi @Tawara</p>\n\n<p>Congrats on 10th place!</p>\n\n<p>Could I have one question?\nWhy did you pad image to different image size then do Rotate/ RandomCrop?</p>\n\n<p>Looking forward to your reply.</p>",
      "rawMarkdown": "Hi @Tawara\n\nCongrats on 10th place!\n\nCould I have one question?\nWhy did you pad image to different image size then do Rotate/ RandomCrop?\n\nLooking forward to your reply."
    },
    {
      "id": 789297,
      "postDate": "2020-03-28T14:33:05.923Z",
      "content": "<p>I updated the result of visualization and  late submissions. The source code is now available on GitHub. <br>\nI think Chris's magic is a really magic😂</p>\n\n<p>Thanks.</p>",
      "rawMarkdown": "I updated the result of visualization and  late submissions. The source code is now available on GitHub.  \nI think Chris's magic is a really magic😂\n\nThanks."
    },
    {
      "id": 780080,
      "postDate": "2020-03-20T00:22:56.677Z",
      "content": "<p>One question, when predict how did you process the images? Directly resize to  128x224?</p>",
      "rawMarkdown": "One question, when predict how did you process the images? Directly resize to  128x224?",
      "replies": [
        {
          "id": 780090,
          "postDate": "2020-03-20T00:59:46.430Z",
          "content": "<p>Sorry, that is represented by <code>Padding + Resize</code> in the overview figure but somewhat ambiguous.</p>\n\n<p>I applied padding original images[137x236] to 140x245, then resized them into 128x224 (and normalized them).</p>\n\n<p>I added this information and corrected the figure. Thanks!</p>",
          "rawMarkdown": "Sorry, that is represented by `Padding + Resize` in the overview figure but somewhat ambiguous.\n\nI applied padding original images[137x236] to 140x245, then resized them into 128x224 (and normalized them).\n\nI added this information and corrected the figure. Thanks!",
          "votes": 1
        }
      ]
    },
    {
      "id": 779414,
      "postDate": "2020-03-19T10:19:03.717Z",
      "content": "<p>Congrats for solo Gold and becoming competition master. From your solution I learned many things that I didn't know. thanks for sharing.</p>",
      "rawMarkdown": "Congrats for solo Gold and becoming competition master. From your solution I learned many things that I didn't know. thanks for sharing.",
      "replies": [
        {
          "id": 779538,
          "postDate": "2020-03-19T13:04:59.473Z",
          "content": "<p>Thanks! I'm glad that my solution gives you many impressions.</p>",
          "rawMarkdown": "Thanks! I'm glad that my solution gives you many impressions."
        }
      ]
    },
    {
      "id": 781646,
      "postDate": "2020-03-21T14:25:19.710Z",
      "rawMarkdown": "",
      "isDeleted": true,
      "replies": [
        {
          "id": 781833,
          "postDate": "2020-03-21T17:19:28.377Z",
          "content": "<p>Thanks!</p>",
          "rawMarkdown": "Thanks!"
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 781819,
      "author_name": "Bibek",
      "author_url": "",
      "post_date": "2020-03-21T17:05:19.840000",
      "content": "<p>Thanks for sharing and congratulations on your solo gold!! Are you passing in the 3-channel image to your model? In your diagram, it's 1-channel; if it's 1-channel then there has to be a conv-layer to convert into 3-channel before passing in to pretrained serex50</p>",
      "votes": 1,
      "replies": [
        {
          "id": 781853,
          "author_name": "Tawara",
          "author_url": "",
          "post_date": "2020-03-21T17:46:26.883000",
          "content": "<p>Thanks!</p>\n\n<p>No, I'm passing <strong><em>1-channel images</em></strong> to SE-ResNeXt50, but this is the pre-trained model(3-channel) on ImageNet.\nI do this in a naive way.</p>\n\n<p>The first conv-layer of SE-ResNeXt50 is <code>Conv2D(in_channels=3, out_channels=64, kernel_size=(7,7), stride=(2,2), pad=(3,3))</code>, of which weight has the shape of <code>(64, 3, 7, 7)</code> as bellow. <br>\n(<strong>NOTE</strong>: I think this shape depends on your framework. I'm using Chainer and ChainerCV.)</p>\n\n<p><code>python\n&amp;gt;&amp;gt;&amp;gt; import numpy as np\n&amp;gt;&amp;gt;&amp;gt; import chainercv\n&amp;gt;&amp;gt;&amp;gt; model = chainercv.links.SEResNet50(pretrained_model=\"imagenet\")\n&amp;gt;&amp;gt;&amp;gt; model.conv1.conv.W.data.shape\n(64, 3, 7, 7)\n</code></p>\n\n<p>I simply average it along <code>axis=1</code> which is corresponded to <code>in_channels</code>, then add a new-axis to this average weight.</p>\n\n<p><code>python\n&amp;gt;&amp;gt;&amp;gt; w_mean = model.conv1.conv.W.data.mean(axis=1)\n&amp;gt;&amp;gt;&amp;gt; w_mean.shape\n(64, 7, 7)\n&amp;gt;&amp;gt;&amp;gt; w_mean = w_mean[:, np.newaxis, :, :]\n&amp;gt;&amp;gt;&amp;gt; w_mean.shape\n(64, 1, 7, 7)\n</code></p>\n\n<p>Finally, I use it as a weight of the first conv-layer of SE-ResNeXt50.</p>\n\n<p><code>\n&amp;gt;&amp;gt;&amp;gt; model.conv1.conv.W.data = w_mean\n&amp;gt;&amp;gt;&amp;gt; model.conv1.conv.W.data.shape\n(64, 1, 7, 7)\n&amp;gt;&amp;gt;&amp;gt; # # passing examples # #\n&amp;gt;&amp;gt;&amp;gt; x = np.random.rand(4, 1, 128, 224).astype(\"f\")\n&amp;gt;&amp;gt;&amp;gt; x.shape  # shape=(batch_size, channel, height, width)\n(4, 1, 128, 224)\n&amp;gt;&amp;gt;&amp;gt; y = model(x)\n&amp;gt;&amp;gt;&amp;gt; y.shape  # shape=(bs, output_class_num)\n(4, 1000)\n</code></p>\n\n<p>I once pushed <code>publish topic</code> by mistake, sorry😅 .</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 781958,
          "author_name": "Bibek",
          "author_url": "",
          "post_date": "2020-03-21T19:38:26.817000",
          "content": "<p>Thanks!! I'm currently trying to replicate your results with pytorch and with 3-channel image. I'll update my results once I'm done training. Just to be clear: you save model checkpoints at epoch 35, 70 and 105 epochs, right?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 781976,
          "author_name": "Tawara",
          "author_url": "",
          "post_date": "2020-03-21T20:05:52.110000",
          "content": "<p>Thank you for trying to reproduce! Yes, I saved the model at epoch 35, 70, and 105.</p>\n\n<p>For your reference, I show you my final model's training curve.</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F473234%2Fc1a76e0dbf058bc81471e746027c8359%2FBengaliAI_CV2019_20200322045207.png?generation=1584820441866694&amp;alt=media\" alt=\"training curve\"></p>\n\n<ul>\n<li>blue line (left axis) : training loss\n<ul><li>It is jaggy because represents the loss of last iteration at each epoch.</li></ul></li>\n<li>pale red line (right axis): learning rate\n<ul><li>cosine anealing scheduling, lr-max=1.5e-02, lr-min=0.0</li></ul></li>\n<li>There is no validation loss <strong>because of training by all train data</strong>.</li>\n</ul>\n\n<p>Regards.</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 778407,
      "author_name": "Kavitha Kavi",
      "author_url": "",
      "post_date": "2020-03-18T12:16:08.607000",
      "content": "<p>Nice approach </p>",
      "votes": 1,
      "replies": [
        {
          "id": 778459,
          "author_name": "Tawara",
          "author_url": "",
          "post_date": "2020-03-18T13:11:56.327000",
          "content": "<p>Thanks!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 777954,
      "author_name": "Gold Retriever",
      "author_url": "",
      "post_date": "2020-03-18T03:23:20.093000",
      "content": "<p>I guess the biggest factor is training on the whole dataset? I tried training on different folds and ensemble but not sure if it does the same thing as training on the whole dataset. I encountered a big shake-down.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 778442,
          "author_name": "CoreyJamesLevinson",
          "author_url": "",
          "post_date": "2020-03-18T12:56:29.510000",
          "content": "<p>We tried a snapshot ensemble by training on the entire dataset, using the number of validation epochs as a guide for early stopping. We still shook down. So I think it has more to do with sSEBlock and RandomErasing. <a href=\"/ttahara\">@ttahara</a> , congratulations on your gold!</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 778456,
          "author_name": "Tawara",
          "author_url": "",
          "post_date": "2020-03-18T13:10:49.823000",
          "content": "<p>Thanks <a href=\"/returnofsputnik\">@returnofsputnik</a> !</p>\n\n<p>I have a similar thought as yours because my each cycle model got silver or gold place.\nIn my intuition, using RandomErasing <strong>but not using CutMix/MixUp</strong> is more important than sSE Block. </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 778464,
          "author_name": "CoreyJamesLevinson",
          "author_url": "",
          "post_date": "2020-03-18T13:18:00.850000",
          "content": "<p><a href=\"/ttahara\">@ttahara</a> Is there any differnce between <code>RandomErasing</code> and <code>Cutout</code> augmentations? Thank you</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 778507,
          "author_name": "YL",
          "author_url": "",
          "post_date": "2020-03-18T14:16:05.627000",
          "content": "<p>I have the opposite intuition. wouldn't erase/cutout force the model to remember combinations / associate different components?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 778925,
          "author_name": "Tawara",
          "author_url": "",
          "post_date": "2020-03-18T21:25:14.580000",
          "content": "<p><a href=\"/returnofsputnik\">@returnofsputnik</a> </p>\n\n<p>As far as I know, these two have the following differnces.( sorry if there is any mistake )</p>\n\n<p>|   property ＼ Method     |  Cutout | RandomErasing |\n|:------------|:--------------------------|:---------------|\n|  mask size  |  fixed (hyper param)  | <strong>random</strong> (decided from ranges of hyper param)|\n| mask aspect ratio | fixed | <strong>random</strong> (decided from ranges of hyper param) |\n| mask pixel value | fixed(mean value) | <strong>random</strong> (or fixed value) |\n| mask location | random (if mask crosses edges, the  overflown is cut ) | random (decided to not cross edges) |\n| mask num| 1 or <strong>more than 1</strong> (hyperparam) | 1 |</p>\n\n<p>if you want to know exact information, I recommend you to read the original papers and codes:</p>\n\n<ul>\n<li><code>Cutout</code>\n<ul><li><a href=\"https://arxiv.org/abs/1708.04552\">Improved Regularization of Convolutional Neural Networks with Cutout (Terrance DeVries, et al., arXiv:1708.04552) </a></li>\n<li><a href=\"https://github.com/uoguelph-mlrg/Cutout/blob/master/util/cutout.py\">Original Implementation(GitHub) </a></li>\n<li><a href=\"https://albumentations.readthedocs.io/en/latest/api/augmentations.html#albumentations.augmentations.transforms.Cutout\">Albumentations</a></li></ul></li>\n<li><code>RandomErasing</code>\n<ul><li><a href=\"https://arxiv.org/abs/1708.04896\">Random Erasing Data Augmentation(Zhun Zhong, et al., arXiv:1708.04896) </a></li>\n<li><a href=\"https://github.com/zhunzhong07/Random-Erasing/blob/master/utils/transforms.py\">Original Implementation(GitHub)</a></li></ul></li>\n</ul>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 778933,
          "author_name": "Tawara",
          "author_url": "",
          "post_date": "2020-03-18T21:35:09.297000",
          "content": "<p><a href=\"/yl1202\">@yl1202</a> Thanks! that's a very useful comment for me.</p>\n\n<p>I think your intuition is right if RandomErasing/CutOut  are <strong>always</strong> applied, but practically, these are  <strong>stochastically</strong> applied (e.g. p=0.5).</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 777911,
      "author_name": "Tian",
      "author_url": "",
      "post_date": "2020-03-18T02:46:20.870000",
      "content": "<p>👍 👍 👍 </p>",
      "votes": 1,
      "replies": [
        {
          "id": 778458,
          "author_name": "Tawara",
          "author_url": "",
          "post_date": "2020-03-18T13:11:40.240000",
          "content": "<p>Thanks!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 777891,
      "author_name": "Appian",
      "author_url": "",
      "post_date": "2020-03-18T02:29:23.500000",
      "content": "<p>Hi, thank you for sharing and congrats on 10th place.\nI'm amazed you achieved that score without postprocessing such as <a href=\"https://www.kaggle.com/c/bengaliai-cv19/discussion/136021\">https://www.kaggle.com/c/bengaliai-cv19/discussion/136021</a></p>",
      "votes": 1,
      "replies": [
        {
          "id": 777988,
          "author_name": "Tawara",
          "author_url": "",
          "post_date": "2020-03-18T04:03:55.437000",
          "content": "<p>Thanks! I didn't think of that post-processing. I'll try it.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 777828,
      "author_name": "mocobt",
      "author_url": "",
      "post_date": "2020-03-18T01:26:43.177000",
      "content": "<p>Thank you for sharing and congratulations!!</p>",
      "votes": 1,
      "replies": [
        {
          "id": 777920,
          "author_name": "Tawara",
          "author_url": "",
          "post_date": "2020-03-18T02:51:09.950000",
          "content": "<p>Thanks!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 779039,
      "author_name": "YuyaYamamoto",
      "author_url": "",
      "post_date": "2020-03-19T01:06:49.547000",
      "content": "<p>Congrats on your 10th place and awesome solo gold! \nOne of the reasons why your model performs so well in Private may be your validation scheme. You used only unseen graphemes for validation until the middle stage of this competition (according to your tweet). I think it made your base model and approach robust to unseen graphemes, even though it might not be so apparent in your final model architecture and settings.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 779531,
          "author_name": "Tawara",
          "author_url": "",
          "post_date": "2020-03-19T12:54:32.750000",
          "content": "<p>Thanks and congrats on your 6th place!</p>\n\n<p>Yes, you are right. For the first two days, I validated models by validation set split in that manner and decided several settings such as parameters of data augmentation.</p>\n\n<p>But validation loss would not decrease more at an early epoch (5 ~ 7). I was very troubled😅 </p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 777888,
      "author_name": "Qishen Ha",
      "author_url": "",
      "post_date": "2020-03-18T02:27:29.290000",
      "content": "<p>Really interesting, you didn't do any optimization on unseen graphemes but your models did it for you!\nThat's impressive!</p>",
      "votes": 2,
      "replies": [
        {
          "id": 777932,
          "author_name": "Tawara",
          "author_url": "",
          "post_date": "2020-03-18T03:00:39.387000",
          "content": "<p>Thanks and congrats on 8th place! I'm really surprised at this result, too.</p>\n\n<p>Although I still don't understand, I'm happy to get this place by very simple solution.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 795959,
      "author_name": "YS",
      "author_url": "",
      "post_date": "2020-04-03T06:41:11.283000",
      "content": "<p>Hi @Tawara</p>\n\n<p>Congrats on 10th place!</p>\n\n<p>Could I have one question?\nWhy did you pad image to different image size then do Rotate/ RandomCrop?</p>\n\n<p>Looking forward to your reply.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 789297,
      "author_name": "Tawara",
      "author_url": "",
      "post_date": "2020-03-28T14:33:05.923000",
      "content": "<p>I updated the result of visualization and  late submissions. The source code is now available on GitHub. <br>\nI think Chris's magic is a really magic😂</p>\n\n<p>Thanks.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 780080,
      "author_name": "Strideradu",
      "author_url": "",
      "post_date": "2020-03-20T00:22:56.677000",
      "content": "<p>One question, when predict how did you process the images? Directly resize to  128x224?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 780090,
          "author_name": "Tawara",
          "author_url": "",
          "post_date": "2020-03-20T00:59:46.430000",
          "content": "<p>Sorry, that is represented by <code>Padding + Resize</code> in the overview figure but somewhat ambiguous.</p>\n\n<p>I applied padding original images[137x236] to 140x245, then resized them into 128x224 (and normalized them).</p>\n\n<p>I added this information and corrected the figure. Thanks!</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 779414,
      "author_name": "Raghawendra Singh",
      "author_url": "",
      "post_date": "2020-03-19T10:19:03.717000",
      "content": "<p>Congrats for solo Gold and becoming competition master. From your solution I learned many things that I didn't know. thanks for sharing.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 779538,
          "author_name": "Tawara",
          "author_url": "",
          "post_date": "2020-03-19T13:04:59.473000",
          "content": "<p>Thanks! I'm glad that my solution gives you many impressions.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 781646,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-03-21T14:25:19.710000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 781833,
          "author_name": "Tawara",
          "author_url": "",
          "post_date": "2020-03-21T17:19:28.377000",
          "content": "<p>Thanks!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "777797": "## UPDATED\nI added the additional study at the end of this post.  \n* link to discussion about DataAugmentation\n* links to discussion and notebook about visualization of sSE Block \n* tables of my late submission results (w/o RandomErasing, w/o sSE Block, w/ common sSE Block, and w/ **_Chris's Magic_**)\n* link to the source code repository on GitHub\n\n----\n\nI was more surprised than pleased when the private leaderboard uncovered. I didn't imagine such a big shake happens.\nI still don't understand why got this place (sorry, since I had only 4 days, I was not able to do enough experiments) , and want to investigate it by late submission.\n\nAlthough there remain some mysteries, this is my first solo gold medal. I'm so glad to share my solution as a gold place one and become Competitions Master!\n\nLastly, congrats to all the teams got the medal and participants finished this competition! And thanks to Kaggle and Bengali.AI for hosting this competition!\n\n<br>\nHere, I'd like to share my solution's summary.\n\n### Resources\nKaggle notebooks and a local machine (GTX1080ti x 1)\n\n## Approach\nAt first, for tackling the unseen  graphemes problem, I split train data by multi-label stratified group K-fold (regarding a grapheme as a group) and tried training models. However, I was not able to train models successfully. I suspect this is because some components are contained only in one or few grapheme(s).\n\nMy approach is very simple － training a model using all train data (because I want to include all components in training). This way, However, has a risk of overfitting. For preventing worsening of generalization performance, I train a model using cosine annealing scheduling and do snapshot ensemble.\n\nThere is no special pre-processing nor post-processing.\n\n## Model\n\nThe model is based on SE-ResNext50 and not so special, but has something unique **in global pooling step**. The figure below is the model overview.\n\n![model overview](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F473234%2F30dfe0e6ec593a3824841290c0b5df98%2F10th_place_model_overview_v2.png?generation=1584668448561081&amp;alt=media)\n\nInstead of simply applying Global Average Pooling to feature map and using its result as **common input** of each component's head, I used **sSE Block** and Global Average Pooling **for each component**.\n\nIf you want to know more details about sSE Block, read the original paper:\n\n[Concurrent Spatial and Channel Squeeze &amp; Excitation in Fully Convolutional Networks(Abhijit Guha Roy, et al., MICCAI 2018)](https://arxiv.org/abs/1803.02579)\n\n**NOTE**: For passing 1-channel images to the pre-trained model(3-channel), I use a simple way. See [the comment bellow](https://www.kaggle.com/c/bengaliai-cv19/discussion/136815#781819) for more details.\n\n## Training\n### Data Augmentation\napplying augmentations  by the following order, implemented by [albumentations](https://albumentations.readthedocs.io/en/latest/)\n\n[Original Image: 137x236]\n-&gt; Padding(to 140x245) -&gt; Rotate(rotate limit=5, p=0.8)\n-&gt; Resize(to 128x224)    -&gt; RandomScale(scale limit=0.1, p=1.0)\n-&gt; Padding(to 146x256) -&gt; RandomCrop(to 128x224, p=1.0)\n-&gt; **RandomErasing(mask-ratio-min=0.02, mask-ratio-max=0.4, p=0.5)**\n-&gt; Normalize(by channel-wise mean and std of train data)\n-&gt; [Model Input: 128x224]\n\nAs you see, **I did not use CutMix/MixUp.**\n\n### Loss\ncalculating softmax cross entropy for each component, and averaging them by weights (`grapheme_root`:`vowel_diacritic`:`consonant_diacritic` = 2 : 1 : 1) \n\n### Optimization\n- Optimizer : SGD + NesterovAG(momentum: 0.9) with weight decay(1e-04)\n- batch size: 64\n- epoch: 105\n- learning schedule: cosine annealing (3 cycle)\n  - 35 epoch per cycle\n  - learning rate: max=1.5e-02, min=0.0\n\n## Inference\n### transform\napplying the following transforms to images before feeding them into models\n\n[Original Image: 137x236]\n-&gt; Padding(to 140x245) -&gt; Resize(to 128x224)\n-&gt; Normalize(by channel-wise mean and std of train data)\n-&gt; [Model Input: 128x224]\n \n### ensemble\nObtaining models(35, 75,  and 105 epoch)' outputs using softmax by component, simply averaging them and applying argmax by component.\n\n## Why got this place ?\nI don't still understand, but have some hypotheses.\n\n### RandomErasing is Good?\nMany participants used CutMix or MixUp. While these techniques seem effective in making combinations of components, not effective in disentangling interdependence of components in original combinations, especially about MixUp.\n\nBecause of this, may be, simple masking method like RandomErasing is good.\n\n### sSE Block is Good?\nI'm not sure whether or not sSE Block is the best. But simply applying global average pooling looks not good for me because I suspect each component has spacial dependence.\n\nI suppose a kind of weighted global average pooling **for each component** is good.\n\n### Training a model by all data and Snapshot Ensemble is Good?\nIn previous competitions, training K-fold and K-fold averaging was the very effective way for me. But in this competition, it was very difficult for me to split K-fold because of unseen graphemes problem and rare components, I trained a model using all train data.\n\nHere are each cycle's and snapshot ensemble's scores.\n\n|       model         |  Public Score (Rank) | Private Score (Rank) |\n|:-------------------:|:-------:|:--------|\n| cycle 1 (35epoch)   | 0.9754 (288th)  |  0.9438 (23rd) |\n| cycle 2 (70epoch)   | 0.9821 (160th) |  0.9499 (14th) |\n| cycle 3 (105epoch)  | **_0.9843 (124th)_** |  0.9497 (14th) |\n| snapshot ensemble   | 0.9840 (131st) |  **_0.9536 (10th)_** |\n\nThe model cycle 3 achieved the best public score, but this slightly overfitted.\nSnapshot ensemble achieved the best private score, which is **0.0037** higher than the best single model(cycle 2).\n\nI think snapshot ensemble is good when you want to train models **using all train data**.\n\n<br>\nThat's all. Thank you for reading!\n\n<br>\n\n----\n\n## Additional Study\n\n### RandomErasing is Good?\n- Discussion: [CutMix/MixUp is **Not** All You Need?](https://www.kaggle.com/c/bengaliai-cv19/discussion/137029)\n\n### sSE Block is Good?\n- Discussion: [Key of the 10th solution? : Where sSE Block looks](https://www.kaggle.com/c/bengaliai-cv19/discussion/137552)\n- Notebook: [Visualize 10th place model: Where sSE Block looks?](https://www.kaggle.com/ttahara/visualize-10th-place-model-where-sse-block-looks)\n\n### Late Submission\n\n#### All the results\n- w/o RandomErasing : not using RandomErasing\n- w/o sSE Module : simply apply GAP to feature map extracted from SE-ResNeXt50 and feed it into each component's head (Dense -&gt; ReLU -&gt; Dropout -&gt; Dense)\n- w/  **_Common_** sSE Module : apply one common sSE-Pooling to feature map and feed it into each component's head\n\n| model |       cycle       | Public Score | Private Score  |\n|:-------:|:-------------------:|:-------:|:--------|\n| w/o RandomErasing |  1 (35epoch)   | 0.9425  |  0.9187  |\n| 〃 | 2 (70epoch)   | 0.9479 |  0.9201  |\n| 〃 | 3 (105epoch)  | 0.9648 |  0.9345  |\n| 〃 | Snapshot Ensemble | 0.9643 |  0.9373 |\n| w/o sSE Module |  1 (35epoch)   | 0.9746 | 0.9431 |\n| 〃 | 2 (70epoch)   | 0.9824 |  0.9490  |\n| 〃 | 3 (105epoch)  | 0.9836 |  0.9485  |\n| 〃 | Snapshot Ensemble   | **0.9843** | 0.9517 |\n| w/ **_Common_** sSE Module | 1 (35epoch) | 0.9739 | 0.9433 |\n| 〃 | 2 (70epoch) | 0.9809 | 0.9474 |\n| 〃 | 3 (105epoch) | 0.9827 | 0.9496 |\n| 〃 | Snapshot Ensemble | 0.9832 | **0.9527** |\n\n#### Compare results using Snapshot Ensemble (&amp; **_Magical_** Post-Processing)\n\n| model | Public Score (rank) | Private Score (rank)  |\n|:--------:|:----------- -:|:---------------|\n| w/o RandomErasing | 0.9643 (1210th) | 0.9373 (66th) |\n| w/o sSE Module | 0.9843 (124th) | 0.9517 (12th) |\n| w/ **_Common_** sSE Module  | 0.9832 (149th) | 0.9527 (12th) |\n| final sub model (**_component-wise_** sSE Module) | 0.9840 (131st) |  0.9536 (10th) |\n| w/ **Chris's Magic**(-0.6,-0.6,-0.4) |  **_0.9848 (112th)_** | 0.9623 (4th) |\n| w/ **Chris's Magic**(-0.8,-0.8,-0.7) |  0.9828 (153rd) | **_0.9653 (3rd)_** |\n\nOMG! It's a really magic!\n\nIf you want to use this magic (it is very easy to use but effective!), check this discussion: https://www.kaggle.com/c/bengaliai-cv19/discussion/136021\n\n----\n\n## code\n\nI've published the source code on GitHub:\nhttps://github.com/tawatawara/kaggle-bengaliai-cv19",
    "781819": "Thanks for sharing and congratulations on your solo gold!! Are you passing in the 3-channel image to your model? In your diagram, it's 1-channel; if it's 1-channel then there has to be a conv-layer to convert into 3-channel before passing in to pretrained serex50",
    "778407": "Nice approach ",
    "777954": "I guess the biggest factor is training on the whole dataset? I tried training on different folds and ensemble but not sure if it does the same thing as training on the whole dataset. I encountered a big shake-down.",
    "777911": "👍 👍 👍 ",
    "777891": "Hi, thank you for sharing and congrats on 10th place.\nI'm amazed you achieved that score without postprocessing such as https://www.kaggle.com/c/bengaliai-cv19/discussion/136021",
    "777828": "Thank you for sharing and congratulations!!",
    "779039": "Congrats on your 10th place and awesome solo gold! \nOne of the reasons why your model performs so well in Private may be your validation scheme. You used only unseen graphemes for validation until the middle stage of this competition (according to your tweet). I think it made your base model and approach robust to unseen graphemes, even though it might not be so apparent in your final model architecture and settings.",
    "777888": "Really interesting, you didn't do any optimization on unseen graphemes but your models did it for you!\nThat's impressive!",
    "795959": "Hi @Tawara\n\nCongrats on 10th place!\n\nCould I have one question?\nWhy did you pad image to different image size then do Rotate/ RandomCrop?\n\nLooking forward to your reply.",
    "789297": "I updated the result of visualization and  late submissions. The source code is now available on GitHub.  \nI think Chris's magic is a really magic😂\n\nThanks.",
    "780080": "One question, when predict how did you process the images? Directly resize to  128x224?",
    "779414": "Congrats for solo Gold and becoming competition master. From your solution I learned many things that I didn't know. thanks for sharing.",
    "781646": ""
  }
}