{
  "id": 136129,
  "title": "5th place solution",
  "url": "/competitions/bengaliai-cv19/discussion/136129",
  "author_name": "Dieter",
  "post_date": "2020-03-17T15:27:08.871000",
  "votes": 90,
  "comment_count": 29,
  "views": 0,
  "content": "<p>Thanks at Bengali.ai and kaggle for hosting this interesting competition as it was more to it than simple ensembling of computer vision models. Thanks to the authors and contributors of pytorch, pytorch-lightning, apex and pytorchcv for making my life easier.</p>\n\n<p><strong>Acknowledgements &amp; notes on GM</strong></p>\n\n<p>I am humbled to finally become competition grandmaster. I learned a tremendous amount of tricks in the recent two years and it would not have been able if it wasn’t for all this generous sharing of top solutions and extremely talented teammates I was lucky to have along the way. </p>\n\n<h3>Short Summary</h3>\n\n<p>My solution is a simple bag (different seeds) of the same model seresnext50 with custom head. Major boost in LB score came from redesign of consonant diacritic target.  Minor LB improvement from postprocessing tweaks such as thresholding and finding closest train examples. </p>\n\n<h3>Preprocessing</h3>\n\n<p>I did not resize the image, the only thing I did was normalize each image by its mean and std, since I experienced a good regularization from that in previous competitions.</p>\n\n<h3>Architecture &amp; Training</h3>\n\n<p><strong>Backbone</strong>\nI used a plain seresnext50 but adjusted the very first layer to replace resizing the image and account for single channel. I did that by changing input units and reducing stride from (2,2) to (1,2).</p>\n\n<p>```\nfrom pytorchcv.model_provider import get_model as ptcv_get_model</p>\n\n<p>backbone = ptcv_get_model('seresnext50_32x4d', pretrained=True)\nbackbone = backbone.features\nbackbone.init_block.conv = nn.Conv2d(1, 64, kernel_size=(7, 7), stride=(1, 2))\n```</p>\n\n<p><strong>neck</strong></p>\n\n<p>same as in maciejsypetkowski solution <a href=\"https://www.kaggle.com/c/recursion-cellular-image-classification/discussion/110543\">https://www.kaggle.com/c/recursion-cellular-image-classification/discussion/110543</a> </p>\n\n<p><strong>heads</strong></p>\n\n<ul>\n<li>3 heads, each for consonant, vowel and root and </li>\n<li>auxiliary head for grapheme with arccos loss</li>\n</ul>\n\n<p>So in total my used architecture looks like this:</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1424766%2F3e968cbc82806b9211bf7609b662a18a%2Fbengali%20arch.001.jpeg?generation=1584456769762003&amp;alt=media\" alt=\"\"></p>\n\n<p><strong>Training</strong></p>\n\n<ul>\n<li>train on all data</li>\n<li>4 cycle cosine annealing with augmentation increase\n<ul><li>cycle 1 &amp; 2 cutmix</li>\n<li>cycle 3 &amp; 4 cutmix, cutout, scale, translate, rotate</li></ul></li>\n<li>Adam </li>\n</ul>\n\n<p><strong>Losses</strong></p>\n\n<ul>\n<li>root loss: CrossEntropy</li>\n<li>consonant loss: Multi label Binary Crossentropy</li>\n<li>vowel loss : CrossEntropy</li>\n<li><p>grapheme loss: ArcCos + CrossEntropy</p></li>\n<li><p>total loss = root loss + consonant loss + vowel loss + grapheme loss</p></li>\n</ul>\n\n<h3>The Magic (jôfôla)</h3>\n\n<p>Now to the key aspect of my solution. I realised all models perform poorly on consonant diacritic 3 (I abbreviate with cons3 in the following) and especially confuse it with cons4. And with poorly I mean game changing poorly: 89% recall compared to 99% for the other consonant classes. At first I just fought the symptoms (instead of the root cause) as several other teams did, e.g. by transferring low confident predicts from high frequent class cons4 to low frequent class cons3 and hence leveraging the definition of macro recall metric and that gave a good LB boost (0.9900 -&gt; 0.9915). That might have been enough for gold but I wanted to improve my chances. So I wanted to understand why the model did so poor on this one and read through discussions, wikipedia and other domain knowledge related things. Luckily, I found 2 things:\n- Wikipedia: As the last member of a conjunct, য jô appears as a wavy vertical line (called যফলা jôfôla) to the right of the previous member: ক্য \"kyô\" খ্য \"khyô\" গ্য \"gyô\" ঘ্য \"ghyô\" etc. In some fonts, certain conjuncts with যফলা jôfôla appear using special fused forms: দ্য \"dyô\" ন্য \"nyô\" শ্য \"shyô\" ষ্য \"ṣyô\" স্য \"syô\" হ্য \"hyô\".</p>\n\n<ul>\n<li>The consonant diacritics are itself composed from lower level components\n<ul><li>0: []</li>\n<li>1:  ['ঁ']</li>\n<li>2: ['র', '্']</li>\n<li>3: ['র', '্', 'য']</li>\n<li>4: ['্', 'য']</li>\n<li>5:['্', 'র']</li>\n<li>6: ['্', 'র', '্', 'য']</li></ul></li>\n</ul>\n\n<p>So I realized that not only 3,4 and 6 have a jô component which might have the special jôfôla form but also that the special forms are in train for <strong>only</strong> cons4 but <strong>not</strong> for cons3, and thats why the models are so bad on cons3. My solution was the following. I recoded the cons classes into a multilabel classification problem using the following function:</p>\n\n<p>```\ncons_components = ['ঁ', 'য', 'র','্']</p>\n\n<p>def is_sub(sub, lst):\n   ln = len(sub)\n   return any(lst[i: i + ln] == sub for i in range(len(lst) - ln + 1))</p>\n\n<p>def label2label_v2(label):\n   res = np.zeros((6,),dtype=int)\n   components = cons2components[label]\n   for i,item in enumerate(cons_components):\n       if item in components:\n           res[i] = 1\n   if is_sub(['্', 'য'], components):\n       res[4] = 1\n   if is_sub(['্', 'র'], components):\n       res[5] = 1\n   return res</p>\n\n<p>```</p>\n\n<p>That function re-codes the 7 consonant class labels into multilabel 6 dimensional vectors</p>\n\n<p><code>\n0 -&gt; [0, 0, 0, 0, 0, 0]\n1 -&gt; [1, 0, 0, 0, 0, 0]\n2 -&gt; [0, 0, 1, 1, 0, 0]\n3 -&gt; [0, 1, 1, 1, 1, 0]\n4 -&gt; [0, 1, 0, 1, 1, 0]\n5 -&gt; [0, 0, 1, 1, 0, 1]\n6 -&gt; [0, 1, 1, 1, 1, 1]\n</code></p>\n\n<p>That schema not only enabled to learn special forms of cons3 from cons4 but also enables to tune via thresholding (i.e. when to set a continuous prediction to 0 or 1) how the model separates the 6 classes. E.g. by setting a very low threshold on the 2nd logit more predictions would be set to 1 and hence more predictions would move from cons4 to cons3 as these two classes only differ in logit2. Purely changing the target for consonant diacritic in this way improved Public LB from 0.990 to 0.993 !</p>\n\n<h3>Postprocessing</h3>\n\n<ul>\n<li>I adjusted threshold for binarizing consonant diacritic. Best was a threshold of 0.1 for all logits classes. I tried not dare to temper which each logit threshold individually as I was afraid of overfitting. Result on Public LB: 9930 -&gt; 9937 </li>\n<li>I used cosine similarity to find closest train sample for root and vowel when top1 probability &lt; 90% and use its label. bestfitting did the same in the 1st place solution of Human Protein competition) as can be found under <a href=\"https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/78109\">https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/78109</a> Result on Public LB: 9937 -&gt; 9942 </li>\n</ul>\n\n<h3>Choosing subs</h3>\n\n<p>I tempered around with binarization and metric learning thresholds and purely checked LB hence I was aware of a high risk to potentially overfit to LB. So I chose my best sub and a failsafe sub that I felt was the least tempered with but still doing ok. (Both were gold on private at the end)</p>\n\n<p>Cheers. </p>",
  "messages": [
    {
      "id": 776709,
      "postDate": "2020-03-17T15:27:08.870Z",
      "content": "<p>Thanks at Bengali.ai and kaggle for hosting this interesting competition as it was more to it than simple ensembling of computer vision models. Thanks to the authors and contributors of pytorch, pytorch-lightning, apex and pytorchcv for making my life easier.</p>\n\n<p><strong>Acknowledgements &amp; notes on GM</strong></p>\n\n<p>I am humbled to finally become competition grandmaster. I learned a tremendous amount of tricks in the recent two years and it would not have been able if it wasn’t for all this generous sharing of top solutions and extremely talented teammates I was lucky to have along the way. </p>\n\n<h3>Short Summary</h3>\n\n<p>My solution is a simple bag (different seeds) of the same model seresnext50 with custom head. Major boost in LB score came from redesign of consonant diacritic target.  Minor LB improvement from postprocessing tweaks such as thresholding and finding closest train examples. </p>\n\n<h3>Preprocessing</h3>\n\n<p>I did not resize the image, the only thing I did was normalize each image by its mean and std, since I experienced a good regularization from that in previous competitions.</p>\n\n<h3>Architecture &amp; Training</h3>\n\n<p><strong>Backbone</strong>\nI used a plain seresnext50 but adjusted the very first layer to replace resizing the image and account for single channel. I did that by changing input units and reducing stride from (2,2) to (1,2).</p>\n\n<p>```\nfrom pytorchcv.model_provider import get_model as ptcv_get_model</p>\n\n<p>backbone = ptcv_get_model('seresnext50_32x4d', pretrained=True)\nbackbone = backbone.features\nbackbone.init_block.conv = nn.Conv2d(1, 64, kernel_size=(7, 7), stride=(1, 2))\n```</p>\n\n<p><strong>neck</strong></p>\n\n<p>same as in maciejsypetkowski solution <a href=\"https://www.kaggle.com/c/recursion-cellular-image-classification/discussion/110543\">https://www.kaggle.com/c/recursion-cellular-image-classification/discussion/110543</a> </p>\n\n<p><strong>heads</strong></p>\n\n<ul>\n<li>3 heads, each for consonant, vowel and root and </li>\n<li>auxiliary head for grapheme with arccos loss</li>\n</ul>\n\n<p>So in total my used architecture looks like this:</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1424766%2F3e968cbc82806b9211bf7609b662a18a%2Fbengali%20arch.001.jpeg?generation=1584456769762003&amp;alt=media\" alt=\"\"></p>\n\n<p><strong>Training</strong></p>\n\n<ul>\n<li>train on all data</li>\n<li>4 cycle cosine annealing with augmentation increase\n<ul><li>cycle 1 &amp; 2 cutmix</li>\n<li>cycle 3 &amp; 4 cutmix, cutout, scale, translate, rotate</li></ul></li>\n<li>Adam </li>\n</ul>\n\n<p><strong>Losses</strong></p>\n\n<ul>\n<li>root loss: CrossEntropy</li>\n<li>consonant loss: Multi label Binary Crossentropy</li>\n<li>vowel loss : CrossEntropy</li>\n<li><p>grapheme loss: ArcCos + CrossEntropy</p></li>\n<li><p>total loss = root loss + consonant loss + vowel loss + grapheme loss</p></li>\n</ul>\n\n<h3>The Magic (jôfôla)</h3>\n\n<p>Now to the key aspect of my solution. I realised all models perform poorly on consonant diacritic 3 (I abbreviate with cons3 in the following) and especially confuse it with cons4. And with poorly I mean game changing poorly: 89% recall compared to 99% for the other consonant classes. At first I just fought the symptoms (instead of the root cause) as several other teams did, e.g. by transferring low confident predicts from high frequent class cons4 to low frequent class cons3 and hence leveraging the definition of macro recall metric and that gave a good LB boost (0.9900 -&gt; 0.9915). That might have been enough for gold but I wanted to improve my chances. So I wanted to understand why the model did so poor on this one and read through discussions, wikipedia and other domain knowledge related things. Luckily, I found 2 things:\n- Wikipedia: As the last member of a conjunct, য jô appears as a wavy vertical line (called যফলা jôfôla) to the right of the previous member: ক্য \"kyô\" খ্য \"khyô\" গ্য \"gyô\" ঘ্য \"ghyô\" etc. In some fonts, certain conjuncts with যফলা jôfôla appear using special fused forms: দ্য \"dyô\" ন্য \"nyô\" শ্য \"shyô\" ষ্য \"ṣyô\" স্য \"syô\" হ্য \"hyô\".</p>\n\n<ul>\n<li>The consonant diacritics are itself composed from lower level components\n<ul><li>0: []</li>\n<li>1:  ['ঁ']</li>\n<li>2: ['র', '্']</li>\n<li>3: ['র', '্', 'য']</li>\n<li>4: ['্', 'য']</li>\n<li>5:['্', 'র']</li>\n<li>6: ['্', 'র', '্', 'য']</li></ul></li>\n</ul>\n\n<p>So I realized that not only 3,4 and 6 have a jô component which might have the special jôfôla form but also that the special forms are in train for <strong>only</strong> cons4 but <strong>not</strong> for cons3, and thats why the models are so bad on cons3. My solution was the following. I recoded the cons classes into a multilabel classification problem using the following function:</p>\n\n<p>```\ncons_components = ['ঁ', 'য', 'র','্']</p>\n\n<p>def is_sub(sub, lst):\n   ln = len(sub)\n   return any(lst[i: i + ln] == sub for i in range(len(lst) - ln + 1))</p>\n\n<p>def label2label_v2(label):\n   res = np.zeros((6,),dtype=int)\n   components = cons2components[label]\n   for i,item in enumerate(cons_components):\n       if item in components:\n           res[i] = 1\n   if is_sub(['্', 'য'], components):\n       res[4] = 1\n   if is_sub(['্', 'র'], components):\n       res[5] = 1\n   return res</p>\n\n<p>```</p>\n\n<p>That function re-codes the 7 consonant class labels into multilabel 6 dimensional vectors</p>\n\n<p><code>\n0 -&gt; [0, 0, 0, 0, 0, 0]\n1 -&gt; [1, 0, 0, 0, 0, 0]\n2 -&gt; [0, 0, 1, 1, 0, 0]\n3 -&gt; [0, 1, 1, 1, 1, 0]\n4 -&gt; [0, 1, 0, 1, 1, 0]\n5 -&gt; [0, 0, 1, 1, 0, 1]\n6 -&gt; [0, 1, 1, 1, 1, 1]\n</code></p>\n\n<p>That schema not only enabled to learn special forms of cons3 from cons4 but also enables to tune via thresholding (i.e. when to set a continuous prediction to 0 or 1) how the model separates the 6 classes. E.g. by setting a very low threshold on the 2nd logit more predictions would be set to 1 and hence more predictions would move from cons4 to cons3 as these two classes only differ in logit2. Purely changing the target for consonant diacritic in this way improved Public LB from 0.990 to 0.993 !</p>\n\n<h3>Postprocessing</h3>\n\n<ul>\n<li>I adjusted threshold for binarizing consonant diacritic. Best was a threshold of 0.1 for all logits classes. I tried not dare to temper which each logit threshold individually as I was afraid of overfitting. Result on Public LB: 9930 -&gt; 9937 </li>\n<li>I used cosine similarity to find closest train sample for root and vowel when top1 probability &lt; 90% and use its label. bestfitting did the same in the 1st place solution of Human Protein competition) as can be found under <a href=\"https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/78109\">https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/78109</a> Result on Public LB: 9937 -&gt; 9942 </li>\n</ul>\n\n<h3>Choosing subs</h3>\n\n<p>I tempered around with binarization and metric learning thresholds and purely checked LB hence I was aware of a high risk to potentially overfit to LB. So I chose my best sub and a failsafe sub that I felt was the least tempered with but still doing ok. (Both were gold on private at the end)</p>\n\n<p>Cheers. </p>",
      "rawMarkdown": "Thanks at Bengali.ai and kaggle for hosting this interesting competition as it was more to it than simple ensembling of computer vision models. Thanks to the authors and contributors of pytorch, pytorch-lightning, apex and pytorchcv for making my life easier.\n\n**Acknowledgements &amp; notes on GM**\n\nI am humbled to finally become competition grandmaster. I learned a tremendous amount of tricks in the recent two years and it would not have been able if it wasn’t for all this generous sharing of top solutions and extremely talented teammates I was lucky to have along the way. \n\n\n### Short Summary\n\nMy solution is a simple bag (different seeds) of the same model seresnext50 with custom head. Major boost in LB score came from redesign of consonant diacritic target.  Minor LB improvement from postprocessing tweaks such as thresholding and finding closest train examples. \n\n\n### Preprocessing\n\nI did not resize the image, the only thing I did was normalize each image by its mean and std, since I experienced a good regularization from that in previous competitions.\n\n### Architecture &amp; Training\n\n**Backbone**\nI used a plain seresnext50 but adjusted the very first layer to replace resizing the image and account for single channel. I did that by changing input units and reducing stride from (2,2) to (1,2).\n\n```\nfrom pytorchcv.model_provider import get_model as ptcv_get_model\n\nbackbone = ptcv_get_model('seresnext50_32x4d', pretrained=True)\nbackbone = backbone.features\nbackbone.init_block.conv = nn.Conv2d(1, 64, kernel_size=(7, 7), stride=(1, 2))\n```\n\n**neck**\n\nsame as in maciejsypetkowski solution https://www.kaggle.com/c/recursion-cellular-image-classification/discussion/110543 \n\n\n**heads**\n\n- 3 heads, each for consonant, vowel and root and \n- auxiliary head for grapheme with arccos loss\n\nSo in total my used architecture looks like this:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1424766%2F3e968cbc82806b9211bf7609b662a18a%2Fbengali%20arch.001.jpeg?generation=1584456769762003&amp;alt=media)\n\n\n**Training**\n\n- train on all data\n- 4 cycle cosine annealing with augmentation increase\n  - cycle 1 &amp; 2 cutmix\n  - cycle 3 &amp; 4 cutmix, cutout, scale, translate, rotate\n- Adam \n\n**Losses**\n\n- root loss: CrossEntropy\n- consonant loss: Multi label Binary Crossentropy\n- vowel loss : CrossEntropy\n- grapheme loss: ArcCos + CrossEntropy\n\n- total loss = root loss + consonant loss + vowel loss + grapheme loss\n\n\n### The Magic (jôfôla) \nNow to the key aspect of my solution. I realised all models perform poorly on consonant diacritic 3 (I abbreviate with cons3 in the following) and especially confuse it with cons4. And with poorly I mean game changing poorly: 89% recall compared to 99% for the other consonant classes. At first I just fought the symptoms (instead of the root cause) as several other teams did, e.g. by transferring low confident predicts from high frequent class cons4 to low frequent class cons3 and hence leveraging the definition of macro recall metric and that gave a good LB boost (0.9900 -&gt; 0.9915). That might have been enough for gold but I wanted to improve my chances. So I wanted to understand why the model did so poor on this one and read through discussions, wikipedia and other domain knowledge related things. Luckily, I found 2 things:\n- Wikipedia: As the last member of a conjunct, য jô appears as a wavy vertical line (called যফলা jôfôla) to the right of the previous member: ক্য \"kyô\" খ্য \"khyô\" গ্য \"gyô\" ঘ্য \"ghyô\" etc. In some fonts, certain conjuncts with যফলা jôfôla appear using special fused forms: দ্য \"dyô\" ন্য \"nyô\" শ্য \"shyô\" ষ্য \"ṣyô\" স্য \"syô\" হ্য \"hyô\".\n \n- The consonant diacritics are itself composed from lower level components\n  - 0: []\n  - 1:  ['ঁ']\n  - 2: ['র', '্']\n  - 3: ['র', '্', 'য']\n  - 4: ['্', 'য']\n  - 5:['্', 'র']\n  - 6: ['্', 'র', '্', 'য']\n \nSo I realized that not only 3,4 and 6 have a jô component which might have the special jôfôla form but also that the special forms are in train for **only** cons4 but **not** for cons3, and thats why the models are so bad on cons3. My solution was the following. I recoded the cons classes into a multilabel classification problem using the following function:\n\n```\ncons_components = ['ঁ', 'য', 'র','্']\n\ndef is_sub(sub, lst):\n   ln = len(sub)\n   return any(lst[i: i + ln] == sub for i in range(len(lst) - ln + 1))\n\ndef label2label_v2(label):\n   res = np.zeros((6,),dtype=int)\n   components = cons2components[label]\n   for i,item in enumerate(cons_components):\n       if item in components:\n           res[i] = 1\n   if is_sub(['্', 'য'], components):\n       res[4] = 1\n   if is_sub(['্', 'র'], components):\n       res[5] = 1\n   return res\n\n```\n\nThat function re-codes the 7 consonant class labels into multilabel 6 dimensional vectors\n\n```\n0 -&gt; [0, 0, 0, 0, 0, 0]\n1 -&gt; [1, 0, 0, 0, 0, 0]\n2 -&gt; [0, 0, 1, 1, 0, 0]\n3 -&gt; [0, 1, 1, 1, 1, 0]\n4 -&gt; [0, 1, 0, 1, 1, 0]\n5 -&gt; [0, 0, 1, 1, 0, 1]\n6 -&gt; [0, 1, 1, 1, 1, 1]\n```\n\nThat schema not only enabled to learn special forms of cons3 from cons4 but also enables to tune via thresholding (i.e. when to set a continuous prediction to 0 or 1) how the model separates the 6 classes. E.g. by setting a very low threshold on the 2nd logit more predictions would be set to 1 and hence more predictions would move from cons4 to cons3 as these two classes only differ in logit2. Purely changing the target for consonant diacritic in this way improved Public LB from 0.990 to 0.993 !\n\n\n### Postprocessing\n\n- I adjusted threshold for binarizing consonant diacritic. Best was a threshold of 0.1 for all logits classes. I tried not dare to temper which each logit threshold individually as I was afraid of overfitting. Result on Public LB: 9930 -&gt; 9937 \n- I used cosine similarity to find closest train sample for root and vowel when top1 probability &lt; 90% and use its label. bestfitting did the same in the 1st place solution of Human Protein competition) as can be found under https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/78109 Result on Public LB: 9937 -&gt; 9942 \n\n\n### Choosing subs\nI tempered around with binarization and metric learning thresholds and purely checked LB hence I was aware of a high risk to potentially overfit to LB. So I chose my best sub and a failsafe sub that I felt was the least tempered with but still doing ok. (Both were gold on private at the end)\n\nCheers. \n",
      "votes": 90
    },
    {
      "id": 780997,
      "postDate": "2020-03-20T19:55:32.437Z",
      "content": "<p>Congratulations! So many admires to you for achieving 3*GM in such a short time! You're my role model now.😊 </p>",
      "rawMarkdown": "Congratulations! So many admires to you for achieving 3*GM in such a short time! You're my role model now.😊 ",
      "votes": 1
    },
    {
      "id": 779331,
      "postDate": "2020-03-19T08:11:42.943Z",
      "content": "<p>Thank you for sharing your magic, and Congrats on triple Grand Master!</p>",
      "rawMarkdown": "Thank you for sharing your magic, and Congrats on triple Grand Master!",
      "votes": 1
    },
    {
      "id": 778614,
      "postDate": "2020-03-18T15:40:12.820Z",
      "content": "<p>Congratulation for the solo gold and being competition GM! You made it in only 5 weeks by solo, that's a great job!\nI learned a lot from your post and I have a question about your re-codes labels.\nIs your re-codes labels doing better than original ones for unseen graphemes? Have you tried to submit it without any post-processing?</p>",
      "rawMarkdown": "Congratulation for the solo gold and being competition GM! You made it in only 5 weeks by solo, that's a great job!\nI learned a lot from your post and I have a question about your re-codes labels.\nIs your re-codes labels doing better than original ones for unseen graphemes? Have you tried to submit it without any post-processing?",
      "votes": 1,
      "replies": [
        {
          "id": 778741,
          "postDate": "2020-03-18T17:28:09.250Z",
          "content": "<p>Private score (roughly) 0.92 -&gt; 0.95,  so a huge improvement</p>",
          "rawMarkdown": "Private score (roughly) 0.92 -&gt; 0.95,  so a huge improvement"
        },
        {
          "id": 778748,
          "postDate": "2020-03-18T17:36:19.840Z",
          "content": "<p>That's huge!\nI'll try it when I have my GPU free ;)</p>",
          "rawMarkdown": "That's huge!\nI'll try it when I have my GPU free ;)"
        }
      ]
    },
    {
      "id": 777874,
      "postDate": "2020-03-18T02:04:16.210Z",
      "content": "<p>Congratulations on  triple GM!</p>",
      "rawMarkdown": "Congratulations on  triple GM!",
      "votes": 1
    },
    {
      "id": 777839,
      "postDate": "2020-03-18T01:42:03.020Z",
      "content": "<p>Congratulations on becoming the 180th Competitions GM😏 😏 </p>",
      "rawMarkdown": "Congratulations on becoming the 180th Competitions GM😏 😏 ",
      "votes": 1
    },
    {
      "id": 777820,
      "postDate": "2020-03-18T01:08:15.080Z",
      "content": "<p>Congrats to triple GM, consonant labeling method is interesting!</p>",
      "rawMarkdown": "Congrats to triple GM, consonant labeling method is interesting!",
      "votes": 1
    },
    {
      "id": 777633,
      "postDate": "2020-03-17T20:35:20.747Z",
      "content": "<p>Congratulations <a href=\"/christofhenkel\">@christofhenkel</a> for the competition and for the Grand-Master status. \nCould you please elaborate and provide some reference links to help understand the following in <strong>Choosing subs</strong>:\n&gt; I tempered around with binarization and metric learning thresholds and purely checked LB hence I was aware of a high risk to potentially overfit to LB...</p>\n\n<p>This could be really helpful for someone like me who has no idea on <em>how to select solutions for final submission.</em></p>",
      "rawMarkdown": "Congratulations @christofhenkel for the competition and for the Grand-Master status. \nCould you please elaborate and provide some reference links to help understand the following in **Choosing subs**:\n&gt; I tempered around with binarization and metric learning thresholds and purely checked LB hence I was aware of a high risk to potentially overfit to LB...\n\nThis could be really helpful for someone like me who has no idea on _how to select solutions for final submission._",
      "votes": 1,
      "replies": [
        {
          "id": 778048,
          "postDate": "2020-03-18T05:29:04.777Z",
          "content": "<p>Every time you use the Public Leaderboard to evaluate if something is working or not, you fit on it and that's risky, since the behavior might not hold on private. So as the second submission, I tried to chose a less fitted to leaderboard one.</p>",
          "rawMarkdown": "Every time you use the Public Leaderboard to evaluate if something is working or not, you fit on it and that's risky, since the behavior might not hold on private. So as the second submission, I tried to chose a less fitted to leaderboard one."
        }
      ]
    },
    {
      "id": 776844,
      "postDate": "2020-03-17T17:12:12.953Z",
      "content": "<p>Brilliant solution! Congrats on gold medal and grandmaster status.</p>\n\n<p>Re-coding the consonants was smart. Cosine similarity is a great trick. That's a great way to use GPU multiplication and/or RAPIDS cuML kNN.</p>",
      "rawMarkdown": "Brilliant solution! Congrats on gold medal and grandmaster status.\n\nRe-coding the consonants was smart. Cosine similarity is a great trick. That's a great way to use GPU multiplication and/or RAPIDS cuML kNN.",
      "votes": 1
    },
    {
      "id": 776794,
      "postDate": "2020-03-17T16:33:08.537Z",
      "content": "<p>Congrats on the result and the method.  Well deserved GM title in prime!</p>",
      "rawMarkdown": "Congrats on the result and the method.  Well deserved GM title in prime!\n",
      "votes": 1
    },
    {
      "id": 776761,
      "postDate": "2020-03-17T16:10:44.567Z",
      "content": "<p>Wow! That's a really interesting and creative way to handle consonant 3 and 6.</p>",
      "rawMarkdown": "Wow! That's a really interesting and creative way to handle consonant 3 and 6.",
      "votes": 1
    },
    {
      "id": 776717,
      "postDate": "2020-03-17T15:33:32.863Z",
      "content": "<p>Handling C=3 was key :) Congrats!</p>",
      "rawMarkdown": "Handling C=3 was key :) Congrats!",
      "votes": 1
    },
    {
      "id": 776712,
      "postDate": "2020-03-17T15:32:21.563Z",
      "content": "<p>Congrats on reaching GM!</p>",
      "rawMarkdown": "Congrats on reaching GM!",
      "votes": 1
    },
    {
      "id": 778050,
      "postDate": "2020-03-18T05:31:49.187Z",
      "content": "<p>Congratulations, Small question about components part, Did you set the threshold search local CV？and if we get vectors other than those seven ( for example : result =[0, 1, 1, 1, 0, 0]), what should we do?</p>",
      "rawMarkdown": "Congratulations, Small question about components part, Did you set the threshold search local CV？and if we get vectors other than those seven ( for example : result =[0, 1, 1, 1, 0, 0]), what should we do?",
      "votes": 2,
      "replies": [
        {
          "id": 778055,
          "postDate": "2020-03-18T05:41:06.903Z",
          "content": "<blockquote>\n  <p>Did you set the threshold search local CV？</p>\n</blockquote>\n\n<p>Yes, but I also probed leaderboard a bit for that. </p>\n\n<blockquote>\n  <p>and if we get vectors other than those seven, what should we do?</p>\n</blockquote>\n\n<p>luckily there were not many of that cases. (Which was also a good check if the model is working properly) You could take the closest of the 7 using some distance metric (e.g. cosine distance). Here it made more sense to directly label as C=3, due to the design of the competition metric</p>",
          "rawMarkdown": "&gt; Did you set the threshold search local CV？\n\nYes, but I also probed leaderboard a bit for that. \n\n&gt; and if we get vectors other than those seven, what should we do?\n\nluckily there were not many of that cases. (Which was also a good check if the model is working properly) You could take the closest of the 7 using some distance metric (e.g. cosine distance). Here it made more sense to directly label as C=3, due to the design of the competition metric",
          "votes": 3
        },
        {
          "id": 778071,
          "postDate": "2020-03-18T05:57:05.337Z",
          "content": "<p>Ok, Thanks for your reply</p>",
          "rawMarkdown": "Ok, Thanks for your reply"
        }
      ]
    },
    {
      "id": 777986,
      "postDate": "2020-03-18T04:00:30.443Z",
      "content": "<p><a href=\"/christofhenkel\">@christofhenkel</a> Congrats to beoming triple GM, really ingenious and problem-specific solution! I remember struggling up the LB and seeing you a few places before me every step of the way during the latter quarter of the competition; chasing \"grandpa\" had been great motivation hahaha.</p>\n\n<p>3 questions\n1) When you said normalizing each image by its mean and std, do you mean you apply (1) the same statistics computed on all training samples or (2) normalize every single image based on its mean and std (different for each image, kind of like InstanceNorm) ?\n2) I have never tried model with two heads (with 2 losses, like CE+arcface in this way). Do you usually get good feedbacks from this design? I have been hesitant because it's hard to find correct weighting of the losses. I wonder what is your chosen weights and how you usually choose them.\n3) Because the model is multi-output, do you leverage the multi-output for the models' prediction before post-processing? In other words, do you only use the RCV prediction or somehow combine that with grapheme prediction for inference?</p>\n\n<p>Thanks in advance</p>",
      "rawMarkdown": "@christofhenkel Congrats to beoming triple GM, really ingenious and problem-specific solution! I remember struggling up the LB and seeing you a few places before me every step of the way during the latter quarter of the competition; chasing \"grandpa\" had been great motivation hahaha.\n\n3 questions\n1) When you said normalizing each image by its mean and std, do you mean you apply (1) the same statistics computed on all training samples or (2) normalize every single image based on its mean and std (different for each image, kind of like InstanceNorm) ?\n2) I have never tried model with two heads (with 2 losses, like CE+arcface in this way). Do you usually get good feedbacks from this design? I have been hesitant because it's hard to find correct weighting of the losses. I wonder what is your chosen weights and how you usually choose them.\n3) Because the model is multi-output, do you leverage the multi-output for the models' prediction before post-processing? In other words, do you only use the RCV prediction or somehow combine that with grapheme prediction for inference?\n\nThanks in advance",
      "votes": 2,
      "replies": [
        {
          "id": 778045,
          "postDate": "2020-03-18T05:20:14.193Z",
          "content": "<p>1)  normalize every single image based on its mean and std, which I thought makes sense, since you couldn't peak into the test images to norm over all samples</p>\n\n<p>2) Several heads  or multi losses normally work well for me. I normally start with very simple weights. and then tune them a bit. I use a simple scaling factor to get the losses to same scale, when using different losses. </p>\n\n<p>3) Sometimes you can leverage from multi output. For example average the predictions, but here I did though using the grapheme prediction might hurt the models performance on unseen graphemes, so I did not try. </p>",
          "rawMarkdown": "1)  normalize every single image based on its mean and std, which I thought makes sense, since you couldn't peak into the test images to norm over all samples\n\n2) Several heads  or multi losses normally work well for me. I normally start with very simple weights. and then tune them a bit. I use a simple scaling factor to get the losses to same scale, when using different losses. \n\n3) Sometimes you can leverage from multi output. For example average the predictions, but here I did though using the grapheme prediction might hurt the models performance on unseen graphemes, so I did not try. ",
          "votes": 2
        }
      ]
    },
    {
      "id": 777698,
      "postDate": "2020-03-17T21:29:37.420Z",
      "content": "<p>Congrats to triple GM!</p>",
      "rawMarkdown": "Congrats to triple GM!",
      "votes": 2
    },
    {
      "id": 776752,
      "postDate": "2020-03-17T15:59:22.947Z",
      "content": "<p>Congrats Dieter!! You really are an inspiration to many amateurs like me. Now that you are Competition GM(and 3xGM), what would be your (<strong>realistic</strong>) advice for someone like me who wants to get better at Kaggle and become GM someday?(What is not realistic? -&gt; don't be afraid kind of thing 😂 )</p>",
      "rawMarkdown": "Congrats Dieter!! You really are an inspiration to many amateurs like me. Now that you are Competition GM(and 3xGM), what would be your (**realistic**) advice for someone like me who wants to get better at Kaggle and become GM someday?(What is not realistic? -&gt; don't be afraid kind of thing 😂 )",
      "votes": 2,
      "replies": [
        {
          "id": 778062,
          "postDate": "2020-03-18T05:47:30.737Z",
          "content": "<p>I think the young imposter in this <a href=\"https://www.youtube.com/watch?v=Q0_Xajic_9U&amp;feature=youtu.be\">video</a> gives some tips. Try to learn long term and as broad as possible (NLP, computer vision, tabular data) you never know when you could use a trick from a different field.</p>",
          "rawMarkdown": "I think the young imposter in this [video](https://www.youtube.com/watch?v=Q0_Xajic_9U&amp;feature=youtu.be ) gives some tips. Try to learn long term and as broad as possible (NLP, computer vision, tabular data) you never know when you could use a trick from a different field.",
          "votes": 6
        }
      ]
    },
    {
      "id": 780925,
      "postDate": "2020-03-20T18:36:06.823Z",
      "content": "<p>Congrats, so surprised to see that you got large improvement by just using multilabel for consonant . I tried similar thing for root but not working. Also glad to see that the old gold solution from former competitions also did well in this competition. </p>",
      "rawMarkdown": "Congrats, so surprised to see that you got large improvement by just using multilabel for consonant . I tried similar thing for root but not working. Also glad to see that the old gold solution from former competitions also did well in this competition. "
    },
    {
      "id": 778210,
      "postDate": "2020-03-18T08:28:54.523Z",
      "content": "<p>Congratulations!</p>",
      "rawMarkdown": "Congratulations!"
    },
    {
      "id": 782222,
      "postDate": "2020-03-22T03:50:58.050Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 781032,
      "postDate": "2020-03-20T20:58:48.410Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 780878,
      "postDate": "2020-03-20T17:41:06.083Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 777507,
      "postDate": "2020-03-17T18:21:34.353Z",
      "rawMarkdown": "",
      "votes": -2,
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 780997,
      "author_name": "Helen",
      "author_url": "",
      "post_date": "2020-03-20T19:55:32.437000",
      "content": "<p>Congratulations! So many admires to you for achieving 3*GM in such a short time! You're my role model now.😊 </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 779331,
      "author_name": "Tawara",
      "author_url": "",
      "post_date": "2020-03-19T08:11:42.943000",
      "content": "<p>Thank you for sharing your magic, and Congrats on triple Grand Master!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 778614,
      "author_name": "Qishen Ha",
      "author_url": "",
      "post_date": "2020-03-18T15:40:12.820000",
      "content": "<p>Congratulation for the solo gold and being competition GM! You made it in only 5 weeks by solo, that's a great job!\nI learned a lot from your post and I have a question about your re-codes labels.\nIs your re-codes labels doing better than original ones for unseen graphemes? Have you tried to submit it without any post-processing?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 778741,
          "author_name": "Dieter",
          "author_url": "",
          "post_date": "2020-03-18T17:28:09.250000",
          "content": "<p>Private score (roughly) 0.92 -&gt; 0.95,  so a huge improvement</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 778748,
          "author_name": "Qishen Ha",
          "author_url": "",
          "post_date": "2020-03-18T17:36:19.840000",
          "content": "<p>That's huge!\nI'll try it when I have my GPU free ;)</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 777874,
      "author_name": "zr",
      "author_url": "",
      "post_date": "2020-03-18T02:04:16.210000",
      "content": "<p>Congratulations on  triple GM!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 777839,
      "author_name": "Gary",
      "author_url": "",
      "post_date": "2020-03-18T01:42:03.020000",
      "content": "<p>Congratulations on becoming the 180th Competitions GM😏 😏 </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 777820,
      "author_name": "corochann",
      "author_url": "",
      "post_date": "2020-03-18T01:08:15.080000",
      "content": "<p>Congrats to triple GM, consonant labeling method is interesting!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 777633,
      "author_name": "Tushar",
      "author_url": "",
      "post_date": "2020-03-17T20:35:20.747000",
      "content": "<p>Congratulations <a href=\"/christofhenkel\">@christofhenkel</a> for the competition and for the Grand-Master status. \nCould you please elaborate and provide some reference links to help understand the following in <strong>Choosing subs</strong>:\n&gt; I tempered around with binarization and metric learning thresholds and purely checked LB hence I was aware of a high risk to potentially overfit to LB...</p>\n\n<p>This could be really helpful for someone like me who has no idea on <em>how to select solutions for final submission.</em></p>",
      "votes": 1,
      "replies": [
        {
          "id": 778048,
          "author_name": "Dieter",
          "author_url": "",
          "post_date": "2020-03-18T05:29:04.777000",
          "content": "<p>Every time you use the Public Leaderboard to evaluate if something is working or not, you fit on it and that's risky, since the behavior might not hold on private. So as the second submission, I tried to chose a less fitted to leaderboard one.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 776844,
      "author_name": "Chris Deotte",
      "author_url": "",
      "post_date": "2020-03-17T17:12:12.953000",
      "content": "<p>Brilliant solution! Congrats on gold medal and grandmaster status.</p>\n\n<p>Re-coding the consonants was smart. Cosine similarity is a great trick. That's a great way to use GPU multiplication and/or RAPIDS cuML kNN.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 776794,
      "author_name": "CPMP",
      "author_url": "",
      "post_date": "2020-03-17T16:33:08.537000",
      "content": "<p>Congrats on the result and the method.  Well deserved GM title in prime!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 776761,
      "author_name": "Dipam Chakraborty",
      "author_url": "",
      "post_date": "2020-03-17T16:10:44.567000",
      "content": "<p>Wow! That's a really interesting and creative way to handle consonant 3 and 6.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 776717,
      "author_name": "Psi",
      "author_url": "",
      "post_date": "2020-03-17T15:33:32.863000",
      "content": "<p>Handling C=3 was key :) Congrats!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 776712,
      "author_name": "GreatGameDota",
      "author_url": "",
      "post_date": "2020-03-17T15:32:21.563000",
      "content": "<p>Congrats on reaching GM!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 778050,
      "author_name": "He",
      "author_url": "",
      "post_date": "2020-03-18T05:31:49.187000",
      "content": "<p>Congratulations, Small question about components part, Did you set the threshold search local CV？and if we get vectors other than those seven ( for example : result =[0, 1, 1, 1, 0, 0]), what should we do?</p>",
      "votes": 2,
      "replies": [
        {
          "id": 778055,
          "author_name": "Dieter",
          "author_url": "",
          "post_date": "2020-03-18T05:41:06.903000",
          "content": "<blockquote>\n  <p>Did you set the threshold search local CV？</p>\n</blockquote>\n\n<p>Yes, but I also probed leaderboard a bit for that. </p>\n\n<blockquote>\n  <p>and if we get vectors other than those seven, what should we do?</p>\n</blockquote>\n\n<p>luckily there were not many of that cases. (Which was also a good check if the model is working properly) You could take the closest of the 7 using some distance metric (e.g. cosine distance). Here it made more sense to directly label as C=3, due to the design of the competition metric</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 778071,
          "author_name": "He",
          "author_url": "",
          "post_date": "2020-03-18T05:57:05.337000",
          "content": "<p>Ok, Thanks for your reply</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 777986,
      "author_name": "Nicholas Lyu",
      "author_url": "",
      "post_date": "2020-03-18T04:00:30.443000",
      "content": "<p><a href=\"/christofhenkel\">@christofhenkel</a> Congrats to beoming triple GM, really ingenious and problem-specific solution! I remember struggling up the LB and seeing you a few places before me every step of the way during the latter quarter of the competition; chasing \"grandpa\" had been great motivation hahaha.</p>\n\n<p>3 questions\n1) When you said normalizing each image by its mean and std, do you mean you apply (1) the same statistics computed on all training samples or (2) normalize every single image based on its mean and std (different for each image, kind of like InstanceNorm) ?\n2) I have never tried model with two heads (with 2 losses, like CE+arcface in this way). Do you usually get good feedbacks from this design? I have been hesitant because it's hard to find correct weighting of the losses. I wonder what is your chosen weights and how you usually choose them.\n3) Because the model is multi-output, do you leverage the multi-output for the models' prediction before post-processing? In other words, do you only use the RCV prediction or somehow combine that with grapheme prediction for inference?</p>\n\n<p>Thanks in advance</p>",
      "votes": 2,
      "replies": [
        {
          "id": 778045,
          "author_name": "Dieter",
          "author_url": "",
          "post_date": "2020-03-18T05:20:14.193000",
          "content": "<p>1)  normalize every single image based on its mean and std, which I thought makes sense, since you couldn't peak into the test images to norm over all samples</p>\n\n<p>2) Several heads  or multi losses normally work well for me. I normally start with very simple weights. and then tune them a bit. I use a simple scaling factor to get the losses to same scale, when using different losses. </p>\n\n<p>3) Sometimes you can leverage from multi output. For example average the predictions, but here I did though using the grapheme prediction might hurt the models performance on unseen graphemes, so I did not try. </p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 777698,
      "author_name": "Max Jeblick",
      "author_url": "",
      "post_date": "2020-03-17T21:29:37.420000",
      "content": "<p>Congrats to triple GM!</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 776752,
      "author_name": "Bibek",
      "author_url": "",
      "post_date": "2020-03-17T15:59:22.947000",
      "content": "<p>Congrats Dieter!! You really are an inspiration to many amateurs like me. Now that you are Competition GM(and 3xGM), what would be your (<strong>realistic</strong>) advice for someone like me who wants to get better at Kaggle and become GM someday?(What is not realistic? -&gt; don't be afraid kind of thing 😂 )</p>",
      "votes": 2,
      "replies": [
        {
          "id": 778062,
          "author_name": "Dieter",
          "author_url": "",
          "post_date": "2020-03-18T05:47:30.737000",
          "content": "<p>I think the young imposter in this <a href=\"https://www.youtube.com/watch?v=Q0_Xajic_9U&amp;feature=youtu.be\">video</a> gives some tips. Try to learn long term and as broad as possible (NLP, computer vision, tabular data) you never know when you could use a trick from a different field.</p>",
          "votes": 6,
          "replies": []
        }
      ]
    },
    {
      "id": 780925,
      "author_name": "Morphy",
      "author_url": "",
      "post_date": "2020-03-20T18:36:06.823000",
      "content": "<p>Congrats, so surprised to see that you got large improvement by just using multilabel for consonant . I tried similar thing for root but not working. Also glad to see that the old gold solution from former competitions also did well in this competition. </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 778210,
      "author_name": "Kranti Kumar",
      "author_url": "",
      "post_date": "2020-03-18T08:28:54.523000",
      "content": "<p>Congratulations!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 782222,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-03-22T03:50:58.050000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 781032,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-03-20T20:58:48.410000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 780878,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-03-20T17:41:06.083000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 777507,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-03-17T18:21:34.353000",
      "content": "",
      "votes": -2,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "776709": "Thanks at Bengali.ai and kaggle for hosting this interesting competition as it was more to it than simple ensembling of computer vision models. Thanks to the authors and contributors of pytorch, pytorch-lightning, apex and pytorchcv for making my life easier.\n\n**Acknowledgements &amp; notes on GM**\n\nI am humbled to finally become competition grandmaster. I learned a tremendous amount of tricks in the recent two years and it would not have been able if it wasn’t for all this generous sharing of top solutions and extremely talented teammates I was lucky to have along the way. \n\n\n### Short Summary\n\nMy solution is a simple bag (different seeds) of the same model seresnext50 with custom head. Major boost in LB score came from redesign of consonant diacritic target.  Minor LB improvement from postprocessing tweaks such as thresholding and finding closest train examples. \n\n\n### Preprocessing\n\nI did not resize the image, the only thing I did was normalize each image by its mean and std, since I experienced a good regularization from that in previous competitions.\n\n### Architecture &amp; Training\n\n**Backbone**\nI used a plain seresnext50 but adjusted the very first layer to replace resizing the image and account for single channel. I did that by changing input units and reducing stride from (2,2) to (1,2).\n\n```\nfrom pytorchcv.model_provider import get_model as ptcv_get_model\n\nbackbone = ptcv_get_model('seresnext50_32x4d', pretrained=True)\nbackbone = backbone.features\nbackbone.init_block.conv = nn.Conv2d(1, 64, kernel_size=(7, 7), stride=(1, 2))\n```\n\n**neck**\n\nsame as in maciejsypetkowski solution https://www.kaggle.com/c/recursion-cellular-image-classification/discussion/110543 \n\n\n**heads**\n\n- 3 heads, each for consonant, vowel and root and \n- auxiliary head for grapheme with arccos loss\n\nSo in total my used architecture looks like this:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1424766%2F3e968cbc82806b9211bf7609b662a18a%2Fbengali%20arch.001.jpeg?generation=1584456769762003&amp;alt=media)\n\n\n**Training**\n\n- train on all data\n- 4 cycle cosine annealing with augmentation increase\n  - cycle 1 &amp; 2 cutmix\n  - cycle 3 &amp; 4 cutmix, cutout, scale, translate, rotate\n- Adam \n\n**Losses**\n\n- root loss: CrossEntropy\n- consonant loss: Multi label Binary Crossentropy\n- vowel loss : CrossEntropy\n- grapheme loss: ArcCos + CrossEntropy\n\n- total loss = root loss + consonant loss + vowel loss + grapheme loss\n\n\n### The Magic (jôfôla) \nNow to the key aspect of my solution. I realised all models perform poorly on consonant diacritic 3 (I abbreviate with cons3 in the following) and especially confuse it with cons4. And with poorly I mean game changing poorly: 89% recall compared to 99% for the other consonant classes. At first I just fought the symptoms (instead of the root cause) as several other teams did, e.g. by transferring low confident predicts from high frequent class cons4 to low frequent class cons3 and hence leveraging the definition of macro recall metric and that gave a good LB boost (0.9900 -&gt; 0.9915). That might have been enough for gold but I wanted to improve my chances. So I wanted to understand why the model did so poor on this one and read through discussions, wikipedia and other domain knowledge related things. Luckily, I found 2 things:\n- Wikipedia: As the last member of a conjunct, য jô appears as a wavy vertical line (called যফলা jôfôla) to the right of the previous member: ক্য \"kyô\" খ্য \"khyô\" গ্য \"gyô\" ঘ্য \"ghyô\" etc. In some fonts, certain conjuncts with যফলা jôfôla appear using special fused forms: দ্য \"dyô\" ন্য \"nyô\" শ্য \"shyô\" ষ্য \"ṣyô\" স্য \"syô\" হ্য \"hyô\".\n \n- The consonant diacritics are itself composed from lower level components\n  - 0: []\n  - 1:  ['ঁ']\n  - 2: ['র', '্']\n  - 3: ['র', '্', 'য']\n  - 4: ['্', 'য']\n  - 5:['্', 'র']\n  - 6: ['্', 'র', '্', 'য']\n \nSo I realized that not only 3,4 and 6 have a jô component which might have the special jôfôla form but also that the special forms are in train for **only** cons4 but **not** for cons3, and thats why the models are so bad on cons3. My solution was the following. I recoded the cons classes into a multilabel classification problem using the following function:\n\n```\ncons_components = ['ঁ', 'য', 'র','্']\n\ndef is_sub(sub, lst):\n   ln = len(sub)\n   return any(lst[i: i + ln] == sub for i in range(len(lst) - ln + 1))\n\ndef label2label_v2(label):\n   res = np.zeros((6,),dtype=int)\n   components = cons2components[label]\n   for i,item in enumerate(cons_components):\n       if item in components:\n           res[i] = 1\n   if is_sub(['্', 'য'], components):\n       res[4] = 1\n   if is_sub(['্', 'র'], components):\n       res[5] = 1\n   return res\n\n```\n\nThat function re-codes the 7 consonant class labels into multilabel 6 dimensional vectors\n\n```\n0 -&gt; [0, 0, 0, 0, 0, 0]\n1 -&gt; [1, 0, 0, 0, 0, 0]\n2 -&gt; [0, 0, 1, 1, 0, 0]\n3 -&gt; [0, 1, 1, 1, 1, 0]\n4 -&gt; [0, 1, 0, 1, 1, 0]\n5 -&gt; [0, 0, 1, 1, 0, 1]\n6 -&gt; [0, 1, 1, 1, 1, 1]\n```\n\nThat schema not only enabled to learn special forms of cons3 from cons4 but also enables to tune via thresholding (i.e. when to set a continuous prediction to 0 or 1) how the model separates the 6 classes. E.g. by setting a very low threshold on the 2nd logit more predictions would be set to 1 and hence more predictions would move from cons4 to cons3 as these two classes only differ in logit2. Purely changing the target for consonant diacritic in this way improved Public LB from 0.990 to 0.993 !\n\n\n### Postprocessing\n\n- I adjusted threshold for binarizing consonant diacritic. Best was a threshold of 0.1 for all logits classes. I tried not dare to temper which each logit threshold individually as I was afraid of overfitting. Result on Public LB: 9930 -&gt; 9937 \n- I used cosine similarity to find closest train sample for root and vowel when top1 probability &lt; 90% and use its label. bestfitting did the same in the 1st place solution of Human Protein competition) as can be found under https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/78109 Result on Public LB: 9937 -&gt; 9942 \n\n\n### Choosing subs\nI tempered around with binarization and metric learning thresholds and purely checked LB hence I was aware of a high risk to potentially overfit to LB. So I chose my best sub and a failsafe sub that I felt was the least tempered with but still doing ok. (Both were gold on private at the end)\n\nCheers. \n",
    "780997": "Congratulations! So many admires to you for achieving 3*GM in such a short time! You're my role model now.😊 ",
    "779331": "Thank you for sharing your magic, and Congrats on triple Grand Master!",
    "778614": "Congratulation for the solo gold and being competition GM! You made it in only 5 weeks by solo, that's a great job!\nI learned a lot from your post and I have a question about your re-codes labels.\nIs your re-codes labels doing better than original ones for unseen graphemes? Have you tried to submit it without any post-processing?",
    "777874": "Congratulations on  triple GM!",
    "777839": "Congratulations on becoming the 180th Competitions GM😏 😏 ",
    "777820": "Congrats to triple GM, consonant labeling method is interesting!",
    "777633": "Congratulations @christofhenkel for the competition and for the Grand-Master status. \nCould you please elaborate and provide some reference links to help understand the following in **Choosing subs**:\n&gt; I tempered around with binarization and metric learning thresholds and purely checked LB hence I was aware of a high risk to potentially overfit to LB...\n\nThis could be really helpful for someone like me who has no idea on _how to select solutions for final submission._",
    "776844": "Brilliant solution! Congrats on gold medal and grandmaster status.\n\nRe-coding the consonants was smart. Cosine similarity is a great trick. That's a great way to use GPU multiplication and/or RAPIDS cuML kNN.",
    "776794": "Congrats on the result and the method.  Well deserved GM title in prime!\n",
    "776761": "Wow! That's a really interesting and creative way to handle consonant 3 and 6.",
    "776717": "Handling C=3 was key :) Congrats!",
    "776712": "Congrats on reaching GM!",
    "778050": "Congratulations, Small question about components part, Did you set the threshold search local CV？and if we get vectors other than those seven ( for example : result =[0, 1, 1, 1, 0, 0]), what should we do?",
    "777986": "@christofhenkel Congrats to beoming triple GM, really ingenious and problem-specific solution! I remember struggling up the LB and seeing you a few places before me every step of the way during the latter quarter of the competition; chasing \"grandpa\" had been great motivation hahaha.\n\n3 questions\n1) When you said normalizing each image by its mean and std, do you mean you apply (1) the same statistics computed on all training samples or (2) normalize every single image based on its mean and std (different for each image, kind of like InstanceNorm) ?\n2) I have never tried model with two heads (with 2 losses, like CE+arcface in this way). Do you usually get good feedbacks from this design? I have been hesitant because it's hard to find correct weighting of the losses. I wonder what is your chosen weights and how you usually choose them.\n3) Because the model is multi-output, do you leverage the multi-output for the models' prediction before post-processing? In other words, do you only use the RCV prediction or somehow combine that with grapheme prediction for inference?\n\nThanks in advance",
    "777698": "Congrats to triple GM!",
    "776752": "Congrats Dieter!! You really are an inspiration to many amateurs like me. Now that you are Competition GM(and 3xGM), what would be your (**realistic**) advice for someone like me who wants to get better at Kaggle and become GM someday?(What is not realistic? -&gt; don't be afraid kind of thing 😂 )",
    "780925": "Congrats, so surprised to see that you got large improvement by just using multilabel for consonant . I tried similar thing for root but not working. Also glad to see that the old gold solution from former competitions also did well in this competition. ",
    "778210": "Congratulations!",
    "782222": "",
    "781032": "",
    "780878": "",
    "777507": ""
  }
}