{
  "id": 136011,
  "title": "6th place solution",
  "url": "/competitions/bengaliai-cv19/discussion/136011",
  "author_name": "Maxwell",
  "post_date": "2020-03-17T03:44:10.579000",
  "votes": 53,
  "comment_count": 17,
  "views": 0,
  "content": "<p>At first, congratulations to the winner <a href=\"/linshokaku\">@linshokaku</a> and all participants who finished this competition. <br>\nWe did not imagine such a big shake ( We estimated that there would not be so much unseen data in private test dataset ).  </p>\n\n<p>Anyway thank you Bengali.AI and kaggle for organizing this competition, also to my teammate <a href=\"/nejumi\">@nejumi</a> !</p>\n\n<hr>\n\n<p>Below figure shows an overview of our solution.  I think our solution is very simple and nothing special. <br>\n( BTW 1st place solution is very unique and impressive for me ;-)  )\nHere I will add naive explanations to some points worth to be noted.  </p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F479538%2F7b541b6993070fbf37181aa631843011%2FBengali_AI.png?generation=1584423861453528&amp;alt=media\" alt=\"\"></p>\n\n<p>NOTE: If you want to see a higher resolution image, please refer this <a href=\"https://speakerdeck.com/hoxomaxwell/kaggle-bengali-dot-ai-6-th-place-solution\">link</a>.</p>\n\n<ol>\n<li>Intensive augmentations\nWe used CutMix, MixUp, Cutout, width/height shift, rotate, Erosion, <a href=\"https://www.kaggle.com/haqishen/gridmask\">GridMask</a> ( credit to <a href=\"/haqishen\">@haqishen</a> ), Zoom... <br>\nCutMix and Cutout worked well for both of us, but in my case, height shift and rotation did bad, maybe due to 137 x 236 size (some graphemes become outside image region :-( ). <br>\n<br></li>\n<li>Co-occurence\nThere are limited combinations between R, V and C in train dataset. But as a host said, in test dataset, some unseen combinations will come. So I did a simple trick; build a model which consider both co-occurence and non-co-occurence of R, V, C to be robust with unseen data in SE-ResNet50 based model. This model has dual paths after SE-ResNet block. One is straightforward path from GeM2D to each top fc layers (R, V and C). And the other is through 2 fc layers (please see above figure).\n<br></li>\n<li>Multi stage learning ( Xentropy, Reduced Focal Loss (credit to <a href=\"/phalanx\">@phalanx</a> ) )\nIn the middle point of this competition, <a href=\"/phalanx\">@phalanx</a> advised us about OHEM in this nice <a href=\"https://www.kaggle.com/c/bengaliai-cv19/discussion/123198#734490\">thread</a>. <br>\nIndeed, our cv score (recall) became better with the same technique(Reduced Focal Loss). \n<br></li>\n<li>Different image size ensemble\nIn some CV competitions, different size image ensemble boosted scores. We thought this will be same in this competition. This ensemble pushes our place from 12th to 6th.\n<br></li>\n<li>Post Processing (big impact for us)\nWe adjusted our prediction to improve evaluation metric in this competition (Recall). Optimization was performed with all classes (168 + 11 + 7). And we used (1/each class counts) as initial values for optimization. The optimal coefficients are almost close to initial values. We were very nervous to use this post processing, but as a result this pushes us to gold range.</li>\n</ol>",
  "messages": [
    {
      "id": 775963,
      "postDate": "2020-03-17T03:44:10.580Z",
      "content": "<p>At first, congratulations to the winner <a href=\"/linshokaku\">@linshokaku</a> and all participants who finished this competition. <br>\nWe did not imagine such a big shake ( We estimated that there would not be so much unseen data in private test dataset ).  </p>\n\n<p>Anyway thank you Bengali.AI and kaggle for organizing this competition, also to my teammate <a href=\"/nejumi\">@nejumi</a> !</p>\n\n<hr>\n\n<p>Below figure shows an overview of our solution.  I think our solution is very simple and nothing special. <br>\n( BTW 1st place solution is very unique and impressive for me ;-)  )\nHere I will add naive explanations to some points worth to be noted.  </p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F479538%2F7b541b6993070fbf37181aa631843011%2FBengali_AI.png?generation=1584423861453528&amp;alt=media\" alt=\"\"></p>\n\n<p>NOTE: If you want to see a higher resolution image, please refer this <a href=\"https://speakerdeck.com/hoxomaxwell/kaggle-bengali-dot-ai-6-th-place-solution\">link</a>.</p>\n\n<ol>\n<li>Intensive augmentations\nWe used CutMix, MixUp, Cutout, width/height shift, rotate, Erosion, <a href=\"https://www.kaggle.com/haqishen/gridmask\">GridMask</a> ( credit to <a href=\"/haqishen\">@haqishen</a> ), Zoom... <br>\nCutMix and Cutout worked well for both of us, but in my case, height shift and rotation did bad, maybe due to 137 x 236 size (some graphemes become outside image region :-( ). <br>\n<br></li>\n<li>Co-occurence\nThere are limited combinations between R, V and C in train dataset. But as a host said, in test dataset, some unseen combinations will come. So I did a simple trick; build a model which consider both co-occurence and non-co-occurence of R, V, C to be robust with unseen data in SE-ResNet50 based model. This model has dual paths after SE-ResNet block. One is straightforward path from GeM2D to each top fc layers (R, V and C). And the other is through 2 fc layers (please see above figure).\n<br></li>\n<li>Multi stage learning ( Xentropy, Reduced Focal Loss (credit to <a href=\"/phalanx\">@phalanx</a> ) )\nIn the middle point of this competition, <a href=\"/phalanx\">@phalanx</a> advised us about OHEM in this nice <a href=\"https://www.kaggle.com/c/bengaliai-cv19/discussion/123198#734490\">thread</a>. <br>\nIndeed, our cv score (recall) became better with the same technique(Reduced Focal Loss). \n<br></li>\n<li>Different image size ensemble\nIn some CV competitions, different size image ensemble boosted scores. We thought this will be same in this competition. This ensemble pushes our place from 12th to 6th.\n<br></li>\n<li>Post Processing (big impact for us)\nWe adjusted our prediction to improve evaluation metric in this competition (Recall). Optimization was performed with all classes (168 + 11 + 7). And we used (1/each class counts) as initial values for optimization. The optimal coefficients are almost close to initial values. We were very nervous to use this post processing, but as a result this pushes us to gold range.</li>\n</ol>",
      "rawMarkdown": "At first, congratulations to the winner @linshokaku and all participants who finished this competition.  \nWe did not imagine such a big shake ( We estimated that there would not be so much unseen data in private test dataset ).  \n  \nAnyway thank you Bengali.AI and kaggle for organizing this competition, also to my teammate @nejumi !\n\n---\n\nBelow figure shows an overview of our solution.  I think our solution is very simple and nothing special.   \n( BTW 1st place solution is very unique and impressive for me ;-)  )\nHere I will add naive explanations to some points worth to be noted.  \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F479538%2F7b541b6993070fbf37181aa631843011%2FBengali_AI.png?generation=1584423861453528&amp;alt=media)\n\nNOTE: If you want to see a higher resolution image, please refer this [link](https://speakerdeck.com/hoxomaxwell/kaggle-bengali-dot-ai-6-th-place-solution).\n\n\n1. Intensive augmentations\nWe used CutMix, MixUp, Cutout, width/height shift, rotate, Erosion, [GridMask](https://www.kaggle.com/haqishen/gridmask) ( credit to @haqishen ), Zoom...  \nCutMix and Cutout worked well for both of us, but in my case, height shift and rotation did bad, maybe due to 137 x 236 size (some graphemes become outside image region :-( ).  \n<br>\n2. Co-occurence\nThere are limited combinations between R, V and C in train dataset. But as a host said, in test dataset, some unseen combinations will come. So I did a simple trick; build a model which consider both co-occurence and non-co-occurence of R, V, C to be robust with unseen data in SE-ResNet50 based model. This model has dual paths after SE-ResNet block. One is straightforward path from GeM2D to each top fc layers (R, V and C). And the other is through 2 fc layers (please see above figure).\n<br>\n3. Multi stage learning ( Xentropy, Reduced Focal Loss (credit to @phalanx ) )\nIn the middle point of this competition, @phalanx advised us about OHEM in this nice [thread](https://www.kaggle.com/c/bengaliai-cv19/discussion/123198#734490).  \nIndeed, our cv score (recall) became better with the same technique(Reduced Focal Loss). \n<br>\n4. Different image size ensemble\nIn some CV competitions, different size image ensemble boosted scores. We thought this will be same in this competition. This ensemble pushes our place from 12th to 6th.\n<br>\n5. Post Processing (big impact for us)\nWe adjusted our prediction to improve evaluation metric in this competition (Recall). Optimization was performed with all classes (168 + 11 + 7). And we used (1/each class counts) as initial values for optimization. The optimal coefficients are almost close to initial values. We were very nervous to use this post processing, but as a result this pushes us to gold range.",
      "votes": 53
    },
    {
      "id": 778616,
      "postDate": "2020-03-18T15:41:37.483Z",
      "content": "<p>Congratulation! I'm glad that my code is useful for you ;)</p>",
      "rawMarkdown": "Congratulation! I'm glad that my code is useful for you ;)",
      "votes": 1
    },
    {
      "id": 777837,
      "postDate": "2020-03-18T01:41:43.007Z",
      "content": "<p>Congrats!!\nThank you for sharing the good pipeline!!</p>",
      "rawMarkdown": "Congrats!!\nThank you for sharing the good pipeline!!",
      "votes": 1
    },
    {
      "id": 777826,
      "postDate": "2020-03-18T01:21:18.477Z",
      "content": "<p>Congrats and thanks for sharing, learned a lot from this summarized solution.</p>",
      "rawMarkdown": "Congrats and thanks for sharing, learned a lot from this summarized solution.",
      "votes": 1,
      "replies": [
        {
          "id": 777832,
          "postDate": "2020-03-18T01:35:55.543Z",
          "content": "<p><a href=\"/corochann\">@corochann</a> \nThanks! I also learned from your kernel posted at the early stage in this competition.\nLooking forward to seeing you in another competition.</p>",
          "rawMarkdown": "@corochann \nThanks! I also learned from your kernel posted at the early stage in this competition.\nLooking forward to seeing you in another competition.",
          "votes": 1
        },
        {
          "id": 777838,
          "postDate": "2020-03-18T01:41:45.803Z",
          "content": "<p>Thanks, yeah I'm writing kernels in other competition too 💪 </p>",
          "rawMarkdown": "Thanks, yeah I'm writing kernels in other competition too 💪 ",
          "votes": 1
        }
      ]
    },
    {
      "id": 776763,
      "postDate": "2020-03-17T16:12:50.770Z",
      "content": "<p>Great!! Thanks for ur fast and kind sharing. Congrats on your shaking up! What a magic post processing and it provides good insights for generalizing the deep learning model.</p>",
      "rawMarkdown": "Great!! Thanks for ur fast and kind sharing. Congrats on your shaking up! What a magic post processing and it provides good insights for generalizing the deep learning model.",
      "votes": 1
    },
    {
      "id": 776727,
      "postDate": "2020-03-17T15:39:32.467Z",
      "content": "<p>Congratulations! A little confused over your post-processing method. Is it like Chris's post-processing? It is over my head how \"initial values\" and \"coefficients\" are involved in predictions. I'm sure this is because I'm not getting the best understanding of your method.</p>",
      "rawMarkdown": "Congratulations! A little confused over your post-processing method. Is it like Chris's post-processing? It is over my head how \"initial values\" and \"coefficients\" are involved in predictions. I'm sure this is because I'm not getting the best understanding of your method.",
      "votes": 1,
      "replies": [
        {
          "id": 776748,
          "postDate": "2020-03-17T15:55:54.023Z",
          "content": "<p>Congratulations, too. <a href=\"/roguekk007\">@roguekk007</a> !\nWe used Nelder-Mead optimizer to do this post-processing. Nelder-Mead takes initial values as an arg. And optimal results, <code>coefficients</code>, will depend on this initial values to some extent ( will depend on situations, iteration times, tolerance, ... ). As I wrote, we used inverse values of counts as initial values and got consistent results (coefficients) in all experiments.  </p>\n\n<p>I think our post-processing is almost same as Chris's one.  </p>",
          "rawMarkdown": "Congratulations, too. @roguekk007 !\nWe used Nelder-Mead optimizer to do this post-processing. Nelder-Mead takes initial values as an arg. And optimal results, `coefficients`, will depend on this initial values to some extent ( will depend on situations, iteration times, tolerance, ... ). As I wrote, we used inverse values of counts as initial values and got consistent results (coefficients) in all experiments.  \n\nI think our post-processing is almost same as Chris's one.  "
        },
        {
          "id": 777968,
          "postDate": "2020-03-18T03:42:33.097Z",
          "content": "<p><a href=\"/maxwell110\">@maxwell110</a> Thank you for the reply. Apologize for making my question clear enough. I am wondering why would any optimization algorithm be needed? When the model outputs probabilities we use argmax (or apply Chris's PP). I don't see where there are coefficients that need to be optimized.</p>\n\n<p>Can you please elaborate a bit how exactly you obtain predictions from model probabilities😃 I think that would clear things up. </p>",
          "rawMarkdown": "@maxwell110 Thank you for the reply. Apologize for making my question clear enough. I am wondering why would any optimization algorithm be needed? When the model outputs probabilities we use argmax (or apply Chris's PP). I don't see where there are coefficients that need to be optimized.\n\nCan you please elaborate a bit how exactly you obtain predictions from model probabilities😃 I think that would clear things up. "
        },
        {
          "id": 778227,
          "postDate": "2020-03-18T08:46:56.147Z",
          "content": "<p><a href=\"/roguekk007\">@roguekk007</a> </p>\n\n<blockquote>\n  <p>I am wondering why would any optimization algorithm be needed?</p>\n</blockquote>\n\n<p>Because the distribution of each class (R + V + C = 186 classes) is not uniform. I mean counts of some classes are higher than others as bellow figure. </p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F479538%2F1a6151694102081712301b8bf3bf49b7%2Ffig.PNG?generation=1584520815962422&amp;alt=media\" alt=\"\"></p>\n\n<p>In this case, evaluation metric, Recall ( <strong>Macro Averaged Recall</strong> ) is not fair for all classes. Just adjusting probability higher for low counted classes and lower for higher counted classes makes Recall better. <br>\nI think this is more elaborately explaned in <a href=\"/cdeotte\">@cdeotte</a> 's <a href=\"https://www.kaggle.com/c/bengaliai-cv19/discussion/136021\">thread</a>. </p>",
          "rawMarkdown": "@roguekk007 \n&gt; I am wondering why would any optimization algorithm be needed?\n\nBecause the distribution of each class (R + V + C = 186 classes) is not uniform. I mean counts of some classes are higher than others as bellow figure. \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F479538%2F1a6151694102081712301b8bf3bf49b7%2Ffig.PNG?generation=1584520815962422&amp;alt=media)\n\nIn this case, evaluation metric, Recall ( **Macro Averaged Recall** ) is not fair for all classes. Just adjusting probability higher for low counted classes and lower for higher counted classes makes Recall better.  \nI think this is more elaborately explaned in @cdeotte 's [thread](https://www.kaggle.com/c/bengaliai-cv19/discussion/136021). ",
          "votes": 1
        },
        {
          "id": 778417,
          "postDate": "2020-03-18T12:24:27.930Z",
          "content": "<p><a href=\"/maxwell110\">@maxwell110</a> I get it now why coefficients are needed (different distribution for each class) and that the coefficients are used to adjust the probabilities for each class/component (probably element-wise multiplication?). I still don't get how the optimization is performed. Is this doing optimization during inference on the fly? What is the thing being optimized when we only have predicted probabilities? Chris did one-step discretizing optimization of expected metric based on predicted values; I fail to connect that with your use of Nelder-mead here. A little confusing to me😧 I am really curious about how this post-processing works, though</p>",
          "rawMarkdown": "@maxwell110 I get it now why coefficients are needed (different distribution for each class) and that the coefficients are used to adjust the probabilities for each class/component (probably element-wise multiplication?). I still don't get how the optimization is performed. Is this doing optimization during inference on the fly? What is the thing being optimized when we only have predicted probabilities? Chris did one-step discretizing optimization of expected metric based on predicted values; I fail to connect that with your use of Nelder-mead here. A little confusing to me😧 I am really curious about how this post-processing works, though"
        },
        {
          "id": 780434,
          "postDate": "2020-03-20T09:04:49.373Z",
          "content": "<p><a href=\"/maxwell110\">@maxwell110</a> congratulation. I had the same question as <a href=\"/roguekk007\">@roguekk007</a>, couldn't fully understand the optimization part; I am new to such tricks in using cv problems. Would you please refer to some articles or kernel or any sort of quick intuitive articles where it demonstrates with a working code? </p>",
          "rawMarkdown": "@maxwell110 congratulation. I had the same question as @roguekk007, couldn't fully understand the optimization part; I am new to such tricks in using cv problems. Would you please refer to some articles or kernel or any sort of quick intuitive articles where it demonstrates with a working code? "
        }
      ]
    },
    {
      "id": 776234,
      "postDate": "2020-03-17T08:20:54.153Z",
      "content": "<p>Congrats. May I ask what's your motivation using other 2fc layers besides each top layer? Looks like it is a bunch of shared fc layers across the three components. What could be the major contributions of this modification from your perspective?  Thanks</p>",
      "rawMarkdown": "Congrats. May I ask what's your motivation using other 2fc layers besides each top layer? Looks like it is a bunch of shared fc layers across the three components. What could be the major contributions of this modification from your perspective?  Thanks",
      "votes": 1,
      "replies": [
        {
          "id": 776423,
          "postDate": "2020-03-17T11:34:11.910Z",
          "content": "<p>Good point, thanks <a href=\"/murphy89\">@murphy89</a> . I think my explanation is a little short. <br>\nThis architecture was motivated by topic 2, <code>Co-occurence</code>.</p>\n\n<p>I wrote, <br>\n&gt;build a model which consider both co-occurence and non-co-occurence of R, V, C to be robust with unseen data in SE-ResNet50 based model. This model has dual paths after SE-ResNet block. One is straightforward path from GeM2D to each top fc layers (R, V and C). And the other is through 2 fc layers (please see above figure).</p>\n\n<p>R, V and C share the same weight in 2 fc layers (1280 fc -&gt; 512 fc). I used this architecture to catch co-occurence, because I thought sharing same weights means considering the relationship of R, V and C: co-occurence. <br>\nCo-occurence of R, V and C is important. But if we make models learn this relationship extremely, models will become unpredictable to unseen data (unseen combination of R, V and C).</p>\n\n<p>Hope this is what you want to know.  </p>",
          "rawMarkdown": "Good point, thanks @murphy89 . I think my explanation is a little short.  \nThis architecture was motivated by topic 2, `Co-occurence`.\n\nI wrote,  \n&gt;build a model which consider both co-occurence and non-co-occurence of R, V, C to be robust with unseen data in SE-ResNet50 based model. This model has dual paths after SE-ResNet block. One is straightforward path from GeM2D to each top fc layers (R, V and C). And the other is through 2 fc layers (please see above figure).\n\nR, V and C share the same weight in 2 fc layers (1280 fc -&gt; 512 fc). I used this architecture to catch co-occurence, because I thought sharing same weights means considering the relationship of R, V and C: co-occurence.  \nCo-occurence of R, V and C is important. But if we make models learn this relationship extremely, models will become unpredictable to unseen data (unseen combination of R, V and C).\n\nHope this is what you want to know.  "
        }
      ]
    },
    {
      "id": 776142,
      "postDate": "2020-03-17T06:32:43.853Z",
      "content": "<p>Congrats! Thanks for sharing! Your pipeline image is awesome. You also use the magic post-processing~ 👍 </p>",
      "rawMarkdown": "Congrats! Thanks for sharing! Your pipeline image is awesome. You also use the magic post-processing~ 👍 ",
      "votes": 1
    },
    {
      "id": 775991,
      "postDate": "2020-03-17T04:13:17.633Z",
      "content": "<p>congrats, really waiting to see your full updates :) </p>",
      "rawMarkdown": "congrats, really waiting to see your full updates :) ",
      "votes": 1
    },
    {
      "id": 779598,
      "postDate": "2020-03-19T14:09:44.137Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 778616,
      "author_name": "Qishen Ha",
      "author_url": "",
      "post_date": "2020-03-18T15:41:37.483000",
      "content": "<p>Congratulation! I'm glad that my code is useful for you ;)</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 777837,
      "author_name": "mocobt",
      "author_url": "",
      "post_date": "2020-03-18T01:41:43.007000",
      "content": "<p>Congrats!!\nThank you for sharing the good pipeline!!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 777826,
      "author_name": "corochann",
      "author_url": "",
      "post_date": "2020-03-18T01:21:18.477000",
      "content": "<p>Congrats and thanks for sharing, learned a lot from this summarized solution.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 777832,
          "author_name": "Maxwell",
          "author_url": "",
          "post_date": "2020-03-18T01:35:55.543000",
          "content": "<p><a href=\"/corochann\">@corochann</a> \nThanks! I also learned from your kernel posted at the early stage in this competition.\nLooking forward to seeing you in another competition.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 777838,
          "author_name": "corochann",
          "author_url": "",
          "post_date": "2020-03-18T01:41:45.803000",
          "content": "<p>Thanks, yeah I'm writing kernels in other competition too 💪 </p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 776763,
      "author_name": "Hunmin Yang",
      "author_url": "",
      "post_date": "2020-03-17T16:12:50.770000",
      "content": "<p>Great!! Thanks for ur fast and kind sharing. Congrats on your shaking up! What a magic post processing and it provides good insights for generalizing the deep learning model.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 776727,
      "author_name": "Nicholas Lyu",
      "author_url": "",
      "post_date": "2020-03-17T15:39:32.467000",
      "content": "<p>Congratulations! A little confused over your post-processing method. Is it like Chris's post-processing? It is over my head how \"initial values\" and \"coefficients\" are involved in predictions. I'm sure this is because I'm not getting the best understanding of your method.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 776748,
          "author_name": "Maxwell",
          "author_url": "",
          "post_date": "2020-03-17T15:55:54.023000",
          "content": "<p>Congratulations, too. <a href=\"/roguekk007\">@roguekk007</a> !\nWe used Nelder-Mead optimizer to do this post-processing. Nelder-Mead takes initial values as an arg. And optimal results, <code>coefficients</code>, will depend on this initial values to some extent ( will depend on situations, iteration times, tolerance, ... ). As I wrote, we used inverse values of counts as initial values and got consistent results (coefficients) in all experiments.  </p>\n\n<p>I think our post-processing is almost same as Chris's one.  </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 777968,
          "author_name": "Nicholas Lyu",
          "author_url": "",
          "post_date": "2020-03-18T03:42:33.097000",
          "content": "<p><a href=\"/maxwell110\">@maxwell110</a> Thank you for the reply. Apologize for making my question clear enough. I am wondering why would any optimization algorithm be needed? When the model outputs probabilities we use argmax (or apply Chris's PP). I don't see where there are coefficients that need to be optimized.</p>\n\n<p>Can you please elaborate a bit how exactly you obtain predictions from model probabilities😃 I think that would clear things up. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 778227,
          "author_name": "Maxwell",
          "author_url": "",
          "post_date": "2020-03-18T08:46:56.147000",
          "content": "<p><a href=\"/roguekk007\">@roguekk007</a> </p>\n\n<blockquote>\n  <p>I am wondering why would any optimization algorithm be needed?</p>\n</blockquote>\n\n<p>Because the distribution of each class (R + V + C = 186 classes) is not uniform. I mean counts of some classes are higher than others as bellow figure. </p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F479538%2F1a6151694102081712301b8bf3bf49b7%2Ffig.PNG?generation=1584520815962422&amp;alt=media\" alt=\"\"></p>\n\n<p>In this case, evaluation metric, Recall ( <strong>Macro Averaged Recall</strong> ) is not fair for all classes. Just adjusting probability higher for low counted classes and lower for higher counted classes makes Recall better. <br>\nI think this is more elaborately explaned in <a href=\"/cdeotte\">@cdeotte</a> 's <a href=\"https://www.kaggle.com/c/bengaliai-cv19/discussion/136021\">thread</a>. </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 778417,
          "author_name": "Nicholas Lyu",
          "author_url": "",
          "post_date": "2020-03-18T12:24:27.930000",
          "content": "<p><a href=\"/maxwell110\">@maxwell110</a> I get it now why coefficients are needed (different distribution for each class) and that the coefficients are used to adjust the probabilities for each class/component (probably element-wise multiplication?). I still don't get how the optimization is performed. Is this doing optimization during inference on the fly? What is the thing being optimized when we only have predicted probabilities? Chris did one-step discretizing optimization of expected metric based on predicted values; I fail to connect that with your use of Nelder-mead here. A little confusing to me😧 I am really curious about how this post-processing works, though</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 780434,
          "author_name": "Innat",
          "author_url": "",
          "post_date": "2020-03-20T09:04:49.373000",
          "content": "<p><a href=\"/maxwell110\">@maxwell110</a> congratulation. I had the same question as <a href=\"/roguekk007\">@roguekk007</a>, couldn't fully understand the optimization part; I am new to such tricks in using cv problems. Would you please refer to some articles or kernel or any sort of quick intuitive articles where it demonstrates with a working code? </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 776234,
      "author_name": "Morphy",
      "author_url": "",
      "post_date": "2020-03-17T08:20:54.153000",
      "content": "<p>Congrats. May I ask what's your motivation using other 2fc layers besides each top layer? Looks like it is a bunch of shared fc layers across the three components. What could be the major contributions of this modification from your perspective?  Thanks</p>",
      "votes": 1,
      "replies": [
        {
          "id": 776423,
          "author_name": "Maxwell",
          "author_url": "",
          "post_date": "2020-03-17T11:34:11.910000",
          "content": "<p>Good point, thanks <a href=\"/murphy89\">@murphy89</a> . I think my explanation is a little short. <br>\nThis architecture was motivated by topic 2, <code>Co-occurence</code>.</p>\n\n<p>I wrote, <br>\n&gt;build a model which consider both co-occurence and non-co-occurence of R, V, C to be robust with unseen data in SE-ResNet50 based model. This model has dual paths after SE-ResNet block. One is straightforward path from GeM2D to each top fc layers (R, V and C). And the other is through 2 fc layers (please see above figure).</p>\n\n<p>R, V and C share the same weight in 2 fc layers (1280 fc -&gt; 512 fc). I used this architecture to catch co-occurence, because I thought sharing same weights means considering the relationship of R, V and C: co-occurence. <br>\nCo-occurence of R, V and C is important. But if we make models learn this relationship extremely, models will become unpredictable to unseen data (unseen combination of R, V and C).</p>\n\n<p>Hope this is what you want to know.  </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 776142,
      "author_name": "Helen",
      "author_url": "",
      "post_date": "2020-03-17T06:32:43.853000",
      "content": "<p>Congrats! Thanks for sharing! Your pipeline image is awesome. You also use the magic post-processing~ 👍 </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 775991,
      "author_name": "Innat",
      "author_url": "",
      "post_date": "2020-03-17T04:13:17.633000",
      "content": "<p>congrats, really waiting to see your full updates :) </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 779598,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-03-19T14:09:44.137000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "775963": "At first, congratulations to the winner @linshokaku and all participants who finished this competition.  \nWe did not imagine such a big shake ( We estimated that there would not be so much unseen data in private test dataset ).  \n  \nAnyway thank you Bengali.AI and kaggle for organizing this competition, also to my teammate @nejumi !\n\n---\n\nBelow figure shows an overview of our solution.  I think our solution is very simple and nothing special.   \n( BTW 1st place solution is very unique and impressive for me ;-)  )\nHere I will add naive explanations to some points worth to be noted.  \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F479538%2F7b541b6993070fbf37181aa631843011%2FBengali_AI.png?generation=1584423861453528&amp;alt=media)\n\nNOTE: If you want to see a higher resolution image, please refer this [link](https://speakerdeck.com/hoxomaxwell/kaggle-bengali-dot-ai-6-th-place-solution).\n\n\n1. Intensive augmentations\nWe used CutMix, MixUp, Cutout, width/height shift, rotate, Erosion, [GridMask](https://www.kaggle.com/haqishen/gridmask) ( credit to @haqishen ), Zoom...  \nCutMix and Cutout worked well for both of us, but in my case, height shift and rotation did bad, maybe due to 137 x 236 size (some graphemes become outside image region :-( ).  \n<br>\n2. Co-occurence\nThere are limited combinations between R, V and C in train dataset. But as a host said, in test dataset, some unseen combinations will come. So I did a simple trick; build a model which consider both co-occurence and non-co-occurence of R, V, C to be robust with unseen data in SE-ResNet50 based model. This model has dual paths after SE-ResNet block. One is straightforward path from GeM2D to each top fc layers (R, V and C). And the other is through 2 fc layers (please see above figure).\n<br>\n3. Multi stage learning ( Xentropy, Reduced Focal Loss (credit to @phalanx ) )\nIn the middle point of this competition, @phalanx advised us about OHEM in this nice [thread](https://www.kaggle.com/c/bengaliai-cv19/discussion/123198#734490).  \nIndeed, our cv score (recall) became better with the same technique(Reduced Focal Loss). \n<br>\n4. Different image size ensemble\nIn some CV competitions, different size image ensemble boosted scores. We thought this will be same in this competition. This ensemble pushes our place from 12th to 6th.\n<br>\n5. Post Processing (big impact for us)\nWe adjusted our prediction to improve evaluation metric in this competition (Recall). Optimization was performed with all classes (168 + 11 + 7). And we used (1/each class counts) as initial values for optimization. The optimal coefficients are almost close to initial values. We were very nervous to use this post processing, but as a result this pushes us to gold range.",
    "778616": "Congratulation! I'm glad that my code is useful for you ;)",
    "777837": "Congrats!!\nThank you for sharing the good pipeline!!",
    "777826": "Congrats and thanks for sharing, learned a lot from this summarized solution.",
    "776763": "Great!! Thanks for ur fast and kind sharing. Congrats on your shaking up! What a magic post processing and it provides good insights for generalizing the deep learning model.",
    "776727": "Congratulations! A little confused over your post-processing method. Is it like Chris's post-processing? It is over my head how \"initial values\" and \"coefficients\" are involved in predictions. I'm sure this is because I'm not getting the best understanding of your method.",
    "776234": "Congrats. May I ask what's your motivation using other 2fc layers besides each top layer? Looks like it is a bunch of shared fc layers across the three components. What could be the major contributions of this modification from your perspective?  Thanks",
    "776142": "Congrats! Thanks for sharing! Your pipeline image is awesome. You also use the magic post-processing~ 👍 ",
    "775991": "congrats, really waiting to see your full updates :) ",
    "779598": ""
  }
}