{
  "id": 136064,
  "title": "24th place solution - Post processing(+0.013 Private LB)",
  "url": "/competitions/bengaliai-cv19/writeups/statsu-24th-place-solution-post-processing-0-013-p",
  "author_name": "",
  "post_date": "2020-03-17T09:20:52.370Z",
  "votes": 15,
  "comment_count": 1,
  "views": 0,
  "content": "<p>First of all, I would like to thank the organizers for holding a very fun competition. Thank you to all kagglers! Through the discussion, I learned a lot.</p>\n\n<p>My solution used a simple ensemble model and post-processing to maximize macro recall. Post-processing adds a bias to the logit output by the model. Without post-processing I am out of medal zone😭</p>\n\n<p>Here is an overview of the solution:</p>\n\n<h2>Inference Label</h2>\n\n<p>First,  the DL model calculates the logits. As a post-process, add optimal bias to logits. Infer estimated labels using argmax for logits.</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2117884%2Fdb0652c31dc1ea8a9abbb306f3ba6bd2%2Finference.png?generation=1584433357383197&amp;alt=media\" alt=\"\"></p>\n\n<h2>Model</h2>\n\n<p>The model is an ensemble of five models. All training data was used for learning. Due to my limited time and computational resources, I've done very little cross-validation.</p>\n\n<ul>\n<li>model 1, 2, 3, 4 <br>\n<ul><li>Seresnext50_32x4d (pretrained model using imagenet)  </li>\n<li>image size 128x128  </li>\n<li>Three linear head (output is grapheme root logit (168), Vowel diacritic logit (168), Consonant diacritic logit (168))  </li>\n<li>Mish activation, Drop block</li>\n<li>Cross Entropy Loss  </li>\n<li>AdaBound  </li>\n<li>Change the following parameters for each model  </li>\n<li>Manifold mixup (alpha, layer), ShiftScaleRotate, Cutout  </li></ul></li>\n<li>model 5\n<ul><li>Seresnext50_32x4d (pretrained model using imagenet)  </li>\n<li>image size 137x236  </li>\n<li>Three linear head (output is grapheme root logit (168), Vowel diacritic logit (168), Consonant diacritic logit (168))  </li>\n<li>Mish activation, Drop block</li>\n<li>Cross Entropy Loss  </li>\n<li>AdaBound  </li>\n<li>Manifold mixup (alpha, layer), ShiftScaleRotate, Cutout  </li></ul></li>\n</ul>\n\n<h2>Calculation of post-processing bias</h2>\n\n<p>Since the model is trained in Cross Entropy Loss, it generally does not maximize macro recall.  So I considered using a bias that would be added to logit to optimize macro recall. The optimal bias was calculated using a real coded genetic algorithm as follows:  </p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2117884%2F1b04e39fcfe7d6862464b81da5bd0878%2Fsearch_optimal_bias.png?generation=1584434262345894&amp;alt=media\" alt=\"\"></p>\n\n<ul>\n<li>Calculate logit for all training data.  </li>\n<li>Calculate bias to maximize macro recall with real-valued genetic algorithm.  </li>\n<li>Calculate logit for all training data with data extension.  </li>\n<li>Calculate bias to maximize macro recall with real-valued genetic algorithm.  </li>\n<li>The above biases are averaged to obtain the final bias.</li>\n</ul>\n\n<p>In estimating the test data, the above bias calculated from the training data is used.</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2117884%2Fb8874f1bef7f5f1e5a0125ab21d3d7d0%2Foptimal_bias.png?generation=1584434294711062&amp;alt=media\" alt=\"\"></p>\n\n<p>There is not much good English literature on real coded genetic algorithms. I think the following is helpful for implementation.\n<a href=\"https://github.com/statsu1990/Real-coded-genetic-algorithm\">https://github.com/statsu1990/Real-coded-genetic-algorithm</a></p>\n\n<p>CMA-ES is similar to real coded genetic algorithm, and there is a lot of English literature.\nTherefore, you may want to try CMA-ES.  </p>\n\n<h2>Score (Public LB / Private LB)</h2>\n\n<p>Single model <br>\n- model1 (wo bias): 0.9689 / 0.9285 <br>\n- model2 (wo bias): 0.9680 / 0.9270 <br>\n- model3 (wo bias): 0.9691 / 0.9317 <br>\n- model4 (wo bias): 0.9681 / 0.9243 <br>\n- model5 (wo bias): 0.9705 / 0.9290  </p>\n\n<p>Ensemble model\n- ensemble1 ~ 5 (TTA, wo bias): 0.9712 / 0.9309 <br>\n- ensemble1 ~ 5 (TTA, with bias): 0.9744 / 0.9435</p>\n\n<p>That's all. thank you for reading!</p>",
  "messages": [
    {
      "id": "776255",
      "postDate": "03/17/2020 08:39:51",
      "content": "<p>First of all, I would like to thank the organizers for holding a very fun competition. Thank you to all kagglers! Through the discussion, I learned a lot.</p>\n\n<p>My solution used a simple ensemble model and post-processing to maximize macro recall. Post-processing adds a bias to the logit output by the model. Without post-processing I am out of medal zone😭</p>\n\n<p>Here is an overview of the solution:</p>\n\n<h2>Inference Label</h2>\n\n<p>First,  the DL model calculates the logits. As a post-process, add optimal bias to logits. Infer estimated labels using argmax for logits.</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2117884%2Fdb0652c31dc1ea8a9abbb306f3ba6bd2%2Finference.png?generation=1584433357383197&amp;alt=media\" alt=\"\"></p>\n\n<h2>Model</h2>\n\n<p>The model is an ensemble of five models. All training data was used for learning. Due to my limited time and computational resources, I've done very little cross-validation.</p>\n\n<ul>\n<li>model 1, 2, 3, 4 <br>\n<ul><li>Seresnext50_32x4d (pretrained model using imagenet)  </li>\n<li>image size 128x128  </li>\n<li>Three linear head (output is grapheme root logit (168), Vowel diacritic logit (168), Consonant diacritic logit (168))  </li>\n<li>Mish activation, Drop block</li>\n<li>Cross Entropy Loss  </li>\n<li>AdaBound  </li>\n<li>Change the following parameters for each model  </li>\n<li>Manifold mixup (alpha, layer), ShiftScaleRotate, Cutout  </li></ul></li>\n<li>model 5\n<ul><li>Seresnext50_32x4d (pretrained model using imagenet)  </li>\n<li>image size 137x236  </li>\n<li>Three linear head (output is grapheme root logit (168), Vowel diacritic logit (168), Consonant diacritic logit (168))  </li>\n<li>Mish activation, Drop block</li>\n<li>Cross Entropy Loss  </li>\n<li>AdaBound  </li>\n<li>Manifold mixup (alpha, layer), ShiftScaleRotate, Cutout  </li></ul></li>\n</ul>\n\n<h2>Calculation of post-processing bias</h2>\n\n<p>Since the model is trained in Cross Entropy Loss, it generally does not maximize macro recall.  So I considered using a bias that would be added to logit to optimize macro recall. The optimal bias was calculated using a real coded genetic algorithm as follows:  </p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2117884%2F1b04e39fcfe7d6862464b81da5bd0878%2Fsearch_optimal_bias.png?generation=1584434262345894&amp;alt=media\" alt=\"\"></p>\n\n<ul>\n<li>Calculate logit for all training data.  </li>\n<li>Calculate bias to maximize macro recall with real-valued genetic algorithm.  </li>\n<li>Calculate logit for all training data with data extension.  </li>\n<li>Calculate bias to maximize macro recall with real-valued genetic algorithm.  </li>\n<li>The above biases are averaged to obtain the final bias.</li>\n</ul>\n\n<p>In estimating the test data, the above bias calculated from the training data is used.</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2117884%2Fb8874f1bef7f5f1e5a0125ab21d3d7d0%2Foptimal_bias.png?generation=1584434294711062&amp;alt=media\" alt=\"\"></p>\n\n<p>There is not much good English literature on real coded genetic algorithms. I think the following is helpful for implementation.\n<a href=\"https://github.com/statsu1990/Real-coded-genetic-algorithm\">https://github.com/statsu1990/Real-coded-genetic-algorithm</a></p>\n\n<p>CMA-ES is similar to real coded genetic algorithm, and there is a lot of English literature.\nTherefore, you may want to try CMA-ES.  </p>\n\n<h2>Score (Public LB / Private LB)</h2>\n\n<p>Single model <br>\n- model1 (wo bias): 0.9689 / 0.9285 <br>\n- model2 (wo bias): 0.9680 / 0.9270 <br>\n- model3 (wo bias): 0.9691 / 0.9317 <br>\n- model4 (wo bias): 0.9681 / 0.9243 <br>\n- model5 (wo bias): 0.9705 / 0.9290  </p>\n\n<p>Ensemble model\n- ensemble1 ~ 5 (TTA, wo bias): 0.9712 / 0.9309 <br>\n- ensemble1 ~ 5 (TTA, with bias): 0.9744 / 0.9435</p>\n\n<p>That's all. thank you for reading!</p>",
      "rawMarkdown": "First of all, I would like to thank the organizers for holding a very fun competition. Thank you to all kagglers! Through the discussion, I learned a lot.\n\n\nMy solution used a simple ensemble model and post-processing to maximize macro recall. Post-processing adds a bias to the logit output by the model. Without post-processing I am out of medal zone😭\n\n\nHere is an overview of the solution:\n\n## Inference Label\nFirst,  the DL model calculates the logits. As a post-process, add optimal bias to logits. Infer estimated labels using argmax for logits.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2117884%2Fdb0652c31dc1ea8a9abbb306f3ba6bd2%2Finference.png?generation=1584433357383197&amp;alt=media)\n\n\n## Model\nThe model is an ensemble of five models. All training data was used for learning. Due to my limited time and computational resources, I've done very little cross-validation.\n\n- model 1, 2, 3, 4  \n  - Seresnext50_32x4d (pretrained model using imagenet)  \n  - image size 128x128  \n  - Three linear head (output is grapheme root logit (168), Vowel diacritic logit (168), Consonant diacritic logit (168))  \n  - Mish activation, Drop block\n  - Cross Entropy Loss  \n  - AdaBound  \n  - Change the following parameters for each model  \n    - Manifold mixup (alpha, layer), ShiftScaleRotate, Cutout  \n- model 5\n  - Seresnext50_32x4d (pretrained model using imagenet)  \n  - image size 137x236  \n  - Three linear head (output is grapheme root logit (168), Vowel diacritic logit (168), Consonant diacritic logit (168))  \n  - Mish activation, Drop block\n  - Cross Entropy Loss  \n  - AdaBound  \n  - Manifold mixup (alpha, layer), ShiftScaleRotate, Cutout  \n\n\n## Calculation of post-processing bias\nSince the model is trained in Cross Entropy Loss, it generally does not maximize macro recall.  So I considered using a bias that would be added to logit to optimize macro recall. The optimal bias was calculated using a real coded genetic algorithm as follows:  \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2117884%2F1b04e39fcfe7d6862464b81da5bd0878%2Fsearch_optimal_bias.png?generation=1584434262345894&amp;alt=media)\n\n- Calculate logit for all training data.  \n- Calculate bias to maximize macro recall with real-valued genetic algorithm.  \n- Calculate logit for all training data with data extension.  \n- Calculate bias to maximize macro recall with real-valued genetic algorithm.  \n- The above biases are averaged to obtain the final bias.\n\nIn estimating the test data, the above bias calculated from the training data is used.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2117884%2Fb8874f1bef7f5f1e5a0125ab21d3d7d0%2Foptimal_bias.png?generation=1584434294711062&amp;alt=media)\n\n\nThere is not much good English literature on real coded genetic algorithms. I think the following is helpful for implementation.\nhttps://github.com/statsu1990/Real-coded-genetic-algorithm\n\nCMA-ES is similar to real coded genetic algorithm, and there is a lot of English literature.\nTherefore, you may want to try CMA-ES.  \n\n\n## Score (Public LB / Private LB)\nSingle model  \n- model1 (wo bias): 0.9689 / 0.9285  \n- model2 (wo bias): 0.9680 / 0.9270  \n- model3 (wo bias): 0.9691 / 0.9317  \n- model4 (wo bias): 0.9681 / 0.9243  \n- model5 (wo bias): 0.9705 / 0.9290  \n\n\nEnsemble model\n- ensemble1 ~ 5 (TTA, wo bias): 0.9712 / 0.9309  \n- ensemble1 ~ 5 (TTA, with bias): 0.9744 / 0.9435\n\n\nThat's all. thank you for reading!",
      "votes": null
    },
    {
      "id": "776257",
      "postDate": "03/17/2020 08:42:57",
      "content": "<p>Thank you for sharing!! Your post-processing part  if very impressive. Are you planning to release your code for that part?</p>",
      "rawMarkdown": "Thank you for sharing!! Your post-processing part  if very impressive. Are you planning to release your code for that part?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 776257,
      "author_name": "bibek777",
      "author_url": "",
      "post_date": "03/17/2020 08:42:57",
      "content": "<p>Thank you for sharing!! Your post-processing part  if very impressive. Are you planning to release your code for that part?</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "776255": "First of all, I would like to thank the organizers for holding a very fun competition. Thank you to all kagglers! Through the discussion, I learned a lot.\n\n\nMy solution used a simple ensemble model and post-processing to maximize macro recall. Post-processing adds a bias to the logit output by the model. Without post-processing I am out of medal zone😭\n\n\nHere is an overview of the solution:\n\n## Inference Label\nFirst,  the DL model calculates the logits. As a post-process, add optimal bias to logits. Infer estimated labels using argmax for logits.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2117884%2Fdb0652c31dc1ea8a9abbb306f3ba6bd2%2Finference.png?generation=1584433357383197&amp;alt=media)\n\n\n## Model\nThe model is an ensemble of five models. All training data was used for learning. Due to my limited time and computational resources, I've done very little cross-validation.\n\n- model 1, 2, 3, 4  \n  - Seresnext50_32x4d (pretrained model using imagenet)  \n  - image size 128x128  \n  - Three linear head (output is grapheme root logit (168), Vowel diacritic logit (168), Consonant diacritic logit (168))  \n  - Mish activation, Drop block\n  - Cross Entropy Loss  \n  - AdaBound  \n  - Change the following parameters for each model  \n    - Manifold mixup (alpha, layer), ShiftScaleRotate, Cutout  \n- model 5\n  - Seresnext50_32x4d (pretrained model using imagenet)  \n  - image size 137x236  \n  - Three linear head (output is grapheme root logit (168), Vowel diacritic logit (168), Consonant diacritic logit (168))  \n  - Mish activation, Drop block\n  - Cross Entropy Loss  \n  - AdaBound  \n  - Manifold mixup (alpha, layer), ShiftScaleRotate, Cutout  \n\n\n## Calculation of post-processing bias\nSince the model is trained in Cross Entropy Loss, it generally does not maximize macro recall.  So I considered using a bias that would be added to logit to optimize macro recall. The optimal bias was calculated using a real coded genetic algorithm as follows:  \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2117884%2F1b04e39fcfe7d6862464b81da5bd0878%2Fsearch_optimal_bias.png?generation=1584434262345894&amp;alt=media)\n\n- Calculate logit for all training data.  \n- Calculate bias to maximize macro recall with real-valued genetic algorithm.  \n- Calculate logit for all training data with data extension.  \n- Calculate bias to maximize macro recall with real-valued genetic algorithm.  \n- The above biases are averaged to obtain the final bias.\n\nIn estimating the test data, the above bias calculated from the training data is used.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2117884%2Fb8874f1bef7f5f1e5a0125ab21d3d7d0%2Foptimal_bias.png?generation=1584434294711062&amp;alt=media)\n\n\nThere is not much good English literature on real coded genetic algorithms. I think the following is helpful for implementation.\nhttps://github.com/statsu1990/Real-coded-genetic-algorithm\n\nCMA-ES is similar to real coded genetic algorithm, and there is a lot of English literature.\nTherefore, you may want to try CMA-ES.  \n\n\n## Score (Public LB / Private LB)\nSingle model  \n- model1 (wo bias): 0.9689 / 0.9285  \n- model2 (wo bias): 0.9680 / 0.9270  \n- model3 (wo bias): 0.9691 / 0.9317  \n- model4 (wo bias): 0.9681 / 0.9243  \n- model5 (wo bias): 0.9705 / 0.9290  \n\n\nEnsemble model\n- ensemble1 ~ 5 (TTA, wo bias): 0.9712 / 0.9309  \n- ensemble1 ~ 5 (TTA, with bias): 0.9744 / 0.9435\n\n\nThat's all. thank you for reading!",
    "776257": "Thank you for sharing!! Your post-processing part  if very impressive. Are you planning to release your code for that part?"
  },
  "source": "meta"
}