{
  "id": 136116,
  "title": "13th place brief writeup (public 107th) ",
  "url": "/competitions/bengaliai-cv19/writeups/yama-13th-place-brief-writeup-public-107th",
  "author_name": "",
  "post_date": "2020-03-18T15:50:53.707Z",
  "votes": 29,
  "comment_count": 28,
  "views": 0,
  "content": "<p>updated 2020/03/19</p>\n\n<hr>\n\n<p>Thanks to the hosts and helpful discussions.\nEspecially, I really appreciated <a href=\"/hengck23\">@hengck23</a>'s kind sharing.</p>\n\n<p>I am so lucky to get gold.\nI expected random-split public/private and did nothing special for unseen grapheme.</p>\n\n<p>It seems postprocessing for macro recall hack is really important.</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2796584%2Fd7d8dab30ab64980f8b92edbf1c29af6%2Fbengali2.png?generation=1584544815730850&amp;alt=media\" alt=\"\"></p>\n\n<h2>Model</h2>\n\n<ul>\n<li>modified SE-ResNeXt50. 2-stride conv is replaced by 1-stride-conv+maxblur-pool\nbasically following <a href=\"/hengck23\">@hengck23</a> 's <a href=\"https://www.kaggle.com/c/bengaliai-cv19/discussion/123757#757358\">advice</a></li>\n<li>phase1: trained with stem-conv(stride=2), phase2: replace stem-conv with 3 convs(stride=1)</li>\n<li>Adam + ReduceLROnPlateau, about 200 epochs</li>\n<li><p>5 fold ensemble</p>\n\n<h2>Loss</h2></li>\n<li><p>CrossEntropy for root, vowel, consonant</p>\n\n<ul><li>2 out of 5 folds are finetuned using ohem loss</li></ul></li>\n<li>Multilabel BinaryCrossEntropy to classify decomposed grapheme exists or not\n<ul><li>graphme (or root or vowel or consotant) can be decomposed as shown in the figure. 61 possible decomposed parts exist (except '0' part). below is my exact code to make multihot label.</li></ul></li>\n</ul>\n\n<p>```</p>\n\n<h1>load data</h1>\n\n<p>train = pd.read_csv(C.datadir/'train.csv')\ntrain_labels = train[['grapheme_root', 'vowel_diacritic', 'consonant_diacritic']].values.astype(np.int64)</p>\n\n<h1>components label</h1>\n\n<p>parts_pre = pd.read_csv(C.datadir/'class_map.csv').component\nparts = np.sort(np.unique(np.concatenate([list(e) for e in parts_pre])))\nparts = parts[parts != '0']  # 0 has no meanning\nprint(\"parts:\", parts)  # 61 parts shown</p>\n\n<p>train_labels_comp = []\nfor grapheme in train['grapheme'].values:\n    train_labels_comp.append([part in list(grapheme) for part in parts])\ntrain_labels_comp = np.array(train_labels_comp).astype(np.int64)\nprint(\"train_labels_comp.shape\", train_labels_comp.shape)\nif True: # debug\n    print(\"train_labels_comp\", train_labels_comp[0].tolist())\n    print(\"train_labels_comp\", train_labels_comp[1].tolist())\n    print(\"train_labels_comp\", train_labels_comp[2].tolist())\n```</p>\n\n<h2>Preprocess and Augmentation</h2>\n\n<p>100% same as <a href=\"/hengck23\">@hengck23</a> 's <a href=\"https://www.kaggle.com/c/bengaliai-cv19/discussion/132898#763633\">advice</a>\nPreprocess :  just resizing with cv2.INTER_AREA\nAugmentation : OneOf(basic augmentations) + DropBlock</p>\n\n<h2>Postprocess</h2>\n\n<ul>\n<li>prediction for decomposed grapheme is not used</li>\n<li><strong>maximize expected recall by\nargmax( softmax(logits) / np.power(class_count, 1) )</strong>. note that final sub use factor of 1.15 for grapheme root(no difference for private lb).\n<ul><li>there may be better threshold optimization since confusion matrix still looks asymmetry after above optimization. But I gave up further improvement due to low sample size.</li>\n<li>below is plot of \"macro recall scores V.S. factor n\" for my model (which is not final sub). You can confirm that the peak is around factor=1. You can also find that the improvement is not huge unlike private LB.\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2796584%2F26fd5a4a563451455a39333fe670aab4%2Fmodel033_recall_hack.png?generation=1584545521502984&amp;alt=media\" alt=\"\"></li></ul></li>\n</ul>\n\n<hr>\n\n<p><strong>What I learned during competition ( note to self )</strong>\n<em>how to use pytorch lightning / how to freeze specified layer / how to ensemble / concept of snapshot ensemble / ohem loss, class-balanced loss / hengck23's model surgery method / maxblur-pool and its implementation / keep high resolution may be important for low resolution image / interpolation method for resizing is sometimes important / how to implement some augmentation methods (mixup, cutmix, gridmask) / public &amp; private is not always split randomly / vast.ai is low-priced compared to GCP</em></p>",
  "messages": [
    {
      "id": "776628",
      "postDate": "03/17/2020 14:19:01",
      "content": "<p>updated 2020/03/19</p>\n\n<hr>\n\n<p>Thanks to the hosts and helpful discussions.\nEspecially, I really appreciated <a href=\"/hengck23\">@hengck23</a>'s kind sharing.</p>\n\n<p>I am so lucky to get gold.\nI expected random-split public/private and did nothing special for unseen grapheme.</p>\n\n<p>It seems postprocessing for macro recall hack is really important.</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2796584%2Fd7d8dab30ab64980f8b92edbf1c29af6%2Fbengali2.png?generation=1584544815730850&amp;alt=media\" alt=\"\"></p>\n\n<h2>Model</h2>\n\n<ul>\n<li>modified SE-ResNeXt50. 2-stride conv is replaced by 1-stride-conv+maxblur-pool\nbasically following <a href=\"/hengck23\">@hengck23</a> 's <a href=\"https://www.kaggle.com/c/bengaliai-cv19/discussion/123757#757358\">advice</a></li>\n<li>phase1: trained with stem-conv(stride=2), phase2: replace stem-conv with 3 convs(stride=1)</li>\n<li>Adam + ReduceLROnPlateau, about 200 epochs</li>\n<li><p>5 fold ensemble</p>\n\n<h2>Loss</h2></li>\n<li><p>CrossEntropy for root, vowel, consonant</p>\n\n<ul><li>2 out of 5 folds are finetuned using ohem loss</li></ul></li>\n<li>Multilabel BinaryCrossEntropy to classify decomposed grapheme exists or not\n<ul><li>graphme (or root or vowel or consotant) can be decomposed as shown in the figure. 61 possible decomposed parts exist (except '0' part). below is my exact code to make multihot label.</li></ul></li>\n</ul>\n\n<p>```</p>\n\n<h1>load data</h1>\n\n<p>train = pd.read_csv(C.datadir/'train.csv')\ntrain_labels = train[['grapheme_root', 'vowel_diacritic', 'consonant_diacritic']].values.astype(np.int64)</p>\n\n<h1>components label</h1>\n\n<p>parts_pre = pd.read_csv(C.datadir/'class_map.csv').component\nparts = np.sort(np.unique(np.concatenate([list(e) for e in parts_pre])))\nparts = parts[parts != '0']  # 0 has no meanning\nprint(\"parts:\", parts)  # 61 parts shown</p>\n\n<p>train_labels_comp = []\nfor grapheme in train['grapheme'].values:\n    train_labels_comp.append([part in list(grapheme) for part in parts])\ntrain_labels_comp = np.array(train_labels_comp).astype(np.int64)\nprint(\"train_labels_comp.shape\", train_labels_comp.shape)\nif True: # debug\n    print(\"train_labels_comp\", train_labels_comp[0].tolist())\n    print(\"train_labels_comp\", train_labels_comp[1].tolist())\n    print(\"train_labels_comp\", train_labels_comp[2].tolist())\n```</p>\n\n<h2>Preprocess and Augmentation</h2>\n\n<p>100% same as <a href=\"/hengck23\">@hengck23</a> 's <a href=\"https://www.kaggle.com/c/bengaliai-cv19/discussion/132898#763633\">advice</a>\nPreprocess :  just resizing with cv2.INTER_AREA\nAugmentation : OneOf(basic augmentations) + DropBlock</p>\n\n<h2>Postprocess</h2>\n\n<ul>\n<li>prediction for decomposed grapheme is not used</li>\n<li><strong>maximize expected recall by\nargmax( softmax(logits) / np.power(class_count, 1) )</strong>. note that final sub use factor of 1.15 for grapheme root(no difference for private lb).\n<ul><li>there may be better threshold optimization since confusion matrix still looks asymmetry after above optimization. But I gave up further improvement due to low sample size.</li>\n<li>below is plot of \"macro recall scores V.S. factor n\" for my model (which is not final sub). You can confirm that the peak is around factor=1. You can also find that the improvement is not huge unlike private LB.\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2796584%2F26fd5a4a563451455a39333fe670aab4%2Fmodel033_recall_hack.png?generation=1584545521502984&amp;alt=media\" alt=\"\"></li></ul></li>\n</ul>\n\n<hr>\n\n<p><strong>What I learned during competition ( note to self )</strong>\n<em>how to use pytorch lightning / how to freeze specified layer / how to ensemble / concept of snapshot ensemble / ohem loss, class-balanced loss / hengck23's model surgery method / maxblur-pool and its implementation / keep high resolution may be important for low resolution image / interpolation method for resizing is sometimes important / how to implement some augmentation methods (mixup, cutmix, gridmask) / public &amp; private is not always split randomly / vast.ai is low-priced compared to GCP</em></p>",
      "rawMarkdown": "updated 2020/03/19\n\n----\n\nThanks to the hosts and helpful discussions.\nEspecially, I really appreciated @hengck23's kind sharing.\n\nI am so lucky to get gold.\nI expected random-split public/private and did nothing special for unseen grapheme.\n\nIt seems postprocessing for macro recall hack is really important.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2796584%2Fd7d8dab30ab64980f8b92edbf1c29af6%2Fbengali2.png?generation=1584544815730850&amp;alt=media)\n## Model\n- modified SE-ResNeXt50. 2-stride conv is replaced by 1-stride-conv+maxblur-pool\n   basically following @hengck23 's [advice](https://www.kaggle.com/c/bengaliai-cv19/discussion/123757#757358)\n- phase1: trained with stem-conv(stride=2), phase2: replace stem-conv with 3 convs(stride=1)\n- Adam + ReduceLROnPlateau, about 200 epochs\n- 5 fold ensemble\n## Loss\n- CrossEntropy for root, vowel, consonant\n  - 2 out of 5 folds are finetuned using ohem loss\n- Multilabel BinaryCrossEntropy to classify decomposed grapheme exists or not\n  - graphme (or root or vowel or consotant) can be decomposed as shown in the figure. 61 possible decomposed parts exist (except '0' part). below is my exact code to make multihot label.\n\n```\n# load data\ntrain = pd.read_csv(C.datadir/'train.csv')\ntrain_labels = train[['grapheme_root', 'vowel_diacritic', 'consonant_diacritic']].values.astype(np.int64)\n\n# components label\nparts_pre = pd.read_csv(C.datadir/'class_map.csv').component\nparts = np.sort(np.unique(np.concatenate([list(e) for e in parts_pre])))\nparts = parts[parts != '0']  # 0 has no meanning\nprint(\"parts:\", parts)  # 61 parts shown\n\ntrain_labels_comp = []\nfor grapheme in train['grapheme'].values:\n    train_labels_comp.append([part in list(grapheme) for part in parts])\ntrain_labels_comp = np.array(train_labels_comp).astype(np.int64)\nprint(\"train_labels_comp.shape\", train_labels_comp.shape)\nif True: # debug\n    print(\"train_labels_comp\", train_labels_comp[0].tolist())\n    print(\"train_labels_comp\", train_labels_comp[1].tolist())\n    print(\"train_labels_comp\", train_labels_comp[2].tolist())\n```\n\n## Preprocess and Augmentation\n100% same as @hengck23 's [advice](https://www.kaggle.com/c/bengaliai-cv19/discussion/132898#763633)\nPreprocess :  just resizing with cv2.INTER_AREA\nAugmentation : OneOf(basic augmentations) + DropBlock\n## Postprocess\n- prediction for decomposed grapheme is not used\n- **maximize expected recall by\n  argmax( softmax(logits) / np.power(class_count, 1) )**. note that final sub use factor of 1.15 for grapheme root(no difference for private lb).\n  - there may be better threshold optimization since confusion matrix still looks asymmetry after above optimization. But I gave up further improvement due to low sample size.\n  - below is plot of \"macro recall scores V.S. factor n\" for my model (which is not final sub). You can confirm that the peak is around factor=1. You can also find that the improvement is not huge unlike private LB.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2796584%2F26fd5a4a563451455a39333fe670aab4%2Fmodel033_recall_hack.png?generation=1584545521502984&amp;alt=media)\n\n-----\n**What I learned during competition ( note to self )**\n*how to use pytorch lightning / how to freeze specified layer / how to ensemble / concept of snapshot ensemble / ohem loss, class-balanced loss / hengck23's model surgery method / maxblur-pool and its implementation / keep high resolution may be important for low resolution image / interpolation method for resizing is sometimes important / how to implement some augmentation methods (mixup, cutmix, gridmask) / public &amp; private is not always split randomly / vast.ai is low-priced compared to GCP*",
      "votes": null
    },
    {
      "id": "776637",
      "postDate": "03/17/2020 14:23:51",
      "content": "<p>Congrats on your gold finish!! I've a question, were you using Heng's code too??</p>",
      "rawMarkdown": "Congrats on your gold finish!! I've a question, were you using Heng's code too??",
      "votes": null
    },
    {
      "id": "776643",
      "postDate": "03/17/2020 14:28:58",
      "content": "<p>most of augmentation and preprocess codes are from his implementation. ( I implement my own, but discarded it since his augmentation was better)</p>",
      "rawMarkdown": "most of augmentation and preprocess codes are from his implementation. ( I implement my own, but discarded it since his augmentation was better)",
      "votes": null
    },
    {
      "id": "776647",
      "postDate": "03/17/2020 14:31:44",
      "content": "<p>Congrats, may I ask how do you use decomposed grapheme when test image?</p>",
      "rawMarkdown": "Congrats, may I ask how do you use decomposed grapheme when test image?",
      "votes": null
    },
    {
      "id": "776671",
      "postDate": "03/17/2020 14:45:55",
      "content": "<p>good work!</p>\n\n<p>i will try to repeat your experiment as well :)</p>",
      "rawMarkdown": "good work!\n\ni will try to repeat your experiment as well :)",
      "votes": null
    },
    {
      "id": "776680",
      "postDate": "03/17/2020 14:49:53",
      "content": "<p>Not used for test. just for auxiliary task</p>",
      "rawMarkdown": "Not used for test. just for auxiliary task",
      "votes": null
    },
    {
      "id": "776700",
      "postDate": "03/17/2020 15:17:39",
      "content": "<p>Ok, Thanks for your sharing</p>",
      "rawMarkdown": "Ok, Thanks for your sharing",
      "votes": null
    },
    {
      "id": "776776",
      "postDate": "03/17/2020 16:18:43",
      "content": "<p>Nice!  I like the unicode prediction block.</p>",
      "rawMarkdown": "Nice!  I like the unicode prediction block.",
      "votes": null
    },
    {
      "id": "776892",
      "postDate": "03/17/2020 17:55:48",
      "content": "<p>Congrats! Thanks for sharing your solution! I am confused about the decomposed 61 channels. Could you please explain what do you do extractly to get them？I didn't get it. Thanks again!</p>",
      "rawMarkdown": "Congrats! Thanks for sharing your solution! I am confused about the decomposed 61 channels. Could you please explain what do you do extractly to get them？I didn't get it. Thanks again!",
      "votes": null
    },
    {
      "id": "777650",
      "postDate": "03/17/2020 20:46:59",
      "content": "<p>co9ngrats.  how do you decompose grapheme?</p>",
      "rawMarkdown": "co9ngrats.  how do you decompose grapheme?",
      "votes": null
    },
    {
      "id": "777851",
      "postDate": "03/18/2020 01:52:01",
      "content": "<p>Congrats and thank you for sharing.</p>",
      "rawMarkdown": "Congrats and thank you for sharing.",
      "votes": null
    },
    {
      "id": "777948",
      "postDate": "03/18/2020 03:15:56",
      "content": "<p>Thanks for your question!\ndecomposition is from  <a href=\"https://www.kaggle.com/c/bengaliai-cv19/discussion/123757#765573\">https://www.kaggle.com/c/bengaliai-cv19/discussion/123757#765573</a></p>\n\n<p>I will add some explanation later.</p>",
      "rawMarkdown": "Thanks for your question!\ndecomposition is from  https://www.kaggle.com/c/bengaliai-cv19/discussion/123757#765573\n\nI will add some explanation later.",
      "votes": null
    },
    {
      "id": "777949",
      "postDate": "03/18/2020 03:16:21",
      "content": "<p>Thanks for your question!\ndecomposition is from  <a href=\"https://www.kaggle.com/c/bengaliai-cv19/discussion/123757#765573\">https://www.kaggle.com/c/bengaliai-cv19/discussion/123757#765573</a></p>\n\n<p>I will add some explanation later.</p>",
      "rawMarkdown": "Thanks for your question!\ndecomposition is from  https://www.kaggle.com/c/bengaliai-cv19/discussion/123757#765573\n\nI will add some explanation later.",
      "votes": null
    },
    {
      "id": "778532",
      "postDate": "03/18/2020 14:33:06",
      "content": "<p>Congrats!\nAfter reading the link you provided, still don't know how to decompose grapheme.\nCan you explain more detail?</p>",
      "rawMarkdown": "Congrats!\nAfter reading the link you provided, still don't know how to decompose grapheme.\nCan you explain more detail?",
      "votes": null
    },
    {
      "id": "778618",
      "postDate": "03/18/2020 15:42:25",
      "content": "<p>I added details. Please feel free to ask me questions!</p>",
      "rawMarkdown": "I added details. Please feel free to ask me questions!",
      "votes": null
    },
    {
      "id": "778619",
      "postDate": "03/18/2020 15:43:23",
      "content": "<p>I added details for decomposition. Please feel free to ask me questions!</p>",
      "rawMarkdown": "I added details for decomposition. Please feel free to ask me questions!",
      "votes": null
    },
    {
      "id": "778629",
      "postDate": "03/18/2020 15:49:06",
      "content": "<p>Thanks!\nI added details for decomposition. Please feel free to ask me questions!</p>",
      "rawMarkdown": "Thanks!\nI added details for decomposition. Please feel free to ask me questions!",
      "votes": null
    },
    {
      "id": "778633",
      "postDate": "03/18/2020 15:52:15",
      "content": "<p>Thank you!\nI was really impressed by your creative posts!</p>",
      "rawMarkdown": "Thank you!\nI was really impressed by your creative posts!",
      "votes": null
    },
    {
      "id": "778638",
      "postDate": "03/18/2020 15:54:14",
      "content": "<p>Thank you! \nActually, I checked and used some of your SEResnext kernel codes for my first step!</p>",
      "rawMarkdown": "Thank you! \nActually, I checked and used some of your SEResnext kernel codes for my first step!",
      "votes": null
    },
    {
      "id": "778640",
      "postDate": "03/18/2020 15:55:27",
      "content": "<p>Thank you! Unicode encoding is really informative.</p>",
      "rawMarkdown": "Thank you! Unicode encoding is really informative.",
      "votes": null
    },
    {
      "id": "779082",
      "postDate": "03/19/2020 02:34:39",
      "content": "<p>Thanks for your reply! I get it this time. ：）</p>",
      "rawMarkdown": "Thanks for your reply! I get it this time. ：）",
      "votes": null
    },
    {
      "id": "779091",
      "postDate": "03/19/2020 02:47:51",
      "content": "<p>you can use the head from <a href=\"https://www.kaggle.com/c/bengaliai-cv19/discussion/135990\">https://www.kaggle.com/c/bengaliai-cv19/discussion/135990</a>,\nwhich include the arcface loss with my modified se-resnext50 backbone.</p>\n\n<p>the result is quite amazing. here are the results i get:</p>\n\n<p>reference (<a href=\"/qishen\">@qishen</a> ha implementation): efficientnetb1-128x128 public lb 0.985\nmy implementation:  efficientnetb1-128x128 local cv 0.995 (estimated lb 0.985)\nmy implementation:  modified se-resnext50-112x112 local cv 0.997 (estimated lb 0.987)</p>\n\n<hr>\n\n<p>increasing 112x112 to 128x128 can add 0.001 in my previous experiements.\nusing arcface head + your decompose head may further improve results?</p>",
      "rawMarkdown": "you can use the head from https://www.kaggle.com/c/bengaliai-cv19/discussion/135990,\nwhich include the arcface loss with my modified se-resnext50 backbone.\n\nthe result is quite amazing. here are the results i get:\n\nreference (@qishen ha implementation): efficientnetb1-128x128 public lb 0.985\nmy implementation:  efficientnetb1-128x128 local cv 0.995 (estimated lb 0.985)\nmy implementation:  modified se-resnext50-112x112 local cv 0.997 (estimated lb 0.987)\n\n---\n\nincreasing 112x112 to 128x128 can add 0.001 in my previous experiements.\nusing arcface head + your decompose head may further improve results?",
      "votes": null
    },
    {
      "id": "779105",
      "postDate": "03/19/2020 03:04:34",
      "content": "<p>Congratulations. Awesome solo gold. I like your decomposed grapheme head. That's smart! I didn't realize that Python could decompose graphemes nor that graphemes could be decomposed. Post processing was very powerful in this comp, we found <code>EXP = -1.2</code> to be best too.</p>",
      "rawMarkdown": "Congratulations. Awesome solo gold. I like your decomposed grapheme head. That's smart! I didn't realize that Python could decompose graphemes nor that graphemes could be decomposed. Post processing was very powerful in this comp, we found `EXP = -1.2` to be best too.",
      "votes": null
    },
    {
      "id": "779701",
      "postDate": "03/19/2020 15:56:16",
      "content": "<p>Thanks for your sharing,and I want to ask what's the loss weight of your final choice?</p>",
      "rawMarkdown": "Thanks for your sharing,and I want to ask what's the loss weight of your final choice?",
      "votes": null
    },
    {
      "id": "779886",
      "postDate": "03/19/2020 19:24:34",
      "content": "<p>Wow, arcface really works!\nIn my experiment, the contribution of decompose head was minor. The boost was around +0.0010:\n  CV(0.9804+0.0011), public LB(0.9751+0.0005), private LB(0.9422+0.0011).</p>\n\n<p>I believe that adding decomposed label (which has richer information) will not harm training under any conditions even though the boost might be small.</p>\n\n<hr>\n\n<p>Actually, I anticipated decomposed head serves as a kind of metric learning.\nCombining decomposed head with arcface may yield better learning of inter-grapheme distance.</p>\n\n<p>(Important note: I am not familiar with metric learning AT ALL. I will start learning acrface from now :)  )</p>",
      "rawMarkdown": "Wow, arcface really works!\nIn my experiment, the contribution of decompose head was minor. The boost was around +0.0010:\n  CV(0.9804+0.0011), public LB(0.9751+0.0005), private LB(0.9422+0.0011).\n\nI believe that adding decomposed label (which has richer information) will not harm training under any conditions even though the boost might be small.\n\n---\n\nActually, I anticipated decomposed head serves as a kind of metric learning.\nCombining decomposed head with arcface may yield better learning of inter-grapheme distance.\n\n(Important note: I am not familiar with metric learning AT ALL. I will start learning acrface from now :)  )",
      "votes": null
    },
    {
      "id": "781194",
      "postDate": "03/21/2020 02:42:32",
      "content": "<p>Hi <a href=\"/lisosia\">@lisosia</a> \nThank you for your reply.\nCould I have further questions?</p>\n\n<ol>\n<li><p>What is your loss weight of the 61 BinaryCrossEntropy and  loss weight of the cross entropy of root, vowel, consotant?</p></li>\n<li><p>Why you divide np.power(class_count, 1) ? for unbalanced classes?\n<code>argmax( softmax(logits) / np.power(class_count, 1) )</code></p></li>\n<li><p>Which layer you inserted DropBlock? and the parameters?</p></li>\n</ol>\n\n<p>Looking forward to your reply</p>",
      "rawMarkdown": "Hi @lisosia \nThank you for your reply.\nCould I have further questions?\n\n1. What is your loss weight of the 61 BinaryCrossEntropy and  loss weight of the cross entropy of root, vowel, consotant?\n\n2. Why you divide np.power(class_count, 1) ? for unbalanced classes?\n`argmax( softmax(logits) / np.power(class_count, 1) )`\n\n3. Which layer you inserted DropBlock? and the parameters?\n\nLooking forward to your reply",
      "votes": null
    },
    {
      "id": "784793",
      "postDate": "03/24/2020 14:05:35",
      "content": "<p>Thank you and Congrats to you too!\nWe followed a similar path in the competition :)</p>",
      "rawMarkdown": "Thank you and Congrats to you too!\nWe followed a similar path in the competition :)",
      "votes": null
    },
    {
      "id": "784803",
      "postDate": "03/24/2020 14:09:25",
      "content": "<p>root : vowel : consotant : sum-of-61-BCEs = 2:1:1:1/40</p>",
      "rawMarkdown": "root : vowel : consotant : sum-of-61-BCEs = 2:1:1:1/40",
      "votes": null
    },
    {
      "id": "784837",
      "postDate": "03/24/2020 14:34:26",
      "content": "<p>1.\nroot : vowel : consotant : sum-of-61-BCEs = 2:1:1:1/40</p>\n\n<p>2.\nwell describeld in <a href=\"https://www.kaggle.com/c/bengaliai-cv19/discussion/136021\">Chris Writeup</a></p>\n\n<p>Suppose you are given sample XX with class i (you donot know the correct class).\nIf you correctly predict class i for sample XX, you will get additional score of <code>1 / (sample number of class i in test-set)</code> which can be approximated by <code>1 / (sample number of class i in train-set)</code>.</p>\n\n<p>softmax() output for sample XX is probabiliy.\nif i-th softmax output for sample XX is <code>p_i</code> , the probabiliy of sample XX is in class i is <code>p_i</code> .\nIf you predict class i for sample XX, \nYour <em>expected gain of score</em> is\n<code>p_i * 1 / (sample number of class i in train-set)</code> = <code>softmax(logit)[i] / (sample number of class i in trainset)</code></p>\n\n<p>So the best prediction for sample XX is \n<code>argmax-of-i  ( softmax(logit)[i] / (sample number of class i in train-set)</code> )</p>\n\n<p>3.\nafter layer0, 1, 2. same as <a href=\"https://www.kaggle.com/c/bengaliai-cv19/discussion/132898#763633\">the post</a> I mentioned in the writeup.</p>",
      "rawMarkdown": "1.\nroot : vowel : consotant : sum-of-61-BCEs = 2:1:1:1/40\n\n2.\nwell describeld in [Chris Writeup](https://www.kaggle.com/c/bengaliai-cv19/discussion/136021)\n\nSuppose you are given sample XX with class i (you donot know the correct class).\nIf you correctly predict class i for sample XX, you will get additional score of `1 / (sample number of class i in test-set)` which can be approximated by `1 / (sample number of class i in train-set)`.\n\nsoftmax() output for sample XX is probabiliy.\nif i-th softmax output for sample XX is `p_i` , the probabiliy of sample XX is in class i is `p_i` .\nIf you predict class i for sample XX, \nYour *expected gain of score* is\n`p_i * 1 / (sample number of class i in train-set)` = `softmax(logit)[i] / (sample number of class i in trainset)`\n\nSo the best prediction for sample XX is \n` argmax-of-i  ( softmax(logit)[i] / (sample number of class i in train-set)` )\n\n3.\nafter layer0, 1, 2. same as [the post](https://www.kaggle.com/c/bengaliai-cv19/discussion/132898#763633) I mentioned in the writeup.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 776637,
      "author_name": "bibek777",
      "author_url": "",
      "post_date": "03/17/2020 14:23:51",
      "content": "<p>Congrats on your gold finish!! I've a question, were you using Heng's code too??</p>",
      "votes": null,
      "replies": [
        {
          "id": 776643,
          "author_name": "lisosia",
          "author_url": "",
          "post_date": "03/17/2020 14:28:58",
          "content": "<p>most of augmentation and preprocess codes are from his implementation. ( I implement my own, but discarded it since his augmentation was better)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 776647,
      "author_name": "hesene",
      "author_url": "",
      "post_date": "03/17/2020 14:31:44",
      "content": "<p>Congrats, may I ask how do you use decomposed grapheme when test image?</p>",
      "votes": null,
      "replies": [
        {
          "id": 776680,
          "author_name": "lisosia",
          "author_url": "",
          "post_date": "03/17/2020 14:49:53",
          "content": "<p>Not used for test. just for auxiliary task</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 776700,
          "author_name": "hesene",
          "author_url": "",
          "post_date": "03/17/2020 15:17:39",
          "content": "<p>Ok, Thanks for your sharing</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 776671,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "03/17/2020 14:45:55",
      "content": "<p>good work!</p>\n\n<p>i will try to repeat your experiment as well :)</p>",
      "votes": null,
      "replies": [
        {
          "id": 778633,
          "author_name": "lisosia",
          "author_url": "",
          "post_date": "03/18/2020 15:52:15",
          "content": "<p>Thank you!\nI was really impressed by your creative posts!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 776776,
      "author_name": "harshpatel1692",
      "author_url": "",
      "post_date": "03/17/2020 16:18:43",
      "content": "<p>Nice!  I like the unicode prediction block.</p>",
      "votes": null,
      "replies": [
        {
          "id": 778640,
          "author_name": "lisosia",
          "author_url": "",
          "post_date": "03/18/2020 15:55:27",
          "content": "<p>Thank you! Unicode encoding is really informative.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 776892,
      "author_name": "yuanlin08",
      "author_url": "",
      "post_date": "03/17/2020 17:55:48",
      "content": "<p>Congrats! Thanks for sharing your solution! I am confused about the decomposed 61 channels. Could you please explain what do you do extractly to get them？I didn't get it. Thanks again!</p>",
      "votes": null,
      "replies": [
        {
          "id": 777948,
          "author_name": "lisosia",
          "author_url": "",
          "post_date": "03/18/2020 03:15:56",
          "content": "<p>Thanks for your question!\ndecomposition is from  <a href=\"https://www.kaggle.com/c/bengaliai-cv19/discussion/123757#765573\">https://www.kaggle.com/c/bengaliai-cv19/discussion/123757#765573</a></p>\n\n<p>I will add some explanation later.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 778618,
          "author_name": "lisosia",
          "author_url": "",
          "post_date": "03/18/2020 15:42:25",
          "content": "<p>I added details. Please feel free to ask me questions!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 779082,
          "author_name": "yuanlin08",
          "author_url": "",
          "post_date": "03/19/2020 02:34:39",
          "content": "<p>Thanks for your reply! I get it this time. ：）</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 777650,
      "author_name": "cpmpml",
      "author_url": "",
      "post_date": "03/17/2020 20:46:59",
      "content": "<p>co9ngrats.  how do you decompose grapheme?</p>",
      "votes": null,
      "replies": [
        {
          "id": 777949,
          "author_name": "lisosia",
          "author_url": "",
          "post_date": "03/18/2020 03:16:21",
          "content": "<p>Thanks for your question!\ndecomposition is from  <a href=\"https://www.kaggle.com/c/bengaliai-cv19/discussion/123757#765573\">https://www.kaggle.com/c/bengaliai-cv19/discussion/123757#765573</a></p>\n\n<p>I will add some explanation later.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 778619,
          "author_name": "lisosia",
          "author_url": "",
          "post_date": "03/18/2020 15:43:23",
          "content": "<p>I added details for decomposition. Please feel free to ask me questions!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 777851,
      "author_name": "corochann",
      "author_url": "",
      "post_date": "03/18/2020 01:52:01",
      "content": "<p>Congrats and thank you for sharing.</p>",
      "votes": null,
      "replies": [
        {
          "id": 778638,
          "author_name": "lisosia",
          "author_url": "",
          "post_date": "03/18/2020 15:54:14",
          "content": "<p>Thank you! \nActually, I checked and used some of your SEResnext kernel codes for my first step!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 778532,
      "author_name": "kasim0226",
      "author_url": "",
      "post_date": "03/18/2020 14:33:06",
      "content": "<p>Congrats!\nAfter reading the link you provided, still don't know how to decompose grapheme.\nCan you explain more detail?</p>",
      "votes": null,
      "replies": [
        {
          "id": 778629,
          "author_name": "lisosia",
          "author_url": "",
          "post_date": "03/18/2020 15:49:06",
          "content": "<p>Thanks!\nI added details for decomposition. Please feel free to ask me questions!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 781194,
          "author_name": "kasim0226",
          "author_url": "",
          "post_date": "03/21/2020 02:42:32",
          "content": "<p>Hi <a href=\"/lisosia\">@lisosia</a> \nThank you for your reply.\nCould I have further questions?</p>\n\n<ol>\n<li><p>What is your loss weight of the 61 BinaryCrossEntropy and  loss weight of the cross entropy of root, vowel, consotant?</p></li>\n<li><p>Why you divide np.power(class_count, 1) ? for unbalanced classes?\n<code>argmax( softmax(logits) / np.power(class_count, 1) )</code></p></li>\n<li><p>Which layer you inserted DropBlock? and the parameters?</p></li>\n</ol>\n\n<p>Looking forward to your reply</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 784837,
          "author_name": "lisosia",
          "author_url": "",
          "post_date": "03/24/2020 14:34:26",
          "content": "<p>1.\nroot : vowel : consotant : sum-of-61-BCEs = 2:1:1:1/40</p>\n\n<p>2.\nwell describeld in <a href=\"https://www.kaggle.com/c/bengaliai-cv19/discussion/136021\">Chris Writeup</a></p>\n\n<p>Suppose you are given sample XX with class i (you donot know the correct class).\nIf you correctly predict class i for sample XX, you will get additional score of <code>1 / (sample number of class i in test-set)</code> which can be approximated by <code>1 / (sample number of class i in train-set)</code>.</p>\n\n<p>softmax() output for sample XX is probabiliy.\nif i-th softmax output for sample XX is <code>p_i</code> , the probabiliy of sample XX is in class i is <code>p_i</code> .\nIf you predict class i for sample XX, \nYour <em>expected gain of score</em> is\n<code>p_i * 1 / (sample number of class i in train-set)</code> = <code>softmax(logit)[i] / (sample number of class i in trainset)</code></p>\n\n<p>So the best prediction for sample XX is \n<code>argmax-of-i  ( softmax(logit)[i] / (sample number of class i in train-set)</code> )</p>\n\n<p>3.\nafter layer0, 1, 2. same as <a href=\"https://www.kaggle.com/c/bengaliai-cv19/discussion/132898#763633\">the post</a> I mentioned in the writeup.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 779091,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "03/19/2020 02:47:51",
      "content": "<p>you can use the head from <a href=\"https://www.kaggle.com/c/bengaliai-cv19/discussion/135990\">https://www.kaggle.com/c/bengaliai-cv19/discussion/135990</a>,\nwhich include the arcface loss with my modified se-resnext50 backbone.</p>\n\n<p>the result is quite amazing. here are the results i get:</p>\n\n<p>reference (<a href=\"/qishen\">@qishen</a> ha implementation): efficientnetb1-128x128 public lb 0.985\nmy implementation:  efficientnetb1-128x128 local cv 0.995 (estimated lb 0.985)\nmy implementation:  modified se-resnext50-112x112 local cv 0.997 (estimated lb 0.987)</p>\n\n<hr>\n\n<p>increasing 112x112 to 128x128 can add 0.001 in my previous experiements.\nusing arcface head + your decompose head may further improve results?</p>",
      "votes": null,
      "replies": [
        {
          "id": 779886,
          "author_name": "lisosia",
          "author_url": "",
          "post_date": "03/19/2020 19:24:34",
          "content": "<p>Wow, arcface really works!\nIn my experiment, the contribution of decompose head was minor. The boost was around +0.0010:\n  CV(0.9804+0.0011), public LB(0.9751+0.0005), private LB(0.9422+0.0011).</p>\n\n<p>I believe that adding decomposed label (which has richer information) will not harm training under any conditions even though the boost might be small.</p>\n\n<hr>\n\n<p>Actually, I anticipated decomposed head serves as a kind of metric learning.\nCombining decomposed head with arcface may yield better learning of inter-grapheme distance.</p>\n\n<p>(Important note: I am not familiar with metric learning AT ALL. I will start learning acrface from now :)  )</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 779105,
      "author_name": "cdeotte",
      "author_url": "",
      "post_date": "03/19/2020 03:04:34",
      "content": "<p>Congratulations. Awesome solo gold. I like your decomposed grapheme head. That's smart! I didn't realize that Python could decompose graphemes nor that graphemes could be decomposed. Post processing was very powerful in this comp, we found <code>EXP = -1.2</code> to be best too.</p>",
      "votes": null,
      "replies": [
        {
          "id": 784793,
          "author_name": "lisosia",
          "author_url": "",
          "post_date": "03/24/2020 14:05:35",
          "content": "<p>Thank you and Congrats to you too!\nWe followed a similar path in the competition :)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 779701,
      "author_name": "thefatcat",
      "author_url": "",
      "post_date": "03/19/2020 15:56:16",
      "content": "<p>Thanks for your sharing,and I want to ask what's the loss weight of your final choice?</p>",
      "votes": null,
      "replies": [
        {
          "id": 784803,
          "author_name": "lisosia",
          "author_url": "",
          "post_date": "03/24/2020 14:09:25",
          "content": "<p>root : vowel : consotant : sum-of-61-BCEs = 2:1:1:1/40</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "776628": "updated 2020/03/19\n\n----\n\nThanks to the hosts and helpful discussions.\nEspecially, I really appreciated @hengck23's kind sharing.\n\nI am so lucky to get gold.\nI expected random-split public/private and did nothing special for unseen grapheme.\n\nIt seems postprocessing for macro recall hack is really important.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2796584%2Fd7d8dab30ab64980f8b92edbf1c29af6%2Fbengali2.png?generation=1584544815730850&amp;alt=media)\n## Model\n- modified SE-ResNeXt50. 2-stride conv is replaced by 1-stride-conv+maxblur-pool\n   basically following @hengck23 's [advice](https://www.kaggle.com/c/bengaliai-cv19/discussion/123757#757358)\n- phase1: trained with stem-conv(stride=2), phase2: replace stem-conv with 3 convs(stride=1)\n- Adam + ReduceLROnPlateau, about 200 epochs\n- 5 fold ensemble\n## Loss\n- CrossEntropy for root, vowel, consonant\n  - 2 out of 5 folds are finetuned using ohem loss\n- Multilabel BinaryCrossEntropy to classify decomposed grapheme exists or not\n  - graphme (or root or vowel or consotant) can be decomposed as shown in the figure. 61 possible decomposed parts exist (except '0' part). below is my exact code to make multihot label.\n\n```\n# load data\ntrain = pd.read_csv(C.datadir/'train.csv')\ntrain_labels = train[['grapheme_root', 'vowel_diacritic', 'consonant_diacritic']].values.astype(np.int64)\n\n# components label\nparts_pre = pd.read_csv(C.datadir/'class_map.csv').component\nparts = np.sort(np.unique(np.concatenate([list(e) for e in parts_pre])))\nparts = parts[parts != '0']  # 0 has no meanning\nprint(\"parts:\", parts)  # 61 parts shown\n\ntrain_labels_comp = []\nfor grapheme in train['grapheme'].values:\n    train_labels_comp.append([part in list(grapheme) for part in parts])\ntrain_labels_comp = np.array(train_labels_comp).astype(np.int64)\nprint(\"train_labels_comp.shape\", train_labels_comp.shape)\nif True: # debug\n    print(\"train_labels_comp\", train_labels_comp[0].tolist())\n    print(\"train_labels_comp\", train_labels_comp[1].tolist())\n    print(\"train_labels_comp\", train_labels_comp[2].tolist())\n```\n\n## Preprocess and Augmentation\n100% same as @hengck23 's [advice](https://www.kaggle.com/c/bengaliai-cv19/discussion/132898#763633)\nPreprocess :  just resizing with cv2.INTER_AREA\nAugmentation : OneOf(basic augmentations) + DropBlock\n## Postprocess\n- prediction for decomposed grapheme is not used\n- **maximize expected recall by\n  argmax( softmax(logits) / np.power(class_count, 1) )**. note that final sub use factor of 1.15 for grapheme root(no difference for private lb).\n  - there may be better threshold optimization since confusion matrix still looks asymmetry after above optimization. But I gave up further improvement due to low sample size.\n  - below is plot of \"macro recall scores V.S. factor n\" for my model (which is not final sub). You can confirm that the peak is around factor=1. You can also find that the improvement is not huge unlike private LB.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2796584%2F26fd5a4a563451455a39333fe670aab4%2Fmodel033_recall_hack.png?generation=1584545521502984&amp;alt=media)\n\n-----\n**What I learned during competition ( note to self )**\n*how to use pytorch lightning / how to freeze specified layer / how to ensemble / concept of snapshot ensemble / ohem loss, class-balanced loss / hengck23's model surgery method / maxblur-pool and its implementation / keep high resolution may be important for low resolution image / interpolation method for resizing is sometimes important / how to implement some augmentation methods (mixup, cutmix, gridmask) / public &amp; private is not always split randomly / vast.ai is low-priced compared to GCP*",
    "776637": "Congrats on your gold finish!! I've a question, were you using Heng's code too??",
    "776643": "most of augmentation and preprocess codes are from his implementation. ( I implement my own, but discarded it since his augmentation was better)",
    "776647": "Congrats, may I ask how do you use decomposed grapheme when test image?",
    "776671": "good work!\n\ni will try to repeat your experiment as well :)",
    "776680": "Not used for test. just for auxiliary task",
    "776700": "Ok, Thanks for your sharing",
    "776776": "Nice!  I like the unicode prediction block.",
    "776892": "Congrats! Thanks for sharing your solution! I am confused about the decomposed 61 channels. Could you please explain what do you do extractly to get them？I didn't get it. Thanks again!",
    "777650": "co9ngrats.  how do you decompose grapheme?",
    "777851": "Congrats and thank you for sharing.",
    "777948": "Thanks for your question!\ndecomposition is from  https://www.kaggle.com/c/bengaliai-cv19/discussion/123757#765573\n\nI will add some explanation later.",
    "777949": "Thanks for your question!\ndecomposition is from  https://www.kaggle.com/c/bengaliai-cv19/discussion/123757#765573\n\nI will add some explanation later.",
    "778532": "Congrats!\nAfter reading the link you provided, still don't know how to decompose grapheme.\nCan you explain more detail?",
    "778618": "I added details. Please feel free to ask me questions!",
    "778619": "I added details for decomposition. Please feel free to ask me questions!",
    "778629": "Thanks!\nI added details for decomposition. Please feel free to ask me questions!",
    "778633": "Thank you!\nI was really impressed by your creative posts!",
    "778638": "Thank you! \nActually, I checked and used some of your SEResnext kernel codes for my first step!",
    "778640": "Thank you! Unicode encoding is really informative.",
    "779082": "Thanks for your reply! I get it this time. ：）",
    "779091": "you can use the head from https://www.kaggle.com/c/bengaliai-cv19/discussion/135990,\nwhich include the arcface loss with my modified se-resnext50 backbone.\n\nthe result is quite amazing. here are the results i get:\n\nreference (@qishen ha implementation): efficientnetb1-128x128 public lb 0.985\nmy implementation:  efficientnetb1-128x128 local cv 0.995 (estimated lb 0.985)\nmy implementation:  modified se-resnext50-112x112 local cv 0.997 (estimated lb 0.987)\n\n---\n\nincreasing 112x112 to 128x128 can add 0.001 in my previous experiements.\nusing arcface head + your decompose head may further improve results?",
    "779105": "Congratulations. Awesome solo gold. I like your decomposed grapheme head. That's smart! I didn't realize that Python could decompose graphemes nor that graphemes could be decomposed. Post processing was very powerful in this comp, we found `EXP = -1.2` to be best too.",
    "779701": "Thanks for your sharing,and I want to ask what's the loss weight of your final choice?",
    "779886": "Wow, arcface really works!\nIn my experiment, the contribution of decompose head was minor. The boost was around +0.0010:\n  CV(0.9804+0.0011), public LB(0.9751+0.0005), private LB(0.9422+0.0011).\n\nI believe that adding decomposed label (which has richer information) will not harm training under any conditions even though the boost might be small.\n\n---\n\nActually, I anticipated decomposed head serves as a kind of metric learning.\nCombining decomposed head with arcface may yield better learning of inter-grapheme distance.\n\n(Important note: I am not familiar with metric learning AT ALL. I will start learning acrface from now :)  )",
    "781194": "Hi @lisosia \nThank you for your reply.\nCould I have further questions?\n\n1. What is your loss weight of the 61 BinaryCrossEntropy and  loss weight of the cross entropy of root, vowel, consotant?\n\n2. Why you divide np.power(class_count, 1) ? for unbalanced classes?\n`argmax( softmax(logits) / np.power(class_count, 1) )`\n\n3. Which layer you inserted DropBlock? and the parameters?\n\nLooking forward to your reply",
    "784793": "Thank you and Congrats to you too!\nWe followed a similar path in the competition :)",
    "784803": "root : vowel : consotant : sum-of-61-BCEs = 2:1:1:1/40",
    "784837": "1.\nroot : vowel : consotant : sum-of-61-BCEs = 2:1:1:1/40\n\n2.\nwell describeld in [Chris Writeup](https://www.kaggle.com/c/bengaliai-cv19/discussion/136021)\n\nSuppose you are given sample XX with class i (you donot know the correct class).\nIf you correctly predict class i for sample XX, you will get additional score of `1 / (sample number of class i in test-set)` which can be approximated by `1 / (sample number of class i in train-set)`.\n\nsoftmax() output for sample XX is probabiliy.\nif i-th softmax output for sample XX is `p_i` , the probabiliy of sample XX is in class i is `p_i` .\nIf you predict class i for sample XX, \nYour *expected gain of score* is\n`p_i * 1 / (sample number of class i in train-set)` = `softmax(logit)[i] / (sample number of class i in trainset)`\n\nSo the best prediction for sample XX is \n` argmax-of-i  ( softmax(logit)[i] / (sample number of class i in train-set)` )\n\n3.\nafter layer0, 1, 2. same as [the post](https://www.kaggle.com/c/bengaliai-cv19/discussion/132898#763633) I mentioned in the writeup."
  },
  "source": "meta"
}