{
  "id": 135982,
  "title": "3rd place solution",
  "url": "/competitions/bengaliai-cv19/discussion/135982",
  "author_name": "phalanx",
  "post_date": "2020-03-17T01:31:18.604000",
  "votes": 85,
  "comment_count": 25,
  "views": 0,
  "content": "<p>Congrats everyone with excellent result.</p>\n\n<p>Here is our solution summary.</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1620223%2Fb2143c50d9c92271e7b871912dd171d7%2Fbengali_pipeline%20.jpg?generation=1584447993610250&amp;alt=media\" alt=\"\"></p>\n\n<h1>Solution summary</h1>\n\n<h2>Dataset</h2>\n\n<ul>\n<li><strong>pretrain</strong>: triple identities by hflip and vflip</li>\n<li>image resolution: 137 x236</li>\n</ul>\n\n<h2>Model</h2>\n\n<p><strong>for seen grapheme and unseen grapheme</strong>\n* (phalanx) encoder -&gt; gempool -&gt; batch_norm -&gt; fc\n* (earhian) encoder -&gt; avgpool -&gt; batch_norm -&gt; dropout -&gt; fc</p>\n\n<p><strong>arcface</strong>\n* encoder -&gt; avgpool -&gt; conv1d -&gt; bn\n* s 32(train), 1.0(test)\n* m 0.5</p>\n\n<p><strong>stacking</strong>\n* conv -&gt; relu -&gt; dropout -&gt; fc</p>\n\n<h2>Augmentation</h2>\n\n<ul>\n<li>cutmix</li>\n<li>shift, scale, rotate, shear</li>\n</ul>\n\n<h2>Training</h2>\n\n<ul>\n<li><strong>seen grapheme</strong>: train model for seen grapheme with pretrain dataset, then finetune it with original dataset</li>\n<li><strong>arcface and unseen grapheme</strong>: train pretrained model for seen grapheme with original dataset</li>\n<li>replace softmax with <a href=\"https://arxiv.org/abs/1911.10688\">pc-softmax</a></li>\n<li>loss function: negative log likelihood</li>\n<li>SGD with CosineAnnealing</li>\n<li>Stochastic Weighted Average</li>\n</ul>\n\n<h2>Inference</h2>\n\n<p><strong>arcface</strong>\n* use cosine similarity between train and test embedding feature\n* threshold: smallest cosine similarity between train and validation embedding feature</p>\n\n<p><strong></strong></p>",
  "messages": [
    {
      "id": 775834,
      "postDate": "2020-03-17T01:31:18.603Z",
      "content": "<p>Congrats everyone with excellent result.</p>\n\n<p>Here is our solution summary.</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1620223%2Fb2143c50d9c92271e7b871912dd171d7%2Fbengali_pipeline%20.jpg?generation=1584447993610250&amp;alt=media\" alt=\"\"></p>\n\n<h1>Solution summary</h1>\n\n<h2>Dataset</h2>\n\n<ul>\n<li><strong>pretrain</strong>: triple identities by hflip and vflip</li>\n<li>image resolution: 137 x236</li>\n</ul>\n\n<h2>Model</h2>\n\n<p><strong>for seen grapheme and unseen grapheme</strong>\n* (phalanx) encoder -&gt; gempool -&gt; batch_norm -&gt; fc\n* (earhian) encoder -&gt; avgpool -&gt; batch_norm -&gt; dropout -&gt; fc</p>\n\n<p><strong>arcface</strong>\n* encoder -&gt; avgpool -&gt; conv1d -&gt; bn\n* s 32(train), 1.0(test)\n* m 0.5</p>\n\n<p><strong>stacking</strong>\n* conv -&gt; relu -&gt; dropout -&gt; fc</p>\n\n<h2>Augmentation</h2>\n\n<ul>\n<li>cutmix</li>\n<li>shift, scale, rotate, shear</li>\n</ul>\n\n<h2>Training</h2>\n\n<ul>\n<li><strong>seen grapheme</strong>: train model for seen grapheme with pretrain dataset, then finetune it with original dataset</li>\n<li><strong>arcface and unseen grapheme</strong>: train pretrained model for seen grapheme with original dataset</li>\n<li>replace softmax with <a href=\"https://arxiv.org/abs/1911.10688\">pc-softmax</a></li>\n<li>loss function: negative log likelihood</li>\n<li>SGD with CosineAnnealing</li>\n<li>Stochastic Weighted Average</li>\n</ul>\n\n<h2>Inference</h2>\n\n<p><strong>arcface</strong>\n* use cosine similarity between train and test embedding feature\n* threshold: smallest cosine similarity between train and validation embedding feature</p>\n\n<p><strong></strong></p>",
      "rawMarkdown": "Congrats everyone with excellent result.\n\nHere is our solution summary.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1620223%2Fb2143c50d9c92271e7b871912dd171d7%2Fbengali_pipeline%20.jpg?generation=1584447993610250&amp;alt=media)\n\n\n# Solution summary\n## Dataset\n* <strong>pretrain</strong>: triple identities by hflip and vflip\n* image resolution: 137 x236\n\n## Model\n<strong>for seen grapheme and unseen grapheme</strong>\n* (phalanx) encoder -&gt; gempool -&gt; batch_norm -&gt; fc\n* (earhian) encoder -&gt; avgpool -&gt; batch_norm -&gt; dropout -&gt; fc\n\n<strong>arcface</strong>\n* encoder -&gt; avgpool -&gt; conv1d -&gt; bn\n* s 32(train), 1.0(test)\n* m 0.5\n\n<strong>stacking</strong>\n* conv -&gt; relu -&gt; dropout -&gt; fc\n\n## Augmentation\n* cutmix\n* shift, scale, rotate, shear\n\n## Training\n* <strong>seen grapheme</strong>: train model for seen grapheme with pretrain dataset, then finetune it with original dataset\n* <strong>arcface and unseen grapheme</strong>: train pretrained model for seen grapheme with original dataset\n* replace softmax with [pc-softmax](https://arxiv.org/abs/1911.10688)\n* loss function: negative log likelihood\n* SGD with CosineAnnealing\n* Stochastic Weighted Average\n\n\n## Inference\n<strong>arcface</strong>\n* use cosine similarity between train and test embedding feature\n* threshold: smallest cosine similarity between train and validation embedding feature\n\n<strong></strong>",
      "votes": 84
    },
    {
      "id": 775883,
      "postDate": "2020-03-17T02:17:57.290Z",
      "content": "<p>if i understand correctly:</p>\n\n<ol>\n<li>use arcface (metric learning) to determine if a input test samples is seen or unseen in train dataset.\n(but how to use arcface to test if seen or unseen? how to set threshold?)</li>\n<li>if seen, use shared feature for multi-task</li>\n<li>if unseen, it is better not to use shared feature. this prevents overfitting. hence train separate models for each root,vowel, const.</li>\n</ol>\n\n<p>is the understanding correct?</p>",
      "rawMarkdown": "if i understand correctly:\n\n1. use arcface (metric learning) to determine if a input test samples is seen or unseen in train dataset.\n(but how to use arcface to test if seen or unseen? how to set threshold?)\n2. if seen, use shared feature for multi-task\n3. if unseen, it is better not to use shared feature. this prevents overfitting. hence train separate models for each root,vowel, const.\n\nis the understanding correct?",
      "votes": 5,
      "replies": [
        {
          "id": 775914,
          "postDate": "2020-03-17T02:48:15.217Z",
          "content": "<ol>\n<li>use cosine similarity between train and test feature. about threshold, we use smallest cosine similarity between train and validation feature.</li>\n</ol>\n\n<p>2., 3. you're right</p>",
          "rawMarkdown": "1. use cosine similarity between train and test feature. about threshold, we use smallest cosine similarity between train and validation feature.\n\n2., 3. you're right"
        },
        {
          "id": 775920,
          "postDate": "2020-03-17T02:55:58.153Z",
          "content": "<p>thanks for the reply!</p>\n\n<p>very nice works and congrats to your team !</p>",
          "rawMarkdown": "thanks for the reply!\n\nvery nice works and congrats to your team !"
        }
      ]
    },
    {
      "id": 780016,
      "postDate": "2020-03-19T22:47:47.943Z",
      "content": "<p>congratulations！\ntwo question:\n1 How does stacking for unseen work? b3/b4/b5 respectively outputs 3 results, 3*3 results are concatenated?\n2 I can not understand the sentence \"arcface and unseen grapheme: train pretrained model for seen grapheme with original dataset\" and how is the model for unseen grapheme trained.\nThanks</p>",
      "rawMarkdown": "congratulations！\ntwo question:\n1 How does stacking for unseen work? b3/b4/b5 respectively outputs 3 results, 3*3 results are concatenated?\n2 I can not understand the sentence \"arcface and unseen grapheme: train pretrained model for seen grapheme with original dataset\" and how is the model for unseen grapheme trained.\nThanks",
      "votes": 1
    },
    {
      "id": 775916,
      "postDate": "2020-03-17T02:49:04.317Z",
      "content": "<p>Congrats Phalanx, your solution is elegant!</p>",
      "rawMarkdown": "Congrats Phalanx, your solution is elegant!",
      "votes": 1
    },
    {
      "id": 789149,
      "postDate": "2020-03-28T12:55:25.917Z",
      "content": "<p>I am very confused with the difference between the seen grapheme and unseen grapheme.\nCould you explain the both definition?</p>",
      "rawMarkdown": "I am very confused with the difference between the seen grapheme and unseen grapheme.\nCould you explain the both definition?"
    },
    {
      "id": 779384,
      "postDate": "2020-03-19T09:40:12.150Z",
      "content": "<p>Hi, congratulations to your team :) I am interested to your solution with perspective to implement it by myself. Will you provide some additional details about the solution? For instance the Arcface? How was it trained with metric learning? What the inputs / labels for it needed? How seen/unseen split was made? Thank for posting your approaches!</p>",
      "rawMarkdown": "Hi, congratulations to your team :) I am interested to your solution with perspective to implement it by myself. Will you provide some additional details about the solution? For instance the Arcface? How was it trained with metric learning? What the inputs / labels for it needed? How seen/unseen split was made? Thank for posting your approaches!"
    },
    {
      "id": 778046,
      "postDate": "2020-03-18T05:23:16.663Z",
      "content": "<p>Congratulations</p>",
      "rawMarkdown": "Congratulations"
    },
    {
      "id": 777815,
      "postDate": "2020-03-18T00:58:45.817Z",
      "content": "<p>Congrats and thanks for sharing. How did you train unseen grapheme, what kind of dataset is used?</p>",
      "rawMarkdown": "Congrats and thanks for sharing. How did you train unseen grapheme, what kind of dataset is used?"
    },
    {
      "id": 777792,
      "postDate": "2020-03-17T23:53:11.337Z",
      "content": "<p>Congratulations !</p>",
      "rawMarkdown": "Congratulations !\n\n"
    },
    {
      "id": 776704,
      "postDate": "2020-03-17T15:24:37.953Z",
      "content": "<p>Hi Phalanx, congratulations! Two questions:\n1) I am wondering if you have considered using grapheme classifier instead of RCV classifier for seen graphemes? Discussions seem to indicate that grapheme classifier is better on seen graphemes. Wondering if your choice of RCV over grapheme classifier for predicted seen graphemes is intentional.\n2) For seen/unseen grapheme arcface model, I take that it is simply RCV classifier with dim==186 Linear layer at the end, which means that you have set equal s and m hyperparameter for arcface loss for R, C, and V?</p>",
      "rawMarkdown": "Hi Phalanx, congratulations! Two questions:\n1) I am wondering if you have considered using grapheme classifier instead of RCV classifier for seen graphemes? Discussions seem to indicate that grapheme classifier is better on seen graphemes. Wondering if your choice of RCV over grapheme classifier for predicted seen graphemes is intentional.\n2) For seen/unseen grapheme arcface model, I take that it is simply RCV classifier with dim==186 Linear layer at the end, which means that you have set equal s and m hyperparameter for arcface loss for R, C, and V?"
    },
    {
      "id": 776695,
      "postDate": "2020-03-17T15:10:06.217Z",
      "content": "<p>Thanks for sharing your pipeline. \n<code>triple identities by hflip and vflip</code> if I am correct, you converted all the n classes to 3n classes by introducing their flipped versions in the dataset? I can't understand how does this help. During inference, wouldn't 2/3 of the predictable classes be useless? or are you using TTA and then mapping flipped classes to originals?</p>\n\n<p>Very interesting idea to use arcface for classification into seen/unseen. </p>",
      "rawMarkdown": "Thanks for sharing your pipeline. \n`triple identities by hflip and vflip` if I am correct, you converted all the n classes to 3n classes by introducing their flipped versions in the dataset? I can't understand how does this help. During inference, wouldn't 2/3 of the predictable classes be useless? or are you using TTA and then mapping flipped classes to originals?\n\nVery interesting idea to use arcface for classification into seen/unseen. "
    },
    {
      "id": 776504,
      "postDate": "2020-03-17T12:58:28.623Z",
      "content": "<p>Awesome, congrats on the prize win.  Lots to learn from this solution.</p>",
      "rawMarkdown": "Awesome, congrats on the prize win.  Lots to learn from this solution."
    },
    {
      "id": 776426,
      "postDate": "2020-03-17T11:38:44.237Z",
      "content": "<p>Very nice solution, congrats! What exactly are your datasets for seen and unseen data?</p>",
      "rawMarkdown": "Very nice solution, congrats! What exactly are your datasets for seen and unseen data?",
      "replies": [
        {
          "id": 776448,
          "postDate": "2020-03-17T11:59:37.113Z",
          "content": "<p>Sorry, that is grapheme, not data. I already modified it.\nSo we predict whether grapheme class of test data is included in train grapheme or not.</p>",
          "rawMarkdown": "Sorry, that is grapheme, not data. I already modified it.\nSo we predict whether grapheme class of test data is included in train grapheme or not."
        },
        {
          "id": 776681,
          "postDate": "2020-03-17T14:50:25.020Z",
          "content": "<p>And on which data do you get the thresholds for that?</p>",
          "rawMarkdown": "And on which data do you get the thresholds for that?"
        }
      ]
    },
    {
      "id": 776418,
      "postDate": "2020-03-17T11:30:38.853Z",
      "content": "<p>write solution summary.\nfeel free to comment if you have question.</p>",
      "rawMarkdown": "write solution summary.\nfeel free to comment if you have question."
    },
    {
      "id": 775966,
      "postDate": "2020-03-17T03:48:21.043Z",
      "content": "<p>Congratulations! Your pipeline is quite similar with mine 😃 </p>",
      "rawMarkdown": "Congratulations! Your pipeline is quite similar with mine 😃 ",
      "replies": [
        {
          "id": 775981,
          "postDate": "2020-03-17T04:02:40.970Z",
          "content": "<p>Thanks for your post 👀 </p>",
          "rawMarkdown": "Thanks for your post 👀 ",
          "votes": 1
        }
      ]
    },
    {
      "id": 775941,
      "postDate": "2020-03-17T03:16:54.563Z",
      "content": "<p>Congratulations !</p>",
      "rawMarkdown": "Congratulations !"
    },
    {
      "id": 775925,
      "postDate": "2020-03-17T02:59:14.707Z",
      "content": "<p>Congrats! Thanks for sharing! Waiting for your details~</p>",
      "rawMarkdown": "Congrats! Thanks for sharing! Waiting for your details~"
    },
    {
      "id": 775859,
      "postDate": "2020-03-17T01:56:42.660Z",
      "content": "<p>congratulations, I learned a lot in this competition..🤓 </p>",
      "rawMarkdown": "congratulations, I learned a lot in this competition..🤓 "
    },
    {
      "id": 781478,
      "postDate": "2020-03-21T10:55:32.063Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 787405,
      "postDate": "2020-03-26T19:02:52.567Z",
      "content": "<p>Thanks for sharing !</p>",
      "rawMarkdown": "Thanks for sharing !"
    },
    {
      "id": 775836,
      "postDate": "2020-03-17T01:33:48.850Z",
      "content": "<p>thanks！</p>",
      "rawMarkdown": "thanks！"
    }
  ],
  "comments": [
    {
      "id": 775883,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2020-03-17T02:17:57.290000",
      "content": "<p>if i understand correctly:</p>\n\n<ol>\n<li>use arcface (metric learning) to determine if a input test samples is seen or unseen in train dataset.\n(but how to use arcface to test if seen or unseen? how to set threshold?)</li>\n<li>if seen, use shared feature for multi-task</li>\n<li>if unseen, it is better not to use shared feature. this prevents overfitting. hence train separate models for each root,vowel, const.</li>\n</ol>\n\n<p>is the understanding correct?</p>",
      "votes": 5,
      "replies": [
        {
          "id": 775914,
          "author_name": "phalanx",
          "author_url": "",
          "post_date": "2020-03-17T02:48:15.217000",
          "content": "<ol>\n<li>use cosine similarity between train and test feature. about threshold, we use smallest cosine similarity between train and validation feature.</li>\n</ol>\n\n<p>2., 3. you're right</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 775920,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2020-03-17T02:55:58.153000",
          "content": "<p>thanks for the reply!</p>\n\n<p>very nice works and congrats to your team !</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 780016,
      "author_name": "Mingxing Liu",
      "author_url": "",
      "post_date": "2020-03-19T22:47:47.943000",
      "content": "<p>congratulations！\ntwo question:\n1 How does stacking for unseen work? b3/b4/b5 respectively outputs 3 results, 3*3 results are concatenated?\n2 I can not understand the sentence \"arcface and unseen grapheme: train pretrained model for seen grapheme with original dataset\" and how is the model for unseen grapheme trained.\nThanks</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 775916,
      "author_name": "Yiheng Wang",
      "author_url": "",
      "post_date": "2020-03-17T02:49:04.317000",
      "content": "<p>Congrats Phalanx, your solution is elegant!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 789149,
      "author_name": "Yeonsu",
      "author_url": "",
      "post_date": "2020-03-28T12:55:25.917000",
      "content": "<p>I am very confused with the difference between the seen grapheme and unseen grapheme.\nCould you explain the both definition?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 779384,
      "author_name": "artur.k.space",
      "author_url": "",
      "post_date": "2020-03-19T09:40:12.150000",
      "content": "<p>Hi, congratulations to your team :) I am interested to your solution with perspective to implement it by myself. Will you provide some additional details about the solution? For instance the Arcface? How was it trained with metric learning? What the inputs / labels for it needed? How seen/unseen split was made? Thank for posting your approaches!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 778046,
      "author_name": "He",
      "author_url": "",
      "post_date": "2020-03-18T05:23:16.663000",
      "content": "<p>Congratulations</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 777815,
      "author_name": "corochann",
      "author_url": "",
      "post_date": "2020-03-18T00:58:45.817000",
      "content": "<p>Congrats and thanks for sharing. How did you train unseen grapheme, what kind of dataset is used?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 777792,
      "author_name": "Naim Houes ",
      "author_url": "",
      "post_date": "2020-03-17T23:53:11.337000",
      "content": "<p>Congratulations !</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 776704,
      "author_name": "Nicholas Lyu",
      "author_url": "",
      "post_date": "2020-03-17T15:24:37.953000",
      "content": "<p>Hi Phalanx, congratulations! Two questions:\n1) I am wondering if you have considered using grapheme classifier instead of RCV classifier for seen graphemes? Discussions seem to indicate that grapheme classifier is better on seen graphemes. Wondering if your choice of RCV over grapheme classifier for predicted seen graphemes is intentional.\n2) For seen/unseen grapheme arcface model, I take that it is simply RCV classifier with dim==186 Linear layer at the end, which means that you have set equal s and m hyperparameter for arcface loss for R, C, and V?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 776695,
      "author_name": "timetraveller",
      "author_url": "",
      "post_date": "2020-03-17T15:10:06.217000",
      "content": "<p>Thanks for sharing your pipeline. \n<code>triple identities by hflip and vflip</code> if I am correct, you converted all the n classes to 3n classes by introducing their flipped versions in the dataset? I can't understand how does this help. During inference, wouldn't 2/3 of the predictable classes be useless? or are you using TTA and then mapping flipped classes to originals?</p>\n\n<p>Very interesting idea to use arcface for classification into seen/unseen. </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 776504,
      "author_name": "CPMP",
      "author_url": "",
      "post_date": "2020-03-17T12:58:28.623000",
      "content": "<p>Awesome, congrats on the prize win.  Lots to learn from this solution.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 776426,
      "author_name": "Psi",
      "author_url": "",
      "post_date": "2020-03-17T11:38:44.237000",
      "content": "<p>Very nice solution, congrats! What exactly are your datasets for seen and unseen data?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 776448,
          "author_name": "phalanx",
          "author_url": "",
          "post_date": "2020-03-17T11:59:37.113000",
          "content": "<p>Sorry, that is grapheme, not data. I already modified it.\nSo we predict whether grapheme class of test data is included in train grapheme or not.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 776681,
          "author_name": "Psi",
          "author_url": "",
          "post_date": "2020-03-17T14:50:25.020000",
          "content": "<p>And on which data do you get the thresholds for that?</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 776418,
      "author_name": "phalanx",
      "author_url": "",
      "post_date": "2020-03-17T11:30:38.853000",
      "content": "<p>write solution summary.\nfeel free to comment if you have question.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 775966,
      "author_name": "Qishen Ha",
      "author_url": "",
      "post_date": "2020-03-17T03:48:21.043000",
      "content": "<p>Congratulations! Your pipeline is quite similar with mine 😃 </p>",
      "votes": 0,
      "replies": [
        {
          "id": 775981,
          "author_name": "earhian",
          "author_url": "",
          "post_date": "2020-03-17T04:02:40.970000",
          "content": "<p>Thanks for your post 👀 </p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 775941,
      "author_name": "Utsav Nandi",
      "author_url": "",
      "post_date": "2020-03-17T03:16:54.563000",
      "content": "<p>Congratulations !</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 775925,
      "author_name": "Helen",
      "author_url": "",
      "post_date": "2020-03-17T02:59:14.707000",
      "content": "<p>Congrats! Thanks for sharing! Waiting for your details~</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 775859,
      "author_name": "Gary",
      "author_url": "",
      "post_date": "2020-03-17T01:56:42.660000",
      "content": "<p>congratulations, I learned a lot in this competition..🤓 </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 781478,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-03-21T10:55:32.063000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 787405,
      "author_name": "Areyana",
      "author_url": "",
      "post_date": "2020-03-26T19:02:52.567000",
      "content": "<p>Thanks for sharing !</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 775836,
      "author_name": "Finlay",
      "author_url": "",
      "post_date": "2020-03-17T01:33:48.850000",
      "content": "<p>thanks！</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "775834": "Congrats everyone with excellent result.\n\nHere is our solution summary.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1620223%2Fb2143c50d9c92271e7b871912dd171d7%2Fbengali_pipeline%20.jpg?generation=1584447993610250&amp;alt=media)\n\n\n# Solution summary\n## Dataset\n* <strong>pretrain</strong>: triple identities by hflip and vflip\n* image resolution: 137 x236\n\n## Model\n<strong>for seen grapheme and unseen grapheme</strong>\n* (phalanx) encoder -&gt; gempool -&gt; batch_norm -&gt; fc\n* (earhian) encoder -&gt; avgpool -&gt; batch_norm -&gt; dropout -&gt; fc\n\n<strong>arcface</strong>\n* encoder -&gt; avgpool -&gt; conv1d -&gt; bn\n* s 32(train), 1.0(test)\n* m 0.5\n\n<strong>stacking</strong>\n* conv -&gt; relu -&gt; dropout -&gt; fc\n\n## Augmentation\n* cutmix\n* shift, scale, rotate, shear\n\n## Training\n* <strong>seen grapheme</strong>: train model for seen grapheme with pretrain dataset, then finetune it with original dataset\n* <strong>arcface and unseen grapheme</strong>: train pretrained model for seen grapheme with original dataset\n* replace softmax with [pc-softmax](https://arxiv.org/abs/1911.10688)\n* loss function: negative log likelihood\n* SGD with CosineAnnealing\n* Stochastic Weighted Average\n\n\n## Inference\n<strong>arcface</strong>\n* use cosine similarity between train and test embedding feature\n* threshold: smallest cosine similarity between train and validation embedding feature\n\n<strong></strong>",
    "775883": "if i understand correctly:\n\n1. use arcface (metric learning) to determine if a input test samples is seen or unseen in train dataset.\n(but how to use arcface to test if seen or unseen? how to set threshold?)\n2. if seen, use shared feature for multi-task\n3. if unseen, it is better not to use shared feature. this prevents overfitting. hence train separate models for each root,vowel, const.\n\nis the understanding correct?",
    "780016": "congratulations！\ntwo question:\n1 How does stacking for unseen work? b3/b4/b5 respectively outputs 3 results, 3*3 results are concatenated?\n2 I can not understand the sentence \"arcface and unseen grapheme: train pretrained model for seen grapheme with original dataset\" and how is the model for unseen grapheme trained.\nThanks",
    "775916": "Congrats Phalanx, your solution is elegant!",
    "789149": "I am very confused with the difference between the seen grapheme and unseen grapheme.\nCould you explain the both definition?",
    "779384": "Hi, congratulations to your team :) I am interested to your solution with perspective to implement it by myself. Will you provide some additional details about the solution? For instance the Arcface? How was it trained with metric learning? What the inputs / labels for it needed? How seen/unseen split was made? Thank for posting your approaches!",
    "778046": "Congratulations",
    "777815": "Congrats and thanks for sharing. How did you train unseen grapheme, what kind of dataset is used?",
    "777792": "Congratulations !\n\n",
    "776704": "Hi Phalanx, congratulations! Two questions:\n1) I am wondering if you have considered using grapheme classifier instead of RCV classifier for seen graphemes? Discussions seem to indicate that grapheme classifier is better on seen graphemes. Wondering if your choice of RCV over grapheme classifier for predicted seen graphemes is intentional.\n2) For seen/unseen grapheme arcface model, I take that it is simply RCV classifier with dim==186 Linear layer at the end, which means that you have set equal s and m hyperparameter for arcface loss for R, C, and V?",
    "776695": "Thanks for sharing your pipeline. \n`triple identities by hflip and vflip` if I am correct, you converted all the n classes to 3n classes by introducing their flipped versions in the dataset? I can't understand how does this help. During inference, wouldn't 2/3 of the predictable classes be useless? or are you using TTA and then mapping flipped classes to originals?\n\nVery interesting idea to use arcface for classification into seen/unseen. ",
    "776504": "Awesome, congrats on the prize win.  Lots to learn from this solution.",
    "776426": "Very nice solution, congrats! What exactly are your datasets for seen and unseen data?",
    "776418": "write solution summary.\nfeel free to comment if you have question.",
    "775966": "Congratulations! Your pipeline is quite similar with mine 😃 ",
    "775941": "Congratulations !",
    "775925": "Congrats! Thanks for sharing! Waiting for your details~",
    "775859": "congratulations, I learned a lot in this competition..🤓 ",
    "781478": "",
    "787405": "Thanks for sharing !",
    "775836": "thanks！"
  }
}