{
  "id": 135990,
  "title": "8th Place Solution /w Code (minimal ver.)",
  "url": "/competitions/bengaliai-cv19/discussion/135990",
  "author_name": "Qishen Ha",
  "post_date": "2020-03-17T01:54:44.166000",
  "votes": 173,
  "comment_count": 70,
  "views": 0,
  "content": "<p>Hi, everyone</p>\n\n<p>First of all, I want to thank Kaggle and the hosts for hosting the competition.</p>\n\n<p>Congratulation to all winners!</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F448347%2Fd37b2cfc0fd3524b5be0ba10e5d8b649%2F1.png?generation=1584409865722874&amp;alt=media\" alt=\"\"></p>\n\n<p>An illustration of my pipeline. It contains 2 models, Seen Model and Unseen Model.</p>\n\n<h1>Summary</h1>\n\n<ul>\n<li><p>NO EXTERNAL DATA</p></li>\n<li><p>I used 2-stage prediction in my pipeline, it contains 2 models as one set (Seen Model and Unseen Model) as shown in the figure above.</p></li>\n<li>I used Arcface to distinguish unseen graphemes. Just like to distinguish unseen faces ;)</li>\n<li>If an image is detected to be seen, I use the output of Seen Model as prediction directly. Else if the image is detected to be unseen, I’ll pass it to Unseen Model to get unseen prediction.</li>\n<li>The different points between Seen Model and Unseen Model are as follow:\n<ul><li>Seen Model only uses very few augmentation to make sure it can ‘overfit’ to seen graphemes, while Unseen Model uses heavy augmentations to make it generalize to unseen ones.</li>\n<li>Seen Model has 5-head outputs, including Arcface output. While Unseen Model has normal 4-head outputs.</li>\n<li>Beside the number of output heads, the architecture of the top is a little bit different as well.</li>\n<li>I decompose the predicted grapheme to 3 components when the prediction is made by Seen Model, while I use predicted 3 components directly when the prediction is made by Unseen Model.</li></ul></li>\n</ul>\n\n<h1>Code</h1>\n\n<p>I’ve published a minimal version of my training &amp; testing pipeline on Kaggle kernel as follow:</p>\n\n<ul>\n<li>Step 1. <a href=\"https://www.kaggle.com/haqishen/bengali-train-seen-model\">Bengali Train Seen Model</a> (trained ~5h on kernel)</li>\n<li>Step 2. <a href=\"https://www.kaggle.com/haqishen/bengali-train-unseen-model\">Bengali Train Unseen Model</a> (trained ~8.5h on kernel)</li>\n<li>Step 3. <a href=\"https://www.kaggle.com/haqishen/bengali-predict-with-seen-unseen-models\">Bengali Predict with Seen &amp; Unseen Models</a> (submission ~20min)</li>\n</ul>\n\n<p>The minimal version with two efficientnet-b1 which are training on 128x128 for 30 epochs can give you public LB 0.985+ or private LB 0.935+ \nIf you find my notebooks helpful, please upvote them! </p>\n\n<hr>\n\n<p>To get over 0.993+ or 0.950+ on public or private  LB, what you need to do is just simply:\n* Change the backbone to efficientnet-b5 / b6 / b7 / b8\n* Train on 224x224\n* Train for 60~90 epochs\n* Emsenble more Seen &amp; Unseen Model</p>\n\n<p>If you have any questions please feel free to leave a message to me.\nThanks!</p>",
  "messages": [
    {
      "id": 775855,
      "postDate": "2020-03-17T01:54:44.167Z",
      "content": "<p>Hi, everyone</p>\n\n<p>First of all, I want to thank Kaggle and the hosts for hosting the competition.</p>\n\n<p>Congratulation to all winners!</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F448347%2Fd37b2cfc0fd3524b5be0ba10e5d8b649%2F1.png?generation=1584409865722874&amp;alt=media\" alt=\"\"></p>\n\n<p>An illustration of my pipeline. It contains 2 models, Seen Model and Unseen Model.</p>\n\n<h1>Summary</h1>\n\n<ul>\n<li><p>NO EXTERNAL DATA</p></li>\n<li><p>I used 2-stage prediction in my pipeline, it contains 2 models as one set (Seen Model and Unseen Model) as shown in the figure above.</p></li>\n<li>I used Arcface to distinguish unseen graphemes. Just like to distinguish unseen faces ;)</li>\n<li>If an image is detected to be seen, I use the output of Seen Model as prediction directly. Else if the image is detected to be unseen, I’ll pass it to Unseen Model to get unseen prediction.</li>\n<li>The different points between Seen Model and Unseen Model are as follow:\n<ul><li>Seen Model only uses very few augmentation to make sure it can ‘overfit’ to seen graphemes, while Unseen Model uses heavy augmentations to make it generalize to unseen ones.</li>\n<li>Seen Model has 5-head outputs, including Arcface output. While Unseen Model has normal 4-head outputs.</li>\n<li>Beside the number of output heads, the architecture of the top is a little bit different as well.</li>\n<li>I decompose the predicted grapheme to 3 components when the prediction is made by Seen Model, while I use predicted 3 components directly when the prediction is made by Unseen Model.</li></ul></li>\n</ul>\n\n<h1>Code</h1>\n\n<p>I’ve published a minimal version of my training &amp; testing pipeline on Kaggle kernel as follow:</p>\n\n<ul>\n<li>Step 1. <a href=\"https://www.kaggle.com/haqishen/bengali-train-seen-model\">Bengali Train Seen Model</a> (trained ~5h on kernel)</li>\n<li>Step 2. <a href=\"https://www.kaggle.com/haqishen/bengali-train-unseen-model\">Bengali Train Unseen Model</a> (trained ~8.5h on kernel)</li>\n<li>Step 3. <a href=\"https://www.kaggle.com/haqishen/bengali-predict-with-seen-unseen-models\">Bengali Predict with Seen &amp; Unseen Models</a> (submission ~20min)</li>\n</ul>\n\n<p>The minimal version with two efficientnet-b1 which are training on 128x128 for 30 epochs can give you public LB 0.985+ or private LB 0.935+ \nIf you find my notebooks helpful, please upvote them! </p>\n\n<hr>\n\n<p>To get over 0.993+ or 0.950+ on public or private  LB, what you need to do is just simply:\n* Change the backbone to efficientnet-b5 / b6 / b7 / b8\n* Train on 224x224\n* Train for 60~90 epochs\n* Emsenble more Seen &amp; Unseen Model</p>\n\n<p>If you have any questions please feel free to leave a message to me.\nThanks!</p>",
      "rawMarkdown": "Hi, everyone\n\nFirst of all, I want to thank Kaggle and the hosts for hosting the competition.\n\nCongratulation to all winners!\n\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F448347%2Fd37b2cfc0fd3524b5be0ba10e5d8b649%2F1.png?generation=1584409865722874&amp;alt=media)\n\nAn illustration of my pipeline. It contains 2 models, Seen Model and Unseen Model.\n\n# Summary\n\n* NO EXTERNAL DATA\n\n* I used 2-stage prediction in my pipeline, it contains 2 models as one set (Seen Model and Unseen Model) as shown in the figure above.\n* I used Arcface to distinguish unseen graphemes. Just like to distinguish unseen faces ;)\n* If an image is detected to be seen, I use the output of Seen Model as prediction directly. Else if the image is detected to be unseen, I’ll pass it to Unseen Model to get unseen prediction.\n* The different points between Seen Model and Unseen Model are as follow:\n  * Seen Model only uses very few augmentation to make sure it can ‘overfit’ to seen graphemes, while Unseen Model uses heavy augmentations to make it generalize to unseen ones.\n  * Seen Model has 5-head outputs, including Arcface output. While Unseen Model has normal 4-head outputs.\n  * Beside the number of output heads, the architecture of the top is a little bit different as well.\n  * I decompose the predicted grapheme to 3 components when the prediction is made by Seen Model, while I use predicted 3 components directly when the prediction is made by Unseen Model.\n\n\n# Code\n\nI’ve published a minimal version of my training &amp; testing pipeline on Kaggle kernel as follow:\n\n* Step 1. [Bengali Train Seen Model](https://www.kaggle.com/haqishen/bengali-train-seen-model) (trained ~5h on kernel)\n* Step 2. [Bengali Train Unseen Model](https://www.kaggle.com/haqishen/bengali-train-unseen-model) (trained ~8.5h on kernel)\n* Step 3. [Bengali Predict with Seen &amp; Unseen Models](https://www.kaggle.com/haqishen/bengali-predict-with-seen-unseen-models) (submission ~20min)\n\nThe minimal version with two efficientnet-b1 which are training on 128x128 for 30 epochs can give you public LB 0.985+ or private LB 0.935+ \nIf you find my notebooks helpful, please upvote them! \n\n---\n\nTo get over 0.993+ or 0.950+ on public or private  LB, what you need to do is just simply:\n* Change the backbone to efficientnet-b5 / b6 / b7 / b8\n* Train on 224x224\n* Train for 60~90 epochs\n* Emsenble more Seen &amp; Unseen Model\n\nIf you have any questions please feel free to leave a message to me.\nThanks!\n\n",
      "votes": 172
    },
    {
      "id": 777919,
      "postDate": "2020-03-18T02:50:42.860Z",
      "content": "<p>Thanks for the sharing! Great idea of combining the power of two models for unseen and seen. I just have one question - could it be your threshold is a bit less aggressive if you rely on LB to select it? As we know that private have more unseen than public, so if you had chosen an even higher threshold to allow more test samples to go through unseen models, you might have a better private score (but a lower public score of couse), however, we never know the true distribution of both sets, so we can never choose a perfect one based on this approach. </p>\n\n<p>I am not sure if my understanding is correct, but curious to learn your opinion after this competition ended :)</p>",
      "rawMarkdown": "Thanks for the sharing! Great idea of combining the power of two models for unseen and seen. I just have one question - could it be your threshold is a bit less aggressive if you rely on LB to select it? As we know that private have more unseen than public, so if you had chosen an even higher threshold to allow more test samples to go through unseen models, you might have a better private score (but a lower public score of couse), however, we never know the true distribution of both sets, so we can never choose a perfect one based on this approach. \n\nI am not sure if my understanding is correct, but curious to learn your opinion after this competition ended :)",
      "votes": 3,
      "replies": [
        {
          "id": 777928,
          "postDate": "2020-03-18T02:57:27.710Z",
          "content": "<p>Yes, 0.7 performs the best in public LB and I guess there's more unseen in private so I tuned it to 0.75 before the end. I got my best private LB score by the final submission but obviously 0.75 is still less aggressive.</p>",
          "rawMarkdown": "Yes, 0.7 performs the best in public LB and I guess there's more unseen in private so I tuned it to 0.75 before the end. I got my best private LB score by the final submission but obviously 0.75 is still less aggressive."
        },
        {
          "id": 777935,
          "postDate": "2020-03-18T03:02:19.953Z",
          "content": "<p>I see! now I have a feeling that for such type of competition, the only way to get a truly robust model is to use extended dataset to cover more unseen (I am not trying to promote my external dataset approach at all :)) ways include GAN, synthestic or external unlabled. However, you did a great job as you seem to have pushed the limit of a method that does not rely on any external dataset in this situation! so congrats again! </p>",
          "rawMarkdown": "I see! now I have a feeling that for such type of competition, the only way to get a truly robust model is to use extended dataset to cover more unseen (I am not trying to promote my external dataset approach at all :)) ways include GAN, synthestic or external unlabled. However, you did a great job as you seem to have pushed the limit of a method that does not rely on any external dataset in this situation! so congrats again! ",
          "votes": 4
        },
        {
          "id": 777940,
          "postDate": "2020-03-18T03:07:39.977Z",
          "content": "<p>Thanks and congratulation to you, too!\nIt's the 3rd players got the best score without external data, not me ;)</p>",
          "rawMarkdown": "Thanks and congratulation to you, too!\nIt's the 3rd players got the best score without external data, not me ;)"
        }
      ]
    },
    {
      "id": 775871,
      "postDate": "2020-03-17T02:08:37.180Z",
      "content": "<p>Congratulations, Ha! Thank you again for all your wonderful posts. I have a question. How do you know if the new image is seen or unseen? through the predictions of which model? You have to make predictions to \"guess\" the roots/diacritics of a sample, and based on my experience it seems that model tends to predict some actually unseen grapheme samples as seen graphemes (biased towards predicting what it has seen often). Ingenious solution</p>",
      "rawMarkdown": "Congratulations, Ha! Thank you again for all your wonderful posts. I have a question. How do you know if the new image is seen or unseen? through the predictions of which model? You have to make predictions to \"guess\" the roots/diacritics of a sample, and based on my experience it seems that model tends to predict some actually unseen grapheme samples as seen graphemes (biased towards predicting what it has seen often). Ingenious solution",
      "votes": 4,
      "replies": [
        {
          "id": 775961,
          "postDate": "2020-03-17T03:42:10.523Z",
          "content": "<p>I use Arcface to find outliers when doing prediction.</p>",
          "rawMarkdown": "I use Arcface to find outliers when doing prediction.",
          "votes": 1
        }
      ]
    },
    {
      "id": 780818,
      "postDate": "2020-03-20T16:26:48.230Z",
      "content": "<p>Congrats man </p>",
      "rawMarkdown": "Congrats man \n",
      "votes": 1
    },
    {
      "id": 780372,
      "postDate": "2020-03-20T07:42:40.080Z",
      "content": "<p>Hello, <a href=\"/haqishen\">@haqishen</a> congrats again and thanks for such a great summary of your pipeline. I am not familiar with Arcface until this competition where I am seeing many top solutions have used it and did a quick googling it seems really promising for such a problem (where we have to predict properly some unseen samples). </p>\n\n<p>If I understand your approach properly, you have two models (one is for seen samples and another is for unseen samples). In the prediction phase, if the test samples are familiar, the seen model would predict otherwise unseen model, is it correct? If so, you said, the seen model has 5 outputs (3 graphene components, 1 grapheme, and Arcface) and unseen model has 4 outputs  (3 grapheme components, and Arcface) - in that case, instead of predicted 3 components directly, why decomposed the predicted grapheme by the seen model? And why use Arcface output for the seen model why not only to the unseen model as it suppose to handle the unseen samples? And how did you get the score of the grapheme components by decomposing the grapheme? And one side question, is there any superiority to use a square image than non-square shape?</p>",
      "rawMarkdown": "Hello, @haqishen congrats again and thanks for such a great summary of your pipeline. I am not familiar with Arcface until this competition where I am seeing many top solutions have used it and did a quick googling it seems really promising for such a problem (where we have to predict properly some unseen samples). \n\nIf I understand your approach properly, you have two models (one is for seen samples and another is for unseen samples). In the prediction phase, if the test samples are familiar, the seen model would predict otherwise unseen model, is it correct? If so, you said, the seen model has 5 outputs (3 graphene components, 1 grapheme, and Arcface) and unseen model has 4 outputs  (3 grapheme components, and Arcface) - in that case, instead of predicted 3 components directly, why decomposed the predicted grapheme by the seen model? And why use Arcface output for the seen model why not only to the unseen model as it suppose to handle the unseen samples? And how did you get the score of the grapheme components by decomposing the grapheme? And one side question, is there any superiority to use a square image than non-square shape?\n\n\n",
      "votes": 1,
      "replies": [
        {
          "id": 780610,
          "postDate": "2020-03-20T12:46:59.593Z",
          "content": "<p>If all test data is seen graphemes, decompose from predicted grapheme will give better score, and decompose from arcface prediction is a little bit better than decompose from softmax.</p>\n\n<p>You can decompose component from graphemes and then calculate the score on it...</p>\n\n<p><code>square image</code> is giving higher local score in my case. But not necessarily better for all CNN.</p>",
          "rawMarkdown": "If all test data is seen graphemes, decompose from predicted grapheme will give better score, and decompose from arcface prediction is a little bit better than decompose from softmax.\n\nYou can decompose component from graphemes and then calculate the score on it...\n\n`square image` is giving higher local score in my case. But not necessarily better for all CNN.",
          "votes": 1
        },
        {
          "id": 780751,
          "postDate": "2020-03-20T15:04:45.397Z",
          "content": "<p>Thanks. One more, why it is not good practice to split the data set randomly (8:2) portion, though it depends on the data set. </p>",
          "rawMarkdown": "Thanks. One more, why it is not good practice to split the data set randomly (8:2) portion, though it depends on the data set. ",
          "votes": 1
        }
      ]
    },
    {
      "id": 779353,
      "postDate": "2020-03-19T08:52:33.883Z",
      "content": "<p>Congrats! and thanks for so many sharing</p>",
      "rawMarkdown": "Congrats! and thanks for so many sharing",
      "votes": 1
    },
    {
      "id": 778349,
      "postDate": "2020-03-18T11:10:18.360Z",
      "content": "<p>Congrax <a href=\"/haqishen\">@haqishen</a>, good approach. I observed you are training a model 5 hrs long. I'm facing <strong>kernel idle</strong> issues. Is there(other than commit) a way to train a model for a long duration with no interruption.</p>",
      "rawMarkdown": "Congrax @haqishen, good approach. I observed you are training a model 5 hrs long. I'm facing **kernel idle** issues. Is there(other than commit) a way to train a model for a long duration with no interruption.",
      "votes": 1,
      "replies": [
        {
          "id": 780604,
          "postDate": "2020-03-20T12:41:34.777Z",
          "content": "<p>I'm sorry but to be honest, without a local GPU it's almost impossible to get a gold medal in CV competition nowadays....</p>",
          "rawMarkdown": "I'm sorry but to be honest, without a local GPU it's almost impossible to get a gold medal in CV competition nowadays....",
          "votes": 2
        }
      ]
    },
    {
      "id": 778262,
      "postDate": "2020-03-18T09:32:02.540Z",
      "content": "<p>Sounds great!</p>",
      "rawMarkdown": "Sounds great!",
      "votes": 1
    },
    {
      "id": 778211,
      "postDate": "2020-03-18T08:29:52.870Z",
      "content": "<p>Hi Qishen. Congratulations on the great work!</p>\n\n<p>I'm wondering where all this \"seen\"/\"unseen\" talk is coming from (I'm relatively new to deep learning). After some googling I realised it might have to do with zero-shot learning, which is also a new concept to me.</p>\n\n<p>Do you (or anyone else who sees this) have any recommendations for a paper or resource to read to get the fundamentals of your underlying strategy here?</p>",
      "rawMarkdown": "Hi Qishen. Congratulations on the great work!\n\nI'm wondering where all this \"seen\"/\"unseen\" talk is coming from (I'm relatively new to deep learning). After some googling I realised it might have to do with zero-shot learning, which is also a new concept to me.\n\nDo you (or anyone else who sees this) have any recommendations for a paper or resource to read to get the fundamentals of your underlying strategy here?",
      "votes": 1,
      "replies": [
        {
          "id": 780602,
          "postDate": "2020-03-20T12:40:18.673Z",
          "content": "<p>Well, maybe you can refer to Arcface paper or some blogs. I highly recommend it.</p>",
          "rawMarkdown": "Well, maybe you can refer to Arcface paper or some blogs. I highly recommend it."
        }
      ]
    },
    {
      "id": 778182,
      "postDate": "2020-03-18T07:58:33.983Z",
      "content": "<p>You have used apex.amp. How did it affect the training?</p>",
      "rawMarkdown": "You have used apex.amp. How did it affect the training?",
      "votes": 1,
      "replies": [
        {
          "id": 780601,
          "postDate": "2020-03-20T12:38:30.540Z",
          "content": "<p>I can set 2x batchsize by using apex. It's better for batch norm.</p>",
          "rawMarkdown": "I can set 2x batchsize by using apex. It's better for batch norm.",
          "votes": 1
        }
      ]
    },
    {
      "id": 777845,
      "postDate": "2020-03-18T01:48:02.707Z",
      "content": "<p>Congrats and thank you for sharing codes too.</p>",
      "rawMarkdown": "Congrats and thank you for sharing codes too.",
      "votes": 1
    },
    {
      "id": 777692,
      "postDate": "2020-03-17T21:22:40.660Z",
      "content": "<p>Congrats! </p>",
      "rawMarkdown": "Congrats! ",
      "votes": 1
    },
    {
      "id": 777638,
      "postDate": "2020-03-17T20:38:09.310Z",
      "content": "<p>congratulations :D</p>",
      "rawMarkdown": "congratulations :D\n",
      "votes": 1
    },
    {
      "id": 777629,
      "postDate": "2020-03-17T20:32:15.670Z",
      "content": "<p>Congrats ! Ty for sharing.</p>",
      "rawMarkdown": "Congrats ! Ty for sharing.",
      "votes": 1
    },
    {
      "id": 777563,
      "postDate": "2020-03-17T19:17:01.700Z",
      "content": "<p>Congratulations and thank you for sharing your solution.</p>",
      "rawMarkdown": "Congratulations and thank you for sharing your solution.",
      "votes": 1
    },
    {
      "id": 777520,
      "postDate": "2020-03-17T18:35:53.203Z",
      "content": "<p>Thanks a lot. I've learned a lot from you. Great work and a very generous winner!</p>",
      "rawMarkdown": "Thanks a lot. I've learned a lot from you. Great work and a very generous winner!",
      "votes": 1
    },
    {
      "id": 776817,
      "postDate": "2020-03-17T16:45:03.637Z",
      "content": "<p>Qishen congratulations, we were all cheering for you. I have some questions:</p>\n\n<p>Question 1:</p>\n\n<p>Hi, i dont understand, in this line in the \"Predict with Seen &amp; Unseen Models\"</p>\n\n<pre><code># fill predictions with Seen Model prediction as first\n# I decode 3 components from predicted grapheme here\npred = df_label_map.iloc[logits_metric.detach().cpu().numpy().argmax(1), :3].values\n</code></pre>\n\n<p>Why are you using <code>logits_metric</code> for the prediction? As far as I understand, this contains arcface information. As far as I understand, <code>logits_1</code> contains the 1295 prediction for grapheme. So why are you using the arcface predictions instead?</p>\n\n<p>Thank you.</p>\n\n<p>Question 2:</p>\n\n<p>Why did you and other top competitors use arcface as opposed to using binary model to predict seen vs unseen? Is it because arcface can perform better because the loss is used with 1295 graphemes, while binary classification would be harder to have \"seen vs unseen\"? Thank you. Also, if you can give some tips on picking the threshold for arcface, I would appreciate it. Maybe just test it using a holdout validation set if you don't have a LB in real life.</p>\n\n<p>Question 3:</p>\n\n<p>The main difference between the Seen and Unseen architecture (besides not including arcface): The Unseen model does not use an extra Linear(4096) before feeding into the MultiDropout, and then it just does Linear(186) instead of doing Linear(512) -&gt; Linear(186). So it makes the unseen model is simpler to help generalize. My question is, why do you still fit to Linear(1295) in the Unseen model, if you are trying to predict unseen graphemes? In fact, you even used it in the loss function. But we know that if they are unseen graphemes then your grapheme cannot appear. Thank you.</p>",
      "rawMarkdown": "Qishen congratulations, we were all cheering for you. I have some questions:\n\nQuestion 1:\n\nHi, i dont understand, in this line in the \"Predict with Seen &amp; Unseen Models\"\n\n    # fill predictions with Seen Model prediction as first\n    # I decode 3 components from predicted grapheme here\n    pred = df_label_map.iloc[logits_metric.detach().cpu().numpy().argmax(1), :3].values\n\nWhy are you using `logits_metric` for the prediction? As far as I understand, this contains arcface information. As far as I understand, `logits_1` contains the 1295 prediction for grapheme. So why are you using the arcface predictions instead?\n\nThank you.\n\nQuestion 2:\n\nWhy did you and other top competitors use arcface as opposed to using binary model to predict seen vs unseen? Is it because arcface can perform better because the loss is used with 1295 graphemes, while binary classification would be harder to have \"seen vs unseen\"? Thank you. Also, if you can give some tips on picking the threshold for arcface, I would appreciate it. Maybe just test it using a holdout validation set if you don't have a LB in real life.\n\nQuestion 3:\n\nThe main difference between the Seen and Unseen architecture (besides not including arcface): The Unseen model does not use an extra Linear(4096) before feeding into the MultiDropout, and then it just does Linear(186) instead of doing Linear(512) -&gt; Linear(186). So it makes the unseen model is simpler to help generalize. My question is, why do you still fit to Linear(1295) in the Unseen model, if you are trying to predict unseen graphemes? In fact, you even used it in the loss function. But we know that if they are unseen graphemes then your grapheme cannot appear. Thank you.",
      "votes": 1,
      "replies": [
        {
          "id": 776822,
          "postDate": "2020-03-17T16:58:22.210Z",
          "content": "<p>Thanks for your question and here's my answer.</p>\n\n<blockquote>\n  <p>Why are you using logits_metric for the prediction?</p>\n</blockquote>\n\n<p>It performs a little bit better than the softmax head in my validation set, like 0.00002 better.</p>\n\n<blockquote>\n  <p>Is it because arcface can perform better because the loss is used with 1295 graphemes, while binary classification would be harder to have \"seen vs unseen\"?</p>\n</blockquote>\n\n<p>Yes you're right.\nAs for the threshold selection, it's based on validation score and LB score.</p>\n\n<blockquote>\n  <p>why do you still fit to Linear(1295) in the Unseen model, if you are trying to predict unseen graphemes?</p>\n</blockquote>\n\n<p>The  Linear(1295) is not parallel with Linear(168) in unseen model, but is after it. It helps the model to converge faster.</p>",
          "rawMarkdown": "Thanks for your question and here's my answer.\n\n&gt; Why are you using logits_metric for the prediction?\n\nIt performs a little bit better than the softmax head in my validation set, like 0.00002 better.\n\n&gt; Is it because arcface can perform better because the loss is used with 1295 graphemes, while binary classification would be harder to have \"seen vs unseen\"?\n\nYes you're right.\nAs for the threshold selection, it's based on validation score and LB score.\n\n&gt; why do you still fit to Linear(1295) in the Unseen model, if you are trying to predict unseen graphemes?\n\nThe  Linear(1295) is not parallel with Linear(168) in unseen model, but is after it. It helps the model to converge faster.",
          "votes": 1
        },
        {
          "id": 784251,
          "postDate": "2020-03-24T03:27:56.423Z",
          "content": "<p><code>As for the threshold selection, it's based on validation score and LB score.</code>\nThat is to say: you tried to comment submissions many times and watched the scores, and then according to the score changes, change the seen threshold. right?</p>",
          "rawMarkdown": "`As for the threshold selection, it's based on validation score and LB score.`\nThat is to say: you tried to comment submissions many times and watched the scores, and then according to the score changes, change the seen threshold. right?"
        }
      ]
    },
    {
      "id": 776524,
      "postDate": "2020-03-17T13:11:06.663Z",
      "content": "<p>Hi qishen, May I ask how do you choose the seen / unseen threshold when test image(in your code, seen_th=0.825 )</p>",
      "rawMarkdown": "Hi qishen, May I ask how do you choose the seen / unseen threshold when test image(in your code, seen_th=0.825 )",
      "votes": 1,
      "replies": [
        {
          "id": 776555,
          "postDate": "2020-03-17T13:23:06.100Z",
          "content": "<p>I choose seen_th by validation and LB, for the public kernel, I tried 0.7, 0.75, 0.8, 0.85, 0.825. You can see it in the older version.</p>",
          "rawMarkdown": "I choose seen_th by validation and LB, for the public kernel, I tried 0.7, 0.75, 0.8, 0.85, 0.825. You can see it in the older version.",
          "votes": 1
        },
        {
          "id": 776705,
          "postDate": "2020-03-17T15:24:39.390Z",
          "content": "<p>Ok, Thanks a lot</p>",
          "rawMarkdown": "Ok, Thanks a lot",
          "votes": 1
        }
      ]
    },
    {
      "id": 776445,
      "postDate": "2020-03-17T11:57:35.527Z",
      "content": "<p>congrts!!It will help us to learn</p>",
      "rawMarkdown": "congrts!!It will help us to learn",
      "votes": 1
    },
    {
      "id": 776392,
      "postDate": "2020-03-17T11:05:57.640Z",
      "content": "<p>Thank you very much for sharing and congratulations on the impressive placing! Your discussion insights have also been valuable and insightful!</p>",
      "rawMarkdown": "Thank you very much for sharing and congratulations on the impressive placing! Your discussion insights have also been valuable and insightful!",
      "votes": 1
    },
    {
      "id": 776279,
      "postDate": "2020-03-17T09:06:19.093Z",
      "content": "<p>Congratulation, learned a lot form your sharing in this competition, thank you!</p>",
      "rawMarkdown": "Congratulation, learned a lot form your sharing in this competition, thank you!",
      "votes": 1
    },
    {
      "id": 776270,
      "postDate": "2020-03-17T08:54:46.810Z",
      "content": "<p>Congrats and thanks for your posts! Learned a lot from your kernels.</p>",
      "rawMarkdown": "Congrats and thanks for your posts! Learned a lot from your kernels.",
      "votes": 1
    },
    {
      "id": 776251,
      "postDate": "2020-03-17T08:34:37.697Z",
      "content": "<p>Well done! Congrats! Thanks for all your contributions 🤓 </p>",
      "rawMarkdown": "Well done! Congrats! Thanks for all your contributions 🤓 ",
      "votes": 1
    },
    {
      "id": 776210,
      "postDate": "2020-03-17T07:51:20.667Z",
      "content": "<p>Congrats! This is a great idea. Learned a lot of thing from your posts.</p>",
      "rawMarkdown": "Congrats! This is a great idea. Learned a lot of thing from your posts.",
      "votes": 1
    },
    {
      "id": 776198,
      "postDate": "2020-03-17T07:34:20.757Z",
      "content": "<p>Congrats. Learned a lot from your solution 😀 </p>",
      "rawMarkdown": "Congrats. Learned a lot from your solution 😀 ",
      "votes": 1
    },
    {
      "id": 776165,
      "postDate": "2020-03-17T06:58:16.213Z",
      "content": "<p>Congrats and thanks for sharing your solution <a href=\"/haqishen\">@haqishen</a> !</p>",
      "rawMarkdown": "Congrats and thanks for sharing your solution @haqishen !",
      "votes": 1
    },
    {
      "id": 776124,
      "postDate": "2020-03-17T06:14:00.797Z",
      "content": "<p>Congrats.\nWhich augmentation you used for your top scoring kernel?\nWill you share some insights about augmentation you have found during the competition?\nIt was great to learn so many things from your findings during the competition.Thanks for sharing.</p>",
      "rawMarkdown": "Congrats.\nWhich augmentation you used for your top scoring kernel?\nWill you share some insights about augmentation you have found during the competition?\nIt was great to learn so many things from your findings during the competition.Thanks for sharing.",
      "votes": 1,
      "replies": [
        {
          "id": 776146,
          "postDate": "2020-03-17T06:37:16Z",
          "content": "<p>Thanks.\nAugmentation is not the key for this competition. I don't suggest you pay too much attention on it.\nI use cutout, gridmask and cutmix for Seen Model, and bunch of other usual augmentations for Unseen Model.\nPlease refer to my training kernel for that.</p>",
          "rawMarkdown": "Thanks.\nAugmentation is not the key for this competition. I don't suggest you pay too much attention on it.\nI use cutout, gridmask and cutmix for Seen Model, and bunch of other usual augmentations for Unseen Model.\nPlease refer to my training kernel for that.",
          "votes": 2
        }
      ]
    },
    {
      "id": 776122,
      "postDate": "2020-03-17T06:13:26.287Z",
      "content": "<p>Congrats! Learned a lot of things today!</p>",
      "rawMarkdown": "Congrats! Learned a lot of things today!",
      "votes": 1
    },
    {
      "id": 776052,
      "postDate": "2020-03-17T05:13:26.220Z",
      "content": "<p>so great !</p>",
      "rawMarkdown": "so great !",
      "votes": 1
    },
    {
      "id": 775987,
      "postDate": "2020-03-17T04:11:33.220Z",
      "content": "<p>Thanks a lot <a href=\"/haqishen\">@haqishen</a>. This was my first competition and i got to learn a lot of things from your kernels. Thanks again :)</p>",
      "rawMarkdown": "Thanks a lot @haqishen. This was my first competition and i got to learn a lot of things from your kernels. Thanks again :)",
      "votes": 1
    },
    {
      "id": 775983,
      "postDate": "2020-03-17T04:06:56.267Z",
      "content": "<p>Congratulations， and I want to know what is the difference between seen and unseen models?</p>",
      "rawMarkdown": "Congratulations， and I want to know what is the difference between seen and unseen models?",
      "votes": 1,
      "replies": [
        {
          "id": 775993,
          "postDate": "2020-03-17T04:14:54.397Z",
          "content": "<p>I updated it in the <code>Summary</code> part, please have a look at it ;)</p>",
          "rawMarkdown": "I updated it in the `Summary` part, please have a look at it ;)",
          "votes": 1
        }
      ]
    },
    {
      "id": 775960,
      "postDate": "2020-03-17T03:41:13.483Z",
      "content": "<p>Thank you guys for the kind words!\nI've update the <code>Summary</code> part.</p>",
      "rawMarkdown": "Thank you guys for the kind words!\nI've update the `Summary` part.",
      "votes": 1
    },
    {
      "id": 775955,
      "postDate": "2020-03-17T03:38:54.830Z",
      "content": "<p><a href=\"/haqishen\">@haqishen</a> congratulation.  Thank you so much for your wonderful contribution. I have learned a lot from your notebooks. I didn't know about grid-mask augmentation and by <a href=\"https://www.kaggle.com/haqishen/gridmask\">your notebook</a>, it was most of the time in my training pipeline.  Thank you, you are awesome. 😊 </p>",
      "rawMarkdown": "@haqishen congratulation.  Thank you so much for your wonderful contribution. I have learned a lot from your notebooks. I didn't know about grid-mask augmentation and by [your notebook](https://www.kaggle.com/haqishen/gridmask), it was most of the time in my training pipeline.  Thank you, you are awesome. 😊 ",
      "votes": 1
    },
    {
      "id": 775918,
      "postDate": "2020-03-17T02:51:58.660Z",
      "content": "<p>Congrats! Learned so much from your posts!</p>",
      "rawMarkdown": "Congrats! Learned so much from your posts!",
      "votes": 1
    },
    {
      "id": 775913,
      "postDate": "2020-03-17T02:48:00.650Z",
      "content": "<p>Thanks Qishen, your selfless posts pushed me and let me learnt a lot during the whole period of this competition! Congrats to your gold medal and you'll definitely win more in the future.</p>",
      "rawMarkdown": "Thanks Qishen, your selfless posts pushed me and let me learnt a lot during the whole period of this competition! Congrats to your gold medal and you'll definitely win more in the future.",
      "votes": 1
    },
    {
      "id": 775899,
      "postDate": "2020-03-17T02:26:36.243Z",
      "content": "<p>Congratulations!👍 </p>",
      "rawMarkdown": "Congratulations!👍 ",
      "votes": 1
    },
    {
      "id": 775898,
      "postDate": "2020-03-17T02:26:19.737Z",
      "content": "<p>Congratulations. Thank you for all your contributions in this competition. With your sharing, I learned gridmask and augmix.  Although they did not help me at this competition,  they helped me improved 2 models at my current job.</p>",
      "rawMarkdown": "Congratulations. Thank you for all your contributions in this competition. With your sharing, I learned gridmask and augmix.  Although they did not help me at this competition,  they helped me improved 2 models at my current job.\n",
      "votes": 1
    },
    {
      "id": 775895,
      "postDate": "2020-03-17T02:23:43.780Z",
      "content": "<p>I've learned a lot from your wonderful posts. Congratulations and thank you!</p>",
      "rawMarkdown": "I've learned a lot from your wonderful posts. Congratulations and thank you!",
      "votes": 1
    },
    {
      "id": 775892,
      "postDate": "2020-03-17T02:22:05.480Z",
      "content": "<p>Congratulations, your posts and kernels were really helpful, thank you.</p>",
      "rawMarkdown": "Congratulations, your posts and kernels were really helpful, thank you.",
      "votes": 1
    },
    {
      "id": 775888,
      "postDate": "2020-03-17T02:20:08.737Z",
      "content": "<p>Congratulations! And Thanks for you great kernel in this competition</p>",
      "rawMarkdown": "Congratulations! And Thanks for you great kernel in this competition",
      "votes": 1
    },
    {
      "id": 775874,
      "postDate": "2020-03-17T02:12:13.877Z",
      "content": "<p>Congratulations, and thank you for your sharing during this competition, I learned a lot from you.</p>",
      "rawMarkdown": "Congratulations, and thank you for your sharing during this competition, I learned a lot from you.",
      "votes": 1
    },
    {
      "id": 775872,
      "postDate": "2020-03-17T02:08:59.397Z",
      "content": "<p>Congratulations <a href=\"/haqishen\">@haqishen</a> ! Your contribution in this competition was very helpful. Thank you 😊 </p>",
      "rawMarkdown": "Congratulations @haqishen ! Your contribution in this competition was very helpful. Thank you 😊 ",
      "votes": 1
    },
    {
      "id": 775867,
      "postDate": "2020-03-17T02:05:19.983Z",
      "content": "<p>666</p>",
      "rawMarkdown": "666",
      "votes": 1
    },
    {
      "id": 775862,
      "postDate": "2020-03-17T01:58:36.383Z",
      "content": "<p>wow，Congratulations! </p>",
      "rawMarkdown": "wow，Congratulations! ",
      "votes": 1
    },
    {
      "id": 775861,
      "postDate": "2020-03-17T01:56:56.787Z",
      "content": "<p>Congrats! Nice solution!</p>",
      "rawMarkdown": "Congrats! Nice solution!",
      "votes": 1
    },
    {
      "id": 775858,
      "postDate": "2020-03-17T01:56:09.717Z",
      "content": "<p>congratulations..😃 </p>",
      "rawMarkdown": "congratulations..😃 ",
      "votes": 1
    },
    {
      "id": 777552,
      "postDate": "2020-03-17T19:07:54.567Z",
      "content": "<p>Congrats <a href=\"/haqishen\">@haqishen</a> and thank you so much for sharing. This is a good one to reproduce and play around with post competition.</p>",
      "rawMarkdown": "Congrats @haqishen and thank you so much for sharing. This is a good one to reproduce and play around with post competition.",
      "votes": 2
    },
    {
      "id": 776484,
      "postDate": "2020-03-17T12:35:21.597Z",
      "content": "<p>Wow, congrats on this very elegant solution.  I was wondering if representation learning (eg arcface) could be useful.</p>\n\n<p>Congrats on the solo gold, this is a great achievement.  I hope you're not too disappointed after leading the competition for so long.  </p>",
      "rawMarkdown": "Wow, congrats on this very elegant solution.  I was wondering if representation learning (eg arcface) could be useful.\n\nCongrats on the solo gold, this is a great achievement.  I hope you're not too disappointed after leading the competition for so long.  ",
      "votes": 2,
      "replies": [
        {
          "id": 776559,
          "postDate": "2020-03-17T13:24:49.667Z",
          "content": "<p>Thanks! \nIf the top players are showing more elegant solution than me, I won't be disappointed. And actually I'm not ;)</p>",
          "rawMarkdown": "Thanks! \nIf the top players are showing more elegant solution than me, I won't be disappointed. And actually I'm not ;)",
          "votes": 1
        }
      ]
    },
    {
      "id": 776287,
      "postDate": "2020-03-17T09:11:46.577Z",
      "content": "<p>Congrats!\nThanks for sharing your great work!\nI leaned a lot of things from you in this competition!</p>",
      "rawMarkdown": "Congrats!\nThanks for sharing your great work!\nI leaned a lot of things from you in this competition!",
      "votes": 2
    },
    {
      "id": 776286,
      "postDate": "2020-03-17T09:11:43.733Z",
      "content": "<p>Impressive, I find it amazing how all top solutions approached the problem of generalizing to unseen graphemes a bit differently. </p>",
      "rawMarkdown": "Impressive, I find it amazing how all top solutions approached the problem of generalizing to unseen graphemes a bit differently. ",
      "votes": 2
    },
    {
      "id": 776004,
      "postDate": "2020-03-17T04:22:38.387Z",
      "content": "<p>Thanks Qishen, your pipeline for seen/unseen data is really interesting. (I planned to join this competition 1 month ago, but I couldn't solve the unseen data problem, so I gave up)</p>",
      "rawMarkdown": "Thanks Qishen, your pipeline for seen/unseen data is really interesting. (I planned to join this competition 1 month ago, but I couldn't solve the unseen data problem, so I gave up)",
      "votes": 2
    },
    {
      "id": 786720,
      "postDate": "2020-03-26T06:27:30.653Z",
      "content": "<p><a href=\"/haqishen\">@haqishen</a> congrats on top-10 position! And thanks for your support and insights throughout the competition. </p>\n\n<p>I just have one question please, why is cutout size set to 0.85*img_size?</p>",
      "rawMarkdown": "@haqishen congrats on top-10 position! And thanks for your support and insights throughout the competition. \n\nI just have one question please, why is cutout size set to 0.85*img_size?"
    },
    {
      "id": 776363,
      "postDate": "2020-03-17T10:34:07.327Z",
      "rawMarkdown": "",
      "votes": 4,
      "isDeleted": true
    },
    {
      "id": 776280,
      "postDate": "2020-03-17T09:06:41.140Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 776470,
      "postDate": "2020-03-17T12:23:49.880Z",
      "content": "<p>Congrats!\nThanks for sharing</p>",
      "rawMarkdown": "Congrats!\nThanks for sharing",
      "votes": 1
    }
  ],
  "comments": [
    {
      "id": 777919,
      "author_name": "Weimin Wang",
      "author_url": "",
      "post_date": "2020-03-18T02:50:42.860000",
      "content": "<p>Thanks for the sharing! Great idea of combining the power of two models for unseen and seen. I just have one question - could it be your threshold is a bit less aggressive if you rely on LB to select it? As we know that private have more unseen than public, so if you had chosen an even higher threshold to allow more test samples to go through unseen models, you might have a better private score (but a lower public score of couse), however, we never know the true distribution of both sets, so we can never choose a perfect one based on this approach. </p>\n\n<p>I am not sure if my understanding is correct, but curious to learn your opinion after this competition ended :)</p>",
      "votes": 3,
      "replies": [
        {
          "id": 777928,
          "author_name": "Qishen Ha",
          "author_url": "",
          "post_date": "2020-03-18T02:57:27.710000",
          "content": "<p>Yes, 0.7 performs the best in public LB and I guess there's more unseen in private so I tuned it to 0.75 before the end. I got my best private LB score by the final submission but obviously 0.75 is still less aggressive.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 777935,
          "author_name": "Weimin Wang",
          "author_url": "",
          "post_date": "2020-03-18T03:02:19.953000",
          "content": "<p>I see! now I have a feeling that for such type of competition, the only way to get a truly robust model is to use extended dataset to cover more unseen (I am not trying to promote my external dataset approach at all :)) ways include GAN, synthestic or external unlabled. However, you did a great job as you seem to have pushed the limit of a method that does not rely on any external dataset in this situation! so congrats again! </p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 777940,
          "author_name": "Qishen Ha",
          "author_url": "",
          "post_date": "2020-03-18T03:07:39.977000",
          "content": "<p>Thanks and congratulation to you, too!\nIt's the 3rd players got the best score without external data, not me ;)</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 775871,
      "author_name": "Nicholas Lyu",
      "author_url": "",
      "post_date": "2020-03-17T02:08:37.180000",
      "content": "<p>Congratulations, Ha! Thank you again for all your wonderful posts. I have a question. How do you know if the new image is seen or unseen? through the predictions of which model? You have to make predictions to \"guess\" the roots/diacritics of a sample, and based on my experience it seems that model tends to predict some actually unseen grapheme samples as seen graphemes (biased towards predicting what it has seen often). Ingenious solution</p>",
      "votes": 4,
      "replies": [
        {
          "id": 775961,
          "author_name": "Qishen Ha",
          "author_url": "",
          "post_date": "2020-03-17T03:42:10.523000",
          "content": "<p>I use Arcface to find outliers when doing prediction.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 780818,
      "author_name": "taber bin zameer",
      "author_url": "",
      "post_date": "2020-03-20T16:26:48.230000",
      "content": "<p>Congrats man </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 780372,
      "author_name": "Innat",
      "author_url": "",
      "post_date": "2020-03-20T07:42:40.080000",
      "content": "<p>Hello, <a href=\"/haqishen\">@haqishen</a> congrats again and thanks for such a great summary of your pipeline. I am not familiar with Arcface until this competition where I am seeing many top solutions have used it and did a quick googling it seems really promising for such a problem (where we have to predict properly some unseen samples). </p>\n\n<p>If I understand your approach properly, you have two models (one is for seen samples and another is for unseen samples). In the prediction phase, if the test samples are familiar, the seen model would predict otherwise unseen model, is it correct? If so, you said, the seen model has 5 outputs (3 graphene components, 1 grapheme, and Arcface) and unseen model has 4 outputs  (3 grapheme components, and Arcface) - in that case, instead of predicted 3 components directly, why decomposed the predicted grapheme by the seen model? And why use Arcface output for the seen model why not only to the unseen model as it suppose to handle the unseen samples? And how did you get the score of the grapheme components by decomposing the grapheme? And one side question, is there any superiority to use a square image than non-square shape?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 780610,
          "author_name": "Qishen Ha",
          "author_url": "",
          "post_date": "2020-03-20T12:46:59.593000",
          "content": "<p>If all test data is seen graphemes, decompose from predicted grapheme will give better score, and decompose from arcface prediction is a little bit better than decompose from softmax.</p>\n\n<p>You can decompose component from graphemes and then calculate the score on it...</p>\n\n<p><code>square image</code> is giving higher local score in my case. But not necessarily better for all CNN.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 780751,
          "author_name": "Innat",
          "author_url": "",
          "post_date": "2020-03-20T15:04:45.397000",
          "content": "<p>Thanks. One more, why it is not good practice to split the data set randomly (8:2) portion, though it depends on the data set. </p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 779353,
      "author_name": "liuzhangzhen",
      "author_url": "",
      "post_date": "2020-03-19T08:52:33.883000",
      "content": "<p>Congrats! and thanks for so many sharing</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 778349,
      "author_name": "Mettu Venkataramireddy",
      "author_url": "",
      "post_date": "2020-03-18T11:10:18.360000",
      "content": "<p>Congrax <a href=\"/haqishen\">@haqishen</a>, good approach. I observed you are training a model 5 hrs long. I'm facing <strong>kernel idle</strong> issues. Is there(other than commit) a way to train a model for a long duration with no interruption.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 780604,
          "author_name": "Qishen Ha",
          "author_url": "",
          "post_date": "2020-03-20T12:41:34.777000",
          "content": "<p>I'm sorry but to be honest, without a local GPU it's almost impossible to get a gold medal in CV competition nowadays....</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 778262,
      "author_name": "Erfan Jalili",
      "author_url": "",
      "post_date": "2020-03-18T09:32:02.540000",
      "content": "<p>Sounds great!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 778211,
      "author_name": "Alexander Soare",
      "author_url": "",
      "post_date": "2020-03-18T08:29:52.870000",
      "content": "<p>Hi Qishen. Congratulations on the great work!</p>\n\n<p>I'm wondering where all this \"seen\"/\"unseen\" talk is coming from (I'm relatively new to deep learning). After some googling I realised it might have to do with zero-shot learning, which is also a new concept to me.</p>\n\n<p>Do you (or anyone else who sees this) have any recommendations for a paper or resource to read to get the fundamentals of your underlying strategy here?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 780602,
          "author_name": "Qishen Ha",
          "author_url": "",
          "post_date": "2020-03-20T12:40:18.673000",
          "content": "<p>Well, maybe you can refer to Arcface paper or some blogs. I highly recommend it.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 778182,
      "author_name": "Shayekh Islam",
      "author_url": "",
      "post_date": "2020-03-18T07:58:33.983000",
      "content": "<p>You have used apex.amp. How did it affect the training?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 780601,
          "author_name": "Qishen Ha",
          "author_url": "",
          "post_date": "2020-03-20T12:38:30.540000",
          "content": "<p>I can set 2x batchsize by using apex. It's better for batch norm.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 777845,
      "author_name": "corochann",
      "author_url": "",
      "post_date": "2020-03-18T01:48:02.707000",
      "content": "<p>Congrats and thank you for sharing codes too.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 777692,
      "author_name": "Xuan Cao",
      "author_url": "",
      "post_date": "2020-03-17T21:22:40.660000",
      "content": "<p>Congrats! </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 777638,
      "author_name": "Jeeten",
      "author_url": "",
      "post_date": "2020-03-17T20:38:09.310000",
      "content": "<p>congratulations :D</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 777629,
      "author_name": "Areyana",
      "author_url": "",
      "post_date": "2020-03-17T20:32:15.670000",
      "content": "<p>Congrats ! Ty for sharing.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 777563,
      "author_name": "Deep Chatterjee",
      "author_url": "",
      "post_date": "2020-03-17T19:17:01.700000",
      "content": "<p>Congratulations and thank you for sharing your solution.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 777520,
      "author_name": "Jingjie Zhang",
      "author_url": "",
      "post_date": "2020-03-17T18:35:53.203000",
      "content": "<p>Thanks a lot. I've learned a lot from you. Great work and a very generous winner!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 776817,
      "author_name": "CoreyJamesLevinson",
      "author_url": "",
      "post_date": "2020-03-17T16:45:03.637000",
      "content": "<p>Qishen congratulations, we were all cheering for you. I have some questions:</p>\n\n<p>Question 1:</p>\n\n<p>Hi, i dont understand, in this line in the \"Predict with Seen &amp; Unseen Models\"</p>\n\n<pre><code># fill predictions with Seen Model prediction as first\n# I decode 3 components from predicted grapheme here\npred = df_label_map.iloc[logits_metric.detach().cpu().numpy().argmax(1), :3].values\n</code></pre>\n\n<p>Why are you using <code>logits_metric</code> for the prediction? As far as I understand, this contains arcface information. As far as I understand, <code>logits_1</code> contains the 1295 prediction for grapheme. So why are you using the arcface predictions instead?</p>\n\n<p>Thank you.</p>\n\n<p>Question 2:</p>\n\n<p>Why did you and other top competitors use arcface as opposed to using binary model to predict seen vs unseen? Is it because arcface can perform better because the loss is used with 1295 graphemes, while binary classification would be harder to have \"seen vs unseen\"? Thank you. Also, if you can give some tips on picking the threshold for arcface, I would appreciate it. Maybe just test it using a holdout validation set if you don't have a LB in real life.</p>\n\n<p>Question 3:</p>\n\n<p>The main difference between the Seen and Unseen architecture (besides not including arcface): The Unseen model does not use an extra Linear(4096) before feeding into the MultiDropout, and then it just does Linear(186) instead of doing Linear(512) -&gt; Linear(186). So it makes the unseen model is simpler to help generalize. My question is, why do you still fit to Linear(1295) in the Unseen model, if you are trying to predict unseen graphemes? In fact, you even used it in the loss function. But we know that if they are unseen graphemes then your grapheme cannot appear. Thank you.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 776822,
          "author_name": "Qishen Ha",
          "author_url": "",
          "post_date": "2020-03-17T16:58:22.210000",
          "content": "<p>Thanks for your question and here's my answer.</p>\n\n<blockquote>\n  <p>Why are you using logits_metric for the prediction?</p>\n</blockquote>\n\n<p>It performs a little bit better than the softmax head in my validation set, like 0.00002 better.</p>\n\n<blockquote>\n  <p>Is it because arcface can perform better because the loss is used with 1295 graphemes, while binary classification would be harder to have \"seen vs unseen\"?</p>\n</blockquote>\n\n<p>Yes you're right.\nAs for the threshold selection, it's based on validation score and LB score.</p>\n\n<blockquote>\n  <p>why do you still fit to Linear(1295) in the Unseen model, if you are trying to predict unseen graphemes?</p>\n</blockquote>\n\n<p>The  Linear(1295) is not parallel with Linear(168) in unseen model, but is after it. It helps the model to converge faster.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 784251,
          "author_name": "Mingxing Liu",
          "author_url": "",
          "post_date": "2020-03-24T03:27:56.423000",
          "content": "<p><code>As for the threshold selection, it's based on validation score and LB score.</code>\nThat is to say: you tried to comment submissions many times and watched the scores, and then according to the score changes, change the seen threshold. right?</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 776524,
      "author_name": "He",
      "author_url": "",
      "post_date": "2020-03-17T13:11:06.663000",
      "content": "<p>Hi qishen, May I ask how do you choose the seen / unseen threshold when test image(in your code, seen_th=0.825 )</p>",
      "votes": 1,
      "replies": [
        {
          "id": 776555,
          "author_name": "Qishen Ha",
          "author_url": "",
          "post_date": "2020-03-17T13:23:06.100000",
          "content": "<p>I choose seen_th by validation and LB, for the public kernel, I tried 0.7, 0.75, 0.8, 0.85, 0.825. You can see it in the older version.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 776705,
          "author_name": "He",
          "author_url": "",
          "post_date": "2020-03-17T15:24:39.390000",
          "content": "<p>Ok, Thanks a lot</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 776445,
      "author_name": "Anurag",
      "author_url": "",
      "post_date": "2020-03-17T11:57:35.527000",
      "content": "<p>congrts!!It will help us to learn</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 776392,
      "author_name": "hirek",
      "author_url": "",
      "post_date": "2020-03-17T11:05:57.640000",
      "content": "<p>Thank you very much for sharing and congratulations on the impressive placing! Your discussion insights have also been valuable and insightful!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 776279,
      "author_name": "zr",
      "author_url": "",
      "post_date": "2020-03-17T09:06:19.093000",
      "content": "<p>Congratulation, learned a lot form your sharing in this competition, thank you!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 776270,
      "author_name": "Yu",
      "author_url": "",
      "post_date": "2020-03-17T08:54:46.810000",
      "content": "<p>Congrats and thanks for your posts! Learned a lot from your kernels.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 776251,
      "author_name": "Chulvi",
      "author_url": "",
      "post_date": "2020-03-17T08:34:37.697000",
      "content": "<p>Well done! Congrats! Thanks for all your contributions 🤓 </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 776210,
      "author_name": "Minh Nguyen",
      "author_url": "",
      "post_date": "2020-03-17T07:51:20.667000",
      "content": "<p>Congrats! This is a great idea. Learned a lot of thing from your posts.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 776198,
      "author_name": "Qingyao Shuai",
      "author_url": "",
      "post_date": "2020-03-17T07:34:20.757000",
      "content": "<p>Congrats. Learned a lot from your solution 😀 </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 776165,
      "author_name": "KhanhVD",
      "author_url": "",
      "post_date": "2020-03-17T06:58:16.213000",
      "content": "<p>Congrats and thanks for sharing your solution <a href=\"/haqishen\">@haqishen</a> !</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 776124,
      "author_name": "Md.Abrar Istiak Akib",
      "author_url": "",
      "post_date": "2020-03-17T06:14:00.797000",
      "content": "<p>Congrats.\nWhich augmentation you used for your top scoring kernel?\nWill you share some insights about augmentation you have found during the competition?\nIt was great to learn so many things from your findings during the competition.Thanks for sharing.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 776146,
          "author_name": "Qishen Ha",
          "author_url": "",
          "post_date": "2020-03-17T06:37:16",
          "content": "<p>Thanks.\nAugmentation is not the key for this competition. I don't suggest you pay too much attention on it.\nI use cutout, gridmask and cutmix for Seen Model, and bunch of other usual augmentations for Unseen Model.\nPlease refer to my training kernel for that.</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 776122,
      "author_name": "Bryce1010",
      "author_url": "",
      "post_date": "2020-03-17T06:13:26.287000",
      "content": "<p>Congrats! Learned a lot of things today!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 776052,
      "author_name": "Izmaylov Konstantin",
      "author_url": "",
      "post_date": "2020-03-17T05:13:26.220000",
      "content": "<p>so great !</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 775987,
      "author_name": "sajwankit",
      "author_url": "",
      "post_date": "2020-03-17T04:11:33.220000",
      "content": "<p>Thanks a lot <a href=\"/haqishen\">@haqishen</a>. This was my first competition and i got to learn a lot of things from your kernels. Thanks again :)</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 775983,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-03-17T04:06:56.267000",
      "content": "",
      "votes": 1,
      "replies": [
        {
          "id": 775993,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-03-17T04:14:54.397000",
          "content": "",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 775960,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-03-17T03:41:13.483000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 775955,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-03-17T03:38:54.830000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 775918,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-03-17T02:51:58.660000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 775913,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-03-17T02:48:00.650000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 775899,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-03-17T02:26:36.243000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 775898,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-03-17T02:26:19.737000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 775895,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-03-17T02:23:43.780000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 775892,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-03-17T02:22:05.480000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 775888,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-03-17T02:20:08.737000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 775874,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-03-17T02:12:13.877000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 775872,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-03-17T02:08:59.397000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 775867,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-03-17T02:05:19.983000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 775862,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-03-17T01:58:36.383000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 775861,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-03-17T01:56:56.787000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 775858,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-03-17T01:56:09.717000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 777552,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-03-17T19:07:54.567000",
      "content": "",
      "votes": 2,
      "replies": []
    },
    {
      "id": 776484,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-03-17T12:35:21.597000",
      "content": "",
      "votes": 2,
      "replies": [
        {
          "id": 776559,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-03-17T13:24:49.667000",
          "content": "",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 776287,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-03-17T09:11:46.577000",
      "content": "",
      "votes": 2,
      "replies": []
    },
    {
      "id": 776286,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-03-17T09:11:43.733000",
      "content": "",
      "votes": 2,
      "replies": []
    },
    {
      "id": 776004,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-03-17T04:22:38.387000",
      "content": "",
      "votes": 2,
      "replies": []
    },
    {
      "id": 786720,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-03-26T06:27:30.653000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 776363,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-03-17T10:34:07.327000",
      "content": "",
      "votes": 4,
      "replies": []
    },
    {
      "id": 776280,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-03-17T09:06:41.140000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 776470,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-03-17T12:23:49.880000",
      "content": "",
      "votes": 1,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "775855": "Hi, everyone\n\nFirst of all, I want to thank Kaggle and the hosts for hosting the competition.\n\nCongratulation to all winners!\n\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F448347%2Fd37b2cfc0fd3524b5be0ba10e5d8b649%2F1.png?generation=1584409865722874&amp;alt=media)\n\nAn illustration of my pipeline. It contains 2 models, Seen Model and Unseen Model.\n\n# Summary\n\n* NO EXTERNAL DATA\n\n* I used 2-stage prediction in my pipeline, it contains 2 models as one set (Seen Model and Unseen Model) as shown in the figure above.\n* I used Arcface to distinguish unseen graphemes. Just like to distinguish unseen faces ;)\n* If an image is detected to be seen, I use the output of Seen Model as prediction directly. Else if the image is detected to be unseen, I’ll pass it to Unseen Model to get unseen prediction.\n* The different points between Seen Model and Unseen Model are as follow:\n  * Seen Model only uses very few augmentation to make sure it can ‘overfit’ to seen graphemes, while Unseen Model uses heavy augmentations to make it generalize to unseen ones.\n  * Seen Model has 5-head outputs, including Arcface output. While Unseen Model has normal 4-head outputs.\n  * Beside the number of output heads, the architecture of the top is a little bit different as well.\n  * I decompose the predicted grapheme to 3 components when the prediction is made by Seen Model, while I use predicted 3 components directly when the prediction is made by Unseen Model.\n\n\n# Code\n\nI’ve published a minimal version of my training &amp; testing pipeline on Kaggle kernel as follow:\n\n* Step 1. [Bengali Train Seen Model](https://www.kaggle.com/haqishen/bengali-train-seen-model) (trained ~5h on kernel)\n* Step 2. [Bengali Train Unseen Model](https://www.kaggle.com/haqishen/bengali-train-unseen-model) (trained ~8.5h on kernel)\n* Step 3. [Bengali Predict with Seen &amp; Unseen Models](https://www.kaggle.com/haqishen/bengali-predict-with-seen-unseen-models) (submission ~20min)\n\nThe minimal version with two efficientnet-b1 which are training on 128x128 for 30 epochs can give you public LB 0.985+ or private LB 0.935+ \nIf you find my notebooks helpful, please upvote them! \n\n---\n\nTo get over 0.993+ or 0.950+ on public or private  LB, what you need to do is just simply:\n* Change the backbone to efficientnet-b5 / b6 / b7 / b8\n* Train on 224x224\n* Train for 60~90 epochs\n* Emsenble more Seen &amp; Unseen Model\n\nIf you have any questions please feel free to leave a message to me.\nThanks!\n\n",
    "777919": "Thanks for the sharing! Great idea of combining the power of two models for unseen and seen. I just have one question - could it be your threshold is a bit less aggressive if you rely on LB to select it? As we know that private have more unseen than public, so if you had chosen an even higher threshold to allow more test samples to go through unseen models, you might have a better private score (but a lower public score of couse), however, we never know the true distribution of both sets, so we can never choose a perfect one based on this approach. \n\nI am not sure if my understanding is correct, but curious to learn your opinion after this competition ended :)",
    "775871": "Congratulations, Ha! Thank you again for all your wonderful posts. I have a question. How do you know if the new image is seen or unseen? through the predictions of which model? You have to make predictions to \"guess\" the roots/diacritics of a sample, and based on my experience it seems that model tends to predict some actually unseen grapheme samples as seen graphemes (biased towards predicting what it has seen often). Ingenious solution",
    "780818": "Congrats man \n",
    "780372": "Hello, @haqishen congrats again and thanks for such a great summary of your pipeline. I am not familiar with Arcface until this competition where I am seeing many top solutions have used it and did a quick googling it seems really promising for such a problem (where we have to predict properly some unseen samples). \n\nIf I understand your approach properly, you have two models (one is for seen samples and another is for unseen samples). In the prediction phase, if the test samples are familiar, the seen model would predict otherwise unseen model, is it correct? If so, you said, the seen model has 5 outputs (3 graphene components, 1 grapheme, and Arcface) and unseen model has 4 outputs  (3 grapheme components, and Arcface) - in that case, instead of predicted 3 components directly, why decomposed the predicted grapheme by the seen model? And why use Arcface output for the seen model why not only to the unseen model as it suppose to handle the unseen samples? And how did you get the score of the grapheme components by decomposing the grapheme? And one side question, is there any superiority to use a square image than non-square shape?\n\n\n",
    "779353": "Congrats! and thanks for so many sharing",
    "778349": "Congrax @haqishen, good approach. I observed you are training a model 5 hrs long. I'm facing **kernel idle** issues. Is there(other than commit) a way to train a model for a long duration with no interruption.",
    "778262": "Sounds great!",
    "778211": "Hi Qishen. Congratulations on the great work!\n\nI'm wondering where all this \"seen\"/\"unseen\" talk is coming from (I'm relatively new to deep learning). After some googling I realised it might have to do with zero-shot learning, which is also a new concept to me.\n\nDo you (or anyone else who sees this) have any recommendations for a paper or resource to read to get the fundamentals of your underlying strategy here?",
    "778182": "You have used apex.amp. How did it affect the training?",
    "777845": "Congrats and thank you for sharing codes too.",
    "777692": "Congrats! ",
    "777638": "congratulations :D\n",
    "777629": "Congrats ! Ty for sharing.",
    "777563": "Congratulations and thank you for sharing your solution.",
    "777520": "Thanks a lot. I've learned a lot from you. Great work and a very generous winner!",
    "776817": "Qishen congratulations, we were all cheering for you. I have some questions:\n\nQuestion 1:\n\nHi, i dont understand, in this line in the \"Predict with Seen &amp; Unseen Models\"\n\n    # fill predictions with Seen Model prediction as first\n    # I decode 3 components from predicted grapheme here\n    pred = df_label_map.iloc[logits_metric.detach().cpu().numpy().argmax(1), :3].values\n\nWhy are you using `logits_metric` for the prediction? As far as I understand, this contains arcface information. As far as I understand, `logits_1` contains the 1295 prediction for grapheme. So why are you using the arcface predictions instead?\n\nThank you.\n\nQuestion 2:\n\nWhy did you and other top competitors use arcface as opposed to using binary model to predict seen vs unseen? Is it because arcface can perform better because the loss is used with 1295 graphemes, while binary classification would be harder to have \"seen vs unseen\"? Thank you. Also, if you can give some tips on picking the threshold for arcface, I would appreciate it. Maybe just test it using a holdout validation set if you don't have a LB in real life.\n\nQuestion 3:\n\nThe main difference between the Seen and Unseen architecture (besides not including arcface): The Unseen model does not use an extra Linear(4096) before feeding into the MultiDropout, and then it just does Linear(186) instead of doing Linear(512) -&gt; Linear(186). So it makes the unseen model is simpler to help generalize. My question is, why do you still fit to Linear(1295) in the Unseen model, if you are trying to predict unseen graphemes? In fact, you even used it in the loss function. But we know that if they are unseen graphemes then your grapheme cannot appear. Thank you.",
    "776524": "Hi qishen, May I ask how do you choose the seen / unseen threshold when test image(in your code, seen_th=0.825 )",
    "776445": "congrts!!It will help us to learn",
    "776392": "Thank you very much for sharing and congratulations on the impressive placing! Your discussion insights have also been valuable and insightful!",
    "776279": "Congratulation, learned a lot form your sharing in this competition, thank you!",
    "776270": "Congrats and thanks for your posts! Learned a lot from your kernels.",
    "776251": "Well done! Congrats! Thanks for all your contributions 🤓 ",
    "776210": "Congrats! This is a great idea. Learned a lot of thing from your posts.",
    "776198": "Congrats. Learned a lot from your solution 😀 ",
    "776165": "Congrats and thanks for sharing your solution @haqishen !",
    "776124": "Congrats.\nWhich augmentation you used for your top scoring kernel?\nWill you share some insights about augmentation you have found during the competition?\nIt was great to learn so many things from your findings during the competition.Thanks for sharing.",
    "776122": "Congrats! Learned a lot of things today!",
    "776052": "so great !",
    "775987": "Thanks a lot @haqishen. This was my first competition and i got to learn a lot of things from your kernels. Thanks again :)",
    "775983": "Congratulations， and I want to know what is the difference between seen and unseen models?",
    "775960": "Thank you guys for the kind words!\nI've update the `Summary` part.",
    "775955": "@haqishen congratulation.  Thank you so much for your wonderful contribution. I have learned a lot from your notebooks. I didn't know about grid-mask augmentation and by [your notebook](https://www.kaggle.com/haqishen/gridmask), it was most of the time in my training pipeline.  Thank you, you are awesome. 😊 ",
    "775918": "Congrats! Learned so much from your posts!",
    "775913": "Thanks Qishen, your selfless posts pushed me and let me learnt a lot during the whole period of this competition! Congrats to your gold medal and you'll definitely win more in the future.",
    "775899": "Congratulations!👍 ",
    "775898": "Congratulations. Thank you for all your contributions in this competition. With your sharing, I learned gridmask and augmix.  Although they did not help me at this competition,  they helped me improved 2 models at my current job.\n",
    "775895": "I've learned a lot from your wonderful posts. Congratulations and thank you!",
    "775892": "Congratulations, your posts and kernels were really helpful, thank you.",
    "775888": "Congratulations! And Thanks for you great kernel in this competition",
    "775874": "Congratulations, and thank you for your sharing during this competition, I learned a lot from you.",
    "775872": "Congratulations @haqishen ! Your contribution in this competition was very helpful. Thank you 😊 ",
    "775867": "666",
    "775862": "wow，Congratulations! ",
    "775861": "Congrats! Nice solution!",
    "775858": "congratulations..😃 ",
    "777552": "Congrats @haqishen and thank you so much for sharing. This is a good one to reproduce and play around with post competition.",
    "776484": "Wow, congrats on this very elegant solution.  I was wondering if representation learning (eg arcface) could be useful.\n\nCongrats on the solo gold, this is a great achievement.  I hope you're not too disappointed after leading the competition for so long.  ",
    "776287": "Congrats!\nThanks for sharing your great work!\nI leaned a lot of things from you in this competition!",
    "776286": "Impressive, I find it amazing how all top solutions approached the problem of generalizing to unseen graphemes a bit differently. ",
    "776004": "Thanks Qishen, your pipeline for seen/unseen data is really interesting. (I planned to join this competition 1 month ago, but I couldn't solve the unseen data problem, so I gave up)",
    "786720": "@haqishen congrats on top-10 position! And thanks for your support and insights throughout the competition. \n\nI just have one question please, why is cutout size set to 0.85*img_size?",
    "776363": "",
    "776280": "",
    "776470": "Congrats!\nThanks for sharing"
  }
}