{
  "id": 136982,
  "title": "4th place solution",
  "url": "/competitions/bengaliai-cv19/writeups/h2o-cv-4th-place-solution",
  "author_name": "",
  "post_date": "2020-03-18T17:46:53.625404900Z",
  "votes": 38,
  "comment_count": 11,
  "views": 0,
  "content": "<p>We appreciate the efforts Bengali.AI and Kaggle spent on this competition, and congrats to all the winners. We are very impressed by your interesting findings, novol methods, deep understanding of the problem, the dataset, as well as the evaluation metric.</p>\n\n<p>Our solution is relatively straightforward and bears similarities with some of the top solutions.</p>\n\n<p>Let's call the 1295 graphemes in the train set ID (in-dictionary), and the unknown graphemes in the test set OOD (out-of-dictionary). The first step was training Arcface models and computing the feature centers of each of the 1295 graphemes. We can then tell if a test image is ID or OOD based on its smallest feature distance to every grapheme centers. In our 4th place submission, we set the classification threshold to 0.15 (cosine distance), which was estimated locally. After this ID/OOD classification, we applied the Arcface models to the ID set only and determined the component classes by the grapheme classes. For the OOD set, we trained another group of 1-head and 3-head component models and applied them to the OOD set only.</p>\n\n<p><strong>More of the Arcface models:</strong>\nnetwork structures: inception_resnet_v2 and seresnext101\ninput image preprocessing: plain resizing\ninput image size: 320x320\nloss: Arcface\ntarget: 1295 grapheme classification\ntotal training epochs: 90\naugmentations: cutout only</p>\n\n<p><strong>More of the OOD models:</strong>\nnetwork structures: inceptions, resnets, densnets, efficientnets ... ... ...\ninput image preprocessing: <a href=\"https://www.kaggle.com/iafoss/image-preprocessing-128x128\">https://www.kaggle.com/iafoss/image-preprocessing-128x128</a>\ninput image size: various sizes from 256x256 to 416x416\nloss: CE\ntarget: 3-head for vowel_diacritic and consonant_diacritic, 1-head for grapheme_root\ntotal training epochs: 10 (more training harmed the OOD performance)\naugmentations: mixup only</p>",
  "messages": [
    {
      "id": "778756",
      "postDate": "03/18/2020 17:46:53",
      "content": "<p>We appreciate the efforts Bengali.AI and Kaggle spent on this competition, and congrats to all the winners. We are very impressed by your interesting findings, novol methods, deep understanding of the problem, the dataset, as well as the evaluation metric.</p>\n\n<p>Our solution is relatively straightforward and bears similarities with some of the top solutions.</p>\n\n<p>Let's call the 1295 graphemes in the train set ID (in-dictionary), and the unknown graphemes in the test set OOD (out-of-dictionary). The first step was training Arcface models and computing the feature centers of each of the 1295 graphemes. We can then tell if a test image is ID or OOD based on its smallest feature distance to every grapheme centers. In our 4th place submission, we set the classification threshold to 0.15 (cosine distance), which was estimated locally. After this ID/OOD classification, we applied the Arcface models to the ID set only and determined the component classes by the grapheme classes. For the OOD set, we trained another group of 1-head and 3-head component models and applied them to the OOD set only.</p>\n\n<p><strong>More of the Arcface models:</strong>\nnetwork structures: inception_resnet_v2 and seresnext101\ninput image preprocessing: plain resizing\ninput image size: 320x320\nloss: Arcface\ntarget: 1295 grapheme classification\ntotal training epochs: 90\naugmentations: cutout only</p>\n\n<p><strong>More of the OOD models:</strong>\nnetwork structures: inceptions, resnets, densnets, efficientnets ... ... ...\ninput image preprocessing: <a href=\"https://www.kaggle.com/iafoss/image-preprocessing-128x128\">https://www.kaggle.com/iafoss/image-preprocessing-128x128</a>\ninput image size: various sizes from 256x256 to 416x416\nloss: CE\ntarget: 3-head for vowel_diacritic and consonant_diacritic, 1-head for grapheme_root\ntotal training epochs: 10 (more training harmed the OOD performance)\naugmentations: mixup only</p>",
      "rawMarkdown": "We appreciate the efforts Bengali.AI and Kaggle spent on this competition, and congrats to all the winners. We are very impressed by your interesting findings, novol methods, deep understanding of the problem, the dataset, as well as the evaluation metric.\n\nOur solution is relatively straightforward and bears similarities with some of the top solutions.\n\nLet's call the 1295 graphemes in the train set ID (in-dictionary), and the unknown graphemes in the test set OOD (out-of-dictionary). The first step was training Arcface models and computing the feature centers of each of the 1295 graphemes. We can then tell if a test image is ID or OOD based on its smallest feature distance to every grapheme centers. In our 4th place submission, we set the classification threshold to 0.15 (cosine distance), which was estimated locally. After this ID/OOD classification, we applied the Arcface models to the ID set only and determined the component classes by the grapheme classes. For the OOD set, we trained another group of 1-head and 3-head component models and applied them to the OOD set only.\n\n**More of the Arcface models:**\nnetwork structures: inception_resnet_v2 and seresnext101\ninput image preprocessing: plain resizing\ninput image size: 320x320\nloss: Arcface\ntarget: 1295 grapheme classification\ntotal training epochs: 90\naugmentations: cutout only\n\n**More of the OOD models:**\nnetwork structures: inceptions, resnets, densnets, efficientnets ... ... ...\ninput image preprocessing: https://www.kaggle.com/iafoss/image-preprocessing-128x128\ninput image size: various sizes from 256x256 to 416x416\nloss: CE\ntarget: 3-head for vowel_diacritic and consonant_diacritic, 1-head for grapheme_root\ntotal training epochs: 10 (more training harmed the OOD performance)\naugmentations: mixup only",
      "votes": null
    },
    {
      "id": "778765",
      "postDate": "03/18/2020 17:57:08",
      "content": "<p><a href=\"/wowfattie\">@wowfattie</a> <a href=\"/ybabakhin\">@ybabakhin</a> I have been humbled and privileged to team up with both of you. \nIt has been quite a journey for me in the image world !\nI'd like to really thank you both for your kindness, patience and cheerful moments we had ! \nI've learnt so much at your side.\nGreat Lesson ;-) </p>",
      "rawMarkdown": "wowfattie @ybabakhin I have been humbled and privileged to team up with both of you. \nIt has been quite a journey for me in the image world !\nI'd like to really thank you both for your kindness, patience and cheerful moments we had ! \nI've learnt so much at your side.\nGreat Lesson ;-)",
      "votes": null
    },
    {
      "id": "779015",
      "postDate": "03/18/2020 23:55:26",
      "content": "<p>Wow, simple solution and powerful. Congrats team. Great job.</p>\n\n<p>In your OOD model, what are each of the 3 heads doing when classifying the 2 components vowel and consonant? In your ID model, how do you compute root, vowel, and consonant from your grapheme prediction?</p>",
      "rawMarkdown": "Wow, simple solution and powerful. Congrats team. Great job.\n\nIn your OOD model, what are each of the 3 heads doing when classifying the 2 components vowel and consonant? In your ID model, how do you compute root, vowel, and consonant from your grapheme prediction?",
      "votes": null
    },
    {
      "id": "779029",
      "postDate": "03/19/2020 00:39:10",
      "content": "<p>Thank you.\nThere was nothing special of our 3-head model at training stage, but during inference, we did not use the predictions of the \"grapheme_root\" head. We built dedicated classifiers (168-class classification, only 1-head) for grapheme_root predictions. \nOur ID (Arcface) models predicted grapheme classes (1295-class), and we could directly lookup the three components from train.csv.</p>",
      "rawMarkdown": "Thank you.\nThere was nothing special of our 3-head model at training stage, but during inference, we did not use the predictions of the \"grapheme_root\" head. We built dedicated classifiers (168-class classification, only 1-head) for grapheme_root predictions. \nOur ID (Arcface) models predicted grapheme classes (1295-class), and we could directly lookup the three components from train.csv.",
      "votes": null
    },
    {
      "id": "779152",
      "postDate": "03/19/2020 04:13:42",
      "content": "<p>Congrats! Thanks for sharing. May I ask a few questions?\n1.  How do you train your OOD? I mean, how do you divide your training set for ID and OOD? \n2. About the 3-head of OOD, are they voweldiacritic, consonantdiacritic and grapheme?</p>",
      "rawMarkdown": "Congrats! Thanks for sharing. May I ask a few questions?\n1.  How do you train your OOD? I mean, how do you divide your training set for ID and OOD? \n2. About the 3-head of OOD, are they voweldiacritic, consonantdiacritic and grapheme?",
      "votes": null
    },
    {
      "id": "779255",
      "postDate": "03/19/2020 06:40:25",
      "content": "<p>Hi <a href=\"/wowfattie\">@wowfattie</a> , Congrats for the medal. From where did you get unknown graphemes to train on ? </p>",
      "rawMarkdown": "Hi @wowfattie , Congrats for the medal. From where did you get unknown graphemes to train on ?",
      "votes": null
    },
    {
      "id": "779674",
      "postDate": "03/19/2020 15:30:41",
      "content": "<p>We only trained on known graphemes.</p>",
      "rawMarkdown": "We only trained on known graphemes.",
      "votes": null
    },
    {
      "id": "779686",
      "postDate": "03/19/2020 15:41:04",
      "content": "<p>For local OOD experiments, we split the 60 rarest graphemes out of the 1295 as validation set. For final LB submission, we simply trained the OOD models on all the train data.\nThe 3-head corresponded to grapheme_root, vowel_diacritic, consonant_diacritic</p>",
      "rawMarkdown": "For local OOD experiments, we split the 60 rarest graphemes out of the 1295 as validation set. For final LB submission, we simply trained the OOD models on all the train data.\nThe 3-head corresponded to grapheme_root, vowel_diacritic, consonant_diacritic",
      "votes": null
    },
    {
      "id": "779753",
      "postDate": "03/19/2020 16:51:27",
      "content": "<p>Oh, i see. Thanks! </p>",
      "rawMarkdown": "Oh, i see. Thanks!",
      "votes": null
    },
    {
      "id": "779831",
      "postDate": "03/19/2020 18:17:48",
      "content": "<p>Congratulation <a href=\"/wowfattie\">@wowfattie</a>, few queries:</p>\n\n<ul>\n<li>Arcface loss, is it one type of metric learning? what's the intuition to choose over CE? </li>\n<li>ID and OOD - If I'm not wrong, it's your validation strategies. You said, there were 60 rarest graphemes in your OOD. Would you please elaborate on them in detail?</li>\n<li>what is about classification threshold (cosine distance)?</li>\n</ul>\n\n<p>I found it very interesting. As you said, you first computed feature centers of each of the 1295 graphemes by the Arcface model and yes if so we can tell a test image whether it is in ID or OOD sets based on the feature distance. I am new to such an approach, would you please elaborate on this or kindly inform some references? </p>\n\n<p>If I understand your approach or pipeline:</p>\n\n<ul>\n<li>Training an Arcface model and compute feature centers of each 1295 graphemes.</li>\n<li>Using only ID sets, an Arcface model is used to classify all graphene classes (1295).</li>\n<li>Using only the OOD set, the model has 3 head (vowel, consonant) and 1 (grapheme_root) to predict the class</li>\n</ul>\n\n<p>How did you get the predicted score from ID sets? You have used a bigger image size. The original image size is a non-square image, a side question, why not chose the non-square image while if you make it square the image file may look strange by getting stretch, isn't it? </p>\n\n<p>It's a bit hard for me to understand your full pipeline, a visual demonstration would be awesome. My confusion comes from your ID and OOD sets and their respective models and outcome and conclusion. Would you please clarify? Thanks. </p>",
      "rawMarkdown": "Congratulation @wowfattie, few queries:\n\n- Arcface loss, is it one type of metric learning? what's the intuition to choose over CE? \n- ID and OOD - If I'm not wrong, it's your validation strategies. You said, there were 60 rarest graphemes in your OOD. Would you please elaborate on them in detail?\n- what is about classification threshold (cosine distance)?\n\nI found it very interesting. As you said, you first computed feature centers of each of the 1295 graphemes by the Arcface model and yes if so we can tell a test image whether it is in ID or OOD sets based on the feature distance. I am new to such an approach, would you please elaborate on this or kindly inform some references? \n\nIf I understand your approach or pipeline:\n\n- Training an Arcface model and compute feature centers of each 1295 graphemes.\n- Using only ID sets, an Arcface model is used to classify all graphene classes (1295).\n-  Using only the OOD set, the model has 3 head (vowel, consonant) and 1 (grapheme_root) to predict the class\n\nHow did you get the predicted score from ID sets? You have used a bigger image size. The original image size is a non-square image, a side question, why not chose the non-square image while if you make it square the image file may look strange by getting stretch, isn't it? \n\nIt's a bit hard for me to understand your full pipeline, a visual demonstration would be awesome. My confusion comes from your ID and OOD sets and their respective models and outcome and conclusion. Would you please clarify? Thanks.",
      "votes": null
    },
    {
      "id": "780340",
      "postDate": "03/20/2020 06:56:49",
      "content": "<p>Congratulation.\none question:\nHow is the feature centers of each of the 1295 graphemes computed?\nThanks</p>",
      "rawMarkdown": "Congratulation.\none question:\nHow is the feature centers of each of the 1295 graphemes computed?\nThanks",
      "votes": null
    },
    {
      "id": "780574",
      "postDate": "03/20/2020 12:13:08",
      "content": "<p>Hi <a href=\"/wowfattie\">@wowfattie</a>! Can u give more details on how to train the arcface? How to make splits on seen/unseen using 1295 graphemes in training dataset?</p>",
      "rawMarkdown": "Hi @wowfattie! Can u give more details on how to train the arcface? How to make splits on seen/unseen using 1295 graphemes in training dataset?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 778765,
      "author_name": "ogrellier",
      "author_url": "",
      "post_date": "03/18/2020 17:57:08",
      "content": "<p><a href=\"/wowfattie\">@wowfattie</a> <a href=\"/ybabakhin\">@ybabakhin</a> I have been humbled and privileged to team up with both of you. \nIt has been quite a journey for me in the image world !\nI'd like to really thank you both for your kindness, patience and cheerful moments we had ! \nI've learnt so much at your side.\nGreat Lesson ;-) </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 779015,
      "author_name": "cdeotte",
      "author_url": "",
      "post_date": "03/18/2020 23:55:26",
      "content": "<p>Wow, simple solution and powerful. Congrats team. Great job.</p>\n\n<p>In your OOD model, what are each of the 3 heads doing when classifying the 2 components vowel and consonant? In your ID model, how do you compute root, vowel, and consonant from your grapheme prediction?</p>",
      "votes": null,
      "replies": [
        {
          "id": 779029,
          "author_name": "wowfattie",
          "author_url": "",
          "post_date": "03/19/2020 00:39:10",
          "content": "<p>Thank you.\nThere was nothing special of our 3-head model at training stage, but during inference, we did not use the predictions of the \"grapheme_root\" head. We built dedicated classifiers (168-class classification, only 1-head) for grapheme_root predictions. \nOur ID (Arcface) models predicted grapheme classes (1295-class), and we could directly lookup the three components from train.csv.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 779152,
      "author_name": "yuanlin08",
      "author_url": "",
      "post_date": "03/19/2020 04:13:42",
      "content": "<p>Congrats! Thanks for sharing. May I ask a few questions?\n1.  How do you train your OOD? I mean, how do you divide your training set for ID and OOD? \n2. About the 3-head of OOD, are they voweldiacritic, consonantdiacritic and grapheme?</p>",
      "votes": null,
      "replies": [
        {
          "id": 779686,
          "author_name": "wowfattie",
          "author_url": "",
          "post_date": "03/19/2020 15:41:04",
          "content": "<p>For local OOD experiments, we split the 60 rarest graphemes out of the 1295 as validation set. For final LB submission, we simply trained the OOD models on all the train data.\nThe 3-head corresponded to grapheme_root, vowel_diacritic, consonant_diacritic</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 779753,
          "author_name": "yuanlin08",
          "author_url": "",
          "post_date": "03/19/2020 16:51:27",
          "content": "<p>Oh, i see. Thanks! </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 779255,
      "author_name": "virajbagal",
      "author_url": "",
      "post_date": "03/19/2020 06:40:25",
      "content": "<p>Hi <a href=\"/wowfattie\">@wowfattie</a> , Congrats for the medal. From where did you get unknown graphemes to train on ? </p>",
      "votes": null,
      "replies": [
        {
          "id": 779674,
          "author_name": "wowfattie",
          "author_url": "",
          "post_date": "03/19/2020 15:30:41",
          "content": "<p>We only trained on known graphemes.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 779831,
      "author_name": "ipythonx",
      "author_url": "",
      "post_date": "03/19/2020 18:17:48",
      "content": "<p>Congratulation <a href=\"/wowfattie\">@wowfattie</a>, few queries:</p>\n\n<ul>\n<li>Arcface loss, is it one type of metric learning? what's the intuition to choose over CE? </li>\n<li>ID and OOD - If I'm not wrong, it's your validation strategies. You said, there were 60 rarest graphemes in your OOD. Would you please elaborate on them in detail?</li>\n<li>what is about classification threshold (cosine distance)?</li>\n</ul>\n\n<p>I found it very interesting. As you said, you first computed feature centers of each of the 1295 graphemes by the Arcface model and yes if so we can tell a test image whether it is in ID or OOD sets based on the feature distance. I am new to such an approach, would you please elaborate on this or kindly inform some references? </p>\n\n<p>If I understand your approach or pipeline:</p>\n\n<ul>\n<li>Training an Arcface model and compute feature centers of each 1295 graphemes.</li>\n<li>Using only ID sets, an Arcface model is used to classify all graphene classes (1295).</li>\n<li>Using only the OOD set, the model has 3 head (vowel, consonant) and 1 (grapheme_root) to predict the class</li>\n</ul>\n\n<p>How did you get the predicted score from ID sets? You have used a bigger image size. The original image size is a non-square image, a side question, why not chose the non-square image while if you make it square the image file may look strange by getting stretch, isn't it? </p>\n\n<p>It's a bit hard for me to understand your full pipeline, a visual demonstration would be awesome. My confusion comes from your ID and OOD sets and their respective models and outcome and conclusion. Would you please clarify? Thanks. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 780340,
      "author_name": "mingxingliu",
      "author_url": "",
      "post_date": "03/20/2020 06:56:49",
      "content": "<p>Congratulation.\none question:\nHow is the feature centers of each of the 1295 graphemes computed?\nThanks</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 780574,
      "author_name": "arturdatascientist",
      "author_url": "",
      "post_date": "03/20/2020 12:13:08",
      "content": "<p>Hi <a href=\"/wowfattie\">@wowfattie</a>! Can u give more details on how to train the arcface? How to make splits on seen/unseen using 1295 graphemes in training dataset?</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "778756": "We appreciate the efforts Bengali.AI and Kaggle spent on this competition, and congrats to all the winners. We are very impressed by your interesting findings, novol methods, deep understanding of the problem, the dataset, as well as the evaluation metric.\n\nOur solution is relatively straightforward and bears similarities with some of the top solutions.\n\nLet's call the 1295 graphemes in the train set ID (in-dictionary), and the unknown graphemes in the test set OOD (out-of-dictionary). The first step was training Arcface models and computing the feature centers of each of the 1295 graphemes. We can then tell if a test image is ID or OOD based on its smallest feature distance to every grapheme centers. In our 4th place submission, we set the classification threshold to 0.15 (cosine distance), which was estimated locally. After this ID/OOD classification, we applied the Arcface models to the ID set only and determined the component classes by the grapheme classes. For the OOD set, we trained another group of 1-head and 3-head component models and applied them to the OOD set only.\n\n**More of the Arcface models:**\nnetwork structures: inception_resnet_v2 and seresnext101\ninput image preprocessing: plain resizing\ninput image size: 320x320\nloss: Arcface\ntarget: 1295 grapheme classification\ntotal training epochs: 90\naugmentations: cutout only\n\n**More of the OOD models:**\nnetwork structures: inceptions, resnets, densnets, efficientnets ... ... ...\ninput image preprocessing: https://www.kaggle.com/iafoss/image-preprocessing-128x128\ninput image size: various sizes from 256x256 to 416x416\nloss: CE\ntarget: 3-head for vowel_diacritic and consonant_diacritic, 1-head for grapheme_root\ntotal training epochs: 10 (more training harmed the OOD performance)\naugmentations: mixup only",
    "778765": "wowfattie @ybabakhin I have been humbled and privileged to team up with both of you. \nIt has been quite a journey for me in the image world !\nI'd like to really thank you both for your kindness, patience and cheerful moments we had ! \nI've learnt so much at your side.\nGreat Lesson ;-)",
    "779015": "Wow, simple solution and powerful. Congrats team. Great job.\n\nIn your OOD model, what are each of the 3 heads doing when classifying the 2 components vowel and consonant? In your ID model, how do you compute root, vowel, and consonant from your grapheme prediction?",
    "779029": "Thank you.\nThere was nothing special of our 3-head model at training stage, but during inference, we did not use the predictions of the \"grapheme_root\" head. We built dedicated classifiers (168-class classification, only 1-head) for grapheme_root predictions. \nOur ID (Arcface) models predicted grapheme classes (1295-class), and we could directly lookup the three components from train.csv.",
    "779152": "Congrats! Thanks for sharing. May I ask a few questions?\n1.  How do you train your OOD? I mean, how do you divide your training set for ID and OOD? \n2. About the 3-head of OOD, are they voweldiacritic, consonantdiacritic and grapheme?",
    "779255": "Hi @wowfattie , Congrats for the medal. From where did you get unknown graphemes to train on ?",
    "779674": "We only trained on known graphemes.",
    "779686": "For local OOD experiments, we split the 60 rarest graphemes out of the 1295 as validation set. For final LB submission, we simply trained the OOD models on all the train data.\nThe 3-head corresponded to grapheme_root, vowel_diacritic, consonant_diacritic",
    "779753": "Oh, i see. Thanks!",
    "779831": "Congratulation @wowfattie, few queries:\n\n- Arcface loss, is it one type of metric learning? what's the intuition to choose over CE? \n- ID and OOD - If I'm not wrong, it's your validation strategies. You said, there were 60 rarest graphemes in your OOD. Would you please elaborate on them in detail?\n- what is about classification threshold (cosine distance)?\n\nI found it very interesting. As you said, you first computed feature centers of each of the 1295 graphemes by the Arcface model and yes if so we can tell a test image whether it is in ID or OOD sets based on the feature distance. I am new to such an approach, would you please elaborate on this or kindly inform some references? \n\nIf I understand your approach or pipeline:\n\n- Training an Arcface model and compute feature centers of each 1295 graphemes.\n- Using only ID sets, an Arcface model is used to classify all graphene classes (1295).\n-  Using only the OOD set, the model has 3 head (vowel, consonant) and 1 (grapheme_root) to predict the class\n\nHow did you get the predicted score from ID sets? You have used a bigger image size. The original image size is a non-square image, a side question, why not chose the non-square image while if you make it square the image file may look strange by getting stretch, isn't it? \n\nIt's a bit hard for me to understand your full pipeline, a visual demonstration would be awesome. My confusion comes from your ID and OOD sets and their respective models and outcome and conclusion. Would you please clarify? Thanks.",
    "780340": "Congratulation.\none question:\nHow is the feature centers of each of the 1295 graphemes computed?\nThanks",
    "780574": "Hi @wowfattie! Can u give more details on how to train the arcface? How to make splits on seen/unseen using 1295 graphemes in training dataset?"
  },
  "source": "meta"
}