{
  "id": 126209,
  "title": "Strange tradeoff",
  "url": "/competitions/bengaliai-cv19/discussion/126209",
  "author_name": "Bibek",
  "post_date": "2020-01-16T09:36:55.911000",
  "votes": 6,
  "comment_count": 27,
  "views": 0,
  "content": "<p>As I see improvements in my local CV, I also notice a strange tradeoff in the CV scores of <code>grapheme_root</code>, <code>vowel_diacritic</code> and <code>consonant_diacritic</code>. </p>\n\n<p>Improvement in the score of <code>grapheme_root</code> leads to relative decrease in the scores of <code>vowel_diacritic</code> and <code>consonant_diacritic</code>, and vice-versa. So model with high score of <code>grapheme_root</code> but with low overall CV may score better on LB compared to the model with high CV and low <code>grapheme_root</code> score. This is what exactly happened to me recently, first time when improvement in CV did not correlate to LB increase. </p>\n\n<p>So make sure you split your CV strategy into overall CV and class scores</p>",
  "messages": [
    {
      "id": 720273,
      "postDate": "2020-01-16T09:36:55.910Z",
      "content": "<p>As I see improvements in my local CV, I also notice a strange tradeoff in the CV scores of <code>grapheme_root</code>, <code>vowel_diacritic</code> and <code>consonant_diacritic</code>. </p>\n\n<p>Improvement in the score of <code>grapheme_root</code> leads to relative decrease in the scores of <code>vowel_diacritic</code> and <code>consonant_diacritic</code>, and vice-versa. So model with high score of <code>grapheme_root</code> but with low overall CV may score better on LB compared to the model with high CV and low <code>grapheme_root</code> score. This is what exactly happened to me recently, first time when improvement in CV did not correlate to LB increase. </p>\n\n<p>So make sure you split your CV strategy into overall CV and class scores</p>",
      "rawMarkdown": "As I see improvements in my local CV, I also notice a strange tradeoff in the CV scores of `grapheme_root`, `vowel_diacritic` and `consonant_diacritic`. \n\nImprovement in the score of `grapheme_root` leads to relative decrease in the scores of `vowel_diacritic` and `consonant_diacritic`, and vice-versa. So model with high score of `grapheme_root` but with low overall CV may score better on LB compared to the model with high CV and low `grapheme_root` score. This is what exactly happened to me recently, first time when improvement in CV did not correlate to LB increase. \n\nSo make sure you split your CV strategy into overall CV and class scores\n",
      "votes": 6
    },
    {
      "id": 723248,
      "postDate": "2020-01-19T18:19:30.753Z",
      "content": "<p><a href=\"/bibek777\">@bibek777</a> it might be due to the learning rate, if you have the same for each head. I saw with a basic resnet18 that training with a high learning rate like 0.1 made the <code>vowel_diacritic</code> and <code>consonant_diacritic</code> learn well on first 5 epochs, while <code>grapheme_root</code> was doing worse than random systematically.</p>\n\n<p>Not sure if this is just because of different number of classes or if it's really something.\nNot sure what can be done either : \n- different learning rates for different heads (but if you have just a linear layer in each head it probably won't help much)\n- freeze two heads and fine tune the last one alternatively on each type (vowel, consonant or grapheme - per batch or per epoch maybe) with different learning rates, this seems a bit weird but I wonder if anyone tried it.</p>",
      "rawMarkdown": "@bibek777 it might be due to the learning rate, if you have the same for each head. I saw with a basic resnet18 that training with a high learning rate like 0.1 made the `vowel_diacritic` and `consonant_diacritic` learn well on first 5 epochs, while `grapheme_root` was doing worse than random systematically.\n\nNot sure if this is just because of different number of classes or if it's really something.\nNot sure what can be done either : \n- different learning rates for different heads (but if you have just a linear layer in each head it probably won't help much)\n- freeze two heads and fine tune the last one alternatively on each type (vowel, consonant or grapheme - per batch or per epoch maybe) with different learning rates, this seems a bit weird but I wonder if anyone tried it.",
      "votes": 1
    },
    {
      "id": 720318,
      "postDate": "2020-01-16T10:17:13.213Z",
      "content": "<p><a href=\"/bibek777\">@bibek777</a> what do you mean by splitting your CV strategy into overall CV and class scores is the way forward</p>",
      "rawMarkdown": "@bibek777 what do you mean by splitting your CV strategy into overall CV and class scores is the way forward",
      "votes": 1,
      "replies": [
        {
          "id": 720323,
          "postDate": "2020-01-16T10:20:33.330Z",
          "content": "<p>something like this:\n<code>\nOverall CV: XXX\ngrapheme_root : xxx\nvowel_diacritic: yyy\nconsonant_diacritic: zzz\n</code></p>",
          "rawMarkdown": "something like this:\n```\nOverall CV: XXX\ngrapheme_root : xxx\nvowel_diacritic: yyy\nconsonant_diacritic: zzz\n```"
        }
      ]
    },
    {
      "id": 720315,
      "postDate": "2020-01-16T10:16:11.820Z",
      "content": "<blockquote>\n  <p>Looks like we have more grapheme_root samples</p>\n</blockquote>\n\n<p>The number of samples of all three components is going to be equal. <br>\nbut looks like grapheme root part is more difficult to classify in the test set as compared to the training set.</p>",
      "rawMarkdown": "&gt; Looks like we have more grapheme_root samples\n\nThe number of samples of all three components is going to be equal.   \nbut looks like grapheme root part is more difficult to classify in the test set as compared to the training set.",
      "votes": 1,
      "replies": [
        {
          "id": 720317,
          "postDate": "2020-01-16T10:17:08.903Z",
          "content": "<blockquote>\n  <p>The number of samples of all three components is going to be equal.</p>\n</blockquote>\n\n<p>what makes you think that??</p>",
          "rawMarkdown": "&gt; The number of samples of all three components is going to be equal.\n\nwhat makes you think that??"
        },
        {
          "id": 720319,
          "postDate": "2020-01-16T10:17:54.647Z",
          "content": "<p>because each test image contains all the three components.</p>",
          "rawMarkdown": "because each test image contains all the three components.",
          "votes": 2
        },
        {
          "id": 720325,
          "postDate": "2020-01-16T10:21:34.567Z",
          "content": "<p>Thanx!! I will update my post</p>",
          "rawMarkdown": "Thanx!! I will update my post"
        },
        {
          "id": 720334,
          "postDate": "2020-01-16T10:28:01.703Z",
          "content": "<p>but this doesn't explain why increase in the scores of  <code>grapheme_root</code> leads to decrease in scores of <code>vowel_diacritic</code> and  <code>consonant_diacritic</code>, and vice-versa?</p>",
          "rawMarkdown": "but this doesn't explain why increase in the scores of  `grapheme_root` leads to decrease in scores of `vowel_diacritic` and  `consonant_diacritic`, and vice-versa?"
        },
        {
          "id": 723625,
          "postDate": "2020-01-20T08:50:21.730Z",
          "content": "<p><a href=\"/dhananjay3\">@dhananjay3</a> I am sorry , how do we know that all test images contain all three component ?  if consonant_diacritics and vowel_diacritics are 0 that means there are no such component to detect right ? Does that not make it \"Class Not Present\" for some samples ?\ne.g If we create a classifier with vowel_diacritic = 0 and non-zero , will it not become a binary classification problem? Whereas , there is always a grapeheme_root present in an image .</p>",
          "rawMarkdown": "@dhananjay3 I am sorry , how do we know that all test images contain all three component ?  if consonant_diacritics and vowel_diacritics are 0 that means there are no such component to detect right ? Does that not make it \"Class Not Present\" for some samples ?\ne.g If we create a classifier with vowel_diacritic = 0 and non-zero , will it not become a binary classification problem? Whereas , there is always a grapeheme_root present in an image ."
        },
        {
          "id": 723692,
          "postDate": "2020-01-20T10:42:36.080Z",
          "content": "<p><a href=\"/phoenix9032\">@phoenix9032</a> 0 does mean that image does not contain consonant/vowel but we still have to predict that it is not there in the image and score for this will be included in the recall. so it is just another class to classify into.</p>",
          "rawMarkdown": "@phoenix9032 0 does mean that image does not contain consonant/vowel but we still have to predict that it is not there in the image and score for this will be included in the recall. so it is just another class to classify into.",
          "votes": 2
        }
      ]
    },
    {
      "id": 720307,
      "postDate": "2020-01-16T10:08:37.100Z",
      "content": "<p>In such case, do you think that bagging models that would respectively perform their best on either <code>grapheme_root</code> or <code>vowel_diacritic</code> and <code>consonant_diacritic</code> could be a good idea ? I have not tested it myself for now !</p>",
      "rawMarkdown": "In such case, do you think that bagging models that would respectively perform their best on either `grapheme_root` or `vowel_diacritic` and `consonant_diacritic` could be a good idea ? I have not tested it myself for now !",
      "votes": 1,
      "replies": [
        {
          "id": 720309,
          "postDate": "2020-01-16T10:10:41.237Z",
          "content": "<p>that could be a nice thing to try once you run out of ideas. for now I would recommend you to improve your single model for all classes</p>",
          "rawMarkdown": "that could be a nice thing to try once you run out of ideas. for now I would recommend you to improve your single model for all classes"
        }
      ]
    },
    {
      "id": 724870,
      "postDate": "2020-01-21T15:30:25.220Z",
      "content": "<p><a href=\"/bibek777\">@bibek777</a> Even I am facing the same problem, I think there are few damaged images (improper crop or improper labelling) which decreases our score !!! as mentioned here - <a href=\"https://www.kaggle.com/c/bengaliai-cv19/discussion/126833\">https://www.kaggle.com/c/bengaliai-cv19/discussion/126833</a> by Nihar. I am personally trying different model together for vowel and consonant and a separate model for graphemes</p>",
      "rawMarkdown": "@bibek777 Even I am facing the same problem, I think there are few damaged images (improper crop or improper labelling) which decreases our score !!! as mentioned here - https://www.kaggle.com/c/bengaliai-cv19/discussion/126833 by Nihar. I am personally trying different model together for vowel and consonant and a separate model for graphemes"
    },
    {
      "id": 723908,
      "postDate": "2020-01-20T15:41:33.653Z",
      "content": "<p>I have also tried to reduce not the <code>val_loss</code> but the <code>grapheme_root_val_loss</code> which seemed to give some interesting results. I didn't necessarily measure everything correctly (which I need to start getting in the habit of doing though) but it might be worth investigating more in-depth.</p>",
      "rawMarkdown": "I have also tried to reduce not the `val_loss` but the `grapheme_root_val_loss` which seemed to give some interesting results. I didn't necessarily measure everything correctly (which I need to start getting in the habit of doing though) but it might be worth investigating more in-depth."
    },
    {
      "id": 720908,
      "postDate": "2020-01-16T21:03:26.583Z",
      "content": "<p>I have been seeing something similar. Maybe adding some more layers which are separate would help (I'm guessing you have a single fully connected layer for each output...)?\n<img src=\"https://i.imgur.com/iB1dxkk.png\" alt=\"\"></p>",
      "rawMarkdown": "I have been seeing something similar. Maybe adding some more layers which are separate would help (I'm guessing you have a single fully connected layer for each output...)?\n![](https://i.imgur.com/iB1dxkk.png)"
    },
    {
      "id": 720755,
      "postDate": "2020-01-16T17:45:02.817Z",
      "content": "<p>why not have two separate models? \nOne for grapheme and other for the vowel and consonant? (I am pretty confident this would yield good results, but would increase submission time drastically)</p>",
      "rawMarkdown": "why not have two separate models? \nOne for grapheme and other for the vowel and consonant? (I am pretty confident this would yield good results, but would increase submission time drastically)",
      "replies": [
        {
          "id": 721187,
          "postDate": "2020-01-17T06:38:04.770Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 745433,
          "postDate": "2020-02-13T20:07:53.163Z",
          "content": "<blockquote>\n  <p>Multi-task learning (MTL) is a subfield of machine learning in which multiple learning tasks are solved at the same time, while exploiting commonalities and differences across tasks. This can result in improved learning efficiency and prediction accuracy for the task-specific models, when compared to training the models separately. Early versions of MTL were called \"hints\". </p>\n</blockquote>\n\n<p><a href=\"https://en.wikipedia.org/wiki/Multi-task_learning\">https://en.wikipedia.org/wiki/Multi-task_learning</a></p>\n\n<p>Plus, as you've noted, increase compute, increased train time, increased inference time, increase ram resources, etc.</p>",
          "rawMarkdown": "&gt; Multi-task learning (MTL) is a subfield of machine learning in which multiple learning tasks are solved at the same time, while exploiting commonalities and differences across tasks. This can result in improved learning efficiency and prediction accuracy for the task-specific models, when compared to training the models separately. Early versions of MTL were called \"hints\". \n\nhttps://en.wikipedia.org/wiki/Multi-task_learning\n\nPlus, as you've noted, increase compute, increased train time, increased inference time, increase ram resources, etc."
        }
      ]
    },
    {
      "id": 720721,
      "postDate": "2020-01-16T17:12:21.990Z",
      "rawMarkdown": "",
      "votes": -1,
      "isDeleted": true,
      "replies": [
        {
          "id": 720764,
          "postDate": "2020-01-16T17:56:10.593Z",
          "content": "<blockquote>\n  <p>you should be able to prove class distributions of popular combinations to see if they are consistently different from the train set</p>\n</blockquote>\n\n<p>things are getting interesting now!!!</p>",
          "rawMarkdown": "&gt; you should be able to prove class distributions of popular combinations to see if they are consistently different from the train set\n\nthings are getting interesting now!!!"
        },
        {
          "id": 720782,
          "postDate": "2020-01-16T18:13:11.513Z",
          "content": "<p>Looks like failed submissions are correctly applied towards daily limits, but aren't properly showing up in the displayed submission count. I'll file a bug report, but unlimited probing should be infeasible.</p>",
          "rawMarkdown": "Looks like failed submissions are correctly applied towards daily limits, but aren't properly showing up in the displayed submission count. I'll file a bug report, but unlimited probing should be infeasible.",
          "votes": 3
        },
        {
          "id": 720786,
          "postDate": "2020-01-16T18:17:44.040Z",
          "rawMarkdown": "",
          "votes": -1,
          "isDeleted": true
        },
        {
          "id": 720791,
          "postDate": "2020-01-16T18:28:34.577Z",
          "content": "<p>I don't think private probing is possible in this competition. I mean, technically, what you wrote is possible, but without a good model, you don't have the information you'd like to leak. (Or am I missing something?)</p>",
          "rawMarkdown": "I don't think private probing is possible in this competition. I mean, technically, what you wrote is possible, but without a good model, you don't have the information you'd like to leak. (Or am I missing something?)"
        },
        {
          "id": 720795,
          "postDate": "2020-01-16T18:39:04.067Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 721083,
          "postDate": "2020-01-17T03:04:06.610Z",
          "content": "<p>I didn't even thought about probing public LB test set.\nThe most important thing is that you might not be able to tell whether private LB test set has a same distribution as public LB test set.\nI think probing public LB can only help you overfitting to public LB and nothing more.</p>",
          "rawMarkdown": "I didn't even thought about probing public LB test set.\nThe most important thing is that you might not be able to tell whether private LB test set has a same distribution as public LB test set.\nI think probing public LB can only help you overfitting to public LB and nothing more.",
          "votes": 6
        },
        {
          "id": 721135,
          "postDate": "2020-01-17T04:57:31.557Z",
          "content": "<p>probing test dataset is a competition trick. i think it is not required here. i estimated that by using ML methods, 0.99+ is possible for private test set as i suggest using normal ML methods is sufficient.</p>\n\n<p>on a side note, you can take a look at the kaggle google doodle competition on how knowing the number of class instances can improve results (i.e. greedy selection of the top results for k classes)</p>",
          "rawMarkdown": "probing test dataset is a competition trick. i think it is not required here. i estimated that by using ML methods, 0.99+ is possible for private test set as i suggest using normal ML methods is sufficient.\n\non a side note, you can take a look at the kaggle google doodle competition on how knowing the number of class instances can improve results (i.e. greedy selection of the top results for k classes)"
        },
        {
          "id": 722607,
          "postDate": "2020-01-18T19:53:55.790Z",
          "content": "<p>This might be an unpopular opinion, but I would hope that we could stay away from such practices. The goal here is still to provide some help to a challenging problem, that has useful applications. While I do get that by having mechanics of gamification and leader-boards with prize money, from a purely game theory point-of-view public LB probing is more and more of a thing; I still believe we shouldn't really encourage this type of practice.</p>\n\n<p>That also goes with the fact that purely from a utilitarian point of view, it probably doesn't help much, because of the fact that your private LB has unseen characters, I agree with <a href=\"/haqishen\">@haqishen</a> there. If you want to do probing like this, you might as well also just look at the distribution of each root, vowel &amp; consonant in the Bengali language, that at least would be beneficial for those hosting the competition.</p>\n\n<p>Don't get me wrong, I think those are all very interesting ideas, and I'm pretty amazed at how smart some can be. I just think it's better if we can also have a discussion about it.</p>",
          "rawMarkdown": "This might be an unpopular opinion, but I would hope that we could stay away from such practices. The goal here is still to provide some help to a challenging problem, that has useful applications. While I do get that by having mechanics of gamification and leader-boards with prize money, from a purely game theory point-of-view public LB probing is more and more of a thing; I still believe we shouldn't really encourage this type of practice.\n\nThat also goes with the fact that purely from a utilitarian point of view, it probably doesn't help much, because of the fact that your private LB has unseen characters, I agree with @haqishen there. If you want to do probing like this, you might as well also just look at the distribution of each root, vowel &amp; consonant in the Bengali language, that at least would be beneficial for those hosting the competition.\n\nDon't get me wrong, I think those are all very interesting ideas, and I'm pretty amazed at how smart some can be. I just think it's better if we can also have a discussion about it.",
          "votes": 5
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 723248,
      "author_name": "Optimo",
      "author_url": "",
      "post_date": "2020-01-19T18:19:30.753000",
      "content": "<p><a href=\"/bibek777\">@bibek777</a> it might be due to the learning rate, if you have the same for each head. I saw with a basic resnet18 that training with a high learning rate like 0.1 made the <code>vowel_diacritic</code> and <code>consonant_diacritic</code> learn well on first 5 epochs, while <code>grapheme_root</code> was doing worse than random systematically.</p>\n\n<p>Not sure if this is just because of different number of classes or if it's really something.\nNot sure what can be done either : \n- different learning rates for different heads (but if you have just a linear layer in each head it probably won't help much)\n- freeze two heads and fine tune the last one alternatively on each type (vowel, consonant or grapheme - per batch or per epoch maybe) with different learning rates, this seems a bit weird but I wonder if anyone tried it.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 720318,
      "author_name": "Dhananjay Raut",
      "author_url": "",
      "post_date": "2020-01-16T10:17:13.213000",
      "content": "<p><a href=\"/bibek777\">@bibek777</a> what do you mean by splitting your CV strategy into overall CV and class scores is the way forward</p>",
      "votes": 1,
      "replies": [
        {
          "id": 720323,
          "author_name": "Bibek",
          "author_url": "",
          "post_date": "2020-01-16T10:20:33.330000",
          "content": "<p>something like this:\n<code>\nOverall CV: XXX\ngrapheme_root : xxx\nvowel_diacritic: yyy\nconsonant_diacritic: zzz\n</code></p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 720315,
      "author_name": "Dhananjay Raut",
      "author_url": "",
      "post_date": "2020-01-16T10:16:11.820000",
      "content": "<blockquote>\n  <p>Looks like we have more grapheme_root samples</p>\n</blockquote>\n\n<p>The number of samples of all three components is going to be equal. <br>\nbut looks like grapheme root part is more difficult to classify in the test set as compared to the training set.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 720317,
          "author_name": "Bibek",
          "author_url": "",
          "post_date": "2020-01-16T10:17:08.903000",
          "content": "<blockquote>\n  <p>The number of samples of all three components is going to be equal.</p>\n</blockquote>\n\n<p>what makes you think that??</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 720319,
          "author_name": "Dhananjay Raut",
          "author_url": "",
          "post_date": "2020-01-16T10:17:54.647000",
          "content": "<p>because each test image contains all the three components.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 720325,
          "author_name": "Bibek",
          "author_url": "",
          "post_date": "2020-01-16T10:21:34.567000",
          "content": "<p>Thanx!! I will update my post</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 720334,
          "author_name": "Bibek",
          "author_url": "",
          "post_date": "2020-01-16T10:28:01.703000",
          "content": "<p>but this doesn't explain why increase in the scores of  <code>grapheme_root</code> leads to decrease in scores of <code>vowel_diacritic</code> and  <code>consonant_diacritic</code>, and vice-versa?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 723625,
          "author_name": "Nirjhar Roy",
          "author_url": "",
          "post_date": "2020-01-20T08:50:21.730000",
          "content": "<p><a href=\"/dhananjay3\">@dhananjay3</a> I am sorry , how do we know that all test images contain all three component ?  if consonant_diacritics and vowel_diacritics are 0 that means there are no such component to detect right ? Does that not make it \"Class Not Present\" for some samples ?\ne.g If we create a classifier with vowel_diacritic = 0 and non-zero , will it not become a binary classification problem? Whereas , there is always a grapeheme_root present in an image .</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 723692,
          "author_name": "Dhananjay Raut",
          "author_url": "",
          "post_date": "2020-01-20T10:42:36.080000",
          "content": "<p><a href=\"/phoenix9032\">@phoenix9032</a> 0 does mean that image does not contain consonant/vowel but we still have to predict that it is not there in the image and score for this will be included in the recall. so it is just another class to classify into.</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 720307,
      "author_name": "Thomas Di Martino",
      "author_url": "",
      "post_date": "2020-01-16T10:08:37.100000",
      "content": "<p>In such case, do you think that bagging models that would respectively perform their best on either <code>grapheme_root</code> or <code>vowel_diacritic</code> and <code>consonant_diacritic</code> could be a good idea ? I have not tested it myself for now !</p>",
      "votes": 1,
      "replies": [
        {
          "id": 720309,
          "author_name": "Bibek",
          "author_url": "",
          "post_date": "2020-01-16T10:10:41.237000",
          "content": "<p>that could be a nice thing to try once you run out of ideas. for now I would recommend you to improve your single model for all classes</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 724870,
      "author_name": "A/C",
      "author_url": "",
      "post_date": "2020-01-21T15:30:25.220000",
      "content": "<p><a href=\"/bibek777\">@bibek777</a> Even I am facing the same problem, I think there are few damaged images (improper crop or improper labelling) which decreases our score !!! as mentioned here - <a href=\"https://www.kaggle.com/c/bengaliai-cv19/discussion/126833\">https://www.kaggle.com/c/bengaliai-cv19/discussion/126833</a> by Nihar. I am personally trying different model together for vowel and consonant and a separate model for graphemes</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 723908,
      "author_name": "Maxime Lenormand",
      "author_url": "",
      "post_date": "2020-01-20T15:41:33.653000",
      "content": "<p>I have also tried to reduce not the <code>val_loss</code> but the <code>grapheme_root_val_loss</code> which seemed to give some interesting results. I didn't necessarily measure everything correctly (which I need to start getting in the habit of doing though) but it might be worth investigating more in-depth.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 720908,
      "author_name": "Ollie Perrée",
      "author_url": "",
      "post_date": "2020-01-16T21:03:26.583000",
      "content": "<p>I have been seeing something similar. Maybe adding some more layers which are separate would help (I'm guessing you have a single fully connected layer for each output...)?\n<img src=\"https://i.imgur.com/iB1dxkk.png\" alt=\"\"></p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 720755,
      "author_name": "timetraveller",
      "author_url": "",
      "post_date": "2020-01-16T17:45:02.817000",
      "content": "<p>why not have two separate models? \nOne for grapheme and other for the vowel and consonant? (I am pretty confident this would yield good results, but would increase submission time drastically)</p>",
      "votes": 0,
      "replies": [
        {
          "id": 721187,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-01-17T06:38:04.770000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 745433,
          "author_name": "عثمان",
          "author_url": "",
          "post_date": "2020-02-13T20:07:53.163000",
          "content": "<blockquote>\n  <p>Multi-task learning (MTL) is a subfield of machine learning in which multiple learning tasks are solved at the same time, while exploiting commonalities and differences across tasks. This can result in improved learning efficiency and prediction accuracy for the task-specific models, when compared to training the models separately. Early versions of MTL were called \"hints\". </p>\n</blockquote>\n\n<p><a href=\"https://en.wikipedia.org/wiki/Multi-task_learning\">https://en.wikipedia.org/wiki/Multi-task_learning</a></p>\n\n<p>Plus, as you've noted, increase compute, increased train time, increased inference time, increase ram resources, etc.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 720721,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-01-16T17:12:21.990000",
      "content": "",
      "votes": -1,
      "replies": [
        {
          "id": 720764,
          "author_name": "Bibek",
          "author_url": "",
          "post_date": "2020-01-16T17:56:10.593000",
          "content": "<blockquote>\n  <p>you should be able to prove class distributions of popular combinations to see if they are consistently different from the train set</p>\n</blockquote>\n\n<p>things are getting interesting now!!!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 720782,
          "author_name": "Sohier Dane",
          "author_url": "",
          "post_date": "2020-01-16T18:13:11.513000",
          "content": "<p>Looks like failed submissions are correctly applied towards daily limits, but aren't properly showing up in the displayed submission count. I'll file a bug report, but unlimited probing should be infeasible.</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 720786,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-01-16T18:17:44.040000",
          "content": "",
          "votes": -1,
          "replies": []
        },
        {
          "id": 720791,
          "author_name": "Peter",
          "author_url": "",
          "post_date": "2020-01-16T18:28:34.577000",
          "content": "<p>I don't think private probing is possible in this competition. I mean, technically, what you wrote is possible, but without a good model, you don't have the information you'd like to leak. (Or am I missing something?)</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 720795,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-01-16T18:39:04.067000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 721083,
          "author_name": "Qishen Ha",
          "author_url": "",
          "post_date": "2020-01-17T03:04:06.610000",
          "content": "<p>I didn't even thought about probing public LB test set.\nThe most important thing is that you might not be able to tell whether private LB test set has a same distribution as public LB test set.\nI think probing public LB can only help you overfitting to public LB and nothing more.</p>",
          "votes": 6,
          "replies": []
        },
        {
          "id": 721135,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2020-01-17T04:57:31.557000",
          "content": "<p>probing test dataset is a competition trick. i think it is not required here. i estimated that by using ML methods, 0.99+ is possible for private test set as i suggest using normal ML methods is sufficient.</p>\n\n<p>on a side note, you can take a look at the kaggle google doodle competition on how knowing the number of class instances can improve results (i.e. greedy selection of the top results for k classes)</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 722607,
          "author_name": "Maxime Lenormand",
          "author_url": "",
          "post_date": "2020-01-18T19:53:55.790000",
          "content": "<p>This might be an unpopular opinion, but I would hope that we could stay away from such practices. The goal here is still to provide some help to a challenging problem, that has useful applications. While I do get that by having mechanics of gamification and leader-boards with prize money, from a purely game theory point-of-view public LB probing is more and more of a thing; I still believe we shouldn't really encourage this type of practice.</p>\n\n<p>That also goes with the fact that purely from a utilitarian point of view, it probably doesn't help much, because of the fact that your private LB has unseen characters, I agree with <a href=\"/haqishen\">@haqishen</a> there. If you want to do probing like this, you might as well also just look at the distribution of each root, vowel &amp; consonant in the Bengali language, that at least would be beneficial for those hosting the competition.</p>\n\n<p>Don't get me wrong, I think those are all very interesting ideas, and I'm pretty amazed at how smart some can be. I just think it's better if we can also have a discussion about it.</p>",
          "votes": 5,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "720273": "As I see improvements in my local CV, I also notice a strange tradeoff in the CV scores of `grapheme_root`, `vowel_diacritic` and `consonant_diacritic`. \n\nImprovement in the score of `grapheme_root` leads to relative decrease in the scores of `vowel_diacritic` and `consonant_diacritic`, and vice-versa. So model with high score of `grapheme_root` but with low overall CV may score better on LB compared to the model with high CV and low `grapheme_root` score. This is what exactly happened to me recently, first time when improvement in CV did not correlate to LB increase. \n\nSo make sure you split your CV strategy into overall CV and class scores\n",
    "723248": "@bibek777 it might be due to the learning rate, if you have the same for each head. I saw with a basic resnet18 that training with a high learning rate like 0.1 made the `vowel_diacritic` and `consonant_diacritic` learn well on first 5 epochs, while `grapheme_root` was doing worse than random systematically.\n\nNot sure if this is just because of different number of classes or if it's really something.\nNot sure what can be done either : \n- different learning rates for different heads (but if you have just a linear layer in each head it probably won't help much)\n- freeze two heads and fine tune the last one alternatively on each type (vowel, consonant or grapheme - per batch or per epoch maybe) with different learning rates, this seems a bit weird but I wonder if anyone tried it.",
    "720318": "@bibek777 what do you mean by splitting your CV strategy into overall CV and class scores is the way forward",
    "720315": "&gt; Looks like we have more grapheme_root samples\n\nThe number of samples of all three components is going to be equal.   \nbut looks like grapheme root part is more difficult to classify in the test set as compared to the training set.",
    "720307": "In such case, do you think that bagging models that would respectively perform their best on either `grapheme_root` or `vowel_diacritic` and `consonant_diacritic` could be a good idea ? I have not tested it myself for now !",
    "724870": "@bibek777 Even I am facing the same problem, I think there are few damaged images (improper crop or improper labelling) which decreases our score !!! as mentioned here - https://www.kaggle.com/c/bengaliai-cv19/discussion/126833 by Nihar. I am personally trying different model together for vowel and consonant and a separate model for graphemes",
    "723908": "I have also tried to reduce not the `val_loss` but the `grapheme_root_val_loss` which seemed to give some interesting results. I didn't necessarily measure everything correctly (which I need to start getting in the habit of doing though) but it might be worth investigating more in-depth.",
    "720908": "I have been seeing something similar. Maybe adding some more layers which are separate would help (I'm guessing you have a single fully connected layer for each output...)?\n![](https://i.imgur.com/iB1dxkk.png)",
    "720755": "why not have two separate models? \nOne for grapheme and other for the vowel and consonant? (I am pretty confident this would yield good results, but would increase submission time drastically)",
    "720721": ""
  }
}