{
  "id": 136106,
  "title": "Selecting submission",
  "url": "/competitions/bengaliai-cv19/discussion/136106",
  "author_name": "CPMP",
  "post_date": "2020-03-17T13:15:34.782000",
  "votes": 32,
  "comment_count": 3,
  "views": 0,
  "content": "<p>It is tempting to say: I could have got this far if only I selected this sub, and make a post about it.  We can all do it.  And yes, it is frustrating to see where we could have been.</p>\n\n<p>It happened to me, in one competition I had 8 subs that were better than the winner team model.  8.  I ended up around 10th rank if I remember well.</p>\n\n<p>So what?</p>\n\n<p>The ability to evaluate properly models before we know the private LB is key.  It is key in Kaggle because producing a great model but not selecting it is exactly as bad as not producing a great model at all.  It is also key in real world because you need to select which models you deploy in production, based on that data you have access to, i.e. training data.</p>\n\n<p>Therefore, when you see you did not select your best model, then ask yourself why you did not select it, and what you could have done to detect it was the best model.  </p>\n\n<p>Instead of complaining publicly, learn how to evaluate properly your models.  It will benefit your Kaggle ranking, but, more importantly, it will improve a lot the way you select your models in real life.</p>",
  "messages": [
    {
      "id": 776538,
      "postDate": "2020-03-17T13:15:34.783Z",
      "content": "<p>It is tempting to say: I could have got this far if only I selected this sub, and make a post about it.  We can all do it.  And yes, it is frustrating to see where we could have been.</p>\n\n<p>It happened to me, in one competition I had 8 subs that were better than the winner team model.  8.  I ended up around 10th rank if I remember well.</p>\n\n<p>So what?</p>\n\n<p>The ability to evaluate properly models before we know the private LB is key.  It is key in Kaggle because producing a great model but not selecting it is exactly as bad as not producing a great model at all.  It is also key in real world because you need to select which models you deploy in production, based on that data you have access to, i.e. training data.</p>\n\n<p>Therefore, when you see you did not select your best model, then ask yourself why you did not select it, and what you could have done to detect it was the best model.  </p>\n\n<p>Instead of complaining publicly, learn how to evaluate properly your models.  It will benefit your Kaggle ranking, but, more importantly, it will improve a lot the way you select your models in real life.</p>",
      "rawMarkdown": "It is tempting to say: I could have got this far if only I selected this sub, and make a post about it.  We can all do it.  And yes, it is frustrating to see where we could have been.\n\nIt happened to me, in one competition I had 8 subs that were better than the winner team model.  8.  I ended up around 10th rank if I remember well.\n\nSo what?\n\nThe ability to evaluate properly models before we know the private LB is key.  It is key in Kaggle because producing a great model but not selecting it is exactly as bad as not producing a great model at all.  It is also key in real world because you need to select which models you deploy in production, based on that data you have access to, i.e. training data.\n\nTherefore, when you see you did not select your best model, then ask yourself why you did not select it, and what you could have done to detect it was the best model.  \n\nInstead of complaining publicly, learn how to evaluate properly your models.  It will benefit your Kaggle ranking, but, more importantly, it will improve a lot the way you select your models in real life.",
      "votes": 32
    },
    {
      "id": 776588,
      "postDate": "2020-03-17T13:44:44.980Z",
      "content": "<p>People made a fatal flaw when complaining about how they selected the 'bad sub' and where they could've ended up: They were assuming no one else made the same mistake they did when, in fact , it's pretty common. \nDo you think you're still going up 1420 places if everyone submitted their private best solution?</p>",
      "rawMarkdown": "People made a fatal flaw when complaining about how they selected the 'bad sub' and where they could've ended up: They were assuming no one else made the same mistake they did when, in fact , it's pretty common. \nDo you think you're still going up 1420 places if everyone submitted their private best solution?",
      "votes": 5
    },
    {
      "id": 777654,
      "postDate": "2020-03-17T20:51:51.037Z",
      "content": "<p>I viewed few posts on selecting \"bad sub\" and solutions, they were not even having good score on public result at start.  The code was super simple, something that can be produced in few hours.  It's not like they were intentionally underfitting the model. They were unable to fit the model well on the training data. It's not like they were on the path to win on way or another.  It's just happened that the private set were drastic different than public set so that many people (including me) who were unable to fit well the training data got a huge jump at the end.  </p>\n\n<p>For me, reading through the good solutions, I realize that:\n- I did not train with a head with 1200+ grapheme classses (this is the major reason I could not even get 0.98+ on local CV, and I was so confused how people are getting 0.99+).\n- I did not use cutmix correctly (small alpha leading to small cutting boxes that are useless for training)\n- I did not think about doing any post processing \n- I did not even know that there were many unseen in the private dataset.  Hence I was not even thinking about distinguishing seen / unseen /  using GAN or other generative method to create synthetic data. </p>\n\n<p>I do think most of people who got a huge jump did not do any of the above, so there is absolutely nothing to complain on.</p>",
      "rawMarkdown": "I viewed few posts on selecting \"bad sub\" and solutions, they were not even having good score on public result at start.  The code was super simple, something that can be produced in few hours.  It's not like they were intentionally underfitting the model. They were unable to fit the model well on the training data. It's not like they were on the path to win on way or another.  It's just happened that the private set were drastic different than public set so that many people (including me) who were unable to fit well the training data got a huge jump at the end.  \n\nFor me, reading through the good solutions, I realize that:\n- I did not train with a head with 1200+ grapheme classses (this is the major reason I could not even get 0.98+ on local CV, and I was so confused how people are getting 0.99+).\n- I did not use cutmix correctly (small alpha leading to small cutting boxes that are useless for training)\n- I did not think about doing any post processing \n- I did not even know that there were many unseen in the private dataset.  Hence I was not even thinking about distinguishing seen / unseen /  using GAN or other generative method to create synthetic data. \n\nI do think most of people who got a huge jump did not do any of the above, so there is absolutely nothing to complain on.",
      "votes": 1
    },
    {
      "id": 776571,
      "postDate": "2020-03-17T13:33:48.887Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 776588,
      "author_name": "Khoi Nguyen",
      "author_url": "",
      "post_date": "2020-03-17T13:44:44.980000",
      "content": "<p>People made a fatal flaw when complaining about how they selected the 'bad sub' and where they could've ended up: They were assuming no one else made the same mistake they did when, in fact , it's pretty common. \nDo you think you're still going up 1420 places if everyone submitted their private best solution?</p>",
      "votes": 5,
      "replies": []
    },
    {
      "id": 777654,
      "author_name": "sandiudiu",
      "author_url": "",
      "post_date": "2020-03-17T20:51:51.037000",
      "content": "<p>I viewed few posts on selecting \"bad sub\" and solutions, they were not even having good score on public result at start.  The code was super simple, something that can be produced in few hours.  It's not like they were intentionally underfitting the model. They were unable to fit the model well on the training data. It's not like they were on the path to win on way or another.  It's just happened that the private set were drastic different than public set so that many people (including me) who were unable to fit well the training data got a huge jump at the end.  </p>\n\n<p>For me, reading through the good solutions, I realize that:\n- I did not train with a head with 1200+ grapheme classses (this is the major reason I could not even get 0.98+ on local CV, and I was so confused how people are getting 0.99+).\n- I did not use cutmix correctly (small alpha leading to small cutting boxes that are useless for training)\n- I did not think about doing any post processing \n- I did not even know that there were many unseen in the private dataset.  Hence I was not even thinking about distinguishing seen / unseen /  using GAN or other generative method to create synthetic data. </p>\n\n<p>I do think most of people who got a huge jump did not do any of the above, so there is absolutely nothing to complain on.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 776571,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-03-17T13:33:48.887000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "776538": "It is tempting to say: I could have got this far if only I selected this sub, and make a post about it.  We can all do it.  And yes, it is frustrating to see where we could have been.\n\nIt happened to me, in one competition I had 8 subs that were better than the winner team model.  8.  I ended up around 10th rank if I remember well.\n\nSo what?\n\nThe ability to evaluate properly models before we know the private LB is key.  It is key in Kaggle because producing a great model but not selecting it is exactly as bad as not producing a great model at all.  It is also key in real world because you need to select which models you deploy in production, based on that data you have access to, i.e. training data.\n\nTherefore, when you see you did not select your best model, then ask yourself why you did not select it, and what you could have done to detect it was the best model.  \n\nInstead of complaining publicly, learn how to evaluate properly your models.  It will benefit your Kaggle ranking, but, more importantly, it will improve a lot the way you select your models in real life.",
    "776588": "People made a fatal flaw when complaining about how they selected the 'bad sub' and where they could've ended up: They were assuming no one else made the same mistake they did when, in fact , it's pretty common. \nDo you think you're still going up 1420 places if everyone submitted their private best solution?",
    "777654": "I viewed few posts on selecting \"bad sub\" and solutions, they were not even having good score on public result at start.  The code was super simple, something that can be produced in few hours.  It's not like they were intentionally underfitting the model. They were unable to fit the model well on the training data. It's not like they were on the path to win on way or another.  It's just happened that the private set were drastic different than public set so that many people (including me) who were unable to fit well the training data got a huge jump at the end.  \n\nFor me, reading through the good solutions, I realize that:\n- I did not train with a head with 1200+ grapheme classses (this is the major reason I could not even get 0.98+ on local CV, and I was so confused how people are getting 0.99+).\n- I did not use cutmix correctly (small alpha leading to small cutting boxes that are useless for training)\n- I did not think about doing any post processing \n- I did not even know that there were many unseen in the private dataset.  Hence I was not even thinking about distinguishing seen / unseen /  using GAN or other generative method to create synthetic data. \n\nI do think most of people who got a huge jump did not do any of the above, so there is absolutely nothing to complain on.",
    "776571": ""
  }
}