{
  "id": 70247,
  "title": "Just Do It!",
  "url": "/competitions/PLAsTiCC-2018/discussion/70247",
  "author_name": "",
  "post_date": "2018-11-01T09:24:42.780976300Z",
  "votes": 22,
  "comment_count": 4,
  "views": 0,
  "content": "<p>I see several questions on the forum from people who have not submitted yet but are asking if their local validation score is good.</p>\n\n<p>There is only one way to know: submit and see how it goes on the LB.  The sooner you submit, the sooner you get feedback on what you do.  Similarly, if you wonder about a new feature, or a new twist in your NN architecture, just do it, run local validation, and submit.</p>\n\n<p>What matters most is not so much the first score you will get. What matters most is if your local validation score moves the same way as the LB.  If that's so then you can focus on local validation and only submit from time to time.</p>\n\n<p>Sure, everyone dreams of having a single sub that appears at the top of the LB,  but there is no bonus on Kaggle for those having few submissions.   The only thing that matters is how your model will score on the private LB.  </p>\n\n<p>To be honest, the ability to select a great model and only use one sub is valuable.  It is close to what is  required in real world, where you often have only one model to deploy in production.  But that's one area where Kaggle differs from real world machine learning...</p>",
  "messages": [
    {
      "id": "413637",
      "postDate": "11/01/2018 09:24:42",
      "content": "<p>I see several questions on the forum from people who have not submitted yet but are asking if their local validation score is good.</p>\n\n<p>There is only one way to know: submit and see how it goes on the LB.  The sooner you submit, the sooner you get feedback on what you do.  Similarly, if you wonder about a new feature, or a new twist in your NN architecture, just do it, run local validation, and submit.</p>\n\n<p>What matters most is not so much the first score you will get. What matters most is if your local validation score moves the same way as the LB.  If that's so then you can focus on local validation and only submit from time to time.</p>\n\n<p>Sure, everyone dreams of having a single sub that appears at the top of the LB,  but there is no bonus on Kaggle for those having few submissions.   The only thing that matters is how your model will score on the private LB.  </p>\n\n<p>To be honest, the ability to select a great model and only use one sub is valuable.  It is close to what is  required in real world, where you often have only one model to deploy in production.  But that's one area where Kaggle differs from real world machine learning...</p>",
      "rawMarkdown": "I see several questions on the forum from people who have not submitted yet but are asking if their local validation score is good.\n\nThere is only one way to know: submit and see how it goes on the LB.  The sooner you submit, the sooner you get feedback on what you do.  Similarly, if you wonder about a new feature, or a new twist in your NN architecture, just do it, run local validation, and submit.\n\nWhat matters most is not so much the first score you will get. What matters most is if your local validation score moves the same way as the LB.  If that's so then you can focus on local validation and only submit from time to time.\n\nSure, everyone dreams of having a single sub that appears at the top of the LB,  but there is no bonus on Kaggle for those having few submissions.   The only thing that matters is how your model will score on the private LB.  \n\nTo be honest, the ability to select a great model and only use one sub is valuable.  It is close to what is  required in real world, where you often have only one model to deploy in production.  But that's one area where Kaggle differs from real world machine learning...",
      "votes": null
    },
    {
      "id": "413721",
      "postDate": "11/01/2018 12:18:28",
      "content": "<blockquote>\n  <p>I see several questions on the forum from people who have not submitted yet but are asking if their local validation score is good.</p>\n</blockquote>\n\n<p>I wonder how one can answer those questions without having a clue about the validation strategy ;)</p>\n\n<p>Anyway, public LB can be considered as extra-fold, thus submission can be part of validation strategy</p>",
      "rawMarkdown": "&gt; I see several questions on the forum from people who have not submitted yet but are asking if their local validation score is good.\n\nI wonder how one can answer those questions without having a clue about the validation strategy ;)\n\n\nAnyway, public LB can be considered as extra-fold, thus submission can be part of validation strategy",
      "votes": null
    },
    {
      "id": "413735",
      "postDate": "11/01/2018 12:55:13",
      "content": "<p>In general I always advise to trust CV and ignore LB, but here it is different as train is not representative of test.  First, class_99 is not present in train.  Second, nothing says the distribution of known classes is similar in train and test.  LB probing will probably play a greater role here than in many competitions ;)</p>",
      "rawMarkdown": "In general I always advise to trust CV and ignore LB, but here it is different as train is not representative of test.  First, class_99 is not present in train.  Second, nothing says the distribution of known classes is similar in train and test.  LB probing will probably play a greater role here than in many competitions ;)",
      "votes": null
    },
    {
      "id": "414114",
      "postDate": "11/02/2018 07:03:40",
      "content": "<p>I strongly agree with your statement. There are no shorts cut in kaggle.</p>",
      "rawMarkdown": "I strongly agree with your statement. There are no shorts cut in kaggle.",
      "votes": null
    },
    {
      "id": "414472",
      "postDate": "11/02/2018 20:34:52",
      "content": "<p>Thank you.</p>",
      "rawMarkdown": "Thank you.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 413721,
      "author_name": "serigne",
      "author_url": "",
      "post_date": "11/01/2018 12:18:28",
      "content": "<blockquote>\n  <p>I see several questions on the forum from people who have not submitted yet but are asking if their local validation score is good.</p>\n</blockquote>\n\n<p>I wonder how one can answer those questions without having a clue about the validation strategy ;)</p>\n\n<p>Anyway, public LB can be considered as extra-fold, thus submission can be part of validation strategy</p>",
      "votes": null,
      "replies": [
        {
          "id": 413735,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "11/01/2018 12:55:13",
          "content": "<p>In general I always advise to trust CV and ignore LB, but here it is different as train is not representative of test.  First, class_99 is not present in train.  Second, nothing says the distribution of known classes is similar in train and test.  LB probing will probably play a greater role here than in many competitions ;)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 414114,
      "author_name": "massy103",
      "author_url": "",
      "post_date": "11/02/2018 07:03:40",
      "content": "<p>I strongly agree with your statement. There are no shorts cut in kaggle.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 414472,
      "author_name": "bycati71",
      "author_url": "",
      "post_date": "11/02/2018 20:34:52",
      "content": "<p>Thank you.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "413637": "I see several questions on the forum from people who have not submitted yet but are asking if their local validation score is good.\n\nThere is only one way to know: submit and see how it goes on the LB.  The sooner you submit, the sooner you get feedback on what you do.  Similarly, if you wonder about a new feature, or a new twist in your NN architecture, just do it, run local validation, and submit.\n\nWhat matters most is not so much the first score you will get. What matters most is if your local validation score moves the same way as the LB.  If that's so then you can focus on local validation and only submit from time to time.\n\nSure, everyone dreams of having a single sub that appears at the top of the LB,  but there is no bonus on Kaggle for those having few submissions.   The only thing that matters is how your model will score on the private LB.  \n\nTo be honest, the ability to select a great model and only use one sub is valuable.  It is close to what is  required in real world, where you often have only one model to deploy in production.  But that's one area where Kaggle differs from real world machine learning...",
    "413721": "&gt; I see several questions on the forum from people who have not submitted yet but are asking if their local validation score is good.\n\nI wonder how one can answer those questions without having a clue about the validation strategy ;)\n\n\nAnyway, public LB can be considered as extra-fold, thus submission can be part of validation strategy",
    "413735": "In general I always advise to trust CV and ignore LB, but here it is different as train is not representative of test.  First, class_99 is not present in train.  Second, nothing says the distribution of known classes is similar in train and test.  LB probing will probably play a greater role here than in many competitions ;)",
    "414114": "I strongly agree with your statement. There are no shorts cut in kaggle.",
    "414472": "Thank you."
  },
  "source": "meta"
}