{
  "id": 69481,
  "title": "What is your validation scheme and local score?",
  "url": "/competitions/PLAsTiCC-2018/discussion/69481",
  "author_name": "",
  "post_date": "2018-10-24T05:11:11.471782100Z",
  "votes": 2,
  "comment_count": 10,
  "views": 0,
  "content": "<p>I am using StratifiedKFold (5fold) with reference to olivier's kernel. (thank you, olivier)</p>\n\n<p>In the kernel(<a href=\"https://www.kaggle.com/ogrellier/plasticc-in-a-kernel-meta-and-data/code\">https://www.kaggle.com/ogrellier/plasticc-in-a-kernel-meta-and-data/code</a>), \nthe local score is 0.83 and LB is 1.71. (My local score is 0.58, but I have not finished predicting for \n the test set. I'm not dealing with the problem of  '15 th class'))</p>\n\n<p>I think that there is a big gap between CV and LB.\nI would like to know how others build a local validation scheme and local score.</p>\n\n<hr>\n\n<p>'15 th class' means class 99</p>",
  "messages": [
    {
      "id": "409307",
      "postDate": "10/24/2018 05:11:11",
      "content": "<p>I am using StratifiedKFold (5fold) with reference to olivier's kernel. (thank you, olivier)</p>\n\n<p>In the kernel(<a href=\"https://www.kaggle.com/ogrellier/plasticc-in-a-kernel-meta-and-data/code\">https://www.kaggle.com/ogrellier/plasticc-in-a-kernel-meta-and-data/code</a>), \nthe local score is 0.83 and LB is 1.71. (My local score is 0.58, but I have not finished predicting for \n the test set. I'm not dealing with the problem of  '15 th class'))</p>\n\n<p>I think that there is a big gap between CV and LB.\nI would like to know how others build a local validation scheme and local score.</p>\n\n<hr>\n\n<p>'15 th class' means class 99</p>",
      "rawMarkdown": "I am using StratifiedKFold (5fold) with reference to olivier's kernel. (thank you, olivier)\n\n\n\nIn the kernel(https://www.kaggle.com/ogrellier/plasticc-in-a-kernel-meta-and-data/code), \nthe local score is 0.83 and LB is 1.71. (My local score is 0.58, but I have not finished predicting for \n the test set. I'm not dealing with the problem of  '15 th class'))\n\n\n\nI think that there is a big gap between CV and LB.\nI would like to know how others build a local validation scheme and local score.\n\n\n\n\n\n----------\n'15 th class' means class 99",
      "votes": null
    },
    {
      "id": "409326",
      "postDate": "10/24/2018 06:04:09",
      "content": "<p>Class 99 can explain a lot of the difference between local cv and public lb.</p>",
      "rawMarkdown": "Class 99 can explain a lot of the difference between local cv and public lb.",
      "votes": null
    },
    {
      "id": "409573",
      "postDate": "10/24/2018 14:04:55",
      "content": "<p>0.71 CV with constant class99 prediction -&gt; LB 1.08</p>",
      "rawMarkdown": "0.71 CV with constant class99 prediction -&gt; LB 1.08",
      "votes": null
    },
    {
      "id": "409588",
      "postDate": "10/24/2018 14:40:44",
      "content": "<p>How did you select the constant?  LB probing?</p>",
      "rawMarkdown": "How did you select the constant?  LB probing?",
      "votes": null
    },
    {
      "id": "409591",
      "postDate": "10/24/2018 14:44:21",
      "content": "<p>I suspect Class 99 is what's causing the issue. How does one reflect class 99 error in the training set when it doesn't exist?</p>",
      "rawMarkdown": "I suspect Class 99 is what's causing the issue. How does one reflect class 99 error in the training set when it doesn't exist?",
      "votes": null
    },
    {
      "id": "409627",
      "postDate": "10/24/2018 16:02:37",
      "content": "<p>0.86 CV and 1.378 LB. I did not prob class 99 and predicting it as it's done in Olivier's kernel. Also, I have always dropped hostgal_specz column for my training even though it improves CV a lot. But as you know, only 4% of test set has this value. I wonder, is there anyone tried submitting just with and without the hostgal_specz included model? Does it still improve LB even with this percentage? I'm so lazy to run submission generation for just this difference again :)</p>",
      "rawMarkdown": "0.86 CV and 1.378 LB. I did not prob class 99 and predicting it as it's done in Olivier's kernel. Also, I have always dropped hostgal_specz column for my training even though it improves CV a lot. But as you know, only 4% of test set has this value. I wonder, is there anyone tried submitting just with and without the hostgal_specz included model? Does it still improve LB even with this percentage? I'm so lazy to run submission generation for just this difference again :)",
      "votes": null
    },
    {
      "id": "409696",
      "postDate": "10/24/2018 17:58:05",
      "content": "<p>Thank you, CPMP. I think so, too.</p>\n\n<p>Should I work on improving the local validation score with simple StratifiedKFold? \nDo you think it will contribute to LB?(of course, I also need a scheme that addresses \"class 99\" is necessary in addition to StratifiedKFold.)</p>",
      "rawMarkdown": "Thank you, CPMP. I think so, too.\n\nShould I work on improving the local validation score with simple StratifiedKFold? \nDo you think it will contribute to LB?(of course, I also need a scheme that addresses \"class 99\" is necessary in addition to StratifiedKFold.)",
      "votes": null
    },
    {
      "id": "409697",
      "postDate": "10/24/2018 17:59:01",
      "content": "<p>I am considering the problem separately.</p>\n\n<ol>\n<li>Problems predicting classes 14</li>\n<li>Problems predicting class 99</li>\n</ol>\n\n<p>As for 1, I'm using StratifiedKFold.\nAs for 2, I have no idea other than just to apply Anomaly Detection to the test set.</p>",
      "rawMarkdown": "I am considering the problem separately.\n\n 1. Problems predicting classes 14\n 2. Problems predicting class 99\n\nAs for 1, I'm using StratifiedKFold.\nAs for 2, I have no idea other than just to apply Anomaly Detection to the test set.",
      "votes": null
    },
    {
      "id": "409698",
      "postDate": "10/24/2018 17:59:31",
      "content": "<p>Oh, it certainly seems very important to dropping hostgal_specz. thank you :)</p>",
      "rawMarkdown": "Oh, it certainly seems very important to dropping hostgal_specz. thank you :)",
      "votes": null
    },
    {
      "id": "409699",
      "postDate": "10/24/2018 17:59:42",
      "content": "<p>My gut feeling is that the better your model, the better.  How to handle class 99 is a bit orthogonal to predicting which of the 14 classes of train is most likely.</p>",
      "rawMarkdown": "My gut feeling is that the better your model, the better.  How to handle class 99 is a bit orthogonal to predicting which of the 14 classes of train is most likely.",
      "votes": null
    },
    {
      "id": "409704",
      "postDate": "10/24/2018 18:11:47",
      "content": "<p>thx :)</p>",
      "rawMarkdown": "thx :)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 409326,
      "author_name": "cpmpml",
      "author_url": "",
      "post_date": "10/24/2018 06:04:09",
      "content": "<p>Class 99 can explain a lot of the difference between local cv and public lb.</p>",
      "votes": null,
      "replies": [
        {
          "id": 409696,
          "author_name": "sugawarya",
          "author_url": "",
          "post_date": "10/24/2018 17:58:05",
          "content": "<p>Thank you, CPMP. I think so, too.</p>\n\n<p>Should I work on improving the local validation score with simple StratifiedKFold? \nDo you think it will contribute to LB?(of course, I also need a scheme that addresses \"class 99\" is necessary in addition to StratifiedKFold.)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 409699,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "10/24/2018 17:59:42",
          "content": "<p>My gut feeling is that the better your model, the better.  How to handle class 99 is a bit orthogonal to predicting which of the 14 classes of train is most likely.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 409704,
          "author_name": "sugawarya",
          "author_url": "",
          "post_date": "10/24/2018 18:11:47",
          "content": "<p>thx :)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 409573,
      "author_name": "aerdem4",
      "author_url": "",
      "post_date": "10/24/2018 14:04:55",
      "content": "<p>0.71 CV with constant class99 prediction -&gt; LB 1.08</p>",
      "votes": null,
      "replies": [
        {
          "id": 409588,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "10/24/2018 14:40:44",
          "content": "<p>How did you select the constant?  LB probing?</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 409591,
      "author_name": "nnnnick",
      "author_url": "",
      "post_date": "10/24/2018 14:44:21",
      "content": "<p>I suspect Class 99 is what's causing the issue. How does one reflect class 99 error in the training set when it doesn't exist?</p>",
      "votes": null,
      "replies": [
        {
          "id": 409697,
          "author_name": "sugawarya",
          "author_url": "",
          "post_date": "10/24/2018 17:59:01",
          "content": "<p>I am considering the problem separately.</p>\n\n<ol>\n<li>Problems predicting classes 14</li>\n<li>Problems predicting class 99</li>\n</ol>\n\n<p>As for 1, I'm using StratifiedKFold.\nAs for 2, I have no idea other than just to apply Anomaly Detection to the test set.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 409627,
      "author_name": "fatihozturk",
      "author_url": "",
      "post_date": "10/24/2018 16:02:37",
      "content": "<p>0.86 CV and 1.378 LB. I did not prob class 99 and predicting it as it's done in Olivier's kernel. Also, I have always dropped hostgal_specz column for my training even though it improves CV a lot. But as you know, only 4% of test set has this value. I wonder, is there anyone tried submitting just with and without the hostgal_specz included model? Does it still improve LB even with this percentage? I'm so lazy to run submission generation for just this difference again :)</p>",
      "votes": null,
      "replies": [
        {
          "id": 409698,
          "author_name": "sugawarya",
          "author_url": "",
          "post_date": "10/24/2018 17:59:31",
          "content": "<p>Oh, it certainly seems very important to dropping hostgal_specz. thank you :)</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "409307": "I am using StratifiedKFold (5fold) with reference to olivier's kernel. (thank you, olivier)\n\n\n\nIn the kernel(https://www.kaggle.com/ogrellier/plasticc-in-a-kernel-meta-and-data/code), \nthe local score is 0.83 and LB is 1.71. (My local score is 0.58, but I have not finished predicting for \n the test set. I'm not dealing with the problem of  '15 th class'))\n\n\n\nI think that there is a big gap between CV and LB.\nI would like to know how others build a local validation scheme and local score.\n\n\n\n\n\n----------\n'15 th class' means class 99",
    "409326": "Class 99 can explain a lot of the difference between local cv and public lb.",
    "409573": "0.71 CV with constant class99 prediction -&gt; LB 1.08",
    "409588": "How did you select the constant?  LB probing?",
    "409591": "I suspect Class 99 is what's causing the issue. How does one reflect class 99 error in the training set when it doesn't exist?",
    "409627": "0.86 CV and 1.378 LB. I did not prob class 99 and predicting it as it's done in Olivier's kernel. Also, I have always dropped hostgal_specz column for my training even though it improves CV a lot. But as you know, only 4% of test set has this value. I wonder, is there anyone tried submitting just with and without the hostgal_specz included model? Does it still improve LB even with this percentage? I'm so lazy to run submission generation for just this difference again :)",
    "409696": "Thank you, CPMP. I think so, too.\n\nShould I work on improving the local validation score with simple StratifiedKFold? \nDo you think it will contribute to LB?(of course, I also need a scheme that addresses \"class 99\" is necessary in addition to StratifiedKFold.)",
    "409697": "I am considering the problem separately.\n\n 1. Problems predicting classes 14\n 2. Problems predicting class 99\n\nAs for 1, I'm using StratifiedKFold.\nAs for 2, I have no idea other than just to apply Anomaly Detection to the test set.",
    "409698": "Oh, it certainly seems very important to dropping hostgal_specz. thank you :)",
    "409699": "My gut feeling is that the better your model, the better.  How to handle class 99 is a bit orthogonal to predicting which of the 14 classes of train is most likely.",
    "409704": "thx :)"
  },
  "source": "meta"
}