{
  "id": 47021,
  "title": "Difference of private and public LB test set: beware of shakeup?",
  "url": "/competitions/tensorflow-speech-recognition-challenge/discussion/47021",
  "author_name": "",
  "post_date": "2018-01-07T01:40:20.730210Z",
  "votes": 10,
  "comment_count": 1,
  "views": 0,
  "content": "<p>Note that the LB unknown samples may be different for  private and public LB test set. We do not know. </p>\n\n<p>While trying to improve accuracy, one should also ensure that the models are \"robust\" and the results are consistent.</p>\n\n<p>It is important to do some error analysis to see where the error comes from and \"probe\"  the LB dataset to \"guess\" the underlying distribution.</p>\n\n<p>To avoid shakeup:</p>\n\n<ol>\n<li>one should measure the \"sensitivity\" of the models  to different unknown samples (e.g. there are many  unknown samples in the validation set. Divide them to N subset by words to measure the accuracy of your models to give accuracy_of_unknown = mean +/- std</li>\n<li>Try to model the distribution of the unknown samples in the  LB set, i.e what words are present in the LB set.</li>\n</ol>\n\n<p>Any other suggestions and comments are welcome.</p>\n\n<p>An example of running simulation to estimate shakeup is at: <a href=\"http://blog.kaggle.com/2017/10/17/planet-understanding-the-amazon-from-space-1st-place-winners-interview/\">http://blog.kaggle.com/2017/10/17/planet-understanding-the-amazon-from-space-1st-place-winners-interview/</a></p>",
  "messages": [
    {
      "id": "265883",
      "postDate": "01/07/2018 01:40:20",
      "content": "<p>Note that the LB unknown samples may be different for  private and public LB test set. We do not know. </p>\n\n<p>While trying to improve accuracy, one should also ensure that the models are \"robust\" and the results are consistent.</p>\n\n<p>It is important to do some error analysis to see where the error comes from and \"probe\"  the LB dataset to \"guess\" the underlying distribution.</p>\n\n<p>To avoid shakeup:</p>\n\n<ol>\n<li>one should measure the \"sensitivity\" of the models  to different unknown samples (e.g. there are many  unknown samples in the validation set. Divide them to N subset by words to measure the accuracy of your models to give accuracy_of_unknown = mean +/- std</li>\n<li>Try to model the distribution of the unknown samples in the  LB set, i.e what words are present in the LB set.</li>\n</ol>\n\n<p>Any other suggestions and comments are welcome.</p>\n\n<p>An example of running simulation to estimate shakeup is at: <a href=\"http://blog.kaggle.com/2017/10/17/planet-understanding-the-amazon-from-space-1st-place-winners-interview/\">http://blog.kaggle.com/2017/10/17/planet-understanding-the-amazon-from-space-1st-place-winners-interview/</a></p>",
      "rawMarkdown": "Note that the LB unknown samples may be different for  private and public LB test set. We do not know. \n\nWhile trying to improve accuracy, one should also ensure that the models are \"robust\" and the results are consistent.\n\nIt is important to do some error analysis to see where the error comes from and \"probe\"  the LB dataset to \"guess\" the underlying distribution.\n\n\nTo avoid shakeup:\n\n 1. one should measure the \"sensitivity\" of the models  to different unknown samples (e.g. there are many  unknown samples in the validation set. Divide them to N subset by words to measure the accuracy of your models to give accuracy_of_unknown = mean +/- std\n 2. Try to model the distribution of the unknown samples in the  LB set, i.e what words are present in the LB set.\n\nAny other suggestions and comments are welcome.\n\nAn example of running simulation to estimate shakeup is at: http://blog.kaggle.com/2017/10/17/planet-understanding-the-amazon-from-space-1st-place-winners-interview/",
      "votes": null
    },
    {
      "id": "266427",
      "postDate": "01/08/2018 18:27:58",
      "content": "<p>Sure,it is always important to analyse  the distributions of the test set predictions of/among  models,by doing so ,we can know the strength and weakness of our models,then we can not only control overfitting but also find some ensemble methods.</p>",
      "rawMarkdown": "Sure,it is always important to analyse  the distributions of the test set predictions of/among  models,by doing so ,we can know the strength and weakness of our models,then we can not only control overfitting but also find some ensemble methods.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 266427,
      "author_name": "bestfitting",
      "author_url": "",
      "post_date": "01/08/2018 18:27:58",
      "content": "<p>Sure,it is always important to analyse  the distributions of the test set predictions of/among  models,by doing so ,we can know the strength and weakness of our models,then we can not only control overfitting but also find some ensemble methods.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "265883": "Note that the LB unknown samples may be different for  private and public LB test set. We do not know. \n\nWhile trying to improve accuracy, one should also ensure that the models are \"robust\" and the results are consistent.\n\nIt is important to do some error analysis to see where the error comes from and \"probe\"  the LB dataset to \"guess\" the underlying distribution.\n\n\nTo avoid shakeup:\n\n 1. one should measure the \"sensitivity\" of the models  to different unknown samples (e.g. there are many  unknown samples in the validation set. Divide them to N subset by words to measure the accuracy of your models to give accuracy_of_unknown = mean +/- std\n 2. Try to model the distribution of the unknown samples in the  LB set, i.e what words are present in the LB set.\n\nAny other suggestions and comments are welcome.\n\nAn example of running simulation to estimate shakeup is at: http://blog.kaggle.com/2017/10/17/planet-understanding-the-amazon-from-space-1st-place-winners-interview/",
    "266427": "Sure,it is always important to analyse  the distributions of the test set predictions of/among  models,by doing so ,we can know the strength and weakness of our models,then we can not only control overfitting but also find some ensemble methods."
  },
  "source": "meta"
}