{
  "id": 66731,
  "title": "Auto evaluation of model",
  "url": "/competitions/rsna-pneumonia-detection-challenge/discussion/66731",
  "author_name": "",
  "post_date": "2018-09-24T19:35:03.861920300Z",
  "votes": null,
  "comment_count": 1,
  "views": 0,
  "content": "<p>Hello all,\nThis is my first competition on Kaggle and I have a question. It seems that there is no way for me to evaluate my model before submitting, since true outcomes are not provided. </p>\n\n<p>Alex EBE.  </p>",
  "messages": [
    {
      "id": "393068",
      "postDate": "09/24/2018 19:35:03",
      "content": "<p>Hello all,\nThis is my first competition on Kaggle and I have a question. It seems that there is no way for me to evaluate my model before submitting, since true outcomes are not provided. </p>\n\n<p>Alex EBE.  </p>",
      "rawMarkdown": "Hello all,\nThis is my first competition on Kaggle and I have a question. It seems that there is no way for me to evaluate my model before submitting, since true outcomes are not provided. \n\nAlex EBE.",
      "votes": null
    },
    {
      "id": "393100",
      "postDate": "09/24/2018 20:23:24",
      "content": "<p>Typically you would split the \"train\" data into a \"training\" set and a \"validation\" set. You can use 10-20% of the data for the validation set.</p>\n\n<p>Then you can \"score\" the validation set yourself to see how you are doing.</p>\n\n<p>There are kernels that have example programs that will score your \"validation\" set. There is no \"official\" scoring program that you can apply. You have to adapt one of the kernels or write your own scoring algorithm.</p>\n\n<p>Since your \"validation\" set might not always be statistically representative of your remaining data (or the test data), one would typically run an algorithm approximately 5 times with different validation sets to get an \"average\".</p>\n\n<p>Other threads have discussed whether the \"test\" data is similar to the \"training\" data. Many find their \"validation\" score doesn't match their \"test\" score on the leaderboard.</p>",
      "rawMarkdown": "Typically you would split the \"train\" data into a \"training\" set and a \"validation\" set. You can use 10-20% of the data for the validation set.\n\nThen you can \"score\" the validation set yourself to see how you are doing.\n\nThere are kernels that have example programs that will score your \"validation\" set. There is no \"official\" scoring program that you can apply. You have to adapt one of the kernels or write your own scoring algorithm.\n\nSince your \"validation\" set might not always be statistically representative of your remaining data (or the test data), one would typically run an algorithm approximately 5 times with different validation sets to get an \"average\".\n\nOther threads have discussed whether the \"test\" data is similar to the \"training\" data. Many find their \"validation\" score doesn't match their \"test\" score on the leaderboard.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 393100,
      "author_name": "richardepstein",
      "author_url": "",
      "post_date": "09/24/2018 20:23:24",
      "content": "<p>Typically you would split the \"train\" data into a \"training\" set and a \"validation\" set. You can use 10-20% of the data for the validation set.</p>\n\n<p>Then you can \"score\" the validation set yourself to see how you are doing.</p>\n\n<p>There are kernels that have example programs that will score your \"validation\" set. There is no \"official\" scoring program that you can apply. You have to adapt one of the kernels or write your own scoring algorithm.</p>\n\n<p>Since your \"validation\" set might not always be statistically representative of your remaining data (or the test data), one would typically run an algorithm approximately 5 times with different validation sets to get an \"average\".</p>\n\n<p>Other threads have discussed whether the \"test\" data is similar to the \"training\" data. Many find their \"validation\" score doesn't match their \"test\" score on the leaderboard.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "393068": "Hello all,\nThis is my first competition on Kaggle and I have a question. It seems that there is no way for me to evaluate my model before submitting, since true outcomes are not provided. \n\nAlex EBE.",
    "393100": "Typically you would split the \"train\" data into a \"training\" set and a \"validation\" set. You can use 10-20% of the data for the validation set.\n\nThen you can \"score\" the validation set yourself to see how you are doing.\n\nThere are kernels that have example programs that will score your \"validation\" set. There is no \"official\" scoring program that you can apply. You have to adapt one of the kernels or write your own scoring algorithm.\n\nSince your \"validation\" set might not always be statistically representative of your remaining data (or the test data), one would typically run an algorithm approximately 5 times with different validation sets to get an \"average\".\n\nOther threads have discussed whether the \"test\" data is similar to the \"training\" data. Many find their \"validation\" score doesn't match their \"test\" score on the leaderboard."
  },
  "source": "meta"
}