{
  "id": 2451,
  "title": "possibility to use test data for training ?",
  "url": "/competitions/predict-closed-questions-on-stack-overflow/discussion/2451",
  "author_name": "",
  "post_date": "2012-08-26T22:42:54.997Z",
  "votes": null,
  "comment_count": 4,
  "views": 2216,
  "content": "<p>Hi,</p>\r\n<p>During the final evaluation: is the prediction model allowed to be trained using unlabeled test evaluation dataset ? for the moment it is possible to do so with the unlabeled leaderboard data, since they are available all at once, so I am curious about the\r\n final evaluation procedure :-)</p>\r\n<p>thanks!</p>",
  "messages": [
    {
      "id": "13485",
      "postDate": "08/26/2012 22:42:54",
      "content": "<p>Hi,</p>\r\n<p>During the final evaluation: is the prediction model allowed to be trained using unlabeled test evaluation dataset ? for the moment it is possible to do so with the unlabeled leaderboard data, since they are available all at once, so I am curious about the\r\n final evaluation procedure :-)</p>\r\n<p>thanks!</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "13489",
      "postDate": "08/27/2012 04:25:37",
      "content": "<p>[quote=Ordurio Benchmark;13485]</p>\r\n<p>Hi,</p>\r\n<p>During the final evaluation: is the prediction model allowed to be trained using unlabeled test evaluation dataset ? for the moment it is possible to do so with the unlabeled leaderboard data, since they are available all at once, so I am curious about the\r\n final evaluation procedure :-)</p>\r\n<p>thanks!</p>\r\n<p>[/quote]No. The training and prediction portions of the model should be separated, and the training portion should only used the training data (more closely emulating a production application of the model).</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "13499",
      "postDate": "08/27/2012 08:20:47",
      "content": "<p>Ok. But, for a given test example with timestamp t, can we still make use of all the posts that were gathered up to t, including those of the test dataset ? i am thinking in particular about the user features you propose (OwnerUndeletedAnswerCountAtPostTime).\r\n Are those features computed on every available data at the given post creation date, or only using training data ?</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "13507",
      "postDate": "08/27/2012 15:11:31",
      "content": "<p>Orduiro,</p>\r\n<p>It seems like your approach would have the affect of making your algorithm learn from data it won't have access to in a true production scenario.&nbsp; I think you could easily end up overfitting the seen data which would have the affect of reducing your real\r\n performance.</p>\r\n<p>&nbsp;</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "13510",
      "postDate": "08/27/2012 16:04:07",
      "content": "<p>&nbsp;</p>\r\n<p>i'm trying to be more precise this time :p</p>\r\n<p>- the feature &quot;OwnerUndeletedAnswerCountAtPostTime&quot; is computed on each post, train posts as well as test posts.</p>\r\n<p>- my question: for a post in the test dataset with creation date t, computing this feature requires all the posts up to time t, including a few ones in the test dataset. Am i correct ?</p>\r\n<p>- this is why i asked before whether we are allowed to build a model which possibly updates its internal state/resources each time it sees a new test data, provided it is tested with chronologically ordered posts to emulate real production timeline.</p>\r\n<p>Well, this question came to my mind reading another post ( <a href=\"http://www.kaggle.com/c/predict-closed-questions-on-stack-overflow/forums/t/2432/feature-ownerundeletedanswercountatposttime\">\r\nhttp://www.kaggle.com/c/predict-closed-questions-on-stack-overflow/forums/t/2432/feature-ownerundeletedanswercountatposttime</a>&nbsp;), Ben suggested to consider another related feature, the number of non closed questions created by the author up to time t, and\r\n the same question arises with this feature. If this feature is relevant, i would like to be able to compute it accurately on train data as well as on test data :-)</p>\r\n<p>best</p>",
      "rawMarkdown": "",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 13489,
      "author_name": "benhamner",
      "author_url": "",
      "post_date": "08/27/2012 04:25:37",
      "content": "<p>[quote=Ordurio Benchmark;13485]</p>\r\n<p>Hi,</p>\r\n<p>During the final evaluation: is the prediction model allowed to be trained using unlabeled test evaluation dataset ? for the moment it is possible to do so with the unlabeled leaderboard data, since they are available all at once, so I am curious about the\r\n final evaluation procedure :-)</p>\r\n<p>thanks!</p>\r\n<p>[/quote]No. The training and prediction portions of the model should be separated, and the training portion should only used the training data (more closely emulating a production application of the model).</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 13499,
      "author_name": "ordurio",
      "author_url": "",
      "post_date": "08/27/2012 08:20:47",
      "content": "<p>Ok. But, for a given test example with timestamp t, can we still make use of all the posts that were gathered up to t, including those of the test dataset ? i am thinking in particular about the user features you propose (OwnerUndeletedAnswerCountAtPostTime).\r\n Are those features computed on every available data at the given post creation date, or only using training data ?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 13507,
      "author_name": "mcstar",
      "author_url": "",
      "post_date": "08/27/2012 15:11:31",
      "content": "<p>Orduiro,</p>\r\n<p>It seems like your approach would have the affect of making your algorithm learn from data it won't have access to in a true production scenario.&nbsp; I think you could easily end up overfitting the seen data which would have the affect of reducing your real\r\n performance.</p>\r\n<p>&nbsp;</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 13510,
      "author_name": "ordurio",
      "author_url": "",
      "post_date": "08/27/2012 16:04:07",
      "content": "<p>&nbsp;</p>\r\n<p>i'm trying to be more precise this time :p</p>\r\n<p>- the feature &quot;OwnerUndeletedAnswerCountAtPostTime&quot; is computed on each post, train posts as well as test posts.</p>\r\n<p>- my question: for a post in the test dataset with creation date t, computing this feature requires all the posts up to time t, including a few ones in the test dataset. Am i correct ?</p>\r\n<p>- this is why i asked before whether we are allowed to build a model which possibly updates its internal state/resources each time it sees a new test data, provided it is tested with chronologically ordered posts to emulate real production timeline.</p>\r\n<p>Well, this question came to my mind reading another post ( <a href=\"http://www.kaggle.com/c/predict-closed-questions-on-stack-overflow/forums/t/2432/feature-ownerundeletedanswercountatposttime\">\r\nhttp://www.kaggle.com/c/predict-closed-questions-on-stack-overflow/forums/t/2432/feature-ownerundeletedanswercountatposttime</a>&nbsp;), Ben suggested to consider another related feature, the number of non closed questions created by the author up to time t, and\r\n the same question arises with this feature. If this feature is relevant, i would like to be able to compute it accurately on train data as well as on test data :-)</p>\r\n<p>best</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "13485": "",
    "13489": "",
    "13499": "",
    "13507": "",
    "13510": ""
  },
  "source": "meta"
}