{
  "id": 1438,
  "title": "A problem about test data",
  "url": "/competitions/kddcup2012-track2/discussion/1438",
  "author_name": "",
  "post_date": "2012-03-03T04:56:45.330Z",
  "votes": null,
  "comment_count": 8,
  "views": 6254,
  "content": "<p>Can anybody tell me that whether the target users, queries, and ads in the test data set for CTR prediction are also involved in the training data set? That is, are we required to predict CTR&nbsp;among users, queries, and ads that have already appeared in the\r\n training data set, or are we required to predict CTR between NEW users, NEW queries and NEW ads? This is important for the design of the model I think. Thanks</p>",
  "messages": [
    {
      "id": "8912",
      "postDate": "03/03/2012 04:56:45",
      "content": "<p>Can anybody tell me that whether the target users, queries, and ads in the test data set for CTR prediction are also involved in the training data set? That is, are we required to predict CTR&nbsp;among users, queries, and ads that have already appeared in the\r\n training data set, or are we required to predict CTR between NEW users, NEW queries and NEW ads? This is important for the design of the model I think. Thanks</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "8924",
      "postDate": "03/03/2012 12:10:24",
      "content": "<p>It is not guaranteed that users (or queries or ads) in the test data set appear in the training data set. I am not sure if there is a test instance that have new user, new query and new ad. However, even if it is true, we can find similarities between an\r\n old user and a new user (or query, or ad) if we use the various _tokensid.txt files.\r\n</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "8933",
      "postDate": "03/03/2012 16:58:08",
      "content": "<p>I think the key question is that how the data were sampled and splited into training set and testing set. It does count! And we may simulate this sample approach to generate local test set.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "8968",
      "postDate": "03/05/2012 01:46:36",
      "content": "<p>There isn't a sample approach. The data were divided by date. The whole data set is a sample of logs in 55 days, where the part of the first 50 days is the training data set, and the part of the rest 5 days is the testing data set.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "8972",
      "postDate": "03/05/2012 02:10:49",
      "content": "<p>Well, it makes sense. However, there is no timestamp information given for us. Shall we assume that the training instance are sorted by date? Thanks!</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "8973",
      "postDate": "03/05/2012 03:01:43",
      "content": "<pre>No, we cannot assume a temporal order, as stated in the Description page: &quot;We aggregate instances with the same user id, ad id, query, and setting in order to reduce the dataset size. &quot;  In other words, each training/testing instance comes from one or more session log messages.</pre>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "8974",
      "postDate": "03/05/2012 03:05:24",
      "content": "<p>Could you give an example of how a test instance would look like?</p>\r\n<p>The same as training except for the impression and click cols?</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "8981",
      "postDate": "03/05/2012 04:19:35",
      "content": "<p>Yes, as you said, same as training instances except for the impression and click columns.</p>\r\n<p>Soon, we will publish a validation data set, which has exactly the same format as the testing data set.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "9024",
      "postDate": "03/05/2012 11:21:13",
      "content": "<p>Did you mean the training data was produced by aggregating the first 50 days' log, and the testing data was produced by aggregating the last 5 days' log? So we would expect statistically identical distribution of users, querys, and ads?</p>",
      "rawMarkdown": "",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 8924,
      "author_name": "yiwang0",
      "author_url": "",
      "post_date": "03/03/2012 12:10:24",
      "content": "<p>It is not guaranteed that users (or queries or ads) in the test data set appear in the training data set. I am not sure if there is a test instance that have new user, new query and new ad. However, even if it is true, we can find similarities between an\r\n old user and a new user (or query, or ad) if we use the various _tokensid.txt files.\r\n</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 8933,
      "author_name": "goldensection",
      "author_url": "",
      "post_date": "03/03/2012 16:58:08",
      "content": "<p>I think the key question is that how the data were sampled and splited into training set and testing set. It does count! And we may simulate this sample approach to generate local test set.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 8968,
      "author_name": "yiwang0",
      "author_url": "",
      "post_date": "03/05/2012 01:46:36",
      "content": "<p>There isn't a sample approach. The data were divided by date. The whole data set is a sample of logs in 55 days, where the part of the first 50 days is the training data set, and the part of the rest 5 days is the testing data set.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 8972,
      "author_name": "goldensection",
      "author_url": "",
      "post_date": "03/05/2012 02:10:49",
      "content": "<p>Well, it makes sense. However, there is no timestamp information given for us. Shall we assume that the training instance are sorted by date? Thanks!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 8973,
      "author_name": "yiwang0",
      "author_url": "",
      "post_date": "03/05/2012 03:01:43",
      "content": "<pre>No, we cannot assume a temporal order, as stated in the Description page: &quot;We aggregate instances with the same user id, ad id, query, and setting in order to reduce the dataset size. &quot;  In other words, each training/testing instance comes from one or more session log messages.</pre>",
      "votes": null,
      "replies": []
    },
    {
      "id": 8974,
      "author_name": "leustagos",
      "author_url": "",
      "post_date": "03/05/2012 03:05:24",
      "content": "<p>Could you give an example of how a test instance would look like?</p>\r\n<p>The same as training except for the impression and click cols?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 8981,
      "author_name": "yiwang0",
      "author_url": "",
      "post_date": "03/05/2012 04:19:35",
      "content": "<p>Yes, as you said, same as training instances except for the impression and click columns.</p>\r\n<p>Soon, we will publish a validation data set, which has exactly the same format as the testing data set.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 9024,
      "author_name": "piiswrong",
      "author_url": "",
      "post_date": "03/05/2012 11:21:13",
      "content": "<p>Did you mean the training data was produced by aggregating the first 50 days' log, and the testing data was produced by aggregating the last 5 days' log? So we would expect statistically identical distribution of users, querys, and ads?</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "8912": "",
    "8924": "",
    "8933": "",
    "8968": "",
    "8972": "",
    "8973": "",
    "8974": "",
    "8981": "",
    "9024": ""
  },
  "source": "meta"
}