{
  "id": 20485,
  "title": "CNT features in test",
  "url": "/competitions/expedia-hotel-recommendations/discussion/20485",
  "author_name": "",
  "post_date": "2016-04-27T16:24:00.830Z",
  "votes": null,
  "comment_count": 2,
  "views": 554,
  "content": "<p>After a half-day exploration into the correlation of 'cnt' with other features, I still have no clues about how this column are aligned with others.\nIs there anyone have an insight into how to generate cnt for testset?</p>",
  "messages": [
    {
      "id": "117155",
      "postDate": "04/27/2016 16:24:00",
      "content": "<p>After a half-day exploration into the correlation of 'cnt' with other features, I still have no clues about how this column are aligned with others.\nIs there anyone have an insight into how to generate cnt for testset?</p>",
      "rawMarkdown": "After a half-day exploration into the correlation of 'cnt' with other features, I still have no clues about how this column are aligned with others.\r\nIs there anyone have an insight into how to generate cnt for testset?",
      "votes": null
    },
    {
      "id": "117167",
      "postDate": "04/27/2016 17:00:11",
      "content": "<p>Hi, </p>\n\n<p>you cannot easily create this column (with valuable content) in test. Because you just don't have that information. Which is explicitly stated.</p>\n\n<p>But you might use this information nevertheless. Maybe depending on your classifier, algorithm or script. The popular script that calculates most often clusters per destination for example uses this information. But not in the test set. They just join data from train and test.</p>\n\n<p>The reasoning is this: Someone who clicked a hotel a hundred times might as well have booked that hotel instead of the one he actually did.  </p>\n\n<p>Gerhard</p>",
      "rawMarkdown": "Hi, \r\n\r\nyou cannot easily create this column (with valuable content) in test. Because you just don't have that information. Which is explicitly stated.\r\n\r\nBut you might use this information nevertheless. Maybe depending on your classifier, algorithm or script. The popular script that calculates most often clusters per destination for example uses this information. But not in the test set. They just join data from train and test.\r\n\r\nThe reasoning is this: Someone who clicked a hotel a hundred times might as well have booked that hotel instead of the one he actually did.  \r\n\r\nGerhard",
      "votes": null
    },
    {
      "id": "117177",
      "postDate": "04/27/2016 17:44:14",
      "content": "<p>Actually, cnt on test should be (almost always) 1, acoording to Adam in \n<a href=\"https://www.kaggle.com/c/expedia-hotel-recommendations/forums/t/20204/data-source-questions\">https://www.kaggle.com/c/expedia-hotel-recommendations/forums/t/20204/data-source-questions</a></p>\n\n<p>I'm planning on simply fill it with ones.</p>",
      "rawMarkdown": "Actually, cnt on test should be (almost always) 1, acoording to Adam in \r\nhttps://www.kaggle.com/c/expedia-hotel-recommendations/forums/t/20204/data-source-questions\r\n\r\nI'm planning on simply fill it with ones.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 117167,
      "author_name": "mightybird",
      "author_url": "",
      "post_date": "04/27/2016 17:00:11",
      "content": "<p>Hi, </p>\n\n<p>you cannot easily create this column (with valuable content) in test. Because you just don't have that information. Which is explicitly stated.</p>\n\n<p>But you might use this information nevertheless. Maybe depending on your classifier, algorithm or script. The popular script that calculates most often clusters per destination for example uses this information. But not in the test set. They just join data from train and test.</p>\n\n<p>The reasoning is this: Someone who clicked a hotel a hundred times might as well have booked that hotel instead of the one he actually did.  </p>\n\n<p>Gerhard</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 117177,
      "author_name": "khaoticmind",
      "author_url": "",
      "post_date": "04/27/2016 17:44:14",
      "content": "<p>Actually, cnt on test should be (almost always) 1, acoording to Adam in \n<a href=\"https://www.kaggle.com/c/expedia-hotel-recommendations/forums/t/20204/data-source-questions\">https://www.kaggle.com/c/expedia-hotel-recommendations/forums/t/20204/data-source-questions</a></p>\n\n<p>I'm planning on simply fill it with ones.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "117155": "After a half-day exploration into the correlation of 'cnt' with other features, I still have no clues about how this column are aligned with others.\r\nIs there anyone have an insight into how to generate cnt for testset?",
    "117167": "Hi, \r\n\r\nyou cannot easily create this column (with valuable content) in test. Because you just don't have that information. Which is explicitly stated.\r\n\r\nBut you might use this information nevertheless. Maybe depending on your classifier, algorithm or script. The popular script that calculates most often clusters per destination for example uses this information. But not in the test set. They just join data from train and test.\r\n\r\nThe reasoning is this: Someone who clicked a hotel a hundred times might as well have booked that hotel instead of the one he actually did.  \r\n\r\nGerhard",
    "117177": "Actually, cnt on test should be (almost always) 1, acoording to Adam in \r\nhttps://www.kaggle.com/c/expedia-hotel-recommendations/forums/t/20204/data-source-questions\r\n\r\nI'm planning on simply fill it with ones."
  },
  "source": "meta"
}