{
  "id": 21131,
  "title": "have you used destinations.csv?",
  "url": "/competitions/expedia-hotel-recommendations/discussion/21131",
  "author_name": "masterLiu",
  "post_date": "2016-05-22T01:56:46.413000",
  "votes": 0,
  "comment_count": 4,
  "views": 1016,
  "content": "<p>I don't see topic on destinations.csv, do you guys used that data? i see it has 149 features and mainly float numbers, is that useful and how to use it?</p>",
  "messages": [
    {
      "id": 120962,
      "postDate": "2016-05-22T08:47:38.573Z",
      "content": "<p>I am using them as additional features for my training sets.</p>\n\n<p>You can include them straight away, or - as a hint - you may notice that the sum of   10^x  for all values in a row always equals to 1.0. With that, you can start preprocessing your destination features.</p>",
      "rawMarkdown": "I am using them as additional features for my training sets.\r\n\r\nYou can include them straight away, or - as a hint - you may notice that the sum of   10^x  for all values in a row always equals to 1.0. With that, you can start preprocessing your destination features.",
      "votes": 1
    },
    {
      "id": 121284,
      "postDate": "2016-05-25T10:47:58.200Z",
      "content": "<p>Sure.</p>\n\n<p>It looks like the last preprocessing step applied the the values was x = log10(x)</p>\n\n<p>If you revert that transformation (by doing x = 10^x) you will get values in [0,1]. These values, since they all sum up to 1.0 (in each row) look <em>suspiciously</em> as probabilities.</p>\n\n<p>Knowing that, you might feel more comfortable working with these values represented as probabilities. Or even further, you may notice that none of these values are 0.0, but instead, the mode of them is also the minimum value - which suggest that some kind of Laplace smoothing was used.</p>\n\n<p>With that, you can try to reverse-engineering the computation of probabilities, and end up with a sparse matrix of destinations X number of times the destination shown a feature. Which is what I am using at the moment.</p>",
      "rawMarkdown": "Sure.\r\n\r\nIt looks like the last preprocessing step applied the the values was x = log10(x)\r\n\r\nIf you revert that transformation (by doing x = 10^x) you will get values in [0,1]. These values, since they all sum up to 1.0 (in each row) look *suspiciously* as probabilities.\r\n\r\nKnowing that, you might feel more comfortable working with these values represented as probabilities. Or even further, you may notice that none of these values are 0.0, but instead, the mode of them is also the minimum value - which suggest that some kind of Laplace smoothing was used.\r\n\r\nWith that, you can try to reverse-engineering the computation of probabilities, and end up with a sparse matrix of destinations X number of times the destination shown a feature. Which is what I am using at the moment.\r\n",
      "votes": 2
    },
    {
      "id": 120948,
      "postDate": "2016-05-22T01:56:46.413Z",
      "content": "<p>I don't see topic on destinations.csv, do you guys used that data? i see it has 149 features and mainly float numbers, is that useful and how to use it?</p>",
      "rawMarkdown": "I don't see topic on destinations.csv, do you guys used that data? i see it has 149 features and mainly float numbers, is that useful and how to use it?"
    },
    {
      "id": 121292,
      "postDate": "2016-05-25T12:14:25.917Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 121278,
      "postDate": "2016-05-25T09:58:16.540Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 120962,
      "author_name": "CarrDelling",
      "author_url": "",
      "post_date": "2016-05-22T08:47:38.573000",
      "content": "<p>I am using them as additional features for my training sets.</p>\n\n<p>You can include them straight away, or - as a hint - you may notice that the sum of   10^x  for all values in a row always equals to 1.0. With that, you can start preprocessing your destination features.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 121284,
      "author_name": "CarrDelling",
      "author_url": "",
      "post_date": "2016-05-25T10:47:58.200000",
      "content": "<p>Sure.</p>\n\n<p>It looks like the last preprocessing step applied the the values was x = log10(x)</p>\n\n<p>If you revert that transformation (by doing x = 10^x) you will get values in [0,1]. These values, since they all sum up to 1.0 (in each row) look <em>suspiciously</em> as probabilities.</p>\n\n<p>Knowing that, you might feel more comfortable working with these values represented as probabilities. Or even further, you may notice that none of these values are 0.0, but instead, the mode of them is also the minimum value - which suggest that some kind of Laplace smoothing was used.</p>\n\n<p>With that, you can try to reverse-engineering the computation of probabilities, and end up with a sparse matrix of destinations X number of times the destination shown a feature. Which is what I am using at the moment.</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 121292,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-05-25T12:14:25.917000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 121278,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-05-25T09:58:16.540000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "120962": "I am using them as additional features for my training sets.\r\n\r\nYou can include them straight away, or - as a hint - you may notice that the sum of   10^x  for all values in a row always equals to 1.0. With that, you can start preprocessing your destination features.",
    "121284": "Sure.\r\n\r\nIt looks like the last preprocessing step applied the the values was x = log10(x)\r\n\r\nIf you revert that transformation (by doing x = 10^x) you will get values in [0,1]. These values, since they all sum up to 1.0 (in each row) look *suspiciously* as probabilities.\r\n\r\nKnowing that, you might feel more comfortable working with these values represented as probabilities. Or even further, you may notice that none of these values are 0.0, but instead, the mode of them is also the minimum value - which suggest that some kind of Laplace smoothing was used.\r\n\r\nWith that, you can try to reverse-engineering the computation of probabilities, and end up with a sparse matrix of destinations X number of times the destination shown a feature. Which is what I am using at the moment.\r\n",
    "120948": "I don't see topic on destinations.csv, do you guys used that data? i see it has 149 features and mainly float numbers, is that useful and how to use it?",
    "121292": "",
    "121278": ""
  }
}