{
  "id": 20682,
  "title": "Map@5 with random forest classifier",
  "url": "/competitions/expedia-hotel-recommendations/discussion/20682",
  "author_name": "",
  "post_date": "2016-05-03T19:04:34.117Z",
  "votes": null,
  "comment_count": 7,
  "views": 1270,
  "content": "<p>Hi, </p>\n\n<p>I use a random forest classifier to predict the hotel cluster, currently I get only one value. I do not know how to get the 4 other hotel cluster (the most closest hotel cluster to the predicted value) using the random forest classifier. I do not if it's possible using a random forest classifier or not.</p>\n\n<p>Regards, </p>",
  "messages": [
    {
      "id": "118454",
      "postDate": "05/03/2016 19:04:34",
      "content": "<p>Hi, </p>\n\n<p>I use a random forest classifier to predict the hotel cluster, currently I get only one value. I do not know how to get the 4 other hotel cluster (the most closest hotel cluster to the predicted value) using the random forest classifier. I do not if it's possible using a random forest classifier or not.</p>\n\n<p>Regards, </p>",
      "rawMarkdown": "Hi, \r\n\r\nI use a random forest classifier to predict the hotel cluster, currently I get only one value. I do not know how to get the 4 other hotel cluster (the most closest hotel cluster to the predicted value) using the random forest classifier. I do not if it's possible using a random forest classifier or not.\r\n\r\nRegards,",
      "votes": null
    },
    {
      "id": "118458",
      "postDate": "05/03/2016 19:39:22",
      "content": "<p>You need to ask for you calssifier output the probabilites of the classification.\nThis output is something on the lines of [0.01, 0.02, 0.3, 0.005 ...] where each value is the probability of the outcome of the class at that index. So in the example, Class 0 (wheetever this may be) has probability 0.01 of happening, Class 1 has probability 0.02 and so forth.</p>\n\n<p>With that you need to sort this array and get the 5 higher probabilites as your prediction.</p>\n\n<p>If you are using Python It is something on the lines of </p>\n\n<pre><code>preds = clf.predict_proba(X)\npredicted = preds.argsort(axis=1)[:,-np.arange(1,6)]\n</code></pre>",
      "rawMarkdown": "You need to ask for you calssifier output the probabilites of the classification.\r\nThis output is something on the lines of [0.01, 0.02, 0.3, 0.005 ...] where each value is the probability of the outcome of the class at that index. So in the example, Class 0 (wheetever this may be) has probability 0.01 of happening, Class 1 has probability 0.02 and so forth.\r\n\r\nWith that you need to sort this array and get the 5 higher probabilites as your prediction.\r\n\r\nIf you are using Python It is something on the lines of \r\n\r\n    preds = clf.predict_proba(X)\r\n    predicted = preds.argsort(axis=1)[:,-np.arange(1,6)]",
      "votes": null
    },
    {
      "id": "118462",
      "postDate": "05/03/2016 19:53:51",
      "content": "<p>If you are using random forest in R, to get the probabilities for each of the 100 hotel clusters you can do something like this:</p>\n\n<p>predict.model &lt;- predict(model.forest, <strong>type=&quot;prob&quot;</strong>, newdata=test.data)</p>",
      "rawMarkdown": "If you are using random forest in R, to get the probabilities for each of the 100 hotel clusters you can do something like this:\r\n\r\npredict.model <- predict(model.forest, **type=\"prob\"**, newdata=test.data)",
      "votes": null
    },
    {
      "id": "118463",
      "postDate": "05/03/2016 19:54:35",
      "content": "<p>Thank you Antonio, I forgot this alternative.</p>",
      "rawMarkdown": "Thank you Antonio, I forgot this alternative.",
      "votes": null
    },
    {
      "id": "118476",
      "postDate": "05/03/2016 20:59:23",
      "content": "<p>Do you have an idea about the most efficient way to create the hotel_cluster string from the following array ? </p>\n\n<pre><code>array([[25, 35, 16, 40, 83],\n   [43, 99, 12, 20, 64],\n   [95, 89, 21, 69, 91],\n   ..., \n   [ 1, 49, 51, 54, 79],\n   [42, 13, 84, 18, 10],\n   [62, 99,  8, 51, 57]])\n</code></pre>\n\n<p>I want to get for the first row for example the following string  : </p>\n\n<pre><code> &quot;25 35 16 40 83&quot;\n</code></pre>",
      "rawMarkdown": "Do you have an idea about the most efficient way to create the hotel_cluster string from the following array ? \r\n\r\n    array([[25, 35, 16, 40, 83],\r\n       [43, 99, 12, 20, 64],\r\n       [95, 89, 21, 69, 91],\r\n       ..., \r\n       [ 1, 49, 51, 54, 79],\r\n       [42, 13, 84, 18, 10],\r\n       [62, 99,  8, 51, 57]])\r\n\r\nI want to get for the first row for example the following string  : \r\n\r\n     \"25 35 16 40 83\"",
      "votes": null
    },
    {
      "id": "118481",
      "postDate": "05/03/2016 21:52:26",
      "content": "<p>[quote=AissaElOuafi;118476]</p>\n\n<p>Do you have an idea about the most efficient way to create the hotel_cluster string from the following array ? </p>\n\n<pre><code>array([[25, 35, 16, 40, 83],\n   [43, 99, 12, 20, 64],\n   [95, 89, 21, 69, 91],\n   ..., \n   [ 1, 49, 51, 54, 79],\n   [42, 13, 84, 18, 10],\n   [62, 99,  8, 51, 57]])\n</code></pre>\n\n<p>I want to get for the first row for example the following string  : </p>\n\n<pre><code> &quot;25 35 16 40 83&quot;\n</code></pre>\n\n<p>[/quote]</p>\n\n<p>To convert each row you can use something like</p>\n\n<pre><code>row_string = &quot; &quot;.join(str(n) for n in row)\n</code></pre>",
      "rawMarkdown": "[quote=AissaElOuafi;118476]\r\n\r\nDo you have an idea about the most efficient way to create the hotel_cluster string from the following array ? \r\n\r\n    array([[25, 35, 16, 40, 83],\r\n       [43, 99, 12, 20, 64],\r\n       [95, 89, 21, 69, 91],\r\n       ..., \r\n       [ 1, 49, 51, 54, 79],\r\n       [42, 13, 84, 18, 10],\r\n       [62, 99,  8, 51, 57]])\r\n\r\nI want to get for the first row for example the following string  : \r\n\r\n     \"25 35 16 40 83\"\r\n\r\n[/quote]\r\n\r\nTo convert each row you can use something like\r\n\r\n    row_string = \" \".join(str(n) for n in row)",
      "votes": null
    },
    {
      "id": "118482",
      "postDate": "05/03/2016 21:57:41",
      "content": "<p>Thank you David, it's work. </p>\n\n<p>I have another question related to the ML algorithm to apply, I try a Random Forest Classifier but I get a very bad result (Map@5 ~0.11) using 180 000 rows of data. I created an AWS instance (64 Gb RAM and 16 core) I will try to use more of data rows, do you have an idea about algorithms that I should try ? </p>",
      "rawMarkdown": "Thank you David, it's work. \r\n\r\nI have another question related to the ML algorithm to apply, I try a Random Forest Classifier but I get a very bad result (Map@5 ~0.11) using 180 000 rows of data. I created an AWS instance (64 Gb RAM and 16 core) I will try to use more of data rows, do you have an idea about algorithms that I should try ?",
      "votes": null
    },
    {
      "id": "119028",
      "postDate": "05/06/2016 19:08:56",
      "content": "<p>What features did you end up using for your random forest, AissaElOuafi?</p>",
      "rawMarkdown": "What features did you end up using for your random forest, AissaElOuafi?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 118458,
      "author_name": "khaoticmind",
      "author_url": "",
      "post_date": "05/03/2016 19:39:22",
      "content": "<p>You need to ask for you calssifier output the probabilites of the classification.\nThis output is something on the lines of [0.01, 0.02, 0.3, 0.005 ...] where each value is the probability of the outcome of the class at that index. So in the example, Class 0 (wheetever this may be) has probability 0.01 of happening, Class 1 has probability 0.02 and so forth.</p>\n\n<p>With that you need to sort this array and get the 5 higher probabilites as your prediction.</p>\n\n<p>If you are using Python It is something on the lines of </p>\n\n<pre><code>preds = clf.predict_proba(X)\npredicted = preds.argsort(axis=1)[:,-np.arange(1,6)]\n</code></pre>",
      "votes": null,
      "replies": []
    },
    {
      "id": 118462,
      "author_name": "mathandbikes",
      "author_url": "",
      "post_date": "05/03/2016 19:53:51",
      "content": "<p>If you are using random forest in R, to get the probabilities for each of the 100 hotel clusters you can do something like this:</p>\n\n<p>predict.model &lt;- predict(model.forest, <strong>type=&quot;prob&quot;</strong>, newdata=test.data)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 118463,
      "author_name": "aissaelouafi",
      "author_url": "",
      "post_date": "05/03/2016 19:54:35",
      "content": "<p>Thank you Antonio, I forgot this alternative.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 118476,
      "author_name": "aissaelouafi",
      "author_url": "",
      "post_date": "05/03/2016 20:59:23",
      "content": "<p>Do you have an idea about the most efficient way to create the hotel_cluster string from the following array ? </p>\n\n<pre><code>array([[25, 35, 16, 40, 83],\n   [43, 99, 12, 20, 64],\n   [95, 89, 21, 69, 91],\n   ..., \n   [ 1, 49, 51, 54, 79],\n   [42, 13, 84, 18, 10],\n   [62, 99,  8, 51, 57]])\n</code></pre>\n\n<p>I want to get for the first row for example the following string  : </p>\n\n<pre><code> &quot;25 35 16 40 83&quot;\n</code></pre>",
      "votes": null,
      "replies": []
    },
    {
      "id": 118481,
      "author_name": "davidtran",
      "author_url": "",
      "post_date": "05/03/2016 21:52:26",
      "content": "<p>[quote=AissaElOuafi;118476]</p>\n\n<p>Do you have an idea about the most efficient way to create the hotel_cluster string from the following array ? </p>\n\n<pre><code>array([[25, 35, 16, 40, 83],\n   [43, 99, 12, 20, 64],\n   [95, 89, 21, 69, 91],\n   ..., \n   [ 1, 49, 51, 54, 79],\n   [42, 13, 84, 18, 10],\n   [62, 99,  8, 51, 57]])\n</code></pre>\n\n<p>I want to get for the first row for example the following string  : </p>\n\n<pre><code> &quot;25 35 16 40 83&quot;\n</code></pre>\n\n<p>[/quote]</p>\n\n<p>To convert each row you can use something like</p>\n\n<pre><code>row_string = &quot; &quot;.join(str(n) for n in row)\n</code></pre>",
      "votes": null,
      "replies": []
    },
    {
      "id": 118482,
      "author_name": "aissaelouafi",
      "author_url": "",
      "post_date": "05/03/2016 21:57:41",
      "content": "<p>Thank you David, it's work. </p>\n\n<p>I have another question related to the ML algorithm to apply, I try a Random Forest Classifier but I get a very bad result (Map@5 ~0.11) using 180 000 rows of data. I created an AWS instance (64 Gb RAM and 16 core) I will try to use more of data rows, do you have an idea about algorithms that I should try ? </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 119028,
      "author_name": "janphysicsmcgill",
      "author_url": "",
      "post_date": "05/06/2016 19:08:56",
      "content": "<p>What features did you end up using for your random forest, AissaElOuafi?</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "118454": "Hi, \r\n\r\nI use a random forest classifier to predict the hotel cluster, currently I get only one value. I do not know how to get the 4 other hotel cluster (the most closest hotel cluster to the predicted value) using the random forest classifier. I do not if it's possible using a random forest classifier or not.\r\n\r\nRegards,",
    "118458": "You need to ask for you calssifier output the probabilites of the classification.\r\nThis output is something on the lines of [0.01, 0.02, 0.3, 0.005 ...] where each value is the probability of the outcome of the class at that index. So in the example, Class 0 (wheetever this may be) has probability 0.01 of happening, Class 1 has probability 0.02 and so forth.\r\n\r\nWith that you need to sort this array and get the 5 higher probabilites as your prediction.\r\n\r\nIf you are using Python It is something on the lines of \r\n\r\n    preds = clf.predict_proba(X)\r\n    predicted = preds.argsort(axis=1)[:,-np.arange(1,6)]",
    "118462": "If you are using random forest in R, to get the probabilities for each of the 100 hotel clusters you can do something like this:\r\n\r\npredict.model <- predict(model.forest, **type=\"prob\"**, newdata=test.data)",
    "118463": "Thank you Antonio, I forgot this alternative.",
    "118476": "Do you have an idea about the most efficient way to create the hotel_cluster string from the following array ? \r\n\r\n    array([[25, 35, 16, 40, 83],\r\n       [43, 99, 12, 20, 64],\r\n       [95, 89, 21, 69, 91],\r\n       ..., \r\n       [ 1, 49, 51, 54, 79],\r\n       [42, 13, 84, 18, 10],\r\n       [62, 99,  8, 51, 57]])\r\n\r\nI want to get for the first row for example the following string  : \r\n\r\n     \"25 35 16 40 83\"",
    "118481": "[quote=AissaElOuafi;118476]\r\n\r\nDo you have an idea about the most efficient way to create the hotel_cluster string from the following array ? \r\n\r\n    array([[25, 35, 16, 40, 83],\r\n       [43, 99, 12, 20, 64],\r\n       [95, 89, 21, 69, 91],\r\n       ..., \r\n       [ 1, 49, 51, 54, 79],\r\n       [42, 13, 84, 18, 10],\r\n       [62, 99,  8, 51, 57]])\r\n\r\nI want to get for the first row for example the following string  : \r\n\r\n     \"25 35 16 40 83\"\r\n\r\n[/quote]\r\n\r\nTo convert each row you can use something like\r\n\r\n    row_string = \" \".join(str(n) for n in row)",
    "118482": "Thank you David, it's work. \r\n\r\nI have another question related to the ML algorithm to apply, I try a Random Forest Classifier but I get a very bad result (Map@5 ~0.11) using 180 000 rows of data. I created an AWS instance (64 Gb RAM and 16 core) I will try to use more of data rows, do you have an idea about algorithms that I should try ?",
    "119028": "What features did you end up using for your random forest, AissaElOuafi?"
  },
  "source": "meta"
}