{
  "id": 20503,
  "title": "Is this competition a 100 classifications issue?",
  "url": "/competitions/expedia-hotel-recommendations/discussion/20503",
  "author_name": "",
  "post_date": "2016-04-28T10:07:17.313Z",
  "votes": null,
  "comment_count": 21,
  "views": 2892,
  "content": "<p>Is this possible to train a logistic regression or any other classification model to get a good result? Or just make some regulations...(for example ,most popular..)</p>",
  "messages": [
    {
      "id": "117275",
      "postDate": "04/28/2016 10:07:17",
      "content": "<p>Is this possible to train a logistic regression or any other classification model to get a good result? Or just make some regulations...(for example ,most popular..)</p>",
      "rawMarkdown": "Is this possible to train a logistic regression or any other classification model to get a good result? Or just make some regulations...(for example ,most popular..)",
      "votes": null
    },
    {
      "id": "117295",
      "postDate": "04/28/2016 12:07:26",
      "content": "<p>It should be possible to solve this with logistic regression. However you need predictions (probabilities) for all 100 hotel_clusters. Then take the top 5 clusters for each line in test.</p>\n\n<p>Gerhard</p>",
      "rawMarkdown": "It should be possible to solve this with logistic regression. However you need predictions (probabilities) for all 100 hotel_clusters. Then take the top 5 clusters for each line in test.\r\n\r\nGerhard",
      "votes": null
    },
    {
      "id": "117304",
      "postDate": "04/28/2016 12:33:58",
      "content": "<p>I fail to see how you can use regression on this problem...\nThe hotel cluster values are not ints, they are IDs, so if a hotel cluster is 1 it is not any closer to 2 than it is from 99.</p>\n\n<p>Also, you have to give five results (that depends on the probability of all clusters) what, again, regression can't do for you.</p>",
      "rawMarkdown": "I fail to see how you can use regression on this problem...\r\nThe hotel cluster values are not ints, they are IDs, so if a hotel cluster is 1 it is not any closer to 2 than it is from 99.\r\n\r\nAlso, you have to give five results (that depends on the probability of all clusters) what, again, regression can't do for you.",
      "votes": null
    },
    {
      "id": "117306",
      "postDate": "04/28/2016 13:11:13",
      "content": "<p>Hi,</p>\n\n<p>let me show you the way...</p>\n\n<p>Let's reduce the problem first. Let's say you want to predict if a booking belongs to cluster 0 or NOT. This can well be solved by LOGISTIC regression. Because it will give you probabilities for just that question. </p>\n\n<p>Now you do this for every cluster value, so 100 times. You end up with 100 probabilities per test row. Each indicating the probablity for this row to result in cluster x. Now you take the 5 results (clusters) with the highest probabilities and submit them.</p>\n\n<p>Gerhard</p>",
      "rawMarkdown": "Hi,\r\n\r\nlet me show you the way...\r\n\r\nLet's reduce the problem first. Let's say you want to predict if a booking belongs to cluster 0 or NOT. This can well be solved by LOGISTIC regression. Because it will give you probabilities for just that question. \r\n\r\nNow you do this for every cluster value, so 100 times. You end up with 100 probabilities per test row. Each indicating the probablity for this row to result in cluster x. Now you take the 5 results (clusters) with the highest probabilities and submit them.\r\n\r\nGerhard",
      "votes": null
    },
    {
      "id": "117310",
      "postDate": "04/28/2016 13:26:05",
      "content": "<p>[quote=Antonio Augusto Santos;117304]</p>\n\n<p>I fail to see how you can use regression on this problem...</p>\n\n<p>[/quote]</p>\n\n<p>Logistic Regression = Classification</p>\n\n<p>If you use a two-class logistic regression, you will need 100 models: ID=1 against all others, ID=2 against all others, ID=3 against all others, etc. Once you have the probabilities for each ID, you rank them by how high the probability is.\nIf you use a multi-class logistic regression, 1 model is sufficient, you only need to sort ranks afterwards.</p>",
      "rawMarkdown": "[quote=Antonio Augusto Santos;117304]\r\n\r\nI fail to see how you can use regression on this problem...\r\n\r\n[/quote]\r\n\r\nLogistic Regression = Classification\r\n\r\nIf you use a two-class logistic regression, you will need 100 models: ID=1 against all others, ID=2 against all others, ID=3 against all others, etc. Once you have the probabilities for each ID, you rank them by how high the probability is.\r\nIf you use a multi-class logistic regression, 1 model is sufficient, you only need to sort ranks afterwards.",
      "votes": null
    },
    {
      "id": "117364",
      "postDate": "04/28/2016 18:16:28",
      "content": "<p>Sorry guys! My mistake :)\nI faield to see the LOGISTIC before the REGRESSION lol! </p>",
      "rawMarkdown": "Sorry guys! My mistake :)\r\nI faield to see the LOGISTIC before the REGRESSION lol!",
      "votes": null
    },
    {
      "id": "122545",
      "postDate": "06/05/2016 01:06:53",
      "content": "<p>Hi,</p>\n\n<p>I took this thread's advice and made 100 classification models.  I now have an array with 2.5 million rows (for each data item in the test sample) and 100 columns with the classification prediction numbers (for each hotel cluster).  Can anyone please suggest a way to get the top 5 clusters?  You don't have to code this for me, but any help would be appreciated.  I tried transposing the numpy matrix and sorting by one column (which represents a data item in the tranposed matrix) at a time, but this gives me the top 5 probabilities, not the hotel cluster numbers that correspond to those top five probabilities.</p>\n\n<p>Thank you</p>",
      "rawMarkdown": "Hi,\r\n\r\nI took this thread's advice and made 100 classification models.  I now have an array with 2.5 million rows (for each data item in the test sample) and 100 columns with the classification prediction numbers (for each hotel cluster).  Can anyone please suggest a way to get the top 5 clusters?  You don't have to code this for me, but any help would be appreciated.  I tried transposing the numpy matrix and sorting by one column (which represents a data item in the tranposed matrix) at a time, but this gives me the top 5 probabilities, not the hotel cluster numbers that correspond to those top five probabilities.\r\n\r\nThank you",
      "votes": null
    },
    {
      "id": "122549",
      "postDate": "06/05/2016 02:39:04",
      "content": "<p>Each column should represent the likelihood of a booking being in that hotel class. So you just need to find the top five values for each row, and return the associated hotel cluster.</p>\n\n<p>Good luck,\nJeff</p>",
      "rawMarkdown": "Each column should represent the likelihood of a booking being in that hotel class. So you just need to find the top five values for each row, and return the associated hotel cluster.\r\n\r\nGood luck,\r\nJeff",
      "votes": null
    },
    {
      "id": "122550",
      "postDate": "06/05/2016 02:50:37",
      "content": "<p>May be this thread will help you.</p>\n\n<p><a href=\"https://www.kaggle.com/c/expedia-hotel-recommendations/forums/t/20556/map5-function-or-eval-metric/120977\">https://www.kaggle.com/c/expedia-hotel-recommendations/forums/t/20556/map5-function-or-eval-metric/120977</a></p>",
      "rawMarkdown": "May be this thread will help you.\r\n\r\nhttps://www.kaggle.com/c/expedia-hotel-recommendations/forums/t/20556/map5-function-or-eval-metric/120977",
      "votes": null
    },
    {
      "id": "122551",
      "postDate": "06/05/2016 02:58:05",
      "content": "<p>JeffH, </p>\n\n<p>Thank you.   I have a numpy matrix with the columns like you said.   I understand that I need to find the top five values for each row, but my problem is that I am having trouble writing code to do that.   I can get the top value, but not the top five. </p>",
      "rawMarkdown": "JeffH, \r\n\r\nThank you.   I have a numpy matrix with the columns like you said.   I understand that I need to find the top five values for each row, but my problem is that I am having trouble writing code to do that.   I can get the top value, but not the top five.",
      "votes": null
    },
    {
      "id": "122552",
      "postDate": "06/05/2016 03:06:33",
      "content": "<p>Subhajit, </p>\n\n<p>Sorry, I forgot to mention that I am using python.   I do not see this in the link you provided. </p>\n\n<p>Thank you though, \nMike</p>",
      "rawMarkdown": "Subhajit, \r\n\r\nSorry, I forgot to mention that I am using python.   I do not see this in the link you provided. \r\n\r\nThank you though, \r\nMike",
      "votes": null
    },
    {
      "id": "122560",
      "postDate": "06/05/2016 04:08:12",
      "content": "<p>Look for dune_dweller's post in the thread. Anyway, let me repost it here for you:</p>\n\n<pre><code>def map5eval(preds, dtrain):\n    actual = dtrain.get_label()\n    predicted = preds.argsort(axis=1)[:,-np.arange(1,6)]\n    metric = 0.\n    for i in range(5):\n        metric += np.sum(actual==predicted[:,i])/(i+1)\n    metric /= actual.shape[0]\n    return 'MAP@5', metric\n</code></pre>",
      "rawMarkdown": "Look for dune_dweller's post in the thread. Anyway, let me repost it here for you:\r\n\r\n    def map5eval(preds, dtrain):\r\n        actual = dtrain.get_label()\r\n        predicted = preds.argsort(axis=1)[:,-np.arange(1,6)]\r\n        metric = 0.\r\n        for i in range(5):\r\n            metric += np.sum(actual==predicted[:,i])/(i+1)\r\n        metric /= actual.shape[0]\r\n        return 'MAP@5', metric",
      "votes": null
    },
    {
      "id": "122563",
      "postDate": "06/05/2016 04:41:01",
      "content": "<p>@Subhajit Mandal   Did you use xgb result in your best submission?  </p>",
      "rawMarkdown": "Subhajit Mandal   Did you use xgb result in your best submission?",
      "votes": null
    },
    {
      "id": "122623",
      "postDate": "06/05/2016 21:57:04",
      "content": "<p>Subhajit,</p>\n\n<p>Thank you so much!  That worked.  I can't tell you how good that feels to try this for 3 hours and not be able to come up with a solution and then to discover this .argsort method you shared and it works!</p>\n\n<p>Thank you again,\nMike</p>",
      "rawMarkdown": "Subhajit,\r\n\r\nThank you so much!  That worked.  I can't tell you how good that feels to try this for 3 hours and not be able to come up with a solution and then to discover this .argsort method you shared and it works!\r\n\r\nThank you again,\r\nMike",
      "votes": null
    },
    {
      "id": "122627",
      "postDate": "06/05/2016 23:24:21",
      "content": "<p>@FengLi, yes I used xgb.</p>\n\n<p>@Mike, you're welcome... although the code is not mine.</p>",
      "rawMarkdown": "FengLi, yes I used xgb.\r\n\r\n@Mike, you're welcome... although the code is not mine.",
      "votes": null
    },
    {
      "id": "122629",
      "postDate": "06/05/2016 23:56:21",
      "content": "<p>I tried this method with 100 models, as explained above by Mighty, Laurae and Subhajit. I used XGBoost (binary logistic), as target is <code>is_booking</code>, so it tries to get likelihood of booking vs clicking for each <code>hotel_cluster</code>. I finished it in 5 days. </p>\n\n<p>All <code>destinations.csv</code> features are merged. To get top 5 for each row, I tried (1) rank based and (2) scale [0..1] based sorting. Logloss is around 0.19 for each xgb model. The result is very disappointing, <strong>under 0.1</strong> public LB score (before adding leakage result). </p>\n\n<p>Why? Is the target (<code>is_booking</code>) wrong? </p>",
      "rawMarkdown": "I tried this method with 100 models, as explained above by Mighty, Laurae and Subhajit. I used XGBoost (binary logistic), as target is `is_booking`, so it tries to get likelihood of booking vs clicking for each `hotel_cluster`. I finished it in 5 days. \r\n\r\nAll `destinations.csv` features are merged. To get top 5 for each row, I tried (1) rank based and (2) scale [0..1] based sorting. Logloss is around 0.19 for each xgb model. The result is very disappointing, **under 0.1** public LB score (before adding leakage result). \r\n\r\nWhy? Is the target (`is_booking`) wrong?",
      "votes": null
    },
    {
      "id": "122645",
      "postDate": "06/06/2016 03:15:01",
      "content": "<p>@Subhajit Mandal Did you use 'multi:softprob' objective function? </p>",
      "rawMarkdown": "Subhajit Mandal Did you use 'multi:softprob' objective function?",
      "votes": null
    },
    {
      "id": "122886",
      "postDate": "06/08/2016 02:35:04",
      "content": "<p>[quote=aldente;122629]\nIs the target (<code>is_booking</code>) wrong? \n[/quote]</p>\n\n<p>The target is the variable &quot;hotel_cluster&quot;</p>",
      "rawMarkdown": "[quote=aldente;122629]\r\nIs the target (`is_booking`) wrong? \r\n[/quote]\r\n\r\nThe target is the variable \"hotel_cluster\"",
      "votes": null
    },
    {
      "id": "122888",
      "postDate": "06/08/2016 02:49:57",
      "content": "<p>@Icaro, </p>\n\n<p>It should be <code>hotel_cluster</code> if we use all rows. However to save memory &amp; time, I divided train data into 100 subdata based on <code>hotel_cluster</code>. This way each subdata only has around 100K-400K rows (although several subdata has 600K, 700K, 1M rows), instead of 30M+ (full). So each subdata only has &quot;one unique value&quot; of <code>hotel_cluster</code>, hence I tried <code>is_booking</code> as target, which is not good in that my experiment. Thank you. </p>",
      "rawMarkdown": "Icaro, \r\n\r\nIt should be `hotel_cluster` if we use all rows. However to save memory & time, I divided train data into 100 subdata based on `hotel_cluster`. This way each subdata only has around 100K-400K rows (although several subdata has 600K, 700K, 1M rows), instead of 30M+ (full). So each subdata only has \"one unique value\" of `hotel_cluster`, hence I tried `is_booking` as target, which is not good in that my experiment. Thank you.",
      "votes": null
    },
    {
      "id": "122896",
      "postDate": "06/08/2016 04:29:58",
      "content": "<p>You should have a mix of clusters in your training set and train each forest to detect it's specific cluster.</p>\n\n<p>Now, you may want to skew the data towards the cluster you're actually training for (so there's more than 10% of positive data in the data set), but that's a separate issue.</p>",
      "rawMarkdown": "You should have a mix of clusters in your training set and train each forest to detect it's specific cluster.\r\n\r\nNow, you may want to skew the data towards the cluster you're actually training for (so there's more than 10% of positive data in the data set), but that's a separate issue.",
      "votes": null
    },
    {
      "id": "246565",
      "postDate": "11/21/2017 10:00:53",
      "content": "<p>Hi,\nI am solving a similar problem.  Finding the age group to which a user belongs to.\nWith reference to above example, consider:-</p>\n\n<p>\"User\" (maps to=&gt;) \"booking\".\n\"Age group\" (maps to=&gt;) \"cluster\".</p>\n\n<p>I have 5 age groups with ids 1,2,3,4,5.\nI have users u1,u2....</p>\n\n<p>At last, for every user's row, i have 5 probabilities one for each age group.\nI take the top 3 age groups for every user.</p>\n\n<p>Now what?\nWhat is the prediction of age group for each user? ( I have 3 top values. Which should be picked up?)</p>",
      "rawMarkdown": "Hi,\nI am solving a similar problem.  Finding the age group to which a user belongs to.\nWith reference to above example, consider:-\n\n\"User\" (maps to=&gt;) \"booking\".\n\"Age group\" (maps to=&gt;) \"cluster\".\n\nI have 5 age groups with ids 1,2,3,4,5.\nI have users u1,u2....\n\nAt last, for every user's row, i have 5 probabilities one for each age group.\nI take the top 3 age groups for every user.\n\nNow what?\nWhat is the prediction of age group for each user? ( I have 3 top values. Which should be picked up?)",
      "votes": null
    },
    {
      "id": "254915",
      "postDate": "12/07/2017 21:27:13",
      "content": "<p>hey subhajit, I am creating 100 binarly logistic regression classifiers and then using the one with highest probability but I am getting less than 0.10 accuracy. Do you know which features I should exclude? Thank You.</p>",
      "rawMarkdown": "hey subhajit, I am creating 100 binarly logistic regression classifiers and then using the one with highest probability but I am getting less than 0.10 accuracy. Do you know which features I should exclude? Thank You.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 117295,
      "author_name": "mightybird",
      "author_url": "",
      "post_date": "04/28/2016 12:07:26",
      "content": "<p>It should be possible to solve this with logistic regression. However you need predictions (probabilities) for all 100 hotel_clusters. Then take the top 5 clusters for each line in test.</p>\n\n<p>Gerhard</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 117304,
      "author_name": "khaoticmind",
      "author_url": "",
      "post_date": "04/28/2016 12:33:58",
      "content": "<p>I fail to see how you can use regression on this problem...\nThe hotel cluster values are not ints, they are IDs, so if a hotel cluster is 1 it is not any closer to 2 than it is from 99.</p>\n\n<p>Also, you have to give five results (that depends on the probability of all clusters) what, again, regression can't do for you.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 117306,
      "author_name": "mightybird",
      "author_url": "",
      "post_date": "04/28/2016 13:11:13",
      "content": "<p>Hi,</p>\n\n<p>let me show you the way...</p>\n\n<p>Let's reduce the problem first. Let's say you want to predict if a booking belongs to cluster 0 or NOT. This can well be solved by LOGISTIC regression. Because it will give you probabilities for just that question. </p>\n\n<p>Now you do this for every cluster value, so 100 times. You end up with 100 probabilities per test row. Each indicating the probablity for this row to result in cluster x. Now you take the 5 results (clusters) with the highest probabilities and submit them.</p>\n\n<p>Gerhard</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 117310,
      "author_name": "laurae2",
      "author_url": "",
      "post_date": "04/28/2016 13:26:05",
      "content": "<p>[quote=Antonio Augusto Santos;117304]</p>\n\n<p>I fail to see how you can use regression on this problem...</p>\n\n<p>[/quote]</p>\n\n<p>Logistic Regression = Classification</p>\n\n<p>If you use a two-class logistic regression, you will need 100 models: ID=1 against all others, ID=2 against all others, ID=3 against all others, etc. Once you have the probabilities for each ID, you rank them by how high the probability is.\nIf you use a multi-class logistic regression, 1 model is sufficient, you only need to sort ranks afterwards.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 117364,
      "author_name": "khaoticmind",
      "author_url": "",
      "post_date": "04/28/2016 18:16:28",
      "content": "<p>Sorry guys! My mistake :)\nI faield to see the LOGISTIC before the REGRESSION lol! </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 122545,
      "author_name": "mikebak",
      "author_url": "",
      "post_date": "06/05/2016 01:06:53",
      "content": "<p>Hi,</p>\n\n<p>I took this thread's advice and made 100 classification models.  I now have an array with 2.5 million rows (for each data item in the test sample) and 100 columns with the classification prediction numbers (for each hotel cluster).  Can anyone please suggest a way to get the top 5 clusters?  You don't have to code this for me, but any help would be appreciated.  I tried transposing the numpy matrix and sorting by one column (which represents a data item in the tranposed matrix) at a time, but this gives me the top 5 probabilities, not the hotel cluster numbers that correspond to those top five probabilities.</p>\n\n<p>Thank you</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 122549,
      "author_name": "jeffhebert",
      "author_url": "",
      "post_date": "06/05/2016 02:39:04",
      "content": "<p>Each column should represent the likelihood of a booking being in that hotel class. So you just need to find the top five values for each row, and return the associated hotel cluster.</p>\n\n<p>Good luck,\nJeff</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 122550,
      "author_name": "mandalsubhajit",
      "author_url": "",
      "post_date": "06/05/2016 02:50:37",
      "content": "<p>May be this thread will help you.</p>\n\n<p><a href=\"https://www.kaggle.com/c/expedia-hotel-recommendations/forums/t/20556/map5-function-or-eval-metric/120977\">https://www.kaggle.com/c/expedia-hotel-recommendations/forums/t/20556/map5-function-or-eval-metric/120977</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 254915,
          "author_name": "vyanktesh",
          "author_url": "",
          "post_date": "12/07/2017 21:27:13",
          "content": "<p>hey subhajit, I am creating 100 binarly logistic regression classifiers and then using the one with highest probability but I am getting less than 0.10 accuracy. Do you know which features I should exclude? Thank You.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 122551,
      "author_name": "mikebak",
      "author_url": "",
      "post_date": "06/05/2016 02:58:05",
      "content": "<p>JeffH, </p>\n\n<p>Thank you.   I have a numpy matrix with the columns like you said.   I understand that I need to find the top five values for each row, but my problem is that I am having trouble writing code to do that.   I can get the top value, but not the top five. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 122552,
      "author_name": "mikebak",
      "author_url": "",
      "post_date": "06/05/2016 03:06:33",
      "content": "<p>Subhajit, </p>\n\n<p>Sorry, I forgot to mention that I am using python.   I do not see this in the link you provided. </p>\n\n<p>Thank you though, \nMike</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 122560,
      "author_name": "mandalsubhajit",
      "author_url": "",
      "post_date": "06/05/2016 04:08:12",
      "content": "<p>Look for dune_dweller's post in the thread. Anyway, let me repost it here for you:</p>\n\n<pre><code>def map5eval(preds, dtrain):\n    actual = dtrain.get_label()\n    predicted = preds.argsort(axis=1)[:,-np.arange(1,6)]\n    metric = 0.\n    for i in range(5):\n        metric += np.sum(actual==predicted[:,i])/(i+1)\n    metric /= actual.shape[0]\n    return 'MAP@5', metric\n</code></pre>",
      "votes": null,
      "replies": []
    },
    {
      "id": 122563,
      "author_name": "beedata",
      "author_url": "",
      "post_date": "06/05/2016 04:41:01",
      "content": "<p>@Subhajit Mandal   Did you use xgb result in your best submission?  </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 122623,
      "author_name": "mikebak",
      "author_url": "",
      "post_date": "06/05/2016 21:57:04",
      "content": "<p>Subhajit,</p>\n\n<p>Thank you so much!  That worked.  I can't tell you how good that feels to try this for 3 hours and not be able to come up with a solution and then to discover this .argsort method you shared and it works!</p>\n\n<p>Thank you again,\nMike</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 122627,
      "author_name": "mandalsubhajit",
      "author_url": "",
      "post_date": "06/05/2016 23:24:21",
      "content": "<p>@FengLi, yes I used xgb.</p>\n\n<p>@Mike, you're welcome... although the code is not mine.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 122629,
      "author_name": "aldente",
      "author_url": "",
      "post_date": "06/05/2016 23:56:21",
      "content": "<p>I tried this method with 100 models, as explained above by Mighty, Laurae and Subhajit. I used XGBoost (binary logistic), as target is <code>is_booking</code>, so it tries to get likelihood of booking vs clicking for each <code>hotel_cluster</code>. I finished it in 5 days. </p>\n\n<p>All <code>destinations.csv</code> features are merged. To get top 5 for each row, I tried (1) rank based and (2) scale [0..1] based sorting. Logloss is around 0.19 for each xgb model. The result is very disappointing, <strong>under 0.1</strong> public LB score (before adding leakage result). </p>\n\n<p>Why? Is the target (<code>is_booking</code>) wrong? </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 122645,
      "author_name": "beedata",
      "author_url": "",
      "post_date": "06/06/2016 03:15:01",
      "content": "<p>@Subhajit Mandal Did you use 'multi:softprob' objective function? </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 122886,
      "author_name": "ibombonato",
      "author_url": "",
      "post_date": "06/08/2016 02:35:04",
      "content": "<p>[quote=aldente;122629]\nIs the target (<code>is_booking</code>) wrong? \n[/quote]</p>\n\n<p>The target is the variable &quot;hotel_cluster&quot;</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 122888,
      "author_name": "aldente",
      "author_url": "",
      "post_date": "06/08/2016 02:49:57",
      "content": "<p>@Icaro, </p>\n\n<p>It should be <code>hotel_cluster</code> if we use all rows. However to save memory &amp; time, I divided train data into 100 subdata based on <code>hotel_cluster</code>. This way each subdata only has around 100K-400K rows (although several subdata has 600K, 700K, 1M rows), instead of 30M+ (full). So each subdata only has &quot;one unique value&quot; of <code>hotel_cluster</code>, hence I tried <code>is_booking</code> as target, which is not good in that my experiment. Thank you. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 122896,
      "author_name": "mfagerlund",
      "author_url": "",
      "post_date": "06/08/2016 04:29:58",
      "content": "<p>You should have a mix of clusters in your training set and train each forest to detect it's specific cluster.</p>\n\n<p>Now, you may want to skew the data towards the cluster you're actually training for (so there's more than 10% of positive data in the data set), but that's a separate issue.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 246565,
      "author_name": "kiransadani",
      "author_url": "",
      "post_date": "11/21/2017 10:00:53",
      "content": "<p>Hi,\nI am solving a similar problem.  Finding the age group to which a user belongs to.\nWith reference to above example, consider:-</p>\n\n<p>\"User\" (maps to=&gt;) \"booking\".\n\"Age group\" (maps to=&gt;) \"cluster\".</p>\n\n<p>I have 5 age groups with ids 1,2,3,4,5.\nI have users u1,u2....</p>\n\n<p>At last, for every user's row, i have 5 probabilities one for each age group.\nI take the top 3 age groups for every user.</p>\n\n<p>Now what?\nWhat is the prediction of age group for each user? ( I have 3 top values. Which should be picked up?)</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "117275": "Is this possible to train a logistic regression or any other classification model to get a good result? Or just make some regulations...(for example ,most popular..)",
    "117295": "It should be possible to solve this with logistic regression. However you need predictions (probabilities) for all 100 hotel_clusters. Then take the top 5 clusters for each line in test.\r\n\r\nGerhard",
    "117304": "I fail to see how you can use regression on this problem...\r\nThe hotel cluster values are not ints, they are IDs, so if a hotel cluster is 1 it is not any closer to 2 than it is from 99.\r\n\r\nAlso, you have to give five results (that depends on the probability of all clusters) what, again, regression can't do for you.",
    "117306": "Hi,\r\n\r\nlet me show you the way...\r\n\r\nLet's reduce the problem first. Let's say you want to predict if a booking belongs to cluster 0 or NOT. This can well be solved by LOGISTIC regression. Because it will give you probabilities for just that question. \r\n\r\nNow you do this for every cluster value, so 100 times. You end up with 100 probabilities per test row. Each indicating the probablity for this row to result in cluster x. Now you take the 5 results (clusters) with the highest probabilities and submit them.\r\n\r\nGerhard",
    "117310": "[quote=Antonio Augusto Santos;117304]\r\n\r\nI fail to see how you can use regression on this problem...\r\n\r\n[/quote]\r\n\r\nLogistic Regression = Classification\r\n\r\nIf you use a two-class logistic regression, you will need 100 models: ID=1 against all others, ID=2 against all others, ID=3 against all others, etc. Once you have the probabilities for each ID, you rank them by how high the probability is.\r\nIf you use a multi-class logistic regression, 1 model is sufficient, you only need to sort ranks afterwards.",
    "117364": "Sorry guys! My mistake :)\r\nI faield to see the LOGISTIC before the REGRESSION lol!",
    "122545": "Hi,\r\n\r\nI took this thread's advice and made 100 classification models.  I now have an array with 2.5 million rows (for each data item in the test sample) and 100 columns with the classification prediction numbers (for each hotel cluster).  Can anyone please suggest a way to get the top 5 clusters?  You don't have to code this for me, but any help would be appreciated.  I tried transposing the numpy matrix and sorting by one column (which represents a data item in the tranposed matrix) at a time, but this gives me the top 5 probabilities, not the hotel cluster numbers that correspond to those top five probabilities.\r\n\r\nThank you",
    "122549": "Each column should represent the likelihood of a booking being in that hotel class. So you just need to find the top five values for each row, and return the associated hotel cluster.\r\n\r\nGood luck,\r\nJeff",
    "122550": "May be this thread will help you.\r\n\r\nhttps://www.kaggle.com/c/expedia-hotel-recommendations/forums/t/20556/map5-function-or-eval-metric/120977",
    "122551": "JeffH, \r\n\r\nThank you.   I have a numpy matrix with the columns like you said.   I understand that I need to find the top five values for each row, but my problem is that I am having trouble writing code to do that.   I can get the top value, but not the top five.",
    "122552": "Subhajit, \r\n\r\nSorry, I forgot to mention that I am using python.   I do not see this in the link you provided. \r\n\r\nThank you though, \r\nMike",
    "122560": "Look for dune_dweller's post in the thread. Anyway, let me repost it here for you:\r\n\r\n    def map5eval(preds, dtrain):\r\n        actual = dtrain.get_label()\r\n        predicted = preds.argsort(axis=1)[:,-np.arange(1,6)]\r\n        metric = 0.\r\n        for i in range(5):\r\n            metric += np.sum(actual==predicted[:,i])/(i+1)\r\n        metric /= actual.shape[0]\r\n        return 'MAP@5', metric",
    "122563": "Subhajit Mandal   Did you use xgb result in your best submission?",
    "122623": "Subhajit,\r\n\r\nThank you so much!  That worked.  I can't tell you how good that feels to try this for 3 hours and not be able to come up with a solution and then to discover this .argsort method you shared and it works!\r\n\r\nThank you again,\r\nMike",
    "122627": "FengLi, yes I used xgb.\r\n\r\n@Mike, you're welcome... although the code is not mine.",
    "122629": "I tried this method with 100 models, as explained above by Mighty, Laurae and Subhajit. I used XGBoost (binary logistic), as target is `is_booking`, so it tries to get likelihood of booking vs clicking for each `hotel_cluster`. I finished it in 5 days. \r\n\r\nAll `destinations.csv` features are merged. To get top 5 for each row, I tried (1) rank based and (2) scale [0..1] based sorting. Logloss is around 0.19 for each xgb model. The result is very disappointing, **under 0.1** public LB score (before adding leakage result). \r\n\r\nWhy? Is the target (`is_booking`) wrong?",
    "122645": "Subhajit Mandal Did you use 'multi:softprob' objective function?",
    "122886": "[quote=aldente;122629]\r\nIs the target (`is_booking`) wrong? \r\n[/quote]\r\n\r\nThe target is the variable \"hotel_cluster\"",
    "122888": "Icaro, \r\n\r\nIt should be `hotel_cluster` if we use all rows. However to save memory & time, I divided train data into 100 subdata based on `hotel_cluster`. This way each subdata only has around 100K-400K rows (although several subdata has 600K, 700K, 1M rows), instead of 30M+ (full). So each subdata only has \"one unique value\" of `hotel_cluster`, hence I tried `is_booking` as target, which is not good in that my experiment. Thank you.",
    "122896": "You should have a mix of clusters in your training set and train each forest to detect it's specific cluster.\r\n\r\nNow, you may want to skew the data towards the cluster you're actually training for (so there's more than 10% of positive data in the data set), but that's a separate issue.",
    "246565": "Hi,\nI am solving a similar problem.  Finding the age group to which a user belongs to.\nWith reference to above example, consider:-\n\n\"User\" (maps to=&gt;) \"booking\".\n\"Age group\" (maps to=&gt;) \"cluster\".\n\nI have 5 age groups with ids 1,2,3,4,5.\nI have users u1,u2....\n\nAt last, for every user's row, i have 5 probabilities one for each age group.\nI take the top 3 age groups for every user.\n\nNow what?\nWhat is the prediction of age group for each user? ( I have 3 top values. Which should be picked up?)",
    "254915": "hey subhajit, I am creating 100 binarly logistic regression classifiers and then using the one with highest probability but I am getting less than 0.10 accuracy. Do you know which features I should exclude? Thank You."
  },
  "source": "meta"
}