{
  "id": 21360,
  "title": "[R] Naive Bayes (e1071 package) - Return not only the more probable hotel_cluster but (for example) the 3 more probable",
  "url": "/competitions/expedia-hotel-recommendations/discussion/21360",
  "author_name": "",
  "post_date": "2016-06-01T11:09:32.777Z",
  "votes": null,
  "comment_count": 10,
  "views": 3496,
  "content": "<p>Hello Kaggle !</p>\n\n<p>I was wondering if you knew how to return not only the more probable hotel_cluster using naiveBayes in R (from package e1071) . In fact, I would like to make a submission with the 3 more probable hotel_cluster to see if it improves my score.</p>\n\n<p>Is there an option to add in my code, or is it impossible in R with the e1071 implementation :\n&quot;\n<strong>classifier&lt;-naiveBayes(as.factor(hotel_cluster)~.,data=train)\nsubmission$hotel_cluster&lt;-predict(classifier, test)</strong>\n&quot;</p>\n\n<p>Thanks :P</p>",
  "messages": [
    {
      "id": "122115",
      "postDate": "06/01/2016 11:09:32",
      "content": "<p>Hello Kaggle !</p>\n\n<p>I was wondering if you knew how to return not only the more probable hotel_cluster using naiveBayes in R (from package e1071) . In fact, I would like to make a submission with the 3 more probable hotel_cluster to see if it improves my score.</p>\n\n<p>Is there an option to add in my code, or is it impossible in R with the e1071 implementation :\n&quot;\n<strong>classifier&lt;-naiveBayes(as.factor(hotel_cluster)~.,data=train)\nsubmission$hotel_cluster&lt;-predict(classifier, test)</strong>\n&quot;</p>\n\n<p>Thanks :P</p>",
      "rawMarkdown": "Hello Kaggle !\r\n\r\nI was wondering if you knew how to return not only the more probable hotel_cluster using naiveBayes in R (from package e1071) . In fact, I would like to make a submission with the 3 more probable hotel_cluster to see if it improves my score.\r\n\r\n Is there an option to add in my code, or is it impossible in R with the e1071 implementation :\r\n\"\r\n**classifier<-naiveBayes(as.factor(hotel_cluster)~.,data=train)\r\nsubmission$hotel_cluster<-predict(classifier, test)**\r\n\"\r\n\r\nThanks :P",
      "votes": null
    },
    {
      "id": "122122",
      "postDate": "06/01/2016 12:23:01",
      "content": "<p>Hi meurett,</p>\n\n<p>Reading the doc at <a href=\"https://cran.r-project.org/web/packages/e1071/e1071.pdf\">https://cran.r-project.org/web/packages/e1071/e1071.pdf</a>, the method takes a 'type' parameter:</p>\n\n<p>type      If &quot;raw&quot;, the conditional a-posterior probabilities for each class are returned, and the class with maximal probability else.</p>\n\n<p>Looks like this is what you needed?</p>\n\n<p>Just looking at using NB myself. But currently with my own hand-coded version, the LB performance was a notch lower than a basic rule based model. And I still don't see a good way to blend them...</p>",
      "rawMarkdown": "Hi meurett,\r\n\r\nReading the doc at https://cran.r-project.org/web/packages/e1071/e1071.pdf, the method takes a 'type' parameter:\r\n\r\ntype      If \"raw\", the conditional a-posterior probabilities for each class are returned, and the class with maximal probability else.\r\n\r\nLooks like this is what you needed?\r\n\r\n\r\nJust looking at using NB myself. But currently with my own hand-coded version, the LB performance was a notch lower than a basic rule based model. And I still don't see a good way to blend them...",
      "votes": null
    },
    {
      "id": "122146",
      "postDate": "06/01/2016 15:47:43",
      "content": "<p>Yup, exactly what I was looking for :)\nIf it can help someone, this is how I managed getting the 5 most probable ones :</p>\n\n<p>&quot;\nclassifier&lt;-naiveBayes(as.factor(hotel_cluster)~.,data=train)</p>\n\n<p>for (i in 1:nrow(test)) {</p>\n\n<pre><code>res&lt;-rev((order(predict(classifier, test[i], type=&quot;raw&quot;))-1)[96:100])\nsubmission$hotel_cluster[i]&lt;-paste(res[1],res[2],res[3],res[4],res[5],sep=&quot; &quot;)\n</code></pre>\n\n<p>} &quot;</p>\n\n<p>I know it is possible to improve the completion of submission data frame: using &quot;apply&quot; instead of using a &quot;for&quot;, but I need to understand how to use &quot;apply&quot; better atm :/</p>",
      "rawMarkdown": "Yup, exactly what I was looking for :)\r\nIf it can help someone, this is how I managed getting the 5 most probable ones :\r\n\r\n\"\r\nclassifier<-naiveBayes(as.factor(hotel_cluster)~.,data=train)\r\n\r\nfor (i in 1:nrow(test)) {\r\n\r\n\tres<-rev((order(predict(classifier, test[i], type=\"raw\"))-1)[96:100])\r\n\tsubmission$hotel_cluster[i]<-paste(res[1],res[2],res[3],res[4],res[5],sep=\" \")\r\n\r\n} \"\r\n\r\nI know it is possible to improve the completion of submission data frame: using \"apply\" instead of using a \"for\", but I need to understand how to use \"apply\" better atm :/",
      "votes": null
    },
    {
      "id": "122174",
      "postDate": "06/01/2016 19:00:52",
      "content": "<p>I implemented a solution using the full train data set - as Can said, use type=&quot;raw&quot; and you've got the idea. You'll need to consider laplace smoothing and eps and threshold settings.  I never scored higher than 0.47 LB using any type of NB, so I moved on to other approaches.  Good luck!</p>",
      "rawMarkdown": "I implemented a solution using the full train data set - as Can said, use type=\"raw\" and you've got the idea. You'll need to consider laplace smoothing and eps and threshold settings.  I never scored higher than 0.47 LB using any type of NB, so I moved on to other approaches.  Good luck!",
      "votes": null
    },
    {
      "id": "122346",
      "postDate": "06/03/2016 07:29:52",
      "content": "<p>@eipiplus1 I was wondering how you succeded in performing that high with a that simple model ? I did compute my algorithm multiple times but I only got ~0.06-0.07 at LB. I send my personnal code with this message, would do mind give a look at it and give some tips and ideas on what I may improve ? (option, pre-calculate some things, ect.) Thanks ! o/</p>",
      "rawMarkdown": "eipiplus1 I was wondering how you succeded in performing that high with a that simple model ? I did compute my algorithm multiple times but I only got ~0.06-0.07 at LB. I send my personnal code with this message, would do mind give a look at it and give some tips and ideas on what I may improve ? (option, pre-calculate some things, ect.) Thanks ! o/",
      "votes": null
    },
    {
      "id": "122348",
      "postDate": "06/03/2016 07:49:50",
      "content": "<p>I read through your code.  Can you give me a synopsis of the features which remain when you train the model via naiveBayes(hotel_cluster~., ...)?  I see you set many to NULL, added some other columns, etc., and my gut reaction after reading your code for 2 minutes is that there are too many date and other features.</p>",
      "rawMarkdown": "I read through your code.  Can you give me a synopsis of the features which remain when you train the model via naiveBayes(hotel_cluster~., ...)?  I see you set many to NULL, added some other columns, etc., and my gut reaction after reading your code for 2 minutes is that there are too many date and other features.",
      "votes": null
    },
    {
      "id": "122350",
      "postDate": "06/03/2016 08:16:00",
      "content": "<p>All these as.date were here to convert string date into integer, and to create a variable (dayofyear_*) that resume both day and month information into one; also added the duration of a trip.</p>\n\n<p>So, to sum it up, at the end of all these calculation it remains 25 variables that you can see in my linked .png</p>\n\n<p>To sum it up :</p>\n\n<ul>\n<li>first 18 ones are from the data</li>\n<li><strong>day/month_action</strong> : integer value of the day(1-31)/month(1-12) of click and booking</li>\n<li><strong>day/month_ci</strong> : same for check-in</li>\n<li><strong>day_of_year_action/ci</strong> : eg. March 3rd = 31 + 28 + 3 = 62</li>\n</ul>",
      "rawMarkdown": "All these as.date were here to convert string date into integer, and to create a variable (dayofyear_*) that resume both day and month information into one; also added the duration of a trip.\r\n\r\nSo, to sum it up, at the end of all these calculation it remains 25 variables that you can see in my linked .png\r\n\r\nTo sum it up :\r\n\r\n * first 18 ones are from the data\r\n * **day/month_action** : integer value of the day(1-31)/month(1-12) of click and booking\r\n * **day/month_ci** : same for check-in\r\n * **day_of_year_action/ci** : eg. March 3rd = 31 + 28 + 3 = 62",
      "votes": null
    },
    {
      "id": "122351",
      "postDate": "06/03/2016 08:19:23",
      "content": "<p>@meurett, my quick suggestion will be to use fewer features to start with. Even if you use only srch_destination_id the score will be higher than this. It will get you the most frequent hotel_cluster for each destination.</p>",
      "rawMarkdown": "meurett, my quick suggestion will be to use fewer features to start with. Even if you use only srch_destination_id the score will be higher than this. It will get you the most frequent hotel_cluster for each destination.",
      "votes": null
    },
    {
      "id": "122352",
      "postDate": "06/03/2016 08:31:09",
      "content": "<p>To my view, it's too many features, agreeing with Can.  Also (1) too many date-oriented features, and (2) I would recommend binning many of those features &quot;native&quot; to the training data set.  For example, look at a plot of days(srch_co - srch_ci) - see anything interesting there?  :-)</p>",
      "rawMarkdown": "To my view, it's too many features, agreeing with Can.  Also (1) too many date-oriented features, and (2) I would recommend binning many of those features \"native\" to the training data set.  For example, look at a plot of days(srch_co - srch_ci) - see anything interesting there?  :-)",
      "votes": null
    },
    {
      "id": "122354",
      "postDate": "06/03/2016 08:47:08",
      "content": "<p>Ok, will try with less features and see how it goes, thanks !</p>\n\n<p>Another question, directly related to my post on loop :</p>\n\n<p>I'm new with R loop with &quot;apply&quot; and I did try a lot of stuff that didn't work (R added all my hotel_cluster integer and put this sum result in all raws, etc.). The fact is that I know I spend (waste :/) most of my time completing the submission DF because either I'm using index on DF (for i in 1 : nrow(test) : submission$hotel_cluster[i]&lt;- ...) which is very slow or an &quot;apply&quot;/&quot;append&quot; (as in my .txt) which is very slow too. To give you an idea, I spend like 10 min max. calculating the classifier but up to 8h to fill the submission DF : really insane in my opinion. </p>\n\n<p>So, do you know how to do it efficiently in R ?</p>",
      "rawMarkdown": "Ok, will try with less features and see how it goes, thanks !\r\n\r\nAnother question, directly related to my post on loop :\r\n\r\n I'm new with R loop with \"apply\" and I did try a lot of stuff that didn't work (R added all my hotel_cluster integer and put this sum result in all raws, etc.). The fact is that I know I spend (waste :/) most of my time completing the submission DF because either I'm using index on DF (for i in 1 : nrow(test) : submission$hotel_cluster[i]<- ...) which is very slow or an \"apply\"/\"append\" (as in my .txt) which is very slow too. To give you an idea, I spend like 10 min max. calculating the classifier but up to 8h to fill the submission DF : really insane in my opinion. \r\n\r\nSo, do you know how to do it efficiently in R ?",
      "votes": null
    },
    {
      "id": "122694",
      "postDate": "06/06/2016 15:26:48",
      "content": "<p>Ok, I worked a lot on improving my score on LB according to what you , @Can Zheg and @eipiplus1, told me; atm I'm working with only the data related to the data leak with the following model : </p>\n\n<p>as.factor(hotel_cluster) ~ as.factor(user_location_city) + orig_destination_distance + as.factor(srch_destination_id)  + as.factor(hotel_market) + as.factor(hotel_country)</p>\n\n<p>More of that, I decided to implement the 3/17 ratio between click and booking to improve my model. The fact is that I keep having trouble reaching more than 0.08 on LB [0.7854 to be precise].\nSo, I was wondering if the way I'm preparing the data for the NB e1071 function was wrong ? Could anyone of you help me on this ?</p>",
      "rawMarkdown": "Ok, I worked a lot on improving my score on LB according to what you , @Can Zheg and @eipiplus1, told me; atm I'm working with only the data related to the data leak with the following model : \r\n\r\nas.factor(hotel_cluster) ~ as.factor(user_location_city) + orig_destination_distance + as.factor(srch_destination_id)  + as.factor(hotel_market) + as.factor(hotel_country)\r\n\r\nMore of that, I decided to implement the 3/17 ratio between click and booking to improve my model. The fact is that I keep having trouble reaching more than 0.08 on LB [0.7854 to be precise].\r\nSo, I was wondering if the way I'm preparing the data for the NB e1071 function was wrong ? Could anyone of you help me on this ?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 122122,
      "author_name": "kenzheng",
      "author_url": "",
      "post_date": "06/01/2016 12:23:01",
      "content": "<p>Hi meurett,</p>\n\n<p>Reading the doc at <a href=\"https://cran.r-project.org/web/packages/e1071/e1071.pdf\">https://cran.r-project.org/web/packages/e1071/e1071.pdf</a>, the method takes a 'type' parameter:</p>\n\n<p>type      If &quot;raw&quot;, the conditional a-posterior probabilities for each class are returned, and the class with maximal probability else.</p>\n\n<p>Looks like this is what you needed?</p>\n\n<p>Just looking at using NB myself. But currently with my own hand-coded version, the LB performance was a notch lower than a basic rule based model. And I still don't see a good way to blend them...</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 122146,
      "author_name": "meurett",
      "author_url": "",
      "post_date": "06/01/2016 15:47:43",
      "content": "<p>Yup, exactly what I was looking for :)\nIf it can help someone, this is how I managed getting the 5 most probable ones :</p>\n\n<p>&quot;\nclassifier&lt;-naiveBayes(as.factor(hotel_cluster)~.,data=train)</p>\n\n<p>for (i in 1:nrow(test)) {</p>\n\n<pre><code>res&lt;-rev((order(predict(classifier, test[i], type=&quot;raw&quot;))-1)[96:100])\nsubmission$hotel_cluster[i]&lt;-paste(res[1],res[2],res[3],res[4],res[5],sep=&quot; &quot;)\n</code></pre>\n\n<p>} &quot;</p>\n\n<p>I know it is possible to improve the completion of submission data frame: using &quot;apply&quot; instead of using a &quot;for&quot;, but I need to understand how to use &quot;apply&quot; better atm :/</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 122174,
      "author_name": "siliconvalley",
      "author_url": "",
      "post_date": "06/01/2016 19:00:52",
      "content": "<p>I implemented a solution using the full train data set - as Can said, use type=&quot;raw&quot; and you've got the idea. You'll need to consider laplace smoothing and eps and threshold settings.  I never scored higher than 0.47 LB using any type of NB, so I moved on to other approaches.  Good luck!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 122346,
      "author_name": "meurett",
      "author_url": "",
      "post_date": "06/03/2016 07:29:52",
      "content": "<p>@eipiplus1 I was wondering how you succeded in performing that high with a that simple model ? I did compute my algorithm multiple times but I only got ~0.06-0.07 at LB. I send my personnal code with this message, would do mind give a look at it and give some tips and ideas on what I may improve ? (option, pre-calculate some things, ect.) Thanks ! o/</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 122348,
      "author_name": "siliconvalley",
      "author_url": "",
      "post_date": "06/03/2016 07:49:50",
      "content": "<p>I read through your code.  Can you give me a synopsis of the features which remain when you train the model via naiveBayes(hotel_cluster~., ...)?  I see you set many to NULL, added some other columns, etc., and my gut reaction after reading your code for 2 minutes is that there are too many date and other features.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 122350,
      "author_name": "meurett",
      "author_url": "",
      "post_date": "06/03/2016 08:16:00",
      "content": "<p>All these as.date were here to convert string date into integer, and to create a variable (dayofyear_*) that resume both day and month information into one; also added the duration of a trip.</p>\n\n<p>So, to sum it up, at the end of all these calculation it remains 25 variables that you can see in my linked .png</p>\n\n<p>To sum it up :</p>\n\n<ul>\n<li>first 18 ones are from the data</li>\n<li><strong>day/month_action</strong> : integer value of the day(1-31)/month(1-12) of click and booking</li>\n<li><strong>day/month_ci</strong> : same for check-in</li>\n<li><strong>day_of_year_action/ci</strong> : eg. March 3rd = 31 + 28 + 3 = 62</li>\n</ul>",
      "votes": null,
      "replies": []
    },
    {
      "id": 122351,
      "author_name": "kenzheng",
      "author_url": "",
      "post_date": "06/03/2016 08:19:23",
      "content": "<p>@meurett, my quick suggestion will be to use fewer features to start with. Even if you use only srch_destination_id the score will be higher than this. It will get you the most frequent hotel_cluster for each destination.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 122352,
      "author_name": "siliconvalley",
      "author_url": "",
      "post_date": "06/03/2016 08:31:09",
      "content": "<p>To my view, it's too many features, agreeing with Can.  Also (1) too many date-oriented features, and (2) I would recommend binning many of those features &quot;native&quot; to the training data set.  For example, look at a plot of days(srch_co - srch_ci) - see anything interesting there?  :-)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 122354,
      "author_name": "meurett",
      "author_url": "",
      "post_date": "06/03/2016 08:47:08",
      "content": "<p>Ok, will try with less features and see how it goes, thanks !</p>\n\n<p>Another question, directly related to my post on loop :</p>\n\n<p>I'm new with R loop with &quot;apply&quot; and I did try a lot of stuff that didn't work (R added all my hotel_cluster integer and put this sum result in all raws, etc.). The fact is that I know I spend (waste :/) most of my time completing the submission DF because either I'm using index on DF (for i in 1 : nrow(test) : submission$hotel_cluster[i]&lt;- ...) which is very slow or an &quot;apply&quot;/&quot;append&quot; (as in my .txt) which is very slow too. To give you an idea, I spend like 10 min max. calculating the classifier but up to 8h to fill the submission DF : really insane in my opinion. </p>\n\n<p>So, do you know how to do it efficiently in R ?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 122694,
      "author_name": "meurett",
      "author_url": "",
      "post_date": "06/06/2016 15:26:48",
      "content": "<p>Ok, I worked a lot on improving my score on LB according to what you , @Can Zheg and @eipiplus1, told me; atm I'm working with only the data related to the data leak with the following model : </p>\n\n<p>as.factor(hotel_cluster) ~ as.factor(user_location_city) + orig_destination_distance + as.factor(srch_destination_id)  + as.factor(hotel_market) + as.factor(hotel_country)</p>\n\n<p>More of that, I decided to implement the 3/17 ratio between click and booking to improve my model. The fact is that I keep having trouble reaching more than 0.08 on LB [0.7854 to be precise].\nSo, I was wondering if the way I'm preparing the data for the NB e1071 function was wrong ? Could anyone of you help me on this ?</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "122115": "Hello Kaggle !\r\n\r\nI was wondering if you knew how to return not only the more probable hotel_cluster using naiveBayes in R (from package e1071) . In fact, I would like to make a submission with the 3 more probable hotel_cluster to see if it improves my score.\r\n\r\n Is there an option to add in my code, or is it impossible in R with the e1071 implementation :\r\n\"\r\n**classifier<-naiveBayes(as.factor(hotel_cluster)~.,data=train)\r\nsubmission$hotel_cluster<-predict(classifier, test)**\r\n\"\r\n\r\nThanks :P",
    "122122": "Hi meurett,\r\n\r\nReading the doc at https://cran.r-project.org/web/packages/e1071/e1071.pdf, the method takes a 'type' parameter:\r\n\r\ntype      If \"raw\", the conditional a-posterior probabilities for each class are returned, and the class with maximal probability else.\r\n\r\nLooks like this is what you needed?\r\n\r\n\r\nJust looking at using NB myself. But currently with my own hand-coded version, the LB performance was a notch lower than a basic rule based model. And I still don't see a good way to blend them...",
    "122146": "Yup, exactly what I was looking for :)\r\nIf it can help someone, this is how I managed getting the 5 most probable ones :\r\n\r\n\"\r\nclassifier<-naiveBayes(as.factor(hotel_cluster)~.,data=train)\r\n\r\nfor (i in 1:nrow(test)) {\r\n\r\n\tres<-rev((order(predict(classifier, test[i], type=\"raw\"))-1)[96:100])\r\n\tsubmission$hotel_cluster[i]<-paste(res[1],res[2],res[3],res[4],res[5],sep=\" \")\r\n\r\n} \"\r\n\r\nI know it is possible to improve the completion of submission data frame: using \"apply\" instead of using a \"for\", but I need to understand how to use \"apply\" better atm :/",
    "122174": "I implemented a solution using the full train data set - as Can said, use type=\"raw\" and you've got the idea. You'll need to consider laplace smoothing and eps and threshold settings.  I never scored higher than 0.47 LB using any type of NB, so I moved on to other approaches.  Good luck!",
    "122346": "eipiplus1 I was wondering how you succeded in performing that high with a that simple model ? I did compute my algorithm multiple times but I only got ~0.06-0.07 at LB. I send my personnal code with this message, would do mind give a look at it and give some tips and ideas on what I may improve ? (option, pre-calculate some things, ect.) Thanks ! o/",
    "122348": "I read through your code.  Can you give me a synopsis of the features which remain when you train the model via naiveBayes(hotel_cluster~., ...)?  I see you set many to NULL, added some other columns, etc., and my gut reaction after reading your code for 2 minutes is that there are too many date and other features.",
    "122350": "All these as.date were here to convert string date into integer, and to create a variable (dayofyear_*) that resume both day and month information into one; also added the duration of a trip.\r\n\r\nSo, to sum it up, at the end of all these calculation it remains 25 variables that you can see in my linked .png\r\n\r\nTo sum it up :\r\n\r\n * first 18 ones are from the data\r\n * **day/month_action** : integer value of the day(1-31)/month(1-12) of click and booking\r\n * **day/month_ci** : same for check-in\r\n * **day_of_year_action/ci** : eg. March 3rd = 31 + 28 + 3 = 62",
    "122351": "meurett, my quick suggestion will be to use fewer features to start with. Even if you use only srch_destination_id the score will be higher than this. It will get you the most frequent hotel_cluster for each destination.",
    "122352": "To my view, it's too many features, agreeing with Can.  Also (1) too many date-oriented features, and (2) I would recommend binning many of those features \"native\" to the training data set.  For example, look at a plot of days(srch_co - srch_ci) - see anything interesting there?  :-)",
    "122354": "Ok, will try with less features and see how it goes, thanks !\r\n\r\nAnother question, directly related to my post on loop :\r\n\r\n I'm new with R loop with \"apply\" and I did try a lot of stuff that didn't work (R added all my hotel_cluster integer and put this sum result in all raws, etc.). The fact is that I know I spend (waste :/) most of my time completing the submission DF because either I'm using index on DF (for i in 1 : nrow(test) : submission$hotel_cluster[i]<- ...) which is very slow or an \"apply\"/\"append\" (as in my .txt) which is very slow too. To give you an idea, I spend like 10 min max. calculating the classifier but up to 8h to fill the submission DF : really insane in my opinion. \r\n\r\nSo, do you know how to do it efficiently in R ?",
    "122694": "Ok, I worked a lot on improving my score on LB according to what you , @Can Zheg and @eipiplus1, told me; atm I'm working with only the data related to the data leak with the following model : \r\n\r\nas.factor(hotel_cluster) ~ as.factor(user_location_city) + orig_destination_distance + as.factor(srch_destination_id)  + as.factor(hotel_market) + as.factor(hotel_country)\r\n\r\nMore of that, I decided to implement the 3/17 ratio between click and booking to improve my model. The fact is that I keep having trouble reaching more than 0.08 on LB [0.7854 to be precise].\r\nSo, I was wondering if the way I'm preparing the data for the NB e1071 function was wrong ? Could anyone of you help me on this ?"
  },
  "source": "meta"
}