{
  "id": 20815,
  "title": "Does Naive Bayes work here?",
  "url": "/competitions/expedia-hotel-recommendations/discussion/20815",
  "author_name": "",
  "post_date": "2016-05-08T23:58:05.970Z",
  "votes": 8,
  "comment_count": 3,
  "views": 1615,
  "content": "<p>I wrote a fast version of Naive Bayes that requires pypy for the base function then I wrapped it for R use (with data.table). I'd like to post it here to see if anyone is using NB and a code check. The probabilities aren't normalized, but I think the code is correct (it may not be). It's fairly fast and should work on the full train data and test data in minutes.</p>\n\n<p>R does not have a good Naive Bayes. They all seem to memory leak or freeze on me. My function requires the input train data.frame to have the first column as target (0...numclasses-1) and the other columns as features (use correct header names), and the test data.frame to have the features with correct header names.</p>",
  "messages": [
    {
      "id": "119299",
      "postDate": "05/08/2016 23:58:05",
      "content": "<p>I wrote a fast version of Naive Bayes that requires pypy for the base function then I wrapped it for R use (with data.table). I'd like to post it here to see if anyone is using NB and a code check. The probabilities aren't normalized, but I think the code is correct (it may not be). It's fairly fast and should work on the full train data and test data in minutes.</p>\n\n<p>R does not have a good Naive Bayes. They all seem to memory leak or freeze on me. My function requires the input train data.frame to have the first column as target (0...numclasses-1) and the other columns as features (use correct header names), and the test data.frame to have the features with correct header names.</p>",
      "rawMarkdown": "I wrote a fast version of Naive Bayes that requires pypy for the base function then I wrapped it for R use (with data.table). I'd like to post it here to see if anyone is using NB and a code check. The probabilities aren't normalized, but I think the code is correct (it may not be). It's fairly fast and should work on the full train data and test data in minutes.\r\n\r\nR does not have a good Naive Bayes. They all seem to memory leak or freeze on me. My function requires the input train data.frame to have the first column as target (0...numclasses-1) and the other columns as features (use correct header names), and the test data.frame to have the features with correct header names.",
      "votes": null
    },
    {
      "id": "119309",
      "postDate": "05/09/2016 04:36:49",
      "content": "<p>My first scoring model is one I deemed hierarchical, though in fact it looks a lot like some of the Python scripts that have been posted.  However, my count-and-sort had some particular nuances.  That got me to 0.487 or so.  In the process, it occurred to me just how much what I had come up with was like NB.</p>\n\n<p>I am using the e1071 package in R, the naiveBayes function (different from NaiveBayes) and, frankly, I think it flies.  However, even varying the feature selection and using Laplace smoothing at l=1, 2, and 3, I haven't yet submitted any worthwhile scores.  I tinkered with eps and threshold in naiveBayes, but they didn't result in any train/validation improvements either.   I'm due for another round of trying NB, and on my 32GB laptop, I'm not hesitant to work in R Studio with the e1071 implementation.  In fact, I did build models on the entire train data set!</p>",
      "rawMarkdown": "My first scoring model is one I deemed hierarchical, though in fact it looks a lot like some of the Python scripts that have been posted.  However, my count-and-sort had some particular nuances.  That got me to 0.487 or so.  In the process, it occurred to me just how much what I had come up with was like NB.\r\n\r\nI am using the e1071 package in R, the naiveBayes function (different from NaiveBayes) and, frankly, I think it flies.  However, even varying the feature selection and using Laplace smoothing at l=1, 2, and 3, I haven't yet submitted any worthwhile scores.  I tinkered with eps and threshold in naiveBayes, but they didn't result in any train/validation improvements either.   I'm due for another round of trying NB, and on my 32GB laptop, I'm not hesitant to work in R Studio with the e1071 implementation.  In fact, I did build models on the entire train data set!",
      "votes": null
    },
    {
      "id": "119353",
      "postDate": "05/09/2016 12:41:26",
      "content": "<p>Thats really cool Mike, thanks for sharing ! I will try it out. (p.s. good to learn how R and python can interact)</p>",
      "rawMarkdown": "Thats really cool Mike, thanks for sharing ! I will try it out. (p.s. good to learn how R and python can interact)",
      "votes": null
    },
    {
      "id": "121036",
      "postDate": "05/23/2016 05:45:43",
      "content": "<p>I have only gotten ~42.5 LB with partial data, but 30.8 MAP local validation.</p>\n\n<p><a href=\"https://www.kaggle.com/yukatherin/expedia-hotel-recommendations/naive-bayes-with-countvectorizer/run/244962\">https://www.kaggle.com/yukatherin/expedia-hotel-recommendations/naive-bayes-with-countvectorizer/run/244962</a></p>",
      "rawMarkdown": "I have only gotten ~42.5 LB with partial data, but 30.8 MAP local validation.\r\n\r\nhttps://www.kaggle.com/yukatherin/expedia-hotel-recommendations/naive-bayes-with-countvectorizer/run/244962",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 119309,
      "author_name": "siliconvalley",
      "author_url": "",
      "post_date": "05/09/2016 04:36:49",
      "content": "<p>My first scoring model is one I deemed hierarchical, though in fact it looks a lot like some of the Python scripts that have been posted.  However, my count-and-sort had some particular nuances.  That got me to 0.487 or so.  In the process, it occurred to me just how much what I had come up with was like NB.</p>\n\n<p>I am using the e1071 package in R, the naiveBayes function (different from NaiveBayes) and, frankly, I think it flies.  However, even varying the feature selection and using Laplace smoothing at l=1, 2, and 3, I haven't yet submitted any worthwhile scores.  I tinkered with eps and threshold in naiveBayes, but they didn't result in any train/validation improvements either.   I'm due for another round of trying NB, and on my 32GB laptop, I'm not hesitant to work in R Studio with the e1071 implementation.  In fact, I did build models on the entire train data set!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 119353,
      "author_name": "darraghdog",
      "author_url": "",
      "post_date": "05/09/2016 12:41:26",
      "content": "<p>Thats really cool Mike, thanks for sharing ! I will try it out. (p.s. good to learn how R and python can interact)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 121036,
      "author_name": "yukatherin",
      "author_url": "",
      "post_date": "05/23/2016 05:45:43",
      "content": "<p>I have only gotten ~42.5 LB with partial data, but 30.8 MAP local validation.</p>\n\n<p><a href=\"https://www.kaggle.com/yukatherin/expedia-hotel-recommendations/naive-bayes-with-countvectorizer/run/244962\">https://www.kaggle.com/yukatherin/expedia-hotel-recommendations/naive-bayes-with-countvectorizer/run/244962</a></p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "119299": "I wrote a fast version of Naive Bayes that requires pypy for the base function then I wrapped it for R use (with data.table). I'd like to post it here to see if anyone is using NB and a code check. The probabilities aren't normalized, but I think the code is correct (it may not be). It's fairly fast and should work on the full train data and test data in minutes.\r\n\r\nR does not have a good Naive Bayes. They all seem to memory leak or freeze on me. My function requires the input train data.frame to have the first column as target (0...numclasses-1) and the other columns as features (use correct header names), and the test data.frame to have the features with correct header names.",
    "119309": "My first scoring model is one I deemed hierarchical, though in fact it looks a lot like some of the Python scripts that have been posted.  However, my count-and-sort had some particular nuances.  That got me to 0.487 or so.  In the process, it occurred to me just how much what I had come up with was like NB.\r\n\r\nI am using the e1071 package in R, the naiveBayes function (different from NaiveBayes) and, frankly, I think it flies.  However, even varying the feature selection and using Laplace smoothing at l=1, 2, and 3, I haven't yet submitted any worthwhile scores.  I tinkered with eps and threshold in naiveBayes, but they didn't result in any train/validation improvements either.   I'm due for another round of trying NB, and on my 32GB laptop, I'm not hesitant to work in R Studio with the e1071 implementation.  In fact, I did build models on the entire train data set!",
    "119353": "Thats really cool Mike, thanks for sharing ! I will try it out. (p.s. good to learn how R and python can interact)",
    "121036": "I have only gotten ~42.5 LB with partial data, but 30.8 MAP local validation.\r\n\r\nhttps://www.kaggle.com/yukatherin/expedia-hotel-recommendations/naive-bayes-with-countvectorizer/run/244962"
  },
  "source": "meta"
}