{
  "id": 24829,
  "title": "Getting lost in the Random Forest",
  "url": "/competitions/outbrain-click-prediction/discussion/24829",
  "author_name": "Bootes",
  "post_date": "2016-10-27T14:34:11.007000",
  "votes": 2,
  "comment_count": 0,
  "views": 382,
  "content": "<p>I have been running various ML classifiers on the training/test set, and found that Pythons RandomForestAlg seems to post plausible results (score of &gt;0.5), but I still am not able to beat Clustifiers cool probability sorting algorithm  (score of &gt;0.63). I have included meta data from some other files (e.g. promoted_content.csv) to potentially improve the fit, but it seems to make little difference . As far as munging goes, I have only done basic .fillna operations. </p>\n\n<p>I am curious if any other groups have used RFs to better results than I have. Any tips for a new DS on implementing this type of classifier?</p>\n\n<p>Here is a Kaggle friendly version of my script, although the real version is not limited to a 20 minute run time: <a href=\"https://www.kaggle.com/franckjay/outbrain-click-prediction/easy-random-forests/run/407620\">Kaggle Script</a></p>",
  "messages": [
    {
      "id": 141600,
      "postDate": "2016-10-27T14:34:11.007Z",
      "content": "<p>I have been running various ML classifiers on the training/test set, and found that Pythons RandomForestAlg seems to post plausible results (score of &gt;0.5), but I still am not able to beat Clustifiers cool probability sorting algorithm  (score of &gt;0.63). I have included meta data from some other files (e.g. promoted_content.csv) to potentially improve the fit, but it seems to make little difference . As far as munging goes, I have only done basic .fillna operations. </p>\n\n<p>I am curious if any other groups have used RFs to better results than I have. Any tips for a new DS on implementing this type of classifier?</p>\n\n<p>Here is a Kaggle friendly version of my script, although the real version is not limited to a 20 minute run time: <a href=\"https://www.kaggle.com/franckjay/outbrain-click-prediction/easy-random-forests/run/407620\">Kaggle Script</a></p>",
      "rawMarkdown": "I have been running various ML classifiers on the training/test set, and found that Pythons RandomForestAlg seems to post plausible results (score of >0.5), but I still am not able to beat Clustifiers cool probability sorting algorithm  (score of >0.63). I have included meta data from some other files (e.g. promoted_content.csv) to potentially improve the fit, but it seems to make little difference . As far as munging goes, I have only done basic .fillna operations. \r\n\r\nI am curious if any other groups have used RFs to better results than I have. Any tips for a new DS on implementing this type of classifier?\r\n\r\nHere is a Kaggle friendly version of my script, although the real version is not limited to a 20 minute run time: [Kaggle Script][1]\r\n\r\n\r\n  [1]: https://www.kaggle.com/franckjay/outbrain-click-prediction/easy-random-forests/run/407620",
      "votes": 2
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "141600": "I have been running various ML classifiers on the training/test set, and found that Pythons RandomForestAlg seems to post plausible results (score of >0.5), but I still am not able to beat Clustifiers cool probability sorting algorithm  (score of >0.63). I have included meta data from some other files (e.g. promoted_content.csv) to potentially improve the fit, but it seems to make little difference . As far as munging goes, I have only done basic .fillna operations. \r\n\r\nI am curious if any other groups have used RFs to better results than I have. Any tips for a new DS on implementing this type of classifier?\r\n\r\nHere is a Kaggle friendly version of my script, although the real version is not limited to a 20 minute run time: [Kaggle Script][1]\r\n\r\n\r\n  [1]: https://www.kaggle.com/franckjay/outbrain-click-prediction/easy-random-forests/run/407620"
  }
}