{
  "id": 10838,
  "title": "Why are random forests performing so poorly?",
  "url": "/competitions/seizure-prediction/discussion/10838",
  "author_name": "",
  "post_date": "2014-11-05T02:44:34.167Z",
  "votes": 1,
  "comment_count": 6,
  "views": 6570,
  "content": "<p>Hey all,</p>\n\n<p>I've managed to extract a bunch of features (between 500-1000 per patient) which I think intuitively ought to make for good classifiers. I saw that the winners of the previous challenge used random forests (or some variation thereof) and decided to try it myself. Much to my dismay, using both the random forests and extra trees packages in R, most or even all of my predictions come out to be 0s. I'm aware that using advanced machine learning algorithms like random forests may overfit the data, but I didn't think it would be this bad, especially since the previous competitors had similar numbers of features.</p>\n\n<p>Are other people having similar experiences? Is there something obvious I might be doing wrong?</p>\n\n<p>Thanks,</p>\n<p>Mike</p>",
  "messages": [
    {
      "id": "57328",
      "postDate": "11/05/2014 02:44:34",
      "content": "<p>Hey all,</p>\n\n<p>I've managed to extract a bunch of features (between 500-1000 per patient) which I think intuitively ought to make for good classifiers. I saw that the winners of the previous challenge used random forests (or some variation thereof) and decided to try it myself. Much to my dismay, using both the random forests and extra trees packages in R, most or even all of my predictions come out to be 0s. I'm aware that using advanced machine learning algorithms like random forests may overfit the data, but I didn't think it would be this bad, especially since the previous competitors had similar numbers of features.</p>\n\n<p>Are other people having similar experiences? Is there something obvious I might be doing wrong?</p>\n\n<p>Thanks,</p>\n<p>Mike</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "57370",
      "postDate": "11/06/2014 02:01:11",
      "content": "<p>I also noticed that Random Forest didn't perform as well as I was expecting it to. Instead I threw all the different classifiers scikit-learn has to offer and picked one that offered better performance in cross-validation.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "57480",
      "postDate": "11/07/2014 19:06:22",
      "content": "<p>Don't know if it is possible in R since I am using Matlab, but you might want to configure your RandomForest to output probabilities instead of binary results, maybe by using the regression version instead of classification version of the algorithm. This could help you learn more about the predictions made by your model, and get a more fine-grained ROC curve.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "57495",
      "postDate": "11/08/2014 00:11:42",
      "content": "<p>I've also noticed that random forest doesn't do a great job. I got around LB 0.7-.72 with a random forest. Hastie et al point out in 'elements of statistical learning' that random forests suffer if there are too few good variables relative to noisy variables; not sure what's going on here, though. TreeBagger in matlab is a bit misleading when using the oobpred option. I found that my ooberror was very low at first but was more realistic when setting prior to uniform</p>\n<p>I haven't had any luck with fitensemble.m either. I have the most luck with an open source fortran package ;)</p>\n<p>TreeBagger will output both class labels and probabilities (average output across trees).&nbsp;</p>\n<p>[labels, posterior] = predict(b); % where b is returned by TreeBagger.m. Second column in posterior corresponds to p(class=1)</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "57894",
      "postDate": "11/12/2014 10:45:51",
      "content": "<p>Random forests seem to do best with a small number of very informative variables. &nbsp;If I had to fit a random forest to hundreds of features, I would be tempted to do dimension reduction first (SVD/PCA) on the training data, and just take the first 50 or so dimensions. &nbsp;Then I would try looking at the Importance of each of these new variables in the forest fit to them (in R, importance of each variable is available as an array in a fitted random forest object). &nbsp;Then I would probably remove all but the 10 most important variables. &nbsp;I think the RF method can soon be degraded by noisy (non-predictive) variables.</p>\n<p>This is all quite heuristic, and hasn't won me any honours yet!</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "58026",
      "postDate": "11/13/2014 20:26:28",
      "content": "<p>I had a similar experiance with Random Forests.&nbsp; However, I belive I had way to many features and by reducing the feature count I was able to avoid over-fitting and improve my leaderboard score. Also, using too many trees in the RF can cause over-fitting.&nbsp;</p>\n<p>Some channels are hightly correlated so I throw out some to keep my dataset size in check.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "96253",
      "postDate": "10/15/2015 09:58:09",
      "content": "<p>The R issue you are facing (all 0's or all 1's)  may be likely because you haven't chosen type='prob' in your <code>randomforest()</code> function call. This function by default returns the majority vote result rather than the fraction of votes.</p>\n\n<p><a href=\"https://cran.r-project.org/web/packages/randomForest/randomForest.pdf\">https://cran.r-project.org/web/packages/randomForest/randomForest.pdf</a></p>",
      "rawMarkdown": "The R issue you are facing (all 0's or all 1's)  may be likely because you haven't chosen type='prob' in your `randomforest()` function call. This function by default returns the majority vote result rather than the fraction of votes.\r\n\r\nhttps://cran.r-project.org/web/packages/randomForest/randomForest.pdf",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 57370,
      "author_name": "michaelhills",
      "author_url": "",
      "post_date": "11/06/2014 02:01:11",
      "content": "<p>I also noticed that Random Forest didn't perform as well as I was expecting it to. Instead I threw all the different classifiers scikit-learn has to offer and picked one that offered better performance in cross-validation.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 57480,
      "author_name": "meditativeape",
      "author_url": "",
      "post_date": "11/07/2014 19:06:22",
      "content": "<p>Don't know if it is possible in R since I am using Matlab, but you might want to configure your RandomForest to output probabilities instead of binary results, maybe by using the regression version instead of classification version of the algorithm. This could help you learn more about the predictions made by your model, and get a more fine-grained ROC curve.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 57495,
      "author_name": "bgeier",
      "author_url": "",
      "post_date": "11/08/2014 00:11:42",
      "content": "<p>I've also noticed that random forest doesn't do a great job. I got around LB 0.7-.72 with a random forest. Hastie et al point out in 'elements of statistical learning' that random forests suffer if there are too few good variables relative to noisy variables; not sure what's going on here, though. TreeBagger in matlab is a bit misleading when using the oobpred option. I found that my ooberror was very low at first but was more realistic when setting prior to uniform</p>\n<p>I haven't had any luck with fitensemble.m either. I have the most luck with an open source fortran package ;)</p>\n<p>TreeBagger will output both class labels and probabilities (average output across trees).&nbsp;</p>\n<p>[labels, posterior] = predict(b); % where b is returned by TreeBagger.m. Second column in posterior corresponds to p(class=1)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 57894,
      "author_name": "jaymoore",
      "author_url": "",
      "post_date": "11/12/2014 10:45:51",
      "content": "<p>Random forests seem to do best with a small number of very informative variables. &nbsp;If I had to fit a random forest to hundreds of features, I would be tempted to do dimension reduction first (SVD/PCA) on the training data, and just take the first 50 or so dimensions. &nbsp;Then I would try looking at the Importance of each of these new variables in the forest fit to them (in R, importance of each variable is available as an array in a fitted random forest object). &nbsp;Then I would probably remove all but the 10 most important variables. &nbsp;I think the RF method can soon be degraded by noisy (non-predictive) variables.</p>\n<p>This is all quite heuristic, and hasn't won me any honours yet!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 58026,
      "author_name": "dantybor",
      "author_url": "",
      "post_date": "11/13/2014 20:26:28",
      "content": "<p>I had a similar experiance with Random Forests.&nbsp; However, I belive I had way to many features and by reducing the feature count I was able to avoid over-fitting and improve my leaderboard score. Also, using too many trees in the RF can cause over-fitting.&nbsp;</p>\n<p>Some channels are hightly correlated so I throw out some to keep my dataset size in check.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 96253,
      "author_name": "zhubarb",
      "author_url": "",
      "post_date": "10/15/2015 09:58:09",
      "content": "<p>The R issue you are facing (all 0's or all 1's)  may be likely because you haven't chosen type='prob' in your <code>randomforest()</code> function call. This function by default returns the majority vote result rather than the fraction of votes.</p>\n\n<p><a href=\"https://cran.r-project.org/web/packages/randomForest/randomForest.pdf\">https://cran.r-project.org/web/packages/randomForest/randomForest.pdf</a></p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "57328": "",
    "57370": "",
    "57480": "",
    "57495": "",
    "57894": "",
    "58026": "",
    "96253": "The R issue you are facing (all 0's or all 1's)  may be likely because you haven't chosen type='prob' in your `randomforest()` function call. This function by default returns the majority vote result rather than the fraction of votes.\r\n\r\nhttps://cran.r-project.org/web/packages/randomForest/randomForest.pdf"
  },
  "source": "meta"
}