{
  "id": 11017,
  "title": "Xgboost: boost from existing predictions",
  "url": "/competitions/inria-bci-challenge/discussion/11017",
  "author_name": "",
  "post_date": "2014-11-22T16:20:25.813Z",
  "votes": 7,
  "comment_count": 12,
  "views": 4957,
  "content": "<p>Hi everyone, I'm toying this new feature of the latest version xgboost: https://github.com/tqchen/xgboost</p>\n<p>&quot;boost from existing predictions&quot;</p>\n<p>https://github.com/tqchen/xgboost/blob/master/demo/guide-python/boost_from_prediction.py</p>\n<p>so I modify Abhishek's great benchmark, http://www.kaggle.com/c/inria-bci-challenge/forums/t/11009/beating-the-benchmark/58669#post58669</p>\n<p>to let xgb boost from random forest predictions. The code is attached.</p>\n<p>To generate a submission, install xgboost first, change the path of xgb wrapper in xgb_classifier.py, then run &quot;python xgb2.py&quot;</p>\n<p>I am curious about how well it can help other people's existing model and please share your observations. Thank you.</p>\n<p>Just one thing, please don't downvote me, I have a fragile heart. o_o</p>\n<p>edit: LB&nbsp;0.56195</p>",
  "messages": [
    {
      "id": "58688",
      "postDate": "11/22/2014 16:20:25",
      "content": "<p>Hi everyone, I'm toying this new feature of the latest version xgboost: https://github.com/tqchen/xgboost</p>\n<p>&quot;boost from existing predictions&quot;</p>\n<p>https://github.com/tqchen/xgboost/blob/master/demo/guide-python/boost_from_prediction.py</p>\n<p>so I modify Abhishek's great benchmark, http://www.kaggle.com/c/inria-bci-challenge/forums/t/11009/beating-the-benchmark/58669#post58669</p>\n<p>to let xgb boost from random forest predictions. The code is attached.</p>\n<p>To generate a submission, install xgboost first, change the path of xgb wrapper in xgb_classifier.py, then run &quot;python xgb2.py&quot;</p>\n<p>I am curious about how well it can help other people's existing model and please share your observations. Thank you.</p>\n<p>Just one thing, please don't downvote me, I have a fragile heart. o_o</p>\n<p>edit: LB&nbsp;0.56195</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "58690",
      "postDate": "11/22/2014 16:43:56",
      "content": "<p>+1, but internet fora are a dangerous place for fragile hearts. fyi Abhishek's code as I just ran it gives&nbsp;0.57488. I'm also interested in understanding better when metafeatures as used on tradeshift are useful.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "58692",
      "postDate": "11/22/2014 17:21:39",
      "content": "<p>[quote=rcarson;58688]</p>\n\n<p>Just one thing, please don't downvote me, I have a fragile heart. o_o</p>\n\n<p>[/quote]</p>\n<p>Ha! I get downvoted a lot because of my benchmark dislike - it makes me laugh though.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "58693",
      "postDate": "11/22/2014 17:34:12",
      "content": "<p>[quote=James King;58690]</p>\n<p>fyi Abhishek's code as I just ran it gives&nbsp;0.57488. I'm also interested in understanding better when metafeatures as used on tradeshift are useful.</p>\n<p>[/quote]</p>\n<p>Thank you. I use very weak rf classifier for base (no particular reason, just playing), not exactly the same as the original benchmark. And I don't think here it is for meta features. Xgb uses existing predictions as &quot;margin&quot; for further training (I don't know exactly what that means). Extracting meta features for this one will be much less straightforward than Tradeshift. But I think this one might help:&nbsp;https://dl.dropboxusercontent.com/u/5755730/ModelDescription.pdf</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "58694",
      "postDate": "11/22/2014 17:35:57",
      "content": "<p>[quote=James King;58690]</p>\n<p>+1, but internet fora are a dangerous place for fragile hearts.&nbsp;</p>\n<p>[/quote]</p>\n<p>[quote=M;58692]</p>\n<p>Ha! I get downvoted a lot because of my benchmark dislike - it makes me laugh though.</p>\n<p>[/quote]</p>\n<p>Thank you! I'm just kidding :D My wife will upvote whatever nonsense I'm posting anyway.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "58756",
      "postDate": "11/24/2014 00:44:24",
      "content": "<p>HI,</p>\n<p>The public score is based on 20% of test data , it may not be a good indicator. Better trust your own AUC in CV.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "58759",
      "postDate": "11/24/2014 02:44:09",
      "content": "<p>On eight fold by subject cv I have AUC ranging between .48 and .75 on individual folds. Leader board is also two subjects..</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "58768",
      "postDate": "11/24/2014 09:40:01",
      "content": "<p>[quote=phalaris;58759]</p>\n<p>On eight fold by subject cv I have AUC ranging between .48 and .75 on individual folds. Leader board is also two subjects..</p>\n<p>[/quote]</p>\n\n<p>Do you mean that the public score is based on the performance on 2 of the 10 test subjects, rather than a mixture of samples from the 10 test subjects?</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "58775",
      "postDate": "11/24/2014 14:37:40",
      "content": "<p>[quote=Jose M.;58768]</p>\n<p>Do you mean that the public score is based on the performance on 2 of the 10 test subjects, rather than a mixture of samples from the 10 test subjects?</p>\n<p>[/quote]</p>\n<p>I'm almost sure the public LB&nbsp;is based on two subjects.</p>\n<p>Don't trust the LB !</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "58776",
      "postDate": "11/24/2014 14:46:30",
      "content": "<p>@Jose, I am not positive that it is by subject, though it is very likely. There are ten test subjects so 20% would be two.&nbsp; AUC is much higher for some subjects on local CV. The paper for this competition mentions that the performance of their classifiers varied by subject.</p>\n<p>An AUC of .75 is at the top of those I observed on local CV and my LB score was about .75. From this I assumed that the leaderboard test set consisted of two test subjects, because a random sample should have a lower AUC.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "58793",
      "postDate": "11/24/2014 17:37:43",
      "content": "<p>I decided to burn a submission to check&nbsp;this - I submitted Prediction=0.1 for S15 and Prediction=0.5 for all other subjects, and got back a score of 0.50000.</p>\n<p>That seems to confirm that only 2 subjects are used for the public leaderboard (and S15 is not one of them).&nbsp;</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "58826",
      "postDate": "11/24/2014 23:00:05",
      "content": "<p>Thanks for sharing, @emolson, I'll let you (and the rest) know If I discover the two subjects (I haven't read anything against this kind of reverse engineering in the rules...),</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "58827",
      "postDate": "11/24/2014 23:00:38",
      "content": "",
      "rawMarkdown": "",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 58690,
      "author_name": "jfkingiii",
      "author_url": "",
      "post_date": "11/22/2014 16:43:56",
      "content": "<p>+1, but internet fora are a dangerous place for fragile hearts. fyi Abhishek's code as I just ran it gives&nbsp;0.57488. I'm also interested in understanding better when metafeatures as used on tradeshift are useful.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 58692,
      "author_name": "domcastro",
      "author_url": "",
      "post_date": "11/22/2014 17:21:39",
      "content": "<p>[quote=rcarson;58688]</p>\n\n<p>Just one thing, please don't downvote me, I have a fragile heart. o_o</p>\n\n<p>[/quote]</p>\n<p>Ha! I get downvoted a lot because of my benchmark dislike - it makes me laugh though.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 58693,
      "author_name": "jiweiliu",
      "author_url": "",
      "post_date": "11/22/2014 17:34:12",
      "content": "<p>[quote=James King;58690]</p>\n<p>fyi Abhishek's code as I just ran it gives&nbsp;0.57488. I'm also interested in understanding better when metafeatures as used on tradeshift are useful.</p>\n<p>[/quote]</p>\n<p>Thank you. I use very weak rf classifier for base (no particular reason, just playing), not exactly the same as the original benchmark. And I don't think here it is for meta features. Xgb uses existing predictions as &quot;margin&quot; for further training (I don't know exactly what that means). Extracting meta features for this one will be much less straightforward than Tradeshift. But I think this one might help:&nbsp;https://dl.dropboxusercontent.com/u/5755730/ModelDescription.pdf</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 58694,
      "author_name": "jiweiliu",
      "author_url": "",
      "post_date": "11/22/2014 17:35:57",
      "content": "<p>[quote=James King;58690]</p>\n<p>+1, but internet fora are a dangerous place for fragile hearts.&nbsp;</p>\n<p>[/quote]</p>\n<p>[quote=M;58692]</p>\n<p>Ha! I get downvoted a lot because of my benchmark dislike - it makes me laugh though.</p>\n<p>[/quote]</p>\n<p>Thank you! I'm just kidding :D My wife will upvote whatever nonsense I'm posting anyway.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 58756,
      "author_name": "stevendu",
      "author_url": "",
      "post_date": "11/24/2014 00:44:24",
      "content": "<p>HI,</p>\n<p>The public score is based on 20% of test data , it may not be a good indicator. Better trust your own AUC in CV.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 58759,
      "author_name": "devinanzelmo",
      "author_url": "",
      "post_date": "11/24/2014 02:44:09",
      "content": "<p>On eight fold by subject cv I have AUC ranging between .48 and .75 on individual folds. Leader board is also two subjects..</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 58768,
      "author_name": "joseleiva",
      "author_url": "",
      "post_date": "11/24/2014 09:40:01",
      "content": "<p>[quote=phalaris;58759]</p>\n<p>On eight fold by subject cv I have AUC ranging between .48 and .75 on individual folds. Leader board is also two subjects..</p>\n<p>[/quote]</p>\n\n<p>Do you mean that the public score is based on the performance on 2 of the 10 test subjects, rather than a mixture of samples from the 10 test subjects?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 58775,
      "author_name": "alexandrebarachant",
      "author_url": "",
      "post_date": "11/24/2014 14:37:40",
      "content": "<p>[quote=Jose M.;58768]</p>\n<p>Do you mean that the public score is based on the performance on 2 of the 10 test subjects, rather than a mixture of samples from the 10 test subjects?</p>\n<p>[/quote]</p>\n<p>I'm almost sure the public LB&nbsp;is based on two subjects.</p>\n<p>Don't trust the LB !</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 58776,
      "author_name": "devinanzelmo",
      "author_url": "",
      "post_date": "11/24/2014 14:46:30",
      "content": "<p>@Jose, I am not positive that it is by subject, though it is very likely. There are ten test subjects so 20% would be two.&nbsp; AUC is much higher for some subjects on local CV. The paper for this competition mentions that the performance of their classifiers varied by subject.</p>\n<p>An AUC of .75 is at the top of those I observed on local CV and my LB score was about .75. From this I assumed that the leaderboard test set consisted of two test subjects, because a random sample should have a lower AUC.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 58793,
      "author_name": "emolson",
      "author_url": "",
      "post_date": "11/24/2014 17:37:43",
      "content": "<p>I decided to burn a submission to check&nbsp;this - I submitted Prediction=0.1 for S15 and Prediction=0.5 for all other subjects, and got back a score of 0.50000.</p>\n<p>That seems to confirm that only 2 subjects are used for the public leaderboard (and S15 is not one of them).&nbsp;</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 58826,
      "author_name": "joseleiva",
      "author_url": "",
      "post_date": "11/24/2014 23:00:05",
      "content": "<p>Thanks for sharing, @emolson, I'll let you (and the rest) know If I discover the two subjects (I haven't read anything against this kind of reverse engineering in the rules...),</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 58827,
      "author_name": "joseleiva",
      "author_url": "",
      "post_date": "11/24/2014 23:00:38",
      "content": "",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "58688": "",
    "58690": "",
    "58692": "",
    "58693": "",
    "58694": "",
    "58756": "",
    "58759": "",
    "58768": "",
    "58775": "",
    "58776": "",
    "58793": "",
    "58826": "",
    "58827": ""
  },
  "source": "meta"
}