{
  "id": 10884,
  "title": "What type of models are people using?",
  "url": "/competitions/seizure-prediction/discussion/10884",
  "author_name": "",
  "post_date": "2014-11-10T12:09:34.273Z",
  "votes": 2,
  "comment_count": 16,
  "views": 3663,
  "content": "<p>Without giving anything away about HOW you are using them, anyone feel like sharing what's working for them? What isn't?</p>\n<p>From my end, I'm finding that SVMs with a radial kernel seem to be providing the most consistent results. Random forests are not working so well, nor are GBMs, but oblique random forests with an SVM decision boundary do somewhat better. </p>",
  "messages": [
    {
      "id": "57593",
      "postDate": "11/10/2014 12:09:34",
      "content": "<p>Without giving anything away about HOW you are using them, anyone feel like sharing what's working for them? What isn't?</p>\n<p>From my end, I'm finding that SVMs with a radial kernel seem to be providing the most consistent results. Random forests are not working so well, nor are GBMs, but oblique random forests with an SVM decision boundary do somewhat better. </p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "57672",
      "postDate": "11/10/2014 20:40:30",
      "content": "<p>fft (see attached file) + logistic regression</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "57676",
      "postDate": "11/10/2014 21:03:36",
      "content": "<p>Many people have reported random forests not working for them, but they (seem to be) working OK for me! &nbsp;;)</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "58144",
      "postDate": "11/16/2014 17:09:57",
      "content": "<p>Hi Ruai,</p>\n<p>I'm new to FFT. I looked at your code but couldn't understand why you created 24 lists for each channel for each file. </p>\n<p>Why 24?</p>\n\n<p>Thanks<br>Chad</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "58145",
      "postDate": "11/16/2014 19:06:04",
      "content": "<p>I am using RandomForest plus FFT features with RFE selection. RF works for me but the result is not good enough though. ROC = 0.73 on public leaderboard.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "58147",
      "postDate": "11/16/2014 19:24:43",
      "content": "<p>[quote=Chuzzelwit;58144]</p>\n<p>Hi Ruai,</p>\n<p>I'm new to FFT. I looked at your code but couldn't understand why you created 24 lists for each channel for each file.</p>\n<p>Why 24?</p>\n<p>Thanks<br>Chad</p>\n<p>[/quote]</p>\n<p>I had tried different pairs of (num, part). <br>And had found that pair (num=24, part=0.03) is not too bad with logistics regression.<br>After that I had used the &#8220;calibration&#8221;. <br>Here calibration is shifting/scaling of dogs/patients predictions to maximize public score. Yes, it is cheating :-). But it seems the rules are not violated ...</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "58148",
      "postDate": "11/16/2014 19:32:59",
      "content": "<p>[quote=ruai;58147]</p>\n<p>[quote=Chuzzelwit;58144]</p>\n<p>Hi Ruai,</p>\n<p>I'm new to FFT. I looked at your code but couldn't understand why you created 24 lists for each channel for each file.</p>\n<p>Why 24?</p>\n<p>Thanks<br>Chad</p>\n<p>[/quote]</p>\n<p>I had tried different pairs of (num, part). <br>And had found that pair (num=24, part=0.03) is not too bad with logistics regression.<br>After that I had used the &#8220;calibration&#8221;. <br>Here calibration is shifting/scaling of dogs/patients predictions to maximize public score. Yes, it is cheating :-). But it seems the rules are not violated ...</p>\n<p>[/quote]</p>\n\n<p>Which calibration method did you use? I tried isotonic probabiliy calibration but it didn't help me much.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "58149",
      "postDate": "11/16/2014 20:03:09",
      "content": "<p>It seems imposable to calibrate with only train data. So, calibration can use test data:<br>http://www.kaggle.com/c/seizure-prediction/forums/t/10790/use-of-test-data</p>\n<p>Example of &#8220;calibration&#8221; Patient_1:</p>\n<p>library(boot)<br>...<br> if (k==6)<br> data[,out] =&nbsp;inv.logit(-1.0+0.95*logit(data[,out]))</p>\n<p>Here -1.0 and 0.95 had be approximately found by test data. If data division 40/60 is based on full time series than private score will be very low.</p>\n<p>Probably prediction posts statistics may be used &#8230;</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "58150",
      "postDate": "11/16/2014 20:12:07",
      "content": "<p>[quote=ruai;58149]</p>\n<p>It seems imposable to calibrate with only train data. So, calibration can use test data:<br>http://www.kaggle.com/c/seizure-prediction/forums/t/10790/use-of-test-data</p>\n<p>Example of &#8220;calibration&#8221; Patient_1:</p>\n<p>library(boot)<br>...<br> if (k==6)<br> data[,out] =&nbsp;inv.logit(-1.0+0.95*logit(data[,out]))</p>\n<p>Here -1.0 and 0.95 had be approximately found by test data. If data division 40/60 is based on full time series than private score will be very low.</p>\n<p>Probably prediction posts statistics may be used &#8230;</p>\n<p>[/quote]</p>\n<p>If I am not understanding it wrong, the point lies in finding the parameters like -1.0 and 0.95 for all subjects. But finding those parameters need you to submit a lot of entries and see the change of leaderboard score? Or just look at statistics of the output of the original classifier and somehow &quot;normalise&quot; it by using those parameters?</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "58151",
      "postDate": "11/16/2014 20:17:45",
      "content": "<p>...post entries and see leaderboard score.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "58152",
      "postDate": "11/16/2014 20:21:31",
      "content": "<p>[quote=ruai;58151]</p>\n<p>...post entries and see leaderboard score.</p>\n<p>[/quote]</p>\n\n<p>If this is the case, because the leaderboard score is calculated based on 40% of the data, there is a potential risk of overfitting. By the way, can you share how much you benefit from calibration? I have tried isotonic regression and platt scaling and got roughly 0.03-0.05 improvement of ROC.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "58153",
      "postDate": "11/16/2014 20:25:35",
      "content": "<p>Yes, there is overfitting risk. But by rules you can select up to 2 submissions.</p>\n<p>Public score was&nbsp;up on&nbsp;~0.06</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "58154",
      "postDate": "11/16/2014 20:32:39",
      "content": "<p>[quote=Eureka222;58152]</p>\n<p>[quote=ruai;58151]</p>\n<p>...post entries and see leaderboard score.</p>\n<p>[/quote]</p>\n<p>If this is the case, because the leaderboard score is calculated based on 40% of the data, there is a potential risk of overfitting. By the way, can you share how much you benefit from calibration? I have tried isotonic regression and platt scaling and got roughly 0.03-0.05 improvement of ROC.</p>\n<p>[/quote]</p>\n<p>For the isotonic regression, did you do cross validation? Or did you just apply the fully trained model on the train set to get the probabilities? I tried some time ago with cross validation but my score actually even dropped.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "58155",
      "postDate": "11/16/2014 20:36:33",
      "content": "",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "58156",
      "postDate": "11/16/2014 20:36:40",
      "content": "<p>[quote=nils;58154]</p>\n<p>[quote=Eureka222;58152]</p>\n<p>[quote=ruai;58151]</p>\n<p>...post entries and see leaderboard score.</p>\n<p>[/quote]</p>\n<p>If this is the case, because the leaderboard score is calculated based on 40% of the data, there is a potential risk of overfitting. By the way, can you share how much you benefit from calibration? I have tried isotonic regression and platt scaling and got roughly 0.03-0.05 improvement of ROC.</p>\n<p>[/quote]</p>\n<p>For the isotonic regression, did you do cross validation? Or did you just apply the fully trained model on the train set to get the probabilities? I tried some time ago with cross validation but my score actually even dropped.</p>\n<p>[/quote]</p>\n\n\n<p>I have tried two ways, one is to hold out a part of unseen data out of the whole training and cv data and do the calibration. The second way is to do a &quot;nested&quot; cross validation. I am now using the first one.&nbsp;</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "58157",
      "postDate": "11/16/2014 20:58:20",
      "content": "<p>Thanks Ruai, I am able to get ~0.74 leaderboard using your features and simple models, without doing any &quot;peeking&quot; adjustments base on the public score.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "58307",
      "postDate": "11/18/2014 12:05:45",
      "content": "<p>[quote=ruai;58149]</p>\n<p>It seems imposable to calibrate with only train data. So, calibration can use test data:<br>http://www.kaggle.com/c/seizure-prediction/forums/t/10790/use-of-test-data</p>\n<p>Example of &#8220;calibration&#8221; Patient_1:</p>\n<p>library(boot)<br>...<br> if (k==6)<br> data[,out] =&nbsp;inv.logit(-1.0+0.95*logit(data[,out]))</p>\n<p>Here -1.0 and 0.95 had be approximately found by test data. If data division 40/60 is based on full time series than private score will be very low.</p>\n<p>Probably prediction posts statistics may be used &#8230;</p>\n<p>[/quote]</p>\n<p>As the competition is finished, can you explain how to get -1.0 and 0.95 from the test set for Patient_1? It is really interesting for me.</p>\n<p>Thanks in advance.</p>\n<p>Alex</p>",
      "rawMarkdown": "",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 57672,
      "author_name": "ruai00",
      "author_url": "",
      "post_date": "11/10/2014 20:40:30",
      "content": "<p>fft (see attached file) + logistic regression</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 57676,
      "author_name": "drewabbot",
      "author_url": "",
      "post_date": "11/10/2014 21:03:36",
      "content": "<p>Many people have reported random forests not working for them, but they (seem to be) working OK for me! &nbsp;;)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 58144,
      "author_name": "chuzzelwit",
      "author_url": "",
      "post_date": "11/16/2014 17:09:57",
      "content": "<p>Hi Ruai,</p>\n<p>I'm new to FFT. I looked at your code but couldn't understand why you created 24 lists for each channel for each file. </p>\n<p>Why 24?</p>\n\n<p>Thanks<br>Chad</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 58145,
      "author_name": "feixia",
      "author_url": "",
      "post_date": "11/16/2014 19:06:04",
      "content": "<p>I am using RandomForest plus FFT features with RFE selection. RF works for me but the result is not good enough though. ROC = 0.73 on public leaderboard.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 58147,
      "author_name": "ruai00",
      "author_url": "",
      "post_date": "11/16/2014 19:24:43",
      "content": "<p>[quote=Chuzzelwit;58144]</p>\n<p>Hi Ruai,</p>\n<p>I'm new to FFT. I looked at your code but couldn't understand why you created 24 lists for each channel for each file.</p>\n<p>Why 24?</p>\n<p>Thanks<br>Chad</p>\n<p>[/quote]</p>\n<p>I had tried different pairs of (num, part). <br>And had found that pair (num=24, part=0.03) is not too bad with logistics regression.<br>After that I had used the &#8220;calibration&#8221;. <br>Here calibration is shifting/scaling of dogs/patients predictions to maximize public score. Yes, it is cheating :-). But it seems the rules are not violated ...</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 58148,
      "author_name": "scatterbrain333",
      "author_url": "",
      "post_date": "11/16/2014 19:32:59",
      "content": "<p>[quote=ruai;58147]</p>\n<p>[quote=Chuzzelwit;58144]</p>\n<p>Hi Ruai,</p>\n<p>I'm new to FFT. I looked at your code but couldn't understand why you created 24 lists for each channel for each file.</p>\n<p>Why 24?</p>\n<p>Thanks<br>Chad</p>\n<p>[/quote]</p>\n<p>I had tried different pairs of (num, part). <br>And had found that pair (num=24, part=0.03) is not too bad with logistics regression.<br>After that I had used the &#8220;calibration&#8221;. <br>Here calibration is shifting/scaling of dogs/patients predictions to maximize public score. Yes, it is cheating :-). But it seems the rules are not violated ...</p>\n<p>[/quote]</p>\n\n<p>Which calibration method did you use? I tried isotonic probabiliy calibration but it didn't help me much.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 58149,
      "author_name": "ruai00",
      "author_url": "",
      "post_date": "11/16/2014 20:03:09",
      "content": "<p>It seems imposable to calibrate with only train data. So, calibration can use test data:<br>http://www.kaggle.com/c/seizure-prediction/forums/t/10790/use-of-test-data</p>\n<p>Example of &#8220;calibration&#8221; Patient_1:</p>\n<p>library(boot)<br>...<br> if (k==6)<br> data[,out] =&nbsp;inv.logit(-1.0+0.95*logit(data[,out]))</p>\n<p>Here -1.0 and 0.95 had be approximately found by test data. If data division 40/60 is based on full time series than private score will be very low.</p>\n<p>Probably prediction posts statistics may be used &#8230;</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 58150,
      "author_name": "feixia",
      "author_url": "",
      "post_date": "11/16/2014 20:12:07",
      "content": "<p>[quote=ruai;58149]</p>\n<p>It seems imposable to calibrate with only train data. So, calibration can use test data:<br>http://www.kaggle.com/c/seizure-prediction/forums/t/10790/use-of-test-data</p>\n<p>Example of &#8220;calibration&#8221; Patient_1:</p>\n<p>library(boot)<br>...<br> if (k==6)<br> data[,out] =&nbsp;inv.logit(-1.0+0.95*logit(data[,out]))</p>\n<p>Here -1.0 and 0.95 had be approximately found by test data. If data division 40/60 is based on full time series than private score will be very low.</p>\n<p>Probably prediction posts statistics may be used &#8230;</p>\n<p>[/quote]</p>\n<p>If I am not understanding it wrong, the point lies in finding the parameters like -1.0 and 0.95 for all subjects. But finding those parameters need you to submit a lot of entries and see the change of leaderboard score? Or just look at statistics of the output of the original classifier and somehow &quot;normalise&quot; it by using those parameters?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 58151,
      "author_name": "ruai00",
      "author_url": "",
      "post_date": "11/16/2014 20:17:45",
      "content": "<p>...post entries and see leaderboard score.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 58152,
      "author_name": "feixia",
      "author_url": "",
      "post_date": "11/16/2014 20:21:31",
      "content": "<p>[quote=ruai;58151]</p>\n<p>...post entries and see leaderboard score.</p>\n<p>[/quote]</p>\n\n<p>If this is the case, because the leaderboard score is calculated based on 40% of the data, there is a potential risk of overfitting. By the way, can you share how much you benefit from calibration? I have tried isotonic regression and platt scaling and got roughly 0.03-0.05 improvement of ROC.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 58153,
      "author_name": "ruai00",
      "author_url": "",
      "post_date": "11/16/2014 20:25:35",
      "content": "<p>Yes, there is overfitting risk. But by rules you can select up to 2 submissions.</p>\n<p>Public score was&nbsp;up on&nbsp;~0.06</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 58154,
      "author_name": "hasselmann",
      "author_url": "",
      "post_date": "11/16/2014 20:32:39",
      "content": "<p>[quote=Eureka222;58152]</p>\n<p>[quote=ruai;58151]</p>\n<p>...post entries and see leaderboard score.</p>\n<p>[/quote]</p>\n<p>If this is the case, because the leaderboard score is calculated based on 40% of the data, there is a potential risk of overfitting. By the way, can you share how much you benefit from calibration? I have tried isotonic regression and platt scaling and got roughly 0.03-0.05 improvement of ROC.</p>\n<p>[/quote]</p>\n<p>For the isotonic regression, did you do cross validation? Or did you just apply the fully trained model on the train set to get the probabilities? I tried some time ago with cross validation but my score actually even dropped.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 58155,
      "author_name": "feixia",
      "author_url": "",
      "post_date": "11/16/2014 20:36:33",
      "content": "",
      "votes": null,
      "replies": []
    },
    {
      "id": 58156,
      "author_name": "feixia",
      "author_url": "",
      "post_date": "11/16/2014 20:36:40",
      "content": "<p>[quote=nils;58154]</p>\n<p>[quote=Eureka222;58152]</p>\n<p>[quote=ruai;58151]</p>\n<p>...post entries and see leaderboard score.</p>\n<p>[/quote]</p>\n<p>If this is the case, because the leaderboard score is calculated based on 40% of the data, there is a potential risk of overfitting. By the way, can you share how much you benefit from calibration? I have tried isotonic regression and platt scaling and got roughly 0.03-0.05 improvement of ROC.</p>\n<p>[/quote]</p>\n<p>For the isotonic regression, did you do cross validation? Or did you just apply the fully trained model on the train set to get the probabilities? I tried some time ago with cross validation but my score actually even dropped.</p>\n<p>[/quote]</p>\n\n\n<p>I have tried two ways, one is to hold out a part of unseen data out of the whole training and cv data and do the calibration. The second way is to do a &quot;nested&quot; cross validation. I am now using the first one.&nbsp;</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 58157,
      "author_name": "jfkingiii",
      "author_url": "",
      "post_date": "11/16/2014 20:58:20",
      "content": "<p>Thanks Ruai, I am able to get ~0.74 leaderboard using your features and simple models, without doing any &quot;peeking&quot; adjustments base on the public score.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 58307,
      "author_name": "alexryzhkov",
      "author_url": "",
      "post_date": "11/18/2014 12:05:45",
      "content": "<p>[quote=ruai;58149]</p>\n<p>It seems imposable to calibrate with only train data. So, calibration can use test data:<br>http://www.kaggle.com/c/seizure-prediction/forums/t/10790/use-of-test-data</p>\n<p>Example of &#8220;calibration&#8221; Patient_1:</p>\n<p>library(boot)<br>...<br> if (k==6)<br> data[,out] =&nbsp;inv.logit(-1.0+0.95*logit(data[,out]))</p>\n<p>Here -1.0 and 0.95 had be approximately found by test data. If data division 40/60 is based on full time series than private score will be very low.</p>\n<p>Probably prediction posts statistics may be used &#8230;</p>\n<p>[/quote]</p>\n<p>As the competition is finished, can you explain how to get -1.0 and 0.95 from the test set for Patient_1? It is really interesting for me.</p>\n<p>Thanks in advance.</p>\n<p>Alex</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "57593": "",
    "57672": "",
    "57676": "",
    "58144": "",
    "58145": "",
    "58147": "",
    "58148": "",
    "58149": "",
    "58150": "",
    "58151": "",
    "58152": "",
    "58153": "",
    "58154": "",
    "58155": "",
    "58156": "",
    "58157": "",
    "58307": ""
  },
  "source": "meta"
}