{
  "id": 268574,
  "title": "Shot Self in Foot(?)",
  "url": "/competitions/g2net-gravitational-wave-detection/discussion/268574",
  "author_name": "",
  "post_date": "2021-08-27T22:27:03.734182700Z",
  "votes": 6,
  "comment_count": 9,
  "views": 0,
  "content": "<p>When I first started this competition, the first thing I did was look at how many samples exist, what their distribution was, and then setup my CV. I split the data into 6 buckets, and have been using four of those to build 4-fold CV models and kept the remaining two buckets for stacking. I had just been doing OOF blending for ensembling. But since I have ~48 trained models now, I tried to give the holdout stacker a go. </p>\n<p>The results were…… not what I was expecting.</p>\n<p>I've tried a variety of model types for the stacker: GPR, LogisticRegression, Linear, Ridge, Lasso, Elastic, SVR (RBF+Linear), MultinomialNB, RF, SGDR. Basically everything under the sun except a GBDT and a NNet. No matter what ML model I use to stack the holdout predictions, the model's validation performance has never beaten a linear combination of weights derived from a straight scipy.optimize.minimize. How is that possible??</p>\n<p>I base my understanding of stacking on <a href=\"https://www.kaggle.com/kazanova\" target=\"_blank\">@kazanova</a>'s great Kaggle coursera series. I believe stacking with a ML model (particularly, one that has more than one parameter per base-model's prediction) is that it should be superior because it can detect \"zones\" where one model preforms good and prefer it those regions; and then prefer other models in regions where their results are better.</p>\n<p>Best Stack:</p>\n<ul>\n<li>Val: 0.874972 (again, this is trained+validated on 2/6th of the data, split in two folds)</li>\n<li>LB: 0.8767</li>\n</ul>\n<p>Random Minimize Model Weights:</p>\n<ul>\n<li>Val 0.875204</li>\n<li>LB: 0.8768</li>\n</ul>\n<p>The best stacker model was cumlSGD. No other models tested came even close (example CVs ~840, ~800, 720, 0.5!, etc).</p>\n<p>Did I shoot myself in the foot? Is better to spend the next full week re-training all of my models on the full data rather than keeping 33% holdout for a stacker that is worthless? Has any one else experimented with stacking their models yet? I'm running experiments now with GBDT and NNet models to finish off the experiment set..</p>",
  "messages": [
    {
      "id": "1493437",
      "postDate": "08/27/2021 22:27:03",
      "content": "<p>When I first started this competition, the first thing I did was look at how many samples exist, what their distribution was, and then setup my CV. I split the data into 6 buckets, and have been using four of those to build 4-fold CV models and kept the remaining two buckets for stacking. I had just been doing OOF blending for ensembling. But since I have ~48 trained models now, I tried to give the holdout stacker a go. </p>\n<p>The results were…… not what I was expecting.</p>\n<p>I've tried a variety of model types for the stacker: GPR, LogisticRegression, Linear, Ridge, Lasso, Elastic, SVR (RBF+Linear), MultinomialNB, RF, SGDR. Basically everything under the sun except a GBDT and a NNet. No matter what ML model I use to stack the holdout predictions, the model's validation performance has never beaten a linear combination of weights derived from a straight scipy.optimize.minimize. How is that possible??</p>\n<p>I base my understanding of stacking on <a href=\"https://www.kaggle.com/kazanova\" target=\"_blank\">@kazanova</a>'s great Kaggle coursera series. I believe stacking with a ML model (particularly, one that has more than one parameter per base-model's prediction) is that it should be superior because it can detect \"zones\" where one model preforms good and prefer it those regions; and then prefer other models in regions where their results are better.</p>\n<p>Best Stack:</p>\n<ul>\n<li>Val: 0.874972 (again, this is trained+validated on 2/6th of the data, split in two folds)</li>\n<li>LB: 0.8767</li>\n</ul>\n<p>Random Minimize Model Weights:</p>\n<ul>\n<li>Val 0.875204</li>\n<li>LB: 0.8768</li>\n</ul>\n<p>The best stacker model was cumlSGD. No other models tested came even close (example CVs ~840, ~800, 720, 0.5!, etc).</p>\n<p>Did I shoot myself in the foot? Is better to spend the next full week re-training all of my models on the full data rather than keeping 33% holdout for a stacker that is worthless? Has any one else experimented with stacking their models yet? I'm running experiments now with GBDT and NNet models to finish off the experiment set..</p>",
      "rawMarkdown": "When I first started this competition, the first thing I did was look at how many samples exist, what their distribution was, and then setup my CV. I split the data into 6 buckets, and have been using four of those to build 4-fold CV models and kept the remaining two buckets for stacking. I had just been doing OOF blending for ensembling. But since I have ~48 trained models now, I tried to give the holdout stacker a go. \n\nThe results were...... not what I was expecting.\n\nI've tried a variety of model types for the stacker: GPR, LogisticRegression, Linear, Ridge, Lasso, Elastic, SVR (RBF+Linear), MultinomialNB, RF, SGDR. Basically everything under the sun except a GBDT and a NNet. No matter what ML model I use to stack the holdout predictions, the model's validation performance has never beaten a linear combination of weights derived from a straight scipy.optimize.minimize. How is that possible??\n\nI base my understanding of stacking on @kazanova's great Kaggle coursera series. I believe stacking with a ML model (particularly, one that has more than one parameter per base-model's prediction) is that it should be superior because it can detect \"zones\" where one model preforms good and prefer it those regions; and then prefer other models in regions where their results are better.\n\nBest Stack:\n- Val: 0.874972 (again, this is trained+validated on 2/6th of the data, split in two folds)\n- LB: 0.8767\n\nRandom Minimize Model Weights:\n- Val 0.875204\n- LB: 0.8768\n\nThe best stacker model was cumlSGD. No other models tested came even close (example CVs ~840, ~800, 720, 0.5!, etc).\n\nDid I shoot myself in the foot? Is better to spend the next full week re-training all of my models on the full data rather than keeping 33% holdout for a stacker that is worthless? Has any one else experimented with stacking their models yet? I'm running experiments now with GBDT and NNet models to finish off the experiment set..",
      "votes": null
    },
    {
      "id": "1493457",
      "postDate": "08/27/2021 23:00:10",
      "content": "<p>From nearly experiments I did, it looked like the dataset was nearly homogeneous: either working on the whole dataset or on 10% of it, would lead to the same conclusions. Hence, similarly to you, I opted to train my models on 90% of the dataset and kept the rest as an holdhout I would use for stacking. As for now, I have trained around 15 models with holdout scores around 0.861-0.872. I used them to feed a second layer. As stackers, so far, I tried logistic regressions, MLP, LGBM and none of them were able to outperform the score I could reach with first layer models (both on holdout score and LB)….</p>\n<p>As you do, I currently reconsider my training methodology and wonder if I shouldn't train my model on 100% of the dataset …</p>",
      "rawMarkdown": "From nearly experiments I did, it looked like the dataset was nearly homogeneous: either working on the whole dataset or on 10% of it, would lead to the same conclusions. Hence, similarly to you, I opted to train my models on 90% of the dataset and kept the rest as an holdhout I would use for stacking. As for now, I have trained around 15 models with holdout scores around 0.861-0.872. I used them to feed a second layer. As stackers, so far, I tried logistic regressions, MLP, LGBM and none of them were able to outperform the score I could reach with first layer models (both on holdout score and LB)....\n\nAs you do, I currently reconsider my training methodology and wonder if I shouldn't train my model on 100% of the dataset ...",
      "votes": null
    },
    {
      "id": "1493478",
      "postDate": "08/27/2021 23:44:14",
      "content": "<p>Thank you for confirming my suspicion.</p>\n<p>It might be beneficial to look at the top ~300 worst FP and top ~300 worst FN by prediction error. If all models are essentially failing on the same samples, that conclusively answers why stacking only produces modest gains VS just training on more data.</p>\n<p>Looking at model correlations, they're all very similar (0.96-0.99999), so I guess they're all essentially learning the same thing. Some GW are just easier for all models to see (?), other GWs are more difficult. </p>\n<p>Now if only we could learn the mapping to make the hard ones to see one's easier…</p>",
      "rawMarkdown": "Thank you for confirming my suspicion.\n\nIt might be beneficial to look at the top ~300 worst FP and top ~300 worst FN by prediction error. If all models are essentially failing on the same samples, that conclusively answers why stacking only produces modest gains VS just training on more data.\n\nLooking at model correlations, they're all very similar (0.96-0.99999), so I guess they're all essentially learning the same thing. Some GW are just easier for all models to see (?), other GWs are more difficult. \n\nNow if only we could learn the mapping to make the hard ones to see one's easier...",
      "votes": null
    },
    {
      "id": "1494498",
      "postDate": "08/28/2021 17:53:56",
      "content": "<p><a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a> stated a week ago that their best model was a single model and that doing ensembling for ROC AUC competitions does not work well often.  </p>",
      "rawMarkdown": "cpmpml stated a week ago that their best model was a single model and that doing ensembling for ROC AUC competitions does not work well often.",
      "votes": null
    },
    {
      "id": "1494581",
      "postDate": "08/28/2021 19:44:22",
      "content": "<p>Not exactly.  I said our best model achieved 0.880x and that ensembling does not help much with roc-auc in general.</p>",
      "rawMarkdown": "Not exactly.  I said our best model achieved 0.880x and that ensembling does not help much with roc-auc in general.",
      "votes": null
    },
    {
      "id": "1494587",
      "postDate": "08/28/2021 19:56:46",
      "content": "<p>I recall <a href=\"https://www.kaggle.com/felipebihaiek\" target=\"_blank\">@felipebihaiek</a>, I was in that conversation :-). OOF Ensembling gives me a little boost—but it's pennies to the dollar compared to my base models. Still worth it at current LB. Stacking, however, hasn't provided me with any return whatsoever, so now I'm burning cycles re-running experiments which is undesired.</p>",
      "rawMarkdown": "I recall @felipebihaiek, I was in that conversation :-). OOF Ensembling gives me a little boost—but it's pennies to the dollar compared to my base models. Still worth it at current LB. Stacking, however, hasn't provided me with any return whatsoever, so now I'm burning cycles re-running experiments which is undesired.",
      "votes": null
    },
    {
      "id": "1495016",
      "postDate": "08/29/2021 08:41:03",
      "content": "<p>Hi Thanks for sharing amazing information..<br>\ni think it is doubt question, but i wonder that why do you using ML model rather than deep learning?<br>\nAs i think that this problem is so deep and complicated, So deep model will be answer.</p>\n<p>And moreover, why are you using only 4-cv without 2 holdout?<br>\ni guess amounts of train data is significant in ml field.</p>\n<p>Thanks.</p>",
      "rawMarkdown": "Hi Thanks for sharing amazing information..\ni think it is doubt question, but i wonder that why do you using ML model rather than deep learning?\nAs i think that this problem is so deep and complicated, So deep model will be answer.\n\nAnd moreover, why are you using only 4-cv without 2 holdout?\ni guess amounts of train data is significant in ml field.\n\nThanks.",
      "votes": null
    },
    {
      "id": "1495231",
      "postDate": "08/29/2021 11:55:01",
      "content": "<p>the 2nd layer stack should be one that optimizes AUC directly …. maybe that will work</p>",
      "rawMarkdown": "the 2nd layer stack should be one that optimizes AUC directly .... maybe that will work",
      "votes": null
    },
    {
      "id": "1496313",
      "postDate": "08/30/2021 09:12:37",
      "content": "<p>Thanks for sharing amazing information.. =))</p>",
      "rawMarkdown": "Thanks for sharing amazing information.. =))",
      "votes": null
    },
    {
      "id": "1561012",
      "postDate": "10/27/2021 09:09:27",
      "content": "<p>Hey All,</p>\n<p>Thank you all for taking part in our competition. The participation has been overwhelmingly positive. We are currently conducting a survey to gauge the demographic and outreach achieved. Kindly spare 2min and fill in this survey <a href=\"https://forms.gle/QP9L16niPexozyhu5\" target=\"_blank\">https://forms.gle/QP9L16niPexozyhu5</a>.</p>\n<p>Thank you all,</p>\n<p>Regards,<br>\nChris</p>",
      "rawMarkdown": "Hey All,\n\nThank you all for taking part in our competition. The participation has been overwhelmingly positive. We are currently conducting a survey to gauge the demographic and outreach achieved. Kindly spare 2min and fill in this survey https://forms.gle/QP9L16niPexozyhu5.\n\nThank you all,\n\nRegards,\nChris",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1493457,
      "author_name": "fabiendaniel",
      "author_url": "",
      "post_date": "08/27/2021 23:00:10",
      "content": "<p>From nearly experiments I did, it looked like the dataset was nearly homogeneous: either working on the whole dataset or on 10% of it, would lead to the same conclusions. Hence, similarly to you, I opted to train my models on 90% of the dataset and kept the rest as an holdhout I would use for stacking. As for now, I have trained around 15 models with holdout scores around 0.861-0.872. I used them to feed a second layer. As stackers, so far, I tried logistic regressions, MLP, LGBM and none of them were able to outperform the score I could reach with first layer models (both on holdout score and LB)….</p>\n<p>As you do, I currently reconsider my training methodology and wonder if I shouldn't train my model on 100% of the dataset …</p>",
      "votes": null,
      "replies": [
        {
          "id": 1493478,
          "author_name": "authman",
          "author_url": "",
          "post_date": "08/27/2021 23:44:14",
          "content": "<p>Thank you for confirming my suspicion.</p>\n<p>It might be beneficial to look at the top ~300 worst FP and top ~300 worst FN by prediction error. If all models are essentially failing on the same samples, that conclusively answers why stacking only produces modest gains VS just training on more data.</p>\n<p>Looking at model correlations, they're all very similar (0.96-0.99999), so I guess they're all essentially learning the same thing. Some GW are just easier for all models to see (?), other GWs are more difficult. </p>\n<p>Now if only we could learn the mapping to make the hard ones to see one's easier…</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1494498,
      "author_name": "felipebihaiek",
      "author_url": "",
      "post_date": "08/28/2021 17:53:56",
      "content": "<p><a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a> stated a week ago that their best model was a single model and that doing ensembling for ROC AUC competitions does not work well often.  </p>",
      "votes": null,
      "replies": [
        {
          "id": 1494581,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "08/28/2021 19:44:22",
          "content": "<p>Not exactly.  I said our best model achieved 0.880x and that ensembling does not help much with roc-auc in general.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1494587,
          "author_name": "authman",
          "author_url": "",
          "post_date": "08/28/2021 19:56:46",
          "content": "<p>I recall <a href=\"https://www.kaggle.com/felipebihaiek\" target=\"_blank\">@felipebihaiek</a>, I was in that conversation :-). OOF Ensembling gives me a little boost—but it's pennies to the dollar compared to my base models. Still worth it at current LB. Stacking, however, hasn't provided me with any return whatsoever, so now I'm burning cycles re-running experiments which is undesired.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1495016,
      "author_name": "leemop",
      "author_url": "",
      "post_date": "08/29/2021 08:41:03",
      "content": "<p>Hi Thanks for sharing amazing information..<br>\ni think it is doubt question, but i wonder that why do you using ML model rather than deep learning?<br>\nAs i think that this problem is so deep and complicated, So deep model will be answer.</p>\n<p>And moreover, why are you using only 4-cv without 2 holdout?<br>\ni guess amounts of train data is significant in ml field.</p>\n<p>Thanks.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1495231,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "08/29/2021 11:55:01",
      "content": "<p>the 2nd layer stack should be one that optimizes AUC directly …. maybe that will work</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1496313,
      "author_name": "firefliesqn",
      "author_url": "",
      "post_date": "08/30/2021 09:12:37",
      "content": "<p>Thanks for sharing amazing information.. =))</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1561012,
      "author_name": "zerafachris",
      "author_url": "",
      "post_date": "10/27/2021 09:09:27",
      "content": "<p>Hey All,</p>\n<p>Thank you all for taking part in our competition. The participation has been overwhelmingly positive. We are currently conducting a survey to gauge the demographic and outreach achieved. Kindly spare 2min and fill in this survey <a href=\"https://forms.gle/QP9L16niPexozyhu5\" target=\"_blank\">https://forms.gle/QP9L16niPexozyhu5</a>.</p>\n<p>Thank you all,</p>\n<p>Regards,<br>\nChris</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1493437": "When I first started this competition, the first thing I did was look at how many samples exist, what their distribution was, and then setup my CV. I split the data into 6 buckets, and have been using four of those to build 4-fold CV models and kept the remaining two buckets for stacking. I had just been doing OOF blending for ensembling. But since I have ~48 trained models now, I tried to give the holdout stacker a go. \n\nThe results were...... not what I was expecting.\n\nI've tried a variety of model types for the stacker: GPR, LogisticRegression, Linear, Ridge, Lasso, Elastic, SVR (RBF+Linear), MultinomialNB, RF, SGDR. Basically everything under the sun except a GBDT and a NNet. No matter what ML model I use to stack the holdout predictions, the model's validation performance has never beaten a linear combination of weights derived from a straight scipy.optimize.minimize. How is that possible??\n\nI base my understanding of stacking on @kazanova's great Kaggle coursera series. I believe stacking with a ML model (particularly, one that has more than one parameter per base-model's prediction) is that it should be superior because it can detect \"zones\" where one model preforms good and prefer it those regions; and then prefer other models in regions where their results are better.\n\nBest Stack:\n- Val: 0.874972 (again, this is trained+validated on 2/6th of the data, split in two folds)\n- LB: 0.8767\n\nRandom Minimize Model Weights:\n- Val 0.875204\n- LB: 0.8768\n\nThe best stacker model was cumlSGD. No other models tested came even close (example CVs ~840, ~800, 720, 0.5!, etc).\n\nDid I shoot myself in the foot? Is better to spend the next full week re-training all of my models on the full data rather than keeping 33% holdout for a stacker that is worthless? Has any one else experimented with stacking their models yet? I'm running experiments now with GBDT and NNet models to finish off the experiment set..",
    "1493457": "From nearly experiments I did, it looked like the dataset was nearly homogeneous: either working on the whole dataset or on 10% of it, would lead to the same conclusions. Hence, similarly to you, I opted to train my models on 90% of the dataset and kept the rest as an holdhout I would use for stacking. As for now, I have trained around 15 models with holdout scores around 0.861-0.872. I used them to feed a second layer. As stackers, so far, I tried logistic regressions, MLP, LGBM and none of them were able to outperform the score I could reach with first layer models (both on holdout score and LB)....\n\nAs you do, I currently reconsider my training methodology and wonder if I shouldn't train my model on 100% of the dataset ...",
    "1493478": "Thank you for confirming my suspicion.\n\nIt might be beneficial to look at the top ~300 worst FP and top ~300 worst FN by prediction error. If all models are essentially failing on the same samples, that conclusively answers why stacking only produces modest gains VS just training on more data.\n\nLooking at model correlations, they're all very similar (0.96-0.99999), so I guess they're all essentially learning the same thing. Some GW are just easier for all models to see (?), other GWs are more difficult. \n\nNow if only we could learn the mapping to make the hard ones to see one's easier...",
    "1494498": "cpmpml stated a week ago that their best model was a single model and that doing ensembling for ROC AUC competitions does not work well often.",
    "1494581": "Not exactly.  I said our best model achieved 0.880x and that ensembling does not help much with roc-auc in general.",
    "1494587": "I recall @felipebihaiek, I was in that conversation :-). OOF Ensembling gives me a little boost—but it's pennies to the dollar compared to my base models. Still worth it at current LB. Stacking, however, hasn't provided me with any return whatsoever, so now I'm burning cycles re-running experiments which is undesired.",
    "1495016": "Hi Thanks for sharing amazing information..\ni think it is doubt question, but i wonder that why do you using ML model rather than deep learning?\nAs i think that this problem is so deep and complicated, So deep model will be answer.\n\nAnd moreover, why are you using only 4-cv without 2 holdout?\ni guess amounts of train data is significant in ml field.\n\nThanks.",
    "1495231": "the 2nd layer stack should be one that optimizes AUC directly .... maybe that will work",
    "1496313": "Thanks for sharing amazing information.. =))",
    "1561012": "Hey All,\n\nThank you all for taking part in our competition. The participation has been overwhelmingly positive. We are currently conducting a survey to gauge the demographic and outreach achieved. Kindly spare 2min and fill in this survey https://forms.gle/QP9L16niPexozyhu5.\n\nThank you all,\n\nRegards,\nChris"
  },
  "source": "meta"
}