{
  "id": 274825,
  "title": "LGB stacking AUC less than the best model's AUC",
  "url": "/competitions/g2net-gravitational-wave-detection/discussion/274825",
  "author_name": "",
  "post_date": "2021-09-27T18:55:32.613923800Z",
  "votes": 15,
  "comment_count": 9,
  "views": 0,
  "content": "<p>Did a dry run of the final stacking pipeline consisting of 3 trained models.</p>\n<p>cnn1d+attention in the frequency domain: CV: 877, LB: 878<br>\ncnn1d base+heavily filtered data: CV: 873, LB: 875<br>\nresnet34d: CV: 870, LB:872</p>\n<p>LightGBM stacking using oofs: CV: 877, LB: 8782</p>\n<p>Did I just encounter a leakage? What happened here? 😂</p>",
  "messages": [
    {
      "id": "1525970",
      "postDate": "09/27/2021 18:55:32",
      "content": "<p>Did a dry run of the final stacking pipeline consisting of 3 trained models.</p>\n<p>cnn1d+attention in the frequency domain: CV: 877, LB: 878<br>\ncnn1d base+heavily filtered data: CV: 873, LB: 875<br>\nresnet34d: CV: 870, LB:872</p>\n<p>LightGBM stacking using oofs: CV: 877, LB: 8782</p>\n<p>Did I just encounter a leakage? What happened here? 😂</p>",
      "rawMarkdown": "Did a dry run of the final stacking pipeline consisting of 3 trained models.\n\ncnn1d+attention in the frequency domain: CV: 877, LB: 878\ncnn1d base+heavily filtered data: CV: 873, LB: 875\nresnet34d: CV: 870, LB:872\n\nLightGBM stacking using oofs: CV: 877, LB: 8782\n\nDid I just encounter a leakage? What happened here? 😂",
      "votes": null
    },
    {
      "id": "1526014",
      "postDate": "09/27/2021 19:27:07",
      "content": "<p>I've never been able to outperform optimized weight blend with ml stacking in any competition, especially with ranking metrics haha </p>",
      "rawMarkdown": "I've never been able to outperform optimized weight blend with ml stacking in any competition, especially with ranking metrics haha",
      "votes": null
    },
    {
      "id": "1526042",
      "postDate": "09/27/2021 19:50:54",
      "content": "<p>Can confirm that stacking hasn't been helpful here, made a thread about it ~month ago.</p>",
      "rawMarkdown": "Can confirm that stacking hasn't been helpful here, made a thread about it ~month ago.",
      "votes": null
    },
    {
      "id": "1526046",
      "postDate": "09/27/2021 19:53:45",
      "content": "<p>i suspect overfitting. <br>\ndo note that you have skewed distribution. </p>\n<p>also, there are noisy +ve samples. it is easy to improve these scores by memorizing their noisy components (instead of signals) but this is actually poor generalisation</p>\n<hr>\n<p>you can try:</p>\n<p>model1 - fold1,2,3<br>\nmodel2 - fold1,2,3<br>\nmodel3 - fold1,2,3</p>\n<p>stack1 = model1-fold1, model2-fold1,  model3-fold1,   + (more augment if you can)<br>\nstack2 = model1-fold2, model2-fold2,  model3-fold2, +  (more augment if you can)<br>\n…</p>\n<p>ensemble stack1,stack2 …</p>",
      "rawMarkdown": "i suspect overfitting. \ndo note that you have skewed distribution. \n\nalso, there are noisy +ve samples. it is easy to improve these scores by memorizing their noisy components (instead of signals) but this is actually poor generalisation\n\n--- \nyou can try:\n\nmodel1 - fold1,2,3\nmodel2 - fold1,2,3\nmodel3 - fold1,2,3\n\nstack1 = model1-fold1, model2-fold1,  model3-fold1,   + (more augment if you can)\nstack2 = model1-fold2, model2-fold2,  model3-fold2, +  (more augment if you can)\n...\n\nensemble stack1,stack2 ...",
      "votes": null
    },
    {
      "id": "1526561",
      "postDate": "09/28/2021 07:41:44",
      "content": "<p>Ensembling with roc-auc metric is always disappointing compared to, say, ensembling when the metric is RMSE.  </p>",
      "rawMarkdown": "Ensembling with roc-auc metric is always disappointing compared to, say, ensembling when the metric is RMSE.",
      "votes": null
    },
    {
      "id": "1526829",
      "postDate": "09/28/2021 11:13:44",
      "content": "<p>calibration issue?</p>",
      "rawMarkdown": "calibration issue?",
      "votes": null
    },
    {
      "id": "1527143",
      "postDate": "09/28/2021 14:40:50",
      "content": "<p>Heng likes to drop hints everywhere, I am rooting for Heng to get gold. 😆</p>",
      "rawMarkdown": "Heng likes to drop hints everywhere, I am rooting for Heng to get gold. 😆",
      "votes": null
    },
    {
      "id": "1527209",
      "postDate": "09/28/2021 15:30:26",
      "content": "<p>Haha, count me in.</p>",
      "rawMarkdown": "Haha, count me in.",
      "votes": null
    },
    {
      "id": "1527426",
      "postDate": "09/28/2021 18:59:10",
      "content": "<blockquote>\n  <p>calibration issue?</p>\n</blockquote>\n<p>No.  Same issue with MAE.</p>\n<p>Say you average two models A and B.  Take a sample with ground truth y, and predictions a,b for the two models.  Let c = (a+b)/2 be the blend prediction.  </p>\n<p>Let's look at different metrics.</p>\n<p><strong>MSE (also RMSE)</strong></p>\n<p>The function<code>f(y) = (y - x)^2</code> is convex, therefore  <code>(y - c)^2 &lt;= [(y - a)^2 + (y - b)^2] / 2</code>. Moreover the inequality is an equality only if a = b.  Therefore the blend error is always smaller than the best of the blended model for that example, except when the two models predict the same value. The improvement is quasi certain.</p>\n<p><strong>MAE</strong><br>\nThe function<code>f(y) = abs(y - x)</code> is convex, therefore  <code>abs(y - c) &lt;= [abs(y - a) + abs(y - b)] / 2</code>.</p>\n<p>However, this is an equality whenever <code>(y -a)</code> and <code>(y - b)</code>have the same sign.  The blend is better than best model only when the model predictions are on opposite side of ground truth.  This may not be true often depending on the case.  The improvement is not quasi certain.</p>\n<p>For AUC we need to consider pairs, but the reasoning is a bit similar than for MAE.  The cases where a blend improve are not quasi certain.</p>",
      "rawMarkdown": "> calibration issue?\n\nNo.  Same issue with MAE.\n\nSay you average two models A and B.  Take a sample with ground truth y, and predictions a,b for the two models.  Let c = (a+b)/2 be the blend prediction.  \n\nLet's look at different metrics.\n\n**MSE (also RMSE)**\n\nThe function` f(y) = (y - x)^2` is convex, therefore  `(y - c)^2 <= [(y - a)^2 + (y - b)^2] / 2`. Moreover the inequality is an equality only if a = b.  Therefore the blend error is always smaller than the best of the blended model for that example, except when the two models predict the same value. The improvement is quasi certain.\n\n**MAE**\nThe function` f(y) = abs(y - x)` is convex, therefore  `abs(y - c) <= [abs(y - a) + abs(y - b)] / 2`.\n\nHowever, this is an equality whenever `(y -a)` and `(y - b) `have the same sign.  The blend is better than best model only when the model predictions are on opposite side of ground truth.  This may not be true often depending on the case.  The improvement is not quasi certain.\n\nFor AUC we need to consider pairs, but the reasoning is a bit similar than for MAE.  The cases where a blend improve are not quasi certain.",
      "votes": null
    },
    {
      "id": "1559738",
      "postDate": "10/27/2021 07:10:38",
      "content": "<p>Hey All,</p>\n<p>Thank you all for taking part in our competition. The participation has been overwhelmingly positive. We are currently conducting a survey to gauge the demographic and outreach achieved. Kindly spare 2min and fill in this survey <a href=\"https://forms.gle/QP9L16niPexozyhu5\" target=\"_blank\">https://forms.gle/QP9L16niPexozyhu5</a>.</p>\n<p>Thank you all,</p>\n<p>Regards,<br>\nChris</p>",
      "rawMarkdown": "Hey All,\n\nThank you all for taking part in our competition. The participation has been overwhelmingly positive. We are currently conducting a survey to gauge the demographic and outreach achieved. Kindly spare 2min and fill in this survey https://forms.gle/QP9L16niPexozyhu5.\n\nThank you all,\n\nRegards,\nChris",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1526014,
      "author_name": "nischaydnk",
      "author_url": "",
      "post_date": "09/27/2021 19:27:07",
      "content": "<p>I've never been able to outperform optimized weight blend with ml stacking in any competition, especially with ranking metrics haha </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1526042,
      "author_name": "authman",
      "author_url": "",
      "post_date": "09/27/2021 19:50:54",
      "content": "<p>Can confirm that stacking hasn't been helpful here, made a thread about it ~month ago.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1526046,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "09/27/2021 19:53:45",
      "content": "<p>i suspect overfitting. <br>\ndo note that you have skewed distribution. </p>\n<p>also, there are noisy +ve samples. it is easy to improve these scores by memorizing their noisy components (instead of signals) but this is actually poor generalisation</p>\n<hr>\n<p>you can try:</p>\n<p>model1 - fold1,2,3<br>\nmodel2 - fold1,2,3<br>\nmodel3 - fold1,2,3</p>\n<p>stack1 = model1-fold1, model2-fold1,  model3-fold1,   + (more augment if you can)<br>\nstack2 = model1-fold2, model2-fold2,  model3-fold2, +  (more augment if you can)<br>\n…</p>\n<p>ensemble stack1,stack2 …</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1526561,
      "author_name": "cpmpml",
      "author_url": "",
      "post_date": "09/28/2021 07:41:44",
      "content": "<p>Ensembling with roc-auc metric is always disappointing compared to, say, ensembling when the metric is RMSE.  </p>",
      "votes": null,
      "replies": [
        {
          "id": 1526829,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "09/28/2021 11:13:44",
          "content": "<p>calibration issue?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1527143,
          "author_name": "scaomath",
          "author_url": "",
          "post_date": "09/28/2021 14:40:50",
          "content": "<p>Heng likes to drop hints everywhere, I am rooting for Heng to get gold. 😆</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1527209,
          "author_name": "wuliaokaola",
          "author_url": "",
          "post_date": "09/28/2021 15:30:26",
          "content": "<p>Haha, count me in.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1527426,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "09/28/2021 18:59:10",
          "content": "<blockquote>\n  <p>calibration issue?</p>\n</blockquote>\n<p>No.  Same issue with MAE.</p>\n<p>Say you average two models A and B.  Take a sample with ground truth y, and predictions a,b for the two models.  Let c = (a+b)/2 be the blend prediction.  </p>\n<p>Let's look at different metrics.</p>\n<p><strong>MSE (also RMSE)</strong></p>\n<p>The function<code>f(y) = (y - x)^2</code> is convex, therefore  <code>(y - c)^2 &lt;= [(y - a)^2 + (y - b)^2] / 2</code>. Moreover the inequality is an equality only if a = b.  Therefore the blend error is always smaller than the best of the blended model for that example, except when the two models predict the same value. The improvement is quasi certain.</p>\n<p><strong>MAE</strong><br>\nThe function<code>f(y) = abs(y - x)</code> is convex, therefore  <code>abs(y - c) &lt;= [abs(y - a) + abs(y - b)] / 2</code>.</p>\n<p>However, this is an equality whenever <code>(y -a)</code> and <code>(y - b)</code>have the same sign.  The blend is better than best model only when the model predictions are on opposite side of ground truth.  This may not be true often depending on the case.  The improvement is not quasi certain.</p>\n<p>For AUC we need to consider pairs, but the reasoning is a bit similar than for MAE.  The cases where a blend improve are not quasi certain.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1559738,
      "author_name": "zerafachris",
      "author_url": "",
      "post_date": "10/27/2021 07:10:38",
      "content": "<p>Hey All,</p>\n<p>Thank you all for taking part in our competition. The participation has been overwhelmingly positive. We are currently conducting a survey to gauge the demographic and outreach achieved. Kindly spare 2min and fill in this survey <a href=\"https://forms.gle/QP9L16niPexozyhu5\" target=\"_blank\">https://forms.gle/QP9L16niPexozyhu5</a>.</p>\n<p>Thank you all,</p>\n<p>Regards,<br>\nChris</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1525970": "Did a dry run of the final stacking pipeline consisting of 3 trained models.\n\ncnn1d+attention in the frequency domain: CV: 877, LB: 878\ncnn1d base+heavily filtered data: CV: 873, LB: 875\nresnet34d: CV: 870, LB:872\n\nLightGBM stacking using oofs: CV: 877, LB: 8782\n\nDid I just encounter a leakage? What happened here? 😂",
    "1526014": "I've never been able to outperform optimized weight blend with ml stacking in any competition, especially with ranking metrics haha",
    "1526042": "Can confirm that stacking hasn't been helpful here, made a thread about it ~month ago.",
    "1526046": "i suspect overfitting. \ndo note that you have skewed distribution. \n\nalso, there are noisy +ve samples. it is easy to improve these scores by memorizing their noisy components (instead of signals) but this is actually poor generalisation\n\n--- \nyou can try:\n\nmodel1 - fold1,2,3\nmodel2 - fold1,2,3\nmodel3 - fold1,2,3\n\nstack1 = model1-fold1, model2-fold1,  model3-fold1,   + (more augment if you can)\nstack2 = model1-fold2, model2-fold2,  model3-fold2, +  (more augment if you can)\n...\n\nensemble stack1,stack2 ...",
    "1526561": "Ensembling with roc-auc metric is always disappointing compared to, say, ensembling when the metric is RMSE.",
    "1526829": "calibration issue?",
    "1527143": "Heng likes to drop hints everywhere, I am rooting for Heng to get gold. 😆",
    "1527209": "Haha, count me in.",
    "1527426": "> calibration issue?\n\nNo.  Same issue with MAE.\n\nSay you average two models A and B.  Take a sample with ground truth y, and predictions a,b for the two models.  Let c = (a+b)/2 be the blend prediction.  \n\nLet's look at different metrics.\n\n**MSE (also RMSE)**\n\nThe function` f(y) = (y - x)^2` is convex, therefore  `(y - c)^2 <= [(y - a)^2 + (y - b)^2] / 2`. Moreover the inequality is an equality only if a = b.  Therefore the blend error is always smaller than the best of the blended model for that example, except when the two models predict the same value. The improvement is quasi certain.\n\n**MAE**\nThe function` f(y) = abs(y - x)` is convex, therefore  `abs(y - c) <= [abs(y - a) + abs(y - b)] / 2`.\n\nHowever, this is an equality whenever `(y -a)` and `(y - b) `have the same sign.  The blend is better than best model only when the model predictions are on opposite side of ground truth.  This may not be true often depending on the case.  The improvement is not quasi certain.\n\nFor AUC we need to consider pairs, but the reasoning is a bit similar than for MAE.  The cases where a blend improve are not quasi certain.",
    "1559738": "Hey All,\n\nThank you all for taking part in our competition. The participation has been overwhelmingly positive. We are currently conducting a survey to gauge the demographic and outreach achieved. Kindly spare 2min and fill in this survey https://forms.gle/QP9L16niPexozyhu5.\n\nThank you all,\n\nRegards,\nChris"
  },
  "source": "meta"
}