{
  "id": 445048,
  "title": "Ideas why ensembles work so well on LB?",
  "url": "/competitions/open-problems-single-cell-perturbations/discussion/445048",
  "author_name": "",
  "post_date": "2023-10-04T23:42:50.231641500Z",
  "votes": 11,
  "comment_count": 9,
  "views": 0,
  "content": "<p>Some of you may have noticed: </p>\n<p>Simply ensembling results from different models (even ones with bad scores) can easily boost the score on LB. Is this normal for regression problems?</p>\n<p>I am trying to understand why. Here are some of my thoughts of the possible reasons:</p>\n<ol>\n<li>Evaluation metric sensitive to outliers -&gt; ensembling smooth out the outlier predictions that have high value.</li>\n<li>The above mentioned reason, plus maybe that public test set has less outliers compared to the train set. So we are just overfitting the LB and big shakeup inevitable when private test results come.</li>\n<li>Not overfitting, there are just a lof of valid signals that one single/type of model alone cannot capture.</li>\n</ol>\n<p>I don't know which one is true, maybe a bit of everything? would really like to know what you think.</p>",
  "messages": [
    {
      "id": "2467854",
      "postDate": "10/04/2023 23:42:50",
      "content": "<p>Some of you may have noticed: </p>\n<p>Simply ensembling results from different models (even ones with bad scores) can easily boost the score on LB. Is this normal for regression problems?</p>\n<p>I am trying to understand why. Here are some of my thoughts of the possible reasons:</p>\n<ol>\n<li>Evaluation metric sensitive to outliers -&gt; ensembling smooth out the outlier predictions that have high value.</li>\n<li>The above mentioned reason, plus maybe that public test set has less outliers compared to the train set. So we are just overfitting the LB and big shakeup inevitable when private test results come.</li>\n<li>Not overfitting, there are just a lof of valid signals that one single/type of model alone cannot capture.</li>\n</ol>\n<p>I don't know which one is true, maybe a bit of everything? would really like to know what you think.</p>",
      "rawMarkdown": "Some of you may have noticed: \n\nSimply ensembling results from different models (even ones with bad scores) can easily boost the score on LB. Is this normal for regression problems?\n\nI am trying to understand why. Here are some of my thoughts of the possible reasons:\n\n1. Evaluation metric sensitive to outliers -> ensembling smooth out the outlier predictions that have high value.\n2. The above mentioned reason, plus maybe that public test set has less outliers compared to the train set. So we are just overfitting the LB and big shakeup inevitable when private test results come.\n3. Not overfitting, there are just a lof of valid signals that one single/type of model alone cannot capture.\n\nI don't know which one is true, maybe a bit of everything? would really like to know what you think.",
      "votes": null
    },
    {
      "id": "2467901",
      "postDate": "10/05/2023 01:40:51",
      "content": "<p>My experience in this competition is that ensembling is not working as well as it has in past competitions.</p>\n<p>Run a decent number and have not yet achieved a LB score better than the best submission values in the group of files I used for the ensemble.  This is without any weights.  </p>",
      "rawMarkdown": "My experience in this competition is that ensembling is not working as well as it has in past competitions.\n\nRun a decent number and have not yet achieved a LB score better than the best submission values in the group of files I used for the ensemble.  This is without any weights.",
      "votes": null
    },
    {
      "id": "2467924",
      "postDate": "10/05/2023 02:53:53",
      "content": "<blockquote>\n  <p>My experience in this competition is that ensembling is not working as well as it has in past competitions.</p>\n  <p>Run a decent number and have not yet achieved a LB score better than the best submission values in the group of files I used for the ensemble.  This is without any weights.</p>\n</blockquote>\n<p>hmm.. that's a very different experience from mine. Maybe my models are overfitting and yours are underfitting.</p>",
      "rawMarkdown": "> My experience in this competition is that ensembling is not working as well as it has in past competitions.\n> \n> Run a decent number and have not yet achieved a LB score better than the best submission values in the group of files I used for the ensemble.  This is without any weights.\n\nhmm.. that's a very different experience from mine. Maybe my models are overfitting and yours are underfitting.",
      "votes": null
    },
    {
      "id": "2468044",
      "postDate": "10/05/2023 06:22:13",
      "content": "<p>Have someone tried to ensemble by CV , not just LB ?<br>\nI have no time to do it yet. </p>",
      "rawMarkdown": "Have someone tried to ensemble by CV , not just LB ?\nI have no time to do it yet.",
      "votes": null
    },
    {
      "id": "2468410",
      "postDate": "10/05/2023 13:23:36",
      "content": "<p>Can you elaborate a bit more about \"ensemble by CV\"? Thanks!</p>",
      "rawMarkdown": "Can you elaborate a bit more about \"ensemble by CV\"? Thanks!",
      "votes": null
    },
    {
      "id": "2468522",
      "postDate": "10/05/2023 15:13:57",
      "content": "<p>I mean to choose weights for ensembling using CV scores , but not just looking on LB scores. Otherwise it is overfit to LB</p>",
      "rawMarkdown": "I mean to choose weights for ensembling using CV scores , but not just looking on LB scores. Otherwise it is overfit to LB",
      "votes": null
    },
    {
      "id": "2468556",
      "postDate": "10/05/2023 15:45:10",
      "content": "<p>Ah, I see. Thank you.</p>\n<p>I was simply taking the average. I didn't think of chossing weights at all, thinking the risk of verfitting would be high.  </p>",
      "rawMarkdown": "Ah, I see. Thank you.\n\nI was simply taking the average. I didn't think of chossing weights at all, thinking the risk of verfitting would be high.",
      "votes": null
    },
    {
      "id": "2483122",
      "postDate": "10/15/2023 13:48:02",
      "content": "<p>I tend to lean towards the second point more. It does seem like we might be over-optimizing for the specific characteristics of the 40% test data, and as a result, our models might struggle to generalize to the remaining 60%. This could explain the relatively lower performance on the public leaderboard compared to what we observe in our own experiments.</p>\n<p>A noteworthy observation is that, even without the inclusion of new features, we tend to achieve relatively lower scores.  This raises questions about whether our models are adept at generalization or if we are potentially overfitting to the public test data.</p>\n<p>In my experience, ensembles do provide a substantial boost in performance, but it's worth mentioning that a well-optimized single complex model can already yield a commendable score, around 0.575.</p>\n<p>I'm curious to learn more about your approaches to mitigate overfitting on the public test set. Managing overfitting is a common challenge in this competition, and exchanging insights on this front could be immensely beneficial. Thank you!</p>",
      "rawMarkdown": "I tend to lean towards the second point more. It does seem like we might be over-optimizing for the specific characteristics of the 40% test data, and as a result, our models might struggle to generalize to the remaining 60%. This could explain the relatively lower performance on the public leaderboard compared to what we observe in our own experiments.\n\nA noteworthy observation is that, even without the inclusion of new features, we tend to achieve relatively lower scores.  This raises questions about whether our models are adept at generalization or if we are potentially overfitting to the public test data.\n\nIn my experience, ensembles do provide a substantial boost in performance, but it's worth mentioning that a well-optimized single complex model can already yield a commendable score, around 0.575.\n\nI'm curious to learn more about your approaches to mitigate overfitting on the public test set. Managing overfitting is a common challenge in this competition, and exchanging insights on this front could be immensely beneficial. Thank you!",
      "votes": null
    },
    {
      "id": "2485135",
      "postDate": "10/17/2023 00:51:34",
      "content": "<blockquote>\n  <p>I tend to lean towards the second point more. It does seem like we might be over-optimizing for the specific characteristics of the 40% test data, and as a result, our models might struggle to generalize to the remaining 60%. This could explain the relatively lower performance on the public leaderboard compared to what we observe in our own experiments.</p>\n  <p>A noteworthy observation is that, even without the inclusion of new features, we tend to achieve relatively lower scores.  This raises questions about whether our models are adept at generalization or if we are potentially overfitting to the public test data.</p>\n  <p>In my experience, ensembles do provide a substantial boost in performance, but it's worth mentioning that a well-optimized single complex model can already yield a commendable score, around 0.575.</p>\n  <p>I'm curious to learn more about your approaches to mitigate overfitting on the public test set. Managing overfitting is a common challenge in this competition, and exchanging insights on this front could be immensely beneficial. Thank you!</p>\n</blockquote>\n<p>Same experience here. Ensembling give an extra boost of about ~0.004 for the LB score beyond the single best model. I am probably the wrong person to answer on advice how not to overfit, because I am literally overfitting the LB… lol</p>\n<p>Stick to one of the two CV strategies proposed by MT and Ambros, depending on your approach. That's probably the safest thing to do. My model training doesn't strictly follows those, so my CV and LB are all over the places, that actually created a mess for me…</p>",
      "rawMarkdown": "> I tend to lean towards the second point more. It does seem like we might be over-optimizing for the specific characteristics of the 40% test data, and as a result, our models might struggle to generalize to the remaining 60%. This could explain the relatively lower performance on the public leaderboard compared to what we observe in our own experiments.\n> \n> A noteworthy observation is that, even without the inclusion of new features, we tend to achieve relatively lower scores.  This raises questions about whether our models are adept at generalization or if we are potentially overfitting to the public test data.\n> \n> In my experience, ensembles do provide a substantial boost in performance, but it's worth mentioning that a well-optimized single complex model can already yield a commendable score, around 0.575.\n> \n> I'm curious to learn more about your approaches to mitigate overfitting on the public test set. Managing overfitting is a common challenge in this competition, and exchanging insights on this front could be immensely beneficial. Thank you!\n\n\nSame experience here. Ensembling give an extra boost of about ~0.004 for the LB score beyond the single best model. I am probably the wrong person to answer on advice how not to overfit, because I am literally overfitting the LB... lol\n\nStick to one of the two CV strategies proposed by MT and Ambros, depending on your approach. That's probably the safest thing to do. My model training doesn't strictly follows those, so my CV and LB are all over the places, that actually created a mess for me...",
      "votes": null
    },
    {
      "id": "2485588",
      "postDate": "10/17/2023 09:38:58",
      "content": "<p>Thank you! I've also considered the idea of using semi-supervised learning based on the best LB score submission. However, this approach could potentially lead to overfitting on the 40% test data, resulting in a significant drop in performance on the full test dataset.</p>",
      "rawMarkdown": "Thank you! I've also considered the idea of using semi-supervised learning based on the best LB score submission. However, this approach could potentially lead to overfitting on the 40% test data, resulting in a significant drop in performance on the full test dataset.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2467901,
      "author_name": "pcjimmmy",
      "author_url": "",
      "post_date": "10/05/2023 01:40:51",
      "content": "<p>My experience in this competition is that ensembling is not working as well as it has in past competitions.</p>\n<p>Run a decent number and have not yet achieved a LB score better than the best submission values in the group of files I used for the ensemble.  This is without any weights.  </p>",
      "votes": null,
      "replies": [
        {
          "id": 2467924,
          "author_name": "qihuaz",
          "author_url": "",
          "post_date": "10/05/2023 02:53:53",
          "content": "<blockquote>\n  <p>My experience in this competition is that ensembling is not working as well as it has in past competitions.</p>\n  <p>Run a decent number and have not yet achieved a LB score better than the best submission values in the group of files I used for the ensemble.  This is without any weights.</p>\n</blockquote>\n<p>hmm.. that's a very different experience from mine. Maybe my models are overfitting and yours are underfitting.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2468044,
      "author_name": "alexandervc",
      "author_url": "",
      "post_date": "10/05/2023 06:22:13",
      "content": "<p>Have someone tried to ensemble by CV , not just LB ?<br>\nI have no time to do it yet. </p>",
      "votes": null,
      "replies": [
        {
          "id": 2468410,
          "author_name": "qihuaz",
          "author_url": "",
          "post_date": "10/05/2023 13:23:36",
          "content": "<p>Can you elaborate a bit more about \"ensemble by CV\"? Thanks!</p>",
          "votes": null,
          "replies": [
            {
              "id": 2468522,
              "author_name": "alexandervc",
              "author_url": "",
              "post_date": "10/05/2023 15:13:57",
              "content": "<p>I mean to choose weights for ensembling using CV scores , but not just looking on LB scores. Otherwise it is overfit to LB</p>",
              "votes": null,
              "replies": [
                {
                  "id": 2468556,
                  "author_name": "qihuaz",
                  "author_url": "",
                  "post_date": "10/05/2023 15:45:10",
                  "content": "<p>Ah, I see. Thank you.</p>\n<p>I was simply taking the average. I didn't think of chossing weights at all, thinking the risk of verfitting would be high.  </p>",
                  "votes": null,
                  "replies": []
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 2483122,
      "author_name": "kishanvavdara",
      "author_url": "",
      "post_date": "10/15/2023 13:48:02",
      "content": "<p>I tend to lean towards the second point more. It does seem like we might be over-optimizing for the specific characteristics of the 40% test data, and as a result, our models might struggle to generalize to the remaining 60%. This could explain the relatively lower performance on the public leaderboard compared to what we observe in our own experiments.</p>\n<p>A noteworthy observation is that, even without the inclusion of new features, we tend to achieve relatively lower scores.  This raises questions about whether our models are adept at generalization or if we are potentially overfitting to the public test data.</p>\n<p>In my experience, ensembles do provide a substantial boost in performance, but it's worth mentioning that a well-optimized single complex model can already yield a commendable score, around 0.575.</p>\n<p>I'm curious to learn more about your approaches to mitigate overfitting on the public test set. Managing overfitting is a common challenge in this competition, and exchanging insights on this front could be immensely beneficial. Thank you!</p>",
      "votes": null,
      "replies": [
        {
          "id": 2485135,
          "author_name": "qihuaz",
          "author_url": "",
          "post_date": "10/17/2023 00:51:34",
          "content": "<blockquote>\n  <p>I tend to lean towards the second point more. It does seem like we might be over-optimizing for the specific characteristics of the 40% test data, and as a result, our models might struggle to generalize to the remaining 60%. This could explain the relatively lower performance on the public leaderboard compared to what we observe in our own experiments.</p>\n  <p>A noteworthy observation is that, even without the inclusion of new features, we tend to achieve relatively lower scores.  This raises questions about whether our models are adept at generalization or if we are potentially overfitting to the public test data.</p>\n  <p>In my experience, ensembles do provide a substantial boost in performance, but it's worth mentioning that a well-optimized single complex model can already yield a commendable score, around 0.575.</p>\n  <p>I'm curious to learn more about your approaches to mitigate overfitting on the public test set. Managing overfitting is a common challenge in this competition, and exchanging insights on this front could be immensely beneficial. Thank you!</p>\n</blockquote>\n<p>Same experience here. Ensembling give an extra boost of about ~0.004 for the LB score beyond the single best model. I am probably the wrong person to answer on advice how not to overfit, because I am literally overfitting the LB… lol</p>\n<p>Stick to one of the two CV strategies proposed by MT and Ambros, depending on your approach. That's probably the safest thing to do. My model training doesn't strictly follows those, so my CV and LB are all over the places, that actually created a mess for me…</p>",
          "votes": null,
          "replies": [
            {
              "id": 2485588,
              "author_name": "kishanvavdara",
              "author_url": "",
              "post_date": "10/17/2023 09:38:58",
              "content": "<p>Thank you! I've also considered the idea of using semi-supervised learning based on the best LB score submission. However, this approach could potentially lead to overfitting on the 40% test data, resulting in a significant drop in performance on the full test dataset.</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2467854": "Some of you may have noticed: \n\nSimply ensembling results from different models (even ones with bad scores) can easily boost the score on LB. Is this normal for regression problems?\n\nI am trying to understand why. Here are some of my thoughts of the possible reasons:\n\n1. Evaluation metric sensitive to outliers -> ensembling smooth out the outlier predictions that have high value.\n2. The above mentioned reason, plus maybe that public test set has less outliers compared to the train set. So we are just overfitting the LB and big shakeup inevitable when private test results come.\n3. Not overfitting, there are just a lof of valid signals that one single/type of model alone cannot capture.\n\nI don't know which one is true, maybe a bit of everything? would really like to know what you think.",
    "2467901": "My experience in this competition is that ensembling is not working as well as it has in past competitions.\n\nRun a decent number and have not yet achieved a LB score better than the best submission values in the group of files I used for the ensemble.  This is without any weights.",
    "2467924": "> My experience in this competition is that ensembling is not working as well as it has in past competitions.\n> \n> Run a decent number and have not yet achieved a LB score better than the best submission values in the group of files I used for the ensemble.  This is without any weights.\n\nhmm.. that's a very different experience from mine. Maybe my models are overfitting and yours are underfitting.",
    "2468044": "Have someone tried to ensemble by CV , not just LB ?\nI have no time to do it yet.",
    "2468410": "Can you elaborate a bit more about \"ensemble by CV\"? Thanks!",
    "2468522": "I mean to choose weights for ensembling using CV scores , but not just looking on LB scores. Otherwise it is overfit to LB",
    "2468556": "Ah, I see. Thank you.\n\nI was simply taking the average. I didn't think of chossing weights at all, thinking the risk of verfitting would be high.",
    "2483122": "I tend to lean towards the second point more. It does seem like we might be over-optimizing for the specific characteristics of the 40% test data, and as a result, our models might struggle to generalize to the remaining 60%. This could explain the relatively lower performance on the public leaderboard compared to what we observe in our own experiments.\n\nA noteworthy observation is that, even without the inclusion of new features, we tend to achieve relatively lower scores.  This raises questions about whether our models are adept at generalization or if we are potentially overfitting to the public test data.\n\nIn my experience, ensembles do provide a substantial boost in performance, but it's worth mentioning that a well-optimized single complex model can already yield a commendable score, around 0.575.\n\nI'm curious to learn more about your approaches to mitigate overfitting on the public test set. Managing overfitting is a common challenge in this competition, and exchanging insights on this front could be immensely beneficial. Thank you!",
    "2485135": "> I tend to lean towards the second point more. It does seem like we might be over-optimizing for the specific characteristics of the 40% test data, and as a result, our models might struggle to generalize to the remaining 60%. This could explain the relatively lower performance on the public leaderboard compared to what we observe in our own experiments.\n> \n> A noteworthy observation is that, even without the inclusion of new features, we tend to achieve relatively lower scores.  This raises questions about whether our models are adept at generalization or if we are potentially overfitting to the public test data.\n> \n> In my experience, ensembles do provide a substantial boost in performance, but it's worth mentioning that a well-optimized single complex model can already yield a commendable score, around 0.575.\n> \n> I'm curious to learn more about your approaches to mitigate overfitting on the public test set. Managing overfitting is a common challenge in this competition, and exchanging insights on this front could be immensely beneficial. Thank you!\n\n\nSame experience here. Ensembling give an extra boost of about ~0.004 for the LB score beyond the single best model. I am probably the wrong person to answer on advice how not to overfit, because I am literally overfitting the LB... lol\n\nStick to one of the two CV strategies proposed by MT and Ambros, depending on your approach. That's probably the safest thing to do. My model training doesn't strictly follows those, so my CV and LB are all over the places, that actually created a mess for me...",
    "2485588": "Thank you! I've also considered the idea of using semi-supervised learning based on the best LB score submission. However, this approach could potentially lead to overfitting on the 40% test data, resulting in a significant drop in performance on the full test dataset."
  },
  "source": "meta"
}