{
  "id": 334100,
  "title": "Ensemble with 3 decimals",
  "url": "/competitions/amex-default-prediction/discussion/334100",
  "author_name": "",
  "post_date": "2022-06-29T19:18:02.523879600Z",
  "votes": 11,
  "comment_count": 8,
  "views": 0,
  "content": "<p>Good day kagglers, </p>\n<p>I've been trying to improve my score by fitting several models and making an ensemble. First, I fit a Catboost or an Xgboost several times with different features or parameters. Then, I submit my best models to the leaderboard and record their score. Finally, i make my weighted ensemble based on the scores I get (the best scores have a higher weight). </p>\n<p>So far it's been good and I've been slowly improving my score over time. However, lately the difference between my models is very small and i can't see it in the leaderboard since we only have 3 decimals. I would like to give a bigger weight to a model even though the score is almost the same (0.7964 vs 0.7967, for example). <strong>I was wondering if there is a recommended way to mix my private validation score with leaderboard score for an ensemble?</strong> Does that approach even make sense at all? </p>\n<p>And an additional question, how many models would you recommend to ensemble? is there a recommended limit? the more the better?</p>\n<p>Many thanks in advance!! </p>\n<p>Fernando</p>",
  "messages": [
    {
      "id": "1837667",
      "postDate": "06/29/2022 19:18:02",
      "content": "<p>Good day kagglers, </p>\n<p>I've been trying to improve my score by fitting several models and making an ensemble. First, I fit a Catboost or an Xgboost several times with different features or parameters. Then, I submit my best models to the leaderboard and record their score. Finally, i make my weighted ensemble based on the scores I get (the best scores have a higher weight). </p>\n<p>So far it's been good and I've been slowly improving my score over time. However, lately the difference between my models is very small and i can't see it in the leaderboard since we only have 3 decimals. I would like to give a bigger weight to a model even though the score is almost the same (0.7964 vs 0.7967, for example). <strong>I was wondering if there is a recommended way to mix my private validation score with leaderboard score for an ensemble?</strong> Does that approach even make sense at all? </p>\n<p>And an additional question, how many models would you recommend to ensemble? is there a recommended limit? the more the better?</p>\n<p>Many thanks in advance!! </p>\n<p>Fernando</p>",
      "rawMarkdown": "Good day kagglers, \n\nI've been trying to improve my score by fitting several models and making an ensemble. First, I fit a Catboost or an Xgboost several times with different features or parameters. Then, I submit my best models to the leaderboard and record their score. Finally, i make my weighted ensemble based on the scores I get (the best scores have a higher weight). \n\nSo far it's been good and I've been slowly improving my score over time. However, lately the difference between my models is very small and i can't see it in the leaderboard since we only have 3 decimals. I would like to give a bigger weight to a model even though the score is almost the same (0.7964 vs 0.7967, for example). **I was wondering if there is a recommended way to mix my private validation score with leaderboard score for an ensemble?** Does that approach even make sense at all? \n\nAnd an additional question, how many models would you recommend to ensemble? is there a recommended limit? the more the better?\n\n\nMany thanks in advance!! \n\nFernando",
      "votes": null
    },
    {
      "id": "1837935",
      "postDate": "06/30/2022 04:11:47",
      "content": "<p>We have 3 decimals - but sorting your LB by public score you can infer a 4th.  </p>",
      "rawMarkdown": "We have 3 decimals - but sorting your LB by public score you can infer a 4th.",
      "votes": null
    },
    {
      "id": "1838379",
      "postDate": "06/30/2022 13:15:54",
      "content": "<p>Welp. I had no idea I could sort my submissions! thanks a lot. I'm still ondering if I could mix the validation scores to make it more stable.. I guess I'm just going to experiment</p>",
      "rawMarkdown": "Welp. I had no idea I could sort my submissions! thanks a lot. I'm still ondering if I could mix the validation scores to make it more stable.. I guess I'm just going to experiment",
      "votes": null
    },
    {
      "id": "1838912",
      "postDate": "07/01/2022 02:07:02",
      "content": "<p>I would suggest focusing on FE and improving one model at a time. Then ensemble the best models close to the end of the competition.</p>",
      "rawMarkdown": "I would suggest focusing on FE and improving one model at a time. Then ensemble the best models close to the end of the competition.",
      "votes": null
    },
    {
      "id": "1839306",
      "postDate": "07/01/2022 09:41:35",
      "content": "<p>Thanks! Makes sense. Would you say the ensemble will make it more stable? I have a feeling the leaderboard will change a lot with the public score</p>",
      "rawMarkdown": "Thanks! Makes sense. Would you say the ensemble will make it more stable? I have a feeling the leaderboard will change a lot with the public score",
      "votes": null
    },
    {
      "id": "1839321",
      "postDate": "07/01/2022 10:09:55",
      "content": "<p>This is my first competition but I have the same feeling. </p>\n<p>If you have time, read Chapter 6 of The Kaggle Book which talks about designing good validation. <br>\nHere is a relevant summary of chapter 6 regarding the shake-ups:</p>\n<ul>\n<li>There is little adaptive overfitting; in other words, public standings usually do hold in the unveiled private leaderboard.</li>\n<li>Most shake-ups are due to random fluctuations and overcrowded rankings where competitors are too near to each other.</li>\n<li>Shake-ups happen when training set is very small or training data is not independent and identically distributed.</li>\n</ul>\n<p>I think the second reason will be the main reason for shake-up in this competition.</p>",
      "rawMarkdown": "This is my first competition but I have the same feeling. \n\nIf you have time, read Chapter 6 of The Kaggle Book which talks about designing good validation. \nHere is a relevant summary of chapter 6 regarding the shake-ups:\n\n- There is little adaptive overfitting; in other words, public standings usually do hold in the unveiled private leaderboard.\n- Most shake-ups are due to random fluctuations and overcrowded rankings where competitors are too near to each other.\n- Shake-ups happen when training set is very small or training data is not independent and identically distributed.\n\nI think the second reason will be the main reason for shake-up in this competition.",
      "votes": null
    },
    {
      "id": "1839323",
      "postDate": "07/01/2022 10:10:39",
      "content": "<p>This is extremely useful. Thank you!</p>",
      "rawMarkdown": "This is extremely useful. Thank you!",
      "votes": null
    },
    {
      "id": "1840481",
      "postDate": "07/02/2022 09:24:22",
      "content": "<p>Thanks! I've heard about the Kaggle book, you think its worth it? I'll give it a try :)</p>\n<p>I also think the second reason will be the main one</p>",
      "rawMarkdown": "Thanks! I've heard about the Kaggle book, you think its worth it? I'll give it a try :)\n\nI also think the second reason will be the main one",
      "votes": null
    },
    {
      "id": "1842323",
      "postDate": "07/03/2022 23:03:20",
      "content": "<p>As you have 2 submissions to choose for final, it's recommended to choose 1 with the highest score on a public LB and 1 with the highest score on your validation set, so in this case you are kind of protected of shake up and still choose best one for the situation when there is no shakeup at all.</p>\n<p>So, as for your question, you might give bigger weights for better submissions on a public LB and for the second submission you can give bigger weights for models that are better on a validation set. But yes, if these are overfitted models, they will make score worse on the test set, so in some cases it will be more safe to just average weights.</p>",
      "rawMarkdown": "As you have 2 submissions to choose for final, it's recommended to choose 1 with the highest score on a public LB and 1 with the highest score on your validation set, so in this case you are kind of protected of shake up and still choose best one for the situation when there is no shakeup at all.\n\nSo, as for your question, you might give bigger weights for better submissions on a public LB and for the second submission you can give bigger weights for models that are better on a validation set. But yes, if these are overfitted models, they will make score worse on the test set, so in some cases it will be more safe to just average weights.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1837935,
      "author_name": "pcjimmmy",
      "author_url": "",
      "post_date": "06/30/2022 04:11:47",
      "content": "<p>We have 3 decimals - but sorting your LB by public score you can infer a 4th.  </p>",
      "votes": null,
      "replies": [
        {
          "id": 1838379,
          "author_name": "heyspaceturtle",
          "author_url": "",
          "post_date": "06/30/2022 13:15:54",
          "content": "<p>Welp. I had no idea I could sort my submissions! thanks a lot. I'm still ondering if I could mix the validation scores to make it more stable.. I guess I'm just going to experiment</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1839323,
          "author_name": "mohammadrahmati",
          "author_url": "",
          "post_date": "07/01/2022 10:10:39",
          "content": "<p>This is extremely useful. Thank you!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1838912,
      "author_name": "mohammadrahmati",
      "author_url": "",
      "post_date": "07/01/2022 02:07:02",
      "content": "<p>I would suggest focusing on FE and improving one model at a time. Then ensemble the best models close to the end of the competition.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1839306,
          "author_name": "heyspaceturtle",
          "author_url": "",
          "post_date": "07/01/2022 09:41:35",
          "content": "<p>Thanks! Makes sense. Would you say the ensemble will make it more stable? I have a feeling the leaderboard will change a lot with the public score</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1839321,
          "author_name": "mohammadrahmati",
          "author_url": "",
          "post_date": "07/01/2022 10:09:55",
          "content": "<p>This is my first competition but I have the same feeling. </p>\n<p>If you have time, read Chapter 6 of The Kaggle Book which talks about designing good validation. <br>\nHere is a relevant summary of chapter 6 regarding the shake-ups:</p>\n<ul>\n<li>There is little adaptive overfitting; in other words, public standings usually do hold in the unveiled private leaderboard.</li>\n<li>Most shake-ups are due to random fluctuations and overcrowded rankings where competitors are too near to each other.</li>\n<li>Shake-ups happen when training set is very small or training data is not independent and identically distributed.</li>\n</ul>\n<p>I think the second reason will be the main reason for shake-up in this competition.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1840481,
          "author_name": "heyspaceturtle",
          "author_url": "",
          "post_date": "07/02/2022 09:24:22",
          "content": "<p>Thanks! I've heard about the Kaggle book, you think its worth it? I'll give it a try :)</p>\n<p>I also think the second reason will be the main one</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1842323,
          "author_name": "manwithaflower",
          "author_url": "",
          "post_date": "07/03/2022 23:03:20",
          "content": "<p>As you have 2 submissions to choose for final, it's recommended to choose 1 with the highest score on a public LB and 1 with the highest score on your validation set, so in this case you are kind of protected of shake up and still choose best one for the situation when there is no shakeup at all.</p>\n<p>So, as for your question, you might give bigger weights for better submissions on a public LB and for the second submission you can give bigger weights for models that are better on a validation set. But yes, if these are overfitted models, they will make score worse on the test set, so in some cases it will be more safe to just average weights.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1837667": "Good day kagglers, \n\nI've been trying to improve my score by fitting several models and making an ensemble. First, I fit a Catboost or an Xgboost several times with different features or parameters. Then, I submit my best models to the leaderboard and record their score. Finally, i make my weighted ensemble based on the scores I get (the best scores have a higher weight). \n\nSo far it's been good and I've been slowly improving my score over time. However, lately the difference between my models is very small and i can't see it in the leaderboard since we only have 3 decimals. I would like to give a bigger weight to a model even though the score is almost the same (0.7964 vs 0.7967, for example). **I was wondering if there is a recommended way to mix my private validation score with leaderboard score for an ensemble?** Does that approach even make sense at all? \n\nAnd an additional question, how many models would you recommend to ensemble? is there a recommended limit? the more the better?\n\n\nMany thanks in advance!! \n\nFernando",
    "1837935": "We have 3 decimals - but sorting your LB by public score you can infer a 4th.",
    "1838379": "Welp. I had no idea I could sort my submissions! thanks a lot. I'm still ondering if I could mix the validation scores to make it more stable.. I guess I'm just going to experiment",
    "1838912": "I would suggest focusing on FE and improving one model at a time. Then ensemble the best models close to the end of the competition.",
    "1839306": "Thanks! Makes sense. Would you say the ensemble will make it more stable? I have a feeling the leaderboard will change a lot with the public score",
    "1839321": "This is my first competition but I have the same feeling. \n\nIf you have time, read Chapter 6 of The Kaggle Book which talks about designing good validation. \nHere is a relevant summary of chapter 6 regarding the shake-ups:\n\n- There is little adaptive overfitting; in other words, public standings usually do hold in the unveiled private leaderboard.\n- Most shake-ups are due to random fluctuations and overcrowded rankings where competitors are too near to each other.\n- Shake-ups happen when training set is very small or training data is not independent and identically distributed.\n\nI think the second reason will be the main reason for shake-up in this competition.",
    "1839323": "This is extremely useful. Thank you!",
    "1840481": "Thanks! I've heard about the Kaggle book, you think its worth it? I'll give it a try :)\n\nI also think the second reason will be the main one",
    "1842323": "As you have 2 submissions to choose for final, it's recommended to choose 1 with the highest score on a public LB and 1 with the highest score on your validation set, so in this case you are kind of protected of shake up and still choose best one for the situation when there is no shakeup at all.\n\nSo, as for your question, you might give bigger weights for better submissions on a public LB and for the second submission you can give bigger weights for models that are better on a validation set. But yes, if these are overfitted models, they will make score worse on the test set, so in some cases it will be more safe to just average weights."
  },
  "source": "meta"
}