{
  "id": 339261,
  "title": "Why model scores are so sensitive to seed?",
  "url": "/competitions/amex-default-prediction/discussion/339261",
  "author_name": "",
  "post_date": "2022-07-24T02:03:14.081796600Z",
  "votes": 7,
  "comment_count": 6,
  "views": 0,
  "content": "<p>Hello Everyone! I've been using seed 42 to do experiment for this competition. However, I changed the seed to 168 today, and I found that my score dropped significantly. I'm using LGBM with the same parameters.<br>\nSeed 42: CV 0.79909, LB 0.799<br>\nSeed 168: CV 0.79819, LB 0.798<br>\nIs it because the seed affects how CV datasets are constructed so that the it worsened my modal performance? So is it better to use multiple seeds and ensemble all the submissions? Does anyone have the same issue? <br>\nHave a good weekend!</p>",
  "messages": [
    {
      "id": "1868460",
      "postDate": "07/24/2022 02:03:14",
      "content": "<p>Hello Everyone! I've been using seed 42 to do experiment for this competition. However, I changed the seed to 168 today, and I found that my score dropped significantly. I'm using LGBM with the same parameters.<br>\nSeed 42: CV 0.79909, LB 0.799<br>\nSeed 168: CV 0.79819, LB 0.798<br>\nIs it because the seed affects how CV datasets are constructed so that the it worsened my modal performance? So is it better to use multiple seeds and ensemble all the submissions? Does anyone have the same issue? <br>\nHave a good weekend!</p>",
      "rawMarkdown": "Hello Everyone! I've been using seed 42 to do experiment for this competition. However, I changed the seed to 168 today, and I found that my score dropped significantly. I'm using LGBM with the same parameters.\nSeed 42: CV 0.79909, LB 0.799\nSeed 168: CV 0.79819, LB 0.798\nIs it because the seed affects how CV datasets are constructed so that the it worsened my modal performance? So is it better to use multiple seeds and ensemble all the submissions? Does anyone have the same issue? \nHave a good weekend!",
      "votes": null
    },
    {
      "id": "1868510",
      "postDate": "07/24/2022 03:10:28",
      "content": "<p><a href=\"https://www.kaggle.com/ambrosm\" target=\"_blank\">@ambrosm</a> conducted some experiments about the correlation between seed and cv score <a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/329787\" target=\"_blank\">here</a></p>",
      "rawMarkdown": "ambrosm conducted some experiments about the correlation between seed and cv score [here](https://www.kaggle.com/competitions/amex-default-prediction/discussion/329787)",
      "votes": null
    },
    {
      "id": "1868738",
      "postDate": "07/24/2022 07:26:14",
      "content": "<p>This analysis is very detailed and highly useful to understand the relation between the state value and the final predictions</p>",
      "rawMarkdown": "This analysis is very detailed and highly useful to understand the relation between the state value and the final predictions",
      "votes": null
    },
    {
      "id": "1869245",
      "postDate": "07/24/2022 15:55:58",
      "content": "<p>I think it is because lgbm randomly selects some features and samples when establishing each tree, and the random seeds determine the randomness. Maybe those features randomly selected of seed 42 are more suitable for the model.<br>\nBut at present, they are all public LB's achievements, which may be different in private.<br>\nensemble random seeds may have a slight improvement, but I think it is best to do feature engineering and parameter tunning instead of paying too much attention to random seeds.</p>",
      "rawMarkdown": "I think it is because lgbm randomly selects some features and samples when establishing each tree, and the random seeds determine the randomness. Maybe those features randomly selected of seed 42 are more suitable for the model.\nBut at present, they are all public LB's achievements, which may be different in private.\nensemble random seeds may have a slight improvement, but I think it is best to do feature engineering and parameter tunning instead of paying too much attention to random seeds.",
      "votes": null
    },
    {
      "id": "1869344",
      "postDate": "07/24/2022 17:08:16",
      "content": "<p>I guess with a robust cross validation scheme randomness shouldnt play a huge role. Thanks for sharing</p>",
      "rawMarkdown": "I guess with a robust cross validation scheme randomness shouldnt play a huge role. Thanks for sharing",
      "votes": null
    },
    {
      "id": "1869385",
      "postDate": "07/24/2022 17:43:50",
      "content": "<blockquote>\n  <p>Is it because the seed affects how CV datasets are constructed so that the it worsened my modal performance?</p>\n</blockquote>\n<p>Absolutely if you're using the seed to select validation folds -- the same model trained and evaluated on different datasets can of course see variance in results. But people have reported significant variance even when holding validation folds fixed but changing the model's random seed, which I would say is unusual for GBT models. I expect them to typically be quite insensitive to seed, and I don't think I've encountered this sensitive behavior before.</p>\n<p>It's hard to know the exact cause, but I'd guess a few factors are at play--</p>\n<ol>\n<li>The metric is very noisy and sensitive to small changes</li>\n<li>The training dataset is pretty small (especially for such a noisy metric, only 450k is not ideal)</li>\n<li>The extreme importance of certain features like <code>P_2</code> causes wonky overfitting in the training process</li>\n</ol>\n<p>I think 3 is most speculative and hardest for me to reason about. But in my experience, \"super features\" can make models act weird and overlook smaller but still useful signals buried in the data. And most models in this competition use not just raw \"super features\", but also multiple derived copies that are similar (agg features etc.)! This is my guess as to why low column sampling values and DART work so well in this competition. They both help to prevent or combat bad adaption to \"super features\", especially early in the training process when the base trees are most influential. </p>\n<p>For me, a big theme of this competition is figuring out how to extract differentiating signal against a backdrop of dominating features and a sea of metric noise. Quite a challenge!  </p>",
      "rawMarkdown": "> Is it because the seed affects how CV datasets are constructed so that the it worsened my modal performance?\n\nAbsolutely if you're using the seed to select validation folds -- the same model trained and evaluated on different datasets can of course see variance in results. But people have reported significant variance even when holding validation folds fixed but changing the model's random seed, which I would say is unusual for GBT models. I expect them to typically be quite insensitive to seed, and I don't think I've encountered this sensitive behavior before.\n\nIt's hard to know the exact cause, but I'd guess a few factors are at play--\n1. The metric is very noisy and sensitive to small changes\n2. The training dataset is pretty small (especially for such a noisy metric, only 450k is not ideal)\n3. The extreme importance of certain features like `P_2` causes wonky overfitting in the training process\n\nI think 3 is most speculative and hardest for me to reason about. But in my experience, \"super features\" can make models act weird and overlook smaller but still useful signals buried in the data. And most models in this competition use not just raw \"super features\", but also multiple derived copies that are similar (agg features etc.)! This is my guess as to why low column sampling values and DART work so well in this competition. They both help to prevent or combat bad adaption to \"super features\", especially early in the training process when the base trees are most influential. \n\nFor me, a big theme of this competition is figuring out how to extract differentiating signal against a backdrop of dominating features and a sea of metric noise. Quite a challenge!",
      "votes": null
    },
    {
      "id": "1871434",
      "postDate": "07/26/2022 09:13:50",
      "content": "<p>Seed of the CV or of the model? <br>\nWhat CV do you use?</p>",
      "rawMarkdown": "Seed of the CV or of the model? \nWhat CV do you use?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1868510,
      "author_name": "raphael1123",
      "author_url": "",
      "post_date": "07/24/2022 03:10:28",
      "content": "<p><a href=\"https://www.kaggle.com/ambrosm\" target=\"_blank\">@ambrosm</a> conducted some experiments about the correlation between seed and cv score <a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/329787\" target=\"_blank\">here</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 1868738,
          "author_name": "ravi20076",
          "author_url": "",
          "post_date": "07/24/2022 07:26:14",
          "content": "<p>This analysis is very detailed and highly useful to understand the relation between the state value and the final predictions</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1869245,
      "author_name": "yichuwang123",
      "author_url": "",
      "post_date": "07/24/2022 15:55:58",
      "content": "<p>I think it is because lgbm randomly selects some features and samples when establishing each tree, and the random seeds determine the randomness. Maybe those features randomly selected of seed 42 are more suitable for the model.<br>\nBut at present, they are all public LB's achievements, which may be different in private.<br>\nensemble random seeds may have a slight improvement, but I think it is best to do feature engineering and parameter tunning instead of paying too much attention to random seeds.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1869344,
      "author_name": "rockford64",
      "author_url": "",
      "post_date": "07/24/2022 17:08:16",
      "content": "<p>I guess with a robust cross validation scheme randomness shouldnt play a huge role. Thanks for sharing</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1869385,
      "author_name": "aquatic",
      "author_url": "",
      "post_date": "07/24/2022 17:43:50",
      "content": "<blockquote>\n  <p>Is it because the seed affects how CV datasets are constructed so that the it worsened my modal performance?</p>\n</blockquote>\n<p>Absolutely if you're using the seed to select validation folds -- the same model trained and evaluated on different datasets can of course see variance in results. But people have reported significant variance even when holding validation folds fixed but changing the model's random seed, which I would say is unusual for GBT models. I expect them to typically be quite insensitive to seed, and I don't think I've encountered this sensitive behavior before.</p>\n<p>It's hard to know the exact cause, but I'd guess a few factors are at play--</p>\n<ol>\n<li>The metric is very noisy and sensitive to small changes</li>\n<li>The training dataset is pretty small (especially for such a noisy metric, only 450k is not ideal)</li>\n<li>The extreme importance of certain features like <code>P_2</code> causes wonky overfitting in the training process</li>\n</ol>\n<p>I think 3 is most speculative and hardest for me to reason about. But in my experience, \"super features\" can make models act weird and overlook smaller but still useful signals buried in the data. And most models in this competition use not just raw \"super features\", but also multiple derived copies that are similar (agg features etc.)! This is my guess as to why low column sampling values and DART work so well in this competition. They both help to prevent or combat bad adaption to \"super features\", especially early in the training process when the base trees are most influential. </p>\n<p>For me, a big theme of this competition is figuring out how to extract differentiating signal against a backdrop of dominating features and a sea of metric noise. Quite a challenge!  </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1871434,
      "author_name": "thedevastator",
      "author_url": "",
      "post_date": "07/26/2022 09:13:50",
      "content": "<p>Seed of the CV or of the model? <br>\nWhat CV do you use?</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1868460": "Hello Everyone! I've been using seed 42 to do experiment for this competition. However, I changed the seed to 168 today, and I found that my score dropped significantly. I'm using LGBM with the same parameters.\nSeed 42: CV 0.79909, LB 0.799\nSeed 168: CV 0.79819, LB 0.798\nIs it because the seed affects how CV datasets are constructed so that the it worsened my modal performance? So is it better to use multiple seeds and ensemble all the submissions? Does anyone have the same issue? \nHave a good weekend!",
    "1868510": "ambrosm conducted some experiments about the correlation between seed and cv score [here](https://www.kaggle.com/competitions/amex-default-prediction/discussion/329787)",
    "1868738": "This analysis is very detailed and highly useful to understand the relation between the state value and the final predictions",
    "1869245": "I think it is because lgbm randomly selects some features and samples when establishing each tree, and the random seeds determine the randomness. Maybe those features randomly selected of seed 42 are more suitable for the model.\nBut at present, they are all public LB's achievements, which may be different in private.\nensemble random seeds may have a slight improvement, but I think it is best to do feature engineering and parameter tunning instead of paying too much attention to random seeds.",
    "1869344": "I guess with a robust cross validation scheme randomness shouldnt play a huge role. Thanks for sharing",
    "1869385": "> Is it because the seed affects how CV datasets are constructed so that the it worsened my modal performance?\n\nAbsolutely if you're using the seed to select validation folds -- the same model trained and evaluated on different datasets can of course see variance in results. But people have reported significant variance even when holding validation folds fixed but changing the model's random seed, which I would say is unusual for GBT models. I expect them to typically be quite insensitive to seed, and I don't think I've encountered this sensitive behavior before.\n\nIt's hard to know the exact cause, but I'd guess a few factors are at play--\n1. The metric is very noisy and sensitive to small changes\n2. The training dataset is pretty small (especially for such a noisy metric, only 450k is not ideal)\n3. The extreme importance of certain features like `P_2` causes wonky overfitting in the training process\n\nI think 3 is most speculative and hardest for me to reason about. But in my experience, \"super features\" can make models act weird and overlook smaller but still useful signals buried in the data. And most models in this competition use not just raw \"super features\", but also multiple derived copies that are similar (agg features etc.)! This is my guess as to why low column sampling values and DART work so well in this competition. They both help to prevent or combat bad adaption to \"super features\", especially early in the training process when the base trees are most influential. \n\nFor me, a big theme of this competition is figuring out how to extract differentiating signal against a backdrop of dominating features and a sea of metric noise. Quite a challenge!",
    "1871434": "Seed of the CV or of the model? \nWhat CV do you use?"
  },
  "source": "meta"
}