{
  "id": 336957,
  "title": "Unstable random seed behaviour on public LB",
  "url": "/competitions/amex-default-prediction/discussion/336957",
  "author_name": "",
  "post_date": "2022-07-13T18:56:32.839862700Z",
  "votes": 28,
  "comment_count": 13,
  "views": 0,
  "content": "<p>Currently I have a single model that give different public results on different random seeds.<br>\nEach model uses the same parameters.</p>\n<p>For example:<br>\nLightGBM (dart) 5 Folds<br>\nSeed XA : CV 0.7994 AMEX; 0.96302 AUC / LB 0.798 <br>\nSeed XB:  CV 0.7984 AMEX; 0.96302 AUC / LB 0.797<br>\nSeed XC:  CV 0.7983 AMEX; 0.96301 AUC / LB 0.799<br>\nAVG(XA, XB, XC) CV 0.8004 AMEX; 0.96331 AUC; LB 0.798 AMEX</p>\n<p>I think it can be a result of a 20x weighted metric and sensitivity to positive objects.<br>\nSo you probably need to improve your public CV, but that doesn't guarantee we'll have a lottery at the end.</p>\n<p>So I think you need to be lucky or try <br>\nSubmission 1.  Your  public overfitting score 🙂<br>\nSubmission 2. Your trusted CV.</p>\n<p>I'll try to trust my CV.</p>",
  "messages": [
    {
      "id": "1854520",
      "postDate": "07/13/2022 18:56:32",
      "content": "<p>Currently I have a single model that give different public results on different random seeds.<br>\nEach model uses the same parameters.</p>\n<p>For example:<br>\nLightGBM (dart) 5 Folds<br>\nSeed XA : CV 0.7994 AMEX; 0.96302 AUC / LB 0.798 <br>\nSeed XB:  CV 0.7984 AMEX; 0.96302 AUC / LB 0.797<br>\nSeed XC:  CV 0.7983 AMEX; 0.96301 AUC / LB 0.799<br>\nAVG(XA, XB, XC) CV 0.8004 AMEX; 0.96331 AUC; LB 0.798 AMEX</p>\n<p>I think it can be a result of a 20x weighted metric and sensitivity to positive objects.<br>\nSo you probably need to improve your public CV, but that doesn't guarantee we'll have a lottery at the end.</p>\n<p>So I think you need to be lucky or try <br>\nSubmission 1.  Your  public overfitting score 🙂<br>\nSubmission 2. Your trusted CV.</p>\n<p>I'll try to trust my CV.</p>",
      "rawMarkdown": "Currently I have a single model that give different public results on different random seeds.\nEach model uses the same parameters.\n\nFor example:\nLightGBM (dart) 5 Folds\nSeed XA : CV 0.7994 AMEX; 0.96302 AUC / LB 0.798 \nSeed XB:  CV 0.7984 AMEX; 0.96302 AUC / LB 0.797\nSeed XC:  CV 0.7983 AMEX; 0.96301 AUC / LB 0.799\nAVG(XA, XB, XC) CV 0.8004 AMEX; 0.96331 AUC; LB 0.798 AMEX\n\nI think it can be a result of a 20x weighted metric and sensitivity to positive objects.\nSo you probably need to improve your public CV, but that doesn't guarantee we'll have a lottery at the end.\n\nSo I think you need to be lucky or try \nSubmission 1.  Your  public overfitting score 🙂\nSubmission 2. Your trusted CV.\n\nI'll try to trust my CV.",
      "votes": null
    },
    {
      "id": "1854537",
      "postDate": "07/13/2022 19:09:03",
      "content": "<p>thanks, random seed, you mean random seed for the models or the random seed for data spliting? I only tried one seed till now, which is 22 for both … Averaging of random seed predictions can sometimes cause slight overfittings, therefore I havn't try to do it. But it may work fine for this dataset, worth trying.</p>",
      "rawMarkdown": "thanks, random seed, you mean random seed for the models or the random seed for data spliting? I only tried one seed till now, which is 22 for both ... Averaging of random seed predictions can sometimes cause slight overfittings, therefore I havn't try to do it. But it may work fine for this dataset, worth trying.",
      "votes": null
    },
    {
      "id": "1854540",
      "postDate": "07/13/2022 19:12:47",
      "content": "<p>i mean splitting random seed with StratifiedKFold</p>",
      "rawMarkdown": "i mean splitting random seed with StratifiedKFold",
      "votes": null
    },
    {
      "id": "1854547",
      "postDate": "07/13/2022 19:25:39",
      "content": "<p>You're right. I use the same fixed number of trees for each fold for all seeds. Another case a public LB get worst because of overfitting on hold-out CV.</p>",
      "rawMarkdown": "You're right. I use the same fixed number of trees for each fold for all seeds. Another case a public LB get worst because of overfitting on hold-out CV.",
      "votes": null
    },
    {
      "id": "1854564",
      "postDate": "07/13/2022 19:46:11",
      "content": "<p>In a normal world, the answer to your question would be not to worry about differences on a third decimal place. That doesn't apply here because it is quite possible that differences at fourth decimal place will decide winners and medals.</p>\n<p>What you are seeing means that your folds have enough difference in their data distributions to produce different scores. You can probably see that by comparing all the metrics you listed at the end of each fold, rather than their average values. Fold stratification as done by most algorithms only assures that we have proportional label representation, but not all data points are equal. In a perfect world without noisy data, label stratification would be enough to give similar folds. If you think about it, that's actually what happens even with this dataset, because the differences you see are miniscule. It is just that even small score differences in a Kaggle environment can hurl one up or down the leaderboard by thousands of places.</p>\n<p>You may want to try <a href=\"https://scikit-learn.org/stable/modules/generated/sklearn.model_selection.RepeatedStratifiedKFold.html\" target=\"_blank\">RepeatedStratifiedKFold</a> which will make several fold splits, and therefore average out differences in fold distributions. Naturally, it will take longer to train.</p>",
      "rawMarkdown": "In a normal world, the answer to your question would be not to worry about differences on a third decimal place. That doesn't apply here because it is quite possible that differences at fourth decimal place will decide winners and medals.\n\nWhat you are seeing means that your folds have enough difference in their data distributions to produce different scores. You can probably see that by comparing all the metrics you listed at the end of each fold, rather than their average values. Fold stratification as done by most algorithms only assures that we have proportional label representation, but not all data points are equal. In a perfect world without noisy data, label stratification would be enough to give similar folds. If you think about it, that's actually what happens even with this dataset, because the differences you see are miniscule. It is just that even small score differences in a Kaggle environment can hurl one up or down the leaderboard by thousands of places.\n\nYou may want to try [RepeatedStratifiedKFold](https://scikit-learn.org/stable/modules/generated/sklearn.model_selection.RepeatedStratifiedKFold.html) which will make several fold splits, and therefore average out differences in fold distributions. Naturally, it will take longer to train.",
      "votes": null
    },
    {
      "id": "1854570",
      "postDate": "07/13/2022 19:57:40",
      "content": "<p>Noise is, perhaps, expected for this metric. LB under CV is a possible warning sign, however?</p>\n<p>Are you using early stopping and save best model? It's possible that early stopping on the best AMEX score is inflating your CV score artificially? Try early stopping on something more consistent (auc) and see if there's still noise, but your CV and your LB are a bit closer? (Because CV is 'worse' by not picking the best 4% score) Or try without early stopping, just pick a number of rounds based on prior inspection and stop there.</p>\n<p>Just a thought, others might have better insight :)</p>",
      "rawMarkdown": "Noise is, perhaps, expected for this metric. LB under CV is a possible warning sign, however?\n\nAre you using early stopping and save best model? It's possible that early stopping on the best AMEX score is inflating your CV score artificially? Try early stopping on something more consistent (auc) and see if there's still noise, but your CV and your LB are a bit closer? (Because CV is 'worse' by not picking the best 4% score) Or try without early stopping, just pick a number of rounds based on prior inspection and stop there.\n\nJust a thought, others might have better insight :)",
      "votes": null
    },
    {
      "id": "1854575",
      "postDate": "07/13/2022 20:03:59",
      "content": "<p>I don't use early_stopping_rounds. First I fit train.cv and evaluate AUC because of AMEX is noisy. Than I used single fixed best numbers of trees  for every hold-out train set without CV during fitting process because I use dart.</p>",
      "rawMarkdown": "I don't use early_stopping_rounds. First I fit train.cv and evaluate AUC because of AMEX is noisy. Than I used single fixed best numbers of trees  for every hold-out train set without CV during fitting process because I use dart.",
      "votes": null
    },
    {
      "id": "1854578",
      "postDate": "07/13/2022 20:07:08",
      "content": "<p>Thank you. I understand it. But it's a competition where we do a lot of efforts to do result better, but finally we also need to be lucky. In real world we don't play with random seeds if we have enough data and label representation😊</p>",
      "rawMarkdown": "Thank you. I understand it. But it's a competition where we do a lot of efforts to do result better, but finally we also need to be lucky. In real world we don't play with random seeds if we have enough data and label representation😊",
      "votes": null
    },
    {
      "id": "1854666",
      "postDate": "07/13/2022 22:21:52",
      "content": "<p>I have also noticed a large variation in local CV overall score <a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/336145#1850687\" target=\"_blank\">here</a>. If you use the same folds with a different XGB random seed, the overall CV has standard deviation <code>0.0012</code>. So it makes sense that LB score would also have large variation similar to standard deviation <code>0.0012</code>.</p>",
      "rawMarkdown": "I have also noticed a large variation in local CV overall score [here][1]. If you use the same folds with a different XGB random seed, the overall CV has standard deviation `0.0012`. So it makes sense that LB score would also have large variation similar to standard deviation `0.0012`.\n\n[1]: https://www.kaggle.com/competitions/amex-default-prediction/discussion/336145#1850687",
      "votes": null
    },
    {
      "id": "1854692",
      "postDate": "07/13/2022 23:15:45",
      "content": "<p>Exactly. Though I will reiterate that <strong>large</strong> is a very relative term and applies only to this competition setting. In most real-world applications, variations on a third decimal place are normal and would be considered insignificant.</p>\n<p>It is also worth adding that XGB/LGB hyper-parameters may exacerbate this variation. During a grid search, I see average fold variations in AmEx scores anywhere between <code>0.001</code> to <code>0.003</code>. Log-loss variations have narrower distribution (<code>0.001-0.002</code>). Variations in accuracy are almost always &lt; <code>0.001</code>.</p>",
      "rawMarkdown": "Exactly. Though I will reiterate that **large** is a very relative term and applies only to this competition setting. In most real-world applications, variations on a third decimal place are normal and would be considered insignificant.\n\nIt is also worth adding that XGB/LGB hyper-parameters may exacerbate this variation. During a grid search, I see average fold variations in AmEx scores anywhere between `0.001` to `0.003`. Log-loss variations have narrower distribution (`0.001-0.002`). Variations in accuracy are almost always < `0.001`.",
      "votes": null
    },
    {
      "id": "1854770",
      "postDate": "07/14/2022 02:03:27",
      "content": "<p>The competition need more luck</p>",
      "rawMarkdown": "The competition need more luck",
      "votes": null
    },
    {
      "id": "1854849",
      "postDate": "07/14/2022 05:22:57",
      "content": "<p>I would like to add that, if using 5-fold CV, each individual OOF score represents the score on 20% of the training data. The metric (specifically the <em>D</em> component) is so noisy that the competition organizers have had to provide just as much test data as training data just in order to obtain a three digit Public LB score (and a similar amount dedicated to calculating the Private LB score) hence each individual OOF score will have a large variance. <br>\nOnly after averaging over the five folds will this make the training data roughly the same size as the amount of data used to calculate the Public LB score (51% of the test data), so the uncertainty in the LB score should be commensurate with the uncertainty in the CV score.</p>",
      "rawMarkdown": "I would like to add that, if using 5-fold CV, each individual OOF score represents the score on 20% of the training data. The metric (specifically the *D* component) is so noisy that the competition organizers have had to provide just as much test data as training data just in order to obtain a three digit Public LB score (and a similar amount dedicated to calculating the Private LB score) hence each individual OOF score will have a large variance. \nOnly after averaging over the five folds will this make the training data roughly the same size as the amount of data used to calculate the Public LB score (51% of the test data), so the uncertainty in the LB score should be commensurate with the uncertainty in the CV score.",
      "votes": null
    },
    {
      "id": "1855770",
      "postDate": "07/14/2022 22:35:02",
      "content": "<p>I am not sure what ensembling strategy is working right now my top model seems to be better than any ensemble I try. Have tried around 10 strategies 😄. Though with low scoring models it was working</p>",
      "rawMarkdown": "I am not sure what ensembling strategy is working right now my top model seems to be better than any ensemble I try. Have tried around 10 strategies 😄. Though with low scoring models it was working",
      "votes": null
    },
    {
      "id": "1861074",
      "postDate": "07/18/2022 18:51:14",
      "content": "<p>I am taking a piece of your luck.</p>",
      "rawMarkdown": "I am taking a piece of your luck.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1854537,
      "author_name": "meli19",
      "author_url": "",
      "post_date": "07/13/2022 19:09:03",
      "content": "<p>thanks, random seed, you mean random seed for the models or the random seed for data spliting? I only tried one seed till now, which is 22 for both … Averaging of random seed predictions can sometimes cause slight overfittings, therefore I havn't try to do it. But it may work fine for this dataset, worth trying.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1854540,
          "author_name": "bogorodvo",
          "author_url": "",
          "post_date": "07/13/2022 19:12:47",
          "content": "<p>i mean splitting random seed with StratifiedKFold</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1854547,
          "author_name": "bogorodvo",
          "author_url": "",
          "post_date": "07/13/2022 19:25:39",
          "content": "<p>You're right. I use the same fixed number of trees for each fold for all seeds. Another case a public LB get worst because of overfitting on hold-out CV.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1854564,
      "author_name": "tilii7",
      "author_url": "",
      "post_date": "07/13/2022 19:46:11",
      "content": "<p>In a normal world, the answer to your question would be not to worry about differences on a third decimal place. That doesn't apply here because it is quite possible that differences at fourth decimal place will decide winners and medals.</p>\n<p>What you are seeing means that your folds have enough difference in their data distributions to produce different scores. You can probably see that by comparing all the metrics you listed at the end of each fold, rather than their average values. Fold stratification as done by most algorithms only assures that we have proportional label representation, but not all data points are equal. In a perfect world without noisy data, label stratification would be enough to give similar folds. If you think about it, that's actually what happens even with this dataset, because the differences you see are miniscule. It is just that even small score differences in a Kaggle environment can hurl one up or down the leaderboard by thousands of places.</p>\n<p>You may want to try <a href=\"https://scikit-learn.org/stable/modules/generated/sklearn.model_selection.RepeatedStratifiedKFold.html\" target=\"_blank\">RepeatedStratifiedKFold</a> which will make several fold splits, and therefore average out differences in fold distributions. Naturally, it will take longer to train.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1854578,
          "author_name": "bogorodvo",
          "author_url": "",
          "post_date": "07/13/2022 20:07:08",
          "content": "<p>Thank you. I understand it. But it's a competition where we do a lot of efforts to do result better, but finally we also need to be lucky. In real world we don't play with random seeds if we have enough data and label representation😊</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1854849,
          "author_name": "carlmcbrideellis",
          "author_url": "",
          "post_date": "07/14/2022 05:22:57",
          "content": "<p>I would like to add that, if using 5-fold CV, each individual OOF score represents the score on 20% of the training data. The metric (specifically the <em>D</em> component) is so noisy that the competition organizers have had to provide just as much test data as training data just in order to obtain a three digit Public LB score (and a similar amount dedicated to calculating the Private LB score) hence each individual OOF score will have a large variance. <br>\nOnly after averaging over the five folds will this make the training data roughly the same size as the amount of data used to calculate the Public LB score (51% of the test data), so the uncertainty in the LB score should be commensurate with the uncertainty in the CV score.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1854570,
      "author_name": "roberthatch",
      "author_url": "",
      "post_date": "07/13/2022 19:57:40",
      "content": "<p>Noise is, perhaps, expected for this metric. LB under CV is a possible warning sign, however?</p>\n<p>Are you using early stopping and save best model? It's possible that early stopping on the best AMEX score is inflating your CV score artificially? Try early stopping on something more consistent (auc) and see if there's still noise, but your CV and your LB are a bit closer? (Because CV is 'worse' by not picking the best 4% score) Or try without early stopping, just pick a number of rounds based on prior inspection and stop there.</p>\n<p>Just a thought, others might have better insight :)</p>",
      "votes": null,
      "replies": [
        {
          "id": 1854575,
          "author_name": "bogorodvo",
          "author_url": "",
          "post_date": "07/13/2022 20:03:59",
          "content": "<p>I don't use early_stopping_rounds. First I fit train.cv and evaluate AUC because of AMEX is noisy. Than I used single fixed best numbers of trees  for every hold-out train set without CV during fitting process because I use dart.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1854666,
      "author_name": "cdeotte",
      "author_url": "",
      "post_date": "07/13/2022 22:21:52",
      "content": "<p>I have also noticed a large variation in local CV overall score <a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/336145#1850687\" target=\"_blank\">here</a>. If you use the same folds with a different XGB random seed, the overall CV has standard deviation <code>0.0012</code>. So it makes sense that LB score would also have large variation similar to standard deviation <code>0.0012</code>.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1854692,
          "author_name": "tilii7",
          "author_url": "",
          "post_date": "07/13/2022 23:15:45",
          "content": "<p>Exactly. Though I will reiterate that <strong>large</strong> is a very relative term and applies only to this competition setting. In most real-world applications, variations on a third decimal place are normal and would be considered insignificant.</p>\n<p>It is also worth adding that XGB/LGB hyper-parameters may exacerbate this variation. During a grid search, I see average fold variations in AmEx scores anywhere between <code>0.001</code> to <code>0.003</code>. Log-loss variations have narrower distribution (<code>0.001-0.002</code>). Variations in accuracy are almost always &lt; <code>0.001</code>.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1854770,
      "author_name": "mahluo",
      "author_url": "",
      "post_date": "07/14/2022 02:03:27",
      "content": "<p>The competition need more luck</p>",
      "votes": null,
      "replies": [
        {
          "id": 1861074,
          "author_name": "zhehaoliang",
          "author_url": "",
          "post_date": "07/18/2022 18:51:14",
          "content": "<p>I am taking a piece of your luck.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1855770,
      "author_name": "gauravbrills",
      "author_url": "",
      "post_date": "07/14/2022 22:35:02",
      "content": "<p>I am not sure what ensembling strategy is working right now my top model seems to be better than any ensemble I try. Have tried around 10 strategies 😄. Though with low scoring models it was working</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1854520": "Currently I have a single model that give different public results on different random seeds.\nEach model uses the same parameters.\n\nFor example:\nLightGBM (dart) 5 Folds\nSeed XA : CV 0.7994 AMEX; 0.96302 AUC / LB 0.798 \nSeed XB:  CV 0.7984 AMEX; 0.96302 AUC / LB 0.797\nSeed XC:  CV 0.7983 AMEX; 0.96301 AUC / LB 0.799\nAVG(XA, XB, XC) CV 0.8004 AMEX; 0.96331 AUC; LB 0.798 AMEX\n\nI think it can be a result of a 20x weighted metric and sensitivity to positive objects.\nSo you probably need to improve your public CV, but that doesn't guarantee we'll have a lottery at the end.\n\nSo I think you need to be lucky or try \nSubmission 1.  Your  public overfitting score 🙂\nSubmission 2. Your trusted CV.\n\nI'll try to trust my CV.",
    "1854537": "thanks, random seed, you mean random seed for the models or the random seed for data spliting? I only tried one seed till now, which is 22 for both ... Averaging of random seed predictions can sometimes cause slight overfittings, therefore I havn't try to do it. But it may work fine for this dataset, worth trying.",
    "1854540": "i mean splitting random seed with StratifiedKFold",
    "1854547": "You're right. I use the same fixed number of trees for each fold for all seeds. Another case a public LB get worst because of overfitting on hold-out CV.",
    "1854564": "In a normal world, the answer to your question would be not to worry about differences on a third decimal place. That doesn't apply here because it is quite possible that differences at fourth decimal place will decide winners and medals.\n\nWhat you are seeing means that your folds have enough difference in their data distributions to produce different scores. You can probably see that by comparing all the metrics you listed at the end of each fold, rather than their average values. Fold stratification as done by most algorithms only assures that we have proportional label representation, but not all data points are equal. In a perfect world without noisy data, label stratification would be enough to give similar folds. If you think about it, that's actually what happens even with this dataset, because the differences you see are miniscule. It is just that even small score differences in a Kaggle environment can hurl one up or down the leaderboard by thousands of places.\n\nYou may want to try [RepeatedStratifiedKFold](https://scikit-learn.org/stable/modules/generated/sklearn.model_selection.RepeatedStratifiedKFold.html) which will make several fold splits, and therefore average out differences in fold distributions. Naturally, it will take longer to train.",
    "1854570": "Noise is, perhaps, expected for this metric. LB under CV is a possible warning sign, however?\n\nAre you using early stopping and save best model? It's possible that early stopping on the best AMEX score is inflating your CV score artificially? Try early stopping on something more consistent (auc) and see if there's still noise, but your CV and your LB are a bit closer? (Because CV is 'worse' by not picking the best 4% score) Or try without early stopping, just pick a number of rounds based on prior inspection and stop there.\n\nJust a thought, others might have better insight :)",
    "1854575": "I don't use early_stopping_rounds. First I fit train.cv and evaluate AUC because of AMEX is noisy. Than I used single fixed best numbers of trees  for every hold-out train set without CV during fitting process because I use dart.",
    "1854578": "Thank you. I understand it. But it's a competition where we do a lot of efforts to do result better, but finally we also need to be lucky. In real world we don't play with random seeds if we have enough data and label representation😊",
    "1854666": "I have also noticed a large variation in local CV overall score [here][1]. If you use the same folds with a different XGB random seed, the overall CV has standard deviation `0.0012`. So it makes sense that LB score would also have large variation similar to standard deviation `0.0012`.\n\n[1]: https://www.kaggle.com/competitions/amex-default-prediction/discussion/336145#1850687",
    "1854692": "Exactly. Though I will reiterate that **large** is a very relative term and applies only to this competition setting. In most real-world applications, variations on a third decimal place are normal and would be considered insignificant.\n\nIt is also worth adding that XGB/LGB hyper-parameters may exacerbate this variation. During a grid search, I see average fold variations in AmEx scores anywhere between `0.001` to `0.003`. Log-loss variations have narrower distribution (`0.001-0.002`). Variations in accuracy are almost always < `0.001`.",
    "1854770": "The competition need more luck",
    "1854849": "I would like to add that, if using 5-fold CV, each individual OOF score represents the score on 20% of the training data. The metric (specifically the *D* component) is so noisy that the competition organizers have had to provide just as much test data as training data just in order to obtain a three digit Public LB score (and a similar amount dedicated to calculating the Private LB score) hence each individual OOF score will have a large variance. \nOnly after averaging over the five folds will this make the training data roughly the same size as the amount of data used to calculate the Public LB score (51% of the test data), so the uncertainty in the LB score should be commensurate with the uncertainty in the CV score.",
    "1855770": "I am not sure what ensembling strategy is working right now my top model seems to be better than any ensemble I try. Have tried around 10 strategies 😄. Though with low scoring models it was working",
    "1861074": "I am taking a piece of your luck."
  },
  "source": "meta"
}