{
  "id": 508276,
  "title": "Why Does the LB Score Fluctuate Significantly Even When Training on the Same Fold?",
  "url": "/competitions/birdclef-2024/discussion/508276",
  "author_name": "Tanuki_boosting",
  "post_date": "2024-05-29T01:02:36.171000",
  "votes": 1,
  "comment_count": 10,
  "views": 0,
  "content": "<p>As you all know, there is a significant difference between the training data and the test data. However, even when training on data from the same fold (i.e., a fold with exactly the same training and validation data), the LB score varies depending on small differences like the number of epochs or learning rates, or even when training under the exact same conditions.</p>\n<p>In other words, the variance in LB scores is so large that it's difficult to trust the public leaderboard scores.</p>\n<p>Why is the variance in LB scores so large?<br>\nAdditionally, while shake might be a possibility, what can we do to mitigate the risk of such shake?</p>",
  "messages": [
    {
      "id": 2842209,
      "postDate": "2024-05-29T01:02:36.170Z",
      "content": "<p>As you all know, there is a significant difference between the training data and the test data. However, even when training on data from the same fold (i.e., a fold with exactly the same training and validation data), the LB score varies depending on small differences like the number of epochs or learning rates, or even when training under the exact same conditions.</p>\n<p>In other words, the variance in LB scores is so large that it's difficult to trust the public leaderboard scores.</p>\n<p>Why is the variance in LB scores so large?<br>\nAdditionally, while shake might be a possibility, what can we do to mitigate the risk of such shake?</p>",
      "rawMarkdown": "As you all know, there is a significant difference between the training data and the test data. However, even when training on data from the same fold (i.e., a fold with exactly the same training and validation data), the LB score varies depending on small differences like the number of epochs or learning rates, or even when training under the exact same conditions.\n\nIn other words, the variance in LB scores is so large that it's difficult to trust the public leaderboard scores.\n\nWhy is the variance in LB scores so large?\nAdditionally, while shake might be a possibility, what can we do to mitigate the risk of such shake?",
      "votes": 1
    },
    {
      "id": 2842272,
      "postDate": "2024-05-29T02:24:13.020Z",
      "content": "<blockquote>\n  <p>small differences like the number of epochs or learning rates</p>\n</blockquote>\n<p>These two can be make or break for training, I wouldn't call it small difference </p>\n<p>I hope I'm not wrong, I recall the past 4 years there has been no shake for bird competition.  Usually it's pretty safe fitting to LB </p>",
      "rawMarkdown": "> small differences like the number of epochs or learning rates\n\nThese two can be make or break for training, I wouldn't call it small difference \n\nI hope I'm not wrong, I recall the past 4 years there has been no shake for bird competition.  Usually it's pretty safe fitting to LB ",
      "votes": 2,
      "replies": [
        {
          "id": 2842293,
          "postDate": "2024-05-29T03:03:31.553Z",
          "content": "<p>In this year's test set, only some partial classes are included. It is not clear whether the classes in public and private are consistent. So if they are not consistent, i think big shake may occur.</p>",
          "rawMarkdown": "In this year's test set, only some partial classes are included. It is not clear whether the classes in public and private are consistent. So if they are not consistent, i think big shake may occur.",
          "votes": 1
        },
        {
          "id": 2842430,
          "postDate": "2024-05-29T04:56:39.667Z",
          "content": "<p>Indeed, you are right that learning rates and epochs are significant differences. I apologize for my oversight. The issue I am concerned about is the extremely high variance in LB scores.</p>\n<p>I believe that even when training under identical conditions, merely changing the random seed results in different LB scores.</p>\n<p>More specifically, I estimate that the LB scores of inference results from models with the same architecture and conditions, differing only in random seed, follow a Gaussian distribution with certain mean and variance parameters. I am asserting that this variance is very large. This means that when this variance is large, it becomes difficult to determine whether the deviations in LB scores caused by different preprocessing methods are statistically significant.</p>",
          "rawMarkdown": "Indeed, you are right that learning rates and epochs are significant differences. I apologize for my oversight. The issue I am concerned about is the extremely high variance in LB scores.\n\nI believe that even when training under identical conditions, merely changing the random seed results in different LB scores.\n\nMore specifically, I estimate that the LB scores of inference results from models with the same architecture and conditions, differing only in random seed, follow a Gaussian distribution with certain mean and variance parameters. I am asserting that this variance is very large. This means that when this variance is large, it becomes difficult to determine whether the deviations in LB scores caused by different preprocessing methods are statistically significant.",
          "votes": 1,
          "replies": [
            {
              "id": 2844065,
              "postDate": "2024-05-29T21:41:30.517Z",
              "content": "<p>I am seeing a similar situation. For all my top models I trained multiple times with different seed. They show an LB difference of about 0.03. This should partly due to the large domain shift -- but given the stable lb/private score relation, we should somewhat trust the public LB? </p>\n<p>The true problem for this, like you said, is it significantly slows down my experiments. For experiments that did well I need to make sure it's not just randomness, and for those that are bad, if they are not too bad, I need to try multiple seeds to confirm. Now I am regretting not using the submission chances in earlier stages of this competition… But at least you can evaluate an experiment based on multiple LB scores with different seeds. Epoch length should be hard to tune because of the weak relation between cv and lb.</p>",
              "rawMarkdown": "I am seeing a similar situation. For all my top models I trained multiple times with different seed. They show an LB difference of about 0.03. This should partly due to the large domain shift -- but given the stable lb/private score relation, we should somewhat trust the public LB? \n\nThe true problem for this, like you said, is it significantly slows down my experiments. For experiments that did well I need to make sure it's not just randomness, and for those that are bad, if they are not too bad, I need to try multiple seeds to confirm. Now I am regretting not using the submission chances in earlier stages of this competition... But at least you can evaluate an experiment based on multiple LB scores with different seeds. Epoch length should be hard to tune because of the weak relation between cv and lb.",
              "votes": 3
            },
            {
              "id": 2846176,
              "postDate": "2024-05-31T02:00:30.903Z",
              "content": "<p>Thank you very much for your comments. I truly believe you are absolutely right.<br>\nTherefore, I am considering what measures would best diversify the risks as part of the second submission strategy.</p>",
              "rawMarkdown": "Thank you very much for your comments. I truly believe you are absolutely right.\nTherefore, I am considering what measures would best diversify the risks as part of the second submission strategy."
            },
            {
              "id": 2846233,
              "postDate": "2024-05-31T03:18:46.293Z",
              "content": "<p>I would like to point out that last year number one is at position 86 holding LB 0.66.  I want to correct myself that even the last few years there is no shake, a grandmaster at that position usually implies massive shake :(   So yeah, now I'm beginning to think perhaps fitting LB is a bad idea </p>",
              "rawMarkdown": "I would like to point out that last year number one is at position 86 holding LB 0.66.  I want to correct myself that even the last few years there is no shake, a grandmaster at that position usually implies massive shake :(   So yeah, now I'm beginning to think perhaps fitting LB is a bad idea "
            },
            {
              "id": 2846249,
              "postDate": "2024-05-31T03:35:36.010Z",
              "content": "<p>I didn't notice that. Thank you for your insight! Another reason that may contribute to the shake is the different metric used this year, which significantly increased the LB/CV gap. But the problem for me is I have not built a reliable CV. I just learned yesterday from a podcast from Dieter that one possible way is to seclude one region of audio as the validation set to simulate the domain shift. But the remaining time and submission chances does not allow me to try this method. </p>\n<p>Therefore my current strategy is to choose 1) my best LB ensemble and 2) my most robustly built model ensemble regardless of LB. </p>\n<p>Regarding last year's first place's position, I am somewhat surpised because, even though this year lb have a lot of randomness, 200+ of submissions should overcome that even just for one submission? I am not sure -- but in the past the skake among top places is mostly not significant enough to change your medal(&lt;20 places).</p>",
              "rawMarkdown": "I didn't notice that. Thank you for your insight! Another reason that may contribute to the shake is the different metric used this year, which significantly increased the LB/CV gap. But the problem for me is I have not built a reliable CV. I just learned yesterday from a podcast from Dieter that one possible way is to seclude one region of audio as the validation set to simulate the domain shift. But the remaining time and submission chances does not allow me to try this method. \n\nTherefore my current strategy is to choose 1) my best LB ensemble and 2) my most robustly built model ensemble regardless of LB. \n\nRegarding last year's first place's position, I am somewhat surpised because, even though this year lb have a lot of randomness, 200+ of submissions should overcome that even just for one submission? I am not sure -- but in the past the skake among top places is mostly not significant enough to change your medal(<20 places)."
            },
            {
              "id": 2846665,
              "postDate": "2024-05-31T07:35:21.790Z",
              "rawMarkdown": "",
              "isDeleted": true
            },
            {
              "id": 2850157,
              "postDate": "2024-06-02T01:59:37.523Z",
              "content": "<p>Hii, still looking for teammate？</p>",
              "rawMarkdown": "Hii, still looking for teammate？"
            }
          ]
        }
      ]
    },
    {
      "id": 2848203,
      "postDate": "2024-05-31T21:58:07.290Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 2842272,
      "author_name": "yukiya",
      "author_url": "",
      "post_date": "2024-05-29T02:24:13.020000",
      "content": "<blockquote>\n  <p>small differences like the number of epochs or learning rates</p>\n</blockquote>\n<p>These two can be make or break for training, I wouldn't call it small difference </p>\n<p>I hope I'm not wrong, I recall the past 4 years there has been no shake for bird competition.  Usually it's pretty safe fitting to LB </p>",
      "votes": 2,
      "replies": [
        {
          "id": 2842293,
          "author_name": "tanxxx",
          "author_url": "",
          "post_date": "2024-05-29T03:03:31.553000",
          "content": "<p>In this year's test set, only some partial classes are included. It is not clear whether the classes in public and private are consistent. So if they are not consistent, i think big shake may occur.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 2842430,
          "author_name": "Tanuki_boosting",
          "author_url": "",
          "post_date": "2024-05-29T04:56:39.667000",
          "content": "<p>Indeed, you are right that learning rates and epochs are significant differences. I apologize for my oversight. The issue I am concerned about is the extremely high variance in LB scores.</p>\n<p>I believe that even when training under identical conditions, merely changing the random seed results in different LB scores.</p>\n<p>More specifically, I estimate that the LB scores of inference results from models with the same architecture and conditions, differing only in random seed, follow a Gaussian distribution with certain mean and variance parameters. I am asserting that this variance is very large. This means that when this variance is large, it becomes difficult to determine whether the deviations in LB scores caused by different preprocessing methods are statistically significant.</p>",
          "votes": 1,
          "replies": [
            {
              "id": 2844065,
              "author_name": "LLLEEEOOOH",
              "author_url": "",
              "post_date": "2024-05-29T21:41:30.517000",
              "content": "<p>I am seeing a similar situation. For all my top models I trained multiple times with different seed. They show an LB difference of about 0.03. This should partly due to the large domain shift -- but given the stable lb/private score relation, we should somewhat trust the public LB? </p>\n<p>The true problem for this, like you said, is it significantly slows down my experiments. For experiments that did well I need to make sure it's not just randomness, and for those that are bad, if they are not too bad, I need to try multiple seeds to confirm. Now I am regretting not using the submission chances in earlier stages of this competition… But at least you can evaluate an experiment based on multiple LB scores with different seeds. Epoch length should be hard to tune because of the weak relation between cv and lb.</p>",
              "votes": 3,
              "replies": []
            },
            {
              "id": 2846176,
              "author_name": "Tanuki_boosting",
              "author_url": "",
              "post_date": "2024-05-31T02:00:30.903000",
              "content": "<p>Thank you very much for your comments. I truly believe you are absolutely right.<br>\nTherefore, I am considering what measures would best diversify the risks as part of the second submission strategy.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2846233,
              "author_name": "yukiya",
              "author_url": "",
              "post_date": "2024-05-31T03:18:46.293000",
              "content": "<p>I would like to point out that last year number one is at position 86 holding LB 0.66.  I want to correct myself that even the last few years there is no shake, a grandmaster at that position usually implies massive shake :(   So yeah, now I'm beginning to think perhaps fitting LB is a bad idea </p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2846249,
              "author_name": "LLLEEEOOOH",
              "author_url": "",
              "post_date": "2024-05-31T03:35:36.010000",
              "content": "<p>I didn't notice that. Thank you for your insight! Another reason that may contribute to the shake is the different metric used this year, which significantly increased the LB/CV gap. But the problem for me is I have not built a reliable CV. I just learned yesterday from a podcast from Dieter that one possible way is to seclude one region of audio as the validation set to simulate the domain shift. But the remaining time and submission chances does not allow me to try this method. </p>\n<p>Therefore my current strategy is to choose 1) my best LB ensemble and 2) my most robustly built model ensemble regardless of LB. </p>\n<p>Regarding last year's first place's position, I am somewhat surpised because, even though this year lb have a lot of randomness, 200+ of submissions should overcome that even just for one submission? I am not sure -- but in the past the skake among top places is mostly not significant enough to change your medal(&lt;20 places).</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2846665,
              "author_name": "",
              "author_url": "",
              "post_date": "2024-05-31T07:35:21.790000",
              "content": "",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2850157,
              "author_name": "SwiftSquirrel",
              "author_url": "",
              "post_date": "2024-06-02T01:59:37.523000",
              "content": "<p>Hii, still looking for teammate？</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2848203,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-05-31T21:58:07.290000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2842209": "As you all know, there is a significant difference between the training data and the test data. However, even when training on data from the same fold (i.e., a fold with exactly the same training and validation data), the LB score varies depending on small differences like the number of epochs or learning rates, or even when training under the exact same conditions.\n\nIn other words, the variance in LB scores is so large that it's difficult to trust the public leaderboard scores.\n\nWhy is the variance in LB scores so large?\nAdditionally, while shake might be a possibility, what can we do to mitigate the risk of such shake?",
    "2842272": "> small differences like the number of epochs or learning rates\n\nThese two can be make or break for training, I wouldn't call it small difference \n\nI hope I'm not wrong, I recall the past 4 years there has been no shake for bird competition.  Usually it's pretty safe fitting to LB ",
    "2848203": ""
  }
}