{
  "id": 583866,
  "title": "Why leaderboard score is worse than cross-validation",
  "url": "/competitions/waveform-inversion/discussion/583866",
  "author_name": "",
  "post_date": "2025-06-10T04:32:40.870207200Z",
  "votes": 7,
  "comment_count": 7,
  "views": 0,
  "content": "<p>So in this competition, leaderboard score (LB) is much closer to local cross-validation (CV) than we're used to. This is because there is no domain shift; the test set is taken from the same population as the train set. But then, why is CV higher than LB at all? For example, in <a href=\"https://www.kaggle.com/brendanartley\" target=\"_blank\">@brendanartley</a>'s famous model, the CV score is 24.4 while the LB score is 28.8.</p>\n<p>It may be because while the prior for each of the 10 families (FlatVelA, FlatVelB, etc.) is the same, the distribution of families is not. It seems the harder families appear somewhat more often. I base this on a reliable classifier that can identify some of the families with near-certainty on the test set. Here's the distribution as far as I can find it:</p>\n<table>\n<thead>\n<tr>\n<th>Family</th>\n<th>Ratio in training set</th>\n<th>Ratio in test set</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>FlatVelA+B</td>\n<td>20%</td>\n<td>18.1%</td>\n</tr>\n<tr>\n<td>StyleA</td>\n<td>10%</td>\n<td>8.6%</td>\n</tr>\n<tr>\n<td>StyleB</td>\n<td>10%</td>\n<td>9.6%</td>\n</tr>\n<tr>\n<td>All others</td>\n<td>60%</td>\n<td>63.6%</td>\n</tr>\n</tbody>\n</table>\n<p>So the easier families seem to occur less often. I'm not sure if this explains everything - just having 2% less FlatVel for example only accounts for an ~0.5 difference. But since all the \"A\" models are much easier then the \"B\" models, it could be that more such imbalances do explain the full CV-LB gap.</p>",
  "messages": [
    {
      "id": "3220861",
      "postDate": "06/10/2025 04:32:40",
      "content": "<p>So in this competition, leaderboard score (LB) is much closer to local cross-validation (CV) than we're used to. This is because there is no domain shift; the test set is taken from the same population as the train set. But then, why is CV higher than LB at all? For example, in <a href=\"https://www.kaggle.com/brendanartley\" target=\"_blank\">@brendanartley</a>'s famous model, the CV score is 24.4 while the LB score is 28.8.</p>\n<p>It may be because while the prior for each of the 10 families (FlatVelA, FlatVelB, etc.) is the same, the distribution of families is not. It seems the harder families appear somewhat more often. I base this on a reliable classifier that can identify some of the families with near-certainty on the test set. Here's the distribution as far as I can find it:</p>\n<table>\n<thead>\n<tr>\n<th>Family</th>\n<th>Ratio in training set</th>\n<th>Ratio in test set</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>FlatVelA+B</td>\n<td>20%</td>\n<td>18.1%</td>\n</tr>\n<tr>\n<td>StyleA</td>\n<td>10%</td>\n<td>8.6%</td>\n</tr>\n<tr>\n<td>StyleB</td>\n<td>10%</td>\n<td>9.6%</td>\n</tr>\n<tr>\n<td>All others</td>\n<td>60%</td>\n<td>63.6%</td>\n</tr>\n</tbody>\n</table>\n<p>So the easier families seem to occur less often. I'm not sure if this explains everything - just having 2% less FlatVel for example only accounts for an ~0.5 difference. But since all the \"A\" models are much easier then the \"B\" models, it could be that more such imbalances do explain the full CV-LB gap.</p>",
      "rawMarkdown": "So in this competition, leaderboard score (LB) is much closer to local cross-validation (CV) than we're used to. This is because there is no domain shift; the test set is taken from the same population as the train set. But then, why is CV higher than LB at all? For example, in @brendanartley's famous model, the CV score is 24.4 while the LB score is 28.8.\n\nIt may be because while the prior for each of the 10 families (FlatVelA, FlatVelB, etc.) is the same, the distribution of families is not. It seems the harder families appear somewhat more often. I base this on a reliable classifier that can identify some of the families with near-certainty on the test set. Here's the distribution as far as I can find it:\n\n| Family | Ratio in training set | Ratio in test set |\n| --- | --- |\n| FlatVelA+B | 20% | 18.1% |\n| StyleA | 10% | 8.6% | \n| StyleB | 10% | 9.6% |\n| All others | 60% | 63.6% |\n\nSo the easier families seem to occur less often. I'm not sure if this explains everything - just having 2% less FlatVel for example only accounts for an ~0.5 difference. But since all the \"A\" models are much easier then the \"B\" models, it could be that more such imbalances do explain the full CV-LB gap.",
      "votes": null
    },
    {
      "id": "3220877",
      "postDate": "06/10/2025 05:15:04",
      "content": "<p>Although unlikely in <a href=\"https://www.kaggle.com/brendanartley\" target=\"_blank\">@brendanartley</a>'s case, It might also be distribution of the validation set to some extent, I had a CV of 23.7 for my 25.3 LB and CV 21.8 for 23.1 LB</p>\n<p>The method families in the test set are distributed differently, but when I calculated the weighted average, that should not impact the score much (less than 1.0 if I can recall), in my validation the trend continues: the more my score gets better, the less the difference between CV and LB remains.</p>",
      "rawMarkdown": "Although unlikely in @brendanartley's case, It might also be distribution of the validation set to some extent, I had a CV of 23.7 for my 25.3 LB and CV 21.8 for 23.1 LB\n\nThe method families in the test set are distributed differently, but when I calculated the weighted average, that should not impact the score much (less than 1.0 if I can recall), in my validation the trend continues: the more my score gets better, the less the difference between CV and LB remains.",
      "votes": null
    },
    {
      "id": "3220927",
      "postDate": "06/10/2025 06:51:47",
      "content": "<p>The ratio is not equal. If you look at the full set, it's 30k for the four simple sets, 54k for the four fault sets and 67k for the last two style sets. Now calculate the weighted average according to this ratio…it will be closer.  </p>",
      "rawMarkdown": "The ratio is not equal. If you look at the full set, it's 30k for the four simple sets, 54k for the four fault sets and 67k for the last two style sets. Now calculate the weighted average according to this ratio...it will be closer.",
      "votes": null
    },
    {
      "id": "3221086",
      "postDate": "06/10/2025 11:33:12",
      "content": "<p>May I ask whether your CV (cross-validation) is calculated using a weighted average or a simple average?</p>",
      "rawMarkdown": "May I ask whether your CV (cross-validation) is calculated using a weighted average or a simple average?",
      "votes": null
    },
    {
      "id": "3221088",
      "postDate": "06/10/2025 11:34:28",
      "content": "<p>I calculate the CV between every sample, that would nearly mimic weighted average, I don't average the CV per method</p>",
      "rawMarkdown": "I calculate the CV between every sample, that would nearly mimic weighted average, I don't average the CV per method",
      "votes": null
    },
    {
      "id": "3221089",
      "postDate": "06/10/2025 11:35:23",
      "content": "<p>ok,thank you</p>",
      "rawMarkdown": "ok,thank you",
      "votes": null
    },
    {
      "id": "3221211",
      "postDate": "06/10/2025 15:20:57",
      "content": "<p><a href=\"https://www.kaggle.com/jeroencottaar\" target=\"_blank\">@jeroencottaar</a> </p>\n<p>Ratios in train set as follow right? </p>\n<table>\n<thead>\n<tr>\n<th>dataset</th>\n<th>files</th>\n<th>sampels</th>\n<th>share</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>CurveVel_A</td>\n<td>60</td>\n<td>30000</td>\n<td>6.38298</td>\n</tr>\n<tr>\n<td>CurveVel_B</td>\n<td>60</td>\n<td>30000</td>\n<td>6.38298</td>\n</tr>\n<tr>\n<td>FlatVel_A</td>\n<td>60</td>\n<td>30000</td>\n<td>6.38298</td>\n</tr>\n<tr>\n<td>FlatVel_B</td>\n<td>60</td>\n<td>30000</td>\n<td>6.38298</td>\n</tr>\n<tr>\n<td>CurveFault_A</td>\n<td>108</td>\n<td>54000</td>\n<td>11.4894</td>\n</tr>\n<tr>\n<td>CurveFault_B</td>\n<td>108</td>\n<td>54000</td>\n<td>11.4894</td>\n</tr>\n<tr>\n<td>FlatFault_A</td>\n<td>108</td>\n<td>54000</td>\n<td>11.4894</td>\n</tr>\n<tr>\n<td>FlatFault_B</td>\n<td>108</td>\n<td>54000</td>\n<td>11.4894</td>\n</tr>\n<tr>\n<td>Style_A</td>\n<td>134</td>\n<td>67000</td>\n<td>14.2553</td>\n</tr>\n<tr>\n<td>Style_B</td>\n<td>134</td>\n<td>67000</td>\n<td>14.2553</td>\n</tr>\n</tbody>\n</table>",
      "rawMarkdown": "jeroencottaar \n\nRatios in train set as follow right? \n| dataset      |   files |   sampels |    share |\n|:-------------|--------:|----------:|---------:|\n| CurveVel_A   |      60 |     30000 |  6.38298 |\n| CurveVel_B   |      60 |     30000 |  6.38298 |\n| FlatVel_A    |      60 |     30000 |  6.38298 |\n| FlatVel_B    |      60 |     30000 |  6.38298 |\n| CurveFault_A |     108 |     54000 | 11.4894  |\n| CurveFault_B |     108 |     54000 | 11.4894  |\n| FlatFault_A  |     108 |     54000 | 11.4894  |\n| FlatFault_B  |     108 |     54000 | 11.4894  |\n| Style_A      |     134 |     67000 | 14.2553  |\n| Style_B      |     134 |     67000 | 14.2553  |",
      "votes": null
    },
    {
      "id": "3221240",
      "postDate": "06/10/2025 16:03:15",
      "content": "<p>I was only considering the train set attached to the competition, not the full one. Those ratios seem even more different from the test set…</p>",
      "rawMarkdown": "I was only considering the train set attached to the competition, not the full one. Those ratios seem even more different from the test set...",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3220877,
      "author_name": "harshitsheoran",
      "author_url": "",
      "post_date": "06/10/2025 05:15:04",
      "content": "<p>Although unlikely in <a href=\"https://www.kaggle.com/brendanartley\" target=\"_blank\">@brendanartley</a>'s case, It might also be distribution of the validation set to some extent, I had a CV of 23.7 for my 25.3 LB and CV 21.8 for 23.1 LB</p>\n<p>The method families in the test set are distributed differently, but when I calculated the weighted average, that should not impact the score much (less than 1.0 if I can recall), in my validation the trend continues: the more my score gets better, the less the difference between CV and LB remains.</p>",
      "votes": null,
      "replies": [
        {
          "id": 3221086,
          "author_name": "guoooooooss",
          "author_url": "",
          "post_date": "06/10/2025 11:33:12",
          "content": "<p>May I ask whether your CV (cross-validation) is calculated using a weighted average or a simple average?</p>",
          "votes": null,
          "replies": [
            {
              "id": 3221088,
              "author_name": "harshitsheoran",
              "author_url": "",
              "post_date": "06/10/2025 11:34:28",
              "content": "<p>I calculate the CV between every sample, that would nearly mimic weighted average, I don't average the CV per method</p>",
              "votes": null,
              "replies": [
                {
                  "id": 3221089,
                  "author_name": "guoooooooss",
                  "author_url": "",
                  "post_date": "06/10/2025 11:35:23",
                  "content": "<p>ok,thank you</p>",
                  "votes": null,
                  "replies": []
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 3220927,
      "author_name": "shlomoron",
      "author_url": "",
      "post_date": "06/10/2025 06:51:47",
      "content": "<p>The ratio is not equal. If you look at the full set, it's 30k for the four simple sets, 54k for the four fault sets and 67k for the last two style sets. Now calculate the weighted average according to this ratio…it will be closer.  </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3221211,
      "author_name": "seshurajup",
      "author_url": "",
      "post_date": "06/10/2025 15:20:57",
      "content": "<p><a href=\"https://www.kaggle.com/jeroencottaar\" target=\"_blank\">@jeroencottaar</a> </p>\n<p>Ratios in train set as follow right? </p>\n<table>\n<thead>\n<tr>\n<th>dataset</th>\n<th>files</th>\n<th>sampels</th>\n<th>share</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>CurveVel_A</td>\n<td>60</td>\n<td>30000</td>\n<td>6.38298</td>\n</tr>\n<tr>\n<td>CurveVel_B</td>\n<td>60</td>\n<td>30000</td>\n<td>6.38298</td>\n</tr>\n<tr>\n<td>FlatVel_A</td>\n<td>60</td>\n<td>30000</td>\n<td>6.38298</td>\n</tr>\n<tr>\n<td>FlatVel_B</td>\n<td>60</td>\n<td>30000</td>\n<td>6.38298</td>\n</tr>\n<tr>\n<td>CurveFault_A</td>\n<td>108</td>\n<td>54000</td>\n<td>11.4894</td>\n</tr>\n<tr>\n<td>CurveFault_B</td>\n<td>108</td>\n<td>54000</td>\n<td>11.4894</td>\n</tr>\n<tr>\n<td>FlatFault_A</td>\n<td>108</td>\n<td>54000</td>\n<td>11.4894</td>\n</tr>\n<tr>\n<td>FlatFault_B</td>\n<td>108</td>\n<td>54000</td>\n<td>11.4894</td>\n</tr>\n<tr>\n<td>Style_A</td>\n<td>134</td>\n<td>67000</td>\n<td>14.2553</td>\n</tr>\n<tr>\n<td>Style_B</td>\n<td>134</td>\n<td>67000</td>\n<td>14.2553</td>\n</tr>\n</tbody>\n</table>",
      "votes": null,
      "replies": [
        {
          "id": 3221240,
          "author_name": "jeroencottaar",
          "author_url": "",
          "post_date": "06/10/2025 16:03:15",
          "content": "<p>I was only considering the train set attached to the competition, not the full one. Those ratios seem even more different from the test set…</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3220861": "So in this competition, leaderboard score (LB) is much closer to local cross-validation (CV) than we're used to. This is because there is no domain shift; the test set is taken from the same population as the train set. But then, why is CV higher than LB at all? For example, in @brendanartley's famous model, the CV score is 24.4 while the LB score is 28.8.\n\nIt may be because while the prior for each of the 10 families (FlatVelA, FlatVelB, etc.) is the same, the distribution of families is not. It seems the harder families appear somewhat more often. I base this on a reliable classifier that can identify some of the families with near-certainty on the test set. Here's the distribution as far as I can find it:\n\n| Family | Ratio in training set | Ratio in test set |\n| --- | --- |\n| FlatVelA+B | 20% | 18.1% |\n| StyleA | 10% | 8.6% | \n| StyleB | 10% | 9.6% |\n| All others | 60% | 63.6% |\n\nSo the easier families seem to occur less often. I'm not sure if this explains everything - just having 2% less FlatVel for example only accounts for an ~0.5 difference. But since all the \"A\" models are much easier then the \"B\" models, it could be that more such imbalances do explain the full CV-LB gap.",
    "3220877": "Although unlikely in @brendanartley's case, It might also be distribution of the validation set to some extent, I had a CV of 23.7 for my 25.3 LB and CV 21.8 for 23.1 LB\n\nThe method families in the test set are distributed differently, but when I calculated the weighted average, that should not impact the score much (less than 1.0 if I can recall), in my validation the trend continues: the more my score gets better, the less the difference between CV and LB remains.",
    "3220927": "The ratio is not equal. If you look at the full set, it's 30k for the four simple sets, 54k for the four fault sets and 67k for the last two style sets. Now calculate the weighted average according to this ratio...it will be closer.",
    "3221086": "May I ask whether your CV (cross-validation) is calculated using a weighted average or a simple average?",
    "3221088": "I calculate the CV between every sample, that would nearly mimic weighted average, I don't average the CV per method",
    "3221089": "ok,thank you",
    "3221211": "jeroencottaar \n\nRatios in train set as follow right? \n| dataset      |   files |   sampels |    share |\n|:-------------|--------:|----------:|---------:|\n| CurveVel_A   |      60 |     30000 |  6.38298 |\n| CurveVel_B   |      60 |     30000 |  6.38298 |\n| FlatVel_A    |      60 |     30000 |  6.38298 |\n| FlatVel_B    |      60 |     30000 |  6.38298 |\n| CurveFault_A |     108 |     54000 | 11.4894  |\n| CurveFault_B |     108 |     54000 | 11.4894  |\n| FlatFault_A  |     108 |     54000 | 11.4894  |\n| FlatFault_B  |     108 |     54000 | 11.4894  |\n| Style_A      |     134 |     67000 | 14.2553  |\n| Style_B      |     134 |     67000 | 14.2553  |",
    "3221240": "I was only considering the train set attached to the competition, not the full one. Those ratios seem even more different from the test set..."
  },
  "source": "meta"
}