{
  "id": 554815,
  "title": "Train vs test set difference (statistical analysis)",
  "url": "/competitions/czii-cryo-et-object-identification/discussion/554815",
  "author_name": "",
  "post_date": "2025-01-03T13:39:59.823232100Z",
  "votes": 26,
  "comment_count": 20,
  "views": 0,
  "content": "<p>I've observed quite a gap between my performance in cross-validation vs. submission, so dove a bit deeper into this. </p>\n<p>I consider a simple template finding model. Both training and inference are fully reproducible. I perform 7-fold cross-validation: the model is evaluated on 1 training dataset after being trained on the 6 others. Here are the results per dataset:</p>\n<table>\n<thead>\n<tr>\n<th>Dataset</th>\n<th>Score</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>TS_5_4</td>\n<td>0.702</td>\n</tr>\n<tr>\n<td>TS_69_2</td>\n<td>0.729</td>\n</tr>\n<tr>\n<td>TS_6_4</td>\n<td>0.540</td>\n</tr>\n<tr>\n<td>TS_6_6</td>\n<td>0.754</td>\n</tr>\n<tr>\n<td>TS_73_6</td>\n<td>0.684</td>\n</tr>\n<tr>\n<td>TS_86_3</td>\n<td>0.720</td>\n</tr>\n<tr>\n<td>TS_99_9</td>\n<td>0.695</td>\n</tr>\n</tbody>\n</table>\n<p>So if the test set is similar to the training set, I'd expect a score of 0.689+/- 0.024 (stepping over the fact that the score is not actually an average per dataset, but I don't think it matters too much here).</p>\n<p>However, the actual score is 0.598. This indicates a very significant difference (p&lt;&lt;0.001). This seems to indicate some serious difference in distribution between the training and test set. However, as I understand it the training set was specifically selected to be similar to the test set. (EDIT: as <a href=\"https://www.kaggle.com/davidlist\" target=\"_blank\">@davidlist</a> pointed out below, my analysis was wrong. The correct p-value is 0.014, still somewhat significant but nowhere near as wild. In addition the assumption of a Gaussian distribution seems rather inappropriate, so the whole analysis is suspect).</p>\n<p>The model in question scales all data such that the median and standard deviation are equal. This doesn't actually help performance, but does ensure in this case that the distribution difference is not related to scale or shift.</p>\n<p>The model has few hyperparameters; they are not overtuned on the training set.</p>\n<p>Anyone have any idea what's going on?</p>",
  "messages": [
    {
      "id": "3087431",
      "postDate": "01/03/2025 13:39:59",
      "content": "<p>I've observed quite a gap between my performance in cross-validation vs. submission, so dove a bit deeper into this. </p>\n<p>I consider a simple template finding model. Both training and inference are fully reproducible. I perform 7-fold cross-validation: the model is evaluated on 1 training dataset after being trained on the 6 others. Here are the results per dataset:</p>\n<table>\n<thead>\n<tr>\n<th>Dataset</th>\n<th>Score</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>TS_5_4</td>\n<td>0.702</td>\n</tr>\n<tr>\n<td>TS_69_2</td>\n<td>0.729</td>\n</tr>\n<tr>\n<td>TS_6_4</td>\n<td>0.540</td>\n</tr>\n<tr>\n<td>TS_6_6</td>\n<td>0.754</td>\n</tr>\n<tr>\n<td>TS_73_6</td>\n<td>0.684</td>\n</tr>\n<tr>\n<td>TS_86_3</td>\n<td>0.720</td>\n</tr>\n<tr>\n<td>TS_99_9</td>\n<td>0.695</td>\n</tr>\n</tbody>\n</table>\n<p>So if the test set is similar to the training set, I'd expect a score of 0.689+/- 0.024 (stepping over the fact that the score is not actually an average per dataset, but I don't think it matters too much here).</p>\n<p>However, the actual score is 0.598. This indicates a very significant difference (p&lt;&lt;0.001). This seems to indicate some serious difference in distribution between the training and test set. However, as I understand it the training set was specifically selected to be similar to the test set. (EDIT: as <a href=\"https://www.kaggle.com/davidlist\" target=\"_blank\">@davidlist</a> pointed out below, my analysis was wrong. The correct p-value is 0.014, still somewhat significant but nowhere near as wild. In addition the assumption of a Gaussian distribution seems rather inappropriate, so the whole analysis is suspect).</p>\n<p>The model in question scales all data such that the median and standard deviation are equal. This doesn't actually help performance, but does ensure in this case that the distribution difference is not related to scale or shift.</p>\n<p>The model has few hyperparameters; they are not overtuned on the training set.</p>\n<p>Anyone have any idea what's going on?</p>",
      "rawMarkdown": "I've observed quite a gap between my performance in cross-validation vs. submission, so dove a bit deeper into this. \n\nI consider a simple template finding model. Both training and inference are fully reproducible. I perform 7-fold cross-validation: the model is evaluated on 1 training dataset after being trained on the 6 others. Here are the results per dataset:\n\n| Dataset | Score | \n| --- | \n| TS_5_4 | 0.702 | \n| TS_69_2 | 0.729 | \n| TS_6_4 | 0.540 | \n| TS_6_6 | 0.754 | \n| TS_73_6 | 0.684 | \n| TS_86_3 | 0.720 | \n| TS_99_9 | 0.695 | \n\n\nSo if the test set is similar to the training set, I'd expect a score of 0.689+/- 0.024 (stepping over the fact that the score is not actually an average per dataset, but I don't think it matters too much here).\n\nHowever, the actual score is 0.598. This indicates a very significant difference (p<<0.001). This seems to indicate some serious difference in distribution between the training and test set. However, as I understand it the training set was specifically selected to be similar to the test set. (EDIT: as @davidlist pointed out below, my analysis was wrong. The correct p-value is 0.014, still somewhat significant but nowhere near as wild. In addition the assumption of a Gaussian distribution seems rather inappropriate, so the whole analysis is suspect).\n\nThe model in question scales all data such that the median and standard deviation are equal. This doesn't actually help performance, but does ensure in this case that the distribution difference is not related to scale or shift.\n\nThe model has few hyperparameters; they are not overtuned on the training set.\n\nAnyone have any idea what's going on?",
      "votes": null
    },
    {
      "id": "3087440",
      "postDate": "01/03/2025 13:46:21",
      "content": "<p>Statistics are my Achilles heel. I'm not going deeper than basic analysis and interpretation but as I see it. The main problem is 7 against 500. Is hard to represent properly a 500 population distribution with only 7 samples. Even more, you can only monitorize public. So I would focus into make a very general model without payng too much attention to public.</p>",
      "rawMarkdown": "Statistics are my Achilles heel. I'm not going deeper than basic analysis and interpretation but as I see it. The main problem is 7 against 500. Is hard to represent properly a 500 population distribution with only 7 samples. Even more, you can only monitorize public. So I would focus into make a very general model without payng too much attention to public.",
      "votes": null
    },
    {
      "id": "3087706",
      "postDate": "01/03/2025 19:20:26",
      "content": "<blockquote>\n  <p>The model has few hyperparameters; they are not overtuned on the training set.</p>\n</blockquote>\n<p>Are you optimizing for f4 score?  At least for what I'm doing that seems to be the main source of overfitting.  If you are taking the best model based on validation score, that would be another area.</p>",
      "rawMarkdown": ">The model has few hyperparameters; they are not overtuned on the training set.\n\nAre you optimizing for f4 score?  At least for what I'm doing that seems to be the main source of overfitting.  If you are taking the best model based on validation score, that would be another area.",
      "votes": null
    },
    {
      "id": "3087712",
      "postDate": "01/03/2025 19:29:13",
      "content": "<p>Thanks for the input. I am indeed optimizing for f4 score. But that's mainly adapting the threshold, which does not lead to improvement on the public leaderboard…</p>",
      "rawMarkdown": "Thanks for the input. I am indeed optimizing for f4 score. But that's mainly adapting the threshold, which does not lead to improvement on the public leaderboard...",
      "votes": null
    },
    {
      "id": "3087756",
      "postDate": "01/03/2025 20:35:04",
      "content": "<p>thanks for sharing the insightful results. would you mind sharing which template matching model you used?</p>",
      "rawMarkdown": "thanks for sharing the insightful results. would you mind sharing which template matching model you used?",
      "votes": null
    },
    {
      "id": "3087763",
      "postDate": "01/03/2025 20:42:39",
      "content": "<p>It's something I cooked together myself; it basically takes the mean over all instances of a given particle in training, and then searches for the best convolutional match (with some additional tricks). Anyway it doesn't work particularly well :(…</p>",
      "rawMarkdown": "It's something I cooked together myself; it basically takes the mean over all instances of a given particle in training, and then searches for the best convolutional match (with some additional tricks). Anyway it doesn't work particularly well :(...",
      "votes": null
    },
    {
      "id": "3087769",
      "postDate": "01/03/2025 21:01:18",
      "content": "<p>Actually, I think a lot of my f4 optimization is related to connected component analysis which it doesn't sound like you're doing.  But I will say, there does appear to be an unusual amount of variance at pretty much every step in this process.  Not sure how much of that explains your results, though.</p>",
      "rawMarkdown": "Actually, I think a lot of my f4 optimization is related to connected component analysis which it doesn't sound like you're doing.  But I will say, there does appear to be an unusual amount of variance at pretty much every step in this process.  Not sure how much of that explains your results, though.",
      "votes": null
    },
    {
      "id": "3087771",
      "postDate": "01/03/2025 21:04:43",
      "content": "<p>I would point out that your LB score of 0.598 is actually better than your 0.540 result from TS_6_4.  Not doing a particularly rigorous statistical analysis here, but wondering if p&lt;&lt;0.001 is reasonable given that fact.</p>",
      "rawMarkdown": "I would point out that your LB score of 0.598 is actually better than your 0.540 result from TS_6_4.  Not doing a particularly rigorous statistical analysis here, but wondering if p<<0.001 is reasonable given that fact.",
      "votes": null
    },
    {
      "id": "3087777",
      "postDate": "01/03/2025 21:35:39",
      "content": "<p>…Okay, got a little curious.  I get a p-value of 0.014 which is still quite significant.  …Although, that comes with some assumptions that may or may not be valid.</p>",
      "rawMarkdown": "...Okay, got a little curious.  I get a p-value of 0.014 which is still quite significant.  ...Although, that comes with some assumptions that may or may not be valid.",
      "votes": null
    },
    {
      "id": "3087783",
      "postDate": "01/03/2025 21:57:24",
      "content": "<p>thanks for sharing!</p>",
      "rawMarkdown": "thanks for sharing!",
      "votes": null
    },
    {
      "id": "3087784",
      "postDate": "01/03/2025 21:58:57",
      "content": "<p>Hm just rechecked and I got 0.00014, not quite \"&lt;&lt;0.001\" but still significant enough. Difference in our calculation is a factor 100, maybe a percentage somewher?</p>\n<p>Anyway, the analysis is only valid if the distribution is Gaussian. It's a fair point that the TS_6_4 result already shows that it is not. A more reasonable explanation might then be that the full set has many more datasets similar to TS_6_4, and we just happened to unluckily get only a few in the training set. That's still a bit unlikely, but definitely plausible. </p>\n<p>Perhaps figuring out why TS_6_4 performs so poorly could be a clue to figuring out the poor LB performance. Thanks for your input!</p>",
      "rawMarkdown": "Hm just rechecked and I got 0.00014, not quite \"<<0.001\" but still significant enough. Difference in our calculation is a factor 100, maybe a percentage somewher?\n\nAnyway, the analysis is only valid if the distribution is Gaussian. It's a fair point that the TS_6_4 result already shows that it is not. A more reasonable explanation might then be that the full set has many more datasets similar to TS_6_4, and we just happened to unluckily get only a few in the training set. That's still a bit unlikely, but definitely plausible. \n\nPerhaps figuring out why TS_6_4 performs so poorly could be a clue to figuring out the poor LB performance. Thanks for your input!",
      "votes": null
    },
    {
      "id": "3087802",
      "postDate": "01/03/2025 22:42:43",
      "content": "<p>I may very well be setting up the problem incorrectly.  I did a t-distribution from your averages and got a test statistic of -3.42…I also got lazy and had perplexity.ai generate the code for me, but it's not obvious to me what's wrong if it is:</p>\n<pre><code> numpy as np\n scipy import stats\n\n = np.array([., ., ., ., ., ., .])  # Your  data points\n = .  # The point you want to test\n\n = np.mean(sample)\n = np.std(sample, ddof=)\n = (test_point - sample_mean) / (sample_std / np.sqrt(len(sample)))\n\n = len(sample) - \n = stats.t.sf(np.abs(t_statistic), degrees_of_freedom) * \n\n(f)\n(f)\n</code></pre>\n<p>There was another post in which they said samples were divided into 16 bins based on Kolmogorov Smirnov.  That would have obviously been reduced further from 16 down to 7 so maybe that explains some of it.</p>",
      "rawMarkdown": "I may very well be setting up the problem incorrectly.  I did a t-distribution from your averages and got a test statistic of -3.42...I also got lazy and had perplexity.ai generate the code for me, but it's not obvious to me what's wrong if it is:\n\n```\nimport numpy as np\nfrom scipy import stats\n\nsample = np.array([0.702, 0.709, 0.540, 0.754, 0.684, 0.720, 0.695])  # Your 7 data points\ntest_point = 0.598  # The point you want to test\n\nsample_mean = np.mean(sample)\nsample_std = np.std(sample, ddof=1)\nt_statistic = (test_point - sample_mean) / (sample_std / np.sqrt(len(sample)))\n\ndegrees_of_freedom = len(sample) - 1\np_value = stats.t.sf(np.abs(t_statistic), degrees_of_freedom) * 2\n\nprint(f\"T-statistic: {t_statistic}\")\nprint(f\"P-value: {p_value}\")\n```\n\nThere was another post in which they said samples were divided into 16 bins based on Kolmogorov Smirnov.  That would have obviously been reduced further from 16 down to 7 so maybe that explains some of it.",
      "votes": null
    },
    {
      "id": "3087832",
      "postDate": "01/03/2025 23:59:10",
      "content": "<p>Have you done by class submissions on the LB? </p>\n<p><code>thyroglobulin</code> and <code>beta-galactosidase</code> are dragging down my LB score more than expected.</p>\n<table>\n<thead>\n<tr>\n<th>Class</th>\n<th>CV</th>\n<th>LB</th>\n<th>Diff</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>apo-ferritin</td>\n<td>0.8741</td>\n<td>0.882</td>\n<td>0.0079</td>\n</tr>\n<tr>\n<td>beta-galactosidase</td>\n<td>0.7012</td>\n<td>0.5845</td>\n<td><strong>-0.1167</strong></td>\n</tr>\n<tr>\n<td>ribosome</td>\n<td>0.8131</td>\n<td>0.798</td>\n<td>-0.0151</td>\n</tr>\n<tr>\n<td>thyroglobulin</td>\n<td>0.6948</td>\n<td>0.5950</td>\n<td><strong>-0.0998</strong></td>\n</tr>\n<tr>\n<td>virus-like-particle</td>\n<td>0.9101</td>\n<td>0.896</td>\n<td>-0.0141</td>\n</tr>\n<tr>\n<td>Overall f-4</td>\n<td>0.7699</td>\n<td>0.7050</td>\n<td>-0.0649</td>\n</tr>\n</tbody>\n</table>",
      "rawMarkdown": "Have you done by class submissions on the LB? \n\n`thyroglobulin` and `beta-galactosidase` are dragging down my LB score more than expected.\n\n| Class          | CV | LB | Diff      |\n|----------------------|-----------------|-----------------|------------------|\n| apo-ferritin      | 0.8741          | 0.882           | 0.0079           |\n| beta-galactosidase | 0.7012          | 0.5845          | **-0.1167**          |\n| ribosome          | 0.8131          | 0.798           | -0.0151          |\n| thyroglobulin     | 0.6948          | 0.5950          | **-0.0998**          |\n| virus-like-particle | 0.9101         | 0.896           | -0.0141          |\n| Overall f-4          | 0.7699          | 0.7050          | -0.0649          |",
      "votes": null
    },
    {
      "id": "3087846",
      "postDate": "01/04/2025 00:21:19",
      "content": "<p>Thanks <a href=\"https://www.kaggle.com/brendanartley\" target=\"_blank\">@brendanartley</a>.  This info is helpful for a whole variety of reasons!</p>",
      "rawMarkdown": "Thanks @brendanartley.  This info is helpful for a whole variety of reasons!",
      "votes": null
    },
    {
      "id": "3087960",
      "postDate": "01/04/2025 06:25:36",
      "content": "<p>You are right, I was wrong. I just assumed I could estimate the sigma accurately from my 7 points and just check the normal distribution, which is wildly inappopriate here. I'll update the original post.</p>",
      "rawMarkdown": "You are right, I was wrong. I just assumed I could estimate the sigma accurately from my 7 points and just check the normal distribution, which is wildly inappopriate here. I'll update the original post.",
      "votes": null
    },
    {
      "id": "3087970",
      "postDate": "01/04/2025 06:52:15",
      "content": "<p>Wow that's generous!</p>\n<p>Here's mine. This is not the same version of the model as the original post, but I expect differences between CV and LB to be similar.</p>\n<table>\n<thead>\n<tr>\n<th>Class</th>\n<th>CV</th>\n<th>LB</th>\n<th>Diff</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>apo-ferritin</td>\n<td>0.898</td>\n<td>0.868</td>\n<td>-0.03</td>\n</tr>\n<tr>\n<td>beta-galactosidase</td>\n<td>0.533</td>\n<td>0.431</td>\n<td>-0.102</td>\n</tr>\n<tr>\n<td>ribosome</td>\n<td>0.855</td>\n<td>0.819</td>\n<td>-0.036</td>\n</tr>\n<tr>\n<td>thyroglobulin</td>\n<td>0.685</td>\n<td>0.501</td>\n<td>-0.184</td>\n</tr>\n<tr>\n<td>VLP</td>\n<td>0.927</td>\n<td>0.812</td>\n<td>-0.115</td>\n</tr>\n</tbody>\n</table>\n<p>Seems somewhat similar, though I suffer more in thyroglobulin and VLPs.</p>",
      "rawMarkdown": "Wow that's generous!\n\nHere's mine. This is not the same version of the model as the original post, but I expect differences between CV and LB to be similar.\n\n| Class | CV    | LB    | Diff   |\n| ----- | ----- | ----- | ------ |\n| apo-ferritin | 0.898 | 0.868 | -0.03  |\n| beta-galactosidase | 0.533 | 0.431 | -0.102 |\n| ribosome | 0.855 | 0.819 | -0.036 |\n| thyroglobulin | 0.685 | 0.501 | -0.184 |\n| VLP | 0.927 | 0.812 | -0.115 |\n\nSeems somewhat similar, though I suffer more in thyroglobulin and VLPs.",
      "votes": null
    },
    {
      "id": "3088009",
      "postDate": "01/04/2025 08:18:13",
      "content": "<p>So going back the Kolmogorov Smirnov test and dividing the full set of data into 16 bins…from which our 7 samples were drawn.  I'm thinking just because there were 16 bins that doesn't mean they're all equally represented.  For instance, only TS_6_4 has the elongated, non-spherical virus particles.  But does that mean 14% or 1/7th of the 500 samples have those same virus particles?  Seems unlikely.  Maybe the 7 samples are representative of the variety of the 500 though not necessarily in the same proportions.</p>",
      "rawMarkdown": "So going back the Kolmogorov Smirnov test and dividing the full set of data into 16 bins...from which our 7 samples were drawn.  I'm thinking just because there were 16 bins that doesn't mean they're all equally represented.  For instance, only TS_6_4 has the elongated, non-spherical virus particles.  But does that mean 14% or 1/7th of the 500 samples have those same virus particles?  Seems unlikely.  Maybe the 7 samples are representative of the variety of the 500 though not necessarily in the same proportions.",
      "votes": null
    },
    {
      "id": "3088230",
      "postDate": "01/04/2025 13:00:04",
      "content": "<p>Interesting, thanks for sharing. </p>\n<p>I wonder if there is something about hard classes on the LB that we don’t see in the 7 training samples. Could be something like shape, size, handedness, label discrepancy, etc?</p>",
      "rawMarkdown": "Interesting, thanks for sharing. \n\nI wonder if there is something about hard classes on the LB that we don’t see in the 7 training samples. Could be something like shape, size, handedness, label discrepancy, etc?",
      "votes": null
    },
    {
      "id": "3088612",
      "postDate": "01/04/2025 22:21:24",
      "content": "<p>Handedness for beta-galactosidase seem likely.  I know you already posted about that.  I saw some similar evidence while training with the Flipd transform in monai.  But I wonder if another issue may be that beta-galactosidase and thyroglobulin don't have spherical symmetry like the three easy ones so they require more examples.</p>",
      "rawMarkdown": "Handedness for beta-galactosidase seem likely.  I know you already posted about that.  I saw some similar evidence while training with the Flipd transform in monai.  But I wonder if another issue may be that beta-galactosidase and thyroglobulin don't have spherical symmetry like the three easy ones so they require more examples.",
      "votes": null
    },
    {
      "id": "3097049",
      "postDate": "01/15/2025 02:03:23",
      "content": "<p>Wow, this is amazing! May I know how do you get the class-specific LB score? </p>",
      "rawMarkdown": "Wow, this is amazing! May I know how do you get the class-specific LB score?",
      "votes": null
    },
    {
      "id": "3097079",
      "postDate": "01/15/2025 02:45:28",
      "content": "<p>Thanks <a href=\"https://www.kaggle.com/huiliu0\" target=\"_blank\">@huiliu0</a>, good question.</p>\n<p>First, you have to make a submission that only contains rows for the class you want to test. This makes the LB score for all other classes 0. </p>\n<p>Since we know the weights per class, we then can reverse engineer the class score by dividing by the sum of the class weights.</p>\n<p>Here is a python code snippet you can use.</p>\n<pre><code>\nweights = {\n    : ,\n    : ,\n    : ,\n    : ,\n    : ,\n}\n\n\nscores = {\n    : ,\n}\n\n\ntotal_weight = (weights.values())\n k  scores.keys():\n    s = (scores[k] * total_weight) / weights[k]\n    ()\n</code></pre>",
      "rawMarkdown": "Thanks @huiliu0, good question.\n\nFirst, you have to make a submission that only contains rows for the class you want to test. This makes the LB score for all other classes 0. \n\nSince we know the weights per class, we then can reverse engineer the class score by dividing by the sum of the class weights.\n\nHere is a python code snippet you can use.\n\n```\n# Weights for each class\nweights = {\n    'apo-ferritin': 1,\n    'beta-galactosidase': 2,\n    'ribosome': 1,\n    'thyroglobulin': 2,\n    'virus-like-particle': 1,\n}\n\n# LB score\nscores = {\n    'virus-like-particle': 0.133,\n}\n\n# Reverse engineer\ntotal_weight = sum(weights.values())\nfor k in scores.keys():\n    s = (scores[k] * total_weight) / weights[k]\n    print(f\"{k} Score: {s}\")\n```",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3087440,
      "author_name": "sacuscreed",
      "author_url": "",
      "post_date": "01/03/2025 13:46:21",
      "content": "<p>Statistics are my Achilles heel. I'm not going deeper than basic analysis and interpretation but as I see it. The main problem is 7 against 500. Is hard to represent properly a 500 population distribution with only 7 samples. Even more, you can only monitorize public. So I would focus into make a very general model without payng too much attention to public.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3087706,
      "author_name": "davidlist",
      "author_url": "",
      "post_date": "01/03/2025 19:20:26",
      "content": "<blockquote>\n  <p>The model has few hyperparameters; they are not overtuned on the training set.</p>\n</blockquote>\n<p>Are you optimizing for f4 score?  At least for what I'm doing that seems to be the main source of overfitting.  If you are taking the best model based on validation score, that would be another area.</p>",
      "votes": null,
      "replies": [
        {
          "id": 3087712,
          "author_name": "jeroencottaar",
          "author_url": "",
          "post_date": "01/03/2025 19:29:13",
          "content": "<p>Thanks for the input. I am indeed optimizing for f4 score. But that's mainly adapting the threshold, which does not lead to improvement on the public leaderboard…</p>",
          "votes": null,
          "replies": [
            {
              "id": 3087769,
              "author_name": "davidlist",
              "author_url": "",
              "post_date": "01/03/2025 21:01:18",
              "content": "<p>Actually, I think a lot of my f4 optimization is related to connected component analysis which it doesn't sound like you're doing.  But I will say, there does appear to be an unusual amount of variance at pretty much every step in this process.  Not sure how much of that explains your results, though.</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3087756,
      "author_name": "maedward",
      "author_url": "",
      "post_date": "01/03/2025 20:35:04",
      "content": "<p>thanks for sharing the insightful results. would you mind sharing which template matching model you used?</p>",
      "votes": null,
      "replies": [
        {
          "id": 3087763,
          "author_name": "jeroencottaar",
          "author_url": "",
          "post_date": "01/03/2025 20:42:39",
          "content": "<p>It's something I cooked together myself; it basically takes the mean over all instances of a given particle in training, and then searches for the best convolutional match (with some additional tricks). Anyway it doesn't work particularly well :(…</p>",
          "votes": null,
          "replies": [
            {
              "id": 3087783,
              "author_name": "maedward",
              "author_url": "",
              "post_date": "01/03/2025 21:57:24",
              "content": "<p>thanks for sharing!</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3087771,
      "author_name": "davidlist",
      "author_url": "",
      "post_date": "01/03/2025 21:04:43",
      "content": "<p>I would point out that your LB score of 0.598 is actually better than your 0.540 result from TS_6_4.  Not doing a particularly rigorous statistical analysis here, but wondering if p&lt;&lt;0.001 is reasonable given that fact.</p>",
      "votes": null,
      "replies": [
        {
          "id": 3087777,
          "author_name": "davidlist",
          "author_url": "",
          "post_date": "01/03/2025 21:35:39",
          "content": "<p>…Okay, got a little curious.  I get a p-value of 0.014 which is still quite significant.  …Although, that comes with some assumptions that may or may not be valid.</p>",
          "votes": null,
          "replies": [
            {
              "id": 3087784,
              "author_name": "jeroencottaar",
              "author_url": "",
              "post_date": "01/03/2025 21:58:57",
              "content": "<p>Hm just rechecked and I got 0.00014, not quite \"&lt;&lt;0.001\" but still significant enough. Difference in our calculation is a factor 100, maybe a percentage somewher?</p>\n<p>Anyway, the analysis is only valid if the distribution is Gaussian. It's a fair point that the TS_6_4 result already shows that it is not. A more reasonable explanation might then be that the full set has many more datasets similar to TS_6_4, and we just happened to unluckily get only a few in the training set. That's still a bit unlikely, but definitely plausible. </p>\n<p>Perhaps figuring out why TS_6_4 performs so poorly could be a clue to figuring out the poor LB performance. Thanks for your input!</p>",
              "votes": null,
              "replies": [
                {
                  "id": 3087802,
                  "author_name": "davidlist",
                  "author_url": "",
                  "post_date": "01/03/2025 22:42:43",
                  "content": "<p>I may very well be setting up the problem incorrectly.  I did a t-distribution from your averages and got a test statistic of -3.42…I also got lazy and had perplexity.ai generate the code for me, but it's not obvious to me what's wrong if it is:</p>\n<pre><code> numpy as np\n scipy import stats\n\n = np.array([., ., ., ., ., ., .])  # Your  data points\n = .  # The point you want to test\n\n = np.mean(sample)\n = np.std(sample, ddof=)\n = (test_point - sample_mean) / (sample_std / np.sqrt(len(sample)))\n\n = len(sample) - \n = stats.t.sf(np.abs(t_statistic), degrees_of_freedom) * \n\n(f)\n(f)\n</code></pre>\n<p>There was another post in which they said samples were divided into 16 bins based on Kolmogorov Smirnov.  That would have obviously been reduced further from 16 down to 7 so maybe that explains some of it.</p>",
                  "votes": null,
                  "replies": [
                    {
                      "id": 3087960,
                      "author_name": "jeroencottaar",
                      "author_url": "",
                      "post_date": "01/04/2025 06:25:36",
                      "content": "<p>You are right, I was wrong. I just assumed I could estimate the sigma accurately from my 7 points and just check the normal distribution, which is wildly inappopriate here. I'll update the original post.</p>",
                      "votes": null,
                      "replies": []
                    }
                  ]
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 3087832,
      "author_name": "brendanartley",
      "author_url": "",
      "post_date": "01/03/2025 23:59:10",
      "content": "<p>Have you done by class submissions on the LB? </p>\n<p><code>thyroglobulin</code> and <code>beta-galactosidase</code> are dragging down my LB score more than expected.</p>\n<table>\n<thead>\n<tr>\n<th>Class</th>\n<th>CV</th>\n<th>LB</th>\n<th>Diff</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>apo-ferritin</td>\n<td>0.8741</td>\n<td>0.882</td>\n<td>0.0079</td>\n</tr>\n<tr>\n<td>beta-galactosidase</td>\n<td>0.7012</td>\n<td>0.5845</td>\n<td><strong>-0.1167</strong></td>\n</tr>\n<tr>\n<td>ribosome</td>\n<td>0.8131</td>\n<td>0.798</td>\n<td>-0.0151</td>\n</tr>\n<tr>\n<td>thyroglobulin</td>\n<td>0.6948</td>\n<td>0.5950</td>\n<td><strong>-0.0998</strong></td>\n</tr>\n<tr>\n<td>virus-like-particle</td>\n<td>0.9101</td>\n<td>0.896</td>\n<td>-0.0141</td>\n</tr>\n<tr>\n<td>Overall f-4</td>\n<td>0.7699</td>\n<td>0.7050</td>\n<td>-0.0649</td>\n</tr>\n</tbody>\n</table>",
      "votes": null,
      "replies": [
        {
          "id": 3087846,
          "author_name": "davidlist",
          "author_url": "",
          "post_date": "01/04/2025 00:21:19",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/brendanartley\" target=\"_blank\">@brendanartley</a>.  This info is helpful for a whole variety of reasons!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 3087970,
          "author_name": "jeroencottaar",
          "author_url": "",
          "post_date": "01/04/2025 06:52:15",
          "content": "<p>Wow that's generous!</p>\n<p>Here's mine. This is not the same version of the model as the original post, but I expect differences between CV and LB to be similar.</p>\n<table>\n<thead>\n<tr>\n<th>Class</th>\n<th>CV</th>\n<th>LB</th>\n<th>Diff</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>apo-ferritin</td>\n<td>0.898</td>\n<td>0.868</td>\n<td>-0.03</td>\n</tr>\n<tr>\n<td>beta-galactosidase</td>\n<td>0.533</td>\n<td>0.431</td>\n<td>-0.102</td>\n</tr>\n<tr>\n<td>ribosome</td>\n<td>0.855</td>\n<td>0.819</td>\n<td>-0.036</td>\n</tr>\n<tr>\n<td>thyroglobulin</td>\n<td>0.685</td>\n<td>0.501</td>\n<td>-0.184</td>\n</tr>\n<tr>\n<td>VLP</td>\n<td>0.927</td>\n<td>0.812</td>\n<td>-0.115</td>\n</tr>\n</tbody>\n</table>\n<p>Seems somewhat similar, though I suffer more in thyroglobulin and VLPs.</p>",
          "votes": null,
          "replies": [
            {
              "id": 3088230,
              "author_name": "brendanartley",
              "author_url": "",
              "post_date": "01/04/2025 13:00:04",
              "content": "<p>Interesting, thanks for sharing. </p>\n<p>I wonder if there is something about hard classes on the LB that we don’t see in the 7 training samples. Could be something like shape, size, handedness, label discrepancy, etc?</p>",
              "votes": null,
              "replies": [
                {
                  "id": 3088612,
                  "author_name": "davidlist",
                  "author_url": "",
                  "post_date": "01/04/2025 22:21:24",
                  "content": "<p>Handedness for beta-galactosidase seem likely.  I know you already posted about that.  I saw some similar evidence while training with the Flipd transform in monai.  But I wonder if another issue may be that beta-galactosidase and thyroglobulin don't have spherical symmetry like the three easy ones so they require more examples.</p>",
                  "votes": null,
                  "replies": []
                }
              ]
            }
          ]
        },
        {
          "id": 3097049,
          "author_name": "huiliu0",
          "author_url": "",
          "post_date": "01/15/2025 02:03:23",
          "content": "<p>Wow, this is amazing! May I know how do you get the class-specific LB score? </p>",
          "votes": null,
          "replies": [
            {
              "id": 3097079,
              "author_name": "brendanartley",
              "author_url": "",
              "post_date": "01/15/2025 02:45:28",
              "content": "<p>Thanks <a href=\"https://www.kaggle.com/huiliu0\" target=\"_blank\">@huiliu0</a>, good question.</p>\n<p>First, you have to make a submission that only contains rows for the class you want to test. This makes the LB score for all other classes 0. </p>\n<p>Since we know the weights per class, we then can reverse engineer the class score by dividing by the sum of the class weights.</p>\n<p>Here is a python code snippet you can use.</p>\n<pre><code>\nweights = {\n    : ,\n    : ,\n    : ,\n    : ,\n    : ,\n}\n\n\nscores = {\n    : ,\n}\n\n\ntotal_weight = (weights.values())\n k  scores.keys():\n    s = (scores[k] * total_weight) / weights[k]\n    ()\n</code></pre>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3088009,
      "author_name": "davidlist",
      "author_url": "",
      "post_date": "01/04/2025 08:18:13",
      "content": "<p>So going back the Kolmogorov Smirnov test and dividing the full set of data into 16 bins…from which our 7 samples were drawn.  I'm thinking just because there were 16 bins that doesn't mean they're all equally represented.  For instance, only TS_6_4 has the elongated, non-spherical virus particles.  But does that mean 14% or 1/7th of the 500 samples have those same virus particles?  Seems unlikely.  Maybe the 7 samples are representative of the variety of the 500 though not necessarily in the same proportions.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3087431": "I've observed quite a gap between my performance in cross-validation vs. submission, so dove a bit deeper into this. \n\nI consider a simple template finding model. Both training and inference are fully reproducible. I perform 7-fold cross-validation: the model is evaluated on 1 training dataset after being trained on the 6 others. Here are the results per dataset:\n\n| Dataset | Score | \n| --- | \n| TS_5_4 | 0.702 | \n| TS_69_2 | 0.729 | \n| TS_6_4 | 0.540 | \n| TS_6_6 | 0.754 | \n| TS_73_6 | 0.684 | \n| TS_86_3 | 0.720 | \n| TS_99_9 | 0.695 | \n\n\nSo if the test set is similar to the training set, I'd expect a score of 0.689+/- 0.024 (stepping over the fact that the score is not actually an average per dataset, but I don't think it matters too much here).\n\nHowever, the actual score is 0.598. This indicates a very significant difference (p<<0.001). This seems to indicate some serious difference in distribution between the training and test set. However, as I understand it the training set was specifically selected to be similar to the test set. (EDIT: as @davidlist pointed out below, my analysis was wrong. The correct p-value is 0.014, still somewhat significant but nowhere near as wild. In addition the assumption of a Gaussian distribution seems rather inappropriate, so the whole analysis is suspect).\n\nThe model in question scales all data such that the median and standard deviation are equal. This doesn't actually help performance, but does ensure in this case that the distribution difference is not related to scale or shift.\n\nThe model has few hyperparameters; they are not overtuned on the training set.\n\nAnyone have any idea what's going on?",
    "3087440": "Statistics are my Achilles heel. I'm not going deeper than basic analysis and interpretation but as I see it. The main problem is 7 against 500. Is hard to represent properly a 500 population distribution with only 7 samples. Even more, you can only monitorize public. So I would focus into make a very general model without payng too much attention to public.",
    "3087706": ">The model has few hyperparameters; they are not overtuned on the training set.\n\nAre you optimizing for f4 score?  At least for what I'm doing that seems to be the main source of overfitting.  If you are taking the best model based on validation score, that would be another area.",
    "3087712": "Thanks for the input. I am indeed optimizing for f4 score. But that's mainly adapting the threshold, which does not lead to improvement on the public leaderboard...",
    "3087756": "thanks for sharing the insightful results. would you mind sharing which template matching model you used?",
    "3087763": "It's something I cooked together myself; it basically takes the mean over all instances of a given particle in training, and then searches for the best convolutional match (with some additional tricks). Anyway it doesn't work particularly well :(...",
    "3087769": "Actually, I think a lot of my f4 optimization is related to connected component analysis which it doesn't sound like you're doing.  But I will say, there does appear to be an unusual amount of variance at pretty much every step in this process.  Not sure how much of that explains your results, though.",
    "3087771": "I would point out that your LB score of 0.598 is actually better than your 0.540 result from TS_6_4.  Not doing a particularly rigorous statistical analysis here, but wondering if p<<0.001 is reasonable given that fact.",
    "3087777": "...Okay, got a little curious.  I get a p-value of 0.014 which is still quite significant.  ...Although, that comes with some assumptions that may or may not be valid.",
    "3087783": "thanks for sharing!",
    "3087784": "Hm just rechecked and I got 0.00014, not quite \"<<0.001\" but still significant enough. Difference in our calculation is a factor 100, maybe a percentage somewher?\n\nAnyway, the analysis is only valid if the distribution is Gaussian. It's a fair point that the TS_6_4 result already shows that it is not. A more reasonable explanation might then be that the full set has many more datasets similar to TS_6_4, and we just happened to unluckily get only a few in the training set. That's still a bit unlikely, but definitely plausible. \n\nPerhaps figuring out why TS_6_4 performs so poorly could be a clue to figuring out the poor LB performance. Thanks for your input!",
    "3087802": "I may very well be setting up the problem incorrectly.  I did a t-distribution from your averages and got a test statistic of -3.42...I also got lazy and had perplexity.ai generate the code for me, but it's not obvious to me what's wrong if it is:\n\n```\nimport numpy as np\nfrom scipy import stats\n\nsample = np.array([0.702, 0.709, 0.540, 0.754, 0.684, 0.720, 0.695])  # Your 7 data points\ntest_point = 0.598  # The point you want to test\n\nsample_mean = np.mean(sample)\nsample_std = np.std(sample, ddof=1)\nt_statistic = (test_point - sample_mean) / (sample_std / np.sqrt(len(sample)))\n\ndegrees_of_freedom = len(sample) - 1\np_value = stats.t.sf(np.abs(t_statistic), degrees_of_freedom) * 2\n\nprint(f\"T-statistic: {t_statistic}\")\nprint(f\"P-value: {p_value}\")\n```\n\nThere was another post in which they said samples were divided into 16 bins based on Kolmogorov Smirnov.  That would have obviously been reduced further from 16 down to 7 so maybe that explains some of it.",
    "3087832": "Have you done by class submissions on the LB? \n\n`thyroglobulin` and `beta-galactosidase` are dragging down my LB score more than expected.\n\n| Class          | CV | LB | Diff      |\n|----------------------|-----------------|-----------------|------------------|\n| apo-ferritin      | 0.8741          | 0.882           | 0.0079           |\n| beta-galactosidase | 0.7012          | 0.5845          | **-0.1167**          |\n| ribosome          | 0.8131          | 0.798           | -0.0151          |\n| thyroglobulin     | 0.6948          | 0.5950          | **-0.0998**          |\n| virus-like-particle | 0.9101         | 0.896           | -0.0141          |\n| Overall f-4          | 0.7699          | 0.7050          | -0.0649          |",
    "3087846": "Thanks @brendanartley.  This info is helpful for a whole variety of reasons!",
    "3087960": "You are right, I was wrong. I just assumed I could estimate the sigma accurately from my 7 points and just check the normal distribution, which is wildly inappopriate here. I'll update the original post.",
    "3087970": "Wow that's generous!\n\nHere's mine. This is not the same version of the model as the original post, but I expect differences between CV and LB to be similar.\n\n| Class | CV    | LB    | Diff   |\n| ----- | ----- | ----- | ------ |\n| apo-ferritin | 0.898 | 0.868 | -0.03  |\n| beta-galactosidase | 0.533 | 0.431 | -0.102 |\n| ribosome | 0.855 | 0.819 | -0.036 |\n| thyroglobulin | 0.685 | 0.501 | -0.184 |\n| VLP | 0.927 | 0.812 | -0.115 |\n\nSeems somewhat similar, though I suffer more in thyroglobulin and VLPs.",
    "3088009": "So going back the Kolmogorov Smirnov test and dividing the full set of data into 16 bins...from which our 7 samples were drawn.  I'm thinking just because there were 16 bins that doesn't mean they're all equally represented.  For instance, only TS_6_4 has the elongated, non-spherical virus particles.  But does that mean 14% or 1/7th of the 500 samples have those same virus particles?  Seems unlikely.  Maybe the 7 samples are representative of the variety of the 500 though not necessarily in the same proportions.",
    "3088230": "Interesting, thanks for sharing. \n\nI wonder if there is something about hard classes on the LB that we don’t see in the 7 training samples. Could be something like shape, size, handedness, label discrepancy, etc?",
    "3088612": "Handedness for beta-galactosidase seem likely.  I know you already posted about that.  I saw some similar evidence while training with the Flipd transform in monai.  But I wonder if another issue may be that beta-galactosidase and thyroglobulin don't have spherical symmetry like the three easy ones so they require more examples.",
    "3097049": "Wow, this is amazing! May I know how do you get the class-specific LB score?",
    "3097079": "Thanks @huiliu0, good question.\n\nFirst, you have to make a submission that only contains rows for the class you want to test. This makes the LB score for all other classes 0. \n\nSince we know the weights per class, we then can reverse engineer the class score by dividing by the sum of the class weights.\n\nHere is a python code snippet you can use.\n\n```\n# Weights for each class\nweights = {\n    'apo-ferritin': 1,\n    'beta-galactosidase': 2,\n    'ribosome': 1,\n    'thyroglobulin': 2,\n    'virus-like-particle': 1,\n}\n\n# LB score\nscores = {\n    'virus-like-particle': 0.133,\n}\n\n# Reverse engineer\ntotal_weight = sum(weights.values())\nfor k in scores.keys():\n    s = (scores[k] * total_weight) / weights[k]\n    print(f\"{k} Score: {s}\")\n```"
  },
  "source": "meta"
}