{
  "id": 579072,
  "title": "How to compare experiments",
  "url": "/competitions/birdclef-2025/discussion/579072",
  "author_name": "",
  "post_date": "2025-05-15T00:58:47.761986600Z",
  "votes": null,
  "comment_count": 2,
  "views": 0,
  "content": "<p>Curious how people are approaching this problem of having large swings on the LB with no correlated local test.</p>\n<p>For example, I have just run two experiments, for each one I trained 5 folds and submitted each fold (each submission uses a single model trained on 4/5 of the dataset). Here are the LB results:</p>\n<table>\n<thead>\n<tr>\n<th></th>\n<th>min</th>\n<th>mean</th>\n<th>max</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>A</td>\n<td>0.812</td>\n<td>0.823</td>\n<td>0.835</td>\n</tr>\n<tr>\n<td>B</td>\n<td>0.815</td>\n<td>0.822</td>\n<td>0.831</td>\n</tr>\n</tbody>\n</table>\n<p>Is this similar to what you are seeing? How do you compare these results when the variance is so high?</p>\n<p>Instead of submitting each fold separately, do you only submit ensembles of all folds? Do you always train on the full data instead of partitioning the dataset?</p>\n<p>I see references on this forum noting things like a \"0.06 improvement\" -- not sure how that can be trusted when repeating exact same training setup has such high variance.</p>",
  "messages": [
    {
      "id": "3202169",
      "postDate": "05/15/2025 00:58:47",
      "content": "<p>Curious how people are approaching this problem of having large swings on the LB with no correlated local test.</p>\n<p>For example, I have just run two experiments, for each one I trained 5 folds and submitted each fold (each submission uses a single model trained on 4/5 of the dataset). Here are the LB results:</p>\n<table>\n<thead>\n<tr>\n<th></th>\n<th>min</th>\n<th>mean</th>\n<th>max</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>A</td>\n<td>0.812</td>\n<td>0.823</td>\n<td>0.835</td>\n</tr>\n<tr>\n<td>B</td>\n<td>0.815</td>\n<td>0.822</td>\n<td>0.831</td>\n</tr>\n</tbody>\n</table>\n<p>Is this similar to what you are seeing? How do you compare these results when the variance is so high?</p>\n<p>Instead of submitting each fold separately, do you only submit ensembles of all folds? Do you always train on the full data instead of partitioning the dataset?</p>\n<p>I see references on this forum noting things like a \"0.06 improvement\" -- not sure how that can be trusted when repeating exact same training setup has such high variance.</p>",
      "rawMarkdown": "Curious how people are approaching this problem of having large swings on the LB with no correlated local test.\n\nFor example, I have just run two experiments, for each one I trained 5 folds and submitted each fold (each submission uses a single model trained on 4/5 of the dataset). Here are the LB results:\n\n|  | min | mean | max\n| --- | --- | --- | --- |\n| A | 0.812 | 0.823 | 0.835\n| B | 0.815 | 0.822 | 0.831\n\n\nIs this similar to what you are seeing? How do you compare these results when the variance is so high?\n\nInstead of submitting each fold separately, do you only submit ensembles of all folds? Do you always train on the full data instead of partitioning the dataset?\n\nI see references on this forum noting things like a \"0.06 improvement\" -- not sure how that can be trusted when repeating exact same training setup has such high variance.",
      "votes": null
    },
    {
      "id": "3203130",
      "postDate": "05/16/2025 10:34:44",
      "content": "<p>I initially started with 5-fold cross-validation but later switched to using the full dataset, as it seemed to give me a better LB score.</p>\n<p>In my experiments, I also encountered a high variance issue. To mitigate this, I ran 5 experiments using the same setting and submitted an ensemble of the 5 models (by averaging their outputs).</p>\n<p>By the way, may I ask how you managed to improve your single model's LB score from 0.812 to 0.835? I'm still struggling to push mine past 0.82.</p>",
      "rawMarkdown": "I initially started with 5-fold cross-validation but later switched to using the full dataset, as it seemed to give me a better LB score.\n\nIn my experiments, I also encountered a high variance issue. To mitigate this, I ran 5 experiments using the same setting and submitted an ensemble of the 5 models (by averaging their outputs).\n\nBy the way, may I ask how you managed to improve your single model's LB score from 0.812 to 0.835? I'm still struggling to push mine past 0.82.",
      "votes": null
    },
    {
      "id": "3203273",
      "postDate": "05/16/2025 14:12:01",
      "content": "<p>Training on the full data seems reasonable, though oddly I've found that models trained on the full data have measurably lower LB score. Also, I'm not sure when to stop training without a validation set.</p>\n<p>Do you also plan to submit those ensembles as your final solution? I've been planning to eventually combine single models from different experiments in my final ensembles, but for that I feel like I need to know individual model performance rather than ensemble performance of an experiment.</p>\n<p>I haven't reliably improved score from 0.812 to 0.835 -- those are just min and max folds (the same exact training set up repeated over 5 folds). Overall I've seen measurable improvements from adding augmentations, incorporating focal loss, tweaking spectrogram params, and smoothing predictions in post processing, similar what others have found.</p>",
      "rawMarkdown": "Training on the full data seems reasonable, though oddly I've found that models trained on the full data have measurably lower LB score. Also, I'm not sure when to stop training without a validation set.\n\nDo you also plan to submit those ensembles as your final solution? I've been planning to eventually combine single models from different experiments in my final ensembles, but for that I feel like I need to know individual model performance rather than ensemble performance of an experiment.\n\nI haven't reliably improved score from 0.812 to 0.835 -- those are just min and max folds (the same exact training set up repeated over 5 folds). Overall I've seen measurable improvements from adding augmentations, incorporating focal loss, tweaking spectrogram params, and smoothing predictions in post processing, similar what others have found.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3203130,
      "author_name": "fangsionfang",
      "author_url": "",
      "post_date": "05/16/2025 10:34:44",
      "content": "<p>I initially started with 5-fold cross-validation but later switched to using the full dataset, as it seemed to give me a better LB score.</p>\n<p>In my experiments, I also encountered a high variance issue. To mitigate this, I ran 5 experiments using the same setting and submitted an ensemble of the 5 models (by averaging their outputs).</p>\n<p>By the way, may I ask how you managed to improve your single model's LB score from 0.812 to 0.835? I'm still struggling to push mine past 0.82.</p>",
      "votes": null,
      "replies": [
        {
          "id": 3203273,
          "author_name": "robbynevels",
          "author_url": "",
          "post_date": "05/16/2025 14:12:01",
          "content": "<p>Training on the full data seems reasonable, though oddly I've found that models trained on the full data have measurably lower LB score. Also, I'm not sure when to stop training without a validation set.</p>\n<p>Do you also plan to submit those ensembles as your final solution? I've been planning to eventually combine single models from different experiments in my final ensembles, but for that I feel like I need to know individual model performance rather than ensemble performance of an experiment.</p>\n<p>I haven't reliably improved score from 0.812 to 0.835 -- those are just min and max folds (the same exact training set up repeated over 5 folds). Overall I've seen measurable improvements from adding augmentations, incorporating focal loss, tweaking spectrogram params, and smoothing predictions in post processing, similar what others have found.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3202169": "Curious how people are approaching this problem of having large swings on the LB with no correlated local test.\n\nFor example, I have just run two experiments, for each one I trained 5 folds and submitted each fold (each submission uses a single model trained on 4/5 of the dataset). Here are the LB results:\n\n|  | min | mean | max\n| --- | --- | --- | --- |\n| A | 0.812 | 0.823 | 0.835\n| B | 0.815 | 0.822 | 0.831\n\n\nIs this similar to what you are seeing? How do you compare these results when the variance is so high?\n\nInstead of submitting each fold separately, do you only submit ensembles of all folds? Do you always train on the full data instead of partitioning the dataset?\n\nI see references on this forum noting things like a \"0.06 improvement\" -- not sure how that can be trusted when repeating exact same training setup has such high variance.",
    "3203130": "I initially started with 5-fold cross-validation but later switched to using the full dataset, as it seemed to give me a better LB score.\n\nIn my experiments, I also encountered a high variance issue. To mitigate this, I ran 5 experiments using the same setting and submitted an ensemble of the 5 models (by averaging their outputs).\n\nBy the way, may I ask how you managed to improve your single model's LB score from 0.812 to 0.835? I'm still struggling to push mine past 0.82.",
    "3203273": "Training on the full data seems reasonable, though oddly I've found that models trained on the full data have measurably lower LB score. Also, I'm not sure when to stop training without a validation set.\n\nDo you also plan to submit those ensembles as your final solution? I've been planning to eventually combine single models from different experiments in my final ensembles, but for that I feel like I need to know individual model performance rather than ensemble performance of an experiment.\n\nI haven't reliably improved score from 0.812 to 0.835 -- those are just min and max folds (the same exact training set up repeated over 5 folds). Overall I've seen measurable improvements from adding augmentations, incorporating focal loss, tweaking spectrogram params, and smoothing predictions in post processing, similar what others have found."
  },
  "source": "meta"
}