{
  "id": 570402,
  "title": "Unstable Experiments and Key Early Takeaways",
  "url": "/competitions/birdclef-2025/discussion/570402",
  "author_name": "",
  "post_date": "2025-03-27T15:19:16.227302900Z",
  "votes": 55,
  "comment_count": 20,
  "views": 0,
  "content": "<p>I am sharing some of the important takeaways from my experiments to improve the LB.<br>\nI wanted to quickly iterate through multiple hyper parameters to rank the best performing ones from start of the competition. I'll try to share more.</p>\n<ul>\n<li><p>I was unable to get my model with RAW Audio above LB: 0.697. (I've tried a lot of ways to make it work.)</p></li>\n<li><p>I've tried multiple N_FFT values but for some reason 2048 works best. It looks like 4096 would be too much zoom into the signal and there maybe more noise, and 1048 most likely miss those features.<br>\nOther hyperparameters  impact the performance but not too much compared to N_FFT.</p></li>\n<li><p>Somehow, EfficientNet-B0 and ECA_NFNET_L1 performs much much better compared to other models like seresnext, resnets, efficientvit. (For the exact same settings of the experiment B0 scored 0.83 and others scored around 0.80) [I tried to look into the reason, but unfortunately there is no appropriate validation set to look into, all of the models have same CV on basic KFold split so that doesn't help either.] [My only assumption is other models might be overfitting  because of the large size. If that's the case then why ECA_NFNET_L1 didn't overfit, might be because it's norm free?]</p></li>\n<li><p>Since there is no other way to properly validate if model is overfitting, I tried to test performance of model from different epochs, My same EfficientNet-B0 performs as the follow: [Epoch 15 LB: 0.811, Epoch 17 LB: 0.796, Epoch 25 LB: 0.790, …. Epoch 30 LB: 0.813]<br>\n[I think the reason might be the highly imbalanced classes, because in epoch K model add more weights to a class C with just 1 or 2 samples than class D with just 1 or 2 samples, and in Epoch K+1  this was reversed. [To handle this I did Checkpoint Averaging, and that helped to make it stable a little.]</p></li>\n<li><p>Now the most interesting thing, for the exact same experiment just by changing the seed and slightly changing  the probability of a single augmentation type, my LB score is changing drastically. </p></li>\n</ul>",
  "messages": [
    {
      "id": "3161140",
      "postDate": "03/27/2025 15:19:16",
      "content": "<p>I am sharing some of the important takeaways from my experiments to improve the LB.<br>\nI wanted to quickly iterate through multiple hyper parameters to rank the best performing ones from start of the competition. I'll try to share more.</p>\n<ul>\n<li><p>I was unable to get my model with RAW Audio above LB: 0.697. (I've tried a lot of ways to make it work.)</p></li>\n<li><p>I've tried multiple N_FFT values but for some reason 2048 works best. It looks like 4096 would be too much zoom into the signal and there maybe more noise, and 1048 most likely miss those features.<br>\nOther hyperparameters  impact the performance but not too much compared to N_FFT.</p></li>\n<li><p>Somehow, EfficientNet-B0 and ECA_NFNET_L1 performs much much better compared to other models like seresnext, resnets, efficientvit. (For the exact same settings of the experiment B0 scored 0.83 and others scored around 0.80) [I tried to look into the reason, but unfortunately there is no appropriate validation set to look into, all of the models have same CV on basic KFold split so that doesn't help either.] [My only assumption is other models might be overfitting  because of the large size. If that's the case then why ECA_NFNET_L1 didn't overfit, might be because it's norm free?]</p></li>\n<li><p>Since there is no other way to properly validate if model is overfitting, I tried to test performance of model from different epochs, My same EfficientNet-B0 performs as the follow: [Epoch 15 LB: 0.811, Epoch 17 LB: 0.796, Epoch 25 LB: 0.790, …. Epoch 30 LB: 0.813]<br>\n[I think the reason might be the highly imbalanced classes, because in epoch K model add more weights to a class C with just 1 or 2 samples than class D with just 1 or 2 samples, and in Epoch K+1  this was reversed. [To handle this I did Checkpoint Averaging, and that helped to make it stable a little.]</p></li>\n<li><p>Now the most interesting thing, for the exact same experiment just by changing the seed and slightly changing  the probability of a single augmentation type, my LB score is changing drastically. </p></li>\n</ul>",
      "rawMarkdown": "I am sharing some of the important takeaways from my experiments to improve the LB.\nI wanted to quickly iterate through multiple hyper parameters to rank the best performing ones from start of the competition. I'll try to share more.\n\n- I was unable to get my model with RAW Audio above LB: 0.697. (I've tried a lot of ways to make it work.)\n\n- I've tried multiple N_FFT values but for some reason 2048 works best. It looks like 4096 would be too much zoom into the signal and there maybe more noise, and 1048 most likely miss those features.\nOther hyperparameters  impact the performance but not too much compared to N_FFT.\n\n- Somehow, EfficientNet-B0 and ECA_NFNET_L1 performs much much better compared to other models like seresnext, resnets, efficientvit. (For the exact same settings of the experiment B0 scored 0.83 and others scored around 0.80) [I tried to look into the reason, but unfortunately there is no appropriate validation set to look into, all of the models have same CV on basic KFold split so that doesn't help either.] [My only assumption is other models might be overfitting  because of the large size. If that's the case then why ECA_NFNET_L1 didn't overfit, might be because it's norm free?]\n\n- Since there is no other way to properly validate if model is overfitting, I tried to test performance of model from different epochs, My same EfficientNet-B0 performs as the follow: [Epoch 15 LB: 0.811, Epoch 17 LB: 0.796, Epoch 25 LB: 0.790, .... Epoch 30 LB: 0.813]\n[I think the reason might be the highly imbalanced classes, because in epoch K model add more weights to a class C with just 1 or 2 samples than class D with just 1 or 2 samples, and in Epoch K+1  this was reversed. [To handle this I did Checkpoint Averaging, and that helped to make it stable a little.]\n\n- Now the most interesting thing, for the exact same experiment just by changing the seed and slightly changing  the probability of a single augmentation type, my LB score is changing drastically.",
      "votes": null
    },
    {
      "id": "3161197",
      "postDate": "03/27/2025 16:21:44",
      "content": "<p>excellent, thank you for sharing. I have noticed similar patterns regarding the number of epochs, do the 30+ epoch tries always perform better? I suppose it is highly correlated of how you feed the dataset as well, eg random samples etc..</p>",
      "rawMarkdown": "excellent, thank you for sharing. I have noticed similar patterns regarding the number of epochs, do the 30+ epoch tries always perform better? I suppose it is highly correlated of how you feed the dataset as well, eg random samples etc..",
      "votes": null
    },
    {
      "id": "3161316",
      "postDate": "03/27/2025 18:26:35",
      "content": "<p>No it’s not always better, but I try to make it work with certain checkpoints averaging. </p>",
      "rawMarkdown": "No it’s not always better, but I try to make it work with certain checkpoints averaging.",
      "votes": null
    },
    {
      "id": "3161568",
      "postDate": "03/28/2025 05:21:40",
      "content": "<p>how long is the duration of your sampling window? for me, 5 sec gave LB 0.76, 10 sec gave 0.803 and 7 sec gave 0.82x. </p>\n<p>haven't played with other parameters or pseudo labeling yet. currently focusing on optimizing Inference pipeline</p>",
      "rawMarkdown": "how long is the duration of your sampling window? for me, 5 sec gave LB 0.76, 10 sec gave 0.803 and 7 sec gave 0.82x. \n\nhaven't played with other parameters or pseudo labeling yet. currently focusing on optimizing Inference pipeline",
      "votes": null
    },
    {
      "id": "3161654",
      "postDate": "03/28/2025 08:08:09",
      "content": "<p>I have been trying with random 5 seconds till now. Will explore other ways to handle this later.</p>",
      "rawMarkdown": "I have been trying with random 5 seconds till now. Will explore other ways to handle this later.",
      "votes": null
    },
    {
      "id": "3161667",
      "postDate": "03/28/2025 08:39:42",
      "content": "<p>What does the window duration you mentioned mean</p>",
      "rawMarkdown": "What does the window duration you mentioned mean",
      "votes": null
    },
    {
      "id": "3161735",
      "postDate": "03/28/2025 10:32:49",
      "content": "<p>Are you using kfold split or simple 80, 20 split ?</p>",
      "rawMarkdown": "Are you using kfold split or simple 80, 20 split ?",
      "votes": null
    },
    {
      "id": "3161849",
      "postDate": "03/28/2025 13:43:50",
      "content": "<p>In my experiments, the <code>random_seed</code> seems to be the most important factor (ha-ha). Using the same training procedure with the same model architecture and data and having almost the same CV results and loss value at training, I can get 0.8+ LB for <code>random_seed==0</code> and ~0.3-0.4 LB for <code>random_seed==1</code>.</p>\n<p>So, I'm working at searching for something more stable and reliable.</p>",
      "rawMarkdown": "In my experiments, the `random_seed` seems to be the most important factor (ha-ha). Using the same training procedure with the same model architecture and data and having almost the same CV results and loss value at training, I can get 0.8+ LB for `random_seed==0` and ~0.3-0.4 LB for `random_seed==1`.\n\nSo, I'm working at searching for something more stable and reliable.",
      "votes": null
    },
    {
      "id": "3161852",
      "postDate": "03/28/2025 13:50:27",
      "content": "<p>wow. huge swings</p>",
      "rawMarkdown": "wow. huge swings",
      "votes": null
    },
    {
      "id": "3161937",
      "postDate": "03/28/2025 15:56:25",
      "content": "<blockquote>\n  <p>for the exact same experiment just by changing the seed and slightly changing the probability of a single augmentation type, my LB score is changing drastically</p>\n</blockquote>\n<p>early LB shake alert 😅</p>",
      "rawMarkdown": ">for the exact same experiment just by changing the seed and slightly changing the probability of a single augmentation type, my LB score is changing drastically\n\nearly LB shake alert 😅",
      "votes": null
    },
    {
      "id": "3162530",
      "postDate": "03/29/2025 10:36:59",
      "content": "<p>Yeah, I was confused with such behavior of some models, but later found a stable way to train. </p>",
      "rawMarkdown": "Yeah, I was confused with such behavior of some models, but later found a stable way to train.",
      "votes": null
    },
    {
      "id": "3162532",
      "postDate": "03/29/2025 10:37:46",
      "content": "<p>If you want to train a single fold then both are same thing.</p>",
      "rawMarkdown": "If you want to train a single fold then both are same thing.",
      "votes": null
    },
    {
      "id": "3162612",
      "postDate": "03/29/2025 13:19:48",
      "content": "<p>When you set seed, do you make everything deterministic too, including torchtorch.backends.cudnn.deterministic = True etc? Otherwise the same seed would result in different models and LB results as well. But I believe doing that can also slow things down</p>",
      "rawMarkdown": "When you set seed, do you make everything deterministic too, including torchtorch.backends.cudnn.deterministic = True etc? Otherwise the same seed would result in different models and LB results as well. But I believe doing that can also slow things down",
      "votes": null
    },
    {
      "id": "3162616",
      "postDate": "03/29/2025 13:24:18",
      "content": "<p>Did you ensemble different b0 and eca-nfnet-l1 to achieve current score?</p>",
      "rawMarkdown": "Did you ensemble different b0 and eca-nfnet-l1 to achieve current score?",
      "votes": null
    },
    {
      "id": "3162990",
      "postDate": "03/30/2025 06:24:15",
      "content": "<p>Thanks! This is a great starting point.</p>",
      "rawMarkdown": "Thanks! This is a great starting point.",
      "votes": null
    },
    {
      "id": "3163174",
      "postDate": "03/30/2025 11:42:55",
      "content": "<p>I am new here so I don't understand much but it seems accuracy problems exist for even professionals😅</p>",
      "rawMarkdown": "I am new here so I don't understand much but it seems accuracy problems exist for even professionals😅",
      "votes": null
    },
    {
      "id": "3171542",
      "postDate": "04/05/2025 19:13:15",
      "content": "<p>interesting to me N_FFT of 1024 worked best for me, could be due to the instability across epochs. Are you just averaging checkpoints at the end or during your training? I would be curious to learn about that method a bit</p>",
      "rawMarkdown": "interesting to me N_FFT of 1024 worked best for me, could be due to the instability across epochs. Are you just averaging checkpoints at the end or during your training? I would be curious to learn about that method a bit",
      "votes": null
    },
    {
      "id": "3171632",
      "postDate": "04/05/2025 21:32:29",
      "content": "<p>There you go: <a href=\"https://github.com/mlfoundations/model-soups\" target=\"_blank\">https://github.com/mlfoundations/model-soups</a></p>",
      "rawMarkdown": "There you go: [https://github.com/mlfoundations/model-soups](https://github.com/mlfoundations/model-soups)",
      "votes": null
    },
    {
      "id": "3173277",
      "postDate": "04/07/2025 17:57:52",
      "content": "<p>interesting for me, the average of weights does not make really make any change</p>",
      "rawMarkdown": "interesting for me, the average of weights does not make really make any change",
      "votes": null
    },
    {
      "id": "3203699",
      "postDate": "05/17/2025 07:27:22",
      "content": "<p>What strategy do you have for selecting which epoch to use?(average different checkpoints?) If I train for 15 epochs it might be epoch 15 with best performanc or epoch 4. It is very time consuming to use 5 days worth of submissions to test only one model..<br>\nI cannot get better performance than 0.83 from a single effnetb0.</p>",
      "rawMarkdown": "What strategy do you have for selecting which epoch to use?(average different checkpoints?) If I train for 15 epochs it might be epoch 15 with best performanc or epoch 4. It is very time consuming to use 5 days worth of submissions to test only one model..\nI cannot get better performance than 0.83 from a single effnetb0.",
      "votes": null
    },
    {
      "id": "3203744",
      "postDate": "05/17/2025 08:57:49",
      "content": "<p>Great insights <a href=\"https://www.kaggle.com/salmanahmedtamu\" target=\"_blank\">@salmanahmedtamu</a> ! The N_FFT value impacts the model performance by a lot, also removing the human audio parts helped me !</p>",
      "rawMarkdown": "Great insights @salmanahmedtamu ! The N_FFT value impacts the model performance by a lot, also removing the human audio parts helped me !",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3161197,
      "author_name": "left13",
      "author_url": "",
      "post_date": "03/27/2025 16:21:44",
      "content": "<p>excellent, thank you for sharing. I have noticed similar patterns regarding the number of epochs, do the 30+ epoch tries always perform better? I suppose it is highly correlated of how you feed the dataset as well, eg random samples etc..</p>",
      "votes": null,
      "replies": [
        {
          "id": 3161316,
          "author_name": "salmanahmedtamu",
          "author_url": "",
          "post_date": "03/27/2025 18:26:35",
          "content": "<p>No it’s not always better, but I try to make it work with certain checkpoints averaging. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3161568,
      "author_name": "nymfree",
      "author_url": "",
      "post_date": "03/28/2025 05:21:40",
      "content": "<p>how long is the duration of your sampling window? for me, 5 sec gave LB 0.76, 10 sec gave 0.803 and 7 sec gave 0.82x. </p>\n<p>haven't played with other parameters or pseudo labeling yet. currently focusing on optimizing Inference pipeline</p>",
      "votes": null,
      "replies": [
        {
          "id": 3161654,
          "author_name": "salmanahmedtamu",
          "author_url": "",
          "post_date": "03/28/2025 08:08:09",
          "content": "<p>I have been trying with random 5 seconds till now. Will explore other ways to handle this later.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 3161667,
          "author_name": "bigbigclever",
          "author_url": "",
          "post_date": "03/28/2025 08:39:42",
          "content": "<p>What does the window duration you mentioned mean</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3161735,
      "author_name": "arunodhayan",
      "author_url": "",
      "post_date": "03/28/2025 10:32:49",
      "content": "<p>Are you using kfold split or simple 80, 20 split ?</p>",
      "votes": null,
      "replies": [
        {
          "id": 3162532,
          "author_name": "salmanahmedtamu",
          "author_url": "",
          "post_date": "03/29/2025 10:37:46",
          "content": "<p>If you want to train a single fold then both are same thing.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3161849,
      "author_name": "kdmitrie",
      "author_url": "",
      "post_date": "03/28/2025 13:43:50",
      "content": "<p>In my experiments, the <code>random_seed</code> seems to be the most important factor (ha-ha). Using the same training procedure with the same model architecture and data and having almost the same CV results and loss value at training, I can get 0.8+ LB for <code>random_seed==0</code> and ~0.3-0.4 LB for <code>random_seed==1</code>.</p>\n<p>So, I'm working at searching for something more stable and reliable.</p>",
      "votes": null,
      "replies": [
        {
          "id": 3161852,
          "author_name": "nymfree",
          "author_url": "",
          "post_date": "03/28/2025 13:50:27",
          "content": "<p>wow. huge swings</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 3162530,
          "author_name": "salmanahmedtamu",
          "author_url": "",
          "post_date": "03/29/2025 10:36:59",
          "content": "<p>Yeah, I was confused with such behavior of some models, but later found a stable way to train. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 3162612,
          "author_name": "robbynevels",
          "author_url": "",
          "post_date": "03/29/2025 13:19:48",
          "content": "<p>When you set seed, do you make everything deterministic too, including torchtorch.backends.cudnn.deterministic = True etc? Otherwise the same seed would result in different models and LB results as well. But I believe doing that can also slow things down</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3161937,
      "author_name": "nyleve",
      "author_url": "",
      "post_date": "03/28/2025 15:56:25",
      "content": "<blockquote>\n  <p>for the exact same experiment just by changing the seed and slightly changing the probability of a single augmentation type, my LB score is changing drastically</p>\n</blockquote>\n<p>early LB shake alert 😅</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3162616,
      "author_name": "yuanzhezhou",
      "author_url": "",
      "post_date": "03/29/2025 13:24:18",
      "content": "<p>Did you ensemble different b0 and eca-nfnet-l1 to achieve current score?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3162990,
      "author_name": "armaankhan2007",
      "author_url": "",
      "post_date": "03/30/2025 06:24:15",
      "content": "<p>Thanks! This is a great starting point.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3163174,
      "author_name": "adityanshsrivastava",
      "author_url": "",
      "post_date": "03/30/2025 11:42:55",
      "content": "<p>I am new here so I don't understand much but it seems accuracy problems exist for even professionals😅</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3171542,
      "author_name": "firlas47",
      "author_url": "",
      "post_date": "04/05/2025 19:13:15",
      "content": "<p>interesting to me N_FFT of 1024 worked best for me, could be due to the instability across epochs. Are you just averaging checkpoints at the end or during your training? I would be curious to learn about that method a bit</p>",
      "votes": null,
      "replies": [
        {
          "id": 3171632,
          "author_name": "salmanahmedtamu",
          "author_url": "",
          "post_date": "04/05/2025 21:32:29",
          "content": "<p>There you go: <a href=\"https://github.com/mlfoundations/model-soups\" target=\"_blank\">https://github.com/mlfoundations/model-soups</a></p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3173277,
      "author_name": "ludovick",
      "author_url": "",
      "post_date": "04/07/2025 17:57:52",
      "content": "<p>interesting for me, the average of weights does not make really make any change</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3203699,
      "author_name": "adgr4416",
      "author_url": "",
      "post_date": "05/17/2025 07:27:22",
      "content": "<p>What strategy do you have for selecting which epoch to use?(average different checkpoints?) If I train for 15 epochs it might be epoch 15 with best performanc or epoch 4. It is very time consuming to use 5 days worth of submissions to test only one model..<br>\nI cannot get better performance than 0.83 from a single effnetb0.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3203744,
      "author_name": "digantabhattacharya",
      "author_url": "",
      "post_date": "05/17/2025 08:57:49",
      "content": "<p>Great insights <a href=\"https://www.kaggle.com/salmanahmedtamu\" target=\"_blank\">@salmanahmedtamu</a> ! The N_FFT value impacts the model performance by a lot, also removing the human audio parts helped me !</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3161140": "I am sharing some of the important takeaways from my experiments to improve the LB.\nI wanted to quickly iterate through multiple hyper parameters to rank the best performing ones from start of the competition. I'll try to share more.\n\n- I was unable to get my model with RAW Audio above LB: 0.697. (I've tried a lot of ways to make it work.)\n\n- I've tried multiple N_FFT values but for some reason 2048 works best. It looks like 4096 would be too much zoom into the signal and there maybe more noise, and 1048 most likely miss those features.\nOther hyperparameters  impact the performance but not too much compared to N_FFT.\n\n- Somehow, EfficientNet-B0 and ECA_NFNET_L1 performs much much better compared to other models like seresnext, resnets, efficientvit. (For the exact same settings of the experiment B0 scored 0.83 and others scored around 0.80) [I tried to look into the reason, but unfortunately there is no appropriate validation set to look into, all of the models have same CV on basic KFold split so that doesn't help either.] [My only assumption is other models might be overfitting  because of the large size. If that's the case then why ECA_NFNET_L1 didn't overfit, might be because it's norm free?]\n\n- Since there is no other way to properly validate if model is overfitting, I tried to test performance of model from different epochs, My same EfficientNet-B0 performs as the follow: [Epoch 15 LB: 0.811, Epoch 17 LB: 0.796, Epoch 25 LB: 0.790, .... Epoch 30 LB: 0.813]\n[I think the reason might be the highly imbalanced classes, because in epoch K model add more weights to a class C with just 1 or 2 samples than class D with just 1 or 2 samples, and in Epoch K+1  this was reversed. [To handle this I did Checkpoint Averaging, and that helped to make it stable a little.]\n\n- Now the most interesting thing, for the exact same experiment just by changing the seed and slightly changing  the probability of a single augmentation type, my LB score is changing drastically.",
    "3161197": "excellent, thank you for sharing. I have noticed similar patterns regarding the number of epochs, do the 30+ epoch tries always perform better? I suppose it is highly correlated of how you feed the dataset as well, eg random samples etc..",
    "3161316": "No it’s not always better, but I try to make it work with certain checkpoints averaging.",
    "3161568": "how long is the duration of your sampling window? for me, 5 sec gave LB 0.76, 10 sec gave 0.803 and 7 sec gave 0.82x. \n\nhaven't played with other parameters or pseudo labeling yet. currently focusing on optimizing Inference pipeline",
    "3161654": "I have been trying with random 5 seconds till now. Will explore other ways to handle this later.",
    "3161667": "What does the window duration you mentioned mean",
    "3161735": "Are you using kfold split or simple 80, 20 split ?",
    "3161849": "In my experiments, the `random_seed` seems to be the most important factor (ha-ha). Using the same training procedure with the same model architecture and data and having almost the same CV results and loss value at training, I can get 0.8+ LB for `random_seed==0` and ~0.3-0.4 LB for `random_seed==1`.\n\nSo, I'm working at searching for something more stable and reliable.",
    "3161852": "wow. huge swings",
    "3161937": ">for the exact same experiment just by changing the seed and slightly changing the probability of a single augmentation type, my LB score is changing drastically\n\nearly LB shake alert 😅",
    "3162530": "Yeah, I was confused with such behavior of some models, but later found a stable way to train.",
    "3162532": "If you want to train a single fold then both are same thing.",
    "3162612": "When you set seed, do you make everything deterministic too, including torchtorch.backends.cudnn.deterministic = True etc? Otherwise the same seed would result in different models and LB results as well. But I believe doing that can also slow things down",
    "3162616": "Did you ensemble different b0 and eca-nfnet-l1 to achieve current score?",
    "3162990": "Thanks! This is a great starting point.",
    "3163174": "I am new here so I don't understand much but it seems accuracy problems exist for even professionals😅",
    "3171542": "interesting to me N_FFT of 1024 worked best for me, could be due to the instability across epochs. Are you just averaging checkpoints at the end or during your training? I would be curious to learn about that method a bit",
    "3171632": "There you go: [https://github.com/mlfoundations/model-soups](https://github.com/mlfoundations/model-soups)",
    "3173277": "interesting for me, the average of weights does not make really make any change",
    "3203699": "What strategy do you have for selecting which epoch to use?(average different checkpoints?) If I train for 15 epochs it might be epoch 15 with best performanc or epoch 4. It is very time consuming to use 5 days worth of submissions to test only one model..\nI cannot get better performance than 0.83 from a single effnetb0.",
    "3203744": "Great insights @salmanahmedtamu ! The N_FFT value impacts the model performance by a lot, also removing the human audio parts helped me !"
  },
  "source": "meta"
}