{
  "id": 568303,
  "title": "CV   vs    LB. (Unreliable CV. Renamed - Initial experiments and LB)",
  "url": "/competitions/birdclef-2025/discussion/568303",
  "author_name": "",
  "post_date": "2025-03-15T03:18:50.451362800Z",
  "votes": 45,
  "comment_count": 33,
  "views": 0,
  "content": "<p>CV vs LB for some of my initial experiments. <br>\n5 R means Random 5 seconds for training, and first 5 seconds for validation. </p>\n<table>\n<thead>\n<tr>\n<th>Datasets</th>\n<th>Fold</th>\n<th>Backbone</th>\n<th>CV mAP</th>\n<th>CV Score</th>\n<th>LB</th>\n<th>Type</th>\n<th>Shape</th>\n<th>Aug</th>\n<th>Model Additional</th>\n<th>Random TTA</th>\n<th>Epochs</th>\n<th>Chunk</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>2025</td>\n<td>1</td>\n<td>hgnet_v2_b0</td>\n<td>0.7</td>\n<td></td>\n<td>0.633</td>\n<td>RAW Signal</td>\n<td>2 x 160 x -1</td>\n<td></td>\n<td></td>\n<td></td>\n<td>30</td>\n<td>5 R</td>\n</tr>\n<tr>\n<td></td>\n<td>1</td>\n<td></td>\n<td></td>\n<td></td>\n<td>0.648</td>\n<td>RAW Signal</td>\n<td>1 x 160 x -1 (3 channels)</td>\n<td></td>\n<td></td>\n<td></td>\n<td>40</td>\n<td>5 R</td>\n</tr>\n<tr>\n<td></td>\n<td>1</td>\n<td>tf_efficientnet_b0</td>\n<td>0.62</td>\n<td>0.9529</td>\n<td>0.697</td>\n<td>RAW Signal</td>\n<td>2 x -1 x 80</td>\n<td>Cutmix</td>\n<td></td>\n<td>Yes</td>\n<td>15</td>\n<td>5 R</td>\n</tr>\n<tr>\n<td></td>\n<td>1</td>\n<td>tf_efficientnet_b0</td>\n<td>0.759</td>\n<td>0.97</td>\n<td>0.768</td>\n<td>MelSpec</td>\n<td></td>\n<td>Mixup</td>\n<td></td>\n<td></td>\n<td>17</td>\n<td>5 R</td>\n</tr>\n<tr>\n<td></td>\n<td>1</td>\n<td>tf_efficientnet_b0</td>\n<td>0.759</td>\n<td>0.97</td>\n<td>0.788</td>\n<td>MelSpec</td>\n<td></td>\n<td>Mixup</td>\n<td></td>\n<td>Yes</td>\n<td>17</td>\n<td>5 R</td>\n</tr>\n<tr>\n<td></td>\n<td>1</td>\n<td>tf_efficientnet_b0</td>\n<td>0.768</td>\n<td>0.97</td>\n<td>0.776</td>\n<td>MelSpec</td>\n<td></td>\n<td>Mixup</td>\n<td></td>\n<td>Yes</td>\n<td>20</td>\n<td>10 R</td>\n</tr>\n<tr>\n<td></td>\n<td>1</td>\n<td>tf_efficientnet_b0</td>\n<td>0.74</td>\n<td>0.97</td>\n<td>0.811</td>\n<td>MelSpec</td>\n<td></td>\n<td>Mixup</td>\n<td></td>\n<td>Yes</td>\n<td>15</td>\n<td>5 R</td>\n</tr>\n<tr>\n<td></td>\n<td>1</td>\n<td>FCAN B0</td>\n<td>0.75</td>\n<td>0.97</td>\n<td>0.80</td>\n<td>MelSpec</td>\n<td></td>\n<td>Mixup</td>\n<td></td>\n<td>Yes</td>\n<td>15</td>\n<td>5 R</td>\n</tr>\n<tr>\n<td></td>\n<td>1</td>\n<td>FCAN prediction B0</td>\n<td>0.76</td>\n<td>0.97</td>\n<td>0.78</td>\n<td>MelSpec</td>\n<td></td>\n<td>Mixup</td>\n<td></td>\n<td>Yes</td>\n<td>15</td>\n<td>5 R</td>\n</tr>\n<tr>\n<td></td>\n<td>4+1</td>\n<td>tf_efficientnet_b0</td>\n<td>0.77</td>\n<td>0.97</td>\n<td>0.826</td>\n<td>MelSpec</td>\n<td></td>\n<td>Mixup</td>\n<td></td>\n<td>No</td>\n<td>40</td>\n<td>5 R</td>\n</tr>\n<tr>\n<td></td>\n<td>4+1</td>\n<td>tf_efficientnet_b0</td>\n<td>0.77</td>\n<td>0.97</td>\n<td>0.825</td>\n<td>MelSpec</td>\n<td></td>\n<td>Mixup</td>\n<td></td>\n<td>Yes</td>\n<td>40</td>\n<td>5 R</td>\n</tr>\n</tbody>\n</table>",
  "messages": [
    {
      "id": "3150065",
      "postDate": "03/15/2025 03:18:50",
      "content": "<p>CV vs LB for some of my initial experiments. <br>\n5 R means Random 5 seconds for training, and first 5 seconds for validation. </p>\n<table>\n<thead>\n<tr>\n<th>Datasets</th>\n<th>Fold</th>\n<th>Backbone</th>\n<th>CV mAP</th>\n<th>CV Score</th>\n<th>LB</th>\n<th>Type</th>\n<th>Shape</th>\n<th>Aug</th>\n<th>Model Additional</th>\n<th>Random TTA</th>\n<th>Epochs</th>\n<th>Chunk</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>2025</td>\n<td>1</td>\n<td>hgnet_v2_b0</td>\n<td>0.7</td>\n<td></td>\n<td>0.633</td>\n<td>RAW Signal</td>\n<td>2 x 160 x -1</td>\n<td></td>\n<td></td>\n<td></td>\n<td>30</td>\n<td>5 R</td>\n</tr>\n<tr>\n<td></td>\n<td>1</td>\n<td></td>\n<td></td>\n<td></td>\n<td>0.648</td>\n<td>RAW Signal</td>\n<td>1 x 160 x -1 (3 channels)</td>\n<td></td>\n<td></td>\n<td></td>\n<td>40</td>\n<td>5 R</td>\n</tr>\n<tr>\n<td></td>\n<td>1</td>\n<td>tf_efficientnet_b0</td>\n<td>0.62</td>\n<td>0.9529</td>\n<td>0.697</td>\n<td>RAW Signal</td>\n<td>2 x -1 x 80</td>\n<td>Cutmix</td>\n<td></td>\n<td>Yes</td>\n<td>15</td>\n<td>5 R</td>\n</tr>\n<tr>\n<td></td>\n<td>1</td>\n<td>tf_efficientnet_b0</td>\n<td>0.759</td>\n<td>0.97</td>\n<td>0.768</td>\n<td>MelSpec</td>\n<td></td>\n<td>Mixup</td>\n<td></td>\n<td></td>\n<td>17</td>\n<td>5 R</td>\n</tr>\n<tr>\n<td></td>\n<td>1</td>\n<td>tf_efficientnet_b0</td>\n<td>0.759</td>\n<td>0.97</td>\n<td>0.788</td>\n<td>MelSpec</td>\n<td></td>\n<td>Mixup</td>\n<td></td>\n<td>Yes</td>\n<td>17</td>\n<td>5 R</td>\n</tr>\n<tr>\n<td></td>\n<td>1</td>\n<td>tf_efficientnet_b0</td>\n<td>0.768</td>\n<td>0.97</td>\n<td>0.776</td>\n<td>MelSpec</td>\n<td></td>\n<td>Mixup</td>\n<td></td>\n<td>Yes</td>\n<td>20</td>\n<td>10 R</td>\n</tr>\n<tr>\n<td></td>\n<td>1</td>\n<td>tf_efficientnet_b0</td>\n<td>0.74</td>\n<td>0.97</td>\n<td>0.811</td>\n<td>MelSpec</td>\n<td></td>\n<td>Mixup</td>\n<td></td>\n<td>Yes</td>\n<td>15</td>\n<td>5 R</td>\n</tr>\n<tr>\n<td></td>\n<td>1</td>\n<td>FCAN B0</td>\n<td>0.75</td>\n<td>0.97</td>\n<td>0.80</td>\n<td>MelSpec</td>\n<td></td>\n<td>Mixup</td>\n<td></td>\n<td>Yes</td>\n<td>15</td>\n<td>5 R</td>\n</tr>\n<tr>\n<td></td>\n<td>1</td>\n<td>FCAN prediction B0</td>\n<td>0.76</td>\n<td>0.97</td>\n<td>0.78</td>\n<td>MelSpec</td>\n<td></td>\n<td>Mixup</td>\n<td></td>\n<td>Yes</td>\n<td>15</td>\n<td>5 R</td>\n</tr>\n<tr>\n<td></td>\n<td>4+1</td>\n<td>tf_efficientnet_b0</td>\n<td>0.77</td>\n<td>0.97</td>\n<td>0.826</td>\n<td>MelSpec</td>\n<td></td>\n<td>Mixup</td>\n<td></td>\n<td>No</td>\n<td>40</td>\n<td>5 R</td>\n</tr>\n<tr>\n<td></td>\n<td>4+1</td>\n<td>tf_efficientnet_b0</td>\n<td>0.77</td>\n<td>0.97</td>\n<td>0.825</td>\n<td>MelSpec</td>\n<td></td>\n<td>Mixup</td>\n<td></td>\n<td>Yes</td>\n<td>40</td>\n<td>5 R</td>\n</tr>\n</tbody>\n</table>",
      "rawMarkdown": "CV vs LB for some of my initial experiments. \n5 R means Random 5 seconds for training, and first 5 seconds for validation. \n\n\n\n| Datasets | Fold  | Backbone              | CV mAP | CV Score | LB    | Type       | Shape                 | Aug   | Model Additional | Random TTA                  | Epochs | Chunk |\n|----------|-------|-----------------------|--------|----------|-------|------------|------------------------|-------|------------------|---------------------------|--------|--------|\n| 2025     | 1     | hgnet_v2_b0           | 0.7    |          | 0.633 | RAW Signal | 2 x 160 x -1           |       |                  |                           | 30     | 5 R    |\n|          | 1     |                       |        |          | 0.648 | RAW Signal | 1 x 160 x -1 (3 channels) |       |                  |                           | 40     | 5 R    |\n|          | 1     | tf_efficientnet_b0    | 0.62   | 0.9529   | 0.697 | RAW Signal | 2 x -1 x 80            | Cutmix |                  |  Yes                      | 15     | 5 R    |\n|          | 1     | tf_efficientnet_b0    | 0.759  | 0.97     | 0.768 | MelSpec    |                        | Mixup |                  |                           | 17     | 5 R    |\n|          | 1     | tf_efficientnet_b0    | 0.759  | 0.97     | 0.788 | MelSpec    |                        | Mixup |                  |     Yes                   | 17     | 5 R    |\n|          | 1     | tf_efficientnet_b0    | 0.768  | 0.97     | 0.776 | MelSpec    |                        | Mixup |                  |     Yes                   | 20     | 10 R    |\n|          | 1     | tf_efficientnet_b0    | 0.74  | 0.97     | 0.811 | MelSpec    |                        | Mixup |                  |     Yes                   | 15     | 5 R    |\n|          | 1     | FCAN B0                  | 0.75  | 0.97     | 0.80 | MelSpec    |                        | Mixup |                  |     Yes                   | 15     | 5 R    |\n|          | 1     | FCAN prediction B0| 0.76  | 0.97     | 0.78 | MelSpec    |                        | Mixup |                  |     Yes                   | 15     | 5 R    |\n|          | 4+1   | tf_efficientnet_b0    | 0.77  | 0.97     | 0.826 | MelSpec    |                        | Mixup |                  |     No                   | 40     | 5 R    |\n|          | 4+1   | tf_efficientnet_b0    | 0.77  | 0.97     | 0.825 | MelSpec    |                        | Mixup |                  |     Yes                  | 40     | 5 R    |",
      "votes": null
    },
    {
      "id": "3150310",
      "postDate": "03/15/2025 10:40:45",
      "content": "<p>May i ask, what is CV mAP?</p>",
      "rawMarkdown": "May i ask, what is CV mAP?",
      "votes": null
    },
    {
      "id": "3150958",
      "postDate": "03/16/2025 06:06:29",
      "content": "<p>average_precision_score</p>",
      "rawMarkdown": "average_precision_score",
      "votes": null
    },
    {
      "id": "3151344",
      "postDate": "03/16/2025 15:29:38",
      "content": "<p>My out of fold: 0.987. LB: 0.815.</p>",
      "rawMarkdown": "My out of fold: 0.987. LB: 0.815.",
      "votes": null
    },
    {
      "id": "3151440",
      "postDate": "03/16/2025 17:29:41",
      "content": "<p>Thanks for posting. How did you compute CV? Some with-held portion of train_audio? </p>",
      "rawMarkdown": "Thanks for posting. How did you compute CV? Some with-held portion of train_audio?",
      "votes": null
    },
    {
      "id": "3151442",
      "postDate": "03/16/2025 17:30:29",
      "content": "<p>Random TTA = test time augmentation? Did you augment the audio, or just spectrograms? </p>",
      "rawMarkdown": "Random TTA = test time augmentation? Did you augment the audio, or just spectrograms?",
      "votes": null
    },
    {
      "id": "3151453",
      "postDate": "03/16/2025 17:53:15",
      "content": "<p><a href=\"https://www.kaggle.com/salmanahmedtamu\" target=\"_blank\">@salmanahmedtamu</a> no of epochs used to train</p>",
      "rawMarkdown": "salmanahmedtamu no of epochs used to train",
      "votes": null
    },
    {
      "id": "3151557",
      "postDate": "03/16/2025 21:02:03",
      "content": "<p>please cna you explain mixup and what is the loss function used <a href=\"https://www.kaggle.com/salmanahmedtamu\" target=\"_blank\">@salmanahmedtamu</a> </p>",
      "rawMarkdown": "please cna you explain mixup and what is the loss function used @salmanahmedtamu",
      "votes": null
    },
    {
      "id": "3151582",
      "postDate": "03/16/2025 22:11:07",
      "content": "<p>simple BCE, Mixup -&gt; <a href=\"https://arxiv.org/abs/1710.09412\" target=\"_blank\">https://arxiv.org/abs/1710.09412</a></p>",
      "rawMarkdown": "simple BCE, Mixup -> [https://arxiv.org/abs/1710.09412](https://arxiv.org/abs/1710.09412)",
      "votes": null
    },
    {
      "id": "3151583",
      "postDate": "03/16/2025 22:11:33",
      "content": "<p>Simple Stratified Fold, </p>",
      "rawMarkdown": "Simple Stratified Fold,",
      "votes": null
    },
    {
      "id": "3151584",
      "postDate": "03/16/2025 22:11:51",
      "content": "<p>Around 20 Epochs</p>",
      "rawMarkdown": "Around 20 Epochs",
      "votes": null
    },
    {
      "id": "3151861",
      "postDate": "03/17/2025 06:57:51",
      "content": "<p>Did you face nan issues, while training? <a href=\"https://www.kaggle.com/salmanahmedtamu\" target=\"_blank\">@salmanahmedtamu</a></p>",
      "rawMarkdown": "Did you face nan issues, while training? @salmanahmedtamu",
      "votes": null
    },
    {
      "id": "3151997",
      "postDate": "03/17/2025 10:12:01",
      "content": "<p>OOF: 0.9898. LB: 0.820</p>",
      "rawMarkdown": "OOF: 0.9898. LB: 0.820",
      "votes": null
    },
    {
      "id": "3152026",
      "postDate": "03/17/2025 10:42:02",
      "content": "<p>Thank you for your sharing!<br>\nmy CV score is 0.956, LB score is 0.813.</p>",
      "rawMarkdown": "Thank you for your sharing!\nmy CV score is 0.956, LB score is 0.813.",
      "votes": null
    },
    {
      "id": "3152916",
      "postDate": "03/18/2025 09:02:27",
      "content": "<p><a href=\"https://www.kaggle.com/salmanahmedtamu\" target=\"_blank\">@salmanahmedtamu</a> may i ask about the difference between last experiment and 0.788 LB experiment is there any difference except epoch number. and thank you so much for your patience  </p>",
      "rawMarkdown": "salmanahmedtamu may i ask about the difference between last experiment and 0.788 LB experiment is there any difference except epoch number. and thank you so much for your patience",
      "votes": null
    },
    {
      "id": "3153120",
      "postDate": "03/18/2025 12:48:18",
      "content": "<p>just epochs. </p>",
      "rawMarkdown": "just epochs.",
      "votes": null
    },
    {
      "id": "3153469",
      "postDate": "03/18/2025 21:17:46",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/salmanahmedtamu\" target=\"_blank\">@salmanahmedtamu</a> , I implemented your approach from previous competition everything similar to last experiment even all metrics except CV maP it's very low around 58.4% everything else is the same but i got much lower LB score please can you tell me how Map correlate to LB score and another question what type of TTA you used for inference.<br>\nthanks again</p>",
      "rawMarkdown": "Hi @salmanahmedtamu , I implemented your approach from previous competition everything similar to last experiment even all metrics except CV maP it's very low around 58.4% everything else is the same but i got much lower LB score please can you tell me how Map correlate to LB score and another question what type of TTA you used for inference.\nthanks again",
      "votes": null
    },
    {
      "id": "3154087",
      "postDate": "03/19/2025 13:58:30",
      "content": "<p>mAP doesn't mean anything, you can ignore that, you might be getting low mAP because of missing classes in validation set. I don't think there is any appropriate way to validate the model, specially not by split the provided data in train and test.</p>",
      "rawMarkdown": "mAP doesn't mean anything, you can ignore that, you might be getting low mAP because of missing classes in validation set. I don't think there is any appropriate way to validate the model, specially not by split the provided data in train and test.",
      "votes": null
    },
    {
      "id": "3154088",
      "postDate": "03/19/2025 13:59:21",
      "content": "<p>Not that kind of TTA, its TTA on predictions.</p>",
      "rawMarkdown": "Not that kind of TTA, its TTA on predictions.",
      "votes": null
    },
    {
      "id": "3154119",
      "postDate": "03/19/2025 14:40:01",
      "content": "<p>i tried the basic script but score is low now i split the data validation is first 5 seconds and from 6th second i crop random 5 seconds to minimize the overlapping between train and validation samples, may i ask what kind of TTA enhance LB i tried hfplip, vflip but no significant enhance</p>",
      "rawMarkdown": "i tried the basic script but score is low now i split the data validation is first 5 seconds and from 6th second i crop random 5 seconds to minimize the overlapping between train and validation samples, may i ask what kind of TTA enhance LB i tried hfplip, vflip but no significant enhance",
      "votes": null
    },
    {
      "id": "3154127",
      "postDate": "03/19/2025 15:03:22",
      "content": "<p>It's TTA on predictions. You can see solutions from past competitions they did this as well.</p>",
      "rawMarkdown": "It's TTA on predictions. You can see solutions from past competitions they did this as well.",
      "votes": null
    },
    {
      "id": "3155915",
      "postDate": "03/21/2025 14:18:44",
      "content": "<p>No, if there are NaN issues, I would suggest to look at the gradients or grad scaler if you are using one. <br>\nMaybe you can try to clip while training.</p>",
      "rawMarkdown": "No, if there are NaN issues, I would suggest to look at the gradients or grad scaler if you are using one. \nMaybe you can try to clip while training.",
      "votes": null
    },
    {
      "id": "3161292",
      "postDate": "03/27/2025 17:56:20",
      "content": "<p>similar to me 0.81X for LB with efficientnet B1. CV is 0.989 on auc</p>",
      "rawMarkdown": "similar to me 0.81X for LB with efficientnet B1. CV is 0.989 on auc",
      "votes": null
    },
    {
      "id": "3161304",
      "postDate": "03/27/2025 18:07:59",
      "content": "<p>Have you experimented with resizing the melSpecs to image net dimensions or increasing the resolution of MelSpecs while also increasing n_ffts and hops? </p>",
      "rawMarkdown": "Have you experimented with resizing the melSpecs to image net dimensions or increasing the resolution of MelSpecs while also increasing n_ffts and hops?",
      "votes": null
    },
    {
      "id": "3162383",
      "postDate": "03/29/2025 06:23:55",
      "content": "<p>If you encounter the \"nan\" issue, it might be because the number of data categories is not 206. I previously had this problem when I reduced the number of data categories.</p>",
      "rawMarkdown": "If you encounter the \"nan\" issue, it might be because the number of data categories is not 206. I previously had this problem when I reduced the number of data categories.",
      "votes": null
    },
    {
      "id": "3163027",
      "postDate": "03/30/2025 07:45:48",
      "content": "<p>How you decide which epoch to stop training at? In my experiments, the CV score just keeps going up to 0.9999…</p>",
      "rawMarkdown": "How you decide which epoch to stop training at? In my experiments, the CV score just keeps going up to 0.9999...",
      "votes": null
    },
    {
      "id": "3163403",
      "postDate": "03/30/2025 18:24:31",
      "content": "<p>Thanks for sharing!</p>\n<p>My cv score is around .95x lb .828</p>",
      "rawMarkdown": "Thanks for sharing!\n\nMy cv score is around .95x lb .828",
      "votes": null
    },
    {
      "id": "3171652",
      "postDate": "04/05/2025 23:01:41",
      "content": "<p>Hi, thanks for sharing! I was wondering are all these runs based on the same training setup? Based on the table, I noticed your LB score improved from 0.788 to 0.811, and the only difference seems to be switching from epoch 17 to epoch 15. Just curious if you’re using different checkpoints from the same training process? Thank you!</p>",
      "rawMarkdown": "Hi, thanks for sharing! I was wondering are all these runs based on the same training setup? Based on the table, I noticed your LB score improved from 0.788 to 0.811, and the only difference seems to be switching from epoch 17 to epoch 15. Just curious if you’re using different checkpoints from the same training process? Thank you!",
      "votes": null
    },
    {
      "id": "3171672",
      "postDate": "04/06/2025 00:24:16",
      "content": "<p>I believe, some classes that have less than 10 labels impact a lot on LB, specially if the score is more than 0.8<br>\nSo, in this case, it really matters which combinations of labels are being passed into each batch and if you have mixup then which combinations are actually being passed to model in the mixup.<br>\nThat's why performance varies epoch to epoch, so, imagine in Epoch K, model gave more importance to class E than F because E had highest loss in those specific batches it was passed during training of that epoch.</p>\n<p>So, I think just by taking average of multiple checkpoints (almost last 1/3 of the training process actually helps is stabilizing it. So if I am training for 100 epochs, I will take average of last 33 epochs.) </p>\n<p>But I don't think, this is a proper solution, but in this case if you try LB probing and think about customizing batches, then you might overfit to the public LB.</p>",
      "rawMarkdown": "I believe, some classes that have less than 10 labels impact a lot on LB, specially if the score is more than 0.8\nSo, in this case, it really matters which combinations of labels are being passed into each batch and if you have mixup then which combinations are actually being passed to model in the mixup.\nThat's why performance varies epoch to epoch, so, imagine in Epoch K, model gave more importance to class E than F because E had highest loss in those specific batches it was passed during training of that epoch.\n\nSo, I think just by taking average of multiple checkpoints (almost last 1/3 of the training process actually helps is stabilizing it. So if I am training for 100 epochs, I will take average of last 33 epochs.) \n\nBut I don't think, this is a proper solution, but in this case if you try LB probing and think about customizing batches, then you might overfit to the public LB.",
      "votes": null
    },
    {
      "id": "3183319",
      "postDate": "04/20/2025 18:33:00",
      "content": "<p>Thanks a lot for sharing your early results, may I ask how are you splitting your data into 5 folds?<br>\nGiven that there can be classes with frequency of only two.</p>",
      "rawMarkdown": "Thanks a lot for sharing your early results, may I ask how are you splitting your data into 5 folds?\nGiven that there can be classes with frequency of only two.",
      "votes": null
    },
    {
      "id": "3183322",
      "postDate": "04/20/2025 18:39:12",
      "content": "<p>I'll try 30 epochs now with regnety_008 and effecientnetb0 ensemble</p>",
      "rawMarkdown": "I'll try 30 epochs now with regnety_008 and effecientnetb0 ensemble",
      "votes": null
    },
    {
      "id": "3190151",
      "postDate": "04/30/2025 07:32:33",
      "content": "<p><a href=\"https://www.kaggle.com/salmanahmedtamu\" target=\"_blank\">@salmanahmedtamu</a><br>\nThanks for sharing! I’ve also noticed that the public score can vary depending on the batch size. I’m still unsure whether it’s better to go with a larger or smaller batch size—both seem to have their pros and cons in my experiments. How are you approaching this? Would love to hear your thoughts.</p>",
      "rawMarkdown": "salmanahmedtamu\nThanks for sharing! I’ve also noticed that the public score can vary depending on the batch size. I’m still unsure whether it’s better to go with a larger or smaller batch size—both seem to have their pros and cons in my experiments. How are you approaching this? Would love to hear your thoughts.",
      "votes": null
    },
    {
      "id": "3209584",
      "postDate": "05/26/2025 04:08:56",
      "content": "<p>How's that, have you found the best batch size to process</p>",
      "rawMarkdown": "How's that, have you found the best batch size to process",
      "votes": null
    },
    {
      "id": "3213392",
      "postDate": "05/29/2025 23:43:19",
      "content": "<p><a href=\"https://www.kaggle.com/salmanahmedtamu\" target=\"_blank\">@salmanahmedtamu</a> I did the same thing the training (100 epochs and 5 Folds) with Early stop to avoid overfitting and the average was 30 (mean CV of 0.97 for tf_effiencient_b0) However 0.8 ~ 0.82 on the LB.</p>",
      "rawMarkdown": "salmanahmedtamu I did the same thing the training (100 epochs and 5 Folds) with Early stop to avoid overfitting and the average was 30 (mean CV of 0.97 for tf_effiencient_b0) However 0.8 ~ 0.82 on the LB.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3150310,
      "author_name": "eikyou",
      "author_url": "",
      "post_date": "03/15/2025 10:40:45",
      "content": "<p>May i ask, what is CV mAP?</p>",
      "votes": null,
      "replies": [
        {
          "id": 3150958,
          "author_name": "salmanahmedtamu",
          "author_url": "",
          "post_date": "03/16/2025 06:06:29",
          "content": "<p>average_precision_score</p>",
          "votes": null,
          "replies": [
            {
              "id": 3151453,
              "author_name": "arunodhayan",
              "author_url": "",
              "post_date": "03/16/2025 17:53:15",
              "content": "<p><a href=\"https://www.kaggle.com/salmanahmedtamu\" target=\"_blank\">@salmanahmedtamu</a> no of epochs used to train</p>",
              "votes": null,
              "replies": [
                {
                  "id": 3151584,
                  "author_name": "salmanahmedtamu",
                  "author_url": "",
                  "post_date": "03/16/2025 22:11:51",
                  "content": "<p>Around 20 Epochs</p>",
                  "votes": null,
                  "replies": []
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 3151344,
      "author_name": "quan0095",
      "author_url": "",
      "post_date": "03/16/2025 15:29:38",
      "content": "<p>My out of fold: 0.987. LB: 0.815.</p>",
      "votes": null,
      "replies": [
        {
          "id": 3151997,
          "author_name": "quan0095",
          "author_url": "",
          "post_date": "03/17/2025 10:12:01",
          "content": "<p>OOF: 0.9898. LB: 0.820</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3151440,
      "author_name": "robbynevels",
      "author_url": "",
      "post_date": "03/16/2025 17:29:41",
      "content": "<p>Thanks for posting. How did you compute CV? Some with-held portion of train_audio? </p>",
      "votes": null,
      "replies": [
        {
          "id": 3151583,
          "author_name": "salmanahmedtamu",
          "author_url": "",
          "post_date": "03/16/2025 22:11:33",
          "content": "<p>Simple Stratified Fold, </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3151442,
      "author_name": "robbynevels",
      "author_url": "",
      "post_date": "03/16/2025 17:30:29",
      "content": "<p>Random TTA = test time augmentation? Did you augment the audio, or just spectrograms? </p>",
      "votes": null,
      "replies": [
        {
          "id": 3154088,
          "author_name": "salmanahmedtamu",
          "author_url": "",
          "post_date": "03/19/2025 13:59:21",
          "content": "<p>Not that kind of TTA, its TTA on predictions.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3151557,
      "author_name": "mohamed3abdelrazik",
      "author_url": "",
      "post_date": "03/16/2025 21:02:03",
      "content": "<p>please cna you explain mixup and what is the loss function used <a href=\"https://www.kaggle.com/salmanahmedtamu\" target=\"_blank\">@salmanahmedtamu</a> </p>",
      "votes": null,
      "replies": [
        {
          "id": 3151582,
          "author_name": "salmanahmedtamu",
          "author_url": "",
          "post_date": "03/16/2025 22:11:07",
          "content": "<p>simple BCE, Mixup -&gt; <a href=\"https://arxiv.org/abs/1710.09412\" target=\"_blank\">https://arxiv.org/abs/1710.09412</a></p>",
          "votes": null,
          "replies": [
            {
              "id": 3151861,
              "author_name": "seshurajup",
              "author_url": "",
              "post_date": "03/17/2025 06:57:51",
              "content": "<p>Did you face nan issues, while training? <a href=\"https://www.kaggle.com/salmanahmedtamu\" target=\"_blank\">@salmanahmedtamu</a></p>",
              "votes": null,
              "replies": [
                {
                  "id": 3155915,
                  "author_name": "salmanahmedtamu",
                  "author_url": "",
                  "post_date": "03/21/2025 14:18:44",
                  "content": "<p>No, if there are NaN issues, I would suggest to look at the gradients or grad scaler if you are using one. <br>\nMaybe you can try to clip while training.</p>",
                  "votes": null,
                  "replies": []
                },
                {
                  "id": 3162383,
                  "author_name": "liuxiaoonline",
                  "author_url": "",
                  "post_date": "03/29/2025 06:23:55",
                  "content": "<p>If you encounter the \"nan\" issue, it might be because the number of data categories is not 206. I previously had this problem when I reduced the number of data categories.</p>",
                  "votes": null,
                  "replies": []
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 3152026,
      "author_name": "kmatsu01",
      "author_url": "",
      "post_date": "03/17/2025 10:42:02",
      "content": "<p>Thank you for your sharing!<br>\nmy CV score is 0.956, LB score is 0.813.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3152916,
      "author_name": "mohamed3abdelrazik",
      "author_url": "",
      "post_date": "03/18/2025 09:02:27",
      "content": "<p><a href=\"https://www.kaggle.com/salmanahmedtamu\" target=\"_blank\">@salmanahmedtamu</a> may i ask about the difference between last experiment and 0.788 LB experiment is there any difference except epoch number. and thank you so much for your patience  </p>",
      "votes": null,
      "replies": [
        {
          "id": 3153120,
          "author_name": "salmanahmedtamu",
          "author_url": "",
          "post_date": "03/18/2025 12:48:18",
          "content": "<p>just epochs. </p>",
          "votes": null,
          "replies": [
            {
              "id": 3153469,
              "author_name": "mohamed3abdelrazik",
              "author_url": "",
              "post_date": "03/18/2025 21:17:46",
              "content": "<p>Hi <a href=\"https://www.kaggle.com/salmanahmedtamu\" target=\"_blank\">@salmanahmedtamu</a> , I implemented your approach from previous competition everything similar to last experiment even all metrics except CV maP it's very low around 58.4% everything else is the same but i got much lower LB score please can you tell me how Map correlate to LB score and another question what type of TTA you used for inference.<br>\nthanks again</p>",
              "votes": null,
              "replies": [
                {
                  "id": 3154087,
                  "author_name": "salmanahmedtamu",
                  "author_url": "",
                  "post_date": "03/19/2025 13:58:30",
                  "content": "<p>mAP doesn't mean anything, you can ignore that, you might be getting low mAP because of missing classes in validation set. I don't think there is any appropriate way to validate the model, specially not by split the provided data in train and test.</p>",
                  "votes": null,
                  "replies": [
                    {
                      "id": 3154119,
                      "author_name": "mohamed3abdelrazik",
                      "author_url": "",
                      "post_date": "03/19/2025 14:40:01",
                      "content": "<p>i tried the basic script but score is low now i split the data validation is first 5 seconds and from 6th second i crop random 5 seconds to minimize the overlapping between train and validation samples, may i ask what kind of TTA enhance LB i tried hfplip, vflip but no significant enhance</p>",
                      "votes": null,
                      "replies": [
                        {
                          "id": 3154127,
                          "author_name": "salmanahmedtamu",
                          "author_url": "",
                          "post_date": "03/19/2025 15:03:22",
                          "content": "<p>It's TTA on predictions. You can see solutions from past competitions they did this as well.</p>",
                          "votes": null,
                          "replies": []
                        }
                      ]
                    }
                  ]
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 3161292,
      "author_name": "ludovick",
      "author_url": "",
      "post_date": "03/27/2025 17:56:20",
      "content": "<p>similar to me 0.81X for LB with efficientnet B1. CV is 0.989 on auc</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3161304,
      "author_name": "hotsonhonet",
      "author_url": "",
      "post_date": "03/27/2025 18:07:59",
      "content": "<p>Have you experimented with resizing the melSpecs to image net dimensions or increasing the resolution of MelSpecs while also increasing n_ffts and hops? </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3163027,
      "author_name": "digiranger",
      "author_url": "",
      "post_date": "03/30/2025 07:45:48",
      "content": "<p>How you decide which epoch to stop training at? In my experiments, the CV score just keeps going up to 0.9999…</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3163403,
      "author_name": "aman1391",
      "author_url": "",
      "post_date": "03/30/2025 18:24:31",
      "content": "<p>Thanks for sharing!</p>\n<p>My cv score is around .95x lb .828</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3171652,
      "author_name": "leonshangguan",
      "author_url": "",
      "post_date": "04/05/2025 23:01:41",
      "content": "<p>Hi, thanks for sharing! I was wondering are all these runs based on the same training setup? Based on the table, I noticed your LB score improved from 0.788 to 0.811, and the only difference seems to be switching from epoch 17 to epoch 15. Just curious if you’re using different checkpoints from the same training process? Thank you!</p>",
      "votes": null,
      "replies": [
        {
          "id": 3171672,
          "author_name": "salmanahmedtamu",
          "author_url": "",
          "post_date": "04/06/2025 00:24:16",
          "content": "<p>I believe, some classes that have less than 10 labels impact a lot on LB, specially if the score is more than 0.8<br>\nSo, in this case, it really matters which combinations of labels are being passed into each batch and if you have mixup then which combinations are actually being passed to model in the mixup.<br>\nThat's why performance varies epoch to epoch, so, imagine in Epoch K, model gave more importance to class E than F because E had highest loss in those specific batches it was passed during training of that epoch.</p>\n<p>So, I think just by taking average of multiple checkpoints (almost last 1/3 of the training process actually helps is stabilizing it. So if I am training for 100 epochs, I will take average of last 33 epochs.) </p>\n<p>But I don't think, this is a proper solution, but in this case if you try LB probing and think about customizing batches, then you might overfit to the public LB.</p>",
          "votes": null,
          "replies": [
            {
              "id": 3213392,
              "author_name": "idrissamdicko",
              "author_url": "",
              "post_date": "05/29/2025 23:43:19",
              "content": "<p><a href=\"https://www.kaggle.com/salmanahmedtamu\" target=\"_blank\">@salmanahmedtamu</a> I did the same thing the training (100 epochs and 5 Folds) with Early stop to avoid overfitting and the average was 30 (mean CV of 0.97 for tf_effiencient_b0) However 0.8 ~ 0.82 on the LB.</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3183319,
      "author_name": "chaudharypriyanshu",
      "author_url": "",
      "post_date": "04/20/2025 18:33:00",
      "content": "<p>Thanks a lot for sharing your early results, may I ask how are you splitting your data into 5 folds?<br>\nGiven that there can be classes with frequency of only two.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3183322,
      "author_name": "devasypatel23",
      "author_url": "",
      "post_date": "04/20/2025 18:39:12",
      "content": "<p>I'll try 30 epochs now with regnety_008 and effecientnetb0 ensemble</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3190151,
      "author_name": "kmatsu01",
      "author_url": "",
      "post_date": "04/30/2025 07:32:33",
      "content": "<p><a href=\"https://www.kaggle.com/salmanahmedtamu\" target=\"_blank\">@salmanahmedtamu</a><br>\nThanks for sharing! I’ve also noticed that the public score can vary depending on the batch size. I’m still unsure whether it’s better to go with a larger or smaller batch size—both seem to have their pros and cons in my experiments. How are you approaching this? Would love to hear your thoughts.</p>",
      "votes": null,
      "replies": [
        {
          "id": 3209584,
          "author_name": "xiayuxuan",
          "author_url": "",
          "post_date": "05/26/2025 04:08:56",
          "content": "<p>How's that, have you found the best batch size to process</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3150065": "CV vs LB for some of my initial experiments. \n5 R means Random 5 seconds for training, and first 5 seconds for validation. \n\n\n\n| Datasets | Fold  | Backbone              | CV mAP | CV Score | LB    | Type       | Shape                 | Aug   | Model Additional | Random TTA                  | Epochs | Chunk |\n|----------|-------|-----------------------|--------|----------|-------|------------|------------------------|-------|------------------|---------------------------|--------|--------|\n| 2025     | 1     | hgnet_v2_b0           | 0.7    |          | 0.633 | RAW Signal | 2 x 160 x -1           |       |                  |                           | 30     | 5 R    |\n|          | 1     |                       |        |          | 0.648 | RAW Signal | 1 x 160 x -1 (3 channels) |       |                  |                           | 40     | 5 R    |\n|          | 1     | tf_efficientnet_b0    | 0.62   | 0.9529   | 0.697 | RAW Signal | 2 x -1 x 80            | Cutmix |                  |  Yes                      | 15     | 5 R    |\n|          | 1     | tf_efficientnet_b0    | 0.759  | 0.97     | 0.768 | MelSpec    |                        | Mixup |                  |                           | 17     | 5 R    |\n|          | 1     | tf_efficientnet_b0    | 0.759  | 0.97     | 0.788 | MelSpec    |                        | Mixup |                  |     Yes                   | 17     | 5 R    |\n|          | 1     | tf_efficientnet_b0    | 0.768  | 0.97     | 0.776 | MelSpec    |                        | Mixup |                  |     Yes                   | 20     | 10 R    |\n|          | 1     | tf_efficientnet_b0    | 0.74  | 0.97     | 0.811 | MelSpec    |                        | Mixup |                  |     Yes                   | 15     | 5 R    |\n|          | 1     | FCAN B0                  | 0.75  | 0.97     | 0.80 | MelSpec    |                        | Mixup |                  |     Yes                   | 15     | 5 R    |\n|          | 1     | FCAN prediction B0| 0.76  | 0.97     | 0.78 | MelSpec    |                        | Mixup |                  |     Yes                   | 15     | 5 R    |\n|          | 4+1   | tf_efficientnet_b0    | 0.77  | 0.97     | 0.826 | MelSpec    |                        | Mixup |                  |     No                   | 40     | 5 R    |\n|          | 4+1   | tf_efficientnet_b0    | 0.77  | 0.97     | 0.825 | MelSpec    |                        | Mixup |                  |     Yes                  | 40     | 5 R    |",
    "3150310": "May i ask, what is CV mAP?",
    "3150958": "average_precision_score",
    "3151344": "My out of fold: 0.987. LB: 0.815.",
    "3151440": "Thanks for posting. How did you compute CV? Some with-held portion of train_audio?",
    "3151442": "Random TTA = test time augmentation? Did you augment the audio, or just spectrograms?",
    "3151453": "salmanahmedtamu no of epochs used to train",
    "3151557": "please cna you explain mixup and what is the loss function used @salmanahmedtamu",
    "3151582": "simple BCE, Mixup -> [https://arxiv.org/abs/1710.09412](https://arxiv.org/abs/1710.09412)",
    "3151583": "Simple Stratified Fold,",
    "3151584": "Around 20 Epochs",
    "3151861": "Did you face nan issues, while training? @salmanahmedtamu",
    "3151997": "OOF: 0.9898. LB: 0.820",
    "3152026": "Thank you for your sharing!\nmy CV score is 0.956, LB score is 0.813.",
    "3152916": "salmanahmedtamu may i ask about the difference between last experiment and 0.788 LB experiment is there any difference except epoch number. and thank you so much for your patience",
    "3153120": "just epochs.",
    "3153469": "Hi @salmanahmedtamu , I implemented your approach from previous competition everything similar to last experiment even all metrics except CV maP it's very low around 58.4% everything else is the same but i got much lower LB score please can you tell me how Map correlate to LB score and another question what type of TTA you used for inference.\nthanks again",
    "3154087": "mAP doesn't mean anything, you can ignore that, you might be getting low mAP because of missing classes in validation set. I don't think there is any appropriate way to validate the model, specially not by split the provided data in train and test.",
    "3154088": "Not that kind of TTA, its TTA on predictions.",
    "3154119": "i tried the basic script but score is low now i split the data validation is first 5 seconds and from 6th second i crop random 5 seconds to minimize the overlapping between train and validation samples, may i ask what kind of TTA enhance LB i tried hfplip, vflip but no significant enhance",
    "3154127": "It's TTA on predictions. You can see solutions from past competitions they did this as well.",
    "3155915": "No, if there are NaN issues, I would suggest to look at the gradients or grad scaler if you are using one. \nMaybe you can try to clip while training.",
    "3161292": "similar to me 0.81X for LB with efficientnet B1. CV is 0.989 on auc",
    "3161304": "Have you experimented with resizing the melSpecs to image net dimensions or increasing the resolution of MelSpecs while also increasing n_ffts and hops?",
    "3162383": "If you encounter the \"nan\" issue, it might be because the number of data categories is not 206. I previously had this problem when I reduced the number of data categories.",
    "3163027": "How you decide which epoch to stop training at? In my experiments, the CV score just keeps going up to 0.9999...",
    "3163403": "Thanks for sharing!\n\nMy cv score is around .95x lb .828",
    "3171652": "Hi, thanks for sharing! I was wondering are all these runs based on the same training setup? Based on the table, I noticed your LB score improved from 0.788 to 0.811, and the only difference seems to be switching from epoch 17 to epoch 15. Just curious if you’re using different checkpoints from the same training process? Thank you!",
    "3171672": "I believe, some classes that have less than 10 labels impact a lot on LB, specially if the score is more than 0.8\nSo, in this case, it really matters which combinations of labels are being passed into each batch and if you have mixup then which combinations are actually being passed to model in the mixup.\nThat's why performance varies epoch to epoch, so, imagine in Epoch K, model gave more importance to class E than F because E had highest loss in those specific batches it was passed during training of that epoch.\n\nSo, I think just by taking average of multiple checkpoints (almost last 1/3 of the training process actually helps is stabilizing it. So if I am training for 100 epochs, I will take average of last 33 epochs.) \n\nBut I don't think, this is a proper solution, but in this case if you try LB probing and think about customizing batches, then you might overfit to the public LB.",
    "3183319": "Thanks a lot for sharing your early results, may I ask how are you splitting your data into 5 folds?\nGiven that there can be classes with frequency of only two.",
    "3183322": "I'll try 30 epochs now with regnety_008 and effecientnetb0 ensemble",
    "3190151": "salmanahmedtamu\nThanks for sharing! I’ve also noticed that the public score can vary depending on the batch size. I’m still unsure whether it’s better to go with a larger or smaller batch size—both seem to have their pros and cons in my experiments. How are you approaching this? Would love to hear your thoughts.",
    "3209584": "How's that, have you found the best batch size to process",
    "3213392": "salmanahmedtamu I did the same thing the training (100 epochs and 5 Folds) with Early stop to avoid overfitting and the average was 30 (mean CV of 0.97 for tf_effiencient_b0) However 0.8 ~ 0.82 on the LB."
  },
  "source": "meta"
}