{
  "id": 573066,
  "title": "Recipe to Public LB 0.872",
  "url": "/competitions/birdclef-2025/discussion/573066",
  "author_name": "Salman Ahmed",
  "post_date": "2025-04-13T10:19:32.400000",
  "votes": 119,
  "comment_count": 88,
  "views": 0,
  "content": "<p>I am sharing what worked for me in early baseline, but still haven't figured out what's working yet, still trying. </p>\n<ol>\n<li>Applying any augmentation or processing on raw waves, hurts the performance.</li>\n<li>When I look at last year's leaderboard, position # 5 was able to achieve a awesome melspec model without any pseudo labels or anything, while other top solutions used pseudo labels. I believe melspec parameters were one of the import things in that solution. (Not sure.) Considering this, I started exploring and I was able to get scores from 0.810 to 0.859 just by changing the melspec parameters. You just have to think about it, visualize and explore more. (5 subs a day to find the best parameters).</li>\n<li>I've tried BCE with multiple types of weights, but FocalBCE works best for me.</li>\n<li>Consider melspecs as images and not the waves, I converted the melspecs to images, applied normalization and augment with RandAug and RandomErazing, Time and Freq Masking + (Mixup with prob of 1.0). Remaining training pipeline is same as my public notebook from last year. <a href=\"https://www.kaggle.com/code/salmanahmedtamu/training-0-65-0-66\" target=\"_blank\">https://www.kaggle.com/code/salmanahmedtamu/training-0-65-0-66</a></li>\n<li>Another important thing, in Efficientnet models when i combined middle features and passed into the global pool that worked better than the conv_head in timm. Still Exploring…..!!!!</li>\n<li>Even simple models starts overfitting pretty easily, and you cannot track because K-Fold doesn't make sense here, your validation set is the public LB. Keep track of the Batch Size, Number of iterations, warmp up and LR. </li>\n<li>Filtering/Processing CSA recordings by Human sound helps.</li>\n<li>Look at predictions of your model and visualize the attentions on Train Soundscape, to see what's the CAM of your model.</li>\n<li>Ensemble to 0.854, 0.856, 0.858 and 0.859 results in 0.872. </li>\n<li>Post Processing is important as I shared in my previous public notebook, There are a lot of other ways to post process as well.</li>\n<li>I am sure, I missed a lot of detail, I'll try to share some public notebook if i get some time this week.</li>\n<li>Training on Random 5 seconds helped compared to first 5 seconds.</li>\n</ol>",
  "messages": [
    {
      "id": 3177838,
      "postDate": "2025-04-13T10:19:32.400Z",
      "content": "<p>I am sharing what worked for me in early baseline, but still haven't figured out what's working yet, still trying. </p>\n<ol>\n<li>Applying any augmentation or processing on raw waves, hurts the performance.</li>\n<li>When I look at last year's leaderboard, position # 5 was able to achieve a awesome melspec model without any pseudo labels or anything, while other top solutions used pseudo labels. I believe melspec parameters were one of the import things in that solution. (Not sure.) Considering this, I started exploring and I was able to get scores from 0.810 to 0.859 just by changing the melspec parameters. You just have to think about it, visualize and explore more. (5 subs a day to find the best parameters).</li>\n<li>I've tried BCE with multiple types of weights, but FocalBCE works best for me.</li>\n<li>Consider melspecs as images and not the waves, I converted the melspecs to images, applied normalization and augment with RandAug and RandomErazing, Time and Freq Masking + (Mixup with prob of 1.0). Remaining training pipeline is same as my public notebook from last year. <a href=\"https://www.kaggle.com/code/salmanahmedtamu/training-0-65-0-66\" target=\"_blank\">https://www.kaggle.com/code/salmanahmedtamu/training-0-65-0-66</a></li>\n<li>Another important thing, in Efficientnet models when i combined middle features and passed into the global pool that worked better than the conv_head in timm. Still Exploring…..!!!!</li>\n<li>Even simple models starts overfitting pretty easily, and you cannot track because K-Fold doesn't make sense here, your validation set is the public LB. Keep track of the Batch Size, Number of iterations, warmp up and LR. </li>\n<li>Filtering/Processing CSA recordings by Human sound helps.</li>\n<li>Look at predictions of your model and visualize the attentions on Train Soundscape, to see what's the CAM of your model.</li>\n<li>Ensemble to 0.854, 0.856, 0.858 and 0.859 results in 0.872. </li>\n<li>Post Processing is important as I shared in my previous public notebook, There are a lot of other ways to post process as well.</li>\n<li>I am sure, I missed a lot of detail, I'll try to share some public notebook if i get some time this week.</li>\n<li>Training on Random 5 seconds helped compared to first 5 seconds.</li>\n</ol>",
      "rawMarkdown": "I am sharing what worked for me in early baseline, but still haven't figured out what's working yet, still trying. \n\n1. Applying any augmentation or processing on raw waves, hurts the performance.\n2. When I look at last year's leaderboard, position # 5 was able to achieve a awesome melspec model without any pseudo labels or anything, while other top solutions used pseudo labels. I believe melspec parameters were one of the import things in that solution. (Not sure.) Considering this, I started exploring and I was able to get scores from 0.810 to 0.859 just by changing the melspec parameters. You just have to think about it, visualize and explore more. (5 subs a day to find the best parameters).\n3. I've tried BCE with multiple types of weights, but FocalBCE works best for me.\n4. Consider melspecs as images and not the waves, I converted the melspecs to images, applied normalization and augment with RandAug and RandomErazing, Time and Freq Masking + (Mixup with prob of 1.0). Remaining training pipeline is same as my public notebook from last year. [https://www.kaggle.com/code/salmanahmedtamu/training-0-65-0-66](https://www.kaggle.com/code/salmanahmedtamu/training-0-65-0-66)\n5. Another important thing, in Efficientnet models when i combined middle features and passed into the global pool that worked better than the conv_head in timm. Still Exploring.....!!!!\n6. Even simple models starts overfitting pretty easily, and you cannot track because K-Fold doesn't make sense here, your validation set is the public LB. Keep track of the Batch Size, Number of iterations, warmp up and LR. \n7. Filtering/Processing CSA recordings by Human sound helps.\n8. Look at predictions of your model and visualize the attentions on Train Soundscape, to see what's the CAM of your model.\n9. Ensemble to 0.854, 0.856, 0.858 and 0.859 results in 0.872. \n10. Post Processing is important as I shared in my previous public notebook, There are a lot of other ways to post process as well.\n11. I am sure, I missed a lot of detail, I'll try to share some public notebook if i get some time this week.\n12. Training on Random 5 seconds helped compared to first 5 seconds.",
      "votes": 119
    },
    {
      "id": 3195341,
      "postDate": "2025-05-06T23:01:50.483Z",
      "content": "<p>After applying TTA using efficientnet_B0 model, every time , notebook throws message as network threw exception, while submission on scoring. Can anyone guide to resolve this issue?</p>",
      "rawMarkdown": "After applying TTA using efficientnet_B0 model, every time , notebook throws message as network threw exception, while submission on scoring. Can anyone guide to resolve this issue?",
      "votes": 1,
      "replies": [
        {
          "id": 3195344,
          "postDate": "2025-05-06T23:12:38.450Z",
          "content": "<p>Can you share that notebook? or the code for TTA? </p>",
          "rawMarkdown": "Can you share that notebook? or the code for TTA? ",
          "replies": [
            {
              "id": 3195349,
              "postDate": "2025-05-06T23:28:23.460Z",
              "rawMarkdown": "",
              "isDeleted": true
            },
            {
              "id": 3195351,
              "postDate": "2025-05-06T23:31:36.143Z",
              "content": "<p><a href=\"https://www.kaggle.com/salmanahmedtamu\" target=\"_blank\">@salmanahmedtamu</a> , this is the attached notebook, I referred. Plz guide me. In this use_tta = False by default, After setting use_tta = true, it is threwing exception, on submission for scoring, even for minimum values.</p>",
              "rawMarkdown": "@salmanahmedtamu , this is the attached notebook, I referred. Plz guide me. In this use_tta = False by default, After setting use_tta = true, it is threwing exception, on submission for scoring, even for minimum values.",
              "votes": 1
            }
          ]
        },
        {
          "id": 3196907,
          "postDate": "2025-05-07T15:31:48.570Z",
          "content": "<p>There is an error using source data, you need to add \".copy\" when you use the apply_tta function or return directly your augmentatied mel.</p>",
          "rawMarkdown": "There is an error using source data, you need to add \".copy\" when you use the apply_tta function or return directly your augmentatied mel.",
          "votes": 1
        },
        {
          "id": 3196910,
          "postDate": "2025-05-07T15:36:36.300Z",
          "content": "<p>I think the error raise because of the numpy.</p>",
          "rawMarkdown": "I think the error raise because of the numpy.",
          "votes": 2,
          "replies": [
            {
              "id": 3198881,
              "postDate": "2025-05-10T06:26:06.267Z",
              "content": "<p>Thanks, now the notebook timed out even i used tta=3(minimum).  Should I increase HOP_LENGTH?</p>",
              "rawMarkdown": "Thanks, now the notebook timed out even i used tta=3(minimum).  Should I increase HOP_LENGTH?"
            },
            {
              "id": 3198917,
              "postDate": "2025-05-10T07:34:45.463Z",
              "content": "<p>To try light-weight model or quantitize the model (onnx or openvino), I think increase HOP will lead to low accuracy and just shorten few time (can be ignored) in my experiment.</p>",
              "rawMarkdown": "To try light-weight model or quantitize the model (onnx or openvino), I think increase HOP will lead to low accuracy and just shorten few time (can be ignored) in my experiment.",
              "votes": 1
            },
            {
              "id": 3198922,
              "postDate": "2025-05-10T07:46:27.177Z",
              "content": "<p>Ok, but increasing tta upto value 4,5 can boost up the score, but my notebook timed out &amp; also increasing HOP will also decreases the accuracy.</p>",
              "rawMarkdown": "Ok, but increasing tta upto value 4,5 can boost up the score, but my notebook timed out & also increasing HOP will also decreases the accuracy."
            }
          ]
        }
      ]
    },
    {
      "id": 3209120,
      "postDate": "2025-05-25T09:07:24.410Z",
      "content": "<p>The posting give me some hint, appreciate!</p>",
      "rawMarkdown": "The posting give me some hint, appreciate!",
      "votes": 2
    },
    {
      "id": 3203448,
      "postDate": "2025-05-16T18:39:33.560Z",
      "content": "<p>Thanks for the insights! Did you find any efficient way to tune the melspec parameters?</p>",
      "rawMarkdown": "Thanks for the insights! Did you find any efficient way to tune the melspec parameters?"
    },
    {
      "id": 3180493,
      "postDate": "2025-04-16T17:44:19.720Z",
      "content": "<p>Could you talk about what you mean by this a bit more? \"Filtering/Processing CSA recordings by Human sound helps.\" And also by post processing I assume you mean things like smoothing predictions, similar to that of some of the top public submission notebooks?</p>",
      "rawMarkdown": "Could you talk about what you mean by this a bit more? \"Filtering/Processing CSA recordings by Human sound helps.\" And also by post processing I assume you mean things like smoothing predictions, similar to that of some of the top public submission notebooks?",
      "votes": 1,
      "replies": [
        {
          "id": 3180629,
          "postDate": "2025-04-16T21:54:07.370Z",
          "content": "<p>I think the better way to describe it is, If I train the model without Filtering, then CAM of my model is always on the Human sound part, and the model is focusing on human sound as discriminative features for that class. We need to handle that.</p>",
          "rawMarkdown": "I think the better way to describe it is, If I train the model without Filtering, then CAM of my model is always on the Human sound part, and the model is focusing on human sound as discriminative features for that class. We need to handle that.",
          "votes": 7
        }
      ]
    },
    {
      "id": 3180253,
      "postDate": "2025-04-16T10:52:57.473Z",
      "content": "<p>I have a wonder that should I change the N_MELS in the train phase same to the inference phase?<br>\nPls reply</p>",
      "rawMarkdown": "I have a wonder that should I change the N_MELS in the train phase same to the inference phase?\nPls reply",
      "votes": 1,
      "replies": [
        {
          "id": 3180255,
          "postDate": "2025-04-16T10:56:40.837Z",
          "content": "<p>You should use the same N_NELS in inference which you used to train the model.</p>",
          "rawMarkdown": "You should use the same N_NELS in inference which you used to train the model."
        }
      ]
    },
    {
      "id": 3178087,
      "postDate": "2025-04-13T17:42:32.550Z",
      "content": "<p>training is too unstable, result varies a lot</p>",
      "rawMarkdown": "training is too unstable, result varies a lot",
      "votes": 1,
      "replies": [
        {
          "id": 3178129,
          "postDate": "2025-04-13T18:44:22.260Z",
          "content": "<p>There's a big domain shift between soundscapes and train data.</p>",
          "rawMarkdown": "There's a big domain shift between soundscapes and train data.",
          "votes": 2
        }
      ]
    },
    {
      "id": 3192188,
      "postDate": "2025-05-02T14:56:14.207Z",
      "content": "<p>With a single EfficientNet-B0 model (single fold), I was able to get up to a 0.843 public score. I haven’t been able to push a single model beyond 0.85, but ensembling results from multiple folds gave me a score of 0.857. For training, I’m not using random 5-second crops—instead, I always use the first 5 seconds. Also, I’ve barely made use of the soundscapes data so far.</p>",
      "rawMarkdown": "With a single EfficientNet-B0 model (single fold), I was able to get up to a 0.843 public score. I haven’t been able to push a single model beyond 0.85, but ensembling results from multiple folds gave me a score of 0.857. For training, I’m not using random 5-second crops—instead, I always use the first 5 seconds. Also, I’ve barely made use of the soundscapes data so far.",
      "votes": 2,
      "replies": [
        {
          "id": 3203134,
          "postDate": "2025-05-16T10:38:42.423Z",
          "content": "<p>Did you use any additional data?</p>",
          "rawMarkdown": "Did you use any additional data?",
          "replies": [
            {
              "id": 3205475,
              "postDate": "2025-05-20T01:35:30.937Z",
              "content": "<p>No. I used only given train data.</p>",
              "rawMarkdown": "No. I used only given train data."
            }
          ]
        }
      ]
    },
    {
      "id": 3177847,
      "postDate": "2025-04-13T10:47:46.437Z",
      "content": "<p>\"2. When I look at last year's leaderboard, position # 5 was able to achieve a awesome melspec model without any pseudo labels or anything, while other top solutions used pseudo labels. I believe melspec parameters were one of the import things in that solution. (Not sure.) Considering this, I started exploring and I was able to get scores from 0.810 to 0.859 just by changing the melspec parameters. You just have to think about it, visualize and explore more. (5 subs a day to find the best parameters).\" This is very promising. However, I am uncertain about the randomness of the results. In my k-fold results, one fold has an AUC of 0.817, while another fold has an AUC of 0.85. </p>",
      "rawMarkdown": "\"2. When I look at last year's leaderboard, position # 5 was able to achieve a awesome melspec model without any pseudo labels or anything, while other top solutions used pseudo labels. I believe melspec parameters were one of the import things in that solution. (Not sure.) Considering this, I started exploring and I was able to get scores from 0.810 to 0.859 just by changing the melspec parameters. You just have to think about it, visualize and explore more. (5 subs a day to find the best parameters).\" This is very promising. However, I am uncertain about the randomness of the results. In my k-fold results, one fold has an AUC of 0.817, while another fold has an AUC of 0.85. ",
      "votes": 2,
      "replies": [
        {
          "id": 3177937,
          "postDate": "2025-04-13T13:48:06.630Z",
          "content": "<p>your kfold seems stable not quickly overfitting to 0.95+. </p>",
          "rawMarkdown": "your kfold seems stable not quickly overfitting to 0.95+. ",
          "replies": [
            {
              "id": 3178093,
              "postDate": "2025-04-13T17:52:15.240Z",
              "content": "<p>My CV is 0.99, 0.817 and 0.85 are public score.</p>",
              "rawMarkdown": "My CV is 0.99, 0.817 and 0.85 are public score."
            },
            {
              "id": 3178485,
              "postDate": "2025-04-14T08:38:50.797Z",
              "content": "<p>thanks for clarifying</p>",
              "rawMarkdown": "thanks for clarifying"
            }
          ]
        },
        {
          "id": 3178004,
          "postDate": "2025-04-13T16:12:00.113Z",
          "content": "<p>Can you tell how many folds do you use?</p>",
          "rawMarkdown": "Can you tell how many folds do you use?",
          "replies": [
            {
              "id": 3178094,
              "postDate": "2025-04-13T17:52:29.070Z",
              "content": "<p>I use 5 folds.</p>",
              "rawMarkdown": "I use 5 folds.",
              "votes": 1
            }
          ]
        },
        {
          "id": 3178368,
          "postDate": "2025-04-14T04:49:36.773Z",
          "content": "<p><a href=\"https://www.kaggle.com/quan0095\" target=\"_blank\">@quan0095</a>, do you only change the parameters during inference, or do you change them during both training and inference?</p>",
          "rawMarkdown": "@quan0095, do you only change the parameters during inference, or do you change them during both training and inference?",
          "replies": [
            {
              "id": 3178701,
              "postDate": "2025-04-14T13:47:28.990Z",
              "content": "<p>Sorry, I am quoting the words of <a href=\"https://www.kaggle.com/salmanahmedtamu\" target=\"_blank\">@salmanahmedtamu</a> </p>",
              "rawMarkdown": "Sorry, I am quoting the words of @salmanahmedtamu "
            }
          ]
        },
        {
          "id": 3180249,
          "postDate": "2025-04-16T10:45:51.040Z",
          "content": "<p>How are you guys, splitting into k-folds.? I tried multilabel kfold creation, but the CV and leaderboard score is less than my kfold creation using only the primary label. </p>",
          "rawMarkdown": "How are you guys, splitting into k-folds.? I tried multilabel kfold creation, but the CV and leaderboard score is less than my kfold creation using only the primary label. ",
          "replies": [
            {
              "id": 3180427,
              "postDate": "2025-04-16T15:57:22.723Z",
              "content": "<p>I use StratifiedKFold. training is too unstable 🤧 CV score is less meaningful. </p>",
              "rawMarkdown": "I use StratifiedKFold. training is too unstable 🤧 CV score is less meaningful. "
            },
            {
              "id": 3180430,
              "postDate": "2025-04-16T16:03:33.870Z",
              "content": "<p>Us bro us (not with ranking but with problems 🤝) </p>",
              "rawMarkdown": "Us bro us (not with ranking but with problems 🤝) "
            }
          ]
        }
      ]
    },
    {
      "id": 3179843,
      "postDate": "2025-04-15T18:47:48.367Z",
      "content": "<p>May I ask what is the range for Mel spec params hop length,nfft n_mels just range to guide me not a specific number off course . And for sure I will understand if you don't want to reveal this info thanks in advance <a href=\"https://www.kaggle.com/salmanahmedtamu\" target=\"_blank\">@salmanahmedtamu</a> </p>",
      "rawMarkdown": "May I ask what is the range for Mel spec params hop length,nfft n_mels just range to guide me not a specific number off course . And for sure I will understand if you don't want to reveal this info thanks in advance @salmanahmedtamu ",
      "replies": [
        {
          "id": 3182234,
          "postDate": "2025-04-19T00:36:31.900Z",
          "content": "<p>set nfft to 2048 and others for 128/256/512 and try. I tested a lot combinations and the score ranging from 0.751 to 0.826.</p>",
          "rawMarkdown": "set nfft to 2048 and others for 128/256/512 and try. I tested a lot combinations and the score ranging from 0.751 to 0.826.",
          "replies": [
            {
              "id": 3183560,
              "postDate": "2025-04-21T02:09:45.707Z",
              "rawMarkdown": "",
              "isDeleted": true
            },
            {
              "id": 3194406,
              "postDate": "2025-05-05T20:43:07.293Z",
              "content": "<p>What other mel spec parameters should I focus on? you said you were able to achieve 0.859 from just these changes, but the previous comment says score ranging from 0.751 to 0.826. I am a little confused :(</p>",
              "rawMarkdown": "What other mel spec parameters should I focus on? you said you were able to achieve 0.859 from just these changes, but the previous comment says score ranging from 0.751 to 0.826. I am a little confused :("
            },
            {
              "id": 3194536,
              "postDate": "2025-05-06T02:59:36.633Z",
              "content": "<p>'achieve 0.859'<br>\nthat's the topic author, not me.😁</p>",
              "rawMarkdown": "'achieve 0.859'\nthat's the topic author, not me.😁"
            }
          ]
        }
      ]
    },
    {
      "id": 3210476,
      "postDate": "2025-05-27T08:16:43.843Z",
      "content": "<p>thanks for notes。Does adding background noise help or not？😃</p>",
      "rawMarkdown": "thanks for notes。Does adding background noise help or not？😃"
    },
    {
      "id": 3193582,
      "postDate": "2025-05-04T16:23:27.597Z",
      "content": "<blockquote>\n  <p>When I look at last year's leaderboard, position # 5 was able to achieve a awesome melspec model without any pseudo labels or anything, while other top solutions used pseudo labels. I believe melspec parameters were one of the import things in that solution. (Not sure.) </p>\n</blockquote>\n<p>The combination with the raw waveform model could also be the reason, not necessarily melspec param.</p>",
      "rawMarkdown": ">When I look at last year's leaderboard, position # 5 was able to achieve a awesome melspec model without any pseudo labels or anything, while other top solutions used pseudo labels. I believe melspec parameters were one of the import things in that solution. (Not sure.) \n\nThe combination with the raw waveform model could also be the reason, not necessarily melspec param."
    },
    {
      "id": 3190316,
      "postDate": "2025-04-30T13:40:44.977Z",
      "content": "<p>Can I ask you for how many epochs do you usually train your models?</p>",
      "rawMarkdown": "Can I ask you for how many epochs do you usually train your models?"
    },
    {
      "id": 3190150,
      "postDate": "2025-04-30T07:26:57.943Z",
      "content": "<p>Hi Salman! I was wondering—did you track the fixed epoch of a single model, or did you use model soup?</p>",
      "rawMarkdown": "Hi Salman! I was wondering—did you track the fixed epoch of a single model, or did you use model soup?",
      "replies": [
        {
          "id": 3190154,
          "postDate": "2025-04-30T07:46:09.150Z",
          "content": "<p>In my current solution, for some models I use final epoch and for some models I use model soup</p>",
          "rawMarkdown": "In my current solution, for some models I use final epoch and for some models I use model soup",
          "votes": 1
        }
      ]
    },
    {
      "id": 3188397,
      "postDate": "2025-04-27T14:08:37.903Z",
      "content": "<p>Hey, Sal!<br>\nHow many in_channels did you choose? and which architecture do you use?</p>",
      "rawMarkdown": "Hey, Sal!\nHow many in_channels did you choose? and which architecture do you use?",
      "replies": [
        {
          "id": 3188404,
          "postDate": "2025-04-27T14:22:33.757Z",
          "content": "<p>3 Channels, efficient net b0</p>",
          "rawMarkdown": "3 Channels, efficient net b0",
          "votes": 2,
          "replies": [
            {
              "id": 3188405,
              "postDate": "2025-04-27T14:25:57.260Z",
              "content": "<p>Thank you 🙏<br>\nHope you win it this year ❤️</p>",
              "rawMarkdown": "Thank you 🙏\nHope you win it this year ❤️"
            },
            {
              "id": 3188997,
              "postDate": "2025-04-28T15:37:22.450Z",
              "content": "<p>Hey, Sal!<br>\nI was working with 1 input channel, and I had an 82.2 LB score (single-fold).<br>\nWhen I adapted my code to the 3 channels, sadly, my score dropped to 79.8. This is because I need to refine the parameters to suit the 3 channels, right?</p>",
              "rawMarkdown": "Hey, Sal!\nI was working with 1 input channel, and I had an 82.2 LB score (single-fold).\nWhen I adapted my code to the 3 channels, sadly, my score dropped to 79.8. This is because I need to refine the parameters to suit the 3 channels, right?"
            },
            {
              "id": 3189031,
              "postDate": "2025-04-28T16:30:32.800Z",
              "content": "<p>From my experiments, models usually starts to overfit pretty easily. <br>\nConsidering models are pretrained on Imagenet dataset, I think, in your case, when you trained with 1 input channel, model didn't overfit that much because it spent a couple of epochs to adjust 1 channel input to it's learned weights.<br>\nBut when you trained on 3-channels, model is already trained to extract pretty good features for 3-channel input so it started overfitting a little earlier.<br>\nI would suggest you to reduce a couple of epochs on 3 channels and look at the LB score, if it's still worse, then there might be something else.<br>\nAlso, are you applying Imagenet normalization on 3-channel input? if yes, then your model might behave differently than single channel. <br>\nTo compare b/w both inputs, make sure normalization and everything is same in the experiment.<br>\nHope it helps.</p>",
              "rawMarkdown": "From my experiments, models usually starts to overfit pretty easily. \nConsidering models are pretrained on Imagenet dataset, I think, in your case, when you trained with 1 input channel, model didn't overfit that much because it spent a couple of epochs to adjust 1 channel input to it's learned weights.\nBut when you trained on 3-channels, model is already trained to extract pretty good features for 3-channel input so it started overfitting a little earlier.\nI would suggest you to reduce a couple of epochs on 3 channels and look at the LB score, if it's still worse, then there might be something else.\nAlso, are you applying Imagenet normalization on 3-channel input? if yes, then your model might behave differently than single channel. \nTo compare b/w both inputs, make sure normalization and everything is same in the experiment.\nHope it helps.",
              "votes": 6
            }
          ]
        }
      ]
    },
    {
      "id": 3188064,
      "postDate": "2025-04-27T02:11:42.273Z",
      "content": "<p>How did you do this Filtering/Processing CSA recordings by Human sound? I have had a ton of problems in my efforts to do the same.</p>",
      "rawMarkdown": "How did you do this Filtering/Processing CSA recordings by Human sound? I have had a ton of problems in my efforts to do the same."
    },
    {
      "id": 3186223,
      "postDate": "2025-04-24T11:09:20.647Z",
      "content": "<p>Hello, I would like to ask if the random 5s is used to convert into a spectrogram, if the audio length is longer, such as 10s, will some information be lost？</p>",
      "rawMarkdown": "Hello, I would like to ask if the random 5s is used to convert into a spectrogram, if the audio length is longer, such as 10s, will some information be lost？",
      "replies": [
        {
          "id": 3188701,
          "postDate": "2025-04-28T04:17:33.507Z",
          "content": "<p>Agreed, but that’s the trade-off. Just test different durations and pick the best one.</p>",
          "rawMarkdown": "Agreed, but that’s the trade-off. Just test different durations and pick the best one."
        }
      ]
    },
    {
      "id": 3185872,
      "postDate": "2025-04-24T00:23:50.997Z",
      "content": "<p>Thanks for sharing! I have the similar trend—changing the melspec parameters can really impact the score. So far, I’ve managed to hit a public score of 0.836 using a single EfficientNet-B0 model (single fold), but it’s quite unstable. Even slight changes to the parameters or switching folds can cause the score to drop a lot (public score is around 0.8). How’s it going on your end?</p>",
      "rawMarkdown": "Thanks for sharing! I have the similar trend—changing the melspec parameters can really impact the score. So far, I’ve managed to hit a public score of 0.836 using a single EfficientNet-B0 model (single fold), but it’s quite unstable. Even slight changes to the parameters or switching folds can cause the score to drop a lot (public score is around 0.8). How’s it going on your end?",
      "replies": [
        {
          "id": 3186231,
          "postDate": "2025-04-24T11:24:33.937Z",
          "content": "<p><a href=\"https://www.kaggle.com/kmatsu01\" target=\"_blank\">@kmatsu01</a> Are you Images using 3 channels or single channel for training ?</p>",
          "rawMarkdown": "@kmatsu01 Are you Images using 3 channels or single channel for training ?",
          "isDeleted": true,
          "replies": [
            {
              "id": 3186309,
              "postDate": "2025-04-24T13:28:41.120Z",
              "content": "<p>3 channels. I haven't tried single channel.</p>",
              "rawMarkdown": "3 channels. I haven't tried single channel."
            },
            {
              "id": 3187296,
              "postDate": "2025-04-25T19:14:01.730Z",
              "content": "<p>Thanks <a href=\"https://www.kaggle.com/kmatsu01\" target=\"_blank\">@kmatsu01</a> </p>",
              "rawMarkdown": "Thanks @kmatsu01 ",
              "votes": 1,
              "isDeleted": true
            },
            {
              "id": 3199030,
              "postDate": "2025-05-10T11:21:15.430Z",
              "content": "<p>can i ask what do 3 channels mean?<br>\nis a spec gram channel duplicate to 3 or sth else?<br>\nThank you.</p>",
              "rawMarkdown": "can i ask what do 3 channels mean?\nis a spec gram channel duplicate to 3 or sth else?\nThank you."
            },
            {
              "id": 3199039,
              "postDate": "2025-05-10T11:26:05.013Z",
              "content": "<p><a href=\"https://www.kaggle.com/qminh211\" target=\"_blank\">@qminh211</a> Yes ,we duplicate single channel into 3 channel .</p>",
              "rawMarkdown": "@qminh211 Yes ,we duplicate single channel into 3 channel .",
              "votes": 1,
              "isDeleted": true
            }
          ]
        },
        {
          "id": 3189671,
          "postDate": "2025-04-29T15:21:35.977Z",
          "rawMarkdown": "",
          "isDeleted": true
        }
      ]
    },
    {
      "id": 3184868,
      "postDate": "2025-04-22T14:57:46.817Z",
      "content": "<p>thanks for sharing and enhancing our knowledge….</p>",
      "rawMarkdown": "thanks for sharing and enhancing our knowledge...."
    },
    {
      "id": 3184495,
      "postDate": "2025-04-22T06:58:02.370Z",
      "content": "<p>Hi, may I ask your score including TTA or any post processing to achieve this score?</p>",
      "rawMarkdown": "Hi, may I ask your score including TTA or any post processing to achieve this score?",
      "replies": [
        {
          "id": 3184593,
          "postDate": "2025-04-22T09:19:18.443Z",
          "content": "<p>Yeah, I applied TTA / Post Processing, in all of these.</p>",
          "rawMarkdown": "Yeah, I applied TTA / Post Processing, in all of these.",
          "votes": 1
        }
      ]
    },
    {
      "id": 3183573,
      "postDate": "2025-04-21T02:47:20.880Z",
      "content": "<p>Thanks for sharing.</p>\n<p>Did you use external data to achieve this .872 mentioned here?</p>",
      "rawMarkdown": "Thanks for sharing.\n\nDid you use external data to achieve this .872 mentioned here?",
      "replies": [
        {
          "id": 3184038,
          "postDate": "2025-04-21T15:05:56.393Z",
          "content": "<p>No, I don't. BirdCLEF 2025 data only, till now.</p>",
          "rawMarkdown": "No, I don't. BirdCLEF 2025 data only, till now.",
          "votes": 3
        }
      ]
    },
    {
      "id": 3182652,
      "postDate": "2025-04-19T17:08:23.430Z",
      "content": "<p>thanks for sharing! I notice that your training pipeline didn't use pre-compute  mel spec. Will it affect your training time? My training process takes 1min40s per epoch with pre-compute mel spec, and computing 5s mel spec for every audio takes 20min. 😅</p>",
      "rawMarkdown": "thanks for sharing! I notice that your training pipeline didn't use pre-compute  mel spec. Will it affect your training time? My training process takes 1min40s per epoch with pre-compute mel spec, and computing 5s mel spec for every audio takes 20min. 😅",
      "replies": [
        {
          "id": 3182757,
          "postDate": "2025-04-19T19:17:04.750Z",
          "content": "<p>I am training on random 5 seconds, that's why didn't save Mel specs. <br>\nIt takes around 2.5 minutes for a single epoch.</p>",
          "rawMarkdown": "I am training on random 5 seconds, that's why didn't save Mel specs. \nIt takes around 2.5 minutes for a single epoch.",
          "votes": 3,
          "replies": [
            {
              "id": 3182771,
              "postDate": "2025-04-19T19:40:32.200Z",
              "content": "<p>Hmm that’s interesting, I checked the old notebook with some of the modifications and I was getting like 20 minutes per epoch. I was thinking it wasn’t even using GPU haha I probably did something wrong </p>",
              "rawMarkdown": "Hmm that’s interesting, I checked the old notebook with some of the modifications and I was getting like 20 minutes per epoch. I was thinking it wasn’t even using GPU haha I probably did something wrong "
            },
            {
              "id": 3188006,
              "postDate": "2025-04-26T21:51:43.033Z",
              "content": "<p>how do you get these fast training times? Without augmentations I am at 4mins and with mixup and specaugment more like 13mins/epoch. I even precomputed melspecs..</p>",
              "rawMarkdown": "how do you get these fast training times? Without augmentations I am at 4mins and with mixup and specaugment more like 13mins/epoch. I even precomputed melspecs.."
            }
          ]
        }
      ]
    },
    {
      "id": 3182172,
      "postDate": "2025-04-18T21:35:02.227Z",
      "content": "<p>Thanks for sharing this — super valuable! </p>\n<p>Your baseline-to-0.872 journey gives a lot of actionable directions to explore. <br>\nTreating MelSpecs as images – totally agree. Applying image-based augmentations (like RandAugment, RandomErasing, Mixup, SpecAugment) feels like a powerful combo. Did you use torchvision or albumentations for this?</p>",
      "rawMarkdown": "Thanks for sharing this — super valuable! \n\nYour baseline-to-0.872 journey gives a lot of actionable directions to explore. \nTreating MelSpecs as images – totally agree. Applying image-based augmentations (like RandAugment, RandomErasing, Mixup, SpecAugment) feels like a powerful combo. Did you use torchvision or albumentations for this?\n\n",
      "replies": [
        {
          "id": 3182232,
          "postDate": "2025-04-19T00:22:12.813Z",
          "content": "<p>torchvision</p>",
          "rawMarkdown": "torchvision"
        }
      ]
    },
    {
      "id": 3181215,
      "postDate": "2025-04-17T15:33:53.540Z",
      "content": "<p>did you get 0.872 without pseudo labels?</p>",
      "rawMarkdown": "did you get 0.872 without pseudo labels?",
      "replies": [
        {
          "id": 3181231,
          "postDate": "2025-04-17T15:53:01.040Z",
          "content": "<p>Yes,     Without Pseudo Labels</p>",
          "rawMarkdown": "Yes,     Without Pseudo Labels"
        }
      ]
    },
    {
      "id": 3178772,
      "postDate": "2025-04-14T15:27:51.060Z",
      "content": "<p>Is FocalBCE a weighting of focal loss and bce loss? Did you have to experiment a lot with different weighting or focal params to get it to work?</p>",
      "rawMarkdown": "Is FocalBCE a weighting of focal loss and bce loss? Did you have to experiment a lot with different weighting or focal params to get it to work?",
      "replies": [
        {
          "id": 3179013,
          "postDate": "2025-04-14T22:20:29.267Z",
          "content": "<p>It's just sum of both losses.</p>",
          "rawMarkdown": "It's just sum of both losses."
        }
      ]
    },
    {
      "id": 3178650,
      "postDate": "2025-04-14T12:41:42.983Z",
      "content": "<p>Quite hard to tell for me whats work or not, because i have contradictory behavior when changing seed</p>",
      "rawMarkdown": "Quite hard to tell for me whats work or not, because i have contradictory behavior when changing seed"
    },
    {
      "id": 3178411,
      "postDate": "2025-04-14T06:10:32.413Z",
      "content": "<p>May I ask how you handled the human voice parts in the CSA recordings? Did you remove the segments containing human voices?</p>",
      "rawMarkdown": "May I ask how you handled the human voice parts in the CSA recordings? Did you remove the segments containing human voices?"
    },
    {
      "id": 3178158,
      "postDate": "2025-04-13T19:38:22.580Z",
      "content": "<p>Has anyone tried using Triplet loss for this purpose?</p>",
      "rawMarkdown": "Has anyone tried using Triplet loss for this purpose?"
    },
    {
      "id": 3177897,
      "postDate": "2025-04-13T12:22:00.103Z",
      "content": "<p>Nice post! I can tell that usage of FocalLoss had also improved my result.</p>",
      "rawMarkdown": "Nice post! I can tell that usage of FocalLoss had also improved my result.",
      "replies": [
        {
          "id": 3184672,
          "postDate": "2025-04-22T11:05:41.143Z",
          "content": "<p>May I ask did you clean your data? Because In my expriments, focal loss hurt performance in my local valid set(both loss and auc)</p>",
          "rawMarkdown": "May I ask did you clean your data? Because In my expriments, focal loss hurt performance in my local valid set(both loss and auc)",
          "replies": [
            {
              "id": 3187924,
              "postDate": "2025-04-26T17:55:29.680Z",
              "content": "<p>Yes, I did.</p>",
              "rawMarkdown": "Yes, I did."
            }
          ]
        }
      ]
    },
    {
      "id": 3177896,
      "postDate": "2025-04-13T12:19:24.557Z",
      "content": "<p>Thanks for sharing your valuable insights <a href=\"https://www.kaggle.com/salmanahmedtamu\" target=\"_blank\">@salmanahmedtamu</a> , this will help to improve everyone's solutions.</p>",
      "rawMarkdown": "Thanks for sharing your valuable insights @salmanahmedtamu , this will help to improve everyone's solutions."
    },
    {
      "id": 3206043,
      "postDate": "2025-05-20T18:58:59.760Z",
      "content": "<p><a href=\"https://www.kaggle.com/salmanahmedtamu\" target=\"_blank\">@salmanahmedtamu</a> I know it is very stupid question but I see many inference notebook where participants using single channel. Is it something like that we have to train with 3 channel (lets assume) and inference on single channel.</p>",
      "rawMarkdown": "@salmanahmedtamu I know it is very stupid question but I see many inference notebook where participants using single channel. Is it something like that we have to train with 3 channel (lets assume) and inference on single channel.",
      "isDeleted": true,
      "replies": [
        {
          "id": 3206216,
          "postDate": "2025-05-21T03:29:34.477Z",
          "content": "<p><a href=\"https://www.kaggle.com/asteyagaur\" target=\"_blank\">@asteyagaur</a> No, I think they trained on a single channel and predict on a single channel. <br>\nFor any audio segment, we create a mel spectogram that represents frequency bins and temporal features. That mel spec is a 2D image which can be represented as a single channel.<br>\nBut some times for model to converge faster, it's better to keep the image as 3 channels, but it shouldn't matter much if trained correctly.</p>",
          "rawMarkdown": "@asteyagaur No, I think they trained on a single channel and predict on a single channel. \nFor any audio segment, we create a mel spectogram that represents frequency bins and temporal features. That mel spec is a 2D image which can be represented as a single channel.\nBut some times for model to converge faster, it's better to keep the image as 3 channels, but it shouldn't matter much if trained correctly."
        }
      ]
    },
    {
      "id": 3205960,
      "postDate": "2025-05-20T16:10:27.013Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 3178659,
      "postDate": "2025-04-14T12:54:46.300Z",
      "content": "<p>I was able to get scores from 0.810 to 0.859 just by changing the melspec parameters. You just have to think about it, visualize and explore more. (5 subs a day to find the best parameters). <a href=\"https://www.kaggle.com/salmanahmedtamu\" target=\"_blank\">@salmanahmedtamu</a> did you change mels params for both train and inference or just in inference </p>",
      "rawMarkdown": "I was able to get scores from 0.810 to 0.859 just by changing the melspec parameters. You just have to think about it, visualize and explore more. (5 subs a day to find the best parameters). @salmanahmedtamu did you change mels params for both train and inference or just in inference ",
      "isDeleted": true,
      "replies": [
        {
          "id": 3179014,
          "postDate": "2025-04-14T22:20:59.657Z",
          "content": "<p>I do inference on the same features / melspec on which the model was trained.</p>",
          "rawMarkdown": "I do inference on the same features / melspec on which the model was trained."
        }
      ]
    },
    {
      "id": 3213546,
      "postDate": "2025-05-30T06:55:31.960Z",
      "content": "<p>Thanks for sharing!</p>",
      "rawMarkdown": "Thanks for sharing!",
      "votes": -1
    },
    {
      "id": 3180189,
      "postDate": "2025-04-16T09:14:00.430Z",
      "content": "<p>Thanks for your sharing !!!</p>",
      "rawMarkdown": "Thanks for your sharing !!!"
    },
    {
      "id": 3179706,
      "postDate": "2025-04-15T16:01:35.843Z",
      "content": "<p>THANKS. FOR SHARING </p>",
      "rawMarkdown": "THANKS. FOR SHARING "
    },
    {
      "id": 3178509,
      "postDate": "2025-04-14T09:20:43.810Z",
      "content": "<p>thanks! very helpful</p>",
      "rawMarkdown": "thanks! very helpful"
    },
    {
      "id": 3182405,
      "postDate": "2025-04-19T08:40:40.553Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 3195341,
      "author_name": "PRASHANT SHUKLA91",
      "author_url": "",
      "post_date": "2025-05-06T23:01:50.483000",
      "content": "<p>After applying TTA using efficientnet_B0 model, every time , notebook throws message as network threw exception, while submission on scoring. Can anyone guide to resolve this issue?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 3195344,
          "author_name": "Salman Ahmed",
          "author_url": "",
          "post_date": "2025-05-06T23:12:38.450000",
          "content": "<p>Can you share that notebook? or the code for TTA? </p>",
          "votes": 0,
          "replies": [
            {
              "id": 3195349,
              "author_name": "",
              "author_url": "",
              "post_date": "2025-05-06T23:28:23.460000",
              "content": "",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3195351,
              "author_name": "PRASHANT SHUKLA91",
              "author_url": "",
              "post_date": "2025-05-06T23:31:36.143000",
              "content": "<p><a href=\"https://www.kaggle.com/salmanahmedtamu\" target=\"_blank\">@salmanahmedtamu</a> , this is the attached notebook, I referred. Plz guide me. In this use_tta = False by default, After setting use_tta = true, it is threwing exception, on submission for scoring, even for minimum values.</p>",
              "votes": 1,
              "replies": []
            }
          ]
        },
        {
          "id": 3196907,
          "author_name": "Kurise",
          "author_url": "",
          "post_date": "2025-05-07T15:31:48.570000",
          "content": "<p>There is an error using source data, you need to add \".copy\" when you use the apply_tta function or return directly your augmentatied mel.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 3196910,
          "author_name": "Kurise",
          "author_url": "",
          "post_date": "2025-05-07T15:36:36.300000",
          "content": "<p>I think the error raise because of the numpy.</p>",
          "votes": 2,
          "replies": [
            {
              "id": 3198881,
              "author_name": "PRASHANT SHUKLA91",
              "author_url": "",
              "post_date": "2025-05-10T06:26:06.267000",
              "content": "<p>Thanks, now the notebook timed out even i used tta=3(minimum).  Should I increase HOP_LENGTH?</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3198917,
              "author_name": "Kurise",
              "author_url": "",
              "post_date": "2025-05-10T07:34:45.463000",
              "content": "<p>To try light-weight model or quantitize the model (onnx or openvino), I think increase HOP will lead to low accuracy and just shorten few time (can be ignored) in my experiment.</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 3198922,
              "author_name": "PRASHANT SHUKLA91",
              "author_url": "",
              "post_date": "2025-05-10T07:46:27.177000",
              "content": "<p>Ok, but increasing tta upto value 4,5 can boost up the score, but my notebook timed out &amp; also increasing HOP will also decreases the accuracy.</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3209120,
      "author_name": "",
      "author_url": "",
      "post_date": "2025-05-25T09:07:24.410000",
      "content": "<p>The posting give me some hint, appreciate!</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 3203448,
      "author_name": "Diganta",
      "author_url": "",
      "post_date": "2025-05-16T18:39:33.560000",
      "content": "<p>Thanks for the insights! Did you find any efficient way to tune the melspec parameters?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3180493,
      "author_name": "Cody_Null",
      "author_url": "",
      "post_date": "2025-04-16T17:44:19.720000",
      "content": "<p>Could you talk about what you mean by this a bit more? \"Filtering/Processing CSA recordings by Human sound helps.\" And also by post processing I assume you mean things like smoothing predictions, similar to that of some of the top public submission notebooks?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 3180629,
          "author_name": "Salman Ahmed",
          "author_url": "",
          "post_date": "2025-04-16T21:54:07.370000",
          "content": "<p>I think the better way to describe it is, If I train the model without Filtering, then CAM of my model is always on the Human sound part, and the model is focusing on human sound as discriminative features for that class. We need to handle that.</p>",
          "votes": 7,
          "replies": []
        }
      ]
    },
    {
      "id": 3180253,
      "author_name": "qminh211",
      "author_url": "",
      "post_date": "2025-04-16T10:52:57.473000",
      "content": "<p>I have a wonder that should I change the N_MELS in the train phase same to the inference phase?<br>\nPls reply</p>",
      "votes": 1,
      "replies": [
        {
          "id": 3180255,
          "author_name": "Salman Ahmed",
          "author_url": "",
          "post_date": "2025-04-16T10:56:40.837000",
          "content": "<p>You should use the same N_NELS in inference which you used to train the model.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 3178087,
      "author_name": "yuanzhe zhou",
      "author_url": "",
      "post_date": "2025-04-13T17:42:32.550000",
      "content": "<p>training is too unstable, result varies a lot</p>",
      "votes": 1,
      "replies": [
        {
          "id": 3178129,
          "author_name": "Araik Tamazian",
          "author_url": "",
          "post_date": "2025-04-13T18:44:22.260000",
          "content": "<p>There's a big domain shift between soundscapes and train data.</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 3192188,
      "author_name": "kmn",
      "author_url": "",
      "post_date": "2025-05-02T14:56:14.207000",
      "content": "<p>With a single EfficientNet-B0 model (single fold), I was able to get up to a 0.843 public score. I haven’t been able to push a single model beyond 0.85, but ensembling results from multiple folds gave me a score of 0.857. For training, I’m not using random 5-second crops—instead, I always use the first 5 seconds. Also, I’ve barely made use of the soundscapes data so far.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 3203134,
          "author_name": "rhythm master",
          "author_url": "",
          "post_date": "2025-05-16T10:38:42.423000",
          "content": "<p>Did you use any additional data?</p>",
          "votes": 0,
          "replies": [
            {
              "id": 3205475,
              "author_name": "kmn",
              "author_url": "",
              "post_date": "2025-05-20T01:35:30.937000",
              "content": "<p>No. I used only given train data.</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3177847,
      "author_name": "Quan Vu",
      "author_url": "",
      "post_date": "2025-04-13T10:47:46.437000",
      "content": "<p>\"2. When I look at last year's leaderboard, position # 5 was able to achieve a awesome melspec model without any pseudo labels or anything, while other top solutions used pseudo labels. I believe melspec parameters were one of the import things in that solution. (Not sure.) Considering this, I started exploring and I was able to get scores from 0.810 to 0.859 just by changing the melspec parameters. You just have to think about it, visualize and explore more. (5 subs a day to find the best parameters).\" This is very promising. However, I am uncertain about the randomness of the results. In my k-fold results, one fold has an AUC of 0.817, while another fold has an AUC of 0.85. </p>",
      "votes": 2,
      "replies": [
        {
          "id": 3177937,
          "author_name": "nymfree",
          "author_url": "",
          "post_date": "2025-04-13T13:48:06.630000",
          "content": "<p>your kfold seems stable not quickly overfitting to 0.95+. </p>",
          "votes": 0,
          "replies": [
            {
              "id": 3178093,
              "author_name": "Quan Vu",
              "author_url": "",
              "post_date": "2025-04-13T17:52:15.240000",
              "content": "<p>My CV is 0.99, 0.817 and 0.85 are public score.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3178485,
              "author_name": "nymfree",
              "author_url": "",
              "post_date": "2025-04-14T08:38:50.797000",
              "content": "<p>thanks for clarifying</p>",
              "votes": 0,
              "replies": []
            }
          ]
        },
        {
          "id": 3178004,
          "author_name": "Araik Tamazian",
          "author_url": "",
          "post_date": "2025-04-13T16:12:00.113000",
          "content": "<p>Can you tell how many folds do you use?</p>",
          "votes": 0,
          "replies": [
            {
              "id": 3178094,
              "author_name": "Quan Vu",
              "author_url": "",
              "post_date": "2025-04-13T17:52:29.070000",
              "content": "<p>I use 5 folds.</p>",
              "votes": 1,
              "replies": []
            }
          ]
        },
        {
          "id": 3178368,
          "author_name": "Wayne_127",
          "author_url": "",
          "post_date": "2025-04-14T04:49:36.773000",
          "content": "<p><a href=\"https://www.kaggle.com/quan0095\" target=\"_blank\">@quan0095</a>, do you only change the parameters during inference, or do you change them during both training and inference?</p>",
          "votes": 0,
          "replies": [
            {
              "id": 3178701,
              "author_name": "Quan Vu",
              "author_url": "",
              "post_date": "2025-04-14T13:47:28.990000",
              "content": "<p>Sorry, I am quoting the words of <a href=\"https://www.kaggle.com/salmanahmedtamu\" target=\"_blank\">@salmanahmedtamu</a> </p>",
              "votes": 0,
              "replies": []
            }
          ]
        },
        {
          "id": 3180249,
          "author_name": "Hotson Honet",
          "author_url": "",
          "post_date": "2025-04-16T10:45:51.040000",
          "content": "<p>How are you guys, splitting into k-folds.? I tried multilabel kfold creation, but the CV and leaderboard score is less than my kfold creation using only the primary label. </p>",
          "votes": 0,
          "replies": [
            {
              "id": 3180427,
              "author_name": "Quan Vu",
              "author_url": "",
              "post_date": "2025-04-16T15:57:22.723000",
              "content": "<p>I use StratifiedKFold. training is too unstable 🤧 CV score is less meaningful. </p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3180430,
              "author_name": "Hotson Honet",
              "author_url": "",
              "post_date": "2025-04-16T16:03:33.870000",
              "content": "<p>Us bro us (not with ranking but with problems 🤝) </p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3179843,
      "author_name": "Mohamed abdelrazik",
      "author_url": "",
      "post_date": "2025-04-15T18:47:48.367000",
      "content": "<p>May I ask what is the range for Mel spec params hop length,nfft n_mels just range to guide me not a specific number off course . And for sure I will understand if you don't want to reveal this info thanks in advance <a href=\"https://www.kaggle.com/salmanahmedtamu\" target=\"_blank\">@salmanahmedtamu</a> </p>",
      "votes": 0,
      "replies": [
        {
          "id": 3182234,
          "author_name": "Mingyao",
          "author_url": "",
          "post_date": "2025-04-19T00:36:31.900000",
          "content": "<p>set nfft to 2048 and others for 128/256/512 and try. I tested a lot combinations and the score ranging from 0.751 to 0.826.</p>",
          "votes": 0,
          "replies": [
            {
              "id": 3183560,
              "author_name": "",
              "author_url": "",
              "post_date": "2025-04-21T02:09:45.707000",
              "content": "",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3194406,
              "author_name": "Vajinder Kaur",
              "author_url": "",
              "post_date": "2025-05-05T20:43:07.293000",
              "content": "<p>What other mel spec parameters should I focus on? you said you were able to achieve 0.859 from just these changes, but the previous comment says score ranging from 0.751 to 0.826. I am a little confused :(</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3194536,
              "author_name": "Mingyao",
              "author_url": "",
              "post_date": "2025-05-06T02:59:36.633000",
              "content": "<p>'achieve 0.859'<br>\nthat's the topic author, not me.😁</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3210476,
      "author_name": "HaogeBob",
      "author_url": "",
      "post_date": "2025-05-27T08:16:43.843000",
      "content": "<p>thanks for notes。Does adding background noise help or not？😃</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3193582,
      "author_name": "yukiya",
      "author_url": "",
      "post_date": "2025-05-04T16:23:27.597000",
      "content": "<blockquote>\n  <p>When I look at last year's leaderboard, position # 5 was able to achieve a awesome melspec model without any pseudo labels or anything, while other top solutions used pseudo labels. I believe melspec parameters were one of the import things in that solution. (Not sure.) </p>\n</blockquote>\n<p>The combination with the raw waveform model could also be the reason, not necessarily melspec param.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3190316,
      "author_name": "Adam Mazur",
      "author_url": "",
      "post_date": "2025-04-30T13:40:44.977000",
      "content": "<p>Can I ask you for how many epochs do you usually train your models?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3190150,
      "author_name": "Ryenhails",
      "author_url": "",
      "post_date": "2025-04-30T07:26:57.943000",
      "content": "<p>Hi Salman! I was wondering—did you track the fixed epoch of a single model, or did you use model soup?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 3190154,
          "author_name": "Salman Ahmed",
          "author_url": "",
          "post_date": "2025-04-30T07:46:09.150000",
          "content": "<p>In my current solution, for some models I use final epoch and for some models I use model soup</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 3188397,
      "author_name": "Ahmed Elsayed",
      "author_url": "",
      "post_date": "2025-04-27T14:08:37.903000",
      "content": "<p>Hey, Sal!<br>\nHow many in_channels did you choose? and which architecture do you use?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 3188404,
          "author_name": "Salman Ahmed",
          "author_url": "",
          "post_date": "2025-04-27T14:22:33.757000",
          "content": "<p>3 Channels, efficient net b0</p>",
          "votes": 2,
          "replies": [
            {
              "id": 3188405,
              "author_name": "Ahmed Elsayed",
              "author_url": "",
              "post_date": "2025-04-27T14:25:57.260000",
              "content": "<p>Thank you 🙏<br>\nHope you win it this year ❤️</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3188997,
              "author_name": "Ahmed Elsayed",
              "author_url": "",
              "post_date": "2025-04-28T15:37:22.450000",
              "content": "<p>Hey, Sal!<br>\nI was working with 1 input channel, and I had an 82.2 LB score (single-fold).<br>\nWhen I adapted my code to the 3 channels, sadly, my score dropped to 79.8. This is because I need to refine the parameters to suit the 3 channels, right?</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3189031,
              "author_name": "Salman Ahmed",
              "author_url": "",
              "post_date": "2025-04-28T16:30:32.800000",
              "content": "<p>From my experiments, models usually starts to overfit pretty easily. <br>\nConsidering models are pretrained on Imagenet dataset, I think, in your case, when you trained with 1 input channel, model didn't overfit that much because it spent a couple of epochs to adjust 1 channel input to it's learned weights.<br>\nBut when you trained on 3-channels, model is already trained to extract pretty good features for 3-channel input so it started overfitting a little earlier.<br>\nI would suggest you to reduce a couple of epochs on 3 channels and look at the LB score, if it's still worse, then there might be something else.<br>\nAlso, are you applying Imagenet normalization on 3-channel input? if yes, then your model might behave differently than single channel. <br>\nTo compare b/w both inputs, make sure normalization and everything is same in the experiment.<br>\nHope it helps.</p>",
              "votes": 6,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3188064,
      "author_name": "Fir_las_47",
      "author_url": "",
      "post_date": "2025-04-27T02:11:42.273000",
      "content": "<p>How did you do this Filtering/Processing CSA recordings by Human sound? I have had a ton of problems in my efforts to do the same.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3186223,
      "author_name": "Yan  Zhang",
      "author_url": "",
      "post_date": "2025-04-24T11:09:20.647000",
      "content": "<p>Hello, I would like to ask if the random 5s is used to convert into a spectrogram, if the audio length is longer, such as 10s, will some information be lost？</p>",
      "votes": 0,
      "replies": [
        {
          "id": 3188701,
          "author_name": "lcx777",
          "author_url": "",
          "post_date": "2025-04-28T04:17:33.507000",
          "content": "<p>Agreed, but that’s the trade-off. Just test different durations and pick the best one.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 3185872,
      "author_name": "kmn",
      "author_url": "",
      "post_date": "2025-04-24T00:23:50.997000",
      "content": "<p>Thanks for sharing! I have the similar trend—changing the melspec parameters can really impact the score. So far, I’ve managed to hit a public score of 0.836 using a single EfficientNet-B0 model (single fold), but it’s quite unstable. Even slight changes to the parameters or switching folds can cause the score to drop a lot (public score is around 0.8). How’s it going on your end?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 3186231,
          "author_name": "",
          "author_url": "",
          "post_date": "2025-04-24T11:24:33.937000",
          "content": "<p><a href=\"https://www.kaggle.com/kmatsu01\" target=\"_blank\">@kmatsu01</a> Are you Images using 3 channels or single channel for training ?</p>",
          "votes": 0,
          "replies": [
            {
              "id": 3186309,
              "author_name": "kmn",
              "author_url": "",
              "post_date": "2025-04-24T13:28:41.120000",
              "content": "<p>3 channels. I haven't tried single channel.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3187296,
              "author_name": "",
              "author_url": "",
              "post_date": "2025-04-25T19:14:01.730000",
              "content": "<p>Thanks <a href=\"https://www.kaggle.com/kmatsu01\" target=\"_blank\">@kmatsu01</a> </p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 3199030,
              "author_name": "qminh211",
              "author_url": "",
              "post_date": "2025-05-10T11:21:15.430000",
              "content": "<p>can i ask what do 3 channels mean?<br>\nis a spec gram channel duplicate to 3 or sth else?<br>\nThank you.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3199039,
              "author_name": "",
              "author_url": "",
              "post_date": "2025-05-10T11:26:05.013000",
              "content": "<p><a href=\"https://www.kaggle.com/qminh211\" target=\"_blank\">@qminh211</a> Yes ,we duplicate single channel into 3 channel .</p>",
              "votes": 1,
              "replies": []
            }
          ]
        },
        {
          "id": 3189671,
          "author_name": "",
          "author_url": "",
          "post_date": "2025-04-29T15:21:35.977000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 3184868,
      "author_name": "Sarah Arshad",
      "author_url": "",
      "post_date": "2025-04-22T14:57:46.817000",
      "content": "<p>thanks for sharing and enhancing our knowledge….</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3184495,
      "author_name": "Mingyao",
      "author_url": "",
      "post_date": "2025-04-22T06:58:02.370000",
      "content": "<p>Hi, may I ask your score including TTA or any post processing to achieve this score?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 3184593,
          "author_name": "Salman Ahmed",
          "author_url": "",
          "post_date": "2025-04-22T09:19:18.443000",
          "content": "<p>Yeah, I applied TTA / Post Processing, in all of these.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 3183573,
      "author_name": "JuseTiz",
      "author_url": "",
      "post_date": "2025-04-21T02:47:20.880000",
      "content": "<p>Thanks for sharing.</p>\n<p>Did you use external data to achieve this .872 mentioned here?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 3184038,
          "author_name": "Salman Ahmed",
          "author_url": "",
          "post_date": "2025-04-21T15:05:56.393000",
          "content": "<p>No, I don't. BirdCLEF 2025 data only, till now.</p>",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 3182652,
      "author_name": "Yu Wu",
      "author_url": "",
      "post_date": "2025-04-19T17:08:23.430000",
      "content": "<p>thanks for sharing! I notice that your training pipeline didn't use pre-compute  mel spec. Will it affect your training time? My training process takes 1min40s per epoch with pre-compute mel spec, and computing 5s mel spec for every audio takes 20min. 😅</p>",
      "votes": 0,
      "replies": [
        {
          "id": 3182757,
          "author_name": "Salman Ahmed",
          "author_url": "",
          "post_date": "2025-04-19T19:17:04.750000",
          "content": "<p>I am training on random 5 seconds, that's why didn't save Mel specs. <br>\nIt takes around 2.5 minutes for a single epoch.</p>",
          "votes": 3,
          "replies": [
            {
              "id": 3182771,
              "author_name": "Cody_Null",
              "author_url": "",
              "post_date": "2025-04-19T19:40:32.200000",
              "content": "<p>Hmm that’s interesting, I checked the old notebook with some of the modifications and I was getting like 20 minutes per epoch. I was thinking it wasn’t even using GPU haha I probably did something wrong </p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3188006,
              "author_name": "AdiG",
              "author_url": "",
              "post_date": "2025-04-26T21:51:43.033000",
              "content": "<p>how do you get these fast training times? Without augmentations I am at 4mins and with mixup and specaugment more like 13mins/epoch. I even precomputed melspecs..</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3182172,
      "author_name": "GURRAM GOPICHANDH",
      "author_url": "",
      "post_date": "2025-04-18T21:35:02.227000",
      "content": "<p>Thanks for sharing this — super valuable! </p>\n<p>Your baseline-to-0.872 journey gives a lot of actionable directions to explore. <br>\nTreating MelSpecs as images – totally agree. Applying image-based augmentations (like RandAugment, RandomErasing, Mixup, SpecAugment) feels like a powerful combo. Did you use torchvision or albumentations for this?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 3182232,
          "author_name": "Salman Ahmed",
          "author_url": "",
          "post_date": "2025-04-19T00:22:12.813000",
          "content": "<p>torchvision</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 3181215,
      "author_name": "Dmitriiy2",
      "author_url": "",
      "post_date": "2025-04-17T15:33:53.540000",
      "content": "<p>did you get 0.872 without pseudo labels?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 3181231,
          "author_name": "Salman Ahmed",
          "author_url": "",
          "post_date": "2025-04-17T15:53:01.040000",
          "content": "<p>Yes,     Without Pseudo Labels</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 3178772,
      "author_name": "thacrobatheskis",
      "author_url": "",
      "post_date": "2025-04-14T15:27:51.060000",
      "content": "<p>Is FocalBCE a weighting of focal loss and bce loss? Did you have to experiment a lot with different weighting or focal params to get it to work?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 3179013,
          "author_name": "Salman Ahmed",
          "author_url": "",
          "post_date": "2025-04-14T22:20:29.267000",
          "content": "<p>It's just sum of both losses.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 3178650,
      "author_name": "Shiro",
      "author_url": "",
      "post_date": "2025-04-14T12:41:42.983000",
      "content": "<p>Quite hard to tell for me whats work or not, because i have contradictory behavior when changing seed</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3178411,
      "author_name": "noshakeplz",
      "author_url": "",
      "post_date": "2025-04-14T06:10:32.413000",
      "content": "<p>May I ask how you handled the human voice parts in the CSA recordings? Did you remove the segments containing human voices?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3178158,
      "author_name": "Devasy Patel23",
      "author_url": "",
      "post_date": "2025-04-13T19:38:22.580000",
      "content": "<p>Has anyone tried using Triplet loss for this purpose?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3177897,
      "author_name": "Araik Tamazian",
      "author_url": "",
      "post_date": "2025-04-13T12:22:00.103000",
      "content": "<p>Nice post! I can tell that usage of FocalLoss had also improved my result.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 3184672,
          "author_name": "XuKong Ji",
          "author_url": "",
          "post_date": "2025-04-22T11:05:41.143000",
          "content": "<p>May I ask did you clean your data? Because In my expriments, focal loss hurt performance in my local valid set(both loss and auc)</p>",
          "votes": 0,
          "replies": [
            {
              "id": 3187924,
              "author_name": "Araik Tamazian",
              "author_url": "",
              "post_date": "2025-04-26T17:55:29.680000",
              "content": "<p>Yes, I did.</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3177896,
      "author_name": "SeshuRaju 🧘‍♂️",
      "author_url": "",
      "post_date": "2025-04-13T12:19:24.557000",
      "content": "<p>Thanks for sharing your valuable insights <a href=\"https://www.kaggle.com/salmanahmedtamu\" target=\"_blank\">@salmanahmedtamu</a> , this will help to improve everyone's solutions.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3206043,
      "author_name": "",
      "author_url": "",
      "post_date": "2025-05-20T18:58:59.760000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 3206216,
          "author_name": "",
          "author_url": "",
          "post_date": "2025-05-21T03:29:34.477000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 3205960,
      "author_name": "",
      "author_url": "",
      "post_date": "2025-05-20T16:10:27.013000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3178659,
      "author_name": "",
      "author_url": "",
      "post_date": "2025-04-14T12:54:46.300000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 3179014,
          "author_name": "",
          "author_url": "",
          "post_date": "2025-04-14T22:20:59.657000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 3213546,
      "author_name": "",
      "author_url": "",
      "post_date": "2025-05-30T06:55:31.960000",
      "content": "",
      "votes": -1,
      "replies": []
    },
    {
      "id": 3180189,
      "author_name": "",
      "author_url": "",
      "post_date": "2025-04-16T09:14:00.430000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3179706,
      "author_name": "",
      "author_url": "",
      "post_date": "2025-04-15T16:01:35.843000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3178509,
      "author_name": "",
      "author_url": "",
      "post_date": "2025-04-14T09:20:43.810000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3182405,
      "author_name": "",
      "author_url": "",
      "post_date": "2025-04-19T08:40:40.553000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3177838": "I am sharing what worked for me in early baseline, but still haven't figured out what's working yet, still trying. \n\n1. Applying any augmentation or processing on raw waves, hurts the performance.\n2. When I look at last year's leaderboard, position # 5 was able to achieve a awesome melspec model without any pseudo labels or anything, while other top solutions used pseudo labels. I believe melspec parameters were one of the import things in that solution. (Not sure.) Considering this, I started exploring and I was able to get scores from 0.810 to 0.859 just by changing the melspec parameters. You just have to think about it, visualize and explore more. (5 subs a day to find the best parameters).\n3. I've tried BCE with multiple types of weights, but FocalBCE works best for me.\n4. Consider melspecs as images and not the waves, I converted the melspecs to images, applied normalization and augment with RandAug and RandomErazing, Time and Freq Masking + (Mixup with prob of 1.0). Remaining training pipeline is same as my public notebook from last year. [https://www.kaggle.com/code/salmanahmedtamu/training-0-65-0-66](https://www.kaggle.com/code/salmanahmedtamu/training-0-65-0-66)\n5. Another important thing, in Efficientnet models when i combined middle features and passed into the global pool that worked better than the conv_head in timm. Still Exploring.....!!!!\n6. Even simple models starts overfitting pretty easily, and you cannot track because K-Fold doesn't make sense here, your validation set is the public LB. Keep track of the Batch Size, Number of iterations, warmp up and LR. \n7. Filtering/Processing CSA recordings by Human sound helps.\n8. Look at predictions of your model and visualize the attentions on Train Soundscape, to see what's the CAM of your model.\n9. Ensemble to 0.854, 0.856, 0.858 and 0.859 results in 0.872. \n10. Post Processing is important as I shared in my previous public notebook, There are a lot of other ways to post process as well.\n11. I am sure, I missed a lot of detail, I'll try to share some public notebook if i get some time this week.\n12. Training on Random 5 seconds helped compared to first 5 seconds.",
    "3195341": "After applying TTA using efficientnet_B0 model, every time , notebook throws message as network threw exception, while submission on scoring. Can anyone guide to resolve this issue?",
    "3209120": "The posting give me some hint, appreciate!",
    "3203448": "Thanks for the insights! Did you find any efficient way to tune the melspec parameters?",
    "3180493": "Could you talk about what you mean by this a bit more? \"Filtering/Processing CSA recordings by Human sound helps.\" And also by post processing I assume you mean things like smoothing predictions, similar to that of some of the top public submission notebooks?",
    "3180253": "I have a wonder that should I change the N_MELS in the train phase same to the inference phase?\nPls reply",
    "3178087": "training is too unstable, result varies a lot",
    "3192188": "With a single EfficientNet-B0 model (single fold), I was able to get up to a 0.843 public score. I haven’t been able to push a single model beyond 0.85, but ensembling results from multiple folds gave me a score of 0.857. For training, I’m not using random 5-second crops—instead, I always use the first 5 seconds. Also, I’ve barely made use of the soundscapes data so far.",
    "3177847": "\"2. When I look at last year's leaderboard, position # 5 was able to achieve a awesome melspec model without any pseudo labels or anything, while other top solutions used pseudo labels. I believe melspec parameters were one of the import things in that solution. (Not sure.) Considering this, I started exploring and I was able to get scores from 0.810 to 0.859 just by changing the melspec parameters. You just have to think about it, visualize and explore more. (5 subs a day to find the best parameters).\" This is very promising. However, I am uncertain about the randomness of the results. In my k-fold results, one fold has an AUC of 0.817, while another fold has an AUC of 0.85. ",
    "3179843": "May I ask what is the range for Mel spec params hop length,nfft n_mels just range to guide me not a specific number off course . And for sure I will understand if you don't want to reveal this info thanks in advance @salmanahmedtamu ",
    "3210476": "thanks for notes。Does adding background noise help or not？😃",
    "3193582": ">When I look at last year's leaderboard, position # 5 was able to achieve a awesome melspec model without any pseudo labels or anything, while other top solutions used pseudo labels. I believe melspec parameters were one of the import things in that solution. (Not sure.) \n\nThe combination with the raw waveform model could also be the reason, not necessarily melspec param.",
    "3190316": "Can I ask you for how many epochs do you usually train your models?",
    "3190150": "Hi Salman! I was wondering—did you track the fixed epoch of a single model, or did you use model soup?",
    "3188397": "Hey, Sal!\nHow many in_channels did you choose? and which architecture do you use?",
    "3188064": "How did you do this Filtering/Processing CSA recordings by Human sound? I have had a ton of problems in my efforts to do the same.",
    "3186223": "Hello, I would like to ask if the random 5s is used to convert into a spectrogram, if the audio length is longer, such as 10s, will some information be lost？",
    "3185872": "Thanks for sharing! I have the similar trend—changing the melspec parameters can really impact the score. So far, I’ve managed to hit a public score of 0.836 using a single EfficientNet-B0 model (single fold), but it’s quite unstable. Even slight changes to the parameters or switching folds can cause the score to drop a lot (public score is around 0.8). How’s it going on your end?",
    "3184868": "thanks for sharing and enhancing our knowledge....",
    "3184495": "Hi, may I ask your score including TTA or any post processing to achieve this score?",
    "3183573": "Thanks for sharing.\n\nDid you use external data to achieve this .872 mentioned here?",
    "3182652": "thanks for sharing! I notice that your training pipeline didn't use pre-compute  mel spec. Will it affect your training time? My training process takes 1min40s per epoch with pre-compute mel spec, and computing 5s mel spec for every audio takes 20min. 😅",
    "3182172": "Thanks for sharing this — super valuable! \n\nYour baseline-to-0.872 journey gives a lot of actionable directions to explore. \nTreating MelSpecs as images – totally agree. Applying image-based augmentations (like RandAugment, RandomErasing, Mixup, SpecAugment) feels like a powerful combo. Did you use torchvision or albumentations for this?\n\n",
    "3181215": "did you get 0.872 without pseudo labels?",
    "3178772": "Is FocalBCE a weighting of focal loss and bce loss? Did you have to experiment a lot with different weighting or focal params to get it to work?",
    "3178650": "Quite hard to tell for me whats work or not, because i have contradictory behavior when changing seed",
    "3178411": "May I ask how you handled the human voice parts in the CSA recordings? Did you remove the segments containing human voices?",
    "3178158": "Has anyone tried using Triplet loss for this purpose?",
    "3177897": "Nice post! I can tell that usage of FocalLoss had also improved my result.",
    "3177896": "Thanks for sharing your valuable insights @salmanahmedtamu , this will help to improve everyone's solutions.",
    "3206043": "@salmanahmedtamu I know it is very stupid question but I see many inference notebook where participants using single channel. Is it something like that we have to train with 3 channel (lets assume) and inference on single channel.",
    "3205960": "",
    "3178659": "I was able to get scores from 0.810 to 0.859 just by changing the melspec parameters. You just have to think about it, visualize and explore more. (5 subs a day to find the best parameters). @salmanahmedtamu did you change mels params for both train and inference or just in inference ",
    "3213546": "Thanks for sharing!",
    "3180189": "Thanks for your sharing !!!",
    "3179706": "THANKS. FOR SHARING ",
    "3178509": "thanks! very helpful",
    "3182405": ""
  }
}