{
  "id": 572159,
  "title": "Worth it this year?",
  "url": "/competitions/birdclef-2025/discussion/572159",
  "author_name": "Cody_Null",
  "post_date": "2025-04-08T01:17:04.631000",
  "votes": 10,
  "comment_count": 11,
  "views": 0,
  "content": "<p>Year over year I know there have been complications with validation in this comp (cv/lb gaps) from the outside looking in it seems a bit better this year? For those in the competition, how has your experience been with correlation/consistency with CV/LB. Or is it the standard instability of better fold AUC but worse LB even between epochs of the same model for the same training run?</p>",
  "messages": [
    {
      "id": 3173484,
      "postDate": "2025-04-08T01:17:04.633Z",
      "content": "<p>Year over year I know there have been complications with validation in this comp (cv/lb gaps) from the outside looking in it seems a bit better this year? For those in the competition, how has your experience been with correlation/consistency with CV/LB. Or is it the standard instability of better fold AUC but worse LB even between epochs of the same model for the same training run?</p>",
      "rawMarkdown": "Year over year I know there have been complications with validation in this comp (cv/lb gaps) from the outside looking in it seems a bit better this year? For those in the competition, how has your experience been with correlation/consistency with CV/LB. Or is it the standard instability of better fold AUC but worse LB even between epochs of the same model for the same training run?",
      "votes": 10
    },
    {
      "id": 3173692,
      "postDate": "2025-04-08T08:12:11.553Z",
      "content": "<p>Same issues as last year from what I've seen. AUCs quickly reaching 0.97+, unstable LB scores.</p>\n<p>The novelty this year is that we have more classes with only very few training samples, which is not good for stability either. You can't really validate classes with &lt;5 samples imo</p>",
      "rawMarkdown": "Same issues as last year from what I've seen. AUCs quickly reaching 0.97+, unstable LB scores.\n\nThe novelty this year is that we have more classes with only very few training samples, which is not good for stability either. You can't really validate classes with <5 samples imo",
      "votes": 6,
      "replies": [
        {
          "id": 3174672,
          "postDate": "2025-04-09T10:31:26.423Z",
          "content": "<p>This time we have not only birds, but other species as well.</p>",
          "rawMarkdown": "This time we have not only birds, but other species as well.",
          "replies": [
            {
              "id": 3201004,
              "postDate": "2025-05-13T10:35:27.777Z",
              "rawMarkdown": "",
              "isDeleted": true
            }
          ]
        }
      ]
    },
    {
      "id": 3177150,
      "postDate": "2025-04-12T10:34:28.127Z",
      "content": "<p>Hi everyone, interesting discussion on single model performance!</p>\n<p>I'm feeling a bit stuck myself. I've been working off a baseline and haven't managed to push the score much beyond that yet (around the low 0.8x range). I tried a few things like using a 3-channel input (mel spec + MFCCs + delta MFCCs) but that didn't seem to help unfortunately.</p>\n<p>Looking at the summary of past top solutions (like the one I compiled <a href=\"https://www.kaggle.com/competitions/birdclef-2025/discussion/572928)\" target=\"_blank\">https://www.kaggle.com/competitions/birdclef-2025/discussion/572928)</a>, it's clear that pseudo-labeling on the unlabeled soundscapes was almost universally beneficial last year. Handling secondary labels effectively also seems important.</p>\n<p>My plan is to try implementing pseudo-labeling next, and then explore ideas like weighted loss for secondary labels or maybe even triplet loss if I get adventurous.</p>\n<p>Given that I'm currently plateauing with the basics, where would you suggest focusing efforts right now for the potentially biggest impact? Is diving straight into pseudo-labeling the best next step, or should I perhaps revisit data augmentation, model architecture choices, or something else first?</p>\n<p>Appreciate any insights from those who've managed to break through similar plateaus!</p>",
      "rawMarkdown": "Hi everyone, interesting discussion on single model performance!\n\nI'm feeling a bit stuck myself. I've been working off a baseline and haven't managed to push the score much beyond that yet (around the low 0.8x range). I tried a few things like using a 3-channel input (mel spec + MFCCs + delta MFCCs) but that didn't seem to help unfortunately.\n\nLooking at the summary of past top solutions (like the one I compiled https://www.kaggle.com/competitions/birdclef-2025/discussion/572928), it's clear that pseudo-labeling on the unlabeled soundscapes was almost universally beneficial last year. Handling secondary labels effectively also seems important.\n\nMy plan is to try implementing pseudo-labeling next, and then explore ideas like weighted loss for secondary labels or maybe even triplet loss if I get adventurous.\n\nGiven that I'm currently plateauing with the basics, where would you suggest focusing efforts right now for the potentially biggest impact? Is diving straight into pseudo-labeling the best next step, or should I perhaps revisit data augmentation, model architecture choices, or something else first?\n\nAppreciate any insights from those who've managed to break through similar plateaus!",
      "votes": 4
    },
    {
      "id": 3173591,
      "postDate": "2025-04-08T05:50:18.593Z",
      "content": "<p>I truly want to know whether the categories of private and public are consistent. If there are minor differences in the categories, I can accept this phenomenon. However, if there are significant differences in category distribution, the reference value of public would be very limited.</p>",
      "rawMarkdown": "I truly want to know whether the categories of private and public are consistent. If there are minor differences in the categories, I can accept this phenomenon. However, if there are significant differences in category distribution, the reference value of public would be very limited.",
      "votes": 1,
      "replies": [
        {
          "id": 3173630,
          "postDate": "2025-04-08T06:53:03.137Z",
          "content": "<p>Private and public LB use a truly random split of the hidden test data - so we expect the distribution to be very similar (aside from a few random biases).</p>",
          "rawMarkdown": "Private and public LB use a truly random split of the hidden test data - so we expect the distribution to be very similar (aside from a few random biases).",
          "votes": 19
        }
      ]
    },
    {
      "id": 3175113,
      "postDate": "2025-04-09T18:22:28.030Z",
      "content": "<p>Although I have last participated in audio competition 2-3 years back the thing that seem problematic here is that certain classes have very few number of instances &lt;= 5 in training data this makes it difficult for training and validation on those classes the validation scores reaches quickly to 0.96+ if we stratify based on primary label . </p>",
      "rawMarkdown": "Although I have last participated in audio competition 2-3 years back the thing that seem problematic here is that certain classes have very few number of instances <= 5 in training data this makes it difficult for training and validation on those classes the validation scores reaches quickly to 0.96+ if we stratify based on primary label . "
    },
    {
      "id": 3174736,
      "postDate": "2025-04-09T12:21:02.367Z",
      "content": "<p>I see my AUC quickly rising, but there is a correlation between cv and lb scoring when using the same pipeline with minor adjustments. So this is some good news, I guess.</p>",
      "rawMarkdown": "I see my AUC quickly rising, but there is a correlation between cv and lb scoring when using the same pipeline with minor adjustments. So this is some good news, I guess.",
      "replies": [
        {
          "id": 3174749,
          "postDate": "2025-04-09T12:44:59.460Z",
          "content": "<p>From the outside looking in, it seems unstable. I saw some comments about scores raising between epochs but not LB. They had used averaging between gradients and still experience this. Just unfortunate we could not just point to the validation dataset and optimize for that </p>",
          "rawMarkdown": "From the outside looking in, it seems unstable. I saw some comments about scores raising between epochs but not LB. They had used averaging between gradients and still experience this. Just unfortunate we could not just point to the validation dataset and optimize for that ",
          "votes": 1,
          "replies": [
            {
              "id": 3174772,
              "postDate": "2025-04-09T13:00:50.537Z",
              "content": "<p>Depends on what you tune I guess. I have some strange behaviour as well. But for now it's the only active not too computationally intensive comp at this moment.</p>",
              "rawMarkdown": "Depends on what you tune I guess. I have some strange behaviour as well. But for now it's the only active not too computationally intensive comp at this moment."
            },
            {
              "id": 3174815,
              "postDate": "2025-04-09T13:26:05.970Z",
              "content": "<p>My experience so far (haven't participated in earlier years): </p>\n<p>Downloaded a few datasets from BirdSet (<a href=\"https://huggingface.co/datasets/DBD-research-group/BirdSet);\" target=\"_blank\">https://huggingface.co/datasets/DBD-research-group/BirdSet);</a> they have the same format as this competition, just different locations, looks ideal for the CV . So surely if I can find something that works consistently over all of those datasets, it should translate to this competition as well, I naively thought. </p>\n<p>Well, no. Realistic looking CV scores, found a consistent 4-5% AUC point increments over simple baseline model for all CV datasets… but public LB score -5%.</p>",
              "rawMarkdown": "My experience so far (haven't participated in earlier years): \n\nDownloaded a few datasets from BirdSet (https://huggingface.co/datasets/DBD-research-group/BirdSet); they have the same format as this competition, just different locations, looks ideal for the CV . So surely if I can find something that works consistently over all of those datasets, it should translate to this competition as well, I naively thought. \n\nWell, no. Realistic looking CV scores, found a consistent 4-5% AUC point increments over simple baseline model for all CV datasets... but public LB score -5%.",
              "votes": 1
            }
          ]
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 3173692,
      "author_name": "Theo Viel",
      "author_url": "",
      "post_date": "2025-04-08T08:12:11.553000",
      "content": "<p>Same issues as last year from what I've seen. AUCs quickly reaching 0.97+, unstable LB scores.</p>\n<p>The novelty this year is that we have more classes with only very few training samples, which is not good for stability either. You can't really validate classes with &lt;5 samples imo</p>",
      "votes": 6,
      "replies": [
        {
          "id": 3174672,
          "author_name": "Araik Tamazian",
          "author_url": "",
          "post_date": "2025-04-09T10:31:26.423000",
          "content": "<p>This time we have not only birds, but other species as well.</p>",
          "votes": 0,
          "replies": [
            {
              "id": 3201004,
              "author_name": "",
              "author_url": "",
              "post_date": "2025-05-13T10:35:27.777000",
              "content": "",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3177150,
      "author_name": "Devasy Patel23",
      "author_url": "",
      "post_date": "2025-04-12T10:34:28.127000",
      "content": "<p>Hi everyone, interesting discussion on single model performance!</p>\n<p>I'm feeling a bit stuck myself. I've been working off a baseline and haven't managed to push the score much beyond that yet (around the low 0.8x range). I tried a few things like using a 3-channel input (mel spec + MFCCs + delta MFCCs) but that didn't seem to help unfortunately.</p>\n<p>Looking at the summary of past top solutions (like the one I compiled <a href=\"https://www.kaggle.com/competitions/birdclef-2025/discussion/572928)\" target=\"_blank\">https://www.kaggle.com/competitions/birdclef-2025/discussion/572928)</a>, it's clear that pseudo-labeling on the unlabeled soundscapes was almost universally beneficial last year. Handling secondary labels effectively also seems important.</p>\n<p>My plan is to try implementing pseudo-labeling next, and then explore ideas like weighted loss for secondary labels or maybe even triplet loss if I get adventurous.</p>\n<p>Given that I'm currently plateauing with the basics, where would you suggest focusing efforts right now for the potentially biggest impact? Is diving straight into pseudo-labeling the best next step, or should I perhaps revisit data augmentation, model architecture choices, or something else first?</p>\n<p>Appreciate any insights from those who've managed to break through similar plateaus!</p>",
      "votes": 4,
      "replies": []
    },
    {
      "id": 3173591,
      "author_name": "tanxxx",
      "author_url": "",
      "post_date": "2025-04-08T05:50:18.593000",
      "content": "<p>I truly want to know whether the categories of private and public are consistent. If there are minor differences in the categories, I can accept this phenomenon. However, if there are significant differences in category distribution, the reference value of public would be very limited.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 3173630,
          "author_name": "Stefan Kahl",
          "author_url": "",
          "post_date": "2025-04-08T06:53:03.137000",
          "content": "<p>Private and public LB use a truly random split of the hidden test data - so we expect the distribution to be very similar (aside from a few random biases).</p>",
          "votes": 19,
          "replies": []
        }
      ]
    },
    {
      "id": 3175113,
      "author_name": "Athar Sayed",
      "author_url": "",
      "post_date": "2025-04-09T18:22:28.030000",
      "content": "<p>Although I have last participated in audio competition 2-3 years back the thing that seem problematic here is that certain classes have very few number of instances &lt;= 5 in training data this makes it difficult for training and validation on those classes the validation scores reaches quickly to 0.96+ if we stratify based on primary label . </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3174736,
      "author_name": "stefanoclss",
      "author_url": "",
      "post_date": "2025-04-09T12:21:02.367000",
      "content": "<p>I see my AUC quickly rising, but there is a correlation between cv and lb scoring when using the same pipeline with minor adjustments. So this is some good news, I guess.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 3174749,
          "author_name": "Cody_Null",
          "author_url": "",
          "post_date": "2025-04-09T12:44:59.460000",
          "content": "<p>From the outside looking in, it seems unstable. I saw some comments about scores raising between epochs but not LB. They had used averaging between gradients and still experience this. Just unfortunate we could not just point to the validation dataset and optimize for that </p>",
          "votes": 1,
          "replies": [
            {
              "id": 3174772,
              "author_name": "stefanoclss",
              "author_url": "",
              "post_date": "2025-04-09T13:00:50.537000",
              "content": "<p>Depends on what you tune I guess. I have some strange behaviour as well. But for now it's the only active not too computationally intensive comp at this moment.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3174815,
              "author_name": "Herra Huu",
              "author_url": "",
              "post_date": "2025-04-09T13:26:05.970000",
              "content": "<p>My experience so far (haven't participated in earlier years): </p>\n<p>Downloaded a few datasets from BirdSet (<a href=\"https://huggingface.co/datasets/DBD-research-group/BirdSet);\" target=\"_blank\">https://huggingface.co/datasets/DBD-research-group/BirdSet);</a> they have the same format as this competition, just different locations, looks ideal for the CV . So surely if I can find something that works consistently over all of those datasets, it should translate to this competition as well, I naively thought. </p>\n<p>Well, no. Realistic looking CV scores, found a consistent 4-5% AUC point increments over simple baseline model for all CV datasets… but public LB score -5%.</p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3173484": "Year over year I know there have been complications with validation in this comp (cv/lb gaps) from the outside looking in it seems a bit better this year? For those in the competition, how has your experience been with correlation/consistency with CV/LB. Or is it the standard instability of better fold AUC but worse LB even between epochs of the same model for the same training run?",
    "3173692": "Same issues as last year from what I've seen. AUCs quickly reaching 0.97+, unstable LB scores.\n\nThe novelty this year is that we have more classes with only very few training samples, which is not good for stability either. You can't really validate classes with <5 samples imo",
    "3177150": "Hi everyone, interesting discussion on single model performance!\n\nI'm feeling a bit stuck myself. I've been working off a baseline and haven't managed to push the score much beyond that yet (around the low 0.8x range). I tried a few things like using a 3-channel input (mel spec + MFCCs + delta MFCCs) but that didn't seem to help unfortunately.\n\nLooking at the summary of past top solutions (like the one I compiled https://www.kaggle.com/competitions/birdclef-2025/discussion/572928), it's clear that pseudo-labeling on the unlabeled soundscapes was almost universally beneficial last year. Handling secondary labels effectively also seems important.\n\nMy plan is to try implementing pseudo-labeling next, and then explore ideas like weighted loss for secondary labels or maybe even triplet loss if I get adventurous.\n\nGiven that I'm currently plateauing with the basics, where would you suggest focusing efforts right now for the potentially biggest impact? Is diving straight into pseudo-labeling the best next step, or should I perhaps revisit data augmentation, model architecture choices, or something else first?\n\nAppreciate any insights from those who've managed to break through similar plateaus!",
    "3173591": "I truly want to know whether the categories of private and public are consistent. If there are minor differences in the categories, I can accept this phenomenon. However, if there are significant differences in category distribution, the reference value of public would be very limited.",
    "3175113": "Although I have last participated in audio competition 2-3 years back the thing that seem problematic here is that certain classes have very few number of instances <= 5 in training data this makes it difficult for training and validation on those classes the validation scores reaches quickly to 0.96+ if we stratify based on primary label . ",
    "3174736": "I see my AUC quickly rising, but there is a correlation between cv and lb scoring when using the same pipeline with minor adjustments. So this is some good news, I guess."
  }
}