{
  "id": 582330,
  "title": "Something test on train_soundscapes",
  "url": "/competitions/birdclef-2025/discussion/582330",
  "author_name": "",
  "post_date": "2025-05-30T14:02:57.822989Z",
  "votes": null,
  "comment_count": 6,
  "views": 0,
  "content": "<p><strong>1.Category Imbalance</strong> <br>\nI use my ensemble model(3CNN + 1 SED) to generate Pseudo labels,and select top 10% of those with the lowest class label entropy,it seems category imbalance.For me,<br>\n<code>np.unique(np.argmax(labels)) = 116</code><br>\nAnyone also try it?Can u share?<br>\n<strong>2.SED model fitting to noise</strong></p>\n<pre><code>\nresult = ((result1 + result2 + result3)/ + result4)/\nmax_label = np.argmax(result)\n(max_label,result[max_label],result4[max_label])\n</code></pre>\n<pre><code>\n  \n  \n  \n  \n  \n  \n  \n  \n  \n  \n  \n  \n\n  \n  \n  \n  \n  \n  \n  \n  \n  \n  \n  \n  \n</code></pre>\n<p>For label 35,you can see that my SED model's prediction is always twice than the ensemble model<br>\n(means that all probabilities of the ensemble model are provided by sed). <br>\nSo, I listened to these two audio files, which are almost composed of noise/rain sound, indicating that my SED model treats noise as label 35.</p>",
  "messages": [
    {
      "id": "3213818",
      "postDate": "05/30/2025 14:02:57",
      "content": "<p><strong>1.Category Imbalance</strong> <br>\nI use my ensemble model(3CNN + 1 SED) to generate Pseudo labels,and select top 10% of those with the lowest class label entropy,it seems category imbalance.For me,<br>\n<code>np.unique(np.argmax(labels)) = 116</code><br>\nAnyone also try it?Can u share?<br>\n<strong>2.SED model fitting to noise</strong></p>\n<pre><code>\nresult = ((result1 + result2 + result3)/ + result4)/\nmax_label = np.argmax(result)\n(max_label,result[max_label],result4[max_label])\n</code></pre>\n<pre><code>\n  \n  \n  \n  \n  \n  \n  \n  \n  \n  \n  \n  \n\n  \n  \n  \n  \n  \n  \n  \n  \n  \n  \n  \n  \n</code></pre>\n<p>For label 35,you can see that my SED model's prediction is always twice than the ensemble model<br>\n(means that all probabilities of the ensemble model are provided by sed). <br>\nSo, I listened to these two audio files, which are almost composed of noise/rain sound, indicating that my SED model treats noise as label 35.</p>",
      "rawMarkdown": "**1.Category Imbalance** \nI use my ensemble model(3CNN + 1 SED) to generate Pseudo labels,and select top 10% of those with the lowest class label entropy,it seems category imbalance.For me,\n` np.unique(np.argmax(labels)) = 116`\nAnyone also try it?Can u share?\n**2.SED model fitting to noise**\n```python\n#result1-3 means 3 CNN models' predictions, result4 means SED model's prediction\nresult = ((result1 + result2 + result3)/3 + result4)/2\nmax_label = np.argmax(result)\nprint(max_label,result[max_label],result4[max_label])\n\n```\n```python\n#/kaggle/input/birdclef-2025/train_soundscapes/H02_20230420_074000.ogg\n35 0.4110214 0.8112731\n35 0.38948205 0.7771356\n35 0.40969712 0.81655526\n35 0.48448178 0.76775\n35 0.39543396 0.7653094\n61 0.20838024 0.37686992\n35 0.4652759 0.89501786\n35 0.39653823 0.7902652\n35 0.43820482 0.8678603\n35 0.30535662 0.6081055\n35 0.38768384 0.77417016\n35 0.43872216 0.83240676\n#/kaggle/input/birdclef-2025/train_soundscapes/H02_20230421_113500.ogg\n35 0.3967473 0.7933539\n35 0.42145443 0.8409543\n35 0.39631504 0.79147595\n35 0.37850812 0.75387836\n35 0.3453101 0.68904823\n35 0.38385478 0.7669082\n35 0.43560475 0.8707274\n35 0.38267577 0.76395994\n35 0.39830032 0.7962671\n35 0.37669384 0.7524742\n35 0.4310157 0.86051077\n35 0.3688453 0.73715603\n```\nFor label 35,you can see that my SED model's prediction is always twice than the ensemble model\n(means that all probabilities of the ensemble model are provided by sed). \nSo, I listened to these two audio files, which are almost composed of noise/rain sound, indicating that my SED model treats noise as label 35.",
      "votes": null
    },
    {
      "id": "3213824",
      "postDate": "05/30/2025 14:12:57",
      "content": "<p>by the way,i try to replace my ensemble model's predictions of label 35 to my CNN models.The methond shows same scores.</p>",
      "rawMarkdown": "by the way,i try to replace my ensemble model's predictions of label 35 to my CNN models.The methond shows same scores.",
      "votes": null
    },
    {
      "id": "3214559",
      "postDate": "05/31/2025 19:14:50",
      "content": "<p>Regarding the class imbalance in train_soundscape, I believe it is expected. Experiments(<a href=\"url\" target=\"_blank\">https://www.kaggle.com/competitions/birdclef-2025/discussion/570837</a>) have confirmed that the public test set contains only 70–100 classes, so <code>np.unique(np.argmax(labels)) = 116</code> falls within a reasonable range.</p>\n<p>Furthermore, considering that this competition focuses on the AUC metric, the argmax(labels) for pseudo labels may not hold that much value. The model only needs to capture the ranking order within each column rather than the exact class predictions.</p>\n<p>For noise-fitting phenomenon, I also discovered a similar case:</p>\n<p>In my experiment, whether using my CNN model or SED model, showed a strong tendency to predict noise as category ID 21211 in train_soundscape.</p>\n<p>However, based on my LB probing results (where I set all predictions for category 21211 to zero), this category did not appear in the public test set. Therefore, I believe that for this specific category, it essentially serves as a nocall label.</p>",
      "rawMarkdown": "Regarding the class imbalance in train_soundscape, I believe it is expected. Experiments([https://www.kaggle.com/competitions/birdclef-2025/discussion/570837](url)) have confirmed that the public test set contains only 70–100 classes, so `np.unique(np.argmax(labels)) = 116` falls within a reasonable range.\n\nFurthermore, considering that this competition focuses on the AUC metric, the argmax(labels) for pseudo labels may not hold that much value. The model only needs to capture the ranking order within each column rather than the exact class predictions.\n\nFor noise-fitting phenomenon, I also discovered a similar case:\n\nIn my experiment, whether using my CNN model or SED model, showed a strong tendency to predict noise as category ID 21211 in train_soundscape.\n\nHowever, based on my LB probing results (where I set all predictions for category 21211 to zero), this category did not appear in the public test set. Therefore, I believe that for this specific category, it essentially serves as a nocall label.",
      "votes": null
    },
    {
      "id": "3214605",
      "postDate": "05/31/2025 20:34:06",
      "content": "<p>Thanks for sharing. Are you sure about this? File 21211/iNat361222.ogg for instance sounds like there is a target there. <br>\nAlso the taxonomy has a species defined:</p>\n<p>21211, Allobates femoralis, Spotted-thighed Poison Frog, Amphibia</p>\n<p>In general Amphibia samples are very small, so suspect they are witholding  from the public test set ( if the PB leaderboard was randomly sampled, I would expect alot of species from this class to be missing…). </p>\n<p>A better approach IMO would be for them to have  a public LB that was stratified by class…</p>",
      "rawMarkdown": "Thanks for sharing. Are you sure about this? File 21211/iNat361222.ogg for instance sounds like there is a target there. \nAlso the taxonomy has a species defined:\n\n21211, Allobates femoralis, Spotted-thighed Poison Frog, Amphibia\n\nIn general Amphibia samples are very small, so suspect they are witholding  from the public test set ( if the PB leaderboard was randomly sampled, I would expect alot of species from this class to be missing...). \n\nA better approach IMO would be for them to have  a public LB that was stratified by class...",
      "votes": null
    },
    {
      "id": "3214613",
      "postDate": "05/31/2025 21:00:32",
      "content": "<p>Perhaps my wording was ambiguous.</p>\n<p>There are indeed instances of this category in the dataset, but for some unclear reason, my model tends to classify noise as 21211 in the soundscape data.</p>\n<p>However, considering that this category does not exist in the public test set, from the perspective of public AUC, 21211 can be regarded as a nocall label.</p>\n<p>In terms of improving public-private correlation, I agree that stratified splitting would be more reasonable. However, this might also increase the risk of being probed. Therefore, random splitting seems to be a balanced approach.</p>",
      "rawMarkdown": "Perhaps my wording was ambiguous.\n\nThere are indeed instances of this category in the dataset, but for some unclear reason, my model tends to classify noise as 21211 in the soundscape data.\n\nHowever, considering that this category does not exist in the public test set, from the perspective of public AUC, 21211 can be regarded as a nocall label.\n\nIn terms of improving public-private correlation, I agree that stratified splitting would be more reasonable. However, this might also increase the risk of being probed. Therefore, random splitting seems to be a balanced approach.",
      "votes": null
    },
    {
      "id": "3218195",
      "postDate": "06/06/2025 01:04:19",
      "content": "<p>Some of the species may not show up on leaderboard and that's the reason.</p>",
      "rawMarkdown": "Some of the species may not show up on leaderboard and that's the reason.",
      "votes": null
    },
    {
      "id": "3218201",
      "postDate": "06/06/2025 01:13:54",
      "content": "<p>Does this happen on private lb as well? Or only on public?</p>",
      "rawMarkdown": "Does this happen on private lb as well? Or only on public?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3213824,
      "author_name": "minatoyukinaxlisa",
      "author_url": "",
      "post_date": "05/30/2025 14:12:57",
      "content": "<p>by the way,i try to replace my ensemble model's predictions of label 35 to my CNN models.The methond shows same scores.</p>",
      "votes": null,
      "replies": [
        {
          "id": 3218195,
          "author_name": "shanzhong8",
          "author_url": "",
          "post_date": "06/06/2025 01:04:19",
          "content": "<p>Some of the species may not show up on leaderboard and that's the reason.</p>",
          "votes": null,
          "replies": [
            {
              "id": 3218201,
              "author_name": "ziyi777",
              "author_url": "",
              "post_date": "06/06/2025 01:13:54",
              "content": "<p>Does this happen on private lb as well? Or only on public?</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3214559,
      "author_name": "ryenhails",
      "author_url": "",
      "post_date": "05/31/2025 19:14:50",
      "content": "<p>Regarding the class imbalance in train_soundscape, I believe it is expected. Experiments(<a href=\"url\" target=\"_blank\">https://www.kaggle.com/competitions/birdclef-2025/discussion/570837</a>) have confirmed that the public test set contains only 70–100 classes, so <code>np.unique(np.argmax(labels)) = 116</code> falls within a reasonable range.</p>\n<p>Furthermore, considering that this competition focuses on the AUC metric, the argmax(labels) for pseudo labels may not hold that much value. The model only needs to capture the ranking order within each column rather than the exact class predictions.</p>\n<p>For noise-fitting phenomenon, I also discovered a similar case:</p>\n<p>In my experiment, whether using my CNN model or SED model, showed a strong tendency to predict noise as category ID 21211 in train_soundscape.</p>\n<p>However, based on my LB probing results (where I set all predictions for category 21211 to zero), this category did not appear in the public test set. Therefore, I believe that for this specific category, it essentially serves as a nocall label.</p>",
      "votes": null,
      "replies": [
        {
          "id": 3214605,
          "author_name": "ikarosilva",
          "author_url": "",
          "post_date": "05/31/2025 20:34:06",
          "content": "<p>Thanks for sharing. Are you sure about this? File 21211/iNat361222.ogg for instance sounds like there is a target there. <br>\nAlso the taxonomy has a species defined:</p>\n<p>21211, Allobates femoralis, Spotted-thighed Poison Frog, Amphibia</p>\n<p>In general Amphibia samples are very small, so suspect they are witholding  from the public test set ( if the PB leaderboard was randomly sampled, I would expect alot of species from this class to be missing…). </p>\n<p>A better approach IMO would be for them to have  a public LB that was stratified by class…</p>",
          "votes": null,
          "replies": [
            {
              "id": 3214613,
              "author_name": "ryenhails",
              "author_url": "",
              "post_date": "05/31/2025 21:00:32",
              "content": "<p>Perhaps my wording was ambiguous.</p>\n<p>There are indeed instances of this category in the dataset, but for some unclear reason, my model tends to classify noise as 21211 in the soundscape data.</p>\n<p>However, considering that this category does not exist in the public test set, from the perspective of public AUC, 21211 can be regarded as a nocall label.</p>\n<p>In terms of improving public-private correlation, I agree that stratified splitting would be more reasonable. However, this might also increase the risk of being probed. Therefore, random splitting seems to be a balanced approach.</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3213818": "**1.Category Imbalance** \nI use my ensemble model(3CNN + 1 SED) to generate Pseudo labels,and select top 10% of those with the lowest class label entropy,it seems category imbalance.For me,\n` np.unique(np.argmax(labels)) = 116`\nAnyone also try it?Can u share?\n**2.SED model fitting to noise**\n```python\n#result1-3 means 3 CNN models' predictions, result4 means SED model's prediction\nresult = ((result1 + result2 + result3)/3 + result4)/2\nmax_label = np.argmax(result)\nprint(max_label,result[max_label],result4[max_label])\n\n```\n```python\n#/kaggle/input/birdclef-2025/train_soundscapes/H02_20230420_074000.ogg\n35 0.4110214 0.8112731\n35 0.38948205 0.7771356\n35 0.40969712 0.81655526\n35 0.48448178 0.76775\n35 0.39543396 0.7653094\n61 0.20838024 0.37686992\n35 0.4652759 0.89501786\n35 0.39653823 0.7902652\n35 0.43820482 0.8678603\n35 0.30535662 0.6081055\n35 0.38768384 0.77417016\n35 0.43872216 0.83240676\n#/kaggle/input/birdclef-2025/train_soundscapes/H02_20230421_113500.ogg\n35 0.3967473 0.7933539\n35 0.42145443 0.8409543\n35 0.39631504 0.79147595\n35 0.37850812 0.75387836\n35 0.3453101 0.68904823\n35 0.38385478 0.7669082\n35 0.43560475 0.8707274\n35 0.38267577 0.76395994\n35 0.39830032 0.7962671\n35 0.37669384 0.7524742\n35 0.4310157 0.86051077\n35 0.3688453 0.73715603\n```\nFor label 35,you can see that my SED model's prediction is always twice than the ensemble model\n(means that all probabilities of the ensemble model are provided by sed). \nSo, I listened to these two audio files, which are almost composed of noise/rain sound, indicating that my SED model treats noise as label 35.",
    "3213824": "by the way,i try to replace my ensemble model's predictions of label 35 to my CNN models.The methond shows same scores.",
    "3214559": "Regarding the class imbalance in train_soundscape, I believe it is expected. Experiments([https://www.kaggle.com/competitions/birdclef-2025/discussion/570837](url)) have confirmed that the public test set contains only 70–100 classes, so `np.unique(np.argmax(labels)) = 116` falls within a reasonable range.\n\nFurthermore, considering that this competition focuses on the AUC metric, the argmax(labels) for pseudo labels may not hold that much value. The model only needs to capture the ranking order within each column rather than the exact class predictions.\n\nFor noise-fitting phenomenon, I also discovered a similar case:\n\nIn my experiment, whether using my CNN model or SED model, showed a strong tendency to predict noise as category ID 21211 in train_soundscape.\n\nHowever, based on my LB probing results (where I set all predictions for category 21211 to zero), this category did not appear in the public test set. Therefore, I believe that for this specific category, it essentially serves as a nocall label.",
    "3214605": "Thanks for sharing. Are you sure about this? File 21211/iNat361222.ogg for instance sounds like there is a target there. \nAlso the taxonomy has a species defined:\n\n21211, Allobates femoralis, Spotted-thighed Poison Frog, Amphibia\n\nIn general Amphibia samples are very small, so suspect they are witholding  from the public test set ( if the PB leaderboard was randomly sampled, I would expect alot of species from this class to be missing...). \n\nA better approach IMO would be for them to have  a public LB that was stratified by class...",
    "3214613": "Perhaps my wording was ambiguous.\n\nThere are indeed instances of this category in the dataset, but for some unclear reason, my model tends to classify noise as 21211 in the soundscape data.\n\nHowever, considering that this category does not exist in the public test set, from the perspective of public AUC, 21211 can be regarded as a nocall label.\n\nIn terms of improving public-private correlation, I agree that stratified splitting would be more reasonable. However, this might also increase the risk of being probed. Therefore, random splitting seems to be a balanced approach.",
    "3218195": "Some of the species may not show up on leaderboard and that's the reason.",
    "3218201": "Does this happen on private lb as well? Or only on public?"
  },
  "source": "meta"
}