{
  "id": 579109,
  "title": "What the keypoints of humvoice removing?",
  "url": "/competitions/birdclef-2025/discussion/579109",
  "author_name": "",
  "post_date": "2025-05-15T09:11:43.780652300Z",
  "votes": 3,
  "comment_count": 3,
  "views": 0,
  "content": "<p>I tried to remove all hum voice in the audio, but this hurt pb score badly. So I think the testset containing hum voice too  in the leaderboard. To keep a same distribution and filter the dirty hum voice at the same time, I keep doing nothing  with the audio in a little probability when training, And this method finally benefit. Any better methods from you guys?</p>",
  "messages": [
    {
      "id": "3202355",
      "postDate": "05/15/2025 09:11:43",
      "content": "<p>I tried to remove all hum voice in the audio, but this hurt pb score badly. So I think the testset containing hum voice too  in the leaderboard. To keep a same distribution and filter the dirty hum voice at the same time, I keep doing nothing  with the audio in a little probability when training, And this method finally benefit. Any better methods from you guys?</p>",
      "rawMarkdown": "I tried to remove all hum voice in the audio, but this hurt pb score badly. So I think the testset containing hum voice too  in the leaderboard. To keep a same distribution and filter the dirty hum voice at the same time, I keep doing nothing  with the audio in a little probability when training, And this method finally benefit. Any better methods from you guys?",
      "votes": null
    },
    {
      "id": "3202402",
      "postDate": "05/15/2025 10:15:30",
      "content": "<p>I just use the first 30s or 60s as training data if file is CSA format. Because in the voice files with human voice, CSA format will have more serious human voice. I also find that removing human voice works better in SED model than CNN model.</p>",
      "rawMarkdown": "I just use the first 30s or 60s as training data if file is CSA format. Because in the voice files with human voice, CSA format will have more serious human voice. I also find that removing human voice works better in SED model than CNN model.",
      "votes": null
    },
    {
      "id": "3202840",
      "postDate": "05/16/2025 02:23:34",
      "content": "<p>Thanks a lot!</p>",
      "rawMarkdown": "Thanks a lot!",
      "votes": null
    },
    {
      "id": "3203277",
      "postDate": "05/16/2025 14:19:04",
      "content": "<p>There is SileroVAD discussion, you may take a look: <a href=\"https://www.kaggle.com/competitions/birdclef-2025/discussion/568886\" target=\"_blank\">https://www.kaggle.com/competitions/birdclef-2025/discussion/568886</a></p>\n<p>Not only CSA files have human voice. </p>",
      "rawMarkdown": "There is SileroVAD discussion, you may take a look: https://www.kaggle.com/competitions/birdclef-2025/discussion/568886\n\nNot only CSA files have human voice.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3202402,
      "author_name": "i2nfinit3y",
      "author_url": "",
      "post_date": "05/15/2025 10:15:30",
      "content": "<p>I just use the first 30s or 60s as training data if file is CSA format. Because in the voice files with human voice, CSA format will have more serious human voice. I also find that removing human voice works better in SED model than CNN model.</p>",
      "votes": null,
      "replies": [
        {
          "id": 3202840,
          "author_name": "xukongji",
          "author_url": "",
          "post_date": "05/16/2025 02:23:34",
          "content": "<p>Thanks a lot!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3203277,
      "author_name": "alexandergremyakov",
      "author_url": "",
      "post_date": "05/16/2025 14:19:04",
      "content": "<p>There is SileroVAD discussion, you may take a look: <a href=\"https://www.kaggle.com/competitions/birdclef-2025/discussion/568886\" target=\"_blank\">https://www.kaggle.com/competitions/birdclef-2025/discussion/568886</a></p>\n<p>Not only CSA files have human voice. </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3202355": "I tried to remove all hum voice in the audio, but this hurt pb score badly. So I think the testset containing hum voice too  in the leaderboard. To keep a same distribution and filter the dirty hum voice at the same time, I keep doing nothing  with the audio in a little probability when training, And this method finally benefit. Any better methods from you guys?",
    "3202402": "I just use the first 30s or 60s as training data if file is CSA format. Because in the voice files with human voice, CSA format will have more serious human voice. I also find that removing human voice works better in SED model than CNN model.",
    "3202840": "Thanks a lot!",
    "3203277": "There is SileroVAD discussion, you may take a look: https://www.kaggle.com/competitions/birdclef-2025/discussion/568886\n\nNot only CSA files have human voice."
  },
  "source": "meta"
}