{
  "id": 574772,
  "title": "Voice removing from audiofiles",
  "url": "/competitions/birdclef-2025/discussion/574772",
  "author_name": "",
  "post_date": "2025-04-23T21:10:42.696295900Z",
  "votes": 2,
  "comment_count": 8,
  "views": 0,
  "content": "<p>In continuation of the topic about voice detecting in train samples (<a href=\"https://www.kaggle.com/competitions/birdclef-2025/discussion/568886)\" target=\"_blank\">https://www.kaggle.com/competitions/birdclef-2025/discussion/568886)</a>.</p>\n<p>I would like to know what useful methods and DL models do you use to remove human voice or make it more noteless?<br>\nI guess this trick could add some score to the solution, but i surfed the internet searching for the similar cases and didn't find any proper information for this topic.</p>\n<p>Maybe you could give some advices on where to start and what will work?</p>",
  "messages": [
    {
      "id": "3185799",
      "postDate": "04/23/2025 21:10:42",
      "content": "<p>In continuation of the topic about voice detecting in train samples (<a href=\"https://www.kaggle.com/competitions/birdclef-2025/discussion/568886)\" target=\"_blank\">https://www.kaggle.com/competitions/birdclef-2025/discussion/568886)</a>.</p>\n<p>I would like to know what useful methods and DL models do you use to remove human voice or make it more noteless?<br>\nI guess this trick could add some score to the solution, but i surfed the internet searching for the similar cases and didn't find any proper information for this topic.</p>\n<p>Maybe you could give some advices on where to start and what will work?</p>",
      "rawMarkdown": "In continuation of the topic about voice detecting in train samples (https://www.kaggle.com/competitions/birdclef-2025/discussion/568886).\n\nI would like to know what useful methods and DL models do you use to remove human voice or make it more noteless?\nI guess this trick could add some score to the solution, but i surfed the internet searching for the similar cases and didn't find any proper information for this topic.\n\nMaybe you could give some advices on where to start and what will work?",
      "votes": null
    },
    {
      "id": "3185890",
      "postDate": "04/24/2025 01:32:24",
      "content": "<p>I filter 5% top loss to make it more robust</p>",
      "rawMarkdown": "I filter 5% top loss to make it more robust",
      "votes": null
    },
    {
      "id": "3185917",
      "postDate": "04/24/2025 02:36:06",
      "content": "<p>I would go ahead and skip timestamps containing vocals when sampling, and experimented with this a number of times, and it didn't end up transforming much. (I didn't do any ablation experiments because the results were so erratic)</p>",
      "rawMarkdown": "I would go ahead and skip timestamps containing vocals when sampling, and experimented with this a number of times, and it didn't end up transforming much. (I didn't do any ablation experiments because the results were so erratic)",
      "votes": null
    },
    {
      "id": "3186507",
      "postDate": "04/24/2025 19:21:56",
      "content": "<p>This notebook <a href=\"https://www.kaggle.com/code/kdmitrie/bc25-separation-voice-from-data/notebook\" target=\"_blank\">https://www.kaggle.com/code/kdmitrie/bc25-separation-voice-from-data/notebook</a> has a pickle table of voice. LB score went from .808 to .812 after using it on<br>\ntorch notebook.</p>",
      "rawMarkdown": "This notebook https://www.kaggle.com/code/kdmitrie/bc25-separation-voice-from-data/notebook has a pickle table of voice. LB score went from .808 to .812 after using it on\ntorch notebook.",
      "votes": null
    },
    {
      "id": "3188078",
      "postDate": "04/27/2025 02:39:51",
      "content": "<p>If I understand you are saying your current solution uses the raw data then? No changes or filters of voice? </p>",
      "rawMarkdown": "If I understand you are saying your current solution uses the raw data then? No changes or filters of voice?",
      "votes": null
    },
    {
      "id": "3188163",
      "postDate": "04/27/2025 06:30:13",
      "content": "<p>May I ask that you drop the hum voice segments or just fill them as zeros?  </p>",
      "rawMarkdown": "May I ask that you drop the hum voice segments or just fill them as zeros?",
      "votes": null
    },
    {
      "id": "3188177",
      "postDate": "04/27/2025 07:09:34",
      "content": "<p>Some of my experiments were filtered for vocals and some used direct random sampling. When I train with multicard, the results also differ when I make sure that all the parameters are the same and the random seeds are the same, so I did not perform the ablation experiments on my machine, but fine tuned the parameters to integrate them. Instead, the ablation experiments were performed in the gpu provided by kaggle, and 30 hours was simply not enough.</p>",
      "rawMarkdown": "Some of my experiments were filtered for vocals and some used direct random sampling. When I train with multicard, the results also differ when I make sure that all the parameters are the same and the random seeds are the same, so I did not perform the ablation experiments on my machine, but fine tuned the parameters to integrate them. Instead, the ablation experiments were performed in the gpu provided by kaggle, and 30 hours was simply not enough.",
      "votes": null
    },
    {
      "id": "3188629",
      "postDate": "04/28/2025 01:04:41",
      "content": "<p>I just deleted from the table in the above notebook. I do not know if it include hum voice.</p>",
      "rawMarkdown": "I just deleted from the table in the above notebook. I do not know if it include hum voice.",
      "votes": null
    },
    {
      "id": "3188666",
      "postDate": "04/28/2025 02:42:46",
      "content": "<p>What does that mean?</p>",
      "rawMarkdown": "What does that mean?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3185890,
      "author_name": "xukongji",
      "author_url": "",
      "post_date": "04/24/2025 01:32:24",
      "content": "<p>I filter 5% top loss to make it more robust</p>",
      "votes": null,
      "replies": [
        {
          "id": 3188666,
          "author_name": "tim6502",
          "author_url": "",
          "post_date": "04/28/2025 02:42:46",
          "content": "<p>What does that mean?</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3185917,
      "author_name": "agcsdedf",
      "author_url": "",
      "post_date": "04/24/2025 02:36:06",
      "content": "<p>I would go ahead and skip timestamps containing vocals when sampling, and experimented with this a number of times, and it didn't end up transforming much. (I didn't do any ablation experiments because the results were so erratic)</p>",
      "votes": null,
      "replies": [
        {
          "id": 3188078,
          "author_name": "firlas47",
          "author_url": "",
          "post_date": "04/27/2025 02:39:51",
          "content": "<p>If I understand you are saying your current solution uses the raw data then? No changes or filters of voice? </p>",
          "votes": null,
          "replies": [
            {
              "id": 3188177,
              "author_name": "agcsdedf",
              "author_url": "",
              "post_date": "04/27/2025 07:09:34",
              "content": "<p>Some of my experiments were filtered for vocals and some used direct random sampling. When I train with multicard, the results also differ when I make sure that all the parameters are the same and the random seeds are the same, so I did not perform the ablation experiments on my machine, but fine tuned the parameters to integrate them. Instead, the ablation experiments were performed in the gpu provided by kaggle, and 30 hours was simply not enough.</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3186507,
      "author_name": "tomkkk",
      "author_url": "",
      "post_date": "04/24/2025 19:21:56",
      "content": "<p>This notebook <a href=\"https://www.kaggle.com/code/kdmitrie/bc25-separation-voice-from-data/notebook\" target=\"_blank\">https://www.kaggle.com/code/kdmitrie/bc25-separation-voice-from-data/notebook</a> has a pickle table of voice. LB score went from .808 to .812 after using it on<br>\ntorch notebook.</p>",
      "votes": null,
      "replies": [
        {
          "id": 3188163,
          "author_name": "xukongji",
          "author_url": "",
          "post_date": "04/27/2025 06:30:13",
          "content": "<p>May I ask that you drop the hum voice segments or just fill them as zeros?  </p>",
          "votes": null,
          "replies": [
            {
              "id": 3188629,
              "author_name": "tomkkk",
              "author_url": "",
              "post_date": "04/28/2025 01:04:41",
              "content": "<p>I just deleted from the table in the above notebook. I do not know if it include hum voice.</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3185799": "In continuation of the topic about voice detecting in train samples (https://www.kaggle.com/competitions/birdclef-2025/discussion/568886).\n\nI would like to know what useful methods and DL models do you use to remove human voice or make it more noteless?\nI guess this trick could add some score to the solution, but i surfed the internet searching for the similar cases and didn't find any proper information for this topic.\n\nMaybe you could give some advices on where to start and what will work?",
    "3185890": "I filter 5% top loss to make it more robust",
    "3185917": "I would go ahead and skip timestamps containing vocals when sampling, and experimented with this a number of times, and it didn't end up transforming much. (I didn't do any ablation experiments because the results were so erratic)",
    "3186507": "This notebook https://www.kaggle.com/code/kdmitrie/bc25-separation-voice-from-data/notebook has a pickle table of voice. LB score went from .808 to .812 after using it on\ntorch notebook.",
    "3188078": "If I understand you are saying your current solution uses the raw data then? No changes or filters of voice?",
    "3188163": "May I ask that you drop the hum voice segments or just fill them as zeros?",
    "3188177": "Some of my experiments were filtered for vocals and some used direct random sampling. When I train with multicard, the results also differ when I make sure that all the parameters are the same and the random seeds are the same, so I did not perform the ablation experiments on my machine, but fine tuned the parameters to integrate them. Instead, the ablation experiments were performed in the gpu provided by kaggle, and 30 hours was simply not enough.",
    "3188629": "I just deleted from the table in the above notebook. I do not know if it include hum voice.",
    "3188666": "What does that mean?"
  },
  "source": "meta"
}