{
  "id": 571706,
  "title": "Shall i clean data with \"google bird vocalization classfier\" ?",
  "url": "/competitions/birdclef-2025/discussion/571706",
  "author_name": "",
  "post_date": "2025-04-05T02:32:06.301755600Z",
  "votes": 2,
  "comment_count": 2,
  "views": 0,
  "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F24077773%2Fd17736d6729b33b0e770336623bcf936%2F2025-04-05%20101852.png?generation=1743819557761498&amp;alt=media\" alt=\"\"></p>\n<h1>In thw first place's notebook of 2024's,they replace the primary lable by the secondary lable if the prediction of google's model match the secondary one,and they drop the chunk if the prediction of google's doesnt match any of the lables.</h1>\n<p>We can deduce from this that,they trust the google's model than the train data,but why would they still use the training data to train theire model?<br>\nAnd,since we probably can't see the lables's of goole's model's dataset,we don't know wheathethe google's model's dataset contain all the species of this compitition,so that the google's model can;t recognize these species,what if they drop some valuebale data by doing this?</p>",
  "messages": [
    {
      "id": "3170789",
      "postDate": "04/05/2025 02:32:06",
      "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F24077773%2Fd17736d6729b33b0e770336623bcf936%2F2025-04-05%20101852.png?generation=1743819557761498&amp;alt=media\" alt=\"\"></p>\n<h1>In thw first place's notebook of 2024's,they replace the primary lable by the secondary lable if the prediction of google's model match the secondary one,and they drop the chunk if the prediction of google's doesnt match any of the lables.</h1>\n<p>We can deduce from this that,they trust the google's model than the train data,but why would they still use the training data to train theire model?<br>\nAnd,since we probably can't see the lables's of goole's model's dataset,we don't know wheathethe google's model's dataset contain all the species of this compitition,so that the google's model can;t recognize these species,what if they drop some valuebale data by doing this?</p>",
      "rawMarkdown": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F24077773%2Fd17736d6729b33b0e770336623bcf936%2F2025-04-05%20101852.png?generation=1743819557761498&alt=media)\n\n# In thw first place's notebook of 2024's,they replace the primary lable by the secondary lable if the prediction of google's model match the secondary one,and they drop the chunk if the prediction of google's doesnt match any of the lables.\nWe can deduce from this that,they trust the google's model than the train data,but why would they still use the training data to train theire model?\nAnd,since we probably can't see the lables's of goole's model's dataset,we don't know wheathethe google's model's dataset contain all the species of this compitition,so that the google's model can;t recognize these species,what if they drop some valuebale data by doing this?",
      "votes": null
    },
    {
      "id": "3172628",
      "postDate": "04/07/2025 03:21:16",
      "content": "<p>Don’t forget.  Birds are only 25% of the creatures we need to identify.    That was not the case last year.  </p>",
      "rawMarkdown": "Don’t forget.  Birds are only 25% of the creatures we need to identify.    That was not the case last year.",
      "votes": null
    },
    {
      "id": "3172634",
      "postDate": "04/07/2025 03:37:52",
      "content": "<p>ohoh,i forgot,thanks!</p>",
      "rawMarkdown": "ohoh,i forgot,thanks!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3172628,
      "author_name": "pcjimmmy",
      "author_url": "",
      "post_date": "04/07/2025 03:21:16",
      "content": "<p>Don’t forget.  Birds are only 25% of the creatures we need to identify.    That was not the case last year.  </p>",
      "votes": null,
      "replies": [
        {
          "id": 3172634,
          "author_name": "zhangyuehan151",
          "author_url": "",
          "post_date": "04/07/2025 03:37:52",
          "content": "<p>ohoh,i forgot,thanks!</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3170789": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F24077773%2Fd17736d6729b33b0e770336623bcf936%2F2025-04-05%20101852.png?generation=1743819557761498&alt=media)\n\n# In thw first place's notebook of 2024's,they replace the primary lable by the secondary lable if the prediction of google's model match the secondary one,and they drop the chunk if the prediction of google's doesnt match any of the lables.\nWe can deduce from this that,they trust the google's model than the train data,but why would they still use the training data to train theire model?\nAnd,since we probably can't see the lables's of goole's model's dataset,we don't know wheathethe google's model's dataset contain all the species of this compitition,so that the google's model can;t recognize these species,what if they drop some valuebale data by doing this?",
    "3172628": "Don’t forget.  Birds are only 25% of the creatures we need to identify.    That was not the case last year.",
    "3172634": "ohoh,i forgot,thanks!"
  },
  "source": "meta"
}