{
  "id": 503042,
  "title": "Train and test data are obviously very different but in what way?",
  "url": "/competitions/birdclef-2024/discussion/503042",
  "author_name": "",
  "post_date": "2024-05-15T21:17:20.326547400Z",
  "votes": 2,
  "comment_count": 11,
  "views": 0,
  "content": "<p>In almost every case I saw the CV and LB of the model/notebook are very different. This implies that the train and test data are very different but in which way? My thought is that test data is much noisier because adding more data (using all 5-second chunks instead of only the first chunk or adding more xeno canto data) makes the model perform worse. This would be the case if the model only trained on non-noisy data and when adding more data it is almost \"overfitting\" on the less noisy data it sees in the train data. Is this theory correct? If it is then if overfitting happens on non-noisy data adding Gaussian noise augs could make it so that extra data could be used without any issues. Please let me know if I am mistaken, thanks!</p>",
  "messages": [
    {
      "id": "2815412",
      "postDate": "05/15/2024 21:17:20",
      "content": "<p>In almost every case I saw the CV and LB of the model/notebook are very different. This implies that the train and test data are very different but in which way? My thought is that test data is much noisier because adding more data (using all 5-second chunks instead of only the first chunk or adding more xeno canto data) makes the model perform worse. This would be the case if the model only trained on non-noisy data and when adding more data it is almost \"overfitting\" on the less noisy data it sees in the train data. Is this theory correct? If it is then if overfitting happens on non-noisy data adding Gaussian noise augs could make it so that extra data could be used without any issues. Please let me know if I am mistaken, thanks!</p>",
      "rawMarkdown": "In almost every case I saw the CV and LB of the model/notebook are very different. This implies that the train and test data are very different but in which way? My thought is that test data is much noisier because adding more data (using all 5-second chunks instead of only the first chunk or adding more xeno canto data) makes the model perform worse. This would be the case if the model only trained on non-noisy data and when adding more data it is almost \"overfitting\" on the less noisy data it sees in the train data. Is this theory correct? If it is then if overfitting happens on non-noisy data adding Gaussian noise augs could make it so that extra data could be used without any issues. Please let me know if I am mistaken, thanks!",
      "votes": null
    },
    {
      "id": "2816541",
      "postDate": "05/16/2024 11:25:27",
      "content": "<p>See <a href=\"https://www.kaggle.com/competitions/birdclef-2024/discussion/498404\" target=\"_blank\">https://www.kaggle.com/competitions/birdclef-2024/discussion/498404</a></p>",
      "rawMarkdown": "See https://www.kaggle.com/competitions/birdclef-2024/discussion/498404",
      "votes": null
    },
    {
      "id": "2816780",
      "postDate": "05/16/2024 14:30:42",
      "content": "<p>Yes, I have already seen this but it doesn't explain what the difference between test and train is.</p>",
      "rawMarkdown": "Yes, I have already seen this but it doesn't explain what the difference between test and train is.",
      "votes": null
    },
    {
      "id": "2816801",
      "postDate": "05/16/2024 14:39:51",
      "content": "<p>OK, I thought it id. Sorry.</p>",
      "rawMarkdown": "OK, I thought it id. Sorry.",
      "votes": null
    },
    {
      "id": "2817060",
      "postDate": "05/16/2024 17:07:37",
      "content": "<p>I think you're correct in general. Test soundscapes are definitely noisier, and that's why we have such large gap between CV and LB. The key to the good model here is a correct set of augmentations, which will convert \"better quality\" train audios to something closer to test soundscapes. I recommend to check previous BirdCLEF competitions - top solutions there had featured many types of such augmentations.</p>",
      "rawMarkdown": "I think you're correct in general. Test soundscapes are definitely noisier, and that's why we have such large gap between CV and LB. The key to the good model here is a correct set of augmentations, which will convert \"better quality\" train audios to something closer to test soundscapes. I recommend to check previous BirdCLEF competitions - top solutions there had featured many types of such augmentations.",
      "votes": null
    },
    {
      "id": "2817071",
      "postDate": "05/16/2024 17:15:47",
      "content": "<p>Its fine do you have any idea though? Again congrats on the #1 so far 😁</p>",
      "rawMarkdown": "Its fine do you have any idea though? Again congrats on the #1 so far 😁",
      "votes": null
    },
    {
      "id": "2817074",
      "postDate": "05/16/2024 17:20:10",
      "content": "<p>Thanks for the insights!</p>",
      "rawMarkdown": "Thanks for the insights!",
      "votes": null
    },
    {
      "id": "2817434",
      "postDate": "05/16/2024 22:35:20",
      "content": "<p>See <a href=\"https://www.kaggle.com/competitions/birdclef-2024/discussion/498404#2817205\" target=\"_blank\">https://www.kaggle.com/competitions/birdclef-2024/discussion/498404#2817205</a></p>",
      "rawMarkdown": "See https://www.kaggle.com/competitions/birdclef-2024/discussion/498404#2817205",
      "votes": null
    },
    {
      "id": "2819654",
      "postDate": "05/17/2024 16:11:14",
      "content": "<p>Thank you!!!</p>",
      "rawMarkdown": "Thank you!!!",
      "votes": null
    },
    {
      "id": "2820792",
      "postDate": "05/17/2024 17:52:17",
      "content": "<p>Glad you it useful. 3 people downvoted, not sure why.</p>",
      "rawMarkdown": "Glad you it useful. 3 people downvoted, not sure why.",
      "votes": null
    },
    {
      "id": "2820830",
      "postDate": "05/17/2024 18:14:39",
      "content": "<p>I think they are salty of your first place 🤣</p>",
      "rawMarkdown": "I think they are salty of your first place 🤣",
      "votes": null
    },
    {
      "id": "2821035",
      "postDate": "05/17/2024 20:24:07",
      "content": "<p>Well, now that it is no longer the case they can relax maybe.</p>",
      "rawMarkdown": "Well, now that it is no longer the case they can relax maybe.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2816541,
      "author_name": "cpmpml",
      "author_url": "",
      "post_date": "05/16/2024 11:25:27",
      "content": "<p>See <a href=\"https://www.kaggle.com/competitions/birdclef-2024/discussion/498404\" target=\"_blank\">https://www.kaggle.com/competitions/birdclef-2024/discussion/498404</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 2816780,
          "author_name": "max1mum",
          "author_url": "",
          "post_date": "05/16/2024 14:30:42",
          "content": "<p>Yes, I have already seen this but it doesn't explain what the difference between test and train is.</p>",
          "votes": null,
          "replies": [
            {
              "id": 2816801,
              "author_name": "cpmpml",
              "author_url": "",
              "post_date": "05/16/2024 14:39:51",
              "content": "<p>OK, I thought it id. Sorry.</p>",
              "votes": null,
              "replies": [
                {
                  "id": 2817071,
                  "author_name": "max1mum",
                  "author_url": "",
                  "post_date": "05/16/2024 17:15:47",
                  "content": "<p>Its fine do you have any idea though? Again congrats on the #1 so far 😁</p>",
                  "votes": null,
                  "replies": [
                    {
                      "id": 2817434,
                      "author_name": "cpmpml",
                      "author_url": "",
                      "post_date": "05/16/2024 22:35:20",
                      "content": "<p>See <a href=\"https://www.kaggle.com/competitions/birdclef-2024/discussion/498404#2817205\" target=\"_blank\">https://www.kaggle.com/competitions/birdclef-2024/discussion/498404#2817205</a></p>",
                      "votes": null,
                      "replies": [
                        {
                          "id": 2819654,
                          "author_name": "max1mum",
                          "author_url": "",
                          "post_date": "05/17/2024 16:11:14",
                          "content": "<p>Thank you!!!</p>",
                          "votes": null,
                          "replies": [
                            {
                              "id": 2820792,
                              "author_name": "cpmpml",
                              "author_url": "",
                              "post_date": "05/17/2024 17:52:17",
                              "content": "<p>Glad you it useful. 3 people downvoted, not sure why.</p>",
                              "votes": null,
                              "replies": [
                                {
                                  "id": 2820830,
                                  "author_name": "max1mum",
                                  "author_url": "",
                                  "post_date": "05/17/2024 18:14:39",
                                  "content": "<p>I think they are salty of your first place 🤣</p>",
                                  "votes": null,
                                  "replies": [
                                    {
                                      "id": 2821035,
                                      "author_name": "cpmpml",
                                      "author_url": "",
                                      "post_date": "05/17/2024 20:24:07",
                                      "content": "<p>Well, now that it is no longer the case they can relax maybe.</p>",
                                      "votes": null,
                                      "replies": []
                                    }
                                  ]
                                }
                              ]
                            }
                          ]
                        }
                      ]
                    }
                  ]
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 2817060,
      "author_name": "atamazian",
      "author_url": "",
      "post_date": "05/16/2024 17:07:37",
      "content": "<p>I think you're correct in general. Test soundscapes are definitely noisier, and that's why we have such large gap between CV and LB. The key to the good model here is a correct set of augmentations, which will convert \"better quality\" train audios to something closer to test soundscapes. I recommend to check previous BirdCLEF competitions - top solutions there had featured many types of such augmentations.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2817074,
          "author_name": "max1mum",
          "author_url": "",
          "post_date": "05/16/2024 17:20:10",
          "content": "<p>Thanks for the insights!</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2815412": "In almost every case I saw the CV and LB of the model/notebook are very different. This implies that the train and test data are very different but in which way? My thought is that test data is much noisier because adding more data (using all 5-second chunks instead of only the first chunk or adding more xeno canto data) makes the model perform worse. This would be the case if the model only trained on non-noisy data and when adding more data it is almost \"overfitting\" on the less noisy data it sees in the train data. Is this theory correct? If it is then if overfitting happens on non-noisy data adding Gaussian noise augs could make it so that extra data could be used without any issues. Please let me know if I am mistaken, thanks!",
    "2816541": "See https://www.kaggle.com/competitions/birdclef-2024/discussion/498404",
    "2816780": "Yes, I have already seen this but it doesn't explain what the difference between test and train is.",
    "2816801": "OK, I thought it id. Sorry.",
    "2817060": "I think you're correct in general. Test soundscapes are definitely noisier, and that's why we have such large gap between CV and LB. The key to the good model here is a correct set of augmentations, which will convert \"better quality\" train audios to something closer to test soundscapes. I recommend to check previous BirdCLEF competitions - top solutions there had featured many types of such augmentations.",
    "2817071": "Its fine do you have any idea though? Again congrats on the #1 so far 😁",
    "2817074": "Thanks for the insights!",
    "2817434": "See https://www.kaggle.com/competitions/birdclef-2024/discussion/498404#2817205",
    "2819654": "Thank you!!!",
    "2820792": "Glad you it useful. 3 people downvoted, not sure why.",
    "2820830": "I think they are salty of your first place 🤣",
    "2821035": "Well, now that it is no longer the case they can relax maybe."
  },
  "source": "meta"
}