{
  "id": 492300,
  "title": "Using additional data significantly improves CV, but has almost no effect on LB",
  "url": "/competitions/birdclef-2024/discussion/492300",
  "author_name": "",
  "post_date": "2024-04-09T07:11:57.160029200Z",
  "votes": 8,
  "comment_count": 10,
  "views": 0,
  "content": "<p>Hello everybody!<br>\nI trained the same model twice with two stages (just ran the training twice):</p>\n<ol>\n<li>Preliminary training - data from previous competitions and XENO were used to train the model.</li>\n<li>After that, training was conducted on the data of the current competition.</li>\n</ol>\n<p>The results after the second stage were about the same val_auc - 0.82, but lb one time was 0.63 the other 0.6.<br>\nI tried to experiment and train the model only at the second stage and got the following results val_auc - 0.72, lb 0.59.<br>\nTo sum up, the spread between two identical models is still too large, it needs to be understood and taken into account.<br>\nSecond, the use of additional data is not so effective, in the initial stages you can experiment only with the original dataset, and only then add additional data for a relatively small improvement.</p>\n<p>P.S. I posted the code I used a few days ago and separate inference (with the best lb so far), but now you can find it at the end of the list of notebooks because kaggle thinks it's not relevant or something, I do not know.</p>",
  "messages": [
    {
      "id": "2743016",
      "postDate": "04/09/2024 07:11:57",
      "content": "<p>Hello everybody!<br>\nI trained the same model twice with two stages (just ran the training twice):</p>\n<ol>\n<li>Preliminary training - data from previous competitions and XENO were used to train the model.</li>\n<li>After that, training was conducted on the data of the current competition.</li>\n</ol>\n<p>The results after the second stage were about the same val_auc - 0.82, but lb one time was 0.63 the other 0.6.<br>\nI tried to experiment and train the model only at the second stage and got the following results val_auc - 0.72, lb 0.59.<br>\nTo sum up, the spread between two identical models is still too large, it needs to be understood and taken into account.<br>\nSecond, the use of additional data is not so effective, in the initial stages you can experiment only with the original dataset, and only then add additional data for a relatively small improvement.</p>\n<p>P.S. I posted the code I used a few days ago and separate inference (with the best lb so far), but now you can find it at the end of the list of notebooks because kaggle thinks it's not relevant or something, I do not know.</p>",
      "rawMarkdown": "Hello everybody!\nI trained the same model twice with two stages (just ran the training twice):\n1. Preliminary training - data from previous competitions and XENO were used to train the model.\n2. After that, training was conducted on the data of the current competition.\n\nThe results after the second stage were about the same val_auc - 0.82, but lb one time was 0.63 the other 0.6.\nI tried to experiment and train the model only at the second stage and got the following results val_auc - 0.72, lb 0.59.\nTo sum up, the spread between two identical models is still too large, it needs to be understood and taken into account.\nSecond, the use of additional data is not so effective, in the initial stages you can experiment only with the original dataset, and only then add additional data for a relatively small improvement.\n\nP.S. I posted the code I used a few days ago and separate inference (with the best lb so far), but now you can find it at the end of the list of notebooks because kaggle thinks it's not relevant or something, I do not know.",
      "votes": null
    },
    {
      "id": "2743136",
      "postDate": "04/09/2024 08:47:21",
      "content": "<p>When you are comparing the val_auc, are you keeping the val sets identical ? Or are you making a new split with this data and comparing the two vals ? Interesting results nevertheless, could it be duplicates ?</p>",
      "rawMarkdown": "When you are comparing the val_auc, are you keeping the val sets identical ? Or are you making a new split with this data and comparing the two vals ? Interesting results nevertheless, could it be duplicates ?",
      "votes": null
    },
    {
      "id": "2743139",
      "postDate": "04/09/2024 08:48:23",
      "content": "<p>What format of the extra data are you using ? mp3 or waw or else ? maybe some compression could have shifted the training distribution</p>",
      "rawMarkdown": "What format of the extra data are you using ? mp3 or waw or else ? maybe some compression could have shifted the training distribution",
      "votes": null
    },
    {
      "id": "2743182",
      "postDate": "04/09/2024 09:09:56",
      "content": "<p>Hello! Here is the code <a href=\"https://www.kaggle.com/code/aikhmelnytskyy/birdclef24-pretraining-is-all-you-need\" target=\"_blank\">https://www.kaggle.com/code/aikhmelnytskyy/birdclef24-pretraining-is-all-you-need</a>. Version 2 features a two-step workout, but if you want to test, copy the workbook from version 3. The difference between the versions is the environment. The latest environments offered by kaggle break tensorflow libraries. I can't do anything about it, since this is a training on TPU, to fix the error, you need to install all the libraries from scratch, which is better not to do in this case</p>",
      "rawMarkdown": "Hello! Here is the code https://www.kaggle.com/code/aikhmelnytskyy/birdclef24-pretraining-is-all-you-need. Version 2 features a two-step workout, but if you want to test, copy the workbook from version 3. The difference between the versions is the environment. The latest environments offered by kaggle break tensorflow libraries. I can't do anything about it, since this is a training on TPU, to fix the error, you need to install all the libraries from scratch, which is better not to do in this case",
      "votes": null
    },
    {
      "id": "2743189",
      "postDate": "04/09/2024 09:13:58",
      "content": "<p>As I wrote in the notebook, this approach worked last year, I think it's something else.</p>",
      "rawMarkdown": "As I wrote in the notebook, this approach worked last year, I think it's something else.",
      "votes": null
    },
    {
      "id": "2743314",
      "postDate": "04/09/2024 11:24:16",
      "content": "<p>interesting. I was planning same approach but using pytorch. If it works i will share the insight but i would do that during the week end. Not much time during weekday</p>",
      "rawMarkdown": "interesting. I was planning same approach but using pytorch. If it works i will share the insight but i would do that during the week end. Not much time during weekday",
      "votes": null
    },
    {
      "id": "2743336",
      "postDate": "04/09/2024 11:42:53",
      "content": "<p>It's actually a great idea. I'm using tensorflow because I don't have a powerful enough PC to compete with the rest of the participants. In this sense, TPU expands my possibilities, but with each competition I realize that pytorch is better.</p>",
      "rawMarkdown": "It's actually a great idea. I'm using tensorflow because I don't have a powerful enough PC to compete with the rest of the participants. In this sense, TPU expands my possibilities, but with each competition I realize that pytorch is better.",
      "votes": null
    },
    {
      "id": "2743613",
      "postDate": "04/09/2024 14:48:59",
      "content": "<p>And there's Jax! :)</p>",
      "rawMarkdown": "And there's Jax! :)",
      "votes": null
    },
    {
      "id": "2744451",
      "postDate": "04/09/2024 23:25:06",
      "content": "<p>I don't have solutions - but if it makes you feel better - I just had the exact same experience.</p>\n<p>Trained on 10x the data. Validation went way up.  LB went down .01.</p>",
      "rawMarkdown": "I don't have solutions - but if it makes you feel better - I just had the exact same experience.\n\nTrained on 10x the data. Validation went way up.  LB went down .01.",
      "votes": null
    },
    {
      "id": "2744953",
      "postDate": "04/10/2024 08:29:00",
      "content": "<p>In my experience, the additional data did improve the LB score, but the improvement was minimal. Possibly because most of the additional data is from unscored species?</p>",
      "rawMarkdown": "In my experience, the additional data did improve the LB score, but the improvement was minimal. Possibly because most of the additional data is from unscored species?",
      "votes": null
    },
    {
      "id": "2744961",
      "postDate": "04/10/2024 08:35:01",
      "content": "<p>Yes, I think that's the point. Adding species that are not evaluated does not help, in addition, they may sound too good, or the recordings have other noises that help to \"classify\", thus artificially raising the quality of the models due to known different samples</p>",
      "rawMarkdown": "Yes, I think that's the point. Adding species that are not evaluated does not help, in addition, they may sound too good, or the recordings have other noises that help to \"classify\", thus artificially raising the quality of the models due to known different samples",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2743136,
      "author_name": "janmpia",
      "author_url": "",
      "post_date": "04/09/2024 08:47:21",
      "content": "<p>When you are comparing the val_auc, are you keeping the val sets identical ? Or are you making a new split with this data and comparing the two vals ? Interesting results nevertheless, could it be duplicates ?</p>",
      "votes": null,
      "replies": [
        {
          "id": 2743139,
          "author_name": "janmpia",
          "author_url": "",
          "post_date": "04/09/2024 08:48:23",
          "content": "<p>What format of the extra data are you using ? mp3 or waw or else ? maybe some compression could have shifted the training distribution</p>",
          "votes": null,
          "replies": [
            {
              "id": 2743182,
              "author_name": "aikhmelnytskyy",
              "author_url": "",
              "post_date": "04/09/2024 09:09:56",
              "content": "<p>Hello! Here is the code <a href=\"https://www.kaggle.com/code/aikhmelnytskyy/birdclef24-pretraining-is-all-you-need\" target=\"_blank\">https://www.kaggle.com/code/aikhmelnytskyy/birdclef24-pretraining-is-all-you-need</a>. Version 2 features a two-step workout, but if you want to test, copy the workbook from version 3. The difference between the versions is the environment. The latest environments offered by kaggle break tensorflow libraries. I can't do anything about it, since this is a training on TPU, to fix the error, you need to install all the libraries from scratch, which is better not to do in this case</p>",
              "votes": null,
              "replies": [
                {
                  "id": 2743189,
                  "author_name": "aikhmelnytskyy",
                  "author_url": "",
                  "post_date": "04/09/2024 09:13:58",
                  "content": "<p>As I wrote in the notebook, this approach worked last year, I think it's something else.</p>",
                  "votes": null,
                  "replies": [
                    {
                      "id": 2743314,
                      "author_name": "ludovick",
                      "author_url": "",
                      "post_date": "04/09/2024 11:24:16",
                      "content": "<p>interesting. I was planning same approach but using pytorch. If it works i will share the insight but i would do that during the week end. Not much time during weekday</p>",
                      "votes": null,
                      "replies": [
                        {
                          "id": 2743336,
                          "author_name": "aikhmelnytskyy",
                          "author_url": "",
                          "post_date": "04/09/2024 11:42:53",
                          "content": "<p>It's actually a great idea. I'm using tensorflow because I don't have a powerful enough PC to compete with the rest of the participants. In this sense, TPU expands my possibilities, but with each competition I realize that pytorch is better.</p>",
                          "votes": null,
                          "replies": [
                            {
                              "id": 2743613,
                              "author_name": "tomdenton",
                              "author_url": "",
                              "post_date": "04/09/2024 14:48:59",
                              "content": "<p>And there's Jax! :)</p>",
                              "votes": null,
                              "replies": []
                            }
                          ]
                        }
                      ]
                    }
                  ]
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 2744451,
      "author_name": "richolson",
      "author_url": "",
      "post_date": "04/09/2024 23:25:06",
      "content": "<p>I don't have solutions - but if it makes you feel better - I just had the exact same experience.</p>\n<p>Trained on 10x the data. Validation went way up.  LB went down .01.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2744953,
      "author_name": "hichambellafkir",
      "author_url": "",
      "post_date": "04/10/2024 08:29:00",
      "content": "<p>In my experience, the additional data did improve the LB score, but the improvement was minimal. Possibly because most of the additional data is from unscored species?</p>",
      "votes": null,
      "replies": [
        {
          "id": 2744961,
          "author_name": "aikhmelnytskyy",
          "author_url": "",
          "post_date": "04/10/2024 08:35:01",
          "content": "<p>Yes, I think that's the point. Adding species that are not evaluated does not help, in addition, they may sound too good, or the recordings have other noises that help to \"classify\", thus artificially raising the quality of the models due to known different samples</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2743016": "Hello everybody!\nI trained the same model twice with two stages (just ran the training twice):\n1. Preliminary training - data from previous competitions and XENO were used to train the model.\n2. After that, training was conducted on the data of the current competition.\n\nThe results after the second stage were about the same val_auc - 0.82, but lb one time was 0.63 the other 0.6.\nI tried to experiment and train the model only at the second stage and got the following results val_auc - 0.72, lb 0.59.\nTo sum up, the spread between two identical models is still too large, it needs to be understood and taken into account.\nSecond, the use of additional data is not so effective, in the initial stages you can experiment only with the original dataset, and only then add additional data for a relatively small improvement.\n\nP.S. I posted the code I used a few days ago and separate inference (with the best lb so far), but now you can find it at the end of the list of notebooks because kaggle thinks it's not relevant or something, I do not know.",
    "2743136": "When you are comparing the val_auc, are you keeping the val sets identical ? Or are you making a new split with this data and comparing the two vals ? Interesting results nevertheless, could it be duplicates ?",
    "2743139": "What format of the extra data are you using ? mp3 or waw or else ? maybe some compression could have shifted the training distribution",
    "2743182": "Hello! Here is the code https://www.kaggle.com/code/aikhmelnytskyy/birdclef24-pretraining-is-all-you-need. Version 2 features a two-step workout, but if you want to test, copy the workbook from version 3. The difference between the versions is the environment. The latest environments offered by kaggle break tensorflow libraries. I can't do anything about it, since this is a training on TPU, to fix the error, you need to install all the libraries from scratch, which is better not to do in this case",
    "2743189": "As I wrote in the notebook, this approach worked last year, I think it's something else.",
    "2743314": "interesting. I was planning same approach but using pytorch. If it works i will share the insight but i would do that during the week end. Not much time during weekday",
    "2743336": "It's actually a great idea. I'm using tensorflow because I don't have a powerful enough PC to compete with the rest of the participants. In this sense, TPU expands my possibilities, but with each competition I realize that pytorch is better.",
    "2743613": "And there's Jax! :)",
    "2744451": "I don't have solutions - but if it makes you feel better - I just had the exact same experience.\n\nTrained on 10x the data. Validation went way up.  LB went down .01.",
    "2744953": "In my experience, the additional data did improve the LB score, but the improvement was minimal. Possibly because most of the additional data is from unscored species?",
    "2744961": "Yes, I think that's the point. Adding species that are not evaluated does not help, in addition, they may sound too good, or the recordings have other noises that help to \"classify\", thus artificially raising the quality of the models due to known different samples"
  },
  "source": "meta"
}