{
  "id": 511569,
  "title": "Time to learn",
  "url": "/competitions/birdclef-2024/discussion/511569",
  "author_name": "",
  "post_date": "2024-06-11T08:34:59.319472500Z",
  "votes": 47,
  "comment_count": 9,
  "views": 0,
  "content": "<p>Past the shock for those who dropped, or the disappointment to see some unselected submissions score way higher, now is the best time to learn from what people share about their solutions. Especially if this was your first Kaggle competition.</p>\n<p>Read the writeups people will share (we'll share ours as well ASAP).</p>\n<p>Look for how people tried to cope with the train/test difference. How they prepared data. Which models they used. How they used unlabelled soundscape, if they did. Etc.</p>\n<p>After that, check what you could have done differently.</p>\n<p>You can also submit new models (not sure if late submissions are already opened, but usually they are available some time after competition ends). However, beware that now you see the private LB, hence you are no longer in the competition conditions. It is now much easier to improve the private LB score.</p>\n<p>Doing the above consistently is how people improve on Kaggle.</p>\n<p>To conclude, here are two thing that many experienced as well, no need to report heavily on these IMHO:</p>\n<ol>\n<li><p>You have a much better sub that you did not select. This happens a lot, especially when there is no reliable local validation. In our team we have many unselected subs that score over 0.70 on private. Yet we had no way to know these would be great, hence no real regret here.</p></li>\n<li><p>Using late subs you can reach extraordinary score. Sure. But as said above, you now have more information than during the competition. Any score obtained after the deadline can't be compared fairly with competition scores. It is not an apple to apple comparison.</p></li>\n</ol>\n<p>With that, I wish all of us to learn from this competition. It applies to me as well of course.</p>",
  "messages": [
    {
      "id": "2866295",
      "postDate": "06/11/2024 08:34:59",
      "content": "<p>Past the shock for those who dropped, or the disappointment to see some unselected submissions score way higher, now is the best time to learn from what people share about their solutions. Especially if this was your first Kaggle competition.</p>\n<p>Read the writeups people will share (we'll share ours as well ASAP).</p>\n<p>Look for how people tried to cope with the train/test difference. How they prepared data. Which models they used. How they used unlabelled soundscape, if they did. Etc.</p>\n<p>After that, check what you could have done differently.</p>\n<p>You can also submit new models (not sure if late submissions are already opened, but usually they are available some time after competition ends). However, beware that now you see the private LB, hence you are no longer in the competition conditions. It is now much easier to improve the private LB score.</p>\n<p>Doing the above consistently is how people improve on Kaggle.</p>\n<p>To conclude, here are two thing that many experienced as well, no need to report heavily on these IMHO:</p>\n<ol>\n<li><p>You have a much better sub that you did not select. This happens a lot, especially when there is no reliable local validation. In our team we have many unselected subs that score over 0.70 on private. Yet we had no way to know these would be great, hence no real regret here.</p></li>\n<li><p>Using late subs you can reach extraordinary score. Sure. But as said above, you now have more information than during the competition. Any score obtained after the deadline can't be compared fairly with competition scores. It is not an apple to apple comparison.</p></li>\n</ol>\n<p>With that, I wish all of us to learn from this competition. It applies to me as well of course.</p>",
      "rawMarkdown": "Past the shock for those who dropped, or the disappointment to see some unselected submissions score way higher, now is the best time to learn from what people share about their solutions. Especially if this was your first Kaggle competition.\n\nRead the writeups people will share (we'll share ours as well ASAP).\n\nLook for how people tried to cope with the train/test difference. How they prepared data. Which models they used. How they used unlabelled soundscape, if they did. Etc.\n\nAfter that, check what you could have done differently.\n\nYou can also submit new models (not sure if late submissions are already opened, but usually they are available some time after competition ends). However, beware that now you see the private LB, hence you are no longer in the competition conditions. It is now much easier to improve the private LB score.\n\nDoing the above consistently is how people improve on Kaggle.\n\nTo conclude, here are two thing that many experienced as well, no need to report heavily on these IMHO:\n\n1. You have a much better sub that you did not select. This happens a lot, especially when there is no reliable local validation. In our team we have many unselected subs that score over 0.70 on private. Yet we had no way to know these would be great, hence no real regret here.\n\n2. Using late subs you can reach extraordinary score. Sure. But as said above, you now have more information than during the competition. Any score obtained after the deadline can't be compared fairly with competition scores. It is not an apple to apple comparison.\n\nWith that, I wish all of us to learn from this competition. It applies to me as well of course.",
      "votes": null
    },
    {
      "id": "2866401",
      "postDate": "06/11/2024 10:05:03",
      "content": "<p>Haha, yes, we had no way to know which submission would be great.<br>\nFor me, the only thing I think I can trust during the competition is whether the methodology is correct or not to prevent choosing a submission overfitting the LB.</p>\n<p>I have experienced an overfitting in another competition which gave me the 1st place in silver medal zone, missing the 5th place gold. So I sweared never to make the same mistake again.</p>",
      "rawMarkdown": "Haha, yes, we had no way to know which submission would be great.\nFor me, the only thing I think I can trust during the competition is whether the methodology is correct or not to prevent choosing a submission overfitting the LB.\n\nI have experienced an overfitting in another competition which gave me the 1st place in silver medal zone, missing the 5th place gold. So I sweared never to make the same mistake again.",
      "votes": null
    },
    {
      "id": "2866638",
      "postDate": "06/11/2024 12:29:16",
      "content": "<p>Same here. In my first or second kaggle comp I moved from 100ish rank public to 2000ish rank private. I got burned so badly that I fear overfititng like hell.</p>\n<p>It sill happens form time to time, but I try to be conservative.</p>",
      "rawMarkdown": "Same here. In my first or second kaggle comp I moved from 100ish rank public to 2000ish rank private. I got burned so badly that I fear overfititng like hell.\n\nIt sill happens form time to time, but I try to be conservative.",
      "votes": null
    },
    {
      "id": "2866756",
      "postDate": "06/11/2024 13:38:00",
      "content": "<p>100% can relate (though I didn’t take part in this comp).</p>\n<p>This is largely how I expanded my knowledge across the past year I have been on Kaggle. The data science field is vast and is impossible to know everything, nor try out every possible solution under the sun during the 3-month competition period…there’s likely to have some creative idea you missed out on (even if you finish in the gold zone)</p>\n<p>The learnings from a comp is more important than the absolute rank one gets. Not doing well in a comp does not stop you from learning/evaluating your mistakes and aim for better in the next one!</p>",
      "rawMarkdown": "100% can relate (though I didn’t take part in this comp).\n\nThis is largely how I expanded my knowledge across the past year I have been on Kaggle. The data science field is vast and is impossible to know everything, nor try out every possible solution under the sun during the 3-month competition period…there’s likely to have some creative idea you missed out on (even if you finish in the gold zone)\n\nThe learnings from a comp is more important than the absolute rank one gets. Not doing well in a comp does not stop you from learning/evaluating your mistakes and aim for better in the next one!",
      "votes": null
    },
    {
      "id": "2867540",
      "postDate": "06/11/2024 22:14:10",
      "content": "<p>First of all congratulations on your achievement.. wanted to ask After seeing the solutions of the winner what do you suggest is the best way to start thinking the way winners think in terms of coming up with a solution or strategy? </p>",
      "rawMarkdown": "First of all congratulations on your achievement.. wanted to ask After seeing the solutions of the winner what do you suggest is the best way to start thinking the way winners think in terms of coming up with a solution or strategy?",
      "votes": null
    },
    {
      "id": "2867787",
      "postDate": "06/12/2024 05:07:44",
      "content": "<p>100% right. Thanks <a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a>. I think your team results is the most impressive ones in this competition. You guys are the only team on the leaderboard who saved same position in private. Curious to read about your validation scheme. Bravo and congarts!</p>",
      "rawMarkdown": "100% right. Thanks @cpmpml. I think your team results is the most impressive ones in this competition. You guys are the only team on the leaderboard who saved same position in private. Curious to read about your validation scheme. Bravo and congarts!",
      "votes": null
    },
    {
      "id": "2868181",
      "postDate": "06/12/2024 08:30:43",
      "content": "<p>Congratulations! I've just started studying, but my scores haven't been as good as I expected, which has made me a bit anxious. However, I feel that I can grow quickly with the help of people like you, who are willing to discuss, share, and provide guidance. Thank you.</p>",
      "rawMarkdown": "Congratulations! I've just started studying, but my scores haven't been as good as I expected, which has made me a bit anxious. However, I feel that I can grow quickly with the help of people like you, who are willing to discuss, share, and provide guidance. Thank you.",
      "votes": null
    },
    {
      "id": "2868230",
      "postDate": "06/12/2024 08:55:24",
      "content": "<p>Thanks for these tips and waiting for your solution writeup. 👍</p>",
      "rawMarkdown": "Thanks for these tips and waiting for your solution writeup. 👍",
      "votes": null
    },
    {
      "id": "2868534",
      "postDate": "06/12/2024 13:03:58",
      "content": "<p>The one advice is to try lots of things and test them rigorously. Test one thing at a time. Here, testing means submitting and looking at LB core given we did not have a good local validation setup.</p>\n<p>Test ideas and keep those who improve the LB.</p>\n<p>This can lead to overfititng the LB however. </p>\n<p>Therefore do not stick to one run. Submit models trained with different random seeds, and consider the mean of their performance to be a more robust evaluation than single runs. You don't want to tune for specific random seeds. That's a recipe for disaster. Same applies to folds if you use kfold cv when training models. If scores across folds vary a lot, then use the mean as a robust score. Do not tune for the best fold.</p>",
      "rawMarkdown": "The one advice is to try lots of things and test them rigorously. Test one thing at a time. Here, testing means submitting and looking at LB core given we did not have a good local validation setup.\n\nTest ideas and keep those who improve the LB.\n\nThis can lead to overfititng the LB however. \n\nTherefore do not stick to one run. Submit models trained with different random seeds, and consider the mean of their performance to be a more robust evaluation than single runs. You don't want to tune for specific random seeds. That's a recipe for disaster. Same applies to folds if you use kfold cv when training models. If scores across folds vary a lot, then use the mean as a robust score. Do not tune for the best fold.",
      "votes": null
    },
    {
      "id": "2869063",
      "postDate": "06/12/2024 20:06:07",
      "content": "<p>Thank you for your advice. These insights are beneficial for me.</p>",
      "rawMarkdown": "Thank you for your advice. These insights are beneficial for me.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2866401,
      "author_name": "honglihang",
      "author_url": "",
      "post_date": "06/11/2024 10:05:03",
      "content": "<p>Haha, yes, we had no way to know which submission would be great.<br>\nFor me, the only thing I think I can trust during the competition is whether the methodology is correct or not to prevent choosing a submission overfitting the LB.</p>\n<p>I have experienced an overfitting in another competition which gave me the 1st place in silver medal zone, missing the 5th place gold. So I sweared never to make the same mistake again.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2866638,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "06/11/2024 12:29:16",
          "content": "<p>Same here. In my first or second kaggle comp I moved from 100ish rank public to 2000ish rank private. I got burned so badly that I fear overfititng like hell.</p>\n<p>It sill happens form time to time, but I try to be conservative.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2866756,
      "author_name": "yeoyunsianggeremie",
      "author_url": "",
      "post_date": "06/11/2024 13:38:00",
      "content": "<p>100% can relate (though I didn’t take part in this comp).</p>\n<p>This is largely how I expanded my knowledge across the past year I have been on Kaggle. The data science field is vast and is impossible to know everything, nor try out every possible solution under the sun during the 3-month competition period…there’s likely to have some creative idea you missed out on (even if you finish in the gold zone)</p>\n<p>The learnings from a comp is more important than the absolute rank one gets. Not doing well in a comp does not stop you from learning/evaluating your mistakes and aim for better in the next one!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2867540,
      "author_name": "arjunm97",
      "author_url": "",
      "post_date": "06/11/2024 22:14:10",
      "content": "<p>First of all congratulations on your achievement.. wanted to ask After seeing the solutions of the winner what do you suggest is the best way to start thinking the way winners think in terms of coming up with a solution or strategy? </p>",
      "votes": null,
      "replies": [
        {
          "id": 2868534,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "06/12/2024 13:03:58",
          "content": "<p>The one advice is to try lots of things and test them rigorously. Test one thing at a time. Here, testing means submitting and looking at LB core given we did not have a good local validation setup.</p>\n<p>Test ideas and keep those who improve the LB.</p>\n<p>This can lead to overfititng the LB however. </p>\n<p>Therefore do not stick to one run. Submit models trained with different random seeds, and consider the mean of their performance to be a more robust evaluation than single runs. You don't want to tune for specific random seeds. That's a recipe for disaster. Same applies to folds if you use kfold cv when training models. If scores across folds vary a lot, then use the mean as a robust score. Do not tune for the best fold.</p>",
          "votes": null,
          "replies": [
            {
              "id": 2869063,
              "author_name": "arjunm97",
              "author_url": "",
              "post_date": "06/12/2024 20:06:07",
              "content": "<p>Thank you for your advice. These insights are beneficial for me.</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2867787,
      "author_name": "samvelkoch",
      "author_url": "",
      "post_date": "06/12/2024 05:07:44",
      "content": "<p>100% right. Thanks <a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a>. I think your team results is the most impressive ones in this competition. You guys are the only team on the leaderboard who saved same position in private. Curious to read about your validation scheme. Bravo and congarts!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2868181,
      "author_name": "youngjunplayer",
      "author_url": "",
      "post_date": "06/12/2024 08:30:43",
      "content": "<p>Congratulations! I've just started studying, but my scores haven't been as good as I expected, which has made me a bit anxious. However, I feel that I can grow quickly with the help of people like you, who are willing to discuss, share, and provide guidance. Thank you.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2868230,
      "author_name": "yassinealouini",
      "author_url": "",
      "post_date": "06/12/2024 08:55:24",
      "content": "<p>Thanks for these tips and waiting for your solution writeup. 👍</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2866295": "Past the shock for those who dropped, or the disappointment to see some unselected submissions score way higher, now is the best time to learn from what people share about their solutions. Especially if this was your first Kaggle competition.\n\nRead the writeups people will share (we'll share ours as well ASAP).\n\nLook for how people tried to cope with the train/test difference. How they prepared data. Which models they used. How they used unlabelled soundscape, if they did. Etc.\n\nAfter that, check what you could have done differently.\n\nYou can also submit new models (not sure if late submissions are already opened, but usually they are available some time after competition ends). However, beware that now you see the private LB, hence you are no longer in the competition conditions. It is now much easier to improve the private LB score.\n\nDoing the above consistently is how people improve on Kaggle.\n\nTo conclude, here are two thing that many experienced as well, no need to report heavily on these IMHO:\n\n1. You have a much better sub that you did not select. This happens a lot, especially when there is no reliable local validation. In our team we have many unselected subs that score over 0.70 on private. Yet we had no way to know these would be great, hence no real regret here.\n\n2. Using late subs you can reach extraordinary score. Sure. But as said above, you now have more information than during the competition. Any score obtained after the deadline can't be compared fairly with competition scores. It is not an apple to apple comparison.\n\nWith that, I wish all of us to learn from this competition. It applies to me as well of course.",
    "2866401": "Haha, yes, we had no way to know which submission would be great.\nFor me, the only thing I think I can trust during the competition is whether the methodology is correct or not to prevent choosing a submission overfitting the LB.\n\nI have experienced an overfitting in another competition which gave me the 1st place in silver medal zone, missing the 5th place gold. So I sweared never to make the same mistake again.",
    "2866638": "Same here. In my first or second kaggle comp I moved from 100ish rank public to 2000ish rank private. I got burned so badly that I fear overfititng like hell.\n\nIt sill happens form time to time, but I try to be conservative.",
    "2866756": "100% can relate (though I didn’t take part in this comp).\n\nThis is largely how I expanded my knowledge across the past year I have been on Kaggle. The data science field is vast and is impossible to know everything, nor try out every possible solution under the sun during the 3-month competition period…there’s likely to have some creative idea you missed out on (even if you finish in the gold zone)\n\nThe learnings from a comp is more important than the absolute rank one gets. Not doing well in a comp does not stop you from learning/evaluating your mistakes and aim for better in the next one!",
    "2867540": "First of all congratulations on your achievement.. wanted to ask After seeing the solutions of the winner what do you suggest is the best way to start thinking the way winners think in terms of coming up with a solution or strategy?",
    "2867787": "100% right. Thanks @cpmpml. I think your team results is the most impressive ones in this competition. You guys are the only team on the leaderboard who saved same position in private. Curious to read about your validation scheme. Bravo and congarts!",
    "2868181": "Congratulations! I've just started studying, but my scores haven't been as good as I expected, which has made me a bit anxious. However, I feel that I can grow quickly with the help of people like you, who are willing to discuss, share, and provide guidance. Thank you.",
    "2868230": "Thanks for these tips and waiting for your solution writeup. 👍",
    "2868534": "The one advice is to try lots of things and test them rigorously. Test one thing at a time. Here, testing means submitting and looking at LB core given we did not have a good local validation setup.\n\nTest ideas and keep those who improve the LB.\n\nThis can lead to overfititng the LB however. \n\nTherefore do not stick to one run. Submit models trained with different random seeds, and consider the mean of their performance to be a more robust evaluation than single runs. You don't want to tune for specific random seeds. That's a recipe for disaster. Same applies to folds if you use kfold cv when training models. If scores across folds vary a lot, then use the mean as a robust score. Do not tune for the best fold.",
    "2869063": "Thank you for your advice. These insights are beneficial for me."
  },
  "source": "meta"
}