{
  "id": 369789,
  "title": "To the people still hacking away on this competition...",
  "url": "/competitions/otto-recommender-system/discussion/369789",
  "author_name": "",
  "post_date": "2022-12-01T12:59:43.611954500Z",
  "votes": 20,
  "comment_count": 14,
  "views": 0,
  "content": "<p>I look at the LB and I see that people are submitting less and less…</p>\n<p>I look at the forums and I see the posting has died down…</p>\n<p>To the people still hacking on this competition, based on what little experience I have, I would like to say this:</p>\n<blockquote>\n  <p>You got this!</p>\n</blockquote>\n<p>One of the tricks, maybe the trick, to doing well in a Kaggle competition is to join early and to keep improving your solution every day. To keep reading the forums and keep trying out things.</p>\n<p>If you follow what some of the KGs say on Twitter, in their blog posts, you will find that even at the highest levels there is no substitute to experimentation and putting in the work.</p>\n<p>Yes, you may have more hardware to run your experiments on. Yes, you may have experience from previous competitions to fall back on. But this only gives you small, incremental improvements.</p>\n<p>You come up at the top by relentlessly hacking away at things.</p>\n<p>Anyhow, kudos to the people who are thinking about the problem every day 🙏 See you at the top when this competition finishes 🙂</p>\n<h3>Other resources you might find useful:</h3>\n<ul>\n<li><a href=\"https://www.kaggle.com/code/radek1/2-methods-how-to-ensemble-predictions\" target=\"_blank\">💡 [2 methods] How-to ensemble predictions 🏅🏅🏅</a></li>\n<li><a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/364991\" target=\"_blank\">local validation tracks public LB perfecty -- here is the setup</a></li>\n<li><a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/368560\" target=\"_blank\">💡 For my friends from Twitter and LinkedIn -- here is how to dive into this competition 🐳</a></li>\n<li><a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/363843\" target=\"_blank\">Full dataset processed to CSV/parquet files with optimized memory footprint</a></li>\n<li><a href=\"https://www.kaggle.com/code/radek1/co-visitation-matrix-simplified-imprvd-logic\" target=\"_blank\">co-visitation matrix - simplified, imprvd logic 🔥</a></li>\n<li><a href=\"https://www.kaggle.com/code/radek1/word2vec-how-to-training-and-submission\" target=\"_blank\">💡 Word2Vec How-to [training and submission]🚀🚀🚀</a></li>\n</ul>",
  "messages": [
    {
      "id": "2051477",
      "postDate": "12/01/2022 12:59:43",
      "content": "<p>I look at the LB and I see that people are submitting less and less…</p>\n<p>I look at the forums and I see the posting has died down…</p>\n<p>To the people still hacking on this competition, based on what little experience I have, I would like to say this:</p>\n<blockquote>\n  <p>You got this!</p>\n</blockquote>\n<p>One of the tricks, maybe the trick, to doing well in a Kaggle competition is to join early and to keep improving your solution every day. To keep reading the forums and keep trying out things.</p>\n<p>If you follow what some of the KGs say on Twitter, in their blog posts, you will find that even at the highest levels there is no substitute to experimentation and putting in the work.</p>\n<p>Yes, you may have more hardware to run your experiments on. Yes, you may have experience from previous competitions to fall back on. But this only gives you small, incremental improvements.</p>\n<p>You come up at the top by relentlessly hacking away at things.</p>\n<p>Anyhow, kudos to the people who are thinking about the problem every day 🙏 See you at the top when this competition finishes 🙂</p>\n<h3>Other resources you might find useful:</h3>\n<ul>\n<li><a href=\"https://www.kaggle.com/code/radek1/2-methods-how-to-ensemble-predictions\" target=\"_blank\">💡 [2 methods] How-to ensemble predictions 🏅🏅🏅</a></li>\n<li><a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/364991\" target=\"_blank\">local validation tracks public LB perfecty -- here is the setup</a></li>\n<li><a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/368560\" target=\"_blank\">💡 For my friends from Twitter and LinkedIn -- here is how to dive into this competition 🐳</a></li>\n<li><a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/363843\" target=\"_blank\">Full dataset processed to CSV/parquet files with optimized memory footprint</a></li>\n<li><a href=\"https://www.kaggle.com/code/radek1/co-visitation-matrix-simplified-imprvd-logic\" target=\"_blank\">co-visitation matrix - simplified, imprvd logic 🔥</a></li>\n<li><a href=\"https://www.kaggle.com/code/radek1/word2vec-how-to-training-and-submission\" target=\"_blank\">💡 Word2Vec How-to [training and submission]🚀🚀🚀</a></li>\n</ul>",
      "rawMarkdown": "I look at the LB and I see that people are submitting less and less...\n\nI look at the forums and I see the posting has died down...\n\nTo the people still hacking on this competition, based on what little experience I have, I would like to say this:\n\n> You got this!\n\nOne of the tricks, maybe the trick, to doing well in a Kaggle competition is to join early and to keep improving your solution every day. To keep reading the forums and keep trying out things.\n\nIf you follow what some of the KGs say on Twitter, in their blog posts, you will find that even at the highest levels there is no substitute to experimentation and putting in the work.\n\nYes, you may have more hardware to run your experiments on. Yes, you may have experience from previous competitions to fall back on. But this only gives you small, incremental improvements.\n\nYou come up at the top by relentlessly hacking away at things.\n\nAnyhow, kudos to the people who are thinking about the problem every day 🙏 See you at the top when this competition finishes 🙂\n\n### Other resources you might find useful:\n\n* [💡 [2 methods] How-to ensemble predictions 🏅🏅🏅](https://www.kaggle.com/code/radek1/2-methods-how-to-ensemble-predictions)\n* [local validation tracks public LB perfecty -- here is the setup](https://www.kaggle.com/competitions/otto-recommender-system/discussion/364991)\n* [💡 For my friends from Twitter and LinkedIn -- here is how to dive into this competition 🐳](https://www.kaggle.com/competitions/otto-recommender-system/discussion/368560)\n* [Full dataset processed to CSV/parquet files with optimized memory footprint](https://www.kaggle.com/competitions/otto-recommender-system/discussion/363843)\n* [co-visitation matrix - simplified, imprvd logic 🔥](https://www.kaggle.com/code/radek1/co-visitation-matrix-simplified-imprvd-logic)\n* [💡 Word2Vec How-to [training and submission]🚀🚀🚀](https://www.kaggle.com/code/radek1/word2vec-how-to-training-and-submission)",
      "votes": null
    },
    {
      "id": "2052050",
      "postDate": "12/01/2022 21:18:20",
      "content": "<p>Totally disagree with you guys! Perhaps, for long-training-runtime competitions, one can apply your advice (because you will run out of time to train models if you join too late). Even here, you can still join not too late. The early-joiners must be very hard working guys, otherwise you will tire and abandon before the end of the competition. Being able to keep motivated for three months is very difficult!</p>\n<p>Late-joiners have the edge of learning from others, starting with kernel prototypes, clean data, great insights from the forum, full list of unveiled pain points … Don't work hard, work smartly :). In fact, working hard is good too, but you have to be very strong then !</p>",
      "rawMarkdown": "Totally disagree with you guys! Perhaps, for long-training-runtime competitions, one can apply your advice (because you will run out of time to train models if you join too late). Even here, you can still join not too late. The early-joiners must be very hard working guys, otherwise you will tire and abandon before the end of the competition. Being able to keep motivated for three months is very difficult!\n\nLate-joiners have the edge of learning from others, starting with kernel prototypes, clean data, great insights from the forum, full list of unveiled pain points ... Don't work hard, work smartly :). In fact, working hard is good too, but you have to be very strong then !",
      "votes": null
    },
    {
      "id": "2052058",
      "postDate": "12/01/2022 21:47:56",
      "content": "<p>Thank you very much <a href=\"https://www.kaggle.com/kneroma\" target=\"_blank\">@kneroma</a> for sharing your perspective, really appreciate it! 🙂 Plus you have the medals to prove what you are talking about, that speaks a lot 🙂</p>\n<p>I suspect there is no one path. Possibly for a very large percentage of people who are maybe not total experts and are learning very actively working on a competition (nearly) every day from start is the way to go.</p>\n<p>Whenever I hear of someone joining a competition and getting to the gold medal range within 3 weeks, that feels to me like an extreme exception 🙂 But maybe that is happening much more broadly than I imagine it to be, would love to be corrected.</p>\n<p>Anyhow, really cool to have your perspective here 🙂 Thank you very much for sharing it 🙏</p>",
      "rawMarkdown": "Thank you very much @kneroma for sharing your perspective, really appreciate it! 🙂 Plus you have the medals to prove what you are talking about, that speaks a lot 🙂\n\nI suspect there is no one path. Possibly for a very large percentage of people who are maybe not total experts and are learning very actively working on a competition (nearly) every day from start is the way to go.\n\nWhenever I hear of someone joining a competition and getting to the gold medal range within 3 weeks, that feels to me like an extreme exception 🙂 But maybe that is happening much more broadly than I imagine it to be, would love to be corrected.\n\nAnyhow, really cool to have your perspective here 🙂 Thank you very much for sharing it 🙏",
      "votes": null
    },
    {
      "id": "2052079",
      "postDate": "12/01/2022 22:24:36",
      "content": "<p>Thank you also for sharing your points. We are all here to learn :). Indeed, there is no real single answer to this question. As I said, in no case should you join too late. \"Lazy\" guys can join a bit after the start and hard workers can join from the very beginning.</p>",
      "rawMarkdown": "Thank you also for sharing your points. We are all here to learn :). Indeed, there is no real single answer to this question. As I said, in no case should you join too late. \"Lazy\" guys can join a bit after the start and hard workers can join from the very beginning.",
      "votes": null
    },
    {
      "id": "2052240",
      "postDate": "12/02/2022 03:17:50",
      "content": "<p>\"The early-joiners must be very hard working guys, otherwise you will tire and abandon before the end of the competition. Being able to keep motivated for three months is very difficult!\"<br>\n<a href=\"https://www.kaggle.com/kneroma\" target=\"_blank\">@kneroma</a> this's somehow both right and wrong to me. In my last competition, I started very early and follow till the end (and we got the gold medal). But for the previous one, I almost \"tire and abandon\" as you said (we only got silver in that comp). The critical difference is whether you have enough ideas to do during that 3 months or not :D <br>\nTo someone newbie like me, I still prefer joining from the beginning, reading all the discussions and don't miss any public ideas as <a href=\"https://www.kaggle.com/radek1\" target=\"_blank\">@radek1</a> suggested. Getting to a high place in the first month is very exciting and motivating</p>",
      "rawMarkdown": "\"The early-joiners must be very hard working guys, otherwise you will tire and abandon before the end of the competition. Being able to keep motivated for three months is very difficult!\"\n@kneroma this's somehow both right and wrong to me. In my last competition, I started very early and follow till the end (and we got the gold medal). But for the previous one, I almost \"tire and abandon\" as you said (we only got silver in that comp). The critical difference is whether you have enough ideas to do during that 3 months or not :D \nTo someone newbie like me, I still prefer joining from the beginning, reading all the discussions and don't miss any public ideas as @radek1 suggested. Getting to a high place in the first month is very exciting and motivating",
      "votes": null
    },
    {
      "id": "2052579",
      "postDate": "12/02/2022 10:06:51",
      "content": "<p>I think competitions like this are an amazing source of inspiration to keep our skills current. Before this competition I had never heard of Polars. Now after playing with it for a couple of weeks I feel pretty competent. I've used it to good effect and now my code is literally 100x faster. Even if I don't rank highly in the end, I've learned something useful, and that's what counts!</p>",
      "rawMarkdown": "I think competitions like this are an amazing source of inspiration to keep our skills current. Before this competition I had never heard of Polars. Now after playing with it for a couple of weeks I feel pretty competent. I've used it to good effect and now my code is literally 100x faster. Even if I don't rank highly in the end, I've learned something useful, and that's what counts!",
      "votes": null
    },
    {
      "id": "2052580",
      "postDate": "12/02/2022 10:10:30",
      "content": "<p>It absolutely does, <a href=\"https://www.kaggle.com/johnwakefield\" target=\"_blank\">@johnwakefield</a>! 🙂 It is such a rare skill set among ML folks to be able to create solutions that can keep up with what people are dong on Kaggle.</p>\n<p>I do honestly feel this is a very valuable skill and am in the same boat as you in seeing Kaggle as an outstanding learning environment 🙂</p>\n<p>Here's to us continuing to learn in this competition! ☕ 🙂</p>",
      "rawMarkdown": "It absolutely does, @johnwakefield! 🙂 It is such a rare skill set among ML folks to be able to create solutions that can keep up with what people are dong on Kaggle.\n\nI do honestly feel this is a very valuable skill and am in the same boat as you in seeing Kaggle as an outstanding learning environment 🙂\n\nHere's to us continuing to learn in this competition! ☕ 🙂",
      "votes": null
    },
    {
      "id": "2055648",
      "postDate": "12/05/2022 09:20:21",
      "content": "<p>Many thanks for sharing !!! I learn so many things from your thread.</p>",
      "rawMarkdown": "Many thanks for sharing !!! I learn so many things from your thread.",
      "votes": null
    },
    {
      "id": "2055708",
      "postDate": "12/05/2022 10:49:02",
      "content": "<p>hey <a href=\"https://www.kaggle.com/marcuskk\" target=\"_blank\">@marcuskk</a>! That is great to hear! 🙂 Thank you for letting me know 🙏 </p>",
      "rawMarkdown": "hey @marcuskk! That is great to hear! 🙂 Thank you for letting me know 🙏",
      "votes": null
    },
    {
      "id": "2055783",
      "postDate": "12/05/2022 12:48:18",
      "content": "<p>if CV is well correlated with LB then there is no need to submit a lot.  Did you nitice a drop in submissions after you described your CV setting?</p>\n<p>And thanks for your sharing so far. I am reading through them little by little.</p>",
      "rawMarkdown": "if CV is well correlated with LB then there is no need to submit a lot.  Did you nitice a drop in submissions after you described your CV setting?\n\nAnd thanks for your sharing so far. I am reading through them little by little.",
      "votes": null
    },
    {
      "id": "2055788",
      "postDate": "12/05/2022 12:50:53",
      "content": "<p>Late joining backfires when training models takes a long time. For instance I joined ai4code 15 days before end, and my best model took 5 days to train. This leaves too little  room for experiments.</p>\n<p>In other competitions, training models takes few minute and late join is a good option.</p>",
      "rawMarkdown": "Late joining backfires when training models takes a long time. For instance I joined ai4code 15 days before end, and my best model took 5 days to train. This leaves too little  room for experiments.\n\nIn other competitions, training models takes few minute and late join is a good option.",
      "votes": null
    },
    {
      "id": "2055807",
      "postDate": "12/05/2022 13:04:56",
      "content": "<p>Thanks, JFP! It genuinely means a lot to me you are finding what I shared useful 🙂 Thank you!</p>\n<p>And very valuable observations here on how a good correlation between CV and LB and how long it takes to train impacts joining a competition early vs late 🙂 That is quite a nuanced take and awesome to be aware of it!</p>\n<p>If I am reading this right, essentially, if there is a good CV &amp; LB correlation, people are happy to use their CV and don't even bother submitting to LB all that much… I guess it doesn't add that much value and one effect I experienced is if you see someone jumping ahead you sort of realize that this is possible and work extra hard to catch up 🙂 Probably not something you generally want to encourage!</p>\n<p>But then again, this also makes me wonder -- in a competition without a strong CV - LB correlation probably submitting to the LB is used as a source of information? Ideally, we might want to modify our validation set to follow the LB… but in the absence of being able to do so, I guess you need to rely on what the LB is telling you… factor this into which submission you chose and what you work on… though relying on public LB is quite a dangerous proposition 🙂</p>\n<p>Anyhow, very interesting to think about this. Thank you for your comment!</p>",
      "rawMarkdown": "Thanks, JFP! It genuinely means a lot to me you are finding what I shared useful 🙂 Thank you!\n\nAnd very valuable observations here on how a good correlation between CV and LB and how long it takes to train impacts joining a competition early vs late 🙂 That is quite a nuanced take and awesome to be aware of it!\n\nIf I am reading this right, essentially, if there is a good CV & LB correlation, people are happy to use their CV and don't even bother submitting to LB all that much... I guess it doesn't add that much value and one effect I experienced is if you see someone jumping ahead you sort of realize that this is possible and work extra hard to catch up 🙂 Probably not something you generally want to encourage!\n\nBut then again, this also makes me wonder -- in a competition without a strong CV - LB correlation probably submitting to the LB is used as a source of information? Ideally, we might want to modify our validation set to follow the LB... but in the absence of being able to do so, I guess you need to rely on what the LB is telling you... factor this into which submission you chose and what you work on... though relying on public LB is quite a dangerous proposition 🙂\n\nAnyhow, very interesting to think about this. Thank you for your comment!",
      "votes": null
    },
    {
      "id": "2055856",
      "postDate": "12/05/2022 13:50:42",
      "content": "<p>I spend time until CV and LB are correlated. Always.</p>\n<p>I never use public LB as the main source for selecting models. This is probably why i resist shakeups usually.  But, it is true that in some competitions, public LB is well correlated with private LB and not correlated with CV score. In these rare cases, LB climbing is the right strategy.</p>",
      "rawMarkdown": "I spend time until CV and LB are correlated. Always.\n\nI never use public LB as the main source for selecting models. This is probably why i resist shakeups usually.  But, it is true that in some competitions, public LB is well correlated with private LB and not correlated with CV score. In these rare cases, LB climbing is the right strategy.",
      "votes": null
    },
    {
      "id": "2055861",
      "postDate": "12/05/2022 13:55:20",
      "content": "<p>Thank you very much for this answer! 🙂</p>",
      "rawMarkdown": "Thank you very much for this answer! 🙂",
      "votes": null
    },
    {
      "id": "2069167",
      "postDate": "12/18/2022 16:59:33",
      "content": "<p><strong>And one last thing:</strong> Remember that it is only \"a game\" at the end. Try to keep some balance in your life.</p>\n<p>The Devastator.</p>",
      "rawMarkdown": "**And one last thing:** Remember that it is only \"a game\" at the end. Try to keep some balance in your life.\n\nThe Devastator.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2052050,
      "author_name": "kneroma",
      "author_url": "",
      "post_date": "12/01/2022 21:18:20",
      "content": "<p>Totally disagree with you guys! Perhaps, for long-training-runtime competitions, one can apply your advice (because you will run out of time to train models if you join too late). Even here, you can still join not too late. The early-joiners must be very hard working guys, otherwise you will tire and abandon before the end of the competition. Being able to keep motivated for three months is very difficult!</p>\n<p>Late-joiners have the edge of learning from others, starting with kernel prototypes, clean data, great insights from the forum, full list of unveiled pain points … Don't work hard, work smartly :). In fact, working hard is good too, but you have to be very strong then !</p>",
      "votes": null,
      "replies": [
        {
          "id": 2052058,
          "author_name": "radek1",
          "author_url": "",
          "post_date": "12/01/2022 21:47:56",
          "content": "<p>Thank you very much <a href=\"https://www.kaggle.com/kneroma\" target=\"_blank\">@kneroma</a> for sharing your perspective, really appreciate it! 🙂 Plus you have the medals to prove what you are talking about, that speaks a lot 🙂</p>\n<p>I suspect there is no one path. Possibly for a very large percentage of people who are maybe not total experts and are learning very actively working on a competition (nearly) every day from start is the way to go.</p>\n<p>Whenever I hear of someone joining a competition and getting to the gold medal range within 3 weeks, that feels to me like an extreme exception 🙂 But maybe that is happening much more broadly than I imagine it to be, would love to be corrected.</p>\n<p>Anyhow, really cool to have your perspective here 🙂 Thank you very much for sharing it 🙏</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 2052079,
          "author_name": "kneroma",
          "author_url": "",
          "post_date": "12/01/2022 22:24:36",
          "content": "<p>Thank you also for sharing your points. We are all here to learn :). Indeed, there is no real single answer to this question. As I said, in no case should you join too late. \"Lazy\" guys can join a bit after the start and hard workers can join from the very beginning.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 2052240,
          "author_name": "minhtu123",
          "author_url": "",
          "post_date": "12/02/2022 03:17:50",
          "content": "<p>\"The early-joiners must be very hard working guys, otherwise you will tire and abandon before the end of the competition. Being able to keep motivated for three months is very difficult!\"<br>\n<a href=\"https://www.kaggle.com/kneroma\" target=\"_blank\">@kneroma</a> this's somehow both right and wrong to me. In my last competition, I started very early and follow till the end (and we got the gold medal). But for the previous one, I almost \"tire and abandon\" as you said (we only got silver in that comp). The critical difference is whether you have enough ideas to do during that 3 months or not :D <br>\nTo someone newbie like me, I still prefer joining from the beginning, reading all the discussions and don't miss any public ideas as <a href=\"https://www.kaggle.com/radek1\" target=\"_blank\">@radek1</a> suggested. Getting to a high place in the first month is very exciting and motivating</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 2055788,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "12/05/2022 12:50:53",
          "content": "<p>Late joining backfires when training models takes a long time. For instance I joined ai4code 15 days before end, and my best model took 5 days to train. This leaves too little  room for experiments.</p>\n<p>In other competitions, training models takes few minute and late join is a good option.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2052579,
      "author_name": "johnwakefield",
      "author_url": "",
      "post_date": "12/02/2022 10:06:51",
      "content": "<p>I think competitions like this are an amazing source of inspiration to keep our skills current. Before this competition I had never heard of Polars. Now after playing with it for a couple of weeks I feel pretty competent. I've used it to good effect and now my code is literally 100x faster. Even if I don't rank highly in the end, I've learned something useful, and that's what counts!</p>",
      "votes": null,
      "replies": [
        {
          "id": 2052580,
          "author_name": "radek1",
          "author_url": "",
          "post_date": "12/02/2022 10:10:30",
          "content": "<p>It absolutely does, <a href=\"https://www.kaggle.com/johnwakefield\" target=\"_blank\">@johnwakefield</a>! 🙂 It is such a rare skill set among ML folks to be able to create solutions that can keep up with what people are dong on Kaggle.</p>\n<p>I do honestly feel this is a very valuable skill and am in the same boat as you in seeing Kaggle as an outstanding learning environment 🙂</p>\n<p>Here's to us continuing to learn in this competition! ☕ 🙂</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2055648,
      "author_name": "marcuskk",
      "author_url": "",
      "post_date": "12/05/2022 09:20:21",
      "content": "<p>Many thanks for sharing !!! I learn so many things from your thread.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2055708,
          "author_name": "radek1",
          "author_url": "",
          "post_date": "12/05/2022 10:49:02",
          "content": "<p>hey <a href=\"https://www.kaggle.com/marcuskk\" target=\"_blank\">@marcuskk</a>! That is great to hear! 🙂 Thank you for letting me know 🙏 </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2055783,
      "author_name": "cpmpml",
      "author_url": "",
      "post_date": "12/05/2022 12:48:18",
      "content": "<p>if CV is well correlated with LB then there is no need to submit a lot.  Did you nitice a drop in submissions after you described your CV setting?</p>\n<p>And thanks for your sharing so far. I am reading through them little by little.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2055807,
          "author_name": "radek1",
          "author_url": "",
          "post_date": "12/05/2022 13:04:56",
          "content": "<p>Thanks, JFP! It genuinely means a lot to me you are finding what I shared useful 🙂 Thank you!</p>\n<p>And very valuable observations here on how a good correlation between CV and LB and how long it takes to train impacts joining a competition early vs late 🙂 That is quite a nuanced take and awesome to be aware of it!</p>\n<p>If I am reading this right, essentially, if there is a good CV &amp; LB correlation, people are happy to use their CV and don't even bother submitting to LB all that much… I guess it doesn't add that much value and one effect I experienced is if you see someone jumping ahead you sort of realize that this is possible and work extra hard to catch up 🙂 Probably not something you generally want to encourage!</p>\n<p>But then again, this also makes me wonder -- in a competition without a strong CV - LB correlation probably submitting to the LB is used as a source of information? Ideally, we might want to modify our validation set to follow the LB… but in the absence of being able to do so, I guess you need to rely on what the LB is telling you… factor this into which submission you chose and what you work on… though relying on public LB is quite a dangerous proposition 🙂</p>\n<p>Anyhow, very interesting to think about this. Thank you for your comment!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 2055856,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "12/05/2022 13:50:42",
          "content": "<p>I spend time until CV and LB are correlated. Always.</p>\n<p>I never use public LB as the main source for selecting models. This is probably why i resist shakeups usually.  But, it is true that in some competitions, public LB is well correlated with private LB and not correlated with CV score. In these rare cases, LB climbing is the right strategy.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 2055861,
          "author_name": "radek1",
          "author_url": "",
          "post_date": "12/05/2022 13:55:20",
          "content": "<p>Thank you very much for this answer! 🙂</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2069167,
      "author_name": "thedevastator",
      "author_url": "",
      "post_date": "12/18/2022 16:59:33",
      "content": "<p><strong>And one last thing:</strong> Remember that it is only \"a game\" at the end. Try to keep some balance in your life.</p>\n<p>The Devastator.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2051477": "I look at the LB and I see that people are submitting less and less...\n\nI look at the forums and I see the posting has died down...\n\nTo the people still hacking on this competition, based on what little experience I have, I would like to say this:\n\n> You got this!\n\nOne of the tricks, maybe the trick, to doing well in a Kaggle competition is to join early and to keep improving your solution every day. To keep reading the forums and keep trying out things.\n\nIf you follow what some of the KGs say on Twitter, in their blog posts, you will find that even at the highest levels there is no substitute to experimentation and putting in the work.\n\nYes, you may have more hardware to run your experiments on. Yes, you may have experience from previous competitions to fall back on. But this only gives you small, incremental improvements.\n\nYou come up at the top by relentlessly hacking away at things.\n\nAnyhow, kudos to the people who are thinking about the problem every day 🙏 See you at the top when this competition finishes 🙂\n\n### Other resources you might find useful:\n\n* [💡 [2 methods] How-to ensemble predictions 🏅🏅🏅](https://www.kaggle.com/code/radek1/2-methods-how-to-ensemble-predictions)\n* [local validation tracks public LB perfecty -- here is the setup](https://www.kaggle.com/competitions/otto-recommender-system/discussion/364991)\n* [💡 For my friends from Twitter and LinkedIn -- here is how to dive into this competition 🐳](https://www.kaggle.com/competitions/otto-recommender-system/discussion/368560)\n* [Full dataset processed to CSV/parquet files with optimized memory footprint](https://www.kaggle.com/competitions/otto-recommender-system/discussion/363843)\n* [co-visitation matrix - simplified, imprvd logic 🔥](https://www.kaggle.com/code/radek1/co-visitation-matrix-simplified-imprvd-logic)\n* [💡 Word2Vec How-to [training and submission]🚀🚀🚀](https://www.kaggle.com/code/radek1/word2vec-how-to-training-and-submission)",
    "2052050": "Totally disagree with you guys! Perhaps, for long-training-runtime competitions, one can apply your advice (because you will run out of time to train models if you join too late). Even here, you can still join not too late. The early-joiners must be very hard working guys, otherwise you will tire and abandon before the end of the competition. Being able to keep motivated for three months is very difficult!\n\nLate-joiners have the edge of learning from others, starting with kernel prototypes, clean data, great insights from the forum, full list of unveiled pain points ... Don't work hard, work smartly :). In fact, working hard is good too, but you have to be very strong then !",
    "2052058": "Thank you very much @kneroma for sharing your perspective, really appreciate it! 🙂 Plus you have the medals to prove what you are talking about, that speaks a lot 🙂\n\nI suspect there is no one path. Possibly for a very large percentage of people who are maybe not total experts and are learning very actively working on a competition (nearly) every day from start is the way to go.\n\nWhenever I hear of someone joining a competition and getting to the gold medal range within 3 weeks, that feels to me like an extreme exception 🙂 But maybe that is happening much more broadly than I imagine it to be, would love to be corrected.\n\nAnyhow, really cool to have your perspective here 🙂 Thank you very much for sharing it 🙏",
    "2052079": "Thank you also for sharing your points. We are all here to learn :). Indeed, there is no real single answer to this question. As I said, in no case should you join too late. \"Lazy\" guys can join a bit after the start and hard workers can join from the very beginning.",
    "2052240": "\"The early-joiners must be very hard working guys, otherwise you will tire and abandon before the end of the competition. Being able to keep motivated for three months is very difficult!\"\n@kneroma this's somehow both right and wrong to me. In my last competition, I started very early and follow till the end (and we got the gold medal). But for the previous one, I almost \"tire and abandon\" as you said (we only got silver in that comp). The critical difference is whether you have enough ideas to do during that 3 months or not :D \nTo someone newbie like me, I still prefer joining from the beginning, reading all the discussions and don't miss any public ideas as @radek1 suggested. Getting to a high place in the first month is very exciting and motivating",
    "2052579": "I think competitions like this are an amazing source of inspiration to keep our skills current. Before this competition I had never heard of Polars. Now after playing with it for a couple of weeks I feel pretty competent. I've used it to good effect and now my code is literally 100x faster. Even if I don't rank highly in the end, I've learned something useful, and that's what counts!",
    "2052580": "It absolutely does, @johnwakefield! 🙂 It is such a rare skill set among ML folks to be able to create solutions that can keep up with what people are dong on Kaggle.\n\nI do honestly feel this is a very valuable skill and am in the same boat as you in seeing Kaggle as an outstanding learning environment 🙂\n\nHere's to us continuing to learn in this competition! ☕ 🙂",
    "2055648": "Many thanks for sharing !!! I learn so many things from your thread.",
    "2055708": "hey @marcuskk! That is great to hear! 🙂 Thank you for letting me know 🙏",
    "2055783": "if CV is well correlated with LB then there is no need to submit a lot.  Did you nitice a drop in submissions after you described your CV setting?\n\nAnd thanks for your sharing so far. I am reading through them little by little.",
    "2055788": "Late joining backfires when training models takes a long time. For instance I joined ai4code 15 days before end, and my best model took 5 days to train. This leaves too little  room for experiments.\n\nIn other competitions, training models takes few minute and late join is a good option.",
    "2055807": "Thanks, JFP! It genuinely means a lot to me you are finding what I shared useful 🙂 Thank you!\n\nAnd very valuable observations here on how a good correlation between CV and LB and how long it takes to train impacts joining a competition early vs late 🙂 That is quite a nuanced take and awesome to be aware of it!\n\nIf I am reading this right, essentially, if there is a good CV & LB correlation, people are happy to use their CV and don't even bother submitting to LB all that much... I guess it doesn't add that much value and one effect I experienced is if you see someone jumping ahead you sort of realize that this is possible and work extra hard to catch up 🙂 Probably not something you generally want to encourage!\n\nBut then again, this also makes me wonder -- in a competition without a strong CV - LB correlation probably submitting to the LB is used as a source of information? Ideally, we might want to modify our validation set to follow the LB... but in the absence of being able to do so, I guess you need to rely on what the LB is telling you... factor this into which submission you chose and what you work on... though relying on public LB is quite a dangerous proposition 🙂\n\nAnyhow, very interesting to think about this. Thank you for your comment!",
    "2055856": "I spend time until CV and LB are correlated. Always.\n\nI never use public LB as the main source for selecting models. This is probably why i resist shakeups usually.  But, it is true that in some competitions, public LB is well correlated with private LB and not correlated with CV score. In these rare cases, LB climbing is the right strategy.",
    "2055861": "Thank you very much for this answer! 🙂",
    "2069167": "**And one last thing:** Remember that it is only \"a game\" at the end. Try to keep some balance in your life.\n\nThe Devastator."
  },
  "source": "meta"
}