{
  "id": 368000,
  "title": "💡 how to move faster on an ML project and achieve more with less compute/time/energy (relevant to this competition)",
  "url": "/competitions/otto-recommender-system/discussion/368000",
  "author_name": "",
  "post_date": "2022-11-22T22:57:58.350075400Z",
  "votes": 31,
  "comment_count": 14,
  "views": 0,
  "content": "<p>Inspired by the very positive reception of my recent posts (<a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/367503\" target=\"_blank\">How to thrive in this competition without going crazy ❤️‍🔥</a>, let me share a couple of things that have been on my mind recently.</p>\n<p>A couple of techniques that will allow you to make faster progress on ML projects and that you can use in this competition!</p>\n<h3>Get out of the waiting hell</h3>\n<p>In any project, your time is the most valuable asset you have!</p>\n<p>So how does it normally play out?</p>\n<p>You tweak a small piece of code, run your notebook, and wait for 20 minutes only to find you made a typo and need to rerun the experiment again! 😱</p>\n<p>What is the solution? The <code>debug</code> flag!</p>\n<p>This is from the top of a notebook I am currently working on:</p>\n<pre><code>VALIDATE = True\nDEBUG = True\n\nSEED = 0\n</code></pre>\n<p>With <code>DEBUG == True</code> the notebook will execute on around 1% of data, if not less. It is blazingly fast.</p>\n<p>I am keeping this variable right at the top of my notebook to get in the habit of using it!</p>\n<p>Whenever I do, I save time. Whenever I don't, I end up rerunning the same code over and over again to fix small issues… This not only means I need to wait longer for the results, but I also need to keep spending my attention on babysitting the notebook! Troubleshooting it over and over again.</p>\n<p>With first running with <code>DEBUG == True</code>, I get to fix all the issues at one go and can only (in most cases) come back to evaluate the results.</p>\n<h3>Separate writing code from running experiments</h3>\n<p>This is a tough one though somewhat related to the point above.</p>\n<p>You need to make your life to stop being GIL locked and embrace paralellism as polars does natively 😄</p>\n<p>Essentially, you don't want your ability to write code (or do other meaningful things in life) to be blocked by waiting for the running of code to finish!</p>\n<p>Using the <code>DEBUG</code> flag is helpful, but it is only a start.</p>\n<p>The mindset is the key.</p>\n<p>I am trying to get in the mindset of separating writing code from running experiments (which is really hard to do when all you have at your disposal is your local machine, it is just so tempting to keep tweaking and running things live, but this changes you into a zombie, your brain is gone after an hour two of such an activity).</p>\n<p>So the general setup is:</p>\n<ol>\n<li>Write code that trains a model and get it running.</li>\n<li>Write code that runs experiments using the code above. Could be looking for better hyperparams, etc.</li>\n<li>Once you have written the code, RUN IT, and go do something else!</li>\n</ol>\n<p>It sounds super simple, but this right here holds the mystery to achieving more on the things that matter.</p>\n<h3>Prioritize relentlessly</h3>\n<p>The search space of an ML solution is immense. For all intents and purposes, it is infinite.</p>\n<p>It is so easy to go down a rabbit hole of tweaking just this one small part of a solution… You can spend weeks developing a better model of a given arch for instance.</p>\n<p>But what you need to do is this. Ask yourself -- what is the highest point of leverage in this project I could be working on? Usually, it will stand out as a sore thumb!</p>\n<p>And you don't even have to get it perfectly right. Just roughly working on something that has a chance of helping your solution is 1000x better than being pulled into solving an interesting but low-impact problem (AKA time sink).</p>\n<h3>Grow the solution organically</h3>\n<p>There were competitions where I got crushed by the complexity of my code. I started working on a part of a solution, then another, but it all became so complex that a) I lost interest in the competition b) it stopped being fun c) the task of integrating the pieces felt overwhelming and not worth it.</p>\n<p>The solution to this is to keep growing your code organically!</p>\n<p>What do I mean by this?</p>\n<p>You probably don't need to go super deep on tweaking our ranking model IF you don't have the code for outputting a submission!</p>\n<p>You probably don't need to integrate this super elaborate score from a super elaborate DL model until you figure out how to ensemble simple solutions.</p>\n<p>You probably don't need to go super duper deep on co-visitation matrices IF you know you are building towards a solution that will involve a ranking model.</p>\n<p>Better put in place an end-to-end pipeline of roughly what you'd like to attempt AND only then keep growing each of the components. And keep integrating the pieces as you go so that it all runs!</p>\n<p>This also guarantees you receive feedback on your changes and know whether you are introducing bugs or moving the right direction.</p>\n<h3>Summary</h3>\n<p>Hoping what I wrote above can be of genuine help to you 🙏 If you find it useful, please upvote, and tomorrow I will post how I am addressing dealing with RAM in this competition. You need so much of it!</p>\n<p>I nearly bought a fun piece of equipment to address this problem 😄 But decided otherwise as not sure where I would keep it plus have a bunch of other expenses so the family budget is not in a happy spot. But it is a good story I feel 🙂 And am using a different solution that I would like to share and that so far works!</p>\n<p>Anyhow, let's see how it goes and whether I am on the right track with these posts 🙂 Would hate to spam people with stuff they don't find interesting.</p>\n<p>So please upvote if you feel this is useful content and hopefully talk to you tomorrow! 🙂</p>\n<h3>Other resources you might find useful:</h3>\n<ul>\n<li><a href=\"https://www.kaggle.com/code/radek1/2-methods-how-to-ensemble-predictions\" target=\"_blank\">💡 [2 methods] How-to ensemble predictions 🏅🏅🏅</a></li>\n<li><a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/364991\" target=\"_blank\">local validation tracks public LB perfecty -- here is the setup</a></li>\n<li><a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/368560\" target=\"_blank\">💡 For my friends from Twitter and LinkedIn -- here is how to dive into this competition 🐳</a></li>\n<li><a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/363843\" target=\"_blank\">Full dataset processed to CSV/parquet files with optimized memory footprint</a></li>\n<li><a href=\"https://www.kaggle.com/code/radek1/co-visitation-matrix-simplified-imprvd-logic\" target=\"_blank\">co-visitation matrix - simplified, imprvd logic 🔥</a></li>\n<li><a href=\"https://www.kaggle.com/code/radek1/word2vec-how-to-training-and-submission\" target=\"_blank\">💡 Word2Vec How-to [training and submission]🚀🚀🚀</a></li>\n</ul>",
  "messages": [
    {
      "id": "2040323",
      "postDate": "11/22/2022 22:57:58",
      "content": "<p>Inspired by the very positive reception of my recent posts (<a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/367503\" target=\"_blank\">How to thrive in this competition without going crazy ❤️‍🔥</a>, let me share a couple of things that have been on my mind recently.</p>\n<p>A couple of techniques that will allow you to make faster progress on ML projects and that you can use in this competition!</p>\n<h3>Get out of the waiting hell</h3>\n<p>In any project, your time is the most valuable asset you have!</p>\n<p>So how does it normally play out?</p>\n<p>You tweak a small piece of code, run your notebook, and wait for 20 minutes only to find you made a typo and need to rerun the experiment again! 😱</p>\n<p>What is the solution? The <code>debug</code> flag!</p>\n<p>This is from the top of a notebook I am currently working on:</p>\n<pre><code>VALIDATE = True\nDEBUG = True\n\nSEED = 0\n</code></pre>\n<p>With <code>DEBUG == True</code> the notebook will execute on around 1% of data, if not less. It is blazingly fast.</p>\n<p>I am keeping this variable right at the top of my notebook to get in the habit of using it!</p>\n<p>Whenever I do, I save time. Whenever I don't, I end up rerunning the same code over and over again to fix small issues… This not only means I need to wait longer for the results, but I also need to keep spending my attention on babysitting the notebook! Troubleshooting it over and over again.</p>\n<p>With first running with <code>DEBUG == True</code>, I get to fix all the issues at one go and can only (in most cases) come back to evaluate the results.</p>\n<h3>Separate writing code from running experiments</h3>\n<p>This is a tough one though somewhat related to the point above.</p>\n<p>You need to make your life to stop being GIL locked and embrace paralellism as polars does natively 😄</p>\n<p>Essentially, you don't want your ability to write code (or do other meaningful things in life) to be blocked by waiting for the running of code to finish!</p>\n<p>Using the <code>DEBUG</code> flag is helpful, but it is only a start.</p>\n<p>The mindset is the key.</p>\n<p>I am trying to get in the mindset of separating writing code from running experiments (which is really hard to do when all you have at your disposal is your local machine, it is just so tempting to keep tweaking and running things live, but this changes you into a zombie, your brain is gone after an hour two of such an activity).</p>\n<p>So the general setup is:</p>\n<ol>\n<li>Write code that trains a model and get it running.</li>\n<li>Write code that runs experiments using the code above. Could be looking for better hyperparams, etc.</li>\n<li>Once you have written the code, RUN IT, and go do something else!</li>\n</ol>\n<p>It sounds super simple, but this right here holds the mystery to achieving more on the things that matter.</p>\n<h3>Prioritize relentlessly</h3>\n<p>The search space of an ML solution is immense. For all intents and purposes, it is infinite.</p>\n<p>It is so easy to go down a rabbit hole of tweaking just this one small part of a solution… You can spend weeks developing a better model of a given arch for instance.</p>\n<p>But what you need to do is this. Ask yourself -- what is the highest point of leverage in this project I could be working on? Usually, it will stand out as a sore thumb!</p>\n<p>And you don't even have to get it perfectly right. Just roughly working on something that has a chance of helping your solution is 1000x better than being pulled into solving an interesting but low-impact problem (AKA time sink).</p>\n<h3>Grow the solution organically</h3>\n<p>There were competitions where I got crushed by the complexity of my code. I started working on a part of a solution, then another, but it all became so complex that a) I lost interest in the competition b) it stopped being fun c) the task of integrating the pieces felt overwhelming and not worth it.</p>\n<p>The solution to this is to keep growing your code organically!</p>\n<p>What do I mean by this?</p>\n<p>You probably don't need to go super deep on tweaking our ranking model IF you don't have the code for outputting a submission!</p>\n<p>You probably don't need to integrate this super elaborate score from a super elaborate DL model until you figure out how to ensemble simple solutions.</p>\n<p>You probably don't need to go super duper deep on co-visitation matrices IF you know you are building towards a solution that will involve a ranking model.</p>\n<p>Better put in place an end-to-end pipeline of roughly what you'd like to attempt AND only then keep growing each of the components. And keep integrating the pieces as you go so that it all runs!</p>\n<p>This also guarantees you receive feedback on your changes and know whether you are introducing bugs or moving the right direction.</p>\n<h3>Summary</h3>\n<p>Hoping what I wrote above can be of genuine help to you 🙏 If you find it useful, please upvote, and tomorrow I will post how I am addressing dealing with RAM in this competition. You need so much of it!</p>\n<p>I nearly bought a fun piece of equipment to address this problem 😄 But decided otherwise as not sure where I would keep it plus have a bunch of other expenses so the family budget is not in a happy spot. But it is a good story I feel 🙂 And am using a different solution that I would like to share and that so far works!</p>\n<p>Anyhow, let's see how it goes and whether I am on the right track with these posts 🙂 Would hate to spam people with stuff they don't find interesting.</p>\n<p>So please upvote if you feel this is useful content and hopefully talk to you tomorrow! 🙂</p>\n<h3>Other resources you might find useful:</h3>\n<ul>\n<li><a href=\"https://www.kaggle.com/code/radek1/2-methods-how-to-ensemble-predictions\" target=\"_blank\">💡 [2 methods] How-to ensemble predictions 🏅🏅🏅</a></li>\n<li><a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/364991\" target=\"_blank\">local validation tracks public LB perfecty -- here is the setup</a></li>\n<li><a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/368560\" target=\"_blank\">💡 For my friends from Twitter and LinkedIn -- here is how to dive into this competition 🐳</a></li>\n<li><a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/363843\" target=\"_blank\">Full dataset processed to CSV/parquet files with optimized memory footprint</a></li>\n<li><a href=\"https://www.kaggle.com/code/radek1/co-visitation-matrix-simplified-imprvd-logic\" target=\"_blank\">co-visitation matrix - simplified, imprvd logic 🔥</a></li>\n<li><a href=\"https://www.kaggle.com/code/radek1/word2vec-how-to-training-and-submission\" target=\"_blank\">💡 Word2Vec How-to [training and submission]🚀🚀🚀</a></li>\n</ul>",
      "rawMarkdown": "Inspired by the very positive reception of my recent posts ([How to thrive in this competition without going crazy ❤️‍🔥](https://www.kaggle.com/competitions/otto-recommender-system/discussion/367503), let me share a couple of things that have been on my mind recently.\n\nA couple of techniques that will allow you to make faster progress on ML projects and that you can use in this competition!\n\n### Get out of the waiting hell\n\nIn any project, your time is the most valuable asset you have!\n\nSo how does it normally play out?\n\nYou tweak a small piece of code, run your notebook, and wait for 20 minutes only to find you made a typo and need to rerun the experiment again! 😱\n\nWhat is the solution? The `debug` flag!\n\nThis is from the top of a notebook I am currently working on:\n```\nVALIDATE = True\nDEBUG = True\n\nSEED = 0\n```\n\nWith `DEBUG == True` the notebook will execute on around 1% of data, if not less. It is blazingly fast.\n\nI am keeping this variable right at the top of my notebook to get in the habit of using it!\n\nWhenever I do, I save time. Whenever I don't, I end up rerunning the same code over and over again to fix small issues... This not only means I need to wait longer for the results, but I also need to keep spending my attention on babysitting the notebook! Troubleshooting it over and over again.\n\nWith first running with `DEBUG == True`, I get to fix all the issues at one go and can only (in most cases) come back to evaluate the results.\n\n### Separate writing code from running experiments\n\nThis is a tough one though somewhat related to the point above.\n\nYou need to make your life to stop being GIL locked and embrace paralellism as polars does natively 😄\n\nEssentially, you don't want your ability to write code (or do other meaningful things in life) to be blocked by waiting for the running of code to finish!\n\nUsing the `DEBUG` flag is helpful, but it is only a start.\n\nThe mindset is the key.\n\nI am trying to get in the mindset of separating writing code from running experiments (which is really hard to do when all you have at your disposal is your local machine, it is just so tempting to keep tweaking and running things live, but this changes you into a zombie, your brain is gone after an hour two of such an activity).\n\nSo the general setup is:\n1. Write code that trains a model and get it running.\n2. Write code that runs experiments using the code above. Could be looking for better hyperparams, etc.\n3. Once you have written the code, RUN IT, and go do something else!\n\nIt sounds super simple, but this right here holds the mystery to achieving more on the things that matter.\n\n### Prioritize relentlessly\n\nThe search space of an ML solution is immense. For all intents and purposes, it is infinite.\n\nIt is so easy to go down a rabbit hole of tweaking just this one small part of a solution... You can spend weeks developing a better model of a given arch for instance.\n\nBut what you need to do is this. Ask yourself -- what is the highest point of leverage in this project I could be working on? Usually, it will stand out as a sore thumb!\n\nAnd you don't even have to get it perfectly right. Just roughly working on something that has a chance of helping your solution is 1000x better than being pulled into solving an interesting but low-impact problem (AKA time sink).\n\n### Grow the solution organically\n\nThere were competitions where I got crushed by the complexity of my code. I started working on a part of a solution, then another, but it all became so complex that a) I lost interest in the competition b) it stopped being fun c) the task of integrating the pieces felt overwhelming and not worth it.\n\nThe solution to this is to keep growing your code organically!\n\nWhat do I mean by this?\n\nYou probably don't need to go super deep on tweaking our ranking model IF you don't have the code for outputting a submission!\n\nYou probably don't need to integrate this super elaborate score from a super elaborate DL model until you figure out how to ensemble simple solutions.\n\nYou probably don't need to go super duper deep on co-visitation matrices IF you know you are building towards a solution that will involve a ranking model.\n\nBetter put in place an end-to-end pipeline of roughly what you'd like to attempt AND only then keep growing each of the components. And keep integrating the pieces as you go so that it all runs!\n\nThis also guarantees you receive feedback on your changes and know whether you are introducing bugs or moving the right direction.\n\n### Summary\n\nHoping what I wrote above can be of genuine help to you 🙏 If you find it useful, please upvote, and tomorrow I will post how I am addressing dealing with RAM in this competition. You need so much of it!\n\nI nearly bought a fun piece of equipment to address this problem 😄 But decided otherwise as not sure where I would keep it plus have a bunch of other expenses so the family budget is not in a happy spot. But it is a good story I feel 🙂 And am using a different solution that I would like to share and that so far works!\n\nAnyhow, let's see how it goes and whether I am on the right track with these posts 🙂 Would hate to spam people with stuff they don't find interesting.\n\nSo please upvote if you feel this is useful content and hopefully talk to you tomorrow! 🙂\n\n### Other resources you might find useful:\n\n* [💡 [2 methods] How-to ensemble predictions 🏅🏅🏅](https://www.kaggle.com/code/radek1/2-methods-how-to-ensemble-predictions)\n* [local validation tracks public LB perfecty -- here is the setup](https://www.kaggle.com/competitions/otto-recommender-system/discussion/364991)\n* [💡 For my friends from Twitter and LinkedIn -- here is how to dive into this competition 🐳](https://www.kaggle.com/competitions/otto-recommender-system/discussion/368560)\n* [Full dataset processed to CSV/parquet files with optimized memory footprint](https://www.kaggle.com/competitions/otto-recommender-system/discussion/363843)\n* [co-visitation matrix - simplified, imprvd logic 🔥](https://www.kaggle.com/code/radek1/co-visitation-matrix-simplified-imprvd-logic)\n* [💡 Word2Vec How-to [training and submission]🚀🚀🚀](https://www.kaggle.com/code/radek1/word2vec-how-to-training-and-submission)",
      "votes": null
    },
    {
      "id": "2040470",
      "postDate": "11/23/2022 04:16:38",
      "content": "<p>Great advice! Thank you Radek, looking forward to your new posts! </p>",
      "rawMarkdown": "Great advice! Thank you Radek, looking forward to your new posts!",
      "votes": null
    },
    {
      "id": "2040475",
      "postDate": "11/23/2022 04:21:36",
      "content": "<p>Thank you very much, <a href=\"https://www.kaggle.com/zachary666\" target=\"_blank\">@zachary666</a>! Really appreciate your kind comment and the voice of confidence! Thank you! 🙏 </p>",
      "rawMarkdown": "Thank you very much, @zachary666! Really appreciate your kind comment and the voice of confidence! Thank you! 🙏",
      "votes": null
    },
    {
      "id": "2044501",
      "postDate": "11/26/2022 15:15:27",
      "content": "<p>Really can't keep up with your amazing sharing! So much great stuff from you! Thank you very much!</p>",
      "rawMarkdown": "Really can't keep up with your amazing sharing! So much great stuff from you! Thank you very much!",
      "votes": null
    },
    {
      "id": "2044924",
      "postDate": "11/26/2022 20:33:03",
      "content": "<p>Thank you for this. This is super useful.</p>",
      "rawMarkdown": "Thank you for this. This is super useful.",
      "votes": null
    },
    {
      "id": "2044964",
      "postDate": "11/26/2022 21:52:33",
      "content": "<p>Great to hear, <a href=\"https://www.kaggle.com/lordxerxes\" target=\"_blank\">@lordxerxes</a>! 🙂 Glad to be of help </p>",
      "rawMarkdown": "Great to hear, @lordxerxes! 🙂 Glad to be of help",
      "votes": null
    },
    {
      "id": "2044965",
      "postDate": "11/26/2022 21:53:12",
      "content": "<p>Haha, thanks <a href=\"https://www.kaggle.com/danielliao\" target=\"_blank\">@danielliao</a>! 🙂 Really glad you are finding this useful! </p>",
      "rawMarkdown": "Haha, thanks @danielliao! 🙂 Really glad you are finding this useful!",
      "votes": null
    },
    {
      "id": "2045229",
      "postDate": "11/27/2022 07:29:58",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/radek1\" target=\"_blank\">@radek1</a>, thank you so much for the advice on fast debugging with just 1% of the dataset. </p>\n<p>I have implemented the DEBUG option for <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> Candidate ReRank CV notebook in order to debug faster, the full dataset notebook runs just under 40 mins, and the DEBUG version is just under 4 mins. This is the best I can do without changing much of the original notebook. Is 4 min fast enough, what do you think?</p>\n<p>Is the DEBUG version solely for speeding up to see whether all codes runs without error? Does it do anything else for us?</p>\n<p>Below is the code added or changed (very little) to the original notebook. Do you mind to have a look to see whether I have done it properly? 🙏 Thanks!</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F850197%2F490c42bcf2869356acc1b9d33c638550%2FScreen%20Shot%202022-11-27%20at%2015.06.37.png?generation=1669533977515751&amp;alt=media\" alt=\"\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F850197%2Fb84e0c710cd9c313b33973021e98ea39%2FScreen%20Shot%202022-11-27%20at%2015.06.50.png?generation=1669533988729535&amp;alt=media\" alt=\"\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F850197%2F36c62b2f43c28941e130889d1560ffdf%2FScreen%20Shot%202022-11-27%20at%2015.06.58.png?generation=1669534000990752&amp;alt=media\" alt=\"\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F850197%2F9262c5a744e69262d26f316dfd87787e%2FScreen%20Shot%202022-11-27%20at%2015.07.07.png?generation=1669534011721440&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "Hi @radek1, thank you so much for the advice on fast debugging with just 1% of the dataset. \n\nI have implemented the DEBUG option for @cdeotte Candidate ReRank CV notebook in order to debug faster, the full dataset notebook runs just under 40 mins, and the DEBUG version is just under 4 mins. This is the best I can do without changing much of the original notebook. Is 4 min fast enough, what do you think?\n\nIs the DEBUG version solely for speeding up to see whether all codes runs without error? Does it do anything else for us?\n\nBelow is the code added or changed (very little) to the original notebook. Do you mind to have a look to see whether I have done it properly? 🙏 Thanks!\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F850197%2F490c42bcf2869356acc1b9d33c638550%2FScreen%20Shot%202022-11-27%20at%2015.06.37.png?generation=1669533977515751&alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F850197%2Fb84e0c710cd9c313b33973021e98ea39%2FScreen%20Shot%202022-11-27%20at%2015.06.50.png?generation=1669533988729535&alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F850197%2F36c62b2f43c28941e130889d1560ffdf%2FScreen%20Shot%202022-11-27%20at%2015.06.58.png?generation=1669534000990752&alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F850197%2F9262c5a744e69262d26f316dfd87787e%2FScreen%20Shot%202022-11-27%20at%2015.07.07.png?generation=1669534011721440&alt=media)",
      "votes": null
    },
    {
      "id": "2045245",
      "postDate": "11/27/2022 07:37:38",
      "content": "<blockquote>\n  <p>Is the DEBUG version solely for speeding up to see whether all codes runs without error? </p>\n</blockquote>\n<p>Yes, that is the way I do it 🙂</p>\n<p>The faster the code can run, the better 🙂 4 minutes is okay but in my experience, this would still interfere with the way I work on stuff. What are you going to do in those 4 minutes? I get distracted, go on Twitter, etc.</p>\n<p>Of course, life isn't perfect, I still often wait for stuff, but if I can just make sure the notebook runs in a couple of seconds before doing the full run, then I can take my attention and do something else, work on other stuff, without having to wait 🙂 But 4 minutes vs 40 minutes and finding out at the 39th minute you have a typo is already an enormous step in the right direction!</p>",
      "rawMarkdown": "> Is the DEBUG version solely for speeding up to see whether all codes runs without error? \n\nYes, that is the way I do it 🙂\n\nThe faster the code can run, the better 🙂 4 minutes is okay but in my experience, this would still interfere with the way I work on stuff. What are you going to do in those 4 minutes? I get distracted, go on Twitter, etc.\n\nOf course, life isn't perfect, I still often wait for stuff, but if I can just make sure the notebook runs in a couple of seconds before doing the full run, then I can take my attention and do something else, work on other stuff, without having to wait 🙂 But 4 minutes vs 40 minutes and finding out at the 39th minute you have a typo is already an enormous step in the right direction!",
      "votes": null
    },
    {
      "id": "2045328",
      "postDate": "11/27/2022 09:03:36",
      "content": "<p>Thank you so much Radek! You are absolutely right,  even 4 minutes we can still easily get distracted. It would be great if it can just take a few seconds to run. I will dig a little more to see whether I can make it faster.</p>",
      "rawMarkdown": "Thank you so much Radek! You are absolutely right,  even 4 minutes we can still easily get distracted. It would be great if it can just take a few seconds to run. I will dig a little more to see whether I can make it faster.",
      "votes": null
    },
    {
      "id": "2045409",
      "postDate": "11/27/2022 10:18:03",
      "content": "<p>Hi Radek, I am happy to report that I have reduced the time to 90s from 230s, adding a few more lines of code.</p>\n<p>Now, I managed to make it run within 34s. </p>\n<p>Without your great feedback I would dream to reduce the time from 4 mins to 30s. Thanks <a href=\"https://www.kaggle.com/radek1\" target=\"_blank\">@radek1</a> 🙏</p>",
      "rawMarkdown": "Hi Radek, I am happy to report that I have reduced the time to 90s from 230s, adding a few more lines of code.\n\nNow, I managed to make it run within 34s. \n\nWithout your great feedback I would dream to reduce the time from 4 mins to 30s. Thanks @radek1 🙏",
      "votes": null
    },
    {
      "id": "2120439",
      "postDate": "01/29/2023 15:44:20",
      "content": "<p>Great ! I didn't know about the first trick, thanks </p>",
      "rawMarkdown": "Great ! I didn't know about the first trick, thanks",
      "votes": null
    },
    {
      "id": "2120583",
      "postDate": "01/29/2023 17:32:21",
      "content": "<p><a href=\"https://www.kaggle.com/danielliao\" target=\"_blank\">@danielliao</a> , how were you able to do it?</p>",
      "rawMarkdown": "danielliao , how were you able to do it?",
      "votes": null
    },
    {
      "id": "2120585",
      "postDate": "01/29/2023 17:33:33",
      "content": "<p>Who later got to know how to run it in much less time?</p>",
      "rawMarkdown": "Who later got to know how to run it in much less time?",
      "votes": null
    },
    {
      "id": "2122254",
      "postDate": "01/30/2023 19:53:21",
      "content": "<p>My trick of that sort is to always round cv metric. It is important to avoid paying attention to those constantly changing multiple digits and tell youself honestly \"those changes didn't improve the model\". </p>",
      "rawMarkdown": "My trick of that sort is to always round cv metric. It is important to avoid paying attention to those constantly changing multiple digits and tell youself honestly \"those changes didn't improve the model\".",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2040470,
      "author_name": "zachary666",
      "author_url": "",
      "post_date": "11/23/2022 04:16:38",
      "content": "<p>Great advice! Thank you Radek, looking forward to your new posts! </p>",
      "votes": null,
      "replies": [
        {
          "id": 2040475,
          "author_name": "radek1",
          "author_url": "",
          "post_date": "11/23/2022 04:21:36",
          "content": "<p>Thank you very much, <a href=\"https://www.kaggle.com/zachary666\" target=\"_blank\">@zachary666</a>! Really appreciate your kind comment and the voice of confidence! Thank you! 🙏 </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2044501,
      "author_name": "danielliao",
      "author_url": "",
      "post_date": "11/26/2022 15:15:27",
      "content": "<p>Really can't keep up with your amazing sharing! So much great stuff from you! Thank you very much!</p>",
      "votes": null,
      "replies": [
        {
          "id": 2044965,
          "author_name": "radek1",
          "author_url": "",
          "post_date": "11/26/2022 21:53:12",
          "content": "<p>Haha, thanks <a href=\"https://www.kaggle.com/danielliao\" target=\"_blank\">@danielliao</a>! 🙂 Really glad you are finding this useful! </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2044924,
      "author_name": "lordxerxes",
      "author_url": "",
      "post_date": "11/26/2022 20:33:03",
      "content": "<p>Thank you for this. This is super useful.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2044964,
          "author_name": "radek1",
          "author_url": "",
          "post_date": "11/26/2022 21:52:33",
          "content": "<p>Great to hear, <a href=\"https://www.kaggle.com/lordxerxes\" target=\"_blank\">@lordxerxes</a>! 🙂 Glad to be of help </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2045229,
      "author_name": "danielliao",
      "author_url": "",
      "post_date": "11/27/2022 07:29:58",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/radek1\" target=\"_blank\">@radek1</a>, thank you so much for the advice on fast debugging with just 1% of the dataset. </p>\n<p>I have implemented the DEBUG option for <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> Candidate ReRank CV notebook in order to debug faster, the full dataset notebook runs just under 40 mins, and the DEBUG version is just under 4 mins. This is the best I can do without changing much of the original notebook. Is 4 min fast enough, what do you think?</p>\n<p>Is the DEBUG version solely for speeding up to see whether all codes runs without error? Does it do anything else for us?</p>\n<p>Below is the code added or changed (very little) to the original notebook. Do you mind to have a look to see whether I have done it properly? 🙏 Thanks!</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F850197%2F490c42bcf2869356acc1b9d33c638550%2FScreen%20Shot%202022-11-27%20at%2015.06.37.png?generation=1669533977515751&amp;alt=media\" alt=\"\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F850197%2Fb84e0c710cd9c313b33973021e98ea39%2FScreen%20Shot%202022-11-27%20at%2015.06.50.png?generation=1669533988729535&amp;alt=media\" alt=\"\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F850197%2F36c62b2f43c28941e130889d1560ffdf%2FScreen%20Shot%202022-11-27%20at%2015.06.58.png?generation=1669534000990752&amp;alt=media\" alt=\"\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F850197%2F9262c5a744e69262d26f316dfd87787e%2FScreen%20Shot%202022-11-27%20at%2015.07.07.png?generation=1669534011721440&amp;alt=media\" alt=\"\"></p>",
      "votes": null,
      "replies": [
        {
          "id": 2045245,
          "author_name": "radek1",
          "author_url": "",
          "post_date": "11/27/2022 07:37:38",
          "content": "<blockquote>\n  <p>Is the DEBUG version solely for speeding up to see whether all codes runs without error? </p>\n</blockquote>\n<p>Yes, that is the way I do it 🙂</p>\n<p>The faster the code can run, the better 🙂 4 minutes is okay but in my experience, this would still interfere with the way I work on stuff. What are you going to do in those 4 minutes? I get distracted, go on Twitter, etc.</p>\n<p>Of course, life isn't perfect, I still often wait for stuff, but if I can just make sure the notebook runs in a couple of seconds before doing the full run, then I can take my attention and do something else, work on other stuff, without having to wait 🙂 But 4 minutes vs 40 minutes and finding out at the 39th minute you have a typo is already an enormous step in the right direction!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 2045328,
          "author_name": "danielliao",
          "author_url": "",
          "post_date": "11/27/2022 09:03:36",
          "content": "<p>Thank you so much Radek! You are absolutely right,  even 4 minutes we can still easily get distracted. It would be great if it can just take a few seconds to run. I will dig a little more to see whether I can make it faster.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 2045409,
          "author_name": "danielliao",
          "author_url": "",
          "post_date": "11/27/2022 10:18:03",
          "content": "<p>Hi Radek, I am happy to report that I have reduced the time to 90s from 230s, adding a few more lines of code.</p>\n<p>Now, I managed to make it run within 34s. </p>\n<p>Without your great feedback I would dream to reduce the time from 4 mins to 30s. Thanks <a href=\"https://www.kaggle.com/radek1\" target=\"_blank\">@radek1</a> 🙏</p>",
          "votes": null,
          "replies": [
            {
              "id": 2120583,
              "author_name": "lordxerxes",
              "author_url": "",
              "post_date": "01/29/2023 17:32:21",
              "content": "<p><a href=\"https://www.kaggle.com/danielliao\" target=\"_blank\">@danielliao</a> , how were you able to do it?</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2120439,
      "author_name": "rayanaay",
      "author_url": "",
      "post_date": "01/29/2023 15:44:20",
      "content": "<p>Great ! I didn't know about the first trick, thanks </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2120585,
      "author_name": "lordxerxes",
      "author_url": "",
      "post_date": "01/29/2023 17:33:33",
      "content": "<p>Who later got to know how to run it in much less time?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2122254,
      "author_name": "artemfedorov",
      "author_url": "",
      "post_date": "01/30/2023 19:53:21",
      "content": "<p>My trick of that sort is to always round cv metric. It is important to avoid paying attention to those constantly changing multiple digits and tell youself honestly \"those changes didn't improve the model\". </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2040323": "Inspired by the very positive reception of my recent posts ([How to thrive in this competition without going crazy ❤️‍🔥](https://www.kaggle.com/competitions/otto-recommender-system/discussion/367503), let me share a couple of things that have been on my mind recently.\n\nA couple of techniques that will allow you to make faster progress on ML projects and that you can use in this competition!\n\n### Get out of the waiting hell\n\nIn any project, your time is the most valuable asset you have!\n\nSo how does it normally play out?\n\nYou tweak a small piece of code, run your notebook, and wait for 20 minutes only to find you made a typo and need to rerun the experiment again! 😱\n\nWhat is the solution? The `debug` flag!\n\nThis is from the top of a notebook I am currently working on:\n```\nVALIDATE = True\nDEBUG = True\n\nSEED = 0\n```\n\nWith `DEBUG == True` the notebook will execute on around 1% of data, if not less. It is blazingly fast.\n\nI am keeping this variable right at the top of my notebook to get in the habit of using it!\n\nWhenever I do, I save time. Whenever I don't, I end up rerunning the same code over and over again to fix small issues... This not only means I need to wait longer for the results, but I also need to keep spending my attention on babysitting the notebook! Troubleshooting it over and over again.\n\nWith first running with `DEBUG == True`, I get to fix all the issues at one go and can only (in most cases) come back to evaluate the results.\n\n### Separate writing code from running experiments\n\nThis is a tough one though somewhat related to the point above.\n\nYou need to make your life to stop being GIL locked and embrace paralellism as polars does natively 😄\n\nEssentially, you don't want your ability to write code (or do other meaningful things in life) to be blocked by waiting for the running of code to finish!\n\nUsing the `DEBUG` flag is helpful, but it is only a start.\n\nThe mindset is the key.\n\nI am trying to get in the mindset of separating writing code from running experiments (which is really hard to do when all you have at your disposal is your local machine, it is just so tempting to keep tweaking and running things live, but this changes you into a zombie, your brain is gone after an hour two of such an activity).\n\nSo the general setup is:\n1. Write code that trains a model and get it running.\n2. Write code that runs experiments using the code above. Could be looking for better hyperparams, etc.\n3. Once you have written the code, RUN IT, and go do something else!\n\nIt sounds super simple, but this right here holds the mystery to achieving more on the things that matter.\n\n### Prioritize relentlessly\n\nThe search space of an ML solution is immense. For all intents and purposes, it is infinite.\n\nIt is so easy to go down a rabbit hole of tweaking just this one small part of a solution... You can spend weeks developing a better model of a given arch for instance.\n\nBut what you need to do is this. Ask yourself -- what is the highest point of leverage in this project I could be working on? Usually, it will stand out as a sore thumb!\n\nAnd you don't even have to get it perfectly right. Just roughly working on something that has a chance of helping your solution is 1000x better than being pulled into solving an interesting but low-impact problem (AKA time sink).\n\n### Grow the solution organically\n\nThere were competitions where I got crushed by the complexity of my code. I started working on a part of a solution, then another, but it all became so complex that a) I lost interest in the competition b) it stopped being fun c) the task of integrating the pieces felt overwhelming and not worth it.\n\nThe solution to this is to keep growing your code organically!\n\nWhat do I mean by this?\n\nYou probably don't need to go super deep on tweaking our ranking model IF you don't have the code for outputting a submission!\n\nYou probably don't need to integrate this super elaborate score from a super elaborate DL model until you figure out how to ensemble simple solutions.\n\nYou probably don't need to go super duper deep on co-visitation matrices IF you know you are building towards a solution that will involve a ranking model.\n\nBetter put in place an end-to-end pipeline of roughly what you'd like to attempt AND only then keep growing each of the components. And keep integrating the pieces as you go so that it all runs!\n\nThis also guarantees you receive feedback on your changes and know whether you are introducing bugs or moving the right direction.\n\n### Summary\n\nHoping what I wrote above can be of genuine help to you 🙏 If you find it useful, please upvote, and tomorrow I will post how I am addressing dealing with RAM in this competition. You need so much of it!\n\nI nearly bought a fun piece of equipment to address this problem 😄 But decided otherwise as not sure where I would keep it plus have a bunch of other expenses so the family budget is not in a happy spot. But it is a good story I feel 🙂 And am using a different solution that I would like to share and that so far works!\n\nAnyhow, let's see how it goes and whether I am on the right track with these posts 🙂 Would hate to spam people with stuff they don't find interesting.\n\nSo please upvote if you feel this is useful content and hopefully talk to you tomorrow! 🙂\n\n### Other resources you might find useful:\n\n* [💡 [2 methods] How-to ensemble predictions 🏅🏅🏅](https://www.kaggle.com/code/radek1/2-methods-how-to-ensemble-predictions)\n* [local validation tracks public LB perfecty -- here is the setup](https://www.kaggle.com/competitions/otto-recommender-system/discussion/364991)\n* [💡 For my friends from Twitter and LinkedIn -- here is how to dive into this competition 🐳](https://www.kaggle.com/competitions/otto-recommender-system/discussion/368560)\n* [Full dataset processed to CSV/parquet files with optimized memory footprint](https://www.kaggle.com/competitions/otto-recommender-system/discussion/363843)\n* [co-visitation matrix - simplified, imprvd logic 🔥](https://www.kaggle.com/code/radek1/co-visitation-matrix-simplified-imprvd-logic)\n* [💡 Word2Vec How-to [training and submission]🚀🚀🚀](https://www.kaggle.com/code/radek1/word2vec-how-to-training-and-submission)",
    "2040470": "Great advice! Thank you Radek, looking forward to your new posts!",
    "2040475": "Thank you very much, @zachary666! Really appreciate your kind comment and the voice of confidence! Thank you! 🙏",
    "2044501": "Really can't keep up with your amazing sharing! So much great stuff from you! Thank you very much!",
    "2044924": "Thank you for this. This is super useful.",
    "2044964": "Great to hear, @lordxerxes! 🙂 Glad to be of help",
    "2044965": "Haha, thanks @danielliao! 🙂 Really glad you are finding this useful!",
    "2045229": "Hi @radek1, thank you so much for the advice on fast debugging with just 1% of the dataset. \n\nI have implemented the DEBUG option for @cdeotte Candidate ReRank CV notebook in order to debug faster, the full dataset notebook runs just under 40 mins, and the DEBUG version is just under 4 mins. This is the best I can do without changing much of the original notebook. Is 4 min fast enough, what do you think?\n\nIs the DEBUG version solely for speeding up to see whether all codes runs without error? Does it do anything else for us?\n\nBelow is the code added or changed (very little) to the original notebook. Do you mind to have a look to see whether I have done it properly? 🙏 Thanks!\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F850197%2F490c42bcf2869356acc1b9d33c638550%2FScreen%20Shot%202022-11-27%20at%2015.06.37.png?generation=1669533977515751&alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F850197%2Fb84e0c710cd9c313b33973021e98ea39%2FScreen%20Shot%202022-11-27%20at%2015.06.50.png?generation=1669533988729535&alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F850197%2F36c62b2f43c28941e130889d1560ffdf%2FScreen%20Shot%202022-11-27%20at%2015.06.58.png?generation=1669534000990752&alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F850197%2F9262c5a744e69262d26f316dfd87787e%2FScreen%20Shot%202022-11-27%20at%2015.07.07.png?generation=1669534011721440&alt=media)",
    "2045245": "> Is the DEBUG version solely for speeding up to see whether all codes runs without error? \n\nYes, that is the way I do it 🙂\n\nThe faster the code can run, the better 🙂 4 minutes is okay but in my experience, this would still interfere with the way I work on stuff. What are you going to do in those 4 minutes? I get distracted, go on Twitter, etc.\n\nOf course, life isn't perfect, I still often wait for stuff, but if I can just make sure the notebook runs in a couple of seconds before doing the full run, then I can take my attention and do something else, work on other stuff, without having to wait 🙂 But 4 minutes vs 40 minutes and finding out at the 39th minute you have a typo is already an enormous step in the right direction!",
    "2045328": "Thank you so much Radek! You are absolutely right,  even 4 minutes we can still easily get distracted. It would be great if it can just take a few seconds to run. I will dig a little more to see whether I can make it faster.",
    "2045409": "Hi Radek, I am happy to report that I have reduced the time to 90s from 230s, adding a few more lines of code.\n\nNow, I managed to make it run within 34s. \n\nWithout your great feedback I would dream to reduce the time from 4 mins to 30s. Thanks @radek1 🙏",
    "2120439": "Great ! I didn't know about the first trick, thanks",
    "2120583": "danielliao , how were you able to do it?",
    "2120585": "Who later got to know how to run it in much less time?",
    "2122254": "My trick of that sort is to always round cv metric. It is important to avoid paying attention to those constantly changing multiple digits and tell youself honestly \"those changes didn't improve the model\"."
  },
  "source": "meta"
}