{
  "id": 367754,
  "title": "How to thrive in this competition without going crazy (part 2 of 2) ❤️‍🔥",
  "url": "/competitions/otto-recommender-system/discussion/367754",
  "author_name": "",
  "post_date": "2022-11-22T00:24:43.048199700Z",
  "votes": 47,
  "comment_count": 3,
  "views": 0,
  "content": "<p>This competition is extremely complex.</p>\n<p>We are presented with this seemingly simple data, but THERE IS SO MUCH you can do with it. Just note the explosion of Kaggle kernels sharing ideas already (co-visitation matrices, matrix factorization, reranking, word2vec, training sequential DL models…)</p>\n<p>In <a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/367503\" target=\"_blank\">part 1 of this two-post series</a> I talked about one aspect of combating complexity -- how to navigate the landscape of ideas to identify what to work on and how to structure your solution. The answer has been simple and could be summed up in a single word: <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> 😄</p>\n<p>In this post let's look at the other type of complexity we need to deal with to be successful - the complexity of the code that we will write/have to work with.</p>\n<p>This type of complexity can crush you extremely quickly. In fact, I only feel I stand a chance in this competition to do something interesting because I got crushed by code complexity in ML projects so many times before!</p>\n<p>What can the pipeline in this competition look like?</p>\n<p><code>create datasets to train on and for evaluation -&gt; train MF/word2vec -&gt; create a covisitation matrix (separate for train with validation and for the full train set for submission) -&gt; create features &amp; diagnostic code (measure hit rate, recall) for ranking models -&gt; train a ranking model -&gt; create submission based on this output</code></p>\n<p>And that is a simplification of what the final pipeline will look like! There are so many things you can do along the way (for instance, what could be the equivalent of training with CV in this competition? what interesting things can you do to the data? what about training ranking models with different seeds? what about training ranking models on subsets of columns? how about doing feature selection? what about combining outputs from multiple ranking models?)</p>\n<p>This is a humongous undertaking… one that the field of Machine Learning doesn't have a solution to!</p>\n<p>With infinite resources, you could use something like <a href=\"https://twitter.com/radekosmulski/status/1556475294084526081?s=20&amp;t=E88ywgWI8RZMQIhOh4_0eQ\" target=\"_blank\">Metaflow</a>. But a) most Kagglers don't have the money to spend thousands of dollars on infra b) most Kagglers don't know where to begin setting up stuff like this.</p>\n<p>So what can we do? We need to solve this problem with the right tool for the job given our circumstances.</p>\n<p>And the answer, at its core, is just a single word. Refactoring.</p>\n<p>YOU HAVE TO TIRELESSLY REFACTOR. It is not enough that your code runs and seems to output something useful.</p>\n<p>You need to make it readable. You (ideally) should make it composable.</p>\n<p>But let's stop with these big words. People create a lot of mystique about relatively straightforward things. Let's not fall into that trap here.</p>\n<p>The act of refactoring is very simple (though <a href=\"https://www.amazon.com.au/Refactoring-Martin-Fowler/dp/0134757599/ref=asc_df_0134757599/?tag=googleshopdsk-22&amp;linkCode=df0&amp;hvadid=341743255824&amp;hvpos=&amp;hvnetw=g&amp;hvrand=1048400666557196734&amp;hvpone=&amp;hvptwo=&amp;hvqmt=&amp;hvdev=c&amp;hvdvcmdl=&amp;hvlocint=&amp;hvlocphy=9069055&amp;hvtargid=pla-464425925893&amp;psc=1\" target=\"_blank\">some absolutely kickass books</a> have been written on it)</p>\n<p>A couple of very powerful components of refactoring techniques are:</p>\n<ul>\n<li>naming</li>\n<li>create functions to communicate intention (related to 👆)</li>\n</ul>\n<p>What is the idea behind naming? Use variable names that communicate what the variable holds. Trade being succinct for readability. Always.</p>\n<p>Here are a couple of variable names from my code:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F83267%2F6683c4e7ae54f9a4b2ba79c6c9f10abe%2Fvariable_names.png?generation=1669071906008304&amp;alt=media\" alt=\"\"></p>\n<p>I really couldn't care less how long my variable names are AS LONG AS I UNDERSTAND what they hold.</p>\n<p>Creating Kaggle kernels is a great exercise in writing readable code. In doing so you learn how to express your ideas so that people with varying levels of experience will be able to follow along.</p>\n<p>And as for \"creating functions to communicate intent\", this is taking the idea of naming to another level.</p>\n<p>This is a function from my code that I am using so far in a single place. Yes, you want to create functions so that you do not repeat yourself (DRY). But even if I had this bit of code in a single place in my notebook, it might still be valuable to wrap it up in a function.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F83267%2F55ef1fe5123c4dd0bede945f4a8fed60%2Ffunctions_to_communicate_intent.png?generation=1669072142790250&amp;alt=media\" alt=\"\"></p>\n<p>How much easier it is to understand what is happening in your code when you encounter <code>cast_to_df_dtypes</code> vs if you were to encounter just the blob of code that is in the body of the method?</p>\n<p>Code is meant to be read by humans, the fact that it runs on a computer is secondary. Code is meant to express (and communicate!) ideas to other human beings. Be that someone reading your code who might have just picked up Python a month ago, or you from the future.</p>\n<p>I am now experimenting with <a href=\"https://twitter.com/radekosmulski/status/1594556992424931328?s=20&amp;t=-7XnTDzGouCnQrvOYKnJJA\" target=\"_blank\">execnb by fast.ai</a> and am thinking this might be another way to reduce the complexity of my code. We'll see.</p>\n<p>But it all starts with refactoring 🙂</p>\n<p>And the wildest story is I have been looking for resources on my Twitter on refactoring to share with you and it turns out I created <a href=\"https://github.com/radekosmulski/refactoring\" target=\"_blank\">a repository on refactoring</a> three years ago and completely forgot about 😳 But some really good stuff there that I accumulated over months if not years.</p>\n<p>Anyhow, hope this can be of help 🙂 All these ideas are experimental at best, but will try to put them into action in this competition (if I manage to find the time and not get pulled into some other direction 😉) and will let you know how it goes</p>\n<h3>Other resources you might find useful:</h3>\n<ul>\n<li><a href=\"https://www.kaggle.com/code/radek1/2-methods-how-to-ensemble-predictions\" target=\"_blank\">💡 [2 methods] How-to ensemble predictions 🏅🏅🏅</a></li>\n<li><a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/364991\" target=\"_blank\">local validation tracks public LB perfecty -- here is the setup</a></li>\n<li><a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/368560\" target=\"_blank\">💡 For my friends from Twitter and LinkedIn -- here is how to dive into this competition 🐳</a></li>\n<li><a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/363843\" target=\"_blank\">Full dataset processed to CSV/parquet files with optimized memory footprint</a></li>\n<li><a href=\"https://www.kaggle.com/code/radek1/co-visitation-matrix-simplified-imprvd-logic\" target=\"_blank\">co-visitation matrix - simplified, imprvd logic 🔥</a></li>\n<li><a href=\"https://www.kaggle.com/code/radek1/word2vec-how-to-training-and-submission\" target=\"_blank\">💡 Word2Vec How-to [training and submission]🚀🚀🚀</a></li>\n</ul>",
  "messages": [
    {
      "id": "2039145",
      "postDate": "11/22/2022 00:24:43",
      "content": "<p>This competition is extremely complex.</p>\n<p>We are presented with this seemingly simple data, but THERE IS SO MUCH you can do with it. Just note the explosion of Kaggle kernels sharing ideas already (co-visitation matrices, matrix factorization, reranking, word2vec, training sequential DL models…)</p>\n<p>In <a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/367503\" target=\"_blank\">part 1 of this two-post series</a> I talked about one aspect of combating complexity -- how to navigate the landscape of ideas to identify what to work on and how to structure your solution. The answer has been simple and could be summed up in a single word: <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> 😄</p>\n<p>In this post let's look at the other type of complexity we need to deal with to be successful - the complexity of the code that we will write/have to work with.</p>\n<p>This type of complexity can crush you extremely quickly. In fact, I only feel I stand a chance in this competition to do something interesting because I got crushed by code complexity in ML projects so many times before!</p>\n<p>What can the pipeline in this competition look like?</p>\n<p><code>create datasets to train on and for evaluation -&gt; train MF/word2vec -&gt; create a covisitation matrix (separate for train with validation and for the full train set for submission) -&gt; create features &amp; diagnostic code (measure hit rate, recall) for ranking models -&gt; train a ranking model -&gt; create submission based on this output</code></p>\n<p>And that is a simplification of what the final pipeline will look like! There are so many things you can do along the way (for instance, what could be the equivalent of training with CV in this competition? what interesting things can you do to the data? what about training ranking models with different seeds? what about training ranking models on subsets of columns? how about doing feature selection? what about combining outputs from multiple ranking models?)</p>\n<p>This is a humongous undertaking… one that the field of Machine Learning doesn't have a solution to!</p>\n<p>With infinite resources, you could use something like <a href=\"https://twitter.com/radekosmulski/status/1556475294084526081?s=20&amp;t=E88ywgWI8RZMQIhOh4_0eQ\" target=\"_blank\">Metaflow</a>. But a) most Kagglers don't have the money to spend thousands of dollars on infra b) most Kagglers don't know where to begin setting up stuff like this.</p>\n<p>So what can we do? We need to solve this problem with the right tool for the job given our circumstances.</p>\n<p>And the answer, at its core, is just a single word. Refactoring.</p>\n<p>YOU HAVE TO TIRELESSLY REFACTOR. It is not enough that your code runs and seems to output something useful.</p>\n<p>You need to make it readable. You (ideally) should make it composable.</p>\n<p>But let's stop with these big words. People create a lot of mystique about relatively straightforward things. Let's not fall into that trap here.</p>\n<p>The act of refactoring is very simple (though <a href=\"https://www.amazon.com.au/Refactoring-Martin-Fowler/dp/0134757599/ref=asc_df_0134757599/?tag=googleshopdsk-22&amp;linkCode=df0&amp;hvadid=341743255824&amp;hvpos=&amp;hvnetw=g&amp;hvrand=1048400666557196734&amp;hvpone=&amp;hvptwo=&amp;hvqmt=&amp;hvdev=c&amp;hvdvcmdl=&amp;hvlocint=&amp;hvlocphy=9069055&amp;hvtargid=pla-464425925893&amp;psc=1\" target=\"_blank\">some absolutely kickass books</a> have been written on it)</p>\n<p>A couple of very powerful components of refactoring techniques are:</p>\n<ul>\n<li>naming</li>\n<li>create functions to communicate intention (related to 👆)</li>\n</ul>\n<p>What is the idea behind naming? Use variable names that communicate what the variable holds. Trade being succinct for readability. Always.</p>\n<p>Here are a couple of variable names from my code:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F83267%2F6683c4e7ae54f9a4b2ba79c6c9f10abe%2Fvariable_names.png?generation=1669071906008304&amp;alt=media\" alt=\"\"></p>\n<p>I really couldn't care less how long my variable names are AS LONG AS I UNDERSTAND what they hold.</p>\n<p>Creating Kaggle kernels is a great exercise in writing readable code. In doing so you learn how to express your ideas so that people with varying levels of experience will be able to follow along.</p>\n<p>And as for \"creating functions to communicate intent\", this is taking the idea of naming to another level.</p>\n<p>This is a function from my code that I am using so far in a single place. Yes, you want to create functions so that you do not repeat yourself (DRY). But even if I had this bit of code in a single place in my notebook, it might still be valuable to wrap it up in a function.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F83267%2F55ef1fe5123c4dd0bede945f4a8fed60%2Ffunctions_to_communicate_intent.png?generation=1669072142790250&amp;alt=media\" alt=\"\"></p>\n<p>How much easier it is to understand what is happening in your code when you encounter <code>cast_to_df_dtypes</code> vs if you were to encounter just the blob of code that is in the body of the method?</p>\n<p>Code is meant to be read by humans, the fact that it runs on a computer is secondary. Code is meant to express (and communicate!) ideas to other human beings. Be that someone reading your code who might have just picked up Python a month ago, or you from the future.</p>\n<p>I am now experimenting with <a href=\"https://twitter.com/radekosmulski/status/1594556992424931328?s=20&amp;t=-7XnTDzGouCnQrvOYKnJJA\" target=\"_blank\">execnb by fast.ai</a> and am thinking this might be another way to reduce the complexity of my code. We'll see.</p>\n<p>But it all starts with refactoring 🙂</p>\n<p>And the wildest story is I have been looking for resources on my Twitter on refactoring to share with you and it turns out I created <a href=\"https://github.com/radekosmulski/refactoring\" target=\"_blank\">a repository on refactoring</a> three years ago and completely forgot about 😳 But some really good stuff there that I accumulated over months if not years.</p>\n<p>Anyhow, hope this can be of help 🙂 All these ideas are experimental at best, but will try to put them into action in this competition (if I manage to find the time and not get pulled into some other direction 😉) and will let you know how it goes</p>\n<h3>Other resources you might find useful:</h3>\n<ul>\n<li><a href=\"https://www.kaggle.com/code/radek1/2-methods-how-to-ensemble-predictions\" target=\"_blank\">💡 [2 methods] How-to ensemble predictions 🏅🏅🏅</a></li>\n<li><a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/364991\" target=\"_blank\">local validation tracks public LB perfecty -- here is the setup</a></li>\n<li><a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/368560\" target=\"_blank\">💡 For my friends from Twitter and LinkedIn -- here is how to dive into this competition 🐳</a></li>\n<li><a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/363843\" target=\"_blank\">Full dataset processed to CSV/parquet files with optimized memory footprint</a></li>\n<li><a href=\"https://www.kaggle.com/code/radek1/co-visitation-matrix-simplified-imprvd-logic\" target=\"_blank\">co-visitation matrix - simplified, imprvd logic 🔥</a></li>\n<li><a href=\"https://www.kaggle.com/code/radek1/word2vec-how-to-training-and-submission\" target=\"_blank\">💡 Word2Vec How-to [training and submission]🚀🚀🚀</a></li>\n</ul>",
      "rawMarkdown": "This competition is extremely complex.\n\nWe are presented with this seemingly simple data, but THERE IS SO MUCH you can do with it. Just note the explosion of Kaggle kernels sharing ideas already (co-visitation matrices, matrix factorization, reranking, word2vec, training sequential DL models...)\n\nIn [part 1 of this two-post series](https://www.kaggle.com/competitions/otto-recommender-system/discussion/367503) I talked about one aspect of combating complexity -- how to navigate the landscape of ideas to identify what to work on and how to structure your solution. The answer has been simple and could be summed up in a single word: @cdeotte 😄\n\nIn this post let's look at the other type of complexity we need to deal with to be successful - the complexity of the code that we will write/have to work with.\n\nThis type of complexity can crush you extremely quickly. In fact, I only feel I stand a chance in this competition to do something interesting because I got crushed by code complexity in ML projects so many times before!\n\nWhat can the pipeline in this competition look like?\n\n```create datasets to train on and for evaluation -> train MF/word2vec -> create a covisitation matrix (separate for train with validation and for the full train set for submission) -> create features & diagnostic code (measure hit rate, recall) for ranking models -> train a ranking model -> create submission based on this output```\n\nAnd that is a simplification of what the final pipeline will look like! There are so many things you can do along the way (for instance, what could be the equivalent of training with CV in this competition? what interesting things can you do to the data? what about training ranking models with different seeds? what about training ranking models on subsets of columns? how about doing feature selection? what about combining outputs from multiple ranking models?)\n\nThis is a humongous undertaking... one that the field of Machine Learning doesn't have a solution to!\n\nWith infinite resources, you could use something like [Metaflow](https://twitter.com/radekosmulski/status/1556475294084526081?s=20&t=E88ywgWI8RZMQIhOh4_0eQ). But a) most Kagglers don't have the money to spend thousands of dollars on infra b) most Kagglers don't know where to begin setting up stuff like this.\n\nSo what can we do? We need to solve this problem with the right tool for the job given our circumstances.\n\nAnd the answer, at its core, is just a single word. Refactoring.\n\nYOU HAVE TO TIRELESSLY REFACTOR. It is not enough that your code runs and seems to output something useful.\n\nYou need to make it readable. You (ideally) should make it composable.\n\nBut let's stop with these big words. People create a lot of mystique about relatively straightforward things. Let's not fall into that trap here.\n\nThe act of refactoring is very simple (though [some absolutely kickass books](https://www.amazon.com.au/Refactoring-Martin-Fowler/dp/0134757599/ref=asc_df_0134757599/?tag=googleshopdsk-22&linkCode=df0&hvadid=341743255824&hvpos=&hvnetw=g&hvrand=1048400666557196734&hvpone=&hvptwo=&hvqmt=&hvdev=c&hvdvcmdl=&hvlocint=&hvlocphy=9069055&hvtargid=pla-464425925893&psc=1) have been written on it)\n\nA couple of very powerful components of refactoring techniques are:\n* naming\n* create functions to communicate intention (related to 👆)\n\nWhat is the idea behind naming? Use variable names that communicate what the variable holds. Trade being succinct for readability. Always.\n\nHere are a couple of variable names from my code:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F83267%2F6683c4e7ae54f9a4b2ba79c6c9f10abe%2Fvariable_names.png?generation=1669071906008304&alt=media)\n\nI really couldn't care less how long my variable names are AS LONG AS I UNDERSTAND what they hold.\n\nCreating Kaggle kernels is a great exercise in writing readable code. In doing so you learn how to express your ideas so that people with varying levels of experience will be able to follow along.\n\nAnd as for \"creating functions to communicate intent\", this is taking the idea of naming to another level.\n\nThis is a function from my code that I am using so far in a single place. Yes, you want to create functions so that you do not repeat yourself (DRY). But even if I had this bit of code in a single place in my notebook, it might still be valuable to wrap it up in a function.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F83267%2F55ef1fe5123c4dd0bede945f4a8fed60%2Ffunctions_to_communicate_intent.png?generation=1669072142790250&alt=media)\n\nHow much easier it is to understand what is happening in your code when you encounter `cast_to_df_dtypes` vs if you were to encounter just the blob of code that is in the body of the method?\n\nCode is meant to be read by humans, the fact that it runs on a computer is secondary. Code is meant to express (and communicate!) ideas to other human beings. Be that someone reading your code who might have just picked up Python a month ago, or you from the future.\n\nI am now experimenting with [execnb by fast.ai](https://twitter.com/radekosmulski/status/1594556992424931328?s=20&t=-7XnTDzGouCnQrvOYKnJJA) and am thinking this might be another way to reduce the complexity of my code. We'll see.\n\nBut it all starts with refactoring 🙂\n\nAnd the wildest story is I have been looking for resources on my Twitter on refactoring to share with you and it turns out I created [a repository on refactoring](https://github.com/radekosmulski/refactoring) three years ago and completely forgot about 😳 But some really good stuff there that I accumulated over months if not years.\n\nAnyhow, hope this can be of help 🙂 All these ideas are experimental at best, but will try to put them into action in this competition (if I manage to find the time and not get pulled into some other direction 😉) and will let you know how it goes\n\n### Other resources you might find useful:\n\n* [💡 [2 methods] How-to ensemble predictions 🏅🏅🏅](https://www.kaggle.com/code/radek1/2-methods-how-to-ensemble-predictions)\n* [local validation tracks public LB perfecty -- here is the setup](https://www.kaggle.com/competitions/otto-recommender-system/discussion/364991)\n* [💡 For my friends from Twitter and LinkedIn -- here is how to dive into this competition 🐳](https://www.kaggle.com/competitions/otto-recommender-system/discussion/368560)\n* [Full dataset processed to CSV/parquet files with optimized memory footprint](https://www.kaggle.com/competitions/otto-recommender-system/discussion/363843)\n* [co-visitation matrix - simplified, imprvd logic 🔥](https://www.kaggle.com/code/radek1/co-visitation-matrix-simplified-imprvd-logic)\n* [💡 Word2Vec How-to [training and submission]🚀🚀🚀](https://www.kaggle.com/code/radek1/word2vec-how-to-training-and-submission)",
      "votes": null
    },
    {
      "id": "2039829",
      "postDate": "11/22/2022 14:52:59",
      "content": "<p>Thanks for sharing! 👏 Such posts and experience sharing definitely have a place in Kaggle.</p>\n<p>Good suggestions:</p>\n<ul>\n<li>to keep naming meaningful and </li>\n<li>functions as a means for explaining code blocks</li>\n</ul>\n<p>Do share your experience with execnb when you have tried it out.</p>\n<blockquote>\n  <p>\"Kaggle is a place to share and learn\"</p>\n</blockquote>",
      "rawMarkdown": "Thanks for sharing! 👏 Such posts and experience sharing definitely have a place in Kaggle.\n\nGood suggestions:\n- to keep naming meaningful and \n- functions as a means for explaining code blocks\n\nDo share your experience with execnb when you have tried it out.\n> \"Kaggle is a place to share and learn\"",
      "votes": null
    },
    {
      "id": "2040218",
      "postDate": "11/22/2022 20:34:46",
      "content": "<p>The easiest thing to do is become crazy and get used to it 😂<br>\nAlso, your casting function for large-scale data would consume a lot of computational power!</p>\n<p>The biggest issue with all kagglers including myself is a machine-learning strategy! <br>\nWe trial and test algorithms without exploring the business context and researching the right al algorithms to use to save time and computing power!</p>",
      "rawMarkdown": "The easiest thing to do is become crazy and get used to it 😂\nAlso, your casting function for large-scale data would consume a lot of computational power!\n\nThe biggest issue with all kagglers including myself is a machine-learning strategy! \nWe trial and test algorithms without exploring the business context and researching the right al algorithms to use to save time and computing power!",
      "votes": null
    },
    {
      "id": "2041384",
      "postDate": "11/23/2022 22:17:57",
      "content": "<p>Thanks for your efforts in this competition. Many thanks for the experience you're sharing with us!<br>\nThis competition is really strict for optimization and computational power usage, love that 😃</p>",
      "rawMarkdown": "Thanks for your efforts in this competition. Many thanks for the experience you're sharing with us!\nThis competition is really strict for optimization and computational power usage, love that 😃",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2039829,
      "author_name": "noobmldude",
      "author_url": "",
      "post_date": "11/22/2022 14:52:59",
      "content": "<p>Thanks for sharing! 👏 Such posts and experience sharing definitely have a place in Kaggle.</p>\n<p>Good suggestions:</p>\n<ul>\n<li>to keep naming meaningful and </li>\n<li>functions as a means for explaining code blocks</li>\n</ul>\n<p>Do share your experience with execnb when you have tried it out.</p>\n<blockquote>\n  <p>\"Kaggle is a place to share and learn\"</p>\n</blockquote>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2040218,
      "author_name": "muhammadammarjamshed",
      "author_url": "",
      "post_date": "11/22/2022 20:34:46",
      "content": "<p>The easiest thing to do is become crazy and get used to it 😂<br>\nAlso, your casting function for large-scale data would consume a lot of computational power!</p>\n<p>The biggest issue with all kagglers including myself is a machine-learning strategy! <br>\nWe trial and test algorithms without exploring the business context and researching the right al algorithms to use to save time and computing power!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2041384,
      "author_name": "r1chardson",
      "author_url": "",
      "post_date": "11/23/2022 22:17:57",
      "content": "<p>Thanks for your efforts in this competition. Many thanks for the experience you're sharing with us!<br>\nThis competition is really strict for optimization and computational power usage, love that 😃</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2039145": "This competition is extremely complex.\n\nWe are presented with this seemingly simple data, but THERE IS SO MUCH you can do with it. Just note the explosion of Kaggle kernels sharing ideas already (co-visitation matrices, matrix factorization, reranking, word2vec, training sequential DL models...)\n\nIn [part 1 of this two-post series](https://www.kaggle.com/competitions/otto-recommender-system/discussion/367503) I talked about one aspect of combating complexity -- how to navigate the landscape of ideas to identify what to work on and how to structure your solution. The answer has been simple and could be summed up in a single word: @cdeotte 😄\n\nIn this post let's look at the other type of complexity we need to deal with to be successful - the complexity of the code that we will write/have to work with.\n\nThis type of complexity can crush you extremely quickly. In fact, I only feel I stand a chance in this competition to do something interesting because I got crushed by code complexity in ML projects so many times before!\n\nWhat can the pipeline in this competition look like?\n\n```create datasets to train on and for evaluation -> train MF/word2vec -> create a covisitation matrix (separate for train with validation and for the full train set for submission) -> create features & diagnostic code (measure hit rate, recall) for ranking models -> train a ranking model -> create submission based on this output```\n\nAnd that is a simplification of what the final pipeline will look like! There are so many things you can do along the way (for instance, what could be the equivalent of training with CV in this competition? what interesting things can you do to the data? what about training ranking models with different seeds? what about training ranking models on subsets of columns? how about doing feature selection? what about combining outputs from multiple ranking models?)\n\nThis is a humongous undertaking... one that the field of Machine Learning doesn't have a solution to!\n\nWith infinite resources, you could use something like [Metaflow](https://twitter.com/radekosmulski/status/1556475294084526081?s=20&t=E88ywgWI8RZMQIhOh4_0eQ). But a) most Kagglers don't have the money to spend thousands of dollars on infra b) most Kagglers don't know where to begin setting up stuff like this.\n\nSo what can we do? We need to solve this problem with the right tool for the job given our circumstances.\n\nAnd the answer, at its core, is just a single word. Refactoring.\n\nYOU HAVE TO TIRELESSLY REFACTOR. It is not enough that your code runs and seems to output something useful.\n\nYou need to make it readable. You (ideally) should make it composable.\n\nBut let's stop with these big words. People create a lot of mystique about relatively straightforward things. Let's not fall into that trap here.\n\nThe act of refactoring is very simple (though [some absolutely kickass books](https://www.amazon.com.au/Refactoring-Martin-Fowler/dp/0134757599/ref=asc_df_0134757599/?tag=googleshopdsk-22&linkCode=df0&hvadid=341743255824&hvpos=&hvnetw=g&hvrand=1048400666557196734&hvpone=&hvptwo=&hvqmt=&hvdev=c&hvdvcmdl=&hvlocint=&hvlocphy=9069055&hvtargid=pla-464425925893&psc=1) have been written on it)\n\nA couple of very powerful components of refactoring techniques are:\n* naming\n* create functions to communicate intention (related to 👆)\n\nWhat is the idea behind naming? Use variable names that communicate what the variable holds. Trade being succinct for readability. Always.\n\nHere are a couple of variable names from my code:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F83267%2F6683c4e7ae54f9a4b2ba79c6c9f10abe%2Fvariable_names.png?generation=1669071906008304&alt=media)\n\nI really couldn't care less how long my variable names are AS LONG AS I UNDERSTAND what they hold.\n\nCreating Kaggle kernels is a great exercise in writing readable code. In doing so you learn how to express your ideas so that people with varying levels of experience will be able to follow along.\n\nAnd as for \"creating functions to communicate intent\", this is taking the idea of naming to another level.\n\nThis is a function from my code that I am using so far in a single place. Yes, you want to create functions so that you do not repeat yourself (DRY). But even if I had this bit of code in a single place in my notebook, it might still be valuable to wrap it up in a function.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F83267%2F55ef1fe5123c4dd0bede945f4a8fed60%2Ffunctions_to_communicate_intent.png?generation=1669072142790250&alt=media)\n\nHow much easier it is to understand what is happening in your code when you encounter `cast_to_df_dtypes` vs if you were to encounter just the blob of code that is in the body of the method?\n\nCode is meant to be read by humans, the fact that it runs on a computer is secondary. Code is meant to express (and communicate!) ideas to other human beings. Be that someone reading your code who might have just picked up Python a month ago, or you from the future.\n\nI am now experimenting with [execnb by fast.ai](https://twitter.com/radekosmulski/status/1594556992424931328?s=20&t=-7XnTDzGouCnQrvOYKnJJA) and am thinking this might be another way to reduce the complexity of my code. We'll see.\n\nBut it all starts with refactoring 🙂\n\nAnd the wildest story is I have been looking for resources on my Twitter on refactoring to share with you and it turns out I created [a repository on refactoring](https://github.com/radekosmulski/refactoring) three years ago and completely forgot about 😳 But some really good stuff there that I accumulated over months if not years.\n\nAnyhow, hope this can be of help 🙂 All these ideas are experimental at best, but will try to put them into action in this competition (if I manage to find the time and not get pulled into some other direction 😉) and will let you know how it goes\n\n### Other resources you might find useful:\n\n* [💡 [2 methods] How-to ensemble predictions 🏅🏅🏅](https://www.kaggle.com/code/radek1/2-methods-how-to-ensemble-predictions)\n* [local validation tracks public LB perfecty -- here is the setup](https://www.kaggle.com/competitions/otto-recommender-system/discussion/364991)\n* [💡 For my friends from Twitter and LinkedIn -- here is how to dive into this competition 🐳](https://www.kaggle.com/competitions/otto-recommender-system/discussion/368560)\n* [Full dataset processed to CSV/parquet files with optimized memory footprint](https://www.kaggle.com/competitions/otto-recommender-system/discussion/363843)\n* [co-visitation matrix - simplified, imprvd logic 🔥](https://www.kaggle.com/code/radek1/co-visitation-matrix-simplified-imprvd-logic)\n* [💡 Word2Vec How-to [training and submission]🚀🚀🚀](https://www.kaggle.com/code/radek1/word2vec-how-to-training-and-submission)",
    "2039829": "Thanks for sharing! 👏 Such posts and experience sharing definitely have a place in Kaggle.\n\nGood suggestions:\n- to keep naming meaningful and \n- functions as a means for explaining code blocks\n\nDo share your experience with execnb when you have tried it out.\n> \"Kaggle is a place to share and learn\"",
    "2040218": "The easiest thing to do is become crazy and get used to it 😂\nAlso, your casting function for large-scale data would consume a lot of computational power!\n\nThe biggest issue with all kagglers including myself is a machine-learning strategy! \nWe trial and test algorithms without exploring the business context and researching the right al algorithms to use to save time and computing power!",
    "2041384": "Thanks for your efforts in this competition. Many thanks for the experience you're sharing with us!\nThis competition is really strict for optimization and computational power usage, love that 😃"
  },
  "source": "meta"
}