{
  "id": 367503,
  "title": "How to thrive in this competition without going crazy -- 1 out of 2 important truths ❤️‍🔥",
  "url": "/competitions/otto-recommender-system/discussion/367503",
  "author_name": "Radek Osmulski",
  "post_date": "2022-11-21T00:27:27.111000",
  "votes": 61,
  "comment_count": 7,
  "views": 0,
  "content": "<p>If you are not cautious, the complexity of this competition will crush you. On the surface, we are presented with just a sequence of actions but once you start going deeper, this competition becomes EXTREMELY complex very quickly.</p>\n<p>What to work on to improve your LB standing? What should be the shape and structure of the winning model? </p>\n<h1>The first essential truth you need to follow to thrive in this competition:</h1>\n<p>Make good use of the fabulous resource that this forum is. In particular,  <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> is on a mission to guide us to the light, and reading his comments is as high an ROI activity as it gets!</p>\n<p>Here are a couple of things Chris said that are worth their weight in gold 🥇:</p>\n<h3>How to structure the solution to this competition:</h3>\n<blockquote>\n  <p>Great post. This is a great way to build a model in this competition!</p>\n</blockquote>\n<p><a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/364721\" target=\"_blank\">Source</a>. You probably should opt to construct a reranker how I show <a href=\"https://www.kaggle.com/code/radek1/polars-proof-of-concept-lgbm-ranker\" target=\"_blank\">here</a>.</p>\n<p>While my code is a good start, there are a lot of conceptual challenges that await you along the way!</p>\n<p>Again, Chris answers some of the most pressing questions 😄</p>\n<h3>Can a reranker beat the manual heuristic (reranking covisitation matrices using handcrafted rules)</h3>\n<blockquote>\n  <p>If the features are designed correctly, the reranker should always beat heuristics.</p>\n</blockquote>\n<p>This is from a discussion <a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/366474#2032493\" target=\"_blank\">here</a>.</p>\n<p>You might be thinking -- \"my reranker HAS some inherent limitations, even as I provide it the same data that the heuristics use, it still doesn't perform better!\". But that would be WRONG!</p>\n<p>I have been in that spot and have nearly given up completely on the reranker approach. <strong>Essentially, I had no clue how to dig myself out of the complexity of creating a good reranker</strong>. But that is where Chris's comment came to the rescue!</p>\n<p>Once I had my north star that a reranker should always beat a manual heuristic, I knew where to focus to fix the problem! I had to fix the data I was feeding my reranker and I were off to the races!</p>\n<h3>How to improve your results</h3>\n<p>Again, Chris drops the absolutely gold in <a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/365369#2036565\" target=\"_blank\">his comment here </a>.</p>\n<p>Do you need to invest a bunch of time into the covisitation matrices? Absolutely no! But you absolutely can adopt that reasoning in thinking what data to feed to your reranker! (though probably spending a bit more time on improving those co-visitation matrices can go a long way, like Chris suggests 🙂)</p>\n<h3>How do I know what features to create for my reranker?</h3>\n<p>If only a person with 19 gold medals on Kaggle explained this to me, that would be great!</p>\n<p>Wait a second! That is exactly what Chris does in his comment <a href=\"https://www.kaggle.com/code/cdeotte/candidate-rerank-model-lb-0-575/comments#2030893\" target=\"_blank\">here</a>. 🔥</p>\n<h1>Summary</h1>\n<p>The comments from Chris are extremely valuable. You will not learn how to think about solving a ML problem by reading how this or that algorithm works. That will only be a part of the solution. This tacit knowledge that Chris shares is super valuable.</p>\n<p>Thank you, Chris! 😄</p>\n<p>And that is truth number 1 (out of 2) that I am using to thrive in this competition. Will post the truth number 2 in a day or so!</p>\n<p>**Would appreciate your upvote on this post if you found it useful 🙏 If you are not cautious, the complexity of this competition will crush you. On the surface, we are presented with just a sequence of actions but once you start going deeper, this competition becomes EXTREMELY complex very quickly.</p>\n<p>What to work on to improve your LB standing? What should be the shape and structure of the winning model? </p>\n<h1>The first essential truth you need to follow to thrive in this competition:</h1>\n<p>Make good use of the fabulous resource that this forum is. In particular,  <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> is on a mission to guide us to the light, and reading his comments is as high an ROI activity as it gets!</p>\n<p>Here are a couple of things Chris said that are worth their weight in gold 🥇:</p>\n<h3>How to structure the solution to this competition:</h3>\n<blockquote>\n  <p>Great post. This is a great way to build a model in this competition!</p>\n</blockquote>\n<p><a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/364721\" target=\"_blank\">Source</a>. You probably should opt to construct a reranker how I show <a href=\"https://www.kaggle.com/code/radek1/polars-proof-of-concept-lgbm-ranker\" target=\"_blank\">here</a>.</p>\n<p>While my code is a good start, there are a lot of conceptual challenges that await you along the way!</p>\n<p>Again, Chris answers some of the most pressing questions 😄</p>\n<h3>Can a reranker beat the manual heuristic (reranking covisitation matrices using handcrafted rules)</h3>\n<blockquote>\n  <p>If the features are designed correctly, the reranker should always beat heuristics.</p>\n</blockquote>\n<p>This is from a discussion <a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/366474#2032493\" target=\"_blank\">here</a>.</p>\n<p>You might be thinking -- \"my reranker HAS some inherent limitations, even as I provide it the same data that the heuristics use, it still doesn't perform better!\". But that would be WRONG!</p>\n<p>I have been in that spot and have nearly given up completely on the reranker approach. <strong>Essentially, I had no clue how to dig myself out of the complexity of creating a good reranker</strong>. But that is where Chris's comment came to the rescue!</p>\n<p>Once I had my north star that a reranker should always beat a manual heuristic, I knew where to focus to fix the problem! I had to fix the data I was feeding my reranker and I were off to the races!</p>\n<h3>How to improve your results</h3>\n<p>Again, Chris drops the absolutely gold in <a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/365369#2036565\" target=\"_blank\">his comment here </a>.</p>\n<p>Do you need to invest a bunch of time into the covisitation matrices? Absolutely no! But you absolutely can adopt that reasoning in thinking what data to feed to your reranker! (though probably spending a bit more time on improving those co-visitation matrices can go a long way, like Chris suggests 🙂)</p>\n<h3>How do I know what features to create for my reranker?</h3>\n<p>If only a person with 19 gold medals on Kaggle explained this to me, that would be great!</p>\n<p>Wait a second! That is exactly what Chris does in his comment <a href=\"https://www.kaggle.com/code/cdeotte/candidate-rerank-model-lb-0-575/comments#2030893\" target=\"_blank\">here</a>. 🔥</p>\n<h1>Summary</h1>\n<p>The comments from Chris are extremely valuable. You will not learn how to think about solving a ML problem by reading how this or that algorithm works. That will only be a part of the solution. This tacit knowledge that Chris shares is super valuable.</p>\n<p>Thank you, Chris! 😄</p>\n<p>Read the 2nd part here: <a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/367754\" target=\"_blank\">How to thrive in this competition without going crazy (part 2 of 2) ❤️‍🔥</a></p>\n<p><strong>A couple of related resources that you may find useful:</strong></p>\n<p><strong>A couple of related resources you might find useful:</strong></p>\n<ul>\n<li><a href=\"https://www.kaggle.com/code/radek1/2-methods-how-to-ensemble-predictions\" target=\"_blank\">💡 [2 methods] How-to ensemble predictions 🏅🏅🏅</a></li>\n<li><a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/364991\" target=\"_blank\">local validation tracks public LB perfecty -- here is the setup</a></li>\n<li><a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/368560\" target=\"_blank\">💡 For my friends from Twitter and LinkedIn -- here is how to dive into this competition 🐳</a></li>\n<li><a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/363843\" target=\"_blank\">Full dataset processed to CSV/parquet files with optimized memory footprint</a></li>\n<li><a href=\"https://www.kaggle.com/code/radek1/co-visitation-matrix-simplified-imprvd-logic\" target=\"_blank\">co-visitation matrix - simplified, imprvd logic 🔥</a></li>\n<li><a href=\"https://www.kaggle.com/code/radek1/word2vec-how-to-training-and-submission\" target=\"_blank\">💡 Word2Vec How-to [training and submission]🚀🚀🚀</a></li>\n</ul>",
  "messages": [
    {
      "id": 2037800,
      "postDate": "2022-11-21T00:27:27.110Z",
      "content": "<p>If you are not cautious, the complexity of this competition will crush you. On the surface, we are presented with just a sequence of actions but once you start going deeper, this competition becomes EXTREMELY complex very quickly.</p>\n<p>What to work on to improve your LB standing? What should be the shape and structure of the winning model? </p>\n<h1>The first essential truth you need to follow to thrive in this competition:</h1>\n<p>Make good use of the fabulous resource that this forum is. In particular,  <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> is on a mission to guide us to the light, and reading his comments is as high an ROI activity as it gets!</p>\n<p>Here are a couple of things Chris said that are worth their weight in gold 🥇:</p>\n<h3>How to structure the solution to this competition:</h3>\n<blockquote>\n  <p>Great post. This is a great way to build a model in this competition!</p>\n</blockquote>\n<p><a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/364721\" target=\"_blank\">Source</a>. You probably should opt to construct a reranker how I show <a href=\"https://www.kaggle.com/code/radek1/polars-proof-of-concept-lgbm-ranker\" target=\"_blank\">here</a>.</p>\n<p>While my code is a good start, there are a lot of conceptual challenges that await you along the way!</p>\n<p>Again, Chris answers some of the most pressing questions 😄</p>\n<h3>Can a reranker beat the manual heuristic (reranking covisitation matrices using handcrafted rules)</h3>\n<blockquote>\n  <p>If the features are designed correctly, the reranker should always beat heuristics.</p>\n</blockquote>\n<p>This is from a discussion <a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/366474#2032493\" target=\"_blank\">here</a>.</p>\n<p>You might be thinking -- \"my reranker HAS some inherent limitations, even as I provide it the same data that the heuristics use, it still doesn't perform better!\". But that would be WRONG!</p>\n<p>I have been in that spot and have nearly given up completely on the reranker approach. <strong>Essentially, I had no clue how to dig myself out of the complexity of creating a good reranker</strong>. But that is where Chris's comment came to the rescue!</p>\n<p>Once I had my north star that a reranker should always beat a manual heuristic, I knew where to focus to fix the problem! I had to fix the data I was feeding my reranker and I were off to the races!</p>\n<h3>How to improve your results</h3>\n<p>Again, Chris drops the absolutely gold in <a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/365369#2036565\" target=\"_blank\">his comment here </a>.</p>\n<p>Do you need to invest a bunch of time into the covisitation matrices? Absolutely no! But you absolutely can adopt that reasoning in thinking what data to feed to your reranker! (though probably spending a bit more time on improving those co-visitation matrices can go a long way, like Chris suggests 🙂)</p>\n<h3>How do I know what features to create for my reranker?</h3>\n<p>If only a person with 19 gold medals on Kaggle explained this to me, that would be great!</p>\n<p>Wait a second! That is exactly what Chris does in his comment <a href=\"https://www.kaggle.com/code/cdeotte/candidate-rerank-model-lb-0-575/comments#2030893\" target=\"_blank\">here</a>. 🔥</p>\n<h1>Summary</h1>\n<p>The comments from Chris are extremely valuable. You will not learn how to think about solving a ML problem by reading how this or that algorithm works. That will only be a part of the solution. This tacit knowledge that Chris shares is super valuable.</p>\n<p>Thank you, Chris! 😄</p>\n<p>And that is truth number 1 (out of 2) that I am using to thrive in this competition. Will post the truth number 2 in a day or so!</p>\n<p>**Would appreciate your upvote on this post if you found it useful 🙏 If you are not cautious, the complexity of this competition will crush you. On the surface, we are presented with just a sequence of actions but once you start going deeper, this competition becomes EXTREMELY complex very quickly.</p>\n<p>What to work on to improve your LB standing? What should be the shape and structure of the winning model? </p>\n<h1>The first essential truth you need to follow to thrive in this competition:</h1>\n<p>Make good use of the fabulous resource that this forum is. In particular,  <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> is on a mission to guide us to the light, and reading his comments is as high an ROI activity as it gets!</p>\n<p>Here are a couple of things Chris said that are worth their weight in gold 🥇:</p>\n<h3>How to structure the solution to this competition:</h3>\n<blockquote>\n  <p>Great post. This is a great way to build a model in this competition!</p>\n</blockquote>\n<p><a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/364721\" target=\"_blank\">Source</a>. You probably should opt to construct a reranker how I show <a href=\"https://www.kaggle.com/code/radek1/polars-proof-of-concept-lgbm-ranker\" target=\"_blank\">here</a>.</p>\n<p>While my code is a good start, there are a lot of conceptual challenges that await you along the way!</p>\n<p>Again, Chris answers some of the most pressing questions 😄</p>\n<h3>Can a reranker beat the manual heuristic (reranking covisitation matrices using handcrafted rules)</h3>\n<blockquote>\n  <p>If the features are designed correctly, the reranker should always beat heuristics.</p>\n</blockquote>\n<p>This is from a discussion <a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/366474#2032493\" target=\"_blank\">here</a>.</p>\n<p>You might be thinking -- \"my reranker HAS some inherent limitations, even as I provide it the same data that the heuristics use, it still doesn't perform better!\". But that would be WRONG!</p>\n<p>I have been in that spot and have nearly given up completely on the reranker approach. <strong>Essentially, I had no clue how to dig myself out of the complexity of creating a good reranker</strong>. But that is where Chris's comment came to the rescue!</p>\n<p>Once I had my north star that a reranker should always beat a manual heuristic, I knew where to focus to fix the problem! I had to fix the data I was feeding my reranker and I were off to the races!</p>\n<h3>How to improve your results</h3>\n<p>Again, Chris drops the absolutely gold in <a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/365369#2036565\" target=\"_blank\">his comment here </a>.</p>\n<p>Do you need to invest a bunch of time into the covisitation matrices? Absolutely no! But you absolutely can adopt that reasoning in thinking what data to feed to your reranker! (though probably spending a bit more time on improving those co-visitation matrices can go a long way, like Chris suggests 🙂)</p>\n<h3>How do I know what features to create for my reranker?</h3>\n<p>If only a person with 19 gold medals on Kaggle explained this to me, that would be great!</p>\n<p>Wait a second! That is exactly what Chris does in his comment <a href=\"https://www.kaggle.com/code/cdeotte/candidate-rerank-model-lb-0-575/comments#2030893\" target=\"_blank\">here</a>. 🔥</p>\n<h1>Summary</h1>\n<p>The comments from Chris are extremely valuable. You will not learn how to think about solving a ML problem by reading how this or that algorithm works. That will only be a part of the solution. This tacit knowledge that Chris shares is super valuable.</p>\n<p>Thank you, Chris! 😄</p>\n<p>Read the 2nd part here: <a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/367754\" target=\"_blank\">How to thrive in this competition without going crazy (part 2 of 2) ❤️‍🔥</a></p>\n<p><strong>A couple of related resources that you may find useful:</strong></p>\n<p><strong>A couple of related resources you might find useful:</strong></p>\n<ul>\n<li><a href=\"https://www.kaggle.com/code/radek1/2-methods-how-to-ensemble-predictions\" target=\"_blank\">💡 [2 methods] How-to ensemble predictions 🏅🏅🏅</a></li>\n<li><a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/364991\" target=\"_blank\">local validation tracks public LB perfecty -- here is the setup</a></li>\n<li><a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/368560\" target=\"_blank\">💡 For my friends from Twitter and LinkedIn -- here is how to dive into this competition 🐳</a></li>\n<li><a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/363843\" target=\"_blank\">Full dataset processed to CSV/parquet files with optimized memory footprint</a></li>\n<li><a href=\"https://www.kaggle.com/code/radek1/co-visitation-matrix-simplified-imprvd-logic\" target=\"_blank\">co-visitation matrix - simplified, imprvd logic 🔥</a></li>\n<li><a href=\"https://www.kaggle.com/code/radek1/word2vec-how-to-training-and-submission\" target=\"_blank\">💡 Word2Vec How-to [training and submission]🚀🚀🚀</a></li>\n</ul>",
      "rawMarkdown": "If you are not cautious, the complexity of this competition will crush you. On the surface, we are presented with just a sequence of actions but once you start going deeper, this competition becomes EXTREMELY complex very quickly.\n\nWhat to work on to improve your LB standing? What should be the shape and structure of the winning model? \n\n# The first essential truth you need to follow to thrive in this competition:\n\nMake good use of the fabulous resource that this forum is. In particular,  @cdeotte is on a mission to guide us to the light, and reading his comments is as high an ROI activity as it gets!\n\nHere are a couple of things Chris said that are worth their weight in gold 🥇:\n\n### How to structure the solution to this competition:\n\n> Great post. This is a great way to build a model in this competition!\n\n[Source](https://www.kaggle.com/competitions/otto-recommender-system/discussion/364721). You probably should opt to construct a reranker how I show [here](https://www.kaggle.com/code/radek1/polars-proof-of-concept-lgbm-ranker).\n\nWhile my code is a good start, there are a lot of conceptual challenges that await you along the way!\n\nAgain, Chris answers some of the most pressing questions 😄\n\n### Can a reranker beat the manual heuristic (reranking covisitation matrices using handcrafted rules)\n\n> If the features are designed correctly, the reranker should always beat heuristics.\n\nThis is from a discussion [here](https://www.kaggle.com/competitions/otto-recommender-system/discussion/366474#2032493).\n\nYou might be thinking -- \"my reranker HAS some inherent limitations, even as I provide it the same data that the heuristics use, it still doesn't perform better!\". But that would be WRONG!\n\nI have been in that spot and have nearly given up completely on the reranker approach. **Essentially, I had no clue how to dig myself out of the complexity of creating a good reranker**. But that is where Chris's comment came to the rescue!\n\nOnce I had my north star that a reranker should always beat a manual heuristic, I knew where to focus to fix the problem! I had to fix the data I was feeding my reranker and I were off to the races!\n\n### How to improve your results\n\nAgain, Chris drops the absolutely gold in [his comment here ](https://www.kaggle.com/competitions/otto-recommender-system/discussion/365369#2036565).\n\nDo you need to invest a bunch of time into the covisitation matrices? Absolutely no! But you absolutely can adopt that reasoning in thinking what data to feed to your reranker! (though probably spending a bit more time on improving those co-visitation matrices can go a long way, like Chris suggests 🙂)\n\n### How do I know what features to create for my reranker?\n\nIf only a person with 19 gold medals on Kaggle explained this to me, that would be great!\n\nWait a second! That is exactly what Chris does in his comment [here](https://www.kaggle.com/code/cdeotte/candidate-rerank-model-lb-0-575/comments#2030893). 🔥\n\n# Summary\n\nThe comments from Chris are extremely valuable. You will not learn how to think about solving a ML problem by reading how this or that algorithm works. That will only be a part of the solution. This tacit knowledge that Chris shares is super valuable.\n\nThank you, Chris! 😄\n\nAnd that is truth number 1 (out of 2) that I am using to thrive in this competition. Will post the truth number 2 in a day or so!\n\n**Would appreciate your upvote on this post if you found it useful 🙏 If you are not cautious, the complexity of this competition will crush you. On the surface, we are presented with just a sequence of actions but once you start going deeper, this competition becomes EXTREMELY complex very quickly.\n\nWhat to work on to improve your LB standing? What should be the shape and structure of the winning model? \n\n# The first essential truth you need to follow to thrive in this competition:\n\nMake good use of the fabulous resource that this forum is. In particular,  @cdeotte is on a mission to guide us to the light, and reading his comments is as high an ROI activity as it gets!\n\nHere are a couple of things Chris said that are worth their weight in gold 🥇:\n\n### How to structure the solution to this competition:\n\n> Great post. This is a great way to build a model in this competition!\n\n[Source](https://www.kaggle.com/competitions/otto-recommender-system/discussion/364721). You probably should opt to construct a reranker how I show [here](https://www.kaggle.com/code/radek1/polars-proof-of-concept-lgbm-ranker).\n\nWhile my code is a good start, there are a lot of conceptual challenges that await you along the way!\n\nAgain, Chris answers some of the most pressing questions 😄\n\n### Can a reranker beat the manual heuristic (reranking covisitation matrices using handcrafted rules)\n\n> If the features are designed correctly, the reranker should always beat heuristics.\n\nThis is from a discussion [here](https://www.kaggle.com/competitions/otto-recommender-system/discussion/366474#2032493).\n\nYou might be thinking -- \"my reranker HAS some inherent limitations, even as I provide it the same data that the heuristics use, it still doesn't perform better!\". But that would be WRONG!\n\nI have been in that spot and have nearly given up completely on the reranker approach. **Essentially, I had no clue how to dig myself out of the complexity of creating a good reranker**. But that is where Chris's comment came to the rescue!\n\nOnce I had my north star that a reranker should always beat a manual heuristic, I knew where to focus to fix the problem! I had to fix the data I was feeding my reranker and I were off to the races!\n\n### How to improve your results\n\nAgain, Chris drops the absolutely gold in [his comment here ](https://www.kaggle.com/competitions/otto-recommender-system/discussion/365369#2036565).\n\nDo you need to invest a bunch of time into the covisitation matrices? Absolutely no! But you absolutely can adopt that reasoning in thinking what data to feed to your reranker! (though probably spending a bit more time on improving those co-visitation matrices can go a long way, like Chris suggests 🙂)\n\n### How do I know what features to create for my reranker?\n\nIf only a person with 19 gold medals on Kaggle explained this to me, that would be great!\n\nWait a second! That is exactly what Chris does in his comment [here](https://www.kaggle.com/code/cdeotte/candidate-rerank-model-lb-0-575/comments#2030893). 🔥\n\n# Summary\n\nThe comments from Chris are extremely valuable. You will not learn how to think about solving a ML problem by reading how this or that algorithm works. That will only be a part of the solution. This tacit knowledge that Chris shares is super valuable.\n\nThank you, Chris! 😄\n\nRead the 2nd part here: [How to thrive in this competition without going crazy (part 2 of 2) ❤️‍🔥](https://www.kaggle.com/competitions/otto-recommender-system/discussion/367754)\n\n**A couple of related resources that you may find useful:**\n\n**A couple of related resources you might find useful:**\n\n* [💡 [2 methods] How-to ensemble predictions 🏅🏅🏅](https://www.kaggle.com/code/radek1/2-methods-how-to-ensemble-predictions)\n* [local validation tracks public LB perfecty -- here is the setup](https://www.kaggle.com/competitions/otto-recommender-system/discussion/364991)\n* [💡 For my friends from Twitter and LinkedIn -- here is how to dive into this competition 🐳](https://www.kaggle.com/competitions/otto-recommender-system/discussion/368560)\n* [Full dataset processed to CSV/parquet files with optimized memory footprint](https://www.kaggle.com/competitions/otto-recommender-system/discussion/363843)\n* [co-visitation matrix - simplified, imprvd logic 🔥](https://www.kaggle.com/code/radek1/co-visitation-matrix-simplified-imprvd-logic)\n* [💡 Word2Vec How-to [training and submission]🚀🚀🚀](https://www.kaggle.com/code/radek1/word2vec-how-to-training-and-submission)\n\n",
      "votes": 60
    },
    {
      "id": 2037879,
      "postDate": "2022-11-21T02:42:09.717Z",
      "content": "<p>haha, thanks for highlighting my posts. The winning solution from Kaggle's last recommender system competition (i.e. H&amp;M Personalized Fashion Recommendations) was a \"candidate rerank\" model described <a href=\"https://www.kaggle.com/competitions/h-and-m-personalized-fashion-recommendations/discussion/324070\" target=\"_blank\">here</a>. (All top solutions were \"candidate rerank\"). Based on the high cardinality of the items in Otto comp (i.e. 1.8 million unique items!), I also believe the winning solution in Otto comp will be a \"candidate rerank\" model.</p>\n<p>An easy and quick way to explore what \"candidates\" are important and what \"rerank features\" are important is to make heuristic models like my public notebook. It is a fun way to explore the Otto data and brainstorm what information is important. Afterward, you can use the same \"candidates\" but then replace heuristics with GBT rerank model and you will have a very strong solution for Otto comp!</p>",
      "rawMarkdown": "haha, thanks for highlighting my posts. The winning solution from Kaggle's last recommender system competition (i.e. H&M Personalized Fashion Recommendations) was a \"candidate rerank\" model described [here][1]. (All top solutions were \"candidate rerank\"). Based on the high cardinality of the items in Otto comp (i.e. 1.8 million unique items!), I also believe the winning solution in Otto comp will be a \"candidate rerank\" model.\n\nAn easy and quick way to explore what \"candidates\" are important and what \"rerank features\" are important is to make heuristic models like my public notebook. It is a fun way to explore the Otto data and brainstorm what information is important. Afterward, you can use the same \"candidates\" but then replace heuristics with GBT rerank model and you will have a very strong solution for Otto comp!\n\n[1]: https://www.kaggle.com/competitions/h-and-m-personalized-fashion-recommendations/discussion/324070",
      "votes": 12,
      "replies": [
        {
          "id": 2037890,
          "postDate": "2022-11-21T03:13:01.480Z",
          "content": "<p>Thank you very much Chris for everything that you are sharing in this competition and the above comment! 🙂🙏</p>\n<p>It is one thing to sort of have a vague idea of how to go about arriving at a solution and a completely different thing to have someone super experienced guide you along the way via their forum posts. It makes a huge difference, thank you so much again! 🙂</p>\n<p>I am <em>somewhat</em> following the overall trajectory of the H&amp;M solution you linked to in your comment. Including monitoring the <code>HR</code> at various stages as mentioned there 😄</p>\n<p>But I am quite surprised how complex of a software engineering project a solution to a Kaggle competition ends up being 🙂 Will continue to chip away at the solution every now and then as time permits, quite fortunate that I joined this competition so early on 😄</p>\n<p>But can already get a feeling for how complex the pipeline is shaping up to be! Oh man. To not be crushed by the conceptual complexity of the solution and the code required to implement it is really something 🙂 Quite a fun challenge!</p>\n<p>Thanks again for all the help that you are giving us 🙏</p>",
          "rawMarkdown": "Thank you very much Chris for everything that you are sharing in this competition and the above comment! 🙂🙏\n\nIt is one thing to sort of have a vague idea of how to go about arriving at a solution and a completely different thing to have someone super experienced guide you along the way via their forum posts. It makes a huge difference, thank you so much again! 🙂\n\nI am *somewhat* following the overall trajectory of the H&M solution you linked to in your comment. Including monitoring the `HR` at various stages as mentioned there 😄\n\nBut I am quite surprised how complex of a software engineering project a solution to a Kaggle competition ends up being 🙂 Will continue to chip away at the solution every now and then as time permits, quite fortunate that I joined this competition so early on 😄\n\nBut can already get a feeling for how complex the pipeline is shaping up to be! Oh man. To not be crushed by the conceptual complexity of the solution and the code required to implement it is really something 🙂 Quite a fun challenge!\n\nThanks again for all the help that you are giving us 🙏",
          "votes": 3
        },
        {
          "id": 2037894,
          "postDate": "2022-11-21T03:21:46.887Z",
          "content": "<p>Yes, everything gets huge. These large scale recommender system models require train data with millions of users where each user has thousands of candidates. It can be overwhelming. One approach is to create training data chunk by chunk and save to disk. Then load from disk to train your GBT reranker. Also during inference, we can infer chunk by chunk.</p>\n<p>Using heuristics is a fun simpler solution which doesn't require all the train pipeline transformation and train model training. So its a nice way to start. But eventually, top solutions will require dealing with the massive amounts of data and feature generation and model training. </p>",
          "rawMarkdown": "Yes, everything gets huge. These large scale recommender system models require train data with millions of users where each user has thousands of candidates. It can be overwhelming. One approach is to create training data chunk by chunk and save to disk. Then load from disk to train your GBT reranker. Also during inference, we can infer chunk by chunk.\n\nUsing heuristics is a fun simpler solution which doesn't require all the train pipeline transformation and train model training. So its a nice way to start. But eventually, top solutions will require dealing with the massive amounts of data and feature generation and model training. ",
          "votes": 11
        }
      ]
    },
    {
      "id": 2052994,
      "postDate": "2022-12-02T17:55:28.217Z",
      "content": "<p>Thank you very much for posting this. I was crushed by this competition very much. <br>\nNow I see the light is coming for your article. <br>\nThank you for you and Chris <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> </p>",
      "rawMarkdown": "Thank you very much for posting this. I was crushed by this competition very much. \nNow I see the light is coming for your article. \nThank you for you and Chris @cdeotte ",
      "votes": 3,
      "replies": [
        {
          "id": 2053006,
          "postDate": "2022-12-02T18:09:04.650Z",
          "content": "<p>That is wonderful to hear 🙂 Thank you very much for your comment, <a href=\"https://www.kaggle.com/leiwong\" target=\"_blank\">@leiwong</a>! 🙂 </p>",
          "rawMarkdown": "That is wonderful to hear 🙂 Thank you very much for your comment, @leiwong! 🙂 ",
          "votes": 1
        }
      ]
    },
    {
      "id": 2040867,
      "postDate": "2022-11-23T13:16:57.207Z",
      "content": "<p>super thanks for you both <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> <a href=\"https://www.kaggle.com/radek1\" target=\"_blank\">@radek1</a>, this is a gold post containing multiple gold posts and comments!</p>",
      "rawMarkdown": "super thanks for you both @cdeotte @radek1, this is a gold post containing multiple gold posts and comments!",
      "votes": 2
    },
    {
      "id": 2038609,
      "postDate": "2022-11-21T14:11:41.060Z",
      "rawMarkdown": "",
      "votes": 2,
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 2037879,
      "author_name": "Chris Deotte",
      "author_url": "",
      "post_date": "2022-11-21T02:42:09.717000",
      "content": "<p>haha, thanks for highlighting my posts. The winning solution from Kaggle's last recommender system competition (i.e. H&amp;M Personalized Fashion Recommendations) was a \"candidate rerank\" model described <a href=\"https://www.kaggle.com/competitions/h-and-m-personalized-fashion-recommendations/discussion/324070\" target=\"_blank\">here</a>. (All top solutions were \"candidate rerank\"). Based on the high cardinality of the items in Otto comp (i.e. 1.8 million unique items!), I also believe the winning solution in Otto comp will be a \"candidate rerank\" model.</p>\n<p>An easy and quick way to explore what \"candidates\" are important and what \"rerank features\" are important is to make heuristic models like my public notebook. It is a fun way to explore the Otto data and brainstorm what information is important. Afterward, you can use the same \"candidates\" but then replace heuristics with GBT rerank model and you will have a very strong solution for Otto comp!</p>",
      "votes": 12,
      "replies": [
        {
          "id": 2037890,
          "author_name": "Radek Osmulski",
          "author_url": "",
          "post_date": "2022-11-21T03:13:01.480000",
          "content": "<p>Thank you very much Chris for everything that you are sharing in this competition and the above comment! 🙂🙏</p>\n<p>It is one thing to sort of have a vague idea of how to go about arriving at a solution and a completely different thing to have someone super experienced guide you along the way via their forum posts. It makes a huge difference, thank you so much again! 🙂</p>\n<p>I am <em>somewhat</em> following the overall trajectory of the H&amp;M solution you linked to in your comment. Including monitoring the <code>HR</code> at various stages as mentioned there 😄</p>\n<p>But I am quite surprised how complex of a software engineering project a solution to a Kaggle competition ends up being 🙂 Will continue to chip away at the solution every now and then as time permits, quite fortunate that I joined this competition so early on 😄</p>\n<p>But can already get a feeling for how complex the pipeline is shaping up to be! Oh man. To not be crushed by the conceptual complexity of the solution and the code required to implement it is really something 🙂 Quite a fun challenge!</p>\n<p>Thanks again for all the help that you are giving us 🙏</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 2037894,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2022-11-21T03:21:46.887000",
          "content": "<p>Yes, everything gets huge. These large scale recommender system models require train data with millions of users where each user has thousands of candidates. It can be overwhelming. One approach is to create training data chunk by chunk and save to disk. Then load from disk to train your GBT reranker. Also during inference, we can infer chunk by chunk.</p>\n<p>Using heuristics is a fun simpler solution which doesn't require all the train pipeline transformation and train model training. So its a nice way to start. But eventually, top solutions will require dealing with the massive amounts of data and feature generation and model training. </p>",
          "votes": 11,
          "replies": []
        }
      ]
    },
    {
      "id": 2052994,
      "author_name": "Lei Wang",
      "author_url": "",
      "post_date": "2022-12-02T17:55:28.217000",
      "content": "<p>Thank you very much for posting this. I was crushed by this competition very much. <br>\nNow I see the light is coming for your article. <br>\nThank you for you and Chris <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> </p>",
      "votes": 3,
      "replies": [
        {
          "id": 2053006,
          "author_name": "Radek Osmulski",
          "author_url": "",
          "post_date": "2022-12-02T18:09:04.650000",
          "content": "<p>That is wonderful to hear 🙂 Thank you very much for your comment, <a href=\"https://www.kaggle.com/leiwong\" target=\"_blank\">@leiwong</a>! 🙂 </p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 2040867,
      "author_name": "danielliao",
      "author_url": "",
      "post_date": "2022-11-23T13:16:57.207000",
      "content": "<p>super thanks for you both <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> <a href=\"https://www.kaggle.com/radek1\" target=\"_blank\">@radek1</a>, this is a gold post containing multiple gold posts and comments!</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 2038609,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-11-21T14:11:41.060000",
      "content": "",
      "votes": 2,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2037800": "If you are not cautious, the complexity of this competition will crush you. On the surface, we are presented with just a sequence of actions but once you start going deeper, this competition becomes EXTREMELY complex very quickly.\n\nWhat to work on to improve your LB standing? What should be the shape and structure of the winning model? \n\n# The first essential truth you need to follow to thrive in this competition:\n\nMake good use of the fabulous resource that this forum is. In particular,  @cdeotte is on a mission to guide us to the light, and reading his comments is as high an ROI activity as it gets!\n\nHere are a couple of things Chris said that are worth their weight in gold 🥇:\n\n### How to structure the solution to this competition:\n\n> Great post. This is a great way to build a model in this competition!\n\n[Source](https://www.kaggle.com/competitions/otto-recommender-system/discussion/364721). You probably should opt to construct a reranker how I show [here](https://www.kaggle.com/code/radek1/polars-proof-of-concept-lgbm-ranker).\n\nWhile my code is a good start, there are a lot of conceptual challenges that await you along the way!\n\nAgain, Chris answers some of the most pressing questions 😄\n\n### Can a reranker beat the manual heuristic (reranking covisitation matrices using handcrafted rules)\n\n> If the features are designed correctly, the reranker should always beat heuristics.\n\nThis is from a discussion [here](https://www.kaggle.com/competitions/otto-recommender-system/discussion/366474#2032493).\n\nYou might be thinking -- \"my reranker HAS some inherent limitations, even as I provide it the same data that the heuristics use, it still doesn't perform better!\". But that would be WRONG!\n\nI have been in that spot and have nearly given up completely on the reranker approach. **Essentially, I had no clue how to dig myself out of the complexity of creating a good reranker**. But that is where Chris's comment came to the rescue!\n\nOnce I had my north star that a reranker should always beat a manual heuristic, I knew where to focus to fix the problem! I had to fix the data I was feeding my reranker and I were off to the races!\n\n### How to improve your results\n\nAgain, Chris drops the absolutely gold in [his comment here ](https://www.kaggle.com/competitions/otto-recommender-system/discussion/365369#2036565).\n\nDo you need to invest a bunch of time into the covisitation matrices? Absolutely no! But you absolutely can adopt that reasoning in thinking what data to feed to your reranker! (though probably spending a bit more time on improving those co-visitation matrices can go a long way, like Chris suggests 🙂)\n\n### How do I know what features to create for my reranker?\n\nIf only a person with 19 gold medals on Kaggle explained this to me, that would be great!\n\nWait a second! That is exactly what Chris does in his comment [here](https://www.kaggle.com/code/cdeotte/candidate-rerank-model-lb-0-575/comments#2030893). 🔥\n\n# Summary\n\nThe comments from Chris are extremely valuable. You will not learn how to think about solving a ML problem by reading how this or that algorithm works. That will only be a part of the solution. This tacit knowledge that Chris shares is super valuable.\n\nThank you, Chris! 😄\n\nAnd that is truth number 1 (out of 2) that I am using to thrive in this competition. Will post the truth number 2 in a day or so!\n\n**Would appreciate your upvote on this post if you found it useful 🙏 If you are not cautious, the complexity of this competition will crush you. On the surface, we are presented with just a sequence of actions but once you start going deeper, this competition becomes EXTREMELY complex very quickly.\n\nWhat to work on to improve your LB standing? What should be the shape and structure of the winning model? \n\n# The first essential truth you need to follow to thrive in this competition:\n\nMake good use of the fabulous resource that this forum is. In particular,  @cdeotte is on a mission to guide us to the light, and reading his comments is as high an ROI activity as it gets!\n\nHere are a couple of things Chris said that are worth their weight in gold 🥇:\n\n### How to structure the solution to this competition:\n\n> Great post. This is a great way to build a model in this competition!\n\n[Source](https://www.kaggle.com/competitions/otto-recommender-system/discussion/364721). You probably should opt to construct a reranker how I show [here](https://www.kaggle.com/code/radek1/polars-proof-of-concept-lgbm-ranker).\n\nWhile my code is a good start, there are a lot of conceptual challenges that await you along the way!\n\nAgain, Chris answers some of the most pressing questions 😄\n\n### Can a reranker beat the manual heuristic (reranking covisitation matrices using handcrafted rules)\n\n> If the features are designed correctly, the reranker should always beat heuristics.\n\nThis is from a discussion [here](https://www.kaggle.com/competitions/otto-recommender-system/discussion/366474#2032493).\n\nYou might be thinking -- \"my reranker HAS some inherent limitations, even as I provide it the same data that the heuristics use, it still doesn't perform better!\". But that would be WRONG!\n\nI have been in that spot and have nearly given up completely on the reranker approach. **Essentially, I had no clue how to dig myself out of the complexity of creating a good reranker**. But that is where Chris's comment came to the rescue!\n\nOnce I had my north star that a reranker should always beat a manual heuristic, I knew where to focus to fix the problem! I had to fix the data I was feeding my reranker and I were off to the races!\n\n### How to improve your results\n\nAgain, Chris drops the absolutely gold in [his comment here ](https://www.kaggle.com/competitions/otto-recommender-system/discussion/365369#2036565).\n\nDo you need to invest a bunch of time into the covisitation matrices? Absolutely no! But you absolutely can adopt that reasoning in thinking what data to feed to your reranker! (though probably spending a bit more time on improving those co-visitation matrices can go a long way, like Chris suggests 🙂)\n\n### How do I know what features to create for my reranker?\n\nIf only a person with 19 gold medals on Kaggle explained this to me, that would be great!\n\nWait a second! That is exactly what Chris does in his comment [here](https://www.kaggle.com/code/cdeotte/candidate-rerank-model-lb-0-575/comments#2030893). 🔥\n\n# Summary\n\nThe comments from Chris are extremely valuable. You will not learn how to think about solving a ML problem by reading how this or that algorithm works. That will only be a part of the solution. This tacit knowledge that Chris shares is super valuable.\n\nThank you, Chris! 😄\n\nRead the 2nd part here: [How to thrive in this competition without going crazy (part 2 of 2) ❤️‍🔥](https://www.kaggle.com/competitions/otto-recommender-system/discussion/367754)\n\n**A couple of related resources that you may find useful:**\n\n**A couple of related resources you might find useful:**\n\n* [💡 [2 methods] How-to ensemble predictions 🏅🏅🏅](https://www.kaggle.com/code/radek1/2-methods-how-to-ensemble-predictions)\n* [local validation tracks public LB perfecty -- here is the setup](https://www.kaggle.com/competitions/otto-recommender-system/discussion/364991)\n* [💡 For my friends from Twitter and LinkedIn -- here is how to dive into this competition 🐳](https://www.kaggle.com/competitions/otto-recommender-system/discussion/368560)\n* [Full dataset processed to CSV/parquet files with optimized memory footprint](https://www.kaggle.com/competitions/otto-recommender-system/discussion/363843)\n* [co-visitation matrix - simplified, imprvd logic 🔥](https://www.kaggle.com/code/radek1/co-visitation-matrix-simplified-imprvd-logic)\n* [💡 Word2Vec How-to [training and submission]🚀🚀🚀](https://www.kaggle.com/code/radek1/word2vec-how-to-training-and-submission)\n\n",
    "2037879": "haha, thanks for highlighting my posts. The winning solution from Kaggle's last recommender system competition (i.e. H&M Personalized Fashion Recommendations) was a \"candidate rerank\" model described [here][1]. (All top solutions were \"candidate rerank\"). Based on the high cardinality of the items in Otto comp (i.e. 1.8 million unique items!), I also believe the winning solution in Otto comp will be a \"candidate rerank\" model.\n\nAn easy and quick way to explore what \"candidates\" are important and what \"rerank features\" are important is to make heuristic models like my public notebook. It is a fun way to explore the Otto data and brainstorm what information is important. Afterward, you can use the same \"candidates\" but then replace heuristics with GBT rerank model and you will have a very strong solution for Otto comp!\n\n[1]: https://www.kaggle.com/competitions/h-and-m-personalized-fashion-recommendations/discussion/324070",
    "2052994": "Thank you very much for posting this. I was crushed by this competition very much. \nNow I see the light is coming for your article. \nThank you for you and Chris @cdeotte ",
    "2040867": "super thanks for you both @cdeotte @radek1, this is a gold post containing multiple gold posts and comments!",
    "2038609": ""
  }
}