{
  "id": 368685,
  "title": "💡What is a good initial goal in the competition? How to improve beyond it? 📈",
  "url": "/competitions/otto-recommender-system/discussion/368685",
  "author_name": "",
  "post_date": "2022-11-27T05:33:21.531883800Z",
  "votes": 23,
  "comment_count": 2,
  "views": 0,
  "content": "<p>You may have followed some of the posts or notebooks that I shared. A good initial goal is to get everything up and running on your end (be that on Kaggle or on your own rig).</p>\n<p>A great first milestone is getting the ranker to beat or output comparative predictions to the ones you would obtain directly from the covisitation matrix!</p>\n<p>But once you have that in place (and mind you, that is a big goal -- it might take you several days of hacking away on the competition, if not more), what do you do next?</p>\n<p>At that point, you will find yourself in a really fun spot 🙂 And there is one thing you absolutely need to do!</p>\n<h6>WRITE. DIAGNOSTIC. CODE.</h6>\n<p>This is a vital step. You need to understand what your model is doing and how. You need to understand the data.</p>\n<p>I started to look at this a bit today and here is something very interesting that I found.</p>\n<p>If you were to take the 20 suggestions from the <code>clicks covisitation matrix</code>, you would get <code>0.21</code> on clicks, <code>0.10</code> on carts and <code>0.06</code> on orders from the competition metric (this is on a specific validation split I am using, you can read more about creating one for yourself here: <a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/364534\" target=\"_blank\">📅 Dataset for local validation created using organizer's repository (parquet files)</a>.</p>\n<p>This is a super interesting and valuable piece of information! It also speaks to the immense value of the covistation matrices as constructed and improved by <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> here: <a href=\"https://www.kaggle.com/code/cdeotte/candidate-rerank-model-lb-0-575\" target=\"_blank\">Candidate ReRank Model - [LB 0.575]</a>.</p>\n<p>And do you know what is most impressive about the candidates from the <code>clicks covisitation matrix</code>? They don't include the last aid from the input sequence! (which is extremely likely to appear among ground truth labels).</p>\n<p>So as you work on additional ways for candidate generation, an exercise as the above gives you a better understanding of whether you are moving in the right direction, and how your results stack up against a very powerful approach in the form of the covistation matrix.</p>\n<p>How much of the ground truth in this competition comes from session history? What else would a customer be buying? How can we find those aids in the data?</p>\n<p>This is a really fun part and will ultimately decide the standing on the final leaderboard 🙂 As you continue to learn about the data, as you continue to understand the patterns (or expose information to your model that can help it find them) you codify what you learn in code and progressively grow your solution.</p>\n<p>And that is the plan 🙂</p>\n<p>PS. I am really excited about this competition! 🙂 Plan to share more models and techniques you will be able to implement into your solution over the next couple of days! 🙂 Please stay tuned! 📺</p>\n<h3>Other resources you might find useful:</h3>\n<ul>\n<li><a href=\"https://www.kaggle.com/code/radek1/2-methods-how-to-ensemble-predictions\" target=\"_blank\">💡 [2 methods] How-to ensemble predictions 🏅🏅🏅</a></li>\n<li><a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/364991\" target=\"_blank\">local validation tracks public LB perfecty -- here is the setup</a></li>\n<li><a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/368560\" target=\"_blank\">💡 For my friends from Twitter and LinkedIn -- here is how to dive into this competition 🐳</a></li>\n<li><a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/363843\" target=\"_blank\">Full dataset processed to CSV/parquet files with optimized memory footprint</a></li>\n<li><a href=\"https://www.kaggle.com/code/radek1/co-visitation-matrix-simplified-imprvd-logic\" target=\"_blank\">co-visitation matrix - simplified, imprvd logic 🔥</a></li>\n<li><a href=\"https://www.kaggle.com/code/radek1/word2vec-how-to-training-and-submission\" target=\"_blank\">💡 Word2Vec How-to [training and submission]🚀🚀🚀</a></li>\n</ul>",
  "messages": [
    {
      "id": "2045126",
      "postDate": "11/27/2022 05:33:21",
      "content": "<p>You may have followed some of the posts or notebooks that I shared. A good initial goal is to get everything up and running on your end (be that on Kaggle or on your own rig).</p>\n<p>A great first milestone is getting the ranker to beat or output comparative predictions to the ones you would obtain directly from the covisitation matrix!</p>\n<p>But once you have that in place (and mind you, that is a big goal -- it might take you several days of hacking away on the competition, if not more), what do you do next?</p>\n<p>At that point, you will find yourself in a really fun spot 🙂 And there is one thing you absolutely need to do!</p>\n<h6>WRITE. DIAGNOSTIC. CODE.</h6>\n<p>This is a vital step. You need to understand what your model is doing and how. You need to understand the data.</p>\n<p>I started to look at this a bit today and here is something very interesting that I found.</p>\n<p>If you were to take the 20 suggestions from the <code>clicks covisitation matrix</code>, you would get <code>0.21</code> on clicks, <code>0.10</code> on carts and <code>0.06</code> on orders from the competition metric (this is on a specific validation split I am using, you can read more about creating one for yourself here: <a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/364534\" target=\"_blank\">📅 Dataset for local validation created using organizer's repository (parquet files)</a>.</p>\n<p>This is a super interesting and valuable piece of information! It also speaks to the immense value of the covistation matrices as constructed and improved by <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> here: <a href=\"https://www.kaggle.com/code/cdeotte/candidate-rerank-model-lb-0-575\" target=\"_blank\">Candidate ReRank Model - [LB 0.575]</a>.</p>\n<p>And do you know what is most impressive about the candidates from the <code>clicks covisitation matrix</code>? They don't include the last aid from the input sequence! (which is extremely likely to appear among ground truth labels).</p>\n<p>So as you work on additional ways for candidate generation, an exercise as the above gives you a better understanding of whether you are moving in the right direction, and how your results stack up against a very powerful approach in the form of the covistation matrix.</p>\n<p>How much of the ground truth in this competition comes from session history? What else would a customer be buying? How can we find those aids in the data?</p>\n<p>This is a really fun part and will ultimately decide the standing on the final leaderboard 🙂 As you continue to learn about the data, as you continue to understand the patterns (or expose information to your model that can help it find them) you codify what you learn in code and progressively grow your solution.</p>\n<p>And that is the plan 🙂</p>\n<p>PS. I am really excited about this competition! 🙂 Plan to share more models and techniques you will be able to implement into your solution over the next couple of days! 🙂 Please stay tuned! 📺</p>\n<h3>Other resources you might find useful:</h3>\n<ul>\n<li><a href=\"https://www.kaggle.com/code/radek1/2-methods-how-to-ensemble-predictions\" target=\"_blank\">💡 [2 methods] How-to ensemble predictions 🏅🏅🏅</a></li>\n<li><a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/364991\" target=\"_blank\">local validation tracks public LB perfecty -- here is the setup</a></li>\n<li><a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/368560\" target=\"_blank\">💡 For my friends from Twitter and LinkedIn -- here is how to dive into this competition 🐳</a></li>\n<li><a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/363843\" target=\"_blank\">Full dataset processed to CSV/parquet files with optimized memory footprint</a></li>\n<li><a href=\"https://www.kaggle.com/code/radek1/co-visitation-matrix-simplified-imprvd-logic\" target=\"_blank\">co-visitation matrix - simplified, imprvd logic 🔥</a></li>\n<li><a href=\"https://www.kaggle.com/code/radek1/word2vec-how-to-training-and-submission\" target=\"_blank\">💡 Word2Vec How-to [training and submission]🚀🚀🚀</a></li>\n</ul>",
      "rawMarkdown": "You may have followed some of the posts or notebooks that I shared. A good initial goal is to get everything up and running on your end (be that on Kaggle or on your own rig).\n\nA great first milestone is getting the ranker to beat or output comparative predictions to the ones you would obtain directly from the covisitation matrix!\n\nBut once you have that in place (and mind you, that is a big goal -- it might take you several days of hacking away on the competition, if not more), what do you do next?\n\nAt that point, you will find yourself in a really fun spot 🙂 And there is one thing you absolutely need to do!\n\n###### WRITE. DIAGNOSTIC. CODE.\n\nThis is a vital step. You need to understand what your model is doing and how. You need to understand the data.\n\nI started to look at this a bit today and here is something very interesting that I found.\n\nIf you were to take the 20 suggestions from the `clicks covisitation matrix`, you would get `0.21` on clicks, `0.10` on carts and `0.06` on orders from the competition metric (this is on a specific validation split I am using, you can read more about creating one for yourself here: [📅 Dataset for local validation created using organizer's repository (parquet files)](https://www.kaggle.com/competitions/otto-recommender-system/discussion/364534).\n\nThis is a super interesting and valuable piece of information! It also speaks to the immense value of the covistation matrices as constructed and improved by @cdeotte here: [Candidate ReRank Model - [LB 0.575]](https://www.kaggle.com/code/cdeotte/candidate-rerank-model-lb-0-575).\n\nAnd do you know what is most impressive about the candidates from the `clicks covisitation matrix`? They don't include the last aid from the input sequence! (which is extremely likely to appear among ground truth labels).\n\nSo as you work on additional ways for candidate generation, an exercise as the above gives you a better understanding of whether you are moving in the right direction, and how your results stack up against a very powerful approach in the form of the covistation matrix.\n\nHow much of the ground truth in this competition comes from session history? What else would a customer be buying? How can we find those aids in the data?\n\nThis is a really fun part and will ultimately decide the standing on the final leaderboard 🙂 As you continue to learn about the data, as you continue to understand the patterns (or expose information to your model that can help it find them) you codify what you learn in code and progressively grow your solution.\n\nAnd that is the plan 🙂\n\nPS. I am really excited about this competition! 🙂 Plan to share more models and techniques you will be able to implement into your solution over the next couple of days! 🙂 Please stay tuned! 📺\n\n### Other resources you might find useful:\n\n* [💡 [2 methods] How-to ensemble predictions 🏅🏅🏅](https://www.kaggle.com/code/radek1/2-methods-how-to-ensemble-predictions)\n* [local validation tracks public LB perfecty -- here is the setup](https://www.kaggle.com/competitions/otto-recommender-system/discussion/364991)\n* [💡 For my friends from Twitter and LinkedIn -- here is how to dive into this competition 🐳](https://www.kaggle.com/competitions/otto-recommender-system/discussion/368560)\n* [Full dataset processed to CSV/parquet files with optimized memory footprint](https://www.kaggle.com/competitions/otto-recommender-system/discussion/363843)\n* [co-visitation matrix - simplified, imprvd logic 🔥](https://www.kaggle.com/code/radek1/co-visitation-matrix-simplified-imprvd-logic)\n* [💡 Word2Vec How-to [training and submission]🚀🚀🚀](https://www.kaggle.com/code/radek1/word2vec-how-to-training-and-submission)",
      "votes": null
    },
    {
      "id": "2060331",
      "postDate": "12/09/2022 18:41:28",
      "content": "<p>Hi, <a href=\"https://www.kaggle.com/radek1\" target=\"_blank\">@radek1</a>! Thanks for making this competition more understandable and fun. It's an overwhelming one to take as a first one 😅</p>",
      "rawMarkdown": "Hi, @radek1! Thanks for making this competition more understandable and fun. It's an overwhelming one to take as a first one 😅",
      "votes": null
    },
    {
      "id": "2061663",
      "postDate": "12/11/2022 11:24:57",
      "content": "<p>Thank you for your kind words, <a href=\"https://www.kaggle.com/jamnik99\" target=\"_blank\">@jamnik99</a>! 🙌 Extremely glad I could be of help 🙂</p>",
      "rawMarkdown": "Thank you for your kind words, @jamnik99! 🙌 Extremely glad I could be of help 🙂",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2060331,
      "author_name": "jamnik99",
      "author_url": "",
      "post_date": "12/09/2022 18:41:28",
      "content": "<p>Hi, <a href=\"https://www.kaggle.com/radek1\" target=\"_blank\">@radek1</a>! Thanks for making this competition more understandable and fun. It's an overwhelming one to take as a first one 😅</p>",
      "votes": null,
      "replies": [
        {
          "id": 2061663,
          "author_name": "radek1",
          "author_url": "",
          "post_date": "12/11/2022 11:24:57",
          "content": "<p>Thank you for your kind words, <a href=\"https://www.kaggle.com/jamnik99\" target=\"_blank\">@jamnik99</a>! 🙌 Extremely glad I could be of help 🙂</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2045126": "You may have followed some of the posts or notebooks that I shared. A good initial goal is to get everything up and running on your end (be that on Kaggle or on your own rig).\n\nA great first milestone is getting the ranker to beat or output comparative predictions to the ones you would obtain directly from the covisitation matrix!\n\nBut once you have that in place (and mind you, that is a big goal -- it might take you several days of hacking away on the competition, if not more), what do you do next?\n\nAt that point, you will find yourself in a really fun spot 🙂 And there is one thing you absolutely need to do!\n\n###### WRITE. DIAGNOSTIC. CODE.\n\nThis is a vital step. You need to understand what your model is doing and how. You need to understand the data.\n\nI started to look at this a bit today and here is something very interesting that I found.\n\nIf you were to take the 20 suggestions from the `clicks covisitation matrix`, you would get `0.21` on clicks, `0.10` on carts and `0.06` on orders from the competition metric (this is on a specific validation split I am using, you can read more about creating one for yourself here: [📅 Dataset for local validation created using organizer's repository (parquet files)](https://www.kaggle.com/competitions/otto-recommender-system/discussion/364534).\n\nThis is a super interesting and valuable piece of information! It also speaks to the immense value of the covistation matrices as constructed and improved by @cdeotte here: [Candidate ReRank Model - [LB 0.575]](https://www.kaggle.com/code/cdeotte/candidate-rerank-model-lb-0-575).\n\nAnd do you know what is most impressive about the candidates from the `clicks covisitation matrix`? They don't include the last aid from the input sequence! (which is extremely likely to appear among ground truth labels).\n\nSo as you work on additional ways for candidate generation, an exercise as the above gives you a better understanding of whether you are moving in the right direction, and how your results stack up against a very powerful approach in the form of the covistation matrix.\n\nHow much of the ground truth in this competition comes from session history? What else would a customer be buying? How can we find those aids in the data?\n\nThis is a really fun part and will ultimately decide the standing on the final leaderboard 🙂 As you continue to learn about the data, as you continue to understand the patterns (or expose information to your model that can help it find them) you codify what you learn in code and progressively grow your solution.\n\nAnd that is the plan 🙂\n\nPS. I am really excited about this competition! 🙂 Plan to share more models and techniques you will be able to implement into your solution over the next couple of days! 🙂 Please stay tuned! 📺\n\n### Other resources you might find useful:\n\n* [💡 [2 methods] How-to ensemble predictions 🏅🏅🏅](https://www.kaggle.com/code/radek1/2-methods-how-to-ensemble-predictions)\n* [local validation tracks public LB perfecty -- here is the setup](https://www.kaggle.com/competitions/otto-recommender-system/discussion/364991)\n* [💡 For my friends from Twitter and LinkedIn -- here is how to dive into this competition 🐳](https://www.kaggle.com/competitions/otto-recommender-system/discussion/368560)\n* [Full dataset processed to CSV/parquet files with optimized memory footprint](https://www.kaggle.com/competitions/otto-recommender-system/discussion/363843)\n* [co-visitation matrix - simplified, imprvd logic 🔥](https://www.kaggle.com/code/radek1/co-visitation-matrix-simplified-imprvd-logic)\n* [💡 Word2Vec How-to [training and submission]🚀🚀🚀](https://www.kaggle.com/code/radek1/word2vec-how-to-training-and-submission)",
    "2060331": "Hi, @radek1! Thanks for making this competition more understandable and fun. It's an overwhelming one to take as a first one 😅",
    "2061663": "Thank you for your kind words, @jamnik99! 🙌 Extremely glad I could be of help 🙂"
  },
  "source": "meta"
}