{
  "id": 368560,
  "title": "💡 For my friends from Twitter and LinkedIn -- here is how to dive into this competition 🐳",
  "url": "/competitions/otto-recommender-system/discussion/368560",
  "author_name": "",
  "post_date": "2022-11-26T07:22:51.348571Z",
  "votes": 136,
  "comment_count": 8,
  "views": 0,
  "content": "<p>I am inviting my friends from Twitter and LinkedIn to join this competition and learn with us 🙂 This competition is spectacularly fun and the dataset (given its size and random sampling of the test set) makes for a great learning environment! (the feedback that you get from local CV / LB is valid, and serves as good feedback for experiments)</p>\n<p>So, if you just arrived from Twitter on LinkedIn, let me show you around 🙂</p>\n<p>First of all, before getting started, head <a href=\"https://www.kaggle.com/code/radek1/eda-an-overview-of-the-full-dataset\" target=\"_blank\">over to this EDA notebook</a> where I share a couple of observations about the competition data! 🙌 Always useful to understand your data a bit better 🙂</p>\n<h3>Use the right data</h3>\n<p>As in every ML project, having a good <code>train - val - test</code> split is essential.</p>\n<ul>\n<li>To make your life easier, use <a href=\"https://www.kaggle.com/datasets/radek1/otto-full-optimized-memory-footprint\" target=\"_blank\">the dataset I preprocessed to <code>parquet</code> files</a> (I also minimized the memory footprint without losing any information).</li>\n<li>As for the validation set, I reverse-engineered how the organizer went about creating the test set, you can find the implementation in this notebook: <a href=\"https://www.kaggle.com/code/radek1/a-robust-local-validation-framework\" target=\"_blank\">💡A robust local validation framework 🚀🚀🚀</a>. The validation set tracks the LB perfectly: <a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/364991\" target=\"_blank\">local validation tracks public LB perfectly  -- here is the setup</a>.</li>\n<li>But maybe you would rather prefer to use a validation set created with the organizer's code that they shared on github? I uploaded the validation set created in such a way <a href=\"https://www.kaggle.com/datasets/radek1/otto-train-and-test-data-for-local-validation\" target=\"_blank\">here</a>.</li>\n</ul>\n<h3>How to start on modeling</h3>\n<ul>\n<li>The covisitation matrices from <a href=\"https://www.kaggle.com/code/cdeotte/candidate-rerank-model-lb-0-575\" target=\"_blank\">Candidate ReRank Model - [LB 0.575]</a> by <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> are invaluable. At least 50% of the solution in the gold medal range will use them as their basis, if not more 🙂 Would be great to start there.</li>\n<li>As <a href=\"https://www.kaggle.com/ravishah1\" target=\"_blank\">@ravishah1</a> outlines in <a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/364721\" target=\"_blank\">Recommendation Systems for Large Datasets</a>, you will likely need a two-stage recommender to fare well in this competition (the post is extremely insightful, a recommended read!)</li>\n<li>I've put this notebook together to help you get started on building a two-stage pipeline with an LGBM ranker: <a href=\"https://www.kaggle.com/code/radek1/polars-proof-of-concept-lgbm-ranker\" target=\"_blank\">💡 [polars] Proof of concept: LGBM Ranker🧪🧪🧪</a></li>\n</ul>\n<h3>How to improve your solution?</h3>\n<p>With better candidate generation! 🙂</p>\n<ul>\n<li><a href=\"https://www.kaggle.com/code/radek1/word2vec-how-to-training-and-submission\" target=\"_blank\">💡 Word2Vec How-to [training and submission]🚀🚀🚀</a> -- this method is amazing! Use it to generate embeddings and nearest neighbor search (as in the notebook) for candidate generation (it is blazingly fast)</li>\n<li><a href=\"https://www.kaggle.com/code/radek1/matrix-factorization-pytorch-merlin-dataloader\" target=\"_blank\">💡Matrix Factorization [PyTorch+Merlin Dataloader]</a> -- I am 100% convinced winning solutions will include creative takes on generating candidates using Matrix Factorization 🙂 Use this notebook as a starter but come up with new approaches (how to segment sessions? how to pair <code>aids</code> for training? these questions are likely to be key)</li>\n</ul>\n<h3>General approach</h3>\n<p>With a good approach, you can do more with less! (less time invested, hardware resources, etc). Here are a couple of threads that can be of help:</p>\n<p><strong>A couple of related resources you might find useful:</strong></p>\n<ul>\n<li><a href=\"https://www.kaggle.com/code/radek1/2-methods-how-to-ensemble-predictions\" target=\"_blank\">💡 [2 methods] How-to ensemble predictions 🏅🏅🏅</a></li>\n<li><a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/364991\" target=\"_blank\">local validation tracks public LB perfecty -- here is the setup</a></li>\n<li><a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/368560\" target=\"_blank\">💡 For my friends from Twitter and LinkedIn -- here is how to dive into this competition 🐳</a></li>\n<li><a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/363843\" target=\"_blank\">Full dataset processed to CSV/parquet files with optimized memory footprint</a></li>\n<li><a href=\"https://www.kaggle.com/code/radek1/co-visitation-matrix-simplified-imprvd-logic\" target=\"_blank\">co-visitation matrix - simplified, imprvd logic 🔥</a></li>\n<li><a href=\"https://www.kaggle.com/code/radek1/word2vec-how-to-training-and-submission\" target=\"_blank\">💡 Word2Vec How-to [training and submission]🚀🚀🚀</a></li>\n</ul>\n<p>Hope these can help you get up to speed quickly! 🙂 I plan to share more advanced models soon, please stay tuned for more!</p>\n<p><strong>I would appreciate it if you could please upvote the posts/notebooks/datasets that you find useful 🙏 Trying to make this competition as fun as I can for as many people as I can 😊 Thank you for your help!</strong></p>",
  "messages": [
    {
      "id": "2044046",
      "postDate": "11/26/2022 07:22:51",
      "content": "<p>I am inviting my friends from Twitter and LinkedIn to join this competition and learn with us 🙂 This competition is spectacularly fun and the dataset (given its size and random sampling of the test set) makes for a great learning environment! (the feedback that you get from local CV / LB is valid, and serves as good feedback for experiments)</p>\n<p>So, if you just arrived from Twitter on LinkedIn, let me show you around 🙂</p>\n<p>First of all, before getting started, head <a href=\"https://www.kaggle.com/code/radek1/eda-an-overview-of-the-full-dataset\" target=\"_blank\">over to this EDA notebook</a> where I share a couple of observations about the competition data! 🙌 Always useful to understand your data a bit better 🙂</p>\n<h3>Use the right data</h3>\n<p>As in every ML project, having a good <code>train - val - test</code> split is essential.</p>\n<ul>\n<li>To make your life easier, use <a href=\"https://www.kaggle.com/datasets/radek1/otto-full-optimized-memory-footprint\" target=\"_blank\">the dataset I preprocessed to <code>parquet</code> files</a> (I also minimized the memory footprint without losing any information).</li>\n<li>As for the validation set, I reverse-engineered how the organizer went about creating the test set, you can find the implementation in this notebook: <a href=\"https://www.kaggle.com/code/radek1/a-robust-local-validation-framework\" target=\"_blank\">💡A robust local validation framework 🚀🚀🚀</a>. The validation set tracks the LB perfectly: <a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/364991\" target=\"_blank\">local validation tracks public LB perfectly  -- here is the setup</a>.</li>\n<li>But maybe you would rather prefer to use a validation set created with the organizer's code that they shared on github? I uploaded the validation set created in such a way <a href=\"https://www.kaggle.com/datasets/radek1/otto-train-and-test-data-for-local-validation\" target=\"_blank\">here</a>.</li>\n</ul>\n<h3>How to start on modeling</h3>\n<ul>\n<li>The covisitation matrices from <a href=\"https://www.kaggle.com/code/cdeotte/candidate-rerank-model-lb-0-575\" target=\"_blank\">Candidate ReRank Model - [LB 0.575]</a> by <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> are invaluable. At least 50% of the solution in the gold medal range will use them as their basis, if not more 🙂 Would be great to start there.</li>\n<li>As <a href=\"https://www.kaggle.com/ravishah1\" target=\"_blank\">@ravishah1</a> outlines in <a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/364721\" target=\"_blank\">Recommendation Systems for Large Datasets</a>, you will likely need a two-stage recommender to fare well in this competition (the post is extremely insightful, a recommended read!)</li>\n<li>I've put this notebook together to help you get started on building a two-stage pipeline with an LGBM ranker: <a href=\"https://www.kaggle.com/code/radek1/polars-proof-of-concept-lgbm-ranker\" target=\"_blank\">💡 [polars] Proof of concept: LGBM Ranker🧪🧪🧪</a></li>\n</ul>\n<h3>How to improve your solution?</h3>\n<p>With better candidate generation! 🙂</p>\n<ul>\n<li><a href=\"https://www.kaggle.com/code/radek1/word2vec-how-to-training-and-submission\" target=\"_blank\">💡 Word2Vec How-to [training and submission]🚀🚀🚀</a> -- this method is amazing! Use it to generate embeddings and nearest neighbor search (as in the notebook) for candidate generation (it is blazingly fast)</li>\n<li><a href=\"https://www.kaggle.com/code/radek1/matrix-factorization-pytorch-merlin-dataloader\" target=\"_blank\">💡Matrix Factorization [PyTorch+Merlin Dataloader]</a> -- I am 100% convinced winning solutions will include creative takes on generating candidates using Matrix Factorization 🙂 Use this notebook as a starter but come up with new approaches (how to segment sessions? how to pair <code>aids</code> for training? these questions are likely to be key)</li>\n</ul>\n<h3>General approach</h3>\n<p>With a good approach, you can do more with less! (less time invested, hardware resources, etc). Here are a couple of threads that can be of help:</p>\n<p><strong>A couple of related resources you might find useful:</strong></p>\n<ul>\n<li><a href=\"https://www.kaggle.com/code/radek1/2-methods-how-to-ensemble-predictions\" target=\"_blank\">💡 [2 methods] How-to ensemble predictions 🏅🏅🏅</a></li>\n<li><a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/364991\" target=\"_blank\">local validation tracks public LB perfecty -- here is the setup</a></li>\n<li><a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/368560\" target=\"_blank\">💡 For my friends from Twitter and LinkedIn -- here is how to dive into this competition 🐳</a></li>\n<li><a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/363843\" target=\"_blank\">Full dataset processed to CSV/parquet files with optimized memory footprint</a></li>\n<li><a href=\"https://www.kaggle.com/code/radek1/co-visitation-matrix-simplified-imprvd-logic\" target=\"_blank\">co-visitation matrix - simplified, imprvd logic 🔥</a></li>\n<li><a href=\"https://www.kaggle.com/code/radek1/word2vec-how-to-training-and-submission\" target=\"_blank\">💡 Word2Vec How-to [training and submission]🚀🚀🚀</a></li>\n</ul>\n<p>Hope these can help you get up to speed quickly! 🙂 I plan to share more advanced models soon, please stay tuned for more!</p>\n<p><strong>I would appreciate it if you could please upvote the posts/notebooks/datasets that you find useful 🙏 Trying to make this competition as fun as I can for as many people as I can 😊 Thank you for your help!</strong></p>",
      "rawMarkdown": "I am inviting my friends from Twitter and LinkedIn to join this competition and learn with us 🙂 This competition is spectacularly fun and the dataset (given its size and random sampling of the test set) makes for a great learning environment! (the feedback that you get from local CV / LB is valid, and serves as good feedback for experiments)\n\nSo, if you just arrived from Twitter on LinkedIn, let me show you around 🙂\n\nFirst of all, before getting started, head [over to this EDA notebook](https://www.kaggle.com/code/radek1/eda-an-overview-of-the-full-dataset) where I share a couple of observations about the competition data! 🙌 Always useful to understand your data a bit better 🙂\n\n### Use the right data\n\nAs in every ML project, having a good `train - val - test` split is essential.\n\n* To make your life easier, use [the dataset I preprocessed to `parquet` files](https://www.kaggle.com/datasets/radek1/otto-full-optimized-memory-footprint) (I also minimized the memory footprint without losing any information).\n* As for the validation set, I reverse-engineered how the organizer went about creating the test set, you can find the implementation in this notebook: [💡A robust local validation framework 🚀🚀🚀](https://www.kaggle.com/code/radek1/a-robust-local-validation-framework). The validation set tracks the LB perfectly: [local validation tracks public LB perfectly  -- here is the setup](https://www.kaggle.com/competitions/otto-recommender-system/discussion/364991).\n* But maybe you would rather prefer to use a validation set created with the organizer's code that they shared on github? I uploaded the validation set created in such a way [here](https://www.kaggle.com/datasets/radek1/otto-train-and-test-data-for-local-validation).\n\n### How to start on modeling\n\n* The covisitation matrices from [Candidate ReRank Model - [LB 0.575]](https://www.kaggle.com/code/cdeotte/candidate-rerank-model-lb-0-575) by @cdeotte are invaluable. At least 50% of the solution in the gold medal range will use them as their basis, if not more 🙂 Would be great to start there.\n* As @ravishah1 outlines in [Recommendation Systems for Large Datasets](https://www.kaggle.com/competitions/otto-recommender-system/discussion/364721), you will likely need a two-stage recommender to fare well in this competition (the post is extremely insightful, a recommended read!)\n* I've put this notebook together to help you get started on building a two-stage pipeline with an LGBM ranker: [💡 [polars] Proof of concept: LGBM Ranker🧪🧪🧪](https://www.kaggle.com/code/radek1/polars-proof-of-concept-lgbm-ranker)\n\n### How to improve your solution?\n\nWith better candidate generation! 🙂\n\n* [💡 Word2Vec How-to [training and submission]🚀🚀🚀](https://www.kaggle.com/code/radek1/word2vec-how-to-training-and-submission) -- this method is amazing! Use it to generate embeddings and nearest neighbor search (as in the notebook) for candidate generation (it is blazingly fast)\n* [💡Matrix Factorization [PyTorch+Merlin Dataloader]](https://www.kaggle.com/code/radek1/matrix-factorization-pytorch-merlin-dataloader) -- I am 100% convinced winning solutions will include creative takes on generating candidates using Matrix Factorization 🙂 Use this notebook as a starter but come up with new approaches (how to segment sessions? how to pair `aids` for training? these questions are likely to be key)\n\n### General approach\n\nWith a good approach, you can do more with less! (less time invested, hardware resources, etc). Here are a couple of threads that can be of help:\n\n**A couple of related resources you might find useful:**\n\n* [💡 [2 methods] How-to ensemble predictions 🏅🏅🏅](https://www.kaggle.com/code/radek1/2-methods-how-to-ensemble-predictions)\n* [local validation tracks public LB perfecty -- here is the setup](https://www.kaggle.com/competitions/otto-recommender-system/discussion/364991)\n* [💡 For my friends from Twitter and LinkedIn -- here is how to dive into this competition 🐳](https://www.kaggle.com/competitions/otto-recommender-system/discussion/368560)\n* [Full dataset processed to CSV/parquet files with optimized memory footprint](https://www.kaggle.com/competitions/otto-recommender-system/discussion/363843)\n* [co-visitation matrix - simplified, imprvd logic 🔥](https://www.kaggle.com/code/radek1/co-visitation-matrix-simplified-imprvd-logic)\n* [💡 Word2Vec How-to [training and submission]🚀🚀🚀](https://www.kaggle.com/code/radek1/word2vec-how-to-training-and-submission)\n\nHope these can help you get up to speed quickly! 🙂 I plan to share more advanced models soon, please stay tuned for more!\n\n**I would appreciate it if you could please upvote the posts/notebooks/datasets that you find useful 🙏 Trying to make this competition as fun as I can for as many people as I can 😊 Thank you for your help!**",
      "votes": null
    },
    {
      "id": "2045953",
      "postDate": "11/27/2022 19:56:42",
      "content": "<p>Great wrap-up as always, <a href=\"https://www.kaggle.com/radek1\" target=\"_blank\">@radek1</a> !</p>\n<p>Do you know if the LB metric calculation changed? I'm only getting LB 0.564 with the awesome kernel \"Candidate ReRank Model - [LB 0.575]\".</p>",
      "rawMarkdown": "Great wrap-up as always, @radek1 !\n\nDo you know if the LB metric calculation changed? I'm only getting LB 0.564 with the awesome kernel \"Candidate ReRank Model - [LB 0.575]\".",
      "votes": null
    },
    {
      "id": "2046046",
      "postDate": "11/27/2022 21:42:31",
      "content": "<p>Hey <a href=\"https://www.kaggle.com/gabrielmoraesbarros\" target=\"_blank\">@gabrielmoraesbarros</a>! Thank you! 🙂</p>\n<p>No, I don't believe anything has changed with the metric, maybe you made some modifications to the notebook?</p>",
      "rawMarkdown": "Hey @gabrielmoraesbarros! Thank you! 🙂\n\nNo, I don't believe anything has changed with the metric, maybe you made some modifications to the notebook?",
      "votes": null
    },
    {
      "id": "2046051",
      "postDate": "11/27/2022 21:52:32",
      "content": "<p>Thanks for the reply.</p>\n<p>I've just tried version 4 and got 0.575, but the v5 was 0.564.</p>\n<p>Now I will check your polars/LightGBM kernel, since I am learning a lot in this competition. 🙏</p>",
      "rawMarkdown": "Thanks for the reply.\n\nI've just tried version 4 and got 0.575, but the v5 was 0.564.\n\nNow I will check your polars/LightGBM kernel, since I am learning a lot in this competition. 🙏",
      "votes": null
    },
    {
      "id": "2046099",
      "postDate": "11/27/2022 23:18:30",
      "content": "<p>Ah, makes sense! 🙂 Awesome to hear you are having a good time in the competition!🙂 I am enjoying it quite a lot myself!</p>",
      "rawMarkdown": "Ah, makes sense! 🙂 Awesome to hear you are having a good time in the competition!🙂 I am enjoying it quite a lot myself!",
      "votes": null
    },
    {
      "id": "2055705",
      "postDate": "12/05/2022 10:41:43",
      "content": "<p>Thank you very much for sharing this, it helps me a lot for OttO. 👍👍👍</p>",
      "rawMarkdown": "Thank you very much for sharing this, it helps me a lot for OttO. 👍👍👍",
      "votes": null
    },
    {
      "id": "2055707",
      "postDate": "12/05/2022 10:48:14",
      "content": "<p>Awesome <a href=\"https://www.kaggle.com/leiwong\" target=\"_blank\">@leiwong</a>, super happy to hear! 🙂 Thank you for your comment! 🙌 </p>",
      "rawMarkdown": "Awesome @leiwong, super happy to hear! 🙂 Thank you for your comment! 🙌",
      "votes": null
    },
    {
      "id": "2090612",
      "postDate": "01/07/2023 14:11:38",
      "content": "<p>Very good introduction pack for this competition. You are doing a very good work to attract more people to contribute and have fun competing in this competition. And the quality of the material you are systematized is very high. <br>\nThank you for your contributions.</p>",
      "rawMarkdown": "Very good introduction pack for this competition. You are doing a very good work to attract more people to contribute and have fun competing in this competition. And the quality of the material you are systematized is very high. \nThank you for your contributions.",
      "votes": null
    },
    {
      "id": "2090982",
      "postDate": "01/07/2023 22:02:58",
      "content": "<p>Hey <a href=\"https://www.kaggle.com/gpreda\" target=\"_blank\">@gpreda</a>! Very nice to see you here 🙂 And thank you so much for your feedback! Really appreciate it! 🙌 </p>",
      "rawMarkdown": "Hey @gpreda! Very nice to see you here 🙂 And thank you so much for your feedback! Really appreciate it! 🙌",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2045953,
      "author_name": "gabrielmoraesbarros",
      "author_url": "",
      "post_date": "11/27/2022 19:56:42",
      "content": "<p>Great wrap-up as always, <a href=\"https://www.kaggle.com/radek1\" target=\"_blank\">@radek1</a> !</p>\n<p>Do you know if the LB metric calculation changed? I'm only getting LB 0.564 with the awesome kernel \"Candidate ReRank Model - [LB 0.575]\".</p>",
      "votes": null,
      "replies": [
        {
          "id": 2046046,
          "author_name": "radek1",
          "author_url": "",
          "post_date": "11/27/2022 21:42:31",
          "content": "<p>Hey <a href=\"https://www.kaggle.com/gabrielmoraesbarros\" target=\"_blank\">@gabrielmoraesbarros</a>! Thank you! 🙂</p>\n<p>No, I don't believe anything has changed with the metric, maybe you made some modifications to the notebook?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 2046051,
          "author_name": "gabrielmoraesbarros",
          "author_url": "",
          "post_date": "11/27/2022 21:52:32",
          "content": "<p>Thanks for the reply.</p>\n<p>I've just tried version 4 and got 0.575, but the v5 was 0.564.</p>\n<p>Now I will check your polars/LightGBM kernel, since I am learning a lot in this competition. 🙏</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 2046099,
          "author_name": "radek1",
          "author_url": "",
          "post_date": "11/27/2022 23:18:30",
          "content": "<p>Ah, makes sense! 🙂 Awesome to hear you are having a good time in the competition!🙂 I am enjoying it quite a lot myself!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2055705,
      "author_name": "leiwong",
      "author_url": "",
      "post_date": "12/05/2022 10:41:43",
      "content": "<p>Thank you very much for sharing this, it helps me a lot for OttO. 👍👍👍</p>",
      "votes": null,
      "replies": [
        {
          "id": 2055707,
          "author_name": "radek1",
          "author_url": "",
          "post_date": "12/05/2022 10:48:14",
          "content": "<p>Awesome <a href=\"https://www.kaggle.com/leiwong\" target=\"_blank\">@leiwong</a>, super happy to hear! 🙂 Thank you for your comment! 🙌 </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2090612,
      "author_name": "gpreda",
      "author_url": "",
      "post_date": "01/07/2023 14:11:38",
      "content": "<p>Very good introduction pack for this competition. You are doing a very good work to attract more people to contribute and have fun competing in this competition. And the quality of the material you are systematized is very high. <br>\nThank you for your contributions.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2090982,
          "author_name": "radek1",
          "author_url": "",
          "post_date": "01/07/2023 22:02:58",
          "content": "<p>Hey <a href=\"https://www.kaggle.com/gpreda\" target=\"_blank\">@gpreda</a>! Very nice to see you here 🙂 And thank you so much for your feedback! Really appreciate it! 🙌 </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2044046": "I am inviting my friends from Twitter and LinkedIn to join this competition and learn with us 🙂 This competition is spectacularly fun and the dataset (given its size and random sampling of the test set) makes for a great learning environment! (the feedback that you get from local CV / LB is valid, and serves as good feedback for experiments)\n\nSo, if you just arrived from Twitter on LinkedIn, let me show you around 🙂\n\nFirst of all, before getting started, head [over to this EDA notebook](https://www.kaggle.com/code/radek1/eda-an-overview-of-the-full-dataset) where I share a couple of observations about the competition data! 🙌 Always useful to understand your data a bit better 🙂\n\n### Use the right data\n\nAs in every ML project, having a good `train - val - test` split is essential.\n\n* To make your life easier, use [the dataset I preprocessed to `parquet` files](https://www.kaggle.com/datasets/radek1/otto-full-optimized-memory-footprint) (I also minimized the memory footprint without losing any information).\n* As for the validation set, I reverse-engineered how the organizer went about creating the test set, you can find the implementation in this notebook: [💡A robust local validation framework 🚀🚀🚀](https://www.kaggle.com/code/radek1/a-robust-local-validation-framework). The validation set tracks the LB perfectly: [local validation tracks public LB perfectly  -- here is the setup](https://www.kaggle.com/competitions/otto-recommender-system/discussion/364991).\n* But maybe you would rather prefer to use a validation set created with the organizer's code that they shared on github? I uploaded the validation set created in such a way [here](https://www.kaggle.com/datasets/radek1/otto-train-and-test-data-for-local-validation).\n\n### How to start on modeling\n\n* The covisitation matrices from [Candidate ReRank Model - [LB 0.575]](https://www.kaggle.com/code/cdeotte/candidate-rerank-model-lb-0-575) by @cdeotte are invaluable. At least 50% of the solution in the gold medal range will use them as their basis, if not more 🙂 Would be great to start there.\n* As @ravishah1 outlines in [Recommendation Systems for Large Datasets](https://www.kaggle.com/competitions/otto-recommender-system/discussion/364721), you will likely need a two-stage recommender to fare well in this competition (the post is extremely insightful, a recommended read!)\n* I've put this notebook together to help you get started on building a two-stage pipeline with an LGBM ranker: [💡 [polars] Proof of concept: LGBM Ranker🧪🧪🧪](https://www.kaggle.com/code/radek1/polars-proof-of-concept-lgbm-ranker)\n\n### How to improve your solution?\n\nWith better candidate generation! 🙂\n\n* [💡 Word2Vec How-to [training and submission]🚀🚀🚀](https://www.kaggle.com/code/radek1/word2vec-how-to-training-and-submission) -- this method is amazing! Use it to generate embeddings and nearest neighbor search (as in the notebook) for candidate generation (it is blazingly fast)\n* [💡Matrix Factorization [PyTorch+Merlin Dataloader]](https://www.kaggle.com/code/radek1/matrix-factorization-pytorch-merlin-dataloader) -- I am 100% convinced winning solutions will include creative takes on generating candidates using Matrix Factorization 🙂 Use this notebook as a starter but come up with new approaches (how to segment sessions? how to pair `aids` for training? these questions are likely to be key)\n\n### General approach\n\nWith a good approach, you can do more with less! (less time invested, hardware resources, etc). Here are a couple of threads that can be of help:\n\n**A couple of related resources you might find useful:**\n\n* [💡 [2 methods] How-to ensemble predictions 🏅🏅🏅](https://www.kaggle.com/code/radek1/2-methods-how-to-ensemble-predictions)\n* [local validation tracks public LB perfecty -- here is the setup](https://www.kaggle.com/competitions/otto-recommender-system/discussion/364991)\n* [💡 For my friends from Twitter and LinkedIn -- here is how to dive into this competition 🐳](https://www.kaggle.com/competitions/otto-recommender-system/discussion/368560)\n* [Full dataset processed to CSV/parquet files with optimized memory footprint](https://www.kaggle.com/competitions/otto-recommender-system/discussion/363843)\n* [co-visitation matrix - simplified, imprvd logic 🔥](https://www.kaggle.com/code/radek1/co-visitation-matrix-simplified-imprvd-logic)\n* [💡 Word2Vec How-to [training and submission]🚀🚀🚀](https://www.kaggle.com/code/radek1/word2vec-how-to-training-and-submission)\n\nHope these can help you get up to speed quickly! 🙂 I plan to share more advanced models soon, please stay tuned for more!\n\n**I would appreciate it if you could please upvote the posts/notebooks/datasets that you find useful 🙏 Trying to make this competition as fun as I can for as many people as I can 😊 Thank you for your help!**",
    "2045953": "Great wrap-up as always, @radek1 !\n\nDo you know if the LB metric calculation changed? I'm only getting LB 0.564 with the awesome kernel \"Candidate ReRank Model - [LB 0.575]\".",
    "2046046": "Hey @gabrielmoraesbarros! Thank you! 🙂\n\nNo, I don't believe anything has changed with the metric, maybe you made some modifications to the notebook?",
    "2046051": "Thanks for the reply.\n\nI've just tried version 4 and got 0.575, but the v5 was 0.564.\n\nNow I will check your polars/LightGBM kernel, since I am learning a lot in this competition. 🙏",
    "2046099": "Ah, makes sense! 🙂 Awesome to hear you are having a good time in the competition!🙂 I am enjoying it quite a lot myself!",
    "2055705": "Thank you very much for sharing this, it helps me a lot for OttO. 👍👍👍",
    "2055707": "Awesome @leiwong, super happy to hear! 🙂 Thank you for your comment! 🙌",
    "2090612": "Very good introduction pack for this competition. You are doing a very good work to attract more people to contribute and have fun competing in this competition. And the quality of the material you are systematized is very high. \nThank you for your contributions.",
    "2090982": "Hey @gpreda! Very nice to see you here 🙂 And thank you so much for your feedback! Really appreciate it! 🙌"
  },
  "source": "meta"
}