{
  "id": 367463,
  "title": "What happened so far? Full list of resources and discussions for a quick start to the competition 🚀",
  "url": "/competitions/otto-recommender-system/discussion/367463",
  "author_name": "",
  "post_date": "2022-11-20T19:13:45.030778Z",
  "votes": 21,
  "comment_count": 5,
  "views": 0,
  "content": "<p>Hi everyone,</p>\n<p>It's been a while since the competition started so I would like to share the all resources that have been discussed/shared so far for people who join the party later:</p>\n<h4>Discussions:</h4>\n<ol>\n<li><a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/364062\" target=\"_blank\">📈 What do we know so far? ⚡Summary with links to relevant resources</a> by <a href=\"https://www.kaggle.com/radek1\" target=\"_blank\">@radek1</a>. Radek summarized the first weeks of the competition quite well already + with additional more recent resources. </li>\n<li><a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/363874\" target=\"_blank\">A top-down perspective on the current metric values</a> by <a href=\"https://www.kaggle.com/narsil\" target=\"_blank\">@narsil</a>. A discussion about how sessions are truncated and how the dataset is created (also confirmed by the organizers).</li>\n<li><a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/363939\" target=\"_blank\">Yes, We Can Use Test Data Leakage</a> by <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a>. A discussion where Chris shows us why it makes sense to use the test data in the training and how (in the second notebook below).</li>\n<li><a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/364721\" target=\"_blank\">Recommendation Systems for Large Datasets</a> by <a href=\"https://www.kaggle.com/ravishah1\" target=\"_blank\">@ravishah1</a>. Ravi points out important parts of large scale recommender systems: candidate generation and ranking. He also gives additional directions on how to tackle those problems. </li>\n<li><a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/364216\" target=\"_blank\">💡A robust local validation framework 🚀🚀🚀</a> by <a href=\"https://www.kaggle.com/radek1\" target=\"_blank\">@radek1</a>. Radek shows us how to create a robust validation pipeline by using the last week of the train data as hold-out data without any inclusion of original test data.</li>\n<li><a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/364991\" target=\"_blank\">local validation tracks public LB perfecty -- here is the setup</a> by <a href=\"https://www.kaggle.com/radek1\" target=\"_blank\">@radek1</a>. Radek shares his local validation setup with code examples as explained above.</li>\n<li><a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/365358\" target=\"_blank\">💡 What is the co-visitation matrix, really?</a> by <a href=\"https://www.kaggle.com/radek1\" target=\"_blank\">@radek1</a>. Radek explains what co-visitation matrix is and how it is useful for this competition so far.</li>\n<li><a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/366138\" target=\"_blank\">Locate Real Users and Real Sessions EDA</a> by <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a>. With the accompanying notebook 4 - Chris shows us how to detect the real user sessions given the full month of data.</li>\n<li><a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/366194\" target=\"_blank\">[Starter pack] LGBMRanker with polars 🚀🚀🚀</a> by <a href=\"https://www.kaggle.com/radek1\" target=\"_blank\">@radek1</a>.</li>\n<li><a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/366893\" target=\"_blank\">[Starter Pack] Matrix Factorization [Pytorch + Merlin Dataloader] 🚀🚀🚀</a> by <a href=\"https://www.kaggle.com/radek1\" target=\"_blank\">@radek1</a>. In this discussion and accompanying notebook 6 - Radek shows us how to train a matrix factorization model using Pytorch - Merlin's Data Loader.</li>\n</ol>\n<h4>Notebooks:</h4>\n<ol>\n<li><a href=\"https://www.kaggle.com/code/vslaykovsky/co-visitation-matrix\" target=\"_blank\">Co-visitation Matrix</a> by <a href=\"https://www.kaggle.com/vslaykovsky\" target=\"_blank\">vslaykovsky</a>. Leveraging the idea of frequently viewed and bought items by creating a co-visitation matrix.</li>\n<li><a href=\"https://www.kaggle.com/code/cdeotte/test-data-leak-lb-boost\" target=\"_blank\">Test Data Leak - LB Boost</a> by <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a>. Chris demonstrates how to use the test data in the training. It also has a good track of experiment logs.</li>\n<li><a href=\"https://www.kaggle.com/code/cdeotte/candidate-rerank-model-lb-0-575?scriptVersionId=111214204\" target=\"_blank\">Candidate ReRank Model - [LB 0.575]</a> by <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a>. In this notebook, Chris shows us how to generate the candidates (most popular clicks, matrices, etc.) and rank them using hand-crafted features.</li>\n<li><a href=\"https://www.kaggle.com/code/cdeotte/time-series-eda-users-and-real-sessions\" target=\"_blank\">Time Series EDA - Users and Real Sessions</a> by <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a>.</li>\n<li><a href=\"https://www.kaggle.com/code/radek1/polars-proof-of-concept-lgbm-ranker/notebook\" target=\"_blank\">💡 [polars] Proof of concept: LGBM Ranker🧪🧪🧪</a> by <a href=\"https://www.kaggle.com/radek1\" target=\"_blank\">@radek1</a>. In this notebook, Radek shows us how to build an LGBMRanker model.</li>\n<li><a href=\"https://www.kaggle.com/code/radek1/matrix-factorization-pytorch-merlin-dataloader\" target=\"_blank\">💡Matrix Factorization [PyTorch+Merlin Dataloader]\n</a> by <a href=\"https://www.kaggle.com/radek1\" target=\"_blank\">@radek1</a>.</li>\n<li><a href=\"https://www.kaggle.com/code/theoviel/pretraining-with-merlin-s-transformers4rec\" target=\"_blank\">Pretraining with Merlin's Transformers4Rec🧙</a> by <a href=\"https://www.kaggle.com/theoviel\" target=\"_blank\">@theoviel</a>. In this notebook, Theo shows us how to use Merlin's Transformers4Rec library to pretrain an XLNet language model using the next item prediction task.</li>\n<li><a href=\"https://www.kaggle.com/code/cdeotte/compute-validation-score-cv-565\" target=\"_blank\">Compute Validation Score - [CV 565]</a> by <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a>. In this notebook, Chris shows us how to compute the validation score given Radek's schema replicating the original test period.</li>\n</ol>",
  "messages": [
    {
      "id": "2037622",
      "postDate": "11/20/2022 19:13:45",
      "content": "<p>Hi everyone,</p>\n<p>It's been a while since the competition started so I would like to share the all resources that have been discussed/shared so far for people who join the party later:</p>\n<h4>Discussions:</h4>\n<ol>\n<li><a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/364062\" target=\"_blank\">📈 What do we know so far? ⚡Summary with links to relevant resources</a> by <a href=\"https://www.kaggle.com/radek1\" target=\"_blank\">@radek1</a>. Radek summarized the first weeks of the competition quite well already + with additional more recent resources. </li>\n<li><a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/363874\" target=\"_blank\">A top-down perspective on the current metric values</a> by <a href=\"https://www.kaggle.com/narsil\" target=\"_blank\">@narsil</a>. A discussion about how sessions are truncated and how the dataset is created (also confirmed by the organizers).</li>\n<li><a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/363939\" target=\"_blank\">Yes, We Can Use Test Data Leakage</a> by <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a>. A discussion where Chris shows us why it makes sense to use the test data in the training and how (in the second notebook below).</li>\n<li><a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/364721\" target=\"_blank\">Recommendation Systems for Large Datasets</a> by <a href=\"https://www.kaggle.com/ravishah1\" target=\"_blank\">@ravishah1</a>. Ravi points out important parts of large scale recommender systems: candidate generation and ranking. He also gives additional directions on how to tackle those problems. </li>\n<li><a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/364216\" target=\"_blank\">💡A robust local validation framework 🚀🚀🚀</a> by <a href=\"https://www.kaggle.com/radek1\" target=\"_blank\">@radek1</a>. Radek shows us how to create a robust validation pipeline by using the last week of the train data as hold-out data without any inclusion of original test data.</li>\n<li><a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/364991\" target=\"_blank\">local validation tracks public LB perfecty -- here is the setup</a> by <a href=\"https://www.kaggle.com/radek1\" target=\"_blank\">@radek1</a>. Radek shares his local validation setup with code examples as explained above.</li>\n<li><a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/365358\" target=\"_blank\">💡 What is the co-visitation matrix, really?</a> by <a href=\"https://www.kaggle.com/radek1\" target=\"_blank\">@radek1</a>. Radek explains what co-visitation matrix is and how it is useful for this competition so far.</li>\n<li><a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/366138\" target=\"_blank\">Locate Real Users and Real Sessions EDA</a> by <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a>. With the accompanying notebook 4 - Chris shows us how to detect the real user sessions given the full month of data.</li>\n<li><a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/366194\" target=\"_blank\">[Starter pack] LGBMRanker with polars 🚀🚀🚀</a> by <a href=\"https://www.kaggle.com/radek1\" target=\"_blank\">@radek1</a>.</li>\n<li><a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/366893\" target=\"_blank\">[Starter Pack] Matrix Factorization [Pytorch + Merlin Dataloader] 🚀🚀🚀</a> by <a href=\"https://www.kaggle.com/radek1\" target=\"_blank\">@radek1</a>. In this discussion and accompanying notebook 6 - Radek shows us how to train a matrix factorization model using Pytorch - Merlin's Data Loader.</li>\n</ol>\n<h4>Notebooks:</h4>\n<ol>\n<li><a href=\"https://www.kaggle.com/code/vslaykovsky/co-visitation-matrix\" target=\"_blank\">Co-visitation Matrix</a> by <a href=\"https://www.kaggle.com/vslaykovsky\" target=\"_blank\">vslaykovsky</a>. Leveraging the idea of frequently viewed and bought items by creating a co-visitation matrix.</li>\n<li><a href=\"https://www.kaggle.com/code/cdeotte/test-data-leak-lb-boost\" target=\"_blank\">Test Data Leak - LB Boost</a> by <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a>. Chris demonstrates how to use the test data in the training. It also has a good track of experiment logs.</li>\n<li><a href=\"https://www.kaggle.com/code/cdeotte/candidate-rerank-model-lb-0-575?scriptVersionId=111214204\" target=\"_blank\">Candidate ReRank Model - [LB 0.575]</a> by <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a>. In this notebook, Chris shows us how to generate the candidates (most popular clicks, matrices, etc.) and rank them using hand-crafted features.</li>\n<li><a href=\"https://www.kaggle.com/code/cdeotte/time-series-eda-users-and-real-sessions\" target=\"_blank\">Time Series EDA - Users and Real Sessions</a> by <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a>.</li>\n<li><a href=\"https://www.kaggle.com/code/radek1/polars-proof-of-concept-lgbm-ranker/notebook\" target=\"_blank\">💡 [polars] Proof of concept: LGBM Ranker🧪🧪🧪</a> by <a href=\"https://www.kaggle.com/radek1\" target=\"_blank\">@radek1</a>. In this notebook, Radek shows us how to build an LGBMRanker model.</li>\n<li><a href=\"https://www.kaggle.com/code/radek1/matrix-factorization-pytorch-merlin-dataloader\" target=\"_blank\">💡Matrix Factorization [PyTorch+Merlin Dataloader]\n</a> by <a href=\"https://www.kaggle.com/radek1\" target=\"_blank\">@radek1</a>.</li>\n<li><a href=\"https://www.kaggle.com/code/theoviel/pretraining-with-merlin-s-transformers4rec\" target=\"_blank\">Pretraining with Merlin's Transformers4Rec🧙</a> by <a href=\"https://www.kaggle.com/theoviel\" target=\"_blank\">@theoviel</a>. In this notebook, Theo shows us how to use Merlin's Transformers4Rec library to pretrain an XLNet language model using the next item prediction task.</li>\n<li><a href=\"https://www.kaggle.com/code/cdeotte/compute-validation-score-cv-565\" target=\"_blank\">Compute Validation Score - [CV 565]</a> by <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a>. In this notebook, Chris shows us how to compute the validation score given Radek's schema replicating the original test period.</li>\n</ol>",
      "rawMarkdown": "Hi everyone,\n\nIt's been a while since the competition started so I would like to share the all resources that have been discussed/shared so far for people who join the party later:\n\n#### Discussions:\n\n1. [📈 What do we know so far? ⚡Summary with links to relevant resources](https://www.kaggle.com/competitions/otto-recommender-system/discussion/364062) by [@radek1](https://www.kaggle.com/radek1). Radek summarized the first weeks of the competition quite well already + with additional more recent resources. \n2. [A top-down perspective on the current metric values](https://www.kaggle.com/competitions/otto-recommender-system/discussion/363874) by [@narsil](https://www.kaggle.com/narsil). A discussion about how sessions are truncated and how the dataset is created (also confirmed by the organizers).\n3. [Yes, We Can Use Test Data Leakage](https://www.kaggle.com/competitions/otto-recommender-system/discussion/363939) by [@cdeotte](https://www.kaggle.com/cdeotte). A discussion where Chris shows us why it makes sense to use the test data in the training and how (in the second notebook below).\n4. [Recommendation Systems for Large Datasets](https://www.kaggle.com/competitions/otto-recommender-system/discussion/364721) by [@ravishah1](https://www.kaggle.com/ravishah1). Ravi points out important parts of large scale recommender systems: candidate generation and ranking. He also gives additional directions on how to tackle those problems. \n5. [💡A robust local validation framework 🚀🚀🚀](https://www.kaggle.com/competitions/otto-recommender-system/discussion/364216) by [@radek1](https://www.kaggle.com/radek1). Radek shows us how to create a robust validation pipeline by using the last week of the train data as hold-out data without any inclusion of original test data.\n6. [local validation tracks public LB perfecty -- here is the setup](https://www.kaggle.com/competitions/otto-recommender-system/discussion/364991) by [@radek1](https://www.kaggle.com/radek1). Radek shares his local validation setup with code examples as explained above.\n7. [💡 What is the co-visitation matrix, really?](https://www.kaggle.com/competitions/otto-recommender-system/discussion/365358) by [@radek1](https://www.kaggle.com/radek1). Radek explains what co-visitation matrix is and how it is useful for this competition so far.\n8. [Locate Real Users and Real Sessions EDA](https://www.kaggle.com/competitions/otto-recommender-system/discussion/366138) by [@cdeotte](https://www.kaggle.com/cdeotte). With the accompanying notebook 4 - Chris shows us how to detect the real user sessions given the full month of data.\n9. [[Starter pack] LGBMRanker with polars 🚀🚀🚀](https://www.kaggle.com/competitions/otto-recommender-system/discussion/366194) by [@radek1](https://www.kaggle.com/radek1).\n10. [[Starter Pack] Matrix Factorization [Pytorch + Merlin Dataloader] 🚀🚀🚀](https://www.kaggle.com/competitions/otto-recommender-system/discussion/366893) by [@radek1](https://www.kaggle.com/radek1). In this discussion and accompanying notebook 6 - Radek shows us how to train a matrix factorization model using Pytorch - Merlin's Data Loader.\n\n#### Notebooks:\n\n1. [Co-visitation Matrix](https://www.kaggle.com/code/vslaykovsky/co-visitation-matrix) by [vslaykovsky](https://www.kaggle.com/vslaykovsky). Leveraging the idea of frequently viewed and bought items by creating a co-visitation matrix.\n2. [Test Data Leak - LB Boost](https://www.kaggle.com/code/cdeotte/test-data-leak-lb-boost) by [@cdeotte](https://www.kaggle.com/cdeotte). Chris demonstrates how to use the test data in the training. It also has a good track of experiment logs.\n3. [Candidate ReRank Model - [LB 0.575]](https://www.kaggle.com/code/cdeotte/candidate-rerank-model-lb-0-575?scriptVersionId=111214204) by [@cdeotte](https://www.kaggle.com/cdeotte). In this notebook, Chris shows us how to generate the candidates (most popular clicks, matrices, etc.) and rank them using hand-crafted features.\n4. [Time Series EDA - Users and Real Sessions](https://www.kaggle.com/code/cdeotte/time-series-eda-users-and-real-sessions) by [@cdeotte](https://www.kaggle.com/cdeotte).\n5. [💡 [polars] Proof of concept: LGBM Ranker🧪🧪🧪](https://www.kaggle.com/code/radek1/polars-proof-of-concept-lgbm-ranker/notebook) by [@radek1](https://www.kaggle.com/radek1). In this notebook, Radek shows us how to build an LGBMRanker model.\n6. [💡Matrix Factorization [PyTorch+Merlin Dataloader]\n](https://www.kaggle.com/code/radek1/matrix-factorization-pytorch-merlin-dataloader) by [@radek1](https://www.kaggle.com/radek1).\n7. [Pretraining with Merlin's Transformers4Rec🧙](https://www.kaggle.com/code/theoviel/pretraining-with-merlin-s-transformers4rec) by [@theoviel](https://www.kaggle.com/theoviel). In this notebook, Theo shows us how to use Merlin's Transformers4Rec library to pretrain an XLNet language model using the next item prediction task.\n8. [Compute Validation Score - [CV 565]](https://www.kaggle.com/code/cdeotte/compute-validation-score-cv-565) by [@cdeotte](https://www.kaggle.com/cdeotte). In this notebook, Chris shows us how to compute the validation score given Radek's schema replicating the original test period.",
      "votes": null
    },
    {
      "id": "2037651",
      "postDate": "11/20/2022 19:30:04",
      "content": "<p>Nice summary</p>",
      "rawMarkdown": "Nice summary",
      "votes": null
    },
    {
      "id": "2037664",
      "postDate": "11/20/2022 19:49:49",
      "content": "<p>Thanks Chris! Thanks for the great resources as always. Half of it directly comes from your discussions/notebooks 😅</p>",
      "rawMarkdown": "Thanks Chris! Thanks for the great resources as always. Half of it directly comes from your discussions/notebooks 😅",
      "votes": null
    },
    {
      "id": "2037682",
      "postDate": "11/20/2022 20:17:06",
      "content": "<p>Thank you, i'm questioning whose solution is better : Work2Vec or Analyse Matrix TF-IDF. If the order of chose of product in one session is important, I'm thinking that work2vec is better than analyse tf-idf.</p>",
      "rawMarkdown": "Thank you, i'm questioning whose solution is better : Work2Vec or Analyse Matrix TF-IDF. If the order of chose of product in one session is important, I'm thinking that work2vec is better than analyse tf-idf.",
      "votes": null
    },
    {
      "id": "2037696",
      "postDate": "11/20/2022 20:33:34",
      "content": "<p>I would suggest using both and many additional others for candidate generation. Then use a ranking model for selection of final 20</p>",
      "rawMarkdown": "I would suggest using both and many additional others for candidate generation. Then use a ranking model for selection of final 20",
      "votes": null
    },
    {
      "id": "2037720",
      "postDate": "11/20/2022 21:36:37",
      "content": "<p>Thanks for putting this up here. Nice to have a list.</p>",
      "rawMarkdown": "Thanks for putting this up here. Nice to have a list.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2037651,
      "author_name": "cdeotte",
      "author_url": "",
      "post_date": "11/20/2022 19:30:04",
      "content": "<p>Nice summary</p>",
      "votes": null,
      "replies": [
        {
          "id": 2037664,
          "author_name": "snnclsr",
          "author_url": "",
          "post_date": "11/20/2022 19:49:49",
          "content": "<p>Thanks Chris! Thanks for the great resources as always. Half of it directly comes from your discussions/notebooks 😅</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2037682,
      "author_name": "lgregory",
      "author_url": "",
      "post_date": "11/20/2022 20:17:06",
      "content": "<p>Thank you, i'm questioning whose solution is better : Work2Vec or Analyse Matrix TF-IDF. If the order of chose of product in one session is important, I'm thinking that work2vec is better than analyse tf-idf.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2037696,
          "author_name": "cdeotte",
          "author_url": "",
          "post_date": "11/20/2022 20:33:34",
          "content": "<p>I would suggest using both and many additional others for candidate generation. Then use a ranking model for selection of final 20</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2037720,
      "author_name": "megan3",
      "author_url": "",
      "post_date": "11/20/2022 21:36:37",
      "content": "<p>Thanks for putting this up here. Nice to have a list.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2037622": "Hi everyone,\n\nIt's been a while since the competition started so I would like to share the all resources that have been discussed/shared so far for people who join the party later:\n\n#### Discussions:\n\n1. [📈 What do we know so far? ⚡Summary with links to relevant resources](https://www.kaggle.com/competitions/otto-recommender-system/discussion/364062) by [@radek1](https://www.kaggle.com/radek1). Radek summarized the first weeks of the competition quite well already + with additional more recent resources. \n2. [A top-down perspective on the current metric values](https://www.kaggle.com/competitions/otto-recommender-system/discussion/363874) by [@narsil](https://www.kaggle.com/narsil). A discussion about how sessions are truncated and how the dataset is created (also confirmed by the organizers).\n3. [Yes, We Can Use Test Data Leakage](https://www.kaggle.com/competitions/otto-recommender-system/discussion/363939) by [@cdeotte](https://www.kaggle.com/cdeotte). A discussion where Chris shows us why it makes sense to use the test data in the training and how (in the second notebook below).\n4. [Recommendation Systems for Large Datasets](https://www.kaggle.com/competitions/otto-recommender-system/discussion/364721) by [@ravishah1](https://www.kaggle.com/ravishah1). Ravi points out important parts of large scale recommender systems: candidate generation and ranking. He also gives additional directions on how to tackle those problems. \n5. [💡A robust local validation framework 🚀🚀🚀](https://www.kaggle.com/competitions/otto-recommender-system/discussion/364216) by [@radek1](https://www.kaggle.com/radek1). Radek shows us how to create a robust validation pipeline by using the last week of the train data as hold-out data without any inclusion of original test data.\n6. [local validation tracks public LB perfecty -- here is the setup](https://www.kaggle.com/competitions/otto-recommender-system/discussion/364991) by [@radek1](https://www.kaggle.com/radek1). Radek shares his local validation setup with code examples as explained above.\n7. [💡 What is the co-visitation matrix, really?](https://www.kaggle.com/competitions/otto-recommender-system/discussion/365358) by [@radek1](https://www.kaggle.com/radek1). Radek explains what co-visitation matrix is and how it is useful for this competition so far.\n8. [Locate Real Users and Real Sessions EDA](https://www.kaggle.com/competitions/otto-recommender-system/discussion/366138) by [@cdeotte](https://www.kaggle.com/cdeotte). With the accompanying notebook 4 - Chris shows us how to detect the real user sessions given the full month of data.\n9. [[Starter pack] LGBMRanker with polars 🚀🚀🚀](https://www.kaggle.com/competitions/otto-recommender-system/discussion/366194) by [@radek1](https://www.kaggle.com/radek1).\n10. [[Starter Pack] Matrix Factorization [Pytorch + Merlin Dataloader] 🚀🚀🚀](https://www.kaggle.com/competitions/otto-recommender-system/discussion/366893) by [@radek1](https://www.kaggle.com/radek1). In this discussion and accompanying notebook 6 - Radek shows us how to train a matrix factorization model using Pytorch - Merlin's Data Loader.\n\n#### Notebooks:\n\n1. [Co-visitation Matrix](https://www.kaggle.com/code/vslaykovsky/co-visitation-matrix) by [vslaykovsky](https://www.kaggle.com/vslaykovsky). Leveraging the idea of frequently viewed and bought items by creating a co-visitation matrix.\n2. [Test Data Leak - LB Boost](https://www.kaggle.com/code/cdeotte/test-data-leak-lb-boost) by [@cdeotte](https://www.kaggle.com/cdeotte). Chris demonstrates how to use the test data in the training. It also has a good track of experiment logs.\n3. [Candidate ReRank Model - [LB 0.575]](https://www.kaggle.com/code/cdeotte/candidate-rerank-model-lb-0-575?scriptVersionId=111214204) by [@cdeotte](https://www.kaggle.com/cdeotte). In this notebook, Chris shows us how to generate the candidates (most popular clicks, matrices, etc.) and rank them using hand-crafted features.\n4. [Time Series EDA - Users and Real Sessions](https://www.kaggle.com/code/cdeotte/time-series-eda-users-and-real-sessions) by [@cdeotte](https://www.kaggle.com/cdeotte).\n5. [💡 [polars] Proof of concept: LGBM Ranker🧪🧪🧪](https://www.kaggle.com/code/radek1/polars-proof-of-concept-lgbm-ranker/notebook) by [@radek1](https://www.kaggle.com/radek1). In this notebook, Radek shows us how to build an LGBMRanker model.\n6. [💡Matrix Factorization [PyTorch+Merlin Dataloader]\n](https://www.kaggle.com/code/radek1/matrix-factorization-pytorch-merlin-dataloader) by [@radek1](https://www.kaggle.com/radek1).\n7. [Pretraining with Merlin's Transformers4Rec🧙](https://www.kaggle.com/code/theoviel/pretraining-with-merlin-s-transformers4rec) by [@theoviel](https://www.kaggle.com/theoviel). In this notebook, Theo shows us how to use Merlin's Transformers4Rec library to pretrain an XLNet language model using the next item prediction task.\n8. [Compute Validation Score - [CV 565]](https://www.kaggle.com/code/cdeotte/compute-validation-score-cv-565) by [@cdeotte](https://www.kaggle.com/cdeotte). In this notebook, Chris shows us how to compute the validation score given Radek's schema replicating the original test period.",
    "2037651": "Nice summary",
    "2037664": "Thanks Chris! Thanks for the great resources as always. Half of it directly comes from your discussions/notebooks 😅",
    "2037682": "Thank you, i'm questioning whose solution is better : Work2Vec or Analyse Matrix TF-IDF. If the order of chose of product in one session is important, I'm thinking that work2vec is better than analyse tf-idf.",
    "2037696": "I would suggest using both and many additional others for candidate generation. Then use a ranking model for selection of final 20",
    "2037720": "Thanks for putting this up here. Nice to have a list."
  },
  "source": "meta"
}