{
  "id": 363814,
  "title": "Resources for getting started with recommender systems",
  "url": "/competitions/otto-recommender-system/discussion/363814",
  "author_name": "",
  "post_date": "2022-11-03T08:22:54.480178200Z",
  "votes": 47,
  "comment_count": 16,
  "views": 0,
  "content": "<p>Hey!</p>\n<p>I am super excited for this competition 😊 Have been perusing the world of recommender systems for the last 6 months or so, will interesting to put all that to a test 🙂 </p>\n<p>In terms of good learning resources to get started with recommender systems, there are not that many, unfortunately. </p>\n<p>I would highly recommend <a href=\"https://youtu.be/bLhq63ygoU8\" target=\"_blank\">this lecture</a> by a former colleague of mine, Xavier Amatriain. It provides the best introduction to thinking about recommendations, what recommender systems are, what can be achieved, that I have ever come across.</p>\n<p>This particular problem that we will be working on in this competition are session based recommendations. Essentially, from a timeseries of actions taken by a user in a single session (single visit to a website), we want to predict what is the likely next action a user might take (and most importantly, on what item the action is likely to be taken!)</p>\n<p>Session recommendations are a very hot topic. They pull in a lot of exciting concepts together such as serving real time predictions (versus predictions calculated offline, usually in batches). And their sequential nature lends themselves well to modelling with RNNs or more recently, Transformers!</p>\n<p>There is a lot of complexity and interesting ideas to all of this 😊 I am hoping we will have a chance to explore many of them in this competition together!</p>\n<p>If you'd like to read more about using Transformers for session based recommendations, here is <a href=\"https://medium.com/nvidia-merlin/transformers4rec-4523cc7d8fa8\" target=\"_blank\">a very nice blog post</a> from my colleagues.</p>\n<p>So looking forward to jumping into this competition 🥳</p>\n<p>Happy Kaggling!</p>\n<h3>Other resources you might find useful:</h3>\n<ul>\n<li><a href=\"https://www.kaggle.com/code/radek1/2-methods-how-to-ensemble-predictions\" target=\"_blank\">💡 [2 methods] How-to ensemble predictions 🏅🏅🏅</a></li>\n<li><a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/364991\" target=\"_blank\">local validation tracks public LB perfecty -- here is the setup</a></li>\n<li><a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/368560\" target=\"_blank\">💡 For my friends from Twitter and LinkedIn -- here is how to dive into this competition 🐳</a></li>\n<li><a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/363843\" target=\"_blank\">Full dataset processed to CSV/parquet files with optimized memory footprint</a></li>\n<li><a href=\"https://www.kaggle.com/code/radek1/co-visitation-matrix-simplified-imprvd-logic\" target=\"_blank\">co-visitation matrix - simplified, imprvd logic 🔥</a></li>\n<li><a href=\"https://www.kaggle.com/code/radek1/word2vec-how-to-training-and-submission\" target=\"_blank\">💡 Word2Vec How-to [training and submission]🚀🚀🚀</a></li>\n</ul>",
  "messages": [
    {
      "id": "2015354",
      "postDate": "11/03/2022 08:22:54",
      "content": "<p>Hey!</p>\n<p>I am super excited for this competition 😊 Have been perusing the world of recommender systems for the last 6 months or so, will interesting to put all that to a test 🙂 </p>\n<p>In terms of good learning resources to get started with recommender systems, there are not that many, unfortunately. </p>\n<p>I would highly recommend <a href=\"https://youtu.be/bLhq63ygoU8\" target=\"_blank\">this lecture</a> by a former colleague of mine, Xavier Amatriain. It provides the best introduction to thinking about recommendations, what recommender systems are, what can be achieved, that I have ever come across.</p>\n<p>This particular problem that we will be working on in this competition are session based recommendations. Essentially, from a timeseries of actions taken by a user in a single session (single visit to a website), we want to predict what is the likely next action a user might take (and most importantly, on what item the action is likely to be taken!)</p>\n<p>Session recommendations are a very hot topic. They pull in a lot of exciting concepts together such as serving real time predictions (versus predictions calculated offline, usually in batches). And their sequential nature lends themselves well to modelling with RNNs or more recently, Transformers!</p>\n<p>There is a lot of complexity and interesting ideas to all of this 😊 I am hoping we will have a chance to explore many of them in this competition together!</p>\n<p>If you'd like to read more about using Transformers for session based recommendations, here is <a href=\"https://medium.com/nvidia-merlin/transformers4rec-4523cc7d8fa8\" target=\"_blank\">a very nice blog post</a> from my colleagues.</p>\n<p>So looking forward to jumping into this competition 🥳</p>\n<p>Happy Kaggling!</p>\n<h3>Other resources you might find useful:</h3>\n<ul>\n<li><a href=\"https://www.kaggle.com/code/radek1/2-methods-how-to-ensemble-predictions\" target=\"_blank\">💡 [2 methods] How-to ensemble predictions 🏅🏅🏅</a></li>\n<li><a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/364991\" target=\"_blank\">local validation tracks public LB perfecty -- here is the setup</a></li>\n<li><a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/368560\" target=\"_blank\">💡 For my friends from Twitter and LinkedIn -- here is how to dive into this competition 🐳</a></li>\n<li><a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/363843\" target=\"_blank\">Full dataset processed to CSV/parquet files with optimized memory footprint</a></li>\n<li><a href=\"https://www.kaggle.com/code/radek1/co-visitation-matrix-simplified-imprvd-logic\" target=\"_blank\">co-visitation matrix - simplified, imprvd logic 🔥</a></li>\n<li><a href=\"https://www.kaggle.com/code/radek1/word2vec-how-to-training-and-submission\" target=\"_blank\">💡 Word2Vec How-to [training and submission]🚀🚀🚀</a></li>\n</ul>",
      "rawMarkdown": "Hey!\n\nI am super excited for this competition 😊 Have been perusing the world of recommender systems for the last 6 months or so, will interesting to put all that to a test 🙂 \n\nIn terms of good learning resources to get started with recommender systems, there are not that many, unfortunately. \n \nI would highly recommend [this lecture](https://youtu.be/bLhq63ygoU8) by a former colleague of mine, Xavier Amatriain. It provides the best introduction to thinking about recommendations, what recommender systems are, what can be achieved, that I have ever come across.\n\nThis particular problem that we will be working on in this competition are session based recommendations. Essentially, from a timeseries of actions taken by a user in a single session (single visit to a website), we want to predict what is the likely next action a user might take (and most importantly, on what item the action is likely to be taken!)\n\nSession recommendations are a very hot topic. They pull in a lot of exciting concepts together such as serving real time predictions (versus predictions calculated offline, usually in batches). And their sequential nature lends themselves well to modelling with RNNs or more recently, Transformers!\n\nThere is a lot of complexity and interesting ideas to all of this 😊 I am hoping we will have a chance to explore many of them in this competition together!\n\nIf you'd like to read more about using Transformers for session based recommendations, here is [a very nice blog post](https://medium.com/nvidia-merlin/transformers4rec-4523cc7d8fa8) from my colleagues.\n\nSo looking forward to jumping into this competition 🥳\n\nHappy Kaggling!\n\n### Other resources you might find useful:\n\n* [💡 [2 methods] How-to ensemble predictions 🏅🏅🏅](https://www.kaggle.com/code/radek1/2-methods-how-to-ensemble-predictions)\n* [local validation tracks public LB perfecty -- here is the setup](https://www.kaggle.com/competitions/otto-recommender-system/discussion/364991)\n* [💡 For my friends from Twitter and LinkedIn -- here is how to dive into this competition 🐳](https://www.kaggle.com/competitions/otto-recommender-system/discussion/368560)\n* [Full dataset processed to CSV/parquet files with optimized memory footprint](https://www.kaggle.com/competitions/otto-recommender-system/discussion/363843)\n* [co-visitation matrix - simplified, imprvd logic 🔥](https://www.kaggle.com/code/radek1/co-visitation-matrix-simplified-imprvd-logic)\n* [💡 Word2Vec How-to [training and submission]🚀🚀🚀](https://www.kaggle.com/code/radek1/word2vec-how-to-training-and-submission)",
      "votes": null
    },
    {
      "id": "2015715",
      "postDate": "11/03/2022 13:47:05",
      "content": "<p>Could you post a notebook on how to get Transformers4Rec running on kaggle?<br>\nI wanted to try out the following notebook:<br>\n<a href=\"https://github.com/NVIDIA-Merlin/Transformers4Rec/blob/main/examples/end-to-end-session-based/01-ETL-with-NVTabular.ipynb\" target=\"_blank\">https://github.com/NVIDIA-Merlin/Transformers4Rec/blob/main/examples/end-to-end-session-based/01-ETL-with-NVTabular.ipynb</a><br>\nbut had problems…</p>",
      "rawMarkdown": "Could you post a notebook on how to get Transformers4Rec running on kaggle?\nI wanted to try out the following notebook:\nhttps://github.com/NVIDIA-Merlin/Transformers4Rec/blob/main/examples/end-to-end-session-based/01-ETL-with-NVTabular.ipynb\nbut had problems...",
      "votes": null
    },
    {
      "id": "2015722",
      "postDate": "11/03/2022 13:52:47",
      "content": "<p>Will look into that! 🙌</p>\n<p>The transformers will soon be coming to Merlin Models (thous will be HuggingFace transformers implemented in Tensorflow) so maybe that will make it a bit easier to run here on Kaggle, we will see 🙂</p>",
      "rawMarkdown": "Will look into that! 🙌\n\nThe transformers will soon be coming to Merlin Models (thous will be HuggingFace transformers implemented in Tensorflow) so maybe that will make it a bit easier to run here on Kaggle, we will see 🙂",
      "votes": null
    },
    {
      "id": "2016161",
      "postDate": "11/03/2022 20:21:46",
      "content": "<p>Thanks a lot for sharing, coincidentally I've also been reading up on it recently because I am interested in trying to apply it on a different domain. Excited for this comp too!</p>",
      "rawMarkdown": "Thanks a lot for sharing, coincidentally I've also been reading up on it recently because I am interested in trying to apply it on a different domain. Excited for this comp too!",
      "votes": null
    },
    {
      "id": "2016249",
      "postDate": "11/03/2022 21:32:25",
      "content": "<p>Thanks, <a href=\"https://www.kaggle.com/ollibolli\" target=\"_blank\">@ollibolli</a>! Really glad you are finding this useful 🙏 BTW what domain are you thinking of applying this to? Just curious, don't need to answer if you don't want to share 🙂</p>",
      "rawMarkdown": "Thanks, @ollibolli! Really glad you are finding this useful 🙏 BTW what domain are you thinking of applying this to? Just curious, don't need to answer if you don't want to share 🙂",
      "votes": null
    },
    {
      "id": "2017263",
      "postDate": "11/04/2022 16:46:42",
      "content": "<p>Thanks Radek, <br>\nno worries, I apologize I was being cryptic I didn't intend to be. One of the things I love most about Kaggle is that everyone is sharing, like you!<br>\nMight sound silly but I've been getting into soccer analytics and I thought about applying it to player scouting and to extension maybe even selection with some tweaks. I am quite the newbie so it might lead nowhere haha.</p>",
      "rawMarkdown": "Thanks Radek, \nno worries, I apologize I was being cryptic I didn't intend to be. One of the things I love most about Kaggle is that everyone is sharing, like you!\nMight sound silly but I've been getting into soccer analytics and I thought about applying it to player scouting and to extension maybe even selection with some tweaks. I am quite the newbie so it might lead nowhere haha.",
      "votes": null
    },
    {
      "id": "2017463",
      "postDate": "11/04/2022 20:30:40",
      "content": "<p>Have you seen the Bundesliga Dataset yet? Its quite interesting. I watched a live coding on prev. competition about this, maybe you want to check it out. <a href=\"https://youtu.be/fNHW7zvi_F8\" target=\"_blank\">https://youtu.be/fNHW7zvi_F8</a></p>",
      "rawMarkdown": "Have you seen the Bundesliga Dataset yet? Its quite interesting. I watched a live coding on prev. competition about this, maybe you want to check it out. https://youtu.be/fNHW7zvi_F8",
      "votes": null
    },
    {
      "id": "2017778",
      "postDate": "11/05/2022 06:30:32",
      "content": "<p>Thank you for your reply, <a href=\"https://www.kaggle.com/oliver\" target=\"_blank\">@oliver</a>! Sounds like an exciting project! 🙂</p>",
      "rawMarkdown": "Thank you for your reply, @oliver! Sounds like an exciting project! 🙂",
      "votes": null
    },
    {
      "id": "2018044",
      "postDate": "11/05/2022 10:56:56",
      "content": "<p>Hi Ian, cheers, yes, I followed this quite closely, there is a group at my school that actually has a few research projects on this. It's a very exciting time :-).</p>",
      "rawMarkdown": "Hi Ian, cheers, yes, I followed this quite closely, there is a group at my school that actually has a few research projects on this. It's a very exciting time :-).",
      "votes": null
    },
    {
      "id": "2019123",
      "postDate": "11/06/2022 11:05:16",
      "content": "<p>Thanks, even I couldn't find out anything interesting. Will definitely have a look.</p>",
      "rawMarkdown": "Thanks, even I couldn't find out anything interesting. Will definitely have a look.",
      "votes": null
    },
    {
      "id": "2020487",
      "postDate": "11/07/2022 14:22:23",
      "content": "<p><a href=\"https://github.com/hidasib/GRU4Rec\" target=\"_blank\">https://github.com/hidasib/GRU4Rec</a> <br>\nimplementation of the algorithm in \"Session-based Recommendations with Recurrent Neural Networks\" paper.</p>",
      "rawMarkdown": "https://github.com/hidasib/GRU4Rec \nimplementation of the algorithm in \"Session-based Recommendations with Recurrent Neural Networks\" paper.",
      "votes": null
    },
    {
      "id": "2020557",
      "postDate": "11/07/2022 15:31:27",
      "content": "<p>Gratitude for sharing!</p>",
      "rawMarkdown": "Gratitude for sharing!",
      "votes": null
    },
    {
      "id": "2020801",
      "postDate": "11/07/2022 19:33:44",
      "content": "<p>Really glad you are finding this useful! 🙂</p>",
      "rawMarkdown": "Really glad you are finding this useful! 🙂",
      "votes": null
    },
    {
      "id": "2023136",
      "postDate": "11/09/2022 14:58:19",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/radek1\" target=\"_blank\">@radek1</a> Thank you so much for the resources to help beginners to get started. </p>\n<p>I have finished watching and <a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/365039\" target=\"_blank\">note taking</a> the getting started videos on recsys from Andrew Ng (which are relatively easy to follow) and I am half way through those of Xavier Amatriain (two more hours to go for lecture 3+4) which I found many concepts are not very familiar and definitely need to rewatch a few more times.</p>\n<p>I wonder how much should I understand Xavier's lectures? I know you would approve that the best way to help understanding is to experiment with the models he talked about, but there are no accompany notebooks or codes available from the lectures right? </p>\n<p>I wonder have you found or implemented those the models or codes when you study the videos? or do you know where I can find them? or  do actually I need to study them now for learning recsys in order to prepare me for participating in OTTO competition? Or maybe focusing on experimenting with your notebooks is already enough for beginners?</p>\n<p>Thanks!</p>",
      "rawMarkdown": "Hi @radek1 Thank you so much for the resources to help beginners to get started. \n\nI have finished watching and [note taking](https://www.kaggle.com/competitions/otto-recommender-system/discussion/365039) the getting started videos on recsys from Andrew Ng (which are relatively easy to follow) and I am half way through those of Xavier Amatriain (two more hours to go for lecture 3+4) which I found many concepts are not very familiar and definitely need to rewatch a few more times.\n\nI wonder how much should I understand Xavier's lectures? I know you would approve that the best way to help understanding is to experiment with the models he talked about, but there are no accompany notebooks or codes available from the lectures right? \n\nI wonder have you found or implemented those the models or codes when you study the videos? or do you know where I can find them? or  do actually I need to study them now for learning recsys in order to prepare me for participating in OTTO competition? Or maybe focusing on experimenting with your notebooks is already enough for beginners?\n\nThanks!",
      "votes": null
    },
    {
      "id": "2023699",
      "postDate": "11/10/2022 01:00:01",
      "content": "<p>I just found a summary of Xavier's lecture 1 + 2 by <a href=\"https://www.kaggle.com/radek1\" target=\"_blank\">@radek1</a> <a href=\"https://twitter.com/radekosmulski/status/1565716248083566592\" target=\"_blank\">here</a>! Thanks Radek</p>",
      "rawMarkdown": "I just found a summary of Xavier's lecture 1 + 2 by @radek1 [here](https://twitter.com/radekosmulski/status/1565716248083566592)! Thanks Radek",
      "votes": null
    },
    {
      "id": "2023718",
      "postDate": "11/10/2022 01:25:12",
      "content": "<p>Xavier does a great job in his lectures of explaining the reasoning behind recommender systems, I think that is where the value lies. Very few people do that and nearly no one has experience comparable to his 🙂 As in, the thinking behind what you want and why you want to make the decisions that you do.</p>\n<p>But on the technical side, I think you will find there is a great disconnect between the ideas and what you can implement in code. Everything is very dataset-specific. In some way, this is another skillset that you have to build up, this ability to transfer those ideas to the problem you are working on. Even one or two ideas can go a long way.</p>\n<p>Here, for instance, this is a recommender system competition, but the curve ball is the high cardinality of action ids. How do you deal with this? There has been some really good discussion on this here: <a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/364722\" target=\"_blank\">🐘 the elephant in the room -- high cardinality of targets and what to do about this</a>.</p>\n<p>Or even more concretely, how do you work with the amount of data that you have available in this competition? This is on a completely different scale than most of the \"toy\" ML problems.</p>\n<p>So I think chipping away at this competition and learning from the conversations and the kernels that people post is a great way of getting hands-on experience 🙂 It really does go a long way in actual business context.</p>\n<p>Though some of the challenges that you might have to solve in the real world will be different, the problem-solving methodology that you can work out while trying to improve at a Kaggle competition goes a really long way 🙂 </p>\n<p>Meaning, I feel if you found your way to Kaggle and this competition, you probably are in a great spot to grow your practical set of abilities by working on this problem 🙂 </p>",
      "rawMarkdown": "Xavier does a great job in his lectures of explaining the reasoning behind recommender systems, I think that is where the value lies. Very few people do that and nearly no one has experience comparable to his 🙂 As in, the thinking behind what you want and why you want to make the decisions that you do.\n\nBut on the technical side, I think you will find there is a great disconnect between the ideas and what you can implement in code. Everything is very dataset-specific. In some way, this is another skillset that you have to build up, this ability to transfer those ideas to the problem you are working on. Even one or two ideas can go a long way.\n\nHere, for instance, this is a recommender system competition, but the curve ball is the high cardinality of action ids. How do you deal with this? There has been some really good discussion on this here: [🐘 the elephant in the room -- high cardinality of targets and what to do about this](https://www.kaggle.com/competitions/otto-recommender-system/discussion/364722).\n\nOr even more concretely, how do you work with the amount of data that you have available in this competition? This is on a completely different scale than most of the \"toy\" ML problems.\n\nSo I think chipping away at this competition and learning from the conversations and the kernels that people post is a great way of getting hands-on experience 🙂 It really does go a long way in actual business context.\n\nThough some of the challenges that you might have to solve in the real world will be different, the problem-solving methodology that you can work out while trying to improve at a Kaggle competition goes a really long way 🙂 \n\nMeaning, I feel if you found your way to Kaggle and this competition, you probably are in a great spot to grow your practical set of abilities by working on this problem 🙂",
      "votes": null
    },
    {
      "id": "2033005",
      "postDate": "11/17/2022 01:23:36",
      "content": "<p><a href=\"https://www.kaggle.com/simonveitner\" target=\"_blank\">@simonveitner</a> Did you ever figure it out on getting Transformers4Rec to work? I was trying to do the same thing with trying the example notebook from the merlin github repo but unfortunately, I've been running into issues after installing nvtabular with the following error <code>ModuleNotFoundError: No module named 'merlin.dag.executors'</code>. Were you also running into this error?</p>\n<p>Update: I was able to successfully install the packages using the following pip install commands that resolved the error i was getting. FYI, kaggle already has cudf install when using a GPU accelerator.</p>\n<p>!pip install transformers4rec[pytorch,nvtabular]<br>\n!pip install -U nvtabular==1.3.3</p>",
      "rawMarkdown": "simonveitner Did you ever figure it out on getting Transformers4Rec to work? I was trying to do the same thing with trying the example notebook from the merlin github repo but unfortunately, I've been running into issues after installing nvtabular with the following error `ModuleNotFoundError: No module named 'merlin.dag.executors'`. Were you also running into this error?\n\nUpdate: I was able to successfully install the packages using the following pip install commands that resolved the error i was getting. FYI, kaggle already has cudf install when using a GPU accelerator.\n\n!pip install transformers4rec[pytorch,nvtabular]\n!pip install -U nvtabular==1.3.3",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2015715,
      "author_name": "simonveitner",
      "author_url": "",
      "post_date": "11/03/2022 13:47:05",
      "content": "<p>Could you post a notebook on how to get Transformers4Rec running on kaggle?<br>\nI wanted to try out the following notebook:<br>\n<a href=\"https://github.com/NVIDIA-Merlin/Transformers4Rec/blob/main/examples/end-to-end-session-based/01-ETL-with-NVTabular.ipynb\" target=\"_blank\">https://github.com/NVIDIA-Merlin/Transformers4Rec/blob/main/examples/end-to-end-session-based/01-ETL-with-NVTabular.ipynb</a><br>\nbut had problems…</p>",
      "votes": null,
      "replies": [
        {
          "id": 2015722,
          "author_name": "radek1",
          "author_url": "",
          "post_date": "11/03/2022 13:52:47",
          "content": "<p>Will look into that! 🙌</p>\n<p>The transformers will soon be coming to Merlin Models (thous will be HuggingFace transformers implemented in Tensorflow) so maybe that will make it a bit easier to run here on Kaggle, we will see 🙂</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 2033005,
          "author_name": "dillonquan",
          "author_url": "",
          "post_date": "11/17/2022 01:23:36",
          "content": "<p><a href=\"https://www.kaggle.com/simonveitner\" target=\"_blank\">@simonveitner</a> Did you ever figure it out on getting Transformers4Rec to work? I was trying to do the same thing with trying the example notebook from the merlin github repo but unfortunately, I've been running into issues after installing nvtabular with the following error <code>ModuleNotFoundError: No module named 'merlin.dag.executors'</code>. Were you also running into this error?</p>\n<p>Update: I was able to successfully install the packages using the following pip install commands that resolved the error i was getting. FYI, kaggle already has cudf install when using a GPU accelerator.</p>\n<p>!pip install transformers4rec[pytorch,nvtabular]<br>\n!pip install -U nvtabular==1.3.3</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2016161,
      "author_name": "ollibolli",
      "author_url": "",
      "post_date": "11/03/2022 20:21:46",
      "content": "<p>Thanks a lot for sharing, coincidentally I've also been reading up on it recently because I am interested in trying to apply it on a different domain. Excited for this comp too!</p>",
      "votes": null,
      "replies": [
        {
          "id": 2016249,
          "author_name": "radek1",
          "author_url": "",
          "post_date": "11/03/2022 21:32:25",
          "content": "<p>Thanks, <a href=\"https://www.kaggle.com/ollibolli\" target=\"_blank\">@ollibolli</a>! Really glad you are finding this useful 🙏 BTW what domain are you thinking of applying this to? Just curious, don't need to answer if you don't want to share 🙂</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 2017263,
          "author_name": "ollibolli",
          "author_url": "",
          "post_date": "11/04/2022 16:46:42",
          "content": "<p>Thanks Radek, <br>\nno worries, I apologize I was being cryptic I didn't intend to be. One of the things I love most about Kaggle is that everyone is sharing, like you!<br>\nMight sound silly but I've been getting into soccer analytics and I thought about applying it to player scouting and to extension maybe even selection with some tweaks. I am quite the newbie so it might lead nowhere haha.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 2017463,
          "author_name": "ianfit",
          "author_url": "",
          "post_date": "11/04/2022 20:30:40",
          "content": "<p>Have you seen the Bundesliga Dataset yet? Its quite interesting. I watched a live coding on prev. competition about this, maybe you want to check it out. <a href=\"https://youtu.be/fNHW7zvi_F8\" target=\"_blank\">https://youtu.be/fNHW7zvi_F8</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 2017778,
          "author_name": "radek1",
          "author_url": "",
          "post_date": "11/05/2022 06:30:32",
          "content": "<p>Thank you for your reply, <a href=\"https://www.kaggle.com/oliver\" target=\"_blank\">@oliver</a>! Sounds like an exciting project! 🙂</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 2018044,
          "author_name": "ollibolli",
          "author_url": "",
          "post_date": "11/05/2022 10:56:56",
          "content": "<p>Hi Ian, cheers, yes, I followed this quite closely, there is a group at my school that actually has a few research projects on this. It's a very exciting time :-).</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2019123,
      "author_name": "priyanshuprajapatis",
      "author_url": "",
      "post_date": "11/06/2022 11:05:16",
      "content": "<p>Thanks, even I couldn't find out anything interesting. Will definitely have a look.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2020487,
      "author_name": "ridwanultanvir",
      "author_url": "",
      "post_date": "11/07/2022 14:22:23",
      "content": "<p><a href=\"https://github.com/hidasib/GRU4Rec\" target=\"_blank\">https://github.com/hidasib/GRU4Rec</a> <br>\nimplementation of the algorithm in \"Session-based Recommendations with Recurrent Neural Networks\" paper.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2020557,
      "author_name": "shreyamishra0307",
      "author_url": "",
      "post_date": "11/07/2022 15:31:27",
      "content": "<p>Gratitude for sharing!</p>",
      "votes": null,
      "replies": [
        {
          "id": 2020801,
          "author_name": "radek1",
          "author_url": "",
          "post_date": "11/07/2022 19:33:44",
          "content": "<p>Really glad you are finding this useful! 🙂</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2023136,
      "author_name": "danielliao",
      "author_url": "",
      "post_date": "11/09/2022 14:58:19",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/radek1\" target=\"_blank\">@radek1</a> Thank you so much for the resources to help beginners to get started. </p>\n<p>I have finished watching and <a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/365039\" target=\"_blank\">note taking</a> the getting started videos on recsys from Andrew Ng (which are relatively easy to follow) and I am half way through those of Xavier Amatriain (two more hours to go for lecture 3+4) which I found many concepts are not very familiar and definitely need to rewatch a few more times.</p>\n<p>I wonder how much should I understand Xavier's lectures? I know you would approve that the best way to help understanding is to experiment with the models he talked about, but there are no accompany notebooks or codes available from the lectures right? </p>\n<p>I wonder have you found or implemented those the models or codes when you study the videos? or do you know where I can find them? or  do actually I need to study them now for learning recsys in order to prepare me for participating in OTTO competition? Or maybe focusing on experimenting with your notebooks is already enough for beginners?</p>\n<p>Thanks!</p>",
      "votes": null,
      "replies": [
        {
          "id": 2023699,
          "author_name": "danielliao",
          "author_url": "",
          "post_date": "11/10/2022 01:00:01",
          "content": "<p>I just found a summary of Xavier's lecture 1 + 2 by <a href=\"https://www.kaggle.com/radek1\" target=\"_blank\">@radek1</a> <a href=\"https://twitter.com/radekosmulski/status/1565716248083566592\" target=\"_blank\">here</a>! Thanks Radek</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 2023718,
          "author_name": "radek1",
          "author_url": "",
          "post_date": "11/10/2022 01:25:12",
          "content": "<p>Xavier does a great job in his lectures of explaining the reasoning behind recommender systems, I think that is where the value lies. Very few people do that and nearly no one has experience comparable to his 🙂 As in, the thinking behind what you want and why you want to make the decisions that you do.</p>\n<p>But on the technical side, I think you will find there is a great disconnect between the ideas and what you can implement in code. Everything is very dataset-specific. In some way, this is another skillset that you have to build up, this ability to transfer those ideas to the problem you are working on. Even one or two ideas can go a long way.</p>\n<p>Here, for instance, this is a recommender system competition, but the curve ball is the high cardinality of action ids. How do you deal with this? There has been some really good discussion on this here: <a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/364722\" target=\"_blank\">🐘 the elephant in the room -- high cardinality of targets and what to do about this</a>.</p>\n<p>Or even more concretely, how do you work with the amount of data that you have available in this competition? This is on a completely different scale than most of the \"toy\" ML problems.</p>\n<p>So I think chipping away at this competition and learning from the conversations and the kernels that people post is a great way of getting hands-on experience 🙂 It really does go a long way in actual business context.</p>\n<p>Though some of the challenges that you might have to solve in the real world will be different, the problem-solving methodology that you can work out while trying to improve at a Kaggle competition goes a really long way 🙂 </p>\n<p>Meaning, I feel if you found your way to Kaggle and this competition, you probably are in a great spot to grow your practical set of abilities by working on this problem 🙂 </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2015354": "Hey!\n\nI am super excited for this competition 😊 Have been perusing the world of recommender systems for the last 6 months or so, will interesting to put all that to a test 🙂 \n\nIn terms of good learning resources to get started with recommender systems, there are not that many, unfortunately. \n \nI would highly recommend [this lecture](https://youtu.be/bLhq63ygoU8) by a former colleague of mine, Xavier Amatriain. It provides the best introduction to thinking about recommendations, what recommender systems are, what can be achieved, that I have ever come across.\n\nThis particular problem that we will be working on in this competition are session based recommendations. Essentially, from a timeseries of actions taken by a user in a single session (single visit to a website), we want to predict what is the likely next action a user might take (and most importantly, on what item the action is likely to be taken!)\n\nSession recommendations are a very hot topic. They pull in a lot of exciting concepts together such as serving real time predictions (versus predictions calculated offline, usually in batches). And their sequential nature lends themselves well to modelling with RNNs or more recently, Transformers!\n\nThere is a lot of complexity and interesting ideas to all of this 😊 I am hoping we will have a chance to explore many of them in this competition together!\n\nIf you'd like to read more about using Transformers for session based recommendations, here is [a very nice blog post](https://medium.com/nvidia-merlin/transformers4rec-4523cc7d8fa8) from my colleagues.\n\nSo looking forward to jumping into this competition 🥳\n\nHappy Kaggling!\n\n### Other resources you might find useful:\n\n* [💡 [2 methods] How-to ensemble predictions 🏅🏅🏅](https://www.kaggle.com/code/radek1/2-methods-how-to-ensemble-predictions)\n* [local validation tracks public LB perfecty -- here is the setup](https://www.kaggle.com/competitions/otto-recommender-system/discussion/364991)\n* [💡 For my friends from Twitter and LinkedIn -- here is how to dive into this competition 🐳](https://www.kaggle.com/competitions/otto-recommender-system/discussion/368560)\n* [Full dataset processed to CSV/parquet files with optimized memory footprint](https://www.kaggle.com/competitions/otto-recommender-system/discussion/363843)\n* [co-visitation matrix - simplified, imprvd logic 🔥](https://www.kaggle.com/code/radek1/co-visitation-matrix-simplified-imprvd-logic)\n* [💡 Word2Vec How-to [training and submission]🚀🚀🚀](https://www.kaggle.com/code/radek1/word2vec-how-to-training-and-submission)",
    "2015715": "Could you post a notebook on how to get Transformers4Rec running on kaggle?\nI wanted to try out the following notebook:\nhttps://github.com/NVIDIA-Merlin/Transformers4Rec/blob/main/examples/end-to-end-session-based/01-ETL-with-NVTabular.ipynb\nbut had problems...",
    "2015722": "Will look into that! 🙌\n\nThe transformers will soon be coming to Merlin Models (thous will be HuggingFace transformers implemented in Tensorflow) so maybe that will make it a bit easier to run here on Kaggle, we will see 🙂",
    "2016161": "Thanks a lot for sharing, coincidentally I've also been reading up on it recently because I am interested in trying to apply it on a different domain. Excited for this comp too!",
    "2016249": "Thanks, @ollibolli! Really glad you are finding this useful 🙏 BTW what domain are you thinking of applying this to? Just curious, don't need to answer if you don't want to share 🙂",
    "2017263": "Thanks Radek, \nno worries, I apologize I was being cryptic I didn't intend to be. One of the things I love most about Kaggle is that everyone is sharing, like you!\nMight sound silly but I've been getting into soccer analytics and I thought about applying it to player scouting and to extension maybe even selection with some tweaks. I am quite the newbie so it might lead nowhere haha.",
    "2017463": "Have you seen the Bundesliga Dataset yet? Its quite interesting. I watched a live coding on prev. competition about this, maybe you want to check it out. https://youtu.be/fNHW7zvi_F8",
    "2017778": "Thank you for your reply, @oliver! Sounds like an exciting project! 🙂",
    "2018044": "Hi Ian, cheers, yes, I followed this quite closely, there is a group at my school that actually has a few research projects on this. It's a very exciting time :-).",
    "2019123": "Thanks, even I couldn't find out anything interesting. Will definitely have a look.",
    "2020487": "https://github.com/hidasib/GRU4Rec \nimplementation of the algorithm in \"Session-based Recommendations with Recurrent Neural Networks\" paper.",
    "2020557": "Gratitude for sharing!",
    "2020801": "Really glad you are finding this useful! 🙂",
    "2023136": "Hi @radek1 Thank you so much for the resources to help beginners to get started. \n\nI have finished watching and [note taking](https://www.kaggle.com/competitions/otto-recommender-system/discussion/365039) the getting started videos on recsys from Andrew Ng (which are relatively easy to follow) and I am half way through those of Xavier Amatriain (two more hours to go for lecture 3+4) which I found many concepts are not very familiar and definitely need to rewatch a few more times.\n\nI wonder how much should I understand Xavier's lectures? I know you would approve that the best way to help understanding is to experiment with the models he talked about, but there are no accompany notebooks or codes available from the lectures right? \n\nI wonder have you found or implemented those the models or codes when you study the videos? or do you know where I can find them? or  do actually I need to study them now for learning recsys in order to prepare me for participating in OTTO competition? Or maybe focusing on experimenting with your notebooks is already enough for beginners?\n\nThanks!",
    "2023699": "I just found a summary of Xavier's lecture 1 + 2 by @radek1 [here](https://twitter.com/radekosmulski/status/1565716248083566592)! Thanks Radek",
    "2023718": "Xavier does a great job in his lectures of explaining the reasoning behind recommender systems, I think that is where the value lies. Very few people do that and nearly no one has experience comparable to his 🙂 As in, the thinking behind what you want and why you want to make the decisions that you do.\n\nBut on the technical side, I think you will find there is a great disconnect between the ideas and what you can implement in code. Everything is very dataset-specific. In some way, this is another skillset that you have to build up, this ability to transfer those ideas to the problem you are working on. Even one or two ideas can go a long way.\n\nHere, for instance, this is a recommender system competition, but the curve ball is the high cardinality of action ids. How do you deal with this? There has been some really good discussion on this here: [🐘 the elephant in the room -- high cardinality of targets and what to do about this](https://www.kaggle.com/competitions/otto-recommender-system/discussion/364722).\n\nOr even more concretely, how do you work with the amount of data that you have available in this competition? This is on a completely different scale than most of the \"toy\" ML problems.\n\nSo I think chipping away at this competition and learning from the conversations and the kernels that people post is a great way of getting hands-on experience 🙂 It really does go a long way in actual business context.\n\nThough some of the challenges that you might have to solve in the real world will be different, the problem-solving methodology that you can work out while trying to improve at a Kaggle competition goes a really long way 🙂 \n\nMeaning, I feel if you found your way to Kaggle and this competition, you probably are in a great spot to grow your practical set of abilities by working on this problem 🙂",
    "2033005": "simonveitner Did you ever figure it out on getting Transformers4Rec to work? I was trying to do the same thing with trying the example notebook from the merlin github repo but unfortunately, I've been running into issues after installing nvtabular with the following error `ModuleNotFoundError: No module named 'merlin.dag.executors'`. Were you also running into this error?\n\nUpdate: I was able to successfully install the packages using the following pip install commands that resolved the error i was getting. FYI, kaggle already has cudf install when using a GPU accelerator.\n\n!pip install transformers4rec[pytorch,nvtabular]\n!pip install -U nvtabular==1.3.3"
  },
  "source": "meta"
}