{
  "id": 367495,
  "title": "The Cold Start Problem",
  "url": "/competitions/otto-recommender-system/discussion/367495",
  "author_name": "Ravi Shah",
  "post_date": "2022-11-20T22:50:20.820000",
  "votes": 16,
  "comment_count": 0,
  "views": 0,
  "content": "<h1>What is the Cold Start Problem</h1>\n<p>When creating recommendation systems, the \"Cold Start\" problem occurs when there is not much data about certain users or products. This is usually because these are new users or products</p>\n<p>There are 2 main types of the Cold Start Problem:</p>\n<ul>\n<li>User (or in our case session) cold start - there is very little data about a session (they haven't clicked on many previous items)</li>\n<li>Item cold start - there is very little information/data about a given product</li>\n</ul>\n<p>If you've looked into the test set, you will notice that a user based cold start problem is definitely present.</p>\n<ul>\n<li>For reference the mean number of events per train sessions is ~16.799. In comparison, the mean number of events per test session was only ~4.144.</li>\n</ul>\n<p>Unfortunately, many solutions to this challenge revolve around changing the data collection process by asking new users questions (representative based). However, our data is already collected, so here are some ideas on addressing this challenge given what we have.</p>\n<h1>Popular Products</h1>\n<p>One idea for addressing this issue is determining which products are generally most popular. On its own, this wouldn't be super effective because you would be recommending everyone the same items. However, combining this approach after a content based filtering or clustering algorithm could work.</p>\n<p>Here's the idea:</p>\n<ol>\n<li>Create clusters of products from the train data</li>\n<li>Take the product(s) that have been looked at and determine which cluster they are a part of</li>\n<li>Recommend the other most products from that cluster</li>\n</ol>\n<h1>Co-Visitation Matrices and Matrix Factorization</h1>\n<p>I won't go into too much detail about this approach here because there are a lot of resources about this already. In fact, most high scoring notebooks are currently using this approach, and I have a <a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/365589\" target=\"_blank\">full discussion</a> dedicated to this topic. This isn't super specific to the cold start problem; however, these approaches do tend to handle this challenge well. In general, look into content based filtering and collaborative based filter.</p>\n<h1>Deep Learning Ideas</h1>\n<p>Do note that these deep learning approaches can be very hard and timely to train (especially given the size of our dataset). Nonetheless, here are two approaches.</p>\n<p><strong>DropoutNet</strong><br>\nThe idea behind this approach is that you simply have a neural network that outputs predictions. However, we can make the network more robust by dropping events. Note that we are dropping features and not nodes in this approach. As a result, we hope that the network is more generalizable to sessions with less events.</p>\n<p><strong>Session-based RNN</strong><br>\nThe idea here is that we create a Recurrent Neural Network in which the current state is the previous event, and the output is the next event. Due to the nature of RNNs, the preceding layer’s hidden state becomes the input for the next layer. In theory, this should work well when we have limited data.</p>\n<h1>References</h1>\n<p><a href=\"https://www.kaggle.com/code/cdeotte/time-series-eda-users-and-real-sessions/notebook\" target=\"_blank\">useful time series EDA by Chris - see observations section</a><br>\n<a href=\"https://kojinoshiba.com/recsys-cold-start/\" target=\"_blank\">recsys cold start article</a><br>\n<a href=\"https://analyticsindiamag.com/cold-start-problem-in-recommender-systems-and-its-mitigation-techniques/\" target=\"_blank\">another cold start article</a></p>",
  "messages": [
    {
      "id": 2037747,
      "postDate": "2022-11-20T22:50:20.820Z",
      "content": "<h1>What is the Cold Start Problem</h1>\n<p>When creating recommendation systems, the \"Cold Start\" problem occurs when there is not much data about certain users or products. This is usually because these are new users or products</p>\n<p>There are 2 main types of the Cold Start Problem:</p>\n<ul>\n<li>User (or in our case session) cold start - there is very little data about a session (they haven't clicked on many previous items)</li>\n<li>Item cold start - there is very little information/data about a given product</li>\n</ul>\n<p>If you've looked into the test set, you will notice that a user based cold start problem is definitely present.</p>\n<ul>\n<li>For reference the mean number of events per train sessions is ~16.799. In comparison, the mean number of events per test session was only ~4.144.</li>\n</ul>\n<p>Unfortunately, many solutions to this challenge revolve around changing the data collection process by asking new users questions (representative based). However, our data is already collected, so here are some ideas on addressing this challenge given what we have.</p>\n<h1>Popular Products</h1>\n<p>One idea for addressing this issue is determining which products are generally most popular. On its own, this wouldn't be super effective because you would be recommending everyone the same items. However, combining this approach after a content based filtering or clustering algorithm could work.</p>\n<p>Here's the idea:</p>\n<ol>\n<li>Create clusters of products from the train data</li>\n<li>Take the product(s) that have been looked at and determine which cluster they are a part of</li>\n<li>Recommend the other most products from that cluster</li>\n</ol>\n<h1>Co-Visitation Matrices and Matrix Factorization</h1>\n<p>I won't go into too much detail about this approach here because there are a lot of resources about this already. In fact, most high scoring notebooks are currently using this approach, and I have a <a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/365589\" target=\"_blank\">full discussion</a> dedicated to this topic. This isn't super specific to the cold start problem; however, these approaches do tend to handle this challenge well. In general, look into content based filtering and collaborative based filter.</p>\n<h1>Deep Learning Ideas</h1>\n<p>Do note that these deep learning approaches can be very hard and timely to train (especially given the size of our dataset). Nonetheless, here are two approaches.</p>\n<p><strong>DropoutNet</strong><br>\nThe idea behind this approach is that you simply have a neural network that outputs predictions. However, we can make the network more robust by dropping events. Note that we are dropping features and not nodes in this approach. As a result, we hope that the network is more generalizable to sessions with less events.</p>\n<p><strong>Session-based RNN</strong><br>\nThe idea here is that we create a Recurrent Neural Network in which the current state is the previous event, and the output is the next event. Due to the nature of RNNs, the preceding layer’s hidden state becomes the input for the next layer. In theory, this should work well when we have limited data.</p>\n<h1>References</h1>\n<p><a href=\"https://www.kaggle.com/code/cdeotte/time-series-eda-users-and-real-sessions/notebook\" target=\"_blank\">useful time series EDA by Chris - see observations section</a><br>\n<a href=\"https://kojinoshiba.com/recsys-cold-start/\" target=\"_blank\">recsys cold start article</a><br>\n<a href=\"https://analyticsindiamag.com/cold-start-problem-in-recommender-systems-and-its-mitigation-techniques/\" target=\"_blank\">another cold start article</a></p>",
      "rawMarkdown": "# What is the Cold Start Problem\n\nWhen creating recommendation systems, the \"Cold Start\" problem occurs when there is not much data about certain users or products. This is usually because these are new users or products\n\nThere are 2 main types of the Cold Start Problem:\n- User (or in our case session) cold start - there is very little data about a session (they haven't clicked on many previous items)\n- Item cold start - there is very little information/data about a given product\n\nIf you've looked into the test set, you will notice that a user based cold start problem is definitely present.\n- For reference the mean number of events per train sessions is ~16.799. In comparison, the mean number of events per test session was only ~4.144.\n\nUnfortunately, many solutions to this challenge revolve around changing the data collection process by asking new users questions (representative based). However, our data is already collected, so here are some ideas on addressing this challenge given what we have.\n\n# Popular Products\n\nOne idea for addressing this issue is determining which products are generally most popular. On its own, this wouldn't be super effective because you would be recommending everyone the same items. However, combining this approach after a content based filtering or clustering algorithm could work.\n\nHere's the idea:\n1. Create clusters of products from the train data\n2. Take the product(s) that have been looked at and determine which cluster they are a part of\n3. Recommend the other most products from that cluster\n\n# Co-Visitation Matrices and Matrix Factorization\n\nI won't go into too much detail about this approach here because there are a lot of resources about this already. In fact, most high scoring notebooks are currently using this approach, and I have a [full discussion](https://www.kaggle.com/competitions/otto-recommender-system/discussion/365589) dedicated to this topic. This isn't super specific to the cold start problem; however, these approaches do tend to handle this challenge well. In general, look into content based filtering and collaborative based filter.\n\n# Deep Learning Ideas\n\nDo note that these deep learning approaches can be very hard and timely to train (especially given the size of our dataset). Nonetheless, here are two approaches.\n\n**DropoutNet**\nThe idea behind this approach is that you simply have a neural network that outputs predictions. However, we can make the network more robust by dropping events. Note that we are dropping features and not nodes in this approach. As a result, we hope that the network is more generalizable to sessions with less events.\n\n**Session-based RNN**\nThe idea here is that we create a Recurrent Neural Network in which the current state is the previous event, and the output is the next event. Due to the nature of RNNs, the preceding layer’s hidden state becomes the input for the next layer. In theory, this should work well when we have limited data.\n\n# References\n[useful time series EDA by Chris - see observations section](https://www.kaggle.com/code/cdeotte/time-series-eda-users-and-real-sessions/notebook)\n[recsys cold start article](https://kojinoshiba.com/recsys-cold-start/)\n[another cold start article](https://analyticsindiamag.com/cold-start-problem-in-recommender-systems-and-its-mitigation-techniques/)",
      "votes": 16
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "2037747": "# What is the Cold Start Problem\n\nWhen creating recommendation systems, the \"Cold Start\" problem occurs when there is not much data about certain users or products. This is usually because these are new users or products\n\nThere are 2 main types of the Cold Start Problem:\n- User (or in our case session) cold start - there is very little data about a session (they haven't clicked on many previous items)\n- Item cold start - there is very little information/data about a given product\n\nIf you've looked into the test set, you will notice that a user based cold start problem is definitely present.\n- For reference the mean number of events per train sessions is ~16.799. In comparison, the mean number of events per test session was only ~4.144.\n\nUnfortunately, many solutions to this challenge revolve around changing the data collection process by asking new users questions (representative based). However, our data is already collected, so here are some ideas on addressing this challenge given what we have.\n\n# Popular Products\n\nOne idea for addressing this issue is determining which products are generally most popular. On its own, this wouldn't be super effective because you would be recommending everyone the same items. However, combining this approach after a content based filtering or clustering algorithm could work.\n\nHere's the idea:\n1. Create clusters of products from the train data\n2. Take the product(s) that have been looked at and determine which cluster they are a part of\n3. Recommend the other most products from that cluster\n\n# Co-Visitation Matrices and Matrix Factorization\n\nI won't go into too much detail about this approach here because there are a lot of resources about this already. In fact, most high scoring notebooks are currently using this approach, and I have a [full discussion](https://www.kaggle.com/competitions/otto-recommender-system/discussion/365589) dedicated to this topic. This isn't super specific to the cold start problem; however, these approaches do tend to handle this challenge well. In general, look into content based filtering and collaborative based filter.\n\n# Deep Learning Ideas\n\nDo note that these deep learning approaches can be very hard and timely to train (especially given the size of our dataset). Nonetheless, here are two approaches.\n\n**DropoutNet**\nThe idea behind this approach is that you simply have a neural network that outputs predictions. However, we can make the network more robust by dropping events. Note that we are dropping features and not nodes in this approach. As a result, we hope that the network is more generalizable to sessions with less events.\n\n**Session-based RNN**\nThe idea here is that we create a Recurrent Neural Network in which the current state is the previous event, and the output is the next event. Due to the nature of RNNs, the preceding layer’s hidden state becomes the input for the next layer. In theory, this should work well when we have limited data.\n\n# References\n[useful time series EDA by Chris - see observations section](https://www.kaggle.com/code/cdeotte/time-series-eda-users-and-real-sessions/notebook)\n[recsys cold start article](https://kojinoshiba.com/recsys-cold-start/)\n[another cold start article](https://analyticsindiamag.com/cold-start-problem-in-recommender-systems-and-its-mitigation-techniques/)"
  }
}