{
  "id": 368541,
  "title": "What are we predicting?",
  "url": "/competitions/otto-recommender-system/discussion/368541",
  "author_name": "",
  "post_date": "2022-11-26T02:26:00.455484600Z",
  "votes": 7,
  "comment_count": 2,
  "views": 0,
  "content": "<p>The competition requires the prediction of user behavior in terms of click, cart, and order of products on a per-session basis.<br>\nOn the other hand, when we consider how the competition data was generated, users are browsing products according to the OTTO recommendation algorithm.<br>\nIn other words, it would be the OTTO recommendation algorithm and user behavior prediction, not user behavior prediction.<br>\n(*Except when a user transitions from the search box or category field to a page on the screen.</p>\n<p>For this reason, I think it is important to look at the actual OTTO site and see what kind of recommendations are being made.</p>\n<p>For example, if a predictions includes a product that does not follow the OTTO recommendation algorithm, the score may be improved by removing the product in post-processing.<br>\nThe OTTO site is <a href=\"https://www.otto.de/\" target=\"_blank\">here</a>. train,test data is from the 1st of August until the 4th of September, so the OTTO recommendation algorithm may have changed slightly between November 1, 2022 and January 31, 2023, the period when the competition was held, but I think it may be helpful.</p>\n<p>I would appreciate it if you could comment on any errors in my comments.</p>\n<p>PS<br>\nThe above may not be very effective, as comments from the host below revealed that in OTTO, page views from recommendations account for 20% of all product page views.</p>",
  "messages": [
    {
      "id": "2043843",
      "postDate": "11/26/2022 02:26:00",
      "content": "<p>The competition requires the prediction of user behavior in terms of click, cart, and order of products on a per-session basis.<br>\nOn the other hand, when we consider how the competition data was generated, users are browsing products according to the OTTO recommendation algorithm.<br>\nIn other words, it would be the OTTO recommendation algorithm and user behavior prediction, not user behavior prediction.<br>\n(*Except when a user transitions from the search box or category field to a page on the screen.</p>\n<p>For this reason, I think it is important to look at the actual OTTO site and see what kind of recommendations are being made.</p>\n<p>For example, if a predictions includes a product that does not follow the OTTO recommendation algorithm, the score may be improved by removing the product in post-processing.<br>\nThe OTTO site is <a href=\"https://www.otto.de/\" target=\"_blank\">here</a>. train,test data is from the 1st of August until the 4th of September, so the OTTO recommendation algorithm may have changed slightly between November 1, 2022 and January 31, 2023, the period when the competition was held, but I think it may be helpful.</p>\n<p>I would appreciate it if you could comment on any errors in my comments.</p>\n<p>PS<br>\nThe above may not be very effective, as comments from the host below revealed that in OTTO, page views from recommendations account for 20% of all product page views.</p>",
      "rawMarkdown": "The competition requires the prediction of user behavior in terms of click, cart, and order of products on a per-session basis.\nOn the other hand, when we consider how the competition data was generated, users are browsing products according to the OTTO recommendation algorithm.\nIn other words, it would be the OTTO recommendation algorithm and user behavior prediction, not user behavior prediction.\n(*Except when a user transitions from the search box or category field to a page on the screen.\n\nFor this reason, I think it is important to look at the actual OTTO site and see what kind of recommendations are being made.\n\nFor example, if a predictions includes a product that does not follow the OTTO recommendation algorithm, the score may be improved by removing the product in post-processing.\nThe OTTO site is [here](https://www.otto.de/). train,test data is ~~from July-September 2022~~from the 1st of August until the 4th of September, so the OTTO recommendation algorithm may have changed slightly between November 1, 2022 and January 31, 2023, the period when the competition was held, but I think it may be helpful.\n\nI would appreciate it if you could comment on any errors in my comments.\n\n\n\nPS\nThe above may not be very effective, as comments from the host below revealed that in OTTO, page views from recommendations account for 20% of all product page views.",
      "votes": null
    },
    {
      "id": "2049775",
      "postDate": "11/30/2022 10:31:32",
      "content": "<p>Hello <a href=\"https://www.kaggle.com/mujrush\" target=\"_blank\">@mujrush</a>, thanks for your interesting thoughts! <br>\nWe noticed three things while reading your comment:</p>\n<ol>\n<li>You mentioned that the train and test set range from July to September 2022, while actually, it ranges from the 1st of August until the 4th of September.</li>\n<li>In our shop, a large variety of recommendation algorithms exist, each with its own purpose, serving an even larger variety of recommendation features. While you are right that these features impact user behavior to a certain degree, only about 20% of all product detail page views are caused by recommendations.</li>\n<li>We assume you are planning to crawl our shop's recommendation algorithms to create an additional dataset for training your models. How do you plan to use this supplementary dataset since we anonymized all IDs for the public dataset?</li>\n</ol>",
      "rawMarkdown": "Hello @mujrush, thanks for your interesting thoughts! \nWe noticed three things while reading your comment:\n\n1. You mentioned that the train and test set range from July to September 2022, while actually, it ranges from the 1st of August until the 4th of September.\n2. In our shop, a large variety of recommendation algorithms exist, each with its own purpose, serving an even larger variety of recommendation features. While you are right that these features impact user behavior to a certain degree, only about 20% of all product detail page views are caused by recommendations.\n3. We assume you are planning to crawl our shop's recommendation algorithms to create an additional dataset for training your models. How do you plan to use this supplementary dataset since we anonymized all IDs for the public dataset?",
      "votes": null
    },
    {
      "id": "2050110",
      "postDate": "11/30/2022 14:30:23",
      "content": "<p><a href=\"https://www.kaggle.com/pnormann\" target=\"_blank\">@pnormann</a> <br>\nThank you for your valuable comments!</p>\n<ol>\n<li><p>Sorry, I thought the first date of the competition data was 7/31, but I missed the discussion <a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/367991\" target=\"_blank\">here</a>. So I actually need to add 2 hours and this results in data from August 1 to September 4.</p></li>\n<li><p>Thanks for the valuable info! <br>\nBy the way, do you know the percentage of page views by recommendation in a single session?<br>\nIn the competition data, I thought the recommendation rate for all product detail page views would be 20% since most of the sessions have a small number of events, but for single sessions with many events, I thought the recommendation rate would be higher.</p></li>\n<li><p>I assume that the OTTO site is checked visually, not by web scraping, etc.<br>\nFor example, suppose we check OTTO's recommendation algorithm (popularity, new release, campaign, etc.) and save a group of predictions based on that recommendation algorithm. Then, if there are no candidates based on the newly created rules at all in the saved prediction group, we thought that the candidates based on the new rules are unlikely to be selected by users, and I can increase the score by removing them in post-processing. This would be possible even if they were anonymized.<br>\nHowever, I thought the above would work if the majority of page views are due to recommendations, but not if the majority of page views are due to user searches, categories, etc.</p></li>\n</ol>",
      "rawMarkdown": "pnormann \nThank you for your valuable comments!\n\n1. Sorry, I thought the first date of the competition data was 7/31, but I missed the discussion [here](https://www.kaggle.com/competitions/otto-recommender-system/discussion/367991). So I actually need to add 2 hours and this results in data from August 1 to September 4.\n\n2. Thanks for the valuable info! \nBy the way, do you know the percentage of page views by recommendation in a single session?\nIn the competition data, I thought the recommendation rate for all product detail page views would be 20% since most of the sessions have a small number of events, but for single sessions with many events, I thought the recommendation rate would be higher.\n\n3. I assume that the OTTO site is checked visually, not by web scraping, etc.\nFor example, suppose we check OTTO's recommendation algorithm (popularity, new release, campaign, etc.) and save a group of predictions based on that recommendation algorithm. Then, if there are no candidates based on the newly created rules at all in the saved prediction group, we thought that the candidates based on the new rules are unlikely to be selected by users, and I can increase the score by removing them in post-processing. This would be possible even if they were anonymized.\nHowever, I thought the above would work if the majority of page views are due to recommendations, but not if the majority of page views are due to user searches, categories, etc.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2049775,
      "author_name": "pnormann",
      "author_url": "",
      "post_date": "11/30/2022 10:31:32",
      "content": "<p>Hello <a href=\"https://www.kaggle.com/mujrush\" target=\"_blank\">@mujrush</a>, thanks for your interesting thoughts! <br>\nWe noticed three things while reading your comment:</p>\n<ol>\n<li>You mentioned that the train and test set range from July to September 2022, while actually, it ranges from the 1st of August until the 4th of September.</li>\n<li>In our shop, a large variety of recommendation algorithms exist, each with its own purpose, serving an even larger variety of recommendation features. While you are right that these features impact user behavior to a certain degree, only about 20% of all product detail page views are caused by recommendations.</li>\n<li>We assume you are planning to crawl our shop's recommendation algorithms to create an additional dataset for training your models. How do you plan to use this supplementary dataset since we anonymized all IDs for the public dataset?</li>\n</ol>",
      "votes": null,
      "replies": [
        {
          "id": 2050110,
          "author_name": "mujrush",
          "author_url": "",
          "post_date": "11/30/2022 14:30:23",
          "content": "<p><a href=\"https://www.kaggle.com/pnormann\" target=\"_blank\">@pnormann</a> <br>\nThank you for your valuable comments!</p>\n<ol>\n<li><p>Sorry, I thought the first date of the competition data was 7/31, but I missed the discussion <a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/367991\" target=\"_blank\">here</a>. So I actually need to add 2 hours and this results in data from August 1 to September 4.</p></li>\n<li><p>Thanks for the valuable info! <br>\nBy the way, do you know the percentage of page views by recommendation in a single session?<br>\nIn the competition data, I thought the recommendation rate for all product detail page views would be 20% since most of the sessions have a small number of events, but for single sessions with many events, I thought the recommendation rate would be higher.</p></li>\n<li><p>I assume that the OTTO site is checked visually, not by web scraping, etc.<br>\nFor example, suppose we check OTTO's recommendation algorithm (popularity, new release, campaign, etc.) and save a group of predictions based on that recommendation algorithm. Then, if there are no candidates based on the newly created rules at all in the saved prediction group, we thought that the candidates based on the new rules are unlikely to be selected by users, and I can increase the score by removing them in post-processing. This would be possible even if they were anonymized.<br>\nHowever, I thought the above would work if the majority of page views are due to recommendations, but not if the majority of page views are due to user searches, categories, etc.</p></li>\n</ol>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2043843": "The competition requires the prediction of user behavior in terms of click, cart, and order of products on a per-session basis.\nOn the other hand, when we consider how the competition data was generated, users are browsing products according to the OTTO recommendation algorithm.\nIn other words, it would be the OTTO recommendation algorithm and user behavior prediction, not user behavior prediction.\n(*Except when a user transitions from the search box or category field to a page on the screen.\n\nFor this reason, I think it is important to look at the actual OTTO site and see what kind of recommendations are being made.\n\nFor example, if a predictions includes a product that does not follow the OTTO recommendation algorithm, the score may be improved by removing the product in post-processing.\nThe OTTO site is [here](https://www.otto.de/). train,test data is ~~from July-September 2022~~from the 1st of August until the 4th of September, so the OTTO recommendation algorithm may have changed slightly between November 1, 2022 and January 31, 2023, the period when the competition was held, but I think it may be helpful.\n\nI would appreciate it if you could comment on any errors in my comments.\n\n\n\nPS\nThe above may not be very effective, as comments from the host below revealed that in OTTO, page views from recommendations account for 20% of all product page views.",
    "2049775": "Hello @mujrush, thanks for your interesting thoughts! \nWe noticed three things while reading your comment:\n\n1. You mentioned that the train and test set range from July to September 2022, while actually, it ranges from the 1st of August until the 4th of September.\n2. In our shop, a large variety of recommendation algorithms exist, each with its own purpose, serving an even larger variety of recommendation features. While you are right that these features impact user behavior to a certain degree, only about 20% of all product detail page views are caused by recommendations.\n3. We assume you are planning to crawl our shop's recommendation algorithms to create an additional dataset for training your models. How do you plan to use this supplementary dataset since we anonymized all IDs for the public dataset?",
    "2050110": "pnormann \nThank you for your valuable comments!\n\n1. Sorry, I thought the first date of the competition data was 7/31, but I missed the discussion [here](https://www.kaggle.com/competitions/otto-recommender-system/discussion/367991). So I actually need to add 2 hours and this results in data from August 1 to September 4.\n\n2. Thanks for the valuable info! \nBy the way, do you know the percentage of page views by recommendation in a single session?\nIn the competition data, I thought the recommendation rate for all product detail page views would be 20% since most of the sessions have a small number of events, but for single sessions with many events, I thought the recommendation rate would be higher.\n\n3. I assume that the OTTO site is checked visually, not by web scraping, etc.\nFor example, suppose we check OTTO's recommendation algorithm (popularity, new release, campaign, etc.) and save a group of predictions based on that recommendation algorithm. Then, if there are no candidates based on the newly created rules at all in the saved prediction group, we thought that the candidates based on the new rules are unlikely to be selected by users, and I can increase the score by removing them in post-processing. This would be possible even if they were anonymized.\nHowever, I thought the above would work if the majority of page views are due to recommendations, but not if the majority of page views are due to user searches, categories, etc."
  },
  "source": "meta"
}