{
  "id": 375224,
  "title": " We're just Rebuilding Otto's Recommendation System?",
  "url": "/competitions/otto-recommender-system/discussion/375224",
  "author_name": "",
  "post_date": "2022-12-31T05:46:30.650892400Z",
  "votes": 14,
  "comment_count": 9,
  "views": 0,
  "content": "<p>I think some users click on the product because it was suggested by Otto instead of searching for them. So our job is just rebuilding the Otto Recommender system ?</p>",
  "messages": [
    {
      "id": "2081373",
      "postDate": "12/31/2022 05:46:30",
      "content": "<p>I think some users click on the product because it was suggested by Otto instead of searching for them. So our job is just rebuilding the Otto Recommender system ?</p>",
      "rawMarkdown": "I think some users click on the product because it was suggested by Otto instead of searching for them. So our job is just rebuilding the Otto Recommender system ?",
      "votes": null
    },
    {
      "id": "2081500",
      "postDate": "12/31/2022 08:43:17",
      "content": "<p>Probably the emphasis has to be on \"some users\" 🙂 We won't know how many of the actions come from recommendations, maybe the fraction is not as high as we think.</p>\n<p>But that is slightly beyond the point. The true value of a competition like this is not only in getting a working solution at the end. Rather, the organizer can draw from the various approaches and use them as a source of inspiration for techniques of working with their data! Instead of replacing their system, they can map out new ways to grow it!</p>\n<p>And yes, you are right, evaluating recommender systems is super tricky. Are we mapping true customer preferences or are we causing change in them through our recommendations?</p>\n<p>But there is really no simple and straightforward answer to this problem! And it is probably beyond the scope of a Kaggle competition to address this (which makes RecSys so fascinating IMO, all these real-world interactions that have an impact on the company's bottom line!) </p>\n<p>Still, having that many worlds-best data scientists as on Kaggle going through your data, have to be super valuable, I would imagine 🙂</p>",
      "rawMarkdown": "Probably the emphasis has to be on \"some users\" 🙂 We won't know how many of the actions come from recommendations, maybe the fraction is not as high as we think.\n\nBut that is slightly beyond the point. The true value of a competition like this is not only in getting a working solution at the end. Rather, the organizer can draw from the various approaches and use them as a source of inspiration for techniques of working with their data! Instead of replacing their system, they can map out new ways to grow it!\n\nAnd yes, you are right, evaluating recommender systems is super tricky. Are we mapping true customer preferences or are we causing change in them through our recommendations?\n\nBut there is really no simple and straightforward answer to this problem! And it is probably beyond the scope of a Kaggle competition to address this (which makes RecSys so fascinating IMO, all these real-world interactions that have an impact on the company's bottom line!) \n\nStill, having that many worlds-best data scientists as on Kaggle going through your data, have to be super valuable, I would imagine 🙂",
      "votes": null
    },
    {
      "id": "2081502",
      "postDate": "12/31/2022 08:47:49",
      "content": "<p>BTW just looked this up, it is a great blog post on some of the intricacies of measuring recommender systems that also mentions your concerns 🙂</p>\n<p><a href=\"https://eugeneyan.com/writing/counterfactual-evaluation/\" target=\"_blank\">Counterfactual Evaluation for Recommendation Systems</a></p>\n<p>A fascinating read by Eugene Yan!</p>",
      "rawMarkdown": "BTW just looked this up, it is a great blog post on some of the intricacies of measuring recommender systems that also mentions your concerns 🙂\n\n[Counterfactual Evaluation for Recommendation Systems](https://eugeneyan.com/writing/counterfactual-evaluation/)\n\nA fascinating read by Eugene Yan!",
      "votes": null
    },
    {
      "id": "2081521",
      "postDate": "12/31/2022 09:28:18",
      "content": "<p>Thank you so much for your information!<br>\nI understand our goal is still to generate value from the dataset, but I think in a contest, if someone knows how the Otto Recommender System approach, they can simulate it again and win. . For example, Otto suggests that by taking the most popular item at a time, simple methods like Rule-based will obviously give good results. This result is simply because it accurately simulates what is happening. However, it is not necessarily the best if implemented in practice. So I think, it's better if we have data that are \"searches\" (by typing on search engine, for example) that actually come from the user, rather than passive suggestions. Here, of course, we don't know how Otto collected the data and how they suggested it.</p>",
      "rawMarkdown": "Thank you so much for your information!\nI understand our goal is still to generate value from the dataset, but I think in a contest, if someone knows how the Otto Recommender System approach, they can simulate it again and win. . For example, Otto suggests that by taking the most popular item at a time, simple methods like Rule-based will obviously give good results. This result is simply because it accurately simulates what is happening. However, it is not necessarily the best if implemented in practice. So I think, it's better if we have data that are \"searches\" (by typing on search engine, for example) that actually come from the user, rather than passive suggestions. Here, of course, we don't know how Otto collected the data and how they suggested it.",
      "votes": null
    },
    {
      "id": "2081696",
      "postDate": "12/31/2022 14:20:57",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/bibanh\" target=\"_blank\">@bibanh</a> <br>\nOnly about 20% of product views in our shop are triggered by recommendations. This applies to the provided dataset, too. These 20% cover all recommendations including complementary products, alternative products and personalized product recommendations.<br>\nThe vast majority of product views are generated by search result pages and product lists. </p>",
      "rawMarkdown": "Hi @bibanh \nOnly about 20% of product views in our shop are triggered by recommendations. This applies to the provided dataset, too. These 20% cover all recommendations including complementary products, alternative products and personalized product recommendations.\nThe vast majority of product views are generated by search result pages and product lists.",
      "votes": null
    },
    {
      "id": "2081734",
      "postDate": "12/31/2022 14:55:03",
      "content": "<p>The host responded with the following response <a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/371678#2062700\" target=\"_blank\">here</a>.</p>\n<blockquote>\n  <p>Hi <a href=\"https://www.kaggle.com/jamnik99\" target=\"_blank\">@jamnik99</a>, although this idea may seem plausible, one should consider that in our case, only about 20% of the product page views are generated by our recommendations, and the majority of the users reach product pages via search results and product lists.</p>\n</blockquote>\n<p>They says that only 20% are generated by Otto recommender. But it seems to me that Otto's recommender would also influence the other 80% too</p>\n<blockquote>\n  <p>and the majority of the users reach product pages via search results and product lists.</p>\n</blockquote>\n<p>I would think <code>product pages via search results</code> and <code>product lists</code> are also a result of some algorithm from Otto too. For example when a user searches, doesn't Otto sort the results by their recommender? And aren't <code>product lists</code> decided by Otto based on popular trends?</p>\n<p>UPDATE: At the same time that I posted this,  <a href=\"https://www.kaggle.com/andreaswand\" target=\"_blank\">@andreaswand</a> reiterated this below.</p>",
      "rawMarkdown": "The host responded with the following response [here][1].\n> Hi @jamnik99, although this idea may seem plausible, one should consider that in our case, only about 20% of the product page views are generated by our recommendations, and the majority of the users reach product pages via search results and product lists.\n\nThey says that only 20% are generated by Otto recommender. But it seems to me that Otto's recommender would also influence the other 80% too\n>and the majority of the users reach product pages via search results and product lists.\n\nI would think `product pages via search results` and `product lists` are also a result of some algorithm from Otto too. For example when a user searches, doesn't Otto sort the results by their recommender? And aren't `product lists` decided by Otto based on popular trends?\n\nUPDATE: At the same time that I posted this,  @andreaswand reiterated this below.\n\n[1]: https://www.kaggle.com/competitions/otto-recommender-system/discussion/371678#2062700",
      "votes": null
    },
    {
      "id": "2081736",
      "postDate": "12/31/2022 14:57:19",
      "content": "<p><a href=\"https://www.kaggle.com/andreaswand\" target=\"_blank\">@andreaswand</a> Are <code>search result pages</code> and <code>product lists</code> also influenced by some model and/or algorithm at Otto?</p>\n<p>For example, if a search results or product list is too long to display, how do you choose which items to display on <code>page 1 of X</code>? And how do you decide the order of display?</p>",
      "rawMarkdown": "andreaswand Are `search result pages` and `product lists` also influenced by some model and/or algorithm at Otto?\n\nFor example, if a search results or product list is too long to display, how do you choose which items to display on `page 1 of X`? And how do you decide the order of display?",
      "votes": null
    },
    {
      "id": "2081771",
      "postDate": "12/31/2022 15:47:15",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> <br>\nfar as I know, the generation of product lists and search results pages is a two-step process where candidates are first selected and then reranked, similar to some approaches discussed in this challenge. The big difference is that the search models work primarily with product attribute data. How well a product has sold recently has an additional impact on the ranking.</p>\n<p>The reranking determines the order of products on these lists. Each page of a list can contain a maximum number of products (e.g. 120), so the product ranked 121 would be the first one on page two.</p>\n<p>In addition, \"sponsored products\" can be displayed, i.e. products that belong to a product list or a search result page  in that they are valid candidates, but which are displayed at the top of the list regardless of their rank because a seller is willing to pay for their promotion. These products are visibly marked as \"sponsored\" (\"gesponsored\" in German).</p>\n<p>I hope this is helpful. If you are interested in more information about the ranking of product lists I need to get in touch with my colleagues. <a href=\"https://www.kaggle.com/pnormann\" target=\"_blank\">@pnormann</a> and I are part of the recommendations team here at OTTO, the search teams belong to a separate unit. As you can imagine, today of all days is not the best time for this consultation, so it might be a few days before I can provide more details.</p>",
      "rawMarkdown": "Hi @cdeotte \nfar as I know, the generation of product lists and search results pages is a two-step process where candidates are first selected and then reranked, similar to some approaches discussed in this challenge. The big difference is that the search models work primarily with product attribute data. How well a product has sold recently has an additional impact on the ranking.\n\nThe reranking determines the order of products on these lists. Each page of a list can contain a maximum number of products (e.g. 120), so the product ranked 121 would be the first one on page two.\n\nIn addition, \"sponsored products\" can be displayed, i.e. products that belong to a product list or a search result page  in that they are valid candidates, but which are displayed at the top of the list regardless of their rank because a seller is willing to pay for their promotion. These products are visibly marked as \"sponsored\" (\"gesponsored\" in German).\n\nI hope this is helpful. If you are interested in more information about the ranking of product lists I need to get in touch with my colleagues. @pnormann and I are part of the recommendations team here at OTTO, the search teams belong to a separate unit. As you can imagine, today of all days is not the best time for this consultation, so it might be a few days before I can provide more details.",
      "votes": null
    },
    {
      "id": "2081907",
      "postDate": "12/31/2022 19:58:24",
      "content": "<p>This is a super interesting discussion 🙂 Thank you so much for sharing these additional thoughts, <a href=\"https://www.kaggle.com/andreaswand\" target=\"_blank\">@andreaswand</a> and <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a>!</p>\n<p>It goes to show how complex and multifaceted RecSys is and how many factors can influence the data and the performance of the system! What a fascinating field! 🙂</p>",
      "rawMarkdown": "This is a super interesting discussion 🙂 Thank you so much for sharing these additional thoughts, @andreaswand and @cdeotte!\n\nIt goes to show how complex and multifaceted RecSys is and how many factors can influence the data and the performance of the system! What a fascinating field! 🙂",
      "votes": null
    },
    {
      "id": "2081910",
      "postDate": "12/31/2022 20:03:15",
      "content": "<p>oh and <a href=\"https://www.kaggle.com/bibanh\" target=\"_blank\">@bibanh</a>, if I may add please, the fact that popularity-based methods provide a good baseline is a feature of many recsys problems/datasets 🙂</p>\n<p>That is also another fun thing about recsys -- you can get to a \"decent\" solution using very simple techniques, but to push the performance forward, you need to do quite a few fancy things 🙂</p>\n<p>Also, to be clear, in this competition popularity-based methods actually do not seem to work to well, hard to say why. This is why the baseline is built around the co-visitation matrix, which is a much more elaborate approach. It might be because of the long-tail of the distribution, that there are just so many less popular items being purchased, or that the data comes from multiple stores. But seems that ranking items by popularity (or using popularity to generate candidates) doesn't really take you that far, at least that has been my experience from what I recall.</p>",
      "rawMarkdown": "oh and @bibanh, if I may add please, the fact that popularity-based methods provide a good baseline is a feature of many recsys problems/datasets 🙂\n\nThat is also another fun thing about recsys -- you can get to a \"decent\" solution using very simple techniques, but to push the performance forward, you need to do quite a few fancy things 🙂\n\nAlso, to be clear, in this competition popularity-based methods actually do not seem to work to well, hard to say why. This is why the baseline is built around the co-visitation matrix, which is a much more elaborate approach. It might be because of the long-tail of the distribution, that there are just so many less popular items being purchased, or that the data comes from multiple stores. But seems that ranking items by popularity (or using popularity to generate candidates) doesn't really take you that far, at least that has been my experience from what I recall.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2081500,
      "author_name": "radek1",
      "author_url": "",
      "post_date": "12/31/2022 08:43:17",
      "content": "<p>Probably the emphasis has to be on \"some users\" 🙂 We won't know how many of the actions come from recommendations, maybe the fraction is not as high as we think.</p>\n<p>But that is slightly beyond the point. The true value of a competition like this is not only in getting a working solution at the end. Rather, the organizer can draw from the various approaches and use them as a source of inspiration for techniques of working with their data! Instead of replacing their system, they can map out new ways to grow it!</p>\n<p>And yes, you are right, evaluating recommender systems is super tricky. Are we mapping true customer preferences or are we causing change in them through our recommendations?</p>\n<p>But there is really no simple and straightforward answer to this problem! And it is probably beyond the scope of a Kaggle competition to address this (which makes RecSys so fascinating IMO, all these real-world interactions that have an impact on the company's bottom line!) </p>\n<p>Still, having that many worlds-best data scientists as on Kaggle going through your data, have to be super valuable, I would imagine 🙂</p>",
      "votes": null,
      "replies": [
        {
          "id": 2081521,
          "author_name": "bibanh",
          "author_url": "",
          "post_date": "12/31/2022 09:28:18",
          "content": "<p>Thank you so much for your information!<br>\nI understand our goal is still to generate value from the dataset, but I think in a contest, if someone knows how the Otto Recommender System approach, they can simulate it again and win. . For example, Otto suggests that by taking the most popular item at a time, simple methods like Rule-based will obviously give good results. This result is simply because it accurately simulates what is happening. However, it is not necessarily the best if implemented in practice. So I think, it's better if we have data that are \"searches\" (by typing on search engine, for example) that actually come from the user, rather than passive suggestions. Here, of course, we don't know how Otto collected the data and how they suggested it.</p>",
          "votes": null,
          "replies": [
            {
              "id": 2081696,
              "author_name": "andreaswand",
              "author_url": "",
              "post_date": "12/31/2022 14:20:57",
              "content": "<p>Hi <a href=\"https://www.kaggle.com/bibanh\" target=\"_blank\">@bibanh</a> <br>\nOnly about 20% of product views in our shop are triggered by recommendations. This applies to the provided dataset, too. These 20% cover all recommendations including complementary products, alternative products and personalized product recommendations.<br>\nThe vast majority of product views are generated by search result pages and product lists. </p>",
              "votes": null,
              "replies": [
                {
                  "id": 2081736,
                  "author_name": "cdeotte",
                  "author_url": "",
                  "post_date": "12/31/2022 14:57:19",
                  "content": "<p><a href=\"https://www.kaggle.com/andreaswand\" target=\"_blank\">@andreaswand</a> Are <code>search result pages</code> and <code>product lists</code> also influenced by some model and/or algorithm at Otto?</p>\n<p>For example, if a search results or product list is too long to display, how do you choose which items to display on <code>page 1 of X</code>? And how do you decide the order of display?</p>",
                  "votes": null,
                  "replies": [
                    {
                      "id": 2081771,
                      "author_name": "andreaswand",
                      "author_url": "",
                      "post_date": "12/31/2022 15:47:15",
                      "content": "<p>Hi <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> <br>\nfar as I know, the generation of product lists and search results pages is a two-step process where candidates are first selected and then reranked, similar to some approaches discussed in this challenge. The big difference is that the search models work primarily with product attribute data. How well a product has sold recently has an additional impact on the ranking.</p>\n<p>The reranking determines the order of products on these lists. Each page of a list can contain a maximum number of products (e.g. 120), so the product ranked 121 would be the first one on page two.</p>\n<p>In addition, \"sponsored products\" can be displayed, i.e. products that belong to a product list or a search result page  in that they are valid candidates, but which are displayed at the top of the list regardless of their rank because a seller is willing to pay for their promotion. These products are visibly marked as \"sponsored\" (\"gesponsored\" in German).</p>\n<p>I hope this is helpful. If you are interested in more information about the ranking of product lists I need to get in touch with my colleagues. <a href=\"https://www.kaggle.com/pnormann\" target=\"_blank\">@pnormann</a> and I are part of the recommendations team here at OTTO, the search teams belong to a separate unit. As you can imagine, today of all days is not the best time for this consultation, so it might be a few days before I can provide more details.</p>",
                      "votes": null,
                      "replies": [
                        {
                          "id": 2081907,
                          "author_name": "radek1",
                          "author_url": "",
                          "post_date": "12/31/2022 19:58:24",
                          "content": "<p>This is a super interesting discussion 🙂 Thank you so much for sharing these additional thoughts, <a href=\"https://www.kaggle.com/andreaswand\" target=\"_blank\">@andreaswand</a> and <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a>!</p>\n<p>It goes to show how complex and multifaceted RecSys is and how many factors can influence the data and the performance of the system! What a fascinating field! 🙂</p>",
                          "votes": null,
                          "replies": [
                            {
                              "id": 2081910,
                              "author_name": "radek1",
                              "author_url": "",
                              "post_date": "12/31/2022 20:03:15",
                              "content": "<p>oh and <a href=\"https://www.kaggle.com/bibanh\" target=\"_blank\">@bibanh</a>, if I may add please, the fact that popularity-based methods provide a good baseline is a feature of many recsys problems/datasets 🙂</p>\n<p>That is also another fun thing about recsys -- you can get to a \"decent\" solution using very simple techniques, but to push the performance forward, you need to do quite a few fancy things 🙂</p>\n<p>Also, to be clear, in this competition popularity-based methods actually do not seem to work to well, hard to say why. This is why the baseline is built around the co-visitation matrix, which is a much more elaborate approach. It might be because of the long-tail of the distribution, that there are just so many less popular items being purchased, or that the data comes from multiple stores. But seems that ranking items by popularity (or using popularity to generate candidates) doesn't really take you that far, at least that has been my experience from what I recall.</p>",
                              "votes": null,
                              "replies": []
                            }
                          ]
                        }
                      ]
                    }
                  ]
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 2081502,
      "author_name": "radek1",
      "author_url": "",
      "post_date": "12/31/2022 08:47:49",
      "content": "<p>BTW just looked this up, it is a great blog post on some of the intricacies of measuring recommender systems that also mentions your concerns 🙂</p>\n<p><a href=\"https://eugeneyan.com/writing/counterfactual-evaluation/\" target=\"_blank\">Counterfactual Evaluation for Recommendation Systems</a></p>\n<p>A fascinating read by Eugene Yan!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2081734,
      "author_name": "cdeotte",
      "author_url": "",
      "post_date": "12/31/2022 14:55:03",
      "content": "<p>The host responded with the following response <a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/371678#2062700\" target=\"_blank\">here</a>.</p>\n<blockquote>\n  <p>Hi <a href=\"https://www.kaggle.com/jamnik99\" target=\"_blank\">@jamnik99</a>, although this idea may seem plausible, one should consider that in our case, only about 20% of the product page views are generated by our recommendations, and the majority of the users reach product pages via search results and product lists.</p>\n</blockquote>\n<p>They says that only 20% are generated by Otto recommender. But it seems to me that Otto's recommender would also influence the other 80% too</p>\n<blockquote>\n  <p>and the majority of the users reach product pages via search results and product lists.</p>\n</blockquote>\n<p>I would think <code>product pages via search results</code> and <code>product lists</code> are also a result of some algorithm from Otto too. For example when a user searches, doesn't Otto sort the results by their recommender? And aren't <code>product lists</code> decided by Otto based on popular trends?</p>\n<p>UPDATE: At the same time that I posted this,  <a href=\"https://www.kaggle.com/andreaswand\" target=\"_blank\">@andreaswand</a> reiterated this below.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2081373": "I think some users click on the product because it was suggested by Otto instead of searching for them. So our job is just rebuilding the Otto Recommender system ?",
    "2081500": "Probably the emphasis has to be on \"some users\" 🙂 We won't know how many of the actions come from recommendations, maybe the fraction is not as high as we think.\n\nBut that is slightly beyond the point. The true value of a competition like this is not only in getting a working solution at the end. Rather, the organizer can draw from the various approaches and use them as a source of inspiration for techniques of working with their data! Instead of replacing their system, they can map out new ways to grow it!\n\nAnd yes, you are right, evaluating recommender systems is super tricky. Are we mapping true customer preferences or are we causing change in them through our recommendations?\n\nBut there is really no simple and straightforward answer to this problem! And it is probably beyond the scope of a Kaggle competition to address this (which makes RecSys so fascinating IMO, all these real-world interactions that have an impact on the company's bottom line!) \n\nStill, having that many worlds-best data scientists as on Kaggle going through your data, have to be super valuable, I would imagine 🙂",
    "2081502": "BTW just looked this up, it is a great blog post on some of the intricacies of measuring recommender systems that also mentions your concerns 🙂\n\n[Counterfactual Evaluation for Recommendation Systems](https://eugeneyan.com/writing/counterfactual-evaluation/)\n\nA fascinating read by Eugene Yan!",
    "2081521": "Thank you so much for your information!\nI understand our goal is still to generate value from the dataset, but I think in a contest, if someone knows how the Otto Recommender System approach, they can simulate it again and win. . For example, Otto suggests that by taking the most popular item at a time, simple methods like Rule-based will obviously give good results. This result is simply because it accurately simulates what is happening. However, it is not necessarily the best if implemented in practice. So I think, it's better if we have data that are \"searches\" (by typing on search engine, for example) that actually come from the user, rather than passive suggestions. Here, of course, we don't know how Otto collected the data and how they suggested it.",
    "2081696": "Hi @bibanh \nOnly about 20% of product views in our shop are triggered by recommendations. This applies to the provided dataset, too. These 20% cover all recommendations including complementary products, alternative products and personalized product recommendations.\nThe vast majority of product views are generated by search result pages and product lists.",
    "2081734": "The host responded with the following response [here][1].\n> Hi @jamnik99, although this idea may seem plausible, one should consider that in our case, only about 20% of the product page views are generated by our recommendations, and the majority of the users reach product pages via search results and product lists.\n\nThey says that only 20% are generated by Otto recommender. But it seems to me that Otto's recommender would also influence the other 80% too\n>and the majority of the users reach product pages via search results and product lists.\n\nI would think `product pages via search results` and `product lists` are also a result of some algorithm from Otto too. For example when a user searches, doesn't Otto sort the results by their recommender? And aren't `product lists` decided by Otto based on popular trends?\n\nUPDATE: At the same time that I posted this,  @andreaswand reiterated this below.\n\n[1]: https://www.kaggle.com/competitions/otto-recommender-system/discussion/371678#2062700",
    "2081736": "andreaswand Are `search result pages` and `product lists` also influenced by some model and/or algorithm at Otto?\n\nFor example, if a search results or product list is too long to display, how do you choose which items to display on `page 1 of X`? And how do you decide the order of display?",
    "2081771": "Hi @cdeotte \nfar as I know, the generation of product lists and search results pages is a two-step process where candidates are first selected and then reranked, similar to some approaches discussed in this challenge. The big difference is that the search models work primarily with product attribute data. How well a product has sold recently has an additional impact on the ranking.\n\nThe reranking determines the order of products on these lists. Each page of a list can contain a maximum number of products (e.g. 120), so the product ranked 121 would be the first one on page two.\n\nIn addition, \"sponsored products\" can be displayed, i.e. products that belong to a product list or a search result page  in that they are valid candidates, but which are displayed at the top of the list regardless of their rank because a seller is willing to pay for their promotion. These products are visibly marked as \"sponsored\" (\"gesponsored\" in German).\n\nI hope this is helpful. If you are interested in more information about the ranking of product lists I need to get in touch with my colleagues. @pnormann and I are part of the recommendations team here at OTTO, the search teams belong to a separate unit. As you can imagine, today of all days is not the best time for this consultation, so it might be a few days before I can provide more details.",
    "2081907": "This is a super interesting discussion 🙂 Thank you so much for sharing these additional thoughts, @andreaswand and @cdeotte!\n\nIt goes to show how complex and multifaceted RecSys is and how many factors can influence the data and the performance of the system! What a fascinating field! 🙂",
    "2081910": "oh and @bibanh, if I may add please, the fact that popularity-based methods provide a good baseline is a feature of many recsys problems/datasets 🙂\n\nThat is also another fun thing about recsys -- you can get to a \"decent\" solution using very simple techniques, but to push the performance forward, you need to do quite a few fancy things 🙂\n\nAlso, to be clear, in this competition popularity-based methods actually do not seem to work to well, hard to say why. This is why the baseline is built around the co-visitation matrix, which is a much more elaborate approach. It might be because of the long-tail of the distribution, that there are just so many less popular items being purchased, or that the data comes from multiple stores. But seems that ranking items by popularity (or using popularity to generate candidates) doesn't really take you that far, at least that has been my experience from what I recall."
  },
  "source": "meta"
}