{
  "id": 373163,
  "title": "Exposure Bias and Its Impact on Recommender Systems",
  "url": "/competitions/otto-recommender-system/discussion/373163",
  "author_name": "The Devastator",
  "post_date": "2022-12-20T00:41:03.983000",
  "votes": 5,
  "comment_count": 0,
  "views": 0,
  "content": "<h3>Exposure Bias and Its Impact on Recommender Systems</h3>\n<p>I found an interesting post <a href=\"https://www.kaggle.com/competitions/h-and-m-personalized-fashion-recommendations/discussion/307153\" target=\"_blank\">here</a> about exposure bias for RecSys messing up with our validation strategies. </p>\n<p>Let me try to summarise it to the best of my understanding since this can be really important for this competition. </p>\n<p><strong>Recommender systems are very sensitive to the way the data is collected.</strong></p>\n<p>This is different from supervised learning where we can control for most of our assumptions about the target variable. (Make sure that the data is balanced, Labels are correct, Samples are iid, etc…)</p>\n<p>For Recsys the data is generated by users and this throws most traditional Machine Learning assumptions out of the window.</p>\n<blockquote>\n  <p>Example: A user cannot interact with an item without seeing it first. -&gt; The collected data heavily depends on the previous exposure of the user (this exposure can come from a previous recommender system, the popularity of the item, a friend recommendation, another platform, or an ad…)</p>\n</blockquote>\n<h4>What is the problem?</h4>\n<p>Exposure bias creates a <strong>Missing Not At Random</strong> problem that biases the estimation of the loss functions, evaluation functions, etc…<br>\nSo we can no longer trust that the loss value you calculated is unbiased and hence it may give us the wrong estimation of the true performance of the model. </p>\n<blockquote>\n  <p>A good paper about the subject can be found <a href=\"https://arxiv.org/abs/1602.05352\" target=\"_blank\">here</a>. (From the original post like above).</p>\n</blockquote>\n<h4>Test Time</h4>\n<p>When speaking about test time - this gets even harder. How do you know that you recommended an item without showing it to the user first and getting their impression?</p>\n<p>Well in practice this can be easily solved:  Online A/B testing is used extensively in the real-world to help with this issue and to gauge the performance of a new model.<br>\nA/B testing is however expensive (And we are also not in the real-world around here). So we try to come up with other ways of offline evaluation.</p>\n<p>These offline evaluations can be very tricky as we don't really know if the user has seen and did not like the item or if the user simply was not exposed to the item.</p>\n<h4>Want To Read More?</h4>\n<p>Many papers discuss this issue and come up with more rigorous evaluation techniques that can mitigate this bias:</p>\n<ul>\n<li><p>A great paper explaining the effect of negative sampling on different evaluation metrics can be found <a href=\"https://www.kdd.org/kdd2020/accepted-papers/view/on-sampled-metrics-for-item-recommendation\" target=\"_blank\">here</a></p></li>\n<li><p>A good nice paper discussing the different offline evaluation methods in Recsys can be found <a href=\"http://adrem.uantwerpen.be/bibrem/pubs/JeunenRecSys19_DoctoralSymposium.pdf\" target=\"_blank\">here</a></p></li>\n</ul>\n<p>Good Luck Everyone1<br>\nDon't Overfit!</p>",
  "messages": [
    {
      "id": 2070394,
      "postDate": "2022-12-20T00:41:03.983Z",
      "content": "<h3>Exposure Bias and Its Impact on Recommender Systems</h3>\n<p>I found an interesting post <a href=\"https://www.kaggle.com/competitions/h-and-m-personalized-fashion-recommendations/discussion/307153\" target=\"_blank\">here</a> about exposure bias for RecSys messing up with our validation strategies. </p>\n<p>Let me try to summarise it to the best of my understanding since this can be really important for this competition. </p>\n<p><strong>Recommender systems are very sensitive to the way the data is collected.</strong></p>\n<p>This is different from supervised learning where we can control for most of our assumptions about the target variable. (Make sure that the data is balanced, Labels are correct, Samples are iid, etc…)</p>\n<p>For Recsys the data is generated by users and this throws most traditional Machine Learning assumptions out of the window.</p>\n<blockquote>\n  <p>Example: A user cannot interact with an item without seeing it first. -&gt; The collected data heavily depends on the previous exposure of the user (this exposure can come from a previous recommender system, the popularity of the item, a friend recommendation, another platform, or an ad…)</p>\n</blockquote>\n<h4>What is the problem?</h4>\n<p>Exposure bias creates a <strong>Missing Not At Random</strong> problem that biases the estimation of the loss functions, evaluation functions, etc…<br>\nSo we can no longer trust that the loss value you calculated is unbiased and hence it may give us the wrong estimation of the true performance of the model. </p>\n<blockquote>\n  <p>A good paper about the subject can be found <a href=\"https://arxiv.org/abs/1602.05352\" target=\"_blank\">here</a>. (From the original post like above).</p>\n</blockquote>\n<h4>Test Time</h4>\n<p>When speaking about test time - this gets even harder. How do you know that you recommended an item without showing it to the user first and getting their impression?</p>\n<p>Well in practice this can be easily solved:  Online A/B testing is used extensively in the real-world to help with this issue and to gauge the performance of a new model.<br>\nA/B testing is however expensive (And we are also not in the real-world around here). So we try to come up with other ways of offline evaluation.</p>\n<p>These offline evaluations can be very tricky as we don't really know if the user has seen and did not like the item or if the user simply was not exposed to the item.</p>\n<h4>Want To Read More?</h4>\n<p>Many papers discuss this issue and come up with more rigorous evaluation techniques that can mitigate this bias:</p>\n<ul>\n<li><p>A great paper explaining the effect of negative sampling on different evaluation metrics can be found <a href=\"https://www.kdd.org/kdd2020/accepted-papers/view/on-sampled-metrics-for-item-recommendation\" target=\"_blank\">here</a></p></li>\n<li><p>A good nice paper discussing the different offline evaluation methods in Recsys can be found <a href=\"http://adrem.uantwerpen.be/bibrem/pubs/JeunenRecSys19_DoctoralSymposium.pdf\" target=\"_blank\">here</a></p></li>\n</ul>\n<p>Good Luck Everyone1<br>\nDon't Overfit!</p>",
      "rawMarkdown": "### Exposure Bias and Its Impact on Recommender Systems\n\nI found an interesting post [here](https://www.kaggle.com/competitions/h-and-m-personalized-fashion-recommendations/discussion/307153) about exposure bias for RecSys messing up with our validation strategies. \n\nLet me try to summarise it to the best of my understanding since this can be really important for this competition. \n\n**Recommender systems are very sensitive to the way the data is collected.**\n\nThis is different from supervised learning where we can control for most of our assumptions about the target variable. (Make sure that the data is balanced, Labels are correct, Samples are iid, etc…)\n\nFor Recsys the data is generated by users and this throws most traditional Machine Learning assumptions out of the window.\n\n> Example: A user cannot interact with an item without seeing it first. -> The collected data heavily depends on the previous exposure of the user (this exposure can come from a previous recommender system, the popularity of the item, a friend recommendation, another platform, or an ad…)\n\n#### What is the problem?\n\nExposure bias creates a **Missing Not At Random** problem that biases the estimation of the loss functions, evaluation functions, etc…\nSo we can no longer trust that the loss value you calculated is unbiased and hence it may give us the wrong estimation of the true performance of the model. \n\n> A good paper about the subject can be found [here](https://arxiv.org/abs/1602.05352). (From the original post like above).\n\n#### Test Time\n\nWhen speaking about test time - this gets even harder. How do you know that you recommended an item without showing it to the user first and getting their impression?\n\n\nWell in practice this can be easily solved:  Online A/B testing is used extensively in the real-world to help with this issue and to gauge the performance of a new model.\nA/B testing is however expensive (And we are also not in the real-world around here). So we try to come up with other ways of offline evaluation.\n\n\n\nThese offline evaluations can be very tricky as we don't really know if the user has seen and did not like the item or if the user simply was not exposed to the item.\n\n#### Want To Read More?\n\nMany papers discuss this issue and come up with more rigorous evaluation techniques that can mitigate this bias:\n\n- A great paper explaining the effect of negative sampling on different evaluation metrics can be found [here](https://www.kdd.org/kdd2020/accepted-papers/view/on-sampled-metrics-for-item-recommendation)\n\n- A good nice paper discussing the different offline evaluation methods in Recsys can be found [here](http://adrem.uantwerpen.be/bibrem/pubs/JeunenRecSys19_DoctoralSymposium.pdf)\n\n\nGood Luck Everyone1\nDon't Overfit!",
      "votes": 5
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "2070394": "### Exposure Bias and Its Impact on Recommender Systems\n\nI found an interesting post [here](https://www.kaggle.com/competitions/h-and-m-personalized-fashion-recommendations/discussion/307153) about exposure bias for RecSys messing up with our validation strategies. \n\nLet me try to summarise it to the best of my understanding since this can be really important for this competition. \n\n**Recommender systems are very sensitive to the way the data is collected.**\n\nThis is different from supervised learning where we can control for most of our assumptions about the target variable. (Make sure that the data is balanced, Labels are correct, Samples are iid, etc…)\n\nFor Recsys the data is generated by users and this throws most traditional Machine Learning assumptions out of the window.\n\n> Example: A user cannot interact with an item without seeing it first. -> The collected data heavily depends on the previous exposure of the user (this exposure can come from a previous recommender system, the popularity of the item, a friend recommendation, another platform, or an ad…)\n\n#### What is the problem?\n\nExposure bias creates a **Missing Not At Random** problem that biases the estimation of the loss functions, evaluation functions, etc…\nSo we can no longer trust that the loss value you calculated is unbiased and hence it may give us the wrong estimation of the true performance of the model. \n\n> A good paper about the subject can be found [here](https://arxiv.org/abs/1602.05352). (From the original post like above).\n\n#### Test Time\n\nWhen speaking about test time - this gets even harder. How do you know that you recommended an item without showing it to the user first and getting their impression?\n\n\nWell in practice this can be easily solved:  Online A/B testing is used extensively in the real-world to help with this issue and to gauge the performance of a new model.\nA/B testing is however expensive (And we are also not in the real-world around here). So we try to come up with other ways of offline evaluation.\n\n\n\nThese offline evaluations can be very tricky as we don't really know if the user has seen and did not like the item or if the user simply was not exposed to the item.\n\n#### Want To Read More?\n\nMany papers discuss this issue and come up with more rigorous evaluation techniques that can mitigate this bias:\n\n- A great paper explaining the effect of negative sampling on different evaluation metrics can be found [here](https://www.kdd.org/kdd2020/accepted-papers/view/on-sampled-metrics-for-item-recommendation)\n\n- A good nice paper discussing the different offline evaluation methods in Recsys can be found [here](http://adrem.uantwerpen.be/bibrem/pubs/JeunenRecSys19_DoctoralSymposium.pdf)\n\n\nGood Luck Everyone1\nDon't Overfit!"
  }
}