{
  "id": 307153,
  "title": "Recommender systems - Evaluation hell",
  "url": "/competitions/h-and-m-personalized-fashion-recommendations/discussion/307153",
  "author_name": "TweakIT",
  "post_date": "2022-02-12T21:44:57.382000",
  "votes": 38,
  "comment_count": 0,
  "views": 0,
  "content": "<p>Hey!<br>\nVery excited to finally see a Recommender systems competition. I hope some creative solutions will give birth to new models. Keep in mind that the last big Recsys competition gave us Matrix Factorization :D<br>\nI wanted to take the opportunity and tickle your minds about how to properly evaluate Recommender systems, both in training and testing.</p>\n<p>A well-known issue in recommender systems is the way the data is collected. In a perfect supervised learning setting, we make sure that the data is balanced, labels are correct, points are iid, etc… In Recsys, data is generated by users and this throws most traditional Machine Learning assumptions out of the window. For instance,  a user cannot interact with an item without seeing it first. therefore, the collected data heavily depends on the previous exposure of the user (this exposure can come from a previous recommender system, the popularity of the item, a friend recommendation, another platform, an ad…) <br>\nWhy is this important? This exposure bias creates a Missing Not At Random problem that biases the estimation of the loss functions, evaluation functions, etc…<br>\nWhat I mean by biasing an estimation is that you can no longer trust that the loss value you calculated is unbiased and hence it may give you the wrong estimation of the true performance of the model.<br>\nTo understand more about how this exposure may become a problem I invite you to read this paper:<br>\n<a href=\"https://arxiv.org/abs/1602.05352\" target=\"_blank\">https://arxiv.org/abs/1602.05352</a> and other papers that cite it.</p>\n<p>In testing, it is even harder. How do you know that you recommended a negative item without showing it to the user first and getting their impression? That is why in industry, online A/B testing is used extensively to gauge the performance of a new model. A/B testing is however expensive. So we try to come up with other ways of offline evaluation. These offline evaluations can be very tricky as we don't really know if the user has seen and did not like the item or the user simply was not exposed to the item. Many papers discuss this issue and come up with more rigorous evaluation techniques that can mitigate this bias:<br>\n<a href=\"https://www.kdd.org/kdd2020/accepted-papers/view/on-sampled-metrics-for-item-recommendation\" target=\"_blank\">https://www.kdd.org/kdd2020/accepted-papers/view/on-sampled-metrics-for-item-recommendation</a> a great paper explaining the effect of negative sampling on different evaluation metrics</p>\n<p>A nice paper discussing the different offline evaluation methods in Recsys<br>\n<a href=\"http://adrem.uantwerpen.be/bibrem/pubs/JeunenRecSys19_DoctoralSymposium.pdf\" target=\"_blank\">http://adrem.uantwerpen.be/bibrem/pubs/JeunenRecSys19_DoctoralSymposium.pdf</a> </p>\n<p>Very excited to see the validation strategies that people are going to use in this competition as it is going to be a major deciding factor.<br>\nGoog luck everyone</p>",
  "messages": [
    {
      "id": 1687451,
      "postDate": "2022-02-12T21:44:57.383Z",
      "content": "<p>Hey!<br>\nVery excited to finally see a Recommender systems competition. I hope some creative solutions will give birth to new models. Keep in mind that the last big Recsys competition gave us Matrix Factorization :D<br>\nI wanted to take the opportunity and tickle your minds about how to properly evaluate Recommender systems, both in training and testing.</p>\n<p>A well-known issue in recommender systems is the way the data is collected. In a perfect supervised learning setting, we make sure that the data is balanced, labels are correct, points are iid, etc… In Recsys, data is generated by users and this throws most traditional Machine Learning assumptions out of the window. For instance,  a user cannot interact with an item without seeing it first. therefore, the collected data heavily depends on the previous exposure of the user (this exposure can come from a previous recommender system, the popularity of the item, a friend recommendation, another platform, an ad…) <br>\nWhy is this important? This exposure bias creates a Missing Not At Random problem that biases the estimation of the loss functions, evaluation functions, etc…<br>\nWhat I mean by biasing an estimation is that you can no longer trust that the loss value you calculated is unbiased and hence it may give you the wrong estimation of the true performance of the model.<br>\nTo understand more about how this exposure may become a problem I invite you to read this paper:<br>\n<a href=\"https://arxiv.org/abs/1602.05352\" target=\"_blank\">https://arxiv.org/abs/1602.05352</a> and other papers that cite it.</p>\n<p>In testing, it is even harder. How do you know that you recommended a negative item without showing it to the user first and getting their impression? That is why in industry, online A/B testing is used extensively to gauge the performance of a new model. A/B testing is however expensive. So we try to come up with other ways of offline evaluation. These offline evaluations can be very tricky as we don't really know if the user has seen and did not like the item or the user simply was not exposed to the item. Many papers discuss this issue and come up with more rigorous evaluation techniques that can mitigate this bias:<br>\n<a href=\"https://www.kdd.org/kdd2020/accepted-papers/view/on-sampled-metrics-for-item-recommendation\" target=\"_blank\">https://www.kdd.org/kdd2020/accepted-papers/view/on-sampled-metrics-for-item-recommendation</a> a great paper explaining the effect of negative sampling on different evaluation metrics</p>\n<p>A nice paper discussing the different offline evaluation methods in Recsys<br>\n<a href=\"http://adrem.uantwerpen.be/bibrem/pubs/JeunenRecSys19_DoctoralSymposium.pdf\" target=\"_blank\">http://adrem.uantwerpen.be/bibrem/pubs/JeunenRecSys19_DoctoralSymposium.pdf</a> </p>\n<p>Very excited to see the validation strategies that people are going to use in this competition as it is going to be a major deciding factor.<br>\nGoog luck everyone</p>",
      "rawMarkdown": "Hey!\nVery excited to finally see a Recommender systems competition. I hope some creative solutions will give birth to new models. Keep in mind that the last big Recsys competition gave us Matrix Factorization :D\nI wanted to take the opportunity and tickle your minds about how to properly evaluate Recommender systems, both in training and testing.\n\nA well-known issue in recommender systems is the way the data is collected. In a perfect supervised learning setting, we make sure that the data is balanced, labels are correct, points are iid, etc... In Recsys, data is generated by users and this throws most traditional Machine Learning assumptions out of the window. For instance,  a user cannot interact with an item without seeing it first. therefore, the collected data heavily depends on the previous exposure of the user (this exposure can come from a previous recommender system, the popularity of the item, a friend recommendation, another platform, an ad...) \nWhy is this important? This exposure bias creates a Missing Not At Random problem that biases the estimation of the loss functions, evaluation functions, etc...\nWhat I mean by biasing an estimation is that you can no longer trust that the loss value you calculated is unbiased and hence it may give you the wrong estimation of the true performance of the model.\nTo understand more about how this exposure may become a problem I invite you to read this paper:\nhttps://arxiv.org/abs/1602.05352 and other papers that cite it.\n\nIn testing, it is even harder. How do you know that you recommended a negative item without showing it to the user first and getting their impression? That is why in industry, online A/B testing is used extensively to gauge the performance of a new model. A/B testing is however expensive. So we try to come up with other ways of offline evaluation. These offline evaluations can be very tricky as we don't really know if the user has seen and did not like the item or the user simply was not exposed to the item. Many papers discuss this issue and come up with more rigorous evaluation techniques that can mitigate this bias:\nhttps://www.kdd.org/kdd2020/accepted-papers/view/on-sampled-metrics-for-item-recommendation a great paper explaining the effect of negative sampling on different evaluation metrics\n\nA nice paper discussing the different offline evaluation methods in Recsys\nhttp://adrem.uantwerpen.be/bibrem/pubs/JeunenRecSys19_DoctoralSymposium.pdf \n\nVery excited to see the validation strategies that people are going to use in this competition as it is going to be a major deciding factor.\nGoog luck everyone\n",
      "votes": 38
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "1687451": "Hey!\nVery excited to finally see a Recommender systems competition. I hope some creative solutions will give birth to new models. Keep in mind that the last big Recsys competition gave us Matrix Factorization :D\nI wanted to take the opportunity and tickle your minds about how to properly evaluate Recommender systems, both in training and testing.\n\nA well-known issue in recommender systems is the way the data is collected. In a perfect supervised learning setting, we make sure that the data is balanced, labels are correct, points are iid, etc... In Recsys, data is generated by users and this throws most traditional Machine Learning assumptions out of the window. For instance,  a user cannot interact with an item without seeing it first. therefore, the collected data heavily depends on the previous exposure of the user (this exposure can come from a previous recommender system, the popularity of the item, a friend recommendation, another platform, an ad...) \nWhy is this important? This exposure bias creates a Missing Not At Random problem that biases the estimation of the loss functions, evaluation functions, etc...\nWhat I mean by biasing an estimation is that you can no longer trust that the loss value you calculated is unbiased and hence it may give you the wrong estimation of the true performance of the model.\nTo understand more about how this exposure may become a problem I invite you to read this paper:\nhttps://arxiv.org/abs/1602.05352 and other papers that cite it.\n\nIn testing, it is even harder. How do you know that you recommended a negative item without showing it to the user first and getting their impression? That is why in industry, online A/B testing is used extensively to gauge the performance of a new model. A/B testing is however expensive. So we try to come up with other ways of offline evaluation. These offline evaluations can be very tricky as we don't really know if the user has seen and did not like the item or the user simply was not exposed to the item. Many papers discuss this issue and come up with more rigorous evaluation techniques that can mitigate this bias:\nhttps://www.kdd.org/kdd2020/accepted-papers/view/on-sampled-metrics-for-item-recommendation a great paper explaining the effect of negative sampling on different evaluation metrics\n\nA nice paper discussing the different offline evaluation methods in Recsys\nhttp://adrem.uantwerpen.be/bibrem/pubs/JeunenRecSys19_DoctoralSymposium.pdf \n\nVery excited to see the validation strategies that people are going to use in this competition as it is going to be a major deciding factor.\nGoog luck everyone\n"
  }
}