{
  "id": 364229,
  "title": "Why best model might not win RecSys competition",
  "url": "/competitions/otto-recommender-system/discussion/364229",
  "author_name": "",
  "post_date": "2022-11-05T10:04:15.792765600Z",
  "votes": 20,
  "comment_count": 1,
  "views": 0,
  "content": "<h1>Online vs Offline</h1>\n<p>Here is a \"little\" difference between offline and online metric value in RecSys competitions:<br>\nGround truth labels are derived by organizers from a result of work of <strong>their current model</strong>.</p>\n<p>So if their model has an average performance and your model has an excellent performance your model still might not beat <strong>their</strong> model on <strong>their</strong> ground truth data. </p>\n<p>Just to go a little deeper:<br>\nReal people see in their recommendations only item that were recommended by current model. If you make a new model that is better and measure an offline metric using ground truth labels that were a result of current model then you might have a worse <strong>offline</strong> metric. But if you would perform a proper <a href=\"https://en.wikipedia.org/wiki/A/B_testing\" target=\"_blank\">A/B test</a> your model might actually outperform a current model.</p>\n<p>Hope my explanation makes sense.</p>\n<h1>How this might affect a competition?</h1>\n<p>If you are developing a model that is too much different from what organizers had when they were collecting data for this competition then your metric might degrade. But if you are using something really similar to organizers model then performance might increase.</p>\n<h1>How to deal with it in real life?</h1>\n<p>There are few ways:</p>\n<ol>\n<li>Measure offline metric just for a sanity check, to make sure that model do produce logical recommendations and then perform a proper A/B test.</li>\n<li>Always keep some % of users with very basic recommendations (like top rated, or most popular) and measure model performance on offline metric against their ground true</li>\n</ol>\n<h1>Which model might have a better chance to have a good performance?</h1>\n<p>Here is my prediction: <a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/364228\" target=\"_blank\">Working as Vanga: which model would win this competition?</a></p>",
  "messages": [
    {
      "id": "2018018",
      "postDate": "11/05/2022 10:04:15",
      "content": "<h1>Online vs Offline</h1>\n<p>Here is a \"little\" difference between offline and online metric value in RecSys competitions:<br>\nGround truth labels are derived by organizers from a result of work of <strong>their current model</strong>.</p>\n<p>So if their model has an average performance and your model has an excellent performance your model still might not beat <strong>their</strong> model on <strong>their</strong> ground truth data. </p>\n<p>Just to go a little deeper:<br>\nReal people see in their recommendations only item that were recommended by current model. If you make a new model that is better and measure an offline metric using ground truth labels that were a result of current model then you might have a worse <strong>offline</strong> metric. But if you would perform a proper <a href=\"https://en.wikipedia.org/wiki/A/B_testing\" target=\"_blank\">A/B test</a> your model might actually outperform a current model.</p>\n<p>Hope my explanation makes sense.</p>\n<h1>How this might affect a competition?</h1>\n<p>If you are developing a model that is too much different from what organizers had when they were collecting data for this competition then your metric might degrade. But if you are using something really similar to organizers model then performance might increase.</p>\n<h1>How to deal with it in real life?</h1>\n<p>There are few ways:</p>\n<ol>\n<li>Measure offline metric just for a sanity check, to make sure that model do produce logical recommendations and then perform a proper A/B test.</li>\n<li>Always keep some % of users with very basic recommendations (like top rated, or most popular) and measure model performance on offline metric against their ground true</li>\n</ol>\n<h1>Which model might have a better chance to have a good performance?</h1>\n<p>Here is my prediction: <a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/364228\" target=\"_blank\">Working as Vanga: which model would win this competition?</a></p>",
      "rawMarkdown": "# Online vs Offline\nHere is a \"little\" difference between offline and online metric value in RecSys competitions:\nGround truth labels are derived by organizers from a result of work of **their current model**.\n\nSo if their model has an average performance and your model has an excellent performance your model still might not beat **their** model on **their** ground truth data. \n\nJust to go a little deeper:\nReal people see in their recommendations only item that were recommended by current model. If you make a new model that is better and measure an offline metric using ground truth labels that were a result of current model then you might have a worse **offline** metric. But if you would perform a proper [A/B test](https://en.wikipedia.org/wiki/A/B_testing) your model might actually outperform a current model.\n\nHope my explanation makes sense.\n\n# How this might affect a competition?\nIf you are developing a model that is too much different from what organizers had when they were collecting data for this competition then your metric might degrade. But if you are using something really similar to organizers model then performance might increase.\n\n# How to deal with it in real life?\nThere are few ways:\n1. Measure offline metric just for a sanity check, to make sure that model do produce logical recommendations and then perform a proper A/B test.\n2. Always keep some % of users with very basic recommendations (like top rated, or most popular) and measure model performance on offline metric against their ground true\n\n# Which model might have a better chance to have a good performance?\nHere is my prediction: [Working as Vanga: which model would win this competition?](https://www.kaggle.com/competitions/otto-recommender-system/discussion/364228)",
      "votes": null
    },
    {
      "id": "2018264",
      "postDate": "11/05/2022 14:37:45",
      "content": "<p>I'm not really familiar with this domain but I noticed that feedback loop too. I don't think it is possible to create a very different model with existing labels. All models will be highly correlated with the existing one.</p>",
      "rawMarkdown": "I'm not really familiar with this domain but I noticed that feedback loop too. I don't think it is possible to create a very different model with existing labels. All models will be highly correlated with the existing one.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2018264,
      "author_name": "gunesevitan",
      "author_url": "",
      "post_date": "11/05/2022 14:37:45",
      "content": "<p>I'm not really familiar with this domain but I noticed that feedback loop too. I don't think it is possible to create a very different model with existing labels. All models will be highly correlated with the existing one.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2018018": "# Online vs Offline\nHere is a \"little\" difference between offline and online metric value in RecSys competitions:\nGround truth labels are derived by organizers from a result of work of **their current model**.\n\nSo if their model has an average performance and your model has an excellent performance your model still might not beat **their** model on **their** ground truth data. \n\nJust to go a little deeper:\nReal people see in their recommendations only item that were recommended by current model. If you make a new model that is better and measure an offline metric using ground truth labels that were a result of current model then you might have a worse **offline** metric. But if you would perform a proper [A/B test](https://en.wikipedia.org/wiki/A/B_testing) your model might actually outperform a current model.\n\nHope my explanation makes sense.\n\n# How this might affect a competition?\nIf you are developing a model that is too much different from what organizers had when they were collecting data for this competition then your metric might degrade. But if you are using something really similar to organizers model then performance might increase.\n\n# How to deal with it in real life?\nThere are few ways:\n1. Measure offline metric just for a sanity check, to make sure that model do produce logical recommendations and then perform a proper A/B test.\n2. Always keep some % of users with very basic recommendations (like top rated, or most popular) and measure model performance on offline metric against their ground true\n\n# Which model might have a better chance to have a good performance?\nHere is my prediction: [Working as Vanga: which model would win this competition?](https://www.kaggle.com/competitions/otto-recommender-system/discussion/364228)",
    "2018264": "I'm not really familiar with this domain but I noticed that feedback loop too. I don't think it is possible to create a very different model with existing labels. All models will be highly correlated with the existing one."
  },
  "source": "meta"
}