{
  "id": 370950,
  "title": "💡 Implicit vs Explicit data -- a useful distinction for recommender systems",
  "url": "/competitions/otto-recommender-system/discussion/370950",
  "author_name": "",
  "post_date": "2022-12-07T09:33:38.860485200Z",
  "votes": 11,
  "comment_count": 3,
  "views": 0,
  "content": "<p>Hey!</p>\n<p>I just wanted to share something that has been quite confusing to me when I started to deep dive into RecSys.</p>\n<p>And this is important -- once you understand this distinction it is much easier to google for relevant blog posts/papers!</p>\n<p>So here is the distinction:</p>\n<ul>\n<li>explicit data -- data that contains label information coming directly from the user (score on a questionnaire, rating, etc)</li>\n<li>implicit data -- data where the label has been derived from user actions</li>\n</ul>\n<p>The type of data we have in this competition is of the <strong>implicit</strong> type. The labels have been generated from the user activity stream!</p>\n<p>There are many issues with explicit data:</p>\n<ul>\n<li>people give aspirational answers (for instance, they might rank a documentary highly just because they heard it is acclaimed and so they feel like they should rank it highly, it makes them look good, but when going to the cinema they will opt to watch something else that they would have given a lower score to!)</li>\n<li>explicit data is quite noisy -- a 7 to me might be something else than a 7 is to you</li>\n<li>explicit data is much harder and more expensive to get</li>\n</ul>\n<p>Implicit data doesn't suffer from the above issues. Also, we get millions of rows like in this competition exactly because we can automatically generate the labels from the data stream!</p>\n<p>Anyhow, hope you have found this useful and that it will help you come up with an even stronger solution to this challenge! 🙂 I know these two terms have eluded me for quite a while and it took me quite a while to actually come across and explanation of them that would click for me 🙂</p>\n<p>Hope this can help you as you research the methods you can use to tackle this challenge 🙏</p>",
  "messages": [
    {
      "id": "2057656",
      "postDate": "12/07/2022 09:33:38",
      "content": "<p>Hey!</p>\n<p>I just wanted to share something that has been quite confusing to me when I started to deep dive into RecSys.</p>\n<p>And this is important -- once you understand this distinction it is much easier to google for relevant blog posts/papers!</p>\n<p>So here is the distinction:</p>\n<ul>\n<li>explicit data -- data that contains label information coming directly from the user (score on a questionnaire, rating, etc)</li>\n<li>implicit data -- data where the label has been derived from user actions</li>\n</ul>\n<p>The type of data we have in this competition is of the <strong>implicit</strong> type. The labels have been generated from the user activity stream!</p>\n<p>There are many issues with explicit data:</p>\n<ul>\n<li>people give aspirational answers (for instance, they might rank a documentary highly just because they heard it is acclaimed and so they feel like they should rank it highly, it makes them look good, but when going to the cinema they will opt to watch something else that they would have given a lower score to!)</li>\n<li>explicit data is quite noisy -- a 7 to me might be something else than a 7 is to you</li>\n<li>explicit data is much harder and more expensive to get</li>\n</ul>\n<p>Implicit data doesn't suffer from the above issues. Also, we get millions of rows like in this competition exactly because we can automatically generate the labels from the data stream!</p>\n<p>Anyhow, hope you have found this useful and that it will help you come up with an even stronger solution to this challenge! 🙂 I know these two terms have eluded me for quite a while and it took me quite a while to actually come across and explanation of them that would click for me 🙂</p>\n<p>Hope this can help you as you research the methods you can use to tackle this challenge 🙏</p>",
      "rawMarkdown": "Hey!\n\nI just wanted to share something that has been quite confusing to me when I started to deep dive into RecSys.\n\nAnd this is important -- once you understand this distinction it is much easier to google for relevant blog posts/papers!\n\nSo here is the distinction:\n\n* explicit data -- data that contains label information coming directly from the user (score on a questionnaire, rating, etc)\n* implicit data -- data where the label has been derived from user actions\n\nThe type of data we have in this competition is of the **implicit** type. The labels have been generated from the user activity stream!\n\nThere are many issues with explicit data:\n\n* people give aspirational answers (for instance, they might rank a documentary highly just because they heard it is acclaimed and so they feel like they should rank it highly, it makes them look good, but when going to the cinema they will opt to watch something else that they would have given a lower score to!)\n* explicit data is quite noisy -- a 7 to me might be something else than a 7 is to you\n* explicit data is much harder and more expensive to get\n\nImplicit data doesn't suffer from the above issues. Also, we get millions of rows like in this competition exactly because we can automatically generate the labels from the data stream!\n\nAnyhow, hope you have found this useful and that it will help you come up with an even stronger solution to this challenge! 🙂 I know these two terms have eluded me for quite a while and it took me quite a while to actually come across and explanation of them that would click for me 🙂\n\nHope this can help you as you research the methods you can use to tackle this challenge 🙏",
      "votes": null
    },
    {
      "id": "2063669",
      "postDate": "12/13/2022 07:21:18",
      "content": "<p>Good point. </p>\n<p>I remember reading somewhere about the big difference between movies people \"like\" on Netflix, e.g. Citizen Kane, and what they actually watch, e.g. Psycho Killer IV.</p>",
      "rawMarkdown": "Good point. \n\nI remember reading somewhere about the big difference between movies people \"like\" on Netflix, e.g. Citizen Kane, and what they actually watch, e.g. Psycho Killer IV.",
      "votes": null
    },
    {
      "id": "2067065",
      "postDate": "12/16/2022 11:02:31",
      "content": "<p>There is another thought on implicit data that could be interesting if you think about negative sampling:<br>\nExplicit data like movie ratings have three states:</p>\n<ol>\n<li>a good rating score indicates a positive relation between the user and the object<br>\n (e.g., a movie rated with five stars is liked by the person who gave the rating)</li>\n<li>a bad rating score indicates a negative relation between the user and the object <br>\n (e.g., a movie rated with one  star is disliked by the person who gave the rating)</li>\n<li>no rating means the relation between the user and the object is unknown <br>\n (e.g., the person has not seen this movie yet)</li>\n</ol>\n<p>In implicit, behavioral data state 2) and 3) are indistinguishable. A click of the user indicates an affinity to the object clicked. An object the user did not click on could be</p>\n<ul>\n<li>not interesting to him at all</li>\n<li>just less interesting than the object shown next to it </li>\n<li>never been seen by the user in the first place</li>\n</ul>\n<p>It could be interesting to consider of objects of all three cases have the same value as negative sample and if not how to increase the usage of beneficial objects.</p>",
      "rawMarkdown": "There is another thought on implicit data that could be interesting if you think about negative sampling:\nExplicit data like movie ratings have three states:\n1. a good rating score indicates a positive relation between the user and the object\n     (e.g., a movie rated with five stars is liked by the person who gave the rating)\n2. a bad rating score indicates a negative relation between the user and the object \n     (e.g., a movie rated with one  star is disliked by the person who gave the rating)\n3. no rating means the relation between the user and the object is unknown \n     (e.g., the person has not seen this movie yet)\n\nIn implicit, behavioral data state 2) and 3) are indistinguishable. A click of the user indicates an affinity to the object clicked. An object the user did not click on could be\n- not interesting to him at all\n- just less interesting than the object shown next to it \n- never been seen by the user in the first place\n\nIt could be interesting to consider of objects of all three cases have the same value as negative sample and if not how to increase the usage of beneficial objects.",
      "votes": null
    },
    {
      "id": "2079146",
      "postDate": "12/29/2022 03:48:42",
      "content": "<p>yup <a href=\"https://www.kaggle.com/nnjjpp\" target=\"_blank\">@nnjjpp</a>!</p>\n<p>There are likely two main components to us liking \"Citizen Kane\" vs \"Psycho Killer IV\" </p>\n<ul>\n<li>We might believe \"cultured\" people like \"Citizen Kane\" and would like to come across as one such person</li>\n<li>We might believe \"Citizen Kane\", being so legendary, is a good movie. In this case, we are being swayed in our judgment by \"social proof\" (by other people speaking highly of \"Citizen Kane\")</li>\n</ul>\n<p>Guess it all goes back to the old adage, \"do not judge a book by the cover\" or people by what they say but rather by what they do! 🙂</p>\n<p>Thanks for your kind comment and interesting example, <a href=\"https://www.kaggle.com/nnjjpp\" target=\"_blank\">@nnjjpp</a>! </p>",
      "rawMarkdown": "yup @nnjjpp!\n\nThere are likely two main components to us liking \"Citizen Kane\" vs \"Psycho Killer IV\" \n\n* We might believe \"cultured\" people like \"Citizen Kane\" and would like to come across as one such person\n* We might believe \"Citizen Kane\", being so legendary, is a good movie. In this case, we are being swayed in our judgment by \"social proof\" (by other people speaking highly of \"Citizen Kane\")\n\nGuess it all goes back to the old adage, \"do not judge a book by the cover\" or people by what they say but rather by what they do! 🙂\n\nThanks for your kind comment and interesting example, @nnjjpp!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2063669,
      "author_name": "nnjjpp",
      "author_url": "",
      "post_date": "12/13/2022 07:21:18",
      "content": "<p>Good point. </p>\n<p>I remember reading somewhere about the big difference between movies people \"like\" on Netflix, e.g. Citizen Kane, and what they actually watch, e.g. Psycho Killer IV.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2079146,
          "author_name": "radek1",
          "author_url": "",
          "post_date": "12/29/2022 03:48:42",
          "content": "<p>yup <a href=\"https://www.kaggle.com/nnjjpp\" target=\"_blank\">@nnjjpp</a>!</p>\n<p>There are likely two main components to us liking \"Citizen Kane\" vs \"Psycho Killer IV\" </p>\n<ul>\n<li>We might believe \"cultured\" people like \"Citizen Kane\" and would like to come across as one such person</li>\n<li>We might believe \"Citizen Kane\", being so legendary, is a good movie. In this case, we are being swayed in our judgment by \"social proof\" (by other people speaking highly of \"Citizen Kane\")</li>\n</ul>\n<p>Guess it all goes back to the old adage, \"do not judge a book by the cover\" or people by what they say but rather by what they do! 🙂</p>\n<p>Thanks for your kind comment and interesting example, <a href=\"https://www.kaggle.com/nnjjpp\" target=\"_blank\">@nnjjpp</a>! </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2067065,
      "author_name": "andreaswand",
      "author_url": "",
      "post_date": "12/16/2022 11:02:31",
      "content": "<p>There is another thought on implicit data that could be interesting if you think about negative sampling:<br>\nExplicit data like movie ratings have three states:</p>\n<ol>\n<li>a good rating score indicates a positive relation between the user and the object<br>\n (e.g., a movie rated with five stars is liked by the person who gave the rating)</li>\n<li>a bad rating score indicates a negative relation between the user and the object <br>\n (e.g., a movie rated with one  star is disliked by the person who gave the rating)</li>\n<li>no rating means the relation between the user and the object is unknown <br>\n (e.g., the person has not seen this movie yet)</li>\n</ol>\n<p>In implicit, behavioral data state 2) and 3) are indistinguishable. A click of the user indicates an affinity to the object clicked. An object the user did not click on could be</p>\n<ul>\n<li>not interesting to him at all</li>\n<li>just less interesting than the object shown next to it </li>\n<li>never been seen by the user in the first place</li>\n</ul>\n<p>It could be interesting to consider of objects of all three cases have the same value as negative sample and if not how to increase the usage of beneficial objects.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2057656": "Hey!\n\nI just wanted to share something that has been quite confusing to me when I started to deep dive into RecSys.\n\nAnd this is important -- once you understand this distinction it is much easier to google for relevant blog posts/papers!\n\nSo here is the distinction:\n\n* explicit data -- data that contains label information coming directly from the user (score on a questionnaire, rating, etc)\n* implicit data -- data where the label has been derived from user actions\n\nThe type of data we have in this competition is of the **implicit** type. The labels have been generated from the user activity stream!\n\nThere are many issues with explicit data:\n\n* people give aspirational answers (for instance, they might rank a documentary highly just because they heard it is acclaimed and so they feel like they should rank it highly, it makes them look good, but when going to the cinema they will opt to watch something else that they would have given a lower score to!)\n* explicit data is quite noisy -- a 7 to me might be something else than a 7 is to you\n* explicit data is much harder and more expensive to get\n\nImplicit data doesn't suffer from the above issues. Also, we get millions of rows like in this competition exactly because we can automatically generate the labels from the data stream!\n\nAnyhow, hope you have found this useful and that it will help you come up with an even stronger solution to this challenge! 🙂 I know these two terms have eluded me for quite a while and it took me quite a while to actually come across and explanation of them that would click for me 🙂\n\nHope this can help you as you research the methods you can use to tackle this challenge 🙏",
    "2063669": "Good point. \n\nI remember reading somewhere about the big difference between movies people \"like\" on Netflix, e.g. Citizen Kane, and what they actually watch, e.g. Psycho Killer IV.",
    "2067065": "There is another thought on implicit data that could be interesting if you think about negative sampling:\nExplicit data like movie ratings have three states:\n1. a good rating score indicates a positive relation between the user and the object\n     (e.g., a movie rated with five stars is liked by the person who gave the rating)\n2. a bad rating score indicates a negative relation between the user and the object \n     (e.g., a movie rated with one  star is disliked by the person who gave the rating)\n3. no rating means the relation between the user and the object is unknown \n     (e.g., the person has not seen this movie yet)\n\nIn implicit, behavioral data state 2) and 3) are indistinguishable. A click of the user indicates an affinity to the object clicked. An object the user did not click on could be\n- not interesting to him at all\n- just less interesting than the object shown next to it \n- never been seen by the user in the first place\n\nIt could be interesting to consider of objects of all three cases have the same value as negative sample and if not how to increase the usage of beneficial objects.",
    "2079146": "yup @nnjjpp!\n\nThere are likely two main components to us liking \"Citizen Kane\" vs \"Psycho Killer IV\" \n\n* We might believe \"cultured\" people like \"Citizen Kane\" and would like to come across as one such person\n* We might believe \"Citizen Kane\", being so legendary, is a good movie. In this case, we are being swayed in our judgment by \"social proof\" (by other people speaking highly of \"Citizen Kane\")\n\nGuess it all goes back to the old adage, \"do not judge a book by the cover\" or people by what they say but rather by what they do! 🙂\n\nThanks for your kind comment and interesting example, @nnjjpp!"
  },
  "source": "meta"
}