{
  "id": 207480,
  "title": "Content_id is not a unique id",
  "url": "/competitions/riiid-test-answer-prediction/discussion/207480",
  "author_name": "",
  "post_date": "2020-12-29T21:02:19.859802300Z",
  "votes": 6,
  "comment_count": 3,
  "views": 0,
  "content": "<p>Just a heads up, in case anyone is relying on content_id to actually uniquely identify a piece of content. </p>\n<p>In reality, content_id is either a question_id of a question or a lecture_id of a lecture. The problem is that <strong>these ids have a large overlap</strong>.</p>\n<p>From a notebook where I checked this:</p>\n<pre><code>repeat_ids = set(lectures_df['lecture_id']).intersection(set(qs_df['question_id']))\nlen(repeat_ids)\n</code></pre>\n<pre><code>158\n</code></pre>\n<p>so <strong>158/418 lectures have ids that overlap seemingly-unrelated question ids</strong></p>\n<p>In particular, I first noticed it when trying to summarize results with content_id=4057. Question with id 4057 has to do with tag 152, whereas lecture with id 4057 has to do with tag 1.</p>\n<p>so <strong>don't group on content_id, it may produce nonsensical results</strong></p>",
  "messages": [
    {
      "id": "1131611",
      "postDate": "12/29/2020 21:02:19",
      "content": "<p>Just a heads up, in case anyone is relying on content_id to actually uniquely identify a piece of content. </p>\n<p>In reality, content_id is either a question_id of a question or a lecture_id of a lecture. The problem is that <strong>these ids have a large overlap</strong>.</p>\n<p>From a notebook where I checked this:</p>\n<pre><code>repeat_ids = set(lectures_df['lecture_id']).intersection(set(qs_df['question_id']))\nlen(repeat_ids)\n</code></pre>\n<pre><code>158\n</code></pre>\n<p>so <strong>158/418 lectures have ids that overlap seemingly-unrelated question ids</strong></p>\n<p>In particular, I first noticed it when trying to summarize results with content_id=4057. Question with id 4057 has to do with tag 152, whereas lecture with id 4057 has to do with tag 1.</p>\n<p>so <strong>don't group on content_id, it may produce nonsensical results</strong></p>",
      "rawMarkdown": "Just a heads up, in case anyone is relying on content_id to actually uniquely identify a piece of content. \n\nIn reality, content_id is either a question_id of a question or a lecture_id of a lecture. The problem is that **these ids have a large overlap**.\n\nFrom a notebook where I checked this:\n```\nrepeat_ids = set(lectures_df['lecture_id']).intersection(set(qs_df['question_id']))\nlen(repeat_ids)\n```\n```\n158\n```\n\nso **158/418 lectures have ids that overlap seemingly-unrelated question ids**\n\nIn particular, I first noticed it when trying to summarize results with content_id=4057. Question with id 4057 has to do with tag 152, whereas lecture with id 4057 has to do with tag 1.\n\nso **don't group on content_id, it may produce nonsensical results**",
      "votes": null
    },
    {
      "id": "1131630",
      "postDate": "12/29/2020 21:27:03",
      "content": "<p>If you filter first  by content_type_id, you can group by content_id</p>",
      "rawMarkdown": "If you filter first  by content_type_id, you can group by content_id",
      "votes": null
    },
    {
      "id": "1131753",
      "postDate": "12/29/2020 23:57:39",
      "content": "<p>True.</p>\n<p>I'm not saying that there are no workarounds, I'm just saying that one has to be careful and make sure apply the workarounds if this is important to them. </p>\n<p>And that this is not obvious from the name (or description) of the field.</p>",
      "rawMarkdown": "True.\n\nI'm not saying that there are no workarounds, I'm just saying that one has to be careful and make sure apply the workarounds if this is important to them. \n\nAnd that this is not obvious from the name (or description) of the field.",
      "votes": null
    },
    {
      "id": "1131759",
      "postDate": "12/30/2020 00:04:04",
      "content": "<p>Aah ok ok, sorry. It's a good point to keep in mind</p>",
      "rawMarkdown": "Aah ok ok, sorry. It's a good point to keep in mind",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1131630,
      "author_name": "maherelouahabi",
      "author_url": "",
      "post_date": "12/29/2020 21:27:03",
      "content": "<p>If you filter first  by content_type_id, you can group by content_id</p>",
      "votes": null,
      "replies": [
        {
          "id": 1131753,
          "author_name": "yanamal",
          "author_url": "",
          "post_date": "12/29/2020 23:57:39",
          "content": "<p>True.</p>\n<p>I'm not saying that there are no workarounds, I'm just saying that one has to be careful and make sure apply the workarounds if this is important to them. </p>\n<p>And that this is not obvious from the name (or description) of the field.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1131759,
          "author_name": "maherelouahabi",
          "author_url": "",
          "post_date": "12/30/2020 00:04:04",
          "content": "<p>Aah ok ok, sorry. It's a good point to keep in mind</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1131611": "Just a heads up, in case anyone is relying on content_id to actually uniquely identify a piece of content. \n\nIn reality, content_id is either a question_id of a question or a lecture_id of a lecture. The problem is that **these ids have a large overlap**.\n\nFrom a notebook where I checked this:\n```\nrepeat_ids = set(lectures_df['lecture_id']).intersection(set(qs_df['question_id']))\nlen(repeat_ids)\n```\n```\n158\n```\n\nso **158/418 lectures have ids that overlap seemingly-unrelated question ids**\n\nIn particular, I first noticed it when trying to summarize results with content_id=4057. Question with id 4057 has to do with tag 152, whereas lecture with id 4057 has to do with tag 1.\n\nso **don't group on content_id, it may produce nonsensical results**",
    "1131630": "If you filter first  by content_type_id, you can group by content_id",
    "1131753": "True.\n\nI'm not saying that there are no workarounds, I'm just saying that one has to be careful and make sure apply the workarounds if this is important to them. \n\nAnd that this is not obvious from the name (or description) of the field.",
    "1131759": "Aah ok ok, sorry. It's a good point to keep in mind"
  },
  "source": "meta"
}