{
  "id": 192839,
  "title": "Content IDs overlap for questions and lectures",
  "url": "/competitions/riiid-test-answer-prediction/discussion/192839",
  "author_name": "Joe Eddy",
  "post_date": "2020-10-23T14:57:15.130000",
  "votes": 27,
  "comment_count": 0,
  "views": 0,
  "content": "<p>Maybe this was easier to catch for most people, but it's a bit of a format quirk and I missed it for a while so figured I'd highlight it in case it helps others. It is implied in the data description but I don't think it's totally obvious. </p>\n<p>While <code>content_id</code> is unique for each individual question, that id can be reused to identify a lecture. So e.g. there is an id 0 for a question and an id 0 for a lecture. The upshot of this is that when you compute statistics at the <code>content_id</code> level, you should first filter down to the appropriate subset (all questions or all lectures based on <code>content_type_id</code>) to make sure that you're not bleeding the two types of content into each other. </p>\n<p>I realized this by checking mean question correctness stats and finding it odd that some of those means were negative…..</p>",
  "messages": [
    {
      "id": 1058310,
      "postDate": "2020-10-23T14:57:15.130Z",
      "content": "<p>Maybe this was easier to catch for most people, but it's a bit of a format quirk and I missed it for a while so figured I'd highlight it in case it helps others. It is implied in the data description but I don't think it's totally obvious. </p>\n<p>While <code>content_id</code> is unique for each individual question, that id can be reused to identify a lecture. So e.g. there is an id 0 for a question and an id 0 for a lecture. The upshot of this is that when you compute statistics at the <code>content_id</code> level, you should first filter down to the appropriate subset (all questions or all lectures based on <code>content_type_id</code>) to make sure that you're not bleeding the two types of content into each other. </p>\n<p>I realized this by checking mean question correctness stats and finding it odd that some of those means were negative…..</p>",
      "rawMarkdown": "Maybe this was easier to catch for most people, but it's a bit of a format quirk and I missed it for a while so figured I'd highlight it in case it helps others. It is implied in the data description but I don't think it's totally obvious. \n\nWhile `content_id` is unique for each individual question, that id can be reused to identify a lecture. So e.g. there is an id 0 for a question and an id 0 for a lecture. The upshot of this is that when you compute statistics at the `content_id` level, you should first filter down to the appropriate subset (all questions or all lectures based on `content_type_id`) to make sure that you're not bleeding the two types of content into each other. \n\nI realized this by checking mean question correctness stats and finding it odd that some of those means were negative.....",
      "votes": 27
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "1058310": "Maybe this was easier to catch for most people, but it's a bit of a format quirk and I missed it for a while so figured I'd highlight it in case it helps others. It is implied in the data description but I don't think it's totally obvious. \n\nWhile `content_id` is unique for each individual question, that id can be reused to identify a lecture. So e.g. there is an id 0 for a question and an id 0 for a lecture. The upshot of this is that when you compute statistics at the `content_id` level, you should first filter down to the appropriate subset (all questions or all lectures based on `content_type_id`) to make sure that you're not bleeding the two types of content into each other. \n\nI realized this by checking mean question correctness stats and finding it odd that some of those means were negative....."
  }
}