{
  "id": 198622,
  "title": "task_container_id: relationship to (test data) batches and (question) bundles",
  "url": "/competitions/riiid-test-answer-prediction/discussion/198622",
  "author_name": "",
  "post_date": "2020-11-22T05:10:33.166079700Z",
  "votes": 1,
  "comment_count": 1,
  "views": 0,
  "content": "<p>I'm just getting started on this challenge. I've read what seem to be the relevant discussisons, but I'm still a bit confused about the task_container_id column in train.csv.<br>\nSpecifically:</p>\n<ol>\n<li>How are the task container batches related to the batches that the test data comes in, if at all? In particular, is it possible that a single user's task container would get split across several test data batches? or is each batch of test data guaranteed to contain some set of \"whole\" task_container_id batches?</li>\n<li>How are task_container_id batches related to the bundles of questions? I understand that the ids themselves aren't guaranteed to match, but can a task_container_id for a single user contain questions from several bundles? or does a task container always represent at most one bundle of questions, with perhaps some lectures added?</li>\n</ol>",
  "messages": [
    {
      "id": "1086833",
      "postDate": "11/22/2020 05:10:33",
      "content": "<p>I'm just getting started on this challenge. I've read what seem to be the relevant discussisons, but I'm still a bit confused about the task_container_id column in train.csv.<br>\nSpecifically:</p>\n<ol>\n<li>How are the task container batches related to the batches that the test data comes in, if at all? In particular, is it possible that a single user's task container would get split across several test data batches? or is each batch of test data guaranteed to contain some set of \"whole\" task_container_id batches?</li>\n<li>How are task_container_id batches related to the bundles of questions? I understand that the ids themselves aren't guaranteed to match, but can a task_container_id for a single user contain questions from several bundles? or does a task container always represent at most one bundle of questions, with perhaps some lectures added?</li>\n</ol>",
      "rawMarkdown": "I'm just getting started on this challenge. I've read what seem to be the relevant discussisons, but I'm still a bit confused about the task_container_id column in train.csv.\n\nSpecifically:\n1. How are the task container batches related to the batches that the test data comes in, if at all? In particular, is it possible that a single user's task container would get split across several test data batches? or is each batch of test data guaranteed to contain some set of \"whole\" task_container_id batches?\n2. How are task_container_id batches related to the bundles of questions? I understand that the ids themselves aren't guaranteed to match, but can a task_container_id for a single user contain questions from several bundles? or does a task container always represent at most one bundle of questions, with perhaps some lectures added?",
      "votes": null
    },
    {
      "id": "1099036",
      "postDate": "12/02/2020 03:22:50",
      "content": "<p>Experimentation suggests that each task_container_id for a given user does map exactly onto some bundle_id from the questions table. Still not sure about question (1).</p>\n<p>I do wish the owners of this competition would clarify these points.</p>",
      "rawMarkdown": "Experimentation suggests that each task_container_id for a given user does map exactly onto some bundle_id from the questions table. Still not sure about question (1).\n\nI do wish the owners of this competition would clarify these points.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1099036,
      "author_name": "yanamal",
      "author_url": "",
      "post_date": "12/02/2020 03:22:50",
      "content": "<p>Experimentation suggests that each task_container_id for a given user does map exactly onto some bundle_id from the questions table. Still not sure about question (1).</p>\n<p>I do wish the owners of this competition would clarify these points.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1086833": "I'm just getting started on this challenge. I've read what seem to be the relevant discussisons, but I'm still a bit confused about the task_container_id column in train.csv.\n\nSpecifically:\n1. How are the task container batches related to the batches that the test data comes in, if at all? In particular, is it possible that a single user's task container would get split across several test data batches? or is each batch of test data guaranteed to contain some set of \"whole\" task_container_id batches?\n2. How are task_container_id batches related to the bundles of questions? I understand that the ids themselves aren't guaranteed to match, but can a task_container_id for a single user contain questions from several bundles? or does a task container always represent at most one bundle of questions, with perhaps some lectures added?",
    "1099036": "Experimentation suggests that each task_container_id for a given user does map exactly onto some bundle_id from the questions table. Still not sure about question (1).\n\nI do wish the owners of this competition would clarify these points."
  },
  "source": "meta"
}