{
  "id": 190828,
  "title": "Question about task_container_id and bundle_id",
  "url": "/competitions/riiid-test-answer-prediction/discussion/190828",
  "author_name": "",
  "post_date": "2020-10-13T13:51:00.842201400Z",
  "votes": 27,
  "comment_count": 8,
  "views": 0,
  "content": "<p>Hi <a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a> and organizers,</p>\n<p>I had a query related to the questions metadata. The question is as follows:</p>\n<p>In train.csv we have a column 'task_container_id'. The description of this column says  \"code for the batch of questions or lectures. For example, a user might see three questions in a row before seeing the explanations for any of them. Those three would all share a task_container_id\"</p>\n<p>Now, the questions.csv has a column 'bundle_id'. The description of this column says \"code for which questions are served together\".</p>\n<p>Thus, for rows in train/test where the content type is question (0), would it be fair to assume that 'task_container_id' is the same as 'bundle_id'?</p>\n<p>I am entering this competition just now, so I apologize if this is a repeat question from some other thread.</p>\n<p>Cheers!!</p>",
  "messages": [
    {
      "id": "1048438",
      "postDate": "10/13/2020 13:51:00",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a> and organizers,</p>\n<p>I had a query related to the questions metadata. The question is as follows:</p>\n<p>In train.csv we have a column 'task_container_id'. The description of this column says  \"code for the batch of questions or lectures. For example, a user might see three questions in a row before seeing the explanations for any of them. Those three would all share a task_container_id\"</p>\n<p>Now, the questions.csv has a column 'bundle_id'. The description of this column says \"code for which questions are served together\".</p>\n<p>Thus, for rows in train/test where the content type is question (0), would it be fair to assume that 'task_container_id' is the same as 'bundle_id'?</p>\n<p>I am entering this competition just now, so I apologize if this is a repeat question from some other thread.</p>\n<p>Cheers!!</p>",
      "rawMarkdown": "Hi @sohier and organizers,\n\nI had a query related to the questions metadata. The question is as follows:\n\nIn train.csv we have a column 'task_container_id'. The description of this column says  \"code for the batch of questions or lectures. For example, a user might see three questions in a row before seeing the explanations for any of them. Those three would all share a task_container_id\"\n\nNow, the questions.csv has a column 'bundle_id'. The description of this column says \"code for which questions are served together\".\n\nThus, for rows in train/test where the content type is question (0), would it be fair to assume that 'task_container_id' is the same as 'bundle_id'?\n\nI am entering this competition just now, so I apologize if this is a repeat question from some other thread.\n\nCheers!!",
      "votes": null
    },
    {
      "id": "1051204",
      "postDate": "10/16/2020 08:59:45",
      "content": "<p><a href=\"https://www.kaggle.com/sandy1112\" target=\"_blank\">@sandy1112</a> I had the very same question in mind. Although the description in questions.csv says so, I think it's much more reliable to use <code>task_container_id</code> instead of <code>bundle_id</code>. </p>\n<p>Reason:<br>\n<code>prior_question_elapsed_time</code> is the avg time a student takes to complete the previous batch (total_time taken / total_no_questions). So it should follow that prior_question_elapsed time must be the same for any row of any user for particular <em>bundle</em> (bundle being a group of questions that a user sees at a time). </p>\n<p>I tried verifying the above statement by grouping all rows and checking the number of unique prior_question_elapsed_time values:</p>\n<ul>\n<li>Grouping by ['user_id', 'task_container_id']:</li>\n</ul>\n<pre><code># unique prior_elapsed_time per bundle is always one (0 for nans)\n(data.groupby(['user_id', 'task_container_id'])\n ['prior_question_elapsed_time']\n .nunique().values &lt;= 1).all()\n\n&gt;&gt;&gt; True\n</code></pre>\n<ul>\n<li>Grouping by ['user_id', 'bundle_id']</li>\n</ul>\n<pre><code># unique prior_elapsed_time per bundle is always one (0 for nans)\n(data.merge(ql, on=['content_id', 'content_type_id'])\n .groupby(['user_id', 'bundle_id'])\n ['prior_question_elapsed_time']\n .nunique().values &lt;= 1).all()\n\n&gt;&gt;&gt; False\n</code></pre>\n<p>My logic is that since the other features match better with the assumption that each bundle have the same task_container_id, it is much better to use this feature instead of using bundle_id. I may be totally wrong. Do let me know your thoughts on this.</p>",
      "rawMarkdown": "sandy1112 I had the very same question in mind. Although the description in questions.csv says so, I think it's much more reliable to use `task_container_id` instead of `bundle_id`. \n\nReason:\n`prior_question_elapsed_time` is the avg time a student takes to complete the previous batch (total_time taken / total_no_questions). So it should follow that prior_question_elapsed time must be the same for any row of any user for particular *bundle* (bundle being a group of questions that a user sees at a time). \n\nI tried verifying the above statement by grouping all rows and checking the number of unique prior_question_elapsed_time values:\n\n- Grouping by ['user_id', 'task_container_id']:\n\n```python3\n# unique prior_elapsed_time per bundle is always one (0 for nans)\n(data.groupby(['user_id', 'task_container_id'])\n ['prior_question_elapsed_time']\n .nunique().values <= 1).all()\n\n>>> True\n```\n\n- Grouping by ['user_id', 'bundle_id']\n\n```python3\n# unique prior_elapsed_time per bundle is always one (0 for nans)\n(data.merge(ql, on=['content_id', 'content_type_id'])\n .groupby(['user_id', 'bundle_id'])\n ['prior_question_elapsed_time']\n .nunique().values <= 1).all()\n\n>>> False\n```\n\nMy logic is that since the other features match better with the assumption that each bundle have the same task_container_id, it is much better to use this feature instead of using bundle_id. I may be totally wrong. Do let me know your thoughts on this.",
      "votes": null
    },
    {
      "id": "1051526",
      "postDate": "10/16/2020 15:50:19",
      "content": "<p>They're totally distinct concepts. Questions with the same bundle ID should be served together, but every time they are served they will be under a different task container id (for a given user). </p>",
      "rawMarkdown": "They're totally distinct concepts. Questions with the same bundle ID should be served together, but every time they are served they will be under a different task container id (for a given user).",
      "votes": null
    },
    {
      "id": "1051533",
      "postDate": "10/16/2020 15:54:49",
      "content": "<p>Thanks for confirming. Appreciate the help on this.</p>",
      "rawMarkdown": "Thanks for confirming. Appreciate the help on this.",
      "votes": null
    },
    {
      "id": "1064479",
      "postDate": "10/30/2020 07:21:45",
      "content": "<p><a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a>  does it means that  bundle id could remain same for set of questions grouped under one.. but they may keep appering under different Task container ids  together as bundle</p>",
      "rawMarkdown": "sohier  does it means that  bundle id could remain same for set of questions grouped under one.. but they may keep appering under different Task container ids  together as bundle",
      "votes": null
    },
    {
      "id": "1109640",
      "postDate": "12/11/2020 22:45:53",
      "content": "<p><a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a> how to explain same bunld_id with same contrainer_id and same user_id?</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3273905%2F4c409b21afcef0f3e5ae172d6090a844%2FWX20201212-0644132x.png?generation=1607726695891832&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "sohier how to explain same bunld_id with same contrainer_id and same user_id?\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3273905%2F4c409b21afcef0f3e5ae172d6090a844%2FWX20201212-0644132x.png?generation=1607726695891832&alt=media)",
      "votes": null
    },
    {
      "id": "1118325",
      "postDate": "12/19/2020 00:20:32",
      "content": "<blockquote>\n  <p>each bundle have the same task_container_id</p>\n</blockquote>\n<p>Does this mean, even for different users, if the questions are from the same bundle, these questions share the same <code>task_container_id</code>?</p>",
      "rawMarkdown": "> each bundle have the same task_container_id\n\nDoes this mean, even for different users, if the questions are from the same bundle, these questions share the same `task_container_id`?",
      "votes": null
    },
    {
      "id": "1118383",
      "postDate": "12/19/2020 02:53:15",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/wuwenmin\" target=\"_blank\">@wuwenmin</a>, questions of the same bundle share the same <code>task_container_id</code> but for different users the <em>shared</em> task_container_ids might still be different. <code>task_container_id</code>'s purpose is to point at the order in which the user had seen the questions. It is not guaranteed to be <em>same</em> for different users. </p>\n<p>For one user the question bundle containing 3 questions might have task_container_ids as [2, 2, 2] whereas for another user the same question bundle might be something like [998, 998, 998]. Hope this helps :) </p>",
      "rawMarkdown": "Hi @wuwenmin, questions of the same bundle share the same `task_container_id` but for different users the *shared* task_container_ids might still be different. `task_container_id`'s purpose is to point at the order in which the user had seen the questions. It is not guaranteed to be *same* for different users. \n\nFor one user the question bundle containing 3 questions might have task_container_ids as [2, 2, 2] whereas for another user the same question bundle might be something like [998, 998, 998]. Hope this helps :)",
      "votes": null
    },
    {
      "id": "1118390",
      "postDate": "12/19/2020 03:01:43",
      "content": "<p>Thanks for your reply. Yes, I just checked the data, the following are the distribution of the number of unique questions of the container task id:</p>\n<pre><code>count    10000.00000\nmean      2766.53760\nstd       3155.41549\nmin        165.00000\n25%        411.75000\n50%       1228.00000\n75%       4263.50000\nmax      11503.00000\nName: num_ques, dtype: float64\n</code></pre>\n<p>I saw some public notebooks use the <code>task_container_id</code> as a feature, so I was wondering whether the <code>task_container_id</code> indicates the group of questions for all users. It's not.</p>",
      "rawMarkdown": "Thanks for your reply. Yes, I just checked the data, the following are the distribution of the number of unique questions of the container task id:\n```\ncount    10000.00000\nmean      2766.53760\nstd       3155.41549\nmin        165.00000\n25%        411.75000\n50%       1228.00000\n75%       4263.50000\nmax      11503.00000\nName: num_ques, dtype: float64\n```\nI saw some public notebooks use the `task_container_id` as a feature, so I was wondering whether the `task_container_id` indicates the group of questions for all users. It's not.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1051204,
      "author_name": "doctorkael",
      "author_url": "",
      "post_date": "10/16/2020 08:59:45",
      "content": "<p><a href=\"https://www.kaggle.com/sandy1112\" target=\"_blank\">@sandy1112</a> I had the very same question in mind. Although the description in questions.csv says so, I think it's much more reliable to use <code>task_container_id</code> instead of <code>bundle_id</code>. </p>\n<p>Reason:<br>\n<code>prior_question_elapsed_time</code> is the avg time a student takes to complete the previous batch (total_time taken / total_no_questions). So it should follow that prior_question_elapsed time must be the same for any row of any user for particular <em>bundle</em> (bundle being a group of questions that a user sees at a time). </p>\n<p>I tried verifying the above statement by grouping all rows and checking the number of unique prior_question_elapsed_time values:</p>\n<ul>\n<li>Grouping by ['user_id', 'task_container_id']:</li>\n</ul>\n<pre><code># unique prior_elapsed_time per bundle is always one (0 for nans)\n(data.groupby(['user_id', 'task_container_id'])\n ['prior_question_elapsed_time']\n .nunique().values &lt;= 1).all()\n\n&gt;&gt;&gt; True\n</code></pre>\n<ul>\n<li>Grouping by ['user_id', 'bundle_id']</li>\n</ul>\n<pre><code># unique prior_elapsed_time per bundle is always one (0 for nans)\n(data.merge(ql, on=['content_id', 'content_type_id'])\n .groupby(['user_id', 'bundle_id'])\n ['prior_question_elapsed_time']\n .nunique().values &lt;= 1).all()\n\n&gt;&gt;&gt; False\n</code></pre>\n<p>My logic is that since the other features match better with the assumption that each bundle have the same task_container_id, it is much better to use this feature instead of using bundle_id. I may be totally wrong. Do let me know your thoughts on this.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1118325,
          "author_name": "wuwenmin",
          "author_url": "",
          "post_date": "12/19/2020 00:20:32",
          "content": "<blockquote>\n  <p>each bundle have the same task_container_id</p>\n</blockquote>\n<p>Does this mean, even for different users, if the questions are from the same bundle, these questions share the same <code>task_container_id</code>?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1118383,
          "author_name": "doctorkael",
          "author_url": "",
          "post_date": "12/19/2020 02:53:15",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/wuwenmin\" target=\"_blank\">@wuwenmin</a>, questions of the same bundle share the same <code>task_container_id</code> but for different users the <em>shared</em> task_container_ids might still be different. <code>task_container_id</code>'s purpose is to point at the order in which the user had seen the questions. It is not guaranteed to be <em>same</em> for different users. </p>\n<p>For one user the question bundle containing 3 questions might have task_container_ids as [2, 2, 2] whereas for another user the same question bundle might be something like [998, 998, 998]. Hope this helps :) </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1118390,
          "author_name": "wuwenmin",
          "author_url": "",
          "post_date": "12/19/2020 03:01:43",
          "content": "<p>Thanks for your reply. Yes, I just checked the data, the following are the distribution of the number of unique questions of the container task id:</p>\n<pre><code>count    10000.00000\nmean      2766.53760\nstd       3155.41549\nmin        165.00000\n25%        411.75000\n50%       1228.00000\n75%       4263.50000\nmax      11503.00000\nName: num_ques, dtype: float64\n</code></pre>\n<p>I saw some public notebooks use the <code>task_container_id</code> as a feature, so I was wondering whether the <code>task_container_id</code> indicates the group of questions for all users. It's not.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1051526,
      "author_name": "sohier",
      "author_url": "",
      "post_date": "10/16/2020 15:50:19",
      "content": "<p>They're totally distinct concepts. Questions with the same bundle ID should be served together, but every time they are served they will be under a different task container id (for a given user). </p>",
      "votes": null,
      "replies": [
        {
          "id": 1051533,
          "author_name": "sandy1112",
          "author_url": "",
          "post_date": "10/16/2020 15:54:49",
          "content": "<p>Thanks for confirming. Appreciate the help on this.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1064479,
          "author_name": "jaideepvalani",
          "author_url": "",
          "post_date": "10/30/2020 07:21:45",
          "content": "<p><a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a>  does it means that  bundle id could remain same for set of questions grouped under one.. but they may keep appering under different Task container ids  together as bundle</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1109640,
          "author_name": "jt120lz",
          "author_url": "",
          "post_date": "12/11/2020 22:45:53",
          "content": "<p><a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a> how to explain same bunld_id with same contrainer_id and same user_id?</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3273905%2F4c409b21afcef0f3e5ae172d6090a844%2FWX20201212-0644132x.png?generation=1607726695891832&amp;alt=media\" alt=\"\"></p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1048438": "Hi @sohier and organizers,\n\nI had a query related to the questions metadata. The question is as follows:\n\nIn train.csv we have a column 'task_container_id'. The description of this column says  \"code for the batch of questions or lectures. For example, a user might see three questions in a row before seeing the explanations for any of them. Those three would all share a task_container_id\"\n\nNow, the questions.csv has a column 'bundle_id'. The description of this column says \"code for which questions are served together\".\n\nThus, for rows in train/test where the content type is question (0), would it be fair to assume that 'task_container_id' is the same as 'bundle_id'?\n\nI am entering this competition just now, so I apologize if this is a repeat question from some other thread.\n\nCheers!!",
    "1051204": "sandy1112 I had the very same question in mind. Although the description in questions.csv says so, I think it's much more reliable to use `task_container_id` instead of `bundle_id`. \n\nReason:\n`prior_question_elapsed_time` is the avg time a student takes to complete the previous batch (total_time taken / total_no_questions). So it should follow that prior_question_elapsed time must be the same for any row of any user for particular *bundle* (bundle being a group of questions that a user sees at a time). \n\nI tried verifying the above statement by grouping all rows and checking the number of unique prior_question_elapsed_time values:\n\n- Grouping by ['user_id', 'task_container_id']:\n\n```python3\n# unique prior_elapsed_time per bundle is always one (0 for nans)\n(data.groupby(['user_id', 'task_container_id'])\n ['prior_question_elapsed_time']\n .nunique().values <= 1).all()\n\n>>> True\n```\n\n- Grouping by ['user_id', 'bundle_id']\n\n```python3\n# unique prior_elapsed_time per bundle is always one (0 for nans)\n(data.merge(ql, on=['content_id', 'content_type_id'])\n .groupby(['user_id', 'bundle_id'])\n ['prior_question_elapsed_time']\n .nunique().values <= 1).all()\n\n>>> False\n```\n\nMy logic is that since the other features match better with the assumption that each bundle have the same task_container_id, it is much better to use this feature instead of using bundle_id. I may be totally wrong. Do let me know your thoughts on this.",
    "1051526": "They're totally distinct concepts. Questions with the same bundle ID should be served together, but every time they are served they will be under a different task container id (for a given user).",
    "1051533": "Thanks for confirming. Appreciate the help on this.",
    "1064479": "sohier  does it means that  bundle id could remain same for set of questions grouped under one.. but they may keep appering under different Task container ids  together as bundle",
    "1109640": "sohier how to explain same bunld_id with same contrainer_id and same user_id?\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3273905%2F4c409b21afcef0f3e5ae172d6090a844%2FWX20201212-0644132x.png?generation=1607726695891832&alt=media)",
    "1118325": "> each bundle have the same task_container_id\n\nDoes this mean, even for different users, if the questions are from the same bundle, these questions share the same `task_container_id`?",
    "1118383": "Hi @wuwenmin, questions of the same bundle share the same `task_container_id` but for different users the *shared* task_container_ids might still be different. `task_container_id`'s purpose is to point at the order in which the user had seen the questions. It is not guaranteed to be *same* for different users. \n\nFor one user the question bundle containing 3 questions might have task_container_ids as [2, 2, 2] whereas for another user the same question bundle might be something like [998, 998, 998]. Hope this helps :)",
    "1118390": "Thanks for your reply. Yes, I just checked the data, the following are the distribution of the number of unique questions of the container task id:\n```\ncount    10000.00000\nmean      2766.53760\nstd       3155.41549\nmin        165.00000\n25%        411.75000\n50%       1228.00000\n75%       4263.50000\nmax      11503.00000\nName: num_ques, dtype: float64\n```\nI saw some public notebooks use the `task_container_id` as a feature, so I was wondering whether the `task_container_id` indicates the group of questions for all users. It's not."
  },
  "source": "meta"
}