{
  "id": 189465,
  "title": "Inconsistency between timestamp and task_container_id?",
  "url": "/competitions/riiid-test-answer-prediction/discussion/189465",
  "author_name": "Zhao Chow",
  "post_date": "2020-10-07T17:12:00.533000",
  "votes": 33,
  "comment_count": 13,
  "views": 0,
  "content": "<p>It seems that there is an inconsistency between the <code>timestamp</code> and <code>task_container_id</code> columns in the training set (or my understanding is incorrect). The definitions for both are:</p>\n<ul>\n<li><code>timestamp</code>: (int64) the time between this user interaction and the first event from that user.</li>\n<li><code>task_container_id</code>:  (int16) Id code for the batch of questions or lectures. For example, a user might see three questions in a row before seeing the explanations for any of them. Those three would all share a <code>task_container_id</code>. Monotonically increasing for each user.</li>\n</ul>\n<p>From my understanding, <code>task_container_id</code> should start from 0 and monotonically increase with the <code>timestamp</code>. However, it is not the case for some users (e.g. user 115).</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5589062%2Fb25f3e890ad72c5d919fcc03f91d891d%2Fkaggle.png?generation=1602090463430829&amp;alt=media\" alt=\"\"></p>\n<p>It would be great to confirm with the hosts of this competition if this is an actual mistake. In the meantime, I would like to share a fix suggested by <a href=\"https://www.kaggle.com/maxhalford\" target=\"_blank\">@maxhalford</a> from the original discussion.</p>\n<blockquote>\n  <p>To fix this, I suggest renumbering the tasks to make sure they're monotonically increasing for each user:</p>\n<pre><code>train['task_container_id'] = (\n    train\n    .groupby('user_id')['task_container_id']\n    .transform(lambda x: pd.factorize(x)[0])\n    .astype('int16')\n)\n</code></pre>\n</blockquote>\n<p>I am also interested to know how much applying this fix will impact the training and score so feel free to give your insights in the comments.</p>",
  "messages": [
    {
      "id": 1041289,
      "postDate": "2020-10-07T17:12:00.533Z",
      "content": "<p>It seems that there is an inconsistency between the <code>timestamp</code> and <code>task_container_id</code> columns in the training set (or my understanding is incorrect). The definitions for both are:</p>\n<ul>\n<li><code>timestamp</code>: (int64) the time between this user interaction and the first event from that user.</li>\n<li><code>task_container_id</code>:  (int16) Id code for the batch of questions or lectures. For example, a user might see three questions in a row before seeing the explanations for any of them. Those three would all share a <code>task_container_id</code>. Monotonically increasing for each user.</li>\n</ul>\n<p>From my understanding, <code>task_container_id</code> should start from 0 and monotonically increase with the <code>timestamp</code>. However, it is not the case for some users (e.g. user 115).</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5589062%2Fb25f3e890ad72c5d919fcc03f91d891d%2Fkaggle.png?generation=1602090463430829&amp;alt=media\" alt=\"\"></p>\n<p>It would be great to confirm with the hosts of this competition if this is an actual mistake. In the meantime, I would like to share a fix suggested by <a href=\"https://www.kaggle.com/maxhalford\" target=\"_blank\">@maxhalford</a> from the original discussion.</p>\n<blockquote>\n  <p>To fix this, I suggest renumbering the tasks to make sure they're monotonically increasing for each user:</p>\n<pre><code>train['task_container_id'] = (\n    train\n    .groupby('user_id')['task_container_id']\n    .transform(lambda x: pd.factorize(x)[0])\n    .astype('int16')\n)\n</code></pre>\n</blockquote>\n<p>I am also interested to know how much applying this fix will impact the training and score so feel free to give your insights in the comments.</p>",
      "rawMarkdown": "It seems that there is an inconsistency between the `timestamp` and `task_container_id` columns in the training set (or my understanding is incorrect). The definitions for both are:\n\n- `timestamp`: (int64) the time between this user interaction and the first event from that user.\n- `task_container_id`:  (int16) Id code for the batch of questions or lectures. For example, a user might see three questions in a row before seeing the explanations for any of them. Those three would all share a `task_container_id`. Monotonically increasing for each user.\n\nFrom my understanding, `task_container_id` should start from 0 and monotonically increase with the `timestamp`. However, it is not the case for some users (e.g. user 115).\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5589062%2Fb25f3e890ad72c5d919fcc03f91d891d%2Fkaggle.png?generation=1602090463430829&alt=media)\n\nIt would be great to confirm with the hosts of this competition if this is an actual mistake. In the meantime, I would like to share a fix suggested by @maxhalford from the original discussion.\n\n> To fix this, I suggest renumbering the tasks to make sure they're monotonically increasing for each user:\n> \n> ```python\n> train['task_container_id'] = (\n>     train\n>     .groupby('user_id')['task_container_id']\n>     .transform(lambda x: pd.factorize(x)[0])\n>     .astype('int16')\n> )\n> ```\n\nI am also interested to know how much applying this fix will impact the training and score so feel free to give your insights in the comments.",
      "votes": 33
    },
    {
      "id": 1041597,
      "postDate": "2020-10-07T20:42:15.730Z",
      "content": "<p>From <a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a> in this other <a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/189351\" target=\"_blank\">thread</a> </p>\n<blockquote>\n  <ul>\n  <li>The <code>timestamp</code> column shows when an activity is finished, not when it started. A user could have spent an arbitrary amount of time working on the first problem before the first <code>timestamp</code> value would get logged.</li>\n  </ul>\n</blockquote>\n<p>To me this suggests that while <code>task_container_id</code> captures the order that a user <strong>first sees</strong> tasks in, <code>timestamp</code> captures the order in which a user actually <strong>completes</strong> tasks. A user can start one task, start a second, and then finish the second before finishing the first -- this results in the later timestamp for the first task.</p>",
      "rawMarkdown": "From @sohier in this other [thread](https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/189351) \n\n> - The `timestamp` column shows when an activity is finished, not when it started. A user could have spent an arbitrary amount of time working on the first problem before the first `timestamp` value would get logged.\n\nTo me this suggests that while `task_container_id` captures the order that a user **first sees** tasks in, `timestamp` captures the order in which a user actually **completes** tasks. A user can start one task, start a second, and then finish the second before finishing the first -- this results in the later timestamp for the first task.",
      "votes": 13,
      "replies": [
        {
          "id": 1041925,
          "postDate": "2020-10-08T01:18:47.657Z",
          "content": "<p>Okay with this additional information on <code>timestamp</code>, it makes sense now! Thanks!</p>",
          "rawMarkdown": "Okay with this additional information on `timestamp`, it makes sense now! Thanks!",
          "votes": 2
        },
        {
          "id": 1100598,
          "postDate": "2020-12-03T07:18:05.673Z",
          "content": "<p><a href=\"https://www.kaggle.com/zhaochow\" target=\"_blank\">@zhaochow</a>  could you help in understanding what do we exactly term as task container id.. is it pre fixed value by  learning platfor, like user will be served some xyz question together so they would belong to some task container id ,  or it is not fixed   then on what does it depends on</p>\n<p>secondly, how can we related bundle id to task container id if there is any relation between two ?</p>",
          "rawMarkdown": "@zhaochow  could you help in understanding what do we exactly term as task container id.. is it pre fixed value by  learning platfor, like user will be served some xyz question together so they would belong to some task container id ,  or it is not fixed   then on what does it depends on\n\nsecondly, how can we related bundle id to task container id if there is any relation between two ?\n\n"
        }
      ]
    },
    {
      "id": 1041603,
      "postDate": "2020-10-07T20:46:31.153Z",
      "content": "<p>I won't have time to dig in to confirm this today, but I believe this is just a data artifact. Imagine someone logging in from both their phone and a tablet; the first batch of questions they started working on (lowest container ID) wouldn't necessarily actually be the first one to finish.</p>\n<p>Generally speaking, time tracked across multiple devices is unlikely to be precisely correct unless you take extreme measures like providing each device <a href=\"https://www.wired.com/2012/11/google-spanner-time/\" target=\"_blank\">with an atomic clock </a>.</p>",
      "rawMarkdown": "I won't have time to dig in to confirm this today, but I believe this is just a data artifact. Imagine someone logging in from both their phone and a tablet; the first batch of questions they started working on (lowest container ID) wouldn't necessarily actually be the first one to finish.\n\nGenerally speaking, time tracked across multiple devices is unlikely to be precisely correct unless you take extreme measures like providing each device [with an atomic clock ](https://www.wired.com/2012/11/google-spanner-time/).",
      "votes": 5,
      "replies": [
        {
          "id": 1041977,
          "postDate": "2020-10-08T02:06:12.117Z",
          "content": "<p>Thank you for your comment! If it is just a data artifact, we can thus leave the <code>task_container_id</code> column untouched. It is as <a href=\"https://www.kaggle.com/aquatic\" target=\"_blank\">@aquatic</a> said, <code>task_container_id</code> captures the order in which a user starts/sees tasks whereas the <code>timestamp</code> captures the order of completion. </p>\n<p>You used logging from multiple devices as an example. Just to be sure, could this situation happen even on a single device? For example, users can start a new task before finishing the previous one?</p>",
          "rawMarkdown": "Thank you for your comment! If it is just a data artifact, we can thus leave the `task_container_id` column untouched. It is as @aquatic said, `task_container_id` captures the order in which a user starts/sees tasks whereas the `timestamp` captures the order of completion. \n\nYou used logging from multiple devices as an example. Just to be sure, could this situation happen even on a single device? For example, users can start a new task before finishing the previous one?"
        },
        {
          "id": 1099684,
          "postDate": "2020-12-02T14:31:59.223Z",
          "content": "<p><a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a>  could you clarify on this doubbt regarding Timestamp.<br>\nI believe it starts from 0 for every user  when they start first their interaction with learning program. </p>\n<p>But what i find that Number of Unique Records where TImestamp=0 is differs by Unique user dis by few thousands .Can this be possiility , i expected them to be equal . </p>",
          "rawMarkdown": "@sohier  could you clarify on this doubbt regarding Timestamp.\nI believe it starts from 0 for every user  when they start first their interaction with learning program. \n\nBut what i find that Number of Unique Records where TImestamp=0 is differs by Unique user dis by few thousands .Can this be possiility , i expected them to be equal . ",
          "votes": 1
        }
      ]
    },
    {
      "id": 1041604,
      "postDate": "2020-10-07T20:46:55.660Z",
      "content": "<p>I guess that users can choose which groups of questions to work on. One user could do task 1, 2, and 4, then comes back to task 3.</p>",
      "rawMarkdown": "I guess that users can choose which groups of questions to work on. One user could do task 1, 2, and 4, then comes back to task 3."
    },
    {
      "id": 1041588,
      "postDate": "2020-10-07T20:35:05.023Z",
      "content": "<p>I also found this issue. The host should clarify this …BTW, where is the discussion you mentioned in <code>a fix suggested by @maxhalford from the original discussion.</code> ? Thanks.</p>",
      "rawMarkdown": "I also found this issue. The host should clarify this ...BTW, where is the discussion you mentioned in ` a fix suggested by @maxhalford from the original discussion.` ? Thanks.",
      "replies": [
        {
          "id": 1041700,
          "postDate": "2020-10-07T21:58:09.393Z",
          "content": "<p>Do you mean this discussion?<br>\n<a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/189351\" target=\"_blank\">https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/189351</a></p>",
          "rawMarkdown": "Do you mean this discussion?\nhttps://www.kaggle.com/c/riiid-test-answer-prediction/discussion/189351",
          "votes": 1
        }
      ]
    },
    {
      "id": 1047038,
      "postDate": "2020-10-12T07:40:05.747Z",
      "content": "<p>Thank you for bringing this issue and I apologize for any inconvenience. Indeed, the non-increasing values of <code>task_container_id</code> is data artifact due to internal workings of the service, and it serves the purpose of distinguishing different bundles without monotonicity assumption. We took this into account by dropping the original note about monotonic increase.</p>",
      "rawMarkdown": "Thank you for bringing this issue and I apologize for any inconvenience. Indeed, the non-increasing values of `task_container_id` is data artifact due to internal workings of the service, and it serves the purpose of distinguishing different bundles without monotonicity assumption. We took this into account by dropping the original note about monotonic increase.",
      "votes": 10,
      "isDeleted": true,
      "replies": [
        {
          "id": 1053821,
          "postDate": "2020-10-19T11:54:55.553Z",
          "content": "<p><a href=\"https://www.kaggle.com/jineonjinbaek\" target=\"_blank\">@jineonjinbaek</a> <a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a> I understood <code>task_container_id</code> is not monotonically increasing in the training set due to internal workings of the service. Then, is <code>task_container_id</code> not monotonically increasing in the test set? I mean, when you provide a group using generator <code>env.iter_test()</code>, is the sample with the smallest <code>task_container_id</code> always selected first, or is the sample with the smallest <code>timestamp</code> selected first? I'm looking forward to the clarification, thanks!</p>",
          "rawMarkdown": "@jineonjinbaek @sohier I understood ``task_container_id`` is not monotonically increasing in the training set due to internal workings of the service. Then, is ``task_container_id`` not monotonically increasing in the test set? I mean, when you provide a group using generator ``env.iter_test()``, is the sample with the smallest ``task_container_id`` always selected first, or is the sample with the smallest ``timestamp`` selected first? I'm looking forward to the clarification, thanks!",
          "votes": 12
        },
        {
          "id": 1065184,
          "postDate": "2020-10-31T01:21:33.710Z",
          "content": "<p>I have the same question!</p>",
          "rawMarkdown": "I have the same question!"
        },
        {
          "id": 1100420,
          "postDate": "2020-12-03T04:41:28.057Z",
          "content": "<p>Hi I am late to the party, has this been clariified in any way?</p>",
          "rawMarkdown": "Hi I am late to the party, has this been clariified in any way?",
          "votes": 1
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 1041597,
      "author_name": "Joe Eddy",
      "author_url": "",
      "post_date": "2020-10-07T20:42:15.730000",
      "content": "<p>From <a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a> in this other <a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/189351\" target=\"_blank\">thread</a> </p>\n<blockquote>\n  <ul>\n  <li>The <code>timestamp</code> column shows when an activity is finished, not when it started. A user could have spent an arbitrary amount of time working on the first problem before the first <code>timestamp</code> value would get logged.</li>\n  </ul>\n</blockquote>\n<p>To me this suggests that while <code>task_container_id</code> captures the order that a user <strong>first sees</strong> tasks in, <code>timestamp</code> captures the order in which a user actually <strong>completes</strong> tasks. A user can start one task, start a second, and then finish the second before finishing the first -- this results in the later timestamp for the first task.</p>",
      "votes": 13,
      "replies": [
        {
          "id": 1041925,
          "author_name": "Zhao Chow",
          "author_url": "",
          "post_date": "2020-10-08T01:18:47.657000",
          "content": "<p>Okay with this additional information on <code>timestamp</code>, it makes sense now! Thanks!</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1100598,
          "author_name": "Jaideep",
          "author_url": "",
          "post_date": "2020-12-03T07:18:05.673000",
          "content": "<p><a href=\"https://www.kaggle.com/zhaochow\" target=\"_blank\">@zhaochow</a>  could you help in understanding what do we exactly term as task container id.. is it pre fixed value by  learning platfor, like user will be served some xyz question together so they would belong to some task container id ,  or it is not fixed   then on what does it depends on</p>\n<p>secondly, how can we related bundle id to task container id if there is any relation between two ?</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1041603,
      "author_name": "Sohier Dane",
      "author_url": "",
      "post_date": "2020-10-07T20:46:31.153000",
      "content": "<p>I won't have time to dig in to confirm this today, but I believe this is just a data artifact. Imagine someone logging in from both their phone and a tablet; the first batch of questions they started working on (lowest container ID) wouldn't necessarily actually be the first one to finish.</p>\n<p>Generally speaking, time tracked across multiple devices is unlikely to be precisely correct unless you take extreme measures like providing each device <a href=\"https://www.wired.com/2012/11/google-spanner-time/\" target=\"_blank\">with an atomic clock </a>.</p>",
      "votes": 5,
      "replies": [
        {
          "id": 1041977,
          "author_name": "Zhao Chow",
          "author_url": "",
          "post_date": "2020-10-08T02:06:12.117000",
          "content": "<p>Thank you for your comment! If it is just a data artifact, we can thus leave the <code>task_container_id</code> column untouched. It is as <a href=\"https://www.kaggle.com/aquatic\" target=\"_blank\">@aquatic</a> said, <code>task_container_id</code> captures the order in which a user starts/sees tasks whereas the <code>timestamp</code> captures the order of completion. </p>\n<p>You used logging from multiple devices as an example. Just to be sure, could this situation happen even on a single device? For example, users can start a new task before finishing the previous one?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1099684,
          "author_name": "Jaideep",
          "author_url": "",
          "post_date": "2020-12-02T14:31:59.223000",
          "content": "<p><a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a>  could you clarify on this doubbt regarding Timestamp.<br>\nI believe it starts from 0 for every user  when they start first their interaction with learning program. </p>\n<p>But what i find that Number of Unique Records where TImestamp=0 is differs by Unique user dis by few thousands .Can this be possiility , i expected them to be equal . </p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1041604,
      "author_name": "statlearning",
      "author_url": "",
      "post_date": "2020-10-07T20:46:55.660000",
      "content": "<p>I guess that users can choose which groups of questions to work on. One user could do task 1, 2, and 4, then comes back to task 3.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1041588,
      "author_name": "Yih-Dar SHIEH",
      "author_url": "",
      "post_date": "2020-10-07T20:35:05.023000",
      "content": "<p>I also found this issue. The host should clarify this …BTW, where is the discussion you mentioned in <code>a fix suggested by @maxhalford from the original discussion.</code> ? Thanks.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1041700,
          "author_name": "n17r",
          "author_url": "",
          "post_date": "2020-10-07T21:58:09.393000",
          "content": "<p>Do you mean this discussion?<br>\n<a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/189351\" target=\"_blank\">https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/189351</a></p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1047038,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-10-12T07:40:05.747000",
      "content": "<p>Thank you for bringing this issue and I apologize for any inconvenience. Indeed, the non-increasing values of <code>task_container_id</code> is data artifact due to internal workings of the service, and it serves the purpose of distinguishing different bundles without monotonicity assumption. We took this into account by dropping the original note about monotonic increase.</p>",
      "votes": 10,
      "replies": [
        {
          "id": 1053821,
          "author_name": "mamas",
          "author_url": "",
          "post_date": "2020-10-19T11:54:55.553000",
          "content": "<p><a href=\"https://www.kaggle.com/jineonjinbaek\" target=\"_blank\">@jineonjinbaek</a> <a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a> I understood <code>task_container_id</code> is not monotonically increasing in the training set due to internal workings of the service. Then, is <code>task_container_id</code> not monotonically increasing in the test set? I mean, when you provide a group using generator <code>env.iter_test()</code>, is the sample with the smallest <code>task_container_id</code> always selected first, or is the sample with the smallest <code>timestamp</code> selected first? I'm looking forward to the clarification, thanks!</p>",
          "votes": 12,
          "replies": []
        },
        {
          "id": 1065184,
          "author_name": "jwc",
          "author_url": "",
          "post_date": "2020-10-31T01:21:33.710000",
          "content": "<p>I have the same question!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1100420,
          "author_name": "Felipe Bivort Haiek",
          "author_url": "",
          "post_date": "2020-12-03T04:41:28.057000",
          "content": "<p>Hi I am late to the party, has this been clariified in any way?</p>",
          "votes": 1,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1041289": "It seems that there is an inconsistency between the `timestamp` and `task_container_id` columns in the training set (or my understanding is incorrect). The definitions for both are:\n\n- `timestamp`: (int64) the time between this user interaction and the first event from that user.\n- `task_container_id`:  (int16) Id code for the batch of questions or lectures. For example, a user might see three questions in a row before seeing the explanations for any of them. Those three would all share a `task_container_id`. Monotonically increasing for each user.\n\nFrom my understanding, `task_container_id` should start from 0 and monotonically increase with the `timestamp`. However, it is not the case for some users (e.g. user 115).\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5589062%2Fb25f3e890ad72c5d919fcc03f91d891d%2Fkaggle.png?generation=1602090463430829&alt=media)\n\nIt would be great to confirm with the hosts of this competition if this is an actual mistake. In the meantime, I would like to share a fix suggested by @maxhalford from the original discussion.\n\n> To fix this, I suggest renumbering the tasks to make sure they're monotonically increasing for each user:\n> \n> ```python\n> train['task_container_id'] = (\n>     train\n>     .groupby('user_id')['task_container_id']\n>     .transform(lambda x: pd.factorize(x)[0])\n>     .astype('int16')\n> )\n> ```\n\nI am also interested to know how much applying this fix will impact the training and score so feel free to give your insights in the comments.",
    "1041597": "From @sohier in this other [thread](https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/189351) \n\n> - The `timestamp` column shows when an activity is finished, not when it started. A user could have spent an arbitrary amount of time working on the first problem before the first `timestamp` value would get logged.\n\nTo me this suggests that while `task_container_id` captures the order that a user **first sees** tasks in, `timestamp` captures the order in which a user actually **completes** tasks. A user can start one task, start a second, and then finish the second before finishing the first -- this results in the later timestamp for the first task.",
    "1041603": "I won't have time to dig in to confirm this today, but I believe this is just a data artifact. Imagine someone logging in from both their phone and a tablet; the first batch of questions they started working on (lowest container ID) wouldn't necessarily actually be the first one to finish.\n\nGenerally speaking, time tracked across multiple devices is unlikely to be precisely correct unless you take extreme measures like providing each device [with an atomic clock ](https://www.wired.com/2012/11/google-spanner-time/).",
    "1041604": "I guess that users can choose which groups of questions to work on. One user could do task 1, 2, and 4, then comes back to task 3.",
    "1041588": "I also found this issue. The host should clarify this ...BTW, where is the discussion you mentioned in ` a fix suggested by @maxhalford from the original discussion.` ? Thanks.",
    "1047038": "Thank you for bringing this issue and I apologize for any inconvenience. Indeed, the non-increasing values of `task_container_id` is data artifact due to internal workings of the service, and it serves the purpose of distinguishing different bundles without monotonicity assumption. We took this into account by dropping the original note about monotonic increase."
  }
}