{
  "id": 200919,
  "title": "User Id  and TImestamp",
  "url": "/competitions/riiid-test-answer-prediction/discussion/200919",
  "author_name": "Jaideep",
  "post_date": "2020-12-02T11:44:55.153000",
  "votes": 4,
  "comment_count": 8,
  "views": 0,
  "content": "<p>Timestamp =0 means if m correct ,time from which the user interaction clock starts . </p>\n<p>However there are some doubts i have<br>\nEach user should have Timestamp staring at 0 .  so number of Timestam=0 should be come as Number of unique users. <br>\nBut i think that is not the case</p>\n<pre><code>train_df.user_id.nunique()\n123365\n</code></pre>\n<pre><code>train_df[train_df.timestamp==0]\n124586 rows × 9 columns\n</code></pre>",
  "messages": [
    {
      "id": 1099473,
      "postDate": "2020-12-02T11:44:55.153Z",
      "content": "<p>Timestamp =0 means if m correct ,time from which the user interaction clock starts . </p>\n<p>However there are some doubts i have<br>\nEach user should have Timestamp staring at 0 .  so number of Timestam=0 should be come as Number of unique users. <br>\nBut i think that is not the case</p>\n<pre><code>train_df.user_id.nunique()\n123365\n</code></pre>\n<pre><code>train_df[train_df.timestamp==0]\n124586 rows × 9 columns\n</code></pre>",
      "rawMarkdown": "Timestamp =0 means if m correct ,time from which the user interaction clock starts . \n\nHowever there are some doubts i have\nEach user should have Timestamp staring at 0 .  so number of Timestam=0 should be come as Number of unique users. \nBut i think that is not the case\n```\ntrain_df.user_id.nunique()\n123365\n```\n\n```\ntrain_df[train_df.timestamp==0]\n124586 rows × 9 columns\n```",
      "votes": 4
    },
    {
      "id": 1099622,
      "postDate": "2020-12-02T13:37:13.497Z",
      "content": "<p>As discussed <a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/189465\" target=\"_blank\">here</a>, timestamp is the time when a task (= unique task_container_id) is completed by a user. So <code>train_df[train_df.timestamp==0]</code> selects all rows corresponding to the first completed interactions - some of them are a bunch of questions from a single user.<br>\nBy the way, we don't always have <code>task_container_id==0</code> when <code>timestamp==0</code>, because container ids are allocated when the user first sees a task (from <a href=\"https://www.kaggle.com/aquatic\" target=\"_blank\">@aquatic</a> in the aboved linked thread).</p>",
      "rawMarkdown": "As discussed [here](https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/189465), timestamp is the time when a task (= unique task_container_id) is completed by a user. So `train_df[train_df.timestamp==0]` selects all rows corresponding to the first completed interactions - some of them are a bunch of questions from a single user.\nBy the way, we don't always have `task_container_id==0` when `timestamp==0`, because container ids are allocated when the user first sees a task (from @aquatic in the aboved linked thread).",
      "votes": 1,
      "replies": [
        {
          "id": 1099690,
          "postDate": "2020-12-02T14:35:54.653Z",
          "content": "<p><a href=\"https://www.kaggle.com/johannhuber\" target=\"_blank\">@johannhuber</a>  i checked the discussion it   talks on Task container id  order more .<br>\nHere the doubt is why Not all User ids have Timestamp starting 0 , so in next task or interaction we see the time user took to complete the set of previous tasks.. </p>",
          "rawMarkdown": "@johannhuber  i checked the discussion it   talks on Task container id  order more .\nHere the doubt is why Not all User ids have Timestamp starting 0 , so in next task or interaction we see the time user took to complete the set of previous tasks.. "
        },
        {
          "id": 1099701,
          "postDate": "2020-12-02T14:47:43.387Z",
          "content": "<p>Then I'm not sure to understand your doubt : what do you mean by \"Not all User ids have Timestamp starting 0\" ?<br>\nActually I got :<br>\n<code>train.groupby(['user_id'])['timestamp'].first().unique() # array([0])</code><br>\n<code>train.groupby(['user_id'])['timestamp'].min().unique() # array([0])</code></p>",
          "rawMarkdown": "Then I'm not sure to understand your doubt : what do you mean by \"Not all User ids have Timestamp starting 0\" ?\nActually I got :\n`train.groupby(['user_id'])['timestamp'].first().unique() # array([0])`\n`train.groupby(['user_id'])['timestamp'].min().unique() # array([0])`\n",
          "votes": 1
        },
        {
          "id": 1099714,
          "postDate": "2020-12-02T14:58:44.817Z",
          "content": "<p>you can check my query  above for DF..  number of unique users , number rows with Timestamp==0 should come as same, please see if you getting it same</p>",
          "rawMarkdown": "you can check my query  above for DF..  number of unique users , number rows with Timestamp==0 should come as same, please see if you getting it same"
        },
        {
          "id": 1099836,
          "postDate": "2020-12-02T16:13:31.733Z",
          "content": "<p>If I understand you well, you're assuming : </p>\n<blockquote>\n  <p>(1) There should be : len(train_df.user_id.unique()) == len(train_df[train_df.timestamp==0])</p>\n</blockquote>\n<p>which comes from this assumption : <code>(2) one user's row =&gt; one unique timestamp</code></p>\n<p>Am I right ?</p>\n<p>If so, I'm arguing that (2) isn't correct, because the data actually verify : <code>one unique timestamp =&gt; one unique task_container_id =&gt; one or several rows</code></p>\n<p>Implying that (1) isn't correct too, as there could be timestamp==0 for several rows corresponding to a unique user. For example, the first interaction of a user could be : answering to 3 questions. That's why I mentioned task_container_id :) </p>\n<p>Tell me if we are on the same page !</p>",
          "rawMarkdown": "If I understand you well, you're assuming : \n> (1) There should be : len(train_df.user_id.unique()) == len(train_df[train_df.timestamp==0])\n\nwhich comes from this assumption : `(2) one user's row => one unique timestamp `\n\nAm I right ?\n\nIf so, I'm arguing that (2) isn't correct, because the data actually verify : `one unique timestamp => one unique task_container_id => one or several rows `\n\nImplying that (1) isn't correct too, as there could be timestamp==0 for several rows corresponding to a unique user. For example, the first interaction of a user could be : answering to 3 questions. That's why I mentioned task_container_id :) \n\nTell me if we are on the same page !",
          "votes": 1
        },
        {
          "id": 1100023,
          "postDate": "2020-12-02T19:13:15.477Z",
          "content": "<pre><code>train_df.user_id.nunique() \n393656\n\ntrain_df.groupby('user_id').agg(ts=('timestamp','min'))-393656\ntrain_df.groupby('user_id').agg(ts=('timestamp','min')) .sum() - 0 \n</code></pre>\n<p>So i think this sets assumption right,maybe some records earlier got filtered.</p>",
          "rawMarkdown": "```\ntrain_df.user_id.nunique() \n393656\n\ntrain_df.groupby('user_id').agg(ts=('timestamp','min'))-393656\ntrain_df.groupby('user_id').agg(ts=('timestamp','min')) .sum() - 0 \n\n```\nSo i think this sets assumption right,maybe some records earlier got filtered."
        },
        {
          "id": 1132857,
          "postDate": "2020-12-30T18:21:55.030Z",
          "content": "<blockquote>\n  <p>As discussed <a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/189465\" target=\"_blank\">here</a>, timestamp is the time when a task (= unique task_container_id) is completed by a user. So <code>train_df[train_df.timestamp==0]</code> selects all rows corresponding to the first completed interactions - some of them are a bunch of questions from a single user.<br>\n  By the way, we don't always have <code>task_container_id==0</code> when <code>timestamp==0</code>, because container ids are allocated when the user first sees a task (from <a href=\"https://www.kaggle.com/aquatic\" target=\"_blank\">@aquatic</a> in the aboved linked thread).</p>\n</blockquote>\n<p><a href=\"https://www.kaggle.com/johannhuber\" target=\"_blank\">@johannhuber</a>  after long time.. i have one question.. do we know why do they set timestamp=0 instead of marking it as Elapsed time +0   as per definition of time stamp. </p>\n<p>However Time stamp definition is still confusing when comparing it with actual meaning</p>\n<p><code>It says  Timestamp in milliseconds between this user interaction and time of completion of first even from user</code></p>\n<p>How does this definition maps to actual meaning.<br>\nthis user interaction -current question ?  then which time period are we talking about here.</p>",
          "rawMarkdown": "> As discussed [here](https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/189465), timestamp is the time when a task (= unique task_container_id) is completed by a user. So `train_df[train_df.timestamp==0]` selects all rows corresponding to the first completed interactions - some of them are a bunch of questions from a single user.\n> By the way, we don't always have `task_container_id==0` when `timestamp==0`, because container ids are allocated when the user first sees a task (from @aquatic in the aboved linked thread).\n\n@johannhuber  after long time.. i have one question.. do we know why do they set timestamp=0 instead of marking it as Elapsed time +0   as per definition of time stamp. \n\nHowever Time stamp definition is still confusing when comparing it with actual meaning\n\n`It says  Timestamp in milliseconds between this user interaction and time of completion of first even from user `\n\nHow does this definition maps to actual meaning.\nthis user interaction -current question ?  then which time period are we talking about here."
        }
      ]
    },
    {
      "id": 1133036,
      "postDate": "2020-12-30T21:40:39.200Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 1099622,
      "author_name": "JT-karl",
      "author_url": "",
      "post_date": "2020-12-02T13:37:13.497000",
      "content": "<p>As discussed <a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/189465\" target=\"_blank\">here</a>, timestamp is the time when a task (= unique task_container_id) is completed by a user. So <code>train_df[train_df.timestamp==0]</code> selects all rows corresponding to the first completed interactions - some of them are a bunch of questions from a single user.<br>\nBy the way, we don't always have <code>task_container_id==0</code> when <code>timestamp==0</code>, because container ids are allocated when the user first sees a task (from <a href=\"https://www.kaggle.com/aquatic\" target=\"_blank\">@aquatic</a> in the aboved linked thread).</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1099690,
          "author_name": "Jaideep",
          "author_url": "",
          "post_date": "2020-12-02T14:35:54.653000",
          "content": "<p><a href=\"https://www.kaggle.com/johannhuber\" target=\"_blank\">@johannhuber</a>  i checked the discussion it   talks on Task container id  order more .<br>\nHere the doubt is why Not all User ids have Timestamp starting 0 , so in next task or interaction we see the time user took to complete the set of previous tasks.. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1099701,
          "author_name": "JT-karl",
          "author_url": "",
          "post_date": "2020-12-02T14:47:43.387000",
          "content": "<p>Then I'm not sure to understand your doubt : what do you mean by \"Not all User ids have Timestamp starting 0\" ?<br>\nActually I got :<br>\n<code>train.groupby(['user_id'])['timestamp'].first().unique() # array([0])</code><br>\n<code>train.groupby(['user_id'])['timestamp'].min().unique() # array([0])</code></p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1099714,
          "author_name": "Jaideep",
          "author_url": "",
          "post_date": "2020-12-02T14:58:44.817000",
          "content": "<p>you can check my query  above for DF..  number of unique users , number rows with Timestamp==0 should come as same, please see if you getting it same</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1099836,
          "author_name": "JT-karl",
          "author_url": "",
          "post_date": "2020-12-02T16:13:31.733000",
          "content": "<p>If I understand you well, you're assuming : </p>\n<blockquote>\n  <p>(1) There should be : len(train_df.user_id.unique()) == len(train_df[train_df.timestamp==0])</p>\n</blockquote>\n<p>which comes from this assumption : <code>(2) one user's row =&gt; one unique timestamp</code></p>\n<p>Am I right ?</p>\n<p>If so, I'm arguing that (2) isn't correct, because the data actually verify : <code>one unique timestamp =&gt; one unique task_container_id =&gt; one or several rows</code></p>\n<p>Implying that (1) isn't correct too, as there could be timestamp==0 for several rows corresponding to a unique user. For example, the first interaction of a user could be : answering to 3 questions. That's why I mentioned task_container_id :) </p>\n<p>Tell me if we are on the same page !</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1100023,
          "author_name": "Jaideep",
          "author_url": "",
          "post_date": "2020-12-02T19:13:15.477000",
          "content": "<pre><code>train_df.user_id.nunique() \n393656\n\ntrain_df.groupby('user_id').agg(ts=('timestamp','min'))-393656\ntrain_df.groupby('user_id').agg(ts=('timestamp','min')) .sum() - 0 \n</code></pre>\n<p>So i think this sets assumption right,maybe some records earlier got filtered.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1132857,
          "author_name": "Jaideep",
          "author_url": "",
          "post_date": "2020-12-30T18:21:55.030000",
          "content": "<blockquote>\n  <p>As discussed <a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/189465\" target=\"_blank\">here</a>, timestamp is the time when a task (= unique task_container_id) is completed by a user. So <code>train_df[train_df.timestamp==0]</code> selects all rows corresponding to the first completed interactions - some of them are a bunch of questions from a single user.<br>\n  By the way, we don't always have <code>task_container_id==0</code> when <code>timestamp==0</code>, because container ids are allocated when the user first sees a task (from <a href=\"https://www.kaggle.com/aquatic\" target=\"_blank\">@aquatic</a> in the aboved linked thread).</p>\n</blockquote>\n<p><a href=\"https://www.kaggle.com/johannhuber\" target=\"_blank\">@johannhuber</a>  after long time.. i have one question.. do we know why do they set timestamp=0 instead of marking it as Elapsed time +0   as per definition of time stamp. </p>\n<p>However Time stamp definition is still confusing when comparing it with actual meaning</p>\n<p><code>It says  Timestamp in milliseconds between this user interaction and time of completion of first even from user</code></p>\n<p>How does this definition maps to actual meaning.<br>\nthis user interaction -current question ?  then which time period are we talking about here.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1133036,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-12-30T21:40:39.200000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1099473": "Timestamp =0 means if m correct ,time from which the user interaction clock starts . \n\nHowever there are some doubts i have\nEach user should have Timestamp staring at 0 .  so number of Timestam=0 should be come as Number of unique users. \nBut i think that is not the case\n```\ntrain_df.user_id.nunique()\n123365\n```\n\n```\ntrain_df[train_df.timestamp==0]\n124586 rows × 9 columns\n```",
    "1099622": "As discussed [here](https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/189465), timestamp is the time when a task (= unique task_container_id) is completed by a user. So `train_df[train_df.timestamp==0]` selects all rows corresponding to the first completed interactions - some of them are a bunch of questions from a single user.\nBy the way, we don't always have `task_container_id==0` when `timestamp==0`, because container ids are allocated when the user first sees a task (from @aquatic in the aboved linked thread).",
    "1133036": ""
  }
}