{
  "id": 205515,
  "title": "Temporal Features Discrepancies ",
  "url": "/competitions/riiid-test-answer-prediction/discussion/205515",
  "author_name": "",
  "post_date": "2020-12-20T14:07:30.035428800Z",
  "votes": 10,
  "comment_count": 7,
  "views": 0,
  "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F933480%2Fc0e4f503358f71fa75e4958d7b0ca5c9%2FScreen%20Shot%202020-12-20%20at%207.56.21%20AM.png?generation=1608473002212557&amp;alt=media\" alt=\"\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F933480%2Fd5de5988313f08d3787ebc95a557ef8b%2FScreen%20Shot%202020-12-20%20at%207.56.30%20AM.png?generation=1608473014102214&amp;alt=media\" alt=\"\"> </p>\n<p>When I attempt to generate these features from the dataset, the distributions look n-o-t-h-i-n-g like this. They are MUCH longer tailed, with less prominent peaks, even after doing the //1000 and //60000. I assume the units of timestamp = milliseconds. Are others experiencing the same thing? I am being careful to take into consideration the difference between lectures and questions for these features, e.g. lag time is forward filled onto questions before removing exercises.</p>",
  "messages": [
    {
      "id": "1119981",
      "postDate": "12/20/2020 14:07:30",
      "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F933480%2Fc0e4f503358f71fa75e4958d7b0ca5c9%2FScreen%20Shot%202020-12-20%20at%207.56.21%20AM.png?generation=1608473002212557&amp;alt=media\" alt=\"\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F933480%2Fd5de5988313f08d3787ebc95a557ef8b%2FScreen%20Shot%202020-12-20%20at%207.56.30%20AM.png?generation=1608473014102214&amp;alt=media\" alt=\"\"> </p>\n<p>When I attempt to generate these features from the dataset, the distributions look n-o-t-h-i-n-g like this. They are MUCH longer tailed, with less prominent peaks, even after doing the //1000 and //60000. I assume the units of timestamp = milliseconds. Are others experiencing the same thing? I am being careful to take into consideration the difference between lectures and questions for these features, e.g. lag time is forward filled onto questions before removing exercises.</p>",
      "rawMarkdown": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F933480%2Fc0e4f503358f71fa75e4958d7b0ca5c9%2FScreen%20Shot%202020-12-20%20at%207.56.21%20AM.png?generation=1608473002212557&alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F933480%2Fd5de5988313f08d3787ebc95a557ef8b%2FScreen%20Shot%202020-12-20%20at%207.56.30%20AM.png?generation=1608473014102214&alt=media) \n\nWhen I attempt to generate these features from the dataset, the distributions look n-o-t-h-i-n-g like this. They are MUCH longer tailed, with less prominent peaks, even after doing the //1000 and //60000. I assume the units of timestamp = milliseconds. Are others experiencing the same thing? I am being careful to take into consideration the difference between lectures and questions for these features, e.g. lag time is forward filled onto questions before removing exercises.",
      "votes": null
    },
    {
      "id": "1120060",
      "postDate": "12/20/2020 15:07:19",
      "content": "<p>This is what my distributions look like on aa sample of 1M Q's when I exclude ==0 and ==300 because there are a lot of values peaked there:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F933480%2F26cb2a65491d90592abe853510185c29%2Fmydistros.png?generation=1608476837488455&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "This is what my distributions look like on aa sample of 1M Q's when I exclude ==0 and ==300 because there are a lot of values peaked there:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F933480%2F26cb2a65491d90592abe853510185c29%2Fmydistros.png?generation=1608476837488455&alt=media)",
      "votes": null
    },
    {
      "id": "1120119",
      "postDate": "12/20/2020 15:23:34",
      "content": "<p>I just realized something. In the SAINT+ paper, elapsed time is the delta between when a question is answered and when the prompt of the question is first shown. This is evidenced not by the text of the paper, but by the very figure I shared above. It has nothing to do with lectures whatsoever, and furthermore, is <strong>not calculable</strong> given the granularity of the dataset we currently have. A few people have mentioned they implemented SAINT+ <code>as-is</code> while reporting CV/LB scores—did you guys miss this, or did you guys have different interpretations of the above features? Similarly, lag time, as defined by the paper would be the delta between the completion of one question and when the next question's prompt is first shown. We don't have the later, so again, this feature in it's SAINT+ form isn't computable given our dataset.</p>\n<p>Timestamp being defined as <code>the time in milliseconds between this user interaction and the first event completion from that user</code>, which starts at 0, the best we can do is calculate deltas between question completions.</p>",
      "rawMarkdown": "I just realized something. In the SAINT+ paper, elapsed time is the delta between when a question is answered and when the prompt of the question is first shown. This is evidenced not by the text of the paper, but by the very figure I shared above. It has nothing to do with lectures whatsoever, and furthermore, is **not calculable** given the granularity of the dataset we currently have. A few people have mentioned they implemented SAINT+ `as-is` while reporting CV/LB scores—did you guys miss this, or did you guys have different interpretations of the above features? Similarly, lag time, as defined by the paper would be the delta between the completion of one question and when the next question's prompt is first shown. We don't have the later, so again, this feature in it's SAINT+ form isn't computable given our dataset.\n\nTimestamp being defined as `the time in milliseconds between this user interaction and the first event completion from that user`, which starts at 0, the best we can do is calculate deltas between question completions.",
      "votes": null
    },
    {
      "id": "1120180",
      "postDate": "12/20/2020 16:19:27",
      "content": "<blockquote>\n  <p>Timestamp being defined as the time in milliseconds between this user interaction and the first event completion from that user, which starts at 0…</p>\n</blockquote>\n<p>What we are given for each user is the first TS (let's say reg time) subtracted from each TS further (future interactions) is my current understanding.</p>",
      "rawMarkdown": ">Timestamp being defined as the time in milliseconds between this user interaction and the first event completion from that user, which starts at 0...\n\nWhat we are given for each user is the first TS (let's say reg time) subtracted from each TS further (future interactions) is my current understanding.",
      "votes": null
    },
    {
      "id": "1120225",
      "postDate": "12/20/2020 17:19:24",
      "content": "<p>The 'un-calculable' of lag time as exactly in the paper is one reason I hesitated for so long to do it, but finally I just decided to do something similar (so lag time in some sense, but not the same as in the paper).</p>",
      "rawMarkdown": "The 'un-calculable' of lag time as exactly in the paper is one reason I hesitated for so long to do it, but finally I just decided to do something similar (so lag time in some sense, but not the same as in the paper).",
      "votes": null
    },
    {
      "id": "1120250",
      "postDate": "12/20/2020 17:35:14",
      "content": "<p><a href=\"https://www.kaggle.com/yihdarshieh\" target=\"_blank\">@yihdarshieh</a> Thank you for the confirmation. If you don't mind me prying: in your implementation, do you discretize and embed your lag variable—or do you transform it with, e.g., a linear layer?</p>",
      "rawMarkdown": "yihdarshieh Thank you for the confirmation. If you don't mind me prying: in your implementation, do you discretize and embed your lag variable—or do you transform it with, e.g., a linear layer?",
      "votes": null
    },
    {
      "id": "1120701",
      "postDate": "12/21/2020 04:09:04",
      "content": "<p>I have not read original paper, but the distribution of elapsed time you show seems to be different from that of prior_question_elapsed_time in <a href=\"https://www.kaggle.com/isaienkov/riiid-answer-correctness-prediction-eda-modeling\" target=\"_blank\">this EDA notebook</a>. <br>\nThe distribution in EDA seems to be similar to that of Fig. 3 you show.<br>\nCorrect me if I am wrong.</p>",
      "rawMarkdown": "I have not read original paper, but the distribution of elapsed time you show seems to be different from that of prior_question_elapsed_time in [this EDA notebook](https://www.kaggle.com/isaienkov/riiid-answer-correctness-prediction-eda-modeling). \nThe distribution in EDA seems to be similar to that of Fig. 3 you show.\nCorrect me if I am wrong.",
      "votes": null
    },
    {
      "id": "1132802",
      "postDate": "12/30/2020 17:40:26",
      "content": "<p><a href=\"https://www.kaggle.com/yihdarshieh\" target=\"_blank\">@yihdarshieh</a> as per above time diagram.. <br>\nI am unable to understand how we can calculate the Lag time which i think currently kinda approximated as  Ts2-Ts1 </p>\n<p>The we can calculate for the current container a feature: prior_container_lag_time = prior_container_ts_diff - prior_container_elapsed_time * prior_container_num_items</p>\n<p>below is my understanding looking at diagram above</p>\n<pre><code>E1- start time T1 - We dont know\nR1 (Response submitted time) - TIme-T2 - We know it as TS1\n\nLagtime Lt1- To calculate\n\nE2-Start TIme T3- We dont know\n\nR2-Stop Time T4 - We know TS2\n</code></pre>",
      "rawMarkdown": "yihdarshieh as per above time diagram.. \nI am unable to understand how we can calculate the Lag time which i think currently kinda approximated as  Ts2-Ts1 \n\nThe we can calculate for the current container a feature: prior_container_lag_time = prior_container_ts_diff - prior_container_elapsed_time * prior_container_num_items\n\nbelow is my understanding looking at diagram above\n\n```\n \nE1- start time T1 - We dont know\nR1 (Response submitted time) - TIme-T2 - We know it as TS1\n\nLagtime Lt1- To calculate\n\nE2-Start TIme T3- We dont know\n\nR2-Stop Time T4 - We know TS2\n\n\n```",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1120060,
      "author_name": "authman",
      "author_url": "",
      "post_date": "12/20/2020 15:07:19",
      "content": "<p>This is what my distributions look like on aa sample of 1M Q's when I exclude ==0 and ==300 because there are a lot of values peaked there:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F933480%2F26cb2a65491d90592abe853510185c29%2Fmydistros.png?generation=1608476837488455&amp;alt=media\" alt=\"\"></p>",
      "votes": null,
      "replies": [
        {
          "id": 1120701,
          "author_name": "tomooinubushi",
          "author_url": "",
          "post_date": "12/21/2020 04:09:04",
          "content": "<p>I have not read original paper, but the distribution of elapsed time you show seems to be different from that of prior_question_elapsed_time in <a href=\"https://www.kaggle.com/isaienkov/riiid-answer-correctness-prediction-eda-modeling\" target=\"_blank\">this EDA notebook</a>. <br>\nThe distribution in EDA seems to be similar to that of Fig. 3 you show.<br>\nCorrect me if I am wrong.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1120119,
      "author_name": "authman",
      "author_url": "",
      "post_date": "12/20/2020 15:23:34",
      "content": "<p>I just realized something. In the SAINT+ paper, elapsed time is the delta between when a question is answered and when the prompt of the question is first shown. This is evidenced not by the text of the paper, but by the very figure I shared above. It has nothing to do with lectures whatsoever, and furthermore, is <strong>not calculable</strong> given the granularity of the dataset we currently have. A few people have mentioned they implemented SAINT+ <code>as-is</code> while reporting CV/LB scores—did you guys miss this, or did you guys have different interpretations of the above features? Similarly, lag time, as defined by the paper would be the delta between the completion of one question and when the next question's prompt is first shown. We don't have the later, so again, this feature in it's SAINT+ form isn't computable given our dataset.</p>\n<p>Timestamp being defined as <code>the time in milliseconds between this user interaction and the first event completion from that user</code>, which starts at 0, the best we can do is calculate deltas between question completions.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1120180,
          "author_name": "adityaecdrid",
          "author_url": "",
          "post_date": "12/20/2020 16:19:27",
          "content": "<blockquote>\n  <p>Timestamp being defined as the time in milliseconds between this user interaction and the first event completion from that user, which starts at 0…</p>\n</blockquote>\n<p>What we are given for each user is the first TS (let's say reg time) subtracted from each TS further (future interactions) is my current understanding.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1120225,
          "author_name": "yihdarshieh",
          "author_url": "",
          "post_date": "12/20/2020 17:19:24",
          "content": "<p>The 'un-calculable' of lag time as exactly in the paper is one reason I hesitated for so long to do it, but finally I just decided to do something similar (so lag time in some sense, but not the same as in the paper).</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1120250,
          "author_name": "authman",
          "author_url": "",
          "post_date": "12/20/2020 17:35:14",
          "content": "<p><a href=\"https://www.kaggle.com/yihdarshieh\" target=\"_blank\">@yihdarshieh</a> Thank you for the confirmation. If you don't mind me prying: in your implementation, do you discretize and embed your lag variable—or do you transform it with, e.g., a linear layer?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1132802,
          "author_name": "jaideepvalani",
          "author_url": "",
          "post_date": "12/30/2020 17:40:26",
          "content": "<p><a href=\"https://www.kaggle.com/yihdarshieh\" target=\"_blank\">@yihdarshieh</a> as per above time diagram.. <br>\nI am unable to understand how we can calculate the Lag time which i think currently kinda approximated as  Ts2-Ts1 </p>\n<p>The we can calculate for the current container a feature: prior_container_lag_time = prior_container_ts_diff - prior_container_elapsed_time * prior_container_num_items</p>\n<p>below is my understanding looking at diagram above</p>\n<pre><code>E1- start time T1 - We dont know\nR1 (Response submitted time) - TIme-T2 - We know it as TS1\n\nLagtime Lt1- To calculate\n\nE2-Start TIme T3- We dont know\n\nR2-Stop Time T4 - We know TS2\n</code></pre>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1119981": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F933480%2Fc0e4f503358f71fa75e4958d7b0ca5c9%2FScreen%20Shot%202020-12-20%20at%207.56.21%20AM.png?generation=1608473002212557&alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F933480%2Fd5de5988313f08d3787ebc95a557ef8b%2FScreen%20Shot%202020-12-20%20at%207.56.30%20AM.png?generation=1608473014102214&alt=media) \n\nWhen I attempt to generate these features from the dataset, the distributions look n-o-t-h-i-n-g like this. They are MUCH longer tailed, with less prominent peaks, even after doing the //1000 and //60000. I assume the units of timestamp = milliseconds. Are others experiencing the same thing? I am being careful to take into consideration the difference between lectures and questions for these features, e.g. lag time is forward filled onto questions before removing exercises.",
    "1120060": "This is what my distributions look like on aa sample of 1M Q's when I exclude ==0 and ==300 because there are a lot of values peaked there:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F933480%2F26cb2a65491d90592abe853510185c29%2Fmydistros.png?generation=1608476837488455&alt=media)",
    "1120119": "I just realized something. In the SAINT+ paper, elapsed time is the delta between when a question is answered and when the prompt of the question is first shown. This is evidenced not by the text of the paper, but by the very figure I shared above. It has nothing to do with lectures whatsoever, and furthermore, is **not calculable** given the granularity of the dataset we currently have. A few people have mentioned they implemented SAINT+ `as-is` while reporting CV/LB scores—did you guys miss this, or did you guys have different interpretations of the above features? Similarly, lag time, as defined by the paper would be the delta between the completion of one question and when the next question's prompt is first shown. We don't have the later, so again, this feature in it's SAINT+ form isn't computable given our dataset.\n\nTimestamp being defined as `the time in milliseconds between this user interaction and the first event completion from that user`, which starts at 0, the best we can do is calculate deltas between question completions.",
    "1120180": ">Timestamp being defined as the time in milliseconds between this user interaction and the first event completion from that user, which starts at 0...\n\nWhat we are given for each user is the first TS (let's say reg time) subtracted from each TS further (future interactions) is my current understanding.",
    "1120225": "The 'un-calculable' of lag time as exactly in the paper is one reason I hesitated for so long to do it, but finally I just decided to do something similar (so lag time in some sense, but not the same as in the paper).",
    "1120250": "yihdarshieh Thank you for the confirmation. If you don't mind me prying: in your implementation, do you discretize and embed your lag variable—or do you transform it with, e.g., a linear layer?",
    "1120701": "I have not read original paper, but the distribution of elapsed time you show seems to be different from that of prior_question_elapsed_time in [this EDA notebook](https://www.kaggle.com/isaienkov/riiid-answer-correctness-prediction-eda-modeling). \nThe distribution in EDA seems to be similar to that of Fig. 3 you show.\nCorrect me if I am wrong.",
    "1132802": "yihdarshieh as per above time diagram.. \nI am unable to understand how we can calculate the Lag time which i think currently kinda approximated as  Ts2-Ts1 \n\nThe we can calculate for the current container a feature: prior_container_lag_time = prior_container_ts_diff - prior_container_elapsed_time * prior_container_num_items\n\nbelow is my understanding looking at diagram above\n\n```\n \nE1- start time T1 - We dont know\nR1 (Response submitted time) - TIme-T2 - We know it as TS1\n\nLagtime Lt1- To calculate\n\nE2-Start TIme T3- We dont know\n\nR2-Stop Time T4 - We know TS2\n\n\n```"
  },
  "source": "meta"
}