{
  "id": 191019,
  "title": "Can we find user sessions?",
  "url": "/competitions/riiid-test-answer-prediction/discussion/191019",
  "author_name": "Aditya Soni",
  "post_date": "2020-10-14T07:43:28.983000",
  "votes": 14,
  "comment_count": 17,
  "views": 0,
  "content": "<p>Hi All,</p>\n<p>So i was trying to find user sessions as to when the student has given the exam. If we will plot the row_id for let's say some user_id \"k\",  it will look like this,</p>\n<p>(non-zoomed variant)</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F835774%2F4251140a07e20d4e5954d99f4e384384%2FScreenshot%202020-10-14%20at%201.11.34%20PM.png?generation=1602661329817583&amp;alt=media\" alt=\"zommed_userid_4421282\"></p>\n<p>(i have zoomed in slightly at the start)</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F835774%2F579fe96788af0f8eb5723a82eb124da2%2FScreenshot%202020-10-14%20at%201.10.29%20PM.png?generation=1602661290414276&amp;alt=media\" alt=\"org_userid_4421282\"></p>\n<p>So, can you notice those sharp kinks? What do people think about those and how can we possibly leverage this out as well?</p>\n<pre><code># for RRP\nimport plotly.express as px \nfig = px.line(df_train[df_train[\"user_id\"] == 4421282], x='row_id', y=\"timestamp\")\nfig.show()\n</code></pre>\n<p>Thanks!</p>\n<p>PS I might be over thinking, so take it lightly.</p>",
  "messages": [
    {
      "id": 1049196,
      "postDate": "2020-10-14T07:43:28.983Z",
      "content": "<p>Hi All,</p>\n<p>So i was trying to find user sessions as to when the student has given the exam. If we will plot the row_id for let's say some user_id \"k\",  it will look like this,</p>\n<p>(non-zoomed variant)</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F835774%2F4251140a07e20d4e5954d99f4e384384%2FScreenshot%202020-10-14%20at%201.11.34%20PM.png?generation=1602661329817583&amp;alt=media\" alt=\"zommed_userid_4421282\"></p>\n<p>(i have zoomed in slightly at the start)</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F835774%2F579fe96788af0f8eb5723a82eb124da2%2FScreenshot%202020-10-14%20at%201.10.29%20PM.png?generation=1602661290414276&amp;alt=media\" alt=\"org_userid_4421282\"></p>\n<p>So, can you notice those sharp kinks? What do people think about those and how can we possibly leverage this out as well?</p>\n<pre><code># for RRP\nimport plotly.express as px \nfig = px.line(df_train[df_train[\"user_id\"] == 4421282], x='row_id', y=\"timestamp\")\nfig.show()\n</code></pre>\n<p>Thanks!</p>\n<p>PS I might be over thinking, so take it lightly.</p>",
      "rawMarkdown": "Hi All,\n\nSo i was trying to find user sessions as to when the student has given the exam. If we will plot the row_id for let's say some user_id \"k\",  it will look like this,\n\n(non-zoomed variant)\n\n![zommed_userid_4421282](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F835774%2F4251140a07e20d4e5954d99f4e384384%2FScreenshot%202020-10-14%20at%201.11.34%20PM.png?generation=1602661329817583&alt=media)\n\n(i have zoomed in slightly at the start)\n\n![org_userid_4421282](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F835774%2F579fe96788af0f8eb5723a82eb124da2%2FScreenshot%202020-10-14%20at%201.10.29%20PM.png?generation=1602661290414276&alt=media)\n\nSo, can you notice those sharp kinks? What do people think about those and how can we possibly leverage this out as well?\n\n```\n# for RRP\nimport plotly.express as px \nfig = px.line(df_train[df_train[\"user_id\"] == 4421282], x='row_id', y=\"timestamp\")\nfig.show()\n```\n\nThanks!\n\nPS I might be over thinking, so take it lightly.",
      "votes": 14
    },
    {
      "id": 1050482,
      "postDate": "2020-10-15T13:00:46.363Z",
      "content": "<p>Grouping interactions into sessions is very subjective in nature. The most common way to create sessions is by using an idle time cutoff. It's simplistic, not perfect but works surprisingly well (based on my past experience in ad tech). The choice of cutoff depends on the data / company / industry / use-case.</p>\n<p>eg: Google Analytics uses a cutoff of 30mins to define session - <a href=\"https://support.google.com/analytics/answer/2731565\" target=\"_blank\">https://support.google.com/analytics/answer/2731565</a></p>",
      "rawMarkdown": "Grouping interactions into sessions is very subjective in nature. The most common way to create sessions is by using an idle time cutoff. It's simplistic, not perfect but works surprisingly well (based on my past experience in ad tech). The choice of cutoff depends on the data / company / industry / use-case.\n\neg: Google Analytics uses a cutoff of 30mins to define session - https://support.google.com/analytics/answer/2731565",
      "votes": 6,
      "replies": [
        {
          "id": 1050488,
          "postDate": "2020-10-15T13:06:48.787Z",
          "content": "<p>Thanks Rohan, I/We would love if you can help and share more info as to how it's done in real time based on your past experiences!<br>\nIn our case, the average time one should ideally spent on any question is &lt;=300secs. </p>\n<p>My current idea is to \"groupby nearby time-stamps\" to mock the session so to say. But then again doing it for test set is another challenge :zip:</p>",
          "rawMarkdown": "Thanks Rohan, I/We would love if you can help and share more info as to how it's done in real time based on your past experiences!\nIn our case, the average time one should ideally spent on any question is <=300secs. \n\nMy current idea is to \"groupby nearby time-stamps\" to mock the session so to say. But then again doing it for test set is another challenge :zip:"
        },
        {
          "id": 1050785,
          "postDate": "2020-10-15T18:14:57.173Z",
          "content": "<p>You can iterate through a user’s interaction chronologically and whenever the time gap is &gt; chosen cutoff, increase session number by 1.</p>\n<p>You can do the same for test set as well but need to be careful about edge cases like if a user is split across batches, the sessions will be split as well even if it doesn’t exceed the cutoff or else you might even have incomplete session information due to batches.</p>",
          "rawMarkdown": "You can iterate through a user’s interaction chronologically and whenever the time gap is > chosen cutoff, increase session number by 1.\n\nYou can do the same for test set as well but need to be careful about edge cases like if a user is split across batches, the sessions will be split as well even if it doesn’t exceed the cutoff or else you might even have incomplete session information due to batches.",
          "votes": 2
        }
      ]
    },
    {
      "id": 1049725,
      "postDate": "2020-10-14T18:11:19.020Z",
      "content": "<p>To me, those just seem like times between play sessions. Some are small, like interruptions within a single setting. Others are longer breaks for even up to weeks. From what I've understood so far of the test set, we shoooould be able to generate this sort of information, since the data is grouped by blocks of users for performance sake. Not sure what benefit it brings though. Maybe if the user hasn't tested in a while, their performance might suffer?</p>",
      "rawMarkdown": "To me, those just seem like times between play sessions. Some are small, like interruptions within a single setting. Others are longer breaks for even up to weeks. From what I've understood so far of the test set, we shoooould be able to generate this sort of information, since the data is grouped by blocks of users for performance sake. Not sure what benefit it brings though. Maybe if the user hasn't tested in a while, their performance might suffer?",
      "votes": 2,
      "replies": [
        {
          "id": 1049731,
          "postDate": "2020-10-14T18:24:50.640Z",
          "content": "<p>I have run a notebook where I calculated the time since the previous action using the timestamp. Prelim results indicate that it was a somewhat important feature (more important that most other user features I have tried to generate). However, I could not perform the same calculation on the test data and received a submission error every time I tried after the notebook ran for 9 hours. There is probably a quicker and better way to do things that captures similar information.</p>",
          "rawMarkdown": "I have run a notebook where I calculated the time since the previous action using the timestamp. Prelim results indicate that it was a somewhat important feature (more important that most other user features I have tried to generate). However, I could not perform the same calculation on the test data and received a submission error every time I tried after the notebook ran for 9 hours. There is probably a quicker and better way to do things that captures similar information.",
          "votes": 2
        },
        {
          "id": 1049746,
          "postDate": "2020-10-14T18:43:26.750Z",
          "content": "<blockquote>\n  <p>Not sure what benefit it brings though. Maybe if the user hasn't tested in a while, their performance might suffer?</p>\n</blockquote>\n<p>Well the performance is likely to take a hit, that's why i want to kinda look out for such scenarios. Assuming me as a student, if i haven't practised the exam prep questions for quite a while, i am pretty sure my scores are going to take a hit (in my next prep_test etc) and it's likely that i will spend little more time on videos/looking at the solution for sure, right? (i am yet to validate this) At-least in real world scenario, it would be the case with many.  And then you have users who just appear for like 30-40 rows and then they vanish out (at least in training data). Thinking some sort of decay should help but if we don't get the user for quite some time, we need to ensure that we decay their stat accordingly to their next show-up on our radar (from test point of view).</p>\n<p>In case you aren't aware authman, we do have access to test-set labels it seems after we have made the prediction for the previous batch. (in my understanding as of now, <a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/190748\" target=\"_blank\">here</a>).</p>\n<blockquote>\n  <p>I could not perform the same calculation on the test data and received a submission error every time I tried after the notebook ran for 9 hours. There is probably a quicker and better way to do things that captures similar information.</p>\n</blockquote>\n<p>Well that's a huge pain point in this comp. The pipeline for test set is not going to be easy for sure to write as many things can happen here and we can only make \"assumptions\" about them. One thing will fail and your whole pipeline will be doomed pretty much here. Which i feel is a huge drawback. Imagine a single nan is sufficient to break your pipeline. You have to account for few things for sure and we cannot really do a shift here and calculate them as well. We have to map the new seq to the users last timestamp we have, if not, then that's a new user for sure.</p>",
          "rawMarkdown": "> Not sure what benefit it brings though. Maybe if the user hasn't tested in a while, their performance might suffer?\n\nWell the performance is likely to take a hit, that's why i want to kinda look out for such scenarios. Assuming me as a student, if i haven't practised the exam prep questions for quite a while, i am pretty sure my scores are going to take a hit (in my next prep_test etc) and it's likely that i will spend little more time on videos/looking at the solution for sure, right? (i am yet to validate this) At-least in real world scenario, it would be the case with many.  And then you have users who just appear for like 30-40 rows and then they vanish out (at least in training data). Thinking some sort of decay should help but if we don't get the user for quite some time, we need to ensure that we decay their stat accordingly to their next show-up on our radar (from test point of view).\n\nIn case you aren't aware authman, we do have access to test-set labels it seems after we have made the prediction for the previous batch. (in my understanding as of now, [here](https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/190748)).\n\n\n>I could not perform the same calculation on the test data and received a submission error every time I tried after the notebook ran for 9 hours. There is probably a quicker and better way to do things that captures similar information.\n\nWell that's a huge pain point in this comp. The pipeline for test set is not going to be easy for sure to write as many things can happen here and we can only make \"assumptions\" about them. One thing will fail and your whole pipeline will be doomed pretty much here. Which i feel is a huge drawback. Imagine a single nan is sufficient to break your pipeline. You have to account for few things for sure and we cannot really do a shift here and calculate them as well. We have to map the new seq to the users last timestamp we have, if not, then that's a new user for sure.",
          "votes": 1
        },
        {
          "id": 1049820,
          "postDate": "2020-10-14T19:56:03.970Z",
          "content": "<p>Thank you for your explorative work and reports <a href=\"https://www.kaggle.com/dwit392\" target=\"_blank\">@dwit392</a>. Based on this <a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/191106\" target=\"_blank\">fresh information</a>, I think it's saafe to say we should be able to do this now for test set.</p>\n<blockquote>\n  <p>The test data follows chronologically after the train data. The test iterations give interactions of users chronologically.</p>\n</blockquote>",
          "rawMarkdown": "Thank you for your explorative work and reports @dwit392. Based on this [fresh information](https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/191106), I think it's saafe to say we should be able to do this now for test set.\n\n> The test data follows chronologically after the train data. The test iterations give interactions of users chronologically.",
          "votes": 1
        },
        {
          "id": 1049884,
          "postDate": "2020-10-14T21:21:49.487Z",
          "content": "<p>I may as well post my Py code that received a submission scoring error after 9 hours of running the notebook to see if someone sees something I am doing something wrong or a better way to do things.</p>\n<p>time_df contains only the last record for each user and just the user_id, content_id and timestamp</p>\n<p>The only thing I can think of that would actually be wrong is maybe I need an (ignore_index = True) in the first concat function</p>\n<pre>time_df = pd.concat([time_df, test_df[['user_id', 'content_id', 'timestamp']]])\n\ntime_df = time_df.sort_values(['user_id','timestamp'])\n\ntime_df['prev_timestamp'] = time_df.timestamp.shift(1)\n\ntime_df['time_since_last_action'] = time_df['timestamp'] - time_df['prev_timestamp']\n\ntime_df.time_since_last_action.replace(to_replace=0, method = 'ffill', inplace = True)\n\ntime_df.loc[time_df.time_since_last_action &lt; 0, 'time_since_last_action'] = np.nan\n\ntime_df.time_since_last_action.fillna(time_since_median, inplace = True)\n\ntest_df = pd.merge(test_df, time_df, on=['user_id', 'content_id', 'timestamp'], how = \"left\")\n</pre>",
          "rawMarkdown": "I may as well post my Py code that received a submission scoring error after 9 hours of running the notebook to see if someone sees something I am doing something wrong or a better way to do things.\n\ntime_df contains only the last record for each user and just the user_id, content_id and timestamp\n\nThe only thing I can think of that would actually be wrong is maybe I need an (ignore_index = True) in the first concat function\n\n<pre>\ntime_df = pd.concat([time_df, test_df[['user_id', 'content_id', 'timestamp']]])\n\ntime_df = time_df.sort_values(['user_id','timestamp'])\n\ntime_df['prev_timestamp'] = time_df.timestamp.shift(1)\n\ntime_df['time_since_last_action'] = time_df['timestamp'] - time_df['prev_timestamp']\n\ntime_df.time_since_last_action.replace(to_replace=0, method = 'ffill', inplace = True)\n\ntime_df.loc[time_df.time_since_last_action < 0, 'time_since_last_action'] = np.nan\n\ntime_df.time_since_last_action.fillna(time_since_median, inplace = True)\n\ntest_df = pd.merge(test_df, time_df, on=['user_id', 'content_id', 'timestamp'], how = \"left\")\n</pre>",
          "votes": 2
        },
        {
          "id": 1049932,
          "postDate": "2020-10-14T22:37:53.913Z",
          "content": "<p>Don't know how to make your code work for now. But I think the logic of <code>time_df.timestamp.shift(1)</code> won't work for bundle questions.  At least you should fix for those situations.  </p>",
          "rawMarkdown": "Don't know how to make your code work for now. But I think the logic of `time_df.timestamp.shift(1)` won't work for bundle questions.  At least you should fix for those situations.  ",
          "votes": 2
        },
        {
          "id": 1049938,
          "postDate": "2020-10-14T22:41:11.157Z",
          "content": "<p>Admittedly, I should have commented more, but I saved that for the training part. That's why I replace 0's. As I mentioned, this did actually appear to be an OK feature in my training.</p>\n<p>EDIT: I finally realized why this logic doesn't make sense for test data, but I still don't know why my submission doesn't work</p>",
          "rawMarkdown": "Admittedly, I should have commented more, but I saved that for the training part. That's why I replace 0's. As I mentioned, this did actually appear to be an OK feature in my training.\n\nEDIT: I finally realized why this logic doesn't make sense for test data, but I still don't know why my submission doesn't work"
        },
        {
          "id": 1050158,
          "postDate": "2020-10-15T06:01:10.610Z",
          "content": "<p>deleted_placeholder</p>",
          "rawMarkdown": "deleted_placeholder"
        },
        {
          "id": 1050665,
          "postDate": "2020-10-15T15:57:05.200Z",
          "content": "<blockquote>\n  <p>EDIT: I finally realized why this logic doesn't make sense for test data, but I still don't know why my submission doesn't work</p>\n</blockquote>\n<p>Great! If you want to do this, you will have to maintain the state of the user first i think and then you can do it. (which is not an easy thing to do/achieve 😅)</p>",
          "rawMarkdown": ">EDIT: I finally realized why this logic doesn't make sense for test data, but I still don't know why my submission doesn't work\n\nGreat! If you want to do this, you will have to maintain the state of the user first i think and then you can do it. (which is not an easy thing to do/achieve 😅)"
        },
        {
          "id": 1050948,
          "postDate": "2020-10-16T00:00:33.403Z",
          "content": "<p>It took me 3 hours to write a function that I think actually works. Hopefully a decent score and no submission error in 9 hours 😬</p>",
          "rawMarkdown": "It took me 3 hours to write a function that I think actually works. Hopefully a decent score and no submission error in 9 hours 😬",
          "votes": 1
        },
        {
          "id": 1050976,
          "postDate": "2020-10-16T01:57:44.057Z",
          "content": "<blockquote>\n  <p>It took me 3 hours to write a function that I think actually works. Hopefully a decent score and no submission error in 9 hours 😬</p>\n</blockquote>\n<p>Congratulations 🎊🎊🎊🎊</p>",
          "rawMarkdown": ">It took me 3 hours to write a function that I think actually works. Hopefully a decent score and no submission error in 9 hours 😬\n\nCongratulations 🎊🎊🎊🎊"
        },
        {
          "id": 1051246,
          "postDate": "2020-10-16T09:55:08.590Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 1051690,
          "postDate": "2020-10-16T18:24:48.770Z",
          "content": "<p>Here was the function I used to calculate time since last action for the test data. It proved to be a decent feature (best one I have currently other than the obvious user accuracy history and question accuracy history). I am sure there is a better one out there that captures a similar idea, but this function could also be applied to other things regarding users' \"state.\" I will warn this was computationally expensive and took 4-5 hours to submit.</p>\n<p>Also, just for reference last_record is a dataframe with just the user_id and last timestamp for each user.</p>\n<p><a href=\"https://www.kaggle.com/dwit392/riiid-challenge-time-since-last-action-for-test\" target=\"_blank\">https://www.kaggle.com/dwit392/riiid-challenge-time-since-last-action-for-test</a></p>",
          "rawMarkdown": "Here was the function I used to calculate time since last action for the test data. It proved to be a decent feature (best one I have currently other than the obvious user accuracy history and question accuracy history). I am sure there is a better one out there that captures a similar idea, but this function could also be applied to other things regarding users' \"state.\" I will warn this was computationally expensive and took 4-5 hours to submit.\n\nAlso, just for reference last_record is a dataframe with just the user_id and last timestamp for each user.\n\nhttps://www.kaggle.com/dwit392/riiid-challenge-time-since-last-action-for-test",
          "votes": 1
        }
      ]
    },
    {
      "id": 1049772,
      "postDate": "2020-10-14T19:06:00.650Z",
      "content": "<p>Nor sure whether we can use some sort of slope as it's will be sensitive to the scaling then and might not be the correct thing as well. (Sharp peaks will considerably have higher yaxis value as compared to X axis value's as by trigo, arctan of a big value is almost ~80+)</p>",
      "rawMarkdown": "Nor sure whether we can use some sort of slope as it's will be sensitive to the scaling then and might not be the correct thing as well. (Sharp peaks will considerably have higher yaxis value as compared to X axis value's as by trigo, arctan of a big value is almost ~80+)"
    }
  ],
  "comments": [
    {
      "id": 1050482,
      "author_name": "Vopani",
      "author_url": "",
      "post_date": "2020-10-15T13:00:46.363000",
      "content": "<p>Grouping interactions into sessions is very subjective in nature. The most common way to create sessions is by using an idle time cutoff. It's simplistic, not perfect but works surprisingly well (based on my past experience in ad tech). The choice of cutoff depends on the data / company / industry / use-case.</p>\n<p>eg: Google Analytics uses a cutoff of 30mins to define session - <a href=\"https://support.google.com/analytics/answer/2731565\" target=\"_blank\">https://support.google.com/analytics/answer/2731565</a></p>",
      "votes": 6,
      "replies": [
        {
          "id": 1050488,
          "author_name": "Aditya Soni",
          "author_url": "",
          "post_date": "2020-10-15T13:06:48.787000",
          "content": "<p>Thanks Rohan, I/We would love if you can help and share more info as to how it's done in real time based on your past experiences!<br>\nIn our case, the average time one should ideally spent on any question is &lt;=300secs. </p>\n<p>My current idea is to \"groupby nearby time-stamps\" to mock the session so to say. But then again doing it for test set is another challenge :zip:</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1050785,
          "author_name": "Vopani",
          "author_url": "",
          "post_date": "2020-10-15T18:14:57.173000",
          "content": "<p>You can iterate through a user’s interaction chronologically and whenever the time gap is &gt; chosen cutoff, increase session number by 1.</p>\n<p>You can do the same for test set as well but need to be careful about edge cases like if a user is split across batches, the sessions will be split as well even if it doesn’t exceed the cutoff or else you might even have incomplete session information due to batches.</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 1049725,
      "author_name": "عثمان",
      "author_url": "",
      "post_date": "2020-10-14T18:11:19.020000",
      "content": "<p>To me, those just seem like times between play sessions. Some are small, like interruptions within a single setting. Others are longer breaks for even up to weeks. From what I've understood so far of the test set, we shoooould be able to generate this sort of information, since the data is grouped by blocks of users for performance sake. Not sure what benefit it brings though. Maybe if the user hasn't tested in a while, their performance might suffer?</p>",
      "votes": 2,
      "replies": [
        {
          "id": 1049731,
          "author_name": "David Witten",
          "author_url": "",
          "post_date": "2020-10-14T18:24:50.640000",
          "content": "<p>I have run a notebook where I calculated the time since the previous action using the timestamp. Prelim results indicate that it was a somewhat important feature (more important that most other user features I have tried to generate). However, I could not perform the same calculation on the test data and received a submission error every time I tried after the notebook ran for 9 hours. There is probably a quicker and better way to do things that captures similar information.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1049746,
          "author_name": "Aditya Soni",
          "author_url": "",
          "post_date": "2020-10-14T18:43:26.750000",
          "content": "<blockquote>\n  <p>Not sure what benefit it brings though. Maybe if the user hasn't tested in a while, their performance might suffer?</p>\n</blockquote>\n<p>Well the performance is likely to take a hit, that's why i want to kinda look out for such scenarios. Assuming me as a student, if i haven't practised the exam prep questions for quite a while, i am pretty sure my scores are going to take a hit (in my next prep_test etc) and it's likely that i will spend little more time on videos/looking at the solution for sure, right? (i am yet to validate this) At-least in real world scenario, it would be the case with many.  And then you have users who just appear for like 30-40 rows and then they vanish out (at least in training data). Thinking some sort of decay should help but if we don't get the user for quite some time, we need to ensure that we decay their stat accordingly to their next show-up on our radar (from test point of view).</p>\n<p>In case you aren't aware authman, we do have access to test-set labels it seems after we have made the prediction for the previous batch. (in my understanding as of now, <a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/190748\" target=\"_blank\">here</a>).</p>\n<blockquote>\n  <p>I could not perform the same calculation on the test data and received a submission error every time I tried after the notebook ran for 9 hours. There is probably a quicker and better way to do things that captures similar information.</p>\n</blockquote>\n<p>Well that's a huge pain point in this comp. The pipeline for test set is not going to be easy for sure to write as many things can happen here and we can only make \"assumptions\" about them. One thing will fail and your whole pipeline will be doomed pretty much here. Which i feel is a huge drawback. Imagine a single nan is sufficient to break your pipeline. You have to account for few things for sure and we cannot really do a shift here and calculate them as well. We have to map the new seq to the users last timestamp we have, if not, then that's a new user for sure.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1049820,
          "author_name": "عثمان",
          "author_url": "",
          "post_date": "2020-10-14T19:56:03.970000",
          "content": "<p>Thank you for your explorative work and reports <a href=\"https://www.kaggle.com/dwit392\" target=\"_blank\">@dwit392</a>. Based on this <a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/191106\" target=\"_blank\">fresh information</a>, I think it's saafe to say we should be able to do this now for test set.</p>\n<blockquote>\n  <p>The test data follows chronologically after the train data. The test iterations give interactions of users chronologically.</p>\n</blockquote>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1049884,
          "author_name": "David Witten",
          "author_url": "",
          "post_date": "2020-10-14T21:21:49.487000",
          "content": "<p>I may as well post my Py code that received a submission scoring error after 9 hours of running the notebook to see if someone sees something I am doing something wrong or a better way to do things.</p>\n<p>time_df contains only the last record for each user and just the user_id, content_id and timestamp</p>\n<p>The only thing I can think of that would actually be wrong is maybe I need an (ignore_index = True) in the first concat function</p>\n<pre>time_df = pd.concat([time_df, test_df[['user_id', 'content_id', 'timestamp']]])\n\ntime_df = time_df.sort_values(['user_id','timestamp'])\n\ntime_df['prev_timestamp'] = time_df.timestamp.shift(1)\n\ntime_df['time_since_last_action'] = time_df['timestamp'] - time_df['prev_timestamp']\n\ntime_df.time_since_last_action.replace(to_replace=0, method = 'ffill', inplace = True)\n\ntime_df.loc[time_df.time_since_last_action &lt; 0, 'time_since_last_action'] = np.nan\n\ntime_df.time_since_last_action.fillna(time_since_median, inplace = True)\n\ntest_df = pd.merge(test_df, time_df, on=['user_id', 'content_id', 'timestamp'], how = \"left\")\n</pre>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1049932,
          "author_name": "Ning Jia",
          "author_url": "",
          "post_date": "2020-10-14T22:37:53.913000",
          "content": "<p>Don't know how to make your code work for now. But I think the logic of <code>time_df.timestamp.shift(1)</code> won't work for bundle questions.  At least you should fix for those situations.  </p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1049938,
          "author_name": "David Witten",
          "author_url": "",
          "post_date": "2020-10-14T22:41:11.157000",
          "content": "<p>Admittedly, I should have commented more, but I saved that for the training part. That's why I replace 0's. As I mentioned, this did actually appear to be an OK feature in my training.</p>\n<p>EDIT: I finally realized why this logic doesn't make sense for test data, but I still don't know why my submission doesn't work</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1050158,
          "author_name": "Aditya Soni",
          "author_url": "",
          "post_date": "2020-10-15T06:01:10.610000",
          "content": "<p>deleted_placeholder</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1050665,
          "author_name": "Aditya Soni",
          "author_url": "",
          "post_date": "2020-10-15T15:57:05.200000",
          "content": "<blockquote>\n  <p>EDIT: I finally realized why this logic doesn't make sense for test data, but I still don't know why my submission doesn't work</p>\n</blockquote>\n<p>Great! If you want to do this, you will have to maintain the state of the user first i think and then you can do it. (which is not an easy thing to do/achieve 😅)</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1050948,
          "author_name": "David Witten",
          "author_url": "",
          "post_date": "2020-10-16T00:00:33.403000",
          "content": "<p>It took me 3 hours to write a function that I think actually works. Hopefully a decent score and no submission error in 9 hours 😬</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1050976,
          "author_name": "Aditya Soni",
          "author_url": "",
          "post_date": "2020-10-16T01:57:44.057000",
          "content": "<blockquote>\n  <p>It took me 3 hours to write a function that I think actually works. Hopefully a decent score and no submission error in 9 hours 😬</p>\n</blockquote>\n<p>Congratulations 🎊🎊🎊🎊</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1051246,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-10-16T09:55:08.590000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1051690,
          "author_name": "David Witten",
          "author_url": "",
          "post_date": "2020-10-16T18:24:48.770000",
          "content": "<p>Here was the function I used to calculate time since last action for the test data. It proved to be a decent feature (best one I have currently other than the obvious user accuracy history and question accuracy history). I am sure there is a better one out there that captures a similar idea, but this function could also be applied to other things regarding users' \"state.\" I will warn this was computationally expensive and took 4-5 hours to submit.</p>\n<p>Also, just for reference last_record is a dataframe with just the user_id and last timestamp for each user.</p>\n<p><a href=\"https://www.kaggle.com/dwit392/riiid-challenge-time-since-last-action-for-test\" target=\"_blank\">https://www.kaggle.com/dwit392/riiid-challenge-time-since-last-action-for-test</a></p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1049772,
      "author_name": "Aditya Soni",
      "author_url": "",
      "post_date": "2020-10-14T19:06:00.650000",
      "content": "<p>Nor sure whether we can use some sort of slope as it's will be sensitive to the scaling then and might not be the correct thing as well. (Sharp peaks will considerably have higher yaxis value as compared to X axis value's as by trigo, arctan of a big value is almost ~80+)</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1049196": "Hi All,\n\nSo i was trying to find user sessions as to when the student has given the exam. If we will plot the row_id for let's say some user_id \"k\",  it will look like this,\n\n(non-zoomed variant)\n\n![zommed_userid_4421282](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F835774%2F4251140a07e20d4e5954d99f4e384384%2FScreenshot%202020-10-14%20at%201.11.34%20PM.png?generation=1602661329817583&alt=media)\n\n(i have zoomed in slightly at the start)\n\n![org_userid_4421282](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F835774%2F579fe96788af0f8eb5723a82eb124da2%2FScreenshot%202020-10-14%20at%201.10.29%20PM.png?generation=1602661290414276&alt=media)\n\nSo, can you notice those sharp kinks? What do people think about those and how can we possibly leverage this out as well?\n\n```\n# for RRP\nimport plotly.express as px \nfig = px.line(df_train[df_train[\"user_id\"] == 4421282], x='row_id', y=\"timestamp\")\nfig.show()\n```\n\nThanks!\n\nPS I might be over thinking, so take it lightly.",
    "1050482": "Grouping interactions into sessions is very subjective in nature. The most common way to create sessions is by using an idle time cutoff. It's simplistic, not perfect but works surprisingly well (based on my past experience in ad tech). The choice of cutoff depends on the data / company / industry / use-case.\n\neg: Google Analytics uses a cutoff of 30mins to define session - https://support.google.com/analytics/answer/2731565",
    "1049725": "To me, those just seem like times between play sessions. Some are small, like interruptions within a single setting. Others are longer breaks for even up to weeks. From what I've understood so far of the test set, we shoooould be able to generate this sort of information, since the data is grouped by blocks of users for performance sake. Not sure what benefit it brings though. Maybe if the user hasn't tested in a while, their performance might suffer?",
    "1049772": "Nor sure whether we can use some sort of slope as it's will be sensitive to the scaling then and might not be the correct thing as well. (Sharp peaks will considerably have higher yaxis value as compared to X axis value's as by trigo, arctan of a big value is almost ~80+)"
  }
}