{
  "id": 194266,
  "title": "An interesting Pattern detected for Repeated Questions",
  "url": "/competitions/riiid-test-answer-prediction/discussion/194266",
  "author_name": "Aravind P",
  "post_date": "2020-10-31T16:31:06.069000",
  "votes": 141,
  "comment_count": 61,
  "views": 0,
  "content": "<p>I have noticed that users are given the same questions multiple times during the course of their learning. The following table shows a question repeated 9 times for a user.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F560450%2F574463e460ee3d42c675690db772ea64%2FScreenshot%202020-10-31%20at%209.45.58%20PM.png?generation=1604161013985800&amp;alt=media\" alt=\"\"></p>\n<p>After the above analysis, my hypothesis was that the probability of a user answering a question correctly given that the question has already seen before is higher than the probability of the user answering the question correctly in the first attempt.<br>\nTo prove this I have plotted the probability distribution of average answer correctness at a user level for two populations(<strong>First attempt Vs Repeated Attempt</strong>). The following figure shows the results<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F560450%2Ff4d8b0e53df65d3877f6ec2d31b00fac%2FScreenshot%202020-10-31%20at%209.51.03%20PM.png?generation=1604161298744094&amp;alt=media\" alt=\"\"><br>\nI have also done a welch's t-test to prove that these two distributions are indeed different.</p>\n<p>For the detailed analysis, please refer to the notebook I have created <br>\n<a href=\"https://www.kaggle.com/aravindpadman/riiid-statistical-analysis-on-repeated-questions\" target=\"_blank\">https://www.kaggle.com/aravindpadman/riiid-statistical-analysis-on-repeated-questions</a></p>\n<p>I am currently trying to incorporate this idea into my modeling, but I am not yet there.</p>\n<p>Let me know your valuable ideas in the comment.</p>",
  "messages": [
    {
      "id": 1065724,
      "postDate": "2020-10-31T16:31:06.070Z",
      "content": "<p>I have noticed that users are given the same questions multiple times during the course of their learning. The following table shows a question repeated 9 times for a user.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F560450%2F574463e460ee3d42c675690db772ea64%2FScreenshot%202020-10-31%20at%209.45.58%20PM.png?generation=1604161013985800&amp;alt=media\" alt=\"\"></p>\n<p>After the above analysis, my hypothesis was that the probability of a user answering a question correctly given that the question has already seen before is higher than the probability of the user answering the question correctly in the first attempt.<br>\nTo prove this I have plotted the probability distribution of average answer correctness at a user level for two populations(<strong>First attempt Vs Repeated Attempt</strong>). The following figure shows the results<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F560450%2Ff4d8b0e53df65d3877f6ec2d31b00fac%2FScreenshot%202020-10-31%20at%209.51.03%20PM.png?generation=1604161298744094&amp;alt=media\" alt=\"\"><br>\nI have also done a welch's t-test to prove that these two distributions are indeed different.</p>\n<p>For the detailed analysis, please refer to the notebook I have created <br>\n<a href=\"https://www.kaggle.com/aravindpadman/riiid-statistical-analysis-on-repeated-questions\" target=\"_blank\">https://www.kaggle.com/aravindpadman/riiid-statistical-analysis-on-repeated-questions</a></p>\n<p>I am currently trying to incorporate this idea into my modeling, but I am not yet there.</p>\n<p>Let me know your valuable ideas in the comment.</p>",
      "rawMarkdown": "I have noticed that users are given the same questions multiple times during the course of their learning. The following table shows a question repeated 9 times for a user.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F560450%2F574463e460ee3d42c675690db772ea64%2FScreenshot%202020-10-31%20at%209.45.58%20PM.png?generation=1604161013985800&alt=media)\n\n\nAfter the above analysis, my hypothesis was that the probability of a user answering a question correctly given that the question has already seen before is higher than the probability of the user answering the question correctly in the first attempt.\nTo prove this I have plotted the probability distribution of average answer correctness at a user level for two populations(**First attempt Vs Repeated Attempt**). The following figure shows the results\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F560450%2Ff4d8b0e53df65d3877f6ec2d31b00fac%2FScreenshot%202020-10-31%20at%209.51.03%20PM.png?generation=1604161298744094&alt=media)\nI have also done a welch's t-test to prove that these two distributions are indeed different.\n\nFor the detailed analysis, please refer to the notebook I have created \nhttps://www.kaggle.com/aravindpadman/riiid-statistical-analysis-on-repeated-questions\n\nI am currently trying to incorporate this idea into my modeling, but I am not yet there.\n\nLet me know your valuable ideas in the comment.",
      "votes": 139
    },
    {
      "id": 1065729,
      "postDate": "2020-10-31T16:44:53.630Z",
      "content": "<p>This is a really nice find <a href=\"https://www.kaggle.com/aravindpadman\" target=\"_blank\">@aravindpadman</a> . Nice!</p>",
      "rawMarkdown": "This is a really nice find @aravindpadman . Nice!",
      "votes": 3
    },
    {
      "id": 1067024,
      "postDate": "2020-11-02T09:59:01.307Z",
      "content": "<p>That's an interesting finding and that seems to me logical. <br>\nI'll add this feature to my model, not sure yet if only a categorical column (first attempt, repeated attempt) will be enough; the number of attempts could provide information too. I would say that more the user see the question and more the user is likely to answer correctly</p>",
      "rawMarkdown": "That's an interesting finding and that seems to me logical. \nI'll add this feature to my model, not sure yet if only a categorical column (first attempt, repeated attempt) will be enough; the number of attempts could provide information too. I would say that more the user see the question and more the user is likely to answer correctly",
      "votes": 4,
      "replies": [
        {
          "id": 1067290,
          "postDate": "2020-11-02T13:00:12.690Z",
          "content": "<p><code>I would say that more the user see the question and more the user is likely to answer correctly</code><br>\nTrue. I have actually observed this trend during my analysis.</p>",
          "rawMarkdown": "```I would say that more the user see the question and more the user is likely to answer correctly```\nTrue. I have actually observed this trend during my analysis.\n",
          "votes": 3
        },
        {
          "id": 1067361,
          "postDate": "2020-11-02T13:49:24.473Z",
          "content": "<p>Simply adding a boolean column for first attempt does boost my validation (+ 0.006)<br>\nHowever I don't know yet how to make it work during inference without reaching 9 hours limit</p>",
          "rawMarkdown": "Simply adding a boolean column for first attempt does boost my validation (+ 0.006)\nHowever I don't know yet how to make it work during inference without reaching 9 hours limit",
          "votes": 3
        },
        {
          "id": 1087141,
          "postDate": "2020-11-22T11:47:00.473Z",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/alexj21\" target=\"_blank\">@alexj21</a> , how do you store the repeated-questions-information of the training set? I stored it as a dict and the key is (user, question) pair, but it's too large to kaggle kernel so that I can't even load it during inference.</p>",
          "rawMarkdown": "Hi @alexj21 , how do you store the repeated-questions-information of the training set? I stored it as a dict and the key is (user, question) pair, but it's too large to kaggle kernel so that I can't even load it during inference."
        },
        {
          "id": 1088059,
          "postDate": "2020-11-23T09:28:32.770Z",
          "content": "<p>I stored it as a dict in which the key is the user_id and the value a bitarray of length 15k. Each bit is assigned to a content ID. If the user saw a given content id then its bit is set to 1. <br>\nYou can find some code examples in the discussion bellow. It works pretty well :) </p>",
          "rawMarkdown": "I stored it as a dict in which the key is the user_id and the value a bitarray of length 15k. Each bit is assigned to a content ID. If the user saw a given content id then its bit is set to 1. \nYou can find some code examples in the discussion bellow. It works pretty well :) \n",
          "votes": 4
        },
        {
          "id": 1088159,
          "postDate": "2020-11-23T11:41:10.500Z",
          "content": "<p>Thanks, really smart method!</p>",
          "rawMarkdown": "Thanks, really smart method!",
          "votes": 1
        },
        {
          "id": 1088224,
          "postDate": "2020-11-23T12:55:01.930Z",
          "content": "<p>The idea was firstly proposed by <a href=\"https://www.kaggle.com/adityaecdrid\" target=\"_blank\">@adityaecdrid</a> </p>",
          "rawMarkdown": "The idea was firstly proposed by @adityaecdrid ",
          "votes": 2
        }
      ]
    },
    {
      "id": 1065809,
      "postDate": "2020-10-31T19:37:22.447Z",
      "content": "<p>That is cool. You could include the cumulative count or sum of <code>answered_correctly</code> by <code>content_id</code> as a single column feature. Here is some sql that should work.</p>\n<pre><code>SELECT user_id, task_container_id, row_id, content_id, COUNT(answered_correctly)\n  OVER (\n    PARTITION BY user_id, content_id\n    ORDER BY task_container_id\n    RANGE BETWEEN UNBOUNDED PRECEDING AND 1 PRECEDING\n  ) calc\nFROM train\nORDER BY user_id, content_id, task_container_id, row_id\n</code></pre>",
      "rawMarkdown": "That is cool. You could include the cumulative count or sum of `answered_correctly` by `content_id` as a single column feature. Here is some sql that should work.\n\n```\nSELECT user_id, task_container_id, row_id, content_id, COUNT(answered_correctly)\n  OVER (\n    PARTITION BY user_id, content_id\n    ORDER BY task_container_id\n    RANGE BETWEEN UNBOUNDED PRECEDING AND 1 PRECEDING\n  ) calc\nFROM train\nORDER BY user_id, content_id, task_container_id, row_id\n```  ",
      "votes": 4,
      "replies": [
        {
          "id": 1065957,
          "postDate": "2020-11-01T05:30:32.577Z",
          "content": "<p>Thank you <a href=\"https://www.kaggle.com/calebeverett\" target=\"_blank\">@calebeverett</a>. I will try this out</p>",
          "rawMarkdown": "Thank you @calebeverett. I will try this out",
          "votes": 1
        },
        {
          "id": 1066739,
          "postDate": "2020-11-02T01:32:03.870Z",
          "content": "<p>Did you use SQL to complete feature engineering? Very cool!!!</p>",
          "rawMarkdown": "Did you use SQL to complete feature engineering? Very cool!!!",
          "votes": 1
        },
        {
          "id": 1066747,
          "postDate": "2020-11-02T01:40:14.140Z",
          "content": "<p>I've got a cool set up going with BigQuery. I was going to share it once I get the prediction pipeline done.</p>",
          "rawMarkdown": "I've got a cool set up going with BigQuery. I was going to share it once I get the prediction pipeline done.",
          "votes": 2
        },
        {
          "id": 1069174,
          "postDate": "2020-11-04T07:09:18.227Z",
          "content": "<p>I got my BigQuery set up cleaned up enough to share.</p>\n<p><a href=\"https://www.kaggle.com/calebeverett/riiid-bigquery-xgboost-end-to-end/\" target=\"_blank\">https://www.kaggle.com/calebeverett/riiid-bigquery-xgboost-end-to-end/</a></p>",
          "rawMarkdown": "I got my BigQuery set up cleaned up enough to share.\n\nhttps://www.kaggle.com/calebeverett/riiid-bigquery-xgboost-end-to-end/",
          "votes": 7
        },
        {
          "id": 1069698,
          "postDate": "2020-11-04T19:17:46.737Z",
          "content": "<p><a href=\"https://www.kaggle.com/calebeverett\" target=\"_blank\">@calebeverett</a> neat!</p>",
          "rawMarkdown": "@calebeverett neat!",
          "votes": 2
        }
      ]
    },
    {
      "id": 1093974,
      "postDate": "2020-11-28T07:27:56.933Z",
      "content": "<p>The data is described as being the result of a \"smart\" training process.  If it's smart than repeated questions should occur in a couple of cases that I can think about.</p>\n<ol>\n<li>Very difficult question (overall low percentage of correct answers for the full set of users).  You would likely want to repeat a question several times to make sure that the user really does know the answer - vs just a lucky guess the first time asked.</li>\n<li>User got it wrong - need to repeat until they get it right.</li>\n</ol>\n<p>But at some point you might stop asking the question after one or more correct responses.  </p>\n<p>So your plots make lots of sense. </p>",
      "rawMarkdown": "The data is described as being the result of a \"smart\" training process.  If it's smart than repeated questions should occur in a couple of cases that I can think about.\n\n1.  Very difficult question (overall low percentage of correct answers for the full set of users).  You would likely want to repeat a question several times to make sure that the user really does know the answer - vs just a lucky guess the first time asked.\n2.  User got it wrong - need to repeat until they get it right.\n\nBut at some point you might stop asking the question after one or more correct responses.  \n\nSo your plots make lots of sense. ",
      "votes": 1
    },
    {
      "id": 1073219,
      "postDate": "2020-11-09T09:39:01.347Z",
      "content": "<p>Great find!</p>",
      "rawMarkdown": "Great find!",
      "votes": 1
    },
    {
      "id": 1070388,
      "postDate": "2020-11-05T17:47:24.140Z",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/aravindpadman\" target=\"_blank\">@aravindpadman</a>. Nice work. That could be useful! Thanks for sharing!</p>",
      "rawMarkdown": "Hi @aravindpadman. Nice work. That could be useful! Thanks for sharing!",
      "votes": 1
    },
    {
      "id": 1070105,
      "postDate": "2020-11-05T11:07:18.170Z",
      "content": "<p>Woah!! Good work there.</p>",
      "rawMarkdown": "Woah!! Good work there.",
      "votes": 1
    },
    {
      "id": 1069484,
      "postDate": "2020-11-04T14:10:31.990Z",
      "content": "<p>Nice work. That could be usefull! Thanks for sharing!</p>",
      "rawMarkdown": "Nice work. That could be usefull! Thanks for sharing!",
      "votes": 1
    },
    {
      "id": 1069253,
      "postDate": "2020-11-04T08:57:48.223Z",
      "content": "<p>I tried to use this in my model(lgbm) and it looks promising and my validation score did go up. <br>\nI used this code to find attempt number.</p>\n<p>train[\"attempt_no\"] =1<br>\ntrain[\"attemtp_no\"] = train[[\"user_id\",\"content_id\",\"attempt_no\"]].groupby([\"user_id\",\"content_id\"])[\"attempt_no\"].cumsum()</p>\n<p>but the problem is that i can not do same in test data as test data comes in batch and we can not see wether all the content_id of a user are in same batch. So it will give inaccurte results.</p>\n<p>So my question is. Is there any way to resolve this issue?</p>",
      "rawMarkdown": "I tried to use this in my model(lgbm) and it looks promising and my validation score did go up. \nI used this code to find attempt number.\n\ntrain[\"attempt_no\"] =1\ntrain[\"attemtp_no\"] = train[[\"user_id\",\"content_id\",\"attempt_no\"]].groupby([\"user_id\",\"content_id\"])[\"attempt_no\"].cumsum()\n\nbut the problem is that i can not do same in test data as test data comes in batch and we can not see wether all the content_id of a user are in same batch. So it will give inaccurte results.\n\nSo my question is. Is there any way to resolve this issue?",
      "votes": 1,
      "replies": [
        {
          "id": 1069291,
          "postDate": "2020-11-04T09:35:36.083Z",
          "content": "<p>I'm experiencing the same issue.<br>\nThe first solution that came to my mind would be to keep your train dataframe and update your test set on the fly with features from the train set. But this is a quite expensive solution. Plus you need to handle unseen users</p>",
          "rawMarkdown": "I'm experiencing the same issue.\nThe first solution that came to my mind would be to keep your train dataframe and update your test set on the fly with features from the train set. But this is a quite expensive solution. Plus you need to handle unseen users"
        },
        {
          "id": 1069347,
          "postDate": "2020-11-04T10:40:02.540Z",
          "content": "<blockquote>\n  <p>i can not do same in test data as test data comes in batch</p>\n</blockquote>\n<p>You have to keep a hash_map for any given user at any given point. If we re-see a particular content_id, then we can use the old attempt_no for that question directly to start with and later as we move ahead, we can re-think whether we want to update that old_no with the new_max number possible.</p>\n<p>And if one wants to efficient with little more efforts than the usual, use something like a bit-array as your data structure e.g. (to store which all content_ids a user has seen already) and it's easy to update them as well on FLY. We can use <a href=\"https://docs.python.org/3/library/array.html\" target=\"_blank\">array</a> instead of the list as it's mem efficient to cache nos..</p>\n<p>All the prime numbers less than 18 can be written as,</p>\n<pre><code>bit 31                           bit 0\n    |                              |\n    00000000000000100010100010101100 \n</code></pre>\n<p>Prime no's are (2, 3, 5, 7, 11, 13, 17)  (and it should be efficient in terms of memory and speed as well) (32 bits is just 4B) <strong>(we might be complicating things)</strong></p>\n<pre><code>from bitarray import bitarray\na = bitarray(31, endian='little') # important to setup a fixed endianess.\na.setall(False)\n\nfor i in [2, 3, 5, 7, 11, 13, 17]:\n    a[i] = 1\n\nprint(a) # bitarray('0011010100010100010000000000000')\n</code></pre>\n<p>To store, 14k content_ids, we will need around 1814B for each user. (rough calculations, not apt precisely)</p>\n<p>In case someone's looking for a team member, let me know! Ty! (PS Just my idea, i haven't implemented it (yet) and it can be extended to other stats about users if they are just numbers by shifting up bits etc)</p>",
          "rawMarkdown": ">i can not do same in test data as test data comes in batch\n\nYou have to keep a hash_map for any given user at any given point. If we re-see a particular content_id, then we can use the old attempt_no for that question directly to start with and later as we move ahead, we can re-think whether we want to update that old_no with the new_max number possible.\n\nAnd if one wants to efficient with little more efforts than the usual, use something like a bit-array as your data structure e.g. (to store which all content_ids a user has seen already) and it's easy to update them as well on FLY. We can use [array](https://docs.python.org/3/library/array.html) instead of the list as it's mem efficient to cache nos..\n\nAll the prime numbers less than 18 can be written as,\n\n```\nbit 31                           bit 0\n    |                              |\n    00000000000000100010100010101100 \n```\n\nPrime no's are (2, 3, 5, 7, 11, 13, 17)  (and it should be efficient in terms of memory and speed as well) (32 bits is just 4B) **(we might be complicating things)**\n\n```\nfrom bitarray import bitarray\na = bitarray(31, endian='little') # important to setup a fixed endianess.\na.setall(False)\n\nfor i in [2, 3, 5, 7, 11, 13, 17]:\n    a[i] = 1\n \nprint(a) # bitarray('0011010100010100010000000000000')\n```\n\nTo store, 14k content_ids, we will need around 1814B for each user. (rough calculations, not apt precisely)\n\nIn case someone's looking for a team member, let me know! Ty! (PS Just my idea, i haven't implemented it (yet) and it can be extended to other stats about users if they are just numbers by shifting up bits etc)",
          "votes": 14
        },
        {
          "id": 1069550,
          "postDate": "2020-11-04T15:35:38.167Z",
          "content": "<p>Does anybody know if the competition api gets loaded on to the gpu and then iterates through batches or do batches have to get moved on as the data is iterated through from the cpu?</p>",
          "rawMarkdown": "Does anybody know if the competition api gets loaded on to the gpu and then iterates through batches or do batches have to get moved on as the data is iterated through from the cpu?",
          "votes": 1
        },
        {
          "id": 1069551,
          "postDate": "2020-11-04T15:41:04.757Z",
          "content": "<p>We have to move data to GPU for sure, by default it shouldn't, right?</p>",
          "rawMarkdown": "We have to move data to GPU for sure, by default it shouldn't, right?",
          "votes": 1
        },
        {
          "id": 1069562,
          "postDate": "2020-11-04T15:50:55.460Z",
          "content": "<p>I haven't tried it yet, but was thinking that the speedup from running inference on the gpu would be a lot faster if the competition api data loaded there before serving batches, otherwise you have the overhead of moving a couple thousand small batches on.</p>",
          "rawMarkdown": "I haven't tried it yet, but was thinking that the speedup from running inference on the gpu would be a lot faster if the competition api data loaded there before serving batches, otherwise you have the overhead of moving a couple thousand small batches on.",
          "votes": 1
        },
        {
          "id": 1069604,
          "postDate": "2020-11-04T17:06:28.583Z",
          "content": "<p>Th competition api serves tuples of pandas dataframes so that obviously means they are on the cpu.</p>",
          "rawMarkdown": "Th competition api serves tuples of pandas dataframes so that obviously means they are on the cpu.",
          "votes": 2
        },
        {
          "id": 1069605,
          "postDate": "2020-11-04T17:06:33.523Z",
          "content": "<p>Yes that overhead is there. Plus lot of code optimisations needed with not a single failure case as it will break the pipeline and we cannot afford that. Also they cannot move data on GPU before hand as we might be doing one thing or the another on it as well. Hence, it's better leave it at the user-side as to how they want to use the information..</p>",
          "rawMarkdown": "Yes that overhead is there. Plus lot of code optimisations needed with not a single failure case as it will break the pipeline and we cannot afford that. Also they cannot move data on GPU before hand as we might be doing one thing or the another on it as well. Hence, it's better leave it at the user-side as to how they want to use the information..",
          "votes": 2
        },
        {
          "id": 1070009,
          "postDate": "2020-11-05T08:17:22.590Z",
          "content": "<p><a href=\"https://www.kaggle.com/adityaecdrid\" target=\"_blank\">@adityaecdrid</a> thanks for providing some ideas. So if I sum-up the first solution would be to create a dictionary for each user where the key would be the content id and the pair the number of attempts ?<br>\nThe second solution looks faster indeed but I'm not sure to understand the link with prime numbers. Do you suggest to assign a bit signifying first attempt for each prime number along the bit array?  They are 11 prime numbers that can fit into a 32-bit array so for 14k contents you need 14K/11 = ~1272 bytes is that right ?  </p>",
          "rawMarkdown": "@adityaecdrid thanks for providing some ideas. So if I sum-up the first solution would be to create a dictionary for each user where the key would be the content id and the pair the number of attempts ?\nThe second solution looks faster indeed but I'm not sure to understand the link with prime numbers. Do you suggest to assign a bit signifying first attempt for each prime number along the bit array?  They are 11 prime numbers that can fit into a 32-bit array so for 14k contents you need 14K/11 = ~1272 bytes is that right ?  ",
          "votes": 1
        },
        {
          "id": 1070048,
          "postDate": "2020-11-05T09:04:28.733Z",
          "content": "<p>We know that any user can at max have 14k contents possible (max content_id is ~14k). So i can represent these information easily by using a bit vector of length 14k where every position signifies whether the user has seen this particular content_id or not.</p>\n<p>e.g. let's say user_id 115 has seen content_ids \"2, 3, 5, 7, 11, 13, 17\" till now, we can represent that information easily like this,</p>\n<pre><code>bit 31                           bit 0\n    |                              |\n    00000000000000100010100010101100 \n</code></pre>\n<p>And that's all for the training set let's say. </p>\n<p>Now when it comes to test set, we can easily add a new_content_id the user hasn't seen in the past but sees it now. I am adding 19 as an e.g. to the same. So the above bit vector becomes,</p>\n<pre><code>bit 31                           bit 0\n    |                              |\n    00000000000010100010100010101100 \n</code></pre>\n<p>i.e. <code>a[19] = 1</code>; Cool, right?</p>\n<p>And now let's say we have received another content_id as 2, now whether we have seen this content_id or not, can easily be deduced. One simple check is <code>a[2] == 1</code> ?</p>\n<p>Hope it's clear now Alex.</p>",
          "rawMarkdown": "We know that any user can at max have 14k contents possible (max content_id is ~14k). So i can represent these information easily by using a bit vector of length 14k where every position signifies whether the user has seen this particular content_id or not.\n\ne.g. let's say user_id 115 has seen content_ids \"2, 3, 5, 7, 11, 13, 17\" till now, we can represent that information easily like this,\n\n```\nbit 31                           bit 0\n    |                              |\n    00000000000000100010100010101100 \n```\n\nAnd that's all for the training set let's say. \n\nNow when it comes to test set, we can easily add a new_content_id the user hasn't seen in the past but sees it now. I am adding 19 as an e.g. to the same. So the above bit vector becomes,\n\n```\nbit 31                           bit 0\n    |                              |\n    00000000000010100010100010101100 \n```\n\ni.e. `a[19] = 1`; Cool, right?\n\nAnd now let's say we have received another content_id as 2, now whether we have seen this content_id or not, can easily be deduced. One simple check is `a[2] == 1` ?\n\nHope it's clear now Alex.",
          "votes": 7
        },
        {
          "id": 1070050,
          "postDate": "2020-11-05T09:18:12.437Z",
          "content": "<p>Yeah it's clear now ! I was lost with the prime numbers, I though there was something like a subtle combination with content id. </p>",
          "rawMarkdown": "Yeah it's clear now ! I was lost with the prime numbers, I though there was something like a subtle combination with content id. ",
          "votes": 2
        },
        {
          "id": 1070370,
          "postDate": "2020-11-05T17:11:40.820Z",
          "content": "<p>try this <a href=\"https://www.kaggle.com/alexj21\" target=\"_blank\">@alexj21</a>,</p>\n<pre><code>from collections import defaultdict\nfrom bitarray import bitarray\n# tqdm.pandas()\n\n# ONLY FOR NON_LECTURE_ROWS\n\nall_users_past_content_ids = {}\n\ndef already_seen_content_id_in_past(user_id:int, content_id:int) -&gt; int :\n    if user_id not in all_users_past_content_ids:\n        obj = bitarray(13550, endian='little')\n        all_users_past_content_ids[user_id] = obj # &lt;edit here&gt;\n        all_users_past_content_ids[user_id].setall(0)\n    if all_users_past_content_ids[user_id][content_id] == 1:\n        # the bit was already set before, so a repeated content\n        return 1\n    # set the content_id bit if seeing for the first time.\n    all_users_past_content_ids[user_id][content_id] = 1\n    # set it up for the first time, so we return a 0 here.\n    return 0\n\ndf_train[\"repeated_content\"] = df_train[[\"user_id\", \"content_id\"]].progress_apply(lambda row: already_seen_content_id_in_past(row[\"user_id\"], row[\"content_id\"]), axis=1)\n\n# 2000K rows in 24 secs, little slow :(\n# NB plz test it's correctness before ctrl+c,v :)\n</code></pre>",
          "rawMarkdown": "try this @alexj21,\n\n```\n\nfrom collections import defaultdict\nfrom bitarray import bitarray\n# tqdm.pandas()\n\n# ONLY FOR NON_LECTURE_ROWS\n\nall_users_past_content_ids = {}\n\ndef already_seen_content_id_in_past(user_id:int, content_id:int) -> int :\n    if user_id not in all_users_past_content_ids:\n        obj = bitarray(13550, endian='little')\n        all_users_past_content_ids[user_id] = obj # <edit here>\n        all_users_past_content_ids[user_id].setall(0)\n    if all_users_past_content_ids[user_id][content_id] == 1:\n        # the bit was already set before, so a repeated content\n        return 1\n    # set the content_id bit if seeing for the first time.\n    all_users_past_content_ids[user_id][content_id] = 1\n    # set it up for the first time, so we return a 0 here.\n    return 0\n\ndf_train[\"repeated_content\"] = df_train[[\"user_id\", \"content_id\"]].progress_apply(lambda row: already_seen_content_id_in_past(row[\"user_id\"], row[\"content_id\"]), axis=1)\n\n# 2000K rows in 24 secs, little slow :(\n# NB plz test it's correctness before ctrl+c,v :)\n```",
          "votes": 11
        },
        {
          "id": 1070426,
          "postDate": "2020-11-05T18:37:16.650Z",
          "content": "<p>Thanks ! I was implementing the same kind of function for the test set !<br>\nThat's how I fill the dictionary with the train set atm:</p>\n<pre><code>g = train[['user_id', 'content_id']].groupby(['user_id'])['content_id'].apply(lambda x: list(np.unique(x)))\n\nusers_dict = {}\nfor u in g.index:\n    a = bitarray(15000, endian='little')\n    a.setall(False)\n    for i in g[u]:\n        a[i] = 1\n    users_dict[u] = a\n\ndel g\n</code></pre>\n<p>I don't know if it's quicker but it works pretty well is a reasonable amount of time (&lt;2min for the whole train set)</p>",
          "rawMarkdown": "Thanks ! I was implementing the same kind of function for the test set !\nThat's how I fill the dictionary with the train set atm:\n```\ng = train[['user_id', 'content_id']].groupby(['user_id'])['content_id'].apply(lambda x: list(np.unique(x)))\n\nusers_dict = {}\nfor u in g.index:\n    a = bitarray(15000, endian='little')\n    a.setall(False)\n    for i in g[u]:\n        a[i] = 1\n    users_dict[u] = a\n    \ndel g\n```\nI don't know if it's quicker but it works pretty well is a reasonable amount of time (<2min for the whole train set)",
          "votes": 4,
          "replies": [
            {
              "id": 1070797,
              "postDate": "2020-11-06T07:27:42.477Z",
              "content": "<blockquote>\n  <p>users_dict = {}<br>\n  for u in g.index:<br>\n      a = bitarray(15000, endian='little')<br>\n      a.setall(False)<br>\n      for i in g[u]:<br>\n          a[i] = 1<br>\n      users_dict[u] = a</p>\n  <p>del g<br>\n  <a href=\"https://www.kaggle.com/alexj21\" target=\"_blank\">@alexj21</a> This code is used to keep track of whether a user has seen particular content or not right? <br>\n  can you suggest how can I keep track of the number of attempts made by user for particular content?<br>\n  and how to update test data using users_dict ?</p>\n</blockquote>",
              "rawMarkdown": " > users_dict = {}\n> for u in g.index:\n>     a = bitarray(15000, endian='little')\n>     a.setall(False)\n>     for i in g[u]:\n>         a[i] = 1\n>     users_dict[u] = a\n>     \n> del g\n\n@alexj21 This code is used to keep track of whether a user has seen particular content or not right? \ncan you suggest how can I keep track of the number of attempts made by user for particular content?\nand how to update test data using users_dict ?",
              "votes": 1
            },
            {
              "id": 1070815,
              "postDate": "2020-11-06T07:51:01.210Z",
              "content": "<p><a href=\"https://www.kaggle.com/maunish\" target=\"_blank\">@maunish</a> you're right<br>\nIt requires more memory for tracking the number of attemps, I did not work on it yet. <br>\nYour test loop should look like this:</p>\n<pre><code>prev_test_df = pd.DataFrame()\n\nfor (current_test_df, current_prediction_df) in iter_test:\n\n    # Update seen content\n    current_test_df[\"content_attempt\"] = current_test_df[[\"user_id\", \"content_id\"]].apply(lambda row: already_seen_content_id_in_past(row[\"user_id\"], row[\"content_id\"]), axis=1)\n\n    # Predict\n    current_test_df[target_col] =  model.predict(current_test_df[feat_col])\n    env.predict(current_test_df.loc[current_test_df['content_type_id'] == 0, ['row_id', 'answered_correctly']])\n</code></pre>\n<p>So that the very good function of <a href=\"https://www.kaggle.com/adityaecdrid\" target=\"_blank\">@adityaecdrid</a> would update both the current_test_df and the dictionnary which tracks the user's attempt. Beforehand you need to be sure that the dictionnary has been filled with stats from the whole train set (eg 10M rows) using functions above. It should work then</p>",
              "rawMarkdown": "@maunish you're right\nIt requires more memory for tracking the number of attemps, I did not work on it yet. \nYour test loop should look like this:\n```\nprev_test_df = pd.DataFrame()\n\nfor (current_test_df, current_prediction_df) in iter_test:\n \n    # Update seen content\n    current_test_df[\"content_attempt\"] = current_test_df[[\"user_id\", \"content_id\"]].apply(lambda row: already_seen_content_id_in_past(row[\"user_id\"], row[\"content_id\"]), axis=1)\n    \n    # Predict\n    current_test_df[target_col] =  model.predict(current_test_df[feat_col])\n    env.predict(current_test_df.loc[current_test_df['content_type_id'] == 0, ['row_id', 'answered_correctly']])\n    \n```\n\nSo that the very good function of @adityaecdrid would update both the current_test_df and the dictionnary which tracks the user's attempt. Beforehand you need to be sure that the dictionnary has been filled with stats from the whole train set (eg 10M rows) using functions above. It should work then",
              "votes": 3
            },
            {
              "id": 1070879,
              "postDate": "2020-11-06T09:04:28.277Z",
              "content": "<p>I have this idea of counting the number of attempts. Not sure if it will work with whole data during test.</p>\n<p><strong>To calaculate attempt number in train data is easy</strong></p>\n<pre><code>train_data[\"attempt_no\"] = 1\ntrain_data[\"attempt_no\"] = train_data[[\"user_id\",\"content_id\",'attempt_no']].groupby([\"user_id\",\"content_id\"])[\"attempt_no\"].cumsum()\n</code></pre>\n<p><strong>Now get the dictionary like tuple(user_id,content_id): max_attempt using this.</strong><br>\n<code>attempt_df = train_data[['user_id', 'content_id','attempt_no']].groupby(['user_id','content_id'])['attempt_no'].max().to_dict()</code></p>\n<p><strong>Now to get \"attempt_no\" in test data use below function in apply.</strong></p>\n<pre><code>def get_max_attempt(user_id,content_id):\n    k = (user_id,content_id)\n\n    if k in attempt_df.keys():\n        attempt_df[k]+=1\n        return attempt_df[k]\n\n    attempt_df[k] = 1\n    return attempt_df[k]\n</code></pre>\n<p><strong>Add this to test loop</strong></p>\n<p><code>test_data[\"attempt_no\"] = test_data[[\"user_id\", \"content_id\"]].apply(lambda row: get_max_attempt(row[\"user_id\"], row[\"content_id\"]), axis=1)</code></p>",
              "rawMarkdown": "I have this idea of counting the number of attempts. Not sure if it will work with whole data during test.\n\n**To calaculate attempt number in train data is easy**\n```\ntrain_data[\"attempt_no\"] = 1\ntrain_data[\"attempt_no\"] = train_data[[\"user_id\",\"content_id\",'attempt_no']].groupby([\"user_id\",\"content_id\"])[\"attempt_no\"].cumsum()\n```\n\n\n**Now get the dictionary like tuple(user_id,content_id): max_attempt using this.**\n`attempt_df = train_data[['user_id', 'content_id','attempt_no']].groupby(['user_id','content_id'])['attempt_no'].max().to_dict()`\n\n**Now to get \"attempt_no\" in test data use below function in apply.**\n```\ndef get_max_attempt(user_id,content_id):\n    k = (user_id,content_id)\n\n    if k in attempt_df.keys():\n        attempt_df[k]+=1\n        return attempt_df[k]\n\n    attempt_df[k] = 1\n    return attempt_df[k]\n```\n**Add this to test loop**\n\n`test_data[\"attempt_no\"] = test_data[[\"user_id\", \"content_id\"]].apply(lambda row: get_max_attempt(row[\"user_id\"], row[\"content_id\"]), axis=1)`\n\n\n",
              "votes": 17
            },
            {
              "id": 1070905,
              "postDate": "2020-11-06T09:49:14.117Z",
              "content": "<p>Nice ! I agree the challenge is to make it work under 9 hours. <br>\nWill definetly add this to my todo list</p>",
              "rawMarkdown": "Nice ! I agree the challenge is to make it work under 9 hours. \nWill definetly add this to my todo list",
              "votes": 1
            },
            {
              "id": 1071056,
              "postDate": "2020-11-06T12:57:48.540Z",
              "content": "<p>Glad it works out of the box :)</p>",
              "rawMarkdown": "Glad it works out of the box :)",
              "votes": 1
            },
            {
              "id": 1071753,
              "postDate": "2020-11-07T11:24:12.347Z",
              "content": "<p></p>\n<p>Patched.</p>",
              "rawMarkdown": "~~I am not sure but the bit array is losing information when i ran it for 100M rows. Will update. But if i run it for samples ~10M , it just runs perfectly... :( ~~\n\nPatched.",
              "votes": 1
            },
            {
              "id": 1071755,
              "postDate": "2020-11-07T11:32:29.240Z",
              "content": "<p>My above solution works perfectly. It did increase my score on lb.</p>",
              "rawMarkdown": "My above solution works perfectly. It did increase my score on lb.",
              "votes": 2
            },
            {
              "id": 1071846,
              "postDate": "2020-11-07T13:24:58.610Z",
              "content": "<p>I found my silly bug. Doing it in pandas might not be efficient IMO, that was the sole reason to find other ways.</p>\n<p>I have updated the code accordingly. Just create new bit array always otherwise you keep referring to the same memory slot. That was the bug. Plz update the same, I have validated the same now.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F835774%2Ff82e9a9e7db286cb81831ecbb64e0c40%2FScreenshot%202020-11-07%20at%206.59.02%20PM.png?generation=1604755777816352&amp;alt=media\" alt=\"Sample\"></p>",
              "rawMarkdown": "I found my silly bug. Doing it in pandas might not be efficient IMO, that was the sole reason to find other ways.\n\nI have updated the code accordingly. Just create new bit array always otherwise you keep referring to the same memory slot. That was the bug. Plz update the same, I have validated the same now.\n\n![Sample](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F835774%2Ff82e9a9e7db286cb81831ecbb64e0c40%2FScreenshot%202020-11-07%20at%206.59.02%20PM.png?generation=1604755777816352&alt=media)",
              "votes": 2
            },
            {
              "id": 1076642,
              "postDate": "2020-11-12T17:55:58.123Z",
              "content": "<p>Nice implementation, I'm having a problem with this line of code, memory is exploding, I'm using CuDF to speed submission so I'm limited to memory. Is there a work around? It looks like it uses alot of memory on the .to_dict() part</p>\n<p><code>attempt_df = train_data[['user_id', 'content_id','attempt_no']].groupby(['user_id','content_id'])['attempt_no'].max().to_dict()</code></p>\n<p>edit: also tried doing the process on 2 seperate notebooks and it still alocates alot of memory</p>",
              "rawMarkdown": "Nice implementation, I'm having a problem with this line of code, memory is exploding, I'm using CuDF to speed submission so I'm limited to memory. Is there a work around? It looks like it uses alot of memory on the .to_dict() part\n\n`attempt_df = train_data[['user_id', 'content_id','attempt_no']].groupby(['user_id','content_id'])['attempt_no'].max().to_dict()`\n\nedit: also tried doing the process on 2 seperate notebooks and it still alocates alot of memory\n",
              "votes": 4
            },
            {
              "id": 1078742,
              "postDate": "2020-11-15T08:52:14.017Z",
              "content": "<p>You can use Series constructed by <code>attempt_series = train_data[['user_id', 'content_id','attempt_no']].groupby(['user_id','content_id'])['attempt_no'].max()</code>. By <code>attempt_series[user_id,content_id]</code> you can get the element you want. It can reduce much of the memory cost.</p>",
              "rawMarkdown": "You can use Series constructed by `attempt_series = train_data[['user_id', 'content_id','attempt_no']].groupby(['user_id','content_id'])['attempt_no'].max()`. By `attempt_series[user_id,content_id]` you can get the element you want. It can reduce much of the memory cost.",
              "votes": 6
            },
            {
              "id": 1083114,
              "postDate": "2020-11-18T16:00:27.283Z",
              "content": "<p>I tried to use defaultdict, it cost me nearly 10g of memory, which is terrible.</p>",
              "rawMarkdown": "I tried to use defaultdict, it cost me nearly 10g of memory, which is terrible.",
              "votes": 1
            },
            {
              "id": 1083147,
              "postDate": "2020-11-18T16:48:57.767Z",
              "rawMarkdown": "",
              "votes": -1,
              "isDeleted": true
            }
          ]
        },
        {
          "id": 1070430,
          "postDate": "2020-11-05T18:47:28.337Z",
          "content": "<p>Nice implementation 🎉</p>",
          "rawMarkdown": "Nice implementation 🎉",
          "votes": 2
        },
        {
          "id": 1070663,
          "postDate": "2020-11-06T03:08:55.477Z",
          "content": "<p>Thanks for sharing this <a href=\"https://www.kaggle.com/adityaecdrid\" target=\"_blank\">@adityaecdrid</a>! </p>",
          "rawMarkdown": "Thanks for sharing this @adityaecdrid! ",
          "votes": 1
        },
        {
          "id": 1084285,
          "postDate": "2020-11-19T22:14:08.887Z",
          "content": "<p>Thanks for sharing <a href=\"https://www.kaggle.com/maunish\" target=\"_blank\">@maunish</a> ! I'm using cuDf which doesn't have .cumsum() unfortunately. I'm struggling to work out an alternative way of doing this without .cumsum(). Can you think of an alternative way of doing this?</p>",
          "rawMarkdown": "Thanks for sharing @maunish ! I'm using cuDf which doesn't have .cumsum() unfortunately. I'm struggling to work out an alternative way of doing this without .cumsum(). Can you think of an alternative way of doing this?"
        },
        {
          "id": 1084326,
          "postDate": "2020-11-19T23:55:02.457Z",
          "content": "<p>I had a similar question, that got answered <a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/193411\" target=\"_blank\">here</a>. I'm not sure if it's fast enough, but may give some ideas.</p>",
          "rawMarkdown": "I had a similar question, that got answered [here](https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/193411). I'm not sure if it's fast enough, but may give some ideas.",
          "votes": 1
        },
        {
          "id": 1084840,
          "postDate": "2020-11-20T12:45:14.933Z",
          "content": "<p>Thanks very much indeed!</p>",
          "rawMarkdown": "Thanks very much indeed!"
        },
        {
          "id": 1103539,
          "postDate": "2020-12-06T02:35:36.290Z",
          "content": "<p>Is anyone successfully implements in testing data?<br>\nI always got out of memory…</p>",
          "rawMarkdown": "Is anyone successfully implements in testing data?\nI always got out of memory...",
          "votes": 1
        },
        {
          "id": 1103826,
          "postDate": "2020-12-06T10:30:37.677Z",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/louieshao\" target=\"_blank\">@louieshao</a> , I tried the same approach as you suggested, helped me in handling memory error, but after submission I got \"Submission Scoring Error\". Moreover while running on test dataframe in notebook, it took around 40mins to complete.<br>\nHow did you add and update this attempt_no in test loop?</p>",
          "rawMarkdown": "Hi @louieshao , I tried the same approach as you suggested, helped me in handling memory error, but after submission I got \"Submission Scoring Error\". Moreover while running on test dataframe in notebook, it took around 40mins to complete.\nHow did you add and update this attempt_no in test loop?",
          "votes": 2
        },
        {
          "id": 1107891,
          "postDate": "2020-12-10T03:39:19.927Z",
          "content": "<p>Hi, <a href=\"https://www.kaggle.com/smritisingh1997\" target=\"_blank\">@smritisingh1997</a> , I also find this issue. Finally I use boolean to contruct this features…. Haven't yet found a way to have a balance between time-cost and memory-cost.</p>",
          "rawMarkdown": "Hi, @smritisingh1997 , I also find this issue. Finally I use boolean to contruct this features.... Haven't yet found a way to have a balance between time-cost and memory-cost."
        }
      ]
    },
    {
      "id": 1067262,
      "postDate": "2020-11-02T12:26:48.170Z",
      "content": "<p>Interesting one</p>",
      "rawMarkdown": "Interesting one",
      "votes": 1
    },
    {
      "id": 1066983,
      "postDate": "2020-11-02T09:19:22.497Z",
      "content": "<p>That's a nice find </p>",
      "rawMarkdown": "That's a nice find ",
      "votes": 1
    },
    {
      "id": 1066105,
      "postDate": "2020-11-01T10:30:24.237Z",
      "content": "<p>Yeah I find this insight too but for me it doesn't really boost my score idk what I do wrong</p>",
      "rawMarkdown": "Yeah I find this insight too but for me it doesn't really boost my score idk what I do wrong",
      "votes": 1
    },
    {
      "id": 1067805,
      "postDate": "2020-11-02T18:17:23.693Z",
      "content": "<p>One thing which i don't understand is the headline in the section, (from their website)</p>\n<blockquote>\n  <p>\"Be sure to give priority to problems and lectures that raise the score. With an AI guide you can learn only what you need to improve your score.\"</p>\n</blockquote>\n<p>In real life, we do learn a lot of things which aren't useful directly but it does helps a lot in the long run IMHO.</p>",
      "rawMarkdown": "One thing which i don't understand is the headline in the section, (from their website)\n\n> \"Be sure to give priority to problems and lectures that raise the score. With an AI guide you can learn only what you need to improve your score.\"\n\nIn real life, we do learn a lot of things which aren't useful directly but it does helps a lot in the long run IMHO.",
      "votes": 2
    },
    {
      "id": 1066061,
      "postDate": "2020-11-01T09:03:03.003Z",
      "content": "<p>I have tried this feature out. My current model doesn't have a very good score but for what it is worth it gave a boost of ~0.005.</p>",
      "rawMarkdown": "I have tried this feature out. My current model doesn't have a very good score but for what it is worth it gave a boost of ~0.005.",
      "votes": 2
    },
    {
      "id": 1067263,
      "postDate": "2020-11-02T12:27:41.317Z",
      "content": "<p>This is nice work, thank a  lot for sharing it</p>",
      "rawMarkdown": "This is nice work, thank a  lot for sharing it",
      "votes": 1
    },
    {
      "id": 1066178,
      "postDate": "2020-11-01T12:22:59.513Z",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/aravindpadman\" target=\"_blank\">@aravindpadman</a>. Thanks a lot for sharing this nice work with us.</p>",
      "rawMarkdown": "Hi @aravindpadman. Thanks a lot for sharing this nice work with us.",
      "votes": 1
    },
    {
      "id": 1093295,
      "postDate": "2020-11-27T15:46:45.580Z",
      "content": "<p>Nice work!<br>\nThis feature, it works</p>",
      "rawMarkdown": "Nice work!\nThis feature, it works"
    },
    {
      "id": 1090249,
      "postDate": "2020-11-25T07:40:51.630Z",
      "content": "<p>Nice job!! Thank for sharing!! It helps us to make features for this situation.</p>",
      "rawMarkdown": "Nice job!! Thank for sharing!! It helps us to make features for this situation."
    }
  ],
  "comments": [
    {
      "id": 1065729,
      "author_name": "Abhimanyu Dikshit",
      "author_url": "",
      "post_date": "2020-10-31T16:44:53.630000",
      "content": "<p>This is a really nice find <a href=\"https://www.kaggle.com/aravindpadman\" target=\"_blank\">@aravindpadman</a> . Nice!</p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 1067024,
      "author_name": "Alex",
      "author_url": "",
      "post_date": "2020-11-02T09:59:01.307000",
      "content": "<p>That's an interesting finding and that seems to me logical. <br>\nI'll add this feature to my model, not sure yet if only a categorical column (first attempt, repeated attempt) will be enough; the number of attempts could provide information too. I would say that more the user see the question and more the user is likely to answer correctly</p>",
      "votes": 4,
      "replies": [
        {
          "id": 1067290,
          "author_name": "Aravind P",
          "author_url": "",
          "post_date": "2020-11-02T13:00:12.690000",
          "content": "<p><code>I would say that more the user see the question and more the user is likely to answer correctly</code><br>\nTrue. I have actually observed this trend during my analysis.</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 1067361,
          "author_name": "Alex",
          "author_url": "",
          "post_date": "2020-11-02T13:49:24.473000",
          "content": "<p>Simply adding a boolean column for first attempt does boost my validation (+ 0.006)<br>\nHowever I don't know yet how to make it work during inference without reaching 9 hours limit</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 1087141,
          "author_name": "Yu Kang",
          "author_url": "",
          "post_date": "2020-11-22T11:47:00.473000",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/alexj21\" target=\"_blank\">@alexj21</a> , how do you store the repeated-questions-information of the training set? I stored it as a dict and the key is (user, question) pair, but it's too large to kaggle kernel so that I can't even load it during inference.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1088059,
          "author_name": "Alex",
          "author_url": "",
          "post_date": "2020-11-23T09:28:32.770000",
          "content": "<p>I stored it as a dict in which the key is the user_id and the value a bitarray of length 15k. Each bit is assigned to a content ID. If the user saw a given content id then its bit is set to 1. <br>\nYou can find some code examples in the discussion bellow. It works pretty well :) </p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 1088159,
          "author_name": "Yu Kang",
          "author_url": "",
          "post_date": "2020-11-23T11:41:10.500000",
          "content": "<p>Thanks, really smart method!</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1088224,
          "author_name": "Alex",
          "author_url": "",
          "post_date": "2020-11-23T12:55:01.930000",
          "content": "<p>The idea was firstly proposed by <a href=\"https://www.kaggle.com/adityaecdrid\" target=\"_blank\">@adityaecdrid</a> </p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 1065809,
      "author_name": "Caleb",
      "author_url": "",
      "post_date": "2020-10-31T19:37:22.447000",
      "content": "<p>That is cool. You could include the cumulative count or sum of <code>answered_correctly</code> by <code>content_id</code> as a single column feature. Here is some sql that should work.</p>\n<pre><code>SELECT user_id, task_container_id, row_id, content_id, COUNT(answered_correctly)\n  OVER (\n    PARTITION BY user_id, content_id\n    ORDER BY task_container_id\n    RANGE BETWEEN UNBOUNDED PRECEDING AND 1 PRECEDING\n  ) calc\nFROM train\nORDER BY user_id, content_id, task_container_id, row_id\n</code></pre>",
      "votes": 4,
      "replies": [
        {
          "id": 1065957,
          "author_name": "Aravind P",
          "author_url": "",
          "post_date": "2020-11-01T05:30:32.577000",
          "content": "<p>Thank you <a href=\"https://www.kaggle.com/calebeverett\" target=\"_blank\">@calebeverett</a>. I will try this out</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1066739,
          "author_name": "Mingjie Wang",
          "author_url": "",
          "post_date": "2020-11-02T01:32:03.870000",
          "content": "<p>Did you use SQL to complete feature engineering? Very cool!!!</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1066747,
          "author_name": "Caleb",
          "author_url": "",
          "post_date": "2020-11-02T01:40:14.140000",
          "content": "<p>I've got a cool set up going with BigQuery. I was going to share it once I get the prediction pipeline done.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1069174,
          "author_name": "Caleb",
          "author_url": "",
          "post_date": "2020-11-04T07:09:18.227000",
          "content": "<p>I got my BigQuery set up cleaned up enough to share.</p>\n<p><a href=\"https://www.kaggle.com/calebeverett/riiid-bigquery-xgboost-end-to-end/\" target=\"_blank\">https://www.kaggle.com/calebeverett/riiid-bigquery-xgboost-end-to-end/</a></p>",
          "votes": 7,
          "replies": []
        },
        {
          "id": 1069698,
          "author_name": "Sohier Dane",
          "author_url": "",
          "post_date": "2020-11-04T19:17:46.737000",
          "content": "<p><a href=\"https://www.kaggle.com/calebeverett\" target=\"_blank\">@calebeverett</a> neat!</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 1093974,
      "author_name": "PC Jimmmy",
      "author_url": "",
      "post_date": "2020-11-28T07:27:56.933000",
      "content": "<p>The data is described as being the result of a \"smart\" training process.  If it's smart than repeated questions should occur in a couple of cases that I can think about.</p>\n<ol>\n<li>Very difficult question (overall low percentage of correct answers for the full set of users).  You would likely want to repeat a question several times to make sure that the user really does know the answer - vs just a lucky guess the first time asked.</li>\n<li>User got it wrong - need to repeat until they get it right.</li>\n</ol>\n<p>But at some point you might stop asking the question after one or more correct responses.  </p>\n<p>So your plots make lots of sense. </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1073219,
      "author_name": "Balraj Gupta",
      "author_url": "",
      "post_date": "2020-11-09T09:39:01.347000",
      "content": "<p>Great find!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1070388,
      "author_name": "Overfitted",
      "author_url": "",
      "post_date": "2020-11-05T17:47:24.140000",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/aravindpadman\" target=\"_blank\">@aravindpadman</a>. Nice work. That could be useful! Thanks for sharing!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1070105,
      "author_name": "Zubair Rahman Tusar",
      "author_url": "",
      "post_date": "2020-11-05T11:07:18.170000",
      "content": "<p>Woah!! Good work there.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1069484,
      "author_name": "Szymon Szostak",
      "author_url": "",
      "post_date": "2020-11-04T14:10:31.990000",
      "content": "<p>Nice work. That could be usefull! Thanks for sharing!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1069253,
      "author_name": "Maunish dave",
      "author_url": "",
      "post_date": "2020-11-04T08:57:48.223000",
      "content": "<p>I tried to use this in my model(lgbm) and it looks promising and my validation score did go up. <br>\nI used this code to find attempt number.</p>\n<p>train[\"attempt_no\"] =1<br>\ntrain[\"attemtp_no\"] = train[[\"user_id\",\"content_id\",\"attempt_no\"]].groupby([\"user_id\",\"content_id\"])[\"attempt_no\"].cumsum()</p>\n<p>but the problem is that i can not do same in test data as test data comes in batch and we can not see wether all the content_id of a user are in same batch. So it will give inaccurte results.</p>\n<p>So my question is. Is there any way to resolve this issue?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1069291,
          "author_name": "Alex",
          "author_url": "",
          "post_date": "2020-11-04T09:35:36.083000",
          "content": "<p>I'm experiencing the same issue.<br>\nThe first solution that came to my mind would be to keep your train dataframe and update your test set on the fly with features from the train set. But this is a quite expensive solution. Plus you need to handle unseen users</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1069347,
          "author_name": "Aditya Soni",
          "author_url": "",
          "post_date": "2020-11-04T10:40:02.540000",
          "content": "<blockquote>\n  <p>i can not do same in test data as test data comes in batch</p>\n</blockquote>\n<p>You have to keep a hash_map for any given user at any given point. If we re-see a particular content_id, then we can use the old attempt_no for that question directly to start with and later as we move ahead, we can re-think whether we want to update that old_no with the new_max number possible.</p>\n<p>And if one wants to efficient with little more efforts than the usual, use something like a bit-array as your data structure e.g. (to store which all content_ids a user has seen already) and it's easy to update them as well on FLY. We can use <a href=\"https://docs.python.org/3/library/array.html\" target=\"_blank\">array</a> instead of the list as it's mem efficient to cache nos..</p>\n<p>All the prime numbers less than 18 can be written as,</p>\n<pre><code>bit 31                           bit 0\n    |                              |\n    00000000000000100010100010101100 \n</code></pre>\n<p>Prime no's are (2, 3, 5, 7, 11, 13, 17)  (and it should be efficient in terms of memory and speed as well) (32 bits is just 4B) <strong>(we might be complicating things)</strong></p>\n<pre><code>from bitarray import bitarray\na = bitarray(31, endian='little') # important to setup a fixed endianess.\na.setall(False)\n\nfor i in [2, 3, 5, 7, 11, 13, 17]:\n    a[i] = 1\n\nprint(a) # bitarray('0011010100010100010000000000000')\n</code></pre>\n<p>To store, 14k content_ids, we will need around 1814B for each user. (rough calculations, not apt precisely)</p>\n<p>In case someone's looking for a team member, let me know! Ty! (PS Just my idea, i haven't implemented it (yet) and it can be extended to other stats about users if they are just numbers by shifting up bits etc)</p>",
          "votes": 14,
          "replies": []
        },
        {
          "id": 1069550,
          "author_name": "Caleb",
          "author_url": "",
          "post_date": "2020-11-04T15:35:38.167000",
          "content": "<p>Does anybody know if the competition api gets loaded on to the gpu and then iterates through batches or do batches have to get moved on as the data is iterated through from the cpu?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1069551,
          "author_name": "Aditya Soni",
          "author_url": "",
          "post_date": "2020-11-04T15:41:04.757000",
          "content": "<p>We have to move data to GPU for sure, by default it shouldn't, right?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1069562,
          "author_name": "Caleb",
          "author_url": "",
          "post_date": "2020-11-04T15:50:55.460000",
          "content": "<p>I haven't tried it yet, but was thinking that the speedup from running inference on the gpu would be a lot faster if the competition api data loaded there before serving batches, otherwise you have the overhead of moving a couple thousand small batches on.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1069604,
          "author_name": "Caleb",
          "author_url": "",
          "post_date": "2020-11-04T17:06:28.583000",
          "content": "<p>Th competition api serves tuples of pandas dataframes so that obviously means they are on the cpu.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1069605,
          "author_name": "Aditya Soni",
          "author_url": "",
          "post_date": "2020-11-04T17:06:33.523000",
          "content": "<p>Yes that overhead is there. Plus lot of code optimisations needed with not a single failure case as it will break the pipeline and we cannot afford that. Also they cannot move data on GPU before hand as we might be doing one thing or the another on it as well. Hence, it's better leave it at the user-side as to how they want to use the information..</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1070009,
          "author_name": "Alex",
          "author_url": "",
          "post_date": "2020-11-05T08:17:22.590000",
          "content": "<p><a href=\"https://www.kaggle.com/adityaecdrid\" target=\"_blank\">@adityaecdrid</a> thanks for providing some ideas. So if I sum-up the first solution would be to create a dictionary for each user where the key would be the content id and the pair the number of attempts ?<br>\nThe second solution looks faster indeed but I'm not sure to understand the link with prime numbers. Do you suggest to assign a bit signifying first attempt for each prime number along the bit array?  They are 11 prime numbers that can fit into a 32-bit array so for 14k contents you need 14K/11 = ~1272 bytes is that right ?  </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1070048,
          "author_name": "Aditya Soni",
          "author_url": "",
          "post_date": "2020-11-05T09:04:28.733000",
          "content": "<p>We know that any user can at max have 14k contents possible (max content_id is ~14k). So i can represent these information easily by using a bit vector of length 14k where every position signifies whether the user has seen this particular content_id or not.</p>\n<p>e.g. let's say user_id 115 has seen content_ids \"2, 3, 5, 7, 11, 13, 17\" till now, we can represent that information easily like this,</p>\n<pre><code>bit 31                           bit 0\n    |                              |\n    00000000000000100010100010101100 \n</code></pre>\n<p>And that's all for the training set let's say. </p>\n<p>Now when it comes to test set, we can easily add a new_content_id the user hasn't seen in the past but sees it now. I am adding 19 as an e.g. to the same. So the above bit vector becomes,</p>\n<pre><code>bit 31                           bit 0\n    |                              |\n    00000000000010100010100010101100 \n</code></pre>\n<p>i.e. <code>a[19] = 1</code>; Cool, right?</p>\n<p>And now let's say we have received another content_id as 2, now whether we have seen this content_id or not, can easily be deduced. One simple check is <code>a[2] == 1</code> ?</p>\n<p>Hope it's clear now Alex.</p>",
          "votes": 7,
          "replies": []
        },
        {
          "id": 1070050,
          "author_name": "Alex",
          "author_url": "",
          "post_date": "2020-11-05T09:18:12.437000",
          "content": "<p>Yeah it's clear now ! I was lost with the prime numbers, I though there was something like a subtle combination with content id. </p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1070370,
          "author_name": "Aditya Soni",
          "author_url": "",
          "post_date": "2020-11-05T17:11:40.820000",
          "content": "<p>try this <a href=\"https://www.kaggle.com/alexj21\" target=\"_blank\">@alexj21</a>,</p>\n<pre><code>from collections import defaultdict\nfrom bitarray import bitarray\n# tqdm.pandas()\n\n# ONLY FOR NON_LECTURE_ROWS\n\nall_users_past_content_ids = {}\n\ndef already_seen_content_id_in_past(user_id:int, content_id:int) -&gt; int :\n    if user_id not in all_users_past_content_ids:\n        obj = bitarray(13550, endian='little')\n        all_users_past_content_ids[user_id] = obj # &lt;edit here&gt;\n        all_users_past_content_ids[user_id].setall(0)\n    if all_users_past_content_ids[user_id][content_id] == 1:\n        # the bit was already set before, so a repeated content\n        return 1\n    # set the content_id bit if seeing for the first time.\n    all_users_past_content_ids[user_id][content_id] = 1\n    # set it up for the first time, so we return a 0 here.\n    return 0\n\ndf_train[\"repeated_content\"] = df_train[[\"user_id\", \"content_id\"]].progress_apply(lambda row: already_seen_content_id_in_past(row[\"user_id\"], row[\"content_id\"]), axis=1)\n\n# 2000K rows in 24 secs, little slow :(\n# NB plz test it's correctness before ctrl+c,v :)\n</code></pre>",
          "votes": 11,
          "replies": []
        },
        {
          "id": 1070426,
          "author_name": "Alex",
          "author_url": "",
          "post_date": "2020-11-05T18:37:16.650000",
          "content": "<p>Thanks ! I was implementing the same kind of function for the test set !<br>\nThat's how I fill the dictionary with the train set atm:</p>\n<pre><code>g = train[['user_id', 'content_id']].groupby(['user_id'])['content_id'].apply(lambda x: list(np.unique(x)))\n\nusers_dict = {}\nfor u in g.index:\n    a = bitarray(15000, endian='little')\n    a.setall(False)\n    for i in g[u]:\n        a[i] = 1\n    users_dict[u] = a\n\ndel g\n</code></pre>\n<p>I don't know if it's quicker but it works pretty well is a reasonable amount of time (&lt;2min for the whole train set)</p>",
          "votes": 4,
          "replies": [
            {
              "id": 1070797,
              "author_name": "Maunish dave",
              "author_url": "",
              "post_date": "2020-11-06T07:27:42.477000",
              "content": "<blockquote>\n  <p>users_dict = {}<br>\n  for u in g.index:<br>\n      a = bitarray(15000, endian='little')<br>\n      a.setall(False)<br>\n      for i in g[u]:<br>\n          a[i] = 1<br>\n      users_dict[u] = a</p>\n  <p>del g<br>\n  <a href=\"https://www.kaggle.com/alexj21\" target=\"_blank\">@alexj21</a> This code is used to keep track of whether a user has seen particular content or not right? <br>\n  can you suggest how can I keep track of the number of attempts made by user for particular content?<br>\n  and how to update test data using users_dict ?</p>\n</blockquote>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 1070815,
              "author_name": "Alex",
              "author_url": "",
              "post_date": "2020-11-06T07:51:01.210000",
              "content": "<p><a href=\"https://www.kaggle.com/maunish\" target=\"_blank\">@maunish</a> you're right<br>\nIt requires more memory for tracking the number of attemps, I did not work on it yet. <br>\nYour test loop should look like this:</p>\n<pre><code>prev_test_df = pd.DataFrame()\n\nfor (current_test_df, current_prediction_df) in iter_test:\n\n    # Update seen content\n    current_test_df[\"content_attempt\"] = current_test_df[[\"user_id\", \"content_id\"]].apply(lambda row: already_seen_content_id_in_past(row[\"user_id\"], row[\"content_id\"]), axis=1)\n\n    # Predict\n    current_test_df[target_col] =  model.predict(current_test_df[feat_col])\n    env.predict(current_test_df.loc[current_test_df['content_type_id'] == 0, ['row_id', 'answered_correctly']])\n</code></pre>\n<p>So that the very good function of <a href=\"https://www.kaggle.com/adityaecdrid\" target=\"_blank\">@adityaecdrid</a> would update both the current_test_df and the dictionnary which tracks the user's attempt. Beforehand you need to be sure that the dictionnary has been filled with stats from the whole train set (eg 10M rows) using functions above. It should work then</p>",
              "votes": 3,
              "replies": []
            },
            {
              "id": 1070879,
              "author_name": "Maunish dave",
              "author_url": "",
              "post_date": "2020-11-06T09:04:28.277000",
              "content": "<p>I have this idea of counting the number of attempts. Not sure if it will work with whole data during test.</p>\n<p><strong>To calaculate attempt number in train data is easy</strong></p>\n<pre><code>train_data[\"attempt_no\"] = 1\ntrain_data[\"attempt_no\"] = train_data[[\"user_id\",\"content_id\",'attempt_no']].groupby([\"user_id\",\"content_id\"])[\"attempt_no\"].cumsum()\n</code></pre>\n<p><strong>Now get the dictionary like tuple(user_id,content_id): max_attempt using this.</strong><br>\n<code>attempt_df = train_data[['user_id', 'content_id','attempt_no']].groupby(['user_id','content_id'])['attempt_no'].max().to_dict()</code></p>\n<p><strong>Now to get \"attempt_no\" in test data use below function in apply.</strong></p>\n<pre><code>def get_max_attempt(user_id,content_id):\n    k = (user_id,content_id)\n\n    if k in attempt_df.keys():\n        attempt_df[k]+=1\n        return attempt_df[k]\n\n    attempt_df[k] = 1\n    return attempt_df[k]\n</code></pre>\n<p><strong>Add this to test loop</strong></p>\n<p><code>test_data[\"attempt_no\"] = test_data[[\"user_id\", \"content_id\"]].apply(lambda row: get_max_attempt(row[\"user_id\"], row[\"content_id\"]), axis=1)</code></p>",
              "votes": 17,
              "replies": []
            },
            {
              "id": 1070905,
              "author_name": "Alex",
              "author_url": "",
              "post_date": "2020-11-06T09:49:14.117000",
              "content": "<p>Nice ! I agree the challenge is to make it work under 9 hours. <br>\nWill definetly add this to my todo list</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 1071056,
              "author_name": "Aditya Soni",
              "author_url": "",
              "post_date": "2020-11-06T12:57:48.540000",
              "content": "<p>Glad it works out of the box :)</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 1071753,
              "author_name": "Aditya Soni",
              "author_url": "",
              "post_date": "2020-11-07T11:24:12.347000",
              "content": "<p></p>\n<p>Patched.</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 1071755,
              "author_name": "Maunish dave",
              "author_url": "",
              "post_date": "2020-11-07T11:32:29.240000",
              "content": "<p>My above solution works perfectly. It did increase my score on lb.</p>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 1071846,
              "author_name": "Aditya Soni",
              "author_url": "",
              "post_date": "2020-11-07T13:24:58.610000",
              "content": "<p>I found my silly bug. Doing it in pandas might not be efficient IMO, that was the sole reason to find other ways.</p>\n<p>I have updated the code accordingly. Just create new bit array always otherwise you keep referring to the same memory slot. That was the bug. Plz update the same, I have validated the same now.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F835774%2Ff82e9a9e7db286cb81831ecbb64e0c40%2FScreenshot%202020-11-07%20at%206.59.02%20PM.png?generation=1604755777816352&amp;alt=media\" alt=\"Sample\"></p>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 1076642,
              "author_name": "Iuryck Santos",
              "author_url": "",
              "post_date": "2020-11-12T17:55:58.123000",
              "content": "<p>Nice implementation, I'm having a problem with this line of code, memory is exploding, I'm using CuDF to speed submission so I'm limited to memory. Is there a work around? It looks like it uses alot of memory on the .to_dict() part</p>\n<p><code>attempt_df = train_data[['user_id', 'content_id','attempt_no']].groupby(['user_id','content_id'])['attempt_no'].max().to_dict()</code></p>\n<p>edit: also tried doing the process on 2 seperate notebooks and it still alocates alot of memory</p>",
              "votes": 4,
              "replies": []
            },
            {
              "id": 1078742,
              "author_name": "Shihao Shao",
              "author_url": "",
              "post_date": "2020-11-15T08:52:14.017000",
              "content": "<p>You can use Series constructed by <code>attempt_series = train_data[['user_id', 'content_id','attempt_no']].groupby(['user_id','content_id'])['attempt_no'].max()</code>. By <code>attempt_series[user_id,content_id]</code> you can get the element you want. It can reduce much of the memory cost.</p>",
              "votes": 6,
              "replies": []
            },
            {
              "id": 1083114,
              "author_name": "MeeCreeps",
              "author_url": "",
              "post_date": "2020-11-18T16:00:27.283000",
              "content": "<p>I tried to use defaultdict, it cost me nearly 10g of memory, which is terrible.</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 1083147,
              "author_name": "",
              "author_url": "",
              "post_date": "2020-11-18T16:48:57.767000",
              "content": "",
              "votes": -1,
              "replies": []
            }
          ]
        },
        {
          "id": 1070430,
          "author_name": "Aditya Soni",
          "author_url": "",
          "post_date": "2020-11-05T18:47:28.337000",
          "content": "<p>Nice implementation 🎉</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1070663,
          "author_name": "Naresh Jagadeesan",
          "author_url": "",
          "post_date": "2020-11-06T03:08:55.477000",
          "content": "<p>Thanks for sharing this <a href=\"https://www.kaggle.com/adityaecdrid\" target=\"_blank\">@adityaecdrid</a>! </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1084285,
          "author_name": "Neil Gibbons",
          "author_url": "",
          "post_date": "2020-11-19T22:14:08.887000",
          "content": "<p>Thanks for sharing <a href=\"https://www.kaggle.com/maunish\" target=\"_blank\">@maunish</a> ! I'm using cuDf which doesn't have .cumsum() unfortunately. I'm struggling to work out an alternative way of doing this without .cumsum(). Can you think of an alternative way of doing this?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1084326,
          "author_name": "Caleb",
          "author_url": "",
          "post_date": "2020-11-19T23:55:02.457000",
          "content": "<p>I had a similar question, that got answered <a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/193411\" target=\"_blank\">here</a>. I'm not sure if it's fast enough, but may give some ideas.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1084840,
          "author_name": "Neil Gibbons",
          "author_url": "",
          "post_date": "2020-11-20T12:45:14.933000",
          "content": "<p>Thanks very much indeed!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1103539,
          "author_name": "Dean",
          "author_url": "",
          "post_date": "2020-12-06T02:35:36.290000",
          "content": "<p>Is anyone successfully implements in testing data?<br>\nI always got out of memory…</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1103826,
          "author_name": "Smriti ",
          "author_url": "",
          "post_date": "2020-12-06T10:30:37.677000",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/louieshao\" target=\"_blank\">@louieshao</a> , I tried the same approach as you suggested, helped me in handling memory error, but after submission I got \"Submission Scoring Error\". Moreover while running on test dataframe in notebook, it took around 40mins to complete.<br>\nHow did you add and update this attempt_no in test loop?</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1107891,
          "author_name": "Shihao Shao",
          "author_url": "",
          "post_date": "2020-12-10T03:39:19.927000",
          "content": "<p>Hi, <a href=\"https://www.kaggle.com/smritisingh1997\" target=\"_blank\">@smritisingh1997</a> , I also find this issue. Finally I use boolean to contruct this features…. Haven't yet found a way to have a balance between time-cost and memory-cost.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1067262,
      "author_name": "sreeram78",
      "author_url": "",
      "post_date": "2020-11-02T12:26:48.170000",
      "content": "<p>Interesting one</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1066983,
      "author_name": "Saurav Anand",
      "author_url": "",
      "post_date": "2020-11-02T09:19:22.497000",
      "content": "<p>That's a nice find </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1066105,
      "author_name": "Marcello Susanto",
      "author_url": "",
      "post_date": "2020-11-01T10:30:24.237000",
      "content": "<p>Yeah I find this insight too but for me it doesn't really boost my score idk what I do wrong</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1067805,
      "author_name": "Aditya Soni",
      "author_url": "",
      "post_date": "2020-11-02T18:17:23.693000",
      "content": "<p>One thing which i don't understand is the headline in the section, (from their website)</p>\n<blockquote>\n  <p>\"Be sure to give priority to problems and lectures that raise the score. With an AI guide you can learn only what you need to improve your score.\"</p>\n</blockquote>\n<p>In real life, we do learn a lot of things which aren't useful directly but it does helps a lot in the long run IMHO.</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 1066061,
      "author_name": "LGreig",
      "author_url": "",
      "post_date": "2020-11-01T09:03:03.003000",
      "content": "<p>I have tried this feature out. My current model doesn't have a very good score but for what it is worth it gave a boost of ~0.005.</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 1067263,
      "author_name": "sreeram78",
      "author_url": "",
      "post_date": "2020-11-02T12:27:41.317000",
      "content": "<p>This is nice work, thank a  lot for sharing it</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1066178,
      "author_name": "Deep Neuron AI",
      "author_url": "",
      "post_date": "2020-11-01T12:22:59.513000",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/aravindpadman\" target=\"_blank\">@aravindpadman</a>. Thanks a lot for sharing this nice work with us.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1093295,
      "author_name": "Karl",
      "author_url": "",
      "post_date": "2020-11-27T15:46:45.580000",
      "content": "<p>Nice work!<br>\nThis feature, it works</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1090249,
      "author_name": "ydl1y17",
      "author_url": "",
      "post_date": "2020-11-25T07:40:51.630000",
      "content": "<p>Nice job!! Thank for sharing!! It helps us to make features for this situation.</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1065724": "I have noticed that users are given the same questions multiple times during the course of their learning. The following table shows a question repeated 9 times for a user.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F560450%2F574463e460ee3d42c675690db772ea64%2FScreenshot%202020-10-31%20at%209.45.58%20PM.png?generation=1604161013985800&alt=media)\n\n\nAfter the above analysis, my hypothesis was that the probability of a user answering a question correctly given that the question has already seen before is higher than the probability of the user answering the question correctly in the first attempt.\nTo prove this I have plotted the probability distribution of average answer correctness at a user level for two populations(**First attempt Vs Repeated Attempt**). The following figure shows the results\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F560450%2Ff4d8b0e53df65d3877f6ec2d31b00fac%2FScreenshot%202020-10-31%20at%209.51.03%20PM.png?generation=1604161298744094&alt=media)\nI have also done a welch's t-test to prove that these two distributions are indeed different.\n\nFor the detailed analysis, please refer to the notebook I have created \nhttps://www.kaggle.com/aravindpadman/riiid-statistical-analysis-on-repeated-questions\n\nI am currently trying to incorporate this idea into my modeling, but I am not yet there.\n\nLet me know your valuable ideas in the comment.",
    "1065729": "This is a really nice find @aravindpadman . Nice!",
    "1067024": "That's an interesting finding and that seems to me logical. \nI'll add this feature to my model, not sure yet if only a categorical column (first attempt, repeated attempt) will be enough; the number of attempts could provide information too. I would say that more the user see the question and more the user is likely to answer correctly",
    "1065809": "That is cool. You could include the cumulative count or sum of `answered_correctly` by `content_id` as a single column feature. Here is some sql that should work.\n\n```\nSELECT user_id, task_container_id, row_id, content_id, COUNT(answered_correctly)\n  OVER (\n    PARTITION BY user_id, content_id\n    ORDER BY task_container_id\n    RANGE BETWEEN UNBOUNDED PRECEDING AND 1 PRECEDING\n  ) calc\nFROM train\nORDER BY user_id, content_id, task_container_id, row_id\n```  ",
    "1093974": "The data is described as being the result of a \"smart\" training process.  If it's smart than repeated questions should occur in a couple of cases that I can think about.\n\n1.  Very difficult question (overall low percentage of correct answers for the full set of users).  You would likely want to repeat a question several times to make sure that the user really does know the answer - vs just a lucky guess the first time asked.\n2.  User got it wrong - need to repeat until they get it right.\n\nBut at some point you might stop asking the question after one or more correct responses.  \n\nSo your plots make lots of sense. ",
    "1073219": "Great find!",
    "1070388": "Hi @aravindpadman. Nice work. That could be useful! Thanks for sharing!",
    "1070105": "Woah!! Good work there.",
    "1069484": "Nice work. That could be usefull! Thanks for sharing!",
    "1069253": "I tried to use this in my model(lgbm) and it looks promising and my validation score did go up. \nI used this code to find attempt number.\n\ntrain[\"attempt_no\"] =1\ntrain[\"attemtp_no\"] = train[[\"user_id\",\"content_id\",\"attempt_no\"]].groupby([\"user_id\",\"content_id\"])[\"attempt_no\"].cumsum()\n\nbut the problem is that i can not do same in test data as test data comes in batch and we can not see wether all the content_id of a user are in same batch. So it will give inaccurte results.\n\nSo my question is. Is there any way to resolve this issue?",
    "1067262": "Interesting one",
    "1066983": "That's a nice find ",
    "1066105": "Yeah I find this insight too but for me it doesn't really boost my score idk what I do wrong",
    "1067805": "One thing which i don't understand is the headline in the section, (from their website)\n\n> \"Be sure to give priority to problems and lectures that raise the score. With an AI guide you can learn only what you need to improve your score.\"\n\nIn real life, we do learn a lot of things which aren't useful directly but it does helps a lot in the long run IMHO.",
    "1066061": "I have tried this feature out. My current model doesn't have a very good score but for what it is worth it gave a boost of ~0.005.",
    "1067263": "This is nice work, thank a  lot for sharing it",
    "1066178": "Hi @aravindpadman. Thanks a lot for sharing this nice work with us.",
    "1093295": "Nice work!\nThis feature, it works",
    "1090249": "Nice job!! Thank for sharing!! It helps us to make features for this situation."
  }
}