{
  "id": 192137,
  "title": "Any feature engineering ideas?",
  "url": "/competitions/riiid-test-answer-prediction/discussion/192137",
  "author_name": "",
  "post_date": "2020-10-20T10:24:01.207455600Z",
  "votes": 67,
  "comment_count": 35,
  "views": 0,
  "content": "<p>What I have tested:</p>\n<ol>\n<li>Flags that show that user watched at least one lecture of each part and each type (11 features)</li>\n<li>Flag that shows that user watched at least one lecture</li>\n<li>Merge 'bundle_id' and 'part' from questions data set</li>\n<li>Fill 'prior_question_elapsed_time' with mean</li>\n</ol>\n<p>BTW: All these features improve score</p>",
  "messages": [
    {
      "id": "1054945",
      "postDate": "10/20/2020 10:24:01",
      "content": "<p>What I have tested:</p>\n<ol>\n<li>Flags that show that user watched at least one lecture of each part and each type (11 features)</li>\n<li>Flag that shows that user watched at least one lecture</li>\n<li>Merge 'bundle_id' and 'part' from questions data set</li>\n<li>Fill 'prior_question_elapsed_time' with mean</li>\n</ol>\n<p>BTW: All these features improve score</p>",
      "rawMarkdown": "What I have tested:\n\n1. Flags that show that user watched at least one lecture of each part and each type (11 features)\n2. Flag that shows that user watched at least one lecture\n3. Merge 'bundle_id' and 'part' from questions data set\n4. Fill 'prior_question_elapsed_time' with mean\n\nBTW: All these features improve score",
      "votes": null
    },
    {
      "id": "1054980",
      "postDate": "10/20/2020 11:37:12",
      "content": "<p>I'm trying </p>\n<ul>\n<li>mean bundle accuracy</li>\n<li>mean part accuracy</li>\n<li>tag, user_id, bundle_id and part embeddings</li>\n</ul>\n<p>Not sure about the feature importance yet.</p>",
      "rawMarkdown": "I'm trying \n- mean bundle accuracy\n- mean part accuracy\n- tag, user_id, bundle_id and part embeddings\n\nNot sure about the feature importance yet.",
      "votes": null
    },
    {
      "id": "1055027",
      "postDate": "10/20/2020 12:22:20",
      "content": "<p>I am using Target Encoding for:</p>\n<ul>\n<li>user_id</li>\n<li>content_id</li>\n<li>task_container_id</li>\n<li>bundle_id</li>\n<li>part</li>\n</ul>",
      "rawMarkdown": "I am using Target Encoding for:\n- user_id\n- content_id\n- task_container_id\n- bundle_id\n- part",
      "votes": null
    },
    {
      "id": "1055053",
      "postDate": "10/20/2020 12:51:22",
      "content": "<p>How about seq2seq model</p>",
      "rawMarkdown": "How about seq2seq model",
      "votes": null
    },
    {
      "id": "1055393",
      "postDate": "10/20/2020 18:12:38",
      "content": "<blockquote>\n  <p>Flag that shows that user watched at least one lecture</p>\n</blockquote>\n<p>You might also want to create one feature per lecture <em>type</em>:</p>\n<ul>\n<li>\"concept\"</li>\n<li>\"solving question\"</li>\n<li>\"intention\"</li>\n<li>\"starter\"</li>\n</ul>",
      "rawMarkdown": "> Flag that shows that user watched at least one lecture\n\nYou might also want to create one feature per lecture _type_:\n- \"concept\"\n- \"solving question\"\n- \"intention\"\n- \"starter\"",
      "votes": null
    },
    {
      "id": "1055598",
      "postDate": "10/21/2020 00:12:59",
      "content": "<p>I'm just trying to make features about congeniality between users and questions.<br>\nSpecifically, I compare the quesion's average correct answer rate for each \"part\" with the user's one.<br>\nI'm just trying now, so Not sure about the feature importance yet.</p>",
      "rawMarkdown": "I'm just trying to make features about congeniality between users and questions.\nSpecifically, I compare the quesion's average correct answer rate for each \"part\" with the user's one.\nI'm just trying now, so Not sure about the feature importance yet.",
      "votes": null
    },
    {
      "id": "1055684",
      "postDate": "10/21/2020 02:26:16",
      "content": "<p>I have 11 features:</p>\n<pre><code>['part_1_boolean',\n 'part_2_boolean',\n 'part_3_boolean',\n 'part_4_boolean',\n 'part_5_boolean',\n 'part_6_boolean',\n 'part_7_boolean',\n 'type_of_concept_boolean',\n 'type_of_intention_boolean',\n 'type_of_solving_question_boolean',\n 'type_of_starter_boolean']\n</code></pre>\n<p>for lecture types and for parts.</p>",
      "rawMarkdown": "I have 11 features:\n\n```\n['part_1_boolean',\n 'part_2_boolean',\n 'part_3_boolean',\n 'part_4_boolean',\n 'part_5_boolean',\n 'part_6_boolean',\n 'part_7_boolean',\n 'type_of_concept_boolean',\n 'type_of_intention_boolean',\n 'type_of_solving_question_boolean',\n 'type_of_starter_boolean']\n```\n\nfor lecture types and for parts.",
      "votes": null
    },
    {
      "id": "1056354",
      "postDate": "10/21/2020 16:13:26",
      "content": "<p>I am trying to split the users to categories: smart ones, with high percentage of correctly answered questions and not. It is also correlating with summary time spent on the platform by the user.</p>",
      "rawMarkdown": "I am trying to split the users to categories: smart ones, with high percentage of correctly answered questions and not. It is also correlating with summary time spent on the platform by the user.",
      "votes": null
    },
    {
      "id": "1056355",
      "postDate": "10/21/2020 16:14:37",
      "content": "<p>Also I found the users which are both in train and test. I think, this could be useful.</p>",
      "rawMarkdown": "Also I found the users which are both in train and test. I think, this could be useful.",
      "votes": null
    },
    {
      "id": "1056422",
      "postDate": "10/21/2020 17:26:52",
      "content": "<p><a href=\"https://www.kaggle.com/anaidashaginian\" target=\"_blank\">@anaidashaginian</a> Related to this, you might find this notebook interesting: <a href=\"https://www.kaggle.com/datafan07/riiid-challenge-eda-baseline-model#Answer-Accuracy---Time-Relations\" target=\"_blank\">https://www.kaggle.com/datafan07/riiid-challenge-eda-baseline-model#Answer-Accuracy---Time-Relations</a></p>",
      "rawMarkdown": "anaidashaginian Related to this, you might find this notebook interesting: https://www.kaggle.com/datafan07/riiid-challenge-eda-baseline-model#Answer-Accuracy---Time-Relations",
      "votes": null
    },
    {
      "id": "1056427",
      "postDate": "10/21/2020 17:32:19",
      "content": "<p><a href=\"https://www.kaggle.com/anaidashaginian\" target=\"_blank\">@anaidashaginian</a> is that feature working for you? It's not working for me: <a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/190217\" target=\"_blank\">https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/190217</a></p>",
      "rawMarkdown": "anaidashaginian is that feature working for you? It's not working for me: https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/190217",
      "votes": null
    },
    {
      "id": "1056748",
      "postDate": "10/22/2020 03:35:50",
      "content": "<p><a href=\"https://www.kaggle.com/anaidashaginian\" target=\"_blank\">@anaidashaginian</a> Both in public and private test data??</p>",
      "rawMarkdown": "anaidashaginian Both in public and private test data??",
      "votes": null
    },
    {
      "id": "1060628",
      "postDate": "10/26/2020 12:18:02",
      "content": "<p>What about trying to figure out a sort of \"learning curves\"? Something like a cumulative number of lectures or questions related to each part or some group of tags (they also can be clustered somehow).<br>\nI'm working on it now.</p>",
      "rawMarkdown": "What about trying to figure out a sort of \"learning curves\"? Something like a cumulative number of lectures or questions related to each part or some group of tags (they also can be clustered somehow).\nI'm working on it now.",
      "votes": null
    },
    {
      "id": "1060813",
      "postDate": "10/26/2020 14:52:10",
      "content": "<p>Will look into it.</p>",
      "rawMarkdown": "Will look into it.",
      "votes": null
    },
    {
      "id": "1063734",
      "postDate": "10/29/2020 08:57:11",
      "content": "<p>Does anyone manage to extract relevant information from the user's answer ? </p>",
      "rawMarkdown": "Does anyone manage to extract relevant information from the user's answer ?",
      "votes": null
    },
    {
      "id": "1063749",
      "postDate": "10/29/2020 09:11:00",
      "content": "<p>because testcase donot contain user's answer</p>",
      "rawMarkdown": "because testcase donot contain user's answer",
      "votes": null
    },
    {
      "id": "1063754",
      "postDate": "10/29/2020 09:19:36",
      "content": "<p>Test data does contain user's answer in the <code>prior_group_responses</code> field.</p>",
      "rawMarkdown": "Test data does contain user's answer in the `prior_group_responses` field.",
      "votes": null
    },
    {
      "id": "1063764",
      "postDate": "10/29/2020 09:37:21",
      "content": "<p>I'm looking to update on the fly user's answer statistics in test set like some of us do for user answered correctly </p>",
      "rawMarkdown": "I'm looking to update on the fly user's answer statistics in test set like some of us do for user answered correctly",
      "votes": null
    },
    {
      "id": "1063765",
      "postDate": "10/29/2020 09:41:25",
      "content": "<p>I'm curious about the team name \"question the answers\"  in LB btw</p>",
      "rawMarkdown": "I'm curious about the team name \"question the answers\"  in LB btw",
      "votes": null
    },
    {
      "id": "1063978",
      "postDate": "10/29/2020 14:58:05",
      "content": "<blockquote>\n  <p>I'm looking to update on the fly user's answer statistics in test set like some of us do for user answered correctly</p>\n</blockquote>\n<p>Had tried a simple decay on past feats but it didn't help as updating the old stats with new will have little effect on the old one unless you compute them over a good sample size is my current understanding .</p>",
      "rawMarkdown": ">I'm looking to update on the fly user's answer statistics in test set like some of us do for user answered correctly\n\nHad tried a simple decay on past feats but it didn't help as updating the old stats with new will have little effect on the old one unless you compute them over a good sample size is my current understanding .",
      "votes": null
    },
    {
      "id": "1065125",
      "postDate": "10/30/2020 21:48:24",
      "content": "<p>I've implemented an expanding window average of target variable by user. I wrapped it with numba ant it runs all 100m rows in a minute.</p>\n<pre><code>import numba\n\n@numba.jit(nopython=True)\ndef expanding_target_mean(data):\n    \"\"\"\n    it takes user_id and answered correctly columns\n    and calculates expanding mean over target\n    \"\"\"\n    res = np.zeros_like(data[:,0])\n    user_id = 0\n    counter = 0\n    for i in range(len(res)):\n        if data[i, 0] != user_id:\n            user_id = data[i, 0]\n            res[i]  = data[i, 1]\n            counter = 0\n            continue\n        if data[i, 0] == user_id:\n            counter += 1\n            res[i] = np.mean(data[i-counter:i, 1])\n    return res\n\ndata[\"expanding_target_mean\"] = expanding_target_mean(data[['user_id', 'answered_correctly']].astype(float).values)\n</code></pre>",
      "rawMarkdown": "I've implemented an expanding window average of target variable by user. I wrapped it with numba ant it runs all 100m rows in a minute.\n```\nimport numba\n\n@numba.jit(nopython=True)\ndef expanding_target_mean(data):\n    \"\"\"\n    it takes user_id and answered correctly columns\n    and calculates expanding mean over target\n    \"\"\"\n    res = np.zeros_like(data[:,0])\n    user_id = 0\n    counter = 0\n    for i in range(len(res)):\n        if data[i, 0] != user_id:\n            user_id = data[i, 0]\n            res[i]  = data[i, 1]\n            counter = 0\n            continue\n        if data[i, 0] == user_id:\n            counter += 1\n            res[i] = np.mean(data[i-counter:i, 1])\n    return res\n\ndata[\"expanding_target_mean\"] = expanding_target_mean(data[['user_id', 'answered_correctly']].astype(float).values)\n\n```",
      "votes": null
    },
    {
      "id": "1065144",
      "postDate": "10/30/2020 22:39:05",
      "content": "<p>What is the <code>res</code> variable?</p>",
      "rawMarkdown": "What is the `res` variable?",
      "votes": null
    },
    {
      "id": "1065454",
      "postDate": "10/31/2020 09:57:32",
      "content": "<p>I have observed that content_id and bundle_id are the same thing if I am not missing something. So bundle_id does not add much?</p>",
      "rawMarkdown": "I have observed that content_id and bundle_id are the same thing if I am not missing something. So bundle_id does not add much?",
      "votes": null
    },
    {
      "id": "1065462",
      "postDate": "10/31/2020 10:10:34",
      "content": "<p>They aren't always the same. Some examples:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113389%2F8d592f77b0469f021312687c9dc76ad8%2FScreenshot%202020-10-31%20at%203.38.32%20PM.png?generation=1604139023264162&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "They aren't always the same. Some examples:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113389%2F8d592f77b0469f021312687c9dc76ad8%2FScreenshot%202020-10-31%20at%203.38.32%20PM.png?generation=1604139023264162&alt=media)",
      "votes": null
    },
    {
      "id": "1066908",
      "postDate": "11/02/2020 07:44:01",
      "content": "<p>Thanks for the detailed table. It seems bundle id can take same value for different question ids but these different values are almost same as question id (&lt;1%). Mathematically they are linearly correalated in the data, you can see it if you plot it. So not sure if it will add any additional info to the model.</p>",
      "rawMarkdown": "Thanks for the detailed table. It seems bundle id can take same value for different question ids but these different values are almost same as question id (<1%). Mathematically they are linearly correalated in the data, you can see it if you plot it. So not sure if it will add any additional info to the model.",
      "votes": null
    },
    {
      "id": "1069157",
      "postDate": "11/04/2020 06:41:00",
      "content": "<p>I feel t's crucial to engineer features which one can translate to the test set optimally as well..</p>",
      "rawMarkdown": "I feel t's crucial to engineer features which one can translate to the test set optimally as well..",
      "votes": null
    },
    {
      "id": "1075040",
      "postDate": "11/11/2020 10:52:23",
      "content": "<p>So I tried the following:</p>\n<pre><code>X[\"correct_user_sum\"] = X[[\"user_id\",'answered_correctly']].groupby([\"user_id\"])['answered_correctly'].cumsum()\nX['correct_user_sum'] = X['correct_user_sum']-1\n</code></pre>\n<p>Last line is to remove leakage<br>\nThis feature performs really really well localy (around 0.80 AUC), but very very poorly on the LB, I have no idea why.<br>\nI've already tried different CV strategies and it always performs the same</p>",
      "rawMarkdown": "So I tried the following:\n\n```\nX[\"correct_user_sum\"] = X[[\"user_id\",'answered_correctly']].groupby([\"user_id\"])['answered_correctly'].cumsum()\nX['correct_user_sum'] = X['correct_user_sum']-1\n```\n\nLast line is to remove leakage\nThis feature performs really really well localy (around 0.80 AUC), but very very poorly on the LB, I have no idea why.\nI've already tried different CV strategies and it always performs the same",
      "votes": null
    },
    {
      "id": "1075050",
      "postDate": "11/11/2020 11:04:41",
      "content": "<p>How are you dealing with the new information and new users in the test set during submission?</p>",
      "rawMarkdown": "How are you dealing with the new information and new users in the test set during submission?",
      "votes": null
    },
    {
      "id": "1075062",
      "postDate": "11/11/2020 11:11:36",
      "content": "<p>Tried simple merge and FillNa, then tried just a simple [correct_user_sum+answered_correctly]. Both cases gave the same result. All my features are put into Series, so it's easy to update and merge</p>",
      "rawMarkdown": "Tried simple merge and FillNa, then tried just a simple [correct_user_sum+answered_correctly]. Both cases gave the same result. All my features are put into Series, so it's easy to update and merge",
      "votes": null
    },
    {
      "id": "1075065",
      "postDate": "11/11/2020 11:14:41",
      "content": "<p>this is my update funtion that gets a Df with User-ID in previous loop and there answers in the next loop</p>\n<pre><code>def registrar(df):\n    reg = cudf.DataFrame()\n    reg['user_id'] = df['user_id']\n    reg['answered_correctly'] = df['answered_correctly']\n    reg['answered_wrong'] = 1-reg['answered_correctly']\n    #reg = cudf.merge(reg, cudf.DataFrame(wrong_user_sum[[\"wrong_user_sum\",'user_id']]), right_on='user_id',left_on='user_id',how='left')\n    #reg = cudf.merge(reg, cudf.DataFrame(corrrect_user_sum[[\"correct_user_sum\",'user_id']]), right_on='user_id',left_on='user_id',how='left')\n    reg = cudf.merge(reg, cudf.DataFrame(answer_wrong_user[['user_id','answer_wrong_user']]), right_on='user_id',left_on='user_id',how='left')\n    reg = cudf.merge(reg, cudf.DataFrame(correct_user[['user_id','answered_correctly_user']]), right_on='user_id',left_on='user_id',how='left')\n    reg = cudf.merge(reg, cudf.DataFrame(user_questions[['user_id','User_questions']]), right_on='user_id',left_on='user_id',how='left')\n\n    reg['User_questions'].fillna(0,inplace=True)\n    reg['answered_correctly_user'].fillna(np.float32(0.643215),inplace=True)\n    reg['answer_wrong_user'].fillna(np.float32(1-0.643215),inplace=True)\n    #reg['wrong_user_sum'].fillna(0,inplace=True)\n    #reg['correct_user_sum'].fillna(0,inplace=True)\n\n    reg['answered_correctly_user'] = ((reg['answered_correctly_user']*reg['User_questions'])+reg['answered_correctly'])/(reg['User_questions']+1)\n    reg['answer_wrong_user'] = 1-reg['answered_correctly_user']\n    #reg['wrong_user_sum'] = reg['wrong_user_sum'] + reg['answered_wrong']\n    #reg['correct_user_sum'] = reg['correct_user_sum'] + reg['answered_correctly']\n    reg['User_questions'] = reg['User_questions']+1\n\n    user_questions.append(reg[['user_id','User_questions']].to_pandas())\n    correct_user.append(reg[['user_id','answered_correctly_user']].to_pandas())\n    answer_wrong_user.append(reg[['user_id','answer_wrong_user']].to_pandas())\n    #corrrect_user_sum.append(reg[[\"correct_user_sum\",'user_id']].to_pandas())\n    #wrong_user_sum.append(reg[[\"wrong_user_sum\",'user_id']].to_pandas())\n\n    user_questions.drop_duplicates(subset=['user_id'], keep='last')\n    correct_user.drop_duplicates(subset=['user_id'], keep='last')\n    answer_wrong_user.drop_duplicates(subset=['user_id'], keep='last')\n    #corrrect_user_sum.drop_duplicates(subset=['user_id'], keep='last')\n    #wrong_user_sum.drop_duplicates(subset=['user_id'], keep='last')\n</code></pre>",
      "rawMarkdown": "this is my update funtion that gets a Df with User-ID in previous loop and there answers in the next loop\n\n```\ndef registrar(df):\n    reg = cudf.DataFrame()\n    reg['user_id'] = df['user_id']\n    reg['answered_correctly'] = df['answered_correctly']\n    reg['answered_wrong'] = 1-reg['answered_correctly']\n    #reg = cudf.merge(reg, cudf.DataFrame(wrong_user_sum[[\"wrong_user_sum\",'user_id']]), right_on='user_id',left_on='user_id',how='left')\n    #reg = cudf.merge(reg, cudf.DataFrame(corrrect_user_sum[[\"correct_user_sum\",'user_id']]), right_on='user_id',left_on='user_id',how='left')\n    reg = cudf.merge(reg, cudf.DataFrame(answer_wrong_user[['user_id','answer_wrong_user']]), right_on='user_id',left_on='user_id',how='left')\n    reg = cudf.merge(reg, cudf.DataFrame(correct_user[['user_id','answered_correctly_user']]), right_on='user_id',left_on='user_id',how='left')\n    reg = cudf.merge(reg, cudf.DataFrame(user_questions[['user_id','User_questions']]), right_on='user_id',left_on='user_id',how='left')\n    \n    reg['User_questions'].fillna(0,inplace=True)\n    reg['answered_correctly_user'].fillna(np.float32(0.643215),inplace=True)\n    reg['answer_wrong_user'].fillna(np.float32(1-0.643215),inplace=True)\n    #reg['wrong_user_sum'].fillna(0,inplace=True)\n    #reg['correct_user_sum'].fillna(0,inplace=True)\n    \n    reg['answered_correctly_user'] = ((reg['answered_correctly_user']*reg['User_questions'])+reg['answered_correctly'])/(reg['User_questions']+1)\n    reg['answer_wrong_user'] = 1-reg['answered_correctly_user']\n    #reg['wrong_user_sum'] = reg['wrong_user_sum'] + reg['answered_wrong']\n    #reg['correct_user_sum'] = reg['correct_user_sum'] + reg['answered_correctly']\n    reg['User_questions'] = reg['User_questions']+1\n    \n    user_questions.append(reg[['user_id','User_questions']].to_pandas())\n    correct_user.append(reg[['user_id','answered_correctly_user']].to_pandas())\n    answer_wrong_user.append(reg[['user_id','answer_wrong_user']].to_pandas())\n    #corrrect_user_sum.append(reg[[\"correct_user_sum\",'user_id']].to_pandas())\n    #wrong_user_sum.append(reg[[\"wrong_user_sum\",'user_id']].to_pandas())\n    \n    user_questions.drop_duplicates(subset=['user_id'], keep='last')\n    correct_user.drop_duplicates(subset=['user_id'], keep='last')\n    answer_wrong_user.drop_duplicates(subset=['user_id'], keep='last')\n    #corrrect_user_sum.drop_duplicates(subset=['user_id'], keep='last')\n    #wrong_user_sum.drop_duplicates(subset=['user_id'], keep='last')\n    \n    \n```",
      "votes": null
    },
    {
      "id": "1075067",
      "postDate": "11/11/2020 11:15:53",
      "content": "<p>Have a look at <a href=\"https://www.kaggle.com/its7171/lgbm-with-loop-feature-engineering\" target=\"_blank\">https://www.kaggle.com/its7171/lgbm-with-loop-feature-engineering</a> for ideas on how to achieve something similar to this. I think <code>count_u</code> would be the same as your <code>correct_user_sum</code></p>",
      "rawMarkdown": "Have a look at https://www.kaggle.com/its7171/lgbm-with-loop-feature-engineering for ideas on how to achieve something similar to this. I think `count_u` would be the same as your `correct_user_sum`",
      "votes": null
    },
    {
      "id": "1075567",
      "postDate": "11/11/2020 19:29:55",
      "content": "<p>X['correct_user_sum']-1 this may not really prevent leakage</p>",
      "rawMarkdown": "X['correct_user_sum']-1 this may not really prevent leakage",
      "votes": null
    },
    {
      "id": "1075595",
      "postDate": "11/11/2020 19:56:23",
      "content": "<p>I suspect it as well. Tho I don't know why. Since its cumsum and done by row, it should theoretically remove answers from the same row. Even looking at the data frame it looks ok  gives the same result if I use a shift(1)</p>",
      "rawMarkdown": "I suspect it as well. Tho I don't know why. Since its cumsum and done by row, it should theoretically remove answers from the same row. Even looking at the data frame it looks ok  gives the same result if I use a shift(1)",
      "votes": null
    },
    {
      "id": "1090501",
      "postDate": "11/25/2020 11:29:33",
      "content": "<p>A leak free way to do this would be:</p>\n<pre><code>X[\"correct_user_sum\"] = X[[\"user_id\",'answered_correctly']].groupby([\"user_id\"])['answered_correctly'].apply(lambda x: x.cumsum().shift())\n</code></pre>\n<p>Doing a <code>X[\"correct_user_sum\"] = X[\"correct_user_sum\"] - 1</code> would still introduce a leak as <code>correct_user_sum</code> for that row would <em>still be greater than previous <code>correct_user_sum</code> by 1 unit</em>. I don't really know how a non-sequence model such as LGB could possibly exploit this leak.</p>",
      "rawMarkdown": "A leak free way to do this would be:\n\n```\nX[\"correct_user_sum\"] = X[[\"user_id\",'answered_correctly']].groupby([\"user_id\"])['answered_correctly'].apply(lambda x: x.cumsum().shift())\n```\n\nDoing a ```X[\"correct_user_sum\"] = X[\"correct_user_sum\"] - 1``` would still introduce a leak as `correct_user_sum` for that row would _still be greater than previous `correct_user_sum` by 1 unit_. I don't really know how a non-sequence model such as LGB could possibly exploit this leak.",
      "votes": null
    },
    {
      "id": "1090536",
      "postDate": "11/25/2020 11:56:22",
      "content": "<p>Is this working within a reasonable time?</p>",
      "rawMarkdown": "Is this working within a reasonable time?",
      "votes": null
    },
    {
      "id": "1150608",
      "postDate": "01/12/2021 17:45:19",
      "content": "<p>Hi,<br>\nThanks. I have used some of the insights here:<br>\n<a href=\"https://www.kaggle.com/kritidoneria/beginner-wids21-feature-engineering-starter\" target=\"_blank\">https://www.kaggle.com/kritidoneria/beginner-wids21-feature-engineering-starter</a></p>",
      "rawMarkdown": "Hi,\nThanks. I have used some of the insights here:\nhttps://www.kaggle.com/kritidoneria/beginner-wids21-feature-engineering-starter",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1054980,
      "author_name": "kaushal2896",
      "author_url": "",
      "post_date": "10/20/2020 11:37:12",
      "content": "<p>I'm trying </p>\n<ul>\n<li>mean bundle accuracy</li>\n<li>mean part accuracy</li>\n<li>tag, user_id, bundle_id and part embeddings</li>\n</ul>\n<p>Not sure about the feature importance yet.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1055027,
          "author_name": "pavelvpster",
          "author_url": "",
          "post_date": "10/20/2020 12:22:20",
          "content": "<p>I am using Target Encoding for:</p>\n<ul>\n<li>user_id</li>\n<li>content_id</li>\n<li>task_container_id</li>\n<li>bundle_id</li>\n<li>part</li>\n</ul>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1055053,
      "author_name": "msharuk589",
      "author_url": "",
      "post_date": "10/20/2020 12:51:22",
      "content": "<p>How about seq2seq model</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1055393,
      "author_name": "nyakaggle",
      "author_url": "",
      "post_date": "10/20/2020 18:12:38",
      "content": "<blockquote>\n  <p>Flag that shows that user watched at least one lecture</p>\n</blockquote>\n<p>You might also want to create one feature per lecture <em>type</em>:</p>\n<ul>\n<li>\"concept\"</li>\n<li>\"solving question\"</li>\n<li>\"intention\"</li>\n<li>\"starter\"</li>\n</ul>",
      "votes": null,
      "replies": [
        {
          "id": 1055684,
          "author_name": "pavelvpster",
          "author_url": "",
          "post_date": "10/21/2020 02:26:16",
          "content": "<p>I have 11 features:</p>\n<pre><code>['part_1_boolean',\n 'part_2_boolean',\n 'part_3_boolean',\n 'part_4_boolean',\n 'part_5_boolean',\n 'part_6_boolean',\n 'part_7_boolean',\n 'type_of_concept_boolean',\n 'type_of_intention_boolean',\n 'type_of_solving_question_boolean',\n 'type_of_starter_boolean']\n</code></pre>\n<p>for lecture types and for parts.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1055598,
      "author_name": "suzukiseiya",
      "author_url": "",
      "post_date": "10/21/2020 00:12:59",
      "content": "<p>I'm just trying to make features about congeniality between users and questions.<br>\nSpecifically, I compare the quesion's average correct answer rate for each \"part\" with the user's one.<br>\nI'm just trying now, so Not sure about the feature importance yet.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1056354,
      "author_name": "anaidashaginian",
      "author_url": "",
      "post_date": "10/21/2020 16:13:26",
      "content": "<p>I am trying to split the users to categories: smart ones, with high percentage of correctly answered questions and not. It is also correlating with summary time spent on the platform by the user.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1056422,
          "author_name": "nyakaggle",
          "author_url": "",
          "post_date": "10/21/2020 17:26:52",
          "content": "<p><a href=\"https://www.kaggle.com/anaidashaginian\" target=\"_blank\">@anaidashaginian</a> Related to this, you might find this notebook interesting: <a href=\"https://www.kaggle.com/datafan07/riiid-challenge-eda-baseline-model#Answer-Accuracy---Time-Relations\" target=\"_blank\">https://www.kaggle.com/datafan07/riiid-challenge-eda-baseline-model#Answer-Accuracy---Time-Relations</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1056427,
          "author_name": "kaushal2896",
          "author_url": "",
          "post_date": "10/21/2020 17:32:19",
          "content": "<p><a href=\"https://www.kaggle.com/anaidashaginian\" target=\"_blank\">@anaidashaginian</a> is that feature working for you? It's not working for me: <a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/190217\" target=\"_blank\">https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/190217</a></p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1056355,
      "author_name": "anaidashaginian",
      "author_url": "",
      "post_date": "10/21/2020 16:14:37",
      "content": "<p>Also I found the users which are both in train and test. I think, this could be useful.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1056748,
          "author_name": "msharuk589",
          "author_url": "",
          "post_date": "10/22/2020 03:35:50",
          "content": "<p><a href=\"https://www.kaggle.com/anaidashaginian\" target=\"_blank\">@anaidashaginian</a> Both in public and private test data??</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1060628,
      "author_name": "valentinart",
      "author_url": "",
      "post_date": "10/26/2020 12:18:02",
      "content": "<p>What about trying to figure out a sort of \"learning curves\"? Something like a cumulative number of lectures or questions related to each part or some group of tags (they also can be clustered somehow).<br>\nI'm working on it now.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1060813,
      "author_name": "jatta3399",
      "author_url": "",
      "post_date": "10/26/2020 14:52:10",
      "content": "<p>Will look into it.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1063734,
      "author_name": "alexj21",
      "author_url": "",
      "post_date": "10/29/2020 08:57:11",
      "content": "<p>Does anyone manage to extract relevant information from the user's answer ? </p>",
      "votes": null,
      "replies": [
        {
          "id": 1063749,
          "author_name": "pp2file",
          "author_url": "",
          "post_date": "10/29/2020 09:11:00",
          "content": "<p>because testcase donot contain user's answer</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1063754,
          "author_name": "rohanrao",
          "author_url": "",
          "post_date": "10/29/2020 09:19:36",
          "content": "<p>Test data does contain user's answer in the <code>prior_group_responses</code> field.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1063764,
          "author_name": "alexj21",
          "author_url": "",
          "post_date": "10/29/2020 09:37:21",
          "content": "<p>I'm looking to update on the fly user's answer statistics in test set like some of us do for user answered correctly </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1063765,
          "author_name": "alexj21",
          "author_url": "",
          "post_date": "10/29/2020 09:41:25",
          "content": "<p>I'm curious about the team name \"question the answers\"  in LB btw</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1063978,
          "author_name": "adityaecdrid",
          "author_url": "",
          "post_date": "10/29/2020 14:58:05",
          "content": "<blockquote>\n  <p>I'm looking to update on the fly user's answer statistics in test set like some of us do for user answered correctly</p>\n</blockquote>\n<p>Had tried a simple decay on past feats but it didn't help as updating the old stats with new will have little effect on the old one unless you compute them over a good sample size is my current understanding .</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1065125,
      "author_name": "bturan19",
      "author_url": "",
      "post_date": "10/30/2020 21:48:24",
      "content": "<p>I've implemented an expanding window average of target variable by user. I wrapped it with numba ant it runs all 100m rows in a minute.</p>\n<pre><code>import numba\n\n@numba.jit(nopython=True)\ndef expanding_target_mean(data):\n    \"\"\"\n    it takes user_id and answered correctly columns\n    and calculates expanding mean over target\n    \"\"\"\n    res = np.zeros_like(data[:,0])\n    user_id = 0\n    counter = 0\n    for i in range(len(res)):\n        if data[i, 0] != user_id:\n            user_id = data[i, 0]\n            res[i]  = data[i, 1]\n            counter = 0\n            continue\n        if data[i, 0] == user_id:\n            counter += 1\n            res[i] = np.mean(data[i-counter:i, 1])\n    return res\n\ndata[\"expanding_target_mean\"] = expanding_target_mean(data[['user_id', 'answered_correctly']].astype(float).values)\n</code></pre>",
      "votes": null,
      "replies": [
        {
          "id": 1065144,
          "author_name": "nyakaggle",
          "author_url": "",
          "post_date": "10/30/2020 22:39:05",
          "content": "<p>What is the <code>res</code> variable?</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1065454,
      "author_name": "onurbaris",
      "author_url": "",
      "post_date": "10/31/2020 09:57:32",
      "content": "<p>I have observed that content_id and bundle_id are the same thing if I am not missing something. So bundle_id does not add much?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1065462,
          "author_name": "rohanrao",
          "author_url": "",
          "post_date": "10/31/2020 10:10:34",
          "content": "<p>They aren't always the same. Some examples:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113389%2F8d592f77b0469f021312687c9dc76ad8%2FScreenshot%202020-10-31%20at%203.38.32%20PM.png?generation=1604139023264162&amp;alt=media\" alt=\"\"></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1066908,
          "author_name": "onurbaris",
          "author_url": "",
          "post_date": "11/02/2020 07:44:01",
          "content": "<p>Thanks for the detailed table. It seems bundle id can take same value for different question ids but these different values are almost same as question id (&lt;1%). Mathematically they are linearly correalated in the data, you can see it if you plot it. So not sure if it will add any additional info to the model.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1069157,
      "author_name": "adityaecdrid",
      "author_url": "",
      "post_date": "11/04/2020 06:41:00",
      "content": "<p>I feel t's crucial to engineer features which one can translate to the test set optimally as well..</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1075040,
      "author_name": "iuryck",
      "author_url": "",
      "post_date": "11/11/2020 10:52:23",
      "content": "<p>So I tried the following:</p>\n<pre><code>X[\"correct_user_sum\"] = X[[\"user_id\",'answered_correctly']].groupby([\"user_id\"])['answered_correctly'].cumsum()\nX['correct_user_sum'] = X['correct_user_sum']-1\n</code></pre>\n<p>Last line is to remove leakage<br>\nThis feature performs really really well localy (around 0.80 AUC), but very very poorly on the LB, I have no idea why.<br>\nI've already tried different CV strategies and it always performs the same</p>",
      "votes": null,
      "replies": [
        {
          "id": 1075050,
          "author_name": "ghostskipper",
          "author_url": "",
          "post_date": "11/11/2020 11:04:41",
          "content": "<p>How are you dealing with the new information and new users in the test set during submission?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1075062,
          "author_name": "iuryck",
          "author_url": "",
          "post_date": "11/11/2020 11:11:36",
          "content": "<p>Tried simple merge and FillNa, then tried just a simple [correct_user_sum+answered_correctly]. Both cases gave the same result. All my features are put into Series, so it's easy to update and merge</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1075065,
          "author_name": "iuryck",
          "author_url": "",
          "post_date": "11/11/2020 11:14:41",
          "content": "<p>this is my update funtion that gets a Df with User-ID in previous loop and there answers in the next loop</p>\n<pre><code>def registrar(df):\n    reg = cudf.DataFrame()\n    reg['user_id'] = df['user_id']\n    reg['answered_correctly'] = df['answered_correctly']\n    reg['answered_wrong'] = 1-reg['answered_correctly']\n    #reg = cudf.merge(reg, cudf.DataFrame(wrong_user_sum[[\"wrong_user_sum\",'user_id']]), right_on='user_id',left_on='user_id',how='left')\n    #reg = cudf.merge(reg, cudf.DataFrame(corrrect_user_sum[[\"correct_user_sum\",'user_id']]), right_on='user_id',left_on='user_id',how='left')\n    reg = cudf.merge(reg, cudf.DataFrame(answer_wrong_user[['user_id','answer_wrong_user']]), right_on='user_id',left_on='user_id',how='left')\n    reg = cudf.merge(reg, cudf.DataFrame(correct_user[['user_id','answered_correctly_user']]), right_on='user_id',left_on='user_id',how='left')\n    reg = cudf.merge(reg, cudf.DataFrame(user_questions[['user_id','User_questions']]), right_on='user_id',left_on='user_id',how='left')\n\n    reg['User_questions'].fillna(0,inplace=True)\n    reg['answered_correctly_user'].fillna(np.float32(0.643215),inplace=True)\n    reg['answer_wrong_user'].fillna(np.float32(1-0.643215),inplace=True)\n    #reg['wrong_user_sum'].fillna(0,inplace=True)\n    #reg['correct_user_sum'].fillna(0,inplace=True)\n\n    reg['answered_correctly_user'] = ((reg['answered_correctly_user']*reg['User_questions'])+reg['answered_correctly'])/(reg['User_questions']+1)\n    reg['answer_wrong_user'] = 1-reg['answered_correctly_user']\n    #reg['wrong_user_sum'] = reg['wrong_user_sum'] + reg['answered_wrong']\n    #reg['correct_user_sum'] = reg['correct_user_sum'] + reg['answered_correctly']\n    reg['User_questions'] = reg['User_questions']+1\n\n    user_questions.append(reg[['user_id','User_questions']].to_pandas())\n    correct_user.append(reg[['user_id','answered_correctly_user']].to_pandas())\n    answer_wrong_user.append(reg[['user_id','answer_wrong_user']].to_pandas())\n    #corrrect_user_sum.append(reg[[\"correct_user_sum\",'user_id']].to_pandas())\n    #wrong_user_sum.append(reg[[\"wrong_user_sum\",'user_id']].to_pandas())\n\n    user_questions.drop_duplicates(subset=['user_id'], keep='last')\n    correct_user.drop_duplicates(subset=['user_id'], keep='last')\n    answer_wrong_user.drop_duplicates(subset=['user_id'], keep='last')\n    #corrrect_user_sum.drop_duplicates(subset=['user_id'], keep='last')\n    #wrong_user_sum.drop_duplicates(subset=['user_id'], keep='last')\n</code></pre>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1075067,
          "author_name": "ghostskipper",
          "author_url": "",
          "post_date": "11/11/2020 11:15:53",
          "content": "<p>Have a look at <a href=\"https://www.kaggle.com/its7171/lgbm-with-loop-feature-engineering\" target=\"_blank\">https://www.kaggle.com/its7171/lgbm-with-loop-feature-engineering</a> for ideas on how to achieve something similar to this. I think <code>count_u</code> would be the same as your <code>correct_user_sum</code></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1075567,
          "author_name": "bturan19",
          "author_url": "",
          "post_date": "11/11/2020 19:29:55",
          "content": "<p>X['correct_user_sum']-1 this may not really prevent leakage</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1075595,
          "author_name": "iuryck",
          "author_url": "",
          "post_date": "11/11/2020 19:56:23",
          "content": "<p>I suspect it as well. Tho I don't know why. Since its cumsum and done by row, it should theoretically remove answers from the same row. Even looking at the data frame it looks ok  gives the same result if I use a shift(1)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1090501,
          "author_name": "doctorkael",
          "author_url": "",
          "post_date": "11/25/2020 11:29:33",
          "content": "<p>A leak free way to do this would be:</p>\n<pre><code>X[\"correct_user_sum\"] = X[[\"user_id\",'answered_correctly']].groupby([\"user_id\"])['answered_correctly'].apply(lambda x: x.cumsum().shift())\n</code></pre>\n<p>Doing a <code>X[\"correct_user_sum\"] = X[\"correct_user_sum\"] - 1</code> would still introduce a leak as <code>correct_user_sum</code> for that row would <em>still be greater than previous <code>correct_user_sum</code> by 1 unit</em>. I don't really know how a non-sequence model such as LGB could possibly exploit this leak.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1090536,
          "author_name": "bturan19",
          "author_url": "",
          "post_date": "11/25/2020 11:56:22",
          "content": "<p>Is this working within a reasonable time?</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1150608,
      "author_name": "kritidoneria",
      "author_url": "",
      "post_date": "01/12/2021 17:45:19",
      "content": "<p>Hi,<br>\nThanks. I have used some of the insights here:<br>\n<a href=\"https://www.kaggle.com/kritidoneria/beginner-wids21-feature-engineering-starter\" target=\"_blank\">https://www.kaggle.com/kritidoneria/beginner-wids21-feature-engineering-starter</a></p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1054945": "What I have tested:\n\n1. Flags that show that user watched at least one lecture of each part and each type (11 features)\n2. Flag that shows that user watched at least one lecture\n3. Merge 'bundle_id' and 'part' from questions data set\n4. Fill 'prior_question_elapsed_time' with mean\n\nBTW: All these features improve score",
    "1054980": "I'm trying \n- mean bundle accuracy\n- mean part accuracy\n- tag, user_id, bundle_id and part embeddings\n\nNot sure about the feature importance yet.",
    "1055027": "I am using Target Encoding for:\n- user_id\n- content_id\n- task_container_id\n- bundle_id\n- part",
    "1055053": "How about seq2seq model",
    "1055393": "> Flag that shows that user watched at least one lecture\n\nYou might also want to create one feature per lecture _type_:\n- \"concept\"\n- \"solving question\"\n- \"intention\"\n- \"starter\"",
    "1055598": "I'm just trying to make features about congeniality between users and questions.\nSpecifically, I compare the quesion's average correct answer rate for each \"part\" with the user's one.\nI'm just trying now, so Not sure about the feature importance yet.",
    "1055684": "I have 11 features:\n\n```\n['part_1_boolean',\n 'part_2_boolean',\n 'part_3_boolean',\n 'part_4_boolean',\n 'part_5_boolean',\n 'part_6_boolean',\n 'part_7_boolean',\n 'type_of_concept_boolean',\n 'type_of_intention_boolean',\n 'type_of_solving_question_boolean',\n 'type_of_starter_boolean']\n```\n\nfor lecture types and for parts.",
    "1056354": "I am trying to split the users to categories: smart ones, with high percentage of correctly answered questions and not. It is also correlating with summary time spent on the platform by the user.",
    "1056355": "Also I found the users which are both in train and test. I think, this could be useful.",
    "1056422": "anaidashaginian Related to this, you might find this notebook interesting: https://www.kaggle.com/datafan07/riiid-challenge-eda-baseline-model#Answer-Accuracy---Time-Relations",
    "1056427": "anaidashaginian is that feature working for you? It's not working for me: https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/190217",
    "1056748": "anaidashaginian Both in public and private test data??",
    "1060628": "What about trying to figure out a sort of \"learning curves\"? Something like a cumulative number of lectures or questions related to each part or some group of tags (they also can be clustered somehow).\nI'm working on it now.",
    "1060813": "Will look into it.",
    "1063734": "Does anyone manage to extract relevant information from the user's answer ?",
    "1063749": "because testcase donot contain user's answer",
    "1063754": "Test data does contain user's answer in the `prior_group_responses` field.",
    "1063764": "I'm looking to update on the fly user's answer statistics in test set like some of us do for user answered correctly",
    "1063765": "I'm curious about the team name \"question the answers\"  in LB btw",
    "1063978": ">I'm looking to update on the fly user's answer statistics in test set like some of us do for user answered correctly\n\nHad tried a simple decay on past feats but it didn't help as updating the old stats with new will have little effect on the old one unless you compute them over a good sample size is my current understanding .",
    "1065125": "I've implemented an expanding window average of target variable by user. I wrapped it with numba ant it runs all 100m rows in a minute.\n```\nimport numba\n\n@numba.jit(nopython=True)\ndef expanding_target_mean(data):\n    \"\"\"\n    it takes user_id and answered correctly columns\n    and calculates expanding mean over target\n    \"\"\"\n    res = np.zeros_like(data[:,0])\n    user_id = 0\n    counter = 0\n    for i in range(len(res)):\n        if data[i, 0] != user_id:\n            user_id = data[i, 0]\n            res[i]  = data[i, 1]\n            counter = 0\n            continue\n        if data[i, 0] == user_id:\n            counter += 1\n            res[i] = np.mean(data[i-counter:i, 1])\n    return res\n\ndata[\"expanding_target_mean\"] = expanding_target_mean(data[['user_id', 'answered_correctly']].astype(float).values)\n\n```",
    "1065144": "What is the `res` variable?",
    "1065454": "I have observed that content_id and bundle_id are the same thing if I am not missing something. So bundle_id does not add much?",
    "1065462": "They aren't always the same. Some examples:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113389%2F8d592f77b0469f021312687c9dc76ad8%2FScreenshot%202020-10-31%20at%203.38.32%20PM.png?generation=1604139023264162&alt=media)",
    "1066908": "Thanks for the detailed table. It seems bundle id can take same value for different question ids but these different values are almost same as question id (<1%). Mathematically they are linearly correalated in the data, you can see it if you plot it. So not sure if it will add any additional info to the model.",
    "1069157": "I feel t's crucial to engineer features which one can translate to the test set optimally as well..",
    "1075040": "So I tried the following:\n\n```\nX[\"correct_user_sum\"] = X[[\"user_id\",'answered_correctly']].groupby([\"user_id\"])['answered_correctly'].cumsum()\nX['correct_user_sum'] = X['correct_user_sum']-1\n```\n\nLast line is to remove leakage\nThis feature performs really really well localy (around 0.80 AUC), but very very poorly on the LB, I have no idea why.\nI've already tried different CV strategies and it always performs the same",
    "1075050": "How are you dealing with the new information and new users in the test set during submission?",
    "1075062": "Tried simple merge and FillNa, then tried just a simple [correct_user_sum+answered_correctly]. Both cases gave the same result. All my features are put into Series, so it's easy to update and merge",
    "1075065": "this is my update funtion that gets a Df with User-ID in previous loop and there answers in the next loop\n\n```\ndef registrar(df):\n    reg = cudf.DataFrame()\n    reg['user_id'] = df['user_id']\n    reg['answered_correctly'] = df['answered_correctly']\n    reg['answered_wrong'] = 1-reg['answered_correctly']\n    #reg = cudf.merge(reg, cudf.DataFrame(wrong_user_sum[[\"wrong_user_sum\",'user_id']]), right_on='user_id',left_on='user_id',how='left')\n    #reg = cudf.merge(reg, cudf.DataFrame(corrrect_user_sum[[\"correct_user_sum\",'user_id']]), right_on='user_id',left_on='user_id',how='left')\n    reg = cudf.merge(reg, cudf.DataFrame(answer_wrong_user[['user_id','answer_wrong_user']]), right_on='user_id',left_on='user_id',how='left')\n    reg = cudf.merge(reg, cudf.DataFrame(correct_user[['user_id','answered_correctly_user']]), right_on='user_id',left_on='user_id',how='left')\n    reg = cudf.merge(reg, cudf.DataFrame(user_questions[['user_id','User_questions']]), right_on='user_id',left_on='user_id',how='left')\n    \n    reg['User_questions'].fillna(0,inplace=True)\n    reg['answered_correctly_user'].fillna(np.float32(0.643215),inplace=True)\n    reg['answer_wrong_user'].fillna(np.float32(1-0.643215),inplace=True)\n    #reg['wrong_user_sum'].fillna(0,inplace=True)\n    #reg['correct_user_sum'].fillna(0,inplace=True)\n    \n    reg['answered_correctly_user'] = ((reg['answered_correctly_user']*reg['User_questions'])+reg['answered_correctly'])/(reg['User_questions']+1)\n    reg['answer_wrong_user'] = 1-reg['answered_correctly_user']\n    #reg['wrong_user_sum'] = reg['wrong_user_sum'] + reg['answered_wrong']\n    #reg['correct_user_sum'] = reg['correct_user_sum'] + reg['answered_correctly']\n    reg['User_questions'] = reg['User_questions']+1\n    \n    user_questions.append(reg[['user_id','User_questions']].to_pandas())\n    correct_user.append(reg[['user_id','answered_correctly_user']].to_pandas())\n    answer_wrong_user.append(reg[['user_id','answer_wrong_user']].to_pandas())\n    #corrrect_user_sum.append(reg[[\"correct_user_sum\",'user_id']].to_pandas())\n    #wrong_user_sum.append(reg[[\"wrong_user_sum\",'user_id']].to_pandas())\n    \n    user_questions.drop_duplicates(subset=['user_id'], keep='last')\n    correct_user.drop_duplicates(subset=['user_id'], keep='last')\n    answer_wrong_user.drop_duplicates(subset=['user_id'], keep='last')\n    #corrrect_user_sum.drop_duplicates(subset=['user_id'], keep='last')\n    #wrong_user_sum.drop_duplicates(subset=['user_id'], keep='last')\n    \n    \n```",
    "1075067": "Have a look at https://www.kaggle.com/its7171/lgbm-with-loop-feature-engineering for ideas on how to achieve something similar to this. I think `count_u` would be the same as your `correct_user_sum`",
    "1075567": "X['correct_user_sum']-1 this may not really prevent leakage",
    "1075595": "I suspect it as well. Tho I don't know why. Since its cumsum and done by row, it should theoretically remove answers from the same row. Even looking at the data frame it looks ok  gives the same result if I use a shift(1)",
    "1090501": "A leak free way to do this would be:\n\n```\nX[\"correct_user_sum\"] = X[[\"user_id\",'answered_correctly']].groupby([\"user_id\"])['answered_correctly'].apply(lambda x: x.cumsum().shift())\n```\n\nDoing a ```X[\"correct_user_sum\"] = X[\"correct_user_sum\"] - 1``` would still introduce a leak as `correct_user_sum` for that row would _still be greater than previous `correct_user_sum` by 1 unit_. I don't really know how a non-sequence model such as LGB could possibly exploit this leak.",
    "1090536": "Is this working within a reasonable time?",
    "1150608": "Hi,\nThanks. I have used some of the insights here:\nhttps://www.kaggle.com/kritidoneria/beginner-wids21-feature-engineering-starter"
  },
  "source": "meta"
}