{
  "id": 209611,
  "title": "Wow!! Fun Comp",
  "url": "/competitions/riiid-test-answer-prediction/discussion/209611",
  "author_name": "عثمان",
  "post_date": "2021-01-08T02:26:28.783000",
  "votes": 11,
  "comment_count": 0,
  "views": 0,
  "content": "<ul>\n<li><a href=\"https://www.kaggle.com/authman/humble-fe-train/\" target=\"_blank\">Just show me the code</a></li>\n<li><a href=\"https://www.kaggle.com/authman/inference-attempt-incomplete/edit\" target=\"_blank\">Inference attempt</a></li>\n</ul>\n<p>3406 teams and an untold number of people competed here. How many won? How do you define a win? Perhaps it's time spent competing ÷ LB position :). Anyhow, I wanted to share some reflections I had about this competition. But first, I echo what <a href=\"https://www.kaggle.com/yanamal\" target=\"_blank\">@yanamal</a> <a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/209598\" target=\"_blank\">wrote here</a>, and really that's all that needs to be said there.</p>\n<p>For my implementation, here are the validation scores using the CV4 fold of the <a href=\"https://www.kaggle.com/marisakamozz/cv-strategy-in-the-kaggle-environment\" target=\"_blank\">unbiased</a> version <a href=\"https://www.kaggle.com/its7171\" target=\"_blank\">@its7171</a>'s <a href=\"https://www.kaggle.com/its7171/cv-strategy\" target=\"_blank\">dataset</a> by <a href=\"https://www.kaggle.com/marisakamozz\" target=\"_blank\">@marisakamozz</a>: 7824,7877,7901,7918,7929,794,7945,7955,7964,7965,7966,7973,7973,7977,7978,798,7982,7985,7986,7987,7991,7988,799,7998,7997,8003,8001,8001,8007,8007,8011,8007,8007,8007,8011,8008,8009,8009,8012. Takes about 2h to train on dual 2080 Ti.</p>\n<p>I'm still waiting on the write-ups from the grand winners, but here are the interesting tidbits I've gleamed in my experimentation.</p>\n<p><strong>My features</strong></p>\n<ol>\n<li>'task_container_id' Only used for masking</li>\n<li>Categorical:  'user_id', 'content_id', 'part_id', 'prior_question_elapsed_time', 'prior_question_had_explanation', 'incorrect_rank', 'content_type_id', 'bundle_id',</li>\n<li>Continuous Group 1: 'answer_ratio1', 'answer_ratio2', 'correct_streak_u', 'incorrect_streak_u', 'correct_streak_alltime_u', 'incorrect_streak_alltime_u',</li>\n<li>Session: 'session_content_num_u', 'session_duration_u', 'session_ans_duration_u', 'lifetime_ans_duration_u',</li>\n<li>Continuous Group 2: 'lag_ts_u_recency', 'last_ts_u_recency1', 'last_ts_u_recency2', 'last_ts_u_recency3', 'last_correct_ts_u_recency', 'last_incorrect_ts_u_recency',</li>\n<li>Correctness: 'correctness_u_recency', 'part_correctness_u_recency', 'session_correctness_u',</li>\n<li>Diagnostic: 'diagnostic_u_recency_1', 'diagnostic_u_recency_2', 'diagnostic_u_recency_3', 'diagnostic_u_recency_4', 'diagnostic_u_recency_5', 'diagnostic_u_recency_6',</li>\n<li>Boolean: 'encountered',</li>\n</ol>\n<h2>User_ID Embedding</h2>\n<p>Why don't people use user_id? It's a fair field to use when you know there is decent overlap between train + submission set. If you have a neural model, you can train it as per normal. Then, freeze all the embeddings and add a user_id embedding and just finetune that and everything upstream (connect it further in the pipeline). If you have a GBDT, set the user_id to nan for any unseen user. This seems like the easiest way to allow the model to perform better with minimal effort. You will have to adjust your validation setup such that the amount of 'unseen' users that you force map to the 0 catchall user embedding is roughly the amount you expect to encounter on the private lb.</p>\n<h2>Incorrect Rank</h2>\n<p>Every single solution I saw including the original saint/+ papers only look to see if the user answered correctly or not. Why? Is there not signal in looking at which incorrect answer the user selected? I ranked the incorrect answers from 1-3 as least likely incorrect to most likely incorrect, and 4 would be the actual correct. Training using this rather than simply binary correct/incorrect embedding as in the saint papers produced a visible boost.</p>\n<pre><code>question_incorrect_ranks = df[['content_id', 'user_answer', 'answered_correctly']]\nquestion_incorrect_ranks = question_incorrect_ranks[question_incorrect_ranks.answered_correctly == 0]\nquestion_incorrect_ranks = question_incorrect_ranks.groupby(['content_id','user_answer'], sort=False).count().reset_index()\nquestion_incorrect_ranks.columns = ['content_id', 'user_answer', 'incorrect_rank']\nquestion_incorrect_ranks.sort_values(['content_id', 'incorrect_rank'], inplace=True)\nquestion_incorrect_ranks['incorrect_rank'] = 3 + question_incorrect_ranks.groupby('content_id', sort=False).cumcount()\nquestion_incorrect_ranks.content_id = question_incorrect_ranks.content_id.astype(np.int32)\nquestion_incorrect_ranks.user_answer = question_incorrect_ranks.user_answer.astype(np.int8)\njoblib.dump(question_incorrect_ranks, f'./question_incorrect_ranks.pkl')\n</code></pre>\n<h2>Answer Ratio</h2>\n<p><a href=\"https://www.kaggle.com/temuujinerdene\" target=\"_blank\">@temuujinerdene</a> mentioned <a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/208244\" target=\"_blank\">on the forums</a> that what the user selects (1,2,3,4) has some signal. I did a check by looking at the mean answered_correctly value in the train set grouped by the user_answer VS grouped by the actual correct_answer from questions DF. Turns out, there is a discernibly measurable difference but only for options 1 and 2. 3 and 4 don't give signal. These ratios are the number of user_answer=1 QUESTIONS / cumlative questions seen to date for user, and same for user_answer = 2.</p>\n<h2>Correct and Incorrect</h2>\n<p>How many correct answers has the user answered to date consecutively. As with ALL features, be careful to bundle align these to prevent leakage. These features reset to 0 when the user answers the opposite of the feature, respective.</p>\n<h2>All time streaks</h2>\n<p>Same as before but storing the longest win or loss streak. This really helps against those annoying users that have 893483489 content_ids but always guess one answer, resulting in them averaging around 25% correctness.</p>\n<h2>Session features</h2>\n<p>The RIIID people have <a href=\"https://www.prnewswire.com/news-releases/riiids-ai-study-on-session-dropout-prediction-in-a-mobile-learning-environment-has-been-accepted-at-csedu-301011369.html\" target=\"_blank\">some papers</a> they published before saint/saint+ where they talk about sessions. They define session as a gap of 1 hour+ between interactions. Easy enough, we can create a session and track some variables across it. I track the cumcount of contents, the duration of it to date (bundle aligned), and the answer duration. Answer duration is defined as the cumsum of prior_question_elapsed_time. I also have an lifetime_ans_duration_u feature which is the cumsum to-date of the prior_question_elapsed_time variable. This is a good opportunity to mention that all time variables are log1p transformed and standard scaled before being fed into the net, with nans being set to 0.</p>\n<h2>Recency features</h2>\n<p>I have a <a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/205515\" target=\"_blank\">number of</a> posts <a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/207434\" target=\"_blank\">where I detail</a> the lag feature. last_ts_u_recency is just timestamp differences, and the last 2 of them. The last_correct_ts_u_recency and the incorrect variant are bundle aligned of course, and tell is how long its been since the last correct/incorrect answer up until this bundle (task container id)</p>\n<h2>Encountered</h2>\n<p>This was a <a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/206620\" target=\"_blank\">great feature discovered and talked about</a> on the forum by <a href=\"https://www.kaggle.com/rodolphelampe\" target=\"_blank\">@rodolphelampe</a>. I used my good friend <a href=\"https://www.kaggle.com/adityaecdrid\" target=\"_blank\">@adityaecdrid</a>'s bitarray implementation to track seen/not seen. This could effortless be extended with 2-3 copies of the bit array to track number of seen, capping it at a common sense level, and there is additional signal in that.</p>\n<h2>Correctness features</h2>\n<p>The correctness_u_recency feature is the bundle aligned cumulative ratio of correct to total content seen. I count all lectures as 'correct'. The part correctness are the same ration but aggregating only over the respective part. I don't have features for all parts, but I compute them separately and populate the column with the value for the part corresponding to the sample. I also do the same thing across the session. There In the thread in <a href=\"https://www.kaggle.com/gaozhanfire\" target=\"_blank\">@gaozhanfire</a>'s kernel that was <a href=\"https://www.kaggle.com/gaozhanfire/riiid-lgbm-val-0-788-offline-and-lsimodel-updated\" target=\"_blank\">mostly in Chinese</a>, they talked about last 5, 10, 20, etc correctness mean. Bundle aligning by session allows us to capture this feature at varying relevant timescales correctly and without leak.</p>\n<h2>Diagnostic features</h2>\n<p>We were told the first 30 or so questions were used to determine the users skill, so I figured these were extremely important. I grouped them by mean target into a few buckets. By default, I take the mean the the means of these questions. Whenever one of the relevant content_id's is encountered, I update the position with 0/1 based on if the user got it right or wrong. Then whatever number of items have their mean recalculated and that becomes the diagnostic feature. There might be 3 or 7 questions in one of these diagnostic features.</p>\n<p>I wasn't able to complete a submission pipeline due to prioritizing FE, huge mistake on my part. I guess I'll learn from the L and team up quickly next time so that others can work on the engineering and I can focus more on what I love, deep FE and playing with architectures. Oh, that reminds me.</p>\n<p>Only the correct/incorrect embedding needs to be 'lagged'. All other embeddings shouldn't be offset because we are able of having these features at inference time, so long as you bundle align. Now how do we mask the history (user answers)? We can't just use bundle alignment, because that will mask all the rest of those juicy features we just created for the whole bundle. I believe the task container id features of other questions are relevant, not just the question you're on. The only thing we want to hide is just this history. So I actually have two encoders and one decoder. The first encoder is for questions and has regular pad mask and triu mask. The decoder also have triu mask and pad mask. The 'history' decoder has task container mask so that we don't look at user answers within the task container.</p>\n<p>This was a fun competition. Tons of data, few features, non-anonymous, really allowed you to explore the data deeper as an analyst to bring out the sunshine. I tried using statistical and ML models to drive my FE pursuits. I'm sure the top LB'ers will share their tales. See you on the flip side!</p>",
  "messages": [
    {
      "id": 1143698,
      "postDate": "2021-01-08T02:26:28.783Z",
      "content": "<ul>\n<li><a href=\"https://www.kaggle.com/authman/humble-fe-train/\" target=\"_blank\">Just show me the code</a></li>\n<li><a href=\"https://www.kaggle.com/authman/inference-attempt-incomplete/edit\" target=\"_blank\">Inference attempt</a></li>\n</ul>\n<p>3406 teams and an untold number of people competed here. How many won? How do you define a win? Perhaps it's time spent competing ÷ LB position :). Anyhow, I wanted to share some reflections I had about this competition. But first, I echo what <a href=\"https://www.kaggle.com/yanamal\" target=\"_blank\">@yanamal</a> <a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/209598\" target=\"_blank\">wrote here</a>, and really that's all that needs to be said there.</p>\n<p>For my implementation, here are the validation scores using the CV4 fold of the <a href=\"https://www.kaggle.com/marisakamozz/cv-strategy-in-the-kaggle-environment\" target=\"_blank\">unbiased</a> version <a href=\"https://www.kaggle.com/its7171\" target=\"_blank\">@its7171</a>'s <a href=\"https://www.kaggle.com/its7171/cv-strategy\" target=\"_blank\">dataset</a> by <a href=\"https://www.kaggle.com/marisakamozz\" target=\"_blank\">@marisakamozz</a>: 7824,7877,7901,7918,7929,794,7945,7955,7964,7965,7966,7973,7973,7977,7978,798,7982,7985,7986,7987,7991,7988,799,7998,7997,8003,8001,8001,8007,8007,8011,8007,8007,8007,8011,8008,8009,8009,8012. Takes about 2h to train on dual 2080 Ti.</p>\n<p>I'm still waiting on the write-ups from the grand winners, but here are the interesting tidbits I've gleamed in my experimentation.</p>\n<p><strong>My features</strong></p>\n<ol>\n<li>'task_container_id' Only used for masking</li>\n<li>Categorical:  'user_id', 'content_id', 'part_id', 'prior_question_elapsed_time', 'prior_question_had_explanation', 'incorrect_rank', 'content_type_id', 'bundle_id',</li>\n<li>Continuous Group 1: 'answer_ratio1', 'answer_ratio2', 'correct_streak_u', 'incorrect_streak_u', 'correct_streak_alltime_u', 'incorrect_streak_alltime_u',</li>\n<li>Session: 'session_content_num_u', 'session_duration_u', 'session_ans_duration_u', 'lifetime_ans_duration_u',</li>\n<li>Continuous Group 2: 'lag_ts_u_recency', 'last_ts_u_recency1', 'last_ts_u_recency2', 'last_ts_u_recency3', 'last_correct_ts_u_recency', 'last_incorrect_ts_u_recency',</li>\n<li>Correctness: 'correctness_u_recency', 'part_correctness_u_recency', 'session_correctness_u',</li>\n<li>Diagnostic: 'diagnostic_u_recency_1', 'diagnostic_u_recency_2', 'diagnostic_u_recency_3', 'diagnostic_u_recency_4', 'diagnostic_u_recency_5', 'diagnostic_u_recency_6',</li>\n<li>Boolean: 'encountered',</li>\n</ol>\n<h2>User_ID Embedding</h2>\n<p>Why don't people use user_id? It's a fair field to use when you know there is decent overlap between train + submission set. If you have a neural model, you can train it as per normal. Then, freeze all the embeddings and add a user_id embedding and just finetune that and everything upstream (connect it further in the pipeline). If you have a GBDT, set the user_id to nan for any unseen user. This seems like the easiest way to allow the model to perform better with minimal effort. You will have to adjust your validation setup such that the amount of 'unseen' users that you force map to the 0 catchall user embedding is roughly the amount you expect to encounter on the private lb.</p>\n<h2>Incorrect Rank</h2>\n<p>Every single solution I saw including the original saint/+ papers only look to see if the user answered correctly or not. Why? Is there not signal in looking at which incorrect answer the user selected? I ranked the incorrect answers from 1-3 as least likely incorrect to most likely incorrect, and 4 would be the actual correct. Training using this rather than simply binary correct/incorrect embedding as in the saint papers produced a visible boost.</p>\n<pre><code>question_incorrect_ranks = df[['content_id', 'user_answer', 'answered_correctly']]\nquestion_incorrect_ranks = question_incorrect_ranks[question_incorrect_ranks.answered_correctly == 0]\nquestion_incorrect_ranks = question_incorrect_ranks.groupby(['content_id','user_answer'], sort=False).count().reset_index()\nquestion_incorrect_ranks.columns = ['content_id', 'user_answer', 'incorrect_rank']\nquestion_incorrect_ranks.sort_values(['content_id', 'incorrect_rank'], inplace=True)\nquestion_incorrect_ranks['incorrect_rank'] = 3 + question_incorrect_ranks.groupby('content_id', sort=False).cumcount()\nquestion_incorrect_ranks.content_id = question_incorrect_ranks.content_id.astype(np.int32)\nquestion_incorrect_ranks.user_answer = question_incorrect_ranks.user_answer.astype(np.int8)\njoblib.dump(question_incorrect_ranks, f'./question_incorrect_ranks.pkl')\n</code></pre>\n<h2>Answer Ratio</h2>\n<p><a href=\"https://www.kaggle.com/temuujinerdene\" target=\"_blank\">@temuujinerdene</a> mentioned <a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/208244\" target=\"_blank\">on the forums</a> that what the user selects (1,2,3,4) has some signal. I did a check by looking at the mean answered_correctly value in the train set grouped by the user_answer VS grouped by the actual correct_answer from questions DF. Turns out, there is a discernibly measurable difference but only for options 1 and 2. 3 and 4 don't give signal. These ratios are the number of user_answer=1 QUESTIONS / cumlative questions seen to date for user, and same for user_answer = 2.</p>\n<h2>Correct and Incorrect</h2>\n<p>How many correct answers has the user answered to date consecutively. As with ALL features, be careful to bundle align these to prevent leakage. These features reset to 0 when the user answers the opposite of the feature, respective.</p>\n<h2>All time streaks</h2>\n<p>Same as before but storing the longest win or loss streak. This really helps against those annoying users that have 893483489 content_ids but always guess one answer, resulting in them averaging around 25% correctness.</p>\n<h2>Session features</h2>\n<p>The RIIID people have <a href=\"https://www.prnewswire.com/news-releases/riiids-ai-study-on-session-dropout-prediction-in-a-mobile-learning-environment-has-been-accepted-at-csedu-301011369.html\" target=\"_blank\">some papers</a> they published before saint/saint+ where they talk about sessions. They define session as a gap of 1 hour+ between interactions. Easy enough, we can create a session and track some variables across it. I track the cumcount of contents, the duration of it to date (bundle aligned), and the answer duration. Answer duration is defined as the cumsum of prior_question_elapsed_time. I also have an lifetime_ans_duration_u feature which is the cumsum to-date of the prior_question_elapsed_time variable. This is a good opportunity to mention that all time variables are log1p transformed and standard scaled before being fed into the net, with nans being set to 0.</p>\n<h2>Recency features</h2>\n<p>I have a <a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/205515\" target=\"_blank\">number of</a> posts <a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/207434\" target=\"_blank\">where I detail</a> the lag feature. last_ts_u_recency is just timestamp differences, and the last 2 of them. The last_correct_ts_u_recency and the incorrect variant are bundle aligned of course, and tell is how long its been since the last correct/incorrect answer up until this bundle (task container id)</p>\n<h2>Encountered</h2>\n<p>This was a <a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/206620\" target=\"_blank\">great feature discovered and talked about</a> on the forum by <a href=\"https://www.kaggle.com/rodolphelampe\" target=\"_blank\">@rodolphelampe</a>. I used my good friend <a href=\"https://www.kaggle.com/adityaecdrid\" target=\"_blank\">@adityaecdrid</a>'s bitarray implementation to track seen/not seen. This could effortless be extended with 2-3 copies of the bit array to track number of seen, capping it at a common sense level, and there is additional signal in that.</p>\n<h2>Correctness features</h2>\n<p>The correctness_u_recency feature is the bundle aligned cumulative ratio of correct to total content seen. I count all lectures as 'correct'. The part correctness are the same ration but aggregating only over the respective part. I don't have features for all parts, but I compute them separately and populate the column with the value for the part corresponding to the sample. I also do the same thing across the session. There In the thread in <a href=\"https://www.kaggle.com/gaozhanfire\" target=\"_blank\">@gaozhanfire</a>'s kernel that was <a href=\"https://www.kaggle.com/gaozhanfire/riiid-lgbm-val-0-788-offline-and-lsimodel-updated\" target=\"_blank\">mostly in Chinese</a>, they talked about last 5, 10, 20, etc correctness mean. Bundle aligning by session allows us to capture this feature at varying relevant timescales correctly and without leak.</p>\n<h2>Diagnostic features</h2>\n<p>We were told the first 30 or so questions were used to determine the users skill, so I figured these were extremely important. I grouped them by mean target into a few buckets. By default, I take the mean the the means of these questions. Whenever one of the relevant content_id's is encountered, I update the position with 0/1 based on if the user got it right or wrong. Then whatever number of items have their mean recalculated and that becomes the diagnostic feature. There might be 3 or 7 questions in one of these diagnostic features.</p>\n<p>I wasn't able to complete a submission pipeline due to prioritizing FE, huge mistake on my part. I guess I'll learn from the L and team up quickly next time so that others can work on the engineering and I can focus more on what I love, deep FE and playing with architectures. Oh, that reminds me.</p>\n<p>Only the correct/incorrect embedding needs to be 'lagged'. All other embeddings shouldn't be offset because we are able of having these features at inference time, so long as you bundle align. Now how do we mask the history (user answers)? We can't just use bundle alignment, because that will mask all the rest of those juicy features we just created for the whole bundle. I believe the task container id features of other questions are relevant, not just the question you're on. The only thing we want to hide is just this history. So I actually have two encoders and one decoder. The first encoder is for questions and has regular pad mask and triu mask. The decoder also have triu mask and pad mask. The 'history' decoder has task container mask so that we don't look at user answers within the task container.</p>\n<p>This was a fun competition. Tons of data, few features, non-anonymous, really allowed you to explore the data deeper as an analyst to bring out the sunshine. I tried using statistical and ML models to drive my FE pursuits. I'm sure the top LB'ers will share their tales. See you on the flip side!</p>",
      "rawMarkdown": "- [Just show me the code](https://www.kaggle.com/authman/humble-fe-train/)\n- [Inference attempt](https://www.kaggle.com/authman/inference-attempt-incomplete/edit)\n\n3406 teams and an untold number of people competed here. How many won? How do you define a win? Perhaps it's time spent competing ÷ LB position :). Anyhow, I wanted to share some reflections I had about this competition. But first, I echo what @yanamal [wrote here](https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/209598), and really that's all that needs to be said there.\n\nFor my implementation, here are the validation scores using the CV4 fold of the [unbiased](https://www.kaggle.com/marisakamozz/cv-strategy-in-the-kaggle-environment) version @its7171's [dataset](https://www.kaggle.com/its7171/cv-strategy) by @marisakamozz: 7824,7877,7901,7918,7929,794,7945,7955,7964,7965,7966,7973,7973,7977,7978,798,7982,7985,7986,7987,7991,7988,799,7998,7997,8003,8001,8001,8007,8007,8011,8007,8007,8007,8011,8008,8009,8009,8012. Takes about 2h to train on dual 2080 Ti.\n\nI'm still waiting on the write-ups from the grand winners, but here are the interesting tidbits I've gleamed in my experimentation.\n\n**My features**\n1. 'task_container_id' Only used for masking\n1. Categorical:  'user_id', 'content_id', 'part_id', 'prior_question_elapsed_time', 'prior_question_had_explanation', 'incorrect_rank', 'content_type_id', 'bundle_id',\n1. Continuous Group 1: 'answer_ratio1', 'answer_ratio2', 'correct_streak_u', 'incorrect_streak_u', 'correct_streak_alltime_u', 'incorrect_streak_alltime_u',\n1. Session: 'session_content_num_u', 'session_duration_u', 'session_ans_duration_u', 'lifetime_ans_duration_u',\n1. Continuous Group 2: 'lag_ts_u_recency', 'last_ts_u_recency1', 'last_ts_u_recency2', 'last_ts_u_recency3', 'last_correct_ts_u_recency', 'last_incorrect_ts_u_recency',\n1. Correctness: 'correctness_u_recency', 'part_correctness_u_recency', 'session_correctness_u',\n1. Diagnostic: 'diagnostic_u_recency_1', 'diagnostic_u_recency_2', 'diagnostic_u_recency_3', 'diagnostic_u_recency_4', 'diagnostic_u_recency_5', 'diagnostic_u_recency_6',\n1. Boolean: 'encountered',\n\n\n## User_ID Embedding\nWhy don't people use user_id? It's a fair field to use when you know there is decent overlap between train + submission set. If you have a neural model, you can train it as per normal. Then, freeze all the embeddings and add a user_id embedding and just finetune that and everything upstream (connect it further in the pipeline). If you have a GBDT, set the user_id to nan for any unseen user. This seems like the easiest way to allow the model to perform better with minimal effort. You will have to adjust your validation setup such that the amount of 'unseen' users that you force map to the 0 catchall user embedding is roughly the amount you expect to encounter on the private lb.\n\n## Incorrect Rank\nEvery single solution I saw including the original saint/+ papers only look to see if the user answered correctly or not. Why? Is there not signal in looking at which incorrect answer the user selected? I ranked the incorrect answers from 1-3 as least likely incorrect to most likely incorrect, and 4 would be the actual correct. Training using this rather than simply binary correct/incorrect embedding as in the saint papers produced a visible boost.\n\n```\nquestion_incorrect_ranks = df[['content_id', 'user_answer', 'answered_correctly']]\nquestion_incorrect_ranks = question_incorrect_ranks[question_incorrect_ranks.answered_correctly == 0]\nquestion_incorrect_ranks = question_incorrect_ranks.groupby(['content_id','user_answer'], sort=False).count().reset_index()\nquestion_incorrect_ranks.columns = ['content_id', 'user_answer', 'incorrect_rank']\nquestion_incorrect_ranks.sort_values(['content_id', 'incorrect_rank'], inplace=True)\nquestion_incorrect_ranks['incorrect_rank'] = 3 + question_incorrect_ranks.groupby('content_id', sort=False).cumcount()\nquestion_incorrect_ranks.content_id = question_incorrect_ranks.content_id.astype(np.int32)\nquestion_incorrect_ranks.user_answer = question_incorrect_ranks.user_answer.astype(np.int8)\njoblib.dump(question_incorrect_ranks, f'./question_incorrect_ranks.pkl')\n```\n\n## Answer Ratio\n@temuujinerdene mentioned [on the forums](https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/208244) that what the user selects (1,2,3,4) has some signal. I did a check by looking at the mean answered_correctly value in the train set grouped by the user_answer VS grouped by the actual correct_answer from questions DF. Turns out, there is a discernibly measurable difference but only for options 1 and 2. 3 and 4 don't give signal. These ratios are the number of user_answer=1 QUESTIONS / cumlative questions seen to date for user, and same for user_answer = 2.\n\n## Correct and Incorrect\nHow many correct answers has the user answered to date consecutively. As with ALL features, be careful to bundle align these to prevent leakage. These features reset to 0 when the user answers the opposite of the feature, respective.\n\n## All time streaks\nSame as before but storing the longest win or loss streak. This really helps against those annoying users that have 893483489 content_ids but always guess one answer, resulting in them averaging around 25% correctness.\n\n## Session features\nThe RIIID people have [some papers](https://www.prnewswire.com/news-releases/riiids-ai-study-on-session-dropout-prediction-in-a-mobile-learning-environment-has-been-accepted-at-csedu-301011369.html) they published before saint/saint+ where they talk about sessions. They define session as a gap of 1 hour+ between interactions. Easy enough, we can create a session and track some variables across it. I track the cumcount of contents, the duration of it to date (bundle aligned), and the answer duration. Answer duration is defined as the cumsum of prior_question_elapsed_time. I also have an lifetime_ans_duration_u feature which is the cumsum to-date of the prior_question_elapsed_time variable. This is a good opportunity to mention that all time variables are log1p transformed and standard scaled before being fed into the net, with nans being set to 0.\n\n## Recency features\nI have a [number of](https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/205515) posts [where I detail](https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/207434) the lag feature. last_ts_u_recency is just timestamp differences, and the last 2 of them. The last_correct_ts_u_recency and the incorrect variant are bundle aligned of course, and tell is how long its been since the last correct/incorrect answer up until this bundle (task container id)\n\n## Encountered\nThis was a [great feature discovered and talked about](https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/206620) on the forum by @rodolphelampe. I used my good friend @adityaecdrid's bitarray implementation to track seen/not seen. This could effortless be extended with 2-3 copies of the bit array to track number of seen, capping it at a common sense level, and there is additional signal in that.\n\n##  Correctness features\nThe correctness_u_recency feature is the bundle aligned cumulative ratio of correct to total content seen. I count all lectures as 'correct'. The part correctness are the same ration but aggregating only over the respective part. I don't have features for all parts, but I compute them separately and populate the column with the value for the part corresponding to the sample. I also do the same thing across the session. There In the thread in @gaozhanfire's kernel that was [mostly in Chinese](https://www.kaggle.com/gaozhanfire/riiid-lgbm-val-0-788-offline-and-lsimodel-updated), they talked about last 5, 10, 20, etc correctness mean. Bundle aligning by session allows us to capture this feature at varying relevant timescales correctly and without leak.\n\n## Diagnostic features\nWe were told the first 30 or so questions were used to determine the users skill, so I figured these were extremely important. I grouped them by mean target into a few buckets. By default, I take the mean the the means of these questions. Whenever one of the relevant content_id's is encountered, I update the position with 0/1 based on if the user got it right or wrong. Then whatever number of items have their mean recalculated and that becomes the diagnostic feature. There might be 3 or 7 questions in one of these diagnostic features.\n\nI wasn't able to complete a submission pipeline due to prioritizing FE, huge mistake on my part. I guess I'll learn from the L and team up quickly next time so that others can work on the engineering and I can focus more on what I love, deep FE and playing with architectures. Oh, that reminds me.\n\nOnly the correct/incorrect embedding needs to be 'lagged'. All other embeddings shouldn't be offset because we are able of having these features at inference time, so long as you bundle align. Now how do we mask the history (user answers)? We can't just use bundle alignment, because that will mask all the rest of those juicy features we just created for the whole bundle. I believe the task container id features of other questions are relevant, not just the question you're on. The only thing we want to hide is just this history. So I actually have two encoders and one decoder. The first encoder is for questions and has regular pad mask and triu mask. The decoder also have triu mask and pad mask. The 'history' decoder has task container mask so that we don't look at user answers within the task container.\n\nThis was a fun competition. Tons of data, few features, non-anonymous, really allowed you to explore the data deeper as an analyst to bring out the sunshine. I tried using statistical and ML models to drive my FE pursuits. I'm sure the top LB'ers will share their tales. See you on the flip side!",
      "votes": 11
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "1143698": "- [Just show me the code](https://www.kaggle.com/authman/humble-fe-train/)\n- [Inference attempt](https://www.kaggle.com/authman/inference-attempt-incomplete/edit)\n\n3406 teams and an untold number of people competed here. How many won? How do you define a win? Perhaps it's time spent competing ÷ LB position :). Anyhow, I wanted to share some reflections I had about this competition. But first, I echo what @yanamal [wrote here](https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/209598), and really that's all that needs to be said there.\n\nFor my implementation, here are the validation scores using the CV4 fold of the [unbiased](https://www.kaggle.com/marisakamozz/cv-strategy-in-the-kaggle-environment) version @its7171's [dataset](https://www.kaggle.com/its7171/cv-strategy) by @marisakamozz: 7824,7877,7901,7918,7929,794,7945,7955,7964,7965,7966,7973,7973,7977,7978,798,7982,7985,7986,7987,7991,7988,799,7998,7997,8003,8001,8001,8007,8007,8011,8007,8007,8007,8011,8008,8009,8009,8012. Takes about 2h to train on dual 2080 Ti.\n\nI'm still waiting on the write-ups from the grand winners, but here are the interesting tidbits I've gleamed in my experimentation.\n\n**My features**\n1. 'task_container_id' Only used for masking\n1. Categorical:  'user_id', 'content_id', 'part_id', 'prior_question_elapsed_time', 'prior_question_had_explanation', 'incorrect_rank', 'content_type_id', 'bundle_id',\n1. Continuous Group 1: 'answer_ratio1', 'answer_ratio2', 'correct_streak_u', 'incorrect_streak_u', 'correct_streak_alltime_u', 'incorrect_streak_alltime_u',\n1. Session: 'session_content_num_u', 'session_duration_u', 'session_ans_duration_u', 'lifetime_ans_duration_u',\n1. Continuous Group 2: 'lag_ts_u_recency', 'last_ts_u_recency1', 'last_ts_u_recency2', 'last_ts_u_recency3', 'last_correct_ts_u_recency', 'last_incorrect_ts_u_recency',\n1. Correctness: 'correctness_u_recency', 'part_correctness_u_recency', 'session_correctness_u',\n1. Diagnostic: 'diagnostic_u_recency_1', 'diagnostic_u_recency_2', 'diagnostic_u_recency_3', 'diagnostic_u_recency_4', 'diagnostic_u_recency_5', 'diagnostic_u_recency_6',\n1. Boolean: 'encountered',\n\n\n## User_ID Embedding\nWhy don't people use user_id? It's a fair field to use when you know there is decent overlap between train + submission set. If you have a neural model, you can train it as per normal. Then, freeze all the embeddings and add a user_id embedding and just finetune that and everything upstream (connect it further in the pipeline). If you have a GBDT, set the user_id to nan for any unseen user. This seems like the easiest way to allow the model to perform better with minimal effort. You will have to adjust your validation setup such that the amount of 'unseen' users that you force map to the 0 catchall user embedding is roughly the amount you expect to encounter on the private lb.\n\n## Incorrect Rank\nEvery single solution I saw including the original saint/+ papers only look to see if the user answered correctly or not. Why? Is there not signal in looking at which incorrect answer the user selected? I ranked the incorrect answers from 1-3 as least likely incorrect to most likely incorrect, and 4 would be the actual correct. Training using this rather than simply binary correct/incorrect embedding as in the saint papers produced a visible boost.\n\n```\nquestion_incorrect_ranks = df[['content_id', 'user_answer', 'answered_correctly']]\nquestion_incorrect_ranks = question_incorrect_ranks[question_incorrect_ranks.answered_correctly == 0]\nquestion_incorrect_ranks = question_incorrect_ranks.groupby(['content_id','user_answer'], sort=False).count().reset_index()\nquestion_incorrect_ranks.columns = ['content_id', 'user_answer', 'incorrect_rank']\nquestion_incorrect_ranks.sort_values(['content_id', 'incorrect_rank'], inplace=True)\nquestion_incorrect_ranks['incorrect_rank'] = 3 + question_incorrect_ranks.groupby('content_id', sort=False).cumcount()\nquestion_incorrect_ranks.content_id = question_incorrect_ranks.content_id.astype(np.int32)\nquestion_incorrect_ranks.user_answer = question_incorrect_ranks.user_answer.astype(np.int8)\njoblib.dump(question_incorrect_ranks, f'./question_incorrect_ranks.pkl')\n```\n\n## Answer Ratio\n@temuujinerdene mentioned [on the forums](https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/208244) that what the user selects (1,2,3,4) has some signal. I did a check by looking at the mean answered_correctly value in the train set grouped by the user_answer VS grouped by the actual correct_answer from questions DF. Turns out, there is a discernibly measurable difference but only for options 1 and 2. 3 and 4 don't give signal. These ratios are the number of user_answer=1 QUESTIONS / cumlative questions seen to date for user, and same for user_answer = 2.\n\n## Correct and Incorrect\nHow many correct answers has the user answered to date consecutively. As with ALL features, be careful to bundle align these to prevent leakage. These features reset to 0 when the user answers the opposite of the feature, respective.\n\n## All time streaks\nSame as before but storing the longest win or loss streak. This really helps against those annoying users that have 893483489 content_ids but always guess one answer, resulting in them averaging around 25% correctness.\n\n## Session features\nThe RIIID people have [some papers](https://www.prnewswire.com/news-releases/riiids-ai-study-on-session-dropout-prediction-in-a-mobile-learning-environment-has-been-accepted-at-csedu-301011369.html) they published before saint/saint+ where they talk about sessions. They define session as a gap of 1 hour+ between interactions. Easy enough, we can create a session and track some variables across it. I track the cumcount of contents, the duration of it to date (bundle aligned), and the answer duration. Answer duration is defined as the cumsum of prior_question_elapsed_time. I also have an lifetime_ans_duration_u feature which is the cumsum to-date of the prior_question_elapsed_time variable. This is a good opportunity to mention that all time variables are log1p transformed and standard scaled before being fed into the net, with nans being set to 0.\n\n## Recency features\nI have a [number of](https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/205515) posts [where I detail](https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/207434) the lag feature. last_ts_u_recency is just timestamp differences, and the last 2 of them. The last_correct_ts_u_recency and the incorrect variant are bundle aligned of course, and tell is how long its been since the last correct/incorrect answer up until this bundle (task container id)\n\n## Encountered\nThis was a [great feature discovered and talked about](https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/206620) on the forum by @rodolphelampe. I used my good friend @adityaecdrid's bitarray implementation to track seen/not seen. This could effortless be extended with 2-3 copies of the bit array to track number of seen, capping it at a common sense level, and there is additional signal in that.\n\n##  Correctness features\nThe correctness_u_recency feature is the bundle aligned cumulative ratio of correct to total content seen. I count all lectures as 'correct'. The part correctness are the same ration but aggregating only over the respective part. I don't have features for all parts, but I compute them separately and populate the column with the value for the part corresponding to the sample. I also do the same thing across the session. There In the thread in @gaozhanfire's kernel that was [mostly in Chinese](https://www.kaggle.com/gaozhanfire/riiid-lgbm-val-0-788-offline-and-lsimodel-updated), they talked about last 5, 10, 20, etc correctness mean. Bundle aligning by session allows us to capture this feature at varying relevant timescales correctly and without leak.\n\n## Diagnostic features\nWe were told the first 30 or so questions were used to determine the users skill, so I figured these were extremely important. I grouped them by mean target into a few buckets. By default, I take the mean the the means of these questions. Whenever one of the relevant content_id's is encountered, I update the position with 0/1 based on if the user got it right or wrong. Then whatever number of items have their mean recalculated and that becomes the diagnostic feature. There might be 3 or 7 questions in one of these diagnostic features.\n\nI wasn't able to complete a submission pipeline due to prioritizing FE, huge mistake on my part. I guess I'll learn from the L and team up quickly next time so that others can work on the engineering and I can focus more on what I love, deep FE and playing with architectures. Oh, that reminds me.\n\nOnly the correct/incorrect embedding needs to be 'lagged'. All other embeddings shouldn't be offset because we are able of having these features at inference time, so long as you bundle align. Now how do we mask the history (user answers)? We can't just use bundle alignment, because that will mask all the rest of those juicy features we just created for the whole bundle. I believe the task container id features of other questions are relevant, not just the question you're on. The only thing we want to hide is just this history. So I actually have two encoders and one decoder. The first encoder is for questions and has regular pad mask and triu mask. The decoder also have triu mask and pad mask. The 'history' decoder has task container mask so that we don't look at user answers within the task container.\n\nThis was a fun competition. Tons of data, few features, non-anonymous, really allowed you to explore the data deeper as an analyst to bring out the sunshine. I tried using statistical and ML models to drive my FE pursuits. I'm sure the top LB'ers will share their tales. See you on the flip side!"
  }
}