{
  "id": 209628,
  "title": "[LB 125th Position - .795] Pure Feature Engineering .(3M pkl files) Comp & Solution",
  "url": "/competitions/riiid-test-answer-prediction/writeups/we-aren-t-saints-lb-125th-position-795-pure-featur",
  "author_name": "",
  "post_date": "2021-01-08T05:29:03.757Z",
  "votes": 26,
  "comment_count": 14,
  "views": 0,
  "content": "<p>First of all, Thanks a lot Kaggle, Riiid and the everyone who participated! I have learnt a lot from you guys in just 3 months and really out of the box solution attempts and FEs! </p>\n<p>And special thanks to my teammates! I have learnt a lot about different topics and FE ideas as well which i definitely couldn't come up on my own.</p>\n<p>So coming to our solutions, it's a very simple ensemble of gbm models (~39/41 Features) and a NN (AKT). CV-LB was amazingly in sync, so no surprise there.</p>\n<p>Features List -:</p>\n<pre><code>['user_resid_rolling',\n 'sum_answered_correctly_user',\n 'count_user',\n 'content_attempt',\n 'lag_1',\n 'lag_2',\n 'lag_3',\n 'lag_4',\n 'lag_5',\n 'lag_sum_5',\n 'part',\n 'tags1',\n 'prior_question_had_explanation_enc',\n 'count_content',\n 'sum_answered_correctly_content',\n 'mean_answered_correctly_content',\n 'prior_question_elapsed_time',\n 'mean_explanation_content',\n 'mean_elapsed_content',\n 'theta_user_ct', # from elo\n 'theta_user_pt', # from elo\n 'theta_user_bd', # from elo\n 'content_id',\n 'last_lec_ts',\n 'diff_to_last_lec_ts',\n 'num_lects_watched',\n 'answer_1',\n 'answer_2',\n 'answer_3',\n 'answer_4',\n 'answer_5',\n 'answer_6',\n 'answer_7',\n 'last_7_answers_sum',\n 'correct_answer',\n 'n_tasks_performed',\n 'timestamp',\n 'prev_answer_content',\n 'diff_content_time',\n 'mean_answered_correctly_user',\n 'mean_user_resid_rolling']\n</code></pre>\n<p>Yes, we did have lecture features like  <code>'last_lec_ts',  'diff_to_last_lec_ts',  'num_lects_watched',</code> and they definitely help us improve both, our CV and LB. Other interesting features that we had were not limited but deserve a mention \"user_resid_rolling\", \"diff_content_time\", \"multiple elo features\", \"last_7_answers_sum\".</p>\n<p>\"user_resid_rolling\" -&gt; </p>\n<pre><code>train['content_mean'] = train.groupby('content_id')['answered_correctly'].transform('mean')\ntrain['resid'] = train['answered_correctly'] - train['content_mean']\ntrain['resid'] = train.groupby('user_id')['resid'].shift()\ntrain['user_resid_rolling'] = train.groupby('user_id')['resid'].agg(['cumsum'])\ntrain['user_resid_rolling'] = train['user_resid_rolling'].fillna(0)\ntrain.drop(['resid', 'content_mean'], axis=1, inplace=True)\n</code></pre>\n<p>\"last_lec_ts\" &amp; \"diff_to_last_lec_ts\" -&gt; </p>\n<pre><code>prev_u, lec_ts = -1, 0\nlast_lec_ts = []\n\nfor u, c_type, t in zip(train['user_id'], train['content_type_id'], train['timestamp']):\n    if u != prev_u:\n        lec_ts = 0\n    if c_type == True:\n        lec_ts = t\n    prev_u = u\n    last_lec_ts.append(lec_ts)\n\ntrain['last_lec_ts'] = np.array(last_lec_ts)\ntrain['diff_to_last_lec_ts'] = train['timestamp'] - train['last_lec_ts']\n</code></pre>\n<p>And the craziest of all FE and a  cool FE, <strong>We saved ~.3M pkl files during runtime for the FE \"diff_content_time\",</strong> (do delete the dir after your inference is complete) (basically for every user difference b/w the TS b/w their current and last TS for a particular content_id)</p>\n<pre><code>!mkdir pkld\ndct = {}\ndct[-1] = 0\nprev_u = -1\nres = np.zeros(3)\nfor u, cid, ctype, ts in tqdm(zip(train['user_id'], train['content_id'], train['content_type_id'], train['timestamp'])):\n    if prev_u != u:\n        pickle.dump(dct[prev_u], open('pkld/' + str(prev_u) + '.pkl', 'wb'), protocol=pickle.HIGHEST_PROTOCOL)\n        del dct[prev_u]\n        dct[u] = {}\n    if ctype == False:\n        dct[u][cid] = ts\n    prev_u = u\npickle.dump(dct[prev_u], open('pkld/' + str(prev_u) + '.pkl', 'wb'), protocol=pickle.HIGHEST_PROTOCOL) #########\n\ndel dct\n</code></pre>\n<p>There were a lot of FE ideas that were tried but didn't help us. To mention few, \"tf-idf\" on content_ids history of the user, rolling_windows using numba, FTRL's, and other hell lot of FE's ideas that gave marginal boost but we had a rule to add +.001 single-single features in place.</p>\n<p>What we regret was definitely not having a SAINT like arch for sure. I messed up badly, should have simply used nn.Transformer rather than handling the encoder/decoder parts myself.</p>\n<p>But until next time,</p>\n<p>Happy Fair Kaggling &amp; Keep Sharing A Ton!</p>\n<p>cc <a href=\"https://www.kaggle.com/alexj21\" target=\"_blank\">@alexj21</a> <a href=\"https://www.kaggle.com/nikhilmishradev\" target=\"_blank\">@nikhilmishradev</a> <a href=\"https://www.kaggle.com/rohanrao\" target=\"_blank\">@rohanrao</a>  [Alphabetical Order]. (Please add anything that i missed)<br>\nThanks a lot for the collaboration guys!</p>",
  "messages": [
    {
      "id": "1143784",
      "postDate": "01/08/2021 04:14:49",
      "content": "<p>First of all, Thanks a lot Kaggle, Riiid and the everyone who participated! I have learnt a lot from you guys in just 3 months and really out of the box solution attempts and FEs! </p>\n<p>And special thanks to my teammates! I have learnt a lot about different topics and FE ideas as well which i definitely couldn't come up on my own.</p>\n<p>So coming to our solutions, it's a very simple ensemble of gbm models (~39/41 Features) and a NN (AKT). CV-LB was amazingly in sync, so no surprise there.</p>\n<p>Features List -:</p>\n<pre><code>['user_resid_rolling',\n 'sum_answered_correctly_user',\n 'count_user',\n 'content_attempt',\n 'lag_1',\n 'lag_2',\n 'lag_3',\n 'lag_4',\n 'lag_5',\n 'lag_sum_5',\n 'part',\n 'tags1',\n 'prior_question_had_explanation_enc',\n 'count_content',\n 'sum_answered_correctly_content',\n 'mean_answered_correctly_content',\n 'prior_question_elapsed_time',\n 'mean_explanation_content',\n 'mean_elapsed_content',\n 'theta_user_ct', # from elo\n 'theta_user_pt', # from elo\n 'theta_user_bd', # from elo\n 'content_id',\n 'last_lec_ts',\n 'diff_to_last_lec_ts',\n 'num_lects_watched',\n 'answer_1',\n 'answer_2',\n 'answer_3',\n 'answer_4',\n 'answer_5',\n 'answer_6',\n 'answer_7',\n 'last_7_answers_sum',\n 'correct_answer',\n 'n_tasks_performed',\n 'timestamp',\n 'prev_answer_content',\n 'diff_content_time',\n 'mean_answered_correctly_user',\n 'mean_user_resid_rolling']\n</code></pre>\n<p>Yes, we did have lecture features like  <code>'last_lec_ts',  'diff_to_last_lec_ts',  'num_lects_watched',</code> and they definitely help us improve both, our CV and LB. Other interesting features that we had were not limited but deserve a mention \"user_resid_rolling\", \"diff_content_time\", \"multiple elo features\", \"last_7_answers_sum\".</p>\n<p>\"user_resid_rolling\" -&gt; </p>\n<pre><code>train['content_mean'] = train.groupby('content_id')['answered_correctly'].transform('mean')\ntrain['resid'] = train['answered_correctly'] - train['content_mean']\ntrain['resid'] = train.groupby('user_id')['resid'].shift()\ntrain['user_resid_rolling'] = train.groupby('user_id')['resid'].agg(['cumsum'])\ntrain['user_resid_rolling'] = train['user_resid_rolling'].fillna(0)\ntrain.drop(['resid', 'content_mean'], axis=1, inplace=True)\n</code></pre>\n<p>\"last_lec_ts\" &amp; \"diff_to_last_lec_ts\" -&gt; </p>\n<pre><code>prev_u, lec_ts = -1, 0\nlast_lec_ts = []\n\nfor u, c_type, t in zip(train['user_id'], train['content_type_id'], train['timestamp']):\n    if u != prev_u:\n        lec_ts = 0\n    if c_type == True:\n        lec_ts = t\n    prev_u = u\n    last_lec_ts.append(lec_ts)\n\ntrain['last_lec_ts'] = np.array(last_lec_ts)\ntrain['diff_to_last_lec_ts'] = train['timestamp'] - train['last_lec_ts']\n</code></pre>\n<p>And the craziest of all FE and a  cool FE, <strong>We saved ~.3M pkl files during runtime for the FE \"diff_content_time\",</strong> (do delete the dir after your inference is complete) (basically for every user difference b/w the TS b/w their current and last TS for a particular content_id)</p>\n<pre><code>!mkdir pkld\ndct = {}\ndct[-1] = 0\nprev_u = -1\nres = np.zeros(3)\nfor u, cid, ctype, ts in tqdm(zip(train['user_id'], train['content_id'], train['content_type_id'], train['timestamp'])):\n    if prev_u != u:\n        pickle.dump(dct[prev_u], open('pkld/' + str(prev_u) + '.pkl', 'wb'), protocol=pickle.HIGHEST_PROTOCOL)\n        del dct[prev_u]\n        dct[u] = {}\n    if ctype == False:\n        dct[u][cid] = ts\n    prev_u = u\npickle.dump(dct[prev_u], open('pkld/' + str(prev_u) + '.pkl', 'wb'), protocol=pickle.HIGHEST_PROTOCOL) #########\n\ndel dct\n</code></pre>\n<p>There were a lot of FE ideas that were tried but didn't help us. To mention few, \"tf-idf\" on content_ids history of the user, rolling_windows using numba, FTRL's, and other hell lot of FE's ideas that gave marginal boost but we had a rule to add +.001 single-single features in place.</p>\n<p>What we regret was definitely not having a SAINT like arch for sure. I messed up badly, should have simply used nn.Transformer rather than handling the encoder/decoder parts myself.</p>\n<p>But until next time,</p>\n<p>Happy Fair Kaggling &amp; Keep Sharing A Ton!</p>\n<p>cc <a href=\"https://www.kaggle.com/alexj21\" target=\"_blank\">@alexj21</a> <a href=\"https://www.kaggle.com/nikhilmishradev\" target=\"_blank\">@nikhilmishradev</a> <a href=\"https://www.kaggle.com/rohanrao\" target=\"_blank\">@rohanrao</a>  [Alphabetical Order]. (Please add anything that i missed)<br>\nThanks a lot for the collaboration guys!</p>",
      "rawMarkdown": "First of all, Thanks a lot Kaggle, Riiid and the everyone who participated! I have learnt a lot from you guys in just 3 months and really out of the box solution attempts and FEs! \n\nAnd special thanks to my teammates! I have learnt a lot about different topics and FE ideas as well which i definitely couldn't come up on my own.\n\nSo coming to our solutions, it's a very simple ensemble of gbm models (~39/41 Features) and a NN (AKT). CV-LB was amazingly in sync, so no surprise there.\n\nFeatures List -:\n\n```\n['user_resid_rolling',\n 'sum_answered_correctly_user',\n 'count_user',\n 'content_attempt',\n 'lag_1',\n 'lag_2',\n 'lag_3',\n 'lag_4',\n 'lag_5',\n 'lag_sum_5',\n 'part',\n 'tags1',\n 'prior_question_had_explanation_enc',\n 'count_content',\n 'sum_answered_correctly_content',\n 'mean_answered_correctly_content',\n 'prior_question_elapsed_time',\n 'mean_explanation_content',\n 'mean_elapsed_content',\n 'theta_user_ct', # from elo\n 'theta_user_pt', # from elo\n 'theta_user_bd', # from elo\n 'content_id',\n 'last_lec_ts',\n 'diff_to_last_lec_ts',\n 'num_lects_watched',\n 'answer_1',\n 'answer_2',\n 'answer_3',\n 'answer_4',\n 'answer_5',\n 'answer_6',\n 'answer_7',\n 'last_7_answers_sum',\n 'correct_answer',\n 'n_tasks_performed',\n 'timestamp',\n 'prev_answer_content',\n 'diff_content_time',\n 'mean_answered_correctly_user',\n 'mean_user_resid_rolling']\n```\n\n\nYes, we did have lecture features like  `'last_lec_ts',  'diff_to_last_lec_ts',  'num_lects_watched',` and they definitely help us improve both, our CV and LB. Other interesting features that we had were not limited but deserve a mention \"user_resid_rolling\", \"diff_content_time\", \"multiple elo features\", \"last_7_answers_sum\".\n\n\n\"user_resid_rolling\" -> \n\n```\ntrain['content_mean'] = train.groupby('content_id')['answered_correctly'].transform('mean')\ntrain['resid'] = train['answered_correctly'] - train['content_mean']\ntrain['resid'] = train.groupby('user_id')['resid'].shift()\ntrain['user_resid_rolling'] = train.groupby('user_id')['resid'].agg(['cumsum'])\ntrain['user_resid_rolling'] = train['user_resid_rolling'].fillna(0)\ntrain.drop(['resid', 'content_mean'], axis=1, inplace=True)\n```\n\n\"last_lec_ts\" & \"diff_to_last_lec_ts\" -> \n\n```\nprev_u, lec_ts = -1, 0\nlast_lec_ts = []\n\nfor u, c_type, t in zip(train['user_id'], train['content_type_id'], train['timestamp']):\n    if u != prev_u:\n        lec_ts = 0\n    if c_type == True:\n        lec_ts = t\n    prev_u = u\n    last_lec_ts.append(lec_ts)\n\ntrain['last_lec_ts'] = np.array(last_lec_ts)\ntrain['diff_to_last_lec_ts'] = train['timestamp'] - train['last_lec_ts']\n```\n\n\nAnd the craziest of all FE and a  cool FE, **We saved ~.3M pkl files during runtime for the FE \"diff_content_time\",** (do delete the dir after your inference is complete) (basically for every user difference b/w the TS b/w their current and last TS for a particular content_id)\n\n```\n!mkdir pkld\ndct = {}\ndct[-1] = 0\nprev_u = -1\nres = np.zeros(3)\nfor u, cid, ctype, ts in tqdm(zip(train['user_id'], train['content_id'], train['content_type_id'], train['timestamp'])):\n    if prev_u != u:\n        pickle.dump(dct[prev_u], open('pkld/' + str(prev_u) + '.pkl', 'wb'), protocol=pickle.HIGHEST_PROTOCOL)\n        del dct[prev_u]\n        dct[u] = {}\n    if ctype == False:\n        dct[u][cid] = ts\n    prev_u = u\npickle.dump(dct[prev_u], open('pkld/' + str(prev_u) + '.pkl', 'wb'), protocol=pickle.HIGHEST_PROTOCOL) #########\n\ndel dct\n```\n\nThere were a lot of FE ideas that were tried but didn't help us. To mention few, \"tf-idf\" on content_ids history of the user, rolling_windows using numba, FTRL's, and other hell lot of FE's ideas that gave marginal boost but we had a rule to add +.001 single-single features in place.\n\nWhat we regret was definitely not having a SAINT like arch for sure. I messed up badly, should have simply used nn.Transformer rather than handling the encoder/decoder parts myself.\n\n\nBut until next time,\n\nHappy Fair Kaggling & Keep Sharing A Ton!\n\ncc @alexj21 @nikhilmishradev @rohanrao  [Alphabetical Order]. (Please add anything that i missed)\nThanks a lot for the collaboration guys!",
      "votes": null
    },
    {
      "id": "1143803",
      "postDate": "01/08/2021 04:29:19",
      "content": "<p>Congratulations Aditya and team. Great feature engineering!</p>",
      "rawMarkdown": "Congratulations Aditya and team. Great feature engineering!",
      "votes": null
    },
    {
      "id": "1143807",
      "postDate": "01/08/2021 04:33:06",
      "content": "<p>Thanks a lot Chris! Great Work for yet another Gold! Any chances you are going to share your team's solution's as well? Ty!</p>",
      "rawMarkdown": "Thanks a lot Chris! Great Work for yet another Gold! Any chances you are going to share your team's solution's as well? Ty!",
      "votes": null
    },
    {
      "id": "1143815",
      "postDate": "01/08/2021 04:36:10",
      "content": "<p>Thanks. We'll share tomorrow. It's a single SAINT+ with 5 features. CV 812 LB 812.</p>",
      "rawMarkdown": "Thanks. We'll share tomorrow. It's a single SAINT+ with 5 features. CV 812 LB 812.",
      "votes": null
    },
    {
      "id": "1143861",
      "postDate": "01/08/2021 05:28:41",
      "content": "<p>Congrats on results <a href=\"https://www.kaggle.com/adityaecdrid\" target=\"_blank\">@adityaecdrid</a> and team</p>",
      "rawMarkdown": "Congrats on results @adityaecdrid and team",
      "votes": null
    },
    {
      "id": "1143864",
      "postDate": "01/08/2021 05:32:14",
      "content": "<p>Thanks a lot!</p>",
      "rawMarkdown": "Thanks a lot!",
      "votes": null
    },
    {
      "id": "1144070",
      "postDate": "01/08/2021 08:10:58",
      "content": "<p>Thanks for sharing especially about your FE!</p>",
      "rawMarkdown": "Thanks for sharing especially about your FE!",
      "votes": null
    },
    {
      "id": "1144130",
      "postDate": "01/08/2021 08:51:16",
      "content": "<p><a href=\"https://www.kaggle.com/adityaecdrid\" target=\"_blank\">@adityaecdrid</a>  Congratulations Aditya ! And thanks for the active discussions!</p>\n<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> Only 5 features …? OMG, that's great - yet It's killing me, as I added a lot of features to improve my SAINT + like model, and yet introduced bugs in the inference time … (CV improved, but LB worse …)</p>",
      "rawMarkdown": "adityaecdrid  Congratulations Aditya ! And thanks for the active discussions!\n\n@cdeotte Only 5 features ...? OMG, that's great - yet It's killing me, as I added a lot of features to improve my SAINT + like model, and yet introduced bugs in the inference time ... (CV improved, but LB worse ...)",
      "votes": null
    },
    {
      "id": "1144391",
      "postDate": "01/08/2021 12:41:08",
      "content": "<p>Congratz for the nice finish ! </p>\n<p>Only one small step before competition master <a href=\"https://www.kaggle.com/adityaecdrid\" target=\"_blank\">@adityaecdrid</a>  :) </p>",
      "rawMarkdown": "Congratz for the nice finish ! \n\nOnly one small step before competition master @adityaecdrid  :)",
      "votes": null
    },
    {
      "id": "1144394",
      "postDate": "01/08/2021 12:42:54",
      "content": "<p>Congratulations to <a href=\"https://www.kaggle.com/adityaecdrid\" target=\"_blank\">@adityaecdrid</a> and team! You were probably the person who engaged with us all the most for this competition :P</p>",
      "rawMarkdown": "Congratulations to @adityaecdrid and team! You were probably the person who engaged with us all the most for this competition :P",
      "votes": null
    },
    {
      "id": "1144471",
      "postDate": "01/08/2021 13:32:47",
      "content": "<p>Thanks a lot for the honour but i would like to share the same with many others as well competing, both after the comp ends and during the comp! It's because of these discussions, we all made it possible!</p>",
      "rawMarkdown": "Thanks a lot for the honour but i would like to share the same with many others as well competing, both after the comp ends and during the comp! It's because of these discussions, we all made it possible!",
      "votes": null
    },
    {
      "id": "1144474",
      "postDate": "01/08/2021 13:33:44",
      "content": "<p>Thanks a lot Theo! Yep, just one more silver/gold medal and finally, I can proudly say, I am a Comp Master!</p>",
      "rawMarkdown": "Thanks a lot Theo! Yep, just one more silver/gold medal and finally, I can proudly say, I am a Comp Master!",
      "votes": null
    },
    {
      "id": "1144476",
      "postDate": "01/08/2021 13:34:31",
      "content": "<p>Thanks a lot! i saw you joined early on but i guess you were busy with other things in // !</p>",
      "rawMarkdown": "Thanks a lot! i saw you joined early on but i guess you were busy with other things in // !",
      "votes": null
    },
    {
      "id": "1147494",
      "postDate": "01/10/2021 14:44:53",
      "content": "<p>So bright method! I am ashamed that I dont know FE. Can you explain more about it or which thread talk about it?</p>",
      "rawMarkdown": "So bright method! I am ashamed that I dont know FE. Can you explain more about it or which thread talk about it?",
      "votes": null
    },
    {
      "id": "1148930",
      "postDate": "01/11/2021 13:35:29",
      "content": "<p>Great work gays, congratulations !</p>\n<ul>\n<li>By curiosity, are <code>answer_1 ... answer_7</code> the selected answer (0,1,2,3) or whether the user was right/wrong ? I'm surprised that such features don't lead to overfitting</li>\n<li>How does your store-to-pkl trick impact the inference time ?</li>\n</ul>\n<p>Cheers</p>",
      "rawMarkdown": "Great work gays, congratulations !\n\n- By curiosity, are `answer_1 ... answer_7` the selected answer (0,1,2,3) or whether the user was right/wrong ? I'm surprised that such features don't lead to overfitting\n- How does your store-to-pkl trick impact the inference time ?\n\nCheers",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1143803,
      "author_name": "cdeotte",
      "author_url": "",
      "post_date": "01/08/2021 04:29:19",
      "content": "<p>Congratulations Aditya and team. Great feature engineering!</p>",
      "votes": null,
      "replies": [
        {
          "id": 1143807,
          "author_name": "adityaecdrid",
          "author_url": "",
          "post_date": "01/08/2021 04:33:06",
          "content": "<p>Thanks a lot Chris! Great Work for yet another Gold! Any chances you are going to share your team's solution's as well? Ty!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1143815,
          "author_name": "cdeotte",
          "author_url": "",
          "post_date": "01/08/2021 04:36:10",
          "content": "<p>Thanks. We'll share tomorrow. It's a single SAINT+ with 5 features. CV 812 LB 812.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1144130,
          "author_name": "yihdarshieh",
          "author_url": "",
          "post_date": "01/08/2021 08:51:16",
          "content": "<p><a href=\"https://www.kaggle.com/adityaecdrid\" target=\"_blank\">@adityaecdrid</a>  Congratulations Aditya ! And thanks for the active discussions!</p>\n<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> Only 5 features …? OMG, that's great - yet It's killing me, as I added a lot of features to improve my SAINT + like model, and yet introduced bugs in the inference time … (CV improved, but LB worse …)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1143861,
      "author_name": "duykhanh99",
      "author_url": "",
      "post_date": "01/08/2021 05:28:41",
      "content": "<p>Congrats on results <a href=\"https://www.kaggle.com/adityaecdrid\" target=\"_blank\">@adityaecdrid</a> and team</p>",
      "votes": null,
      "replies": [
        {
          "id": 1143864,
          "author_name": "adityaecdrid",
          "author_url": "",
          "post_date": "01/08/2021 05:32:14",
          "content": "<p>Thanks a lot!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1144070,
      "author_name": "sishihara",
      "author_url": "",
      "post_date": "01/08/2021 08:10:58",
      "content": "<p>Thanks for sharing especially about your FE!</p>",
      "votes": null,
      "replies": [
        {
          "id": 1144476,
          "author_name": "adityaecdrid",
          "author_url": "",
          "post_date": "01/08/2021 13:34:31",
          "content": "<p>Thanks a lot! i saw you joined early on but i guess you were busy with other things in // !</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1144391,
      "author_name": "theoviel",
      "author_url": "",
      "post_date": "01/08/2021 12:41:08",
      "content": "<p>Congratz for the nice finish ! </p>\n<p>Only one small step before competition master <a href=\"https://www.kaggle.com/adityaecdrid\" target=\"_blank\">@adityaecdrid</a>  :) </p>",
      "votes": null,
      "replies": [
        {
          "id": 1144474,
          "author_name": "adityaecdrid",
          "author_url": "",
          "post_date": "01/08/2021 13:33:44",
          "content": "<p>Thanks a lot Theo! Yep, just one more silver/gold medal and finally, I can proudly say, I am a Comp Master!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1144394,
      "author_name": "doctorkael",
      "author_url": "",
      "post_date": "01/08/2021 12:42:54",
      "content": "<p>Congratulations to <a href=\"https://www.kaggle.com/adityaecdrid\" target=\"_blank\">@adityaecdrid</a> and team! You were probably the person who engaged with us all the most for this competition :P</p>",
      "votes": null,
      "replies": [
        {
          "id": 1144471,
          "author_name": "adityaecdrid",
          "author_url": "",
          "post_date": "01/08/2021 13:32:47",
          "content": "<p>Thanks a lot for the honour but i would like to share the same with many others as well competing, both after the comp ends and during the comp! It's because of these discussions, we all made it possible!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1147494,
      "author_name": "zjjszj2",
      "author_url": "",
      "post_date": "01/10/2021 14:44:53",
      "content": "<p>So bright method! I am ashamed that I dont know FE. Can you explain more about it or which thread talk about it?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1148930,
      "author_name": "johannhuber",
      "author_url": "",
      "post_date": "01/11/2021 13:35:29",
      "content": "<p>Great work gays, congratulations !</p>\n<ul>\n<li>By curiosity, are <code>answer_1 ... answer_7</code> the selected answer (0,1,2,3) or whether the user was right/wrong ? I'm surprised that such features don't lead to overfitting</li>\n<li>How does your store-to-pkl trick impact the inference time ?</li>\n</ul>\n<p>Cheers</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1143784": "First of all, Thanks a lot Kaggle, Riiid and the everyone who participated! I have learnt a lot from you guys in just 3 months and really out of the box solution attempts and FEs! \n\nAnd special thanks to my teammates! I have learnt a lot about different topics and FE ideas as well which i definitely couldn't come up on my own.\n\nSo coming to our solutions, it's a very simple ensemble of gbm models (~39/41 Features) and a NN (AKT). CV-LB was amazingly in sync, so no surprise there.\n\nFeatures List -:\n\n```\n['user_resid_rolling',\n 'sum_answered_correctly_user',\n 'count_user',\n 'content_attempt',\n 'lag_1',\n 'lag_2',\n 'lag_3',\n 'lag_4',\n 'lag_5',\n 'lag_sum_5',\n 'part',\n 'tags1',\n 'prior_question_had_explanation_enc',\n 'count_content',\n 'sum_answered_correctly_content',\n 'mean_answered_correctly_content',\n 'prior_question_elapsed_time',\n 'mean_explanation_content',\n 'mean_elapsed_content',\n 'theta_user_ct', # from elo\n 'theta_user_pt', # from elo\n 'theta_user_bd', # from elo\n 'content_id',\n 'last_lec_ts',\n 'diff_to_last_lec_ts',\n 'num_lects_watched',\n 'answer_1',\n 'answer_2',\n 'answer_3',\n 'answer_4',\n 'answer_5',\n 'answer_6',\n 'answer_7',\n 'last_7_answers_sum',\n 'correct_answer',\n 'n_tasks_performed',\n 'timestamp',\n 'prev_answer_content',\n 'diff_content_time',\n 'mean_answered_correctly_user',\n 'mean_user_resid_rolling']\n```\n\n\nYes, we did have lecture features like  `'last_lec_ts',  'diff_to_last_lec_ts',  'num_lects_watched',` and they definitely help us improve both, our CV and LB. Other interesting features that we had were not limited but deserve a mention \"user_resid_rolling\", \"diff_content_time\", \"multiple elo features\", \"last_7_answers_sum\".\n\n\n\"user_resid_rolling\" -> \n\n```\ntrain['content_mean'] = train.groupby('content_id')['answered_correctly'].transform('mean')\ntrain['resid'] = train['answered_correctly'] - train['content_mean']\ntrain['resid'] = train.groupby('user_id')['resid'].shift()\ntrain['user_resid_rolling'] = train.groupby('user_id')['resid'].agg(['cumsum'])\ntrain['user_resid_rolling'] = train['user_resid_rolling'].fillna(0)\ntrain.drop(['resid', 'content_mean'], axis=1, inplace=True)\n```\n\n\"last_lec_ts\" & \"diff_to_last_lec_ts\" -> \n\n```\nprev_u, lec_ts = -1, 0\nlast_lec_ts = []\n\nfor u, c_type, t in zip(train['user_id'], train['content_type_id'], train['timestamp']):\n    if u != prev_u:\n        lec_ts = 0\n    if c_type == True:\n        lec_ts = t\n    prev_u = u\n    last_lec_ts.append(lec_ts)\n\ntrain['last_lec_ts'] = np.array(last_lec_ts)\ntrain['diff_to_last_lec_ts'] = train['timestamp'] - train['last_lec_ts']\n```\n\n\nAnd the craziest of all FE and a  cool FE, **We saved ~.3M pkl files during runtime for the FE \"diff_content_time\",** (do delete the dir after your inference is complete) (basically for every user difference b/w the TS b/w their current and last TS for a particular content_id)\n\n```\n!mkdir pkld\ndct = {}\ndct[-1] = 0\nprev_u = -1\nres = np.zeros(3)\nfor u, cid, ctype, ts in tqdm(zip(train['user_id'], train['content_id'], train['content_type_id'], train['timestamp'])):\n    if prev_u != u:\n        pickle.dump(dct[prev_u], open('pkld/' + str(prev_u) + '.pkl', 'wb'), protocol=pickle.HIGHEST_PROTOCOL)\n        del dct[prev_u]\n        dct[u] = {}\n    if ctype == False:\n        dct[u][cid] = ts\n    prev_u = u\npickle.dump(dct[prev_u], open('pkld/' + str(prev_u) + '.pkl', 'wb'), protocol=pickle.HIGHEST_PROTOCOL) #########\n\ndel dct\n```\n\nThere were a lot of FE ideas that were tried but didn't help us. To mention few, \"tf-idf\" on content_ids history of the user, rolling_windows using numba, FTRL's, and other hell lot of FE's ideas that gave marginal boost but we had a rule to add +.001 single-single features in place.\n\nWhat we regret was definitely not having a SAINT like arch for sure. I messed up badly, should have simply used nn.Transformer rather than handling the encoder/decoder parts myself.\n\n\nBut until next time,\n\nHappy Fair Kaggling & Keep Sharing A Ton!\n\ncc @alexj21 @nikhilmishradev @rohanrao  [Alphabetical Order]. (Please add anything that i missed)\nThanks a lot for the collaboration guys!",
    "1143803": "Congratulations Aditya and team. Great feature engineering!",
    "1143807": "Thanks a lot Chris! Great Work for yet another Gold! Any chances you are going to share your team's solution's as well? Ty!",
    "1143815": "Thanks. We'll share tomorrow. It's a single SAINT+ with 5 features. CV 812 LB 812.",
    "1143861": "Congrats on results @adityaecdrid and team",
    "1143864": "Thanks a lot!",
    "1144070": "Thanks for sharing especially about your FE!",
    "1144130": "adityaecdrid  Congratulations Aditya ! And thanks for the active discussions!\n\n@cdeotte Only 5 features ...? OMG, that's great - yet It's killing me, as I added a lot of features to improve my SAINT + like model, and yet introduced bugs in the inference time ... (CV improved, but LB worse ...)",
    "1144391": "Congratz for the nice finish ! \n\nOnly one small step before competition master @adityaecdrid  :)",
    "1144394": "Congratulations to @adityaecdrid and team! You were probably the person who engaged with us all the most for this competition :P",
    "1144471": "Thanks a lot for the honour but i would like to share the same with many others as well competing, both after the comp ends and during the comp! It's because of these discussions, we all made it possible!",
    "1144474": "Thanks a lot Theo! Yep, just one more silver/gold medal and finally, I can proudly say, I am a Comp Master!",
    "1144476": "Thanks a lot! i saw you joined early on but i guess you were busy with other things in // !",
    "1147494": "So bright method! I am ashamed that I dont know FE. Can you explain more about it or which thread talk about it?",
    "1148930": "Great work gays, congratulations !\n\n- By curiosity, are `answer_1 ... answer_7` the selected answer (0,1,2,3) or whether the user was right/wrong ? I'm surprised that such features don't lead to overfitting\n- How does your store-to-pkl trick impact the inference time ?\n\nCheers"
  },
  "source": "meta"
}