{
  "id": 191856,
  "title": "A possible solution to partial_fit on chunks rather than on a single batch!",
  "url": "/competitions/riiid-test-answer-prediction/discussion/191856",
  "author_name": "Aditya Soni",
  "post_date": "2020-10-19T03:41:34.088000",
  "votes": 8,
  "comment_count": 10,
  "views": 0,
  "content": "<p>Well in all the current public kernels that uses some sort of active/online learning (e,g kernels <a href=\"https://www.kaggle.com/spacelx/2020-r3id-incremental-learning-pytorch-creme\" target=\"_blank\">kernel_1 by </a><a href=\"https://www.kaggle.com/spacelx\" target=\"_blank\">@spacelx</a>, <a href=\"https://www.kaggle.com/rohanrao/riiid-ftrl-ftw\" target=\"_blank\">kernel_2 by </a><a href=\"https://www.kaggle.com/rohanrao\" target=\"_blank\">@rohanrao</a>, <a href=\"https://www.kaggle.com/dwit392/expanding-on-simple-lgbm\" target=\"_blank\">multiple baseline kernels by </a><a href=\"https://www.kaggle.com/dwit392\" target=\"_blank\">@dwit392</a> and others as well), they all do it lazily, that's they all re-fit on every batch of the test_df which we get which might be a sub optimal thing but it's what you should do when getting started and making a successful sub first as that's important in this comp. After that, you should try and complicate things slightly so as to get better performance and get better at coding yet ensuring the logic is dead-simple!</p>\n<p>Also, i would like to thanks <a href=\"https://www.kaggle.com/rously\" target=\"_blank\">@rously</a> for this <a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/190439\" target=\"_blank\">discussion</a> fake-test-df-api as well; There are some differences wrt the original but they can be easily figured out, like the nan one and in this one, the author just has a blank when it's videos..</p>\n<p></p>\n<p>Update-&gt; <strong>The below works and we have a successful submission for 1k chunks! 🎊🎊🎊🎉🎉🎉</strong></p>\n<pre><code>import riiideducation\n# imports that i have missed goes here\n\ndef chunking():\n\n    last_k_batches = pd.DataFrame()\n    mod_by = 3 # tuneable as nothings hard-coded here...\n    prev_answers = [] # caching the answer labels we get at next batch of the data\n    env = riiideducation.make_env()\n    iter_test = env.iter_test()\n\n    for current_group_num, (test_df, sample_prediction_df) in enumerate(iter_test, 0):\n        if current_group_num % mod_by == 0 and current_group_num != 0:\n            # re-training blocks\n            prev_answers = np.array(prev_answers) # remember to change this back to list..\n            video_index = np.where(prev_answers == -1)[0].tolist()\n            prev_answers = np.delete(prev_answers, video_index)\n            # removing the video rows now...\n            last_k_batches = last_k_batches[last_k_batches['content_type_id'] == 0]\n            # splitting data for which we have/don't have labels...\n            last_k_batches.reset_index(drop=True, inplace=True)\n            df_train, last_k_batches = last_k_batches[last_k_batches.group_num &lt; current_group_num-1], last_k_batches[last_k_batches.group_num &gt;= current_group_num-1]\n\n            ###################################\n            #######  re-training logic goes here #######\n            ####### have added dummy idea #########\n            ###################################\n\n            # print(\"re-training-the-model-now\")\n            df_train = df_train[df_train['content_type_id'] == 0] # remove lecture rows in re-training data\n            # assign the correct labels\n            df_train.loc[:, \"prior_group_answers_correct\"] = prev_answers\n            ########################\n            # re-training logic goes here...#\n            ########################\n            # reset prev_answers\n            prev_answers = []\n            del df_train, video_index (thanks for the tip!)\n            gc.collect()\n\n        if str(test_df.iloc[0][\"prior_group_answers_correct\"]) == \"nan\":\n            # logic to make it compatible with the fake_test api as well...\n            test_df.loc[0][\"prior_group_answers_correct\"] = '[]'\n\n        # make a copy of the current batch test_df\n        _test_df = test_df.copy(deep=True)\n        # getting the answers of the current batch and merging it\n        # NB it still has videos in it, we will handle it later in training part..\n        prev_answers += eval(_test_df.iloc[0][\"prior_group_answers_correct\"])\n        last_k_batches = pd.concat([last_k_batches, _test_df.reset_index()], ignore_index=True)\n        # removing the video rows.\n        _test_df = _test_df[_test_df['content_type_id'] == 0]\n        # making predictions on the same\n        _test_df['answered_correctly'] = 0.5 # dummy\n        env.predict(_test_df.loc[:,['row_id', 'answered_correctly']])\n        del _test_df\n        gc.collect()\n</code></pre>\n<p>Credits,</p>\n<p>I loved doing this and realised that i should be more confident as to what every single line of code i am writing is going to do in future.. Apologies to kaggle for my -ve thoughts earlier regarding the comp's format as it's little painful, but I guess, I am going to love it going forward irrespective of where I land into final LB. <strong>And I would suggest to every participant, spend a day reading the data_desc, it's very important that you understand everything there rather than just throwing fancy modelling ideas.</strong></p>\n<p>Also I feel that people should kinda <strong>merge little earlier</strong> in this comp as opposed to last 2 weeks as they might have their own logic to handle the test_df's, at the very end it might be a mess!</p>\n<p>Lastly, i would like to thanks many kagglers who are active in forums wrt this comp as of now, the discussions have all the answers you are looking for, so search first before posting a new topic…</p>\n<p>Best,<br>\nAditya.</p>",
  "messages": [
    {
      "id": 1053485,
      "postDate": "2020-10-19T03:41:34.090Z",
      "content": "<p>Well in all the current public kernels that uses some sort of active/online learning (e,g kernels <a href=\"https://www.kaggle.com/spacelx/2020-r3id-incremental-learning-pytorch-creme\" target=\"_blank\">kernel_1 by </a><a href=\"https://www.kaggle.com/spacelx\" target=\"_blank\">@spacelx</a>, <a href=\"https://www.kaggle.com/rohanrao/riiid-ftrl-ftw\" target=\"_blank\">kernel_2 by </a><a href=\"https://www.kaggle.com/rohanrao\" target=\"_blank\">@rohanrao</a>, <a href=\"https://www.kaggle.com/dwit392/expanding-on-simple-lgbm\" target=\"_blank\">multiple baseline kernels by </a><a href=\"https://www.kaggle.com/dwit392\" target=\"_blank\">@dwit392</a> and others as well), they all do it lazily, that's they all re-fit on every batch of the test_df which we get which might be a sub optimal thing but it's what you should do when getting started and making a successful sub first as that's important in this comp. After that, you should try and complicate things slightly so as to get better performance and get better at coding yet ensuring the logic is dead-simple!</p>\n<p>Also, i would like to thanks <a href=\"https://www.kaggle.com/rously\" target=\"_blank\">@rously</a> for this <a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/190439\" target=\"_blank\">discussion</a> fake-test-df-api as well; There are some differences wrt the original but they can be easily figured out, like the nan one and in this one, the author just has a blank when it's videos..</p>\n<p></p>\n<p>Update-&gt; <strong>The below works and we have a successful submission for 1k chunks! 🎊🎊🎊🎉🎉🎉</strong></p>\n<pre><code>import riiideducation\n# imports that i have missed goes here\n\ndef chunking():\n\n    last_k_batches = pd.DataFrame()\n    mod_by = 3 # tuneable as nothings hard-coded here...\n    prev_answers = [] # caching the answer labels we get at next batch of the data\n    env = riiideducation.make_env()\n    iter_test = env.iter_test()\n\n    for current_group_num, (test_df, sample_prediction_df) in enumerate(iter_test, 0):\n        if current_group_num % mod_by == 0 and current_group_num != 0:\n            # re-training blocks\n            prev_answers = np.array(prev_answers) # remember to change this back to list..\n            video_index = np.where(prev_answers == -1)[0].tolist()\n            prev_answers = np.delete(prev_answers, video_index)\n            # removing the video rows now...\n            last_k_batches = last_k_batches[last_k_batches['content_type_id'] == 0]\n            # splitting data for which we have/don't have labels...\n            last_k_batches.reset_index(drop=True, inplace=True)\n            df_train, last_k_batches = last_k_batches[last_k_batches.group_num &lt; current_group_num-1], last_k_batches[last_k_batches.group_num &gt;= current_group_num-1]\n\n            ###################################\n            #######  re-training logic goes here #######\n            ####### have added dummy idea #########\n            ###################################\n\n            # print(\"re-training-the-model-now\")\n            df_train = df_train[df_train['content_type_id'] == 0] # remove lecture rows in re-training data\n            # assign the correct labels\n            df_train.loc[:, \"prior_group_answers_correct\"] = prev_answers\n            ########################\n            # re-training logic goes here...#\n            ########################\n            # reset prev_answers\n            prev_answers = []\n            del df_train, video_index (thanks for the tip!)\n            gc.collect()\n\n        if str(test_df.iloc[0][\"prior_group_answers_correct\"]) == \"nan\":\n            # logic to make it compatible with the fake_test api as well...\n            test_df.loc[0][\"prior_group_answers_correct\"] = '[]'\n\n        # make a copy of the current batch test_df\n        _test_df = test_df.copy(deep=True)\n        # getting the answers of the current batch and merging it\n        # NB it still has videos in it, we will handle it later in training part..\n        prev_answers += eval(_test_df.iloc[0][\"prior_group_answers_correct\"])\n        last_k_batches = pd.concat([last_k_batches, _test_df.reset_index()], ignore_index=True)\n        # removing the video rows.\n        _test_df = _test_df[_test_df['content_type_id'] == 0]\n        # making predictions on the same\n        _test_df['answered_correctly'] = 0.5 # dummy\n        env.predict(_test_df.loc[:,['row_id', 'answered_correctly']])\n        del _test_df\n        gc.collect()\n</code></pre>\n<p>Credits,</p>\n<p>I loved doing this and realised that i should be more confident as to what every single line of code i am writing is going to do in future.. Apologies to kaggle for my -ve thoughts earlier regarding the comp's format as it's little painful, but I guess, I am going to love it going forward irrespective of where I land into final LB. <strong>And I would suggest to every participant, spend a day reading the data_desc, it's very important that you understand everything there rather than just throwing fancy modelling ideas.</strong></p>\n<p>Also I feel that people should kinda <strong>merge little earlier</strong> in this comp as opposed to last 2 weeks as they might have their own logic to handle the test_df's, at the very end it might be a mess!</p>\n<p>Lastly, i would like to thanks many kagglers who are active in forums wrt this comp as of now, the discussions have all the answers you are looking for, so search first before posting a new topic…</p>\n<p>Best,<br>\nAditya.</p>",
      "rawMarkdown": "Well in all the current public kernels that uses some sort of active/online learning (e,g kernels [kernel_1 by @spacelx](https://www.kaggle.com/spacelx/2020-r3id-incremental-learning-pytorch-creme), [kernel_2 by @rohanrao](https://www.kaggle.com/rohanrao/riiid-ftrl-ftw), [multiple baseline kernels by @dwit392](https://www.kaggle.com/dwit392/expanding-on-simple-lgbm) and others as well), they all do it lazily, that's they all re-fit on every batch of the test_df which we get which might be a sub optimal thing but it's what you should do when getting started and making a successful sub first as that's important in this comp. After that, you should try and complicate things slightly so as to get better performance and get better at coding yet ensuring the logic is dead-simple!\n\nAlso, i would like to thanks @rously for this [discussion](https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/190439) fake-test-df-api as well; There are some differences wrt the original but they can be easily figured out, like the nan one and in this one, the author just has a blank when it's videos..\n\n\n~~**NB the submission I made using below is still running, so I will update if it crashes!**\n**Plus I have made 2 subs, one with 1000 chunks cached and other one with just 3 chunks cached, so results might take some time to come.**~~\n\n\nUpdate-> **The below works and we have a successful submission for 1k chunks! 🎊🎊🎊🎉🎉🎉**\n\n```\nimport riiideducation\n# imports that i have missed goes here\n\ndef chunking():\n    \n    last_k_batches = pd.DataFrame()\n    mod_by = 3 # tuneable as nothings hard-coded here...\n    prev_answers = [] # caching the answer labels we get at next batch of the data\n    env = riiideducation.make_env()\n    iter_test = env.iter_test()\n\n    for current_group_num, (test_df, sample_prediction_df) in enumerate(iter_test, 0):\n        if current_group_num % mod_by == 0 and current_group_num != 0:\n            # re-training blocks\n            prev_answers = np.array(prev_answers) # remember to change this back to list..\n            video_index = np.where(prev_answers == -1)[0].tolist()\n            prev_answers = np.delete(prev_answers, video_index)\n            # removing the video rows now...\n            last_k_batches = last_k_batches[last_k_batches['content_type_id'] == 0]\n            # splitting data for which we have/don't have labels...\n            last_k_batches.reset_index(drop=True, inplace=True)\n            df_train, last_k_batches = last_k_batches[last_k_batches.group_num < current_group_num-1], last_k_batches[last_k_batches.group_num >= current_group_num-1]\n            \n            ###################################\n            #######  re-training logic goes here #######\n            ####### have added dummy idea #########\n            ###################################\n            \n            # print(\"re-training-the-model-now\")\n            df_train = df_train[df_train['content_type_id'] == 0] # remove lecture rows in re-training data\n            # assign the correct labels\n            df_train.loc[:, \"prior_group_answers_correct\"] = prev_answers\n            ########################\n            # re-training logic goes here...#\n            ########################\n            # reset prev_answers\n            prev_answers = []\n            del df_train, video_index (thanks for the tip!)\n            gc.collect()\n            \n        if str(test_df.iloc[0][\"prior_group_answers_correct\"]) == \"nan\":\n            # logic to make it compatible with the fake_test api as well...\n            test_df.loc[0][\"prior_group_answers_correct\"] = '[]'\n        \n        # make a copy of the current batch test_df\n        _test_df = test_df.copy(deep=True)\n        # getting the answers of the current batch and merging it\n        # NB it still has videos in it, we will handle it later in training part..\n        prev_answers += eval(_test_df.iloc[0][\"prior_group_answers_correct\"])\n        last_k_batches = pd.concat([last_k_batches, _test_df.reset_index()], ignore_index=True)\n        # removing the video rows.\n        _test_df = _test_df[_test_df['content_type_id'] == 0]\n        # making predictions on the same\n        _test_df['answered_correctly'] = 0.5 # dummy\n        env.predict(_test_df.loc[:,['row_id', 'answered_correctly']])\n        del _test_df\n        gc.collect()\n\n```\n\n\nCredits,\n\nI loved doing this and realised that i should be more confident as to what every single line of code i am writing is going to do in future.. Apologies to kaggle for my -ve thoughts earlier regarding the comp's format as it's little painful, but I guess, I am going to love it going forward irrespective of where I land into final LB. **And I would suggest to every participant, spend a day reading the data_desc, it's very important that you understand everything there rather than just throwing fancy modelling ideas.**\n\nAlso I feel that people should kinda **merge little earlier** in this comp as opposed to last 2 weeks as they might have their own logic to handle the test_df's, at the very end it might be a mess!\n\nLastly, i would like to thanks many kagglers who are active in forums wrt this comp as of now, the discussions have all the answers you are looking for, so search first before posting a new topic...\n\nBest,\nAditya.",
      "votes": 8
    },
    {
      "id": 1053539,
      "postDate": "2020-10-19T05:32:26.997Z",
      "content": "<p>Good job! So I do get that it may be considered sub-optimal to re-fit the model on every batch from a perspective of speed and efficiency - is there any downside from a logical perspective which I'm not aware of? <br>\nI would have thought it beneficial to update with every chunk we get so that we can update features (say recent user correctness for example, although that could be done separately from updating the model) and so that the model can quickly pick up on dynamic changes (although these are probably going to be small in this challenge).</p>",
      "rawMarkdown": "Good job! So I do get that it may be considered sub-optimal to re-fit the model on every batch from a perspective of speed and efficiency - is there any downside from a logical perspective which I'm not aware of? \nI would have thought it beneficial to update with every chunk we get so that we can update features (say recent user correctness for example, although that could be done separately from updating the model) and so that the model can quickly pick up on dynamic changes (although these are probably going to be small in this challenge).",
      "votes": 2,
      "replies": [
        {
          "id": 1053587,
          "postDate": "2020-10-19T06:07:41.363Z",
          "content": "<p>Yep, i believe the time to run this will be similar as you will be using the time saved by this way to <strong>update</strong> your stats pretty much if you want to do it making it more dynamic as a humans' performance does shift. E.g. if i haven't practised Competitive Programming for like a month, and in the next month i do it, it's possible that i will have a lower rank on the points table let's say or will take little more time than i used to take for solving such problems as the idea was fresh in the past but it kinda is clouded after 1-2 months. Plus now we can update our historical stats over past users and compute it for new ones as well though that also needs expertise and add's to complexity overall involved…</p>\n<p>But this can open up new doors, so i just did it nonetheless :)</p>",
          "rawMarkdown": "Yep, i believe the time to run this will be similar as you will be using the time saved by this way to **update** your stats pretty much if you want to do it making it more dynamic as a humans' performance does shift. E.g. if i haven't practised Competitive Programming for like a month, and in the next month i do it, it's possible that i will have a lower rank on the points table let's say or will take little more time than i used to take for solving such problems as the idea was fresh in the past but it kinda is clouded after 1-2 months. Plus now we can update our historical stats over past users and compute it for new ones as well though that also needs expertise and add's to complexity overall involved...\n\nBut this can open up new doors, so i just did it nonetheless :)",
          "votes": 1
        }
      ]
    },
    {
      "id": 1053594,
      "postDate": "2020-10-19T06:21:40.857Z",
      "content": "<p><a href=\"https://www.kaggle.com/spacelx\" target=\"_blank\">@spacelx</a> Let's say you want to compute a feature as depicted in this <a href=\"https://www.kaggle.com/dwit392/riiid-challenge-time-since-last-action-for-test\" target=\"_blank\">kernel</a>, How would you do it at individual batch level?</p>",
      "rawMarkdown": "@spacelx Let's say you want to compute a feature as depicted in this [kernel](https://www.kaggle.com/dwit392/riiid-challenge-time-since-last-action-for-test), How would you do it at individual batch level?",
      "replies": [
        {
          "id": 1053600,
          "postDate": "2020-10-19T06:46:52.537Z",
          "content": "<p>I'm not sure I see the difference between doing this for an individual batch or for a concatenation of some batches… <br>\n<code>last_record</code> seems to be coming from a table holding a list of all users with their last recorded timestamp, and whether we use this information to calculate the time since the last interaction for one user or a series of users doesn't make a difference. <br>\nWhere it <em>might</em> make a difference though is when you have several entries of the same user in your concatenated larger batch - in the given code example the calculation is done entry by entry through a loop which works fine if a user's records within the large batch are sorted by time, but it fails if they aren't - you'd soon hit cases where the latest recorded timestamp is larger than the timestamp of the record you're currently looking at.<br>\nIn terms of efficiency, it would make sense to use some pandas logic like groupby/merge and such anyway to do this calculation, but then you'd also have to specifically handle the problem of having several records of the same user in your batch.</p>\n<p>So I'd strongly prefer to do this at individual batch level rather than on large combined batches. Sure, you <em>can</em> do it on larger batches too if you take care in the code, but I simply don't see any benefit in that.</p>\n<p>Am I missing something?</p>",
          "rawMarkdown": "I'm not sure I see the difference between doing this for an individual batch or for a concatenation of some batches... \n`last_record` seems to be coming from a table holding a list of all users with their last recorded timestamp, and whether we use this information to calculate the time since the last interaction for one user or a series of users doesn't make a difference. \nWhere it *might* make a difference though is when you have several entries of the same user in your concatenated larger batch - in the given code example the calculation is done entry by entry through a loop which works fine if a user's records within the large batch are sorted by time, but it fails if they aren't - you'd soon hit cases where the latest recorded timestamp is larger than the timestamp of the record you're currently looking at.\nIn terms of efficiency, it would make sense to use some pandas logic like groupby/merge and such anyway to do this calculation, but then you'd also have to specifically handle the problem of having several records of the same user in your batch.\n\nSo I'd strongly prefer to do this at individual batch level rather than on large combined batches. Sure, you *can* do it on larger batches too if you take care in the code, but I simply don't see any benefit in that.\n\nAm I missing something?",
          "votes": 2
        },
        {
          "id": 1053602,
          "postDate": "2020-10-19T06:51:46.907Z",
          "content": "<p>Cool! Thanks; Guess i invested time on something that's not useful then (as of now) :( Thanks a lot!</p>\n<blockquote>\n  <p>at individual batch level rather than on large combined batches.</p>\n</blockquote>\n<p>Then I don't think we will be able to \"detect\" continued-sessions directly (without having any other logic) as there will be more or less only one task_container_id in any given batch… TO do so, you need to keep track of last_max and then check if some threshold + that value has become less now.. (Kadane's algo)</p>",
          "rawMarkdown": "Cool! Thanks; Guess i invested time on something that's not useful then (as of now) :( Thanks a lot!\n\n>at individual batch level rather than on large combined batches.\n\nThen I don't think we will be able to \"detect\" continued-sessions directly (without having any other logic) as there will be more or less only one task_container_id in any given batch... TO do so, you need to keep track of last_max and then check if some threshold + that value has become less now.. (Kadane's algo)"
        },
        {
          "id": 1053608,
          "postDate": "2020-10-19T06:56:48.513Z",
          "content": "<p>I mean it could come in handy if you do a lot of feature engineering and your kernel has trouble finishing within 9 hours. And if your feature engineering has less than O(N) complexity.</p>",
          "rawMarkdown": "I mean it could come in handy if you do a lot of feature engineering and your kernel has trouble finishing within 9 hours. And if your feature engineering has less than O(N) complexity."
        }
      ]
    },
    {
      "id": 1053530,
      "postDate": "2020-10-19T05:18:00.970Z",
      "content": "<blockquote>\n  <p>re-fit on every batch of the test_df which we get which might be a sub optimal thing</p>\n</blockquote>\n<p>Can you explain why is this sub-optimal? How else could it be done?</p>",
      "rawMarkdown": "> re-fit on every batch of the test_df which we get which might be a sub optimal thing\n\nCan you explain why is this sub-optimal? How else could it be done?",
      "replies": [
        {
          "id": 1053548,
          "postDate": "2020-10-19T05:45:39.873Z",
          "content": "<p>The below is completely my understanding, please correct me!</p>\n<p>As per the API's desc, </p>\n<blockquote>\n  <p>The API provides user interactions groups in the order in which they occurred. Each group will contain interactions from many different users, but no more than one task_container_id of questions from any single user. Each group has between 1 and 1000 users.</p>\n</blockquote>\n<p>So, that means a particular user will only have one \"task_container_id\" in the batch we receive. And it's likely that we will get another record from the same user in the next batch if the user was active and had new interactions. So if we want to capture this information, we are bound to accumulate test set over time for \"k\" chunks let's say, then we have a better chance that this will also come and your model will see it as well (with the old one's). Plus speed and efficiency might be better in this case, though i need to run some tests to make a statement on that..</p>\n<p>Obviously one can argue that this also suffers from the same thing mentioned above but i feel it does slightly a better job in handling the same. Assumption being made here is that people do want to update their new feat's value's over time depending on the students performance in future rather than using the historical statistics…</p>\n<p>Let me know if my understanding is incorrect,</p>\n<p>Best,<br>\nAditya.</p>",
          "rawMarkdown": "The below is completely my understanding, please correct me!\n\nAs per the API's desc, \n\n>The API provides user interactions groups in the order in which they occurred. Each group will contain interactions from many different users, but no more than one task_container_id of questions from any single user. Each group has between 1 and 1000 users.\n\nSo, that means a particular user will only have one \"task_container_id\" in the batch we receive. And it's likely that we will get another record from the same user in the next batch if the user was active and had new interactions. So if we want to capture this information, we are bound to accumulate test set over time for \"k\" chunks let's say, then we have a better chance that this will also come and your model will see it as well (with the old one's). Plus speed and efficiency might be better in this case, though i need to run some tests to make a statement on that..\n\nObviously one can argue that this also suffers from the same thing mentioned above but i feel it does slightly a better job in handling the same. Assumption being made here is that people do want to update their new feat's value's over time depending on the students performance in future rather than using the historical statistics...\n\nLet me know if my understanding is incorrect,\n\nBest,\nAditya."
        },
        {
          "id": 1053585,
          "postDate": "2020-10-19T06:07:16.070Z",
          "content": "<p>I don't think the number of observations of a user per group matters.</p>\n<p>My understanding is aligned with <a href=\"https://www.kaggle.com/spacelx\" target=\"_blank\">@spacelx</a> 's comment above:</p>\n<blockquote>\n  <p>I would have thought it beneficial to update with every chunk we get so that we can update features and so that the model can quickly pick up on dynamic changes</p>\n</blockquote>\n<p>Let's look at the first three batches: B1, B2, B3.</p>\n<ul>\n<li><p><strong>Method 1:</strong> Create features for B1. Score B1. Increment model. Create features for B2. Score B2. Increment model. Create features for B3 (Using only original or original + B1 or original + B1 + B2). Score B3.</p></li>\n<li><p><strong>Method 2:</strong> Create features for B1. Score B1. Create features for B2. Score B2. Increment model. Create features for B3 (Using only original or original + B1 or original + B1 + B2). Score B3.</p></li>\n</ul>\n<p>I can't think of a good reason why you would use Method 2 except if the runtime of incrementing model is high. Note that the feature values in Method 1 and Method 2 will be exactly same (even while creating features for B3, you can choose to only use B1 or use B1+B2 or use none for both methods), only the model weights change between the two methods.</p>",
          "rawMarkdown": "I don't think the number of observations of a user per group matters.\n\nMy understanding is aligned with @spacelx 's comment above:\n> I would have thought it beneficial to update with every chunk we get so that we can update features and so that the model can quickly pick up on dynamic changes\n\nLet's look at the first three batches: B1, B2, B3.\n\n* **Method 1:** Create features for B1. Score B1. Increment model. Create features for B2. Score B2. Increment model. Create features for B3 (Using only original or original + B1 or original + B1 + B2). Score B3.\n\n* **Method 2:** Create features for B1. Score B1. Create features for B2. Score B2. Increment model. Create features for B3 (Using only original or original + B1 or original + B1 + B2). Score B3.\n\nI can't think of a good reason why you would use Method 2 except if the runtime of incrementing model is high. Note that the feature values in Method 1 and Method 2 will be exactly same (even while creating features for B3, you can choose to only use B1 or use B1+B2 or use none for both methods), only the model weights change between the two methods.",
          "votes": 2
        },
        {
          "id": 1053590,
          "postDate": "2020-10-19T06:13:36.987Z",
          "rawMarkdown": "",
          "votes": -1,
          "isDeleted": true
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 1053539,
      "author_name": "Alex Bader",
      "author_url": "",
      "post_date": "2020-10-19T05:32:26.997000",
      "content": "<p>Good job! So I do get that it may be considered sub-optimal to re-fit the model on every batch from a perspective of speed and efficiency - is there any downside from a logical perspective which I'm not aware of? <br>\nI would have thought it beneficial to update with every chunk we get so that we can update features (say recent user correctness for example, although that could be done separately from updating the model) and so that the model can quickly pick up on dynamic changes (although these are probably going to be small in this challenge).</p>",
      "votes": 2,
      "replies": [
        {
          "id": 1053587,
          "author_name": "Aditya Soni",
          "author_url": "",
          "post_date": "2020-10-19T06:07:41.363000",
          "content": "<p>Yep, i believe the time to run this will be similar as you will be using the time saved by this way to <strong>update</strong> your stats pretty much if you want to do it making it more dynamic as a humans' performance does shift. E.g. if i haven't practised Competitive Programming for like a month, and in the next month i do it, it's possible that i will have a lower rank on the points table let's say or will take little more time than i used to take for solving such problems as the idea was fresh in the past but it kinda is clouded after 1-2 months. Plus now we can update our historical stats over past users and compute it for new ones as well though that also needs expertise and add's to complexity overall involved…</p>\n<p>But this can open up new doors, so i just did it nonetheless :)</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1053594,
      "author_name": "Aditya Soni",
      "author_url": "",
      "post_date": "2020-10-19T06:21:40.857000",
      "content": "<p><a href=\"https://www.kaggle.com/spacelx\" target=\"_blank\">@spacelx</a> Let's say you want to compute a feature as depicted in this <a href=\"https://www.kaggle.com/dwit392/riiid-challenge-time-since-last-action-for-test\" target=\"_blank\">kernel</a>, How would you do it at individual batch level?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1053600,
          "author_name": "Alex Bader",
          "author_url": "",
          "post_date": "2020-10-19T06:46:52.537000",
          "content": "<p>I'm not sure I see the difference between doing this for an individual batch or for a concatenation of some batches… <br>\n<code>last_record</code> seems to be coming from a table holding a list of all users with their last recorded timestamp, and whether we use this information to calculate the time since the last interaction for one user or a series of users doesn't make a difference. <br>\nWhere it <em>might</em> make a difference though is when you have several entries of the same user in your concatenated larger batch - in the given code example the calculation is done entry by entry through a loop which works fine if a user's records within the large batch are sorted by time, but it fails if they aren't - you'd soon hit cases where the latest recorded timestamp is larger than the timestamp of the record you're currently looking at.<br>\nIn terms of efficiency, it would make sense to use some pandas logic like groupby/merge and such anyway to do this calculation, but then you'd also have to specifically handle the problem of having several records of the same user in your batch.</p>\n<p>So I'd strongly prefer to do this at individual batch level rather than on large combined batches. Sure, you <em>can</em> do it on larger batches too if you take care in the code, but I simply don't see any benefit in that.</p>\n<p>Am I missing something?</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1053602,
          "author_name": "Aditya Soni",
          "author_url": "",
          "post_date": "2020-10-19T06:51:46.907000",
          "content": "<p>Cool! Thanks; Guess i invested time on something that's not useful then (as of now) :( Thanks a lot!</p>\n<blockquote>\n  <p>at individual batch level rather than on large combined batches.</p>\n</blockquote>\n<p>Then I don't think we will be able to \"detect\" continued-sessions directly (without having any other logic) as there will be more or less only one task_container_id in any given batch… TO do so, you need to keep track of last_max and then check if some threshold + that value has become less now.. (Kadane's algo)</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1053608,
          "author_name": "Alex Bader",
          "author_url": "",
          "post_date": "2020-10-19T06:56:48.513000",
          "content": "<p>I mean it could come in handy if you do a lot of feature engineering and your kernel has trouble finishing within 9 hours. And if your feature engineering has less than O(N) complexity.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1053530,
      "author_name": "Vopani",
      "author_url": "",
      "post_date": "2020-10-19T05:18:00.970000",
      "content": "<blockquote>\n  <p>re-fit on every batch of the test_df which we get which might be a sub optimal thing</p>\n</blockquote>\n<p>Can you explain why is this sub-optimal? How else could it be done?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1053548,
          "author_name": "Aditya Soni",
          "author_url": "",
          "post_date": "2020-10-19T05:45:39.873000",
          "content": "<p>The below is completely my understanding, please correct me!</p>\n<p>As per the API's desc, </p>\n<blockquote>\n  <p>The API provides user interactions groups in the order in which they occurred. Each group will contain interactions from many different users, but no more than one task_container_id of questions from any single user. Each group has between 1 and 1000 users.</p>\n</blockquote>\n<p>So, that means a particular user will only have one \"task_container_id\" in the batch we receive. And it's likely that we will get another record from the same user in the next batch if the user was active and had new interactions. So if we want to capture this information, we are bound to accumulate test set over time for \"k\" chunks let's say, then we have a better chance that this will also come and your model will see it as well (with the old one's). Plus speed and efficiency might be better in this case, though i need to run some tests to make a statement on that..</p>\n<p>Obviously one can argue that this also suffers from the same thing mentioned above but i feel it does slightly a better job in handling the same. Assumption being made here is that people do want to update their new feat's value's over time depending on the students performance in future rather than using the historical statistics…</p>\n<p>Let me know if my understanding is incorrect,</p>\n<p>Best,<br>\nAditya.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1053585,
          "author_name": "Vopani",
          "author_url": "",
          "post_date": "2020-10-19T06:07:16.070000",
          "content": "<p>I don't think the number of observations of a user per group matters.</p>\n<p>My understanding is aligned with <a href=\"https://www.kaggle.com/spacelx\" target=\"_blank\">@spacelx</a> 's comment above:</p>\n<blockquote>\n  <p>I would have thought it beneficial to update with every chunk we get so that we can update features and so that the model can quickly pick up on dynamic changes</p>\n</blockquote>\n<p>Let's look at the first three batches: B1, B2, B3.</p>\n<ul>\n<li><p><strong>Method 1:</strong> Create features for B1. Score B1. Increment model. Create features for B2. Score B2. Increment model. Create features for B3 (Using only original or original + B1 or original + B1 + B2). Score B3.</p></li>\n<li><p><strong>Method 2:</strong> Create features for B1. Score B1. Create features for B2. Score B2. Increment model. Create features for B3 (Using only original or original + B1 or original + B1 + B2). Score B3.</p></li>\n</ul>\n<p>I can't think of a good reason why you would use Method 2 except if the runtime of incrementing model is high. Note that the feature values in Method 1 and Method 2 will be exactly same (even while creating features for B3, you can choose to only use B1 or use B1+B2 or use none for both methods), only the model weights change between the two methods.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1053590,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-10-19T06:13:36.987000",
          "content": "",
          "votes": -1,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1053485": "Well in all the current public kernels that uses some sort of active/online learning (e,g kernels [kernel_1 by @spacelx](https://www.kaggle.com/spacelx/2020-r3id-incremental-learning-pytorch-creme), [kernel_2 by @rohanrao](https://www.kaggle.com/rohanrao/riiid-ftrl-ftw), [multiple baseline kernels by @dwit392](https://www.kaggle.com/dwit392/expanding-on-simple-lgbm) and others as well), they all do it lazily, that's they all re-fit on every batch of the test_df which we get which might be a sub optimal thing but it's what you should do when getting started and making a successful sub first as that's important in this comp. After that, you should try and complicate things slightly so as to get better performance and get better at coding yet ensuring the logic is dead-simple!\n\nAlso, i would like to thanks @rously for this [discussion](https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/190439) fake-test-df-api as well; There are some differences wrt the original but they can be easily figured out, like the nan one and in this one, the author just has a blank when it's videos..\n\n\n~~**NB the submission I made using below is still running, so I will update if it crashes!**\n**Plus I have made 2 subs, one with 1000 chunks cached and other one with just 3 chunks cached, so results might take some time to come.**~~\n\n\nUpdate-> **The below works and we have a successful submission for 1k chunks! 🎊🎊🎊🎉🎉🎉**\n\n```\nimport riiideducation\n# imports that i have missed goes here\n\ndef chunking():\n    \n    last_k_batches = pd.DataFrame()\n    mod_by = 3 # tuneable as nothings hard-coded here...\n    prev_answers = [] # caching the answer labels we get at next batch of the data\n    env = riiideducation.make_env()\n    iter_test = env.iter_test()\n\n    for current_group_num, (test_df, sample_prediction_df) in enumerate(iter_test, 0):\n        if current_group_num % mod_by == 0 and current_group_num != 0:\n            # re-training blocks\n            prev_answers = np.array(prev_answers) # remember to change this back to list..\n            video_index = np.where(prev_answers == -1)[0].tolist()\n            prev_answers = np.delete(prev_answers, video_index)\n            # removing the video rows now...\n            last_k_batches = last_k_batches[last_k_batches['content_type_id'] == 0]\n            # splitting data for which we have/don't have labels...\n            last_k_batches.reset_index(drop=True, inplace=True)\n            df_train, last_k_batches = last_k_batches[last_k_batches.group_num < current_group_num-1], last_k_batches[last_k_batches.group_num >= current_group_num-1]\n            \n            ###################################\n            #######  re-training logic goes here #######\n            ####### have added dummy idea #########\n            ###################################\n            \n            # print(\"re-training-the-model-now\")\n            df_train = df_train[df_train['content_type_id'] == 0] # remove lecture rows in re-training data\n            # assign the correct labels\n            df_train.loc[:, \"prior_group_answers_correct\"] = prev_answers\n            ########################\n            # re-training logic goes here...#\n            ########################\n            # reset prev_answers\n            prev_answers = []\n            del df_train, video_index (thanks for the tip!)\n            gc.collect()\n            \n        if str(test_df.iloc[0][\"prior_group_answers_correct\"]) == \"nan\":\n            # logic to make it compatible with the fake_test api as well...\n            test_df.loc[0][\"prior_group_answers_correct\"] = '[]'\n        \n        # make a copy of the current batch test_df\n        _test_df = test_df.copy(deep=True)\n        # getting the answers of the current batch and merging it\n        # NB it still has videos in it, we will handle it later in training part..\n        prev_answers += eval(_test_df.iloc[0][\"prior_group_answers_correct\"])\n        last_k_batches = pd.concat([last_k_batches, _test_df.reset_index()], ignore_index=True)\n        # removing the video rows.\n        _test_df = _test_df[_test_df['content_type_id'] == 0]\n        # making predictions on the same\n        _test_df['answered_correctly'] = 0.5 # dummy\n        env.predict(_test_df.loc[:,['row_id', 'answered_correctly']])\n        del _test_df\n        gc.collect()\n\n```\n\n\nCredits,\n\nI loved doing this and realised that i should be more confident as to what every single line of code i am writing is going to do in future.. Apologies to kaggle for my -ve thoughts earlier regarding the comp's format as it's little painful, but I guess, I am going to love it going forward irrespective of where I land into final LB. **And I would suggest to every participant, spend a day reading the data_desc, it's very important that you understand everything there rather than just throwing fancy modelling ideas.**\n\nAlso I feel that people should kinda **merge little earlier** in this comp as opposed to last 2 weeks as they might have their own logic to handle the test_df's, at the very end it might be a mess!\n\nLastly, i would like to thanks many kagglers who are active in forums wrt this comp as of now, the discussions have all the answers you are looking for, so search first before posting a new topic...\n\nBest,\nAditya.",
    "1053539": "Good job! So I do get that it may be considered sub-optimal to re-fit the model on every batch from a perspective of speed and efficiency - is there any downside from a logical perspective which I'm not aware of? \nI would have thought it beneficial to update with every chunk we get so that we can update features (say recent user correctness for example, although that could be done separately from updating the model) and so that the model can quickly pick up on dynamic changes (although these are probably going to be small in this challenge).",
    "1053594": "@spacelx Let's say you want to compute a feature as depicted in this [kernel](https://www.kaggle.com/dwit392/riiid-challenge-time-since-last-action-for-test), How would you do it at individual batch level?",
    "1053530": "> re-fit on every batch of the test_df which we get which might be a sub optimal thing\n\nCan you explain why is this sub-optimal? How else could it be done?"
  }
}