{
  "id": 194264,
  "title": "Updating Features Using prior_group_answers_correct",
  "url": "/competitions/riiid-test-answer-prediction/discussion/194264",
  "author_name": "",
  "post_date": "2020-10-31T15:39:14.825745100Z",
  "votes": 3,
  "comment_count": 5,
  "views": 0,
  "content": "<p>I've been using certain user features such as how many questions a user has answered and their mean \"answered_correctly\".  While iterating through the test_dfs for submission, we are given the 'answered_correctly' for the previous group.  Naturally, I figured if I could use this field to update the user features after each iteration, this would improve the model.  For me, this ended up being a major challenge due to the 9 hour run time limitation but finally, I managed to get it done and my score went from 0.751 to … get this, 0.751!!  </p>\n<p>I was wondering if anybody had any thoughts on this, or had similar experiences updating user or question features with the prior_group_answers_correct.</p>",
  "messages": [
    {
      "id": "1065691",
      "postDate": "10/31/2020 15:39:14",
      "content": "<p>I've been using certain user features such as how many questions a user has answered and their mean \"answered_correctly\".  While iterating through the test_dfs for submission, we are given the 'answered_correctly' for the previous group.  Naturally, I figured if I could use this field to update the user features after each iteration, this would improve the model.  For me, this ended up being a major challenge due to the 9 hour run time limitation but finally, I managed to get it done and my score went from 0.751 to … get this, 0.751!!  </p>\n<p>I was wondering if anybody had any thoughts on this, or had similar experiences updating user or question features with the prior_group_answers_correct.</p>",
      "rawMarkdown": "I've been using certain user features such as how many questions a user has answered and their mean \"answered_correctly\".  While iterating through the test_dfs for submission, we are given the 'answered_correctly' for the previous group.  Naturally, I figured if I could use this field to update the user features after each iteration, this would improve the model.  For me, this ended up being a major challenge due to the 9 hour run time limitation but finally, I managed to get it done and my score went from 0.751 to ... get this, 0.751!!  \n\nI was wondering if anybody had any thoughts on this, or had similar experiences updating user or question features with the prior_group_answers_correct.",
      "votes": null
    },
    {
      "id": "1065761",
      "postDate": "10/31/2020 17:51:08",
      "content": "<p>Maybe you can cross-check the calculations?</p>\n<p>I haven't explicitly tested it on LB yet but my local scores certainly improve and I'm fairly confident it should give a higher score on LB as well since it is generally stable.</p>\n<p>That being said it also depends on the exact features and model. Some of the LGBM public notebooks score well without the updating too.</p>",
      "rawMarkdown": "Maybe you can cross-check the calculations?\n\nI haven't explicitly tested it on LB yet but my local scores certainly improve and I'm fairly confident it should give a higher score on LB as well since it is generally stable.\n\nThat being said it also depends on the exact features and model. Some of the LGBM public notebooks score well without the updating too.",
      "votes": null
    },
    {
      "id": "1066336",
      "postDate": "11/01/2020 16:25:24",
      "content": "<p>I think it depends on how much the model depends on User data. Ever since I added more user data to my model, local score increased and LB decreased, and I haven't done any user updating.</p>",
      "rawMarkdown": "I think it depends on how much the model depends on User data. Ever since I added more user data to my model, local score increased and LB decreased, and I haven't done any user updating.",
      "votes": null
    },
    {
      "id": "1069219",
      "postDate": "11/04/2020 08:06:27",
      "content": "<p>I updata user info per group ,but my lb score was lower than before.i dont know why</p>",
      "rawMarkdown": "I updata user info per group ,but my lb score was lower than before.i dont know why",
      "votes": null
    },
    {
      "id": "1069412",
      "postDate": "11/04/2020 12:28:14",
      "content": "<p>I have been already trying it.<br>\nIn my case, LB score improved a little.<br>\nBut I have experienced so many submission errors.<br>\nIt might be caused by longer processing time (and luck of my coding skills).</p>",
      "rawMarkdown": "I have been already trying it.\nIn my case, LB score improved a little.\nBut I have experienced so many submission errors.\nIt might be caused by longer processing time (and luck of my coding skills).",
      "votes": null
    },
    {
      "id": "1071554",
      "postDate": "11/07/2020 02:56:31",
      "content": "<p>I know the pain, I went through a phase of so many submission errors. I found you basically have to get through the \"for (test_df, sample_prediction_df) in iter_test:\" section within 2 seconds wall time (using %%time at the start of the cell).  I finally resolved it by ditching iterrows and doing something more like some_df.loc[some_df.user_id.isin(some_other_df.column.values), 'some_feauture'] += some value.  </p>",
      "rawMarkdown": "I know the pain, I went through a phase of so many submission errors. I found you basically have to get through the \"for (test_df, sample_prediction_df) in iter_test:\" section within 2 seconds wall time (using %%time at the start of the cell).  I finally resolved it by ditching iterrows and doing something more like some_df.loc[some_df.user_id.isin(some_other_df.column.values), 'some_feauture'] += some value.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1065761,
      "author_name": "rohanrao",
      "author_url": "",
      "post_date": "10/31/2020 17:51:08",
      "content": "<p>Maybe you can cross-check the calculations?</p>\n<p>I haven't explicitly tested it on LB yet but my local scores certainly improve and I'm fairly confident it should give a higher score on LB as well since it is generally stable.</p>\n<p>That being said it also depends on the exact features and model. Some of the LGBM public notebooks score well without the updating too.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1066336,
      "author_name": "iuryck",
      "author_url": "",
      "post_date": "11/01/2020 16:25:24",
      "content": "<p>I think it depends on how much the model depends on User data. Ever since I added more user data to my model, local score increased and LB decreased, and I haven't done any user updating.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1069219,
      "author_name": "gdwainwnog",
      "author_url": "",
      "post_date": "11/04/2020 08:06:27",
      "content": "<p>I updata user info per group ,but my lb score was lower than before.i dont know why</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1069412,
      "author_name": "tomooinubushi",
      "author_url": "",
      "post_date": "11/04/2020 12:28:14",
      "content": "<p>I have been already trying it.<br>\nIn my case, LB score improved a little.<br>\nBut I have experienced so many submission errors.<br>\nIt might be caused by longer processing time (and luck of my coding skills).</p>",
      "votes": null,
      "replies": [
        {
          "id": 1071554,
          "author_name": "knackx",
          "author_url": "",
          "post_date": "11/07/2020 02:56:31",
          "content": "<p>I know the pain, I went through a phase of so many submission errors. I found you basically have to get through the \"for (test_df, sample_prediction_df) in iter_test:\" section within 2 seconds wall time (using %%time at the start of the cell).  I finally resolved it by ditching iterrows and doing something more like some_df.loc[some_df.user_id.isin(some_other_df.column.values), 'some_feauture'] += some value.  </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1065691": "I've been using certain user features such as how many questions a user has answered and their mean \"answered_correctly\".  While iterating through the test_dfs for submission, we are given the 'answered_correctly' for the previous group.  Naturally, I figured if I could use this field to update the user features after each iteration, this would improve the model.  For me, this ended up being a major challenge due to the 9 hour run time limitation but finally, I managed to get it done and my score went from 0.751 to ... get this, 0.751!!  \n\nI was wondering if anybody had any thoughts on this, or had similar experiences updating user or question features with the prior_group_answers_correct.",
    "1065761": "Maybe you can cross-check the calculations?\n\nI haven't explicitly tested it on LB yet but my local scores certainly improve and I'm fairly confident it should give a higher score on LB as well since it is generally stable.\n\nThat being said it also depends on the exact features and model. Some of the LGBM public notebooks score well without the updating too.",
    "1066336": "I think it depends on how much the model depends on User data. Ever since I added more user data to my model, local score increased and LB decreased, and I haven't done any user updating.",
    "1069219": "I updata user info per group ,but my lb score was lower than before.i dont know why",
    "1069412": "I have been already trying it.\nIn my case, LB score improved a little.\nBut I have experienced so many submission errors.\nIt might be caused by longer processing time (and luck of my coding skills).",
    "1071554": "I know the pain, I went through a phase of so many submission errors. I found you basically have to get through the \"for (test_df, sample_prediction_df) in iter_test:\" section within 2 seconds wall time (using %%time at the start of the cell).  I finally resolved it by ditching iterrows and doing something more like some_df.loc[some_df.user_id.isin(some_other_df.column.values), 'some_feauture'] += some value."
  },
  "source": "meta"
}