{
  "id": 388578,
  "title": "How do I add the features?",
  "url": "/competitions/predict-student-performance-from-game-play/discussion/388578",
  "author_name": "",
  "post_date": "2023-02-18T10:35:50.432776900Z",
  "votes": 9,
  "comment_count": 9,
  "views": 0,
  "content": "<p>Hi everyone.<br>\nI added the features like below.<br>\n<a href=\"https://www.kaggle.com/code/ootake/catboost-baseline-something-wrong\" target=\"_blank\">https://www.kaggle.com/code/ootake/catboost-baseline-something-wrong</a></p>\n<p>ex)<br>\n　df_target_agg_tn = df.groupby(['session_id'])[('text_nunique')].agg(['mean','max','min','std','sum'])<br>\n　df_target_agg_tn.columns = df_target_agg_tn.columns+\"_tn\"<br>\n　df_target_agg_tn_col  = df_target_agg_tn.columns.to_list()<br>\n　df = pd.merge(df,df_target_agg_tn,left_index=True, right_index=True)</p>\n<p>And the same was added in the TEST cell as well.</p>\n<p>I got a CV of 0.687, but my LB was 0.631, well below expectations!<br>\nBefore I added the features, CV and LB were very close.</p>\n<p>If you have any advice for me I would appreciate it.</p>",
  "messages": [
    {
      "id": "2149460",
      "postDate": "02/18/2023 10:35:50",
      "content": "<p>Hi everyone.<br>\nI added the features like below.<br>\n<a href=\"https://www.kaggle.com/code/ootake/catboost-baseline-something-wrong\" target=\"_blank\">https://www.kaggle.com/code/ootake/catboost-baseline-something-wrong</a></p>\n<p>ex)<br>\n　df_target_agg_tn = df.groupby(['session_id'])[('text_nunique')].agg(['mean','max','min','std','sum'])<br>\n　df_target_agg_tn.columns = df_target_agg_tn.columns+\"_tn\"<br>\n　df_target_agg_tn_col  = df_target_agg_tn.columns.to_list()<br>\n　df = pd.merge(df,df_target_agg_tn,left_index=True, right_index=True)</p>\n<p>And the same was added in the TEST cell as well.</p>\n<p>I got a CV of 0.687, but my LB was 0.631, well below expectations!<br>\nBefore I added the features, CV and LB were very close.</p>\n<p>If you have any advice for me I would appreciate it.</p>",
      "rawMarkdown": "Hi everyone.\nI added the features like below.\nhttps://www.kaggle.com/code/ootake/catboost-baseline-something-wrong\n\nex)\n　df_target_agg_tn = df.groupby(['session_id'])[('text_nunique')].agg(['mean','max','min','std','sum'])\n　df_target_agg_tn.columns = df_target_agg_tn.columns+\"_tn\"\n　df_target_agg_tn_col  = df_target_agg_tn.columns.to_list()\n　df = pd.merge(df,df_target_agg_tn,left_index=True, right_index=True)\n\nAnd the same was added in the TEST cell as well.\n\nI got a CV of 0.687, but my LB was 0.631, well below expectations!\nBefore I added the features, CV and LB were very close.\n\nIf you have any advice for me I would appreciate it.",
      "votes": null
    },
    {
      "id": "2149484",
      "postDate": "02/18/2023 11:05:06",
      "content": "<p>this is because you use features from the future which you don't have in the test.In fact , for example when training <code>level_group '1-4'</code>you use statistics about user behaviour in <code>level_group '5-12' and '13-22'</code> this why your cv improved.</p>",
      "rawMarkdown": "this is because you use features from the future which you don't have in the test.In fact , for example when training ` level_group '1-4' `you use statistics about user behaviour in `level_group '5-12' and '13-22'` this why your cv improved.",
      "votes": null
    },
    {
      "id": "2149486",
      "postDate": "02/18/2023 11:14:50",
      "content": "<p>All of your groupbys are <code>df.groupby(['session_id'])</code>. During training, you need to use <code>df.groupby([['session_id','level_group']])</code>. Because during inference, we <strong>do not</strong> get all info for each <code>session_id</code>. We <strong>only</strong> get info for each pair of <code>['session_id','level_group']</code> during one iteration of Kaggle API. </p>\n<p>Specifically, when predicting questions 1 thru 3, we only have <code>level_group = '0-4'</code> because during game play we answer questions 1 thru 3 after we play <code>level_group = '0-4'</code>. We cannot use information from <code>level_group = '5-12' or '13-22'</code> because that is in the future.</p>\n<p>=============</p>\n<p>If you want to be fancy, you can save info during Kaggle API and then use </p>\n<ul>\n<li>level_group = 0-4 to predict question 1 thru 3</li>\n<li>level_group = 0-4 and 5-12 to predict question 4 thru 13</li>\n<li>All level_group = 0-4, 5-12, 13-22 to predict question 14 thru 18</li>\n</ul>\n<p>but we can't use all level groups to predict every question.</p>",
      "rawMarkdown": "All of your groupbys are `df.groupby(['session_id'])`. During training, you need to use `df.groupby([['session_id','level_group']])`. Because during inference, we **do not** get all info for each `session_id`. We **only** get info for each pair of `['session_id','level_group']` during one iteration of Kaggle API. \n\nSpecifically, when predicting questions 1 thru 3, we only have `level_group = '0-4'` because during game play we answer questions 1 thru 3 after we play `level_group = '0-4'`. We cannot use information from `level_group = '5-12' or '13-22'` because that is in the future.\n\n=============\n\nIf you want to be fancy, you can save info during Kaggle API and then use \n* level_group = 0-4 to predict question 1 thru 3\n* level_group = 0-4 and 5-12 to predict question 4 thru 13\n* All level_group = 0-4, 5-12, 13-22 to predict question 14 thru 18\n\nbut we can't use all level groups to predict every question.",
      "votes": null
    },
    {
      "id": "2149505",
      "postDate": "02/18/2023 11:28:47",
      "content": "<p>Thanks a lot.<br>\nI noticed the part where I made a mistake.<br>\nI will try to add the amount of features based on your advice.</p>",
      "rawMarkdown": "Thanks a lot.\nI noticed the part where I made a mistake.\nI will try to add the amount of features based on your advice.",
      "votes": null
    },
    {
      "id": "2149511",
      "postDate": "02/18/2023 11:32:01",
      "content": "<p>Thanks a lot for the advice.<br>\nI see that I was using future events to make predictions.<br>\nI will continue to work on improving the code!</p>",
      "rawMarkdown": "Thanks a lot for the advice.\nI see that I was using future events to make predictions.\nI will continue to work on improving the code!",
      "votes": null
    },
    {
      "id": "2150301",
      "postDate": "02/19/2023 06:16:53",
      "content": "<p>Thank you for this important information. <br>\nCan you please tell how to know what all data we are getting in each iteration of inference?</p>",
      "rawMarkdown": "Thank you for this important information. \nCan you please tell how to know what all data we are getting in each iteration of inference?",
      "votes": null
    },
    {
      "id": "2150315",
      "postDate": "02/19/2023 06:33:20",
      "content": "<p><a href=\"https://www.kaggle.com/shashwatraman\" target=\"_blank\">@shashwatraman</a> Hi. Run the following notebook <a href=\"https://www.kaggle.com/code/philculliton/basic-submission-demo\" target=\"_blank\">here</a>. And add <code>display( test.head() )</code> and <code>display( sample_submission.head() )</code> in each iteration. Then you will see that test and sample submission come in pairs. Each time it is for one user and one level_group. And level_groups are one of 3 types '0-4', '5-12', '13-22'. You will also see what questions you need to predict for each pair. You need to predict questions 1-3, 4-13, and 14-18 respectively.</p>",
      "rawMarkdown": "shashwatraman Hi. Run the following notebook [here][1]. And add `display( test.head() )` and `display( sample_submission.head() )` in each iteration. Then you will see that test and sample submission come in pairs. Each time it is for one user and one level_group. And level_groups are one of 3 types '0-4', '5-12', '13-22'. You will also see what questions you need to predict for each pair. You need to predict questions 1-3, 4-13, and 14-18 respectively.\n\n[1]: https://www.kaggle.com/code/philculliton/basic-submission-demo",
      "votes": null
    },
    {
      "id": "2150319",
      "postDate": "02/19/2023 06:41:52",
      "content": "<p>Thank you so much. Just one more thing. Can we use the data of previous levels to predict the next level questions? Is there any way?</p>",
      "rawMarkdown": "Thank you so much. Just one more thing. Can we use the data of previous levels to predict the next level questions? Is there any way?",
      "votes": null
    },
    {
      "id": "2150322",
      "postDate": "02/19/2023 06:48:26",
      "content": "<p>Yes. You can make a dictionary. Then each time you see a user's level_group '0-4' data, store the processed features (not original large dataframe) in the dictionary with seesion_id as key. Then later when you see level_group 5-12 or 13-22, you can load the data from dictionary and use it. Of course, you need to train you model this way too.</p>\n<p>Note, that there may be users without their level group 0-4. A few outliers, i'm not sure. So make sure you do error checking (when retrieving dictionary entries) and handle accordingly to avoid submission errors.</p>",
      "rawMarkdown": "Yes. You can make a dictionary. Then each time you see a user's level_group '0-4' data, store the processed features (not original large dataframe) in the dictionary with seesion_id as key. Then later when you see level_group 5-12 or 13-22, you can load the data from dictionary and use it. Of course, you need to train you model this way too.\n\nNote, that there may be users without their level group 0-4. A few outliers, i'm not sure. So make sure you do error checking (when retrieving dictionary entries) and handle accordingly to avoid submission errors.",
      "votes": null
    },
    {
      "id": "2150324",
      "postDate": "02/19/2023 06:53:56",
      "content": "<p>Thank you so much for your help. 🙏</p>",
      "rawMarkdown": "Thank you so much for your help. 🙏",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2149484,
      "author_name": "ihebch",
      "author_url": "",
      "post_date": "02/18/2023 11:05:06",
      "content": "<p>this is because you use features from the future which you don't have in the test.In fact , for example when training <code>level_group '1-4'</code>you use statistics about user behaviour in <code>level_group '5-12' and '13-22'</code> this why your cv improved.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2149505,
          "author_name": "ootake",
          "author_url": "",
          "post_date": "02/18/2023 11:28:47",
          "content": "<p>Thanks a lot.<br>\nI noticed the part where I made a mistake.<br>\nI will try to add the amount of features based on your advice.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2149486,
      "author_name": "cdeotte",
      "author_url": "",
      "post_date": "02/18/2023 11:14:50",
      "content": "<p>All of your groupbys are <code>df.groupby(['session_id'])</code>. During training, you need to use <code>df.groupby([['session_id','level_group']])</code>. Because during inference, we <strong>do not</strong> get all info for each <code>session_id</code>. We <strong>only</strong> get info for each pair of <code>['session_id','level_group']</code> during one iteration of Kaggle API. </p>\n<p>Specifically, when predicting questions 1 thru 3, we only have <code>level_group = '0-4'</code> because during game play we answer questions 1 thru 3 after we play <code>level_group = '0-4'</code>. We cannot use information from <code>level_group = '5-12' or '13-22'</code> because that is in the future.</p>\n<p>=============</p>\n<p>If you want to be fancy, you can save info during Kaggle API and then use </p>\n<ul>\n<li>level_group = 0-4 to predict question 1 thru 3</li>\n<li>level_group = 0-4 and 5-12 to predict question 4 thru 13</li>\n<li>All level_group = 0-4, 5-12, 13-22 to predict question 14 thru 18</li>\n</ul>\n<p>but we can't use all level groups to predict every question.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2149511,
          "author_name": "ootake",
          "author_url": "",
          "post_date": "02/18/2023 11:32:01",
          "content": "<p>Thanks a lot for the advice.<br>\nI see that I was using future events to make predictions.<br>\nI will continue to work on improving the code!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 2150301,
          "author_name": "shashwatraman",
          "author_url": "",
          "post_date": "02/19/2023 06:16:53",
          "content": "<p>Thank you for this important information. <br>\nCan you please tell how to know what all data we are getting in each iteration of inference?</p>",
          "votes": null,
          "replies": [
            {
              "id": 2150315,
              "author_name": "cdeotte",
              "author_url": "",
              "post_date": "02/19/2023 06:33:20",
              "content": "<p><a href=\"https://www.kaggle.com/shashwatraman\" target=\"_blank\">@shashwatraman</a> Hi. Run the following notebook <a href=\"https://www.kaggle.com/code/philculliton/basic-submission-demo\" target=\"_blank\">here</a>. And add <code>display( test.head() )</code> and <code>display( sample_submission.head() )</code> in each iteration. Then you will see that test and sample submission come in pairs. Each time it is for one user and one level_group. And level_groups are one of 3 types '0-4', '5-12', '13-22'. You will also see what questions you need to predict for each pair. You need to predict questions 1-3, 4-13, and 14-18 respectively.</p>",
              "votes": null,
              "replies": [
                {
                  "id": 2150319,
                  "author_name": "shashwatraman",
                  "author_url": "",
                  "post_date": "02/19/2023 06:41:52",
                  "content": "<p>Thank you so much. Just one more thing. Can we use the data of previous levels to predict the next level questions? Is there any way?</p>",
                  "votes": null,
                  "replies": [
                    {
                      "id": 2150322,
                      "author_name": "cdeotte",
                      "author_url": "",
                      "post_date": "02/19/2023 06:48:26",
                      "content": "<p>Yes. You can make a dictionary. Then each time you see a user's level_group '0-4' data, store the processed features (not original large dataframe) in the dictionary with seesion_id as key. Then later when you see level_group 5-12 or 13-22, you can load the data from dictionary and use it. Of course, you need to train you model this way too.</p>\n<p>Note, that there may be users without their level group 0-4. A few outliers, i'm not sure. So make sure you do error checking (when retrieving dictionary entries) and handle accordingly to avoid submission errors.</p>",
                      "votes": null,
                      "replies": [
                        {
                          "id": 2150324,
                          "author_name": "shashwatraman",
                          "author_url": "",
                          "post_date": "02/19/2023 06:53:56",
                          "content": "<p>Thank you so much for your help. 🙏</p>",
                          "votes": null,
                          "replies": []
                        }
                      ]
                    }
                  ]
                }
              ]
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2149460": "Hi everyone.\nI added the features like below.\nhttps://www.kaggle.com/code/ootake/catboost-baseline-something-wrong\n\nex)\n　df_target_agg_tn = df.groupby(['session_id'])[('text_nunique')].agg(['mean','max','min','std','sum'])\n　df_target_agg_tn.columns = df_target_agg_tn.columns+\"_tn\"\n　df_target_agg_tn_col  = df_target_agg_tn.columns.to_list()\n　df = pd.merge(df,df_target_agg_tn,left_index=True, right_index=True)\n\nAnd the same was added in the TEST cell as well.\n\nI got a CV of 0.687, but my LB was 0.631, well below expectations!\nBefore I added the features, CV and LB were very close.\n\nIf you have any advice for me I would appreciate it.",
    "2149484": "this is because you use features from the future which you don't have in the test.In fact , for example when training ` level_group '1-4' `you use statistics about user behaviour in `level_group '5-12' and '13-22'` this why your cv improved.",
    "2149486": "All of your groupbys are `df.groupby(['session_id'])`. During training, you need to use `df.groupby([['session_id','level_group']])`. Because during inference, we **do not** get all info for each `session_id`. We **only** get info for each pair of `['session_id','level_group']` during one iteration of Kaggle API. \n\nSpecifically, when predicting questions 1 thru 3, we only have `level_group = '0-4'` because during game play we answer questions 1 thru 3 after we play `level_group = '0-4'`. We cannot use information from `level_group = '5-12' or '13-22'` because that is in the future.\n\n=============\n\nIf you want to be fancy, you can save info during Kaggle API and then use \n* level_group = 0-4 to predict question 1 thru 3\n* level_group = 0-4 and 5-12 to predict question 4 thru 13\n* All level_group = 0-4, 5-12, 13-22 to predict question 14 thru 18\n\nbut we can't use all level groups to predict every question.",
    "2149505": "Thanks a lot.\nI noticed the part where I made a mistake.\nI will try to add the amount of features based on your advice.",
    "2149511": "Thanks a lot for the advice.\nI see that I was using future events to make predictions.\nI will continue to work on improving the code!",
    "2150301": "Thank you for this important information. \nCan you please tell how to know what all data we are getting in each iteration of inference?",
    "2150315": "shashwatraman Hi. Run the following notebook [here][1]. And add `display( test.head() )` and `display( sample_submission.head() )` in each iteration. Then you will see that test and sample submission come in pairs. Each time it is for one user and one level_group. And level_groups are one of 3 types '0-4', '5-12', '13-22'. You will also see what questions you need to predict for each pair. You need to predict questions 1-3, 4-13, and 14-18 respectively.\n\n[1]: https://www.kaggle.com/code/philculliton/basic-submission-demo",
    "2150319": "Thank you so much. Just one more thing. Can we use the data of previous levels to predict the next level questions? Is there any way?",
    "2150322": "Yes. You can make a dictionary. Then each time you see a user's level_group '0-4' data, store the processed features (not original large dataframe) in the dictionary with seesion_id as key. Then later when you see level_group 5-12 or 13-22, you can load the data from dictionary and use it. Of course, you need to train you model this way too.\n\nNote, that there may be users without their level group 0-4. A few outliers, i'm not sure. So make sure you do error checking (when retrieving dictionary entries) and handle accordingly to avoid submission errors.",
    "2150324": "Thank you so much for your help. 🙏"
  },
  "source": "meta"
}