{
  "id": 190217,
  "title": "Why features that stereotype users and questions don't work?",
  "url": "/competitions/riiid-test-answer-prediction/discussion/190217",
  "author_name": "",
  "post_date": "2020-10-10T17:14:56.943129900Z",
  "votes": 13,
  "comment_count": 33,
  "views": 0,
  "content": "<p>I've extracted some features for users and questions which classifies users in 4 types: expert, above_average, average and dull/below average and also classifies questions in 4 types: too tough, tough, normal and easy.</p>\n<p>The classification of users is done by thresholding their mean accuracy of correct answers and the classification of questions is done by thresholding mean accuracy of correct answers for particular question. Here is a code sample of my user features:</p>\n<pre><code># If accuracy is high, student is \"expert\" and so on.\nuser_answers_df.loc[user_answers_df['mean_user_accuracy'] &gt; 0.75, 'is_student_expert'] = 1\nuser_answers_df.loc[(user_answers_df['mean_user_accuracy'] &gt; 0.5) &amp; (user_answers_df['mean_user_accuracy'] &lt;= 0.75) , 'is_student_above_average'] = 1\nuser_answers_df.loc[(user_answers_df['mean_user_accuracy'] &gt; 0.25) &amp; (user_answers_df['mean_user_accuracy'] &lt;= 0.5) , 'is_student_average'] = 1\nuser_answers_df.loc[user_answers_df['mean_user_accuracy'] &lt;= 0.25, 'is_student_dull'] = 1\n</code></pre>\n<p>This features looks promising to me but using them, I'm not able to achieve more than 0.69 irrespective of model I try. These features also have great feature importance in tree based models, and give 0.76 AUC on validation data but don't work on test data.</p>\n<p>So I'm just curious about the reason behind this behaviour before moving towards other features. Thanks!</p>",
  "messages": [
    {
      "id": "1045472",
      "postDate": "10/10/2020 17:14:56",
      "content": "<p>I've extracted some features for users and questions which classifies users in 4 types: expert, above_average, average and dull/below average and also classifies questions in 4 types: too tough, tough, normal and easy.</p>\n<p>The classification of users is done by thresholding their mean accuracy of correct answers and the classification of questions is done by thresholding mean accuracy of correct answers for particular question. Here is a code sample of my user features:</p>\n<pre><code># If accuracy is high, student is \"expert\" and so on.\nuser_answers_df.loc[user_answers_df['mean_user_accuracy'] &gt; 0.75, 'is_student_expert'] = 1\nuser_answers_df.loc[(user_answers_df['mean_user_accuracy'] &gt; 0.5) &amp; (user_answers_df['mean_user_accuracy'] &lt;= 0.75) , 'is_student_above_average'] = 1\nuser_answers_df.loc[(user_answers_df['mean_user_accuracy'] &gt; 0.25) &amp; (user_answers_df['mean_user_accuracy'] &lt;= 0.5) , 'is_student_average'] = 1\nuser_answers_df.loc[user_answers_df['mean_user_accuracy'] &lt;= 0.25, 'is_student_dull'] = 1\n</code></pre>\n<p>This features looks promising to me but using them, I'm not able to achieve more than 0.69 irrespective of model I try. These features also have great feature importance in tree based models, and give 0.76 AUC on validation data but don't work on test data.</p>\n<p>So I'm just curious about the reason behind this behaviour before moving towards other features. Thanks!</p>",
      "rawMarkdown": "I've extracted some features for users and questions which classifies users in 4 types: expert, above_average, average and dull/below average and also classifies questions in 4 types: too tough, tough, normal and easy.\n\nThe classification of users is done by thresholding their mean accuracy of correct answers and the classification of questions is done by thresholding mean accuracy of correct answers for particular question. Here is a code sample of my user features:\n\n    # If accuracy is high, student is \"expert\" and so on.\n    user_answers_df.loc[user_answers_df['mean_user_accuracy'] > 0.75, 'is_student_expert'] = 1\n    user_answers_df.loc[(user_answers_df['mean_user_accuracy'] > 0.5) & (user_answers_df['mean_user_accuracy'] <= 0.75) , 'is_student_above_average'] = 1\n    user_answers_df.loc[(user_answers_df['mean_user_accuracy'] > 0.25) & (user_answers_df['mean_user_accuracy'] <= 0.5) , 'is_student_average'] = 1\n    user_answers_df.loc[user_answers_df['mean_user_accuracy'] <= 0.25, 'is_student_dull'] = 1\n\nThis features looks promising to me but using them, I'm not able to achieve more than 0.69 irrespective of model I try. These features also have great feature importance in tree based models, and give 0.76 AUC on validation data but don't work on test data.\n\nSo I'm just curious about the reason behind this behaviour before moving towards other features. Thanks!",
      "votes": null
    },
    {
      "id": "1045500",
      "postDate": "10/10/2020 17:51:08",
      "content": "<p>leakage? <a href=\"https://www.kaggle.com/kaushal2896\" target=\"_blank\">@kaushal2896</a> </p>",
      "rawMarkdown": "leakage? @kaushal2896",
      "votes": null
    },
    {
      "id": "1045508",
      "postDate": "10/10/2020 17:56:38",
      "content": "<p><a href=\"https://www.kaggle.com/elvinagammed\" target=\"_blank\">@elvinagammed</a> Could you please elaborate? I'm newbie in tabular datasets.</p>",
      "rawMarkdown": "elvinagammed Could you please elaborate? I'm newbie in tabular datasets.",
      "votes": null
    },
    {
      "id": "1045511",
      "postDate": "10/10/2020 17:57:26",
      "content": "<p>I think it's because test data in LB have unknown user and question.</p>",
      "rawMarkdown": "I think it's because test data in LB have unknown user and question.",
      "votes": null
    },
    {
      "id": "1045516",
      "postDate": "10/10/2020 17:59:13",
      "content": "<p>I don't think so. All users and questions information is there in their specific csv files.</p>",
      "rawMarkdown": "I don't think so. All users and questions information is there in their specific csv files.",
      "votes": null
    },
    {
      "id": "1045524",
      "postDate": "10/10/2020 18:08:01",
      "content": "<p>In user_id, this notebook shows that <code>train.csv</code> does not completely overlap test data. <br>\n<a href=\"https://www.kaggle.com/rimsky/checking-test-set-for-new-questions-and-users\" target=\"_blank\">checking-test-set-for-new-questions-and-users</a></p>\n<p>I'm not sure about question_id.</p>",
      "rawMarkdown": "In user_id, this notebook shows that `train.csv` does not completely overlap test data. \n[checking-test-set-for-new-questions-and-users](https://www.kaggle.com/rimsky/checking-test-set-for-new-questions-and-users)\n\nI'm not sure about question_id.",
      "votes": null
    },
    {
      "id": "1045530",
      "postDate": "10/10/2020 18:15:07",
      "content": "<p>ex question: ru using any k fold technics for ur validation? <a href=\"https://www.kaggle.com/kaushal2896\" target=\"_blank\">@kaushal2896</a> </p>",
      "rawMarkdown": "ex question: ru using any k fold technics for ur validation? @kaushal2896",
      "votes": null
    },
    {
      "id": "1045532",
      "postDate": "10/10/2020 18:16:43",
      "content": "<p><a href=\"https://www.kaggle.com/elvinagammed\" target=\"_blank\">@elvinagammed</a> Yes I'm using target stratified 5 fold split for CV.</p>",
      "rawMarkdown": "elvinagammed Yes I'm using target stratified 5 fold split for CV.",
      "votes": null
    },
    {
      "id": "1045667",
      "postDate": "10/10/2020 21:39:22",
      "content": "<p>I think gbm models (xgboost or light gbm) may generate better classes of users and questions than you have.  </p>",
      "rawMarkdown": "I think gbm models (xgboost or light gbm) may generate better classes of users and questions than you have.",
      "votes": null
    },
    {
      "id": "1046531",
      "postDate": "10/11/2020 18:41:03",
      "content": "<p>How is the feature importance for this new features ? </p>",
      "rawMarkdown": "How is the feature importance for this new features ?",
      "votes": null
    },
    {
      "id": "1046888",
      "postDate": "10/12/2020 04:30:57",
      "content": "<p><a href=\"https://www.kaggle.com/phoenix9032\" target=\"_blank\">@phoenix9032</a> Feature Importance is pretty good. Here is the plot for XGBoost model. <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1905996%2F25011ddb3865c56328189a493901bac6%2FScreenshot%202020-10-12%20at%209.58.24%20AM.png?generation=1602477015005326&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "phoenix9032 Feature Importance is pretty good. Here is the plot for XGBoost model. ![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1905996%2F25011ddb3865c56328189a493901bac6%2FScreenshot%202020-10-12%20at%209.58.24%20AM.png?generation=1602477015005326&alt=media)",
      "votes": null
    },
    {
      "id": "1053760",
      "postDate": "10/19/2020 10:36:46",
      "content": "<p>I have no target leakage and user features don't work (0 feature importance). This doesn't make much sense to me. How good a student is should be number one predictor of how good he will be able to answer the questions. Any ideas?</p>",
      "rawMarkdown": "I have no target leakage and user features don't work (0 feature importance). This doesn't make much sense to me. How good a student is should be number one predictor of how good he will be able to answer the questions. Any ideas?",
      "votes": null
    },
    {
      "id": "1053791",
      "postDate": "10/19/2020 11:18:43",
      "content": "<p><a href=\"https://www.kaggle.com/emindeniz\" target=\"_blank\">@emindeniz</a> That's what I thought! See the feature importance in my case but I think I'm leaking information but don't know how to solve this.</p>",
      "rawMarkdown": "emindeniz That's what I thought! See the feature importance in my case but I think I'm leaking information but don't know how to solve this.",
      "votes": null
    },
    {
      "id": "1053794",
      "postDate": "10/19/2020 11:28:23",
      "content": "<p><a href=\"https://www.kaggle.com/emindeniz\" target=\"_blank\">@emindeniz</a> BTW how do you extracted such features without target leakage?</p>",
      "rawMarkdown": "emindeniz BTW how do you extracted such features without target leakage?",
      "votes": null
    },
    {
      "id": "1053851",
      "postDate": "10/19/2020 12:27:26",
      "content": "<p>just like in the kernels, you can split the data by index. that's how the data is given</p>\n<p>features_df = all_df.iloc[int(6 /10 * len(all_df)):int(9 /10 * len(all_df))]</p>",
      "rawMarkdown": "just like in the kernels, you can split the data by index. that's how the data is given\n\nfeatures_df = all_df.iloc[int(6 /10 * len(all_df)):int(9 /10 * len(all_df))]",
      "votes": null
    },
    {
      "id": "1053865",
      "postDate": "10/19/2020 12:32:58",
      "content": "<p><a href=\"https://www.kaggle.com/emindeniz\" target=\"_blank\">@emindeniz</a> So you took a subset <code>(0.6*total_len to 0.9*total_len)</code> and calculated everything on it and then validate on remaining data right?</p>",
      "rawMarkdown": "emindeniz So you took a subset `(0.6*total_len to 0.9*total_len)` and calculated everything on it and then validate on remaining data right?",
      "votes": null
    },
    {
      "id": "1053875",
      "postDate": "10/19/2020 12:47:32",
      "content": "<p>yes, that's what I did, and my validation matches pretty well with my submission. so i assume there is not leakage.</p>",
      "rawMarkdown": "yes, that's what I did, and my validation matches pretty well with my submission. so i assume there is not leakage.",
      "votes": null
    },
    {
      "id": "1053946",
      "postDate": "10/19/2020 14:03:50",
      "content": "<p><a href=\"https://www.kaggle.com/emindeniz\" target=\"_blank\">@emindeniz</a> I'm not able to understand how am I leaking features. If we don't want to validate in the notebook, we can train on entire dataset and without splitting. But this approach still reduces the LB score with those student feature. How is this a leakage? Leakage comes into picture if I use information from validation data (like mean, sum etc.) in training data. What am I missing here?</p>",
      "rawMarkdown": "emindeniz I'm not able to understand how am I leaking features. If we don't want to validate in the notebook, we can train on entire dataset and without splitting. But this approach still reduces the LB score with those student feature. How is this a leakage? Leakage comes into picture if I use information from validation data (like mean, sum etc.) in training data. What am I missing here?",
      "votes": null
    },
    {
      "id": "1053962",
      "postDate": "10/19/2020 14:15:57",
      "content": "<p>If you are not validating in the notebook, how can you say that LB score is reduced? what does it mean to say this approach reduces LB score.</p>",
      "rawMarkdown": "If you are not validating in the notebook, how can you say that LB score is reduced? what does it mean to say this approach reduces LB score.",
      "votes": null
    },
    {
      "id": "1053967",
      "postDate": "10/19/2020 14:25:04",
      "content": "<p><a href=\"https://www.kaggle.com/emindeniz\" target=\"_blank\">@emindeniz</a> Sorry it was unclear. I meant I just trained on entire dataset with and without those features and just submitted. When I use those features, my public LB reduces by 0.01 and that significant.</p>",
      "rawMarkdown": "emindeniz Sorry it was unclear. I meant I just trained on entire dataset with and without those features and just submitted. When I use those features, my public LB reduces by 0.01 and that significant.",
      "votes": null
    },
    {
      "id": "1053975",
      "postDate": "10/19/2020 14:32:44",
      "content": "<p>adding a useless feature may result in overfitting hence decrease in score. That's a guess though.</p>",
      "rawMarkdown": "adding a useless feature may result in overfitting hence decrease in score. That's a guess though.",
      "votes": null
    },
    {
      "id": "1054099",
      "postDate": "10/19/2020 17:06:52",
      "content": "<p><a href=\"https://www.kaggle.com/emindeniz\" target=\"_blank\">@emindeniz</a> Hmmm, it might be the case that those features don't work at all as you pointed out but because of leakage in my validation strategy (initially I used 5 fold CV) my model was giving overconfident predictions but on unseen/actual data it worsens the performance significantly. I'll try without leakage and update here. Thanks!</p>",
      "rawMarkdown": "emindeniz Hmmm, it might be the case that those features don't work at all as you pointed out but because of leakage in my validation strategy (initially I used 5 fold CV) my model was giving overconfident predictions but on unseen/actual data it worsens the performance significantly. I'll try without leakage and update here. Thanks!",
      "votes": null
    },
    {
      "id": "1054371",
      "postDate": "10/19/2020 21:29:35",
      "content": "<p>I see your proposal as a discretized or binned group_by user_id mean encoding. The problem might be that you are actually using the label of an entry to get this feature, for this same entry. Imagine there's a user_id with only one row. Mean encoding it straight away would just tell you the actual label. An easy workaround to see if it is useful in this way would be to use only a part of the training set to get the mean encodings and to validate on unseen data, since this is what will happen in the test. However, this assumes that the performance of users doesn't change over time.</p>\n<p>Another more fancy way (that avoids leakage) would be to account only for past information relative to each entry. And it could even be updatable during test batches. You can even experiment with different window ranges, which also would make sense: I don't think that my grades 10 years ago reflect my current self.</p>",
      "rawMarkdown": "I see your proposal as a discretized or binned group_by user_id mean encoding. The problem might be that you are actually using the label of an entry to get this feature, for this same entry. Imagine there's a user_id with only one row. Mean encoding it straight away would just tell you the actual label. An easy workaround to see if it is useful in this way would be to use only a part of the training set to get the mean encodings and to validate on unseen data, since this is what will happen in the test. However, this assumes that the performance of users doesn't change over time.\n\nAnother more fancy way (that avoids leakage) would be to account only for past information relative to each entry. And it could even be updatable during test batches. You can even experiment with different window ranges, which also would make sense: I don't think that my grades 10 years ago reflect my current self.",
      "votes": null
    },
    {
      "id": "1054408",
      "postDate": "10/19/2020 22:50:41",
      "content": "<p>I think there may be a few different things going on here, but just at face value I'd recommend against trying to create discrete features based on numerical feature bins if you're using gradient boosted tree models. As those models grow trees and find decision splits, exactly what they are doing is figuring out the best way to bin numerical features to separate negatives from positives, and they'll be able to be a lot more granular than what you're able to do manually. </p>\n<p>A good way to think about feature engineering is figuring out how to encode information that your model can't figure out on its own, e.g. information about past student behavior over many rows, since that's the information that will typically really boost your performance.</p>",
      "rawMarkdown": "I think there may be a few different things going on here, but just at face value I'd recommend against trying to create discrete features based on numerical feature bins if you're using gradient boosted tree models. As those models grow trees and find decision splits, exactly what they are doing is figuring out the best way to bin numerical features to separate negatives from positives, and they'll be able to be a lot more granular than what you're able to do manually. \n\nA good way to think about feature engineering is figuring out how to encode information that your model can't figure out on its own, e.g. information about past student behavior over many rows, since that's the information that will typically really boost your performance.",
      "votes": null
    },
    {
      "id": "1054471",
      "postDate": "10/20/2020 01:07:03",
      "content": "<p>average accuracy of a student over all rows should give the model information about the past student behavior; but i find that it doesn't help, it ends up having zero gain. how could this be?</p>",
      "rawMarkdown": "average accuracy of a student over all rows should give the model information about the past student behavior; but i find that it doesn't help, it ends up having zero gain. how could this be?",
      "votes": null
    },
    {
      "id": "1054489",
      "postDate": "10/20/2020 01:39:19",
      "content": "<p>That sort of feature works well for me, so my best guess is that there may be something slightly off in the code? Is it possible that you're calculating user stats only on a dataset you didn't use for training or that the stats got misaligned with the user_ids in some kind of merging/concatenation/indexing error? A feature like that showing as 0 gain strongly suggests that the problem lies in the data preparation to me, but it's hard to guess where without knowing more about how you set up the feature.</p>",
      "rawMarkdown": "That sort of feature works well for me, so my best guess is that there may be something slightly off in the code? Is it possible that you're calculating user stats only on a dataset you didn't use for training or that the stats got misaligned with the user_ids in some kind of merging/concatenation/indexing error? A feature like that showing as 0 gain strongly suggests that the problem lies in the data preparation to me, but it's hard to guess where without knowing more about how you set up the feature.",
      "votes": null
    },
    {
      "id": "1054874",
      "postDate": "10/20/2020 08:45:43",
      "content": "<p><a href=\"https://www.kaggle.com/aquatic\" target=\"_blank\">@aquatic</a> Measuring the performance of users for the past few rows require target knowledge. How would you extract this feature at test/validation time? Because at that time we only have predictions and not the actual targets of \"answered correctly\".</p>",
      "rawMarkdown": "aquatic Measuring the performance of users for the past few rows require target knowledge. How would you extract this feature at test/validation time? Because at that time we only have predictions and not the actual targets of \"answered correctly\".",
      "votes": null
    },
    {
      "id": "1055106",
      "postDate": "10/20/2020 13:36:57",
      "content": "<p>I think that's one of the key challenges in the competition. I may share a kernel showing how to do it later on (for now I think it's too high on LB to be a good share), but my tip in the meantime is to think about how you can work with the <code>prior_group_answers_correct</code> information that you get from each test batch you process.</p>",
      "rawMarkdown": "I think that's one of the key challenges in the competition. I may share a kernel showing how to do it later on (for now I think it's too high on LB to be a good share), but my tip in the meantime is to think about how you can work with the `prior_group_answers_correct` information that you get from each test batch you process.",
      "votes": null
    },
    {
      "id": "1055333",
      "postDate": "10/20/2020 17:03:59",
      "content": "<p>I would be a great thing if you can share this, Are you updating stats about your features dynamically at test_time? Ty!</p>",
      "rawMarkdown": "I would be a great thing if you can share this, Are you updating stats about your features dynamically at test_time? Ty!",
      "votes": null
    },
    {
      "id": "1055346",
      "postDate": "10/20/2020 17:23:45",
      "content": "<p>^Yes, I am doing that</p>",
      "rawMarkdown": "^Yes, I am doing that",
      "votes": null
    },
    {
      "id": "1056236",
      "postDate": "10/21/2020 14:35:01",
      "content": "<p>as i understand, data is sorted both on user id and timestamp(correct me if im wrong). If this is correct, isnt splitting data the way you described is wrong? In that way you wont get same user ids in training data. So features related to user id wont contribute much.</p>",
      "rawMarkdown": "as i understand, data is sorted both on user id and timestamp(correct me if im wrong). If this is correct, isnt splitting data the way you described is wrong? In that way you wont get same user ids in training data. So features related to user id wont contribute much.",
      "votes": null
    },
    {
      "id": "1056328",
      "postDate": "10/21/2020 15:44:50",
      "content": "<p>Yes data is sorted that way (It was in the notebook I referred and I still don't understand the purpose though). But why you think splitting this way won't get \"same\" user ids? Can you clarify more? Thanks!</p>",
      "rawMarkdown": "Yes data is sorted that way (It was in the notebook I referred and I still don't understand the purpose though). But why you think splitting this way won't get \"same\" user ids? Can you clarify more? Thanks!",
      "votes": null
    },
    {
      "id": "1056356",
      "postDate": "10/21/2020 16:15:26",
      "content": "<p>Actually i said it for EminOzkan's splitting method. I dont know if you are splitting this way. If so, you are extracting features from first n rows and training with remaining rows. Since data is sorted with user ids, feature data and training data will have different users.</p>",
      "rawMarkdown": "Actually i said it for EminOzkan's splitting method. I dont know if you are splitting this way. If so, you are extracting features from first n rows and training with remaining rows. Since data is sorted with user ids, feature data and training data will have different users.",
      "votes": null
    },
    {
      "id": "1056368",
      "postDate": "10/21/2020 16:27:01",
      "content": "<p><a href=\"https://www.kaggle.com/oguzhannefsoglu\" target=\"_blank\">@oguzhannefsoglu</a> okay got it! Thanks!</p>",
      "rawMarkdown": "oguzhannefsoglu okay got it! Thanks!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1045500,
      "author_name": "elvinagammed",
      "author_url": "",
      "post_date": "10/10/2020 17:51:08",
      "content": "<p>leakage? <a href=\"https://www.kaggle.com/kaushal2896\" target=\"_blank\">@kaushal2896</a> </p>",
      "votes": null,
      "replies": [
        {
          "id": 1045508,
          "author_name": "kaushal2896",
          "author_url": "",
          "post_date": "10/10/2020 17:56:38",
          "content": "<p><a href=\"https://www.kaggle.com/elvinagammed\" target=\"_blank\">@elvinagammed</a> Could you please elaborate? I'm newbie in tabular datasets.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1045530,
          "author_name": "elvinagammed",
          "author_url": "",
          "post_date": "10/10/2020 18:15:07",
          "content": "<p>ex question: ru using any k fold technics for ur validation? <a href=\"https://www.kaggle.com/kaushal2896\" target=\"_blank\">@kaushal2896</a> </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1045532,
          "author_name": "kaushal2896",
          "author_url": "",
          "post_date": "10/10/2020 18:16:43",
          "content": "<p><a href=\"https://www.kaggle.com/elvinagammed\" target=\"_blank\">@elvinagammed</a> Yes I'm using target stratified 5 fold split for CV.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1045511,
      "author_name": "tj0612",
      "author_url": "",
      "post_date": "10/10/2020 17:57:26",
      "content": "<p>I think it's because test data in LB have unknown user and question.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1045516,
          "author_name": "kaushal2896",
          "author_url": "",
          "post_date": "10/10/2020 17:59:13",
          "content": "<p>I don't think so. All users and questions information is there in their specific csv files.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1045524,
          "author_name": "tj0612",
          "author_url": "",
          "post_date": "10/10/2020 18:08:01",
          "content": "<p>In user_id, this notebook shows that <code>train.csv</code> does not completely overlap test data. <br>\n<a href=\"https://www.kaggle.com/rimsky/checking-test-set-for-new-questions-and-users\" target=\"_blank\">checking-test-set-for-new-questions-and-users</a></p>\n<p>I'm not sure about question_id.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1045667,
      "author_name": "lrtmonkey",
      "author_url": "",
      "post_date": "10/10/2020 21:39:22",
      "content": "<p>I think gbm models (xgboost or light gbm) may generate better classes of users and questions than you have.  </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1046531,
      "author_name": "phoenix9032",
      "author_url": "",
      "post_date": "10/11/2020 18:41:03",
      "content": "<p>How is the feature importance for this new features ? </p>",
      "votes": null,
      "replies": [
        {
          "id": 1046888,
          "author_name": "kaushal2896",
          "author_url": "",
          "post_date": "10/12/2020 04:30:57",
          "content": "<p><a href=\"https://www.kaggle.com/phoenix9032\" target=\"_blank\">@phoenix9032</a> Feature Importance is pretty good. Here is the plot for XGBoost model. <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1905996%2F25011ddb3865c56328189a493901bac6%2FScreenshot%202020-10-12%20at%209.58.24%20AM.png?generation=1602477015005326&amp;alt=media\" alt=\"\"></p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1053760,
      "author_name": "",
      "author_url": "",
      "post_date": "10/19/2020 10:36:46",
      "content": "<p>I have no target leakage and user features don't work (0 feature importance). This doesn't make much sense to me. How good a student is should be number one predictor of how good he will be able to answer the questions. Any ideas?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1053791,
          "author_name": "kaushal2896",
          "author_url": "",
          "post_date": "10/19/2020 11:18:43",
          "content": "<p><a href=\"https://www.kaggle.com/emindeniz\" target=\"_blank\">@emindeniz</a> That's what I thought! See the feature importance in my case but I think I'm leaking information but don't know how to solve this.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1053794,
          "author_name": "kaushal2896",
          "author_url": "",
          "post_date": "10/19/2020 11:28:23",
          "content": "<p><a href=\"https://www.kaggle.com/emindeniz\" target=\"_blank\">@emindeniz</a> BTW how do you extracted such features without target leakage?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1053851,
          "author_name": "",
          "author_url": "",
          "post_date": "10/19/2020 12:27:26",
          "content": "<p>just like in the kernels, you can split the data by index. that's how the data is given</p>\n<p>features_df = all_df.iloc[int(6 /10 * len(all_df)):int(9 /10 * len(all_df))]</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1053865,
          "author_name": "kaushal2896",
          "author_url": "",
          "post_date": "10/19/2020 12:32:58",
          "content": "<p><a href=\"https://www.kaggle.com/emindeniz\" target=\"_blank\">@emindeniz</a> So you took a subset <code>(0.6*total_len to 0.9*total_len)</code> and calculated everything on it and then validate on remaining data right?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1053875,
          "author_name": "",
          "author_url": "",
          "post_date": "10/19/2020 12:47:32",
          "content": "<p>yes, that's what I did, and my validation matches pretty well with my submission. so i assume there is not leakage.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1053946,
          "author_name": "kaushal2896",
          "author_url": "",
          "post_date": "10/19/2020 14:03:50",
          "content": "<p><a href=\"https://www.kaggle.com/emindeniz\" target=\"_blank\">@emindeniz</a> I'm not able to understand how am I leaking features. If we don't want to validate in the notebook, we can train on entire dataset and without splitting. But this approach still reduces the LB score with those student feature. How is this a leakage? Leakage comes into picture if I use information from validation data (like mean, sum etc.) in training data. What am I missing here?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1053962,
          "author_name": "",
          "author_url": "",
          "post_date": "10/19/2020 14:15:57",
          "content": "<p>If you are not validating in the notebook, how can you say that LB score is reduced? what does it mean to say this approach reduces LB score.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1053967,
          "author_name": "kaushal2896",
          "author_url": "",
          "post_date": "10/19/2020 14:25:04",
          "content": "<p><a href=\"https://www.kaggle.com/emindeniz\" target=\"_blank\">@emindeniz</a> Sorry it was unclear. I meant I just trained on entire dataset with and without those features and just submitted. When I use those features, my public LB reduces by 0.01 and that significant.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1053975,
          "author_name": "",
          "author_url": "",
          "post_date": "10/19/2020 14:32:44",
          "content": "<p>adding a useless feature may result in overfitting hence decrease in score. That's a guess though.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1054099,
          "author_name": "kaushal2896",
          "author_url": "",
          "post_date": "10/19/2020 17:06:52",
          "content": "<p><a href=\"https://www.kaggle.com/emindeniz\" target=\"_blank\">@emindeniz</a> Hmmm, it might be the case that those features don't work at all as you pointed out but because of leakage in my validation strategy (initially I used 5 fold CV) my model was giving overconfident predictions but on unseen/actual data it worsens the performance significantly. I'll try without leakage and update here. Thanks!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1056236,
          "author_name": "oguzhannefsoglu",
          "author_url": "",
          "post_date": "10/21/2020 14:35:01",
          "content": "<p>as i understand, data is sorted both on user id and timestamp(correct me if im wrong). If this is correct, isnt splitting data the way you described is wrong? In that way you wont get same user ids in training data. So features related to user id wont contribute much.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1056328,
          "author_name": "kaushal2896",
          "author_url": "",
          "post_date": "10/21/2020 15:44:50",
          "content": "<p>Yes data is sorted that way (It was in the notebook I referred and I still don't understand the purpose though). But why you think splitting this way won't get \"same\" user ids? Can you clarify more? Thanks!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1056356,
          "author_name": "oguzhannefsoglu",
          "author_url": "",
          "post_date": "10/21/2020 16:15:26",
          "content": "<p>Actually i said it for EminOzkan's splitting method. I dont know if you are splitting this way. If so, you are extracting features from first n rows and training with remaining rows. Since data is sorted with user ids, feature data and training data will have different users.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1056368,
          "author_name": "kaushal2896",
          "author_url": "",
          "post_date": "10/21/2020 16:27:01",
          "content": "<p><a href=\"https://www.kaggle.com/oguzhannefsoglu\" target=\"_blank\">@oguzhannefsoglu</a> okay got it! Thanks!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1054371,
      "author_name": "isaacllorente",
      "author_url": "",
      "post_date": "10/19/2020 21:29:35",
      "content": "<p>I see your proposal as a discretized or binned group_by user_id mean encoding. The problem might be that you are actually using the label of an entry to get this feature, for this same entry. Imagine there's a user_id with only one row. Mean encoding it straight away would just tell you the actual label. An easy workaround to see if it is useful in this way would be to use only a part of the training set to get the mean encodings and to validate on unseen data, since this is what will happen in the test. However, this assumes that the performance of users doesn't change over time.</p>\n<p>Another more fancy way (that avoids leakage) would be to account only for past information relative to each entry. And it could even be updatable during test batches. You can even experiment with different window ranges, which also would make sense: I don't think that my grades 10 years ago reflect my current self.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1054408,
      "author_name": "aquatic",
      "author_url": "",
      "post_date": "10/19/2020 22:50:41",
      "content": "<p>I think there may be a few different things going on here, but just at face value I'd recommend against trying to create discrete features based on numerical feature bins if you're using gradient boosted tree models. As those models grow trees and find decision splits, exactly what they are doing is figuring out the best way to bin numerical features to separate negatives from positives, and they'll be able to be a lot more granular than what you're able to do manually. </p>\n<p>A good way to think about feature engineering is figuring out how to encode information that your model can't figure out on its own, e.g. information about past student behavior over many rows, since that's the information that will typically really boost your performance.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1054471,
          "author_name": "",
          "author_url": "",
          "post_date": "10/20/2020 01:07:03",
          "content": "<p>average accuracy of a student over all rows should give the model information about the past student behavior; but i find that it doesn't help, it ends up having zero gain. how could this be?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1054489,
          "author_name": "aquatic",
          "author_url": "",
          "post_date": "10/20/2020 01:39:19",
          "content": "<p>That sort of feature works well for me, so my best guess is that there may be something slightly off in the code? Is it possible that you're calculating user stats only on a dataset you didn't use for training or that the stats got misaligned with the user_ids in some kind of merging/concatenation/indexing error? A feature like that showing as 0 gain strongly suggests that the problem lies in the data preparation to me, but it's hard to guess where without knowing more about how you set up the feature.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1054874,
          "author_name": "kaushal2896",
          "author_url": "",
          "post_date": "10/20/2020 08:45:43",
          "content": "<p><a href=\"https://www.kaggle.com/aquatic\" target=\"_blank\">@aquatic</a> Measuring the performance of users for the past few rows require target knowledge. How would you extract this feature at test/validation time? Because at that time we only have predictions and not the actual targets of \"answered correctly\".</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1055106,
          "author_name": "aquatic",
          "author_url": "",
          "post_date": "10/20/2020 13:36:57",
          "content": "<p>I think that's one of the key challenges in the competition. I may share a kernel showing how to do it later on (for now I think it's too high on LB to be a good share), but my tip in the meantime is to think about how you can work with the <code>prior_group_answers_correct</code> information that you get from each test batch you process.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1055333,
          "author_name": "adityaecdrid",
          "author_url": "",
          "post_date": "10/20/2020 17:03:59",
          "content": "<p>I would be a great thing if you can share this, Are you updating stats about your features dynamically at test_time? Ty!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1055346,
          "author_name": "aquatic",
          "author_url": "",
          "post_date": "10/20/2020 17:23:45",
          "content": "<p>^Yes, I am doing that</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1045472": "I've extracted some features for users and questions which classifies users in 4 types: expert, above_average, average and dull/below average and also classifies questions in 4 types: too tough, tough, normal and easy.\n\nThe classification of users is done by thresholding their mean accuracy of correct answers and the classification of questions is done by thresholding mean accuracy of correct answers for particular question. Here is a code sample of my user features:\n\n    # If accuracy is high, student is \"expert\" and so on.\n    user_answers_df.loc[user_answers_df['mean_user_accuracy'] > 0.75, 'is_student_expert'] = 1\n    user_answers_df.loc[(user_answers_df['mean_user_accuracy'] > 0.5) & (user_answers_df['mean_user_accuracy'] <= 0.75) , 'is_student_above_average'] = 1\n    user_answers_df.loc[(user_answers_df['mean_user_accuracy'] > 0.25) & (user_answers_df['mean_user_accuracy'] <= 0.5) , 'is_student_average'] = 1\n    user_answers_df.loc[user_answers_df['mean_user_accuracy'] <= 0.25, 'is_student_dull'] = 1\n\nThis features looks promising to me but using them, I'm not able to achieve more than 0.69 irrespective of model I try. These features also have great feature importance in tree based models, and give 0.76 AUC on validation data but don't work on test data.\n\nSo I'm just curious about the reason behind this behaviour before moving towards other features. Thanks!",
    "1045500": "leakage? @kaushal2896",
    "1045508": "elvinagammed Could you please elaborate? I'm newbie in tabular datasets.",
    "1045511": "I think it's because test data in LB have unknown user and question.",
    "1045516": "I don't think so. All users and questions information is there in their specific csv files.",
    "1045524": "In user_id, this notebook shows that `train.csv` does not completely overlap test data. \n[checking-test-set-for-new-questions-and-users](https://www.kaggle.com/rimsky/checking-test-set-for-new-questions-and-users)\n\nI'm not sure about question_id.",
    "1045530": "ex question: ru using any k fold technics for ur validation? @kaushal2896",
    "1045532": "elvinagammed Yes I'm using target stratified 5 fold split for CV.",
    "1045667": "I think gbm models (xgboost or light gbm) may generate better classes of users and questions than you have.",
    "1046531": "How is the feature importance for this new features ?",
    "1046888": "phoenix9032 Feature Importance is pretty good. Here is the plot for XGBoost model. ![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1905996%2F25011ddb3865c56328189a493901bac6%2FScreenshot%202020-10-12%20at%209.58.24%20AM.png?generation=1602477015005326&alt=media)",
    "1053760": "I have no target leakage and user features don't work (0 feature importance). This doesn't make much sense to me. How good a student is should be number one predictor of how good he will be able to answer the questions. Any ideas?",
    "1053791": "emindeniz That's what I thought! See the feature importance in my case but I think I'm leaking information but don't know how to solve this.",
    "1053794": "emindeniz BTW how do you extracted such features without target leakage?",
    "1053851": "just like in the kernels, you can split the data by index. that's how the data is given\n\nfeatures_df = all_df.iloc[int(6 /10 * len(all_df)):int(9 /10 * len(all_df))]",
    "1053865": "emindeniz So you took a subset `(0.6*total_len to 0.9*total_len)` and calculated everything on it and then validate on remaining data right?",
    "1053875": "yes, that's what I did, and my validation matches pretty well with my submission. so i assume there is not leakage.",
    "1053946": "emindeniz I'm not able to understand how am I leaking features. If we don't want to validate in the notebook, we can train on entire dataset and without splitting. But this approach still reduces the LB score with those student feature. How is this a leakage? Leakage comes into picture if I use information from validation data (like mean, sum etc.) in training data. What am I missing here?",
    "1053962": "If you are not validating in the notebook, how can you say that LB score is reduced? what does it mean to say this approach reduces LB score.",
    "1053967": "emindeniz Sorry it was unclear. I meant I just trained on entire dataset with and without those features and just submitted. When I use those features, my public LB reduces by 0.01 and that significant.",
    "1053975": "adding a useless feature may result in overfitting hence decrease in score. That's a guess though.",
    "1054099": "emindeniz Hmmm, it might be the case that those features don't work at all as you pointed out but because of leakage in my validation strategy (initially I used 5 fold CV) my model was giving overconfident predictions but on unseen/actual data it worsens the performance significantly. I'll try without leakage and update here. Thanks!",
    "1054371": "I see your proposal as a discretized or binned group_by user_id mean encoding. The problem might be that you are actually using the label of an entry to get this feature, for this same entry. Imagine there's a user_id with only one row. Mean encoding it straight away would just tell you the actual label. An easy workaround to see if it is useful in this way would be to use only a part of the training set to get the mean encodings and to validate on unseen data, since this is what will happen in the test. However, this assumes that the performance of users doesn't change over time.\n\nAnother more fancy way (that avoids leakage) would be to account only for past information relative to each entry. And it could even be updatable during test batches. You can even experiment with different window ranges, which also would make sense: I don't think that my grades 10 years ago reflect my current self.",
    "1054408": "I think there may be a few different things going on here, but just at face value I'd recommend against trying to create discrete features based on numerical feature bins if you're using gradient boosted tree models. As those models grow trees and find decision splits, exactly what they are doing is figuring out the best way to bin numerical features to separate negatives from positives, and they'll be able to be a lot more granular than what you're able to do manually. \n\nA good way to think about feature engineering is figuring out how to encode information that your model can't figure out on its own, e.g. information about past student behavior over many rows, since that's the information that will typically really boost your performance.",
    "1054471": "average accuracy of a student over all rows should give the model information about the past student behavior; but i find that it doesn't help, it ends up having zero gain. how could this be?",
    "1054489": "That sort of feature works well for me, so my best guess is that there may be something slightly off in the code? Is it possible that you're calculating user stats only on a dataset you didn't use for training or that the stats got misaligned with the user_ids in some kind of merging/concatenation/indexing error? A feature like that showing as 0 gain strongly suggests that the problem lies in the data preparation to me, but it's hard to guess where without knowing more about how you set up the feature.",
    "1054874": "aquatic Measuring the performance of users for the past few rows require target knowledge. How would you extract this feature at test/validation time? Because at that time we only have predictions and not the actual targets of \"answered correctly\".",
    "1055106": "I think that's one of the key challenges in the competition. I may share a kernel showing how to do it later on (for now I think it's too high on LB to be a good share), but my tip in the meantime is to think about how you can work with the `prior_group_answers_correct` information that you get from each test batch you process.",
    "1055333": "I would be a great thing if you can share this, Are you updating stats about your features dynamically at test_time? Ty!",
    "1055346": "^Yes, I am doing that",
    "1056236": "as i understand, data is sorted both on user id and timestamp(correct me if im wrong). If this is correct, isnt splitting data the way you described is wrong? In that way you wont get same user ids in training data. So features related to user id wont contribute much.",
    "1056328": "Yes data is sorted that way (It was in the notebook I referred and I still don't understand the purpose though). But why you think splitting this way won't get \"same\" user ids? Can you clarify more? Thanks!",
    "1056356": "Actually i said it for EminOzkan's splitting method. I dont know if you are splitting this way. If so, you are extracting features from first n rows and training with remaining rows. Since data is sorted with user ids, feature data and training data will have different users.",
    "1056368": "oguzhannefsoglu okay got it! Thanks!"
  },
  "source": "meta"
}