{
  "id": 202086,
  "title": "Make smart features to reduce RAM usage and overfitting",
  "url": "/competitions/riiid-test-answer-prediction/discussion/202086",
  "author_name": "",
  "post_date": "2020-12-08T08:50:55.672134900Z",
  "votes": 31,
  "comment_count": 18,
  "views": 0,
  "content": "<p>This competition is all about smart feature engineering.</p>\n<p>I see there and there people using one hot encoding of the categorical features (categories, parts, tags, …). This is a terrible idea.</p>\n<p>First of all, it leads to explosition of memory usage, and second, it build models that are easy to overfit because of the number of dimensions.</p>\n<p>So instead of doing so be smart on the way you create your features. For exemple, if you want to create a feature that average the answers of each id for each part. Instead of building 7 features with some that might be unrelevant, make just one that map the relevant average depending of the part. </p>\n<p>By doing so not only you will reduce the overfitting due to the complexity of your data, but you will be able also to train on much more data.</p>",
  "messages": [
    {
      "id": "1105842",
      "postDate": "12/08/2020 08:50:55",
      "content": "<p>This competition is all about smart feature engineering.</p>\n<p>I see there and there people using one hot encoding of the categorical features (categories, parts, tags, …). This is a terrible idea.</p>\n<p>First of all, it leads to explosition of memory usage, and second, it build models that are easy to overfit because of the number of dimensions.</p>\n<p>So instead of doing so be smart on the way you create your features. For exemple, if you want to create a feature that average the answers of each id for each part. Instead of building 7 features with some that might be unrelevant, make just one that map the relevant average depending of the part. </p>\n<p>By doing so not only you will reduce the overfitting due to the complexity of your data, but you will be able also to train on much more data.</p>",
      "rawMarkdown": "This competition is all about smart feature engineering.\n\nI see there and there people using one hot encoding of the categorical features (categories, parts, tags, ...). This is a terrible idea.\n\nFirst of all, it leads to explosition of memory usage, and second, it build models that are easy to overfit because of the number of dimensions.\n\nSo instead of doing so be smart on the way you create your features. For exemple, if you want to create a feature that average the answers of each id for each part. Instead of building 7 features with some that might be unrelevant, make just one that map the relevant average depending of the part. \n\nBy doing so not only you will reduce the overfitting due to the complexity of your data, but you will be able also to train on much more data.",
      "votes": null
    },
    {
      "id": "1105862",
      "postDate": "12/08/2020 09:15:36",
      "content": "<p>Great Insight !</p>",
      "rawMarkdown": "Great Insight !",
      "votes": null
    },
    {
      "id": "1105939",
      "postDate": "12/08/2020 10:55:27",
      "content": "<p>thanks for sharing you experience, <br>\nI reached same conclusion, and was working on building smarter, smaller sized features (to be able to add more features).</p>",
      "rawMarkdown": "thanks for sharing you experience, \nI reached same conclusion, and was working on building smarter, smaller sized features (to be able to add more features).",
      "votes": null
    },
    {
      "id": "1106084",
      "postDate": "12/08/2020 13:55:33",
      "content": "<p>It seems hard for me to understand the solution( why we need to building 7 features ?), could you explain it more specifically? </p>",
      "rawMarkdown": "It seems hard for me to understand the solution( why we need to building 7 features ?), could you explain it more specifically?",
      "votes": null
    },
    {
      "id": "1106415",
      "postDate": "12/08/2020 20:10:12",
      "content": "<p>Good to see you here, <a href=\"https://www.kaggle.com/bowaka\" target=\"_blank\">@bowaka</a> !<br>\nYou're absolutely right. Actually I put one-hot-encoded tags in the beginning of this competition. <br>\nNow I'm thinking how to utilize tags in other way. </p>",
      "rawMarkdown": "Good to see you here, @bowaka !\nYou're absolutely right. Actually I put one-hot-encoded tags in the beginning of this competition. \nNow I'm thinking how to utilize tags in other way.",
      "votes": null
    },
    {
      "id": "1106498",
      "postDate": "12/08/2020 22:06:24",
      "content": "<p>Aha same !<br>\nI started also with one-hot-encoding before seing it was way overkill for this competition. <br>\nMy current model is a simple LGBoost that perfom quite well with only ~30 \"simple\" features.</p>\n<p>Personnaly I have 6 columns for the tags (as a question has at maximum 6 tags) + some columns where I make the average of some metrics for all tags.</p>\n<p>For this competition I also completly abandonned pandas and dataframe, and I am using simple for loop iteration and dictionnaries.</p>",
      "rawMarkdown": "Aha same !\nI started also with one-hot-encoding before seing it was way overkill for this competition. \nMy current model is a simple LGBoost that perfom quite well with only ~30 \"simple\" features.\n\nPersonnaly I have 6 columns for the tags (as a question has at maximum 6 tags) + some columns where I make the average of some metrics for all tags.\n\nFor this competition I also completly abandonned pandas and dataframe, and I am using simple for loop iteration and dictionnaries.",
      "votes": null
    },
    {
      "id": "1106505",
      "postDate": "12/08/2020 22:17:11",
      "content": "<p>For example, assume you want to keep track for each user_id of the average score for each part. That   would make  a total of 7 features to add in your training set, which is a lot. </p>\n<p>Instead, you can just have one feature that select the average score corresponding to the part corresponding to the question. You don't show to your classifier all the information (ie: the 7 scores), but just the relevant one (the score of the user_id regarding that given part).</p>\n<p>Now this work for scores, reaction time, parts tags, etc… </p>\n<p>I see some public notebooks with more that 1000 features 😨 Not only it is impossible to keep in RAM but it also overfit.</p>\n<p>I have a model trained with only 10 features that can already score 0.66+. <br>\nMy current model have 44 features, but they \"adapt\" to the context of the sample</p>",
      "rawMarkdown": "For example, assume you want to keep track for each user_id of the average score for each part. That   would make  a total of 7 features to add in your training set, which is a lot. \n\nInstead, you can just have one feature that select the average score corresponding to the part corresponding to the question. You don't show to your classifier all the information (ie: the 7 scores), but just the relevant one (the score of the user_id regarding that given part).\n\nNow this work for scores, reaction time, parts tags, etc... \n\nI see some public notebooks with more that 1000 features 😨 Not only it is impossible to keep in RAM but it also overfit.\n\nI have a model trained with only 10 features that can already score 0.66+. \nMy current model have 44 features, but they \"adapt\" to the context of the sample",
      "votes": null
    },
    {
      "id": "1106578",
      "postDate": "12/08/2020 23:44:14",
      "content": "<p>Hello <a href=\"https://www.kaggle.com/Bowaka\" target=\"_blank\">@Bowaka</a>, thanks for sharing your knowledge. Great sturff and well done.</p>",
      "rawMarkdown": "Hello @Bowaka, thanks for sharing your knowledge. Great sturff and well done.",
      "votes": null
    },
    {
      "id": "1106677",
      "postDate": "12/09/2020 03:12:52",
      "content": "<p>I understand it, thank you very much for sharing this! 😁</p>",
      "rawMarkdown": "I understand it, thank you very much for sharing this! 😁",
      "votes": null
    },
    {
      "id": "1106976",
      "postDate": "12/09/2020 09:08:50",
      "content": "<p>Of course, exponential increase of one hot features definitely problematic, but how about the chance that the one-hot encoded features overwhelm the increase of dimension? <br>\nIn your part features example, I thought 7 features may have some correlations each other. For instance, part 3 and part 4 feature somehow related, then part 4 feature could be helpful when we predict some part 3 question rows, or vice versa. We can utilize it in one-hot encoded settings and such relevance cannot be captured when a feature exists for each part in the row. We can take them all at first, and removing non-contributed features later.</p>",
      "rawMarkdown": "Of course, exponential increase of one hot features definitely problematic, but how about the chance that the one-hot encoded features overwhelm the increase of dimension? \nIn your part features example, I thought 7 features may have some correlations each other. For instance, part 3 and part 4 feature somehow related, then part 4 feature could be helpful when we predict some part 3 question rows, or vice versa. We can utilize it in one-hot encoded settings and such relevance cannot be captured when a feature exists for each part in the row. We can take them all at first, and removing non-contributed features later.",
      "votes": null
    },
    {
      "id": "1106984",
      "postDate": "12/09/2020 09:16:04",
      "content": "<p>I kind of agree, if you take my example isolated in make only 7 features.</p>\n<p>Now, I personnaly have one \"score\" metric, two lagged score average metrics, reaction time of the user, count of question, etc…</p>\n<p>If I have to use 7x times each of those, my ram would explode (and it did, at first, actually 😃 )</p>",
      "rawMarkdown": "I kind of agree, if you take my example isolated in make only 7 features.\n\nNow, I personnaly have one \"score\" metric, two lagged score average metrics, reaction time of the user, count of question, etc...\n\nIf I have to use 7x times each of those, my ram would explode (and it did, at first, actually 😃 )",
      "votes": null
    },
    {
      "id": "1107133",
      "postDate": "12/09/2020 12:21:01",
      "content": "<p>Totally agree. We should be careful when we take too large dimension and keep the number of feature optimal.</p>",
      "rawMarkdown": "Totally agree. We should be careful when we take too large dimension and keep the number of feature optimal.",
      "votes": null
    },
    {
      "id": "1108098",
      "postDate": "12/10/2020 09:11:34",
      "content": "<p>Great advice, but I wonder to know if there are some methods for feature selection ? Or identification of feature with the highest importance </p>",
      "rawMarkdown": "Great advice, but I wonder to know if there are some methods for feature selection ? Or identification of feature with the highest importance",
      "votes": null
    },
    {
      "id": "1108116",
      "postDate": "12/10/2020 09:30:43",
      "content": "<p>You can use feature importance tools such as <a href=\"https://www.kaggle.com/dansbecker/shap-values\" target=\"_blank\">the shap values</a>, or the feature importance parameter of your classifier.</p>\n<p>For feature selection you can also apply method such as using a Lasso to set some coefs to 0, etc…</p>\n<p>And of course, self judgment and logic 😃</p>",
      "rawMarkdown": "You can use feature importance tools such as [the shap values](https://www.kaggle.com/dansbecker/shap-values), or the feature importance parameter of your classifier.\n\nFor feature selection you can also apply method such as using a Lasso to set some coefs to 0, etc...\n\nAnd of course, self judgment and logic 😃",
      "votes": null
    },
    {
      "id": "1108122",
      "postDate": "12/10/2020 09:40:32",
      "content": "<p>Thanks for the hints 😄</p>",
      "rawMarkdown": "Thanks for the hints 😄",
      "votes": null
    },
    {
      "id": "1109210",
      "postDate": "12/11/2020 12:34:21",
      "content": "<p>Thanks Alot ! For the Hints ! Will try this out ! Generating Features Via Loops is much more feasible for this dataset as compared to generating Features using DataFrame operations ! </p>",
      "rawMarkdown": "Thanks Alot ! For the Hints ! Will try this out ! Generating Features Via Loops is much more feasible for this dataset as compared to generating Features using DataFrame operations !",
      "votes": null
    },
    {
      "id": "1109213",
      "postDate": "12/11/2020 12:39:55",
      "content": "<p>Yes I do that also.<br>\nAnd rather than iterating on the pandas dataframe, I convert it to numpy array and iterate on the rows. It goes much faster.</p>",
      "rawMarkdown": "Yes I do that also.\nAnd rather than iterating on the pandas dataframe, I convert it to numpy array and iterate on the rows. It goes much faster.",
      "votes": null
    },
    {
      "id": "1109220",
      "postDate": "12/11/2020 12:45:20",
      "content": "<p>Thanks ! Will also Try that !</p>",
      "rawMarkdown": "Thanks ! Will also Try that !",
      "votes": null
    },
    {
      "id": "1112100",
      "postDate": "12/14/2020 09:25:33",
      "content": "<p><a href=\"https://www.kaggle.com/bowaka\" target=\"_blank\">@bowaka</a>  feature based on averages   start getting in accurate when more test data comes in.Although we do update the feature stats every iteration but then model was trained only for certain Avg value.ut over  period of time during inference this value starts  getting different n different for model so its prediction might start getting off track</p>",
      "rawMarkdown": "bowaka  feature based on averages   start getting in accurate when more test data comes in.Although we do update the feature stats every iteration but then model was trained only for certain Avg value.ut over  period of time during inference this value starts  getting different n different for model so its prediction might start getting off track",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1105862,
      "author_name": "mamasinkgs",
      "author_url": "",
      "post_date": "12/08/2020 09:15:36",
      "content": "<p>Great Insight !</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1105939,
      "author_name": "mohamadnawfal",
      "author_url": "",
      "post_date": "12/08/2020 10:55:27",
      "content": "<p>thanks for sharing you experience, <br>\nI reached same conclusion, and was working on building smarter, smaller sized features (to be able to add more features).</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1106084,
      "author_name": "changyulve",
      "author_url": "",
      "post_date": "12/08/2020 13:55:33",
      "content": "<p>It seems hard for me to understand the solution( why we need to building 7 features ?), could you explain it more specifically? </p>",
      "votes": null,
      "replies": [
        {
          "id": 1106505,
          "author_name": "bowaka",
          "author_url": "",
          "post_date": "12/08/2020 22:17:11",
          "content": "<p>For example, assume you want to keep track for each user_id of the average score for each part. That   would make  a total of 7 features to add in your training set, which is a lot. </p>\n<p>Instead, you can just have one feature that select the average score corresponding to the part corresponding to the question. You don't show to your classifier all the information (ie: the 7 scores), but just the relevant one (the score of the user_id regarding that given part).</p>\n<p>Now this work for scores, reaction time, parts tags, etc… </p>\n<p>I see some public notebooks with more that 1000 features 😨 Not only it is impossible to keep in RAM but it also overfit.</p>\n<p>I have a model trained with only 10 features that can already score 0.66+. <br>\nMy current model have 44 features, but they \"adapt\" to the context of the sample</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1106677,
          "author_name": "changyulve",
          "author_url": "",
          "post_date": "12/09/2020 03:12:52",
          "content": "<p>I understand it, thank you very much for sharing this! 😁</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1106415,
      "author_name": "kokitanisaka",
      "author_url": "",
      "post_date": "12/08/2020 20:10:12",
      "content": "<p>Good to see you here, <a href=\"https://www.kaggle.com/bowaka\" target=\"_blank\">@bowaka</a> !<br>\nYou're absolutely right. Actually I put one-hot-encoded tags in the beginning of this competition. <br>\nNow I'm thinking how to utilize tags in other way. </p>",
      "votes": null,
      "replies": [
        {
          "id": 1106498,
          "author_name": "bowaka",
          "author_url": "",
          "post_date": "12/08/2020 22:06:24",
          "content": "<p>Aha same !<br>\nI started also with one-hot-encoding before seing it was way overkill for this competition. <br>\nMy current model is a simple LGBoost that perfom quite well with only ~30 \"simple\" features.</p>\n<p>Personnaly I have 6 columns for the tags (as a question has at maximum 6 tags) + some columns where I make the average of some metrics for all tags.</p>\n<p>For this competition I also completly abandonned pandas and dataframe, and I am using simple for loop iteration and dictionnaries.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1106578,
      "author_name": "puzuwe",
      "author_url": "",
      "post_date": "12/08/2020 23:44:14",
      "content": "<p>Hello <a href=\"https://www.kaggle.com/Bowaka\" target=\"_blank\">@Bowaka</a>, thanks for sharing your knowledge. Great sturff and well done.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1106976,
      "author_name": "jihunlorenzopark",
      "author_url": "",
      "post_date": "12/09/2020 09:08:50",
      "content": "<p>Of course, exponential increase of one hot features definitely problematic, but how about the chance that the one-hot encoded features overwhelm the increase of dimension? <br>\nIn your part features example, I thought 7 features may have some correlations each other. For instance, part 3 and part 4 feature somehow related, then part 4 feature could be helpful when we predict some part 3 question rows, or vice versa. We can utilize it in one-hot encoded settings and such relevance cannot be captured when a feature exists for each part in the row. We can take them all at first, and removing non-contributed features later.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1106984,
          "author_name": "bowaka",
          "author_url": "",
          "post_date": "12/09/2020 09:16:04",
          "content": "<p>I kind of agree, if you take my example isolated in make only 7 features.</p>\n<p>Now, I personnaly have one \"score\" metric, two lagged score average metrics, reaction time of the user, count of question, etc…</p>\n<p>If I have to use 7x times each of those, my ram would explode (and it did, at first, actually 😃 )</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1107133,
          "author_name": "jihunlorenzopark",
          "author_url": "",
          "post_date": "12/09/2020 12:21:01",
          "content": "<p>Totally agree. We should be careful when we take too large dimension and keep the number of feature optimal.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1108098,
      "author_name": "houssemayed",
      "author_url": "",
      "post_date": "12/10/2020 09:11:34",
      "content": "<p>Great advice, but I wonder to know if there are some methods for feature selection ? Or identification of feature with the highest importance </p>",
      "votes": null,
      "replies": [
        {
          "id": 1108116,
          "author_name": "bowaka",
          "author_url": "",
          "post_date": "12/10/2020 09:30:43",
          "content": "<p>You can use feature importance tools such as <a href=\"https://www.kaggle.com/dansbecker/shap-values\" target=\"_blank\">the shap values</a>, or the feature importance parameter of your classifier.</p>\n<p>For feature selection you can also apply method such as using a Lasso to set some coefs to 0, etc…</p>\n<p>And of course, self judgment and logic 😃</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1108122,
          "author_name": "houssemayed",
          "author_url": "",
          "post_date": "12/10/2020 09:40:32",
          "content": "<p>Thanks for the hints 😄</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1109210,
      "author_name": "sayedathar11",
      "author_url": "",
      "post_date": "12/11/2020 12:34:21",
      "content": "<p>Thanks Alot ! For the Hints ! Will try this out ! Generating Features Via Loops is much more feasible for this dataset as compared to generating Features using DataFrame operations ! </p>",
      "votes": null,
      "replies": [
        {
          "id": 1109213,
          "author_name": "bowaka",
          "author_url": "",
          "post_date": "12/11/2020 12:39:55",
          "content": "<p>Yes I do that also.<br>\nAnd rather than iterating on the pandas dataframe, I convert it to numpy array and iterate on the rows. It goes much faster.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1109220,
          "author_name": "sayedathar11",
          "author_url": "",
          "post_date": "12/11/2020 12:45:20",
          "content": "<p>Thanks ! Will also Try that !</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1112100,
      "author_name": "jaideepvalani",
      "author_url": "",
      "post_date": "12/14/2020 09:25:33",
      "content": "<p><a href=\"https://www.kaggle.com/bowaka\" target=\"_blank\">@bowaka</a>  feature based on averages   start getting in accurate when more test data comes in.Although we do update the feature stats every iteration but then model was trained only for certain Avg value.ut over  period of time during inference this value starts  getting different n different for model so its prediction might start getting off track</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1105842": "This competition is all about smart feature engineering.\n\nI see there and there people using one hot encoding of the categorical features (categories, parts, tags, ...). This is a terrible idea.\n\nFirst of all, it leads to explosition of memory usage, and second, it build models that are easy to overfit because of the number of dimensions.\n\nSo instead of doing so be smart on the way you create your features. For exemple, if you want to create a feature that average the answers of each id for each part. Instead of building 7 features with some that might be unrelevant, make just one that map the relevant average depending of the part. \n\nBy doing so not only you will reduce the overfitting due to the complexity of your data, but you will be able also to train on much more data.",
    "1105862": "Great Insight !",
    "1105939": "thanks for sharing you experience, \nI reached same conclusion, and was working on building smarter, smaller sized features (to be able to add more features).",
    "1106084": "It seems hard for me to understand the solution( why we need to building 7 features ?), could you explain it more specifically?",
    "1106415": "Good to see you here, @bowaka !\nYou're absolutely right. Actually I put one-hot-encoded tags in the beginning of this competition. \nNow I'm thinking how to utilize tags in other way.",
    "1106498": "Aha same !\nI started also with one-hot-encoding before seing it was way overkill for this competition. \nMy current model is a simple LGBoost that perfom quite well with only ~30 \"simple\" features.\n\nPersonnaly I have 6 columns for the tags (as a question has at maximum 6 tags) + some columns where I make the average of some metrics for all tags.\n\nFor this competition I also completly abandonned pandas and dataframe, and I am using simple for loop iteration and dictionnaries.",
    "1106505": "For example, assume you want to keep track for each user_id of the average score for each part. That   would make  a total of 7 features to add in your training set, which is a lot. \n\nInstead, you can just have one feature that select the average score corresponding to the part corresponding to the question. You don't show to your classifier all the information (ie: the 7 scores), but just the relevant one (the score of the user_id regarding that given part).\n\nNow this work for scores, reaction time, parts tags, etc... \n\nI see some public notebooks with more that 1000 features 😨 Not only it is impossible to keep in RAM but it also overfit.\n\nI have a model trained with only 10 features that can already score 0.66+. \nMy current model have 44 features, but they \"adapt\" to the context of the sample",
    "1106578": "Hello @Bowaka, thanks for sharing your knowledge. Great sturff and well done.",
    "1106677": "I understand it, thank you very much for sharing this! 😁",
    "1106976": "Of course, exponential increase of one hot features definitely problematic, but how about the chance that the one-hot encoded features overwhelm the increase of dimension? \nIn your part features example, I thought 7 features may have some correlations each other. For instance, part 3 and part 4 feature somehow related, then part 4 feature could be helpful when we predict some part 3 question rows, or vice versa. We can utilize it in one-hot encoded settings and such relevance cannot be captured when a feature exists for each part in the row. We can take them all at first, and removing non-contributed features later.",
    "1106984": "I kind of agree, if you take my example isolated in make only 7 features.\n\nNow, I personnaly have one \"score\" metric, two lagged score average metrics, reaction time of the user, count of question, etc...\n\nIf I have to use 7x times each of those, my ram would explode (and it did, at first, actually 😃 )",
    "1107133": "Totally agree. We should be careful when we take too large dimension and keep the number of feature optimal.",
    "1108098": "Great advice, but I wonder to know if there are some methods for feature selection ? Or identification of feature with the highest importance",
    "1108116": "You can use feature importance tools such as [the shap values](https://www.kaggle.com/dansbecker/shap-values), or the feature importance parameter of your classifier.\n\nFor feature selection you can also apply method such as using a Lasso to set some coefs to 0, etc...\n\nAnd of course, self judgment and logic 😃",
    "1108122": "Thanks for the hints 😄",
    "1109210": "Thanks Alot ! For the Hints ! Will try this out ! Generating Features Via Loops is much more feasible for this dataset as compared to generating Features using DataFrame operations !",
    "1109213": "Yes I do that also.\nAnd rather than iterating on the pandas dataframe, I convert it to numpy array and iterate on the rows. It goes much faster.",
    "1109220": "Thanks ! Will also Try that !",
    "1112100": "bowaka  feature based on averages   start getting in accurate when more test data comes in.Although we do update the feature stats every iteration but then model was trained only for certain Avg value.ut over  period of time during inference this value starts  getting different n different for model so its prediction might start getting off track"
  },
  "source": "meta"
}