{
  "id": 336704,
  "title": "Need Hints on Feature Engineering",
  "url": "/competitions/amex-default-prediction/discussion/336704",
  "author_name": "",
  "post_date": "2022-07-12T15:18:41.033200100Z",
  "votes": 3,
  "comment_count": 5,
  "views": 0,
  "content": "<p>Hello Everyone,</p>\n<p>I have tried lots of aggregated features and lag features, removed highly correlated features, and tried several models both tree and non tree based. My CV metric value is not breaching 0.78 (not made my submission on test data yet).</p>\n<p>Any pointers on more features will be greatly appreciated. </p>\n<p>Thanks</p>",
  "messages": [
    {
      "id": "1853073",
      "postDate": "07/12/2022 15:18:41",
      "content": "<p>Hello Everyone,</p>\n<p>I have tried lots of aggregated features and lag features, removed highly correlated features, and tried several models both tree and non tree based. My CV metric value is not breaching 0.78 (not made my submission on test data yet).</p>\n<p>Any pointers on more features will be greatly appreciated. </p>\n<p>Thanks</p>",
      "rawMarkdown": "Hello Everyone,\n\nI have tried lots of aggregated features and lag features, removed highly correlated features, and tried several models both tree and non tree based. My CV metric value is not breaching 0.78 (not made my submission on test data yet).\n\nAny pointers on more features will be greatly appreciated. \n\nThanks",
      "votes": null
    },
    {
      "id": "1853111",
      "postDate": "07/12/2022 15:56:08",
      "content": "<p>That is interesting, what type of CV and models are you using?<br>\nI would suggest trying to replicate some of the XGBoost or LGBM baselines such as <a href=\"https://www.kaggle.com/code/cdeotte/xgboost-starter-0-793\" target=\"_blank\">https://www.kaggle.com/code/cdeotte/xgboost-starter-0-793</a> or <a href=\"https://www.kaggle.com/code/ambrosm/amex-lightgbm-quickstart\" target=\"_blank\">https://www.kaggle.com/code/ambrosm/amex-lightgbm-quickstart</a><br>\nBoth of these two can easily get you above 0.790</p>",
      "rawMarkdown": "That is interesting, what type of CV and models are you using?\nI would suggest trying to replicate some of the XGBoost or LGBM baselines such as https://www.kaggle.com/code/cdeotte/xgboost-starter-0-793 or https://www.kaggle.com/code/ambrosm/amex-lightgbm-quickstart\nBoth of these two can easily get you above 0.790",
      "votes": null
    },
    {
      "id": "1853480",
      "postDate": "07/12/2022 23:15:08",
      "content": "<p>How many rounds of boosting are used? I suggest 10,000 rounds. What are the features and parameters you are using?</p>",
      "rawMarkdown": "How many rounds of boosting are used? I suggest 10,000 rounds. What are the features and parameters you are using?",
      "votes": null
    },
    {
      "id": "1853769",
      "postDate": "07/13/2022 05:59:24",
      "content": "<p>Thanks…I tried using XGBoost but I was using max 2k rounds. Perhaps I thought having too many rounds will make algo overfit to training data. Let me try going through the notebooks here and try with higher n_estimator value.</p>",
      "rawMarkdown": "Thanks...I tried using XGBoost but I was using max 2k rounds. Perhaps I thought having too many rounds will make algo overfit to training data. Let me try going through the notebooks here and try with higher n_estimator value.",
      "votes": null
    },
    {
      "id": "1853780",
      "postDate": "07/13/2022 06:12:06",
      "content": "<p>My features are basic aggregation (median/mean/max/min) across Num Vars and OH encoding of cat vars and mean on top of OH encodings. If a customer does not change any cat var in the entire 13 month window, the mean for that OH column will be 1. If it does, then two columns will have values for that customer proportional to time spent on that category label. For e.g. - if Income Band is a category, and customer was in Band 1 for 8 out of 13 months and Band 2 for 5 out of 13 months, then my column IB_1 will have 8/13 value and IB_2 will have 5/13 value in the final dataset. </p>\n<p>Apart from this, I have also taken value for Num Vars as on Last Month (L0) and other significant months as per consumer cycles (0,1,2,3,6,12) taking difference with L0</p>",
      "rawMarkdown": "My features are basic aggregation (median/mean/max/min) across Num Vars and OH encoding of cat vars and mean on top of OH encodings. If a customer does not change any cat var in the entire 13 month window, the mean for that OH column will be 1. If it does, then two columns will have values for that customer proportional to time spent on that category label. For e.g. - if Income Band is a category, and customer was in Band 1 for 8 out of 13 months and Band 2 for 5 out of 13 months, then my column IB_1 will have 8/13 value and IB_2 will have 5/13 value in the final dataset. \n\nApart from this, I have also taken value for Num Vars as on Last Month (L0) and other significant months as per consumer cycles (0,1,2,3,6,12) taking difference with L0",
      "votes": null
    },
    {
      "id": "1853885",
      "postDate": "07/13/2022 08:58:21",
      "content": "<p>You can just use early stopping to detect overfitting and cancel your training accordingly. For XGBoost you need to declare in your fit method following:</p>\n<p><code>early_stopping_rounds=100</code></p>\n<p>This is the patience of the early stopping algorithm. I use 100 but you can choose any other value you like. For sure you need validation data for it to work too:</p>\n<p><code>eval_set=[(X_valid_gb, y_valid_gb)]</code></p>",
      "rawMarkdown": "You can just use early stopping to detect overfitting and cancel your training accordingly. For XGBoost you need to declare in your fit method following:\n\n`early_stopping_rounds=100`\n\nThis is the patience of the early stopping algorithm. I use 100 but you can choose any other value you like. For sure you need validation data for it to work too:\n\n`eval_set=[(X_valid_gb, y_valid_gb)]`",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1853111,
      "author_name": "mrandri19",
      "author_url": "",
      "post_date": "07/12/2022 15:56:08",
      "content": "<p>That is interesting, what type of CV and models are you using?<br>\nI would suggest trying to replicate some of the XGBoost or LGBM baselines such as <a href=\"https://www.kaggle.com/code/cdeotte/xgboost-starter-0-793\" target=\"_blank\">https://www.kaggle.com/code/cdeotte/xgboost-starter-0-793</a> or <a href=\"https://www.kaggle.com/code/ambrosm/amex-lightgbm-quickstart\" target=\"_blank\">https://www.kaggle.com/code/ambrosm/amex-lightgbm-quickstart</a><br>\nBoth of these two can easily get you above 0.790</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1853480,
      "author_name": "mohammadrahmati",
      "author_url": "",
      "post_date": "07/12/2022 23:15:08",
      "content": "<p>How many rounds of boosting are used? I suggest 10,000 rounds. What are the features and parameters you are using?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1853769,
      "author_name": "jayeshgokhale",
      "author_url": "",
      "post_date": "07/13/2022 05:59:24",
      "content": "<p>Thanks…I tried using XGBoost but I was using max 2k rounds. Perhaps I thought having too many rounds will make algo overfit to training data. Let me try going through the notebooks here and try with higher n_estimator value.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1853885,
          "author_name": "aliabdin1",
          "author_url": "",
          "post_date": "07/13/2022 08:58:21",
          "content": "<p>You can just use early stopping to detect overfitting and cancel your training accordingly. For XGBoost you need to declare in your fit method following:</p>\n<p><code>early_stopping_rounds=100</code></p>\n<p>This is the patience of the early stopping algorithm. I use 100 but you can choose any other value you like. For sure you need validation data for it to work too:</p>\n<p><code>eval_set=[(X_valid_gb, y_valid_gb)]</code></p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1853780,
      "author_name": "jayeshgokhale",
      "author_url": "",
      "post_date": "07/13/2022 06:12:06",
      "content": "<p>My features are basic aggregation (median/mean/max/min) across Num Vars and OH encoding of cat vars and mean on top of OH encodings. If a customer does not change any cat var in the entire 13 month window, the mean for that OH column will be 1. If it does, then two columns will have values for that customer proportional to time spent on that category label. For e.g. - if Income Band is a category, and customer was in Band 1 for 8 out of 13 months and Band 2 for 5 out of 13 months, then my column IB_1 will have 8/13 value and IB_2 will have 5/13 value in the final dataset. </p>\n<p>Apart from this, I have also taken value for Num Vars as on Last Month (L0) and other significant months as per consumer cycles (0,1,2,3,6,12) taking difference with L0</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1853073": "Hello Everyone,\n\nI have tried lots of aggregated features and lag features, removed highly correlated features, and tried several models both tree and non tree based. My CV metric value is not breaching 0.78 (not made my submission on test data yet).\n\nAny pointers on more features will be greatly appreciated. \n\nThanks",
    "1853111": "That is interesting, what type of CV and models are you using?\nI would suggest trying to replicate some of the XGBoost or LGBM baselines such as https://www.kaggle.com/code/cdeotte/xgboost-starter-0-793 or https://www.kaggle.com/code/ambrosm/amex-lightgbm-quickstart\nBoth of these two can easily get you above 0.790",
    "1853480": "How many rounds of boosting are used? I suggest 10,000 rounds. What are the features and parameters you are using?",
    "1853769": "Thanks...I tried using XGBoost but I was using max 2k rounds. Perhaps I thought having too many rounds will make algo overfit to training data. Let me try going through the notebooks here and try with higher n_estimator value.",
    "1853780": "My features are basic aggregation (median/mean/max/min) across Num Vars and OH encoding of cat vars and mean on top of OH encodings. If a customer does not change any cat var in the entire 13 month window, the mean for that OH column will be 1. If it does, then two columns will have values for that customer proportional to time spent on that category label. For e.g. - if Income Band is a category, and customer was in Band 1 for 8 out of 13 months and Band 2 for 5 out of 13 months, then my column IB_1 will have 8/13 value and IB_2 will have 5/13 value in the final dataset. \n\nApart from this, I have also taken value for Num Vars as on Last Month (L0) and other significant months as per consumer cycles (0,1,2,3,6,12) taking difference with L0",
    "1853885": "You can just use early stopping to detect overfitting and cancel your training accordingly. For XGBoost you need to declare in your fit method following:\n\n`early_stopping_rounds=100`\n\nThis is the patience of the early stopping algorithm. I use 100 but you can choose any other value you like. For sure you need validation data for it to work too:\n\n`eval_set=[(X_valid_gb, y_valid_gb)]`"
  },
  "source": "meta"
}