{
  "id": 348058,
  "title": "12th Place Gold (2/2) - LGBM + XGBoost + Catboost",
  "url": "/competitions/amex-default-prediction/discussion/348058",
  "author_name": "",
  "post_date": "2022-08-26T16:57:20.308527400Z",
  "votes": 47,
  "comment_count": 13,
  "views": 0,
  "content": "<p>Thanks to AMEX, Kaggle, and all participants for this great competition!</p>\n<p>This post is the second part of our team's ensemble solution for 15th place. The first part written by my great teammate <a href=\"https://www.kaggle.com/joseantonioalatorre\" target=\"_blank\">@joseantonioalatorre</a> can be seen <a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/347740\" target=\"_blank\">here</a>.</p>\n<p>TLDR: My part of the ensemble consists of LGBM, XGBoost, and Catboost models. In addition to the standard features, my most important features were model predictions for each individual statement and aggregations of only the last 3 statements of each customer. The best model was a dart-LGBM with early stopping.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3715340%2F6dcd770ec0417b3695f87747a0725585%2Famex_models.png?generation=1661532812385651&amp;alt=media\" alt=\"\"></p>\n<h1>Features</h1>\n<p>I only used the integer dataset provided by <a href=\"https://www.kaggle.com/raddar\" target=\"_blank\">@raddar</a>. Here are some additional features that I didn't see in public notebooks:</p>\n<ul>\n<li>As each customer has several statements, I trained a model on the individual statements and created a dataset based on the aggregated predictions for each customer (mean, std, last…). (similar to the approach <a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/347786\" target=\"_blank\">here</a>)</li>\n<li>Aggregations of only on the last 3 statements for each customer</li>\n<li>Aggregations for the time between two statements</li>\n<li>10 PCA dimensions</li>\n<li>Coefficient of variation (<a href=\"https://en.wikipedia.org/wiki/Coefficient_of_variation\" target=\"_blank\">https://en.wikipedia.org/wiki/Coefficient_of_variation</a>)</li>\n</ul>\n<h1>Models</h1>\n<p>I used a diverse set of models, with the dart-LGBM performing best overall. For all models, I used 10-fold cross-validation and trained on binary cross-entropy. Overall my dataset consisted of about 2-3 thousand features. Hyperparameter-tuning or feature selection did not seem to help much. <br>\nI should note that a huge part of my time went into finding ways to prevent memory errors from training with that many features.</p>\n<h2>LGBM</h2>\n<p>I used dart with a custom early stopping callback on the validation metric. I trained several models with small changes in features and parameters. The parameters are mostly similar to the public kernels. I did however not transform the features to 16bit.</p>\n<h2>XGBoost</h2>\n<p>This model was trained on the GPU in a similar style as the LGBM.</p>\n<h2>Catboost</h2>\n<p>The Catboost was the hardest to get right, especially since no early stopping or callbacks were available for GPU training. This model's predictions have very little weight in the ensemble, however, they brought a significant boost to the final score.</p>\n<h3>What I tried that did not work:</h3>\n<ul>\n<li>Neural Networks<br>\nCouldn't get it working, and when I thought I did, the public score was abysmal.</li>\n<li>Autoencoder-Features</li>\n<li>Cluster aggregations</li>\n<li>Using a model to predict whether a statement is the last of a customer</li>\n<li>….</li>\n</ul>\n<p>I learned a lot in this competition, met new friends and got a very pleasant result. Huge success for me.</p>",
  "messages": [
    {
      "id": "1915099",
      "postDate": "08/26/2022 16:57:20",
      "content": "<p>Thanks to AMEX, Kaggle, and all participants for this great competition!</p>\n<p>This post is the second part of our team's ensemble solution for 15th place. The first part written by my great teammate <a href=\"https://www.kaggle.com/joseantonioalatorre\" target=\"_blank\">@joseantonioalatorre</a> can be seen <a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/347740\" target=\"_blank\">here</a>.</p>\n<p>TLDR: My part of the ensemble consists of LGBM, XGBoost, and Catboost models. In addition to the standard features, my most important features were model predictions for each individual statement and aggregations of only the last 3 statements of each customer. The best model was a dart-LGBM with early stopping.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3715340%2F6dcd770ec0417b3695f87747a0725585%2Famex_models.png?generation=1661532812385651&amp;alt=media\" alt=\"\"></p>\n<h1>Features</h1>\n<p>I only used the integer dataset provided by <a href=\"https://www.kaggle.com/raddar\" target=\"_blank\">@raddar</a>. Here are some additional features that I didn't see in public notebooks:</p>\n<ul>\n<li>As each customer has several statements, I trained a model on the individual statements and created a dataset based on the aggregated predictions for each customer (mean, std, last…). (similar to the approach <a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/347786\" target=\"_blank\">here</a>)</li>\n<li>Aggregations of only on the last 3 statements for each customer</li>\n<li>Aggregations for the time between two statements</li>\n<li>10 PCA dimensions</li>\n<li>Coefficient of variation (<a href=\"https://en.wikipedia.org/wiki/Coefficient_of_variation\" target=\"_blank\">https://en.wikipedia.org/wiki/Coefficient_of_variation</a>)</li>\n</ul>\n<h1>Models</h1>\n<p>I used a diverse set of models, with the dart-LGBM performing best overall. For all models, I used 10-fold cross-validation and trained on binary cross-entropy. Overall my dataset consisted of about 2-3 thousand features. Hyperparameter-tuning or feature selection did not seem to help much. <br>\nI should note that a huge part of my time went into finding ways to prevent memory errors from training with that many features.</p>\n<h2>LGBM</h2>\n<p>I used dart with a custom early stopping callback on the validation metric. I trained several models with small changes in features and parameters. The parameters are mostly similar to the public kernels. I did however not transform the features to 16bit.</p>\n<h2>XGBoost</h2>\n<p>This model was trained on the GPU in a similar style as the LGBM.</p>\n<h2>Catboost</h2>\n<p>The Catboost was the hardest to get right, especially since no early stopping or callbacks were available for GPU training. This model's predictions have very little weight in the ensemble, however, they brought a significant boost to the final score.</p>\n<h3>What I tried that did not work:</h3>\n<ul>\n<li>Neural Networks<br>\nCouldn't get it working, and when I thought I did, the public score was abysmal.</li>\n<li>Autoencoder-Features</li>\n<li>Cluster aggregations</li>\n<li>Using a model to predict whether a statement is the last of a customer</li>\n<li>….</li>\n</ul>\n<p>I learned a lot in this competition, met new friends and got a very pleasant result. Huge success for me.</p>",
      "rawMarkdown": "Thanks to AMEX, Kaggle, and all participants for this great competition!\n\nThis post is the second part of our team's ensemble solution for 15th place. The first part written by my great teammate @joseantonioalatorre can be seen [here](https://www.kaggle.com/competitions/amex-default-prediction/discussion/347740).\n\nTLDR: My part of the ensemble consists of LGBM, XGBoost, and Catboost models. In addition to the standard features, my most important features were model predictions for each individual statement and aggregations of only the last 3 statements of each customer. The best model was a dart-LGBM with early stopping.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3715340%2F6dcd770ec0417b3695f87747a0725585%2Famex_models.png?generation=1661532812385651&alt=media)\n\n# Features\nI only used the integer dataset provided by @raddar. Here are some additional features that I didn't see in public notebooks:\n\n- As each customer has several statements, I trained a model on the individual statements and created a dataset based on the aggregated predictions for each customer (mean, std, last...). (similar to the approach [here](https://www.kaggle.com/competitions/amex-default-prediction/discussion/347786))\n- Aggregations of only on the last 3 statements for each customer\n- Aggregations for the time between two statements\n- 10 PCA dimensions\n- Coefficient of variation (https://en.wikipedia.org/wiki/Coefficient_of_variation)\n\n# Models\nI used a diverse set of models, with the dart-LGBM performing best overall. For all models, I used 10-fold cross-validation and trained on binary cross-entropy. Overall my dataset consisted of about 2-3 thousand features. Hyperparameter-tuning or feature selection did not seem to help much. \nI should note that a huge part of my time went into finding ways to prevent memory errors from training with that many features.\n\n## LGBM\nI used dart with a custom early stopping callback on the validation metric. I trained several models with small changes in features and parameters. The parameters are mostly similar to the public kernels. I did however not transform the features to 16bit.\n\n## XGBoost\nThis model was trained on the GPU in a similar style as the LGBM.\n\n## Catboost\nThe Catboost was the hardest to get right, especially since no early stopping or callbacks were available for GPU training. This model's predictions have very little weight in the ensemble, however, they brought a significant boost to the final score.\n\n\n### What I tried that did not work:\n- Neural Networks\nCouldn't get it working, and when I thought I did, the public score was abysmal.\n- Autoencoder-Features\n- Cluster aggregations\n- Using a model to predict whether a statement is the last of a customer\n- ....\n\nI learned a lot in this competition, met new friends and got a very pleasant result. Huge success for me.",
      "votes": null
    },
    {
      "id": "1915126",
      "postDate": "08/26/2022 17:23:12",
      "content": "<p>Felicidades Daniel. ¡Buen trabajo!</p>",
      "rawMarkdown": "Felicidades Daniel. ¡Buen trabajo!",
      "votes": null
    },
    {
      "id": "1915173",
      "postDate": "08/26/2022 18:11:02",
      "content": "<p>Amazing approach! Keep up the great work <a href=\"https://www.kaggle.com/gandagorn\" target=\"_blank\">@gandagorn</a> !! Hearty congratulations for the result as well!</p>",
      "rawMarkdown": "Amazing approach! Keep up the great work @gandagorn !! Hearty congratulations for the result as well!",
      "votes": null
    },
    {
      "id": "1915174",
      "postDate": "08/26/2022 18:11:14",
      "content": "<p>Muchas gracias!</p>",
      "rawMarkdown": "Muchas gracias!",
      "votes": null
    },
    {
      "id": "1915180",
      "postDate": "08/26/2022 18:13:52",
      "content": "<p>Thank you very much!</p>",
      "rawMarkdown": "Thank you very much!",
      "votes": null
    },
    {
      "id": "1915281",
      "postDate": "08/26/2022 20:22:22",
      "content": "<p>Congrats and thanks for the thorough desrciption! <br>\nCould you please say bit more about Auto encoder features? Was it a DAE? Used both train/test data? In their raw or aggregated (by customer id) form? Thanks!</p>",
      "rawMarkdown": "Congrats and thanks for the thorough desrciption! \nCould you please say bit more about Auto encoder features? Was it a DAE? Used both train/test data? In their raw or aggregated (by customer id) form? Thanks!",
      "votes": null
    },
    {
      "id": "1915329",
      "postDate": "08/26/2022 21:54:34",
      "content": "<p>Great work and approach - congrats!</p>",
      "rawMarkdown": "Great work and approach - congrats!",
      "votes": null
    },
    {
      "id": "1915354",
      "postDate": "08/26/2022 22:55:34",
      "content": "<p>Thanks for sharing your approach, flowchart make it clear</p>",
      "rawMarkdown": "Thanks for sharing your approach, flowchart make it clear",
      "votes": null
    },
    {
      "id": "1915720",
      "postDate": "08/27/2022 08:43:45",
      "content": "<p>No only with bottleneck layer, with both train and test in their aggregated form. It didn't improve the model, so I focused on other things.</p>",
      "rawMarkdown": "No only with bottleneck layer, with both train and test in their aggregated form. It didn't improve the model, so I focused on other things.",
      "votes": null
    },
    {
      "id": "1918496",
      "postDate": "08/29/2022 16:04:31",
      "content": "<p>Great post and description. Thanks for sharing this. Learning a lot ..🙏</p>",
      "rawMarkdown": "Great post and description. Thanks for sharing this. Learning a lot ..🙏",
      "votes": null
    },
    {
      "id": "1920889",
      "postDate": "08/31/2022 13:06:55",
      "content": "<p>Thanks for sharing this solution! May I invite you to participate in this survey regarding your experience on Kaggle (10 min)? This is not a scam. We are a group of researchers at the City University of Hong Kong. The survey link is: <a href=\"https://cityuhk.questionpro.com/survey-of-kaggle-contestants\" target=\"_blank\">https://cityuhk.questionpro.com/survey-of-kaggle-contestants</a></p>",
      "rawMarkdown": "Thanks for sharing this solution! May I invite you to participate in this survey regarding your experience on Kaggle (10 min)? This is not a scam. We are a group of researchers at the City University of Hong Kong. The survey link is: https://cityuhk.questionpro.com/survey-of-kaggle-contestants",
      "votes": null
    },
    {
      "id": "1926534",
      "postDate": "09/04/2022 23:11:39",
      "content": "<p>Thanks for sharing it! Congrats!!</p>",
      "rawMarkdown": "Thanks for sharing it! Congrats!!",
      "votes": null
    },
    {
      "id": "1928216",
      "postDate": "09/06/2022 10:33:24",
      "content": "<p>Congratulations and thanks for sharing! Same as Jose, happy to see a spanish kaggler (are you spanish, right?😄) in the gold zone. </p>\n<p>I have some questions about your approach:</p>\n<p><strong>1.</strong> This one is to be sure that I understood correctly how you calculated some features:</p>\n<ul>\n<li><em>Aggregations for the time between two statements</em>:  If I am not mistaken time between statements was almost constant (~month). What type of features did you create here?</li>\n<li><em>10 PCA dimensions</em>: Over which features did you compute the PCA?</li>\n<li><em>Coefficient of variation</em>: Is this one another aggregated feature by customer_ID? Because it is essentially the same as calculating the std, right?</li>\n</ul>\n<p><strong>2.</strong> My second question is about how you did the FE/FS part. Do you started with standard aggregation features and sequentially adding the rest of the feature subsets (e.g. add all the PCA features all at once)? Did you perform any kind of FS selection method or just adding feature subsets one by one and checking whether the CV/LB scores improved above a threshold?</p>",
      "rawMarkdown": "Congratulations and thanks for sharing! Same as Jose, happy to see a spanish kaggler (are you spanish, right?😄) in the gold zone. \n\nI have some questions about your approach:\n\n**1.** This one is to be sure that I understood correctly how you calculated some features:\n- *Aggregations for the time between two statements*:  If I am not mistaken time between statements was almost constant (~month). What type of features did you create here?\n- *10 PCA dimensions*: Over which features did you compute the PCA?\n- *Coefficient of variation*: Is this one another aggregated feature by customer_ID? Because it is essentially the same as calculating the std, right?\n\n**2.** My second question is about how you did the FE/FS part. Do you started with standard aggregation features and sequentially adding the rest of the feature subsets (e.g. add all the PCA features all at once)? Did you perform any kind of FS selection method or just adding feature subsets one by one and checking whether the CV/LB scores improved above a threshold?",
      "votes": null
    },
    {
      "id": "1929143",
      "postDate": "09/06/2022 21:15:35",
      "content": "<p>Hi, thanks for your comment! Actually I'm from Austria :D</p>\n<ul>\n<li>Most of the time it was, but not for all samples, and I wanted to include that information. I used the diff of days between statements and build aggregations (mean, std …). Overall it wasn't a very important feature class.</li>\n<li>Over the aggregation features</li>\n<li>It's the ratio between std and mean. I basically used every pair of features XX_mean and XX_std to create a feature XX_cov.</li>\n<li>I created the different classes of features and iteratively added them to the training. I did not use much FS, however, I removed classes of features that seemed to decrease the models performance (which was tricky, as the results had a high variance). Went with my gut feeling most of the time.</li>\n</ul>",
      "rawMarkdown": "Hi, thanks for your comment! Actually I'm from Austria :D\n- Most of the time it was, but not for all samples, and I wanted to include that information. I used the diff of days between statements and build aggregations (mean, std ...). Overall it wasn't a very important feature class.\n- Over the aggregation features\n- It's the ratio between std and mean. I basically used every pair of features XX_mean and XX_std to create a feature XX_cov.\n- I created the different classes of features and iteratively added them to the training. I did not use much FS, however, I removed classes of features that seemed to decrease the models performance (which was tricky, as the results had a high variance). Went with my gut feeling most of the time.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1915126,
      "author_name": "santiagomota",
      "author_url": "",
      "post_date": "08/26/2022 17:23:12",
      "content": "<p>Felicidades Daniel. ¡Buen trabajo!</p>",
      "votes": null,
      "replies": [
        {
          "id": 1915174,
          "author_name": "gandagorn",
          "author_url": "",
          "post_date": "08/26/2022 18:11:14",
          "content": "<p>Muchas gracias!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1915173,
      "author_name": "ravi20076",
      "author_url": "",
      "post_date": "08/26/2022 18:11:02",
      "content": "<p>Amazing approach! Keep up the great work <a href=\"https://www.kaggle.com/gandagorn\" target=\"_blank\">@gandagorn</a> !! Hearty congratulations for the result as well!</p>",
      "votes": null,
      "replies": [
        {
          "id": 1915180,
          "author_name": "gandagorn",
          "author_url": "",
          "post_date": "08/26/2022 18:13:52",
          "content": "<p>Thank you very much!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1915281,
      "author_name": "imeintanis",
      "author_url": "",
      "post_date": "08/26/2022 20:22:22",
      "content": "<p>Congrats and thanks for the thorough desrciption! <br>\nCould you please say bit more about Auto encoder features? Was it a DAE? Used both train/test data? In their raw or aggregated (by customer id) form? Thanks!</p>",
      "votes": null,
      "replies": [
        {
          "id": 1915720,
          "author_name": "gandagorn",
          "author_url": "",
          "post_date": "08/27/2022 08:43:45",
          "content": "<p>No only with bottleneck layer, with both train and test in their aggregated form. It didn't improve the model, so I focused on other things.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1915329,
      "author_name": "alberteinsten",
      "author_url": "",
      "post_date": "08/26/2022 21:54:34",
      "content": "<p>Great work and approach - congrats!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1915354,
      "author_name": "cv13j0",
      "author_url": "",
      "post_date": "08/26/2022 22:55:34",
      "content": "<p>Thanks for sharing your approach, flowchart make it clear</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1918496,
      "author_name": "cid007",
      "author_url": "",
      "post_date": "08/29/2022 16:04:31",
      "content": "<p>Great post and description. Thanks for sharing this. Learning a lot ..🙏</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1920889,
      "author_name": "lystriving",
      "author_url": "",
      "post_date": "08/31/2022 13:06:55",
      "content": "<p>Thanks for sharing this solution! May I invite you to participate in this survey regarding your experience on Kaggle (10 min)? This is not a scam. We are a group of researchers at the City University of Hong Kong. The survey link is: <a href=\"https://cityuhk.questionpro.com/survey-of-kaggle-contestants\" target=\"_blank\">https://cityuhk.questionpro.com/survey-of-kaggle-contestants</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1926534,
      "author_name": "mahyararani",
      "author_url": "",
      "post_date": "09/04/2022 23:11:39",
      "content": "<p>Thanks for sharing it! Congrats!!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1928216,
      "author_name": "delai50",
      "author_url": "",
      "post_date": "09/06/2022 10:33:24",
      "content": "<p>Congratulations and thanks for sharing! Same as Jose, happy to see a spanish kaggler (are you spanish, right?😄) in the gold zone. </p>\n<p>I have some questions about your approach:</p>\n<p><strong>1.</strong> This one is to be sure that I understood correctly how you calculated some features:</p>\n<ul>\n<li><em>Aggregations for the time between two statements</em>:  If I am not mistaken time between statements was almost constant (~month). What type of features did you create here?</li>\n<li><em>10 PCA dimensions</em>: Over which features did you compute the PCA?</li>\n<li><em>Coefficient of variation</em>: Is this one another aggregated feature by customer_ID? Because it is essentially the same as calculating the std, right?</li>\n</ul>\n<p><strong>2.</strong> My second question is about how you did the FE/FS part. Do you started with standard aggregation features and sequentially adding the rest of the feature subsets (e.g. add all the PCA features all at once)? Did you perform any kind of FS selection method or just adding feature subsets one by one and checking whether the CV/LB scores improved above a threshold?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1929143,
          "author_name": "gandagorn",
          "author_url": "",
          "post_date": "09/06/2022 21:15:35",
          "content": "<p>Hi, thanks for your comment! Actually I'm from Austria :D</p>\n<ul>\n<li>Most of the time it was, but not for all samples, and I wanted to include that information. I used the diff of days between statements and build aggregations (mean, std …). Overall it wasn't a very important feature class.</li>\n<li>Over the aggregation features</li>\n<li>It's the ratio between std and mean. I basically used every pair of features XX_mean and XX_std to create a feature XX_cov.</li>\n<li>I created the different classes of features and iteratively added them to the training. I did not use much FS, however, I removed classes of features that seemed to decrease the models performance (which was tricky, as the results had a high variance). Went with my gut feeling most of the time.</li>\n</ul>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1915099": "Thanks to AMEX, Kaggle, and all participants for this great competition!\n\nThis post is the second part of our team's ensemble solution for 15th place. The first part written by my great teammate @joseantonioalatorre can be seen [here](https://www.kaggle.com/competitions/amex-default-prediction/discussion/347740).\n\nTLDR: My part of the ensemble consists of LGBM, XGBoost, and Catboost models. In addition to the standard features, my most important features were model predictions for each individual statement and aggregations of only the last 3 statements of each customer. The best model was a dart-LGBM with early stopping.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3715340%2F6dcd770ec0417b3695f87747a0725585%2Famex_models.png?generation=1661532812385651&alt=media)\n\n# Features\nI only used the integer dataset provided by @raddar. Here are some additional features that I didn't see in public notebooks:\n\n- As each customer has several statements, I trained a model on the individual statements and created a dataset based on the aggregated predictions for each customer (mean, std, last...). (similar to the approach [here](https://www.kaggle.com/competitions/amex-default-prediction/discussion/347786))\n- Aggregations of only on the last 3 statements for each customer\n- Aggregations for the time between two statements\n- 10 PCA dimensions\n- Coefficient of variation (https://en.wikipedia.org/wiki/Coefficient_of_variation)\n\n# Models\nI used a diverse set of models, with the dart-LGBM performing best overall. For all models, I used 10-fold cross-validation and trained on binary cross-entropy. Overall my dataset consisted of about 2-3 thousand features. Hyperparameter-tuning or feature selection did not seem to help much. \nI should note that a huge part of my time went into finding ways to prevent memory errors from training with that many features.\n\n## LGBM\nI used dart with a custom early stopping callback on the validation metric. I trained several models with small changes in features and parameters. The parameters are mostly similar to the public kernels. I did however not transform the features to 16bit.\n\n## XGBoost\nThis model was trained on the GPU in a similar style as the LGBM.\n\n## Catboost\nThe Catboost was the hardest to get right, especially since no early stopping or callbacks were available for GPU training. This model's predictions have very little weight in the ensemble, however, they brought a significant boost to the final score.\n\n\n### What I tried that did not work:\n- Neural Networks\nCouldn't get it working, and when I thought I did, the public score was abysmal.\n- Autoencoder-Features\n- Cluster aggregations\n- Using a model to predict whether a statement is the last of a customer\n- ....\n\nI learned a lot in this competition, met new friends and got a very pleasant result. Huge success for me.",
    "1915126": "Felicidades Daniel. ¡Buen trabajo!",
    "1915173": "Amazing approach! Keep up the great work @gandagorn !! Hearty congratulations for the result as well!",
    "1915174": "Muchas gracias!",
    "1915180": "Thank you very much!",
    "1915281": "Congrats and thanks for the thorough desrciption! \nCould you please say bit more about Auto encoder features? Was it a DAE? Used both train/test data? In their raw or aggregated (by customer id) form? Thanks!",
    "1915329": "Great work and approach - congrats!",
    "1915354": "Thanks for sharing your approach, flowchart make it clear",
    "1915720": "No only with bottleneck layer, with both train and test in their aggregated form. It didn't improve the model, so I focused on other things.",
    "1918496": "Great post and description. Thanks for sharing this. Learning a lot ..🙏",
    "1920889": "Thanks for sharing this solution! May I invite you to participate in this survey regarding your experience on Kaggle (10 min)? This is not a scam. We are a group of researchers at the City University of Hong Kong. The survey link is: https://cityuhk.questionpro.com/survey-of-kaggle-contestants",
    "1926534": "Thanks for sharing it! Congrats!!",
    "1928216": "Congratulations and thanks for sharing! Same as Jose, happy to see a spanish kaggler (are you spanish, right?😄) in the gold zone. \n\nI have some questions about your approach:\n\n**1.** This one is to be sure that I understood correctly how you calculated some features:\n- *Aggregations for the time between two statements*:  If I am not mistaken time between statements was almost constant (~month). What type of features did you create here?\n- *10 PCA dimensions*: Over which features did you compute the PCA?\n- *Coefficient of variation*: Is this one another aggregated feature by customer_ID? Because it is essentially the same as calculating the std, right?\n\n**2.** My second question is about how you did the FE/FS part. Do you started with standard aggregation features and sequentially adding the rest of the feature subsets (e.g. add all the PCA features all at once)? Did you perform any kind of FS selection method or just adding feature subsets one by one and checking whether the CV/LB scores improved above a threshold?",
    "1929143": "Hi, thanks for your comment! Actually I'm from Austria :D\n- Most of the time it was, but not for all samples, and I wanted to include that information. I used the diff of days between statements and build aggregations (mean, std ...). Overall it wasn't a very important feature class.\n- Over the aggregation features\n- It's the ratio between std and mean. I basically used every pair of features XX_mean and XX_std to create a feature XX_cov.\n- I created the different classes of features and iteratively added them to the training. I did not use much FS, however, I removed classes of features that seemed to decrease the models performance (which was tricky, as the results had a high variance). Went with my gut feeling most of the time."
  },
  "source": "meta"
}