{
  "id": 347882,
  "title": "Silver (211 th) Long story short",
  "url": "/competitions/amex-default-prediction/writeups/slightly-conscious-silver-211-th-long-story-short",
  "author_name": "",
  "post_date": "2023-06-24T04:40:40.050Z",
  "votes": 20,
  "comment_count": 3,
  "views": 0,
  "content": "<p>Thank you all and congrats to the winners!</p>\n<p>My solution is very simple and straight-forward:</p>\n<ul>\n<li>Created ~6000 features: aggregations, diff(1 to 5), moving average and statistics, lags divisions, trends, many interactions among the top 100 features, KMeans, PCA etc.</li>\n<li>Selected features by permutation importances on the validation set (ended up with ~5 different features sets with ~650 to ~1400 features)</li>\n<li>Used XGB, LGBM, Catboost models with different preprocessing strategies, feature sets, hyperparameters</li>\n<li>Ensembled everything with different seeds (simple average of ~40 models)</li>\n<li>XGB was always the best model for me</li>\n<li>time spent: 10% struggling with data volume, 10% algorithms and ensemble, 80% feature engineering and selection</li>\n</ul>\n<p>Single models metrics (5 fold):</p>\n<ul>\n<li>amex ~ 0.797xx</li>\n<li>roc_auc ~ 0.877xx</li>\n<li>accuracy ~ 0.905xx</li>\n<li>f1 ~ 0.816xx</li>\n<li>precision ~ 0.813xx</li>\n<li>recall ~ 0.819xx</li>\n</ul>\n<p>What did not work:</p>\n<ul>\n<li>I could not get dart boosting to outperform gbtree on my setup</li>\n<li>Stacking features</li>\n<li>Knowledge distillation</li>\n<li>OOF and confusion matrix analysis (false negatives and false positives): could not find any insight that I could be sure it led to better performance (I blame the anonymized features for that 😆)</li>\n<li>Denoising autoencoder (not sure if I implemented it correctly)</li>\n<li>TabNet</li>\n<li>GRU</li>\n<li>MLP (keras)</li>\n</ul>",
  "messages": [
    {
      "id": "1914150",
      "postDate": "08/25/2022 19:18:17",
      "content": "<p>Thank you all and congrats to the winners!</p>\n<p>My solution is very simple and straight-forward:</p>\n<ul>\n<li>Created ~6000 features: aggregations, diff(1 to 5), moving average and statistics, lags divisions, trends, many interactions among the top 100 features, KMeans, PCA etc.</li>\n<li>Selected features by permutation importances on the validation set (ended up with ~5 different features sets with ~650 to ~1400 features)</li>\n<li>Used XGB, LGBM, Catboost models with different preprocessing strategies, feature sets, hyperparameters</li>\n<li>Ensembled everything with different seeds (simple average of ~40 models)</li>\n<li>XGB was always the best model for me</li>\n<li>time spent: 10% struggling with data volume, 10% algorithms and ensemble, 80% feature engineering and selection</li>\n</ul>\n<p>Single models metrics (5 fold):</p>\n<ul>\n<li>amex ~ 0.797xx</li>\n<li>roc_auc ~ 0.877xx</li>\n<li>accuracy ~ 0.905xx</li>\n<li>f1 ~ 0.816xx</li>\n<li>precision ~ 0.813xx</li>\n<li>recall ~ 0.819xx</li>\n</ul>\n<p>What did not work:</p>\n<ul>\n<li>I could not get dart boosting to outperform gbtree on my setup</li>\n<li>Stacking features</li>\n<li>Knowledge distillation</li>\n<li>OOF and confusion matrix analysis (false negatives and false positives): could not find any insight that I could be sure it led to better performance (I blame the anonymized features for that 😆)</li>\n<li>Denoising autoencoder (not sure if I implemented it correctly)</li>\n<li>TabNet</li>\n<li>GRU</li>\n<li>MLP (keras)</li>\n</ul>",
      "rawMarkdown": "Thank you all and congrats to the winners!\n\nMy solution is very simple and straight-forward:\n\n- Created ~6000 features: aggregations, diff(1 to 5), moving average and statistics, lags divisions, trends, many interactions among the top 100 features, KMeans, PCA etc.\n- Selected features by permutation importances on the validation set (ended up with ~5 different features sets with ~650 to ~1400 features)\n- Used XGB, LGBM, Catboost models with different preprocessing strategies, feature sets, hyperparameters\n- Ensembled everything with different seeds (simple average of ~40 models)\n- XGB was always the best model for me\n- time spent: 10% struggling with data volume, 10% algorithms and ensemble, 80% feature engineering and selection\n\nSingle models metrics (5 fold):\n- amex ~ 0.797xx\n- roc_auc ~ 0.877xx\n- accuracy ~ 0.905xx\n- f1 ~ 0.816xx\n- precision ~ 0.813xx\n- recall ~ 0.819xx\n\nWhat did not work:\n\n- I could not get dart boosting to outperform gbtree on my setup\n- Stacking features\n- Knowledge distillation\n- OOF and confusion matrix analysis (false negatives and false positives): could not find any insight that I could be sure it led to better performance (I blame the anonymized features for that 😆)\n- Denoising autoencoder (not sure if I implemented it correctly)\n- TabNet\n- GRU\n- MLP (keras)",
      "votes": null
    },
    {
      "id": "1914302",
      "postDate": "08/26/2022 01:23:43",
      "content": "<p>Congratulations !!</p>",
      "rawMarkdown": "Congratulations !!",
      "votes": null
    },
    {
      "id": "1914452",
      "postDate": "08/26/2022 05:34:45",
      "content": "<p>Hearty congratulations for the results and thanks for a detailed summary of the approach!!</p>",
      "rawMarkdown": "Hearty congratulations for the results and thanks for a detailed summary of the approach!!",
      "votes": null
    },
    {
      "id": "1922127",
      "postDate": "09/01/2022 09:30:31",
      "content": "<p>Congratulations <a href=\"https://www.kaggle.com/hinepo\" target=\"_blank\">@hinepo</a> May I invite you to participate in this survey regarding your experience on Kaggle (10 min)? This is not a scam. We are a group of researchers at the City University of Hong Kong. The survey link is: <a href=\"https://cityuhk.questionpro.com/survey-of-kaggle-contestants\" target=\"_blank\">https://cityuhk.questionpro.com/survey-of-kaggle-contestants</a></p>",
      "rawMarkdown": "Congratulations @hinepo May I invite you to participate in this survey regarding your experience on Kaggle (10 min)? This is not a scam. We are a group of researchers at the City University of Hong Kong. The survey link is: https://cityuhk.questionpro.com/survey-of-kaggle-contestants",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1914302,
      "author_name": "gabrielvinicius",
      "author_url": "",
      "post_date": "08/26/2022 01:23:43",
      "content": "<p>Congratulations !!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1914452,
      "author_name": "ravi20076",
      "author_url": "",
      "post_date": "08/26/2022 05:34:45",
      "content": "<p>Hearty congratulations for the results and thanks for a detailed summary of the approach!!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1922127,
      "author_name": "lystriving",
      "author_url": "",
      "post_date": "09/01/2022 09:30:31",
      "content": "<p>Congratulations <a href=\"https://www.kaggle.com/hinepo\" target=\"_blank\">@hinepo</a> May I invite you to participate in this survey regarding your experience on Kaggle (10 min)? This is not a scam. We are a group of researchers at the City University of Hong Kong. The survey link is: <a href=\"https://cityuhk.questionpro.com/survey-of-kaggle-contestants\" target=\"_blank\">https://cityuhk.questionpro.com/survey-of-kaggle-contestants</a></p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1914150": "Thank you all and congrats to the winners!\n\nMy solution is very simple and straight-forward:\n\n- Created ~6000 features: aggregations, diff(1 to 5), moving average and statistics, lags divisions, trends, many interactions among the top 100 features, KMeans, PCA etc.\n- Selected features by permutation importances on the validation set (ended up with ~5 different features sets with ~650 to ~1400 features)\n- Used XGB, LGBM, Catboost models with different preprocessing strategies, feature sets, hyperparameters\n- Ensembled everything with different seeds (simple average of ~40 models)\n- XGB was always the best model for me\n- time spent: 10% struggling with data volume, 10% algorithms and ensemble, 80% feature engineering and selection\n\nSingle models metrics (5 fold):\n- amex ~ 0.797xx\n- roc_auc ~ 0.877xx\n- accuracy ~ 0.905xx\n- f1 ~ 0.816xx\n- precision ~ 0.813xx\n- recall ~ 0.819xx\n\nWhat did not work:\n\n- I could not get dart boosting to outperform gbtree on my setup\n- Stacking features\n- Knowledge distillation\n- OOF and confusion matrix analysis (false negatives and false positives): could not find any insight that I could be sure it led to better performance (I blame the anonymized features for that 😆)\n- Denoising autoencoder (not sure if I implemented it correctly)\n- TabNet\n- GRU\n- MLP (keras)",
    "1914302": "Congratulations !!",
    "1914452": "Hearty congratulations for the results and thanks for a detailed summary of the approach!!",
    "1922127": "Congratulations @hinepo May I invite you to participate in this survey regarding your experience on Kaggle (10 min)? This is not a scam. We are a group of researchers at the City University of Hong Kong. The survey link is: https://cityuhk.questionpro.com/survey-of-kaggle-contestants"
  },
  "source": "meta"
}