{
  "id": 350538,
  "title": "9th Place Solution ( XGBoost+LGBM+NN )",
  "url": "/competitions/amex-default-prediction/writeups/george-reus-9th-place-solution-xgboost-lgbm-nn",
  "author_name": "",
  "post_date": "2022-09-06T08:08:18.446536200Z",
  "votes": 32,
  "comment_count": 8,
  "views": 0,
  "content": "<p>Many thanks to AMEX,Kaggle and all contributors of discussion during the entire competition ( <a href=\"https://www.kaggle.com/raddar\" target=\"_blank\">@raddar</a> , <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> , <a href=\"https://www.kaggle.com/ragnar123\" target=\"_blank\">@ragnar123</a> ,…). Congratulations to all winners and new Experts, Masters and Grandmaster!</p>\n<h2>Score &amp; Result</h2>\n<p>My best submission<br>\nCV: 0.799106        Public: 0.80062        Private:0.80875<br>\nMy result<br>\nCV: 0.799194        Public: 0.80057        Private:0.80868</p>\n<h1>Feature Engineering</h1>\n<p>I only used the integer dataset provided by <a href=\"https://www.kaggle.com/raddar\" target=\"_blank\">@raddar</a>. </p>\n<ul>\n<li>Base features<br>\naggregated features like mean,max,min,std,sum,medium,last,first</li>\n<li>other rate and diff features with datediff<br>\nlast-first,last1-last2,last1-last3,last-mean,max/last,sum/last  and so on</li>\n<li>Date features<br>\nis it a holiday </li>\n</ul>\n<h1>Model</h1>\n<ul>\n<li>LightGBM <br>\n3 models of LGBM with different data representation &amp; parameters give CV in the range [0.796-0.799] and LB in [0.797-0.799]  (2 model with dart-LGBM , 1 model with goss-LGBM)</li>\n<li>XGBoost<br>\n6 models of XGB, with different data representation &amp; parameters give CV in the range [0.794-0.796] and LB in [0.795-0.796]</li>\n<li>NeuralNet<br>\n4 models  of NeuralNet  with different parameters give CV in the range<br>\n[0.788-0.790] and LB in [0.790-0.792]<br>\n(I am not so proud of NNs &nbsp;Thank again&nbsp;@cdeotte for sharing his great public kernel NN.)</li>\n</ul>\n<h1>Ensemble(stacking)</h1>\n<p>Using 13 models to stack with 10-fold cross-validation , Hyperparameter-tuning and appropriate early stopping can give   Private in the range [0.80853-0.80875] </p>\n<h1>some ideas</h1>\n<p>Predict if a customer will default  when customers have already used credit to consume ,trending of the features changed over time(like consumption frequency ,Change in consumption amount) is very important ,especially in the last few months.</p>\n<p>I am very grateful to this competition. I learned a lot in this competition for a newbie in kaggle .Thanks you all😎</p>",
  "messages": [
    {
      "id": "1928061",
      "postDate": "09/06/2022 08:08:18",
      "content": "<p>Many thanks to AMEX,Kaggle and all contributors of discussion during the entire competition ( <a href=\"https://www.kaggle.com/raddar\" target=\"_blank\">@raddar</a> , <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> , <a href=\"https://www.kaggle.com/ragnar123\" target=\"_blank\">@ragnar123</a> ,…). Congratulations to all winners and new Experts, Masters and Grandmaster!</p>\n<h2>Score &amp; Result</h2>\n<p>My best submission<br>\nCV: 0.799106        Public: 0.80062        Private:0.80875<br>\nMy result<br>\nCV: 0.799194        Public: 0.80057        Private:0.80868</p>\n<h1>Feature Engineering</h1>\n<p>I only used the integer dataset provided by <a href=\"https://www.kaggle.com/raddar\" target=\"_blank\">@raddar</a>. </p>\n<ul>\n<li>Base features<br>\naggregated features like mean,max,min,std,sum,medium,last,first</li>\n<li>other rate and diff features with datediff<br>\nlast-first,last1-last2,last1-last3,last-mean,max/last,sum/last  and so on</li>\n<li>Date features<br>\nis it a holiday </li>\n</ul>\n<h1>Model</h1>\n<ul>\n<li>LightGBM <br>\n3 models of LGBM with different data representation &amp; parameters give CV in the range [0.796-0.799] and LB in [0.797-0.799]  (2 model with dart-LGBM , 1 model with goss-LGBM)</li>\n<li>XGBoost<br>\n6 models of XGB, with different data representation &amp; parameters give CV in the range [0.794-0.796] and LB in [0.795-0.796]</li>\n<li>NeuralNet<br>\n4 models  of NeuralNet  with different parameters give CV in the range<br>\n[0.788-0.790] and LB in [0.790-0.792]<br>\n(I am not so proud of NNs &nbsp;Thank again&nbsp;@cdeotte for sharing his great public kernel NN.)</li>\n</ul>\n<h1>Ensemble(stacking)</h1>\n<p>Using 13 models to stack with 10-fold cross-validation , Hyperparameter-tuning and appropriate early stopping can give   Private in the range [0.80853-0.80875] </p>\n<h1>some ideas</h1>\n<p>Predict if a customer will default  when customers have already used credit to consume ,trending of the features changed over time(like consumption frequency ,Change in consumption amount) is very important ,especially in the last few months.</p>\n<p>I am very grateful to this competition. I learned a lot in this competition for a newbie in kaggle .Thanks you all😎</p>",
      "rawMarkdown": "Many thanks to AMEX,Kaggle and all contributors of discussion during the entire competition ( @raddar , @cdeotte , @ragnar123 ,…). Congratulations to all winners and new Experts, Masters and Grandmaster!\n## Score & Result\nMy best submission\nCV: 0.799106        Public: 0.80062        Private:0.80875\nMy result\nCV: 0.799194        Public: 0.80057        Private:0.80868\n# Feature Engineering\nI only used the integer dataset provided by @raddar. \n- Base features\naggregated features like mean,max,min,std,sum,medium,last,first\n- other rate and diff features with datediff\nlast-first,last1-last2,last1-last3,last-mean,max/last,sum/last  and so on\n- Date features\n is it a holiday \n# Model\n- LightGBM \n3 models of LGBM with different data representation & parameters give CV in the range [0.796-0.799] and LB in [0.797-0.799]  (2 model with dart-LGBM , 1 model with goss-LGBM)\n- XGBoost\n6 models of XGB, with different data representation & parameters give CV in the range [0.794-0.796] and LB in [0.795-0.796]\n-  NeuralNet\n4 models  of NeuralNet  with different parameters give CV in the range\n[0.788-0.790] and LB in [0.790-0.792]\n(I am not so proud of NNs  Thank again @cdeotte for sharing his great public kernel NN.)\n# Ensemble(stacking)\nUsing 13 models to stack with 10-fold cross-validation , Hyperparameter-tuning and appropriate early stopping can give   Private in the range [0.80853-0.80875] \n# some ideas\nPredict if a customer will default  when customers have already used credit to consume ,trending of the features changed over time(like consumption frequency ,Change in consumption amount) is very important ,especially in the last few months.\n\nI am very grateful to this competition. I learned a lot in this competition for a newbie in kaggle .Thanks you all😎",
      "votes": null
    },
    {
      "id": "1928086",
      "postDate": "09/06/2022 08:41:44",
      "content": "<p>Well done. Congrats</p>",
      "rawMarkdown": "Well done. Congrats",
      "votes": null
    },
    {
      "id": "1928270",
      "postDate": "09/06/2022 11:29:13",
      "content": "<p>Congratulations for the gold in your first Kaggle comp! </p>\n<p>How did you find useful features besides the base features? I started with almost the same base features as you (Martin's notebook features) but was sooo hard to find new features that improved the existing ones. On the other hand, what are the differences between different data representations?</p>\n<p>Thanks!</p>",
      "rawMarkdown": "Congratulations for the gold in your first Kaggle comp! \n\nHow did you find useful features besides the base features? I started with almost the same base features as you (Martin's notebook features) but was sooo hard to find new features that improved the existing ones. On the other hand, what are the differences between different data representations?\n\nThanks!",
      "votes": null
    },
    {
      "id": "1929316",
      "postDate": "09/07/2022 01:33:05",
      "content": "<p>Thanks delai!<br>\nI didn't use all features in one big model, because it's difficult to find useful new features from feature importance and mabye some new features is useful but the result is not significantly improved. Actually, every single-model is created by different base features and other new features. For example, XGBoost-model-1 is created by base features like 'mean', 'max', 'sum' and diff/rate/date features between first month and last month, XGBoost-model-2 is created by base features like 'medium','max,'min' and diff/rate/date features between the last two months.So every model has different data representations and some new useful features can give a relatively obvious improvemen.</p>",
      "rawMarkdown": "Thanks delai!\nI didn't use all features in one big model, because it's difficult to find useful new features from feature importance and mabye some new features is useful but the result is not significantly improved. Actually, every single-model is created by different base features and other new features. For example, XGBoost-model-1 is created by base features like 'mean', 'max', 'sum' and diff/rate/date features between first month and last month, XGBoost-model-2 is created by base features like 'medium','max,'min' and diff/rate/date features between the last two months.So every model has different data representations and some new useful features can give a relatively obvious improvemen.",
      "votes": null
    },
    {
      "id": "1929320",
      "postDate": "09/07/2022 01:35:25",
      "content": "<p>Thanks Santiago !</p>",
      "rawMarkdown": "Thanks Santiago !",
      "votes": null
    },
    {
      "id": "1929710",
      "postDate": "09/07/2022 09:53:41",
      "content": "<p>congrats…great job</p>",
      "rawMarkdown": "congrats...great job",
      "votes": null
    },
    {
      "id": "1930664",
      "postDate": "09/08/2022 05:41:18",
      "content": "<p>Congrats 🎉<br>\nI learned how important it is to create various models with different features from your solution. Actually, I could not make the most of stacking successfully in this competition.<br>\nI am very curious about</p>\n<ol>\n<li>What was the CV gain when you added \"is it holiday feature\"?</li>\n<li>What was the model for 2nd stage stacking? Was it Ridge or Lasso?<br>\nThanks in advance!</li>\n</ol>",
      "rawMarkdown": "Congrats 🎉\nI learned how important it is to create various models with different features from your solution. Actually, I could not make the most of stacking successfully in this competition.\nI am very curious about\n1. What was the CV gain when you added \"is it holiday feature\"?\n2. What was the model for 2nd stage stacking? Was it Ridge or Lasso?\nThanks in advance!",
      "votes": null
    },
    {
      "id": "1931279",
      "postDate": "09/08/2022 15:18:41",
      "content": "<p>Great job!</p>",
      "rawMarkdown": "Great job!",
      "votes": null
    },
    {
      "id": "1931737",
      "postDate": "09/09/2022 01:34:46",
      "content": "<p>Thanks Shibata !</p>\n<ol>\n<li>Using Holiday features can  give about +0.0002 boost in CV.</li>\n<li>The model is LinearRegression for stacking with rounded oof score ( oof=oof.round(5) )</li>\n</ol>",
      "rawMarkdown": "Thanks Shibata !\n1. Using Holiday features can  give about +0.0002 boost in CV.\n2. The model is LinearRegression for stacking with rounded oof score ( oof=oof.round(5) )",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1928086,
      "author_name": "santiagomota",
      "author_url": "",
      "post_date": "09/06/2022 08:41:44",
      "content": "<p>Well done. Congrats</p>",
      "votes": null,
      "replies": [
        {
          "id": 1929320,
          "author_name": "dodwww",
          "author_url": "",
          "post_date": "09/07/2022 01:35:25",
          "content": "<p>Thanks Santiago !</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1928270,
      "author_name": "delai50",
      "author_url": "",
      "post_date": "09/06/2022 11:29:13",
      "content": "<p>Congratulations for the gold in your first Kaggle comp! </p>\n<p>How did you find useful features besides the base features? I started with almost the same base features as you (Martin's notebook features) but was sooo hard to find new features that improved the existing ones. On the other hand, what are the differences between different data representations?</p>\n<p>Thanks!</p>",
      "votes": null,
      "replies": [
        {
          "id": 1929316,
          "author_name": "dodwww",
          "author_url": "",
          "post_date": "09/07/2022 01:33:05",
          "content": "<p>Thanks delai!<br>\nI didn't use all features in one big model, because it's difficult to find useful new features from feature importance and mabye some new features is useful but the result is not significantly improved. Actually, every single-model is created by different base features and other new features. For example, XGBoost-model-1 is created by base features like 'mean', 'max', 'sum' and diff/rate/date features between first month and last month, XGBoost-model-2 is created by base features like 'medium','max,'min' and diff/rate/date features between the last two months.So every model has different data representations and some new useful features can give a relatively obvious improvemen.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1929710,
      "author_name": "amirdhavarshinis",
      "author_url": "",
      "post_date": "09/07/2022 09:53:41",
      "content": "<p>congrats…great job</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1930664,
      "author_name": "yutoshibata",
      "author_url": "",
      "post_date": "09/08/2022 05:41:18",
      "content": "<p>Congrats 🎉<br>\nI learned how important it is to create various models with different features from your solution. Actually, I could not make the most of stacking successfully in this competition.<br>\nI am very curious about</p>\n<ol>\n<li>What was the CV gain when you added \"is it holiday feature\"?</li>\n<li>What was the model for 2nd stage stacking? Was it Ridge or Lasso?<br>\nThanks in advance!</li>\n</ol>",
      "votes": null,
      "replies": [
        {
          "id": 1931737,
          "author_name": "dodwww",
          "author_url": "",
          "post_date": "09/09/2022 01:34:46",
          "content": "<p>Thanks Shibata !</p>\n<ol>\n<li>Using Holiday features can  give about +0.0002 boost in CV.</li>\n<li>The model is LinearRegression for stacking with rounded oof score ( oof=oof.round(5) )</li>\n</ol>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1931279,
      "author_name": "ruslansadykov",
      "author_url": "",
      "post_date": "09/08/2022 15:18:41",
      "content": "<p>Great job!</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1928061": "Many thanks to AMEX,Kaggle and all contributors of discussion during the entire competition ( @raddar , @cdeotte , @ragnar123 ,…). Congratulations to all winners and new Experts, Masters and Grandmaster!\n## Score & Result\nMy best submission\nCV: 0.799106        Public: 0.80062        Private:0.80875\nMy result\nCV: 0.799194        Public: 0.80057        Private:0.80868\n# Feature Engineering\nI only used the integer dataset provided by @raddar. \n- Base features\naggregated features like mean,max,min,std,sum,medium,last,first\n- other rate and diff features with datediff\nlast-first,last1-last2,last1-last3,last-mean,max/last,sum/last  and so on\n- Date features\n is it a holiday \n# Model\n- LightGBM \n3 models of LGBM with different data representation & parameters give CV in the range [0.796-0.799] and LB in [0.797-0.799]  (2 model with dart-LGBM , 1 model with goss-LGBM)\n- XGBoost\n6 models of XGB, with different data representation & parameters give CV in the range [0.794-0.796] and LB in [0.795-0.796]\n-  NeuralNet\n4 models  of NeuralNet  with different parameters give CV in the range\n[0.788-0.790] and LB in [0.790-0.792]\n(I am not so proud of NNs  Thank again @cdeotte for sharing his great public kernel NN.)\n# Ensemble(stacking)\nUsing 13 models to stack with 10-fold cross-validation , Hyperparameter-tuning and appropriate early stopping can give   Private in the range [0.80853-0.80875] \n# some ideas\nPredict if a customer will default  when customers have already used credit to consume ,trending of the features changed over time(like consumption frequency ,Change in consumption amount) is very important ,especially in the last few months.\n\nI am very grateful to this competition. I learned a lot in this competition for a newbie in kaggle .Thanks you all😎",
    "1928086": "Well done. Congrats",
    "1928270": "Congratulations for the gold in your first Kaggle comp! \n\nHow did you find useful features besides the base features? I started with almost the same base features as you (Martin's notebook features) but was sooo hard to find new features that improved the existing ones. On the other hand, what are the differences between different data representations?\n\nThanks!",
    "1929316": "Thanks delai!\nI didn't use all features in one big model, because it's difficult to find useful new features from feature importance and mabye some new features is useful but the result is not significantly improved. Actually, every single-model is created by different base features and other new features. For example, XGBoost-model-1 is created by base features like 'mean', 'max', 'sum' and diff/rate/date features between first month and last month, XGBoost-model-2 is created by base features like 'medium','max,'min' and diff/rate/date features between the last two months.So every model has different data representations and some new useful features can give a relatively obvious improvemen.",
    "1929320": "Thanks Santiago !",
    "1929710": "congrats...great job",
    "1930664": "Congrats 🎉\nI learned how important it is to create various models with different features from your solution. Actually, I could not make the most of stacking successfully in this competition.\nI am very curious about\n1. What was the CV gain when you added \"is it holiday feature\"?\n2. What was the model for 2nd stage stacking? Was it Ridge or Lasso?\nThanks in advance!",
    "1931279": "Great job!",
    "1931737": "Thanks Shibata !\n1. Using Holiday features can  give about +0.0002 boost in CV.\n2. The model is LinearRegression for stacking with rounded oof score ( oof=oof.round(5) )"
  },
  "source": "meta"
}