{
  "id": 349250,
  "title": "18th Place Gold",
  "url": "/competitions/amex-default-prediction/discussion/349250",
  "author_name": "MartinBarus",
  "post_date": "2022-08-31T19:42:15.495000",
  "votes": 18,
  "comment_count": 0,
  "views": 0,
  "content": "<h2>Overview</h2>\n<p>This competition got me my first (solo) gold medal, so I am sure you can imagine how happy I am and how much I enjoyed it. Big thanks to the organizers and the kaggle team!</p>\n<p>My work was also based on other people's great effort, so big thanks and shoutout to <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> <a href=\"https://www.kaggle.com/raddar\" target=\"_blank\">@raddar</a> <a href=\"https://www.kaggle.com/ragnar123\" target=\"_blank\">@ragnar123</a> and <a href=\"https://www.kaggle.com/thedevastator\" target=\"_blank\">@thedevastator</a> I will mention all their contributions that I used further.</p>\n<p>I built a 3 stage model - 39 stage one base models, 2 stage two ensemble models and stage 3 is a simple average of the 2 stage two models.</p>\n<h2>First Stage - 39 base models</h2>\n<p>I used same CV strategy for every model <code>fold = argsort(customer_id)%5</code>, for every model I generated out-of-fold predictions and test predictions averaged across folds.</p>\n<h3>1. lightgbm with <a href=\"https://www.kaggle.com/datasets/raddar/amex-data-integer-dtypes-parquet-format\" target=\"_blank\">dataset</a> from <a href=\"https://www.kaggle.com/raddar\" target=\"_blank\">@raddar</a></h3>\n<ul>\n<li>computed simple features, fractions of \"first / last\" feature value and \"mean / last\" value when features are nonzero in train and test, else difference of these combinations</li>\n<li>simple hand tuning of hyper-parameters</li>\n<li>13 models altogether</li>\n</ul>\n<h3>2. <a href=\"https://www.kaggle.com/code/ragnar123/amex-lgbm-dart-cv-0-7977\" target=\"_blank\">lightgbm</a> from <a href=\"https://www.kaggle.com/ragnar123\" target=\"_blank\">@ragnar123</a></h3>\n<ul>\n<li>8 models altogether</li>\n<li>1 original model</li>\n<li>5 models with tuned learning rates</li>\n<li>2 scaled models: original_num_trees * N, original_learning_rate/N for N in [2,4]</li>\n</ul>\n<h3>3. <a href=\"https://www.kaggle.com/code/thedevastator/amex-bruteforce-feature-engineering\" target=\"_blank\">lightgbm</a> from <a href=\"https://www.kaggle.com/thedevastator\" target=\"_blank\">@thedevastator</a></h3>\n<ul>\n<li>3 models altogether </li>\n<li>1 original model</li>\n<li>2 scaled models: original_num_trees * N, original_learning_rate/N for N in [2,4]</li>\n</ul>\n<h3>4. my custom CNN implementation with custom dataset</h3>\n<ul>\n<li>13 models altogether</li>\n<li>different architectures (number of filters and convolution layers)</li>\n</ul>\n<h3>5. <a href=\"https://www.kaggle.com/code/cdeotte/tensorflow-transformer-0-790\" target=\"_blank\">transformer</a> from <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a></h3>\n<ul>\n<li>only 1 original model</li>\n</ul>\n<h3>6. gaussian naive bayes</h3>\n<ul>\n<li>only 1 model, using same dataset as 1. lightgbm</li>\n</ul>\n<h2>Second Stage - 2 MLPRegressors</h2>\n<p>I tried different ensembles of different groups of base models. These were my findings:</p>\n<p>Average of 3 models of the same parameters scaled with factor n=1,2,4 for the public lightgbms (stage 1, models 3 and 4) worked pretty well.</p>\n<p>Average of \"few better models\" and their ensemble showed better CV, but LB improvements didn't correspond.</p>\n<p>This led me to build ensembles using all the models I built, even Naive Bayes with CV 0.55.</p>\n<p>I tried different approaches - ElasticNet, BayesianRidge, LogisticRegression, My custom non-negative linear model (weights are either zero or positive, max 1, add up to 1), KNN, Lightgbm/XGboost and MLPRegressor.</p>\n<p>I found out MLPRegressor works the best, so I ran a random gird search for 100 models: randomly choose number of hidden layers from 1 to 3, for each layer select randomly number of neurons up to 100.</p>\n<p>Since MLPRegressors are non-linear, I tried not only single best models, but also average of few best MLPRegressors.</p>\n<p>My final 2 second stage models were 2 MLPRegressors both with 2 hidden layers, first one had 52 and 94 neurones in the hidden layers, second one had 10 and 20 neurones.</p>\n<h2>Third Stage - simple average</h2>\n<p>I averaged the outputs of the second stage models.</p>\n<p>I saw discrepancy between my CV and LB (higher CV had lower LB score) so I computed CV two ways - mean of the 5 scores of each split: <em>CV1</em>, and single out-of-fold score, where probabilities where min/max scaled per fold: <em>CV2</em>.</p>\n<p>Firstly I selected best LB submission with CV1 0.79861 and CV2 0.79871, which was the better out of the two final submissions, with 0.80833 private and 0.80074 public score.</p>\n<p>For second submission I conservatively selected submission with the highest and most similar CV computed both ways with CV1 0.79928 and CV2 0.79924 which scored 0.80814 on private and 0.8004 on public LB.</p>\n<p>My best submission which I didn't select with CV1 0.79947 and CV2 0.79871 scored 0.80849 on private and 0.80060 on public and was exactly the same as the described solution, with one more MLPRegressor with single hidden layer of size 54.</p>",
  "messages": [
    {
      "id": 1921386,
      "postDate": "2022-08-31T19:42:15.497Z",
      "content": "<h2>Overview</h2>\n<p>This competition got me my first (solo) gold medal, so I am sure you can imagine how happy I am and how much I enjoyed it. Big thanks to the organizers and the kaggle team!</p>\n<p>My work was also based on other people's great effort, so big thanks and shoutout to <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> <a href=\"https://www.kaggle.com/raddar\" target=\"_blank\">@raddar</a> <a href=\"https://www.kaggle.com/ragnar123\" target=\"_blank\">@ragnar123</a> and <a href=\"https://www.kaggle.com/thedevastator\" target=\"_blank\">@thedevastator</a> I will mention all their contributions that I used further.</p>\n<p>I built a 3 stage model - 39 stage one base models, 2 stage two ensemble models and stage 3 is a simple average of the 2 stage two models.</p>\n<h2>First Stage - 39 base models</h2>\n<p>I used same CV strategy for every model <code>fold = argsort(customer_id)%5</code>, for every model I generated out-of-fold predictions and test predictions averaged across folds.</p>\n<h3>1. lightgbm with <a href=\"https://www.kaggle.com/datasets/raddar/amex-data-integer-dtypes-parquet-format\" target=\"_blank\">dataset</a> from <a href=\"https://www.kaggle.com/raddar\" target=\"_blank\">@raddar</a></h3>\n<ul>\n<li>computed simple features, fractions of \"first / last\" feature value and \"mean / last\" value when features are nonzero in train and test, else difference of these combinations</li>\n<li>simple hand tuning of hyper-parameters</li>\n<li>13 models altogether</li>\n</ul>\n<h3>2. <a href=\"https://www.kaggle.com/code/ragnar123/amex-lgbm-dart-cv-0-7977\" target=\"_blank\">lightgbm</a> from <a href=\"https://www.kaggle.com/ragnar123\" target=\"_blank\">@ragnar123</a></h3>\n<ul>\n<li>8 models altogether</li>\n<li>1 original model</li>\n<li>5 models with tuned learning rates</li>\n<li>2 scaled models: original_num_trees * N, original_learning_rate/N for N in [2,4]</li>\n</ul>\n<h3>3. <a href=\"https://www.kaggle.com/code/thedevastator/amex-bruteforce-feature-engineering\" target=\"_blank\">lightgbm</a> from <a href=\"https://www.kaggle.com/thedevastator\" target=\"_blank\">@thedevastator</a></h3>\n<ul>\n<li>3 models altogether </li>\n<li>1 original model</li>\n<li>2 scaled models: original_num_trees * N, original_learning_rate/N for N in [2,4]</li>\n</ul>\n<h3>4. my custom CNN implementation with custom dataset</h3>\n<ul>\n<li>13 models altogether</li>\n<li>different architectures (number of filters and convolution layers)</li>\n</ul>\n<h3>5. <a href=\"https://www.kaggle.com/code/cdeotte/tensorflow-transformer-0-790\" target=\"_blank\">transformer</a> from <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a></h3>\n<ul>\n<li>only 1 original model</li>\n</ul>\n<h3>6. gaussian naive bayes</h3>\n<ul>\n<li>only 1 model, using same dataset as 1. lightgbm</li>\n</ul>\n<h2>Second Stage - 2 MLPRegressors</h2>\n<p>I tried different ensembles of different groups of base models. These were my findings:</p>\n<p>Average of 3 models of the same parameters scaled with factor n=1,2,4 for the public lightgbms (stage 1, models 3 and 4) worked pretty well.</p>\n<p>Average of \"few better models\" and their ensemble showed better CV, but LB improvements didn't correspond.</p>\n<p>This led me to build ensembles using all the models I built, even Naive Bayes with CV 0.55.</p>\n<p>I tried different approaches - ElasticNet, BayesianRidge, LogisticRegression, My custom non-negative linear model (weights are either zero or positive, max 1, add up to 1), KNN, Lightgbm/XGboost and MLPRegressor.</p>\n<p>I found out MLPRegressor works the best, so I ran a random gird search for 100 models: randomly choose number of hidden layers from 1 to 3, for each layer select randomly number of neurons up to 100.</p>\n<p>Since MLPRegressors are non-linear, I tried not only single best models, but also average of few best MLPRegressors.</p>\n<p>My final 2 second stage models were 2 MLPRegressors both with 2 hidden layers, first one had 52 and 94 neurones in the hidden layers, second one had 10 and 20 neurones.</p>\n<h2>Third Stage - simple average</h2>\n<p>I averaged the outputs of the second stage models.</p>\n<p>I saw discrepancy between my CV and LB (higher CV had lower LB score) so I computed CV two ways - mean of the 5 scores of each split: <em>CV1</em>, and single out-of-fold score, where probabilities where min/max scaled per fold: <em>CV2</em>.</p>\n<p>Firstly I selected best LB submission with CV1 0.79861 and CV2 0.79871, which was the better out of the two final submissions, with 0.80833 private and 0.80074 public score.</p>\n<p>For second submission I conservatively selected submission with the highest and most similar CV computed both ways with CV1 0.79928 and CV2 0.79924 which scored 0.80814 on private and 0.8004 on public LB.</p>\n<p>My best submission which I didn't select with CV1 0.79947 and CV2 0.79871 scored 0.80849 on private and 0.80060 on public and was exactly the same as the described solution, with one more MLPRegressor with single hidden layer of size 54.</p>",
      "rawMarkdown": "## Overview\nThis competition got me my first (solo) gold medal, so I am sure you can imagine how happy I am and how much I enjoyed it. Big thanks to the organizers and the kaggle team!\n\nMy work was also based on other people's great effort, so big thanks and shoutout to @cdeotte @raddar @ragnar123 and @thedevastator I will mention all their contributions that I used further.\n\nI built a 3 stage model - 39 stage one base models, 2 stage two ensemble models and stage 3 is a simple average of the 2 stage two models.\n\n\n## First Stage - 39 base models\n\nI used same CV strategy for every model `fold = argsort(customer_id)%5`, for every model I generated out-of-fold predictions and test predictions averaged across folds.\n\n### 1. lightgbm with [dataset](https://www.kaggle.com/datasets/raddar/amex-data-integer-dtypes-parquet-format) from @raddar \n- computed simple features, fractions of \"first / last\" feature value and \"mean / last\" value when features are nonzero in train and test, else difference of these combinations\n- simple hand tuning of hyper-parameters\n- 13 models altogether\n\n### 2. [lightgbm](https://www.kaggle.com/code/ragnar123/amex-lgbm-dart-cv-0-7977) from @ragnar123 \n- 8 models altogether\n- 1 original model\n- 5 models with tuned learning rates\n- 2 scaled models: original_num_trees * N, original_learning_rate/N for N in [2,4]\n\n### 3. [lightgbm](https://www.kaggle.com/code/thedevastator/amex-bruteforce-feature-engineering) from @thedevastator \n- 3 models altogether \n- 1 original model\n- 2 scaled models: original_num_trees * N, original_learning_rate/N for N in [2,4]\n\n\n### 4. my custom CNN implementation with custom dataset\n- 13 models altogether\n- different architectures (number of filters and convolution layers)\n\n\n### 5. [transformer](https://www.kaggle.com/code/cdeotte/tensorflow-transformer-0-790) from @cdeotte \n- only 1 original model\n\n### 6. gaussian naive bayes\n- only 1 model, using same dataset as 1. lightgbm\n\n## Second Stage - 2 MLPRegressors\n\nI tried different ensembles of different groups of base models. These were my findings:\n\nAverage of 3 models of the same parameters scaled with factor n=1,2,4 for the public lightgbms (stage 1, models 3 and 4) worked pretty well.\n\nAverage of \"few better models\" and their ensemble showed better CV, but LB improvements didn't correspond.\n\nThis led me to build ensembles using all the models I built, even Naive Bayes with CV 0.55.\n\nI tried different approaches - ElasticNet, BayesianRidge, LogisticRegression, My custom non-negative linear model (weights are either zero or positive, max 1, add up to 1), KNN, Lightgbm/XGboost and MLPRegressor.\n\nI found out MLPRegressor works the best, so I ran a random gird search for 100 models: randomly choose number of hidden layers from 1 to 3, for each layer select randomly number of neurons up to 100.\n\nSince MLPRegressors are non-linear, I tried not only single best models, but also average of few best MLPRegressors.\n\nMy final 2 second stage models were 2 MLPRegressors both with 2 hidden layers, first one had 52 and 94 neurones in the hidden layers, second one had 10 and 20 neurones.\n\n\n## Third Stage - simple average\n\nI averaged the outputs of the second stage models.\n\nI saw discrepancy between my CV and LB (higher CV had lower LB score) so I computed CV two ways - mean of the 5 scores of each split: *CV1*, and single out-of-fold score, where probabilities where min/max scaled per fold: *CV2*.\n\nFirstly I selected best LB submission with CV1 0.79861 and CV2 0.79871, which was the better out of the two final submissions, with 0.80833 private and 0.80074 public score.\n\nFor second submission I conservatively selected submission with the highest and most similar CV computed both ways with CV1 0.79928 and CV2 0.79924 which scored 0.80814 on private and 0.8004 on public LB.\n\nMy best submission which I didn't select with CV1 0.79947 and CV2 0.79871 scored 0.80849 on private and 0.80060 on public and was exactly the same as the described solution, with one more MLPRegressor with single hidden layer of size 54.\n",
      "votes": 18
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "1921386": "## Overview\nThis competition got me my first (solo) gold medal, so I am sure you can imagine how happy I am and how much I enjoyed it. Big thanks to the organizers and the kaggle team!\n\nMy work was also based on other people's great effort, so big thanks and shoutout to @cdeotte @raddar @ragnar123 and @thedevastator I will mention all their contributions that I used further.\n\nI built a 3 stage model - 39 stage one base models, 2 stage two ensemble models and stage 3 is a simple average of the 2 stage two models.\n\n\n## First Stage - 39 base models\n\nI used same CV strategy for every model `fold = argsort(customer_id)%5`, for every model I generated out-of-fold predictions and test predictions averaged across folds.\n\n### 1. lightgbm with [dataset](https://www.kaggle.com/datasets/raddar/amex-data-integer-dtypes-parquet-format) from @raddar \n- computed simple features, fractions of \"first / last\" feature value and \"mean / last\" value when features are nonzero in train and test, else difference of these combinations\n- simple hand tuning of hyper-parameters\n- 13 models altogether\n\n### 2. [lightgbm](https://www.kaggle.com/code/ragnar123/amex-lgbm-dart-cv-0-7977) from @ragnar123 \n- 8 models altogether\n- 1 original model\n- 5 models with tuned learning rates\n- 2 scaled models: original_num_trees * N, original_learning_rate/N for N in [2,4]\n\n### 3. [lightgbm](https://www.kaggle.com/code/thedevastator/amex-bruteforce-feature-engineering) from @thedevastator \n- 3 models altogether \n- 1 original model\n- 2 scaled models: original_num_trees * N, original_learning_rate/N for N in [2,4]\n\n\n### 4. my custom CNN implementation with custom dataset\n- 13 models altogether\n- different architectures (number of filters and convolution layers)\n\n\n### 5. [transformer](https://www.kaggle.com/code/cdeotte/tensorflow-transformer-0-790) from @cdeotte \n- only 1 original model\n\n### 6. gaussian naive bayes\n- only 1 model, using same dataset as 1. lightgbm\n\n## Second Stage - 2 MLPRegressors\n\nI tried different ensembles of different groups of base models. These were my findings:\n\nAverage of 3 models of the same parameters scaled with factor n=1,2,4 for the public lightgbms (stage 1, models 3 and 4) worked pretty well.\n\nAverage of \"few better models\" and their ensemble showed better CV, but LB improvements didn't correspond.\n\nThis led me to build ensembles using all the models I built, even Naive Bayes with CV 0.55.\n\nI tried different approaches - ElasticNet, BayesianRidge, LogisticRegression, My custom non-negative linear model (weights are either zero or positive, max 1, add up to 1), KNN, Lightgbm/XGboost and MLPRegressor.\n\nI found out MLPRegressor works the best, so I ran a random gird search for 100 models: randomly choose number of hidden layers from 1 to 3, for each layer select randomly number of neurons up to 100.\n\nSince MLPRegressors are non-linear, I tried not only single best models, but also average of few best MLPRegressors.\n\nMy final 2 second stage models were 2 MLPRegressors both with 2 hidden layers, first one had 52 and 94 neurones in the hidden layers, second one had 10 and 20 neurones.\n\n\n## Third Stage - simple average\n\nI averaged the outputs of the second stage models.\n\nI saw discrepancy between my CV and LB (higher CV had lower LB score) so I computed CV two ways - mean of the 5 scores of each split: *CV1*, and single out-of-fold score, where probabilities where min/max scaled per fold: *CV2*.\n\nFirstly I selected best LB submission with CV1 0.79861 and CV2 0.79871, which was the better out of the two final submissions, with 0.80833 private and 0.80074 public score.\n\nFor second submission I conservatively selected submission with the highest and most similar CV computed both ways with CV1 0.79928 and CV2 0.79924 which scored 0.80814 on private and 0.8004 on public LB.\n\nMy best submission which I didn't select with CV1 0.79947 and CV2 0.79871 scored 0.80849 on private and 0.80060 on public and was exactly the same as the described solution, with one more MLPRegressor with single hidden layer of size 54.\n"
  }
}