{
  "id": 348530,
  "title": "21st Solution and Code Sharing",
  "url": "/competitions/amex-default-prediction/discussion/348530",
  "author_name": "Joe Eddy",
  "post_date": "2022-08-28T21:33:11.852000",
  "votes": 57,
  "comment_count": 18,
  "views": 0,
  "content": "<p>A huge huge thank you to <a href=\"https://www.kaggle.com/andrew60909\" target=\"_blank\">@andrew60909</a> and <a href=\"https://www.kaggle.com/ryotak12\" target=\"_blank\">@ryotak12</a> for an excellent team experience. We learned a lot and had fun, and I appreciate their patience with my extended, largely unsuccessful efforts to make a transformer work really well 😂. Thanks also to Amex and Kaggle for hosting this interesting competition and to <a href=\"https://www.kaggle.com/raddar\" target=\"_blank\">@raddar</a> and everyone else who shared resources that made this competition much more accessible. </p>\n<p>I've made my personal code public here: <a href=\"https://github.com/JEddy92/amex_default_kaggle\" target=\"_blank\">https://github.com/JEddy92/amex_default_kaggle</a>. I tried to create a reasonably general structure for division of FE and modeling, with reusable functionality for training and logging results across a variety of models. There is still some clean up to do and it's not as polished as would be ideal!</p>\n<p>Feel free to ask us anything if you have questions about our approach!</p>\n<h2>Solution Summary</h2>\n<p>We used logistic regression to ensemble ~60 sets of predictions with diversity obtained from different models, feature sets, and data views. Models included the usual mix of gradient boosters (LGBM, LGBM dart, XGBoost, CatBoost), MLPs, TabNet, and Transformers with a few special tricks. Data views included aggregated data, fully flattened data (13 raw features per customer), sequential data (Transformer), and augmented data (shifted statements forward by 1 and appended to original). Very detailed feature information to follow -- they are split across our 3 different team members' different models, but this aims to be a complete compilation of what we all used. </p>\n<hr>\n<h2>Datasets &amp; Notation</h2>\n<ul>\n<li>Everything is derived from Raddar's dataset</li>\n</ul>\n<p>\\(\\mathcal{D}_{fulltrain}\\) : Original full size train-set (5531451 rows)   </p>\n<p>\\(\\mathcal{D}_{train}\\) : Train-set (458913 rows)    </p>\n<p>\\(\\mathcal{D}_{fulltest}\\) : Original full size test-set   </p>\n<p>\\(\\mathcal{D}_{test}\\) : Test-set  </p>\n<hr>\n<h2>Primary Features</h2>\n<p>The train-sets with the best CV score for each of us are generated by \\(\\mathcal{D}_{fulltrain}\\) with these features:  </p>\n<ul>\n<li><p><strong>Numeric</strong>     </p>\n<ul>\n<li>Mean, std, max, min, first, last  </li>\n<li>Last-mean, last/mean, last/first, last/std, max/min, linear-weighted mean </li>\n<li>Similar aggregates as above but limited to last 3, 5, etc. statements</li>\n<li>Recent diff features and aggregates of all diffs </li>\n<li>Monthly-based ranking</li>\n<li>Multiplication and Division: create all possible pairs for all features, except for categorical features and features highly correlated with other features. Multiplication and division is done after scaling with MinMaxScaler(0-1) and clipping outliers</li>\n<li>Null aggregates (including starting streak vs. at random)</li>\n<li>Date features (e.g. stats on customer's statement gaps)</li></ul></li>\n<li><p><strong>Trend</strong> :  Oldest month as 1 and the newest month as 13, take the average over the following periods   </p>\n<ul>\n<li>mean(13,12) - mean(11,10)</li>\n<li>mean(13,12,11) - mean(10,9,8)</li>\n<li>mean(13,12,11) - mean(3,2,1)</li>\n<li>mean(13,12,11,10,9,8) - mean(7,6,5,4,3,2)</li></ul></li>\n<li><p><strong>Category</strong>  </p>\n<ul>\n<li>first, last  </li>\n<li>entropy: Shannon entropy on frequency table   </li>\n<li>Nunique: Number of unique category values</li>\n<li>Counts of each value in each category (both raw counts and tf-idf counts)</li>\n<li><code>P_2</code> mean (across train and test) of each category</li>\n<li>SVD factorization over time of each category</li>\n<li>freq1name : the number of most frequent category  </li>\n<li>freq1ratio : the number of most frequent category / group size  </li>\n<li>freq_last1name : the number of least frequent category  </li>\n<li>freq_last1ratio : the number of least frequent category / group size  </li></ul></li>\n<li><p><strong>PCA</strong></p>\n<ul>\n<li>(N, 13*raw_features) → (N, 64)</li></ul></li>\n<li><p><strong>KNN-based target encoding methods</strong></p>\n<ul>\n<li>We selected the nearest 500 points for each sample with euclidean distance(with 2 groups of most important features, each has 14 features. One with last statement of each customer, another one with mean of all statements), calculate the average of target as features</li>\n<li>Follow the above method, using nearest 500 points but applying average with weight by the distance matrix</li>\n<li>Aggregate the distance matrix, by mean, max, min, std</li>\n<li>We also calculated the cosine similarity of each customer's past N months of data in one dimension.(N=1,3,6,13) The average of the 500 nearest neighbor customers' targets is used as the feature. This feature was binned because the distribution is somewhat different between Train and Test</li></ul></li>\n<li><p><strong>Nested model</strong></p>\n<ul>\n<li><p>We added the labels to \\(\\mathcal{D}_{fulltrain}\\), then trained a LGBM model. The purpose of doing this is to capture what kinds of records, behavior and attributes will cause default. Instead of aggregating features into 458913 rows, this method can let the model learn some extra information from a more \"base\" level. After that, for each <code>customer_ID</code> we will have a \"predicted target\" that has the same length as the number of records in each <code>customer_ID</code>.   </p></li>\n<li><p>Finally, to merge it into \\(\\mathcal{D}_{train}\\) : Train-set (458913 rows), we aggregate similarly to other features: mean, std, max, min, first, last, man-min, last-mean, last/std.  </p></li>\n<li><p>An interesting point here is that the last value of our \"predicted value\" with \\(\\mathcal{D}_{fulltrain}\\) scores 0.785 on the Amex metric, while the score of the max value is around 0.62 -- indicating the importance of the last value.</p></li></ul></li>\n</ul>\n<hr>       \n<h2>Other Feature Methods:</h2>\n<ul>\n<li><p><strong>Target/Count encoding</strong>: On \\(\\mathcal{D}_{train}\\)</p></li>\n<li><p><strong>Target encoding</strong> with \\(\\mathcal{D}_{fulltrain}\\)</p></li>\n<li><p><strong>Other functions to aggregate</strong> \\(\\mathcal{D}_{fulltrain}\\)</p>\n<ul>\n<li>Exponential weight average</li>\n<li>Skewness, Kurtosis</li>\n<li>The ratio beyond 1std </li>\n<li>Median Absolute Deviation</li>\n<li>Mean, max, min, std on 1st derivative(t=1,2,…13), timestamp of max and min also used </li>\n<li>Shannon entropy: estimate of the spectral density of sequence</li>\n<li>Stability, Lumpiness: variance of the means and variance of the variances on tiles of windows</li>\n<li>Crossing points: number of times a sequence crosses the median line</li>\n<li>Some measurements from <a href=\"http://isadoranun.github.io/tsfeat/FeaturesDocumentation.html\" target=\"_blank\">here</a> </li></ul></li>\n<li><p><strong>Clustering methods</strong>: Kmeans, DBSCAN on important feature(e.g. <code>P_2_last</code> or <code>B_1_last</code>(including other measurements)). Using raw clusters largely degraded our CV and we're not sure about the reason, so count/target encoding was used here.</p></li>\n<li><p><strong>pred0 feature</strong>: using full data to predict specific important features via 2 datasets: \\(\\mathcal{D}_{train}\\) and the fullsize one. For example, if we want to predict <code>P_2</code> in the former case or predict <code>P_2_last</code> in the latter case, we will excluded all \"P-related\" feature and train a lgbm model on it.</p></li>\n<li><p><strong>Probit, Logit models with regularization</strong> on different parts of features (e.g. S-related, B-related sets), like how we make pred0 features, we also exclude related features when we construct the predictions</p></li>\n<li><p><strong>Time-series based methods</strong> to predict important features in the next statement(e.g. <code>P_2</code>, <code>B_1</code>….)</p></li>\n<li><p><strong>Curve fitting, smoothing, denoising</strong>:</p>\n<ul>\n<li>Linear fit: return coefficient and the value of the next 1,3,6,12 months as features</li>\n<li>Poly fit(2,3,4): take the coefficient, and the prediction in next 1 month</li>\n<li>LOWESS fit: only with the <code>customer_ID</code> that have &gt;10 statements.</li>\n<li>Wavelet based method of denoising: Least Asymmetric, Haar, Daubechies(16) then construct some agg features, this method only applies on several important features</li></ul></li>\n<li><p><strong>Leaf-embedding</strong>: An old trick. We mixed the approaches from <a href=\"https://scontent-tpe1-1.xx.fbcdn.net/v/t39.8562-6/240842589_204052295113548_74168590424110542_n.pdf?_nc_cat=109&amp;ccb=1-7&amp;_nc_sid=ad8a9d&amp;_nc_ohc=nd2mJEAuPSkAX-EVY6m&amp;_nc_ht=scontent-tpe1-1.xx&amp;oh=00_AT9NGACeKxD36CIc-jVHSZ_hyBN5NsVETQihOZqmm7nZ1A&amp;oe=630C598A\" target=\"_blank\">this paper</a> and <a href=\"https://www.kaggle.com/code/mmueller/categorical-embedding-with-xgb\" target=\"_blank\">this kaggle code</a>. We use lgbm with 1round + huge amount of leaves + super high regularization to embed the \"whole\" data, which has only 1 dimension. And of course using only categorical data or expanding the embedding matrix to the dimension of <code>nrounds</code> is also possible. </p></li>\n<li><p>Most of the above methods are not really able to improve the best CV in a single model, so we turn to utilize it in different models to generate more diversity. Most of our lgbm models with the above selected methods can have a score ranging from CV.796~.801/ PbLB.797~.80 / PriLB.804~.806,  and some of them can generate pretty good diversity. But we are not going to do a detailed ablation study to check the improvement in the PrivateLB of each method :P</p></li>\n</ul>\n<hr>\n<h2>Feature Selection Strategies</h2>\n<ul>\n<li><strong>LightGBM w/ permutation importance</strong>: During the training process, we observed that the metric D is quite noisy and unstable. Probably as a result, permutation importance selection with Amex-Score didn't work here. So we selected features by monitoring G (Gini score) alone and it worked much better! For example, Ryota's FE generated over 20k features, but the final number of features utilized was about 1200.</li>\n</ul>\n<hr>    \n<h2>Model Details</h2>\n<ul>\n<li><p><strong>GBDT</strong></p>\n<ul>\n<li>LightGBM (gbdt, dart) - classification, regression: CV 0.796~0.801, Public 0.799, Private 0.805~.806</li>\n<li>XGBoost - classification: CV 0.797~0.800, Public 0.798, Private 0.805</li>\n<li>CatBoost - classification, regression CV 0.796~0.799, Public 0.797, Private 0.804</li>\n<li>LGB linear_tree, ExtraTrees, RF, RGF and rmse objective to generate some diversity</li>\n<li>Some LGB dart models trained on the fully flattened view of the data (don't score well individually but ensemble well)</li>\n<li>Some LGB dart models trained on features derived from augmented data view </li></ul></li>\n<li><p><strong>TabNet</strong></p>\n<ul>\n<li>TabNet: CV 0.793 Public 0.794 Private 0.801</li>\n<li>Residual Learning (LightGBM): CV 0.795, Public 0.796 Private 0.802 (specifically, we predict <code>ground truth - tabnet’s prediction</code> with LightGBM and then finalize predictions as <code>tabnet's prediction + resid prediction</code></li></ul></li>\n<li><p><strong>MLP</strong></p>\n<ul>\n<li>MLP: CV 0.796 Public 0.794 Private 0.802</li></ul></li>\n<li><p><strong>Transformer</strong></p>\n<ul>\n<li>CV: .793 -&gt; .794 with residual learning    </li>\n<li>All nulls imputed with LGB trained on all other features, categories encoded as mean <code>P_2</code> value from entire train+test. Data Augmented ~2x by statement shift strategy</li></ul></li>\n</ul>\n<p>Example Dart hyper-parameters from Angus, no special sauce here:)  :</p>\n<pre><code>lgb_param &lt;- list(boosting_type = 'dart',\n                  objective = \"binary\",\n                  metric = amex,\n                  learning_rate = 0.02,\n                  num_leaves = 48,\n                  feature_fraction = 0.1,\n                  #bagging_freq = 1,\n                  #bagging_fraction = 1,\n                  min_child_weight = 1,\n                  lambda_l1 = 1,\n                  lambda_l2 = 64,\n                  skip_drop = 0.8\n)\n</code></pre>\n<hr>\n<h2>Some Other Interesting Things we Tried</h2>\n<ol>\n<li><p><strong>Private set specialization</strong>: We tried using the KNN result mentioned above to estimate which samples are closer to Private LB data. In particular, we calculated the nearest 500 points for each sample, then  the ratio of each set in those 500 points. After that, we took the top 30% of samples which are closer to private LB and ran a logistic regression to decide the weight of our final ensemble. But unfortunately, the CV score dropped a lot. The CV of the top 30% data is ~0.88, but overall(450k data) Amex-score is ~0.7812. We also tried the top 30% \"public ratio\" one. The CV is ~0.7259 and the overall CV is 0.7896711. It's possible that using adversarial validation predictions trained on top importance features would be a more reliable option.</p></li>\n<li><p><strong>Transfer learning on <code>P_2</code></strong>: this idea unfortunately came up very close to the end before it could pay off and probably was worth pursuing further. This was an attempt at exploiting the test data - pretrain a transformer on train+test to predict <code>P_2</code> (because it's strongly related to the target), then fine-tune on the actual target with train. This is like a worse version of the knowledge distillation techniques that were used to exploit the test data more for neural net training. </p></li>\n<li><p><strong>Psuedo labeling</strong>: We only got 0.00005 improvement on CV (LB doesn't seem tio change), and we have tuned the size and what proportion of test-set can be used for it but didn't have real luck. However, log loss and AUC may be high enough to do this more optimally and get a considerable improvement.</p></li>\n<li><p><strong>Ensembling methods</strong>, we tried lgbm with 50 seeds bagging, LR and optimization methods like <code>md1*par1 + md2*par2....</code> with Nelder-Mead or L-BFGS-B solver. LR performed the best (CV.8031), and lgbm seemed to be severely overfitting. </p></li>\n</ol>\n<hr>    \n<h2>Reflections</h2>\n<p>We fell hard from 3rd Public -&gt; 21st Private, with 6 submissions that would have scored gold. In hindsight, CV here did not seem to be a great measurement for the private LB, even though CV improvement aligned very well with public LB. We probably had some bad luck, but 2 things we may have benefited from doing differently are:</p>\n<ul>\n<li>Exploiting the test data more with knowledge distillation, pretraining, etc.</li>\n<li>Hedged our submission strategy more - the two submissions we chose (best Public LB and best CV) were not that different from each other, and choosing a different backup could have been a better way to game the noisiness here</li>\n</ul>\n<p><br></p>",
  "messages": [
    {
      "id": 1917580,
      "postDate": "2022-08-28T21:33:11.853Z",
      "content": "<p>A huge huge thank you to <a href=\"https://www.kaggle.com/andrew60909\" target=\"_blank\">@andrew60909</a> and <a href=\"https://www.kaggle.com/ryotak12\" target=\"_blank\">@ryotak12</a> for an excellent team experience. We learned a lot and had fun, and I appreciate their patience with my extended, largely unsuccessful efforts to make a transformer work really well 😂. Thanks also to Amex and Kaggle for hosting this interesting competition and to <a href=\"https://www.kaggle.com/raddar\" target=\"_blank\">@raddar</a> and everyone else who shared resources that made this competition much more accessible. </p>\n<p>I've made my personal code public here: <a href=\"https://github.com/JEddy92/amex_default_kaggle\" target=\"_blank\">https://github.com/JEddy92/amex_default_kaggle</a>. I tried to create a reasonably general structure for division of FE and modeling, with reusable functionality for training and logging results across a variety of models. There is still some clean up to do and it's not as polished as would be ideal!</p>\n<p>Feel free to ask us anything if you have questions about our approach!</p>\n<h2>Solution Summary</h2>\n<p>We used logistic regression to ensemble ~60 sets of predictions with diversity obtained from different models, feature sets, and data views. Models included the usual mix of gradient boosters (LGBM, LGBM dart, XGBoost, CatBoost), MLPs, TabNet, and Transformers with a few special tricks. Data views included aggregated data, fully flattened data (13 raw features per customer), sequential data (Transformer), and augmented data (shifted statements forward by 1 and appended to original). Very detailed feature information to follow -- they are split across our 3 different team members' different models, but this aims to be a complete compilation of what we all used. </p>\n<hr>\n<h2>Datasets &amp; Notation</h2>\n<ul>\n<li>Everything is derived from Raddar's dataset</li>\n</ul>\n<p>\\(\\mathcal{D}_{fulltrain}\\) : Original full size train-set (5531451 rows)   </p>\n<p>\\(\\mathcal{D}_{train}\\) : Train-set (458913 rows)    </p>\n<p>\\(\\mathcal{D}_{fulltest}\\) : Original full size test-set   </p>\n<p>\\(\\mathcal{D}_{test}\\) : Test-set  </p>\n<hr>\n<h2>Primary Features</h2>\n<p>The train-sets with the best CV score for each of us are generated by \\(\\mathcal{D}_{fulltrain}\\) with these features:  </p>\n<ul>\n<li><p><strong>Numeric</strong>     </p>\n<ul>\n<li>Mean, std, max, min, first, last  </li>\n<li>Last-mean, last/mean, last/first, last/std, max/min, linear-weighted mean </li>\n<li>Similar aggregates as above but limited to last 3, 5, etc. statements</li>\n<li>Recent diff features and aggregates of all diffs </li>\n<li>Monthly-based ranking</li>\n<li>Multiplication and Division: create all possible pairs for all features, except for categorical features and features highly correlated with other features. Multiplication and division is done after scaling with MinMaxScaler(0-1) and clipping outliers</li>\n<li>Null aggregates (including starting streak vs. at random)</li>\n<li>Date features (e.g. stats on customer's statement gaps)</li></ul></li>\n<li><p><strong>Trend</strong> :  Oldest month as 1 and the newest month as 13, take the average over the following periods   </p>\n<ul>\n<li>mean(13,12) - mean(11,10)</li>\n<li>mean(13,12,11) - mean(10,9,8)</li>\n<li>mean(13,12,11) - mean(3,2,1)</li>\n<li>mean(13,12,11,10,9,8) - mean(7,6,5,4,3,2)</li></ul></li>\n<li><p><strong>Category</strong>  </p>\n<ul>\n<li>first, last  </li>\n<li>entropy: Shannon entropy on frequency table   </li>\n<li>Nunique: Number of unique category values</li>\n<li>Counts of each value in each category (both raw counts and tf-idf counts)</li>\n<li><code>P_2</code> mean (across train and test) of each category</li>\n<li>SVD factorization over time of each category</li>\n<li>freq1name : the number of most frequent category  </li>\n<li>freq1ratio : the number of most frequent category / group size  </li>\n<li>freq_last1name : the number of least frequent category  </li>\n<li>freq_last1ratio : the number of least frequent category / group size  </li></ul></li>\n<li><p><strong>PCA</strong></p>\n<ul>\n<li>(N, 13*raw_features) → (N, 64)</li></ul></li>\n<li><p><strong>KNN-based target encoding methods</strong></p>\n<ul>\n<li>We selected the nearest 500 points for each sample with euclidean distance(with 2 groups of most important features, each has 14 features. One with last statement of each customer, another one with mean of all statements), calculate the average of target as features</li>\n<li>Follow the above method, using nearest 500 points but applying average with weight by the distance matrix</li>\n<li>Aggregate the distance matrix, by mean, max, min, std</li>\n<li>We also calculated the cosine similarity of each customer's past N months of data in one dimension.(N=1,3,6,13) The average of the 500 nearest neighbor customers' targets is used as the feature. This feature was binned because the distribution is somewhat different between Train and Test</li></ul></li>\n<li><p><strong>Nested model</strong></p>\n<ul>\n<li><p>We added the labels to \\(\\mathcal{D}_{fulltrain}\\), then trained a LGBM model. The purpose of doing this is to capture what kinds of records, behavior and attributes will cause default. Instead of aggregating features into 458913 rows, this method can let the model learn some extra information from a more \"base\" level. After that, for each <code>customer_ID</code> we will have a \"predicted target\" that has the same length as the number of records in each <code>customer_ID</code>.   </p></li>\n<li><p>Finally, to merge it into \\(\\mathcal{D}_{train}\\) : Train-set (458913 rows), we aggregate similarly to other features: mean, std, max, min, first, last, man-min, last-mean, last/std.  </p></li>\n<li><p>An interesting point here is that the last value of our \"predicted value\" with \\(\\mathcal{D}_{fulltrain}\\) scores 0.785 on the Amex metric, while the score of the max value is around 0.62 -- indicating the importance of the last value.</p></li></ul></li>\n</ul>\n<hr>       \n<h2>Other Feature Methods:</h2>\n<ul>\n<li><p><strong>Target/Count encoding</strong>: On \\(\\mathcal{D}_{train}\\)</p></li>\n<li><p><strong>Target encoding</strong> with \\(\\mathcal{D}_{fulltrain}\\)</p></li>\n<li><p><strong>Other functions to aggregate</strong> \\(\\mathcal{D}_{fulltrain}\\)</p>\n<ul>\n<li>Exponential weight average</li>\n<li>Skewness, Kurtosis</li>\n<li>The ratio beyond 1std </li>\n<li>Median Absolute Deviation</li>\n<li>Mean, max, min, std on 1st derivative(t=1,2,…13), timestamp of max and min also used </li>\n<li>Shannon entropy: estimate of the spectral density of sequence</li>\n<li>Stability, Lumpiness: variance of the means and variance of the variances on tiles of windows</li>\n<li>Crossing points: number of times a sequence crosses the median line</li>\n<li>Some measurements from <a href=\"http://isadoranun.github.io/tsfeat/FeaturesDocumentation.html\" target=\"_blank\">here</a> </li></ul></li>\n<li><p><strong>Clustering methods</strong>: Kmeans, DBSCAN on important feature(e.g. <code>P_2_last</code> or <code>B_1_last</code>(including other measurements)). Using raw clusters largely degraded our CV and we're not sure about the reason, so count/target encoding was used here.</p></li>\n<li><p><strong>pred0 feature</strong>: using full data to predict specific important features via 2 datasets: \\(\\mathcal{D}_{train}\\) and the fullsize one. For example, if we want to predict <code>P_2</code> in the former case or predict <code>P_2_last</code> in the latter case, we will excluded all \"P-related\" feature and train a lgbm model on it.</p></li>\n<li><p><strong>Probit, Logit models with regularization</strong> on different parts of features (e.g. S-related, B-related sets), like how we make pred0 features, we also exclude related features when we construct the predictions</p></li>\n<li><p><strong>Time-series based methods</strong> to predict important features in the next statement(e.g. <code>P_2</code>, <code>B_1</code>….)</p></li>\n<li><p><strong>Curve fitting, smoothing, denoising</strong>:</p>\n<ul>\n<li>Linear fit: return coefficient and the value of the next 1,3,6,12 months as features</li>\n<li>Poly fit(2,3,4): take the coefficient, and the prediction in next 1 month</li>\n<li>LOWESS fit: only with the <code>customer_ID</code> that have &gt;10 statements.</li>\n<li>Wavelet based method of denoising: Least Asymmetric, Haar, Daubechies(16) then construct some agg features, this method only applies on several important features</li></ul></li>\n<li><p><strong>Leaf-embedding</strong>: An old trick. We mixed the approaches from <a href=\"https://scontent-tpe1-1.xx.fbcdn.net/v/t39.8562-6/240842589_204052295113548_74168590424110542_n.pdf?_nc_cat=109&amp;ccb=1-7&amp;_nc_sid=ad8a9d&amp;_nc_ohc=nd2mJEAuPSkAX-EVY6m&amp;_nc_ht=scontent-tpe1-1.xx&amp;oh=00_AT9NGACeKxD36CIc-jVHSZ_hyBN5NsVETQihOZqmm7nZ1A&amp;oe=630C598A\" target=\"_blank\">this paper</a> and <a href=\"https://www.kaggle.com/code/mmueller/categorical-embedding-with-xgb\" target=\"_blank\">this kaggle code</a>. We use lgbm with 1round + huge amount of leaves + super high regularization to embed the \"whole\" data, which has only 1 dimension. And of course using only categorical data or expanding the embedding matrix to the dimension of <code>nrounds</code> is also possible. </p></li>\n<li><p>Most of the above methods are not really able to improve the best CV in a single model, so we turn to utilize it in different models to generate more diversity. Most of our lgbm models with the above selected methods can have a score ranging from CV.796~.801/ PbLB.797~.80 / PriLB.804~.806,  and some of them can generate pretty good diversity. But we are not going to do a detailed ablation study to check the improvement in the PrivateLB of each method :P</p></li>\n</ul>\n<hr>\n<h2>Feature Selection Strategies</h2>\n<ul>\n<li><strong>LightGBM w/ permutation importance</strong>: During the training process, we observed that the metric D is quite noisy and unstable. Probably as a result, permutation importance selection with Amex-Score didn't work here. So we selected features by monitoring G (Gini score) alone and it worked much better! For example, Ryota's FE generated over 20k features, but the final number of features utilized was about 1200.</li>\n</ul>\n<hr>    \n<h2>Model Details</h2>\n<ul>\n<li><p><strong>GBDT</strong></p>\n<ul>\n<li>LightGBM (gbdt, dart) - classification, regression: CV 0.796~0.801, Public 0.799, Private 0.805~.806</li>\n<li>XGBoost - classification: CV 0.797~0.800, Public 0.798, Private 0.805</li>\n<li>CatBoost - classification, regression CV 0.796~0.799, Public 0.797, Private 0.804</li>\n<li>LGB linear_tree, ExtraTrees, RF, RGF and rmse objective to generate some diversity</li>\n<li>Some LGB dart models trained on the fully flattened view of the data (don't score well individually but ensemble well)</li>\n<li>Some LGB dart models trained on features derived from augmented data view </li></ul></li>\n<li><p><strong>TabNet</strong></p>\n<ul>\n<li>TabNet: CV 0.793 Public 0.794 Private 0.801</li>\n<li>Residual Learning (LightGBM): CV 0.795, Public 0.796 Private 0.802 (specifically, we predict <code>ground truth - tabnet’s prediction</code> with LightGBM and then finalize predictions as <code>tabnet's prediction + resid prediction</code></li></ul></li>\n<li><p><strong>MLP</strong></p>\n<ul>\n<li>MLP: CV 0.796 Public 0.794 Private 0.802</li></ul></li>\n<li><p><strong>Transformer</strong></p>\n<ul>\n<li>CV: .793 -&gt; .794 with residual learning    </li>\n<li>All nulls imputed with LGB trained on all other features, categories encoded as mean <code>P_2</code> value from entire train+test. Data Augmented ~2x by statement shift strategy</li></ul></li>\n</ul>\n<p>Example Dart hyper-parameters from Angus, no special sauce here:)  :</p>\n<pre><code>lgb_param &lt;- list(boosting_type = 'dart',\n                  objective = \"binary\",\n                  metric = amex,\n                  learning_rate = 0.02,\n                  num_leaves = 48,\n                  feature_fraction = 0.1,\n                  #bagging_freq = 1,\n                  #bagging_fraction = 1,\n                  min_child_weight = 1,\n                  lambda_l1 = 1,\n                  lambda_l2 = 64,\n                  skip_drop = 0.8\n)\n</code></pre>\n<hr>\n<h2>Some Other Interesting Things we Tried</h2>\n<ol>\n<li><p><strong>Private set specialization</strong>: We tried using the KNN result mentioned above to estimate which samples are closer to Private LB data. In particular, we calculated the nearest 500 points for each sample, then  the ratio of each set in those 500 points. After that, we took the top 30% of samples which are closer to private LB and ran a logistic regression to decide the weight of our final ensemble. But unfortunately, the CV score dropped a lot. The CV of the top 30% data is ~0.88, but overall(450k data) Amex-score is ~0.7812. We also tried the top 30% \"public ratio\" one. The CV is ~0.7259 and the overall CV is 0.7896711. It's possible that using adversarial validation predictions trained on top importance features would be a more reliable option.</p></li>\n<li><p><strong>Transfer learning on <code>P_2</code></strong>: this idea unfortunately came up very close to the end before it could pay off and probably was worth pursuing further. This was an attempt at exploiting the test data - pretrain a transformer on train+test to predict <code>P_2</code> (because it's strongly related to the target), then fine-tune on the actual target with train. This is like a worse version of the knowledge distillation techniques that were used to exploit the test data more for neural net training. </p></li>\n<li><p><strong>Psuedo labeling</strong>: We only got 0.00005 improvement on CV (LB doesn't seem tio change), and we have tuned the size and what proportion of test-set can be used for it but didn't have real luck. However, log loss and AUC may be high enough to do this more optimally and get a considerable improvement.</p></li>\n<li><p><strong>Ensembling methods</strong>, we tried lgbm with 50 seeds bagging, LR and optimization methods like <code>md1*par1 + md2*par2....</code> with Nelder-Mead or L-BFGS-B solver. LR performed the best (CV.8031), and lgbm seemed to be severely overfitting. </p></li>\n</ol>\n<hr>    \n<h2>Reflections</h2>\n<p>We fell hard from 3rd Public -&gt; 21st Private, with 6 submissions that would have scored gold. In hindsight, CV here did not seem to be a great measurement for the private LB, even though CV improvement aligned very well with public LB. We probably had some bad luck, but 2 things we may have benefited from doing differently are:</p>\n<ul>\n<li>Exploiting the test data more with knowledge distillation, pretraining, etc.</li>\n<li>Hedged our submission strategy more - the two submissions we chose (best Public LB and best CV) were not that different from each other, and choosing a different backup could have been a better way to game the noisiness here</li>\n</ul>\n<p><br></p>",
      "rawMarkdown": "A huge huge thank you to @andrew60909 and @ryotak12 for an excellent team experience. We learned a lot and had fun, and I appreciate their patience with my extended, largely unsuccessful efforts to make a transformer work really well 😂. Thanks also to Amex and Kaggle for hosting this interesting competition and to @raddar and everyone else who shared resources that made this competition much more accessible. \n\nI've made my personal code public here: https://github.com/JEddy92/amex_default_kaggle. I tried to create a reasonably general structure for division of FE and modeling, with reusable functionality for training and logging results across a variety of models. There is still some clean up to do and it's not as polished as would be ideal!\n\nFeel free to ask us anything if you have questions about our approach!\n\n## Solution Summary\n\nWe used logistic regression to ensemble ~60 sets of predictions with diversity obtained from different models, feature sets, and data views. Models included the usual mix of gradient boosters (LGBM, LGBM dart, XGBoost, CatBoost), MLPs, TabNet, and Transformers with a few special tricks. Data views included aggregated data, fully flattened data (13 raw features per customer), sequential data (Transformer), and augmented data (shifted statements forward by 1 and appended to original). Very detailed feature information to follow -- they are split across our 3 different team members' different models, but this aims to be a complete compilation of what we all used. \n\n\n<hr style=\"border:2px solid gray\">\n\n## Datasets & Notation\n* Everything is derived from Raddar's dataset\n\n\\\\(\\mathcal{D}_{fulltrain}\\\\) : Original full size train-set (5531451 rows)   \n\n\\\\(\\mathcal{D}_{train}\\\\) : Train-set (458913 rows)    \n\n\\\\(\\mathcal{D}_{fulltest}\\\\) : Original full size test-set   \n\n\\\\(\\mathcal{D}_{test}\\\\) : Test-set  \n\n\n<hr style=\"border:2px solid gray\">\n\n## Primary Features\nThe train-sets with the best CV score for each of us are generated by \\\\(\\mathcal{D}_{fulltrain}\\\\) with these features:  \n  \n* **Numeric**     \n  + Mean, std, max, min, first, last  \n  + Last-mean, last/mean, last/first, last/std, max/min, linear-weighted mean \n  + Similar aggregates as above but limited to last 3, 5, etc. statements\n  + Recent diff features and aggregates of all diffs \n  + Monthly-based ranking\n  + Multiplication and Division: create all possible pairs for all features, except for categorical features and features highly correlated with other features. Multiplication and division is done after scaling with MinMaxScaler(0-1) and clipping outliers\n  + Null aggregates (including starting streak vs. at random)\n  + Date features (e.g. stats on customer's statement gaps)\n  \n* **Trend** :  Oldest month as 1 and the newest month as 13, take the average over the following periods   \n    - mean(13,12) - mean(11,10)\n    + mean(13,12,11) - mean(10,9,8)\n    + mean(13,12,11) - mean(3,2,1)\n    + mean(13,12,11,10,9,8) - mean(7,6,5,4,3,2)\n\n* **Category**  \n    + first, last  \n    + entropy: Shannon entropy on frequency table   \n    + Nunique: Number of unique category values\n    + Counts of each value in each category (both raw counts and tf-idf counts)\n    + `P_2` mean (across train and test) of each category\n    + SVD factorization over time of each category\n    + freq1name : the number of most frequent category  \n    + freq1ratio : the number of most frequent category / group size  \n    + freq_last1name : the number of least frequent category  \n    + freq_last1ratio : the number of least frequent category / group size  \n  \n* **PCA**\n    + (N, 13*raw_features) → (N, 64)\n  \n* **KNN-based target encoding methods**\n    + We selected the nearest 500 points for each sample with euclidean distance(with 2 groups of most important features, each has 14 features. One with last statement of each customer, another one with mean of all statements), calculate the average of target as features\n    + Follow the above method, using nearest 500 points but applying average with weight by the distance matrix\n    + Aggregate the distance matrix, by mean, max, min, std\n    + We also calculated the cosine similarity of each customer's past N months of data in one dimension.(N=1,3,6,13) The average of the 500 nearest neighbor customers' targets is used as the feature. This feature was binned because the distribution is somewhat different between Train and Test\n    \n    \n* **Nested model**\n\n    + We added the labels to \\\\(\\mathcal{D}_{fulltrain}\\\\), then trained a LGBM model. The purpose of doing this is to capture what kinds of records, behavior and attributes will cause default. Instead of aggregating features into 458913 rows, this method can let the model learn some extra information from a more \"base\" level. After that, for each `customer_ID` we will have a \"predicted target\" that has the same length as the number of records in each `customer_ID`.   \n    \n    + Finally, to merge it into \\\\(\\mathcal{D}_{train}\\\\) : Train-set (458913 rows), we aggregate similarly to other features: mean, std, max, min, first, last, man-min, last-mean, last/std.  \n    \n    + An interesting point here is that the last value of our \"predicted value\" with \\\\(\\mathcal{D}_{fulltrain}\\\\) scores 0.785 on the Amex metric, while the score of the max value is around 0.62 -- indicating the importance of the last value.\n      \n<hr style=\"border:2px solid gray\">       \n\n## Other Feature Methods: \n\n* **Target/Count encoding**: On \\\\(\\mathcal{D}_{train}\\\\)\n\n* **Target encoding** with \\\\(\\mathcal{D}_{fulltrain}\\\\)\n\n* **Other functions to aggregate** \\\\(\\mathcal{D}_{fulltrain}\\\\)\n    + Exponential weight average\n    + Skewness, Kurtosis\n    + The ratio beyond 1std \n    + Median Absolute Deviation\n    + Mean, max, min, std on 1st derivative(t=1,2,...13), timestamp of max and min also used \n    + Shannon entropy: estimate of the spectral density of sequence\n    + Stability, Lumpiness: variance of the means and variance of the variances on tiles of windows\n    + Crossing points: number of times a sequence crosses the median line\n    + Some measurements from [here](http://isadoranun.github.io/tsfeat/FeaturesDocumentation.html) \n    \n* **Clustering methods**: Kmeans, DBSCAN on important feature(e.g. `P_2_last` or `B_1_last`(including other measurements)). Using raw clusters largely degraded our CV and we're not sure about the reason, so count/target encoding was used here.\n\n* **pred0 feature**: using full data to predict specific important features via 2 datasets: \\\\(\\mathcal{D}_{train}\\\\) and the fullsize one. For example, if we want to predict `P_2` in the former case or predict `P_2_last` in the latter case, we will excluded all \"P-related\" feature and train a lgbm model on it.\n\n* **Probit, Logit models with regularization** on different parts of features (e.g. S-related, B-related sets), like how we make pred0 features, we also exclude related features when we construct the predictions\n\n* **Time-series based methods** to predict important features in the next statement(e.g. `P_2`, `B_1`....)\n\n* **Curve fitting, smoothing, denoising**:\n    + Linear fit: return coefficient and the value of the next 1,3,6,12 months as features\n    + Poly fit(2,3,4): take the coefficient, and the prediction in next 1 month\n    + LOWESS fit: only with the `customer_ID` that have >10 statements.\n    + Wavelet based method of denoising: Least Asymmetric, Haar, Daubechies(16) then construct some agg features, this method only applies on several important features\n\n* **Leaf-embedding**: An old trick. We mixed the approaches from [this paper](https://scontent-tpe1-1.xx.fbcdn.net/v/t39.8562-6/240842589_204052295113548_74168590424110542_n.pdf?_nc_cat=109&ccb=1-7&_nc_sid=ad8a9d&_nc_ohc=nd2mJEAuPSkAX-EVY6m&_nc_ht=scontent-tpe1-1.xx&oh=00_AT9NGACeKxD36CIc-jVHSZ_hyBN5NsVETQihOZqmm7nZ1A&oe=630C598A) and [this kaggle code](https://www.kaggle.com/code/mmueller/categorical-embedding-with-xgb). We use lgbm with 1round + huge amount of leaves + super high regularization to embed the \"whole\" data, which has only 1 dimension. And of course using only categorical data or expanding the embedding matrix to the dimension of `nrounds` is also possible. \n        \n* Most of the above methods are not really able to improve the best CV in a single model, so we turn to utilize it in different models to generate more diversity. Most of our lgbm models with the above selected methods can have a score ranging from CV.796~.801/ PbLB.797~.80 / PriLB.804~.806,  and some of them can generate pretty good diversity. But we are not going to do a detailed ablation study to check the improvement in the PrivateLB of each method :P\n\n\n<hr style=\"border:2px solid gray\">\n\n## Feature Selection Strategies\n\n* **LightGBM w/ permutation importance**: During the training process, we observed that the metric D is quite noisy and unstable. Probably as a result, permutation importance selection with Amex-Score didn't work here. So we selected features by monitoring G (Gini score) alone and it worked much better! For example, Ryota's FE generated over 20k features, but the final number of features utilized was about 1200.\n\n<hr style=\"border:2px solid gray\">    \n\n## Model Details\n\n* **GBDT**\n    - LightGBM (gbdt, dart) - classification, regression: CV 0.796~0.801, Public 0.799, Private 0.805~.806\n    - XGBoost - classification: CV 0.797~0.800, Public 0.798, Private 0.805\n    - CatBoost - classification, regression CV 0.796~0.799, Public 0.797, Private 0.804\n    - LGB linear_tree, ExtraTrees, RF, RGF and rmse objective to generate some diversity\n    - Some LGB dart models trained on the fully flattened view of the data (don't score well individually but ensemble well)\n    - Some LGB dart models trained on features derived from augmented data view \n\n* **TabNet**\n    - TabNet: CV 0.793 Public 0.794 Private 0.801\n    - Residual Learning (LightGBM): CV 0.795, Public 0.796 Private 0.802 (specifically, we predict `ground truth - tabnet’s prediction` with LightGBM and then finalize predictions as `tabnet's prediction + resid prediction`\n    \n* **MLP**\n    - MLP: CV 0.796 Public 0.794 Private 0.802\n\n* **Transformer**\n\t- CV: .793  -> .794 with residual learning\t\n\t- All nulls imputed with LGB trained on all other features, categories encoded as mean `P_2` value from entire train+test. Data Augmented ~2x by statement shift strategy\n\nExample Dart hyper-parameters from Angus, no special sauce here:)  :\n\n```\nlgb_param <- list(boosting_type = 'dart',\n                  objective = \"binary\",\n                  metric = amex,\n                  learning_rate = 0.02,\n                  num_leaves = 48,\n                  feature_fraction = 0.1,\n                  #bagging_freq = 1,\n                  #bagging_fraction = 1,\n                  min_child_weight = 1,\n                  lambda_l1 = 1,\n                  lambda_l2 = 64,\n                  skip_drop = 0.8\n)\n```\n\n<hr style=\"border:2px solid gray\">\n\n## Some Other Interesting Things we Tried\n\n1. **Private set specialization**: We tried using the KNN result mentioned above to estimate which samples are closer to Private LB data. In particular, we calculated the nearest 500 points for each sample, then  the ratio of each set in those 500 points. After that, we took the top 30% of samples which are closer to private LB and ran a logistic regression to decide the weight of our final ensemble. But unfortunately, the CV score dropped a lot. The CV of the top 30% data is ~0.88, but overall(450k data) Amex-score is ~0.7812. We also tried the top 30% \"public ratio\" one. The CV is ~0.7259 and the overall CV is 0.7896711. It's possible that using adversarial validation predictions trained on top importance features would be a more reliable option.\n\n2. **Transfer learning on `P_2`**: this idea unfortunately came up very close to the end before it could pay off and probably was worth pursuing further. This was an attempt at exploiting the test data - pretrain a transformer on train+test to predict `P_2` (because it's strongly related to the target), then fine-tune on the actual target with train. This is like a worse version of the knowledge distillation techniques that were used to exploit the test data more for neural net training. \n\n2. **Psuedo labeling**: We only got 0.00005 improvement on CV (LB doesn't seem tio change), and we have tuned the size and what proportion of test-set can be used for it but didn't have real luck. However, log loss and AUC may be high enough to do this more optimally and get a considerable improvement.\n\n3. **Ensembling methods**, we tried lgbm with 50 seeds bagging, LR and optimization methods like `md1*par1 + md2*par2....` with Nelder-Mead or L-BFGS-B solver. LR performed the best (CV.8031), and lgbm seemed to be severely overfitting. \n\n<hr style=\"border:2px solid gray\">    \n\n## Reflections\n\nWe fell hard from 3rd Public -> 21st Private, with 6 submissions that would have scored gold. In hindsight, CV here did not seem to be a great measurement for the private LB, even though CV improvement aligned very well with public LB. We probably had some bad luck, but 2 things we may have benefited from doing differently are:\n- Exploiting the test data more with knowledge distillation, pretraining, etc.\n- Hedged our submission strategy more - the two submissions we chose (best Public LB and best CV) were not that different from each other, and choosing a different backup could have been a better way to game the noisiness here\n\n\n<br/>\n\n\n",
      "votes": 57
    },
    {
      "id": 1921355,
      "postDate": "2022-08-31T18:59:21.167Z",
      "content": "<p>thanks! its very helpful for beginner how to approach problem</p>",
      "rawMarkdown": "thanks! its very helpful for beginner how to approach problem",
      "votes": 1
    },
    {
      "id": 1920201,
      "postDate": "2022-08-31T01:52:54.433Z",
      "content": "<p>Thanks for sharing your solution <a href=\"https://www.kaggle.com/aquatic\" target=\"_blank\">@aquatic</a>! Very insightful and it shows the amount of work that you put in!</p>",
      "rawMarkdown": "Thanks for sharing your solution @aquatic! Very insightful and it shows the amount of work that you put in!",
      "votes": 1
    },
    {
      "id": 1918478,
      "postDate": "2022-08-29T15:50:58.550Z",
      "content": "<p>Wow, you did a lot of great work!</p>\n<p>I know you varied the data for your models, but was it the same 1.2k base after feature selection for all models using the \"normal\" set of data? I wonder if \"less aggressive feature selection\" is another possible thing that, in retrospect, could've helped the private lb score?</p>",
      "rawMarkdown": "Wow, you did a lot of great work!\n\nI know you varied the data for your models, but was it the same 1.2k base after feature selection for all models using the \"normal\" set of data? I wonder if \"less aggressive feature selection\" is another possible thing that, in retrospect, could've helped the private lb score?",
      "votes": 1,
      "replies": [
        {
          "id": 1918500,
          "postDate": "2022-08-29T16:10:03.987Z",
          "content": "<p>Note I'm asking because I see so few differences that seem critical (that I noticed) between your solution and top 2 solutions. </p>\n<p>Hmm, maybe top two had more emphasis on a normalization factor? You had monthly based ranking, but top solution had ranking of a single customer's statements across months AND ranking across customers for a given month, AND aggregated those things. (Also could feature selection have removed some of the monthly ranking or any other normalization that you did do?)</p>",
          "rawMarkdown": "Note I'm asking because I see so few differences that seem critical (that I noticed) between your solution and top 2 solutions. \n\nHmm, maybe top two had more emphasis on a normalization factor? You had monthly based ranking, but top solution had ranking of a single customer's statements across months AND ranking across customers for a given month, AND aggregated those things. (Also could feature selection have removed some of the monthly ranking or any other normalization that you did do?)",
          "votes": 1
        },
        {
          "id": 1918718,
          "postDate": "2022-08-29T19:17:40.697Z",
          "content": "<p>Thanks!</p>\n<p>That method was specifically what <a href=\"https://www.kaggle.com/ryotak12\" target=\"_blank\">@ryotak12</a> used for selection (especially for some of the feature sets that would otherwise just be enormous). For my models, I actually didn't use selection, more like feed-forward additions (and actually included different/older versions of models that had a subset of the total feature set I used, these made the cut in the ensemble). Feature selection is hard, especially on this problem, and I thought it'd be more time efficient to prioritize diversity for my work. </p>\n<p>If I'm remembering correctly, top 2 solutions both didn't use stacking so that may be one differentiating point. Perhaps our stacking overfits to CV / Public LB relative to the highest scoring Private solutions.</p>",
          "rawMarkdown": "Thanks!\n\nThat method was specifically what @ryotak12 used for selection (especially for some of the feature sets that would otherwise just be enormous). For my models, I actually didn't use selection, more like feed-forward additions (and actually included different/older versions of models that had a subset of the total feature set I used, these made the cut in the ensemble). Feature selection is hard, especially on this problem, and I thought it'd be more time efficient to prioritize diversity for my work. \n\nIf I'm remembering correctly, top 2 solutions both didn't use stacking so that may be one differentiating point. Perhaps our stacking overfits to CV / Public LB relative to the highest scoring Private solutions.",
          "votes": 2
        },
        {
          "id": 1919011,
          "postDate": "2022-08-30T04:09:45.923Z",
          "content": "<p>As Joe said, only my feature set was applied to this selection method. <br>\nI generated a large number of features at beginning and could not train them all 20k over features.<br>\nSo, I had to select features due to training time.</p>",
          "rawMarkdown": "As Joe said, only my feature set was applied to this selection method. \nI generated a large number of features at beginning and could not train them all 20k over features.\nSo, I had to select features due to training time.",
          "votes": 1
        }
      ]
    },
    {
      "id": 1917789,
      "postDate": "2022-08-29T04:39:57.480Z",
      "content": "<p>Congratulations Joe!<br>\nCould you please share more about feature selection part? After you get the permutation importance of 20k features, how do you select 1200 features? Did you just pick the top 1200 features? If then, how do you decide how many features? Thanks </p>",
      "rawMarkdown": "Congratulations Joe!\nCould you please share more about feature selection part? After you get the permutation importance of 20k features, how do you select 1200 features? Did you just pick the top 1200 features? If then, how do you decide how many features? Thanks \n",
      "votes": 1,
      "replies": [
        {
          "id": 1917830,
          "postDate": "2022-08-29T05:24:58.427Z",
          "content": "<p>I will answer your question.</p>\n<p>I could not train 20k features at once, because it takes too long time and memory.<br>\nSo, at first, I divided it into 15 subsets and experimented with 1500 features each.<br>\nAt this stage, I selected features that decrease by more than 1e-5 in G-metric.<br>\nThe reason that this threshold was chosen is that experiments had shown that features that do not contribute to improving metric can also results -1e-5.<br>\nThese are repeated and finally the experiment is conducted with a single set of features. And 1200 features were obtained.</p>",
          "rawMarkdown": "I will answer your question.\n\nI could not train 20k features at once, because it takes too long time and memory.\nSo, at first, I divided it into 15 subsets and experimented with 1500 features each.\nAt this stage, I selected features that decrease by more than 1e-5 in G-metric.\nThe reason that this threshold was chosen is that experiments had shown that features that do not contribute to improving metric can also results -1e-5.\nThese are repeated and finally the experiment is conducted with a single set of features. And 1200 features were obtained.",
          "votes": 6
        },
        {
          "id": 1918356,
          "postDate": "2022-08-29T14:25:40.043Z",
          "content": "<p>Thanks for your reply.<br>\nChosing negative 1e-5 by experiments is really a fantastic idea! Usually people drop features with negative G-metric.<br>\nMay I ask when training using 1500 features, is it necessary to train 5 models if I use 5 folds? In that case, it would be very time-consuming. <br>\nAlso do you permutate feature on the oof data? It seems to me permutation importance on the oof would be better, cause it represents the ability a feature can generalize on unseen data or not.</p>",
          "rawMarkdown": "Thanks for your reply.\nChosing negative 1e-5 by experiments is really a fantastic idea! Usually people drop features with negative G-metric.\nMay I ask when training using 1500 features, is it necessary to train 5 models if I use 5 folds? In that case, it would be very time-consuming. \nAlso do you permutate feature on the oof data? It seems to me permutation importance on the oof would be better, cause it represents the ability a feature can generalize on unseen data or not.",
          "votes": 1
        },
        {
          "id": 1918373,
          "postDate": "2022-08-29T14:44:37.483Z",
          "content": "<p>I used 3 folds, and learning_rate=0.1.<br>\nThis still took a huge amount of time for feature selection.<br>\nIf I had used RAPIDs FIL, it would have been faster. But I didn't know.<br>\n(Chris introduces it in his solution. <a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/347641\" target=\"_blank\">https://www.kaggle.com/competitions/amex-default-prediction/discussion/347641</a>)</p>\n<p>Yes, oof data.</p>",
          "rawMarkdown": "I used 3 folds, and learning_rate=0.1.\nThis still took a huge amount of time for feature selection.\nIf I had used RAPIDs FIL, it would have been faster. But I didn't know.\n(Chris introduces it in his solution. https://www.kaggle.com/competitions/amex-default-prediction/discussion/347641)\n\nYes, oof data.",
          "votes": 2
        }
      ]
    },
    {
      "id": 1917636,
      "postDate": "2022-08-28T23:57:33.703Z",
      "content": "<p>Hello <a href=\"https://www.kaggle.com/aquatic\" target=\"_blank\">@aquatic</a>, impressive model architecture; thanks for sharing</p>",
      "rawMarkdown": "Hello @aquatic, impressive model architecture; thanks for sharing",
      "votes": 1
    },
    {
      "id": 1917630,
      "postDate": "2022-08-28T23:37:00.770Z",
      "content": "<p>Congratulations Joe and team. I'm sorry to see you drop on private. I was excited to see your huge climb on public LB in the last days.</p>\n<p>Can you explain \"residual learning\" more? Is this training a model on the error of another model?</p>",
      "rawMarkdown": "Congratulations Joe and team. I'm sorry to see you drop on private. I was excited to see your huge climb on public LB in the last days.\n\nCan you explain \"residual learning\" more? Is this training a model on the error of another model?",
      "votes": 1,
      "replies": [
        {
          "id": 1917645,
          "postDate": "2022-08-29T00:18:24.850Z",
          "content": "<p>Thanks for the kind words Chris!</p>\n<p><a href=\"https://www.kaggle.com/andrew60909\" target=\"_blank\">@andrew60909</a> brought this method to the team and referenced the <a href=\"https://www.kaggle.com/c/mercari-price-suggestion-challenge/discussion/50256\" target=\"_blank\">winning Mercari solution</a>. That's exactly it, we took OOF predictions, trained a regression LGBM model on some of our feature sets to predict <code>resid = y_train - oof_pred_train</code>, and then generated the test predictions as <code>pred_test + pred_resid_test</code> (and new OOF as <code>oof_pred_val + pred_resid_val</code>). This improved our TabNet and Transformer models by .001-.002 while still yielding a diverse final model. We did find it to be a little finicky - it overfits quickly so needs to be well regularized, and I think only really good features help. </p>\n<p>Would love to hear if Angus has more thoughts, but to me it's almost like a mini post-processor, or a hybrid gradient boosting model where you're boosting off of a strong prediction initial prediction instead of the usual best constant prediction starting point. </p>",
          "rawMarkdown": "Thanks for the kind words Chris!\n\n@andrew60909 brought this method to the team and referenced the [winning Mercari solution](https://www.kaggle.com/c/mercari-price-suggestion-challenge/discussion/50256). That's exactly it, we took OOF predictions, trained a regression LGBM model on some of our feature sets to predict `resid = y_train - oof_pred_train`, and then generated the test predictions as `pred_test + pred_resid_test` (and new OOF as `oof_pred_val + pred_resid_val`). This improved our TabNet and Transformer models by .001-.002 while still yielding a diverse final model. We did find it to be a little finicky - it overfits quickly so needs to be well regularized, and I think only really good features help. \n\nWould love to hear if Angus has more thoughts, but to me it's almost like a mini post-processor, or a hybrid gradient boosting model where you're boosting off of a strong prediction initial prediction instead of the usual best constant prediction starting point. ",
          "votes": 6
        },
        {
          "id": 1917725,
          "postDate": "2022-08-29T02:47:59.393Z",
          "content": "<p>Just as Joe said - it's actually more like post-processing, but it can generate pretty good diversity and boost the score of the relatively weak learner. We can also consider it as a kind of non-linear stacking. As for final ensembling, having these 2 models simultaneously will be a good idea too.</p>",
          "rawMarkdown": "Just as Joe said - it's actually more like post-processing, but it can generate pretty good diversity and boost the score of the relatively weak learner. We can also consider it as a kind of non-linear stacking. As for final ensembling, having these 2 models simultaneously will be a good idea too.",
          "votes": 4
        },
        {
          "id": 1918493,
          "postDate": "2022-08-29T16:02:23.080Z",
          "content": "<p>If I understand correctly, you can also look at it as what LGBM (or XGB) is doing every round anyways. Each tree is predicting to eliminate the residual up to that point. </p>\n<p>Joe, Angus, do you know if this has been tried much at all on the output of the entire final ensemble? Can it be effective as the final step?</p>",
          "rawMarkdown": "If I understand correctly, you can also look at it as what LGBM (or XGB) is doing every round anyways. Each tree is predicting to eliminate the residual up to that point. \n\nJoe, Angus, do you know if this has been tried much at all on the output of the entire final ensemble? Can it be effective as the final step?",
          "votes": 2
        },
        {
          "id": 1918723,
          "postDate": "2022-08-29T19:26:23.230Z",
          "content": "<p>Haven't tried this, but I wouldn't expect it to help an ensemble. It should work if there is a systematic way to use features and a different model/view to explain errors made by a specific model/view, but I would think everything systematic is already in an ensemble that includes the GBT methods as base models. Also the error margins that the method can work with just become slimmer and slimmer as the model its applied to improves. I think part of why we saw it work for Transformer and TabNet but not MLP is that the data view for the former two is further removed from the flat aggregated data + GBT view.</p>",
          "rawMarkdown": "Haven't tried this, but I wouldn't expect it to help an ensemble. It should work if there is a systematic way to use features and a different model/view to explain errors made by a specific model/view, but I would think everything systematic is already in an ensemble that includes the GBT methods as base models. Also the error margins that the method can work with just become slimmer and slimmer as the model its applied to improves. I think part of why we saw it work for Transformer and TabNet but not MLP is that the data view for the former two is further removed from the flat aggregated data + GBT view.",
          "votes": 2
        }
      ]
    },
    {
      "id": 1920878,
      "postDate": "2022-08-31T13:04:42.617Z",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/aquatic\" target=\"_blank\">@aquatic</a>, May I invite you to participate in this survey regarding your experience on Kaggle (10 min)? This is not a scam. We are a group of researchers at the City University of Hong Kong. The survey link is: <a href=\"https://cityuhk.questionpro.com/survey-of-kaggle-contestants\" target=\"_blank\">https://cityuhk.questionpro.com/survey-of-kaggle-contestants</a></p>",
      "rawMarkdown": "Hi @aquatic, May I invite you to participate in this survey regarding your experience on Kaggle (10 min)? This is not a scam. We are a group of researchers at the City University of Hong Kong. The survey link is: https://cityuhk.questionpro.com/survey-of-kaggle-contestants",
      "votes": -3
    },
    {
      "id": 1917646,
      "postDate": "2022-08-29T00:21:06.407Z",
      "rawMarkdown": "",
      "votes": 1,
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 1921355,
      "author_name": "asjad2024",
      "author_url": "",
      "post_date": "2022-08-31T18:59:21.167000",
      "content": "<p>thanks! its very helpful for beginner how to approach problem</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1920201,
      "author_name": "Oscar Aguilar",
      "author_url": "",
      "post_date": "2022-08-31T01:52:54.433000",
      "content": "<p>Thanks for sharing your solution <a href=\"https://www.kaggle.com/aquatic\" target=\"_blank\">@aquatic</a>! Very insightful and it shows the amount of work that you put in!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1918478,
      "author_name": "Robert Hatch",
      "author_url": "",
      "post_date": "2022-08-29T15:50:58.550000",
      "content": "<p>Wow, you did a lot of great work!</p>\n<p>I know you varied the data for your models, but was it the same 1.2k base after feature selection for all models using the \"normal\" set of data? I wonder if \"less aggressive feature selection\" is another possible thing that, in retrospect, could've helped the private lb score?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1918500,
          "author_name": "Robert Hatch",
          "author_url": "",
          "post_date": "2022-08-29T16:10:03.987000",
          "content": "<p>Note I'm asking because I see so few differences that seem critical (that I noticed) between your solution and top 2 solutions. </p>\n<p>Hmm, maybe top two had more emphasis on a normalization factor? You had monthly based ranking, but top solution had ranking of a single customer's statements across months AND ranking across customers for a given month, AND aggregated those things. (Also could feature selection have removed some of the monthly ranking or any other normalization that you did do?)</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1918718,
          "author_name": "Joe Eddy",
          "author_url": "",
          "post_date": "2022-08-29T19:17:40.697000",
          "content": "<p>Thanks!</p>\n<p>That method was specifically what <a href=\"https://www.kaggle.com/ryotak12\" target=\"_blank\">@ryotak12</a> used for selection (especially for some of the feature sets that would otherwise just be enormous). For my models, I actually didn't use selection, more like feed-forward additions (and actually included different/older versions of models that had a subset of the total feature set I used, these made the cut in the ensemble). Feature selection is hard, especially on this problem, and I thought it'd be more time efficient to prioritize diversity for my work. </p>\n<p>If I'm remembering correctly, top 2 solutions both didn't use stacking so that may be one differentiating point. Perhaps our stacking overfits to CV / Public LB relative to the highest scoring Private solutions.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1919011,
          "author_name": "Ryota",
          "author_url": "",
          "post_date": "2022-08-30T04:09:45.923000",
          "content": "<p>As Joe said, only my feature set was applied to this selection method. <br>\nI generated a large number of features at beginning and could not train them all 20k over features.<br>\nSo, I had to select features due to training time.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1917789,
      "author_name": "Evan",
      "author_url": "",
      "post_date": "2022-08-29T04:39:57.480000",
      "content": "<p>Congratulations Joe!<br>\nCould you please share more about feature selection part? After you get the permutation importance of 20k features, how do you select 1200 features? Did you just pick the top 1200 features? If then, how do you decide how many features? Thanks </p>",
      "votes": 1,
      "replies": [
        {
          "id": 1917830,
          "author_name": "Ryota",
          "author_url": "",
          "post_date": "2022-08-29T05:24:58.427000",
          "content": "<p>I will answer your question.</p>\n<p>I could not train 20k features at once, because it takes too long time and memory.<br>\nSo, at first, I divided it into 15 subsets and experimented with 1500 features each.<br>\nAt this stage, I selected features that decrease by more than 1e-5 in G-metric.<br>\nThe reason that this threshold was chosen is that experiments had shown that features that do not contribute to improving metric can also results -1e-5.<br>\nThese are repeated and finally the experiment is conducted with a single set of features. And 1200 features were obtained.</p>",
          "votes": 6,
          "replies": []
        },
        {
          "id": 1918356,
          "author_name": "Evan",
          "author_url": "",
          "post_date": "2022-08-29T14:25:40.043000",
          "content": "<p>Thanks for your reply.<br>\nChosing negative 1e-5 by experiments is really a fantastic idea! Usually people drop features with negative G-metric.<br>\nMay I ask when training using 1500 features, is it necessary to train 5 models if I use 5 folds? In that case, it would be very time-consuming. <br>\nAlso do you permutate feature on the oof data? It seems to me permutation importance on the oof would be better, cause it represents the ability a feature can generalize on unseen data or not.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1918373,
          "author_name": "Ryota",
          "author_url": "",
          "post_date": "2022-08-29T14:44:37.483000",
          "content": "<p>I used 3 folds, and learning_rate=0.1.<br>\nThis still took a huge amount of time for feature selection.<br>\nIf I had used RAPIDs FIL, it would have been faster. But I didn't know.<br>\n(Chris introduces it in his solution. <a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/347641\" target=\"_blank\">https://www.kaggle.com/competitions/amex-default-prediction/discussion/347641</a>)</p>\n<p>Yes, oof data.</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 1917636,
      "author_name": "C4rl05/V",
      "author_url": "",
      "post_date": "2022-08-28T23:57:33.703000",
      "content": "<p>Hello <a href=\"https://www.kaggle.com/aquatic\" target=\"_blank\">@aquatic</a>, impressive model architecture; thanks for sharing</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1917630,
      "author_name": "Chris Deotte",
      "author_url": "",
      "post_date": "2022-08-28T23:37:00.770000",
      "content": "<p>Congratulations Joe and team. I'm sorry to see you drop on private. I was excited to see your huge climb on public LB in the last days.</p>\n<p>Can you explain \"residual learning\" more? Is this training a model on the error of another model?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1917645,
          "author_name": "Joe Eddy",
          "author_url": "",
          "post_date": "2022-08-29T00:18:24.850000",
          "content": "<p>Thanks for the kind words Chris!</p>\n<p><a href=\"https://www.kaggle.com/andrew60909\" target=\"_blank\">@andrew60909</a> brought this method to the team and referenced the <a href=\"https://www.kaggle.com/c/mercari-price-suggestion-challenge/discussion/50256\" target=\"_blank\">winning Mercari solution</a>. That's exactly it, we took OOF predictions, trained a regression LGBM model on some of our feature sets to predict <code>resid = y_train - oof_pred_train</code>, and then generated the test predictions as <code>pred_test + pred_resid_test</code> (and new OOF as <code>oof_pred_val + pred_resid_val</code>). This improved our TabNet and Transformer models by .001-.002 while still yielding a diverse final model. We did find it to be a little finicky - it overfits quickly so needs to be well regularized, and I think only really good features help. </p>\n<p>Would love to hear if Angus has more thoughts, but to me it's almost like a mini post-processor, or a hybrid gradient boosting model where you're boosting off of a strong prediction initial prediction instead of the usual best constant prediction starting point. </p>",
          "votes": 6,
          "replies": []
        },
        {
          "id": 1917725,
          "author_name": "Angus Chang",
          "author_url": "",
          "post_date": "2022-08-29T02:47:59.393000",
          "content": "<p>Just as Joe said - it's actually more like post-processing, but it can generate pretty good diversity and boost the score of the relatively weak learner. We can also consider it as a kind of non-linear stacking. As for final ensembling, having these 2 models simultaneously will be a good idea too.</p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 1918493,
          "author_name": "Robert Hatch",
          "author_url": "",
          "post_date": "2022-08-29T16:02:23.080000",
          "content": "<p>If I understand correctly, you can also look at it as what LGBM (or XGB) is doing every round anyways. Each tree is predicting to eliminate the residual up to that point. </p>\n<p>Joe, Angus, do you know if this has been tried much at all on the output of the entire final ensemble? Can it be effective as the final step?</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1918723,
          "author_name": "Joe Eddy",
          "author_url": "",
          "post_date": "2022-08-29T19:26:23.230000",
          "content": "<p>Haven't tried this, but I wouldn't expect it to help an ensemble. It should work if there is a systematic way to use features and a different model/view to explain errors made by a specific model/view, but I would think everything systematic is already in an ensemble that includes the GBT methods as base models. Also the error margins that the method can work with just become slimmer and slimmer as the model its applied to improves. I think part of why we saw it work for Transformer and TabNet but not MLP is that the data view for the former two is further removed from the flat aggregated data + GBT view.</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 1920878,
      "author_name": "Yang Liu",
      "author_url": "",
      "post_date": "2022-08-31T13:04:42.617000",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/aquatic\" target=\"_blank\">@aquatic</a>, May I invite you to participate in this survey regarding your experience on Kaggle (10 min)? This is not a scam. We are a group of researchers at the City University of Hong Kong. The survey link is: <a href=\"https://cityuhk.questionpro.com/survey-of-kaggle-contestants\" target=\"_blank\">https://cityuhk.questionpro.com/survey-of-kaggle-contestants</a></p>",
      "votes": -3,
      "replies": []
    },
    {
      "id": 1917646,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-08-29T00:21:06.407000",
      "content": "",
      "votes": 1,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1917580": "A huge huge thank you to @andrew60909 and @ryotak12 for an excellent team experience. We learned a lot and had fun, and I appreciate their patience with my extended, largely unsuccessful efforts to make a transformer work really well 😂. Thanks also to Amex and Kaggle for hosting this interesting competition and to @raddar and everyone else who shared resources that made this competition much more accessible. \n\nI've made my personal code public here: https://github.com/JEddy92/amex_default_kaggle. I tried to create a reasonably general structure for division of FE and modeling, with reusable functionality for training and logging results across a variety of models. There is still some clean up to do and it's not as polished as would be ideal!\n\nFeel free to ask us anything if you have questions about our approach!\n\n## Solution Summary\n\nWe used logistic regression to ensemble ~60 sets of predictions with diversity obtained from different models, feature sets, and data views. Models included the usual mix of gradient boosters (LGBM, LGBM dart, XGBoost, CatBoost), MLPs, TabNet, and Transformers with a few special tricks. Data views included aggregated data, fully flattened data (13 raw features per customer), sequential data (Transformer), and augmented data (shifted statements forward by 1 and appended to original). Very detailed feature information to follow -- they are split across our 3 different team members' different models, but this aims to be a complete compilation of what we all used. \n\n\n<hr style=\"border:2px solid gray\">\n\n## Datasets & Notation\n* Everything is derived from Raddar's dataset\n\n\\\\(\\mathcal{D}_{fulltrain}\\\\) : Original full size train-set (5531451 rows)   \n\n\\\\(\\mathcal{D}_{train}\\\\) : Train-set (458913 rows)    \n\n\\\\(\\mathcal{D}_{fulltest}\\\\) : Original full size test-set   \n\n\\\\(\\mathcal{D}_{test}\\\\) : Test-set  \n\n\n<hr style=\"border:2px solid gray\">\n\n## Primary Features\nThe train-sets with the best CV score for each of us are generated by \\\\(\\mathcal{D}_{fulltrain}\\\\) with these features:  \n  \n* **Numeric**     \n  + Mean, std, max, min, first, last  \n  + Last-mean, last/mean, last/first, last/std, max/min, linear-weighted mean \n  + Similar aggregates as above but limited to last 3, 5, etc. statements\n  + Recent diff features and aggregates of all diffs \n  + Monthly-based ranking\n  + Multiplication and Division: create all possible pairs for all features, except for categorical features and features highly correlated with other features. Multiplication and division is done after scaling with MinMaxScaler(0-1) and clipping outliers\n  + Null aggregates (including starting streak vs. at random)\n  + Date features (e.g. stats on customer's statement gaps)\n  \n* **Trend** :  Oldest month as 1 and the newest month as 13, take the average over the following periods   \n    - mean(13,12) - mean(11,10)\n    + mean(13,12,11) - mean(10,9,8)\n    + mean(13,12,11) - mean(3,2,1)\n    + mean(13,12,11,10,9,8) - mean(7,6,5,4,3,2)\n\n* **Category**  \n    + first, last  \n    + entropy: Shannon entropy on frequency table   \n    + Nunique: Number of unique category values\n    + Counts of each value in each category (both raw counts and tf-idf counts)\n    + `P_2` mean (across train and test) of each category\n    + SVD factorization over time of each category\n    + freq1name : the number of most frequent category  \n    + freq1ratio : the number of most frequent category / group size  \n    + freq_last1name : the number of least frequent category  \n    + freq_last1ratio : the number of least frequent category / group size  \n  \n* **PCA**\n    + (N, 13*raw_features) → (N, 64)\n  \n* **KNN-based target encoding methods**\n    + We selected the nearest 500 points for each sample with euclidean distance(with 2 groups of most important features, each has 14 features. One with last statement of each customer, another one with mean of all statements), calculate the average of target as features\n    + Follow the above method, using nearest 500 points but applying average with weight by the distance matrix\n    + Aggregate the distance matrix, by mean, max, min, std\n    + We also calculated the cosine similarity of each customer's past N months of data in one dimension.(N=1,3,6,13) The average of the 500 nearest neighbor customers' targets is used as the feature. This feature was binned because the distribution is somewhat different between Train and Test\n    \n    \n* **Nested model**\n\n    + We added the labels to \\\\(\\mathcal{D}_{fulltrain}\\\\), then trained a LGBM model. The purpose of doing this is to capture what kinds of records, behavior and attributes will cause default. Instead of aggregating features into 458913 rows, this method can let the model learn some extra information from a more \"base\" level. After that, for each `customer_ID` we will have a \"predicted target\" that has the same length as the number of records in each `customer_ID`.   \n    \n    + Finally, to merge it into \\\\(\\mathcal{D}_{train}\\\\) : Train-set (458913 rows), we aggregate similarly to other features: mean, std, max, min, first, last, man-min, last-mean, last/std.  \n    \n    + An interesting point here is that the last value of our \"predicted value\" with \\\\(\\mathcal{D}_{fulltrain}\\\\) scores 0.785 on the Amex metric, while the score of the max value is around 0.62 -- indicating the importance of the last value.\n      \n<hr style=\"border:2px solid gray\">       \n\n## Other Feature Methods: \n\n* **Target/Count encoding**: On \\\\(\\mathcal{D}_{train}\\\\)\n\n* **Target encoding** with \\\\(\\mathcal{D}_{fulltrain}\\\\)\n\n* **Other functions to aggregate** \\\\(\\mathcal{D}_{fulltrain}\\\\)\n    + Exponential weight average\n    + Skewness, Kurtosis\n    + The ratio beyond 1std \n    + Median Absolute Deviation\n    + Mean, max, min, std on 1st derivative(t=1,2,...13), timestamp of max and min also used \n    + Shannon entropy: estimate of the spectral density of sequence\n    + Stability, Lumpiness: variance of the means and variance of the variances on tiles of windows\n    + Crossing points: number of times a sequence crosses the median line\n    + Some measurements from [here](http://isadoranun.github.io/tsfeat/FeaturesDocumentation.html) \n    \n* **Clustering methods**: Kmeans, DBSCAN on important feature(e.g. `P_2_last` or `B_1_last`(including other measurements)). Using raw clusters largely degraded our CV and we're not sure about the reason, so count/target encoding was used here.\n\n* **pred0 feature**: using full data to predict specific important features via 2 datasets: \\\\(\\mathcal{D}_{train}\\\\) and the fullsize one. For example, if we want to predict `P_2` in the former case or predict `P_2_last` in the latter case, we will excluded all \"P-related\" feature and train a lgbm model on it.\n\n* **Probit, Logit models with regularization** on different parts of features (e.g. S-related, B-related sets), like how we make pred0 features, we also exclude related features when we construct the predictions\n\n* **Time-series based methods** to predict important features in the next statement(e.g. `P_2`, `B_1`....)\n\n* **Curve fitting, smoothing, denoising**:\n    + Linear fit: return coefficient and the value of the next 1,3,6,12 months as features\n    + Poly fit(2,3,4): take the coefficient, and the prediction in next 1 month\n    + LOWESS fit: only with the `customer_ID` that have >10 statements.\n    + Wavelet based method of denoising: Least Asymmetric, Haar, Daubechies(16) then construct some agg features, this method only applies on several important features\n\n* **Leaf-embedding**: An old trick. We mixed the approaches from [this paper](https://scontent-tpe1-1.xx.fbcdn.net/v/t39.8562-6/240842589_204052295113548_74168590424110542_n.pdf?_nc_cat=109&ccb=1-7&_nc_sid=ad8a9d&_nc_ohc=nd2mJEAuPSkAX-EVY6m&_nc_ht=scontent-tpe1-1.xx&oh=00_AT9NGACeKxD36CIc-jVHSZ_hyBN5NsVETQihOZqmm7nZ1A&oe=630C598A) and [this kaggle code](https://www.kaggle.com/code/mmueller/categorical-embedding-with-xgb). We use lgbm with 1round + huge amount of leaves + super high regularization to embed the \"whole\" data, which has only 1 dimension. And of course using only categorical data or expanding the embedding matrix to the dimension of `nrounds` is also possible. \n        \n* Most of the above methods are not really able to improve the best CV in a single model, so we turn to utilize it in different models to generate more diversity. Most of our lgbm models with the above selected methods can have a score ranging from CV.796~.801/ PbLB.797~.80 / PriLB.804~.806,  and some of them can generate pretty good diversity. But we are not going to do a detailed ablation study to check the improvement in the PrivateLB of each method :P\n\n\n<hr style=\"border:2px solid gray\">\n\n## Feature Selection Strategies\n\n* **LightGBM w/ permutation importance**: During the training process, we observed that the metric D is quite noisy and unstable. Probably as a result, permutation importance selection with Amex-Score didn't work here. So we selected features by monitoring G (Gini score) alone and it worked much better! For example, Ryota's FE generated over 20k features, but the final number of features utilized was about 1200.\n\n<hr style=\"border:2px solid gray\">    \n\n## Model Details\n\n* **GBDT**\n    - LightGBM (gbdt, dart) - classification, regression: CV 0.796~0.801, Public 0.799, Private 0.805~.806\n    - XGBoost - classification: CV 0.797~0.800, Public 0.798, Private 0.805\n    - CatBoost - classification, regression CV 0.796~0.799, Public 0.797, Private 0.804\n    - LGB linear_tree, ExtraTrees, RF, RGF and rmse objective to generate some diversity\n    - Some LGB dart models trained on the fully flattened view of the data (don't score well individually but ensemble well)\n    - Some LGB dart models trained on features derived from augmented data view \n\n* **TabNet**\n    - TabNet: CV 0.793 Public 0.794 Private 0.801\n    - Residual Learning (LightGBM): CV 0.795, Public 0.796 Private 0.802 (specifically, we predict `ground truth - tabnet’s prediction` with LightGBM and then finalize predictions as `tabnet's prediction + resid prediction`\n    \n* **MLP**\n    - MLP: CV 0.796 Public 0.794 Private 0.802\n\n* **Transformer**\n\t- CV: .793  -> .794 with residual learning\t\n\t- All nulls imputed with LGB trained on all other features, categories encoded as mean `P_2` value from entire train+test. Data Augmented ~2x by statement shift strategy\n\nExample Dart hyper-parameters from Angus, no special sauce here:)  :\n\n```\nlgb_param <- list(boosting_type = 'dart',\n                  objective = \"binary\",\n                  metric = amex,\n                  learning_rate = 0.02,\n                  num_leaves = 48,\n                  feature_fraction = 0.1,\n                  #bagging_freq = 1,\n                  #bagging_fraction = 1,\n                  min_child_weight = 1,\n                  lambda_l1 = 1,\n                  lambda_l2 = 64,\n                  skip_drop = 0.8\n)\n```\n\n<hr style=\"border:2px solid gray\">\n\n## Some Other Interesting Things we Tried\n\n1. **Private set specialization**: We tried using the KNN result mentioned above to estimate which samples are closer to Private LB data. In particular, we calculated the nearest 500 points for each sample, then  the ratio of each set in those 500 points. After that, we took the top 30% of samples which are closer to private LB and ran a logistic regression to decide the weight of our final ensemble. But unfortunately, the CV score dropped a lot. The CV of the top 30% data is ~0.88, but overall(450k data) Amex-score is ~0.7812. We also tried the top 30% \"public ratio\" one. The CV is ~0.7259 and the overall CV is 0.7896711. It's possible that using adversarial validation predictions trained on top importance features would be a more reliable option.\n\n2. **Transfer learning on `P_2`**: this idea unfortunately came up very close to the end before it could pay off and probably was worth pursuing further. This was an attempt at exploiting the test data - pretrain a transformer on train+test to predict `P_2` (because it's strongly related to the target), then fine-tune on the actual target with train. This is like a worse version of the knowledge distillation techniques that were used to exploit the test data more for neural net training. \n\n2. **Psuedo labeling**: We only got 0.00005 improvement on CV (LB doesn't seem tio change), and we have tuned the size and what proportion of test-set can be used for it but didn't have real luck. However, log loss and AUC may be high enough to do this more optimally and get a considerable improvement.\n\n3. **Ensembling methods**, we tried lgbm with 50 seeds bagging, LR and optimization methods like `md1*par1 + md2*par2....` with Nelder-Mead or L-BFGS-B solver. LR performed the best (CV.8031), and lgbm seemed to be severely overfitting. \n\n<hr style=\"border:2px solid gray\">    \n\n## Reflections\n\nWe fell hard from 3rd Public -> 21st Private, with 6 submissions that would have scored gold. In hindsight, CV here did not seem to be a great measurement for the private LB, even though CV improvement aligned very well with public LB. We probably had some bad luck, but 2 things we may have benefited from doing differently are:\n- Exploiting the test data more with knowledge distillation, pretraining, etc.\n- Hedged our submission strategy more - the two submissions we chose (best Public LB and best CV) were not that different from each other, and choosing a different backup could have been a better way to game the noisiness here\n\n\n<br/>\n\n\n",
    "1921355": "thanks! its very helpful for beginner how to approach problem",
    "1920201": "Thanks for sharing your solution @aquatic! Very insightful and it shows the amount of work that you put in!",
    "1918478": "Wow, you did a lot of great work!\n\nI know you varied the data for your models, but was it the same 1.2k base after feature selection for all models using the \"normal\" set of data? I wonder if \"less aggressive feature selection\" is another possible thing that, in retrospect, could've helped the private lb score?",
    "1917789": "Congratulations Joe!\nCould you please share more about feature selection part? After you get the permutation importance of 20k features, how do you select 1200 features? Did you just pick the top 1200 features? If then, how do you decide how many features? Thanks \n",
    "1917636": "Hello @aquatic, impressive model architecture; thanks for sharing",
    "1917630": "Congratulations Joe and team. I'm sorry to see you drop on private. I was excited to see your huge climb on public LB in the last days.\n\nCan you explain \"residual learning\" more? Is this training a model on the error of another model?",
    "1920878": "Hi @aquatic, May I invite you to participate in this survey regarding your experience on Kaggle (10 min)? This is not a scam. We are a group of researchers at the City University of Hong Kong. The survey link is: https://cityuhk.questionpro.com/survey-of-kaggle-contestants",
    "1917646": ""
  }
}