{
  "id": 347966,
  "title": "45th place with XGBoost in first Kaggle competition",
  "url": "/competitions/amex-default-prediction/writeups/robert-hatch-45th-place-with-xgboost-in-first-kagg",
  "author_name": "",
  "post_date": "2022-08-30T21:02:09.223Z",
  "votes": 40,
  "comment_count": 4,
  "views": 0,
  "content": "<p>Some luck and some innovation lead to a surprising top 1% in my first competition!</p>\n<p>49th overall, 5th out of solo rookies. 3rd out of amateur solo rookies. (7th and 10th place overall were soloing their first kaggle competition, but profile shows they are professional data scientists.) To say the least, I'm very happy with my result! Who knows, but I might've had one of the top 10 or top 20 single models with my 0.80798 XGB score on private.</p>\n<p>I'm a professional software engineer at Intel, but only a ML hobbyist (started with stock market related some years ago), so professional problem solver, amateur at ML. I spent far too much time on this competition, but a lot of it struggling with basics like \"Oh, that's not a pandas df, it's a cudf df. Now I can look at the right documentation. …Oh, the kaggle version of rapids cudf (20.x) doesn't have this function call that I'm staring at the cudf v21.x documentation of right this second.\" Only more painful and recurring small issues even than it sounds. I didn't touch NN. I touched LGBM along with XGB, but even spending time on two models was a bit more than I could easily do.</p>\n<p>I'll cover my submission report first, and hopefully later add my general takeaways from my first experience with kaggle competitions. I like semi-stream of consciousness long-winded posts, so buckle up! :)</p>\n<h1>Submission Report</h1>\n<h2>Score and Result</h2>\n<p>49th place. 0.80798 private LB. 0.79889 public.</p>\n<p>Went from 406th (ensemble of two of mine with LGBM dart public) -&gt; 49th (my standalone model submission).</p>\n<h2>Solution Overview</h2>\n<p>My solution was pure XGBoost model for each step. At the high level, ignoring the timeline of my journey of discovery, I did these things:</p>\n<ul>\n<li>Note: I used auc score not logloss nor amex metric for any optimization or tuning along the way</li>\n<li>Created the XGB Pyramid via pathfinding and pure theorycrafting.<ul>\n<li><a href=\"https://www.kaggle.com/code/roberthatch/pyramid-api-for-easy-deployment\" target=\"_blank\">https://www.kaggle.com/code/roberthatch/pyramid-api-for-easy-deployment</a></li></ul></li>\n<li>Basic popular feature aggregation and a few of my own. Most notably moving averages, though I only used them on all statements. Hull moving average, and exponential averages. 16 total aggregations per base numerical feature.<ul>\n<li><a href=\"https://www.kaggle.com/code/roberthatch/amex-feature-engg-gpu-or-cpu-process-in-chunks\" target=\"_blank\">https://www.kaggle.com/code/roberthatch/amex-feature-engg-gpu-or-cpu-process-in-chunks</a> </li>\n<li><a href=\"https://www.kaggle.com/code/roberthatch/exponential-averages-amex-feature-engineering\" target=\"_blank\">https://www.kaggle.com/code/roberthatch/exponential-averages-amex-feature-engineering</a> </li>\n<li>I dropped B_29 and never looked back!</li></ul></li>\n<li>Meta-feature: Predict for every row in train and every row in test, the chance of missing the next month's payment, meaning days overdue increase by a large value, and ending at or over 28.<ul>\n<li><a href=\"https://www.kaggle.com/code/roberthatch/amex-fe-02-days-overdue-label\" target=\"_blank\">https://www.kaggle.com/code/roberthatch/amex-fe-02-days-overdue-label</a> </li>\n<li>Inspired by: <a href=\"https://www.kaggle.com/code/raddar/deanonymized-days-overdue-feat-amex\" target=\"_blank\">https://www.kaggle.com/code/raddar/deanonymized-days-overdue-feat-amex</a> </li>\n<li>Not just using that row's data to predict, I wanted to use that row AND all past data via backwards aggregation. To fit in memory, forward feature selection (on normal target predictions) using various shortcuts, ending with 280 features.</li>\n<li>I also predict the chance that next month's days overdue will be non-zero as an independent meta-feature. </li>\n<li>Low on GPU and time ~48 hours to go, just do single five fold oof predictions. Better to average 3-5 models. Could predict last statement with all models, since they couldn't be used for training, but for simplicity only predict \"oof\" last.</li></ul></li>\n<li>Main model was step two, taking the aggregated meta-features (2*16) and the 3000 other features, convert to float 16 for memory, and run the XGBoost Pyramid.<ul>\n<li>Train 4 models on entire train dataset with no CV using a set number of rounds based on inspecting when early stopping happened on CV models.</li></ul></li>\n</ul>\n<h2>Notes</h2>\n<p>I had the hardest time doing (and failing at) permutation importance. My biggest bottleneck was my own time, and I didn't want to write it from scratch, but I couldn't use sklearn with my xgb model, and the things I tried kept failing for various reasons (including attempting to do an sklearn version of my xgb model, and including trying to leverage a non-sklearn version of permutation importance library). And I think I gave up on the non-sklearn one just because it was so slow.</p>\n<p>In any case, thats why I ended up doing homebrew forward feature selection. Which I spent way WAY too much time on (but was kinda fun). At first I did selection only from last. Then added max, then e7. (from other experiments I had done forward feature selection of entire aggregation styles, and if forced to do only 3x aggregations, last, max, e7 was best for me. I got 70 features individually before finally stopping that slow approach. That allowed me to split the base columns into three groups, the \"good\" the \"decent\" and the \"didn't seem good\". So I did grouped forward feature selection based on the 16 aggregations, done on 1 of the three base feature subgroups, so 48 groups in total of 40-80 features each. I got to 3xx, but ran out of memory and to keep moving forward backed up to 280 features total. Not all base features were represented at all.</p>\n<p>I had a ton of other random ideas, but all along was convinced that predicting using the test sets and days overdue was my best single shot at a good idea and great score. I didn't have time to try other good ideas. I may still keep going on this competition for more fun, even if it's over. :)</p>\n<p>I used really small learning rate 0.005. Combined with XGB pyramid meant I was doing 10 rounds of 200 tree boosted forests to start, and along with a lot in the middle, ending with 0.0025 learning rate for last 9600 rounds. So each single model run was a bit of an ensemble by itself.</p>\n<p>Train 5 times on entire dataset vs 5 fold CV didn't seem much different, maybe marginally better, when I tried it with prior model, but I didn't have much time at the end so just went with it. Maybe 10 fold CV would be better, good diversity and still train on almost all data.</p>\n<p>I trained two models at the end, one that dropped R_1, S_11, D_59, S_9. I was hoping that days overdue mega-feature would reduce the need for them at step two, and thus eliminate the concern of private LB shakeup on those features. However, it got much worse score on public LB, so I wasn't willing to try it, and did submit my best model in the end.</p>\n<p>The next thing I really wanted to try was inspired by reading about a prior credit kaggle competition, I wanted to use KNN to create features.</p>\n<p>Especially after reading other people's great ideas, I really wonder if multiple ways of using diverse approaches to extract features from the 13 statements would be ideal.</p>\n<p>In other words, for each customer and base feature col with 13 statements, do things like:</p>\n<ul>\n<li>Predict a few different things:<ul>\n<li>'days overdue' (or predict target, but then you have to be careful with nested folds)</li>\n<li>P_2</li>\n<li>predict next in sequence</li></ul></li>\n<li>Using a few approaches:<ul>\n<li>XGB or LGBM</li>\n<li>LSTM, RNN or something</li>\n<li>KNN</li>\n<li>linear regression</li>\n<li>predict next in sequence with simple least squares line</li></ul></li>\n<li>Using data diversity:<ul>\n<li>besides using all 13 statements, what about using 3, 4, 5, and 6 statements to predict <em>next</em> on: days overdue, P_2, and next in sequence? This gets a lot more than 1 training sample per customer for training each of those models, and should do well with the \"why not both?\" approach that throws all features to the model to figure out.</li></ul></li>\n</ul>\n<p>Other thoughts:</p>\n<ul>\n<li><p>Maybe a good approach would've been to spend a lot more time trying to create any and all aggregate features…. on P_2 only, and see how well I could predict using ONLY P_2? Then create meta features aka mega features like days overdue prediction, and use what worked well on P_2 on that and maybe on all feature columns.</p></li>\n<li><p>I probably made a couple interesting mistakes:</p>\n<ul>\n<li>I should've allowed B_29 in the days overdue model, I think. I forgot to try it.</li>\n<li>I allowed everything to be float16. I think that might've really hurt on my super feature for days overdue prediction, I should've allowed the top few features (at least) to be float32.</li></ul></li>\n</ul>",
  "messages": [
    {
      "id": "1914518",
      "postDate": "08/26/2022 06:43:39",
      "content": "<p>Some luck and some innovation lead to a surprising top 1% in my first competition!</p>\n<p>49th overall, 5th out of solo rookies. 3rd out of amateur solo rookies. (7th and 10th place overall were soloing their first kaggle competition, but profile shows they are professional data scientists.) To say the least, I'm very happy with my result! Who knows, but I might've had one of the top 10 or top 20 single models with my 0.80798 XGB score on private.</p>\n<p>I'm a professional software engineer at Intel, but only a ML hobbyist (started with stock market related some years ago), so professional problem solver, amateur at ML. I spent far too much time on this competition, but a lot of it struggling with basics like \"Oh, that's not a pandas df, it's a cudf df. Now I can look at the right documentation. …Oh, the kaggle version of rapids cudf (20.x) doesn't have this function call that I'm staring at the cudf v21.x documentation of right this second.\" Only more painful and recurring small issues even than it sounds. I didn't touch NN. I touched LGBM along with XGB, but even spending time on two models was a bit more than I could easily do.</p>\n<p>I'll cover my submission report first, and hopefully later add my general takeaways from my first experience with kaggle competitions. I like semi-stream of consciousness long-winded posts, so buckle up! :)</p>\n<h1>Submission Report</h1>\n<h2>Score and Result</h2>\n<p>49th place. 0.80798 private LB. 0.79889 public.</p>\n<p>Went from 406th (ensemble of two of mine with LGBM dart public) -&gt; 49th (my standalone model submission).</p>\n<h2>Solution Overview</h2>\n<p>My solution was pure XGBoost model for each step. At the high level, ignoring the timeline of my journey of discovery, I did these things:</p>\n<ul>\n<li>Note: I used auc score not logloss nor amex metric for any optimization or tuning along the way</li>\n<li>Created the XGB Pyramid via pathfinding and pure theorycrafting.<ul>\n<li><a href=\"https://www.kaggle.com/code/roberthatch/pyramid-api-for-easy-deployment\" target=\"_blank\">https://www.kaggle.com/code/roberthatch/pyramid-api-for-easy-deployment</a></li></ul></li>\n<li>Basic popular feature aggregation and a few of my own. Most notably moving averages, though I only used them on all statements. Hull moving average, and exponential averages. 16 total aggregations per base numerical feature.<ul>\n<li><a href=\"https://www.kaggle.com/code/roberthatch/amex-feature-engg-gpu-or-cpu-process-in-chunks\" target=\"_blank\">https://www.kaggle.com/code/roberthatch/amex-feature-engg-gpu-or-cpu-process-in-chunks</a> </li>\n<li><a href=\"https://www.kaggle.com/code/roberthatch/exponential-averages-amex-feature-engineering\" target=\"_blank\">https://www.kaggle.com/code/roberthatch/exponential-averages-amex-feature-engineering</a> </li>\n<li>I dropped B_29 and never looked back!</li></ul></li>\n<li>Meta-feature: Predict for every row in train and every row in test, the chance of missing the next month's payment, meaning days overdue increase by a large value, and ending at or over 28.<ul>\n<li><a href=\"https://www.kaggle.com/code/roberthatch/amex-fe-02-days-overdue-label\" target=\"_blank\">https://www.kaggle.com/code/roberthatch/amex-fe-02-days-overdue-label</a> </li>\n<li>Inspired by: <a href=\"https://www.kaggle.com/code/raddar/deanonymized-days-overdue-feat-amex\" target=\"_blank\">https://www.kaggle.com/code/raddar/deanonymized-days-overdue-feat-amex</a> </li>\n<li>Not just using that row's data to predict, I wanted to use that row AND all past data via backwards aggregation. To fit in memory, forward feature selection (on normal target predictions) using various shortcuts, ending with 280 features.</li>\n<li>I also predict the chance that next month's days overdue will be non-zero as an independent meta-feature. </li>\n<li>Low on GPU and time ~48 hours to go, just do single five fold oof predictions. Better to average 3-5 models. Could predict last statement with all models, since they couldn't be used for training, but for simplicity only predict \"oof\" last.</li></ul></li>\n<li>Main model was step two, taking the aggregated meta-features (2*16) and the 3000 other features, convert to float 16 for memory, and run the XGBoost Pyramid.<ul>\n<li>Train 4 models on entire train dataset with no CV using a set number of rounds based on inspecting when early stopping happened on CV models.</li></ul></li>\n</ul>\n<h2>Notes</h2>\n<p>I had the hardest time doing (and failing at) permutation importance. My biggest bottleneck was my own time, and I didn't want to write it from scratch, but I couldn't use sklearn with my xgb model, and the things I tried kept failing for various reasons (including attempting to do an sklearn version of my xgb model, and including trying to leverage a non-sklearn version of permutation importance library). And I think I gave up on the non-sklearn one just because it was so slow.</p>\n<p>In any case, thats why I ended up doing homebrew forward feature selection. Which I spent way WAY too much time on (but was kinda fun). At first I did selection only from last. Then added max, then e7. (from other experiments I had done forward feature selection of entire aggregation styles, and if forced to do only 3x aggregations, last, max, e7 was best for me. I got 70 features individually before finally stopping that slow approach. That allowed me to split the base columns into three groups, the \"good\" the \"decent\" and the \"didn't seem good\". So I did grouped forward feature selection based on the 16 aggregations, done on 1 of the three base feature subgroups, so 48 groups in total of 40-80 features each. I got to 3xx, but ran out of memory and to keep moving forward backed up to 280 features total. Not all base features were represented at all.</p>\n<p>I had a ton of other random ideas, but all along was convinced that predicting using the test sets and days overdue was my best single shot at a good idea and great score. I didn't have time to try other good ideas. I may still keep going on this competition for more fun, even if it's over. :)</p>\n<p>I used really small learning rate 0.005. Combined with XGB pyramid meant I was doing 10 rounds of 200 tree boosted forests to start, and along with a lot in the middle, ending with 0.0025 learning rate for last 9600 rounds. So each single model run was a bit of an ensemble by itself.</p>\n<p>Train 5 times on entire dataset vs 5 fold CV didn't seem much different, maybe marginally better, when I tried it with prior model, but I didn't have much time at the end so just went with it. Maybe 10 fold CV would be better, good diversity and still train on almost all data.</p>\n<p>I trained two models at the end, one that dropped R_1, S_11, D_59, S_9. I was hoping that days overdue mega-feature would reduce the need for them at step two, and thus eliminate the concern of private LB shakeup on those features. However, it got much worse score on public LB, so I wasn't willing to try it, and did submit my best model in the end.</p>\n<p>The next thing I really wanted to try was inspired by reading about a prior credit kaggle competition, I wanted to use KNN to create features.</p>\n<p>Especially after reading other people's great ideas, I really wonder if multiple ways of using diverse approaches to extract features from the 13 statements would be ideal.</p>\n<p>In other words, for each customer and base feature col with 13 statements, do things like:</p>\n<ul>\n<li>Predict a few different things:<ul>\n<li>'days overdue' (or predict target, but then you have to be careful with nested folds)</li>\n<li>P_2</li>\n<li>predict next in sequence</li></ul></li>\n<li>Using a few approaches:<ul>\n<li>XGB or LGBM</li>\n<li>LSTM, RNN or something</li>\n<li>KNN</li>\n<li>linear regression</li>\n<li>predict next in sequence with simple least squares line</li></ul></li>\n<li>Using data diversity:<ul>\n<li>besides using all 13 statements, what about using 3, 4, 5, and 6 statements to predict <em>next</em> on: days overdue, P_2, and next in sequence? This gets a lot more than 1 training sample per customer for training each of those models, and should do well with the \"why not both?\" approach that throws all features to the model to figure out.</li></ul></li>\n</ul>\n<p>Other thoughts:</p>\n<ul>\n<li><p>Maybe a good approach would've been to spend a lot more time trying to create any and all aggregate features…. on P_2 only, and see how well I could predict using ONLY P_2? Then create meta features aka mega features like days overdue prediction, and use what worked well on P_2 on that and maybe on all feature columns.</p></li>\n<li><p>I probably made a couple interesting mistakes:</p>\n<ul>\n<li>I should've allowed B_29 in the days overdue model, I think. I forgot to try it.</li>\n<li>I allowed everything to be float16. I think that might've really hurt on my super feature for days overdue prediction, I should've allowed the top few features (at least) to be float32.</li></ul></li>\n</ul>",
      "rawMarkdown": "Some luck and some innovation lead to a surprising top 1% in my first competition!\n\n49th overall, 5th out of solo rookies. 3rd out of amateur solo rookies. (7th and 10th place overall were soloing their first kaggle competition, but profile shows they are professional data scientists.) To say the least, I'm very happy with my result! Who knows, but I might've had one of the top 10 or top 20 single models with my 0.80798 XGB score on private.\n\nI'm a professional software engineer at Intel, but only a ML hobbyist (started with stock market related some years ago), so professional problem solver, amateur at ML. I spent far too much time on this competition, but a lot of it struggling with basics like \"Oh, that's not a pandas df, it's a cudf df. Now I can look at the right documentation. ...Oh, the kaggle version of rapids cudf (20.x) doesn't have this function call that I'm staring at the cudf v21.x documentation of right this second.\" Only more painful and recurring small issues even than it sounds. I didn't touch NN. I touched LGBM along with XGB, but even spending time on two models was a bit more than I could easily do.\n\nI'll cover my submission report first, and hopefully later add my general takeaways from my first experience with kaggle competitions. I like semi-stream of consciousness long-winded posts, so buckle up! :)\n\n# Submission Report\n## Score and Result\n49th place. 0.80798 private LB. 0.79889 public.\n\nWent from 406th (ensemble of two of mine with LGBM dart public) -> 49th (my standalone model submission).\n\n## Solution Overview\nMy solution was pure XGBoost model for each step. At the high level, ignoring the timeline of my journey of discovery, I did these things:\n* Note: I used auc score not logloss nor amex metric for any optimization or tuning along the way\n* Created the XGB Pyramid via pathfinding and pure theorycrafting.\n  * https://www.kaggle.com/code/roberthatch/pyramid-api-for-easy-deployment\n* Basic popular feature aggregation and a few of my own. Most notably moving averages, though I only used them on all statements. Hull moving average, and exponential averages. 16 total aggregations per base numerical feature.\n  * https://www.kaggle.com/code/roberthatch/amex-feature-engg-gpu-or-cpu-process-in-chunks \n  * https://www.kaggle.com/code/roberthatch/exponential-averages-amex-feature-engineering \n  * I dropped B_29 and never looked back!\n* Meta-feature: Predict for every row in train and every row in test, the chance of missing the next month's payment, meaning days overdue increase by a large value, and ending at or over 28.\n  * https://www.kaggle.com/code/roberthatch/amex-fe-02-days-overdue-label \n  * Inspired by: https://www.kaggle.com/code/raddar/deanonymized-days-overdue-feat-amex \n  * Not just using that row's data to predict, I wanted to use that row AND all past data via backwards aggregation. To fit in memory, forward feature selection (on normal target predictions) using various shortcuts, ending with 280 features.\n  * I also predict the chance that next month's days overdue will be non-zero as an independent meta-feature. \n  * Low on GPU and time ~48 hours to go, just do single five fold oof predictions. Better to average 3-5 models. Could predict last statement with all models, since they couldn't be used for training, but for simplicity only predict \"oof\" last.\n* Main model was step two, taking the aggregated meta-features (2*16) and the 3000 other features, convert to float 16 for memory, and run the XGBoost Pyramid.\n  * Train 4 models on entire train dataset with no CV using a set number of rounds based on inspecting when early stopping happened on CV models.\n\n## Notes\nI had the hardest time doing (and failing at) permutation importance. My biggest bottleneck was my own time, and I didn't want to write it from scratch, but I couldn't use sklearn with my xgb model, and the things I tried kept failing for various reasons (including attempting to do an sklearn version of my xgb model, and including trying to leverage a non-sklearn version of permutation importance library). And I think I gave up on the non-sklearn one just because it was so slow.\n\nIn any case, thats why I ended up doing homebrew forward feature selection. Which I spent way WAY too much time on (but was kinda fun). At first I did selection only from last. Then added max, then e7. (from other experiments I had done forward feature selection of entire aggregation styles, and if forced to do only 3x aggregations, last, max, e7 was best for me. I got 70 features individually before finally stopping that slow approach. That allowed me to split the base columns into three groups, the \"good\" the \"decent\" and the \"didn't seem good\". So I did grouped forward feature selection based on the 16 aggregations, done on 1 of the three base feature subgroups, so 48 groups in total of 40-80 features each. I got to 3xx, but ran out of memory and to keep moving forward backed up to 280 features total. Not all base features were represented at all.\n\nI had a ton of other random ideas, but all along was convinced that predicting using the test sets and days overdue was my best single shot at a good idea and great score. I didn't have time to try other good ideas. I may still keep going on this competition for more fun, even if it's over. :)\n\nI used really small learning rate 0.005. Combined with XGB pyramid meant I was doing 10 rounds of 200 tree boosted forests to start, and along with a lot in the middle, ending with 0.0025 learning rate for last 9600 rounds. So each single model run was a bit of an ensemble by itself.\n\nTrain 5 times on entire dataset vs 5 fold CV didn't seem much different, maybe marginally better, when I tried it with prior model, but I didn't have much time at the end so just went with it. Maybe 10 fold CV would be better, good diversity and still train on almost all data.\n\nI trained two models at the end, one that dropped R_1, S_11, D_59, S_9. I was hoping that days overdue mega-feature would reduce the need for them at step two, and thus eliminate the concern of private LB shakeup on those features. However, it got much worse score on public LB, so I wasn't willing to try it, and did submit my best model in the end.\n\nThe next thing I really wanted to try was inspired by reading about a prior credit kaggle competition, I wanted to use KNN to create features.\n\nEspecially after reading other people's great ideas, I really wonder if multiple ways of using diverse approaches to extract features from the 13 statements would be ideal.\n\nIn other words, for each customer and base feature col with 13 statements, do things like:\n* Predict a few different things:\n  * 'days overdue' (or predict target, but then you have to be careful with nested folds)\n  * P_2\n  * predict next in sequence\n* Using a few approaches:\n  * XGB or LGBM\n  * LSTM, RNN or something\n  * KNN\n  * linear regression\n  * predict next in sequence with simple least squares line\n* Using data diversity:\n  * besides using all 13 statements, what about using 3, 4, 5, and 6 statements to predict *next* on: days overdue, P_2, and next in sequence? This gets a lot more than 1 training sample per customer for training each of those models, and should do well with the \"why not both?\" approach that throws all features to the model to figure out.\n\nOther thoughts:\n* Maybe a good approach would've been to spend a lot more time trying to create any and all aggregate features.... on P_2 only, and see how well I could predict using ONLY P_2? Then create meta features aka mega features like days overdue prediction, and use what worked well on P_2 on that and maybe on all feature columns.\n\n* I probably made a couple interesting mistakes:\n  * I should've allowed B_29 in the days overdue model, I think. I forgot to try it.\n  * I allowed everything to be float16. I think that might've really hurt on my super feature for days overdue prediction, I should've allowed the top few features (at least) to be float32.",
      "votes": null
    },
    {
      "id": "1915435",
      "postDate": "08/27/2022 01:30:10",
      "content": "<p>Thanks for sharing！</p>",
      "rawMarkdown": "Thanks for sharing！",
      "votes": null
    },
    {
      "id": "1919067",
      "postDate": "08/30/2022 05:25:21",
      "content": "<p>Moving averages is something I wish I had more time to explore (joined with &lt;month to go), I did try using compound annual growth rate (CAGR), and it gave me a slight bit of lift on my XG Boost models, but tended to hurt my LGBM models. That being said, given how some of my lower rated models (based on the private LB) did really well on the private LB, I wish I had spent more time refining that approach a bit more. Either way, congrats on winning silver…… and I'm not sure you can still call yourself an amateur…. </p>",
      "rawMarkdown": "Moving averages is something I wish I had more time to explore (joined with <month to go), I did try using compound annual growth rate (CAGR), and it gave me a slight bit of lift on my XG Boost models, but tended to hurt my LGBM models. That being said, given how some of my lower rated models (based on the private LB) did really well on the private LB, I wish I had spent more time refining that approach a bit more. Either way, congrats on winning silver...... and I'm not sure you can still call yourself an amateur....",
      "votes": null
    },
    {
      "id": "1920066",
      "postDate": "08/30/2022 21:11:40",
      "content": "<p>Thanks!</p>\n<p>Amateur I'm only using in the technical sense, as contrast with, for example, second place team of 5 people with day jobs as data scientists. Even there, as I highlighted, I am a professional engineer, a professional problem solver. Just not in the realm of data science. </p>\n<p>Rookie I don't think I can claim anymore, though ;) Though I should probably tune my first NN of my life before I completely give up all claim to rookie status, haha. </p>\n<p>I spent countless hours wrestling with python ML basics this competition, it did have a big impact, even though I did well and got lucky final result. (+0.009 on private LB with my best submission, most scores all go up right about +0.007)</p>",
      "rawMarkdown": "Thanks!\n\nAmateur I'm only using in the technical sense, as contrast with, for example, second place team of 5 people with day jobs as data scientists. Even there, as I highlighted, I am a professional engineer, a professional problem solver. Just not in the realm of data science. \n\nRookie I don't think I can claim anymore, though ;) Though I should probably tune my first NN of my life before I completely give up all claim to rookie status, haha. \n\nI spent countless hours wrestling with python ML basics this competition, it did have a big impact, even though I did well and got lucky final result. (+0.009 on private LB with my best submission, most scores all go up right about +0.007)",
      "votes": null
    },
    {
      "id": "1920881",
      "postDate": "08/31/2022 13:05:28",
      "content": "<p>Hi, thanks for sharing! May I invite you to participate in this survey regarding your experience on Kaggle (10 min)? This is not a scam. We are a group of researchers at the City University of Hong Kong. The survey link is: <a href=\"https://cityuhk.questionpro.com/survey-of-kaggle-contestants\" target=\"_blank\">https://cityuhk.questionpro.com/survey-of-kaggle-contestants</a></p>",
      "rawMarkdown": "Hi, thanks for sharing! May I invite you to participate in this survey regarding your experience on Kaggle (10 min)? This is not a scam. We are a group of researchers at the City University of Hong Kong. The survey link is: https://cityuhk.questionpro.com/survey-of-kaggle-contestants",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1915435,
      "author_name": "iliiiiiili",
      "author_url": "",
      "post_date": "08/27/2022 01:30:10",
      "content": "<p>Thanks for sharing！</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1919067,
      "author_name": "markhamlee",
      "author_url": "",
      "post_date": "08/30/2022 05:25:21",
      "content": "<p>Moving averages is something I wish I had more time to explore (joined with &lt;month to go), I did try using compound annual growth rate (CAGR), and it gave me a slight bit of lift on my XG Boost models, but tended to hurt my LGBM models. That being said, given how some of my lower rated models (based on the private LB) did really well on the private LB, I wish I had spent more time refining that approach a bit more. Either way, congrats on winning silver…… and I'm not sure you can still call yourself an amateur…. </p>",
      "votes": null,
      "replies": [
        {
          "id": 1920066,
          "author_name": "roberthatch",
          "author_url": "",
          "post_date": "08/30/2022 21:11:40",
          "content": "<p>Thanks!</p>\n<p>Amateur I'm only using in the technical sense, as contrast with, for example, second place team of 5 people with day jobs as data scientists. Even there, as I highlighted, I am a professional engineer, a professional problem solver. Just not in the realm of data science. </p>\n<p>Rookie I don't think I can claim anymore, though ;) Though I should probably tune my first NN of my life before I completely give up all claim to rookie status, haha. </p>\n<p>I spent countless hours wrestling with python ML basics this competition, it did have a big impact, even though I did well and got lucky final result. (+0.009 on private LB with my best submission, most scores all go up right about +0.007)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1920881,
      "author_name": "lystriving",
      "author_url": "",
      "post_date": "08/31/2022 13:05:28",
      "content": "<p>Hi, thanks for sharing! May I invite you to participate in this survey regarding your experience on Kaggle (10 min)? This is not a scam. We are a group of researchers at the City University of Hong Kong. The survey link is: <a href=\"https://cityuhk.questionpro.com/survey-of-kaggle-contestants\" target=\"_blank\">https://cityuhk.questionpro.com/survey-of-kaggle-contestants</a></p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1914518": "Some luck and some innovation lead to a surprising top 1% in my first competition!\n\n49th overall, 5th out of solo rookies. 3rd out of amateur solo rookies. (7th and 10th place overall were soloing their first kaggle competition, but profile shows they are professional data scientists.) To say the least, I'm very happy with my result! Who knows, but I might've had one of the top 10 or top 20 single models with my 0.80798 XGB score on private.\n\nI'm a professional software engineer at Intel, but only a ML hobbyist (started with stock market related some years ago), so professional problem solver, amateur at ML. I spent far too much time on this competition, but a lot of it struggling with basics like \"Oh, that's not a pandas df, it's a cudf df. Now I can look at the right documentation. ...Oh, the kaggle version of rapids cudf (20.x) doesn't have this function call that I'm staring at the cudf v21.x documentation of right this second.\" Only more painful and recurring small issues even than it sounds. I didn't touch NN. I touched LGBM along with XGB, but even spending time on two models was a bit more than I could easily do.\n\nI'll cover my submission report first, and hopefully later add my general takeaways from my first experience with kaggle competitions. I like semi-stream of consciousness long-winded posts, so buckle up! :)\n\n# Submission Report\n## Score and Result\n49th place. 0.80798 private LB. 0.79889 public.\n\nWent from 406th (ensemble of two of mine with LGBM dart public) -> 49th (my standalone model submission).\n\n## Solution Overview\nMy solution was pure XGBoost model for each step. At the high level, ignoring the timeline of my journey of discovery, I did these things:\n* Note: I used auc score not logloss nor amex metric for any optimization or tuning along the way\n* Created the XGB Pyramid via pathfinding and pure theorycrafting.\n  * https://www.kaggle.com/code/roberthatch/pyramid-api-for-easy-deployment\n* Basic popular feature aggregation and a few of my own. Most notably moving averages, though I only used them on all statements. Hull moving average, and exponential averages. 16 total aggregations per base numerical feature.\n  * https://www.kaggle.com/code/roberthatch/amex-feature-engg-gpu-or-cpu-process-in-chunks \n  * https://www.kaggle.com/code/roberthatch/exponential-averages-amex-feature-engineering \n  * I dropped B_29 and never looked back!\n* Meta-feature: Predict for every row in train and every row in test, the chance of missing the next month's payment, meaning days overdue increase by a large value, and ending at or over 28.\n  * https://www.kaggle.com/code/roberthatch/amex-fe-02-days-overdue-label \n  * Inspired by: https://www.kaggle.com/code/raddar/deanonymized-days-overdue-feat-amex \n  * Not just using that row's data to predict, I wanted to use that row AND all past data via backwards aggregation. To fit in memory, forward feature selection (on normal target predictions) using various shortcuts, ending with 280 features.\n  * I also predict the chance that next month's days overdue will be non-zero as an independent meta-feature. \n  * Low on GPU and time ~48 hours to go, just do single five fold oof predictions. Better to average 3-5 models. Could predict last statement with all models, since they couldn't be used for training, but for simplicity only predict \"oof\" last.\n* Main model was step two, taking the aggregated meta-features (2*16) and the 3000 other features, convert to float 16 for memory, and run the XGBoost Pyramid.\n  * Train 4 models on entire train dataset with no CV using a set number of rounds based on inspecting when early stopping happened on CV models.\n\n## Notes\nI had the hardest time doing (and failing at) permutation importance. My biggest bottleneck was my own time, and I didn't want to write it from scratch, but I couldn't use sklearn with my xgb model, and the things I tried kept failing for various reasons (including attempting to do an sklearn version of my xgb model, and including trying to leverage a non-sklearn version of permutation importance library). And I think I gave up on the non-sklearn one just because it was so slow.\n\nIn any case, thats why I ended up doing homebrew forward feature selection. Which I spent way WAY too much time on (but was kinda fun). At first I did selection only from last. Then added max, then e7. (from other experiments I had done forward feature selection of entire aggregation styles, and if forced to do only 3x aggregations, last, max, e7 was best for me. I got 70 features individually before finally stopping that slow approach. That allowed me to split the base columns into three groups, the \"good\" the \"decent\" and the \"didn't seem good\". So I did grouped forward feature selection based on the 16 aggregations, done on 1 of the three base feature subgroups, so 48 groups in total of 40-80 features each. I got to 3xx, but ran out of memory and to keep moving forward backed up to 280 features total. Not all base features were represented at all.\n\nI had a ton of other random ideas, but all along was convinced that predicting using the test sets and days overdue was my best single shot at a good idea and great score. I didn't have time to try other good ideas. I may still keep going on this competition for more fun, even if it's over. :)\n\nI used really small learning rate 0.005. Combined with XGB pyramid meant I was doing 10 rounds of 200 tree boosted forests to start, and along with a lot in the middle, ending with 0.0025 learning rate for last 9600 rounds. So each single model run was a bit of an ensemble by itself.\n\nTrain 5 times on entire dataset vs 5 fold CV didn't seem much different, maybe marginally better, when I tried it with prior model, but I didn't have much time at the end so just went with it. Maybe 10 fold CV would be better, good diversity and still train on almost all data.\n\nI trained two models at the end, one that dropped R_1, S_11, D_59, S_9. I was hoping that days overdue mega-feature would reduce the need for them at step two, and thus eliminate the concern of private LB shakeup on those features. However, it got much worse score on public LB, so I wasn't willing to try it, and did submit my best model in the end.\n\nThe next thing I really wanted to try was inspired by reading about a prior credit kaggle competition, I wanted to use KNN to create features.\n\nEspecially after reading other people's great ideas, I really wonder if multiple ways of using diverse approaches to extract features from the 13 statements would be ideal.\n\nIn other words, for each customer and base feature col with 13 statements, do things like:\n* Predict a few different things:\n  * 'days overdue' (or predict target, but then you have to be careful with nested folds)\n  * P_2\n  * predict next in sequence\n* Using a few approaches:\n  * XGB or LGBM\n  * LSTM, RNN or something\n  * KNN\n  * linear regression\n  * predict next in sequence with simple least squares line\n* Using data diversity:\n  * besides using all 13 statements, what about using 3, 4, 5, and 6 statements to predict *next* on: days overdue, P_2, and next in sequence? This gets a lot more than 1 training sample per customer for training each of those models, and should do well with the \"why not both?\" approach that throws all features to the model to figure out.\n\nOther thoughts:\n* Maybe a good approach would've been to spend a lot more time trying to create any and all aggregate features.... on P_2 only, and see how well I could predict using ONLY P_2? Then create meta features aka mega features like days overdue prediction, and use what worked well on P_2 on that and maybe on all feature columns.\n\n* I probably made a couple interesting mistakes:\n  * I should've allowed B_29 in the days overdue model, I think. I forgot to try it.\n  * I allowed everything to be float16. I think that might've really hurt on my super feature for days overdue prediction, I should've allowed the top few features (at least) to be float32.",
    "1915435": "Thanks for sharing！",
    "1919067": "Moving averages is something I wish I had more time to explore (joined with <month to go), I did try using compound annual growth rate (CAGR), and it gave me a slight bit of lift on my XG Boost models, but tended to hurt my LGBM models. That being said, given how some of my lower rated models (based on the private LB) did really well on the private LB, I wish I had spent more time refining that approach a bit more. Either way, congrats on winning silver...... and I'm not sure you can still call yourself an amateur....",
    "1920066": "Thanks!\n\nAmateur I'm only using in the technical sense, as contrast with, for example, second place team of 5 people with day jobs as data scientists. Even there, as I highlighted, I am a professional engineer, a professional problem solver. Just not in the realm of data science. \n\nRookie I don't think I can claim anymore, though ;) Though I should probably tune my first NN of my life before I completely give up all claim to rookie status, haha. \n\nI spent countless hours wrestling with python ML basics this competition, it did have a big impact, even though I did well and got lucky final result. (+0.009 on private LB with my best submission, most scores all go up right about +0.007)",
    "1920881": "Hi, thanks for sharing! May I invite you to participate in this survey regarding your experience on Kaggle (10 min)? This is not a scam. We are a group of researchers at the City University of Hong Kong. The survey link is: https://cityuhk.questionpro.com/survey-of-kaggle-contestants"
  },
  "source": "meta"
}