{
  "id": 347809,
  "title": "What didn't work (and why ?)",
  "url": "/competitions/amex-default-prediction/discussion/347809",
  "author_name": "",
  "post_date": "2022-08-25T13:24:15.012618900Z",
  "votes": 19,
  "comment_count": 3,
  "views": 0,
  "content": "<p>First I wanted to say thanks to everyone that shared insightfull stuff. Kaggle is a very good place to learn from others. </p>\n<p>As I spent quite some time trying different stuff I tought I could share some of them that seemed interesting. As you can see from my final ranking, none of them really worked. Maybe you'll find this interesting ayway, maybe you can learn from my mistakes or share how you made one of the appraoch work. </p>\n<p>As i am generally not interested in the crazy deep ensemble stuff i am not surprised to be nowhere near the medal zone. I am still pretty happy with my results (3 notebook golds / lot of learning). By the way, thanks to every one that took the time to upvote my work.</p>\n<p>In relatively chronological order:</p>\n<p><strong>Clustering</strong></p>\n<p>There was pretty obvious clusters in the data. Below are  the clusters observed (notebook <a href=\"https://www.kaggle.com/code/lucasmorin/amex-umap-hdbscan-data-exploration\" target=\"_blank\">here</a>), colored by average default rates.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4037336%2Fb08be8204ef0c45829ca4ec721ae1e56%2Fclusters.png?generation=1661427754786958&amp;alt=media\" alt=\"\"></p>\n<p>I tought it would be very helpfull to the model. It accelerate learning a little bit but it didn't improve performance. After investigation, it turns out that: 1) clusters are very easy to learn from the data alone (1-2 levels trees are enough to learn each individual cluster) and what matter from error analysis is variance inside clusters.</p>\n<p><strong>Feature engineering: Transform functions</strong></p>\n<p>From the public notebooks it was clear that a lot of customer wise features were being shared and used. But with only 13 statments you can't really do an infinite amount of feature engineering without introducing correlation. I tought some edge would come from building features across customers. I've shared a baseline for this kind of feature engineering (<a href=\"https://www.kaggle.com/code/lucasmorin/amex-feature-engineering-3-transform-functions\" target=\"_blank\">here</a>).</p>\n<p>It seemed to bring performance on training set but not on validation. I figured that it introduce leakage… I still think such feature engineering could work, but need to be done fold-wise to avoid such leakage. For me it was too late to redesign my pipeline… next time :-)</p>\n<p><strong>Finger-printing clients with S_2</strong></p>\n<p>There was some notebook shared on how to use S_2. My idea was to use S_2 for \"finger-printing\" clients. That is looking at the pattern of payments (concatenating all dates) as a single categorical features. It turns out the default rate vary wildly depending on if a lot of customer share the same pattern or not (default rate by number of client with same pattern binned in 15 quantiles):</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4037336%2F5cbc2bc8a45464fec7fee176ed8a70cc%2Ffinger-printing.png?generation=1661430408110353&amp;alt=media\" alt=\"\"></p>\n<p>Unfortunately it didn't bring performance. I suspect an unique pattern would be linked to a difficulty that would appears otherwise in the dataset. Still, I had a lot of fun exploring this and it was a good idea for…</p>\n<p><strong>Studying Public / Private Overlap</strong></p>\n<p>Both test set were overlapping, I tried to deanonymise the overlap (finding clients in both sets) with categorical features, but they weren't enough. Using S_2 payments patterns helped me deanonymise the overlap. In the end I found only 300 customers that seemed to overlap… either AMEX limited overlap during their data selection process or they introduced some swap noise in the categorical features. In the end it didn't seems possible to exploit this overlap.</p>\n<p><strong>Modelling</strong></p>\n<p>I mainly stuck to my lgbm baseline… and tried some of what people shared regarding gbdt. I am very happy to have learned about dart / some parameters I usually don't use and improve that simple ensemble to make it relatively competitive. </p>\n<p>Some other things that can be mentionned as there was some discussion about feature selection / explainability:</p>\n<ul>\n<li><p>L1 regularisation initially felt like the natural way to do feature selection… it turns out removing more than 50 features this way immediately start to deteriorate perfomance. As I was already out of my comfort zone with models with 1000+ features I didn't push in that direction.</p></li>\n<li><p>Shapley values: lgbm now includes Shapley approximation with the pred contrib option. It is extremely fast. As a paper just got out about comparing shapley / shuffling approaches (<a href=\"https://www.bde.es/f/webbde/SES/Secciones/Publicaciones/PublicacionesSeriadas/DocumentosTrabajo/22/Files/dt2222e.pdf\" target=\"_blank\">ACCURACY OF EXPLANATIONS OF MACHINE LEARNING MODELS FOR CREDIT DECISIONS</a>) I tried to do feature selection this way. That is looking at contributions on a validation set and trying to get which feature influenced the prediction the wrong way. Results were inconsistent with other approaches (L1 reg / shuffling approaches) and not significantly different than doing nothing.</p></li>\n<li><p>Focal loss: I too tried this without much sucess. Nothing much to add to what has already been said except to mention this blog post of a rare quality: <a href=\"https://maxhalford.github.io/blog/lightgbm-focal-loss/\" target=\"_blank\">Focal Loss Implementation for lgbm</a></p></li>\n<li><p>NN ranking: I tried to implement a NN ranking framework (shared architecture + triplet loss). I tought playing the way to sample anchors would be a good way to help the model discriminate around the 4% threshold … I learned a lot in the process but couldn't really get out of the \"nan nan nan\" loss problem.</p></li>\n</ul>\n<p><strong>Post processing with external data</strong></p>\n<p>I suspected the LB shake-up would come from the covid impacting the private targets. As covid would not be contained in any of the data set available, relying on external data about the target seemed a good way to get some edge… I spent most of the last month of the competion trying to get external data about the target / trying to deanonymise some categorical features / shift predictions accordingly. </p>\n<p>It was very interesting to try to get external info on AMEX default rate. I mostly found info from <a href=\"https://www.americanexpress.com/en-us/business/trends-and-insights/keywords/public-relations/\" target=\"_blank\">AMEX public relations</a>. </p>\n<p>Some exemple: from the financial table you can see that there should be a drift between us and non-us default rate between train  / public / private. Then I tried to apply a similar shift in predictions based on categorical features with the same repartition as the us/non-us repartition (70%/30% with minimal difference in default rate on the training set - best candidate was D_126.). These manual shifts costed me 1200 ranks over a base ranking that wasn't very good. No hard feelings as I was already far out of the medal zone.</p>\n<hr>\n<p>All in all I really enjoyed this competition, trying new stuff even if it didn't always work. Thanks to Kaggle, AMEX and all participants.</p>",
  "messages": [
    {
      "id": "1913730",
      "postDate": "08/25/2022 13:24:15",
      "content": "<p>First I wanted to say thanks to everyone that shared insightfull stuff. Kaggle is a very good place to learn from others. </p>\n<p>As I spent quite some time trying different stuff I tought I could share some of them that seemed interesting. As you can see from my final ranking, none of them really worked. Maybe you'll find this interesting ayway, maybe you can learn from my mistakes or share how you made one of the appraoch work. </p>\n<p>As i am generally not interested in the crazy deep ensemble stuff i am not surprised to be nowhere near the medal zone. I am still pretty happy with my results (3 notebook golds / lot of learning). By the way, thanks to every one that took the time to upvote my work.</p>\n<p>In relatively chronological order:</p>\n<p><strong>Clustering</strong></p>\n<p>There was pretty obvious clusters in the data. Below are  the clusters observed (notebook <a href=\"https://www.kaggle.com/code/lucasmorin/amex-umap-hdbscan-data-exploration\" target=\"_blank\">here</a>), colored by average default rates.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4037336%2Fb08be8204ef0c45829ca4ec721ae1e56%2Fclusters.png?generation=1661427754786958&amp;alt=media\" alt=\"\"></p>\n<p>I tought it would be very helpfull to the model. It accelerate learning a little bit but it didn't improve performance. After investigation, it turns out that: 1) clusters are very easy to learn from the data alone (1-2 levels trees are enough to learn each individual cluster) and what matter from error analysis is variance inside clusters.</p>\n<p><strong>Feature engineering: Transform functions</strong></p>\n<p>From the public notebooks it was clear that a lot of customer wise features were being shared and used. But with only 13 statments you can't really do an infinite amount of feature engineering without introducing correlation. I tought some edge would come from building features across customers. I've shared a baseline for this kind of feature engineering (<a href=\"https://www.kaggle.com/code/lucasmorin/amex-feature-engineering-3-transform-functions\" target=\"_blank\">here</a>).</p>\n<p>It seemed to bring performance on training set but not on validation. I figured that it introduce leakage… I still think such feature engineering could work, but need to be done fold-wise to avoid such leakage. For me it was too late to redesign my pipeline… next time :-)</p>\n<p><strong>Finger-printing clients with S_2</strong></p>\n<p>There was some notebook shared on how to use S_2. My idea was to use S_2 for \"finger-printing\" clients. That is looking at the pattern of payments (concatenating all dates) as a single categorical features. It turns out the default rate vary wildly depending on if a lot of customer share the same pattern or not (default rate by number of client with same pattern binned in 15 quantiles):</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4037336%2F5cbc2bc8a45464fec7fee176ed8a70cc%2Ffinger-printing.png?generation=1661430408110353&amp;alt=media\" alt=\"\"></p>\n<p>Unfortunately it didn't bring performance. I suspect an unique pattern would be linked to a difficulty that would appears otherwise in the dataset. Still, I had a lot of fun exploring this and it was a good idea for…</p>\n<p><strong>Studying Public / Private Overlap</strong></p>\n<p>Both test set were overlapping, I tried to deanonymise the overlap (finding clients in both sets) with categorical features, but they weren't enough. Using S_2 payments patterns helped me deanonymise the overlap. In the end I found only 300 customers that seemed to overlap… either AMEX limited overlap during their data selection process or they introduced some swap noise in the categorical features. In the end it didn't seems possible to exploit this overlap.</p>\n<p><strong>Modelling</strong></p>\n<p>I mainly stuck to my lgbm baseline… and tried some of what people shared regarding gbdt. I am very happy to have learned about dart / some parameters I usually don't use and improve that simple ensemble to make it relatively competitive. </p>\n<p>Some other things that can be mentionned as there was some discussion about feature selection / explainability:</p>\n<ul>\n<li><p>L1 regularisation initially felt like the natural way to do feature selection… it turns out removing more than 50 features this way immediately start to deteriorate perfomance. As I was already out of my comfort zone with models with 1000+ features I didn't push in that direction.</p></li>\n<li><p>Shapley values: lgbm now includes Shapley approximation with the pred contrib option. It is extremely fast. As a paper just got out about comparing shapley / shuffling approaches (<a href=\"https://www.bde.es/f/webbde/SES/Secciones/Publicaciones/PublicacionesSeriadas/DocumentosTrabajo/22/Files/dt2222e.pdf\" target=\"_blank\">ACCURACY OF EXPLANATIONS OF MACHINE LEARNING MODELS FOR CREDIT DECISIONS</a>) I tried to do feature selection this way. That is looking at contributions on a validation set and trying to get which feature influenced the prediction the wrong way. Results were inconsistent with other approaches (L1 reg / shuffling approaches) and not significantly different than doing nothing.</p></li>\n<li><p>Focal loss: I too tried this without much sucess. Nothing much to add to what has already been said except to mention this blog post of a rare quality: <a href=\"https://maxhalford.github.io/blog/lightgbm-focal-loss/\" target=\"_blank\">Focal Loss Implementation for lgbm</a></p></li>\n<li><p>NN ranking: I tried to implement a NN ranking framework (shared architecture + triplet loss). I tought playing the way to sample anchors would be a good way to help the model discriminate around the 4% threshold … I learned a lot in the process but couldn't really get out of the \"nan nan nan\" loss problem.</p></li>\n</ul>\n<p><strong>Post processing with external data</strong></p>\n<p>I suspected the LB shake-up would come from the covid impacting the private targets. As covid would not be contained in any of the data set available, relying on external data about the target seemed a good way to get some edge… I spent most of the last month of the competion trying to get external data about the target / trying to deanonymise some categorical features / shift predictions accordingly. </p>\n<p>It was very interesting to try to get external info on AMEX default rate. I mostly found info from <a href=\"https://www.americanexpress.com/en-us/business/trends-and-insights/keywords/public-relations/\" target=\"_blank\">AMEX public relations</a>. </p>\n<p>Some exemple: from the financial table you can see that there should be a drift between us and non-us default rate between train  / public / private. Then I tried to apply a similar shift in predictions based on categorical features with the same repartition as the us/non-us repartition (70%/30% with minimal difference in default rate on the training set - best candidate was D_126.). These manual shifts costed me 1200 ranks over a base ranking that wasn't very good. No hard feelings as I was already far out of the medal zone.</p>\n<hr>\n<p>All in all I really enjoyed this competition, trying new stuff even if it didn't always work. Thanks to Kaggle, AMEX and all participants.</p>",
      "rawMarkdown": "First I wanted to say thanks to everyone that shared insightfull stuff. Kaggle is a very good place to learn from others. \n\nAs I spent quite some time trying different stuff I tought I could share some of them that seemed interesting. As you can see from my final ranking, none of them really worked. Maybe you'll find this interesting ayway, maybe you can learn from my mistakes or share how you made one of the appraoch work. \n\nAs i am generally not interested in the crazy deep ensemble stuff i am not surprised to be nowhere near the medal zone. I am still pretty happy with my results (3 notebook golds / lot of learning). By the way, thanks to every one that took the time to upvote my work.\n\nIn relatively chronological order:\n\n**Clustering**\n\nThere was pretty obvious clusters in the data. Below are  the clusters observed (notebook [here](https://www.kaggle.com/code/lucasmorin/amex-umap-hdbscan-data-exploration)), colored by average default rates.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4037336%2Fb08be8204ef0c45829ca4ec721ae1e56%2Fclusters.png?generation=1661427754786958&alt=media)\n\nI tought it would be very helpfull to the model. It accelerate learning a little bit but it didn't improve performance. After investigation, it turns out that: 1) clusters are very easy to learn from the data alone (1-2 levels trees are enough to learn each individual cluster) and what matter from error analysis is variance inside clusters.\n\n**Feature engineering: Transform functions**\n\nFrom the public notebooks it was clear that a lot of customer wise features were being shared and used. But with only 13 statments you can't really do an infinite amount of feature engineering without introducing correlation. I tought some edge would come from building features across customers. I've shared a baseline for this kind of feature engineering ([here](https://www.kaggle.com/code/lucasmorin/amex-feature-engineering-3-transform-functions)).\n\nIt seemed to bring performance on training set but not on validation. I figured that it introduce leakage... I still think such feature engineering could work, but need to be done fold-wise to avoid such leakage. For me it was too late to redesign my pipeline... next time :-)\n\n**Finger-printing clients with S_2**\n\nThere was some notebook shared on how to use S_2. My idea was to use S_2 for \"finger-printing\" clients. That is looking at the pattern of payments (concatenating all dates) as a single categorical features. It turns out the default rate vary wildly depending on if a lot of customer share the same pattern or not (default rate by number of client with same pattern binned in 15 quantiles):\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4037336%2F5cbc2bc8a45464fec7fee176ed8a70cc%2Ffinger-printing.png?generation=1661430408110353&alt=media)\n\nUnfortunately it didn't bring performance. I suspect an unique pattern would be linked to a difficulty that would appears otherwise in the dataset. Still, I had a lot of fun exploring this and it was a good idea for...\n\n**Studying Public / Private Overlap**\n\nBoth test set were overlapping, I tried to deanonymise the overlap (finding clients in both sets) with categorical features, but they weren't enough. Using S_2 payments patterns helped me deanonymise the overlap. In the end I found only 300 customers that seemed to overlap... either AMEX limited overlap during their data selection process or they introduced some swap noise in the categorical features. In the end it didn't seems possible to exploit this overlap.\n\n**Modelling**\n\nI mainly stuck to my lgbm baseline... and tried some of what people shared regarding gbdt. I am very happy to have learned about dart / some parameters I usually don't use and improve that simple ensemble to make it relatively competitive. \n\nSome other things that can be mentionned as there was some discussion about feature selection / explainability:\n\n- L1 regularisation initially felt like the natural way to do feature selection... it turns out removing more than 50 features this way immediately start to deteriorate perfomance. As I was already out of my comfort zone with models with 1000+ features I didn't push in that direction.\n\n- Shapley values: lgbm now includes Shapley approximation with the pred contrib option. It is extremely fast. As a paper just got out about comparing shapley / shuffling approaches ([ACCURACY OF EXPLANATIONS OF MACHINE LEARNING MODELS FOR CREDIT DECISIONS](https://www.bde.es/f/webbde/SES/Secciones/Publicaciones/PublicacionesSeriadas/DocumentosTrabajo/22/Files/dt2222e.pdf)) I tried to do feature selection this way. That is looking at contributions on a validation set and trying to get which feature influenced the prediction the wrong way. Results were inconsistent with other approaches (L1 reg / shuffling approaches) and not significantly different than doing nothing.\n\n- Focal loss: I too tried this without much sucess. Nothing much to add to what has already been said except to mention this blog post of a rare quality: [Focal Loss Implementation for lgbm](https://maxhalford.github.io/blog/lightgbm-focal-loss/)\n\n- NN ranking: I tried to implement a NN ranking framework (shared architecture + triplet loss). I tought playing the way to sample anchors would be a good way to help the model discriminate around the 4% threshold ... I learned a lot in the process but couldn't really get out of the \"nan nan nan\" loss problem.\n\n**Post processing with external data**\n\nI suspected the LB shake-up would come from the covid impacting the private targets. As covid would not be contained in any of the data set available, relying on external data about the target seemed a good way to get some edge... I spent most of the last month of the competion trying to get external data about the target / trying to deanonymise some categorical features / shift predictions accordingly. \n\nIt was very interesting to try to get external info on AMEX default rate. I mostly found info from [AMEX public relations](https://www.americanexpress.com/en-us/business/trends-and-insights/keywords/public-relations/). \n\nSome exemple: from the financial table you can see that there should be a drift between us and non-us default rate between train  / public / private. Then I tried to apply a similar shift in predictions based on categorical features with the same repartition as the us/non-us repartition (70%/30% with minimal difference in default rate on the training set - best candidate was D_126.). These manual shifts costed me 1200 ranks over a base ranking that wasn't very good. No hard feelings as I was already far out of the medal zone.\n\n----\n\nAll in all I really enjoyed this competition, trying new stuff even if it didn't always work. Thanks to Kaggle, AMEX and all participants.",
      "votes": null
    },
    {
      "id": "1913743",
      "postDate": "08/25/2022 13:34:08",
      "content": "<p>This was indeed a great challenge! Thanks for the competition and hearty congratulations to the winners!</p>",
      "rawMarkdown": "This was indeed a great challenge! Thanks for the competition and hearty congratulations to the winners!",
      "votes": null
    },
    {
      "id": "1913758",
      "postDate": "08/25/2022 13:43:48",
      "content": "<ol>\n<li>Trained a CNN with focal loss + 'Mish' as an activation function during the last stages of the competition, the network itself scores 0.8034 on the private leaderboard, but when ensembled with lgbm/xgboost there was very minimal boost to the overall cv/lb score.</li>\n<li>Tried my hands on Tabnet and I just couldn't get my oof cv beyond 0.790, lb 0.791, so I gave up on it. Later read raddar's post on how multiple of them combined gave a significant boost to his cv score, and I realized I probably should have stuck with it.</li>\n</ol>\n<p>There are a million other things that didn't work but these two came to mind the quickest. </p>",
      "rawMarkdown": "1. Trained a CNN with focal loss + 'Mish' as an activation function during the last stages of the competition, the network itself scores 0.8034 on the private leaderboard, but when ensembled with lgbm/xgboost there was very minimal boost to the overall cv/lb score.\n2. Tried my hands on Tabnet and I just couldn't get my oof cv beyond 0.790, lb 0.791, so I gave up on it. Later read raddar's post on how multiple of them combined gave a significant boost to his cv score, and I realized I probably should have stuck with it.\n\nThere are a million other things that didn't work but these two came to mind the quickest.",
      "votes": null
    },
    {
      "id": "1914949",
      "postDate": "08/26/2022 14:59:32",
      "content": "<p>Valuable sharing of what didn't work giving that many share the mostly the best results (in any career too).</p>",
      "rawMarkdown": "Valuable sharing of what didn't work giving that many share the mostly the best results (in any career too).",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1913743,
      "author_name": "ravi20076",
      "author_url": "",
      "post_date": "08/25/2022 13:34:08",
      "content": "<p>This was indeed a great challenge! Thanks for the competition and hearty congratulations to the winners!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1913758,
      "author_name": "ol0fmeister",
      "author_url": "",
      "post_date": "08/25/2022 13:43:48",
      "content": "<ol>\n<li>Trained a CNN with focal loss + 'Mish' as an activation function during the last stages of the competition, the network itself scores 0.8034 on the private leaderboard, but when ensembled with lgbm/xgboost there was very minimal boost to the overall cv/lb score.</li>\n<li>Tried my hands on Tabnet and I just couldn't get my oof cv beyond 0.790, lb 0.791, so I gave up on it. Later read raddar's post on how multiple of them combined gave a significant boost to his cv score, and I realized I probably should have stuck with it.</li>\n</ol>\n<p>There are a million other things that didn't work but these two came to mind the quickest. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1914949,
      "author_name": "mpwolke",
      "author_url": "",
      "post_date": "08/26/2022 14:59:32",
      "content": "<p>Valuable sharing of what didn't work giving that many share the mostly the best results (in any career too).</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1913730": "First I wanted to say thanks to everyone that shared insightfull stuff. Kaggle is a very good place to learn from others. \n\nAs I spent quite some time trying different stuff I tought I could share some of them that seemed interesting. As you can see from my final ranking, none of them really worked. Maybe you'll find this interesting ayway, maybe you can learn from my mistakes or share how you made one of the appraoch work. \n\nAs i am generally not interested in the crazy deep ensemble stuff i am not surprised to be nowhere near the medal zone. I am still pretty happy with my results (3 notebook golds / lot of learning). By the way, thanks to every one that took the time to upvote my work.\n\nIn relatively chronological order:\n\n**Clustering**\n\nThere was pretty obvious clusters in the data. Below are  the clusters observed (notebook [here](https://www.kaggle.com/code/lucasmorin/amex-umap-hdbscan-data-exploration)), colored by average default rates.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4037336%2Fb08be8204ef0c45829ca4ec721ae1e56%2Fclusters.png?generation=1661427754786958&alt=media)\n\nI tought it would be very helpfull to the model. It accelerate learning a little bit but it didn't improve performance. After investigation, it turns out that: 1) clusters are very easy to learn from the data alone (1-2 levels trees are enough to learn each individual cluster) and what matter from error analysis is variance inside clusters.\n\n**Feature engineering: Transform functions**\n\nFrom the public notebooks it was clear that a lot of customer wise features were being shared and used. But with only 13 statments you can't really do an infinite amount of feature engineering without introducing correlation. I tought some edge would come from building features across customers. I've shared a baseline for this kind of feature engineering ([here](https://www.kaggle.com/code/lucasmorin/amex-feature-engineering-3-transform-functions)).\n\nIt seemed to bring performance on training set but not on validation. I figured that it introduce leakage... I still think such feature engineering could work, but need to be done fold-wise to avoid such leakage. For me it was too late to redesign my pipeline... next time :-)\n\n**Finger-printing clients with S_2**\n\nThere was some notebook shared on how to use S_2. My idea was to use S_2 for \"finger-printing\" clients. That is looking at the pattern of payments (concatenating all dates) as a single categorical features. It turns out the default rate vary wildly depending on if a lot of customer share the same pattern or not (default rate by number of client with same pattern binned in 15 quantiles):\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4037336%2F5cbc2bc8a45464fec7fee176ed8a70cc%2Ffinger-printing.png?generation=1661430408110353&alt=media)\n\nUnfortunately it didn't bring performance. I suspect an unique pattern would be linked to a difficulty that would appears otherwise in the dataset. Still, I had a lot of fun exploring this and it was a good idea for...\n\n**Studying Public / Private Overlap**\n\nBoth test set were overlapping, I tried to deanonymise the overlap (finding clients in both sets) with categorical features, but they weren't enough. Using S_2 payments patterns helped me deanonymise the overlap. In the end I found only 300 customers that seemed to overlap... either AMEX limited overlap during their data selection process or they introduced some swap noise in the categorical features. In the end it didn't seems possible to exploit this overlap.\n\n**Modelling**\n\nI mainly stuck to my lgbm baseline... and tried some of what people shared regarding gbdt. I am very happy to have learned about dart / some parameters I usually don't use and improve that simple ensemble to make it relatively competitive. \n\nSome other things that can be mentionned as there was some discussion about feature selection / explainability:\n\n- L1 regularisation initially felt like the natural way to do feature selection... it turns out removing more than 50 features this way immediately start to deteriorate perfomance. As I was already out of my comfort zone with models with 1000+ features I didn't push in that direction.\n\n- Shapley values: lgbm now includes Shapley approximation with the pred contrib option. It is extremely fast. As a paper just got out about comparing shapley / shuffling approaches ([ACCURACY OF EXPLANATIONS OF MACHINE LEARNING MODELS FOR CREDIT DECISIONS](https://www.bde.es/f/webbde/SES/Secciones/Publicaciones/PublicacionesSeriadas/DocumentosTrabajo/22/Files/dt2222e.pdf)) I tried to do feature selection this way. That is looking at contributions on a validation set and trying to get which feature influenced the prediction the wrong way. Results were inconsistent with other approaches (L1 reg / shuffling approaches) and not significantly different than doing nothing.\n\n- Focal loss: I too tried this without much sucess. Nothing much to add to what has already been said except to mention this blog post of a rare quality: [Focal Loss Implementation for lgbm](https://maxhalford.github.io/blog/lightgbm-focal-loss/)\n\n- NN ranking: I tried to implement a NN ranking framework (shared architecture + triplet loss). I tought playing the way to sample anchors would be a good way to help the model discriminate around the 4% threshold ... I learned a lot in the process but couldn't really get out of the \"nan nan nan\" loss problem.\n\n**Post processing with external data**\n\nI suspected the LB shake-up would come from the covid impacting the private targets. As covid would not be contained in any of the data set available, relying on external data about the target seemed a good way to get some edge... I spent most of the last month of the competion trying to get external data about the target / trying to deanonymise some categorical features / shift predictions accordingly. \n\nIt was very interesting to try to get external info on AMEX default rate. I mostly found info from [AMEX public relations](https://www.americanexpress.com/en-us/business/trends-and-insights/keywords/public-relations/). \n\nSome exemple: from the financial table you can see that there should be a drift between us and non-us default rate between train  / public / private. Then I tried to apply a similar shift in predictions based on categorical features with the same repartition as the us/non-us repartition (70%/30% with minimal difference in default rate on the training set - best candidate was D_126.). These manual shifts costed me 1200 ranks over a base ranking that wasn't very good. No hard feelings as I was already far out of the medal zone.\n\n----\n\nAll in all I really enjoyed this competition, trying new stuff even if it didn't always work. Thanks to Kaggle, AMEX and all participants.",
    "1913743": "This was indeed a great challenge! Thanks for the competition and hearty congratulations to the winners!",
    "1913758": "1. Trained a CNN with focal loss + 'Mish' as an activation function during the last stages of the competition, the network itself scores 0.8034 on the private leaderboard, but when ensembled with lgbm/xgboost there was very minimal boost to the overall cv/lb score.\n2. Tried my hands on Tabnet and I just couldn't get my oof cv beyond 0.790, lb 0.791, so I gave up on it. Later read raddar's post on how multiple of them combined gave a significant boost to his cv score, and I realized I probably should have stuck with it.\n\nThere are a million other things that didn't work but these two came to mind the quickest.",
    "1914949": "Valuable sharing of what didn't work giving that many share the mostly the best results (in any career too)."
  },
  "source": "meta"
}