{
  "id": 347685,
  "title": "Bronze Medal Solution",
  "url": "/competitions/amex-default-prediction/discussion/347685",
  "author_name": "",
  "post_date": "2022-08-25T05:35:13.556424600Z",
  "votes": 33,
  "comment_count": 2,
  "views": 0,
  "content": "<p>Hello everyone,<br>\nReally thanks Amex and Kaggle for hosting such an amazing competition! And Thank you Kagglers for sharing many helpful discussions and notebooks. especially <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> , <a href=\"https://www.kaggle.com/raddar\" target=\"_blank\">@raddar</a> and <a href=\"https://www.kaggle.com/ragnar123\" target=\"_blank\">@ragnar123</a>. This competition would be impossible without you guys, really thanks!</p>\n<p>I am really happy to get a bronze medal in my first Kaggle competition! My Solution was just an average between my model (single catboost) and two public notebooks. <br>\nThe mean score of my CV was 0.795. One of the folds reached 0.799 score. It scored 0.795 In the public LB. Averaging it with public notebooks boosted it to reach 0.799, and 0.807 in private LB.</p>\n<p>My features were not special. Most of them were from the published kernels. However, I added some ideas as well.<br>\nI thought that in addition to having the aggregations of each feature alone, I need features that can capture the relations between multiple features in the same time. Here what I added:</p>\n<p>1- The Polar Coordinates between features: here I focused especially in P_2 and added the polar coordinates between it and all the other features. They appeared to be powerful in predicting the target as that the highest 20 features that catboost used were all Polars features and all of them were used equally (they roughly have the same importance value in catboost).</p>\n<p>2- Features Groups: as that the features were anonymized, I used two approaches to utilize them:<br>\nFirstly: I added statistics and interactions for features that had the same mean value.<br>\nSecondly: I added statistics and interactions for features that had really high correlation.</p>\n<p>In addition, I added target encoding, PCA, binning of some features and some aggregations based on lagged data.</p>\n<p>In terms of features selection, I used the null importance approach. It reduced the number of features from almost 2200 to only 1200.</p>\n<p>Lastly, I used a Catboost model with eta 0.01, depth 4 and about 40000 iterations.</p>\n<p>You can check the code here:<br>\n<a href=\"https://www.kaggle.com/code/mohammad2012191/catboost-unique-feature-engineering-ideas\" target=\"_blank\">https://www.kaggle.com/code/mohammad2012191/catboost-unique-feature-engineering-ideas</a></p>",
  "messages": [
    {
      "id": "1913058",
      "postDate": "08/25/2022 05:35:13",
      "content": "<p>Hello everyone,<br>\nReally thanks Amex and Kaggle for hosting such an amazing competition! And Thank you Kagglers for sharing many helpful discussions and notebooks. especially <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> , <a href=\"https://www.kaggle.com/raddar\" target=\"_blank\">@raddar</a> and <a href=\"https://www.kaggle.com/ragnar123\" target=\"_blank\">@ragnar123</a>. This competition would be impossible without you guys, really thanks!</p>\n<p>I am really happy to get a bronze medal in my first Kaggle competition! My Solution was just an average between my model (single catboost) and two public notebooks. <br>\nThe mean score of my CV was 0.795. One of the folds reached 0.799 score. It scored 0.795 In the public LB. Averaging it with public notebooks boosted it to reach 0.799, and 0.807 in private LB.</p>\n<p>My features were not special. Most of them were from the published kernels. However, I added some ideas as well.<br>\nI thought that in addition to having the aggregations of each feature alone, I need features that can capture the relations between multiple features in the same time. Here what I added:</p>\n<p>1- The Polar Coordinates between features: here I focused especially in P_2 and added the polar coordinates between it and all the other features. They appeared to be powerful in predicting the target as that the highest 20 features that catboost used were all Polars features and all of them were used equally (they roughly have the same importance value in catboost).</p>\n<p>2- Features Groups: as that the features were anonymized, I used two approaches to utilize them:<br>\nFirstly: I added statistics and interactions for features that had the same mean value.<br>\nSecondly: I added statistics and interactions for features that had really high correlation.</p>\n<p>In addition, I added target encoding, PCA, binning of some features and some aggregations based on lagged data.</p>\n<p>In terms of features selection, I used the null importance approach. It reduced the number of features from almost 2200 to only 1200.</p>\n<p>Lastly, I used a Catboost model with eta 0.01, depth 4 and about 40000 iterations.</p>\n<p>You can check the code here:<br>\n<a href=\"https://www.kaggle.com/code/mohammad2012191/catboost-unique-feature-engineering-ideas\" target=\"_blank\">https://www.kaggle.com/code/mohammad2012191/catboost-unique-feature-engineering-ideas</a></p>",
      "rawMarkdown": "Hello everyone,\nReally thanks Amex and Kaggle for hosting such an amazing competition! And Thank you Kagglers for sharing many helpful discussions and notebooks. especially @cdeotte , @raddar and @ragnar123. This competition would be impossible without you guys, really thanks!\n\nI am really happy to get a bronze medal in my first Kaggle competition! My Solution was just an average between my model (single catboost) and two public notebooks. \nThe mean score of my CV was 0.795. One of the folds reached 0.799 score. It scored 0.795 In the public LB. Averaging it with public notebooks boosted it to reach 0.799, and 0.807 in private LB.\n\nMy features were not special. Most of them were from the published kernels. However, I added some ideas as well.\nI thought that in addition to having the aggregations of each feature alone, I need features that can capture the relations between multiple features in the same time. Here what I added:\n\n1- The Polar Coordinates between features: here I focused especially in P_2 and added the polar coordinates between it and all the other features. They appeared to be powerful in predicting the target as that the highest 20 features that catboost used were all Polars features and all of them were used equally (they roughly have the same importance value in catboost).\n\n2- Features Groups: as that the features were anonymized, I used two approaches to utilize them:\nFirstly: I added statistics and interactions for features that had the same mean value.\nSecondly: I added statistics and interactions for features that had really high correlation.\n\nIn addition, I added target encoding, PCA, binning of some features and some aggregations based on lagged data.\n\nIn terms of features selection, I used the null importance approach. It reduced the number of features from almost 2200 to only 1200.\n\nLastly, I used a Catboost model with eta 0.01, depth 4 and about 40000 iterations.\n\nYou can check the code here:\nhttps://www.kaggle.com/code/mohammad2012191/catboost-unique-feature-engineering-ideas",
      "votes": null
    },
    {
      "id": "1913059",
      "postDate": "08/25/2022 05:35:51",
      "content": "<p>Hearty congratulations! This is remarkable! Wishing you success in the future as well!</p>",
      "rawMarkdown": "Hearty congratulations! This is remarkable! Wishing you success in the future as well!",
      "votes": null
    },
    {
      "id": "1913506",
      "postDate": "08/25/2022 10:46:05",
      "content": "<p>Congratulations! This was my first competition as well. I learned a lot from it and enjoyed kaggling! <br>\nThank you for sharing some ideas here! </p>",
      "rawMarkdown": "Congratulations! This was my first competition as well. I learned a lot from it and enjoyed kaggling! \nThank you for sharing some ideas here!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1913059,
      "author_name": "ravi20076",
      "author_url": "",
      "post_date": "08/25/2022 05:35:51",
      "content": "<p>Hearty congratulations! This is remarkable! Wishing you success in the future as well!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1913506,
      "author_name": "dlakshma",
      "author_url": "",
      "post_date": "08/25/2022 10:46:05",
      "content": "<p>Congratulations! This was my first competition as well. I learned a lot from it and enjoyed kaggling! <br>\nThank you for sharing some ideas here! </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1913058": "Hello everyone,\nReally thanks Amex and Kaggle for hosting such an amazing competition! And Thank you Kagglers for sharing many helpful discussions and notebooks. especially @cdeotte , @raddar and @ragnar123. This competition would be impossible without you guys, really thanks!\n\nI am really happy to get a bronze medal in my first Kaggle competition! My Solution was just an average between my model (single catboost) and two public notebooks. \nThe mean score of my CV was 0.795. One of the folds reached 0.799 score. It scored 0.795 In the public LB. Averaging it with public notebooks boosted it to reach 0.799, and 0.807 in private LB.\n\nMy features were not special. Most of them were from the published kernels. However, I added some ideas as well.\nI thought that in addition to having the aggregations of each feature alone, I need features that can capture the relations between multiple features in the same time. Here what I added:\n\n1- The Polar Coordinates between features: here I focused especially in P_2 and added the polar coordinates between it and all the other features. They appeared to be powerful in predicting the target as that the highest 20 features that catboost used were all Polars features and all of them were used equally (they roughly have the same importance value in catboost).\n\n2- Features Groups: as that the features were anonymized, I used two approaches to utilize them:\nFirstly: I added statistics and interactions for features that had the same mean value.\nSecondly: I added statistics and interactions for features that had really high correlation.\n\nIn addition, I added target encoding, PCA, binning of some features and some aggregations based on lagged data.\n\nIn terms of features selection, I used the null importance approach. It reduced the number of features from almost 2200 to only 1200.\n\nLastly, I used a Catboost model with eta 0.01, depth 4 and about 40000 iterations.\n\nYou can check the code here:\nhttps://www.kaggle.com/code/mohammad2012191/catboost-unique-feature-engineering-ideas",
    "1913059": "Hearty congratulations! This is remarkable! Wishing you success in the future as well!",
    "1913506": "Congratulations! This was my first competition as well. I learned a lot from it and enjoyed kaggling! \nThank you for sharing some ideas here!"
  },
  "source": "meta"
}