{
  "id": 335944,
  "title": "Permutation importance XGB",
  "url": "/competitions/amex-default-prediction/discussion/335944",
  "author_name": "Noir.s",
  "post_date": "2022-07-08T16:07:49.775000",
  "votes": 18,
  "comment_count": 7,
  "views": 0,
  "content": "<p>I found the discussion post on feature importance by <a href=\"https://www.kaggle.com/ambrosm\" target=\"_blank\">@ambrosm</a> very interesting. So I wanted to try to implement permutation feature importance and see the results</p>\n<p>I have explored this in my notebook and it looks like there is potential to improve on the baseline<br>\n<a href=\"https://www.kaggle.com/code/illidan7/amex-permutation-feature-importance\" target=\"_blank\">https://www.kaggle.com/code/illidan7/amex-permutation-feature-importance</a></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F471225%2F2ad61716f63933662f3e1a644ed06bf3%2FPermImp.png?generation=1657550240342914&amp;alt=media\" alt=\"\"></p>\n<p>If we consider the XGB split feature importance from <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> 's XGB Starter notebook, we see that feature 'P_3_last' is the 15th most important feature</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F471225%2Faa205a644a05e6bf8dbeb0136d908fed%2FXGBImp.png?generation=1657550262055064&amp;alt=media\" alt=\"\"></p>\n<p>However, based on the results from my notebook</p>\n<p>137       P_3_last  0.791860</p>\n<p>367       BASELINE  0.791720</p>\n<p>We can see that the score can potentially improve by not using this variable</p>\n<p>Very similar exploration to that done by <a href=\"https://www.kaggle.com/ambrosm\" target=\"_blank\">@ambrosm</a>. Just sharing my own observations here</p>\n<p>Please check out the following links:</p>\n<p><a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/331131\" target=\"_blank\">https://www.kaggle.com/competitions/amex-default-prediction/discussion/331131</a><br>\n<a href=\"https://www.kaggle.com/code/cdeotte/xgboost-starter-0-793\" target=\"_blank\">https://www.kaggle.com/code/cdeotte/xgboost-starter-0-793</a><br>\n<a href=\"https://www.kaggle.com/code/cdeotte/lstm-feature-importance\" target=\"_blank\">https://www.kaggle.com/code/cdeotte/lstm-feature-importance</a><br>\n<a href=\"https://www.kaggle.com/datasets/raddar/amex-data-integer-dtypes-parquet-format\" target=\"_blank\">https://www.kaggle.com/datasets/raddar/amex-data-integer-dtypes-parquet-format</a></p>",
  "messages": [
    {
      "id": 1848429,
      "postDate": "2022-07-08T16:07:49.777Z",
      "content": "<p>I found the discussion post on feature importance by <a href=\"https://www.kaggle.com/ambrosm\" target=\"_blank\">@ambrosm</a> very interesting. So I wanted to try to implement permutation feature importance and see the results</p>\n<p>I have explored this in my notebook and it looks like there is potential to improve on the baseline<br>\n<a href=\"https://www.kaggle.com/code/illidan7/amex-permutation-feature-importance\" target=\"_blank\">https://www.kaggle.com/code/illidan7/amex-permutation-feature-importance</a></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F471225%2F2ad61716f63933662f3e1a644ed06bf3%2FPermImp.png?generation=1657550240342914&amp;alt=media\" alt=\"\"></p>\n<p>If we consider the XGB split feature importance from <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> 's XGB Starter notebook, we see that feature 'P_3_last' is the 15th most important feature</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F471225%2Faa205a644a05e6bf8dbeb0136d908fed%2FXGBImp.png?generation=1657550262055064&amp;alt=media\" alt=\"\"></p>\n<p>However, based on the results from my notebook</p>\n<p>137       P_3_last  0.791860</p>\n<p>367       BASELINE  0.791720</p>\n<p>We can see that the score can potentially improve by not using this variable</p>\n<p>Very similar exploration to that done by <a href=\"https://www.kaggle.com/ambrosm\" target=\"_blank\">@ambrosm</a>. Just sharing my own observations here</p>\n<p>Please check out the following links:</p>\n<p><a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/331131\" target=\"_blank\">https://www.kaggle.com/competitions/amex-default-prediction/discussion/331131</a><br>\n<a href=\"https://www.kaggle.com/code/cdeotte/xgboost-starter-0-793\" target=\"_blank\">https://www.kaggle.com/code/cdeotte/xgboost-starter-0-793</a><br>\n<a href=\"https://www.kaggle.com/code/cdeotte/lstm-feature-importance\" target=\"_blank\">https://www.kaggle.com/code/cdeotte/lstm-feature-importance</a><br>\n<a href=\"https://www.kaggle.com/datasets/raddar/amex-data-integer-dtypes-parquet-format\" target=\"_blank\">https://www.kaggle.com/datasets/raddar/amex-data-integer-dtypes-parquet-format</a></p>",
      "rawMarkdown": "I found the discussion post on feature importance by @ambrosm very interesting. So I wanted to try to implement permutation feature importance and see the results\n\nI have explored this in my notebook and it looks like there is potential to improve on the baseline\nhttps://www.kaggle.com/code/illidan7/amex-permutation-feature-importance\n\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F471225%2F2ad61716f63933662f3e1a644ed06bf3%2FPermImp.png?generation=1657550240342914&alt=media)\n\n\nIf we consider the XGB split feature importance from @cdeotte 's XGB Starter notebook, we see that feature 'P_3_last' is the 15th most important feature\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F471225%2Faa205a644a05e6bf8dbeb0136d908fed%2FXGBImp.png?generation=1657550262055064&alt=media)\n\n\nHowever, based on the results from my notebook\n\n137       P_3_last  0.791860\n\n367       BASELINE  0.791720\n\nWe can see that the score can potentially improve by not using this variable\n\nVery similar exploration to that done by @ambrosm. Just sharing my own observations here\n\nPlease check out the following links:\n\nhttps://www.kaggle.com/competitions/amex-default-prediction/discussion/331131\nhttps://www.kaggle.com/code/cdeotte/xgboost-starter-0-793\nhttps://www.kaggle.com/code/cdeotte/lstm-feature-importance\nhttps://www.kaggle.com/datasets/raddar/amex-data-integer-dtypes-parquet-format",
      "votes": 17
    },
    {
      "id": 1851357,
      "postDate": "2022-07-11T07:26:00.370Z",
      "content": "<p>This is great! thanks for sharing your work!</p>",
      "rawMarkdown": "This is great! thanks for sharing your work!",
      "votes": 1
    },
    {
      "id": 1850995,
      "postDate": "2022-07-11T01:05:03.453Z",
      "content": "<p>Nice work! 2 questions:</p>\n<ul>\n<li>Did you try retraining without these features? I also wonder what threshold we should use to decide to keep/reject features.<br>\nThe best boost in your example is D_52_std: +0.000716538. It is less than the variation we get just by changing the seed so it's tricky.</li>\n<li>Do you think we should apply this feature selection to each fold independently? Or get a common list to apply to each fold?</li>\n</ul>",
      "rawMarkdown": "Nice work! 2 questions:\n- Did you try retraining without these features? I also wonder what threshold we should use to decide to keep/reject features.\nThe best boost in your example is D_52_std: +0.000716538. It is less than the variation we get just by changing the seed so it's tricky.\n- Do you think we should apply this feature selection to each fold independently? Or get a common list to apply to each fold?",
      "votes": 1,
      "replies": [
        {
          "id": 1852122,
          "postDate": "2022-07-11T19:36:47.103Z",
          "content": "<p>Yeah choosing features to eliminate/improve the model is definitely proving tricky. I have not specifically tried to retrain using these results and this set of features. I am trying different feature engineering ideas at the moment</p>\n<p>I have not tried to apply it to more than one fold and compare results. But that could be useful too</p>\n<p>I found this discussion interesting where <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> has explained an approach to selecting features which impact by more than 2x standard deviation of CV score change based on different seeds</p>\n<p><a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/336145\" target=\"_blank\">https://www.kaggle.com/competitions/amex-default-prediction/discussion/336145</a></p>",
          "rawMarkdown": "Yeah choosing features to eliminate/improve the model is definitely proving tricky. I have not specifically tried to retrain using these results and this set of features. I am trying different feature engineering ideas at the moment\n\nI have not tried to apply it to more than one fold and compare results. But that could be useful too\n\nI found this discussion interesting where @cdeotte has explained an approach to selecting features which impact by more than 2x standard deviation of CV score change based on different seeds\n\nhttps://www.kaggle.com/competitions/amex-default-prediction/discussion/336145"
        }
      ]
    },
    {
      "id": 1849911,
      "postDate": "2022-07-09T23:41:34.997Z",
      "content": "<p>Thanks for sharing your observations!<br>\nThe images seems broken on your post.. </p>",
      "rawMarkdown": "Thanks for sharing your observations!\nThe images seems broken on your post.. \n",
      "votes": 1,
      "replies": [
        {
          "id": 1850841,
          "postDate": "2022-07-10T20:15:31.197Z",
          "content": "<p>Thanks for pointing it out. Looks ok on my end now!</p>",
          "rawMarkdown": "Thanks for pointing it out. Looks ok on my end now!"
        }
      ]
    },
    {
      "id": 1848439,
      "postDate": "2022-07-08T16:17:59.030Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 1848850,
      "postDate": "2022-07-09T01:34:59.587Z",
      "content": "<p>Thank you <a href=\"https://www.kaggle.com/illidan7\" target=\"_blank\">@illidan7</a> </p>",
      "rawMarkdown": "Thank you @illidan7 ",
      "votes": 1
    }
  ],
  "comments": [
    {
      "id": 1851357,
      "author_name": "1110Ra",
      "author_url": "",
      "post_date": "2022-07-11T07:26:00.370000",
      "content": "<p>This is great! thanks for sharing your work!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1850995,
      "author_name": "k3vla",
      "author_url": "",
      "post_date": "2022-07-11T01:05:03.453000",
      "content": "<p>Nice work! 2 questions:</p>\n<ul>\n<li>Did you try retraining without these features? I also wonder what threshold we should use to decide to keep/reject features.<br>\nThe best boost in your example is D_52_std: +0.000716538. It is less than the variation we get just by changing the seed so it's tricky.</li>\n<li>Do you think we should apply this feature selection to each fold independently? Or get a common list to apply to each fold?</li>\n</ul>",
      "votes": 1,
      "replies": [
        {
          "id": 1852122,
          "author_name": "Noir.s",
          "author_url": "",
          "post_date": "2022-07-11T19:36:47.103000",
          "content": "<p>Yeah choosing features to eliminate/improve the model is definitely proving tricky. I have not specifically tried to retrain using these results and this set of features. I am trying different feature engineering ideas at the moment</p>\n<p>I have not tried to apply it to more than one fold and compare results. But that could be useful too</p>\n<p>I found this discussion interesting where <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> has explained an approach to selecting features which impact by more than 2x standard deviation of CV score change based on different seeds</p>\n<p><a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/336145\" target=\"_blank\">https://www.kaggle.com/competitions/amex-default-prediction/discussion/336145</a></p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1849911,
      "author_name": "The Devastator",
      "author_url": "",
      "post_date": "2022-07-09T23:41:34.997000",
      "content": "<p>Thanks for sharing your observations!<br>\nThe images seems broken on your post.. </p>",
      "votes": 1,
      "replies": [
        {
          "id": 1850841,
          "author_name": "Noir.s",
          "author_url": "",
          "post_date": "2022-07-10T20:15:31.197000",
          "content": "<p>Thanks for pointing it out. Looks ok on my end now!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1848439,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-07-08T16:17:59.030000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1848850,
      "author_name": "Muhammed Tausif",
      "author_url": "",
      "post_date": "2022-07-09T01:34:59.587000",
      "content": "<p>Thank you <a href=\"https://www.kaggle.com/illidan7\" target=\"_blank\">@illidan7</a> </p>",
      "votes": 1,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1848429": "I found the discussion post on feature importance by @ambrosm very interesting. So I wanted to try to implement permutation feature importance and see the results\n\nI have explored this in my notebook and it looks like there is potential to improve on the baseline\nhttps://www.kaggle.com/code/illidan7/amex-permutation-feature-importance\n\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F471225%2F2ad61716f63933662f3e1a644ed06bf3%2FPermImp.png?generation=1657550240342914&alt=media)\n\n\nIf we consider the XGB split feature importance from @cdeotte 's XGB Starter notebook, we see that feature 'P_3_last' is the 15th most important feature\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F471225%2Faa205a644a05e6bf8dbeb0136d908fed%2FXGBImp.png?generation=1657550262055064&alt=media)\n\n\nHowever, based on the results from my notebook\n\n137       P_3_last  0.791860\n\n367       BASELINE  0.791720\n\nWe can see that the score can potentially improve by not using this variable\n\nVery similar exploration to that done by @ambrosm. Just sharing my own observations here\n\nPlease check out the following links:\n\nhttps://www.kaggle.com/competitions/amex-default-prediction/discussion/331131\nhttps://www.kaggle.com/code/cdeotte/xgboost-starter-0-793\nhttps://www.kaggle.com/code/cdeotte/lstm-feature-importance\nhttps://www.kaggle.com/datasets/raddar/amex-data-integer-dtypes-parquet-format",
    "1851357": "This is great! thanks for sharing your work!",
    "1850995": "Nice work! 2 questions:\n- Did you try retraining without these features? I also wonder what threshold we should use to decide to keep/reject features.\nThe best boost in your example is D_52_std: +0.000716538. It is less than the variation we get just by changing the seed so it's tricky.\n- Do you think we should apply this feature selection to each fold independently? Or get a common list to apply to each fold?",
    "1849911": "Thanks for sharing your observations!\nThe images seems broken on your post.. \n",
    "1848439": "",
    "1848850": "Thank you @illidan7 "
  }
}