{
  "id": 346775,
  "title": "Feature Selection: Random Bar Method⭐⭐⭐",
  "url": "/competitions/amex-default-prediction/discussion/346775",
  "author_name": "",
  "post_date": "2022-08-21T09:26:08.686314Z",
  "votes": 17,
  "comment_count": 6,
  "views": 0,
  "content": "<p>Hello everyone,<br>\nAs that this competition requires very good features selection, here is a very elegant and simple method for feature selection that I found once used in Kaggle. Not sure who invented it (Please mention him in the comments if you know him).</p>\n<p>The method as follow:<br>\n1- Make a feature that has random values e.g.<br>\n<code>df['Random'] = np.random.normal(size=(len(df.index),1))</code><br>\n2- Make a model using your original features in addition to this one.<br>\n3- Find the features importances.<br>\n4- We know that our feature consists of just noise. So, any feature that has an importance value less than the importance of our feature is nothing but noise!</p>\n<p>Very elegant and simple!</p>",
  "messages": [
    {
      "id": "1908012",
      "postDate": "08/21/2022 09:26:08",
      "content": "<p>Hello everyone,<br>\nAs that this competition requires very good features selection, here is a very elegant and simple method for feature selection that I found once used in Kaggle. Not sure who invented it (Please mention him in the comments if you know him).</p>\n<p>The method as follow:<br>\n1- Make a feature that has random values e.g.<br>\n<code>df['Random'] = np.random.normal(size=(len(df.index),1))</code><br>\n2- Make a model using your original features in addition to this one.<br>\n3- Find the features importances.<br>\n4- We know that our feature consists of just noise. So, any feature that has an importance value less than the importance of our feature is nothing but noise!</p>\n<p>Very elegant and simple!</p>",
      "rawMarkdown": "Hello everyone,\nAs that this competition requires very good features selection, here is a very elegant and simple method for feature selection that I found once used in Kaggle. Not sure who invented it (Please mention him in the comments if you know him).\n\nThe method as follow:\n1- Make a feature that has random values e.g.\n`df['Random'] = np.random.normal(size=(len(df.index),1))`\n2- Make a model using your original features in addition to this one.\n3- Find the features importances.\n4- We know that our feature consists of just noise. So, any feature that has an importance value less than the importance of our feature is nothing but noise!\n\nVery elegant and simple!",
      "votes": null
    },
    {
      "id": "1908100",
      "postDate": "08/21/2022 11:19:34",
      "content": "<p>It sounds interesting.</p>",
      "rawMarkdown": "It sounds interesting.",
      "votes": null
    },
    {
      "id": "1908800",
      "postDate": "08/22/2022 03:34:31",
      "content": "<p><a href=\"https://towardsdatascience.com/feature-selection-with-borutapy-f0ea84c9366\" target=\"_blank\">Feature Selection With BorutaPy</a><br>\nThis is the advanced version of your idea. Boruta is a very powerful method for feature selection, but it's very computationally expensive when applied on huge datasets</p>",
      "rawMarkdown": "[Feature Selection With BorutaPy](https://towardsdatascience.com/feature-selection-with-borutapy-f0ea84c9366)\nThis is the advanced version of your idea. Boruta is a very powerful method for feature selection, but it's very computationally expensive when applied on huge datasets",
      "votes": null
    },
    {
      "id": "1909028",
      "postDate": "08/22/2022 08:26:32",
      "content": "<p>Hi Mohamed</p>\n<p>Your idea is interesting and simple, but it doesn't seem to work as is. I have checked it. This random feature got an importance of 1463, which is quite a lot for a random variable, and it turned out that it cut off too many useful features, so that the metric on test dropped dramatically. The problem with this approach, I suppose, is that the gradient boosting is overfitting on this useless feature, and we cannot distinguish this feature from those that have good generalization power.</p>",
      "rawMarkdown": "Hi Mohamed\n\nYour idea is interesting and simple, but it doesn't seem to work as is. I have checked it. This random feature got an importance of 1463, which is quite a lot for a random variable, and it turned out that it cut off too many useful features, so that the metric on test dropped dramatically. The problem with this approach, I suppose, is that the gradient boosting is overfitting on this useless feature, and we cannot distinguish this feature from those that have good generalization power.",
      "votes": null
    },
    {
      "id": "1909255",
      "postDate": "08/22/2022 13:16:32",
      "content": "<p>Looks interesting, maybe give it a try.Thanks～</p>",
      "rawMarkdown": "Looks interesting, maybe give it a try.Thanks～",
      "votes": null
    },
    {
      "id": "1909774",
      "postDate": "08/22/2022 23:27:59",
      "content": "<p>Interesting idea <a href=\"https://www.kaggle.com/mohammad2012191\" target=\"_blank\">@mohammad2012191</a>. I'll try it. Thanks for sharing.</p>",
      "rawMarkdown": "Interesting idea @mohammad2012191. I'll try it. Thanks for sharing.",
      "votes": null
    },
    {
      "id": "1922354",
      "postDate": "09/01/2022 12:35:53",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/mohammad2012191\" target=\"_blank\">@mohammad2012191</a>. May I invite you to participate in this survey regarding your experience on Kaggle (10 min)? This is not a scam. We are a group of researchers at the City University of Hong Kong. The survey link is: <a href=\"https://cityuhk.questionpro.com/survey-of-kaggle-contestants\" target=\"_blank\">https://cityuhk.questionpro.com/survey-of-kaggle-contestants</a></p>",
      "rawMarkdown": "Hi @mohammad2012191. May I invite you to participate in this survey regarding your experience on Kaggle (10 min)? This is not a scam. We are a group of researchers at the City University of Hong Kong. The survey link is: https://cityuhk.questionpro.com/survey-of-kaggle-contestants",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1908100,
      "author_name": "pabuoro",
      "author_url": "",
      "post_date": "08/21/2022 11:19:34",
      "content": "<p>It sounds interesting.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1908800,
      "author_name": "pavelvod",
      "author_url": "",
      "post_date": "08/22/2022 03:34:31",
      "content": "<p><a href=\"https://towardsdatascience.com/feature-selection-with-borutapy-f0ea84c9366\" target=\"_blank\">Feature Selection With BorutaPy</a><br>\nThis is the advanced version of your idea. Boruta is a very powerful method for feature selection, but it's very computationally expensive when applied on huge datasets</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1909028,
      "author_name": "andrejvetrov",
      "author_url": "",
      "post_date": "08/22/2022 08:26:32",
      "content": "<p>Hi Mohamed</p>\n<p>Your idea is interesting and simple, but it doesn't seem to work as is. I have checked it. This random feature got an importance of 1463, which is quite a lot for a random variable, and it turned out that it cut off too many useful features, so that the metric on test dropped dramatically. The problem with this approach, I suppose, is that the gradient boosting is overfitting on this useless feature, and we cannot distinguish this feature from those that have good generalization power.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1909255,
      "author_name": "",
      "author_url": "",
      "post_date": "08/22/2022 13:16:32",
      "content": "<p>Looks interesting, maybe give it a try.Thanks～</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1909774,
      "author_name": "oscarm524",
      "author_url": "",
      "post_date": "08/22/2022 23:27:59",
      "content": "<p>Interesting idea <a href=\"https://www.kaggle.com/mohammad2012191\" target=\"_blank\">@mohammad2012191</a>. I'll try it. Thanks for sharing.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1922354,
      "author_name": "lystriving",
      "author_url": "",
      "post_date": "09/01/2022 12:35:53",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/mohammad2012191\" target=\"_blank\">@mohammad2012191</a>. May I invite you to participate in this survey regarding your experience on Kaggle (10 min)? This is not a scam. We are a group of researchers at the City University of Hong Kong. The survey link is: <a href=\"https://cityuhk.questionpro.com/survey-of-kaggle-contestants\" target=\"_blank\">https://cityuhk.questionpro.com/survey-of-kaggle-contestants</a></p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1908012": "Hello everyone,\nAs that this competition requires very good features selection, here is a very elegant and simple method for feature selection that I found once used in Kaggle. Not sure who invented it (Please mention him in the comments if you know him).\n\nThe method as follow:\n1- Make a feature that has random values e.g.\n`df['Random'] = np.random.normal(size=(len(df.index),1))`\n2- Make a model using your original features in addition to this one.\n3- Find the features importances.\n4- We know that our feature consists of just noise. So, any feature that has an importance value less than the importance of our feature is nothing but noise!\n\nVery elegant and simple!",
    "1908100": "It sounds interesting.",
    "1908800": "[Feature Selection With BorutaPy](https://towardsdatascience.com/feature-selection-with-borutapy-f0ea84c9366)\nThis is the advanced version of your idea. Boruta is a very powerful method for feature selection, but it's very computationally expensive when applied on huge datasets",
    "1909028": "Hi Mohamed\n\nYour idea is interesting and simple, but it doesn't seem to work as is. I have checked it. This random feature got an importance of 1463, which is quite a lot for a random variable, and it turned out that it cut off too many useful features, so that the metric on test dropped dramatically. The problem with this approach, I suppose, is that the gradient boosting is overfitting on this useless feature, and we cannot distinguish this feature from those that have good generalization power.",
    "1909255": "Looks interesting, maybe give it a try.Thanks～",
    "1909774": "Interesting idea @mohammad2012191. I'll try it. Thanks for sharing.",
    "1922354": "Hi @mohammad2012191. May I invite you to participate in this survey regarding your experience on Kaggle (10 min)? This is not a scam. We are a group of researchers at the City University of Hong Kong. The survey link is: https://cityuhk.questionpro.com/survey-of-kaggle-contestants"
  },
  "source": "meta"
}