{
  "id": 487696,
  "title": "Creating a weapon of math destruction",
  "url": "/competitions/home-credit-credit-risk-model-stability/discussion/487696",
  "author_name": "",
  "post_date": "2024-03-30T07:05:58.008416600Z",
  "votes": 1,
  "comment_count": 2,
  "views": 0,
  "content": "<p>So I've been looking at the feature importance of my model and the top two most important features are, by an order of magnitude, basically <em>where you live</em>. I don't think it's too much of a stretch to suspect that these features are just correlates to class. If so, I'm worried that any model that uses them will do the opposite of what is suggested is a goal of the competition, i.e., \"If data science could help better predict one’s repayment capabilities, loans might become more accessible to those who may benefit from them the most.\"</p>\n<p>Interestingly, the model doesn't seem to suffer if I don't use these two features. However, there might be other features it is using to get at the same class information.</p>\n<p>So I think I'll try not to use the features and be on the lookout for others but it's possible that doing well in the competition will require them.</p>\n<p>(also, the title comes from the book <a href=\"url\" target=\"_blank\">https://en.wikipedia.org/wiki/Weapons_of_Math_Destruction</a> if anyone was curious)</p>",
  "messages": [
    {
      "id": "2723393",
      "postDate": "03/30/2024 07:05:58",
      "content": "<p>So I've been looking at the feature importance of my model and the top two most important features are, by an order of magnitude, basically <em>where you live</em>. I don't think it's too much of a stretch to suspect that these features are just correlates to class. If so, I'm worried that any model that uses them will do the opposite of what is suggested is a goal of the competition, i.e., \"If data science could help better predict one’s repayment capabilities, loans might become more accessible to those who may benefit from them the most.\"</p>\n<p>Interestingly, the model doesn't seem to suffer if I don't use these two features. However, there might be other features it is using to get at the same class information.</p>\n<p>So I think I'll try not to use the features and be on the lookout for others but it's possible that doing well in the competition will require them.</p>\n<p>(also, the title comes from the book <a href=\"url\" target=\"_blank\">https://en.wikipedia.org/wiki/Weapons_of_Math_Destruction</a> if anyone was curious)</p>",
      "rawMarkdown": "So I've been looking at the feature importance of my model and the top two most important features are, by an order of magnitude, basically *where you live*. I don't think it's too much of a stretch to suspect that these features are just correlates to class. If so, I'm worried that any model that uses them will do the opposite of what is suggested is a goal of the competition, i.e., \"If data science could help better predict one’s repayment capabilities, loans might become more accessible to those who may benefit from them the most.\"\n\nInterestingly, the model doesn't seem to suffer if I don't use these two features. However, there might be other features it is using to get at the same class information.\n\nSo I think I'll try not to use the features and be on the lookout for others but it's possible that doing well in the competition will require them.\n\n(also, the title comes from the book [https://en.wikipedia.org/wiki/Weapons_of_Math_Destruction](url) if anyone was curious)",
      "votes": null
    },
    {
      "id": "2723423",
      "postDate": "03/30/2024 07:51:55",
      "content": "<p>Dear <a href=\"https://www.kaggle.com/caelhasse\" target=\"_blank\">@caelhasse</a> </p>\n<blockquote>\n  <p>\"<em>…the top two most important features are,…</em></p>\n</blockquote>\n<p>However,</p>\n<blockquote>\n  <p>\"<em>…the model doesn't seem to suffer if I don't use these two features</em>\"</p>\n</blockquote>\n<p>I think there is, almost by definition, something amiss with your feature importance calculation.</p>\n<p>All the best,<br>\ncarl</p>",
      "rawMarkdown": "Dear @caelhasse \n\n> \"*...the top two most important features are,...*\n\nHowever,\n\n> \"*...the model doesn't seem to suffer if I don't use these two features*\"\n\nI think there is, almost by definition, something amiss with your feature importance calculation.\n\nAll the best,\ncarl",
      "votes": null
    },
    {
      "id": "2723561",
      "postDate": "03/30/2024 10:11:50",
      "content": "<p>I'm just using a LightGBM model with the .feature_importance(importance_type=\"gain\") method. My guess was that there isn't a unique good choice of parameters and the importance calc is dependent on the precise way it is trained. Indeed, I just tested calculating importance from trainings with different seeds and get slightly different results.</p>",
      "rawMarkdown": "I'm just using a LightGBM model with the .feature_importance(importance_type=\"gain\") method. My guess was that there isn't a unique good choice of parameters and the importance calc is dependent on the precise way it is trained. Indeed, I just tested calculating importance from trainings with different seeds and get slightly different results.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2723423,
      "author_name": "carlmcbrideellis",
      "author_url": "",
      "post_date": "03/30/2024 07:51:55",
      "content": "<p>Dear <a href=\"https://www.kaggle.com/caelhasse\" target=\"_blank\">@caelhasse</a> </p>\n<blockquote>\n  <p>\"<em>…the top two most important features are,…</em></p>\n</blockquote>\n<p>However,</p>\n<blockquote>\n  <p>\"<em>…the model doesn't seem to suffer if I don't use these two features</em>\"</p>\n</blockquote>\n<p>I think there is, almost by definition, something amiss with your feature importance calculation.</p>\n<p>All the best,<br>\ncarl</p>",
      "votes": null,
      "replies": [
        {
          "id": 2723561,
          "author_name": "caelhasse",
          "author_url": "",
          "post_date": "03/30/2024 10:11:50",
          "content": "<p>I'm just using a LightGBM model with the .feature_importance(importance_type=\"gain\") method. My guess was that there isn't a unique good choice of parameters and the importance calc is dependent on the precise way it is trained. Indeed, I just tested calculating importance from trainings with different seeds and get slightly different results.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2723393": "So I've been looking at the feature importance of my model and the top two most important features are, by an order of magnitude, basically *where you live*. I don't think it's too much of a stretch to suspect that these features are just correlates to class. If so, I'm worried that any model that uses them will do the opposite of what is suggested is a goal of the competition, i.e., \"If data science could help better predict one’s repayment capabilities, loans might become more accessible to those who may benefit from them the most.\"\n\nInterestingly, the model doesn't seem to suffer if I don't use these two features. However, there might be other features it is using to get at the same class information.\n\nSo I think I'll try not to use the features and be on the lookout for others but it's possible that doing well in the competition will require them.\n\n(also, the title comes from the book [https://en.wikipedia.org/wiki/Weapons_of_Math_Destruction](url) if anyone was curious)",
    "2723423": "Dear @caelhasse \n\n> \"*...the top two most important features are,...*\n\nHowever,\n\n> \"*...the model doesn't seem to suffer if I don't use these two features*\"\n\nI think there is, almost by definition, something amiss with your feature importance calculation.\n\nAll the best,\ncarl",
    "2723561": "I'm just using a LightGBM model with the .feature_importance(importance_type=\"gain\") method. My guess was that there isn't a unique good choice of parameters and the importance calc is dependent on the precise way it is trained. Indeed, I just tested calculating importance from trainings with different seeds and get slightly different results."
  },
  "source": "meta"
}