{
  "id": 491692,
  "title": "How to find unstable features?",
  "url": "/competitions/home-credit-credit-risk-model-stability/discussion/491692",
  "author_name": "",
  "post_date": "2024-04-06T23:46:11.045223700Z",
  "votes": 12,
  "comment_count": 4,
  "views": 0,
  "content": "<p>We all know that the evaluation metric of this competition place high demands on the stability of the model. When the 'mean_gini' of the model increases, the stability of the model may decrease, ultimately leading to a low score.</p>\n<p>Compared to improving the effectiveness of the model, identifying these features that affect stability and removing them may improve the score.</p>\n<p>My current approach is to divide the training data into two equal parts and compare the pearson correlation coefficients between the features and the target in the two parts. If there is too much fluctuation in the correlation and this feature is not particularly important, remove it.</p>\n<p>If you have any good ideas, please feel free to share them.</p>",
  "messages": [
    {
      "id": "2739287",
      "postDate": "04/06/2024 23:46:11",
      "content": "<p>We all know that the evaluation metric of this competition place high demands on the stability of the model. When the 'mean_gini' of the model increases, the stability of the model may decrease, ultimately leading to a low score.</p>\n<p>Compared to improving the effectiveness of the model, identifying these features that affect stability and removing them may improve the score.</p>\n<p>My current approach is to divide the training data into two equal parts and compare the pearson correlation coefficients between the features and the target in the two parts. If there is too much fluctuation in the correlation and this feature is not particularly important, remove it.</p>\n<p>If you have any good ideas, please feel free to share them.</p>",
      "rawMarkdown": "We all know that the evaluation metric of this competition place high demands on the stability of the model. When the 'mean_gini' of the model increases, the stability of the model may decrease, ultimately leading to a low score.\n\nCompared to improving the effectiveness of the model, identifying these features that affect stability and removing them may improve the score.\n\nMy current approach is to divide the training data into two equal parts and compare the pearson correlation coefficients between the features and the target in the two parts. If there is too much fluctuation in the correlation and this feature is not particularly important, remove it.\n\n If you have any good ideas, please feel free to share them.",
      "votes": null
    },
    {
      "id": "2739825",
      "postDate": "04/07/2024 10:36:27",
      "content": "<p>Adversarial Validation</p>",
      "rawMarkdown": "Adversarial Validation",
      "votes": null
    },
    {
      "id": "2739871",
      "postDate": "04/07/2024 11:17:21",
      "content": "<p>I know this, this competition is based on temporal features, is it really feasible to use this?</p>",
      "rawMarkdown": "I know this, this competition is based on temporal features, is it really feasible to use this?",
      "votes": null
    },
    {
      "id": "2739956",
      "postDate": "04/07/2024 12:27:33",
      "content": "<p>some of the simple things that can be done is by removing the redudent features, or by removing the features that have very less discriminatory power. then you can also check the split histogram for a feature to understand how stable that particular feature is for a tree based algorithm, etc.</p>",
      "rawMarkdown": "some of the simple things that can be done is by removing the redudent features, or by removing the features that have very less discriminatory power. then you can also check the split histogram for a feature to understand how stable that particular feature is for a tree based algorithm, etc.",
      "votes": null
    },
    {
      "id": "2741669",
      "postDate": "04/08/2024 14:21:41",
      "content": "<p><a href=\"https://www.kaggle.com/greysky\" target=\"_blank\">@greysky</a> Posted a nice notebook. You may have a look <a href=\"https://www.kaggle.com/code/greysky/home-credit-deepchecks\" target=\"_blank\">https://www.kaggle.com/code/greysky/home-credit-deepchecks</a> </p>",
      "rawMarkdown": "greysky Posted a nice notebook. You may have a look https://www.kaggle.com/code/greysky/home-credit-deepchecks",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2739825,
      "author_name": "denisplatonovua",
      "author_url": "",
      "post_date": "04/07/2024 10:36:27",
      "content": "<p>Adversarial Validation</p>",
      "votes": null,
      "replies": [
        {
          "id": 2739871,
          "author_name": "yunsuxiaozi",
          "author_url": "",
          "post_date": "04/07/2024 11:17:21",
          "content": "<p>I know this, this competition is based on temporal features, is it really feasible to use this?</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2739956,
      "author_name": "shreyas9181",
      "author_url": "",
      "post_date": "04/07/2024 12:27:33",
      "content": "<p>some of the simple things that can be done is by removing the redudent features, or by removing the features that have very less discriminatory power. then you can also check the split histogram for a feature to understand how stable that particular feature is for a tree based algorithm, etc.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2741669,
      "author_name": "lucamtb",
      "author_url": "",
      "post_date": "04/08/2024 14:21:41",
      "content": "<p><a href=\"https://www.kaggle.com/greysky\" target=\"_blank\">@greysky</a> Posted a nice notebook. You may have a look <a href=\"https://www.kaggle.com/code/greysky/home-credit-deepchecks\" target=\"_blank\">https://www.kaggle.com/code/greysky/home-credit-deepchecks</a> </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2739287": "We all know that the evaluation metric of this competition place high demands on the stability of the model. When the 'mean_gini' of the model increases, the stability of the model may decrease, ultimately leading to a low score.\n\nCompared to improving the effectiveness of the model, identifying these features that affect stability and removing them may improve the score.\n\nMy current approach is to divide the training data into two equal parts and compare the pearson correlation coefficients between the features and the target in the two parts. If there is too much fluctuation in the correlation and this feature is not particularly important, remove it.\n\n If you have any good ideas, please feel free to share them.",
    "2739825": "Adversarial Validation",
    "2739871": "I know this, this competition is based on temporal features, is it really feasible to use this?",
    "2739956": "some of the simple things that can be done is by removing the redudent features, or by removing the features that have very less discriminatory power. then you can also check the split histogram for a feature to understand how stable that particular feature is for a tree based algorithm, etc.",
    "2741669": "greysky Posted a nice notebook. You may have a look https://www.kaggle.com/code/greysky/home-credit-deepchecks"
  },
  "source": "meta"
}