{
  "id": 494087,
  "title": "Max features is all you need?",
  "url": "/competitions/home-credit-credit-risk-model-stability/discussion/494087",
  "author_name": "yunsuxiaozi",
  "post_date": "2024-04-15T23:50:28.685000",
  "votes": 15,
  "comment_count": 10,
  "views": 0,
  "content": "<p>This morning, I saw a high scoring open-source code that only extracted the 'max' feature and achieved a high score of 0.586.</p>\n<p><a href=\"https://www.kaggle.com/code/harrychan123/home-credit-lgb-cat-ensemble\" target=\"_blank\">https://www.kaggle.com/code/harrychan123/home-credit-lgb-cat-ensemble</a></p>\n<p>Through my extensive experiments, I have come to the conclusion that if we construct too few features, the model's gini may not be high but has good stability, resulting in a high score. If we construct too many features, the gini will be high but the stability will be poor, so the score will not be higher than the former.</p>\n<p>I currently prefer to construct notebooks with fewer features and better stability.</p>",
  "messages": [
    {
      "id": 2754290,
      "postDate": "2024-04-15T23:50:28.687Z",
      "content": "<p>This morning, I saw a high scoring open-source code that only extracted the 'max' feature and achieved a high score of 0.586.</p>\n<p><a href=\"https://www.kaggle.com/code/harrychan123/home-credit-lgb-cat-ensemble\" target=\"_blank\">https://www.kaggle.com/code/harrychan123/home-credit-lgb-cat-ensemble</a></p>\n<p>Through my extensive experiments, I have come to the conclusion that if we construct too few features, the model's gini may not be high but has good stability, resulting in a high score. If we construct too many features, the gini will be high but the stability will be poor, so the score will not be higher than the former.</p>\n<p>I currently prefer to construct notebooks with fewer features and better stability.</p>",
      "rawMarkdown": "This morning, I saw a high scoring open-source code that only extracted the 'max' feature and achieved a high score of 0.586.\n\nhttps://www.kaggle.com/code/harrychan123/home-credit-lgb-cat-ensemble\n\nThrough my extensive experiments, I have come to the conclusion that if we construct too few features, the model's gini may not be high but has good stability, resulting in a high score. If we construct too many features, the gini will be high but the stability will be poor, so the score will not be higher than the former.\n\n\nI currently prefer to construct notebooks with fewer features and better stability.\n\n\n",
      "votes": 15
    },
    {
      "id": 2756765,
      "postDate": "2024-04-17T06:55:10.897Z",
      "content": "<p>the cv of that model is very very poor and the things they do obviously do not make much sense at all, so they are just indirectly exploiting the metric by making a bad model.</p>\n<p>the sad thing is that this will probably just carry over to the private LB aswell, so the host will just get a bunch of bad models and will learn nothing at all, but they've been warned so much about the term being way too large, so i guess that's just what they want.</p>",
      "rawMarkdown": "the cv of that model is very very poor and the things they do obviously do not make much sense at all, so they are just indirectly exploiting the metric by making a bad model.\n\nthe sad thing is that this will probably just carry over to the private LB aswell, so the host will just get a bunch of bad models and will learn nothing at all, but they've been warned so much about the term being way too large, so i guess that's just what they want.",
      "votes": 4,
      "replies": [
        {
          "id": 2757556,
          "postDate": "2024-04-17T15:21:01.890Z",
          "content": "<p>Are you trying to say that the issue of stability metrics remains unresolved?<br>\nI think using max alone is a meaningless statistical feature.<br>\nWhat baffles me is that a single model like theirs that simply uses only statistical features can get such a high LB score.</p>",
          "rawMarkdown": "Are you trying to say that the issue of stability metrics remains unresolved?\nI think using max alone is a meaningless statistical feature.\nWhat baffles me is that a single model like theirs that simply uses only statistical features can get such a high LB score."
        }
      ]
    },
    {
      "id": 2756209,
      "postDate": "2024-04-17T00:08:30.827Z",
      "content": "<p>It seems that the high score was achieved with version 1, where all features were included (min, last, mean etc)</p>",
      "rawMarkdown": "It seems that the high score was achieved with version 1, where all features were included (min, last, mean etc)",
      "votes": 1
    },
    {
      "id": 2754360,
      "postDate": "2024-04-16T01:39:43.590Z",
      "content": "<p>I also tried adding more features, but the results became worse.</p>",
      "rawMarkdown": "I also tried adding more features, but the results became worse.",
      "votes": 1
    },
    {
      "id": 2754305,
      "postDate": "2024-04-16T00:26:58.920Z",
      "content": "<p>I also have this doubt, because theoretically speaking, there doesn't seem to be much logic in using 'max' for str. I tried replacing 'max' with 'last' for str, and the effect was very poor.</p>",
      "rawMarkdown": "I also have this doubt, because theoretically speaking, there doesn't seem to be much logic in using 'max' for str. I tried replacing 'max' with 'last' for str, and the effect was very poor.",
      "votes": 1,
      "replies": [
        {
          "id": 2754315,
          "postDate": "2024-04-16T00:47:05.597Z",
          "content": "<p>IMO, just like our maximum pooling in CNN, this competition aims to capture anomalies and expose more obvious features as much as possible. As for strings, the time string will retain the latest time, and non categorical strings will also be removed.</p>",
          "rawMarkdown": "IMO, just like our maximum pooling in CNN, this competition aims to capture anomalies and expose more obvious features as much as possible. As for strings, the time string will retain the latest time, and non categorical strings will also be removed."
        },
        {
          "id": 2755439,
          "postDate": "2024-04-16T14:06:20.697Z",
          "rawMarkdown": "",
          "isDeleted": true
        }
      ]
    },
    {
      "id": 2766900,
      "postDate": "2024-04-22T02:27:48.607Z",
      "content": "<p>There really isn't any logic in using the max function for aggregation of category features, so I tried converting the category features to numeric values in some way (e.g. count encoding) and then aggregating them, but it resulted in a drop in the score</p>",
      "rawMarkdown": "There really isn't any logic in using the max function for aggregation of category features, so I tried converting the category features to numeric values in some way (e.g. count encoding) and then aggregating them, but it resulted in a drop in the score",
      "replies": [
        {
          "id": 2787947,
          "postDate": "2024-05-02T02:59:44.640Z",
          "content": "<p>Do you understand now why the score dropped? From my cognitive perspective, these are all characteristics that can lead to overfitting.</p>",
          "rawMarkdown": "Do you understand now why the score dropped? From my cognitive perspective, these are all characteristics that can lead to overfitting.\n\n",
          "isDeleted": true
        }
      ]
    },
    {
      "id": 2755193,
      "postDate": "2024-04-16T12:13:59.683Z",
      "content": "<p>Interesting, I believe addition of high time sensitive features and variability are leading to the stability loss and would have been inducing bias during learning and tight coupling with target variable, can you please share what were some of the variables that on addition reduced stability scores, would be helpful to further dig down.</p>",
      "rawMarkdown": "Interesting, I believe addition of high time sensitive features and variability are leading to the stability loss and would have been inducing bias during learning and tight coupling with target variable, can you please share what were some of the variables that on addition reduced stability scores, would be helpful to further dig down."
    }
  ],
  "comments": [
    {
      "id": 2756765,
      "author_name": "at7459",
      "author_url": "",
      "post_date": "2024-04-17T06:55:10.897000",
      "content": "<p>the cv of that model is very very poor and the things they do obviously do not make much sense at all, so they are just indirectly exploiting the metric by making a bad model.</p>\n<p>the sad thing is that this will probably just carry over to the private LB aswell, so the host will just get a bunch of bad models and will learn nothing at all, but they've been warned so much about the term being way too large, so i guess that's just what they want.</p>",
      "votes": 4,
      "replies": [
        {
          "id": 2757556,
          "author_name": "WenJhuang",
          "author_url": "",
          "post_date": "2024-04-17T15:21:01.890000",
          "content": "<p>Are you trying to say that the issue of stability metrics remains unresolved?<br>\nI think using max alone is a meaningless statistical feature.<br>\nWhat baffles me is that a single model like theirs that simply uses only statistical features can get such a high LB score.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2756209,
      "author_name": "Evgeniia Grigoreva",
      "author_url": "",
      "post_date": "2024-04-17T00:08:30.827000",
      "content": "<p>It seems that the high score was achieved with version 1, where all features were included (min, last, mean etc)</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2754360,
      "author_name": "Luck is what I need",
      "author_url": "",
      "post_date": "2024-04-16T01:39:43.590000",
      "content": "<p>I also tried adding more features, but the results became worse.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2754305,
      "author_name": "renxiaohan",
      "author_url": "",
      "post_date": "2024-04-16T00:26:58.920000",
      "content": "<p>I also have this doubt, because theoretically speaking, there doesn't seem to be much logic in using 'max' for str. I tried replacing 'max' with 'last' for str, and the effect was very poor.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2754315,
          "author_name": "yunsuxiaozi",
          "author_url": "",
          "post_date": "2024-04-16T00:47:05.597000",
          "content": "<p>IMO, just like our maximum pooling in CNN, this competition aims to capture anomalies and expose more obvious features as much as possible. As for strings, the time string will retain the latest time, and non categorical strings will also be removed.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 2755439,
          "author_name": "",
          "author_url": "",
          "post_date": "2024-04-16T14:06:20.697000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2766900,
      "author_name": "samuel tan",
      "author_url": "",
      "post_date": "2024-04-22T02:27:48.607000",
      "content": "<p>There really isn't any logic in using the max function for aggregation of category features, so I tried converting the category features to numeric values in some way (e.g. count encoding) and then aggregating them, but it resulted in a drop in the score</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2787947,
          "author_name": "",
          "author_url": "",
          "post_date": "2024-05-02T02:59:44.640000",
          "content": "<p>Do you understand now why the score dropped? From my cognitive perspective, these are all characteristics that can lead to overfitting.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2755193,
      "author_name": "VedangBhardwaj_0114",
      "author_url": "",
      "post_date": "2024-04-16T12:13:59.683000",
      "content": "<p>Interesting, I believe addition of high time sensitive features and variability are leading to the stability loss and would have been inducing bias during learning and tight coupling with target variable, can you please share what were some of the variables that on addition reduced stability scores, would be helpful to further dig down.</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2754290": "This morning, I saw a high scoring open-source code that only extracted the 'max' feature and achieved a high score of 0.586.\n\nhttps://www.kaggle.com/code/harrychan123/home-credit-lgb-cat-ensemble\n\nThrough my extensive experiments, I have come to the conclusion that if we construct too few features, the model's gini may not be high but has good stability, resulting in a high score. If we construct too many features, the gini will be high but the stability will be poor, so the score will not be higher than the former.\n\n\nI currently prefer to construct notebooks with fewer features and better stability.\n\n\n",
    "2756765": "the cv of that model is very very poor and the things they do obviously do not make much sense at all, so they are just indirectly exploiting the metric by making a bad model.\n\nthe sad thing is that this will probably just carry over to the private LB aswell, so the host will just get a bunch of bad models and will learn nothing at all, but they've been warned so much about the term being way too large, so i guess that's just what they want.",
    "2756209": "It seems that the high score was achieved with version 1, where all features were included (min, last, mean etc)",
    "2754360": "I also tried adding more features, but the results became worse.",
    "2754305": "I also have this doubt, because theoretically speaking, there doesn't seem to be much logic in using 'max' for str. I tried replacing 'max' with 'last' for str, and the effect was very poor.",
    "2766900": "There really isn't any logic in using the max function for aggregation of category features, so I tried converting the category features to numeric values in some way (e.g. count encoding) and then aggregating them, but it resulted in a drop in the score",
    "2755193": "Interesting, I believe addition of high time sensitive features and variability are leading to the stability loss and would have been inducing bias during learning and tight coupling with target variable, can you please share what were some of the variables that on addition reduced stability scores, would be helpful to further dig down."
  }
}