{
  "id": 484134,
  "title": "Be careful with time features",
  "url": "/competitions/home-credit-credit-risk-model-stability/discussion/484134",
  "author_name": "",
  "post_date": "2024-03-15T13:54:25.034684100Z",
  "votes": 51,
  "comment_count": 6,
  "views": 0,
  "content": "<p>I noticed that some features without the D suffix represent the month and year, so their distribution changes depending on time. Since decision tree-based models cannot perform extrapolation, such features harm model performance. You can find them by searching for \"year\" and \"month\" in the column names.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4516327%2F85f599536307412872f18989105ee547%2Fnewplot%20(2).png?generation=1710510274233825&amp;alt=media\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4516327%2F507070a839a91c21297dc74d52c81bb5%2Fnewplot%20(1).png?generation=1710510526060588&amp;alt=media\"></p>",
  "messages": [
    {
      "id": "2698418",
      "postDate": "03/15/2024 13:54:25",
      "content": "<p>I noticed that some features without the D suffix represent the month and year, so their distribution changes depending on time. Since decision tree-based models cannot perform extrapolation, such features harm model performance. You can find them by searching for \"year\" and \"month\" in the column names.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4516327%2F85f599536307412872f18989105ee547%2Fnewplot%20(2).png?generation=1710510274233825&amp;alt=media\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4516327%2F507070a839a91c21297dc74d52c81bb5%2Fnewplot%20(1).png?generation=1710510526060588&amp;alt=media\"></p>",
      "rawMarkdown": "I noticed that some features without the D suffix represent the month and year, so their distribution changes depending on time. Since decision tree-based models cannot perform extrapolation, such features harm model performance. You can find them by searching for \"year\" and \"month\" in the column names.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4516327%2F85f599536307412872f18989105ee547%2Fnewplot%20(2).png?generation=1710510274233825&alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4516327%2F507070a839a91c21297dc74d52c81bb5%2Fnewplot%20(1).png?generation=1710510526060588&alt=media)",
      "votes": null
    },
    {
      "id": "2703826",
      "postDate": "03/18/2024 12:30:43",
      "content": "<p>How should we deal with such features?</p>",
      "rawMarkdown": "How should we deal with such features?",
      "votes": null
    },
    {
      "id": "2704664",
      "postDate": "03/18/2024 21:00:52",
      "content": "<p>You can subtract from prediction date (date_decision). In such cases, you can prevent feature drift by setting the date of the prediction as base point</p>",
      "rawMarkdown": "You can subtract from prediction date (date_decision). In such cases, you can prevent feature drift by setting the date of the prediction as base point",
      "votes": null
    },
    {
      "id": "2715418",
      "postDate": "03/25/2024 13:34:28",
      "content": "<p>Does the \"Test dataset\" in first figure refer to public test table?</p>",
      "rawMarkdown": "Does the \"Test dataset\" in first figure refer to public test table?",
      "votes": null
    },
    {
      "id": "2715977",
      "postDate": "03/25/2024 18:58:58",
      "content": "<p>I configured weeks 0-45 of the public data as train, and weeks 45-90 as test data set</p>",
      "rawMarkdown": "I configured weeks 0-45 of the public data as train, and weeks 45-90 as test data set",
      "votes": null
    },
    {
      "id": "2716028",
      "postDate": "03/25/2024 19:27:48",
      "content": "<p>Excuse me for interrupting. Did you find any correlation between  Local Score and LB when implementing the Time Series data split method?</p>\n<p>I added feature engineering to the provided notebook (using 5-fold cross-validation without considering time series), but CV and LB did not correlate.</p>",
      "rawMarkdown": "Excuse me for interrupting. Did you find any correlation between  Local Score and LB when implementing the Time Series data split method?\n\nI added feature engineering to the provided notebook (using 5-fold cross-validation without considering time series), but CV and LB did not correlate.",
      "votes": null
    },
    {
      "id": "2717113",
      "postDate": "03/26/2024 11:31:20",
      "content": "<p>Unfortunately I couldn't find any correlation. Now I'm starting from scratch. Scores over 0.540 can be achieved with static tables</p>",
      "rawMarkdown": "Unfortunately I couldn't find any correlation. Now I'm starting from scratch. Scores over 0.540 can be achieved with static tables",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2703826,
      "author_name": "uiosun",
      "author_url": "",
      "post_date": "03/18/2024 12:30:43",
      "content": "<p>How should we deal with such features?</p>",
      "votes": null,
      "replies": [
        {
          "id": 2704664,
          "author_name": "greysky",
          "author_url": "",
          "post_date": "03/18/2024 21:00:52",
          "content": "<p>You can subtract from prediction date (date_decision). In such cases, you can prevent feature drift by setting the date of the prediction as base point</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2715418,
      "author_name": "dimakoshman",
      "author_url": "",
      "post_date": "03/25/2024 13:34:28",
      "content": "<p>Does the \"Test dataset\" in first figure refer to public test table?</p>",
      "votes": null,
      "replies": [
        {
          "id": 2715977,
          "author_name": "greysky",
          "author_url": "",
          "post_date": "03/25/2024 18:58:58",
          "content": "<p>I configured weeks 0-45 of the public data as train, and weeks 45-90 as test data set</p>",
          "votes": null,
          "replies": [
            {
              "id": 2716028,
              "author_name": "youheitomio",
              "author_url": "",
              "post_date": "03/25/2024 19:27:48",
              "content": "<p>Excuse me for interrupting. Did you find any correlation between  Local Score and LB when implementing the Time Series data split method?</p>\n<p>I added feature engineering to the provided notebook (using 5-fold cross-validation without considering time series), but CV and LB did not correlate.</p>",
              "votes": null,
              "replies": [
                {
                  "id": 2717113,
                  "author_name": "greysky",
                  "author_url": "",
                  "post_date": "03/26/2024 11:31:20",
                  "content": "<p>Unfortunately I couldn't find any correlation. Now I'm starting from scratch. Scores over 0.540 can be achieved with static tables</p>",
                  "votes": null,
                  "replies": []
                }
              ]
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2698418": "I noticed that some features without the D suffix represent the month and year, so their distribution changes depending on time. Since decision tree-based models cannot perform extrapolation, such features harm model performance. You can find them by searching for \"year\" and \"month\" in the column names.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4516327%2F85f599536307412872f18989105ee547%2Fnewplot%20(2).png?generation=1710510274233825&alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4516327%2F507070a839a91c21297dc74d52c81bb5%2Fnewplot%20(1).png?generation=1710510526060588&alt=media)",
    "2703826": "How should we deal with such features?",
    "2704664": "You can subtract from prediction date (date_decision). In such cases, you can prevent feature drift by setting the date of the prediction as base point",
    "2715418": "Does the \"Test dataset\" in first figure refer to public test table?",
    "2715977": "I configured weeks 0-45 of the public data as train, and weeks 45-90 as test data set",
    "2716028": "Excuse me for interrupting. Did you find any correlation between  Local Score and LB when implementing the Time Series data split method?\n\nI added feature engineering to the provided notebook (using 5-fold cross-validation without considering time series), but CV and LB did not correlate.",
    "2717113": "Unfortunately I couldn't find any correlation. Now I'm starting from scratch. Scores over 0.540 can be achieved with static tables"
  },
  "source": "meta"
}