{
  "id": 476603,
  "title": "Post Decision Date Events",
  "url": "/competitions/home-credit-credit-risk-model-stability/discussion/476603",
  "author_name": "",
  "post_date": "2024-02-13T02:40:18.960987200Z",
  "votes": 1,
  "comment_count": 4,
  "views": 0,
  "content": "<p>Just checking if I got it right. </p>\n<p>It appears that the date columns in the datasets train_static_0_0 and train_static_0_1, such as validfrom_1069D, lastrejectdate_50D, and others, mostly contain dates that occur after the decision date. In contexts like loan approvals, this implies that these columns represent events happening after the decision was made. That means that this data refers to a period posterior of the decision date and, therefore, doesn't make sense to used it for prediction purposes, right? </p>",
  "messages": [
    {
      "id": "2649619",
      "postDate": "02/13/2024 02:40:18",
      "content": "<p>Just checking if I got it right. </p>\n<p>It appears that the date columns in the datasets train_static_0_0 and train_static_0_1, such as validfrom_1069D, lastrejectdate_50D, and others, mostly contain dates that occur after the decision date. In contexts like loan approvals, this implies that these columns represent events happening after the decision was made. That means that this data refers to a period posterior of the decision date and, therefore, doesn't make sense to used it for prediction purposes, right? </p>",
      "rawMarkdown": "Just checking if I got it right. \n\nIt appears that the date columns in the datasets train_static_0_0 and train_static_0_1, such as validfrom_1069D, lastrejectdate_50D, and others, mostly contain dates that occur after the decision date. In contexts like loan approvals, this implies that these columns represent events happening after the decision was made. That means that this data refers to a period posterior of the decision date and, therefore, doesn't make sense to used it for prediction purposes, right?",
      "votes": null
    },
    {
      "id": "2650328",
      "postDate": "02/13/2024 11:50:12",
      "content": "<p>Thanks for noticing that. All date columns were transformed. You can use those dates freely :)</p>",
      "rawMarkdown": "Thanks for noticing that. All date columns were transformed. You can use those dates freely :)",
      "votes": null
    },
    {
      "id": "2798676",
      "postDate": "05/07/2024 11:00:39",
      "content": "<p>Hi Daniel, could you explain why we can use the transformed dates when these dates actually represent data post decision as <a href=\"https://www.kaggle.com/JOS\" target=\"_blank\">@JOS</a>ÉAUGUSTONETO  mentioned. </p>",
      "rawMarkdown": "Hi Daniel, could you explain why we can use the transformed dates when these dates actually represent data post decision as @JOSÉAUGUSTONETO  mentioned.",
      "votes": null
    },
    {
      "id": "2798787",
      "postDate": "05/07/2024 12:22:11",
      "content": "<p><a href=\"https://www.kaggle.com/manojmangam\" target=\"_blank\">@manojmangam</a> When the client is processed internally on a new request for a loan we collect information and most of that is present to you in the dataset. Some of the data at the date of application we create from already collected information or we also have external sources of information about the client. Some of the dates might be in future to the respect of application date, but they are collected at the date of application, so there is no target leakage. </p>\n<p>A made-up example could be a column that tells you when is your next birthday going to be. It is generated at the date of application when the default - target is not known, so there is no information about the future target indicated by the date of next birthday.</p>\n<p>Hope it is clear now. </p>",
      "rawMarkdown": "manojmangam When the client is processed internally on a new request for a loan we collect information and most of that is present to you in the dataset. Some of the data at the date of application we create from already collected information or we also have external sources of information about the client. Some of the dates might be in future to the respect of application date, but they are collected at the date of application, so there is no target leakage. \n\nA made-up example could be a column that tells you when is your next birthday going to be. It is generated at the date of application when the default - target is not known, so there is no information about the future target indicated by the date of next birthday.\n\nHope it is clear now.",
      "votes": null
    },
    {
      "id": "2803336",
      "postDate": "05/09/2024 12:36:36",
      "content": "<p>Thanks Daniel.</p>",
      "rawMarkdown": "Thanks Daniel.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2650328,
      "author_name": "jetakow",
      "author_url": "",
      "post_date": "02/13/2024 11:50:12",
      "content": "<p>Thanks for noticing that. All date columns were transformed. You can use those dates freely :)</p>",
      "votes": null,
      "replies": [
        {
          "id": 2798676,
          "author_name": "manojmangam",
          "author_url": "",
          "post_date": "05/07/2024 11:00:39",
          "content": "<p>Hi Daniel, could you explain why we can use the transformed dates when these dates actually represent data post decision as <a href=\"https://www.kaggle.com/JOS\" target=\"_blank\">@JOS</a>ÉAUGUSTONETO  mentioned. </p>",
          "votes": null,
          "replies": [
            {
              "id": 2798787,
              "author_name": "jetakow",
              "author_url": "",
              "post_date": "05/07/2024 12:22:11",
              "content": "<p><a href=\"https://www.kaggle.com/manojmangam\" target=\"_blank\">@manojmangam</a> When the client is processed internally on a new request for a loan we collect information and most of that is present to you in the dataset. Some of the data at the date of application we create from already collected information or we also have external sources of information about the client. Some of the dates might be in future to the respect of application date, but they are collected at the date of application, so there is no target leakage. </p>\n<p>A made-up example could be a column that tells you when is your next birthday going to be. It is generated at the date of application when the default - target is not known, so there is no information about the future target indicated by the date of next birthday.</p>\n<p>Hope it is clear now. </p>",
              "votes": null,
              "replies": [
                {
                  "id": 2803336,
                  "author_name": "manojmangam",
                  "author_url": "",
                  "post_date": "05/09/2024 12:36:36",
                  "content": "<p>Thanks Daniel.</p>",
                  "votes": null,
                  "replies": []
                }
              ]
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2649619": "Just checking if I got it right. \n\nIt appears that the date columns in the datasets train_static_0_0 and train_static_0_1, such as validfrom_1069D, lastrejectdate_50D, and others, mostly contain dates that occur after the decision date. In contexts like loan approvals, this implies that these columns represent events happening after the decision was made. That means that this data refers to a period posterior of the decision date and, therefore, doesn't make sense to used it for prediction purposes, right?",
    "2650328": "Thanks for noticing that. All date columns were transformed. You can use those dates freely :)",
    "2798676": "Hi Daniel, could you explain why we can use the transformed dates when these dates actually represent data post decision as @JOSÉAUGUSTONETO  mentioned.",
    "2798787": "manojmangam When the client is processed internally on a new request for a loan we collect information and most of that is present to you in the dataset. Some of the data at the date of application we create from already collected information or we also have external sources of information about the client. Some of the dates might be in future to the respect of application date, but they are collected at the date of application, so there is no target leakage. \n\nA made-up example could be a column that tells you when is your next birthday going to be. It is generated at the date of application when the default - target is not known, so there is no information about the future target indicated by the date of next birthday.\n\nHope it is clear now.",
    "2803336": "Thanks Daniel."
  },
  "source": "meta"
}