{
  "id": 491556,
  "title": "dtypes of Date columns using light LGB",
  "url": "/competitions/home-credit-credit-risk-model-stability/discussion/491556",
  "author_name": "",
  "post_date": "2024-04-06T09:38:34.553240800Z",
  "votes": 1,
  "comment_count": 6,
  "views": 0,
  "content": "<p>Hi Everyone, </p>\n<p>I copied the notebook <a href=\"https://www.kaggle.com/code/daviddirethucus/home-credit-risk-lightgbm\" target=\"_blank\">https://www.kaggle.com/code/daviddirethucus/home-credit-risk-lightgbm</a>     to practice, after run <br>\ndf_train = feature_eng(**data_store)<br>\nprint(\"train data shape:\\t\", df_train.shape)<br>\nI noticed that the date columns ('D')dtypes are not change into pl.Date.  Here is my run results (output 14):<a href=\"https://www.kaggle.com/code/yanzhang743/home-credit-risk-lightgbm-f74c38?scriptVersionId=170613289\" target=\"_blank\">https://www.kaggle.com/code/yanzhang743/home-credit-risk-lightgbm-f74c38?scriptVersionId=170613289</a>. <br>\nany suggestions to fix it?</p>",
  "messages": [
    {
      "id": "2738357",
      "postDate": "04/06/2024 09:38:34",
      "content": "<p>Hi Everyone, </p>\n<p>I copied the notebook <a href=\"https://www.kaggle.com/code/daviddirethucus/home-credit-risk-lightgbm\" target=\"_blank\">https://www.kaggle.com/code/daviddirethucus/home-credit-risk-lightgbm</a>     to practice, after run <br>\ndf_train = feature_eng(**data_store)<br>\nprint(\"train data shape:\\t\", df_train.shape)<br>\nI noticed that the date columns ('D')dtypes are not change into pl.Date.  Here is my run results (output 14):<a href=\"https://www.kaggle.com/code/yanzhang743/home-credit-risk-lightgbm-f74c38?scriptVersionId=170613289\" target=\"_blank\">https://www.kaggle.com/code/yanzhang743/home-credit-risk-lightgbm-f74c38?scriptVersionId=170613289</a>. <br>\nany suggestions to fix it?</p>",
      "rawMarkdown": "Hi Everyone, \n\nI copied the notebook https://www.kaggle.com/code/daviddirethucus/home-credit-risk-lightgbm     to practice, after run \ndf_train = feature_eng(**data_store)\nprint(\"train data shape:\\t\", df_train.shape)\nI noticed that the date columns ('D')dtypes are not change into pl.Date.  Here is my run results (output 14):https://www.kaggle.com/code/yanzhang743/home-credit-risk-lightgbm-f74c38?scriptVersionId=170613289. \nany suggestions to fix it?",
      "votes": null
    },
    {
      "id": "2738388",
      "postDate": "04/06/2024 10:11:19",
      "content": "<p>In the above-mentioned notebook, there is a class Pipeline. in there, there is a function set_table_dtypes. In there, there is a condition that checks if the last letter of the column is 'D' than assign Date dtype to it.</p>",
      "rawMarkdown": "In the above-mentioned notebook, there is a class Pipeline. in there, there is a function set_table_dtypes. In there, there is a condition that checks if the last letter of the column is 'D' than assign Date dtype to it.",
      "votes": null
    },
    {
      "id": "2738452",
      "postDate": "04/06/2024 10:48:29",
      "content": "<p>Thank you! It didn’t work. As you see in data_store step, the read_file will handle the function set _table_dtypes. But in output 14, the dtype  “D” columns didn’t show the right data types.</p>",
      "rawMarkdown": "Thank you! It didn’t work. As you see in data_store step, the read_file will handle the function set _table_dtypes. But in output 14, the dtype  “D” columns didn’t show the right data types.",
      "votes": null
    },
    {
      "id": "2738802",
      "postDate": "04/06/2024 16:09:37",
      "content": "<p>I am nable to open the second notebook, is it public or private?</p>",
      "rawMarkdown": "I am nable to open the second notebook, is it public or private?",
      "votes": null
    },
    {
      "id": "2738822",
      "postDate": "04/06/2024 16:28:17",
      "content": "<p>Please check it now, it it public.</p>",
      "rawMarkdown": "Please check it now, it it public.",
      "votes": null
    },
    {
      "id": "2739192",
      "postDate": "04/06/2024 21:56:30",
      "content": "<p>In feature_eng function, Pipeline.handle_dates is called. In that handle_dates function every column with date tpe is substracted from date of decision. so now it no longer date but a period. then dt.total_days() extracts the number of days in that period. making it a int type. thus the final col that is present has been transformed from date to the int type.</p>",
      "rawMarkdown": "In feature_eng function, Pipeline.handle_dates is called. In that handle_dates function every column with date tpe is substracted from date of decision. so now it no longer date but a period. then dt.total_days() extracts the number of days in that period. making it a int type. thus the final col that is present has been transformed from date to the int type.",
      "votes": null
    },
    {
      "id": "2745080",
      "postDate": "04/10/2024 11:03:43",
      "content": "<p>Thank you!</p>",
      "rawMarkdown": "Thank you!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2738388,
      "author_name": "shreyas9181",
      "author_url": "",
      "post_date": "04/06/2024 10:11:19",
      "content": "<p>In the above-mentioned notebook, there is a class Pipeline. in there, there is a function set_table_dtypes. In there, there is a condition that checks if the last letter of the column is 'D' than assign Date dtype to it.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2738452,
          "author_name": "yanzhang743",
          "author_url": "",
          "post_date": "04/06/2024 10:48:29",
          "content": "<p>Thank you! It didn’t work. As you see in data_store step, the read_file will handle the function set _table_dtypes. But in output 14, the dtype  “D” columns didn’t show the right data types.</p>",
          "votes": null,
          "replies": [
            {
              "id": 2738802,
              "author_name": "shreyas9181",
              "author_url": "",
              "post_date": "04/06/2024 16:09:37",
              "content": "<p>I am nable to open the second notebook, is it public or private?</p>",
              "votes": null,
              "replies": [
                {
                  "id": 2738822,
                  "author_name": "yanzhang743",
                  "author_url": "",
                  "post_date": "04/06/2024 16:28:17",
                  "content": "<p>Please check it now, it it public.</p>",
                  "votes": null,
                  "replies": [
                    {
                      "id": 2739192,
                      "author_name": "shreyas9181",
                      "author_url": "",
                      "post_date": "04/06/2024 21:56:30",
                      "content": "<p>In feature_eng function, Pipeline.handle_dates is called. In that handle_dates function every column with date tpe is substracted from date of decision. so now it no longer date but a period. then dt.total_days() extracts the number of days in that period. making it a int type. thus the final col that is present has been transformed from date to the int type.</p>",
                      "votes": null,
                      "replies": [
                        {
                          "id": 2745080,
                          "author_name": "yanzhang743",
                          "author_url": "",
                          "post_date": "04/10/2024 11:03:43",
                          "content": "<p>Thank you!</p>",
                          "votes": null,
                          "replies": []
                        }
                      ]
                    }
                  ]
                }
              ]
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2738357": "Hi Everyone, \n\nI copied the notebook https://www.kaggle.com/code/daviddirethucus/home-credit-risk-lightgbm     to practice, after run \ndf_train = feature_eng(**data_store)\nprint(\"train data shape:\\t\", df_train.shape)\nI noticed that the date columns ('D')dtypes are not change into pl.Date.  Here is my run results (output 14):https://www.kaggle.com/code/yanzhang743/home-credit-risk-lightgbm-f74c38?scriptVersionId=170613289. \nany suggestions to fix it?",
    "2738388": "In the above-mentioned notebook, there is a class Pipeline. in there, there is a function set_table_dtypes. In there, there is a condition that checks if the last letter of the column is 'D' than assign Date dtype to it.",
    "2738452": "Thank you! It didn’t work. As you see in data_store step, the read_file will handle the function set _table_dtypes. But in output 14, the dtype  “D” columns didn’t show the right data types.",
    "2738802": "I am nable to open the second notebook, is it public or private?",
    "2738822": "Please check it now, it it public.",
    "2739192": "In feature_eng function, Pipeline.handle_dates is called. In that handle_dates function every column with date tpe is substracted from date of decision. so now it no longer date but a period. then dt.total_days() extracts the number of days in that period. making it a int type. thus the final col that is present has been transformed from date to the int type.",
    "2745080": "Thank you!"
  },
  "source": "meta"
}