{
  "id": 472866,
  "title": "Welcome note from Home Credit",
  "url": "/competitions/home-credit-credit-risk-model-stability/discussion/472866",
  "author_name": "Tomas Jelinek",
  "post_date": "2024-02-02T12:46:54.961000",
  "votes": 55,
  "comment_count": 40,
  "views": 0,
  "content": "<p>Dear Kagglers,</p>\n<p>Home Credit’s back! After a 5-year hiatus, I’m delighted to welcome you to another of our data-packed competitions “Home Credit: Credit Risk Model Stability”, where the winning teams split a hefty jackpot and their solutions inspire our lending approval process. </p>\n<p>About the contest, many things remain same as in our previous competition \"Home Credit Default Risk\": the goal is to assess the credit risk quality of clients and predict their future payment behavior. We’ve provided data from a wide range of sources: application forms, social-demographic data, previous credit behavior data, etc. We’ve endeavored to keep the data in a raw form, and where necessary (e.g., with personal data, business sensitive information) we’ve applied data masking instead of aggregations to retain as much information value as possible. We’ve even kept effects of business decisions in the dataset intact, for example, a change in usage of certain data sources, (please read the data description carefully, so you’re aware of all potential data pitfalls). I can honestly state that the data sample presented mirrors almost exactly the datasets we use for real machine learning problems in Home Credit.</p>\n<p>Though some aspects of this contest are the same as before, this is no rehashing of our previous competition—we’ve got lots of new things to share.   Two of the most important are the vast size of the dataset and the evaluation metric. As to the first, we’ve decided to provide you with data on the majority of our portfolio so you will have enough ‘food’ even for your most data hungry algorithms and ML approaches. Measured by raw size, the current sample is roughly 10 times bigger than the previous one. At the same time, we’ve tried to keep the structure of data sample simple and readable. As to the second, the evaluation metric, though we still use AUC (Gini coefficient) to measure the accuracy of model predictions, we’ve added additional criteria related to model stability.  We want to see solutions that produce sound predictions also far in the future since this is a fundamental business requirement. There’s significant delay in target observations and limited capacities of R&amp;D teams, so there’s a strong preference for models that last over a long period of time. We’re very keen to see how you’ll handle and incorporate this requirement into your solutions.</p>\n<p>Last but not least, as mentioned above, we’ve significantly boosted the prize money and also increased the number of winning places that can claim a cut in the cash!</p>\n<p>If you have any questions regarding the data, competition, R&amp;D team or our company, please feel free to ask me and my colleague Daniel Herman.  We’re here to answer all your inquiries and support you as best we can.</p>\n<p>Good luck; and we’re looking forward to hearing from you in the discussion.</p>\n<p>On behalf of the entire Home Credit R&amp;D team,</p>\n<p>Tomas Jelinek</p>",
  "messages": [
    {
      "id": 2632545,
      "postDate": "2024-02-02T12:46:54.963Z",
      "content": "<p>Dear Kagglers,</p>\n<p>Home Credit’s back! After a 5-year hiatus, I’m delighted to welcome you to another of our data-packed competitions “Home Credit: Credit Risk Model Stability”, where the winning teams split a hefty jackpot and their solutions inspire our lending approval process. </p>\n<p>About the contest, many things remain same as in our previous competition \"Home Credit Default Risk\": the goal is to assess the credit risk quality of clients and predict their future payment behavior. We’ve provided data from a wide range of sources: application forms, social-demographic data, previous credit behavior data, etc. We’ve endeavored to keep the data in a raw form, and where necessary (e.g., with personal data, business sensitive information) we’ve applied data masking instead of aggregations to retain as much information value as possible. We’ve even kept effects of business decisions in the dataset intact, for example, a change in usage of certain data sources, (please read the data description carefully, so you’re aware of all potential data pitfalls). I can honestly state that the data sample presented mirrors almost exactly the datasets we use for real machine learning problems in Home Credit.</p>\n<p>Though some aspects of this contest are the same as before, this is no rehashing of our previous competition—we’ve got lots of new things to share.   Two of the most important are the vast size of the dataset and the evaluation metric. As to the first, we’ve decided to provide you with data on the majority of our portfolio so you will have enough ‘food’ even for your most data hungry algorithms and ML approaches. Measured by raw size, the current sample is roughly 10 times bigger than the previous one. At the same time, we’ve tried to keep the structure of data sample simple and readable. As to the second, the evaluation metric, though we still use AUC (Gini coefficient) to measure the accuracy of model predictions, we’ve added additional criteria related to model stability.  We want to see solutions that produce sound predictions also far in the future since this is a fundamental business requirement. There’s significant delay in target observations and limited capacities of R&amp;D teams, so there’s a strong preference for models that last over a long period of time. We’re very keen to see how you’ll handle and incorporate this requirement into your solutions.</p>\n<p>Last but not least, as mentioned above, we’ve significantly boosted the prize money and also increased the number of winning places that can claim a cut in the cash!</p>\n<p>If you have any questions regarding the data, competition, R&amp;D team or our company, please feel free to ask me and my colleague Daniel Herman.  We’re here to answer all your inquiries and support you as best we can.</p>\n<p>Good luck; and we’re looking forward to hearing from you in the discussion.</p>\n<p>On behalf of the entire Home Credit R&amp;D team,</p>\n<p>Tomas Jelinek</p>",
      "rawMarkdown": "Dear Kagglers,\n\nHome Credit’s back! After a 5-year hiatus, I’m delighted to welcome you to another of our data-packed competitions “Home Credit: Credit Risk Model Stability”, where the winning teams split a hefty jackpot and their solutions inspire our lending approval process. \n\nAbout the contest, many things remain same as in our previous competition \"Home Credit Default Risk\": the goal is to assess the credit risk quality of clients and predict their future payment behavior. We’ve provided data from a wide range of sources: application forms, social-demographic data, previous credit behavior data, etc. We’ve endeavored to keep the data in a raw form, and where necessary (e.g., with personal data, business sensitive information) we’ve applied data masking instead of aggregations to retain as much information value as possible. We’ve even kept effects of business decisions in the dataset intact, for example, a change in usage of certain data sources, (please read the data description carefully, so you’re aware of all potential data pitfalls). I can honestly state that the data sample presented mirrors almost exactly the datasets we use for real machine learning problems in Home Credit.\n \nThough some aspects of this contest are the same as before, this is no rehashing of our previous competition—we’ve got lots of new things to share.   Two of the most important are the vast size of the dataset and the evaluation metric. As to the first, we’ve decided to provide you with data on the majority of our portfolio so you will have enough ‘food’ even for your most data hungry algorithms and ML approaches. Measured by raw size, the current sample is roughly 10 times bigger than the previous one. At the same time, we’ve tried to keep the structure of data sample simple and readable. As to the second, the evaluation metric, though we still use AUC (Gini coefficient) to measure the accuracy of model predictions, we’ve added additional criteria related to model stability.  We want to see solutions that produce sound predictions also far in the future since this is a fundamental business requirement. There’s significant delay in target observations and limited capacities of R&D teams, so there’s a strong preference for models that last over a long period of time. We’re very keen to see how you’ll handle and incorporate this requirement into your solutions.\n \nLast but not least, as mentioned above, we’ve significantly boosted the prize money and also increased the number of winning places that can claim a cut in the cash!\n \nIf you have any questions regarding the data, competition, R&D team or our company, please feel free to ask me and my colleague Daniel Herman.  We’re here to answer all your inquiries and support you as best we can.\n \nGood luck; and we’re looking forward to hearing from you in the discussion.\n \nOn behalf of the entire Home Credit R&D team,\n \nTomas Jelinek",
      "votes": 54
    },
    {
      "id": 2637309,
      "postDate": "2024-02-05T16:34:45Z",
      "content": "<p>I am looking forward to seeing all the creative solutions! Best of luck to everyone joining us this year! 🚀🚀🚀</p>",
      "rawMarkdown": "I am looking forward to seeing all the creative solutions! Best of luck to everyone joining us this year! 🚀🚀🚀",
      "votes": 20,
      "replies": [
        {
          "id": 2637364,
          "postDate": "2024-02-05T17:06:31.010Z",
          "content": "<p>So excited to see the second round!!</p>",
          "rawMarkdown": "So excited to see the second round!!",
          "votes": 4
        },
        {
          "id": 2648537,
          "postDate": "2024-02-12T09:32:26.413Z",
          "content": "<p>I'm so happy to have found this competition!</p>",
          "rawMarkdown": "I'm so happy to have found this competition!",
          "votes": 3
        }
      ]
    },
    {
      "id": 2644081,
      "postDate": "2024-02-09T09:22:08.807Z",
      "content": "<p>Thank you for organizing this competition! I wonder how large is the test set for the private leaderboard? </p>\n<p>The data section says \"test_base.csv contains approximately 90% of the numbers of case_id values of train_base.csv\". I assume this refers to the size of test set for the public leaderboard. How about the size of test set for the private leaderboard? Will it be twice the size of the public leaderboard test set?</p>\n<p>Thank you!</p>",
      "rawMarkdown": "Thank you for organizing this competition! I wonder how large is the test set for the private leaderboard? \n\nThe data section says \"test_base.csv contains approximately 90% of the numbers of case_id values of train_base.csv\". I assume this refers to the size of test set for the public leaderboard. How about the size of test set for the private leaderboard? Will it be twice the size of the public leaderboard test set?\n\nThank you!",
      "votes": 3
    },
    {
      "id": 2638001,
      "postDate": "2024-02-06T03:44:14.467Z",
      "content": "<p>Can we have a public notebook for metric calculation script? Can we use external data source?</p>",
      "rawMarkdown": "Can we have a public notebook for metric calculation script? Can we use external data source?",
      "votes": 3,
      "replies": [
        {
          "id": 2638256,
          "postDate": "2024-02-06T07:06:56.943Z",
          "content": "<p>Metric calculation script is already included in the <a href=\"https://www.kaggle.com/code/jetakow/home-credit-2024-starter-notebook\" target=\"_blank\">notebook</a> I prepared for you. Regarding external data source see 7. COMPETITION DATA. - C. External Data. </p>\n<blockquote>\n  <p>You may use data other than the Competition Data (“External Data”) to develop and test your Submissions. However, you will ensure the External Data is publicly available and equally accessible to use by all participants of the Competition for purposes of the competition at no cost to the other participants. The ability to use External Data under this Section 7.C (External Data) does not limit your other obligations under these Competition Rules, including but not limited to Section 11 (Winners Obligations).</p>\n</blockquote>",
          "rawMarkdown": "Metric calculation script is already included in the [notebook](https://www.kaggle.com/code/jetakow/home-credit-2024-starter-notebook) I prepared for you. Regarding external data source see 7. COMPETITION DATA. - C. External Data. \n\n>You may use data other than the Competition Data (“External Data”) to develop and test your Submissions. However, you will ensure the External Data is publicly available and equally accessible to use by all participants of the Competition for purposes of the competition at no cost to the other participants. The ability to use External Data under this Section 7.C (External Data) does not limit your other obligations under these Competition Rules, including but not limited to Section 11 (Winners Obligations).",
          "votes": 6,
          "replies": [
            {
              "id": 2638336,
              "postDate": "2024-02-06T08:21:14.803Z",
              "content": "<p>It seems that test dataset is not from the future. that's why I asked about external dataset. There might be winner solutions using look ahead bias. Is this type of solutions acceptable? It seems to be impossible to prevent look ahead bias. </p>",
              "rawMarkdown": "It seems that test dataset is not from the future. that's why I asked about external dataset. There might be winner solutions using look ahead bias. Is this type of solutions acceptable? It seems to be impossible to prevent look ahead bias. ",
              "votes": 2
            },
            {
              "id": 2638777,
              "postDate": "2024-02-06T13:29:28.740Z",
              "content": "<p>Hi,<br>\nas Daniel mentioned, external data are allowed and conditions are described in the Rules.<br>\nRegarding look ahead bias, we had similar rules for 1st Home Credit competition and it was not a problem. We will not disclose details about time period of test sample, so Kagglers will not have direct way how to match external data to correct time period. </p>",
              "rawMarkdown": "Hi,\nas Daniel mentioned, external data are allowed and conditions are described in the Rules.\nRegarding look ahead bias, we had similar rules for 1st Home Credit competition and it was not a problem. We will not disclose details about time period of test sample, so Kagglers will not have direct way how to match external data to correct time period. ",
              "votes": 5
            },
            {
              "id": 2640447,
              "postDate": "2024-02-07T00:15:49.150Z",
              "content": "<blockquote>\n  <p>We will not disclose details about time period of test sample, so Kagglers will not have direct way how to match external data to correct time period.</p>\n</blockquote>\n<p>Nice. And I'm hoping the test set includes very recent data for preventing it.</p>",
              "rawMarkdown": ">We will not disclose details about time period of test sample, so Kagglers will not have direct way how to match external data to correct time period.\n\nNice. And I'm hoping the test set includes very recent data for preventing it.",
              "votes": 2
            }
          ]
        }
      ]
    },
    {
      "id": 2826427,
      "postDate": "2024-05-20T23:01:37.457Z",
      "content": "<p>I am getting this \"TypeError: the truth value of a Series is ambiguous\" when trying to use the Polars library for data cleaning, anyone has idea on how to deal with this. Thanks</p>",
      "rawMarkdown": "I am getting this \"TypeError: the truth value of a Series is ambiguous\" when trying to use the Polars library for data cleaning, anyone has idea on how to deal with this. Thanks",
      "votes": 1
    },
    {
      "id": 2842179,
      "postDate": "2024-05-28T23:35:24.517Z",
      "content": "<p>thank you for organizing this competition. I had a good time doing this.</p>",
      "rawMarkdown": "thank you for organizing this competition. I had a good time doing this.",
      "votes": 2
    },
    {
      "id": 2656081,
      "postDate": "2024-02-17T12:47:00.117Z",
      "content": "<p>Thank you so much for organizing such an amazing competition. I'm trying to understand the data and I don't exactly understand what NumGroup1,2 means. </p>\n<p>The data in Depth 1 is easy to understand because it only has the column \"num_group1\",<br>\nbut the data in Depth 2 has \"num_group1\" , \"num_group2\", which is confusing.  Does having more than one type of num_group in Depth 2 mean that there are multiple branches within a case?</p>",
      "rawMarkdown": "Thank you so much for organizing such an amazing competition. I'm trying to understand the data and I don't exactly understand what NumGroup1,2 means. \n\nThe data in Depth 1 is easy to understand because it only has the column \"num_group1\",\nbut the data in Depth 2 has \"num_group1\" , \"num_group2\", which is confusing.  Does having more than one type of num_group in Depth 2 mean that there are multiple branches within a case?",
      "votes": 1
    },
    {
      "id": 2649736,
      "postDate": "2024-02-13T04:34:09.887Z",
      "content": "<p>Hi, <br>\nI have a question about the date. Can you explain why 'recorddate_4527225D' in tax_registry_a is always larger than 'date_decision' in base table? </p>\n<p>case_id    amount_4527230A name_4527232M   num_group1  recorddate_4527225D date_decision   MONTH   WEEK_NUM    target<br>\n28631    1946.0  \"f980a1ea\"  2   \"2019-09-13\"    \"2019-08-30\"    201908  34  0<br>\n28631    711.0   \"f980a1ea\"  3   \"2019-09-13\"    \"2019-08-30\"    201908  34  0</p>",
      "rawMarkdown": "Hi, \nI have a question about the date. Can you explain why 'recorddate_4527225D' in tax_registry_a is always larger than 'date_decision' in base table? \n\ncase_id\tamount_4527230A\tname_4527232M\tnum_group1\trecorddate_4527225D\tdate_decision\tMONTH\tWEEK_NUM\ttarget\n28631\t1946.0\t\"f980a1ea\"\t2\t\"2019-09-13\"\t\"2019-08-30\"\t201908\t34\t0\n28631\t711.0\t\"f980a1ea\"\t3\t\"2019-09-13\"\t\"2019-08-30\"\t201908\t34\t0",
      "votes": 1,
      "replies": [
        {
          "id": 2650325,
          "postDate": "2024-02-13T11:46:49.350Z",
          "content": "<p>Hello, for the column recorddate_4527225D we have the definition \"Date of tax deduction record.\". This date is from the future as it should be. All date columns were transformed. </p>",
          "rawMarkdown": "Hello, for the column recorddate_4527225D we have the definition \"Date of tax deduction record.\". This date is from the future as it should be. All date columns were transformed. ",
          "votes": 1,
          "replies": [
            {
              "id": 2651136,
              "postDate": "2024-02-14T01:16:16.710Z",
              "content": "<p>Does the prediction of loan default happen before the decision date or the first due date of the loan? Besides, if the internal/external data after the decision date is also available for model prediction, will there be an information leak? </p>",
              "rawMarkdown": "Does the prediction of loan default happen before the decision date or the first due date of the loan? Besides, if the internal/external data after the decision date is also available for model prediction, will there be an information leak? ",
              "votes": 1
            },
            {
              "id": 2651538,
              "postDate": "2024-02-14T08:27:45.747Z",
              "content": "<p>Hi Sophie,<br>\nprediction is done at the time of application (decision date). All data are collected as of that date (including external data).</p>",
              "rawMarkdown": "Hi Sophie,\nprediction is done at the time of application (decision date). All data are collected as of that date (including external data).",
              "votes": 3
            }
          ]
        }
      ]
    },
    {
      "id": 2643157,
      "postDate": "2024-02-08T16:51:25.217Z",
      "content": "<p>Can you explain the content of person_1 file? What is the meaning of role_1084L</p>",
      "rawMarkdown": "Can you explain the content of person_1 file? What is the meaning of role_1084L",
      "votes": 1,
      "replies": [
        {
          "id": 2643976,
          "postDate": "2024-02-09T07:55:39.317Z",
          "content": "<p>Each credit application can have information about several persons (e.g. client him/her-self, contact references). Role describe type of connection to client.</p>",
          "rawMarkdown": "Each credit application can have information about several persons (e.g. client him/her-self, contact references). Role describe type of connection to client.",
          "votes": 2,
          "replies": [
            {
              "id": 2643981,
              "postDate": "2024-02-09T08:03:43.563Z",
              "content": "<p>Thank you. Can you explain the meaning of the categories in role_1084L</p>",
              "rawMarkdown": "Thank you. Can you explain the meaning of the categories in role_1084L",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 2641740,
      "postDate": "2024-02-07T16:33:19.477Z",
      "content": "<p>hi,</p>\n<p>in columns 'credtype_587L', what is the abbreviation of : </p>\n<p>COL, CAL, REL ?</p>",
      "rawMarkdown": "hi,\n\nin columns 'credtype_587L', what is the abbreviation of : \n\nCOL, CAL, REL ?",
      "votes": 1,
      "replies": [
        {
          "id": 2642543,
          "postDate": "2024-02-08T08:22:27.567Z",
          "content": "<p>Hello. Those stand for different types of loans.</p>",
          "rawMarkdown": "Hello. Those stand for different types of loans.",
          "votes": 1
        }
      ]
    },
    {
      "id": 2641584,
      "postDate": "2024-02-07T15:08:31.297Z",
      "content": "<p>Hello! Quick question: since this is an approvals model, I know there are some regulation issues that would prevent data scientists from using less interpretable models (although more predictive). Should we use the same principles here?</p>",
      "rawMarkdown": "Hello! Quick question: since this is an approvals model, I know there are some regulation issues that would prevent data scientists from using less interpretable models (although more predictive). Should we use the same principles here?",
      "votes": 1,
      "replies": [
        {
          "id": 2641597,
          "postDate": "2024-02-07T15:11:46.300Z",
          "content": "<p>Hello! We don't prevent you from using any black box model you can think of. The only limitations are the computational resources and 12 hours for one submission. </p>",
          "rawMarkdown": "Hello! We don't prevent you from using any black box model you can think of. The only limitations are the computational resources and 12 hours for one submission. ",
          "votes": 1
        }
      ]
    },
    {
      "id": 2637330,
      "postDate": "2024-02-05T16:47:50.073Z",
      "content": "<p>heyy all and good luck!</p>",
      "rawMarkdown": "heyy all and good luck!",
      "votes": 1,
      "replies": [
        {
          "id": 2637362,
          "postDate": "2024-02-05T17:06:16.207Z",
          "rawMarkdown": "",
          "isDeleted": true
        }
      ]
    },
    {
      "id": 2650845,
      "postDate": "2024-02-13T18:11:15.127Z",
      "content": "<p>Any advice for out of memory issues? My notebook runs fine but just during the scoring it throws the error</p>",
      "rawMarkdown": "Any advice for out of memory issues? My notebook runs fine but just during the scoring it throws the error",
      "votes": 2
    },
    {
      "id": 2837995,
      "postDate": "2024-05-26T18:36:17.987Z",
      "content": "<p>Do I need to use all the datasets provided during the model implementation?</p>",
      "rawMarkdown": "Do I need to use all the datasets provided during the model implementation?"
    },
    {
      "id": 2799042,
      "postDate": "2024-05-07T15:11:15.897Z",
      "content": "<p>I didn't have the chance of compete , but I think It was a grate challenge </p>",
      "rawMarkdown": "I didn't have the chance of compete , but I think It was a grate challenge "
    },
    {
      "id": 2749282,
      "postDate": "2024-04-13T01:29:09.760Z",
      "content": "<p><a href=\"https://www.kaggle.com/jetakow\" target=\"_blank\">@jetakow</a> -  I am a high school junior passionate about Machine Learning. I noticed that there is a minimum age (18 years old) limit to enter this competition. We have already submitted the parental consent form to enter Kaggle competitions. Is there any additional parental consent form to enter this competition? Thanks!</p>",
      "rawMarkdown": "@jetakow -  I am a high school junior passionate about Machine Learning. I noticed that there is a minimum age (18 years old) limit to enter this competition. We have already submitted the parental consent form to enter Kaggle competitions. Is there any additional parental consent form to enter this competition? Thanks!",
      "replies": [
        {
          "id": 2752258,
          "postDate": "2024-04-14T20:20:33.510Z",
          "content": "<p>Please create a new post in the Discussion for this topic. Thanks.</p>",
          "rawMarkdown": "Please create a new post in the Discussion for this topic. Thanks."
        },
        {
          "id": 2752931,
          "postDate": "2024-04-15T08:18:56.530Z",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/darshjaganmohan\" target=\"_blank\">@darshjaganmohan</a>,<br>\nI think this is question to Kaggle staff, please try to contact them. Perhaps <a href=\"https://www.kaggle.com/addisonhoward\" target=\"_blank\">@addisonhoward</a> can help?</p>",
          "rawMarkdown": "Hi @darshjaganmohan,\nI think this is question to Kaggle staff, please try to contact them. Perhaps @addisonhoward can help?"
        },
        {
          "id": 2753498,
          "postDate": "2024-04-15T14:24:29.847Z",
          "content": "<p>Hi Darsh! Welcome to Kaggle. There are two consent forms to submit. One for <a href=\"https://www.kaggle.com/guardian-consent-minor-use\" target=\"_blank\">using Kaggle</a> and one for entering <a href=\"https://www.kaggle.com/consent-minors-process\" target=\"_blank\">each Kaggle Competition</a>. Please ensure you are compliant with local laws regarding participation as a minor as it applies to you.</p>",
          "rawMarkdown": "Hi Darsh! Welcome to Kaggle. There are two consent forms to submit. One for [using Kaggle](https://www.kaggle.com/guardian-consent-minor-use) and one for entering [each Kaggle Competition](https://www.kaggle.com/consent-minors-process). Please ensure you are compliant with local laws regarding participation as a minor as it applies to you.",
          "votes": 1,
          "replies": [
            {
              "id": 2754029,
              "postDate": "2024-04-15T18:58:36.847Z",
              "content": "<p>Thank you everyone for the responses.</p>",
              "rawMarkdown": "Thank you everyone for the responses."
            }
          ]
        }
      ]
    },
    {
      "id": 2686980,
      "postDate": "2024-03-08T07:26:51.127Z",
      "content": "<p>Thank you for organizing this competition. </p>",
      "rawMarkdown": "Thank you for organizing this competition. "
    },
    {
      "id": 2668627,
      "postDate": "2024-02-25T19:05:49.730Z",
      "content": "<p>Hello All, I need advise how to submit solutions for this competition. I tried everything but I still receive \"Cannot submit<br>\nSubmissions have been disabled for this competition.\" Any help is appreciated.</p>",
      "rawMarkdown": "Hello All, I need advise how to submit solutions for this competition. I tried everything but I still receive \"Cannot submit\nSubmissions have been disabled for this competition.\" Any help is appreciated.",
      "replies": [
        {
          "id": 2669201,
          "postDate": "2024-02-26T06:21:36.630Z",
          "content": "<p>Submitting has been paused while they workout issues with the scoring metric.</p>",
          "rawMarkdown": "Submitting has been paused while they workout issues with the scoring metric."
        }
      ]
    },
    {
      "id": 2810337,
      "postDate": "2024-05-13T08:17:26.727Z",
      "content": "<p>thanks for this funny comp</p>",
      "rawMarkdown": "thanks for this funny comp"
    },
    {
      "id": 2796369,
      "postDate": "2024-05-06T07:39:28.607Z",
      "content": "<p>Thanks for sharing</p>",
      "rawMarkdown": "Thanks for sharing"
    },
    {
      "id": 2774038,
      "postDate": "2024-04-25T03:24:02.827Z",
      "content": "<p>thanks a lot!</p>",
      "rawMarkdown": "thanks a lot!"
    },
    {
      "id": 2757219,
      "postDate": "2024-04-17T11:54:44.647Z",
      "content": "<p>Thank you!</p>",
      "rawMarkdown": "Thank you!"
    }
  ],
  "comments": [
    {
      "id": 2637309,
      "author_name": "Daniel Herman",
      "author_url": "",
      "post_date": "2024-02-05T16:34:45",
      "content": "<p>I am looking forward to seeing all the creative solutions! Best of luck to everyone joining us this year! 🚀🚀🚀</p>",
      "votes": 20,
      "replies": [
        {
          "id": 2637364,
          "author_name": "narsil (jobs-in-data.com)",
          "author_url": "",
          "post_date": "2024-02-05T17:06:31.010000",
          "content": "<p>So excited to see the second round!!</p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 2648537,
          "author_name": "Doraking",
          "author_url": "",
          "post_date": "2024-02-12T09:32:26.413000",
          "content": "<p>I'm so happy to have found this competition!</p>",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 2644081,
      "author_name": "WindClimber",
      "author_url": "",
      "post_date": "2024-02-09T09:22:08.807000",
      "content": "<p>Thank you for organizing this competition! I wonder how large is the test set for the private leaderboard? </p>\n<p>The data section says \"test_base.csv contains approximately 90% of the numbers of case_id values of train_base.csv\". I assume this refers to the size of test set for the public leaderboard. How about the size of test set for the private leaderboard? Will it be twice the size of the public leaderboard test set?</p>\n<p>Thank you!</p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 2638001,
      "author_name": "yuanzhe zhou",
      "author_url": "",
      "post_date": "2024-02-06T03:44:14.467000",
      "content": "<p>Can we have a public notebook for metric calculation script? Can we use external data source?</p>",
      "votes": 3,
      "replies": [
        {
          "id": 2638256,
          "author_name": "Daniel Herman",
          "author_url": "",
          "post_date": "2024-02-06T07:06:56.943000",
          "content": "<p>Metric calculation script is already included in the <a href=\"https://www.kaggle.com/code/jetakow/home-credit-2024-starter-notebook\" target=\"_blank\">notebook</a> I prepared for you. Regarding external data source see 7. COMPETITION DATA. - C. External Data. </p>\n<blockquote>\n  <p>You may use data other than the Competition Data (“External Data”) to develop and test your Submissions. However, you will ensure the External Data is publicly available and equally accessible to use by all participants of the Competition for purposes of the competition at no cost to the other participants. The ability to use External Data under this Section 7.C (External Data) does not limit your other obligations under these Competition Rules, including but not limited to Section 11 (Winners Obligations).</p>\n</blockquote>",
          "votes": 6,
          "replies": [
            {
              "id": 2638336,
              "author_name": "yuanzhe zhou",
              "author_url": "",
              "post_date": "2024-02-06T08:21:14.803000",
              "content": "<p>It seems that test dataset is not from the future. that's why I asked about external dataset. There might be winner solutions using look ahead bias. Is this type of solutions acceptable? It seems to be impossible to prevent look ahead bias. </p>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 2638777,
              "author_name": "Tomas Jelinek",
              "author_url": "",
              "post_date": "2024-02-06T13:29:28.740000",
              "content": "<p>Hi,<br>\nas Daniel mentioned, external data are allowed and conditions are described in the Rules.<br>\nRegarding look ahead bias, we had similar rules for 1st Home Credit competition and it was not a problem. We will not disclose details about time period of test sample, so Kagglers will not have direct way how to match external data to correct time period. </p>",
              "votes": 5,
              "replies": []
            },
            {
              "id": 2640447,
              "author_name": "DongYK",
              "author_url": "",
              "post_date": "2024-02-07T00:15:49.150000",
              "content": "<blockquote>\n  <p>We will not disclose details about time period of test sample, so Kagglers will not have direct way how to match external data to correct time period.</p>\n</blockquote>\n<p>Nice. And I'm hoping the test set includes very recent data for preventing it.</p>",
              "votes": 2,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2826427,
      "author_name": "Abdullahi Dahir",
      "author_url": "",
      "post_date": "2024-05-20T23:01:37.457000",
      "content": "<p>I am getting this \"TypeError: the truth value of a Series is ambiguous\" when trying to use the Polars library for data cleaning, anyone has idea on how to deal with this. Thanks</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2842179,
      "author_name": "Tomoki Takata",
      "author_url": "",
      "post_date": "2024-05-28T23:35:24.517000",
      "content": "<p>thank you for organizing this competition. I had a good time doing this.</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 2656081,
      "author_name": "jerem Kim",
      "author_url": "",
      "post_date": "2024-02-17T12:47:00.117000",
      "content": "<p>Thank you so much for organizing such an amazing competition. I'm trying to understand the data and I don't exactly understand what NumGroup1,2 means. </p>\n<p>The data in Depth 1 is easy to understand because it only has the column \"num_group1\",<br>\nbut the data in Depth 2 has \"num_group1\" , \"num_group2\", which is confusing.  Does having more than one type of num_group in Depth 2 mean that there are multiple branches within a case?</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2649736,
      "author_name": "Sophie",
      "author_url": "",
      "post_date": "2024-02-13T04:34:09.887000",
      "content": "<p>Hi, <br>\nI have a question about the date. Can you explain why 'recorddate_4527225D' in tax_registry_a is always larger than 'date_decision' in base table? </p>\n<p>case_id    amount_4527230A name_4527232M   num_group1  recorddate_4527225D date_decision   MONTH   WEEK_NUM    target<br>\n28631    1946.0  \"f980a1ea\"  2   \"2019-09-13\"    \"2019-08-30\"    201908  34  0<br>\n28631    711.0   \"f980a1ea\"  3   \"2019-09-13\"    \"2019-08-30\"    201908  34  0</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2650325,
          "author_name": "Daniel Herman",
          "author_url": "",
          "post_date": "2024-02-13T11:46:49.350000",
          "content": "<p>Hello, for the column recorddate_4527225D we have the definition \"Date of tax deduction record.\". This date is from the future as it should be. All date columns were transformed. </p>",
          "votes": 1,
          "replies": [
            {
              "id": 2651136,
              "author_name": "Sophie",
              "author_url": "",
              "post_date": "2024-02-14T01:16:16.710000",
              "content": "<p>Does the prediction of loan default happen before the decision date or the first due date of the loan? Besides, if the internal/external data after the decision date is also available for model prediction, will there be an information leak? </p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2651538,
              "author_name": "Tomas Jelinek",
              "author_url": "",
              "post_date": "2024-02-14T08:27:45.747000",
              "content": "<p>Hi Sophie,<br>\nprediction is done at the time of application (decision date). All data are collected as of that date (including external data).</p>",
              "votes": 3,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2643157,
      "author_name": "LucaMTB",
      "author_url": "",
      "post_date": "2024-02-08T16:51:25.217000",
      "content": "<p>Can you explain the content of person_1 file? What is the meaning of role_1084L</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2643976,
          "author_name": "Tomas Jelinek",
          "author_url": "",
          "post_date": "2024-02-09T07:55:39.317000",
          "content": "<p>Each credit application can have information about several persons (e.g. client him/her-self, contact references). Role describe type of connection to client.</p>",
          "votes": 2,
          "replies": [
            {
              "id": 2643981,
              "author_name": "LucaMTB",
              "author_url": "",
              "post_date": "2024-02-09T08:03:43.563000",
              "content": "<p>Thank you. Can you explain the meaning of the categories in role_1084L</p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2641740,
      "author_name": "Rendra A",
      "author_url": "",
      "post_date": "2024-02-07T16:33:19.477000",
      "content": "<p>hi,</p>\n<p>in columns 'credtype_587L', what is the abbreviation of : </p>\n<p>COL, CAL, REL ?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2642543,
          "author_name": "Daniel Herman",
          "author_url": "",
          "post_date": "2024-02-08T08:22:27.567000",
          "content": "<p>Hello. Those stand for different types of loans.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 2641584,
      "author_name": "Elizabeth Mejia-RicarG",
      "author_url": "",
      "post_date": "2024-02-07T15:08:31.297000",
      "content": "<p>Hello! Quick question: since this is an approvals model, I know there are some regulation issues that would prevent data scientists from using less interpretable models (although more predictive). Should we use the same principles here?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2641597,
          "author_name": "Daniel Herman",
          "author_url": "",
          "post_date": "2024-02-07T15:11:46.300000",
          "content": "<p>Hello! We don't prevent you from using any black box model you can think of. The only limitations are the computational resources and 12 hours for one submission. </p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 2637330,
      "author_name": "itsRAWRtime007",
      "author_url": "",
      "post_date": "2024-02-05T16:47:50.073000",
      "content": "<p>heyy all and good luck!</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2637362,
          "author_name": "",
          "author_url": "",
          "post_date": "2024-02-05T17:06:16.207000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2650845,
      "author_name": "Meagan Burkhart",
      "author_url": "",
      "post_date": "2024-02-13T18:11:15.127000",
      "content": "<p>Any advice for out of memory issues? My notebook runs fine but just during the scoring it throws the error</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 2837995,
      "author_name": "Ting Ellis",
      "author_url": "",
      "post_date": "2024-05-26T18:36:17.987000",
      "content": "<p>Do I need to use all the datasets provided during the model implementation?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2799042,
      "author_name": "Azadeh Razmi",
      "author_url": "",
      "post_date": "2024-05-07T15:11:15.897000",
      "content": "<p>I didn't have the chance of compete , but I think It was a grate challenge </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2749282,
      "author_name": "Darsh Jaganmohan",
      "author_url": "",
      "post_date": "2024-04-13T01:29:09.760000",
      "content": "<p><a href=\"https://www.kaggle.com/jetakow\" target=\"_blank\">@jetakow</a> -  I am a high school junior passionate about Machine Learning. I noticed that there is a minimum age (18 years old) limit to enter this competition. We have already submitted the parental consent form to enter Kaggle competitions. Is there any additional parental consent form to enter this competition? Thanks!</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2752258,
          "author_name": "Daniel Herman",
          "author_url": "",
          "post_date": "2024-04-14T20:20:33.510000",
          "content": "<p>Please create a new post in the Discussion for this topic. Thanks.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 2752931,
          "author_name": "Tomas Jelinek",
          "author_url": "",
          "post_date": "2024-04-15T08:18:56.530000",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/darshjaganmohan\" target=\"_blank\">@darshjaganmohan</a>,<br>\nI think this is question to Kaggle staff, please try to contact them. Perhaps <a href=\"https://www.kaggle.com/addisonhoward\" target=\"_blank\">@addisonhoward</a> can help?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 2753498,
          "author_name": "Addison Howard",
          "author_url": "",
          "post_date": "2024-04-15T14:24:29.847000",
          "content": "<p>Hi Darsh! Welcome to Kaggle. There are two consent forms to submit. One for <a href=\"https://www.kaggle.com/guardian-consent-minor-use\" target=\"_blank\">using Kaggle</a> and one for entering <a href=\"https://www.kaggle.com/consent-minors-process\" target=\"_blank\">each Kaggle Competition</a>. Please ensure you are compliant with local laws regarding participation as a minor as it applies to you.</p>",
          "votes": 1,
          "replies": [
            {
              "id": 2754029,
              "author_name": "Darsh Jaganmohan",
              "author_url": "",
              "post_date": "2024-04-15T18:58:36.847000",
              "content": "<p>Thank you everyone for the responses.</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2686980,
      "author_name": "Arun",
      "author_url": "",
      "post_date": "2024-03-08T07:26:51.127000",
      "content": "<p>Thank you for organizing this competition. </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2668627,
      "author_name": "Eva Szin Takacs",
      "author_url": "",
      "post_date": "2024-02-25T19:05:49.730000",
      "content": "<p>Hello All, I need advise how to submit solutions for this competition. I tried everything but I still receive \"Cannot submit<br>\nSubmissions have been disabled for this competition.\" Any help is appreciated.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2669201,
          "author_name": "Branden Murray",
          "author_url": "",
          "post_date": "2024-02-26T06:21:36.630000",
          "content": "<p>Submitting has been paused while they workout issues with the scoring metric.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2810337,
      "author_name": "Xia Yang",
      "author_url": "",
      "post_date": "2024-05-13T08:17:26.727000",
      "content": "<p>thanks for this funny comp</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2796369,
      "author_name": "Monish Ostwal",
      "author_url": "",
      "post_date": "2024-05-06T07:39:28.607000",
      "content": "<p>Thanks for sharing</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2774038,
      "author_name": "LunaYi233",
      "author_url": "",
      "post_date": "2024-04-25T03:24:02.827000",
      "content": "<p>thanks a lot!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2757219,
      "author_name": "Daniil Selivanov",
      "author_url": "",
      "post_date": "2024-04-17T11:54:44.647000",
      "content": "<p>Thank you!</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2632545": "Dear Kagglers,\n\nHome Credit’s back! After a 5-year hiatus, I’m delighted to welcome you to another of our data-packed competitions “Home Credit: Credit Risk Model Stability”, where the winning teams split a hefty jackpot and their solutions inspire our lending approval process. \n\nAbout the contest, many things remain same as in our previous competition \"Home Credit Default Risk\": the goal is to assess the credit risk quality of clients and predict their future payment behavior. We’ve provided data from a wide range of sources: application forms, social-demographic data, previous credit behavior data, etc. We’ve endeavored to keep the data in a raw form, and where necessary (e.g., with personal data, business sensitive information) we’ve applied data masking instead of aggregations to retain as much information value as possible. We’ve even kept effects of business decisions in the dataset intact, for example, a change in usage of certain data sources, (please read the data description carefully, so you’re aware of all potential data pitfalls). I can honestly state that the data sample presented mirrors almost exactly the datasets we use for real machine learning problems in Home Credit.\n \nThough some aspects of this contest are the same as before, this is no rehashing of our previous competition—we’ve got lots of new things to share.   Two of the most important are the vast size of the dataset and the evaluation metric. As to the first, we’ve decided to provide you with data on the majority of our portfolio so you will have enough ‘food’ even for your most data hungry algorithms and ML approaches. Measured by raw size, the current sample is roughly 10 times bigger than the previous one. At the same time, we’ve tried to keep the structure of data sample simple and readable. As to the second, the evaluation metric, though we still use AUC (Gini coefficient) to measure the accuracy of model predictions, we’ve added additional criteria related to model stability.  We want to see solutions that produce sound predictions also far in the future since this is a fundamental business requirement. There’s significant delay in target observations and limited capacities of R&D teams, so there’s a strong preference for models that last over a long period of time. We’re very keen to see how you’ll handle and incorporate this requirement into your solutions.\n \nLast but not least, as mentioned above, we’ve significantly boosted the prize money and also increased the number of winning places that can claim a cut in the cash!\n \nIf you have any questions regarding the data, competition, R&D team or our company, please feel free to ask me and my colleague Daniel Herman.  We’re here to answer all your inquiries and support you as best we can.\n \nGood luck; and we’re looking forward to hearing from you in the discussion.\n \nOn behalf of the entire Home Credit R&D team,\n \nTomas Jelinek",
    "2637309": "I am looking forward to seeing all the creative solutions! Best of luck to everyone joining us this year! 🚀🚀🚀",
    "2644081": "Thank you for organizing this competition! I wonder how large is the test set for the private leaderboard? \n\nThe data section says \"test_base.csv contains approximately 90% of the numbers of case_id values of train_base.csv\". I assume this refers to the size of test set for the public leaderboard. How about the size of test set for the private leaderboard? Will it be twice the size of the public leaderboard test set?\n\nThank you!",
    "2638001": "Can we have a public notebook for metric calculation script? Can we use external data source?",
    "2826427": "I am getting this \"TypeError: the truth value of a Series is ambiguous\" when trying to use the Polars library for data cleaning, anyone has idea on how to deal with this. Thanks",
    "2842179": "thank you for organizing this competition. I had a good time doing this.",
    "2656081": "Thank you so much for organizing such an amazing competition. I'm trying to understand the data and I don't exactly understand what NumGroup1,2 means. \n\nThe data in Depth 1 is easy to understand because it only has the column \"num_group1\",\nbut the data in Depth 2 has \"num_group1\" , \"num_group2\", which is confusing.  Does having more than one type of num_group in Depth 2 mean that there are multiple branches within a case?",
    "2649736": "Hi, \nI have a question about the date. Can you explain why 'recorddate_4527225D' in tax_registry_a is always larger than 'date_decision' in base table? \n\ncase_id\tamount_4527230A\tname_4527232M\tnum_group1\trecorddate_4527225D\tdate_decision\tMONTH\tWEEK_NUM\ttarget\n28631\t1946.0\t\"f980a1ea\"\t2\t\"2019-09-13\"\t\"2019-08-30\"\t201908\t34\t0\n28631\t711.0\t\"f980a1ea\"\t3\t\"2019-09-13\"\t\"2019-08-30\"\t201908\t34\t0",
    "2643157": "Can you explain the content of person_1 file? What is the meaning of role_1084L",
    "2641740": "hi,\n\nin columns 'credtype_587L', what is the abbreviation of : \n\nCOL, CAL, REL ?",
    "2641584": "Hello! Quick question: since this is an approvals model, I know there are some regulation issues that would prevent data scientists from using less interpretable models (although more predictive). Should we use the same principles here?",
    "2637330": "heyy all and good luck!",
    "2650845": "Any advice for out of memory issues? My notebook runs fine but just during the scoring it throws the error",
    "2837995": "Do I need to use all the datasets provided during the model implementation?",
    "2799042": "I didn't have the chance of compete , but I think It was a grate challenge ",
    "2749282": "@jetakow -  I am a high school junior passionate about Machine Learning. I noticed that there is a minimum age (18 years old) limit to enter this competition. We have already submitted the parental consent form to enter Kaggle competitions. Is there any additional parental consent form to enter this competition? Thanks!",
    "2686980": "Thank you for organizing this competition. ",
    "2668627": "Hello All, I need advise how to submit solutions for this competition. I tried everything but I still receive \"Cannot submit\nSubmissions have been disabled for this competition.\" Any help is appreciated.",
    "2810337": "thanks for this funny comp",
    "2796369": "Thanks for sharing",
    "2774038": "thanks a lot!",
    "2757219": "Thank you!"
  }
}