{
  "id": 508163,
  "title": "End of Competition - note from Home Credit",
  "url": "/competitions/home-credit-credit-risk-model-stability/discussion/508163",
  "author_name": "Tomas Jelinek",
  "post_date": "2024-05-28T13:30:48.775000",
  "votes": 34,
  "comment_count": 54,
  "views": 0,
  "content": "<p>Dear Competitors,</p>\n<p>As we approach the conclusion of the Home Credit - Credit Risk Model Stability competition, we want to extend our gratitude to all enthusiastic participants who contributed to this competition with their hard work. We are looking forward to dive deep into your solutions and explore all the innovative approaches tackling the credit risk model.</p>\n<p>We hope that this competition has provided valuable learning opportunities and insights for all participants. Whether you're a seasoned data scientist or new to the field, we hope you mastered new skills and gained knowledge that will benefit your future endeavors. It has also been inspiring to witness the collaborative spirit within the Kaggle community, where members support each other and share valuable insights.</p>\n<p>The topic of model stability has proven to be challenging, and we've all learned a great deal throughout this competition. The project brought moments of joy as well as periods of painful struggle, and we appreciate your perseverance and commitment to tackling these complex issues.</p>\n<p>As the competition comes to an end, the Kaggle team will finalize the standings and share them with you through a separate communication. Please stay tuned for updates on the timeline for these announcements.</p>\n<p>Throughout the competition, as hosts, we exercised caution in responding to queries to prevent unintended advantages (e.g., disclosing differences between the private and public test sets). If you still have unanswered questions, please don't hesitate to ask them here.</p>\n<p>Once again, thank you for your enthusiastic participation. A big congratulations to the winning teams! We eagerly anticipate reviewing your final submissions and exploring the innovative strategies employed by the top performers.</p>\n<p>On behalf of the entire Home Credit R&amp;D team,</p>\n<p>Tomas Jelinek</p>",
  "messages": [
    {
      "id": 2841246,
      "postDate": "2024-05-28T13:30:48.777Z",
      "content": "<p>Dear Competitors,</p>\n<p>As we approach the conclusion of the Home Credit - Credit Risk Model Stability competition, we want to extend our gratitude to all enthusiastic participants who contributed to this competition with their hard work. We are looking forward to dive deep into your solutions and explore all the innovative approaches tackling the credit risk model.</p>\n<p>We hope that this competition has provided valuable learning opportunities and insights for all participants. Whether you're a seasoned data scientist or new to the field, we hope you mastered new skills and gained knowledge that will benefit your future endeavors. It has also been inspiring to witness the collaborative spirit within the Kaggle community, where members support each other and share valuable insights.</p>\n<p>The topic of model stability has proven to be challenging, and we've all learned a great deal throughout this competition. The project brought moments of joy as well as periods of painful struggle, and we appreciate your perseverance and commitment to tackling these complex issues.</p>\n<p>As the competition comes to an end, the Kaggle team will finalize the standings and share them with you through a separate communication. Please stay tuned for updates on the timeline for these announcements.</p>\n<p>Throughout the competition, as hosts, we exercised caution in responding to queries to prevent unintended advantages (e.g., disclosing differences between the private and public test sets). If you still have unanswered questions, please don't hesitate to ask them here.</p>\n<p>Once again, thank you for your enthusiastic participation. A big congratulations to the winning teams! We eagerly anticipate reviewing your final submissions and exploring the innovative strategies employed by the top performers.</p>\n<p>On behalf of the entire Home Credit R&amp;D team,</p>\n<p>Tomas Jelinek</p>",
      "rawMarkdown": "Dear Competitors,\n \nAs we approach the conclusion of the Home Credit - Credit Risk Model Stability competition, we want to extend our gratitude to all enthusiastic participants who contributed to this competition with their hard work. We are looking forward to dive deep into your solutions and explore all the innovative approaches tackling the credit risk model.\n \nWe hope that this competition has provided valuable learning opportunities and insights for all participants. Whether you're a seasoned data scientist or new to the field, we hope you mastered new skills and gained knowledge that will benefit your future endeavors. It has also been inspiring to witness the collaborative spirit within the Kaggle community, where members support each other and share valuable insights.\n \nThe topic of model stability has proven to be challenging, and we've all learned a great deal throughout this competition. The project brought moments of joy as well as periods of painful struggle, and we appreciate your perseverance and commitment to tackling these complex issues.\n \nAs the competition comes to an end, the Kaggle team will finalize the standings and share them with you through a separate communication. Please stay tuned for updates on the timeline for these announcements.\n \nThroughout the competition, as hosts, we exercised caution in responding to queries to prevent unintended advantages (e.g., disclosing differences between the private and public test sets). If you still have unanswered questions, please don't hesitate to ask them here.\n \nOnce again, thank you for your enthusiastic participation. A big congratulations to the winning teams! We eagerly anticipate reviewing your final submissions and exploring the innovative strategies employed by the top performers.\n \nOn behalf of the entire Home Credit R&D team,\n \nTomas Jelinek",
      "votes": 34
    },
    {
      "id": 2847359,
      "postDate": "2024-05-31T14:10:59.097Z",
      "content": "<p>The ideas, views and opinions expressed in this comment represent my own views and not those of any of my current or previous employer. Maybe some of you noticed that I went mostly silent after March, there was a reason for that. As of 1.4.2024 I am no longer employed by&nbsp;Home Credit or EmbedIT and my current affiliation is with Similarweb. The decision to change my employer has little to do with this competition, it was done before the start of the competition and it was purely mine.&nbsp;</p>\n<p>To put into perspective what was the amount of time put into the preparations for this competition, it has been exactly a year since we began with the data preparation. We didn't work on this project continuously throughout the year, but if we sum all the bits and pieces it was about several months of work (longer than the whole competition). Every time I read the positive feedback here in Discussion I feel like it was worth the pain and effort. Thank you all for that.</p>\n<p>There was some amount of very frustrated and angry reactions in Discussion. I noticed many people managed to stay professional even despite that and I tried to remain professional as well. I continued reading the posts and comments in the Discussion even after I stopped working for my previous employer. It probably wasn't a healthy thing to do to read all of it. I would really wish if those&nbsp;people had taken the competition less personally and tried to enjoy the challenge that was given to them. I know it is also about winning the money and the fame, but that is not all. The venting usually resolves nothing.&nbsp;Besides that, you can always step aside when you don't like the challenge and forget about the whole competition. <br>\n&nbsp;<br>\nI want to thank you all for participating. I hope most of you enjoyed the challenge and this competition served as a great opportunity for many of you to work on your data science skills. I want to thank also Kaggle for the&nbsp;data part and for helping us to resolve the issues we faced.&nbsp;</p>",
      "rawMarkdown": "The ideas, views and opinions expressed in this comment represent my own views and not those of any of my current or previous employer. Maybe some of you noticed that I went mostly silent after March, there was a reason for that. As of 1.4.2024 I am no longer employed by Home Credit or EmbedIT and my current affiliation is with Similarweb. The decision to change my employer has little to do with this competition, it was done before the start of the competition and it was purely mine. \n\nTo put into perspective what was the amount of time put into the preparations for this competition, it has been exactly a year since we began with the data preparation. We didn't work on this project continuously throughout the year, but if we sum all the bits and pieces it was about several months of work (longer than the whole competition). Every time I read the positive feedback here in Discussion I feel like it was worth the pain and effort. Thank you all for that.\n\nThere was some amount of very frustrated and angry reactions in Discussion. I noticed many people managed to stay professional even despite that and I tried to remain professional as well. I continued reading the posts and comments in the Discussion even after I stopped working for my previous employer. It probably wasn't a healthy thing to do to read all of it. I would really wish if those people had taken the competition less personally and tried to enjoy the challenge that was given to them. I know it is also about winning the money and the fame, but that is not all. The venting usually resolves nothing. Besides that, you can always step aside when you don't like the challenge and forget about the whole competition. \n \nI want to thank you all for participating. I hope most of you enjoyed the challenge and this competition served as a great opportunity for many of you to work on your data science skills. I want to thank also Kaggle for the data part and for helping us to resolve the issues we faced. ",
      "votes": 11
    },
    {
      "id": 2841301,
      "postDate": "2024-05-28T14:08:17.037Z",
      "content": "<p>Hi,</p>\n<p>Thank you for organizing this competition. It was both fun and challenging, especially in terms of modeling and computational constraints.</p>\n<p>I wanted to ask if the full test set will be made available for analysis. It would be fascinating to examine the post-COVID effects on variables like the average days past due.</p>",
      "rawMarkdown": "Hi,\n\nThank you for organizing this competition. It was both fun and challenging, especially in terms of modeling and computational constraints.\n\nI wanted to ask if the full test set will be made available for analysis. It would be fascinating to examine the post-COVID effects on variables like the average days past due.",
      "votes": 9,
      "replies": [
        {
          "id": 2841757,
          "postDate": "2024-05-28T17:20:47.267Z",
          "content": "<p>Yeah, I would also like to have access to full test set to check why more than half of my submissions failed!</p>",
          "rawMarkdown": "Yeah, I would also like to have access to full test set to check why more than half of my submissions failed!"
        },
        {
          "id": 2841771,
          "postDate": "2024-05-28T17:28:11.110Z",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/alessandrobenetti\" target=\"_blank\">@alessandrobenetti</a>,<br>\npersonally I am okay to share the test set, but let me first check with Kaggle, not sure whether there are not some restrictions (e.g., late submissions will become pointless)</p>",
          "rawMarkdown": "Hi @alessandrobenetti,\npersonally I am okay to share the test set, but let me first check with Kaggle, not sure whether there are not some restrictions (e.g., late submissions will become pointless)",
          "votes": 4,
          "replies": [
            {
              "id": 2841791,
              "postDate": "2024-05-28T17:35:17.263Z",
              "content": "<p>Thank you!</p>",
              "rawMarkdown": "Thank you!"
            },
            {
              "id": 2846072,
              "postDate": "2024-05-30T21:09:05.623Z",
              "content": "<p>That would be great, I think it will be likely used for research in future. </p>",
              "rawMarkdown": "That would be great, I think it will be likely used for research in future. "
            },
            {
              "id": 2935665,
              "postDate": "2024-07-25T13:37:41.210Z",
              "content": "<p>Hi <a href=\"https://www.kaggle.com/jetakow\" target=\"_blank\">@jetakow</a> <a href=\"https://www.kaggle.com/tomasjeline2\" target=\"_blank\">@tomasjeline2</a> , I fully agree that the dataset is incredibly valuable for research. Is it possible to use it for academic research in the future, possibly just the training set? Looking forward to your reply.</p>",
              "rawMarkdown": "Hi @jetakow @tomasjeline2 , I fully agree that the dataset is incredibly valuable for research. Is it possible to use it for academic research in the future, possibly just the training set? Looking forward to your reply.\n"
            }
          ]
        },
        {
          "id": 2846071,
          "postDate": "2024-05-30T21:08:12.760Z",
          "content": "<p>The best reward for me is when I see people dealing with the challenges and enjoying the challenge. It's exactly one year since I started the data preparation. There was a lot of effort put into the preparation of this competition also on the kaggle side and I am happy to see the results of it. </p>",
          "rawMarkdown": "The best reward for me is when I see people dealing with the challenges and enjoying the challenge. It's exactly one year since I started the data preparation. There was a lot of effort put into the preparation of this competition also on the kaggle side and I am happy to see the results of it. "
        }
      ]
    },
    {
      "id": 2841327,
      "postDate": "2024-05-28T14:26:34.717Z",
      "content": "<p>While your decision to keep the metric was probably not the best choice, your speed and availability to answer questions were greatly appreciated.</p>",
      "rawMarkdown": "While your decision to keep the metric was probably not the best choice, your speed and availability to answer questions were greatly appreciated.",
      "votes": 8,
      "replies": [
        {
          "id": 2841363,
          "postDate": "2024-05-28T14:55:35.990Z",
          "rawMarkdown": "",
          "votes": -3,
          "isDeleted": true
        },
        {
          "id": 2846074,
          "postDate": "2024-05-30T21:09:40.707Z",
          "content": "<p>Thanks, I tried my best during the time I worked at Home Credit.</p>",
          "rawMarkdown": "Thanks, I tried my best during the time I worked at Home Credit."
        }
      ]
    },
    {
      "id": 2844176,
      "postDate": "2024-05-30T01:55:10.887Z",
      "content": "<p>Unfortunately, I agree with most of <a href=\"https://www.kaggle.com/at7459\" target=\"_blank\">@at7459</a> points.</p>\n<p>1.For the first time there was a hack, I think the organizers did a good job of changing the data and extending the competition(Whatever the result, the decision was good). When the second time there was a hacker, I was very disappointed that host didn't do anything,  silence was your biggest mistake.</p>\n<p>2.Obviously, top solutions are almost metric hack.Because you promised to punish the Metric Hack, many kaggler didn't choose the public Metric Hack notebook. If you let the metric hack get prize or medal, I think it's a betrayal of us.If you think that the current solution does not have metrik hack, As I just said, silence was your biggest mistake.</p>\n<p>3.I'm glad I didn't invest too much time in the final stage.Although most of Kaggler are volunteers, we can't blame others even if we waste time,but after this competition I believe that many kaggler will be more careful in choosing the competition. Kaggle should give the host more professional advice to avoid the current situation,or kaggle have already done it?</p>",
      "rawMarkdown": "Unfortunately, I agree with most of @at7459 points.\n\n1.For the first time there was a hack, I think the organizers did a good job of changing the data and extending the competition(Whatever the result, the decision was good). When the second time there was a hacker, I was very disappointed that host didn't do anything,  silence was your biggest mistake.\n\n2.Obviously, top solutions are almost metric hack.Because you promised to punish the Metric Hack, many kaggler didn't choose the public Metric Hack notebook. If you let the metric hack get prize or medal, I think it's a betrayal of us.If you think that the current solution does not have metrik hack, As I just said, silence was your biggest mistake.\n\n3.I'm glad I didn't invest too much time in the final stage.Although most of Kaggler are volunteers, we can't blame others even if we waste time,but after this competition I believe that many kaggler will be more careful in choosing the competition. Kaggle should give the host more professional advice to avoid the current situation,or kaggle have already done it?\n",
      "votes": 6
    },
    {
      "id": 2842244,
      "postDate": "2024-05-29T01:27:09.820Z",
      "content": "<p>Hi Tomas!<br>\nThank you very much for organizing this competition, as a practitioner related to the field of banking risk control, I have always wanted to participate in competitions on related topics at kaggle.<br>\nCurrently, both credit risk control and credit card risk control face the problem of instability under the influence of epidemics or economic factors, and the traditional logistic regression scorecard using metrics such as PSI and CSI has certain limitations in stability monitoring.<br>\nThe evaluation metrics of this competition as well as the alternatives provided by other competitors gave me great inspiration, and I will seriously discuss within my team how to embed the above metrics into the model development process.<br>\nLast but not least, I would like to thank you for providing an opportunity to participate in a structured data competition when unstructured data competitions are in vogue, which has improved my personal programming skills and understanding of stability to a certain level.</p>",
      "rawMarkdown": "Hi Tomas!\nThank you very much for organizing this competition, as a practitioner related to the field of banking risk control, I have always wanted to participate in competitions on related topics at kaggle.\nCurrently, both credit risk control and credit card risk control face the problem of instability under the influence of epidemics or economic factors, and the traditional logistic regression scorecard using metrics such as PSI and CSI has certain limitations in stability monitoring.\nThe evaluation metrics of this competition as well as the alternatives provided by other competitors gave me great inspiration, and I will seriously discuss within my team how to embed the above metrics into the model development process.\nLast but not least, I would like to thank you for providing an opportunity to participate in a structured data competition when unstructured data competitions are in vogue, which has improved my personal programming skills and understanding of stability to a certain level.",
      "votes": 3,
      "replies": [
        {
          "id": 2842632,
          "postDate": "2024-05-29T07:35:37.027Z",
          "content": "<p>Thank you <a href=\"https://www.kaggle.com/simingtan\" target=\"_blank\">@simingtan</a> and congratulations!</p>",
          "rawMarkdown": "Thank you @simingtan and congratulations!"
        }
      ]
    },
    {
      "id": 2841919,
      "postDate": "2024-05-28T18:52:29.833Z",
      "content": "<p>Leaving aside the problems we encountered, it was the most comprehensive data set that I enjoyed working on</p>",
      "rawMarkdown": "Leaving aside the problems we encountered, it was the most comprehensive data set that I enjoyed working on",
      "votes": 3,
      "replies": [
        {
          "id": 2842080,
          "postDate": "2024-05-28T20:49:27.233Z",
          "content": "<p>I completely agree. I found the dataset was so different but extensive, than what I have seen in my short time in the field and every moment I spent on the dataset was a new lesson for me.</p>",
          "rawMarkdown": "I completely agree. I found the dataset was so different but extensive, than what I have seen in my short time in the field and every moment I spent on the dataset was a new lesson for me.",
          "votes": 1
        },
        {
          "id": 2845998,
          "postDate": "2024-05-30T20:00:12.803Z",
          "content": "<p>I agree. It broke my mind every time I wanted to understand why some features degraded my score.</p>",
          "rawMarkdown": "I agree. It broke my mind every time I wanted to understand why some features degraded my score.",
          "votes": 1
        },
        {
          "id": 2846079,
          "postDate": "2024-05-30T21:16:13.890Z",
          "content": "<p>I am glad you see that way. There was a lot of energy and many weeks put into the data preparation so the kaggle community doesn't have to deal with it and also we can publish it. To give you some perspective for how long it took, it is exactly year since the whole activity started and I was asked to start preparing the data, we have to stop this activity for several months, but yeah we and kaggle put a lot of energy into this to make it happen. </p>",
          "rawMarkdown": "I am glad you see that way. There was a lot of energy and many weeks put into the data preparation so the kaggle community doesn't have to deal with it and also we can publish it. To give you some perspective for how long it took, it is exactly year since the whole activity started and I was asked to start preparing the data, we have to stop this activity for several months, but yeah we and kaggle put a lot of energy into this to make it happen. ",
          "votes": 3
        }
      ]
    },
    {
      "id": 2844113,
      "postDate": "2024-05-30T00:03:02.573Z",
      "content": "<p><a href=\"https://www.kaggle.com/tomasjeline2\" target=\"_blank\">@tomasjeline2</a> &amp; <a href=\"https://www.kaggle.com/jetakow\" target=\"_blank\">@jetakow</a>, I congratulate you both for your splendid work in good an difficult times. I did not participate, but watched the competition (not so) closely. It really was outstanding  the way you sorted out problems. Congrats again and to all participants!</p>",
      "rawMarkdown": "@tomasjeline2 & @jetakow, I congratulate you both for your splendid work in good an difficult times. I did not participate, but watched the competition (not so) closely. It really was outstanding  the way you sorted out problems. Congrats again and to all participants!",
      "votes": 2,
      "replies": [
        {
          "id": 2846070,
          "postDate": "2024-05-30T21:04:16.527Z",
          "content": "<p>Thanks, I did my best I could in the given time to help resolve the questions and issues and I tried the best to make this competition enjoyable for everyone. The sad reality is that you can't make everyone happy. </p>",
          "rawMarkdown": "Thanks, I did my best I could in the given time to help resolve the questions and issues and I tried the best to make this competition enjoyable for everyone. The sad reality is that you can't make everyone happy. ",
          "votes": 2
        }
      ]
    },
    {
      "id": 2844908,
      "postDate": "2024-05-30T09:33:46.677Z",
      "content": "<p>Thank you for organizing this competition. It was both fun and challenging, especially in terms of modeling and computational constraints.</p>",
      "rawMarkdown": "Thank you for organizing this competition. It was both fun and challenging, especially in terms of modeling and computational constraints.",
      "votes": 1,
      "replies": [
        {
          "id": 2846081,
          "postDate": "2024-05-30T21:17:26.467Z",
          "content": "<p>It is the best reward for me personally to see people enjoying the challenge and taking it as a learning experience. </p>",
          "rawMarkdown": "It is the best reward for me personally to see people enjoying the challenge and taking it as a learning experience. ",
          "votes": 1
        }
      ]
    },
    {
      "id": 2842867,
      "postDate": "2024-05-29T09:46:21.797Z",
      "content": "<p>Hi Tomas, may I ask how you transform the date_decision column? I also found some '_T' columns have a strong correlation with week_num after I aggregate them by date_decision as shown below. But if I don't aggregate them by date, it looks like there will be some fluctuations within each week.<br>\nSo I test the date_decision column of the test dataset for groupby, and then used the aggregated value of the '_T' column for hacking, but found that the LB was significantly worse than the result without the date_decision aggregation. So it seems that after you transform the date_decision column, I can't even use it for aggregation operations. <br>\nDid you add some random noise on it?😂<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F19087078%2F874de32eeed1c44d3a3c8b3865df6819%2F1716975952722.jpg?generation=1716975963601860&amp;alt=media\"><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F19087078%2F55b015c508738fb451d06a81e4c4e155%2FSnipaste_2024-05-29_17-58-59.png?generation=1716976752681331&amp;alt=media\"></p>",
      "rawMarkdown": "Hi Tomas, may I ask how you transform the date_decision column? I also found some '_T' columns have a strong correlation with week_num after I aggregate them by date_decision as shown below. But if I don't aggregate them by date, it looks like there will be some fluctuations within each week.\nSo I test the date_decision column of the test dataset for groupby, and then used the aggregated value of the '_T' column for hacking, but found that the LB was significantly worse than the result without the date_decision aggregation. So it seems that after you transform the date_decision column, I can't even use it for aggregation operations. \nDid you add some random noise on it?😂\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F19087078%2F874de32eeed1c44d3a3c8b3865df6819%2F1716975952722.jpg?generation=1716975963601860&alt=media)![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F19087078%2F55b015c508738fb451d06a81e4c4e155%2FSnipaste_2024-05-29_17-58-59.png?generation=1716976752681331&alt=media)",
      "votes": 1,
      "replies": [
        {
          "id": 2842942,
          "postDate": "2024-05-29T10:53:00.843Z",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/wispvale\" target=\"_blank\">@wispvale</a>,<br>\ndata_decision transformation: we reshuffled it by random number of weeks into the future or history in a way that we keep target rate for each week same as it was in original data sample.<br>\nThen all other date columns were transformed in a same way. It means that all date differences should remain same (e.g., age = date_decision - date of birth). Our big mistake here was that we overlooked these month and year attributes that were as integers in data sample :-(</p>",
          "rawMarkdown": "Hi @wispvale,\ndata_decision transformation: we reshuffled it by random number of weeks into the future or history in a way that we keep target rate for each week same as it was in original data sample.\nThen all other date columns were transformed in a same way. It means that all date differences should remain same (e.g., age = date_decision - date of birth). Our big mistake here was that we overlooked these month and year attributes that were as integers in data sample :-(",
          "votes": 2
        }
      ]
    },
    {
      "id": 2841794,
      "postDate": "2024-05-28T17:36:34.710Z",
      "content": "<p>Could you please share details about how the test set was split into public/private? Did you change it after the hack was found?</p>",
      "rawMarkdown": "Could you please share details about how the test set was split into public/private? Did you change it after the hack was found?",
      "votes": 1,
      "replies": [
        {
          "id": 2841830,
          "postDate": "2024-05-28T17:50:13.080Z",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/eivolkova\" target=\"_blank\">@eivolkova</a>,<br>\nThank you for your active participation in discussions!</p>\n<p>Public/private is 30/70 random split.<br>\nAfter hack / before restart, besides changes we have communicated (hiding week_num etc.), we removed first 30 weeks in private test set, this was not originally planned. Since restart there was no change in test set.</p>",
          "rawMarkdown": "Hi @eivolkova,\nThank you for your active participation in discussions!\n\nPublic/private is 30/70 random split.\nAfter hack / before restart, besides changes we have communicated (hiding week_num etc.), we removed first 30 weeks in private test set, this was not originally planned. Since restart there was no change in test set.",
          "votes": 3,
          "replies": [
            {
              "id": 2842097,
              "postDate": "2024-05-28T21:05:28.907Z",
              "content": "<p>Did you reveal this to the community: \"we removed first 30 weeks in private test set\"? So just private test data, not public leaderboard data?</p>",
              "rawMarkdown": "Did you reveal this to the community: \"we removed first 30 weeks in private test set\"? So just private test data, not public leaderboard data?"
            },
            {
              "id": 2842618,
              "postDate": "2024-05-29T07:22:47.423Z",
              "content": "<p>No, information about 30 weeks in private test was not revealed. </p>\n<p>My opinion is that private test data should cover same problem/use case, but not necessarily same data or period as public test set. Otherwise winning strategy is about overfitting public test. </p>\n<p>We haven't changed public leaderboard data (besides masking column like week_num).</p>",
              "rawMarkdown": "No, information about 30 weeks in private test was not revealed. \n\nMy opinion is that private test data should cover same problem/use case, but not necessarily same data or period as public test set. Otherwise winning strategy is about overfitting public test. \n\nWe haven't changed public leaderboard data (besides masking column like week_num).",
              "votes": -4
            },
            {
              "id": 2842734,
              "postDate": "2024-05-29T08:34:31.477Z",
              "content": "<p>no, please do not hold us for stupid, that is not the reason at all why you did it, or you would have done so at the very start of the competition instead of when you realized that your metric does not work.</p>\n<p>instead, this was purely a very naive and hubris-driven way of trying to punish anyone who attempts hack the metric, which obviously did not work out at all and just ended up making the competition exponentially more random than leaving the private test data unchanged.</p>\n<p>instead of focusing on properly transorming the data set, you focused on making the competition as random as possible, and missed all the date columns which are not explicitly in datetime format effectively leaving the date_decision there to see for anyone who looks a bit into the data. which all the more shows that you do not even understand the data you gave to this competition.</p>\n<p>you repeatedly (specifically your colleague Daniel Herman) said that you would not allow a metric hack by way of looking through the top solutions,</p>\n<p><em>\"If the top 100 solutions will be using some way of week num prediction, yes we will go over those hundred solutions in the top private leaderboard until we find a solution that doesn't have this approach \"</em></p>\n<p>and then later in the <strong><em>LAST TWO WEEKS OF THE COMPETITION</em></strong> completely backtracked on that promise, leaving everyone scrambling to probe the leaderboard and find the best way to hack the metric.</p>",
              "rawMarkdown": "no, please do not hold us for stupid, that is not the reason at all why you did it, or you would have done so at the very start of the competition instead of when you realized that your metric does not work.\n\ninstead, this was purely a very naive and hubris-driven way of trying to punish anyone who attempts hack the metric, which obviously did not work out at all and just ended up making the competition exponentially more random than leaving the private test data unchanged.\n\ninstead of focusing on properly transorming the data set, you focused on making the competition as random as possible, and missed all the date columns which are not explicitly in datetime format effectively leaving the date_decision there to see for anyone who looks a bit into the data. which all the more shows that you do not even understand the data you gave to this competition.\n\nyou repeatedly (specifically your colleague Daniel Herman) said that you would not allow a metric hack by way of looking through the top solutions,\n\n*\"If the top 100 solutions will be using some way of week num prediction, yes we will go over those hundred solutions in the top private leaderboard until we find a solution that doesn't have this approach \"*\n\nand then later in the ***LAST TWO WEEKS OF THE COMPETITION*** completely backtracked on that promise, leaving everyone scrambling to probe the leaderboard and find the best way to hack the metric.\n\n",
              "votes": 13
            },
            {
              "id": 2842779,
              "postDate": "2024-05-29T08:58:24.960Z",
              "content": "<p><a href=\"https://www.kaggle.com/tomasjeline2\" target=\"_blank\">@tomasjeline2</a> If you delete the data from the top 30 weeks of the private test, the ratio of public to private data will still be 3:7? If it weren't for 3:7, this erroneous message would greatly affect the final result of this competition</p>",
              "rawMarkdown": "@tomasjeline2 If you delete the data from the top 30 weeks of the private test, the ratio of public to private data will still be 3:7? If it weren't for 3:7, this erroneous message would greatly affect the final result of this competition",
              "votes": 2
            },
            {
              "id": 2842876,
              "postDate": "2024-05-29T09:52:28.700Z",
              "content": "<p>In my opinion, if the organizers alter the test data during the competition, particularly the private test data, they must synchronize such changes with all participants.</p>\n<p><strong>Failure to do so could raise some significant honest risk</strong>. Given that the organizers can view all the submissions of all competitors, in future competitions, if hosts want a competitor to win, they can privately change a private dataset that fits the model of the competitor most.</p>",
              "rawMarkdown": "In my opinion, if the organizers alter the test data during the competition, particularly the private test data, they must synchronize such changes with all participants.\n\n**Failure to do so could raise some significant honest risk**. Given that the organizers can view all the submissions of all competitors, in future competitions, if hosts want a competitor to win, they can privately change a private dataset that fits the model of the competitor most.",
              "votes": 4
            },
            {
              "id": 2842920,
              "postDate": "2024-05-29T10:39:36.797Z",
              "content": "<p>Hi <a href=\"https://www.kaggle.com/at7459\" target=\"_blank\">@at7459</a>,<br>\nI don't see any question here, so I don't know what answer do you want from me. <br>\nYou made your frustration from the competition very clear on many other places in the discussion, so I really don't understand what do you want to achieve with your post here.</p>",
              "rawMarkdown": "Hi @at7459,\nI don't see any question here, so I don't know what answer do you want from me. \nYou made your frustration from the competition very clear on many other places in the discussion, so I really don't understand what do you want to achieve with your post here.",
              "votes": -5
            },
            {
              "id": 2842930,
              "postDate": "2024-05-29T10:43:08.660Z",
              "content": "<p>Hi <a href=\"https://www.kaggle.com/forcewithme\" target=\"_blank\">@forcewithme</a>,<br>\nOrganizers can't make changes to test set. In our case it was done after consultation with Kaggle, and restart was necessary, so all submissions were deleted.<br>\nUnresolved problem I am aware of is score of previously submitted notebooks in Code section, as it is not visible whether code was submitted before or after test set changes</p>",
              "rawMarkdown": "Hi @forcewithme,\nOrganizers can't make changes to test set. In our case it was done after consultation with Kaggle, and restart was necessary, so all submissions were deleted.\nUnresolved problem I am aware of is score of previously submitted notebooks in Code section, as it is not visible whether code was submitted before or after test set changes",
              "votes": -1
            },
            {
              "id": 2844038,
              "postDate": "2024-05-29T21:06:39.333Z",
              "content": "<p>yeah, i am not sure what i want from you at this point and it's not healthy to get that mad, so i will stop now lol</p>\n<p>but your team literally destroyed the entire competition with all these terrible decisions</p>\n<p>it is not even like these were difficult decisions to get right…</p>\n<p>the most important points are:</p>\n<ol>\n<li>just be clear about what is allowed and what is not allowed</li>\n<li>be very clear about how the test set got transformed and that you removed the first 30 weeks in the private test set</li>\n<li>listen to the people when they tell you how you are messing up and how to do things correctly</li>\n</ol>\n<p>it is not like people are out here to get you, and you need to set traps and hide things from everyone, this is not benficial for either you or the participants of this competition as you can see from the outcome of this mess</p>",
              "rawMarkdown": "yeah, i am not sure what i want from you at this point and it's not healthy to get that mad, so i will stop now lol\n\nbut your team literally destroyed the entire competition with all these terrible decisions\n\nit is not even like these were difficult decisions to get right...\n\nthe most important points are:\n\n1. just be clear about what is allowed and what is not allowed\n2. be very clear about how the test set got transformed and that you removed the first 30 weeks in the private test set\n3. listen to the people when they tell you how you are messing up and how to do things correctly\n\nit is not like people are out here to get you, and you need to set traps and hide things from everyone, this is not benficial for either you or the participants of this competition as you can see from the outcome of this mess",
              "votes": 4
            },
            {
              "id": 2844070,
              "postDate": "2024-05-29T21:58:01.190Z",
              "rawMarkdown": "",
              "votes": 4,
              "isDeleted": true
            },
            {
              "id": 2844620,
              "postDate": "2024-05-30T06:52:33.770Z",
              "content": "<p>Hi <a href=\"https://www.kaggle.com/tomasjeline2\" target=\"_blank\">@tomasjeline2</a> Most of us are frustrated by your changes at the end of the competition, and I'm sure you understand. But we should not be rude because the data were very interesting and the competition should be taken as it is, only a game and a sharing of ideas. Do you intend to come back later with a new metric that would be more resilient ?</p>",
              "rawMarkdown": "Hi @tomasjeline2 Most of us are frustrated by your changes at the end of the competition, and I'm sure you understand. But we should not be rude because the data were very interesting and the competition should be taken as it is, only a game and a sharing of ideas. Do you intend to come back later with a new metric that would be more resilient ?",
              "votes": 2
            }
          ]
        }
      ]
    },
    {
      "id": 2841770,
      "postDate": "2024-05-28T17:28:09.540Z",
      "content": "<p>Hi Tomas!</p>\n<p>First of all, thanks for hosting this competition. We believe that any result is also a result for the organizers, and they can make some useful conclusions from the competition.</p>\n<p>Our team have one question about dates. We tried to restore <code>WEEK_NUM</code>, and during this process we discovered that the following code successfully produces a submission file for the test set:</p>\n<pre><code>t = (- / ) * df_test[]\n\n t.shape[] *  &lt; t.isna().() &lt; t.shape[] * \n  &lt; t.quantile() &lt; \n  &lt;= t.quantile() &lt;= \n\nt1 = t[t &lt; ]\nt2 = t[(t &gt;= ) &amp; (t &lt;= )]\n\n - &lt; t1.() &lt; \n  &lt; t1.() &lt; \n\n  &lt; t1.shape[] / t.shape[] &lt; \n  &lt; t2.shape[] / t.shape[] &lt; \n</code></pre>\n<p>It seems that a significant part of the restored week nums are less than 30 (remember that the training period contains week nums between 0 and 91) or even negative. Could you tell us why it happened? :)</p>",
      "rawMarkdown": "Hi Tomas!\n\nFirst of all, thanks for hosting this competition. We believe that any result is also a result for the organizers, and they can make some useful conclusions from the competition.\n\nOur team have one question about dates. We tried to restore `WEEK_NUM`, and during this process we discovered that the following code successfully produces a submission file for the test set:\n\n```python\nt = (-1.0 / 7.0) * df_test['min_refreshdate_3813885D']\n\nassert t.shape[0] * 0.15 < t.isna().sum() < t.shape[0] * 0.25\nassert 0.4 < t.quantile(0.1) < 25.0\nassert 115.0 <= t.quantile(0.9) <= 140.0\n        \nt1 = t[t < 91.0]\nt2 = t[(t >= 91.0) & (t <= 135.0)]\n\nassert -30.0 < t1.min() < 0.0\nassert 0.0 < t1.max() < 30.0\n\nassert 0.2 < t1.shape[0] / t.shape[0] < 0.4\nassert 0.4 < t2.shape[0] / t.shape[0] < 0.5\n```\n\nIt seems that a significant part of the restored week nums are less than 30 (remember that the training period contains week nums between 0 and 91) or even negative. Could you tell us why it happened? :)",
      "votes": 1,
      "replies": [
        {
          "id": 2841782,
          "postDate": "2024-05-28T17:31:41.580Z",
          "content": "<p>It can be due to the fact that this method is not perfect. For some case_ids refreshdate can be different from 2019/1/3</p>",
          "rawMarkdown": "It can be due to the fact that this method is not perfect. For some case_ids refreshdate can be different from 2019/1/3",
          "replies": [
            {
              "id": 2841801,
              "postDate": "2024-05-28T17:38:32.103Z",
              "content": "<p>But for the training set it works muuuch better. Almost all week_nums (~90%, primarily due to NaNs in the <code>min_refreshdate_3813885D</code>, otherwise it would be almost 100%) are accurately restored. In the test set only ~50% are restored. Also, negative restored week_nums was some kind of surprise.</p>",
              "rawMarkdown": "But for the training set it works muuuch better. Almost all week_nums (~90%, primarily due to NaNs in the `min_refreshdate_3813885D`, otherwise it would be almost 100%) are accurately restored. In the test set only ~50% are restored. Also, negative restored week_nums was some kind of surprise."
            },
            {
              "id": 2841806,
              "postDate": "2024-05-28T17:41:00.757Z",
              "content": "<p>Yeah, I thought that the reason may be that 2019/1/3 was some kind of a filler for missing values, and the more data we have, the less number of missing values. Another option is that, for example, from year 2022 the filler value changed to i.e. 2022/1/1</p>\n<p>We can check this if the test set is published:)</p>",
              "rawMarkdown": "Yeah, I thought that the reason may be that 2019/1/3 was some kind of a filler for missing values, and the more data we have, the less number of missing values. Another option is that, for example, from year 2022 the filler value changed to i.e. 2022/1/1\n\nWe can check this if the test set is published:)"
            }
          ]
        },
        {
          "id": 2841816,
          "postDate": "2024-05-28T17:43:53.230Z",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/gromml\" target=\"_blank\">@gromml</a>,<br>\nwe will have to look into why min_refreshdate_3813885D works in week restoration.</p>\n<p>But to answer your question, test sample contains only weeks after train (91+), so there was no overlap with train and all test cases happen after train period.</p>",
          "rawMarkdown": "Hi @gromml,\nwe will have to look into why min_refreshdate_3813885D works in week restoration.\n\nBut to answer your question, test sample contains only weeks after train (91+), so there was no overlap with train and all test cases happen after train period.",
          "replies": [
            {
              "id": 2843261,
              "postDate": "2024-05-29T13:35:00.433Z",
              "content": "<p>Hi <a href=\"https://www.kaggle.com/gromml\" target=\"_blank\">@gromml</a> again,<br>\nquick update regarding min_refreshdate_3813885D, our hypothesis is following: this is credit bureau attribute, contains (probably) dates when client's report was updated. And for some reason data provider assigned 1/2019 to all clients who had CB report before that date. Once you discovered that, it's straightforward how to reconstruct date_decision, because same transformation was applied…</p>\n<p>Some clients apply for first loan and don't have previous CB report - you have missing value. And some clients have first credit after 1/2019 - that's why you have different value than 1/2019. Share of such clients is increasing in time.</p>",
              "rawMarkdown": "Hi @gromml again,\nquick update regarding min_refreshdate_3813885D, our hypothesis is following: this is credit bureau attribute, contains (probably) dates when client's report was updated. And for some reason data provider assigned 1/2019 to all clients who had CB report before that date. Once you discovered that, it's straightforward how to reconstruct date_decision, because same transformation was applied...\n\nSome clients apply for first loan and don't have previous CB report - you have missing value. And some clients have first credit after 1/2019 - that's why you have different value than 1/2019. Share of such clients is increasing in time.\n",
              "votes": 5
            }
          ]
        },
        {
          "id": 2842270,
          "postDate": "2024-05-29T02:21:13.193Z",
          "content": "<p>According to my experiments, the min_refreshdate_3813885D keeps decresing until around -1100, but goes up even to be larger than -640 (the minimum in train), and then decresing again. My guess is the refreshdate_3813885D has been updated at some point, it does not always be 2019/1. </p>",
          "rawMarkdown": "According to my experiments, the min_refreshdate_3813885D keeps decresing until around -1100, but goes up even to be larger than -640 (the minimum in train), and then decresing again. My guess is the refreshdate_3813885D has been updated at some point, it does not always be 2019/1. ",
          "votes": 1
        }
      ]
    },
    {
      "id": 2845751,
      "postDate": "2024-05-30T16:58:48.770Z",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/tomasjeline2\" target=\"_blank\">@tomasjeline2</a>, Could you explain the meaning of 'price_1097A'? I'm not sure what is 'Credit price', does it means something like interest?</p>",
      "rawMarkdown": "Hi @tomasjeline2, Could you explain the meaning of 'price_1097A'? I'm not sure what is 'Credit price', does it means something like interest?",
      "replies": [
        {
          "id": 2852171,
          "postDate": "2024-06-03T07:00:35.030Z",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/polsyuus\" target=\"_blank\">@polsyuus</a>,<br>\nI have checked with underwriting team, according to them it is price of goods that is financed by credit.</p>",
          "rawMarkdown": "Hi @polsyuus,\nI have checked with underwriting team, according to them it is price of goods that is financed by credit."
        }
      ]
    },
    {
      "id": 2842296,
      "postDate": "2024-05-29T03:10:04.470Z",
      "content": "<p>Hi,Thomas,<br>\nI wonder in practice how Home Credit deals with this model stability issue. Do you train on a rolling basis with new data every time to make the model as effective and can you explain further how the model is used ,for credit approval or screening?</p>",
      "rawMarkdown": "Hi,Thomas,\nI wonder in practice how Home Credit deals with this model stability issue. Do you train on a rolling basis with new data every time to make the model as effective and can you explain further how the model is used ,for credit approval or screening?",
      "replies": [
        {
          "id": 2842662,
          "postDate": "2024-05-29T07:57:30.693Z",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/godgod3\" target=\"_blank\">@godgod3</a>,</p>\n<p>Regarding the stability, in credit risk scorecards there is issue with delay of target observability - you can't say whether client is good/bad based on payment of first installment, so you have to observe first few payments, and delay can be half a year or even more, plus time for model development, implementation, adjustment of underwriting strategies, etc. So continuous re-development does not solve the stability problem, you always train on \"outdated\" data. Different story is for other use cases, for example sales, where you can observe results almost immediately.<br>\nWe check stability of predictors, both from perspective of population structure, and relationship feature-target. But decision what is acceptable threshold for stability is expert-based. And before the competition we were thinking whether there is better approach…</p>\n<p>Regarding model usage. First of all, I read some concerns in discussion regarding bias/discrimination of models, or their interpretability - we've never intended to use Kaggle models in production. Purpose of this competition is to learn new approaches or techniques we can learn from and to improve our model development process.<br>\nHow model is used - for approval, and it can be actual credit application, but also credit offer (I guess this is what you meant by screening). Model is not the only tool used in underwriting, there are many other conditions for approval (e.g. legal/regulatory - min age of applicant &gt; X, affordability - ratio if disposable income and instalment payment &gt; Y, etc.). Besides approval model decides quality of client, which has impact on credit parameters (typically better client = higher limit / credit amount, or lower interest)</p>",
          "rawMarkdown": "Hi @godgod3,\n\nRegarding the stability, in credit risk scorecards there is issue with delay of target observability - you can't say whether client is good/bad based on payment of first installment, so you have to observe first few payments, and delay can be half a year or even more, plus time for model development, implementation, adjustment of underwriting strategies, etc. So continuous re-development does not solve the stability problem, you always train on \"outdated\" data. Different story is for other use cases, for example sales, where you can observe results almost immediately.\nWe check stability of predictors, both from perspective of population structure, and relationship feature-target. But decision what is acceptable threshold for stability is expert-based. And before the competition we were thinking whether there is better approach...\n\nRegarding model usage. First of all, I read some concerns in discussion regarding bias/discrimination of models, or their interpretability - we've never intended to use Kaggle models in production. Purpose of this competition is to learn new approaches or techniques we can learn from and to improve our model development process.\nHow model is used - for approval, and it can be actual credit application, but also credit offer (I guess this is what you meant by screening). Model is not the only tool used in underwriting, there are many other conditions for approval (e.g. legal/regulatory - min age of applicant > X, affordability - ratio if disposable income and instalment payment > Y, etc.). Besides approval model decides quality of client, which has impact on credit parameters (typically better client = higher limit / credit amount, or lower interest)",
          "votes": 3
        }
      ]
    },
    {
      "id": 2920846,
      "postDate": "2024-07-13T20:47:37.850Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 2852368,
      "postDate": "2024-06-03T08:41:21.957Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 2842394,
      "postDate": "2024-05-29T04:22:09.520Z",
      "rawMarkdown": "",
      "isDeleted": true,
      "replies": [
        {
          "id": 2842685,
          "postDate": "2024-05-29T08:04:29.840Z",
          "content": "<p>Hi,<br>\nI don't have exact information, as this activity is driven by Kaggle team. But high level - this week they should validate LB results (detection of fake accounts, etc.), then winners will be notified, winners should then have 2 weeks to officially submit their models, then we as hosts should have some time to review the models.</p>",
          "rawMarkdown": "Hi,\nI don't have exact information, as this activity is driven by Kaggle team. But high level - this week they should validate LB results (detection of fake accounts, etc.), then winners will be notified, winners should then have 2 weeks to officially submit their models, then we as hosts should have some time to review the models."
        }
      ]
    },
    {
      "id": 2852005,
      "postDate": "2024-06-03T04:40:19.917Z",
      "content": "<p>thanks for organizing this competition</p>",
      "rawMarkdown": "thanks for organizing this competition"
    },
    {
      "id": 2844865,
      "postDate": "2024-05-30T09:03:47.707Z",
      "content": "<p>Thank you for amazing idea</p>",
      "rawMarkdown": "Thank you for amazing idea"
    },
    {
      "id": 2844857,
      "postDate": "2024-05-30T09:02:00.503Z",
      "content": "<p>Thank you very much for idea </p>",
      "rawMarkdown": "Thank you very much for idea "
    }
  ],
  "comments": [
    {
      "id": 2847359,
      "author_name": "Daniel Herman",
      "author_url": "",
      "post_date": "2024-05-31T14:10:59.097000",
      "content": "<p>The ideas, views and opinions expressed in this comment represent my own views and not those of any of my current or previous employer. Maybe some of you noticed that I went mostly silent after March, there was a reason for that. As of 1.4.2024 I am no longer employed by&nbsp;Home Credit or EmbedIT and my current affiliation is with Similarweb. The decision to change my employer has little to do with this competition, it was done before the start of the competition and it was purely mine.&nbsp;</p>\n<p>To put into perspective what was the amount of time put into the preparations for this competition, it has been exactly a year since we began with the data preparation. We didn't work on this project continuously throughout the year, but if we sum all the bits and pieces it was about several months of work (longer than the whole competition). Every time I read the positive feedback here in Discussion I feel like it was worth the pain and effort. Thank you all for that.</p>\n<p>There was some amount of very frustrated and angry reactions in Discussion. I noticed many people managed to stay professional even despite that and I tried to remain professional as well. I continued reading the posts and comments in the Discussion even after I stopped working for my previous employer. It probably wasn't a healthy thing to do to read all of it. I would really wish if those&nbsp;people had taken the competition less personally and tried to enjoy the challenge that was given to them. I know it is also about winning the money and the fame, but that is not all. The venting usually resolves nothing.&nbsp;Besides that, you can always step aside when you don't like the challenge and forget about the whole competition. <br>\n&nbsp;<br>\nI want to thank you all for participating. I hope most of you enjoyed the challenge and this competition served as a great opportunity for many of you to work on your data science skills. I want to thank also Kaggle for the&nbsp;data part and for helping us to resolve the issues we faced.&nbsp;</p>",
      "votes": 11,
      "replies": []
    },
    {
      "id": 2841301,
      "author_name": "Alessandro Benetti",
      "author_url": "",
      "post_date": "2024-05-28T14:08:17.037000",
      "content": "<p>Hi,</p>\n<p>Thank you for organizing this competition. It was both fun and challenging, especially in terms of modeling and computational constraints.</p>\n<p>I wanted to ask if the full test set will be made available for analysis. It would be fascinating to examine the post-COVID effects on variables like the average days past due.</p>",
      "votes": 9,
      "replies": [
        {
          "id": 2841757,
          "author_name": "Basu Hela",
          "author_url": "",
          "post_date": "2024-05-28T17:20:47.267000",
          "content": "<p>Yeah, I would also like to have access to full test set to check why more than half of my submissions failed!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 2841771,
          "author_name": "Tomas Jelinek",
          "author_url": "",
          "post_date": "2024-05-28T17:28:11.110000",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/alessandrobenetti\" target=\"_blank\">@alessandrobenetti</a>,<br>\npersonally I am okay to share the test set, but let me first check with Kaggle, not sure whether there are not some restrictions (e.g., late submissions will become pointless)</p>",
          "votes": 4,
          "replies": [
            {
              "id": 2841791,
              "author_name": "Alessandro Benetti",
              "author_url": "",
              "post_date": "2024-05-28T17:35:17.263000",
              "content": "<p>Thank you!</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2846072,
              "author_name": "Daniel Herman",
              "author_url": "",
              "post_date": "2024-05-30T21:09:05.623000",
              "content": "<p>That would be great, I think it will be likely used for research in future. </p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2935665,
              "author_name": "Lg",
              "author_url": "",
              "post_date": "2024-07-25T13:37:41.210000",
              "content": "<p>Hi <a href=\"https://www.kaggle.com/jetakow\" target=\"_blank\">@jetakow</a> <a href=\"https://www.kaggle.com/tomasjeline2\" target=\"_blank\">@tomasjeline2</a> , I fully agree that the dataset is incredibly valuable for research. Is it possible to use it for academic research in the future, possibly just the training set? Looking forward to your reply.</p>",
              "votes": 0,
              "replies": []
            }
          ]
        },
        {
          "id": 2846071,
          "author_name": "Daniel Herman",
          "author_url": "",
          "post_date": "2024-05-30T21:08:12.760000",
          "content": "<p>The best reward for me is when I see people dealing with the challenges and enjoying the challenge. It's exactly one year since I started the data preparation. There was a lot of effort put into the preparation of this competition also on the kaggle side and I am happy to see the results of it. </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2841327,
      "author_name": "Laurent Pourchot",
      "author_url": "",
      "post_date": "2024-05-28T14:26:34.717000",
      "content": "<p>While your decision to keep the metric was probably not the best choice, your speed and availability to answer questions were greatly appreciated.</p>",
      "votes": 8,
      "replies": [
        {
          "id": 2841363,
          "author_name": "",
          "author_url": "",
          "post_date": "2024-05-28T14:55:35.990000",
          "content": "",
          "votes": -3,
          "replies": []
        },
        {
          "id": 2846074,
          "author_name": "Daniel Herman",
          "author_url": "",
          "post_date": "2024-05-30T21:09:40.707000",
          "content": "<p>Thanks, I tried my best during the time I worked at Home Credit.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2844176,
      "author_name": "cooperation",
      "author_url": "",
      "post_date": "2024-05-30T01:55:10.887000",
      "content": "<p>Unfortunately, I agree with most of <a href=\"https://www.kaggle.com/at7459\" target=\"_blank\">@at7459</a> points.</p>\n<p>1.For the first time there was a hack, I think the organizers did a good job of changing the data and extending the competition(Whatever the result, the decision was good). When the second time there was a hacker, I was very disappointed that host didn't do anything,  silence was your biggest mistake.</p>\n<p>2.Obviously, top solutions are almost metric hack.Because you promised to punish the Metric Hack, many kaggler didn't choose the public Metric Hack notebook. If you let the metric hack get prize or medal, I think it's a betrayal of us.If you think that the current solution does not have metrik hack, As I just said, silence was your biggest mistake.</p>\n<p>3.I'm glad I didn't invest too much time in the final stage.Although most of Kaggler are volunteers, we can't blame others even if we waste time,but after this competition I believe that many kaggler will be more careful in choosing the competition. Kaggle should give the host more professional advice to avoid the current situation,or kaggle have already done it?</p>",
      "votes": 6,
      "replies": []
    },
    {
      "id": 2842244,
      "author_name": "samuel tan",
      "author_url": "",
      "post_date": "2024-05-29T01:27:09.820000",
      "content": "<p>Hi Tomas!<br>\nThank you very much for organizing this competition, as a practitioner related to the field of banking risk control, I have always wanted to participate in competitions on related topics at kaggle.<br>\nCurrently, both credit risk control and credit card risk control face the problem of instability under the influence of epidemics or economic factors, and the traditional logistic regression scorecard using metrics such as PSI and CSI has certain limitations in stability monitoring.<br>\nThe evaluation metrics of this competition as well as the alternatives provided by other competitors gave me great inspiration, and I will seriously discuss within my team how to embed the above metrics into the model development process.<br>\nLast but not least, I would like to thank you for providing an opportunity to participate in a structured data competition when unstructured data competitions are in vogue, which has improved my personal programming skills and understanding of stability to a certain level.</p>",
      "votes": 3,
      "replies": [
        {
          "id": 2842632,
          "author_name": "Tomas Jelinek",
          "author_url": "",
          "post_date": "2024-05-29T07:35:37.027000",
          "content": "<p>Thank you <a href=\"https://www.kaggle.com/simingtan\" target=\"_blank\">@simingtan</a> and congratulations!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2841919,
      "author_name": "Farukcan Saglam",
      "author_url": "",
      "post_date": "2024-05-28T18:52:29.833000",
      "content": "<p>Leaving aside the problems we encountered, it was the most comprehensive data set that I enjoyed working on</p>",
      "votes": 3,
      "replies": [
        {
          "id": 2842080,
          "author_name": "Varuni Rao",
          "author_url": "",
          "post_date": "2024-05-28T20:49:27.233000",
          "content": "<p>I completely agree. I found the dataset was so different but extensive, than what I have seen in my short time in the field and every moment I spent on the dataset was a new lesson for me.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 2845998,
          "author_name": "Eduardo Toloza",
          "author_url": "",
          "post_date": "2024-05-30T20:00:12.803000",
          "content": "<p>I agree. It broke my mind every time I wanted to understand why some features degraded my score.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 2846079,
          "author_name": "Daniel Herman",
          "author_url": "",
          "post_date": "2024-05-30T21:16:13.890000",
          "content": "<p>I am glad you see that way. There was a lot of energy and many weeks put into the data preparation so the kaggle community doesn't have to deal with it and also we can publish it. To give you some perspective for how long it took, it is exactly year since the whole activity started and I was asked to start preparing the data, we have to stop this activity for several months, but yeah we and kaggle put a lot of energy into this to make it happen. </p>",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 2844113,
      "author_name": "gmobaz",
      "author_url": "",
      "post_date": "2024-05-30T00:03:02.573000",
      "content": "<p><a href=\"https://www.kaggle.com/tomasjeline2\" target=\"_blank\">@tomasjeline2</a> &amp; <a href=\"https://www.kaggle.com/jetakow\" target=\"_blank\">@jetakow</a>, I congratulate you both for your splendid work in good an difficult times. I did not participate, but watched the competition (not so) closely. It really was outstanding  the way you sorted out problems. Congrats again and to all participants!</p>",
      "votes": 2,
      "replies": [
        {
          "id": 2846070,
          "author_name": "Daniel Herman",
          "author_url": "",
          "post_date": "2024-05-30T21:04:16.527000",
          "content": "<p>Thanks, I did my best I could in the given time to help resolve the questions and issues and I tried the best to make this competition enjoyable for everyone. The sad reality is that you can't make everyone happy. </p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 2844908,
      "author_name": "Ramya Rajula",
      "author_url": "",
      "post_date": "2024-05-30T09:33:46.677000",
      "content": "<p>Thank you for organizing this competition. It was both fun and challenging, especially in terms of modeling and computational constraints.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2846081,
          "author_name": "Daniel Herman",
          "author_url": "",
          "post_date": "2024-05-30T21:17:26.467000",
          "content": "<p>It is the best reward for me personally to see people enjoying the challenge and taking it as a learning experience. </p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 2842867,
      "author_name": "Wisp Vale",
      "author_url": "",
      "post_date": "2024-05-29T09:46:21.797000",
      "content": "<p>Hi Tomas, may I ask how you transform the date_decision column? I also found some '_T' columns have a strong correlation with week_num after I aggregate them by date_decision as shown below. But if I don't aggregate them by date, it looks like there will be some fluctuations within each week.<br>\nSo I test the date_decision column of the test dataset for groupby, and then used the aggregated value of the '_T' column for hacking, but found that the LB was significantly worse than the result without the date_decision aggregation. So it seems that after you transform the date_decision column, I can't even use it for aggregation operations. <br>\nDid you add some random noise on it?😂<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F19087078%2F874de32eeed1c44d3a3c8b3865df6819%2F1716975952722.jpg?generation=1716975963601860&amp;alt=media\"><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F19087078%2F55b015c508738fb451d06a81e4c4e155%2FSnipaste_2024-05-29_17-58-59.png?generation=1716976752681331&amp;alt=media\"></p>",
      "votes": 1,
      "replies": [
        {
          "id": 2842942,
          "author_name": "Tomas Jelinek",
          "author_url": "",
          "post_date": "2024-05-29T10:53:00.843000",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/wispvale\" target=\"_blank\">@wispvale</a>,<br>\ndata_decision transformation: we reshuffled it by random number of weeks into the future or history in a way that we keep target rate for each week same as it was in original data sample.<br>\nThen all other date columns were transformed in a same way. It means that all date differences should remain same (e.g., age = date_decision - date of birth). Our big mistake here was that we overlooked these month and year attributes that were as integers in data sample :-(</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 2841794,
      "author_name": "Evgeniia Grigoreva",
      "author_url": "",
      "post_date": "2024-05-28T17:36:34.710000",
      "content": "<p>Could you please share details about how the test set was split into public/private? Did you change it after the hack was found?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2841830,
          "author_name": "Tomas Jelinek",
          "author_url": "",
          "post_date": "2024-05-28T17:50:13.080000",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/eivolkova\" target=\"_blank\">@eivolkova</a>,<br>\nThank you for your active participation in discussions!</p>\n<p>Public/private is 30/70 random split.<br>\nAfter hack / before restart, besides changes we have communicated (hiding week_num etc.), we removed first 30 weeks in private test set, this was not originally planned. Since restart there was no change in test set.</p>",
          "votes": 3,
          "replies": [
            {
              "id": 2842097,
              "author_name": "S J",
              "author_url": "",
              "post_date": "2024-05-28T21:05:28.907000",
              "content": "<p>Did you reveal this to the community: \"we removed first 30 weeks in private test set\"? So just private test data, not public leaderboard data?</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2842618,
              "author_name": "Tomas Jelinek",
              "author_url": "",
              "post_date": "2024-05-29T07:22:47.423000",
              "content": "<p>No, information about 30 weeks in private test was not revealed. </p>\n<p>My opinion is that private test data should cover same problem/use case, but not necessarily same data or period as public test set. Otherwise winning strategy is about overfitting public test. </p>\n<p>We haven't changed public leaderboard data (besides masking column like week_num).</p>",
              "votes": -4,
              "replies": []
            },
            {
              "id": 2842734,
              "author_name": "at7459",
              "author_url": "",
              "post_date": "2024-05-29T08:34:31.477000",
              "content": "<p>no, please do not hold us for stupid, that is not the reason at all why you did it, or you would have done so at the very start of the competition instead of when you realized that your metric does not work.</p>\n<p>instead, this was purely a very naive and hubris-driven way of trying to punish anyone who attempts hack the metric, which obviously did not work out at all and just ended up making the competition exponentially more random than leaving the private test data unchanged.</p>\n<p>instead of focusing on properly transorming the data set, you focused on making the competition as random as possible, and missed all the date columns which are not explicitly in datetime format effectively leaving the date_decision there to see for anyone who looks a bit into the data. which all the more shows that you do not even understand the data you gave to this competition.</p>\n<p>you repeatedly (specifically your colleague Daniel Herman) said that you would not allow a metric hack by way of looking through the top solutions,</p>\n<p><em>\"If the top 100 solutions will be using some way of week num prediction, yes we will go over those hundred solutions in the top private leaderboard until we find a solution that doesn't have this approach \"</em></p>\n<p>and then later in the <strong><em>LAST TWO WEEKS OF THE COMPETITION</em></strong> completely backtracked on that promise, leaving everyone scrambling to probe the leaderboard and find the best way to hack the metric.</p>",
              "votes": 13,
              "replies": []
            },
            {
              "id": 2842779,
              "author_name": "按59",
              "author_url": "",
              "post_date": "2024-05-29T08:58:24.960000",
              "content": "<p><a href=\"https://www.kaggle.com/tomasjeline2\" target=\"_blank\">@tomasjeline2</a> If you delete the data from the top 30 weeks of the private test, the ratio of public to private data will still be 3:7? If it weren't for 3:7, this erroneous message would greatly affect the final result of this competition</p>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 2842876,
              "author_name": "ForcewithMe",
              "author_url": "",
              "post_date": "2024-05-29T09:52:28.700000",
              "content": "<p>In my opinion, if the organizers alter the test data during the competition, particularly the private test data, they must synchronize such changes with all participants.</p>\n<p><strong>Failure to do so could raise some significant honest risk</strong>. Given that the organizers can view all the submissions of all competitors, in future competitions, if hosts want a competitor to win, they can privately change a private dataset that fits the model of the competitor most.</p>",
              "votes": 4,
              "replies": []
            },
            {
              "id": 2842920,
              "author_name": "Tomas Jelinek",
              "author_url": "",
              "post_date": "2024-05-29T10:39:36.797000",
              "content": "<p>Hi <a href=\"https://www.kaggle.com/at7459\" target=\"_blank\">@at7459</a>,<br>\nI don't see any question here, so I don't know what answer do you want from me. <br>\nYou made your frustration from the competition very clear on many other places in the discussion, so I really don't understand what do you want to achieve with your post here.</p>",
              "votes": -5,
              "replies": []
            },
            {
              "id": 2842930,
              "author_name": "Tomas Jelinek",
              "author_url": "",
              "post_date": "2024-05-29T10:43:08.660000",
              "content": "<p>Hi <a href=\"https://www.kaggle.com/forcewithme\" target=\"_blank\">@forcewithme</a>,<br>\nOrganizers can't make changes to test set. In our case it was done after consultation with Kaggle, and restart was necessary, so all submissions were deleted.<br>\nUnresolved problem I am aware of is score of previously submitted notebooks in Code section, as it is not visible whether code was submitted before or after test set changes</p>",
              "votes": -1,
              "replies": []
            },
            {
              "id": 2844038,
              "author_name": "at7459",
              "author_url": "",
              "post_date": "2024-05-29T21:06:39.333000",
              "content": "<p>yeah, i am not sure what i want from you at this point and it's not healthy to get that mad, so i will stop now lol</p>\n<p>but your team literally destroyed the entire competition with all these terrible decisions</p>\n<p>it is not even like these were difficult decisions to get right…</p>\n<p>the most important points are:</p>\n<ol>\n<li>just be clear about what is allowed and what is not allowed</li>\n<li>be very clear about how the test set got transformed and that you removed the first 30 weeks in the private test set</li>\n<li>listen to the people when they tell you how you are messing up and how to do things correctly</li>\n</ol>\n<p>it is not like people are out here to get you, and you need to set traps and hide things from everyone, this is not benficial for either you or the participants of this competition as you can see from the outcome of this mess</p>",
              "votes": 4,
              "replies": []
            },
            {
              "id": 2844070,
              "author_name": "",
              "author_url": "",
              "post_date": "2024-05-29T21:58:01.190000",
              "content": "",
              "votes": 4,
              "replies": []
            },
            {
              "id": 2844620,
              "author_name": "Laurent Pourchot",
              "author_url": "",
              "post_date": "2024-05-30T06:52:33.770000",
              "content": "<p>Hi <a href=\"https://www.kaggle.com/tomasjeline2\" target=\"_blank\">@tomasjeline2</a> Most of us are frustrated by your changes at the end of the competition, and I'm sure you understand. But we should not be rude because the data were very interesting and the competition should be taken as it is, only a game and a sharing of ideas. Do you intend to come back later with a new metric that would be more resilient ?</p>",
              "votes": 2,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2841770,
      "author_name": "gromml",
      "author_url": "",
      "post_date": "2024-05-28T17:28:09.540000",
      "content": "<p>Hi Tomas!</p>\n<p>First of all, thanks for hosting this competition. We believe that any result is also a result for the organizers, and they can make some useful conclusions from the competition.</p>\n<p>Our team have one question about dates. We tried to restore <code>WEEK_NUM</code>, and during this process we discovered that the following code successfully produces a submission file for the test set:</p>\n<pre><code>t = (- / ) * df_test[]\n\n t.shape[] *  &lt; t.isna().() &lt; t.shape[] * \n  &lt; t.quantile() &lt; \n  &lt;= t.quantile() &lt;= \n\nt1 = t[t &lt; ]\nt2 = t[(t &gt;= ) &amp; (t &lt;= )]\n\n - &lt; t1.() &lt; \n  &lt; t1.() &lt; \n\n  &lt; t1.shape[] / t.shape[] &lt; \n  &lt; t2.shape[] / t.shape[] &lt; \n</code></pre>\n<p>It seems that a significant part of the restored week nums are less than 30 (remember that the training period contains week nums between 0 and 91) or even negative. Could you tell us why it happened? :)</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2841782,
          "author_name": "Evgeniia Grigoreva",
          "author_url": "",
          "post_date": "2024-05-28T17:31:41.580000",
          "content": "<p>It can be due to the fact that this method is not perfect. For some case_ids refreshdate can be different from 2019/1/3</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2841801,
              "author_name": "gromml",
              "author_url": "",
              "post_date": "2024-05-28T17:38:32.103000",
              "content": "<p>But for the training set it works muuuch better. Almost all week_nums (~90%, primarily due to NaNs in the <code>min_refreshdate_3813885D</code>, otherwise it would be almost 100%) are accurately restored. In the test set only ~50% are restored. Also, negative restored week_nums was some kind of surprise.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2841806,
              "author_name": "Evgeniia Grigoreva",
              "author_url": "",
              "post_date": "2024-05-28T17:41:00.757000",
              "content": "<p>Yeah, I thought that the reason may be that 2019/1/3 was some kind of a filler for missing values, and the more data we have, the less number of missing values. Another option is that, for example, from year 2022 the filler value changed to i.e. 2022/1/1</p>\n<p>We can check this if the test set is published:)</p>",
              "votes": 0,
              "replies": []
            }
          ]
        },
        {
          "id": 2841816,
          "author_name": "Tomas Jelinek",
          "author_url": "",
          "post_date": "2024-05-28T17:43:53.230000",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/gromml\" target=\"_blank\">@gromml</a>,<br>\nwe will have to look into why min_refreshdate_3813885D works in week restoration.</p>\n<p>But to answer your question, test sample contains only weeks after train (91+), so there was no overlap with train and all test cases happen after train period.</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2843261,
              "author_name": "Tomas Jelinek",
              "author_url": "",
              "post_date": "2024-05-29T13:35:00.433000",
              "content": "<p>Hi <a href=\"https://www.kaggle.com/gromml\" target=\"_blank\">@gromml</a> again,<br>\nquick update regarding min_refreshdate_3813885D, our hypothesis is following: this is credit bureau attribute, contains (probably) dates when client's report was updated. And for some reason data provider assigned 1/2019 to all clients who had CB report before that date. Once you discovered that, it's straightforward how to reconstruct date_decision, because same transformation was applied…</p>\n<p>Some clients apply for first loan and don't have previous CB report - you have missing value. And some clients have first credit after 1/2019 - that's why you have different value than 1/2019. Share of such clients is increasing in time.</p>",
              "votes": 5,
              "replies": []
            }
          ]
        },
        {
          "id": 2842270,
          "author_name": "Evan",
          "author_url": "",
          "post_date": "2024-05-29T02:21:13.193000",
          "content": "<p>According to my experiments, the min_refreshdate_3813885D keeps decresing until around -1100, but goes up even to be larger than -640 (the minimum in train), and then decresing again. My guess is the refreshdate_3813885D has been updated at some point, it does not always be 2019/1. </p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 2845751,
      "author_name": "Mz",
      "author_url": "",
      "post_date": "2024-05-30T16:58:48.770000",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/tomasjeline2\" target=\"_blank\">@tomasjeline2</a>, Could you explain the meaning of 'price_1097A'? I'm not sure what is 'Credit price', does it means something like interest?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2852171,
          "author_name": "Tomas Jelinek",
          "author_url": "",
          "post_date": "2024-06-03T07:00:35.030000",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/polsyuus\" target=\"_blank\">@polsyuus</a>,<br>\nI have checked with underwriting team, according to them it is price of goods that is financed by credit.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2842296,
      "author_name": "GodGod3",
      "author_url": "",
      "post_date": "2024-05-29T03:10:04.470000",
      "content": "<p>Hi,Thomas,<br>\nI wonder in practice how Home Credit deals with this model stability issue. Do you train on a rolling basis with new data every time to make the model as effective and can you explain further how the model is used ,for credit approval or screening?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2842662,
          "author_name": "Tomas Jelinek",
          "author_url": "",
          "post_date": "2024-05-29T07:57:30.693000",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/godgod3\" target=\"_blank\">@godgod3</a>,</p>\n<p>Regarding the stability, in credit risk scorecards there is issue with delay of target observability - you can't say whether client is good/bad based on payment of first installment, so you have to observe first few payments, and delay can be half a year or even more, plus time for model development, implementation, adjustment of underwriting strategies, etc. So continuous re-development does not solve the stability problem, you always train on \"outdated\" data. Different story is for other use cases, for example sales, where you can observe results almost immediately.<br>\nWe check stability of predictors, both from perspective of population structure, and relationship feature-target. But decision what is acceptable threshold for stability is expert-based. And before the competition we were thinking whether there is better approach…</p>\n<p>Regarding model usage. First of all, I read some concerns in discussion regarding bias/discrimination of models, or their interpretability - we've never intended to use Kaggle models in production. Purpose of this competition is to learn new approaches or techniques we can learn from and to improve our model development process.<br>\nHow model is used - for approval, and it can be actual credit application, but also credit offer (I guess this is what you meant by screening). Model is not the only tool used in underwriting, there are many other conditions for approval (e.g. legal/regulatory - min age of applicant &gt; X, affordability - ratio if disposable income and instalment payment &gt; Y, etc.). Besides approval model decides quality of client, which has impact on credit parameters (typically better client = higher limit / credit amount, or lower interest)</p>",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 2920846,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-07-13T20:47:37.850000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2852368,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-06-03T08:41:21.957000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2842394,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-05-29T04:22:09.520000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 2842685,
          "author_name": "Tomas Jelinek",
          "author_url": "",
          "post_date": "2024-05-29T08:04:29.840000",
          "content": "<p>Hi,<br>\nI don't have exact information, as this activity is driven by Kaggle team. But high level - this week they should validate LB results (detection of fake accounts, etc.), then winners will be notified, winners should then have 2 weeks to officially submit their models, then we as hosts should have some time to review the models.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2852005,
      "author_name": "Mubashir Ahmed Siddiqui",
      "author_url": "",
      "post_date": "2024-06-03T04:40:19.917000",
      "content": "<p>thanks for organizing this competition</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2844865,
      "author_name": "Nipanshu .K. Patel",
      "author_url": "",
      "post_date": "2024-05-30T09:03:47.707000",
      "content": "<p>Thank you for amazing idea</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2844857,
      "author_name": "Nipanshu .K. Patel",
      "author_url": "",
      "post_date": "2024-05-30T09:02:00.503000",
      "content": "<p>Thank you very much for idea </p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2841246": "Dear Competitors,\n \nAs we approach the conclusion of the Home Credit - Credit Risk Model Stability competition, we want to extend our gratitude to all enthusiastic participants who contributed to this competition with their hard work. We are looking forward to dive deep into your solutions and explore all the innovative approaches tackling the credit risk model.\n \nWe hope that this competition has provided valuable learning opportunities and insights for all participants. Whether you're a seasoned data scientist or new to the field, we hope you mastered new skills and gained knowledge that will benefit your future endeavors. It has also been inspiring to witness the collaborative spirit within the Kaggle community, where members support each other and share valuable insights.\n \nThe topic of model stability has proven to be challenging, and we've all learned a great deal throughout this competition. The project brought moments of joy as well as periods of painful struggle, and we appreciate your perseverance and commitment to tackling these complex issues.\n \nAs the competition comes to an end, the Kaggle team will finalize the standings and share them with you through a separate communication. Please stay tuned for updates on the timeline for these announcements.\n \nThroughout the competition, as hosts, we exercised caution in responding to queries to prevent unintended advantages (e.g., disclosing differences between the private and public test sets). If you still have unanswered questions, please don't hesitate to ask them here.\n \nOnce again, thank you for your enthusiastic participation. A big congratulations to the winning teams! We eagerly anticipate reviewing your final submissions and exploring the innovative strategies employed by the top performers.\n \nOn behalf of the entire Home Credit R&D team,\n \nTomas Jelinek",
    "2847359": "The ideas, views and opinions expressed in this comment represent my own views and not those of any of my current or previous employer. Maybe some of you noticed that I went mostly silent after March, there was a reason for that. As of 1.4.2024 I am no longer employed by Home Credit or EmbedIT and my current affiliation is with Similarweb. The decision to change my employer has little to do with this competition, it was done before the start of the competition and it was purely mine. \n\nTo put into perspective what was the amount of time put into the preparations for this competition, it has been exactly a year since we began with the data preparation. We didn't work on this project continuously throughout the year, but if we sum all the bits and pieces it was about several months of work (longer than the whole competition). Every time I read the positive feedback here in Discussion I feel like it was worth the pain and effort. Thank you all for that.\n\nThere was some amount of very frustrated and angry reactions in Discussion. I noticed many people managed to stay professional even despite that and I tried to remain professional as well. I continued reading the posts and comments in the Discussion even after I stopped working for my previous employer. It probably wasn't a healthy thing to do to read all of it. I would really wish if those people had taken the competition less personally and tried to enjoy the challenge that was given to them. I know it is also about winning the money and the fame, but that is not all. The venting usually resolves nothing. Besides that, you can always step aside when you don't like the challenge and forget about the whole competition. \n \nI want to thank you all for participating. I hope most of you enjoyed the challenge and this competition served as a great opportunity for many of you to work on your data science skills. I want to thank also Kaggle for the data part and for helping us to resolve the issues we faced. ",
    "2841301": "Hi,\n\nThank you for organizing this competition. It was both fun and challenging, especially in terms of modeling and computational constraints.\n\nI wanted to ask if the full test set will be made available for analysis. It would be fascinating to examine the post-COVID effects on variables like the average days past due.",
    "2841327": "While your decision to keep the metric was probably not the best choice, your speed and availability to answer questions were greatly appreciated.",
    "2844176": "Unfortunately, I agree with most of @at7459 points.\n\n1.For the first time there was a hack, I think the organizers did a good job of changing the data and extending the competition(Whatever the result, the decision was good). When the second time there was a hacker, I was very disappointed that host didn't do anything,  silence was your biggest mistake.\n\n2.Obviously, top solutions are almost metric hack.Because you promised to punish the Metric Hack, many kaggler didn't choose the public Metric Hack notebook. If you let the metric hack get prize or medal, I think it's a betrayal of us.If you think that the current solution does not have metrik hack, As I just said, silence was your biggest mistake.\n\n3.I'm glad I didn't invest too much time in the final stage.Although most of Kaggler are volunteers, we can't blame others even if we waste time,but after this competition I believe that many kaggler will be more careful in choosing the competition. Kaggle should give the host more professional advice to avoid the current situation,or kaggle have already done it?\n",
    "2842244": "Hi Tomas!\nThank you very much for organizing this competition, as a practitioner related to the field of banking risk control, I have always wanted to participate in competitions on related topics at kaggle.\nCurrently, both credit risk control and credit card risk control face the problem of instability under the influence of epidemics or economic factors, and the traditional logistic regression scorecard using metrics such as PSI and CSI has certain limitations in stability monitoring.\nThe evaluation metrics of this competition as well as the alternatives provided by other competitors gave me great inspiration, and I will seriously discuss within my team how to embed the above metrics into the model development process.\nLast but not least, I would like to thank you for providing an opportunity to participate in a structured data competition when unstructured data competitions are in vogue, which has improved my personal programming skills and understanding of stability to a certain level.",
    "2841919": "Leaving aside the problems we encountered, it was the most comprehensive data set that I enjoyed working on",
    "2844113": "@tomasjeline2 & @jetakow, I congratulate you both for your splendid work in good an difficult times. I did not participate, but watched the competition (not so) closely. It really was outstanding  the way you sorted out problems. Congrats again and to all participants!",
    "2844908": "Thank you for organizing this competition. It was both fun and challenging, especially in terms of modeling and computational constraints.",
    "2842867": "Hi Tomas, may I ask how you transform the date_decision column? I also found some '_T' columns have a strong correlation with week_num after I aggregate them by date_decision as shown below. But if I don't aggregate them by date, it looks like there will be some fluctuations within each week.\nSo I test the date_decision column of the test dataset for groupby, and then used the aggregated value of the '_T' column for hacking, but found that the LB was significantly worse than the result without the date_decision aggregation. So it seems that after you transform the date_decision column, I can't even use it for aggregation operations. \nDid you add some random noise on it?😂\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F19087078%2F874de32eeed1c44d3a3c8b3865df6819%2F1716975952722.jpg?generation=1716975963601860&alt=media)![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F19087078%2F55b015c508738fb451d06a81e4c4e155%2FSnipaste_2024-05-29_17-58-59.png?generation=1716976752681331&alt=media)",
    "2841794": "Could you please share details about how the test set was split into public/private? Did you change it after the hack was found?",
    "2841770": "Hi Tomas!\n\nFirst of all, thanks for hosting this competition. We believe that any result is also a result for the organizers, and they can make some useful conclusions from the competition.\n\nOur team have one question about dates. We tried to restore `WEEK_NUM`, and during this process we discovered that the following code successfully produces a submission file for the test set:\n\n```python\nt = (-1.0 / 7.0) * df_test['min_refreshdate_3813885D']\n\nassert t.shape[0] * 0.15 < t.isna().sum() < t.shape[0] * 0.25\nassert 0.4 < t.quantile(0.1) < 25.0\nassert 115.0 <= t.quantile(0.9) <= 140.0\n        \nt1 = t[t < 91.0]\nt2 = t[(t >= 91.0) & (t <= 135.0)]\n\nassert -30.0 < t1.min() < 0.0\nassert 0.0 < t1.max() < 30.0\n\nassert 0.2 < t1.shape[0] / t.shape[0] < 0.4\nassert 0.4 < t2.shape[0] / t.shape[0] < 0.5\n```\n\nIt seems that a significant part of the restored week nums are less than 30 (remember that the training period contains week nums between 0 and 91) or even negative. Could you tell us why it happened? :)",
    "2845751": "Hi @tomasjeline2, Could you explain the meaning of 'price_1097A'? I'm not sure what is 'Credit price', does it means something like interest?",
    "2842296": "Hi,Thomas,\nI wonder in practice how Home Credit deals with this model stability issue. Do you train on a rolling basis with new data every time to make the model as effective and can you explain further how the model is used ,for credit approval or screening?",
    "2920846": "",
    "2852368": "",
    "2842394": "",
    "2852005": "thanks for organizing this competition",
    "2844865": "Thank you for amazing idea",
    "2844857": "Thank you very much for idea "
  }
}