{
  "id": 482474,
  "title": "We are back - Submissions open on Monday, 11th March 2024",
  "url": "/competitions/home-credit-credit-risk-model-stability/discussion/482474",
  "author_name": "Daniel Herman",
  "post_date": "2024-03-07T21:43:08.211000",
  "votes": 96,
  "comment_count": 125,
  "views": 0,
  "content": "<p>Dear kagglers,</p>\n<p>as promised we are back with an update. During the last two weeks, the submissions were paused due to metric exploitation and we decided to make changes to prevent participants from gaining a huge advantage by using this hack, thereby steering the focus of Kagglers back to the intended topic, which is the long-term stability of the model.</p>\n<p>As we indicated previously, we will be changing date columns in test_base table, namely date_decision, MONTH and WEEK_NUM. The MONTH and WEEK_NUM columns will now only have one constant value. We decided to keep them to minimize the amount of time you'll need to spend altering your scripts to get the submissions working again. Column date_decision is no longer the same. The remaining data were also transformed. Don't worry; the transformation will preserve the meaning of features. Data types will remain the same, and there shouldn't be significant changes in the distributions. If you are aiming to develop a stable well-performing model, this change should not concern you at all. </p>\n<p>We discussed metric change internally and with kagglers here. We received many great ideas and we are foremost very thankful for your time and showed passion for the challenge. We appreciate the effort that the community is willing to make for the sake of a more fair metric within this competition. The stability metric currently has an issue: it is possible to artificially worsen the score on some observations and yet improve the overall stability score. This might seem undesirable or strange to many of you, but there are reasons for this behaviour. Ultimately, we don't want participants to focus on metric hacking; however, we also want a metric that closely approximates how models are assessed in production. When looking at the performance of models in production the predictable stable decrease in performance is undesirable even if it is not yet observed, but expected to come. We didn’t find a metric that would satisfy both requirements and that is why we decided to keep the metric as it is now.</p>\n<p>The fix is prepared and will be applied to the competition on Monday, 11th March 2024 together with resuming submissions. Consequently, the entire competition will be extended by three weeks. We want to encourage you again to focus on the stability topic. We wish everyone good luck with the data exploration, modelling and submitting results. </p>\n<p>Tomas and Daniel</p>",
  "messages": [
    {
      "id": 2686523,
      "postDate": "2024-03-07T21:43:08.210Z",
      "content": "<p>Dear kagglers,</p>\n<p>as promised we are back with an update. During the last two weeks, the submissions were paused due to metric exploitation and we decided to make changes to prevent participants from gaining a huge advantage by using this hack, thereby steering the focus of Kagglers back to the intended topic, which is the long-term stability of the model.</p>\n<p>As we indicated previously, we will be changing date columns in test_base table, namely date_decision, MONTH and WEEK_NUM. The MONTH and WEEK_NUM columns will now only have one constant value. We decided to keep them to minimize the amount of time you'll need to spend altering your scripts to get the submissions working again. Column date_decision is no longer the same. The remaining data were also transformed. Don't worry; the transformation will preserve the meaning of features. Data types will remain the same, and there shouldn't be significant changes in the distributions. If you are aiming to develop a stable well-performing model, this change should not concern you at all. </p>\n<p>We discussed metric change internally and with kagglers here. We received many great ideas and we are foremost very thankful for your time and showed passion for the challenge. We appreciate the effort that the community is willing to make for the sake of a more fair metric within this competition. The stability metric currently has an issue: it is possible to artificially worsen the score on some observations and yet improve the overall stability score. This might seem undesirable or strange to many of you, but there are reasons for this behaviour. Ultimately, we don't want participants to focus on metric hacking; however, we also want a metric that closely approximates how models are assessed in production. When looking at the performance of models in production the predictable stable decrease in performance is undesirable even if it is not yet observed, but expected to come. We didn’t find a metric that would satisfy both requirements and that is why we decided to keep the metric as it is now.</p>\n<p>The fix is prepared and will be applied to the competition on Monday, 11th March 2024 together with resuming submissions. Consequently, the entire competition will be extended by three weeks. We want to encourage you again to focus on the stability topic. We wish everyone good luck with the data exploration, modelling and submitting results. </p>\n<p>Tomas and Daniel</p>",
      "rawMarkdown": "Dear kagglers,\n\nas promised we are back with an update. During the last two weeks, the submissions were paused due to metric exploitation and we decided to make changes to prevent participants from gaining a huge advantage by using this hack, thereby steering the focus of Kagglers back to the intended topic, which is the long-term stability of the model.\n\nAs we indicated previously, we will be changing date columns in test_base table, namely date_decision, MONTH and WEEK_NUM. The MONTH and WEEK_NUM columns will now only have one constant value. We decided to keep them to minimize the amount of time you'll need to spend altering your scripts to get the submissions working again. Column date_decision is no longer the same. The remaining data were also transformed. Don't worry; the transformation will preserve the meaning of features. Data types will remain the same, and there shouldn't be significant changes in the distributions. If you are aiming to develop a stable well-performing model, this change should not concern you at all. \n\nWe discussed metric change internally and with kagglers here. We received many great ideas and we are foremost very thankful for your time and showed passion for the challenge. We appreciate the effort that the community is willing to make for the sake of a more fair metric within this competition. The stability metric currently has an issue: it is possible to artificially worsen the score on some observations and yet improve the overall stability score. This might seem undesirable or strange to many of you, but there are reasons for this behaviour. Ultimately, we don't want participants to focus on metric hacking; however, we also want a metric that closely approximates how models are assessed in production. When looking at the performance of models in production the predictable stable decrease in performance is undesirable even if it is not yet observed, but expected to come. We didn’t find a metric that would satisfy both requirements and that is why we decided to keep the metric as it is now.\n\nThe fix is prepared and will be applied to the competition on Monday, 11th March 2024 together with resuming submissions. Consequently, the entire competition will be extended by three weeks. We want to encourage you again to focus on the stability topic. We wish everyone good luck with the data exploration, modelling and submitting results. \n\nTomas and Daniel\n",
      "votes": 96
    },
    {
      "id": 2689496,
      "postDate": "2024-03-09T21:42:39.613Z",
      "content": "<p>I don't see the concept of manually checking submissions for hacking being tenable.  The metric is still hackable, and the gains from hacking seem to be vastly bigger than the gains from legitimate model development.  I think a lot of competitors will call the bluff of the organizers, because if the choice is between having a submission in the top 100 and hoping it won't be disqualified, and having an uncompetitive legitimate solution, it's not really a choice at all.  You can't have a good Kaggle competition with a bad metric, no matter how many kludges you try to patch it with.</p>",
      "rawMarkdown": "I don't see the concept of manually checking submissions for hacking being tenable.  The metric is still hackable, and the gains from hacking seem to be vastly bigger than the gains from legitimate model development.  I think a lot of competitors will call the bluff of the organizers, because if the choice is between having a submission in the top 100 and hoping it won't be disqualified, and having an uncompetitive legitimate solution, it's not really a choice at all.  You can't have a good Kaggle competition with a bad metric, no matter how many kludges you try to patch it with.",
      "votes": 19,
      "replies": [
        {
          "id": 2690198,
          "postDate": "2024-03-10T11:22:06.420Z",
          "content": "<p>With all due respect, I am tired of the discussion regarding metric. I don't think this topic deserves all that attention. We considered changes in metric, but we didn't find an option that would not deviate from our initial intentions. Saying the metric is bad is little of an overstatement as it approximates well the business decisions. We encourage you to contribute to this competition. With that said this decision of keeping the metric is final.</p>",
          "rawMarkdown": "With all due respect, I am tired of the discussion regarding metric. I don't think this topic deserves all that attention. We considered changes in metric, but we didn't find an option that would not deviate from our initial intentions. Saying the metric is bad is little of an overstatement as it approximates well the business decisions. We encourage you to contribute to this competition. With that said this decision of keeping the metric is final.",
          "votes": -25,
          "replies": [
            {
              "id": 2691057,
              "postDate": "2024-03-11T02:32:29.593Z",
              "content": "<p>With all due respect, I'm sure you also thought the metric was fine before it was hacked the first time. Practices that work well enough in production aren't always suitable in competitions, and those of us who have experience with kaggle have seen competitions hacked that didn't have half the glaring vulnerabilities that this one does.  </p>",
              "rawMarkdown": "With all due respect, I'm sure you also thought the metric was fine before it was hacked the first time. Practices that work well enough in production aren't always suitable in competitions, and those of us who have experience with kaggle have seen competitions hacked that didn't have half the glaring vulnerabilities that this one does.  ",
              "votes": 22
            },
            {
              "id": 2692139,
              "postDate": "2024-03-11T17:46:38.633Z",
              "content": "<p>i seriously don't understand how it could happen that the metric still stays like it is although for everyone with experience in kaggle its clear what the outcome will be and there was provably unhackable metric suggested. for me this competition is over.</p>",
              "rawMarkdown": "i seriously don't understand how it could happen that the metric still stays like it is although for everyone with experience in kaggle its clear what the outcome will be and there was provably unhackable metric suggested. for me this competition is over.",
              "votes": 9
            }
          ]
        }
      ]
    },
    {
      "id": 2704098,
      "postDate": "2024-03-18T15:01:23.410Z",
      "content": "<p>This info should be in data description as the new participants that didn't read the discussions to know about this, otherwise they can use date_decision, MONTH and WEEK_NUM without a clue that in the test these are modified.</p>",
      "rawMarkdown": "This info should be in data description as the new participants that didn't read the discussions to know about this, otherwise they can use date_decision, MONTH and WEEK_NUM without a clue that in the test these are modified.",
      "votes": 13,
      "replies": [
        {
          "id": 2705816,
          "postDate": "2024-03-19T14:39:13.807Z",
          "content": "<p>Thank you for noticing that, I will add it there. </p>",
          "rawMarkdown": "Thank you for noticing that, I will add it there. ",
          "votes": 1
        }
      ]
    },
    {
      "id": 2715402,
      "postDate": "2024-03-25T13:27:44.610Z",
      "content": "<p>Hi, I'm interested about these tables and columns:</p>\n<pre><code>{\n        : [\n            ,\n            ,\n            ,\n            ,\n            ,\n            ,\n            ,\n            ,\n        ],\n        : [\n            ,\n            ,\n            ,\n            ,\n        ],\n        : [\n            ,\n            ,\n            ,\n            ,\n        ],\n    }\n</code></pre>\n<p>These columns also indicate date, although they are not marked as such. Were they transformed also? If yes, are time deltas like <code>dpdmaxdateyear_742T</code> - <code>date_decision</code> preserved?</p>",
      "rawMarkdown": "Hi, I'm interested about these tables and columns:\n\n```\n{\n        \"credit_bureau_a_1\": [\n            \"dpdmaxdatemonth_442T\",\n            \"dpdmaxdatemonth_89T\",\n            \"dpdmaxdateyear_596T\",\n            \"dpdmaxdateyear_896T\",\n            \"overdueamountmaxdatemonth_284T\",\n            \"overdueamountmaxdatemonth_365T\",\n            \"overdueamountmaxdateyear_2T\",\n            \"overdueamountmaxdateyear_994T\",\n        ],\n        \"credit_bureau_b_1\": [\n            \"dpdmaxdatemonth_804T\",\n            \"dpdmaxdateyear_742T\",\n            \"overdueamountmaxdatemonth_494T\",\n            \"overdueamountmaxdateyear_432T\",\n        ],\n        \"credit_bureau_a_2\": [\n            \"pmts_month_158T\",\n            \"pmts_month_706T\",\n            \"pmts_year_1139T\",\n            \"pmts_year_507T\",\n        ],\n    }\n```\n\nThese columns also indicate date, although they are not marked as such. Were they transformed also? If yes, are time deltas like `dpdmaxdateyear_742T` - `date_decision` preserved?",
      "votes": 8,
      "replies": [
        {
          "id": 2726414,
          "postDate": "2024-04-01T07:13:47.630Z",
          "content": "<p><a href=\"https://www.kaggle.com/jetakow\" target=\"_blank\">@jetakow</a> <a href=\"https://www.kaggle.com/tomasjeline2\" target=\"_blank\">@tomasjeline2</a> can anyone answer this?</p>",
          "rawMarkdown": "@jetakow @tomasjeline2 can anyone answer this?",
          "votes": 4,
          "replies": [
            {
              "id": 2830136,
              "postDate": "2024-05-23T03:01:31.737Z",
              "content": "<p><a href=\"https://www.kaggle.com/jetakow\" target=\"_blank\">@jetakow</a> <a href=\"https://www.kaggle.com/tomasjeline2\" target=\"_blank\">@tomasjeline2</a> already answered this somewhere?</p>",
              "rawMarkdown": "@jetakow @tomasjeline2 already answered this somewhere?"
            },
            {
              "id": 2832872,
              "postDate": "2024-05-23T22:58:04.880Z",
              "content": "<p>Please tag active competition hosts <a href=\"https://www.kaggle.com/tomasjeline2\" target=\"_blank\">@tomasjeline2</a> <a href=\"https://www.kaggle.com/vojtechpacak\" target=\"_blank\">@vojtechpacak</a> </p>",
              "rawMarkdown": "Please tag active competition hosts @tomasjeline2 @vojtechpacak "
            }
          ]
        }
      ]
    },
    {
      "id": 2688147,
      "postDate": "2024-03-09T03:43:13.823Z",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/jetakow\" target=\"_blank\">@jetakow</a> </p>\n<p>I'm really happy to hear that the competition will be coming back.<br>\nI appreciate the proactive replacement of the test data by the host.<br>\nHowever, I think there are still two problems with not changing the current metric.</p>\n<h2>1. Indirect hacking is possible</h2>\n<p>Even if direct hacking is not possible with the correction of the test data, it doesn't change the fact that this metric is hackable. <br>\nFor example, one could create a model to predict <code>NUM_WEEK</code> in test data and intentionally worsen the scores of smaller <code>NUM_WEEK</code>. <br>\nI don't know whether such hacking is truly possible, but most participants who are genuinely improving their models in legitimate ways would feel uneasy about continuing to use this metric.</p>\n<h2>2. Problems with the metric itself</h2>\n<p>The image in the following post by <a href=\"https://www.kaggle.com/chumajin\" target=\"_blank\">@chumajin</a> is easy to explain, which is also attached:</p>\n<p><a href=\"https://www.kaggle.com/competitions/home-credit-credit-risk-model-stability/discussion/476867#2651259\" target=\"_blank\">https://www.kaggle.com/competitions/home-credit-credit-risk-model-stability/discussion/476867#2651259</a></p>\n<p>Comparing case2 and case3, it's clear that case3 is the superior model, yet under the current metric, the winner would be case2. I think this is not what the host truly wants to get.</p>\n<p>I feel that the current metric is too focused on stability, neglecting the original goal of improving AUC.</p>\n<p>Therefore, honestly, I'm concerned about these issues.</p>",
      "rawMarkdown": "Hi @jetakow \n\nI'm really happy to hear that the competition will be coming back.\nI appreciate the proactive replacement of the test data by the host.\nHowever, I think there are still two problems with not changing the current metric.\n\n## 1. Indirect hacking is possible\n\nEven if direct hacking is not possible with the correction of the test data, it doesn't change the fact that this metric is hackable. \nFor example, one could create a model to predict `NUM_WEEK` in test data and intentionally worsen the scores of smaller `NUM_WEEK`. \nI don't know whether such hacking is truly possible, but most participants who are genuinely improving their models in legitimate ways would feel uneasy about continuing to use this metric.\n\n## 2. Problems with the metric itself\nThe image in the following post by @chumajin is easy to explain, which is also attached:\n\nhttps://www.kaggle.com/competitions/home-credit-credit-risk-model-stability/discussion/476867#2651259\n\nComparing case2 and case3, it's clear that case3 is the superior model, yet under the current metric, the winner would be case2. I think this is not what the host truly wants to get.\n\nI feel that the current metric is too focused on stability, neglecting the original goal of improving AUC.\n\nTherefore, honestly, I'm concerned about these issues.",
      "votes": 8,
      "replies": [
        {
          "id": 2688708,
          "postDate": "2024-03-09T12:12:41.833Z",
          "content": "<blockquote>\n  <p>Even if direct hacking is not possible with the correction of the test data, it doesn't change the fact that this metric is hackable.</p>\n</blockquote>\n<p>This is valid point, what you may consider is that there is limited time in this competition and as of now, the hacking is no longer feasible from a time-use stance. Besides that we discourage direct hacking and we will be reviewing the potential winners' work. </p>\n<blockquote>\n  <p>I feel that the current metric is too focused on stability, neglecting the original goal of improving AUC.</p>\n</blockquote>\n<p>Maybe we didn't communicate it clearly in previous posts, but the focus on stability is intended. Exploring the stability in model preparation is the target focus in this competition and it has a great importance in Home Credit. In fact, it is so important we decided to stick to the current metric because it covers the near-term performance extrapolation. </p>",
          "rawMarkdown": ">Even if direct hacking is not possible with the correction of the test data, it doesn't change the fact that this metric is hackable.\n\nThis is valid point, what you may consider is that there is limited time in this competition and as of now, the hacking is no longer feasible from a time-use stance. Besides that we discourage direct hacking and we will be reviewing the potential winners' work. \n\n>I feel that the current metric is too focused on stability, neglecting the original goal of improving AUC.\n\nMaybe we didn't communicate it clearly in previous posts, but the focus on stability is intended. Exploring the stability in model preparation is the target focus in this competition and it has a great importance in Home Credit. In fact, it is so important we decided to stick to the current metric because it covers the near-term performance extrapolation. ",
          "votes": -5,
          "replies": [
            {
              "id": 2688785,
              "postDate": "2024-03-09T13:12:34.880Z",
              "content": "<p><a href=\"https://www.kaggle.com/jetakow\" target=\"_blank\">@jetakow</a> <br>\nThank you for your reply.<br>\nWhat you explained makes sense to me.<br>\nLet me ask one more question.</p>\n<pre><code>we will be reviewing the potential winners' work.\n</code></pre>\n<p>If you review winner's solution and find out that the solution was hacked or gamed, will they lose their prize? or in the worst case, will they be banned?</p>",
              "rawMarkdown": "@jetakow \nThank you for your reply.\nWhat you explained makes sense to me.\nLet me ask one more question.\n\n```\nwe will be reviewing the potential winners' work.\n```\n\nIf you review winner's solution and find out that the solution was hacked or gamed, will they lose their prize? or in the worst case, will they be banned?"
            },
            {
              "id": 2688805,
              "postDate": "2024-03-09T13:30:18.137Z",
              "content": "<p>We need to discuss that internally, but I think the most likely scenario as of now it that they won't be able to win the prize if the winning is based on the metric hack. </p>",
              "rawMarkdown": "We need to discuss that internally, but I think the most likely scenario as of now it that they won't be able to win the prize if the winning is based on the metric hack. "
            },
            {
              "id": 2688810,
              "postDate": "2024-03-09T13:41:27.210Z",
              "content": "<p>Okay.<br>\nI hope all the potential winners will follow the legitimate way to present a great solution.<br>\nAnd of course I will enjoy this competition as well.<br>\nThank you.</p>",
              "rawMarkdown": "Okay.\nI hope all the potential winners will follow the legitimate way to present a great solution.\nAnd of course I will enjoy this competition as well.\nThank you."
            },
            {
              "id": 2691381,
              "postDate": "2024-03-11T07:49:18.513Z",
              "rawMarkdown": "",
              "isDeleted": true
            },
            {
              "id": 2748434,
              "postDate": "2024-04-12T12:51:49.763Z",
              "rawMarkdown": "",
              "isDeleted": true
            }
          ]
        }
      ]
    },
    {
      "id": 2693302,
      "postDate": "2024-03-12T10:30:57.250Z",
      "content": "<p>Dear Kagglers,</p>\n<p>After many days our competition is running again. I wish you all success and want to thank you again for your patience.</p>\n<p>There were many comments lately with suggestions how the competition, metric, data, etc. should be designed or changed. I am thankful for many of your suggestions - really last weeks were very inspiring and at the same time very challenging for us. I feel I have to address some of your comments…</p>\n<p>When we were preparing this competition we knew from the beginning that we don't want to just repeat Home Credit competition from 2018. We have decided to add stability into competition objectives. Why? Because we think we optimize stability let's say expertly in our corporation and that there should be better way how to address this issue.</p>\n<p>We understood that our metric was more complex than plain vanilla AUC and there is risk that Kaggle community will find approaches how to optimize it in a way we haven't thought about. Unfortunately this risk materialized and we decided to react. There were basically 2 options: a) change metric so it's not hackable or b) change data so it does not contain information needed for hack. We have investigated many potential metrics (thank you Kaggle community for proposals), however we were not able to find a metric that is good enough to describe \"good model\", moreover many metrics were too complex and would be very hard to interpret. Therefore we have decided for option to adjust test data sample.</p>\n<p>Test data sample is adjusted, we can not guarantee that the changes we made in the fix will prevent all attempts of hacking, we acknowledge that Kagglers are very creative and smart, but we believe that the changes in test data make the hack attempts a magnitude harder to pull off. Obviously we don't want to disclose much information about new test data sample to limit the risk. What we wanted to say was already said in discussion. I would like to recommend you to switch your focus from hacking the metric :-)</p>\n<p>Keep in mind that competitions are here for Kagglers, but also for the Host - please respect that. We have also objectives that we want to achieve. We did our best to adjust data sample in a way that minimize risk of future hack and at the same time keep as much information for your models. We will keep the metric as it is.</p>\n<p>Tomas</p>",
      "rawMarkdown": "Dear Kagglers,\n\nAfter many days our competition is running again. I wish you all success and want to thank you again for your patience.\n\nThere were many comments lately with suggestions how the competition, metric, data, etc. should be designed or changed. I am thankful for many of your suggestions - really last weeks were very inspiring and at the same time very challenging for us. I feel I have to address some of your comments...\n\nWhen we were preparing this competition we knew from the beginning that we don't want to just repeat Home Credit competition from 2018. We have decided to add stability into competition objectives. Why? Because we think we optimize stability let's say expertly in our corporation and that there should be better way how to address this issue.\n\nWe understood that our metric was more complex than plain vanilla AUC and there is risk that Kaggle community will find approaches how to optimize it in a way we haven't thought about. Unfortunately this risk materialized and we decided to react. There were basically 2 options: a) change metric so it's not hackable or b) change data so it does not contain information needed for hack. We have investigated many potential metrics (thank you Kaggle community for proposals), however we were not able to find a metric that is good enough to describe \"good model\", moreover many metrics were too complex and would be very hard to interpret. Therefore we have decided for option to adjust test data sample.\n\nTest data sample is adjusted, we can not guarantee that the changes we made in the fix will prevent all attempts of hacking, we acknowledge that Kagglers are very creative and smart, but we believe that the changes in test data make the hack attempts a magnitude harder to pull off. Obviously we don't want to disclose much information about new test data sample to limit the risk. What we wanted to say was already said in discussion. I would like to recommend you to switch your focus from hacking the metric :-)\n\nKeep in mind that competitions are here for Kagglers, but also for the Host - please respect that. We have also objectives that we want to achieve. We did our best to adjust data sample in a way that minimize risk of future hack and at the same time keep as much information for your models. We will keep the metric as it is.\n\nTomas\n",
      "votes": 6,
      "replies": [
        {
          "id": 2693581,
          "postDate": "2024-03-12T14:31:45.153Z",
          "content": "<p>Can you tell us your current working model what score have if trained on the current training set? Or at least a segment of what a good score can represent if compared of what you have now? It's interesting to compare our results with it, not just between us on LB.</p>",
          "rawMarkdown": "Can you tell us your current working model what score have if trained on the current training set? Or at least a segment of what a good score can represent if compared of what you have now? It's interesting to compare our results with it, not just between us on LB."
        },
        {
          "id": 2695530,
          "postDate": "2024-03-13T18:04:01.640Z",
          "rawMarkdown": "",
          "isDeleted": true
        }
      ]
    },
    {
      "id": 2796995,
      "postDate": "2024-05-06T13:55:39.020Z",
      "content": "<p>Will solution using metric trick will be disqualified?</p>",
      "rawMarkdown": "Will solution using metric trick will be disqualified?",
      "votes": 3
    },
    {
      "id": 2691384,
      "postDate": "2024-03-11T07:51:06.767Z",
      "content": "<p>Thank you for your update and for your continued effort to improve this competition.</p>\n<p>When you review top solutions manually, and you find out that a solution involves metric hacking, which prizes is such a solution ineligible for?<br>\n1/ money prize<br>\n2/ Kaggle medals<br>\n3/ Kaggle ranking points</p>\n<p>FYI, 2/ and 3/ are the most valuable currency on Kaggle, definitely not 1/.</p>",
      "rawMarkdown": "Thank you for your update and for your continued effort to improve this competition.\n\nWhen you review top solutions manually, and you find out that a solution involves metric hacking, which prizes is such a solution ineligible for?\n1/ money prize\n2/ Kaggle medals\n3/ Kaggle ranking points\n\nFYI, 2/ and 3/ are the most valuable currency on Kaggle, definitely not 1/.",
      "votes": 5,
      "replies": [
        {
          "id": 2691392,
          "postDate": "2024-03-11T07:56:04.177Z",
          "content": "<p>Yeah, I'll start this competition based on this. I also mentioned the same thing but got no reply.</p>",
          "rawMarkdown": "Yeah, I'll start this competition based on this. I also mentioned the same thing but got no reply.",
          "votes": 3,
          "replies": [
            {
              "id": 2692142,
              "postDate": "2024-03-11T17:48:33.353Z",
              "content": "<p>even if every submission in top 100 will get checked there will be many people who don't recieve a medal because places &gt; 100 will contain large amount of people using metric hacking i think.</p>",
              "rawMarkdown": "even if every submission in top 100 will get checked there will be many people who don't recieve a medal because places > 100 will contain large amount of people using metric hacking i think.",
              "votes": 2
            },
            {
              "id": 2693312,
              "postDate": "2024-03-12T10:42:32.127Z",
              "content": "<p>Yes, and at the top positions people will use two subs for</p>\n<ol>\n<li>best score utilizing metric hacking for medals</li>\n<li>best score without hacking for the price money</li>\n</ol>",
              "rawMarkdown": "Yes, and at the top positions people will use two subs for\n1. best score utilizing metric hacking for medals\n2. best score without hacking for the price money\n",
              "votes": 3
            }
          ]
        },
        {
          "id": 2698128,
          "postDate": "2024-03-15T10:08:10.010Z",
          "content": "<p><a href=\"https://www.kaggle.com/jetakow\" target=\"_blank\">@jetakow</a> Would you like to comment?</p>",
          "rawMarkdown": "@jetakow Would you like to comment?",
          "votes": -1,
          "replies": [
            {
              "id": 2786545,
              "postDate": "2024-05-01T10:49:54.003Z",
              "content": "<p>just try to comment something</p>",
              "rawMarkdown": "just try to comment something\n",
              "votes": -1
            }
          ]
        }
      ]
    },
    {
      "id": 2728422,
      "postDate": "2024-04-02T08:32:50.643Z",
      "content": "<p>May I ask a question? If column 'date_decision' in test_base table is changed, I can no longer transform date (end with D) columns to days by minusing 'date_decision', right? </p>",
      "rawMarkdown": "May I ask a question? If column 'date_decision' in test_base table is changed, I can no longer transform date (end with D) columns to days by minusing 'date_decision', right? ",
      "votes": 3
    },
    {
      "id": 2686584,
      "postDate": "2024-03-07T23:18:47.077Z",
      "content": "<p>Just to clarify: WEEK_NUM in the training data is constant but it does vary in the test data, right? That would allow the usage of the same metric for evaluation of the submitted results. Unfortunately, a constant WEEK_NUM in the training data would not allow one to perform a fit on the Gini coefficient , get the slope (the coefficient \"a\") and, therefore, priorly compute the stability scores for that data. But I do understand that by making WEEK_NUM in the training data, virtually, there will be no way to make our models be \"manually\" tweaked to better predict for future weeks.</p>",
      "rawMarkdown": "Just to clarify: WEEK_NUM in the training data is constant but it does vary in the test data, right? That would allow the usage of the same metric for evaluation of the submitted results. Unfortunately, a constant WEEK_NUM in the training data would not allow one to perform a fit on the Gini coefficient , get the slope (the coefficient \"a\") and, therefore, priorly compute the stability scores for that data. But I do understand that by making WEEK_NUM in the training data, virtually, there will be no way to make our models be \"manually\" tweaked to better predict for future weeks.",
      "votes": 3,
      "replies": [
        {
          "id": 2687010,
          "postDate": "2024-03-08T07:51:53.490Z",
          "content": "<p>Hi,<br>\nthere is no change in train data - WEEK_NUM is still there and contains correct value, i.e. assignment of case_id to weekly time window - you can compute metric including stability part…<br>\nChange is in test data - week_num is constant so it does not carry any information (we did not delete it completely so you don't need to adjust your codes). But for evaluation in leaderboards we use correct WEEK_NUM, it's just hidden from Kagglers.</p>",
          "rawMarkdown": "Hi,\nthere is no change in train data - WEEK_NUM is still there and contains correct value, i.e. assignment of case_id to weekly time window - you can compute metric including stability part...\nChange is in test data - week_num is constant so it does not carry any information (we did not delete it completely so you don't need to adjust your codes). But for evaluation in leaderboards we use correct WEEK_NUM, it's just hidden from Kagglers.",
          "votes": 5,
          "replies": [
            {
              "id": 2687161,
              "postDate": "2024-03-08T10:23:39.840Z",
              "content": "<p>It is more clear to me now. Thanks!</p>",
              "rawMarkdown": "It is more clear to me now. Thanks!",
              "votes": 2
            },
            {
              "id": 2693559,
              "postDate": "2024-03-12T14:15:20.137Z",
              "content": "<blockquote>\n  <p>week_num is constant so it does not carry any information<br>\n  But for evaluation in leaderboards we use correct WEEK_NUM</p>\n</blockquote>\n<p>I don't understand how this makes sense. Is WEEK_NUM a constant in the test set or not?</p>",
              "rawMarkdown": ">week_num is constant so it does not carry any information\n>But for evaluation in leaderboards we use correct WEEK_NUM\n\nI don't understand how this makes sense. Is WEEK_NUM a constant in the test set or not?"
            },
            {
              "id": 2693570,
              "postDate": "2024-03-12T14:20:00.583Z",
              "content": "<p><a href=\"https://www.kaggle.com/alexanderholmberg\" target=\"_blank\">@alexanderholmberg</a> The fact that the WEEK_NUM is no longer in the test set doesn't imply we don't have this information while evaluating your submissions. It is just not present in test set you can load. </p>",
              "rawMarkdown": "@alexanderholmberg The fact that the WEEK_NUM is no longer in the test set doesn't imply we don't have this information while evaluating your submissions. It is just not present in test set you can load. ",
              "votes": 2
            }
          ]
        }
      ]
    },
    {
      "id": 2694308,
      "postDate": "2024-03-13T01:39:59.567Z",
      "content": "<p>From the <a href=\"https://www.kaggle.com/competitions/home-credit-credit-risk-model-stability/data\" target=\"_blank\">Data</a> tab:</p>\n<blockquote>\n  <p>Each group of tables can comprise one or more individual tables. If a group contains more than one table, they are divided based on <code>WEEK_NUM</code>.</p>\n</blockquote>\n<p>Isn't it possible to guess <code>WEEK_NUM</code> from how large tables such as <code>credit_bureau_a_2</code> are split?</p>",
      "rawMarkdown": "From the [Data](https://www.kaggle.com/competitions/home-credit-credit-risk-model-stability/data) tab:\n\n> Each group of tables can comprise one or more individual tables. If a group contains more than one table, they are divided based on `WEEK_NUM`.\n\nIsn't it possible to guess `WEEK_NUM` from how large tables such as `credit_bureau_a_2` are split?",
      "votes": 4,
      "replies": [
        {
          "id": 2696018,
          "postDate": "2024-03-14T02:30:46.373Z",
          "content": "<p>I experimented a bit by adding noise to the predictions for <code>case_id</code>'s that appear in the first half of <code>test_credit_bureau_a_2_*</code> files. The public score increased from 0.5 to around 0.51.</p>\n<p><a href=\"https://www.kaggle.com/jetakow\" target=\"_blank\">@jetakow</a> This hack is not possible if the test data is shuffled. I recommend shuffling unless it's already done.</p>",
          "rawMarkdown": "I experimented a bit by adding noise to the predictions for `case_id`'s that appear in the first half of `test_credit_bureau_a_2_*` files. The public score increased from 0.5 to around 0.51.\n\n@jetakow This hack is not possible if the test data is shuffled. I recommend shuffling unless it's already done.",
          "votes": 7,
          "replies": [
            {
              "id": 2698115,
              "postDate": "2024-03-15T09:56:13.363Z",
              "content": "<p><a href=\"https://www.kaggle.com/taichiuemura\" target=\"_blank\">@taichiuemura</a> Thank you for noticing, would you try to reproduce the idea once again now?</p>",
              "rawMarkdown": "@taichiuemura Thank you for noticing, would you try to reproduce the idea once again now?"
            },
            {
              "id": 2698175,
              "postDate": "2024-03-15T10:34:56.250Z",
              "content": "<p>I did. The score decreases now.</p>",
              "rawMarkdown": "I did. The score decreases now.",
              "votes": 2
            }
          ]
        }
      ]
    },
    {
      "id": 2688776,
      "postDate": "2024-03-09T13:02:46.817Z",
      "content": "<p>Will the 'date_decision' column still hold meaningful information about actual dates? If yes, then I believe hacking the metric is still possible, and it can even be done silently, making it unclear whether it’s a hack or just a meaningful model-choice decision. If the hypothesis about a structural break due to COVID is correct and the negative slope in the test data is related to this structural break, then tweaking models to perform worse on the COVID sample (for instance, by training the model exclusively on non-COVID data) becomes feasible. This would particularly benefit those who probe LB, which is concerning. Especially if 'date_decision' has been altered by adding a constant or something similar.</p>",
      "rawMarkdown": "Will the 'date_decision' column still hold meaningful information about actual dates? If yes, then I believe hacking the metric is still possible, and it can even be done silently, making it unclear whether it’s a hack or just a meaningful model-choice decision. If the hypothesis about a structural break due to COVID is correct and the negative slope in the test data is related to this structural break, then tweaking models to perform worse on the COVID sample (for instance, by training the model exclusively on non-COVID data) becomes feasible. This would particularly benefit those who probe LB, which is concerning. Especially if 'date_decision' has been altered by adding a constant or something similar.",
      "votes": 4,
      "replies": [
        {
          "id": 2690200,
          "postDate": "2024-03-10T11:24:08.117Z",
          "content": "<p><a href=\"https://www.kaggle.com/eivolkova\" target=\"_blank\">@eivolkova</a> Good point, but we of course thought about it. I can't tell you what we did exactly. </p>",
          "rawMarkdown": "@eivolkova Good point, but we of course thought about it. I can't tell you what we did exactly. ",
          "votes": -4,
          "replies": [
            {
              "id": 2690520,
              "postDate": "2024-03-10T16:01:24.210Z",
              "content": "<p>Good to hear that, thank you for the answer!</p>",
              "rawMarkdown": "Good to hear that, thank you for the answer!"
            }
          ]
        }
      ]
    },
    {
      "id": 2686628,
      "postDate": "2024-03-08T00:14:48.397Z",
      "content": "<blockquote>\n  <p>Column date_decision is no longer the same.</p>\n</blockquote>\n<p>Does that mean 'date_decision' will be meaningless, so we'd better not use it for training? Or has it been already dropped?</p>",
      "rawMarkdown": ">Column date_decision is no longer the same.\n\nDoes that mean 'date_decision' will be meaningless, so we'd better not use it for training? Or has it been already dropped?",
      "votes": 4,
      "replies": [
        {
          "id": 2687007,
          "postDate": "2024-03-08T07:46:26.293Z",
          "content": "<p>No, date_decision still exist, it can be used in models or in feature engineering. We did transformation that should make it impossible to hack metric in a way that was discussed after competition launch.</p>",
          "rawMarkdown": "No, date_decision still exist, it can be used in models or in feature engineering. We did transformation that should make it impossible to hack metric in a way that was discussed after competition launch.",
          "votes": 6,
          "replies": [
            {
              "id": 2745454,
              "postDate": "2024-04-10T16:34:08.670Z",
              "content": "<p>yes yes yes</p>",
              "rawMarkdown": "yes yes yes"
            }
          ]
        }
      ]
    },
    {
      "id": 2829594,
      "postDate": "2024-05-22T17:00:52.107Z",
      "content": "<p>Some thoughts on stablility metrics:</p>\n<p>1) when using prediciton models on production, we acknowledge a hypothesis. The future == The past .<br>\n2) but the real life is : in the past 4 years, we have expericed a greate lock down of Covid -19.  The world has been changed a lot. People's behavior also changed a lot. So The future ^=The past already.<br>\n3) Looking at the decison date on train set, we notice most of them are during the Covid-19 epidemic.<br>\n4) It can be predicted the vintage curve of those applicants will be worse than the population before Covid-19,even though they has same characteristic on demographic and bureau. </p>\n<p>5) the 3rd party data regarding macro economics might be useful.<br>\n6) Other method like risk table , which is typically used on anti-fraud, mighe be useful for risk guys to adjust model and strategy on time manner.  But it requires a target label at time point.</p>",
      "rawMarkdown": "Some thoughts on stablility metrics:\n\n1) when using prediciton models on production, we acknowledge a hypothesis. The future == The past .\n2) but the real life is : in the past 4 years, we have expericed a greate lock down of Covid -19.  The world has been changed a lot. People's behavior also changed a lot. So The future ^=The past already.\n3) Looking at the decison date on train set, we notice most of them are during the Covid-19 epidemic.\n4) It can be predicted the vintage curve of those applicants will be worse than the population before Covid-19,even though they has same characteristic on demographic and bureau. \n\n5) the 3rd party data regarding macro economics might be useful.\n6) Other method like risk table , which is typically used on anti-fraud, mighe be useful for risk guys to adjust model and strategy on time manner.  But it requires a target label at time point.\n\n\n\n\n",
      "votes": 2
    },
    {
      "id": 2734509,
      "postDate": "2024-04-04T06:49:10.863Z",
      "content": "<p>Hi, I am not sure if I understand it correctly, but isn't that the fact that MONTH and WEEK_NUM columns can be directly calculated from date_decision?<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F18899077%2Ff9916bef40ff638c6d0322a90c569e4b%2F2024-04-04%20144508.png?generation=1712213197463571&amp;alt=media\" alt=\"MONTH\"><br>\nif so, what is the meaning of setting WEEK_NUM and MONTH as constant in the given test data?<br>\nWhat's more, do the date_decision corresponds to real date when the decision is made? I am thinking about important relevant US interest rate and other data to facilitate my modeling?<br>\nCan anyone help me out? many thanks.</p>",
      "rawMarkdown": "Hi, I am not sure if I understand it correctly, but isn't that the fact that MONTH and WEEK_NUM columns can be directly calculated from date_decision?\n![MONTH](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F18899077%2Ff9916bef40ff638c6d0322a90c569e4b%2F2024-04-04%20144508.png?generation=1712213197463571&alt=media)\nif so, what is the meaning of setting WEEK_NUM and MONTH as constant in the given test data?\nWhat's more, do the date_decision corresponds to real date when the decision is made? I am thinking about important relevant US interest rate and other data to facilitate my modeling?\nCan anyone help me out? many thanks.",
      "votes": 1,
      "replies": [
        {
          "id": 2736550,
          "postDate": "2024-04-05T09:10:28.487Z",
          "rawMarkdown": "",
          "isDeleted": true
        }
      ]
    },
    {
      "id": 2715040,
      "postDate": "2024-03-25T08:31:02.643Z",
      "content": "<p>Thanks for your update, I appreciate the effort from the host and community to improve the quality of this competition.<br>\nAccording to the changes you mentioned:</p>\n<blockquote>\n  <p>As we indicated previously, we will be changing date columns in test_base table, namely date_decision, MONTH and WEEK_NUM. The MONTH and WEEK_NUM columns will now only have one constant value. We decided to keep them to minimize the amount of time you'll need to spend altering your scripts to get the submissions working again. <strong>Column date_decision is no longer the same</strong></p>\n</blockquote>\n<p>May I ask how did you change the \"date_decision\" column? For MONTH and WEEK_NUM, we know that they will be constant and we can't use them as information for the model. But for \"date_decision\" (and also other datetime columns), I have no idea whether I can use it anymore.<br>\nAs the most voted notebooks use \"date_decision\" as an anchor to process all other datetime columns (e.g. <code>df = df.with_columns(pl.col(col) - pl.col(\"date_decision\"))</code>), this pre-processing may be wrong if the behavior of \"date_decision\" is different in the test dataset.</p>",
      "rawMarkdown": "Thanks for your update, I appreciate the effort from the host and community to improve the quality of this competition.\nAccording to the changes you mentioned:\n>As we indicated previously, we will be changing date columns in test_base table, namely date_decision, MONTH and WEEK_NUM. The MONTH and WEEK_NUM columns will now only have one constant value. We decided to keep them to minimize the amount of time you'll need to spend altering your scripts to get the submissions working again. **Column date_decision is no longer the same**\n\nMay I ask how did you change the \"date_decision\" column? For MONTH and WEEK_NUM, we know that they will be constant and we can't use them as information for the model. But for \"date_decision\" (and also other datetime columns), I have no idea whether I can use it anymore.\nAs the most voted notebooks use \"date_decision\" as an anchor to process all other datetime columns (e.g. `df = df.with_columns(pl.col(col) - pl.col(\"date_decision\"))`), this pre-processing may be wrong if the behavior of \"date_decision\" is different in the test dataset.",
      "votes": 1,
      "replies": [
        {
          "id": 2715068,
          "postDate": "2024-03-25T08:55:05.357Z",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/minhtu123\" target=\"_blank\">@minhtu123</a> ,<br>\nWe don't want to disclose details about how date attributes were changed, but we did it in a way that keeps most of the information (at least from our point of view). Your example with date differences should not be affected by the change we have done.</p>",
          "rawMarkdown": "Hi @minhtu123 ,\nWe don't want to disclose details about how date attributes were changed, but we did it in a way that keeps most of the information (at least from our point of view). Your example with date differences should not be affected by the change we have done.",
          "replies": [
            {
              "id": 2715139,
              "postDate": "2024-03-25T09:53:30.067Z",
              "content": "<p>Thanks for your reply.</p>",
              "rawMarkdown": "Thanks for your reply."
            }
          ]
        }
      ]
    },
    {
      "id": 2691979,
      "postDate": "2024-03-11T15:52:10.223Z",
      "content": "<p>Thanks for the updates, I have a couple of questions.</p>\n<ol>\n<li><p>When 'Submissions' tab (next to 'Team' tab) will appear? I can't access my previous submissions anymore.</p></li>\n<li><p>May be I missed this information somewhere, but what will happen with our previous submissions and scores? Will they be left as they are? Or will they be corrected considering new metric updates? How this will affect ranking? </p></li>\n</ol>\n<p>Thanks for replies and sorry if I am asking such questions as they might be asked somewhere in the 'Discussion' tab.</p>",
      "rawMarkdown": "Thanks for the updates, I have a couple of questions.\n\n1. When 'Submissions' tab (next to 'Team' tab) will appear? I can't access my previous submissions anymore.\n\n2. May be I missed this information somewhere, but what will happen with our previous submissions and scores? Will they be left as they are? Or will they be corrected considering new metric updates? How this will affect ranking? \n\nThanks for replies and sorry if I am asking such questions as they might be asked somewhere in the 'Discussion' tab.",
      "votes": 1,
      "replies": [
        {
          "id": 2693303,
          "postDate": "2024-03-12T10:31:39.617Z",
          "content": "<blockquote>\n  <p>When 'Submissions' tab (next to 'Team' tab) will appear? I can't access my previous submissions anymore.</p>\n</blockquote>\n<p>It is there.</p>\n<blockquote>\n  <p>May be I missed this information somewhere, but what will happen with our previous submissions and scores? Will they be left as they are? Or will they be corrected considering new metric updates? How this will affect ranking?</p>\n</blockquote>\n<p>The LB was reset, your old submissions will work, but the scores changed completely. We encourage you to use the notebooks to generate a new submission table with new test data.</p>",
          "rawMarkdown": "> When 'Submissions' tab (next to 'Team' tab) will appear? I can't access my previous submissions anymore.\n\nIt is there.\n\n> May be I missed this information somewhere, but what will happen with our previous submissions and scores? Will they be left as they are? Or will they be corrected considering new metric updates? How this will affect ranking?\n\nThe LB was reset, your old submissions will work, but the scores changed completely. We encourage you to use the notebooks to generate a new submission table with new test data.",
          "votes": -1
        }
      ]
    },
    {
      "id": 2686959,
      "postDate": "2024-03-08T07:18:35.760Z",
      "content": "<blockquote>\n  <p>Column date_decision is no longer the same.</p>\n</blockquote>\n<p>It’s also not clear to me whether this column can be used to calculate new features or not? Let's say the number of loans for the last year?…</p>",
      "rawMarkdown": ">Column date_decision is no longer the same.\n\nIt’s also not clear to me whether this column can be used to calculate new features or not? Let's say the number of loans for the last year?...",
      "votes": 1,
      "replies": [
        {
          "id": 2687004,
          "postDate": "2024-03-08T07:44:04.890Z",
          "content": "<p>Hi dima,<br>\nyes, such features are possible to calculate</p>",
          "rawMarkdown": "Hi dima,\nyes, such features are possible to calculate",
          "replies": [
            {
              "id": 2687104,
              "postDate": "2024-03-08T09:29:34.607Z",
              "content": "<p>can you do, for example, \"date_decision\" - \"birthdate_87D\"? is that still the same number as it was before the transformation?</p>",
              "rawMarkdown": "can you do, for example, \"date_decision\" - \"birthdate_87D\"? is that still the same number as it was before the transformation?\n",
              "votes": 3
            },
            {
              "id": 2687162,
              "postDate": "2024-03-08T10:25:20.603Z",
              "content": "<p>yes, difference remains same</p>",
              "rawMarkdown": "yes, difference remains same",
              "votes": 2
            },
            {
              "id": 2687203,
              "postDate": "2024-03-08T11:03:56.890Z",
              "content": "<p><a href=\"https://www.kaggle.com/tomasjeline2\" target=\"_blank\">@tomasjeline2</a>  This is very confusing. </p>\n<p>How come the difference is still the same. Does this mean that you have also changed other dates in the test data like 'birthdate_87D' ? I was under the impressions that the only changed columns are date_decision, MONTH and WEEK_NUM as per the announcement.</p>\n<p>Can you please explain how did you change date_decision and not birthdate_87D and the difference is the same ?</p>",
              "rawMarkdown": "@tomasjeline2  This is very confusing. \n\nHow come the difference is still the same. Does this mean that you have also changed other dates in the test data like 'birthdate_87D' ? I was under the impressions that the only changed columns are date_decision, MONTH and WEEK_NUM as per the announcement.\n\nCan you please explain how did you change date_decision and not birthdate_87D and the difference is the same ?",
              "votes": 3
            },
            {
              "id": 2687400,
              "postDate": "2024-03-08T14:34:26.060Z",
              "content": "<p>Hi Armada,<br>\nWe don't want to explain details about how data were transformed.<br>\nFrom the point of view of single case_id, we design transformation in a way that date attributes should contain same/similar information, features like client's age (date_app-birthdate) are not affected.</p>",
              "rawMarkdown": "Hi Armada,\nWe don't want to explain details about how data were transformed.\nFrom the point of view of single case_id, we design transformation in a way that date attributes should contain same/similar information, features like client's age (date_app-birthdate) are not affected.",
              "votes": -2
            },
            {
              "id": 2687930,
              "postDate": "2024-03-08T21:40:17.660Z",
              "content": "<p><a href=\"https://www.kaggle.com/tomasjeline2\" target=\"_blank\">@tomasjeline2</a>  Yes I fully understand that you cannot disclose how the transformation is done. My question was whether only these three columns are transformed : date_decision, MONTH and WEEK_NUM  or were there other date columns transformed and if it is the latter can you please list which other columns were transformed in the new test data.</p>",
              "rawMarkdown": "@tomasjeline2  Yes I fully understand that you cannot disclose how the transformation is done. My question was whether only these three columns are transformed : date_decision, MONTH and WEEK_NUM  or were there other date columns transformed and if it is the latter can you please list which other columns were transformed in the new test data.",
              "votes": 1
            },
            {
              "id": 2688703,
              "postDate": "2024-03-09T12:05:49.400Z",
              "content": "<p>No we don't want to list what was exactly transformed and how and we will not do it. Please understand it that it is crucial for the sake of the fairness in this competition and at the same time it should not be a major concern of yours if you are working on feature engineering. </p>",
              "rawMarkdown": "No we don't want to list what was exactly transformed and how and we will not do it. Please understand it that it is crucial for the sake of the fairness in this competition and at the same time it should not be a major concern of yours if you are working on feature engineering. ",
              "votes": -1
            },
            {
              "id": 2690521,
              "postDate": "2024-03-10T16:03:58.043Z",
              "content": "<blockquote>\n  <p>From the point of view of single case_id</p>\n</blockquote>\n<p>What about aggregations based on time periods? For example, is it possible to calculate a rolling monthly average of some feature for the test data?</p>",
              "rawMarkdown": ">From the point of view of single case_id\n\nWhat about aggregations based on time periods? For example, is it possible to calculate a rolling monthly average of some feature for the test data?",
              "votes": 1
            },
            {
              "id": 2693944,
              "postDate": "2024-03-12T18:57:03.927Z",
              "content": "<p><a href=\"https://www.kaggle.com/jetakow\" target=\"_blank\">@jetakow</a> <a href=\"https://www.kaggle.com/tomasjeline2\" target=\"_blank\">@tomasjeline2</a> hey, thanks for continuing to be so responsive and all. tomas mentioned that \"date_decision\" - \"birthdate_87D\" is the same value as it was before, but i am not sure now if he was speaking only about this one example.</p>\n<p>is this true for EVERY column with D in it, so that i can still do \"date_decision\"-\"arbitrary D column\" and still get the same value as before the dataset change? i hope you can tell us about this, because these features are somewhat important for the model and checking each one with a new submission would be very time consuming…</p>",
              "rawMarkdown": "@jetakow @tomasjeline2 hey, thanks for continuing to be so responsive and all. tomas mentioned that \"date_decision\" - \"birthdate_87D\" is the same value as it was before, but i am not sure now if he was speaking only about this one example.\n\nis this true for EVERY column with D in it, so that i can still do \"date_decision\"-\"arbitrary D column\" and still get the same value as before the dataset change? i hope you can tell us about this, because these features are somewhat important for the model and checking each one with a new submission would be very time consuming...",
              "votes": 2
            },
            {
              "id": 2694148,
              "postDate": "2024-03-12T22:25:18.027Z",
              "content": "<p><a href=\"https://www.kaggle.com/at7459\" target=\"_blank\">@at7459</a> once again, we don't want to list what was exactly transformed and how and we will not do it. Please understand it that it is crucial for the sake of the fairness in this competition. You can try the date diff features and see what is working and what is not. That is afraid is the whole information we can give you. The FE is also part of the competition. </p>",
              "rawMarkdown": "@at7459 once again, we don't want to list what was exactly transformed and how and we will not do it. Please understand it that it is crucial for the sake of the fairness in this competition. You can try the date diff features and see what is working and what is not. That is afraid is the whole information we can give you. The FE is also part of the competition. ",
              "votes": -7
            },
            {
              "id": 2694149,
              "postDate": "2024-03-12T22:27:13.890Z",
              "content": "<p><a href=\"https://www.kaggle.com/eivolkova\" target=\"_blank\">@eivolkova</a> You can try those aggregations and see if any fo that is working or not. FE is also part of the competition and it should be left to be explored by kagglers. I can not guarantee you however that rolling monthly average will be working on test data. That is I am afraid all I can say.</p>",
              "rawMarkdown": "@eivolkova You can try those aggregations and see if any fo that is working or not. FE is also part of the competition and it should be left to be explored by kagglers. I can not guarantee you however that rolling monthly average will be working on test data. That is I am afraid all I can say.",
              "votes": -3
            },
            {
              "id": 2694630,
              "postDate": "2024-03-13T07:13:44.557Z",
              "content": "<p>that is a bad decision, checking what is working and what is not realistically feasible and turns this competition into largely guessing</p>\n<p>for example, i can see by submissions that all these \"dateyear\" \"datemonth\" etc. were not properly transformed, i used to be able to subtract these from date decision and get a score boost, now it lowers my score by 0.005. you can't seriously expect me to go through every date feature and check by submission whether they were properly transformed or not…</p>",
              "rawMarkdown": "that is a bad decision, checking what is working and what is not realistically feasible and turns this competition into largely guessing\n\nfor example, i can see by submissions that all these \"dateyear\" \"datemonth\" etc. were not properly transformed, i used to be able to subtract these from date decision and get a score boost, now it lowers my score by 0.005. you can't seriously expect me to go through every date feature and check by submission whether they were properly transformed or not...",
              "votes": 13
            },
            {
              "id": 2694669,
              "postDate": "2024-03-13T07:38:00.550Z",
              "content": "<p>you said that the meaning of the features is preserved, but this clearly is not the case…</p>\n<p>i mean if you're going to manually check everything anyway i don't know why you would even want to transform anything in the first place</p>",
              "rawMarkdown": "you said that the meaning of the features is preserved, but this clearly is not the case...\n\ni mean if you're going to manually check everything anyway i don't know why you would even want to transform anything in the first place",
              "votes": 6
            }
          ]
        }
      ]
    },
    {
      "id": 2686533,
      "postDate": "2024-03-07T21:52:08.940Z",
      "content": "<p>Thanks for the update.</p>\n<blockquote>\n  <p>Ultimately, we don't want participants to focus on metric hacking; however, we also want a metric that closely approximates how models are assessed in production. When looking at the performance of models in production the predictable stable decrease in performance is undesirable even if it is not yet observed, but expected to come. We didn’t find a metric that would satisfy both requirements and that is why we decided to keep the metric as it is now.</p>\n</blockquote>\n<p>I’m not sure I understand this paragraph. How is it possible to keep the same metric and have no issues with metric hacking?</p>",
      "rawMarkdown": "Thanks for the update.\n\n>Ultimately, we don't want participants to focus on metric hacking; however, we also want a metric that closely approximates how models are assessed in production. When looking at the performance of models in production the predictable stable decrease in performance is undesirable even if it is not yet observed, but expected to come. We didn’t find a metric that would satisfy both requirements and that is why we decided to keep the metric as it is now.\n\nI’m not sure I understand this paragraph. How is it possible to keep the same metric and have no issues with metric hacking?",
      "votes": 1,
      "replies": [
        {
          "id": 2686548,
          "postDate": "2024-03-07T21:59:39.153Z",
          "content": "<p>We believe the changes in test data will minimize metric hacking in terms of the gain.</p>",
          "rawMarkdown": "We believe the changes in test data will minimize metric hacking in terms of the gain.",
          "votes": -2,
          "replies": [
            {
              "id": 2686552,
              "postDate": "2024-03-07T22:06:54.870Z",
              "content": "<p>I see, so “data transformation” means “data shuffling” in this context?</p>",
              "rawMarkdown": "I see, so “data transformation” means “data shuffling” in this context?",
              "votes": 2
            },
            {
              "id": 2686562,
              "postDate": "2024-03-07T22:29:21.573Z",
              "content": "<p>We won't give more detail on what exactly was changed in test set.</p>",
              "rawMarkdown": "We won't give more detail on what exactly was changed in test set.",
              "votes": -2
            }
          ]
        }
      ]
    },
    {
      "id": 2726400,
      "postDate": "2024-04-01T06:58:37.980Z",
      "content": "<p>Is weekday from date decision still makes sense in the test set?</p>",
      "rawMarkdown": "Is weekday from date decision still makes sense in the test set?\n\n\n",
      "votes": 2
    },
    {
      "id": 2693347,
      "postDate": "2024-03-12T11:23:19.217Z",
      "content": "<p>Is weekday from date decision still makes sense in the test set?</p>",
      "rawMarkdown": "Is weekday from date decision still makes sense in the test set?",
      "votes": 2
    },
    {
      "id": 2692770,
      "postDate": "2024-03-12T04:13:12.090Z",
      "content": "<p>Thanks for the update. Had trouble submitting yesterday but seems to work fine now. Excited! May the best win.</p>",
      "rawMarkdown": "Thanks for the update. Had trouble submitting yesterday but seems to work fine now. Excited! May the best win.",
      "votes": 2
    },
    {
      "id": 2690789,
      "postDate": "2024-03-10T19:39:21.503Z",
      "content": "<p>Can someone help ELI5 what's the problem with the WEEK_NUM?</p>",
      "rawMarkdown": "Can someone help ELI5 what's the problem with the WEEK_NUM?",
      "votes": 2
    },
    {
      "id": 2688747,
      "postDate": "2024-03-09T12:36:36.733Z",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/jetakow\" target=\"_blank\">@jetakow</a>!<br>\nWhat do you think about add hash(date_decision.weekday) as new feature in train &amp; test dataset?</p>",
      "rawMarkdown": "Hi @jetakow!\nWhat do you think about add hash(date_decision.weekday) as new feature in train & test dataset?",
      "votes": 2,
      "replies": [
        {
          "id": 2688778,
          "postDate": "2024-03-09T13:02:57.770Z",
          "content": "<p>Sorry to jump in.<br>\nI think it's a great idea and I was thinking about the same thing.<br>\nAggregation like <code>xxx group by date_decision</code> is one of the good way to create new features, which can also be measured for actual industrial ML models.<br>\nHowever, obviously we can't use this way if all <code>date_decision</code> in test data become constant values.<br>\nThe problem is that we can hack and game metrics by directly using date/time information, so it would not be problem if we were to add hashed date features as you said.</p>",
          "rawMarkdown": "Sorry to jump in.\nI think it's a great idea and I was thinking about the same thing.\nAggregation like `xxx group by date_decision` is one of the good way to create new features, which can also be measured for actual industrial ML models.\nHowever, obviously we can't use this way if all `date_decision` in test data become constant values.\nThe problem is that we can hack and game metrics by directly using date/time information, so it would not be problem if we were to add hashed date features as you said.",
          "replies": [
            {
              "id": 2689870,
              "postDate": "2024-03-10T07:08:16.190Z",
              "content": "<p>I think, they preserve correct weekday, while  transform the data</p>",
              "rawMarkdown": "I think, they preserve correct weekday, while ~~shift dates~~ transform the data"
            }
          ]
        }
      ]
    },
    {
      "id": 2686629,
      "postDate": "2024-03-08T00:15:59.050Z",
      "content": "<p>Hi Daniel,</p>\n<blockquote>\n  <p>The remaining data were also transformed. Don't worry; the transformation will preserve the meaning of features. Data types will remain the same, and there shouldn't be significant changes in the distributions.</p>\n</blockquote>\n<p>Will the training dataset also be update with this transformations?</p>\n<p>Thanks.</p>",
      "rawMarkdown": "Hi Daniel,\n\n>The remaining data were also transformed. Don't worry; the transformation will preserve the meaning of features. Data types will remain the same, and there shouldn't be significant changes in the distributions.\n\nWill the training dataset also be update with this transformations?\n\nThanks.",
      "votes": 1,
      "replies": [
        {
          "id": 2687949,
          "postDate": "2024-03-08T22:03:26.147Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 2688697,
          "postDate": "2024-03-09T12:01:43.850Z",
          "content": "<p>Training data were not a subject of this announcement implying they will remain untouched. </p>",
          "rawMarkdown": "Training data were not a subject of this announcement implying they will remain untouched. ",
          "votes": -1
        }
      ]
    },
    {
      "id": 2748444,
      "postDate": "2024-04-12T12:53:56.267Z",
      "content": "<p>请问您当前的工作模型如果在当前训练集上训练的话得分是多少？好分数可以代表什么？</p>",
      "rawMarkdown": "请问您当前的工作模型如果在当前训练集上训练的话得分是多少？好分数可以代表什么？",
      "votes": -2
    },
    {
      "id": 2748605,
      "postDate": "2024-04-12T14:02:34.607Z",
      "content": "<p>Hello dear host, I know that there is a data file in this data set called train_applprev_1_0.csv, which represents the user's historical loans and the status of these loans. The feature: <code>status_219L</code>, which has multiple values, represented by letters A, T, D, etc. However, the state represented by each letter does not seem to be described, which makes me confused. Because I want to consider the stability of the model from the perspective of historical loans, because I need to know the specific loan status represented by these letters, especially the order of the letters. This is particularly important to me. Can you provide a specific explanation of these feature value?</p>",
      "rawMarkdown": "Hello dear host, I know that there is a data file in this data set called train_applprev_1_0.csv, which represents the user's historical loans and the status of these loans. The feature: `status_219L`, which has multiple values, represented by letters A, T, D, etc. However, the state represented by each letter does not seem to be described, which makes me confused. Because I want to consider the stability of the model from the perspective of historical loans, because I need to know the specific loan status represented by these letters, especially the order of the letters. This is particularly important to me. Can you provide a specific explanation of these feature value?",
      "votes": -1
    },
    {
      "id": 2845430,
      "postDate": "2024-05-30T14:59:18.077Z",
      "content": "<p>Good Solution</p>",
      "rawMarkdown": "Good Solution"
    },
    {
      "id": 2826426,
      "postDate": "2024-05-20T23:01:17.030Z",
      "content": "<p>I am getting this \"TypeError: the truth value of a Series is ambiguous\" when trying to use the Polars library for data cleaning, anyone has idea on how to deal with this.</p>",
      "rawMarkdown": "I am getting this \"TypeError: the truth value of a Series is ambiguous\" when trying to use the Polars library for data cleaning, anyone has idea on how to deal with this."
    },
    {
      "id": 2825947,
      "postDate": "2024-05-20T16:33:34.773Z",
      "content": "<p>May I know if we can still upload our models even after the deadline just to check our score? I am aware that we will not be allowed to change our final result of the competition, just asking if we could upload.</p>",
      "rawMarkdown": "May I know if we can still upload our models even after the deadline just to check our score? I am aware that we will not be allowed to change our final result of the competition, just asking if we could upload."
    },
    {
      "id": 2814641,
      "postDate": "2024-05-15T13:17:17.703Z",
      "content": "<p>good n awesome!</p>",
      "rawMarkdown": "good n awesome!"
    },
    {
      "id": 2810703,
      "postDate": "2024-05-13T12:09:18.603Z",
      "content": "<p>Good Solution</p>",
      "rawMarkdown": "Good Solution"
    },
    {
      "id": 2805625,
      "postDate": "2024-05-10T16:44:29.117Z",
      "content": "<p>Thank you for your work, I have learned  a lot from it.</p>",
      "rawMarkdown": "Thank you for your work, I have learned  a lot from it.",
      "replies": [
        {
          "id": 2806283,
          "postDate": "2024-05-11T02:27:00.567Z",
          "content": "<p>Me too. This really helps.</p>",
          "rawMarkdown": "Me too. This really helps."
        }
      ]
    },
    {
      "id": 2795787,
      "postDate": "2024-05-06T01:13:20.133Z",
      "content": "<p>Hello, since this competition is on hold for couple weeks,  I would like to know what is the maximum number of submissions allowed for a team.</p>",
      "rawMarkdown": "Hello, since this competition is on hold for couple weeks,  I would like to know what is the maximum number of submissions allowed for a team.",
      "replies": [
        {
          "id": 2796911,
          "postDate": "2024-05-06T13:13:17.260Z",
          "content": "<p>Usually, it is the length of the competition day multiplied by 5.</p>",
          "rawMarkdown": "Usually, it is the length of the competition day multiplied by 5.",
          "replies": [
            {
              "id": 2797860,
              "postDate": "2024-05-07T01:41:23.630Z",
              "content": "<p>Will the weeks of suspension be counted?</p>",
              "rawMarkdown": "Will the weeks of suspension be counted?"
            }
          ]
        }
      ]
    },
    {
      "id": 2786620,
      "postDate": "2024-05-01T11:34:32.337Z",
      "content": "<p>Great competition</p>",
      "rawMarkdown": "Great competition"
    },
    {
      "id": 2778367,
      "postDate": "2024-04-27T06:08:11.340Z",
      "content": "<p>How do I download the dataset? </p>",
      "rawMarkdown": "How do I download the dataset? "
    },
    {
      "id": 2767508,
      "postDate": "2024-04-22T11:06:35.030Z",
      "content": "<p>great work</p>",
      "rawMarkdown": "great work",
      "replies": [
        {
          "id": 2797090,
          "postDate": "2024-05-06T14:27:28.833Z",
          "content": "<p>gooood work</p>",
          "rawMarkdown": "gooood work"
        }
      ]
    },
    {
      "id": 2763720,
      "postDate": "2024-04-20T17:47:25.823Z",
      "content": "<p>Thanks for Your response are really helpful here.</p>",
      "rawMarkdown": "Thanks for Your response are really helpful here."
    },
    {
      "id": 2754653,
      "postDate": "2024-04-16T06:19:45.507Z",
      "content": "<p>Hello,<br>\nThank you very much for your comprehensive response regarding the metrics. <br>\nI have a question concerning WEEK_NUM in the test data. I understand that WEEK_NUM has been modified, suggesting that it has been encrypted in a way that disrupts its original order. If my understanding is correct, this encryption would allow those with decryption capabilities, such as yourselves, to perform calculations as intended. However, for us who cannot decrypt it, we are no longer able to do metric hacking. Could you please confirm if my understanding is correct?</p>",
      "rawMarkdown": "\nHello,\n\nThank you very much for your comprehensive response regarding the metrics. \n\nI have a question concerning WEEK_NUM in the test data. I understand that WEEK_NUM has been modified, suggesting that it has been encrypted in a way that disrupts its original order. If my understanding is correct, this encryption would allow those with decryption capabilities, such as yourselves, to perform calculations as intended. However, for us who cannot decrypt it, we are no longer able to do metric hacking. Could you please confirm if my understanding is correct?\n",
      "replies": [
        {
          "id": 2754783,
          "postDate": "2024-04-16T07:45:11.623Z",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/ariafschneider\" target=\"_blank\">@ariafschneider</a> ,<br>\nYes, your understanding is correct. <br>\nWEEK_NUM in test sample is not encrypted, it just have constant for all records, therefore no information value.<br>\nFor evaluation, we use real (not transformed) week_num that is hidden from test sample.</p>",
          "rawMarkdown": "Hi @ariafschneider ,\nYes, your understanding is correct. \nWEEK_NUM in test sample is not encrypted, it just have constant for all records, therefore no information value.\nFor evaluation, we use real (not transformed) week_num that is hidden from test sample.",
          "replies": [
            {
              "id": 2756787,
              "postDate": "2024-04-17T07:16:30.697Z",
              "content": "<p>I see! thank you!Good luck with your management!</p>",
              "rawMarkdown": "I see! thank you!Good luck with your management!"
            }
          ]
        }
      ]
    },
    {
      "id": 2743368,
      "postDate": "2024-04-09T12:16:33.597Z",
      "content": "<p>Realizing I have so much to learn.</p>",
      "rawMarkdown": "Realizing I have so much to learn."
    },
    {
      "id": 2719179,
      "postDate": "2024-03-27T14:55:07.410Z",
      "content": "<p>Hi,</p>\n<p>I use an open source machine learning tool called khiops. There is no internet acces, so Is it possible to install it in your notebook ?</p>\n<p>conda install -y -c conda-forge -c khiops khiops</p>\n<p>khiops use a  BSD 3-Clause-clear License</p>\n<p>Nicolas</p>",
      "rawMarkdown": "Hi,\n\nI use an open source machine learning tool called khiops. There is no internet acces, so Is it possible to install it in your notebook ?\n\nconda install -y -c conda-forge -c khiops khiops\n\nkhiops use a  BSD 3-Clause-clear License\n\nNicolas",
      "replies": [
        {
          "id": 2719183,
          "postDate": "2024-03-27T14:56:42.243Z",
          "content": "<p>Here is your potential solution <a href=\"https://www.kaggle.com/code/samuelepino/pip-installing-packages-with-no-internet\" target=\"_blank\">https://www.kaggle.com/code/samuelepino/pip-installing-packages-with-no-internet</a></p>",
          "rawMarkdown": "Here is your potential solution https://www.kaggle.com/code/samuelepino/pip-installing-packages-with-no-internet",
          "votes": 2
        }
      ]
    },
    {
      "id": 2712073,
      "postDate": "2024-03-23T09:33:34.753Z",
      "content": "<p>I think this is an interesting idea</p>",
      "rawMarkdown": "I think this is an interesting idea"
    },
    {
      "id": 2697631,
      "postDate": "2024-03-15T03:24:26.857Z",
      "content": "<p>Hi, I have an unrelated question is that if the data is used for application scoring, I have seen that the record date in the train_tax_registry_table_a is always later than the decision date? This means that the info won't be available at the time of the scoring. How would the table be useful? or it's just that the data has been transformed and the true date is actual before the decision date so I can continue to use it in my model without any issues?</p>",
      "rawMarkdown": "Hi, I have an unrelated question is that if the data is used for application scoring, I have seen that the record date in the train_tax_registry_table_a is always later than the decision date? This means that the info won't be available at the time of the scoring. How would the table be useful? or it's just that the data has been transformed and the true date is actual before the decision date so I can continue to use it in my model without any issues?",
      "replies": [
        {
          "id": 2698201,
          "postDate": "2024-03-15T10:49:04.903Z",
          "content": "<p>Hi FlyingCamel,<br>\nall data are as of decision date, so there is no future information. We know that some dates appear as future, but you can use them in model</p>",
          "rawMarkdown": "Hi FlyingCamel,\nall data are as of decision date, so there is no future information. We know that some dates appear as future, but you can use them in model",
          "votes": 1
        }
      ]
    },
    {
      "id": 2693752,
      "postDate": "2024-03-12T16:21:54.637Z",
      "content": "<p>Does that mean we should be downloading the data again? </p>",
      "rawMarkdown": "Does that mean we should be downloading the data again? ",
      "replies": [
        {
          "id": 2693835,
          "postDate": "2024-03-12T17:48:08.930Z",
          "content": "<p>Train data remained the same.</p>",
          "rawMarkdown": "Train data remained the same.",
          "votes": 2
        }
      ]
    },
    {
      "id": 2691252,
      "postDate": "2024-03-11T06:35:22.027Z",
      "content": "<p>Firstly, my regards to the competition hosts. It can be a very tough time for all of you.</p>\n<p>Taking a step back, I'd say perhaps the best way to emphasize stability is to work on the data horizon, e.g. predict future 5 years with only 1 year-worth of data, possibly giving higher weigh to more recent data. But it is too late now.</p>",
      "rawMarkdown": "Firstly, my regards to the competition hosts. It can be a very tough time for all of you.\n\nTaking a step back, I'd say perhaps the best way to emphasize stability is to work on the data horizon, e.g. predict future 5 years with only 1 year-worth of data, possibly giving higher weigh to more recent data. But it is too late now."
    },
    {
      "id": 2687656,
      "postDate": "2024-03-08T17:13:52.587Z",
      "content": "<p>Hi Daniel, </p>\n<p>While you are updating the data in the test set (the one with only 10 rows), can you please fix the parquet files in the test set?  They actually have a different structure than the train set.</p>\n<ul>\n<li>For example, file parquet_files/test/test_credit_bureau_a_1_3.parquet is not the correct format.  You get a failure when loading all the test_credit_bureau_a_1 files together.  I doubt that this will happen when we run on the hidden test set, but it would be great to be able to test our process on the small test set provided.  </li>\n</ul>\n<p>Below, I illustrate an example where file credit_bureau_a_1_1 and credit_bureau_a_1_3 can't union together.  </p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F522718%2F7cc98e88b2981c3df17d065ed4724a87%2FScreenshot%202024-03-08%20120230.png?generation=1709917399648920&amp;alt=media\"></p>\n<p>The above shows a failure when trying to union those files together.  <br>\nI would suggest to decouple the CSV and PARQUET process completely.  It feels like perhaps you created the PARQUETS after creating the CSV files?  <br>\nParquets are terrific as I am sure you know … less space and retains column types…  and I like that this competition has them.  </p>\n<p>If you load like this -- </p>\n<pre><code>LOAD DATA OVERWRITE `home_credit_comp.test_credit_bureau_a_1_1` FROM FILES (\n                 = \n                ,uris = []);\n</code></pre>\n<pre><code>     DATA OVERWRITE `home_credit_comp.test_credit_bureau_a_1_3`  FILES (\n             = ,\n            uris = []);\n\n     *  `home_credit_comp.test_credit_bureau_a_1_1`\n      \n     *  `home_credit_comp.test_credit_bureau_a_1_3`\n    ; \n</code></pre>\n<p>This also happened when importing the test_static_0 parquets…</p>\n<pre><code>        LOAD DATA OVERWRITE `concrete-acre-home_credit_comp.test_static_0`\n            FROM FILES (\n                 = ,\n                uris = []\n            )\n        ;\n</code></pre>",
      "rawMarkdown": "Hi Daniel, \n\nWhile you are updating the data in the test set (the one with only 10 rows), can you please fix the parquet files in the test set?  They actually have a different structure than the train set.\n\n- For example, file parquet_files/test/test_credit_bureau_a_1_3.parquet is not the correct format.  You get a failure when loading all the test_credit_bureau_a_1 files together.  I doubt that this will happen when we run on the hidden test set, but it would be great to be able to test our process on the small test set provided.  \n\nBelow, I illustrate an example where file credit_bureau_a_1_1 and credit_bureau_a_1_3 can't union together.  \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F522718%2F7cc98e88b2981c3df17d065ed4724a87%2FScreenshot%202024-03-08%20120230.png?generation=1709917399648920&alt=media)\n\nThe above shows a failure when trying to union those files together.  \nI would suggest to decouple the CSV and PARQUET process completely.  It feels like perhaps you created the PARQUETS after creating the CSV files?  \nParquets are terrific as I am sure you know ... less space and retains column types...  and I like that this competition has them.  \n\nIf you load like this -- \n```python\n\nLOAD DATA OVERWRITE `home_credit_comp.test_credit_bureau_a_1_1` FROM FILES (\n                format = 'PARQUET'\n                ,uris = ['gs://kds-2b8b4dc180f8d5e733a579005c44bbc1835fccdef4262654e2e5b687/parquet_files/test/test_credit_bureau_a_1_1.parquet']);\n        \n```\n\n        LOAD DATA OVERWRITE `home_credit_comp.test_credit_bureau_a_1_3` FROM FILES (\n                format = 'PARQUET',\n                uris = ['gs://kds-2b8b4dc180f8d5e733a579005c44bbc1835fccdef4262654e2e5b687/parquet_files/test/test_credit_bureau_a_1_3.parquet']);\n                \n        select * from `home_credit_comp.test_credit_bureau_a_1_1`\n        union all \n        select * from `home_credit_comp.test_credit_bureau_a_1_3`\n        ; \n\nThis also happened when importing the test_static_0 parquets...\n\n```python\n        LOAD DATA OVERWRITE `concrete-acre-416420.home_credit_comp.test_static_0`\n            FROM FILES (\n                format = 'PARQUET',\n                uris = ['gs://kds-2b8b4dc180f8d5e733a579005c44bbc1835fccdef4262654e2e5b687/parquet_files/test/test_static_0*.parquet']\n            )\n        ;\n```\n        ",
      "replies": [
        {
          "id": 2688701,
          "postDate": "2024-03-09T12:04:04.723Z",
          "content": "<p>I believe you can always load tables one by one and then manually change their schema so that you can concatenate them. The schema can be inferred from the values and from the training set. </p>",
          "rawMarkdown": "I believe you can always load tables one by one and then manually change their schema so that you can concatenate them. The schema can be inferred from the values and from the training set. ",
          "votes": -5
        }
      ]
    },
    {
      "id": 2686555,
      "postDate": "2024-03-07T22:11:38.967Z",
      "content": "<p>You mentioned in an earlier post that you will be reviewing the top 100 solutions manually in order to make sure that they are not scoring high by hacking the metric. Is this still the case or will you just accept the score as is even if people found other ways to hack it ?</p>",
      "rawMarkdown": "You mentioned in an earlier post that you will be reviewing the top 100 solutions manually in order to make sure that they are not scoring high by hacking the metric. Is this still the case or will you just accept the score as is even if people found other ways to hack it ?",
      "replies": [
        {
          "id": 2686564,
          "postDate": "2024-03-07T22:31:59.150Z",
          "content": "<p>It is still the case. If it wasn't clear already, direct metric hacking by altering the model scores in a way that artificially worsens selected case_ids is undesirable in this competition. </p>",
          "rawMarkdown": "It is still the case. If it wasn't clear already, direct metric hacking by altering the model scores in a way that artificially worsens selected case_ids is undesirable in this competition. ",
          "replies": [
            {
              "id": 2686910,
              "postDate": "2024-03-08T06:17:04.017Z",
              "content": "<p>But people can still get medals by doing that?</p>",
              "rawMarkdown": "But people can still get medals by doing that?",
              "votes": 5
            }
          ]
        }
      ]
    },
    {
      "id": 2771798,
      "postDate": "2024-04-24T12:10:20.580Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 2748598,
      "postDate": "2024-04-12T14:00:48.200Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 2705487,
      "postDate": "2024-03-19T11:21:11.270Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 2692378,
      "postDate": "2024-03-11T20:28:37.677Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 2686795,
      "postDate": "2024-03-08T05:05:35.267Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 2803335,
      "postDate": "2024-05-09T12:36:06.693Z",
      "content": "<p>Thank you!</p>",
      "rawMarkdown": "Thank you!"
    },
    {
      "id": 2802854,
      "postDate": "2024-05-09T07:40:22.507Z",
      "content": "<p>thanks for hosters</p>",
      "rawMarkdown": "thanks for hosters"
    },
    {
      "id": 2800204,
      "postDate": "2024-05-08T05:55:17.760Z",
      "content": "<p>Thanks for your efforts</p>",
      "rawMarkdown": "Thanks for your efforts"
    },
    {
      "id": 2780164,
      "postDate": "2024-04-28T04:53:11.983Z",
      "content": "<p>thanks your efforts</p>",
      "rawMarkdown": "thanks your efforts"
    },
    {
      "id": 2766318,
      "postDate": "2024-04-21T16:17:48.063Z",
      "content": "<p>Thanks for the update.</p>",
      "rawMarkdown": "Thanks for the update."
    },
    {
      "id": 2762769,
      "postDate": "2024-04-20T04:49:31.233Z",
      "content": "<p>Thanks for sharing</p>",
      "rawMarkdown": "Thanks for sharing"
    },
    {
      "id": 2724298,
      "postDate": "2024-03-30T19:28:05.297Z",
      "content": "<p>Sure, Thank you for the update!</p>",
      "rawMarkdown": "Sure, Thank you for the update!"
    }
  ],
  "comments": [
    {
      "id": 2689496,
      "author_name": "Dmitriy Guller",
      "author_url": "",
      "post_date": "2024-03-09T21:42:39.613000",
      "content": "<p>I don't see the concept of manually checking submissions for hacking being tenable.  The metric is still hackable, and the gains from hacking seem to be vastly bigger than the gains from legitimate model development.  I think a lot of competitors will call the bluff of the organizers, because if the choice is between having a submission in the top 100 and hoping it won't be disqualified, and having an uncompetitive legitimate solution, it's not really a choice at all.  You can't have a good Kaggle competition with a bad metric, no matter how many kludges you try to patch it with.</p>",
      "votes": 19,
      "replies": [
        {
          "id": 2690198,
          "author_name": "Daniel Herman",
          "author_url": "",
          "post_date": "2024-03-10T11:22:06.420000",
          "content": "<p>With all due respect, I am tired of the discussion regarding metric. I don't think this topic deserves all that attention. We considered changes in metric, but we didn't find an option that would not deviate from our initial intentions. Saying the metric is bad is little of an overstatement as it approximates well the business decisions. We encourage you to contribute to this competition. With that said this decision of keeping the metric is final.</p>",
          "votes": -25,
          "replies": [
            {
              "id": 2691057,
              "author_name": "Jacoby Jaeger",
              "author_url": "",
              "post_date": "2024-03-11T02:32:29.593000",
              "content": "<p>With all due respect, I'm sure you also thought the metric was fine before it was hacked the first time. Practices that work well enough in production aren't always suitable in competitions, and those of us who have experience with kaggle have seen competitions hacked that didn't have half the glaring vulnerabilities that this one does.  </p>",
              "votes": 22,
              "replies": []
            },
            {
              "id": 2692139,
              "author_name": "Simon Veitner",
              "author_url": "",
              "post_date": "2024-03-11T17:46:38.633000",
              "content": "<p>i seriously don't understand how it could happen that the metric still stays like it is although for everyone with experience in kaggle its clear what the outcome will be and there was provably unhackable metric suggested. for me this competition is over.</p>",
              "votes": 9,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2704098,
      "author_name": "Danu A.",
      "author_url": "",
      "post_date": "2024-03-18T15:01:23.410000",
      "content": "<p>This info should be in data description as the new participants that didn't read the discussions to know about this, otherwise they can use date_decision, MONTH and WEEK_NUM without a clue that in the test these are modified.</p>",
      "votes": 13,
      "replies": [
        {
          "id": 2705816,
          "author_name": "Daniel Herman",
          "author_url": "",
          "post_date": "2024-03-19T14:39:13.807000",
          "content": "<p>Thank you for noticing that, I will add it there. </p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 2715402,
      "author_name": "Dima",
      "author_url": "",
      "post_date": "2024-03-25T13:27:44.610000",
      "content": "<p>Hi, I'm interested about these tables and columns:</p>\n<pre><code>{\n        : [\n            ,\n            ,\n            ,\n            ,\n            ,\n            ,\n            ,\n            ,\n        ],\n        : [\n            ,\n            ,\n            ,\n            ,\n        ],\n        : [\n            ,\n            ,\n            ,\n            ,\n        ],\n    }\n</code></pre>\n<p>These columns also indicate date, although they are not marked as such. Were they transformed also? If yes, are time deltas like <code>dpdmaxdateyear_742T</code> - <code>date_decision</code> preserved?</p>",
      "votes": 8,
      "replies": [
        {
          "id": 2726414,
          "author_name": "minhtu.mt.mt",
          "author_url": "",
          "post_date": "2024-04-01T07:13:47.630000",
          "content": "<p><a href=\"https://www.kaggle.com/jetakow\" target=\"_blank\">@jetakow</a> <a href=\"https://www.kaggle.com/tomasjeline2\" target=\"_blank\">@tomasjeline2</a> can anyone answer this?</p>",
          "votes": 4,
          "replies": [
            {
              "id": 2830136,
              "author_name": "ano",
              "author_url": "",
              "post_date": "2024-05-23T03:01:31.737000",
              "content": "<p><a href=\"https://www.kaggle.com/jetakow\" target=\"_blank\">@jetakow</a> <a href=\"https://www.kaggle.com/tomasjeline2\" target=\"_blank\">@tomasjeline2</a> already answered this somewhere?</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2832872,
              "author_name": "Daniel Herman",
              "author_url": "",
              "post_date": "2024-05-23T22:58:04.880000",
              "content": "<p>Please tag active competition hosts <a href=\"https://www.kaggle.com/tomasjeline2\" target=\"_blank\">@tomasjeline2</a> <a href=\"https://www.kaggle.com/vojtechpacak\" target=\"_blank\">@vojtechpacak</a> </p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2688147,
      "author_name": "tetsuro731",
      "author_url": "",
      "post_date": "2024-03-09T03:43:13.823000",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/jetakow\" target=\"_blank\">@jetakow</a> </p>\n<p>I'm really happy to hear that the competition will be coming back.<br>\nI appreciate the proactive replacement of the test data by the host.<br>\nHowever, I think there are still two problems with not changing the current metric.</p>\n<h2>1. Indirect hacking is possible</h2>\n<p>Even if direct hacking is not possible with the correction of the test data, it doesn't change the fact that this metric is hackable. <br>\nFor example, one could create a model to predict <code>NUM_WEEK</code> in test data and intentionally worsen the scores of smaller <code>NUM_WEEK</code>. <br>\nI don't know whether such hacking is truly possible, but most participants who are genuinely improving their models in legitimate ways would feel uneasy about continuing to use this metric.</p>\n<h2>2. Problems with the metric itself</h2>\n<p>The image in the following post by <a href=\"https://www.kaggle.com/chumajin\" target=\"_blank\">@chumajin</a> is easy to explain, which is also attached:</p>\n<p><a href=\"https://www.kaggle.com/competitions/home-credit-credit-risk-model-stability/discussion/476867#2651259\" target=\"_blank\">https://www.kaggle.com/competitions/home-credit-credit-risk-model-stability/discussion/476867#2651259</a></p>\n<p>Comparing case2 and case3, it's clear that case3 is the superior model, yet under the current metric, the winner would be case2. I think this is not what the host truly wants to get.</p>\n<p>I feel that the current metric is too focused on stability, neglecting the original goal of improving AUC.</p>\n<p>Therefore, honestly, I'm concerned about these issues.</p>",
      "votes": 8,
      "replies": [
        {
          "id": 2688708,
          "author_name": "Daniel Herman",
          "author_url": "",
          "post_date": "2024-03-09T12:12:41.833000",
          "content": "<blockquote>\n  <p>Even if direct hacking is not possible with the correction of the test data, it doesn't change the fact that this metric is hackable.</p>\n</blockquote>\n<p>This is valid point, what you may consider is that there is limited time in this competition and as of now, the hacking is no longer feasible from a time-use stance. Besides that we discourage direct hacking and we will be reviewing the potential winners' work. </p>\n<blockquote>\n  <p>I feel that the current metric is too focused on stability, neglecting the original goal of improving AUC.</p>\n</blockquote>\n<p>Maybe we didn't communicate it clearly in previous posts, but the focus on stability is intended. Exploring the stability in model preparation is the target focus in this competition and it has a great importance in Home Credit. In fact, it is so important we decided to stick to the current metric because it covers the near-term performance extrapolation. </p>",
          "votes": -5,
          "replies": [
            {
              "id": 2688785,
              "author_name": "tetsuro731",
              "author_url": "",
              "post_date": "2024-03-09T13:12:34.880000",
              "content": "<p><a href=\"https://www.kaggle.com/jetakow\" target=\"_blank\">@jetakow</a> <br>\nThank you for your reply.<br>\nWhat you explained makes sense to me.<br>\nLet me ask one more question.</p>\n<pre><code>we will be reviewing the potential winners' work.\n</code></pre>\n<p>If you review winner's solution and find out that the solution was hacked or gamed, will they lose their prize? or in the worst case, will they be banned?</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2688805,
              "author_name": "Daniel Herman",
              "author_url": "",
              "post_date": "2024-03-09T13:30:18.137000",
              "content": "<p>We need to discuss that internally, but I think the most likely scenario as of now it that they won't be able to win the prize if the winning is based on the metric hack. </p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2688810,
              "author_name": "tetsuro731",
              "author_url": "",
              "post_date": "2024-03-09T13:41:27.210000",
              "content": "<p>Okay.<br>\nI hope all the potential winners will follow the legitimate way to present a great solution.<br>\nAnd of course I will enjoy this competition as well.<br>\nThank you.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2691381,
              "author_name": "",
              "author_url": "",
              "post_date": "2024-03-11T07:49:18.513000",
              "content": "",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2748434,
              "author_name": "",
              "author_url": "",
              "post_date": "2024-04-12T12:51:49.763000",
              "content": "",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2693302,
      "author_name": "Tomas Jelinek",
      "author_url": "",
      "post_date": "2024-03-12T10:30:57.250000",
      "content": "<p>Dear Kagglers,</p>\n<p>After many days our competition is running again. I wish you all success and want to thank you again for your patience.</p>\n<p>There were many comments lately with suggestions how the competition, metric, data, etc. should be designed or changed. I am thankful for many of your suggestions - really last weeks were very inspiring and at the same time very challenging for us. I feel I have to address some of your comments…</p>\n<p>When we were preparing this competition we knew from the beginning that we don't want to just repeat Home Credit competition from 2018. We have decided to add stability into competition objectives. Why? Because we think we optimize stability let's say expertly in our corporation and that there should be better way how to address this issue.</p>\n<p>We understood that our metric was more complex than plain vanilla AUC and there is risk that Kaggle community will find approaches how to optimize it in a way we haven't thought about. Unfortunately this risk materialized and we decided to react. There were basically 2 options: a) change metric so it's not hackable or b) change data so it does not contain information needed for hack. We have investigated many potential metrics (thank you Kaggle community for proposals), however we were not able to find a metric that is good enough to describe \"good model\", moreover many metrics were too complex and would be very hard to interpret. Therefore we have decided for option to adjust test data sample.</p>\n<p>Test data sample is adjusted, we can not guarantee that the changes we made in the fix will prevent all attempts of hacking, we acknowledge that Kagglers are very creative and smart, but we believe that the changes in test data make the hack attempts a magnitude harder to pull off. Obviously we don't want to disclose much information about new test data sample to limit the risk. What we wanted to say was already said in discussion. I would like to recommend you to switch your focus from hacking the metric :-)</p>\n<p>Keep in mind that competitions are here for Kagglers, but also for the Host - please respect that. We have also objectives that we want to achieve. We did our best to adjust data sample in a way that minimize risk of future hack and at the same time keep as much information for your models. We will keep the metric as it is.</p>\n<p>Tomas</p>",
      "votes": 6,
      "replies": [
        {
          "id": 2693581,
          "author_name": "Danu A.",
          "author_url": "",
          "post_date": "2024-03-12T14:31:45.153000",
          "content": "<p>Can you tell us your current working model what score have if trained on the current training set? Or at least a segment of what a good score can represent if compared of what you have now? It's interesting to compare our results with it, not just between us on LB.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 2695530,
          "author_name": "",
          "author_url": "",
          "post_date": "2024-03-13T18:04:01.640000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2796995,
      "author_name": "yuanzhe zhou",
      "author_url": "",
      "post_date": "2024-05-06T13:55:39.020000",
      "content": "<p>Will solution using metric trick will be disqualified?</p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 2691384,
      "author_name": "narsil (jobs-in-data.com)",
      "author_url": "",
      "post_date": "2024-03-11T07:51:06.767000",
      "content": "<p>Thank you for your update and for your continued effort to improve this competition.</p>\n<p>When you review top solutions manually, and you find out that a solution involves metric hacking, which prizes is such a solution ineligible for?<br>\n1/ money prize<br>\n2/ Kaggle medals<br>\n3/ Kaggle ranking points</p>\n<p>FYI, 2/ and 3/ are the most valuable currency on Kaggle, definitely not 1/.</p>",
      "votes": 5,
      "replies": [
        {
          "id": 2691392,
          "author_name": "Gunes Evitan",
          "author_url": "",
          "post_date": "2024-03-11T07:56:04.177000",
          "content": "<p>Yeah, I'll start this competition based on this. I also mentioned the same thing but got no reply.</p>",
          "votes": 3,
          "replies": [
            {
              "id": 2692142,
              "author_name": "Simon Veitner",
              "author_url": "",
              "post_date": "2024-03-11T17:48:33.353000",
              "content": "<p>even if every submission in top 100 will get checked there will be many people who don't recieve a medal because places &gt; 100 will contain large amount of people using metric hacking i think.</p>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 2693312,
              "author_name": "Gunes Evitan",
              "author_url": "",
              "post_date": "2024-03-12T10:42:32.127000",
              "content": "<p>Yes, and at the top positions people will use two subs for</p>\n<ol>\n<li>best score utilizing metric hacking for medals</li>\n<li>best score without hacking for the price money</li>\n</ol>",
              "votes": 3,
              "replies": []
            }
          ]
        },
        {
          "id": 2698128,
          "author_name": "narsil (jobs-in-data.com)",
          "author_url": "",
          "post_date": "2024-03-15T10:08:10.010000",
          "content": "<p><a href=\"https://www.kaggle.com/jetakow\" target=\"_blank\">@jetakow</a> Would you like to comment?</p>",
          "votes": -1,
          "replies": [
            {
              "id": 2786545,
              "author_name": "fanhanxiao",
              "author_url": "",
              "post_date": "2024-05-01T10:49:54.003000",
              "content": "<p>just try to comment something</p>",
              "votes": -1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2728422,
      "author_name": "zhamihong",
      "author_url": "",
      "post_date": "2024-04-02T08:32:50.643000",
      "content": "<p>May I ask a question? If column 'date_decision' in test_base table is changed, I can no longer transform date (end with D) columns to days by minusing 'date_decision', right? </p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 2686584,
      "author_name": "Élio Pereira",
      "author_url": "",
      "post_date": "2024-03-07T23:18:47.077000",
      "content": "<p>Just to clarify: WEEK_NUM in the training data is constant but it does vary in the test data, right? That would allow the usage of the same metric for evaluation of the submitted results. Unfortunately, a constant WEEK_NUM in the training data would not allow one to perform a fit on the Gini coefficient , get the slope (the coefficient \"a\") and, therefore, priorly compute the stability scores for that data. But I do understand that by making WEEK_NUM in the training data, virtually, there will be no way to make our models be \"manually\" tweaked to better predict for future weeks.</p>",
      "votes": 3,
      "replies": [
        {
          "id": 2687010,
          "author_name": "Tomas Jelinek",
          "author_url": "",
          "post_date": "2024-03-08T07:51:53.490000",
          "content": "<p>Hi,<br>\nthere is no change in train data - WEEK_NUM is still there and contains correct value, i.e. assignment of case_id to weekly time window - you can compute metric including stability part…<br>\nChange is in test data - week_num is constant so it does not carry any information (we did not delete it completely so you don't need to adjust your codes). But for evaluation in leaderboards we use correct WEEK_NUM, it's just hidden from Kagglers.</p>",
          "votes": 5,
          "replies": [
            {
              "id": 2687161,
              "author_name": "Élio Pereira",
              "author_url": "",
              "post_date": "2024-03-08T10:23:39.840000",
              "content": "<p>It is more clear to me now. Thanks!</p>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 2693559,
              "author_name": "Alexander Holmberg",
              "author_url": "",
              "post_date": "2024-03-12T14:15:20.137000",
              "content": "<blockquote>\n  <p>week_num is constant so it does not carry any information<br>\n  But for evaluation in leaderboards we use correct WEEK_NUM</p>\n</blockquote>\n<p>I don't understand how this makes sense. Is WEEK_NUM a constant in the test set or not?</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2693570,
              "author_name": "Daniel Herman",
              "author_url": "",
              "post_date": "2024-03-12T14:20:00.583000",
              "content": "<p><a href=\"https://www.kaggle.com/alexanderholmberg\" target=\"_blank\">@alexanderholmberg</a> The fact that the WEEK_NUM is no longer in the test set doesn't imply we don't have this information while evaluating your submissions. It is just not present in test set you can load. </p>",
              "votes": 2,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2694308,
      "author_name": "Taichi Uemura",
      "author_url": "",
      "post_date": "2024-03-13T01:39:59.567000",
      "content": "<p>From the <a href=\"https://www.kaggle.com/competitions/home-credit-credit-risk-model-stability/data\" target=\"_blank\">Data</a> tab:</p>\n<blockquote>\n  <p>Each group of tables can comprise one or more individual tables. If a group contains more than one table, they are divided based on <code>WEEK_NUM</code>.</p>\n</blockquote>\n<p>Isn't it possible to guess <code>WEEK_NUM</code> from how large tables such as <code>credit_bureau_a_2</code> are split?</p>",
      "votes": 4,
      "replies": [
        {
          "id": 2696018,
          "author_name": "Taichi Uemura",
          "author_url": "",
          "post_date": "2024-03-14T02:30:46.373000",
          "content": "<p>I experimented a bit by adding noise to the predictions for <code>case_id</code>'s that appear in the first half of <code>test_credit_bureau_a_2_*</code> files. The public score increased from 0.5 to around 0.51.</p>\n<p><a href=\"https://www.kaggle.com/jetakow\" target=\"_blank\">@jetakow</a> This hack is not possible if the test data is shuffled. I recommend shuffling unless it's already done.</p>",
          "votes": 7,
          "replies": [
            {
              "id": 2698115,
              "author_name": "Daniel Herman",
              "author_url": "",
              "post_date": "2024-03-15T09:56:13.363000",
              "content": "<p><a href=\"https://www.kaggle.com/taichiuemura\" target=\"_blank\">@taichiuemura</a> Thank you for noticing, would you try to reproduce the idea once again now?</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2698175,
              "author_name": "Taichi Uemura",
              "author_url": "",
              "post_date": "2024-03-15T10:34:56.250000",
              "content": "<p>I did. The score decreases now.</p>",
              "votes": 2,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2688776,
      "author_name": "Evgeniia Grigoreva",
      "author_url": "",
      "post_date": "2024-03-09T13:02:46.817000",
      "content": "<p>Will the 'date_decision' column still hold meaningful information about actual dates? If yes, then I believe hacking the metric is still possible, and it can even be done silently, making it unclear whether it’s a hack or just a meaningful model-choice decision. If the hypothesis about a structural break due to COVID is correct and the negative slope in the test data is related to this structural break, then tweaking models to perform worse on the COVID sample (for instance, by training the model exclusively on non-COVID data) becomes feasible. This would particularly benefit those who probe LB, which is concerning. Especially if 'date_decision' has been altered by adding a constant or something similar.</p>",
      "votes": 4,
      "replies": [
        {
          "id": 2690200,
          "author_name": "Daniel Herman",
          "author_url": "",
          "post_date": "2024-03-10T11:24:08.117000",
          "content": "<p><a href=\"https://www.kaggle.com/eivolkova\" target=\"_blank\">@eivolkova</a> Good point, but we of course thought about it. I can't tell you what we did exactly. </p>",
          "votes": -4,
          "replies": [
            {
              "id": 2690520,
              "author_name": "Evgeniia Grigoreva",
              "author_url": "",
              "post_date": "2024-03-10T16:01:24.210000",
              "content": "<p>Good to hear that, thank you for the answer!</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2686628,
      "author_name": "Jack Lee",
      "author_url": "",
      "post_date": "2024-03-08T00:14:48.397000",
      "content": "<blockquote>\n  <p>Column date_decision is no longer the same.</p>\n</blockquote>\n<p>Does that mean 'date_decision' will be meaningless, so we'd better not use it for training? Or has it been already dropped?</p>",
      "votes": 4,
      "replies": [
        {
          "id": 2687007,
          "author_name": "Tomas Jelinek",
          "author_url": "",
          "post_date": "2024-03-08T07:46:26.293000",
          "content": "<p>No, date_decision still exist, it can be used in models or in feature engineering. We did transformation that should make it impossible to hack metric in a way that was discussed after competition launch.</p>",
          "votes": 6,
          "replies": [
            {
              "id": 2745454,
              "author_name": "sahith kumar yedakula",
              "author_url": "",
              "post_date": "2024-04-10T16:34:08.670000",
              "content": "<p>yes yes yes</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2829594,
      "author_name": "Leon Zhou",
      "author_url": "",
      "post_date": "2024-05-22T17:00:52.107000",
      "content": "<p>Some thoughts on stablility metrics:</p>\n<p>1) when using prediciton models on production, we acknowledge a hypothesis. The future == The past .<br>\n2) but the real life is : in the past 4 years, we have expericed a greate lock down of Covid -19.  The world has been changed a lot. People's behavior also changed a lot. So The future ^=The past already.<br>\n3) Looking at the decison date on train set, we notice most of them are during the Covid-19 epidemic.<br>\n4) It can be predicted the vintage curve of those applicants will be worse than the population before Covid-19,even though they has same characteristic on demographic and bureau. </p>\n<p>5) the 3rd party data regarding macro economics might be useful.<br>\n6) Other method like risk table , which is typically used on anti-fraud, mighe be useful for risk guys to adjust model and strategy on time manner.  But it requires a target label at time point.</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 2734509,
      "author_name": "JasonHuang20040812",
      "author_url": "",
      "post_date": "2024-04-04T06:49:10.863000",
      "content": "<p>Hi, I am not sure if I understand it correctly, but isn't that the fact that MONTH and WEEK_NUM columns can be directly calculated from date_decision?<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F18899077%2Ff9916bef40ff638c6d0322a90c569e4b%2F2024-04-04%20144508.png?generation=1712213197463571&amp;alt=media\" alt=\"MONTH\"><br>\nif so, what is the meaning of setting WEEK_NUM and MONTH as constant in the given test data?<br>\nWhat's more, do the date_decision corresponds to real date when the decision is made? I am thinking about important relevant US interest rate and other data to facilitate my modeling?<br>\nCan anyone help me out? many thanks.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2736550,
          "author_name": "",
          "author_url": "",
          "post_date": "2024-04-05T09:10:28.487000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2715040,
      "author_name": "minhtu.mt.mt",
      "author_url": "",
      "post_date": "2024-03-25T08:31:02.643000",
      "content": "<p>Thanks for your update, I appreciate the effort from the host and community to improve the quality of this competition.<br>\nAccording to the changes you mentioned:</p>\n<blockquote>\n  <p>As we indicated previously, we will be changing date columns in test_base table, namely date_decision, MONTH and WEEK_NUM. The MONTH and WEEK_NUM columns will now only have one constant value. We decided to keep them to minimize the amount of time you'll need to spend altering your scripts to get the submissions working again. <strong>Column date_decision is no longer the same</strong></p>\n</blockquote>\n<p>May I ask how did you change the \"date_decision\" column? For MONTH and WEEK_NUM, we know that they will be constant and we can't use them as information for the model. But for \"date_decision\" (and also other datetime columns), I have no idea whether I can use it anymore.<br>\nAs the most voted notebooks use \"date_decision\" as an anchor to process all other datetime columns (e.g. <code>df = df.with_columns(pl.col(col) - pl.col(\"date_decision\"))</code>), this pre-processing may be wrong if the behavior of \"date_decision\" is different in the test dataset.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2715068,
          "author_name": "Tomas Jelinek",
          "author_url": "",
          "post_date": "2024-03-25T08:55:05.357000",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/minhtu123\" target=\"_blank\">@minhtu123</a> ,<br>\nWe don't want to disclose details about how date attributes were changed, but we did it in a way that keeps most of the information (at least from our point of view). Your example with date differences should not be affected by the change we have done.</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2715139,
              "author_name": "minhtu.mt.mt",
              "author_url": "",
              "post_date": "2024-03-25T09:53:30.067000",
              "content": "<p>Thanks for your reply.</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2691979,
      "author_name": "Ablay Dosmaganbetov",
      "author_url": "",
      "post_date": "2024-03-11T15:52:10.223000",
      "content": "<p>Thanks for the updates, I have a couple of questions.</p>\n<ol>\n<li><p>When 'Submissions' tab (next to 'Team' tab) will appear? I can't access my previous submissions anymore.</p></li>\n<li><p>May be I missed this information somewhere, but what will happen with our previous submissions and scores? Will they be left as they are? Or will they be corrected considering new metric updates? How this will affect ranking? </p></li>\n</ol>\n<p>Thanks for replies and sorry if I am asking such questions as they might be asked somewhere in the 'Discussion' tab.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2693303,
          "author_name": "Daniel Herman",
          "author_url": "",
          "post_date": "2024-03-12T10:31:39.617000",
          "content": "<blockquote>\n  <p>When 'Submissions' tab (next to 'Team' tab) will appear? I can't access my previous submissions anymore.</p>\n</blockquote>\n<p>It is there.</p>\n<blockquote>\n  <p>May be I missed this information somewhere, but what will happen with our previous submissions and scores? Will they be left as they are? Or will they be corrected considering new metric updates? How this will affect ranking?</p>\n</blockquote>\n<p>The LB was reset, your old submissions will work, but the scores changed completely. We encourage you to use the notebooks to generate a new submission table with new test data.</p>",
          "votes": -1,
          "replies": []
        }
      ]
    },
    {
      "id": 2686959,
      "author_name": "dima",
      "author_url": "",
      "post_date": "2024-03-08T07:18:35.760000",
      "content": "<blockquote>\n  <p>Column date_decision is no longer the same.</p>\n</blockquote>\n<p>It’s also not clear to me whether this column can be used to calculate new features or not? Let's say the number of loans for the last year?…</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2687004,
          "author_name": "Tomas Jelinek",
          "author_url": "",
          "post_date": "2024-03-08T07:44:04.890000",
          "content": "<p>Hi dima,<br>\nyes, such features are possible to calculate</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2687104,
              "author_name": "at7459",
              "author_url": "",
              "post_date": "2024-03-08T09:29:34.607000",
              "content": "<p>can you do, for example, \"date_decision\" - \"birthdate_87D\"? is that still the same number as it was before the transformation?</p>",
              "votes": 3,
              "replies": []
            },
            {
              "id": 2687162,
              "author_name": "Tomas Jelinek",
              "author_url": "",
              "post_date": "2024-03-08T10:25:20.603000",
              "content": "<p>yes, difference remains same</p>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 2687203,
              "author_name": "ARMADA",
              "author_url": "",
              "post_date": "2024-03-08T11:03:56.890000",
              "content": "<p><a href=\"https://www.kaggle.com/tomasjeline2\" target=\"_blank\">@tomasjeline2</a>  This is very confusing. </p>\n<p>How come the difference is still the same. Does this mean that you have also changed other dates in the test data like 'birthdate_87D' ? I was under the impressions that the only changed columns are date_decision, MONTH and WEEK_NUM as per the announcement.</p>\n<p>Can you please explain how did you change date_decision and not birthdate_87D and the difference is the same ?</p>",
              "votes": 3,
              "replies": []
            },
            {
              "id": 2687400,
              "author_name": "Tomas Jelinek",
              "author_url": "",
              "post_date": "2024-03-08T14:34:26.060000",
              "content": "<p>Hi Armada,<br>\nWe don't want to explain details about how data were transformed.<br>\nFrom the point of view of single case_id, we design transformation in a way that date attributes should contain same/similar information, features like client's age (date_app-birthdate) are not affected.</p>",
              "votes": -2,
              "replies": []
            },
            {
              "id": 2687930,
              "author_name": "ARMADA",
              "author_url": "",
              "post_date": "2024-03-08T21:40:17.660000",
              "content": "<p><a href=\"https://www.kaggle.com/tomasjeline2\" target=\"_blank\">@tomasjeline2</a>  Yes I fully understand that you cannot disclose how the transformation is done. My question was whether only these three columns are transformed : date_decision, MONTH and WEEK_NUM  or were there other date columns transformed and if it is the latter can you please list which other columns were transformed in the new test data.</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2688703,
              "author_name": "Daniel Herman",
              "author_url": "",
              "post_date": "2024-03-09T12:05:49.400000",
              "content": "<p>No we don't want to list what was exactly transformed and how and we will not do it. Please understand it that it is crucial for the sake of the fairness in this competition and at the same time it should not be a major concern of yours if you are working on feature engineering. </p>",
              "votes": -1,
              "replies": []
            },
            {
              "id": 2690521,
              "author_name": "Evgeniia Grigoreva",
              "author_url": "",
              "post_date": "2024-03-10T16:03:58.043000",
              "content": "<blockquote>\n  <p>From the point of view of single case_id</p>\n</blockquote>\n<p>What about aggregations based on time periods? For example, is it possible to calculate a rolling monthly average of some feature for the test data?</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2693944,
              "author_name": "at7459",
              "author_url": "",
              "post_date": "2024-03-12T18:57:03.927000",
              "content": "<p><a href=\"https://www.kaggle.com/jetakow\" target=\"_blank\">@jetakow</a> <a href=\"https://www.kaggle.com/tomasjeline2\" target=\"_blank\">@tomasjeline2</a> hey, thanks for continuing to be so responsive and all. tomas mentioned that \"date_decision\" - \"birthdate_87D\" is the same value as it was before, but i am not sure now if he was speaking only about this one example.</p>\n<p>is this true for EVERY column with D in it, so that i can still do \"date_decision\"-\"arbitrary D column\" and still get the same value as before the dataset change? i hope you can tell us about this, because these features are somewhat important for the model and checking each one with a new submission would be very time consuming…</p>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 2694148,
              "author_name": "Daniel Herman",
              "author_url": "",
              "post_date": "2024-03-12T22:25:18.027000",
              "content": "<p><a href=\"https://www.kaggle.com/at7459\" target=\"_blank\">@at7459</a> once again, we don't want to list what was exactly transformed and how and we will not do it. Please understand it that it is crucial for the sake of the fairness in this competition. You can try the date diff features and see what is working and what is not. That is afraid is the whole information we can give you. The FE is also part of the competition. </p>",
              "votes": -7,
              "replies": []
            },
            {
              "id": 2694149,
              "author_name": "Daniel Herman",
              "author_url": "",
              "post_date": "2024-03-12T22:27:13.890000",
              "content": "<p><a href=\"https://www.kaggle.com/eivolkova\" target=\"_blank\">@eivolkova</a> You can try those aggregations and see if any fo that is working or not. FE is also part of the competition and it should be left to be explored by kagglers. I can not guarantee you however that rolling monthly average will be working on test data. That is I am afraid all I can say.</p>",
              "votes": -3,
              "replies": []
            },
            {
              "id": 2694630,
              "author_name": "at7459",
              "author_url": "",
              "post_date": "2024-03-13T07:13:44.557000",
              "content": "<p>that is a bad decision, checking what is working and what is not realistically feasible and turns this competition into largely guessing</p>\n<p>for example, i can see by submissions that all these \"dateyear\" \"datemonth\" etc. were not properly transformed, i used to be able to subtract these from date decision and get a score boost, now it lowers my score by 0.005. you can't seriously expect me to go through every date feature and check by submission whether they were properly transformed or not…</p>",
              "votes": 13,
              "replies": []
            },
            {
              "id": 2694669,
              "author_name": "at7459",
              "author_url": "",
              "post_date": "2024-03-13T07:38:00.550000",
              "content": "<p>you said that the meaning of the features is preserved, but this clearly is not the case…</p>\n<p>i mean if you're going to manually check everything anyway i don't know why you would even want to transform anything in the first place</p>",
              "votes": 6,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2686533,
      "author_name": "Oleksiy Kononenko",
      "author_url": "",
      "post_date": "2024-03-07T21:52:08.940000",
      "content": "<p>Thanks for the update.</p>\n<blockquote>\n  <p>Ultimately, we don't want participants to focus on metric hacking; however, we also want a metric that closely approximates how models are assessed in production. When looking at the performance of models in production the predictable stable decrease in performance is undesirable even if it is not yet observed, but expected to come. We didn’t find a metric that would satisfy both requirements and that is why we decided to keep the metric as it is now.</p>\n</blockquote>\n<p>I’m not sure I understand this paragraph. How is it possible to keep the same metric and have no issues with metric hacking?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2686548,
          "author_name": "Daniel Herman",
          "author_url": "",
          "post_date": "2024-03-07T21:59:39.153000",
          "content": "<p>We believe the changes in test data will minimize metric hacking in terms of the gain.</p>",
          "votes": -2,
          "replies": [
            {
              "id": 2686552,
              "author_name": "Oleksiy Kononenko",
              "author_url": "",
              "post_date": "2024-03-07T22:06:54.870000",
              "content": "<p>I see, so “data transformation” means “data shuffling” in this context?</p>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 2686562,
              "author_name": "Daniel Herman",
              "author_url": "",
              "post_date": "2024-03-07T22:29:21.573000",
              "content": "<p>We won't give more detail on what exactly was changed in test set.</p>",
              "votes": -2,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2726400,
      "author_name": "Wanglaoji",
      "author_url": "",
      "post_date": "2024-04-01T06:58:37.980000",
      "content": "<p>Is weekday from date decision still makes sense in the test set?</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 2693347,
      "author_name": "Daryna Ronska",
      "author_url": "",
      "post_date": "2024-03-12T11:23:19.217000",
      "content": "<p>Is weekday from date decision still makes sense in the test set?</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 2692770,
      "author_name": "Abilash Reddy",
      "author_url": "",
      "post_date": "2024-03-12T04:13:12.090000",
      "content": "<p>Thanks for the update. Had trouble submitting yesterday but seems to work fine now. Excited! May the best win.</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 2690789,
      "author_name": "Roen Chiew",
      "author_url": "",
      "post_date": "2024-03-10T19:39:21.503000",
      "content": "<p>Can someone help ELI5 what's the problem with the WEEK_NUM?</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 2688747,
      "author_name": "skrrydg",
      "author_url": "",
      "post_date": "2024-03-09T12:36:36.733000",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/jetakow\" target=\"_blank\">@jetakow</a>!<br>\nWhat do you think about add hash(date_decision.weekday) as new feature in train &amp; test dataset?</p>",
      "votes": 2,
      "replies": [
        {
          "id": 2688778,
          "author_name": "tetsuro731",
          "author_url": "",
          "post_date": "2024-03-09T13:02:57.770000",
          "content": "<p>Sorry to jump in.<br>\nI think it's a great idea and I was thinking about the same thing.<br>\nAggregation like <code>xxx group by date_decision</code> is one of the good way to create new features, which can also be measured for actual industrial ML models.<br>\nHowever, obviously we can't use this way if all <code>date_decision</code> in test data become constant values.<br>\nThe problem is that we can hack and game metrics by directly using date/time information, so it would not be problem if we were to add hashed date features as you said.</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2689870,
              "author_name": "skrrydg",
              "author_url": "",
              "post_date": "2024-03-10T07:08:16.190000",
              "content": "<p>I think, they preserve correct weekday, while  transform the data</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2686629,
      "author_name": "Antonio Félix",
      "author_url": "",
      "post_date": "2024-03-08T00:15:59.050000",
      "content": "<p>Hi Daniel,</p>\n<blockquote>\n  <p>The remaining data were also transformed. Don't worry; the transformation will preserve the meaning of features. Data types will remain the same, and there shouldn't be significant changes in the distributions.</p>\n</blockquote>\n<p>Will the training dataset also be update with this transformations?</p>\n<p>Thanks.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2687949,
          "author_name": "",
          "author_url": "",
          "post_date": "2024-03-08T22:03:26.147000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 2688697,
          "author_name": "Daniel Herman",
          "author_url": "",
          "post_date": "2024-03-09T12:01:43.850000",
          "content": "<p>Training data were not a subject of this announcement implying they will remain untouched. </p>",
          "votes": -1,
          "replies": []
        }
      ]
    },
    {
      "id": 2748444,
      "author_name": "RUISHAN QU",
      "author_url": "",
      "post_date": "2024-04-12T12:53:56.267000",
      "content": "<p>请问您当前的工作模型如果在当前训练集上训练的话得分是多少？好分数可以代表什么？</p>",
      "votes": -2,
      "replies": []
    },
    {
      "id": 2748605,
      "author_name": "SUNYang",
      "author_url": "",
      "post_date": "2024-04-12T14:02:34.607000",
      "content": "<p>Hello dear host, I know that there is a data file in this data set called train_applprev_1_0.csv, which represents the user's historical loans and the status of these loans. The feature: <code>status_219L</code>, which has multiple values, represented by letters A, T, D, etc. However, the state represented by each letter does not seem to be described, which makes me confused. Because I want to consider the stability of the model from the perspective of historical loans, because I need to know the specific loan status represented by these letters, especially the order of the letters. This is particularly important to me. Can you provide a specific explanation of these feature value?</p>",
      "votes": -1,
      "replies": []
    },
    {
      "id": 2845430,
      "author_name": "杨成",
      "author_url": "",
      "post_date": "2024-05-30T14:59:18.077000",
      "content": "<p>Good Solution</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2826426,
      "author_name": "Abdullahi Dahir",
      "author_url": "",
      "post_date": "2024-05-20T23:01:17.030000",
      "content": "<p>I am getting this \"TypeError: the truth value of a Series is ambiguous\" when trying to use the Polars library for data cleaning, anyone has idea on how to deal with this.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2825947,
      "author_name": "Wong334",
      "author_url": "",
      "post_date": "2024-05-20T16:33:34.773000",
      "content": "<p>May I know if we can still upload our models even after the deadline just to check our score? I am aware that we will not be allowed to change our final result of the competition, just asking if we could upload.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2814641,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-05-15T13:17:17.703000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2810703,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-05-13T12:09:18.603000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2805625,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-05-10T16:44:29.117000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 2806283,
          "author_name": "",
          "author_url": "",
          "post_date": "2024-05-11T02:27:00.567000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2795787,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-05-06T01:13:20.133000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 2796911,
          "author_name": "",
          "author_url": "",
          "post_date": "2024-05-06T13:13:17.260000",
          "content": "",
          "votes": 0,
          "replies": [
            {
              "id": 2797860,
              "author_name": "",
              "author_url": "",
              "post_date": "2024-05-07T01:41:23.630000",
              "content": "",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2786620,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-05-01T11:34:32.337000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2778367,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-04-27T06:08:11.340000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2767508,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-04-22T11:06:35.030000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 2797090,
          "author_name": "",
          "author_url": "",
          "post_date": "2024-05-06T14:27:28.833000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2763720,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-04-20T17:47:25.823000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2754653,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-04-16T06:19:45.507000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 2754783,
          "author_name": "",
          "author_url": "",
          "post_date": "2024-04-16T07:45:11.623000",
          "content": "",
          "votes": 0,
          "replies": [
            {
              "id": 2756787,
              "author_name": "",
              "author_url": "",
              "post_date": "2024-04-17T07:16:30.697000",
              "content": "",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2743368,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-04-09T12:16:33.597000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2719179,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-03-27T14:55:07.410000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 2719183,
          "author_name": "",
          "author_url": "",
          "post_date": "2024-03-27T14:56:42.243000",
          "content": "",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 2712073,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-03-23T09:33:34.753000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2697631,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-03-15T03:24:26.857000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 2698201,
          "author_name": "",
          "author_url": "",
          "post_date": "2024-03-15T10:49:04.903000",
          "content": "",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 2693752,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-03-12T16:21:54.637000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 2693835,
          "author_name": "",
          "author_url": "",
          "post_date": "2024-03-12T17:48:08.930000",
          "content": "",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 2691252,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-03-11T06:35:22.027000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2687656,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-03-08T17:13:52.587000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 2688701,
          "author_name": "",
          "author_url": "",
          "post_date": "2024-03-09T12:04:04.723000",
          "content": "",
          "votes": -5,
          "replies": []
        }
      ]
    },
    {
      "id": 2686555,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-03-07T22:11:38.967000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 2686564,
          "author_name": "",
          "author_url": "",
          "post_date": "2024-03-07T22:31:59.150000",
          "content": "",
          "votes": 0,
          "replies": [
            {
              "id": 2686910,
              "author_name": "",
              "author_url": "",
              "post_date": "2024-03-08T06:17:04.017000",
              "content": "",
              "votes": 5,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2771798,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-04-24T12:10:20.580000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2748598,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-04-12T14:00:48.200000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2705487,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-03-19T11:21:11.270000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2692378,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-03-11T20:28:37.677000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2686795,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-03-08T05:05:35.267000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2803335,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-05-09T12:36:06.693000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2802854,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-05-09T07:40:22.507000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2800204,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-05-08T05:55:17.760000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2780164,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-04-28T04:53:11.983000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2766318,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-04-21T16:17:48.063000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2762769,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-04-20T04:49:31.233000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2724298,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-03-30T19:28:05.297000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2686523": "Dear kagglers,\n\nas promised we are back with an update. During the last two weeks, the submissions were paused due to metric exploitation and we decided to make changes to prevent participants from gaining a huge advantage by using this hack, thereby steering the focus of Kagglers back to the intended topic, which is the long-term stability of the model.\n\nAs we indicated previously, we will be changing date columns in test_base table, namely date_decision, MONTH and WEEK_NUM. The MONTH and WEEK_NUM columns will now only have one constant value. We decided to keep them to minimize the amount of time you'll need to spend altering your scripts to get the submissions working again. Column date_decision is no longer the same. The remaining data were also transformed. Don't worry; the transformation will preserve the meaning of features. Data types will remain the same, and there shouldn't be significant changes in the distributions. If you are aiming to develop a stable well-performing model, this change should not concern you at all. \n\nWe discussed metric change internally and with kagglers here. We received many great ideas and we are foremost very thankful for your time and showed passion for the challenge. We appreciate the effort that the community is willing to make for the sake of a more fair metric within this competition. The stability metric currently has an issue: it is possible to artificially worsen the score on some observations and yet improve the overall stability score. This might seem undesirable or strange to many of you, but there are reasons for this behaviour. Ultimately, we don't want participants to focus on metric hacking; however, we also want a metric that closely approximates how models are assessed in production. When looking at the performance of models in production the predictable stable decrease in performance is undesirable even if it is not yet observed, but expected to come. We didn’t find a metric that would satisfy both requirements and that is why we decided to keep the metric as it is now.\n\nThe fix is prepared and will be applied to the competition on Monday, 11th March 2024 together with resuming submissions. Consequently, the entire competition will be extended by three weeks. We want to encourage you again to focus on the stability topic. We wish everyone good luck with the data exploration, modelling and submitting results. \n\nTomas and Daniel\n",
    "2689496": "I don't see the concept of manually checking submissions for hacking being tenable.  The metric is still hackable, and the gains from hacking seem to be vastly bigger than the gains from legitimate model development.  I think a lot of competitors will call the bluff of the organizers, because if the choice is between having a submission in the top 100 and hoping it won't be disqualified, and having an uncompetitive legitimate solution, it's not really a choice at all.  You can't have a good Kaggle competition with a bad metric, no matter how many kludges you try to patch it with.",
    "2704098": "This info should be in data description as the new participants that didn't read the discussions to know about this, otherwise they can use date_decision, MONTH and WEEK_NUM without a clue that in the test these are modified.",
    "2715402": "Hi, I'm interested about these tables and columns:\n\n```\n{\n        \"credit_bureau_a_1\": [\n            \"dpdmaxdatemonth_442T\",\n            \"dpdmaxdatemonth_89T\",\n            \"dpdmaxdateyear_596T\",\n            \"dpdmaxdateyear_896T\",\n            \"overdueamountmaxdatemonth_284T\",\n            \"overdueamountmaxdatemonth_365T\",\n            \"overdueamountmaxdateyear_2T\",\n            \"overdueamountmaxdateyear_994T\",\n        ],\n        \"credit_bureau_b_1\": [\n            \"dpdmaxdatemonth_804T\",\n            \"dpdmaxdateyear_742T\",\n            \"overdueamountmaxdatemonth_494T\",\n            \"overdueamountmaxdateyear_432T\",\n        ],\n        \"credit_bureau_a_2\": [\n            \"pmts_month_158T\",\n            \"pmts_month_706T\",\n            \"pmts_year_1139T\",\n            \"pmts_year_507T\",\n        ],\n    }\n```\n\nThese columns also indicate date, although they are not marked as such. Were they transformed also? If yes, are time deltas like `dpdmaxdateyear_742T` - `date_decision` preserved?",
    "2688147": "Hi @jetakow \n\nI'm really happy to hear that the competition will be coming back.\nI appreciate the proactive replacement of the test data by the host.\nHowever, I think there are still two problems with not changing the current metric.\n\n## 1. Indirect hacking is possible\n\nEven if direct hacking is not possible with the correction of the test data, it doesn't change the fact that this metric is hackable. \nFor example, one could create a model to predict `NUM_WEEK` in test data and intentionally worsen the scores of smaller `NUM_WEEK`. \nI don't know whether such hacking is truly possible, but most participants who are genuinely improving their models in legitimate ways would feel uneasy about continuing to use this metric.\n\n## 2. Problems with the metric itself\nThe image in the following post by @chumajin is easy to explain, which is also attached:\n\nhttps://www.kaggle.com/competitions/home-credit-credit-risk-model-stability/discussion/476867#2651259\n\nComparing case2 and case3, it's clear that case3 is the superior model, yet under the current metric, the winner would be case2. I think this is not what the host truly wants to get.\n\nI feel that the current metric is too focused on stability, neglecting the original goal of improving AUC.\n\nTherefore, honestly, I'm concerned about these issues.",
    "2693302": "Dear Kagglers,\n\nAfter many days our competition is running again. I wish you all success and want to thank you again for your patience.\n\nThere were many comments lately with suggestions how the competition, metric, data, etc. should be designed or changed. I am thankful for many of your suggestions - really last weeks were very inspiring and at the same time very challenging for us. I feel I have to address some of your comments...\n\nWhen we were preparing this competition we knew from the beginning that we don't want to just repeat Home Credit competition from 2018. We have decided to add stability into competition objectives. Why? Because we think we optimize stability let's say expertly in our corporation and that there should be better way how to address this issue.\n\nWe understood that our metric was more complex than plain vanilla AUC and there is risk that Kaggle community will find approaches how to optimize it in a way we haven't thought about. Unfortunately this risk materialized and we decided to react. There were basically 2 options: a) change metric so it's not hackable or b) change data so it does not contain information needed for hack. We have investigated many potential metrics (thank you Kaggle community for proposals), however we were not able to find a metric that is good enough to describe \"good model\", moreover many metrics were too complex and would be very hard to interpret. Therefore we have decided for option to adjust test data sample.\n\nTest data sample is adjusted, we can not guarantee that the changes we made in the fix will prevent all attempts of hacking, we acknowledge that Kagglers are very creative and smart, but we believe that the changes in test data make the hack attempts a magnitude harder to pull off. Obviously we don't want to disclose much information about new test data sample to limit the risk. What we wanted to say was already said in discussion. I would like to recommend you to switch your focus from hacking the metric :-)\n\nKeep in mind that competitions are here for Kagglers, but also for the Host - please respect that. We have also objectives that we want to achieve. We did our best to adjust data sample in a way that minimize risk of future hack and at the same time keep as much information for your models. We will keep the metric as it is.\n\nTomas\n",
    "2796995": "Will solution using metric trick will be disqualified?",
    "2691384": "Thank you for your update and for your continued effort to improve this competition.\n\nWhen you review top solutions manually, and you find out that a solution involves metric hacking, which prizes is such a solution ineligible for?\n1/ money prize\n2/ Kaggle medals\n3/ Kaggle ranking points\n\nFYI, 2/ and 3/ are the most valuable currency on Kaggle, definitely not 1/.",
    "2728422": "May I ask a question? If column 'date_decision' in test_base table is changed, I can no longer transform date (end with D) columns to days by minusing 'date_decision', right? ",
    "2686584": "Just to clarify: WEEK_NUM in the training data is constant but it does vary in the test data, right? That would allow the usage of the same metric for evaluation of the submitted results. Unfortunately, a constant WEEK_NUM in the training data would not allow one to perform a fit on the Gini coefficient , get the slope (the coefficient \"a\") and, therefore, priorly compute the stability scores for that data. But I do understand that by making WEEK_NUM in the training data, virtually, there will be no way to make our models be \"manually\" tweaked to better predict for future weeks.",
    "2694308": "From the [Data](https://www.kaggle.com/competitions/home-credit-credit-risk-model-stability/data) tab:\n\n> Each group of tables can comprise one or more individual tables. If a group contains more than one table, they are divided based on `WEEK_NUM`.\n\nIsn't it possible to guess `WEEK_NUM` from how large tables such as `credit_bureau_a_2` are split?",
    "2688776": "Will the 'date_decision' column still hold meaningful information about actual dates? If yes, then I believe hacking the metric is still possible, and it can even be done silently, making it unclear whether it’s a hack or just a meaningful model-choice decision. If the hypothesis about a structural break due to COVID is correct and the negative slope in the test data is related to this structural break, then tweaking models to perform worse on the COVID sample (for instance, by training the model exclusively on non-COVID data) becomes feasible. This would particularly benefit those who probe LB, which is concerning. Especially if 'date_decision' has been altered by adding a constant or something similar.",
    "2686628": ">Column date_decision is no longer the same.\n\nDoes that mean 'date_decision' will be meaningless, so we'd better not use it for training? Or has it been already dropped?",
    "2829594": "Some thoughts on stablility metrics:\n\n1) when using prediciton models on production, we acknowledge a hypothesis. The future == The past .\n2) but the real life is : in the past 4 years, we have expericed a greate lock down of Covid -19.  The world has been changed a lot. People's behavior also changed a lot. So The future ^=The past already.\n3) Looking at the decison date on train set, we notice most of them are during the Covid-19 epidemic.\n4) It can be predicted the vintage curve of those applicants will be worse than the population before Covid-19,even though they has same characteristic on demographic and bureau. \n\n5) the 3rd party data regarding macro economics might be useful.\n6) Other method like risk table , which is typically used on anti-fraud, mighe be useful for risk guys to adjust model and strategy on time manner.  But it requires a target label at time point.\n\n\n\n\n",
    "2734509": "Hi, I am not sure if I understand it correctly, but isn't that the fact that MONTH and WEEK_NUM columns can be directly calculated from date_decision?\n![MONTH](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F18899077%2Ff9916bef40ff638c6d0322a90c569e4b%2F2024-04-04%20144508.png?generation=1712213197463571&alt=media)\nif so, what is the meaning of setting WEEK_NUM and MONTH as constant in the given test data?\nWhat's more, do the date_decision corresponds to real date when the decision is made? I am thinking about important relevant US interest rate and other data to facilitate my modeling?\nCan anyone help me out? many thanks.",
    "2715040": "Thanks for your update, I appreciate the effort from the host and community to improve the quality of this competition.\nAccording to the changes you mentioned:\n>As we indicated previously, we will be changing date columns in test_base table, namely date_decision, MONTH and WEEK_NUM. The MONTH and WEEK_NUM columns will now only have one constant value. We decided to keep them to minimize the amount of time you'll need to spend altering your scripts to get the submissions working again. **Column date_decision is no longer the same**\n\nMay I ask how did you change the \"date_decision\" column? For MONTH and WEEK_NUM, we know that they will be constant and we can't use them as information for the model. But for \"date_decision\" (and also other datetime columns), I have no idea whether I can use it anymore.\nAs the most voted notebooks use \"date_decision\" as an anchor to process all other datetime columns (e.g. `df = df.with_columns(pl.col(col) - pl.col(\"date_decision\"))`), this pre-processing may be wrong if the behavior of \"date_decision\" is different in the test dataset.",
    "2691979": "Thanks for the updates, I have a couple of questions.\n\n1. When 'Submissions' tab (next to 'Team' tab) will appear? I can't access my previous submissions anymore.\n\n2. May be I missed this information somewhere, but what will happen with our previous submissions and scores? Will they be left as they are? Or will they be corrected considering new metric updates? How this will affect ranking? \n\nThanks for replies and sorry if I am asking such questions as they might be asked somewhere in the 'Discussion' tab.",
    "2686959": ">Column date_decision is no longer the same.\n\nIt’s also not clear to me whether this column can be used to calculate new features or not? Let's say the number of loans for the last year?...",
    "2686533": "Thanks for the update.\n\n>Ultimately, we don't want participants to focus on metric hacking; however, we also want a metric that closely approximates how models are assessed in production. When looking at the performance of models in production the predictable stable decrease in performance is undesirable even if it is not yet observed, but expected to come. We didn’t find a metric that would satisfy both requirements and that is why we decided to keep the metric as it is now.\n\nI’m not sure I understand this paragraph. How is it possible to keep the same metric and have no issues with metric hacking?",
    "2726400": "Is weekday from date decision still makes sense in the test set?\n\n\n",
    "2693347": "Is weekday from date decision still makes sense in the test set?",
    "2692770": "Thanks for the update. Had trouble submitting yesterday but seems to work fine now. Excited! May the best win.",
    "2690789": "Can someone help ELI5 what's the problem with the WEEK_NUM?",
    "2688747": "Hi @jetakow!\nWhat do you think about add hash(date_decision.weekday) as new feature in train & test dataset?",
    "2686629": "Hi Daniel,\n\n>The remaining data were also transformed. Don't worry; the transformation will preserve the meaning of features. Data types will remain the same, and there shouldn't be significant changes in the distributions.\n\nWill the training dataset also be update with this transformations?\n\nThanks.",
    "2748444": "请问您当前的工作模型如果在当前训练集上训练的话得分是多少？好分数可以代表什么？",
    "2748605": "Hello dear host, I know that there is a data file in this data set called train_applprev_1_0.csv, which represents the user's historical loans and the status of these loans. The feature: `status_219L`, which has multiple values, represented by letters A, T, D, etc. However, the state represented by each letter does not seem to be described, which makes me confused. Because I want to consider the stability of the model from the perspective of historical loans, because I need to know the specific loan status represented by these letters, especially the order of the letters. This is particularly important to me. Can you provide a specific explanation of these feature value?",
    "2845430": "Good Solution",
    "2826426": "I am getting this \"TypeError: the truth value of a Series is ambiguous\" when trying to use the Polars library for data cleaning, anyone has idea on how to deal with this.",
    "2825947": "May I know if we can still upload our models even after the deadline just to check our score? I am aware that we will not be allowed to change our final result of the competition, just asking if we could upload.",
    "2814641": "good n awesome!",
    "2810703": "Good Solution",
    "2805625": "Thank you for your work, I have learned  a lot from it.",
    "2795787": "Hello, since this competition is on hold for couple weeks,  I would like to know what is the maximum number of submissions allowed for a team.",
    "2786620": "Great competition",
    "2778367": "How do I download the dataset? ",
    "2767508": "great work",
    "2763720": "Thanks for Your response are really helpful here.",
    "2754653": "\nHello,\n\nThank you very much for your comprehensive response regarding the metrics. \n\nI have a question concerning WEEK_NUM in the test data. I understand that WEEK_NUM has been modified, suggesting that it has been encrypted in a way that disrupts its original order. If my understanding is correct, this encryption would allow those with decryption capabilities, such as yourselves, to perform calculations as intended. However, for us who cannot decrypt it, we are no longer able to do metric hacking. Could you please confirm if my understanding is correct?\n",
    "2743368": "Realizing I have so much to learn.",
    "2719179": "Hi,\n\nI use an open source machine learning tool called khiops. There is no internet acces, so Is it possible to install it in your notebook ?\n\nconda install -y -c conda-forge -c khiops khiops\n\nkhiops use a  BSD 3-Clause-clear License\n\nNicolas",
    "2712073": "I think this is an interesting idea",
    "2697631": "Hi, I have an unrelated question is that if the data is used for application scoring, I have seen that the record date in the train_tax_registry_table_a is always later than the decision date? This means that the info won't be available at the time of the scoring. How would the table be useful? or it's just that the data has been transformed and the true date is actual before the decision date so I can continue to use it in my model without any issues?",
    "2693752": "Does that mean we should be downloading the data again? ",
    "2691252": "Firstly, my regards to the competition hosts. It can be a very tough time for all of you.\n\nTaking a step back, I'd say perhaps the best way to emphasize stability is to work on the data horizon, e.g. predict future 5 years with only 1 year-worth of data, possibly giving higher weigh to more recent data. But it is too late now.",
    "2687656": "Hi Daniel, \n\nWhile you are updating the data in the test set (the one with only 10 rows), can you please fix the parquet files in the test set?  They actually have a different structure than the train set.\n\n- For example, file parquet_files/test/test_credit_bureau_a_1_3.parquet is not the correct format.  You get a failure when loading all the test_credit_bureau_a_1 files together.  I doubt that this will happen when we run on the hidden test set, but it would be great to be able to test our process on the small test set provided.  \n\nBelow, I illustrate an example where file credit_bureau_a_1_1 and credit_bureau_a_1_3 can't union together.  \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F522718%2F7cc98e88b2981c3df17d065ed4724a87%2FScreenshot%202024-03-08%20120230.png?generation=1709917399648920&alt=media)\n\nThe above shows a failure when trying to union those files together.  \nI would suggest to decouple the CSV and PARQUET process completely.  It feels like perhaps you created the PARQUETS after creating the CSV files?  \nParquets are terrific as I am sure you know ... less space and retains column types...  and I like that this competition has them.  \n\nIf you load like this -- \n```python\n\nLOAD DATA OVERWRITE `home_credit_comp.test_credit_bureau_a_1_1` FROM FILES (\n                format = 'PARQUET'\n                ,uris = ['gs://kds-2b8b4dc180f8d5e733a579005c44bbc1835fccdef4262654e2e5b687/parquet_files/test/test_credit_bureau_a_1_1.parquet']);\n        \n```\n\n        LOAD DATA OVERWRITE `home_credit_comp.test_credit_bureau_a_1_3` FROM FILES (\n                format = 'PARQUET',\n                uris = ['gs://kds-2b8b4dc180f8d5e733a579005c44bbc1835fccdef4262654e2e5b687/parquet_files/test/test_credit_bureau_a_1_3.parquet']);\n                \n        select * from `home_credit_comp.test_credit_bureau_a_1_1`\n        union all \n        select * from `home_credit_comp.test_credit_bureau_a_1_3`\n        ; \n\nThis also happened when importing the test_static_0 parquets...\n\n```python\n        LOAD DATA OVERWRITE `concrete-acre-416420.home_credit_comp.test_static_0`\n            FROM FILES (\n                format = 'PARQUET',\n                uris = ['gs://kds-2b8b4dc180f8d5e733a579005c44bbc1835fccdef4262654e2e5b687/parquet_files/test/test_static_0*.parquet']\n            )\n        ;\n```\n        ",
    "2686555": "You mentioned in an earlier post that you will be reviewing the top 100 solutions manually in order to make sure that they are not scoring high by hacking the metric. Is this still the case or will you just accept the score as is even if people found other ways to hack it ?",
    "2771798": "",
    "2748598": "",
    "2705487": "",
    "2692378": "",
    "2686795": "",
    "2803335": "Thank you!",
    "2802854": "thanks for hosters",
    "2800204": "Thanks for your efforts",
    "2780164": "thanks your efforts",
    "2766318": "Thanks for the update.",
    "2762769": "Thanks for sharing",
    "2724298": "Sure, Thank you for the update!"
  }
}