{
  "id": 478716,
  "title": "Announcement regarding competition metric - follow-up",
  "url": "/competitions/home-credit-credit-risk-model-stability/discussion/478716",
  "author_name": "",
  "post_date": "2024-02-21T21:48:35.606421300Z",
  "votes": 34,
  "comment_count": 30,
  "views": 0,
  "content": "<p>Dear Kagglers, </p>\n<p>We prepared this year’s competition with an intention to present you something else and more challenging than the previous Home Credit kaggle competition. We decided to select stability of models in production as the main objective of this competition. Therefore, we will keep stability and won’t remove it from the competition.  </p>\n<p>We would like to further clarify the meaning of stability. Our ideal stable model has constant performance by weeks in time and as high of a mean gini over weeks as possible. That is the reason why we have three terms in the stability metric as of now since this is desirable in business and not only to maximize AUC. The metric itself is a good approximation of the decisions made in-house when assessing the performance of models in production. Despite different opinions here within the competition, removing the terms in the stability metric would significantly deviate from the way we rate performance internally. When people leverage the second term and nullify it by carefully choosing scores on the test sample, it does not represent how we deploy models in the production. This is undesirable and we are forced to make steps that will discourage hacking of the metric.  </p>\n<p>We will make a change to the test dataset in such a way that it will represent better client scoring in production. When the model is running in production, it can only see its past performance and predictions. This would be quite difficult to reproduce in a Kaggle competition, but we have an alternative solution. We will remove the date decision, month and week number from the test base table. This part will take two weeks to amend, and we will prolong the competition by this period. Because of this, it is very likely that we won’t be able to rescore the old submissions. We apologize in advance for this inconvenience. </p>\n<p>Besides the change of test data, we have a second plan regarding metric. At this moment we are still considering metric change, and we welcome suggestions that would be aligned with how we view the value of model stability and in a way that would prevent further hacking. We are unlikely to go with a new metric that would prevent this type of hacking and would significantly deviate from the stability part. There is very little guarantee that a new metric would not open doors to a new type of hacking. This part will take at least a week.</p>\n<p>We would like to apologize for the delay and inconvenience caused by the issues connected to the stability metric. We also want to thank the community for many proposals on how to fix the issue. We have tested and analyzed many of them, thank you for all the suggestions! We will disable submissions for two weeks from now on and prolong the competition by two weeks.  </p>\n<p>Tomas &amp; Daniel</p>",
  "messages": [
    {
      "id": "2662418",
      "postDate": "02/21/2024 21:48:35",
      "content": "<p>Dear Kagglers, </p>\n<p>We prepared this year’s competition with an intention to present you something else and more challenging than the previous Home Credit kaggle competition. We decided to select stability of models in production as the main objective of this competition. Therefore, we will keep stability and won’t remove it from the competition.  </p>\n<p>We would like to further clarify the meaning of stability. Our ideal stable model has constant performance by weeks in time and as high of a mean gini over weeks as possible. That is the reason why we have three terms in the stability metric as of now since this is desirable in business and not only to maximize AUC. The metric itself is a good approximation of the decisions made in-house when assessing the performance of models in production. Despite different opinions here within the competition, removing the terms in the stability metric would significantly deviate from the way we rate performance internally. When people leverage the second term and nullify it by carefully choosing scores on the test sample, it does not represent how we deploy models in the production. This is undesirable and we are forced to make steps that will discourage hacking of the metric.  </p>\n<p>We will make a change to the test dataset in such a way that it will represent better client scoring in production. When the model is running in production, it can only see its past performance and predictions. This would be quite difficult to reproduce in a Kaggle competition, but we have an alternative solution. We will remove the date decision, month and week number from the test base table. This part will take two weeks to amend, and we will prolong the competition by this period. Because of this, it is very likely that we won’t be able to rescore the old submissions. We apologize in advance for this inconvenience. </p>\n<p>Besides the change of test data, we have a second plan regarding metric. At this moment we are still considering metric change, and we welcome suggestions that would be aligned with how we view the value of model stability and in a way that would prevent further hacking. We are unlikely to go with a new metric that would prevent this type of hacking and would significantly deviate from the stability part. There is very little guarantee that a new metric would not open doors to a new type of hacking. This part will take at least a week.</p>\n<p>We would like to apologize for the delay and inconvenience caused by the issues connected to the stability metric. We also want to thank the community for many proposals on how to fix the issue. We have tested and analyzed many of them, thank you for all the suggestions! We will disable submissions for two weeks from now on and prolong the competition by two weeks.  </p>\n<p>Tomas &amp; Daniel</p>",
      "rawMarkdown": "Dear Kagglers, \n \nWe prepared this year’s competition with an intention to present you something else and more challenging than the previous Home Credit kaggle competition. We decided to select stability of models in production as the main objective of this competition. Therefore, we will keep stability and won’t remove it from the competition.  \n\nWe would like to further clarify the meaning of stability. Our ideal stable model has constant performance by weeks in time and as high of a mean gini over weeks as possible. That is the reason why we have three terms in the stability metric as of now since this is desirable in business and not only to maximize AUC. The metric itself is a good approximation of the decisions made in-house when assessing the performance of models in production. Despite different opinions here within the competition, removing the terms in the stability metric would significantly deviate from the way we rate performance internally. When people leverage the second term and nullify it by carefully choosing scores on the test sample, it does not represent how we deploy models in the production. This is undesirable and we are forced to make steps that will discourage hacking of the metric.  \n \nWe will make a change to the test dataset in such a way that it will represent better client scoring in production. When the model is running in production, it can only see its past performance and predictions. This would be quite difficult to reproduce in a Kaggle competition, but we have an alternative solution. We will remove the date decision, month and week number from the test base table. This part will take two weeks to amend, and we will prolong the competition by this period. Because of this, it is very likely that we won’t be able to rescore the old submissions. We apologize in advance for this inconvenience. \n \nBesides the change of test data, we have a second plan regarding metric. At this moment we are still considering metric change, and we welcome suggestions that would be aligned with how we view the value of model stability and in a way that would prevent further hacking. We are unlikely to go with a new metric that would prevent this type of hacking and would significantly deviate from the stability part. There is very little guarantee that a new metric would not open doors to a new type of hacking. This part will take at least a week.\n \nWe would like to apologize for the delay and inconvenience caused by the issues connected to the stability metric. We also want to thank the community for many proposals on how to fix the issue. We have tested and analyzed many of them, thank you for all the suggestions! We will disable submissions for two weeks from now on and prolong the competition by two weeks.  \n\nTomas & Daniel",
      "votes": null
    },
    {
      "id": "2662696",
      "postDate": "02/22/2024 04:49:44",
      "content": "<p>If you remove some fields and transform the most of the fields in data, you should restart it with 3 months. It has already been 3 weeks wasted. There are too many files and most people waited for the metric fix for given dataset but now people probably will not be able to use most of their previous work. I think it would be more fair if you restart the timeline with 3 months.</p>",
      "rawMarkdown": "If you remove some fields and transform the most of the fields in data, you should restart it with 3 months. It has already been 3 weeks wasted. There are too many files and most people waited for the metric fix for given dataset but now people probably will not be able to use most of their previous work. I think it would be more fair if you restart the timeline with 3 months.",
      "votes": null
    },
    {
      "id": "2662815",
      "postDate": "02/22/2024 06:20:45",
      "content": "<blockquote>\n  <p>There is very little guarantee that a new metric would not open doors to a new type of hacking. This part will take at least a week.</p>\n</blockquote>\n<p>I've explained what makes the metric hackable: it is the fact that the overall score is not monotonic wrt to the sample scores. This is easy to test for, one just has to check whether the gradient of the the final score wrt to the round scores has strictly one sign. Only when this rule is violated is it possible to make the overall score better by making round scores worse. </p>\n<p>It seems absurd to me not to fix this problem on the excuse that it might create new problems. The fact is that most competitions manage to find metrics that aren't in any way hackable. It seems extremely unlikely to me that any reasonably designed metric that is provably monotonic to round scores is somehow hackable in another way.</p>",
      "rawMarkdown": ">There is very little guarantee that a new metric would not open doors to a new type of hacking. This part will take at least a week.\n\nI've explained what makes the metric hackable: it is the fact that the overall score is not monotonic wrt to the sample scores. This is easy to test for, one just has to check whether the gradient of the the final score wrt to the round scores has strictly one sign. Only when this rule is violated is it possible to make the overall score better by making round scores worse. \n\nIt seems absurd to me not to fix this problem on the excuse that it might create new problems. The fact is that most competitions manage to find metrics that aren't in any way hackable. It seems extremely unlikely to me that any reasonably designed metric that is provably monotonic to round scores is somehow hackable in another way.",
      "votes": null
    },
    {
      "id": "2663009",
      "postDate": "02/22/2024 09:23:06",
      "content": "<p>But there are a lot of date columns in the data other than date_decision which might be close enough to point to the approximate date of decision. I have not checked it myself yet, but something tells it might be the case</p>",
      "rawMarkdown": "But there are a lot of date columns in the data other than date_decision which might be close enough to point to the approximate date of decision. I have not checked it myself yet, but something tells it might be the case",
      "votes": null
    },
    {
      "id": "2663013",
      "postDate": "02/22/2024 09:25:05",
      "content": "<p>I appreciate the effort you put into the discussion and yes it would be ideal to have a metric that is monotonic in respect to every observation score and yet at the same time it can take into account the \"extrapolation\" of near-future performance as it is now with the gini stability metric. In case you can think of any such metric that would not be only monotonic wrt to the sample scores, but also preserves the idea of long-term decline of model. Please share it in <a href=\"https://www.kaggle.com/competitions/home-credit-credit-risk-model-stability/discussion/478699\" target=\"_blank\">https://www.kaggle.com/competitions/home-credit-credit-risk-model-stability/discussion/478699</a></p>",
      "rawMarkdown": "I appreciate the effort you put into the discussion and yes it would be ideal to have a metric that is monotonic in respect to every observation score and yet at the same time it can take into account the \"extrapolation\" of near-future performance as it is now with the gini stability metric. In case you can think of any such metric that would not be only monotonic wrt to the sample scores, but also preserves the idea of long-term decline of model. Please share it in https://www.kaggle.com/competitions/home-credit-credit-risk-model-stability/discussion/478699",
      "votes": null
    },
    {
      "id": "2663038",
      "postDate": "02/22/2024 09:33:14",
      "content": "<p>We are aware of that. Data in test will be transformed.</p>",
      "rawMarkdown": "We are aware of that. Data in test will be transformed.",
      "votes": null
    },
    {
      "id": "2663048",
      "postDate": "02/22/2024 09:35:35",
      "content": "<p>Hi Davut,<br>\nI agree with you. We will adjust deadline of competition accordingly to compensate your lost time. We have agreement with Kaggle team on this.</p>",
      "rawMarkdown": "Hi Davut,\nI agree with you. We will adjust deadline of competition accordingly to compensate your lost time. We have agreement with Kaggle team on this.",
      "votes": null
    },
    {
      "id": "2663126",
      "postDate": "02/22/2024 10:22:34",
      "content": "<p>But the data in train stays like it is?</p>",
      "rawMarkdown": "But the data in train stays like it is?",
      "votes": null
    },
    {
      "id": "2663290",
      "postDate": "02/22/2024 12:05:38",
      "content": "<p>Exactly, the train data stay the same.</p>",
      "rawMarkdown": "Exactly, the train data stay the same.",
      "votes": null
    },
    {
      "id": "2664954",
      "postDate": "02/23/2024 10:06:18",
      "content": "<p>If WEEK_NUM is removed, we won't be able to perform certain feature engineering tasks based on WEEK_NUM, such as quantile calculations.</p>",
      "rawMarkdown": "If WEEK_NUM is removed, we won't be able to perform certain feature engineering tasks based on WEEK_NUM, such as quantile calculations.",
      "votes": null
    },
    {
      "id": "2665286",
      "postDate": "02/23/2024 15:18:05",
      "content": "<p>请问测试数据中的日期将会如何转换？还是说日期字段会全部去掉？</p>",
      "rawMarkdown": "请问测试数据中的日期将会如何转换？还是说日期字段会全部去掉？",
      "votes": null
    },
    {
      "id": "2665386",
      "postDate": "02/23/2024 16:23:29",
      "content": "<p><a href=\"https://www.kaggle.com/boristown\" target=\"_blank\">@boristown</a> โปรดลองใช้นักแปลภาษาอังกฤษเพื่อให้ผู้ใช้เข้าใจได้ง่ายขึ้น</p>",
      "rawMarkdown": "boristown โปรดลองใช้นักแปลภาษาอังกฤษเพื่อให้ผู้ใช้เข้าใจได้ง่ายขึ้น",
      "votes": null
    },
    {
      "id": "2665546",
      "postDate": "02/23/2024 18:00:01",
      "content": "<p>I agree WEEK_NUM is crucial, at least for my models. Consider hashing WEEK_NUM instead of delete it, please. Removing the temporal component of WEEK_NUM seems enough for prevent time hacking.</p>",
      "rawMarkdown": "I agree WEEK_NUM is crucial, at least for my models. Consider hashing WEEK_NUM instead of delete it, please. Removing the temporal component of WEEK_NUM seems enough for prevent time hacking.",
      "votes": null
    },
    {
      "id": "2665835",
      "postDate": "02/23/2024 21:18:03",
      "content": "<p>We can make it constant and let the column there. However, the values of WEEK_NUM will be removed.</p>",
      "rawMarkdown": "We can make it constant and let the column there. However, the values of WEEK_NUM will be removed.",
      "votes": null
    },
    {
      "id": "2666394",
      "postDate": "02/24/2024 11:16:07",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/jetakow\" target=\"_blank\">@jetakow</a> &amp; <a href=\"https://www.kaggle.com/tomasjeline2\" target=\"_blank\">@tomasjeline2</a> <br>\nWhat do you mean by transformed ? Every column with 'D' at the end  which is not nan is a string that can be converted to date format. Do you mean your transformation will avoid the conversion to date format, for every column 'D' ? In that case it will be replaced by nan or a string <strong>whitout any meaning</strong> ? Thank you for your information</p>",
      "rawMarkdown": "Hi @jetakow & @tomasjeline2 \nWhat do you mean by transformed ? Every column with 'D' at the end  which is not nan is a string that can be converted to date format. Do you mean your transformation will avoid the conversion to date format, for every column 'D' ? In that case it will be replaced by nan or a string **whitout any meaning** ? Thank you for your information",
      "votes": null
    },
    {
      "id": "2666649",
      "postDate": "02/24/2024 14:15:14",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/pourchot\" target=\"_blank\">@pourchot</a>, all the columns in test data will stay, the only columns that may be removed are only in test_base, but after seeing responses in Discussion we will most likely leave them, just replace the values with constant (date decision, month and week number). Dtypes will remain same everywhere. Transformation of test data will remain unspecified. </p>",
      "rawMarkdown": "Hi @pourchot, all the columns in test data will stay, the only columns that may be removed are only in test_base, but after seeing responses in Discussion we will most likely leave them, just replace the values with constant (date decision, month and week number). Dtypes will remain same everywhere. Transformation of test data will remain unspecified.",
      "votes": null
    },
    {
      "id": "2666684",
      "postDate": "02/24/2024 14:44:28",
      "content": "<p>I'm trying to figure out how a feature can be transformed \"differently\" in the training set and the testing set and still remain useful. My understanding is that data has to be both structurally and \"semantically\" consistent across both training and testing sets to be useful. Nevertheless I'm sure it will be all clarified when it's ready.</p>",
      "rawMarkdown": "I'm trying to figure out how a feature can be transformed \"differently\" in the training set and the testing set and still remain useful. My understanding is that data has to be both structurally and \"semantically\" consistent across both training and testing sets to be useful. Nevertheless I'm sure it will be all clarified when it's ready.",
      "votes": null
    },
    {
      "id": "2666705",
      "postDate": "02/24/2024 15:07:27",
      "content": "<p><a href=\"https://www.kaggle.com/lohmaa\" target=\"_blank\">@lohmaa</a> The data in train are also not raw data. Test and train were both transformed before the competition. Please understand that we can't say exactly what we are going to do. We will make sure that the column names and dtypes stay the same. In that case, you don't have to worry about the underlying data as it will be the same change for everyone. </p>",
      "rawMarkdown": "lohmaa The data in train are also not raw data. Test and train were both transformed before the competition. Please understand that we can't say exactly what we are going to do. We will make sure that the column names and dtypes stay the same. In that case, you don't have to worry about the underlying data as it will be the same change for everyone.",
      "votes": null
    },
    {
      "id": "2666796",
      "postDate": "02/24/2024 16:31:29",
      "content": "<p>i think the question is, if the \"meaning\" of the date columns stays the same as for train.<br>\nif in train we would like to create feature D1 - D2, can we expect it to behave the same for test?</p>",
      "rawMarkdown": "i think the question is, if the \"meaning\" of the date columns stays the same as for train.\nif in train we would like to create feature D1 - D2, can we expect it to behave the same for test?",
      "votes": null
    },
    {
      "id": "2666818",
      "postDate": "02/24/2024 16:50:52",
      "content": "<p>Thank you <a href=\"https://www.kaggle.com/simonveitner\" target=\"_blank\">@simonveitner</a>, Exactly, It was my \"underlying\" question. We can expect sometimes nan but when there is a value, what can we expect ? I'm not only talking about the dtype. Will the transformation provide a kind of random string in any case, or only sometimes… In a real world, this information would be known by the project team who build a model : what is the expected behavior, even if we always can expect nan or wrong string sometimes. <a href=\"https://www.kaggle.com/jetakow\" target=\"_blank\">@jetakow</a> Thank you for your reactivity to our questions 😉</p>",
      "rawMarkdown": "Thank you @simonveitner, Exactly, It was my \"underlying\" question. We can expect sometimes nan but when there is a value, what can we expect ? I'm not only talking about the dtype. Will the transformation provide a kind of random string in any case, or only sometimes... In a real world, this information would be known by the project team who build a model : what is the expected behavior, even if we always can expect nan or wrong string sometimes. @jetakow Thank you for your reactivity to our questions 😉",
      "votes": null
    },
    {
      "id": "2666884",
      "postDate": "02/24/2024 17:55:20",
      "content": "<p><a href=\"https://www.kaggle.com/pourchot\" target=\"_blank\">@pourchot</a> <a href=\"https://www.kaggle.com/simonveitner\" target=\"_blank\">@simonveitner</a> we don't want to make any change that would make the test data much more different from what we have now, because we want to still interpret the results in our business. I can not give you more details</p>",
      "rawMarkdown": "pourchot @simonveitner we don't want to make any change that would make the test data much more different from what we have now, because we want to still interpret the results in our business. I can not give you more details",
      "votes": null
    },
    {
      "id": "2667483",
      "postDate": "02/25/2024 06:10:09",
      "content": "<p>Hey, I am not able to submit my notebook, any solutions?</p>",
      "rawMarkdown": "Hey, I am not able to submit my notebook, any solutions?",
      "votes": null
    },
    {
      "id": "2667493",
      "postDate": "02/25/2024 06:17:30",
      "content": "<p>Yes, the solution is to read the announcement ,)</p>",
      "rawMarkdown": "Yes, the solution is to read the announcement ,)",
      "votes": null
    },
    {
      "id": "2668141",
      "postDate": "02/25/2024 14:26:47",
      "content": "<p>😂</p>\n<blockquote>\n  <p>We will disable submissions for two weeks from now on and prolong the competition by two weeks.</p>\n</blockquote>",
      "rawMarkdown": "😂\n>We will disable submissions for two weeks from now on and prolong the competition by two weeks.",
      "votes": null
    },
    {
      "id": "2669621",
      "postDate": "02/26/2024 11:46:22",
      "content": "<p>Without WEEK_NUM we can't check the metric in training. Please leave WEEK_NUM in training dataset as now, and delete it ONLY in test set. In this way we can split training by time and eval properly the metric.</p>",
      "rawMarkdown": "Without WEEK_NUM we can't check the metric in training. Please leave WEEK_NUM in training dataset as now, and delete it ONLY in test set. In this way we can split training by time and eval properly the metric.",
      "votes": null
    },
    {
      "id": "2669661",
      "postDate": "02/26/2024 12:14:46",
      "content": "<p>WEEK_NUM will stay in train sample. Daniel was writing about test sample (also in announcement we said we plan to do adjustment in test dataset, nothing about train)</p>",
      "rawMarkdown": "WEEK_NUM will stay in train sample. Daniel was writing about test sample (also in announcement we said we plan to do adjustment in test dataset, nothing about train)",
      "votes": null
    },
    {
      "id": "2671196",
      "postDate": "02/27/2024 11:31:10",
      "content": "<p><a href=\"https://www.kaggle.com/simonveitner\" target=\"_blank\">@simonveitner</a> I would like to add additional information to my previous statement. The transformation will be done in a way that it won't change the meaning of columns and the features created from them should be performing about the same as now. </p>",
      "rawMarkdown": "simonveitner I would like to add additional information to my previous statement. The transformation will be done in a way that it won't change the meaning of columns and the features created from them should be performing about the same as now.",
      "votes": null
    },
    {
      "id": "2671411",
      "postDate": "02/27/2024 13:45:56",
      "content": "<p>Just wait for 1 week.</p>",
      "rawMarkdown": "Just wait for 1 week.",
      "votes": null
    },
    {
      "id": "2672644",
      "postDate": "02/28/2024 07:53:04",
      "content": "<p>I don’t quite understand, is it possible to generate new features based on dates? Or after changing the data on the test, this will not need to be done?<br>\nExample feature:<br>\ndatediff(day,datefirstoffer_1144D,date_decision)</p>",
      "rawMarkdown": "I don’t quite understand, is it possible to generate new features based on dates? Or after changing the data on the test, this will not need to be done?\nExample feature:\ndatediff(day,datefirstoffer_1144D,date_decision)",
      "votes": null
    },
    {
      "id": "2684982",
      "postDate": "03/06/2024 22:19:46",
      "content": "<p>Excuse me, is there an exact time when the submissions will be reopened? Will the public leaderboard be reset?</p>",
      "rawMarkdown": "Excuse me, is there an exact time when the submissions will be reopened? Will the public leaderboard be reset?",
      "votes": null
    },
    {
      "id": "2685529",
      "postDate": "03/07/2024 08:13:55",
      "content": "<p>Excuse me,how long will it take to start submitting the result?</p>",
      "rawMarkdown": "Excuse me,how long will it take to start submitting the result?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2662696,
      "author_name": "davutpolat",
      "author_url": "",
      "post_date": "02/22/2024 04:49:44",
      "content": "<p>If you remove some fields and transform the most of the fields in data, you should restart it with 3 months. It has already been 3 weeks wasted. There are too many files and most people waited for the metric fix for given dataset but now people probably will not be able to use most of their previous work. I think it would be more fair if you restart the timeline with 3 months.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2663048,
          "author_name": "tomasjeline2",
          "author_url": "",
          "post_date": "02/22/2024 09:35:35",
          "content": "<p>Hi Davut,<br>\nI agree with you. We will adjust deadline of competition accordingly to compensate your lost time. We have agreement with Kaggle team on this.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2662815,
      "author_name": "jacobyjaeger",
      "author_url": "",
      "post_date": "02/22/2024 06:20:45",
      "content": "<blockquote>\n  <p>There is very little guarantee that a new metric would not open doors to a new type of hacking. This part will take at least a week.</p>\n</blockquote>\n<p>I've explained what makes the metric hackable: it is the fact that the overall score is not monotonic wrt to the sample scores. This is easy to test for, one just has to check whether the gradient of the the final score wrt to the round scores has strictly one sign. Only when this rule is violated is it possible to make the overall score better by making round scores worse. </p>\n<p>It seems absurd to me not to fix this problem on the excuse that it might create new problems. The fact is that most competitions manage to find metrics that aren't in any way hackable. It seems extremely unlikely to me that any reasonably designed metric that is provably monotonic to round scores is somehow hackable in another way.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2663013,
          "author_name": "jetakow",
          "author_url": "",
          "post_date": "02/22/2024 09:25:05",
          "content": "<p>I appreciate the effort you put into the discussion and yes it would be ideal to have a metric that is monotonic in respect to every observation score and yet at the same time it can take into account the \"extrapolation\" of near-future performance as it is now with the gini stability metric. In case you can think of any such metric that would not be only monotonic wrt to the sample scores, but also preserves the idea of long-term decline of model. Please share it in <a href=\"https://www.kaggle.com/competitions/home-credit-credit-risk-model-stability/discussion/478699\" target=\"_blank\">https://www.kaggle.com/competitions/home-credit-credit-risk-model-stability/discussion/478699</a></p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2663009,
      "author_name": "darynarr",
      "author_url": "",
      "post_date": "02/22/2024 09:23:06",
      "content": "<p>But there are a lot of date columns in the data other than date_decision which might be close enough to point to the approximate date of decision. I have not checked it myself yet, but something tells it might be the case</p>",
      "votes": null,
      "replies": [
        {
          "id": 2663038,
          "author_name": "tomasjeline2",
          "author_url": "",
          "post_date": "02/22/2024 09:33:14",
          "content": "<p>We are aware of that. Data in test will be transformed.</p>",
          "votes": null,
          "replies": [
            {
              "id": 2663126,
              "author_name": "simonveitner",
              "author_url": "",
              "post_date": "02/22/2024 10:22:34",
              "content": "<p>But the data in train stays like it is?</p>",
              "votes": null,
              "replies": [
                {
                  "id": 2663290,
                  "author_name": "jetakow",
                  "author_url": "",
                  "post_date": "02/22/2024 12:05:38",
                  "content": "<p>Exactly, the train data stay the same.</p>",
                  "votes": null,
                  "replies": []
                }
              ]
            },
            {
              "id": 2665286,
              "author_name": "boristown",
              "author_url": "",
              "post_date": "02/23/2024 15:18:05",
              "content": "<p>请问测试数据中的日期将会如何转换？还是说日期字段会全部去掉？</p>",
              "votes": null,
              "replies": [
                {
                  "id": 2665386,
                  "author_name": "lohmaa",
                  "author_url": "",
                  "post_date": "02/23/2024 16:23:29",
                  "content": "<p><a href=\"https://www.kaggle.com/boristown\" target=\"_blank\">@boristown</a> โปรดลองใช้นักแปลภาษาอังกฤษเพื่อให้ผู้ใช้เข้าใจได้ง่ายขึ้น</p>",
                  "votes": null,
                  "replies": []
                }
              ]
            },
            {
              "id": 2666394,
              "author_name": "pourchot",
              "author_url": "",
              "post_date": "02/24/2024 11:16:07",
              "content": "<p>Hi <a href=\"https://www.kaggle.com/jetakow\" target=\"_blank\">@jetakow</a> &amp; <a href=\"https://www.kaggle.com/tomasjeline2\" target=\"_blank\">@tomasjeline2</a> <br>\nWhat do you mean by transformed ? Every column with 'D' at the end  which is not nan is a string that can be converted to date format. Do you mean your transformation will avoid the conversion to date format, for every column 'D' ? In that case it will be replaced by nan or a string <strong>whitout any meaning</strong> ? Thank you for your information</p>",
              "votes": null,
              "replies": [
                {
                  "id": 2666649,
                  "author_name": "jetakow",
                  "author_url": "",
                  "post_date": "02/24/2024 14:15:14",
                  "content": "<p>Hi <a href=\"https://www.kaggle.com/pourchot\" target=\"_blank\">@pourchot</a>, all the columns in test data will stay, the only columns that may be removed are only in test_base, but after seeing responses in Discussion we will most likely leave them, just replace the values with constant (date decision, month and week number). Dtypes will remain same everywhere. Transformation of test data will remain unspecified. </p>",
                  "votes": null,
                  "replies": [
                    {
                      "id": 2666684,
                      "author_name": "lohmaa",
                      "author_url": "",
                      "post_date": "02/24/2024 14:44:28",
                      "content": "<p>I'm trying to figure out how a feature can be transformed \"differently\" in the training set and the testing set and still remain useful. My understanding is that data has to be both structurally and \"semantically\" consistent across both training and testing sets to be useful. Nevertheless I'm sure it will be all clarified when it's ready.</p>",
                      "votes": null,
                      "replies": [
                        {
                          "id": 2666705,
                          "author_name": "jetakow",
                          "author_url": "",
                          "post_date": "02/24/2024 15:07:27",
                          "content": "<p><a href=\"https://www.kaggle.com/lohmaa\" target=\"_blank\">@lohmaa</a> The data in train are also not raw data. Test and train were both transformed before the competition. Please understand that we can't say exactly what we are going to do. We will make sure that the column names and dtypes stay the same. In that case, you don't have to worry about the underlying data as it will be the same change for everyone. </p>",
                          "votes": null,
                          "replies": [
                            {
                              "id": 2666796,
                              "author_name": "simonveitner",
                              "author_url": "",
                              "post_date": "02/24/2024 16:31:29",
                              "content": "<p>i think the question is, if the \"meaning\" of the date columns stays the same as for train.<br>\nif in train we would like to create feature D1 - D2, can we expect it to behave the same for test?</p>",
                              "votes": null,
                              "replies": [
                                {
                                  "id": 2666818,
                                  "author_name": "pourchot",
                                  "author_url": "",
                                  "post_date": "02/24/2024 16:50:52",
                                  "content": "<p>Thank you <a href=\"https://www.kaggle.com/simonveitner\" target=\"_blank\">@simonveitner</a>, Exactly, It was my \"underlying\" question. We can expect sometimes nan but when there is a value, what can we expect ? I'm not only talking about the dtype. Will the transformation provide a kind of random string in any case, or only sometimes… In a real world, this information would be known by the project team who build a model : what is the expected behavior, even if we always can expect nan or wrong string sometimes. <a href=\"https://www.kaggle.com/jetakow\" target=\"_blank\">@jetakow</a> Thank you for your reactivity to our questions 😉</p>",
                                  "votes": null,
                                  "replies": [
                                    {
                                      "id": 2666884,
                                      "author_name": "jetakow",
                                      "author_url": "",
                                      "post_date": "02/24/2024 17:55:20",
                                      "content": "<p><a href=\"https://www.kaggle.com/pourchot\" target=\"_blank\">@pourchot</a> <a href=\"https://www.kaggle.com/simonveitner\" target=\"_blank\">@simonveitner</a> we don't want to make any change that would make the test data much more different from what we have now, because we want to still interpret the results in our business. I can not give you more details</p>",
                                      "votes": null,
                                      "replies": [
                                        {
                                          "id": 2671196,
                                          "author_name": "jetakow",
                                          "author_url": "",
                                          "post_date": "02/27/2024 11:31:10",
                                          "content": "<p><a href=\"https://www.kaggle.com/simonveitner\" target=\"_blank\">@simonveitner</a> I would like to add additional information to my previous statement. The transformation will be done in a way that it won't change the meaning of columns and the features created from them should be performing about the same as now. </p>",
                                          "votes": null,
                                          "replies": []
                                        }
                                      ]
                                    }
                                  ]
                                }
                              ]
                            }
                          ]
                        }
                      ]
                    }
                  ]
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 2664954,
      "author_name": "evan918",
      "author_url": "",
      "post_date": "02/23/2024 10:06:18",
      "content": "<p>If WEEK_NUM is removed, we won't be able to perform certain feature engineering tasks based on WEEK_NUM, such as quantile calculations.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2665546,
          "author_name": "blindape",
          "author_url": "",
          "post_date": "02/23/2024 18:00:01",
          "content": "<p>I agree WEEK_NUM is crucial, at least for my models. Consider hashing WEEK_NUM instead of delete it, please. Removing the temporal component of WEEK_NUM seems enough for prevent time hacking.</p>",
          "votes": null,
          "replies": [
            {
              "id": 2665835,
              "author_name": "jetakow",
              "author_url": "",
              "post_date": "02/23/2024 21:18:03",
              "content": "<p>We can make it constant and let the column there. However, the values of WEEK_NUM will be removed.</p>",
              "votes": null,
              "replies": [
                {
                  "id": 2669621,
                  "author_name": "blindape",
                  "author_url": "",
                  "post_date": "02/26/2024 11:46:22",
                  "content": "<p>Without WEEK_NUM we can't check the metric in training. Please leave WEEK_NUM in training dataset as now, and delete it ONLY in test set. In this way we can split training by time and eval properly the metric.</p>",
                  "votes": null,
                  "replies": [
                    {
                      "id": 2669661,
                      "author_name": "tomasjeline2",
                      "author_url": "",
                      "post_date": "02/26/2024 12:14:46",
                      "content": "<p>WEEK_NUM will stay in train sample. Daniel was writing about test sample (also in announcement we said we plan to do adjustment in test dataset, nothing about train)</p>",
                      "votes": null,
                      "replies": []
                    }
                  ]
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 2667483,
      "author_name": "shivanshkumar752",
      "author_url": "",
      "post_date": "02/25/2024 06:10:09",
      "content": "<p>Hey, I am not able to submit my notebook, any solutions?</p>",
      "votes": null,
      "replies": [
        {
          "id": 2667493,
          "author_name": "lohmaa",
          "author_url": "",
          "post_date": "02/25/2024 06:17:30",
          "content": "<p>Yes, the solution is to read the announcement ,)</p>",
          "votes": null,
          "replies": [
            {
              "id": 2668141,
              "author_name": "waechter",
              "author_url": "",
              "post_date": "02/25/2024 14:26:47",
              "content": "<p>😂</p>\n<blockquote>\n  <p>We will disable submissions for two weeks from now on and prolong the competition by two weeks.</p>\n</blockquote>",
              "votes": null,
              "replies": []
            }
          ]
        },
        {
          "id": 2671411,
          "author_name": "boristown",
          "author_url": "",
          "post_date": "02/27/2024 13:45:56",
          "content": "<p>Just wait for 1 week.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2672644,
      "author_name": "dima1992",
      "author_url": "",
      "post_date": "02/28/2024 07:53:04",
      "content": "<p>I don’t quite understand, is it possible to generate new features based on dates? Or after changing the data on the test, this will not need to be done?<br>\nExample feature:<br>\ndatediff(day,datefirstoffer_1144D,date_decision)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2684982,
      "author_name": "co000l",
      "author_url": "",
      "post_date": "03/06/2024 22:19:46",
      "content": "<p>Excuse me, is there an exact time when the submissions will be reopened? Will the public leaderboard be reset?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2685529,
      "author_name": "dxh2022",
      "author_url": "",
      "post_date": "03/07/2024 08:13:55",
      "content": "<p>Excuse me,how long will it take to start submitting the result?</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2662418": "Dear Kagglers, \n \nWe prepared this year’s competition with an intention to present you something else and more challenging than the previous Home Credit kaggle competition. We decided to select stability of models in production as the main objective of this competition. Therefore, we will keep stability and won’t remove it from the competition.  \n\nWe would like to further clarify the meaning of stability. Our ideal stable model has constant performance by weeks in time and as high of a mean gini over weeks as possible. That is the reason why we have three terms in the stability metric as of now since this is desirable in business and not only to maximize AUC. The metric itself is a good approximation of the decisions made in-house when assessing the performance of models in production. Despite different opinions here within the competition, removing the terms in the stability metric would significantly deviate from the way we rate performance internally. When people leverage the second term and nullify it by carefully choosing scores on the test sample, it does not represent how we deploy models in the production. This is undesirable and we are forced to make steps that will discourage hacking of the metric.  \n \nWe will make a change to the test dataset in such a way that it will represent better client scoring in production. When the model is running in production, it can only see its past performance and predictions. This would be quite difficult to reproduce in a Kaggle competition, but we have an alternative solution. We will remove the date decision, month and week number from the test base table. This part will take two weeks to amend, and we will prolong the competition by this period. Because of this, it is very likely that we won’t be able to rescore the old submissions. We apologize in advance for this inconvenience. \n \nBesides the change of test data, we have a second plan regarding metric. At this moment we are still considering metric change, and we welcome suggestions that would be aligned with how we view the value of model stability and in a way that would prevent further hacking. We are unlikely to go with a new metric that would prevent this type of hacking and would significantly deviate from the stability part. There is very little guarantee that a new metric would not open doors to a new type of hacking. This part will take at least a week.\n \nWe would like to apologize for the delay and inconvenience caused by the issues connected to the stability metric. We also want to thank the community for many proposals on how to fix the issue. We have tested and analyzed many of them, thank you for all the suggestions! We will disable submissions for two weeks from now on and prolong the competition by two weeks.  \n\nTomas & Daniel",
    "2662696": "If you remove some fields and transform the most of the fields in data, you should restart it with 3 months. It has already been 3 weeks wasted. There are too many files and most people waited for the metric fix for given dataset but now people probably will not be able to use most of their previous work. I think it would be more fair if you restart the timeline with 3 months.",
    "2662815": ">There is very little guarantee that a new metric would not open doors to a new type of hacking. This part will take at least a week.\n\nI've explained what makes the metric hackable: it is the fact that the overall score is not monotonic wrt to the sample scores. This is easy to test for, one just has to check whether the gradient of the the final score wrt to the round scores has strictly one sign. Only when this rule is violated is it possible to make the overall score better by making round scores worse. \n\nIt seems absurd to me not to fix this problem on the excuse that it might create new problems. The fact is that most competitions manage to find metrics that aren't in any way hackable. It seems extremely unlikely to me that any reasonably designed metric that is provably monotonic to round scores is somehow hackable in another way.",
    "2663009": "But there are a lot of date columns in the data other than date_decision which might be close enough to point to the approximate date of decision. I have not checked it myself yet, but something tells it might be the case",
    "2663013": "I appreciate the effort you put into the discussion and yes it would be ideal to have a metric that is monotonic in respect to every observation score and yet at the same time it can take into account the \"extrapolation\" of near-future performance as it is now with the gini stability metric. In case you can think of any such metric that would not be only monotonic wrt to the sample scores, but also preserves the idea of long-term decline of model. Please share it in https://www.kaggle.com/competitions/home-credit-credit-risk-model-stability/discussion/478699",
    "2663038": "We are aware of that. Data in test will be transformed.",
    "2663048": "Hi Davut,\nI agree with you. We will adjust deadline of competition accordingly to compensate your lost time. We have agreement with Kaggle team on this.",
    "2663126": "But the data in train stays like it is?",
    "2663290": "Exactly, the train data stay the same.",
    "2664954": "If WEEK_NUM is removed, we won't be able to perform certain feature engineering tasks based on WEEK_NUM, such as quantile calculations.",
    "2665286": "请问测试数据中的日期将会如何转换？还是说日期字段会全部去掉？",
    "2665386": "boristown โปรดลองใช้นักแปลภาษาอังกฤษเพื่อให้ผู้ใช้เข้าใจได้ง่ายขึ้น",
    "2665546": "I agree WEEK_NUM is crucial, at least for my models. Consider hashing WEEK_NUM instead of delete it, please. Removing the temporal component of WEEK_NUM seems enough for prevent time hacking.",
    "2665835": "We can make it constant and let the column there. However, the values of WEEK_NUM will be removed.",
    "2666394": "Hi @jetakow & @tomasjeline2 \nWhat do you mean by transformed ? Every column with 'D' at the end  which is not nan is a string that can be converted to date format. Do you mean your transformation will avoid the conversion to date format, for every column 'D' ? In that case it will be replaced by nan or a string **whitout any meaning** ? Thank you for your information",
    "2666649": "Hi @pourchot, all the columns in test data will stay, the only columns that may be removed are only in test_base, but after seeing responses in Discussion we will most likely leave them, just replace the values with constant (date decision, month and week number). Dtypes will remain same everywhere. Transformation of test data will remain unspecified.",
    "2666684": "I'm trying to figure out how a feature can be transformed \"differently\" in the training set and the testing set and still remain useful. My understanding is that data has to be both structurally and \"semantically\" consistent across both training and testing sets to be useful. Nevertheless I'm sure it will be all clarified when it's ready.",
    "2666705": "lohmaa The data in train are also not raw data. Test and train were both transformed before the competition. Please understand that we can't say exactly what we are going to do. We will make sure that the column names and dtypes stay the same. In that case, you don't have to worry about the underlying data as it will be the same change for everyone.",
    "2666796": "i think the question is, if the \"meaning\" of the date columns stays the same as for train.\nif in train we would like to create feature D1 - D2, can we expect it to behave the same for test?",
    "2666818": "Thank you @simonveitner, Exactly, It was my \"underlying\" question. We can expect sometimes nan but when there is a value, what can we expect ? I'm not only talking about the dtype. Will the transformation provide a kind of random string in any case, or only sometimes... In a real world, this information would be known by the project team who build a model : what is the expected behavior, even if we always can expect nan or wrong string sometimes. @jetakow Thank you for your reactivity to our questions 😉",
    "2666884": "pourchot @simonveitner we don't want to make any change that would make the test data much more different from what we have now, because we want to still interpret the results in our business. I can not give you more details",
    "2667483": "Hey, I am not able to submit my notebook, any solutions?",
    "2667493": "Yes, the solution is to read the announcement ,)",
    "2668141": "😂\n>We will disable submissions for two weeks from now on and prolong the competition by two weeks.",
    "2669621": "Without WEEK_NUM we can't check the metric in training. Please leave WEEK_NUM in training dataset as now, and delete it ONLY in test set. In this way we can split training by time and eval properly the metric.",
    "2669661": "WEEK_NUM will stay in train sample. Daniel was writing about test sample (also in announcement we said we plan to do adjustment in test dataset, nothing about train)",
    "2671196": "simonveitner I would like to add additional information to my previous statement. The transformation will be done in a way that it won't change the meaning of columns and the features created from them should be performing about the same as now.",
    "2671411": "Just wait for 1 week.",
    "2672644": "I don’t quite understand, is it possible to generate new features based on dates? Or after changing the data on the test, this will not need to be done?\nExample feature:\ndatediff(day,datefirstoffer_1144D,date_decision)",
    "2684982": "Excuse me, is there an exact time when the submissions will be reopened? Will the public leaderboard be reset?",
    "2685529": "Excuse me,how long will it take to start submitting the result?"
  },
  "source": "meta"
}