{
  "id": 483262,
  "title": "Submissions are now enabled",
  "url": "/competitions/home-credit-credit-risk-model-stability/discussion/483262",
  "author_name": "",
  "post_date": "2024-03-11T17:38:08.060162900Z",
  "votes": 18,
  "comment_count": 28,
  "views": 0,
  "content": "<p>Please report any problems here.</p>\n<p>Good luck!</p>",
  "messages": [
    {
      "id": "2692133",
      "postDate": "03/11/2024 17:38:08",
      "content": "<p>Please report any problems here.</p>\n<p>Good luck!</p>",
      "rawMarkdown": "Please report any problems here.\n\nGood luck!",
      "votes": null
    },
    {
      "id": "2692594",
      "postDate": "03/11/2024 23:54:59",
      "content": "<p>What <em>exactly</em> changed in the train &amp; test data?</p>\n<p>I picked some random submissions and they work and I got a score.</p>\n<p>When I try to run (live notebook) those same notebooks (that somehow worked to submit) they crash, e.g. crashing with \"numberofoverdueinstlmaxdat_641D\".</p>\n<p>Update: if the local test is a random sample of the real test, and the test changed (particularly, its order), so the small test changed as well, and its so insanely small (like, WAY too small) that it adds all sort of weird patterns, e.g. a feature might change data-types from local-test to real-test. I had fixed those, but they need fix again.</p>",
      "rawMarkdown": "What *exactly* changed in the train & test data?\n\nI picked some random submissions and they work and I got a score.\n\nWhen I try to run (live notebook) those same notebooks (that somehow worked to submit) they crash, e.g. crashing with \"numberofoverdueinstlmaxdat_641D\".\n\nUpdate: if the local test is a random sample of the real test, and the test changed (particularly, its order), so the small test changed as well, and its so insanely small (like, WAY too small) that it adds all sort of weird patterns, e.g. a feature might change data-types from local-test to real-test. I had fixed those, but they need fix again.",
      "votes": null
    },
    {
      "id": "2692633",
      "postDate": "03/12/2024 01:13:09",
      "content": "<p>May I ask if it is possible to exclude errors from the number of submissions? This competition has restarted, and I have made two more errors when submitting the original code.</p>",
      "rawMarkdown": "May I ask if it is possible to exclude errors from the number of submissions? This competition has restarted, and I have made two more errors when submitting the original code.",
      "votes": null
    },
    {
      "id": "2693163",
      "postDate": "03/12/2024 09:05:00",
      "content": "<p>I checked and realized that all my previous submissions failed, now the status is 'Unranked'. Why? Is it technical error? Or is it because metric was changed and applied to the previous submissions? Or only new submissions will be accepted with new metric from now and onward? </p>\n<p>That's a bit scary as many effort had been made and now it seems that it is worthless… </p>\n<p>Can someone, please, explain if I missed something about it?</p>",
      "rawMarkdown": "I checked and realized that all my previous submissions failed, now the status is 'Unranked'. Why? Is it technical error? Or is it because metric was changed and applied to the previous submissions? Or only new submissions will be accepted with new metric from now and onward? \n\nThat's a bit scary as many effort had been made and now it seems that it is worthless... \n\nCan someone, please, explain if I missed something about it?",
      "votes": null
    },
    {
      "id": "2693179",
      "postDate": "03/12/2024 09:24:12",
      "content": "<p>You need to resubmit to have a ranking.</p>",
      "rawMarkdown": "You need to resubmit to have a ranking.",
      "votes": null
    },
    {
      "id": "2693186",
      "postDate": "03/12/2024 09:30:57",
      "content": "<p>Noted, thanks but do you know how resubmissions will affect ranking considering new metric? Will they be adjusted relative to the new metric or ranked as they are?</p>",
      "rawMarkdown": "Noted, thanks but do you know how resubmissions will affect ranking considering new metric? Will they be adjusted relative to the new metric or ranked as they are?",
      "votes": null
    },
    {
      "id": "2693298",
      "postDate": "03/12/2024 10:29:22",
      "content": "<blockquote>\n  <p>What exactly changed in the train &amp; test data?</p>\n</blockquote>\n<p>I am afraid we can't say that for the sake of making the competition more fair for everyone. </p>\n<blockquote>\n  <p>I picked some random submissions and they work and I got a score.</p>\n</blockquote>\n<p>That is expected, yet we encourage you to generate a new submissions.</p>\n<blockquote>\n  <p>Update: if the local test is a random sample of the real test, and the test changed (particularly, its order), so the small test changed as well, and its so insanely small (like, WAY too small) that it adds all sort of weird patterns, e.g. a feature might change data-types from local-test to real-test. I had fixed those, but they need fix again.</p>\n</blockquote>\n<p>You can always use .csv tables to infer the schema first in the test sample, then load the test parquet tables and change the dtypes in schema and then concatenate them. Another approach would be to use the dtypes from train data and use them in schema for test tables. We understand that this poses a minor inconvenience for kagglers in this competition. I believe the data were stored in a way they don't consume more space than needed on a disk. Best of luck.</p>",
      "rawMarkdown": "> What exactly changed in the train & test data?\n\nI am afraid we can't say that for the sake of making the competition more fair for everyone. \n\n> I picked some random submissions and they work and I got a score.\n\nThat is expected, yet we encourage you to generate a new submissions.\n\n> Update: if the local test is a random sample of the real test, and the test changed (particularly, its order), so the small test changed as well, and its so insanely small (like, WAY too small) that it adds all sort of weird patterns, e.g. a feature might change data-types from local-test to real-test. I had fixed those, but they need fix again.\n\nYou can always use .csv tables to infer the schema first in the test sample, then load the test parquet tables and change the dtypes in schema and then concatenate them. Another approach would be to use the dtypes from train data and use them in schema for test tables. We understand that this poses a minor inconvenience for kagglers in this competition. I believe the data were stored in a way they don't consume more space than needed on a disk. Best of luck.",
      "votes": null
    },
    {
      "id": "2693467",
      "postDate": "03/12/2024 12:55:55",
      "content": "<blockquote>\n  <p>You can always use .csv tables to infer the schema first in the test sample, then load the test parquet tables and change the dtypes in schema and then concatenate them.</p>\n</blockquote>\n<p>Polars can read the schema of a parquet without reading the data, <a href=\"https://docs.pola.rs/py-polars/html/reference/api/polars.read_parquet_schema.html\" target=\"_blank\">polars.read_parquet_schema</a></p>\n<blockquote>\n  <p>We understand that this poses a minor inconvenience for kagglers in this competition.</p>\n</blockquote>\n<p>I'm dealing with this problem right now and I personally enjoy this part of the competitions 😀.</p>",
      "rawMarkdown": ">You can always use .csv tables to infer the schema first in the test sample, then load the test parquet tables and change the dtypes in schema and then concatenate them.\n\nPolars can read the schema of a parquet without reading the data, [polars.read_parquet_schema](https://docs.pola.rs/py-polars/html/reference/api/polars.read_parquet_schema.html)\n\n>We understand that this poses a minor inconvenience for kagglers in this competition.\n\nI'm dealing with this problem right now and I personally enjoy this part of the competitions 😀.",
      "votes": null
    },
    {
      "id": "2693525",
      "postDate": "03/12/2024 13:42:45",
      "content": "<blockquote>\n  <p>Polars can read the schema of a parquet without reading the data, polars.read_parquet_schema</p>\n</blockquote>\n<p>Awesome!</p>\n<blockquote>\n  <p>I'm dealing with this problem right now and I personally enjoy this part of the competitions 😀.</p>\n</blockquote>\n<p>We already removed a huge portion of data preparation. I am glad to see someone with positive feedback 😀 and I am glad you enjoy that. </p>",
      "rawMarkdown": "> Polars can read the schema of a parquet without reading the data, polars.read_parquet_schema\n\nAwesome!\n\n> I'm dealing with this problem right now and I personally enjoy this part of the competitions 😀.\n\nWe already removed a huge portion of data preparation. I am glad to see someone with positive feedback 😀 and I am glad you enjoy that.",
      "votes": null
    },
    {
      "id": "2694360",
      "postDate": "03/13/2024 03:15:42",
      "content": "<p>the new metric</p>",
      "rawMarkdown": "the new metric",
      "votes": null
    },
    {
      "id": "2696900",
      "postDate": "03/14/2024 15:59:24",
      "content": "<p>Did the train set change in <strong><em>any</em></strong> way at all? I am checking as if so, we need to re-download if experiments are done in local machines.</p>",
      "rawMarkdown": "Did the train set change in ***any*** way at all? I am checking as if so, we need to re-download if experiments are done in local machines.",
      "votes": null
    },
    {
      "id": "2697399",
      "postDate": "03/14/2024 21:55:49",
      "content": "<p>No change in the training data.</p>",
      "rawMarkdown": "No change in the training data.",
      "votes": null
    },
    {
      "id": "2698022",
      "postDate": "03/15/2024 08:31:15",
      "content": "<p>Hi, I'm trying to make a dummy submission to check if the process works, but keep getting this error:<br>\n<em>Submission Scoring Error\nYour notebook generated a submission file with incorrect format. Some examples causing this are: wrong number of rows or columns, empty values, an incorrect data type for a value, or invalid submission values from what is expected. See more debugging tips</em></p>\n<p>While I have the following output file: <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4157236%2F591a9d809472bba0760837939a9436d2%2FScreenshot%202024-03-15%20092918.png?generation=1710491446315645&amp;alt=media\"></p>\n<p>This does contain all the criteria that is stated on the overview pages. What am I doing wrong here?</p>",
      "rawMarkdown": "Hi, I'm trying to make a dummy submission to check if the process works, but keep getting this error:\n*Submission Scoring Error\nYour notebook generated a submission file with incorrect format. Some examples causing this are: wrong number of rows or columns, empty values, an incorrect data type for a value, or invalid submission values from what is expected. See more debugging tips*\n\nWhile I have the following output file: ![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4157236%2F591a9d809472bba0760837939a9436d2%2FScreenshot%202024-03-15%20092918.png?generation=1710491446315645&alt=media)\n\nThis does contain all the criteria that is stated on the overview pages. What am I doing wrong here?",
      "votes": null
    },
    {
      "id": "2698152",
      "postDate": "03/15/2024 10:26:38",
      "content": "<p>Are you copying the sample submission or using case ids from the test_base table?</p>",
      "rawMarkdown": "Are you copying the sample submission or using case ids from the test_base table?",
      "votes": null
    },
    {
      "id": "2698191",
      "postDate": "03/15/2024 10:43:55",
      "content": "<p>Here I'm using the sample submission, but also got the same error when using the test_base table and using a dummy model. See this post: <a href=\"https://www.kaggle.com/competitions/home-credit-credit-risk-model-stability/discussion/483635\" target=\"_blank\">https://www.kaggle.com/competitions/home-credit-credit-risk-model-stability/discussion/483635</a></p>",
      "rawMarkdown": "Here I'm using the sample submission, but also got the same error when using the test_base table and using a dummy model. See this post: https://www.kaggle.com/competitions/home-credit-credit-risk-model-stability/discussion/483635",
      "votes": null
    },
    {
      "id": "2703320",
      "postDate": "03/18/2024 05:26:54",
      "content": "<p>From the post <a href=\"https://www.kaggle.com/competitions/home-credit-credit-risk-model-stability/discussion/482474\" target=\"_blank\">here</a>.</p>\n<blockquote>\n  <p>Column date_decision is no longer the same</p>\n</blockquote>\n<p>\"no longer the same\" but the fact that train has not changed, makes the situation funny. I am fine doing any transformation, but whats the point if the train is not using those?</p>",
      "rawMarkdown": "From the post [here](https://www.kaggle.com/competitions/home-credit-credit-risk-model-stability/discussion/482474).\n\n>Column date_decision is no longer the same\n\n\"no longer the same\" but the fact that train has not changed, makes the situation funny. I am fine doing any transformation, but whats the point if the train is not using those?",
      "votes": null
    },
    {
      "id": "2703610",
      "postDate": "03/18/2024 09:37:09",
      "content": "<p>There was a request made several times that we should preserve the meaning of features such as column_with_date - date_decision, that is mostly preserved and that is also a reason why we left it in test base table. </p>",
      "rawMarkdown": "There was a request made several times that we should preserve the meaning of features such as column_with_date - date_decision, that is mostly preserved and that is also a reason why we left it in test base table.",
      "votes": null
    },
    {
      "id": "2703924",
      "postDate": "03/18/2024 13:30:43",
      "content": "<p>According to rule 7C, external data might be Ok (granted some other aspects).</p>\n<p>I assume this no longer applies? There is no way to know if I am leaking into the future without knowing the date I should constraint myself.</p>\n<p>What was the transformation? Some random date such as the delta of dates is preserved but nothing else?</p>",
      "rawMarkdown": "According to rule 7C, external data might be Ok (granted some other aspects).\n\nI assume this no longer applies? There is no way to know if I am leaking into the future without knowing the date I should constraint myself.\n\nWhat was the transformation? Some random date such as the delta of dates is preserved but nothing else?",
      "votes": null
    },
    {
      "id": "2705434",
      "postDate": "03/19/2024 10:39:31",
      "content": "<blockquote>\n  <p>According to rule 7C, external data might be Ok</p>\n</blockquote>\n<p>It's all about the licence, if we can not use it internally because of a no-commerce clause, then it is out of the scope.  </p>\n<blockquote>\n  <p>I assume this no longer applies? </p>\n</blockquote>\n<p>It still does.</p>\n<blockquote>\n  <p>What was the transformation?</p>\n</blockquote>\n<p>Please understand, that we can not reveal it for the sake of making competition more fair for everyone.</p>",
      "rawMarkdown": "> According to rule 7C, external data might be Ok\n\nIt's all about the licence, if we can not use it internally because of a no-commerce clause, then it is out of the scope.  \n\n> I assume this no longer applies? \n\nIt still does.\n\n> What was the transformation?\n\nPlease understand, that we can not reveal it for the sake of making competition more fair for everyone.",
      "votes": null
    },
    {
      "id": "2705672",
      "postDate": "03/19/2024 13:06:36",
      "content": "<p>Still very confusing. Let me be explicit.</p>\n<p>1) I grab some free data source, available to everyone, blah blah, all good for the competition.</p>\n<p>2) That data is dependent of time, let's say, SP500 performance at the time of application (actually, at the day before application is better)</p>\n<p>3) I can perfectly use in train. What column do you suggest to use in test to avoid leakage? </p>\n<p>Thats my point, without knowing the transformation its unclear how to build something useful.</p>\n<p>Now if this means no external data is allowed anymore, then thats it.</p>",
      "rawMarkdown": "Still very confusing. Let me be explicit.\n\n1) I grab some free data source, available to everyone, blah blah, all good for the competition.\n\n2) That data is dependent of time, let's say, SP500 performance at the time of application (actually, at the day before application is better)\n\n3) I can perfectly use in train. What column do you suggest to use in test to avoid leakage? \n\nThats my point, without knowing the transformation its unclear how to build something useful.\n\nNow if this means no external data is allowed anymore, then thats it.",
      "votes": null
    },
    {
      "id": "2705889",
      "postDate": "03/19/2024 15:09:27",
      "content": "<p>There is no restriction on using S&amp;P500 or any other time series data that you think might help you improve the prediction. Before the transformation of test data we were ready for kagglers to freely use such data sources and improve their models based on that. We somewhat doubt that S&amp;P500 (and similar) will significantly improve the stability score (before the transformation). What could have a real impact on the stability score is covid period and we didn't want to forbid it, because there is no way to fully control the use of covid related time series (or any). </p>\n<p>After the transformation of data, I completely understand your concern and the motivation behind your post. Please consider following two scenarios 1) I don't reveal the transformation and you won't be able to use fully the, e.g. SP500 time series 2) I reveal the transformation and possibly making the whole \"fix\" and all the effort around it useless. We aimed for making the competition more fair for everyone after it was revealed that the gain from metric hack is significant. Our fix should neutralize almost 100% of the discovered hack. Revealing you the transformation might jeopardize all the effort we put into the fix and I can not allow that. Transformation of test data will not be revealed. </p>",
      "rawMarkdown": "There is no restriction on using S&P500 or any other time series data that you think might help you improve the prediction. Before the transformation of test data we were ready for kagglers to freely use such data sources and improve their models based on that. We somewhat doubt that S&P500 (and similar) will significantly improve the stability score (before the transformation). What could have a real impact on the stability score is covid period and we didn't want to forbid it, because there is no way to fully control the use of covid related time series (or any). \n\nAfter the transformation of data, I completely understand your concern and the motivation behind your post. Please consider following two scenarios 1) I don't reveal the transformation and you won't be able to use fully the, e.g. SP500 time series 2) I reveal the transformation and possibly making the whole \"fix\" and all the effort around it useless. We aimed for making the competition more fair for everyone after it was revealed that the gain from metric hack is significant. Our fix should neutralize almost 100% of the discovered hack. Revealing you the transformation might jeopardize all the effort we put into the fix and I can not allow that. Transformation of test data will not be revealed.",
      "votes": null
    },
    {
      "id": "2712183",
      "postDate": "03/23/2024 11:05:56",
      "content": "<p><a href=\"https://www.kaggle.com/jetakow\" target=\"_blank\">@jetakow</a> is also aggregation over a single date decision of other column preserved or it does not work anymore?</p>",
      "rawMarkdown": "jetakow is also aggregation over a single date decision of other column preserved or it does not work anymore?",
      "votes": null
    },
    {
      "id": "2712400",
      "postDate": "03/23/2024 14:38:22",
      "content": "<p><a href=\"https://www.kaggle.com/stenford23\" target=\"_blank\">@stenford23</a> that would also not be preserved. What we tried to preserve are date differences such as date_diff_feature = date_decision - aggregation_fun(date_col).</p>",
      "rawMarkdown": "stenford23 that would also not be preserved. What we tried to preserve are date differences such as date_diff_feature = date_decision - aggregation_fun(date_col).",
      "votes": null
    },
    {
      "id": "2712459",
      "postDate": "03/23/2024 15:11:38",
      "content": "<p>tyvm <a href=\"https://www.kaggle.com/jetakow\" target=\"_blank\">@jetakow</a> .<br>\nIs also date_col1 - date_col2 preserved? (difference not with date decision but between other date on same case id)</p>",
      "rawMarkdown": "tyvm @jetakow .\nIs also date_col1 - date_col2 preserved? (difference not with date decision but between other date on same case id)",
      "votes": null
    },
    {
      "id": "2712953",
      "postDate": "03/23/2024 21:17:11",
      "content": "<p>The answer is again, in most cases yes. And I can't comment on the actual difference, but about the usability of such a feature.</p>",
      "rawMarkdown": "The answer is again, in most cases yes. And I can't comment on the actual difference, but about the usability of such a feature.",
      "votes": null
    },
    {
      "id": "2715501",
      "postDate": "03/25/2024 14:31:12",
      "content": "<p>Hi guys,<br>\nI'm still not able to subbmit my results. I'm doing this project on my own laptop and then I only want to submit the predictions, is this a wrong approach?</p>\n<p>See my previous issue:<br>\n<a href=\"https://www.kaggle.com/competitions/home-credit-credit-risk-model-stability/discussion/483262#2698022\" target=\"_blank\">https://www.kaggle.com/competitions/home-credit-credit-risk-model-stability/discussion/483262#2698022</a></p>\n<p>I could really use some support in this.<br>\nFurhtermore this is the repo of my code: <a href=\"https://github.com/BartvanWoesik/home_credit\" target=\"_blank\">https://github.com/BartvanWoesik/home_credit</a></p>",
      "rawMarkdown": "Hi guys,\nI'm still not able to subbmit my results. I'm doing this project on my own laptop and then I only want to submit the predictions, is this a wrong approach?\n\nSee my previous issue:\nhttps://www.kaggle.com/competitions/home-credit-credit-risk-model-stability/discussion/483262#2698022\n\nI could really use some support in this.\nFurhtermore this is the repo of my code: https://github.com/BartvanWoesik/home_credit",
      "votes": null
    },
    {
      "id": "2715563",
      "postDate": "03/25/2024 15:07:53",
      "content": "<p>I've taken a look and noticed that in train.py predictions are saved in \"predictions.csv\", while the expected filename is \"submission.csv\". May that be the reason? Also, you can't submit just the predictions, this is a code competition.</p>",
      "rawMarkdown": "I've taken a look and noticed that in train.py predictions are saved in \"predictions.csv\", while the expected filename is \"submission.csv\". May that be the reason? Also, you can't submit just the predictions, this is a code competition.",
      "votes": null
    },
    {
      "id": "2834162",
      "postDate": "05/24/2024 15:29:11",
      "content": "<p>Hi, <a href=\"https://www.kaggle.com/inversion\" target=\"_blank\">@inversion</a>, for some reason previously successfully submitted code can now fail with OOM errors if you submit them again, I have created a discussion about it, a few ppl said they encounter the same. <a href=\"https://www.kaggle.com/competitions/home-credit-credit-risk-model-stability/discussion/506542\" target=\"_blank\">https://www.kaggle.com/competitions/home-credit-credit-risk-model-stability/discussion/506542</a></p>\n<p>What can be the root of such behaviour? </p>",
      "rawMarkdown": "Hi, @inversion, for some reason previously successfully submitted code can now fail with OOM errors if you submit them again, I have created a discussion about it, a few ppl said they encounter the same. https://www.kaggle.com/competitions/home-credit-credit-risk-model-stability/discussion/506542\n\nWhat can be the root of such behaviour?",
      "votes": null
    },
    {
      "id": "2837364",
      "postDate": "05/26/2024 12:51:38",
      "content": "<p>Hi, this is my first competition ever and i had just wonder the dataset and the puzzle so hard thanks to all of you who provide public notebook that's why i can understand and create some feature engineering.🙂</p>",
      "rawMarkdown": "Hi, this is my first competition ever and i had just wonder the dataset and the puzzle so hard thanks to all of you who provide public notebook that's why i can understand and create some feature engineering.🙂",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2692594,
      "author_name": "carloshuertas",
      "author_url": "",
      "post_date": "03/11/2024 23:54:59",
      "content": "<p>What <em>exactly</em> changed in the train &amp; test data?</p>\n<p>I picked some random submissions and they work and I got a score.</p>\n<p>When I try to run (live notebook) those same notebooks (that somehow worked to submit) they crash, e.g. crashing with \"numberofoverdueinstlmaxdat_641D\".</p>\n<p>Update: if the local test is a random sample of the real test, and the test changed (particularly, its order), so the small test changed as well, and its so insanely small (like, WAY too small) that it adds all sort of weird patterns, e.g. a feature might change data-types from local-test to real-test. I had fixed those, but they need fix again.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2693298,
          "author_name": "jetakow",
          "author_url": "",
          "post_date": "03/12/2024 10:29:22",
          "content": "<blockquote>\n  <p>What exactly changed in the train &amp; test data?</p>\n</blockquote>\n<p>I am afraid we can't say that for the sake of making the competition more fair for everyone. </p>\n<blockquote>\n  <p>I picked some random submissions and they work and I got a score.</p>\n</blockquote>\n<p>That is expected, yet we encourage you to generate a new submissions.</p>\n<blockquote>\n  <p>Update: if the local test is a random sample of the real test, and the test changed (particularly, its order), so the small test changed as well, and its so insanely small (like, WAY too small) that it adds all sort of weird patterns, e.g. a feature might change data-types from local-test to real-test. I had fixed those, but they need fix again.</p>\n</blockquote>\n<p>You can always use .csv tables to infer the schema first in the test sample, then load the test parquet tables and change the dtypes in schema and then concatenate them. Another approach would be to use the dtypes from train data and use them in schema for test tables. We understand that this poses a minor inconvenience for kagglers in this competition. I believe the data were stored in a way they don't consume more space than needed on a disk. Best of luck.</p>",
          "votes": null,
          "replies": [
            {
              "id": 2693467,
              "author_name": "enriquezaf",
              "author_url": "",
              "post_date": "03/12/2024 12:55:55",
              "content": "<blockquote>\n  <p>You can always use .csv tables to infer the schema first in the test sample, then load the test parquet tables and change the dtypes in schema and then concatenate them.</p>\n</blockquote>\n<p>Polars can read the schema of a parquet without reading the data, <a href=\"https://docs.pola.rs/py-polars/html/reference/api/polars.read_parquet_schema.html\" target=\"_blank\">polars.read_parquet_schema</a></p>\n<blockquote>\n  <p>We understand that this poses a minor inconvenience for kagglers in this competition.</p>\n</blockquote>\n<p>I'm dealing with this problem right now and I personally enjoy this part of the competitions 😀.</p>",
              "votes": null,
              "replies": [
                {
                  "id": 2693525,
                  "author_name": "jetakow",
                  "author_url": "",
                  "post_date": "03/12/2024 13:42:45",
                  "content": "<blockquote>\n  <p>Polars can read the schema of a parquet without reading the data, polars.read_parquet_schema</p>\n</blockquote>\n<p>Awesome!</p>\n<blockquote>\n  <p>I'm dealing with this problem right now and I personally enjoy this part of the competitions 😀.</p>\n</blockquote>\n<p>We already removed a huge portion of data preparation. I am glad to see someone with positive feedback 😀 and I am glad you enjoy that. </p>",
                  "votes": null,
                  "replies": []
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 2692633,
      "author_name": "yunsuxiaozi",
      "author_url": "",
      "post_date": "03/12/2024 01:13:09",
      "content": "<p>May I ask if it is possible to exclude errors from the number of submissions? This competition has restarted, and I have made two more errors when submitting the original code.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2693163,
      "author_name": "ablaydosmaganbetov",
      "author_url": "",
      "post_date": "03/12/2024 09:05:00",
      "content": "<p>I checked and realized that all my previous submissions failed, now the status is 'Unranked'. Why? Is it technical error? Or is it because metric was changed and applied to the previous submissions? Or only new submissions will be accepted with new metric from now and onward? </p>\n<p>That's a bit scary as many effort had been made and now it seems that it is worthless… </p>\n<p>Can someone, please, explain if I missed something about it?</p>",
      "votes": null,
      "replies": [
        {
          "id": 2693179,
          "author_name": "yunsuxiaozi",
          "author_url": "",
          "post_date": "03/12/2024 09:24:12",
          "content": "<p>You need to resubmit to have a ranking.</p>",
          "votes": null,
          "replies": [
            {
              "id": 2693186,
              "author_name": "ablaydosmaganbetov",
              "author_url": "",
              "post_date": "03/12/2024 09:30:57",
              "content": "<p>Noted, thanks but do you know how resubmissions will affect ranking considering new metric? Will they be adjusted relative to the new metric or ranked as they are?</p>",
              "votes": null,
              "replies": [
                {
                  "id": 2694360,
                  "author_name": "fengsicheng",
                  "author_url": "",
                  "post_date": "03/13/2024 03:15:42",
                  "content": "<p>the new metric</p>",
                  "votes": null,
                  "replies": []
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 2696900,
      "author_name": "carloshuertas",
      "author_url": "",
      "post_date": "03/14/2024 15:59:24",
      "content": "<p>Did the train set change in <strong><em>any</em></strong> way at all? I am checking as if so, we need to re-download if experiments are done in local machines.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2697399,
          "author_name": "inversion",
          "author_url": "",
          "post_date": "03/14/2024 21:55:49",
          "content": "<p>No change in the training data.</p>",
          "votes": null,
          "replies": [
            {
              "id": 2703320,
              "author_name": "carloshuertas",
              "author_url": "",
              "post_date": "03/18/2024 05:26:54",
              "content": "<p>From the post <a href=\"https://www.kaggle.com/competitions/home-credit-credit-risk-model-stability/discussion/482474\" target=\"_blank\">here</a>.</p>\n<blockquote>\n  <p>Column date_decision is no longer the same</p>\n</blockquote>\n<p>\"no longer the same\" but the fact that train has not changed, makes the situation funny. I am fine doing any transformation, but whats the point if the train is not using those?</p>",
              "votes": null,
              "replies": [
                {
                  "id": 2703610,
                  "author_name": "jetakow",
                  "author_url": "",
                  "post_date": "03/18/2024 09:37:09",
                  "content": "<p>There was a request made several times that we should preserve the meaning of features such as column_with_date - date_decision, that is mostly preserved and that is also a reason why we left it in test base table. </p>",
                  "votes": null,
                  "replies": [
                    {
                      "id": 2703924,
                      "author_name": "carloshuertas",
                      "author_url": "",
                      "post_date": "03/18/2024 13:30:43",
                      "content": "<p>According to rule 7C, external data might be Ok (granted some other aspects).</p>\n<p>I assume this no longer applies? There is no way to know if I am leaking into the future without knowing the date I should constraint myself.</p>\n<p>What was the transformation? Some random date such as the delta of dates is preserved but nothing else?</p>",
                      "votes": null,
                      "replies": [
                        {
                          "id": 2705434,
                          "author_name": "jetakow",
                          "author_url": "",
                          "post_date": "03/19/2024 10:39:31",
                          "content": "<blockquote>\n  <p>According to rule 7C, external data might be Ok</p>\n</blockquote>\n<p>It's all about the licence, if we can not use it internally because of a no-commerce clause, then it is out of the scope.  </p>\n<blockquote>\n  <p>I assume this no longer applies? </p>\n</blockquote>\n<p>It still does.</p>\n<blockquote>\n  <p>What was the transformation?</p>\n</blockquote>\n<p>Please understand, that we can not reveal it for the sake of making competition more fair for everyone.</p>",
                          "votes": null,
                          "replies": [
                            {
                              "id": 2705672,
                              "author_name": "carloshuertas",
                              "author_url": "",
                              "post_date": "03/19/2024 13:06:36",
                              "content": "<p>Still very confusing. Let me be explicit.</p>\n<p>1) I grab some free data source, available to everyone, blah blah, all good for the competition.</p>\n<p>2) That data is dependent of time, let's say, SP500 performance at the time of application (actually, at the day before application is better)</p>\n<p>3) I can perfectly use in train. What column do you suggest to use in test to avoid leakage? </p>\n<p>Thats my point, without knowing the transformation its unclear how to build something useful.</p>\n<p>Now if this means no external data is allowed anymore, then thats it.</p>",
                              "votes": null,
                              "replies": [
                                {
                                  "id": 2705889,
                                  "author_name": "jetakow",
                                  "author_url": "",
                                  "post_date": "03/19/2024 15:09:27",
                                  "content": "<p>There is no restriction on using S&amp;P500 or any other time series data that you think might help you improve the prediction. Before the transformation of test data we were ready for kagglers to freely use such data sources and improve their models based on that. We somewhat doubt that S&amp;P500 (and similar) will significantly improve the stability score (before the transformation). What could have a real impact on the stability score is covid period and we didn't want to forbid it, because there is no way to fully control the use of covid related time series (or any). </p>\n<p>After the transformation of data, I completely understand your concern and the motivation behind your post. Please consider following two scenarios 1) I don't reveal the transformation and you won't be able to use fully the, e.g. SP500 time series 2) I reveal the transformation and possibly making the whole \"fix\" and all the effort around it useless. We aimed for making the competition more fair for everyone after it was revealed that the gain from metric hack is significant. Our fix should neutralize almost 100% of the discovered hack. Revealing you the transformation might jeopardize all the effort we put into the fix and I can not allow that. Transformation of test data will not be revealed. </p>",
                                  "votes": null,
                                  "replies": []
                                }
                              ]
                            }
                          ]
                        }
                      ]
                    },
                    {
                      "id": 2712183,
                      "author_name": "stenford23",
                      "author_url": "",
                      "post_date": "03/23/2024 11:05:56",
                      "content": "<p><a href=\"https://www.kaggle.com/jetakow\" target=\"_blank\">@jetakow</a> is also aggregation over a single date decision of other column preserved or it does not work anymore?</p>",
                      "votes": null,
                      "replies": [
                        {
                          "id": 2712400,
                          "author_name": "jetakow",
                          "author_url": "",
                          "post_date": "03/23/2024 14:38:22",
                          "content": "<p><a href=\"https://www.kaggle.com/stenford23\" target=\"_blank\">@stenford23</a> that would also not be preserved. What we tried to preserve are date differences such as date_diff_feature = date_decision - aggregation_fun(date_col).</p>",
                          "votes": null,
                          "replies": [
                            {
                              "id": 2712459,
                              "author_name": "stenford23",
                              "author_url": "",
                              "post_date": "03/23/2024 15:11:38",
                              "content": "<p>tyvm <a href=\"https://www.kaggle.com/jetakow\" target=\"_blank\">@jetakow</a> .<br>\nIs also date_col1 - date_col2 preserved? (difference not with date decision but between other date on same case id)</p>",
                              "votes": null,
                              "replies": [
                                {
                                  "id": 2712953,
                                  "author_name": "jetakow",
                                  "author_url": "",
                                  "post_date": "03/23/2024 21:17:11",
                                  "content": "<p>The answer is again, in most cases yes. And I can't comment on the actual difference, but about the usability of such a feature.</p>",
                                  "votes": null,
                                  "replies": []
                                }
                              ]
                            }
                          ]
                        }
                      ]
                    }
                  ]
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 2698022,
      "author_name": "bartvanwoesik",
      "author_url": "",
      "post_date": "03/15/2024 08:31:15",
      "content": "<p>Hi, I'm trying to make a dummy submission to check if the process works, but keep getting this error:<br>\n<em>Submission Scoring Error\nYour notebook generated a submission file with incorrect format. Some examples causing this are: wrong number of rows or columns, empty values, an incorrect data type for a value, or invalid submission values from what is expected. See more debugging tips</em></p>\n<p>While I have the following output file: <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4157236%2F591a9d809472bba0760837939a9436d2%2FScreenshot%202024-03-15%20092918.png?generation=1710491446315645&amp;alt=media\"></p>\n<p>This does contain all the criteria that is stated on the overview pages. What am I doing wrong here?</p>",
      "votes": null,
      "replies": [
        {
          "id": 2698152,
          "author_name": "dimakoshman",
          "author_url": "",
          "post_date": "03/15/2024 10:26:38",
          "content": "<p>Are you copying the sample submission or using case ids from the test_base table?</p>",
          "votes": null,
          "replies": [
            {
              "id": 2698191,
              "author_name": "bartvanwoesik",
              "author_url": "",
              "post_date": "03/15/2024 10:43:55",
              "content": "<p>Here I'm using the sample submission, but also got the same error when using the test_base table and using a dummy model. See this post: <a href=\"https://www.kaggle.com/competitions/home-credit-credit-risk-model-stability/discussion/483635\" target=\"_blank\">https://www.kaggle.com/competitions/home-credit-credit-risk-model-stability/discussion/483635</a></p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2715501,
      "author_name": "bartvanwoesik",
      "author_url": "",
      "post_date": "03/25/2024 14:31:12",
      "content": "<p>Hi guys,<br>\nI'm still not able to subbmit my results. I'm doing this project on my own laptop and then I only want to submit the predictions, is this a wrong approach?</p>\n<p>See my previous issue:<br>\n<a href=\"https://www.kaggle.com/competitions/home-credit-credit-risk-model-stability/discussion/483262#2698022\" target=\"_blank\">https://www.kaggle.com/competitions/home-credit-credit-risk-model-stability/discussion/483262#2698022</a></p>\n<p>I could really use some support in this.<br>\nFurhtermore this is the repo of my code: <a href=\"https://github.com/BartvanWoesik/home_credit\" target=\"_blank\">https://github.com/BartvanWoesik/home_credit</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 2715563,
          "author_name": "dimakoshman",
          "author_url": "",
          "post_date": "03/25/2024 15:07:53",
          "content": "<p>I've taken a look and noticed that in train.py predictions are saved in \"predictions.csv\", while the expected filename is \"submission.csv\". May that be the reason? Also, you can't submit just the predictions, this is a code competition.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2834162,
      "author_name": "bluepill",
      "author_url": "",
      "post_date": "05/24/2024 15:29:11",
      "content": "<p>Hi, <a href=\"https://www.kaggle.com/inversion\" target=\"_blank\">@inversion</a>, for some reason previously successfully submitted code can now fail with OOM errors if you submit them again, I have created a discussion about it, a few ppl said they encounter the same. <a href=\"https://www.kaggle.com/competitions/home-credit-credit-risk-model-stability/discussion/506542\" target=\"_blank\">https://www.kaggle.com/competitions/home-credit-credit-risk-model-stability/discussion/506542</a></p>\n<p>What can be the root of such behaviour? </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2837364,
      "author_name": "raghavmaheshwarii",
      "author_url": "",
      "post_date": "05/26/2024 12:51:38",
      "content": "<p>Hi, this is my first competition ever and i had just wonder the dataset and the puzzle so hard thanks to all of you who provide public notebook that's why i can understand and create some feature engineering.🙂</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2692133": "Please report any problems here.\n\nGood luck!",
    "2692594": "What *exactly* changed in the train & test data?\n\nI picked some random submissions and they work and I got a score.\n\nWhen I try to run (live notebook) those same notebooks (that somehow worked to submit) they crash, e.g. crashing with \"numberofoverdueinstlmaxdat_641D\".\n\nUpdate: if the local test is a random sample of the real test, and the test changed (particularly, its order), so the small test changed as well, and its so insanely small (like, WAY too small) that it adds all sort of weird patterns, e.g. a feature might change data-types from local-test to real-test. I had fixed those, but they need fix again.",
    "2692633": "May I ask if it is possible to exclude errors from the number of submissions? This competition has restarted, and I have made two more errors when submitting the original code.",
    "2693163": "I checked and realized that all my previous submissions failed, now the status is 'Unranked'. Why? Is it technical error? Or is it because metric was changed and applied to the previous submissions? Or only new submissions will be accepted with new metric from now and onward? \n\nThat's a bit scary as many effort had been made and now it seems that it is worthless... \n\nCan someone, please, explain if I missed something about it?",
    "2693179": "You need to resubmit to have a ranking.",
    "2693186": "Noted, thanks but do you know how resubmissions will affect ranking considering new metric? Will they be adjusted relative to the new metric or ranked as they are?",
    "2693298": "> What exactly changed in the train & test data?\n\nI am afraid we can't say that for the sake of making the competition more fair for everyone. \n\n> I picked some random submissions and they work and I got a score.\n\nThat is expected, yet we encourage you to generate a new submissions.\n\n> Update: if the local test is a random sample of the real test, and the test changed (particularly, its order), so the small test changed as well, and its so insanely small (like, WAY too small) that it adds all sort of weird patterns, e.g. a feature might change data-types from local-test to real-test. I had fixed those, but they need fix again.\n\nYou can always use .csv tables to infer the schema first in the test sample, then load the test parquet tables and change the dtypes in schema and then concatenate them. Another approach would be to use the dtypes from train data and use them in schema for test tables. We understand that this poses a minor inconvenience for kagglers in this competition. I believe the data were stored in a way they don't consume more space than needed on a disk. Best of luck.",
    "2693467": ">You can always use .csv tables to infer the schema first in the test sample, then load the test parquet tables and change the dtypes in schema and then concatenate them.\n\nPolars can read the schema of a parquet without reading the data, [polars.read_parquet_schema](https://docs.pola.rs/py-polars/html/reference/api/polars.read_parquet_schema.html)\n\n>We understand that this poses a minor inconvenience for kagglers in this competition.\n\nI'm dealing with this problem right now and I personally enjoy this part of the competitions 😀.",
    "2693525": "> Polars can read the schema of a parquet without reading the data, polars.read_parquet_schema\n\nAwesome!\n\n> I'm dealing with this problem right now and I personally enjoy this part of the competitions 😀.\n\nWe already removed a huge portion of data preparation. I am glad to see someone with positive feedback 😀 and I am glad you enjoy that.",
    "2694360": "the new metric",
    "2696900": "Did the train set change in ***any*** way at all? I am checking as if so, we need to re-download if experiments are done in local machines.",
    "2697399": "No change in the training data.",
    "2698022": "Hi, I'm trying to make a dummy submission to check if the process works, but keep getting this error:\n*Submission Scoring Error\nYour notebook generated a submission file with incorrect format. Some examples causing this are: wrong number of rows or columns, empty values, an incorrect data type for a value, or invalid submission values from what is expected. See more debugging tips*\n\nWhile I have the following output file: ![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4157236%2F591a9d809472bba0760837939a9436d2%2FScreenshot%202024-03-15%20092918.png?generation=1710491446315645&alt=media)\n\nThis does contain all the criteria that is stated on the overview pages. What am I doing wrong here?",
    "2698152": "Are you copying the sample submission or using case ids from the test_base table?",
    "2698191": "Here I'm using the sample submission, but also got the same error when using the test_base table and using a dummy model. See this post: https://www.kaggle.com/competitions/home-credit-credit-risk-model-stability/discussion/483635",
    "2703320": "From the post [here](https://www.kaggle.com/competitions/home-credit-credit-risk-model-stability/discussion/482474).\n\n>Column date_decision is no longer the same\n\n\"no longer the same\" but the fact that train has not changed, makes the situation funny. I am fine doing any transformation, but whats the point if the train is not using those?",
    "2703610": "There was a request made several times that we should preserve the meaning of features such as column_with_date - date_decision, that is mostly preserved and that is also a reason why we left it in test base table.",
    "2703924": "According to rule 7C, external data might be Ok (granted some other aspects).\n\nI assume this no longer applies? There is no way to know if I am leaking into the future without knowing the date I should constraint myself.\n\nWhat was the transformation? Some random date such as the delta of dates is preserved but nothing else?",
    "2705434": "> According to rule 7C, external data might be Ok\n\nIt's all about the licence, if we can not use it internally because of a no-commerce clause, then it is out of the scope.  \n\n> I assume this no longer applies? \n\nIt still does.\n\n> What was the transformation?\n\nPlease understand, that we can not reveal it for the sake of making competition more fair for everyone.",
    "2705672": "Still very confusing. Let me be explicit.\n\n1) I grab some free data source, available to everyone, blah blah, all good for the competition.\n\n2) That data is dependent of time, let's say, SP500 performance at the time of application (actually, at the day before application is better)\n\n3) I can perfectly use in train. What column do you suggest to use in test to avoid leakage? \n\nThats my point, without knowing the transformation its unclear how to build something useful.\n\nNow if this means no external data is allowed anymore, then thats it.",
    "2705889": "There is no restriction on using S&P500 or any other time series data that you think might help you improve the prediction. Before the transformation of test data we were ready for kagglers to freely use such data sources and improve their models based on that. We somewhat doubt that S&P500 (and similar) will significantly improve the stability score (before the transformation). What could have a real impact on the stability score is covid period and we didn't want to forbid it, because there is no way to fully control the use of covid related time series (or any). \n\nAfter the transformation of data, I completely understand your concern and the motivation behind your post. Please consider following two scenarios 1) I don't reveal the transformation and you won't be able to use fully the, e.g. SP500 time series 2) I reveal the transformation and possibly making the whole \"fix\" and all the effort around it useless. We aimed for making the competition more fair for everyone after it was revealed that the gain from metric hack is significant. Our fix should neutralize almost 100% of the discovered hack. Revealing you the transformation might jeopardize all the effort we put into the fix and I can not allow that. Transformation of test data will not be revealed.",
    "2712183": "jetakow is also aggregation over a single date decision of other column preserved or it does not work anymore?",
    "2712400": "stenford23 that would also not be preserved. What we tried to preserve are date differences such as date_diff_feature = date_decision - aggregation_fun(date_col).",
    "2712459": "tyvm @jetakow .\nIs also date_col1 - date_col2 preserved? (difference not with date decision but between other date on same case id)",
    "2712953": "The answer is again, in most cases yes. And I can't comment on the actual difference, but about the usability of such a feature.",
    "2715501": "Hi guys,\nI'm still not able to subbmit my results. I'm doing this project on my own laptop and then I only want to submit the predictions, is this a wrong approach?\n\nSee my previous issue:\nhttps://www.kaggle.com/competitions/home-credit-credit-risk-model-stability/discussion/483262#2698022\n\nI could really use some support in this.\nFurhtermore this is the repo of my code: https://github.com/BartvanWoesik/home_credit",
    "2715563": "I've taken a look and noticed that in train.py predictions are saved in \"predictions.csv\", while the expected filename is \"submission.csv\". May that be the reason? Also, you can't submit just the predictions, this is a code competition.",
    "2834162": "Hi, @inversion, for some reason previously successfully submitted code can now fail with OOM errors if you submit them again, I have created a discussion about it, a few ppl said they encounter the same. https://www.kaggle.com/competitions/home-credit-credit-risk-model-stability/discussion/506542\n\nWhat can be the root of such behaviour?",
    "2837364": "Hi, this is my first competition ever and i had just wonder the dataset and the puzzle so hard thanks to all of you who provide public notebook that's why i can understand and create some feature engineering.🙂"
  },
  "source": "meta"
}