{
  "id": 483639,
  "title": "Suggestion on not calculating submission counts for failed submissions",
  "url": "/competitions/home-credit-credit-risk-model-stability/discussion/483639",
  "author_name": "",
  "post_date": "2024-03-13T10:18:04.958176100Z",
  "votes": 6,
  "comment_count": 5,
  "views": 0,
  "content": "<p>There are about 30 files in this competition, and the memory is very large. I guess currently, no contestant has used all the files provided by the hoster.</p>\n<p>Due to inconsistent 'dtype' of training and testing data, as well as memory issues, it is believed that most contestants have encountered errors in submitting programs.</p>\n<p>For myself, every time I submit a file, I add a file based on the previous successfully submitted file to see if there are any errors. If one file reports an error, it may take five submissions a day to address the issue of program errors. It may take a week or even several weeks to import all the files into the program in this competition.</p>\n<p>I believe the competition hoster also does not want the contestants to spend a lot of time reading files and lack time thinking about algorithms.</p>\n<p>So, that's why I came up with the idea in the title.</p>",
  "messages": [
    {
      "id": "2694860",
      "postDate": "03/13/2024 10:18:04",
      "content": "<p>There are about 30 files in this competition, and the memory is very large. I guess currently, no contestant has used all the files provided by the hoster.</p>\n<p>Due to inconsistent 'dtype' of training and testing data, as well as memory issues, it is believed that most contestants have encountered errors in submitting programs.</p>\n<p>For myself, every time I submit a file, I add a file based on the previous successfully submitted file to see if there are any errors. If one file reports an error, it may take five submissions a day to address the issue of program errors. It may take a week or even several weeks to import all the files into the program in this competition.</p>\n<p>I believe the competition hoster also does not want the contestants to spend a lot of time reading files and lack time thinking about algorithms.</p>\n<p>So, that's why I came up with the idea in the title.</p>",
      "rawMarkdown": "There are about 30 files in this competition, and the memory is very large. I guess currently, no contestant has used all the files provided by the hoster.\n\nDue to inconsistent 'dtype' of training and testing data, as well as memory issues, it is believed that most contestants have encountered errors in submitting programs.\n\nFor myself, every time I submit a file, I add a file based on the previous successfully submitted file to see if there are any errors. If one file reports an error, it may take five submissions a day to address the issue of program errors. It may take a week or even several weeks to import all the files into the program in this competition.\n\nI believe the competition hoster also does not want the contestants to spend a lot of time reading files and lack time thinking about algorithms.\n\nSo, that's why I came up with the idea in the title.",
      "votes": null
    },
    {
      "id": "2694892",
      "postDate": "03/13/2024 10:32:23",
      "content": "<p>I hope this is implemented across all code competitions and not just here <a href=\"https://www.kaggle.com/yunsuxiaozi\" target=\"_blank\">@yunsuxiaozi</a> </p>",
      "rawMarkdown": "I hope this is implemented across all code competitions and not just here @yunsuxiaozi",
      "votes": null
    },
    {
      "id": "2695048",
      "postDate": "03/13/2024 12:43:56",
      "content": "<p>Yeah the workflow for this competition is annoying. Errors can occur for weird reasons when they swap the testsets, and it's very hard to debug because of the daily limit. I don't think it should be possible for a notebook to run fine and generate a submissions file in the correct format to then NOT run fine when evaluated.</p>",
      "rawMarkdown": "Yeah the workflow for this competition is annoying. Errors can occur for weird reasons when they swap the testsets, and it's very hard to debug because of the daily limit. I don't think it should be possible for a notebook to run fine and generate a submissions file in the correct format to then NOT run fine when evaluated.",
      "votes": null
    },
    {
      "id": "2695556",
      "postDate": "03/13/2024 18:24:15",
      "content": "<p>It's a difficult choice for sure. There has to be balance between the number of submissions and how hard are the data to process. We considered a higher number of submissions, but only with the number that we have now, we can be sure the competition won't be focused on LB probing. Regarding inconsistent dtypes, well I believe everyone can read the schema of test parquets and compare to the schema in merged tables from train. There are only few types depending on the type of column. One strategy would be to prepare your own dictionary of column name and dtype and just work with that on test. If there is an error, you can filter the values or replace them with \"new\" category values. </p>",
      "rawMarkdown": "It's a difficult choice for sure. There has to be balance between the number of submissions and how hard are the data to process. We considered a higher number of submissions, but only with the number that we have now, we can be sure the competition won't be focused on LB probing. Regarding inconsistent dtypes, well I believe everyone can read the schema of test parquets and compare to the schema in merged tables from train. There are only few types depending on the type of column. One strategy would be to prepare your own dictionary of column name and dtype and just work with that on test. If there is an error, you can filter the values or replace them with \"new\" category values.",
      "votes": null
    },
    {
      "id": "2695898",
      "postDate": "03/13/2024 23:54:28",
      "content": "<p>Thank you for the official response. I have noticed that many people use 'parquet' files instead of 'csv' files. Is it because it can avoid some errors?</p>",
      "rawMarkdown": "Thank you for the official response. I have noticed that many people use 'parquet' files instead of 'csv' files. Is it because it can avoid some errors?",
      "votes": null
    },
    {
      "id": "2696389",
      "postDate": "03/14/2024 09:42:44",
      "content": "<p>I can't possibly know what is the reason. The reason why we included it was that we believe not everyone is familiar with .parquet table format. You could save potentially a little of work using .csv instead of .parquet, but I would guess that it also makes your whole pipeline a lot slower. The choice is up to you, what is more convenient or you to use. </p>",
      "rawMarkdown": "I can't possibly know what is the reason. The reason why we included it was that we believe not everyone is familiar with .parquet table format. You could save potentially a little of work using .csv instead of .parquet, but I would guess that it also makes your whole pipeline a lot slower. The choice is up to you, what is more convenient or you to use.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2694892,
      "author_name": "ravi20076",
      "author_url": "",
      "post_date": "03/13/2024 10:32:23",
      "content": "<p>I hope this is implemented across all code competitions and not just here <a href=\"https://www.kaggle.com/yunsuxiaozi\" target=\"_blank\">@yunsuxiaozi</a> </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2695048,
      "author_name": "alexanderholmberg",
      "author_url": "",
      "post_date": "03/13/2024 12:43:56",
      "content": "<p>Yeah the workflow for this competition is annoying. Errors can occur for weird reasons when they swap the testsets, and it's very hard to debug because of the daily limit. I don't think it should be possible for a notebook to run fine and generate a submissions file in the correct format to then NOT run fine when evaluated.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2695556,
      "author_name": "jetakow",
      "author_url": "",
      "post_date": "03/13/2024 18:24:15",
      "content": "<p>It's a difficult choice for sure. There has to be balance between the number of submissions and how hard are the data to process. We considered a higher number of submissions, but only with the number that we have now, we can be sure the competition won't be focused on LB probing. Regarding inconsistent dtypes, well I believe everyone can read the schema of test parquets and compare to the schema in merged tables from train. There are only few types depending on the type of column. One strategy would be to prepare your own dictionary of column name and dtype and just work with that on test. If there is an error, you can filter the values or replace them with \"new\" category values. </p>",
      "votes": null,
      "replies": [
        {
          "id": 2695898,
          "author_name": "yunsuxiaozi",
          "author_url": "",
          "post_date": "03/13/2024 23:54:28",
          "content": "<p>Thank you for the official response. I have noticed that many people use 'parquet' files instead of 'csv' files. Is it because it can avoid some errors?</p>",
          "votes": null,
          "replies": [
            {
              "id": 2696389,
              "author_name": "jetakow",
              "author_url": "",
              "post_date": "03/14/2024 09:42:44",
              "content": "<p>I can't possibly know what is the reason. The reason why we included it was that we believe not everyone is familiar with .parquet table format. You could save potentially a little of work using .csv instead of .parquet, but I would guess that it also makes your whole pipeline a lot slower. The choice is up to you, what is more convenient or you to use. </p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2694860": "There are about 30 files in this competition, and the memory is very large. I guess currently, no contestant has used all the files provided by the hoster.\n\nDue to inconsistent 'dtype' of training and testing data, as well as memory issues, it is believed that most contestants have encountered errors in submitting programs.\n\nFor myself, every time I submit a file, I add a file based on the previous successfully submitted file to see if there are any errors. If one file reports an error, it may take five submissions a day to address the issue of program errors. It may take a week or even several weeks to import all the files into the program in this competition.\n\nI believe the competition hoster also does not want the contestants to spend a lot of time reading files and lack time thinking about algorithms.\n\nSo, that's why I came up with the idea in the title.",
    "2694892": "I hope this is implemented across all code competitions and not just here @yunsuxiaozi",
    "2695048": "Yeah the workflow for this competition is annoying. Errors can occur for weird reasons when they swap the testsets, and it's very hard to debug because of the daily limit. I don't think it should be possible for a notebook to run fine and generate a submissions file in the correct format to then NOT run fine when evaluated.",
    "2695556": "It's a difficult choice for sure. There has to be balance between the number of submissions and how hard are the data to process. We considered a higher number of submissions, but only with the number that we have now, we can be sure the competition won't be focused on LB probing. Regarding inconsistent dtypes, well I believe everyone can read the schema of test parquets and compare to the schema in merged tables from train. There are only few types depending on the type of column. One strategy would be to prepare your own dictionary of column name and dtype and just work with that on test. If there is an error, you can filter the values or replace them with \"new\" category values.",
    "2695898": "Thank you for the official response. I have noticed that many people use 'parquet' files instead of 'csv' files. Is it because it can avoid some errors?",
    "2696389": "I can't possibly know what is the reason. The reason why we included it was that we believe not everyone is familiar with .parquet table format. You could save potentially a little of work using .csv instead of .parquet, but I would guess that it also makes your whole pipeline a lot slower. The choice is up to you, what is more convenient or you to use."
  },
  "source": "meta"
}