{
  "id": 473831,
  "title": "Be wary of memory issues",
  "url": "/competitions/home-credit-credit-risk-model-stability/discussion/473831",
  "author_name": "",
  "post_date": "2024-02-06T08:05:38.725999600Z",
  "votes": 6,
  "comment_count": 3,
  "views": 0,
  "content": "<p>There are many files in the dataset of this competition, which is likely to encounter memory issues. For example, when running normally offline, the online memory exceeds the limit. We assume that the hidden test data is the same size as the training set, and we can find this problem by changing the path of the test data to the training data. Wish everyone a smooth completion of the competition.</p>",
  "messages": [
    {
      "id": "2638300",
      "postDate": "02/06/2024 08:05:38",
      "content": "<p>There are many files in the dataset of this competition, which is likely to encounter memory issues. For example, when running normally offline, the online memory exceeds the limit. We assume that the hidden test data is the same size as the training set, and we can find this problem by changing the path of the test data to the training data. Wish everyone a smooth completion of the competition.</p>",
      "rawMarkdown": "There are many files in the dataset of this competition, which is likely to encounter memory issues. For example, when running normally offline, the online memory exceeds the limit. We assume that the hidden test data is the same size as the training set, and we can find this problem by changing the path of the test data to the training data. Wish everyone a smooth completion of the competition.",
      "votes": null
    },
    {
      "id": "2638329",
      "postDate": "02/06/2024 08:19:06",
      "content": "<p>Please consider that test tables can be bigger in size than train or smaller. That is also part of the competition to deal with this unknown variable. The only thing that you know is \"Note, the hidden test_base.csv contains approximately 90% of the numbers of case_id values of train_base.csv\".</p>",
      "rawMarkdown": "Please consider that test tables can be bigger in size than train or smaller. That is also part of the competition to deal with this unknown variable. The only thing that you know is \"Note, the hidden test_base.csv contains approximately 90% of the numbers of case_id values of train_base.csv\".",
      "votes": null
    },
    {
      "id": "2640952",
      "postDate": "02/07/2024 07:18:34",
      "content": "<p><a href=\"https://www.kaggle.com/jetakow\" target=\"_blank\">@jetakow</a> Thanks for your comment. From a timeline perspective, is the data in the hidden test set in the future (relative to train)?</p>",
      "rawMarkdown": "jetakow Thanks for your comment. From a timeline perspective, is the data in the hidden test set in the future (relative to train)?",
      "votes": null
    },
    {
      "id": "2641077",
      "postDate": "02/07/2024 09:15:33",
      "content": "<p>There are definitely data from the future on the test sample. I am afraid I can't give you more details than that.</p>",
      "rawMarkdown": "There are definitely data from the future on the test sample. I am afraid I can't give you more details than that.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2638329,
      "author_name": "jetakow",
      "author_url": "",
      "post_date": "02/06/2024 08:19:06",
      "content": "<p>Please consider that test tables can be bigger in size than train or smaller. That is also part of the competition to deal with this unknown variable. The only thing that you know is \"Note, the hidden test_base.csv contains approximately 90% of the numbers of case_id values of train_base.csv\".</p>",
      "votes": null,
      "replies": [
        {
          "id": 2640952,
          "author_name": "narsil",
          "author_url": "",
          "post_date": "02/07/2024 07:18:34",
          "content": "<p><a href=\"https://www.kaggle.com/jetakow\" target=\"_blank\">@jetakow</a> Thanks for your comment. From a timeline perspective, is the data in the hidden test set in the future (relative to train)?</p>",
          "votes": null,
          "replies": [
            {
              "id": 2641077,
              "author_name": "jetakow",
              "author_url": "",
              "post_date": "02/07/2024 09:15:33",
              "content": "<p>There are definitely data from the future on the test sample. I am afraid I can't give you more details than that.</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2638300": "There are many files in the dataset of this competition, which is likely to encounter memory issues. For example, when running normally offline, the online memory exceeds the limit. We assume that the hidden test data is the same size as the training set, and we can find this problem by changing the path of the test data to the training data. Wish everyone a smooth completion of the competition.",
    "2638329": "Please consider that test tables can be bigger in size than train or smaller. That is also part of the competition to deal with this unknown variable. The only thing that you know is \"Note, the hidden test_base.csv contains approximately 90% of the numbers of case_id values of train_base.csv\".",
    "2640952": "jetakow Thanks for your comment. From a timeline perspective, is the data in the hidden test set in the future (relative to train)?",
    "2641077": "There are definitely data from the future on the test sample. I am afraid I can't give you more details than that."
  },
  "source": "meta"
}