{
  "id": 552792,
  "title": "Public notebooks are stuck because of the memory issues",
  "url": "/competitions/jane-street-real-time-market-data-forecasting/discussion/552792",
  "author_name": "",
  "post_date": "2024-12-21T15:16:16.093931900Z",
  "votes": null,
  "comment_count": 2,
  "views": 0,
  "content": "<p>So if we don't have enough memory - we can use other sources to retrain models outside of the Kaggle. PAY MONEY? 🤬😃 (even 8$ is too much?)<br>\nYes, I tried to fix these issues by code - reducing memory usage, but it's keeps crashing. Maybe I missed some approaches?</p>\n<p>And else these top public LB scores - because of the new data with online retrain or I misunderstand something?<br>\nSorry for these questions.</p>",
  "messages": [
    {
      "id": "3077939",
      "postDate": "12/21/2024 15:16:16",
      "content": "<p>So if we don't have enough memory - we can use other sources to retrain models outside of the Kaggle. PAY MONEY? 🤬😃 (even 8$ is too much?)<br>\nYes, I tried to fix these issues by code - reducing memory usage, but it's keeps crashing. Maybe I missed some approaches?</p>\n<p>And else these top public LB scores - because of the new data with online retrain or I misunderstand something?<br>\nSorry for these questions.</p>",
      "rawMarkdown": "So if we don't have enough memory - we can use other sources to retrain models outside of the Kaggle. PAY MONEY? 🤬😃 (even 8$ is too much?)\nYes, I tried to fix these issues by code - reducing memory usage, but it's keeps crashing. Maybe I missed some approaches?\n\nAnd else these top public LB scores - because of the new data with online retrain or I misunderstand something?\nSorry for these questions.",
      "votes": null
    },
    {
      "id": "3078083",
      "postDate": "12/21/2024 18:52:23",
      "content": "<p>Something that helped me was to create my own pytorch dataloader class where I just collect samples from the polars lazyframe of the dataset. That way, you won't need to load the whole dataset into memory. Might add a bit of an I/O overhead but it definitely overcomes OOM.</p>",
      "rawMarkdown": "Something that helped me was to create my own pytorch dataloader class where I just collect samples from the polars lazyframe of the dataset. That way, you won't need to load the whole dataset into memory. Might add a bit of an I/O overhead but it definitely overcomes OOM.",
      "votes": null
    },
    {
      "id": "3081335",
      "postDate": "12/26/2024 16:06:40",
      "content": "<p>It works for NN's but not for GBDT models.</p>",
      "rawMarkdown": "It works for NN's but not for GBDT models.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3078083,
      "author_name": "yeraypabongonzalez",
      "author_url": "",
      "post_date": "12/21/2024 18:52:23",
      "content": "<p>Something that helped me was to create my own pytorch dataloader class where I just collect samples from the polars lazyframe of the dataset. That way, you won't need to load the whole dataset into memory. Might add a bit of an I/O overhead but it definitely overcomes OOM.</p>",
      "votes": null,
      "replies": [
        {
          "id": 3081335,
          "author_name": "zeroday19",
          "author_url": "",
          "post_date": "12/26/2024 16:06:40",
          "content": "<p>It works for NN's but not for GBDT models.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3077939": "So if we don't have enough memory - we can use other sources to retrain models outside of the Kaggle. PAY MONEY? 🤬😃 (even 8$ is too much?)\nYes, I tried to fix these issues by code - reducing memory usage, but it's keeps crashing. Maybe I missed some approaches?\n\nAnd else these top public LB scores - because of the new data with online retrain or I misunderstand something?\nSorry for these questions.",
    "3078083": "Something that helped me was to create my own pytorch dataloader class where I just collect samples from the polars lazyframe of the dataset. That way, you won't need to load the whole dataset into memory. Might add a bit of an I/O overhead but it definitely overcomes OOM.",
    "3081335": "It works for NN's but not for GBDT models."
  },
  "source": "meta"
}