{
  "id": 553884,
  "title": "Problem in training LSTM due to large dataset",
  "url": "/competitions/jane-street-real-time-market-data-forecasting/discussion/553884",
  "author_name": "",
  "post_date": "2024-12-29T03:58:53.218198700Z",
  "votes": null,
  "comment_count": 3,
  "views": 0,
  "content": "<p>Hello everyone, I've successfully preprocessed my dataset, but I'm facing challenges training my model due to its large size. My system's memory isn't sufficient to handle the entire dataset at once. Could anyone guide me on efficient strategies to train the model incrementally or manage large datasets effectively? Any advice or best practices would be greatly appreciated!****</p>",
  "messages": [
    {
      "id": "3083121",
      "postDate": "12/29/2024 03:58:53",
      "content": "<p>Hello everyone, I've successfully preprocessed my dataset, but I'm facing challenges training my model due to its large size. My system's memory isn't sufficient to handle the entire dataset at once. Could anyone guide me on efficient strategies to train the model incrementally or manage large datasets effectively? Any advice or best practices would be greatly appreciated!****</p>",
      "rawMarkdown": "Hello everyone, I've successfully preprocessed my dataset, but I'm facing challenges training my model due to its large size. My system's memory isn't sufficient to handle the entire dataset at once. Could anyone guide me on efficient strategies to train the model incrementally or manage large datasets effectively? Any advice or best practices would be greatly appreciated!****",
      "votes": null
    },
    {
      "id": "3083791",
      "postDate": "12/30/2024 03:12:43",
      "content": "<p>Adjust the batch size.</p>",
      "rawMarkdown": "Adjust the batch size.",
      "votes": null
    },
    {
      "id": "3083934",
      "postDate": "12/30/2024 07:24:08",
      "content": "<p>One strategy is to use data loaders to load the batches of data on the fly. <a href=\"https://pytorch.org/tutorials/beginner/basics/data_tutorial.html\" target=\"_blank\">Here </a>is the pytorch doc for their data loaders: </p>\n<p>If you are using tensorflow they have a great <a href=\"https://www.tensorflow.org/guide/data\" target=\"_blank\">guide </a>on ther tf.data class which can do similar things:</p>\n<p>As <a href=\"https://www.kaggle.com/junhanzangai\" target=\"_blank\">@junhanzangai</a> mentioned, you can adjust the batch size (anywhere between 16-1024 depending on your data) so that your RAM usage doesn't explode. </p>",
      "rawMarkdown": "One strategy is to use data loaders to load the batches of data on the fly. [Here ](https://pytorch.org/tutorials/beginner/basics/data_tutorial.html )is the pytorch doc for their data loaders: \n\nIf you are using tensorflow they have a great [guide ](https://www.tensorflow.org/guide/data)on ther tf.data class which can do similar things:\n\nAs @junhanzangai mentioned, you can adjust the batch size (anywhere between 16-1024 depending on your data) so that your RAM usage doesn't explode.",
      "votes": null
    },
    {
      "id": "3091056",
      "postDate": "01/08/2025 01:22:35",
      "content": "<p>Write the generator to yield the data will save your memory.<br>\nAnd used np.float32 instead of np.float64.</p>",
      "rawMarkdown": "Write the generator to yield the data will save your memory.\nAnd used np.float32 instead of np.float64.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3083791,
      "author_name": "junhanzangai",
      "author_url": "",
      "post_date": "12/30/2024 03:12:43",
      "content": "<p>Adjust the batch size.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3083934,
      "author_name": "alexzag",
      "author_url": "",
      "post_date": "12/30/2024 07:24:08",
      "content": "<p>One strategy is to use data loaders to load the batches of data on the fly. <a href=\"https://pytorch.org/tutorials/beginner/basics/data_tutorial.html\" target=\"_blank\">Here </a>is the pytorch doc for their data loaders: </p>\n<p>If you are using tensorflow they have a great <a href=\"https://www.tensorflow.org/guide/data\" target=\"_blank\">guide </a>on ther tf.data class which can do similar things:</p>\n<p>As <a href=\"https://www.kaggle.com/junhanzangai\" target=\"_blank\">@junhanzangai</a> mentioned, you can adjust the batch size (anywhere between 16-1024 depending on your data) so that your RAM usage doesn't explode. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3091056,
      "author_name": "yiconghuangkaitong",
      "author_url": "",
      "post_date": "01/08/2025 01:22:35",
      "content": "<p>Write the generator to yield the data will save your memory.<br>\nAnd used np.float32 instead of np.float64.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3083121": "Hello everyone, I've successfully preprocessed my dataset, but I'm facing challenges training my model due to its large size. My system's memory isn't sufficient to handle the entire dataset at once. Could anyone guide me on efficient strategies to train the model incrementally or manage large datasets effectively? Any advice or best practices would be greatly appreciated!****",
    "3083791": "Adjust the batch size.",
    "3083934": "One strategy is to use data loaders to load the batches of data on the fly. [Here ](https://pytorch.org/tutorials/beginner/basics/data_tutorial.html )is the pytorch doc for their data loaders: \n\nIf you are using tensorflow they have a great [guide ](https://www.tensorflow.org/guide/data)on ther tf.data class which can do similar things:\n\nAs @junhanzangai mentioned, you can adjust the batch size (anywhere between 16-1024 depending on your data) so that your RAM usage doesn't explode.",
    "3091056": "Write the generator to yield the data will save your memory.\nAnd used np.float32 instead of np.float64."
  },
  "source": "meta"
}