{
  "id": 394841,
  "title": "How to train such huge data?",
  "url": "/competitions/icecube-neutrinos-in-deep-ice/discussion/394841",
  "author_name": "",
  "post_date": "2023-03-15T02:33:18.412995500Z",
  "votes": 5,
  "comment_count": 2,
  "views": 0,
  "content": "<p>Hi, everyone:</p>\n<p>We know the the dataset is huge and probably bigger than your RAM, it's hard to load them when we want to train our model,  like graphnet.</p>\n<p>In my thoughts, we can use <strong>three iterations</strong> when train our model. Here is how we can do:</p>\n<pre><code> epoch_num  (epochs):\n     batch_id  batchid_list:\n        dataset =   \n        dataloader =  \n         _  dataloader:\n            ......  \n</code></pre>",
  "messages": [
    {
      "id": "2182220",
      "postDate": "03/15/2023 02:33:18",
      "content": "<p>Hi, everyone:</p>\n<p>We know the the dataset is huge and probably bigger than your RAM, it's hard to load them when we want to train our model,  like graphnet.</p>\n<p>In my thoughts, we can use <strong>three iterations</strong> when train our model. Here is how we can do:</p>\n<pre><code> epoch_num  (epochs):\n     batch_id  batchid_list:\n        dataset =   \n        dataloader =  \n         _  dataloader:\n            ......  \n</code></pre>",
      "rawMarkdown": "Hi, everyone:\n\nWe know the the dataset is huge and probably bigger than your RAM, it's hard to load them when we want to train our model,  like graphnet.\n\nIn my thoughts, we can use **three iterations** when train our model. Here is how we can do:\n\n```python\n\nfor epoch_num in range(epochs):\n    for batch_id in batchid_list:\n        dataset = ''  # generate dataset from 'batch_x.parquet' one by one using func from \"graphnet.data.dataset.Dataset\"; \"from torch.utils.data.Dataset\" or others\n        dataloader = '' #  gen from dataset \n        for _ in dataloader:\n            ......  # train model\n\n```",
      "votes": null
    },
    {
      "id": "2182434",
      "postDate": "03/15/2023 05:34:50",
      "content": "<p>thanks for that informations👀</p>",
      "rawMarkdown": "thanks for that informations👀",
      "votes": null
    },
    {
      "id": "2209156",
      "postDate": "04/04/2023 13:55:25",
      "content": "<p>thanks for sharing</p>",
      "rawMarkdown": "thanks for sharing",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2182434,
      "author_name": "siliva1799",
      "author_url": "",
      "post_date": "03/15/2023 05:34:50",
      "content": "<p>thanks for that informations👀</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2209156,
      "author_name": "liuxin22",
      "author_url": "",
      "post_date": "04/04/2023 13:55:25",
      "content": "<p>thanks for sharing</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2182220": "Hi, everyone:\n\nWe know the the dataset is huge and probably bigger than your RAM, it's hard to load them when we want to train our model,  like graphnet.\n\nIn my thoughts, we can use **three iterations** when train our model. Here is how we can do:\n\n```python\n\nfor epoch_num in range(epochs):\n    for batch_id in batchid_list:\n        dataset = ''  # generate dataset from 'batch_x.parquet' one by one using func from \"graphnet.data.dataset.Dataset\"; \"from torch.utils.data.Dataset\" or others\n        dataloader = '' #  gen from dataset \n        for _ in dataloader:\n            ......  # train model\n\n```",
    "2182434": "thanks for that informations👀",
    "2209156": "thanks for sharing"
  },
  "source": "meta"
}