{
  "id": 446795,
  "title": "GPU RAM issues for inference in Kaggle kernel",
  "url": "/competitions/predict-ai-model-runtime/discussion/446795",
  "author_name": "",
  "post_date": "2023-10-13T05:16:16.302203900Z",
  "votes": 3,
  "comment_count": 3,
  "views": 0,
  "content": "<p>Hi everyone,</p>\n<p>Upon developing models for layout collections, it seems crucial to device a good algorithm that takes into account the data loading and network structure. It's fairly easy to get the good-old 'CUDA out of memory' in Kaggle kernel if the GNN is too large or too many graphs are moved to GPU. And it's also tiring to find that you have to redesign and retrain your model only after your model explodes in Kaggle kernel (if you are training on your own machine).</p>\n<p>If you are running the training and/or inferencing notebooks on kaggle, what tips could you give on managing GPU memory for this competition and for general GNN or other deep learning projects?</p>\n<p>Thanks in advance!</p>",
  "messages": [
    {
      "id": "2480042",
      "postDate": "10/13/2023 05:16:16",
      "content": "<p>Hi everyone,</p>\n<p>Upon developing models for layout collections, it seems crucial to device a good algorithm that takes into account the data loading and network structure. It's fairly easy to get the good-old 'CUDA out of memory' in Kaggle kernel if the GNN is too large or too many graphs are moved to GPU. And it's also tiring to find that you have to redesign and retrain your model only after your model explodes in Kaggle kernel (if you are training on your own machine).</p>\n<p>If you are running the training and/or inferencing notebooks on kaggle, what tips could you give on managing GPU memory for this competition and for general GNN or other deep learning projects?</p>\n<p>Thanks in advance!</p>",
      "rawMarkdown": "Hi everyone,\n\nUpon developing models for layout collections, it seems crucial to device a good algorithm that takes into account the data loading and network structure. It's fairly easy to get the good-old 'CUDA out of memory' in Kaggle kernel if the GNN is too large or too many graphs are moved to GPU. And it's also tiring to find that you have to redesign and retrain your model only after your model explodes in Kaggle kernel (if you are training on your own machine).\n\nIf you are running the training and/or inferencing notebooks on kaggle, what tips could you give on managing GPU memory for this competition and for general GNN or other deep learning projects?\n\nThanks in advance!",
      "votes": null
    },
    {
      "id": "2481036",
      "postDate": "10/13/2023 17:55:04",
      "content": "<p>You can checkout out <a href=\"https://arxiv.org/pdf/2305.12322.pdf\" target=\"_blank\">paper</a> on how we avoid GPU out of memory when training on large graphs.</p>",
      "rawMarkdown": "You can checkout out [paper](https://arxiv.org/pdf/2305.12322.pdf) on how we avoid GPU out of memory when training on large graphs.",
      "votes": null
    },
    {
      "id": "2481411",
      "postDate": "10/14/2023 05:29:18",
      "content": "<p>I think we should rent more powerful GPUs to train the model, only use P100 or T4 is not enough.</p>",
      "rawMarkdown": "I think we should rent more powerful GPUs to train the model, only use P100 or T4 is not enough.",
      "votes": null
    },
    {
      "id": "2481940",
      "postDate": "10/14/2023 14:31:36",
      "content": "<p>As a matter of fact, I have tried and succeeded. In Dataloader you can store file names with a batchsize of 1, and then you can read all the information from one file ata time without exceeding the memory limit.  <a href=\"https://www.kaggle.com/chenboluo/layout-nlp-default-train-and-infer\" target=\"_blank\">https://www.kaggle.com/chenboluo/layout-nlp-default-train-and-infer</a></p>",
      "rawMarkdown": "As a matter of fact, I have tried and succeeded. In Dataloader you can store file names with a batchsize of 1, and then you can read all the information from one file ata time without exceeding the memory limit.  https://www.kaggle.com/chenboluo/layout-nlp-default-train-and-infer",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2481036,
      "author_name": "mangpophothilimthana",
      "author_url": "",
      "post_date": "10/13/2023 17:55:04",
      "content": "<p>You can checkout out <a href=\"https://arxiv.org/pdf/2305.12322.pdf\" target=\"_blank\">paper</a> on how we avoid GPU out of memory when training on large graphs.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2481411,
      "author_name": "lizhecheng",
      "author_url": "",
      "post_date": "10/14/2023 05:29:18",
      "content": "<p>I think we should rent more powerful GPUs to train the model, only use P100 or T4 is not enough.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2481940,
      "author_name": "chenboluo",
      "author_url": "",
      "post_date": "10/14/2023 14:31:36",
      "content": "<p>As a matter of fact, I have tried and succeeded. In Dataloader you can store file names with a batchsize of 1, and then you can read all the information from one file ata time without exceeding the memory limit.  <a href=\"https://www.kaggle.com/chenboluo/layout-nlp-default-train-and-infer\" target=\"_blank\">https://www.kaggle.com/chenboluo/layout-nlp-default-train-and-infer</a></p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2480042": "Hi everyone,\n\nUpon developing models for layout collections, it seems crucial to device a good algorithm that takes into account the data loading and network structure. It's fairly easy to get the good-old 'CUDA out of memory' in Kaggle kernel if the GNN is too large or too many graphs are moved to GPU. And it's also tiring to find that you have to redesign and retrain your model only after your model explodes in Kaggle kernel (if you are training on your own machine).\n\nIf you are running the training and/or inferencing notebooks on kaggle, what tips could you give on managing GPU memory for this competition and for general GNN or other deep learning projects?\n\nThanks in advance!",
    "2481036": "You can checkout out [paper](https://arxiv.org/pdf/2305.12322.pdf) on how we avoid GPU out of memory when training on large graphs.",
    "2481411": "I think we should rent more powerful GPUs to train the model, only use P100 or T4 is not enough.",
    "2481940": "As a matter of fact, I have tried and succeeded. In Dataloader you can store file names with a batchsize of 1, and then you can read all the information from one file ata time without exceeding the memory limit.  https://www.kaggle.com/chenboluo/layout-nlp-default-train-and-infer"
  },
  "source": "meta"
}