{
  "id": 224380,
  "title": "Limits on Training Time for GPU and TPU Notebooks",
  "url": "/competitions/ranzcr-clip-catheter-line-classification/discussion/224380",
  "author_name": "",
  "post_date": "2021-03-08T08:39:12.739060100Z",
  "votes": 2,
  "comment_count": 4,
  "views": 0,
  "content": "<p>Hi,</p>\n<p>Usually, only TPU notebooks are stopped after 9 hours of training but my GPU notebook is also showing Cancelled after 9 hours. Are there any limits like that for training notebooks? </p>\n<p>Were you guys able to train a 5 fold cv model with 512 image size using Kaggle resources, can you share the batch size and highest possible resolution?</p>\n<p>Thank you</p>",
  "messages": [
    {
      "id": "1230562",
      "postDate": "03/08/2021 08:39:12",
      "content": "<p>Hi,</p>\n<p>Usually, only TPU notebooks are stopped after 9 hours of training but my GPU notebook is also showing Cancelled after 9 hours. Are there any limits like that for training notebooks? </p>\n<p>Were you guys able to train a 5 fold cv model with 512 image size using Kaggle resources, can you share the batch size and highest possible resolution?</p>\n<p>Thank you</p>",
      "rawMarkdown": "Hi,\n\nUsually, only TPU notebooks are stopped after 9 hours of training but my GPU notebook is also showing Cancelled after 9 hours. Are there any limits like that for training notebooks? \n\nWere you guys able to train a 5 fold cv model with 512 image size using Kaggle resources, can you share the batch size and highest possible resolution?\n\nThank you",
      "votes": null
    },
    {
      "id": "1230606",
      "postDate": "03/08/2021 09:05:23",
      "content": "<p>I think GPU notebooks do indeed have (and for quite a while have had) a 9 hour limit. With smaller models, I've managed to train 5 fold across 2 notebooks with 600 x 600. With many larger models I've done either 1 fold per notebook or more typically used my own machine.</p>",
      "rawMarkdown": "I think GPU notebooks do indeed have (and for quite a while have had) a 9 hour limit. With smaller models, I've managed to train 5 fold across 2 notebooks with 600 x 600. With many larger models I've done either 1 fold per notebook or more typically used my own machine.",
      "votes": null
    },
    {
      "id": "1230798",
      "postDate": "03/08/2021 13:14:31",
      "content": "<p>yes, <a href=\"https://www.kaggle.com/abh001\" target=\"_blank\">@abh001</a> both GPU and TPU based running have a limit of running a maximum of 9 hrs, Either you can create one notebook for one fold. That way you could able to train a larger model but for fewer epochs as the data size is huge.</p>",
      "rawMarkdown": "yes, @abh001 both GPU and TPU based running have a limit of running a maximum of 9 hrs, Either you can create one notebook for one fold. That way you could able to train a larger model but for fewer epochs as the data size is huge.",
      "votes": null
    },
    {
      "id": "1231121",
      "postDate": "03/08/2021 17:22:26",
      "content": "<p>Thank you for the reply. Yeah, I am also thinking of splitting the folds into separate notebooks. I am using 512 x 512 with resnet200d.</p>",
      "rawMarkdown": "Thank you for the reply. Yeah, I am also thinking of splitting the folds into separate notebooks. I am using 512 x 512 with resnet200d.",
      "votes": null
    },
    {
      "id": "1231123",
      "postDate": "03/08/2021 17:23:50",
      "content": "<p>Thanks, I thought I was able to run GPU notebooks for more hours on the cassava competition. I will split the folds into separate notebooks. </p>",
      "rawMarkdown": "Thanks, I thought I was able to run GPU notebooks for more hours on the cassava competition. I will split the folds into separate notebooks.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1230606,
      "author_name": "bjoernholzhauer",
      "author_url": "",
      "post_date": "03/08/2021 09:05:23",
      "content": "<p>I think GPU notebooks do indeed have (and for quite a while have had) a 9 hour limit. With smaller models, I've managed to train 5 fold across 2 notebooks with 600 x 600. With many larger models I've done either 1 fold per notebook or more typically used my own machine.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1231121,
          "author_name": "abh001",
          "author_url": "",
          "post_date": "03/08/2021 17:22:26",
          "content": "<p>Thank you for the reply. Yeah, I am also thinking of splitting the folds into separate notebooks. I am using 512 x 512 with resnet200d.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1230798,
      "author_name": "vpkprasanna",
      "author_url": "",
      "post_date": "03/08/2021 13:14:31",
      "content": "<p>yes, <a href=\"https://www.kaggle.com/abh001\" target=\"_blank\">@abh001</a> both GPU and TPU based running have a limit of running a maximum of 9 hrs, Either you can create one notebook for one fold. That way you could able to train a larger model but for fewer epochs as the data size is huge.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1231123,
          "author_name": "abh001",
          "author_url": "",
          "post_date": "03/08/2021 17:23:50",
          "content": "<p>Thanks, I thought I was able to run GPU notebooks for more hours on the cassava competition. I will split the folds into separate notebooks. </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1230562": "Hi,\n\nUsually, only TPU notebooks are stopped after 9 hours of training but my GPU notebook is also showing Cancelled after 9 hours. Are there any limits like that for training notebooks? \n\nWere you guys able to train a 5 fold cv model with 512 image size using Kaggle resources, can you share the batch size and highest possible resolution?\n\nThank you",
    "1230606": "I think GPU notebooks do indeed have (and for quite a while have had) a 9 hour limit. With smaller models, I've managed to train 5 fold across 2 notebooks with 600 x 600. With many larger models I've done either 1 fold per notebook or more typically used my own machine.",
    "1230798": "yes, @abh001 both GPU and TPU based running have a limit of running a maximum of 9 hrs, Either you can create one notebook for one fold. That way you could able to train a larger model but for fewer epochs as the data size is huge.",
    "1231121": "Thank you for the reply. Yeah, I am also thinking of splitting the folds into separate notebooks. I am using 512 x 512 with resnet200d.",
    "1231123": "Thanks, I thought I was able to run GPU notebooks for more hours on the cassava competition. I will split the folds into separate notebooks."
  },
  "source": "meta"
}