{
  "id": 163629,
  "title": "Using Kaggle GPU for training",
  "url": "/competitions/alaska2-image-steganalysis/discussion/163629",
  "author_name": "",
  "post_date": "2020-07-02T19:48:16.657232600Z",
  "votes": 4,
  "comment_count": 3,
  "views": 0,
  "content": "<p>Hi,\nThis is my first Kaggle competition. I am using the GPU provided in Kaggle for training my models. However, each epoch takes about 2 hours during training. Ideally, it would allow me to run 15 epochs every week since the total weekly quota is 30 hours, but since the committed notebook only runs for 9 hours, I cannot train for more than 4 epochs at a time.</p>\n\n<p>This makes experimenting with different approaches very difficult. Is it possible to increase the runtime and GPU quota for this competition? Otherwise, this competition is only limited to those with rich resources.</p>\n\n<p>PS: It would be great if anyone could give some good advice in this context.</p>",
  "messages": [
    {
      "id": "912896",
      "postDate": "07/02/2020 19:48:16",
      "content": "<p>Hi,\nThis is my first Kaggle competition. I am using the GPU provided in Kaggle for training my models. However, each epoch takes about 2 hours during training. Ideally, it would allow me to run 15 epochs every week since the total weekly quota is 30 hours, but since the committed notebook only runs for 9 hours, I cannot train for more than 4 epochs at a time.</p>\n\n<p>This makes experimenting with different approaches very difficult. Is it possible to increase the runtime and GPU quota for this competition? Otherwise, this competition is only limited to those with rich resources.</p>\n\n<p>PS: It would be great if anyone could give some good advice in this context.</p>",
      "rawMarkdown": "Hi,\nThis is my first Kaggle competition. I am using the GPU provided in Kaggle for training my models. However, each epoch takes about 2 hours during training. Ideally, it would allow me to run 15 epochs every week since the total weekly quota is 30 hours, but since the committed notebook only runs for 9 hours, I cannot train for more than 4 epochs at a time.\n\nThis makes experimenting with different approaches very difficult. Is it possible to increase the runtime and GPU quota for this competition? Otherwise, this competition is only limited to those with rich resources.\n\nPS: It would be great if anyone could give some good advice in this context.",
      "votes": null
    },
    {
      "id": "913072",
      "postDate": "07/03/2020 00:01:17",
      "content": "<p>I have the same problem with GPU and TPU. Seems like they're just not working... 😏 </p>",
      "rawMarkdown": "I have the same problem with GPU and TPU. Seems like they're just not working... 😏",
      "votes": null
    },
    {
      "id": "913191",
      "postDate": "07/03/2020 03:49:01",
      "content": "<p>You can train 4 epoch each run, save the results, and continue the training in the next run.\nTry google colab for (i think) 12 hours run </p>",
      "rawMarkdown": "You can train 4 epoch each run, save the results, and continue the training in the next run.\nTry google colab for (i think) 12 hours run",
      "votes": null
    },
    {
      "id": "918157",
      "postDate": "07/07/2020 03:13:27",
      "content": "<ol>\n<li>You cannot download dataset this large on colab as the disk space is ~30GB. This makes downloading and unzipping impossible.</li>\n<li>Saving intermittently is one option. But it hurts the model performance when using one cycle policy for fitting (training several times means using several cycles). Also, you cannot experiment like this running 2-3 epochs at a time.</li>\n</ol>\n\n<p>Maybe there is a better way to train on Kaggle, and I am too much of a noob to see it.</p>",
      "rawMarkdown": "1. You cannot download dataset this large on colab as the disk space is ~30GB. This makes downloading and unzipping impossible.\n2. Saving intermittently is one option. But it hurts the model performance when using one cycle policy for fitting (training several times means using several cycles). Also, you cannot experiment like this running 2-3 epochs at a time.\n\nMaybe there is a better way to train on Kaggle, and I am too much of a noob to see it.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 913072,
      "author_name": "mariapushkareva",
      "author_url": "",
      "post_date": "07/03/2020 00:01:17",
      "content": "<p>I have the same problem with GPU and TPU. Seems like they're just not working... 😏 </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 913191,
      "author_name": "nyleve",
      "author_url": "",
      "post_date": "07/03/2020 03:49:01",
      "content": "<p>You can train 4 epoch each run, save the results, and continue the training in the next run.\nTry google colab for (i think) 12 hours run </p>",
      "votes": null,
      "replies": [
        {
          "id": 918157,
          "author_name": "abhinavtripathi9",
          "author_url": "",
          "post_date": "07/07/2020 03:13:27",
          "content": "<ol>\n<li>You cannot download dataset this large on colab as the disk space is ~30GB. This makes downloading and unzipping impossible.</li>\n<li>Saving intermittently is one option. But it hurts the model performance when using one cycle policy for fitting (training several times means using several cycles). Also, you cannot experiment like this running 2-3 epochs at a time.</li>\n</ol>\n\n<p>Maybe there is a better way to train on Kaggle, and I am too much of a noob to see it.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "912896": "Hi,\nThis is my first Kaggle competition. I am using the GPU provided in Kaggle for training my models. However, each epoch takes about 2 hours during training. Ideally, it would allow me to run 15 epochs every week since the total weekly quota is 30 hours, but since the committed notebook only runs for 9 hours, I cannot train for more than 4 epochs at a time.\n\nThis makes experimenting with different approaches very difficult. Is it possible to increase the runtime and GPU quota for this competition? Otherwise, this competition is only limited to those with rich resources.\n\nPS: It would be great if anyone could give some good advice in this context.",
    "913072": "I have the same problem with GPU and TPU. Seems like they're just not working... 😏",
    "913191": "You can train 4 epoch each run, save the results, and continue the training in the next run.\nTry google colab for (i think) 12 hours run",
    "918157": "1. You cannot download dataset this large on colab as the disk space is ~30GB. This makes downloading and unzipping impossible.\n2. Saving intermittently is one option. But it hurts the model performance when using one cycle policy for fitting (training several times means using several cycles). Also, you cannot experiment like this running 2-3 epochs at a time.\n\nMaybe there is a better way to train on Kaggle, and I am too much of a noob to see it."
  },
  "source": "meta"
}