{
  "id": 168186,
  "title": "Training time for 384x384 images on tpu(pytorch)",
  "url": "/competitions/siim-isic-melanoma-classification/discussion/168186",
  "author_name": "",
  "post_date": "2020-07-19T14:48:23.294594400Z",
  "votes": null,
  "comment_count": 6,
  "views": 0,
  "content": "<p>I have been experenting with the 384x384 centre cropped images posted by Chris Deotte on GPU but it the training time on gpu came out to be very slow and so decided to try on tpus. Now basically this is the first time am trying to use tpus on pytorch and if anyone has used the 386 images what's the average training time per epoch and batch size?\nI was trying batch size of 16  and after around 400 batches the notebook throws memory error on ram and restarts. Is this due to some bottleneck in my code or I need to reduce the batch size?</p>",
  "messages": [
    {
      "id": "935661",
      "postDate": "07/19/2020 14:48:23",
      "content": "<p>I have been experenting with the 384x384 centre cropped images posted by Chris Deotte on GPU but it the training time on gpu came out to be very slow and so decided to try on tpus. Now basically this is the first time am trying to use tpus on pytorch and if anyone has used the 386 images what's the average training time per epoch and batch size?\nI was trying batch size of 16  and after around 400 batches the notebook throws memory error on ram and restarts. Is this due to some bottleneck in my code or I need to reduce the batch size?</p>",
      "rawMarkdown": "I have been experenting with the 384x384 centre cropped images posted by Chris Deotte on GPU but it the training time on gpu came out to be very slow and so decided to try on tpus. Now basically this is the first time am trying to use tpus on pytorch and if anyone has used the 386 images what's the average training time per epoch and batch size?\nI was trying batch size of 16  and after around 400 batches the notebook throws memory error on ram and restarts. Is this due to some bottleneck in my code or I need to reduce the batch size?",
      "votes": null
    },
    {
      "id": "935670",
      "postDate": "07/19/2020 14:56:30",
      "content": "<p><a href=\"/sayakdasgupta\">@sayakdasgupta</a>, do you mean 384x384? I have not seen Chris having dataset with 386 resolution for this competition.</p>",
      "rawMarkdown": "sayakdasgupta, do you mean 384x384? I have not seen Chris having dataset with 386 resolution for this competition.",
      "votes": null
    },
    {
      "id": "935673",
      "postDate": "07/19/2020 14:59:06",
      "content": "<p>yes sorry. I will  edit.</p>",
      "rawMarkdown": "yes sorry. I will  edit.",
      "votes": null
    },
    {
      "id": "935677",
      "postDate": "07/19/2020 15:02:19",
      "content": "<p>It could be crashing because your model is too big.\nI am using the batch size of 8*TPU_CORES on EfficientNetB7 taking 6 minutes per an epoch. (image size: 384x384)</p>",
      "rawMarkdown": "It could be crashing because your model is too big.\nI am using the batch size of 8*TPU_CORES on EfficientNetB7 taking 6 minutes per an epoch. (image size: 384x384)",
      "votes": null
    },
    {
      "id": "935679",
      "postDate": "07/19/2020 15:04:13",
      "content": "<p>I was experimenting on B4 only with parallel loader on 8 cores. Single core does fine no issue I have tested my code,moving to 8 cores is throwing the error.</p>",
      "rawMarkdown": "I was experimenting on B4 only with parallel loader on 8 cores. Single core does fine no issue I have tested my code,moving to 8 cores is throwing the error.",
      "votes": null
    },
    {
      "id": "935687",
      "postDate": "07/19/2020 15:11:10",
      "content": "<p>Not sure. I guess there could be something wrong in a pipeline.</p>",
      "rawMarkdown": "Not sure. I guess there could be something wrong in a pipeline.",
      "votes": null
    },
    {
      "id": "935697",
      "postDate": "07/19/2020 15:13:15",
      "content": "<p>You may need to use Colab if you want to run Pytorch on TPU. </p>\n\n<p>Kaggle TPU config is ideal for TF but not for Pytorch. </p>",
      "rawMarkdown": "You may need to use Colab if you want to run Pytorch on TPU. \n\nKaggle TPU config is ideal for TF but not for Pytorch.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 935670,
      "author_name": "sheriytm",
      "author_url": "",
      "post_date": "07/19/2020 14:56:30",
      "content": "<p><a href=\"/sayakdasgupta\">@sayakdasgupta</a>, do you mean 384x384? I have not seen Chris having dataset with 386 resolution for this competition.</p>",
      "votes": null,
      "replies": [
        {
          "id": 935673,
          "author_name": "sayakdasgupta",
          "author_url": "",
          "post_date": "07/19/2020 14:59:06",
          "content": "<p>yes sorry. I will  edit.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 935677,
      "author_name": "bayartsogtya",
      "author_url": "",
      "post_date": "07/19/2020 15:02:19",
      "content": "<p>It could be crashing because your model is too big.\nI am using the batch size of 8*TPU_CORES on EfficientNetB7 taking 6 minutes per an epoch. (image size: 384x384)</p>",
      "votes": null,
      "replies": [
        {
          "id": 935679,
          "author_name": "sayakdasgupta",
          "author_url": "",
          "post_date": "07/19/2020 15:04:13",
          "content": "<p>I was experimenting on B4 only with parallel loader on 8 cores. Single core does fine no issue I have tested my code,moving to 8 cores is throwing the error.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 935687,
          "author_name": "bayartsogtya",
          "author_url": "",
          "post_date": "07/19/2020 15:11:10",
          "content": "<p>Not sure. I guess there could be something wrong in a pipeline.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 935697,
      "author_name": "serigne",
      "author_url": "",
      "post_date": "07/19/2020 15:13:15",
      "content": "<p>You may need to use Colab if you want to run Pytorch on TPU. </p>\n\n<p>Kaggle TPU config is ideal for TF but not for Pytorch. </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "935661": "I have been experenting with the 384x384 centre cropped images posted by Chris Deotte on GPU but it the training time on gpu came out to be very slow and so decided to try on tpus. Now basically this is the first time am trying to use tpus on pytorch and if anyone has used the 386 images what's the average training time per epoch and batch size?\nI was trying batch size of 16  and after around 400 batches the notebook throws memory error on ram and restarts. Is this due to some bottleneck in my code or I need to reduce the batch size?",
    "935670": "sayakdasgupta, do you mean 384x384? I have not seen Chris having dataset with 386 resolution for this competition.",
    "935673": "yes sorry. I will  edit.",
    "935677": "It could be crashing because your model is too big.\nI am using the batch size of 8*TPU_CORES on EfficientNetB7 taking 6 minutes per an epoch. (image size: 384x384)",
    "935679": "I was experimenting on B4 only with parallel loader on 8 cores. Single core does fine no issue I have tested my code,moving to 8 cores is throwing the error.",
    "935687": "Not sure. I guess there could be something wrong in a pipeline.",
    "935697": "You may need to use Colab if you want to run Pytorch on TPU. \n\nKaggle TPU config is ideal for TF but not for Pytorch."
  },
  "source": "meta"
}