{
  "id": 174117,
  "title": "Tips to increase TPU efficiency? ",
  "url": "/competitions/siim-isic-melanoma-classification/discussion/174117",
  "author_name": "",
  "post_date": "2020-08-12T10:26:20.555214100Z",
  "votes": 2,
  "comment_count": 5,
  "views": 0,
  "content": "<p>I am struggling with 3-hour time limit and resource exhaustion.</p>",
  "messages": [
    {
      "id": "967551",
      "postDate": "08/12/2020 10:26:20",
      "content": "<p>I am struggling with 3-hour time limit and resource exhaustion.</p>",
      "rawMarkdown": "I am struggling with 3-hour time limit and resource exhaustion.",
      "votes": null
    },
    {
      "id": "967723",
      "postDate": "08/12/2020 13:07:28",
      "content": "<p>Check out this (<a href=\"https://cloud.google.com/tpu/docs/performance-guide\" target=\"_blank\">https://cloud.google.com/tpu/docs/performance-guide</a>)</p>",
      "rawMarkdown": "Check out this (https://cloud.google.com/tpu/docs/performance-guide)",
      "votes": null
    },
    {
      "id": "967867",
      "postDate": "08/12/2020 14:47:32",
      "content": "<p>I have the same issue, i'm trying to train following Model:</p>\n<p>Model: EfficientNet-B4<br>\nImage-Size: 384x384<br>\nFolds: 5-Folds<br>\nNo external-data<br>\nDownsampled Dataset (512x512-JPEG)<br>\nEpoch (atleast what I would like): 10</p>\n<p>But I have no chance with the 3 hours restriction..</p>\n<p>Did you train such a model already within 3 hours?</p>",
      "rawMarkdown": "I have the same issue, i'm trying to train following Model:\n\nModel: EfficientNet-B4\nImage-Size: 384x384\nFolds: 5-Folds\nNo external-data\nDownsampled Dataset (512x512-JPEG)\nEpoch (atleast what I would like): 10\n\nBut I have no chance with the 3 hours restriction..\n\nDid you train such a model already within 3 hours?",
      "votes": null
    },
    {
      "id": "968213",
      "postDate": "08/12/2020 19:18:59",
      "content": "<p>I have actually put in a fail safe in case I am exceeding the 3 hour limit by putting a timer and halting my N-fold cross validation. </p>",
      "rawMarkdown": "I have actually put in a fail safe in case I am exceeding the 3 hour limit by putting a timer and halting my N-fold cross validation.",
      "votes": null
    },
    {
      "id": "969081",
      "postDate": "08/13/2020 12:53:13",
      "content": "<p><a href=\"https://www.kaggle.com/aliabdin1\" target=\"_blank\">@aliabdin1</a> how big it's your <code>batch size</code> ? When you use TPU your <code>batch size</code> must be bigger than when you use GPU or the TPU it's gonna \"starve\".</p>",
      "rawMarkdown": "aliabdin1 how big it's your `batch size` ? When you use TPU your `batch size` must be bigger than when you use GPU or the TPU it's gonna \"starve\".",
      "votes": null
    },
    {
      "id": "969095",
      "postDate": "08/13/2020 13:03:08",
      "content": "<p>I use 8 cores, each core gets a batch of size 5 -&gt; 40 images at once, the memory does not allow more or it canceles the training.</p>",
      "rawMarkdown": "I use 8 cores, each core gets a batch of size 5 -> 40 images at once, the memory does not allow more or it canceles the training.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 967723,
      "author_name": "msharuk589",
      "author_url": "",
      "post_date": "08/12/2020 13:07:28",
      "content": "<p>Check out this (<a href=\"https://cloud.google.com/tpu/docs/performance-guide\" target=\"_blank\">https://cloud.google.com/tpu/docs/performance-guide</a>)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 967867,
      "author_name": "aliabdin1",
      "author_url": "",
      "post_date": "08/12/2020 14:47:32",
      "content": "<p>I have the same issue, i'm trying to train following Model:</p>\n<p>Model: EfficientNet-B4<br>\nImage-Size: 384x384<br>\nFolds: 5-Folds<br>\nNo external-data<br>\nDownsampled Dataset (512x512-JPEG)<br>\nEpoch (atleast what I would like): 10</p>\n<p>But I have no chance with the 3 hours restriction..</p>\n<p>Did you train such a model already within 3 hours?</p>",
      "votes": null,
      "replies": [
        {
          "id": 969081,
          "author_name": "hiramcho",
          "author_url": "",
          "post_date": "08/13/2020 12:53:13",
          "content": "<p><a href=\"https://www.kaggle.com/aliabdin1\" target=\"_blank\">@aliabdin1</a> how big it's your <code>batch size</code> ? When you use TPU your <code>batch size</code> must be bigger than when you use GPU or the TPU it's gonna \"starve\".</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 969095,
          "author_name": "aliabdin1",
          "author_url": "",
          "post_date": "08/13/2020 13:03:08",
          "content": "<p>I use 8 cores, each core gets a batch of size 5 -&gt; 40 images at once, the memory does not allow more or it canceles the training.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 968213,
      "author_name": "realsid",
      "author_url": "",
      "post_date": "08/12/2020 19:18:59",
      "content": "<p>I have actually put in a fail safe in case I am exceeding the 3 hour limit by putting a timer and halting my N-fold cross validation. </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "967551": "I am struggling with 3-hour time limit and resource exhaustion.",
    "967723": "Check out this (https://cloud.google.com/tpu/docs/performance-guide)",
    "967867": "I have the same issue, i'm trying to train following Model:\n\nModel: EfficientNet-B4\nImage-Size: 384x384\nFolds: 5-Folds\nNo external-data\nDownsampled Dataset (512x512-JPEG)\nEpoch (atleast what I would like): 10\n\nBut I have no chance with the 3 hours restriction..\n\nDid you train such a model already within 3 hours?",
    "968213": "I have actually put in a fail safe in case I am exceeding the 3 hour limit by putting a timer and halting my N-fold cross validation.",
    "969081": "aliabdin1 how big it's your `batch size` ? When you use TPU your `batch size` must be bigger than when you use GPU or the TPU it's gonna \"starve\".",
    "969095": "I use 8 cores, each core gets a batch of size 5 -> 40 images at once, the memory does not allow more or it canceles the training."
  },
  "source": "meta"
}