{
  "id": 123766,
  "title": "Requesting effective ways to reduce GPU usage to <=2 hrs",
  "url": "/competitions/bengaliai-cv19/discussion/123766",
  "author_name": "",
  "post_date": "2019-12-30T09:03:43.711019600Z",
  "votes": 1,
  "comment_count": 6,
  "views": 0,
  "content": "<p>I am facing issues regarding GPU usage. If I increase number of epochs or reduce batch size or use data augmentation, GPU usage shoots above 2hrs.</p>",
  "messages": [
    {
      "id": "706337",
      "postDate": "12/30/2019 09:03:43",
      "content": "<p>I am facing issues regarding GPU usage. If I increase number of epochs or reduce batch size or use data augmentation, GPU usage shoots above 2hrs.</p>",
      "rawMarkdown": "I am facing issues regarding GPU usage. If I increase number of epochs or reduce batch size or use data augmentation, GPU usage shoots above 2hrs.",
      "votes": null
    },
    {
      "id": "706342",
      "postDate": "12/30/2019 09:05:18",
      "content": "<p>Train model in one kernel and make inference in a different kernel.</p>",
      "rawMarkdown": "Train model in one kernel and make inference in a different kernel.",
      "votes": null
    },
    {
      "id": "706389",
      "postDate": "12/30/2019 10:14:27",
      "content": "<p>A few ideas on how to reduce the training time:\n- Reduce the size of the dataset: Not necessarily reducing the number of images, but the images themselves. A very basic approach could be to simply resize them to a lower resolution.\n- Use a shallower model. The deeper the model, the longer it is going to train (roughly).\n- Also do things in different notebooks, as <a href=\"/artgor\">@artgor</a> mentionned: Do your EDA in one noteboook that doesn't need a GPU. If you do pre-processing, you can do it in another notebook and save the new images as a dataset. Then do only the training in another notebook, and save either the weights or predictions, or even both; that will allow you to then make predictions in yet another notebook.</p>\n\n<p>Hope this might help!</p>",
      "rawMarkdown": "A few ideas on how to reduce the training time:\n- Reduce the size of the dataset: Not necessarily reducing the number of images, but the images themselves. A very basic approach could be to simply resize them to a lower resolution.\n- Use a shallower model. The deeper the model, the longer it is going to train (roughly).\n- Also do things in different notebooks, as @artgor mentionned: Do your EDA in one noteboook that doesn't need a GPU. If you do pre-processing, you can do it in another notebook and save the new images as a dataset. Then do only the training in another notebook, and save either the weights or predictions, or even both; that will allow you to then make predictions in yet another notebook.\n\nHope this might help!",
      "votes": null
    },
    {
      "id": "706475",
      "postDate": "12/30/2019 12:47:23",
      "content": "<p>Great idea !\nThanks <a href=\"/artgor\">@artgor</a> </p>",
      "rawMarkdown": "Great idea !\nThanks @artgor",
      "votes": null
    },
    {
      "id": "706476",
      "postDate": "12/30/2019 12:48:18",
      "content": "<p>Yupp gotta try these now !\nThanks for your help 😊 </p>",
      "rawMarkdown": "Yupp gotta try these now !\nThanks for your help 😊",
      "votes": null
    },
    {
      "id": "707111",
      "postDate": "12/31/2019 09:04:12",
      "content": "<p>Good idea man! :)</p>",
      "rawMarkdown": "Good idea man! :)",
      "votes": null
    },
    {
      "id": "770254",
      "postDate": "03/12/2020 18:33:49",
      "content": "<p>The most direct method is to simply write a tensorflow callback to stop training before the timer\n<code>python3\n   history = model.fit(\n        dataset.X[\"train\"], dataset.Y[\"train\"],\n        batch_size=128,\n        epochs=999,\n        verbose=False,\n        validation_data=(dataset.X[\"valid\"], dataset.Y[\"valid\"]),\n        callbacks=[\n            EarlyStopping(),\n            ModelCheckpoint(),\n            KaggleTimeoutCallback( \"115m\", verbose=True ),\n        ]\n    )\n</code></p>\n\n<p>```python3\nimport math\nimport re\nimport time\nfrom typing import Union</p>\n\n<p>import tensorflow as tf</p>\n\n<p>class KaggleTimeoutCallback(tf.keras.callbacks.Callback):\n    start_python = time.time()</p>\n\n<pre><code>def __init__(self, timeout: Union[int, float, str], from_now=False, verbose=False):\n    super().__init__()\n    self.verbose           = verbose\n    self.from_now          = from_now\n    self.start_time        = self.start_python if not self.from_now else time.time()\n    self.timeout_seconds   = self.parse_seconds(timeout)\n\n    self.last_epoch_start  = time.time()\n    self.last_epoch_end    = time.time()\n    self.last_epoch_time   = self.last_epoch_end - self.last_epoch_start\n    self.current_runtime   = self.last_epoch_end - self.start_time\n\n\ndef on_train_begin(self, logs=None):\n    self.check_timeout()  # timeout before first epoch if model.fit() is called again\n\n\ndef on_epoch_begin(self, epoch, logs=None):\n    self.last_epoch_start = time.time()\n\n\ndef on_epoch_end(self, epoch, logs=None):\n    self.last_epoch_end  = time.time()\n    self.last_epoch_time = self.last_epoch_end - self.last_epoch_start\n    self.check_timeout()\n\n\ndef check_timeout(self):\n    self.current_runtime = self.last_epoch_end - self.start_time\n    if self.verbose:\n        print(f'\\nKaggleTimeoutCallback({self.format(self.timeout_seconds)}) runtime {self.format(self.current_runtime)}')\n\n    # Give timeout leeway of 2 * last_epoch_time\n    if (self.current_runtime + self.last_epoch_time*2) &amp;gt;= self.timeout_seconds:\n        print(f\"\\nKaggleTimeoutCallback({self.format(self.timeout_seconds)}) stopped after {self.format(self.current_runtime)}\")\n        self.model.stop_training = True\n\n\n@staticmethod\ndef parse_seconds(timeout) -&amp;gt; int:\n    if isinstance(timeout, (float,int)): return int(timeout)\n    seconds = 0\n    for (number, unit) in re.findall(r\"(\\d+(?:\\.\\d+)?)\\s*([dhms])?\", str(timeout)):\n        if   unit == 'd': seconds += float(number) * 60 * 60 * 24\n        elif unit == 'h': seconds += float(number) * 60 * 60\n        elif unit == 'm': seconds += float(number) * 60\n        else:             seconds += float(number)\n    return int(seconds)\n\n\n@staticmethod\ndef format(seconds: Union[int,float]) -&amp;gt; str:\n    runtime = {\n        \"d\":   math.floor(seconds / (60*60*24) ),\n        \"h\":   math.floor(seconds / (60*60)    ) % 24,\n        \"m\":   math.floor(seconds / (60)       ) % 60,\n        \"s\":   math.floor(seconds              ) % 60,\n    }\n    return \" \".join([ f\"{runtime[unit]}{unit}\" for unit in [\"d\", \"h\", \"m\", \"s\"] if runtime[unit] != 0 ])\n</code></pre>\n\n<p>```</p>",
      "rawMarkdown": "The most direct method is to simply write a tensorflow callback to stop training before the timer\n```python3\n   history = model.fit(\n        dataset.X[\"train\"], dataset.Y[\"train\"],\n        batch_size=128,\n        epochs=999,\n        verbose=False,\n        validation_data=(dataset.X[\"valid\"], dataset.Y[\"valid\"]),\n        callbacks=[\n            EarlyStopping(),\n            ModelCheckpoint(),\n            KaggleTimeoutCallback( \"115m\", verbose=True ),\n        ]\n    )\n```\n\n```python3\nimport math\nimport re\nimport time\nfrom typing import Union\n\nimport tensorflow as tf\n\n\nclass KaggleTimeoutCallback(tf.keras.callbacks.Callback):\n    start_python = time.time()\n\n\n    def __init__(self, timeout: Union[int, float, str], from_now=False, verbose=False):\n        super().__init__()\n        self.verbose           = verbose\n        self.from_now          = from_now\n        self.start_time        = self.start_python if not self.from_now else time.time()\n        self.timeout_seconds   = self.parse_seconds(timeout)\n\n        self.last_epoch_start  = time.time()\n        self.last_epoch_end    = time.time()\n        self.last_epoch_time   = self.last_epoch_end - self.last_epoch_start\n        self.current_runtime   = self.last_epoch_end - self.start_time\n\n\n    def on_train_begin(self, logs=None):\n        self.check_timeout()  # timeout before first epoch if model.fit() is called again\n\n\n    def on_epoch_begin(self, epoch, logs=None):\n        self.last_epoch_start = time.time()\n\n\n    def on_epoch_end(self, epoch, logs=None):\n        self.last_epoch_end  = time.time()\n        self.last_epoch_time = self.last_epoch_end - self.last_epoch_start\n        self.check_timeout()\n\n\n    def check_timeout(self):\n        self.current_runtime = self.last_epoch_end - self.start_time\n        if self.verbose:\n            print(f'\\nKaggleTimeoutCallback({self.format(self.timeout_seconds)}) runtime {self.format(self.current_runtime)}')\n\n        # Give timeout leeway of 2 * last_epoch_time\n        if (self.current_runtime + self.last_epoch_time*2) &gt;= self.timeout_seconds:\n            print(f\"\\nKaggleTimeoutCallback({self.format(self.timeout_seconds)}) stopped after {self.format(self.current_runtime)}\")\n            self.model.stop_training = True\n\n\n    @staticmethod\n    def parse_seconds(timeout) -&gt; int:\n        if isinstance(timeout, (float,int)): return int(timeout)\n        seconds = 0\n        for (number, unit) in re.findall(r\"(\\d+(?:\\.\\d+)?)\\s*([dhms])?\", str(timeout)):\n            if   unit == 'd': seconds += float(number) * 60 * 60 * 24\n            elif unit == 'h': seconds += float(number) * 60 * 60\n            elif unit == 'm': seconds += float(number) * 60\n            else:             seconds += float(number)\n        return int(seconds)\n\n\n    @staticmethod\n    def format(seconds: Union[int,float]) -&gt; str:\n        runtime = {\n            \"d\":   math.floor(seconds / (60*60*24) ),\n            \"h\":   math.floor(seconds / (60*60)    ) % 24,\n            \"m\":   math.floor(seconds / (60)       ) % 60,\n            \"s\":   math.floor(seconds              ) % 60,\n        }\n        return \" \".join([ f\"{runtime[unit]}{unit}\" for unit in [\"d\", \"h\", \"m\", \"s\"] if runtime[unit] != 0 ])\n```",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 706342,
      "author_name": "artgor",
      "author_url": "",
      "post_date": "12/30/2019 09:05:18",
      "content": "<p>Train model in one kernel and make inference in a different kernel.</p>",
      "votes": null,
      "replies": [
        {
          "id": 706475,
          "author_name": "namanj27",
          "author_url": "",
          "post_date": "12/30/2019 12:47:23",
          "content": "<p>Great idea !\nThanks <a href=\"/artgor\">@artgor</a> </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 707111,
          "author_name": "dasmehdixtr",
          "author_url": "",
          "post_date": "12/31/2019 09:04:12",
          "content": "<p>Good idea man! :)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 706389,
      "author_name": "maxlenormand",
      "author_url": "",
      "post_date": "12/30/2019 10:14:27",
      "content": "<p>A few ideas on how to reduce the training time:\n- Reduce the size of the dataset: Not necessarily reducing the number of images, but the images themselves. A very basic approach could be to simply resize them to a lower resolution.\n- Use a shallower model. The deeper the model, the longer it is going to train (roughly).\n- Also do things in different notebooks, as <a href=\"/artgor\">@artgor</a> mentionned: Do your EDA in one noteboook that doesn't need a GPU. If you do pre-processing, you can do it in another notebook and save the new images as a dataset. Then do only the training in another notebook, and save either the weights or predictions, or even both; that will allow you to then make predictions in yet another notebook.</p>\n\n<p>Hope this might help!</p>",
      "votes": null,
      "replies": [
        {
          "id": 706476,
          "author_name": "namanj27",
          "author_url": "",
          "post_date": "12/30/2019 12:48:18",
          "content": "<p>Yupp gotta try these now !\nThanks for your help 😊 </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 770254,
      "author_name": "jamesmcguigan",
      "author_url": "",
      "post_date": "03/12/2020 18:33:49",
      "content": "<p>The most direct method is to simply write a tensorflow callback to stop training before the timer\n<code>python3\n   history = model.fit(\n        dataset.X[\"train\"], dataset.Y[\"train\"],\n        batch_size=128,\n        epochs=999,\n        verbose=False,\n        validation_data=(dataset.X[\"valid\"], dataset.Y[\"valid\"]),\n        callbacks=[\n            EarlyStopping(),\n            ModelCheckpoint(),\n            KaggleTimeoutCallback( \"115m\", verbose=True ),\n        ]\n    )\n</code></p>\n\n<p>```python3\nimport math\nimport re\nimport time\nfrom typing import Union</p>\n\n<p>import tensorflow as tf</p>\n\n<p>class KaggleTimeoutCallback(tf.keras.callbacks.Callback):\n    start_python = time.time()</p>\n\n<pre><code>def __init__(self, timeout: Union[int, float, str], from_now=False, verbose=False):\n    super().__init__()\n    self.verbose           = verbose\n    self.from_now          = from_now\n    self.start_time        = self.start_python if not self.from_now else time.time()\n    self.timeout_seconds   = self.parse_seconds(timeout)\n\n    self.last_epoch_start  = time.time()\n    self.last_epoch_end    = time.time()\n    self.last_epoch_time   = self.last_epoch_end - self.last_epoch_start\n    self.current_runtime   = self.last_epoch_end - self.start_time\n\n\ndef on_train_begin(self, logs=None):\n    self.check_timeout()  # timeout before first epoch if model.fit() is called again\n\n\ndef on_epoch_begin(self, epoch, logs=None):\n    self.last_epoch_start = time.time()\n\n\ndef on_epoch_end(self, epoch, logs=None):\n    self.last_epoch_end  = time.time()\n    self.last_epoch_time = self.last_epoch_end - self.last_epoch_start\n    self.check_timeout()\n\n\ndef check_timeout(self):\n    self.current_runtime = self.last_epoch_end - self.start_time\n    if self.verbose:\n        print(f'\\nKaggleTimeoutCallback({self.format(self.timeout_seconds)}) runtime {self.format(self.current_runtime)}')\n\n    # Give timeout leeway of 2 * last_epoch_time\n    if (self.current_runtime + self.last_epoch_time*2) &amp;gt;= self.timeout_seconds:\n        print(f\"\\nKaggleTimeoutCallback({self.format(self.timeout_seconds)}) stopped after {self.format(self.current_runtime)}\")\n        self.model.stop_training = True\n\n\n@staticmethod\ndef parse_seconds(timeout) -&amp;gt; int:\n    if isinstance(timeout, (float,int)): return int(timeout)\n    seconds = 0\n    for (number, unit) in re.findall(r\"(\\d+(?:\\.\\d+)?)\\s*([dhms])?\", str(timeout)):\n        if   unit == 'd': seconds += float(number) * 60 * 60 * 24\n        elif unit == 'h': seconds += float(number) * 60 * 60\n        elif unit == 'm': seconds += float(number) * 60\n        else:             seconds += float(number)\n    return int(seconds)\n\n\n@staticmethod\ndef format(seconds: Union[int,float]) -&amp;gt; str:\n    runtime = {\n        \"d\":   math.floor(seconds / (60*60*24) ),\n        \"h\":   math.floor(seconds / (60*60)    ) % 24,\n        \"m\":   math.floor(seconds / (60)       ) % 60,\n        \"s\":   math.floor(seconds              ) % 60,\n    }\n    return \" \".join([ f\"{runtime[unit]}{unit}\" for unit in [\"d\", \"h\", \"m\", \"s\"] if runtime[unit] != 0 ])\n</code></pre>\n\n<p>```</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "706337": "I am facing issues regarding GPU usage. If I increase number of epochs or reduce batch size or use data augmentation, GPU usage shoots above 2hrs.",
    "706342": "Train model in one kernel and make inference in a different kernel.",
    "706389": "A few ideas on how to reduce the training time:\n- Reduce the size of the dataset: Not necessarily reducing the number of images, but the images themselves. A very basic approach could be to simply resize them to a lower resolution.\n- Use a shallower model. The deeper the model, the longer it is going to train (roughly).\n- Also do things in different notebooks, as @artgor mentionned: Do your EDA in one noteboook that doesn't need a GPU. If you do pre-processing, you can do it in another notebook and save the new images as a dataset. Then do only the training in another notebook, and save either the weights or predictions, or even both; that will allow you to then make predictions in yet another notebook.\n\nHope this might help!",
    "706475": "Great idea !\nThanks @artgor",
    "706476": "Yupp gotta try these now !\nThanks for your help 😊",
    "707111": "Good idea man! :)",
    "770254": "The most direct method is to simply write a tensorflow callback to stop training before the timer\n```python3\n   history = model.fit(\n        dataset.X[\"train\"], dataset.Y[\"train\"],\n        batch_size=128,\n        epochs=999,\n        verbose=False,\n        validation_data=(dataset.X[\"valid\"], dataset.Y[\"valid\"]),\n        callbacks=[\n            EarlyStopping(),\n            ModelCheckpoint(),\n            KaggleTimeoutCallback( \"115m\", verbose=True ),\n        ]\n    )\n```\n\n```python3\nimport math\nimport re\nimport time\nfrom typing import Union\n\nimport tensorflow as tf\n\n\nclass KaggleTimeoutCallback(tf.keras.callbacks.Callback):\n    start_python = time.time()\n\n\n    def __init__(self, timeout: Union[int, float, str], from_now=False, verbose=False):\n        super().__init__()\n        self.verbose           = verbose\n        self.from_now          = from_now\n        self.start_time        = self.start_python if not self.from_now else time.time()\n        self.timeout_seconds   = self.parse_seconds(timeout)\n\n        self.last_epoch_start  = time.time()\n        self.last_epoch_end    = time.time()\n        self.last_epoch_time   = self.last_epoch_end - self.last_epoch_start\n        self.current_runtime   = self.last_epoch_end - self.start_time\n\n\n    def on_train_begin(self, logs=None):\n        self.check_timeout()  # timeout before first epoch if model.fit() is called again\n\n\n    def on_epoch_begin(self, epoch, logs=None):\n        self.last_epoch_start = time.time()\n\n\n    def on_epoch_end(self, epoch, logs=None):\n        self.last_epoch_end  = time.time()\n        self.last_epoch_time = self.last_epoch_end - self.last_epoch_start\n        self.check_timeout()\n\n\n    def check_timeout(self):\n        self.current_runtime = self.last_epoch_end - self.start_time\n        if self.verbose:\n            print(f'\\nKaggleTimeoutCallback({self.format(self.timeout_seconds)}) runtime {self.format(self.current_runtime)}')\n\n        # Give timeout leeway of 2 * last_epoch_time\n        if (self.current_runtime + self.last_epoch_time*2) &gt;= self.timeout_seconds:\n            print(f\"\\nKaggleTimeoutCallback({self.format(self.timeout_seconds)}) stopped after {self.format(self.current_runtime)}\")\n            self.model.stop_training = True\n\n\n    @staticmethod\n    def parse_seconds(timeout) -&gt; int:\n        if isinstance(timeout, (float,int)): return int(timeout)\n        seconds = 0\n        for (number, unit) in re.findall(r\"(\\d+(?:\\.\\d+)?)\\s*([dhms])?\", str(timeout)):\n            if   unit == 'd': seconds += float(number) * 60 * 60 * 24\n            elif unit == 'h': seconds += float(number) * 60 * 60\n            elif unit == 'm': seconds += float(number) * 60\n            else:             seconds += float(number)\n        return int(seconds)\n\n\n    @staticmethod\n    def format(seconds: Union[int,float]) -&gt; str:\n        runtime = {\n            \"d\":   math.floor(seconds / (60*60*24) ),\n            \"h\":   math.floor(seconds / (60*60)    ) % 24,\n            \"m\":   math.floor(seconds / (60)       ) % 60,\n            \"s\":   math.floor(seconds              ) % 60,\n        }\n        return \" \".join([ f\"{runtime[unit]}{unit}\" for unit in [\"d\", \"h\", \"m\", \"s\"] if runtime[unit] != 0 ])\n```"
  },
  "source": "meta"
}