{
  "id": 262354,
  "title": "Speed up computational time for training my model",
  "url": "/competitions/g2net-gravitational-wave-detection/discussion/262354",
  "author_name": "",
  "post_date": "2021-08-06T12:34:55.193033600Z",
  "votes": null,
  "comment_count": 11,
  "views": 0,
  "content": "<p>Hello everyone, I am a beginner in the field of Data Science, I want to train a very simple CNN model on the given training set (dim=448000). However, while training the dataset my code is not using  GPU. I have selected the Accelerator to GPU and I have 29 hrs of usage CPU time left. In the following code section, I have created the custom DataGenerator and also the simple CNN model which I have written,</p>\n<pre><code>class DataGenerator(Sequence):\n\n    def __init__(self, list_IDs, data, batch_size=32, dim=(3,4096)):\n        'Initialization'\n        self.list_IDs = list_IDs\n        self.data=data\n        self.batch_size=batch_size\n        self.dim=dim\n        self.indexes = np.arange(len(self.list_IDs))\n\n    def __len__(self):\n        'Denotes the number of batches per epoch'\n        return int(np.floor(len(self.list_IDs) / self.batch_size))\n\n    def __getitem__(self, index):\n        'Generate one batch of data'\n        # Generate indexes of the batch\n        indexes = self.indexes[index*self.batch_size:\\\n                                               (index+1)*self.batch_size]\n\n        # Find list of IDs\n        list_IDs_temp = [self.list_IDs[k] for k in indexes]\n\n        # Generate data\n        X = self.__generateX(list_IDs_temp)\n        y = self.__generatey(list_IDs_temp)\n\n        return X, y\n\n    def __generateX(self,list_temp):\n        'Generate the X (time series data)'\n        X = np.zeros((self.batch_size, 3, 4096))\n        i=0\n        for elem in list_temp:\n            # path for X data set for the particular element\n            path=self.data[self.data['id']==elem]['path'].values[0]\n            X[i, ]=np.load(path)\n            i=i+1\n\n        return X\n\n    def __generatey(self, list_temp):\n        'Generate the y (target values)'\n        y=np.zeros((self.batch_size,1))\n        i=0\n        for elem in list_temp:\n            y[i, ]=self.data[self.data['id']==elem]['target'].values[0]\n            i=i+1\n\n        return y       \n\nmodel = Sequential()\nmodel.add(Conv1D(64, input_shape=(3, 4096,), kernel_size=3, activation='relu'))\nmodel.add(BatchNormalization())\nmodel.add(Flatten())\nmodel.add(Dense(64, activation='relu'))\nmodel.add(Dense(1, activation='sigmoid'))\n\ntrain_generator=DataGenerator(list(X_train_id),df,batch_size=256)\nval_generator=DataGenerator(list(X_val_id),df,batch_size=256)\n</code></pre>\n<p><strong>df</strong> variable is the data frame, which has the following columns: <strong>'id'</strong>, <strong>'path'</strong>, <strong>'target'</strong>.<br>\nCould please provide me the right direction on how can I use the GPU power and is there any problem with my code?</p>",
  "messages": [
    {
      "id": "1455053",
      "postDate": "08/06/2021 12:34:55",
      "content": "<p>Hello everyone, I am a beginner in the field of Data Science, I want to train a very simple CNN model on the given training set (dim=448000). However, while training the dataset my code is not using  GPU. I have selected the Accelerator to GPU and I have 29 hrs of usage CPU time left. In the following code section, I have created the custom DataGenerator and also the simple CNN model which I have written,</p>\n<pre><code>class DataGenerator(Sequence):\n\n    def __init__(self, list_IDs, data, batch_size=32, dim=(3,4096)):\n        'Initialization'\n        self.list_IDs = list_IDs\n        self.data=data\n        self.batch_size=batch_size\n        self.dim=dim\n        self.indexes = np.arange(len(self.list_IDs))\n\n    def __len__(self):\n        'Denotes the number of batches per epoch'\n        return int(np.floor(len(self.list_IDs) / self.batch_size))\n\n    def __getitem__(self, index):\n        'Generate one batch of data'\n        # Generate indexes of the batch\n        indexes = self.indexes[index*self.batch_size:\\\n                                               (index+1)*self.batch_size]\n\n        # Find list of IDs\n        list_IDs_temp = [self.list_IDs[k] for k in indexes]\n\n        # Generate data\n        X = self.__generateX(list_IDs_temp)\n        y = self.__generatey(list_IDs_temp)\n\n        return X, y\n\n    def __generateX(self,list_temp):\n        'Generate the X (time series data)'\n        X = np.zeros((self.batch_size, 3, 4096))\n        i=0\n        for elem in list_temp:\n            # path for X data set for the particular element\n            path=self.data[self.data['id']==elem]['path'].values[0]\n            X[i, ]=np.load(path)\n            i=i+1\n\n        return X\n\n    def __generatey(self, list_temp):\n        'Generate the y (target values)'\n        y=np.zeros((self.batch_size,1))\n        i=0\n        for elem in list_temp:\n            y[i, ]=self.data[self.data['id']==elem]['target'].values[0]\n            i=i+1\n\n        return y       \n\nmodel = Sequential()\nmodel.add(Conv1D(64, input_shape=(3, 4096,), kernel_size=3, activation='relu'))\nmodel.add(BatchNormalization())\nmodel.add(Flatten())\nmodel.add(Dense(64, activation='relu'))\nmodel.add(Dense(1, activation='sigmoid'))\n\ntrain_generator=DataGenerator(list(X_train_id),df,batch_size=256)\nval_generator=DataGenerator(list(X_val_id),df,batch_size=256)\n</code></pre>\n<p><strong>df</strong> variable is the data frame, which has the following columns: <strong>'id'</strong>, <strong>'path'</strong>, <strong>'target'</strong>.<br>\nCould please provide me the right direction on how can I use the GPU power and is there any problem with my code?</p>",
      "rawMarkdown": "Hello everyone, I am a beginner in the field of Data Science, I want to train a very simple CNN model on the given training set (dim=448000). However, while training the dataset my code is not using  GPU. I have selected the Accelerator to GPU and I have 29 hrs of usage CPU time left. In the following code section, I have created the custom DataGenerator and also the simple CNN model which I have written,\n\n```\nclass DataGenerator(Sequence):\n\n    def __init__(self, list_IDs, data, batch_size=32, dim=(3,4096)):\n        'Initialization'\n        self.list_IDs = list_IDs\n        self.data=data\n        self.batch_size=batch_size\n        self.dim=dim\n        self.indexes = np.arange(len(self.list_IDs))\n\n    def __len__(self):\n        'Denotes the number of batches per epoch'\n        return int(np.floor(len(self.list_IDs) / self.batch_size))\n    \n    def __getitem__(self, index):\n        'Generate one batch of data'\n        # Generate indexes of the batch\n        indexes = self.indexes[index*self.batch_size:\\\n                                               (index+1)*self.batch_size]\n\n        # Find list of IDs\n        list_IDs_temp = [self.list_IDs[k] for k in indexes]\n\n        # Generate data\n        X = self.__generateX(list_IDs_temp)\n        y = self.__generatey(list_IDs_temp)\n        \n        return X, y\n    \n    def __generateX(self,list_temp):\n        'Generate the X (time series data)'\n        X = np.zeros((self.batch_size, 3, 4096))\n        i=0\n        for elem in list_temp:\n            # path for X data set for the particular element\n            path=self.data[self.data['id']==elem]['path'].values[0]\n            X[i, ]=np.load(path)\n            i=i+1\n            \n        return X\n    \n    def __generatey(self, list_temp):\n        'Generate the y (target values)'\n        y=np.zeros((self.batch_size,1))\n        i=0\n        for elem in list_temp:\n            y[i, ]=self.data[self.data['id']==elem]['target'].values[0]\n            i=i+1\n            \n        return y       \n\nmodel = Sequential()\nmodel.add(Conv1D(64, input_shape=(3, 4096,), kernel_size=3, activation='relu'))\nmodel.add(BatchNormalization())\nmodel.add(Flatten())\nmodel.add(Dense(64, activation='relu'))\nmodel.add(Dense(1, activation='sigmoid'))\n\ntrain_generator=DataGenerator(list(X_train_id),df,batch_size=256)\nval_generator=DataGenerator(list(X_val_id),df,batch_size=256)\n\n```\n\n\n**df** variable is the data frame, which has the following columns: **'id'**, **'path'**, **'target'**.\nCould please provide me the right direction on how can I use the GPU power and is there any problem with my code?",
      "votes": null
    },
    {
      "id": "1455300",
      "postDate": "08/06/2021 14:17:14",
      "content": "<p>Can you format your post so that the code is readable?  If you don't then you may get way less answers.</p>\n<p>Your questions seems to be why keras does not use the GPU.  But how do you now the GPU is not used in the first place?</p>",
      "rawMarkdown": "Can you format your post so that the code is readable?  If you don't then you may get way less answers.\n\nYour questions seems to be why keras does not use the GPU.  But how do you now the GPU is not used in the first place?",
      "votes": null
    },
    {
      "id": "1456239",
      "postDate": "08/06/2021 20:37:56",
      "content": "<p>Looks like you're using Tensorflow. It utilizes GPU automatically if available. The problem in your case is that your Dataset class is numpy-based, meaning that it will use CPU to read the data. This drastically slows down the training - data preparation for batch takes more time than batch training. You might want to use TFRecords - this will speed up data preparation, hence decreasing a training time.</p>",
      "rawMarkdown": "Looks like you're using Tensorflow. It utilizes GPU automatically if available. The problem in your case is that your Dataset class is numpy-based, meaning that it will use CPU to read the data. This drastically slows down the training - data preparation for batch takes more time than batch training. You might want to use TFRecords - this will speed up data preparation, hence decreasing a training time.",
      "votes": null
    },
    {
      "id": "1457964",
      "postDate": "08/07/2021 16:40:18",
      "content": "<p>I have tried to create the .tfrecord file, but the file size exceeded 19.5 GB therefore I was not able to store the file in the <strong>/kaggle/working</strong> folder. Is there any other way to circumvent this issue?</p>",
      "rawMarkdown": "I have tried to create the .tfrecord file, but the file size exceeded 19.5 GB therefore I was not able to store the file in the **/kaggle/working** folder. Is there any other way to circumvent this issue?",
      "votes": null
    },
    {
      "id": "1458063",
      "postDate": "08/07/2021 17:28:25",
      "content": "<p>As suggested by you, I edited my comment and made the code more readable. Thank you for the advice😄, and I will definitely keep it in my mind for the future. In respect to the question,<br>\n <strong>But how do you now the GPU is not used in the first place?</strong><br>\nI was closely monitoring <strong><em>Draft session</em></strong> of the Kaggle notebook which showed GPU percentage usage to be equal to zero, and very rarely showed 1%. For the above code, ETA for each epoch showed is around 18:07:44, which is a long time. Could please suggest the way forward.</p>",
      "rawMarkdown": "As suggested by you, I edited my comment and made the code more readable. Thank you for the advice😄, and I will definitely keep it in my mind for the future. In respect to the question,\n **But how do you now the GPU is not used in the first place?**\nI was closely monitoring ***Draft session*** of the Kaggle notebook which showed GPU percentage usage to be equal to zero, and very rarely showed 1%. For the above code, ETA for each epoch showed is around 18:07:44, which is a long time. Could please suggest the way forward.",
      "votes": null
    },
    {
      "id": "1459212",
      "postDate": "08/08/2021 09:20:23",
      "content": "<p>Don't put all images in a single TFRecord. Create multiple smaller TFRecords.</p>",
      "rawMarkdown": "Don't put all images in a single TFRecord. Create multiple smaller TFRecords.",
      "votes": null
    },
    {
      "id": "1459296",
      "postDate": "08/08/2021 10:08:44",
      "content": "<p>Thanks, <a href=\"https://www.kaggle.com/atamazian\" target=\"_blank\">@atamazian</a> for your suggestion because of it I was able to find a neat solution to my problem in the stack overflow. I have attached the link below,<br>\n <a href=\"https://stackoverflow.com/questions/48889482/feeding-npy-numpy-files-into-tensorflow-data-pipeline\" target=\"_blank\">https://stackoverflow.com/questions/48889482/feeding-npy-numpy-files-into-tensorflow-data-pipeline</a></p>",
      "rawMarkdown": "Thanks, @atamazian for your suggestion because of it I was able to find a neat solution to my problem in the stack overflow. I have attached the link below,\n https://stackoverflow.com/questions/48889482/feeding-npy-numpy-files-into-tensorflow-data-pipeline",
      "votes": null
    },
    {
      "id": "1459301",
      "postDate": "08/08/2021 10:09:55",
      "content": "<p><a href=\"https://www.kaggle.com/abhishekprajapat\" target=\"_blank\">@abhishekprajapat</a> thank you suggesting another possible approach to the above problem</p>",
      "rawMarkdown": "abhishekprajapat thank you suggesting another possible approach to the above problem",
      "votes": null
    },
    {
      "id": "1459896",
      "postDate": "08/08/2021 15:46:39",
      "content": "<p><a href=\"https://www.kaggle.com/pranay1990\" target=\"_blank\">@pranay1990</a> I faced a similar situation and the solution was to define the model within the strategy scope (see code below). Now the GPU utilization is 100%. That being said, an epoch still takes about 50 min. So I think I will be using the TPU next.</p>\n<pre><code># Detect and init the TPU\ntry: # detect TPUs\n    tpu = tf.distribute.cluster_resolver.TPUClusterResolver.connect() # TPU detection\n    strategy = tf.distribute.TPUStrategy(tpu)\nexcept ValueError: # detect GPUs\n    strategy = tf.distribute.get_strategy() # default strategy that works on CPU and single GPU\nprint(\"REPLICAS: \", strategy.num_replicas_in_sync)\n\nwith strategy.scope():\n    # define model here\n</code></pre>",
      "rawMarkdown": "pranay1990 I faced a similar situation and the solution was to define the model within the strategy scope (see code below). Now the GPU utilization is 100%. That being said, an epoch still takes about 50 min. So I think I will be using the TPU next.\n\n```\n# Detect and init the TPU\ntry: # detect TPUs\n    tpu = tf.distribute.cluster_resolver.TPUClusterResolver.connect() # TPU detection\n    strategy = tf.distribute.TPUStrategy(tpu)\nexcept ValueError: # detect GPUs\n    strategy = tf.distribute.get_strategy() # default strategy that works on CPU and single GPU\nprint(\"REPLICAS: \", strategy.num_replicas_in_sync)\n\nwith strategy.scope():\n    # define model here\n```",
      "votes": null
    },
    {
      "id": "1459940",
      "postDate": "08/08/2021 16:05:52",
      "content": "<p>You can use <code>strategy = tf.distribute.MirroredStrategy()</code> instead of <code>strategy = tf.distribute.get_strategy()</code> if you have a machine with multiple GPUs (not Kaggle case).</p>",
      "rawMarkdown": "You can use ` strategy = tf.distribute.MirroredStrategy()` instead of `strategy = tf.distribute.get_strategy()` if you have a machine with multiple GPUs (not Kaggle case).",
      "votes": null
    },
    {
      "id": "1459946",
      "postDate": "08/08/2021 16:07:05",
      "content": "<p>Here's an <a href=\"https://www.kaggle.com/itsuki9180/g2net-create-image-tfrecords/output\" target=\"_blank\">example notebook </a> of creating multiple small TFRecords.</p>",
      "rawMarkdown": "Here's an [example notebook ](https://www.kaggle.com/itsuki9180/g2net-create-image-tfrecords/output) of creating multiple small TFRecords.",
      "votes": null
    },
    {
      "id": "1561198",
      "postDate": "10/27/2021 12:19:11",
      "content": "<p>Hey All,</p>\n<p>Thank you all for taking part in our competition. The participation has been overwhelmingly positive. We are currently conducting a survey to gauge the demographic and outreach achieved. Kindly spare 2min and fill in this survey <a href=\"https://forms.gle/QP9L16niPexozyhu5\" target=\"_blank\">https://forms.gle/QP9L16niPexozyhu5</a>.</p>\n<p>Thank you all,</p>\n<p>Regards,<br>\nChris</p>",
      "rawMarkdown": "Hey All,\n\nThank you all for taking part in our competition. The participation has been overwhelmingly positive. We are currently conducting a survey to gauge the demographic and outreach achieved. Kindly spare 2min and fill in this survey https://forms.gle/QP9L16niPexozyhu5.\n\nThank you all,\n\nRegards,\nChris",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1455300,
      "author_name": "cpmpml",
      "author_url": "",
      "post_date": "08/06/2021 14:17:14",
      "content": "<p>Can you format your post so that the code is readable?  If you don't then you may get way less answers.</p>\n<p>Your questions seems to be why keras does not use the GPU.  But how do you now the GPU is not used in the first place?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1458063,
          "author_name": "pranay1990",
          "author_url": "",
          "post_date": "08/07/2021 17:28:25",
          "content": "<p>As suggested by you, I edited my comment and made the code more readable. Thank you for the advice😄, and I will definitely keep it in my mind for the future. In respect to the question,<br>\n <strong>But how do you now the GPU is not used in the first place?</strong><br>\nI was closely monitoring <strong><em>Draft session</em></strong> of the Kaggle notebook which showed GPU percentage usage to be equal to zero, and very rarely showed 1%. For the above code, ETA for each epoch showed is around 18:07:44, which is a long time. Could please suggest the way forward.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1456239,
      "author_name": "atamazian",
      "author_url": "",
      "post_date": "08/06/2021 20:37:56",
      "content": "<p>Looks like you're using Tensorflow. It utilizes GPU automatically if available. The problem in your case is that your Dataset class is numpy-based, meaning that it will use CPU to read the data. This drastically slows down the training - data preparation for batch takes more time than batch training. You might want to use TFRecords - this will speed up data preparation, hence decreasing a training time.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1457964,
          "author_name": "pranay1990",
          "author_url": "",
          "post_date": "08/07/2021 16:40:18",
          "content": "<p>I have tried to create the .tfrecord file, but the file size exceeded 19.5 GB therefore I was not able to store the file in the <strong>/kaggle/working</strong> folder. Is there any other way to circumvent this issue?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1459212,
          "author_name": "abhishekprajapat",
          "author_url": "",
          "post_date": "08/08/2021 09:20:23",
          "content": "<p>Don't put all images in a single TFRecord. Create multiple smaller TFRecords.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1459296,
          "author_name": "pranay1990",
          "author_url": "",
          "post_date": "08/08/2021 10:08:44",
          "content": "<p>Thanks, <a href=\"https://www.kaggle.com/atamazian\" target=\"_blank\">@atamazian</a> for your suggestion because of it I was able to find a neat solution to my problem in the stack overflow. I have attached the link below,<br>\n <a href=\"https://stackoverflow.com/questions/48889482/feeding-npy-numpy-files-into-tensorflow-data-pipeline\" target=\"_blank\">https://stackoverflow.com/questions/48889482/feeding-npy-numpy-files-into-tensorflow-data-pipeline</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1459301,
          "author_name": "pranay1990",
          "author_url": "",
          "post_date": "08/08/2021 10:09:55",
          "content": "<p><a href=\"https://www.kaggle.com/abhishekprajapat\" target=\"_blank\">@abhishekprajapat</a> thank you suggesting another possible approach to the above problem</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1459946,
          "author_name": "atamazian",
          "author_url": "",
          "post_date": "08/08/2021 16:07:05",
          "content": "<p>Here's an <a href=\"https://www.kaggle.com/itsuki9180/g2net-create-image-tfrecords/output\" target=\"_blank\">example notebook </a> of creating multiple small TFRecords.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1459896,
      "author_name": "arnabsaha",
      "author_url": "",
      "post_date": "08/08/2021 15:46:39",
      "content": "<p><a href=\"https://www.kaggle.com/pranay1990\" target=\"_blank\">@pranay1990</a> I faced a similar situation and the solution was to define the model within the strategy scope (see code below). Now the GPU utilization is 100%. That being said, an epoch still takes about 50 min. So I think I will be using the TPU next.</p>\n<pre><code># Detect and init the TPU\ntry: # detect TPUs\n    tpu = tf.distribute.cluster_resolver.TPUClusterResolver.connect() # TPU detection\n    strategy = tf.distribute.TPUStrategy(tpu)\nexcept ValueError: # detect GPUs\n    strategy = tf.distribute.get_strategy() # default strategy that works on CPU and single GPU\nprint(\"REPLICAS: \", strategy.num_replicas_in_sync)\n\nwith strategy.scope():\n    # define model here\n</code></pre>",
      "votes": null,
      "replies": [
        {
          "id": 1459940,
          "author_name": "atamazian",
          "author_url": "",
          "post_date": "08/08/2021 16:05:52",
          "content": "<p>You can use <code>strategy = tf.distribute.MirroredStrategy()</code> instead of <code>strategy = tf.distribute.get_strategy()</code> if you have a machine with multiple GPUs (not Kaggle case).</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1561198,
      "author_name": "zerafachris",
      "author_url": "",
      "post_date": "10/27/2021 12:19:11",
      "content": "<p>Hey All,</p>\n<p>Thank you all for taking part in our competition. The participation has been overwhelmingly positive. We are currently conducting a survey to gauge the demographic and outreach achieved. Kindly spare 2min and fill in this survey <a href=\"https://forms.gle/QP9L16niPexozyhu5\" target=\"_blank\">https://forms.gle/QP9L16niPexozyhu5</a>.</p>\n<p>Thank you all,</p>\n<p>Regards,<br>\nChris</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1455053": "Hello everyone, I am a beginner in the field of Data Science, I want to train a very simple CNN model on the given training set (dim=448000). However, while training the dataset my code is not using  GPU. I have selected the Accelerator to GPU and I have 29 hrs of usage CPU time left. In the following code section, I have created the custom DataGenerator and also the simple CNN model which I have written,\n\n```\nclass DataGenerator(Sequence):\n\n    def __init__(self, list_IDs, data, batch_size=32, dim=(3,4096)):\n        'Initialization'\n        self.list_IDs = list_IDs\n        self.data=data\n        self.batch_size=batch_size\n        self.dim=dim\n        self.indexes = np.arange(len(self.list_IDs))\n\n    def __len__(self):\n        'Denotes the number of batches per epoch'\n        return int(np.floor(len(self.list_IDs) / self.batch_size))\n    \n    def __getitem__(self, index):\n        'Generate one batch of data'\n        # Generate indexes of the batch\n        indexes = self.indexes[index*self.batch_size:\\\n                                               (index+1)*self.batch_size]\n\n        # Find list of IDs\n        list_IDs_temp = [self.list_IDs[k] for k in indexes]\n\n        # Generate data\n        X = self.__generateX(list_IDs_temp)\n        y = self.__generatey(list_IDs_temp)\n        \n        return X, y\n    \n    def __generateX(self,list_temp):\n        'Generate the X (time series data)'\n        X = np.zeros((self.batch_size, 3, 4096))\n        i=0\n        for elem in list_temp:\n            # path for X data set for the particular element\n            path=self.data[self.data['id']==elem]['path'].values[0]\n            X[i, ]=np.load(path)\n            i=i+1\n            \n        return X\n    \n    def __generatey(self, list_temp):\n        'Generate the y (target values)'\n        y=np.zeros((self.batch_size,1))\n        i=0\n        for elem in list_temp:\n            y[i, ]=self.data[self.data['id']==elem]['target'].values[0]\n            i=i+1\n            \n        return y       \n\nmodel = Sequential()\nmodel.add(Conv1D(64, input_shape=(3, 4096,), kernel_size=3, activation='relu'))\nmodel.add(BatchNormalization())\nmodel.add(Flatten())\nmodel.add(Dense(64, activation='relu'))\nmodel.add(Dense(1, activation='sigmoid'))\n\ntrain_generator=DataGenerator(list(X_train_id),df,batch_size=256)\nval_generator=DataGenerator(list(X_val_id),df,batch_size=256)\n\n```\n\n\n**df** variable is the data frame, which has the following columns: **'id'**, **'path'**, **'target'**.\nCould please provide me the right direction on how can I use the GPU power and is there any problem with my code?",
    "1455300": "Can you format your post so that the code is readable?  If you don't then you may get way less answers.\n\nYour questions seems to be why keras does not use the GPU.  But how do you now the GPU is not used in the first place?",
    "1456239": "Looks like you're using Tensorflow. It utilizes GPU automatically if available. The problem in your case is that your Dataset class is numpy-based, meaning that it will use CPU to read the data. This drastically slows down the training - data preparation for batch takes more time than batch training. You might want to use TFRecords - this will speed up data preparation, hence decreasing a training time.",
    "1457964": "I have tried to create the .tfrecord file, but the file size exceeded 19.5 GB therefore I was not able to store the file in the **/kaggle/working** folder. Is there any other way to circumvent this issue?",
    "1458063": "As suggested by you, I edited my comment and made the code more readable. Thank you for the advice😄, and I will definitely keep it in my mind for the future. In respect to the question,\n **But how do you now the GPU is not used in the first place?**\nI was closely monitoring ***Draft session*** of the Kaggle notebook which showed GPU percentage usage to be equal to zero, and very rarely showed 1%. For the above code, ETA for each epoch showed is around 18:07:44, which is a long time. Could please suggest the way forward.",
    "1459212": "Don't put all images in a single TFRecord. Create multiple smaller TFRecords.",
    "1459296": "Thanks, @atamazian for your suggestion because of it I was able to find a neat solution to my problem in the stack overflow. I have attached the link below,\n https://stackoverflow.com/questions/48889482/feeding-npy-numpy-files-into-tensorflow-data-pipeline",
    "1459301": "abhishekprajapat thank you suggesting another possible approach to the above problem",
    "1459896": "pranay1990 I faced a similar situation and the solution was to define the model within the strategy scope (see code below). Now the GPU utilization is 100%. That being said, an epoch still takes about 50 min. So I think I will be using the TPU next.\n\n```\n# Detect and init the TPU\ntry: # detect TPUs\n    tpu = tf.distribute.cluster_resolver.TPUClusterResolver.connect() # TPU detection\n    strategy = tf.distribute.TPUStrategy(tpu)\nexcept ValueError: # detect GPUs\n    strategy = tf.distribute.get_strategy() # default strategy that works on CPU and single GPU\nprint(\"REPLICAS: \", strategy.num_replicas_in_sync)\n\nwith strategy.scope():\n    # define model here\n```",
    "1459940": "You can use ` strategy = tf.distribute.MirroredStrategy()` instead of `strategy = tf.distribute.get_strategy()` if you have a machine with multiple GPUs (not Kaggle case).",
    "1459946": "Here's an [example notebook ](https://www.kaggle.com/itsuki9180/g2net-create-image-tfrecords/output) of creating multiple small TFRecords.",
    "1561198": "Hey All,\n\nThank you all for taking part in our competition. The participation has been overwhelmingly positive. We are currently conducting a survey to gauge the demographic and outreach achieved. Kindly spare 2min and fill in this survey https://forms.gle/QP9L16niPexozyhu5.\n\nThank you all,\n\nRegards,\nChris"
  },
  "source": "meta"
}