{
  "id": 229230,
  "title": "Keras model.fit overloading CPU ",
  "url": "/competitions/plant-pathology-2021-fgvc8/discussion/229230",
  "author_name": "",
  "post_date": "2021-03-29T03:45:10.658746Z",
  "votes": 3,
  "comment_count": 2,
  "views": 0,
  "content": "<p>Can't get past training my data using model.fit with a train generator using flow_from_dataframe. Here's the code I use for training:</p>\n<p>STEP_SIZE_TRAIN=train_generator.n//train_generator.batch_size<br>\nSTEP_SIZE_VALID=validation_generator.n//validation_generator.batch_size</p>\n<p>history = model.fit(train_generator, <br>\n                              epochs = 20, <br>\n                              steps_per_epoch = STEP_SIZE_TRAIN,<br>\n                              validation_data = validation_generator,<br>\n                              validation_steps = STEP_SIZE_VALID,<br>\n                              callbacks=[callbacks], <br>\n                              verbose=1)</p>\n<p>As soon as I execute the code, the output never goes past epoch 1/20 and does not even display the training status bar.  Even though I'm using the GPU accelerator, the GPU usage does not even go up at all when running this code block but the CPU usage jumps up to 99% and once it gets to over 100%, the kernel just stops. </p>\n<p>Here's the link to my notebook for reference: <a href=\"https://www.kaggle.com/chummicrisologo/plant-pathology/edit\" target=\"_blank\">https://www.kaggle.com/chummicrisologo/plant-pathology/edit</a>   </p>\n<p>*excuse the messy code as I've yet to clean up portions from my previous versions</p>",
  "messages": [
    {
      "id": "1255669",
      "postDate": "03/29/2021 03:45:10",
      "content": "<p>Can't get past training my data using model.fit with a train generator using flow_from_dataframe. Here's the code I use for training:</p>\n<p>STEP_SIZE_TRAIN=train_generator.n//train_generator.batch_size<br>\nSTEP_SIZE_VALID=validation_generator.n//validation_generator.batch_size</p>\n<p>history = model.fit(train_generator, <br>\n                              epochs = 20, <br>\n                              steps_per_epoch = STEP_SIZE_TRAIN,<br>\n                              validation_data = validation_generator,<br>\n                              validation_steps = STEP_SIZE_VALID,<br>\n                              callbacks=[callbacks], <br>\n                              verbose=1)</p>\n<p>As soon as I execute the code, the output never goes past epoch 1/20 and does not even display the training status bar.  Even though I'm using the GPU accelerator, the GPU usage does not even go up at all when running this code block but the CPU usage jumps up to 99% and once it gets to over 100%, the kernel just stops. </p>\n<p>Here's the link to my notebook for reference: <a href=\"https://www.kaggle.com/chummicrisologo/plant-pathology/edit\" target=\"_blank\">https://www.kaggle.com/chummicrisologo/plant-pathology/edit</a>   </p>\n<p>*excuse the messy code as I've yet to clean up portions from my previous versions</p>",
      "rawMarkdown": "Can't get past training my data using model.fit with a train generator using flow_from_dataframe. Here's the code I use for training:\n\nSTEP_SIZE_TRAIN=train_generator.n//train_generator.batch_size\nSTEP_SIZE_VALID=validation_generator.n//validation_generator.batch_size\n\nhistory = model.fit(train_generator, \n                              epochs = 20, \n                              steps_per_epoch = STEP_SIZE_TRAIN,\n                              validation_data = validation_generator,\n                              validation_steps = STEP_SIZE_VALID,\n                              callbacks=[callbacks], \n                              verbose=1)\n\nAs soon as I execute the code, the output never goes past epoch 1/20 and does not even display the training status bar.  Even though I'm using the GPU accelerator, the GPU usage does not even go up at all when running this code block but the CPU usage jumps up to 99% and once it gets to over 100%, the kernel just stops. \n\nHere's the link to my notebook for reference: https://www.kaggle.com/chummicrisologo/plant-pathology/edit   \n\n*excuse the messy code as I've yet to clean up portions from my previous versions",
      "votes": null
    },
    {
      "id": "1259451",
      "postDate": "04/01/2021 12:23:30",
      "content": "<p>The reason for 0% gpu utilization is because it is idle most of the time. While the cpu utilization is full. This is because the cpu is busy fetching and preparing the data whereas the gpu remains idle. And when the gpu is busy the cpu is idle but for a short time.This is due to inefficient input pipeline. In order to increase the speed of training process, you need to create an input pipeline such that when the gpu is busy, the cpu is preparing the next batches in advance to avoid or reduce the idle time for gpu. This can be done with the tf.data.Dataset API.</p>",
      "rawMarkdown": "The reason for 0% gpu utilization is because it is idle most of the time. While the cpu utilization is full. This is because the cpu is busy fetching and preparing the data whereas the gpu remains idle. And when the gpu is busy the cpu is idle but for a short time.This is due to inefficient input pipeline. In order to increase the speed of training process, you need to create an input pipeline such that when the gpu is busy, the cpu is preparing the next batches in advance to avoid or reduce the idle time for gpu. This can be done with the tf.data.Dataset API.",
      "votes": null
    },
    {
      "id": "1277562",
      "postDate": "04/18/2021 23:04:31",
      "content": "<p>Sorry for the late response. Will try this out. Thanks brother!</p>",
      "rawMarkdown": "Sorry for the late response. Will try this out. Thanks brother!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1259451,
      "author_name": "mohammadasimbluemoon",
      "author_url": "",
      "post_date": "04/01/2021 12:23:30",
      "content": "<p>The reason for 0% gpu utilization is because it is idle most of the time. While the cpu utilization is full. This is because the cpu is busy fetching and preparing the data whereas the gpu remains idle. And when the gpu is busy the cpu is idle but for a short time.This is due to inefficient input pipeline. In order to increase the speed of training process, you need to create an input pipeline such that when the gpu is busy, the cpu is preparing the next batches in advance to avoid or reduce the idle time for gpu. This can be done with the tf.data.Dataset API.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1277562,
          "author_name": "chummicrisologo",
          "author_url": "",
          "post_date": "04/18/2021 23:04:31",
          "content": "<p>Sorry for the late response. Will try this out. Thanks brother!</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1255669": "Can't get past training my data using model.fit with a train generator using flow_from_dataframe. Here's the code I use for training:\n\nSTEP_SIZE_TRAIN=train_generator.n//train_generator.batch_size\nSTEP_SIZE_VALID=validation_generator.n//validation_generator.batch_size\n\nhistory = model.fit(train_generator, \n                              epochs = 20, \n                              steps_per_epoch = STEP_SIZE_TRAIN,\n                              validation_data = validation_generator,\n                              validation_steps = STEP_SIZE_VALID,\n                              callbacks=[callbacks], \n                              verbose=1)\n\nAs soon as I execute the code, the output never goes past epoch 1/20 and does not even display the training status bar.  Even though I'm using the GPU accelerator, the GPU usage does not even go up at all when running this code block but the CPU usage jumps up to 99% and once it gets to over 100%, the kernel just stops. \n\nHere's the link to my notebook for reference: https://www.kaggle.com/chummicrisologo/plant-pathology/edit   \n\n*excuse the messy code as I've yet to clean up portions from my previous versions",
    "1259451": "The reason for 0% gpu utilization is because it is idle most of the time. While the cpu utilization is full. This is because the cpu is busy fetching and preparing the data whereas the gpu remains idle. And when the gpu is busy the cpu is idle but for a short time.This is due to inefficient input pipeline. In order to increase the speed of training process, you need to create an input pipeline such that when the gpu is busy, the cpu is preparing the next batches in advance to avoid or reduce the idle time for gpu. This can be done with the tf.data.Dataset API.",
    "1277562": "Sorry for the late response. Will try this out. Thanks brother!"
  },
  "source": "meta"
}