{
  "id": 222643,
  "title": "Using GPU for training models",
  "url": "/competitions/hpa-single-cell-image-classification/discussion/222643",
  "author_name": "GitMach",
  "post_date": "2021-02-28T11:51:26.555000",
  "votes": 1,
  "comment_count": 17,
  "views": 0,
  "content": "<p>Hi, </p>\n<p>I've activated the GPU but it doesn't show GPU activity during train and predicting the model.<br>\nHow can use my GPU during training and prediction?</p>\n<p>Thank you for your feedback.</p>",
  "messages": [
    {
      "id": 1221064,
      "postDate": "2021-02-28T16:36:09.597Z",
      "content": "<p>Ghost - this is pretty open ended question that might not get a fix - there is a ton of things you could have wrong and GPU activity does depend on lots of things.</p>\n<p>If you get no answer in the next day suggest that you make your kernel public and come back with a link so folks can see what specific things your doing.</p>",
      "rawMarkdown": "Ghost - this is pretty open ended question that might not get a fix - there is a ton of things you could have wrong and GPU activity does depend on lots of things.\n\nIf you get no answer in the next day suggest that you make your kernel public and come back with a link so folks can see what specific things your doing.",
      "votes": 1,
      "replies": [
        {
          "id": 1221341,
          "postDate": "2021-02-28T22:45:08.727Z",
          "content": "<p><a href=\"https://www.kaggle.com/pcjimmmy\" target=\"_blank\">@pcjimmmy</a> <br>\nHere 's the link<br>\n<a href=\"https://www.kaggle.com/centorit/hpa-keras-model1\" target=\"_blank\">https://www.kaggle.com/centorit/hpa-keras-model1</a></p>\n<p>As you can see, training the model with only 100 images on train and 50 images on test takes to long, <br>\nBesides, even when GPU is activated you won't see any activity in the top right box \"GPU\"<br>\nthanks for any help</p>",
          "rawMarkdown": "@pcjimmmy \nHere 's the link\nhttps://www.kaggle.com/centorit/hpa-keras-model1\n\nAs you can see, training the model with only 100 images on train and 50 images on test takes to long, \nBesides, even when GPU is activated you won't see any activity in the top right box \"GPU\"\nthanks for any help\n\n\n"
        },
        {
          "id": 1221782,
          "postDate": "2021-03-01T09:17:54.377Z",
          "content": "<p>Tried to fork and run your kernel - but your csv file is in a private dataset.    Can you make that data set public also??</p>\n<p>Once I forked the kernel and stepped thru or Save/Run I got the same device indications - that a CPU and GPU existed. (I did have to turn on the GPU)</p>\n<p>The GPU should pretty much only be active during the last cell where you do the training.   Without your csv file this cell of course does not run for me - the only thing I see that does not seem right is the  batch_size.  Since your data is being fed by your generator you should not have a batch_size specified in the fit - and 4096 is too big.  Comment out that line in fit.  You also specify this size in your generator - should be more like 24 if your reading the image files from train.  But maybe since you seem to be doing a 100x100 image size 4096 might work - I would be very surprised if a 100x100 image ends up with any accuracy.  </p>\n<p>You may also be thinking that the GPU not working when the issue might be the CPU trying to build a 4096 batch - not sure how many cpu cores Kaggle is giving us these days but you might have bottle neck of CPU feeding the GPU.  At one time I think we only got a single core.</p>\n<pre><code>model.fit(generator_wrapper(training_set),\n                  #(np.asarray(training_set).astype(\"float32\")),\n                    #steps_per_epoch=STEP_SIZE_TRAIN,\n                    validation_data=generator_wrapper(test_set),\n                    batch_size=4096, \n                    #validation_steps=STEP_SIZE_VALID,\n                    epochs=3,\n                    callbacks = [rlr,ckp,es],\n                    #verbose=2\n         )\n</code></pre>",
          "rawMarkdown": "Tried to fork and run your kernel - but your csv file is in a private dataset.    Can you make that data set public also??\n\nOnce I forked the kernel and stepped thru or Save/Run I got the same device indications - that a CPU and GPU existed. (I did have to turn on the GPU)\n\nThe GPU should pretty much only be active during the last cell where you do the training.   Without your csv file this cell of course does not run for me - the only thing I see that does not seem right is the  batch_size.  Since your data is being fed by your generator you should not have a batch_size specified in the fit - and 4096 is too big.  Comment out that line in fit.  You also specify this size in your generator - should be more like 24 if your reading the image files from train.  But maybe since you seem to be doing a 100x100 image size 4096 might work - I would be very surprised if a 100x100 image ends up with any accuracy.  \n\nYou may also be thinking that the GPU not working when the issue might be the CPU trying to build a 4096 batch - not sure how many cpu cores Kaggle is giving us these days but you might have bottle neck of CPU feeding the GPU.  At one time I think we only got a single core.\n\n\n```\nmodel.fit(generator_wrapper(training_set),\n                  #(np.asarray(training_set).astype(\"float32\")),\n                    #steps_per_epoch=STEP_SIZE_TRAIN,\n                    validation_data=generator_wrapper(test_set),\n                    batch_size=4096, \n                    #validation_steps=STEP_SIZE_VALID,\n                    epochs=3,\n                    callbacks = [rlr,ckp,es],\n                    #verbose=2\n         )\n```\n\n",
          "votes": 2
        },
        {
          "id": 1221977,
          "postDate": "2021-03-01T12:59:38.650Z",
          "content": "<p><a href=\"https://www.kaggle.com/centorit\" target=\"_blank\">@centorit</a> – The GPU activity, shown in the top right, will be near - (if not non-existent) if the Input/Output portion of your pipeline takes MUCH longer than the processing part. IO is often the limiting factor/bottleneck. I would examine how you are loading/opening/preprocessing the data and find ways to speed this up (or parallelize it).</p>\n<p>Additionally, in frameworks like <strong><code>tf.data</code></strong>, you can cache certain preprocessing steps/transformations so they only have to be performed once. <a href=\"https://www.tensorflow.org/guide/data_performance\" target=\"_blank\"><strong>See here for tips/tricks related to tf.data</strong></a></p>",
          "rawMarkdown": "@centorit – The GPU activity, shown in the top right, will be near - (if not non-existent) if the Input/Output portion of your pipeline takes MUCH longer than the processing part. IO is often the limiting factor/bottleneck. I would examine how you are loading/opening/preprocessing the data and find ways to speed this up (or parallelize it).\n\nAdditionally, in frameworks like **`tf.data`**, you can cache certain preprocessing steps/transformations so they only have to be performed once. [**See here for tips/tricks related to tf.data**](https://www.tensorflow.org/guide/data_performance)",
          "votes": 1
        },
        {
          "id": 1222029,
          "postDate": "2021-03-01T13:37:13.900Z",
          "content": "<p>Sorry, I thought sharing notebook with dataset included, would also share the dataset.<br>\nDataset it's public, here's the link</p>\n<pre><code>www.kaggle.com/dataset/c4bcf19bca717e2ecdee8f9bc4b4bc26387efb2cb1d3e84e3b84d448f98d6af3\n</code></pre>\n<p>I'm going to test your suggestions and build a simple Conv to start.</p>",
          "rawMarkdown": "Sorry, I thought sharing notebook with dataset included, would also share the dataset.\nDataset it's public, here's the link\n```\nwww.kaggle.com/dataset/c4bcf19bca717e2ecdee8f9bc4b4bc26387efb2cb1d3e84e3b84d448f98d6af3\n\n```\nI'm going to test your suggestions and build a simple Conv to start."
        },
        {
          "id": 1222322,
          "postDate": "2021-03-01T17:48:13.633Z",
          "content": "<p>Forked your kernel and stepped thru it.</p>\n<p>You are for sure creating a bottleneck - CPU was running 99% during training but GPU had only very infrequent display of usage, with biggest number I saw 14%.</p>\n<p>Not sure why you have so many loss functions defined but that combined with the 4096 batch means your CPU is working like crazy but the GPU - WHICH IS WORKING only uses a very small slice of the time.   Get down to a single loss function and smaller batch and you will see the GPU doing it's thing.</p>",
          "rawMarkdown": "Forked your kernel and stepped thru it.\n\nYou are for sure creating a bottleneck - CPU was running 99% during training but GPU had only very infrequent display of usage, with biggest number I saw 14%.\n\nNot sure why you have so many loss functions defined but that combined with the 4096 batch means your CPU is working like crazy but the GPU - WHICH IS WORKING only uses a very small slice of the time.   Get down to a single loss function and smaller batch and you will see the GPU doing it's thing."
        }
      ]
    },
    {
      "id": 1220836,
      "postDate": "2021-02-28T11:51:26.557Z",
      "content": "<p>Hi, </p>\n<p>I've activated the GPU but it doesn't show GPU activity during train and predicting the model.<br>\nHow can use my GPU during training and prediction?</p>\n<p>Thank you for your feedback.</p>",
      "rawMarkdown": "Hi, \n\nI've activated the GPU but it doesn't show GPU activity during train and predicting the model.\nHow can use my GPU during training and prediction?\n\nThank you for your feedback.",
      "votes": 1
    },
    {
      "id": 1221073,
      "postDate": "2021-02-28T16:45:44.203Z",
      "content": "<p>I think that may be specific to the framework you're using? E.g. to('cuda') in Pytorch?</p>",
      "rawMarkdown": "I think that may be specific to the framework you're using? E.g. to('cuda') in Pytorch?",
      "replies": [
        {
          "id": 1221181,
          "postDate": "2021-02-28T18:20:22.483Z",
          "content": "<p>I'm just using Keras, and Tensorflow Keras.<br>\nWhen I activate GPU I can check it's activated as result :</p>\n<pre><code>[name: \"/device:CPU:0\"\ndevice_type: \"CPU\"\nmemory_limit: 268435456\nlocality {\n}\n**incarnation: 10302733618983122258\n, name: \"/device:GPU:0\"\ndevice_type: \"GPU\"\nmemory_limit: 15685569792\nlocality {\n  bus_id: 1\n  links {\n  }\n}\nincarnation: 12635735069039391878\nphysical_device_desc: \"device: 0, name: Tesla P100-PCIE-16GB, pci bus id: 0000:00:04.0, compute capability: 6.0\"**\n]\n</code></pre>\n<p>Previous competition \"Jane Street Market\" I could use GPU just using Keras back end instead of tensorflow.keras.<br>\nAnd then <br>\nBut in this competition none of them are working.</p>",
          "rawMarkdown": "I'm just using Keras, and Tensorflow Keras.\nWhen I activate GPU I can check it's activated as result :\n\n```\n[name: \"/device:CPU:0\"\ndevice_type: \"CPU\"\nmemory_limit: 268435456\nlocality {\n}\n**incarnation: 10302733618983122258\n, name: \"/device:GPU:0\"\ndevice_type: \"GPU\"\nmemory_limit: 15685569792\nlocality {\n  bus_id: 1\n  links {\n  }\n}\nincarnation: 12635735069039391878\nphysical_device_desc: \"device: 0, name: Tesla P100-PCIE-16GB, pci bus id: 0000:00:04.0, compute capability: 6.0\"**\n]\n```\nPrevious competition \"Jane Street Market\" I could use GPU just using Keras back end instead of tensorflow.keras.\nAnd then \nBut in this competition none of them are working."
        },
        {
          "id": 1221415,
          "postDate": "2021-03-01T00:38:52.243Z",
          "content": "<p>Do you have this problem when you run the notebook locally? Or only when you submit it?<br>\nI  am also using Keras and so far I had no problem with GPU or CPU. Note that when you submit there is an option to run it with or without an accelerator. </p>",
          "rawMarkdown": "Do you have this problem when you run the notebook locally? Or only when you submit it?\nI  am also using Keras and so far I had no problem with GPU or CPU. Note that when you submit there is an option to run it with or without an accelerator. ",
          "isDeleted": true
        },
        {
          "id": 1221426,
          "postDate": "2021-03-01T00:53:27.950Z",
          "content": "<p>It is also possible that you are using GPU but there is a bottleneck somewhere else for example at dataloading. It might make sense to run a simple program from the Keras tutorials (something like a MNIST classifier where you can fit using a numpy array) just to check if the issue is really the GPU.  </p>",
          "rawMarkdown": "It is also possible that you are using GPU but there is a bottleneck somewhere else for example at dataloading. It might make sense to run a simple program from the Keras tutorials (something like a MNIST classifier where you can fit using a numpy array) just to check if the issue is really the GPU.  ",
          "votes": 1,
          "isDeleted": true
        },
        {
          "id": 1221772,
          "postDate": "2021-03-01T09:05:41.867Z",
          "content": "<p>When running the notebook, I can't submit yet if I can't even get the results. </p>",
          "rawMarkdown": "When running the notebook, I can't submit yet if I can't even get the results. "
        },
        {
          "id": 1221942,
          "postDate": "2021-03-01T12:21:37.887Z",
          "content": "<p>I see. I have looked at the public notebook. The GPU part seems fine, the GPU turns on when activated. I don't quite understand the reason why it would run slow, however the data generator is a bit complex. In order to find the bottleneck, I would suggest first experimenting with  a simpler pipeline to see if that works OK , and then make it more complex:</p>\n<p>As an experiment to check the timing : perhaps prepare first  small numpy arrays X, y, where X has shape (40000, 100, 100, 1), (40000, 100, 100, 3) or (40000, 100, 100, 4) , depending on how many colors you use, and y has shape (40000, 19) and run the convolutional model with the last layer <br>\nDense(19, activation = 'sigmoid'). (loss = binary_crossentropy) I would suggest not to use any data augmentation at the first experiment. If you fit using <br>\nmodel.fit(X,y, epochs = 5) <br>\nthen there should be  a large speed up with GPU vs CPU. </p>",
          "rawMarkdown": "I see. I have looked at the public notebook. The GPU part seems fine, the GPU turns on when activated. I don't quite understand the reason why it would run slow, however the data generator is a bit complex. In order to find the bottleneck, I would suggest first experimenting with  a simpler pipeline to see if that works OK , and then make it more complex:\n\nAs an experiment to check the timing : perhaps prepare first  small numpy arrays X, y, where X has shape (40000, 100, 100, 1), (40000, 100, 100, 3) or (40000, 100, 100, 4) , depending on how many colors you use, and y has shape (40000, 19) and run the convolutional model with the last layer \nDense(19, activation = 'sigmoid'). (loss = binary_crossentropy) I would suggest not to use any data augmentation at the first experiment. If you fit using \nmodel.fit(X,y, epochs = 5) \nthen there should be  a large speed up with GPU vs CPU. ",
          "votes": 1,
          "isDeleted": true
        },
        {
          "id": 1222434,
          "postDate": "2021-03-01T19:38:53.040Z",
          "content": "<p>I've finally got a result. But I had to change some parameters:<br>\n1 - I've reduced the size of the batch from 4096 o 64<br>\n2 - Change class_mode from 'raw' to  'multi_output' <br>\n3 -  No use \"generator_wrapper\" function <br>\n4 - Suppress batch_size during fit </p>\n<p>The model compiled and fit this time with 10 epochs, within acceptable time, even if it didn't change between GPU active and GPU not active, (something around 10s/epoch)</p>\n<p>My preds returned a matrix  dim (50,19) 😳</p>\n<pre><code>len([[0.24326178],\n[0.24545057],\n[0.18905358],\n[0.22454323],\n[0.17228684],\n[0.23446988],\n[0.23484272],\n[0.19753459],\n[0.26074097],\n[0.2010675 ],\n[0.10761576],\n[0.22031462],\n[0.23193006],\n[0.22713315],\n[0.11388983],\n[0.23554248],\n[0.1785873 ],\n[0.25215697],\n[0.18093413],\n[0.1990953 ],\n[0.20176446],\n[0.19727392],\n[0.2491296 ],\n[0.11087943],\n[0.22272176],\n[0.21852216],\n[0.207871  ],\n[0.1856019 ],\n[0.2678131 ],\n[0.20906556],\n[0.27924448],\n[0.27829096],\n[0.13388641],\n[0.2648097 ],\n[0.24399373],\n[0.24952464],\n[0.157573  ],\n[0.23329581],\n[0.23566657],\n[0.27109227],\n[0.22023298],\n[0.21330912],\n[0.28557703],\n[0.17392349],\n[0.17297237],\n[0.24998474],\n[0.24437207],\n[0.06390683],\n[0.20857255],\n[0.23521458]])\n</code></pre>\n<p>How can I interpret this? </p>\n<p>Anyway, I'm going to start from scratch, with a simple ConvNet model but if you could help to understand what's wrong with this approach, </p>",
          "rawMarkdown": "I've finally got a result. But I had to change some parameters:\n1 - I've reduced the size of the batch from 4096 o 64\n2 - Change class_mode from 'raw' to  'multi_output' \n3 -  No use \"generator_wrapper\" function \n4 - Suppress batch_size during fit \n\nThe model compiled and fit this time with 10 epochs, within acceptable time, even if it didn't change between GPU active and GPU not active, (something around 10s/epoch)\n\nMy preds returned a matrix  dim (50,19) 😳\n\n```\nlen([[0.24326178],\n[0.24545057],\n[0.18905358],\n[0.22454323],\n[0.17228684],\n[0.23446988],\n[0.23484272],\n[0.19753459],\n[0.26074097],\n[0.2010675 ],\n[0.10761576],\n[0.22031462],\n[0.23193006],\n[0.22713315],\n[0.11388983],\n[0.23554248],\n[0.1785873 ],\n[0.25215697],\n[0.18093413],\n[0.1990953 ],\n[0.20176446],\n[0.19727392],\n[0.2491296 ],\n[0.11087943],\n[0.22272176],\n[0.21852216],\n[0.207871  ],\n[0.1856019 ],\n[0.2678131 ],\n[0.20906556],\n[0.27924448],\n[0.27829096],\n[0.13388641],\n[0.2648097 ],\n[0.24399373],\n[0.24952464],\n[0.157573  ],\n[0.23329581],\n[0.23566657],\n[0.27109227],\n[0.22023298],\n[0.21330912],\n[0.28557703],\n[0.17392349],\n[0.17297237],\n[0.24998474],\n[0.24437207],\n[0.06390683],\n[0.20857255],\n[0.23521458]])\n```\nHow can I interpret this? \n\nAnyway, I'm going to start from scratch, with a simple ConvNet model but if you could help to understand what's wrong with this approach, \n"
        },
        {
          "id": 1222441,
          "postDate": "2021-03-01T19:49:17.113Z",
          "content": "<p>I don't think there is anything wrong with this approach.<br>\nIt seems these are the class probabilities: pred[n,k] gives the probability that cell n has protein k </p>",
          "rawMarkdown": "I don't think there is anything wrong with this approach.\nIt seems these are the class probabilities: pred[n,k] gives the probability that cell n has protein k ",
          "isDeleted": true
        },
        {
          "id": 1222577,
          "postDate": "2021-03-01T23:06:37.737Z",
          "content": "<p><a href=\"https://www.kaggle.com/Zoltan\" target=\"_blank\">@Zoltan</a> <br>\nI'm going to open another topic for this subject because I dont get how I can compare all this probabilities for a given class with the true values.<br>\nAnyway, now I can train my model, only for \"green\" .png files and a small sample. <br>\nWhen I tried to train the 16000 images it takes 3000s per epoch and with GPU or not, doesn't change anything. </p>",
          "rawMarkdown": "@Zoltan \nI'm going to open another topic for this subject because I dont get how I can compare all this probabilities for a given class with the true values.\nAnyway, now I can train my model, only for \"green\" .png files and a small sample. \nWhen I tried to train the 16000 images it takes 3000s per epoch and with GPU or not, doesn't change anything. "
        },
        {
          "id": 1222586,
          "postDate": "2021-03-01T23:39:31.037Z",
          "content": "<p>That is way too long indeed. Ideally you should aim to process 1 epoch in a couple of minutes. There are a few notebooks that first  preprocess the data: take the large picture, find the boxes for the cells, make a tiled image of a cell. See the notebooks by Darek Kleczek    <a href=\"https://www.kaggle.com/thedrcat\" target=\"_blank\">@thedrcat</a>  for example.<br>\nHere the preprocessing part takes long, but after that the files are saved in one notebook, then the learning can be done in another notebook more quickly. Anyway, Darek's notebooks give a good way to start. Good luck! </p>",
          "rawMarkdown": "That is way too long indeed. Ideally you should aim to process 1 epoch in a couple of minutes. There are a few notebooks that first  preprocess the data: take the large picture, find the boxes for the cells, make a tiled image of a cell. See the notebooks by Darek Kleczek    @thedrcat  for example.\nHere the preprocessing part takes long, but after that the files are saved in one notebook, then the learning can be done in another notebook more quickly. Anyway, Darek's notebooks give a good way to start. Good luck! ",
          "votes": 1,
          "isDeleted": true
        },
        {
          "id": 1224382,
          "postDate": "2021-03-02T17:08:04.777Z",
          "content": "<p>This thread has helped me understand the issue with my notebook as well, thanks! :)</p>",
          "rawMarkdown": "This thread has helped me understand the issue with my notebook as well, thanks! :)"
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 1221064,
      "author_name": "PC Jimmmy",
      "author_url": "",
      "post_date": "2021-02-28T16:36:09.597000",
      "content": "<p>Ghost - this is pretty open ended question that might not get a fix - there is a ton of things you could have wrong and GPU activity does depend on lots of things.</p>\n<p>If you get no answer in the next day suggest that you make your kernel public and come back with a link so folks can see what specific things your doing.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1221341,
          "author_name": "GitMach",
          "author_url": "",
          "post_date": "2021-02-28T22:45:08.727000",
          "content": "<p><a href=\"https://www.kaggle.com/pcjimmmy\" target=\"_blank\">@pcjimmmy</a> <br>\nHere 's the link<br>\n<a href=\"https://www.kaggle.com/centorit/hpa-keras-model1\" target=\"_blank\">https://www.kaggle.com/centorit/hpa-keras-model1</a></p>\n<p>As you can see, training the model with only 100 images on train and 50 images on test takes to long, <br>\nBesides, even when GPU is activated you won't see any activity in the top right box \"GPU\"<br>\nthanks for any help</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1221782,
          "author_name": "PC Jimmmy",
          "author_url": "",
          "post_date": "2021-03-01T09:17:54.377000",
          "content": "<p>Tried to fork and run your kernel - but your csv file is in a private dataset.    Can you make that data set public also??</p>\n<p>Once I forked the kernel and stepped thru or Save/Run I got the same device indications - that a CPU and GPU existed. (I did have to turn on the GPU)</p>\n<p>The GPU should pretty much only be active during the last cell where you do the training.   Without your csv file this cell of course does not run for me - the only thing I see that does not seem right is the  batch_size.  Since your data is being fed by your generator you should not have a batch_size specified in the fit - and 4096 is too big.  Comment out that line in fit.  You also specify this size in your generator - should be more like 24 if your reading the image files from train.  But maybe since you seem to be doing a 100x100 image size 4096 might work - I would be very surprised if a 100x100 image ends up with any accuracy.  </p>\n<p>You may also be thinking that the GPU not working when the issue might be the CPU trying to build a 4096 batch - not sure how many cpu cores Kaggle is giving us these days but you might have bottle neck of CPU feeding the GPU.  At one time I think we only got a single core.</p>\n<pre><code>model.fit(generator_wrapper(training_set),\n                  #(np.asarray(training_set).astype(\"float32\")),\n                    #steps_per_epoch=STEP_SIZE_TRAIN,\n                    validation_data=generator_wrapper(test_set),\n                    batch_size=4096, \n                    #validation_steps=STEP_SIZE_VALID,\n                    epochs=3,\n                    callbacks = [rlr,ckp,es],\n                    #verbose=2\n         )\n</code></pre>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1221977,
          "author_name": "Darien Schettler",
          "author_url": "",
          "post_date": "2021-03-01T12:59:38.650000",
          "content": "<p><a href=\"https://www.kaggle.com/centorit\" target=\"_blank\">@centorit</a> – The GPU activity, shown in the top right, will be near - (if not non-existent) if the Input/Output portion of your pipeline takes MUCH longer than the processing part. IO is often the limiting factor/bottleneck. I would examine how you are loading/opening/preprocessing the data and find ways to speed this up (or parallelize it).</p>\n<p>Additionally, in frameworks like <strong><code>tf.data</code></strong>, you can cache certain preprocessing steps/transformations so they only have to be performed once. <a href=\"https://www.tensorflow.org/guide/data_performance\" target=\"_blank\"><strong>See here for tips/tricks related to tf.data</strong></a></p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1222029,
          "author_name": "GitMach",
          "author_url": "",
          "post_date": "2021-03-01T13:37:13.900000",
          "content": "<p>Sorry, I thought sharing notebook with dataset included, would also share the dataset.<br>\nDataset it's public, here's the link</p>\n<pre><code>www.kaggle.com/dataset/c4bcf19bca717e2ecdee8f9bc4b4bc26387efb2cb1d3e84e3b84d448f98d6af3\n</code></pre>\n<p>I'm going to test your suggestions and build a simple Conv to start.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1222322,
          "author_name": "PC Jimmmy",
          "author_url": "",
          "post_date": "2021-03-01T17:48:13.633000",
          "content": "<p>Forked your kernel and stepped thru it.</p>\n<p>You are for sure creating a bottleneck - CPU was running 99% during training but GPU had only very infrequent display of usage, with biggest number I saw 14%.</p>\n<p>Not sure why you have so many loss functions defined but that combined with the 4096 batch means your CPU is working like crazy but the GPU - WHICH IS WORKING only uses a very small slice of the time.   Get down to a single loss function and smaller batch and you will see the GPU doing it's thing.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1221073,
      "author_name": "Darek Kłeczek",
      "author_url": "",
      "post_date": "2021-02-28T16:45:44.203000",
      "content": "<p>I think that may be specific to the framework you're using? E.g. to('cuda') in Pytorch?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1221181,
          "author_name": "GitMach",
          "author_url": "",
          "post_date": "2021-02-28T18:20:22.483000",
          "content": "<p>I'm just using Keras, and Tensorflow Keras.<br>\nWhen I activate GPU I can check it's activated as result :</p>\n<pre><code>[name: \"/device:CPU:0\"\ndevice_type: \"CPU\"\nmemory_limit: 268435456\nlocality {\n}\n**incarnation: 10302733618983122258\n, name: \"/device:GPU:0\"\ndevice_type: \"GPU\"\nmemory_limit: 15685569792\nlocality {\n  bus_id: 1\n  links {\n  }\n}\nincarnation: 12635735069039391878\nphysical_device_desc: \"device: 0, name: Tesla P100-PCIE-16GB, pci bus id: 0000:00:04.0, compute capability: 6.0\"**\n]\n</code></pre>\n<p>Previous competition \"Jane Street Market\" I could use GPU just using Keras back end instead of tensorflow.keras.<br>\nAnd then <br>\nBut in this competition none of them are working.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1221415,
          "author_name": "",
          "author_url": "",
          "post_date": "2021-03-01T00:38:52.243000",
          "content": "<p>Do you have this problem when you run the notebook locally? Or only when you submit it?<br>\nI  am also using Keras and so far I had no problem with GPU or CPU. Note that when you submit there is an option to run it with or without an accelerator. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1221426,
          "author_name": "",
          "author_url": "",
          "post_date": "2021-03-01T00:53:27.950000",
          "content": "<p>It is also possible that you are using GPU but there is a bottleneck somewhere else for example at dataloading. It might make sense to run a simple program from the Keras tutorials (something like a MNIST classifier where you can fit using a numpy array) just to check if the issue is really the GPU.  </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1221772,
          "author_name": "GitMach",
          "author_url": "",
          "post_date": "2021-03-01T09:05:41.867000",
          "content": "<p>When running the notebook, I can't submit yet if I can't even get the results. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1221942,
          "author_name": "",
          "author_url": "",
          "post_date": "2021-03-01T12:21:37.887000",
          "content": "<p>I see. I have looked at the public notebook. The GPU part seems fine, the GPU turns on when activated. I don't quite understand the reason why it would run slow, however the data generator is a bit complex. In order to find the bottleneck, I would suggest first experimenting with  a simpler pipeline to see if that works OK , and then make it more complex:</p>\n<p>As an experiment to check the timing : perhaps prepare first  small numpy arrays X, y, where X has shape (40000, 100, 100, 1), (40000, 100, 100, 3) or (40000, 100, 100, 4) , depending on how many colors you use, and y has shape (40000, 19) and run the convolutional model with the last layer <br>\nDense(19, activation = 'sigmoid'). (loss = binary_crossentropy) I would suggest not to use any data augmentation at the first experiment. If you fit using <br>\nmodel.fit(X,y, epochs = 5) <br>\nthen there should be  a large speed up with GPU vs CPU. </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1222434,
          "author_name": "GitMach",
          "author_url": "",
          "post_date": "2021-03-01T19:38:53.040000",
          "content": "<p>I've finally got a result. But I had to change some parameters:<br>\n1 - I've reduced the size of the batch from 4096 o 64<br>\n2 - Change class_mode from 'raw' to  'multi_output' <br>\n3 -  No use \"generator_wrapper\" function <br>\n4 - Suppress batch_size during fit </p>\n<p>The model compiled and fit this time with 10 epochs, within acceptable time, even if it didn't change between GPU active and GPU not active, (something around 10s/epoch)</p>\n<p>My preds returned a matrix  dim (50,19) 😳</p>\n<pre><code>len([[0.24326178],\n[0.24545057],\n[0.18905358],\n[0.22454323],\n[0.17228684],\n[0.23446988],\n[0.23484272],\n[0.19753459],\n[0.26074097],\n[0.2010675 ],\n[0.10761576],\n[0.22031462],\n[0.23193006],\n[0.22713315],\n[0.11388983],\n[0.23554248],\n[0.1785873 ],\n[0.25215697],\n[0.18093413],\n[0.1990953 ],\n[0.20176446],\n[0.19727392],\n[0.2491296 ],\n[0.11087943],\n[0.22272176],\n[0.21852216],\n[0.207871  ],\n[0.1856019 ],\n[0.2678131 ],\n[0.20906556],\n[0.27924448],\n[0.27829096],\n[0.13388641],\n[0.2648097 ],\n[0.24399373],\n[0.24952464],\n[0.157573  ],\n[0.23329581],\n[0.23566657],\n[0.27109227],\n[0.22023298],\n[0.21330912],\n[0.28557703],\n[0.17392349],\n[0.17297237],\n[0.24998474],\n[0.24437207],\n[0.06390683],\n[0.20857255],\n[0.23521458]])\n</code></pre>\n<p>How can I interpret this? </p>\n<p>Anyway, I'm going to start from scratch, with a simple ConvNet model but if you could help to understand what's wrong with this approach, </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1222441,
          "author_name": "",
          "author_url": "",
          "post_date": "2021-03-01T19:49:17.113000",
          "content": "<p>I don't think there is anything wrong with this approach.<br>\nIt seems these are the class probabilities: pred[n,k] gives the probability that cell n has protein k </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1222577,
          "author_name": "GitMach",
          "author_url": "",
          "post_date": "2021-03-01T23:06:37.737000",
          "content": "<p><a href=\"https://www.kaggle.com/Zoltan\" target=\"_blank\">@Zoltan</a> <br>\nI'm going to open another topic for this subject because I dont get how I can compare all this probabilities for a given class with the true values.<br>\nAnyway, now I can train my model, only for \"green\" .png files and a small sample. <br>\nWhen I tried to train the 16000 images it takes 3000s per epoch and with GPU or not, doesn't change anything. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1222586,
          "author_name": "",
          "author_url": "",
          "post_date": "2021-03-01T23:39:31.037000",
          "content": "<p>That is way too long indeed. Ideally you should aim to process 1 epoch in a couple of minutes. There are a few notebooks that first  preprocess the data: take the large picture, find the boxes for the cells, make a tiled image of a cell. See the notebooks by Darek Kleczek    <a href=\"https://www.kaggle.com/thedrcat\" target=\"_blank\">@thedrcat</a>  for example.<br>\nHere the preprocessing part takes long, but after that the files are saved in one notebook, then the learning can be done in another notebook more quickly. Anyway, Darek's notebooks give a good way to start. Good luck! </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1224382,
          "author_name": "Arka Saha",
          "author_url": "",
          "post_date": "2021-03-02T17:08:04.777000",
          "content": "<p>This thread has helped me understand the issue with my notebook as well, thanks! :)</p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1221064": "Ghost - this is pretty open ended question that might not get a fix - there is a ton of things you could have wrong and GPU activity does depend on lots of things.\n\nIf you get no answer in the next day suggest that you make your kernel public and come back with a link so folks can see what specific things your doing.",
    "1220836": "Hi, \n\nI've activated the GPU but it doesn't show GPU activity during train and predicting the model.\nHow can use my GPU during training and prediction?\n\nThank you for your feedback.",
    "1221073": "I think that may be specific to the framework you're using? E.g. to('cuda') in Pytorch?"
  }
}