{
  "id": 48258,
  "title": "How to check if you are maximizing your GPU",
  "url": "/competitions/sp-society-camera-model-identification/discussion/48258",
  "author_name": "",
  "post_date": "2018-01-25T09:37:54.277164700Z",
  "votes": 9,
  "comment_count": 7,
  "views": 0,
  "content": "<p>One of the things I'm learning on Kaggle competitions is that experimenting speed is key. </p>\n\n<p>Advice you get from top Kagglers is to first try with reduced complexity networks (in our case simple networks and/or reduced crop size) to get an idea of which hyper-parameters work best (although when you scale up these may/will change); but I also think we need to maximize the code efficiency too.</p>\n\n<p>I often run my training and run <code>watch -n1 nvidia-smi</code>. It's a good sign if the GPUs being maxxed out:</p>\n\n<p><img src=\"http://i.imgur.com/k6aT8FR.png\" alt=\"enter image description here\"></p>\n\n<p>also running <code>htop</code> to see if the CPU cores are being effectively used:</p>\n\n<p><img src=\"http://i.imgur.com/k15JyzW.png\" alt=\"enter image description here\"></p>\n\n<p>In this competition the images are huge so in my first attempt I tried running a simple Keras generator that loaded images and (after processing) sent them to the GPUs. It took roughly ~20+ minutes per epoch (using 2 x 1080  Tis + AMD Threadripper 1950X with 16 cores / 32 threads). Even though Keras generators supposedly support multiprocessing with an argument to <code>.fit_generator</code> I was unable to make them work correctly, so my next attempt was to cache images into a dict so epochs after the first one will just take them from the dict.</p>\n\n<p>The issue is that I cannot keep all uncompressed images in memory (I have 64 Gb RAM but all uncompressed images take more), so I just cached cropped versions of images. While OK for a first try, it's an obvious compromise. Also, trying to saved a pickled version of the dictionary with images ~1024x1024 would result in memory errors (go figure). So another approach was needed:</p>\n\n<p>I then used Python multiprocessing module inside the Keras generator and now the code runs one epoch in ~220 seconds.</p>\n\n<p>What code / monitoring tools/tricks do you use to maximize performance?</p>",
  "messages": [
    {
      "id": "273843",
      "postDate": "01/25/2018 09:37:54",
      "content": "<p>One of the things I'm learning on Kaggle competitions is that experimenting speed is key. </p>\n\n<p>Advice you get from top Kagglers is to first try with reduced complexity networks (in our case simple networks and/or reduced crop size) to get an idea of which hyper-parameters work best (although when you scale up these may/will change); but I also think we need to maximize the code efficiency too.</p>\n\n<p>I often run my training and run <code>watch -n1 nvidia-smi</code>. It's a good sign if the GPUs being maxxed out:</p>\n\n<p><img src=\"http://i.imgur.com/k6aT8FR.png\" alt=\"enter image description here\"></p>\n\n<p>also running <code>htop</code> to see if the CPU cores are being effectively used:</p>\n\n<p><img src=\"http://i.imgur.com/k15JyzW.png\" alt=\"enter image description here\"></p>\n\n<p>In this competition the images are huge so in my first attempt I tried running a simple Keras generator that loaded images and (after processing) sent them to the GPUs. It took roughly ~20+ minutes per epoch (using 2 x 1080  Tis + AMD Threadripper 1950X with 16 cores / 32 threads). Even though Keras generators supposedly support multiprocessing with an argument to <code>.fit_generator</code> I was unable to make them work correctly, so my next attempt was to cache images into a dict so epochs after the first one will just take them from the dict.</p>\n\n<p>The issue is that I cannot keep all uncompressed images in memory (I have 64 Gb RAM but all uncompressed images take more), so I just cached cropped versions of images. While OK for a first try, it's an obvious compromise. Also, trying to saved a pickled version of the dictionary with images ~1024x1024 would result in memory errors (go figure). So another approach was needed:</p>\n\n<p>I then used Python multiprocessing module inside the Keras generator and now the code runs one epoch in ~220 seconds.</p>\n\n<p>What code / monitoring tools/tricks do you use to maximize performance?</p>",
      "rawMarkdown": "One of the things I'm learning on Kaggle competitions is that experimenting speed is key. \n\nAdvice you get from top Kagglers is to first try with reduced complexity networks (in our case simple networks and/or reduced crop size) to get an idea of which hyper-parameters work best (although when you scale up these may/will change); but I also think we need to maximize the code efficiency too.\n\nI often run my training and run `watch -n1 nvidia-smi`. It's a good sign if the GPUs being maxxed out:\n\n![enter image description here][1]\n\nalso running `htop` to see if the CPU cores are being effectively used:\n\n![enter image description here][2]\n\nIn this competition the images are huge so in my first attempt I tried running a simple Keras generator that loaded images and (after processing) sent them to the GPUs. It took roughly ~20+ minutes per epoch (using 2 x 1080  Tis + AMD Threadripper 1950X with 16 cores / 32 threads). Even though Keras generators supposedly support multiprocessing with an argument to `.fit_generator` I was unable to make them work correctly, so my next attempt was to cache images into a dict so epochs after the first one will just take them from the dict.\n\nThe issue is that I cannot keep all uncompressed images in memory (I have 64 Gb RAM but all uncompressed images take more), so I just cached cropped versions of images. While OK for a first try, it's an obvious compromise. Also, trying to saved a pickled version of the dictionary with images ~1024x1024 would result in memory errors (go figure). So another approach was needed:\n\nI then used Python multiprocessing module inside the Keras generator and now the code runs one epoch in ~220 seconds.\n\nWhat code / monitoring tools/tricks do you use to maximize performance?\n\n  [1]: http://i.imgur.com/k6aT8FR.png\n  [2]: http://i.imgur.com/k15JyzW.png",
      "votes": null
    },
    {
      "id": "274590",
      "postDate": "01/26/2018 20:55:24",
      "content": "<p>Small hint: there is built-in switch to make nvidia-smi repeat its output: add -l seconds or -lms milliseconds. I usually use <code>nvidia-smi -lms 200</code> to monitor GPU usage more granularly. </p>",
      "rawMarkdown": "Small hint: there is built-in switch to make nvidia-smi repeat its output: add -l seconds or -lms milliseconds. I usually use `nvidia-smi -lms 200` to monitor GPU usage more granularly.",
      "votes": null
    },
    {
      "id": "274593",
      "postDate": "01/26/2018 20:57:36",
      "content": "<p>Thanks. How to do make it so it stays at the top of the console instead of scrolling indefinitely?</p>",
      "rawMarkdown": "Thanks. How to do make it so it stays at the top of the console instead of scrolling indefinitely?",
      "votes": null
    },
    {
      "id": "274601",
      "postDate": "01/26/2018 21:26:27",
      "content": "<p>It is scrolling the same amount each update so you see just static current output if console size is big enough :) </p>",
      "rawMarkdown": "It is scrolling the same amount each update so you see just static current output if console size is big enough :)",
      "votes": null
    },
    {
      "id": "274818",
      "postDate": "01/27/2018 12:31:19",
      "content": "<p><code>watch -n 0.2 nvidia-smi</code> prevents scrolling. <code>watch</code> works for commands in general (e.g. cat, echo).</p>",
      "rawMarkdown": "`watch -n 0.2 nvidia-smi` prevents scrolling. `watch` works for commands in general (e.g. cat, echo).",
      "votes": null
    },
    {
      "id": "275240",
      "postDate": "01/28/2018 14:52:59",
      "content": "<p>Hi everyone, I personally prefer sth like this: <code>nvidia-smi --query-gpu=index, utilization.gpu --format=csv -l</code>. It gives you the true gpu utilization since running <code>nvidia-smi</code> doesn't show the true gpu utilization in percentage of time over the past sample period during which one or more kernels was executing on the gpu. Another simple alternative for overall system monitoring is using <code>glances</code> which is quite nice as an overall system monitor utility.</p>",
      "rawMarkdown": "Hi everyone, I personally prefer sth like this: `nvidia-smi --query-gpu=index, utilization.gpu --format=csv -l`. It gives you the true gpu utilization since running `nvidia-smi` doesn't show the true gpu utilization in percentage of time over the past sample period during which one or more kernels was executing on the gpu. Another simple alternative for overall system monitoring is using `glances` which is quite nice as an overall system monitor utility.",
      "votes": null
    },
    {
      "id": "275245",
      "postDate": "01/28/2018 15:13:37",
      "content": "<p>Thanks for <code>nvidia-smi --query-gpu=index,utilization.gpu --format=csv -l</code>. Note: I had to remove the space after the comma to get it to display properly.</p>",
      "rawMarkdown": "Thanks for `nvidia-smi --query-gpu=index,utilization.gpu --format=csv -l`. Note: I had to remove the space after the comma to get it to display properly.",
      "votes": null
    },
    {
      "id": "275270",
      "postDate": "01/28/2018 16:11:10",
      "content": "<p>Oh yeah, it also reminds me the kaggle errors that I get when submitting results, somehow if there's any space in submission.csv kaggle fails big time. And then it makes you believe that somehow your submission file is not valid when in reality it only requires removing that extra space :(. It happens to me all  the time.</p>",
      "rawMarkdown": "Oh yeah, it also reminds me the kaggle errors that I get when submitting results, somehow if there's any space in submission.csv kaggle fails big time. And then it makes you believe that somehow your submission file is not valid when in reality it only requires removing that extra space :(. It happens to me all  the time.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 274590,
      "author_name": "ceperaang",
      "author_url": "",
      "post_date": "01/26/2018 20:55:24",
      "content": "<p>Small hint: there is built-in switch to make nvidia-smi repeat its output: add -l seconds or -lms milliseconds. I usually use <code>nvidia-smi -lms 200</code> to monitor GPU usage more granularly. </p>",
      "votes": null,
      "replies": [
        {
          "id": 274593,
          "author_name": "antorsae",
          "author_url": "",
          "post_date": "01/26/2018 20:57:36",
          "content": "<p>Thanks. How to do make it so it stays at the top of the console instead of scrolling indefinitely?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 274601,
          "author_name": "ceperaang",
          "author_url": "",
          "post_date": "01/26/2018 21:26:27",
          "content": "<p>It is scrolling the same amount each update so you see just static current output if console size is big enough :) </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 274818,
          "author_name": "kleinsmith",
          "author_url": "",
          "post_date": "01/27/2018 12:31:19",
          "content": "<p><code>watch -n 0.2 nvidia-smi</code> prevents scrolling. <code>watch</code> works for commands in general (e.g. cat, echo).</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 275240,
      "author_name": "kirk86",
      "author_url": "",
      "post_date": "01/28/2018 14:52:59",
      "content": "<p>Hi everyone, I personally prefer sth like this: <code>nvidia-smi --query-gpu=index, utilization.gpu --format=csv -l</code>. It gives you the true gpu utilization since running <code>nvidia-smi</code> doesn't show the true gpu utilization in percentage of time over the past sample period during which one or more kernels was executing on the gpu. Another simple alternative for overall system monitoring is using <code>glances</code> which is quite nice as an overall system monitor utility.</p>",
      "votes": null,
      "replies": [
        {
          "id": 275245,
          "author_name": "kleinsmith",
          "author_url": "",
          "post_date": "01/28/2018 15:13:37",
          "content": "<p>Thanks for <code>nvidia-smi --query-gpu=index,utilization.gpu --format=csv -l</code>. Note: I had to remove the space after the comma to get it to display properly.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 275270,
          "author_name": "kirk86",
          "author_url": "",
          "post_date": "01/28/2018 16:11:10",
          "content": "<p>Oh yeah, it also reminds me the kaggle errors that I get when submitting results, somehow if there's any space in submission.csv kaggle fails big time. And then it makes you believe that somehow your submission file is not valid when in reality it only requires removing that extra space :(. It happens to me all  the time.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "273843": "One of the things I'm learning on Kaggle competitions is that experimenting speed is key. \n\nAdvice you get from top Kagglers is to first try with reduced complexity networks (in our case simple networks and/or reduced crop size) to get an idea of which hyper-parameters work best (although when you scale up these may/will change); but I also think we need to maximize the code efficiency too.\n\nI often run my training and run `watch -n1 nvidia-smi`. It's a good sign if the GPUs being maxxed out:\n\n![enter image description here][1]\n\nalso running `htop` to see if the CPU cores are being effectively used:\n\n![enter image description here][2]\n\nIn this competition the images are huge so in my first attempt I tried running a simple Keras generator that loaded images and (after processing) sent them to the GPUs. It took roughly ~20+ minutes per epoch (using 2 x 1080  Tis + AMD Threadripper 1950X with 16 cores / 32 threads). Even though Keras generators supposedly support multiprocessing with an argument to `.fit_generator` I was unable to make them work correctly, so my next attempt was to cache images into a dict so epochs after the first one will just take them from the dict.\n\nThe issue is that I cannot keep all uncompressed images in memory (I have 64 Gb RAM but all uncompressed images take more), so I just cached cropped versions of images. While OK for a first try, it's an obvious compromise. Also, trying to saved a pickled version of the dictionary with images ~1024x1024 would result in memory errors (go figure). So another approach was needed:\n\nI then used Python multiprocessing module inside the Keras generator and now the code runs one epoch in ~220 seconds.\n\nWhat code / monitoring tools/tricks do you use to maximize performance?\n\n  [1]: http://i.imgur.com/k6aT8FR.png\n  [2]: http://i.imgur.com/k15JyzW.png",
    "274590": "Small hint: there is built-in switch to make nvidia-smi repeat its output: add -l seconds or -lms milliseconds. I usually use `nvidia-smi -lms 200` to monitor GPU usage more granularly.",
    "274593": "Thanks. How to do make it so it stays at the top of the console instead of scrolling indefinitely?",
    "274601": "It is scrolling the same amount each update so you see just static current output if console size is big enough :)",
    "274818": "`watch -n 0.2 nvidia-smi` prevents scrolling. `watch` works for commands in general (e.g. cat, echo).",
    "275240": "Hi everyone, I personally prefer sth like this: `nvidia-smi --query-gpu=index, utilization.gpu --format=csv -l`. It gives you the true gpu utilization since running `nvidia-smi` doesn't show the true gpu utilization in percentage of time over the past sample period during which one or more kernels was executing on the gpu. Another simple alternative for overall system monitoring is using `glances` which is quite nice as an overall system monitor utility.",
    "275245": "Thanks for `nvidia-smi --query-gpu=index,utilization.gpu --format=csv -l`. Note: I had to remove the space after the comma to get it to display properly.",
    "275270": "Oh yeah, it also reminds me the kaggle errors that I get when submitting results, somehow if there's any space in submission.csv kaggle fails big time. And then it makes you believe that somehow your submission file is not valid when in reality it only requires removing that extra space :(. It happens to me all  the time."
  },
  "source": "meta"
}