{
  "id": 42107,
  "title": "Trying to train inception_resnet_v2 on aws",
  "url": "/competitions/cdiscount-image-classification-challenge/discussion/42107",
  "author_name": "",
  "post_date": "2017-10-26T19:00:40.655287800Z",
  "votes": null,
  "comment_count": 3,
  "views": 0,
  "content": "<p>Hi,</p>\n\n<p>It's my first deep learning project with such a big dataset, and I'd like to play with tensorflow, but with a single aws instance with Tesla K80.</p>\n\n<p>Can someone please tell me the expected speed of such GPU when training inception_resnet_v2 ?</p>\n\n<p>I just started an experiment and I am having 25 images/second.</p>\n\n<p>Is it ok or am I missing something ?</p>\n\n<p>Thanks a lot for sharing your thoughts and experience !</p>\n\n<p>Looks like CPU computations are not optimized but GPU seems up and running. Here's what appears on tensorflow log on startup :</p>\n\n<p>2017-10-26 17:51:20.832551: W tensorflow/core/platform/cpu_feature_guard.cc:45] The TensorFlow library wasn't compiled to use SSE4.1 instructions, but these are available on your machine and could speed up CPU computations.\n2017-10-26 17:51:20.832580: W tensorflow/core/platform/cpu_feature_guard.cc:45] The TensorFlow library wasn't compiled to use SSE4.2 instructions, but these are available on your machine and could speed up CPU computations.\n2017-10-26 17:51:20.832586: W tensorflow/core/platform/cpu_feature_guard.cc:45] The TensorFlow library wasn't compiled to use AVX instructions, but these are available on your machine and could speed up CPU computations.\n2017-10-26 17:51:20.832591: W tensorflow/core/platform/cpu_feature_guard.cc:45] The TensorFlow library wasn't compiled to use AVX2 instructions, but these are available on your machine and could speed up CPU computations.\n2017-10-26 17:51:20.832595: W tensorflow/core/platform/cpu_feature_guard.cc:45] The TensorFlow library wasn't compiled to use FMA instructions, but these are available on your machine and could speed up CPU computations.\n2017-10-26 17:51:21.097779: I tensorflow/stream_executor/cuda/cuda_gpu_executor.cc:893] successful NUMA node read from SysFS had negative value (-1), but there must be at least one NUMA node, so returning NUMA node zero\n2017-10-26 17:51:21.098294: I tensorflow/core/common_runtime/gpu/gpu_device.cc:955] Found device 0 with properties: \nname: Tesla K80\nmajor: 3 minor: 7 memoryClockRate (GHz) 0.8235\npciBusID 0000:00:1e.0\nTotal memory: 11.17GiB\nFree memory: 11.11GiB\n2017-10-26 17:51:21.098323: I tensorflow/core/common_runtime/gpu/gpu_device.cc:976] DMA: 0 \n2017-10-26 17:51:21.098334: I tensorflow/core/common_runtime/gpu/gpu_device.cc:986] 0:   Y \n2017-10-26 17:51:21.098346: I tensorflow/core/common_runtime/gpu/gpu_device.cc:1045] Creating TensorFlow device (/gpu:0) -&gt; (device: 0, name: Tesla K80, pci bus id: 0000:00:1e.0)</p>",
  "messages": [
    {
      "id": "236155",
      "postDate": "10/26/2017 19:00:40",
      "content": "<p>Hi,</p>\n\n<p>It's my first deep learning project with such a big dataset, and I'd like to play with tensorflow, but with a single aws instance with Tesla K80.</p>\n\n<p>Can someone please tell me the expected speed of such GPU when training inception_resnet_v2 ?</p>\n\n<p>I just started an experiment and I am having 25 images/second.</p>\n\n<p>Is it ok or am I missing something ?</p>\n\n<p>Thanks a lot for sharing your thoughts and experience !</p>\n\n<p>Looks like CPU computations are not optimized but GPU seems up and running. Here's what appears on tensorflow log on startup :</p>\n\n<p>2017-10-26 17:51:20.832551: W tensorflow/core/platform/cpu_feature_guard.cc:45] The TensorFlow library wasn't compiled to use SSE4.1 instructions, but these are available on your machine and could speed up CPU computations.\n2017-10-26 17:51:20.832580: W tensorflow/core/platform/cpu_feature_guard.cc:45] The TensorFlow library wasn't compiled to use SSE4.2 instructions, but these are available on your machine and could speed up CPU computations.\n2017-10-26 17:51:20.832586: W tensorflow/core/platform/cpu_feature_guard.cc:45] The TensorFlow library wasn't compiled to use AVX instructions, but these are available on your machine and could speed up CPU computations.\n2017-10-26 17:51:20.832591: W tensorflow/core/platform/cpu_feature_guard.cc:45] The TensorFlow library wasn't compiled to use AVX2 instructions, but these are available on your machine and could speed up CPU computations.\n2017-10-26 17:51:20.832595: W tensorflow/core/platform/cpu_feature_guard.cc:45] The TensorFlow library wasn't compiled to use FMA instructions, but these are available on your machine and could speed up CPU computations.\n2017-10-26 17:51:21.097779: I tensorflow/stream_executor/cuda/cuda_gpu_executor.cc:893] successful NUMA node read from SysFS had negative value (-1), but there must be at least one NUMA node, so returning NUMA node zero\n2017-10-26 17:51:21.098294: I tensorflow/core/common_runtime/gpu/gpu_device.cc:955] Found device 0 with properties: \nname: Tesla K80\nmajor: 3 minor: 7 memoryClockRate (GHz) 0.8235\npciBusID 0000:00:1e.0\nTotal memory: 11.17GiB\nFree memory: 11.11GiB\n2017-10-26 17:51:21.098323: I tensorflow/core/common_runtime/gpu/gpu_device.cc:976] DMA: 0 \n2017-10-26 17:51:21.098334: I tensorflow/core/common_runtime/gpu/gpu_device.cc:986] 0:   Y \n2017-10-26 17:51:21.098346: I tensorflow/core/common_runtime/gpu/gpu_device.cc:1045] Creating TensorFlow device (/gpu:0) -&gt; (device: 0, name: Tesla K80, pci bus id: 0000:00:1e.0)</p>",
      "rawMarkdown": "Hi,\n\nIt's my first deep learning project with such a big dataset, and I'd like to play with tensorflow, but with a single aws instance with Tesla K80.\n\nCan someone please tell me the expected speed of such GPU when training inception_resnet_v2 ?\n\nI just started an experiment and I am having 25 images/second.\n\nIs it ok or am I missing something ?\n\nThanks a lot for sharing your thoughts and experience !\n\n\nLooks like CPU computations are not optimized but GPU seems up and running. Here's what appears on tensorflow log on startup :\n\n2017-10-26 17:51:20.832551: W tensorflow/core/platform/cpu_feature_guard.cc:45] The TensorFlow library wasn't compiled to use SSE4.1 instructions, but these are available on your machine and could speed up CPU computations.\n2017-10-26 17:51:20.832580: W tensorflow/core/platform/cpu_feature_guard.cc:45] The TensorFlow library wasn't compiled to use SSE4.2 instructions, but these are available on your machine and could speed up CPU computations.\n2017-10-26 17:51:20.832586: W tensorflow/core/platform/cpu_feature_guard.cc:45] The TensorFlow library wasn't compiled to use AVX instructions, but these are available on your machine and could speed up CPU computations.\n2017-10-26 17:51:20.832591: W tensorflow/core/platform/cpu_feature_guard.cc:45] The TensorFlow library wasn't compiled to use AVX2 instructions, but these are available on your machine and could speed up CPU computations.\n2017-10-26 17:51:20.832595: W tensorflow/core/platform/cpu_feature_guard.cc:45] The TensorFlow library wasn't compiled to use FMA instructions, but these are available on your machine and could speed up CPU computations.\n2017-10-26 17:51:21.097779: I tensorflow/stream_executor/cuda/cuda_gpu_executor.cc:893] successful NUMA node read from SysFS had negative value (-1), but there must be at least one NUMA node, so returning NUMA node zero\n2017-10-26 17:51:21.098294: I tensorflow/core/common_runtime/gpu/gpu_device.cc:955] Found device 0 with properties: \nname: Tesla K80\nmajor: 3 minor: 7 memoryClockRate (GHz) 0.8235\npciBusID 0000:00:1e.0\nTotal memory: 11.17GiB\nFree memory: 11.11GiB\n2017-10-26 17:51:21.098323: I tensorflow/core/common_runtime/gpu/gpu_device.cc:976] DMA: 0 \n2017-10-26 17:51:21.098334: I tensorflow/core/common_runtime/gpu/gpu_device.cc:986] 0:   Y \n2017-10-26 17:51:21.098346: I tensorflow/core/common_runtime/gpu/gpu_device.cc:1045] Creating TensorFlow device (/gpu:0) -&gt; (device: 0, name: Tesla K80, pci bus id: 0000:00:1e.0)",
      "votes": null
    },
    {
      "id": "236165",
      "postDate": "10/26/2017 19:15:35",
      "content": "<p>I can do around 63 images/second using GTX 1070, takes more than 2 days for a single epoch. K80 is much slower than GTX 1070 and AWS IO performance is worse than having your own SSD. You can try the newly introduced p3 instances which have the much faster Volta series GPUs from today <a href=\"https://aws.amazon.com/about-aws/whats-new/2017/10/introducing-amazon-ec2-p3-instances/\">https://aws.amazon.com/about-aws/whats-new/2017/10/introducing-amazon-ec2-p3-instances/</a>  You might even experiment with fp16 training which can give you training time improvement since the V100 have native support for them.</p>",
      "rawMarkdown": "I can do around 63 images/second using GTX 1070, takes more than 2 days for a single epoch. K80 is much slower than GTX 1070 and AWS IO performance is worse than having your own SSD. You can try the newly introduced p3 instances which have the much faster Volta series GPUs from today https://aws.amazon.com/about-aws/whats-new/2017/10/introducing-amazon-ec2-p3-instances/  You might even experiment with fp16 training which can give you training time improvement since the V100 have native support for them.",
      "votes": null
    },
    {
      "id": "236194",
      "postDate": "10/26/2017 20:17:45",
      "content": "<p>Tks for the tip !</p>",
      "rawMarkdown": "Tks for the tip !",
      "votes": null
    },
    {
      "id": "236386",
      "postDate": "10/27/2017 07:10:56",
      "content": "<p>We can use 8 GPUs on p3.16 large. Using more GPUs is great advantage in this competition.\n<a href=\"https://aws.amazon.com/jp/blogs/aws/new-amazon-ec2-instances-with-up-to-8-nvidia-tesla-v100-gpus-p3/\">https://aws.amazon.com/jp/blogs/aws/new-amazon-ec2-instances-with-up-to-8-nvidia-tesla-v100-gpus-p3/</a></p>",
      "rawMarkdown": "We can use 8 GPUs on p3.16 large. Using more GPUs is great advantage in this competition.\nhttps://aws.amazon.com/jp/blogs/aws/new-amazon-ec2-instances-with-up-to-8-nvidia-tesla-v100-gpus-p3/",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 236165,
      "author_name": "datasec",
      "author_url": "",
      "post_date": "10/26/2017 19:15:35",
      "content": "<p>I can do around 63 images/second using GTX 1070, takes more than 2 days for a single epoch. K80 is much slower than GTX 1070 and AWS IO performance is worse than having your own SSD. You can try the newly introduced p3 instances which have the much faster Volta series GPUs from today <a href=\"https://aws.amazon.com/about-aws/whats-new/2017/10/introducing-amazon-ec2-p3-instances/\">https://aws.amazon.com/about-aws/whats-new/2017/10/introducing-amazon-ec2-p3-instances/</a>  You might even experiment with fp16 training which can give you training time improvement since the V100 have native support for them.</p>",
      "votes": null,
      "replies": [
        {
          "id": 236194,
          "author_name": "randombishop",
          "author_url": "",
          "post_date": "10/26/2017 20:17:45",
          "content": "<p>Tks for the tip !</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 236386,
          "author_name": "toshik",
          "author_url": "",
          "post_date": "10/27/2017 07:10:56",
          "content": "<p>We can use 8 GPUs on p3.16 large. Using more GPUs is great advantage in this competition.\n<a href=\"https://aws.amazon.com/jp/blogs/aws/new-amazon-ec2-instances-with-up-to-8-nvidia-tesla-v100-gpus-p3/\">https://aws.amazon.com/jp/blogs/aws/new-amazon-ec2-instances-with-up-to-8-nvidia-tesla-v100-gpus-p3/</a></p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "236155": "Hi,\n\nIt's my first deep learning project with such a big dataset, and I'd like to play with tensorflow, but with a single aws instance with Tesla K80.\n\nCan someone please tell me the expected speed of such GPU when training inception_resnet_v2 ?\n\nI just started an experiment and I am having 25 images/second.\n\nIs it ok or am I missing something ?\n\nThanks a lot for sharing your thoughts and experience !\n\n\nLooks like CPU computations are not optimized but GPU seems up and running. Here's what appears on tensorflow log on startup :\n\n2017-10-26 17:51:20.832551: W tensorflow/core/platform/cpu_feature_guard.cc:45] The TensorFlow library wasn't compiled to use SSE4.1 instructions, but these are available on your machine and could speed up CPU computations.\n2017-10-26 17:51:20.832580: W tensorflow/core/platform/cpu_feature_guard.cc:45] The TensorFlow library wasn't compiled to use SSE4.2 instructions, but these are available on your machine and could speed up CPU computations.\n2017-10-26 17:51:20.832586: W tensorflow/core/platform/cpu_feature_guard.cc:45] The TensorFlow library wasn't compiled to use AVX instructions, but these are available on your machine and could speed up CPU computations.\n2017-10-26 17:51:20.832591: W tensorflow/core/platform/cpu_feature_guard.cc:45] The TensorFlow library wasn't compiled to use AVX2 instructions, but these are available on your machine and could speed up CPU computations.\n2017-10-26 17:51:20.832595: W tensorflow/core/platform/cpu_feature_guard.cc:45] The TensorFlow library wasn't compiled to use FMA instructions, but these are available on your machine and could speed up CPU computations.\n2017-10-26 17:51:21.097779: I tensorflow/stream_executor/cuda/cuda_gpu_executor.cc:893] successful NUMA node read from SysFS had negative value (-1), but there must be at least one NUMA node, so returning NUMA node zero\n2017-10-26 17:51:21.098294: I tensorflow/core/common_runtime/gpu/gpu_device.cc:955] Found device 0 with properties: \nname: Tesla K80\nmajor: 3 minor: 7 memoryClockRate (GHz) 0.8235\npciBusID 0000:00:1e.0\nTotal memory: 11.17GiB\nFree memory: 11.11GiB\n2017-10-26 17:51:21.098323: I tensorflow/core/common_runtime/gpu/gpu_device.cc:976] DMA: 0 \n2017-10-26 17:51:21.098334: I tensorflow/core/common_runtime/gpu/gpu_device.cc:986] 0:   Y \n2017-10-26 17:51:21.098346: I tensorflow/core/common_runtime/gpu/gpu_device.cc:1045] Creating TensorFlow device (/gpu:0) -&gt; (device: 0, name: Tesla K80, pci bus id: 0000:00:1e.0)",
    "236165": "I can do around 63 images/second using GTX 1070, takes more than 2 days for a single epoch. K80 is much slower than GTX 1070 and AWS IO performance is worse than having your own SSD. You can try the newly introduced p3 instances which have the much faster Volta series GPUs from today https://aws.amazon.com/about-aws/whats-new/2017/10/introducing-amazon-ec2-p3-instances/  You might even experiment with fp16 training which can give you training time improvement since the V100 have native support for them.",
    "236194": "Tks for the tip !",
    "236386": "We can use 8 GPUs on p3.16 large. Using more GPUs is great advantage in this competition.\nhttps://aws.amazon.com/jp/blogs/aws/new-amazon-ec2-instances-with-up-to-8-nvidia-tesla-v100-gpus-p3/"
  },
  "source": "meta"
}