{
  "id": 65299,
  "title": "How should we spend our GCP credits?",
  "url": "/competitions/rsna-pneumonia-detection-challenge/discussion/65299",
  "author_name": "",
  "post_date": "2018-09-08T23:47:04.305061300Z",
  "votes": 4,
  "comment_count": 14,
  "views": 0,
  "content": "<p>A VM instance with a GPU costs at least $1/hr. Based on what I've seen so far, I'm not sure how much progress one could make on this project with just 300 hours of GPU time. Is there a more economical way to use Google Cloud for deep learning?</p>",
  "messages": [
    {
      "id": "383541",
      "postDate": "09/08/2018 23:47:04",
      "content": "<p>A VM instance with a GPU costs at least $1/hr. Based on what I've seen so far, I'm not sure how much progress one could make on this project with just 300 hours of GPU time. Is there a more economical way to use Google Cloud for deep learning?</p>",
      "rawMarkdown": "A VM instance with a GPU costs at least $1/hr. Based on what I've seen so far, I'm not sure how much progress one could make on this project with just 300 hours of GPU time. Is there a more economical way to use Google Cloud for deep learning?",
      "votes": null
    },
    {
      "id": "383546",
      "postDate": "09/09/2018 00:06:09",
      "content": "<p>Preemptive instances. They are a nag but much cheaper</p>",
      "rawMarkdown": "Preemptive instances. They are a nag but much cheaper",
      "votes": null
    },
    {
      "id": "383557",
      "postDate": "09/09/2018 00:57:00",
      "content": "<p>@Andy Harless, good question. I would like to know the answer to that too.</p>",
      "rawMarkdown": "Andy Harless, good question. I would like to know the answer to that too.",
      "votes": null
    },
    {
      "id": "383607",
      "postDate": "09/09/2018 04:30:02",
      "content": "<p>You can try older GPUs like the Nvidia K80, they are obviously slower but the difference is probably not that much and they are much cheaper ($0,35/hr in us-west)</p>",
      "rawMarkdown": "You can try older GPUs like the Nvidia K80, they are obviously slower but the difference is probably not that much and they are much cheaper ($0,35/hr in us-west)",
      "votes": null
    },
    {
      "id": "383615",
      "postDate": "09/09/2018 04:53:34",
      "content": "<p>Also, you can always try to grab a gpu instance at colab. Its only 12h but free and if you design your notebook well you can continue from where you stopped. Requires luck (for getting gou backend), some skill (with writing the notebook) and persistence. However, its really great to see if you are heading the right way for free... </p>",
      "rawMarkdown": "Also, you can always try to grab a gpu instance at colab. Its only 12h but free and if you design your notebook well you can continue from where you stopped. Requires luck (for getting gou backend), some skill (with writing the notebook) and persistence. However, its really great to see if you are heading the right way for free...",
      "votes": null
    },
    {
      "id": "383772",
      "postDate": "09/09/2018 15:31:05",
      "content": "<p>Add to Moshel's comment, to make things easy there's <a href=\"https://colab.research.google.com/github/mdai/ml-lessons/blob/master/lesson3-rsna-pneumonia-detection-kaggle.ipynb\">this colab notebook</a> provided by the competition organizers which could be a good place to start.</p>",
      "rawMarkdown": "Add to Moshel's comment, to make things easy there's [this colab notebook][1] provided by the competition organizers which could be a good place to start.\n\n\n  [1]: https://colab.research.google.com/github/mdai/ml-lessons/blob/master/lesson3-rsna-pneumonia-detection-kaggle.ipynb",
      "votes": null
    },
    {
      "id": "383935",
      "postDate": "09/10/2018 01:42:07",
      "content": "<p>I looked at some benchmarks and decided that the K80 isn't necessarily more cost-effective than the P100 for this kind of task.</p>",
      "rawMarkdown": "I looked at some benchmarks and decided that the K80 isn't necessarily more cost-effective than the P100 for this kind of task.",
      "votes": null
    },
    {
      "id": "383944",
      "postDate": "09/10/2018 02:07:25",
      "content": "<p>Well, my suggestion would be to request quota for both. Use k80 (you can even get preemptive k80) to verify that you qre heading in the right way, didn't make any stupid mistake in the bbox conversion etc. Once you see that everything is going smoothly, use the p100. My experience (limited, i admit) with the p100 was disappointing. Note also that some frameworks requires compilation from source to take advantage of the features of the p100 and also you get most benefits from late versions of cuda and cudnn. Anyways, good luck! </p>",
      "rawMarkdown": "Well, my suggestion would be to request quota for both. Use k80 (you can even get preemptive k80) to verify that you qre heading in the right way, didn't make any stupid mistake in the bbox conversion etc. Once you see that everything is going smoothly, use the p100. My experience (limited, i admit) with the p100 was disappointing. Note also that some frameworks requires compilation from source to take advantage of the features of the p100 and also you get most benefits from late versions of cuda and cudnn. Anyways, good luck!",
      "votes": null
    },
    {
      "id": "383981",
      "postDate": "09/10/2018 04:12:48",
      "content": "<p>I'd try to run a small model with both GPUs and also with the V100. If you look at nvidia benchmarks the best choice would actually be the V100 ($1,7/hr):</p>\n\n<p><img src=\"https://www.nvidia.com/content/dam/en-zz/es_em/Solutions/Data-Center/tesla-v100/data-center-tesla-v100-inference-performance-chart-4-18-update-625-ud.png\" alt=\"P100 vs V100 Nvidia site\"></p>",
      "rawMarkdown": "I'd try to run a small model with both GPUs and also with the V100. If you look at nvidia benchmarks the best choice would actually be the V100 ($1,7/hr):\n\n![P100 vs V100 Nvidia site][1]\n\n\n  [1]: https://www.nvidia.com/content/dam/en-zz/es_em/Solutions/Data-Center/tesla-v100/data-center-tesla-v100-inference-performance-chart-4-18-update-625-ud.png",
      "votes": null
    },
    {
      "id": "384042",
      "postDate": "09/10/2018 07:44:33",
      "content": "<p>Mmmm... Things are not as simple as nvidia (naturally, objective) would like you to believe. In order to get this performance, the pipeline must be top performance. That means fast ssd to read the data fast enough to feed it into the gpu, powerful and many core to preprocess and move the data onwards. So, you end up paying much much more. Also, the setup is not trivial and needs tuning.\nNaturally, ymmv and you might get amazing performance on first try.</p>",
      "rawMarkdown": "Mmmm... Things are not as simple as nvidia (naturally, objective) would like you to believe. In order to get this performance, the pipeline must be top performance. That means fast ssd to read the data fast enough to feed it into the gpu, powerful and many core to preprocess and move the data onwards. So, you end up paying much much more. Also, the setup is not trivial and needs tuning.\nNaturally, ymmv and you might get amazing performance on first try.",
      "votes": null
    },
    {
      "id": "385176",
      "postDate": "09/10/2018 13:48:38",
      "content": "<p>I'm wondering if it's a good idea to create a preemptible instance of one of Google's standard deep learning images.</p>",
      "rawMarkdown": "I'm wondering if it's a good idea to create a preemptible instance of one of Google's standard deep learning images.",
      "votes": null
    },
    {
      "id": "387640",
      "postDate": "09/15/2018 10:35:39",
      "content": "<p>For AWS spot instances, there's <a href=\"https://github.com/samuelreh/spotr\">spotr</a> which does the root volume swap for you everytime you start your spot instance with the recent snapshot of the system. For external data like model,train,test one can maintain an external peristent disk. This setup is usually cheaper. Implementing spotr for GCP would be great</p>",
      "rawMarkdown": "For AWS spot instances, there's [spotr][1] which does the root volume swap for you everytime you start your spot instance with the recent snapshot of the system. For external data like model,train,test one can maintain an external peristent disk. This setup is usually cheaper. Implementing spotr for GCP would be great\n\n\n  [1]: https://github.com/samuelreh/spotr",
      "votes": null
    },
    {
      "id": "387655",
      "postDate": "09/15/2018 11:02:27",
      "content": "<p>It seems that the main cost is GPU. And it doesn't depend on a preemptive/normal instance, right?</p>",
      "rawMarkdown": "It seems that the main cost is GPU. And it doesn't depend on a preemptive/normal instance, right?",
      "votes": null
    },
    {
      "id": "387726",
      "postDate": "09/15/2018 14:50:50",
      "content": "<p>@Sergey It does. GCP offers Pre-emptied and normal GPU instances.</p>",
      "rawMarkdown": "Sergey It does. GCP offers Pre-emptied and normal GPU instances.",
      "votes": null
    },
    {
      "id": "393554",
      "postDate": "09/25/2018 15:05:20",
      "content": "<p>My performance comparison (using a model similar to Jonne's model, with 320x320 input, batch size 16):</p>\n\n<pre><code>V100 ~530 sec/epoch\nP100 ~640 sec/epoch\nK80 ~1080 sec/epoch\n</code></pre>\n\n<p>(With each of them, the first epoch takes longer if it's reading from a hard disk, about the same if from a SSD.)  Given the price comparison, the risk of additional costs imposed by preemption of a slow run, and the general advantage of getting things done quickly, P100 looks like the best choice.  Not sure why the V100 looks so good on nVidia's benchmark.  I might use a K80 for overnight runs if I'm going to be asleep and won't have a chance to shut down the instance when it's done.  I increased my quota to 2 GPUs, but I've discovered it's quite feasible to create and delete instances on the fly (by running a custom setup script and attaching to a shared read-only disk)—and it's necessary to do so if using a local SSD without running continuously—so the quota isn't a big deal unless I want to run several things at once.</p>",
      "rawMarkdown": "My performance comparison (using a model similar to Jonne's model, with 320x320 input, batch size 16):\n\n    V100 ~530 sec/epoch\n    P100 ~640 sec/epoch\n    K80 ~1080 sec/epoch\n\n(With each of them, the first epoch takes longer if it's reading from a hard disk, about the same if from a SSD.)  Given the price comparison, the risk of additional costs imposed by preemption of a slow run, and the general advantage of getting things done quickly, P100 looks like the best choice.  Not sure why the V100 looks so good on nVidia's benchmark.  I might use a K80 for overnight runs if I'm going to be asleep and won't have a chance to shut down the instance when it's done.  I increased my quota to 2 GPUs, but I've discovered it's quite feasible to create and delete instances on the fly (by running a custom setup script and attaching to a shared read-only disk)—and it's necessary to do so if using a local SSD without running continuously—so the quota isn't a big deal unless I want to run several things at once.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 383546,
      "author_name": "moshel",
      "author_url": "",
      "post_date": "09/09/2018 00:06:09",
      "content": "<p>Preemptive instances. They are a nag but much cheaper</p>",
      "votes": null,
      "replies": [
        {
          "id": 387655,
          "author_name": "sergeyzlobin",
          "author_url": "",
          "post_date": "09/15/2018 11:02:27",
          "content": "<p>It seems that the main cost is GPU. And it doesn't depend on a preemptive/normal instance, right?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 387726,
          "author_name": "trekkerthemaker",
          "author_url": "",
          "post_date": "09/15/2018 14:50:50",
          "content": "<p>@Sergey It does. GCP offers Pre-emptied and normal GPU instances.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 383557,
      "author_name": "sheriytm",
      "author_url": "",
      "post_date": "09/09/2018 00:57:00",
      "content": "<p>@Andy Harless, good question. I would like to know the answer to that too.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 383607,
      "author_name": "jsaguiar",
      "author_url": "",
      "post_date": "09/09/2018 04:30:02",
      "content": "<p>You can try older GPUs like the Nvidia K80, they are obviously slower but the difference is probably not that much and they are much cheaper ($0,35/hr in us-west)</p>",
      "votes": null,
      "replies": [
        {
          "id": 383615,
          "author_name": "moshel",
          "author_url": "",
          "post_date": "09/09/2018 04:53:34",
          "content": "<p>Also, you can always try to grab a gpu instance at colab. Its only 12h but free and if you design your notebook well you can continue from where you stopped. Requires luck (for getting gou backend), some skill (with writing the notebook) and persistence. However, its really great to see if you are heading the right way for free... </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 383772,
          "author_name": "keshan",
          "author_url": "",
          "post_date": "09/09/2018 15:31:05",
          "content": "<p>Add to Moshel's comment, to make things easy there's <a href=\"https://colab.research.google.com/github/mdai/ml-lessons/blob/master/lesson3-rsna-pneumonia-detection-kaggle.ipynb\">this colab notebook</a> provided by the competition organizers which could be a good place to start.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 383935,
          "author_name": "aharless",
          "author_url": "",
          "post_date": "09/10/2018 01:42:07",
          "content": "<p>I looked at some benchmarks and decided that the K80 isn't necessarily more cost-effective than the P100 for this kind of task.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 383944,
          "author_name": "moshel",
          "author_url": "",
          "post_date": "09/10/2018 02:07:25",
          "content": "<p>Well, my suggestion would be to request quota for both. Use k80 (you can even get preemptive k80) to verify that you qre heading in the right way, didn't make any stupid mistake in the bbox conversion etc. Once you see that everything is going smoothly, use the p100. My experience (limited, i admit) with the p100 was disappointing. Note also that some frameworks requires compilation from source to take advantage of the features of the p100 and also you get most benefits from late versions of cuda and cudnn. Anyways, good luck! </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 383981,
          "author_name": "jsaguiar",
          "author_url": "",
          "post_date": "09/10/2018 04:12:48",
          "content": "<p>I'd try to run a small model with both GPUs and also with the V100. If you look at nvidia benchmarks the best choice would actually be the V100 ($1,7/hr):</p>\n\n<p><img src=\"https://www.nvidia.com/content/dam/en-zz/es_em/Solutions/Data-Center/tesla-v100/data-center-tesla-v100-inference-performance-chart-4-18-update-625-ud.png\" alt=\"P100 vs V100 Nvidia site\"></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 384042,
          "author_name": "moshel",
          "author_url": "",
          "post_date": "09/10/2018 07:44:33",
          "content": "<p>Mmmm... Things are not as simple as nvidia (naturally, objective) would like you to believe. In order to get this performance, the pipeline must be top performance. That means fast ssd to read the data fast enough to feed it into the gpu, powerful and many core to preprocess and move the data onwards. So, you end up paying much much more. Also, the setup is not trivial and needs tuning.\nNaturally, ymmv and you might get amazing performance on first try.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 385176,
      "author_name": "aharless",
      "author_url": "",
      "post_date": "09/10/2018 13:48:38",
      "content": "<p>I'm wondering if it's a good idea to create a preemptible instance of one of Google's standard deep learning images.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 387640,
      "author_name": "trekkerthemaker",
      "author_url": "",
      "post_date": "09/15/2018 10:35:39",
      "content": "<p>For AWS spot instances, there's <a href=\"https://github.com/samuelreh/spotr\">spotr</a> which does the root volume swap for you everytime you start your spot instance with the recent snapshot of the system. For external data like model,train,test one can maintain an external peristent disk. This setup is usually cheaper. Implementing spotr for GCP would be great</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 393554,
      "author_name": "aharless",
      "author_url": "",
      "post_date": "09/25/2018 15:05:20",
      "content": "<p>My performance comparison (using a model similar to Jonne's model, with 320x320 input, batch size 16):</p>\n\n<pre><code>V100 ~530 sec/epoch\nP100 ~640 sec/epoch\nK80 ~1080 sec/epoch\n</code></pre>\n\n<p>(With each of them, the first epoch takes longer if it's reading from a hard disk, about the same if from a SSD.)  Given the price comparison, the risk of additional costs imposed by preemption of a slow run, and the general advantage of getting things done quickly, P100 looks like the best choice.  Not sure why the V100 looks so good on nVidia's benchmark.  I might use a K80 for overnight runs if I'm going to be asleep and won't have a chance to shut down the instance when it's done.  I increased my quota to 2 GPUs, but I've discovered it's quite feasible to create and delete instances on the fly (by running a custom setup script and attaching to a shared read-only disk)—and it's necessary to do so if using a local SSD without running continuously—so the quota isn't a big deal unless I want to run several things at once.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "383541": "A VM instance with a GPU costs at least $1/hr. Based on what I've seen so far, I'm not sure how much progress one could make on this project with just 300 hours of GPU time. Is there a more economical way to use Google Cloud for deep learning?",
    "383546": "Preemptive instances. They are a nag but much cheaper",
    "383557": "Andy Harless, good question. I would like to know the answer to that too.",
    "383607": "You can try older GPUs like the Nvidia K80, they are obviously slower but the difference is probably not that much and they are much cheaper ($0,35/hr in us-west)",
    "383615": "Also, you can always try to grab a gpu instance at colab. Its only 12h but free and if you design your notebook well you can continue from where you stopped. Requires luck (for getting gou backend), some skill (with writing the notebook) and persistence. However, its really great to see if you are heading the right way for free...",
    "383772": "Add to Moshel's comment, to make things easy there's [this colab notebook][1] provided by the competition organizers which could be a good place to start.\n\n\n  [1]: https://colab.research.google.com/github/mdai/ml-lessons/blob/master/lesson3-rsna-pneumonia-detection-kaggle.ipynb",
    "383935": "I looked at some benchmarks and decided that the K80 isn't necessarily more cost-effective than the P100 for this kind of task.",
    "383944": "Well, my suggestion would be to request quota for both. Use k80 (you can even get preemptive k80) to verify that you qre heading in the right way, didn't make any stupid mistake in the bbox conversion etc. Once you see that everything is going smoothly, use the p100. My experience (limited, i admit) with the p100 was disappointing. Note also that some frameworks requires compilation from source to take advantage of the features of the p100 and also you get most benefits from late versions of cuda and cudnn. Anyways, good luck!",
    "383981": "I'd try to run a small model with both GPUs and also with the V100. If you look at nvidia benchmarks the best choice would actually be the V100 ($1,7/hr):\n\n![P100 vs V100 Nvidia site][1]\n\n\n  [1]: https://www.nvidia.com/content/dam/en-zz/es_em/Solutions/Data-Center/tesla-v100/data-center-tesla-v100-inference-performance-chart-4-18-update-625-ud.png",
    "384042": "Mmmm... Things are not as simple as nvidia (naturally, objective) would like you to believe. In order to get this performance, the pipeline must be top performance. That means fast ssd to read the data fast enough to feed it into the gpu, powerful and many core to preprocess and move the data onwards. So, you end up paying much much more. Also, the setup is not trivial and needs tuning.\nNaturally, ymmv and you might get amazing performance on first try.",
    "385176": "I'm wondering if it's a good idea to create a preemptible instance of one of Google's standard deep learning images.",
    "387640": "For AWS spot instances, there's [spotr][1] which does the root volume swap for you everytime you start your spot instance with the recent snapshot of the system. For external data like model,train,test one can maintain an external peristent disk. This setup is usually cheaper. Implementing spotr for GCP would be great\n\n\n  [1]: https://github.com/samuelreh/spotr",
    "387655": "It seems that the main cost is GPU. And it doesn't depend on a preemptive/normal instance, right?",
    "387726": "Sergey It does. GCP offers Pre-emptied and normal GPU instances.",
    "393554": "My performance comparison (using a model similar to Jonne's model, with 320x320 input, batch size 16):\n\n    V100 ~530 sec/epoch\n    P100 ~640 sec/epoch\n    K80 ~1080 sec/epoch\n\n(With each of them, the first epoch takes longer if it's reading from a hard disk, about the same if from a SSD.)  Given the price comparison, the risk of additional costs imposed by preemption of a slow run, and the general advantage of getting things done quickly, P100 looks like the best choice.  Not sure why the V100 looks so good on nVidia's benchmark.  I might use a K80 for overnight runs if I'm going to be asleep and won't have a chance to shut down the instance when it's done.  I increased my quota to 2 GPUs, but I've discovered it's quite feasible to create and delete instances on the fly (by running a custom setup script and attaching to a shared read-only disk)—and it's necessary to do so if using a local SSD without running continuously—so the quota isn't a big deal unless I want to run several things at once."
  },
  "source": "meta"
}