{
  "id": 45487,
  "title": "My workflow",
  "url": "/competitions/passenger-screening-algorithm-challenge/discussion/45487",
  "author_name": "",
  "post_date": "2017-12-12T00:17:12.916260Z",
  "votes": 8,
  "comment_count": 10,
  "views": 0,
  "content": "<p>Making neural nets, learning them, testing, analyzing the results, analyzing the data and going back. Cycling through this many times and learn (your own head here) is at the basis of creating successful entries. It’s a corollary of the <a href=\"https://i.pinimg.com/originals/2c/68/f6/2c68f617f7cf8ef8e84e3c5b22873d9f.jpg\">fail fast</a> proposition.</p>\n\n<p>To fail fast, it’s worth putting some thought into your workflow, making sure none of your tools is preventing you from failing, or better said from finding out where you got it wrong. I put some thought into my workflow and decided to share my findings, maybe it inspires some of you or maybe it breaks the ice and gets you to share your practices. And not to the least, I’m waiting anxiously for the stage 2 test data to be released and thus had some time at my hands!</p>\n\n<h2>Server less</h2>\n\n<p>I don’t have a computer, I actually don’t like to have humming heating machines in my home. As a kid I made my own computers from parts, I soldered my own extensions to PCI slots, I hacked Modems from tape-recorders and setup Fidonet network branches using VHF radio. I’m no stranger to the smell of heating and outgassing electronic parts. But now I’m done with it all. I removed all hardware from my life and only have a clean minimalistic laptop on an empty wooden table in an empty room (I do allow for a cup of coffee occasionally).</p>\n\n<p>So the cloud it is, it has to be. Originally I was planning on using a VM on Google or Amazon, install Tensorflow, other packages needed for my scripts and some method to sync the code between my laptop and the VM. While this is not too difficult to setup there are a few details to be worked out. I browsed around and stumbled upon a new paradigm, called “server less”.</p>\n\n<blockquote>\n  <h3>Server:</h3>\n  \n  <p>A machine you own or rent, is dusted off by you or someone you\n  pay and can be connected to either remotely or directly.</p>\n  \n  <h3>Cloud or VM:</h3>\n  \n  <p>The look and feel of a server, but its physical location is up to the\n  provider, it may be a small part of a bigger computer, it may be\n  relocated it may even be spread out over multiple physical computers\n  and geographic regions.</p>\n  \n  <h3>Server less:</h3>\n  \n  <p>The look and feel of a webpage,\n  not a server. No command line, no setting up firewalls and configuring\n  basic SSH or HTTP service. It will just run your scripts.</p>\n</blockquote>\n\n<p>So why bother with anything but server less? There are a few good reasons thinkable, vendor lock-in is something I’d worry about. Can I go to a different provider with my code? Can I find a job, with my acquired skills, at a company that uses a different provider? Or will I be forever enslaved to a single super-power?</p>\n\n<p>Google Cloud ML is such a “server less” service offered specifically for running tensorflow scripts on a single machine or cluster of machines with GPU’s. The service has most standard Python packages available and most others can be installed with pip install. That makes the platform fairly open and there isn’t much “skills” to acquire, it just works.</p>\n\n<p>The fact that the TSA data already resided on a Google Storage Bucket and a $500 credit from the Google marketing department sweetened the deal to go and embrace the “server less” paradigm. </p>\n\n<h2>Job submission</h2>\n\n<p>For Cloud ML you prepare your Python Tensorflow script locally, may test-run it locally and eventually you’ll post it on the cloud engine to run it. You configure the type of machine or cluster of machines you’d like to use in a config.yaml file, number of CPU’s and GPU’s amount of RAM, number of machines etc. Finally you post the job with a simple command, like so:</p>\n\n<pre>    gcloud ml-engine jobs submit training \"jobname\" \\\n        --scale-tier CUSTOM \\\n        --config config.yaml \\\n        --module-name train.train \\\n        --package-path train \\\n        --runtime-version=1.2 \\\n</pre>\n\n<p>This will package your script locally, stage the whole package on your cloud storage bucket and launch the job with Cloud ML. For the jobname I use a bash script to automatically generate 5 letter names in alphabetical order, everything remains ordered under those jobnames.</p>\n\n<p>In my version of this I used a global DEBUG flag to essentially create two versions of my script: a simple version that can run (and finish) on my laptop and the real one. This way I could quickly test new development and don’t need to wait for Cloud ML response just to discover a simple syntax error.</p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/256451/8055/jobs.png\" alt=\"Cloud ML organizes its task by jobname\" title=\"\"></p>\n\n<p>Then on a productive Sunday afternoon, when multiple ideas come to mind or when you want to try multiple variations of the same idea, multiple jobs can easily be launched concurrently. Close your laptop and check back in later when they are done. You don’t need to bother spinning up multiple VM’s and you don’t need to make sure they get ended. If you don’t have any inspiration, simply don’t submit any jobs, no worries “should I keep the VM for next time? Or should I park it and save cost?”</p>\n\n<p>Cloud ML has a special offering for inference, where you pay per infer I believe. While this may turn out to be very power full in online high throughput high demand kind of situations I didn’t find it too useful. For the once in a lifetime type of infer that we need in this competition it isn’t worth all the extra effort to setup an online system. Instead I run infer jobs just like training jobs, just this time the input is the weights and the output is a Kaggle submission file.</p>\n\n<h2>Cloud storage</h2>\n\n<p>The inputs and outputs of your script are typically to a cloud storage bucket. When the training data is very large, such as with this competition, it is convenient -necessary even- to not need to download all the data. Besides the actual training job you could also do data preprocessing jobs and store the processed data back onto the cloud bucket.</p>\n\n<p>The results get stored on the cloud as well, the time traces, the weights, the evaluation cross-entropy etc.. You can download them using simple command line tools provided through <a href=\"https://cloud.google.com/storage/docs/gsutil\">gsutils</a>. But it’s not even needed to download the results. The weights can be re-loaded by your infer script job or by your training continuation jobs.</p>\n\n<p>For the time traces of cross entropy, weight sparsity, bias distributions etc I used Tensorboard.</p>\n\n<h2>Tensorboard</h2>\n\n<p>Tensorboard can be started from the cloud shell, directly pointing to the storage bucket containing the training transients. Doing so provides a nice and direct way to review the training process while it is working. Sorted by the same jobnames as the training jobs themselves. Using the tensorboard build in Regex filter we can slice through the data in virtually any way thinkable.</p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/256451/8056/tb.png\" alt=\"Straight from the cloud storage, web based results viewing\" title=\"\"></p>\n\n<h2>Stackdriver</h2>\n\n<p>Finally, I used stackdriver extensively. Log messages generated in your scripts and send to <code>tf.logging.info(...)</code> are routed straight to the online log message review system. It’s the same as wandering through your plain text logging transcripts, but online instead. The lines are folded to one-liners, everything is searchable and there are some really good slicing methods available.</p>\n\n<p>All together, I have managed to stay away from any overhead work related to setting up servers and further did not need to download most of the data (just a small sample for testing purposes). I believe this saved a fair amount of time as there is always more hacking involved in setting up the new server with all the communication channels than I anticipate at the start of such a project. But who knows, maybe that became really easy as well,..</p>",
  "messages": [
    {
      "id": "256451",
      "postDate": "12/12/2017 00:17:12",
      "content": "<p>Making neural nets, learning them, testing, analyzing the results, analyzing the data and going back. Cycling through this many times and learn (your own head here) is at the basis of creating successful entries. It’s a corollary of the <a href=\"https://i.pinimg.com/originals/2c/68/f6/2c68f617f7cf8ef8e84e3c5b22873d9f.jpg\">fail fast</a> proposition.</p>\n\n<p>To fail fast, it’s worth putting some thought into your workflow, making sure none of your tools is preventing you from failing, or better said from finding out where you got it wrong. I put some thought into my workflow and decided to share my findings, maybe it inspires some of you or maybe it breaks the ice and gets you to share your practices. And not to the least, I’m waiting anxiously for the stage 2 test data to be released and thus had some time at my hands!</p>\n\n<h2>Server less</h2>\n\n<p>I don’t have a computer, I actually don’t like to have humming heating machines in my home. As a kid I made my own computers from parts, I soldered my own extensions to PCI slots, I hacked Modems from tape-recorders and setup Fidonet network branches using VHF radio. I’m no stranger to the smell of heating and outgassing electronic parts. But now I’m done with it all. I removed all hardware from my life and only have a clean minimalistic laptop on an empty wooden table in an empty room (I do allow for a cup of coffee occasionally).</p>\n\n<p>So the cloud it is, it has to be. Originally I was planning on using a VM on Google or Amazon, install Tensorflow, other packages needed for my scripts and some method to sync the code between my laptop and the VM. While this is not too difficult to setup there are a few details to be worked out. I browsed around and stumbled upon a new paradigm, called “server less”.</p>\n\n<blockquote>\n  <h3>Server:</h3>\n  \n  <p>A machine you own or rent, is dusted off by you or someone you\n  pay and can be connected to either remotely or directly.</p>\n  \n  <h3>Cloud or VM:</h3>\n  \n  <p>The look and feel of a server, but its physical location is up to the\n  provider, it may be a small part of a bigger computer, it may be\n  relocated it may even be spread out over multiple physical computers\n  and geographic regions.</p>\n  \n  <h3>Server less:</h3>\n  \n  <p>The look and feel of a webpage,\n  not a server. No command line, no setting up firewalls and configuring\n  basic SSH or HTTP service. It will just run your scripts.</p>\n</blockquote>\n\n<p>So why bother with anything but server less? There are a few good reasons thinkable, vendor lock-in is something I’d worry about. Can I go to a different provider with my code? Can I find a job, with my acquired skills, at a company that uses a different provider? Or will I be forever enslaved to a single super-power?</p>\n\n<p>Google Cloud ML is such a “server less” service offered specifically for running tensorflow scripts on a single machine or cluster of machines with GPU’s. The service has most standard Python packages available and most others can be installed with pip install. That makes the platform fairly open and there isn’t much “skills” to acquire, it just works.</p>\n\n<p>The fact that the TSA data already resided on a Google Storage Bucket and a $500 credit from the Google marketing department sweetened the deal to go and embrace the “server less” paradigm. </p>\n\n<h2>Job submission</h2>\n\n<p>For Cloud ML you prepare your Python Tensorflow script locally, may test-run it locally and eventually you’ll post it on the cloud engine to run it. You configure the type of machine or cluster of machines you’d like to use in a config.yaml file, number of CPU’s and GPU’s amount of RAM, number of machines etc. Finally you post the job with a simple command, like so:</p>\n\n<pre>    gcloud ml-engine jobs submit training \"jobname\" \\\n        --scale-tier CUSTOM \\\n        --config config.yaml \\\n        --module-name train.train \\\n        --package-path train \\\n        --runtime-version=1.2 \\\n</pre>\n\n<p>This will package your script locally, stage the whole package on your cloud storage bucket and launch the job with Cloud ML. For the jobname I use a bash script to automatically generate 5 letter names in alphabetical order, everything remains ordered under those jobnames.</p>\n\n<p>In my version of this I used a global DEBUG flag to essentially create two versions of my script: a simple version that can run (and finish) on my laptop and the real one. This way I could quickly test new development and don’t need to wait for Cloud ML response just to discover a simple syntax error.</p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/256451/8055/jobs.png\" alt=\"Cloud ML organizes its task by jobname\" title=\"\"></p>\n\n<p>Then on a productive Sunday afternoon, when multiple ideas come to mind or when you want to try multiple variations of the same idea, multiple jobs can easily be launched concurrently. Close your laptop and check back in later when they are done. You don’t need to bother spinning up multiple VM’s and you don’t need to make sure they get ended. If you don’t have any inspiration, simply don’t submit any jobs, no worries “should I keep the VM for next time? Or should I park it and save cost?”</p>\n\n<p>Cloud ML has a special offering for inference, where you pay per infer I believe. While this may turn out to be very power full in online high throughput high demand kind of situations I didn’t find it too useful. For the once in a lifetime type of infer that we need in this competition it isn’t worth all the extra effort to setup an online system. Instead I run infer jobs just like training jobs, just this time the input is the weights and the output is a Kaggle submission file.</p>\n\n<h2>Cloud storage</h2>\n\n<p>The inputs and outputs of your script are typically to a cloud storage bucket. When the training data is very large, such as with this competition, it is convenient -necessary even- to not need to download all the data. Besides the actual training job you could also do data preprocessing jobs and store the processed data back onto the cloud bucket.</p>\n\n<p>The results get stored on the cloud as well, the time traces, the weights, the evaluation cross-entropy etc.. You can download them using simple command line tools provided through <a href=\"https://cloud.google.com/storage/docs/gsutil\">gsutils</a>. But it’s not even needed to download the results. The weights can be re-loaded by your infer script job or by your training continuation jobs.</p>\n\n<p>For the time traces of cross entropy, weight sparsity, bias distributions etc I used Tensorboard.</p>\n\n<h2>Tensorboard</h2>\n\n<p>Tensorboard can be started from the cloud shell, directly pointing to the storage bucket containing the training transients. Doing so provides a nice and direct way to review the training process while it is working. Sorted by the same jobnames as the training jobs themselves. Using the tensorboard build in Regex filter we can slice through the data in virtually any way thinkable.</p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/256451/8056/tb.png\" alt=\"Straight from the cloud storage, web based results viewing\" title=\"\"></p>\n\n<h2>Stackdriver</h2>\n\n<p>Finally, I used stackdriver extensively. Log messages generated in your scripts and send to <code>tf.logging.info(...)</code> are routed straight to the online log message review system. It’s the same as wandering through your plain text logging transcripts, but online instead. The lines are folded to one-liners, everything is searchable and there are some really good slicing methods available.</p>\n\n<p>All together, I have managed to stay away from any overhead work related to setting up servers and further did not need to download most of the data (just a small sample for testing purposes). I believe this saved a fair amount of time as there is always more hacking involved in setting up the new server with all the communication channels than I anticipate at the start of such a project. But who knows, maybe that became really easy as well,..</p>",
      "rawMarkdown": "Making neural nets, learning them, testing, analyzing the results, analyzing the data and going back. Cycling through this many times and learn (your own head here) is at the basis of creating successful entries. It’s a corollary of the [fail fast][1] proposition.\n\nTo fail fast, it’s worth putting some thought into your workflow, making sure none of your tools is preventing you from failing, or better said from finding out where you got it wrong. I put some thought into my workflow and decided to share my findings, maybe it inspires some of you or maybe it breaks the ice and gets you to share your practices. And not to the least, I’m waiting anxiously for the stage 2 test data to be released and thus had some time at my hands!\n\nServer less\n-----------\nI don’t have a computer, I actually don’t like to have humming heating machines in my home. As a kid I made my own computers from parts, I soldered my own extensions to PCI slots, I hacked Modems from tape-recorders and setup Fidonet network branches using VHF radio. I’m no stranger to the smell of heating and outgassing electronic parts. But now I’m done with it all. I removed all hardware from my life and only have a clean minimalistic laptop on an empty wooden table in an empty room (I do allow for a cup of coffee occasionally).\n\nSo the cloud it is, it has to be. Originally I was planning on using a VM on Google or Amazon, install Tensorflow, other packages needed for my scripts and some method to sync the code between my laptop and the VM. While this is not too difficult to setup there are a few details to be worked out. I browsed around and stumbled upon a new paradigm, called “server less”.\n\n&gt; <h3>Server:</h3>\n&gt; A machine you own or rent, is dusted off by you or someone you\n&gt; pay and can be connected to either remotely or directly.\n&gt; <h3>Cloud or VM:</h3>\n&gt; The look and feel of a server, but its physical location is up to the\n&gt; provider, it may be a small part of a bigger computer, it may be\n&gt; relocated it may even be spread out over multiple physical computers\n&gt; and geographic regions.\n&gt; <h3>Server less:</h3>\n&gt; The look and feel of a webpage,\n&gt; not a server. No command line, no setting up firewalls and configuring\n&gt; basic SSH or HTTP service. It will just run your scripts.\n\nSo why bother with anything but server less? There are a few good reasons thinkable, vendor lock-in is something I’d worry about. Can I go to a different provider with my code? Can I find a job, with my acquired skills, at a company that uses a different provider? Or will I be forever enslaved to a single super-power?\n\nGoogle Cloud ML is such a “server less” service offered specifically for running tensorflow scripts on a single machine or cluster of machines with GPU’s. The service has most standard Python packages available and most others can be installed with pip install. That makes the platform fairly open and there isn’t much “skills” to acquire, it just works.\n\nThe fact that the TSA data already resided on a Google Storage Bucket and a $500 credit from the Google marketing department sweetened the deal to go and embrace the “server less” paradigm. \n\nJob submission\n--------------\nFor Cloud ML you prepare your Python Tensorflow script locally, may test-run it locally and eventually you’ll post it on the cloud engine to run it. You configure the type of machine or cluster of machines you’d like to use in a config.yaml file, number of CPU’s and GPU’s amount of RAM, number of machines etc. Finally you post the job with a simple command, like so:\n<pre>    gcloud ml-engine jobs submit training \"jobname\" \\\n        --scale-tier CUSTOM \\\n        --config config.yaml \\\n        --module-name train.train \\\n        --package-path train \\\n        --runtime-version=1.2 \\\n</pre>\nThis will package your script locally, stage the whole package on your cloud storage bucket and launch the job with Cloud ML. For the jobname I use a bash script to automatically generate 5 letter names in alphabetical order, everything remains ordered under those jobnames.\n\nIn my version of this I used a global DEBUG flag to essentially create two versions of my script: a simple version that can run (and finish) on my laptop and the real one. This way I could quickly test new development and don’t need to wait for Cloud ML response just to discover a simple syntax error.\n\n![Cloud ML organizes its task by jobname][2]\n\nThen on a productive Sunday afternoon, when multiple ideas come to mind or when you want to try multiple variations of the same idea, multiple jobs can easily be launched concurrently. Close your laptop and check back in later when they are done. You don’t need to bother spinning up multiple VM’s and you don’t need to make sure they get ended. If you don’t have any inspiration, simply don’t submit any jobs, no worries “should I keep the VM for next time? Or should I park it and save cost?”\n\nCloud ML has a special offering for inference, where you pay per infer I believe. While this may turn out to be very power full in online high throughput high demand kind of situations I didn’t find it too useful. For the once in a lifetime type of infer that we need in this competition it isn’t worth all the extra effort to setup an online system. Instead I run infer jobs just like training jobs, just this time the input is the weights and the output is a Kaggle submission file.\n\nCloud storage\n-------------\n\nThe inputs and outputs of your script are typically to a cloud storage bucket. When the training data is very large, such as with this competition, it is convenient -necessary even- to not need to download all the data. Besides the actual training job you could also do data preprocessing jobs and store the processed data back onto the cloud bucket.\n\nThe results get stored on the cloud as well, the time traces, the weights, the evaluation cross-entropy etc.. You can download them using simple command line tools provided through [gsutils][3]. But it’s not even needed to download the results. The weights can be re-loaded by your infer script job or by your training continuation jobs.\n\nFor the time traces of cross entropy, weight sparsity, bias distributions etc I used Tensorboard.\n\nTensorboard\n-----------\n\nTensorboard can be started from the cloud shell, directly pointing to the storage bucket containing the training transients. Doing so provides a nice and direct way to review the training process while it is working. Sorted by the same jobnames as the training jobs themselves. Using the tensorboard build in Regex filter we can slice through the data in virtually any way thinkable.\n\n![Straight from the cloud storage, web based results viewing][4]\n\nStackdriver\n-----------\n\nFinally, I used stackdriver extensively. Log messages generated in your scripts and send to <code>tf.logging.info(...)</code> are routed straight to the online log message review system. It’s the same as wandering through your plain text logging transcripts, but online instead. The lines are folded to one-liners, everything is searchable and there are some really good slicing methods available.\n\nAll together, I have managed to stay away from any overhead work related to setting up servers and further did not need to download most of the data (just a small sample for testing purposes). I believe this saved a fair amount of time as there is always more hacking involved in setting up the new server with all the communication channels than I anticipate at the start of such a project. But who knows, maybe that became really easy as well,..\n\n  [1]: https://i.pinimg.com/originals/2c/68/f6/2c68f617f7cf8ef8e84e3c5b22873d9f.jpg\n  [2]: https://kaggle2.blob.core.windows.net/forum-message-attachments/256451/8055/jobs.png\n  [3]: https://cloud.google.com/storage/docs/gsutil\n  [4]: https://kaggle2.blob.core.windows.net/forum-message-attachments/256451/8056/tb.png",
      "votes": null
    },
    {
      "id": "256468",
      "postDate": "12/12/2017 02:26:41",
      "content": "<p>How much have you spent for cloud computations?</p>",
      "rawMarkdown": "How much have you spent for cloud computations?",
      "votes": null
    },
    {
      "id": "256502",
      "postDate": "12/12/2017 04:35:15",
      "content": "<p>$200 plus the Google credit, and counting,..</p>",
      "rawMarkdown": "$200 plus the Google credit, and counting,..",
      "votes": null
    },
    {
      "id": "258306",
      "postDate": "12/15/2017 22:57:49",
      "content": "<p>H(o)i Bastiaan,</p>\n\n<p>Thanks for sharing. I’ve played with Google Cloud VMs a bit, but haven’t tried the Google ML cloud service as yet. Perhaps I’ll give it a try given your good experiences.</p>\n\n<p>On my side, taking more or less the opposite approach, I assembled what’s possibly “the last V8” of high-end Personal Computers suitable for DL. A 4-GPU 1080ti system with all the bells and whistles. It’s water cooled (“hybrid”) so it “only” has 16 fans in it... each of which contributes greatly to it’s super-silent (as compared to a vacuum cleaner) operation...</p>\n\n<p>Indeed it sighs, buzzes, zooms, puffs, sweats and let’s of the occasional whiff of steam. The other day I turned it off and I swear it started snoring! As long as it doesn’t start smoking all is OK though! </p>\n\n<p>Serious though: I’m running InceptionResnetV2 training on 3 of the GPUs leaving the 4th for periodically running inference towards the end of the training to guard against overfitting. From-scratch training my TSA image set (about 1 million 299x299x3 images “extracted” from the TSA data using several image manipulation steps) for the binary classifier takes about 100-125hrs. I tried adding the 4th GPU for training also but don’t see much performance gain anymore. I’m guessing the CPU needing to “serve” the GPU tensor-networks with the training images then becomes a bottleneck (note that I’m running a synchronous multi-GPU TF setup).</p>\n\n<p>Are you able to run a multi-GPU synchronous setup in Google Cloud ML? If so, is the bandwidth from the storage bucket high enough to serve the neural net with training data? Or do you apply an asynchronous setup?</p>\n\n<p>-Hans.</p>",
      "rawMarkdown": "H(o)i Bastiaan,\n\nThanks for sharing. I’ve played with Google Cloud VMs a bit, but haven’t tried the Google ML cloud service as yet. Perhaps I’ll give it a try given your good experiences.\n\nOn my side, taking more or less the opposite approach, I assembled what’s possibly “the last V8” of high-end Personal Computers suitable for DL. A 4-GPU 1080ti system with all the bells and whistles. It’s water cooled (“hybrid”) so it “only” has 16 fans in it... each of which contributes greatly to it’s super-silent (as compared to a vacuum cleaner) operation...\n\nIndeed it sighs, buzzes, zooms, puffs, sweats and let’s of the occasional whiff of steam. The other day I turned it off and I swear it started snoring! As long as it doesn’t start smoking all is OK though! \n\nSerious though: I’m running InceptionResnetV2 training on 3 of the GPUs leaving the 4th for periodically running inference towards the end of the training to guard against overfitting. From-scratch training my TSA image set (about 1 million 299x299x3 images “extracted” from the TSA data using several image manipulation steps) for the binary classifier takes about 100-125hrs. I tried adding the 4th GPU for training also but don’t see much performance gain anymore. I’m guessing the CPU needing to “serve” the GPU tensor-networks with the training images then becomes a bottleneck (note that I’m running a synchronous multi-GPU TF setup).\n\nAre you able to run a multi-GPU synchronous setup in Google Cloud ML? If so, is the bandwidth from the storage bucket high enough to serve the neural net with training data? Or do you apply an asynchronous setup?\n\n-Hans.",
      "votes": null
    },
    {
      "id": "258389",
      "postDate": "12/16/2017 02:43:06",
      "content": "<p>I used personal computer specially built for the competition:</p>\n\n<ul>\n<li><p>Intel Extreme i7-7820X 3.6/4.3GHz, 8-Core </p></li>\n<li><p>Gigabyte GTX1080Ti G1BLK 11GB memory</p></li>\n<li><p>DDR4 4x16GB, 3000MHz</p></li>\n<li><p>Samsung 960 EVO, 500GB, PCI-Express</p></li>\n</ul>\n\n<p>CPU was loaded only by 10% most of the time. So, I would prefer to buy some cheaper model. But I used 100% of the RAM and SSD for 192 GB swap. PCI-Express SSD worked really fast with the swap and I recommend using it when you can't afford or add more memory.</p>",
      "rawMarkdown": "I used personal computer specially built for the competition:\n\n - Intel Extreme i7-7820X 3.6/4.3GHz, 8-Core \n\n - Gigabyte GTX1080Ti G1BLK 11GB memory\n\n - DDR4 4x16GB, 3000MHz\n\n - Samsung 960 EVO, 500GB, PCI-Express\n\nCPU was loaded only by 10% most of the time. So, I would prefer to buy some cheaper model. But I used 100% of the RAM and SSD for 192 GB swap. PCI-Express SSD worked really fast with the swap and I recommend using it when you can't afford or add more memory.",
      "votes": null
    },
    {
      "id": "258415",
      "postDate": "12/16/2017 03:37:49",
      "content": "<p>Thanks Bastiaan for the information, my hardware usage was similar to yours. I ended up exhausting $300 of Google Cloud credit and a few hundred dollars on top of it. Unlike other competitions this one was more data heavy, it stretched(and improved) by data and image munging skills.</p>",
      "rawMarkdown": "Thanks Bastiaan for the information, my hardware usage was similar to yours. I ended up exhausting $300 of Google Cloud credit and a few hundred dollars on top of it. Unlike other competitions this one was more data heavy, it stretched(and improved) by data and image munging skills.",
      "votes": null
    },
    {
      "id": "258706",
      "postDate": "12/16/2017 20:42:21",
      "content": "<p>Hans, +1 on having local GPUs with water cooling.  I have a water-cooled 3-GPU machine and an air-cooled 3-GPU machine (1080 Ti's).  What I have noticed, besides the addressing with the noise factor (which is significant in itself), is GPUs not being subject to thermal throttling under water cooling.  The water-cooled GPUs do not go above 60C when all 3 GPUs are running under full load, where as air-cooled ones can get super hot (my EVGA 1080 Ti SC's get up to 91C) for the ones in the PCI slots that don't get a lot of airflow due to the cards being stacked one against another.  When the cards are under thermal throttling, they pull around 150W or less (as opposed to ~250W) and slow down <em>significantly</em>.  So I'm glad I made investments in my water-cooled setup.\nI have also experimented with Google Cloud.  The K80's were pretty useless.  I don't have the numbers written down but same model trained on my local 1080 Ti were taking something like 5-10x as long per epoch on the K80's.  Then I tried P100's and they were pretty comparable to the 1080 Ti's.  But P100's are really expensive to be running on the cloud (running 4 P100's in parallel would cost about $6/hr for the GPUs alone + the VM cost.)  </p>",
      "rawMarkdown": "Hans, +1 on having local GPUs with water cooling.  I have a water-cooled 3-GPU machine and an air-cooled 3-GPU machine (1080 Ti's).  What I have noticed, besides the addressing with the noise factor (which is significant in itself), is GPUs not being subject to thermal throttling under water cooling.  The water-cooled GPUs do not go above 60C when all 3 GPUs are running under full load, where as air-cooled ones can get super hot (my EVGA 1080 Ti SC's get up to 91C) for the ones in the PCI slots that don't get a lot of airflow due to the cards being stacked one against another.  When the cards are under thermal throttling, they pull around 150W or less (as opposed to ~250W) and slow down *significantly*.  So I'm glad I made investments in my water-cooled setup.\nI have also experimented with Google Cloud.  The K80's were pretty useless.  I don't have the numbers written down but same model trained on my local 1080 Ti were taking something like 5-10x as long per epoch on the K80's.  Then I tried P100's and they were pretty comparable to the 1080 Ti's.  But P100's are really expensive to be running on the cloud (running 4 P100's in parallel would cost about $6/hr for the GPUs alone + the VM cost.)",
      "votes": null
    },
    {
      "id": "258740",
      "postDate": "12/16/2017 23:25:06",
      "content": "<p>Is it possible to use 24GB on Tesla K80 for one training process? Can you explain, why it works 5-10x slower than 1080 Ti?</p>",
      "rawMarkdown": "Is it possible to use 24GB on Tesla K80 for one training process? Can you explain, why it works 5-10x slower than 1080 Ti?",
      "votes": null
    },
    {
      "id": "258756",
      "postDate": "12/17/2017 00:40:50",
      "content": "<p>Maybe the K80's are aircooled? I noticed a drop in speed roughly after 20 epoch's. Never figured out why as I didn't have access to diagnostic tools.</p>",
      "rawMarkdown": "Maybe the K80's are aircooled? I noticed a drop in speed roughly after 20 epoch's. Never figured out why as I didn't have access to diagnostic tools.",
      "votes": null
    },
    {
      "id": "258757",
      "postDate": "12/17/2017 00:45:05",
      "content": "<p>Hoi Hans!\nYes, You can do up to 8 GPU's on a computer, and you can train with them synchronically as well as asynchrone, however you please. You can also use multiple computers each with multiple GPU's. Also here either way, but I'm not sure if anything else then asynchron would make sense as you do have network delays. Data fetched from the Storage bucket in 2 minutes (all APS data), I'd do that for the start of the training job and store it locally for the duration of the job.</p>",
      "rawMarkdown": "Hoi Hans!\nYes, You can do up to 8 GPU's on a computer, and you can train with them synchronically as well as asynchrone, however you please. You can also use multiple computers each with multiple GPU's. Also here either way, but I'm not sure if anything else then asynchron would make sense as you do have network delays. Data fetched from the Storage bucket in 2 minutes (all APS data), I'd do that for the start of the training job and store it locally for the duration of the job.",
      "votes": null
    },
    {
      "id": "258777",
      "postDate": "12/17/2017 02:02:47",
      "content": "<p>The K80 is actually 2 GPUs with 12GB each.  So when you provision a VM on the cloud with an \"instance\" of K80, you are getting a GPU w/ 12GB.\nHere's an article that compares K80 (on Amazon) vs 1080 Ti, and it is pretty consistent with what I saw: <a href=\"https://medium.com/initialized-capital/benchmarking-tensorflow-performance-and-cost-across-different-gpu-options-69bd85fe5d58\">https://medium.com/initialized-capital/benchmarking-tensorflow-performance-and-cost-across-different-gpu-options-69bd85fe5d58</a> (the numbers shown there is K80 vs 1080 Ti is roughly 4x.)</p>",
      "rawMarkdown": "The K80 is actually 2 GPUs with 12GB each.  So when you provision a VM on the cloud with an \"instance\" of K80, you are getting a GPU w/ 12GB.\nHere's an article that compares K80 (on Amazon) vs 1080 Ti, and it is pretty consistent with what I saw: https://medium.com/initialized-capital/benchmarking-tensorflow-performance-and-cost-across-different-gpu-options-69bd85fe5d58 (the numbers shown there is K80 vs 1080 Ti is roughly 4x.)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 256468,
      "author_name": "dmitrykovba",
      "author_url": "",
      "post_date": "12/12/2017 02:26:41",
      "content": "<p>How much have you spent for cloud computations?</p>",
      "votes": null,
      "replies": [
        {
          "id": 256502,
          "author_name": "bastiaanbergman",
          "author_url": "",
          "post_date": "12/12/2017 04:35:15",
          "content": "<p>$200 plus the Google credit, and counting,..</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 258306,
      "author_name": "gra55h0pper",
      "author_url": "",
      "post_date": "12/15/2017 22:57:49",
      "content": "<p>H(o)i Bastiaan,</p>\n\n<p>Thanks for sharing. I’ve played with Google Cloud VMs a bit, but haven’t tried the Google ML cloud service as yet. Perhaps I’ll give it a try given your good experiences.</p>\n\n<p>On my side, taking more or less the opposite approach, I assembled what’s possibly “the last V8” of high-end Personal Computers suitable for DL. A 4-GPU 1080ti system with all the bells and whistles. It’s water cooled (“hybrid”) so it “only” has 16 fans in it... each of which contributes greatly to it’s super-silent (as compared to a vacuum cleaner) operation...</p>\n\n<p>Indeed it sighs, buzzes, zooms, puffs, sweats and let’s of the occasional whiff of steam. The other day I turned it off and I swear it started snoring! As long as it doesn’t start smoking all is OK though! </p>\n\n<p>Serious though: I’m running InceptionResnetV2 training on 3 of the GPUs leaving the 4th for periodically running inference towards the end of the training to guard against overfitting. From-scratch training my TSA image set (about 1 million 299x299x3 images “extracted” from the TSA data using several image manipulation steps) for the binary classifier takes about 100-125hrs. I tried adding the 4th GPU for training also but don’t see much performance gain anymore. I’m guessing the CPU needing to “serve” the GPU tensor-networks with the training images then becomes a bottleneck (note that I’m running a synchronous multi-GPU TF setup).</p>\n\n<p>Are you able to run a multi-GPU synchronous setup in Google Cloud ML? If so, is the bandwidth from the storage bucket high enough to serve the neural net with training data? Or do you apply an asynchronous setup?</p>\n\n<p>-Hans.</p>",
      "votes": null,
      "replies": [
        {
          "id": 258706,
          "author_name": "u39kun",
          "author_url": "",
          "post_date": "12/16/2017 20:42:21",
          "content": "<p>Hans, +1 on having local GPUs with water cooling.  I have a water-cooled 3-GPU machine and an air-cooled 3-GPU machine (1080 Ti's).  What I have noticed, besides the addressing with the noise factor (which is significant in itself), is GPUs not being subject to thermal throttling under water cooling.  The water-cooled GPUs do not go above 60C when all 3 GPUs are running under full load, where as air-cooled ones can get super hot (my EVGA 1080 Ti SC's get up to 91C) for the ones in the PCI slots that don't get a lot of airflow due to the cards being stacked one against another.  When the cards are under thermal throttling, they pull around 150W or less (as opposed to ~250W) and slow down <em>significantly</em>.  So I'm glad I made investments in my water-cooled setup.\nI have also experimented with Google Cloud.  The K80's were pretty useless.  I don't have the numbers written down but same model trained on my local 1080 Ti were taking something like 5-10x as long per epoch on the K80's.  Then I tried P100's and they were pretty comparable to the 1080 Ti's.  But P100's are really expensive to be running on the cloud (running 4 P100's in parallel would cost about $6/hr for the GPUs alone + the VM cost.)  </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 258740,
          "author_name": "dmitrykovba",
          "author_url": "",
          "post_date": "12/16/2017 23:25:06",
          "content": "<p>Is it possible to use 24GB on Tesla K80 for one training process? Can you explain, why it works 5-10x slower than 1080 Ti?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 258756,
          "author_name": "bastiaanbergman",
          "author_url": "",
          "post_date": "12/17/2017 00:40:50",
          "content": "<p>Maybe the K80's are aircooled? I noticed a drop in speed roughly after 20 epoch's. Never figured out why as I didn't have access to diagnostic tools.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 258757,
          "author_name": "bastiaanbergman",
          "author_url": "",
          "post_date": "12/17/2017 00:45:05",
          "content": "<p>Hoi Hans!\nYes, You can do up to 8 GPU's on a computer, and you can train with them synchronically as well as asynchrone, however you please. You can also use multiple computers each with multiple GPU's. Also here either way, but I'm not sure if anything else then asynchron would make sense as you do have network delays. Data fetched from the Storage bucket in 2 minutes (all APS data), I'd do that for the start of the training job and store it locally for the duration of the job.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 258777,
          "author_name": "u39kun",
          "author_url": "",
          "post_date": "12/17/2017 02:02:47",
          "content": "<p>The K80 is actually 2 GPUs with 12GB each.  So when you provision a VM on the cloud with an \"instance\" of K80, you are getting a GPU w/ 12GB.\nHere's an article that compares K80 (on Amazon) vs 1080 Ti, and it is pretty consistent with what I saw: <a href=\"https://medium.com/initialized-capital/benchmarking-tensorflow-performance-and-cost-across-different-gpu-options-69bd85fe5d58\">https://medium.com/initialized-capital/benchmarking-tensorflow-performance-and-cost-across-different-gpu-options-69bd85fe5d58</a> (the numbers shown there is K80 vs 1080 Ti is roughly 4x.)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 258389,
      "author_name": "dmitrykovba",
      "author_url": "",
      "post_date": "12/16/2017 02:43:06",
      "content": "<p>I used personal computer specially built for the competition:</p>\n\n<ul>\n<li><p>Intel Extreme i7-7820X 3.6/4.3GHz, 8-Core </p></li>\n<li><p>Gigabyte GTX1080Ti G1BLK 11GB memory</p></li>\n<li><p>DDR4 4x16GB, 3000MHz</p></li>\n<li><p>Samsung 960 EVO, 500GB, PCI-Express</p></li>\n</ul>\n\n<p>CPU was loaded only by 10% most of the time. So, I would prefer to buy some cheaper model. But I used 100% of the RAM and SSD for 192 GB swap. PCI-Express SSD worked really fast with the swap and I recommend using it when you can't afford or add more memory.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 258415,
      "author_name": "shivgowda",
      "author_url": "",
      "post_date": "12/16/2017 03:37:49",
      "content": "<p>Thanks Bastiaan for the information, my hardware usage was similar to yours. I ended up exhausting $300 of Google Cloud credit and a few hundred dollars on top of it. Unlike other competitions this one was more data heavy, it stretched(and improved) by data and image munging skills.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "256451": "Making neural nets, learning them, testing, analyzing the results, analyzing the data and going back. Cycling through this many times and learn (your own head here) is at the basis of creating successful entries. It’s a corollary of the [fail fast][1] proposition.\n\nTo fail fast, it’s worth putting some thought into your workflow, making sure none of your tools is preventing you from failing, or better said from finding out where you got it wrong. I put some thought into my workflow and decided to share my findings, maybe it inspires some of you or maybe it breaks the ice and gets you to share your practices. And not to the least, I’m waiting anxiously for the stage 2 test data to be released and thus had some time at my hands!\n\nServer less\n-----------\nI don’t have a computer, I actually don’t like to have humming heating machines in my home. As a kid I made my own computers from parts, I soldered my own extensions to PCI slots, I hacked Modems from tape-recorders and setup Fidonet network branches using VHF radio. I’m no stranger to the smell of heating and outgassing electronic parts. But now I’m done with it all. I removed all hardware from my life and only have a clean minimalistic laptop on an empty wooden table in an empty room (I do allow for a cup of coffee occasionally).\n\nSo the cloud it is, it has to be. Originally I was planning on using a VM on Google or Amazon, install Tensorflow, other packages needed for my scripts and some method to sync the code between my laptop and the VM. While this is not too difficult to setup there are a few details to be worked out. I browsed around and stumbled upon a new paradigm, called “server less”.\n\n&gt; <h3>Server:</h3>\n&gt; A machine you own or rent, is dusted off by you or someone you\n&gt; pay and can be connected to either remotely or directly.\n&gt; <h3>Cloud or VM:</h3>\n&gt; The look and feel of a server, but its physical location is up to the\n&gt; provider, it may be a small part of a bigger computer, it may be\n&gt; relocated it may even be spread out over multiple physical computers\n&gt; and geographic regions.\n&gt; <h3>Server less:</h3>\n&gt; The look and feel of a webpage,\n&gt; not a server. No command line, no setting up firewalls and configuring\n&gt; basic SSH or HTTP service. It will just run your scripts.\n\nSo why bother with anything but server less? There are a few good reasons thinkable, vendor lock-in is something I’d worry about. Can I go to a different provider with my code? Can I find a job, with my acquired skills, at a company that uses a different provider? Or will I be forever enslaved to a single super-power?\n\nGoogle Cloud ML is such a “server less” service offered specifically for running tensorflow scripts on a single machine or cluster of machines with GPU’s. The service has most standard Python packages available and most others can be installed with pip install. That makes the platform fairly open and there isn’t much “skills” to acquire, it just works.\n\nThe fact that the TSA data already resided on a Google Storage Bucket and a $500 credit from the Google marketing department sweetened the deal to go and embrace the “server less” paradigm. \n\nJob submission\n--------------\nFor Cloud ML you prepare your Python Tensorflow script locally, may test-run it locally and eventually you’ll post it on the cloud engine to run it. You configure the type of machine or cluster of machines you’d like to use in a config.yaml file, number of CPU’s and GPU’s amount of RAM, number of machines etc. Finally you post the job with a simple command, like so:\n<pre>    gcloud ml-engine jobs submit training \"jobname\" \\\n        --scale-tier CUSTOM \\\n        --config config.yaml \\\n        --module-name train.train \\\n        --package-path train \\\n        --runtime-version=1.2 \\\n</pre>\nThis will package your script locally, stage the whole package on your cloud storage bucket and launch the job with Cloud ML. For the jobname I use a bash script to automatically generate 5 letter names in alphabetical order, everything remains ordered under those jobnames.\n\nIn my version of this I used a global DEBUG flag to essentially create two versions of my script: a simple version that can run (and finish) on my laptop and the real one. This way I could quickly test new development and don’t need to wait for Cloud ML response just to discover a simple syntax error.\n\n![Cloud ML organizes its task by jobname][2]\n\nThen on a productive Sunday afternoon, when multiple ideas come to mind or when you want to try multiple variations of the same idea, multiple jobs can easily be launched concurrently. Close your laptop and check back in later when they are done. You don’t need to bother spinning up multiple VM’s and you don’t need to make sure they get ended. If you don’t have any inspiration, simply don’t submit any jobs, no worries “should I keep the VM for next time? Or should I park it and save cost?”\n\nCloud ML has a special offering for inference, where you pay per infer I believe. While this may turn out to be very power full in online high throughput high demand kind of situations I didn’t find it too useful. For the once in a lifetime type of infer that we need in this competition it isn’t worth all the extra effort to setup an online system. Instead I run infer jobs just like training jobs, just this time the input is the weights and the output is a Kaggle submission file.\n\nCloud storage\n-------------\n\nThe inputs and outputs of your script are typically to a cloud storage bucket. When the training data is very large, such as with this competition, it is convenient -necessary even- to not need to download all the data. Besides the actual training job you could also do data preprocessing jobs and store the processed data back onto the cloud bucket.\n\nThe results get stored on the cloud as well, the time traces, the weights, the evaluation cross-entropy etc.. You can download them using simple command line tools provided through [gsutils][3]. But it’s not even needed to download the results. The weights can be re-loaded by your infer script job or by your training continuation jobs.\n\nFor the time traces of cross entropy, weight sparsity, bias distributions etc I used Tensorboard.\n\nTensorboard\n-----------\n\nTensorboard can be started from the cloud shell, directly pointing to the storage bucket containing the training transients. Doing so provides a nice and direct way to review the training process while it is working. Sorted by the same jobnames as the training jobs themselves. Using the tensorboard build in Regex filter we can slice through the data in virtually any way thinkable.\n\n![Straight from the cloud storage, web based results viewing][4]\n\nStackdriver\n-----------\n\nFinally, I used stackdriver extensively. Log messages generated in your scripts and send to <code>tf.logging.info(...)</code> are routed straight to the online log message review system. It’s the same as wandering through your plain text logging transcripts, but online instead. The lines are folded to one-liners, everything is searchable and there are some really good slicing methods available.\n\nAll together, I have managed to stay away from any overhead work related to setting up servers and further did not need to download most of the data (just a small sample for testing purposes). I believe this saved a fair amount of time as there is always more hacking involved in setting up the new server with all the communication channels than I anticipate at the start of such a project. But who knows, maybe that became really easy as well,..\n\n  [1]: https://i.pinimg.com/originals/2c/68/f6/2c68f617f7cf8ef8e84e3c5b22873d9f.jpg\n  [2]: https://kaggle2.blob.core.windows.net/forum-message-attachments/256451/8055/jobs.png\n  [3]: https://cloud.google.com/storage/docs/gsutil\n  [4]: https://kaggle2.blob.core.windows.net/forum-message-attachments/256451/8056/tb.png",
    "256468": "How much have you spent for cloud computations?",
    "256502": "$200 plus the Google credit, and counting,..",
    "258306": "H(o)i Bastiaan,\n\nThanks for sharing. I’ve played with Google Cloud VMs a bit, but haven’t tried the Google ML cloud service as yet. Perhaps I’ll give it a try given your good experiences.\n\nOn my side, taking more or less the opposite approach, I assembled what’s possibly “the last V8” of high-end Personal Computers suitable for DL. A 4-GPU 1080ti system with all the bells and whistles. It’s water cooled (“hybrid”) so it “only” has 16 fans in it... each of which contributes greatly to it’s super-silent (as compared to a vacuum cleaner) operation...\n\nIndeed it sighs, buzzes, zooms, puffs, sweats and let’s of the occasional whiff of steam. The other day I turned it off and I swear it started snoring! As long as it doesn’t start smoking all is OK though! \n\nSerious though: I’m running InceptionResnetV2 training on 3 of the GPUs leaving the 4th for periodically running inference towards the end of the training to guard against overfitting. From-scratch training my TSA image set (about 1 million 299x299x3 images “extracted” from the TSA data using several image manipulation steps) for the binary classifier takes about 100-125hrs. I tried adding the 4th GPU for training also but don’t see much performance gain anymore. I’m guessing the CPU needing to “serve” the GPU tensor-networks with the training images then becomes a bottleneck (note that I’m running a synchronous multi-GPU TF setup).\n\nAre you able to run a multi-GPU synchronous setup in Google Cloud ML? If so, is the bandwidth from the storage bucket high enough to serve the neural net with training data? Or do you apply an asynchronous setup?\n\n-Hans.",
    "258389": "I used personal computer specially built for the competition:\n\n - Intel Extreme i7-7820X 3.6/4.3GHz, 8-Core \n\n - Gigabyte GTX1080Ti G1BLK 11GB memory\n\n - DDR4 4x16GB, 3000MHz\n\n - Samsung 960 EVO, 500GB, PCI-Express\n\nCPU was loaded only by 10% most of the time. So, I would prefer to buy some cheaper model. But I used 100% of the RAM and SSD for 192 GB swap. PCI-Express SSD worked really fast with the swap and I recommend using it when you can't afford or add more memory.",
    "258415": "Thanks Bastiaan for the information, my hardware usage was similar to yours. I ended up exhausting $300 of Google Cloud credit and a few hundred dollars on top of it. Unlike other competitions this one was more data heavy, it stretched(and improved) by data and image munging skills.",
    "258706": "Hans, +1 on having local GPUs with water cooling.  I have a water-cooled 3-GPU machine and an air-cooled 3-GPU machine (1080 Ti's).  What I have noticed, besides the addressing with the noise factor (which is significant in itself), is GPUs not being subject to thermal throttling under water cooling.  The water-cooled GPUs do not go above 60C when all 3 GPUs are running under full load, where as air-cooled ones can get super hot (my EVGA 1080 Ti SC's get up to 91C) for the ones in the PCI slots that don't get a lot of airflow due to the cards being stacked one against another.  When the cards are under thermal throttling, they pull around 150W or less (as opposed to ~250W) and slow down *significantly*.  So I'm glad I made investments in my water-cooled setup.\nI have also experimented with Google Cloud.  The K80's were pretty useless.  I don't have the numbers written down but same model trained on my local 1080 Ti were taking something like 5-10x as long per epoch on the K80's.  Then I tried P100's and they were pretty comparable to the 1080 Ti's.  But P100's are really expensive to be running on the cloud (running 4 P100's in parallel would cost about $6/hr for the GPUs alone + the VM cost.)",
    "258740": "Is it possible to use 24GB on Tesla K80 for one training process? Can you explain, why it works 5-10x slower than 1080 Ti?",
    "258756": "Maybe the K80's are aircooled? I noticed a drop in speed roughly after 20 epoch's. Never figured out why as I didn't have access to diagnostic tools.",
    "258757": "Hoi Hans!\nYes, You can do up to 8 GPU's on a computer, and you can train with them synchronically as well as asynchrone, however you please. You can also use multiple computers each with multiple GPU's. Also here either way, but I'm not sure if anything else then asynchron would make sense as you do have network delays. Data fetched from the Storage bucket in 2 minutes (all APS data), I'd do that for the start of the training job and store it locally for the duration of the job.",
    "258777": "The K80 is actually 2 GPUs with 12GB each.  So when you provision a VM on the cloud with an \"instance\" of K80, you are getting a GPU w/ 12GB.\nHere's an article that compares K80 (on Amazon) vs 1080 Ti, and it is pretty consistent with what I saw: https://medium.com/initialized-capital/benchmarking-tensorflow-performance-and-cost-across-different-gpu-options-69bd85fe5d58 (the numbers shown there is K80 vs 1080 Ti is roughly 4x.)"
  },
  "source": "meta"
}