{
  "id": 128150,
  "title": "Frustrating \"cloud\" training",
  "url": "/competitions/deepfake-detection-challenge/discussion/128150",
  "author_name": "dagnelies",
  "post_date": "2020-01-29T08:39:18.026000",
  "votes": 4,
  "comment_count": 10,
  "views": 0,
  "content": "<p>Hi,\nare you successfully preprocessing/training your data with AWS/GCP?\nMy overall feeling is like:\n- On Kaggle: \"Oh wonderful, everything is so simple and productive!\"\n- On AWS/GCP: \"Oh my god, why is everything always so complicated?! I spend way more time to set things up than to try models!\"</p>\n\n<p>For example, I saw SageMaker on AWS so I thought it'd be similar to Kaggle and I gladly used.\nBut it turns out to be another productivity loss.</p>\n\n<p>Their SageMaker uses Tensorflow 1.x, so it's an incompatible version from Kaggle here which is 2.x. Despite amazon announced this January that SageMaker now supports TF 2.x, it is a mystery to me how to use it. Every notebook is still TF 1.x. So I thought I'd upgrade it manually within the notebook itself with \"pip --upgrade\" and it did. But then it simply failed at runtime with errors like:\n<code>\nUnknownError:  Failed to get convolution algorithm. This is probably because cuDNN failed to initialize, so try looking to see if a warning log message was printed above.\n     [[node conv2d_1/convolution (defined at /home/ec2-user/anaconda3/envs/tensorflow_p36/lib/python3.6/site-packages/keras/backend/tensorflow_backend.py:3009) ]]\n</code>\nSo basically, you spend a lot of time, just for having your code not working in the end. And this is just one example among many. It's really frustrating. And if you use GPUs, your credits melt like ice cream in the sun. </p>\n\n<p>Do you have smoother experiences? Do you set up VMs from scratch?</p>",
  "messages": [
    {
      "id": 731947,
      "postDate": "2020-01-29T09:34:40.590Z",
      "content": "<p>I am using Amazon EC2 and my experience has been quite good. I'm not using SageMaker and Tensorflow though. I'm just setting up a normal VM and setting up my own Jupyter service, and using PyTorch.</p>\n\n<p>I wrote a script with everything I need to setup the VM (download Anaconda, install CUDA, update the packages I need and so on) that I run when after the VM starts. Since I'm using spot instances I made the script so everything I terminate one VM it becomes easier to start a new one.</p>\n\n<p>I haven't used SageMaker so I don't know how it works, but I guess using a normal VM gives more control over the setup, setting up your own Jupyter server should be easy, just google \"how to setup jupyter server on EC2\", there are many tutorials out there. After coming up with a script to config the stuff as you need it should become straightforward to run your workload without problems.</p>",
      "rawMarkdown": "I am using Amazon EC2 and my experience has been quite good. I'm not using SageMaker and Tensorflow though. I'm just setting up a normal VM and setting up my own Jupyter service, and using PyTorch.\n\nI wrote a script with everything I need to setup the VM (download Anaconda, install CUDA, update the packages I need and so on) that I run when after the VM starts. Since I'm using spot instances I made the script so everything I terminate one VM it becomes easier to start a new one.\n\nI haven't used SageMaker so I don't know how it works, but I guess using a normal VM gives more control over the setup, setting up your own Jupyter server should be easy, just google \"how to setup jupyter server on EC2\", there are many tutorials out there. After coming up with a script to config the stuff as you need it should become straightforward to run your workload without problems.",
      "votes": 8,
      "replies": [
        {
          "id": 732048,
          "postDate": "2020-01-29T12:26:08.453Z",
          "content": "<p><a href=\"/pedromb\">@pedromb</a>  ...but, if you are running spot instances, doesn't it imply that the training might be interrupted in the middle? How do you deal with that?</p>\n\n<p>Although spot instances sounded nice, I wonder how to persist state and recover. Do you use periodic model saves during training? There is also the wonderful \"Batch\" jobs possibility, but you have to build Docker images and so on, so it is another layer of complexity. 😓 </p>\n\n<p>I guess there is no easy way around this.</p>\n\n<p>I am also curious about the google's TPU instances, but they also bring their share of complexity. Sadly, I applied a bit too late because I didn't read there were two rounds. 🙄 </p>",
          "rawMarkdown": "@pedromb  ...but, if you are running spot instances, doesn't it imply that the training might be interrupted in the middle? How do you deal with that?\n\nAlthough spot instances sounded nice, I wonder how to persist state and recover. Do you use periodic model saves during training? There is also the wonderful \"Batch\" jobs possibility, but you have to build Docker images and so on, so it is another layer of complexity. 😓 \n\nI guess there is no easy way around this.\n\nI am also curious about the google's TPU instances, but they also bring their share of complexity. Sadly, I applied a bit too late because I didn't read there were two rounds. 🙄 "
        },
        {
          "id": 732208,
          "postDate": "2020-01-29T15:17:19.883Z",
          "content": "<p>I do save checkpoints of the model during training, but, to be honest, so far I didn't have an issue with the spot instance being interrupted.</p>\n\n<p>For persistence I'm using an EBS volume on AWS. It's a bit more expensive than S3 but at least I can just unmount and mount back when I change instances, so it works well.</p>",
          "rawMarkdown": "I do save checkpoints of the model during training, but, to be honest, so far I didn't have an issue with the spot instance being interrupted.\n\nFor persistence I'm using an EBS volume on AWS. It's a bit more expensive than S3 but at least I can just unmount and mount back when I change instances, so it works well.",
          "votes": 4
        },
        {
          "id": 747148,
          "postDate": "2020-02-16T03:00:07.550Z",
          "content": "<p><a href=\"/pedromb\">@pedromb</a> like your approach, could you share what EC2 instance type to use for this competition (preprocessing and training)</p>",
          "rawMarkdown": "@pedromb like your approach, could you share what EC2 instance type to use for this competition (preprocessing and training)"
        },
        {
          "id": 747378,
          "postDate": "2020-02-16T11:06:11.667Z",
          "content": "<p>I'm using p3.2xlarge, the normal price is $3.06/h, but, using spot instances it goes for around $0.97/h. </p>",
          "rawMarkdown": "I'm using p3.2xlarge, the normal price is $3.06/h, but, using spot instances it goes for around $0.97/h. ",
          "votes": 1
        },
        {
          "id": 747398,
          "postDate": "2020-02-16T11:32:39.033Z",
          "content": "<p><a href=\"/pedromb\">@pedromb</a> Thanks that helps to setup AWS well use resources.</p>",
          "rawMarkdown": "@pedromb Thanks that helps to setup AWS well use resources."
        }
      ]
    },
    {
      "id": 731916,
      "postDate": "2020-01-29T08:39:18.027Z",
      "content": "<p>Hi,\nare you successfully preprocessing/training your data with AWS/GCP?\nMy overall feeling is like:\n- On Kaggle: \"Oh wonderful, everything is so simple and productive!\"\n- On AWS/GCP: \"Oh my god, why is everything always so complicated?! I spend way more time to set things up than to try models!\"</p>\n\n<p>For example, I saw SageMaker on AWS so I thought it'd be similar to Kaggle and I gladly used.\nBut it turns out to be another productivity loss.</p>\n\n<p>Their SageMaker uses Tensorflow 1.x, so it's an incompatible version from Kaggle here which is 2.x. Despite amazon announced this January that SageMaker now supports TF 2.x, it is a mystery to me how to use it. Every notebook is still TF 1.x. So I thought I'd upgrade it manually within the notebook itself with \"pip --upgrade\" and it did. But then it simply failed at runtime with errors like:\n<code>\nUnknownError:  Failed to get convolution algorithm. This is probably because cuDNN failed to initialize, so try looking to see if a warning log message was printed above.\n     [[node conv2d_1/convolution (defined at /home/ec2-user/anaconda3/envs/tensorflow_p36/lib/python3.6/site-packages/keras/backend/tensorflow_backend.py:3009) ]]\n</code>\nSo basically, you spend a lot of time, just for having your code not working in the end. And this is just one example among many. It's really frustrating. And if you use GPUs, your credits melt like ice cream in the sun. </p>\n\n<p>Do you have smoother experiences? Do you set up VMs from scratch?</p>",
      "rawMarkdown": "Hi,\nare you successfully preprocessing/training your data with AWS/GCP?\nMy overall feeling is like:\n- On Kaggle: \"Oh wonderful, everything is so simple and productive!\"\n- On AWS/GCP: \"Oh my god, why is everything always so complicated?! I spend way more time to set things up than to try models!\"\n\nFor example, I saw SageMaker on AWS so I thought it'd be similar to Kaggle and I gladly used.\nBut it turns out to be another productivity loss.\n\nTheir SageMaker uses Tensorflow 1.x, so it's an incompatible version from Kaggle here which is 2.x. Despite amazon announced this January that SageMaker now supports TF 2.x, it is a mystery to me how to use it. Every notebook is still TF 1.x. So I thought I'd upgrade it manually within the notebook itself with \"pip --upgrade\" and it did. But then it simply failed at runtime with errors like:\n```\nUnknownError:  Failed to get convolution algorithm. This is probably because cuDNN failed to initialize, so try looking to see if a warning log message was printed above.\n\t [[node conv2d_1/convolution (defined at /home/ec2-user/anaconda3/envs/tensorflow_p36/lib/python3.6/site-packages/keras/backend/tensorflow_backend.py:3009) ]]\n```\nSo basically, you spend a lot of time, just for having your code not working in the end. And this is just one example among many. It's really frustrating. And if you use GPUs, your credits melt like ice cream in the sun. \n\nDo you have smoother experiences? Do you set up VMs from scratch?",
      "votes": 3
    },
    {
      "id": 732555,
      "postDate": "2020-01-29T23:45:58.057Z",
      "content": "<p>As a very accurate rule - \"life on kaggle is very much easier than real life.\"</p>\n\n<p>When I first used Goggle Cloud Platform last year my 'Curse Jar' filled up pretty fast.  The whole subject of quota's can get me ranting for hours.  At the end I said \"never again\".  For this competition my team mate is filling his jar on AWS.</p>\n\n<p>ML on the cloud would seem to be an absolute skill set needed for a data scientist.   I am in a position were I can avoid learning that skill set (very old, not looking for a job and have some nice PC's) but for most it's a rite of passage that I think you need to take.</p>\n\n<p>See if you can't team up with someone who has gotten past the early part of the learning curve on this or another competition.  I got thru the ramp up on GCP thanks to a great partner in the Molecules challenge.  Put an ad in the team wanted discussion post - maybe some souls out there with the experience will join up with you.</p>",
      "rawMarkdown": "As a very accurate rule - \"life on kaggle is very much easier than real life.\"\n\nWhen I first used Goggle Cloud Platform last year my 'Curse Jar' filled up pretty fast.  The whole subject of quota's can get me ranting for hours.  At the end I said \"never again\".  For this competition my team mate is filling his jar on AWS.\n\nML on the cloud would seem to be an absolute skill set needed for a data scientist.   I am in a position were I can avoid learning that skill set (very old, not looking for a job and have some nice PC's) but for most it's a rite of passage that I think you need to take.\n\nSee if you can't team up with someone who has gotten past the early part of the learning curve on this or another competition.  I got thru the ramp up on GCP thanks to a great partner in the Molecules challenge.  Put an ad in the team wanted discussion post - maybe some souls out there with the experience will join up with you.\n\n\n\n",
      "votes": 4,
      "replies": [
        {
          "id": 746012,
          "postDate": "2020-02-14T13:39:20.867Z",
          "content": "<p>Very accurate. </p>",
          "rawMarkdown": "Very accurate. ",
          "isDeleted": true
        }
      ]
    },
    {
      "id": 732547,
      "postDate": "2020-01-29T23:34:30.767Z",
      "content": "<p>The error you are getting because of incompatible cuDNN version and tensorflow version. Upgrade cuda and cuDNN version to the corresponding version will solve the problem.</p>",
      "rawMarkdown": "The error you are getting because of incompatible cuDNN version and tensorflow version. Upgrade cuda and cuDNN version to the corresponding version will solve the problem.",
      "votes": 1
    },
    {
      "id": 732667,
      "postDate": "2020-01-30T04:37:47.413Z",
      "content": "<p><a href=\"/dagnelies\">@dagnelies</a> if you have some GCP credits. You can use <a href=\"https://cloud.google.com/deep-learning-vm/\">GCP Deep Learning VM</a> Spin it up with a Jupyterlab option enabled. \nYou can then go to the AI platform &gt;&gt; Open a Notebook. TF2.1 is pre-installed. If you need you can also spin a Pytorch based framework, </p>\n\n<p>I have some limited experience with AWS SageMaker. But I was not impressed due to its severe limitations &amp; high costs.</p>",
      "rawMarkdown": "@dagnelies if you have some GCP credits. You can use [GCP Deep Learning VM](https://cloud.google.com/deep-learning-vm/) Spin it up with a Jupyterlab option enabled. \nYou can then go to the AI platform &gt;&gt; Open a Notebook. TF2.1 is pre-installed. If you need you can also spin a Pytorch based framework, \n\nI have some limited experience with AWS SageMaker. But I was not impressed due to its severe limitations &amp; high costs."
    }
  ],
  "comments": [
    {
      "id": 731947,
      "author_name": "Pedro Bernardo",
      "author_url": "",
      "post_date": "2020-01-29T09:34:40.590000",
      "content": "<p>I am using Amazon EC2 and my experience has been quite good. I'm not using SageMaker and Tensorflow though. I'm just setting up a normal VM and setting up my own Jupyter service, and using PyTorch.</p>\n\n<p>I wrote a script with everything I need to setup the VM (download Anaconda, install CUDA, update the packages I need and so on) that I run when after the VM starts. Since I'm using spot instances I made the script so everything I terminate one VM it becomes easier to start a new one.</p>\n\n<p>I haven't used SageMaker so I don't know how it works, but I guess using a normal VM gives more control over the setup, setting up your own Jupyter server should be easy, just google \"how to setup jupyter server on EC2\", there are many tutorials out there. After coming up with a script to config the stuff as you need it should become straightforward to run your workload without problems.</p>",
      "votes": 8,
      "replies": [
        {
          "id": 732048,
          "author_name": "dagnelies",
          "author_url": "",
          "post_date": "2020-01-29T12:26:08.453000",
          "content": "<p><a href=\"/pedromb\">@pedromb</a>  ...but, if you are running spot instances, doesn't it imply that the training might be interrupted in the middle? How do you deal with that?</p>\n\n<p>Although spot instances sounded nice, I wonder how to persist state and recover. Do you use periodic model saves during training? There is also the wonderful \"Batch\" jobs possibility, but you have to build Docker images and so on, so it is another layer of complexity. 😓 </p>\n\n<p>I guess there is no easy way around this.</p>\n\n<p>I am also curious about the google's TPU instances, but they also bring their share of complexity. Sadly, I applied a bit too late because I didn't read there were two rounds. 🙄 </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 732208,
          "author_name": "Pedro Bernardo",
          "author_url": "",
          "post_date": "2020-01-29T15:17:19.883000",
          "content": "<p>I do save checkpoints of the model during training, but, to be honest, so far I didn't have an issue with the spot instance being interrupted.</p>\n\n<p>For persistence I'm using an EBS volume on AWS. It's a bit more expensive than S3 but at least I can just unmount and mount back when I change instances, so it works well.</p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 747148,
          "author_name": "SeshuRaju 🧘‍♂️",
          "author_url": "",
          "post_date": "2020-02-16T03:00:07.550000",
          "content": "<p><a href=\"/pedromb\">@pedromb</a> like your approach, could you share what EC2 instance type to use for this competition (preprocessing and training)</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 747378,
          "author_name": "Pedro Bernardo",
          "author_url": "",
          "post_date": "2020-02-16T11:06:11.667000",
          "content": "<p>I'm using p3.2xlarge, the normal price is $3.06/h, but, using spot instances it goes for around $0.97/h. </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 747398,
          "author_name": "SeshuRaju 🧘‍♂️",
          "author_url": "",
          "post_date": "2020-02-16T11:32:39.033000",
          "content": "<p><a href=\"/pedromb\">@pedromb</a> Thanks that helps to setup AWS well use resources.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 732555,
      "author_name": "PC Jimmmy",
      "author_url": "",
      "post_date": "2020-01-29T23:45:58.057000",
      "content": "<p>As a very accurate rule - \"life on kaggle is very much easier than real life.\"</p>\n\n<p>When I first used Goggle Cloud Platform last year my 'Curse Jar' filled up pretty fast.  The whole subject of quota's can get me ranting for hours.  At the end I said \"never again\".  For this competition my team mate is filling his jar on AWS.</p>\n\n<p>ML on the cloud would seem to be an absolute skill set needed for a data scientist.   I am in a position were I can avoid learning that skill set (very old, not looking for a job and have some nice PC's) but for most it's a rite of passage that I think you need to take.</p>\n\n<p>See if you can't team up with someone who has gotten past the early part of the learning curve on this or another competition.  I got thru the ramp up on GCP thanks to a great partner in the Molecules challenge.  Put an ad in the team wanted discussion post - maybe some souls out there with the experience will join up with you.</p>",
      "votes": 4,
      "replies": [
        {
          "id": 746012,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-02-14T13:39:20.867000",
          "content": "<p>Very accurate. </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 732547,
      "author_name": "Shangqiu Li",
      "author_url": "",
      "post_date": "2020-01-29T23:34:30.767000",
      "content": "<p>The error you are getting because of incompatible cuDNN version and tensorflow version. Upgrade cuda and cuDNN version to the corresponding version will solve the problem.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 732667,
      "author_name": "SkyLord",
      "author_url": "",
      "post_date": "2020-01-30T04:37:47.413000",
      "content": "<p><a href=\"/dagnelies\">@dagnelies</a> if you have some GCP credits. You can use <a href=\"https://cloud.google.com/deep-learning-vm/\">GCP Deep Learning VM</a> Spin it up with a Jupyterlab option enabled. \nYou can then go to the AI platform &gt;&gt; Open a Notebook. TF2.1 is pre-installed. If you need you can also spin a Pytorch based framework, </p>\n\n<p>I have some limited experience with AWS SageMaker. But I was not impressed due to its severe limitations &amp; high costs.</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "731947": "I am using Amazon EC2 and my experience has been quite good. I'm not using SageMaker and Tensorflow though. I'm just setting up a normal VM and setting up my own Jupyter service, and using PyTorch.\n\nI wrote a script with everything I need to setup the VM (download Anaconda, install CUDA, update the packages I need and so on) that I run when after the VM starts. Since I'm using spot instances I made the script so everything I terminate one VM it becomes easier to start a new one.\n\nI haven't used SageMaker so I don't know how it works, but I guess using a normal VM gives more control over the setup, setting up your own Jupyter server should be easy, just google \"how to setup jupyter server on EC2\", there are many tutorials out there. After coming up with a script to config the stuff as you need it should become straightforward to run your workload without problems.",
    "731916": "Hi,\nare you successfully preprocessing/training your data with AWS/GCP?\nMy overall feeling is like:\n- On Kaggle: \"Oh wonderful, everything is so simple and productive!\"\n- On AWS/GCP: \"Oh my god, why is everything always so complicated?! I spend way more time to set things up than to try models!\"\n\nFor example, I saw SageMaker on AWS so I thought it'd be similar to Kaggle and I gladly used.\nBut it turns out to be another productivity loss.\n\nTheir SageMaker uses Tensorflow 1.x, so it's an incompatible version from Kaggle here which is 2.x. Despite amazon announced this January that SageMaker now supports TF 2.x, it is a mystery to me how to use it. Every notebook is still TF 1.x. So I thought I'd upgrade it manually within the notebook itself with \"pip --upgrade\" and it did. But then it simply failed at runtime with errors like:\n```\nUnknownError:  Failed to get convolution algorithm. This is probably because cuDNN failed to initialize, so try looking to see if a warning log message was printed above.\n\t [[node conv2d_1/convolution (defined at /home/ec2-user/anaconda3/envs/tensorflow_p36/lib/python3.6/site-packages/keras/backend/tensorflow_backend.py:3009) ]]\n```\nSo basically, you spend a lot of time, just for having your code not working in the end. And this is just one example among many. It's really frustrating. And if you use GPUs, your credits melt like ice cream in the sun. \n\nDo you have smoother experiences? Do you set up VMs from scratch?",
    "732555": "As a very accurate rule - \"life on kaggle is very much easier than real life.\"\n\nWhen I first used Goggle Cloud Platform last year my 'Curse Jar' filled up pretty fast.  The whole subject of quota's can get me ranting for hours.  At the end I said \"never again\".  For this competition my team mate is filling his jar on AWS.\n\nML on the cloud would seem to be an absolute skill set needed for a data scientist.   I am in a position were I can avoid learning that skill set (very old, not looking for a job and have some nice PC's) but for most it's a rite of passage that I think you need to take.\n\nSee if you can't team up with someone who has gotten past the early part of the learning curve on this or another competition.  I got thru the ramp up on GCP thanks to a great partner in the Molecules challenge.  Put an ad in the team wanted discussion post - maybe some souls out there with the experience will join up with you.\n\n\n\n",
    "732547": "The error you are getting because of incompatible cuDNN version and tensorflow version. Upgrade cuda and cuDNN version to the corresponding version will solve the problem.",
    "732667": "@dagnelies if you have some GCP credits. You can use [GCP Deep Learning VM](https://cloud.google.com/deep-learning-vm/) Spin it up with a Jupyterlab option enabled. \nYou can then go to the AI platform &gt;&gt; Open a Notebook. TF2.1 is pre-installed. If you need you can also spin a Pytorch based framework, \n\nI have some limited experience with AWS SageMaker. But I was not impressed due to its severe limitations &amp; high costs."
  }
}