{
  "id": 122842,
  "title": "AWS GPU instances",
  "url": "/competitions/deepfake-detection-challenge/discussion/122842",
  "author_name": "",
  "post_date": "2019-12-23T07:00:16.388402400Z",
  "votes": 8,
  "comment_count": 10,
  "views": 0,
  "content": "<h1>4 types of GPU Instances - p3, p2, g4, g3</h1>\n\n<h2>NVIDIA Tesla V100</h2>\n\n<p><strong>Price of On-Demand Instances in US West(Oregon)</strong>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3876174%2F8eda302fa1a476160c304c1670216b31%2F2019-12-22_201350.png?generation=1577017171295898&amp;alt=media\" alt=\"\"></p>\n\n<h2>NVIDIA K80</h2>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3876174%2F9cad9050b51217568e85d13d4726e2f6%2F2019-12-22_202243.png?generation=1577017554194915&amp;alt=media\" alt=\"\"></p>\n\n<h2>NVIDIA T4</h2>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3876174%2F304ce27842cd956c334f35c6d36dbea2%2FFastStoneEditor.png?generation=1577019025827016&amp;alt=media\" alt=\"\"></p>\n\n<h2>NVIDIA Tesla M60</h2>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3876174%2Fadd110075bc164f228671e9b19c8f960%2FFastStoneEditor.png?generation=1577018731950254&amp;alt=media\" alt=\"\"></p>\n\n<p>Please upvote, if you think it's useful.</p>",
  "messages": [
    {
      "id": "701160",
      "postDate": "12/23/2019 07:00:16",
      "content": "<h1>4 types of GPU Instances - p3, p2, g4, g3</h1>\n\n<h2>NVIDIA Tesla V100</h2>\n\n<p><strong>Price of On-Demand Instances in US West(Oregon)</strong>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3876174%2F8eda302fa1a476160c304c1670216b31%2F2019-12-22_201350.png?generation=1577017171295898&amp;alt=media\" alt=\"\"></p>\n\n<h2>NVIDIA K80</h2>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3876174%2F9cad9050b51217568e85d13d4726e2f6%2F2019-12-22_202243.png?generation=1577017554194915&amp;alt=media\" alt=\"\"></p>\n\n<h2>NVIDIA T4</h2>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3876174%2F304ce27842cd956c334f35c6d36dbea2%2FFastStoneEditor.png?generation=1577019025827016&amp;alt=media\" alt=\"\"></p>\n\n<h2>NVIDIA Tesla M60</h2>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3876174%2Fadd110075bc164f228671e9b19c8f960%2FFastStoneEditor.png?generation=1577018731950254&amp;alt=media\" alt=\"\"></p>\n\n<p>Please upvote, if you think it's useful.</p>",
      "rawMarkdown": "# 4 types of GPU Instances - p3, p2, g4, g3\n## NVIDIA Tesla V100 \n**Price of On-Demand Instances in US West(Oregon)**\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3876174%2F8eda302fa1a476160c304c1670216b31%2F2019-12-22_201350.png?generation=1577017171295898&amp;alt=media)\n\n## NVIDIA K80\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3876174%2F9cad9050b51217568e85d13d4726e2f6%2F2019-12-22_202243.png?generation=1577017554194915&amp;alt=media)\n\n## NVIDIA T4\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3876174%2F304ce27842cd956c334f35c6d36dbea2%2FFastStoneEditor.png?generation=1577019025827016&amp;alt=media)\n\n## NVIDIA Tesla M60\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3876174%2Fadd110075bc164f228671e9b19c8f960%2FFastStoneEditor.png?generation=1577018731950254&amp;alt=media)\n\nPlease upvote, if you think it's useful.",
      "votes": null
    },
    {
      "id": "701453",
      "postDate": "12/23/2019 14:02:35",
      "content": "<p>any experience on what tiers you need to get good results? my feeling is the lowest tiers are more than enough to do any kaggle contest except video stuff</p>",
      "rawMarkdown": "any experience on what tiers you need to get good results? my feeling is the lowest tiers are more than enough to do any kaggle contest except video stuff",
      "votes": null
    },
    {
      "id": "701480",
      "postDate": "12/23/2019 14:26:08",
      "content": "<p>It's hard to choose, I'm not sure how much Memory we need, anyone give some suggestion?</p>",
      "rawMarkdown": "It's hard to choose, I'm not sure how much Memory we need, anyone give some suggestion?",
      "votes": null
    },
    {
      "id": "701850",
      "postDate": "12/24/2019 01:27:11",
      "content": "<p>p3.2xlarge was $0.91 in Spot instances these days</p>",
      "rawMarkdown": "p3.2xlarge was $0.91 in Spot instances these days",
      "votes": null
    },
    {
      "id": "701867",
      "postDate": "12/24/2019 01:55:42",
      "content": "<p>Yeah, it's much cheaper but floating, so I don't list it.</p>",
      "rawMarkdown": "Yeah, it's much cheaper but floating, so I don't list it.",
      "votes": null
    },
    {
      "id": "705328",
      "postDate": "12/28/2019 19:25:26",
      "content": "<p>This is my first time using AWS. Will AWS use credits by default or I have to set it up somewhere.</p>",
      "rawMarkdown": "This is my first time using AWS. Will AWS use credits by default or I have to set it up somewhere.",
      "votes": null
    },
    {
      "id": "709600",
      "postDate": "01/03/2020 17:55:02",
      "content": "<p>I tested all options in different use cases. My choices:\n- training: p3-2xlarge\n- inference: g4dn-4xlarge</p>\n\n<p>Important: I’m using multiprocessing in inference, all vCPUs sharing the single GPU.\nAnd I’m always using spot instances. They are great!</p>",
      "rawMarkdown": "I tested all options in different use cases. My choices:\n- training: p3-2xlarge\n- inference: g4dn-4xlarge\n\nImportant: I’m using multiprocessing in inference, all vCPUs sharing the single GPU.\nAnd I’m always using spot instances. They are great!",
      "votes": null
    },
    {
      "id": "709872",
      "postDate": "01/04/2020 03:05:27",
      "content": "<p>Thanks for sharing your experience.</p>",
      "rawMarkdown": "Thanks for sharing your experience.",
      "votes": null
    },
    {
      "id": "710483",
      "postDate": "01/04/2020 19:32:27",
      "content": "<p>thanks Carlos, how many hours did you need to buy to train it?</p>",
      "rawMarkdown": "thanks Carlos, how many hours did you need to buy to train it?",
      "votes": null
    },
    {
      "id": "710496",
      "postDate": "01/04/2020 19:59:26",
      "content": "<p>I got the credits from AWS, and I'm currently using it... so, I'm not spending anything now.\nBefore I got AWS credits (which happened on Dec-30th), I was using Google Cloud. Spent ~$80 there, but could have spent much less... I did several mistakes in setting up VMs, a lot of re-work (especially in installing/compiling OpenCV with GPU support).</p>\n\n<p>Training depends on a lot of variables: the network you are using, how many epochs, how much data you are using in training, how fast your algorithm actually is, etc... So far I'm using really deep NNs, pre-trained models, trained for no longer than 20 epochs, and currently using %5 of the data: my training procedures last approx 4 hours using EC2 p3-2xlarge. Their spot cost are approx $1/h, so approx $4/trained model. Considering we got $500 in AWS credit, we can run it without any problems. But the cost will increase as we start to use more data. I'm planning to move from 5% to 15-20% in the next days :)</p>",
      "rawMarkdown": "I got the credits from AWS, and I'm currently using it... so, I'm not spending anything now.\nBefore I got AWS credits (which happened on Dec-30th), I was using Google Cloud. Spent ~$80 there, but could have spent much less... I did several mistakes in setting up VMs, a lot of re-work (especially in installing/compiling OpenCV with GPU support).\n\nTraining depends on a lot of variables: the network you are using, how many epochs, how much data you are using in training, how fast your algorithm actually is, etc... So far I'm using really deep NNs, pre-trained models, trained for no longer than 20 epochs, and currently using %5 of the data: my training procedures last approx 4 hours using EC2 p3-2xlarge. Their spot cost are approx $1/h, so approx $4/trained model. Considering we got $500 in AWS credit, we can run it without any problems. But the cost will increase as we start to use more data. I'm planning to move from 5% to 15-20% in the next days :)",
      "votes": null
    },
    {
      "id": "723070",
      "postDate": "01/19/2020 13:15:27",
      "content": "<p>Thanks Carlos, did you end up with any storage limitation problems, if you move to the full data set I think you can't use the smaller VMs and they are slow to upload or so I have been told.</p>",
      "rawMarkdown": "Thanks Carlos, did you end up with any storage limitation problems, if you move to the full data set I think you can't use the smaller VMs and they are slow to upload or so I have been told.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 701453,
      "author_name": "joe1dataset",
      "author_url": "",
      "post_date": "12/23/2019 14:02:35",
      "content": "<p>any experience on what tiers you need to get good results? my feeling is the lowest tiers are more than enough to do any kaggle contest except video stuff</p>",
      "votes": null,
      "replies": [
        {
          "id": 701480,
          "author_name": "diegojohnson",
          "author_url": "",
          "post_date": "12/23/2019 14:26:08",
          "content": "<p>It's hard to choose, I'm not sure how much Memory we need, anyone give some suggestion?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 709600,
          "author_name": "carlossouza",
          "author_url": "",
          "post_date": "01/03/2020 17:55:02",
          "content": "<p>I tested all options in different use cases. My choices:\n- training: p3-2xlarge\n- inference: g4dn-4xlarge</p>\n\n<p>Important: I’m using multiprocessing in inference, all vCPUs sharing the single GPU.\nAnd I’m always using spot instances. They are great!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 709872,
          "author_name": "diegojohnson",
          "author_url": "",
          "post_date": "01/04/2020 03:05:27",
          "content": "<p>Thanks for sharing your experience.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 710483,
          "author_name": "joe1dataset",
          "author_url": "",
          "post_date": "01/04/2020 19:32:27",
          "content": "<p>thanks Carlos, how many hours did you need to buy to train it?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 710496,
          "author_name": "carlossouza",
          "author_url": "",
          "post_date": "01/04/2020 19:59:26",
          "content": "<p>I got the credits from AWS, and I'm currently using it... so, I'm not spending anything now.\nBefore I got AWS credits (which happened on Dec-30th), I was using Google Cloud. Spent ~$80 there, but could have spent much less... I did several mistakes in setting up VMs, a lot of re-work (especially in installing/compiling OpenCV with GPU support).</p>\n\n<p>Training depends on a lot of variables: the network you are using, how many epochs, how much data you are using in training, how fast your algorithm actually is, etc... So far I'm using really deep NNs, pre-trained models, trained for no longer than 20 epochs, and currently using %5 of the data: my training procedures last approx 4 hours using EC2 p3-2xlarge. Their spot cost are approx $1/h, so approx $4/trained model. Considering we got $500 in AWS credit, we can run it without any problems. But the cost will increase as we start to use more data. I'm planning to move from 5% to 15-20% in the next days :)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 723070,
          "author_name": "joe1dataset",
          "author_url": "",
          "post_date": "01/19/2020 13:15:27",
          "content": "<p>Thanks Carlos, did you end up with any storage limitation problems, if you move to the full data set I think you can't use the smaller VMs and they are slow to upload or so I have been told.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 701850,
      "author_name": "bguberfain",
      "author_url": "",
      "post_date": "12/24/2019 01:27:11",
      "content": "<p>p3.2xlarge was $0.91 in Spot instances these days</p>",
      "votes": null,
      "replies": [
        {
          "id": 701867,
          "author_name": "diegojohnson",
          "author_url": "",
          "post_date": "12/24/2019 01:55:42",
          "content": "<p>Yeah, it's much cheaper but floating, so I don't list it.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 705328,
      "author_name": "unkownhihi",
      "author_url": "",
      "post_date": "12/28/2019 19:25:26",
      "content": "<p>This is my first time using AWS. Will AWS use credits by default or I have to set it up somewhere.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "701160": "# 4 types of GPU Instances - p3, p2, g4, g3\n## NVIDIA Tesla V100 \n**Price of On-Demand Instances in US West(Oregon)**\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3876174%2F8eda302fa1a476160c304c1670216b31%2F2019-12-22_201350.png?generation=1577017171295898&amp;alt=media)\n\n## NVIDIA K80\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3876174%2F9cad9050b51217568e85d13d4726e2f6%2F2019-12-22_202243.png?generation=1577017554194915&amp;alt=media)\n\n## NVIDIA T4\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3876174%2F304ce27842cd956c334f35c6d36dbea2%2FFastStoneEditor.png?generation=1577019025827016&amp;alt=media)\n\n## NVIDIA Tesla M60\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3876174%2Fadd110075bc164f228671e9b19c8f960%2FFastStoneEditor.png?generation=1577018731950254&amp;alt=media)\n\nPlease upvote, if you think it's useful.",
    "701453": "any experience on what tiers you need to get good results? my feeling is the lowest tiers are more than enough to do any kaggle contest except video stuff",
    "701480": "It's hard to choose, I'm not sure how much Memory we need, anyone give some suggestion?",
    "701850": "p3.2xlarge was $0.91 in Spot instances these days",
    "701867": "Yeah, it's much cheaper but floating, so I don't list it.",
    "705328": "This is my first time using AWS. Will AWS use credits by default or I have to set it up somewhere.",
    "709600": "I tested all options in different use cases. My choices:\n- training: p3-2xlarge\n- inference: g4dn-4xlarge\n\nImportant: I’m using multiprocessing in inference, all vCPUs sharing the single GPU.\nAnd I’m always using spot instances. They are great!",
    "709872": "Thanks for sharing your experience.",
    "710483": "thanks Carlos, how many hours did you need to buy to train it?",
    "710496": "I got the credits from AWS, and I'm currently using it... so, I'm not spending anything now.\nBefore I got AWS credits (which happened on Dec-30th), I was using Google Cloud. Spent ~$80 there, but could have spent much less... I did several mistakes in setting up VMs, a lot of re-work (especially in installing/compiling OpenCV with GPU support).\n\nTraining depends on a lot of variables: the network you are using, how many epochs, how much data you are using in training, how fast your algorithm actually is, etc... So far I'm using really deep NNs, pre-trained models, trained for no longer than 20 epochs, and currently using %5 of the data: my training procedures last approx 4 hours using EC2 p3-2xlarge. Their spot cost are approx $1/h, so approx $4/trained model. Considering we got $500 in AWS credit, we can run it without any problems. But the cost will increase as we start to use more data. I'm planning to move from 5% to 15-20% in the next days :)",
    "723070": "Thanks Carlos, did you end up with any storage limitation problems, if you move to the full data set I think you can't use the smaller VMs and they are slow to upload or so I have been told."
  },
  "source": "meta"
}