{
  "id": 131981,
  "title": "Training in Kaggle vs Colab vs SageMaker (ml.p2.xlarge)",
  "url": "/competitions/deepfake-detection-challenge/discussion/131981",
  "author_name": "Debanga Raj Neog",
  "post_date": "2020-02-23T04:54:03.957000",
  "votes": 3,
  "comment_count": 18,
  "views": 0,
  "content": "<p>I am running my same training script (pytorch) in 3 platforms: Kaggle, SageMaker (ml.p2.xlarge), and Google Colab. Interestingly, the training speed is:\n<code>\nKaggle &amp;gt; Colab &amp;gt; SageMaker\n</code>\nIs it normal? So, <code>ml.p2.xlarge</code> is literally slower than free Colab GPU, while AWS charges $1.26/hr? :D</p>",
  "messages": [
    {
      "id": 754130,
      "postDate": "2020-02-23T04:54:03.957Z",
      "content": "<p>I am running my same training script (pytorch) in 3 platforms: Kaggle, SageMaker (ml.p2.xlarge), and Google Colab. Interestingly, the training speed is:\n<code>\nKaggle &amp;gt; Colab &amp;gt; SageMaker\n</code>\nIs it normal? So, <code>ml.p2.xlarge</code> is literally slower than free Colab GPU, while AWS charges $1.26/hr? :D</p>",
      "rawMarkdown": "I am running my same training script (pytorch) in 3 platforms: Kaggle, SageMaker (ml.p2.xlarge), and Google Colab. Interestingly, the training speed is:\n```\nKaggle &gt; Colab &gt; SageMaker\n```\nIs it normal? So, ```ml.p2.xlarge``` is literally slower than free Colab GPU, while AWS charges $1.26/hr? :D\n\n\n\n\n",
      "votes": 3
    },
    {
      "id": 755835,
      "postDate": "2020-02-25T06:56:58.343Z",
      "content": "<p><a href=\"/debanga\">@debanga</a>  I would strongly suggest you'd get Colab Pro ... no disconnects for 24 hours and guaranteed P100 or T4</p>",
      "rawMarkdown": "@debanga  I would strongly suggest you'd get Colab Pro ... no disconnects for 24 hours and guaranteed P100 or T4",
      "votes": 1,
      "replies": [
        {
          "id": 756296,
          "postDate": "2020-02-25T15:38:28.430Z",
          "content": "<p>Thank you for the \"Pro\"tip! :D</p>",
          "rawMarkdown": "Thank you for the \"Pro\"tip! :D"
        },
        {
          "id": 756302,
          "postDate": "2020-02-25T15:44:10.687Z",
          "content": "<p>I am very interested, could you please share your experience with Colab Pro?</p>",
          "rawMarkdown": "I am very interested, could you please share your experience with Colab Pro?"
        },
        {
          "id": 756305,
          "postDate": "2020-02-25T15:45:27.003Z",
          "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2279276%2F2736963b13c88de65bdc93d130e4f554%2FAnnotation%202020-02-25%20104324.png?generation=1582645513162931&amp;alt=media\" alt=\"\"></p>\n\n<p>This is what they have.</p>",
          "rawMarkdown": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2279276%2F2736963b13c88de65bdc93d130e4f554%2FAnnotation%202020-02-25%20104324.png?generation=1582645513162931&amp;alt=media)\n\nThis is what they have."
        },
        {
          "id": 756395,
          "postDate": "2020-02-25T17:22:16.287Z",
          "content": "<p>Oh no! Only available in US. We poor Canadians.\n<a href=\"https://stackoverflow.com/questions/60120703/is-google-colab-pro-available-in-any-state\">https://stackoverflow.com/questions/60120703/is-google-colab-pro-available-in-any-state</a></p>",
          "rawMarkdown": "Oh no! Only available in US. We poor Canadians.\nhttps://stackoverflow.com/questions/60120703/is-google-colab-pro-available-in-any-state"
        }
      ]
    },
    {
      "id": 754614,
      "postDate": "2020-02-23T20:41:59.557Z",
      "content": "<p>I'll suggest to go with ec2. I use g4dn.xlarge(around $0.6/hr) wich has T100 and to me it is faster then kaggle's gpu. </p>",
      "rawMarkdown": "I'll suggest to go with ec2. I use g4dn.xlarge(around $0.6/hr) wich has T100 and to me it is faster then kaggle's gpu. ",
      "votes": 1
    },
    {
      "id": 754158,
      "postDate": "2020-02-23T06:07:41.317Z",
      "content": "<p>As far as I know, Kaggle Notebook's GPU-accelerator uses P100. The P100 is significantly faster than the K80(which SageMaker ml.p2.xlarge instance uses).</p>",
      "rawMarkdown": "As far as I know, Kaggle Notebook's GPU-accelerator uses P100. The P100 is significantly faster than the K80(which SageMaker ml.p2.xlarge instance uses).",
      "votes": 1,
      "replies": [
        {
          "id": 754159,
          "postDate": "2020-02-23T06:09:59.723Z",
          "content": "<p>Yes, 1 epoch=20 min in Kaggle vs 1 epoch = 45 min in SageMaker!</p>",
          "rawMarkdown": "Yes, 1 epoch=20 min in Kaggle vs 1 epoch = 45 min in SageMaker!"
        },
        {
          "id": 754163,
          "postDate": "2020-02-23T06:21:05.627Z",
          "content": "<p>1 epoch with how much data?\nMy single epoch takes hours with 2 million images.</p>",
          "rawMarkdown": "1 epoch with how much data?\nMy single epoch takes hours with 2 million images."
        },
        {
          "id": 754164,
          "postDate": "2020-02-23T06:24:10.927Z",
          "content": "<p>Kaggle 20 min for 120K/120K real/fake for 1 epoch.</p>",
          "rawMarkdown": "Kaggle 20 min for 120K/120K real/fake for 1 epoch.",
          "votes": 1
        },
        {
          "id": 754380,
          "postDate": "2020-02-23T13:33:49.897Z",
          "content": "<p>what is the size of the images? way too fast training speed :)</p>",
          "rawMarkdown": "what is the size of the images? way too fast training speed :)"
        },
        {
          "id": 754392,
          "postDate": "2020-02-23T13:47:18.807Z",
          "content": "<p>224, Xception.</p>",
          "rawMarkdown": "224, Xception."
        },
        {
          "id": 755061,
          "postDate": "2020-02-24T11:46:43.020Z",
          "content": "<p>total 240K images training takes 20min per epoch?.. how is that possible? are you loading all the images to the dataloader or something?</p>",
          "rawMarkdown": "total 240K images training takes 20min per epoch?.. how is that possible? are you loading all the images to the dataloader or something?"
        },
        {
          "id": 756271,
          "postDate": "2020-02-25T15:17:50.487Z",
          "content": "<p>I am using a dataloader (from a single folder)</p>\n\n<p>Also, I guess training time also depends on size of model, how many layers you are training etc.</p>",
          "rawMarkdown": "I am using a dataloader (from a single folder)\n\nAlso, I guess training time also depends on size of model, how many layers you are training etc."
        },
        {
          "id": 757759,
          "postDate": "2020-02-27T04:38:39.817Z",
          "content": "<p>ah maybe you are freezing layers :)</p>",
          "rawMarkdown": "ah maybe you are freezing layers :)"
        }
      ]
    },
    {
      "id": 754137,
      "postDate": "2020-02-23T05:18:44.123Z",
      "content": "<p>My main training environment is my own PC(a $1500 machine with gtx 1080 ti card). So far it is good enough. The difficulties I have faced was: 1- downloading the data, 2- GPU fans are noisy while training:)</p>",
      "rawMarkdown": "My main training environment is my own PC(a $1500 machine with gtx 1080 ti card). So far it is good enough. The difficulties I have faced was: 1- downloading the data, 2- GPU fans are noisy while training:)",
      "replies": [
        {
          "id": 754143,
          "postDate": "2020-02-23T05:29:46.207Z",
          "content": "<p>Never thought about the noise :D Since, I don't own any GPU PC, I bet I would have hard time sleeping in my dorm room during the overnight training. Haha.</p>",
          "rawMarkdown": "Never thought about the noise :D Since, I don't own any GPU PC, I bet I would have hard time sleeping in my dorm room during the overnight training. Haha."
        }
      ]
    },
    {
      "id": 754155,
      "postDate": "2020-02-23T06:02:48.777Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 755835,
      "author_name": "mjsML",
      "author_url": "",
      "post_date": "2020-02-25T06:56:58.343000",
      "content": "<p><a href=\"/debanga\">@debanga</a>  I would strongly suggest you'd get Colab Pro ... no disconnects for 24 hours and guaranteed P100 or T4</p>",
      "votes": 1,
      "replies": [
        {
          "id": 756296,
          "author_name": "Debanga Raj Neog",
          "author_url": "",
          "post_date": "2020-02-25T15:38:28.430000",
          "content": "<p>Thank you for the \"Pro\"tip! :D</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 756302,
          "author_name": "Debanga Raj Neog",
          "author_url": "",
          "post_date": "2020-02-25T15:44:10.687000",
          "content": "<p>I am very interested, could you please share your experience with Colab Pro?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 756305,
          "author_name": "Debanga Raj Neog",
          "author_url": "",
          "post_date": "2020-02-25T15:45:27.003000",
          "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2279276%2F2736963b13c88de65bdc93d130e4f554%2FAnnotation%202020-02-25%20104324.png?generation=1582645513162931&amp;alt=media\" alt=\"\"></p>\n\n<p>This is what they have.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 756395,
          "author_name": "Debanga Raj Neog",
          "author_url": "",
          "post_date": "2020-02-25T17:22:16.287000",
          "content": "<p>Oh no! Only available in US. We poor Canadians.\n<a href=\"https://stackoverflow.com/questions/60120703/is-google-colab-pro-available-in-any-state\">https://stackoverflow.com/questions/60120703/is-google-colab-pro-available-in-any-state</a></p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 754614,
      "author_name": "Ankit Saini",
      "author_url": "",
      "post_date": "2020-02-23T20:41:59.557000",
      "content": "<p>I'll suggest to go with ec2. I use g4dn.xlarge(around $0.6/hr) wich has T100 and to me it is faster then kaggle's gpu. </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 754158,
      "author_name": "Caffeinism🤔",
      "author_url": "",
      "post_date": "2020-02-23T06:07:41.317000",
      "content": "<p>As far as I know, Kaggle Notebook's GPU-accelerator uses P100. The P100 is significantly faster than the K80(which SageMaker ml.p2.xlarge instance uses).</p>",
      "votes": 1,
      "replies": [
        {
          "id": 754159,
          "author_name": "Debanga Raj Neog",
          "author_url": "",
          "post_date": "2020-02-23T06:09:59.723000",
          "content": "<p>Yes, 1 epoch=20 min in Kaggle vs 1 epoch = 45 min in SageMaker!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 754163,
          "author_name": "Emre Bayram",
          "author_url": "",
          "post_date": "2020-02-23T06:21:05.627000",
          "content": "<p>1 epoch with how much data?\nMy single epoch takes hours with 2 million images.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 754164,
          "author_name": "Debanga Raj Neog",
          "author_url": "",
          "post_date": "2020-02-23T06:24:10.927000",
          "content": "<p>Kaggle 20 min for 120K/120K real/fake for 1 epoch.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 754380,
          "author_name": "saewonYang",
          "author_url": "",
          "post_date": "2020-02-23T13:33:49.897000",
          "content": "<p>what is the size of the images? way too fast training speed :)</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 754392,
          "author_name": "Debanga Raj Neog",
          "author_url": "",
          "post_date": "2020-02-23T13:47:18.807000",
          "content": "<p>224, Xception.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 755061,
          "author_name": "saewonYang",
          "author_url": "",
          "post_date": "2020-02-24T11:46:43.020000",
          "content": "<p>total 240K images training takes 20min per epoch?.. how is that possible? are you loading all the images to the dataloader or something?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 756271,
          "author_name": "Debanga Raj Neog",
          "author_url": "",
          "post_date": "2020-02-25T15:17:50.487000",
          "content": "<p>I am using a dataloader (from a single folder)</p>\n\n<p>Also, I guess training time also depends on size of model, how many layers you are training etc.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 757759,
          "author_name": "saewonYang",
          "author_url": "",
          "post_date": "2020-02-27T04:38:39.817000",
          "content": "<p>ah maybe you are freezing layers :)</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 754137,
      "author_name": "Emre Bayram",
      "author_url": "",
      "post_date": "2020-02-23T05:18:44.123000",
      "content": "<p>My main training environment is my own PC(a $1500 machine with gtx 1080 ti card). So far it is good enough. The difficulties I have faced was: 1- downloading the data, 2- GPU fans are noisy while training:)</p>",
      "votes": 0,
      "replies": [
        {
          "id": 754143,
          "author_name": "Debanga Raj Neog",
          "author_url": "",
          "post_date": "2020-02-23T05:29:46.207000",
          "content": "<p>Never thought about the noise :D Since, I don't own any GPU PC, I bet I would have hard time sleeping in my dorm room during the overnight training. Haha.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 754155,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-02-23T06:02:48.777000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "754130": "I am running my same training script (pytorch) in 3 platforms: Kaggle, SageMaker (ml.p2.xlarge), and Google Colab. Interestingly, the training speed is:\n```\nKaggle &gt; Colab &gt; SageMaker\n```\nIs it normal? So, ```ml.p2.xlarge``` is literally slower than free Colab GPU, while AWS charges $1.26/hr? :D\n\n\n\n\n",
    "755835": "@debanga  I would strongly suggest you'd get Colab Pro ... no disconnects for 24 hours and guaranteed P100 or T4",
    "754614": "I'll suggest to go with ec2. I use g4dn.xlarge(around $0.6/hr) wich has T100 and to me it is faster then kaggle's gpu. ",
    "754158": "As far as I know, Kaggle Notebook's GPU-accelerator uses P100. The P100 is significantly faster than the K80(which SageMaker ml.p2.xlarge instance uses).",
    "754137": "My main training environment is my own PC(a $1500 machine with gtx 1080 ti card). So far it is good enough. The difficulties I have faced was: 1- downloading the data, 2- GPU fans are noisy while training:)",
    "754155": ""
  }
}