{
  "id": 98528,
  "title": "Ways to try different things faster?",
  "url": "/competitions/aptos2019-blindness-detection/discussion/98528",
  "author_name": "",
  "post_date": "2019-07-04T11:16:01.994935100Z",
  "votes": 1,
  "comment_count": 5,
  "views": 0,
  "content": "<p>I am trying to check various things like different models, image augmentations, etc. For checking purpose only, I am training for 2 epochs and it takes more than 12 Min for ResNet50 to train. Any ideas to speed up things.</p>",
  "messages": [
    {
      "id": "568078",
      "postDate": "07/04/2019 11:16:01",
      "content": "<p>I am trying to check various things like different models, image augmentations, etc. For checking purpose only, I am training for 2 epochs and it takes more than 12 Min for ResNet50 to train. Any ideas to speed up things.</p>",
      "rawMarkdown": "I am trying to check various things like different models, image augmentations, etc. For checking purpose only, I am training for 2 epochs and it takes more than 12 Min for ResNet50 to train. Any ideas to speed up things.",
      "votes": null
    },
    {
      "id": "568232",
      "postDate": "07/04/2019 15:15:39",
      "content": "<p>Do you use kernel for training? If yes, you can try make I/O overhead smaller. I mean:\n1. Read all images to memory (of course resize them). Not sure if will fir to memory\n2. Download dataset locally, resize it, upload to you kernel. </p>\n\n<p>I use already resized image locally and then I have 4x speed up of training. One epoch take 30s at 1080.</p>",
      "rawMarkdown": "Do you use kernel for training? If yes, you can try make I/O overhead smaller. I mean:\n1. Read all images to memory (of course resize them). Not sure if will fir to memory\n2. Download dataset locally, resize it, upload to you kernel. \n\nI use already resized image locally and then I have 4x speed up of training. One epoch take 30s at 1080.",
      "votes": null
    },
    {
      "id": "568670",
      "postDate": "07/05/2019 09:29:08",
      "content": "<p>After downloading the dataset locally, to what size you are resizing to it?</p>",
      "rawMarkdown": "After downloading the dataset locally, to what size you are resizing to it?",
      "votes": null
    },
    {
      "id": "568826",
      "postDate": "07/05/2019 13:50:09",
      "content": "<p>It helped me by bringing 1 epoch to 60 sec from 400+ sec. Thanks</p>",
      "rawMarkdown": "It helped me by bringing 1 epoch to 60 sec from 400+ sec. Thanks",
      "votes": null
    },
    {
      "id": "569509",
      "postDate": "07/06/2019 19:25:57",
      "content": "<p>Hi! <a href=\"/nitin29\">@nitin29</a>, I'm already thinking to run GCP ML Engine or AWS SageMaker with Hyperparameter tunning in a powerfull VM with GPUs, other options are reducing the image size, datasets size, or the number of epochs but they can generate different results and point you in the wrong direction.</p>",
      "rawMarkdown": "Hi! @nitin29, I'm already thinking to run GCP ML Engine or AWS SageMaker with Hyperparameter tunning in a powerfull VM with GPUs, other options are reducing the image size, datasets size, or the number of epochs but they can generate different results and point you in the wrong direction.",
      "votes": null
    },
    {
      "id": "569556",
      "postDate": "07/06/2019 23:03:34",
      "content": "<p>Pytorch and Apex for training with mixed precision can be a big time saver, if you have appropriate GPU :)</p>\n\n<p><a href=\"https://github.com/NVIDIA/apex\">https://github.com/NVIDIA/apex</a></p>",
      "rawMarkdown": "Pytorch and Apex for training with mixed precision can be a big time saver, if you have appropriate GPU :)\n\nhttps://github.com/NVIDIA/apex",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 568232,
      "author_name": "melgor",
      "author_url": "",
      "post_date": "07/04/2019 15:15:39",
      "content": "<p>Do you use kernel for training? If yes, you can try make I/O overhead smaller. I mean:\n1. Read all images to memory (of course resize them). Not sure if will fir to memory\n2. Download dataset locally, resize it, upload to you kernel. </p>\n\n<p>I use already resized image locally and then I have 4x speed up of training. One epoch take 30s at 1080.</p>",
      "votes": null,
      "replies": [
        {
          "id": 568670,
          "author_name": "nitin29",
          "author_url": "",
          "post_date": "07/05/2019 09:29:08",
          "content": "<p>After downloading the dataset locally, to what size you are resizing to it?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 568826,
          "author_name": "nitin29",
          "author_url": "",
          "post_date": "07/05/2019 13:50:09",
          "content": "<p>It helped me by bringing 1 epoch to 60 sec from 400+ sec. Thanks</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 569509,
      "author_name": "cv13j0",
      "author_url": "",
      "post_date": "07/06/2019 19:25:57",
      "content": "<p>Hi! <a href=\"/nitin29\">@nitin29</a>, I'm already thinking to run GCP ML Engine or AWS SageMaker with Hyperparameter tunning in a powerfull VM with GPUs, other options are reducing the image size, datasets size, or the number of epochs but they can generate different results and point you in the wrong direction.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 569556,
      "author_name": "taindow",
      "author_url": "",
      "post_date": "07/06/2019 23:03:34",
      "content": "<p>Pytorch and Apex for training with mixed precision can be a big time saver, if you have appropriate GPU :)</p>\n\n<p><a href=\"https://github.com/NVIDIA/apex\">https://github.com/NVIDIA/apex</a></p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "568078": "I am trying to check various things like different models, image augmentations, etc. For checking purpose only, I am training for 2 epochs and it takes more than 12 Min for ResNet50 to train. Any ideas to speed up things.",
    "568232": "Do you use kernel for training? If yes, you can try make I/O overhead smaller. I mean:\n1. Read all images to memory (of course resize them). Not sure if will fir to memory\n2. Download dataset locally, resize it, upload to you kernel. \n\nI use already resized image locally and then I have 4x speed up of training. One epoch take 30s at 1080.",
    "568670": "After downloading the dataset locally, to what size you are resizing to it?",
    "568826": "It helped me by bringing 1 epoch to 60 sec from 400+ sec. Thanks",
    "569509": "Hi! @nitin29, I'm already thinking to run GCP ML Engine or AWS SageMaker with Hyperparameter tunning in a powerfull VM with GPUs, other options are reducing the image size, datasets size, or the number of epochs but they can generate different results and point you in the wrong direction.",
    "569556": "Pytorch and Apex for training with mixed precision can be a big time saver, if you have appropriate GPU :)\n\nhttps://github.com/NVIDIA/apex"
  },
  "source": "meta"
}