{
  "id": 102503,
  "title": "I would like to ask how does everybody deal with large scale image data",
  "url": "/competitions/aptos2019-blindness-detection/discussion/102503",
  "author_name": "",
  "post_date": "2019-08-02T09:17:57.923342900Z",
  "votes": null,
  "comment_count": 4,
  "views": 0,
  "content": "<p>I think this is a dumb question, but I would like to ask how does everyone deal with large scale image data. I'm using a laptop that only contains CPU, so I use an cloud GPU machine floydhub to help run deep learning models so far. However, the expense is expensive and the cloud storage is small, so I'm in trouble of finding a way to store large scale dataset and run the program on it. </p>\n\n<p>I've found two methods that may help: one is to use free trial usage like google cloud, it seems to provide limit usage of VM with GPU; and another is to use Spark, that also need to build multiple VM to fulfill distribute computing. </p>\n\n<p>I would like to ask for advice, as a use that don't have powerful GPU equipped on my PC, what is the easiest way to run big dataset like this competition, and run program on them successfully? </p>",
  "messages": [
    {
      "id": "590514",
      "postDate": "08/02/2019 09:17:57",
      "content": "<p>I think this is a dumb question, but I would like to ask how does everyone deal with large scale image data. I'm using a laptop that only contains CPU, so I use an cloud GPU machine floydhub to help run deep learning models so far. However, the expense is expensive and the cloud storage is small, so I'm in trouble of finding a way to store large scale dataset and run the program on it. </p>\n\n<p>I've found two methods that may help: one is to use free trial usage like google cloud, it seems to provide limit usage of VM with GPU; and another is to use Spark, that also need to build multiple VM to fulfill distribute computing. </p>\n\n<p>I would like to ask for advice, as a use that don't have powerful GPU equipped on my PC, what is the easiest way to run big dataset like this competition, and run program on them successfully? </p>",
      "rawMarkdown": "I think this is a dumb question, but I would like to ask how does everyone deal with large scale image data. I'm using a laptop that only contains CPU, so I use an cloud GPU machine floydhub to help run deep learning models so far. However, the expense is expensive and the cloud storage is small, so I'm in trouble of finding a way to store large scale dataset and run the program on it. \n\nI've found two methods that may help: one is to use free trial usage like google cloud, it seems to provide limit usage of VM with GPU; and another is to use Spark, that also need to build multiple VM to fulfill distribute computing. \n\nI would like to ask for advice, as a use that don't have powerful GPU equipped on my PC, what is the easiest way to run big dataset like this competition, and run program on them successfully?",
      "votes": null
    },
    {
      "id": "590519",
      "postDate": "08/02/2019 09:33:05",
      "content": "<p>You are aware that you can use Kaggle kernels directly on this platform?</p>",
      "rawMarkdown": "You are aware that you can use Kaggle kernels directly on this platform?",
      "votes": null
    },
    {
      "id": "590659",
      "postDate": "08/02/2019 12:58:27",
      "content": "<p>You can use free GPU compute through Kaggle kernels or Google Colab. Kaggle kernels is probably the most convenient for you since you can directly link datasets to the kernel instance. You don't have to worry about small cloud storage. 😉 </p>\n\n<p>Hope this helps!</p>\n\n<p>Intro to Kaggle kernels:\n<a href=\"https://towardsdatascience.com/introduction-to-kaggle-kernels-2ad754ebf77\">https://towardsdatascience.com/introduction-to-kaggle-kernels-2ad754ebf77</a></p>\n\n<p>Google Colab:\n<a href=\"https://colab.research.google.com\">https://colab.research.google.com</a></p>",
      "rawMarkdown": "You can use free GPU compute through Kaggle kernels or Google Colab. Kaggle kernels is probably the most convenient for you since you can directly link datasets to the kernel instance. You don't have to worry about small cloud storage. 😉 \n\nHope this helps!\n\nIntro to Kaggle kernels:\nhttps://towardsdatascience.com/introduction-to-kaggle-kernels-2ad754ebf77\n\nGoogle Colab:\nhttps://colab.research.google.com",
      "votes": null
    },
    {
      "id": "591284",
      "postDate": "08/03/2019 12:53:13",
      "content": "<p>Thanks, I'll try it!</p>",
      "rawMarkdown": "Thanks, I'll try it!",
      "votes": null
    },
    {
      "id": "591286",
      "postDate": "08/03/2019 12:54:43",
      "content": "<p>Actually I forgot, I'll try later!</p>",
      "rawMarkdown": "Actually I forgot, I'll try later!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 590519,
      "author_name": "philippsinger",
      "author_url": "",
      "post_date": "08/02/2019 09:33:05",
      "content": "<p>You are aware that you can use Kaggle kernels directly on this platform?</p>",
      "votes": null,
      "replies": [
        {
          "id": 591286,
          "author_name": "laurencelin",
          "author_url": "",
          "post_date": "08/03/2019 12:54:43",
          "content": "<p>Actually I forgot, I'll try later!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 590659,
      "author_name": "carlolepelaars",
      "author_url": "",
      "post_date": "08/02/2019 12:58:27",
      "content": "<p>You can use free GPU compute through Kaggle kernels or Google Colab. Kaggle kernels is probably the most convenient for you since you can directly link datasets to the kernel instance. You don't have to worry about small cloud storage. 😉 </p>\n\n<p>Hope this helps!</p>\n\n<p>Intro to Kaggle kernels:\n<a href=\"https://towardsdatascience.com/introduction-to-kaggle-kernels-2ad754ebf77\">https://towardsdatascience.com/introduction-to-kaggle-kernels-2ad754ebf77</a></p>\n\n<p>Google Colab:\n<a href=\"https://colab.research.google.com\">https://colab.research.google.com</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 591284,
          "author_name": "laurencelin",
          "author_url": "",
          "post_date": "08/03/2019 12:53:13",
          "content": "<p>Thanks, I'll try it!</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "590514": "I think this is a dumb question, but I would like to ask how does everyone deal with large scale image data. I'm using a laptop that only contains CPU, so I use an cloud GPU machine floydhub to help run deep learning models so far. However, the expense is expensive and the cloud storage is small, so I'm in trouble of finding a way to store large scale dataset and run the program on it. \n\nI've found two methods that may help: one is to use free trial usage like google cloud, it seems to provide limit usage of VM with GPU; and another is to use Spark, that also need to build multiple VM to fulfill distribute computing. \n\nI would like to ask for advice, as a use that don't have powerful GPU equipped on my PC, what is the easiest way to run big dataset like this competition, and run program on them successfully?",
    "590519": "You are aware that you can use Kaggle kernels directly on this platform?",
    "590659": "You can use free GPU compute through Kaggle kernels or Google Colab. Kaggle kernels is probably the most convenient for you since you can directly link datasets to the kernel instance. You don't have to worry about small cloud storage. 😉 \n\nHope this helps!\n\nIntro to Kaggle kernels:\nhttps://towardsdatascience.com/introduction-to-kaggle-kernels-2ad754ebf77\n\nGoogle Colab:\nhttps://colab.research.google.com",
    "591284": "Thanks, I'll try it!",
    "591286": "Actually I forgot, I'll try later!"
  },
  "source": "meta"
}