{
  "id": 99099,
  "title": "What kind of computation power are you using to train your model?",
  "url": "/competitions/recursion-cellular-image-classification/discussion/99099",
  "author_name": "",
  "post_date": "2019-07-08T21:08:17.821716100Z",
  "votes": 1,
  "comment_count": 3,
  "views": 0,
  "content": "<p>As a deep learning newcomer, I've been sticking to training in a Kaggle kernel. I find that performance plateaus right around the time the kernel times out, so my submissions are coming from ~9 hours on the included NVIDIA K80. Of course, since RAM/CPU is limited, I think a good chunk of that is lost to IO.</p>\n\n<p>How much training time are people using on their models? If I want to keep improving my model, should I expect to move to a set up with more power available?</p>",
  "messages": [
    {
      "id": "570861",
      "postDate": "07/08/2019 21:08:17",
      "content": "<p>As a deep learning newcomer, I've been sticking to training in a Kaggle kernel. I find that performance plateaus right around the time the kernel times out, so my submissions are coming from ~9 hours on the included NVIDIA K80. Of course, since RAM/CPU is limited, I think a good chunk of that is lost to IO.</p>\n\n<p>How much training time are people using on their models? If I want to keep improving my model, should I expect to move to a set up with more power available?</p>",
      "rawMarkdown": "As a deep learning newcomer, I've been sticking to training in a Kaggle kernel. I find that performance plateaus right around the time the kernel times out, so my submissions are coming from ~9 hours on the included NVIDIA K80. Of course, since RAM/CPU is limited, I think a good chunk of that is lost to IO.\n\nHow much training time are people using on their models? If I want to keep improving my model, should I expect to move to a set up with more power available?",
      "votes": null
    },
    {
      "id": "570913",
      "postDate": "07/08/2019 23:42:19",
      "content": "<p>Hey Max, from my experience the problem is I/O. I've recently seen a kernel that compressed the data into zip. I think it may reduce the training time if you try to run on Kaggle kernels. </p>",
      "rawMarkdown": "Hey Max, from my experience the problem is I/O. I've recently seen a kernel that compressed the data into zip. I think it may reduce the training time if you try to run on Kaggle kernels.",
      "votes": null
    },
    {
      "id": "570982",
      "postDate": "07/09/2019 02:14:03",
      "content": "<p>I think it's best to have your own local GPUs machine, so that you can run multiples experiments, no matter how dumb you think it is, just to not wasting free GPU time. \nIf the deadline is near or local machine is not feasible, use cloud service such as GCP to train. It's on  demand and you can spin up very expensive machine as well, but beware for the cost. It's also easier to setup tracking and saving models as well, be sure to apply for the GCP coupons.</p>\n\n<p>I only used Kaggle kernels so far in this competition, and will switch to GCP soon.</p>",
      "rawMarkdown": "I think it's best to have your own local GPUs machine, so that you can run multiples experiments, no matter how dumb you think it is, just to not wasting free GPU time. \nIf the deadline is near or local machine is not feasible, use cloud service such as GCP to train. It's on  demand and you can spin up very expensive machine as well, but beware for the cost. It's also easier to setup tracking and saving models as well, be sure to apply for the GCP coupons.\n\nI only used Kaggle kernels so far in this competition, and will switch to GCP soon.",
      "votes": null
    },
    {
      "id": "571549",
      "postDate": "07/09/2019 18:06:46",
      "content": "<p>Just one correction, kernels are using P100 instead of K80 now. </p>",
      "rawMarkdown": "Just one correction, kernels are using P100 instead of K80 now.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 570913,
      "author_name": "gidutz",
      "author_url": "",
      "post_date": "07/08/2019 23:42:19",
      "content": "<p>Hey Max, from my experience the problem is I/O. I've recently seen a kernel that compressed the data into zip. I think it may reduce the training time if you try to run on Kaggle kernels. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 570982,
      "author_name": "lkhphuc",
      "author_url": "",
      "post_date": "07/09/2019 02:14:03",
      "content": "<p>I think it's best to have your own local GPUs machine, so that you can run multiples experiments, no matter how dumb you think it is, just to not wasting free GPU time. \nIf the deadline is near or local machine is not feasible, use cloud service such as GCP to train. It's on  demand and you can spin up very expensive machine as well, but beware for the cost. It's also easier to setup tracking and saving models as well, be sure to apply for the GCP coupons.</p>\n\n<p>I only used Kaggle kernels so far in this competition, and will switch to GCP soon.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 571549,
      "author_name": "naivelamb",
      "author_url": "",
      "post_date": "07/09/2019 18:06:46",
      "content": "<p>Just one correction, kernels are using P100 instead of K80 now. </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "570861": "As a deep learning newcomer, I've been sticking to training in a Kaggle kernel. I find that performance plateaus right around the time the kernel times out, so my submissions are coming from ~9 hours on the included NVIDIA K80. Of course, since RAM/CPU is limited, I think a good chunk of that is lost to IO.\n\nHow much training time are people using on their models? If I want to keep improving my model, should I expect to move to a set up with more power available?",
    "570913": "Hey Max, from my experience the problem is I/O. I've recently seen a kernel that compressed the data into zip. I think it may reduce the training time if you try to run on Kaggle kernels.",
    "570982": "I think it's best to have your own local GPUs machine, so that you can run multiples experiments, no matter how dumb you think it is, just to not wasting free GPU time. \nIf the deadline is near or local machine is not feasible, use cloud service such as GCP to train. It's on  demand and you can spin up very expensive machine as well, but beware for the cost. It's also easier to setup tracking and saving models as well, be sure to apply for the GCP coupons.\n\nI only used Kaggle kernels so far in this competition, and will switch to GCP soon.",
    "571549": "Just one correction, kernels are using P100 instead of K80 now."
  },
  "source": "meta"
}