{
  "id": 171476,
  "title": "Kaggle Infrastructure",
  "url": "/competitions/landmark-retrieval-2020/discussion/171476",
  "author_name": "Nawid Sayed",
  "post_date": "2020-07-31T23:38:06.094000",
  "votes": 0,
  "comment_count": 2,
  "views": 0,
  "content": "<p>Hello,</p>\n\n<p>this is my first CV competition on kaggle and after playing around with <a href=\"https://www.kaggle.com/sandy1112/create-and-train-resnet50-from-scratch\">this notebook</a> I noticed that data loading is a huge bottleneck for model training. </p>\n\n<p>I am not particularly surprised given that GPU kernels only have 2 CPU cores available.</p>\n\n<p>Is there an obvious solution to this problem that I missed? The notebook uses Keras DataGenerators with multiple workers. Can you perform data augmentation on the GPU?</p>\n\n<p>I am also curious about whether top teams use Kaggle infrastructure or their own machines? \nIs it pointless to seriously compete in this competition solely with Kaggle infrastructure? </p>\n\n<p>Thanks in advance!</p>",
  "messages": [
    {
      "id": 953756,
      "postDate": "2020-08-01T04:33:06.893Z",
      "content": "<p>Yes, the bottleneck is I/O. I suggest you use TFRecord. Here is a notebook about this. <a href=\"https://www.kaggle.com/wuliaokaola/alaska2-some-tips-for-using-tpu\">https://www.kaggle.com/wuliaokaola/alaska2-some-tips-for-using-tpu</a></p>",
      "rawMarkdown": "Yes, the bottleneck is I/O. I suggest you use TFRecord. Here is a notebook about this. https://www.kaggle.com/wuliaokaola/alaska2-some-tips-for-using-tpu"
    },
    {
      "id": 953626,
      "postDate": "2020-07-31T23:38:06.093Z",
      "content": "<p>Hello,</p>\n\n<p>this is my first CV competition on kaggle and after playing around with <a href=\"https://www.kaggle.com/sandy1112/create-and-train-resnet50-from-scratch\">this notebook</a> I noticed that data loading is a huge bottleneck for model training. </p>\n\n<p>I am not particularly surprised given that GPU kernels only have 2 CPU cores available.</p>\n\n<p>Is there an obvious solution to this problem that I missed? The notebook uses Keras DataGenerators with multiple workers. Can you perform data augmentation on the GPU?</p>\n\n<p>I am also curious about whether top teams use Kaggle infrastructure or their own machines? \nIs it pointless to seriously compete in this competition solely with Kaggle infrastructure? </p>\n\n<p>Thanks in advance!</p>",
      "rawMarkdown": "Hello,\n\nthis is my first CV competition on kaggle and after playing around with [this notebook](https://www.kaggle.com/sandy1112/create-and-train-resnet50-from-scratch) I noticed that data loading is a huge bottleneck for model training. \n\nI am not particularly surprised given that GPU kernels only have 2 CPU cores available.\n\nIs there an obvious solution to this problem that I missed? The notebook uses Keras DataGenerators with multiple workers. Can you perform data augmentation on the GPU?\n\nI am also curious about whether top teams use Kaggle infrastructure or their own machines? \nIs it pointless to seriously compete in this competition solely with Kaggle infrastructure? \n\nThanks in advance!"
    },
    {
      "id": 954326,
      "postDate": "2020-08-01T16:25:46.127Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 953756,
      "author_name": "Johnny Lee",
      "author_url": "",
      "post_date": "2020-08-01T04:33:06.893000",
      "content": "<p>Yes, the bottleneck is I/O. I suggest you use TFRecord. Here is a notebook about this. <a href=\"https://www.kaggle.com/wuliaokaola/alaska2-some-tips-for-using-tpu\">https://www.kaggle.com/wuliaokaola/alaska2-some-tips-for-using-tpu</a></p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 954326,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-08-01T16:25:46.127000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "953756": "Yes, the bottleneck is I/O. I suggest you use TFRecord. Here is a notebook about this. https://www.kaggle.com/wuliaokaola/alaska2-some-tips-for-using-tpu",
    "953626": "Hello,\n\nthis is my first CV competition on kaggle and after playing around with [this notebook](https://www.kaggle.com/sandy1112/create-and-train-resnet50-from-scratch) I noticed that data loading is a huge bottleneck for model training. \n\nI am not particularly surprised given that GPU kernels only have 2 CPU cores available.\n\nIs there an obvious solution to this problem that I missed? The notebook uses Keras DataGenerators with multiple workers. Can you perform data augmentation on the GPU?\n\nI am also curious about whether top teams use Kaggle infrastructure or their own machines? \nIs it pointless to seriously compete in this competition solely with Kaggle infrastructure? \n\nThanks in advance!",
    "954326": ""
  }
}