{
  "id": 60529,
  "title": "Reading images on kaggle kernel",
  "url": "/competitions/google-ai-open-images-object-detection-track/discussion/60529",
  "author_name": "",
  "post_date": "2018-07-06T03:14:57.294219100Z",
  "votes": 1,
  "comment_count": 3,
  "views": 0,
  "content": "<p>How to read the images on Kaggle Kernel directly ?</p>",
  "messages": [
    {
      "id": "353161",
      "postDate": "07/06/2018 03:14:57",
      "content": "<p>How to read the images on Kaggle Kernel directly ?</p>",
      "rawMarkdown": "How to read the images on Kaggle Kernel directly ?",
      "votes": null
    },
    {
      "id": "354115",
      "postDate": "07/08/2018 19:17:12",
      "content": "<p>You can access the test images very easily. I would use this:</p>\n\n<pre><code>from skimage.io import imread\nnumpy_array = imread(\"PATH TO TEST IMAGE(ACT LIKE ZIP IS DIRECTORY)\")\n</code></pre>\n\n<p>The training images cannot be used altogether in a kernel. There are just too many. You can train on quite a lot of them by creating a dataset of them &lt;50 GB. I would recommend using google cloud compute engine for training. The training images are located at URLs, so I just use imread and point it to the URL. This is slower, but does save space.</p>",
      "rawMarkdown": "You can access the test images very easily. I would use this:\n\n    from skimage.io import imread\n    numpy_array = imread(\"PATH TO TEST IMAGE(ACT LIKE ZIP IS DIRECTORY)\")\n\nThe training images cannot be used altogether in a kernel. There are just too many. You can train on quite a lot of them by creating a dataset of them &lt;50 GB. I would recommend using google cloud compute engine for training. The training images are located at URLs, so I just use imread and point it to the URL. This is slower, but does save space.",
      "votes": null
    },
    {
      "id": "355721",
      "postDate": "07/12/2018 08:50:23",
      "content": "<p>Arpan would you mind sharing a sample of code on how to use this trick?</p>",
      "rawMarkdown": "Arpan would you mind sharing a sample of code on how to use this trick?",
      "votes": null
    },
    {
      "id": "356035",
      "postDate": "07/12/2018 19:13:39",
      "content": "<p>The imread function is pretty straightforward. It needs to be pointed to either an Image path or URL. If you were to point it towards \n<code>\"../input/google-ai-open-images-object-detection-track/test/imageID.jpg\"</code>\nit will read the image from the folder. However, you cannot collect data from URLs in a kaggle kernel. They do not have internet access, so I used a google cloud compute engine. What I did first was collect the URLs in <a href=\"https://www.figure-eight.com/dataset/open-images-annotated-with-bounding-boxes/\">a CSV located at figure-eight</a>. The CSV had the ID of each image as the first column and the URL in the second column. The URLs were all pointed at figure eight so I used time.sleep to slow down the request rate. Eventually I realized it was way to slow so I just downloaded all the images to a google cloud storage bucket and now I use that instead since transferring information from a cloud bucket to a VM is free provided they are in the same region. It also has an unlimited request rate. I use pytorch so I used multiple workers to read the data very quickly.</p>",
      "rawMarkdown": "The imread function is pretty straightforward. It needs to be pointed to either an Image path or URL. If you were to point it towards \n`\"../input/google-ai-open-images-object-detection-track/test/imageID.jpg\"`\nit will read the image from the folder. However, you cannot collect data from URLs in a kaggle kernel. They do not have internet access, so I used a google cloud compute engine. What I did first was collect the URLs in [a CSV located at figure-eight][1]. The CSV had the ID of each image as the first column and the URL in the second column. The URLs were all pointed at figure eight so I used time.sleep to slow down the request rate. Eventually I realized it was way to slow so I just downloaded all the images to a google cloud storage bucket and now I use that instead since transferring information from a cloud bucket to a VM is free provided they are in the same region. It also has an unlimited request rate. I use pytorch so I used multiple workers to read the data very quickly.\n\n\n  [1]: https://www.figure-eight.com/dataset/open-images-annotated-with-bounding-boxes/",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 354115,
      "author_name": "arpandhatt",
      "author_url": "",
      "post_date": "07/08/2018 19:17:12",
      "content": "<p>You can access the test images very easily. I would use this:</p>\n\n<pre><code>from skimage.io import imread\nnumpy_array = imread(\"PATH TO TEST IMAGE(ACT LIKE ZIP IS DIRECTORY)\")\n</code></pre>\n\n<p>The training images cannot be used altogether in a kernel. There are just too many. You can train on quite a lot of them by creating a dataset of them &lt;50 GB. I would recommend using google cloud compute engine for training. The training images are located at URLs, so I just use imread and point it to the URL. This is slower, but does save space.</p>",
      "votes": null,
      "replies": [
        {
          "id": 355721,
          "author_name": "adamsfei",
          "author_url": "",
          "post_date": "07/12/2018 08:50:23",
          "content": "<p>Arpan would you mind sharing a sample of code on how to use this trick?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 356035,
          "author_name": "arpandhatt",
          "author_url": "",
          "post_date": "07/12/2018 19:13:39",
          "content": "<p>The imread function is pretty straightforward. It needs to be pointed to either an Image path or URL. If you were to point it towards \n<code>\"../input/google-ai-open-images-object-detection-track/test/imageID.jpg\"</code>\nit will read the image from the folder. However, you cannot collect data from URLs in a kaggle kernel. They do not have internet access, so I used a google cloud compute engine. What I did first was collect the URLs in <a href=\"https://www.figure-eight.com/dataset/open-images-annotated-with-bounding-boxes/\">a CSV located at figure-eight</a>. The CSV had the ID of each image as the first column and the URL in the second column. The URLs were all pointed at figure eight so I used time.sleep to slow down the request rate. Eventually I realized it was way to slow so I just downloaded all the images to a google cloud storage bucket and now I use that instead since transferring information from a cloud bucket to a VM is free provided they are in the same region. It also has an unlimited request rate. I use pytorch so I used multiple workers to read the data very quickly.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "353161": "How to read the images on Kaggle Kernel directly ?",
    "354115": "You can access the test images very easily. I would use this:\n\n    from skimage.io import imread\n    numpy_array = imread(\"PATH TO TEST IMAGE(ACT LIKE ZIP IS DIRECTORY)\")\n\nThe training images cannot be used altogether in a kernel. There are just too many. You can train on quite a lot of them by creating a dataset of them &lt;50 GB. I would recommend using google cloud compute engine for training. The training images are located at URLs, so I just use imread and point it to the URL. This is slower, but does save space.",
    "355721": "Arpan would you mind sharing a sample of code on how to use this trick?",
    "356035": "The imread function is pretty straightforward. It needs to be pointed to either an Image path or URL. If you were to point it towards \n`\"../input/google-ai-open-images-object-detection-track/test/imageID.jpg\"`\nit will read the image from the folder. However, you cannot collect data from URLs in a kaggle kernel. They do not have internet access, so I used a google cloud compute engine. What I did first was collect the URLs in [a CSV located at figure-eight][1]. The CSV had the ID of each image as the first column and the URL in the second column. The URLs were all pointed at figure eight so I used time.sleep to slow down the request rate. Eventually I realized it was way to slow so I just downloaded all the images to a google cloud storage bucket and now I use that instead since transferring information from a cloud bucket to a VM is free provided they are in the same region. It also has an unlimited request rate. I use pytorch so I used multiple workers to read the data very quickly.\n\n\n  [1]: https://www.figure-eight.com/dataset/open-images-annotated-with-bounding-boxes/"
  },
  "source": "meta"
}