{
  "id": 60543,
  "title": "Couldn't Download Datasets ",
  "url": "/competitions/google-ai-open-images-object-detection-track/discussion/60543",
  "author_name": "Xiang Zhang",
  "post_date": "2018-07-06T10:15:12.172000",
  "votes": 13,
  "comment_count": 26,
  "views": 0,
  "content": "<p>I tried to download the open images v4 datasets in Figure Eight and CVDF, both failed. Is there any other way or mirror to download? \nBy the way, I'm in China.</p>",
  "messages": [
    {
      "id": 353266,
      "postDate": "2018-07-06T10:15:12.173Z",
      "content": "<p>I tried to download the open images v4 datasets in Figure Eight and CVDF, both failed. Is there any other way or mirror to download? \nBy the way, I'm in China.</p>",
      "rawMarkdown": "I tried to download the open images v4 datasets in Figure Eight and CVDF, both failed. Is there any other way or mirror to download? \nBy the way, I'm in China.",
      "votes": 13
    },
    {
      "id": 353368,
      "postDate": "2018-07-06T14:45:16.180Z",
      "content": "<p>Same problem, impossible to download. Can someone share files with torrent, please? </p>",
      "rawMarkdown": "Same problem, impossible to download. Can someone share files with torrent, please? ",
      "votes": 5
    },
    {
      "id": 356271,
      "postDate": "2018-07-13T08:23:31.620Z",
      "content": "<p>The download is now available from AWS: please visit this page for the instructions:</p>\n\n<p><a href=\"https://github.com/cvdfoundation/open-images-dataset\">https://github.com/cvdfoundation/open-images-dataset</a></p>",
      "rawMarkdown": "The download is now available from AWS: please visit this page for the instructions:\n\n[https://github.com/cvdfoundation/open-images-dataset][1]\n\n  [1]: https://github.com/cvdfoundation/open-images-dataset",
      "votes": 1,
      "replies": [
        {
          "id": 356628,
          "postDate": "2018-07-14T04:46:55.570Z",
          "content": "<p>AWS works! Thank you!</p>",
          "rawMarkdown": "AWS works! Thank you!"
        },
        {
          "id": 356882,
          "postDate": "2018-07-14T17:42:23.260Z",
          "content": "<p>Nice job. I am downloading now. It's very fast.</p>",
          "rawMarkdown": "Nice job. I am downloading now. It's very fast."
        }
      ]
    },
    {
      "id": 355274,
      "postDate": "2018-07-11T10:58:24.443Z",
      "content": "<p>Hi, everyone,</p>\n\n<p>Thank you all for raising this issue. We are working hard to resolve the problem and the download should be again possible and easy within a couple of days. We apologize for the inconvenience. I will post updates in this thread. </p>",
      "rawMarkdown": "Hi, everyone,\n\nThank you all for raising this issue. We are working hard to resolve the problem and the download should be again possible and easy within a couple of days. We apologize for the inconvenience. I will post updates in this thread. "
    },
    {
      "id": 859056,
      "postDate": "2020-05-24T05:06:20.293Z",
      "content": "<p>can anyone suggest me source of PlantVillage Dataset. I can't download from kaggle \nlink: <a href=\"https://www.kaggle.com/emmarex/plantdisease/download\">https://www.kaggle.com/emmarex/plantdisease/download</a></p>",
      "rawMarkdown": "can anyone suggest me source of PlantVillage Dataset. I can't download from kaggle \nlink: https://www.kaggle.com/emmarex/plantdisease/download"
    },
    {
      "id": 358530,
      "postDate": "2018-07-18T10:32:09.253Z",
      "content": "<p>When I download the dataset from AWS to my local directory via awscli,  I always run into the \"fatal error: ('The read operation timed out',)\" or \"fatal error: ('Connection aborted.', error(110, 'Connection timed out')) \". I also tried \"--cli-read-timeout 0\" and \"--cli-connect-timeout 0\", then it just blocks. Did anyone run into the same problem? Could anyone help split the train set (513G file) into several partitions? Thanks a lot!</p>",
      "rawMarkdown": "When I download the dataset from AWS to my local directory via awscli,  I always run into the \"fatal error: ('The read operation timed out',)\" or \"fatal error: ('Connection aborted.', error(110, 'Connection timed out')) \". I also tried \"--cli-read-timeout 0\" and \"--cli-connect-timeout 0\", then it just blocks. Did anyone run into the same problem? Could anyone help split the train set (513G file) into several partitions? Thanks a lot!",
      "replies": [
        {
          "id": 360389,
          "postDate": "2018-07-22T11:20:17.080Z",
          "content": "<p>Hi Jinglei,\nCVDF is working on packing the training images into different files, I'll let you know when they are ready.\nSorry about the inconvenience, we are working hard to solve the issues.\nBest,</p>",
          "rawMarkdown": "Hi Jinglei,\nCVDF is working on packing the training images into different files, I'll let you know when they are ready.\nSorry about the inconvenience, we are working hard to solve the issues.\nBest,",
          "votes": 1
        },
        {
          "id": 365276,
          "postDate": "2018-08-02T08:42:23.007Z",
          "content": "<p>Hi Jinglei, </p>\n\n<p>CVDF made zipped files available to download, dividing the train set in 16 files, e.g.:\naws s3 --no-sign-request cp s3://open-images-dataset/tar/train_0.tar.gz [target_dir] (46G)\nLet's hope this makes downloads easier.</p>\n\n<p>Best,</p>",
          "rawMarkdown": "Hi Jinglei, \n\nCVDF made zipped files available to download, dividing the train set in 16 files, e.g.:\naws s3 --no-sign-request cp s3://open-images-dataset/tar/train_0.tar.gz [target_dir] (46G)\nLet's hope this makes downloads easier.\n\nBest,"
        },
        {
          "id": 365593,
          "postDate": "2018-08-03T01:21:15.300Z",
          "rawMarkdown": "",
          "isDeleted": true
        }
      ]
    },
    {
      "id": 358333,
      "postDate": "2018-07-18T00:38:00.887Z",
      "content": "<p>I've been trying to download the data via AWS from CVDF and it either doesn't grant access or it keeps failing.</p>\n\n<p>I'm trying to put it into my Google Cloud Bucket, so I first tried Google's Transfer Job tool, but it seems that the permissions on <code>S3://open-images-dataset</code> are not set to grant everyone (got this error: <code>Invalid access key. Make sure the access key for your S3 bucket is correct, or set the bucket permissions to Grant Everyone.</code>).  Next, I tried using <code>gsutil</code>, but similarly, I kept running into access errors.  Then I tried using <code>awscli</code> with the --no-sign-request which would start to work, but then start getting <code>[Errno 5] Input/output error</code>.  Is there something I'm missing here?  Or does anyone know the access credentials I should use?</p>\n\n<p>I know that the GitHub has links TSV files for the whole dataset, but I can't afford to store the whole 18TB dataset.  Does anyone know which of those 10 TSVs are just the competition dataset (fully annotated with bounding boxes, etc.), or does anyone know of another source to transfer the competition data?</p>\n\n<p>Thanks!</p>",
      "rawMarkdown": "I've been trying to download the data via AWS from CVDF and it either doesn't grant access or it keeps failing.\n\nI'm trying to put it into my Google Cloud Bucket, so I first tried Google's Transfer Job tool, but it seems that the permissions on `S3://open-images-dataset` are not set to grant everyone (got this error: `Invalid access key. Make sure the access key for your S3 bucket is correct, or set the bucket permissions to Grant Everyone.`).  Next, I tried using `gsutil`, but similarly, I kept running into access errors.  Then I tried using `awscli` with the --no-sign-request which would start to work, but then start getting `[Errno 5] Input/output error`.  Is there something I'm missing here?  Or does anyone know the access credentials I should use?\n\nI know that the GitHub has links TSV files for the whole dataset, but I can't afford to store the whole 18TB dataset.  Does anyone know which of those 10 TSVs are just the competition dataset (fully annotated with bounding boxes, etc.), or does anyone know of another source to transfer the competition data?\n\nThanks!",
      "replies": [
        {
          "id": 358335,
          "postDate": "2018-07-18T00:56:38.467Z",
          "content": "<p>I haven't tried it myself but tsv files are just glorified csv so you can actually cross reference the csv with the imageid field from the bbox csv file from the smaller dataset and produce csv that will only download the smaller set</p>",
          "rawMarkdown": "I haven't tried it myself but tsv files are just glorified csv so you can actually cross reference the csv with the imageid field from the bbox csv file from the smaller dataset and produce csv that will only download the smaller set"
        },
        {
          "id": 358449,
          "postDate": "2018-07-18T07:03:09.733Z",
          "content": "<p>Hi Nathaniel,</p>\n\n<ul>\n<li>Can you point me to the exact command you're using to transfer from AWS to Google Cloud? Maybe you're missing the --no-sign-request equivalent there?</li>\n<li>If downloading directly to your machine, have you used the command \"aws s3 --no-sign-request sync s3://open-images-dataset/train train\" ? This should work.</li>\n<li>You can find the CSV files with the exact Image IDs you need in the download section of the <a href=\"https://storage.googleapis.com/openimages/web/challenge.html\">Open Images website</a>. For training <a href=\"https://storage.googleapis.com/openimages/2018_04/train/train-images-boxable-with-rotation.csv\">this</a> would be the file, which has the links to the Flickr images, which, however, are not all available.</li>\n</ul>\n\n<p>Best,</p>",
          "rawMarkdown": "Hi Nathaniel,\n\n- Can you point me to the exact command you're using to transfer from AWS to Google Cloud? Maybe you're missing the --no-sign-request equivalent there?\n- If downloading directly to your machine, have you used the command \"aws s3 --no-sign-request sync s3://open-images-dataset/train train\" ? This should work.\n- You can find the CSV files with the exact Image IDs you need in the download section of the [Open Images website][1]. For training [this][2] would be the file, which has the links to the Flickr images, which, however, are not all available.\n\nBest,\n\n  [1]: https://storage.googleapis.com/openimages/web/challenge.html\n  [2]: https://storage.googleapis.com/openimages/2018_04/train/train-images-boxable-with-rotation.csv"
        },
        {
          "id": 358769,
          "postDate": "2018-07-18T20:32:45.007Z",
          "content": "<p>Thanks for the reply Jordi,</p>\n\n<ul>\n<li>When trying the transfer via <a href=\"https://cloud.google.com/storage/docs/gsutil\">gsutil</a>, I used this command:\n<ul><li><code>gsutil cp -r s3://open-images-dataset/validation ~/data/validation</code></li>\n<li>Which returns me this error:</li>\n<li><code>ERROR 0718 13:13:00.083411 utils.py] Unable to read instance data, giving up\nFailure: No handler was ready to authenticate. 4 handlers were checked. ['HmacAuthV1Handler', 'DevshellAuth', 'OAuth2Auth', 'OAuth2ServiceAccountAuth'] Check your credentials.</code></li>\n<li>I had also trying adding <code>--no-sign-request</code> in there, but that just confused gsutil.  However, the documentation made it sound like any request without keys basically defaults to a no sign request.</li></ul></li>\n<li><p>Next, when using awscli, I used the same command that you mentioned:</p>\n\n<ul><li><code>aws s3 --no-sign-request sync s3://open-images-dataset/train train</code></li>\n<li>But after a while I started getting:</li>\n<li><code>[Errno 5] Input/output error</code></li></ul></li>\n<li><p>I'm not downloading directly to my machine as I don't have the harddrive space nor the extra bandwidth for it, plus I'd likely need to put it to sleep and move it before the download finishes.  I'm doing this from a VM within Google Cloud with my storage bucket mounted as the current working directory.</p></li>\n<li><p>Yeah, I guess I could download all the images from Flickr using the CSV, but I figured that Flickr would't appreciate so many people pulling such large amounts of data from them all at once.  And so I'd really rather get everything from a cloud storage solution that is set up for this kind of thing.</p></li>\n</ul>",
          "rawMarkdown": "Thanks for the reply Jordi,\n\n- When trying the transfer via [gsutil][1], I used this command:\n - `gsutil cp -r s3://open-images-dataset/validation ~/data/validation`\n - Which returns me this error:\n - `ERROR 0718 13:13:00.083411 utils.py] Unable to read instance data, giving up\nFailure: No handler was ready to authenticate. 4 handlers were checked. ['HmacAuthV1Handler', 'DevshellAuth', 'OAuth2Auth', 'OAuth2ServiceAccountAuth'] Check your credentials.`\n - I had also trying adding `--no-sign-request` in there, but that just confused gsutil.  However, the documentation made it sound like any request without keys basically defaults to a no sign request.\n- Next, when using awscli, I used the same command that you mentioned:\n - `aws s3 --no-sign-request sync s3://open-images-dataset/train train`\n - But after a while I started getting:\n - `[Errno 5] Input/output error`\n\n- I'm not downloading directly to my machine as I don't have the harddrive space nor the extra bandwidth for it, plus I'd likely need to put it to sleep and move it before the download finishes.  I'm doing this from a VM within Google Cloud with my storage bucket mounted as the current working directory.\n\n- Yeah, I guess I could download all the images from Flickr using the CSV, but I figured that Flickr would't appreciate so many people pulling such large amounts of data from them all at once.  And so I'd really rather get everything from a cloud storage solution that is set up for this kind of thing.\n\n  [1]: https://cloud.google.com/storage/docs/gsutil"
        },
        {
          "id": 358979,
          "postDate": "2018-07-19T09:20:06.403Z",
          "content": "<p>The only thing that comes to my mind is to try it from another computer and do:</p>\n\n<p><code>\ngsutil cp -r s3://open-images-dataset/validation gs://your-bucket\n</code></p>",
          "rawMarkdown": "The only thing that comes to my mind is to try it from another computer and do:\n\n```\ngsutil cp -r s3://open-images-dataset/validation gs://your-bucket\n```\n"
        },
        {
          "id": 359019,
          "postDate": "2018-07-19T10:29:31.917Z",
          "content": "<p>Nathaniel, maybe you just have some problems with access to your storage? Have you tried just copy or create file to your <code>train</code> storage? For example:\n<code>touch train/test.txt</code>\nor\n<code>echo \"Hello\" &gt; train/test.txt</code></p>",
          "rawMarkdown": "Nathaniel, maybe you just have some problems with access to your storage? Have you tried just copy or create file to your `train` storage? For example:\n`touch train/test.txt`\nor\n`echo \"Hello\" &gt; train/test.txt`\n\n"
        },
        {
          "id": 359285,
          "postDate": "2018-07-19T20:03:42.557Z",
          "content": "<p>Thanks Jordi, I tried it from another computer (not a VM) and also received the authentication error.  It was worth a shot though.</p>\n\n<p>Thank you Vadym, there don't seem to be any problems with accessing my storage as I'm able to do other transfers just fine.</p>",
          "rawMarkdown": "Thanks Jordi, I tried it from another computer (not a VM) and also received the authentication error.  It was worth a shot though.\n\nThank you Vadym, there don't seem to be any problems with accessing my storage as I'm able to do other transfers just fine."
        },
        {
          "id": 362721,
          "postDate": "2018-07-27T01:29:59.253Z",
          "rawMarkdown": "",
          "isDeleted": true
        }
      ]
    },
    {
      "id": 356898,
      "postDate": "2018-07-14T18:17:10.493Z",
      "content": "<p>Hey guys! I managed that we could mount AWS as partition (on Ubuntu 18.04).</p>\n\n<pre><code>sudo apt install s3fs\nmkdir -p ~/Downloads/google_io_dataset_mnt\ns3fs open-images-dataset ~/Downloads/google_io_dataset_mnt -o public_bucket=1,umask=0007,uid=1001\n</code></pre>\n\n<p>Here you are:\n<img src=\"https://i.imgur.com/3GMb8mY.png\" alt=\"enter image description here\"></p>",
      "rawMarkdown": "Hey guys! I managed that we could mount AWS as partition (on Ubuntu 18.04).\n\n    sudo apt install s3fs\n    mkdir -p ~/Downloads/google_io_dataset_mnt\n    s3fs open-images-dataset ~/Downloads/google_io_dataset_mnt -o public_bucket=1,umask=0007,uid=1001\n\nHere you are:\n![enter image description here][1]\n\n\n  [1]: https://i.imgur.com/3GMb8mY.png",
      "replies": [
        {
          "id": 358808,
          "postDate": "2018-07-19T00:28:38.103Z",
          "content": "<p>Thank you Vadym, this almost worked for me.  But I ended up getting <code>Input/Output Error</code>s after copying just 102 MB from the mounted partition.</p>\n\n<p>Did you use that mount just to access the partition as if it were one of your own drives (training the model directly from the mount and never downloading it)?  Or did you then copy the data somewhere else after that?   Either way, how's that been working for you so far?</p>\n\n<p>Also note, that I had to change <code>umask</code> to <code>0000</code> because I didn't have permission to copy the subfolders otherwise.</p>",
          "rawMarkdown": "Thank you Vadym, this almost worked for me.  But I ended up getting `Input/Output Error`s after copying just 102 MB from the mounted partition.\n\nDid you use that mount just to access the partition as if it were one of your own drives (training the model directly from the mount and never downloading it)?  Or did you then copy the data somewhere else after that?   Either way, how's that been working for you so far?\n\nAlso note, that I had to change `umask` to `0000` because I didn't have permission to copy the subfolders otherwise.\n\n\n"
        },
        {
          "id": 359016,
          "postDate": "2018-07-19T10:26:00.777Z",
          "content": "<p>I used it just to access w/o copying. I copyed dataset with <code>aws s3 --no-sign-request sync s3://open-images-dataset/train train</code></p>",
          "rawMarkdown": "I used it just to access w/o copying. I copyed dataset with `aws s3 --no-sign-request sync s3://open-images-dataset/train train`"
        }
      ]
    },
    {
      "id": 355218,
      "postDate": "2018-07-11T08:12:59.530Z",
      "content": "<p>I can download, but too slow.. need more then more one month to finish.. so sad.</p>",
      "rawMarkdown": "I can download, but too slow.. need more then more one month to finish.. so sad."
    },
    {
      "id": 355160,
      "postDate": "2018-07-11T05:19:29.927Z",
      "content": "<p>Same here. Please provide more mirror sites to download the dataset.</p>",
      "rawMarkdown": "Same here. Please provide more mirror sites to download the dataset."
    },
    {
      "id": 355033,
      "postDate": "2018-07-10T18:40:39.293Z",
      "content": "<p>Can the Organizers do something about this? Its been more than 2 days !! </p>",
      "rawMarkdown": "Can the Organizers do something about this? Its been more than 2 days !! "
    },
    {
      "id": 354359,
      "postDate": "2018-07-09T11:33:39.897Z",
      "content": "<p>I also cannot download dataset, T_T</p>",
      "rawMarkdown": "I also cannot download dataset, T_T"
    },
    {
      "id": 354264,
      "postDate": "2018-07-09T07:46:14.933Z",
      "content": "<p>Second that</p>",
      "rawMarkdown": "Second that\n"
    }
  ],
  "comments": [
    {
      "id": 353368,
      "author_name": "Azat Akhtyamov",
      "author_url": "",
      "post_date": "2018-07-06T14:45:16.180000",
      "content": "<p>Same problem, impossible to download. Can someone share files with torrent, please? </p>",
      "votes": 5,
      "replies": []
    },
    {
      "id": 356271,
      "author_name": "Alina",
      "author_url": "",
      "post_date": "2018-07-13T08:23:31.620000",
      "content": "<p>The download is now available from AWS: please visit this page for the instructions:</p>\n\n<p><a href=\"https://github.com/cvdfoundation/open-images-dataset\">https://github.com/cvdfoundation/open-images-dataset</a></p>",
      "votes": 1,
      "replies": [
        {
          "id": 356628,
          "author_name": "Xiang Zhang",
          "author_url": "",
          "post_date": "2018-07-14T04:46:55.570000",
          "content": "<p>AWS works! Thank you!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 356882,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-07-14T17:42:23.260000",
          "content": "<p>Nice job. I am downloading now. It's very fast.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 355274,
      "author_name": "Alina",
      "author_url": "",
      "post_date": "2018-07-11T10:58:24.443000",
      "content": "<p>Hi, everyone,</p>\n\n<p>Thank you all for raising this issue. We are working hard to resolve the problem and the download should be again possible and easy within a couple of days. We apologize for the inconvenience. I will post updates in this thread. </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 859056,
      "author_name": "shourov40",
      "author_url": "",
      "post_date": "2020-05-24T05:06:20.293000",
      "content": "<p>can anyone suggest me source of PlantVillage Dataset. I can't download from kaggle \nlink: <a href=\"https://www.kaggle.com/emmarex/plantdisease/download\">https://www.kaggle.com/emmarex/plantdisease/download</a></p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 358530,
      "author_name": "Jinglei",
      "author_url": "",
      "post_date": "2018-07-18T10:32:09.253000",
      "content": "<p>When I download the dataset from AWS to my local directory via awscli,  I always run into the \"fatal error: ('The read operation timed out',)\" or \"fatal error: ('Connection aborted.', error(110, 'Connection timed out')) \". I also tried \"--cli-read-timeout 0\" and \"--cli-connect-timeout 0\", then it just blocks. Did anyone run into the same problem? Could anyone help split the train set (513G file) into several partitions? Thanks a lot!</p>",
      "votes": 0,
      "replies": [
        {
          "id": 360389,
          "author_name": "Jordi Pont-Tuset",
          "author_url": "",
          "post_date": "2018-07-22T11:20:17.080000",
          "content": "<p>Hi Jinglei,\nCVDF is working on packing the training images into different files, I'll let you know when they are ready.\nSorry about the inconvenience, we are working hard to solve the issues.\nBest,</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 365276,
          "author_name": "Jordi Pont-Tuset",
          "author_url": "",
          "post_date": "2018-08-02T08:42:23.007000",
          "content": "<p>Hi Jinglei, </p>\n\n<p>CVDF made zipped files available to download, dividing the train set in 16 files, e.g.:\naws s3 --no-sign-request cp s3://open-images-dataset/tar/train_0.tar.gz [target_dir] (46G)\nLet's hope this makes downloads easier.</p>\n\n<p>Best,</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 365593,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-08-03T01:21:15.300000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 358333,
      "author_name": "Nathaniel Watkins",
      "author_url": "",
      "post_date": "2018-07-18T00:38:00.887000",
      "content": "<p>I've been trying to download the data via AWS from CVDF and it either doesn't grant access or it keeps failing.</p>\n\n<p>I'm trying to put it into my Google Cloud Bucket, so I first tried Google's Transfer Job tool, but it seems that the permissions on <code>S3://open-images-dataset</code> are not set to grant everyone (got this error: <code>Invalid access key. Make sure the access key for your S3 bucket is correct, or set the bucket permissions to Grant Everyone.</code>).  Next, I tried using <code>gsutil</code>, but similarly, I kept running into access errors.  Then I tried using <code>awscli</code> with the --no-sign-request which would start to work, but then start getting <code>[Errno 5] Input/output error</code>.  Is there something I'm missing here?  Or does anyone know the access credentials I should use?</p>\n\n<p>I know that the GitHub has links TSV files for the whole dataset, but I can't afford to store the whole 18TB dataset.  Does anyone know which of those 10 TSVs are just the competition dataset (fully annotated with bounding boxes, etc.), or does anyone know of another source to transfer the competition data?</p>\n\n<p>Thanks!</p>",
      "votes": 0,
      "replies": [
        {
          "id": 358335,
          "author_name": "Moshel",
          "author_url": "",
          "post_date": "2018-07-18T00:56:38.467000",
          "content": "<p>I haven't tried it myself but tsv files are just glorified csv so you can actually cross reference the csv with the imageid field from the bbox csv file from the smaller dataset and produce csv that will only download the smaller set</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 358449,
          "author_name": "Jordi",
          "author_url": "",
          "post_date": "2018-07-18T07:03:09.733000",
          "content": "<p>Hi Nathaniel,</p>\n\n<ul>\n<li>Can you point me to the exact command you're using to transfer from AWS to Google Cloud? Maybe you're missing the --no-sign-request equivalent there?</li>\n<li>If downloading directly to your machine, have you used the command \"aws s3 --no-sign-request sync s3://open-images-dataset/train train\" ? This should work.</li>\n<li>You can find the CSV files with the exact Image IDs you need in the download section of the <a href=\"https://storage.googleapis.com/openimages/web/challenge.html\">Open Images website</a>. For training <a href=\"https://storage.googleapis.com/openimages/2018_04/train/train-images-boxable-with-rotation.csv\">this</a> would be the file, which has the links to the Flickr images, which, however, are not all available.</li>\n</ul>\n\n<p>Best,</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 358769,
          "author_name": "Nathaniel Watkins",
          "author_url": "",
          "post_date": "2018-07-18T20:32:45.007000",
          "content": "<p>Thanks for the reply Jordi,</p>\n\n<ul>\n<li>When trying the transfer via <a href=\"https://cloud.google.com/storage/docs/gsutil\">gsutil</a>, I used this command:\n<ul><li><code>gsutil cp -r s3://open-images-dataset/validation ~/data/validation</code></li>\n<li>Which returns me this error:</li>\n<li><code>ERROR 0718 13:13:00.083411 utils.py] Unable to read instance data, giving up\nFailure: No handler was ready to authenticate. 4 handlers were checked. ['HmacAuthV1Handler', 'DevshellAuth', 'OAuth2Auth', 'OAuth2ServiceAccountAuth'] Check your credentials.</code></li>\n<li>I had also trying adding <code>--no-sign-request</code> in there, but that just confused gsutil.  However, the documentation made it sound like any request without keys basically defaults to a no sign request.</li></ul></li>\n<li><p>Next, when using awscli, I used the same command that you mentioned:</p>\n\n<ul><li><code>aws s3 --no-sign-request sync s3://open-images-dataset/train train</code></li>\n<li>But after a while I started getting:</li>\n<li><code>[Errno 5] Input/output error</code></li></ul></li>\n<li><p>I'm not downloading directly to my machine as I don't have the harddrive space nor the extra bandwidth for it, plus I'd likely need to put it to sleep and move it before the download finishes.  I'm doing this from a VM within Google Cloud with my storage bucket mounted as the current working directory.</p></li>\n<li><p>Yeah, I guess I could download all the images from Flickr using the CSV, but I figured that Flickr would't appreciate so many people pulling such large amounts of data from them all at once.  And so I'd really rather get everything from a cloud storage solution that is set up for this kind of thing.</p></li>\n</ul>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 358979,
          "author_name": "Jordi",
          "author_url": "",
          "post_date": "2018-07-19T09:20:06.403000",
          "content": "<p>The only thing that comes to my mind is to try it from another computer and do:</p>\n\n<p><code>\ngsutil cp -r s3://open-images-dataset/validation gs://your-bucket\n</code></p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 359019,
          "author_name": "Vadym Stupakov",
          "author_url": "",
          "post_date": "2018-07-19T10:29:31.917000",
          "content": "<p>Nathaniel, maybe you just have some problems with access to your storage? Have you tried just copy or create file to your <code>train</code> storage? For example:\n<code>touch train/test.txt</code>\nor\n<code>echo \"Hello\" &gt; train/test.txt</code></p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 359285,
          "author_name": "Nathaniel Watkins",
          "author_url": "",
          "post_date": "2018-07-19T20:03:42.557000",
          "content": "<p>Thanks Jordi, I tried it from another computer (not a VM) and also received the authentication error.  It was worth a shot though.</p>\n\n<p>Thank you Vadym, there don't seem to be any problems with accessing my storage as I'm able to do other transfers just fine.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 362721,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-07-27T01:29:59.253000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 356898,
      "author_name": "Vadym Stupakov",
      "author_url": "",
      "post_date": "2018-07-14T18:17:10.493000",
      "content": "<p>Hey guys! I managed that we could mount AWS as partition (on Ubuntu 18.04).</p>\n\n<pre><code>sudo apt install s3fs\nmkdir -p ~/Downloads/google_io_dataset_mnt\ns3fs open-images-dataset ~/Downloads/google_io_dataset_mnt -o public_bucket=1,umask=0007,uid=1001\n</code></pre>\n\n<p>Here you are:\n<img src=\"https://i.imgur.com/3GMb8mY.png\" alt=\"enter image description here\"></p>",
      "votes": 0,
      "replies": [
        {
          "id": 358808,
          "author_name": "Nathaniel Watkins",
          "author_url": "",
          "post_date": "2018-07-19T00:28:38.103000",
          "content": "<p>Thank you Vadym, this almost worked for me.  But I ended up getting <code>Input/Output Error</code>s after copying just 102 MB from the mounted partition.</p>\n\n<p>Did you use that mount just to access the partition as if it were one of your own drives (training the model directly from the mount and never downloading it)?  Or did you then copy the data somewhere else after that?   Either way, how's that been working for you so far?</p>\n\n<p>Also note, that I had to change <code>umask</code> to <code>0000</code> because I didn't have permission to copy the subfolders otherwise.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 359016,
          "author_name": "Vadym Stupakov",
          "author_url": "",
          "post_date": "2018-07-19T10:26:00.777000",
          "content": "<p>I used it just to access w/o copying. I copyed dataset with <code>aws s3 --no-sign-request sync s3://open-images-dataset/train train</code></p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 355218,
      "author_name": "gezi",
      "author_url": "",
      "post_date": "2018-07-11T08:12:59.530000",
      "content": "<p>I can download, but too slow.. need more then more one month to finish.. so sad.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 355160,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-07-11T05:19:29.927000",
      "content": "<p>Same here. Please provide more mirror sites to download the dataset.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 355033,
      "author_name": "DecentMakeover",
      "author_url": "",
      "post_date": "2018-07-10T18:40:39.293000",
      "content": "<p>Can the Organizers do something about this? Its been more than 2 days !! </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 354359,
      "author_name": "Numphy",
      "author_url": "",
      "post_date": "2018-07-09T11:33:39.897000",
      "content": "<p>I also cannot download dataset, T_T</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 354264,
      "author_name": "DecentMakeover",
      "author_url": "",
      "post_date": "2018-07-09T07:46:14.933000",
      "content": "<p>Second that</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "353266": "I tried to download the open images v4 datasets in Figure Eight and CVDF, both failed. Is there any other way or mirror to download? \nBy the way, I'm in China.",
    "353368": "Same problem, impossible to download. Can someone share files with torrent, please? ",
    "356271": "The download is now available from AWS: please visit this page for the instructions:\n\n[https://github.com/cvdfoundation/open-images-dataset][1]\n\n  [1]: https://github.com/cvdfoundation/open-images-dataset",
    "355274": "Hi, everyone,\n\nThank you all for raising this issue. We are working hard to resolve the problem and the download should be again possible and easy within a couple of days. We apologize for the inconvenience. I will post updates in this thread. ",
    "859056": "can anyone suggest me source of PlantVillage Dataset. I can't download from kaggle \nlink: https://www.kaggle.com/emmarex/plantdisease/download",
    "358530": "When I download the dataset from AWS to my local directory via awscli,  I always run into the \"fatal error: ('The read operation timed out',)\" or \"fatal error: ('Connection aborted.', error(110, 'Connection timed out')) \". I also tried \"--cli-read-timeout 0\" and \"--cli-connect-timeout 0\", then it just blocks. Did anyone run into the same problem? Could anyone help split the train set (513G file) into several partitions? Thanks a lot!",
    "358333": "I've been trying to download the data via AWS from CVDF and it either doesn't grant access or it keeps failing.\n\nI'm trying to put it into my Google Cloud Bucket, so I first tried Google's Transfer Job tool, but it seems that the permissions on `S3://open-images-dataset` are not set to grant everyone (got this error: `Invalid access key. Make sure the access key for your S3 bucket is correct, or set the bucket permissions to Grant Everyone.`).  Next, I tried using `gsutil`, but similarly, I kept running into access errors.  Then I tried using `awscli` with the --no-sign-request which would start to work, but then start getting `[Errno 5] Input/output error`.  Is there something I'm missing here?  Or does anyone know the access credentials I should use?\n\nI know that the GitHub has links TSV files for the whole dataset, but I can't afford to store the whole 18TB dataset.  Does anyone know which of those 10 TSVs are just the competition dataset (fully annotated with bounding boxes, etc.), or does anyone know of another source to transfer the competition data?\n\nThanks!",
    "356898": "Hey guys! I managed that we could mount AWS as partition (on Ubuntu 18.04).\n\n    sudo apt install s3fs\n    mkdir -p ~/Downloads/google_io_dataset_mnt\n    s3fs open-images-dataset ~/Downloads/google_io_dataset_mnt -o public_bucket=1,umask=0007,uid=1001\n\nHere you are:\n![enter image description here][1]\n\n\n  [1]: https://i.imgur.com/3GMb8mY.png",
    "355218": "I can download, but too slow.. need more then more one month to finish.. so sad.",
    "355160": "Same here. Please provide more mirror sites to download the dataset.",
    "355033": "Can the Organizers do something about this? Its been more than 2 days !! ",
    "354359": "I also cannot download dataset, T_T",
    "354264": "Second that\n"
  }
}