{
  "id": 61537,
  "title": "About aws downloading images",
  "url": "/competitions/google-ai-open-images-object-detection-track/discussion/61537",
  "author_name": "peterlau",
  "post_date": "2018-07-20T16:01:24.674000",
  "votes": 0,
  "comment_count": 14,
  "views": 0,
  "content": "<p>Hello,kagglers:\nI am currently downloading training images through  aws s3 as stated before.However,when I have downloaded about  204G training data,I meet this problem:\nfatal error:('The read operation timed out').\nSo T manually tried the command 's3 --no-sign-request sync s3://open-images-dataset/train' again and again ,sometimes it works , but it fails most of the time.How can I solve this issue and  download the whole training dataset  sucecessfully?Thanks! </p>",
  "messages": [
    {
      "id": 365275,
      "postDate": "2018-08-02T08:41:51.097Z",
      "content": "<p>Hi all, </p>\n\n<p>CVDF made zipped files available to download, dividing the train set in 16 files, e.g.:\naws s3 --no-sign-request cp s3://open-images-dataset/tar/train_0.tar.gz [target_dir] (46G)\nLet's hope this makes downloads easier.</p>\n\n<p>Best,</p>",
      "rawMarkdown": "Hi all, \n\nCVDF made zipped files available to download, dividing the train set in 16 files, e.g.:\naws s3 --no-sign-request cp s3://open-images-dataset/tar/train_0.tar.gz [target_dir] (46G)\nLet's hope this makes downloads easier.\n\nBest,",
      "votes": 1,
      "replies": [
        {
          "id": 365621,
          "postDate": "2018-08-03T03:40:45.323Z",
          "content": "<p>Hi Jordi,</p>\n\n<p>I have been transferring the training images to a bucket in google cloud, it took more than 12 hours to transfer 1700 images to my bucket. At this rate I wont be able to even finish transferring in a month. I am not sure if I am doing something wrong that the transfer speed is this slow. I was wondering if there is any way to speed up the transfer, been looking at the documents online but have trouble finding something that works.</p>\n\n<p>I have set up the bucket in us-central1 btw</p>",
          "rawMarkdown": "Hi Jordi,\n\nI have been transferring the training images to a bucket in google cloud, it took more than 12 hours to transfer 1700 images to my bucket. At this rate I wont be able to even finish transferring in a month. I am not sure if I am doing something wrong that the transfer speed is this slow. I was wondering if there is any way to speed up the transfer, been looking at the documents online but have trouble finding something that works.\n\nI have set up the bucket in us-central1 btw"
        },
        {
          "id": 365679,
          "postDate": "2018-08-03T07:15:18.090Z",
          "content": "<p>Can you try downloading the ZIP files instead?\nThere is not much more we can do from our side...</p>",
          "rawMarkdown": "Can you try downloading the ZIP files instead?\nThere is not much more we can do from our side..."
        },
        {
          "id": 365967,
          "postDate": "2018-08-03T18:56:57.513Z",
          "content": "<p>Thats alright Jordi, figured something else, will see if it works.</p>\n\n<p>Cheers</p>",
          "rawMarkdown": "Thats alright Jordi, figured something else, will see if it works.\n\nCheers"
        }
      ]
    },
    {
      "id": 394891,
      "postDate": "2018-09-27T15:56:29.810Z",
      "content": "<p>Has anyone been able to get the entire training set (513gb) into a google cloud bucket?  I have tried the google transfer function (permissions didn't work), using awscli with a google instance and the subsets (failed due to 'device is full', which isn't true as my instance has 50gb of memory), and downloading directly to my computer then uploading (prohibitivly slow).</p>",
      "rawMarkdown": "Has anyone been able to get the entire training set (513gb) into a google cloud bucket?  I have tried the google transfer function (permissions didn't work), using awscli with a google instance and the subsets (failed due to 'device is full', which isn't true as my instance has 50gb of memory), and downloading directly to my computer then uploading (prohibitivly slow)."
    },
    {
      "id": 386637,
      "postDate": "2018-09-13T09:25:50.147Z",
      "content": "<p>The S3 buckets for images with bounding box are - </p>\n\n<p>s3://open-images-dataset/train (513GB)\ns3://open-images-dataset/validation (12GB)\ns3://open-images-dataset/test (36GB)</p>\n\n<p>One way to get them to GCS bucket is by creating GCS Transfer Job function. All my attempt to transfer from S3 bucket to GCS has failed as with error - \"Invalid access key\". I have tried to use my own AWS IAM credentials but they don't work. </p>\n\n<p>How can the data be transferred directly from S3 to GCS bucket directly??</p>",
      "rawMarkdown": "The S3 buckets for images with bounding box are - \n\ns3://open-images-dataset/train (513GB)\ns3://open-images-dataset/validation (12GB)\ns3://open-images-dataset/test (36GB)\n\nOne way to get them to GCS bucket is by creating GCS Transfer Job function. All my attempt to transfer from S3 bucket to GCS has failed as with error - \"Invalid access key\". I have tried to use my own AWS IAM credentials but they don't work. \n\nHow can the data be transferred directly from S3 to GCS bucket directly??"
    },
    {
      "id": 360744,
      "postDate": "2018-07-23T06:30:31.220Z",
      "content": "<p>Thanks for the link. Your link states, \"Google Storage provides a \"storage transfer\" function to transfer online files into a storage bucket ... The size of the whole dataset is around 18TB. Please note that user needs to pay for hosting the dataset on Google Cloud storage...\" My data size estimate of 16TB is from this site: <a href=\"https://blog.algorithmia.com/deep-dive-into-object-detection-with-open-images-using-tensorflow/\">https://blog.algorithmia.com/deep-dive-into-object-detection-with-open-images-using-tensorflow/</a>. I was looking for an option to access the dataset free of cost, an option that also allows reading the data by pointing to a single URL. Any help is greatly appreciated. </p>",
      "rawMarkdown": "Thanks for the link. Your link states, \"Google Storage provides a \"storage transfer\" function to transfer online files into a storage bucket ... The size of the whole dataset is around 18TB. Please note that user needs to pay for hosting the dataset on Google Cloud storage...\" My data size estimate of 16TB is from this site: https://blog.algorithmia.com/deep-dive-into-object-detection-with-open-images-using-tensorflow/. I was looking for an option to access the dataset free of cost, an option that also allows reading the data by pointing to a single URL. Any help is greatly appreciated. ",
      "replies": [
        {
          "id": 360778,
          "postDate": "2018-07-23T08:06:44.453Z",
          "content": "<p>This is for the full database. Its easy to make this mistake. Google has 2 kinds of database. One has heen bboxed by humans and we should use it for the training. It has 1.7m images iirc and is around 530gb (tiny!). The other, which is the full database, has been annotated by machines and has 19m images iirc.\nThe aws link will download the database you want. The Google one will load the huge one.\nI haven't tried it, but I think you can filter the tsv using the annotations and download only the training set we need. It would be ultra super nice if Google could actually put the database in tfrecord format somewhere as even loading this much data into tfrecord would be a very lengthy process, to say nothing of twice the amount of local storage needed. </p>",
          "rawMarkdown": "This is for the full database. Its easy to make this mistake. Google has 2 kinds of database. One has heen bboxed by humans and we should use it for the training. It has 1.7m images iirc and is around 530gb (tiny!). The other, which is the full database, has been annotated by machines and has 19m images iirc.\nThe aws link will download the database you want. The Google one will load the huge one.\nI haven't tried it, but I think you can filter the tsv using the annotations and download only the training set we need. It would be ultra super nice if Google could actually put the database in tfrecord format somewhere as even loading this much data into tfrecord would be a very lengthy process, to say nothing of twice the amount of local storage needed. ",
          "votes": 1
        },
        {
          "id": 360789,
          "postDate": "2018-07-23T08:37:43.783Z",
          "content": "<p>Dear Avantis,\nOpen Images has indeed different subsets:</p>\n\n<ol>\n<li>Subset annotated with bounding boxes (1,743,042 training images)</li>\n<li>Subset with human-verified image-level labels (5,655,108 training images)</li>\n<li>Subset with machine-generated image-level labels (8,853,429 training images)</li>\n<li>Full dataset (9,178,275 images)</li>\n</ol>\n\n<p>The images of subset 1. take up approximately 500GB and can be downloaded from AWS via CVDF or from Open Images. The full dataset takes up approximately 19TB and right now can only be downloaded directly from Flickr using the links provided by CVDF it TSV files.</p>\n\n<p>All this information (and much more) in the Open Images Website.</p>",
          "rawMarkdown": "Dear Avantis,\nOpen Images has indeed different subsets:\n\n1. Subset annotated with bounding boxes (1,743,042 training images)\n2. Subset with human-verified image-level labels (5,655,108 training images)\n3. Subset with machine-generated image-level labels (8,853,429 training images)\n4. Full dataset (9,178,275 images)\n\nThe images of subset 1. take up approximately 500GB and can be downloaded from AWS via CVDF or from Open Images. The full dataset takes up approximately 19TB and right now can only be downloaded directly from Flickr using the links provided by CVDF it TSV files.\n\nAll this information (and much more) in the Open Images Website.",
          "votes": 1
        }
      ]
    },
    {
      "id": 360515,
      "postDate": "2018-07-22T16:22:36.777Z",
      "content": "<p>The complete dataset, downloaded and uncompressed, is around 6TB. Is that about right? Given the vast storage requirement, is the entire dataset accessible via Google Storage?</p>",
      "rawMarkdown": "The complete dataset, downloaded and uncompressed, is around 6TB. Is that about right? Given the vast storage requirement, is the entire dataset accessible via Google Storage?",
      "replies": [
        {
          "id": 360565,
          "postDate": "2018-07-22T17:54:10.560Z",
          "content": "<p>The full subset of Open Images where there are bounding boxes is:\n- Train: <strong>513GB</strong>\n- Test-Challenge 2018: <strong>10GB</strong></p>\n\n<p>More info in the CVDF download site:\n<a href=\"https://github.com/cvdfoundation/open-images-dataset#download-images-with-bounding-boxes-annotations\">https://github.com/cvdfoundation/open-images-dataset#download-images-with-bounding-boxes-annotations</a></p>\n\n<p>Not sure where these 6TB come from.</p>",
          "rawMarkdown": "The full subset of Open Images where there are bounding boxes is:\n- Train: **513GB**\n- Test-Challenge 2018: **10GB**\n\nMore info in the CVDF download site:\nhttps://github.com/cvdfoundation/open-images-dataset#download-images-with-bounding-boxes-annotations\n\nNot sure where these 6TB come from.\n"
        },
        {
          "id": 361490,
          "postDate": "2018-07-24T15:45:03.123Z",
          "rawMarkdown": "",
          "isDeleted": true
        }
      ]
    },
    {
      "id": 359698,
      "postDate": "2018-07-20T16:03:11.217Z",
      "content": "<p>I would appreciate someone who could share the downloaded training dataset on the Internet.Downloading this huge training data is very time-consuming .</p>",
      "rawMarkdown": "I would appreciate someone who could share the downloaded training dataset on the Internet.Downloading this huge training data is very time-consuming .",
      "replies": [
        {
          "id": 360386,
          "postDate": "2018-07-22T11:11:11.547Z",
          "content": "<p>Dear peterlau,\nCVDF is working on packing the images into ZIP files, maybe that will make downloads easier. \nSorry for the inconvenience.</p>",
          "rawMarkdown": "Dear peterlau,\nCVDF is working on packing the images into ZIP files, maybe that will make downloads easier. \nSorry for the inconvenience."
        }
      ]
    },
    {
      "id": 359696,
      "postDate": "2018-07-20T16:01:24.673Z",
      "content": "<p>Hello,kagglers:\nI am currently downloading training images through  aws s3 as stated before.However,when I have downloaded about  204G training data,I meet this problem:\nfatal error:('The read operation timed out').\nSo T manually tried the command 's3 --no-sign-request sync s3://open-images-dataset/train' again and again ,sometimes it works , but it fails most of the time.How can I solve this issue and  download the whole training dataset  sucecessfully?Thanks! </p>",
      "rawMarkdown": "Hello,kagglers:\nI am currently downloading training images through  aws s3 as stated before.However,when I have downloaded about  204G training data,I meet this problem:\nfatal error:('The read operation timed out').\nSo T manually tried the command 's3 --no-sign-request sync s3://open-images-dataset/train' again and again ,sometimes it works , but it fails most of the time.How can I solve this issue and  download the whole training dataset  sucecessfully?Thanks! "
    }
  ],
  "comments": [
    {
      "id": 365275,
      "author_name": "Jordi Pont-Tuset",
      "author_url": "",
      "post_date": "2018-08-02T08:41:51.097000",
      "content": "<p>Hi all, </p>\n\n<p>CVDF made zipped files available to download, dividing the train set in 16 files, e.g.:\naws s3 --no-sign-request cp s3://open-images-dataset/tar/train_0.tar.gz [target_dir] (46G)\nLet's hope this makes downloads easier.</p>\n\n<p>Best,</p>",
      "votes": 1,
      "replies": [
        {
          "id": 365621,
          "author_name": "Mukesh",
          "author_url": "",
          "post_date": "2018-08-03T03:40:45.323000",
          "content": "<p>Hi Jordi,</p>\n\n<p>I have been transferring the training images to a bucket in google cloud, it took more than 12 hours to transfer 1700 images to my bucket. At this rate I wont be able to even finish transferring in a month. I am not sure if I am doing something wrong that the transfer speed is this slow. I was wondering if there is any way to speed up the transfer, been looking at the documents online but have trouble finding something that works.</p>\n\n<p>I have set up the bucket in us-central1 btw</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 365679,
          "author_name": "Jordi Pont-Tuset",
          "author_url": "",
          "post_date": "2018-08-03T07:15:18.090000",
          "content": "<p>Can you try downloading the ZIP files instead?\nThere is not much more we can do from our side...</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 365967,
          "author_name": "Mukesh",
          "author_url": "",
          "post_date": "2018-08-03T18:56:57.513000",
          "content": "<p>Thats alright Jordi, figured something else, will see if it works.</p>\n\n<p>Cheers</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 394891,
      "author_name": "AR.Regel",
      "author_url": "",
      "post_date": "2018-09-27T15:56:29.810000",
      "content": "<p>Has anyone been able to get the entire training set (513gb) into a google cloud bucket?  I have tried the google transfer function (permissions didn't work), using awscli with a google instance and the subsets (failed due to 'device is full', which isn't true as my instance has 50gb of memory), and downloading directly to my computer then uploading (prohibitivly slow).</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 386637,
      "author_name": "av.in",
      "author_url": "",
      "post_date": "2018-09-13T09:25:50.147000",
      "content": "<p>The S3 buckets for images with bounding box are - </p>\n\n<p>s3://open-images-dataset/train (513GB)\ns3://open-images-dataset/validation (12GB)\ns3://open-images-dataset/test (36GB)</p>\n\n<p>One way to get them to GCS bucket is by creating GCS Transfer Job function. All my attempt to transfer from S3 bucket to GCS has failed as with error - \"Invalid access key\". I have tried to use my own AWS IAM credentials but they don't work. </p>\n\n<p>How can the data be transferred directly from S3 to GCS bucket directly??</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 360744,
      "author_name": "Avantis",
      "author_url": "",
      "post_date": "2018-07-23T06:30:31.220000",
      "content": "<p>Thanks for the link. Your link states, \"Google Storage provides a \"storage transfer\" function to transfer online files into a storage bucket ... The size of the whole dataset is around 18TB. Please note that user needs to pay for hosting the dataset on Google Cloud storage...\" My data size estimate of 16TB is from this site: <a href=\"https://blog.algorithmia.com/deep-dive-into-object-detection-with-open-images-using-tensorflow/\">https://blog.algorithmia.com/deep-dive-into-object-detection-with-open-images-using-tensorflow/</a>. I was looking for an option to access the dataset free of cost, an option that also allows reading the data by pointing to a single URL. Any help is greatly appreciated. </p>",
      "votes": 0,
      "replies": [
        {
          "id": 360778,
          "author_name": "Moshel",
          "author_url": "",
          "post_date": "2018-07-23T08:06:44.453000",
          "content": "<p>This is for the full database. Its easy to make this mistake. Google has 2 kinds of database. One has heen bboxed by humans and we should use it for the training. It has 1.7m images iirc and is around 530gb (tiny!). The other, which is the full database, has been annotated by machines and has 19m images iirc.\nThe aws link will download the database you want. The Google one will load the huge one.\nI haven't tried it, but I think you can filter the tsv using the annotations and download only the training set we need. It would be ultra super nice if Google could actually put the database in tfrecord format somewhere as even loading this much data into tfrecord would be a very lengthy process, to say nothing of twice the amount of local storage needed. </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 360789,
          "author_name": "Jordi Pont-Tuset",
          "author_url": "",
          "post_date": "2018-07-23T08:37:43.783000",
          "content": "<p>Dear Avantis,\nOpen Images has indeed different subsets:</p>\n\n<ol>\n<li>Subset annotated with bounding boxes (1,743,042 training images)</li>\n<li>Subset with human-verified image-level labels (5,655,108 training images)</li>\n<li>Subset with machine-generated image-level labels (8,853,429 training images)</li>\n<li>Full dataset (9,178,275 images)</li>\n</ol>\n\n<p>The images of subset 1. take up approximately 500GB and can be downloaded from AWS via CVDF or from Open Images. The full dataset takes up approximately 19TB and right now can only be downloaded directly from Flickr using the links provided by CVDF it TSV files.</p>\n\n<p>All this information (and much more) in the Open Images Website.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 360515,
      "author_name": "Avantis",
      "author_url": "",
      "post_date": "2018-07-22T16:22:36.777000",
      "content": "<p>The complete dataset, downloaded and uncompressed, is around 6TB. Is that about right? Given the vast storage requirement, is the entire dataset accessible via Google Storage?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 360565,
          "author_name": "Jordi Pont-Tuset",
          "author_url": "",
          "post_date": "2018-07-22T17:54:10.560000",
          "content": "<p>The full subset of Open Images where there are bounding boxes is:\n- Train: <strong>513GB</strong>\n- Test-Challenge 2018: <strong>10GB</strong></p>\n\n<p>More info in the CVDF download site:\n<a href=\"https://github.com/cvdfoundation/open-images-dataset#download-images-with-bounding-boxes-annotations\">https://github.com/cvdfoundation/open-images-dataset#download-images-with-bounding-boxes-annotations</a></p>\n\n<p>Not sure where these 6TB come from.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 361490,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-07-24T15:45:03.123000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 359698,
      "author_name": "peterlau",
      "author_url": "",
      "post_date": "2018-07-20T16:03:11.217000",
      "content": "<p>I would appreciate someone who could share the downloaded training dataset on the Internet.Downloading this huge training data is very time-consuming .</p>",
      "votes": 0,
      "replies": [
        {
          "id": 360386,
          "author_name": "Jordi Pont-Tuset",
          "author_url": "",
          "post_date": "2018-07-22T11:11:11.547000",
          "content": "<p>Dear peterlau,\nCVDF is working on packing the images into ZIP files, maybe that will make downloads easier. \nSorry for the inconvenience.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "365275": "Hi all, \n\nCVDF made zipped files available to download, dividing the train set in 16 files, e.g.:\naws s3 --no-sign-request cp s3://open-images-dataset/tar/train_0.tar.gz [target_dir] (46G)\nLet's hope this makes downloads easier.\n\nBest,",
    "394891": "Has anyone been able to get the entire training set (513gb) into a google cloud bucket?  I have tried the google transfer function (permissions didn't work), using awscli with a google instance and the subsets (failed due to 'device is full', which isn't true as my instance has 50gb of memory), and downloading directly to my computer then uploading (prohibitivly slow).",
    "386637": "The S3 buckets for images with bounding box are - \n\ns3://open-images-dataset/train (513GB)\ns3://open-images-dataset/validation (12GB)\ns3://open-images-dataset/test (36GB)\n\nOne way to get them to GCS bucket is by creating GCS Transfer Job function. All my attempt to transfer from S3 bucket to GCS has failed as with error - \"Invalid access key\". I have tried to use my own AWS IAM credentials but they don't work. \n\nHow can the data be transferred directly from S3 to GCS bucket directly??",
    "360744": "Thanks for the link. Your link states, \"Google Storage provides a \"storage transfer\" function to transfer online files into a storage bucket ... The size of the whole dataset is around 18TB. Please note that user needs to pay for hosting the dataset on Google Cloud storage...\" My data size estimate of 16TB is from this site: https://blog.algorithmia.com/deep-dive-into-object-detection-with-open-images-using-tensorflow/. I was looking for an option to access the dataset free of cost, an option that also allows reading the data by pointing to a single URL. Any help is greatly appreciated. ",
    "360515": "The complete dataset, downloaded and uncompressed, is around 6TB. Is that about right? Given the vast storage requirement, is the entire dataset accessible via Google Storage?",
    "359698": "I would appreciate someone who could share the downloaded training dataset on the Internet.Downloading this huge training data is very time-consuming .",
    "359696": "Hello,kagglers:\nI am currently downloading training images through  aws s3 as stated before.However,when I have downloaded about  204G training data,I meet this problem:\nfatal error:('The read operation timed out').\nSo T manually tried the command 's3 --no-sign-request sync s3://open-images-dataset/train' again and again ,sometimes it works , but it fails most of the time.How can I solve this issue and  download the whole training dataset  sucecessfully?Thanks! "
  }
}