{
  "id": 60938,
  "title": "Can google preload the dataset in the cloud?",
  "url": "/competitions/google-ai-open-images-object-detection-track/discussion/60938",
  "author_name": "Moshel",
  "post_date": "2018-07-12T09:37:43.953000",
  "votes": 6,
  "comment_count": 32,
  "views": 0,
  "content": "<p>Hi,\nI was wondering.... as the dataset (open images) is humongous and will take forever to pull and consume resources, etc etc... Is it possible for google to preload it into a cloud bucket that is accessible to us? That would save us soooo much time and effort!\nNot trying to push it, google has gone out of their way already but as the time frame for the competition and the effort involved are so immense this will make us progress much faster.</p>",
  "messages": [
    {
      "id": 356014,
      "postDate": "2018-07-12T18:09:54.553Z",
      "content": "<p>Yes please, I do agree on your suggestion. Hoping that google representative/s will response to this concern.</p>",
      "rawMarkdown": "Yes please, I do agree on your suggestion. Hoping that google representative/s will response to this concern.",
      "votes": 8
    },
    {
      "id": 355753,
      "postDate": "2018-07-12T09:37:43.953Z",
      "content": "<p>Hi,\nI was wondering.... as the dataset (open images) is humongous and will take forever to pull and consume resources, etc etc... Is it possible for google to preload it into a cloud bucket that is accessible to us? That would save us soooo much time and effort!\nNot trying to push it, google has gone out of their way already but as the time frame for the competition and the effort involved are so immense this will make us progress much faster.</p>",
      "rawMarkdown": "Hi,\nI was wondering.... as the dataset (open images) is humongous and will take forever to pull and consume resources, etc etc... Is it possible for google to preload it into a cloud bucket that is accessible to us? That would save us soooo much time and effort!\nNot trying to push it, google has gone out of their way already but as the time frame for the competition and the effort involved are so immense this will make us progress much faster.",
      "votes": 5
    },
    {
      "id": 361223,
      "postDate": "2018-07-24T03:53:17.097Z",
      "content": "<p>i just posted a kernel that trim the tsv files so they contain only the images from the training set. I can't test it but the output looks good.</p>",
      "rawMarkdown": "i just posted a kernel that trim the tsv files so they contain only the images from the training set. I can't test it but the output looks good.",
      "votes": 1
    },
    {
      "id": 360966,
      "postDate": "2018-07-23T15:31:45.057Z",
      "content": "<p>Hi, I found this <a href=\"https://cloud.google.com/bigquery/public-data/openimages\">https://cloud.google.com/bigquery/public-data/openimages</a> It looks like data already in Google Cloud Storage. The problem is how to query those data into our created Bucket, isn't it?</p>",
      "rawMarkdown": "Hi, I found this https://cloud.google.com/bigquery/public-data/openimages It looks like data already in Google Cloud Storage. The problem is how to query those data into our created Bucket, isn't it?",
      "replies": [
        {
          "id": 361059,
          "postDate": "2018-07-23T19:18:56.397Z",
          "content": "<p>It looks like that's the V3 version of the dataset, which doesn't have the bounding box annotations that this competition focuses on.</p>",
          "rawMarkdown": "It looks like that's the V3 version of the dataset, which doesn't have the bounding box annotations that this competition focuses on.",
          "votes": 1
        },
        {
          "id": 361310,
          "postDate": "2018-07-24T08:16:01.277Z",
          "content": "<p>Thank you Nathaniel. I am also reading the other thread \"Cannot download data\". I see your plan to use Google Cloud Computing with the competition, do you? That's exactly what I have in my mind. Do you manage to get the data to, somehow, work with your cloud computing engine?  If you succeed, please enlight me. I have been trying for 3 days already. Thank you.</p>",
          "rawMarkdown": "Thank you Nathaniel. I am also reading the other thread \"Cannot download data\". I see your plan to use Google Cloud Computing with the competition, do you? That's exactly what I have in my mind. Do you manage to get the data to, somehow, work with your cloud computing engine?  If you succeed, please enlight me. I have been trying for 3 days already. Thank you."
        },
        {
          "id": 361632,
          "postDate": "2018-07-24T20:19:31.043Z",
          "content": "<p>Not yet.  I've contacted Google support who is claiming that they need CVDF to provide an \"access/secret key pair\" even if they have \"grant everyone\" set on the S3 bucket.</p>\n\n<p>I was also counting on using the $500 Google Cloud credit that they were giving out, but since I couldn't get the data and then make my first submission, I missed out on that too.  So, I'm probably going to sit this one out.</p>",
          "rawMarkdown": "Not yet.  I've contacted Google support who is claiming that they need CVDF to provide an \"access/secret key pair\" even if they have \"grant everyone\" set on the S3 bucket.\n\nI was also counting on using the $500 Google Cloud credit that they were giving out, but since I couldn't get the data and then make my first submission, I missed out on that too.  So, I'm probably going to sit this one out."
        },
        {
          "id": 361885,
          "postDate": "2018-07-25T08:54:32.420Z",
          "content": "<p>It seems that we  can use Amazon Bucket with Google Compute Engine. Have you ever think of this combination?</p>",
          "rawMarkdown": "It seems that we  can use Amazon Bucket with Google Compute Engine. Have you ever think of this combination?"
        },
        {
          "id": 362172,
          "postDate": "2018-07-25T20:19:42.407Z",
          "content": "<p>I thought about that, as I was able to mount the S3 bucket into my VM using <code>s3fs</code>, but since simply copying the data ran into errors relatively quickly, I didn't want to risk relying on direct access just to have it fail me.  But if you're willing to risk it, go for it!</p>\n\n<p>Here's a comment from another Kaggler explaining how to mount it: <a href=\"https://www.kaggle.com/c/google-ai-open-images-object-detection-track/discussion/60543#356898\">https://www.kaggle.com/c/google-ai-open-images-object-detection-track/discussion/60543#356898</a></p>",
          "rawMarkdown": "I thought about that, as I was able to mount the S3 bucket into my VM using `s3fs`, but since simply copying the data ran into errors relatively quickly, I didn't want to risk relying on direct access just to have it fail me.  But if you're willing to risk it, go for it!\n\nHere's a comment from another Kaggler explaining how to mount it: https://www.kaggle.com/c/google-ai-open-images-object-detection-track/discussion/60543#356898"
        },
        {
          "id": 362175,
          "postDate": "2018-07-25T20:40:00.743Z",
          "content": "<p>I still think the competition organizers should have pre loaded the dataset in a bucket, preferably in tfrecord format. We are spending so much time on problems that are not related to the competition. In a perfect world you would be able to just mount the bucket read only. That would be so nice... Actually, now that the kaggle datasets are mounted that way, they probably could put it on kaggle and make playing with it even easier for us. </p>",
          "rawMarkdown": "I still think the competition organizers should have pre loaded the dataset in a bucket, preferably in tfrecord format. We are spending so much time on problems that are not related to the competition. In a perfect world you would be able to just mount the bucket read only. That would be so nice... Actually, now that the kaggle datasets are mounted that way, they probably could put it on kaggle and make playing with it even easier for us. "
        },
        {
          "id": 362359,
          "postDate": "2018-07-26T07:42:43.947Z",
          "content": "<p>Why do you need to copy the data? Can we just mount s3:image_backet to google compute engine Ubuntu, and then just Read only? I manage to mount the s3 image bucket to google compute engine. But I have been warned that doing so might slower my calculation.</p>",
          "rawMarkdown": "Why do you need to copy the data? Can we just mount s3:image_backet to google compute engine Ubuntu, and then just Read only? I manage to mount the s3 image bucket to google compute engine. But I have been warned that doing so might slower my calculation."
        },
        {
          "id": 362360,
          "postDate": "2018-07-26T07:43:29.327Z",
          "content": "<p>@Moshel have you figure out how to use google compute engine with S3 bucket?</p>",
          "rawMarkdown": "@Moshel have you figure out how to use google compute engine with S3 bucket?"
        },
        {
          "id": 362404,
          "postDate": "2018-07-26T09:23:44.863Z",
          "content": "<p>There is no point in that, as these are separate infrastructures (google and amazon) so any file access will actually happen over the network (slow and unreliable). Google cloud bucket are like local disks, at least in theory. I ended up using the awscli to download. It goes well so far (half downloaded)</p>",
          "rawMarkdown": "There is no point in that, as these are separate infrastructures (google and amazon) so any file access will actually happen over the network (slow and unreliable). Google cloud bucket are like local disks, at least in theory. I ended up using the awscli to download. It goes well so far (half downloaded)"
        },
        {
          "id": 362509,
          "postDate": "2018-07-26T14:12:26.340Z",
          "content": "<p>Thank you Moshel. So after you use aws to download to your local, Are you planning to upload them to google cloud bucket? </p>",
          "rawMarkdown": "Thank you Moshel. So after you use aws to download to your local, Are you planning to upload them to google cloud bucket? "
        },
        {
          "id": 362624,
          "postDate": "2018-07-26T18:55:27.677Z",
          "content": "<p>No, i downloaded the data to a cloud compter. I'll run my stuff there. It's less limited then bucket as you can use any method easily</p>",
          "rawMarkdown": "No, i downloaded the data to a cloud compter. I'll run my stuff there. It's less limited then bucket as you can use any method easily"
        },
        {
          "id": 362722,
          "postDate": "2018-07-27T01:30:39.733Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 362723,
          "postDate": "2018-07-27T01:36:49.933Z",
          "content": "<p>I have created a compute cloud unix instance with 2tb ssd and gpu. The awscli command downloaded the whole thing in under 24h. I had to run it twice as 2 images didn't download the first time. </p>",
          "rawMarkdown": "I have created a compute cloud unix instance with 2tb ssd and gpu. The awscli command downloaded the whole thing in under 24h. I had to run it twice as 2 images didn't download the first time. "
        },
        {
          "id": 362735,
          "postDate": "2018-07-27T02:14:17.380Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 362738,
          "postDate": "2018-07-27T02:42:39.730Z",
          "content": "<p>Not really complicated. And it gives you options that are not tensorflow (blesphamy!). Anyways, after this worked so well i didn't try anything else. I didn't even try my own tsv trimmer.... Although it should work really well. </p>",
          "rawMarkdown": "Not really complicated. And it gives you options that are not tensorflow (blesphamy!). Anyways, after this worked so well i didn't try anything else. I didn't even try my own tsv trimmer.... Although it should work really well. "
        },
        {
          "id": 362772,
          "postDate": "2018-07-27T05:14:37.023Z",
          "content": "<p>I found a way to download the dataset to google cloud bucket</p>\n\n<ol>\n<li><p>Mount google bucket to google cloud computing VM's file system with this command\nYou need .json file. The following stackoverflow reference shows how to get your json file. Upload the json file to the vm's file system.  Please change these parameters to your own:gid, uid, key-file is your json location on the vm file system. openimagedataset is my bucket name, and gs_openimages is my mount point.</p>\n\n<p>mkdir gs_openimages # Mount point</p>\n\n<p>gcsfuse -o allow_other --gid 1002 --uid 1002 --file-mode 777 --dir-mode 777 --key-file $PWD/open.json openimagedataset $PWD/gs_openimages</p></li>\n</ol>\n\n<p><a href=\"https://stackoverflow.com/questions/42630856/how-to-mount-google-bucket-as-local-disk-on-linux-instance-with-full-access-righ\">https://stackoverflow.com/questions/42630856/how-to-mount-google-bucket-as-local-disk-on-linux-instance-with-full-access-righ</a></p>\n\n<p>Testing: If you can create a file in the mounted folder, then its ok</p>\n\n<ol>\n<li>use aws comamnd to download. Installation of aws command please follow  <a href=\"https://github.com/cvdfoundation/open-images-dataset#download-images-with-bounding-boxes-annotations\">https://github.com/cvdfoundation/open-images-dataset#download-images-with-bounding-boxes-annotations</a>. </li>\n</ol>\n\n<p>aws s3 --no-sign-request sync s3://open-images-dataset/train gs_openimages/OpenImages/train</p>",
          "rawMarkdown": "I found a way to download the dataset to google cloud bucket\n\n1.  Mount google bucket to google cloud computing VM's file system with this command\nYou need .json file. The following stackoverflow reference shows how to get your json file. Upload the json file to the vm's file system.  Please change these parameters to your own:gid, uid, key-file is your json location on the vm file system. openimagedataset is my bucket name, and gs_openimages is my mount point.\n\n    mkdir gs_openimages # Mount point\n    \n    gcsfuse -o allow_other --gid 1002 --uid 1002 --file-mode 777 --dir-mode 777 --key-file $PWD/open.json openimagedataset $PWD/gs_openimages\n\nhttps://stackoverflow.com/questions/42630856/how-to-mount-google-bucket-as-local-disk-on-linux-instance-with-full-access-righ\n\nTesting: If you can create a file in the mounted folder, then its ok\n\n2. use aws comamnd to download. Installation of aws command please follow  https://github.com/cvdfoundation/open-images-dataset#download-images-with-bounding-boxes-annotations. \n\naws s3 --no-sign-request sync s3://open-images-dataset/train gs_openimages/OpenImages/train\n"
        },
        {
          "id": 363179,
          "postDate": "2018-07-28T05:07:41.520Z",
          "content": "<p>Would anyone be so kind as to <a href=\"https://cloud.google.com/bigquery/docs/loading-data-cloud-storage\">updating / loading</a> their data to BigQuery and make it public dataset as the OpenImage v3 BigQuery dataset? <br>\nI will be grateful if anyone who already succeeded to copy from s3 to your own google storage bucket could do this? <br>\nAnd maybe follow <a href=\"https://www.kaggle.com/product-feedback/48573#latest-308483\">this post</a>, we can treat this public dataset as BigQuery Datasource used in Kaggle kernel? <br>\nWould this work with everyone here?    </p>",
          "rawMarkdown": "Would anyone be so kind as to [updating / loading](https://cloud.google.com/bigquery/docs/loading-data-cloud-storage) their data to BigQuery and make it public dataset as the OpenImage v3 BigQuery dataset?     \nI will be grateful if anyone who already succeeded to copy from s3 to your own google storage bucket could do this?      \nAnd maybe follow [this post](https://www.kaggle.com/product-feedback/48573#latest-308483), we can treat this public dataset as BigQuery Datasource used in Kaggle kernel?     \nWould this work with everyone here?    "
        }
      ]
    },
    {
      "id": 355951,
      "postDate": "2018-07-12T16:10:53.737Z",
      "content": "<p>Sounds good to me, let's hope Google engineers read this!</p>",
      "rawMarkdown": "Sounds good to me, let's hope Google engineers read this!"
    },
    {
      "id": 355756,
      "postDate": "2018-07-12T09:38:54.477Z",
      "content": "<p>Just to be extra clear - I am talking about the train/valuation dataset, not the test</p>",
      "rawMarkdown": "Just to be extra clear - I am talking about the train/valuation dataset, not the test",
      "replies": [
        {
          "id": 356322,
          "postDate": "2018-07-13T10:54:17.913Z",
          "content": "<p>Yes thats good mate. Does the competition hosts do involve in hosts or only they read only masters or grand masters posts?</p>",
          "rawMarkdown": "Yes thats good mate. Does the competition hosts do involve in hosts or only they read only masters or grand masters posts?"
        },
        {
          "id": 356530,
          "postDate": "2018-07-13T20:11:49.080Z",
          "content": "<p>They did it, to an extent, but it was very cleverly hidden in a post about aws availability (shrug). Here is the link to the link with the information towards the end <a href=\"https://github.com/cvdfoundation/open-images-dataset/blob/master/README.md\">https://github.com/cvdfoundation/open-images-dataset/blob/master/README.md</a></p>",
          "rawMarkdown": "They did it, to an extent, but it was very cleverly hidden in a post about aws availability (shrug). Here is the link to the link with the information towards the end https://github.com/cvdfoundation/open-images-dataset/blob/master/README.md",
          "votes": 2
        },
        {
          "id": 356532,
          "postDate": "2018-07-13T20:15:33.020Z",
          "content": "<p>Woww I am checking for this information dude . Well done mate</p>",
          "rawMarkdown": "Woww I am checking for this information dude . Well done mate"
        },
        {
          "id": 356559,
          "postDate": "2018-07-13T21:27:35.277Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 356709,
          "postDate": "2018-07-14T09:31:00.830Z",
          "content": "<p>Hiii Moshel ,  I have loved your work on datascience . Can we be a team?</p>",
          "rawMarkdown": "Hiii Moshel ,  I have loved your work on datascience . Can we be a team?"
        },
        {
          "id": 356736,
          "postDate": "2018-07-14T11:01:48.163Z",
          "content": "<p>Honestly, I don't think it is in your best interest... I am truly a beginner and the time I can invest in this is very limited.</p>",
          "rawMarkdown": "Honestly, I don't think it is in your best interest... I am truly a beginner and the time I can invest in this is very limited."
        },
        {
          "id": 357757,
          "postDate": "2018-07-16T19:36:39.597Z",
          "content": "<p>Thanks for sharing the info about storage transfer.  I noticed that the Google storage transfer option listed was only for the whole 18TB dataset.  Does anyone know if there's a Google storage transfer option for just the challenge portion of the dataset 561GB?  I know I can get it from AWS, but I'd prefer not to pull it from a 3rd party if the option is available from within GCP.</p>",
          "rawMarkdown": "Thanks for sharing the info about storage transfer.  I noticed that the Google storage transfer option listed was only for the whole 18TB dataset.  Does anyone know if there's a Google storage transfer option for just the challenge portion of the dataset 561GB?  I know I can get it from AWS, but I'd prefer not to pull it from a 3rd party if the option is available from within GCP."
        }
      ]
    },
    {
      "id": 362769,
      "postDate": "2018-07-27T04:59:22.903Z",
      "rawMarkdown": "",
      "isDeleted": true,
      "replies": [
        {
          "id": 362773,
          "postDate": "2018-07-27T05:16:35.683Z",
          "content": "<p>flooksta. Perhaps you could try my method above. If we both are ok with this method, we could share it publicly.</p>",
          "rawMarkdown": "flooksta. Perhaps you could try my method above. If we both are ok with this method, we could share it publicly."
        },
        {
          "id": 362788,
          "postDate": "2018-07-27T05:57:30.970Z",
          "rawMarkdown": "",
          "isDeleted": true
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 356014,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-07-12T18:09:54.553000",
      "content": "<p>Yes please, I do agree on your suggestion. Hoping that google representative/s will response to this concern.</p>",
      "votes": 8,
      "replies": []
    },
    {
      "id": 361223,
      "author_name": "Moshel",
      "author_url": "",
      "post_date": "2018-07-24T03:53:17.097000",
      "content": "<p>i just posted a kernel that trim the tsv files so they contain only the images from the training set. I can't test it but the output looks good.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 360966,
      "author_name": "Peerajak Witoonchart",
      "author_url": "",
      "post_date": "2018-07-23T15:31:45.057000",
      "content": "<p>Hi, I found this <a href=\"https://cloud.google.com/bigquery/public-data/openimages\">https://cloud.google.com/bigquery/public-data/openimages</a> It looks like data already in Google Cloud Storage. The problem is how to query those data into our created Bucket, isn't it?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 361059,
          "author_name": "Nathaniel Watkins",
          "author_url": "",
          "post_date": "2018-07-23T19:18:56.397000",
          "content": "<p>It looks like that's the V3 version of the dataset, which doesn't have the bounding box annotations that this competition focuses on.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 361310,
          "author_name": "Peerajak Witoonchart",
          "author_url": "",
          "post_date": "2018-07-24T08:16:01.277000",
          "content": "<p>Thank you Nathaniel. I am also reading the other thread \"Cannot download data\". I see your plan to use Google Cloud Computing with the competition, do you? That's exactly what I have in my mind. Do you manage to get the data to, somehow, work with your cloud computing engine?  If you succeed, please enlight me. I have been trying for 3 days already. Thank you.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 361632,
          "author_name": "Nathaniel Watkins",
          "author_url": "",
          "post_date": "2018-07-24T20:19:31.043000",
          "content": "<p>Not yet.  I've contacted Google support who is claiming that they need CVDF to provide an \"access/secret key pair\" even if they have \"grant everyone\" set on the S3 bucket.</p>\n\n<p>I was also counting on using the $500 Google Cloud credit that they were giving out, but since I couldn't get the data and then make my first submission, I missed out on that too.  So, I'm probably going to sit this one out.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 361885,
          "author_name": "Peerajak Witoonchart",
          "author_url": "",
          "post_date": "2018-07-25T08:54:32.420000",
          "content": "<p>It seems that we  can use Amazon Bucket with Google Compute Engine. Have you ever think of this combination?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 362172,
          "author_name": "Nathaniel Watkins",
          "author_url": "",
          "post_date": "2018-07-25T20:19:42.407000",
          "content": "<p>I thought about that, as I was able to mount the S3 bucket into my VM using <code>s3fs</code>, but since simply copying the data ran into errors relatively quickly, I didn't want to risk relying on direct access just to have it fail me.  But if you're willing to risk it, go for it!</p>\n\n<p>Here's a comment from another Kaggler explaining how to mount it: <a href=\"https://www.kaggle.com/c/google-ai-open-images-object-detection-track/discussion/60543#356898\">https://www.kaggle.com/c/google-ai-open-images-object-detection-track/discussion/60543#356898</a></p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 362175,
          "author_name": "Moshel",
          "author_url": "",
          "post_date": "2018-07-25T20:40:00.743000",
          "content": "<p>I still think the competition organizers should have pre loaded the dataset in a bucket, preferably in tfrecord format. We are spending so much time on problems that are not related to the competition. In a perfect world you would be able to just mount the bucket read only. That would be so nice... Actually, now that the kaggle datasets are mounted that way, they probably could put it on kaggle and make playing with it even easier for us. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 362359,
          "author_name": "Peerajak Witoonchart",
          "author_url": "",
          "post_date": "2018-07-26T07:42:43.947000",
          "content": "<p>Why do you need to copy the data? Can we just mount s3:image_backet to google compute engine Ubuntu, and then just Read only? I manage to mount the s3 image bucket to google compute engine. But I have been warned that doing so might slower my calculation.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 362360,
          "author_name": "Peerajak Witoonchart",
          "author_url": "",
          "post_date": "2018-07-26T07:43:29.327000",
          "content": "<p>@Moshel have you figure out how to use google compute engine with S3 bucket?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 362404,
          "author_name": "Moshel",
          "author_url": "",
          "post_date": "2018-07-26T09:23:44.863000",
          "content": "<p>There is no point in that, as these are separate infrastructures (google and amazon) so any file access will actually happen over the network (slow and unreliable). Google cloud bucket are like local disks, at least in theory. I ended up using the awscli to download. It goes well so far (half downloaded)</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 362509,
          "author_name": "Peerajak Witoonchart",
          "author_url": "",
          "post_date": "2018-07-26T14:12:26.340000",
          "content": "<p>Thank you Moshel. So after you use aws to download to your local, Are you planning to upload them to google cloud bucket? </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 362624,
          "author_name": "Moshel",
          "author_url": "",
          "post_date": "2018-07-26T18:55:27.677000",
          "content": "<p>No, i downloaded the data to a cloud compter. I'll run my stuff there. It's less limited then bucket as you can use any method easily</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 362722,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-07-27T01:30:39.733000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 362723,
          "author_name": "Moshel",
          "author_url": "",
          "post_date": "2018-07-27T01:36:49.933000",
          "content": "<p>I have created a compute cloud unix instance with 2tb ssd and gpu. The awscli command downloaded the whole thing in under 24h. I had to run it twice as 2 images didn't download the first time. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 362735,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-07-27T02:14:17.380000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 362738,
          "author_name": "Moshel",
          "author_url": "",
          "post_date": "2018-07-27T02:42:39.730000",
          "content": "<p>Not really complicated. And it gives you options that are not tensorflow (blesphamy!). Anyways, after this worked so well i didn't try anything else. I didn't even try my own tsv trimmer.... Although it should work really well. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 362772,
          "author_name": "Peerajak Witoonchart",
          "author_url": "",
          "post_date": "2018-07-27T05:14:37.023000",
          "content": "<p>I found a way to download the dataset to google cloud bucket</p>\n\n<ol>\n<li><p>Mount google bucket to google cloud computing VM's file system with this command\nYou need .json file. The following stackoverflow reference shows how to get your json file. Upload the json file to the vm's file system.  Please change these parameters to your own:gid, uid, key-file is your json location on the vm file system. openimagedataset is my bucket name, and gs_openimages is my mount point.</p>\n\n<p>mkdir gs_openimages # Mount point</p>\n\n<p>gcsfuse -o allow_other --gid 1002 --uid 1002 --file-mode 777 --dir-mode 777 --key-file $PWD/open.json openimagedataset $PWD/gs_openimages</p></li>\n</ol>\n\n<p><a href=\"https://stackoverflow.com/questions/42630856/how-to-mount-google-bucket-as-local-disk-on-linux-instance-with-full-access-righ\">https://stackoverflow.com/questions/42630856/how-to-mount-google-bucket-as-local-disk-on-linux-instance-with-full-access-righ</a></p>\n\n<p>Testing: If you can create a file in the mounted folder, then its ok</p>\n\n<ol>\n<li>use aws comamnd to download. Installation of aws command please follow  <a href=\"https://github.com/cvdfoundation/open-images-dataset#download-images-with-bounding-boxes-annotations\">https://github.com/cvdfoundation/open-images-dataset#download-images-with-bounding-boxes-annotations</a>. </li>\n</ol>\n\n<p>aws s3 --no-sign-request sync s3://open-images-dataset/train gs_openimages/OpenImages/train</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 363179,
          "author_name": "Rene Wang",
          "author_url": "",
          "post_date": "2018-07-28T05:07:41.520000",
          "content": "<p>Would anyone be so kind as to <a href=\"https://cloud.google.com/bigquery/docs/loading-data-cloud-storage\">updating / loading</a> their data to BigQuery and make it public dataset as the OpenImage v3 BigQuery dataset? <br>\nI will be grateful if anyone who already succeeded to copy from s3 to your own google storage bucket could do this? <br>\nAnd maybe follow <a href=\"https://www.kaggle.com/product-feedback/48573#latest-308483\">this post</a>, we can treat this public dataset as BigQuery Datasource used in Kaggle kernel? <br>\nWould this work with everyone here?    </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 355951,
      "author_name": "Nazim Girach",
      "author_url": "",
      "post_date": "2018-07-12T16:10:53.737000",
      "content": "<p>Sounds good to me, let's hope Google engineers read this!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 355756,
      "author_name": "Moshel",
      "author_url": "",
      "post_date": "2018-07-12T09:38:54.477000",
      "content": "<p>Just to be extra clear - I am talking about the train/valuation dataset, not the test</p>",
      "votes": 0,
      "replies": [
        {
          "id": 356322,
          "author_name": "dineshbarri",
          "author_url": "",
          "post_date": "2018-07-13T10:54:17.913000",
          "content": "<p>Yes thats good mate. Does the competition hosts do involve in hosts or only they read only masters or grand masters posts?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 356530,
          "author_name": "Moshel",
          "author_url": "",
          "post_date": "2018-07-13T20:11:49.080000",
          "content": "<p>They did it, to an extent, but it was very cleverly hidden in a post about aws availability (shrug). Here is the link to the link with the information towards the end <a href=\"https://github.com/cvdfoundation/open-images-dataset/blob/master/README.md\">https://github.com/cvdfoundation/open-images-dataset/blob/master/README.md</a></p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 356532,
          "author_name": "dineshbarri",
          "author_url": "",
          "post_date": "2018-07-13T20:15:33.020000",
          "content": "<p>Woww I am checking for this information dude . Well done mate</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 356559,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-07-13T21:27:35.277000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 356709,
          "author_name": "dineshbarri",
          "author_url": "",
          "post_date": "2018-07-14T09:31:00.830000",
          "content": "<p>Hiii Moshel ,  I have loved your work on datascience . Can we be a team?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 356736,
          "author_name": "Moshel",
          "author_url": "",
          "post_date": "2018-07-14T11:01:48.163000",
          "content": "<p>Honestly, I don't think it is in your best interest... I am truly a beginner and the time I can invest in this is very limited.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 357757,
          "author_name": "Nathaniel Watkins",
          "author_url": "",
          "post_date": "2018-07-16T19:36:39.597000",
          "content": "<p>Thanks for sharing the info about storage transfer.  I noticed that the Google storage transfer option listed was only for the whole 18TB dataset.  Does anyone know if there's a Google storage transfer option for just the challenge portion of the dataset 561GB?  I know I can get it from AWS, but I'd prefer not to pull it from a 3rd party if the option is available from within GCP.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 362769,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-07-27T04:59:22.903000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 362773,
          "author_name": "Peerajak Witoonchart",
          "author_url": "",
          "post_date": "2018-07-27T05:16:35.683000",
          "content": "<p>flooksta. Perhaps you could try my method above. If we both are ok with this method, we could share it publicly.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 362788,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-07-27T05:57:30.970000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "356014": "Yes please, I do agree on your suggestion. Hoping that google representative/s will response to this concern.",
    "355753": "Hi,\nI was wondering.... as the dataset (open images) is humongous and will take forever to pull and consume resources, etc etc... Is it possible for google to preload it into a cloud bucket that is accessible to us? That would save us soooo much time and effort!\nNot trying to push it, google has gone out of their way already but as the time frame for the competition and the effort involved are so immense this will make us progress much faster.",
    "361223": "i just posted a kernel that trim the tsv files so they contain only the images from the training set. I can't test it but the output looks good.",
    "360966": "Hi, I found this https://cloud.google.com/bigquery/public-data/openimages It looks like data already in Google Cloud Storage. The problem is how to query those data into our created Bucket, isn't it?",
    "355951": "Sounds good to me, let's hope Google engineers read this!",
    "355756": "Just to be extra clear - I am talking about the train/valuation dataset, not the test",
    "362769": ""
  }
}