{
  "id": 66955,
  "title": "Help in Transferring Data to GCP Bucket from AWS",
  "url": "/competitions/inclusive-images-challenge/discussion/66955",
  "author_name": "",
  "post_date": "2018-09-27T02:54:33.867278200Z",
  "votes": 1,
  "comment_count": 10,
  "views": 0,
  "content": "<p>I received <a href=\"https://www.kaggle.com/c/inclusive-images-challenge/discussion/65157\">Google Cloud Platform coupon credits</a> which is great, but now I've been scratching my head on how to get the data from AWS to GCP for the past couple days. </p>\n\n<p>I'm relatively new in dealing with large datasets (&gt;10GB) like this and I can't seem to figure out an effective way to upload the data to my Google Cloud bucket. I've slogged through a lot of documentation and I've gotten a little overwhelmed since I just can't seem to find a solution that works for me. </p>\n\n<p><a href=\"https://www.kaggle.com/sepeagb2\">Sam Epeagba</a> had some great <a href=\"https://www.kaggle.com/c/inclusive-images-challenge/discussion/66215\">instructions</a> on using <strong>gcsfuse</strong> but I found the transfer speed was much too slow. (I'm sure there's a way to specify faster transfer speeds, I just don't see how to do that properly.)</p>\n\n<p>I've also tried the Google Console's <a href=\"https://console.cloud.google.com/storage/transfer\">transfer page</a> to create a transfer job. Using AWS as the source (<code>open-images-dataset/validation</code> for example) I get <code>Invalid access key. Make sure the access key for your S3 bucket is correct.</code> error even after creating an IAM for AWS (AmazonS3FullAccess permission). I've also looked at <a href=\"https://github.com/cvdfoundation/open-images-dataset#download-full-dataset-with-google-storage-transfer\">CVDF's instructions</a> in using Google's Storage transfer via the tsv files, but of course that is the full 18TB dataset and I just want the training, validation, and testing sets (~550GB) for this challenge.</p>\n\n<p>I've seen people suggest downloading the data to their local machines and then uploading the files to a Google Cloud bucket from there but I don't have a local machine that could handle that in regards to storage space and upload speed.</p>\n\n<p>Any help would be greatly appreciated. I'm hoping that people with more experience can point me in some right directions. Maybe there are some other inexperienced teams like me that could benefit from this too.</p>",
  "messages": [
    {
      "id": "394530",
      "postDate": "09/27/2018 02:54:33",
      "content": "<p>I received <a href=\"https://www.kaggle.com/c/inclusive-images-challenge/discussion/65157\">Google Cloud Platform coupon credits</a> which is great, but now I've been scratching my head on how to get the data from AWS to GCP for the past couple days. </p>\n\n<p>I'm relatively new in dealing with large datasets (&gt;10GB) like this and I can't seem to figure out an effective way to upload the data to my Google Cloud bucket. I've slogged through a lot of documentation and I've gotten a little overwhelmed since I just can't seem to find a solution that works for me. </p>\n\n<p><a href=\"https://www.kaggle.com/sepeagb2\">Sam Epeagba</a> had some great <a href=\"https://www.kaggle.com/c/inclusive-images-challenge/discussion/66215\">instructions</a> on using <strong>gcsfuse</strong> but I found the transfer speed was much too slow. (I'm sure there's a way to specify faster transfer speeds, I just don't see how to do that properly.)</p>\n\n<p>I've also tried the Google Console's <a href=\"https://console.cloud.google.com/storage/transfer\">transfer page</a> to create a transfer job. Using AWS as the source (<code>open-images-dataset/validation</code> for example) I get <code>Invalid access key. Make sure the access key for your S3 bucket is correct.</code> error even after creating an IAM for AWS (AmazonS3FullAccess permission). I've also looked at <a href=\"https://github.com/cvdfoundation/open-images-dataset#download-full-dataset-with-google-storage-transfer\">CVDF's instructions</a> in using Google's Storage transfer via the tsv files, but of course that is the full 18TB dataset and I just want the training, validation, and testing sets (~550GB) for this challenge.</p>\n\n<p>I've seen people suggest downloading the data to their local machines and then uploading the files to a Google Cloud bucket from there but I don't have a local machine that could handle that in regards to storage space and upload speed.</p>\n\n<p>Any help would be greatly appreciated. I'm hoping that people with more experience can point me in some right directions. Maybe there are some other inexperienced teams like me that could benefit from this too.</p>",
      "rawMarkdown": "I received [Google Cloud Platform coupon credits][1] which is great, but now I've been scratching my head on how to get the data from AWS to GCP for the past couple days. \n\nI'm relatively new in dealing with large datasets (&gt;10GB) like this and I can't seem to figure out an effective way to upload the data to my Google Cloud bucket. I've slogged through a lot of documentation and I've gotten a little overwhelmed since I just can't seem to find a solution that works for me. \n\n[Sam Epeagba][2] had some great [instructions][3] on using **gcsfuse** but I found the transfer speed was much too slow. (I'm sure there's a way to specify faster transfer speeds, I just don't see how to do that properly.)\n\nI've also tried the Google Console's [transfer page][4] to create a transfer job. Using AWS as the source (`open-images-dataset/validation` for example) I get `Invalid access key. Make sure the access key for your S3 bucket is correct.` error even after creating an IAM for AWS (AmazonS3FullAccess permission). I've also looked at [CVDF's instructions][5] in using Google's Storage transfer via the tsv files, but of course that is the full 18TB dataset and I just want the training, validation, and testing sets (~550GB) for this challenge.\n\nI've seen people suggest downloading the data to their local machines and then uploading the files to a Google Cloud bucket from there but I don't have a local machine that could handle that in regards to storage space and upload speed.\n\nAny help would be greatly appreciated. I'm hoping that people with more experience can point me in some right directions. Maybe there are some other inexperienced teams like me that could benefit from this too.\n\n  [1]: https://www.kaggle.com/c/inclusive-images-challenge/discussion/65157\n  [2]: https://www.kaggle.com/sepeagb2\n  [3]: https://www.kaggle.com/c/inclusive-images-challenge/discussion/66215\n  [4]: https://console.cloud.google.com/storage/transfer\n  [5]: https://github.com/cvdfoundation/open-images-dataset#download-full-dataset-with-google-storage-transfer",
      "votes": null
    },
    {
      "id": "394753",
      "postDate": "09/27/2018 11:41:21",
      "content": "<p>Ya, me too got stuck with this for a week, but here is the solution--</p>\n\n<p>Make a  basic VM  instance on google cloud.Download the files onto that instance .Then upload it to a cloud bucket of your choice. </p>",
      "rawMarkdown": "Ya, me too got stuck with this for a week, but here is the solution--\n\nMake a  basic VM  instance on google cloud.Download the files onto that instance .Then upload it to a cloud bucket of your choice.",
      "votes": null
    },
    {
      "id": "394758",
      "postDate": "09/27/2018 11:57:53",
      "content": "<p>You can use\ngsutil -m -o GSUtil:parallel_composite_upload_threshold=150M cp -r bigfilename gs://bucket-name\n to copy files in parallel.</p>",
      "rawMarkdown": "You can use\ngsutil -m -o GSUtil:parallel_composite_upload_threshold=150M cp -r bigfilename gs://bucket-name\n to copy files in parallel.",
      "votes": null
    },
    {
      "id": "394807",
      "postDate": "09/27/2018 13:44:28",
      "content": "<p>That makes sense but would that mean you'd have to essentially pay for the storage twice, once for the VM instance and once for the bucket? I figure you could delete the files from the VM after uploading to the bucket; I'm pretty new with using cloud storage and I don't know what costs to expect in this case.</p>\n\n<p>I might have to do it this way if nothing else works. I'll probably drop back into it after work today. Thanks for the advice!</p>",
      "rawMarkdown": "That makes sense but would that mean you'd have to essentially pay for the storage twice, once for the VM instance and once for the bucket? I figure you could delete the files from the VM after uploading to the bucket; I'm pretty new with using cloud storage and I don't know what costs to expect in this case.\n\nI might have to do it this way if nothing else works. I'll probably drop back into it after work today. Thanks for the advice!",
      "votes": null
    },
    {
      "id": "394814",
      "postDate": "09/27/2018 13:57:11",
      "content": "<p>I've seen using <strong>gsutil</strong> this way but can you replace the <code>bigfilename</code> (from your example) with a reference to a ASW S3 bucket? I think this is what someone also <a href=\"https://www.kaggle.com/c/inclusive-images-challenge/discussion/66232#394567\">suggested</a> in another discussion. I'll have to explore more and try this out after work.</p>\n\n<p>Thanks!</p>",
      "rawMarkdown": "I've seen using **gsutil** this way but can you replace the `bigfilename` (from your example) with a reference to a ASW S3 bucket? I think this is what someone also [suggested][1] in another discussion. I'll have to explore more and try this out after work.\n\nThanks!\n\n  [1]: https://www.kaggle.com/c/inclusive-images-challenge/discussion/66232#394567",
      "votes": null
    },
    {
      "id": "395596",
      "postDate": "09/28/2018 22:49:34",
      "content": "<p>Hey Victor, did this work for you? </p>",
      "rawMarkdown": "Hey Victor, did this work for you?",
      "votes": null
    },
    {
      "id": "395625",
      "postDate": "09/29/2018 02:08:48",
      "content": "<p>It did! Well at least much better than what I was able to achieve before. I think I can get better speeds with a different instance.</p>\n\n<p>I wrote a little more in another thread about what I did: <a href=\"https://www.kaggle.com/c/inclusive-images-challenge/discussion/66232#394811\">https://www.kaggle.com/c/inclusive-images-challenge/discussion/66232#394811</a></p>\n\n<p>I figured it might help someone else in a similar position as I was in.</p>",
      "rawMarkdown": "It did! Well at least much better than what I was able to achieve before. I think I can get better speeds with a different instance.\n\nI wrote a little more in another thread about what I did: https://www.kaggle.com/c/inclusive-images-challenge/discussion/66232#394811\n\nI figured it might help someone else in a similar position as I was in.",
      "votes": null
    },
    {
      "id": "399148",
      "postDate": "10/05/2018 09:38:30",
      "content": "<p>Just a  thought .. Can Google/Kaggle host the total data set on their Google drive or Google cloud bucket and provide means to access directly to work in google cloud.  will that be possible ?<br>\nWondering whether this can help all those who wish to use Google cloud and avoid creating additional account on aws just to download. <br> </p>",
      "rawMarkdown": "Just a  thought .. Can Google/Kaggle host the total data set on their Google drive or Google cloud bucket and provide means to access directly to work in google cloud.  will that be possible ?<br>\nWondering whether this can help all those who wish to use Google cloud and avoid creating additional account on aws just to download. <br>",
      "votes": null
    },
    {
      "id": "399462",
      "postDate": "10/05/2018 20:40:21",
      "content": "<p>This would really help....looking forward for a reply from the concerned people.</p>",
      "rawMarkdown": "This would really help....looking forward for a reply from the concerned people.",
      "votes": null
    },
    {
      "id": "400925",
      "postDate": "10/09/2018 06:36:22",
      "content": "<p>That would definitely be the best solution.</p>",
      "rawMarkdown": "That would definitely be the best solution.",
      "votes": null
    },
    {
      "id": "400926",
      "postDate": "10/09/2018 06:37:25",
      "content": "<p>I did this as well. </p>",
      "rawMarkdown": "I did this as well.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 394753,
      "author_name": "rajath25",
      "author_url": "",
      "post_date": "09/27/2018 11:41:21",
      "content": "<p>Ya, me too got stuck with this for a week, but here is the solution--</p>\n\n<p>Make a  basic VM  instance on google cloud.Download the files onto that instance .Then upload it to a cloud bucket of your choice. </p>",
      "votes": null,
      "replies": [
        {
          "id": 394807,
          "author_name": "mrgeislinger",
          "author_url": "",
          "post_date": "09/27/2018 13:44:28",
          "content": "<p>That makes sense but would that mean you'd have to essentially pay for the storage twice, once for the VM instance and once for the bucket? I figure you could delete the files from the VM after uploading to the bucket; I'm pretty new with using cloud storage and I don't know what costs to expect in this case.</p>\n\n<p>I might have to do it this way if nothing else works. I'll probably drop back into it after work today. Thanks for the advice!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 400926,
          "author_name": "flagshipdynamics",
          "author_url": "",
          "post_date": "10/09/2018 06:37:25",
          "content": "<p>I did this as well. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 394758,
      "author_name": "pbaljeka",
      "author_url": "",
      "post_date": "09/27/2018 11:57:53",
      "content": "<p>You can use\ngsutil -m -o GSUtil:parallel_composite_upload_threshold=150M cp -r bigfilename gs://bucket-name\n to copy files in parallel.</p>",
      "votes": null,
      "replies": [
        {
          "id": 394814,
          "author_name": "mrgeislinger",
          "author_url": "",
          "post_date": "09/27/2018 13:57:11",
          "content": "<p>I've seen using <strong>gsutil</strong> this way but can you replace the <code>bigfilename</code> (from your example) with a reference to a ASW S3 bucket? I think this is what someone also <a href=\"https://www.kaggle.com/c/inclusive-images-challenge/discussion/66232#394567\">suggested</a> in another discussion. I'll have to explore more and try this out after work.</p>\n\n<p>Thanks!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 395596,
          "author_name": "aliabdalla",
          "author_url": "",
          "post_date": "09/28/2018 22:49:34",
          "content": "<p>Hey Victor, did this work for you? </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 395625,
          "author_name": "mrgeislinger",
          "author_url": "",
          "post_date": "09/29/2018 02:08:48",
          "content": "<p>It did! Well at least much better than what I was able to achieve before. I think I can get better speeds with a different instance.</p>\n\n<p>I wrote a little more in another thread about what I did: <a href=\"https://www.kaggle.com/c/inclusive-images-challenge/discussion/66232#394811\">https://www.kaggle.com/c/inclusive-images-challenge/discussion/66232#394811</a></p>\n\n<p>I figured it might help someone else in a similar position as I was in.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 399148,
      "author_name": "reachkishore",
      "author_url": "",
      "post_date": "10/05/2018 09:38:30",
      "content": "<p>Just a  thought .. Can Google/Kaggle host the total data set on their Google drive or Google cloud bucket and provide means to access directly to work in google cloud.  will that be possible ?<br>\nWondering whether this can help all those who wish to use Google cloud and avoid creating additional account on aws just to download. <br> </p>",
      "votes": null,
      "replies": [
        {
          "id": 399462,
          "author_name": "srihaindavi",
          "author_url": "",
          "post_date": "10/05/2018 20:40:21",
          "content": "<p>This would really help....looking forward for a reply from the concerned people.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 400925,
          "author_name": "flagshipdynamics",
          "author_url": "",
          "post_date": "10/09/2018 06:36:22",
          "content": "<p>That would definitely be the best solution.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "394530": "I received [Google Cloud Platform coupon credits][1] which is great, but now I've been scratching my head on how to get the data from AWS to GCP for the past couple days. \n\nI'm relatively new in dealing with large datasets (&gt;10GB) like this and I can't seem to figure out an effective way to upload the data to my Google Cloud bucket. I've slogged through a lot of documentation and I've gotten a little overwhelmed since I just can't seem to find a solution that works for me. \n\n[Sam Epeagba][2] had some great [instructions][3] on using **gcsfuse** but I found the transfer speed was much too slow. (I'm sure there's a way to specify faster transfer speeds, I just don't see how to do that properly.)\n\nI've also tried the Google Console's [transfer page][4] to create a transfer job. Using AWS as the source (`open-images-dataset/validation` for example) I get `Invalid access key. Make sure the access key for your S3 bucket is correct.` error even after creating an IAM for AWS (AmazonS3FullAccess permission). I've also looked at [CVDF's instructions][5] in using Google's Storage transfer via the tsv files, but of course that is the full 18TB dataset and I just want the training, validation, and testing sets (~550GB) for this challenge.\n\nI've seen people suggest downloading the data to their local machines and then uploading the files to a Google Cloud bucket from there but I don't have a local machine that could handle that in regards to storage space and upload speed.\n\nAny help would be greatly appreciated. I'm hoping that people with more experience can point me in some right directions. Maybe there are some other inexperienced teams like me that could benefit from this too.\n\n  [1]: https://www.kaggle.com/c/inclusive-images-challenge/discussion/65157\n  [2]: https://www.kaggle.com/sepeagb2\n  [3]: https://www.kaggle.com/c/inclusive-images-challenge/discussion/66215\n  [4]: https://console.cloud.google.com/storage/transfer\n  [5]: https://github.com/cvdfoundation/open-images-dataset#download-full-dataset-with-google-storage-transfer",
    "394753": "Ya, me too got stuck with this for a week, but here is the solution--\n\nMake a  basic VM  instance on google cloud.Download the files onto that instance .Then upload it to a cloud bucket of your choice.",
    "394758": "You can use\ngsutil -m -o GSUtil:parallel_composite_upload_threshold=150M cp -r bigfilename gs://bucket-name\n to copy files in parallel.",
    "394807": "That makes sense but would that mean you'd have to essentially pay for the storage twice, once for the VM instance and once for the bucket? I figure you could delete the files from the VM after uploading to the bucket; I'm pretty new with using cloud storage and I don't know what costs to expect in this case.\n\nI might have to do it this way if nothing else works. I'll probably drop back into it after work today. Thanks for the advice!",
    "394814": "I've seen using **gsutil** this way but can you replace the `bigfilename` (from your example) with a reference to a ASW S3 bucket? I think this is what someone also [suggested][1] in another discussion. I'll have to explore more and try this out after work.\n\nThanks!\n\n  [1]: https://www.kaggle.com/c/inclusive-images-challenge/discussion/66232#394567",
    "395596": "Hey Victor, did this work for you?",
    "395625": "It did! Well at least much better than what I was able to achieve before. I think I can get better speeds with a different instance.\n\nI wrote a little more in another thread about what I did: https://www.kaggle.com/c/inclusive-images-challenge/discussion/66232#394811\n\nI figured it might help someone else in a similar position as I was in.",
    "399148": "Just a  thought .. Can Google/Kaggle host the total data set on their Google drive or Google cloud bucket and provide means to access directly to work in google cloud.  will that be possible ?<br>\nWondering whether this can help all those who wish to use Google cloud and avoid creating additional account on aws just to download. <br>",
    "399462": "This would really help....looking forward for a reply from the concerned people.",
    "400925": "That would definitely be the best solution.",
    "400926": "I did this as well."
  },
  "source": "meta"
}