{
  "id": 61944,
  "title": "Downloading train set using aws s3",
  "url": "/competitions/google-ai-open-images-object-detection-track/discussion/61944",
  "author_name": "",
  "post_date": "2018-07-25T11:35:37.445009800Z",
  "votes": null,
  "comment_count": 7,
  "views": 0,
  "content": "<p>I downloaded 420.2 GB of train set for about 30 hrs using \"aws s3 --no-sign-request sync s3://open-images-dataset/train [target_dir/train] \" command from CVDF , but it hit \"fatal error: ('The read operation timed out',)\" and stopped downloading.</p>\n\n<p>Method 1 : Is there a way to continue downloading from 420.2 GB mark? \nMethod 2: Should I download last ~100 GB from figure eight (train_07.zip and train_08.zip)?\nMethod 3: Try to download whole 512 GB again through CVDF....?! </p>\n\n<p>Method 2 assumes images downloaded from CVDF are in same order as how files are split in 9 zip files from figure eight. Any help would be greatly appreciated!</p>",
  "messages": [
    {
      "id": "361961",
      "postDate": "07/25/2018 11:35:37",
      "content": "<p>I downloaded 420.2 GB of train set for about 30 hrs using \"aws s3 --no-sign-request sync s3://open-images-dataset/train [target_dir/train] \" command from CVDF , but it hit \"fatal error: ('The read operation timed out',)\" and stopped downloading.</p>\n\n<p>Method 1 : Is there a way to continue downloading from 420.2 GB mark? \nMethod 2: Should I download last ~100 GB from figure eight (train_07.zip and train_08.zip)?\nMethod 3: Try to download whole 512 GB again through CVDF....?! </p>\n\n<p>Method 2 assumes images downloaded from CVDF are in same order as how files are split in 9 zip files from figure eight. Any help would be greatly appreciated!</p>",
      "rawMarkdown": "I downloaded 420.2 GB of train set for about 30 hrs using \"aws s3 --no-sign-request sync s3://open-images-dataset/train [target_dir/train] \" command from CVDF , but it hit \"fatal error: ('The read operation timed out',)\" and stopped downloading.\n\nMethod 1 : Is there a way to continue downloading from 420.2 GB mark? \nMethod 2: Should I download last ~100 GB from figure eight (train_07.zip and train_08.zip)?\nMethod 3: Try to download whole 512 GB again through CVDF....?! \n\nMethod 2 assumes images downloaded from CVDF are in same order as how files are split in 9 zip files from figure eight. Any help would be greatly appreciated!",
      "votes": null
    },
    {
      "id": "361983",
      "postDate": "07/25/2018 12:33:32",
      "content": "<p>If you run the sync again, it'll continue downloading from where you left off (and also download any missing / error ones from earlier sequence).</p>\n\n<p>You'll need to wait awhile for the sync to scan through your existing files before it resumes downloading the newest ones.</p>\n\n<p>The fatal error occurs very frequently for me. What I did was create a script with multiple of the same aws sync command, so it auto resumes every time it is cut off.</p>\n\n<p>I tried downloading the 9 zip files on figure eight too, but it is extremely slow (like 200kb+/s), so I resorted back to aws individual file syncing.</p>",
      "rawMarkdown": "If you run the sync again, it'll continue downloading from where you left off (and also download any missing / error ones from earlier sequence).\n\nYou'll need to wait awhile for the sync to scan through your existing files before it resumes downloading the newest ones.\n\nThe fatal error occurs very frequently for me. What I did was create a script with multiple of the same aws sync command, so it auto resumes every time it is cut off.\n\nI tried downloading the 9 zip files on figure eight too, but it is extremely slow (like 200kb+/s), so I resorted back to aws individual file syncing.",
      "votes": null
    },
    {
      "id": "361985",
      "postDate": "07/25/2018 12:44:35",
      "content": "<p>Great!! I didn't know it had such a feature implemented! I actually started looking into \"--include\" and \"--exclude\" from AWS Command Line Interface, but seems like no need to do that!</p>\n\n<p>Thank you Mongrel with sharing your knowledge on this matter and lifesaver tip. Figure eight is very very slow for me as well, and wanted to avoid it at all cost! </p>",
      "rawMarkdown": "Great!! I didn't know it had such a feature implemented! I actually started looking into \"--include\" and \"--exclude\" from AWS Command Line Interface, but seems like no need to do that!\n\nThank you Mongrel with sharing your knowledge on this matter and lifesaver tip. Figure eight is very very slow for me as well, and wanted to avoid it at all cost!",
      "votes": null
    },
    {
      "id": "361986",
      "postDate": "07/25/2018 12:49:17",
      "content": "<p>Welcome! I've been cracking my head on syncing the files for the past many days too, and reported some related AWS syncing errors to the competition host.</p>\n\n<p>Since you are at 400+Gb, I'm sure you'll finish the remaining 20% in no time.</p>\n\n<p>I suggest running the sync a couple more times even after you finish though, as I found that there are still a few missing files here and there due to sync error.</p>\n\n<p>Good luck! :)</p>",
      "rawMarkdown": "Welcome! I've been cracking my head on syncing the files for the past many days too, and reported some related AWS syncing errors to the competition host.\n\nSince you are at 400+Gb, I'm sure you'll finish the remaining 20% in no time.\n\nI suggest running the sync a couple more times even after you finish though, as I found that there are still a few missing files here and there due to sync error.\n\nGood luck! :)",
      "votes": null
    },
    {
      "id": "362206",
      "postDate": "07/25/2018 22:57:27",
      "content": "<p>I realized that it takes a very long time for AWS CLI to go through already downloaded 420.2 GB to figure out where to start downloading again. I wish there was a bucket with same file, but sorted backward so that you can start downloading backward in case download was interrupted toward the end like me.</p>",
      "rawMarkdown": "I realized that it takes a very long time for AWS CLI to go through already downloaded 420.2 GB to figure out where to start downloading again. I wish there was a bucket with same file, but sorted backward so that you can start downloading backward in case download was interrupted toward the end like me.",
      "votes": null
    },
    {
      "id": "362260",
      "postDate": "07/26/2018 02:52:24",
      "content": "<p>Yeah, that would be nice. I think part of the long delay is due to the sync checking and comparing the local files to see which are valid / successfully downloaded.</p>",
      "rawMarkdown": "Yeah, that would be nice. I think part of the long delay is due to the sync checking and comparing the local files to see which are valid / successfully downloaded.",
      "votes": null
    },
    {
      "id": "362307",
      "postDate": "07/26/2018 06:22:29",
      "content": "<p>Thanks Mongrel for stepping in.\nI hope you all finally managed to download the images.</p>",
      "rawMarkdown": "Thanks Mongrel for stepping in.\nI hope you all finally managed to download the images.",
      "votes": null
    },
    {
      "id": "362627",
      "postDate": "07/26/2018 19:01:17",
      "content": "<p>Thanks for checking in, I was able to download everything successfully.</p>",
      "rawMarkdown": "Thanks for checking in, I was able to download everything successfully.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 361983,
      "author_name": "andykoh",
      "author_url": "",
      "post_date": "07/25/2018 12:33:32",
      "content": "<p>If you run the sync again, it'll continue downloading from where you left off (and also download any missing / error ones from earlier sequence).</p>\n\n<p>You'll need to wait awhile for the sync to scan through your existing files before it resumes downloading the newest ones.</p>\n\n<p>The fatal error occurs very frequently for me. What I did was create a script with multiple of the same aws sync command, so it auto resumes every time it is cut off.</p>\n\n<p>I tried downloading the 9 zip files on figure eight too, but it is extremely slow (like 200kb+/s), so I resorted back to aws individual file syncing.</p>",
      "votes": null,
      "replies": [
        {
          "id": 361985,
          "author_name": "silvernine",
          "author_url": "",
          "post_date": "07/25/2018 12:44:35",
          "content": "<p>Great!! I didn't know it had such a feature implemented! I actually started looking into \"--include\" and \"--exclude\" from AWS Command Line Interface, but seems like no need to do that!</p>\n\n<p>Thank you Mongrel with sharing your knowledge on this matter and lifesaver tip. Figure eight is very very slow for me as well, and wanted to avoid it at all cost! </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 361986,
          "author_name": "andykoh",
          "author_url": "",
          "post_date": "07/25/2018 12:49:17",
          "content": "<p>Welcome! I've been cracking my head on syncing the files for the past many days too, and reported some related AWS syncing errors to the competition host.</p>\n\n<p>Since you are at 400+Gb, I'm sure you'll finish the remaining 20% in no time.</p>\n\n<p>I suggest running the sync a couple more times even after you finish though, as I found that there are still a few missing files here and there due to sync error.</p>\n\n<p>Good luck! :)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 362206,
          "author_name": "silvernine",
          "author_url": "",
          "post_date": "07/25/2018 22:57:27",
          "content": "<p>I realized that it takes a very long time for AWS CLI to go through already downloaded 420.2 GB to figure out where to start downloading again. I wish there was a bucket with same file, but sorted backward so that you can start downloading backward in case download was interrupted toward the end like me.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 362260,
          "author_name": "andykoh",
          "author_url": "",
          "post_date": "07/26/2018 02:52:24",
          "content": "<p>Yeah, that would be nice. I think part of the long delay is due to the sync checking and comparing the local files to see which are valid / successfully downloaded.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 362307,
          "author_name": "jponttusset",
          "author_url": "",
          "post_date": "07/26/2018 06:22:29",
          "content": "<p>Thanks Mongrel for stepping in.\nI hope you all finally managed to download the images.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 362627,
          "author_name": "silvernine",
          "author_url": "",
          "post_date": "07/26/2018 19:01:17",
          "content": "<p>Thanks for checking in, I was able to download everything successfully.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "361961": "I downloaded 420.2 GB of train set for about 30 hrs using \"aws s3 --no-sign-request sync s3://open-images-dataset/train [target_dir/train] \" command from CVDF , but it hit \"fatal error: ('The read operation timed out',)\" and stopped downloading.\n\nMethod 1 : Is there a way to continue downloading from 420.2 GB mark? \nMethod 2: Should I download last ~100 GB from figure eight (train_07.zip and train_08.zip)?\nMethod 3: Try to download whole 512 GB again through CVDF....?! \n\nMethod 2 assumes images downloaded from CVDF are in same order as how files are split in 9 zip files from figure eight. Any help would be greatly appreciated!",
    "361983": "If you run the sync again, it'll continue downloading from where you left off (and also download any missing / error ones from earlier sequence).\n\nYou'll need to wait awhile for the sync to scan through your existing files before it resumes downloading the newest ones.\n\nThe fatal error occurs very frequently for me. What I did was create a script with multiple of the same aws sync command, so it auto resumes every time it is cut off.\n\nI tried downloading the 9 zip files on figure eight too, but it is extremely slow (like 200kb+/s), so I resorted back to aws individual file syncing.",
    "361985": "Great!! I didn't know it had such a feature implemented! I actually started looking into \"--include\" and \"--exclude\" from AWS Command Line Interface, but seems like no need to do that!\n\nThank you Mongrel with sharing your knowledge on this matter and lifesaver tip. Figure eight is very very slow for me as well, and wanted to avoid it at all cost!",
    "361986": "Welcome! I've been cracking my head on syncing the files for the past many days too, and reported some related AWS syncing errors to the competition host.\n\nSince you are at 400+Gb, I'm sure you'll finish the remaining 20% in no time.\n\nI suggest running the sync a couple more times even after you finish though, as I found that there are still a few missing files here and there due to sync error.\n\nGood luck! :)",
    "362206": "I realized that it takes a very long time for AWS CLI to go through already downloaded 420.2 GB to figure out where to start downloading again. I wish there was a bucket with same file, but sorted backward so that you can start downloading backward in case download was interrupted toward the end like me.",
    "362260": "Yeah, that would be nice. I think part of the long delay is due to the sync checking and comparing the local files to see which are valid / successfully downloaded.",
    "362307": "Thanks Mongrel for stepping in.\nI hope you all finally managed to download the images.",
    "362627": "Thanks for checking in, I was able to download everything successfully."
  },
  "source": "meta"
}