{
  "id": 2441,
  "title": "How to download 6GB file with guaranteed interruption?",
  "url": "/competitions/predict-closed-questions-on-stack-overflow/discussion/2441",
  "author_name": "",
  "post_date": "2012-08-25T16:38:22.050Z",
  "votes": null,
  "comment_count": 6,
  "views": 4438,
  "content": "<p>Hello,</p>\r\n<p>&nbsp;</p>\r\n<p>&nbsp; Embarassing question, but here we go ;-) how to download that 6GB file? My problem is, I am not able to download it in one pass (it is technical issue, and there is nothing I can do about it) -- and since server adds countermeasures against free download\r\n all my standard methods so far failed.</p>\r\n<p>&nbsp;</p>\r\n<p>&nbsp; I will be grateful for hints (my OS is Linux).</p>\r\n<p>&nbsp;</p>\r\n<p>Kind regards,</p>",
  "messages": [
    {
      "id": "13449",
      "postDate": "08/25/2012 16:38:22",
      "content": "<p>Hello,</p>\r\n<p>&nbsp;</p>\r\n<p>&nbsp; Embarassing question, but here we go ;-) how to download that 6GB file? My problem is, I am not able to download it in one pass (it is technical issue, and there is nothing I can do about it) -- and since server adds countermeasures against free download\r\n all my standard methods so far failed.</p>\r\n<p>&nbsp;</p>\r\n<p>&nbsp; I will be grateful for hints (my OS is Linux).</p>\r\n<p>&nbsp;</p>\r\n<p>Kind regards,</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "13451",
      "postDate": "08/25/2012 16:55:45",
      "content": "<p>Maybe use curl and add a custom Range: header (along with the cookies from the site) to download chunks at a time, and then cat them together later?</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "13452",
      "postDate": "08/25/2012 17:01:38",
      "content": "<p>Thank you, so far I tried wget, aria2c and curl. All of them fail, the server adds a timestamp to a filename (dynamically), on reconnection, and it is very effective in misleading those tools. I.e. on each interruption, given tool sees (for it) different\r\n file and downloading begins from 0.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "13456",
      "postDate": "08/25/2012 18:24:13",
      "content": "<p>The 6GB file is our standard XML datadump, which is also available as a <a href=\"http://www.clearbits.net/torrents/2076-aug-2012\">\r\ntorrent here</a>.</p>\r\n<p>The torrent version includes all graduated Stack Exchange sites (not just Stack Overflow), and splits Stack Overflow into 700mb chunks (7zip will spit it back out as one file after decompressing) for technical reasons but otherwise it's basically the same\r\n data.</p>\r\n<p>If you use a torrent client that lets you pick and choose files to download you can grab just Stack Overflow, and you should be able to deal gracefully with interruptions.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "13459",
      "postDate": "08/25/2012 18:51:03",
      "content": "<p>Kevin, thank you very much. I see some movement already, so I hope I will wake up tomorrow with full set (of SO). Thanks once again.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "13487",
      "postDate": "08/26/2012 23:41:49",
      "content": "<p>Hi, I faced the same problem. Here's how I managed to get aria2 to work.<br>\r\nhttp://yehzheng.blogspot.sg/2012/08/download-using-aria2-with-link.html?m=1</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "13772",
      "postDate": "09/01/2012 19:25:25",
      "content": "<p>For the record, this is piece of cake, the problem is with interruption. Kaggle adds new timestamp each time, so if you have just a single interruption, your download will resume from the start. And since (because of my bandwidth) I have at least one interruption,\r\n aria2c is not capable to deal with it. And since I am already posting this, wget is -- here is how (SE, surprise, surprise :-D):\r\n<br>\r\nhttp://unix.stackexchange.com/questions/46317/download-tool-needed-with-custom-headers-resume-retry-custom-filename-outp</p>\r\n<p>Good to know, I downloaded all the data using torrent as Kevin described.</p>",
      "rawMarkdown": "",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 13451,
      "author_name": "andysloane",
      "author_url": "",
      "post_date": "08/25/2012 16:55:45",
      "content": "<p>Maybe use curl and add a custom Range: header (along with the cookies from the site) to download chunks at a time, and then cat them together later?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 13452,
      "author_name": "macias",
      "author_url": "",
      "post_date": "08/25/2012 17:01:38",
      "content": "<p>Thank you, so far I tried wget, aria2c and curl. All of them fail, the server adds a timestamp to a filename (dynamically), on reconnection, and it is very effective in misleading those tools. I.e. on each interruption, given tool sees (for it) different\r\n file and downloading begins from 0.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 13456,
      "author_name": "kevinmontrose",
      "author_url": "",
      "post_date": "08/25/2012 18:24:13",
      "content": "<p>The 6GB file is our standard XML datadump, which is also available as a <a href=\"http://www.clearbits.net/torrents/2076-aug-2012\">\r\ntorrent here</a>.</p>\r\n<p>The torrent version includes all graduated Stack Exchange sites (not just Stack Overflow), and splits Stack Overflow into 700mb chunks (7zip will spit it back out as one file after decompressing) for technical reasons but otherwise it's basically the same\r\n data.</p>\r\n<p>If you use a torrent client that lets you pick and choose files to download you can grab just Stack Overflow, and you should be able to deal gracefully with interruptions.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 13459,
      "author_name": "macias",
      "author_url": "",
      "post_date": "08/25/2012 18:51:03",
      "content": "<p>Kevin, thank you very much. I see some movement already, so I hope I will wake up tomorrow with full set (of SO). Thanks once again.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 13487,
      "author_name": "yehzhengtan",
      "author_url": "",
      "post_date": "08/26/2012 23:41:49",
      "content": "<p>Hi, I faced the same problem. Here's how I managed to get aria2 to work.<br>\r\nhttp://yehzheng.blogspot.sg/2012/08/download-using-aria2-with-link.html?m=1</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 13772,
      "author_name": "macias",
      "author_url": "",
      "post_date": "09/01/2012 19:25:25",
      "content": "<p>For the record, this is piece of cake, the problem is with interruption. Kaggle adds new timestamp each time, so if you have just a single interruption, your download will resume from the start. And since (because of my bandwidth) I have at least one interruption,\r\n aria2c is not capable to deal with it. And since I am already posting this, wget is -- here is how (SE, surprise, surprise :-D):\r\n<br>\r\nhttp://unix.stackexchange.com/questions/46317/download-tool-needed-with-custom-headers-resume-retry-custom-filename-outp</p>\r\n<p>Good to know, I downloaded all the data using torrent as Kevin described.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "13449": "",
    "13451": "",
    "13452": "",
    "13456": "",
    "13459": "",
    "13487": "",
    "13772": ""
  },
  "source": "meta"
}