{
  "id": 39601,
  "title": "You nhave 72 hours to download",
  "url": "/competitions/cdiscount-image-classification-challenge/discussion/39601",
  "author_name": "",
  "post_date": "2017-09-17T09:26:24.064598900Z",
  "votes": 2,
  "comment_count": 12,
  "views": 0,
  "content": "<p>My http download failed.  Looking at the download link I see that Expires is set to a value approx 72 hours in the future.  That's why my download failed (I'm downloading from home with a (slow) DSL).</p>\n\n<p>Can we either get a torrent download or not have this expire setting?  I'm retrying without expire to see what happens.</p>",
  "messages": [
    {
      "id": "222028",
      "postDate": "09/17/2017 09:26:24",
      "content": "<p>My http download failed.  Looking at the download link I see that Expires is set to a value approx 72 hours in the future.  That's why my download failed (I'm downloading from home with a (slow) DSL).</p>\n\n<p>Can we either get a torrent download or not have this expire setting?  I'm retrying without expire to see what happens.</p>",
      "rawMarkdown": "My http download failed.  Looking at the download link I see that Expires is set to a value approx 72 hours in the future.  That's why my download failed (I'm downloading from home with a (slow) DSL).\n\nCan we either get a torrent download or not have this expire setting?  I'm retrying without expire to see what happens.",
      "votes": null
    },
    {
      "id": "222075",
      "postDate": "09/17/2017 12:03:11",
      "content": "<p>True, such huge data should be provided via torrent download. One of the option I am thinking about is moving to could. </p>\n\n<p>BTW what plans do you have for such a large dataset ? ;) </p>",
      "rawMarkdown": "True, such huge data should be provided via torrent download. One of the option I am thinking about is moving to could. \n\nBTW what plans do you have for such a large dataset ? ;)",
      "votes": null
    },
    {
      "id": "222081",
      "postDate": "09/17/2017 12:31:53",
      "content": "<p>@cpmp - I won't be able to check into a torrent tomorrow. Just to be sure, did you restart the download after the train file was moved to the Google Cloud Storage bucket?</p>",
      "rawMarkdown": "cpmp - I won't be able to check into a torrent tomorrow. Just to be sure, did you restart the download after the train file was moved to the Google Cloud Storage bucket?",
      "votes": null
    },
    {
      "id": "222085",
      "postDate": "09/17/2017 12:46:40",
      "content": "<p>Yes, I restarted download after the move.  I'm now trying with wget.</p>",
      "rawMarkdown": "Yes, I restarted download after the move.  I'm now trying with wget.",
      "votes": null
    },
    {
      "id": "222177",
      "postDate": "09/17/2017 23:34:34",
      "content": "<p>I've had success with downloading 'train.bson' spanning a number of days using the 'curl' command provided by Chrome's developer tools and add arguments like this to the end of it:</p>\n\n<pre><code>-o train.bson -C - --limit-rate 500k\n</code></pre>\n\n<p>This command can be stopped and re-issued whenever you like - it will resume from the correct place.</p>",
      "rawMarkdown": "I've had success with downloading 'train.bson' spanning a number of days using the 'curl' command provided by Chrome's developer tools and add arguments like this to the end of it:\n\n    -o train.bson -C - --limit-rate 500k\n\nThis command can be stopped and re-issued whenever you like - it will resume from the correct place.",
      "votes": null
    },
    {
      "id": "222194",
      "postDate": "09/18/2017 02:20:06",
      "content": "<p>Using 'Internet download Manager' ypu shpuld be able to pause and resume and also resume from last point on failure.</p>",
      "rawMarkdown": "Using 'Internet download Manager' ypu shpuld be able to pause and resume and also resume from last point on failure.",
      "votes": null
    },
    {
      "id": "222266",
      "postDate": "09/18/2017 07:56:15",
      "content": "<p>Could download test in 13 hours.  Will now try train.</p>",
      "rawMarkdown": "Could download test in 13 hours.  Will now try train.",
      "votes": null
    },
    {
      "id": "222270",
      "postDate": "09/18/2017 08:27:56",
      "content": "<p>@cpmp  My education network also has limited speed and the connection turn off every six hours, but I used kaggle-cli to successfully download train data in more than 24 hours.</p>",
      "rawMarkdown": "cpmp  My education network also has limited speed and the connection turn off every six hours, but I used kaggle-cli to successfully download train data in more than 24 hours.",
      "votes": null
    },
    {
      "id": "222272",
      "postDate": "09/18/2017 08:34:55",
      "content": "<p>Thanks.  Ths issue is not the http client, the issue is the Expires parameter set to 72 hours.  I've tried to tamper it but this is not accepted.  I must pause my download when family needs internet, which means it is unlikely I can download within 72 hours (it is scheduled to take 42 hours uninterrupted).  If I use curl, then either the expires is kept the same, and I have 3 days no more, or I reissue a download, and curl will see a different url.  The solution is torrent, not http download.  Or to break train into smaller pieces.</p>",
      "rawMarkdown": "Thanks.  Ths issue is not the http client, the issue is the Expires parameter set to 72 hours.  I've tried to tamper it but this is not accepted.  I must pause my download when family needs internet, which means it is unlikely I can download within 72 hours (it is scheduled to take 42 hours uninterrupted).  If I use curl, then either the expires is kept the same, and I have 3 days no more, or I reissue a download, and curl will see a different url.  The solution is torrent, not http download.  Or to break train into smaller pieces.",
      "votes": null
    },
    {
      "id": "225373",
      "postDate": "09/28/2017 21:20:01",
      "content": "<p>Torrent files for train and test are now available on the data page. Please let me know how they work out for you. Thanks!</p>",
      "rawMarkdown": "Torrent files for train and test are now available on the data page. Please let me know how they work out for you. Thanks!",
      "votes": null
    },
    {
      "id": "225522",
      "postDate": "09/29/2017 11:10:59",
      "content": "<p>Download started, will tell you in few days how it went ;)</p>\n\n<p>Thanks for making this possible.</p>",
      "rawMarkdown": "Download started, will tell you in few days how it went ;)\n\nThanks for making this possible.",
      "votes": null
    },
    {
      "id": "229343",
      "postDate": "10/09/2017 12:26:19",
      "content": "<p>Thanks, download completed this week end.  It took long as I can only download during nights from home as my ISP is used by family during days.  Torrent was the only way to go.</p>\n\n<p>I tried downloading from work but it failed after a couple of hours, before the ETA of 4 hours. I tried several times, don't know why it fails each time.  Torrent is way more robust.</p>",
      "rawMarkdown": "Thanks, download completed this week end.  It took long as I can only download during nights from home as my ISP is used by family during days.  Torrent was the only way to go.\n\nI tried downloading from work but it failed after a couple of hours, before the ETA of 4 hours. I tried several times, don't know why it fails each time.  Torrent is way more robust.",
      "votes": null
    },
    {
      "id": "229500",
      "postDate": "10/09/2017 19:19:27",
      "content": "<p>@CPMP - Thanks for reporting back.</p>",
      "rawMarkdown": "CPMP - Thanks for reporting back.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 222075,
      "author_name": "faamin",
      "author_url": "",
      "post_date": "09/17/2017 12:03:11",
      "content": "<p>True, such huge data should be provided via torrent download. One of the option I am thinking about is moving to could. </p>\n\n<p>BTW what plans do you have for such a large dataset ? ;) </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 222081,
      "author_name": "inversion",
      "author_url": "",
      "post_date": "09/17/2017 12:31:53",
      "content": "<p>@cpmp - I won't be able to check into a torrent tomorrow. Just to be sure, did you restart the download after the train file was moved to the Google Cloud Storage bucket?</p>",
      "votes": null,
      "replies": [
        {
          "id": 222085,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "09/17/2017 12:46:40",
          "content": "<p>Yes, I restarted download after the move.  I'm now trying with wget.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 222266,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "09/18/2017 07:56:15",
          "content": "<p>Could download test in 13 hours.  Will now try train.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 222177,
      "author_name": "stainsby",
      "author_url": "",
      "post_date": "09/17/2017 23:34:34",
      "content": "<p>I've had success with downloading 'train.bson' spanning a number of days using the 'curl' command provided by Chrome's developer tools and add arguments like this to the end of it:</p>\n\n<pre><code>-o train.bson -C - --limit-rate 500k\n</code></pre>\n\n<p>This command can be stopped and re-issued whenever you like - it will resume from the correct place.</p>",
      "votes": null,
      "replies": [
        {
          "id": 222272,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "09/18/2017 08:34:55",
          "content": "<p>Thanks.  Ths issue is not the http client, the issue is the Expires parameter set to 72 hours.  I've tried to tamper it but this is not accepted.  I must pause my download when family needs internet, which means it is unlikely I can download within 72 hours (it is scheduled to take 42 hours uninterrupted).  If I use curl, then either the expires is kept the same, and I have 3 days no more, or I reissue a download, and curl will see a different url.  The solution is torrent, not http download.  Or to break train into smaller pieces.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 222194,
      "author_name": "restlessrational",
      "author_url": "",
      "post_date": "09/18/2017 02:20:06",
      "content": "<p>Using 'Internet download Manager' ypu shpuld be able to pause and resume and also resume from last point on failure.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 222270,
      "author_name": "zonemercy",
      "author_url": "",
      "post_date": "09/18/2017 08:27:56",
      "content": "<p>@cpmp  My education network also has limited speed and the connection turn off every six hours, but I used kaggle-cli to successfully download train data in more than 24 hours.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 225373,
      "author_name": "inversion",
      "author_url": "",
      "post_date": "09/28/2017 21:20:01",
      "content": "<p>Torrent files for train and test are now available on the data page. Please let me know how they work out for you. Thanks!</p>",
      "votes": null,
      "replies": [
        {
          "id": 225522,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "09/29/2017 11:10:59",
          "content": "<p>Download started, will tell you in few days how it went ;)</p>\n\n<p>Thanks for making this possible.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 229343,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "10/09/2017 12:26:19",
          "content": "<p>Thanks, download completed this week end.  It took long as I can only download during nights from home as my ISP is used by family during days.  Torrent was the only way to go.</p>\n\n<p>I tried downloading from work but it failed after a couple of hours, before the ETA of 4 hours. I tried several times, don't know why it fails each time.  Torrent is way more robust.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 229500,
          "author_name": "inversion",
          "author_url": "",
          "post_date": "10/09/2017 19:19:27",
          "content": "<p>@CPMP - Thanks for reporting back.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "222028": "My http download failed.  Looking at the download link I see that Expires is set to a value approx 72 hours in the future.  That's why my download failed (I'm downloading from home with a (slow) DSL).\n\nCan we either get a torrent download or not have this expire setting?  I'm retrying without expire to see what happens.",
    "222075": "True, such huge data should be provided via torrent download. One of the option I am thinking about is moving to could. \n\nBTW what plans do you have for such a large dataset ? ;)",
    "222081": "cpmp - I won't be able to check into a torrent tomorrow. Just to be sure, did you restart the download after the train file was moved to the Google Cloud Storage bucket?",
    "222085": "Yes, I restarted download after the move.  I'm now trying with wget.",
    "222177": "I've had success with downloading 'train.bson' spanning a number of days using the 'curl' command provided by Chrome's developer tools and add arguments like this to the end of it:\n\n    -o train.bson -C - --limit-rate 500k\n\nThis command can be stopped and re-issued whenever you like - it will resume from the correct place.",
    "222194": "Using 'Internet download Manager' ypu shpuld be able to pause and resume and also resume from last point on failure.",
    "222266": "Could download test in 13 hours.  Will now try train.",
    "222270": "cpmp  My education network also has limited speed and the connection turn off every six hours, but I used kaggle-cli to successfully download train data in more than 24 hours.",
    "222272": "Thanks.  Ths issue is not the http client, the issue is the Expires parameter set to 72 hours.  I've tried to tamper it but this is not accepted.  I must pause my download when family needs internet, which means it is unlikely I can download within 72 hours (it is scheduled to take 42 hours uninterrupted).  If I use curl, then either the expires is kept the same, and I have 3 days no more, or I reissue a download, and curl will see a different url.  The solution is torrent, not http download.  Or to break train into smaller pieces.",
    "225373": "Torrent files for train and test are now available on the data page. Please let me know how they work out for you. Thanks!",
    "225522": "Download started, will tell you in few days how it went ;)\n\nThanks for making this possible.",
    "229343": "Thanks, download completed this week end.  It took long as I can only download during nights from home as my ISP is used by family during days.  Torrent was the only way to go.\n\nI tried downloading from work but it failed after a couple of hours, before the ETA of 4 hours. I tried several times, don't know why it fails each time.  Torrent is way more robust.",
    "229500": "CPMP - Thanks for reporting back."
  },
  "source": "meta"
}