{
  "id": 24137,
  "title": "Is it possible to upload the file page_views.csv.zip to a torrent service?",
  "url": "/competitions/outbrain-click-prediction/discussion/24137",
  "author_name": "",
  "post_date": "2016-10-06T11:04:20.563Z",
  "votes": null,
  "comment_count": 6,
  "views": 900,
  "content": "<p>I want to download file page_views.csv.zip, but my connection fails at some point. Is it possible to upload the file in parts or upload this file to a torrent service?</p>",
  "messages": [
    {
      "id": "137973",
      "postDate": "10/06/2016 11:04:20",
      "content": "<p>I want to download file page_views.csv.zip, but my connection fails at some point. Is it possible to upload the file in parts or upload this file to a torrent service?</p>",
      "rawMarkdown": "I want to download file page_views.csv.zip, but my connection fails at some point. Is it possible to upload the file in parts or upload this file to a torrent service?",
      "votes": null
    },
    {
      "id": "138033",
      "postDate": "10/06/2016 18:29:09",
      "content": "<p>Hi, I'm afraid torrent is not going to be an option. From the &quot;<a href=\"https://www.kaggle.com/c/outbrain-click-prediction/forums/t/24107/welcome\">Welcome</a>&quot; post you can read:</p>\n\n<p>&quot;Expect a traffic jam to download the 30 GB file during the launch gold rush (sorry!). Use a download manager with resuming capabilities and we promise you'll eventually get the file. There is also a small sample available. Please read our <a href=\"https://www.kaggle.com/wiki/ANoteOnTorrents\">note</a> on torrents before you suggest using torrents.&quot;</p>\n\n<p>Cheers!</p>",
      "rawMarkdown": "Hi, I'm afraid torrent is not going to be an option. From the \"[Welcome][1]\" post you can read:\r\n\r\n\"Expect a traffic jam to download the 30 GB file during the launch gold rush (sorry!). Use a download manager with resuming capabilities and we promise you'll eventually get the file. There is also a small sample available. Please read our [note][2] on torrents before you suggest using torrents.\"\r\n\r\nCheers!\r\n\r\n\r\n  [1]: https://www.kaggle.com/c/outbrain-click-prediction/forums/t/24107/welcome\r\n  [2]: https://www.kaggle.com/wiki/ANoteOnTorrents",
      "votes": null
    },
    {
      "id": "138185",
      "postDate": "10/07/2016 15:33:16",
      "content": "<p>I'm pretty sure one could torrent an encrypted version of the data and then force users to download the key directly from kaggle. Versioning could be handled in this way too. Given the failure rate, it will take me probably a week to d/l this file. </p>\n\n<p>Maybe we should make a &quot;kaggle&quot; a verb for something that's impossible to d/l, a la &quot;I tried to download the new iOS on launch but download system is totally kaggled&quot;.</p>",
      "rawMarkdown": "I'm pretty sure one could torrent an encrypted version of the data and then force users to download the key directly from kaggle. Versioning could be handled in this way too. Given the failure rate, it will take me probably a week to d/l this file. \r\n\r\nMaybe we should make a \"kaggle\" a verb for something that's impossible to d/l, a la \"I tried to download the new iOS on launch but download system is totally kaggled\".",
      "votes": null
    },
    {
      "id": "138247",
      "postDate": "10/07/2016 18:50:43",
      "content": "<p>Well, I've been able to download it via <code>curl</code>:</p>\n\n<ol>\n<li>Download some of the light files so in your browser you have to accept the data agreement and this action is stored in your <code>cookies</code></li>\n<li>In your browser (in my case it is Chrome) right click and open the Inspector</li>\n<li>Click on the super heavy file and in the Inspector check the petition in the Network tab</li>\n<li>Right click on the network petition and <code>copy as cURL</code></li>\n<li>Finally you can open the terminal and paste the <code>curl</code> command</li>\n</ol>\n\n<p>I was able to download it at once, with a pretty shitty connexion. But the nice thing about curl is that you can resume the download by setting an offset:</p>\n\n<pre><code>curl -C offset url\n</code></pre>\n\n<p>So if it has stopped, you can check for the size of the current file, and set it (+1 I think) as the offset parameter.</p>\n\n<p>Hope this is helpful to somebody, I know it shouldn't be that elaborated, but it's feasible ;)</p>",
      "rawMarkdown": "Well, I've been able to download it via `curl`:\r\n\r\n 1. Download some of the light files so in your browser you have to accept the data agreement and this action is stored in your `cookies`\r\n 2. In your browser (in my case it is Chrome) right click and open the Inspector\r\n 3. Click on the super heavy file and in the Inspector check the petition in the Network tab\r\n 4. Right click on the network petition and `copy as cURL`\r\n 5. Finally you can open the terminal and paste the `curl` command\r\n\r\nI was able to download it at once, with a pretty shitty connexion. But the nice thing about curl is that you can resume the download by setting an offset:\r\n\r\n    curl -C offset url\r\n\r\nSo if it has stopped, you can check for the size of the current file, and set it (+1 I think) as the offset parameter.\r\n\r\nHope this is helpful to somebody, I know it shouldn't be that elaborated, but it's feasible ;)",
      "votes": null
    },
    {
      "id": "138585",
      "postDate": "10/09/2016 17:00:59",
      "content": "<p>Most likely, that's against the contest rules. But just as a test, no, it's not possible. After several days of retrying:</p>\n\n<p>&quot;\nwget '<a href=\"https://kaggle2.blob.core.windows.net/competitions-data/kaggle/5497/page_views.csv.zip\">https://kaggle2.blob.core.windows.net/competitions-data/kaggle/5497/page_views.csv.zip</a>? ...'</p>\n\n<p>Resolving kaggle2.blob.core.windows.net (kaggle2.blob.core.windows.net)... 40.116.120.24\nConnecting to kaggle2.blob.core.windows.net (kaggle2.blob.core.windows.net)|40.116.120.24|:443... connected.\nHTTP request sent, awaiting response... 200 OK\nLength: 31905475875 (30G) [application/zip]\nSaving to: 'page_views.csv.zip... '\npage_views.csv.zip@sv=2012-02-12&amp;se=201  35%[==========================&gt;                                                  ]  10.45G   277KB/s    in 14h 12m</p>\n\n\n\n<p>2016-10-08 07:46:50 (214 KB/s) - Read error at byte 11221434368/31905475875 (Resource temporarily unavailable, try again.). Retrying.</p>\n\n<p>HTTP request sent, awaiting response... 403 Server failed to authenticate the request. Make sure the value of Authorization header is formed correctly including the signature.\n2016-10-09 09:22:35 ERROR 403: Server failed to authenticate the request. Make sure the value of Authorization header is formed correctly including the signature..\n&quot;</p>",
      "rawMarkdown": "Most likely, that's against the contest rules. But just as a test, no, it's not possible. After several days of retrying:\r\n\r\n\"\r\nwget 'https://kaggle2.blob.core.windows.net/competitions-data/kaggle/5497/page_views.csv.zip? ...'<snip>\r\n\r\nResolving kaggle2.blob.core.windows.net (kaggle2.blob.core.windows.net)... 40.116.120.24\r\nConnecting to kaggle2.blob.core.windows.net (kaggle2.blob.core.windows.net)|40.116.120.24|:443... connected.\r\nHTTP request sent, awaiting response... 200 OK\r\nLength: 31905475875 (30G) [application/zip]\r\nSaving to: 'page_views.csv.zip... <snip>'\r\npage_views.csv.zip@sv=2012-02-12&se=201  35%[==========================>                                                  ]  10.45G   277KB/s    in 14h 12m\r\n\r\n<snip several download/error/download more etc>\r\n\r\n2016-10-08 07:46:50 (214 KB/s) - Read error at byte 11221434368/31905475875 (Resource temporarily unavailable, try again.). Retrying.\r\n\r\nHTTP request sent, awaiting response... 403 Server failed to authenticate the request. Make sure the value of Authorization header is formed correctly including the signature.\r\n2016-10-09 09:22:35 ERROR 403: Server failed to authenticate the request. Make sure the value of Authorization header is formed correctly including the signature..\r\n\"",
      "votes": null
    },
    {
      "id": "138648",
      "postDate": "10/10/2016 05:13:12",
      "content": "<p>Firefox allows restarting if the download fails.  I found I couldn't use a command-line utility because I needed to have an authenticated browser session, and I couldn't work out how to do that.</p>\n\n<p>I've been able to resume the download four times so far...</p>",
      "rawMarkdown": "Firefox allows restarting if the download fails.  I found I couldn't use a command-line utility because I needed to have an authenticated browser session, and I couldn't work out how to do that.\r\n\r\nI've been able to resume the download four times so far...",
      "votes": null
    },
    {
      "id": "138841",
      "postDate": "10/11/2016 06:35:27",
      "content": "<p>Chrome has the same. However it seems the azure-hosted files have a token that times out after a few days? So even if you keep restarting, you have to get it in that period or the token (on their end) will timeout, apparently. Retrying therefore may not work in many circumstances. </p>\n\n<p>Also the dl speed is limited to about 200k/s (observed) which further increases the likelihood of failure.</p>\n\n<p>I finally d/l'ed the file with some wget-foo.</p>",
      "rawMarkdown": "Chrome has the same. However it seems the azure-hosted files have a token that times out after a few days? So even if you keep restarting, you have to get it in that period or the token (on their end) will timeout, apparently. Retrying therefore may not work in many circumstances. \r\n\r\nAlso the dl speed is limited to about 200k/s (observed) which further increases the likelihood of failure.\r\n\r\nI finally d/l'ed the file with some wget-foo.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 138033,
      "author_name": "guiemb",
      "author_url": "",
      "post_date": "10/06/2016 18:29:09",
      "content": "<p>Hi, I'm afraid torrent is not going to be an option. From the &quot;<a href=\"https://www.kaggle.com/c/outbrain-click-prediction/forums/t/24107/welcome\">Welcome</a>&quot; post you can read:</p>\n\n<p>&quot;Expect a traffic jam to download the 30 GB file during the launch gold rush (sorry!). Use a download manager with resuming capabilities and we promise you'll eventually get the file. There is also a small sample available. Please read our <a href=\"https://www.kaggle.com/wiki/ANoteOnTorrents\">note</a> on torrents before you suggest using torrents.&quot;</p>\n\n<p>Cheers!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 138185,
      "author_name": "rgwood",
      "author_url": "",
      "post_date": "10/07/2016 15:33:16",
      "content": "<p>I'm pretty sure one could torrent an encrypted version of the data and then force users to download the key directly from kaggle. Versioning could be handled in this way too. Given the failure rate, it will take me probably a week to d/l this file. </p>\n\n<p>Maybe we should make a &quot;kaggle&quot; a verb for something that's impossible to d/l, a la &quot;I tried to download the new iOS on launch but download system is totally kaggled&quot;.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 138247,
      "author_name": "guiemb",
      "author_url": "",
      "post_date": "10/07/2016 18:50:43",
      "content": "<p>Well, I've been able to download it via <code>curl</code>:</p>\n\n<ol>\n<li>Download some of the light files so in your browser you have to accept the data agreement and this action is stored in your <code>cookies</code></li>\n<li>In your browser (in my case it is Chrome) right click and open the Inspector</li>\n<li>Click on the super heavy file and in the Inspector check the petition in the Network tab</li>\n<li>Right click on the network petition and <code>copy as cURL</code></li>\n<li>Finally you can open the terminal and paste the <code>curl</code> command</li>\n</ol>\n\n<p>I was able to download it at once, with a pretty shitty connexion. But the nice thing about curl is that you can resume the download by setting an offset:</p>\n\n<pre><code>curl -C offset url\n</code></pre>\n\n<p>So if it has stopped, you can check for the size of the current file, and set it (+1 I think) as the offset parameter.</p>\n\n<p>Hope this is helpful to somebody, I know it shouldn't be that elaborated, but it's feasible ;)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 138585,
      "author_name": "rgwood",
      "author_url": "",
      "post_date": "10/09/2016 17:00:59",
      "content": "<p>Most likely, that's against the contest rules. But just as a test, no, it's not possible. After several days of retrying:</p>\n\n<p>&quot;\nwget '<a href=\"https://kaggle2.blob.core.windows.net/competitions-data/kaggle/5497/page_views.csv.zip\">https://kaggle2.blob.core.windows.net/competitions-data/kaggle/5497/page_views.csv.zip</a>? ...'</p>\n\n<p>Resolving kaggle2.blob.core.windows.net (kaggle2.blob.core.windows.net)... 40.116.120.24\nConnecting to kaggle2.blob.core.windows.net (kaggle2.blob.core.windows.net)|40.116.120.24|:443... connected.\nHTTP request sent, awaiting response... 200 OK\nLength: 31905475875 (30G) [application/zip]\nSaving to: 'page_views.csv.zip... '\npage_views.csv.zip@sv=2012-02-12&amp;se=201  35%[==========================&gt;                                                  ]  10.45G   277KB/s    in 14h 12m</p>\n\n\n\n<p>2016-10-08 07:46:50 (214 KB/s) - Read error at byte 11221434368/31905475875 (Resource temporarily unavailable, try again.). Retrying.</p>\n\n<p>HTTP request sent, awaiting response... 403 Server failed to authenticate the request. Make sure the value of Authorization header is formed correctly including the signature.\n2016-10-09 09:22:35 ERROR 403: Server failed to authenticate the request. Make sure the value of Authorization header is formed correctly including the signature..\n&quot;</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 138648,
      "author_name": "terribletadpole",
      "author_url": "",
      "post_date": "10/10/2016 05:13:12",
      "content": "<p>Firefox allows restarting if the download fails.  I found I couldn't use a command-line utility because I needed to have an authenticated browser session, and I couldn't work out how to do that.</p>\n\n<p>I've been able to resume the download four times so far...</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 138841,
      "author_name": "rgwood",
      "author_url": "",
      "post_date": "10/11/2016 06:35:27",
      "content": "<p>Chrome has the same. However it seems the azure-hosted files have a token that times out after a few days? So even if you keep restarting, you have to get it in that period or the token (on their end) will timeout, apparently. Retrying therefore may not work in many circumstances. </p>\n\n<p>Also the dl speed is limited to about 200k/s (observed) which further increases the likelihood of failure.</p>\n\n<p>I finally d/l'ed the file with some wget-foo.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "137973": "I want to download file page_views.csv.zip, but my connection fails at some point. Is it possible to upload the file in parts or upload this file to a torrent service?",
    "138033": "Hi, I'm afraid torrent is not going to be an option. From the \"[Welcome][1]\" post you can read:\r\n\r\n\"Expect a traffic jam to download the 30 GB file during the launch gold rush (sorry!). Use a download manager with resuming capabilities and we promise you'll eventually get the file. There is also a small sample available. Please read our [note][2] on torrents before you suggest using torrents.\"\r\n\r\nCheers!\r\n\r\n\r\n  [1]: https://www.kaggle.com/c/outbrain-click-prediction/forums/t/24107/welcome\r\n  [2]: https://www.kaggle.com/wiki/ANoteOnTorrents",
    "138185": "I'm pretty sure one could torrent an encrypted version of the data and then force users to download the key directly from kaggle. Versioning could be handled in this way too. Given the failure rate, it will take me probably a week to d/l this file. \r\n\r\nMaybe we should make a \"kaggle\" a verb for something that's impossible to d/l, a la \"I tried to download the new iOS on launch but download system is totally kaggled\".",
    "138247": "Well, I've been able to download it via `curl`:\r\n\r\n 1. Download some of the light files so in your browser you have to accept the data agreement and this action is stored in your `cookies`\r\n 2. In your browser (in my case it is Chrome) right click and open the Inspector\r\n 3. Click on the super heavy file and in the Inspector check the petition in the Network tab\r\n 4. Right click on the network petition and `copy as cURL`\r\n 5. Finally you can open the terminal and paste the `curl` command\r\n\r\nI was able to download it at once, with a pretty shitty connexion. But the nice thing about curl is that you can resume the download by setting an offset:\r\n\r\n    curl -C offset url\r\n\r\nSo if it has stopped, you can check for the size of the current file, and set it (+1 I think) as the offset parameter.\r\n\r\nHope this is helpful to somebody, I know it shouldn't be that elaborated, but it's feasible ;)",
    "138585": "Most likely, that's against the contest rules. But just as a test, no, it's not possible. After several days of retrying:\r\n\r\n\"\r\nwget 'https://kaggle2.blob.core.windows.net/competitions-data/kaggle/5497/page_views.csv.zip? ...'<snip>\r\n\r\nResolving kaggle2.blob.core.windows.net (kaggle2.blob.core.windows.net)... 40.116.120.24\r\nConnecting to kaggle2.blob.core.windows.net (kaggle2.blob.core.windows.net)|40.116.120.24|:443... connected.\r\nHTTP request sent, awaiting response... 200 OK\r\nLength: 31905475875 (30G) [application/zip]\r\nSaving to: 'page_views.csv.zip... <snip>'\r\npage_views.csv.zip@sv=2012-02-12&se=201  35%[==========================>                                                  ]  10.45G   277KB/s    in 14h 12m\r\n\r\n<snip several download/error/download more etc>\r\n\r\n2016-10-08 07:46:50 (214 KB/s) - Read error at byte 11221434368/31905475875 (Resource temporarily unavailable, try again.). Retrying.\r\n\r\nHTTP request sent, awaiting response... 403 Server failed to authenticate the request. Make sure the value of Authorization header is formed correctly including the signature.\r\n2016-10-09 09:22:35 ERROR 403: Server failed to authenticate the request. Make sure the value of Authorization header is formed correctly including the signature..\r\n\"",
    "138648": "Firefox allows restarting if the download fails.  I found I couldn't use a command-line utility because I needed to have an authenticated browser session, and I couldn't work out how to do that.\r\n\r\nI've been able to resume the download four times so far...",
    "138841": "Chrome has the same. However it seems the azure-hosted files have a token that times out after a few days? So even if you keep restarting, you have to get it in that period or the token (on their end) will timeout, apparently. Retrying therefore may not work in many circumstances. \r\n\r\nAlso the dl speed is limited to about 200k/s (observed) which further increases the likelihood of failure.\r\n\r\nI finally d/l'ed the file with some wget-foo."
  },
  "source": "meta"
}