{
  "id": 10311,
  "title": "Downloading the data",
  "url": "/competitions/seizure-prediction/discussion/10311",
  "author_name": "",
  "post_date": "2014-09-13T02:17:20.077Z",
  "votes": null,
  "comment_count": 7,
  "views": 2102,
  "content": "<p>Hi everyone:</p>\n<p>I want to download the data to my school's HPC, because my own laptop does not have such large memory to deal with GB data. I am new to unix, so I tried some simple command like lynx, wget and curl, noting works.</p>\n<p>I have tried some method decribed in:</p>\n<p>http://www.kaggle.com/c/belkin-energy-disaggregation-competition/forums/t/5118/downloading-data-via-wget/49609</p>\n<p>Still not work.&nbsp;</p>\n<p>I constantly get the download url www.kaggle.com/account/login?Returnurl=%2fc%2fseizure-prediction%2fdownload%2fDog_1.tar.gz</p>\n<p>I guess maybe this is a login problem, that is, I have not correctly logined.</p>\n<p>Is there any possible solution?</p>\n<p>Best Regards,</p>",
  "messages": [
    {
      "id": "53604",
      "postDate": "09/13/2014 02:17:20",
      "content": "<p>Hi everyone:</p>\n<p>I want to download the data to my school's HPC, because my own laptop does not have such large memory to deal with GB data. I am new to unix, so I tried some simple command like lynx, wget and curl, noting works.</p>\n<p>I have tried some method decribed in:</p>\n<p>http://www.kaggle.com/c/belkin-energy-disaggregation-competition/forums/t/5118/downloading-data-via-wget/49609</p>\n<p>Still not work.&nbsp;</p>\n<p>I constantly get the download url www.kaggle.com/account/login?Returnurl=%2fc%2fseizure-prediction%2fdownload%2fDog_1.tar.gz</p>\n<p>I guess maybe this is a login problem, that is, I have not correctly logined.</p>\n<p>Is there any possible solution?</p>\n<p>Best Regards,</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "53606",
      "postDate": "09/13/2014 02:39:15",
      "content": "<p>Did you accept the terms of the challenge ?</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "53608",
      "postDate": "09/13/2014 02:53:19",
      "content": "<p>Yes, I have accepted them.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "53657",
      "postDate": "09/13/2014 18:47:16",
      "content": "<p>You also need to first get the cookies in a browser while being authenticated and then pass them to wget. Take a look here:</p>\n<p>http://www.kaggle.com/c/belkin-energy-disaggregation-competition/forums/t/5118/downloading-data-via-wget/49609</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "53664",
      "postDate": "09/13/2014 21:20:57",
      "content": "<p>I was able to use the lynx as described in the link elyase provided to store the kaggle cookies and to use wget to download while using those cookies&nbsp; ...</p>\n<p><code>wget --load-cookies ~/.lynx_cookies -p http://www.kaggle.com/the/rest/of/the/data/url<br></code></p>\n\n<p>... as also described in that link</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "53668",
      "postDate": "09/13/2014 22:36:08",
      "content": "<p>Here's another way...</p>\n<p>This method is for downloading files using curl (wget should also work fine) without using cookie files or any special setup.</p>\n<p>What I did was start the download in a regular web browser, e.g. Chrome. As soon as the download starts, pause/stop it. Then, go to the Chrome Downloads tab&nbsp;and copy the URL of the dataset you were downloading (the Chrome URL will be quite a lot longer and complex than the one you are currently using, as it contains your cookies and login info in it). Copy this new&nbsp;personalized URL from Chrome and use it with curl just as you did before and&nbsp;your files&nbsp;should&nbsp;start downloading without a hitch. Repeat the process for each dataset you want to download. It worked for me.</p>\n<p>I hope this helps!</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "54330",
      "postDate": "09/20/2014 04:40:46",
      "content": "<p>Did you check the storage limitation per user with your HPC ? Ask your administrator to increase it.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "54331",
      "postDate": "09/20/2014 05:22:30",
      "content": "<p>Hi ,</p>\n<p>I tried using curl with the copied URL, but getting the following error from the server:</p>\n<p>&quot;The specified resource does not exist&quot;</p>\n<p>any help is highly appreciated !</p>\n<p>Thanks,</p>\n<p>Sanjeev</p>\n<p>[quote=spacemanspiff;53668]</p>\n<p>Here's another way...</p>\n<p>This method is for downloading files using curl (wget should also work fine) without using cookie files or any special setup.</p>\n<p>What I did was start the download in a regular web browser, e.g. Chrome. As soon as the download starts, pause/stop it. Then, go to the Chrome Downloads tab&nbsp;and copy the URL of the dataset you were downloading (the Chrome URL will be quite a lot longer and complex than the one you are currently using, as it contains your cookies and login info in it). Copy this new&nbsp;personalized URL from Chrome and use it with curl just as you did before and&nbsp;your files&nbsp;should&nbsp;start downloading without a hitch. Repeat the process for each dataset you want to download. It worked for me.</p>\n<p>I hope this helps!</p>\n<p>[/quote]</p>",
      "rawMarkdown": "",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 53606,
      "author_name": "franklyn",
      "author_url": "",
      "post_date": "09/13/2014 02:39:15",
      "content": "<p>Did you accept the terms of the challenge ?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 53608,
      "author_name": "zhongxiaolong",
      "author_url": "",
      "post_date": "09/13/2014 02:53:19",
      "content": "<p>Yes, I have accepted them.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 53657,
      "author_name": "elyase",
      "author_url": "",
      "post_date": "09/13/2014 18:47:16",
      "content": "<p>You also need to first get the cookies in a browser while being authenticated and then pass them to wget. Take a look here:</p>\n<p>http://www.kaggle.com/c/belkin-energy-disaggregation-competition/forums/t/5118/downloading-data-via-wget/49609</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 53664,
      "author_name": "rbroberg",
      "author_url": "",
      "post_date": "09/13/2014 21:20:57",
      "content": "<p>I was able to use the lynx as described in the link elyase provided to store the kaggle cookies and to use wget to download while using those cookies&nbsp; ...</p>\n<p><code>wget --load-cookies ~/.lynx_cookies -p http://www.kaggle.com/the/rest/of/the/data/url<br></code></p>\n\n<p>... as also described in that link</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 53668,
      "author_name": "jordicohen",
      "author_url": "",
      "post_date": "09/13/2014 22:36:08",
      "content": "<p>Here's another way...</p>\n<p>This method is for downloading files using curl (wget should also work fine) without using cookie files or any special setup.</p>\n<p>What I did was start the download in a regular web browser, e.g. Chrome. As soon as the download starts, pause/stop it. Then, go to the Chrome Downloads tab&nbsp;and copy the URL of the dataset you were downloading (the Chrome URL will be quite a lot longer and complex than the one you are currently using, as it contains your cookies and login info in it). Copy this new&nbsp;personalized URL from Chrome and use it with curl just as you did before and&nbsp;your files&nbsp;should&nbsp;start downloading without a hitch. Repeat the process for each dataset you want to download. It worked for me.</p>\n<p>I hope this helps!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 54330,
      "author_name": "heman35593",
      "author_url": "",
      "post_date": "09/20/2014 04:40:46",
      "content": "<p>Did you check the storage limitation per user with your HPC ? Ask your administrator to increase it.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 54331,
      "author_name": "sanjeevsingh",
      "author_url": "",
      "post_date": "09/20/2014 05:22:30",
      "content": "<p>Hi ,</p>\n<p>I tried using curl with the copied URL, but getting the following error from the server:</p>\n<p>&quot;The specified resource does not exist&quot;</p>\n<p>any help is highly appreciated !</p>\n<p>Thanks,</p>\n<p>Sanjeev</p>\n<p>[quote=spacemanspiff;53668]</p>\n<p>Here's another way...</p>\n<p>This method is for downloading files using curl (wget should also work fine) without using cookie files or any special setup.</p>\n<p>What I did was start the download in a regular web browser, e.g. Chrome. As soon as the download starts, pause/stop it. Then, go to the Chrome Downloads tab&nbsp;and copy the URL of the dataset you were downloading (the Chrome URL will be quite a lot longer and complex than the one you are currently using, as it contains your cookies and login info in it). Copy this new&nbsp;personalized URL from Chrome and use it with curl just as you did before and&nbsp;your files&nbsp;should&nbsp;start downloading without a hitch. Repeat the process for each dataset you want to download. It worked for me.</p>\n<p>I hope this helps!</p>\n<p>[/quote]</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "53604": "",
    "53606": "",
    "53608": "",
    "53657": "",
    "53664": "",
    "53668": "",
    "54330": "",
    "54331": ""
  },
  "source": "meta"
}