{
  "id": 5118,
  "title": "Downloading data via wget",
  "url": "/competitions/belkin-energy-disaggregation-competition/discussion/5118",
  "author_name": "",
  "post_date": "2013-07-16T18:19:41.107Z",
  "votes": 6,
  "comment_count": 17,
  "views": 20365,
  "content": "<p>Hi all, I wanted to download the data to a remote server using wget. The download didn't work because you first have to accept the competition rules before downloading. Is there any way around this issue?</p>",
  "messages": [
    {
      "id": "27284",
      "postDate": "07/16/2013 18:19:41",
      "content": "<p>Hi all, I wanted to download the data to a remote server using wget. The download didn't work because you first have to accept the competition rules before downloading. Is there any way around this issue?</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "27286",
      "postDate": "07/16/2013 18:25:15",
      "content": "<p>I've run into this issue before and haven't found a solution and have always had to scp from a local machine to the remote. Would also like to know if there's a better way!</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "27322",
      "postDate": "07/17/2013 03:47:06",
      "content": "<p>I did exactly this because the downloads were annoying me via Chrome and I trust wgets http resume transfer more.</p>\r\n<p>The trick is to export your cookies from your browser and save them in a file (i use https://chrome.google.com/webstore/detail/lopabhfecdfhgogdbojmaicoicjekelh)&nbsp;and then use wget's&nbsp;<span>--load-cookies option.</span></p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "27344",
      "postDate": "07/17/2013 13:46:46",
      "content": "<p>Amazing!</p>\r\n<p>I saved the cookies from the data page after accepting the rules, and then entered:</p>\r\n<p>wget -x --load-cookies cookies.txt http://www.kaggle.com/c/belkin-energy-disaggregation-competition/download/H1_CSV.zip</p>\r\n<p>I wanted to work with the files on Clemson's Palmetto Cluster. It was going to take 24 hours to transfer the files over Filezilla.&nbsp;I was able to download all the files to the cluster in under 10 minutes!&nbsp;</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "27436",
      "postDate": "07/19/2013 09:29:12",
      "content": "<p>Hmmm, this isn't working for me. I've copied and pasted the text in the chrome extension to cookies.txt and tried wget --load-cookies cookies.txt&nbsp;http://www.kaggle.com/c/belkin-energy-disaggregation-competition/download/H2.zip with no luck.</p>\n<p>Downloading via browser is timing out for me, so I have no way of getting the data at the moment. Any ideas?</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "27486",
      "postDate": "07/21/2013 09:54:40",
      "content": "<p>[quote=sayhey69;27436]</p>\n<p>Hmmm, this isn't working for me. I've copied and pasted the text in the chrome extension to cookies.txt and tried wget --load-cookies cookies.txt&nbsp;http://www.kaggle.com/c/belkin-energy-disaggregation-competition/download/H2.zip with no luck.</p>\n<p>Downloading via browser is timing out for me, so I have no way of getting the data at the moment. Any ideas?</p>\n<p>[/quote]</p>\n<p>&nbsp;</p>\n<p>First I also had an error, saying it was impossible to check certificate. I added the &quot;--no-check-certificate&quot; option on wget command line, and now it seems to work.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "27720",
      "postDate": "07/27/2013 18:13:03",
      "content": "<p>I used lynx (text based web browser) to download the files from a remote EC2 instance. It is a bit cumbersome to login and navigate at first, manual intensive, but it works:</p>\n<p>1) Install Lynx, if you don't have it</p>\n<p>2) Create a ~/.lynxrc configuration file such as:</p>\n<p><code>SET_COOKIES:TRUE<br>ACCEPT_ALL_COOKIES:TRUE<br>PERSISTENT_COOKIES:TRUE<br>COOKIE_FILE:~/.lynx_cookies<br>COOKIE_SAVE_FILE:~/.lynx_cookies</code><code></code></p>\n<p>3) Call the browser</p>\n<p><code>lynx -cfg=~/.lynxrc www.kaggle.com</code></p>\n<p>4) Log in, browse to the competition data page and accept the terms and permissions (if you haven't yet)</p>\n<p>5) Select the link to the file you want to download, and press &quot;d&quot;. The download will start.</p>\n<p>6) Once the download is finished, select &quot;save file to disk&quot; and provide filename/destination where you want to store the data</p>\n<p>7) Repeat for other files</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "30858",
      "postDate": "09/13/2013 15:51:44",
      "content": "<p>Hi, all</p>\n<p>&nbsp;</p>\n<p>Recently I experienced the download failure using chrom (due to slow download speed). I figure out a way to use wget and lynx.</p>\n<p><strong>(1) start lynx, open kaggle.com, login to your user ID. Remember to always accept cookies. and the cookies will be stored at ~/.lynx_cookies. move it to some location and rename it to cookies.&nbsp;</strong></p>\n<p><strong>(2) wget -c --load-cookies=cookies&nbsp;http://www.kaggle.com/c/belkin-energy-disaggregation-competition/download/H3.zip</strong></p>\n<p>then, you can resume the downloading process any time you want.</p>\n<p>&nbsp;</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "32067",
      "postDate": "10/06/2013 19:48:13",
      "content": "<p>I tried to go via&nbsp;scp -r xxxx/yyyy@www.kaggle.com:/c/belkin-energy-disaggregation-competition/download/H1.zip &nbsp; /Users/xxxxx</p>\n<p>but it is not working. It times out. do you have a better server to tie to?</p>\n<p>X</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "32111",
      "postDate": "10/08/2013 12:27:28",
      "content": "<p>I have only success with lynx + wget.&nbsp;</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "49609",
      "postDate": "06/26/2014 09:16:08",
      "content": "<p>Thanks, lynx worked really.</p>\n\n<p>By the way, I want to download the data of&nbsp;Display Advertising Challenge just now.</p>\n<p>I find the url of file contains https. So, lynx should be installed with https support.</p>\n<p>It is very easy to make it with the SSL configure option (--with-ssl). Please reference&nbsp;</p>\n<p>http://lynx.isc.org/current/README.ssl.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "61563",
      "postDate": "12/20/2014 04:22:19",
      "content": "<p>l used lynx to download the files from a remote EC2 instance .<br>when i login by google , pass the verify<br>it show's<br>&quot; Kaggle.com and Google will use this information in accordance with their respective terms of</p>\n<p>service and privacy policies.</p>\n<p>(BUTTON) Accept</p>\n<p>DISABLED form submit button. &quot;</p>\n<p><br>but isn't working to enter the accept button .</p>\n<p>i found the problem is &quot;Lynx doesn't support JavaScript&quot; , is there another way to pass through the problem??</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "63585",
      "postDate": "02/04/2015 14:27:21",
      "content": "<p>I am stuck on the same problem Chifang Jang. I found the best answer was to use Chrome, and copy and the paste the cookies from Chrome Preferences into a cookies file, then use than in wget.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "80431",
      "postDate": "05/30/2015 03:19:56",
      "content": "<p>Okay, I don't know if you peeps are making the&nbsp;same mistake, but at first I didn't notice that when using the Chrome plug-in, you have to click on the small &quot;cookie.txt export&quot; icon, while the browser is on the kaggle.com domain ! Otherwise it will pick up the cookies from whatever domain is showing (I happened to have amazon.com while I clicked on the icon, and there were so many cookies that I assumed it exported all cookies from all sites).</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "96790",
      "postDate": "10/20/2015 10:35:05",
      "content": "<p>Thanks @Eduardo.\nIt was exactly my case! </p>",
      "rawMarkdown": "Thanks @Eduardo.\r\nIt was exactly my case!",
      "votes": null
    },
    {
      "id": "112602",
      "postDate": "03/22/2016 09:22:20",
      "content": "<p><a href=\"http://yasermartinez.com/blog/posts/web-scraping-iii.html\">http://yasermartinez.com/blog/posts/web-scraping-iii.html</a>  This blog give detailed steps to download dataset from Kaggle. I tried it, and it works.</p>",
      "rawMarkdown": "http://yasermartinez.com/blog/posts/web-scraping-iii.html  This blog give detailed steps to download dataset from Kaggle. I tried it, and it works.",
      "votes": null
    },
    {
      "id": "152505",
      "postDate": "12/26/2016 19:24:40",
      "content": "<p>HI Wallace - </p>\n\n<p>I am using EC2. The wget command - You run that on the remote instance. how do you get the cookie.txt on to the remote instance.</p>\n\n<p>Thanks,</p>",
      "rawMarkdown": "HI Wallace - \r\n\r\nI am using EC2. The wget command - You run that on the remote instance. how do you get the cookie.txt on to the remote instance.\r\n\r\nThanks,",
      "votes": null
    },
    {
      "id": "161785",
      "postDate": "02/15/2017 19:45:47",
      "content": "<p>I am trying to download dataset thru wget and cookies option to google cloud but only 15.3 KB file is getting downloaded in place of 7.3 GB file ...Can you please suggest what is going wrong </p>",
      "rawMarkdown": "I am trying to download dataset thru wget and cookies option to google cloud but only 15.3 KB file is getting downloaded in place of 7.3 GB file ...Can you please suggest what is going wrong",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 27286,
      "author_name": "alecradford",
      "author_url": "",
      "post_date": "07/16/2013 18:25:15",
      "content": "<p>I've run into this issue before and haven't found a solution and have always had to scp from a local machine to the remote. Would also like to know if there's a better way!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 27322,
      "author_name": "zacstewart",
      "author_url": "",
      "post_date": "07/17/2013 03:47:06",
      "content": "<p>I did exactly this because the downloads were annoying me via Chrome and I trust wgets http resume transfer more.</p>\r\n<p>The trick is to export your cookies from your browser and save them in a file (i use https://chrome.google.com/webstore/detail/lopabhfecdfhgogdbojmaicoicjekelh)&nbsp;and then use wget's&nbsp;<span>--load-cookies option.</span></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 27344,
      "author_name": "wallacecampbell",
      "author_url": "",
      "post_date": "07/17/2013 13:46:46",
      "content": "<p>Amazing!</p>\r\n<p>I saved the cookies from the data page after accepting the rules, and then entered:</p>\r\n<p>wget -x --load-cookies cookies.txt http://www.kaggle.com/c/belkin-energy-disaggregation-competition/download/H1_CSV.zip</p>\r\n<p>I wanted to work with the files on Clemson's Palmetto Cluster. It was going to take 24 hours to transfer the files over Filezilla.&nbsp;I was able to download all the files to the cluster in under 10 minutes!&nbsp;</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 27436,
      "author_name": "foobarbazquux",
      "author_url": "",
      "post_date": "07/19/2013 09:29:12",
      "content": "<p>Hmmm, this isn't working for me. I've copied and pasted the text in the chrome extension to cookies.txt and tried wget --load-cookies cookies.txt&nbsp;http://www.kaggle.com/c/belkin-energy-disaggregation-competition/download/H2.zip with no luck.</p>\n<p>Downloading via browser is timing out for me, so I have no way of getting the data at the moment. Any ideas?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 27486,
      "author_name": "cbourguignat",
      "author_url": "",
      "post_date": "07/21/2013 09:54:40",
      "content": "<p>[quote=sayhey69;27436]</p>\n<p>Hmmm, this isn't working for me. I've copied and pasted the text in the chrome extension to cookies.txt and tried wget --load-cookies cookies.txt&nbsp;http://www.kaggle.com/c/belkin-energy-disaggregation-competition/download/H2.zip with no luck.</p>\n<p>Downloading via browser is timing out for me, so I have no way of getting the data at the moment. Any ideas?</p>\n<p>[/quote]</p>\n<p>&nbsp;</p>\n<p>First I also had an error, saying it was impossible to check certificate. I added the &quot;--no-check-certificate&quot; option on wget command line, and now it seems to work.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 27720,
      "author_name": "mobster",
      "author_url": "",
      "post_date": "07/27/2013 18:13:03",
      "content": "<p>I used lynx (text based web browser) to download the files from a remote EC2 instance. It is a bit cumbersome to login and navigate at first, manual intensive, but it works:</p>\n<p>1) Install Lynx, if you don't have it</p>\n<p>2) Create a ~/.lynxrc configuration file such as:</p>\n<p><code>SET_COOKIES:TRUE<br>ACCEPT_ALL_COOKIES:TRUE<br>PERSISTENT_COOKIES:TRUE<br>COOKIE_FILE:~/.lynx_cookies<br>COOKIE_SAVE_FILE:~/.lynx_cookies</code><code></code></p>\n<p>3) Call the browser</p>\n<p><code>lynx -cfg=~/.lynxrc www.kaggle.com</code></p>\n<p>4) Log in, browse to the competition data page and accept the terms and permissions (if you haven't yet)</p>\n<p>5) Select the link to the file you want to download, and press &quot;d&quot;. The download will start.</p>\n<p>6) Once the download is finished, select &quot;save file to disk&quot; and provide filename/destination where you want to store the data</p>\n<p>7) Repeat for other files</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 30858,
      "author_name": "liubenyuan",
      "author_url": "",
      "post_date": "09/13/2013 15:51:44",
      "content": "<p>Hi, all</p>\n<p>&nbsp;</p>\n<p>Recently I experienced the download failure using chrom (due to slow download speed). I figure out a way to use wget and lynx.</p>\n<p><strong>(1) start lynx, open kaggle.com, login to your user ID. Remember to always accept cookies. and the cookies will be stored at ~/.lynx_cookies. move it to some location and rename it to cookies.&nbsp;</strong></p>\n<p><strong>(2) wget -c --load-cookies=cookies&nbsp;http://www.kaggle.com/c/belkin-energy-disaggregation-competition/download/H3.zip</strong></p>\n<p>then, you can resume the downloading process any time you want.</p>\n<p>&nbsp;</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 32067,
      "author_name": "dm131478",
      "author_url": "",
      "post_date": "10/06/2013 19:48:13",
      "content": "<p>I tried to go via&nbsp;scp -r xxxx/yyyy@www.kaggle.com:/c/belkin-energy-disaggregation-competition/download/H1.zip &nbsp; /Users/xxxxx</p>\n<p>but it is not working. It times out. do you have a better server to tie to?</p>\n<p>X</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 32111,
      "author_name": "liubenyuan",
      "author_url": "",
      "post_date": "10/08/2013 12:27:28",
      "content": "<p>I have only success with lynx + wget.&nbsp;</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 49609,
      "author_name": "lifematrix",
      "author_url": "",
      "post_date": "06/26/2014 09:16:08",
      "content": "<p>Thanks, lynx worked really.</p>\n\n<p>By the way, I want to download the data of&nbsp;Display Advertising Challenge just now.</p>\n<p>I find the url of file contains https. So, lynx should be installed with https support.</p>\n<p>It is very easy to make it with the SSL configure option (--with-ssl). Please reference&nbsp;</p>\n<p>http://lynx.isc.org/current/README.ssl.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 61563,
      "author_name": "chifangjang",
      "author_url": "",
      "post_date": "12/20/2014 04:22:19",
      "content": "<p>l used lynx to download the files from a remote EC2 instance .<br>when i login by google , pass the verify<br>it show's<br>&quot; Kaggle.com and Google will use this information in accordance with their respective terms of</p>\n<p>service and privacy policies.</p>\n<p>(BUTTON) Accept</p>\n<p>DISABLED form submit button. &quot;</p>\n<p><br>but isn't working to enter the accept button .</p>\n<p>i found the problem is &quot;Lynx doesn't support JavaScript&quot; , is there another way to pass through the problem??</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 63585,
      "author_name": "jaymoore",
      "author_url": "",
      "post_date": "02/04/2015 14:27:21",
      "content": "<p>I am stuck on the same problem Chifang Jang. I found the best answer was to use Chrome, and copy and the paste the cookies from Chrome Preferences into a cookies file, then use than in wget.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 80431,
      "author_name": "evalle",
      "author_url": "",
      "post_date": "05/30/2015 03:19:56",
      "content": "<p>Okay, I don't know if you peeps are making the&nbsp;same mistake, but at first I didn't notice that when using the Chrome plug-in, you have to click on the small &quot;cookie.txt export&quot; icon, while the browser is on the kaggle.com domain ! Otherwise it will pick up the cookies from whatever domain is showing (I happened to have amazon.com while I clicked on the icon, and there were so many cookies that I assumed it exported all cookies from all sites).</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 96790,
      "author_name": "tsereda",
      "author_url": "",
      "post_date": "10/20/2015 10:35:05",
      "content": "<p>Thanks @Eduardo.\nIt was exactly my case! </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 112602,
      "author_name": "qiuwuyiye",
      "author_url": "",
      "post_date": "03/22/2016 09:22:20",
      "content": "<p><a href=\"http://yasermartinez.com/blog/posts/web-scraping-iii.html\">http://yasermartinez.com/blog/posts/web-scraping-iii.html</a>  This blog give detailed steps to download dataset from Kaggle. I tried it, and it works.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 152505,
      "author_name": "sonads",
      "author_url": "",
      "post_date": "12/26/2016 19:24:40",
      "content": "<p>HI Wallace - </p>\n\n<p>I am using EC2. The wget command - You run that on the remote instance. how do you get the cookie.txt on to the remote instance.</p>\n\n<p>Thanks,</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 161785,
      "author_name": "santanuds",
      "author_url": "",
      "post_date": "02/15/2017 19:45:47",
      "content": "<p>I am trying to download dataset thru wget and cookies option to google cloud but only 15.3 KB file is getting downloaded in place of 7.3 GB file ...Can you please suggest what is going wrong </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "27284": "",
    "27286": "",
    "27322": "",
    "27344": "",
    "27436": "",
    "27486": "",
    "27720": "",
    "30858": "",
    "32067": "",
    "32111": "",
    "49609": "",
    "61563": "",
    "63585": "",
    "80431": "",
    "96790": "Thanks @Eduardo.\r\nIt was exactly my case!",
    "112602": "http://yasermartinez.com/blog/posts/web-scraping-iii.html  This blog give detailed steps to download dataset from Kaggle. I tried it, and it works.",
    "152505": "HI Wallace - \r\n\r\nI am using EC2. The wget command - You run that on the remote instance. how do you get the cookie.txt on to the remote instance.\r\n\r\nThanks,",
    "161785": "I am trying to download dataset thru wget and cookies option to google cloud but only 15.3 KB file is getting downloaded in place of 7.3 GB file ...Can you please suggest what is going wrong"
  },
  "source": "meta"
}