{
  "id": 12407,
  "title": "Anyone managed to use wget from a Google login?",
  "url": "/competitions/malware-classification/discussion/12407",
  "author_name": "",
  "post_date": "2015-02-04T14:22:25.770Z",
  "votes": 1,
  "comment_count": 18,
  "views": 6354,
  "content": "<p>I'm trying to wget the data, which requires Kaggle credentials. &nbsp;I usually log into Kaggle using a Google login, anyone any idea how to get that to work in wget?</p>\n<p>I've tried going down the lynx route, to capture the cookies and pass them to wget, and I can log into Google OK through lynx, but can never get as far as getting Kaggle to log me in using lynx.</p>\n<p>Do I need to make a specific Kaggle login with its own username and password? &nbsp;I know this is against Kaggle rules in any case (using multiple logins).</p>",
  "messages": [
    {
      "id": "63583",
      "postDate": "02/04/2015 14:22:25",
      "content": "<p>I'm trying to wget the data, which requires Kaggle credentials. &nbsp;I usually log into Kaggle using a Google login, anyone any idea how to get that to work in wget?</p>\n<p>I've tried going down the lynx route, to capture the cookies and pass them to wget, and I can log into Google OK through lynx, but can never get as far as getting Kaggle to log me in using lynx.</p>\n<p>Do I need to make a specific Kaggle login with its own username and password? &nbsp;I know this is against Kaggle rules in any case (using multiple logins).</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "63584",
      "postDate": "02/04/2015 14:25:45",
      "content": "<p>download the cookies of kaggle and use the following:</p>\n<p>wget -x --load-cookies cookie_file.txt <em>url_to_dataset</em></p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "63586",
      "postDate": "02/04/2015 14:30:34",
      "content": "<p>Abishek, this is what I am trying to do. &nbsp;I use lynx to log in through Google, I get through the Google login fine, but it never redirects back to Kaggle, I just get a page back that says 'use a javascript enabled browser', and no kaggle cookies get stored in the cookie file. &nbsp;Are you managing to log in via Google in lynx?</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "63587",
      "postDate": "02/04/2015 14:36:33",
      "content": "<p>Well, Im not using lynx. I also have google login to kaggle and the above command works perfectly fine for me.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "63588",
      "postDate": "02/04/2015 14:46:29",
      "content": "<p>I found another way to get the cookies now, rather than lynx. &nbsp;Thanks anyway.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "63602",
      "postDate": "02/04/2015 19:38:36",
      "content": "<p>@Jay, could you please share your way? Thanks in advance :)&nbsp;</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "63603",
      "postDate": "02/04/2015 19:40:35",
      "content": "<p>If you guys plan on using wget or similar tools the easiest way to get cookies is to login into your account and use</p>\n<p>https://addons.mozilla.org/en-US/firefox/addon/export-cookies/ on Firefox</p>\n<p>or</p>\n<p>https://chrome.google.com/webstore/detail/cookietxt-export/lopabhfecdfhgogdbojmaicoicjekelh on Chrome</p>\n<p>and then use what Abhishek suggested</p>\n<p>wget -x --load-cookies cookie_file.txt url_to_dataset</p>\n<p>Don't forget to accept the rules first otherwise it won't work</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "63628",
      "postDate": "02/05/2015 03:05:17",
      "content": "<p>Or copy as curl from Chrome:</p>\n<p>http://www.lornajane.net/posts/2013/chrome-feature-copy-as-curl</p>\n<p>delete the detritus after the long single quoted https string.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "63634",
      "postDate": "02/05/2015 10:46:47",
      "content": "<p>@Amika I used Chrome, where I had logged into kaggle previously and accepted the terms of the competition. &nbsp;I went to Preferences/Settings and clicked Show advanced settings&nbsp;</p>\n<p>This brought up the Privacy section and I clicked Content Settings and then All cookies and site data.</p>\n<p>Then I searched for kaggle.com among the cookies list, and selected the cookie called .ASPXAUTH</p>\n<p>I took the information from this cookie, and pasted it into the cookies file that lynx had previously created (I had previously used lynx to try to do this, and had logged into Google through lynx). &nbsp;</p>\n<p>Then I called wget using this cookie file, and it worked first time, to my surprise :).</p>\n<p>Good luck!</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "64323",
      "postDate": "02/16/2015 08:29:13",
      "content": "<p>Hi</p>\n<p>Can anyone tell me how to resolve the error below? I downloaded the cookie file from chrome, and am logged into Kaggle using google...</p>\n<p>C:\\Program Files (x86)\\GnuWin32\\bin&gt;wget -x --load-cookies cookies.txt https<br>ww.kaggle.com/c/malware-classification/download/train.7z<br>SYSTEM_WGETRC = c:/progra~1/wget/etc/wgetrc<br>syswgetrc = C:\\Program Files (x86)\\GnuWin32/etc/wgetrc<br>--2015-02-16 15:57:14-- https://www.kaggle.com/c/malware-classification/dow<br>d/train.7z<br>Resolving www.kaggle.com... 168.62.224.124<br>Connecting to www.kaggle.com|168.62.224.124|:443... <strong>connected</strong>.<br><strong>ERROR: cannot verify www.kaggle.com's certificate, issued by `/C=US/O=GeoTru</strong><br><strong>nc./CN=RapidSSL SHA256 CA - G3':</strong><br> Unable to locally verify the issuer's authority.<br>To connect to www.kaggle.com insecurely, use `--no-check-certificate'.<br>Unable to establish SSL connection.</p>\n<p>Thanks!</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "64324",
      "postDate": "02/16/2015 08:33:36",
      "content": "<p>did you try this:&nbsp;</p>\n\n<p>wget -x --load-cookies cookies.txt --no-check-certificate&nbsp;https://www.kaggle.com/c/malware-classification/download/train.7z</p>\n\n<p>or simply:</p>\n<p>wget -x --load-cookies cookies.txt http://www.kaggle.com/c/malware-classification/download/train.7z</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "64326",
      "postDate": "02/16/2015 08:39:47",
      "content": "<p>as above... after I got my nas working tonight... wget with the syntax above worked perfectly.</p>\n<p>Once I get nfs mounting properly I may even be able to unzip the archives.</p>\n<p>By the time everyone else has results I will have a working environment.&nbsp; :)</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "64328",
      "postDate": "02/16/2015 08:45:34",
      "content": "<p>Hey guys</p>\n\n<p>Thanks for the speedy reply....just got it working by adding<strong>&nbsp;--no-check-certificate</strong> to the cmd line ....but you guys replied before I can update my status! (Yes, you guys rock)</p>\n<p>@Robert-dont worry, there are enough noobs around (like me), who will take forever to set up an environment.....you are doing alright ;)</p>\n<p>(1% downloaded....yippiee :D )</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "64337",
      "postDate": "02/16/2015 11:43:48",
      "content": "<p>eh..umm...my wget doesnt resume download from where it left off.....rather it is restarting download from 0%. (I use -c to continue)....can some1 help? Apologies for wasting your time ...I'm sure I'm doing something stupid....</p>\n\n<p>Edit: finally got the resume working..phew.. I had to use -O to change the name of the saved file for wget to continue saving to the same file</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "145556",
      "postDate": "11/19/2016 08:45:39",
      "content": "<p>[quote=wacax;63603]</p>\n<p>If you guys plan on using wget or similar tools the easiest way to get cookies is to login into your account and use</p>\n<p>https://addons.mozilla.org/en-US/firefox/addon/export-cookies/ on Firefox</p>\n<p>or</p>\n<p>https://chrome.google.com/webstore/detail/cookietxt-export/lopabhfecdfhgogdbojmaicoicjekelh on Chrome</p>\n<p>and then use what Abhishek suggested</p>\n<p>wget -x --load-cookies cookie_file.txt url_to_dataset</p>\n<p>Don't forget to accept the rules first otherwise it won't work</p>\n<p>[/quote]</p>\n\n<p>Hi Dear&nbsp;</p>\n<p>The text were you wrote :&nbsp;wget -x --load-cookies cookie_file.txt url_to_dataset</p>\n<p>Is it exactly the same text? if not exact text can you explain the parameters above?&nbsp;</p>\n<p>&nbsp;&nbsp;</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "145558",
      "postDate": "11/19/2016 09:00:32",
      "content": "<p>It is not work I go to the Directory where files should be located and I didn't found any thing.</p>\n\n<p>I'm sure the coockie.txt is saved on chrom .</p>",
      "rawMarkdown": "It is not work I go to the Directory where files should be located and I didn't found any thing.\r\n\r\nI'm sure the coockie.txt is saved on chrom .",
      "votes": null
    },
    {
      "id": "145559",
      "postDate": "11/19/2016 09:02:57",
      "content": "<p>Hi wacax\nYou said : \"Don't forget to accept the rules first otherwise it won't work\"</p>\n\n<p>Where this rules dialog is displayed is it through lynx or on windows ? , need clarification.</p>\n\n<p>Thank you</p>",
      "rawMarkdown": "Hi wacax\r\nYou said : \"Don't forget to accept the rules first otherwise it won't work\"\r\n\r\nWhere this rules dialog is displayed is it through lynx or on windows ? , need clarification.\r\n\r\nThank you",
      "votes": null
    },
    {
      "id": "145561",
      "postDate": "11/19/2016 09:06:54",
      "content": "<p>Hi wacax</p>\n\n<p>If you means the competition terms it is accepted.</p>",
      "rawMarkdown": "Hi wacax\r\n\r\n If you means the competition terms it is accepted.",
      "votes": null
    },
    {
      "id": "170596",
      "postDate": "03/26/2017 16:20:37",
      "content": "<p>It is possible to download it via wget, if you necessarily must use that (eg, in the scenario where you have too ssh to some machine to download on the larger storage capacity).\nHow to do that is to first trigger a download on your browser - necessarily using your login credentials.</p>\n\n<p>Then copy the URL in the download page while it is downloading.   The HTTP exchange mechanism's final point is to let the browser encode all the cookie information into a single GET request, and thus all download information necessarily have to be in that URL, not in the browser environment (which is true in the case of POST request).   So that URL can be copied to any machines and trigger a download.   Then you can stop the download on the browser as soon as the wget start downloading. </p>",
      "rawMarkdown": "It is possible to download it via wget, if you necessarily must use that (eg, in the scenario where you have too ssh to some machine to download on the larger storage capacity).\nHow to do that is to first trigger a download on your browser - necessarily using your login credentials.\n\nThen copy the URL in the download page while it is downloading.   The HTTP exchange mechanism's final point is to let the browser encode all the cookie information into a single GET request, and thus all download information necessarily have to be in that URL, not in the browser environment (which is true in the case of POST request).   So that URL can be copied to any machines and trigger a download.   Then you can stop the download on the browser as soon as the wget start downloading.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 63584,
      "author_name": "abhishek",
      "author_url": "",
      "post_date": "02/04/2015 14:25:45",
      "content": "<p>download the cookies of kaggle and use the following:</p>\n<p>wget -x --load-cookies cookie_file.txt <em>url_to_dataset</em></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 63586,
      "author_name": "jaymoore",
      "author_url": "",
      "post_date": "02/04/2015 14:30:34",
      "content": "<p>Abishek, this is what I am trying to do. &nbsp;I use lynx to log in through Google, I get through the Google login fine, but it never redirects back to Kaggle, I just get a page back that says 'use a javascript enabled browser', and no kaggle cookies get stored in the cookie file. &nbsp;Are you managing to log in via Google in lynx?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 63587,
      "author_name": "abhishek",
      "author_url": "",
      "post_date": "02/04/2015 14:36:33",
      "content": "<p>Well, Im not using lynx. I also have google login to kaggle and the above command works perfectly fine for me.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 63588,
      "author_name": "jaymoore",
      "author_url": "",
      "post_date": "02/04/2015 14:46:29",
      "content": "<p>I found another way to get the cookies now, rather than lynx. &nbsp;Thanks anyway.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 63602,
      "author_name": "amikadewit1",
      "author_url": "",
      "post_date": "02/04/2015 19:38:36",
      "content": "<p>@Jay, could you please share your way? Thanks in advance :)&nbsp;</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 63603,
      "author_name": "wacaxx",
      "author_url": "",
      "post_date": "02/04/2015 19:40:35",
      "content": "<p>If you guys plan on using wget or similar tools the easiest way to get cookies is to login into your account and use</p>\n<p>https://addons.mozilla.org/en-US/firefox/addon/export-cookies/ on Firefox</p>\n<p>or</p>\n<p>https://chrome.google.com/webstore/detail/cookietxt-export/lopabhfecdfhgogdbojmaicoicjekelh on Chrome</p>\n<p>and then use what Abhishek suggested</p>\n<p>wget -x --load-cookies cookie_file.txt url_to_dataset</p>\n<p>Don't forget to accept the rules first otherwise it won't work</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 63628,
      "author_name": "jfkingiii",
      "author_url": "",
      "post_date": "02/05/2015 03:05:17",
      "content": "<p>Or copy as curl from Chrome:</p>\n<p>http://www.lornajane.net/posts/2013/chrome-feature-copy-as-curl</p>\n<p>delete the detritus after the long single quoted https string.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 63634,
      "author_name": "jaymoore",
      "author_url": "",
      "post_date": "02/05/2015 10:46:47",
      "content": "<p>@Amika I used Chrome, where I had logged into kaggle previously and accepted the terms of the competition. &nbsp;I went to Preferences/Settings and clicked Show advanced settings&nbsp;</p>\n<p>This brought up the Privacy section and I clicked Content Settings and then All cookies and site data.</p>\n<p>Then I searched for kaggle.com among the cookies list, and selected the cookie called .ASPXAUTH</p>\n<p>I took the information from this cookie, and pasted it into the cookies file that lynx had previously created (I had previously used lynx to try to do this, and had logged into Google through lynx). &nbsp;</p>\n<p>Then I called wget using this cookie file, and it worked first time, to my surprise :).</p>\n<p>Good luck!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 64323,
      "author_name": "udatta",
      "author_url": "",
      "post_date": "02/16/2015 08:29:13",
      "content": "<p>Hi</p>\n<p>Can anyone tell me how to resolve the error below? I downloaded the cookie file from chrome, and am logged into Kaggle using google...</p>\n<p>C:\\Program Files (x86)\\GnuWin32\\bin&gt;wget -x --load-cookies cookies.txt https<br>ww.kaggle.com/c/malware-classification/download/train.7z<br>SYSTEM_WGETRC = c:/progra~1/wget/etc/wgetrc<br>syswgetrc = C:\\Program Files (x86)\\GnuWin32/etc/wgetrc<br>--2015-02-16 15:57:14-- https://www.kaggle.com/c/malware-classification/dow<br>d/train.7z<br>Resolving www.kaggle.com... 168.62.224.124<br>Connecting to www.kaggle.com|168.62.224.124|:443... <strong>connected</strong>.<br><strong>ERROR: cannot verify www.kaggle.com's certificate, issued by `/C=US/O=GeoTru</strong><br><strong>nc./CN=RapidSSL SHA256 CA - G3':</strong><br> Unable to locally verify the issuer's authority.<br>To connect to www.kaggle.com insecurely, use `--no-check-certificate'.<br>Unable to establish SSL connection.</p>\n<p>Thanks!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 64324,
      "author_name": "abhishek",
      "author_url": "",
      "post_date": "02/16/2015 08:33:36",
      "content": "<p>did you try this:&nbsp;</p>\n\n<p>wget -x --load-cookies cookies.txt --no-check-certificate&nbsp;https://www.kaggle.com/c/malware-classification/download/train.7z</p>\n\n<p>or simply:</p>\n<p>wget -x --load-cookies cookies.txt http://www.kaggle.com/c/malware-classification/download/train.7z</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 64326,
      "author_name": "robertfontaine",
      "author_url": "",
      "post_date": "02/16/2015 08:39:47",
      "content": "<p>as above... after I got my nas working tonight... wget with the syntax above worked perfectly.</p>\n<p>Once I get nfs mounting properly I may even be able to unzip the archives.</p>\n<p>By the time everyone else has results I will have a working environment.&nbsp; :)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 64328,
      "author_name": "udatta",
      "author_url": "",
      "post_date": "02/16/2015 08:45:34",
      "content": "<p>Hey guys</p>\n\n<p>Thanks for the speedy reply....just got it working by adding<strong>&nbsp;--no-check-certificate</strong> to the cmd line ....but you guys replied before I can update my status! (Yes, you guys rock)</p>\n<p>@Robert-dont worry, there are enough noobs around (like me), who will take forever to set up an environment.....you are doing alright ;)</p>\n<p>(1% downloaded....yippiee :D )</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 64337,
      "author_name": "udatta",
      "author_url": "",
      "post_date": "02/16/2015 11:43:48",
      "content": "<p>eh..umm...my wget doesnt resume download from where it left off.....rather it is restarting download from 0%. (I use -c to continue)....can some1 help? Apologies for wasting your time ...I'm sure I'm doing something stupid....</p>\n\n<p>Edit: finally got the resume working..phew.. I had to use -O to change the name of the saved file for wget to continue saving to the same file</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 145556,
      "author_name": "maabdullah47",
      "author_url": "",
      "post_date": "11/19/2016 08:45:39",
      "content": "<p>[quote=wacax;63603]</p>\n<p>If you guys plan on using wget or similar tools the easiest way to get cookies is to login into your account and use</p>\n<p>https://addons.mozilla.org/en-US/firefox/addon/export-cookies/ on Firefox</p>\n<p>or</p>\n<p>https://chrome.google.com/webstore/detail/cookietxt-export/lopabhfecdfhgogdbojmaicoicjekelh on Chrome</p>\n<p>and then use what Abhishek suggested</p>\n<p>wget -x --load-cookies cookie_file.txt url_to_dataset</p>\n<p>Don't forget to accept the rules first otherwise it won't work</p>\n<p>[/quote]</p>\n\n<p>Hi Dear&nbsp;</p>\n<p>The text were you wrote :&nbsp;wget -x --load-cookies cookie_file.txt url_to_dataset</p>\n<p>Is it exactly the same text? if not exact text can you explain the parameters above?&nbsp;</p>\n<p>&nbsp;&nbsp;</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 145558,
      "author_name": "maabdullah47",
      "author_url": "",
      "post_date": "11/19/2016 09:00:32",
      "content": "<p>It is not work I go to the Directory where files should be located and I didn't found any thing.</p>\n\n<p>I'm sure the coockie.txt is saved on chrom .</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 145559,
      "author_name": "maabdullah47",
      "author_url": "",
      "post_date": "11/19/2016 09:02:57",
      "content": "<p>Hi wacax\nYou said : \"Don't forget to accept the rules first otherwise it won't work\"</p>\n\n<p>Where this rules dialog is displayed is it through lynx or on windows ? , need clarification.</p>\n\n<p>Thank you</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 145561,
      "author_name": "maabdullah47",
      "author_url": "",
      "post_date": "11/19/2016 09:06:54",
      "content": "<p>Hi wacax</p>\n\n<p>If you means the competition terms it is accepted.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 170596,
      "author_name": "tthtlc",
      "author_url": "",
      "post_date": "03/26/2017 16:20:37",
      "content": "<p>It is possible to download it via wget, if you necessarily must use that (eg, in the scenario where you have too ssh to some machine to download on the larger storage capacity).\nHow to do that is to first trigger a download on your browser - necessarily using your login credentials.</p>\n\n<p>Then copy the URL in the download page while it is downloading.   The HTTP exchange mechanism's final point is to let the browser encode all the cookie information into a single GET request, and thus all download information necessarily have to be in that URL, not in the browser environment (which is true in the case of POST request).   So that URL can be copied to any machines and trigger a download.   Then you can stop the download on the browser as soon as the wget start downloading. </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "63583": "",
    "63584": "",
    "63586": "",
    "63587": "",
    "63588": "",
    "63602": "",
    "63603": "",
    "63628": "",
    "63634": "",
    "64323": "",
    "64324": "",
    "64326": "",
    "64328": "",
    "64337": "",
    "145556": "",
    "145558": "It is not work I go to the Directory where files should be located and I didn't found any thing.\r\n\r\nI'm sure the coockie.txt is saved on chrom .",
    "145559": "Hi wacax\r\nYou said : \"Don't forget to accept the rules first otherwise it won't work\"\r\n\r\nWhere this rules dialog is displayed is it through lynx or on windows ? , need clarification.\r\n\r\nThank you",
    "145561": "Hi wacax\r\n\r\n If you means the competition terms it is accepted.",
    "170596": "It is possible to download it via wget, if you necessarily must use that (eg, in the scenario where you have too ssh to some machine to download on the larger storage capacity).\nHow to do that is to first trigger a download on your browser - necessarily using your login credentials.\n\nThen copy the URL in the download page while it is downloading.   The HTTP exchange mechanism's final point is to let the browser encode all the cookie information into a single GET request, and thus all download information necessarily have to be in that URL, not in the browser environment (which is true in the case of POST request).   So that URL can be copied to any machines and trigger a download.   Then you can stop the download on the browser as soon as the wget start downloading."
  },
  "source": "meta"
}