{
  "id": 6060,
  "title": "download data",
  "url": "/competitions/yandex-personalized-web-search-challenge/discussion/6060",
  "author_name": "",
  "post_date": "2013-10-20T12:17:59.837Z",
  "votes": 1,
  "comment_count": 15,
  "views": 5844,
  "content": "<p>how did you manage to download the data? i tried multiple time via browser or wget, but without success...</p>",
  "messages": [
    {
      "id": "32459",
      "postDate": "10/20/2013 12:17:59",
      "content": "<p>how did you manage to download the data? i tried multiple time via browser or wget, but without success...</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "32460",
      "postDate": "10/20/2013 12:56:01",
      "content": "<p>I used DownThemAll!, a firefox extension quite similar to a download manager.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "32465",
      "postDate": "10/20/2013 16:47:10",
      "content": "<p>Many thanks for this hint. I will give it a try</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "32473",
      "postDate": "10/21/2013 02:36:13",
      "content": "<p>Thanks so much EGO for the tip on using &quot;DownLoadThemAll!&quot;. I am in Australia and had tried half a dozen times unsuccessfully to download the data but finally got it working with this tool.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "32734",
      "postDate": "10/27/2013 11:07:40",
      "content": "<p>Hi,</p>\n<p>I cannot download the dataset due to two problems: download managers cannot resume and I have low internet connection to your site. &nbsp;Can you please provide an ftp or Torent link for data?</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "32745",
      "postDate": "10/27/2013 19:50:30",
      "content": "<p>I am not sure who you are asking for ftp or torrent link to. In my case, considering I have small upload speed and the time I can keep my computer on, it would take more time than the rest of the competition. So, it does make sense. Hope admins can help you.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "32776",
      "postDate": "10/28/2013 13:46:34",
      "content": "<p>What I was going to try, is to use wget in screen on a Linux server. But then I need the direct link to the train.gz file, and apparently this is not it:&nbsp;https://www.kaggle.com/c/yandex-personalized-web-search-challenge/download/train.gz Did anyone find the direct link to the file?</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "32777",
      "postDate": "10/28/2013 15:00:14",
      "content": "<p>We don't have support for torrents at this time. This suggestion comes up often, but torrents are not an ideal mechanism for us for two reasons:</p>\n<ul>\n<li><span style=\"line-height: 1.4\">We need to be able to control the data source. We often have to re-release modified datasets and don't want outdated torrents polluting the web and confusing people</span></li>\n<li><span style=\"line-height: 1.4\">We enforce that you accept the rules before downloading. With torrents this is more difficult.</span></li>\n</ul>\n<p>You should be able to resume downloads for up to 3 days after starting them, regardless of browser. There may be combinations of browsers/managers where this doesn't work, but it <em>should</em> work in most cases.</p>\n<p><strong>If you want to use a server or the command line to download a file, you must&nbsp;export your Kaggle cookies from your browser</strong> (this <a href=\"https://chrome.google.com/webstore/detail/lopabhfecdfhgogdbojmaicoicjekelh\">chrome extension</a> is the easiest way) and then call wget's --load-cookies option. &nbsp;Because of the rules clause above, we cannot have naked download links - you need to be logged in and have accepted the rules. After passing your Kaggle cookies, wget should work fine. &nbsp;Note that the file download links will redirect to something like &quot;https://kaggle2.blob.core.windows.net&quot; after you've clicked the download link. This is the URL you should give to wget.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "32823",
      "postDate": "10/29/2013 11:21:23",
      "content": "<p>Thanks, I succeeded in downloading the data!</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "33738",
      "postDate": "11/15/2013 04:48:39",
      "content": "<p>How can one open this dataset file?</p>\n<p>Any software or do we have to write a code?</p>\n<p>I am new in this field so might ask stupid questions!</p>\n<p>Please help</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "33743",
      "postDate": "11/15/2013 07:52:45",
      "content": "<p>You can use gunzip on linux or 7-zip on Windows to extract the .gz files.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "33755",
      "postDate": "11/15/2013 16:30:06",
      "content": "<p>[quote=Suzan Verberne;33743]</p>\n<p>You can use gunzip on linux or 7-zip on Windows to extract the .gz files.</p>\n<p>[/quote]</p>\n<p>&nbsp;</p>\n<p>I have extracted the file using Winrar. But now how can I manipulate or use this data.</p>\n<p>Kindly help.</p>\n<p>I am doing a research on click modelling.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "33790",
      "postDate": "11/16/2013 05:00:47",
      "content": "<p>[quote=William Cukierski;32777]</p>\n<p><span style=\"line-height: 1.4\">Note that the file download links will redirect to something like &quot;https://kaggle2.blob.core.windows.net&quot; after you've clicked the download link. This is the URL you should give to wget.</span></p>\n<p>[/quote]</p>\n<p>&nbsp;</p>\n<p>Just a quick note to others in this situation - using the &quot;.windows.net&quot; addresses gave me 404's for some reason, but using the &quot;http://www.kaggle.com/c/yandex-personalized-web-search-challenge/download/xyz.gz&quot; urls from the download page did work.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "34849",
      "postDate": "11/18/2013 10:00:57",
      "content": "<p>[quote=Sheikh Adnan Ahmed Usmani;33755]</p>\n<p>[quote=Suzan Verberne;33743]</p>\n<p>You can use gunzip on linux or 7-zip on Windows to extract the .gz files.</p>\n<p>[/quote]</p>\n<p>&nbsp;</p>\n<p>I have extracted the file using Winrar. But now how can I manipulate or use this data.</p>\n<p>Kindly help.</p>\n<p>I am doing a research on click modelling.</p>\n<p>[/quote]</p>\n<p>&nbsp;</p>\n<p>I am not sure what kind of answer you expect. Here is a description of the data:&nbsp;https://www.kaggle.com/c/yandex-personalized-web-search-challenge/details/logs-format&nbsp;</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "36061",
      "postDate": "12/11/2013 12:53:31",
      "content": "<p><span style=\"line-height: 1.4\">I frequently run into the wget problem, and I am tired of relying on workarounds each time. The reason for this is that - though I *can wait* &nbsp;for a disproportionate time for the 1st download, the next time round when I run my code against another machine, I need the train file again - and having to transfer the (huge) train file from one machine to another is pain !</span></p>\n<p>Here's what I am doing. Would appreciate if someone could point out what am I missing ?&nbsp;</p>\n<p><span style=\"line-height: 1.4\"># Log in to the server and save the cookies the traditional way- this can also be done by Chrome extension as mentioned by @</span><strong style=\"line-height: 1.4\">William Cukierski</strong><span style=\"line-height: 1.4\">&nbsp; above.</span></p>\n<p><span style=\"line-height: 1.4\"># username &amp; pwd masked obviously</span></p>\n<p><span style=\"line-height: 1.4\">wget --save-cookies cookies.txt&nbsp;</span>--post-data 'user=masked_kaggle_email_address&amp;password=masked'&nbsp;http://www.kaggle.com/</p>\n<p><br># Grab the download page <br> wget --load-cookies cookies.txt \\<br> -p <em>http://www.kaggle.com/c/yandex-personalized-web-search-challenge/forums/t/6060/download-data</em></p>\n<p>I Also tried the link <em>https://kaggle2.blob.core.windows.net</em> and <em>http://www.kaggle.com/c/yandex-personalized-web-search-challenge/download/train.gz</em> - neither of them works for downloading the data-set. <br><strong>All this is saving is some js, html &amp; png files - what am I doing wrong ?</strong><br>following is the structure of the folders it is saving , and the files thereof (See the link to the saved files)&nbsp;</p>\n<p><a href=\"http://findsomethingnewtoday.files.wordpress.com/2013/12/kaggle_yandex.png\">Link to directory structure fetched using wget with load-cookies</a>&nbsp;</p>\n<p>&nbsp;</p>\n<p>PPS : On giving the redirect link as&nbsp;https://kaggle2.blob.core.windows.net , I see the following.&nbsp;</p>\n<p><span style=\"line-height: 1.4\">-- https://kaggle2.blob.core.windows.net/</span><br>Resolving kaggle2.blob.core.windows.net (kaggle2.blob.core.windows.net)... 65.52.106.46<br>Connecting to kaggle2.blob.core.windows.net (kaggle2.blob.core.windows.net)|65.52.106.46|:443... connected.<br>HTTP request sent, awaiting response... 400 Value for one of the query parameters specified in the request URI is invalid.<br>2013-12-11 18:18:41 ERROR 400: Value for one of the query parameters specified in the request URI is invalid..</p>\n<p>&nbsp;</p>\n<p>btw, if I plainly try to access the link -&nbsp;https://kaggle2.blob.core.windows.net/ - it reads &quot;This XML file does not appear to have any style information associated with it. The document tree is shown below.&quot;&nbsp;</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "36072",
      "postDate": "12/11/2013 15:57:21",
      "content": "<p>&nbsp;use DownloadThemAll!, a firefox extension quite similar to a download manager. It will take five to six hours to get dowlnoad.</p>\n<p>But extracting your downloaded train file will take exactly 14 to 15 minutes :)</p>",
      "rawMarkdown": "",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 32460,
      "author_name": "emiliogortizg",
      "author_url": "",
      "post_date": "10/20/2013 12:56:01",
      "content": "<p>I used DownThemAll!, a firefox extension quite similar to a download manager.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 32465,
      "author_name": "mabehles",
      "author_url": "",
      "post_date": "10/20/2013 16:47:10",
      "content": "<p>Many thanks for this hint. I will give it a try</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 32473,
      "author_name": "soates",
      "author_url": "",
      "post_date": "10/21/2013 02:36:13",
      "content": "<p>Thanks so much EGO for the tip on using &quot;DownLoadThemAll!&quot;. I am in Australia and had tried half a dozen times unsuccessfully to download the data but finally got it working with this tool.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 32734,
      "author_name": "ashkansami",
      "author_url": "",
      "post_date": "10/27/2013 11:07:40",
      "content": "<p>Hi,</p>\n<p>I cannot download the dataset due to two problems: download managers cannot resume and I have low internet connection to your site. &nbsp;Can you please provide an ftp or Torent link for data?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 32745,
      "author_name": "emiliogortizg",
      "author_url": "",
      "post_date": "10/27/2013 19:50:30",
      "content": "<p>I am not sure who you are asking for ftp or torrent link to. In my case, considering I have small upload speed and the time I can keep my computer on, it would take more time than the rest of the competition. So, it does make sense. Hope admins can help you.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 32776,
      "author_name": "suzanv",
      "author_url": "",
      "post_date": "10/28/2013 13:46:34",
      "content": "<p>What I was going to try, is to use wget in screen on a Linux server. But then I need the direct link to the train.gz file, and apparently this is not it:&nbsp;https://www.kaggle.com/c/yandex-personalized-web-search-challenge/download/train.gz Did anyone find the direct link to the file?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 32777,
      "author_name": "wcukierski",
      "author_url": "",
      "post_date": "10/28/2013 15:00:14",
      "content": "<p>We don't have support for torrents at this time. This suggestion comes up often, but torrents are not an ideal mechanism for us for two reasons:</p>\n<ul>\n<li><span style=\"line-height: 1.4\">We need to be able to control the data source. We often have to re-release modified datasets and don't want outdated torrents polluting the web and confusing people</span></li>\n<li><span style=\"line-height: 1.4\">We enforce that you accept the rules before downloading. With torrents this is more difficult.</span></li>\n</ul>\n<p>You should be able to resume downloads for up to 3 days after starting them, regardless of browser. There may be combinations of browsers/managers where this doesn't work, but it <em>should</em> work in most cases.</p>\n<p><strong>If you want to use a server or the command line to download a file, you must&nbsp;export your Kaggle cookies from your browser</strong> (this <a href=\"https://chrome.google.com/webstore/detail/lopabhfecdfhgogdbojmaicoicjekelh\">chrome extension</a> is the easiest way) and then call wget's --load-cookies option. &nbsp;Because of the rules clause above, we cannot have naked download links - you need to be logged in and have accepted the rules. After passing your Kaggle cookies, wget should work fine. &nbsp;Note that the file download links will redirect to something like &quot;https://kaggle2.blob.core.windows.net&quot; after you've clicked the download link. This is the URL you should give to wget.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 32823,
      "author_name": "suzanv",
      "author_url": "",
      "post_date": "10/29/2013 11:21:23",
      "content": "<p>Thanks, I succeeded in downloading the data!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 33738,
      "author_name": "sheikhadnanahmedusmani",
      "author_url": "",
      "post_date": "11/15/2013 04:48:39",
      "content": "<p>How can one open this dataset file?</p>\n<p>Any software or do we have to write a code?</p>\n<p>I am new in this field so might ask stupid questions!</p>\n<p>Please help</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 33743,
      "author_name": "suzanv",
      "author_url": "",
      "post_date": "11/15/2013 07:52:45",
      "content": "<p>You can use gunzip on linux or 7-zip on Windows to extract the .gz files.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 33755,
      "author_name": "sheikhadnanahmedusmani",
      "author_url": "",
      "post_date": "11/15/2013 16:30:06",
      "content": "<p>[quote=Suzan Verberne;33743]</p>\n<p>You can use gunzip on linux or 7-zip on Windows to extract the .gz files.</p>\n<p>[/quote]</p>\n<p>&nbsp;</p>\n<p>I have extracted the file using Winrar. But now how can I manipulate or use this data.</p>\n<p>Kindly help.</p>\n<p>I am doing a research on click modelling.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 33790,
      "author_name": "richardwalsh",
      "author_url": "",
      "post_date": "11/16/2013 05:00:47",
      "content": "<p>[quote=William Cukierski;32777]</p>\n<p><span style=\"line-height: 1.4\">Note that the file download links will redirect to something like &quot;https://kaggle2.blob.core.windows.net&quot; after you've clicked the download link. This is the URL you should give to wget.</span></p>\n<p>[/quote]</p>\n<p>&nbsp;</p>\n<p>Just a quick note to others in this situation - using the &quot;.windows.net&quot; addresses gave me 404's for some reason, but using the &quot;http://www.kaggle.com/c/yandex-personalized-web-search-challenge/download/xyz.gz&quot; urls from the download page did work.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 34849,
      "author_name": "suzanv",
      "author_url": "",
      "post_date": "11/18/2013 10:00:57",
      "content": "<p>[quote=Sheikh Adnan Ahmed Usmani;33755]</p>\n<p>[quote=Suzan Verberne;33743]</p>\n<p>You can use gunzip on linux or 7-zip on Windows to extract the .gz files.</p>\n<p>[/quote]</p>\n<p>&nbsp;</p>\n<p>I have extracted the file using Winrar. But now how can I manipulate or use this data.</p>\n<p>Kindly help.</p>\n<p>I am doing a research on click modelling.</p>\n<p>[/quote]</p>\n<p>&nbsp;</p>\n<p>I am not sure what kind of answer you expect. Here is a description of the data:&nbsp;https://www.kaggle.com/c/yandex-personalized-web-search-challenge/details/logs-format&nbsp;</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 36061,
      "author_name": "ektagrover1",
      "author_url": "",
      "post_date": "12/11/2013 12:53:31",
      "content": "<p><span style=\"line-height: 1.4\">I frequently run into the wget problem, and I am tired of relying on workarounds each time. The reason for this is that - though I *can wait* &nbsp;for a disproportionate time for the 1st download, the next time round when I run my code against another machine, I need the train file again - and having to transfer the (huge) train file from one machine to another is pain !</span></p>\n<p>Here's what I am doing. Would appreciate if someone could point out what am I missing ?&nbsp;</p>\n<p><span style=\"line-height: 1.4\"># Log in to the server and save the cookies the traditional way- this can also be done by Chrome extension as mentioned by @</span><strong style=\"line-height: 1.4\">William Cukierski</strong><span style=\"line-height: 1.4\">&nbsp; above.</span></p>\n<p><span style=\"line-height: 1.4\"># username &amp; pwd masked obviously</span></p>\n<p><span style=\"line-height: 1.4\">wget --save-cookies cookies.txt&nbsp;</span>--post-data 'user=masked_kaggle_email_address&amp;password=masked'&nbsp;http://www.kaggle.com/</p>\n<p><br># Grab the download page <br> wget --load-cookies cookies.txt \\<br> -p <em>http://www.kaggle.com/c/yandex-personalized-web-search-challenge/forums/t/6060/download-data</em></p>\n<p>I Also tried the link <em>https://kaggle2.blob.core.windows.net</em> and <em>http://www.kaggle.com/c/yandex-personalized-web-search-challenge/download/train.gz</em> - neither of them works for downloading the data-set. <br><strong>All this is saving is some js, html &amp; png files - what am I doing wrong ?</strong><br>following is the structure of the folders it is saving , and the files thereof (See the link to the saved files)&nbsp;</p>\n<p><a href=\"http://findsomethingnewtoday.files.wordpress.com/2013/12/kaggle_yandex.png\">Link to directory structure fetched using wget with load-cookies</a>&nbsp;</p>\n<p>&nbsp;</p>\n<p>PPS : On giving the redirect link as&nbsp;https://kaggle2.blob.core.windows.net , I see the following.&nbsp;</p>\n<p><span style=\"line-height: 1.4\">-- https://kaggle2.blob.core.windows.net/</span><br>Resolving kaggle2.blob.core.windows.net (kaggle2.blob.core.windows.net)... 65.52.106.46<br>Connecting to kaggle2.blob.core.windows.net (kaggle2.blob.core.windows.net)|65.52.106.46|:443... connected.<br>HTTP request sent, awaiting response... 400 Value for one of the query parameters specified in the request URI is invalid.<br>2013-12-11 18:18:41 ERROR 400: Value for one of the query parameters specified in the request URI is invalid..</p>\n<p>&nbsp;</p>\n<p>btw, if I plainly try to access the link -&nbsp;https://kaggle2.blob.core.windows.net/ - it reads &quot;This XML file does not appear to have any style information associated with it. The document tree is shown below.&quot;&nbsp;</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 36072,
      "author_name": "parthibangowthaman",
      "author_url": "",
      "post_date": "12/11/2013 15:57:21",
      "content": "<p>&nbsp;use DownloadThemAll!, a firefox extension quite similar to a download manager. It will take five to six hours to get dowlnoad.</p>\n<p>But extracting your downloaded train file will take exactly 14 to 15 minutes :)</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "32459": "",
    "32460": "",
    "32465": "",
    "32473": "",
    "32734": "",
    "32745": "",
    "32776": "",
    "32777": "",
    "32823": "",
    "33738": "",
    "33743": "",
    "33755": "",
    "33790": "",
    "34849": "",
    "36061": "",
    "36072": ""
  },
  "source": "meta"
}