{
  "id": 12447,
  "title": "Any suggestions on downloading the dataset? ",
  "url": "/competitions/malware-classification/discussion/12447",
  "author_name": "",
  "post_date": "2015-02-09T02:18:29.937Z",
  "votes": null,
  "comment_count": 17,
  "views": 7072,
  "content": "<p>It is sooo HUGE that my chrome says &quot;7 days left&quot; for downloading this, but once I finish my work shift the download would be stop and I will never finish downloading it.&nbsp;</p>\n\n<p>Any tricks/ hacks for getting the downloading job done?&nbsp;</p>\n\n<p>Thanks!&nbsp;</p>",
  "messages": [
    {
      "id": "63790",
      "postDate": "02/09/2015 02:18:29",
      "content": "<p>It is sooo HUGE that my chrome says &quot;7 days left&quot; for downloading this, but once I finish my work shift the download would be stop and I will never finish downloading it.&nbsp;</p>\n\n<p>Any tricks/ hacks for getting the downloading job done?&nbsp;</p>\n\n<p>Thanks!&nbsp;</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "63794",
      "postDate": "02/09/2015 03:24:28",
      "content": "<p>OK,&nbsp;folks, it seems that the wget being discussed in the forum could do the continuing trick with the follwing command:&nbsp;</p>\n<p>wget -c url<br>wget --continue url<br>wget --continue [options] url</p>\n<p>I am going to give it a shot when the downloading need to be resumed&nbsp;after a black out.&nbsp;</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "63821",
      "postDate": "02/09/2015 18:36:10",
      "content": "<p>Get an AWS EC2 instance and their storage product. The dataset is more than likely too large to load into your machine's memory anyway if you are using R. &nbsp;</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "63884",
      "postDate": "02/10/2015 06:35:17",
      "content": "<p>Downloading to AWS instance took about 15 minutes. And I picked a cheap type of instance with the low network performance.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "63899",
      "postDate": "02/10/2015 09:07:11",
      "content": "<p>Downloading with <a href=\"http://aria2.sourceforge.net/\">Aria2</a>&nbsp;with&nbsp;10 concurrent connections upped my download speed from ~160kB/s to 1.5MB/s (although it oscillates much).</p>\n<p>-c can be used to resume downloading.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "63923",
      "postDate": "02/10/2015 14:52:11",
      "content": "<p>Hey all - Looks like the AWS instance is the most obvious choice, especially when there is more than one person working on the team. For those of you using this option, could you please help, byt letting me know:</p>\n<ol>\n<li>What type of instance you are using?</li>\n<li>Roughly what the cost (per month) is?</li>\n</ol>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "63997",
      "postDate": "02/11/2015 04:10:59",
      "content": "<p>[quote=jsardinha;63923]</p>\n<p>Hey all - Looks like the AWS instance is the most obvious choice, especially when there is more than one person working on the team. For those of you using this option, could you please help, byt letting me know:</p>\n<ol>\n<li>What type of instance you are using?</li>\n<li>Roughly what the cost (per month) is?</li>\n</ol>\n<p>[/quote]</p>\n\n<p>I took t2.medium which costs&nbsp;$0.052 per Hour = $39 per month or you can pay upfront and get it a little cheaper. Plus you'll have to add storage. For EBS General Purpose it is 0.1 per 1Gb per month. I took 1000Gb but it was an overkill :/ So in total that instance should cost me $139 per month.</p>\n<p>With t2.medium you get 4Gb of RAM but in reality there will be only 100-150Mb free, so you won't be able to run anything serious there. I'm planning to use this instance just to preprocess the files and extract features to create a small dataset.&nbsp;Will see how slow it's going to be ;)</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "64024",
      "postDate": "02/11/2015 13:52:13",
      "content": "<p>To expensive for me to play with AWS. &nbsp; &nbsp; Can I be a bottom feeder with R and ff and work from file on an ssd? &nbsp; &nbsp; Fitting the initial 17GB sample data in ram would be doable if I had more ram. &nbsp;I'm going to order more but the xeon phi promotion used up my allowance this week. &nbsp;:). &nbsp; &nbsp;Is there an R package that will run more effectively from a sql rdbms?</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "64026",
      "postDate": "02/11/2015 14:06:32",
      "content": "<p>I think Abishek's steer is right here. &nbsp;The thing that will work is to process the files one at a time, and extract some useful features from each file, and store them away. &nbsp;Once you have a matrix of files x features, it will measure in the MB rather than GB, then it's time for the other part of machine learning, the clever algorithms and cross-validation. &nbsp;If you don't have the free half TB to store the initial data, then it might be a case of using AWS for a few hours to expand the zip files, and extract the features. &nbsp;But before all that the first thing is to look at the sample data and decide what you think might be some good features, I guess!</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "64027",
      "postDate": "02/11/2015 14:07:57",
      "content": "<p>By the way you might like the R package sqldf, it lets you interrogate R data frames using SQL</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "64028",
      "postDate": "02/11/2015 14:12:12",
      "content": "<p>Also, you can take a look at spot instances in AWS. there is a risk of losing data (you can always backup) but prices are way cheaper...</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "64759",
      "postDate": "02/23/2015 10:34:31",
      "content": "<p>Well, frankly I could have done without the space wasted on running IDA on the binaries. I could do that if I thought there was *any* value to it. At least separate it out in to the 10GB it takes up.</p>\n\n<p>The other niggle is that it'd be much smaller if they were in binary rather than being ascii hex dumps.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "64867",
      "postDate": "02/25/2015 14:48:21",
      "content": "<p>After 17Gbs of downloading, wget provides me with a corrupted file, and 7z refuses to even try to open it. Not going to bother continuing with this competition. At least the others have had the decency to split their datasets into pieces.</p>\n<p>*Especially* if this was the result of some update of the files.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "66040",
      "postDate": "03/12/2015 14:08:40",
      "content": "<p>Can anyone list the url to download the data. I am trying to run the following&nbsp;</p>\n<p>wget&nbsp;http://www.kaggle.com/c/malware-classification/download/train.7z</p>\n<p>which is only downloading the html page</p>\n<p>Thanks</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "66045",
      "postDate": "03/12/2015 15:05:29",
      "content": "<p>[quote=Dipanjan Paul;66040]</p>\n<p>Can anyone list the url to download the data. I am trying to run the following&nbsp;</p>\n<p>wget&nbsp;http://www.kaggle.com/c/malware-classification/download/train.7z</p>\n<p>which is only downloading the html page</p>\n<p>Thanks</p>\n<p>[/quote]</p>\n<p>See&nbsp;http://www.kaggle.com/forums/f/15/kaggle-forum/t/6604/downloading-data-via-command-line</p>\n<p>(You also need to save and use a cookie).</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "66057",
      "postDate": "03/12/2015 19:25:28",
      "content": "<p>Hi there,</p>\n<p>I have fully downloaded the dataset using the Free Download Manager on a Windows machine.</p>\n<p>1- I have logged in the system using my browser, e.g. Firefox,</p>\n<p>2- I have clicked on the link, accepted the terms, and then the link was automatically added to FDM,</p>\n<p>3- Unfortunately, at some point in downloading the data set, the switch connecting us to the Internet went down :((((( but the download was resumable :)))))</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "66066",
      "postDate": "03/12/2015 21:53:10",
      "content": "<p>Thanks Triskelion, I will try it. I do want it to download using CLI into a cloud,..</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "66936",
      "postDate": "03/18/2015 10:42:53",
      "content": "<p>Hi,</p>\n<p>I got a 200$ free trial of microsoft azure services for this competition. I have knowledge of r and sql. I have no idea what services i will need to complete this project. Can anyone suggest me&nbsp;services i need to complete this project.</p>\n<p>Thanks,</p>\n<p>Ajay</p>",
      "rawMarkdown": "",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 63794,
      "author_name": "oliverhan",
      "author_url": "",
      "post_date": "02/09/2015 03:24:28",
      "content": "<p>OK,&nbsp;folks, it seems that the wget being discussed in the forum could do the continuing trick with the follwing command:&nbsp;</p>\n<p>wget -c url<br>wget --continue url<br>wget --continue [options] url</p>\n<p>I am going to give it a shot when the downloading need to be resumed&nbsp;after a black out.&nbsp;</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 63821,
      "author_name": "shaynewang",
      "author_url": "",
      "post_date": "02/09/2015 18:36:10",
      "content": "<p>Get an AWS EC2 instance and their storage product. The dataset is more than likely too large to load into your machine's memory anyway if you are using R. &nbsp;</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 63884,
      "author_name": "yankov",
      "author_url": "",
      "post_date": "02/10/2015 06:35:17",
      "content": "<p>Downloading to AWS instance took about 15 minutes. And I picked a cheap type of instance with the low network performance.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 63899,
      "author_name": "pzawal",
      "author_url": "",
      "post_date": "02/10/2015 09:07:11",
      "content": "<p>Downloading with <a href=\"http://aria2.sourceforge.net/\">Aria2</a>&nbsp;with&nbsp;10 concurrent connections upped my download speed from ~160kB/s to 1.5MB/s (although it oscillates much).</p>\n<p>-c can be used to resume downloading.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 63923,
      "author_name": "jsardinha",
      "author_url": "",
      "post_date": "02/10/2015 14:52:11",
      "content": "<p>Hey all - Looks like the AWS instance is the most obvious choice, especially when there is more than one person working on the team. For those of you using this option, could you please help, byt letting me know:</p>\n<ol>\n<li>What type of instance you are using?</li>\n<li>Roughly what the cost (per month) is?</li>\n</ol>",
      "votes": null,
      "replies": []
    },
    {
      "id": 63997,
      "author_name": "yankov",
      "author_url": "",
      "post_date": "02/11/2015 04:10:59",
      "content": "<p>[quote=jsardinha;63923]</p>\n<p>Hey all - Looks like the AWS instance is the most obvious choice, especially when there is more than one person working on the team. For those of you using this option, could you please help, byt letting me know:</p>\n<ol>\n<li>What type of instance you are using?</li>\n<li>Roughly what the cost (per month) is?</li>\n</ol>\n<p>[/quote]</p>\n\n<p>I took t2.medium which costs&nbsp;$0.052 per Hour = $39 per month or you can pay upfront and get it a little cheaper. Plus you'll have to add storage. For EBS General Purpose it is 0.1 per 1Gb per month. I took 1000Gb but it was an overkill :/ So in total that instance should cost me $139 per month.</p>\n<p>With t2.medium you get 4Gb of RAM but in reality there will be only 100-150Mb free, so you won't be able to run anything serious there. I'm planning to use this instance just to preprocess the files and extract features to create a small dataset.&nbsp;Will see how slow it's going to be ;)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 64024,
      "author_name": "robertfontaine",
      "author_url": "",
      "post_date": "02/11/2015 13:52:13",
      "content": "<p>To expensive for me to play with AWS. &nbsp; &nbsp; Can I be a bottom feeder with R and ff and work from file on an ssd? &nbsp; &nbsp; Fitting the initial 17GB sample data in ram would be doable if I had more ram. &nbsp;I'm going to order more but the xeon phi promotion used up my allowance this week. &nbsp;:). &nbsp; &nbsp;Is there an R package that will run more effectively from a sql rdbms?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 64026,
      "author_name": "jaymoore",
      "author_url": "",
      "post_date": "02/11/2015 14:06:32",
      "content": "<p>I think Abishek's steer is right here. &nbsp;The thing that will work is to process the files one at a time, and extract some useful features from each file, and store them away. &nbsp;Once you have a matrix of files x features, it will measure in the MB rather than GB, then it's time for the other part of machine learning, the clever algorithms and cross-validation. &nbsp;If you don't have the free half TB to store the initial data, then it might be a case of using AWS for a few hours to expand the zip files, and extract the features. &nbsp;But before all that the first thing is to look at the sample data and decide what you think might be some good features, I guess!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 64027,
      "author_name": "jaymoore",
      "author_url": "",
      "post_date": "02/11/2015 14:07:57",
      "content": "<p>By the way you might like the R package sqldf, it lets you interrogate R data frames using SQL</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 64028,
      "author_name": "abhishek",
      "author_url": "",
      "post_date": "02/11/2015 14:12:12",
      "content": "<p>Also, you can take a look at spot instances in AWS. there is a risk of losing data (you can always backup) but prices are way cheaper...</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 64759,
      "author_name": "",
      "author_url": "",
      "post_date": "02/23/2015 10:34:31",
      "content": "<p>Well, frankly I could have done without the space wasted on running IDA on the binaries. I could do that if I thought there was *any* value to it. At least separate it out in to the 10GB it takes up.</p>\n\n<p>The other niggle is that it'd be much smaller if they were in binary rather than being ascii hex dumps.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 64867,
      "author_name": "",
      "author_url": "",
      "post_date": "02/25/2015 14:48:21",
      "content": "<p>After 17Gbs of downloading, wget provides me with a corrupted file, and 7z refuses to even try to open it. Not going to bother continuing with this competition. At least the others have had the decency to split their datasets into pieces.</p>\n<p>*Especially* if this was the result of some update of the files.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 66040,
      "author_name": "dipanjanpaul",
      "author_url": "",
      "post_date": "03/12/2015 14:08:40",
      "content": "<p>Can anyone list the url to download the data. I am trying to run the following&nbsp;</p>\n<p>wget&nbsp;http://www.kaggle.com/c/malware-classification/download/train.7z</p>\n<p>which is only downloading the html page</p>\n<p>Thanks</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 66045,
      "author_name": "triskelion",
      "author_url": "",
      "post_date": "03/12/2015 15:05:29",
      "content": "<p>[quote=Dipanjan Paul;66040]</p>\n<p>Can anyone list the url to download the data. I am trying to run the following&nbsp;</p>\n<p>wget&nbsp;http://www.kaggle.com/c/malware-classification/download/train.7z</p>\n<p>which is only downloading the html page</p>\n<p>Thanks</p>\n<p>[/quote]</p>\n<p>See&nbsp;http://www.kaggle.com/forums/f/15/kaggle-forum/t/6604/downloading-data-via-command-line</p>\n<p>(You also need to save and use a cookie).</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 66057,
      "author_name": "alishakiba",
      "author_url": "",
      "post_date": "03/12/2015 19:25:28",
      "content": "<p>Hi there,</p>\n<p>I have fully downloaded the dataset using the Free Download Manager on a Windows machine.</p>\n<p>1- I have logged in the system using my browser, e.g. Firefox,</p>\n<p>2- I have clicked on the link, accepted the terms, and then the link was automatically added to FDM,</p>\n<p>3- Unfortunately, at some point in downloading the data set, the switch connecting us to the Internet went down :((((( but the download was resumable :)))))</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 66066,
      "author_name": "dipanjanpaul",
      "author_url": "",
      "post_date": "03/12/2015 21:53:10",
      "content": "<p>Thanks Triskelion, I will try it. I do want it to download using CLI into a cloud,..</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 66936,
      "author_name": "runforrestrun",
      "author_url": "",
      "post_date": "03/18/2015 10:42:53",
      "content": "<p>Hi,</p>\n<p>I got a 200$ free trial of microsoft azure services for this competition. I have knowledge of r and sql. I have no idea what services i will need to complete this project. Can anyone suggest me&nbsp;services i need to complete this project.</p>\n<p>Thanks,</p>\n<p>Ajay</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "63790": "",
    "63794": "",
    "63821": "",
    "63884": "",
    "63899": "",
    "63923": "",
    "63997": "",
    "64024": "",
    "64026": "",
    "64027": "",
    "64028": "",
    "64759": "",
    "64867": "",
    "66040": "",
    "66045": "",
    "66057": "",
    "66066": "",
    "66936": ""
  },
  "source": "meta"
}