{
  "id": 12468,
  "title": "Complete Newbie Questions...",
  "url": "/competitions/malware-classification/discussion/12468",
  "author_name": "",
  "post_date": "2015-02-10T19:12:30.317Z",
  "votes": null,
  "comment_count": 12,
  "views": 3120,
  "content": "<p>I'm currently collecting some hardware so that I can play along but I thought it would be good to start thinking about these. &nbsp;I drew a picture on the window at work this morning and tried using a little intuition.</p>\n<p>This one is interesting before any analysis takes place. &nbsp; &nbsp;Ignoring what the appropriate algorithms for discovering features I'd like to start with...</p>\n<p>Big Data, Small Compute: &nbsp;What the !@##$ do you do with 400GB of data.</p>\n<p>Compression and transformations that make this a manageable size but are inherently lossy.</p>\n<p>I can't get all this data into main memory or&nbsp;gpu&nbsp;memory, and I don't have the money to throw it all into one big data structure in the cloud so...</p>\n<p>1. How do I resolve an appropriate sample size? &nbsp; Can I measure the randomness of the data and get some kind of metric? &nbsp;&nbsp;</p>\n<p>2. Can I get some gross measures from the data that let me categorize semi-intelligently?</p>\n<p>3. Given that memory, processor, time, and effort &nbsp;are all&nbsp;constrained I could&nbsp;start with small random sample and grow them progressively and see what falls out. &nbsp; Take what falls out and massage it in and see if anything else falls out when you mix them together.</p>\n<p>Not quite spray and pray but not terribly deterministic either...</p>\n\n<p>Am I asking the right questions?</p>\n<p>...&nbsp;</p>",
  "messages": [
    {
      "id": "63953",
      "postDate": "02/10/2015 19:12:30",
      "content": "<p>I'm currently collecting some hardware so that I can play along but I thought it would be good to start thinking about these. &nbsp;I drew a picture on the window at work this morning and tried using a little intuition.</p>\n<p>This one is interesting before any analysis takes place. &nbsp; &nbsp;Ignoring what the appropriate algorithms for discovering features I'd like to start with...</p>\n<p>Big Data, Small Compute: &nbsp;What the !@##$ do you do with 400GB of data.</p>\n<p>Compression and transformations that make this a manageable size but are inherently lossy.</p>\n<p>I can't get all this data into main memory or&nbsp;gpu&nbsp;memory, and I don't have the money to throw it all into one big data structure in the cloud so...</p>\n<p>1. How do I resolve an appropriate sample size? &nbsp; Can I measure the randomness of the data and get some kind of metric? &nbsp;&nbsp;</p>\n<p>2. Can I get some gross measures from the data that let me categorize semi-intelligently?</p>\n<p>3. Given that memory, processor, time, and effort &nbsp;are all&nbsp;constrained I could&nbsp;start with small random sample and grow them progressively and see what falls out. &nbsp; Take what falls out and massage it in and see if anything else falls out when you mix them together.</p>\n<p>Not quite spray and pray but not terribly deterministic either...</p>\n\n<p>Am I asking the right questions?</p>\n<p>...&nbsp;</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "63977",
      "postDate": "02/10/2015 22:29:16",
      "content": "<p>[quote=Robert Fontaine;63953]</p>\n<p>Am I asking the right questions?</p>\n<p>...&nbsp;</p>\n<p>[/quote]</p>\n<p>nop!</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "64022",
      "postDate": "02/11/2015 13:46:00",
      "content": "<p>Thanks :),</p>\n<p>Well then how about SVM ,R, ff?</p>\n<p>I'm not&nbsp;sure that&nbsp;loading the sample data into a dbms&nbsp;buys me anything for the solution but I have to establish a development environment so I will probably install and configure an rdbms tonight &nbsp;mysql/progres. &nbsp;</p>\n<p>Hopefully,&nbsp;kaggle won't go out of business before I learn how to data mine.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "64023",
      "postDate": "02/11/2015 13:49:19",
      "content": "<p>you dont have to load any data to dbms or anything. create a script that mines information from individual files and writes it to a separate file. getting a logloss of less than 0.1 will take you only a couple of hours using R or python.</p>\n<p>P.S. I dont think kaggle will go out of business anytime soon :)</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "64032",
      "postDate": "02/11/2015 14:45:33",
      "content": "<p>[quote=Abhishek;64023]</p>\n<p>you dont have to load any data to dbms or anything. create a script that mines information from individual files and writes it to a separate file. getting a logloss of less than 0.1 will take you only a couple of hours using R or python.</p>\n<p>P.S. I dont think kaggle will go out of business anytime soon :)</p>\n<p>[/quote]</p>\n<p>Thank you for your post. This is the first time that I cope with such a large volume data. Is it&nbsp;necessary to unzip the file that downloaded from kaggle? If so, I wouldn't have so much disk space : (</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "64035",
      "postDate": "02/11/2015 15:10:08",
      "content": "<p>I dont know if you can do that with 7z. I had enough space to decompress the dataset.&nbsp;</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "64068",
      "postDate": "02/11/2015 21:21:04",
      "content": "<p>You can directly pipe files out of the archive. I tried that and it works, you need something along the line of &quot;/cygdrive/c/Program\\ Files/7-Zip/7z.exe e -so dataSample.7z $FILEHASH.asm | grep 'interessing stuff'&quot;, but it's horible slow as you can not have random&nbsp;access in a 7z archive (http://stackoverflow.com/questions/4457997/indexing-random-access-to-7zip-7z-archives).</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "64069",
      "postDate": "02/11/2015 21:26:32",
      "content": "<p>Seagate 1TB Barracuda is&nbsp;about $60CAD. &nbsp; Not terribly good for reliability but a stack of them striped should be able to max out pcie 3.0 bandwidth. &nbsp; I'm going to grab 1 or 2 this weekend. &nbsp; &nbsp; A Multiple ZFS groups of Hitachis in RAIDZ2 or better are far more reliable but data is disposable on my test bench and I can't afford the number of disks that you need to stripe across groups with ZFS.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "64112",
      "postDate": "02/12/2015 13:35:09",
      "content": "<p>Hi,</p>\n<p>I have downloaded and unzipped the dataSample.7z. The folder has 4 files of<span>.</span>bytes and<span> .</span><span>asm</span> types. As two files have the same hash so I concluded that both are different format of the same file.</p>\n<p>Do the train and test folder also have the same types of files?</p>\n<p>Due to storage limitations, I am not unzipping those two zipped files. If they also have both formats, then, can I work with only one format?</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "64179",
      "postDate": "02/13/2015 18:58:44",
      "content": "<p>If space is a concern, you can always unzip portions of the training set as needed.</p>\n<p>.asm files usually contain everything that's in .bytes + extras</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "64379",
      "postDate": "02/17/2015 03:01:44",
      "content": "<p>8 cheap 256TB SSDs on a raid card would make this a lot quicker.&nbsp;&nbsp; Just unzipping on my little ARM based NAS is likely to take a few days.&nbsp;&nbsp; Good thing I have lots of math to study.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "64380",
      "postDate": "02/17/2015 04:41:04",
      "content": "<p>Before you spend money on fancy hardware, please re-read carefully what Abhishek said. You don't need SSDs, RAID, data striping or anything like that. A standard laptop with 4 GB RAM and 500 GB of free hard drive space is all you need to process the data and generate a submission in a couple of hours of computing time.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "64383",
      "postDate": "02/17/2015 05:16:44",
      "content": "<p>Understood but it's nice to dream... I've got 2TB on an old dns323 nas with 2TB of disk at the moment using a 500mhz Marvell ARM and 64MB of RAM.&nbsp;&nbsp;&nbsp;&nbsp; On the plus side the 7zip utility doesn't seem to be bleeding memory so in 4 or 5 days if the NAS doesn't crash I may have the data locally unzipped.&nbsp;</p>\n<p>I may try to find some remote disk tomorrow unzip and download it.&nbsp;&nbsp;&nbsp; Not going to buy new hardware in the next couple of weeks.</p>\n<p>Hopefully sed -i 's/??//&nbsp; *.asm&nbsp; is a little quicker.&nbsp;&nbsp;&nbsp;&nbsp; For me this kaggle is going to be about figuring out the environment, tools and&nbsp; requirements.&nbsp;&nbsp;</p>\n<p>I do wonder whether a less smart unsupervised approach wouldn't work better for discovering features than trying to be a really smart statistician.&nbsp;&nbsp;&nbsp;&nbsp; Curious to see if anyone takes the big data brute force approach and whether it will outperform smart people.&nbsp;&nbsp;&nbsp; There is a lot of data with patterns in it that a visual recognition type approach might work well with.</p>",
      "rawMarkdown": "",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 63977,
      "author_name": "abhishek",
      "author_url": "",
      "post_date": "02/10/2015 22:29:16",
      "content": "<p>[quote=Robert Fontaine;63953]</p>\n<p>Am I asking the right questions?</p>\n<p>...&nbsp;</p>\n<p>[/quote]</p>\n<p>nop!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 64022,
      "author_name": "robertfontaine",
      "author_url": "",
      "post_date": "02/11/2015 13:46:00",
      "content": "<p>Thanks :),</p>\n<p>Well then how about SVM ,R, ff?</p>\n<p>I'm not&nbsp;sure that&nbsp;loading the sample data into a dbms&nbsp;buys me anything for the solution but I have to establish a development environment so I will probably install and configure an rdbms tonight &nbsp;mysql/progres. &nbsp;</p>\n<p>Hopefully,&nbsp;kaggle won't go out of business before I learn how to data mine.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 64023,
      "author_name": "abhishek",
      "author_url": "",
      "post_date": "02/11/2015 13:49:19",
      "content": "<p>you dont have to load any data to dbms or anything. create a script that mines information from individual files and writes it to a separate file. getting a logloss of less than 0.1 will take you only a couple of hours using R or python.</p>\n<p>P.S. I dont think kaggle will go out of business anytime soon :)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 64032,
      "author_name": "yejiming",
      "author_url": "",
      "post_date": "02/11/2015 14:45:33",
      "content": "<p>[quote=Abhishek;64023]</p>\n<p>you dont have to load any data to dbms or anything. create a script that mines information from individual files and writes it to a separate file. getting a logloss of less than 0.1 will take you only a couple of hours using R or python.</p>\n<p>P.S. I dont think kaggle will go out of business anytime soon :)</p>\n<p>[/quote]</p>\n<p>Thank you for your post. This is the first time that I cope with such a large volume data. Is it&nbsp;necessary to unzip the file that downloaded from kaggle? If so, I wouldn't have so much disk space : (</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 64035,
      "author_name": "abhishek",
      "author_url": "",
      "post_date": "02/11/2015 15:10:08",
      "content": "<p>I dont know if you can do that with 7z. I had enough space to decompress the dataset.&nbsp;</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 64068,
      "author_name": "leobuettiker",
      "author_url": "",
      "post_date": "02/11/2015 21:21:04",
      "content": "<p>You can directly pipe files out of the archive. I tried that and it works, you need something along the line of &quot;/cygdrive/c/Program\\ Files/7-Zip/7z.exe e -so dataSample.7z $FILEHASH.asm | grep 'interessing stuff'&quot;, but it's horible slow as you can not have random&nbsp;access in a 7z archive (http://stackoverflow.com/questions/4457997/indexing-random-access-to-7zip-7z-archives).</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 64069,
      "author_name": "robertfontaine",
      "author_url": "",
      "post_date": "02/11/2015 21:26:32",
      "content": "<p>Seagate 1TB Barracuda is&nbsp;about $60CAD. &nbsp; Not terribly good for reliability but a stack of them striped should be able to max out pcie 3.0 bandwidth. &nbsp; I'm going to grab 1 or 2 this weekend. &nbsp; &nbsp; A Multiple ZFS groups of Hitachis in RAIDZ2 or better are far more reliable but data is disposable on my test bench and I can't afford the number of disks that you need to stripe across groups with ZFS.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 64112,
      "author_name": "ajitkumar",
      "author_url": "",
      "post_date": "02/12/2015 13:35:09",
      "content": "<p>Hi,</p>\n<p>I have downloaded and unzipped the dataSample.7z. The folder has 4 files of<span>.</span>bytes and<span> .</span><span>asm</span> types. As two files have the same hash so I concluded that both are different format of the same file.</p>\n<p>Do the train and test folder also have the same types of files?</p>\n<p>Due to storage limitations, I am not unzipping those two zipped files. If they also have both formats, then, can I work with only one format?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 64179,
      "author_name": "datarat",
      "author_url": "",
      "post_date": "02/13/2015 18:58:44",
      "content": "<p>If space is a concern, you can always unzip portions of the training set as needed.</p>\n<p>.asm files usually contain everything that's in .bytes + extras</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 64379,
      "author_name": "robertfontaine1",
      "author_url": "",
      "post_date": "02/17/2015 03:01:44",
      "content": "<p>8 cheap 256TB SSDs on a raid card would make this a lot quicker.&nbsp;&nbsp; Just unzipping on my little ARM based NAS is likely to take a few days.&nbsp;&nbsp; Good thing I have lots of math to study.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 64380,
      "author_name": "davidtran",
      "author_url": "",
      "post_date": "02/17/2015 04:41:04",
      "content": "<p>Before you spend money on fancy hardware, please re-read carefully what Abhishek said. You don't need SSDs, RAID, data striping or anything like that. A standard laptop with 4 GB RAM and 500 GB of free hard drive space is all you need to process the data and generate a submission in a couple of hours of computing time.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 64383,
      "author_name": "robertfontaine1",
      "author_url": "",
      "post_date": "02/17/2015 05:16:44",
      "content": "<p>Understood but it's nice to dream... I've got 2TB on an old dns323 nas with 2TB of disk at the moment using a 500mhz Marvell ARM and 64MB of RAM.&nbsp;&nbsp;&nbsp;&nbsp; On the plus side the 7zip utility doesn't seem to be bleeding memory so in 4 or 5 days if the NAS doesn't crash I may have the data locally unzipped.&nbsp;</p>\n<p>I may try to find some remote disk tomorrow unzip and download it.&nbsp;&nbsp;&nbsp; Not going to buy new hardware in the next couple of weeks.</p>\n<p>Hopefully sed -i 's/??//&nbsp; *.asm&nbsp; is a little quicker.&nbsp;&nbsp;&nbsp;&nbsp; For me this kaggle is going to be about figuring out the environment, tools and&nbsp; requirements.&nbsp;&nbsp;</p>\n<p>I do wonder whether a less smart unsupervised approach wouldn't work better for discovering features than trying to be a really smart statistician.&nbsp;&nbsp;&nbsp;&nbsp; Curious to see if anyone takes the big data brute force approach and whether it will outperform smart people.&nbsp;&nbsp;&nbsp; There is a lot of data with patterns in it that a visual recognition type approach might work well with.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "63953": "",
    "63977": "",
    "64022": "",
    "64023": "",
    "64032": "",
    "64035": "",
    "64068": "",
    "64069": "",
    "64112": "",
    "64179": "",
    "64379": "",
    "64380": "",
    "64383": ""
  },
  "source": "meta"
}