{
  "id": 12490,
  "title": "Beat the benchmark (~0.182) with RandomForest ",
  "url": "/competitions/malware-classification/discussion/12490",
  "author_name": "",
  "post_date": "2015-02-12T07:11:20.770Z",
  "votes": 17,
  "comment_count": 31,
  "views": 14632,
  "content": "<p>Hi Kagglers,&nbsp;</p>\n<p>Here is my github repository for the solution that has scored&nbsp;0.1826662 on leader board.</p>\n<p>https://github.com/vrajs5/Microsoft-Malware-Classification-Challenge</p>\n\n<p>Solution is quite simple, tiresome part is data preparation. It used only .byte files to predict category. It calculate frequency of two-byte-codes (00 to FF) along with ?? and use that information&nbsp;for prediction.</p>\n\n<p>Before using these files you have to follow this step:</p>\n<ol>\n<li><span style=\"line-height: 1.4\">Extract .byte files from train and test 7z</span></li>\n<li><span style=\"line-height: 1.4\">Gzip .byte files to .byte.gz format and move to train_gz / test_gz file.</span></li>\n</ol>\n\n<p>I know these two steps will take hell lot of time, for me 6 hours. :)</p>\n\n<p>Once you have 10868 train files and 10873 test files in gz format, run following commands</p>\n<p><strong>python&nbsp;data_consolidation.py</strong></p>\n<p><strong>python&nbsp;solution.py</strong></p>\n\n<p>Use it, tune it and score as low as you can.</p>\n\n<p>This script should run with Python-2 and Python-3 both.&nbsp;Let me know if you face any problems.</p>\n\n<p>Vishnu</p>",
  "messages": [
    {
      "id": "64097",
      "postDate": "02/12/2015 07:11:20",
      "content": "<p>Hi Kagglers,&nbsp;</p>\n<p>Here is my github repository for the solution that has scored&nbsp;0.1826662 on leader board.</p>\n<p>https://github.com/vrajs5/Microsoft-Malware-Classification-Challenge</p>\n\n<p>Solution is quite simple, tiresome part is data preparation. It used only .byte files to predict category. It calculate frequency of two-byte-codes (00 to FF) along with ?? and use that information&nbsp;for prediction.</p>\n\n<p>Before using these files you have to follow this step:</p>\n<ol>\n<li><span style=\"line-height: 1.4\">Extract .byte files from train and test 7z</span></li>\n<li><span style=\"line-height: 1.4\">Gzip .byte files to .byte.gz format and move to train_gz / test_gz file.</span></li>\n</ol>\n\n<p>I know these two steps will take hell lot of time, for me 6 hours. :)</p>\n\n<p>Once you have 10868 train files and 10873 test files in gz format, run following commands</p>\n<p><strong>python&nbsp;data_consolidation.py</strong></p>\n<p><strong>python&nbsp;solution.py</strong></p>\n\n<p>Use it, tune it and score as low as you can.</p>\n\n<p>This script should run with Python-2 and Python-3 both.&nbsp;Let me know if you face any problems.</p>\n\n<p>Vishnu</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "64098",
      "postDate": "02/12/2015 07:33:42",
      "content": "<p>Good Job. Did you try not using &quot;??&quot; ?</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "64101",
      "postDate": "02/12/2015 07:52:55",
      "content": "<p>[quote=Abhishek;64098]</p>\n<p>Good Job. Did you try not using &quot;??&quot; ?</p>\n<p>[/quote]</p>\n\n<p>Without '??' 0.1733829 .. Better :)</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "64102",
      "postDate": "02/12/2015 07:55:01",
      "content": "<p>;)</p>\n<p>P.S. SVNIT rocks! :P&nbsp;</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "64311",
      "postDate": "02/16/2015 04:58:01",
      "content": "<p>I have reached 0.0964 in LB according to your strategy, by just adding&nbsp;n_estimators to 200.</p>\n<p>Now I'm ranking one position ahead you. :)</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "64318",
      "postDate": "02/16/2015 06:30:27",
      "content": "<p>[quote=Jiming Ye;64311]</p>\n<p>I have reached 0.0964 in LB according to your strategy, by just adding&nbsp;n_estimators to 200.</p>\n<p>Now I'm ranking one position ahead you. :)</p>\n<p>[/quote]</p>\n\n<p>I am happy to know that...&nbsp;</p>\n<p>That's why I said 'use it, tune it, score as low as you can... '</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "64559",
      "postDate": "02/19/2015 05:41:24",
      "content": "<p>[quote=vrajs5;64097]</p>\n<p>Before using these files you have to follow this step:</p>\n<ol>\n<li><span style=\"line-height: 1.4\">Extract .byte files from train and test 7z</span></li>\n<li><span style=\"line-height: 1.4\">Gzip .byte files to .byte.gz format and move to train_gz / test_gz file.</span></li>\n</ol>\n\n<p>Vishnu</p>\n<p>[/quote]</p>\n\n<p>Hi Vishnu</p>\n<p>First, thank you for such easy-to-follow code and instructions :)</p>\n<p>Could you let me know how to compress into .byte.gz format...I'm only getting any option to .tar.gz (using 7z)....have I misunderstood something?</p>\n<p>Thanks!!</p>\n<p>UD</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "64561",
      "postDate": "02/19/2015 06:19:41",
      "content": "<p>[quote=UD1989;64559]</p>\n<p>[quote=vrajs5;64097]</p>\n<p>Before using these files you have to follow this step:</p>\n<ol>\n<li><span style=\"line-height: 1.4\">Extract .byte files from train and test 7z</span></li>\n<li><span style=\"line-height: 1.4\">Gzip .byte files to .byte.gz format and move to train_gz / test_gz file.</span></li>\n</ol>\n<p>Vishnu</p>\n<p>[/quote]</p>\n<p>Hi Vishnu</p>\n<p>First, thank you for such easy-to-follow code and instructions :)</p>\n<p>Could you let me know how to compress into .byte.gz format...I'm only getting any option to .tar.gz (using 7z)....have I misunderstood something?</p>\n<p>Thanks!!</p>\n<p>UD</p>\n<p>[/quote]</p>\n\n<p>If you are using windows you can make batch file using python that can generate following command</p>\n<p>7z a -tgzip train_gz/[file_name].gz train/[file_name]</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "64644",
      "postDate": "02/20/2015 13:21:15",
      "content": "<p>Thanks a lot for the reply....however the job kept failing so I finally followed the instruction enlisted here:&nbsp;http://getlevelten.com/tip/how-gzip-tar-file-7-zip</p>\n<p>it says:</p>\n<p><em>&quot;The trick is that 7-Zip will only gzip a single file. So creating a tar.gz is a two step process. First create the tar archive, then use 7-Zip to select the tar and you will get an option to gzip it.&quot;</em></p>\n<p>So that's what I did..... (I have not yet been able to run your code because of this change, but I'm on it)</p>\n<p>----------------------------------------------------</p>\n<p>Another clarification needed: byteFiles = [i for i in Files if '.byte.gz' in i]</p>\n<p>here 'bytefiles' is a list of .gz files....or the .byte files compressed inside .gz files? Sorry if my questions are stupid....I've not really handled files of this nature before...</p>\n<p>-------------------------------------------------------</p>\n<p>BTW, in you code you comment&nbsp;for the 'consolidate' function&nbsp;&quot;<em>This function reads each asm files (stored in gzip format)and prepare summary. asm gzip files are stored in train_gz and test_gz locating.&quot;</em></p>\n<p>Dont you mean the byte files? Since asm files are as yet untouched....</p>\n\n<p>Thanks for your clarifications !</p>\n<p>UD</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "64907",
      "postDate": "02/26/2015 02:51:28",
      "content": "<p>Here is a code snippet in C# you can use to convert all those .bytes files to .bytes.gz</p>\n<p>string _filesOrigin = &quot;\\test\\\\test_bytes\\\\&quot;; &nbsp;// got all those bytes files to a separate folder here</p>\n<p>string _filesDest = &quot;\\\\test\\\\test_gz\\\\&quot;;</p>\n<p>String[] files = Directory.GetFiles(_filesOrigin);</p>\n<p>foreach (String f in files)<br> {<br> string fileToBeCompressed = f;<br> string compressedFile = _filesDest + f.Replace(_filesOrigin,&quot;&quot;) + &quot;.gz&quot;;<br> <br> using (FileStream target = new FileStream(compressedFile, FileMode.Create, FileAccess.Write))<br> using (GZipStream alg = new GZipStream(target, CompressionMode.Compress))<br> {<br> byte[] data = File.ReadAllBytes(fileToBeCompressed);<br> alg.Write(data, 0, data.Length);<br> alg.Flush();<br> }<br> }</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "64986",
      "postDate": "02/26/2015 19:42:19",
      "content": "<p>[quote=vrajs5;64097]</p>\n<p>Hi Kagglers,&nbsp;</p>\n<p>Here is my github repository for the solution that has scored&nbsp;0.1826662 on leader board.</p>\n<p>https://github.com/vrajs5/Microsoft-Malware-Classification-Challenge</p>\n<p>Solution is quite simple, tiresome part is data preparation. It used only .byte files to predict category. It calculate frequency of two-byte-codes (00 to FF) along with ?? and use that information&nbsp;for prediction.</p>\n<p>Before using these files you have to follow this step:</p>\n<ol>\n<li><span style=\"line-height: 1.4\">Extract .byte files from train and test 7z</span></li>\n<li><span style=\"line-height: 1.4\">Gzip .byte files to .byte.gz format and move to train_gz / test_gz file.</span></li>\n</ol>\n<p>I know these two steps will take hell lot of time, for me 6 hours. :)</p>\n<p>Once you have 10868 train files and 10873 test files in gz format, run following commands</p>\n<p><strong>python&nbsp;data_consolidation.py</strong></p>\n<p><strong>python&nbsp;solution.py</strong></p>\n<p>Use it, tune it and score as low as you can.</p>\n<p>This script should run with Python-2 and Python-3 both.&nbsp;Let me know if you face any problems.</p>\n<p>Vishnu</p>\n<p>[/quote]</p>\n<p>The data&nbsp;consolidation script only counts the single byte and not&nbsp;the two bytes, right?</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "65014",
      "postDate": "02/27/2015 09:03:25",
      "content": "<p>[quote=clustifier;64986]</p>\n<p>The data&nbsp;consolidation script only counts the single byte and not&nbsp;the two bytes, right?</p>\n<p>[/quote]</p>\n<p>Right, not two bytes.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "65017",
      "postDate": "02/27/2015 10:27:50",
      "content": "<p>and the&nbsp;benchmark score (0.182) was achieved with the provided code or after adding the two bytes?</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "65019",
      "postDate": "02/27/2015 10:42:30",
      "content": "<p>[quote=clustifier;65017]</p>\n<p>and the&nbsp;benchmark score (0.182) was achieved with the provided code or after adding the two bytes?</p>\n<p>[/quote]</p>\n<p>Have you checked code?</p>\n<p>And if 0.182 is with 2bytes then I would have added to code. Can you at least look at code please?</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "65020",
      "postDate": "02/27/2015 11:10:13",
      "content": "<p>Sure I checked the code but you initialized the header with one and two byte, but then counted only the one byte so I wasn't sure what you meant to do.</p>\n<p>When I&nbsp;tried the code&nbsp;I&nbsp;got really bad score.</p>\n<p>I'm probably missing something here.</p>\n<p>And thanks for the code, BTW.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "65022",
      "postDate": "02/27/2015 12:46:04",
      "content": "<p>hi guys how much time does the data consoldilation script take to run? Its been on for about 3 hours now and still no results....I'm thinking there must be some issues in the changes i made to the code....</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "65025",
      "postDate": "02/27/2015 13:06:16",
      "content": "<p>hey guys how much time does the data consolidation script take to run ? Its been on for 3 hours now.....wondering if the changes i made to the scripts are causing the problems.....</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "65028",
      "postDate": "02/27/2015 13:49:07",
      "content": "<p>[quote=UD1989;65025]</p>\n<p>hey guys how much time does the data consolidation script take to run ? Its been on for 3 hours now.....wondering if the changes i made to the scripts are causing the problems.....</p>\n<p>[/quote]</p>\n<p>if you read the very first post carefully, you must know that it takes around 6 hours! Come on, please read the post instead of blindly using the benchmark scripts!!!</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "65029",
      "postDate": "02/27/2015 14:04:54",
      "content": "<p>[quote=Abhishek;65028]</p>\n<p>[quote=UD1989;65025]</p>\n<p>hey guys how much time does the data consolidation script take to run ? Its been on for 3 hours now.....wondering if the changes i made to the scripts are causing the problems.....</p>\n<p>[/quote]</p>\n<p>if you read the very first post carefully, you must know that it takes around 6 hours! Come on, please read the post instead of blindly using the benchmark scripts!!!</p>\n<p>[/quote]</p>\n<p>@Abhishek, the 6 hours mentioned was referring to the two steps before running the scripts, @UD1989 was&nbsp;referring to the first consolidation script.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "65047",
      "postDate": "02/27/2015 16:44:14",
      "content": "<p>[quote=clustifier;65029]</p>\n<p>[quote=Abhishek;65028]</p>\n<p>[quote=UD1989;65025]</p>\n<p>hey guys how much time does the data consolidation script take to run ? Its been on for 3 hours now.....wondering if the changes i made to the scripts are causing the problems.....</p>\n<p>[/quote]</p>\n<p>if you read the very first post carefully, you must know that it takes around 6 hours! Come on, please read the post instead of blindly using the benchmark scripts!!!</p>\n<p>[/quote]</p>\n<p>@Abhishek, the 6 hours mentioned was referring to the two steps before running the scripts, @UD1989 was&nbsp;referring to the first consolidation script.</p>\n<p>[/quote]</p>\n\n<p>@UD1989 - there is a logger prints how many files processed... If you are not getting those things printed something is wrong....&nbsp;</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "65058",
      "postDate": "02/27/2015 18:31:24",
      "content": "<p>[quote=vrajs5;65047]</p>\n<p>[quote=clustifier;65029]</p>\n<p>[quote=Abhishek;65028]</p>\n<p>[quote=UD1989;65025]</p>\n<p>hey guys how much time does the data consolidation script take to run ? Its been on for 3 hours now.....wondering if the changes i made to the scripts are causing the problems.....</p>\n<p>[/quote]</p>\n<p>if you read the very first post carefully, you must know that it takes around 6 hours! Come on, please read the post instead of blindly using the benchmark scripts!!!</p>\n<p>[/quote]</p>\n<p>@Abhishek, the 6 hours mentioned was referring to the two steps before running the scripts, @UD1989 was&nbsp;referring to the first consolidation script.</p>\n<p>[/quote]</p>\n<p>@UD1989 - there is a logger prints how many files processed... If you are not getting those things printed something is wrong....&nbsp;</p>\n<p>[/quote]</p>\n\n<p>Thats what I thought. Thanks. I made significant changes to the code to accomodate some differences in my file formats.....so that must be the problem. Anyway, thanks</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "65061",
      "postDate": "02/27/2015 19:04:04",
      "content": "<p>[quote=Abhishek;65028]</p>\n<p>[quote=UD1989;65025]</p>\n<p>hey guys how much time does the data consolidation script take to run ? Its been on for 3 hours now.....wondering if the changes i made to the scripts are causing the problems.....</p>\n<p>[/quote]</p>\n<p>if you read the very first post carefully, you must know that it takes around 6 hours! Come on, please read the post instead of blindly using the benchmark scripts!!!</p>\n<p>[/quote]</p>\n<p>As the person below said, the 6 hours was for extracting and re gzipping the data...not for running the code. I DID read . <em>Very</em> carefully. And have been changing the code for the last 3 days......so I am not blindly running it either. I know I can sometimes ask stupid questions cause I dont have a&nbsp; programming background so I get stuck in wierd places-but I am determined to complete this challenge and can only hope the Kaggle Kings like you would help me along the way ;)</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "65478",
      "postDate": "03/05/2015 07:10:22",
      "content": "<p>i run on bytes files, use about 1 hour to run two bytes feature extraction.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "65922",
      "postDate": "03/10/2015 22:48:54",
      "content": "<p>I was wondering that, its great that we are achieving good accuracy but can you explain me the thought behind using frequency of those bytes for classification? Why a particular byte is repeating in a particular category of malware?&nbsp;</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "65933",
      "postDate": "03/11/2015 04:04:46",
      "content": "<p>[quote=Harshaneel Gokhale;65922]</p>\n<p>I was wondering that, its great that we are achieving good accuracy but can you explain me the thought behind using frequency of those bytes for classification? Why a particular byte is repeating in a particular category of malware?&nbsp;</p>\n<p>[/quote]</p>\n<p>Logic is simple.&nbsp;</p>\n<p>Byte codes are like DNA of application. Each class of Malware&nbsp;tries to do specific task, which&nbsp;can be observed using byte code. ;)</p>\n<p>I know we can add more feature for better predictability.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "65936",
      "postDate": "03/11/2015 04:12:44",
      "content": "<p>[quote=vrajs5;65933]</p>\n<p>[quote=Harshaneel Gokhale;65922]</p>\n<p>I was wondering that, its great that we are achieving good accuracy but can you explain me the thought behind using frequency of those bytes for classification? Why a particular byte is repeating in a particular category of malware?&nbsp;</p>\n<p>[/quote]</p>\n<p>Logic is simple.&nbsp;</p>\n<p>Byte codes are like DNA of application. Each class of Malware&nbsp;tries to do specific task, which&nbsp;can be observed using byte code. ;)</p>\n<p>I know we can add more feature for better predictability.</p>\n<p>[/quote]</p>\n<p>So why out of so many other bytecodes you chose to go with these? Sorry but I am not trying to mock anyone here but just curious to know why these particular two bytecodes affect category. Also I feel flow of a program (.asm code) does affect category of malware so we can take order of instructions in consideration to generate features.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "65939",
      "postDate": "03/11/2015 05:06:49",
      "content": "<p>[quote=Harshaneel Gokhale;65936]</p>\n<p>[quote=vrajs5;65933]</p>\n<p>[quote=Harshaneel Gokhale;65922]</p>\n<p>I was wondering that, its great that we are achieving good accuracy but can you explain me the thought behind using frequency of those bytes for classification? Why a particular byte is repeating in a particular category of malware?&nbsp;</p>\n<p>[/quote]</p>\n<p>Logic is simple.&nbsp;</p>\n<p>Byte codes are like DNA of application. Each class of Malware&nbsp;tries to do specific task, which&nbsp;can be observed using byte code. ;)</p>\n<p>I know we can add more feature for better predictability.</p>\n<p>[/quote]</p>\n<p>So why out of so many other bytecodes you chose to go with these? Sorry but I am not trying to mock anyone here but just curious to know why these particular two bytecodes affect category. Also I feel flow of a program (.asm code) does affect category of malware so we can take order of instructions in consideration to generate features.</p>\n<p>[/quote]</p>\n<p>I think I clearly mentioned your questions. Just to clarify it further,</p>\n<p>Just like DNA if we use bytecode there are multiple ways to detect malware.&nbsp;Simplest is take frequency distribution, you can take combination or order of bytecodes.</p>\n<p>Using asm you can take approach of type call function calls used.&nbsp;You can even take combination of both asm and byte to predict category of malware.&nbsp;</p>\n\n<p><strong>I gave solution with simplest way, now it's up to you to go further and add as many&nbsp;feature as you want to improve efficiency.&nbsp;</strong></p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "65941",
      "postDate": "03/11/2015 05:43:59",
      "content": "<p>[quote=Harshaneel Gokhale;65936]</p>\n<p>[quote=vrajs5;65933]</p>\n<p>[quote=Harshaneel Gokhale;65922]</p>\n<p>I was wondering that, its great that we are achieving good accuracy but can you explain me the thought behind using frequency of those bytes for classification? Why a particular byte is repeating in a particular category of malware?&nbsp;</p>\n<p>[/quote]</p>\n<p>Logic is simple.&nbsp;</p>\n<p>Byte codes are like DNA of application. Each class of Malware&nbsp;tries to do specific task, which&nbsp;can be observed using byte code. ;)</p>\n<p>I know we can add more feature for better predictability.</p>\n<p>[/quote]</p>\n<p>So why out of so many other bytecodes you chose to go with these? Sorry but I am not trying to mock anyone here but just curious to know why these particular two bytecodes affect category. Also I feel flow of a program (.asm code) does affect category of malware so we can take order of instructions in consideration to generate features.</p>\n<p>[/quote]</p>\n\n<p>What do you mean by &quot;particular two bytecodes&quot; . The model is &nbsp;using Frequency of all byte codes as far as I can see. (hence the 256 columns giving the counts of all possible bytecodes)</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "66981",
      "postDate": "03/18/2015 15:02:07",
      "content": "<p>I've limited knowledge of malware and assembly code. What is the meaning of '??' in the asm codes.&nbsp;</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "66984",
      "postDate": "03/18/2015 15:18:11",
      "content": "<p>its a reserved space</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "68860",
      "postDate": "03/29/2015 17:10:03",
      "content": "<p>If you want to convert the .bytes files to .bytes.gz , please follow the following:</p>\n<p>import os<br>for file in glob.glob(&quot;*.bytes&quot;):<br>&nbsp;&nbsp; &nbsp; destFile=file+&quot;.gz&quot;<br>&nbsp;&nbsp;&nbsp;&nbsp; executableString = &quot;7z a -tgzip test_gz/&quot;+destFile+&quot; &quot;+file<br>&nbsp;&nbsp;&nbsp;&nbsp; print 'Executable is :',executableString<br>&nbsp;&nbsp;&nbsp;&nbsp; os.system(executableString)</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "78444",
      "postDate": "05/14/2015 22:06:58",
      "content": "<p>Hey Friends,</p>\n<p>I am trying to understand the significance of the frequencies of this byte count classwise.So, is there any way to deduct the pattern of each class with respect to the count of the byte codes? For example, I want to know that class 1 files show the occurrence of 'F8' tokens to be x number of times or say that if 'F8 tokens' appear in file x number of times, it could be considered to be Class 1 file.</p>\n<p>Please share any information.</p>\n<p>Thanks &amp; Regards</p>",
      "rawMarkdown": "",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 64098,
      "author_name": "abhishek",
      "author_url": "",
      "post_date": "02/12/2015 07:33:42",
      "content": "<p>Good Job. Did you try not using &quot;??&quot; ?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 64101,
      "author_name": "chevli",
      "author_url": "",
      "post_date": "02/12/2015 07:52:55",
      "content": "<p>[quote=Abhishek;64098]</p>\n<p>Good Job. Did you try not using &quot;??&quot; ?</p>\n<p>[/quote]</p>\n\n<p>Without '??' 0.1733829 .. Better :)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 64102,
      "author_name": "abhishek",
      "author_url": "",
      "post_date": "02/12/2015 07:55:01",
      "content": "<p>;)</p>\n<p>P.S. SVNIT rocks! :P&nbsp;</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 64311,
      "author_name": "yejiming",
      "author_url": "",
      "post_date": "02/16/2015 04:58:01",
      "content": "<p>I have reached 0.0964 in LB according to your strategy, by just adding&nbsp;n_estimators to 200.</p>\n<p>Now I'm ranking one position ahead you. :)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 64318,
      "author_name": "chevli",
      "author_url": "",
      "post_date": "02/16/2015 06:30:27",
      "content": "<p>[quote=Jiming Ye;64311]</p>\n<p>I have reached 0.0964 in LB according to your strategy, by just adding&nbsp;n_estimators to 200.</p>\n<p>Now I'm ranking one position ahead you. :)</p>\n<p>[/quote]</p>\n\n<p>I am happy to know that...&nbsp;</p>\n<p>That's why I said 'use it, tune it, score as low as you can... '</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 64559,
      "author_name": "udatta",
      "author_url": "",
      "post_date": "02/19/2015 05:41:24",
      "content": "<p>[quote=vrajs5;64097]</p>\n<p>Before using these files you have to follow this step:</p>\n<ol>\n<li><span style=\"line-height: 1.4\">Extract .byte files from train and test 7z</span></li>\n<li><span style=\"line-height: 1.4\">Gzip .byte files to .byte.gz format and move to train_gz / test_gz file.</span></li>\n</ol>\n\n<p>Vishnu</p>\n<p>[/quote]</p>\n\n<p>Hi Vishnu</p>\n<p>First, thank you for such easy-to-follow code and instructions :)</p>\n<p>Could you let me know how to compress into .byte.gz format...I'm only getting any option to .tar.gz (using 7z)....have I misunderstood something?</p>\n<p>Thanks!!</p>\n<p>UD</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 64561,
      "author_name": "chevli",
      "author_url": "",
      "post_date": "02/19/2015 06:19:41",
      "content": "<p>[quote=UD1989;64559]</p>\n<p>[quote=vrajs5;64097]</p>\n<p>Before using these files you have to follow this step:</p>\n<ol>\n<li><span style=\"line-height: 1.4\">Extract .byte files from train and test 7z</span></li>\n<li><span style=\"line-height: 1.4\">Gzip .byte files to .byte.gz format and move to train_gz / test_gz file.</span></li>\n</ol>\n<p>Vishnu</p>\n<p>[/quote]</p>\n<p>Hi Vishnu</p>\n<p>First, thank you for such easy-to-follow code and instructions :)</p>\n<p>Could you let me know how to compress into .byte.gz format...I'm only getting any option to .tar.gz (using 7z)....have I misunderstood something?</p>\n<p>Thanks!!</p>\n<p>UD</p>\n<p>[/quote]</p>\n\n<p>If you are using windows you can make batch file using python that can generate following command</p>\n<p>7z a -tgzip train_gz/[file_name].gz train/[file_name]</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 64644,
      "author_name": "udatta",
      "author_url": "",
      "post_date": "02/20/2015 13:21:15",
      "content": "<p>Thanks a lot for the reply....however the job kept failing so I finally followed the instruction enlisted here:&nbsp;http://getlevelten.com/tip/how-gzip-tar-file-7-zip</p>\n<p>it says:</p>\n<p><em>&quot;The trick is that 7-Zip will only gzip a single file. So creating a tar.gz is a two step process. First create the tar archive, then use 7-Zip to select the tar and you will get an option to gzip it.&quot;</em></p>\n<p>So that's what I did..... (I have not yet been able to run your code because of this change, but I'm on it)</p>\n<p>----------------------------------------------------</p>\n<p>Another clarification needed: byteFiles = [i for i in Files if '.byte.gz' in i]</p>\n<p>here 'bytefiles' is a list of .gz files....or the .byte files compressed inside .gz files? Sorry if my questions are stupid....I've not really handled files of this nature before...</p>\n<p>-------------------------------------------------------</p>\n<p>BTW, in you code you comment&nbsp;for the 'consolidate' function&nbsp;&quot;<em>This function reads each asm files (stored in gzip format)and prepare summary. asm gzip files are stored in train_gz and test_gz locating.&quot;</em></p>\n<p>Dont you mean the byte files? Since asm files are as yet untouched....</p>\n\n<p>Thanks for your clarifications !</p>\n<p>UD</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 64907,
      "author_name": "bgopalakrishnan",
      "author_url": "",
      "post_date": "02/26/2015 02:51:28",
      "content": "<p>Here is a code snippet in C# you can use to convert all those .bytes files to .bytes.gz</p>\n<p>string _filesOrigin = &quot;\\test\\\\test_bytes\\\\&quot;; &nbsp;// got all those bytes files to a separate folder here</p>\n<p>string _filesDest = &quot;\\\\test\\\\test_gz\\\\&quot;;</p>\n<p>String[] files = Directory.GetFiles(_filesOrigin);</p>\n<p>foreach (String f in files)<br> {<br> string fileToBeCompressed = f;<br> string compressedFile = _filesDest + f.Replace(_filesOrigin,&quot;&quot;) + &quot;.gz&quot;;<br> <br> using (FileStream target = new FileStream(compressedFile, FileMode.Create, FileAccess.Write))<br> using (GZipStream alg = new GZipStream(target, CompressionMode.Compress))<br> {<br> byte[] data = File.ReadAllBytes(fileToBeCompressed);<br> alg.Write(data, 0, data.Length);<br> alg.Flush();<br> }<br> }</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 64986,
      "author_name": "clustifier",
      "author_url": "",
      "post_date": "02/26/2015 19:42:19",
      "content": "<p>[quote=vrajs5;64097]</p>\n<p>Hi Kagglers,&nbsp;</p>\n<p>Here is my github repository for the solution that has scored&nbsp;0.1826662 on leader board.</p>\n<p>https://github.com/vrajs5/Microsoft-Malware-Classification-Challenge</p>\n<p>Solution is quite simple, tiresome part is data preparation. It used only .byte files to predict category. It calculate frequency of two-byte-codes (00 to FF) along with ?? and use that information&nbsp;for prediction.</p>\n<p>Before using these files you have to follow this step:</p>\n<ol>\n<li><span style=\"line-height: 1.4\">Extract .byte files from train and test 7z</span></li>\n<li><span style=\"line-height: 1.4\">Gzip .byte files to .byte.gz format and move to train_gz / test_gz file.</span></li>\n</ol>\n<p>I know these two steps will take hell lot of time, for me 6 hours. :)</p>\n<p>Once you have 10868 train files and 10873 test files in gz format, run following commands</p>\n<p><strong>python&nbsp;data_consolidation.py</strong></p>\n<p><strong>python&nbsp;solution.py</strong></p>\n<p>Use it, tune it and score as low as you can.</p>\n<p>This script should run with Python-2 and Python-3 both.&nbsp;Let me know if you face any problems.</p>\n<p>Vishnu</p>\n<p>[/quote]</p>\n<p>The data&nbsp;consolidation script only counts the single byte and not&nbsp;the two bytes, right?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 65014,
      "author_name": "mikhailtrofimov",
      "author_url": "",
      "post_date": "02/27/2015 09:03:25",
      "content": "<p>[quote=clustifier;64986]</p>\n<p>The data&nbsp;consolidation script only counts the single byte and not&nbsp;the two bytes, right?</p>\n<p>[/quote]</p>\n<p>Right, not two bytes.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 65017,
      "author_name": "clustifier",
      "author_url": "",
      "post_date": "02/27/2015 10:27:50",
      "content": "<p>and the&nbsp;benchmark score (0.182) was achieved with the provided code or after adding the two bytes?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 65019,
      "author_name": "chevli",
      "author_url": "",
      "post_date": "02/27/2015 10:42:30",
      "content": "<p>[quote=clustifier;65017]</p>\n<p>and the&nbsp;benchmark score (0.182) was achieved with the provided code or after adding the two bytes?</p>\n<p>[/quote]</p>\n<p>Have you checked code?</p>\n<p>And if 0.182 is with 2bytes then I would have added to code. Can you at least look at code please?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 65020,
      "author_name": "clustifier",
      "author_url": "",
      "post_date": "02/27/2015 11:10:13",
      "content": "<p>Sure I checked the code but you initialized the header with one and two byte, but then counted only the one byte so I wasn't sure what you meant to do.</p>\n<p>When I&nbsp;tried the code&nbsp;I&nbsp;got really bad score.</p>\n<p>I'm probably missing something here.</p>\n<p>And thanks for the code, BTW.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 65022,
      "author_name": "udatta",
      "author_url": "",
      "post_date": "02/27/2015 12:46:04",
      "content": "<p>hi guys how much time does the data consoldilation script take to run? Its been on for about 3 hours now and still no results....I'm thinking there must be some issues in the changes i made to the code....</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 65025,
      "author_name": "udatta",
      "author_url": "",
      "post_date": "02/27/2015 13:06:16",
      "content": "<p>hey guys how much time does the data consolidation script take to run ? Its been on for 3 hours now.....wondering if the changes i made to the scripts are causing the problems.....</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 65028,
      "author_name": "abhishek",
      "author_url": "",
      "post_date": "02/27/2015 13:49:07",
      "content": "<p>[quote=UD1989;65025]</p>\n<p>hey guys how much time does the data consolidation script take to run ? Its been on for 3 hours now.....wondering if the changes i made to the scripts are causing the problems.....</p>\n<p>[/quote]</p>\n<p>if you read the very first post carefully, you must know that it takes around 6 hours! Come on, please read the post instead of blindly using the benchmark scripts!!!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 65029,
      "author_name": "clustifier",
      "author_url": "",
      "post_date": "02/27/2015 14:04:54",
      "content": "<p>[quote=Abhishek;65028]</p>\n<p>[quote=UD1989;65025]</p>\n<p>hey guys how much time does the data consolidation script take to run ? Its been on for 3 hours now.....wondering if the changes i made to the scripts are causing the problems.....</p>\n<p>[/quote]</p>\n<p>if you read the very first post carefully, you must know that it takes around 6 hours! Come on, please read the post instead of blindly using the benchmark scripts!!!</p>\n<p>[/quote]</p>\n<p>@Abhishek, the 6 hours mentioned was referring to the two steps before running the scripts, @UD1989 was&nbsp;referring to the first consolidation script.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 65047,
      "author_name": "chevli",
      "author_url": "",
      "post_date": "02/27/2015 16:44:14",
      "content": "<p>[quote=clustifier;65029]</p>\n<p>[quote=Abhishek;65028]</p>\n<p>[quote=UD1989;65025]</p>\n<p>hey guys how much time does the data consolidation script take to run ? Its been on for 3 hours now.....wondering if the changes i made to the scripts are causing the problems.....</p>\n<p>[/quote]</p>\n<p>if you read the very first post carefully, you must know that it takes around 6 hours! Come on, please read the post instead of blindly using the benchmark scripts!!!</p>\n<p>[/quote]</p>\n<p>@Abhishek, the 6 hours mentioned was referring to the two steps before running the scripts, @UD1989 was&nbsp;referring to the first consolidation script.</p>\n<p>[/quote]</p>\n\n<p>@UD1989 - there is a logger prints how many files processed... If you are not getting those things printed something is wrong....&nbsp;</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 65058,
      "author_name": "udatta",
      "author_url": "",
      "post_date": "02/27/2015 18:31:24",
      "content": "<p>[quote=vrajs5;65047]</p>\n<p>[quote=clustifier;65029]</p>\n<p>[quote=Abhishek;65028]</p>\n<p>[quote=UD1989;65025]</p>\n<p>hey guys how much time does the data consolidation script take to run ? Its been on for 3 hours now.....wondering if the changes i made to the scripts are causing the problems.....</p>\n<p>[/quote]</p>\n<p>if you read the very first post carefully, you must know that it takes around 6 hours! Come on, please read the post instead of blindly using the benchmark scripts!!!</p>\n<p>[/quote]</p>\n<p>@Abhishek, the 6 hours mentioned was referring to the two steps before running the scripts, @UD1989 was&nbsp;referring to the first consolidation script.</p>\n<p>[/quote]</p>\n<p>@UD1989 - there is a logger prints how many files processed... If you are not getting those things printed something is wrong....&nbsp;</p>\n<p>[/quote]</p>\n\n<p>Thats what I thought. Thanks. I made significant changes to the code to accomodate some differences in my file formats.....so that must be the problem. Anyway, thanks</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 65061,
      "author_name": "udatta",
      "author_url": "",
      "post_date": "02/27/2015 19:04:04",
      "content": "<p>[quote=Abhishek;65028]</p>\n<p>[quote=UD1989;65025]</p>\n<p>hey guys how much time does the data consolidation script take to run ? Its been on for 3 hours now.....wondering if the changes i made to the scripts are causing the problems.....</p>\n<p>[/quote]</p>\n<p>if you read the very first post carefully, you must know that it takes around 6 hours! Come on, please read the post instead of blindly using the benchmark scripts!!!</p>\n<p>[/quote]</p>\n<p>As the person below said, the 6 hours was for extracting and re gzipping the data...not for running the code. I DID read . <em>Very</em> carefully. And have been changing the code for the last 3 days......so I am not blindly running it either. I know I can sometimes ask stupid questions cause I dont have a&nbsp; programming background so I get stuck in wierd places-but I am determined to complete this challenge and can only hope the Kaggle Kings like you would help me along the way ;)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 65478,
      "author_name": "zhaogang",
      "author_url": "",
      "post_date": "03/05/2015 07:10:22",
      "content": "<p>i run on bytes files, use about 1 hour to run two bytes feature extraction.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 65922,
      "author_name": "harshaneel",
      "author_url": "",
      "post_date": "03/10/2015 22:48:54",
      "content": "<p>I was wondering that, its great that we are achieving good accuracy but can you explain me the thought behind using frequency of those bytes for classification? Why a particular byte is repeating in a particular category of malware?&nbsp;</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 65933,
      "author_name": "chevli",
      "author_url": "",
      "post_date": "03/11/2015 04:04:46",
      "content": "<p>[quote=Harshaneel Gokhale;65922]</p>\n<p>I was wondering that, its great that we are achieving good accuracy but can you explain me the thought behind using frequency of those bytes for classification? Why a particular byte is repeating in a particular category of malware?&nbsp;</p>\n<p>[/quote]</p>\n<p>Logic is simple.&nbsp;</p>\n<p>Byte codes are like DNA of application. Each class of Malware&nbsp;tries to do specific task, which&nbsp;can be observed using byte code. ;)</p>\n<p>I know we can add more feature for better predictability.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 65936,
      "author_name": "harshaneel",
      "author_url": "",
      "post_date": "03/11/2015 04:12:44",
      "content": "<p>[quote=vrajs5;65933]</p>\n<p>[quote=Harshaneel Gokhale;65922]</p>\n<p>I was wondering that, its great that we are achieving good accuracy but can you explain me the thought behind using frequency of those bytes for classification? Why a particular byte is repeating in a particular category of malware?&nbsp;</p>\n<p>[/quote]</p>\n<p>Logic is simple.&nbsp;</p>\n<p>Byte codes are like DNA of application. Each class of Malware&nbsp;tries to do specific task, which&nbsp;can be observed using byte code. ;)</p>\n<p>I know we can add more feature for better predictability.</p>\n<p>[/quote]</p>\n<p>So why out of so many other bytecodes you chose to go with these? Sorry but I am not trying to mock anyone here but just curious to know why these particular two bytecodes affect category. Also I feel flow of a program (.asm code) does affect category of malware so we can take order of instructions in consideration to generate features.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 65939,
      "author_name": "chevli",
      "author_url": "",
      "post_date": "03/11/2015 05:06:49",
      "content": "<p>[quote=Harshaneel Gokhale;65936]</p>\n<p>[quote=vrajs5;65933]</p>\n<p>[quote=Harshaneel Gokhale;65922]</p>\n<p>I was wondering that, its great that we are achieving good accuracy but can you explain me the thought behind using frequency of those bytes for classification? Why a particular byte is repeating in a particular category of malware?&nbsp;</p>\n<p>[/quote]</p>\n<p>Logic is simple.&nbsp;</p>\n<p>Byte codes are like DNA of application. Each class of Malware&nbsp;tries to do specific task, which&nbsp;can be observed using byte code. ;)</p>\n<p>I know we can add more feature for better predictability.</p>\n<p>[/quote]</p>\n<p>So why out of so many other bytecodes you chose to go with these? Sorry but I am not trying to mock anyone here but just curious to know why these particular two bytecodes affect category. Also I feel flow of a program (.asm code) does affect category of malware so we can take order of instructions in consideration to generate features.</p>\n<p>[/quote]</p>\n<p>I think I clearly mentioned your questions. Just to clarify it further,</p>\n<p>Just like DNA if we use bytecode there are multiple ways to detect malware.&nbsp;Simplest is take frequency distribution, you can take combination or order of bytecodes.</p>\n<p>Using asm you can take approach of type call function calls used.&nbsp;You can even take combination of both asm and byte to predict category of malware.&nbsp;</p>\n\n<p><strong>I gave solution with simplest way, now it's up to you to go further and add as many&nbsp;feature as you want to improve efficiency.&nbsp;</strong></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 65941,
      "author_name": "udatta",
      "author_url": "",
      "post_date": "03/11/2015 05:43:59",
      "content": "<p>[quote=Harshaneel Gokhale;65936]</p>\n<p>[quote=vrajs5;65933]</p>\n<p>[quote=Harshaneel Gokhale;65922]</p>\n<p>I was wondering that, its great that we are achieving good accuracy but can you explain me the thought behind using frequency of those bytes for classification? Why a particular byte is repeating in a particular category of malware?&nbsp;</p>\n<p>[/quote]</p>\n<p>Logic is simple.&nbsp;</p>\n<p>Byte codes are like DNA of application. Each class of Malware&nbsp;tries to do specific task, which&nbsp;can be observed using byte code. ;)</p>\n<p>I know we can add more feature for better predictability.</p>\n<p>[/quote]</p>\n<p>So why out of so many other bytecodes you chose to go with these? Sorry but I am not trying to mock anyone here but just curious to know why these particular two bytecodes affect category. Also I feel flow of a program (.asm code) does affect category of malware so we can take order of instructions in consideration to generate features.</p>\n<p>[/quote]</p>\n\n<p>What do you mean by &quot;particular two bytecodes&quot; . The model is &nbsp;using Frequency of all byte codes as far as I can see. (hence the 256 columns giving the counts of all possible bytecodes)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 66981,
      "author_name": "azizulhakim",
      "author_url": "",
      "post_date": "03/18/2015 15:02:07",
      "content": "<p>I've limited knowledge of malware and assembly code. What is the meaning of '??' in the asm codes.&nbsp;</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 66984,
      "author_name": "abhishek",
      "author_url": "",
      "post_date": "03/18/2015 15:18:11",
      "content": "<p>its a reserved space</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 68860,
      "author_name": "sravanb",
      "author_url": "",
      "post_date": "03/29/2015 17:10:03",
      "content": "<p>If you want to convert the .bytes files to .bytes.gz , please follow the following:</p>\n<p>import os<br>for file in glob.glob(&quot;*.bytes&quot;):<br>&nbsp;&nbsp; &nbsp; destFile=file+&quot;.gz&quot;<br>&nbsp;&nbsp;&nbsp;&nbsp; executableString = &quot;7z a -tgzip test_gz/&quot;+destFile+&quot; &quot;+file<br>&nbsp;&nbsp;&nbsp;&nbsp; print 'Executable is :',executableString<br>&nbsp;&nbsp;&nbsp;&nbsp; os.system(executableString)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 78444,
      "author_name": "malnuggets",
      "author_url": "",
      "post_date": "05/14/2015 22:06:58",
      "content": "<p>Hey Friends,</p>\n<p>I am trying to understand the significance of the frequencies of this byte count classwise.So, is there any way to deduct the pattern of each class with respect to the count of the byte codes? For example, I want to know that class 1 files show the occurrence of 'F8' tokens to be x number of times or say that if 'F8 tokens' appear in file x number of times, it could be considered to be Class 1 file.</p>\n<p>Please share any information.</p>\n<p>Thanks &amp; Regards</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "64097": "",
    "64098": "",
    "64101": "",
    "64102": "",
    "64311": "",
    "64318": "",
    "64559": "",
    "64561": "",
    "64644": "",
    "64907": "",
    "64986": "",
    "65014": "",
    "65017": "",
    "65019": "",
    "65020": "",
    "65022": "",
    "65025": "",
    "65028": "",
    "65029": "",
    "65047": "",
    "65058": "",
    "65061": "",
    "65478": "",
    "65922": "",
    "65933": "",
    "65936": "",
    "65939": "",
    "65941": "",
    "66981": "",
    "66984": "",
    "68860": "",
    "78444": ""
  },
  "source": "meta"
}