{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"<h1 style=\"font-size:250%; font-family:cursive; color:#ff6666;\"><b>Microsoft Malware Detection</b><a id=\"0\"></a></h1><a href=\"https://youtu.be/QI0qjDfMtAw?list=PLxqBkZuBynVS8mDTc8ZGermXiS-32pR2y\"><h1 style=\"font-size:250%; font-family:cursive; color:#ff6666;\"><b>Link to my Detailed YouTube Video Explaining the whole Notebook</b></h1></a>\n\n\n[![IMAGE ALT TEXT](https://imgur.com/0Ke8XUi.png){:.some-css-class style=\"width: 200px\"}](https://www.youtube.com/watch?v=QI0qjDfMtAw&list=PLxqBkZuBynVS8mDTc8ZGermXiS-32pR2y&ab_channel=RohanPaul \"Microsoft Malware Detection Kaggle Competition - 0.007 LogLoss with XGBoost\")\n\n<iframe  title=\"YouTube video player\" width=\"480\" height=\"390\" src=\"https://youtu.be/QI0qjDfMtAw?list=PLxqBkZuBynVS8mDTc8ZGermXiS-32pR2y\" frameborder=\"0\" allowfullscreen></iframe>\n<h1 style=\"font-size:250%; font-family:cursive; color:#ff6666;\"><b>Microsoft Malware Detection</b><a id=\"0\"></a></h1>\n\n## One drawback of this Notebook is - I can NOT show output of any of the cell, because to see the output of my kernel I have to necessarily execute/run the notebook inside Kaggle Editor. But given the huge 500GB dataset, this notebook can NOT be run inside Kaggle. Hence, I had to do the entire work of this notebook in my local Machine and then upload the NB to Kaggle.\n\n<h1 style=\"font-size:100%; font-family:cursive; color:#ff6666;\"><b>Hence for all the plot-outputs, I have actually copy-pasted the image of the respective plot/visualization from my local machine's notebook to here in this Notebook in Kaggle </b></h1>\n\n\n### [The actual Kaggle Challenge](https://www.kaggle.com/c/malware-classification/overview)\n\nIdentify whether a given piece of file/software is a malware. \n\n### In this Notebook, I achieved a test log loss of 0.0070458 with XGBoost\n\n(This Test Dataset is based on splitting only train.7z (which is ~200GB after extraction) into Train, Test and CV )\n\n<font size=5 color='blue'> What is in this kernel</font>\n\n1. [Data Description](#1)\n2. [Issues I encountered for this large dataset](#2)\n3. [My Final overall approach to handle this huge dataset](#3)\n4. [Data Overview](#4)\n5. [Performance Metric](#5)\n6. [Machine Learing Objectives and Constraints](#6)\n7. [Exploratory Data Analysis](#7)\n8. [Distribution of malware classes in whole data set](#8)\n9. [File size  of byte files as a feature](#9)\n10. [box plots of file size (.byte files) feature](#10)\n11. [Uni-Gram Byte Feature extraction from byte files](#11)\n12. [Multivariate Analysis on byte files](#12)\n13. [Train Test split of only Byte Files Features](#13)\n14. [Random Model ONLY on bytes files](#14)\n15. [K Nearest Neighbour Classification ONLY on bytes files](#15)\n16. [Logistic Regression ONLY on bytes files](#16)\n17. [Random Forest Classifier ONLY on bytes files](#17)\n18. [XgBoost Classification ONLY on bytes files](#18)\n19. [XgBoost Classification with best hyper parameters using RandomSearch ONLY on bytes files](#19)\n20. [Modeling with .asm files](#20)\n21. [Feature extraction from asm files](#21)\n22. [Files sizes of each .asm file as a feature](#22)\n23. [Univariate analysis ONLY on .asm file features](#23)\n24. [Multivariate Analysis ONLY on .asm file features](#24)\n25. [Conclusion on EDA ( ONLY on .asm file features)](#25)\n26. [Train and test split ( ONLY on .asm file featues )](#26)\n27. [K-Nearest Neigbors ONLY on .asm file features](#27)\n28. [Logistic Regression ONLY on .asm file features](#28)\n29. [Random Forest Classifier ONLY on .asm file features](#29)\n30. [XgBoost Classifier ONLY on .asm file features](#30)\n31. [Xgboost Classifier with best hyperparameters ( ONLY on .asm file features )](#31)\n32. [FINAL FEATURIZATION STEPS FOR THE FINAL XGBOOST MODEL TRAINING](#32)\n33. [Uni-Gram Byte Feature extraction from byte files - For FINAL Model Train](#33)\n34. [File sizes of Byte files - Feature Extraction -For FINAL Model Train](#34)\n35. [Creating some important Files and Folders, which I shall use later for saving Featuarized versions of .csv files](#35)\n36. [Merging Unigram of Byte Files + Size of Byte Files to create uni_gram_byte_features__with_size](#36)\n37. [Bi-Gram Byte Feature extraction from byte files](#37)\n38. [Extracting the 2000 Most Important Features from Byte bigrams using SelectKBest with Chi-Square Test](#38)\n39. [ASM Unigram - Top 52 Unigram Features from ASM Files - Final Model Training](#39)\n40. [File Size of ASM Files - Feature Extraction - Final Model Training](#40)\n41. [Merging ASM Unigram + ASM File Size](#41)\n42. [ASM Files - Convert the ASM files to images](#42)\n43. [Extract the first 800 pixel data from ASM File Images](#43)\n44. [Extracting Opcodes Bigrams from ASM Files](#44)\n45. [Calcualte opcodes bigram with above defined function and make them a feature and then save the data matrix of feature as a .csv file](#45)\n46. [ASM File - Top Important 500 features from Opcodes Bigrams](#46)\n47. [Opcodes Trigrams ASM Files - Feature extraction ](#47)\n48. [SM File - Top Important 800 features from Opcodes Trigrams](#48)\n49. [Final Merging of all Features for the Final XGBOOST Training](#49)\n50. [Final Train Test Split. 64% Train, 16% Cross Validation, 20% Test](#50)\n51. [Final XGBoost Training - Hyperparameter tuning with on Final Merged Data-Matrix](#51)\n52. [Final running of XGBoost with the Best HyperParams that we got from above RandomizedSearchCV](#52)\n53. [Possibiliy of Further Analysis and Featurizition](#53)\n\n---\n\n\n\n<h1 style=\"font-size:250%; font-family:cursive; color:#ff6666;\"><b>1.Data Description</b><a id=\"1\"></a></h1>\n\n\n#### [Back to the top](#0)\n\n\nYou are provided with a set of known malware files representing a mix of 9 different families. Each malware file has an Id, a 20 character hash value uniquely identifying the file, and a Class, an integer representing one of 9 family names to which the malware may belong:\n\n![Imgur](https://imgur.com/xyRX60l.png)\n\nFor each file, the raw data contains the hexadecimal representation of the file's binary content, without the PE header (to ensure sterility).  You are also provided a metadata manifest, which is a log containing various metadata information extracted from the binary, such as function calls, strings, etc. This was generated using the IDA disassembler tool. Your task is to develop the best mechanism for classifying files in the test set into their respective family affiliations.\n\nThe dataset contains the following files:\n\n* train.7z - the raw data for the training set (MD5 hash = 4fedb0899fc2210a6c843889a70952ed)\n* trainLabels.csv - the class labels associated with the training set\n* test.7z - the raw data for the test set (MD5 hash = 84b6fbfb9df3c461ed2cbbfa371ffb43)\n* sampleSubmission.csv - a file showing the valid submission format\n* dataSample.csv - a sample of the dataset to preview before downloading\n\n\nHere we are provided with raw data and no pre-extracted features were available\n\n\n#### Total train dataset consist of 200GB data out of which 50Gb of data is .bytes files and 150GB of data is .asm files:\n\n\n---\n\n<h1 style=\"font-size:250%; font-family:cursive; color:#ff6666;\"><b> 2. Issues I encountered for this large dataset <a id=\"2\"></a></b></h1>\n\n#### [Back to the top](#0)\n\n\nDue to the large size of the dataset (500GB), I had real issues fitting the data into memory during runtime (Colab Pro Failed, and I definitely could not run it in Kaggle)\n\nIn Kaggle I got out of disk-space when trying to extract only the train.7z file.\n\nThe below code started extracting and only at around 5% I got the out of disk error\n\n```py\n!pip install py7zr\n\n!python -m py7zr x full_path_of_7z_file\n\n!python -m py7zr x /content/gdrive/MyDrive/MS_Malware_Kaggle_to_Gdrive/train.7z\n\n```\n\nIt could definitely have been done in Google-Cloud or AWS but have not tried these option.\n\n---\n\n<h1 style=\"font-size:250%; font-family:cursive; color:#ff6666;\"><b>3. My Final overall approach to handle this huge dataset <a id=\"3\"></a></b></h1>\n\n#### [Back to the top](#0)\n\n\nTherefore, I ONLY extracted train.7z (which is ~200GB after extraction) in my local machine and then split this set into Train, Test and CV set. And from this split dataset, I did my entire analysis on the train and did the validation part on test and cv set.\n\nFurther to be able to accomodate it in my local Machine (which is not too high-end )\n\nFirst did all my calculations and experimentations and featuriazation ONLY on a sample of 50 files  (i.e. 50 each from byteFiles and asmFile ) -\n\nAfter ONLY I saw that all the featuriazation calculations and xgBoost is running on these 50 samples, then only I ran the same notebook on the full Dataset of 200GB with 20,000+ files\n\n\nAnd here's my approach for calculating all the featuriazations (both for 50-samples and full-dataset).\n\n\n1. Did all the file processing (which are CPU based) part in local machine, it took around 25 to 26 hours. \n\ni.e. This includes calculating the below features in local machine.\n\n\n- Unigram of Byte Files + Size of Byte Files + Top 52 Unigram of ASM Files (These are alrady given by AML)\n\nAdded following extra features.\n\n- Size of ASM Files\n- Top 2000 Bi-Gram of Byte files +  \n- Top 500 Bigram of Opcodes of ASM Files\n- Top 800 Trigram of Opcodes of ASM Files\n- Top 800 ASM Image Features\n\nAfter merging all the above features, the merged dataframe that I got, I created to a .csv file from that (i.e. with the regular to_csv() function ).\n\nThis .csv file with the final merged dataset was just about 170MB. Then Uploaded this file to google-drive. \n\nAnd then from Colab just imported that same final merged .csv file and saved that in a pandas dataframe > do train test and cv split on this and  > \n\nNow below 2 steps with Colab Pro's Tesla V100 16GB GPU \n\nRandomizedSearchCV for hyper param tuning > and after ran XGBoost with best params.\n\n\nAnd these final RandomizedSearchCV and XGBoost took only like 30 mints.\n\n---\n","metadata":{"id":"DD3IQYLbbKzO","editable":false}},{"cell_type":"markdown","source":"<h1 style=\"font-size:250%; font-family:cursive; color:#ff6666;\"><b>4. Data Overview <a id=\"4\"></a></b></h1>\n\n#### [Back to the top](#0)\n\n","metadata":{"id":"t-29OcVabK0H","editable":false}},{"cell_type":"markdown","source":"</h1>2.1.2. Example Data Point</h3>","metadata":{"id":"qwDMzZ9ibK0L","editable":false}},{"cell_type":"markdown","source":"<p style = \"font-size:18px\"><b> .asm file</b></p>\n<pre>\n.text:00401000\t\t\t\t\t\t\t\t       assume es:nothing, ss:nothing, ds:_data,\tfs:nothing, gs:nothing\n.text:00401000 56\t\t\t\t\t\t\t       push    esi\n.text:00401001 8D 44 24\t08\t\t\t\t\t\t       lea     eax, [esp+8]\n.text:00401005 50\t\t\t\t\t\t\t       push    eax\n.text:00401006 8B F1\t\t\t\t\t\t\t       mov     esi, ecx\n.text:00401008 E8 1C 1B\t00 00\t\t\t\t\t\t       call    ??0exception@std@@QAE@ABQBD@Z ; std::exception::exception(char const * const &)\n.text:0040100D C7 06 08\tBB 42 00\t\t\t\t\t       mov     dword ptr [esi],\toffset off_42BB08\n.text:00401013 8B C6\t\t\t\t\t\t\t       mov     eax, esi\n.text:00401015 5E\t\t\t\t\t\t\t       pop     esi\n.text:00401016 C2 04 00\t\t\t\t\t\t\t       retn    4\n.text:00401016\t\t\t\t\t\t       ; ---------------------------------------------------------------------------\n.text:00401019 CC CC CC\tCC CC CC CC\t\t\t\t\t       align 10h\n.text:00401020 C7 01 08\tBB 42 00\t\t\t\t\t       mov     dword ptr [ecx],\toffset off_42BB08\n.text:00401026 E9 26 1C\t00 00\t\t\t\t\t\t       jmp     sub_402C51\n.text:00401026\t\t\t\t\t\t       ; ---------------------------------------------------------------------------\n.text:0040102B CC CC CC\tCC CC\t\t\t\t\t\t       align 10h\n.text:00401030 56\t\t\t\t\t\t\t       push    esi\n.text:00401031 8B F1\t\t\t\t\t\t\t       mov     esi, ecx\n.text:00401033 C7 06 08\tBB 42 00\t\t\t\t\t       mov     dword ptr [esi],\toffset off_42BB08\n.text:00401039 E8 13 1C\t00 00\t\t\t\t\t\t       call    sub_402C51\n.text:0040103E F6 44 24\t08 01\t\t\t\t\t\t       test    byte ptr\t[esp+8], 1\n.text:00401043 74 09\t\t\t\t\t\t\t       jz      short loc_40104E\n.text:00401045 56\t\t\t\t\t\t\t       push    esi\n.text:00401046 E8 6C 1E\t00 00\t\t\t\t\t\t       call    ??3@YAXPAX@Z    ; operator delete(void *)\n.text:0040104B 83 C4 04\t\t\t\t\t\t\t       add     esp, 4\n.text:0040104E\n.text:0040104E\t\t\t\t\t\t       loc_40104E:\t\t\t       ; CODE XREF: .text:00401043\u0018j\n.text:0040104E 8B C6\t\t\t\t\t\t\t       mov     eax, esi\n.text:00401050 5E\t\t\t\t\t\t\t       pop     esi\n.text:00401051 C2 04 00\t\t\t\t\t\t\t       retn    4\n.text:00401051\t\t\t\t\t\t       ; ---------------------------------------------------------------------------\n</pre>\n<p style = \"font-size:18px\"><b> .bytes file</b></p>\n<pre>\n00401000 00 00 80 40 40 28 00 1C 02 42 00 C4 00 20 04 20\n00401010 00 00 20 09 2A 02 00 00 00 00 8E 10 41 0A 21 01\n00401020 40 00 02 01 00 90 21 00 32 40 00 1C 01 40 C8 18\n00401030 40 82 02 63 20 00 00 09 10 01 02 21 00 82 00 04\n00401040 82 20 08 83 00 08 00 00 00 00 02 00 60 80 10 80\n00401050 18 00 00 20 A9 00 00 00 00 04 04 78 01 02 70 90\n00401060 00 02 00 08 20 12 00 00 00 40 10 00 80 00 40 19\n00401070 00 00 00 00 11 20 80 04 80 10 00 20 00 00 25 00\n00401080 00 00 01 00 00 04 00 10 02 C1 80 80 00 20 20 00\n00401090 08 A0 01 01 44 28 00 00 08 10 20 00 02 08 00 00\n004010A0 00 40 00 00 00 34 40 40 00 04 00 08 80 08 00 08\n004010B0 10 00 40 00 68 02 40 04 E1 00 28 14 00 08 20 0A\n004010C0 06 01 02 00 40 00 00 00 00 00 00 20 00 02 00 04\n004010D0 80 18 90 00 00 10 A0 00 45 09 00 10 04 40 44 82\n004010E0 90 00 26 10 00 00 04 00 82 00 00 00 20 40 00 00\n004010F0 B4 00 00 40 00 02 20 25 08 00 00 00 00 00 00 00\n00401100 08 00 00 50 00 08 40 50 00 02 06 22 08 85 30 00\n00401110 00 80 00 80 60 00 09 00 04 20 00 00 00 00 00 00\n00401120 00 82 40 02 00 11 46 01 4A 01 8C 01 E6 00 86 10\n00401130 4C 01 22 00 64 00 AE 01 EA 01 2A 11 E8 10 26 11\n00401140 4E 11 8E 11 C2 00 6C 00 0C 11 60 01 CA 00 62 10\n00401150 6C 01 A0 11 CE 10 2C 11 4E 10 8C 00 CE 01 AE 01\n00401160 6C 10 6C 11 A2 01 AE 00 46 11 EE 10 22 00 A8 00\n00401170 EC 01 08 11 A2 01 AE 10 6C 00 6E 00 AC 11 8C 00\n00401180 EC 01 2A 10 2A 01 AE 00 40 00 C8 10 48 01 4E 11\n00401190 0E 00 EC 11 24 10 4A 10 04 01 C8 11 E6 01 C2 00\n\n</pre>","metadata":{"id":"PXY55eyvbK0L","editable":false}},{"cell_type":"markdown","source":"<h1 style=\"font-size:250%; font-family:cursive; color:#ff6666;\"><b>5. Performance Metric <a id=\"5\"></a></b></h1>\n\n\n#### [Back to the top](#0)\n\nSource: https://www.kaggle.com/c/malware-classification#evaluation\n\nMetric(s): \n* Multi class log-loss \n* Confusion matrix \n\n\n<h1 style=\"font-size:250%; font-family:cursive; color:#ff6666;\"><b>6. Machine Learing Objectives and Constraints <a id=\"6\"></a></b></h1>\n\n#### [Back to the top](#0)\n\nObjective: Predict the probability of each data-point belonging to each of the nine classes.\n\nConstraints:\n\n* Class probabilities are needed.\n* Penalize the errors in class probabilites => Metric is Log-loss.\n* Some Latency constraints.\n\n\n<h1 style=\"font-size:250%; font-family:cursive; color:#ff6666;\"><b>7. Exploratory Data Analysis <a id=\"7\"></a> </b></h1>\n\n\n#### [Back to the top](#0)\n","metadata":{"id":"iuLPW5HZbK0S","editable":false}},{"cell_type":"code","source":"%%time\n\n%pip install -U tornado\n%pip install \"dask[complete]\"\n\nimport warnings\nwarnings.filterwarnings(\"ignore\")\nimport shutil\nimport os\nimport pandas as pd\nimport matplotlib\nmatplotlib.use(u'nbAgg')\nimport matplotlib.pyplot as plt\nimport seaborn as sns\nimport numpy as np\nfrom tqdm import tqdm\nimport pickle\nfrom sklearn.manifold import TSNE\nfrom sklearn import preprocessing\nimport pandas as pd\nfrom multiprocessing import Process# this is used for multithreading\nimport multiprocessing\nimport codecs# this is used for file operations \nimport random as r\nfrom xgboost import XGBClassifier\nfrom sklearn.model_selection import RandomizedSearchCV\nfrom sklearn.tree import DecisionTreeClassifier\nfrom sklearn.calibration import CalibratedClassifierCV\nfrom sklearn.neighbors import KNeighborsClassifier\nfrom sklearn.metrics import log_loss\nfrom sklearn.metrics import confusion_matrix\nfrom sklearn.model_selection import train_test_split\nfrom sklearn.linear_model import LogisticRegression\nfrom sklearn.ensemble import RandomForestClassifier\nimport re\nfrom nltk.util import ngrams\nfrom sklearn.feature_selection import SelectKBest, chi2, f_regression\n\nimport scipy.sparse\nimport gc\nimport pickle as pkl\nfrom datetime import datetime as dt\nimport dask.dataframe as dd","metadata":{"id":"9b6VlHUobK0i","editable":false},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# separating byte files and asm files \n# Below is from AML Assignment file\nfrom google.colab import drive\ndrive.mount('/content/gdrive')\n\nroot_path = '/content/gdrive/MyDrive/AML_Malware/Full_data/'\n# root_path = '../../LARGE_Datasets/'","metadata":{"editable":false},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#separating byte files and asm files \n\nsource = 'train'\ndestination_1 = 'byteFiles'\ndestination_2 = 'asmFiles'\n\n# we will check if the folder 'byteFiles' exists if it not there we will create a folder with the same name\nif not os.path.isdir(destination_1):\n    os.makedirs(destination_1)\nif not os.path.isdir(destination_2):\n    os.makedirs(destination_2)\n\n# if we have folder called 'train' (train folder contains both .asm files and .bytes files) we will rename it 'asmFiles'\n# for every file that we have in our 'asmFiles' directory we check if it is ending with .bytes, if yes we will move it to\n# 'byteFiles' folder\n\n# so by the end of this snippet we will separate all the .byte files and .asm files\nif os.path.isdir(source):\n    data_files = os.listdir(source)\n    for file in data_files:\n        print(file)\n        if (file.endswith(\"bytes\")):\n            shutil.move(source+'\\\\'+file,destination_1)\n        if (file.endswith(\"asm\")):\n            shutil.move(source+'\\\\'+file,destination_2)","metadata":{"id":"zeDuPii-bK0o","editable":false},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<h1 style=\"font-size:250%; font-family:cursive; color:#ff6666;\"><b> 8. Distribution of malware classes in whole data set <a id=\"8\"></a></b></h1>\n\n#### [Back to the top](#0)\n","metadata":{"id":"63z32BWibK0u","editable":false}},{"cell_type":"code","source":"Y=pd.read_csv(\"trainLabels.csv\")\ntotal = len(Y)*1.\nax=sns.countplot(x=\"Class\", data=Y)\nfor p in ax.patches:\n        ax.annotate('{:.1f}%'.format(100*p.get_height()/total), (p.get_x()+0.1, p.get_height()+5))\n\n#put 11 ticks (therefore 10 steps), from 0 to the total number of rows in the dataframe\nax.yaxis.set_ticks(np.linspace(0, total, 11))\n\n#adjust the ticklabel to the desired format, without changing the position of the ticks. \nax.set_yticklabels(map('{:.1f}%'.format, 100*ax.yaxis.get_majorticklocs()/total))\nplt.show()","metadata":{"id":"8-I5OxlfbK0v","outputId":"d45f650c-b673-49df-f9e7-5808349fa448","editable":false},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<h1 style=\"font-size:100%; font-family:cursive; color:#ff6666;\">[As mentioned earlier below plot is the image from my local machine's Notebook as this NB can NOT be run inside Kaggle Editor for the huge size of the dataset]</h1>\n\n\n![Imgur](https://imgur.com/lHoi1F0.png)\n\n<h1 style=\"font-size:250%; font-family:cursive; color:#ff6666;\"><b> 9. File size  of byte files as a feature <a id=\"9\"></a> </b></h1>\n\n#### [Back to the top](#0)\n","metadata":{"id":"MQUJaIigbK03","editable":false}},{"cell_type":"code","source":"#file sizes of byte files\n\nfiles=os.listdir('byteFiles')\nfilenames=Y['Id'].tolist()\nclass_y=Y['Class'].tolist()\nclass_bytes=[]\nsizebytes=[]\nfnames=[]\nfor file in files:\n    # print(os.stat('byteFiles/0A32eTdBKayjCWhZqDOQ.txt'))\n    # os.stat_result(st_mode=33206, st_ino=1125899906874507, st_dev=3561571700, st_nlink=1, st_uid=0, st_gid=0, \n    # st_size=3680109, st_atime=1519638522, st_mtime=1519638522, st_ctime=1519638522)\n    # read more about os.stat: here https://www.tutorialspoint.com/python/os_stat.htm\n    statinfo=os.stat('byteFiles/'+file)\n    # split the file name at '.' and take the first part of it i.e the file name\n    file=file.split('.')[0]\n    if any(file == filename for filename in filenames):\n        i=filenames.index(file)\n        class_bytes.append(class_y[i])\n        # converting into Mb's\n        sizebytes.append(statinfo.st_size/(1024.0*1024.0))\n        fnames.append(file)\ndata_size_byte=pd.DataFrame({'ID':fnames,'size':sizebytes,'Class':class_bytes})\nprint (data_size_byte.head())","metadata":{"id":"v8ux6wZPbK05","outputId":"21709fbb-4642-4ef8-881b-efbc4fea96ed","editable":false},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<h1 style=\"font-size:250%; font-family:cursive; color:#ff6666;\"><b> 10. box plots of file size (.byte files) feature <a id=\"10\"></a></b></h1>\n\n#### [Back to the top](#0)\n","metadata":{"id":"lB6lza8ybK0-","editable":false}},{"cell_type":"code","source":"#boxplot of byte files\nax = sns.boxplot(x=\"Class\", y=\"size\", data=data_size_byte)\nplt.title(\"boxplot of .bytes file sizes\")\nplt.show()","metadata":{"id":"IO2Y4mZIbK0_","outputId":"4d18a0d9-9dc9-44d3-fb92-3dc14601811a","editable":false},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<h1 style=\"font-size:100%; font-family:cursive; color:#ff6666;\">[As mentioned earlier below plot is the image from my local machine's Notebook as this NB can NOT be run inside Kaggle Editor for the huge size of the dataset]</h1>\n\n![Imgur](https://imgur.com/RgY5g3h.png)\n\n<h1 style=\"font-size:250%; font-family:cursive; color:#ff6666;\"><b> 11. Uni-Gram Byte Feature extraction from byte files <a id=\"11\"></a> </b></h1>\n\n#### [Back to the top](#0)\n\n","metadata":{"id":"4YQk-0mibK1H","editable":false}},{"cell_type":"code","source":"#removal of addres from byte files\n# contents of .byte files\n# ----------------\n#00401000 56 8D 44 24 08 50 8B F1 E8 1C 1B 00 00 C7 06 08 \n#-------------------\n#we remove the starting address 00401000\n\nfiles = os.listdir('byteFiles')\nfilenames=[]\narray=[]\nfor file in files:\n    if(file.endswith(\"bytes\")):\n        file=file.split('.')[0]\n        text_file = open('byteFiles/'+file+\".txt\", 'w+')\n        with open('byteFiles/'+file+\".bytes\",\"r\") as fp:\n            lines=\"\"\n            for line in fp:\n                a=line.rstrip().split(\" \")[1:]\n                b=' '.join(a)\n                b=b+\"\\n\"\n                text_file.write(b)\n            fp.close()\n            os.remove('byteFiles/'+file+\".bytes\")\n        text_file.close()\n\nfiles = os.listdir('byteFiles')\nfilenames2=[]\nfeature_matrix = np.zeros((len(files),257),dtype=int)\nk=0\n\n\n\n# program to convert into bag of words of bytefiles\n# this is custom-built bag of words this is unigram bag of words\n# This is a Custom Implementation of CountVectorizer as CountVectorizer will NOT suport working on such huge file system of 50GB\n# For this Uni-Gram feature creating and writing to a file named 'result.csv'\n\nbyte_feature_file=open('result.csv','w+')\nbyte_feature_file.write(\"ID,0,1,2,3,4,5,6,7,8,9,0a,0b,0c,0d,0e,0f,10,11,12,13,14,15,16,17,18,19,1a,1b,1c,1d,1e,1f,20,21,22,23,24,25,26,27,28,29,2a,2b,2c,2d,2e,2f,30,31,32,33,34,35,36,37,38,39,3a,3b,3c,3d,3e,3f,40,41,42,43,44,45,46,47,48,49,4a,4b,4c,4d,4e,4f,50,51,52,53,54,55,56,57,58,59,5a,5b,5c,5d,5e,5f,60,61,62,63,64,65,66,67,68,69,6a,6b,6c,6d,6e,6f,70,71,72,73,74,75,76,77,78,79,7a,7b,7c,7d,7e,7f,80,81,82,83,84,85,86,87,88,89,8a,8b,8c,8d,8e,8f,90,91,92,93,94,95,96,97,98,99,9a,9b,9c,9d,9e,9f,a0,a1,a2,a3,a4,a5,a6,a7,a8,a9,aa,ab,ac,ad,ae,af,b0,b1,b2,b3,b4,b5,b6,b7,b8,b9,ba,bb,bc,bd,be,bf,c0,c1,c2,c3,c4,c5,c6,c7,c8,c9,ca,cb,cc,cd,ce,cf,d0,d1,d2,d3,d4,d5,d6,d7,d8,d9,da,db,dc,dd,de,df,e0,e1,e2,e3,e4,e5,e6,e7,e8,e9,ea,eb,ec,ed,ee,ef,f0,f1,f2,f3,f4,f5,f6,f7,f8,f9,fa,fb,fc,fd,fe,ff,??\")\n\nbyte_feature_file.write(\"\\n\")\n\nfor file in files:\n    filenames2.append(file)\n    byte_feature_file.write(file+\",\")\n    if(file.endswith(\"txt\")):\n        with open('byteFiles/'+file,\"r\") as byte_flie:\n            for lines in byte_flie:\n                line=lines.rstrip().split(\" \")\n                for hex_code in line:\n                    if hex_code=='??':\n                        feature_matrix[k][256]+=1\n                    else:\n                        feature_matrix[k][int(hex_code,16)]+=1\n        byte_flie.close()\n    for i, row in enumerate(feature_matrix[k]):\n        if i!=len(feature_matrix[k])-1:\n            byte_feature_file.write(str(row)+\",\")\n        else:\n            byte_feature_file.write(str(row))\n    byte_feature_file.write(\"\\n\")\n    \n    k += 1\n\nbyte_feature_file.close()","metadata":{"id":"3WbQC_adbK1I","editable":false},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"byte_features=pd.read_csv(\"result.csv\")\nbyte_features['ID']  = byte_features['ID'].str.split('.').str[0]\nbyte_features.head(2)","metadata":{"id":"kHqmlixFbK1P","outputId":"58e9db6c-3dea-4d4e-9717-4c35cdb11d21","scrolled":true,"editable":false},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"data_size_byte.head(2)","metadata":{"id":"SNgJeQMpbK1T","outputId":"d46a8446-0cac-4720-d4ec-443f5694c408","editable":false},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"byte_features_with_size = byte_features.merge(data_size_byte, on='ID')\nbyte_features_with_size.to_csv(\"result_with_size.csv\")\nbyte_features_with_size.head(2)","metadata":{"id":"RoBSAHBIbK1a","outputId":"cf66cef2-daa7-46d1-c9f9-88777c15a1a4","editable":false},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# https://stackoverflow.com/a/29651514\ndef normalize(df):\n    result1 = df.copy()\n    for feature_name in df.columns:\n        if (str(feature_name) != str('ID') and str(feature_name)!=str('Class')):\n            max_value = df[feature_name].max()\n            min_value = df[feature_name].min()\n            result1[feature_name] = (df[feature_name] - min_value) / (max_value - min_value)\n    return result1\n\nresult = normalize(byte_features_with_size)","metadata":{"id":"tcdXq2hCbK1f","editable":false},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"result.head(2)","metadata":{"id":"PisxPTh1bK1l","outputId":"30d2dffd-7946-4e57-c24c-1038ddfafa81","editable":false},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"data_y = result['Class']\nresult.head()","metadata":{"id":"bZVENxs7bK1p","outputId":"de2bb83d-07e1-46b3-89ff-859064a0a4af","editable":false},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<h1 style=\"font-size:250%; font-family:cursive; color:#ff6666;\"><b> 12. Multivariate Analysis on byte files <a id=\"12\"></a></b></h1>\n\n#### [Back to the top](#0)\n","metadata":{"id":"U97CgVJtbK1u","editable":false}},{"cell_type":"code","source":"#multivariate analysis on byte files\n#this is with perplexity 50\nxtsne=TSNE(perplexity=50)\nresults=xtsne.fit_transform(result.drop(['ID','Class'], axis=1))\nvis_x = results[:, 0]\nvis_y = results[:, 1]\nplt.scatter(vis_x, vis_y, c=data_y, cmap=plt.cm.get_cmap(\"jet\", 9))\nplt.colorbar(ticks=range(10))\nplt.clim(0.5, 9)\nplt.show()","metadata":{"id":"YYrPdpnBbK1w","outputId":"4cc04874-a7dc-4810-bcd9-9b912e78bb5d","editable":false},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"![Imgur](https://imgur.com/kmzbnyv.png)","metadata":{"editable":false}},{"cell_type":"code","source":"#this is with perplexity 30\nxtsne=TSNE(perplexity=30)\nresults=xtsne.fit_transform(result.drop(['ID','Class'], axis=1))\nvis_x = results[:, 0]\nvis_y = results[:, 1]\nplt.scatter(vis_x, vis_y, c=data_y, cmap=plt.cm.get_cmap(\"jet\", 9))\nplt.colorbar(ticks=range(10))\nplt.clim(0.5, 9)\nplt.show()","metadata":{"id":"U55OLqnkbK1z","outputId":"3fda0508-ee8f-466c-db56-dd1b2167d586","editable":false},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<h1 style=\"font-size:100%; font-family:cursive; color:#ff6666;\">[As mentioned earlier below plot is the image from my local machine's Notebook as this NB can NOT be run inside Kaggle Editor for the huge size of the dataset]</h1>\n\n![Imgur](https://imgur.com/UDN49hh.png)\n\n\n<h1 style=\"font-size:250%; font-family:cursive; color:#ff6666;\"><b>13.  Train Test split of only Byte Files Features <a id=\"13\"></a> </b></h1>\n\n#### [Back to the top](#0)\n\n","metadata":{"id":"gOdPUILgbK12","editable":false}},{"cell_type":"code","source":"data_y = result['Class']\n# split the data into test and train by maintaining same distribution of output varaible 'y_true' [stratify=y_true]\nX_train, X_test, y_train, y_test = train_test_split(result.drop(['ID','Class'], axis=1), data_y,stratify=data_y,test_size=0.20)\n# split the train data into train and cross validation by maintaining same distribution of output varaible 'y_train' [stratify=y_train]\nX_train, X_cv, y_train, y_cv = train_test_split(X_train, y_train,stratify=y_train,test_size=0.20)","metadata":{"id":"55ctX0LubK13","editable":false},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print('Number of data points in train data:', X_train.shape[0])\nprint('Number of data points in test data:', X_test.shape[0])\nprint('Number of data points in cross validation data:', X_cv.shape[0])","metadata":{"id":"NJWMxDLVbK16","outputId":"b4c206f1-6dea-4681-fcc1-910e23f843de","editable":false},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# it returns a dict, keys as class labels and values as the number of data points in that class\ntrain_class_distribution = y_train.value_counts().sortlevel()\ntest_class_distribution = y_test.value_counts().sortlevel()\ncv_class_distribution = y_cv.value_counts().sortlevel()\n\nmy_colors = 'rgbkymc'\ntrain_class_distribution.plot(kind='bar', color=my_colors)\nplt.xlabel('Class')\nplt.ylabel('Data points per Class')\nplt.title('Distribution of yi in train data')\nplt.grid()\nplt.show()\n\n# ref: argsort https://docs.scipy.org/doc/numpy/reference/generated/numpy.argsort.html\n# -(train_class_distribution.values): the minus sign will give us in decreasing order\nsorted_yi = np.argsort(-train_class_distribution.values)\nfor i in sorted_yi:\n    print('Number of data points in class', i+1, ':',train_class_distribution.values[i], '(', np.round((train_class_distribution.values[i]/y_train.shape[0]*100), 3), '%)')\n\n    \nprint('-'*80)\nmy_colors = 'rgbkymc'\ntest_class_distribution.plot(kind='bar', color=my_colors)\nplt.xlabel('Class')\nplt.ylabel('Data points per Class')\nplt.title('Distribution of yi in test data')\nplt.grid()\nplt.show()\n\n# ref: argsort https://docs.scipy.org/doc/numpy/reference/generated/numpy.argsort.html\n# -(train_class_distribution.values): the minus sign will give us in decreasing order\nsorted_yi = np.argsort(-test_class_distribution.values)\nfor i in sorted_yi:\n    print('Number of data points in class', i+1, ':',test_class_distribution.values[i], '(', np.round((test_class_distribution.values[i]/y_test.shape[0]*100), 3), '%)')\n\nprint('-'*80)\nmy_colors = 'rgbkymc'\ncv_class_distribution.plot(kind='bar', color=my_colors)\nplt.xlabel('Class')\nplt.ylabel('Data points per Class')\nplt.title('Distribution of yi in cross validation data')\nplt.grid()\nplt.show()\n\n# ref: argsort https://docs.scipy.org/doc/numpy/reference/generated/numpy.argsort.html\n# -(train_class_distribution.values): the minus sign will give us in decreasing order\nsorted_yi = np.argsort(-train_class_distribution.values)\nfor i in sorted_yi:\n    print('Number of data points in class', i+1, ':',cv_class_distribution.values[i], '(', np.round((cv_class_distribution.values[i]/y_cv.shape[0]*100), 3), '%)')\n","metadata":{"id":"LqTqe7NMbK1-","outputId":"e12e3bd3-2fa4-41bd-c1df-57ce546850a3","editable":false},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"![Imgur](https://imgur.com/7OB6Ate.png)\n\n\n![Imgur](https://imgur.com/otn6pBx.png)\n\n\n![Imgur](https://imgur.com/6cfaSYq.png)","metadata":{"editable":false}},{"cell_type":"code","source":"def plot_confusion_matrix(test_y, predict_y):\n    C = confusion_matrix(test_y, predict_y)\n    print(\"Number of misclassified points \",(len(test_y)-np.trace(C))/len(test_y)*100)\n    # C = 9,9 matrix, each cell (i,j) represents number of points of class i are predicted class j\n    \n    A =(((C.T)/(C.sum(axis=1))).T)\n    #divid each element of the confusion matrix with the sum of elements in that column\n    \n    # C = [[1, 2],\n    #     [3, 4]]\n    # C.T = [[1, 3],\n    #        [2, 4]]\n    # C.sum(axis = 1)  axis=0 corresonds to columns and axis=1 corresponds to rows in two diamensional array\n    # C.sum(axix =1) = [[3, 7]]\n    # ((C.T)/(C.sum(axis=1))) = [[1/3, 3/7]\n    #                           [2/3, 4/7]]\n\n    # ((C.T)/(C.sum(axis=1))).T = [[1/3, 2/3]\n    #                           [3/7, 4/7]]\n    # sum of row elements = 1\n    \n    B =(C/C.sum(axis=0))\n    #divid each element of the confusion matrix with the sum of elements in that row\n    # C = [[1, 2],\n    #     [3, 4]]\n    # C.sum(axis = 0)  axis=0 corresonds to columns and axis=1 corresponds to rows in two diamensional array\n    # C.sum(axix =0) = [[4, 6]]\n    # (C/C.sum(axis=0)) = [[1/4, 2/6],\n    #                      [3/4, 4/6]] \n    \n    labels = [1,2,3,4,5,6,7,8,9]\n    cmap=sns.light_palette(\"green\")\n    # representing A in heatmap format\n    print(\"-\"*50, \"Confusion matrix\", \"-\"*50)\n    plt.figure(figsize=(10,5))\n    sns.heatmap(C, annot=True, cmap=cmap, fmt=\".3f\", xticklabels=labels, yticklabels=labels)\n    plt.xlabel('Predicted Class')\n    plt.ylabel('Original Class')\n    plt.show()\n\n    print(\"-\"*50, \"Precision matrix\", \"-\"*50)\n    plt.figure(figsize=(10,5))\n    sns.heatmap(B, annot=True, cmap=cmap, fmt=\".3f\", xticklabels=labels, yticklabels=labels)\n    plt.xlabel('Predicted Class')\n    plt.ylabel('Original Class')\n    plt.show()\n    print(\"Sum of columns in precision matrix\",B.sum(axis=0))\n    \n    # representing B in heatmap format\n    print(\"-\"*50, \"Recall matrix\"    , \"-\"*50)\n    plt.figure(figsize=(10,5))\n    sns.heatmap(A, annot=True, cmap=cmap, fmt=\".3f\", xticklabels=labels, yticklabels=labels)\n    plt.xlabel('Predicted Class')\n    plt.ylabel('Original Class')\n    plt.show()\n    print(\"Sum of rows in precision matrix\",A.sum(axis=1))","metadata":{"id":"qTWDK8s9bK2E","editable":false},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<h1 style=\"font-size:250%; font-family:cursive; color:#ff6666;\"><b> 14. Random Model ONLY on bytes files <a id=\"14\"></a> </b></h1>\n\n\n#### [Back to the top](#0)\n","metadata":{"id":"DTwTzxyabK2K","editable":false}},{"cell_type":"code","source":"\n# we need to generate 9 numbers and the sum of numbers should be 1\n# one solution is to genarate 9 numbers and divide each of the numbers by their sum\n# ref: https://stackoverflow.com/a/18662466/4084039\n\ntest_data_len = X_test.shape[0]\ncv_data_len = X_cv.shape[0]\n\n# we create a output array that has exactly same size as the CV data\ncv_predicted_y = np.zeros((cv_data_len,9))\nfor i in range(cv_data_len):\n    rand_probs = np.random.rand(1,9)\n    cv_predicted_y[i] = ((rand_probs/sum(sum(rand_probs)))[0])\nprint(\"Log loss on Cross Validation Data using Random Model\",log_loss(y_cv,cv_predicted_y, eps=1e-15))\n\n\n# Test-Set error.\n#we create a output array that has exactly same as the test data\ntest_predicted_y = np.zeros((test_data_len,9))\nfor i in range(test_data_len):\n    rand_probs = np.random.rand(1,9)\n    test_predicted_y[i] = ((rand_probs/sum(sum(rand_probs)))[0])\nprint(\"Log loss on Test Data using Random Model\",log_loss(y_test,test_predicted_y, eps=1e-15))\n\npredicted_y =np.argmax(test_predicted_y, axis=1)\nplot_confusion_matrix(y_test, predicted_y+1)","metadata":{"id":"yYfgTfvsbK2L","outputId":"0b50b55e-e9c7-4ff1-e1aa-676176a34124","editable":false},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"![Imgur](https://imgur.com/pbu6HJa.png)\n\n![Imgur](https://imgur.com/aIwGMMf.png)\n\n\n\n<h1 style=\"font-size:250%; font-family:cursive; color:#ff6666;\"><b> 15. K Nearest Neighbour Classification ONLY on bytes files <a id=\"15\"></a> </b></h1>\n\n#### [Back to the top](#0)\n","metadata":{"id":"7vM-Ea5WbK2P","editable":false}},{"cell_type":"code","source":"# find more about KNeighborsClassifier() here http://scikit-learn.org/stable/modules/generated/sklearn.neighbors.KNeighborsClassifier.html\n# -------------------------\n# default parameter\n# KNeighborsClassifier(n_neighbors=5, weights=’uniform’, algorithm=’auto’, leaf_size=30, p=2, \n# metric=’minkowski’, metric_params=None, n_jobs=1, **kwargs)\n\n# methods of\n# fit(X, y) : Fit the model using X as training data and y as target values\n# predict(X):Predict the class labels for the provided data\n# predict_proba(X):Return probability estimates for the test data X.\n\n# find more about CalibratedClassifierCV here at http://scikit-learn.org/stable/modules/generated/sklearn.calibration.CalibratedClassifierCV.html\n# ----------------------------\n# default paramters\n# sklearn.calibration.CalibratedClassifierCV(base_estimator=None, method=’sigmoid’, cv=3)\n#\n# some of the methods of CalibratedClassifierCV()\n# fit(X, y[, sample_weight])\tFit the calibrated model\n# get_params([deep])\tGet parameters for this estimator.\n# predict(X)\tPredict the target of new samples.\n# predict_proba(X)\tPosterior probabilities of classification\n\n  \nalpha = [x for x in range(1, 15, 2)]\ncv_log_error_array=[]\nfor i in alpha:\n    k_cfl=KNeighborsClassifier(n_neighbors=i)\n    k_cfl.fit(X_train,y_train)\n    sig_clf = CalibratedClassifierCV(k_cfl, method=\"sigmoid\")\n    sig_clf.fit(X_train, y_train)\n    predict_y = sig_clf.predict_proba(X_cv)\n    cv_log_error_array.append(log_loss(y_cv, predict_y, labels=k_cfl.classes_, eps=1e-15))\n    \nfor i in range(len(cv_log_error_array)):\n    print ('log_loss for k = ',alpha[i],'is',cv_log_error_array[i])\n\nbest_alpha = np.argmin(cv_log_error_array)\n    \nfig, ax = plt.subplots()\nax.plot(alpha, cv_log_error_array,c='g')\nfor i, txt in enumerate(np.round(cv_log_error_array,3)):\n    ax.annotate((alpha[i],np.round(txt,3)), (alpha[i],cv_log_error_array[i]))\nplt.grid()\nplt.title(\"Cross Validation Error for each alpha\")\nplt.xlabel(\"Alpha i's\")\nplt.ylabel(\"Error measure\")\nplt.show()\n\nk_cfl=KNeighborsClassifier(n_neighbors=alpha[best_alpha])\nk_cfl.fit(X_train,y_train)\nsig_clf = CalibratedClassifierCV(k_cfl, method=\"sigmoid\")\nsig_clf.fit(X_train, y_train)\n    \npredict_y = sig_clf.predict_proba(X_train)\nprint ('For values of best alpha = ', alpha[best_alpha], \"The train log loss is:\",log_loss(y_train, predict_y))\npredict_y = sig_clf.predict_proba(X_cv)\nprint('For values of best alpha = ', alpha[best_alpha], \"The cross validation log loss is:\",log_loss(y_cv, predict_y))\npredict_y = sig_clf.predict_proba(X_test)\nprint('For values of best alpha = ', alpha[best_alpha], \"The test log loss is:\",log_loss(y_test, predict_y))\nplot_confusion_matrix(y_test, sig_clf.predict(X_test))","metadata":{"id":"CtATBhZvbK2Q","outputId":"afa84b1d-f460-4a5b-ee0e-0987cfb13526","editable":false},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"![Imgur](https://imgur.com/H8r48hd.png)\n\n![Imgur](https://imgur.com/pmPMwBy.png)\n\n\n\n<h1 style=\"font-size:250%; font-family:cursive; color:#ff6666;\"><b> 16. Logistic Regression ONLY on bytes files <a id=\"16\"></a> </b></h1>\n\n#### [Back to the top](#0)\n","metadata":{"id":"svi0dvkRbK2V","editable":false}},{"cell_type":"code","source":"# read more about SGDClassifier() at http://scikit-learn.org/stable/modules/generated/sklearn.linear_model.SGDClassifier.html\n# ------------------------------\n# default parameters\n# SGDClassifier(loss=’hinge’, penalty=’l2’, alpha=0.0001, l1_ratio=0.15, fit_intercept=True, max_iter=None, tol=None, \n# shuffle=True, verbose=0, epsilon=0.1, n_jobs=1, random_state=None, learning_rate=’optimal’, eta0=0.0, power_t=0.5, \n# class_weight=None, warm_start=False, average=False, n_iter=None)\n\n# some of methods\n# fit(X, y[, coef_init, intercept_init, …])\tFit linear model with Stochastic Gradient Descent.\n# predict(X)\tPredict class labels for samples in X.\n\n\nalpha = [10 ** x for x in range(-5, 4)]\ncv_log_error_array=[]\nfor i in alpha:\n    logisticR=LogisticRegression(penalty='l2',C=i,class_weight='balanced')\n    logisticR.fit(X_train,y_train)\n    sig_clf = CalibratedClassifierCV(logisticR, method=\"sigmoid\")\n    sig_clf.fit(X_train, y_train)\n    predict_y = sig_clf.predict_proba(X_cv)\n    cv_log_error_array.append(log_loss(y_cv, predict_y, labels=logisticR.classes_, eps=1e-15))\n    \nfor i in range(len(cv_log_error_array)):\n    print ('log_loss for c = ',alpha[i],'is',cv_log_error_array[i])\n\nbest_alpha = np.argmin(cv_log_error_array)\n    \nfig, ax = plt.subplots()\nax.plot(alpha, cv_log_error_array,c='g')\nfor i, txt in enumerate(np.round(cv_log_error_array,3)):\n    ax.annotate((alpha[i],np.round(txt,3)), (alpha[i],cv_log_error_array[i]))\nplt.grid()\nplt.title(\"Cross Validation Error for each alpha\")\nplt.xlabel(\"Alpha i's\")\nplt.ylabel(\"Error measure\")\nplt.show()\n\nlogisticR=LogisticRegression(penalty='l2',C=alpha[best_alpha],class_weight='balanced')\nlogisticR.fit(X_train,y_train)\nsig_clf = CalibratedClassifierCV(logisticR, method=\"sigmoid\")\nsig_clf.fit(X_train, y_train)\npred_y=sig_clf.predict(X_test)\n\npredict_y = sig_clf.predict_proba(X_train)\nprint ('log loss for train data',log_loss(y_train, predict_y, labels=logisticR.classes_, eps=1e-15))\npredict_y = sig_clf.predict_proba(X_cv)\nprint ('log loss for cv data',log_loss(y_cv, predict_y, labels=logisticR.classes_, eps=1e-15))\npredict_y = sig_clf.predict_proba(X_test)\nprint ('log loss for test data',log_loss(y_test, predict_y, labels=logisticR.classes_, eps=1e-15))\nplot_confusion_matrix(y_test, sig_clf.predict(X_test))","metadata":{"id":"iETzW7stbK2V","outputId":"79e669d4-e0df-446c-ea18-7d146e9573e0","editable":false},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"\n![Imgur](https://imgur.com/WyHO5BE.png)\n\n![Imgur](https://imgur.com/c6z0Etn.png)\n\n\n<h1 style=\"font-size:250%; font-family:cursive; color:#ff6666;\"><b> 17. Random Forest Classifier ONLY on bytes files <a id=\"17\"></a> </b></h1>\n\n\n#### [Back to the top](#0)\n","metadata":{"id":"OHX3bYaDbK2c","editable":false}},{"cell_type":"code","source":"# --------------------------------\n# default parameters \n# sklearn.ensemble.RandomForestClassifier(n_estimators=10, criterion=’gini’, max_depth=None, min_samples_split=2, \n# min_samples_leaf=1, min_weight_fraction_leaf=0.0, max_features=’auto’, max_leaf_nodes=None, min_impurity_decrease=0.0, \n# min_impurity_split=None, bootstrap=True, oob_score=False, n_jobs=1, random_state=None, verbose=0, warm_start=False, \n# class_weight=None)\n\n# Some of methods of RandomForestClassifier()\n# fit(X, y, [sample_weight])\tFit the SVM model according to the given training data.\n# predict(X)\tPerform classification on samples in X.\n# predict_proba (X)\tPerform classification on samples in X.\n\n# some of attributes of  RandomForestClassifier()\n# feature_importances_ : array of shape = [n_features]\n# The feature importances (the higher, the more important the feature).\n\nalpha=[10,50,100,500,1000,2000,3000]\ncv_log_error_array=[]\ntrain_log_error_array=[]\nfrom sklearn.ensemble import RandomForestClassifier\nfor i in alpha:\n    r_cfl=RandomForestClassifier(n_estimators=i,random_state=42,n_jobs=-1)\n    r_cfl.fit(X_train,y_train)\n    sig_clf = CalibratedClassifierCV(r_cfl, method=\"sigmoid\")\n    sig_clf.fit(X_train, y_train)\n    predict_y = sig_clf.predict_proba(X_cv)\n    cv_log_error_array.append(log_loss(y_cv, predict_y, labels=r_cfl.classes_, eps=1e-15))\n\nfor i in range(len(cv_log_error_array)):\n    print ('log_loss for c = ',alpha[i],'is',cv_log_error_array[i])\n\n\nbest_alpha = np.argmin(cv_log_error_array)\n\nfig, ax = plt.subplots()\nax.plot(alpha, cv_log_error_array,c='g')\nfor i, txt in enumerate(np.round(cv_log_error_array,3)):\n    ax.annotate((alpha[i],np.round(txt,3)), (alpha[i],cv_log_error_array[i]))\nplt.grid()\nplt.title(\"Cross Validation Error for each alpha\")\nplt.xlabel(\"Alpha i's\")\nplt.ylabel(\"Error measure\")\nplt.show()\n\n\nr_cfl=RandomForestClassifier(n_estimators=alpha[best_alpha],random_state=42,n_jobs=-1)\nr_cfl.fit(X_train,y_train)\nsig_clf = CalibratedClassifierCV(r_cfl, method=\"sigmoid\")\nsig_clf.fit(X_train, y_train)\n\npredict_y = sig_clf.predict_proba(X_train)\nprint('For values of best alpha = ', alpha[best_alpha], \"The train log loss is:\",log_loss(y_train, predict_y))\npredict_y = sig_clf.predict_proba(X_cv)\nprint('For values of best alpha = ', alpha[best_alpha], \"The cross validation log loss is:\",log_loss(y_cv, predict_y))\npredict_y = sig_clf.predict_proba(X_test)\nprint('For values of best alpha = ', alpha[best_alpha], \"The test log loss is:\",log_loss(y_test, predict_y))\nplot_confusion_matrix(y_test, sig_clf.predict(X_test))","metadata":{"id":"xuZmPOKLbK2d","outputId":"59220231-75ac-43ee-b0a9-8d5d463c9389","editable":false},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"![Imgur](https://imgur.com/9sbkLUj.png)\n\n![Imgur](https://imgur.com/erFv4t6.png)\n\n<h1 style=\"font-size:250%; font-family:cursive; color:#ff6666;\"><b> 18. XgBoost Classification ONLY on bytes files <a id=\"18\"></a> </b></h1>\n","metadata":{"id":"bD30_F4cbK2g","editable":false}},{"cell_type":"code","source":"# Training a hyper-parameter tuned Xg-Boost regressor on our train data\n\n# find more about XGBClassifier function here http://xgboost.readthedocs.io/en/latest/python/python_api.html?#xgboost.XGBClassifier\n# -------------------------\n# default paramters\n# class xgboost.XGBClassifier(max_depth=3, learning_rate=0.1, n_estimators=100, silent=True, \n# objective='binary:logistic', booster='gbtree', n_jobs=1, nthread=None, gamma=0, min_child_weight=1, \n# max_delta_step=0, subsample=1, colsample_bytree=1, colsample_bylevel=1, reg_alpha=0, reg_lambda=1, \n# scale_pos_weight=1, base_score=0.5, random_state=0, seed=None, missing=None, **kwargs)\n\n# some of methods of RandomForestRegressor()\n# fit(X, y, sample_weight=None, eval_set=None, eval_metric=None, early_stopping_rounds=None, verbose=True, xgb_model=None)\n# get_params([deep])\tGet parameters for this estimator.\n# predict(data, output_margin=False, ntree_limit=0) : Predict with data. NOTE: This function is not thread safe.\n# get_score(importance_type='weight') -> get the feature importance\n\n\nalpha=[10,50,100,500,1000,2000]\ncv_log_error_array=[]\nfor i in alpha:\n    x_cfl=XGBClassifier(n_estimators=i,nthread=-1)\n    x_cfl.fit(X_train,y_train)\n    sig_clf = CalibratedClassifierCV(x_cfl, method=\"sigmoid\")\n    sig_clf.fit(X_train, y_train)\n    predict_y = sig_clf.predict_proba(X_cv)\n    cv_log_error_array.append(log_loss(y_cv, predict_y, labels=x_cfl.classes_, eps=1e-15))\n\nfor i in range(len(cv_log_error_array)):\n    print ('log_loss for c = ',alpha[i],'is',cv_log_error_array[i])\n\n\nbest_alpha = np.argmin(cv_log_error_array)\n\nfig, ax = plt.subplots()\nax.plot(alpha, cv_log_error_array,c='g')\nfor i, txt in enumerate(np.round(cv_log_error_array,3)):\n    ax.annotate((alpha[i],np.round(txt,3)), (alpha[i],cv_log_error_array[i]))\nplt.grid()\nplt.title(\"Cross Validation Error for each alpha\")\nplt.xlabel(\"Alpha i's\")\nplt.ylabel(\"Error measure\")\nplt.show()\n\nx_cfl=XGBClassifier(n_estimators=alpha[best_alpha],nthread=-1)\nx_cfl.fit(X_train,y_train)\nsig_clf = CalibratedClassifierCV(x_cfl, method=\"sigmoid\")\nsig_clf.fit(X_train, y_train)\n    \npredict_y = sig_clf.predict_proba(X_train)\nprint ('For values of best alpha = ', alpha[best_alpha], \"The train log loss is:\",log_loss(y_train, predict_y))\npredict_y = sig_clf.predict_proba(X_cv)\nprint('For values of best alpha = ', alpha[best_alpha], \"The cross validation log loss is:\",log_loss(y_cv, predict_y))\npredict_y = sig_clf.predict_proba(X_test)\nprint('For values of best alpha = ', alpha[best_alpha], \"The test log loss is:\",log_loss(y_test, predict_y))\nplot_confusion_matrix(y_test, sig_clf.predict(X_test))","metadata":{"id":"rWufRH6jbK2h","outputId":"8e5a099d-09c9-4645-d248-f57353ec9711","editable":false},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"![Imgur](https://imgur.com/t2ZrzLU.png)\n\n![Imgur](https://imgur.com/tDXxmiF.png)\n\n<h1 style=\"font-size:250%; font-family:cursive; color:#ff6666;\"><b>19. XgBoost Classification with best hyper parameters using RandomSearch ONLY on bytes files <a id=\"19\"></a> </b></h1>\n\n#### [Back to the top](#0)\n","metadata":{"id":"SmzFwq53bK2k","editable":false}},{"cell_type":"code","source":"# https://www.analyticsvidhya.com/blog/2016/03/complete-guide-parameter-tuning-xgboost-with-codes-python/\nx_cfl=XGBClassifier()\n\nprams={\n    'learning_rate':[0.01,0.03,0.05,0.1,0.15,0.2],\n     'n_estimators':[100,200,500,1000,2000],\n     'max_depth':[3,5,10],\n    'colsample_bytree':[0.1,0.3,0.5,1],\n    'subsample':[0.1,0.3,0.5,1]\n}\nrandom_cfl1=RandomizedSearchCV(x_cfl,param_distributions=prams,verbose=10,n_jobs=-1,)\nrandom_cfl1.fit(X_train,y_train)","metadata":{"id":"H7uj0MitbK2l","outputId":"f9e12da4-42aa-4a93-c69f-a97f1754a336","editable":false},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print (random_cfl1.best_params_)","metadata":{"id":"-KoC3ZfYbK2n","outputId":"1f68e74e-e2f4-4bac-9ff4-7a6f2a46b51f","editable":false},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Training a hyper-parameter tuned Xg-Boost regressor on our train data\n\n# find more about XGBClassifier function here http://xgboost.readthedocs.io/en/latest/python/python_api.html?#xgboost.XGBClassifier\n# -------------------------\n# default paramters\n# class xgboost.XGBClassifier(max_depth=3, learning_rate=0.1, n_estimators=100, silent=True, \n# objective='binary:logistic', booster='gbtree', n_jobs=1, nthread=None, gamma=0, min_child_weight=1, \n# max_delta_step=0, subsample=1, colsample_bytree=1, colsample_bylevel=1, reg_alpha=0, reg_lambda=1, \n# scale_pos_weight=1, base_score=0.5, random_state=0, seed=None, missing=None, **kwargs)\n\n# some of methods of RandomForestRegressor()\n# fit(X, y, sample_weight=None, eval_set=None, eval_metric=None, early_stopping_rounds=None, verbose=True, xgb_model=None)\n# get_params([deep])\tGet parameters for this estimator.\n# predict(data, output_margin=False, ntree_limit=0) : Predict with data. NOTE: This function is not thread safe.\n# get_score(importance_type='weight') -> get the feature importance\n\n\nx_cfl=XGBClassifier(n_estimators=2000, learning_rate=0.05, colsample_bytree=1, max_depth=3)\nx_cfl.fit(X_train,y_train)\nc_cfl=CalibratedClassifierCV(x_cfl,method='sigmoid')\nc_cfl.fit(X_train,y_train)\n\npredict_y = c_cfl.predict_proba(X_train)\nprint ('train loss',log_loss(y_train, predict_y))\npredict_y = c_cfl.predict_proba(X_cv)\nprint ('cv loss',log_loss(y_cv, predict_y))\npredict_y = c_cfl.predict_proba(X_test)\nprint ('test loss',log_loss(y_test, predict_y))","metadata":{"id":"w6YasjdnbK2p","outputId":"e0f1c456-9eef-4126-b35f-c814f30a8cce","editable":false},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<h1 style=\"font-size:250%; font-family:cursive; color:#ff6666;\"><b> 20. Modeling with .asm files <a id=\"20\"></a> </b></h1>\n\n#### [Back to the top](#0)\n\n\nThere are 10868 files of asm \nAll the files make up about 150 GB\nThe asm files contains :\n1. Address\n2. Segments\n3. Opcodes\n4. Registers\n5. function calls\n6. APIs\nWith the help of parallel processing we extracted all the features.In parallel we can use all the cores that are present in our computer.\n\n\nHere we extracted 52 features from all the asm files which are important.\n\nWe read the top solutions and handpicked the features from those papers/videos/blogs. <br> Refer:https://www.kaggle.com/c/malware-classification/discussion\n\n## A note on opcode\n\n“opcode” is short for operational code. These are the bytes stored in memory that the computer actually runs.\n\n[Here](https://www.intel.com/content/dam/www/public/us/en/documents/manuals/64-ia-32-architectures-software-developer-instruction-set-reference-manual-325383.pdf) you will see all the opcodes that a processor supports. An assembler basically takes text and does some relatively simple conversion of it into a file of opcodes that the computer can read and run directly. Most assemble translate very directly to opcodes. And often the 3 to 5 character assembler names for the opcodes are called opcodes. Very technically the opcodes are the binary numbers stored in memory. The names for them in assembler are the opcode names. Also technically not all the binary numbers stored are the opcodes. The opcodes tell the CPU what to do. Often the number right after the opcodes are parameters for the opcode instructions.\n\nAlso Check [this very complete table of x86 opcodes on x86asm.net](http://ref.x86asm.net/coder32.html) and [this](https://www.sandpile.org/) as well.\n\nThere is also [asmjit/asmdb][1] project, which provides public domain [X86/X64 database][2] in a JSON-like format \n\n[1]: https://github.com/asmjit/asmdb\n[2]: https://github.com/asmjit/asmdb/blob/master/x86data.js\n\n[Referring this Paper](https://www.semanticscholar.org/paper/Malware-detection-through-opcode-sequence-analysis-Bragen/62f796c19ffa2ee70fc5ee7aec0fe41fae26f191)\n\nIn this thesis they used reverse engineering to extract the assembly instructions from a given executable file and chose to use only the opcodes, which are the part of the instruction that specifies the operation to be performed, an example \"mov\".\n\n\nBy performing statistical analysis on the datasets, a significant difference between the opcodes in malware and benign files was found. Due to this, supervised and unsupervised machine learning approaches like artificial neural network, support vector machine, bayes net, random forest, k nearest neighbours, and self organizing map was used to look at the sequences of these\ninstructions. The unknown files were classified as either malware or benign depending on the presence of, and number of occurrences of different sequences. We show that by using only opcodes without operands (the rest of the instruction), malware can be distinguished from benign files. By using a sequence length of up to four opcodes, a classification accuracy of 95,58% was achieved. \n\n<h1 style=\"font-size:250%; font-family:cursive; color:#ff6666;\"><b> 21. Feature extraction from asm files <a id=\"21\"></a> </b></h1>\n\n#### [Back to the top](#0)\n\n<p>\n<li> To extract the unigram features from the .asm files we need to process ~150GB of data </li>\n<li style=\"font-size:18px\"><b>Note: Below two cells will take lot of time (over 48 hours to complete)</b></li>\n<li> We will provide you the output file of these two cells, which you can directly use it </li>\n</p>","metadata":{"id":"T_-Au021bK2t","editable":false}},{"cell_type":"code","source":"# This code taken from https://github.com/kunwar-vikrant/Microsoft-Malware-Detection\n#intially create five folders\n#first \n#second\n#thrid\n#fourth\n#fifth\n#this code tells us about random split of files into five folders\nfolder_1 ='first'\nfolder_2 ='second'\nfolder_3 ='third'\nfolder_4 ='fourth'\nfolder_5 ='fifth'\nfolder_6 = 'output'\nfor i in [folder_1,folder_2,folder_3,folder_4,folder_5,folder_6]:\n    if not os.path.isdir(i):\n        os.makedirs(i)\n\nsource='train/'\nfiles = os.listdir('train')\nID=df['Id'].tolist()\ndata=range(0,10868)\nr.shuffle(data)\ncount=0\nfor i in range(0,10868):\n    if i % 5==0:\n        shutil.move(source+files[data[i]],'first')\n    elif i%5==1:\n        shutil.move(source+files[data[i]],'second')\n    elif i%5 ==2:\n        shutil.move(source+files[data[i]],'thrid')\n    elif i%5 ==3:\n        shutil.move(source+files[data[i]],'fourth')\n    elif i%5==4:\n        shutil.move(source+files[data[i]],'fifth')","metadata":{"id":"DHaogo_fbK2w","editable":false},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"\n# http://flint.cs.yale.edu/cs421/papers/x86-asm/asm.html\n\ndef firstprocess():\n    #The prefixes tells about the segments that are present in the asm files\n    #There are 450 segments(approx) present in all asm files.\n    #this prefixes are best segments that gives us best values.\n    #https://en.wikipedia.org/wiki/Data_segment\n    \n    prefixes = ['HEADER:','.text:','.Pav:','.idata:','.data:','.bss:','.rdata:','.edata:','.rsrc:','.tls:','.reloc:','.BSS:','.CODE']\n    #this are opcodes that are used to get best results\n    #https://en.wikipedia.org/wiki/X86_instruction_listings\n    \n    opcodes = ['jmp', 'mov', 'retf', 'push', 'pop', 'xor', 'retn', 'nop', 'sub', 'inc', 'dec', 'add','imul', 'xchg', 'or', 'shr', 'cmp', 'call', 'shl', 'ror', 'rol', 'jnb','jz','rtn','lea','movzx']\n    #best keywords that are taken from different blogs\n    keywords = ['.dll','std::',':dword']\n    #Below taken registers are general purpose registers and special registers\n    #All the registers which are taken are best \n    registers=['edx','esi','eax','ebx','ecx','edi','ebp','esp','eip']\n    file1=open(\"output\\asmsmallfile.txt\",\"w+\")\n    files = os.listdir('first')\n    for f in files:\n        #filling the values with zeros into the arrays\n        prefixescount=np.zeros(len(prefixes),dtype=int)\n        opcodescount=np.zeros(len(opcodes),dtype=int)\n        keywordcount=np.zeros(len(keywords),dtype=int)\n        registerscount=np.zeros(len(registers),dtype=int)\n        features=[]\n        f2=f.split('.')[0]\n        file1.write(f2+\",\")\n        opcodefile.write(f2+\" \")\n        # https://docs.python.org/3/library/codecs.html#codecs.ignore_errors\n        # https://docs.python.org/3/library/codecs.html#codecs.Codec.encode\n        with codecs.open('first/'+f,encoding='cp1252',errors ='replace') as fli:\n            for lines in fli:\n                # https://www.tutorialspoint.com/python3/string_rstrip.htm\n                line=lines.rstrip().split()\n                l=line[0]\n                #counting the prefixs in each and every line\n                for i in range(len(prefixes)):\n                    if prefixes[i] in line[0]:\n                        prefixescount[i]+=1\n                line=line[1:]\n                #counting the opcodes in each and every line\n                for i in range(len(opcodes)):\n                    if any(opcodes[i]==li for li in line):\n                        features.append(opcodes[i])\n                        opcodescount[i]+=1\n                #counting registers in the line\n                for i in range(len(registers)):\n                    for li in line:\n                        # we will use registers only in 'text' and 'CODE' segments\n                        if registers[i] in li and ('text' in l or 'CODE' in l):\n                            registerscount[i]+=1\n                #counting keywords in the line\n                for i in range(len(keywords)):\n                    for li in line:\n                        if keywords[i] in li:\n                            keywordcount[i]+=1\n        #pushing the values into the file after reading whole file\n        for prefix in prefixescount:\n            file1.write(str(prefix)+\",\")\n        for opcode in opcodescount:\n            file1.write(str(opcode)+\",\")\n        for register in registerscount:\n            file1.write(str(register)+\",\")\n        for key in keywordcount:\n            file1.write(str(key)+\",\")\n        file1.write(\"\\n\")\n    file1.close()\n\n\n#same as above \ndef secondprocess():\n    prefixes = ['HEADER:','.text:','.Pav:','.idata:','.data:','.bss:','.rdata:','.edata:','.rsrc:','.tls:','.reloc:','.BSS:','.CODE']\n    opcodes = ['jmp', 'mov', 'retf', 'push', 'pop', 'xor', 'retn', 'nop', 'sub', 'inc', 'dec', 'add','imul', 'xchg', 'or', 'shr', 'cmp', 'call', 'shl', 'ror', 'rol', 'jnb','jz','rtn','lea','movzx']\n    keywords = ['.dll','std::',':dword']\n    registers=['edx','esi','eax','ebx','ecx','edi','ebp','esp','eip']\n    file1=open(\"output\\mediumasmfile.txt\",\"w+\")\n    files = os.listdir('second')\n    for f in files:\n        prefixescount=np.zeros(len(prefixes),dtype=int)\n        opcodescount=np.zeros(len(opcodes),dtype=int)\n        keywordcount=np.zeros(len(keywords),dtype=int)\n        registerscount=np.zeros(len(registers),dtype=int)\n        features=[]\n        f2=f.split('.')[0]\n        file1.write(f2+\",\")\n        opcodefile.write(f2+\" \")\n        with codecs.open('second/'+f,encoding='cp1252',errors ='replace') as fli:\n            for lines in fli:\n                line=lines.rstrip().split()\n                l=line[0]\n                for i in range(len(prefixes)):\n                    if prefixes[i] in line[0]:\n                        prefixescount[i]+=1\n                line=line[1:]\n                for i in range(len(opcodes)):\n                    if any(opcodes[i]==li for li in line):\n                        features.append(opcodes[i])\n                        opcodescount[i]+=1\n                for i in range(len(registers)):\n                    for li in line:\n                        if registers[i] in li and ('text' in l or 'CODE' in l):\n                            registerscount[i]+=1\n                for i in range(len(keywords)):\n                    for li in line:\n                        if keywords[i] in li:\n                            keywordcount[i]+=1\n        for prefix in prefixescount:\n            file1.write(str(prefix)+\",\")\n        for opcode in opcodescount:\n            file1.write(str(opcode)+\",\")\n        for register in registerscount:\n            file1.write(str(register)+\",\")\n        for key in keywordcount:\n            file1.write(str(key)+\",\")\n        file1.write(\"\\n\")\n    file1.close()\n\n# same as smallprocess() functions\ndef thirdprocess():\n    prefixes = ['HEADER:','.text:','.Pav:','.idata:','.data:','.bss:','.rdata:','.edata:','.rsrc:','.tls:','.reloc:','.BSS:','.CODE']\n    opcodes = ['jmp', 'mov', 'retf', 'push', 'pop', 'xor', 'retn', 'nop', 'sub', 'inc', 'dec', 'add','imul', 'xchg', 'or', 'shr', 'cmp', 'call', 'shl', 'ror', 'rol', 'jnb','jz','rtn','lea','movzx']\n    keywords = ['.dll','std::',':dword']\n    registers=['edx','esi','eax','ebx','ecx','edi','ebp','esp','eip']\n    file1=open(\"output\\largeasmfile.txt\",\"w+\")\n    files = os.listdir('thrid')\n    for f in files:\n        prefixescount=np.zeros(len(prefixes),dtype=int)\n        opcodescount=np.zeros(len(opcodes),dtype=int)\n        keywordcount=np.zeros(len(keywords),dtype=int)\n        registerscount=np.zeros(len(registers),dtype=int)\n        features=[]\n        f2=f.split('.')[0]\n        file1.write(f2+\",\")\n        opcodefile.write(f2+\" \")\n        with codecs.open('thrid/'+f,encoding='cp1252',errors ='replace') as fli:\n            for lines in fli:\n                line=lines.rstrip().split()\n                l=line[0]\n                for i in range(len(prefixes)):\n                    if prefixes[i] in line[0]:\n                        prefixescount[i]+=1\n                line=line[1:]\n                for i in range(len(opcodes)):\n                    if any(opcodes[i]==li for li in line):\n                        features.append(opcodes[i])\n                        opcodescount[i]+=1\n                for i in range(len(registers)):\n                    for li in line:\n                        if registers[i] in li and ('text' in l or 'CODE' in l):\n                            registerscount[i]+=1\n                for i in range(len(keywords)):\n                    for li in line:\n                        if keywords[i] in li:\n                            keywordcount[i]+=1\n        for prefix in prefixescount:\n            file1.write(str(prefix)+\",\")\n        for opcode in opcodescount:\n            file1.write(str(opcode)+\",\")\n        for register in registerscount:\n            file1.write(str(register)+\",\")\n        for key in keywordcount:\n            file1.write(str(key)+\",\")\n        file1.write(\"\\n\")\n    file1.close()\n\n\ndef fourthprocess():\n    prefixes = ['HEADER:','.text:','.Pav:','.idata:','.data:','.bss:','.rdata:','.edata:','.rsrc:','.tls:','.reloc:','.BSS:','.CODE']\n    opcodes = ['jmp', 'mov', 'retf', 'push', 'pop', 'xor', 'retn', 'nop', 'sub', 'inc', 'dec', 'add','imul', 'xchg', 'or', 'shr', 'cmp', 'call', 'shl', 'ror', 'rol', 'jnb','jz','rtn','lea','movzx']\n    keywords = ['.dll','std::',':dword']\n    registers=['edx','esi','eax','ebx','ecx','edi','ebp','esp','eip']\n    file1=open(\"output\\hugeasmfile.txt\",\"w+\")\n    files = os.listdir('fourth/')\n    for f in files:\n        prefixescount=np.zeros(len(prefixes),dtype=int)\n        opcodescount=np.zeros(len(opcodes),dtype=int)\n        keywordcount=np.zeros(len(keywords),dtype=int)\n        registerscount=np.zeros(len(registers),dtype=int)\n        features=[]\n        f2=f.split('.')[0]\n        file1.write(f2+\",\")\n        opcodefile.write(f2+\" \")\n        with codecs.open('fourth/'+f,encoding='cp1252',errors ='replace') as fli:\n            for lines in fli:\n                line=lines.rstrip().split()\n                l=line[0]\n                for i in range(len(prefixes)):\n                    if prefixes[i] in line[0]:\n                        prefixescount[i]+=1\n                line=line[1:]\n                for i in range(len(opcodes)):\n                    if any(opcodes[i]==li for li in line):\n                        features.append(opcodes[i])\n                        opcodescount[i]+=1\n                for i in range(len(registers)):\n                    for li in line:\n                        if registers[i] in li and ('text' in l or 'CODE' in l):\n                            registerscount[i]+=1\n                for i in range(len(keywords)):\n                    for li in line:\n                        if keywords[i] in li:\n                            keywordcount[i]+=1\n        for prefix in prefixescount:\n            file1.write(str(prefix)+\",\")\n        for opcode in opcodescount:\n            file1.write(str(opcode)+\",\")\n        for register in registerscount:\n            file1.write(str(register)+\",\")\n        for key in keywordcount:\n            file1.write(str(key)+\",\")\n        file1.write(\"\\n\")\n    file1.close()\n\n\ndef fifthprocess():\n    prefixes = ['HEADER:','.text:','.Pav:','.idata:','.data:','.bss:','.rdata:','.edata:','.rsrc:','.tls:','.reloc:','.BSS:','.CODE']\n    opcodes = ['jmp', 'mov', 'retf', 'push', 'pop', 'xor', 'retn', 'nop', 'sub', 'inc', 'dec', 'add','imul', 'xchg', 'or', 'shr', 'cmp', 'call', 'shl', 'ror', 'rol', 'jnb','jz','rtn','lea','movzx']\n    keywords = ['.dll','std::',':dword']\n    registers=['edx','esi','eax','ebx','ecx','edi','ebp','esp','eip']\n    file1=open(\"output\\trainasmfile.txt\",\"w+\")\n    files = os.listdir('fifth/')\n    for f in files:\n        prefixescount=np.zeros(len(prefixes),dtype=int)\n        opcodescount=np.zeros(len(opcodes),dtype=int)\n        keywordcount=np.zeros(len(keywords),dtype=int)\n        registerscount=np.zeros(len(registers),dtype=int)\n        features=[]\n        f2=f.split('.')[0]\n        file1.write(f2+\",\")\n        opcodefile.write(f2+\" \")\n        with codecs.open('fifth/'+f,encoding='cp1252',errors ='replace') as fli:\n            for lines in fli:\n                line=lines.rstrip().split()\n                l=line[0]\n                for i in range(len(prefixes)):\n                    if prefixes[i] in line[0]:\n                        prefixescount[i]+=1\n                line=line[1:]\n                for i in range(len(opcodes)):\n                    if any(opcodes[i]==li for li in line):\n                        features.append(opcodes[i])\n                        opcodescount[i]+=1\n                for i in range(len(registers)):\n                    for li in line:\n                        if registers[i] in li and ('text' in l or 'CODE' in l):\n                            registerscount[i]+=1\n                for i in range(len(keywords)):\n                    for li in line:\n                        if keywords[i] in li:\n                            keywordcount[i]+=1\n        for prefix in prefixescount:\n            file1.write(str(prefix)+\",\")\n        for opcode in opcodescount:\n            file1.write(str(opcode)+\",\")\n        for register in registerscount:\n            file1.write(str(register)+\",\")\n        for key in keywordcount:\n            file1.write(str(key)+\",\")\n        file1.write(\"\\n\")\n    file1.close()\n\n\ndef main():\n    #the below code is used for multiprogramming\n    #the number of process depends upon the number of cores present System\n    #process is used to call multiprogramming\n    manager=multiprocessing.Manager() \t\n    p1=Process(target=firstprocess)\n    p2=Process(target=secondprocess)\n    p3=Process(target=thirdprocess)\n    p4=Process(target=fourthprocess)\n    p5=Process(target=fifthprocess)\n    #p1.start() is used to start the thread execution\n    p1.start()\n    p2.start()\n    p3.start()\n    p4.start()\n    p5.start()\n    #After completion all the threads are joined\n    p1.join()\n    p2.join()\n    p3.join()\n    p4.join()\n    p5.join()\n\nif __name__==\"__main__\":\n    main()","metadata":{"id":"roJFLqd9bK2y","editable":false},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# asmoutputfile.csv(output genarated from the above two cells) will contain all the extracted features from .asm files\n# we will use this file directly\ndfasm=pd.read_csv(\"asmoutputfile.csv\")\nY.columns = ['ID', 'Class']\nresult_asm = pd.merge(dfasm, Y,on='ID', how='left')\nresult_asm.head()","metadata":{"id":"hmlRQ9HkbK24","outputId":"1aeef80a-1a5a-4fa0-f885-74ee9a8b64ef","editable":false},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<h1 style=\"font-size:250%; font-family:cursive; color:#ff6666;\"><b> 22. Files sizes of each .asm file as a feature <a id=\"22\"></a> </b></h1>\n\n#### [Back to the top](#0)\n","metadata":{"id":"Bjk8MHL3bK2_","editable":false}},{"cell_type":"code","source":"# file sizes of asm files\n\nfiles=os.listdir('asmFiles')\nfilenames=Y['ID'].tolist()\nclass_y=Y['Class'].tolist()\nclass_bytes=[]\nsizebytes=[]\nfnames=[]\nfor file in files:\n    # print(os.stat('byteFiles/0A32eTdBKayjCWhZqDOQ.txt'))\n    # os.stat_result(st_mode=33206, st_ino=1125899906874507, st_dev=3561571700, st_nlink=1, st_uid=0, st_gid=0, \n    # st_size=3680109, st_atime=1519638522, st_mtime=1519638522, st_ctime=1519638522)\n    # read more about os.stat: here https://www.tutorialspoint.com/python/os_stat.htm\n    statinfo=os.stat('asmFiles/'+file)\n    # split the file name at '.' and take the first part of it i.e the file name\n    file=file.split('.')[0]\n    if any(file == filename for filename in filenames):\n        i=filenames.index(file)\n        class_bytes.append(class_y[i])\n        # converting into Mb's\n        sizebytes.append(statinfo.st_size/(1024.0*1024.0))\n        fnames.append(file)\nasm_size_byte=pd.DataFrame({'ID':fnames,'size':sizebytes,'Class':class_bytes})\nprint (asm_size_byte.head())","metadata":{"id":"GWSAKuydbK3A","outputId":"52eccff1-f6a5-46de-da70-42b3bf533e7f","editable":false},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<h4> 4.2.1.2 Distribution of .asm file sizes</h4>","metadata":{"id":"AwhygtOBbK3D","editable":false}},{"cell_type":"code","source":"#boxplot of asm files\nax = sns.boxplot(x=\"Class\", y=\"size\", data=asm_size_byte)\nplt.title(\"boxplot of .bytes file sizes\")\nplt.show()","metadata":{"id":"EQzOsTAcbK3F","outputId":"9a575cd6-90d1-42c0-bccb-be79f11a1770","editable":false},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"![Imgur](https://imgur.com/egYeXAJ.png)","metadata":{"editable":false}},{"cell_type":"code","source":"# add the file size feature to previous extracted features\nprint(result_asm.shape)\nprint(asm_size_byte.shape)\nresult_asm = pd.merge(result_asm, asm_size_byte.drop(['Class'], axis=1),on='ID', how='left')\nresult_asm.head()","metadata":{"id":"hJzB8df0bK3I","outputId":"1ba1693e-705e-4d39-a230-1ce7888d2a13","editable":false},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# we normalize the data each column \nresult_asm = normalize(result_asm)\nresult_asm.head()","metadata":{"id":"eyLZyAxqbK3K","outputId":"fd6dcfc2-60c7-4193-f985-5dabdd53d15b","editable":false},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<h1 style=\"font-size:250%; font-family:cursive; color:#ff6666;\"><b> 23. Univariate analysis ONLY on .asm file features <a id=\"23\"></a></b></h1>\n\n#### [Back to the top](#0)\n","metadata":{"id":"yfFvqYLrbK3N","editable":false}},{"cell_type":"code","source":"ax = sns.boxplot(x=\"Class\", y=\".text:\", data=result_asm)\nplt.title(\"boxplot of .asm text segment\")\nplt.show()","metadata":{"id":"yDExe1AxbK3N","outputId":"1db168a4-4bce-4d82-d963-58410628b0da","editable":false},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"\n![Imgur](https://imgur.com/5jWiNtY.png)\n\n<pre>\nThe plot is between Text and class \nClass 1,2 and 9 can be easly separated\n</pre>","metadata":{"id":"EwBFZVcYbK3Q","editable":false}},{"cell_type":"code","source":"ax = sns.boxplot(x=\"Class\", y=\".Pav:\", data=result_asm)\nplt.title(\"boxplot of .asm pav segment\")\nplt.show()","metadata":{"id":"MXy7gj93bK3Q","outputId":"f4fd2614-9c20-4e19-e24c-f72bc34f5d06","editable":false},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"![Imgur](https://imgur.com/clvpMB9.png)","metadata":{"editable":false}},{"cell_type":"code","source":"ax = sns.boxplot(x=\"Class\", y=\".data:\", data=result_asm)\nplt.title(\"boxplot of .asm data segment\")\nplt.show()","metadata":{"id":"tpCU33S5bK3S","outputId":"67693b33-dea5-49ef-f7d3-87edff01e0dd","editable":false},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"![Imgur](https://imgur.com/CqJhugg.png)\n\n<pre>\nThe plot is between data segment and class label \nclass 6 and class 9 can be easily separated from given points\n</pre>","metadata":{"id":"CMU2PS-NbK3V","editable":false}},{"cell_type":"code","source":"ax = sns.boxplot(x=\"Class\", y=\".bss:\", data=result_asm)\nplt.title(\"boxplot of .asm bss segment\")\nplt.show()","metadata":{"id":"mGWYFGRobK3W","outputId":"2ab0745b-1bca-4b1f-fbd1-5f50a9b70a8e","editable":false},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"![Imgur](https://imgur.com/GKa73JO.png)\n\n<pre>\nplot between bss segment and class label\nvery less number of files are having bss segment\n</pre>","metadata":{"id":"UANZiDq-bK3X","editable":false}},{"cell_type":"code","source":"ax = sns.boxplot(x=\"Class\", y=\".rdata:\", data=result_asm)\nplt.title(\"boxplot of .asm rdata segment\")\nplt.show()","metadata":{"id":"q4TeHYTUbK3X","outputId":"526879e1-cda0-4bf6-af92-7178a6b79d8a","editable":false},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"\n[Imgur](https://imgur.com/SPZxLJL.png)\n\n<pre>\nPlot between rdata segment and Class segment\nClass 2 can be easily separated 75 pecentile files are having 1M rdata lines\n</pre>","metadata":{"id":"U_RYjUIVbK3a","editable":false}},{"cell_type":"code","source":"ax = sns.boxplot(x=\"Class\", y=\"jmp\", data=result_asm)\nplt.title(\"boxplot of .asm jmp opcode\")\nplt.show()","metadata":{"id":"1-KJ8vIVbK3b","outputId":"8656e511-f460-4787-e54b-bc2be6ad3b17","editable":false},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"![Imgur](https://imgur.com/0e0ylU2.png)\n\n\n<pre>\nplot between jmp and Class label\nClass 1 is having frequency of 2000 approx in 75 perentile of files\n</pre>","metadata":{"id":"I3-zVf-VbK3i","editable":false}},{"cell_type":"code","source":"ax = sns.boxplot(x=\"Class\", y=\"mov\", data=result_asm)\nplt.title(\"boxplot of .asm mov opcode\")\nplt.show()","metadata":{"id":"pCHwFsOqbK3i","outputId":"77a09e12-f462-403f-95e1-bf722a52b9dc","editable":false},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"\n![Imgur](https://imgur.com/Jr5dOJk.png)\n\n\n<pre>\nplot between Class label and mov opcode\nClass 1 is having frequency of 2000 approx in 75 perentile of files\n</pre>","metadata":{"id":"pNA246YLbK3k","editable":false}},{"cell_type":"code","source":"ax = sns.boxplot(x=\"Class\", y=\"retf\", data=result_asm)\nplt.title(\"boxplot of .asm retf opcode\")\nplt.show()","metadata":{"id":"XiJLkwBGbK3l","outputId":"c4dc803d-509a-48b6-bc6f-c26846be24df","editable":false},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"![Imgur](https://imgur.com/VQ25RTI.png)\n\n\n<pre>\nplot between Class label and retf\nClass 6 can be easily separated with opcode retf\nThe frequency of retf is approx of 250.\n</pre>","metadata":{"id":"5NPAvXqPbK3o","editable":false}},{"cell_type":"code","source":"ax = sns.boxplot(x=\"Class\", y=\"push\", data=result_asm)\nplt.title(\"boxplot of .asm push opcode\")\nplt.show()","metadata":{"id":"EesJq3mBbK3p","outputId":"86b1013c-b792-4d1a-ebbb-1c51ec32dac7","editable":false},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"\n![Imgur](https://imgur.com/FLpSOdK.png)\n\n<pre>\nplot between push opcode and Class label\nClass 1 is having 75 precentile files with push opcodes of frequency 1000\n</pre>","metadata":{"id":"fp-g2xzEbK3r","editable":false}},{"cell_type":"markdown","source":"<h1 style=\"font-size:250%; font-family:cursive; color:#ff6666;\"><b>24. Multivariate Analysis ONLY on .asm file features <a id=\"24\"></a> </b></h1>\n\n#### [Back to the top](#0)\n","metadata":{"id":"OhWGpkG_bK3s","editable":false}},{"cell_type":"code","source":"#multivariate analysis on asm files\n#this is with perplexity 50\nxtsne=TSNE(perplexity=50)\nresults=xtsne.fit_transform(result_asm.drop(['ID','Class'], axis=1).fillna(0))\n\nvis_x = results[:, 0]\nvis_y = results[:, 1   ]\nplt.scatter(vis_x, vis_y, c=data_y, cmap=plt.cm.get_cmap(\"jet\", 9))\nplt.colorbar(ticks=range(10))\nplt.clim(0.5, 9)\nplt.show()","metadata":{"id":"fEr7P2kpbK3s","outputId":"67148157-714a-4c33-a171-d171da86fe8e","editable":false},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"![Imgur](https://imgur.com/tR4nhGB.png)","metadata":{"editable":false}},{"cell_type":"code","source":"# by univariate analysis on the .asm file features we are getting very negligible information from \n# 'rtn', '.BSS:' '.CODE' features, so heare we are trying multivariate analysis after removing those features\n# the plot looks very messy\n\nxtsne=TSNE(perplexity=30)\nresults=xtsne.fit_transform(result_asm.drop(['ID','Class', 'rtn', '.BSS:', '.CODE','size'], axis=1))\nvis_x = results[:, 0]\nvis_y = results[:, 1]\nplt.scatter(vis_x, vis_y, c=data_y, cmap=plt.cm.get_cmap(\"jet\", 9))\nplt.colorbar(ticks=range(10))\nplt.clim(0.5, 9)\nplt.show()","metadata":{"id":"xCxF_wFzbK3w","outputId":"52fc3c43-862d-40bd-e849-7398dffd54bb","editable":false},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"![Imgur](https://imgur.com/3Fevxnl.png)\n\n<pre>\nTSNE for asm data with perplexity 50\n</pre>","metadata":{"id":"_QqjACR_bK3y","editable":false}},{"cell_type":"markdown","source":"<h1 style=\"font-size:250%; font-family:cursive; color:#ff6666;\"><b>25. Conclusion on EDA ( ONLY on .asm file features) <a id=\"25\"></a> </b></h1>\n\n#### [Back to the top](#0)\n","metadata":{"id":"cojyqmsEbK3z","editable":false}},{"cell_type":"markdown","source":"<p>\n<li>We have taken only 52 features from asm files (after reading through many blogs and research papers) </li>\n<li>The univariate analysis was done only on few important features.</li>\n<li>Take-aways\n<ul>\n<li>1. Class 3 can be easily separated because of the frequency of segments,opcodes and keywords being less </li>\n<li>2. Each feature has its unique importance in separating the Class labels.</li>\n</ul>\n</li>\n</p>","metadata":{"id":"u4zccERUbK30","editable":false}},{"cell_type":"markdown","source":"<h1 style=\"font-size:250%; font-family:cursive; color:#ff6666;\"><b>26. Train and test split ( ONLY on .asm file featues ) <a id=\"26\"></a> </b></h1>\n\n#### [Back to the top](#0)\n","metadata":{"id":"n5cvXzMmbK30","editable":false}},{"cell_type":"code","source":"asm_y = result_asm['Class']\nasm_x = result_asm.drop(['ID','Class','.BSS:','rtn','.CODE'], axis=1)","metadata":{"id":"gI2CEWvJbK32","editable":false},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"X_train_asm, X_test_asm, y_train_asm, y_test_asm = train_test_split(asm_x,asm_y ,stratify=asm_y,test_size=0.20)\nX_train_asm, X_cv_asm, y_train_asm, y_cv_asm = train_test_split(X_train_asm, y_train_asm,stratify=y_train_asm,test_size=0.20)","metadata":{"id":"Ec7f5HLkbK33","editable":false},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print( X_cv_asm.isnull().all())","metadata":{"id":"dUwoDBQLbK36","outputId":"d504ea98-a06d-4397-ba9a-a899fcf8f0a0","editable":false},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<h1 style=\"font-size:250%; font-family:cursive; color:#ff6666;\"><b>27. K-Nearest Neigbors ONLY on .asm file features <a id=\"27\"></a> </b></h1>\n\n#### [Back to the top](#0)\n","metadata":{"id":"_1XtQupvbK38","editable":false}},{"cell_type":"code","source":"# find more about KNeighborsClassifier() here http://scikit-learn.org/stable/modules/generated/sklearn.neighbors.KNeighborsClassifier.html\n# -------------------------\n# default parameter\n# KNeighborsClassifier(n_neighbors=5, weights=’uniform’, algorithm=’auto’, leaf_size=30, p=2, \n# metric=’minkowski’, metric_params=None, n_jobs=1, **kwargs)\n\n# methods of\n# fit(X, y) : Fit the model using X as training data and y as target values\n# predict(X):Predict the class labels for the provided data\n# predict_proba(X):Return probability estimates for the test data X.\n\n\n# find more about CalibratedClassifierCV here at http://scikit-learn.org/stable/modules/generated/sklearn.calibration.CalibratedClassifierCV.html\n# ----------------------------\n# default paramters\n# sklearn.calibration.CalibratedClassifierCV(base_estimator=None, method=’sigmoid’, cv=3)\n#\n# some of the methods of CalibratedClassifierCV()\n# fit(X, y[, sample_weight])\tFit the calibrated model\n# get_params([deep])\tGet parameters for this estimator.\n# predict(X)\tPredict the target of new samples.\n# predict_proba(X)\tPosterior probabilities of classification\n\n\nalpha = [x for x in range(1, 21,2)]\ncv_log_error_array=[]\nfor i in alpha:\n    k_cfl=KNeighborsClassifier(n_neighbors=i)\n    k_cfl.fit(X_train_asm,y_train_asm)\n    sig_clf = CalibratedClassifierCV(k_cfl, method=\"sigmoid\")\n    sig_clf.fit(X_train_asm, y_train_asm)\n    predict_y = sig_clf.predict_proba(X_cv_asm)\n    cv_log_error_array.append(log_loss(y_cv_asm, predict_y, labels=k_cfl.classes_, eps=1e-15))\n    \nfor i in range(len(cv_log_error_array)):\n    print ('log_loss for k = ',alpha[i],'is',cv_log_error_array[i])\n\nbest_alpha = np.argmin(cv_log_error_array)\n    \nfig, ax = plt.subplots()\nax.plot(alpha, cv_log_error_array,c='g')\nfor i, txt in enumerate(np.round(cv_log_error_array,3)):\n    ax.annotate((alpha[i],np.round(txt,3)), (alpha[i],cv_log_error_array[i]))\nplt.grid()\nplt.title(\"Cross Validation Error for each alpha\")\nplt.xlabel(\"Alpha i's\")\nplt.ylabel(\"Error measure\")\nplt.show()\n\nk_cfl=KNeighborsClassifier(n_neighbors=alpha[best_alpha])\nk_cfl.fit(X_train_asm,y_train_asm)\nsig_clf = CalibratedClassifierCV(k_cfl, method=\"sigmoid\")\nsig_clf.fit(X_train_asm, y_train_asm)\npred_y=sig_clf.predict(X_test_asm)\n\n\npredict_y = sig_clf.predict_proba(X_train_asm)\nprint ('log loss for train data',log_loss(y_train_asm, predict_y))\npredict_y = sig_clf.predict_proba(X_cv_asm)\nprint ('log loss for cv data',log_loss(y_cv_asm, predict_y))\npredict_y = sig_clf.predict_proba(X_test_asm)\nprint ('log loss for test data',log_loss(y_test_asm, predict_y))\nplot_confusion_matrix(y_test_asm,sig_clf.predict(X_test_asm))","metadata":{"id":"UJqB8USobK39","outputId":"6f3d2797-ee88-4f37-cb18-c7b6d1d98618","editable":false},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"![Imgur](https://imgur.com/xtCOdJi.png)\n\n![Imgur](https://imgur.com/vTUky0K.png)\n\n\n\n<h1 style=\"font-size:250%; font-family:cursive; color:#ff6666;\"><b>28. Logistic Regression ONLY on .asm file features <a id=\"28\"></a> </b></h1>\n\n\n#### [Back to the top](#0)\n\n","metadata":{"id":"yA1WorbYbK4C","editable":false}},{"cell_type":"code","source":"# read more about SGDClassifier() at http://scikit-learn.org/stable/modules/generated/sklearn.linear_model.SGDClassifier.html\n# ------------------------------\n# default parameters\n# SGDClassifier(loss=’hinge’, penalty=’l2’, alpha=0.0001, l1_ratio=0.15, fit_intercept=True, max_iter=None, tol=None, \n# shuffle=True, verbose=0, epsilon=0.1, n_jobs=1, random_state=None, learning_rate=’optimal’, eta0=0.0, power_t=0.5, \n# class_weight=None, warm_start=False, average=False, n_iter=None)\n\n# some of methods\n# fit(X, y[, coef_init, intercept_init, …])\tFit linear model with Stochastic Gradient Descent.\n# predict(X)\tPredict class labels for samples in X.\n\n\nalpha = [10 ** x for x in range(-5, 4)]\ncv_log_error_array=[]\nfor i in alpha:\n    logisticR=LogisticRegression(penalty='l2',C=i,class_weight='balanced')\n    logisticR.fit(X_train_asm,y_train_asm)\n    sig_clf = CalibratedClassifierCV(logisticR, method=\"sigmoid\")\n    sig_clf.fit(X_train_asm, y_train_asm)\n    predict_y = sig_clf.predict_proba(X_cv_asm)\n    cv_log_error_array.append(log_loss(y_cv_asm, predict_y, labels=logisticR.classes_, eps=1e-15))\n    \nfor i in range(len(cv_log_error_array)):\n    print ('log_loss for c = ',alpha[i],'is',cv_log_error_array[i])\n\nbest_alpha = np.argmin(cv_log_error_array)\n    \nfig, ax = plt.subplots()\nax.plot(alpha, cv_log_error_array,c='g')\nfor i, txt in enumerate(np.round(cv_log_error_array,3)):\n    ax.annotate((alpha[i],np.round(txt,3)), (alpha[i],cv_log_error_array[i]))\nplt.grid()\nplt.title(\"Cross Validation Error for each alpha\")\nplt.xlabel(\"Alpha i's\")\nplt.ylabel(\"Error measure\")\nplt.show()\n\nlogisticR=LogisticRegression(penalty='l2',C=alpha[best_alpha],class_weight='balanced')\nlogisticR.fit(X_train_asm,y_train_asm)\nsig_clf = CalibratedClassifierCV(logisticR, method=\"sigmoid\")\nsig_clf.fit(X_train_asm, y_train_asm)\n\npredict_y = sig_clf.predict_proba(X_train_asm)\nprint ('log loss for train data',(log_loss(y_train_asm, predict_y, labels=logisticR.classes_, eps=1e-15)))\npredict_y = sig_clf.predict_proba(X_cv_asm)\nprint ('log loss for cv data',(log_loss(y_cv_asm, predict_y, labels=logisticR.classes_, eps=1e-15)))\npredict_y = sig_clf.predict_proba(X_test_asm)\nprint ('log loss for test data',(log_loss(y_test_asm, predict_y, labels=logisticR.classes_, eps=1e-15)))\nplot_confusion_matrix(y_test_asm,sig_clf.predict(X_test_asm))","metadata":{"id":"0nkoXRHFbK4D","outputId":"8c15e841-1013-4746-8085-752e55c0ba4f","editable":false},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"![Imgur](https://imgur.com/8uIh7cZ.png)\n\n\n![Imgur](https://imgur.com/wV4w7Er.png)\n\n\n<h1 style=\"font-size:250%; font-family:cursive; color:#ff6666;\"><b>29. Random Forest Classifier ONLY on .asm file features <a id=\"29\"></a> </b></h1>\n\n#### [Back to the top](#0)\n","metadata":{"id":"g-vyA00GbK4P","editable":false}},{"cell_type":"code","source":"# --------------------------------\n# default parameters \n# sklearn.ensemble.RandomForestClassifier(n_estimators=10, criterion=’gini’, max_depth=None, min_samples_split=2, \n# min_samples_leaf=1, min_weight_fraction_leaf=0.0, max_features=’auto’, max_leaf_nodes=None, min_impurity_decrease=0.0, \n# min_impurity_split=None, bootstrap=True, oob_score=False, n_jobs=1, random_state=None, verbose=0, warm_start=False, \n# class_weight=None)\n\n# Some of methods of RandomForestClassifier()\n# fit(X, y, [sample_weight])\tFit the SVM model according to the given training data.\n# predict(X)\tPerform classification on samples in X.\n# predict_proba (X)\tPerform classification on samples in X.\n\n# some of attributes of  RandomForestClassifier()\n# feature_importances_ : array of shape = [n_features]\n# The feature importances (the higher, the more important the feature).\n\nalpha=[10,50,100,500,1000,2000,3000]\ncv_log_error_array=[]\nfor i in alpha:\n    r_cfl=RandomForestClassifier(n_estimators=i,random_state=42,n_jobs=-1)\n    r_cfl.fit(X_train_asm,y_train_asm)\n    sig_clf = CalibratedClassifierCV(r_cfl, method=\"sigmoid\")\n    sig_clf.fit(X_train_asm, y_train_asm)\n    predict_y = sig_clf.predict_proba(X_cv_asm)\n    cv_log_error_array.append(log_loss(y_cv_asm, predict_y, labels=r_cfl.classes_, eps=1e-15))\n\nfor i in range(len(cv_log_error_array)):\n    print ('log_loss for c = ',alpha[i],'is',cv_log_error_array[i])\n\n\nbest_alpha = np.argmin(cv_log_error_array)\n\nfig, ax = plt.subplots()\nax.plot(alpha, cv_log_error_array,c='g')\nfor i, txt in enumerate(np.round(cv_log_error_array,3)):\n    ax.annotate((alpha[i],np.round(txt,3)), (alpha[i],cv_log_error_array[i]))\nplt.grid()\nplt.title(\"Cross Validation Error for each alpha\")\nplt.xlabel(\"Alpha i's\")\nplt.ylabel(\"Error measure\")\nplt.show()\n\nr_cfl=RandomForestClassifier(n_estimators=alpha[best_alpha],random_state=42,n_jobs=-1)\nr_cfl.fit(X_train_asm,y_train_asm)\nsig_clf = CalibratedClassifierCV(r_cfl, method=\"sigmoid\")\nsig_clf.fit(X_train_asm, y_train_asm)\npredict_y = sig_clf.predict_proba(X_train_asm)\nprint ('log loss for train data',(log_loss(y_train_asm, predict_y, labels=sig_clf.classes_, eps=1e-15)))\npredict_y = sig_clf.predict_proba(X_cv_asm)\nprint ('log loss for cv data',(log_loss(y_cv_asm, predict_y, labels=sig_clf.classes_, eps=1e-15)))\npredict_y = sig_clf.predict_proba(X_test_asm)\nprint ('log loss for test data',(log_loss(y_test_asm, predict_y, labels=sig_clf.classes_, eps=1e-15)))\nplot_confusion_matrix(y_test_asm,sig_clf.predict(X_test_asm))","metadata":{"id":"utII0h75bK4P","outputId":"b84d2f62-fe98-44ec-a4f8-acdea6ccf617","editable":false},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"![Imgur](https://imgur.com/C431Dn7.png)\n\n![Imgur](https://imgur.com/RwZwWtJ.png)\n\n\n<h1 style=\"font-size:250%; font-family:cursive; color:#ff6666;\"><b>30. XgBoost Classifier ONLY on .asm file features <a id=\"30\"></a> </b></h1>\n\n#### [Back to the top](#0)\n","metadata":{"id":"A0Nb2KsmbK4R","editable":false}},{"cell_type":"code","source":"# Training a hyper-parameter tuned Xg-Boost regressor on our train data\n\n# find more about XGBClassifier function here http://xgboost.readthedocs.io/en/latest/python/python_api.html?#xgboost.XGBClassifier\n# -------------------------\n# default paramters\n# class xgboost.XGBClassifier(max_depth=3, learning_rate=0.1, n_estimators=100, silent=True, \n# objective='binary:logistic', booster='gbtree', n_jobs=1, nthread=None, gamma=0, min_child_weight=1, \n# max_delta_step=0, subsample=1, colsample_bytree=1, colsample_bylevel=1, reg_alpha=0, reg_lambda=1, \n# scale_pos_weight=1, base_score=0.5, random_state=0, seed=None, missing=None, **kwargs)\n\n# some of methods of RandomForestRegressor()\n# fit(X, y, sample_weight=None, eval_set=None, eval_metric=None, early_stopping_rounds=None, verbose=True, xgb_model=None)\n# get_params([deep])\tGet parameters for this estimator.\n# predict(data, output_margin=False, ntree_limit=0) : Predict with data. NOTE: This function is not thread safe.\n# get_score(importance_type='weight') -> get the feature importance\n\nalpha=[10,50,100,500,1000,2000,3000]\ncv_log_error_array=[]\nfor i in alpha:\n    x_cfl=XGBClassifier(n_estimators=i,nthread=-1)\n    x_cfl.fit(X_train_asm,y_train_asm)\n    sig_clf = CalibratedClassifierCV(x_cfl, method=\"sigmoid\")\n    sig_clf.fit(X_train_asm, y_train_asm)\n    predict_y = sig_clf.predict_proba(X_cv_asm)\n    cv_log_error_array.append(log_loss(y_cv_asm, predict_y, labels=x_cfl.classes_, eps=1e-15))\n\nfor i in range(len(cv_log_error_array)):\n    print ('log_loss for c = ',alpha[i],'is',cv_log_error_array[i])\n\n\nbest_alpha = np.argmin(cv_log_error_array)\n\nfig, ax = plt.subplots()\nax.plot(alpha, cv_log_error_array,c='g')\nfor i, txt in enumerate(np.round(cv_log_error_array,3)):\n    ax.annotate((alpha[i],np.round(txt,3)), (alpha[i],cv_log_error_array[i]))\nplt.grid()\nplt.title(\"Cross Validation Error for each alpha\")\nplt.xlabel(\"Alpha i's\")\nplt.ylabel(\"Error measure\")\nplt.show()\n\nx_cfl=XGBClassifier(n_estimators=alpha[best_alpha],nthread=-1)\nx_cfl.fit(X_train_asm,y_train_asm)\nsig_clf = CalibratedClassifierCV(x_cfl, method=\"sigmoid\")\nsig_clf.fit(X_train_asm, y_train_asm)\n    \npredict_y = sig_clf.predict_proba(X_train_asm)\n\nprint ('For values of best alpha = ', alpha[best_alpha], \"The train log loss is:\",log_loss(y_train_asm, predict_y))\npredict_y = sig_clf.predict_proba(X_cv_asm)\nprint('For values of best alpha = ', alpha[best_alpha], \"The cross validation log loss is:\",log_loss(y_cv_asm, predict_y))\npredict_y = sig_clf.predict_proba(X_test_asm)\nprint('For values of best alpha = ', alpha[best_alpha], \"The test log loss is:\",log_loss(y_test_asm, predict_y))\nplot_confusion_matrix(y_test_asm,sig_clf.predict(X_test_asm))","metadata":{"id":"v8i8RjYwbK4R","outputId":"1f04050c-aab0-49eb-ea1c-4a15bf6f7775","editable":false},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"![Imgur](https://imgur.com/JMb1GDQ.png)\n\n\n![Imgur](https://imgur.com/mp296Le.png)\n\n\n<h1 style=\"font-size:250%; font-family:cursive; color:#ff6666;\"><b>31. Xgboost Classifier with best hyperparameters ( ONLY on .asm file features ) <a id=\"31\"></a></b></h1>\n\n#### [Back to the top](#0)\n","metadata":{"id":"ShSLaaXHbK4V","editable":false}},{"cell_type":"code","source":"x_cfl=XGBClassifier()\n\nprams={\n    'learning_rate':[0.01,0.03,0.05,0.1,0.15,0.2],\n     'n_estimators':[100,200,500,1000,2000],\n     'max_depth':[3,5,10],\n    'colsample_bytree':[0.1,0.3,0.5,1],\n    'subsample':[0.1,0.3,0.5,1]\n}\nrandom_cfl=RandomizedSearchCV(x_cfl,param_distributions=prams,verbose=10,n_jobs=-1,)\nrandom_cfl.fit(X_train_asm,y_train_asm)","metadata":{"id":"1qa426fabK4W","outputId":"8ef2c632-0950-4f36-90bc-616e7c7f50a4","editable":false},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print (random_cfl.best_params_)","metadata":{"id":"rLLHdCXnbK4Z","outputId":"4719d48f-4d0b-4f86-9f44-22839bfee240","editable":false},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Training a hyper-parameter tuned Xg-Boost regressor on our train data\n\n# find more about XGBClassifier function here http://xgboost.readthedocs.io/en/latest/python/python_api.html?#xgboost.XGBClassifier\n# -------------------------\n# default paramters\n# class xgboost.XGBClassifier(max_depth=3, learning_rate=0.1, n_estimators=100, silent=True, \n# objective='binary:logistic', booster='gbtree', n_jobs=1, nthread=None, gamma=0, min_child_weight=1, \n# max_delta_step=0, subsample=1, colsample_bytree=1, colsample_bylevel=1, reg_alpha=0, reg_lambda=1, \n# scale_pos_weight=1, base_score=0.5, random_state=0, seed=None, missing=None, **kwargs)\n\n# some of methods of RandomForestRegressor()\n# fit(X, y, sample_weight=None, eval_set=None, eval_metric=None, early_stopping_rounds=None, verbose=True, xgb_model=None)\n# get_params([deep])\tGet parameters for this estimator.\n# predict(data, output_margin=False, ntree_limit=0) : Predict with data. NOTE: This function is not thread safe.\n# get_score(importance_type='weight') -> get the feature importance\n\nx_cfl=XGBClassifier(n_estimators=200,subsample=0.5,learning_rate=0.15,colsample_bytree=0.5,max_depth=3)\nx_cfl.fit(X_train_asm,y_train_asm)\nc_cfl=CalibratedClassifierCV(x_cfl,method='sigmoid')\nc_cfl.fit(X_train_asm,y_train_asm)\n\npredict_y = c_cfl.predict_proba(X_train_asm)\nprint ('train loss',log_loss(y_train_asm, predict_y))\npredict_y = c_cfl.predict_proba(X_cv_asm)\nprint ('cv loss',log_loss(y_cv_asm, predict_y))\npredict_y = c_cfl.predict_proba(X_test_asm)\nprint ('test loss',log_loss(y_test_asm, predict_y))","metadata":{"id":"Zmp3qSUobK4b","outputId":"21bc0acf-b634-4714-d243-1ff8005f98c1","editable":false},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"\n## Till now, all the above portion of the code was the basic experimentation with the very basic features of byte and asm file. Now comes the final part of this project\n\n---\n\n<h1 style=\"font-size:250%; font-family:cursive; color:#ff6666;\"><b> 32. FINAL FEATURIZATION STEPS FOR THE FINAL XGBOOST MODEL TRAINING <a id=\"32\"></a> </b></h1>\n\n#### [Back to the top](#0)\n","metadata":{"id":"YTXV2PUrV-tp","editable":false}},{"cell_type":"code","source":"# separating byte files and asm files \n# I am doing slight re-arrangement of the files for this FINAL run of Featurization and model training\nfrom google.colab import drive\ndrive.mount('/content/gdrive')\n\nroot_path = '/content/gdrive/MyDrive/Malware/Full_data/'\n# root_path = '../../LARGE_Datasets/'\n\ndestination_1 = root_path+'byteFiles'\ndestination_2 = root_path+'asmFiles'\n","metadata":{"id":"k98OdXkcECOJ","outputId":"e372fde3-b47e-46d7-c60b-06e3c55c5d6a","editable":false},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<h1 style=\"font-size:250%; font-family:cursive; color:#ff6666;\"><b> 33. Uni-Gram Byte Feature extraction from byte files - For FINAL Model Train <a id=\"33\"></a> </b></h1>\n\n#### [Back to the top](#0)\n\nThis cell's code is what we have already ran earlier in the experimentation part, including below here again for the sake of completeness","metadata":{"id":"H4K2jjgSqc2B","editable":false}},{"cell_type":"code","source":"%%time\n\n# This cell's code is what we have already ran earlier in the experimentation part\n# Including here again for the sake of completeness\n# removal of addres from byte files\n# contents of .byte files\n# ----------------\n#00401000 56 8D 44 24 08 50 8B F1 E8 1C 1B 00 00 C7 06 08 \n#-------------------\n#we remove the starting address 00401000\n\nfiles = os.listdir(root_path+'byteFiles/')\nfilenames=[]\narray=[]\nfor file in tqdm(files):\n    if(file.endswith(\"bytes\")):\n        file=file.split('.')[0]\n        text_file = open(root_path+'byteFiles/'+file+\".txt\", 'w+')\n        with open(root_path+'byteFiles/' + file + '.bytes', 'r') as fp:\n            lines=\"\"\n            for line in fp:\n                # rstrip()=> Return a copy of the string with trailing characters removed.\n                # Once we have removed trailing characters, invoke split() to return the list of string which are separated by \",\"\n                # split() specifies the separator to use when splitting the string. By default any whitespace is a separator\n                a=line.rstrip().split(\" \")[1:] # [1:] is equivalent to \"1 to end\" as we are removing 0-th element of address from byte files\n                b=' '.join(a)\n                b=b+\"\\n\" # Python doesn't automatically add line breaks, you need to do that manually\n                text_file.write(b)\n            fp.close()\n            os.remove(root_path+'byteFiles/'+file+\".bytes\")\n        text_file.close()\n\nfiles = os.listdir(root_path+'byteFiles/')\nfilenames2=[]\nfeature_matrix = np.zeros((len(files),257),dtype=int)\nk=0\n\n\n# program to convert into bag of words of bytefiles\n# this is custom-built bag of words this is unigram bag of words\n# This is a Custom Implementation of CountVectorizer as CountVectorizer will NOT suport working on such huge file system of 50GB\n# For this Uni-Gram feature creating and writing to a file named 'result.csv'\n\nbyte_feature_file=open(root_path + 'result.csv','w+')\n\nbyte_feature_file.write(\"ID,0,1,2,3,4,5,6,7,8,9,0a,0b,0c,0d,0e,0f,10,11,12,13,14,15,16,17,18,19,1a,1b,1c,1d,1e,1f,20,21,22,23,24,25,26,27,28,29,2a,2b,2c,2d,2e,2f,30,31,32,33,34,35,36,37,38,39,3a,3b,3c,3d,3e,3f,40,41,42,43,44,45,46,47,48,49,4a,4b,4c,4d,4e,4f,50,51,52,53,54,55,56,57,58,59,5a,5b,5c,5d,5e,5f,60,61,62,63,64,65,66,67,68,69,6a,6b,6c,6d,6e,6f,70,71,72,73,74,75,76,77,78,79,7a,7b,7c,7d,7e,7f,80,81,82,83,84,85,86,87,88,89,8a,8b,8c,8d,8e,8f,90,91,92,93,94,95,96,97,98,99,9a,9b,9c,9d,9e,9f,a0,a1,a2,a3,a4,a5,a6,a7,a8,a9,aa,ab,ac,ad,ae,af,b0,b1,b2,b3,b4,b5,b6,b7,b8,b9,ba,bb,bc,bd,be,bf,c0,c1,c2,c3,c4,c5,c6,c7,c8,c9,ca,cb,cc,cd,ce,cf,d0,d1,d2,d3,d4,d5,d6,d7,d8,d9,da,db,dc,dd,de,df,e0,e1,e2,e3,e4,e5,e6,e7,e8,e9,ea,eb,ec,ed,ee,ef,f0,f1,f2,f3,f4,f5,f6,f7,f8,f9,fa,fb,fc,fd,fe,ff,??\")\nbyte_feature_file.write(\"\\n\")\n\nfor file in tqdm(files):\n    filenames2.append(file)\n    byte_feature_file.write(file+\",\")\n    if(file.endswith(\"txt\")):\n        with open(root_path+'byteFiles/'+file,\"r\") as byte_flie:\n            for lines in byte_flie:\n                line=lines.rstrip().split(\" \")\n                for hex_code in line:\n                    if hex_code=='??':\n                        feature_matrix[k][256]+=1\n                    else:\n                        feature_matrix[k][int(hex_code,16)]+=1\n        byte_flie.close()\n    for i, row in enumerate(feature_matrix[k]):\n        if i!=len(feature_matrix[k])-1:\n            byte_feature_file.write(str(row)+\",\")\n        else:\n            byte_feature_file.write(str(row))\n    byte_feature_file.write(\"\\n\")\n    \n    k += 1\n\nbyte_feature_file.close()","metadata":{"id":"ZqXFjFxdqTE8","outputId":"4e1368b5-03b5-4011-8b84-f5f38b352d21","editable":false},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\n\nuni_gram_byte_features = pd.read_csv(root_path + \"result.csv\")\n\nuni_gram_byte_features['ID']  = uni_gram_byte_features['ID'].str.split('.').str[0]\n\nprint('Unigram byte_featues shape ', uni_gram_byte_features.shape)\n\nuni_gram_byte_features.head(2)","metadata":{"id":"RqcXeIvE5Dlj","outputId":"6f78bda0-f361-4470-f84b-9eb1687fda19","editable":false},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<h1 style=\"font-size:250%; font-family:cursive; color:#ff6666;\"><b> 34. File sizes of Byte files - Feature Extraction -For FINAL Model Train <a id=\"34\"></a></b></h1>\n\n#### [Back to the top](#0)\n\n","metadata":{"id":"wm26kvmFBSy3","editable":false}},{"cell_type":"code","source":"%%time\n\n# This cell's code is what we have already ran earlier in the experimentation part\n# Including here again for the sake of completeness\nY=pd.read_csv(root_path + \"trainLabels.csv\")\n\nfiles=os.listdir(root_path + 'byteFiles')\n\nfilenames=Y['Id'].tolist()\n\nclass_y=Y['Class'].tolist()\n\nclass_bytes=[]\n\nsizebytes=[]\n\nfnames=[]\n\nfor file in tqdm(files):\n    # print(os.stat('byteFiles/0A32eTdBKayjCWhZqDOQ.txt'))\n    # os.stat_result(st_mode=33206, st_ino=1125899906874507, st_dev=3561571700, st_nlink=1, st_uid=0, st_gid=0, \n    # st_size=3680109, st_atime=1519638522, st_mtime=1519638522, st_ctime=1519638522)\n    # read more about os.stat: here https://www.tutorialspoint.com/python/os_stat.htm\n    statinfo=os.stat(root_path+'byteFiles/'+file)\n    # split the file name at '.' and take the first part of it i.e the file name\n    file=file.split('.')[0]\n    if any(file == filename for filename in filenames):\n        i=filenames.index(file)\n        class_bytes.append(class_y[i])\n        # converting into Mb's\n        sizebytes.append(statinfo.st_size/(1024.0*1024.0))\n        fnames.append(file)\n\nbyte_feature_size=pd.DataFrame({'ID':fnames, 'size':sizebytes,'Class':class_bytes})\n\nprint (byte_feature_size.head())","metadata":{"id":"PBN_nBv0BSy3","outputId":"7c43d917-fd0b-4d00-af62-d4d44dd8eb24","editable":false},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<h1 style=\"font-size:250%; font-family:cursive; color:#ff6666;\"><b> 35. Creating some important Files and Folders, which I shall use later for saving Featuarized versions of .csv files <a id=\"35\"></a></b></h1>\n\n#### [Back to the top](#0)\n\n","metadata":{"id":"EsuuWBkPV-tv","editable":false}},{"cell_type":"code","source":"if not os.path.isdir(root_path + \"featurization\"):\n    os.makedirs(root_path + \"featurization\")\n\n\nif not os.path.isdir(root_path + \"featurization/featurization_final\"):\n    os.mkdir(root_path + \"featurization/featurization_final\")\n\n\n# Creating and writing to a file named \"class_labels.pkl\" to get class class_labels and ID from byte unigrams dataframe and save it for later use\n\nclass_labels=byte_feature_size[\"Class\"]\n\nwith open(root_path+'featurization/class_labels.pkl', 'wb') as file:\n    pkl.dump(class_labels, file)\n\n'''\nhttps://www.datacamp.com/community/tutorials/pickle-python-tutorial\n\nTo open the file for writing, simply use the open() function. The first argument should be the name of your file. The second argument is 'wb'. The w means that you'll be writing to the file, and b refers to binary mode. This means that the data will be written in the form of byte objects.\n'''\n\n# Load the class class_labels for training with random forest feature selector\n\nwith open(root_path+'featurization/class_labels.pkl', 'rb') as file:\n    class_labels=pkl.load(file)","metadata":{"id":"tKo74LEQV-tv","editable":false},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<h1 style=\"font-size:250%; font-family:cursive; color:#ff6666;\"><b> 36. Merging Unigram of Byte Files + Size of Byte Files to create uni_gram_byte_features__with_size <a id=\"36\"></a></b></h1>\n\n#### [Back to the top](#0)\n\n","metadata":{"id":"d0NG3x-eV-tv","editable":false}},{"cell_type":"code","source":"%%time\n\nuni_gram_byte_features__with_size = uni_gram_byte_features.merge(byte_feature_size, on=\"ID\")\n\nuni_gram_byte_features__with_size.to_csv(root_path + \"featurization/uni_gram_byte_features__with_size.csv\", index=False)\n\nuni_gram_byte_features__with_size = normalize(uni_gram_byte_features__with_size)","metadata":{"id":"1ObEVXSYV-tw","editable":false},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\n\nfrom sklearn.feature_extraction.text import CountVectorizer\n\nbigram_tokens=\"00,01,02,03,04,05,06,07,08,09,0a,0b,0c,0d,0e,0f,10,11,12,13,14,15,16,17,18,19,1a,1b,1c,1d,1e,1f,20,21,22,23,24,25,26,27,28,29,2a,\\\n2b,2c,2d,2e,2f,30,31,32,33,34,35,36,37,38,39,3a,3b,3c,3d,3e,3f,40,41,42,43,44,45,46,47,48,49,4a,4b,4c,4d,4e,4f,50,51,52,53,54,55,56,57,58,\\\n59,5a,5b,5c,5d,5e,5f,60,61,62,63,64,65,66,67,68,69,6a,6b,6c,6d,6e,6f,70,71,72,73,74,75,76,77,78,79,7a,7b,7c,7d,7e,7f,80,81,82,83,84,85,86,\\\n87,88,89,8a,8b,8c,8d,8e,8f,90,91,92,93,94,95,96,97,98,99,9a,9b,9c,9d,9e,9f,a0,a1,a2,a3,a4,a5,a6,a7,a8,a9,aa,ab,ac,ad,ae,af,b0,b1,b2,b3,b4,b5,\\\nb6,b7,b8,b9,ba,bb,bc,bd,be,bf,c0,c1,c2,c3,c4,c5,c6,c7,c8,c9,ca,cb,cc,cd,ce,cf,d0,d1,d2,d3,d4,d5,d6,d7,d8,d9,da,db,dc,dd,de,df,e0,e1,e2,e3,e4,\\\ne5,e6,e7,e8,e9,ea,eb,ec,ed,ee,ef,f0,f1,f2,f3,f4,f5,f6,f7,f8,f9,fa,fb,fc,fd,fe,ff,??\"\n\nbigram_tokens=bigram_tokens.split(\",\")\n\n# Between 00 and FF there are 256 unique values, so if we take each pair of Hexadecimal Values as one word, \n# we are dealing with 256 unique values. \n# Hence below Function will extract all the possible combinations of bigrams_counts\ndef calculate_bigram(bigram_tokens):\n    sentence=\"\"\n    vocabulary_list_for_byte_bigrams=[]\n    for i in tqdm(range(len(bigram_tokens))):\n        for j in range(len(bigram_tokens)):\n            bigram=bigram_tokens[i]+\" \"+bigram_tokens[j]\n            sentence=sentence+bigram+\",\"\n            vocabulary_list_for_byte_bigrams.append(bigram)\n    return vocabulary_list_for_byte_bigrams\n\nvocabulary_list_for_byte_bigrams = calculate_bigram(bigram_tokens) ","metadata":{"id":"j-CNv41CV-tw","outputId":"cd6be432-b994-46e4-cf4b-6b27f2ee7ee8","editable":false},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<h1 style=\"font-size:250%; font-family:cursive; color:#ff6666;\"><b> 37. Bi-Gram Byte Feature extraction from byte files <a id=\"37\"></a> </b></h1>\n\n\n#### [Back to the top](#0)\n\n","metadata":{"id":"fygGcKCcqTjK","editable":false}},{"cell_type":"code","source":"%%time\n\nimport scipy\nvectorizer = CountVectorizer(tokenizer=lambda x: x.split(),lowercase=False, ngram_range=(2,2),vocabulary=vocabulary_list_for_byte_bigrams) \n\n# For Explanations on \"tokenizer=lambda x: x.split()\"\n# Refer - https://stackoverflow.com/a/37884104/1902852\n# Without this \"??\" was not getting vectorized properly\n\nfile_list_byte_files=os.listdir(root_path + 'byteFiles')\n\nfeatures=[\"ID\"]+vectorizer.get_feature_names()\n\nbyte_file_bigram_df=pd.DataFrame(columns=features)\n\n# Creating \"featurization/byte_files_bigram_df.csv\" and writng to it the full bi-gram data frame\nwith open(root_path + \"featurization/byte_files_bigram_df.csv\", mode='w') as byte_file_bigram_df:\n    byte_file_bigram_df.write(','.join(map(str, features)))\n    byte_file_bigram_df.write('\\n')\n    for _, file in tqdm(enumerate(file_list_byte_files)):\n        file_id=file.split(\".\")[0] #ID of each file\n        file = open(root_path + 'byteFiles/' + file)\n        corpus_byte_codes=[file.read().replace('\\n', ' ').lower()] # corpus_byte_codes holds all the byte codes for a given file\n        bigrams_counts = vectorizer.transform(corpus_byte_codes) # Returning a sparse vector containing all the bigram counts from the corpus_byte_codes\n        \n        # Update each row of our dataframe with the bigram counts of the respective file\n        row = scipy.sparse.csr_matrix(bigrams_counts).toarray() \n        \n        # Write a single row in the CSV file\n        byte_file_bigram_df.write(','.join(map(str, [file_id]+list(row[0]))))\n        \n        byte_file_bigram_df.write('\\n')\n        \n        file.close()\n","metadata":{"id":"4qoHkpayDon2","outputId":"31cef728-3222-43ec-ea95-5c5818115bb9","editable":false},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<h1 style=\"font-size:250%; font-family:cursive; color:#ff6666;\"><b> 38. Extracting the 2000 Most Important Features from Byte bigrams using SelectKBest with Chi-Square Test <a id=\"38\"></a> </b></h1>\n\n\n#### [Back to the top](#0)\n","metadata":{"id":"GmKq_ZHvDon3","editable":false}},{"cell_type":"code","source":"%%time\n\n# Load the byte_files_bigram_df.csv file which is NOT normalized dataset for the byte file's bigrams\n# that we created in the previous cell\nX_byte_bigram_all_df = pd.read_csv(root_path + \"featurization/byte_files_bigram_df.csv\")\n\nX_byte_bigram_all_df.head(2)","metadata":{"id":"BR7R5czhDon3","outputId":"56ac155d-a7ce-424a-bd06-456089b195b3","editable":false},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\n\n\nfrom sklearn.feature_selection import SelectKBest, chi2, f_regression\n\nselect_kbest_object = SelectKBest(score_func=chi2, k=2000)\n# SelectKBest scores the features using a function, which is chi2 here\n# Then \"removes all but the k highest scoring features\"\n\n# Need to remove \"ID\" column, else will get below error \n# \"SelectKBest fit: ValueError: could not convert string to float\"\n\nmost_imp_features_byte_bigram = select_kbest_object.fit(X_byte_bigram_all_df.drop(\"ID\", axis=1), class_labels)\n\n# most_imp_features_byte_bigram.scores_ => gives an array of form \n# array([9.79531407e+05, 4.26642398e+04, 1.78812060e+04, ..., 4.33426736e+07])\n# So now creating a df from this array\nmost_imp_byte_bigram_feature_score_df = pd.DataFrame(most_imp_features_byte_bigram.scores_)\n\n# Creating a df from all the column names from the original full X_byte_bigram_all_df df\nmost_imp_byte_bigram_columns_df = pd.DataFrame(X_byte_bigram_all_df.columns)","metadata":{"id":"9M5dOXS0Don4","outputId":"0da9aa4d-e773-4e74-8b97-e840821f3cdc","editable":false},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Concat the feature scores along with the feature names in a byte_bigram_df_important_feature_score, \n# From this we will get all feature names later, to be matched against X_byte_bigram_all_df - to extract ONLY the best features from the bigrams df data\nbyte_bigram_df_important_feature_score = pd.concat([most_imp_byte_bigram_columns_df, most_imp_byte_bigram_feature_score_df],axis=1)\n\nbyte_bigram_df_important_feature_score.columns = [\"Byte Bigram Top 2000 Feature Names\",\"Byte Bigram Top 2000 Feature Score\"]\n\n# Find the top 2000 features along with their scores\n\n# byte_bigram_df_important_feature_score=byte_bigram_df_important_feature_score.nlargest(1000, \"Byte Bigram Top 2000 Feature Score\")\n\n# Return the first 2000 rows with the largest values in the specified column ( \"Byte Bigram Top 2000 Feature Score\" )\n# in descending order. The columns that are not specified are returned as well, but not used for ordering.\nbyte_bigram_df_important_feature_score = byte_bigram_df_important_feature_score.nlargest(10, \"Byte Bigram Top 2000 Feature Score\")\n\nbyte_bigram_df_important_feature_score.head(2)","metadata":{"id":"CaKdgRmyG78Y","outputId":"41ede10a-470d-4fa5-91c1-edc09cda7b2b","editable":false},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Getting the list of first 2000 feature names\ntop_2000_most_imp_byte_bigram_feature_names = list(byte_bigram_df_important_feature_score[\"Byte Bigram Top 2000 Feature Names\"])\n\n# top_2000_byte_bigram_features = dd.concat([X_byte_bigram_all_df[\"ID\"], X_byte_bigram_all_df[top_2000_most_imp_byte_bigram]], axis=1)\ntop_2000_byte_bigram_features = pd.concat([X_byte_bigram_all_df[\"ID\"], X_byte_bigram_all_df[top_2000_most_imp_byte_bigram_feature_names]], axis=1)\n\ntop_2000_byte_bigram_features.to_csv(root_path + \"featurization/featurization_final/top_2000_imp_byte_bigram_df.csv\",index=None)\n\nprint(top_2000_byte_bigram_features.shape)\ntop_2000_byte_bigram_features.head(2)","metadata":{"id":"rNphU52XV-ty","outputId":"e2f07f50-db97-4a99-e179-2930aadbc996","editable":false},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<h1 style=\"font-size:250%; font-family:cursive; color:#ff6666;\"><b> 39. ASM Unigram - Top 52 Unigram Features from ASM Files - Final Model Training <a id=\"39\"></a> </b></h1>\n\n\n#### [Back to the top](#0)\n\nThere are 10868 files of asm files which make up about 150 GB\n\nThe asm files contains :\n\n1. Address\n2. Segments\n3. Opcodes\n4. Registers\n5. function calls\n6. APIs\n\nEarlier we already extracted all the features with the help of parallel processing. Here we extracted 52 features from all the asm files which are important.\n\n#### The asmoutputfile.csv was generated after extracting the unigram features from the .asm files which was ~150GB of data. Took around 48 Hours to process, and we can directly use this file here.\n","metadata":{"id":"rXCQQr2AV-tz","editable":false}},{"cell_type":"code","source":"# First read the file that was generated above code\n# Meaning, the code that ran for around 48 hours as mentioned above.\ndfasm=pd.read_csv(root_path + \"/asmoutputfile.csv\")\n\nY.columns = ['ID', 'Class'] \n# Note, Y is all the Train Labels of 0 to 9 which has been defined earlier as below\n# Y = pd.read_csv(root_path + \"trainLabels.csv\")\n\nunigram_asm = pd.merge(dfasm, Y, on='ID', how='left')\n\nunigram_asm = normalize(unigram_asm)\n\nunigram_asm.head()","metadata":{"id":"f4YrXGr4V-tz","outputId":"fa8a7eec-2c08-433f-b84a-05f292d81c28","editable":false},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<h1 style=\"font-size:250%; font-family:cursive; color:#ff6666;\"><b> 40. File Size of ASM Files - Feature Extraction - Final Model Training <a id=\"40\"></a> </b></h1>\n\n\n#### [Back to the top](#0)\n\nThis cell's code is what we have already ran earlier in the experimentation part, including below here again for the sake of completeness\n","metadata":{"id":"Tlb6idyzV-tz","editable":false}},{"cell_type":"code","source":"# file sizes of byte files\n# This code is very much similar to what has been used to extract sizes of \n# byte files earlier.\nfiles=os.listdir(root_path + 'asmFiles')\n\nfilenames=Y['ID'].tolist()\n\nclass_y=Y['Class'].tolist()\n\nclass_bytes=[]\n\nsizebytes=[]\n\nfnames=[]\n\nfor file in tqdm(files):\n    # print(os.stat('byteFiles/0A32eTdBKayjCWhZqDOQ.txt'))\n    # os.stat_result(st_mode=33206, st_ino=1125899906874507, st_dev=3561571700, st_nlink=1, st_uid=0, st_gid=0, \n    # st_size=3680109, st_atime=1519638522, st_mtime=1519638522, st_ctime=1519638522)\n    # read more about os.stat: here https://www.tutorialspoint.com/python/os_stat.htm\n    statinfo=os.stat(root_path + 'asmFiles/'+file)\n    # split the file name at '.' and take the first part of it i.e the file name\n    file=file.split('.')[0]\n    if any(file == filename for filename in filenames):\n        i=filenames.index(file)\n        class_bytes.append(class_y[i])\n        # converting into Mb's\n        sizebytes.append(statinfo.st_size/(1024.0*1024.0))\n        fnames.append(file)\n\nasm_file_size=pd.DataFrame({'ID':fnames,'size':sizebytes,'Class':class_bytes})\n\n# asm_file_size.to_csv(root_path + \"featurization/asm_file_size.csv\", index=False)\n\nasm_file_size.head()","metadata":{"id":"1bTlFISWV-t0","outputId":"d1247d57-faa9-49b5-a4bd-218ce4a019bf","editable":false},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<h1 style=\"font-size:250%; font-family:cursive; color:#ff6666;\"><b> 41. Merging ASM Unigram + ASM File Size <a id=\"41\"></a></b></h1>\n\n\n#### [Back to the top](#0)\n\n","metadata":{"id":"SN6Esog-V-t0","editable":false}},{"cell_type":"code","source":"unigram_asm_feature__with_size=pd.merge(asm_file_size, unigram_asm.drop(columns=[\"Class\"]),on='ID', how='left')\n\nunigram_asm_feature__with_size.to_csv(root_path + \"featurization/unigram_asm_feature__with_size\")\n\nunigram_asm_feature__with_size.head()","metadata":{"id":"nUMfi-5bV-t0","outputId":"3185383d-f13a-4abd-e670-e5b8b3b4cbcb","editable":false},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<h1 style=\"font-size:250%; font-family:cursive; color:#ff6666;\"><b> 42. ASM Files - Convert the ASM files to images. <a id=\"42\"></a> </b></h1>\n\n#### [Back to the top](#0)\n\n","metadata":{"id":"fkvxJIjZV-t0","editable":false}},{"cell_type":"code","source":"%%time\n\nimport numpy as np\nimport os\nimport codecs\nimport imageio\nimport array\nfrom datetime import datetime as dt\n\nif not os.path.isdir(root_path + \"image_file_asm\"):\n    os.mkdir(root_path + \"image_file_asm\")\n\nasmfile_list=os.listdir(root_path + \"asmFiles/\")\n\n# Function to extract images from ASM files and save them to a specified folder (the second arg to the func)\ndef extract_images_from_text(arr_of_filenames, folder_to_save_generated_images):  \n    for file_name in tqdm(arr_of_filenames):\n        \n        if(file_name.endswith(\"asm\")):\n            this_file = codecs.open(root_path + \"asmFiles/\" + file_name, 'rb')\n            size_of_current_asm_file = os.path.getsize(root_path + \"asmFiles/\"+file_name)        \n        \n        width_of_file = int(size_of_current_asm_file**0.5)\n        \n        remainder = size_of_current_asm_file % width_of_file\n        \n        # To create array of single bytes, passing type code 'B'\n        # \"B\" is for unsigned characters\n        array_of_image = array.array('B')\n        \n        array_of_image.fromfile(this_file, size_of_current_asm_file-remainder)\n        \n        this_file.close()\n        \n        arr_of_generated_image = np.reshape(array_of_image[:width_of_file * width_of_file], (width_of_file, width_of_file))\n        \n        arr_of_generated_image = np.uint8(arr_of_generated_image)\n        \n        imageio.imwrite(folder_to_save_generated_images+'/' + file_name.split(\".\")[0] + '.png', arr_of_generated_image)\n        \n        \n# Now invoke the above function\n\ndirectory_to_save_generated_image = root_path + 'image_file_asm'\n\nextract_images_from_text(asmfile_list, directory_to_save_generated_image)","metadata":{"id":"lYZNEnBHV-t0","outputId":"e9d39c31-72b1-4626-c65c-191243b0e6ee","editable":false},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<h1 style=\"font-size:250%; font-family:cursive; color:#ff6666;\"><b>43. Extract the first 800 pixel data from ASM File Images <a id=\"43\"></a> </b></h1>\n\n\n#### [Back to the top](#0)\n\n\n\n### Load each ASM image > Convert it into a numpy array > Take the first 800 pixels from each image.","metadata":{"id":"cSJeoQndV-t1","editable":false}},{"cell_type":"code","source":"file_list_asm_files=os.listdir(root_path + 'image_file_asm/')\n\nwith open(root_path + \"featurization/top_800_image_asm_df.csv\", mode='w') as top_800_image_asm_df: #file_list_asm_files = 10868, top_800_image_asm_df=800\n    # top_800_image_asm_df.write(','.join(map(str, [\"ID\"]+[\"pixel_asm{}\".format(i) for i in range(800)])))\n    top_800_image_asm_df.write(','.join(map(str, [\"ID\"]+[\"pixel_asm{}\".format(i) for i in range(10)])))\n    top_800_image_asm_df.write('\\n')\n    \n    for image in tqdm(file_list_asm_files):\n        file_id_asm_files=image.split(\".\")[0]\n        \n         # Create a 2 Matrix to contain the image matrix in 2D format\n        asm_image_array=imageio.imread(root_path + \"image_file_asm/\"+image)\n        \n        # Extracting from flattened array the first 800 pixels \n        # asm_image_array=asm_image_array.flatten()[:800]\n        asm_image_array=asm_image_array.flatten()[:10]\n        top_800_image_asm_df.write(','.join(map(str, [file_id_asm_files]+list(asm_image_array))))\n        top_800_image_asm_df.write('\\n')","metadata":{"id":"WgD93YllV-t1","outputId":"1f1263a8-4b60-42fb-e32e-a0a673ae5d55","editable":false},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\n\ntop_800_image_asm_df=pd.read_csv(root_path + \"featurization/top_800_image_asm_df.csv\")\ntop_800_image_asm_df.head()","metadata":{"id":"vq58IrspV-t1","outputId":"a04ff0da-4f81-447b-8512-b7ab53b68ed4","editable":false},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<h1 style=\"font-size:250%; font-family:cursive; color:#ff6666;\"><b> 44. Extracting Opcodes Bigrams from ASM Files <a id=\"44\"></a></b></h1>\n\n#### [Back to the top](#0)\n\n\n\nWe know that the asm files contain assembly language code which comprises keywords, opcodes, registers, APIs.","metadata":{"id":"F285OokXV-t2","editable":false}},{"cell_type":"code","source":"%%time\n\nopcodes_for_bigram = ['jmp', 'mov', 'retf', 'push', 'pop', 'xor', 'retn', 'nop', 'sub', 'inc', 'dec', 'add','imul', 'xchg', 'or', 'shr', 'cmp', 'call', 'shl', 'ror', 'rol', 'jnb','jz','rtn','lea','movzx']\n\n# Converting list to dictionary for faster runtime\ndict_asm_opcodes = dict(zip(opcodes_for_bigram, [1 for i in range(len(opcodes_for_bigram))]))\n\nif not os.path.isdir(root_path + \"opcodes_asm_files\"):\n    os.mkdir(root_path + 'opcodes_asm_files')\n\n'''\n\nNoting first that the asm files contains :\n\n1. Address\n2. Segments\n3. Opcodes\n4. Registers\n5. function calls\n6. APIs\n\nCalculating opcode sequences for each asm file and save in form of a text file, so that we can process the ASM files as text files\n\nIn that text file, each row corresponds to respective file. \n\nNoting, in asm files the opcodes_for_bigram were not placed side by side, instead there are few words between two opcodes. i.e. The Opcodes occurs with an interval.\n\nSo during extraction of opcodes_for_bigram we need to preserve the sequence information. \n\ne.g. which opcode prcede another opcode or which opcode is followed is followed by another opcode.\n\nBased on this, a bigram data-matrix of vectors is to be derived containing the bigram sequence info on each file.\n\n'''\n\ndef calculate_sequence_of_opcodes():\n    asm_file_names=os.listdir(root_path + 'asmFiles')\n    for this_asm_file in tqdm(asm_file_names):\n        each_asm_opcode_file = open(root_path + \"opcodes_asm_files/{}_opcode_asm_bi_grams.txt\".format(this_asm_file.split('.')[0]), \"w+\")\n        sequence_of_opcodes = \"\"\n        with codecs.open(root_path + 'asmFiles/' + this_asm_file, encoding='cp1252', errors ='replace') as asm_file:\n            for lines in asm_file:\n                \n                line = lines.rstrip().split()            \n                \n                for word in line:\n                    if dict_asm_opcodes.get(word)==1:\n                        sequence_of_opcodes += word + ' '\n        each_asm_opcode_file.write(sequence_of_opcodes + \"\\n\")\n        each_asm_opcode_file.close()\n    \ncalculate_sequence_of_opcodes()\n\nopcodes_asm__bigram_vocabulary = calculate_bigram(opcodes_for_bigram)","metadata":{"id":"dXeWEHsPV-t2","outputId":"be59d895-b698-4d22-eef9-d56ab5e41a04","editable":false},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<h1 style=\"font-size:250%; font-family:cursive; color:#ff6666;\"><b> 45. Calcualte opcodes bigram with above defined function and make them a feature and then save the data matrix of feature as a .csv file <a id=\"45\"></a></b></h1>\n\n\n#### [Back to the top](#0)\n","metadata":{"id":"TivI8uo3V-t2","editable":false}},{"cell_type":"code","source":"vectorizer_opcode = CountVectorizer(\n    tokenizer=lambda x: x.split(),\n    lowercase=False,\n    ngram_range=(2, 2),\n    vocabulary=opcodes_asm__bigram_vocabulary,\n)  # Noting, without \"tokenizer=lambda x: x.split()\", \"??\" would not get vectorized correctly\n\nfile_list_opcode = os.listdir(root_path + \"opcodes_asm_files\")\n\nopcode_features = [\"ID\"] + vectorizer_opcode.get_feature_names()\n\nopcodes_asm_bigram_df = pd.DataFrame(columns=opcode_features)\n\nwith open(\n    root_path + \"featurization/opcodes_asm_bigram_df.csv\", mode=\"w\"\n) as opcodes_asm_bigram_df:\n\n    opcodes_asm_bigram_df.write(\",\".join(map(str, opcode_features)))\n\n    opcodes_asm_bigram_df.write(\"\\n\")\n\n    for _, this_asm_file in tqdm(enumerate(file_list_opcode)):\n\n        this_file_id = this_asm_file.split(\"_\")[0]  # ID of each this_asm_file\n\n        this_asm_file = open(root_path + \"opcodes_asm_files/\" + this_asm_file)\n\n        corpus_opcodes_from_this_asm_file = [\n            this_asm_file.read().replace(\"\\n\", \" \").lower()\n        ]  # Variable to hold all opcodes for a given this_asm_file\n\n        bigrams_opcodes_asm = vectorizer_opcode.transform(\n            corpus_opcodes_from_this_asm_file\n        )  # Returning a sparse vector holding all bigram counts from corpus_opcodes_from_this_asm_file\n\n        # Update each row of the dataframe with the bigram counts of the respective this_asm_file\n        # And return a dense ndarray representation of this matrix. Because,\n        # CountVectorizer produces a sparse representation of the counts using scipy.sparse.csr_matrix\n        row = scipy.sparse.csr_matrix(bigrams_opcodes_asm).toarray()\n\n        opcodes_asm_bigram_df.write(\n            \",\".join(map(str, [this_file_id] + list(row[0])))\n        )  # Write a single row in the CSV this_asm_file\n\n        opcodes_asm_bigram_df.write(\"\\n\")\n\n        this_asm_file.close()\n\n\nopcodes_asm_bigram_df = pd.read_csv(\n    root_path + \"featurization/opcodes_asm_bigram_df.csv\"\n)\n\nopcodes_asm_bigram_df.head()","metadata":{"id":"M8il_TFoV-t2","outputId":"07118624-3e4d-4848-c1bc-28f52d8f76ef","editable":false},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<h1 style=\"font-size:250%; font-family:cursive; color:#ff6666;\"><b> 46. ASM File - Top Important 500 features from Opcodes Bigrams <a id=\"46\"></a> </b></h1>\n\n\n#### [Back to the top](#0)\n","metadata":{"id":"0GGVQLgiV-t3","editable":false}},{"cell_type":"code","source":"X_opcode_asm_bigram = opcodes_asm_bigram_df\ny = class_labels\n# X_opcode_asm_bigram.head()\n\n#Get the best 500 features using SelectKBest. \n\n\nkbest_object = SelectKBest(score_func=chi2, k=500)\n\ntop_features=kbest_object.fit(X_opcode_asm_bigram.drop(\"ID\", axis=1), y)\n\n# Save a dataframe with the feature scores along with the feature names.\n# And we will get the best fetures from this dataframe use to \ntop_features_scores=pd.DataFrame(top_features.scores_)\n\n# Now to get the original features names i.e. the names of all the columns we will need\n# `X_opcode_asm_bigram.columns`\nX_opcode_columns=pd.DataFrame(X_opcode_asm_bigram.columns)\n\n# Now concat all  original features names as a column with another column\n# which is \"top_features_scores\"\ntop_asm_opcode_bigram_df=pd.concat([X_opcode_columns, top_features_scores],axis=1)\n\n# Give 2 Names for these 2 columns of data for this newly creaetd dataframe\ntop_asm_opcode_bigram_df.columns=[\"ASM_Opcode_Bigram_Top_Feature_Name\",\"ASM_Opcode_Bigram_Top_Feature_Score\"]\n\n# Extract the largest 500 from this dataframw based on the values of \"top_features_scores\"\ntop_asm_opcode_bigram_df=top_asm_opcode_bigram_df.nlargest(500,\"ASM_Opcode_Bigram_Top_Feature_Score\")\n\ntop_asm_opcode_bigram_df.head()","metadata":{"id":"rrJZbMKhV-t3","outputId":"49f44a12-1ff5-4ff3-e47a-0c00404086c4","editable":false},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"top_500_asm_bigram_features=list(top_asm_opcode_bigram_df[\"ASM_Opcode_Bigram_Top_Feature_Name\"])\n\ntop_500_asm_bigram_df=pd.concat([X_opcode_asm_bigram[\"ID\"], X_opcode_asm_bigram[top_500_asm_bigram_features]], axis=1)\n\n# The \"ID\" column was being duplicated, hence need to remove that, and also the possibility of any other duplicated column\ntop_500_asm_bigram_df = top_500_asm_bigram_df.loc[:,~top_500_asm_bigram_df.columns.duplicated()]\n\ntop_500_asm_bigram_df.to_csv(root_path + \"featurization/featurization_final/top_500_asm_opcodes_bigram_df.csv\",index=None)\n\ntop_500_asm_bigram_df.head()\n","metadata":{"id":"z79p3mkZV-t3","outputId":"69c73da9-503d-40e7-fee7-2ec027844da3","editable":false},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<h1 style=\"font-size:250%; font-family:cursive; color:#ff6666;\"><b> 47. Opcodes Trigrams ASM Files - Feature extraction <a id=\"47\"></a> </b></h1>\n\n\n#### [Back to the top](#0)\n\n","metadata":{"id":"C24aXD1UV-t3","editable":false}},{"cell_type":"code","source":"# Function to return all possible n*n*n combinations of trigrams\ndef calculate_trigram(tokens):\n    sent = \"\"\n    trigram_result = []\n    for i in range(len(tokens)):\n        for j in range(len(tokens)):\n            for k in range(len(tokens)):\n                trigram = tokens[i] + \" \" + tokens[j] + \" \" + tokens[k]\n                trigram_result.append(trigram)\n    return trigram_result\n  \n\n# test_tokens=['edx','esi','eax']\n# trigram_result = calculate_trigram(test_tokens)\n# trigram_result","metadata":{"id":"dEcZyMTPV-t4","editable":false},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Putting the list in a dictionary to decrease time complexity\nopcodes_trigram = ['jmp', 'mov', 'retf', 'push', 'pop', 'xor', 'retn', 'nop', 'sub', 'inc', 'dec', 'add','imul', 'xchg', 'or', 'shr', 'cmp', 'call', 'shl', 'ror', 'rol', 'jnb','jz','rtn','lea','movzx']\n\nopcodes_trigram_asm_vocabulary = calculate_trigram(\n    opcodes_trigram\n)  # Holding all n*n*n possible combinations of trigrams_from_asm_files\n\nvectorizer = CountVectorizer(\n    tokenizer=lambda x: x.split(),\n    lowercase=False,\n    ngram_range=(3, 3),\n    vocabulary=opcodes_trigram_asm_vocabulary,\n)  # NOTE: without \"tokenizer=lambda x: x.split()\", \"??\" would not get vectorized properly\n\nfile_lists_asm_opcodes = os.listdir(root_path + \"opcodes_asm_files\")\n\nfeatures = [\"ID\"] + vectorizer.get_feature_names()\n\nopcodes_asm_trigram_df = pd.DataFrame(columns=features)\n\nwith open(\n    root_path + \"featurization/opcodes_asm_trigram_df.csv\", mode=\"w\"\n) as opcodes_asm_trigram_df:\n    \n    opcodes_asm_trigram_df.write(\",\".join(map(str, features)))\n    \n    opcodes_asm_trigram_df.write(\"\\n\")\n    \n    for _, current_asm_textized_file in tqdm(enumerate(file_lists_asm_opcodes)):\n        each_file_id = current_asm_textized_file.split(\"_\")[0]\n        current_asm_textized_file = open(\n            root_path + \"opcodes_asm_files/\" + current_asm_textized_file\n        )\n        corpus_for_asm_files_opcodes = [\n            current_asm_textized_file.read().replace(\"\\n\", \" \").lower()\n        ]  # This will contain all the opcodes_trigram codes for a given current_asm_textized_file\n\n        # CountVectorizer produces a sparse representation of the counts using scipy.sparse.csr_matrix.\n        # Hence below is a sparse vector of all trigram counts from corpus_for_asm_files_opcodes\n        trigrams_from_asm_files = vectorizer.transform(corpus_for_asm_files_opcodes)\n\n        # So now return a dense ndarray representation of this matrix\n        # Updating each row_trigram_count of the dataframe with trigram counts\n        # of corresponding current_asm_textized_file\n        row_trigram_count = scipy.sparse.csr_matrix(trigrams_from_asm_files).toarray()\n\n        # Write that single row in the CSV for current_asm_textized_file\n        opcodes_asm_trigram_df.write(\n            \",\".join(map(str, [each_file_id] + list(row_trigram_count[0])))\n        )\n\n        opcodes_asm_trigram_df.write(\"\\n\")\n\n        current_asm_textized_file.close()\n\n\nopcodes_asm_trigram_df = pd.read_csv(\n    root_path + \"featurization/opcodes_asm_trigram_df.csv\"\n)\nopcodes_asm_trigram_df.head()","metadata":{"id":"Ht-sFvvyV-t4","outputId":"9d029078-cc37-4eb6-ad9a-b31630bab4fd","editable":false},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<h1 style=\"font-size:250%; font-family:cursive; color:#ff6666;\"><b> 48. ASM File - Top Important 800 features from Opcodes Trigrams <a id=\"48\"></a> </b></h1>\n\n\n#### [Back to the top](#0)\n\n\n\n\n\nThis will be the same sequence of steps what we applied earlier for extracting top 500 Features from ASM bigrams.","metadata":{"id":"IH8eUlQ6V-t4","editable":false}},{"cell_type":"code","source":"%%time \n\nX_opcode_asm_trigram = opcodes_asm_trigram_df\ny = class_labels\n# X_opcode_asm_trigram.head()\n\n#Get the best 500 features using SelectKBest. Save the feature scores along with the feature names in a feature_score_df_df, which we will use to get the best fetures from the bigrams df data\n\nkbest_object = SelectKBest(score_func=chi2, k=800)\n\ntop_features=kbest_object.fit(X_opcode_asm_trigram.drop(\"ID\", axis=1), y)\n\ntop_features_scores=pd.DataFrame(top_features.scores_)\n\nX_opcode_columns=pd.DataFrame(X_opcode_asm_trigram.columns)\n\ntop_asm_opcode_trigram_df=pd.concat([X_opcode_columns,top_features_scores],axis=1)\n\ntop_asm_opcode_trigram_df.columns=[\"ASM_Opcode_Top_Feature_Name\",\"ASM_Opcode_Top_Feature_Score\"]\n\ntop_asm_opcode_trigram_df=top_asm_opcode_trigram_df.nlargest(800,\"ASM_Opcode_Top_Feature_Score\")\n\ntop_asm_opcode_trigram_df.head()","metadata":{"id":"7bsdl4NgV-t4","outputId":"80e505cc-8c42-4d96-cae7-970bf33fcbd9","editable":false},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\n\n# Get List of the 800 top features\ntop_800_asm_trigram_features=list(top_asm_opcode_trigram_df[\"ASM_Opcode_Top_Feature_Name\"])\n\ntop_800_asm_trigam_df=pd.concat([X_opcode_asm_trigram[\"ID\"], X_opcode_asm_trigram[top_800_asm_trigram_features]], axis=1)\n\n# The \"ID\" column was being duplicated, hence need to remove that, and also the possibility of any other duplicated column\ntop_800_asm_trigam_df = top_800_asm_trigam_df.loc[:,~top_800_asm_trigam_df.columns.duplicated()]\n\ntop_800_asm_trigam_df.to_csv(root_path + \"featurization/featurization_final/top_800_asm_opcodes_trigram_df.csv\",index=None)\n\ntop_800_asm_trigam_df.head()\n\n","metadata":{"id":"3ywmsPWJV-t5","outputId":"04239281-779b-4039-96e3-ee75b8c9cc95","editable":false},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<h1 style=\"font-size:250%; font-family:cursive; color:#ff6666;\"><b> 49. Final Merging of all Features for the Final XGBOOST Training <a id=\"49\"></a> </b></h1>\n\n#### [Back to the top](#0)\n\n\n\n## - Unigram of Byte Files + Size of Byte Files + \n\n## - Top 52 Unigram of ASM Files + Size of ASM Files\n\n## - Top 2000 Bi-Gram of Byte files +  \n\n## - Top 500 Bigram of Opcodes of ASM Files\n\n## - Top 800 Trigram of Opcodes of ASM Files\n\n## - Top 800 ASM Image Features","metadata":{"id":"Txo2rq59V-t5","editable":false}},{"cell_type":"code","source":"%%time\n\n# Unigram of Byte Files + Size of Byte Files + \nuni_gram_byte_features__with_size = pd.read_csv(\n    root_path + \"featurization/uni_gram_byte_features__with_size.csv\"\n)\n\n# Top 52 Unigram of ASM Files  + Size of ASM Files\n# Droping .BSS, .rtn, .CODE features from the unigram_asm_feature__with_size (which is the unigram of asm files) dataset\n# As we earlier saw that these features were not much important in separating class labels\nunigram_asm_feature__with_size = pd.read_csv(\n    root_path + \"featurization/unigram_asm_feature__with_size\"\n).drop([\"Class\", \"rtn\", \".BSS:\", \".CODE\"], axis=1)\n\n# Top 2000 Bi-Gram of Byte files\n# top_2000_imp_byte_bigram_df = pd.read_csv(\n#     root_path + \"featurization/featurization_final/top_2000_imp_byte_bigram_df.csv\"\n# ).drop(columns=[\"ID.1\"])\n\ntop_2000_imp_byte_bigram_df = pd.read_csv(\n    root_path + \"featurization/featurization_final/top_2000_imp_byte_bigram_df.csv\"\n)\n\n# Top 500 Bigram of Opcodes of ASM Files\ntop_500_asm_bigram_df = pd.read_csv(root_path + \"featurization/featurization_final/top_500_asm_opcodes_bigram_df.csv\")\n\n\n# Top 800 Trigram of Opcodes of ASM Files\ntop_800_asm_trigam_df = pd.read_csv(root_path + \"featurization/featurization_final/top_800_asm_opcodes_trigram_df.csv\")\n\n# Top 800 ASM Image Features\ntop_800_image_asm_df = pd.read_csv(root_path + \"featurization/top_800_image_asm_df.csv\")\n\n","metadata":{"id":"LIguBsBRV-t5","outputId":"050d8881-8010-45a4-b959-5e4f6cea7ef7","editable":false},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\n\n# Initiate a dataframe for representing the Combined Features\n# and set it equal to uni_gram_byte_features__with_size\ncombined_features_final_df = uni_gram_byte_features__with_size\n\nindividual_featuarized_dfs = [\n    unigram_asm_feature__with_size,\n    top_800_image_asm_df,\n    top_2000_imp_byte_bigram_df,\n    top_500_asm_bigram_df,\n    top_800_asm_trigam_df\n]\n\nfor df in tqdm(individual_featuarized_dfs):\n    # combined_features_final_df = pd.merge(combined_features_final_df, df, on=\"ID\", how=\"left\")\n    combined_features_final_df = pd.merge(combined_features_final_df, df, on=\"ID\")\n\ncombined_features_final_df.to_csv(\n    root_path + \"featurization/featurization_final/combined_features_final_df.csv\",\n    index=None,\n)\n\ncombined_features_final_df.head()","metadata":{"id":"tpXNZmWqV-t6","outputId":"fb21dd27-512d-4009-e150-7c5024be67d9","editable":false},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<h1 style=\"font-size:250%; font-family:cursive; color:#ff6666;\"><b> 50. Final Train Test Split. 64% Train, 16% Cross Validation, 20% Test <a id=\"50\"></a></b></h1>\n\n\n#### [Back to the top](#0)\n\n","metadata":{"id":"8oLxkOdeV-t6","editable":false}},{"cell_type":"code","source":"combined_features_final_df = pd.read_csv(root_path + \"featurization/featurization_final/combined_features_final_df.csv\")\n\ncombined_features_final_df_normalized = normalize(combined_features_final_df)\n\ncombined_features_final_df_normalized.to_csv(root_path + \"featurization/featurization_final/combined_features_final_df_normalized.csv\", index=None)","metadata":{"editable":false},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\n\nfinal_X = pd.read_csv(root_path + \"featurization/featurization_final/combined_features_final_df_normalized.csv\").fillna(0).drop(['ID'], axis=1)\n\nfinal_y = pd.read_csv(root_path + \"featurization/featurization_final/combined_features_final_df_normalized.csv\")[\"Class\"]\n\n\n# Splitting - Keep same distribution of class label 'y_true' with [stratify=final_y]\nX_train, X_test_final_merged, y_train, y_test_final_merged = train_test_split(final_X, final_y, stratify=final_y, test_size=0.20, random_state=42)\n\nX_train_final_merged, X_cv_final_merged, y_train_final_merged, y_cv_final_merged = train_test_split(X_train, y_train, stratify=y_train, test_size=0.20, random_state=42)\n\nprint('Shape of X_train_final_merged and y_train_final_merged: ', X_train_final_merged.shape, y_train_final_merged.shape)\n\nprint('Shape of X_test_final_merged and y_test_final_merged: ', X_test_final_merged.shape, y_test_final_merged.shape)\n\nprint('Shape of X_cv_final_merged and y_cv_final_merged ', X_cv_final_merged.shape, y_cv_final_merged.shape)","metadata":{"id":"_yhf9F7nV-t6","outputId":"8ceb5802-ceae-4082-df35-439cd32db49b","editable":false},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<h1 style=\"font-size:250%; font-family:cursive; color:#ff6666;\"><b> 51. Final XGBoost Training - Hyperparameter tuning with on Final Merged Data-Matrix <a id=\"51\"></a> </b></h1>\n\n\n#### [Back to the top](#0)\n\n","metadata":{"id":"QPxSh_y6V-t7","editable":false}},{"cell_type":"code","source":"%%time\n\nxgb_clf=XGBClassifier()\n\nprams={\n    'learning_rate':[0.01,0.03,0.05,0.1,0.15,0.2],\n     'n_estimators':[100,200,500,1000,2000],\n     'max_depth':[3,5,10],\n    'colsample_bytree':[0.1,0.3,0.5,1],\n    'subsample':[0.1,0.3,0.5,1],\n    'tree_method':['gpu_hist']\n}\n\nrandom_clf=RandomizedSearchCV(xgb_clf, param_distributions=prams, verbose=10, n_jobs=-1)\n\nrandom_clf.fit(X_train_final_merged, y_train_final_merged)\n\nprint(random_clf.best_params_)","metadata":{"id":"zJuNvlmnV-t7","outputId":"cfb38ec6-7af6-4b19-c13c-a43c16e046ae","editable":false},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Best Params we got from above RandomizedSearchCV\n\n<h1 style=\"font-size:250%; font-family:cursive; color:#ff6666;\"><b> 52. Final running of XGBoost with the Best HyperParams that we got from above RandomizedSearchCV <a id=\"52\"></a> </b></h1>\n\n\n#### [Back to the top](#0)\n\n\n","metadata":{"id":"sLvs6EU-V-t7","editable":false}},{"cell_type":"code","source":"%%time\n\nn_estimators = random_clf.best_params_['n_estimators']\nsubsample = random_clf.best_params_['subsample']\nmax_depth = random_clf.best_params_['max_depth']\nlearning_rate = random_clf.best_params_['learning_rate']\ncolsample_bytree = random_clf.best_params_['colsample_bytree']\ntree_method = random_clf.best_params_['tree_method']\n\n# print(tree_method)\n\nx_clf_with_best_hyper_param=XGBClassifier(n_estimators=n_estimators, max_depth=max_depth, learning_rate= learning_rate, colsample_bytree=colsample_bytree, subsample=subsample, tree_method=tree_method, nthread=-1)\n\nx_clf_with_best_hyper_param.fit(X_train_final_merged, y_train_final_merged, verbose=True)\n\nsig_clf = CalibratedClassifierCV(x_clf_with_best_hyper_param, method=\"sigmoid\")\n\nsig_clf.fit(X_train_final_merged, y_train_final_merged)","metadata":{"id":"-pwQEWF0V-t8","outputId":"72c5819e-a651-4008-c4ef-8a6381fcc33b","editable":false},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\n\nn_estimators = random_clf.best_params_['n_estimators']\n\n# LOGLOSS FOR TRAIN\n\npredict_y_train = sig_clf.predict_proba(X_train_final_merged)\n\nprint ('With best number of estimators = ', n_estimators, \"Our train log loss is:\", log_loss(y_train_final_merged, predict_y_train))\n\n\n# LOGLOSS FOR TEST\n\npredict_y_test = sig_clf.predict_proba(X_test_final_merged)\n\nprint('For values of best number of estimators = ', n_estimators, \"The test log loss is:\", log_loss(y_test_final_merged, predict_y_test))\n\n\n# LOGLOSS FOR CV\n\npredict_y_cv = sig_clf.predict_proba(X_cv_final_merged)\n\nprint('With best number of estimators = ', n_estimators, \"Our cross validation log loss is:\", log_loss(y_cv_final_merged, predict_y_cv))\n\n","metadata":{"id":"qc2mhk6RV-t8","outputId":"d9a03bc2-e9a3-4c3e-df78-de9de251dedd","editable":false},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"![Imgur](https://imgur.com/l4j90Uj.png)\n\n\n<h1 style=\"font-size:250%; font-family:cursive; color:#ff6666;\"><b> 53. Possibiliy of Further Analysis and Featurizition <a id=\"53\"></a> </b></h1>\n\n\n#### [Back to the top](#0)\n\n\nWe could experiment further with following features.\n\n- names of the imported functions\n- Libraries used\n- the number of procedures used.\n- Computing the ”constitionality” of the executable: the number of ”loc *” references in .asm file.\n- According to some papers, the malware content is often encrypted inside the binary, so we could introduce some encryption related features. \n- Computing entropy over a sliding window of half-byte sequences,\n- Extracting some statistics of its distribution (20 quantiles, 20 percentiles, mean, median, std, max, min, max-min), \n- Statistics of first order differences distribution and parts of the entropy sequence. \n- Computing compression ratio (as an approximation of Kolmogorov complexity).\n","metadata":{"editable":false}}]}