{"metadata":{"kaggle":{"accelerator":"none","dataSources":[{"sourceId":59865,"databundleVersionId":6660280,"sourceType":"competition"}],"dockerImageVersionId":30635,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":false},"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"},"language_info":{"name":"python","version":"3.10.12","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"Intro\n    I have a background in networking through a job and a college class, so I started this project as an \"easy\" intro to data science due to it pertaining to a subject I'm familar with. Some ideas (TCP/ UDP counts, ports attacked, etc.) felt natural to look out for, especially those regarding the number of attacks in a given time frame. Much of my work was based on the assumption those attacking with a VPN/ Proxy will do so sparingly so as to not draw attention to themselves (otherwise they wouldn't use VPN/ Proxy) and those spamming likley don't care about being seen (or the attacks are automated). I was also given some advice from CrowdSec's comment about a month into the challenge, timed events. This spurred on the creation of many valuable features (number of specific attacks made by an IP in a given window, number of attacker_ip_enums using a given an attacker_as_num in a time range, etc.). With a great deal of experimentation I found the best window for my model was around eight days, and I stuck with it throughout all my measurements. \n    As uninteresting as it sounds I found CatBoost to work the best through trial and error. Boosted Forest appears to be the king in many binary classification challenges on Kaggle so it's where I started, and it's where I stayed after trying quite a few other models. I then tested between CatBoost, XGBoost, and RandomForestClassifier until I found CatBoost consistently worked better than the other two. After that I poured over the CatBoost docs finding out what each parameter did and thinking on how it could help me. Hyperparamter tuning, much like finding the right model, came down to experimentation and general recommendations from online communities (basically reading StackExchange forums that gave a range (ex. learning_rate should ideally be between 0.001 and 0.1)).\n    My original script could be run without a problem on an M1 Mac, but when attempting to run on Kaggle I constantly found myself running out of available memory. If you're interested in seeing the original, feel free to reach out, I'll share it and maybe you'll have a few more suggestions you can send my way! I probably (most certainly) overused python's garbage collecting module (gc.collect()) but once I got the script in a place where it could run on Kaggle I stopped working on it and didn't risk removing anything. I learned the value of functions both because of avoiding repetition and because they discard values that aren't returned (ensuring no garbage variables stuck around). If I had more time I'd also go back through and remove creating new dataframes to perform calculations. I worked on those functions that used the most memory first and once it was in a place to run on Kaggle I stopped touching everything.","metadata":{}},{"cell_type":"code","source":"#Imports\n#These aren't cleaned up. Many of these modules/ libraries aren't used in this notebook, but I thought it may give you a better idea of what I used to test before my final submission\nimport datetime\nimport gc\nimport lightgbm as lgb\nimport math\nimport matplotlib.patches as mpatches\nimport matplotlib.pyplot as plt\nimport numpy as np\nimport os\nimport pandas as pd\nimport pyarrow as pya\nimport re\nimport scipy.stats as stats\nimport sklearn\nimport seaborn as sns\nfrom catboost import CatBoostClassifier, Pool\nfrom category_encoders import *\nfrom datetime import datetime, timedelta\nfrom scipy.stats import zscore\nfrom scipy.special import logsumexp\nfrom sklearn import linear_model\nfrom sklearn import metrics\nfrom sklearn.compose import ColumnTransformer\nfrom sklearn.ensemble import RandomForestRegressor, RandomForestClassifier, GradientBoostingClassifier, BaggingClassifier, AdaBoostClassifier, StackingClassifier, VotingClassifier, HistGradientBoostingClassifier, ExtraTreesClassifier\nfrom sklearn.feature_selection import mutual_info_regression\nfrom sklearn.impute import SimpleImputer\nfrom sklearn.linear_model import LogisticRegression, RidgeClassifier, RidgeClassifierCV\nfrom sklearn.metrics import accuracy_score, mean_squared_error, mean_absolute_error, cohen_kappa_score, f1_score, confusion_matrix\nfrom sklearn.model_selection import GroupShuffleSplit, train_test_split, cross_val_score, TimeSeriesSplit\nfrom sklearn.multiclass import OneVsRestClassifier\nfrom sklearn.naive_bayes import GaussianNB\nfrom sklearn.neighbors import KNeighborsClassifier, NearestCentroid\nfrom sklearn.neural_network import MLPClassifier, BernoulliRBM\nfrom sklearn.pipeline import Pipeline\nfrom sklearn.preprocessing import OrdinalEncoder, OneHotEncoder, LabelEncoder, MinMaxScaler, StandardScaler, Normalizer, power_transform\nfrom sklearn.svm import SVC, LinearSVC, NuSVC\nfrom sklearn.tree import DecisionTreeRegressor, DecisionTreeClassifier\nfrom sklearn.svm import SVC\nfrom xgboost import XGBClassifier","metadata":{"execution":{"iopub.status.busy":"2024-01-27T16:18:11.000846Z","iopub.execute_input":"2024-01-27T16:18:11.001472Z","iopub.status.idle":"2024-01-27T16:18:11.015219Z","shell.execute_reply.started":"2024-01-27T16:18:11.001374Z","shell.execute_reply":"2024-01-27T16:18:11.014086Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"shodan_df_hashed Feature Engineering\nThis function is used to feature engineer the values in shodan hashed, most notably the fingerprints. I made the assumption that those IP's sharing fingerprints (perhaps through an assignment based on their MAC address) were likely using VPN's/ Proxies. Therefore, the most valuable data here is in the IP's shared among each fingerprint. Besides that I simply counted the number of times each protocol (TCP or UDP) was used when targeting a specific port.","metadata":{}},{"cell_type":"code","source":"#shodanHashedAlt\ndef shodanHashedAlt(shodanCSV):\n    \n    print(\"\\tStarting MultiIndex of Fingerprints...\")\n    #Adds three columns with jarm, headers_hash, and ja3s fingerprints (held as lists)\n    shodanCSV = shodanCSV.replace([\"'\", \",\", \"{\", \"}\", \"None\"], \"\", regex=True) #Removing the following characters to make regex easier, since shodan_info is imported as a string\n    #Creating the pattern to collect the fingerprints, port numbers, and connection type used (storing lists in a dictionary)\n    patternDict = {\"jarm\":\"jarm: (\\S+)\", \"headers_hash\":\"headers_hash: (\\S+)\", \"ja3s\":\"ja3s: (\\S+)\", \"TCPCount\":\"(\\S*/tcp)\", \"UDPCount\":\"(\\S*/udp)\", \"ports\": \"(\\d+)\\/\"}\n    for i in patternDict:   #Extracting each of these patterns and placing them in their own columns\n        tempMultiIndex = shodanCSV[\"shodan_info\"].str.extractall(patternDict[i])    #Extracting the value pattern based on the key used in the iterator i\n        tempMultiIndex.columns = [i]    #Renaming the (only) column in the MultiIndex for readability during testing based on the key in the iterator i\n        tempMultiIndex = tempMultiIndex.groupby(level=0)[i].apply(list)    #Grouping by the first level of the MultiIndex and turning it into a list\n        shodanCSV = pd.concat([shodanCSV, tempMultiIndex], axis= 1)    #Combining shodanCSV with the MultiIndex (because indeces are saved during the extraction this is easy to do)\n        tempMultiIndex = None\n        \"\"\"\n        If this part is confusing no worries, there probably is an easier way to do it, but here's an explanation for the way this runs\n        tempMultiIndex = shodanCSV[\"shodan_info\"].str.extractall(patternDict[i])\n            Creates a MultiIndex (from 0-3) like so (we'll assume it's itterating on headers_hash)...\n                index row headers_hash\n                1     0   ###################\n                      1   ###################\n                3     0   ###################\n            Notice indeces 0 and 2 are missing? This is due to the regex extractall not finding a matching pattern in their shodan_df_hashed row.\n            So we have attacker_ip_enum's basically kept track of by their indices and all of their headers_hash from the shodan_info column in the multi-index\n        tempMultiIndex.columns = [i]\n            Simply ensures (by setting it (even redudently)) the column name is the same as the key currently itterated on\n        tempMultiIndex = tempMultiIndex.groupby(level=0)[i].apply(list)\n            Merges the rows for each index into a list, basically turning the MultiIndex into a series or single column dataframe\n            index row headers_hash\n            1     0   [###################, ###################]\n            3     0   [###################]\n        shodanCSV = pd.concat([shodanCSV, tempMultiIndex], axis= 1)\n            Joins shodanCSV with the MultiIndex. Those indeces without a list are given np.NaN as their value\n        \"\"\"\n        \n        #So now shodanCSV has the following columns: shodan_info (unchanged), attacker_ip_enum (unchanged), jarm, headers_hash, ja3s, TCPCount, UDPCount, and ports\n\n    #Here's the plan: take in each set ONLY if it's not empty, make it a series (single column dataframe), then check each index in each set in each row for duplicates against the newly created dataframe\n    fingerprintList = [\"jarm\", \"headers_hash\", \"ja3s\"]  #Creating a list of strings to reference each specific column of the dataframe in a for loop\n\n    for i in fingerprintList:   #Turning each column of jarm, ja3s, and headers hash into a set of each fingerprint (some columns are empty others have multiple)            \n        shodanCSV[i] = shodanCSV[i].fillna(\"\").apply(set)   #Turned into a set so the duplicates from one IP address are removed\n        #I.E. If one IP attacked multiple ports and has the same header hash twice, it will be knocked down to one so duplicates aren't found in cases where one IP attacks a port twice but no other IP uses the same header hash\n\n    print(\"\\tFrequency of each print...\")\n    #Have to make seperate dataframes due to their varying length (each column in a dataframe must have the same length, that conflicts with what I'm doing)\n    #This creates a single column data frame where each row is a fingerprint. Duplicates ARE allowed, but only so long as they didn't come from the same IP because they're indexed from sets (above), which don't allow duplicates\n    JARM = pd.DataFrame(list([a for b in shodanCSV[\"jarm\"].tolist() for a in b])) #Doesn't bring over empty lists, which is ideal\n    HH = pd.DataFrame(list([a for b in shodanCSV[\"headers_hash\"].tolist() for a in b]))\n    JATS = pd.DataFrame(list([a for b in shodanCSV[\"ja3s\"].tolist() for a in b]))\n    fingerprintBook = [JARM, HH, JATS]\n\n    #Creates a column for the frequency of each fingerprint in tracking style\n    #Value counts removes duplicates from the column, which is fine because the column created (count) shows how many there were\n    for i in range(0, len(fingerprintBook)):\n        fingerprintBook[i] = pd.DataFrame(fingerprintBook[i].value_counts())\n        fingerprintBook[i] = fingerprintBook[i].reset_index(names=[\"unique_values\"])\n        fingerprintBook[i][\"count\"] = fingerprintBook[i][\"count\"] - 1     #Subtracting one because I must take into consideration the IP's fingerprint itself (so if an IP is the only one with a fingerprint, it's shared said fingerprint with no one else, resulting in a 0)\n    \n    #Creates a dataframe with the IPs in one column and fingerprints in the other\n    #It's exploded, so there can be multiple rows with the same IP and fingerprint combo\n    print(\"\\tExploding each fingerprint type...\")\n    for i in range(len(fingerprintBook)):\n        exploded_SHFP = pd.DataFrame(shodanCSV[fingerprintList[i]].explode()).reset_index()    #Creates an exploded dataframe of the IP and a fingerprint type (JARM, JA3s, or Headers Hash)\n        \n        fingerprintBook[i].set_index(\"unique_values\", inplace= True)    #Sets the index of the fingerprint dataframe to the column of unique fingerprints (leaving the new index and a column of counts)\n        exploded_SHFP = exploded_SHFP.join(fingerprintBook[i], on= fingerprintList[i])    #Joins the column of counts to the exploded fingerprint dataframe based on the fingerprints\n        exploded_SHFP.drop([fingerprintList[i]], axis= 1, inplace= True)    #Drops the unique fingerprints column (no longer necessary, only need the IP and counts column)\n        \n        exploded_SHFP = exploded_SHFP.groupby([\"index\"]).agg({\"count\":\"sum\"})    #Groups by the unique IP and aggregate sums the counts column (creating a total of shared fingerprints)\n        shodanCSV = shodanCSV.join(exploded_SHFP)   #Joins the counts column to shodanCSV (where all the data is stored)\n        shodanCSV = shodanCSV.rename(columns= {\"count\": f\"{fingerprintList[i]}_sum\"})   #Renames the counts column to the fingerprintType_sum, so it's not overwritten when the loop reitterates\n        shodanCSV[f\"{fingerprintList[i]}_sum\"] = shodanCSV[f\"{fingerprintList[i]}_sum\"].astype(\"float32\")\n    \n    \"\"\"\n    #Iterates through the rows of shodanCSV assigning a value to fingerprint_shared column based on the count from a dictionary\n    #The dictionary is created per fingerprint book and contains the unique value (fingerprint) and the count\n    #The calculate_sum function finds the unique fingerprints in each row and sums their count, adding it to the fingerprint_shared column\n    #This is how it was performed previously. I often ran out of memory, which I assume was due to itterating over every row coupled with incredibly long strings from the fingerprints\n    print(fingerprintBook[i].info())\n    \n    for i in range(0, len(fingerprintBook)):\n        #multi_index_sh = pd.DataFrame(shodanCSV[f\"{fingerprintList[i]}\"].explode()).reset_index()\n        exit()\n        count_mapping = dict(zip(fingerprintBook[i][\"unique_values\"], fingerprintBook[i][\"count\"]))\n        shodanCSV[f\"{fingerprintList[i]}_shared\"] = np.vectorize(calculate_sum)(shodanCSV[fingerprintList[i]])\n    \"\"\"\n    \n    print(\"\\tStarting Protocol Count...\")\n    protocolC = [\"TCPCount\", \"UDPCount\"]    #Acquiring the count of UDP and TCP protocols used in an attack (those with more are more likely to be spamming and unlikely to use a VPN (assumed out of laws reach))\n    for i in protocolC:\n        shodanCSV[i] = shodanCSV[i].fillna(\"\").apply(list)    #Fills the empty values from the lists created in the loop above with my own empty value, just so I know what to reference\n        shodanCSV[i] = shodanCSV[i].str.len()    #Counts the length of the list, giving the number of TCP ports and UDP ports attacked\n        shodanCSV[i] = shodanCSV[i].astype(\"int32\")\n        \n    #Checks for those who have attacked ports but don't have any fingerprints (this is both those with \"\" and None, although they are different and probably should've been treated differently)\n    shodanCSV[\"NoPrintWithAttack\"] = np.where(((shodanCSV[\"TCPCount\"] > 0) | (shodanCSV[\"UDPCount\"] > 0))  & ((shodanCSV[\"jarm\"].str.len() == 0) & (shodanCSV[\"headers_hash\"].str.len() == 0) & (shodanCSV[\"ja3s\"].str.len() == 0)), 1, 0)\n    shodanCSV[\"NoPrintWithAttack\"] = shodanCSV[\"NoPrintWithAttack\"].astype(\"int32\")\n    shodanCSV = shodanCSV.drop([\"shodan_info\", \"jarm\", \"headers_hash\", \"ja3s\", \"ports\"], axis= 1)\n    \n    \n    return shodanCSV","metadata":{"execution":{"iopub.status.busy":"2024-01-27T16:18:11.110721Z","iopub.execute_input":"2024-01-27T16:18:11.111394Z","iopub.status.idle":"2024-01-27T16:18:11.143696Z","shell.execute_reply.started":"2024-01-27T16:18:11.111350Z","shell.execute_reply":"2024-01-27T16:18:11.142181Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"General Feature Engineering\nThis is where most of the feature engineering takes place. More detail on each function can be found in comments around them, including their purpose.","metadata":{}},{"cell_type":"code","source":"def FindingAverageAttackerIPs(X, window):\n    print(\"Finding average attacker IP's shared among\")\n    #Finding the average attacker IP shared amoung attacker_as_num's\n    X[\"attacker_as_num\"].replace(np.NaN, -1, inplace= True)\n    X.reset_index(inplace= True)\n    #Finding the unique count of attacker_ip_enum per attacker_as_num and merging the resulting column into the original dataframe\n    X = pd.merge(left= X, right= X.groupby(\"attacker_as_num\")[\"attacker_ip_enum\"].nunique().reset_index().rename(columns= {\"attacker_ip_enum\": \"attNumNUniqueIPs\"}), on= \"attacker_as_num\", how= \"inner\")\n    \n    df = X[[\"attacker_as_num\", \"attacker_ip_enum\", \"attNumNUniqueIPs\"]].copy(deep= True)  #Remaking copy with necessary columns\n    X.drop([\"attNumNUniqueIPs\"], axis= 1, inplace= True)    #Dropping, now unneeded, attNumNUniqueIPs column\n    df.drop_duplicates(subset= [\"attacker_ip_enum\", \"attacker_as_num\"], inplace= True)   #Dropping duplicates of attacker_as_num for each attacker_ip_enum\n    df.drop([\"attacker_as_num\"], axis= 1, inplace= True) #Dropping the now unnecessary column (still segregated by unique attNumNUniqueIPs). Done in this way because some attacker_ip_enums may have multiple attacker_as_nums with the same attNumNUniqueIPs\n    df = df.groupby(\"attacker_ip_enum\")[\"attNumNUniqueIPs\"].mean().reset_index().rename(columns= {\"attNumNUniqueIPs\": \"avgSharedIPsWithAttHost\"})  #Grabbing the mean of attNumNUniqueIPs for each attacker_ip_enum\n    \n    #Adding the new column to the original data frame\n    X = pd.merge(left= X, right= df, on= \"attacker_ip_enum\", how= \"inner\")    #Merging the average calculated in df with the original dataframe, X\n    X[\"avgSharedIPsWithAttHost\"] = X[\"avgSharedIPsWithAttHost\"].astype(\"float32\")\n\n    df = None   #Clearing for mem\n    X.set_index(\"index\", inplace= True) #Reinstating \"index\" column as the index\n    X.sort_index(inplace= True)   #Sorting by \"index\" column\n    return X\n\ndef FindingCountUniqueWatAttCom(X, window):\n    print(\"Finding a count of unique\")\n    #Finding a count of unique Watcher_as_Num and attack_type combos in a given time frame\n    X.reset_index(inplace= True)    #Resetting index to keep ordering when merging at the end\n    \n    #X[\"attack_type\"] = X[\"attack_type\"].astype(\"int32\") #Attack_type is categorical during the XTe run through unless explicitly changed it need to be int for the unique ID made...\n    X[\"CombWatNumAttType\"] = X[\"watcher_as_num\"] + (X[\"attack_type\"]/1000)    #here. This takes watcher_as_num and the categorical attack_type and makes a unique ID in a float32 format (etc. watcher_as_num.0attack_type)\n    df = X[[\"CombWatNumAttType\", \"attacker_ip_enum\", \"attack_time\"]].copy(deep= True)  #Making copy of the original dataframe to perform calculations on\n    \n    df = df.groupby([\"attacker_ip_enum\", pd.Grouper(key=\"attack_time\", freq= window)])[\"CombWatNumAttType\"].nunique().reset_index(name= \"NUniqueWattAtt\")   #Creating a new value measuring the number of unique watcher_as_num+attack_type combinations in a given window\n    \n    X[\"rounded_time\"] = X[\"attack_time\"].dt.floor(window)   #Creating a new column with the value being the floor of the window of attack_time\n    X = pd.merge(X, df, left_on= [\"attacker_ip_enum\", \"rounded_time\"], right_on=[\"attacker_ip_enum\", \"attack_time\"])    #Merging the new column into the original dataframe based on attacker_ip_enum and the new rounded_time\n    df = None   #Clearing for mem\n    X.drop([\"rounded_time\", \"attack_time_y\", \"CombWatNumAttType\"], axis= 1, inplace= True)   #Dropping the rounded_time and attack_time column from the copied dataframe\n    X = X.rename(columns= {\"attack_time_x\": \"attack_time\"}) #Renaming attack_time_x (happened during the merge) to attack_time for easier reference later\n    X.set_index(\"index\", inplace= True) #Resetting the index as it was sorted when merged\n    X[\"NUniqueWattAtt\"] = X[\"NUniqueWattAtt\"].astype(\"int32\")\n    return X\n\ndef AddingCountAtt(X, window):\n    print(\"Adding a count of attacks carried out\")\n    #Simply adding a count of attacks carried out in a given time frame (10 seconds)\n    #Making a copy of the original dataframe (this will soon only contain the count by the end of this process)\n    df = X.filter([\"attacker_ip_enum\", \"attack_time\"], axis= 1).copy() #Currently only contains attacker_ip_enum and attack_time, the two values it'll group and roll a window over\n    \n    X.to_parquet(\"XTemp.parq\")\n    X = None\n    gc.collect()\n    \n    #Adding a column to the dataframe with the original index as the new index will be adjusted after grouping, rolling, and sorting\n    df = df.reset_index()\n    #Sorting the dataframe by time. To perform a rolling window the data must be chronological\n    df.sort_values(by= \"attack_time\", inplace= True)\n    #Grouping by attacker_ip_enum, then rolling a window over a period of time, and taking the count. As well as resetting the index (because groupby makes a multi-index) and makes \"level_0\" the new index\n    #So the column called \"index\" becomes the count and the column \"level_0\" becomes the old index\n    df = df.groupby(by= \"attacker_ip_enum\", sort= False, as_index= False).rolling(\"10S\", center= True, on= \"attack_time\").count().reset_index().set_index(\"level_0\")\n    #Sorting the index (level_0, the old, original index) back into order\n    df = df.sort_index(axis= 0)\n    df = df.drop([\"attack_time\", \"attacker_ip_enum\"],axis= 1)   #All that remains is the index (level_0) and the count (index)\n    #At this point df is an index (length of X) and a single column (the count of shared IPs in the rolling window time frame)\n    \n    X = pd.read_parquet(\"XTemp.parq\")\n    \n    #Moving the column from df to X\n    X[\"SharedIPsInRange\"] = df[\"index\"].values\n    X[\"SharedIPsInRange\"] = X[\"SharedIPsInRange\"].astype(\"float32\")\n    #Clear mem space\n    df = None\n    return X\n\ndef NumberSpecificAtt(X, window, uniqueList= None):\n    print(\"Number of specific types\")\n    #Counting the number of times an IP address issued a specific attack in a given time frame (8 Days) (for each attack except those in the list below)\n    #unnecessaryAttTypes = [\"windows:bruteforce_8D\", \"sip:bruteforce_8D\", \"telnet:bruteforce_8D\", \"smb:bruteforce\", \"database:bruteforce\", \"ftp:bruteforce\"]    #Added after testing in hopes of limiting unnecessary memory usage\n    unnecessaryAttTypes = [14, 9, 12, 8, 0, 1]\n    \n    if uniqueList is None:    #Unique List is created here for the Training data...*\n        uniqueList = X[\"attack_type\"].unique()    #This creates a list of all the unique attack types http:bruteforce, tcp:scan, etc. to iterate through and find how many occurances there are in the given window\n    for i in range(len(uniqueList)):    #*... and used here for the Testing data. This is done to ensure the column ordering stays the same throughout (as is required for many models)\n        if uniqueList[i] in unnecessaryAttTypes:\n            pass\n        else:\n            print(f\"\\t{uniqueList[i]}\")\n            df = pd.DataFrame    #Creating an empty dataframe\n            df = X[X[\"attack_type\"] == uniqueList[i]].filter([\"attack_time\", \"attacker_ip_enum\"], axis= 1)  #Creates a dataframe in the list of only the attack_time and attacker_ip_enum\n\n            X.to_parquet(\"XTemp.parq\")\n            X = None\n            gc.collect()\n\n            df.reset_index(inplace= True)    #Keeping the original index as a column as the index will be reset when groupby is used\n            df.sort_values(by= \"attack_time\", inplace= True)    #Sorting for the sake of the rolling window\n            df[f\"{uniqueList[i]}_{window}\"] = df.groupby(by= \"attacker_ip_enum\", sort= False, as_index= False)[[\"attacker_ip_enum\", \"attack_time\"]].rolling(window= window, center= True, on= \"attack_time\").count()\n            df.sort_values(by= \"index\", inplace= True)    #Putting the dataframe back into it's original order\n            df.drop([\"attack_time\", \"attacker_ip_enum\"], axis= 1, inplace= True)   #Dropping the unecessary columns (the index and calculated data is all that matters), it's all still there in the main dataframe\n            df.set_index(\"index\", inplace= True)    #Setting the index back to the true index\n\n            X = pd.read_parquet(\"XTemp.parq\")\n\n            X = pd.concat([X, df], axis= 1)    #Joining the calculated data with the original dataframe\n            X[f\"{uniqueList[i]}_{window}\"].fillna(0, inplace= True)    #Renaming the column to attacktype_window (ex. tcp:scan_8D)\n            X[f\"{uniqueList[i]}_{window}\"] = X[f\"{uniqueList[i]}_{window}\"].astype(\"float32\")\n            df = None    #Clearing for mem\n    return X, uniqueList\n\ndef FindingTimeDiff(X, window):\n    print(\"Finding the time diff between the\")\n    #Finding the difference in time between each attack from each attacker_ip_enum (the first attack is 0)\n    df = X[[\"attacker_ip_enum\", \"attack_time\"]].copy(deep= True)    #Making a copy df to perform calculations on\n    \n    X.to_parquet(\"XTemp.parq\")\n    X = None\n    gc.collect()\n    \n    df[\"time_diff\"] = df.sort_values([\"attack_time\"]).groupby(\"attacker_ip_enum\")[\"attack_time\"].diff().dt.total_seconds()  #Creating a column that measures the difference in time between one attack and the one preceding it (those without a predecessor are marked N/A)\n    df[\"time_diff\"].replace(value= 0 , to_replace= np.NaN, inplace= True)   #Changing N/A's to 0's for model's sake\n    df.drop([\"attacker_ip_enum\", \"attack_time\"], axis= 1, inplace= True)    #Dropping now unnecessary columns\n    \n    X = pd.read_parquet(\"XTemp.parq\")\n    \n    X = X.join(df, how=\"outer\")  #Joining the calculated column with the original dataframe\n    X[\"time_diff\"] = X[\"time_diff\"].astype(\"float32\")\n    df = None    #Clearing for mem\n    return X\n\ndef AttIPNumWat(X, window):\n    print(\"Just the number of watchers each\")\n    #Just the unique number of watcher_as_num each attacker_ip_enum attacked in total\n    df = X[[\"attacker_ip_enum\", \"watcher_as_num\"]].copy(deep= True)    #Making a copy to preserve the original dataframe\n    \n    X.to_parquet(\"XTemp.parq\")\n    X = None\n    gc.collect()\n    \n    df = df.groupby(\"attacker_ip_enum\", as_index= False)[\"watcher_as_num\"].nunique().reset_index()    #Grouping by attacker_ip_enum and including only watcher_as_num in the data, then getting a unique count\n    \n    print(\"\\t\\tCollected all Nunique...\")\n    \n    X = pd.read_parquet(\"XTemp.parq\")\n    \n    X = pd.merge(X, df, how= \"left\", on= \"attacker_ip_enum\")    #Combining the calculated data from df with the original dataframe, X\n    X = X.rename(columns= {\"watcher_as_num_x\": \"watcher_as_num\", \"watcher_as_num_y\": \"WatPerAttIP\"})    #Renaming the columns changed by the merge above\n    X.drop(\"index\", axis= 1, inplace= True)    #Dropping the \"index\" column created during the groupby\n    df = None    #Clearing for mem\n    return X\n    \ndef CountAvgUniqueAttType(X, window):\n    print(\"Counting, and then averaging, the number\")\n    #An average count of unique attack_types used by each attacker_ip_enum in a given window (8 minutes)\n    smallWindow = \"8min\"    #Setting small window of 8 minutes\n    X.reset_index(inplace= True)    #Resetting index incase something messes with it when merging\n    df = X[[\"attack_type\", \"attacker_ip_enum\", \"attack_time\"]].copy(deep= True) #Making a copy so changes aren't made on the original dataframe\n    \n    X.to_parquet(\"XTemp.parq\")\n    X = None\n    gc.collect()\n    \n    df[\"attack_time\"] = df[\"attack_time\"].dt.floor(smallWindow) #Changing the time to a rounded floor based on the small window variable\n    #Grouping by attacker_ip_enum and the floored time range to find the number of unique attack_type's used. Basically the count of unique attack_type's used by each attacker_ip_enum in each 8 minute window\n    df = df.groupby([\"attacker_ip_enum\", pd.Grouper(key=\"attack_time\", freq= smallWindow)])[\"attack_type\"].nunique().reset_index(name= f\"NUniqueAttackType_{smallWindow}\")\n    df.drop([\"attack_time\"], axis= 1, inplace= True)    #Dropping the, now unnecessary, rounded attack_time column from the dataframe being worked on\n    df = df.groupby([\"attacker_ip_enum\"])[f\"NUniqueAttackType_{smallWindow}\"].mean().reset_index()  #Grouping by attacker_ip_enum and averaging every 8 minute window where they attacked\n    \n    X = pd.read_parquet(\"XTemp.parq\")\n    \n    X = pd.merge(X, df, left_on=[\"attacker_ip_enum\"], right_on=[\"attacker_ip_enum\"])    #Merging the worked on dataframe with the original\n    X.set_index(\"index\", inplace= True) #As expected the index was adjusted, but it can re-added to it's orignal state with this line\n    X[\"NUniqueAttackType_8min\"] = X[\"NUniqueAttackType_8min\"].astype(\"float32\")\n    df = None   #Of course, clearing the worked on dataframe to save some memory\n    return X","metadata":{"execution":{"iopub.status.busy":"2024-01-27T16:18:11.149845Z","iopub.execute_input":"2024-01-27T16:18:11.150855Z","iopub.status.idle":"2024-01-27T16:18:11.197910Z","shell.execute_reply.started":"2024-01-27T16:18:11.150815Z","shell.execute_reply":"2024-01-27T16:18:11.196955Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"FEAll Data Function\nThis simply calls all the functions above (including shodanHashedAlt)","metadata":{}},{"cell_type":"code","source":"#Feature Engineering the main dataset\ndef FEAllData(X, uniqueList = None):\n    #Dropping unused features\n    \n    #https://www.kaggle.com/code/pavansanagapati/14-simple-tips-to-save-ram-memory-for-1-gb-dataset  \n    \n    X = X.drop([\"watcher_as_name\", \"attacker_as_name\", \"watcher_uuid_enum\", \"watcher_country\", \"attacker_country\"], axis=1)\n    window = \"8D\"   #Setting window frame for time measurements\n\n##########################################################################################################################################################################################################\n\n    #Performing data engineering with shodan_df_hashed throught the shodanHashedAlt function\n    print(\"Starting shodanHashed...\")\n    if(uniqueList is None):    #This means shodanHashedAlt hasn't run yet because it's the first time this function (FEAllData) is being used (since uniqueList is filled in when FEAllData is run once)\n        X.to_parquet(\"XTemp.parq\")   #Storing the dataframe in the harddrive to save on RAM\n        \n        shodanCSV = pd.read_csv(\"/kaggle/input/vpn-classification/dataset_v2/shodan_df_hashed.csv\")\n        shodanCSV = shodanHashedAlt(shodanCSV)    #Gathering the features from shodan_df_hashed\n\n        X = pd.read_parquet(\"XTemp.parq\")    #Re-reading in the main dataframe\n    \n    else:    #Meaning shodanCSV has already had feature engineering performed on it\n        shodanCSV = pd.read_parquet(\"shodanHashedTemp.parq\")\n    \n    #Combining ShodanCSV DF and the main features DF based on attacker_ip_enum\n    X = X.reset_index().merge(shodanCSV, how = \"left\", on=\"attacker_ip_enum\").set_index(\"index\")\n    shodanCSV.to_parquet(\"shodanHashedTemp.parq\")    #Storing shodanCSV to a parq for future use/ reference\n    shodanCSV = None    #Saving memory\n    gc.collect()\n\n##########################################################################################################################################################################################################\n    \n    #Counting, and then averaging, the number of unique attack_type's used by attacker_ip_enum's in a given (smaller) window\n    X = CountAvgUniqueAttType(X, window)\n    gc.collect()\n    \n##########################################################################################################################################################################################################    \n\n    #Finding average attacker IP's shared among attacker_as_num's used by each attacker_ip_enum\n    X = FindingAverageAttackerIPs(X, window)\n    gc.collect()\n\n##########################################################################################################################################################################################################\n    \n    #Finding a count of unique watcher_as_num+attack_type combinations in a given time frame (as was suggested by the article posted by CrowdSec in the community tab (but slightly adjusted))\n    X = FindingCountUniqueWatAttCom(X, window)\n    gc.collect()\n    \n##########################################################################################################################################################################################################\n\n    #Adding a count of attacks carried out, with the same IP, in a designated window\n    X = AddingCountAtt(X, window)\n    gc.collect()\n    \n##########################################################################################################################################################################################################\n\n    #Finding the time diff between the attack and the one before\n    X = FindingTimeDiff(X, window)\n    gc.collect()\n    \n##########################################################################################################################################################################################################\n    \n    #Just the number of watchers each attacker_ip_enum has attacked in the whole dataset\n    X = AttIPNumWat(X, window)\n    gc.collect()\n    \n##########################################################################################################################################################################################################\n    \n    #Number of specific types of attacks in a time range\n    X, uniqueList = NumberSpecificAtt(X, window, uniqueList)\n    gc.collect()\n    \n##########################################################################################################################################################################################################\n    \n    #Dropping (now) unnecessary columns\n    X = X.drop([\"attack_time\", \"watcher_as_num\", \"attacker_as_num\"], axis= 1)\n#https://stackoverflow.com/questions/52693482/merging-pandas-data-frames-uses-way-too-much-memory\n    return X, uniqueList","metadata":{"execution":{"iopub.status.busy":"2024-01-27T16:18:11.199592Z","iopub.execute_input":"2024-01-27T16:18:11.200466Z","iopub.status.idle":"2024-01-27T16:18:11.217203Z","shell.execute_reply.started":"2024-01-27T16:18:11.200390Z","shell.execute_reply":"2024-01-27T16:18:11.215527Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"One Hot Encoding Function\nUsed to transform the Ordinally Encoded attack_type column into multiple binary columns. I did this to avoid implied ranking from the Ordinal Encoding previously performed on the attack_type column. I also drop columns that, through experimentation, catboost doesn't consider useful to save memory and (in some cases) increase the score.","metadata":{}},{"cell_type":"code","source":"#One Hot Encoding Protocol and Attack Type Together, done in this way so rank isn't assumed in the OrdinalEncoded attack_type column\ndef OneHotEncoding(X, OHEncoder = None):\n    print(\"Starting One Hot Encoding...\")\n    toOHE = [\"attack_type\"]   #Specifying what column to encode\n    if(OHEncoder == None):\n        OHEncoder = OneHotEncoder(handle_unknown= \"ignore\", sparse_output= False)    #Setting one hot encoder\n        OHEncoder.fit(X[toOHE])    #Fitting the encoder on the training data\n    \n    OHXT = pd.DataFrame(OHEncoder.transform(X[toOHE]))    #Encoding the column from the training data\n    OHXT = OHXT.astype(\"bool\")      #Ensuring the columns are bool types (taking up less memory than int64 or string)\n    X = X.drop(toOHE, axis= 1)   #Dropping the one hot encoded column from the training data\n    OHXT.index = X.index    #Ensuring the indeces between the two dataframes are the same\n    X = pd.concat([X, OHXT], axis= 1)   #Moving the data from the one hot encoding dataframe to the original\n\n    #Setting the column names as strings (or, rather, ensuring they are) (as is required for some models during testing)\n    X.columns = X.columns.astype(str)\n    #Dropping Columns\n    print(\"Dropping unnecessary columns...\")\n\n    toDrop = [\"14\", \"9\", \"12\", \"8\", \"0\", \"1\"]    #Creating list of columns to drop\n    X.drop(toDrop, axis= 1, inplace= True)\n    \n    return X, OHEncoder","metadata":{"execution":{"iopub.status.busy":"2024-01-27T16:18:11.220667Z","iopub.execute_input":"2024-01-27T16:18:11.221132Z","iopub.status.idle":"2024-01-27T16:18:11.234353Z","shell.execute_reply.started":"2024-01-27T16:18:11.221096Z","shell.execute_reply":"2024-01-27T16:18:11.233065Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Ordinal Encoding Function\nI used ordinal encoding on attack_type only as it was the sole feature missing a numerical representation. It came in handy for functions like FindingCountUniqueAttWat because it could be added as a decimal place, assisting in the creation of unique ID's using two numerical features. The downside is, as Kaggle's tutorials point out, so models consider this a ranking. To alleviate this, the final step before averaging all the data, is transforming this column into many with a One Hot Encoder.","metadata":{}},{"cell_type":"code","source":"def OrdinalEncoding(X, OrdEncode = None):\n    print(\"Starting Ordinal Encoding...\")   #Fit on attack_type so the values between each are smaller than strings and can be synced between XTr/ XTe\n    if(OrdEncode == None):  #Assuming OrdEncode isn't being passed into the argument (XTr) it will be created. If it is being passed in (XTe) then it doesn't need to be created and is already fit to XTr\n        OrdEncode = OrdinalEncoder()    #Creating the OrdinalEncoder\n        OrdEncode.fit(X[[\"attack_type\"]])   #Fitting it to attack_type alone\n    X[\"attack_type\"] = OrdEncode.transform(X[[\"attack_type\"]]).astype(int)  #Transforming the attack_type column to the fit OrdinalEncoder\n    X[\"attack_type\"] = X[\"attack_type\"].astype(\"int32\") #Ensuring the new attack_type column is int32 to avoid unnecessary memory usage\n    return X, OrdEncode","metadata":{"execution":{"iopub.status.busy":"2024-01-27T16:18:11.236022Z","iopub.execute_input":"2024-01-27T16:18:11.237089Z","iopub.status.idle":"2024-01-27T16:18:11.250313Z","shell.execute_reply.started":"2024-01-27T16:18:11.237040Z","shell.execute_reply":"2024-01-27T16:18:11.248900Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Averaging and Slimming Function","metadata":{}},{"cell_type":"code","source":"def averageSlimming(X, y = None):\n    #Averaging/ Counting everything (making each attacker_ip_enum unique to the attacker_ip_enum column, i.e. no duplicates (easier on memory and how the submission was requested (i.e. not classifying every instance, just every IP)))\n    print(\"Finding Averages...\")\n    toAverage = list(X.columns)    #Creating a list of all columns present in the X dataframe\n    stayingSame = [\"attacker_ip_enum\", \"TCPCount\", \"UDPCount\", \"jarm_sum\", \"ja3s_sum\", \"headers_hash_sum\", \"NoPrintWithAttack\"]    #Specifying which columns should not be averaged\n    toAverage = [column for column in toAverage if column not in stayingSame]   #Creating a list of columns to be averaged based on all columns that are present in the dataframe but not in the stayingSame list\n    colDict = dict.fromkeys(stayingSame, \"first\")    #Creating a dictionary indicating that all columns in stayingSame should only take the first value present...\n    colDict.update(dict.fromkeys(toAverage, \"mean\")) #While those in the toAverage column will have the mean taken and saved\n\n    X = X.groupby(\"attacker_ip_enum\", as_index= False).agg(colDict)    #Grouping by attacker_ip_enum, then keeping the first of the stayingSame columns and the average of the toAverage columns\n    X.sort_values(by= [\"attacker_ip_enum\"], inplace= True)    #Sorting by attacker_ip_enum to ensure consistency between indexes\n\n    if(y is not None):    #If y came into the function as a dataframe (meaning it's XTr's pair)\n        y.drop_duplicates(subset= \"attacker_ip_enum\", inplace= True, keep= \"first\")    #Dropping duplicate values in the y training dataframe\n        y.sort_values(by= [\"attacker_ip_enum\"], inplace= True)    #Sorting by attacker_ip_enum to ensure consistency between indexes (X and y, similar to above)\n        y.drop([\"attacker_ip_enum\"], axis= 1, inplace= True)    #Dropping attacker_ip_enum so it doesn't interfere with the models fit\n\n    else:    #If no y was input into the dataframe (i.e. XTe is being used and a predictions dataframe needs to be created)\n        y = X[\"attacker_ip_enum\"].copy(deep= True)    #The yTe predictions dataframe is created\n    \n    X.drop([\"attacker_ip_enum\"], axis= 1, inplace= True)    #attacker_ip_enum is not useful for the model so it's dropped\n    \n    print(X.shape)\n    print(y.shape)\n    \n    return X, y   #Returning X and y dataframes (XTr and yTr (for model fitting) or XTe and yTe (for model predicting))\n    ","metadata":{"execution":{"iopub.status.busy":"2024-01-27T16:18:11.252329Z","iopub.execute_input":"2024-01-27T16:18:11.252777Z","iopub.status.idle":"2024-01-27T16:18:11.267441Z","shell.execute_reply.started":"2024-01-27T16:18:11.252740Z","shell.execute_reply":"2024-01-27T16:18:11.265844Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Reading in Training Data and calling Feature Engineering functions on it","metadata":{}},{"cell_type":"code","source":"#Could try changing the order so datetime column is dropped near the start (equivalent size as float64)\n\n#Reading in the data\nXTr = pd.read_parquet(\"/kaggle/input/vpn-classification/dataset_v2/train.parq\")\ngc.enable()\n#Creating XTrain, yTrain, etc. as well as calling feature engineering function\nyTr = XTr[[\"attacker_ip_enum\", \"label\"]].copy(deep= True)    #Keeping attacker_ip_enum in there so when I average XTr later the same adjustment can be made to yTr. This ensure the indeces aren't out of order\nXTr.drop([\"label\"], axis= 1, inplace= True)    #Dropping \"label\" from XTr\n\nXTr, OrdEnc = OrdinalEncoding(XTr)  #Ordinal Encoding attack_type\ngc.collect()\n\nXTr, uniqueList = FEAllData(XTr)    #Performing feature engineering on XTr and creating a unique list\ngc.collect()\nXTr, OHEnc = OneHotEncoding(XTr)    #OneHotEncoding attack_type\ngc.collect()\n\nprint(XTr.info())\n\nXTr, yTr = averageSlimming(XTr, yTr)    #Averaging data based on attack_ip_enum and slimming it down to one instance of each attacker_ip_enum per dataframe","metadata":{"execution":{"iopub.status.busy":"2024-01-27T16:18:11.271075Z","iopub.execute_input":"2024-01-27T16:18:11.271618Z","iopub.status.idle":"2024-01-27T16:42:13.780264Z","shell.execute_reply.started":"2024-01-27T16:18:11.271569Z","shell.execute_reply":"2024-01-27T16:42:13.778994Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Fitting the model and reading in/ Feature Engineering the testing data\nDisplaying feature importance was valuable when cutting out data to meet memory requirements (even before attempting to fit into Kaggle's standards). I kept the columns out as catboost determined them unnecessary (with 0 or extremely close 0 value).","metadata":{}},{"cell_type":"code","source":"#Fitting the model and displaying feature importance (as chosen by CatBoost)\nprint(\"Fitting the model...\")\nmodel = CatBoostClassifier(depth=8, early_stopping_rounds= 5, learning_rate= 0.05, subsample= 0.70, colsample_bylevel= 0.65, min_data_in_leaf= 50, scale_pos_weight= 4.0, silent= True)\nprint(\"\\t Fitting Now\")\nmodel.fit(XTr,yTr)\nprint(\"\\tFinished Fitting\")\n\n#Determing and displaying feature importance (Catboost only)\nimportances = model.get_feature_importance(type= \"PredictionValuesChange\")\nfeature_importances = pd.Series(importances, index= XTr.columns).sort_values()\nplt.barh(feature_importances.index, feature_importances.values)\nplt.title(f\"CatBoost: Tr= {len(XTr)}\")\nplt.xlabel(\"Importance\")\nplt.ylabel(\"Features\")\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2024-01-27T16:42:13.781791Z","iopub.execute_input":"2024-01-27T16:42:13.782155Z","iopub.status.idle":"2024-01-27T16:43:13.462775Z","shell.execute_reply.started":"2024-01-27T16:42:13.782125Z","shell.execute_reply":"2024-01-27T16:43:13.461508Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"XTe = pd.read_parquet(\"/kaggle/input/vpn-classification/dataset_v2/test.parq\")\n#Same process as that seen with XTr BUT with the fit for Ordinal Encoder, One Hot Encoder, and the uniqueList used in NumberSpecificAtt function being present and passed into the functions\nXTe, OrdEnc = OrdinalEncoding(XTe, OrdEnc)\nOrdEnc = None\ngc.collect()\n\nXTe, uniqueList = FEAllData(XTe, uniqueList)\nuniqueList = None\nShodanHashed = None\ngc.collect()\n\nXTe, OHEnc = OneHotEncoding(XTe, OHEnc)\nOHEnc = None\ngc.collect()\n\nprint(XTe.info)\n\nXTe, XTe_AttIPEnum = averageSlimming(XTe)","metadata":{"execution":{"iopub.status.busy":"2024-01-27T16:43:13.464610Z","iopub.execute_input":"2024-01-27T16:43:13.465078Z","iopub.status.idle":"2024-01-27T16:49:19.596273Z","shell.execute_reply.started":"2024-01-27T16:43:13.465032Z","shell.execute_reply":"2024-01-27T16:49:19.594998Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Writing out predictions","metadata":{}},{"cell_type":"code","source":"#Predicting yTe and saving it to a CSV\npredicts = pd.DataFrame(model.predict(XTe.values))\n\nsubmission = pd.concat([XTe_AttIPEnum, predicts], axis= 1, join= \"inner\")    #Combining the attacker_ip_enum only dataframe (with no duplicate IPs) with the prediction made above\nsubmission.columns = [\"attacker_ip_enum\", \"label\"]    #Ensuring the columns are properly labeled\nprint(submission.info)\n#submission.sort_values(by= [\"attacker_ip_enum\"], axis= 1, inplace= True)\n\nsubmission.to_csv(\"submission.csv\", index= False)\nprint(XTe.head(5))\nprint(\"Number of each: \", submission.value_counts(\"label\"))    #To get an idea of how many are positive/ negative (assuming ~3-5% should be positive)","metadata":{"execution":{"iopub.status.busy":"2024-01-27T16:49:19.598050Z","iopub.execute_input":"2024-01-27T16:49:19.598939Z","iopub.status.idle":"2024-01-27T16:49:19.812792Z","shell.execute_reply.started":"2024-01-27T16:49:19.598893Z","shell.execute_reply":"2024-01-27T16:49:19.811468Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Due to the stochasticity of col_sample and sub_sample in Catboost results may vary. They generally come out to around the same f1 score with positive/ negative predictions fluctuating by +/- 20","metadata":{}}]}