{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"code","source":"# This Python 3 environment comes with many helpful analytics libraries installed\n# It is defined by the kaggle/python Docker image: https://github.com/kaggle/docker-python\n# For example, here's several helpful packages to load\n\nimport numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\n\n# Input data files are available in the read-only \"../input/\" directory\n# For example, running this (by clicking run or pressing Shift+Enter) will list all files under the input directory\n\nimport os\nfor dirname, _, filenames in os.walk('/kaggle/input'):\n    for filename in filenames:\n        print(os.path.join(dirname, filename))\n\n# You can write up to 20GB to the current directory (/kaggle/working/) that gets preserved as output when you create a version using \"Save & Run All\" \n# You can also write temporary files to /kaggle/temp/, but they won't be saved outside of the current session","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","execution":{"iopub.status.busy":"2023-10-01T05:30:39.665443Z","iopub.execute_input":"2023-10-01T05:30:39.666097Z","iopub.status.idle":"2023-10-01T05:30:40.162243Z","shell.execute_reply.started":"2023-10-01T05:30:39.666049Z","shell.execute_reply":"2023-10-01T05:30:40.161128Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import ast\n\nimport matplotlib.pyplot as plt\nimport numpy as np\nimport pandas as pd\nimport seaborn as sns","metadata":{"execution":{"iopub.status.busy":"2023-10-01T05:31:24.167769Z","iopub.execute_input":"2023-10-01T05:31:24.168385Z","iopub.status.idle":"2023-10-01T05:31:25.017761Z","shell.execute_reply.started":"2023-10-01T05:31:24.168348Z","shell.execute_reply":"2023-10-01T05:31:25.016463Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Introduction\n\n## Goal\n\nThe goal of this competition is to identify users of annoymization services (proxies, VPNs) by analyzing the abuse patterns coming from their associated ips.\n\n\n## Dataset\n\nThe dataset included in this competition is a slice of signals we receive from crowdsec security engines installed all over the world. Each line is a security engine detecting an attack using one of our many scenarios. Each line contains three different sections of data. \n\n* Information about the security engine (called watcher) reporting the attack\n* Information about the ip (called attacker) doing the attack\n* Information about the attack itself\n\nData about the security engine can be found in the columns `watcher_country, watcher_as_num, watcher_as_name, watcher_uuid_enum`  \nData about the attacker can be found in the columns `attacker_country, attacker_as_num, attacker_as_name, attacker_ip_enum`  \nData about the attack can be found in the columns `attack_time, attack_type`\n\nAdditionally, you can find the results of portscans and respective service hashes in `shodan_info_hashed.csv`. Note that this dataset contains information about attacker IPs in both train and test set.","metadata":{}},{"cell_type":"markdown","source":"### CSV format","metadata":{}},{"cell_type":"code","source":"#dtypes = {\n#    # 'attack_time': 'datetime64[ns]',\n#    \"watcher_country\": \"object\",\n#    \"watcher_as_num\": \"float32\",\n#    \"watcher_as_name\": \"object\",\n#    \"attacker_country\": \"category\",\n#    \"attacker_as_num\": \"float32\",\n#    \"attacker_as_name\": \"object\",\n#    \"attack_type\": \"category\",\n#    \"watcher_uuid_enum\": \"int32\",\n#    \"attacker_ip_enum\": \"int32\",\n#    \"label\": \"int8\",\n#}\n\n#df = pd.read_csv(\"/kaggle/input/vpn-classification/dataset_v2/train.csv\", dtype=dtypes, parse_dates=[\"attack_time\"])\n#df.dtypes\n#df.head()","metadata":{"execution":{"iopub.status.busy":"2023-10-01T05:32:31.907740Z","iopub.execute_input":"2023-10-01T05:32:31.908893Z","iopub.status.idle":"2023-10-01T05:35:21.308070Z","shell.execute_reply.started":"2023-10-01T05:32:31.908846Z","shell.execute_reply":"2023-10-01T05:35:21.306070Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Parquet format\n\nAlternatively you can use parquet format, which loads significantly faster.","metadata":{}},{"cell_type":"code","source":"df = pd.read_parquet(\"/kaggle/input/vpn-classification/dataset_v2/train.parq\")\ndf.dtypes\ndf.head()","metadata":{"execution":{"iopub.status.busy":"2023-10-01T05:35:24.173033Z","iopub.execute_input":"2023-10-01T05:35:24.173429Z","iopub.status.idle":"2023-10-01T05:35:43.288930Z","shell.execute_reply.started":"2023-10-01T05:35:24.173399Z","shell.execute_reply":"2023-10-01T05:35:43.287478Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Useful concepts\n\nThis section serves to explain some useful concepts to people unfamilliar with networks and security topics. \n\n## AS (Autonomous systems)\n\nThe dataset contains AS information for both the attacker and the security engine. An autonomous system is a top-level owner of an IP address. For end users it is usually the internet service provider. If you host an instance somewhere in the cloud, the IP is usually owned by the hoster itself (DIGITALOCEAN, AMAZON AWS, GOOGLE CLOUD, HETZNER etc.). Every AS can be identified by a unique number, which then can be used to find more information about it on internet. A convenient website such as IP Info easily provide complementary information of a given AS based on its number: [AS1406](https://ipinfo.io/AS1406)\n\nIPs belonging to the same AS can present the same specificities in terms of attack types and patterns. Some AS are also more likely to host VPN IPs than others. \nThe ASN could for instance be used to compute trust scores for watchers or attackers as not all hosters are equally vigilant with cleaning up malicious actors.\n\n## Attack Description\n\nThe `attack_type` column contains information about the attack that was detected, e.g. `ssh:bruteforce`. It is split into two parts of information: the attack service (http, ssh, windows etc.) and the attack type (bruteforce, exploit, spam etc.). You may find the full list in the [Behavior section of CrowdSec CTI Documentation](https://docs.crowdsec.net/docs/next/cti_api/taxonomy/#behaviors)\n\nThe attack service is mostly self-explanatory with a lot of common known protocols. `http` is for web based attacks, `ssh` for secure shell (remote login for linux basically) `windows` is the umbrella category for all attacks that are detected in windows logs, `tcp` is for network attacks that operate below the `http` layer. There are a few rare ones, `pop3/imap` collects attacks targeting mail servers and `sip` collects attacks targeting VoIP (internet telephone).  \n\nThe attack type covers the basic [MITRE](https://attack.mitre.org/) classifications of cyber attacks. `bruteforce` covers attacks with work through enumeration (trying passwords at random for instance), `exploit` covers attacks that abuse specific vulnerabilities in software like [log4shell](https://en.wikipedia.org/wiki/Log4Shell), `spam` covers spamming attacks, usually directed at some form (input field) on a website, `scan` covers attacks that try to detected services and open ports, usually as reconnaissance for more targeted attacks and `crawl` is refers to attacks the try to generate sitemaps and the like on the http layer. \n\nThe scenario information can for instance be used as a feature directly to separate different types of annonymization services. Importantly, proxies can only be used at the HTTP layer, so a proxy IP is not supposed to be reported for SSH attacks.\nAnother way of using the attack type can be to group watchers by the kind of services they expose. \n\n## Shodan info\n\nThere is an additional dataset that comes with our trainset called `shodan_info_hashed`. The `shodan_info` column contains data we enrich from the third-party threat intelligence provider [shodan.io](https://www.shodan.io/). The column contains a dictionary that contains the open ports of a given attacker and the protocol these ports support.  \nFor instance take the following example: \n```\n{'443/tcp': {'headers_hash': -1035061394,\n  'jarm': '2ad2ad16d00000022c2ad2ad2ad2adbc07bc352a188303b518c8a273c71220',\n  'ja3s': '986571066668055ae9481cb84fda634a'}}\n```\nThe key `443/tcp` tells us that this attacker has port `443` open and accessible using `tcp`. For most ports there are conventions on what services are usually found on it. [Wikipedia](https://en.wikipedia.org/wiki/List_of_TCP_and_UDP_port_numbers) has a good list. If the port in question supports `http` we include a `headers_hash` which is a somewhat unique fingerprint. For connections that support `ssl` we include the `jarm` and `ja3s` which both offer ways to fingerprint specific device configurations, and can certainly be useful for the challenge.\n","metadata":{}},{"cell_type":"markdown","source":"## Viz\n\nWe plot the distribution of attack types and the label. The dataset is highly imbalanced","metadata":{}},{"cell_type":"code","source":"_ = plt.figure(figsize=(14, 7))\n_ = sns.countplot(x=df[\"attack_type\"])\n_ = plt.xticks(rotation=90)\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-10-01T05:35:43.292513Z","iopub.execute_input":"2023-10-01T05:35:43.293785Z","iopub.status.idle":"2023-10-01T05:35:46.371369Z","shell.execute_reply.started":"2023-10-01T05:35:43.293737Z","shell.execute_reply":"2023-10-01T05:35:46.369912Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"_ = plt.figure(figsize=(14, 7))\n# Counting the number of labels for each IP (hence the groupby)\n_ = sns.countplot(x=df.groupby(\"attacker_ip_enum\")[\"label\"].max())\n_ = plt.xticks(rotation=90)\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-10-01T05:35:46.373181Z","iopub.execute_input":"2023-10-01T05:35:46.373903Z","iopub.status.idle":"2023-10-01T05:35:47.925875Z","shell.execute_reply.started":"2023-10-01T05:35:46.373867Z","shell.execute_reply":"2023-10-01T05:35:47.924508Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Basic Model\n\nWe train a very simple RandomForestClassifier on the dataset using a standard train-test split to verify the results. \n\nAs the prediction is on a per-attacker basis, we build features using pandas groupby and aggregation operations.\n\n## Feature Engineering\n\nAs features we use:\n\n* a one-hot encoding of the attack service\n* a one-hot encoding of the attack type (normalized)\n* a one-hot encoding of some of the most common ports\n* total number of open ports","metadata":{}},{"cell_type":"markdown","source":"### One-hot attack types\nWe first split the `attack_type` into `service` and `type` and use `pd.get_dummies()`","metadata":{}},{"cell_type":"code","source":"attack_types_df = (\n    df.attack_type.str.split(\":\", expand=True)\n    .rename(columns={0: \"service\", 1: \"type\"})\n    .set_index(df[\"attacker_ip_enum\"])\n)","metadata":{"execution":{"iopub.status.busy":"2023-10-01T05:35:53.187880Z","iopub.execute_input":"2023-10-01T05:35:53.188320Z","iopub.status.idle":"2023-10-01T05:36:32.693078Z","shell.execute_reply.started":"2023-10-01T05:35:53.188288Z","shell.execute_reply":"2023-10-01T05:36:32.691444Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"one_hot_attack_service_df = pd.get_dummies(\n    # Dropping duplicated service before calling get dummies\n    attack_types_df.reset_index()\n    .drop_duplicates(subset=[\"attacker_ip_enum\", \"service\"])\n    .set_index(\"attacker_ip_enum\")[\"service\"]\n    # ,sparse=True\n)\none_hot_attack_service_df = one_hot_attack_service_df.groupby(\"attacker_ip_enum\").sum()\n# We group by ip_enum and keep only strictly positive values\none_hot_attack_service_df = (one_hot_attack_service_df >= 1).astype(int)\none_hot_attack_service_df.head()","metadata":{"execution":{"iopub.status.busy":"2023-10-01T05:37:01.572749Z","iopub.execute_input":"2023-10-01T05:37:01.573193Z","iopub.status.idle":"2023-10-01T05:37:10.917706Z","shell.execute_reply.started":"2023-10-01T05:37:01.573159Z","shell.execute_reply":"2023-10-01T05:37:10.916289Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"one_hot_attack_types_df = pd.get_dummies(\n    attack_types_df[\"type\"],\n    # sparse=True\n)\none_hot_attack_types_df = one_hot_attack_types_df.groupby(\"attacker_ip_enum\").sum()\n# We group by ip_enum and normalized by the number of attack to get a distribyution\none_hot_attack_types_df = one_hot_attack_types_df / one_hot_attack_types_df.sum(\n    1\n).values.reshape(-1, 1)\none_hot_attack_types_df.head(10)","metadata":{"execution":{"iopub.status.busy":"2023-10-01T05:36:41.804875Z","iopub.execute_input":"2023-10-01T05:36:41.805205Z","iopub.status.idle":"2023-10-01T05:36:54.370291Z","shell.execute_reply.started":"2023-10-01T05:36:41.805176Z","shell.execute_reply":"2023-10-01T05:36:54.368935Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Open ports Feature\n\nAs the port information is collected on a per-attacker basis we only have one row per attacker IP, so we don't need to use `pd.groupby()` this time.","metadata":{}},{"cell_type":"code","source":"REFERENCE_PORTS = set((\"22\", \"80\", \"443\", \"7777\"))\n\n\ndef extract_port_numbers(value: dict):\n    return set(key.split(\"/\")[0] for key in value.keys())\n\n\nshodan_df = pd.read_csv(\n    \"/kaggle/input/vpn-classification/dataset_v2/shodan_df_hashed.csv\",\n    dtype={\"attacker_ip_enum\": \"int32\"},\n    index_col=\"attacker_ip_enum\",\n)\nshodan_df[\"shodan_info\"] = shodan_df[\"shodan_info\"].map(ast.literal_eval)\nshodan_df[\"shodan_open_ports\"] = shodan_df[\"shodan_info\"].map(extract_port_numbers)","metadata":{"execution":{"iopub.status.busy":"2023-10-01T05:38:06.502414Z","iopub.execute_input":"2023-10-01T05:38:06.502920Z","iopub.status.idle":"2023-10-01T05:38:14.512448Z","shell.execute_reply.started":"2023-10-01T05:38:06.502881Z","shell.execute_reply":"2023-10-01T05:38:14.510947Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"open_ports_count = shodan_df[\"shodan_open_ports\"].map(len).rename(\"open_ports_count\")\n# We select the reference port by doing a set intersection\nreference_ports_df = shodan_df[\"shodan_open_ports\"].map(lambda x: REFERENCE_PORTS & x)\n\none_hot_reference_ports_df = pd.get_dummies(reference_ports_df.explode(), prefix=\"port\")\n# Aggregate the dummies by IP\none_hot_reference_ports_df = one_hot_reference_ports_df.groupby(\n    \"attacker_ip_enum\"\n).sum()\n# Final feature for port: count and one hot table of reference port\nports_features_df = pd.concat([one_hot_reference_ports_df, open_ports_count], axis=1)\n# Checking results\nports_features_df[(ports_features_df > 0).any(axis=1)].head()","metadata":{"execution":{"iopub.status.busy":"2023-10-01T05:38:22.589838Z","iopub.execute_input":"2023-10-01T05:38:22.590328Z","iopub.status.idle":"2023-10-01T05:38:23.584545Z","shell.execute_reply.started":"2023-10-01T05:38:22.590290Z","shell.execute_reply":"2023-10-01T05:38:23.583159Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Feature Aggregation\n\nWe merge our engineered features together with the labels from the `label` column to get a tabular dataset with can start training on. Note that we use `join=\"inner\"` because  `ports_df` also includes the attackers in the test set (leaderboard)","metadata":{}},{"cell_type":"code","source":"# We create a df label df which contains one row per attacker IP\nlabel_df = df.drop_duplicates([\"attacker_ip_enum\", \"label\"]).set_index(\n    \"attacker_ip_enum\"\n)[\"label\"]\n\ndataset = pd.concat(\n    [\n        one_hot_attack_types_df,\n        one_hot_attack_service_df,\n        ports_features_df,\n        label_df,\n    ],\n    axis=1,\n    join=\"inner\",\n)","metadata":{"execution":{"iopub.status.busy":"2023-10-01T05:38:51.035955Z","iopub.execute_input":"2023-10-01T05:38:51.036369Z","iopub.status.idle":"2023-10-01T05:38:53.621134Z","shell.execute_reply.started":"2023-10-01T05:38:51.036338Z","shell.execute_reply":"2023-10-01T05:38:53.620133Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Upsampling and Downsampling\n\nSomething needs to be done to address the label imbalance. There is a lot of literature on how to address labeling imbalance, so pick your poison.  \n\nIn the test set for the leaderboard this is addressed by the F1 scoring that we used. This is to ensure that we have proper data to predict and that the leaderboard cannot be gamed by models which simply predict 0 for every label. \n\nFor the demo we opted to use a RandomForest without any up or downsamplig, but we have included code to show an example of the two approaches just in case.\n","metadata":{}},{"cell_type":"code","source":"from imblearn.over_sampling import ADASYN\nfrom imblearn.under_sampling import RandomUnderSampler\nfrom sklearn.model_selection import train_test_split","metadata":{"execution":{"iopub.status.busy":"2023-10-01T05:40:38.229606Z","iopub.execute_input":"2023-10-01T05:40:38.230131Z","iopub.status.idle":"2023-10-01T05:40:39.369907Z","shell.execute_reply.started":"2023-10-01T05:40:38.230097Z","shell.execute_reply":"2023-10-01T05:40:39.368551Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train, test = train_test_split(dataset, test_size=0.3)","metadata":{"execution":{"iopub.status.busy":"2023-10-01T05:40:44.361160Z","iopub.execute_input":"2023-10-01T05:40:44.361974Z","iopub.status.idle":"2023-10-01T05:40:44.434062Z","shell.execute_reply.started":"2023-10-01T05:40:44.361934Z","shell.execute_reply":"2023-10-01T05:40:44.432489Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"X_train_raw = train.drop([\"label\"], axis=1)\ny_train_raw = train[\"label\"]\nX_test_raw = test.drop([\"label\"], axis=1)\ny_test_raw = test[\"label\"]","metadata":{"execution":{"iopub.status.busy":"2023-10-01T05:40:49.302548Z","iopub.execute_input":"2023-10-01T05:40:49.303380Z","iopub.status.idle":"2023-10-01T05:40:49.326878Z","shell.execute_reply.started":"2023-10-01T05:40:49.303336Z","shell.execute_reply":"2023-10-01T05:40:49.325625Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# under = RandomUnderSampler(replacement=False)\n# over = ADASYN(n_neighbors = 10, n_jobs=-1)","metadata":{"execution":{"iopub.status.busy":"2023-10-01T05:40:55.061625Z","iopub.execute_input":"2023-10-01T05:40:55.062094Z","iopub.status.idle":"2023-10-01T05:40:55.067503Z","shell.execute_reply.started":"2023-10-01T05:40:55.062061Z","shell.execute_reply":"2023-10-01T05:40:55.066233Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# X_train, y_train = over.fit_resample(X_train_raw, y_train_raw)\n# X_test, y_test = under.fit_resample(X_test_raw, y_test_raw)","metadata":{"execution":{"iopub.status.busy":"2023-10-01T05:41:00.167909Z","iopub.execute_input":"2023-10-01T05:41:00.168374Z","iopub.status.idle":"2023-10-01T05:41:00.173612Z","shell.execute_reply.started":"2023-10-01T05:41:00.168339Z","shell.execute_reply":"2023-10-01T05:41:00.172182Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"X_train, y_train = X_train_raw, y_train_raw\nX_test, y_test = X_test_raw, y_test_raw","metadata":{"execution":{"iopub.status.busy":"2023-10-01T05:41:06.022814Z","iopub.execute_input":"2023-10-01T05:41:06.023308Z","iopub.status.idle":"2023-10-01T05:41:06.029366Z","shell.execute_reply.started":"2023-10-01T05:41:06.023271Z","shell.execute_reply":"2023-10-01T05:41:06.027955Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Model training\n\nWe train a very simple random forest model and optimize the hyperparameters over a grid search.  \n\nDepending on the model used, you might have to do some standardization to get nice convergence. For decision tree based models this is not required as it the scale of the axes does not impact the discrete decision boundaries. We nonetheless add it as a demonstration.","metadata":{}},{"cell_type":"code","source":"from sklearn.preprocessing import StandardScaler","metadata":{"execution":{"iopub.status.busy":"2023-10-01T05:41:20.220621Z","iopub.execute_input":"2023-10-01T05:41:20.221116Z","iopub.status.idle":"2023-10-01T05:41:20.227488Z","shell.execute_reply.started":"2023-10-01T05:41:20.221081Z","shell.execute_reply":"2023-10-01T05:41:20.226153Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sc = StandardScaler()\nX_train = sc.fit_transform(X_train)\nX_test = sc.transform(X_test)","metadata":{"execution":{"iopub.status.busy":"2023-10-01T05:41:26.727975Z","iopub.execute_input":"2023-10-01T05:41:26.728419Z","iopub.status.idle":"2023-10-01T05:41:26.877069Z","shell.execute_reply.started":"2023-10-01T05:41:26.728387Z","shell.execute_reply":"2023-10-01T05:41:26.875528Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from sklearn.ensemble import RandomForestClassifier\nfrom sklearn.metrics import f1_score\nfrom sklearn.model_selection import GridSearchCV","metadata":{"execution":{"iopub.status.busy":"2023-10-01T05:41:33.871494Z","iopub.execute_input":"2023-10-01T05:41:33.873558Z","iopub.status.idle":"2023-10-01T05:41:33.881281Z","shell.execute_reply.started":"2023-10-01T05:41:33.873493Z","shell.execute_reply":"2023-10-01T05:41:33.878784Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"parameters = {\n    \"max_depth\": [None, 10, 20],\n    \"max_features\": [\"sqrt\", None],\n    \"class_weight\": [\"balanced\", None],\n}\nest = RandomForestClassifier()\nclf = GridSearchCV(est, parameters, scoring=\"f1\", n_jobs=-1)","metadata":{"execution":{"iopub.status.busy":"2023-10-01T05:41:38.648588Z","iopub.execute_input":"2023-10-01T05:41:38.649077Z","iopub.status.idle":"2023-10-01T05:41:38.657749Z","shell.execute_reply.started":"2023-10-01T05:41:38.649043Z","shell.execute_reply":"2023-10-01T05:41:38.655245Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"clf.fit(X_train, y_train)","metadata":{"execution":{"iopub.status.busy":"2023-10-01T05:41:44.006564Z","iopub.execute_input":"2023-10-01T05:41:44.007031Z","iopub.status.idle":"2023-10-01T05:49:03.808666Z","shell.execute_reply.started":"2023-10-01T05:41:44.006998Z","shell.execute_reply":"2023-10-01T05:49:03.806909Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"clf.best_params_","metadata":{"execution":{"iopub.status.busy":"2023-10-01T05:49:03.811367Z","iopub.execute_input":"2023-10-01T05:49:03.811787Z","iopub.status.idle":"2023-10-01T05:49:03.820789Z","shell.execute_reply.started":"2023-10-01T05:49:03.811746Z","shell.execute_reply":"2023-10-01T05:49:03.819408Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"pd.DataFrame(clf.cv_results_).sort_values(by=\"rank_test_score\").head(10)","metadata":{"execution":{"iopub.status.busy":"2023-10-01T05:49:03.822531Z","iopub.execute_input":"2023-10-01T05:49:03.823017Z","iopub.status.idle":"2023-10-01T05:49:03.870621Z","shell.execute_reply.started":"2023-10-01T05:49:03.822987Z","shell.execute_reply":"2023-10-01T05:49:03.869055Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"y_pred = clf.predict(X_test)","metadata":{"execution":{"iopub.status.busy":"2023-10-01T05:49:03.873264Z","iopub.execute_input":"2023-10-01T05:49:03.873584Z","iopub.status.idle":"2023-10-01T05:49:04.503438Z","shell.execute_reply.started":"2023-10-01T05:49:03.873558Z","shell.execute_reply":"2023-10-01T05:49:04.502160Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"f1_score(y_pred=y_pred, y_true=y_test)","metadata":{"execution":{"iopub.status.busy":"2023-10-01T05:49:04.505350Z","iopub.execute_input":"2023-10-01T05:49:04.505713Z","iopub.status.idle":"2023-10-01T05:49:04.531116Z","shell.execute_reply.started":"2023-10-01T05:49:04.505668Z","shell.execute_reply":"2023-10-01T05:49:04.529657Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"model = RandomForestClassifier(n_jobs=-1, **clf.best_params_)\nmodel.fit(sc.transform(dataset.drop([\"label\"], axis=1)), dataset[\"label\"])","metadata":{"execution":{"iopub.status.busy":"2023-10-01T05:49:04.533054Z","iopub.execute_input":"2023-10-01T05:49:04.533549Z","iopub.status.idle":"2023-10-01T05:49:11.155152Z","shell.execute_reply.started":"2023-10-01T05:49:04.533492Z","shell.execute_reply":"2023-10-01T05:49:11.153831Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Submitting predictions\n\nUsing the trained models we can now do predictions for the leaderboard.","metadata":{}},{"cell_type":"code","source":"df = pd.read_parquet(\"/kaggle/input/vpn-classification/dataset_v2/test.parq\")\ndf.head()","metadata":{"execution":{"iopub.status.busy":"2023-10-01T05:49:11.156617Z","iopub.execute_input":"2023-10-01T05:49:11.157006Z","iopub.status.idle":"2023-10-01T05:49:15.598327Z","shell.execute_reply.started":"2023-10-01T05:49:11.156978Z","shell.execute_reply":"2023-10-01T05:49:15.596801Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"attack_types_df = (\n    df.attack_type.str.split(\":\", expand=True)\n    .rename(columns={0: \"service\", 1: \"type\"})\n    .set_index(df[\"attacker_ip_enum\"])\n)","metadata":{"execution":{"iopub.status.busy":"2023-10-01T05:49:15.599830Z","iopub.execute_input":"2023-10-01T05:49:15.600457Z","iopub.status.idle":"2023-10-01T05:49:27.725937Z","shell.execute_reply.started":"2023-10-01T05:49:15.600403Z","shell.execute_reply":"2023-10-01T05:49:27.724563Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"one_hot_attack_service_df = pd.get_dummies(\n    # Dropping duplicated service before calling get dummies\n    attack_types_df.reset_index()\n    .drop_duplicates(subset=[\"attacker_ip_enum\", \"service\"])\n    .set_index(\"attacker_ip_enum\")[\"service\"]\n    # ,sparse=True\n)\none_hot_attack_service_df = one_hot_attack_service_df.groupby(\"attacker_ip_enum\").sum()\none_hot_attack_service_df = (one_hot_attack_service_df >= 1).astype(int)","metadata":{"execution":{"iopub.status.busy":"2023-10-01T05:49:29.845598Z","iopub.execute_input":"2023-10-01T05:49:29.846633Z","iopub.status.idle":"2023-10-01T05:49:31.674260Z","shell.execute_reply.started":"2023-10-01T05:49:29.846601Z","shell.execute_reply":"2023-10-01T05:49:31.672932Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"one_hot_attack_types_df = pd.get_dummies(\n    attack_types_df[\"type\"],\n    # sparse=True\n)\none_hot_attack_types_df = one_hot_attack_types_df.groupby(\"attacker_ip_enum\").sum()\n# We group by ip_enum and normalized by the number of attack to get a distribyution\none_hot_attack_types_df = one_hot_attack_types_df / one_hot_attack_types_df.sum(\n    1\n).values.reshape(-1, 1)","metadata":{"execution":{"iopub.status.busy":"2023-10-01T05:49:31.676024Z","iopub.execute_input":"2023-10-01T05:49:31.677211Z","iopub.status.idle":"2023-10-01T05:49:34.669387Z","shell.execute_reply.started":"2023-10-01T05:49:31.677169Z","shell.execute_reply":"2023-10-01T05:49:34.668203Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"X_val_df = pd.concat(\n    [\n                one_hot_attack_types_df,\n        one_hot_attack_service_df,\n        ports_features_df,\n    ],\n    axis=1,\n    join=\"inner\",\n)\nX_val = sc.transform(X_val_df)","metadata":{"execution":{"iopub.status.busy":"2023-10-01T05:49:34.671028Z","iopub.execute_input":"2023-10-01T05:49:34.672267Z","iopub.status.idle":"2023-10-01T05:49:34.717814Z","shell.execute_reply.started":"2023-10-01T05:49:34.672225Z","shell.execute_reply":"2023-10-01T05:49:34.716315Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"prediction = pd.Series(model.predict(X_val), index=X_val_df.index).rename(\"prediction\")","metadata":{"execution":{"iopub.status.busy":"2023-10-01T05:49:34.719155Z","iopub.execute_input":"2023-10-01T05:49:34.719497Z","iopub.status.idle":"2023-10-01T05:49:35.035665Z","shell.execute_reply.started":"2023-10-01T05:49:34.719468Z","shell.execute_reply":"2023-10-01T05:49:35.034638Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Now we can save the predition as usual","metadata":{}},{"cell_type":"code","source":"#prediction.to_csv(\"predictions.csv\")","metadata":{"execution":{"iopub.status.busy":"2023-10-01T05:49:35.036796Z","iopub.execute_input":"2023-10-01T05:49:35.037076Z","iopub.status.idle":"2023-10-01T05:49:35.042002Z","shell.execute_reply.started":"2023-10-01T05:49:35.037053Z","shell.execute_reply":"2023-10-01T05:49:35.040796Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]}]}