{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Introduction\nIn this notebook we will present how to load and analyse (WIP) the data in the format of a graph. Since the Internet is gigantic set of nodes made of servers connected to each others, this is a pretty convenient representation.\n\n## BiPartite Graph\n\nIn the first section we will work on a bipartite Graph made of (U, V, E) where \n* U: Is the set of node of type Watcher\n* V: Is the set of node of type Attacker\n* E: Is an edge connecting a Watcher to and Attacker\n\n## Construct\n\nWe will load the data and use a group by to compute all the edges, as well as their number of occurence to use it later as a weight. \n\nFor this purpose we will use the `polars` library, not because of the hype around it, but because it's way faster than pandas for this operation. Then we will load the data into a graph using `networkx` library, more for convenience, but you can look also at `cuGRAPH` (https://github.com/rapidsai/cugraph) library from rapidsAI which offers support for graph algo on GPU \n\n## Analysis [WIP]","metadata":{}},{"cell_type":"markdown","source":"# Load ","metadata":{}},{"cell_type":"code","source":"import ast\n\nimport matplotlib.pyplot as plt\nimport numpy as np\nimport polars as pl\nimport seaborn as sns\nimport networkx as nx\n\nsns.set()","metadata":{"execution":{"iopub.status.busy":"2023-10-05T16:10:47.810314Z","iopub.execute_input":"2023-10-05T16:10:47.810754Z","iopub.status.idle":"2023-10-05T16:10:47.908984Z","shell.execute_reply.started":"2023-10-05T16:10:47.810703Z","shell.execute_reply":"2023-10-05T16:10:47.907970Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"dtypes = {\n    # 'attack_time': 'datetime64[ns]',\n    \"watcher_country\": \"object\",\n    \"watcher_as_num\": \"float32\",\n    \"watcher_as_name\": \"object\",\n    \"attacker_country\": \"category\",\n    \"attacker_as_num\": \"float32\",\n    \"attacker_as_name\": \"object\",\n    \"attack_type\": \"category\",\n    \"watcher_uuid_enum\": \"int32\",\n    \"attacker_ip_enum\": \"int32\",\n    \"label\": \"int8\",\n}\ndf = pl.read_parquet(\"/kaggle/input/vpn-classification/dataset_v2/train.parq\")\n# df = pd.read_csv(\"data/train.csv\", dtype=dtypes, parse_dates=[\"attack_time\"])\ndf.dtypes\ndf.shape\ndf.head()","metadata":{"execution":{"iopub.status.busy":"2023-10-05T16:09:28.727534Z","iopub.execute_input":"2023-10-05T16:09:28.727911Z","iopub.status.idle":"2023-10-05T16:09:37.084986Z","shell.execute_reply.started":"2023-10-05T16:09:28.727881Z","shell.execute_reply":"2023-10-05T16:09:37.083901Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Graph","metadata":{}},{"cell_type":"code","source":"edge_cols = [\"watcher_uuid_enum\",\"attacker_ip_enum\"]\n\n# edges_df = df.unique(subset=[\"watcher_uuid_enum\",\"attacker_ip_enum\"])\nunique_watchers = list(sorted(df.select(pl.col(edge_cols[0])).unique().to_numpy().flatten()))\nunique_attackers = list(sorted(df.select(pl.col(edge_cols[1])).unique().to_numpy().flatten()))\nlen(unique_attackers), len(unique_watchers)","metadata":{"execution":{"iopub.status.busy":"2023-10-05T16:09:56.027106Z","iopub.execute_input":"2023-10-05T16:09:56.027508Z","iopub.status.idle":"2023-10-05T16:09:57.882827Z","shell.execute_reply.started":"2023-10-05T16:09:56.027477Z","shell.execute_reply":"2023-10-05T16:09:57.881667Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"len(set(unique_watchers) & set(unique_attackers))","metadata":{"execution":{"iopub.status.busy":"2023-10-05T16:10:07.213743Z","iopub.execute_input":"2023-10-05T16:10:07.214125Z","iopub.status.idle":"2023-10-05T16:10:07.257525Z","shell.execute_reply.started":"2023-10-05T16:10:07.214096Z","shell.execute_reply":"2023-10-05T16:10:07.256598Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Watch out common id are shared between node types, which will likely mess up netwrokx\n# For this purpose we do remap to watchers id negative to avoid overlap\ndf = df.with_columns(\n    (pl.col(edge_cols[0]) * -1 - 1),\n)\nunique_watchers = list(sorted(df.select(pl.col(edge_cols[0])).unique().to_numpy().flatten()))\n","metadata":{"execution":{"iopub.status.busy":"2023-10-05T16:10:24.653170Z","iopub.execute_input":"2023-10-05T16:10:24.653533Z","iopub.status.idle":"2023-10-05T16:10:25.743879Z","shell.execute_reply.started":"2023-10-05T16:10:24.653505Z","shell.execute_reply":"2023-10-05T16:10:25.742657Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"len(set(unique_watchers) & set(unique_attackers))","metadata":{"execution":{"iopub.status.busy":"2023-10-05T16:10:26.545522Z","iopub.execute_input":"2023-10-05T16:10:26.545915Z","iopub.status.idle":"2023-10-05T16:10:26.578173Z","shell.execute_reply.started":"2023-10-05T16:10:26.545883Z","shell.execute_reply":"2023-10-05T16:10:26.577449Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"weighted_edges_df = df.group_by(edge_cols).agg(pl.count().alias(\"weight\")).sort(pl.col(\"weight\"), descending=True)","metadata":{"execution":{"iopub.status.busy":"2023-10-05T16:10:29.585273Z","iopub.execute_input":"2023-10-05T16:10:29.585683Z","iopub.status.idle":"2023-10-05T16:10:33.286709Z","shell.execute_reply.started":"2023-10-05T16:10:29.585653Z","shell.execute_reply":"2023-10-05T16:10:33.285649Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"G = nx.Graph()\nG.add_nodes_from(weighted_edges_df.select(pl.col(edge_cols[0])).to_numpy().flatten(), bipartite=0)\nG.add_nodes_from(weighted_edges_df.select(pl.col(edge_cols[1])).to_numpy().flatten(), bipartite=1)\n# G.add_edges_from(weighted_edges_df.select(pl.col(edge_cols)).iter_rows())\nG.add_weighted_edges_from(weighted_edges_df.iter_rows())\nnx.is_connected(G)","metadata":{"execution":{"iopub.status.busy":"2023-10-05T16:10:51.279710Z","iopub.execute_input":"2023-10-05T16:10:51.280518Z","iopub.status.idle":"2023-10-05T16:12:16.483698Z","shell.execute_reply.started":"2023-10-05T16:10:51.280482Z","shell.execute_reply":"2023-10-05T16:12:16.482560Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Scipy [WIP]\n\nSimilaryl we can construct the adjacency matrix of the bipartite graph - using scipy sparse matrix. Otherwise it won't fit in memory unless you have 4TB of RAM.\n\nOne should look at : https://docs.scipy.org/doc/scipy-0.14.0/reference/generated/scipy.sparse.csgraph.html\n","metadata":{}},{"cell_type":"code","source":"import scipy\nfrom tqdm.auto import tqdm\n\nadjacency = scipy.sparse.lil_array((len(unique_watchers), len(unique_attackers)),\n                                   dtype=np.int32\n                                  )\n\n\n# for row in tqdm(weighted_edges_df.sort(edge_cols).iter_rows()):\n#     # Too slow\n#     adjacency[unique_watchers.index(row[0]), unique_attackers.index(row[1])] = np.int32(row[2]) # We set the weight","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Graph Analysis [WIP]","metadata":{}}]}