{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"### In this notebook, I will share an approach to prepare data for a recommendation system as a graph and build a Graph Convolutional model based on Pytorch Geometric library.","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","execution":{"iopub.status.busy":"2022-11-13T18:33:42.596685Z","iopub.execute_input":"2022-11-13T18:33:42.597224Z","iopub.status.idle":"2022-11-13T18:33:42.607222Z","shell.execute_reply.started":"2022-11-13T18:33:42.597182Z","shell.execute_reply":"2022-11-13T18:33:42.605846Z"}}},{"cell_type":"code","source":"import numpy as np \nimport pandas as pd\n\nimport os\nfor dirname, _, filenames in os.walk('/kaggle/input'):\n    for filename in filenames:\n        print(os.path.join(dirname, filename))","metadata":{"execution":{"iopub.status.busy":"2022-11-13T19:35:59.277345Z","iopub.execute_input":"2022-11-13T19:35:59.277749Z","iopub.status.idle":"2022-11-13T19:35:59.289772Z","shell.execute_reply.started":"2022-11-13T19:35:59.277718Z","shell.execute_reply":"2022-11-13T19:35:59.287927Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### We need to install here PyG library and required components. Please take a coffee, because it is not a quick process :)","metadata":{}},{"cell_type":"code","source":"!pip install torch-geometric\n!pip install torch-scatter\n!pip install torch-sparse\n!pip install torch-cluster\n!pip install torch-spline-conv","metadata":{"execution":{"iopub.status.busy":"2022-11-13T18:34:54.405884Z","iopub.execute_input":"2022-11-13T18:34:54.406366Z","iopub.status.idle":"2022-11-13T19:05:35.937047Z","shell.execute_reply.started":"2022-11-13T18:34:54.406332Z","shell.execute_reply":"2022-11-13T19:05:35.935686Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### To be able to train a Graph Convolutional model for recommedation system we need to convert the data into a suitable form - PyG HeteroData. Basically, HeteroData is a graph with several different types of entities and links inside.","metadata":{}},{"cell_type":"code","source":"from torch_geometric.data import HeteroData","metadata":{"execution":{"iopub.status.busy":"2022-11-13T19:06:20.518898Z","iopub.execute_input":"2022-11-13T19:06:20.519417Z","iopub.status.idle":"2022-11-13T19:06:25.939138Z","shell.execute_reply.started":"2022-11-13T19:06:20.519368Z","shell.execute_reply":"2022-11-13T19:06:25.936985Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"TRAIN_PATH = '/kaggle/input/otto-recommender-system/train.jsonl'","metadata":{"execution":{"iopub.status.busy":"2022-11-13T19:06:25.941831Z","iopub.execute_input":"2022-11-13T19:06:25.943348Z","iopub.status.idle":"2022-11-13T19:06:25.949017Z","shell.execute_reply.started":"2022-11-13T19:06:25.943286Z","shell.execute_reply":"2022-11-13T19:06:25.947004Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### Just for demonstration purposes we will take a small part of the whole dataset.","metadata":{}},{"cell_type":"code","source":"chunks = pd.read_json(TRAIN_PATH, lines=True, chunksize=10000)","metadata":{"execution":{"iopub.status.busy":"2022-11-13T19:06:25.950691Z","iopub.execute_input":"2022-11-13T19:06:25.951234Z","iopub.status.idle":"2022-11-13T19:06:25.969716Z","shell.execute_reply.started":"2022-11-13T19:06:25.951183Z","shell.execute_reply":"2022-11-13T19:06:25.968575Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"for chunk in chunks:\n    df = chunk\n    break","metadata":{"execution":{"iopub.status.busy":"2022-11-13T19:06:27.261218Z","iopub.execute_input":"2022-11-13T19:06:27.262258Z","iopub.status.idle":"2022-11-13T19:06:28.923765Z","shell.execute_reply.started":"2022-11-13T19:06:27.262210Z","shell.execute_reply":"2022-11-13T19:06:28.921899Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df","metadata":{"execution":{"iopub.status.busy":"2022-11-13T19:06:31.143249Z","iopub.execute_input":"2022-11-13T19:06:31.143757Z","iopub.status.idle":"2022-11-13T19:06:31.319793Z","shell.execute_reply.started":"2022-11-13T19:06:31.143715Z","shell.execute_reply":"2022-11-13T19:06:31.318325Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### What is happening here. \n#### 1. We need to have all entities from sessions (it is easy, all of them have values starting from 0).\n#### 2. Need to have all entities of aids (all ids should be scaled to 0-n for the HeteroData).\n#### 3. Need to have all connections (for usability will use pandas dataframe).","metadata":{}},{"cell_type":"code","source":"all_aids = []\nall_sessions = []\nall_connections = []\n\nfor index, row in df.iterrows():\n    all_sessions.append(row['session'])\n    json_row = row['events']\n    for sub_row in json_row:\n        all_aids.append(sub_row['aid'])\n        all_connections.append([row['session'], sub_row['aid'], sub_row['type']])\n\nconnections = pd.DataFrame(all_connections)\nall_aids = list(set(all_aids))\nconnections.columns = ['session', 'aid', 'link']\nconnections","metadata":{"execution":{"iopub.status.busy":"2022-11-13T19:21:10.822917Z","iopub.execute_input":"2022-11-13T19:21:10.823458Z","iopub.status.idle":"2022-11-13T19:21:18.232180Z","shell.execute_reply.started":"2022-11-13T19:21:10.823407Z","shell.execute_reply":"2022-11-13T19:21:18.230898Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### We just map here new index for aids.","metadata":{}},{"cell_type":"code","source":"aids_dict = dict(zip(all_aids, [i for i in range(len(all_aids))]))\nconnections['aid'] = connections['aid'].map(aids_dict)\nconnections","metadata":{"execution":{"iopub.status.busy":"2022-11-13T19:21:21.868891Z","iopub.execute_input":"2022-11-13T19:21:21.869454Z","iopub.status.idle":"2022-11-13T19:21:22.244549Z","shell.execute_reply.started":"2022-11-13T19:21:21.869412Z","shell.execute_reply":"2022-11-13T19:21:22.243263Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### Finaly we are preparing HeteroData object.\n#### 1. Each type of entity should be present in a graph (session, aid).\n#### 2. Entity should mandatory has set of features required for the training. In this dataset we don't have them, so let's use dummy features (0 for sessions, 1 for aids).\n#### 3. Each link type should be present in a graph as a separate dictionary. It is specifying by ['source_entity', 'connection_type', 'target_entity'].","metadata":{}},{"cell_type":"code","source":"import torch\n\ndata = HeteroData()\nsession_x = torch.tensor([0 for _ in range(10000)])\nsession_x = torch.reshape(session_x, (session_x.shape[0], 1))\ndata['session'].x = session_x\n\naid_x = torch.tensor([1 for _ in range(len(all_aids))])\naid_x = torch.reshape(aid_x, (aid_x.shape[0], 1))\ndata['aid'].x = aid_x\ndata['aid'].legend = all_aids\n\nfor category in connections['link'].unique():\n    sub_category = connections[connections['link'] == category]\n    sub_category = sub_category[['session', 'aid']]\n    sub_category = sub_category[['session', 'aid']].values.T\n    links = torch.tensor(sub_category)\n    data['session', category, 'aid'].edge_index = links\n    \ndata","metadata":{"execution":{"iopub.status.busy":"2022-11-13T20:18:37.713870Z","iopub.execute_input":"2022-11-13T20:18:37.714395Z","iopub.status.idle":"2022-11-13T20:18:38.142323Z","shell.execute_reply.started":"2022-11-13T20:18:37.714356Z","shell.execute_reply":"2022-11-13T20:18:38.140861Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Give me some time to prepare Graph Convolutional model :)","metadata":{}},{"cell_type":"markdown","source":"#### For more details you can visit [PyG documentation](https://pytorch-geometric.readthedocs.io/en/latest/).\n#### You can also visit my [personal blog](https://www.data-science-factory.com/) and I will be happy :)","metadata":{}},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]}]}