{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"Implicit ALS model\n\n[Implicit](https://github.com/benfred/implicit/) is a library for recommender models. In theory, it supports GPU out-of-the-box, but I haven't tried it yet.\n\nIn this notebook we use ALS (Alternating Least Squares), but the library supports a lot of other models with not many changes.\n\nALS is one of the most used ML models for recommender systems. It's a matrix factorization method based on SVD (it's actually an approximated, numerical version of SVD). Basically, ALS factorizes the interaction matrix (user x items) into two smaller matrices, one for item embeddings and one for user embeddings. These new matrices are built in a manner such that the multiplication of a user and an item gives (approximately) it's interaction score. This build embeddings for items and for users that live in the same vector space, allowing the implementation of recommendations as simple cosine distances between users and items. This is, the 3 items we recommend for a given user are the 3 items with their embedding vectors closer to the user embedding vector.\n\nThere are a lot of online resources explaining it. For example, [here](https://towardsdatascience.com/prototyping-a-recommender-system-step-by-step-part-2-alternating-least-square-als-matrix-4a76c58714a1).","metadata":{}},{"cell_type":"markdown","source":"# Imports","metadata":{}},{"cell_type":"code","source":"# FYI:\n# This pip command takes a lot with GPU enabled (~15 min)\n# It works though. And GPU accelerates the process *a lot*.\n# I am developing with GPU turned off and submitting with GPU turned on\n!pip install --upgrade implicit","metadata":{"execution":{"iopub.status.busy":"2022-09-28T10:51:19.340953Z","iopub.execute_input":"2022-09-28T10:51:19.341580Z","iopub.status.idle":"2022-09-28T10:51:29.807935Z","shell.execute_reply.started":"2022-09-28T10:51:19.341480Z","shell.execute_reply":"2022-09-28T10:51:29.807103Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import os; os.environ['OPENBLAS_NUM_THREADS']='1'\nimport numpy as np\nimport pandas as pd\nimport implicit\nfrom scipy.sparse import coo_matrix\nfrom implicit.evaluation import mean_average_precision_at_k\nfrom sklearn.model_selection import train_test_split","metadata":{"execution":{"iopub.status.busy":"2022-09-28T10:51:49.435199Z","iopub.execute_input":"2022-09-28T10:51:49.435467Z","iopub.status.idle":"2022-09-28T10:51:49.440746Z","shell.execute_reply.started":"2022-09-28T10:51:49.435434Z","shell.execute_reply":"2022-09-28T10:51:49.439979Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Load dataframes","metadata":{}},{"cell_type":"code","source":"base_path = '../input/visit-data/'\ndata_path = f'{base_path}visit data.csv'\n\ndf = pd.read_csv(data_path)","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","execution":{"iopub.status.busy":"2022-09-28T10:51:53.599493Z","iopub.execute_input":"2022-09-28T10:51:53.600375Z","iopub.status.idle":"2022-09-28T10:52:01.768473Z","shell.execute_reply.started":"2022-09-28T10:51:53.600334Z","shell.execute_reply":"2022-09-28T10:52:01.767641Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df['rating'] = 1\ndf.shape","metadata":{"execution":{"iopub.status.busy":"2022-09-28T10:52:22.475816Z","iopub.execute_input":"2022-09-28T10:52:22.476114Z","iopub.status.idle":"2022-09-28T10:52:22.515925Z","shell.execute_reply.started":"2022-09-28T10:52:22.476080Z","shell.execute_reply":"2022-09-28T10:52:22.515114Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Assign autoincrementing ids starting from 0 to both users and items","metadata":{}},{"cell_type":"code","source":"ALL_USERS = df['player_roblox_id'].unique().tolist()\nALL_ITEMS = df['game_id'].unique().tolist()\n\nuser_ids = dict(list(enumerate(ALL_USERS)))\nitem_ids = dict(list(enumerate(ALL_ITEMS)))\n\nuser_map = {u: uidx for uidx, u in user_ids.items()}\nitem_map = {i: iidx for iidx, i in item_ids.items()}\n\ndf['user_id'] = df['player_roblox_id'].map(user_map)\ndf['item_id'] = df['game_id'].map(item_map)","metadata":{"execution":{"iopub.status.busy":"2022-09-28T11:06:45.694000Z","iopub.execute_input":"2022-09-28T11:06:45.694596Z","iopub.status.idle":"2022-09-28T11:07:00.122761Z","shell.execute_reply.started":"2022-09-28T11:06:45.694558Z","shell.execute_reply":"2022-09-28T11:07:00.122030Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"row = df['user_id'].values\ncol = df['item_id'].values\ndata = np.ones(df.shape[0])\ncoo_train = coo_matrix((data, (row, col)), shape=(len(ALL_USERS), len(ALL_ITEMS)))\ncoo_train","metadata":{"execution":{"iopub.status.busy":"2022-09-28T11:07:06.426633Z","iopub.execute_input":"2022-09-28T11:07:06.429005Z","iopub.status.idle":"2022-09-28T11:07:06.573412Z","shell.execute_reply.started":"2022-09-28T11:07:06.428957Z","shell.execute_reply":"2022-09-28T11:07:06.572758Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Check that model works ok with data","metadata":{}},{"cell_type":"code","source":"%%time\nmodel = implicit.als.AlternatingLeastSquares(factors=10, iterations=2)\nmodel.fit(coo_train)","metadata":{"execution":{"iopub.status.busy":"2022-09-27T14:28:20.820248Z","iopub.execute_input":"2022-09-27T14:28:20.821061Z","iopub.status.idle":"2022-09-27T14:28:32.910084Z","shell.execute_reply.started":"2022-09-27T14:28:20.821010Z","shell.execute_reply":"2022-09-27T14:28:32.909357Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Validation","metadata":{}},{"cell_type":"markdown","source":"## Functions required for validation","metadata":{}},{"cell_type":"code","source":"def to_user_item_coo(df):\n    \"\"\" Turn a dataframe with transactions into a COO sparse items x users matrix\"\"\"\n    row = df['user_id'].values\n    col = df['item_id'].values\n    data = np.ones(df.shape[0])\n    coo = coo_matrix((data, (row, col)), shape=(len(ALL_USERS), len(ALL_ITEMS)))\n    return coo\n\n\ndef split_data(df):\n    df_train, df_val = train_test_split(df, test_size=0.2)\n    return df_train, df_val\n\ndef get_val_matrices(df):\n    \"\"\" Split into training and validation and create various matrices\n        \n        Returns a dictionary with the following keys:\n            coo_train: training data in COO sparse format and as (users x items)\n            csr_train: training data in CSR sparse format and as (users x items)\n            csr_val:  validation data in CSR sparse format and as (users x items)\n    \n    \"\"\"\n    df_train, df_val = split_data(df)\n    coo_train = to_user_item_coo(df_train)\n    coo_val = to_user_item_coo(df_val)\n\n    csr_train = coo_train.tocsr()\n    csr_val = coo_val.tocsr()\n    \n    return {'coo_train': coo_train,\n            'csr_train': csr_train,\n            'csr_val': csr_val\n          }\n\n\ndef validate(matrices, factors=200, iterations=20, regularization=0.01, show_progress=True):\n    \"\"\" Train an ALS model with <<factors>> (embeddings dimension) \n    for <<iterations>> over matrices and validate with MAP@3\n    \"\"\"\n    coo_train, csr_train, csr_val = matrices['coo_train'], matrices['csr_train'], matrices['csr_val']\n    \n    model = implicit.als.AlternatingLeastSquares(factors=factors, \n                                                 iterations=iterations, \n                                                 regularization=regularization, \n                                                 random_state=42)\n    model.fit(coo_train, show_progress=show_progress)\n    \n    # The MAPK by implicit doesn't allow to calculate allowing repeated items, which is the case.\n    map3 = mean_average_precision_at_k(model, csr_train, csr_val, K=3, show_progress=show_progress, num_threads=4)\n    print(f\"Factors: {factors:>3} - Iterations: {iterations:>2} - Regularization: {regularization:4.3f} ==> MAP@3: {map3:6.5f}\")\n    return map3","metadata":{"execution":{"iopub.status.busy":"2022-09-28T11:07:37.244970Z","iopub.execute_input":"2022-09-28T11:07:37.245296Z","iopub.status.idle":"2022-09-28T11:07:37.262727Z","shell.execute_reply.started":"2022-09-28T11:07:37.245261Z","shell.execute_reply":"2022-09-28T11:07:37.261803Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"matrices = get_val_matrices(df)","metadata":{"execution":{"iopub.status.busy":"2022-09-27T14:29:14.708198Z","iopub.execute_input":"2022-09-27T14:29:14.708514Z","iopub.status.idle":"2022-09-27T14:29:19.487488Z","shell.execute_reply.started":"2022-09-27T14:29:14.708469Z","shell.execute_reply":"2022-09-27T14:29:19.486659Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\nbest_map3 = 0\n'''\nLimited by the memory usage of this notebook, \nthe maximum factors we can have is 100\nshould run for factors of 200, 500, 1000 if possible\n'''\nfor factors in [50,100]:\n    for iterations in [12,15,20]:\n        for regularization in [0.01]:\n            map3 = validate(matrices, factors, iterations, regularization, show_progress=False)\n            if map3 > best_map3:\n                best_map3 = map3\n                best_params = {'factors': factors, 'iterations': iterations, 'regularization': regularization}\n                print(f\"Best MAP@3 found. Updating: {best_params}\")","metadata":{"execution":{"iopub.status.busy":"2022-09-27T14:29:24.924911Z","iopub.execute_input":"2022-09-27T14:29:24.925220Z","iopub.status.idle":"2022-09-27T14:34:41.659965Z","shell.execute_reply.started":"2022-09-27T14:29:24.925188Z","shell.execute_reply":"2022-09-27T14:34:41.658189Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"del matrices","metadata":{"execution":{"iopub.status.busy":"2022-09-27T14:36:43.157032Z","iopub.execute_input":"2022-09-27T14:36:43.157299Z","iopub.status.idle":"2022-09-27T14:36:43.163737Z","shell.execute_reply.started":"2022-09-27T14:36:43.157269Z","shell.execute_reply":"2022-09-27T14:36:43.162806Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Training over the full dataset","metadata":{}},{"cell_type":"code","source":"coo_train = to_user_item_coo(df)\ncsr_train = coo_train.tocsr()","metadata":{"execution":{"iopub.status.busy":"2022-09-28T11:07:50.867852Z","iopub.execute_input":"2022-09-28T11:07:50.868142Z","iopub.status.idle":"2022-09-28T11:07:51.416486Z","shell.execute_reply.started":"2022-09-28T11:07:50.868109Z","shell.execute_reply":"2022-09-28T11:07:51.415647Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def train(coo_train, factors=100, iterations=15, regularization=0.01, show_progress=True):\n    model = implicit.als.AlternatingLeastSquares(factors=factors, \n                                                 iterations=iterations, \n                                                 regularization=regularization, \n                                                 random_state=42)\n    model.fit(coo_train, show_progress=show_progress)\n    return model","metadata":{"execution":{"iopub.status.busy":"2022-09-28T11:07:53.896689Z","iopub.execute_input":"2022-09-28T11:07:53.896944Z","iopub.status.idle":"2022-09-28T11:07:53.901433Z","shell.execute_reply.started":"2022-09-28T11:07:53.896914Z","shell.execute_reply":"2022-09-28T11:07:53.900726Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"best_params","metadata":{"execution":{"iopub.status.busy":"2022-09-28T11:08:03.804107Z","iopub.execute_input":"2022-09-28T11:08:03.804403Z","iopub.status.idle":"2022-09-28T11:08:03.810933Z","shell.execute_reply.started":"2022-09-28T11:08:03.804368Z","shell.execute_reply":"2022-09-28T11:08:03.810095Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"model = train(coo_train, **best_params)","metadata":{"execution":{"iopub.status.busy":"2022-09-28T11:08:07.463476Z","iopub.execute_input":"2022-09-28T11:08:07.464195Z","iopub.status.idle":"2022-09-28T11:10:09.305519Z","shell.execute_reply.started":"2022-09-28T11:08:07.464159Z","shell.execute_reply":"2022-09-28T11:10:09.304779Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Recommendation Example","metadata":{}},{"cell_type":"code","source":"# recommend items for a user\n# can handle cold start problem by setting recalculate_user as True\nUser_id = 286281035\n\nresult_id, result_score = model.recommend(user_map[User_id], csr_train[user_map[User_id]], N=3, filter_already_liked_items=True, recalculate_user = False)\nresult_game_id = [item_ids[x] for x in result_id]\nresult_table = pd.DataFrame({\"game_id\": result_game_id, \"score\": result_score})\nresult_table","metadata":{"execution":{"iopub.status.busy":"2022-09-28T11:29:53.063837Z","iopub.execute_input":"2022-09-28T11:29:53.064122Z","iopub.status.idle":"2022-09-28T11:29:53.077934Z","shell.execute_reply.started":"2022-09-28T11:29:53.064090Z","shell.execute_reply":"2022-09-28T11:29:53.077084Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#Calculates a list of similar items\nGame_id = 1892760830\n\nrelated_id, related_score = model.similar_items(item_map[Game_id],filter_items = [item_map[Game_id]])\nrelated_game_id = [item_ids[x] for x in related_id]\nrelated_table = pd.DataFrame({\"game_id\": related_game_id, \"score\": related_score})\nrelated_table","metadata":{"execution":{"iopub.status.busy":"2022-09-28T11:26:35.246673Z","iopub.execute_input":"2022-09-28T11:26:35.246949Z","iopub.status.idle":"2022-09-28T11:26:35.260246Z","shell.execute_reply.started":"2022-09-28T11:26:35.246915Z","shell.execute_reply":"2022-09-28T11:26:35.259027Z"},"trusted":true},"execution_count":null,"outputs":[]}]}