{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# H&M - Implicit ALS model\n\nhttp://ethen8181.github.io/machine-learning/recsys/2_implicit.html#Mean-Average-Precision (main theory here for it overall (mean average precision highlighted here)\n\n\n[Implicit](https://github.com/benfred/implicit/) is a library for recommender models. \n\nUsing alternative least squares method to generate a baseline recommender system model. ALS is an extremely common matrix factorization method used for recommender systems, with its theoretical base being SVD (in that it is an approximated numerical version of Singular value decomposition). ALS approach factorizes the matrix (interaction matrix of user and items) into two separate smaller matrices one being for item embeddings and the other for user embeddings. These new matrices are created such that the multiplcation of a user and an item gives an approximate value for its interaction score.\n\nThe advantage of this system is that it allows for the building of embeddings for both items and for users that live in the same vector space, therefore enabling implementation of recommendations through simplistic approaches such as cosine distances between users and items. \n\nFor example the 12 recommended items for a given user is extracted by computing the 12 closest items in terms of their vector embedding distances between them and the user.\n\n\n\n(for more information look here) [here](https://towardsdatascience.com/prototyping-a-recommender-system-step-by-step-part-2-alternating-least-square-als-matrix-4a76c58714a1).\n\n\n\n\n\n---\n","metadata":{}},{"cell_type":"markdown","source":"# Imports","metadata":{}},{"cell_type":"code","source":"# FYI:\n# This pip command takes a lot with GPU enabled (~15 min)\n# It works though. And GPU accelerates the process *a lot*.\n# I am developing with GPU turned off and submitting with GPU turned on\n!pip install --upgrade implicit","metadata":{"execution":{"iopub.status.busy":"2022-09-09T00:46:55.208975Z","iopub.execute_input":"2022-09-09T00:46:55.209716Z","iopub.status.idle":"2022-09-09T00:47:06.540315Z","shell.execute_reply.started":"2022-09-09T00:46:55.209677Z","shell.execute_reply":"2022-09-09T00:47:06.539393Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import os; os.environ['OPENBLAS_NUM_THREADS']='1'\nimport numpy as np\nimport pandas as pd\nimport implicit\nfrom scipy.sparse import coo_matrix\nfrom implicit.evaluation import mean_average_precision_at_k","metadata":{"execution":{"iopub.status.busy":"2022-09-09T00:47:06.543944Z","iopub.execute_input":"2022-09-09T00:47:06.544164Z","iopub.status.idle":"2022-09-09T00:47:06.898835Z","shell.execute_reply.started":"2022-09-09T00:47:06.544135Z","shell.execute_reply":"2022-09-09T00:47:06.898078Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Load dataframes","metadata":{}},{"cell_type":"code","source":"%%time\n\nbase_path = '../input/h-and-m-personalized-fashion-recommendations/'\ncsv_train = f'{base_path}transactions_train.csv'\ncsv_sub = f'{base_path}sample_submission.csv'\ncsv_users = f'{base_path}customers.csv'\ncsv_items = f'{base_path}articles.csv'\n\ndf = pd.read_csv(csv_train, dtype={'article_id': str}, parse_dates=['t_dat'])\ndf_sub = pd.read_csv(csv_sub)\ndfu = pd.read_csv(csv_users)\ndfi = pd.read_csv(csv_items, dtype={'article_id': str})","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","execution":{"iopub.status.busy":"2022-09-09T00:47:06.900168Z","iopub.execute_input":"2022-09-09T00:47:06.900431Z","iopub.status.idle":"2022-09-09T00:48:25.238838Z","shell.execute_reply.started":"2022-09-09T00:47:06.900395Z","shell.execute_reply":"2022-09-09T00:48:25.237923Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Trying with less data:\n# https://www.kaggle.com/tomooinubushi/folk-of-time-is-our-best-friend/notebook\ndf = df[df['t_dat'] > '2020-08-21']\ndf.shape","metadata":{"execution":{"iopub.status.busy":"2022-09-09T00:48:25.240853Z","iopub.execute_input":"2022-09-09T00:48:25.241727Z","iopub.status.idle":"2022-09-09T00:48:26.126862Z","shell.execute_reply.started":"2022-09-09T00:48:25.241678Z","shell.execute_reply":"2022-09-09T00:48:26.125307Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# For validation this means 3 weeks of training and 1 week for validation\n# For submission, it means 4 weeks of training\ndf['t_dat'].max()","metadata":{"execution":{"iopub.status.busy":"2022-09-09T00:48:26.128170Z","iopub.execute_input":"2022-09-09T00:48:26.128953Z","iopub.status.idle":"2022-09-09T00:48:26.139488Z","shell.execute_reply.started":"2022-09-09T00:48:26.128914Z","shell.execute_reply":"2022-09-09T00:48:26.138601Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Assign autoincrementing ids starting from 0 to both users and items","metadata":{}},{"cell_type":"code","source":"ALL_USERS = dfu['customer_id'].unique().tolist()\nALL_ITEMS = dfi['article_id'].unique().tolist()\n\nuser_ids = dict(list(enumerate(ALL_USERS)))\nitem_ids = dict(list(enumerate(ALL_ITEMS)))\n\nuser_map = {u: uidx for uidx, u in user_ids.items()}\nitem_map = {i: iidx for iidx, i in item_ids.items()}\n\ndf['user_id'] = df['customer_id'].map(user_map)\ndf['item_id'] = df['article_id'].map(item_map)\n\ndel dfu, dfi","metadata":{"execution":{"iopub.status.busy":"2022-09-09T00:48:26.140983Z","iopub.execute_input":"2022-09-09T00:48:26.142404Z","iopub.status.idle":"2022-09-09T00:48:29.087566Z","shell.execute_reply.started":"2022-09-09T00:48:26.142362Z","shell.execute_reply":"2022-09-09T00:48:29.086761Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Create coo_matrix (user x item) and csr matrix (user x item)\nCoo matrix= cordinate format matrices\nCSR matrix= compressed sparse row matrices\n\nRecommender systems typically use sparse matrices to represent the core problem of user and item interaction, wherein each of the values represent if the user purchased,liked or reviewed an item. Since users only interact and purchase a small subsection of the complete catalog of products, (user X item) matrix is sparse in nature (full of zero values)\n\nTherefore we create the matrix in the Coo sparse matrix format \n\n\n### About COO matrices\nCOO matrices are a kind of sparse matrix.\nThey store their values as tuples of `(row, column, value)` (the coordinates)\n\nYou can read more about them here: \n* https://en.wikipedia.org/wiki/Sparse_matrix#Coordinate_list_(COO)\n* https://scipy-lectures.org/advanced/scipy_sparse/coo_matrix.html\n* https://het.as.utexas.edu/HET/Software/Scipy/generated/scipy.sparse.coo_matrix.html\n\nwe then convert the COO matrix into A CSR/CRS matrix, which is a compressed representation of the COO matrix, so as to speed up the process of computation\n\n## About CSR matrices\n* https://en.wikipedia.org/wiki/Sparse_matrix#Compressed_sparse_row_(CSR,_CRS_or_Yale_format)\n\n\n```python\n>>> row  = np.array([0,3,1,0]) # user_ids\n>>> col  = np.array([0,3,1,2]) # item_ids\n>>> data = np.array([4,5,7,9]) # a bunch of ones of lenght unique(user) x unique(items)\n>>> coo_matrix((data,(row,col)), shape=(4,4)).todense()\nmatrix([[4, 0, 9, 0],\n        [0, 7, 0, 0],\n        [0, 0, 0, 0],\n        [0, 0, 0, 5]])\n```\n\n","metadata":{}},{"cell_type":"code","source":"row = df['user_id'].values\ncol = df['item_id'].values\ndata = np.ones(df.shape[0])\ncoo_train = coo_matrix((data, (row, col)), shape=(len(ALL_USERS), len(ALL_ITEMS)))\ncoo_train","metadata":{"execution":{"iopub.status.busy":"2022-09-09T00:56:51.416489Z","iopub.execute_input":"2022-09-09T00:56:51.417355Z","iopub.status.idle":"2022-09-09T00:56:51.437181Z","shell.execute_reply.started":"2022-09-09T00:56:51.417296Z","shell.execute_reply":"2022-09-09T00:56:51.435491Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Testing the model, confirming it works perfectly with the data source","metadata":{}},{"cell_type":"code","source":"%%time\nmodel = implicit.als.AlternatingLeastSquares(factors=10, iterations=2)\nmodel.fit(coo_train)","metadata":{"execution":{"iopub.status.busy":"2022-09-09T00:56:58.005078Z","iopub.execute_input":"2022-09-09T00:56:58.005364Z","iopub.status.idle":"2022-09-09T00:56:59.258368Z","shell.execute_reply.started":"2022-09-09T00:56:58.005332Z","shell.execute_reply":"2022-09-09T00:56:59.257596Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Validation","metadata":{}},{"cell_type":"markdown","source":"## Creating functions required for validation","metadata":{}},{"cell_type":"code","source":"#creates a COO sparse matrix for transactions data\ndef to_user_item_coo(df):\n    \"\"\" Turn a dataframe with transactions into a COO sparse items x users matrix\"\"\"\n    row = df['user_id'].values\n    col = df['item_id'].values\n    data = np.ones(df.shape[0])\n    coo = coo_matrix((data, (row, col)), shape=(len(ALL_USERS), len(ALL_ITEMS)))\n    return coo\n\n#splits the whole transactions pandas dataframe into training and validation data, with\n#validation period of 7 days\ndef split_data(df, validation_days=7):\n    \"\"\" Split a pandas dataframe into training and validation data, using <<validation_days>>\n    \"\"\"\n    validation_cut = df['t_dat'].max() - pd.Timedelta(validation_days)\n\n    df_train = df[df['t_dat'] < validation_cut]\n    df_val = df[df['t_dat'] >= validation_cut]\n    return df_train, df_val\n\n#creates COO matrix and CSR matrix for training data as well as validation data after \n#the split of the pandas dataframe.\ndef get_val_matrices(df, validation_days=7):\n    \"\"\" Split into training and validation and create various matrices\n        \n        Returns a dictionary with the following keys:\n            coo_train: training data in COO sparse format and as (users x items)\n            csr_train: training data in CSR sparse format and as (users x items)\n            csr_val:  validation data in CSR sparse format and as (users x items)\n    \n    \"\"\"\n    df_train, df_val = split_data(df, validation_days=validation_days)\n    coo_train = to_user_item_coo(df_train)\n    coo_val = to_user_item_coo(df_val)\n\n    csr_train = coo_train.tocsr()\n    csr_val = coo_val.tocsr()\n    \n    return {'coo_train': coo_train,\n            'csr_train': csr_train,\n            'csr_val': csr_val\n          }\n\n#training an ALS model with embedding dimension = factors, and using regularization\n\n#validating using MAP@12\n#i.e mean average precision for K=12\n\n\ndef validate(matrices, factors=200, iterations=20, regularization=0.01, show_progress=True):\n    \"\"\" Train an ALS model with <<factors>> (embeddings dimension) \n    for <<iterations>> over matrices and validate with MAP@12\n    \"\"\"\n    coo_train, csr_train, csr_val = matrices['coo_train'], matrices['csr_train'], matrices['csr_val']\n    \n    model = implicit.als.AlternatingLeastSquares(factors=factors, \n                                                 iterations=iterations, \n                                                 regularization=regularization, \n                                                 random_state=42)\n    model.fit(coo_train, show_progress=show_progress)\n    \n    # The MAPK by implicit doesn't allow to calculate allowing repeated items, which is the case.\n    # TODO: change MAP@12 to a library that allows repeated items in prediction\n    \n    \n    map12 = mean_average_precision_at_k(model, csr_train, csr_val, K=12, show_progress=show_progress, num_threads=4)\n    print(f\"Factors: {factors:>3} - Iterations: {iterations:>2} - Regularization: {regularization:4.3f} ==> MAP@12: {map12:6.5f}\")\n    return map12","metadata":{"execution":{"iopub.status.busy":"2022-09-09T00:57:47.334393Z","iopub.execute_input":"2022-09-09T00:57:47.334972Z","iopub.status.idle":"2022-09-09T00:57:47.346115Z","shell.execute_reply.started":"2022-09-09T00:57:47.334937Z","shell.execute_reply":"2022-09-09T00:57:47.345338Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"matrices = get_val_matrices(df)","metadata":{"execution":{"iopub.status.busy":"2022-09-09T00:57:49.076967Z","iopub.execute_input":"2022-09-09T00:57:49.077519Z","iopub.status.idle":"2022-09-09T00:57:49.290274Z","shell.execute_reply.started":"2022-09-09T00:57:49.077484Z","shell.execute_reply":"2022-09-09T00:57:49.289499Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\nbest_map12 = 0\nfor factors in [40, 50, 60, 100, 200, 500, 1000]:\n    for iterations in [3, 12, 14, 15, 20]:\n        for regularization in [0.01]:\n            map12 = validate(matrices, factors, iterations, regularization, show_progress=False)\n            if map12 > best_map12:\n                best_map12 = map12\n                best_params = {'factors': factors, 'iterations': iterations, 'regularization': regularization}\n                print(f\"Best MAP@12 found. Updating: {best_params}\")","metadata":{"execution":{"iopub.status.busy":"2022-09-09T00:57:50.932766Z","iopub.execute_input":"2022-09-09T00:57:50.933433Z","iopub.status.idle":"2022-09-09T01:04:50.452067Z","shell.execute_reply.started":"2022-09-09T00:57:50.933391Z","shell.execute_reply":"2022-09-09T01:04:50.451308Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"del matrices","metadata":{"execution":{"iopub.status.busy":"2022-09-09T01:04:50.453628Z","iopub.execute_input":"2022-09-09T01:04:50.453872Z","iopub.status.idle":"2022-09-09T01:04:50.459550Z","shell.execute_reply.started":"2022-09-09T01:04:50.453840Z","shell.execute_reply":"2022-09-09T01:04:50.458838Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Training over the full dataset","metadata":{}},{"cell_type":"code","source":"coo_train = to_user_item_coo(df)\ncsr_train = coo_train.tocsr()","metadata":{"execution":{"iopub.status.busy":"2022-09-09T01:04:50.460716Z","iopub.execute_input":"2022-09-09T01:04:50.461131Z","iopub.status.idle":"2022-09-09T01:04:50.539749Z","shell.execute_reply.started":"2022-09-09T01:04:50.461094Z","shell.execute_reply":"2022-09-09T01:04:50.538921Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def train(coo_train, factors=200, iterations=15, regularization=0.01, show_progress=True):\n    model = implicit.als.AlternatingLeastSquares(factors=factors, \n                                                 iterations=iterations, \n                                                 regularization=regularization, \n                                                 random_state=42)\n    model.fit(coo_train, show_progress=show_progress)\n    return model","metadata":{"execution":{"iopub.status.busy":"2022-09-09T01:04:50.542192Z","iopub.execute_input":"2022-09-09T01:04:50.542543Z","iopub.status.idle":"2022-09-09T01:04:50.548296Z","shell.execute_reply.started":"2022-09-09T01:04:50.542507Z","shell.execute_reply":"2022-09-09T01:04:50.547308Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"best_params","metadata":{"execution":{"iopub.status.busy":"2022-09-09T01:04:50.549511Z","iopub.execute_input":"2022-09-09T01:04:50.549772Z","iopub.status.idle":"2022-09-09T01:04:50.563768Z","shell.execute_reply.started":"2022-09-09T01:04:50.549734Z","shell.execute_reply":"2022-09-09T01:04:50.562930Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"model = train(coo_train, **best_params)","metadata":{"execution":{"iopub.status.busy":"2022-09-09T01:04:50.565111Z","iopub.execute_input":"2022-09-09T01:04:50.565740Z","iopub.status.idle":"2022-09-09T01:04:53.798807Z","shell.execute_reply.started":"2022-09-09T01:04:50.565620Z","shell.execute_reply":"2022-09-09T01:04:53.798029Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Submission","metadata":{}},{"cell_type":"markdown","source":"## Submission function","metadata":{}},{"cell_type":"code","source":"def submit(model, csr_train, submission_name=\"submissions.csv\"):\n    preds = []\n    batch_size = 2000\n    to_generate = np.arange(len(ALL_USERS))\n    for startidx in range(0, len(to_generate), batch_size):\n        batch = to_generate[startidx : startidx + batch_size]\n        ids, scores = model.recommend(batch, csr_train[batch], N=12, filter_already_liked_items=False)\n        for i, userid in enumerate(batch):\n            customer_id = user_ids[userid]\n            user_items = ids[i]\n            article_ids = [item_ids[item_id] for item_id in user_items]\n            preds.append((customer_id, ' '.join(article_ids)))\n\n    df_preds = pd.DataFrame(preds, columns=['customer_id', 'prediction'])\n    df_preds.to_csv(submission_name, index=False)\n    \n    display(df_preds.head())\n    print(df_preds.shape)\n    \n    return df_preds","metadata":{"execution":{"iopub.status.busy":"2022-09-09T01:04:53.800091Z","iopub.execute_input":"2022-09-09T01:04:53.800488Z","iopub.status.idle":"2022-09-09T01:04:53.809953Z","shell.execute_reply.started":"2022-09-09T01:04:53.800450Z","shell.execute_reply":"2022-09-09T01:04:53.809185Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\ndf_preds = submit(model, csr_train);","metadata":{"execution":{"iopub.status.busy":"2022-09-09T01:04:53.811304Z","iopub.execute_input":"2022-09-09T01:04:53.811740Z","iopub.status.idle":"2022-09-09T01:05:47.368066Z","shell.execute_reply.started":"2022-09-09T01:04:53.811703Z","shell.execute_reply":"2022-09-09T01:05:47.367205Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]}]}