{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"Note:forked","metadata":{}},{"cell_type":"code","source":"import numpy as np\nimport pandas as pd\nimport os\nimport implicit\n\nfrom sklearn.preprocessing import MinMaxScaler, LabelEncoder, OrdinalEncoder\n\nimport scipy.sparse as sparse","metadata":{"execution":{"iopub.status.busy":"2022-02-08T21:43:25.018907Z","iopub.execute_input":"2022-02-08T21:43:25.019668Z","iopub.status.idle":"2022-02-08T21:43:26.175771Z","shell.execute_reply.started":"2022-02-08T21:43:25.019531Z","shell.execute_reply":"2022-02-08T21:43:26.175102Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"It is usually a good idea to approach any problem by establishing a baseline - logistic regression for classification, multi-armed bandit for reinforcement learning etc. Dealing with recommendations, such basic approach would be collaborative filtering: what's the simplest non-trivial recommendation you can make to a user? Suggest things purchased by others with similar history. What follows is a crash introduction.\n\nFirst important distinction is between explicit and implicit data: the former describes a situation where we have some sort of rating - think of stars on MovieLens; based on this information, we can infer preferences of hte user (likes / dislikes). Implicit data is the info we gather from user behavior, with no ratings or specific actions. We know that a purchase indicates someone probably liked what they bought - but there can be different reasons why they did *not* buy something else (maybe they disliked it, maybe they were not even aware of it). In this competition, we are dealing with implicit data.\n\nhttps://datasciencemadesimpler.wordpress.com/tag/alternating-least-squares/]\n\nA seminal paper introducing the idea of using Alternating Least Squares  - *the* first recommendation algo that went massive - can be found here: http://yifanhu.net/PUB/cf.pdf. \n\nFor this notebook, I am using a fantastic [`implicit`](https://implicit.readthedocs.io/en/latest/) package written by [Ben Frederickson](https://github.com/benfred).","metadata":{}},{"cell_type":"markdown","source":"# Data","metadata":{}},{"cell_type":"code","source":"df = pd.read_csv('../input/h-and-m-personalized-fashion-recommendations/transactions_train.csv', dtype={'article_id': str})\n\n# at the moment we are assuming all transactions are equally important => all get the same weight in terms of event strength (how important do we deem the signal)\ndf['event_strength'] = 1\ndf.head()\n\n# this will help us later ;-) \nxlist = list(np.unique(df['customer_id']))","metadata":{"execution":{"iopub.status.busy":"2022-02-08T21:43:26.177073Z","iopub.execute_input":"2022-02-08T21:43:26.177324Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Model","metadata":{}},{"cell_type":"code","source":"# The implicit package implementation does not play nice with non-integer identifiers and we need to be reversible for predictions \n# at the end; unfortunately we don't have sklearn >= 0.24, where new functionality in CardinalEncoder handles unknown values :-( \nle1 = LabelEncoder()\nle1.fit(df['customer_id']  )\ndf['customer_id'] = le1.transform(df['customer_id'])\n\n\nle2 = LabelEncoder()\nle2.fit(df['article_id'] )\ndf['article_id'] = le2.transform(df['article_id'])\n","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# prepare the (sparse) customer and article matrices \nsparse_article_customer = sparse.csr_matrix((df['event_strength'].astype(float), (df['article_id'], df['customer_id'])))\nsparse_customer_article = sparse.csr_matrix((df['event_strength'].astype(float), (df['customer_id'], df['article_id'])))\n","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# instantiate an ALS model\n# model = implicit.als.AlternatingLeastSquares(factors=25, regularization=0.1, iterations=25,num_threads=0)\nmodel = implicit.bpr.BayesianPersonalizedRanking(factors=25,  iterations=20,num_threads=0)\nalpha = 15\ndata = sparse_article_customer.astype('double')\nmodel.fit(data)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Prediction","metadata":{}},{"cell_type":"code","source":"customer_vecs = sparse.csr_matrix(model.user_factors)\narticle_vecs = sparse.csr_matrix(model.item_factors)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# generate recommendations for a specific user\ndef recommendations_per_user(custid, sparse_customer_article, person_vecs, content_vecs, num_contents=10):\n\n    # Get the interactions scores from the sparse person content matrix\n    customer_interactions = sparse_customer_article[custid,:].toarray()    \n    customer_interactions = customer_interactions.reshape(-1) + 1\n\n    # Make articles already interacted zero\n    customer_interactions[customer_interactions > 1] = 0    \n\n    rec_vector = customer_vecs[custid,:].dot(article_vecs.T).toarray()\n\n    # Scale this recommendation vector between 0 and 1\n    min_max = MinMaxScaler()\n    rec_vector_scaled = min_max.fit_transform(rec_vector.reshape(-1,1))[:,0]\n    # Content already interacted have their recommendation multiplied by zero\n    recommend_vector = customer_interactions * rec_vector_scaled\n    # Sort the indices of the content into order of best recommendations\n    content_idx = np.argsort(recommend_vector)[::-1][:num_contents]\n\n  \n    res_str = ' '.join([str(f) for f in list(le2.inverse_transform(content_idx) )])\n    \n    return res_str\n","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_sub = pd.read_csv('../input/h-and-m-personalized-fashion-recommendations/sample_submission.csv')\ndf_sub.head(10)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# we need a workaround to account for new values - and yes, I know it's ugly\n\ndf_old = df_sub[df_sub['customer_id'].isin(xlist)]\ndf_new = df_sub[~df_sub['customer_id'].isin(xlist)]\n","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%time\n\n# testing execution time - so generating proper predictions only for some rows\nhowmany = 10000000\nidlist = le1.transform(df_old['customer_id'])\npredlist = [recommendations_per_user(f, sparse_customer_article, customer_vecs, article_vecs, num_contents=10) for f in idlist[0:howmany]]\ndf_old['prediction'][0:howmany] = predlist","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_sub2 = pd.concat([df_new, df_old], axis = 0)\ndf_sub2.to_csv('als_submission.csv', index = False)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# The best is yet to come\n\nThis is a very simple notebook, written to introduce the idea of collaborative filtering as a way of solving recommendation problems. There is A LOT of space of improvement:\n* clean up the ugly workaround for new values - seriously, it wouldn't all be necessary if we had sklearn >= 0.24 with its functionality for OrdinalEncoder :-) \n* proper validation\n* speedup the prediction generation\n* tuning the parameters for ALS\n\n\nSmash the upvote if you found this useful.\n","metadata":{}}]}