{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"## Now Generating recommendations based on this strategy as well as combining 2 more strategies\n\nCombining the following strategies :\n\n1) Previously purchased items\n\n2) items purchased together \n\n3) Last weeks most popular items\n\n## Strategy explained.\n\nBasically what we are doing here is finding out the customer's last week's purchases, and then generating 12 recommendations for that user by recommending articles using previously purchased items, items purchased together, and last weeks most popular items in that order of priority.\n\nfirst we would recommend them the most repurchased articles by the customer, then articles similar to most repurchased articles (customers also bought logic) and if still there are not 12 articles for the customer, select the most popular articles from that week (ordered based on popularity, and then predict those as a recommendation for the customer)\n\n\n#### Recommend Items Frequently Purchased Together\n\nThis notebook attempts to create a baseline recommender system by using the hypothesis that items purchased together, historic purchases of the user and most popular items together form a crude but effective recommender system.","metadata":{}},{"cell_type":"markdown","source":"## Accelrating computation using cuDF","metadata":{}},{"cell_type":"code","source":"import cudf\nprint('RAPIDS version',cudf.__version__)","metadata":{"execution":{"iopub.status.busy":"2022-09-08T23:45:40.840310Z","iopub.execute_input":"2022-09-08T23:45:40.840773Z","iopub.status.idle":"2022-09-08T23:45:42.227268Z","shell.execute_reply.started":"2022-09-08T23:45:40.840702Z","shell.execute_reply":"2022-09-08T23:45:42.225743Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Loading transactions data and reducing the memory usage through manual typsetting\n","metadata":{}},{"cell_type":"code","source":"train = cudf.read_csv('../input/h-and-m-personalized-fashion-recommendations/transactions_train.csv')\ntrain['customer_id'] = train['customer_id'].str[-16:].str.hex_to_int().astype('int64')\ntrain['article_id'] = train.article_id.astype('int32')\ntrain.t_dat = cudf.to_datetime(train.t_dat)\ntrain = train[['t_dat','customer_id','article_id']]\ntrain.to_parquet('train.pqt',index=False)\nprint( train.shape )\ntrain.head()","metadata":{"execution":{"iopub.status.busy":"2022-09-08T23:45:42.232471Z","iopub.execute_input":"2022-09-08T23:45:42.232980Z","iopub.status.idle":"2022-09-08T23:45:47.033023Z","shell.execute_reply.started":"2022-09-08T23:45:42.232942Z","shell.execute_reply":"2022-09-08T23:45:47.030892Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Finding each customer's last week's purchases, \n\ni.e finding out which week was the last week of purchases for each customer","metadata":{}},{"cell_type":"code","source":"tmp = train.groupby('customer_id').t_dat.max().reset_index()\ntmp.columns = ['customer_id','max_dat']\ntrain = train.merge(tmp,on=['customer_id'],how='left')\ntrain['diff_dat'] = (train.max_dat - train.t_dat).dt.days\ntrain = train.loc[train['diff_dat']<=6]\nprint('Train shape:',train.shape)\ntrain.head(20)","metadata":{"execution":{"iopub.status.busy":"2022-09-08T23:45:47.034221Z","iopub.execute_input":"2022-09-08T23:45:47.034563Z","iopub.status.idle":"2022-09-08T23:45:47.241049Z","shell.execute_reply.started":"2022-09-08T23:45:47.034525Z","shell.execute_reply":"2022-09-08T23:45:47.240347Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Creating Recommendation strategies for most often previously purchased items\n\nBasically grouping the data by customer id and article id, and then aggregating the same. (we'd get a count of the most repurchased articles for each customer after sorting the same in descending order)","metadata":{}},{"cell_type":"code","source":"tmp = train.groupby(['customer_id','article_id'])['t_dat'].agg('count').reset_index()\ntmp.columns = ['customer_id','article_id','ct']\ntrain = train.merge(tmp,on=['customer_id','article_id'],how='left')\ntrain = train.sort_values(['ct','t_dat'],ascending=False)\ntrain = train.drop_duplicates(['customer_id','article_id'])\ntrain = train.sort_values(['ct','t_dat'],ascending=False)\ntrain.head(20)","metadata":{"execution":{"iopub.status.busy":"2022-09-08T23:45:50.406018Z","iopub.execute_input":"2022-09-08T23:45:50.406852Z","iopub.status.idle":"2022-09-08T23:45:50.663541Z","shell.execute_reply.started":"2022-09-08T23:45:50.406776Z","shell.execute_reply":"2022-09-08T23:45:50.662580Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## creating recommendations for items purchased together\n\ncomputing a dictionary of items frequently purchased together (as done above to create the visuals and dropping the duplicate articles to not recommend articles user has already bought and we've already recommended.","metadata":{}},{"cell_type":"markdown","source":"Note that we use the command `drop_duplicates` so that we don't recommend an item that the user has already bought and we have already recommended above.\n\nWe will need to use Pandas for some commands because RAPIDS cuDF doesn't have two conveinent commands, (1) create new column from dictionary map of another column (2) groupby aggregate strings sum.\n\nWe concatenate these rows after the rows containing customers' previous purchases. Therefore we will recommend previous items first and then items purchased together second. Note the trick to convert a column of int32 into a prediction string (using groupby agg str sum) is from notebook.","metadata":{}},{"cell_type":"code","source":"# USE PANDAS TO MAP COLUMN WITH DICTIONARY\nimport pandas as pd, numpy as np\ntrain = train.to_pandas()\npairs = np.load('../input/hmitempairs/pairs_cudf.npy',allow_pickle=True).item()\ntrain['article_id2'] = train.article_id.map(pairs)\n\ntrain.head(20)","metadata":{"execution":{"iopub.status.busy":"2022-09-08T23:52:02.609409Z","iopub.execute_input":"2022-09-08T23:52:02.610213Z","iopub.status.idle":"2022-09-08T23:52:02.846644Z","shell.execute_reply.started":"2022-09-08T23:52:02.610172Z","shell.execute_reply":"2022-09-08T23:52:02.845763Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# RECOMMENDATION OF PAIRED ITEMS\ntrain2 = train[['customer_id','article_id2']].copy()\ntrain2 = train2.loc[train2.article_id2.notnull()]\n\ntrain2 = train2.drop_duplicates(['customer_id','article_id2'])\ntrain2 = train2.rename({'article_id2':'article_id'},axis=1)\n\ntrain2","metadata":{"execution":{"iopub.status.busy":"2022-09-08T23:52:29.557835Z","iopub.execute_input":"2022-09-08T23:52:29.558488Z","iopub.status.idle":"2022-09-08T23:52:30.844388Z","shell.execute_reply.started":"2022-09-08T23:52:29.558448Z","shell.execute_reply":"2022-09-08T23:52:30.843674Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# CONCATENATE PAIRED ITEM RECOMMENDATION AFTER PREVIOUS PURCHASED RECOMMENDATIONS\ntrain = train[['customer_id','article_id']]\ntrain = pd.concat([train,train2],axis=0,ignore_index=True)\ntrain.article_id = train.article_id.astype('int32')\ntrain = train.drop_duplicates(['customer_id','article_id'])\ntrain.head(10)","metadata":{"execution":{"iopub.status.busy":"2022-09-09T00:00:57.688655Z","iopub.execute_input":"2022-09-09T00:00:57.688954Z","iopub.status.idle":"2022-09-09T00:00:59.630591Z","shell.execute_reply.started":"2022-09-09T00:00:57.688918Z","shell.execute_reply":"2022-09-09T00:00:59.629893Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# CONVERT RECOMMENDATIONS INTO SINGLE STRING\ntrain.article_id = ' 0' + train.article_id.astype('str')\npreds = cudf.DataFrame( train.groupby('customer_id').article_id.sum().reset_index() )\npreds.columns = ['customer_id','prediction']\npreds.head(10)","metadata":{"execution":{"iopub.status.busy":"2022-09-09T00:01:03.427648Z","iopub.execute_input":"2022-09-09T00:01:03.427926Z","iopub.status.idle":"2022-09-09T00:01:12.465407Z","shell.execute_reply.started":"2022-09-09T00:01:03.427896Z","shell.execute_reply":"2022-09-09T00:01:12.464691Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# (3) Recommend Last Week's Most Popular Items\nAfter recommending previous purchases and items purchased together we will then recommend the 12 most popular items. Therefore if our previous recommendations did not fill up a customer's 12 recommendations, then it will be filled by popular items.","metadata":{}},{"cell_type":"code","source":"train = cudf.read_parquet('train.pqt')\ntrain.t_dat = cudf.to_datetime(train.t_dat)\ntrain = train.loc[train.t_dat >= cudf.to_datetime('2020-09-16')]\ntop12 = ' 0' + ' 0'.join(train.article_id.value_counts().to_pandas().index.astype('str')[:12])\nprint(\"Last week's top 12 popular items:\")\nprint( top12 )","metadata":{"execution":{"iopub.status.busy":"2022-09-09T00:01:12.466997Z","iopub.execute_input":"2022-09-09T00:01:12.467231Z","iopub.status.idle":"2022-09-09T00:01:12.823401Z","shell.execute_reply.started":"2022-09-09T00:01:12.467198Z","shell.execute_reply":"2022-09-09T00:01:12.822514Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Write Submission CSV\nWe will merge our predictions onto `sample_submission.csv` and submit to Kaggle.","metadata":{}},{"cell_type":"code","source":"sub = cudf.read_csv('../input/h-and-m-personalized-fashion-recommendations/sample_submission.csv')\nsub = sub[['customer_id']]\nsub['customer_id_2'] = sub['customer_id'].str[-16:].str.hex_to_int().astype('int64')\nsub = sub.merge(preds.rename({'customer_id':'customer_id_2'},axis=1),\\\n    on='customer_id_2', how='left').fillna('')\ndel sub['customer_id_2']\nsub.prediction = sub.prediction + top12\nsub.prediction = sub.prediction.str.strip()\nsub.prediction = sub.prediction.str[:131]\nsub.to_csv(f'submission.csv',index=False)\nsub.head()","metadata":{"execution":{"iopub.status.busy":"2022-09-09T00:01:15.868361Z","iopub.execute_input":"2022-09-09T00:01:15.868640Z","iopub.status.idle":"2022-09-09T00:01:21.645277Z","shell.execute_reply.started":"2022-09-09T00:01:15.868609Z","shell.execute_reply":"2022-09-09T00:01:21.644490Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# prediction score [0.021]","metadata":{}}]}