{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# <p style=\"background-color:#682F2F; font-family:newtimeroman; color:#FFF9ED; font-size:150%; text-align:center; border-radius:10px 10px; padding: 15px; \">Recommend Clothes\n</p>\n\nIn this project, I will make product recommendations for buyers. With data from the contest [H&M Personalized Fashion Recommendations](https://www.kaggle.com/competitions/h-and-m-personalized-fashion-recommendations)\n","metadata":{}},{"cell_type":"markdown","source":"<p style=\"background-color:#682F2F;font-family:newtimeroman;color:#FFF9ED;font-size:150%;text-align:center;border-radius:10px 10px; padding: 10px\">TABLE OF CONTENTS</p>   \n    \n* [1. IMPORTING LIBRARIES](#1)\n    \n* [2. LOADING DATA](#2)\n    \n* [3. REDUCE MEMORY](#3)\n    \n* [4. NUMBER OF PURCHASES OF EACH PRODUCT](#4)   \n    \n* [5. PRODUCTS OFTEN PURCHASED TOGETHER](#5) \n      \n* [6. RECOMMEND LAST WEEK'S MOST POPULAR ITEMS](#6)\n    \n* [7. PREDICT FOR TEST DATA](#7)\n    \n* [8. CONCLUSION](#8)\n    \n* [9. END](#9)\n","metadata":{}},{"cell_type":"markdown","source":"<a id=\"1\"></a>\n# <p style=\"background-color:#682F2F;font-family:newtimeroman;color:#FFF9ED;font-size:150%;text-align:center;border-radius:10px 10px; padding: 10px\">IMPORTING LIBRARIES</p>","metadata":{}},{"cell_type":"code","source":"!pip install cudf","metadata":{"execution":{"iopub.status.busy":"2023-08-26T02:27:45.156182Z","iopub.execute_input":"2023-08-26T02:27:45.156781Z","iopub.status.idle":"2023-08-26T02:28:09.146919Z","shell.execute_reply.started":"2023-08-26T02:27:45.156737Z","shell.execute_reply":"2023-08-26T02:28:09.145271Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import numpy as np\nimport pandas as pd\nimport datetime\nimport matplotlib.pyplot as plt\nfrom matplotlib import colors\nimport seaborn as sns\nimport cudf\nfrom os.path import exists\nimport cv2\nimport gc","metadata":{"execution":{"iopub.status.busy":"2023-08-26T02:28:09.149587Z","iopub.execute_input":"2023-08-26T02:28:09.149995Z","iopub.status.idle":"2023-08-26T02:28:12.684423Z","shell.execute_reply.started":"2023-08-26T02:28:09.149964Z","shell.execute_reply":"2023-08-26T02:28:12.683203Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print('Rapids version:' + cudf.__version__)","metadata":{"execution":{"iopub.status.busy":"2023-08-26T02:28:12.686270Z","iopub.execute_input":"2023-08-26T02:28:12.686703Z","iopub.status.idle":"2023-08-26T02:28:12.695834Z","shell.execute_reply.started":"2023-08-26T02:28:12.686662Z","shell.execute_reply":"2023-08-26T02:28:12.694581Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id=\"2\"></a>\n# <p style=\"background-color:#682F2F;font-family:newtimeroman;color:#FFF9ED;font-size:150%;text-align:center;border-radius:10px 10px; padding: 10px\">LOADING DATA</p>","metadata":{}},{"cell_type":"code","source":"df = cudf.read_csv('/kaggle/input/h-and-m-personalized-fashion-recommendations/transactions_train.csv')\ndf.head()","metadata":{"execution":{"iopub.status.busy":"2023-08-26T02:28:12.699202Z","iopub.execute_input":"2023-08-26T02:28:12.700209Z","iopub.status.idle":"2023-08-26T02:28:54.534987Z","shell.execute_reply.started":"2023-08-26T02:28:12.700166Z","shell.execute_reply":"2023-08-26T02:28:54.533780Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df.info()","metadata":{"execution":{"iopub.status.busy":"2023-08-26T02:28:54.536687Z","iopub.execute_input":"2023-08-26T02:28:54.537640Z","iopub.status.idle":"2023-08-26T02:28:54.564189Z","shell.execute_reply.started":"2023-08-26T02:28:54.537598Z","shell.execute_reply":"2023-08-26T02:28:54.562839Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df.shape","metadata":{"execution":{"iopub.status.busy":"2023-08-26T02:28:54.566404Z","iopub.execute_input":"2023-08-26T02:28:54.566895Z","iopub.status.idle":"2023-08-26T02:28:54.575274Z","shell.execute_reply.started":"2023-08-26T02:28:54.566834Z","shell.execute_reply":"2023-08-26T02:28:54.573836Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id=\"3\"></a>\n# <p style=\"background-color:#682F2F;font-family:newtimeroman;color:#FFF9ED;font-size:150%;text-align:center;border-radius:10px 10px; padding: 10px\">REDUCE MEMORY</p>","metadata":{}},{"cell_type":"markdown","source":"The `customer_id` is string and it use 64 bytes to store.This code below will covert this column to `int64` which mean it only use 8 bytes. Also the mapping is 1:1 which means each customer gets a unique `int64`. Therefore the memory you need to save the above df will be reduced by 8 times. Because we only use columns: `customer_id`, `article_id` , `t_dat`, we will delete the other columns.","metadata":{}},{"cell_type":"code","source":"df['customer_id'] = df['customer_id'].str[-16:].str.hex_to_int().astype('int64')\ndf['article_id'] = df.article_id.astype('int32')\ndf.t_dat = cudf.to_datetime(df.t_dat)\ndf = df[['t_dat','customer_id','article_id']]\n_ = gc.collect()","metadata":{"execution":{"iopub.status.busy":"2023-08-26T02:28:54.577636Z","iopub.execute_input":"2023-08-26T02:28:54.578131Z","iopub.status.idle":"2023-08-26T02:28:55.176476Z","shell.execute_reply.started":"2023-08-26T02:28:54.578064Z","shell.execute_reply":"2023-08-26T02:28:55.175269Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id=\"4\"></a>\n# <p style=\"background-color:#682F2F;font-family:newtimeroman;color:#FFF9ED;font-size:150%;text-align:center;border-radius:10px 10px; padding: 10px\">NUMBER OF PURCHASES OF EACH PRODUCT</p>","metadata":{}},{"cell_type":"code","source":"temp = df.groupby('article_id')['t_dat'].agg('count').reset_index()\ntemp.columns = ['article_id', 'count']\ntemp = temp.sort_values('count', ascending=False)\ntemp.head()","metadata":{"execution":{"iopub.status.busy":"2023-08-26T02:28:55.178237Z","iopub.execute_input":"2023-08-26T02:28:55.178958Z","iopub.status.idle":"2023-08-26T02:28:55.256719Z","shell.execute_reply.started":"2023-08-26T02:28:55.178914Z","shell.execute_reply":"2023-08-26T02:28:55.255197Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#GET TOP 20 product\ntemp = temp[:20]\ntemp['article_id'] = temp['article_id'].astype('str')","metadata":{"execution":{"iopub.status.busy":"2023-08-26T02:28:55.259079Z","iopub.execute_input":"2023-08-26T02:28:55.259596Z","iopub.status.idle":"2023-08-26T02:28:55.269430Z","shell.execute_reply.started":"2023-08-26T02:28:55.259552Z","shell.execute_reply":"2023-08-26T02:28:55.267753Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#tOP 20 PRODUCT \nBASE = '../input/h-and-m-personalized-fashion-recommendations/images/'\nfor i in range(0, 20, 4):\n    name1 = BASE+'0'+temp.iloc[i, 0][:2]+'/0'+temp.iloc[i, 0]+'.jpg'\n    name2 = BASE+'0'+temp.iloc[i + 1, 0][:2]+'/0'+temp.iloc[i + 1, 0]+'.jpg'\n    name3 = BASE+'0'+temp.iloc[i + 2, 0][:2]+'/0'+temp.iloc[i + 2, 0]+'.jpg'\n    name4 = BASE+'0'+temp.iloc[i + 3, 0][:2]+'/0'+temp.iloc[i + 3, 0]+'.jpg'\n#     if exists(name1) & exists(name2) & exists(name3) & exists(name4):\n    plt.figure(figsize=(20,5))\n    \n    plt.subplot(1,4,1)\n    plt.title( f\"{temp.iloc[i, 0]}: {temp.iloc[i, 1]}\",size=14)\n    plt.gca().set_axis_off()\n    if exists(name1):\n        img1 = cv2.imread(name1)[:,:,::-1]\n        plt.imshow(img1)\n        \n        \n    plt.subplot(1,4,2)\n    plt.title(f\"{temp.iloc[i + 1, 0]}: {temp.iloc[i+1, 1]}\",size=14)\n    plt.gca().set_axis_off()\n    if exists(name2):        \n        img2 = cv2.imread(name2)[:,:,::-1]\n        plt.imshow(img2)\n        \n        \n    plt.subplot(1,4,3)\n    plt.title(f\"{temp.iloc[i+2, 0]}: {temp.iloc[i+2, 1]}\",size=14)\n    plt.gca().set_axis_off()\n    if exists(name3):                \n        img3 = cv2.imread(name3)[:,:,::-1]\n        plt.imshow(img3)\n        \n    plt.subplot(1,4,4)\n    plt.title(f\"{temp.iloc[i+3, 0]}: {temp.iloc[i+3, 1]}\",size=14)\n    plt.gca().set_axis_off()\n    if exists(name4):                    \n        img4 = cv2.imread(name4)[:,:,::-1]\n        plt.imshow(img4)\n        \n    plt.show()\n#     else:\n#         print('False')","metadata":{"execution":{"iopub.status.busy":"2023-08-26T02:28:55.276303Z","iopub.execute_input":"2023-08-26T02:28:55.276748Z","iopub.status.idle":"2023-08-26T02:29:04.696454Z","shell.execute_reply.started":"2023-08-26T02:28:55.276715Z","shell.execute_reply":"2023-08-26T02:29:04.695318Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id=\"5\"></a>\n# <p style=\"background-color:#682F2F;font-family:newtimeroman;color:#FFF9ED;font-size:150%;text-align:center;border-radius:10px 10px; padding: 10px\">PRODUCTS OFTEN PURCHASED TOGETHER</p>","metadata":{}},{"cell_type":"markdown","source":"We will find product pairs where, if a user buys product A, it is more likely to buy product B.\n\n**Idea:**\nWe will take out the products that are bought by the most people. Then for each product we find the customers who bought it. With the above list of customers, we will list other products that these people also buy. Here we only take 1 product with the most buyers.","metadata":{}},{"cell_type":"markdown","source":"Because we only need to know the number of people buying a product instead of the number of a product purchased, we will delete the lines with the same `customer_id` and `article_id`.","metadata":{}},{"cell_type":"code","source":"df = df.drop_duplicates(['customer_id', 'article_id'])\ndf.shape","metadata":{"execution":{"iopub.status.busy":"2023-08-26T02:29:04.697962Z","iopub.execute_input":"2023-08-26T02:29:04.699157Z","iopub.status.idle":"2023-08-26T02:29:04.825583Z","shell.execute_reply.started":"2023-08-26T02:29:04.699104Z","shell.execute_reply":"2023-08-26T02:29:04.824231Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# mask = df.article_id.value_counts() >= 50\n# article_ndelete = df.article_id.value_counts()[mask].index\n# # article_delete\n# df = df.loc[df.article_id.isin(article_ndelete),:]\n# df.shape","metadata":{"execution":{"iopub.status.busy":"2023-08-26T02:29:04.827833Z","iopub.execute_input":"2023-08-26T02:29:04.828325Z","iopub.status.idle":"2023-08-26T02:29:04.834048Z","shell.execute_reply.started":"2023-08-26T02:29:04.828281Z","shell.execute_reply":"2023-08-26T02:29:04.832714Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Create `pairs_t` is a copy of `df`.","metadata":{}},{"cell_type":"code","source":"df = df[['customer_id','article_id']]\npairs_t = df.copy()\npairs_t.columns = ['customer_id','pair_id']","metadata":{"execution":{"iopub.status.busy":"2023-08-26T02:29:04.836108Z","iopub.execute_input":"2023-08-26T02:29:04.837043Z","iopub.status.idle":"2023-08-26T02:29:04.860293Z","shell.execute_reply.started":"2023-08-26T02:29:04.836998Z","shell.execute_reply":"2023-08-26T02:29:04.858806Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Get article_id and number customer have bought it. Except article which have been bought by only one customer.","metadata":{}},{"cell_type":"code","source":"gr = df.groupby('article_id')['customer_id'].nunique()  \nunique_articles = gr[gr != 1].index","metadata":{"execution":{"iopub.status.busy":"2023-08-26T02:29:04.862415Z","iopub.execute_input":"2023-08-26T02:29:04.862917Z","iopub.status.idle":"2023-08-26T02:29:05.276963Z","shell.execute_reply.started":"2023-08-26T02:29:04.862875Z","shell.execute_reply":"2023-08-26T02:29:05.275660Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"With `unique_articles` being the list of article_ids, we will traverse in batches. For each batch, we will find out who has purchased the products in this batch. Then combine with `pairs_t` based on the `customer_id` column to create a new DataFrame that will contain the product pairs that this person has purchased. Now we will group the product pairs and calculate the number of times they are purchased together. Finally, for each product in the batch under consideration, we will find a proposed product to go with.","metadata":{}},{"cell_type":"code","source":"%%time\n\nbatch_size = 5000\n\nbatch_pairs_dfs = []\n\nfor i in range(0, len(unique_articles), batch_size):\n    print(f\"processing article #{i:,} to #{i+batch_size:,}\")\n    \n\n    # take batch of articles\n    batch_articles = unique_articles[i:i+batch_size]\n\n    # get all pairs for those articles (other articles those customers bought)\n    batch_t = df[df[\"article_id\"].isin(batch_articles)]\n    batch_pairs_df = batch_t.merge(pairs_t, on=\"customer_id\")\n\n    \n    # delete same-article pairs\n    same_article_row_idxs = batch_pairs_df.query(\"article_id==pair_id\").index\n    batch_pairs_df = batch_pairs_df.drop(same_article_row_idxs)\n    \n    # delete single customer articles\n    c1s = (\n        batch_pairs_df.groupby(\"article_id\")[[\"customer_id\"]].nunique()\n        .query(\"customer_id==1\").index\n    )\n    single_customer_row_idxs = batch_pairs_df[batch_pairs_df[\"article_id\"].isin(c1s)].index\n    batch_pairs_df = batch_pairs_df.drop(single_customer_row_idxs)\n    \n    # get sorted counts of article-pair occurences\n    batch_pairs_df = batch_pairs_df.groupby([\"article_id\", \"pair_id\"])[[\"customer_id\"]].count()\n    batch_pairs_df.columns = [\"pair_counts\"]\n    batch_pairs_df = batch_pairs_df.reset_index()\n    batch_pairs_df = batch_pairs_df.sort_values([\"article_id\", \"pair_counts\"], ascending=False)\n\n    # get top one for each article (need pandas)\n    batch_pairs_df = batch_pairs_df.to_pandas().groupby(\"article_id\").head(1)\n    # back to cudf\n    batch_pairs_df = cudf.DataFrame(batch_pairs_df.set_index(\"article_id\")[[\"pair_id\"]])\n    batch_pairs_dfs.append(batch_pairs_df)\n\n    \nall_article_pairs_df = cudf.concat(batch_pairs_dfs)\nprint(len(all_article_pairs_df))","metadata":{"execution":{"iopub.status.busy":"2023-08-26T02:29:05.278399Z","iopub.execute_input":"2023-08-26T02:29:05.285149Z","iopub.status.idle":"2023-08-26T02:31:52.539762Z","shell.execute_reply.started":"2023-08-26T02:29:05.285087Z","shell.execute_reply":"2023-08-26T02:31:52.538609Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Find pairs of products that go together.","metadata":{}},{"cell_type":"code","source":"# pairs = {}\n# for index, row in enumerate(article_seri.index.values):\n#     user = df.loc[df['article_id'] == row.item(), 'customer_id'].unique()\n#     mask = (df.customer_id.isin(user)) & (df.article_id != row.item())\n#     article_together = df.loc[mask, 'article_id'].value_counts()\n#     del mask\n    \n#     pairs[row.item()] = article_together.index[0]\n#     del article_together\n#     _ = gc.collect()","metadata":{"execution":{"iopub.status.busy":"2023-08-26T02:31:52.541540Z","iopub.execute_input":"2023-08-26T02:31:52.542259Z","iopub.status.idle":"2023-08-26T02:31:52.547377Z","shell.execute_reply.started":"2023-08-26T02:31:52.542220Z","shell.execute_reply":"2023-08-26T02:31:52.546100Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"pairs = all_article_pairs_df[\"pair_id\"].to_pandas().to_dict()","metadata":{"execution":{"iopub.status.busy":"2023-08-26T02:31:52.549090Z","iopub.execute_input":"2023-08-26T02:31:52.549803Z","iopub.status.idle":"2023-08-26T02:31:52.685170Z","shell.execute_reply.started":"2023-08-26T02:31:52.549764Z","shell.execute_reply":"2023-08-26T02:31:52.683939Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Predict products customers will buy with train data.","metadata":{}},{"cell_type":"code","source":"del df\n_ = gc.collect()\ndf = cudf.read_csv('/kaggle/input/h-and-m-personalized-fashion-recommendations/transactions_train.csv'\n                  ,usecols=[\"customer_id\", \"article_id\"])\ndf['customer_id'] = df['customer_id'].str[-16:].str.hex_to_int().astype('int64')\ndf['article_id'] = df.article_id.astype('int32')\n\ndf = df.to_pandas()\ndf['article_tog'] = df.article_id.map(pairs)\n_ = gc.collect()","metadata":{"execution":{"iopub.status.busy":"2023-08-26T02:31:52.686745Z","iopub.execute_input":"2023-08-26T02:31:52.687493Z","iopub.status.idle":"2023-08-26T02:31:57.707209Z","shell.execute_reply.started":"2023-08-26T02:31:52.687453Z","shell.execute_reply":"2023-08-26T02:31:57.705995Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"pred = df[['customer_id', 'article_tog']]\npred = pred.loc[pred.article_tog.notnull()]\npred = pred.drop_duplicates()\npred = pred.rename({'article_tog' : 'article_id'}, axis = 1)","metadata":{"execution":{"iopub.status.busy":"2023-08-26T02:31:57.708770Z","iopub.execute_input":"2023-08-26T02:31:57.709172Z","iopub.status.idle":"2023-08-26T02:32:06.061728Z","shell.execute_reply.started":"2023-08-26T02:31:57.709135Z","shell.execute_reply":"2023-08-26T02:32:06.060105Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"pred1 = df[['customer_id', 'article_id']].copy()\npred = pd.concat([pred, pred1], axis=0, ignore_index=True)\npred.article_id = pred.article_id.astype('int32')\n\ndel pred1\ngc.collect()\npred.head()","metadata":{"execution":{"iopub.status.busy":"2023-08-26T02:32:06.063257Z","iopub.execute_input":"2023-08-26T02:32:06.063655Z","iopub.status.idle":"2023-08-26T02:32:07.273255Z","shell.execute_reply.started":"2023-08-26T02:32:06.063619Z","shell.execute_reply":"2023-08-26T02:32:07.271928Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Bring the predictions for each customer into a string.","metadata":{}},{"cell_type":"code","source":"pred.article_id = ' 0' + pred.article_id.astype('str')\npred = cudf.DataFrame( pred.groupby('customer_id').article_id.sum().reset_index() )\npred.columns = ['customer_id','prediction']\n\n_ = gc.collect()\npred.head()","metadata":{"execution":{"iopub.status.busy":"2023-08-26T02:32:07.275204Z","iopub.execute_input":"2023-08-26T02:32:07.275972Z","iopub.status.idle":"2023-08-26T02:33:34.201087Z","shell.execute_reply.started":"2023-08-26T02:32:07.275929Z","shell.execute_reply":"2023-08-26T02:33:34.199916Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id=\"6\"></a>\n# <p style=\"background-color:#682F2F;font-family:newtimeroman;color:#FFF9ED;font-size:150%;text-align:center;border-radius:10px 10px; padding: 10px\">RECOMMEND LAST WEEK'S MOST POPULAR ITEMS</p>","metadata":{}},{"cell_type":"code","source":"# df = cudf.DataFrame(df)\ndel df\n_ = gc.collect()\ndf = cudf.read_csv('/kaggle/input/h-and-m-personalized-fashion-recommendations/transactions_train.csv')\n_ = gc.collect()\ndf.t_dat = cudf.to_datetime(df.t_dat)\ndf = df.sort_values(by='t_dat', ascending = False)\ndf.head()","metadata":{"execution":{"iopub.status.busy":"2023-08-26T02:33:34.202780Z","iopub.execute_input":"2023-08-26T02:33:34.203230Z","iopub.status.idle":"2023-08-26T02:33:46.502533Z","shell.execute_reply.started":"2023-08-26T02:33:34.203195Z","shell.execute_reply":"2023-08-26T02:33:46.501312Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Last date is: 2020-09-22. We will find the most purchased products in the 1 week before this date.","metadata":{}},{"cell_type":"code","source":"#top 12 products\ntop = df.loc[df.t_dat >= cudf.to_datetime('2020-09-16'),'article_id'].value_counts().to_pandas().index.astype('str')\ntop = top[:12]\nres = ' 0' + ' 0'.join(top)\nres","metadata":{"execution":{"iopub.status.busy":"2023-08-26T02:33:46.504219Z","iopub.execute_input":"2023-08-26T02:33:46.504679Z","iopub.status.idle":"2023-08-26T02:33:46.582537Z","shell.execute_reply.started":"2023-08-26T02:33:46.504648Z","shell.execute_reply":"2023-08-26T02:33:46.581273Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id=\"7\"></a>\n# <p style=\"background-color:#682F2F;font-family:newtimeroman;color:#FFF9ED;font-size:150%;text-align:center;border-radius:10px 10px; padding: 10px\">PREDICT FOR TEST DATA</p>","metadata":{}},{"cell_type":"code","source":"sub = cudf.read_csv('/kaggle/input/h-and-m-personalized-fashion-recommendations/sample_submission.csv')\nsub = sub[['customer_id']]\nsub['customer_id2'] = sub.customer_id.str[-16:].str.hex_to_int().astype('int64')\n\npred = pred.rename({'customer_id':'customer_id2', 'article_id':'prediction'},axis=1)\nsub = sub.merge(pred,on='customer_id2', how='left').fillna('')\nsub = sub.drop(columns=['customer_id2'], axis = 1)\n\nsub.prediction = sub.prediction + res\nsub.prediction = sub.prediction.str.strip()\nsub.prediction = sub.prediction.str[:131]\nsub","metadata":{"execution":{"iopub.status.busy":"2023-08-26T02:33:46.584540Z","iopub.execute_input":"2023-08-26T02:33:46.584961Z","iopub.status.idle":"2023-08-26T02:33:50.456682Z","shell.execute_reply.started":"2023-08-26T02:33:46.584924Z","shell.execute_reply":"2023-08-26T02:33:50.455492Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Save result to file","metadata":{}},{"cell_type":"code","source":"sub.to_csv(f'submission.csv',index=False)","metadata":{"execution":{"iopub.status.busy":"2023-08-26T02:33:50.458311Z","iopub.execute_input":"2023-08-26T02:33:50.459252Z","iopub.status.idle":"2023-08-26T02:33:51.263476Z","shell.execute_reply.started":"2023-08-26T02:33:50.459212Z","shell.execute_reply":"2023-08-26T02:33:51.262026Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id=\"8\"></a>\n# <p style=\"background-color:#682F2F;font-family:newtimeroman;color:#FFF9ED;font-size:150%;text-align:center;border-radius:10px 10px; padding: 10px\">CONCLUSION</p>","metadata":{}},{"cell_type":"markdown","source":"In this project I made product recommendations to clients. Besides, there are statistics on data such as the top best-selling products, the most purchased products last week.\n\n<span style=\"color:#682F2F; font-weight:500\"> If you have any questions, don't hesitate to comment. </span>","metadata":{}},{"cell_type":"markdown","source":"<a id=\"9\"></a>\n# <p style=\"background-color:#682F2F;font-family:newtimeroman;color:#FFF9ED;font-size:150%;text-align:center;border-radius:10px 10px; padding: 10px\">END</p>","metadata":{}},{"cell_type":"markdown","source":"**Reference:\n- [Customers Who Bought This Frequently Buy This!](https://www.kaggle.com/code/cdeotte/customers-who-bought-this-frequently-buy-this/notebook)\n- [Customer Segmentation: Clustering 🛍️🛒🛒](https://www.kaggle.com/code/karnikakapoor/customer-segmentation-clustering)\n- [Recommend Items Purchased Together - [0.021]](https://www.kaggle.com/code/cdeotte/recommend-items-purchased-together-0-021/notebook)","metadata":{}}]}