{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# H&M Recommender","metadata":{}},{"cell_type":"markdown","source":"In this competition, we must develop product recommendations based on data from previous transactions, as well as from customer and product meta data. Following, we will import necessay libraries and read the data to see what can we do with it.","metadata":{}},{"cell_type":"code","source":"import pandas as pd","metadata":{"execution":{"iopub.status.busy":"2022-03-03T14:58:51.423443Z","iopub.execute_input":"2022-03-03T14:58:51.424515Z","iopub.status.idle":"2022-03-03T14:58:51.453293Z","shell.execute_reply.started":"2022-03-03T14:58:51.424385Z","shell.execute_reply":"2022-03-03T14:58:51.452514Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Path depends on where data is\ndata_path = '../input/h-and-m-personalized-fashion-recommendations/'","metadata":{"execution":{"iopub.status.busy":"2022-03-03T14:58:51.454725Z","iopub.execute_input":"2022-03-03T14:58:51.455633Z","iopub.status.idle":"2022-03-03T14:58:51.460297Z","shell.execute_reply.started":"2022-03-03T14:58:51.455597Z","shell.execute_reply":"2022-03-03T14:58:51.459407Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Importing and Reading Data","metadata":{}},{"cell_type":"markdown","source":"- Transaction Data","metadata":{}},{"cell_type":"code","source":"df_trans = pd.read_csv(data_path+\"transactions_train.csv\")\ndf_trans.head()","metadata":{"execution":{"iopub.status.busy":"2022-03-03T14:58:51.48561Z","iopub.execute_input":"2022-03-03T14:58:51.48659Z","iopub.status.idle":"2022-03-03T14:59:34.591628Z","shell.execute_reply.started":"2022-03-03T14:58:51.486543Z","shell.execute_reply":"2022-03-03T14:59:34.588183Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"This will be our main training data. We can play around with dates and articles to see which articles are trending at the moment","metadata":{}},{"cell_type":"markdown","source":"- Articles metadata","metadata":{}},{"cell_type":"code","source":"df_articles = pd.read_csv(data_path+\"articles.csv\")\ndf_articles.head()","metadata":{"execution":{"iopub.status.busy":"2022-03-03T14:59:34.592678Z","iopub.status.idle":"2022-03-03T14:59:34.593243Z","shell.execute_reply.started":"2022-03-03T14:59:34.59303Z","shell.execute_reply":"2022-03-03T14:59:34.593056Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_articles[\"index_group_name\"].value_counts()","metadata":{"execution":{"iopub.status.busy":"2022-03-03T14:59:34.594789Z","iopub.status.idle":"2022-03-03T14:59:34.595509Z","shell.execute_reply.started":"2022-03-03T14:59:34.595266Z","shell.execute_reply":"2022-03-03T14:59:34.595293Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We can see some features might be usefull like index_group_name for detecting if the article is meant for women or men or baby, etc. This could help for recommending articles with same index_group_name to customers who already bought something in that group (sort of detecting its sex)","metadata":{}},{"cell_type":"markdown","source":"- Customer Data","metadata":{}},{"cell_type":"code","source":"df_customers = pd.read_csv(data_path+\"customers.csv\")\ndf_customers.head()","metadata":{"execution":{"iopub.status.busy":"2022-03-03T14:59:34.596927Z","iopub.status.idle":"2022-03-03T14:59:34.597272Z","shell.execute_reply.started":"2022-03-03T14:59:34.597084Z","shell.execute_reply":"2022-03-03T14:59:34.59711Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Don't see anything that might be usefull to increase performance","metadata":{}},{"cell_type":"markdown","source":"## 1. Setup training data","metadata":{}},{"cell_type":"markdown","source":"We will use df_trans as our main training set following these steps:\n- Format t_dat column as datetime to be able to sort and filter by date\n- keep only important columns","metadata":{}},{"cell_type":"code","source":"df_trans[\"t_dat\"] = pd.to_datetime(df_trans[\"t_dat\"])\ndf_trans = df_trans[['t_dat','customer_id','article_id']]\ndf_trans.head()","metadata":{"execution":{"iopub.status.busy":"2022-03-03T14:59:34.598516Z","iopub.status.idle":"2022-03-03T14:59:34.59885Z","shell.execute_reply.started":"2022-03-03T14:59:34.598668Z","shell.execute_reply":"2022-03-03T14:59:34.598695Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Now, we will add index_group_name feature from df_articles to our training set","metadata":{}},{"cell_type":"code","source":"product_merge = df_trans.merge(df_articles,on=['article_id'],how='left')\nproduct_merge = product_merge[['t_dat','customer_id','article_id','index_group_name']]\nproduct_merge.head()","metadata":{"execution":{"iopub.status.busy":"2022-03-03T14:59:34.600036Z","iopub.status.idle":"2022-03-03T14:59:34.600409Z","shell.execute_reply.started":"2022-03-03T14:59:34.600191Z","shell.execute_reply":"2022-03-03T14:59:34.600214Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Find customers SEX!","metadata":{}},{"cell_type":"markdown","source":"Finding customer sex will be difficult but with the table we have above, we can see the transaction quantity by index_group_name","metadata":{}},{"cell_type":"code","source":"product_merge[\"index_group_name\"].value_counts()","metadata":{"execution":{"iopub.status.busy":"2022-03-03T14:59:34.601501Z","iopub.status.idle":"2022-03-03T14:59:34.601827Z","shell.execute_reply.started":"2022-03-03T14:59:34.601653Z","shell.execute_reply":"2022-03-03T14:59:34.601676Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(\"Ladieswear Transactions: {:.0%}\".format(product_merge[\"index_group_name\"].value_counts()[0]/product_merge[\"index_group_name\"].value_counts().sum()))\nprint(\"Divided Transactions: {:.0%}\".format(product_merge[\"index_group_name\"].value_counts()[1]/product_merge[\"index_group_name\"].value_counts().sum()))","metadata":{"execution":{"iopub.status.busy":"2022-03-03T14:59:34.603382Z","iopub.status.idle":"2022-03-03T14:59:34.603903Z","shell.execute_reply.started":"2022-03-03T14:59:34.603711Z","shell.execute_reply":"2022-03-03T14:59:34.60374Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"- We can see that Ladieswear transaction account for 64% of all transactions! On the other hand, \"Divided\" transactions (accessories or shoes), which has the second highest quantity, account for only 22%. Since Ladieswear is the most common transaction by far, we will split our our customer base by \"female\" users and \"others\". ","metadata":{}},{"cell_type":"markdown","source":"To do this, we first need to group all the purchases a customer did by index_group_name. We will assume that if the customer purchased a Ladiesware article, the customer is a female","metadata":{}},{"cell_type":"code","source":"sex_user = pd.DataFrame(product_merge.groupby(['customer_id'])['index_group_name'].apply(list)).reset_index()\nsex_user.head()","metadata":{"execution":{"iopub.status.busy":"2022-03-03T14:59:34.605088Z","iopub.status.idle":"2022-03-03T14:59:34.605589Z","shell.execute_reply.started":"2022-03-03T14:59:34.605382Z","shell.execute_reply":"2022-03-03T14:59:34.605409Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sex_user = sex_user.explode('index_group_name')\nsex_user = sex_user.loc[sex_user[\"index_group_name\"]==\"Ladieswear\"]\nsex_user = sex_user.drop_duplicates(['customer_id'])\nsex_user.head()","metadata":{"execution":{"iopub.status.busy":"2022-03-03T14:59:34.606875Z","iopub.status.idle":"2022-03-03T14:59:34.607497Z","shell.execute_reply.started":"2022-03-03T14:59:34.607267Z","shell.execute_reply":"2022-03-03T14:59:34.607293Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"This final sex_user table are the customers which are catalogued as female as part of our test. Next, We can then see how many female are in our data set and how it compares with all our customers","metadata":{}},{"cell_type":"code","source":"sex_user[\"customer_id\"].value_counts().sum()","metadata":{"execution":{"iopub.status.busy":"2022-03-03T14:59:34.608993Z","iopub.status.idle":"2022-03-03T14:59:34.609357Z","shell.execute_reply.started":"2022-03-03T14:59:34.609147Z","shell.execute_reply":"2022-03-03T14:59:34.609171Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(\"Female Customers: {:.0%}\".format(sex_user[\"customer_id\"].value_counts().sum()/df_customers[\"customer_id\"].value_counts().sum()))","metadata":{"execution":{"iopub.status.busy":"2022-03-03T14:59:34.610448Z","iopub.status.idle":"2022-03-03T14:59:34.610765Z","shell.execute_reply.started":"2022-03-03T14:59:34.610596Z","shell.execute_reply":"2022-03-03T14:59:34.610618Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"WOW, 87% are female customers. Does this make any sense? If you go to an H&M I think 9 out of 10 people are usually women so it definitely sounds possible","metadata":{}},{"cell_type":"markdown","source":"## 2. Find Each Customer's Last Week of Purchases","metadata":{}},{"cell_type":"markdown","source":"Now we will find each customers last week of purchases. This can easily be done as following:\n- keeping max t_dat per customer id\n- merging back to our main training table\n- get the difference between actual date and max date per customer\n- filter ones that has more than 6-7 days","metadata":{}},{"cell_type":"code","source":"tmp = product_merge.groupby('customer_id').t_dat.max().reset_index()\ntmp.columns = ['customer_id','max_dat']\ntrain = product_merge.merge(tmp,on=['customer_id'],how='left')\ntrain.head()","metadata":{"execution":{"iopub.status.busy":"2022-03-03T14:59:34.611809Z","iopub.status.idle":"2022-03-03T14:59:34.612125Z","shell.execute_reply.started":"2022-03-03T14:59:34.611952Z","shell.execute_reply":"2022-03-03T14:59:34.611974Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print('Train shape before:',train.shape[0])","metadata":{"execution":{"iopub.status.busy":"2022-03-03T14:59:34.614053Z","iopub.status.idle":"2022-03-03T14:59:34.614436Z","shell.execute_reply.started":"2022-03-03T14:59:34.614225Z","shell.execute_reply":"2022-03-03T14:59:34.61425Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train['diff_dat'] = (train.max_dat - train.t_dat).dt.days\ntrain = train.loc[train['diff_dat']<=6]\nprint('Train shape after:',train.shape[0])","metadata":{"execution":{"iopub.status.busy":"2022-03-03T14:59:34.615632Z","iopub.status.idle":"2022-03-03T14:59:34.615955Z","shell.execute_reply.started":"2022-03-03T14:59:34.615783Z","shell.execute_reply":"2022-03-03T14:59:34.615806Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We can see how we significantly reduced the training set size. Apart for faster computation time, this also works for fast fashion since we want to recommend product customers will buy, so this will be trendy products or the ones that are being marketed and new in the store","metadata":{}},{"cell_type":"markdown","source":"## 3. Most Often Previously Purchased Items","metadata":{}},{"cell_type":"markdown","source":"Here, we will do the main part of our recommender. The steps are the following:\n- Get count of previously purchased item to be able to sort by most trendy item\n- Get pairs of articles bought frequently with each other","metadata":{}},{"cell_type":"markdown","source":"### Count of previoulsy purchased items","metadata":{}},{"cell_type":"code","source":"tmp = train.groupby(['customer_id','article_id'])['t_dat'].agg('count').reset_index()\ntmp.columns = ['customer_id','article_id','count']\ntmp.head()","metadata":{"execution":{"iopub.status.busy":"2022-03-03T14:59:34.617101Z","iopub.status.idle":"2022-03-03T14:59:34.61751Z","shell.execute_reply.started":"2022-03-03T14:59:34.617272Z","shell.execute_reply":"2022-03-03T14:59:34.617296Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train = train.merge(tmp,on=['customer_id','article_id'],how='left')\ntrain = train.sort_values(['count','t_dat'],ascending=False)\ntrain = train.drop_duplicates(['customer_id','article_id'])\ntrain = train.sort_values(['count','t_dat'],ascending=False)\ntrain.head()","metadata":{"execution":{"iopub.status.busy":"2022-03-03T14:59:34.618444Z","iopub.status.idle":"2022-03-03T14:59:34.618754Z","shell.execute_reply.started":"2022-03-03T14:59:34.618586Z","shell.execute_reply":"2022-03-03T14:59:34.618607Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Pairs of items frequently purchased together","metadata":{}},{"cell_type":"markdown","source":"This is the most imporant part of the recommendation. Here we will do the following:\n- Use the main traning transactional dataset (with all the transactions) to calcualte paired articles\n- Get the value counts of female articles and others.\n- Create a dictionary with paired items most frequently bought together for female and for others","metadata":{}},{"cell_type":"code","source":"df_trans1 = product_merge[['customer_id','article_id','index_group_name']]","metadata":{"execution":{"iopub.status.busy":"2022-03-03T14:59:34.620371Z","iopub.status.idle":"2022-03-03T14:59:34.620737Z","shell.execute_reply.started":"2022-03-03T14:59:34.620546Z","shell.execute_reply":"2022-03-03T14:59:34.620567Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#This will sort articles by purchase count \nvc_female = df_trans1.loc[df_trans1[\"index_group_name\"]==\"Ladieswear\"].article_id.value_counts()\nvc_else = df_trans1.loc[df_trans1[\"index_group_name\"]!=\"Ladieswear\"].article_id.value_counts()","metadata":{"execution":{"iopub.status.busy":"2022-03-03T14:59:34.622157Z","iopub.status.idle":"2022-03-03T14:59:34.622508Z","shell.execute_reply.started":"2022-03-03T14:59:34.62231Z","shell.execute_reply":"2022-03-03T14:59:34.622351Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"pairs_f = {}\nfor j,i in enumerate(vc_female.index.values[1000:1032]):\n    #if j%10==0: print(j,', ',end='')\n    USERS = df_trans1.loc[df_trans1.article_id==i.item(),'customer_id'].unique()\n    vc2 = df_trans1.loc[(df_trans1.customer_id.isin(USERS))&(df_trans1.article_id!=i.item()),'article_id'].value_counts()\n    pairs_f[i.item()] = [vc2.index[0], vc2.index[1], vc2.index[2]]","metadata":{"execution":{"iopub.status.busy":"2022-03-03T14:59:34.623911Z","iopub.status.idle":"2022-03-03T14:59:34.62423Z","shell.execute_reply.started":"2022-03-03T14:59:34.624059Z","shell.execute_reply":"2022-03-03T14:59:34.624081Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"pairs_e = {}\nfor j,i in enumerate(vc_else.index.values[1000:1032]):\n    #if j%10==0: print(j,', ',end='')\n    USERS = df_trans1.loc[df_trans1.article_id==i.item(),'customer_id'].unique()\n    vc2 = df_trans1.loc[(df_trans1.customer_id.isin(USERS))&(df_trans1.article_id!=i.item()),'article_id'].value_counts()\n    pairs_e[i.item()] = [vc2.index[0], vc2.index[1], vc2.index[2]]","metadata":{"execution":{"iopub.status.busy":"2022-03-03T14:59:34.626208Z","iopub.status.idle":"2022-03-03T14:59:34.626562Z","shell.execute_reply.started":"2022-03-03T14:59:34.62638Z","shell.execute_reply":"2022-03-03T14:59:34.626398Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"pairs_f","metadata":{"execution":{"iopub.status.busy":"2022-03-03T14:59:34.627484Z","iopub.status.idle":"2022-03-03T14:59:34.627841Z","shell.execute_reply.started":"2022-03-03T14:59:34.62768Z","shell.execute_reply":"2022-03-03T14:59:34.627696Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"pairs_e","metadata":{"execution":{"iopub.status.busy":"2022-03-03T14:59:34.629818Z","iopub.status.idle":"2022-03-03T14:59:34.630353Z","shell.execute_reply.started":"2022-03-03T14:59:34.630133Z","shell.execute_reply":"2022-03-03T14:59:34.630154Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Above, we can see the pairs of articles. Now we can just add up the 2 dictionaries and map them in our trainning set","metadata":{}},{"cell_type":"code","source":"pairs_f.update(pairs_e)","metadata":{"execution":{"iopub.status.busy":"2022-03-03T14:59:34.631454Z","iopub.status.idle":"2022-03-03T14:59:34.6319Z","shell.execute_reply.started":"2022-03-03T14:59:34.631711Z","shell.execute_reply":"2022-03-03T14:59:34.63173Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"pairs_f","metadata":{"execution":{"iopub.status.busy":"2022-03-03T14:59:34.633157Z","iopub.status.idle":"2022-03-03T14:59:34.633508Z","shell.execute_reply.started":"2022-03-03T14:59:34.633301Z","shell.execute_reply":"2022-03-03T14:59:34.633336Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train['article_id2'] = train.article_id.map(pairs_f)\ntrain.head()","metadata":{"execution":{"iopub.status.busy":"2022-03-03T14:59:34.634577Z","iopub.status.idle":"2022-03-03T14:59:34.63489Z","shell.execute_reply.started":"2022-03-03T14:59:34.634729Z","shell.execute_reply":"2022-03-03T14:59:34.634745Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## 4. Recommendation of paired items","metadata":{}},{"cell_type":"markdown","source":"Now we will filter our data and keep only important features. This trainning data set will become our prediction submission. We will do the following:\n- keep only customer id and our paired articles feature\n- remove null values\n- format our new paired articles feature to be something match the submission","metadata":{}},{"cell_type":"code","source":"train2 = train[['customer_id','article_id2']].copy()\ntrain2 = train2.loc[train2.article_id2.notnull()]\ntrain2 = train2.rename({'article_id2':'article_id'},axis=1)\ntrain2.head()","metadata":{"execution":{"iopub.status.busy":"2022-03-03T14:59:34.636005Z","iopub.status.idle":"2022-03-03T14:59:34.636339Z","shell.execute_reply.started":"2022-03-03T14:59:34.636152Z","shell.execute_reply":"2022-03-03T14:59:34.636167Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train2[['team1','team2', \"team3\"]] = pd.DataFrame(train2.article_id.tolist(), index= train2.index)\ntrain2[\"join\"] = train2.team1.astype(str) + \" 0\" + train2.team2.astype(str) + \" 0\" + train2.team3.astype(str)\ntrain2 = train2.drop(['article_id', 'team1','team2', \"team3\"], axis=1)\ntrain2 = train2.rename({'join':'article_id'},axis=1)\ntrain2 = train2.drop_duplicates(['customer_id','article_id'])\ntrain2.head()","metadata":{"execution":{"iopub.status.busy":"2022-03-03T14:59:34.637867Z","iopub.status.idle":"2022-03-03T14:59:34.638185Z","shell.execute_reply.started":"2022-03-03T14:59:34.638018Z","shell.execute_reply":"2022-03-03T14:59:34.638033Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train = train[['customer_id','article_id']]\ntrain = pd.concat([train,train2],axis=0,ignore_index=True)\ntrain = train.drop_duplicates(['customer_id','article_id'])\ntrain","metadata":{"execution":{"iopub.status.busy":"2022-03-03T14:59:34.639915Z","iopub.status.idle":"2022-03-03T14:59:34.64022Z","shell.execute_reply.started":"2022-03-03T14:59:34.640061Z","shell.execute_reply":"2022-03-03T14:59:34.640076Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train.article_id = ' 0' + train.article_id.astype('str')","metadata":{"execution":{"iopub.status.busy":"2022-03-03T14:59:34.641747Z","iopub.status.idle":"2022-03-03T14:59:34.642052Z","shell.execute_reply.started":"2022-03-03T14:59:34.641894Z","shell.execute_reply":"2022-03-03T14:59:34.64191Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train","metadata":{"execution":{"iopub.status.busy":"2022-03-03T14:59:34.643046Z","iopub.status.idle":"2022-03-03T14:59:34.643374Z","shell.execute_reply.started":"2022-03-03T14:59:34.643185Z","shell.execute_reply":"2022-03-03T14:59:34.643201Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"preds = pd.DataFrame(train.groupby('customer_id').article_id.sum().reset_index())\npreds.columns = ['customer_id','prediction']\npreds.head()","metadata":{"execution":{"iopub.status.busy":"2022-03-03T14:59:34.644568Z","iopub.status.idle":"2022-03-03T14:59:34.644883Z","shell.execute_reply.started":"2022-03-03T14:59:34.644724Z","shell.execute_reply":"2022-03-03T14:59:34.644739Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Almost there. Table already in its final format. Now, we need to add the top 12 popular items for each customer (female or others)","metadata":{}},{"cell_type":"markdown","source":"## 5. Recommend Last Week's Most Popular Items","metadata":{}},{"cell_type":"markdown","source":"Here, we will use our first training set with all data and filter by the last week","metadata":{}},{"cell_type":"code","source":"df_trans2 = product_merge.loc[product_merge.t_dat >= pd.to_datetime('2020-09-16')]\ndf_trans2.head()","metadata":{"execution":{"iopub.status.busy":"2022-03-03T14:59:34.64629Z","iopub.status.idle":"2022-03-03T14:59:34.646998Z","shell.execute_reply.started":"2022-03-03T14:59:34.646817Z","shell.execute_reply":"2022-03-03T14:59:34.646837Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We now filter by index_group_name to see if it is female or not by filtering with Ladieswear","metadata":{}},{"cell_type":"code","source":"df_trans2_female = df_trans2.loc[df_trans2[\"index_group_name\"]==\"Ladieswear\"]\ndf_trans2_else = df_trans2.loc[df_trans2[\"index_group_name\"]!=\"Ladieswear\"]","metadata":{"execution":{"iopub.status.busy":"2022-03-03T14:59:34.647999Z","iopub.status.idle":"2022-03-03T14:59:34.648292Z","shell.execute_reply.started":"2022-03-03T14:59:34.64814Z","shell.execute_reply":"2022-03-03T14:59:34.648156Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Next, we get our top12 articles for females and for others, in the format needed","metadata":{}},{"cell_type":"code","source":"top12_female = '0' + ' 0'.join(df_trans2_female[\"article_id\"].value_counts().index.astype('str')[:12])\ntop12_female","metadata":{"execution":{"iopub.status.busy":"2022-03-03T14:59:34.649798Z","iopub.status.idle":"2022-03-03T14:59:34.650096Z","shell.execute_reply.started":"2022-03-03T14:59:34.649942Z","shell.execute_reply":"2022-03-03T14:59:34.649957Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"top12_else = '0' + ' 0'.join(df_trans2_else[\"article_id\"].value_counts().index.astype('str')[:12])\ntop12_else","metadata":{"execution":{"iopub.status.busy":"2022-03-03T14:59:34.651598Z","iopub.status.idle":"2022-03-03T14:59:34.65198Z","shell.execute_reply.started":"2022-03-03T14:59:34.65182Z","shell.execute_reply":"2022-03-03T14:59:34.651837Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Submission","metadata":{}},{"cell_type":"markdown","source":"First, lets read the sample submission, keep all customers and merge our pred table with customer_id as key","metadata":{}},{"cell_type":"code","source":"sub = pd.read_csv(data_path+\"sample_submission.csv\")\nsub.head()","metadata":{"execution":{"iopub.status.busy":"2022-03-03T14:59:34.653328Z","iopub.status.idle":"2022-03-03T14:59:34.653631Z","shell.execute_reply.started":"2022-03-03T14:59:34.653473Z","shell.execute_reply":"2022-03-03T14:59:34.653489Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sub = sub[['customer_id']]\nsub = sub.merge(preds,on='customer_id', how='left').fillna('')\nsub.head()","metadata":{"execution":{"iopub.status.busy":"2022-03-03T14:59:34.65546Z","iopub.status.idle":"2022-03-03T14:59:34.65604Z","shell.execute_reply.started":"2022-03-03T14:59:34.655732Z","shell.execute_reply":"2022-03-03T14:59:34.655763Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Now, we can split our customer_id by female and others (by using the sex_user table we created in step 1)","metadata":{}},{"cell_type":"code","source":"sex_female = sex_user[[\"customer_id\"]]\nsub_female = sub.loc[(sub.customer_id.isin(sex_female.customer_id))]\nsub_else = sub.loc[~sub.customer_id.isin(sex_female.customer_id)]","metadata":{"execution":{"iopub.status.busy":"2022-03-03T14:59:34.657087Z","iopub.status.idle":"2022-03-03T14:59:34.657622Z","shell.execute_reply.started":"2022-03-03T14:59:34.657348Z","shell.execute_reply":"2022-03-03T14:59:34.657376Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Now, lets append the top 12 predictions for females and for others and concat them together to get the submission with all customers.","metadata":{}},{"cell_type":"code","source":"sub_female.prediction = sub_female.prediction + \" \" +top12_female\nsub_else.prediction = sub_else.prediction + \" \" +top12_else\nsub_final = pd.concat([sub_female,sub_else])\nsub_final.prediction = sub_final.prediction.str.strip()\nsub_final.prediction = sub_final.prediction.str[:131]\nsub_final.to_csv(f'submission.csv',index=False)\nsub_final.head()","metadata":{"execution":{"iopub.status.busy":"2022-03-03T14:59:34.661949Z","iopub.status.idle":"2022-03-03T14:59:34.662484Z","shell.execute_reply.started":"2022-03-03T14:59:34.662186Z","shell.execute_reply":"2022-03-03T14:59:34.662213Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Disclaimer - Most of the notebook methodology was taken from: https://www.kaggle.com/cdeotte/recommend-items-purchased-together-0-021.","metadata":{}}]}