{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"<h1><center> H&M Personalized Fashion Recommendations</center></h1>\n<center><img src=\"https://static.euronews.com/articles/stories/04/98/67/36/1440x810_cmsv2_4f7e599c-edce-5354-b734-d1e960153442-4986736.jpg\" width=\"400px\"><center>\n<h3><center>Provide product recommendations based on previous purchases</center></h3>\n    Develop product recommendations based on data from previous transactions, as well as from customer and product meta data.","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19"}},{"cell_type":"markdown","source":"#### Data Description\n\n##### Files\n\n**images/** - a folder of images corresponding to each article_id; images are placed in subfolders starting with the first three digits of the article_id; note, not all article_id values have a corresponding image.\n\n**articles.csv** - detailed metadata for each article_id available for purchase\n\n**customers.csv** - metadata for each customer_id in dataset\n\n**sample_submission.csv** - a sample submission file in the correct format\n\n**transactions_train.csv** - the training data, consisting of the purchases each customer for each date, as well as additional information. Duplicate rows correspond to multiple purchases of the same item. Your task is to predict the article_ids each customer will purchase during the 7-day period immediately after the training data period.\n\n**NOTE:** You must make predictions for all customer_id values found in the sample submission. All customers who made purchases during the test period are scored, regardless of whether they had purchase history in the training data.","metadata":{}},{"cell_type":"markdown","source":"<h2><center> Diving Into DATA <center>","metadata":{}},{"cell_type":"code","source":"## Utility Functions\n\n\ndef plot_percentage(ax, total): \n    for p in ax.patches:\n        percentage = '{:.2f}%'.format(100*p.get_height()/total)\n        x = p.get_x() + p.get_width() / 2 - 0.25\n        y = p.get_y() + p.get_height()\n        ax.annotate(percentage, (x, y), size = 12)\n    \ndef plot_distribution(dataframe, columns, ylabel):\n    for column in columns:\n        result= dataframe.groupby(column).size()\n        plt.figure(figsize=(20,10))\n        plt.xticks(rotation=90)\n        plt.title(\"Distribution of \"+column, fontsize=20)\n        plt.ylabel(ylabel)\n        ax = sns.barplot(x=result.index, y=result.values)\n        total= dataframe.shape[0]\n        plot_percentage(ax, total)\n        ","metadata":{"execution":{"iopub.status.busy":"2022-02-09T06:14:06.179247Z","iopub.execute_input":"2022-02-09T06:14:06.180099Z","iopub.status.idle":"2022-02-09T06:14:06.210714Z","shell.execute_reply.started":"2022-02-09T06:14:06.179994Z","shell.execute_reply":"2022-02-09T06:14:06.209630Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import pandas as pd\nimport matplotlib.pyplot as plt\nimport seaborn as sns\nimport cv2","metadata":{"execution":{"iopub.status.busy":"2022-02-09T06:14:06.814819Z","iopub.execute_input":"2022-02-09T06:14:06.815105Z","iopub.status.idle":"2022-02-09T06:14:08.381480Z","shell.execute_reply.started":"2022-02-09T06:14:06.815078Z","shell.execute_reply":"2022-02-09T06:14:08.380437Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"articles= pd.read_csv('../input/h-and-m-personalized-fashion-recommendations/articles.csv')\npd.set_option('display.max_colwidth', None)\narticles.head()","metadata":{"execution":{"iopub.status.busy":"2022-02-09T06:14:09.070931Z","iopub.execute_input":"2022-02-09T06:14:09.071285Z","iopub.status.idle":"2022-02-09T06:14:10.482450Z","shell.execute_reply.started":"2022-02-09T06:14:09.071250Z","shell.execute_reply":"2022-02-09T06:14:10.481350Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"articles.nunique()\n#Unique values for all the columns","metadata":{"execution":{"iopub.status.busy":"2022-02-09T06:14:10.484884Z","iopub.execute_input":"2022-02-09T06:14:10.485278Z","iopub.status.idle":"2022-02-09T06:14:10.671283Z","shell.execute_reply.started":"2022-02-09T06:14:10.485227Z","shell.execute_reply":"2022-02-09T06:14:10.670553Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Lets deep dive into Relevant columns to understand more about the characteristics of articles.\n\ncolumns= ['product_group_name','index_name','index_group_name','garment_group_name']\n\nplot_distribution(articles, columns,ylabel=\"Number of Articles\")","metadata":{"execution":{"iopub.status.busy":"2022-02-09T06:14:11.117507Z","iopub.execute_input":"2022-02-09T06:14:11.118439Z","iopub.status.idle":"2022-02-09T06:14:13.183640Z","shell.execute_reply.started":"2022-02-09T06:14:11.118386Z","shell.execute_reply":"2022-02-09T06:14:13.180529Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### OBSERVATIONS:\n1) Based on product_group_name, H&M insuatry products are dominated by **Garment Upper Body** with **40.5%** of products, followd by **Garment Lower Body, Garment Full Body and Accessories** with **19.7 and 12.5% and 10.5%** respectively. All others products are 5% or less than that.\n\n2) **Ladieswear and Babywear** segment is close to **70% together** and rest is the Menswear, Divided and Sports category.\n\n3) Looking into the **Garment Groups**, we can see that the **large proportion** of the products are **Woollen wear(Jersey, Knitwear)** followed by Accessories.\n","metadata":{}},{"cell_type":"code","source":"customers= pd.read_csv('../input/h-and-m-personalized-fashion-recommendations/customers.csv')\n\ncustomers.head()","metadata":{"execution":{"iopub.status.busy":"2022-02-09T06:14:13.185485Z","iopub.execute_input":"2022-02-09T06:14:13.186341Z","iopub.status.idle":"2022-02-09T06:14:19.699999Z","shell.execute_reply.started":"2022-02-09T06:14:13.186284Z","shell.execute_reply":"2022-02-09T06:14:19.698880Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"columns= ['club_member_status','fashion_news_frequency','age']\nplot_distribution(customers, columns,ylabel=\"Number of Customers\")","metadata":{"execution":{"iopub.status.busy":"2022-02-09T06:14:19.702680Z","iopub.execute_input":"2022-02-09T06:14:19.703068Z","iopub.status.idle":"2022-02-09T06:14:22.875935Z","shell.execute_reply.started":"2022-02-09T06:14:19.703020Z","shell.execute_reply":"2022-02-09T06:14:22.875220Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### OBSERVATIONS:\n\n1) 92.75% of customers have their member status **Active** and very few members have actually left the club. But I think that does not mean that all the 92% customers are active in terms of purchases. Their account could be active but in dormant state.\n\n2) There is an interesting insight from the **Age Distribution**. We can see that the maximum customers for H&M are from **Age Group 18-25** and then customer proportion starts decreasing which is obvious. But the **45-50+ age group** also has a good proportion and the demand increases for the products. \n\n#### We can further deep dive to understand the type of products that are prominent in both the age groups. \n\n#### Is there any differentiating factor in the product groups or H&M is targeting both the segments in the same product groups?? ---- We will see that further in the transactions data","metadata":{}},{"cell_type":"code","source":"transactions_train= pd.read_csv('../input/h-and-m-personalized-fashion-recommendations/transactions_train.csv')\ntransactions_train.head()","metadata":{"execution":{"iopub.status.busy":"2022-02-09T06:14:22.877101Z","iopub.execute_input":"2022-02-09T06:14:22.878185Z","iopub.status.idle":"2022-02-09T06:15:32.434317Z","shell.execute_reply.started":"2022-02-09T06:14:22.878115Z","shell.execute_reply":"2022-02-09T06:15:32.432510Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"transactions_train.shape","metadata":{"execution":{"iopub.status.busy":"2022-02-09T06:15:32.436671Z","iopub.execute_input":"2022-02-09T06:15:32.436967Z","iopub.status.idle":"2022-02-09T06:15:32.443837Z","shell.execute_reply.started":"2022-02-09T06:15:32.436932Z","shell.execute_reply":"2022-02-09T06:15:32.443211Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Join transaction data with customers and articles data to Understand the transactional behavior\nmerged= (transactions_train.merge(customers, on='customer_id', how='left'))\nmerged.head()\n","metadata":{"execution":{"iopub.status.busy":"2022-02-09T04:42:19.883928Z","iopub.execute_input":"2022-02-09T04:42:19.884272Z","iopub.status.idle":"2022-02-09T04:42:50.278379Z","shell.execute_reply.started":"2022-02-09T04:42:19.88424Z","shell.execute_reply":"2022-02-09T04:42:50.277416Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"merged2= (merged.merge(articles, on='article_id', how='left'))\nmerged2.head()","metadata":{"execution":{"iopub.status.busy":"2022-02-09T04:45:15.825993Z","iopub.execute_input":"2022-02-09T04:45:15.826728Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"## Since the trasactional data is quite big. Merging the customers and articles both dataframes is\n## giving me error as RAM crashes. So the alternate solution:\n## Reference :https://stackoverflow.com/questions/47386405/memoryerror-when-i-merge-two-pandas-data-frames\n\ndf1 = transactions_train\ndf2 = customers\ndf3= articles\n\n# creating a empty bucket to save result\ndf_result = pd.DataFrame(columns=(df1.columns.append(df2.columns.append(df3.columns))).unique())\n\n# deleting df to save memory\ndel(df1)\n\ndef preprocess(x,df_result):\n    inter_res=(x.merge(df2, on = 'customer_id', how='left')).merge(df3,on='article_id', how='left')\n    df_result=pd.concat([df_result,inter_res], axis=0)\n    \n\nreader = pd.read_csv('../input/h-and-m-personalized-fashion-recommendations/transactions_train.csv', chunksize=1000) # chunksize depends with you colsize\n\nfor r in reader:\n    df_result= preprocess(r,df_result)\n","metadata":{"execution":{"iopub.status.busy":"2022-02-09T06:33:08.858534Z","iopub.execute_input":"2022-02-09T06:33:08.858840Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_result.shape","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_result.to_csv('../input/h-and-m-personalized-fashion-recommendations/merged_data.csv')","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<center><h2> Understanding transactional behaviour </h2></center>","metadata":{}},{"cell_type":"markdown","source":"<CENTER><H2> WORK IN PROGRESS............  </H2></CENTER>","metadata":{}},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]}]}