{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.7.12","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"gpu","dataSources":[{"sourceId":31254,"databundleVersionId":3103714,"sourceType":"competition"}],"dockerImageVersionId":30177,"isInternetEnabled":false,"language":"python","sourceType":"notebook","isGpuEnabled":true}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Introduction","metadata":{}},{"cell_type":"markdown","source":"In this notebook we will perform a basic association analysis as a practice of the recommendation system.  \nThe goal is to explore the association rules between products using data from the last two months of the H&M Group.","metadata":{}},{"cell_type":"markdown","source":"# Transaction data (import and processing)","metadata":{}},{"cell_type":"markdown","source":"For this quick basic association analysis, we will focus our data on August-September 2020.  \nOnly three transaction data are used: \"t_dat\" (date-related data), \"customer_id\", and \"article_id\".","metadata":{}},{"cell_type":"code","source":"import numpy as np\nimport pandas as pd\n\nfrom sklearn import preprocessing\nle = preprocessing.LabelEncoder()\n\ndf = pd.read_csv(\"../input/h-and-m-personalized-fashion-recommendations/transactions_train.csv\", dtype={\"article_id\": str})\nprint(df.shape)\n\ndf['day'] = pd.to_datetime(df['t_dat'])\ndf = df.query('day > 20200731')\ndf = df[['t_dat','customer_id','article_id']]\ndf.head()\n\nprint(df.shape)","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","execution":{"iopub.status.busy":"2023-12-14T12:42:08.434738Z","iopub.execute_input":"2023-12-14T12:42:08.435517Z","iopub.status.idle":"2023-12-14T12:43:27.301795Z","shell.execute_reply.started":"2023-12-14T12:42:08.435480Z","shell.execute_reply":"2023-12-14T12:43:27.300961Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df.dtypes","metadata":{"execution":{"iopub.status.busy":"2023-12-14T12:43:27.303534Z","iopub.execute_input":"2023-12-14T12:43:27.303770Z","iopub.status.idle":"2023-12-14T12:43:27.313351Z","shell.execute_reply.started":"2023-12-14T12:43:27.303740Z","shell.execute_reply":"2023-12-14T12:43:27.312476Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df.isnull().sum()","metadata":{"execution":{"iopub.status.busy":"2023-12-14T12:43:27.314630Z","iopub.execute_input":"2023-12-14T12:43:27.314858Z","iopub.status.idle":"2023-12-14T12:43:27.934169Z","shell.execute_reply.started":"2023-12-14T12:43:27.314829Z","shell.execute_reply":"2023-12-14T12:43:27.933341Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Article data(import and processing)","metadata":{}},{"cell_type":"code","source":"df_artic = pd.read_csv(\"../input/h-and-m-personalized-fashion-recommendations/articles.csv\", dtype={\"article_id\": str})\nprint(df_artic.shape)\ndf_artic.head()","metadata":{"execution":{"iopub.status.busy":"2023-12-14T12:43:27.935440Z","iopub.execute_input":"2023-12-14T12:43:27.935728Z","iopub.status.idle":"2023-12-14T12:43:29.105126Z","shell.execute_reply.started":"2023-12-14T12:43:27.935692Z","shell.execute_reply":"2023-12-14T12:43:29.104258Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The items are divided into detailed categories.  \nIn this article, we will consider combinations by color and product type, so let's review those unique values.","metadata":{}},{"cell_type":"code","source":"print(df_artic['perceived_colour_master_name'].nunique())\nprint(df_artic['perceived_colour_master_name'].unique())\nprint()\nprint(df_artic['product_type_name'].nunique())\nprint(df_artic['product_type_name'].unique())","metadata":{"execution":{"iopub.status.busy":"2023-12-14T12:43:29.107723Z","iopub.execute_input":"2023-12-14T12:43:29.108153Z","iopub.status.idle":"2023-12-14T12:43:29.173851Z","shell.execute_reply.started":"2023-12-14T12:43:29.108106Z","shell.execute_reply":"2023-12-14T12:43:29.172733Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Combine color and product type data in one column.","metadata":{}},{"cell_type":"code","source":"df_artic = df_artic[['article_id','perceived_colour_master_name','product_type_name']]\ndf_artic['item'] = df_artic['perceived_colour_master_name'] + ['-'] + df_artic['product_type_name']\ndf_artic.head()","metadata":{"execution":{"iopub.status.busy":"2023-12-14T12:43:29.225047Z","iopub.execute_input":"2023-12-14T12:43:29.225375Z","iopub.status.idle":"2023-12-14T12:43:29.272909Z","shell.execute_reply.started":"2023-12-14T12:43:29.225328Z","shell.execute_reply":"2023-12-14T12:43:29.271855Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_artic = df_artic[['article_id','item']]\ndf_artic.head()","metadata":{"execution":{"iopub.status.busy":"2023-12-14T12:43:29.274157Z","iopub.execute_input":"2023-12-14T12:43:29.274407Z","iopub.status.idle":"2023-12-14T12:43:29.297595Z","shell.execute_reply.started":"2023-12-14T12:43:29.274375Z","shell.execute_reply":"2023-12-14T12:43:29.296807Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_artic.isnull().sum()","metadata":{"execution":{"iopub.status.busy":"2023-12-14T12:43:29.298872Z","iopub.execute_input":"2023-12-14T12:43:29.299166Z","iopub.status.idle":"2023-12-14T12:43:29.329677Z","shell.execute_reply.started":"2023-12-14T12:43:29.299134Z","shell.execute_reply":"2023-12-14T12:43:29.328827Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Data Merge","metadata":{}},{"cell_type":"code","source":"df = pd.merge(df,df_artic,how=\"left\",on='article_id')\ndf.head()","metadata":{"execution":{"iopub.status.busy":"2023-12-14T12:43:29.330887Z","iopub.execute_input":"2023-12-14T12:43:29.331145Z","iopub.status.idle":"2023-12-14T12:43:30.075256Z","shell.execute_reply.started":"2023-12-14T12:43:29.331115Z","shell.execute_reply":"2023-12-14T12:43:30.074491Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Each ID is very long for about 300,000 customer data.    \nSo, we'll transform it.   \nThen merge that data with the date data.","metadata":{}},{"cell_type":"code","source":"df['customer_id'].nunique()","metadata":{"execution":{"iopub.status.busy":"2023-12-14T12:43:30.076544Z","iopub.execute_input":"2023-12-14T12:43:30.076823Z","iopub.status.idle":"2023-12-14T12:43:30.554732Z","shell.execute_reply.started":"2023-12-14T12:43:30.076784Z","shell.execute_reply":"2023-12-14T12:43:30.553988Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df['customer_id'] = le.fit_transform(df['customer_id'])\ndf.head()","metadata":{"execution":{"iopub.status.busy":"2023-12-14T12:43:30.556047Z","iopub.execute_input":"2023-12-14T12:43:30.556709Z","iopub.status.idle":"2023-12-14T12:43:32.842997Z","shell.execute_reply.started":"2023-12-14T12:43:30.556666Z","shell.execute_reply":"2023-12-14T12:43:32.842193Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df['customer_id'].nunique()","metadata":{"execution":{"iopub.status.busy":"2023-12-14T12:43:32.844062Z","iopub.execute_input":"2023-12-14T12:43:32.845533Z","iopub.status.idle":"2023-12-14T12:43:32.874897Z","shell.execute_reply.started":"2023-12-14T12:43:32.845501Z","shell.execute_reply":"2023-12-14T12:43:32.874194Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df['customer_id'] = df['customer_id'].astype(str)\ndf['day-id'] = df['t_dat'] + ['-'] + df['customer_id']\n\ndf.head()","metadata":{"execution":{"iopub.status.busy":"2023-12-14T12:43:32.878013Z","iopub.execute_input":"2023-12-14T12:43:32.878246Z","iopub.status.idle":"2023-12-14T12:43:35.499390Z","shell.execute_reply.started":"2023-12-14T12:43:32.878217Z","shell.execute_reply":"2023-12-14T12:43:35.498584Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df['item'].nunique()","metadata":{"execution":{"iopub.status.busy":"2023-12-14T12:43:35.500756Z","iopub.execute_input":"2023-12-14T12:43:35.501058Z","iopub.status.idle":"2023-12-14T12:43:35.753576Z","shell.execute_reply.started":"2023-12-14T12:43:35.501014Z","shell.execute_reply":"2023-12-14T12:43:35.752753Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df.shape","metadata":{"execution":{"iopub.status.busy":"2023-12-14T12:43:35.754763Z","iopub.execute_input":"2023-12-14T12:43:35.755262Z","iopub.status.idle":"2023-12-14T12:43:35.761844Z","shell.execute_reply.started":"2023-12-14T12:43:35.755220Z","shell.execute_reply":"2023-12-14T12:43:35.760983Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"As a pre-processing step for association analysis, first narrow down the columns to be used.  \nNext, we convert the product type to come in the column name.","metadata":{}},{"cell_type":"code","source":"df = df.groupby(['day-id','item'])['article_id'].count()\ndf.head()","metadata":{"execution":{"iopub.status.busy":"2023-12-14T12:43:35.763042Z","iopub.execute_input":"2023-12-14T12:43:35.763308Z","iopub.status.idle":"2023-12-14T12:43:38.090027Z","shell.execute_reply.started":"2023-12-14T12:43:35.763260Z","shell.execute_reply":"2023-12-14T12:43:38.089292Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df2 = df.unstack()\ndf2.head()","metadata":{"execution":{"iopub.status.busy":"2023-12-14T12:43:38.334441Z","iopub.execute_input":"2023-12-14T12:43:38.334681Z","iopub.status.idle":"2023-12-14T12:43:42.506861Z","shell.execute_reply.started":"2023-12-14T12:43:38.334652Z","shell.execute_reply":"2023-12-14T12:43:42.506069Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df2.shape","metadata":{"execution":{"iopub.status.busy":"2023-12-14T12:43:42.507927Z","iopub.execute_input":"2023-12-14T12:43:42.508161Z","iopub.status.idle":"2023-12-14T12:43:42.513929Z","shell.execute_reply.started":"2023-12-14T12:43:42.508132Z","shell.execute_reply":"2023-12-14T12:43:42.512986Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Convert each cell to True or False format for analysis.   \nThis completes the previous step!","metadata":{}},{"cell_type":"code","source":"df2 = df2.fillna(0)\ndf2.head()","metadata":{"execution":{"iopub.status.busy":"2023-12-14T12:43:42.515045Z","iopub.execute_input":"2023-12-14T12:43:42.515301Z","iopub.status.idle":"2023-12-14T12:43:59.604190Z","shell.execute_reply.started":"2023-12-14T12:43:42.515260Z","shell.execute_reply":"2023-12-14T12:43:59.603264Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df3 = df2.apply(lambda x:x > 0)\ndf3.head()","metadata":{"execution":{"iopub.status.busy":"2023-12-14T12:43:59.605493Z","iopub.execute_input":"2023-12-14T12:43:59.605767Z","iopub.status.idle":"2023-12-14T12:44:01.152196Z","shell.execute_reply.started":"2023-12-14T12:43:59.605731Z","shell.execute_reply":"2023-12-14T12:44:01.151445Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# association analysis","metadata":{}},{"cell_type":"markdown","source":"Importing the library to be used","metadata":{}},{"cell_type":"code","source":"from mlxtend.frequent_patterns import apriori\nfrom mlxtend.frequent_patterns import association_rules\n\npd.set_option('display.max_colwidth',None)","metadata":{"execution":{"iopub.status.busy":"2023-12-14T12:44:01.153567Z","iopub.execute_input":"2023-12-14T12:44:01.154119Z","iopub.status.idle":"2023-12-14T12:44:01.158642Z","shell.execute_reply.started":"2023-12-14T12:44:01.154073Z","shell.execute_reply":"2023-12-14T12:44:01.157782Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Once here, we will summarize three important indicators used in association analysis   \n\n--------------------------------------------\n### (1) Support\n* The probability that an item has been purchased.\n* The degree of support (A) can be calculated by **n(A)/n(U)**.\n* When both A and B are purchased at the same time, n(A∩B)/n(U) can be calculated.  \n<br>\n\n### (2) Lift\n* Indicates the correlation between purchases of Item A and Item B.\n* The calculation of B under the condition that A is purchased is considered below.(B | A)   \n<br>  \n\n   n(A∩B)/n(U)  \n  ――――――――  \nn(A)/n(U) x n(B)/n(U)  \n<br>\n* In general, lift values higher than \"1\" are positively correlated with the occurrence of items A and B.  \n\n※it is better to think basically in terms of the size of the **lift** value while taking the **confidence** value into account.  \n<br>\n### (3) Confidence  \n* The probability that item B will also be purchased when item A is purchased.  \n* It can be calculated in this way.\n* **n(A∩B)/n(A)**\n* The higher the confidence level, the more likely it is that A and B will sell at the same time.  \n\n※**Confidence** levels are high even when the product is not selling very well in the first place, so the value of **support** should be checked.　　\n\n--------------------------------------------","metadata":{}},{"cell_type":"markdown","source":"### (1) Support  \nFirst, let's check the data for support values higher than 0.01.","metadata":{}},{"cell_type":"code","source":"frequent_items = apriori(df3,min_support=0.01,use_colnames=True)\nfrequent_items","metadata":{"execution":{"iopub.status.busy":"2023-12-14T12:44:01.159692Z","iopub.execute_input":"2023-12-14T12:44:01.159910Z","iopub.status.idle":"2023-12-14T12:44:04.788791Z","shell.execute_reply.started":"2023-12-14T12:44:01.159882Z","shell.execute_reply":"2023-12-14T12:44:04.788003Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The itemsets column also contains data for only one item, so check with the data for two items.","metadata":{}},{"cell_type":"code","source":"frequent_items['length'] = frequent_items['itemsets'].apply(lambda x: len(x))\nfrequent_items.query(\"length == 2 & support >= 0.01\").sort_values(by='support',ascending=False)","metadata":{"execution":{"iopub.status.busy":"2023-12-14T12:44:04.811263Z","iopub.execute_input":"2023-12-14T12:44:04.811516Z","iopub.status.idle":"2023-12-14T12:44:04.832387Z","shell.execute_reply.started":"2023-12-14T12:44:04.811484Z","shell.execute_reply":"2023-12-14T12:44:04.831492Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### (2) Lift & (3) Confidence   \nFinally, we will look at the results of the calculation, which also includes Lift and Confidence.  \nLift values are for data higher than \"1\",  \nWe will look at the relationship between one product and one product only.","metadata":{"execution":{"iopub.status.busy":"2023-12-14T12:44:04.833715Z","iopub.execute_input":"2023-12-14T12:44:04.833987Z","iopub.status.idle":"2023-12-14T12:44:04.841689Z","shell.execute_reply.started":"2023-12-14T12:44:04.833930Z","shell.execute_reply":"2023-12-14T12:44:04.840842Z"}}},{"cell_type":"code","source":"df_rule = association_rules(frequent_items,metric='lift',min_threshold=1)\n\ndf_rule['antecedents_length'] = df_rule['antecedents'].apply(lambda x: len(x))\ndf_rule['consequents_length'] = df_rule['consequents'].apply(lambda x: len(x))\n\ndf_rule.query('antecedents_length == 1 & consequents_length ==1').sort_values(by='lift',ascending=False)","metadata":{"execution":{"iopub.status.busy":"2023-12-14T12:44:04.842881Z","iopub.execute_input":"2023-12-14T12:44:04.843173Z","iopub.status.idle":"2023-12-14T12:44:04.878498Z","shell.execute_reply.started":"2023-12-14T12:44:04.843139Z","shell.execute_reply":"2023-12-14T12:44:04.877575Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The results show that white T-shirts and black T-shirts are most frequently purchased at the same time.  \nOther than that, blue, gray, and black trousers also seem to be related to each other, so it would be better to display them close together when selling them.  \n<br>\nThank you so much for reading!","metadata":{}}]}