{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"code","source":"import numpy as np\nimport pandas as pd\nimport seaborn as sns\nimport matplotlib.pyplot as plt\nimport scipy.stats as stats\nimport os, sys\nimport glob\nfrom PIL import Image\nimport matplotlib.pyplot as plt\nfrom tqdm import tqdm\nfrom typing import List, Dict\nfrom matplotlib.ticker import PercentFormatter","metadata":{"execution":{"iopub.status.busy":"2023-08-22T07:00:14.377168Z","iopub.execute_input":"2023-08-22T07:00:14.377672Z","iopub.status.idle":"2023-08-22T07:00:15.767398Z","shell.execute_reply.started":"2023-08-22T07:00:14.377540Z","shell.execute_reply":"2023-08-22T07:00:15.765962Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Contents\n\n* [<font size=4>1. Read Data</font>](#1)\n    \n* [<font size=4>2. Articles</font>](#2)\n    * [2.1 What product do they have the most?](#2.1)\n    * [2.2 Pareto Analysis](#2.2)\n    * [2.3 Univariate Analysis](#2.3)\n    * [2.4 What type of index that is accounted for the most?](#2.4)\n    * [2.5 What's the portion of index group name for each garment group?](#2.5)\n    \n\n* [<font size=4>3. Customer</font>](#3)\n    * [3.1 How many Customers for each of club member status?](#3.1)\n    * [3.2 How many customers receive the fashion news by frequency?](#3.2)\n    * [3.3 How old customers overall?](#3.3)\n    * [3.4 How many customers for each of club member status](#3.4)\n    * [3.5 Correlation : age & club member status](#3.5)\n    * [3.6 Correlation : age & fashion news frequency](#3.6)\n    \n    \n* [<font size=4>4. Transactions</font>](#4)\n    * [4.1 Price Distribution](#4.1)\n    * [4.2 Which products do they buy the most?](#4.2)\n    * [4.3 How's the mean price for each product by index_name](#4.3)\n\n\n* [<font size=4>5. Image</font>](#5)","metadata":{}},{"cell_type":"markdown","source":"# <center><font size=6><a id=1>1. Read Data</a><font><center>","metadata":{}},{"cell_type":"code","source":"# image = \"../input/h-and-m-personalized-fashion-recommendations/images\"\narticles = \"../input/h-and-m-personalized-fashion-recommendations/articles.csv\"\ncustomers = \"../input/h-and-m-personalized-fashion-recommendations/customers.csv\"\ntransactions = \"../input/h-and-m-personalized-fashion-recommendations/transactions_train.csv\"","metadata":{"execution":{"iopub.status.busy":"2023-08-22T07:02:58.456604Z","iopub.execute_input":"2023-08-22T07:02:58.457089Z","iopub.status.idle":"2023-08-22T07:02:58.463240Z","shell.execute_reply.started":"2023-08-22T07:02:58.457049Z","shell.execute_reply":"2023-08-22T07:02:58.461777Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# <center><font size=6><a id=2>2. Articles</a></font></center>","metadata":{}},{"cell_type":"code","source":"def create_df(url:str) -> pd.DataFrame:\n    df = pd.read_csv(url)\n    return df\n\ndef get_article_id(df:pd.DataFrame, ids:pd.Series) -> pd.DataFrame:\n    df['article_id'] = [\"0\" + str(id) for id in ids]\n    return df","metadata":{"execution":{"iopub.status.busy":"2023-08-22T07:02:59.952641Z","iopub.execute_input":"2023-08-22T07:02:59.953070Z","iopub.status.idle":"2023-08-22T07:02:59.960616Z","shell.execute_reply.started":"2023-08-22T07:02:59.953035Z","shell.execute_reply":"2023-08-22T07:02:59.959226Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"articles = create_df(articles)\nids = articles['article_id']\narticles = get_article_id(articles, ids)","metadata":{"execution":{"iopub.status.busy":"2023-08-22T07:03:01.583779Z","iopub.execute_input":"2023-08-22T07:03:01.584271Z","iopub.status.idle":"2023-08-22T07:03:02.566252Z","shell.execute_reply.started":"2023-08-22T07:03:01.584235Z","shell.execute_reply":"2023-08-22T07:03:02.565152Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### 2.1 What product do they have the most?<a id=2.1></a>","metadata":{}},{"cell_type":"code","source":"def val_cnt(df:pd.DataFrame, col:str, top_n:int):\n    plt.figure(figsize = (8, 6))\n    sns.set(style = 'whitegrid')\n    _order = df[col].value_counts()[:top_n].index\n    viz = sns.countplot(x = col, data = df,\n                        order = _order)\n    plt.xticks(rotation = 45)\n    plt.show()","metadata":{"execution":{"iopub.status.busy":"2023-08-22T07:03:03.886214Z","iopub.execute_input":"2023-08-22T07:03:03.886812Z","iopub.status.idle":"2023-08-22T07:03:03.897146Z","shell.execute_reply.started":"2023-08-22T07:03:03.886760Z","shell.execute_reply":"2023-08-22T07:03:03.895596Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"val_cnt(articles, 'prod_name', 10)","metadata":{"execution":{"iopub.status.busy":"2023-08-22T07:03:05.399296Z","iopub.execute_input":"2023-08-22T07:03:05.399882Z","iopub.status.idle":"2023-08-22T07:03:06.032934Z","shell.execute_reply.started":"2023-08-22T07:03:05.399831Z","shell.execute_reply":"2023-08-22T07:03:06.031506Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Since they have many products, I filtered out the ‘Top10’ products. As a result, they have almost 100 ‘Dragonfly dress’ while the ‘DANTE set’ is the product they carry the least amount of. Make sure to know that the product they have the most of doesn’t necessarily mean that product is the most popular. We don’t know exactly why they have products the most.","metadata":{}},{"cell_type":"markdown","source":"### 2.2 Pareto Analysis<a id=2.2></a>","metadata":{}},{"cell_type":"code","source":"def create_dict(lst_1:list, lst_2:list) -> dict:\n    res = dict(zip(lst_1, lst_2))\n    return res\n\ndef id_name_price(id_prod_name_dict:Dict[str,str], id_price_dict:Dict[str, int]):\n    \n    set_1 = set(id_prod_name_dict)\n    set_2 = set(id_price_dict)\n    shared_id = set_1.intersection(set_2)\n    \n    res_lst = []\n    for _id in id_price_dict:\n        if _id in shared_id:\n            res_prod_name = id_prod_name_dict[_id]\n            res_price = id_price_dict[_id]\n            res_lst.append([_id, res_prod_name, res_price])\n    return res_lst","metadata":{"execution":{"iopub.status.busy":"2023-08-22T07:03:28.560308Z","iopub.execute_input":"2023-08-22T07:03:28.560796Z","iopub.status.idle":"2023-08-22T07:03:28.571951Z","shell.execute_reply.started":"2023-08-22T07:03:28.560755Z","shell.execute_reply":"2023-08-22T07:03:28.570507Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df1 = articles[['article_id', 'prod_name']]\nlst_1 = [_id for _id in articles['article_id']]\nlst_2 = [name for name in articles['prod_name']]\nid_name_dict = create_dict(lst_1, lst_2)","metadata":{"execution":{"iopub.status.busy":"2023-08-22T07:03:30.287650Z","iopub.execute_input":"2023-08-22T07:03:30.288138Z","iopub.status.idle":"2023-08-22T07:03:30.424412Z","shell.execute_reply.started":"2023-08-22T07:03:30.288096Z","shell.execute_reply":"2023-08-22T07:03:30.422894Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"transactions = create_df(transactions)\nids = transactions['article_id']\ntransactions = get_article_id(transactions, ids)\n\ndf2 = transactions[['article_id', 'price']]\nt_article_ids = [_id for _id in transactions['article_id']]\nt_prices = [price for price in transactions['price']]\nid_price_dict = create_dict(t_article_ids, t_prices)","metadata":{"execution":{"iopub.status.busy":"2023-08-22T07:03:31.854174Z","iopub.execute_input":"2023-08-22T07:03:31.854638Z","iopub.status.idle":"2023-08-22T07:05:54.572677Z","shell.execute_reply.started":"2023-08-22T07:03:31.854583Z","shell.execute_reply":"2023-08-22T07:05:54.570267Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"res_lst = id_name_price(id_name_dict, id_price_dict)\nids = [lst[0] for lst in res_lst]\nprod_names = [lst[1] for lst in res_lst]\nprices = [lst[2] for lst in res_lst]","metadata":{"execution":{"iopub.status.busy":"2023-08-22T07:07:29.321054Z","iopub.execute_input":"2023-08-22T07:07:29.321819Z","iopub.status.idle":"2023-08-22T07:07:32.654848Z","shell.execute_reply.started":"2023-08-22T07:07:29.321742Z","shell.execute_reply":"2023-08-22T07:07:32.653558Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"final = pd.DataFrame({'article_id' : ids,\n                    'product_name' : prod_names,\n                    'product_price' : prices})\nfinal.head()","metadata":{"execution":{"iopub.status.busy":"2023-08-22T07:07:39.234954Z","iopub.execute_input":"2023-08-22T07:07:39.235425Z","iopub.status.idle":"2023-08-22T07:07:39.411068Z","shell.execute_reply.started":"2023-08-22T07:07:39.235385Z","shell.execute_reply":"2023-08-22T07:07:39.409647Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# pareto chart - 1 : number of purchase\n\ndef df_for_pareto_chart(agg_df:pd.DataFrame, col:str) -> pd.DataFrame:    \n    agg_df['cum_value'] = (agg_df[col].cumsum() / agg_df[col].sum()) * 100\n    return agg_df\n\ndef pareto_chart_viz(df:pd.DataFrame, tar_col:str, left_y_axis:str, right_y_asix:str):\n    \n    color1 = 'steelblue'\n    color2 = 'red'\n    line_size = 5\n\n    fig, ax = plt.subplots()\n    ax.bar(df[tar_col],df[left_y_axis], color = color1)\n    fig.autofmt_xdate(rotation = 45)\n\n    ax2 = ax.twinx()\n    ax2.plot(df[tar_col], df[right_y_asix], color = color2, marker = \"D\", ms = line_size)\n    ax2.yaxis.set_major_formatter(PercentFormatter())\n\n    ax.tick_params(axis = 'y', colors = color1)\n    ax.tick_params(axis = 'y', colors = color2)\n\n    plt.show()","metadata":{"execution":{"iopub.status.busy":"2023-08-22T07:07:47.724837Z","iopub.execute_input":"2023-08-22T07:07:47.725284Z","iopub.status.idle":"2023-08-22T07:07:47.738349Z","shell.execute_reply.started":"2023-08-22T07:07:47.725250Z","shell.execute_reply":"2023-08-22T07:07:47.736705Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"agg_df = final.groupby('product_name')['product_name'].count().reset_index(name = 'num_of_purchase')\nagg_df = agg_df.sort_values(by = 'num_of_purchase', ascending = False)[:10]\ndf = df_for_pareto_chart(agg_df, 'num_of_purchase')\ndf","metadata":{"execution":{"iopub.status.busy":"2023-08-22T07:07:50.543485Z","iopub.execute_input":"2023-08-22T07:07:50.544810Z","iopub.status.idle":"2023-08-22T07:07:50.713612Z","shell.execute_reply.started":"2023-08-22T07:07:50.544751Z","shell.execute_reply":"2023-08-22T07:07:50.712355Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"pareto_chart_viz(df, 'product_name', 'num_of_purchase', 'cum_value')","metadata":{"execution":{"iopub.status.busy":"2023-08-22T07:07:53.889590Z","iopub.execute_input":"2023-08-22T07:07:53.890075Z","iopub.status.idle":"2023-08-22T07:07:54.348986Z","shell.execute_reply.started":"2023-08-22T07:07:53.890034Z","shell.execute_reply":"2023-08-22T07:07:54.347683Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Pareto analysis is that 80% of consequences comes from 20% of causes. Then how apply this method to my project. First I figured out which product did customers purchase. I filtered out top10 products. Then I calculated the cumulative number of purchase to create pareto chart.\n\nInterpretation : Over 80% of all the number of purchase are from the first 7 products. That means we need to order these products more than any other products.","metadata":{}},{"cell_type":"code","source":"# pareto chart 2 - price\n\nagg_df_2 = final.groupby('product_name')['product_price'].sum().reset_index(name = 'sum_price')\nagg_df_2 = agg_df_2.sort_values(by = 'sum_price', ascending = False)[:10]\ndf_2 = df_for_pareto_chart(agg_df_2, 'sum_price')","metadata":{"execution":{"iopub.status.busy":"2023-08-22T07:08:11.061234Z","iopub.execute_input":"2023-08-22T07:08:11.061704Z","iopub.status.idle":"2023-08-22T07:08:11.206317Z","shell.execute_reply.started":"2023-08-22T07:08:11.061659Z","shell.execute_reply":"2023-08-22T07:08:11.205157Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"pareto_chart_viz(df_2, 'product_name', 'sum_price', 'cum_value')","metadata":{"execution":{"iopub.status.busy":"2023-08-22T07:08:13.637261Z","iopub.execute_input":"2023-08-22T07:08:13.637767Z","iopub.status.idle":"2023-08-22T07:08:14.083259Z","shell.execute_reply.started":"2023-08-22T07:08:13.637727Z","shell.execute_reply":"2023-08-22T07:08:14.081997Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### 2.3 Univariate Analysis - Price<a id=2.3></a>","metadata":{}},{"cell_type":"code","source":"data = final['product_price'].tolist()\nfig = plt.figure(figsize =(10, 7))\nplt.boxplot(data)\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-08-22T07:08:28.820364Z","iopub.execute_input":"2023-08-22T07:08:28.820814Z","iopub.status.idle":"2023-08-22T07:08:29.633498Z","shell.execute_reply.started":"2023-08-22T07:08:28.820779Z","shell.execute_reply":"2023-08-22T07:08:29.632141Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"interpretation\n- max : almost 0.6\n- min : 0.0\n- median : 0.03\n- interquartile range : 0.03 - 0.01 = 0.02","metadata":{}},{"cell_type":"markdown","source":"### 2.4 What type of index that is accounted for the most?<a id=2.4></a>","metadata":{}},{"cell_type":"code","source":"def pie_chart(df:pd.DataFrame, col:str):\n    _cnt_of_col = df[col].value_counts()\n    _name_cnt = [tuple((x, y)) for x, y in _cnt_of_col.items()]\n    _vals = [val[1] for val in _name_cnt]\n    _label = [val[0] for val in _name_cnt]\n    plt.pie(_vals, labels = _label,\n            radius = 1.5, autopct = \"%0.2f%%\")\n    plt.show()","metadata":{"execution":{"iopub.status.busy":"2023-08-22T07:14:01.342356Z","iopub.execute_input":"2023-08-22T07:14:01.342890Z","iopub.status.idle":"2023-08-22T07:14:01.353754Z","shell.execute_reply.started":"2023-08-22T07:14:01.342849Z","shell.execute_reply":"2023-08-22T07:14:01.352064Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"pie_chart(articles, 'index_name')","metadata":{"execution":{"iopub.status.busy":"2023-08-22T07:14:03.268762Z","iopub.execute_input":"2023-08-22T07:14:03.269287Z","iopub.status.idle":"2023-08-22T07:14:03.556018Z","shell.execute_reply.started":"2023-08-22T07:14:03.269244Z","shell.execute_reply":"2023-08-22T07:14:03.554836Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"In this dataset, H&M tracks a hierarchy of its products. It seems that the index name is the main category which implies the “sub-categories”, or “index group names”. I’d like to look at this whole pie chart to visualize how the it’s separated by index name.\n\nIt is clear that ‘Ladieswear’ is the largest portion of the index, whereas the ‘Sport’ is undoubtedly the smallest. It is also clear that ‘Ladieswear’ is twice as large a portion as ‘Menswear’.","metadata":{}},{"cell_type":"markdown","source":"### 2.5 What's the portion of index group name for each garment group?<a id=2.5></a>","metadata":{}},{"cell_type":"code","source":"def portion(df:pd.DataFrame, y:str, hue:str):\n    _f, _ax = plt.subplots(figsize = (10, 10))\n    _ax = sns.histplot(data=df, y=y, hue=hue, multiple=\"stack\")\n    plt.show()","metadata":{"execution":{"iopub.status.busy":"2023-08-22T07:14:48.916010Z","iopub.execute_input":"2023-08-22T07:14:48.916892Z","iopub.status.idle":"2023-08-22T07:14:48.924136Z","shell.execute_reply.started":"2023-08-22T07:14:48.916847Z","shell.execute_reply":"2023-08-22T07:14:48.922660Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"portion(articles, 'garment_group_name', 'index_group_name')","metadata":{"execution":{"iopub.status.busy":"2023-08-22T07:14:50.569398Z","iopub.execute_input":"2023-08-22T07:14:50.570576Z","iopub.status.idle":"2023-08-22T07:14:51.666918Z","shell.execute_reply.started":"2023-08-22T07:14:50.570529Z","shell.execute_reply":"2023-08-22T07:14:51.665666Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"As you can see each of garment group has different portion of product index group. You would guess most garment groups have a large portion of Ladieswear. It’s quite obvious that you identified the fact that Ladieswear has the largest portion of the product index.","metadata":{}},{"cell_type":"markdown","source":"# <center><font size=6><a id=3>3. Customer</a></font></center>","metadata":{}},{"cell_type":"code","source":"customers = create_df(customers)\ncustomers.head()","metadata":{"execution":{"iopub.status.busy":"2023-08-22T07:15:07.865765Z","iopub.execute_input":"2023-08-22T07:15:07.866242Z","iopub.status.idle":"2023-08-22T07:15:15.193882Z","shell.execute_reply.started":"2023-08-22T07:15:07.866206Z","shell.execute_reply":"2023-08-22T07:15:15.192094Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### 3.1 How many Customers for each of club member status?(percentage)<a id=3.1></a>","metadata":{}},{"cell_type":"code","source":"def cust_ratio(df:pd.DataFrame, col:str, check_val:str) -> float: \n    _total = df.shape[0]\n    _target = df[df[col] == check_val].shape[0]\n    return round((_target/_total) * 100, 2)","metadata":{"execution":{"iopub.status.busy":"2023-08-22T07:15:45.235915Z","iopub.execute_input":"2023-08-22T07:15:45.236602Z","iopub.status.idle":"2023-08-22T07:15:45.247807Z","shell.execute_reply.started":"2023-08-22T07:15:45.236403Z","shell.execute_reply":"2023-08-22T07:15:45.245898Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"cust_ratio(customers, 'club_member_status', 'ACTIVE')","metadata":{"execution":{"iopub.status.busy":"2023-08-22T07:15:48.028235Z","iopub.execute_input":"2023-08-22T07:15:48.028708Z","iopub.status.idle":"2023-08-22T07:15:48.577444Z","shell.execute_reply.started":"2023-08-22T07:15:48.028670Z","shell.execute_reply":"2023-08-22T07:15:48.576072Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"In this step, you would know what percentage for each of club member status. For example, 'ACTIVE' status in club member status is almost 93%.  ","metadata":{}},{"cell_type":"markdown","source":"### 3.2 How many customers receive the fashion news by frequency?<a id=3.2></a>","metadata":{}},{"cell_type":"code","source":"customers['fashion_news_frequency'] = customers['fashion_news_frequency'].replace('None', 'NONE')","metadata":{"execution":{"iopub.status.busy":"2023-08-22T07:16:07.344441Z","iopub.execute_input":"2023-08-22T07:16:07.345371Z","iopub.status.idle":"2023-08-22T07:16:07.477904Z","shell.execute_reply.started":"2023-08-22T07:16:07.345313Z","shell.execute_reply":"2023-08-22T07:16:07.476420Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def pie_chart(df:pd.DataFrame, col:str):\n    _cnt_of_col = df[col].value_counts()\n    _colname_cnt = [tuple((x, y)) for x, y in _cnt_of_col.items()]\n    \n    _cnt_no_response = df[col].isna().sum()\n    tuple_val = tuple(('No Response', _cnt_no_response))\n    _colname_cnt.append(tuple_val)\n\n    _vals = [val[1] for val in _colname_cnt]\n    _labels = [val[0] for val in _colname_cnt]\n    \n    plt.pie(_vals, labels = _labels, radius = 2, autopct = \"%0.2f%%\")\n    plt.show()","metadata":{"execution":{"iopub.status.busy":"2023-08-22T07:21:26.593749Z","iopub.execute_input":"2023-08-22T07:21:26.594329Z","iopub.status.idle":"2023-08-22T07:21:26.604703Z","shell.execute_reply.started":"2023-08-22T07:21:26.594281Z","shell.execute_reply":"2023-08-22T07:21:26.603216Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"pie_chart(customers, 'fashion_news_frequency')","metadata":{"execution":{"iopub.status.busy":"2023-08-22T07:21:28.010396Z","iopub.execute_input":"2023-08-22T07:21:28.010855Z","iopub.status.idle":"2023-08-22T07:21:28.545207Z","shell.execute_reply.started":"2023-08-22T07:21:28.010820Z","shell.execute_reply":"2023-08-22T07:21:28.543496Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"In this step, I needed to handle the 'fashion_news_frequency' because they have some vague value that may be confusing. It originally had 'None' and 'none' value simultaneously. The 'none' value is supposed to be same value of 'None' so that I replace the 'none' with 'NONE'. The other possible error that you may encountered is that they have 'Nan' value to handle. I normally ignore the 'Nan' value, but this time I decided to keep this because it is possible that some customers didn't answer the question even though they receive the news letter frequently. With that reason, I converted the 'Nan' value into 'No response' in case they missed the chance to answer.","metadata":{}},{"cell_type":"markdown","source":"### 3.3 How old customers overall?<a id=3.3></a>","metadata":{}},{"cell_type":"code","source":"customers = customers.dropna(axis = 0)","metadata":{"execution":{"iopub.status.busy":"2023-08-22T07:22:02.060063Z","iopub.execute_input":"2023-08-22T07:22:02.060506Z","iopub.status.idle":"2023-08-22T07:22:02.804328Z","shell.execute_reply.started":"2023-08-22T07:22:02.060468Z","shell.execute_reply":"2023-08-22T07:22:02.802927Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def distribution(df:pd.DataFrame, col:str):\n    df[col] = df[col].astype(int)\n    df[col].plot.hist()\n    plt.show()","metadata":{"execution":{"iopub.status.busy":"2023-08-22T07:22:03.779272Z","iopub.execute_input":"2023-08-22T07:22:03.779735Z","iopub.status.idle":"2023-08-22T07:22:03.787079Z","shell.execute_reply.started":"2023-08-22T07:22:03.779695Z","shell.execute_reply":"2023-08-22T07:22:03.785577Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"distribution(customers, 'age')","metadata":{"execution":{"iopub.status.busy":"2023-08-22T07:22:05.507844Z","iopub.execute_input":"2023-08-22T07:22:05.508788Z","iopub.status.idle":"2023-08-22T07:22:05.881667Z","shell.execute_reply.started":"2023-08-22T07:22:05.508745Z","shell.execute_reply":"2023-08-22T07:22:05.880690Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The chart shows hte age distribution. Most of people are 20's and 30's age. ","metadata":{}},{"cell_type":"markdown","source":"### 3.4 How many customers for each of club member status(chart)<a id=3.4></a>","metadata":{}},{"cell_type":"code","source":"def bar_plot_for_club_member_status(status : List[str], number_of_customers : List[int]):\n    fig = plt.figure(figsize = (10, 5))\n    plt.bar(status, number_of_customers)","metadata":{"execution":{"iopub.status.busy":"2023-08-22T07:23:08.672345Z","iopub.execute_input":"2023-08-22T07:23:08.672881Z","iopub.status.idle":"2023-08-22T07:23:08.680051Z","shell.execute_reply.started":"2023-08-22T07:23:08.672841Z","shell.execute_reply":"2023-08-22T07:23:08.678844Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"data = customers['club_member_status'].value_counts()\nstatus = list(data.index)\nnumber_of_customers = list(data.values)","metadata":{"execution":{"iopub.status.busy":"2023-08-22T07:23:20.647684Z","iopub.execute_input":"2023-08-22T07:23:20.648119Z","iopub.status.idle":"2023-08-22T07:23:20.732521Z","shell.execute_reply.started":"2023-08-22T07:23:20.648078Z","shell.execute_reply":"2023-08-22T07:23:20.731109Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"bar_plot_for_club_member_status(status, number_of_customers)","metadata":{"execution":{"iopub.status.busy":"2023-08-22T07:23:55.257519Z","iopub.execute_input":"2023-08-22T07:23:55.257990Z","iopub.status.idle":"2023-08-22T07:23:55.491234Z","shell.execute_reply.started":"2023-08-22T07:23:55.257946Z","shell.execute_reply":"2023-08-22T07:23:55.489559Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### 3.5 Correlation : age & club member status<a id=3.5></a>","metadata":{}},{"cell_type":"markdown","source":"In this step, I wanted to figure out how customer age affects club member status. I assumed that the more they're young, the more they're likely to subscribe club member status. For that I did correlation analysis.\n\nCorrelation show the proportionality of two data sets. Simply, I could say y = mx + b. A positive correlation exists between variables 'x' and 'y' if 'm' is a positive value and an increase in 'x' results in an increase in 'y'. Conversely, if two variables have a negative correlation, 'm' will be a negative value, and 'y' will decrease as 'x' increases.\n\nFor me, variable 'x' means 'age' and the other variable is 'number of club member status'","metadata":{}},{"cell_type":"code","source":"active_cust = customers[customers['club_member_status'] == 'ACTIVE']\ndf = active_cust.groupby('age')['club_member_status'].value_counts().reset_index(name = 'number of status')\ndf.head()","metadata":{"execution":{"iopub.status.busy":"2023-08-22T07:26:29.620017Z","iopub.execute_input":"2023-08-22T07:26:29.620420Z","iopub.status.idle":"2023-08-22T07:26:29.936162Z","shell.execute_reply.started":"2023-08-22T07:26:29.620384Z","shell.execute_reply":"2023-08-22T07:26:29.934386Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"val_1 = df['age'].tolist()\nval_2 = df['number of status'].tolist()\n\ndef correlation(df:pd.DataFrame, val_1:list, val_2:list):\n    \n    # number\n    corr, _ = stats.pearsonr(val_1, val_2)\n    res_corr = corr\n    print('correlation score is',res_corr)\n    \n    # viz\n    sns.heatmap(df.corr(), vmin = -1, vmax = 1, annot = True, cmap = 'rocket_r')\n    plt.show()","metadata":{"execution":{"iopub.status.busy":"2023-08-22T07:26:51.547898Z","iopub.execute_input":"2023-08-22T07:26:51.548329Z","iopub.status.idle":"2023-08-22T07:26:51.557165Z","shell.execute_reply.started":"2023-08-22T07:26:51.548295Z","shell.execute_reply":"2023-08-22T07:26:51.555852Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"correlation(df, val_1, val_2)","metadata":{"execution":{"iopub.status.busy":"2023-08-22T07:26:53.892171Z","iopub.execute_input":"2023-08-22T07:26:53.892639Z","iopub.status.idle":"2023-08-22T07:26:54.217859Z","shell.execute_reply.started":"2023-08-22T07:26:53.892588Z","shell.execute_reply":"2023-08-22T07:26:54.216549Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"You may be confused about how to interpret the result of correlation. There's no clear-cut threshold to determine whether data sets are correlated or not. While there's a lot of opinion on this, I used the threshold that is down below.\n\n- -1 to -0.7 : Strong negative correlation\n- -0.7 to -0.5 : Negative correlation\n- -0.5 to 0.5 : No correlation\n- 0.5 to 0.7 : Positive correlation\n- 0.7 to 1 : Strong positive correlation\n\nWe got -0.79 correlation score of previous analysis. That means age and number of club member status have a strong negative correlation, which means as you get older, you're not likely to join the club member.","metadata":{}},{"cell_type":"markdown","source":"### 3.6 Correlation : age & fashion news frequency<a id=3.6></a>","metadata":{}},{"cell_type":"code","source":"age_and_frequency = customers[['age', 'fashion_news_frequency']]\nage_and_frequency.dropna(axis = 0, subset = ['age'], inplace = True)\nage_and_frequency.fillna('No Response', inplace = True)\nage_and_frequency['fashion_news_frequency'] = age_and_frequency['fashion_news_frequency'].replace('NONE', 'None')","metadata":{"execution":{"iopub.status.busy":"2023-08-22T07:28:12.686517Z","iopub.execute_input":"2023-08-22T07:28:12.687062Z","iopub.status.idle":"2023-08-22T07:28:12.818694Z","shell.execute_reply.started":"2023-08-22T07:28:12.687019Z","shell.execute_reply":"2023-08-22T07:28:12.816786Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"age_and_frequency = age_and_frequency.groupby(['age'])['fashion_news_frequency'].value_counts().reset_index(name = 'frequency_status')","metadata":{"execution":{"iopub.status.busy":"2023-08-22T07:28:16.844516Z","iopub.execute_input":"2023-08-22T07:28:16.845026Z","iopub.status.idle":"2023-08-22T07:28:16.971159Z","shell.execute_reply.started":"2023-08-22T07:28:16.844984Z","shell.execute_reply":"2023-08-22T07:28:16.969728Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"a = age_and_frequency['age'].tolist()\nb = age_and_frequency['frequency_status'].tolist()\ncorrelation(age_and_frequency, a, b)","metadata":{"execution":{"iopub.status.busy":"2023-08-22T07:28:28.333512Z","iopub.execute_input":"2023-08-22T07:28:28.334253Z","iopub.status.idle":"2023-08-22T07:28:30.071657Z","shell.execute_reply.started":"2023-08-22T07:28:28.334209Z","shell.execute_reply":"2023-08-22T07:28:30.070287Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Similarly, I calculated the correlation between age and fashion news frequency. I guess that younger people are more subscribe fashion news regularly than the older people. As I assumed, the correlation score is almost -0.8, meaning that as the customer is younger they tend to subscribe the fashion news regularly.","metadata":{}},{"cell_type":"markdown","source":"# <center><font size=6><a id=4>4. Transaction</a><font><center>","metadata":{}},{"cell_type":"code","source":"create_df(transactions)","metadata":{"execution":{"iopub.status.busy":"2023-08-22T07:30:04.231786Z","iopub.execute_input":"2023-08-22T07:30:04.232286Z","iopub.status.idle":"2023-08-22T07:30:04.282510Z","shell.execute_reply.started":"2023-08-22T07:30:04.232243Z","shell.execute_reply":"2023-08-22T07:30:04.280733Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"ids = transactions['article_id']\ntransactions = get_article_id(transactions, ids)","metadata":{"execution":{"iopub.status.busy":"2023-08-22T07:30:10.199208Z","iopub.execute_input":"2023-08-22T07:30:10.200333Z","iopub.status.idle":"2023-08-22T07:30:30.072134Z","shell.execute_reply.started":"2023-08-22T07:30:10.200282Z","shell.execute_reply":"2023-08-22T07:30:30.070908Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"ids = transactions['article_id']\ntransactions = get_article_id(transactions, ids)","metadata":{"execution":{"iopub.status.busy":"2023-08-22T07:30:36.965486Z","iopub.execute_input":"2023-08-22T07:30:36.966447Z","iopub.status.idle":"2023-08-22T07:30:56.934143Z","shell.execute_reply.started":"2023-08-22T07:30:36.966395Z","shell.execute_reply":"2023-08-22T07:30:56.932678Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"transactions.head()","metadata":{"execution":{"iopub.status.busy":"2023-08-22T07:31:03.206558Z","iopub.execute_input":"2023-08-22T07:31:03.207194Z","iopub.status.idle":"2023-08-22T07:31:03.223693Z","shell.execute_reply.started":"2023-08-22T07:31:03.207148Z","shell.execute_reply":"2023-08-22T07:31:03.222390Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### 4.1 Price Distribution<a id=4.1></a>","metadata":{}},{"cell_type":"code","source":"def price_distribution(df:pd.DataFrame, col:str):\n    df[col].plot.hist()\n    plt.show()","metadata":{"execution":{"iopub.status.busy":"2023-08-22T07:31:11.426334Z","iopub.execute_input":"2023-08-22T07:31:11.426785Z","iopub.status.idle":"2023-08-22T07:31:11.433858Z","shell.execute_reply.started":"2023-08-22T07:31:11.426738Z","shell.execute_reply":"2023-08-22T07:31:11.432426Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"price_distribution(transactions, 'price')","metadata":{"execution":{"iopub.status.busy":"2023-08-22T07:31:14.529118Z","iopub.execute_input":"2023-08-22T07:31:14.529553Z","iopub.status.idle":"2023-08-22T07:31:20.455753Z","shell.execute_reply.started":"2023-08-22T07:31:14.529520Z","shell.execute_reply":"2023-08-22T07:31:20.454724Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### 4.2 Which products do they buy the most?<a id=4.2></a>","metadata":{}},{"cell_type":"code","source":"def loyal_cust(df:pd.DataFrame, col:str, top_n:int):\n    top_n_cust = df[col].value_counts()[:top_n]\n    return top_n_cust","metadata":{"execution":{"iopub.status.busy":"2023-08-22T07:32:32.641851Z","iopub.execute_input":"2023-08-22T07:32:32.642351Z","iopub.status.idle":"2023-08-22T07:32:32.649952Z","shell.execute_reply.started":"2023-08-22T07:32:32.642312Z","shell.execute_reply":"2023-08-22T07:32:32.648437Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"loyal_cust(transactions, 'article_id', 10)","metadata":{"execution":{"iopub.status.busy":"2023-08-22T07:32:35.011784Z","iopub.execute_input":"2023-08-22T07:32:35.012232Z","iopub.status.idle":"2023-08-22T07:32:48.549353Z","shell.execute_reply.started":"2023-08-22T07:32:35.012197Z","shell.execute_reply":"2023-08-22T07:32:48.548019Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We can basically answer to the question by counting each of article_id that is represented the unique product id. The product id ‘706016001’ is the most popular.","metadata":{}},{"cell_type":"markdown","source":"### 4.3 What is the mean price for each product by index_name?<a id=4.3></a>","metadata":{}},{"cell_type":"code","source":"articles_df = articles[['article_id','index_name', 'product_group_name']]","metadata":{"execution":{"iopub.status.busy":"2023-08-22T07:33:16.698790Z","iopub.execute_input":"2023-08-22T07:33:16.699272Z","iopub.status.idle":"2023-08-22T07:33:16.716817Z","shell.execute_reply.started":"2023-08-22T07:33:16.699236Z","shell.execute_reply":"2023-08-22T07:33:16.715214Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"transactions_df = transactions[['article_id', 'price']]\nids = transactions_df['article_id']\ntransactions_df = get_article_id(transactions_df, ids)","metadata":{"execution":{"iopub.status.busy":"2023-08-22T07:33:18.378701Z","iopub.execute_input":"2023-08-22T07:33:18.379241Z","iopub.status.idle":"2023-08-22T07:33:36.977735Z","shell.execute_reply.started":"2023-08-22T07:33:18.379198Z","shell.execute_reply":"2023-08-22T07:33:36.976380Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"merge_df = transactions_df.merge(articles_df, on='article_id')\nmerge_df.head()","metadata":{"execution":{"iopub.status.busy":"2023-08-22T07:33:38.955209Z","iopub.execute_input":"2023-08-22T07:33:38.955696Z","iopub.status.idle":"2023-08-22T07:33:51.901308Z","shell.execute_reply.started":"2023-08-22T07:33:38.955652Z","shell.execute_reply":"2023-08-22T07:33:51.900296Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def mean_price(df:pd.DataFrame, groupby_col:str):\n    \n    res = df.groupby(groupby_col)['price'].mean().reset_index()\n    res = res.sort_values(by = 'price', ascending = False)\n\n    sns.set_style('darkgrid')\n    f,ax = plt.subplots(figsize = (10, 5))\n    ax = sns.barplot(x = res.price, y = res.index_name, color = 'pink', alpha = 0.8)\n    plt.show()","metadata":{"execution":{"iopub.status.busy":"2023-08-22T07:34:16.178973Z","iopub.execute_input":"2023-08-22T07:34:16.179522Z","iopub.status.idle":"2023-08-22T07:34:16.190726Z","shell.execute_reply.started":"2023-08-22T07:34:16.179484Z","shell.execute_reply":"2023-08-22T07:34:16.188961Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"mean_price(merge_df, 'index_name')","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"It can be seen that the mean price of Ladies wear is the highest that is over 0.03. Compared to this, the lowest mean price is Baby size clothes which is between 0.015 and 0.020.","metadata":{}},{"cell_type":"markdown","source":"# <center><font size=6><a id=5>5. Image</a><font><center>","metadata":{}},{"cell_type":"code","source":"path = '../input/h-and-m-personalized-fashion-recommendations/images'\n\nf_lst = []\nfor filename in os.listdir(path):\n    res = os.path.join(path, filename)\n    for file in os.listdir(res):\n        f_lst.append(file)\n\nfor article_id in articles['article_id']:\n    folder = article_id[:3]\n    img = f'{article_id}.jpg'\n    if img in f_lst:\n        path = f'../input/h-and-m-personalized-fashion-recommendations/images/{folder}/{img}'\n        res = Image.open(path)\n        plt.imshow(res)\n        break","metadata":{"trusted":true},"execution_count":null,"outputs":[]}]}