{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"![](https://media-cldnry.s-nbcnews.com/image/upload/t_social_share_1200x630_center,f_auto,q_auto:best/newscms/2017_24/1222336/hm-today-170616-tease.jpg)","metadata":{}},{"cell_type":"markdown","source":"Hennes & Mauritz AB is a Swedish multinational clothing company headquartered in Stockholm. It is known for its fast-fashion clothing for men, women, teenagers, and children","metadata":{}},{"cell_type":"markdown","source":"### Data background","metadata":{}},{"cell_type":"markdown","source":"The dataset contains 4 csv files and one folder with several subfolders, each with a different number of images.\n\nIn this Exploratory Data Analysis Notebook we will look to the data, will analyze the content of each csv file, check for missing data, understand the data distribution, see what are the relations between data in various files. There are three tabular data files.\n\n* Customer Data\n* Article Data\n* Transaction Data\n","metadata":{}},{"cell_type":"markdown","source":"### Importing Libraries\n\nWe will include here the required packages for reading, parsing, filtering, processing, visualizing the data, both tabular and image.","metadata":{}},{"cell_type":"code","source":"import numpy as np\nimport pandas as pd\nimport os\nfrom tqdm import tqdm\nimport plotly.express as px\nimport matplotlib.pyplot as plt\nimport seaborn as sns\nfrom wordcloud import WordCloud, STOPWORDS\nfrom datetime import datetime","metadata":{"execution":{"iopub.status.busy":"2022-02-24T13:12:46.849327Z","iopub.execute_input":"2022-02-24T13:12:46.850160Z","iopub.status.idle":"2022-02-24T13:12:46.857303Z","shell.execute_reply.started":"2022-02-24T13:12:46.850125Z","shell.execute_reply":"2022-02-24T13:12:46.856125Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Creating a class with data locations\n","metadata":{}},{"cell_type":"code","source":"class DataLocations:\n    article_csv = '../input/h-and-m-personalized-fashion-recommendations/articles.csv'\n    customer_csv = '../input/h-and-m-personalized-fashion-recommendations/customers.csv'\n    tx_csv = '../input/h-and-m-personalized-fashion-recommendations/transactions_train.csv'\n    sub_csv = '../input/h-and-m-personalized-fashion-recommendations/sample_submission.csv'","metadata":{"execution":{"iopub.status.busy":"2022-02-24T13:04:57.166198Z","iopub.execute_input":"2022-02-24T13:04:57.166393Z","iopub.status.idle":"2022-02-24T13:04:57.170796Z","shell.execute_reply.started":"2022-02-24T13:04:57.166370Z","shell.execute_reply":"2022-02-24T13:04:57.170012Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Customer Data\nWe will start analysis with customer data","metadata":{}},{"cell_type":"code","source":"customer_df = pd.read_csv(DataLocations.customer_csv)\ncustomer_df.head()","metadata":{"execution":{"iopub.status.busy":"2022-02-24T13:04:57.171802Z","iopub.execute_input":"2022-02-24T13:04:57.172481Z","iopub.status.idle":"2022-02-24T13:05:03.353817Z","shell.execute_reply.started":"2022-02-24T13:04:57.172451Z","shell.execute_reply":"2022-02-24T13:05:03.353114Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"it is obvious that there are some null values. We will insect more into this","metadata":{}},{"cell_type":"code","source":"print(\"shape of data Customer data\",customer_df.shape)","metadata":{"execution":{"iopub.status.busy":"2022-02-24T13:05:03.355919Z","iopub.execute_input":"2022-02-24T13:05:03.356197Z","iopub.status.idle":"2022-02-24T13:05:03.362884Z","shell.execute_reply.started":"2022-02-24T13:05:03.356159Z","shell.execute_reply":"2022-02-24T13:05:03.361626Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"we can see that we have `1371980` rows and `7` columns","metadata":{}},{"cell_type":"code","source":"customer_df.isnull().sum()","metadata":{"execution":{"iopub.status.busy":"2022-02-24T13:05:03.363976Z","iopub.execute_input":"2022-02-24T13:05:03.364148Z","iopub.status.idle":"2022-02-24T13:05:03.955815Z","shell.execute_reply.started":"2022-02-24T13:05:03.364124Z","shell.execute_reply":"2022-02-24T13:05:03.954938Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"we can see that there are many null values in the dataset. To have a clear view lets take these into a plot as precentages","metadata":{}},{"cell_type":"code","source":"# Function to plot the Nan percentages of each columns\ndef plot_nas(df):\n    if df.isnull().sum().sum() != 0:\n        na_df = (df.isnull().sum() / len(df)) * 100      \n        na_df = na_df.drop(na_df[na_df == 0].index).sort_values(ascending=False)\n        missing_data = pd.DataFrame({'Missing Ratio %' :na_df})\n        missing_data.plot(kind = \"barh\")\n        plt.show()\n    else:\n        print('No NAs found')\n\nprint(\"Checking Null's in Customer data \")\nplot_nas(customer_df)","metadata":{"execution":{"iopub.status.busy":"2022-02-24T13:05:03.956926Z","iopub.execute_input":"2022-02-24T13:05:03.957178Z","iopub.status.idle":"2022-02-24T13:05:05.294800Z","shell.execute_reply.started":"2022-02-24T13:05:03.957146Z","shell.execute_reply":"2022-02-24T13:05:05.293905Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"As shown in the plot `FN` and `Active` columns have large number of null values. We have to decide whether we can use these columns to build the model. So we have two options. \n\n* Fill null values and use the colums \n* Remove the columns and go with the rest\n\nHowever to take one of above options we have to analyze other data as well.\nLet's see how many uniques values are in the columns","metadata":{}},{"cell_type":"code","source":"print(customer_df.nunique())","metadata":{"execution":{"iopub.status.busy":"2022-02-24T13:05:05.295957Z","iopub.execute_input":"2022-02-24T13:05:05.296199Z","iopub.status.idle":"2022-02-24T13:05:06.607338Z","shell.execute_reply.started":"2022-02-24T13:05:05.296163Z","shell.execute_reply":"2022-02-24T13:05:06.606582Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"As shown above in Active column, there is only one value which is `Active`. We can assume that NaN values of Active ","metadata":{}},{"cell_type":"code","source":"temp = customer_df.groupby([\"age\"])[\"customer_id\"].count()\ndf = pd.DataFrame({'age':temp.index,'count':temp.values})\ndf = df.sort_values(['age'],ascending=False)\nplt.figure(figsize=(35,7))\nplt.title(\"Number of Customers by Age\")\nsns.set_color_codes(\"pastel\")\ns = sns.barplot(x = 'age', y=\"count\", data=df)\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-02-24T13:05:06.608527Z","iopub.execute_input":"2022-02-24T13:05:06.608828Z","iopub.status.idle":"2022-02-24T13:05:07.895248Z","shell.execute_reply.started":"2022-02-24T13:05:06.608802Z","shell.execute_reply":"2022-02-24T13:05:07.894178Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We can identify that the most of customers are between age of 20 to 30.","metadata":{}},{"cell_type":"code","source":"temp = customer_df.groupby([\"fashion_news_frequency\"])[\"customer_id\"].count()\ndf = pd.DataFrame({'Fashion News Frequency': temp.index,\n                   'Customers': temp.values\n                  })\ndf = df.sort_values(['Customers'], ascending=False)\nplt.figure(figsize = (6,6))\nplt.title(f'Number of Customers per each Fashion News Frequency')\nsns.set_color_codes(\"pastel\")\ns = sns.barplot(x = 'Fashion News Frequency', y=\"Customers\", data=df)\ns.set_xticklabels(s.get_xticklabels(),rotation=90)\nlocs, labels = plt.xticks()\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-02-24T13:05:07.896942Z","iopub.execute_input":"2022-02-24T13:05:07.897619Z","iopub.status.idle":"2022-02-24T13:05:08.277779Z","shell.execute_reply.started":"2022-02-24T13:05:07.897594Z","shell.execute_reply":"2022-02-24T13:05:08.276801Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(temp)\ncount = 0\nfor i in temp:\n    count = count + i\nprint(\"Precentage of customers who have subscribed to regulary news : \",round(temp[3]/count*100,2) ,\"%\")","metadata":{"execution":{"iopub.status.busy":"2022-02-24T13:05:08.280790Z","iopub.execute_input":"2022-02-24T13:05:08.281603Z","iopub.status.idle":"2022-02-24T13:05:08.293216Z","shell.execute_reply.started":"2022-02-24T13:05:08.281544Z","shell.execute_reply":"2022-02-24T13:05:08.292429Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"And most of the customer have not subscribed to fashion news. However 35% of customers have subscribed to regularly news","metadata":{}},{"cell_type":"code","source":"temp = customer_df.groupby([\"club_member_status\"])[\"customer_id\"].count()\ndf = pd.DataFrame({'Club Member Status': temp.index,\n                   'Customers': temp.values\n                  })\ndf = df.sort_values(['Customers'], ascending=False)\nplt.figure(figsize = (6,6))\nplt.title(f'Number of Customers per each Club Member Status')\nsns.set_color_codes(\"pastel\")\ns = sns.barplot(x = 'Club Member Status', y=\"Customers\", data=df)\ns.set_xticklabels(s.get_xticklabels(),rotation=90)\nlocs, labels = plt.xticks()\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-02-24T13:05:08.296782Z","iopub.execute_input":"2022-02-24T13:05:08.297682Z","iopub.status.idle":"2022-02-24T13:05:08.814545Z","shell.execute_reply.started":"2022-02-24T13:05:08.297610Z","shell.execute_reply":"2022-02-24T13:05:08.813293Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"As shown in the bar chart most of the customers has an Active club member status. Let's see how many has subscribed to the news regulary from active members.","metadata":{}},{"cell_type":"code","source":"# method to plot club member status bar chart\ndef plot_bar(df, column):\n    long_df = pd.DataFrame(df.groupby(column)['customer_id'].count().reset_index().rename({'customer_id': 'count'}, axis=1))\n    fig = px.bar(long_df, x=column, y=\"count\", color=column, title=\"bar plot for {column} \")\n    fig.show()\n    ","metadata":{"execution":{"iopub.status.busy":"2022-02-24T13:05:08.815786Z","iopub.execute_input":"2022-02-24T13:05:08.815986Z","iopub.status.idle":"2022-02-24T13:05:08.822237Z","shell.execute_reply.started":"2022-02-24T13:05:08.815959Z","shell.execute_reply":"2022-02-24T13:05:08.821615Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plot_bar( customer_df, 'club_member_status')","metadata":{"execution":{"iopub.status.busy":"2022-02-24T13:05:08.823426Z","iopub.execute_input":"2022-02-24T13:05:08.823633Z","iopub.status.idle":"2022-02-24T13:05:10.118172Z","shell.execute_reply.started":"2022-02-24T13:05:08.823609Z","shell.execute_reply":"2022-02-24T13:05:10.117133Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"temp = customer_df.groupby([\"club_member_status\",\"fashion_news_frequency\"])[\"customer_id\"].count()\ntemp\nprint(\"The precentage of customers who have active member status from all the customers who have active status is \",round(471304/477416,2),\"%\")","metadata":{"execution":{"iopub.status.busy":"2022-02-24T13:05:10.119054Z","iopub.execute_input":"2022-02-24T13:05:10.119236Z","iopub.status.idle":"2022-02-24T13:05:10.720782Z","shell.execute_reply.started":"2022-02-24T13:05:10.119206Z","shell.execute_reply":"2022-02-24T13:05:10.719917Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We can see that 99% of the customers who have subscribed to news have an Active club member status.","metadata":{}},{"cell_type":"code","source":"temp_df = temp.to_frame()\n#sns.barplot(x=\"club_member_status\", y=\"customer_id\", hue=\"fashion_news_frequency\", data=temp_df)","metadata":{"execution":{"iopub.status.busy":"2022-02-24T13:05:10.721885Z","iopub.execute_input":"2022-02-24T13:05:10.722525Z","iopub.status.idle":"2022-02-24T13:05:10.727602Z","shell.execute_reply.started":"2022-02-24T13:05:10.722488Z","shell.execute_reply":"2022-02-24T13:05:10.726473Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fig = plt.figure(figsize=(20, 7))\nsns.histplot(customer_df.postal_code.value_counts()[1:], bins=250, kde=False)\nplt.xlim(0, 50)\nplt.tight_layout()\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-02-24T13:05:10.728993Z","iopub.execute_input":"2022-02-24T13:05:10.729230Z","iopub.status.idle":"2022-02-24T13:05:12.540633Z","shell.execute_reply.started":"2022-02-24T13:05:10.729199Z","shell.execute_reply":"2022-02-24T13:05:12.539987Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We can see here that most of the customers are from one particular postal code. ","metadata":{}},{"cell_type":"markdown","source":"### Article Data\nWe will start analysis with article data","metadata":{"execution":{"iopub.status.busy":"2022-02-24T03:59:09.483662Z","iopub.execute_input":"2022-02-24T03:59:09.483981Z","iopub.status.idle":"2022-02-24T03:59:09.490625Z","shell.execute_reply.started":"2022-02-24T03:59:09.483938Z","shell.execute_reply":"2022-02-24T03:59:09.489138Z"}}},{"cell_type":"code","source":"articles_df = pd.read_csv(DataLocations.article_csv)\narticles_df.head()","metadata":{"execution":{"iopub.status.busy":"2022-02-24T13:05:12.541758Z","iopub.execute_input":"2022-02-24T13:05:12.542125Z","iopub.status.idle":"2022-02-24T13:05:13.817492Z","shell.execute_reply.started":"2022-02-24T13:05:12.542094Z","shell.execute_reply":"2022-02-24T13:05:13.816469Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"articles_df.shape","metadata":{"execution":{"iopub.status.busy":"2022-02-24T13:05:13.818850Z","iopub.execute_input":"2022-02-24T13:05:13.819054Z","iopub.status.idle":"2022-02-24T13:05:13.825899Z","shell.execute_reply.started":"2022-02-24T13:05:13.819026Z","shell.execute_reply":"2022-02-24T13:05:13.824570Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"article_csv has 25 columns and 105542 rows","metadata":{}},{"cell_type":"code","source":"temp = articles_df.groupby([\"product_group_name\"])[\"product_type_name\"].nunique()\ndf = pd.DataFrame({'Product Group': temp.index,\n                   'Product Types': temp.values\n                  })\ndf = df.sort_values(['Product Types'], ascending=False)\nplt.figure(figsize = (8,6))\nplt.title('Number of Product Types per each Product Group')\nsns.set_color_codes(\"pastel\")\ns = sns.barplot(x = 'Product Group', y=\"Product Types\", data=df,palette=\"cubehelix\")\ns.set_xticklabels(s.get_xticklabels(),rotation=90)\nlocs, labels = plt.xticks()\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-02-24T13:21:38.667157Z","iopub.execute_input":"2022-02-24T13:21:38.668073Z","iopub.status.idle":"2022-02-24T13:21:38.965031Z","shell.execute_reply.started":"2022-02-24T13:21:38.667977Z","shell.execute_reply":"2022-02-24T13:21:38.964165Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Most of the products are from accessories.And there is also an unknown category.","metadata":{}},{"cell_type":"code","source":"stopwords = set(STOPWORDS)\n\ndef show_wordcloud(data, title = None):\n    wordcloud = WordCloud(\n        background_color='white',\n        stopwords=stopwords,\n        max_words=200,\n        max_font_size=40, \n        scale=5,\n        random_state=1\n    ).generate(str(data))\n\n    fig = plt.figure(1, figsize=(10,10))\n    plt.axis('off')\n    if title: \n        fig.suptitle(title, fontsize=14)\n        fig.subplots_adjust(top=2.3)\n\n    plt.imshow(wordcloud)\n    plt.show()","metadata":{"execution":{"iopub.status.busy":"2022-02-24T13:05:14.408785Z","iopub.execute_input":"2022-02-24T13:05:14.409193Z","iopub.status.idle":"2022-02-24T13:05:14.416792Z","shell.execute_reply.started":"2022-02-24T13:05:14.409166Z","shell.execute_reply":"2022-02-24T13:05:14.416002Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"show_wordcloud(articles_df[\"prod_name\"], \"Wordcloud from product name\")","metadata":{"execution":{"iopub.status.busy":"2022-02-24T13:05:14.418102Z","iopub.execute_input":"2022-02-24T13:05:14.419127Z","iopub.status.idle":"2022-02-24T13:05:15.062223Z","shell.execute_reply.started":"2022-02-24T13:05:14.418954Z","shell.execute_reply":"2022-02-24T13:05:15.061371Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"temp = articles_df.groupby([\"product_group_name\"])[\"article_id\"].nunique()\ndf = pd.DataFrame({'Product Group': temp.index,\n                   'Articles': temp.values\n                  })\ndf = df.sort_values(['Articles'], ascending=False)\nplt.figure(figsize = (8,6))\nplt.title('Number of Articles per each Product Group')\nsns.set_color_codes(\"pastel\")\ns = sns.barplot(x = 'Product Group', y=\"Articles\", data=df)\ns.set_xticklabels(s.get_xticklabels(),rotation=90)\nlocs, labels = plt.xticks()\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-02-24T13:14:41.904270Z","iopub.execute_input":"2022-02-24T13:14:41.905067Z","iopub.status.idle":"2022-02-24T13:14:42.202656Z","shell.execute_reply.started":"2022-02-24T13:14:41.905031Z","shell.execute_reply":"2022-02-24T13:14:42.201881Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Most of the articles are from the `Garmenr Upper Body` product group. Let's see it as a percenetage","metadata":{}},{"cell_type":"code","source":"temp = articles_df.groupby([\"product_group_name\"])['article_id'].nunique().sort_values(ascending=False)\ntemp","metadata":{"execution":{"iopub.status.busy":"2022-02-24T13:05:15.394820Z","iopub.execute_input":"2022-02-24T13:05:15.395136Z","iopub.status.idle":"2022-02-24T13:05:15.430836Z","shell.execute_reply.started":"2022-02-24T13:05:15.395084Z","shell.execute_reply":"2022-02-24T13:05:15.429932Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(\"Garmenr Upper Body articles as a percenetage of total articles : \",\n   round(temp[0]/articles_df['article_id'].\n         count()*100,2),\"%\")","metadata":{"execution":{"iopub.status.busy":"2022-02-24T13:05:15.431811Z","iopub.execute_input":"2022-02-24T13:05:15.432005Z","iopub.status.idle":"2022-02-24T13:05:15.438498Z","shell.execute_reply.started":"2022-02-24T13:05:15.431978Z","shell.execute_reply":"2022-02-24T13:05:15.437715Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"temp = articles_df.groupby([\"product_type_name\"])[\"article_id\"].nunique()\ndf = pd.DataFrame({'Product Type': temp.index,\n                   'Articles': temp.values\n                  })\ntotal_types = len(df['Product Type'].unique())\n\n#getting top 50 \ndf = df.sort_values(['Articles'], ascending=False)[0:50]\nplt.figure(figsize = (16,6))\nplt.title(f'Number of Articles per each Product Type (top 50 from total: {total_types})')\ns = sns.barplot(x = 'Product Type', y=\"Articles\", data=df,palette=\"rocket\")\ns.set_xticklabels(s.get_xticklabels(),rotation=90)\nlocs, labels = plt.xticks()\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-02-24T13:16:28.799039Z","iopub.execute_input":"2022-02-24T13:16:28.799260Z","iopub.status.idle":"2022-02-24T13:16:29.789055Z","shell.execute_reply.started":"2022-02-24T13:16:28.799238Z","shell.execute_reply":"2022-02-24T13:16:29.787771Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Most of the articles are from four product types\n* trousers\n* Dress\n* Sweater\n* T-shirt\n\n","metadata":{}},{"cell_type":"code","source":"temp = articles_df.groupby([\"department_name\"])[\"article_id\"].nunique()\ndf = pd.DataFrame({'Department Name': temp.index,\n                   'Articles': temp.values\n                  })\ntotal_depts = len(df['Department Name'].unique())\ndf = df.sort_values(['Articles'], ascending=False).head(50)\nplt.figure(figsize = (16,6))\nplt.title(f'Number of Articles per each Department (top 50 from total: {total_depts})')\nsns.set_color_codes(\"pastel\")\ns = sns.barplot(x = 'Department Name', y=\"Articles\", data=df,palette=\"CMRmap\")\ns.set_xticklabels(s.get_xticklabels(),rotation=90)\nlocs, labels = plt.xticks()\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-02-24T13:19:45.107572Z","iopub.execute_input":"2022-02-24T13:19:45.108023Z","iopub.status.idle":"2022-02-24T13:19:45.780511Z","shell.execute_reply.started":"2022-02-24T13:19:45.107995Z","shell.execute_reply":"2022-02-24T13:19:45.779646Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"It can be identified that most of the articles are from `jersey` deparment.","metadata":{}},{"cell_type":"code","source":"temp = articles_df.groupby([\"graphical_appearance_name\"])[\"article_id\"].nunique()\ndf = pd.DataFrame({'Graphical Appearance Name': temp.index,\n                   'Articles': temp.values\n                  })\ndf = df.sort_values(['Articles'], ascending=False).head(50)\nplt.figure(figsize = (16,6))\nplt.title(f'Number of Articles per each Graphical Appearance Name')\nsns.set_color_codes(\"pastel\")\ns = sns.barplot(x = 'Graphical Appearance Name', y=\"Articles\", data=df,palette=\"crest\")\ns.set_xticklabels(s.get_xticklabels(),rotation=90)\nlocs, labels = plt.xticks()\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-02-24T13:22:40.162444Z","iopub.execute_input":"2022-02-24T13:22:40.162689Z","iopub.status.idle":"2022-02-24T13:22:40.537530Z","shell.execute_reply.started":"2022-02-24T13:22:40.162665Z","shell.execute_reply":"2022-02-24T13:22:40.535856Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"temp = articles_df.groupby([\"index_group_name\"])[\"article_id\"].nunique()\ndf = pd.DataFrame({'Index Group Name': temp.index,\n                   'Articles': temp.values\n                  })\ndf = df.sort_values(['Articles'], ascending=False)\nplt.figure(figsize = (6,6))\nplt.title(f'Number of Articles per each Index Group Name')\nsns.set_color_codes(\"pastel\")\ns = sns.barplot(x = 'Index Group Name', y=\"Articles\", data=df)\ns.set_xticklabels(s.get_xticklabels(),rotation=90)\nlocs, labels = plt.xticks()\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-02-24T13:24:02.273198Z","iopub.execute_input":"2022-02-24T13:24:02.273664Z","iopub.status.idle":"2022-02-24T13:24:02.464477Z","shell.execute_reply.started":"2022-02-24T13:24:02.273634Z","shell.execute_reply":"2022-02-24T13:24:02.463422Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"`Ladieswear` has the most of the aricles and `Babychildren` also has a large number of articles","metadata":{}},{"cell_type":"code","source":"temp = articles_df.groupby([\"colour_group_name\"])[\"article_id\"].nunique()\ndf = pd.DataFrame({'Colour Group Name': temp.index,\n                   'Articles': temp.values\n                  })\ndf = df.sort_values(['Articles'], ascending=False)\nplt.figure(figsize = (12,6))\nplt.title(f'Number of Articles per each Colour Group Name')\nsns.set_color_codes(\"pastel\")\ns = sns.barplot(x = 'Colour Group Name', y=\"Articles\", data=df)\ns.set_xticklabels(s.get_xticklabels(),rotation=90)\nlocs, labels = plt.xticks()\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-02-24T13:24:15.923720Z","iopub.execute_input":"2022-02-24T13:24:15.924003Z","iopub.status.idle":"2022-02-24T13:24:16.594580Z","shell.execute_reply.started":"2022-02-24T13:24:15.923969Z","shell.execute_reply":"2022-02-24T13:24:16.593510Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We can see that most of the products are `black`","metadata":{}},{"cell_type":"code","source":"temp = articles_df.groupby([\"perceived_colour_value_name\"])[\"article_id\"].nunique()\ndf = pd.DataFrame({'Perceived Colour Group Name': temp.index,\n                   'Articles': temp.values\n                  })\ndf = df.sort_values(['Articles'], ascending=False)\nplt.figure(figsize = (6,6))\nplt.title(f'Number of Articles per each Perceived Colour Group Name')\nsns.set_color_codes(\"pastel\")\ns = sns.barplot(x = 'Perceived Colour Group Name', y=\"Articles\", data=df)\ns.set_xticklabels(s.get_xticklabels(),rotation=90)\nlocs, labels = plt.xticks()\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-02-24T13:27:20.738743Z","iopub.execute_input":"2022-02-24T13:27:20.738998Z","iopub.status.idle":"2022-02-24T13:27:20.939522Z","shell.execute_reply.started":"2022-02-24T13:27:20.738974Z","shell.execute_reply":"2022-02-24T13:27:20.938716Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We can see that most of the products are Dark color. We also identifed that Black is the most famous color in products","metadata":{}},{"cell_type":"markdown","source":"","metadata":{}},{"cell_type":"markdown","source":"","metadata":{}},{"cell_type":"markdown","source":"","metadata":{}},{"cell_type":"markdown","source":"","metadata":{}},{"cell_type":"markdown","source":"","metadata":{}},{"cell_type":"markdown","source":"","metadata":{}}]}