{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"## Step 1: Imports and Reading Data","metadata":{}},{"cell_type":"code","source":"import pandas as pd\nimport numpy as np\nimport matplotlib.pylab as plt #for visualizing dataset\nimport seaborn as sns\n#for using style in matplotlib and seaborn style need to inslall these packages\nplt.style.use('ggplot')","metadata":{"execution":{"iopub.status.busy":"2023-08-03T19:02:42.123135Z","iopub.execute_input":"2023-08-03T19:02:42.123697Z","iopub.status.idle":"2023-08-03T19:02:42.133349Z","shell.execute_reply.started":"2023-08-03T19:02:42.123653Z","shell.execute_reply":"2023-08-03T19:02:42.131792Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#Read the dataset\n\ndf = pd.read_csv('/kaggle/input/h-and-m-personalized-fashion-recommendations/articles.csv')","metadata":{"execution":{"iopub.status.busy":"2023-08-03T19:02:42.136286Z","iopub.execute_input":"2023-08-03T19:02:42.136848Z","iopub.status.idle":"2023-08-03T19:02:42.968345Z","shell.execute_reply.started":"2023-08-03T19:02:42.136801Z","shell.execute_reply":"2023-08-03T19:02:42.967262Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Step 2: Data Understanding\nWhat I can do here:\n* Shape of the Dataframe\n* Head and tail of Dataframe\n* Data types\n* Description of Dataset","metadata":{}},{"cell_type":"code","source":"#Shape of Dataframe\ndf.shape","metadata":{"execution":{"iopub.status.busy":"2023-08-03T19:02:42.970166Z","iopub.execute_input":"2023-08-03T19:02:42.970630Z","iopub.status.idle":"2023-08-03T19:02:42.980542Z","shell.execute_reply.started":"2023-08-03T19:02:42.970586Z","shell.execute_reply":"2023-08-03T19:02:42.979318Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The dataset has 105542 rows and 25 Columns","metadata":{}},{"cell_type":"code","source":"#To see the first couple of rows of the dataset\ndf.head(10)","metadata":{"execution":{"iopub.status.busy":"2023-08-03T19:02:42.984300Z","iopub.execute_input":"2023-08-03T19:02:42.985201Z","iopub.status.idle":"2023-08-03T19:02:43.016730Z","shell.execute_reply.started":"2023-08-03T19:02:42.985154Z","shell.execute_reply":"2023-08-03T19:02:43.015311Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"* By default, it shows us the first 5 columns of the dataset, but if we want to see more, we can put the desired number in the bracket","metadata":{}},{"cell_type":"code","source":"#Due to the higher number of columns in the dataset, it does not show us all the columns in the \"HEAD\" function. To see all of them we can load this function:\npd.set_option('display.max_columns', 200)\ndf.head()","metadata":{"execution":{"iopub.status.busy":"2023-08-03T19:02:43.018805Z","iopub.execute_input":"2023-08-03T19:02:43.019281Z","iopub.status.idle":"2023-08-03T19:02:43.052219Z","shell.execute_reply.started":"2023-08-03T19:02:43.019238Z","shell.execute_reply":"2023-08-03T19:02:43.050862Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#To see all the columns name\ndf.columns","metadata":{"execution":{"iopub.status.busy":"2023-08-03T19:02:43.054156Z","iopub.execute_input":"2023-08-03T19:02:43.054792Z","iopub.status.idle":"2023-08-03T19:02:43.064193Z","shell.execute_reply.started":"2023-08-03T19:02:43.054741Z","shell.execute_reply":"2023-08-03T19:02:43.062899Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Articles\n* **article_id** : A unique identifier of every article.\n* **product_code, prod_name** : A unique identifier of every product and its name\n* **product_type, product_type_name** : The group of product_code and its name\n* **graphical_appearance_no, graphical_appearance_name** : The group of graphics and its name\n* **colour_group_code, colour_group_name** : The group of color and its name\n* **perceived_colour_value_id, perceived_colour_value_name, perceived_colour_master_id, perceived_colour_master_name** : The added color info\n* **department_no, department_name**: A unique identifier of every dep and its name\n* **index_code, index_name**: A unique identifier of every index and its name\n* **index_group_no, index_group_name**: A group of indices and its name\n* **section_no, section_name**: A unique identifier of every section and its name\n* **garment_group_no, garment_group_name**: A unique identifier of every garment and its name\n* **detail_desc**: Details","metadata":{}},{"cell_type":"code","source":"#To see what types of columns the dataset has\ndf.dtypes","metadata":{"execution":{"iopub.status.busy":"2023-08-03T19:02:43.066285Z","iopub.execute_input":"2023-08-03T19:02:43.066858Z","iopub.status.idle":"2023-08-03T19:02:43.079737Z","shell.execute_reply.started":"2023-08-03T19:02:43.066815Z","shell.execute_reply":"2023-08-03T19:02:43.078334Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#To see some statistical information regarding numeric columns of the dataset\ndf.describe()","metadata":{"execution":{"iopub.status.busy":"2023-08-03T19:02:43.081107Z","iopub.execute_input":"2023-08-03T19:02:43.081711Z","iopub.status.idle":"2023-08-03T19:02:43.174841Z","shell.execute_reply.started":"2023-08-03T19:02:43.081650Z","shell.execute_reply":"2023-08-03T19:02:43.173531Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#To see only the numeric columns in the dataset\ndf.select_dtypes(include = 'int64')","metadata":{"execution":{"iopub.status.busy":"2023-08-03T19:02:43.181200Z","iopub.execute_input":"2023-08-03T19:02:43.181925Z","iopub.status.idle":"2023-08-03T19:02:43.205366Z","shell.execute_reply.started":"2023-08-03T19:02:43.181875Z","shell.execute_reply":"2023-08-03T19:02:43.204043Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#To see the columns information\ndf.info()","metadata":{"execution":{"iopub.status.busy":"2023-08-03T19:02:43.207154Z","iopub.execute_input":"2023-08-03T19:02:43.207496Z","iopub.status.idle":"2023-08-03T19:02:43.702356Z","shell.execute_reply.started":"2023-08-03T19:02:43.207468Z","shell.execute_reply":"2023-08-03T19:02:43.701009Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Step 3: Data Preparation/Cleaning","metadata":{}},{"cell_type":"code","source":"##Dropping some columns\n\ndf = df[[#'article_id',\n    'product_code', 'prod_name',\n    #'product_type_no',\n       'product_type_name', 'product_group_name', \n    #'graphical_appearance_no', 'graphical_appearance_name', 'colour_group_code', \n    'colour_group_name',\n       #'perceived_colour_value_id', 'perceived_colour_value_name',\n       #'perceived_colour_master_id', 'perceived_colour_master_name',\n       #'department_no', 'department_name', 'index_code',\n    'index_name',\n       #'index_group_no', \n    'index_group_name']].copy() #\".copy()\" function allows you to have no connection from the original dataset\n    #'section_no', 'section_name', 'garment_group_no', 'garment_group_name', 'detail_desc'","metadata":{"execution":{"iopub.status.busy":"2023-08-03T19:02:43.703904Z","iopub.execute_input":"2023-08-03T19:02:43.704511Z","iopub.status.idle":"2023-08-03T19:02:43.722758Z","shell.execute_reply.started":"2023-08-03T19:02:43.704476Z","shell.execute_reply":"2023-08-03T19:02:43.721491Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"##Reanme column name\ndf = df.rename(columns = {'product_code':'Product_Code',\n                    'prod_name':'Product_Name',\n                    'product_type_name':'Product_Type_Name',\n                    'product_group_name':'Product_Group_Name',\n                    'colour_group_name':'Colour_Group_Name',\n                    'index_name':'Index_Name',\n                    'index_group_name':'Index_Group_Name'})","metadata":{"execution":{"iopub.status.busy":"2023-08-03T19:02:43.724472Z","iopub.execute_input":"2023-08-03T19:02:43.724886Z","iopub.status.idle":"2023-08-03T19:02:43.737304Z","shell.execute_reply.started":"2023-08-03T19:02:43.724853Z","shell.execute_reply":"2023-08-03T19:02:43.736017Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#Examining null or missing values in the dataset\n\ndf.isna()","metadata":{"execution":{"iopub.status.busy":"2023-08-03T19:02:43.739367Z","iopub.execute_input":"2023-08-03T19:02:43.739881Z","iopub.status.idle":"2023-08-03T19:02:43.976500Z","shell.execute_reply.started":"2023-08-03T19:02:43.739839Z","shell.execute_reply":"2023-08-03T19:02:43.975597Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df.isna().sum()","metadata":{"execution":{"iopub.status.busy":"2023-08-03T19:02:43.978320Z","iopub.execute_input":"2023-08-03T19:02:43.979190Z","iopub.status.idle":"2023-08-03T19:02:44.203506Z","shell.execute_reply.started":"2023-08-03T19:02:43.979146Z","shell.execute_reply":"2023-08-03T19:02:44.202273Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#Another way of finding missing data by creating function:\n\ndef missing_data(data):\n    total = data.isnull().sum().sort_values(ascending = False)\n    percent = ((data.isnull().sum()/data.isnull().count())*100).sort_values(ascending = False)\n    return pd.concat([total, percent], axis = 1, keys = ['Total', 'Percent'])\n\nmissing_data(df)","metadata":{"execution":{"iopub.status.busy":"2023-08-03T19:02:44.205151Z","iopub.execute_input":"2023-08-03T19:02:44.205615Z","iopub.status.idle":"2023-08-03T19:02:44.849586Z","shell.execute_reply.started":"2023-08-03T19:02:44.205569Z","shell.execute_reply":"2023-08-03T19:02:44.848355Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"there is no null or missing values","metadata":{}},{"cell_type":"code","source":"df.loc[df.duplicated(subset=['Product_Code'])]","metadata":{"execution":{"iopub.status.busy":"2023-08-03T19:02:44.851035Z","iopub.execute_input":"2023-08-03T19:02:44.851474Z","iopub.status.idle":"2023-08-03T19:02:44.882687Z","shell.execute_reply.started":"2023-08-03T19:02:44.851432Z","shell.execute_reply":"2023-08-03T19:02:44.881369Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#An example of checking duplicate\ndf.query('Product_Code == \"214844\"')","metadata":{"execution":{"iopub.status.busy":"2023-08-03T19:02:44.884442Z","iopub.execute_input":"2023-08-03T19:02:44.885207Z","iopub.status.idle":"2023-08-03T19:02:44.913649Z","shell.execute_reply.started":"2023-08-03T19:02:44.885158Z","shell.execute_reply":"2023-08-03T19:02:44.912252Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"pd.value_counts(df['Index_Group_Name'])","metadata":{"execution":{"iopub.status.busy":"2023-08-03T19:02:44.915527Z","iopub.execute_input":"2023-08-03T19:02:44.916042Z","iopub.status.idle":"2023-08-03T19:02:44.944376Z","shell.execute_reply.started":"2023-08-03T19:02:44.915981Z","shell.execute_reply":"2023-08-03T19:02:44.943236Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Step 4: Data analysing with different charts\n\nTrying to answer this questions:\n* which product type is the most in the product type? #Same question will be tried for other columns also\n* what and how many of the product types are there in different product group? \n* How many articles are there in different product group and product type?\n","metadata":{}},{"cell_type":"code","source":"articles_df = pd.read_csv('/kaggle/input/h-and-m-personalized-fashion-recommendations/articles.csv')","metadata":{"execution":{"iopub.status.busy":"2023-08-03T19:02:44.946012Z","iopub.execute_input":"2023-08-03T19:02:44.947911Z","iopub.status.idle":"2023-08-03T19:02:45.766436Z","shell.execute_reply.started":"2023-08-03T19:02:44.947852Z","shell.execute_reply":"2023-08-03T19:02:45.765142Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#What are the product types and which is the most product type\narticles_df['product_type_name'].value_counts().plot(kind = 'bar', color = 'green', figsize = (28,6))","metadata":{"execution":{"iopub.status.busy":"2023-08-03T19:02:45.768432Z","iopub.execute_input":"2023-08-03T19:02:45.768935Z","iopub.status.idle":"2023-08-03T19:02:47.838281Z","shell.execute_reply.started":"2023-08-03T19:02:45.768892Z","shell.execute_reply":"2023-08-03T19:02:47.836909Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"articles_df['product_group_name'].value_counts().plot(kind = 'bar', color = 'green', width = 0.5, figsize = (8,6))","metadata":{"execution":{"iopub.status.busy":"2023-08-03T19:02:47.840246Z","iopub.execute_input":"2023-08-03T19:02:47.840702Z","iopub.status.idle":"2023-08-03T19:02:48.337461Z","shell.execute_reply.started":"2023-08-03T19:02:47.840661Z","shell.execute_reply":"2023-08-03T19:02:48.335629Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"\"Garment upper body\" is the most product group name ","metadata":{}},{"cell_type":"code","source":"articles_df['index_group_name'].value_counts().plot(kind = 'bar', color = 'green', width = 0.5, figsize = (8,6))","metadata":{"execution":{"iopub.status.busy":"2023-08-03T19:02:48.339134Z","iopub.execute_input":"2023-08-03T19:02:48.340326Z","iopub.status.idle":"2023-08-03T19:02:48.686048Z","shell.execute_reply.started":"2023-08-03T19:02:48.340274Z","shell.execute_reply":"2023-08-03T19:02:48.684630Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Most of the products are ladieswear","metadata":{}},{"cell_type":"code","source":"##what and how many of the product types are there in different product group? \n\ntemp = articles_df.groupby([\"product_group_name\"])[\"product_type_name\"].nunique()\ndf = pd.DataFrame({'Product Group': temp.index,\n                   'Product Types': temp.values\n                  })\ndf = df.sort_values(['Product Types'], ascending = False)\nplt.figure(figsize = (8,6))\nplt.title('Number of Product Types per each Product Group')\nsns.set_color_codes(\"muted\")\ns = sns.barplot(x = 'Product Group', y=\"Product Types\", data=df)\ns.set_xticklabels(s.get_xticklabels(),rotation=90)\nlocs, labels = plt.xticks()\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-08-03T19:02:48.687754Z","iopub.execute_input":"2023-08-03T19:02:48.688868Z","iopub.status.idle":"2023-08-03T19:02:49.225932Z","shell.execute_reply.started":"2023-08-03T19:02:48.688825Z","shell.execute_reply":"2023-08-03T19:02:49.224698Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"##How many articles are there in different product group\n\ntemp = articles_df.groupby(['product_group_name'])['article_id'].nunique()\ndf = pd.DataFrame({'Product Group': temp.index,\n                   'Articles': temp.values\n                  })\ndf = df.sort_values(['Articles'], ascending = False)\nplt.figure(figsize = (8,6))\nplt.title('Number of articles in different product groups')\nsns.set_color_codes('pastel')\ns = sns.barplot(x = 'Product Group', y = 'Articles', data = df)\ns.set_xticklabels(s.get_xticklabels(), rotation = 90)\nlocs, labels = plt.xticks()\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-08-03T19:02:49.233478Z","iopub.execute_input":"2023-08-03T19:02:49.233925Z","iopub.status.idle":"2023-08-03T19:02:49.745086Z","shell.execute_reply.started":"2023-08-03T19:02:49.233883Z","shell.execute_reply":"2023-08-03T19:02:49.743800Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"temp = articles_df.groupby(['product_type_name'])['article_id'].nunique()\ndf = pd.DataFrame({'Product Type': temp.index, 'Articles': temp.values})\ndf = df.sort_values(['Articles'], ascending = False)\nplt.figure(figsize = (25,6))\nplt.title(\"Number of Articles in different product type\")\nsns.set_color_codes(\"pastel\")\ns = sns.barplot(x = 'Product Type', y = 'Articles', data = df)\ns.set_xticklabels(s.get_xticklabels(), rotation = 90)\nlocs, labels = plt.xticks()\nplt.show()\n","metadata":{"execution":{"iopub.status.busy":"2023-08-03T19:02:49.746811Z","iopub.execute_input":"2023-08-03T19:02:49.747220Z","iopub.status.idle":"2023-08-03T19:02:51.481743Z","shell.execute_reply.started":"2023-08-03T19:02:49.747185Z","shell.execute_reply":"2023-08-03T19:02:51.480467Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"temp = articles_df.groupby(['department_name'])['article_id'].nunique()\ndf = pd.DataFrame({\"Department\": temp.index, \"Articles\": temp.values})\ndf = df.sort_values(['Articles'], ascending = False)\nplt.figure(figsize = (25,6))\nplt.title(\"Number of articles in different department\")\nsns.set_color_codes(\"pastel\")\ns = sns.barplot(x = \"Department\", y = \"Articles\", data =df)\ns.set_xticklabels(s.get_xticklabels(), rotation = 90)\nlocs, labels = plt.xticks()\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-08-03T19:02:51.483249Z","iopub.execute_input":"2023-08-03T19:02:51.483635Z","iopub.status.idle":"2023-08-03T19:02:54.757878Z","shell.execute_reply.started":"2023-08-03T19:02:51.483604Z","shell.execute_reply":"2023-08-03T19:02:54.756866Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"articles_df[\"department_name\"].value_counts()","metadata":{"execution":{"iopub.status.busy":"2023-08-03T19:02:54.759455Z","iopub.execute_input":"2023-08-03T19:02:54.760142Z","iopub.status.idle":"2023-08-03T19:02:54.787193Z","shell.execute_reply.started":"2023-08-03T19:02:54.760091Z","shell.execute_reply":"2023-08-03T19:02:54.785921Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"len(articles_df[\"department_name\"].unique())","metadata":{"execution":{"iopub.status.busy":"2023-08-03T19:02:54.788836Z","iopub.execute_input":"2023-08-03T19:02:54.789314Z","iopub.status.idle":"2023-08-03T19:02:54.812014Z","shell.execute_reply.started":"2023-08-03T19:02:54.789273Z","shell.execute_reply":"2023-08-03T19:02:54.811072Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#Take top 50 department\n\ntemp = articles_df.groupby(['department_name'])['article_id'].nunique()\ndf = pd.DataFrame({\"Department\": temp.index, \"Articles\": temp.values})\ndf = df.sort_values(['Articles'], ascending = False).head(50)\nplt.figure(figsize = (15,6))\nplt.title(\"Number of articles in different department (top 50 from total 250)\")\nsns.set_color_codes(\"pastel\")\ns = sns.barplot(x = \"Department\", y = \"Articles\", data =df)\ns.set_xticklabels(s.get_xticklabels(), rotation = 90)\nlocs, labels = plt.xticks()\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-08-03T19:02:54.813682Z","iopub.execute_input":"2023-08-03T19:02:54.814022Z","iopub.status.idle":"2023-08-03T19:02:55.653526Z","shell.execute_reply.started":"2023-08-03T19:02:54.813993Z","shell.execute_reply":"2023-08-03T19:02:55.652343Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"articles_df.groupby([\"department_name\"])[\"product_type_name\"].nunique().sort_values(ascending = False)","metadata":{"execution":{"iopub.status.busy":"2023-08-03T19:02:55.654887Z","iopub.execute_input":"2023-08-03T19:02:55.655352Z","iopub.status.idle":"2023-08-03T19:02:55.701892Z","shell.execute_reply.started":"2023-08-03T19:02:55.655306Z","shell.execute_reply":"2023-08-03T19:02:55.700499Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#Lets see the chart on department name and product type\n\ntemp = articles_df.groupby([\"department_name\"])[\"product_type_name\"].nunique()\ndf = pd.DataFrame({\"Department\": temp.index, \"Product Type\": temp.values})\ndf = df.sort_values([\"Product Type\"], ascending = False).head(50)\nplt.figure(figsize = (15,6))\nplt.title(\"Number of product type in different product department (Top 50)\")\nsns.set_color_codes(\"pastel\")\ns = sns.barplot(x = \"Department\", y = \"Product Type\", data = df)\ns.set_xticklabels(s.get_xticklabels(), rotation = 90)\nlocs, labels = plt.xticks()\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-08-03T19:02:55.703358Z","iopub.execute_input":"2023-08-03T19:02:55.703816Z","iopub.status.idle":"2023-08-03T19:02:56.596260Z","shell.execute_reply.started":"2023-08-03T19:02:55.703773Z","shell.execute_reply":"2023-08-03T19:02:56.595295Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#Lets work on graphical appearance\n\ntemp = articles_df.groupby([\"graphical_appearance_name\"])[\"article_id\"].nunique()\ndf = pd.DataFrame({\"Graphical appearance\": temp.index, \"Articles\": temp.values})\ndf = df.sort_values([\"Articles\"], ascending = False)\nplt.figure(figsize = (15,6))\nplt.title(\"Number of articles based on different graphical appearance\")\nsns.set_color_codes(\"pastel\")\ns = sns.barplot(x = \"Graphical appearance\", y = \"Articles\", data = df)\ns.set_xticklabels(s.get_xticklabels(), rotation = 90)\nlocs, labels = plt.xticks()\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-08-03T19:02:56.597471Z","iopub.execute_input":"2023-08-03T19:02:56.597820Z","iopub.status.idle":"2023-08-03T19:02:57.209106Z","shell.execute_reply.started":"2023-08-03T19:02:56.597791Z","shell.execute_reply":"2023-08-03T19:02:57.207977Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"articles_df.columns","metadata":{"execution":{"iopub.status.busy":"2023-08-03T19:02:57.210616Z","iopub.execute_input":"2023-08-03T19:02:57.210980Z","iopub.status.idle":"2023-08-03T19:02:57.219638Z","shell.execute_reply.started":"2023-08-03T19:02:57.210950Z","shell.execute_reply":"2023-08-03T19:02:57.218269Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"articles_df.groupby([\"index_group_name\"])[\"article_id\"].nunique().sort_values(ascending = False)","metadata":{"execution":{"iopub.status.busy":"2023-08-03T19:02:57.221404Z","iopub.execute_input":"2023-08-03T19:02:57.221969Z","iopub.status.idle":"2023-08-03T19:02:57.257844Z","shell.execute_reply.started":"2023-08-03T19:02:57.221927Z","shell.execute_reply":"2023-08-03T19:02:57.256611Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"temp = articles_df.groupby([\"index_group_name\"])[\"article_id\"].nunique()\ndf = pd.DataFrame({\"Index name\": temp.index, \"Articles\": temp.values})\ndf = df.sort_values([\"Articles\"], ascending = False)\nplt.figure(figsize=(6,6))\nplt.title(\"Number of products in each index group\")\nsns.set_color_codes(\"pastel\")\ns = sns.barplot(x = \"Index name\", y = \"Articles\", width = 0.5, data = df)\ns.set_xticklabels(s.get_xticklabels(), rotation = 90)\nlocs, labels = plt.xticks()\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-08-03T19:02:57.260531Z","iopub.execute_input":"2023-08-03T19:02:57.261256Z","iopub.status.idle":"2023-08-03T19:02:57.605825Z","shell.execute_reply.started":"2023-08-03T19:02:57.261212Z","shell.execute_reply":"2023-08-03T19:02:57.604507Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"articles_df.garment_group_name","metadata":{"execution":{"iopub.status.busy":"2023-08-03T19:02:57.607174Z","iopub.execute_input":"2023-08-03T19:02:57.607534Z","iopub.status.idle":"2023-08-03T19:02:57.616392Z","shell.execute_reply.started":"2023-08-03T19:02:57.607501Z","shell.execute_reply":"2023-08-03T19:02:57.615309Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"articles_df.groupby([\"garment_group_name\"])[\"article_id\"].nunique().sort_values(ascending = False)","metadata":{"execution":{"iopub.status.busy":"2023-08-03T19:02:57.618100Z","iopub.execute_input":"2023-08-03T19:02:57.618413Z","iopub.status.idle":"2023-08-03T19:02:57.654803Z","shell.execute_reply.started":"2023-08-03T19:02:57.618386Z","shell.execute_reply":"2023-08-03T19:02:57.653741Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"temp = articles_df.groupby([\"garment_group_name\"])[\"article_id\"].nunique()\ndf = pd.DataFrame({\"Garment group\": temp.index, \"Articles\": temp.values})\ndf = df.sort_values([\"Articles\"], ascending = False)\nplt.figure(figsize = (8,6))\nplt.title(\"Number of products in different garments group\")\nsns.set_color_codes(\"pastel\")\ns = sns.barplot(x = \"Garment group\", y = \"Articles\", data = df)\ns.set_xticklabels(s.get_xticklabels(), rotation = 90)\nlocs, labels = plt.xticks()\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-08-03T19:02:57.656704Z","iopub.execute_input":"2023-08-03T19:02:57.657131Z","iopub.status.idle":"2023-08-03T19:02:58.119045Z","shell.execute_reply.started":"2023-08-03T19:02:57.657091Z","shell.execute_reply":"2023-08-03T19:02:58.117757Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Working with Customer Data","metadata":{}},{"cell_type":"code","source":"customers_df = pd.read_csv(\"/kaggle/input/h-and-m-personalized-fashion-recommendations/customers.csv\")","metadata":{"execution":{"iopub.status.busy":"2023-08-03T19:02:58.120483Z","iopub.execute_input":"2023-08-03T19:02:58.120885Z","iopub.status.idle":"2023-08-03T19:03:02.698310Z","shell.execute_reply.started":"2023-08-03T19:02:58.120853Z","shell.execute_reply":"2023-08-03T19:03:02.697369Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"customers_df.head(100)","metadata":{"execution":{"iopub.status.busy":"2023-08-03T19:03:02.699636Z","iopub.execute_input":"2023-08-03T19:03:02.700178Z","iopub.status.idle":"2023-08-03T19:03:02.723538Z","shell.execute_reply.started":"2023-08-03T19:03:02.700147Z","shell.execute_reply":"2023-08-03T19:03:02.722243Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"postal code is not clear. looks like same as customer id\n\n* **customer_id** : A unique identifier of every customer\n* **FN** : 1 or missed (not clear what it is)\n* **Active** : 1 or missed (need to check if it is linked to making purchases)\n* **club_member_status** : Status in club\n* **fashion_news_frequency** : How often H&M may send news to customer\n* **age** : The current age\n* **postal_code** : Postal code of customer","metadata":{}},{"cell_type":"code","source":"missing_data(customers_df)","metadata":{"execution":{"iopub.status.busy":"2023-08-03T19:03:02.725262Z","iopub.execute_input":"2023-08-03T19:03:02.725868Z","iopub.status.idle":"2023-08-03T19:03:08.158784Z","shell.execute_reply.started":"2023-08-03T19:03:02.725826Z","shell.execute_reply":"2023-08-03T19:03:08.157440Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def unique_values(data):\n    total = data.count()\n    tt = pd.DataFrame(total)\n    tt.columns = [\"Total\"]\n    uniques = []\n    for col in data.columns:\n        unique = data[col].nunique()\n        uniques.append(unique)\n    tt['Uniques'] = uniques\n    return tt","metadata":{"execution":{"iopub.status.busy":"2023-08-03T19:03:08.161533Z","iopub.execute_input":"2023-08-03T19:03:08.162578Z","iopub.status.idle":"2023-08-03T19:03:08.169290Z","shell.execute_reply.started":"2023-08-03T19:03:08.162510Z","shell.execute_reply":"2023-08-03T19:03:08.167849Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"unique_values(customers_df)","metadata":{"execution":{"iopub.status.busy":"2023-08-03T19:03:08.170955Z","iopub.execute_input":"2023-08-03T19:03:08.171330Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"customers_df.club_member_status","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"customers_df.groupby([\"club_member_status\"]).nunique()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"customers_df.groupby([\"club_member_status\"])[\"customer_id\"].nunique()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"customers_df.fashion_news_frequency","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"customers_df.groupby([\"fashion_news_frequency\"]).nunique()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"customers_df['fashion_news_frequency'].replace({\"None\": \"NONE\"}, inplace = True)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"customers_df.groupby([\"fashion_news_frequency\"]).nunique()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#Let's look into club member status in perspective with customers\n\ntemp = customers_df.groupby([\"club_member_status\"])[\"customer_id\"].nunique()\ndf = pd.DataFrame({\"Club membership status\": temp.index, \"Customers\": temp.values})\ndf = df.sort_values([\"Customers\"], ascending = False)\nplt.figure(figsize = (6,6))\nplt.title(\"Number of customers per different club membership status\")\nsns.set_color_codes(\"pastel\")\ns = sns.barplot(x = \"Club membership status\", y = \"Customers\", width = 0.5, data = df)\ns.set_xticklabels(s.get_xticklabels())\nlocs, labels = plt.xticks()\nplt.show()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Looks like most of the members are active. lets look what type of age group are active or left.","metadata":{}},{"cell_type":"code","source":"customers_df.groupby([\"club_member_status\"])[\"age\"].nunique()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"customers_df.club_member_status.nunique()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"temp = customers_df.groupby([\"club_member_status\"])[\"age\"].nunique()\ndf = pd.DataFrame({\"Club membership status\": temp.index, \"Age\": temp.values})\ndf = df.sort_values([\"Age\"], ascending = False)\nplt.figure(figsize = (6,6))\nplt.title(\"Club Membership status in different age group\")\nsns.set_color_codes(\"pastel\")\ns = sns.barplot(x = \"Club membership status\", y = \"Age\", width = 0.5, data = df)\ns.set_xticklabels(s.get_xticklabels())\nlocs, labels = plt.xticks()\nplt.show()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"It is actually counting the number of membership status in according with the age.\n\n* **Want to see which age of people has what type of membership status. need to figure out what sort of chart can visualise this**\n* **Also want to figure out which age of people has what types of fashion news frequency. Do H&M send the fashion news to the older people more frequently than the younger one or not?**","metadata":{}},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]}]}