{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"## H&M: Personalized Fashion Recommendations EDA\n\n<div>    \n<img src=\"https://static.instyle.de/1920x1080/focal_1440x810:1441x811/images/2017-12/hmnyden.jpg\" width=\"600\", align=\"center\"/>    \n</div>\n\n\n#### About H&M: \n\nH&M Group is a family of brands and businesses with 53 online markets and approximately 4,850 stores. Their online store offers shoppers an extensive selection of products to browse through.\n\n#### About the comp\n\nGiven the purchase history of customers across time, along with supporting metadata, our challenge is to predict what articles each customer will purchase in the 7-day period immediately after the training data ends. Customer who did not make any purchase during that time are excluded from the scoring.\n\n#### Evaluation metrics\n\nSubmissions are evaluated according to the Mean Average Precision @ 12 (MAP@12):\n\n$$MAP@12 = \\frac{1}{U}\\sum_{u=1}^{U} \\sum_{k=1}^{min(n, 12)} P(k) \\times rel(k)$$\n\nwhere $U$ is the number of customers, $P(x)$ is the precision at cutoff $k$, $n$ is the number predictions per customer, and $rel(k)$ is an indicator function equaling 1 if the item at rank $k$ is a relevant (correct) label, zero otherwise.\n\n##### Notes:\n\n- You will be making purchase predictions for all customer_id values provided, regardless of whether these customers made purchases in the training data.\n- Customer that did not make any purchase during test period are excluded from the scoring.\n- There is never a penalty for using the full 12 predictions for a customer that ordered fewer than 12 items; thus, it's advantageous to make 12 predictions for each customer.\n\n---\n\n#### **Table of Contents**\n\n[1. Data Overview](#1)\n\n[2. Articles](#2)\n\n[3. Customers](#3)\n\n[4. Sample Product Pictures](#4)\n\n[5. Transactions](#5)\n\n[6. Reference & credits](#6)\n \n ---","metadata":{}},{"cell_type":"code","source":"import os\n\nimport numpy as np\nimport datatable as dt\n\nimport pandas as pd\nfrom scipy import stats\nimport seaborn as sns\nimport matplotlib.pyplot as plt\nimport plotly.io as pio\nimport plotly.express as px\nimport plotly.figure_factory as ff\nimport plotly.graph_objects as go\n\nfrom plotly.subplots import make_subplots\nfrom plotly.offline import init_notebook_mode, iplot\n\ninit_notebook_mode(connected=True)\npio.templates.default = \"none\"\nimport cv2\n\n\nimport warnings\nwarnings.filterwarnings('ignore')\n\npath = '/kaggle/input/h-and-m-personalized-fashion-recommendations'","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","_kg_hide-output":true,"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-02-08T12:29:04.665000Z","iopub.execute_input":"2022-02-08T12:29:04.665355Z","iopub.status.idle":"2022-02-08T12:29:09.043170Z","shell.execute_reply.started":"2022-02-08T12:29:04.665321Z","shell.execute_reply":"2022-02-08T12:29:09.042093Z"},"jupyter":{"source_hidden":true},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sample_sumbission = dt.fread(path+'/sample_submission.csv').to_pandas()\nart = dt.fread(path+'/articles.csv').to_pandas()\ntrans_train = dt.fread(path+'/transactions_train.csv').to_pandas()\ncust = dt.fread(path+ '/customers.csv').to_pandas()","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-02-08T12:29:09.045090Z","iopub.execute_input":"2022-02-08T12:29:09.045328Z","iopub.status.idle":"2022-02-08T12:30:06.669065Z","shell.execute_reply.started":"2022-02-08T12:29:09.045302Z","shell.execute_reply":"2022-02-08T12:30:06.668229Z"},"jupyter":{"source_hidden":true},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### 1. Data Overview <a class=\"anchor\" id=\"1\"></a>\n\n##### Articles\n\n- Articles data has 105542 rows and 25 cols\n- No null values in article dataset\n- `article_id` has the largest unique values with 105542, and `index_group` the lowest with 5 unique values\n- Top three product name: `trousers, dress, sweaters`\n- Top three product group manes: garment upper body, garment lower body, garment full body\n- Top three article index name: ladies-wear, divided, menswear\n- Top three article index group name: ladies-wear, baby/children, divided\n- Top three graphical appearance name: solid, all-over patter, melange\n- Top three garment group name: jersey fancy, accessories, jersey basic\n- Top three perceived colour value name: dark, dusty light, light\n- Top three section name: women's everyday collection, divided collection, baby essentials & complements \n\n##### Customers\n- Customers data has 1371980 rows and 7 cols\n- Columns `FN`  and `Active` have the most NA values with 65 and 66% respectively. `Age` also has around 1% NA values.\n- Age distribution of the customers of both (active members and non-members) is similar with two peaks around early-mid 20s and 50.\n- Majority (92%) of the customers are club member\n- Customers who follow fashion news regularity are almost entirely club members","metadata":{}},{"cell_type":"markdown","source":"### 2. Articles <a class=\"anchor\" id=\"2\"></a>","metadata":{}},{"cell_type":"code","source":"display(art.shape)\ndisplay(art.head(5))","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-02-08T12:30:06.670404Z","iopub.execute_input":"2022-02-08T12:30:06.670691Z","iopub.status.idle":"2022-02-08T12:30:06.708047Z","shell.execute_reply.started":"2022-02-08T12:30:06.670657Z","shell.execute_reply":"2022-02-08T12:30:06.707079Z"},"jupyter":{"source_hidden":true},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"art.info()","metadata":{"_kg_hide-input":true,"jupyter":{"source_hidden":true},"execution":{"iopub.status.busy":"2022-02-08T12:30:06.709348Z","iopub.execute_input":"2022-02-08T12:30:06.709665Z","iopub.status.idle":"2022-02-08T12:30:06.813902Z","shell.execute_reply.started":"2022-02-08T12:30:06.709598Z","shell.execute_reply":"2022-02-08T12:30:06.812887Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### Uniqness of columns in articles dataframe","metadata":{}},{"cell_type":"code","source":"from termcolor import colored\nfor col in art.columns:\n    x = art[col].nunique()    \n    print(\"{}: {} unique values\".format(col, colored(x, 'white')))","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-02-08T12:30:06.816423Z","iopub.execute_input":"2022-02-08T12:30:06.816712Z","iopub.status.idle":"2022-02-08T12:30:07.004209Z","shell.execute_reply.started":"2022-02-08T12:30:06.816676Z","shell.execute_reply":"2022-02-08T12:30:07.002512Z"},"jupyter":{"source_hidden":true},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### 2.1 Product Type Name","metadata":{}},{"cell_type":"code","source":"fig = px.histogram(art, x=\"product_type_name\",\n                   width=900, \n                   height=400,\n                   histnorm='percent',\n                   template=\"simple_white\"\n                   )\n\nfig.update_layout(title=\"Product Type Name \", \n                  font_family=\"San Serif\",\n                  titlefont={'size': 20},\n                  legend=dict(\n                  orientation=\"v\", y=1, yanchor=\"top\", x=1.0, xanchor=\"right\" )                 \n                 ).update_xaxes(categoryorder='total descending')# ordering the x-axis values\n\ncolors = ['lightgray'] * 100  \ncolors[0] = 'crimson' \ncolors[1] = 'crimson' \ncolors[2] = 'crimson' \n\n\nfig.update_traces(marker_color=colors, \n                )\nfig.show()","metadata":{"_kg_hide-input":true,"jupyter":{"source_hidden":true},"execution":{"iopub.status.busy":"2022-02-08T12:30:07.005531Z","iopub.execute_input":"2022-02-08T12:30:07.006409Z","iopub.status.idle":"2022-02-08T12:30:08.979740Z","shell.execute_reply.started":"2022-02-08T12:30:07.006367Z","shell.execute_reply":"2022-02-08T12:30:08.978523Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### 2.2 Product Group Name","metadata":{}},{"cell_type":"code","source":"fig = px.histogram(art, x=\"product_group_name\",\n                   width=900, \n                   height=400,\n                   histnorm='percent',\n                   template=\"simple_white\"\n                   )\n\nfig.update_layout(title=\"Product Group Name \", \n                  font_family=\"San Serif\",\n                  titlefont={'size': 20},\n                  legend=dict(\n                  orientation=\"v\", y=1, yanchor=\"top\", x=1.0, xanchor=\"right\" )                 \n                 ).update_xaxes(categoryorder='total descending')# ordering the x-axis values\n\ncolors = ['lightgray'] * 100  \ncolors[0] = 'crimson' \ncolors[1] = 'crimson' \ncolors[2] = 'crimson' \n\n\nfig.update_traces(marker_color=colors, \n                )\nfig.show()","metadata":{"_kg_hide-input":true,"jupyter":{"source_hidden":true},"execution":{"iopub.status.busy":"2022-02-08T12:30:08.981045Z","iopub.execute_input":"2022-02-08T12:30:08.981274Z","iopub.status.idle":"2022-02-08T12:30:09.963452Z","shell.execute_reply.started":"2022-02-08T12:30:08.981247Z","shell.execute_reply":"2022-02-08T12:30:09.958772Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### 2.3 Product Index Name/ Index Group Name","metadata":{}},{"cell_type":"code","source":"fig = px.histogram(art, x=\"index_name\",\n                   width=600, \n                   height=400,\n                   histnorm='percent',\n                   template=\"simple_white\"\n                   )\n\nfig.update_layout(title=\"Index Name \", \n                  font_family=\"San Serif\",\n                  titlefont={'size': 20},\n                  legend=dict(\n                  orientation=\"v\", y=1, yanchor=\"top\", x=1.0, xanchor=\"right\" )                 \n                 ).update_xaxes(categoryorder='total descending') # ordering the x-axis values\n\ncolors = ['lightgray'] * 100  \ncolors[0] = 'crimson' \ncolors[1] = 'crimson' \ncolors[2] = 'crimson' \n\n\nfig.update_traces(marker_color=colors, \n                )\n\nfig.show()\n\n\nfig = px.histogram(art, x=\"index_group_name\",\n                   width=600, \n                   height=400,\n                   histnorm='percent',\n                   template=\"simple_white\"\n                   )\n\nfig.update_layout(title=\"Index Group Name \", \n                  font_family=\"San Serif\",\n                  titlefont={'size': 20},\n                  legend=dict(\n                  orientation=\"v\", y=1, yanchor=\"top\", x=1.0, xanchor=\"right\" )                 \n                 ).update_xaxes(categoryorder='total descending') # ordering the x-axis values\n\ncolors = ['lightgray'] * 100  \ncolors[0] = 'crimson' \ncolors[1] = 'crimson' \ncolors[2] = 'crimson' \n\n\nfig.update_traces(marker_color=colors, \n                )\n\nfig.show()","metadata":{"_kg_hide-input":true,"jupyter":{"source_hidden":true},"execution":{"iopub.status.busy":"2022-02-08T12:30:09.964823Z","iopub.execute_input":"2022-02-08T12:30:09.965682Z","iopub.status.idle":"2022-02-08T12:30:11.778329Z","shell.execute_reply.started":"2022-02-08T12:30:09.965641Z","shell.execute_reply":"2022-02-08T12:30:11.776907Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### 2.4 Graphical Appearance Name","metadata":{}},{"cell_type":"code","source":"fig = px.histogram(art, x=\"graphical_appearance_name\",\n                   width=700, \n                   height=400,\n                   histnorm='percent',\n                   template=\"simple_white\"\n                   )\n\nfig.update_layout(title=\"Graphical Appearance Name\", \n                  font_family=\"San Serif\",\n                  titlefont={'size': 20},\n                  legend=dict(\n                  orientation=\"v\", y=1, yanchor=\"top\", x=1.0, xanchor=\"right\" )                 \n                 ).update_xaxes(categoryorder='total descending')\n\ncolors = ['lightgray'] * 100  \ncolors[0] = 'crimson' \ncolors[1] = 'crimson' \ncolors[2] = 'crimson' \nfig.update_traces(marker_color=colors, \n                )\n\nfig.show()","metadata":{"_kg_hide-input":true,"jupyter":{"source_hidden":true},"execution":{"iopub.status.busy":"2022-02-08T12:30:11.779805Z","iopub.execute_input":"2022-02-08T12:30:11.780589Z","iopub.status.idle":"2022-02-08T12:30:12.694317Z","shell.execute_reply.started":"2022-02-08T12:30:11.780541Z","shell.execute_reply":"2022-02-08T12:30:12.693040Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### 2.5 Garment Group Name","metadata":{}},{"cell_type":"code","source":"fig = px.histogram(art, x=\"garment_group_name\",\n                   width=700, \n                   height=400,\n                   histnorm='percent',\n                   template=\"simple_white\"\n                   )\n\nfig.update_layout(title=\"Garment Group Name\", \n                  font_family=\"San Serif\",\n                  titlefont={'size': 20},\n                  legend=dict(\n                  orientation=\"v\", y=1, yanchor=\"top\", x=1.0, xanchor=\"right\" )                 \n                 ).update_xaxes(categoryorder='total descending')\n\ncolors = ['lightgray'] * 100  \ncolors[0] = 'crimson' \ncolors[1] = 'crimson' \ncolors[2] = 'crimson' \n\n\nfig.update_traces(marker_color=colors, \n                )\n\nfig.show()","metadata":{"_kg_hide-input":true,"jupyter":{"source_hidden":true},"execution":{"iopub.status.busy":"2022-02-08T12:30:12.695750Z","iopub.execute_input":"2022-02-08T12:30:12.696047Z","iopub.status.idle":"2022-02-08T12:30:13.751281Z","shell.execute_reply.started":"2022-02-08T12:30:12.696015Z","shell.execute_reply":"2022-02-08T12:30:13.750310Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### 2.6 Perceived Ccolour Vvalue Name","metadata":{}},{"cell_type":"code","source":"fig = px.histogram(art, x=\"perceived_colour_value_name\",\n                   width=700, \n                   height=400,\n                   histnorm='percent',\n                   template=\"simple_white\"\n                   )\n\nfig.update_layout(title=\"Perceived Colour Value Name\", \n                  font_family=\"San Serif\",\n                  titlefont={'size': 20},\n                  legend=dict(\n                  orientation=\"v\", y=1, yanchor=\"top\", x=1.0, xanchor=\"right\" )                 \n                 ).update_xaxes(categoryorder='total descending')\n\ncolors = ['lightgray'] * 100  \ncolors[0] = 'crimson' \ncolors[1] = 'crimson' \ncolors[2] = 'crimson' \n\n\nfig.update_traces(marker_color=colors, \n                )\n\nfig.show()","metadata":{"_kg_hide-input":true,"jupyter":{"source_hidden":true},"execution":{"iopub.status.busy":"2022-02-08T12:30:13.752558Z","iopub.execute_input":"2022-02-08T12:30:13.752830Z","iopub.status.idle":"2022-02-08T12:30:14.474679Z","shell.execute_reply.started":"2022-02-08T12:30:13.752794Z","shell.execute_reply":"2022-02-08T12:30:14.473421Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### 1.7 Section Name","metadata":{}},{"cell_type":"code","source":"fig = px.histogram(art, x=\"section_name\",\n                   width=700, \n                   height=400,\n                   histnorm='percent',\n                   template=\"simple_white\"\n                   )\n\nfig.update_layout(title=\"Section Name\", \n                  font_family=\"San Serif\",\n                  titlefont={'size': 20},\n                  legend=dict(\n                  orientation=\"v\", y=1, yanchor=\"top\", x=1.0, xanchor=\"right\" )                 \n                 ).update_xaxes(categoryorder='total descending')\n\ncolors = ['lightgray'] * 100  \ncolors[0] = 'crimson' \ncolors[1] = 'crimson' \ncolors[2] = 'crimson' \n\n\nfig.update_traces(marker_color=colors, \n                )\n\nfig.show()","metadata":{"_kg_hide-input":true,"jupyter":{"source_hidden":true},"execution":{"iopub.status.busy":"2022-02-08T12:30:14.476245Z","iopub.execute_input":"2022-02-08T12:30:14.476653Z","iopub.status.idle":"2022-02-08T12:30:15.390446Z","shell.execute_reply.started":"2022-02-08T12:30:14.476591Z","shell.execute_reply":"2022-02-08T12:30:15.389296Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### 3. Customers <a class=\"anchor\" id=\"3\"></a>","metadata":{}},{"cell_type":"code","source":"display(cust.shape)\ndisplay(cust.info())","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-02-08T12:30:15.391764Z","iopub.execute_input":"2022-02-08T12:30:15.392034Z","iopub.status.idle":"2022-02-08T12:30:16.441553Z","shell.execute_reply.started":"2022-02-08T12:30:15.391995Z","shell.execute_reply":"2022-02-08T12:30:16.436413Z"},"jupyter":{"source_hidden":true},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"for col in cust.columns:\n    x = cust[col].nunique()    \n    print(\"{}: ======> {} unique values\".format(col, colored(x, 'blue')))","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-02-08T12:30:16.479664Z","iopub.execute_input":"2022-02-08T12:30:16.481682Z","iopub.status.idle":"2022-02-08T12:30:21.224348Z","shell.execute_reply.started":"2022-02-08T12:30:16.481599Z","shell.execute_reply":"2022-02-08T12:30:21.220915Z"},"jupyter":{"source_hidden":true},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"cust.head()","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-02-08T12:30:21.229057Z","iopub.execute_input":"2022-02-08T12:30:21.230253Z","iopub.status.idle":"2022-02-08T12:30:21.292936Z","shell.execute_reply.started":"2022-02-08T12:30:21.230060Z","shell.execute_reply":"2022-02-08T12:30:21.289824Z"},"jupyter":{"source_hidden":true},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# fill missing values\ncust['FN'] = cust['FN'].fillna(0) # dtype float is missing\ncust['Active'] = cust['Active'].fillna(0) # dtype float is missing \ncust['club_member_status'] = cust['club_member_status'].fillna('na') # dtype object is missing, na will do for now ","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-02-08T12:30:21.298824Z","iopub.execute_input":"2022-02-08T12:30:21.300401Z","iopub.status.idle":"2022-02-08T12:30:21.717723Z","shell.execute_reply.started":"2022-02-08T12:30:21.300351Z","shell.execute_reply":"2022-02-08T12:30:21.714113Z"},"jupyter":{"source_hidden":true},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### 3.1 Age of Customers","metadata":{}},{"cell_type":"code","source":"fig = px.histogram(cust, x=\"age\",\n                   width=700, \n                   height=400,\n                   histnorm='percent',\n                   template=\"simple_white\",\n                   color='Active',\n                   color_discrete_sequence =['gray', 'crimson']\n                   )\n\nfig.update_layout(title=\"Age of customers\", \n                  font_family=\"San Serif\",\n                  titlefont={'size': 20},\n                  legend=dict(\n                  orientation=\"v\", y=1, yanchor=\"top\", x=1.0, xanchor=\"right\" )                 \n                 ).update_yaxes(categoryorder='total ascending') \n\nfig.show()","metadata":{"_kg_hide-input":true,"jupyter":{"source_hidden":true},"execution":{"iopub.status.busy":"2022-02-08T12:30:21.723776Z","iopub.execute_input":"2022-02-08T12:30:21.724899Z","iopub.status.idle":"2022-02-08T12:30:31.940171Z","shell.execute_reply.started":"2022-02-08T12:30:21.724784Z","shell.execute_reply":"2022-02-08T12:30:31.939140Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### 3.2 Club Member Status","metadata":{}},{"cell_type":"code","source":"fig = px.histogram(cust, x=\"club_member_status\",\n                   width=600, \n                   height=350,\n                   histnorm='percent',\n                   template=\"simple_white\"\n                   )\n\nfig.update_layout(title=\"Club Member Status\", \n                  font_family=\"San Serif\",\n                  titlefont={'size': 20},\n                  legend=dict(\n                  orientation=\"v\", y=1, yanchor=\"top\", x=1.0, xanchor=\"right\" )                 \n                 ).update_xaxes(categoryorder='total descending')\n\ncolors = ['lightgray'] * 10  \ncolors[0] = 'crimson'\n\nfig.update_traces(marker_color=colors, \n                )\n\nfig.show()","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-02-08T12:30:31.941309Z","iopub.execute_input":"2022-02-08T12:30:31.942008Z","iopub.status.idle":"2022-02-08T12:30:42.047447Z","shell.execute_reply.started":"2022-02-08T12:30:31.941964Z","shell.execute_reply":"2022-02-08T12:30:42.045071Z"},"jupyter":{"source_hidden":true},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### 3.3 Fashion News Frequency","metadata":{}},{"cell_type":"code","source":"fig = px.histogram(cust, x=\"fashion_news_frequency\",\n                   width=700, \n                   height=350,\n                   histnorm='percent',\n                   template=\"simple_white\",\n                   color='Active',\n                   barmode='group',\n                   color_discrete_sequence =['gray', 'crimson']\n                   )\n\nfig.update_layout(title=\"Fashion News Frequency\", \n                  font_family=\"San Serif\",\n                  titlefont={'size': 20},\n                  legend=dict(\n                  orientation=\"v\", y=1, yanchor=\"top\", x=1.0, xanchor=\"right\" )                 \n                 ).update_xaxes(categoryorder='total descending')\n\nfig.show()","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-02-08T12:30:47.636928Z","iopub.execute_input":"2022-02-08T12:30:47.637431Z","iopub.status.idle":"2022-02-08T12:30:57.708128Z","shell.execute_reply.started":"2022-02-08T12:30:47.637358Z","shell.execute_reply":"2022-02-08T12:30:57.704756Z"},"jupyter":{"source_hidden":true},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### 4. Sample product pictures <a class=\"anchor\" id=\"4\"></a>","metadata":{}},{"cell_type":"code","source":"# copied from https://www.kaggle.com/ruchi798/shopee-eda-rapids-preprocessing-w-b/notebook\n\ndef getImagePaths(path):\n    image_names = []\n    for dirname, _, filenames in os.walk(path):\n        for filename in filenames:\n            fullpath = os.path.join(dirname, filename)\n            image_names.append(fullpath)\n    return image_names\n\ndef display_multiple_img(images_paths, rows, cols,title):\n    \n    figure, ax = plt.subplots(nrows=rows,ncols=cols,figsize=(16,8))\n    plt.suptitle(title, fontsize=20)\n    for ind,image_path in enumerate(images_paths):\n        image = cv2.imread(image_path)\n        image = cv2.cvtColor(image, cv2.COLOR_BGR2RGB) \n        try:\n            ax.ravel()[ind].imshow(image)\n            ax.ravel()[ind].set_axis_off()\n        except:\n            continue;\n    plt.tight_layout()\n    plt.show()","metadata":{"_kg_hide-input":true,"jupyter":{"source_hidden":true},"execution":{"iopub.status.busy":"2022-02-08T12:30:57.710101Z","iopub.execute_input":"2022-02-08T12:30:57.710614Z","iopub.status.idle":"2022-02-08T12:30:57.722966Z","shell.execute_reply.started":"2022-02-08T12:30:57.710532Z","shell.execute_reply":"2022-02-08T12:30:57.720762Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"images_path = getImagePaths('../input/h-and-m-personalized-fashion-recommendations/images/')","metadata":{"_kg_hide-input":true,"jupyter":{"source_hidden":true},"execution":{"iopub.status.busy":"2022-02-08T12:30:57.725021Z","iopub.execute_input":"2022-02-08T12:30:57.725830Z","iopub.status.idle":"2022-02-08T12:32:43.019104Z","shell.execute_reply.started":"2022-02-08T12:30:57.725769Z","shell.execute_reply":"2022-02-08T12:32:43.017760Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"display_multiple_img(images_path[0:25], 5, 5,\"Sample product images\")","metadata":{"_kg_hide-input":true,"jupyter":{"source_hidden":true},"execution":{"iopub.status.busy":"2022-02-08T12:32:43.020955Z","iopub.execute_input":"2022-02-08T12:32:43.021363Z","iopub.status.idle":"2022-02-08T12:32:50.200825Z","shell.execute_reply.started":"2022-02-08T12:32:43.021308Z","shell.execute_reply":"2022-02-08T12:32:50.199580Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"display_multiple_img(images_path[200:220], 5, 4,\"Sample product images\")","metadata":{"_kg_hide-input":true,"jupyter":{"source_hidden":true},"execution":{"iopub.status.busy":"2022-02-08T12:32:50.202018Z","iopub.execute_input":"2022-02-08T12:32:50.202245Z","iopub.status.idle":"2022-02-08T12:32:55.470294Z","shell.execute_reply.started":"2022-02-08T12:32:50.202217Z","shell.execute_reply":"2022-02-08T12:32:55.469684Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### 5. Transactions <a class=\"anchor\" id=\"5\"></a>","metadata":{}},{"cell_type":"code","source":"trans_train.head()","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-02-08T12:32:55.471324Z","iopub.execute_input":"2022-02-08T12:32:55.471730Z","iopub.status.idle":"2022-02-08T12:32:55.485604Z","shell.execute_reply.started":"2022-02-08T12:32:55.471684Z","shell.execute_reply":"2022-02-08T12:32:55.484707Z"},"jupyter":{"source_hidden":true},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"for col in trans_train.columns:\n    x = trans_train[col].nunique()    \n    print(\"{}: ======> {} unique\".format(col, colored(x, 'red')))","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-02-08T12:32:55.487142Z","iopub.execute_input":"2022-02-08T12:32:55.487984Z","iopub.status.idle":"2022-02-08T12:33:09.321039Z","shell.execute_reply.started":"2022-02-08T12:32:55.487935Z","shell.execute_reply":"2022-02-08T12:33:09.319667Z"},"jupyter":{"source_hidden":true},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"trans_train.info()","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-02-08T12:33:09.322765Z","iopub.execute_input":"2022-02-08T12:33:09.323108Z","iopub.status.idle":"2022-02-08T12:33:09.336169Z","shell.execute_reply.started":"2022-02-08T12:33:09.323059Z","shell.execute_reply":"2022-02-08T12:33:09.334982Z"},"jupyter":{"source_hidden":true},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<!-- # fig = px.histogram(trans_train, x=\"price\",\n#                    width=600, \n#                    height=400,\n#                    histnorm='percent',\n#                    template=\"simple_white\",\n#                    color='sales_channel_id'\n#                    )\n\n# fig.update_layout(title=\"Price \", \n#                   font_family=\"San Serif\",\n#                   titlefont={'size': 20},\n#                   legend=dict(\n#                   orientation=\"v\", y=1, yanchor=\"top\", x=1.0, xanchor=\"right\" )                 \n#                  ).update_yaxes(categoryorder='total ascending') \n\n# fig.show() -->","metadata":{"execution":{"iopub.status.busy":"2022-02-07T23:27:20.879959Z","iopub.execute_input":"2022-02-07T23:27:20.880548Z","iopub.status.idle":"2022-02-07T23:27:20.892123Z","shell.execute_reply.started":"2022-02-07T23:27:20.880473Z","shell.execute_reply":"2022-02-07T23:27:20.891196Z"}}},{"cell_type":"markdown","source":"### 6. Reference & credits <a class=\"anchor\" id=\"6\"></a>\n- https://www.kaggle.com/ruchi798/shopee-eda-rapids-preprocessing-w-b/notebook","metadata":{}},{"cell_type":"markdown","source":"#### ...Work in progress...","metadata":{}},{"cell_type":"code","source":"","metadata":{"jupyter":{"source_hidden":true}},"execution_count":null,"outputs":[]}]}