{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Exploratory Data Analytics\n### Exploring the data in detail for modelling insights\n","metadata":{}},{"cell_type":"markdown","source":"#### The data provided by H&M to build the recommender engine is multi-modal in nature, with 3 different modalities of information.  \n\n1) numeric and categorical modality regarding transaction data and article data and customer data  \n2) Natural language data (Article descriptions)  \n3) Image data (Article images)  ","metadata":{}},{"cell_type":"markdown","source":"##### The structure of this notebook is as follows:  \n\nI will be exploring dataset separately,and exploring the potential significance of each modality in detail,as well as any potential insights gathered from each dataset provided.The goal here is this to explore the datasets in detail to the extent where the potential opportunities and strategies are clearly visible.","metadata":{}},{"cell_type":"markdown","source":"____________________________________________________________________\n## Importing all the necessary libraries","metadata":{}},{"cell_type":"code","source":"import pandas as pd\nimport numpy as np\nimport seaborn as sns\nimport matplotlib.pyplot as plt\nimport matplotlib.ticker as mtick\nimport datetime as dt\nimport gc\nimport plotly.express as px\nimport plotly.io as pio\nfrom os import walk\nimport matplotlib.image as mpimg\n","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Exploring Customer Data","metadata":{}},{"cell_type":"markdown","source":"This dataset contains metadata of every customer in detail, collected by H&M during the time period. This data is notably limited in terms of typical modern online recommender systems, as there is no user interaction data with customers. This maybe attribtued to the fact that the data collection for H&M is both physical and online store data combined, which may make storing both data together comparatively difficult.  \n\nThis may potentially reduce the potential of customer data as we posess no information regarding the customer viewed certain articles, if they interacted with the images, how long did they stay on the article before leaving to a different article, did they switch to a different article based on the recommendation engine or did they escape back to the search menu and so on. These interactions would help the model potentially gauge the user's interest in a particular article and therefore adjust accordingly.  \n\n#### The data contained in this table ar as follows :  \n \n1) Customer_id: 64 character unique customer ID (Hex valued)  \n \n2) FN: FN is if a customer gets a Fashion News newsletter (N/A implies NO 1 implies yes, sales channel id, 2 is online and 1 store.   \n \n3) Active: Notifier to determine if the customer is active for communications from H&M\n4) Club_member_status: If the customer is a member at H&M (free membership tier)  \n \n5) Fashion_news_frequency: how frequently are the fashion news letters sent to the customer (none,regularly and monthly)  \n \n6) age: age of the customer  \n \n7) Postal_code: Hash coded 64 bit postal codes for all regions of transactions (potentially includes worldwide postcodes)  \n \n\n##### Note that postal codes are anonymized to avoid customer location data leak,same is the case with not storing customer name as an identifier.","metadata":{}},{"cell_type":"code","source":"#loading customer data and viewing the first 5 elements\ntransactions_data= pd.read_csv('../input/h-and-m-personalized-fashion-recommendations/transactions_train.csv')\ncustomer_data= pd.read_csv('../input/h-and-m-personalized-fashion-recommendations/customers.csv')\nprint('number of rows and columns ',customer_data.shape)\ncustomer_data.head()","metadata":{"execution":{"iopub.status.busy":"2022-08-26T05:08:02.733159Z","iopub.execute_input":"2022-08-26T05:08:02.733649Z","iopub.status.idle":"2022-08-26T05:08:08.626560Z","shell.execute_reply.started":"2022-08-26T05:08:02.733584Z","shell.execute_reply":"2022-08-26T05:08:08.625323Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"1,371980 rows, 7 set of fetures, with 3 text columns and 3 numeric columns ","metadata":{}},{"cell_type":"code","source":"#cheecking for duplicate customers in customer_data\ncustomer_data.shape[0] - customer_data['customer_id'].nunique()\n#basically subtracting number of rows with number of unique values of customer id","metadata":{"execution":{"iopub.status.busy":"2022-08-26T05:08:08.628522Z","iopub.execute_input":"2022-08-26T05:08:08.629055Z","iopub.status.idle":"2022-08-26T05:08:09.316723Z","shell.execute_reply.started":"2022-08-26T05:08:08.629020Z","shell.execute_reply":"2022-08-26T05:08:09.315604Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The data is relatively clean, as there are no duplicate customer Ids in this case.","metadata":{}},{"cell_type":"code","source":"#checking for nulls\ncustomer_data.isnull().sum()","metadata":{"execution":{"iopub.status.busy":"2022-08-26T05:08:09.319148Z","iopub.execute_input":"2022-08-26T05:08:09.319592Z","iopub.status.idle":"2022-08-26T05:08:09.604400Z","shell.execute_reply.started":"2022-08-26T05:08:09.319559Z","shell.execute_reply":"2022-08-26T05:08:09.603208Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#fixing NaN columns in FN and active by replacing it with zero\ncustomer_data_fixed=customer_data\ncustomer_data_fixed[['FN','Active']]=customer_data_fixed[['FN','Active']].fillna(0)\ncustomer_data_fixed['club_member_status']=customer_data_fixed['club_member_status'].fillna('unknown')\ncustomer_data_fixed['fashion_news_frequency']=customer_data_fixed['fashion_news_frequency'].fillna('unknown')","metadata":{"execution":{"iopub.status.busy":"2022-08-26T05:08:09.606686Z","iopub.execute_input":"2022-08-26T05:08:09.607066Z","iopub.status.idle":"2022-08-26T05:08:09.861216Z","shell.execute_reply.started":"2022-08-26T05:08:09.607033Z","shell.execute_reply":"2022-08-26T05:08:09.859903Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Exploring how many customers are get Fashion news letters and how many customers are active for communications","metadata":{}},{"cell_type":"code","source":"\nfig, ax = plt.subplots(figsize=(5,5))\nexplode = (0, 0.1)\ncolors = sns.color_palette('Paired')\nax.pie(customer_data_fixed['Active'].value_counts(), explode=explode, labels=['Not-active','Active'],\n       autopct='%1.1f%%',shadow=True, startangle=90, colors=colors)\nax.axis('equal')\nplt.show()\n\n\nfig, ax = plt.subplots(figsize=(5,5))\nexplode = (0, 0.1)\ncolors = sns.color_palette('Paired')\nax.pie(customer_data_fixed['FN'].value_counts(), explode=explode, labels=['Not-FN','FN'],\n       autopct='%1.1f%%',shadow=True, startangle=90, colors=colors)\nax.axis('equal')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-08-26T05:08:09.863863Z","iopub.execute_input":"2022-08-26T05:08:09.865775Z","iopub.status.idle":"2022-08-26T05:08:10.011281Z","shell.execute_reply.started":"2022-08-26T05:08:09.865714Z","shell.execute_reply":"2022-08-26T05:08:10.010370Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Interesting thing to note, that total number of customers that are active for communication is slightly smaller than total number of customers that are willing to recive a fashion letter.\n\nthere is a marked difference of 0.9% who are willing to recieve a newsletter but do not want active communications from H&M","metadata":{}},{"cell_type":"markdown","source":"**Exploring customer membership data spread**","metadata":{}},{"cell_type":"code","source":"sns.catplot(data=customer_data_fixed,x='club_member_status',kind='count')","metadata":{"execution":{"iopub.status.busy":"2022-08-26T05:08:10.012477Z","iopub.execute_input":"2022-08-26T05:08:10.013190Z","iopub.status.idle":"2022-08-26T05:08:12.375950Z","shell.execute_reply.started":"2022-08-26T05:08:10.013157Z","shell.execute_reply":"2022-08-26T05:08:12.374792Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"active=customer_data_fixed['club_member_status'].loc[customer_data_fixed['club_member_status']=='ACTIVE']\nunknown=customer_data_fixed['club_member_status'].loc[customer_data_fixed['club_member_status']=='unknown']\npre_create=customer_data_fixed['club_member_status'].loc[customer_data_fixed['club_member_status']=='PRE-CREATE']\nactive=active.size\nunknown=unknown.size\npre_create=pre_create.size\ntotal=active+unknown+pre_create\n\nactive_per=(active*100)/total\nunknown_per=(unknown*100)/total\npre_create_per=(pre_create*100)/total\n\nprint('Percentage of active club member customers =', active_per)\nprint('Percentage of unknown club member customers =', unknown_per)\nprint('Percentage of Pre-create club member customers =', pre_create_per)","metadata":{"execution":{"iopub.status.busy":"2022-08-26T05:08:12.377383Z","iopub.execute_input":"2022-08-26T05:08:12.377727Z","iopub.status.idle":"2022-08-26T05:08:12.691647Z","shell.execute_reply.started":"2022-08-26T05:08:12.377698Z","shell.execute_reply":"2022-08-26T05:08:12.690407Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"\n##code to create a histogram plot for the distribution of customers age\n##with a marked median value represented on the graph\n##(using median to avoid the problem of outliers)\n\nfig, ax = plt.subplots(figsize=(10,5))\nax = sns.histplot(data=customer_data_fixed, x='age', bins=customer_data_fixed['age'].nunique(), color='orange', stat=\"percent\")\nax.set_xlabel('Distribution of the customers age')\nfor loc in ['bottom', 'left']:\n    ax.spines[loc].set_visible(True)\n    ax.spines[loc].set_linewidth(2)\n    ax.spines[loc].set_color('black')\nax.yaxis.set_major_formatter(mtick.PercentFormatter())\nmedian = customer_data_fixed['age'].median()\nax.axvline(x=median, color=\"green\", ls=\"--\")\nax.text(median, 3.5, 'median: {}'.format(round(median,1)), rotation='vertical', ha='right')\nax.text(12, 5.5, 'Distribution of customers age', color='black', fontsize=10, ha='left', va='bottom', weight='bold', style='italic')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-08-26T05:08:12.952861Z","iopub.execute_input":"2022-08-26T05:08:12.953219Z","iopub.status.idle":"2022-08-26T05:08:13.436057Z","shell.execute_reply.started":"2022-08-26T05:08:12.953187Z","shell.execute_reply":"2022-08-26T05:08:13.434882Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Interestingly,we can notice that there are two main groups of customers, one main group between 20-30 year old (young customers which is expected for H&M, and 45-55 years old which would potentially be older customers buying for younger people)","metadata":{}},{"cell_type":"code","source":"## plotting a similar graph for active vs inactive customers \nfig, ax = plt.subplots(figsize=(10,5))\nax = sns.histplot(data=customer_data_fixed, x='age', bins=customer_data_fixed['age'].nunique(), hue='Active', stat=\"percent\")\nax.set_xlabel('Distribution of the customers age')\nfor loc in ['bottom', 'left']:\n    ax.spines[loc].set_visible(True)\n    ax.spines[loc].set_linewidth(2)\n    ax.spines[loc].set_color('black')\nax.yaxis.set_major_formatter(mtick.PercentFormatter())\nplt.show()\n","metadata":{"execution":{"iopub.status.busy":"2022-08-26T05:08:13.883978Z","iopub.execute_input":"2022-08-26T05:08:13.884882Z","iopub.status.idle":"2022-08-26T05:08:14.769394Z","shell.execute_reply.started":"2022-08-26T05:08:13.884846Z","shell.execute_reply":"2022-08-26T05:08:14.767906Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The above graph shows that customers that have activated communication follow the same trend, although there are much more inactive customers  in comparison. ","metadata":{}},{"cell_type":"markdown","source":"#### Checking for how many customers are in each postcode","metadata":{}},{"cell_type":"code","source":"postal_data=customer_data_fixed.groupby('postal_code',as_index=False).count().sort_values('customer_id',ascending=False)\n#grouping by postcode, and counting the values and sorting the count in descending order\npostal_data.head()","metadata":{"execution":{"iopub.status.busy":"2022-08-26T03:56:51.146587Z","iopub.execute_input":"2022-08-26T03:56:51.147115Z","iopub.status.idle":"2022-08-26T03:56:52.300464Z","shell.execute_reply.started":"2022-08-26T03:56:51.147082Z","shell.execute_reply":"2022-08-26T03:56:52.299564Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Exploring this postcode further, to see if the customer characteristics are different at this postcode or are they the same (which marks it as a wholesale retailer).","metadata":{}},{"cell_type":"code","source":"customer_data_fixed[customer_data_fixed['postal_code']=='2c29ae653a9282cce4151bd87643c907644e09541abc28ae87dea0d1f6603b1c'].head(15)","metadata":{"execution":{"iopub.status.busy":"2022-08-26T05:08:19.220139Z","iopub.execute_input":"2022-08-26T05:08:19.220544Z","iopub.status.idle":"2022-08-26T05:08:19.389682Z","shell.execute_reply.started":"2022-08-26T05:08:19.220512Z","shell.execute_reply":"2022-08-26T05:08:19.388396Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Note that any model created here would have to deal with nulls in the age category, or alternatively the age category would have to be cleaned with imputation methods to prevent loss of useful data.\n","metadata":{}},{"cell_type":"code","source":"#creating a code for imputation of mean value for age\n#two main imputation approaches \n\n#https://towardsdatascience.com/imputing-missing-data-with-simple-and-advanced-techniques-f5c7b157fb87\nfrom sklearn.impute import KNNImputer\nfrom sklearn.preprocessing import MinMaxScaler\nfrom sklearn.experimental import enable_iterative_imputer\nfrom sklearn.impute import IterativeImputer\nfrom sklearn import linear_model\nfrom sklearn.preprocessing import OrdinalEncoder\n# 1 KNN imputer\ncustomer_data_temp=customer_data_fixed\nencoder=OrdinalEncoder()\ncustomer_data_temp=pd.DataFrame(encoder.fit_transform(customer_data_temp))\n#minmax scaling the data to avoid any bias\n\n\nknn_imputer= KNNImputer(n_neighbors=10,weights='uniform',metric='nan_euclidean')\ncustomer_data_Knn=pd.DataFrame(knn_imputer.fit_transform(customer_data_temp),columns=customer_data_fixed.columns)\n\ncustomer_data_fixed_knn=customer_data_fixed\ncustomer_data_fixed_knn['age']=customer_data_Knn['age']\n#2  Mice Imputer (regression based imputation)\n\n","metadata":{"execution":{"iopub.status.busy":"2022-08-26T05:08:45.866064Z","iopub.execute_input":"2022-08-26T05:08:45.866811Z","iopub.status.idle":"2022-08-26T05:08:59.619589Z","shell.execute_reply.started":"2022-08-26T05:08:45.866767Z","shell.execute_reply":"2022-08-26T05:08:59.618058Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"customer_data_transformed_knn=pd.DataFrame(encoder.inverse_transform(customer_data_Knn))\ncustomer_data_transformed_knn.columns=customer_data_fixed.columns\ncustomer_data_transformed_knn","metadata":{"execution":{"iopub.status.busy":"2022-08-26T04:11:12.943623Z","iopub.execute_input":"2022-08-26T04:11:12.944049Z","iopub.status.idle":"2022-08-26T04:11:13.812628Z","shell.execute_reply.started":"2022-08-26T04:11:12.944017Z","shell.execute_reply":"2022-08-26T04:11:13.811754Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from sklearn.experimental import enable_iterative_imputer\nfrom sklearn.impute import IterativeImputer\nfrom sklearn import linear_model\nfrom sklearn.preprocessing import OrdinalEncoder\n#note that mice imputation works much much faster in comparision to KNN imputation\nmice_imputer = IterativeImputer(estimator=linear_model.BayesianRidge(), n_nearest_features=None, imputation_order='ascending')\n\ncustomer_data_m_imputed= pd.DataFrame(mice_imputer.fit_transform(customer_data_temp), columns=customer_data_fixed.columns)\ncustomer_data_m_imputed=pd.DataFrame(encoder.inverse_transform(customer_data_m_imputed))\ncustomer_data_m_imputed.columns=customer_data_fixed.columns\ncustomer_data_fixed_mice=customer_data_fixed\ncustomer_data_fixed_mice['age']=customer_data_m_imputed['age']\n\ncustomer_data_fixed_mice","metadata":{"execution":{"iopub.status.busy":"2022-08-26T05:09:03.286256Z","iopub.execute_input":"2022-08-26T05:09:03.286944Z","iopub.status.idle":"2022-08-26T05:09:09.631521Z","shell.execute_reply.started":"2022-08-26T05:09:03.286908Z","shell.execute_reply":"2022-08-26T05:09:09.630576Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"customer_data_fixed_mice['age'].astype('int32')","metadata":{"execution":{"iopub.status.busy":"2022-08-26T05:09:15.597829Z","iopub.execute_input":"2022-08-26T05:09:15.598909Z","iopub.status.idle":"2022-08-26T05:09:15.686634Z","shell.execute_reply.started":"2022-08-26T05:09:15.598858Z","shell.execute_reply":"2022-08-26T05:09:15.685827Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# exploring article data","metadata":{}},{"cell_type":"markdown","source":"The articles dataset contains detailed descriptions and parameters for the articles being sold/added to the catalogue. There are multiple features present dataset with 25 columns and these \nfeatures can be classified which can be split into these unique categories as follows: \n\n\n### 1) Unique indentifier of an article:  \n\na) article_id (int64) - an unique 9-digit identifier of the article, 105 542 unique values (as the length of the database)  \n\n### 2) 5 product related columns:  \n\na) product_code (int64) - 6-digit product code (the first 6 digits of article_id, 47 224 unique values \n\nb) prod_name (object) - name of a product, 45 875 unique values\n \nc) product_type_no (int64) - product type number, 131 unique values \n\nd) product_type_name (object) - name of a product type, equivalent of product_type_no\n \ne) product_group_name (object) - name of a product group, in total 19 groups\n\n### 3) 2 columns related to the pattern:\n\na) graphical_appearance_no (int64) - code of a pattern, 30 unique values \n\nb) graphical_appearance_name (object) - name of a pattern, 30 unique values\n\n### 4) 2 columns related to the color:\n\na) colour_group_code (int64) - code of a color, 50 unique values \n\nb) colour_group_name (object) - name of a color, 50 unique values\n\n### 5) 4 columns related to perceived colour (general tone):\n\na) perceived_colour_value_id - perceived color id, 8 unique values \n\nb) perceived_colour_value_name - perceived color name, 8 unique values \n\nc) perceived_colour_master_id - perceived master color id, 20 unique values \n\ne) perceived_colour_master_name - perceived master color name, 20 unique values\n\n### 6) 2 columns related to the department:\n\na) department_no - department number, 299 unique values \n\nb) department_name - department name, 299 unique values \n\n### 7) 4 columns related to the index, which is actually a top-level category:\n\na) index_code - index code, 10 unique values \n\nb) index_name - index name, 10 unique values\n \nc) index_group_no - index group code, 5 unique values\n \nd) index_group_name - index group code, 5 unique values \n\n### 8) 2 columns related to the section:\n\na) section_no - section number, 56 unique values\n \nb) section_name - section name, 56 unique values\n \n\n### 9) 2 columns related to the garment group:\n\na) garment_group_n - section number, 56 unique values\n \nb) garment_group_name - section name, 56 unique values\n \n\n### 10) 1 column with a detailed description of the article:\n\na) detail_desc - 43 404 unique values\n\n####  Note :  The garment detail_desc column is the one that has a detailed description of the article which can and will be used for natural language processing, to explore the textual detail as a modality that influences users decisions","metadata":{}},{"cell_type":"code","source":"article_data= pd.read_csv('../input/h-and-m-personalized-fashion-recommendations/articles.csv')\narticle_data.head()","metadata":{"execution":{"iopub.status.busy":"2022-08-16T12:06:13.095782Z","iopub.status.idle":"2022-08-16T12:06:13.096431Z","shell.execute_reply.started":"2022-08-16T12:06:13.096218Z","shell.execute_reply":"2022-08-16T12:06:13.096239Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#printing the percentage of null data per column in the dataset\narticle_data.isnull().sum()/len(article_data)*100","metadata":{"execution":{"iopub.status.busy":"2022-08-16T12:06:13.097648Z","iopub.status.idle":"2022-08-16T12:06:13.098297Z","shell.execute_reply.started":"2022-08-16T12:06:13.098085Z","shell.execute_reply":"2022-08-16T12:06:13.098106Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Notable that there's about 0.4% of nulls in the details dataset, which is extremely minimal and potentially can be ignored/dealt with (unless these are high sale items).","metadata":{}},{"cell_type":"markdown","source":"Note : not all articles have an image as described in the competitino description. Creating a function to show a selected amount of images from any given folder, with its article name. ","metadata":{}},{"cell_type":"code","source":"#Note that the file names do not correspond to the first digit of the article id and have a leading 0\n#stripping the 0 is necessary to get the right product name.\n\ndef show_articles(folder, no_images=3):\n    folder_path = '../input/h-and-m-personalized-fashion-recommendations/images/{}/'.format(folder)\n    # extracting all image names from a folder\n    files = []\n    for _, _, filenames in walk(folder_path):\n        files.extend(filenames)\n    no_files = len(files)\n    if no_images > no_files:\n        no_images = no_files\n        print(\"Warning! In the folder there are less images than requested.\")\n        \n    # plotting selected number of pictures\n    images = files[:no_images]\n    fig, ax = plt.subplots(1,no_images, figsize=(12,4))\n    for i, img in enumerate(images):\n        art_id = img.split('.')[0]\n        img = plt.imread(folder_path+img)\n        ax[i].imshow(img, aspect='equal')\n        ax[i].grid(False)\n        ax[i].set_xticks([], [])\n        ax[i].set_yticks([], [])\n        ax[i].set_xlabel(article_data[article_data['article_id']==int(art_id[1:])]['prod_name'].iloc[0])\n    plt.show()\n    \nshow_articles('020',4)\n","metadata":{"execution":{"iopub.status.busy":"2022-08-16T12:06:13.099468Z","iopub.status.idle":"2022-08-16T12:06:13.100105Z","shell.execute_reply.started":"2022-08-16T12:06:13.099902Z","shell.execute_reply":"2022-08-16T12:06:13.099923Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"this visualisation is interesting, but now let us try to combine both the photos, with the article description to get the photos, description and price at the same time.","metadata":{}},{"cell_type":"markdown","source":"#### Viewing photos with description as well as price for most expensive articles as well as least expensive articles","metadata":{}},{"cell_type":"code","source":"#creating transactions data sorted by values of price and article ids\nmax_price_ids = transactions_data[transactions_data.t_dat==transactions_data.t_dat.max()].sort_values('price', ascending=False).iloc[:5][['article_id', 'price']]\nmin_price_ids = transactions_data[transactions_data.t_dat==transactions_data.t_dat.min()].sort_values('price', ascending=True).iloc[:5][['article_id', 'price']]\n","metadata":{"execution":{"iopub.status.busy":"2022-08-16T12:06:13.101305Z","iopub.status.idle":"2022-08-16T12:06:13.101986Z","shell.execute_reply.started":"2022-08-16T12:06:13.101786Z","shell.execute_reply":"2022-08-16T12:06:13.101807Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"f, ax = plt.subplots(1, 5, figsize=(20,10))\ni = 0\nfor _, data in max_price_ids.iterrows():\n    desc = article_data[article_data['article_id'] == data['article_id']]['detail_desc'].iloc[0]\n    desc_list = desc.split(' ')\n    for j, elem in enumerate(desc_list):\n        if j > 0 and j % 5 == 0:\n            desc_list[j] = desc_list[j] + '\\n'\n    desc = ' '.join(desc_list)\n    img = mpimg.imread(f'../input/h-and-m-personalized-fashion-recommendations/images/0{str(data.article_id)[:2]}/0{int(data.article_id)}.jpg')\n    ax[i].imshow(img)\n    ax[i].set_title(f'price: {data.price:.2f}')\n    ax[i].imshow(img)\n    ax[i].set_title(f'price: {data.price:.2f}')\n    ax[i].set_xticks([], [])\n    ax[i].set_yticks([], [])\n    ax[i].grid(False)\n    ax[i].set_xlabel(desc, fontsize=10)\n    i += 1\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-08-16T12:06:13.103164Z","iopub.status.idle":"2022-08-16T12:06:13.103823Z","shell.execute_reply.started":"2022-08-16T12:06:13.103621Z","shell.execute_reply":"2022-08-16T12:06:13.103641Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"we can make an interesting note that its generally leather items that come in the most expensive price bracket in H&M's diverse product range (leather apparel specifically)\n","metadata":{}},{"cell_type":"code","source":"f, ax = plt.subplots(1, 5, figsize=(20,10))\ni = 0\nfor _, data in min_price_ids.iterrows():\n    desc = article_data[article_data['article_id'] == data['article_id']]['detail_desc'].iloc[0]\n    desc_list = desc.split(' ')\n    for j, elem in enumerate(desc_list):\n        if j > 0 and j % 4 == 0:\n            desc_list[j] = desc_list[j] + '\\n'\n    desc = ' '.join(desc_list)\n    img = mpimg.imread(f'../input/h-and-m-personalized-fashion-recommendations/images/0{str(data.article_id)[:2]}/0{int(data.article_id)}.jpg')\n    ax[i].imshow(img)\n    ax[i].set_title(f'price: {data.price:.2f}')\n    ax[i].imshow(img)\n    ax[i].set_title(f'price: {data.price:.2f}')\n    ax[i].set_xticks([], [])\n    ax[i].set_yticks([], [])\n    ax[i].grid(False)\n    ax[i].set_xlabel(desc, fontsize=10)\n    i += 1\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-08-16T12:06:13.105035Z","iopub.status.idle":"2022-08-16T12:06:13.105436Z","shell.execute_reply.started":"2022-08-16T12:06:13.105243Z","shell.execute_reply":"2022-08-16T12:06:13.105264Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"\nInteresting point to note that the cheapest items are as expected the trivial accesory items in H&Ms catalog","metadata":{}},{"cell_type":"markdown","source":"### Viewing the top 5 and the bottom 5 most sold articles ","metadata":{}},{"cell_type":"code","source":"z=np.arange(104547)\nmost_sold_articles=transactions_data[['article_id']].value_counts().rename_axis('article_id').reset_index(name='counts')\nleast_sold_articles=transactions_data[['article_id']].value_counts().sort_values().rename_axis('article_id').reset_index(name='counts')\n\n\nmost_sold_articles=most_sold_articles.iloc[:6][['article_id', 'counts']]\nleast_sold_articles=least_sold_articles.iloc[:5][['article_id', 'counts']]\n\nmost_sold_articles=most_sold_articles.drop([3])","metadata":{"execution":{"iopub.status.busy":"2022-08-16T12:06:13.107038Z","iopub.status.idle":"2022-08-16T12:06:13.107459Z","shell.execute_reply.started":"2022-08-16T12:06:13.107265Z","shell.execute_reply":"2022-08-16T12:06:13.107284Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"interesting thing to note, the 4th most sold item does not have a picture available on the site (curious, as one would expect articles without pictures to not sell well/not be listed on the site) maybe these articles are sold in store instead","metadata":{}},{"cell_type":"markdown","source":"#### top 5 most sold articles","metadata":{}},{"cell_type":"code","source":"f, ax = plt.subplots(1, 5, figsize=(20,10))\ni = 0\nfor _, data in most_sold_articles.iterrows():\n    desc = article_data[article_data['article_id'] == data['article_id']]['detail_desc'].iloc[0]\n    desc_list = desc.split(' ')\n    for j, elem in enumerate(desc_list):\n        if j > 0 and j % 4 == 0:\n            desc_list[j] = desc_list[j] + '\\n'\n    desc = ' '.join(desc_list)\n    img = mpimg.imread(f'../input/h-and-m-personalized-fashion-recommendations/images/0{str(data.article_id)[:2]}/0{int(data.article_id)}.jpg')\n    ax[i].imshow(img)\n    ax[i].set_title(f'count: {data.counts:.2f}')\n    ax[i].imshow(img)\n    ax[i].set_title(f'count: {data.counts:.2f}')\n    ax[i].set_xticks([], [])\n    ax[i].set_yticks([], [])\n    ax[i].grid(False)\n    ax[i].set_xlabel(desc, fontsize=10)\n    i += 1\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-08-16T12:06:13.108822Z","iopub.status.idle":"2022-08-16T12:06:13.109227Z","shell.execute_reply.started":"2022-08-16T12:06:13.109028Z","shell.execute_reply":"2022-08-16T12:06:13.109048Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"interestingly the top 5 most sold articles are all either black or white and are what you would expect (basic essential clothing items like jeans, socks and camis)","metadata":{}},{"cell_type":"markdown","source":"### Checking the next 5 most sold items to confirm that the trend of essentials being the most sold items","metadata":{}},{"cell_type":"code","source":"most_sold_articles=transactions_data[['article_id']].value_counts().rename_axis('article_id').reset_index(name='counts')\nmost_sold_articles=most_sold_articles.iloc[6:12][['article_id', 'counts']]\n\nmost_sold_articles=most_sold_articles.drop([7])\n\nf, ax = plt.subplots(1, 5, figsize=(20,10))\ni = 0\nfor _, data in most_sold_articles.iterrows():\n    desc = article_data[article_data['article_id'] == data['article_id']]['detail_desc'].iloc[0]\n    desc_list = desc.split(' ')\n    for j, elem in enumerate(desc_list):\n        if j > 0 and j % 4 == 0:\n            desc_list[j] = desc_list[j] + '\\n'\n    desc = ' '.join(desc_list)\n    img = mpimg.imread(f'../input/h-and-m-personalized-fashion-recommendations/images/0{str(data.article_id)[:2]}/0{int(data.article_id)}.jpg')\n    ax[i].imshow(img)\n    ax[i].set_title(f'count: {data.counts:.2f}')\n    ax[i].imshow(img)\n    ax[i].set_title(f'count: {data.counts:.2f}')\n    ax[i].set_xticks([], [])\n    ax[i].set_yticks([], [])\n    ax[i].grid(False)\n    ax[i].set_xlabel(desc, fontsize=10)\n    i += 1\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-08-16T12:06:13.110466Z","iopub.status.idle":"2022-08-16T12:06:13.110900Z","shell.execute_reply.started":"2022-08-16T12:06:13.110701Z","shell.execute_reply":"2022-08-16T12:06:13.110721Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"viewing the above confirms the same, the most sold articles seem to be women's jeans and leggings/tights as well as womens essentials (cami tops, socks, mesh tights, underwear) in the most standard colours (black, white and blue in case of jeans)","metadata":{}},{"cell_type":"markdown","source":"#### Bottom 5 most sold articles","metadata":{}},{"cell_type":"code","source":"f, ax = plt.subplots(1, 5, figsize=(20,10))\ni = 0\nfor _, data in least_sold_articles.iterrows():\n    desc = article_data[article_data['article_id'] == data['article_id']]['detail_desc'].iloc[0]\n    desc_list = desc.split(' ')\n    for j, elem in enumerate(desc_list):\n        if j > 0 and j % 4 == 0:\n            desc_list[j] = desc_list[j] + '\\n'\n    desc = ' '.join(desc_list)\n    img = mpimg.imread(f'../input/h-and-m-personalized-fashion-recommendations/images/0{str(data.article_id)[:2]}/0{int(data.article_id)}.jpg')\n    ax[i].imshow(img)\n    ax[i].set_title(f'count: {data.counts:.2f}')\n    ax[i].imshow(img)\n    ax[i].set_title(f'count: {data.counts:.2f}')\n    ax[i].set_xticks([], [])\n    ax[i].set_yticks([], [])\n    ax[i].grid(False)\n    ax[i].set_xlabel(desc, fontsize=10)\n    i += 1\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-08-16T12:06:13.113132Z","iopub.status.idle":"2022-08-16T12:06:13.113596Z","shell.execute_reply.started":"2022-08-16T12:06:13.113358Z","shell.execute_reply":"2022-08-16T12:06:13.113378Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"interestingly, it seems like the lest sold items are all niche products or recently added items and there doesn't seem to be any exact trend for the same.","metadata":{}},{"cell_type":"markdown","source":"### Viewing the next 5 least sold items as well to confirm","metadata":{}},{"cell_type":"code","source":"least_sold_articles=transactions_data[['article_id']].value_counts().sort_values().rename_axis('article_id').reset_index(name='counts')\nleast_sold_articles=least_sold_articles.iloc[6:11][['article_id', 'counts']]\n\nf, ax = plt.subplots(1, 5, figsize=(20,10))\ni = 0\nfor _, data in least_sold_articles.iterrows():\n    desc = article_data[article_data['article_id'] == data['article_id']]['detail_desc'].iloc[0]\n    desc_list = desc.split(' ')\n    for j, elem in enumerate(desc_list):\n        if j > 0 and j % 4 == 0:\n            desc_list[j] = desc_list[j] + '\\n'\n    desc = ' '.join(desc_list)\n    img = mpimg.imread(f'../input/h-and-m-personalized-fashion-recommendations/images/0{str(data.article_id)[:2]}/0{int(data.article_id)}.jpg')\n    ax[i].imshow(img)\n    ax[i].set_title(f'count: {data.counts:.2f}')\n    ax[i].imshow(img)\n    ax[i].set_title(f'count: {data.counts:.2f}')\n    ax[i].set_xticks([], [])\n    ax[i].set_yticks([], [])\n    ax[i].grid(False)\n    ax[i].set_xlabel(desc, fontsize=10)\n    i += 1\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-08-16T12:06:13.115772Z","iopub.status.idle":"2022-08-16T12:06:13.116555Z","shell.execute_reply.started":"2022-08-16T12:06:13.116279Z","shell.execute_reply":"2022-08-16T12:06:13.116316Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Interesting thing to note, there are no sales numbers for the size, therefore its only article sales in itself. further it seems that some articles that look like high volume items don't seem to sell well (recently added? or some alternative explanation?)","metadata":{}},{"cell_type":"markdown","source":"## Exploring the Classification of articles by grouping of articles by collection","metadata":{}},{"cell_type":"markdown","source":"Let us explore the grouping of articles by their collection index names (for example ladieswear, baby collection, divided collection, menswear collection, sport etc).\n\nGraphing it by percentage of total amount of articles to get a better idea of the ratio of the collections with H&M","metadata":{}},{"cell_type":"code","source":"#creating a simple function to get good barplots \n#in percentages, organised in descending order.\n\ndef plot_bar(database, col, figsize=(13,5), pct=False, label='articles'):\n    fig, ax = plt.subplots(figsize=figsize, facecolor='#f6f6f6')\n    for loc in ['bottom', 'left']:\n        ax.spines[loc].set_visible(True)\n        ax.spines[loc].set_linewidth(2)\n        ax.spines[loc].set_color('black')\n    ax.spines['right'].set_visible(False)\n    ax.spines['top'].set_visible(False)\n    \n    if pct:\n        data = database[col].value_counts()\n        data = data.div(data.sum()).mul(100)\n        data = data.reset_index()\n        ax = sns.barplot(data=data, x=col, y='index', color='#2693d7', lw=1.5, ec='black', zorder=2)\n        ax.set_xlabel('% of ' + label, fontsize=10, weight='bold')\n        ax.xaxis.set_major_formatter(mtick.PercentFormatter())\n    else:\n        data = database[col].value_counts().reset_index()\n        ax = sns.barplot(data=data, x=col, y='index', color='#2693d7', lw=1.5, ec='black', zorder=2)        \n        ax.set_xlabel('# of articles' + label)\n        \n    ax.grid(zorder=0)\n    ax.text(0, -0.75, col, color='black', fontsize=10, ha='left', va='bottom', weight='bold', style='italic')\n    ax.set_ylabel('')\n        \n    plt.show()\n    \n\nplot_bar(article_data, 'index_name', pct=True)","metadata":{"execution":{"iopub.status.busy":"2022-08-16T12:06:13.118485Z","iopub.status.idle":"2022-08-16T12:06:13.118957Z","shell.execute_reply.started":"2022-08-16T12:06:13.118740Z","shell.execute_reply":"2022-08-16T12:06:13.118760Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Printing the percentages for the report\ndata = article_data['index_name'].value_counts()\ndata = data.div(data.sum()).mul(100)\ndata","metadata":{"execution":{"iopub.status.busy":"2022-08-16T12:06:13.120973Z","iopub.status.idle":"2022-08-16T12:06:13.121398Z","shell.execute_reply.started":"2022-08-16T12:06:13.121189Z","shell.execute_reply":"2022-08-16T12:06:13.121210Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We can note that most amount of articles are present in the ladies wear category, but surprisingly there is a second category called 'divided' which while seems vague, upon checking comes out to be the teenagers clothing line (including both girls and boys clothing). \n\nthe least amount of items seem to be in sport.\n\n##### note that the index_name and index_group_name columns together form a multi-index table, therefore let us view the graph for the cumulative index_name collection as well ","metadata":{}},{"cell_type":"code","source":"plot_bar(article_data, 'index_group_name', pct=True)","metadata":{"execution":{"iopub.status.busy":"2022-08-16T12:06:13.122789Z","iopub.status.idle":"2022-08-16T12:06:13.123182Z","shell.execute_reply.started":"2022-08-16T12:06:13.122991Z","shell.execute_reply":"2022-08-16T12:06:13.123010Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Printing the percentages for the report\ndata = article_data['index_group_name'].value_counts()\ndata = data.div(data.sum()).mul(100)\ndata","metadata":{"execution":{"iopub.status.busy":"2022-08-16T12:06:13.124712Z","iopub.status.idle":"2022-08-16T12:06:13.125135Z","shell.execute_reply.started":"2022-08-16T12:06:13.124939Z","shell.execute_reply":"2022-08-16T12:06:13.124959Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"we can note that when cumulating all the articles designated for ladies and children as a broader category, the ratios change, making children the second biggest category instead of divided and increasing the proportion of ladies wear clothing in total.\n\n(ladieswear now includes ladies undergarments, accesories as well)","metadata":{}},{"cell_type":"markdown","source":"#### Note that there's an even finer category called product_group_name which splits the clothing into smaller categories such as upper body, lower body, fullbody etc. \n\n## Plotting a graph that includes all the collection categories (baby,children, divided etc) ","metadata":{}},{"cell_type":"code","source":"plot_bar(article_data, 'product_group_name', pct=True)","metadata":{"execution":{"iopub.status.busy":"2022-08-16T12:06:13.126828Z","iopub.status.idle":"2022-08-16T12:06:13.127225Z","shell.execute_reply.started":"2022-08-16T12:06:13.127030Z","shell.execute_reply":"2022-08-16T12:06:13.127050Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Since there are so many categories, let us try to see it in terms of a list of percentages (also useful for reporting)","metadata":{}},{"cell_type":"code","source":"data = article_data['product_group_name'].value_counts()\ndata = data.div(data.sum()).mul(100)\ndata","metadata":{"execution":{"iopub.status.busy":"2022-08-16T12:06:13.130448Z","iopub.status.idle":"2022-08-16T12:06:13.131105Z","shell.execute_reply.started":"2022-08-16T12:06:13.130806Z","shell.execute_reply":"2022-08-16T12:06:13.130833Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### Exploring the sub cateogories for each product group (there are way too many to reasonably plot and get a proportion therefore printing the count for each cateory as well as printing them out with count of articles in each category as a table for better understanding","metadata":{}},{"cell_type":"code","source":"for group in article_data['product_group_name'].unique():\n    print('Number of subcategories in \"{}\"\" is {}.'.format(group, len(article_data.groupby(['product_group_name', 'product_type_name']).size()[group])))","metadata":{"execution":{"iopub.status.busy":"2022-08-16T12:06:13.132668Z","iopub.status.idle":"2022-08-16T12:06:13.133223Z","shell.execute_reply.started":"2022-08-16T12:06:13.132931Z","shell.execute_reply":"2022-08-16T12:06:13.132959Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"pd.options.display.max_rows = None\npd.DataFrame(article_data.groupby(['product_group_name', 'product_type_name']).count()['article_id']).sort_values(['product_group_name','article_id'])","metadata":{"execution":{"iopub.status.busy":"2022-08-16T12:06:13.134748Z","iopub.status.idle":"2022-08-16T12:06:13.135285Z","shell.execute_reply.started":"2022-08-16T12:06:13.135008Z","shell.execute_reply":"2022-08-16T12:06:13.135035Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Key takeaways in indexing\n\n#### The hierarchy of indexing in the article data seems to be as index group -> index -> group -> type\n\n#### over 80% of products lie in 4 product groups out of a total of 19","metadata":{}},{"cell_type":"markdown","source":"## Viewing more article data, this time exploring the colours instead","metadata":{}},{"cell_type":"markdown","source":"it seems like H&M categorise their colours in a very interesting manner, by having an actual 'colour group name' which goes into more detailed colour descriptions, as well as having a 'percived colour value range' which is a more generic categorisation of the colours present.\n\nfurther they have an additional category 'percived colour master name' which is the super category of the percived colour value range\n\nFurther, they also categorise the colour/graphical appearance of the data by 'patterns' which refer to the graphical patterning on the articles such as printed, embroidered etc.","metadata":{}},{"cell_type":"markdown","source":"#### Viewing the broad colour group names","metadata":{}},{"cell_type":"code","source":"plot_bar(article_data,'colour_group_name',figsize=(15,15),pct=True)","metadata":{"execution":{"iopub.status.busy":"2022-08-16T12:06:13.136778Z","iopub.status.idle":"2022-08-16T12:06:13.137189Z","shell.execute_reply.started":"2022-08-16T12:06:13.136988Z","shell.execute_reply":"2022-08-16T12:06:13.137008Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"data = article_data['colour_group_name'].value_counts()\ndata = data.div(data.sum()).mul(100)\ndata","metadata":{"execution":{"iopub.status.busy":"2022-08-16T12:06:13.138528Z","iopub.status.idle":"2022-08-16T12:06:13.139061Z","shell.execute_reply.started":"2022-08-16T12:06:13.138782Z","shell.execute_reply":"2022-08-16T12:06:13.138809Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The most common colours are what you would typically expect/think as the most common colours in fashion such as black,dark blue, white i.e the safe colour choices ( curious as you'd think something like H&M being in fast fashion would have more bright colours dominating)\n\ninterestingly, we can see that a very small fraction of the colours are categorized as unkown as well","metadata":{}},{"cell_type":"markdown","source":"### Viewing the percived colour master name section (the grouped version of the broader colour group names)","metadata":{}},{"cell_type":"code","source":"plot_bar(article_data,'perceived_colour_master_name',pct=True)","metadata":{"execution":{"iopub.status.busy":"2022-08-16T12:06:13.140368Z","iopub.status.idle":"2022-08-16T12:06:13.140923Z","shell.execute_reply.started":"2022-08-16T12:06:13.140645Z","shell.execute_reply":"2022-08-16T12:06:13.140670Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"data = article_data['perceived_colour_master_name'].value_counts()\ndata = data.div(data.sum()).mul(100)\ndata","metadata":{"execution":{"iopub.status.busy":"2022-08-16T12:06:13.142255Z","iopub.status.idle":"2022-08-16T12:06:13.142760Z","shell.execute_reply.started":"2022-08-16T12:06:13.142553Z","shell.execute_reply":"2022-08-16T12:06:13.142576Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"we can note that this reduces and groups the more varied colour descriptions to a much smaller list without mentioning colour gradients and other finer details.","metadata":{}},{"cell_type":"markdown","source":"#### Viewing Perceived colour value name, which categorises the way we see the colour in terms of dark, light, dusty, bright etc. note this is includes all colours","metadata":{}},{"cell_type":"code","source":"plot_bar (article_data,'perceived_colour_value_name',pct=True)","metadata":{"execution":{"iopub.status.busy":"2022-08-16T12:06:13.145779Z","iopub.status.idle":"2022-08-16T12:06:13.146138Z","shell.execute_reply.started":"2022-08-16T12:06:13.145959Z","shell.execute_reply":"2022-08-16T12:06:13.145976Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"data = article_data['perceived_colour_value_name'].value_counts()\ndata = data.div(data.sum()).mul(100)\ndata","metadata":{"execution":{"iopub.status.busy":"2022-08-16T12:06:13.147290Z","iopub.status.idle":"2022-08-16T12:06:13.147723Z","shell.execute_reply.started":"2022-08-16T12:06:13.147522Z","shell.execute_reply":"2022-08-16T12:06:13.147545Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"interestingly, dusty is a the name used to represent certain colours, let us visualise these articles to get a better idea","metadata":{}},{"cell_type":"code","source":"## creating a small base visualisation function so as to help view images\n##without having to write the code each time.\n\ndef show_items_in_category(column, value, no_imgs=4, title=None):\n    data = article_data[article_data[column]==value]\n    cat_ids = data['article_id'].iloc[:no_imgs].to_list()\n    \n    fig, ax = plt.subplots(1, no_imgs, figsize=(12,4))\n\n    for i, prod_id in enumerate(cat_ids):\n        folder = str(prod_id)[:2]\n        file_path = '../input/h-and-m-personalized-fashion-recommendations/images/0{}/0{}.jpg'.format(folder, prod_id)\n\n        img = plt.imread(file_path)       \n        ax[i].imshow(img, aspect='equal')\n        ax[i].imshow(img, aspect='equal')\n        ax[i].grid(False)\n        ax[i].set_xticks([], [])\n        ax[i].set_yticks([], [])\n        ax[i].set_xlabel(article_data[article_data['article_id']==int(prod_id)]['prod_name'].iloc[0])\n    \n    fig.suptitle(title)\n    plt.show()\n\nshow_items_in_category('perceived_colour_value_name', 'Dusty Light', 5, 'Dusty Light articles')\nshow_items_in_category('perceived_colour_value_name', 'Medium Dusty', 5, 'Dusty Light articles')","metadata":{"execution":{"iopub.status.busy":"2022-08-16T12:06:13.149623Z","iopub.status.idle":"2022-08-16T12:06:13.150047Z","shell.execute_reply.started":"2022-08-16T12:06:13.149827Z","shell.execute_reply":"2022-08-16T12:06:13.149854Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"viewing these we can understand that dusty refers to the slightly 'sandy'/'sanded' textured look on clothes quite common in current fashion trends (also similar to the melange texture category)","metadata":{}},{"cell_type":"markdown","source":"## Viewing all the different types of patterns present","metadata":{}},{"cell_type":"code","source":"plot_bar(article_data,'graphical_appearance_name',figsize=(15,15),pct=True)","metadata":{"execution":{"iopub.status.busy":"2022-08-16T12:06:13.151603Z","iopub.status.idle":"2022-08-16T12:06:13.152045Z","shell.execute_reply.started":"2022-08-16T12:06:13.151846Z","shell.execute_reply":"2022-08-16T12:06:13.151867Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"data = article_data['graphical_appearance_name'].value_counts()\ndata = data.div(data.sum()).mul(100)\ndata","metadata":{"execution":{"iopub.status.busy":"2022-08-16T12:06:13.153224Z","iopub.status.idle":"2022-08-16T12:06:13.153632Z","shell.execute_reply.started":"2022-08-16T12:06:13.153407Z","shell.execute_reply":"2022-08-16T12:06:13.153424Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"we can note that the patterns are stored as graphical apperance and its dominated by solid patterns as epxected, then something that is vague called all over pattern (probably universal patterns across the clothing item that are too niche to be categorized by unique names eg polka dots)\n\nThe least common articles seem to be holograms and argyle cotton clothing style","metadata":{}},{"cell_type":"markdown","source":"## Viewing some of the top patterns and some of the more unusual patterns to get a better idea of them","metadata":{}},{"cell_type":"code","source":"show_items_in_category('graphical_appearance_name', 'All over pattern', 5,  'All over pattern')\nshow_items_in_category('graphical_appearance_name', 'Melange', 5,  'Melange')\nshow_items_in_category('graphical_appearance_name', 'Placement print', 5,  'Placement print')\nshow_items_in_category('graphical_appearance_name', 'Other structure', 5,  'Other structure')\nshow_items_in_category('graphical_appearance_name', 'Transparent', 5,  'Transparent')\nshow_items_in_category('graphical_appearance_name', 'Mesh', 5,  'Mesh')\nshow_items_in_category('graphical_appearance_name', 'Unknown', 5,  'Unknown')\nshow_items_in_category('graphical_appearance_name', 'Hologram', 5,  'Hologram')","metadata":{"execution":{"iopub.status.busy":"2022-08-16T12:06:13.154820Z","iopub.status.idle":"2022-08-16T12:06:13.155185Z","shell.execute_reply.started":"2022-08-16T12:06:13.155008Z","shell.execute_reply":"2022-08-16T12:06:13.155025Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Lets View and visualise some of the product descriptions (visualising using a wordcloud)","metadata":{}},{"cell_type":"code","source":"article_data['detail_desc'].drop_duplicates().to_list()[0:3]","metadata":{"execution":{"iopub.status.busy":"2022-08-16T12:06:13.156343Z","iopub.status.idle":"2022-08-16T12:06:13.156739Z","shell.execute_reply.started":"2022-08-16T12:06:13.156551Z","shell.execute_reply":"2022-08-16T12:06:13.156570Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import nltk\nfrom nltk.corpus import stopwords\nfrom wordcloud import WordCloud, STOPWORDS # library to create a wordcloud\nfrom PIL import Image\n\n'''\nThe goal here is to use the NLTK library's wordcloud\nfeature and its stopwords feature to generate a wordcloud\nfree of commonly used words.\n\nwe're first dropping all the NA values and splitting\nthe descriptions into a bag of words so as to get the\nmost repetative words for the wordcloud's size element \n(all this is handled by the NLTK wordcloud function) and \nusing the stopwords list in NLTK to remove all the common \nstopwords used in description to prevent them from dominating\nthe wordcloud\n\nsome data maybe lost due to the stopwords including some common\ndescriptors, but it still ensures a good enough idea of the common \nwords used to describe the clothing present at H&M\n'''\n\n# creating cloud of words\nwords_raw = article_data['detail_desc'].dropna().apply(nltk.word_tokenize)\nbag_of_words = \" \".join(words_raw.explode())\nstopwords = set(STOPWORDS)\n\n# creating cloud of words\nfig, ax1 = plt.subplots(figsize=(10,10))\n#ensuring we remove all the stopwords before building\n#the wordcloud\nwordcloud = WordCloud(stopwords=stopwords, background_color=\"white\", height=300, contour_width=3).generate(bag_of_words)\nplt.imshow(wordcloud)\nplt.axis(\"off\")\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-08-16T12:06:13.157972Z","iopub.status.idle":"2022-08-16T12:06:13.158329Z","shell.execute_reply.started":"2022-08-16T12:06:13.158152Z","shell.execute_reply":"2022-08-16T12:06:13.158169Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Transactions data exploration","metadata":{}},{"cell_type":"markdown","source":"This dataset contains the transaction history for the customers, and would be the the dataset to use for predicting users future sales behaviour. the dataset interestingly contains a parameter for sales channel, therefore the data contained here is not restricted to their online accounts alone. Further the article price paid is included as well. Upon checking with the competition hosts H&M, it is noted that the transactions here contain all transaction data, and does not exclude returns data. Therefore predicitons maybe skewed on the basis of returns. For example certain users may purchase the same item in multiple sizes and therefore bias their buying patterns towards those items.\n\n#### Columns description:\n\n1) t_dat - date of a transaction in format YYYY-MM-DD but provided as a string\n\n2) customer_id - identifier of the customer which can be mapped to the customer_id column in the customers table\n\n3) article_id - identifier of the product which can be mapped to the article_id column in the articles table\n\n4) price - price paid\n\n5) sales_channel_id - sales channel, 2 unique values\n","metadata":{}},{"cell_type":"code","source":"transactions_data.head()","metadata":{"execution":{"iopub.status.busy":"2022-08-16T12:06:13.159707Z","iopub.status.idle":"2022-08-16T12:06:13.160091Z","shell.execute_reply.started":"2022-08-16T12:06:13.159906Z","shell.execute_reply":"2022-08-16T12:06:13.159924Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"transactions_data.isnull().sum()","metadata":{"execution":{"iopub.status.busy":"2022-08-16T12:06:13.161944Z","iopub.status.idle":"2022-08-16T12:06:13.162428Z","shell.execute_reply.started":"2022-08-16T12:06:13.162190Z","shell.execute_reply":"2022-08-16T12:06:13.162209Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Let us explore the transaction data period, so as to get a better idea of if the data is of a specific time period or specific to seasons alone","metadata":{}},{"cell_type":"code","source":"#converting date data to pandas datetime format\ntransactions_data['t_dat']=pd.to_datetime(transactions_data['t_dat'])\n#getting the date range\n\nstart=transactions_data['t_dat'].min()\nend=transactions_data['t_dat'].max()\n\nprint('The date ranges for the dataset is from {} to {}'.format(start.date(),end.date()))","metadata":{"execution":{"iopub.status.busy":"2022-08-16T12:06:13.163919Z","iopub.status.idle":"2022-08-16T12:06:13.164323Z","shell.execute_reply.started":"2022-08-16T12:06:13.164124Z","shell.execute_reply":"2022-08-16T12:06:13.164144Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"note that we have complete 2 years data from 2018 to 2020. Interesting fact about the time period is that 1 year of the transaction data is during the covid-19 pandemic, and maybe highly skewed to both online retail and potentially lower sales? (pandemic did reduce sales for certain outlets)\n","metadata":{}},{"cell_type":"markdown","source":"lets group the data by dates so as to see transaction performance with time on a plot","metadata":{}},{"cell_type":"code","source":"transactions_per_day=transactions_data.groupby('t_dat', as_index=False).count()\n\n#creating a time series plot with an dotted line for new years and \n#red markers for max and min sales respectively\n\nfig, ax = plt.subplots(figsize=(16,8))\n\nsns.lineplot(data=transactions_per_day, x='t_dat',y='customer_id')\n\nax.set_xlabel('date')\nax.set_ylabel('number of transactions')\n\nax.axvline(x=dt.datetime(2019,1,1), c='green')\nax.axvline(x=dt.datetime(2020,1,1), c='green')\n\nmax_t = transactions_per_day['customer_id'].max()\nmax_t_date = transactions_per_day[transactions_per_day['customer_id']==max_t]['t_dat']\nax.scatter(max_t_date, max_t, c='red')\nax.text(max_t_date+pd.DateOffset(days=5), max_t-4000, '{}\\n{:,d}'.format(max_t_date.iloc[0].date(), max_t))\n\nmin_t = transactions_per_day['customer_id'].min()\nmin_t_date = transactions_per_day[transactions_per_day['customer_id']==min_t]['t_dat']\nax.scatter(min_t_date, min_t, c='red')\nax.text(min_t_date+pd.DateOffset(days=5), min_t-4000, '{}\\n{:,d}'.format(min_t_date.iloc[0].date(), min_t))\nax.set_xlim(transactions_data['t_dat'].min(),transactions_data['t_dat'].max())\n\nfor loc in ['bottom', 'left']:\n    ax.spines[loc].set_visible(True)\n    ax.spines[loc].set_linewidth(2)\n    ax.spines[loc].set_color('black')\n\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-08-16T12:06:13.166185Z","iopub.status.idle":"2022-08-16T12:06:13.166608Z","shell.execute_reply.started":"2022-08-16T12:06:13.166374Z","shell.execute_reply":"2022-08-16T12:06:13.166392Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"we can note that there seem to be clear seasonal spikes in sales in the data, and this can be better visualised using a barchart or a box plot instead.","metadata":{}},{"cell_type":"code","source":"trans_gr_month = transactions_data.groupby('t_dat').size().rename(\"no_transactions\")\ntrans_gr_month = trans_gr_month.reset_index()\ntrans_gr_month['month_year'] = trans_gr_month['t_dat'].dt.to_period('M')","metadata":{"execution":{"iopub.status.busy":"2022-08-16T12:06:13.168012Z","iopub.status.idle":"2022-08-16T12:06:13.168409Z","shell.execute_reply.started":"2022-08-16T12:06:13.168216Z","shell.execute_reply":"2022-08-16T12:06:13.168235Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fig, ax = plt.subplots(figsize=(16,8))\nax = sns.boxplot(x=\"month_year\", y='no_transactions', data=trans_gr_month)\nplt.xticks(rotation=90)\nfor loc in ['bottom', 'left']:\n    ax.spines[loc].set_visible(True)\n    ax.spines[loc].set_linewidth(2)\n    ax.spines[loc].set_color('black')\nax.set_xlabel('Month-Year')\nax.set_ylabel('Number of transactions')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-08-16T12:06:13.169842Z","iopub.status.idle":"2022-08-16T12:06:13.170233Z","shell.execute_reply.started":"2022-08-16T12:06:13.170043Z","shell.execute_reply":"2022-08-16T12:06:13.170061Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"the box plot above shows that there is a consistent range for daily transactions in a month, with sales spikes occuring during summer time and interestingly sales dropping during winter (not enough christmas sales?)","metadata":{}},{"cell_type":"markdown","source":"let us understand how many transactions an average customer does as well","metadata":{}},{"cell_type":"code","source":"t_by_customer = transactions_data.groupby('customer_id', as_index=False).size()\n\nfig, ax = plt.subplots(figsize=(10,5))\nax = sns.histplot(data=t_by_customer, x='size', bins=50, stat=\"percent\")\nax.set_xlabel('Distribution of total transactions per customer')\nfor loc in ['bottom', 'left']:\n    ax.spines[loc].set_visible(True)\n    ax.spines[loc].set_linewidth(2)\n    ax.spines[loc].set_color('black')\nax.yaxis.set_major_formatter(mtick.PercentFormatter())\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-08-16T12:06:13.171463Z","iopub.status.idle":"2022-08-16T12:06:13.171896Z","shell.execute_reply.started":"2022-08-16T12:06:13.171690Z","shell.execute_reply":"2022-08-16T12:06:13.171709Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"notably, there is a large number of outliers in the data and therefore we'd require to threshold the data to explore the typical trends in the customer data instead","metadata":{}},{"cell_type":"code","source":"t_by_customer_50tr = t_by_customer[t_by_customer['size'] < 50]\n\nfig, ax = plt.subplots(figsize=(10,5))\nax = sns.histplot(data=t_by_customer_50tr, x='size', bins=50, stat=\"percent\")\nax.set_xlabel('Distribution of total transactions per customer')\nfor loc in ['bottom', 'left']:\n    ax.spines[loc].set_visible(True)\n    ax.spines[loc].set_linewidth(2)\n    ax.spines[loc].set_color('black')\n\n    ax.yaxis.set_major_formatter(mtick.PercentFormatter())\nplt.show()\n","metadata":{"execution":{"iopub.status.busy":"2022-08-16T12:06:13.173218Z","iopub.status.idle":"2022-08-16T12:06:13.173636Z","shell.execute_reply.started":"2022-08-16T12:06:13.173407Z","shell.execute_reply":"2022-08-16T12:06:13.173424Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sns.set_style(\"darkgrid\", {\"axes.facecolor\": \".9\"})\nfig, ax = plt.subplots(figsize=(10,5))\nax = sns.histplot(data=transactions_data, x='price', bins=50, stat=\"percent\")\nax.set_xlabel('Distribution of the price')\nfor loc in ['bottom', 'left']:\n    ax.spines[loc].set_visible(True)\n    ax.spines[loc].set_linewidth(2)\n    ax.spines[loc].set_color('black')\nax.yaxis.set_major_formatter(mtick.PercentFormatter())\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-08-16T12:06:13.174941Z","iopub.status.idle":"2022-08-16T12:06:13.175556Z","shell.execute_reply.started":"2022-08-16T12:06:13.175329Z","shell.execute_reply":"2022-08-16T12:06:13.175354Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Note that thehistogram seems to stretch the x axis to a large extent due to a very large percentage of outliers in the data which we can attempt to visualise using a scatterplot ","metadata":{}},{"cell_type":"code","source":"sns.boxplot(transactions_data['price'])","metadata":{"execution":{"iopub.status.busy":"2022-08-16T12:06:13.177082Z","iopub.status.idle":"2022-08-16T12:06:13.177436Z","shell.execute_reply.started":"2022-08-16T12:06:13.177259Z","shell.execute_reply":"2022-08-16T12:06:13.177276Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Since there are so many outliers in the data beyond 0.1, let us visualise the majority of the data without outliers to view the trend for the data","metadata":{}},{"cell_type":"code","source":"fig, ax = plt.subplots(figsize=(10,5), facecolor='#f6f5f5')\noutlier_free_data = transactions_data[transactions_data['price']<0.1]\nax = sns.histplot(data=outlier_free_data, x='price', bins=20, stat=\"percent\")\nax.set_xlabel('Distribution of the price')\nfor loc in ['bottom', 'left']:\n    ax.spines[loc].set_visible(True)\n    ax.spines[loc].set_linewidth(2)\n    ax.spines[loc].set_color('black')\nax.yaxis.set_major_formatter(mtick.PercentFormatter())\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-08-16T12:06:13.178983Z","iopub.status.idle":"2022-08-16T12:06:13.179381Z","shell.execute_reply.started":"2022-08-16T12:06:13.179190Z","shell.execute_reply":"2022-08-16T12:06:13.179209Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fig, ax = plt.subplots(figsize=(5,5))\nexplode = (0, 0.1)\ncolors = sns.color_palette('Paired')\nax.pie(transactions_data['sales_channel_id'].value_counts(), explode=explode, labels=['1','2'],\n       autopct='%1.1f%%',shadow=True, startangle=90, colors=colors)\nax.axis('equal')\nax.set_title('Sale channel')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-08-16T12:06:13.180968Z","iopub.status.idle":"2022-08-16T12:06:13.181367Z","shell.execute_reply.started":"2022-08-16T12:06:13.181169Z","shell.execute_reply":"2022-08-16T12:06:13.181188Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"note that 1 here marks online sales and 2 here marks offline sales. most of the sales 70% are online and almost 30% is offline.","metadata":{}},{"cell_type":"markdown","source":"## Creating a weekly transaction animation to understand sales performance","metadata":{}},{"cell_type":"code","source":"#combinding article data and transaction data to create a new combined dataset\narticles = article_data[['article_id', 'section_name']]\ntransactions = transactions_data[['t_dat', 'article_id']]\ntransactions = transactions.merge(articles, on='article_id')\ngc.collect()","metadata":{"execution":{"iopub.status.busy":"2022-08-16T12:06:13.182419Z","iopub.status.idle":"2022-08-16T12:06:13.182880Z","shell.execute_reply.started":"2022-08-16T12:06:13.182684Z","shell.execute_reply":"2022-08-16T12:06:13.182704Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#converting the time to year,week and year of the week notations\ntransactions['t_dat'] = pd.to_datetime(transactions['t_dat'])\ntransactions['year'] = transactions['t_dat'].dt.isocalendar().year\ntransactions['week'] = transactions['t_dat'].dt.isocalendar().week\ntransactions['year_week'] = transactions['year'].astype(str) + '-' + transactions['week'].astype(str)\ngc.collect()","metadata":{"execution":{"iopub.status.busy":"2022-08-16T12:06:13.184038Z","iopub.status.idle":"2022-08-16T12:06:13.184397Z","shell.execute_reply.started":"2022-08-16T12:06:13.184215Z","shell.execute_reply":"2022-08-16T12:06:13.184231Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#creating the graph with an animation for weekly sales\ncounted = transactions.groupby(['year_week', 'section_name'])['article_id'].count()\n\ncounted_df = counted.reset_index()\ncounted_df = counted_df.rename(columns={'article_id':'count'})\n\norder = counted_df.section_name.unique().tolist()\nfig = px.bar(counted_df, x='section_name', y='count',\n             animation_frame='year_week', animation_group='section_name',\n             range_y=[0, 110000], \n             template='simple_white', title='Weekly Sales (Units) by Section')\nfig.update_xaxes(categoryorder='array', categoryarray=order)\nfig['layout']['updatemenus'][0]['pad']=dict(r= 10, t= 240)\nfig['layout']['sliders'][0]['pad']=dict(r= 10, t= 220)\nfig.show(renderer='notebook_connected')","metadata":{"execution":{"iopub.status.busy":"2022-08-16T12:06:13.185805Z","iopub.status.idle":"2022-08-16T12:06:13.186204Z","shell.execute_reply.started":"2022-08-16T12:06:13.186011Z","shell.execute_reply":"2022-08-16T12:06:13.186030Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Interestingly, the weekly sales do not fluctuate significantly and the categories especially do not change much by far. seems like the sales grow exponentially and later shrinks away in slower decrements.\n\nWhat is observable can be linked to the fact that fast fashion trends change weekly, bi-weekly and new catalogues are launched every week or every 2 weeks. Therefore we can see a proportional sales growth as well.\n\n#### this leaves me with the reasonable explanation that one of the most important factors for article sales would be the articles sold last week itself.~","metadata":{}},{"cell_type":"markdown","source":"# Notes from Exploring the Article Images provided by H&M (no code here, manually explored)","metadata":{}},{"cell_type":"markdown","source":"Before using the images for any sort of pre processing or any sort of feature vector/vector embeddings being created from the images, I've explored through the images to understand any potential challenges faced\n\nSummary of my thoughts are as follows :\n\n* **Image quality :** Generally most of the images are extremely clean, with the articles themselves being clearly displayed over a light/white background, shot using professional setups\n\n* **File/Folder Structure :** Images folder ranges from 010 to 095 and don't seem to be structured in any particular manner, and articles shuffle across folders chaotically, not grouped by article type, age,season, colour or any categories (look like a dump of images with their corresponding article name, as even the number of images vary based on folders)\n\n* **Image resolutions :** Image resolutions seem to vary from one to another, as well as the size of the article itself in the image (i.e image to background ratio) also seems to vary from one to another. This would necessitate the standardisation of the image's size upon loading through image matrix manipulation (adjusting resolution). This may have some impact on image quality when processing.\n\n* **Article bundles :** Its noted that some articles are being sold as bundled groups and would result in a potential challenge if colour is used as a potential extracted feature)\n\n* **Close up shots (indistinguishable images):** There are a lot of photos that are taken from very much upclose to represent the texture of the article, but differ a lot from the wider shot images resulting in hard to identify the type of product (would stump some image based models), but we can combine them with the NLP as descriptions are mostly available.\n\n* **Clarity of a small percentage of images:** similar to the closeup shots issue, some images are very hard to understand what article they are by looking at them, and would only be saved by using NLP on their descriptions.\n\n* **Seasonal items:** interestingly there are seasonal items such as christmas and halloween items and maybe image recogntion can be used to detect these articles and recommend them during their typical sales period (october to december)","metadata":{}},{"cell_type":"markdown","source":"# Core thoughts on the data:\n\n* The trend in sales seems to be not as complicated or driven by a more complex underlying trend in itself, but rather a weekly sales high in same high volume product categories (could be from weekend shopping or from new releases as fast fashion releases can be nearly weekly in nature) with a more general seasonal trend to it (higher sales in summer with summer clothing being more varied and also cheaper than winter jackets and other types of clothing).\n\n* Seems like the best predictor for a user would be to predict the top selling articles using a rolling window method due to the same.\n\n* Gender age and other customer characteristics would determine volume of sales and recommendations as there is a huge divide in the number of articles in each category\n\n* The highest volume of sales by age group is in the expected regions for a fast fashion trendy company like H&M for example with young spending populations and the older parents buying for their kids population.\n\n* there is a huge divide between online and offline sales, notably 2/3rds of the sales volume is online. Also it can be considered that online sales in general is more predictable and easier to recommend for as offline stores have additional variables such as current stock availability, size availability and therefore sometimes users tend to buy the next best thing. The accuracy for such a system that attempts to predict user purchasing behaviour both online and offline at once would suffer in terms of its accuracy for the same reasons\n\n* There are some articles with image descriptions and images themselves missing, and the data in general does contain missing values and NaN values which would have to be dealt with before modelling the same.\n\n* The price column in the transactions data seem to be normalized overall, and that has resulted in some interesting pricing values, as the lowest priced items are getting rounded down to 0.00 while not being the case. Rectifying this would be extremely difficult/impractical and this would have to be considered as a factor. \n\n* Further, H&M's pricing in itself has a unique quirk, as the price range seems to have a major amount of outliers as the typical price is much lower than some of the highest priced articles (typically leather items)\n\n* if we look at the customer trends statistic, there seems to be an anomolous postcode with many many sales which does not seem to add up but seem to be all from legitimate customers with varying customer data. This may skew the data heavily towards the preferences of that postcode in itself, not to mention the whole thing in itself is a unique oddity as how can one postcode have such a significantly higher volume of sales?\n\n* Further note on customer statistics, we can see that most of the customers lie in the lower volume of purchases bracket, and that there is a very clear negative exponential trend in sales volume per customer and yet, there seem to be a large number of outliers which have extremely high sales numbers (potential wholesalers)\n\n* the articles themeselves are very well categorized and sub categorized through multilevel indexing for both product type and colour types, as well as being well described. Certain items may lack images or descriptions but in general the data is quite clean in regards to the same. \n\n* there is potential to use a word embedding detail based predictor model as well as an image based recommendation predictor model, a combination of the two as well as a model that would use the customer data, the transaction data as well as the article categories data to generate predictions. A combination of the embeddings and images together with the tabular data may have the best chance in avoiding the cold start problem.\n\n* additional features may have to be created in terms of weekly best selling items, monthly best selling items and other rolling windows based on the time series based trends we noted in the data regarding article sales, as predicting the most recently highly sold article is tending to have a high liklihood of being a good prediction.","metadata":{}}]}