{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# H&M EDA and Baseline (WIP)\n\nThis notebook is a quick EDA and Baseline for the new H&M Personalized Fashion Recommendations competetion. If you find the notebook helpful please give an upvote :)\n\n#### forked - added time filter and changed settings + get output + will work for anyone now without private data - Dan\n\n\n\n### Contents:\n[Load in the data ⏳](#first-bullet)\n    \n[Articles EDA 📚](#second-bullet)   \n  \n[Customers EDA 🛍](#third-bullet)\n    \n[Transaction EDA 💸](#fourth-bullet)\n    \n[Imagery 📸](#fith-bullet)\n   \n[Baseline 📈](#sixth-bullet)","metadata":{}},{"cell_type":"markdown","source":"### Imports","metadata":{}},{"cell_type":"code","source":"import pandas as pd\nfrom pathlib import Path\nimport matplotlib.pyplot as plt\nimport seaborn as sns\nimport datetime as dt\nfrom termcolor import colored\nfrom PIL import Image\nimport os\nimport random\n\nimport warnings\nwarnings.filterwarnings('ignore')","metadata":{"execution":{"iopub.status.busy":"2022-02-09T09:48:12.314448Z","iopub.execute_input":"2022-02-09T09:48:12.315443Z","iopub.status.idle":"2022-02-09T09:48:13.654389Z","shell.execute_reply.started":"2022-02-09T09:48:12.315284Z","shell.execute_reply":"2022-02-09T09:48:13.653229Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Load in the data ⏳","metadata":{}},{"cell_type":"code","source":"DATA_PATH = Path('../input/h-and-m-personalized-fashion-recommendations')\n!ls $DATA_PATH","metadata":{"execution":{"iopub.status.busy":"2022-02-09T09:48:13.656198Z","iopub.execute_input":"2022-02-09T09:48:13.656475Z","iopub.status.idle":"2022-02-09T09:48:14.473752Z","shell.execute_reply.started":"2022-02-09T09:48:13.656446Z","shell.execute_reply":"2022-02-09T09:48:14.472243Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Data Overview\n- `images/` - a folder of images corresponding to each `article_id`; images are placed in subfolders starting with the first three digits of the `article_id`; note, not all `article_id` values have a corresponding image.\n- `articles.csv` - detailed metadata for each article_id available for purchase\n- `customers.csv` - metadata for each `customer_id` in dataset\n- `sample_submission.csv` - a sample submission file in the correct format\n- `transactions_train.csv` - the training data, consisting of the purchases each customer for each date, as well as additional information. Duplicate rows correspond to multiple purchases of the same item. Your task is to predict the article_ids each customer will purchase during the 7-day period immediately after the training data period.","metadata":{}},{"cell_type":"code","source":"articles = pd.read_csv(DATA_PATH/'articles.csv')\ncustomers = pd.read_csv(DATA_PATH/'customers.csv')\ntransactions_train = pd.read_csv(DATA_PATH/'transactions_train.csv')\nsamp_sub = pd.read_csv(DATA_PATH/'sample_submission.csv')","metadata":{"execution":{"iopub.status.busy":"2022-02-09T09:48:14.476656Z","iopub.execute_input":"2022-02-09T09:48:14.477048Z","iopub.status.idle":"2022-02-09T09:49:44.374405Z","shell.execute_reply.started":"2022-02-09T09:48:14.477004Z","shell.execute_reply":"2022-02-09T09:49:44.373258Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"From: https://www.kaggle.com/hengzheng/time-is-our-best-friend","metadata":{}},{"cell_type":"code","source":"### From: https://www.kaggle.com/hengzheng/time-is-our-best-friend\n# only use data after 2020-06-01\n\n# transactions['t_dat'] = pd.to_datetime(transactions['t_dat'])\n# transactions = transactions[transactions['t_dat'] > pd.to_datetime('2020-06-01')]\n\n\ntransactions_train['t_dat'] = pd.to_datetime(transactions_train['t_dat'])\nt_cut = pd.to_datetime('2020-06-01')\ntransactions_train = transactions_train.loc[transactions_train['t_dat'] > t_cut]","metadata":{"execution":{"iopub.status.busy":"2022-02-09T09:49:44.375976Z","iopub.execute_input":"2022-02-09T09:49:44.376265Z","iopub.status.idle":"2022-02-09T09:49:52.271370Z","shell.execute_reply.started":"2022-02-09T09:49:44.376229Z","shell.execute_reply":"2022-02-09T09:49:52.270249Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Articles EDA 📚","metadata":{}},{"cell_type":"code","source":"articles.head()","metadata":{"execution":{"iopub.status.busy":"2022-02-09T09:49:52.273875Z","iopub.execute_input":"2022-02-09T09:49:52.274122Z","iopub.status.idle":"2022-02-09T09:49:52.320367Z","shell.execute_reply.started":"2022-02-09T09:49:52.274092Z","shell.execute_reply":"2022-02-09T09:49:52.319277Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"num_articles = len(articles)\nnum_unique_id = len(articles['article_id'].unique())\nprint(f'We have {num_articles} rows in the df and {num_unique_id} unique article IDs') ","metadata":{"execution":{"iopub.status.busy":"2022-02-09T09:49:52.321700Z","iopub.execute_input":"2022-02-09T09:49:52.321961Z","iopub.status.idle":"2022-02-09T09:49:52.337093Z","shell.execute_reply.started":"2022-02-09T09:49:52.321926Z","shell.execute_reply":"2022-02-09T09:49:52.336311Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"num_prod_codes = len(articles['product_code'].unique())\nprint(f'Each article has a product_code, some articles have the same product_code with a total of {num_prod_codes} unique values')","metadata":{"execution":{"iopub.status.busy":"2022-02-09T09:49:52.338832Z","iopub.execute_input":"2022-02-09T09:49:52.339116Z","iopub.status.idle":"2022-02-09T09:49:52.355870Z","shell.execute_reply.started":"2022-02-09T09:49:52.339076Z","shell.execute_reply":"2022-02-09T09:49:52.354799Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"num_prod_name = len(articles['prod_name'].unique())\nprint(f'Each article also has a prod_name, with a total of {num_prod_name} unique values')","metadata":{"execution":{"iopub.status.busy":"2022-02-09T09:49:52.358342Z","iopub.execute_input":"2022-02-09T09:49:52.359523Z","iopub.status.idle":"2022-02-09T09:49:52.391469Z","shell.execute_reply.started":"2022-02-09T09:49:52.359458Z","shell.execute_reply":"2022-02-09T09:49:52.390463Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Interestingly there are a different number of unique `product_code` values and `prod_name` values meaning there isn't a 1 to 1 mapping betwen them..","metadata":{}},{"cell_type":"code","source":"num_prod_type_no = len(articles['product_type_no'].unique())\nnum_prod_type = len(articles['product_type_name'].unique())\nprint(f'We have {num_prod_type_no} unique product_type_no values and {num_prod_type} unique product_type_name values, with each number mapping to a name')","metadata":{"execution":{"iopub.status.busy":"2022-02-09T09:49:52.392603Z","iopub.execute_input":"2022-02-09T09:49:52.392853Z","iopub.status.idle":"2022-02-09T09:49:52.413842Z","shell.execute_reply.started":"2022-02-09T09:49:52.392821Z","shell.execute_reply":"2022-02-09T09:49:52.412158Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def plot_bar_chart(df, feature, x_lim):\n    feature_count  = df[feature].value_counts()\n    feature_count = feature_count[:x_lim,]\n    plt.figure(figsize=(30,10))\n    sns.barplot(feature_count.index, feature_count.values, alpha=0.7)\n    sns.set(font_scale = 2)\n    plt.title(f'Frequency of top {x_lim} {feature}', fontsize=30)\n    plt.ylabel('Count', fontsize=30)\n    plt.xlabel(feature.replace('_', ' '), fontsize=30)\n    sns.set(font_scale=1.2)\n    plt.show()","metadata":{"execution":{"iopub.status.busy":"2022-02-09T09:49:52.415168Z","iopub.execute_input":"2022-02-09T09:49:52.415720Z","iopub.status.idle":"2022-02-09T09:49:52.424276Z","shell.execute_reply.started":"2022-02-09T09:49:52.415680Z","shell.execute_reply":"2022-02-09T09:49:52.423058Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plot_bar_chart(articles, 'product_type_name', 15)","metadata":{"execution":{"iopub.status.busy":"2022-02-09T09:49:52.425553Z","iopub.execute_input":"2022-02-09T09:49:52.425865Z","iopub.status.idle":"2022-02-09T09:49:52.961120Z","shell.execute_reply.started":"2022-02-09T09:49:52.425829Z","shell.execute_reply":"2022-02-09T09:49:52.959854Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"num_prod_group = len(articles['product_group_name'].unique())\nprint(f'We have {num_prod_group} unique product_group_names values')","metadata":{"execution":{"iopub.status.busy":"2022-02-09T09:49:52.963239Z","iopub.execute_input":"2022-02-09T09:49:52.963909Z","iopub.status.idle":"2022-02-09T09:49:52.981504Z","shell.execute_reply.started":"2022-02-09T09:49:52.963857Z","shell.execute_reply":"2022-02-09T09:49:52.980457Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plot_bar_chart(articles, 'product_group_name', 10)","metadata":{"execution":{"iopub.status.busy":"2022-02-09T09:49:52.983453Z","iopub.execute_input":"2022-02-09T09:49:52.984606Z","iopub.status.idle":"2022-02-09T09:49:53.393683Z","shell.execute_reply.started":"2022-02-09T09:49:52.984549Z","shell.execute_reply":"2022-02-09T09:49:53.392503Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We have six columns related to colour:\n\n- `colour_group_code` \n- `colour_group_name`\n- `perceived_colour_value_id`\n- `perceived_colour_value_name`\n- `perceived_colour_master_id`\n- `perceived_colour_master_name`\n\nFor the sake of brevity we only plot `perceived_colour_master_name`","metadata":{"execution":{"iopub.status.busy":"2022-02-07T23:18:23.306751Z","iopub.execute_input":"2022-02-07T23:18:23.307105Z","iopub.status.idle":"2022-02-07T23:18:23.31299Z","shell.execute_reply.started":"2022-02-07T23:18:23.307066Z","shell.execute_reply":"2022-02-07T23:18:23.311692Z"}}},{"cell_type":"code","source":"plot_bar_chart(articles, 'perceived_colour_master_name', 18)","metadata":{"execution":{"iopub.status.busy":"2022-02-09T09:49:53.398391Z","iopub.execute_input":"2022-02-09T09:49:53.399302Z","iopub.status.idle":"2022-02-09T09:49:53.875000Z","shell.execute_reply.started":"2022-02-09T09:49:53.399181Z","shell.execute_reply":"2022-02-09T09:49:53.873780Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"articles['garment_group_name'].value_counts()","metadata":{"execution":{"iopub.status.busy":"2022-02-09T09:49:53.876474Z","iopub.execute_input":"2022-02-09T09:49:53.876772Z","iopub.status.idle":"2022-02-09T09:49:53.907935Z","shell.execute_reply.started":"2022-02-09T09:49:53.876737Z","shell.execute_reply":"2022-02-09T09:49:53.906939Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We then have the following peices of meta data along with their codes:\n- `department_name` e.g 'Kids Girl Swimwear'\n- `index_name` e.g 'Children Accessories, Swimwear'\n- `index_group_name` e.g 'Baby/Children'\n- `section_name` 'Baby Essentials & Complements'\n- `garment_group_name` e.g. 'Swimwear'","metadata":{}},{"cell_type":"code","source":"display(articles['department_name'].value_counts().head(10))","metadata":{"execution":{"iopub.status.busy":"2022-02-09T09:49:53.909014Z","iopub.execute_input":"2022-02-09T09:49:53.909429Z","iopub.status.idle":"2022-02-09T09:49:53.939107Z","shell.execute_reply.started":"2022-02-09T09:49:53.909389Z","shell.execute_reply":"2022-02-09T09:49:53.937920Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"display(articles['index_name'].value_counts().head(10))","metadata":{"execution":{"iopub.status.busy":"2022-02-09T09:49:53.940568Z","iopub.execute_input":"2022-02-09T09:49:53.940859Z","iopub.status.idle":"2022-02-09T09:49:53.969764Z","shell.execute_reply.started":"2022-02-09T09:49:53.940825Z","shell.execute_reply":"2022-02-09T09:49:53.968377Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"display(articles['index_group_name'].value_counts().head(10))","metadata":{"execution":{"iopub.status.busy":"2022-02-09T09:49:53.971998Z","iopub.execute_input":"2022-02-09T09:49:53.972403Z","iopub.status.idle":"2022-02-09T09:49:54.008042Z","shell.execute_reply.started":"2022-02-09T09:49:53.972355Z","shell.execute_reply":"2022-02-09T09:49:54.007326Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"display(articles['section_name'].value_counts().head(10))","metadata":{"execution":{"iopub.status.busy":"2022-02-09T09:49:54.009740Z","iopub.execute_input":"2022-02-09T09:49:54.010003Z","iopub.status.idle":"2022-02-09T09:49:54.048482Z","shell.execute_reply.started":"2022-02-09T09:49:54.009963Z","shell.execute_reply":"2022-02-09T09:49:54.047215Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"display(articles['garment_group_name'].value_counts().head(10))","metadata":{"execution":{"iopub.status.busy":"2022-02-09T09:49:54.049968Z","iopub.execute_input":"2022-02-09T09:49:54.050261Z","iopub.status.idle":"2022-02-09T09:49:54.087364Z","shell.execute_reply.started":"2022-02-09T09:49:54.050188Z","shell.execute_reply":"2022-02-09T09:49:54.085929Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print('Each articles has a detailed description:\\n')\nfor i, (index, row) in enumerate(articles.sample(5).iterrows()):\n    description = row['detail_desc']\n    print(f'{i+1}. {description} \\n')","metadata":{"execution":{"iopub.status.busy":"2022-02-09T09:49:54.089237Z","iopub.execute_input":"2022-02-09T09:49:54.089625Z","iopub.status.idle":"2022-02-09T09:49:54.104579Z","shell.execute_reply.started":"2022-02-09T09:49:54.089577Z","shell.execute_reply":"2022-02-09T09:49:54.103703Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Customers EDA 🛍","metadata":{}},{"cell_type":"code","source":"customers.head()","metadata":{"execution":{"iopub.status.busy":"2022-02-09T09:49:54.106134Z","iopub.execute_input":"2022-02-09T09:49:54.106435Z","iopub.status.idle":"2022-02-09T09:49:54.125996Z","shell.execute_reply.started":"2022-02-09T09:49:54.106401Z","shell.execute_reply":"2022-02-09T09:49:54.124531Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"num_customers = len(customers)\nnum_customer_id = len(customers['customer_id'].unique())\nprint(f'We have {num_customers} rows and {num_customer_id} unique customer_ids')","metadata":{"execution":{"iopub.status.busy":"2022-02-09T09:49:54.127415Z","iopub.execute_input":"2022-02-09T09:49:54.128036Z","iopub.status.idle":"2022-02-09T09:49:54.818530Z","shell.execute_reply.started":"2022-02-09T09:49:54.127990Z","shell.execute_reply":"2022-02-09T09:49:54.817437Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"customers['club_member_status'].value_counts()","metadata":{"execution":{"iopub.status.busy":"2022-02-09T09:49:54.820115Z","iopub.execute_input":"2022-02-09T09:49:54.820447Z","iopub.status.idle":"2022-02-09T09:49:55.048067Z","shell.execute_reply.started":"2022-02-09T09:49:54.820414Z","shell.execute_reply":"2022-02-09T09:49:55.047030Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"customers['fashion_news_frequency'].value_counts()","metadata":{"execution":{"iopub.status.busy":"2022-02-09T09:49:55.049458Z","iopub.execute_input":"2022-02-09T09:49:55.049761Z","iopub.status.idle":"2022-02-09T09:49:55.295814Z","shell.execute_reply.started":"2022-02-09T09:49:55.049727Z","shell.execute_reply":"2022-02-09T09:49:55.295001Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"colors = sns.color_palette('pastel')[0:5]\nfig, ax = plt.subplots(1,2, figsize=(15, 6))\nfor i, feature in enumerate(['club_member_status', 'fashion_news_frequency']):\n    fashion_news = customers[feature].value_counts()\n    data = fashion_news.to_list()\n    labels = fashion_news.index.to_list()\n    ax[i].pie(data, labels = labels, colors = colors, autopct='%.0f%%')\n    ax[i].set_title(feature)\nfig.show()","metadata":{"execution":{"iopub.status.busy":"2022-02-09T09:49:55.297422Z","iopub.execute_input":"2022-02-09T09:49:55.297777Z","iopub.status.idle":"2022-02-09T09:49:56.041904Z","shell.execute_reply.started":"2022-02-09T09:49:55.297730Z","shell.execute_reply":"2022-02-09T09:49:56.040855Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plt.figure(figsize=(10,6))\np = sns.distplot(customers['age'], color=\"y\")\np.set_xlabel(\"Age\", fontsize = 20)\np.set_ylabel(\"Density\", fontsize = 20)\np.set_title(\"Age of customers\")\np.axvline(customers['age'].mean(), color='r', linestyle='--', label=\"Mean\")\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-02-09T09:49:56.043859Z","iopub.execute_input":"2022-02-09T09:49:56.045505Z","iopub.status.idle":"2022-02-09T09:50:01.845617Z","shell.execute_reply.started":"2022-02-09T09:49:56.045426Z","shell.execute_reply":"2022-02-09T09:50:01.844273Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Transaction EDA 💸","metadata":{}},{"cell_type":"code","source":"transactions_train.head(5)","metadata":{"execution":{"iopub.status.busy":"2022-02-09T09:50:01.846781Z","iopub.execute_input":"2022-02-09T09:50:01.847014Z","iopub.status.idle":"2022-02-09T09:50:01.864943Z","shell.execute_reply.started":"2022-02-09T09:50:01.846986Z","shell.execute_reply":"2022-02-09T09:50:01.863486Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"## add in for compatability - \ntransactions_train['t_dat'] = transactions_train['t_dat'].astype(str)","metadata":{"execution":{"iopub.status.busy":"2022-02-09T09:50:01.867080Z","iopub.execute_input":"2022-02-09T09:50:01.867462Z","iopub.status.idle":"2022-02-09T09:50:42.579992Z","shell.execute_reply.started":"2022-02-09T09:50:01.867424Z","shell.execute_reply":"2022-02-09T09:50:42.578648Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\ndates_list = [dt.datetime.strptime(date, '%Y-%m-%d').date() for date in transactions_train['t_dat'].to_list()]\ntransactions_train['new_date'] = dates_list\ntransactions_train['new_date'] = transactions_train['new_date'] - transactions_train['new_date'].min()\ntransactions_train[\"new_date\"] = transactions_train[\"new_date\"].apply(lambda x: x.days)","metadata":{"execution":{"iopub.status.busy":"2022-02-09T09:50:42.582111Z","iopub.execute_input":"2022-02-09T09:50:42.582456Z","iopub.status.idle":"2022-02-09T09:53:19.766921Z","shell.execute_reply.started":"2022-02-09T09:50:42.582419Z","shell.execute_reply":"2022-02-09T09:53:19.765988Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fig, axes = plt.subplots(1, 2, figsize=(15, 5))\nsns.distplot(transactions_train['new_date'], ax=axes[0])\naxes[0].set_xlabel('Day')\naxes[0].set_title('Date of sale')\n\nsns.distplot(transactions_train['price'], ax=axes[1], color='g')\naxes[1].set_title('Price')\naxes[1].set_xlim(0, 0.2)\nplt.show()","metadata":{"jupyter":{"source_hidden":true},"execution":{"iopub.status.busy":"2022-02-09T09:53:19.768769Z","iopub.execute_input":"2022-02-09T09:53:19.769100Z","iopub.status.idle":"2022-02-09T09:53:57.984860Z","shell.execute_reply.started":"2022-02-09T09:53:19.769056Z","shell.execute_reply":"2022-02-09T09:53:57.983677Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Imagery 📸","metadata":{}},{"cell_type":"markdown","source":"There is a folder of images corresponding to each `article_id`; images are placed in subfolders starting with the first three digits of the `article_id`; note, not all `article_id` values have a corresponding image.\n\n- The directory in which an image is stored is named 0 + the first two characters of the `article_id`\n- The image file name is then the whole `article_id` as a jpg file","metadata":{}},{"cell_type":"code","source":"IMAGE_PATH = Path('../input/h-and-m-personalized-fashion-recommendations/images')","metadata":{"execution":{"iopub.status.busy":"2022-02-09T09:53:57.986316Z","iopub.execute_input":"2022-02-09T09:53:57.986683Z","iopub.status.idle":"2022-02-09T09:53:57.991158Z","shell.execute_reply.started":"2022-02-09T09:53:57.986650Z","shell.execute_reply":"2022-02-09T09:53:57.990291Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def plot_imgs(ids, rows, cols):\n    figure, ax = plt.subplots(nrows=rows,ncols=cols,figsize=(16,8))\n    for ind, id_ in enumerate(ids):\n        fn = f'{IMAGE_PATH}/0{str(id_)[:2]}/0{id_}.jpg'\n        try:\n            img = Image.open(fn)\n        except:\n            pass\n        ax.ravel()[ind].imshow(img)\n        ax.ravel()[ind].set_axis_off()\n    plt.tight_layout()\n    plt.show()","metadata":{"execution":{"iopub.status.busy":"2022-02-09T09:53:57.992845Z","iopub.execute_input":"2022-02-09T09:53:57.993469Z","iopub.status.idle":"2022-02-09T09:53:58.009452Z","shell.execute_reply.started":"2022-02-09T09:53:57.993403Z","shell.execute_reply":"2022-02-09T09:53:58.008311Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Random sample of 5 images\nplot_imgs(articles.sample(5)['article_id'], 1, 5)","metadata":{"execution":{"iopub.status.busy":"2022-02-09T09:53:58.011181Z","iopub.execute_input":"2022-02-09T09:53:58.012310Z","iopub.status.idle":"2022-02-09T09:54:00.033000Z","shell.execute_reply.started":"2022-02-09T09:53:58.012247Z","shell.execute_reply":"2022-02-09T09:54:00.032115Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# 5 Sweaters\nplot_imgs(articles[articles['product_type_name']=='Sweater'].sample(5)['article_id'], 1, 5)","metadata":{"execution":{"iopub.status.busy":"2022-02-09T09:54:00.034676Z","iopub.execute_input":"2022-02-09T09:54:00.035748Z","iopub.status.idle":"2022-02-09T09:54:02.110636Z","shell.execute_reply.started":"2022-02-09T09:54:00.035642Z","shell.execute_reply":"2022-02-09T09:54:02.109339Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Baseline 📈","metadata":{}},{"cell_type":"markdown","source":"Submissions are evaluated according to the Mean Average Precision:\n\n## $\\frac{1}{U} \\sum_{u=1}^{U}  \\sum_{k=1}^{min(n,12)} P(k) \\times rel(k)$\n\nwhere 𝑈 is the number of customers, 𝑃(𝑘) is the precision at cutoff 𝑘, 𝑛 is the number predictions per image, and 𝑟𝑒𝑙(𝑘) is an indicator function equaling 1 if the item at rank 𝑘 is a relevant (correct) label, zero otherwise.\n\nNotes:\n\nYou will be making purchase predictions for all `customer_id` values provided, regardless of whether these customers made purchases in the training data.\nCustomer that did not make any purchase during test period are excluded from the scoring.\nThere is never a penalty for using the full 12 predictions for a customer that ordered fewer than 12 items; thus, it's advantageous to make 12 predictions for each customer.","metadata":{}},{"cell_type":"code","source":"# Due to low memory\n%reset -f","metadata":{"execution":{"iopub.status.busy":"2022-02-09T09:54:02.112081Z","iopub.execute_input":"2022-02-09T09:54:02.112439Z","iopub.status.idle":"2022-02-09T09:54:03.059651Z","shell.execute_reply.started":"2022-02-09T09:54:02.112401Z","shell.execute_reply":"2022-02-09T09:54:03.058971Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import pandas as pd\nfrom pathlib import Path\nfrom collections import Counter\nfrom itertools import chain, combinations\nimport random\nimport pprint\nfrom tqdm import tqdm\n\nimport warnings\nwarnings.filterwarnings('ignore')","metadata":{"execution":{"iopub.status.busy":"2022-02-09T09:54:03.060827Z","iopub.execute_input":"2022-02-09T09:54:03.061796Z","iopub.status.idle":"2022-02-09T09:54:03.067879Z","shell.execute_reply.started":"2022-02-09T09:54:03.061749Z","shell.execute_reply":"2022-02-09T09:54:03.066944Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"DATA_PATH = Path('../input/h-and-m-personalized-fashion-recommendations')\n\ndef add_value(dict_obj, key, value):\n    if key not in dict_obj:\n        dict_obj[key] = value\n    elif isinstance(dict_obj[key], list):\n        dict_obj[key].append(value)\n    else:\n        dict_obj[key] = [dict_obj[key], value]\n        \n# Messy code to get prediction into correct format ..\ndef format_prediction(pred, old_pred):\n    pred = [str(i) for i in pred[:12]]\n    old_list = old_pred.split()\n    del old_list[-len(pred):]\n    old_list = old_list + pred\n    pred =' '.join(old_list)\n    return pred","metadata":{"execution":{"iopub.status.busy":"2022-02-09T09:54:03.069434Z","iopub.execute_input":"2022-02-09T09:54:03.069946Z","iopub.status.idle":"2022-02-09T09:54:03.083231Z","shell.execute_reply.started":"2022-02-09T09:54:03.069897Z","shell.execute_reply":"2022-02-09T09:54:03.082268Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The following Baseline suggests articles to the customer which other customers with similar purchase history have bought. It also uses the most-common-benchmark to fill in some gaps..","metadata":{}},{"cell_type":"code","source":"# Load in the data\ntransactions_train = pd.read_csv(DATA_PATH/'transactions_train.csv',\n                                 dtype={'article_id': str},parse_dates=['t_dat'],infer_datetime_format=True\n                                ,usecols=['t_dat', 'customer_id', 'article_id'])\n### From: https://www.kaggle.com/hengzheng/time-is-our-best-friend\n# only use data after 2020-06-01\n\n# transactions['t_dat'] = pd.to_datetime(transactions['t_dat'])\n# transactions = transactions[transactions['t_dat'] > pd.to_datetime('2020-06-01')]\n\n\n# transactions_train['t_dat'] = pd.to_datetime(transactions_train['t_dat'])\nt_cut = pd.to_datetime('2020-02-01') ## '2020-06-01'\ntransactions_train = transactions_train.loc[(transactions_train['t_dat'] > t_cut) | (transactions_train['t_dat'].dt.month==9)]","metadata":{"execution":{"iopub.status.busy":"2022-02-09T09:57:09.399887Z","iopub.execute_input":"2022-02-09T09:57:09.400297Z","iopub.status.idle":"2022-02-09T09:58:01.265655Z","shell.execute_reply.started":"2022-02-09T09:57:09.400261Z","shell.execute_reply":"2022-02-09T09:58:01.264662Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\n# Select required columns\ntransactions_train = transactions_train[['customer_id', 'article_id']]\n\n# Groupby customer ID and collect all the articles purchased into a list\ntransactions_train = transactions_train.groupby('customer_id')['article_id'].apply(list)\n\n# Due to low memory select a subset of transactions -- to be improved ..\ntransactions_train = transactions_train[:180000] # 150000 - orig , trying more\n\n# Find all pairs of articles which are purchased by the same customer\ntransactions_train = Counter(chain.from_iterable(combinations(customer, 2) for customer in transactions_train.to_list()))\n\n# Remove pairs of articles which do not occur frequently together\nfrequent_pairs = {k: v for k, v in transactions_train.items() if v > 5}\n\n# Sort by frequency\nsorted_pairs = {k: v for k, v in sorted(frequent_pairs.items(), key=lambda item: item[1], reverse=True)}","metadata":{"execution":{"iopub.status.busy":"2022-02-09T09:58:01.268158Z","iopub.execute_input":"2022-02-09T09:58:01.268579Z","iopub.status.idle":"2022-02-09T09:59:11.994487Z","shell.execute_reply.started":"2022-02-09T09:58:01.268530Z","shell.execute_reply":"2022-02-09T09:59:11.993428Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"dict(list(sorted_pairs.items())[0: 5])","metadata":{"execution":{"iopub.status.busy":"2022-02-09T09:59:11.995707Z","iopub.execute_input":"2022-02-09T09:59:11.995972Z","iopub.status.idle":"2022-02-09T09:59:12.308489Z","shell.execute_reply.started":"2022-02-09T09:59:11.995928Z","shell.execute_reply":"2022-02-09T09:59:12.307416Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We can see that `sorted pairs` is a dictionary containing article pairs and their corresponding co occurence frequency within a single customer purchase history (summed over all customers).","metadata":{}},{"cell_type":"code","source":"# Generate final pairing dictionary\nfinal_pairs = {}\nfor k, v in sorted_pairs.items():\n    add_value(final_pairs, k[0], k[1])","metadata":{"execution":{"iopub.status.busy":"2022-02-09T09:59:12.313630Z","iopub.execute_input":"2022-02-09T09:59:12.314040Z","iopub.status.idle":"2022-02-09T09:59:13.585083Z","shell.execute_reply.started":"2022-02-09T09:59:12.313984Z","shell.execute_reply":"2022-02-09T09:59:13.583989Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### generate 12 top myself\n* based on:","metadata":{}},{"cell_type":"code","source":"# top_12_items = df.groupby('article_id')['customer_id'].nunique().sort_values(ascending=False).head(12).index.tolist()\ntop_12_items =  ['0706016001',\n '0372860001',\n '0706016002',\n '0610776002',\n '0759871002',\n '0372860002',\n '0464297007',\n '0720125001',\n '0673396002',\n '0610776001',\n '0673677002',\n '0706016003']","metadata":{"execution":{"iopub.status.busy":"2022-02-09T09:59:13.586624Z","iopub.execute_input":"2022-02-09T09:59:13.586922Z","iopub.status.idle":"2022-02-09T09:59:13.593293Z","shell.execute_reply.started":"2022-02-09T09:59:13.586879Z","shell.execute_reply":"2022-02-09T09:59:13.591760Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"samp_sub = pd.read_csv('../input/h-and-m-personalized-fashion-recommendations/sample_submission.csv')\nsamp_sub['prediction'] =  ' '.join(top_12_items)","metadata":{"execution":{"iopub.status.busy":"2022-02-09T09:59:13.594676Z","iopub.execute_input":"2022-02-09T09:59:13.595243Z","iopub.status.idle":"2022-02-09T09:59:17.995157Z","shell.execute_reply.started":"2022-02-09T09:59:13.595177Z","shell.execute_reply":"2022-02-09T09:59:17.988392Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# # Load in the most-common-benchmark baseline submission ### private datasource\n# samp_sub = pd.read_csv('../input/most-common-bench/most_common_b.csv')\n\n# Convert to dictionary for speed\nsub_dict = samp_sub.to_dict('records')\n\ntransactions_train = pd.read_csv(DATA_PATH/'transactions_train.csv', dtype={'article_id': str},usecols=['t_dat', 'customer_id', 'article_id'])\ntransactions_train = transactions_train[['t_dat', 'customer_id', 'article_id']]\n\n# Sort by date\ntransactions_train['date'] =  pd.to_datetime(transactions_train[\"t_dat\"])\ntransactions_train = transactions_train.sort_values(by=\"date\")\n\ncustomer_dict = dict(zip(transactions_train['customer_id'], transactions_train['article_id']))","metadata":{"execution":{"iopub.status.busy":"2022-02-09T09:59:18.007268Z","iopub.execute_input":"2022-02-09T09:59:18.008437Z","iopub.status.idle":"2022-02-09T10:00:51.657910Z","shell.execute_reply.started":"2022-02-09T09:59:18.008270Z","shell.execute_reply":"2022-02-09T10:00:51.656687Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"dict(list(customer_dict.items())[0: 5])","metadata":{"execution":{"iopub.status.busy":"2022-02-09T10:00:51.659714Z","iopub.execute_input":"2022-02-09T10:00:51.660003Z","iopub.status.idle":"2022-02-09T10:00:52.076871Z","shell.execute_reply.started":"2022-02-09T10:00:51.659969Z","shell.execute_reply":"2022-02-09T10:00:52.076140Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We can see that `customer_dict` is a dictionary with `customer_id` and their most recent purchase `article_id`","metadata":{}},{"cell_type":"code","source":"preds = []\nfor i, row in tqdm(enumerate(sub_dict)):\n    old_pred = row['prediction']\n    customer = row['customer_id']\n    try:\n        most_recent_article = customer_dict[customer]\n        pred = final_pairs[most_recent_article]\n        if type(pred)==str:\n            preds.append(old_pred[:-10] + str(pred))\n            continue\n        preds.append(format_prediction(pred, old_pred))\n    except:\n        preds.append(old_pred)","metadata":{"execution":{"iopub.status.busy":"2022-02-09T10:00:52.077953Z","iopub.execute_input":"2022-02-09T10:00:52.078687Z","iopub.status.idle":"2022-02-09T10:01:03.754064Z","shell.execute_reply.started":"2022-02-09T10:00:52.078647Z","shell.execute_reply":"2022-02-09T10:01:03.753020Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"samp_sub['prediction'] = preds\nsamp_sub.to_csv('submission.csv', index=False)","metadata":{"execution":{"iopub.status.busy":"2022-02-09T10:01:03.756661Z","iopub.execute_input":"2022-02-09T10:01:03.756962Z","iopub.status.idle":"2022-02-09T10:01:17.471138Z","shell.execute_reply.started":"2022-02-09T10:01:03.756928Z","shell.execute_reply":"2022-02-09T10:01:17.470139Z"},"trusted":true},"execution_count":null,"outputs":[]}]}