{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"## Jashanjot Singh Bindra\n### 101903159\n### 3COE16","metadata":{}},{"cell_type":"markdown","source":"## Table of Contents\n**Problem Name** -> H&M Personalized Fashion Recommendations\n\n**Problem Link** -> https://www.kaggle.com/c/h-and-m-personalized-fashion-recommendations/overview\n\n**Problem Type** -> Classification  and Data Analysis\n\n**Libraries used** -> cuDF, cuPy, cuML\n\n**Models Implemented** -> KNearestNeighbours Classifier (using minkowski distance)\n\n**Evaluation Metrics Used** -> Mean Average Precision @ 12\n\n**Kaggle Rank Achieved with total number of teams (if applicable)** -> 482 rank out of 1231\n\n**Tasks done in code:-**\n* Loading training datasets\n* Data Evaluation\n* Pre-processing Training dataset\n* Finding Items that are purchased most often and then sorting  them by date\n* Finding Items that were most popular last week\n* Recommending items by age of customer and other features of article\n* Applying KNN\n* Creating Submission File\n","metadata":{}},{"cell_type":"code","source":"import cudf\nimport cupy as cp\nimport cuml","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","execution":{"iopub.status.busy":"2022-03-16T09:28:30.378829Z","iopub.execute_input":"2022-03-16T09:28:30.379132Z","iopub.status.idle":"2022-03-16T09:28:31.859985Z","shell.execute_reply.started":"2022-03-16T09:28:30.379043Z","shell.execute_reply":"2022-03-16T09:28:31.859214Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Loading training datasets","metadata":{}},{"cell_type":"code","source":"df_train = cudf.read_csv('../input/h-and-m-personalized-fashion-recommendations/transactions_train.csv')\ndf_train.head()","metadata":{"execution":{"iopub.status.busy":"2022-03-16T09:28:31.861774Z","iopub.execute_input":"2022-03-16T09:28:31.862183Z","iopub.status.idle":"2022-03-16T09:28:36.163255Z","shell.execute_reply.started":"2022-03-16T09:28:31.862145Z","shell.execute_reply":"2022-03-16T09:28:36.162563Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"cust = cudf.read_csv('../input/h-and-m-personalized-fashion-recommendations/customers.csv')\ncust.head()","metadata":{"execution":{"iopub.status.busy":"2022-03-16T09:28:36.164484Z","iopub.execute_input":"2022-03-16T09:28:36.165116Z","iopub.status.idle":"2022-03-16T09:28:36.411368Z","shell.execute_reply.started":"2022-03-16T09:28:36.165076Z","shell.execute_reply":"2022-03-16T09:28:36.410631Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"articles=cudf.read_csv('../input/h-and-m-personalized-fashion-recommendations/articles.csv')\narticles.head()","metadata":{"execution":{"iopub.status.busy":"2022-03-16T09:28:36.413936Z","iopub.execute_input":"2022-03-16T09:28:36.414194Z","iopub.status.idle":"2022-03-16T09:28:36.585197Z","shell.execute_reply.started":"2022-03-16T09:28:36.414158Z","shell.execute_reply":"2022-03-16T09:28:36.584559Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Preprocessing Training dataset\n* Here we are trying to reduce the memory consumption of training dataset by storing customer id as int64 which takes 8 bytes instead of string which takes 64 bytes\n* We further reduce memory consumption by storing article id as int32 which takes 4 bytes instead of string which takes 64 bytes\n* We also remove unnecessary columns from training dataset","metadata":{}},{"cell_type":"code","source":"df_train['customer_id'] = df_train['customer_id'].str[-16:].str.hex_to_int().astype('int64')\ndf_train['article_id'] = df_train.article_id.astype('int32')\ndf_train.t_dat = cudf.to_datetime(df_train.t_dat)\ndf_train = df_train[['t_dat','customer_id','article_id']]\ndf_train_original = df_train\nprint( df_train.shape )\ndf_train.head()","metadata":{"execution":{"iopub.status.busy":"2022-03-16T09:28:36.586333Z","iopub.execute_input":"2022-03-16T09:28:36.586664Z","iopub.status.idle":"2022-03-16T09:28:37.004668Z","shell.execute_reply.started":"2022-03-16T09:28:36.586626Z","shell.execute_reply":"2022-03-16T09:28:37.003991Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"cust = cust[['customer_id','age']]\ncust['customer_id'] = cust['customer_id'].str[-16:].str.hex_to_int().astype('int64')\ncust.head()","metadata":{"execution":{"iopub.status.busy":"2022-03-16T09:28:37.006111Z","iopub.execute_input":"2022-03-16T09:28:37.006601Z","iopub.status.idle":"2022-03-16T09:28:37.040845Z","shell.execute_reply.started":"2022-03-16T09:28:37.006562Z","shell.execute_reply":"2022-03-16T09:28:37.040047Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"articles=articles[['article_id','product_type_no','graphical_appearance_no','colour_group_code']]\narticles['article_id'] = articles.article_id.astype('int32')\narticles.head()","metadata":{"execution":{"iopub.status.busy":"2022-03-16T09:28:37.04255Z","iopub.execute_input":"2022-03-16T09:28:37.042933Z","iopub.status.idle":"2022-03-16T09:28:37.063403Z","shell.execute_reply.started":"2022-03-16T09:28:37.042896Z","shell.execute_reply":"2022-03-16T09:28:37.062687Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Finding Customer's Last 2 weeks Purchases\n* We are keeping only those purchases of each customer that are 2 weeks older than his most recent purchase date","metadata":{}},{"cell_type":"code","source":"temp = df_train.groupby('customer_id').t_dat.max().reset_index() #Finding most recent purchase of the customer\ntemp.columns = ['customer_id','max_dat']\ntemp","metadata":{"execution":{"iopub.status.busy":"2022-03-16T09:28:37.06467Z","iopub.execute_input":"2022-03-16T09:28:37.064939Z","iopub.status.idle":"2022-03-16T09:28:37.153718Z","shell.execute_reply.started":"2022-03-16T09:28:37.064906Z","shell.execute_reply":"2022-03-16T09:28:37.152944Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_train = df_train.merge(temp,on=['customer_id'],how='left')\ndf_train['diff_dat'] = (df_train.max_dat - df_train.t_dat).dt.days\ndf_train = df_train.loc[df_train['diff_dat']<=14]","metadata":{"execution":{"iopub.status.busy":"2022-03-16T09:28:37.155082Z","iopub.execute_input":"2022-03-16T09:28:37.155332Z","iopub.status.idle":"2022-03-16T09:28:37.275572Z","shell.execute_reply.started":"2022-03-16T09:28:37.155297Z","shell.execute_reply":"2022-03-16T09:28:37.274872Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_train['diff_dat'].unique() # checking whether all differences are present or not","metadata":{"execution":{"iopub.status.busy":"2022-03-16T09:28:37.276954Z","iopub.execute_input":"2022-03-16T09:28:37.277207Z","iopub.status.idle":"2022-03-16T09:28:37.304085Z","shell.execute_reply.started":"2022-03-16T09:28:37.277173Z","shell.execute_reply":"2022-03-16T09:28:37.303319Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_train","metadata":{"execution":{"iopub.status.busy":"2022-03-16T09:28:37.305423Z","iopub.execute_input":"2022-03-16T09:28:37.305693Z","iopub.status.idle":"2022-03-16T09:28:37.367187Z","shell.execute_reply.started":"2022-03-16T09:28:37.305656Z","shell.execute_reply":"2022-03-16T09:28:37.366419Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## 1) Finding Items that are purchased most often and then sorting them by date\n* If a person purchases an item quite often he is more likely to purchase it again\n* Further we store the most recent of the most frequently purchased items first to further improve predictions","metadata":{}},{"cell_type":"code","source":"temp = df_train.groupby(['customer_id','article_id'])['t_dat'].agg('count').reset_index() # Finding number of times a particular item is purchased by a particular customer\ntemp.columns = ['customer_id','article_id','count']\ntemp","metadata":{"execution":{"iopub.status.busy":"2022-03-16T09:28:37.368372Z","iopub.execute_input":"2022-03-16T09:28:37.368683Z","iopub.status.idle":"2022-03-16T09:28:37.436966Z","shell.execute_reply.started":"2022-03-16T09:28:37.368646Z","shell.execute_reply":"2022-03-16T09:28:37.43619Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_train = df_train.merge(temp,on=['customer_id','article_id'],how='left')\ndf_train = df_train.sort_values(['count','t_dat'],ascending=False)\ndf_train","metadata":{"execution":{"iopub.status.busy":"2022-03-16T09:28:37.440651Z","iopub.execute_input":"2022-03-16T09:28:37.440923Z","iopub.status.idle":"2022-03-16T09:28:37.586201Z","shell.execute_reply.started":"2022-03-16T09:28:37.44089Z","shell.execute_reply":"2022-03-16T09:28:37.58554Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_train = df_train.drop_duplicates(['customer_id','article_id'])\ndf_train = df_train.sort_values(['count','t_dat'],ascending=False)\ndf_train=df_train.reset_index(drop=True)\ndf_train","metadata":{"execution":{"iopub.status.busy":"2022-03-16T09:28:37.58725Z","iopub.execute_input":"2022-03-16T09:28:37.587614Z","iopub.status.idle":"2022-03-16T09:28:37.828837Z","shell.execute_reply.started":"2022-03-16T09:28:37.58758Z","shell.execute_reply":"2022-03-16T09:28:37.827801Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_train=df_train.reset_index(drop=False)\ndf_train","metadata":{"execution":{"iopub.status.busy":"2022-03-16T09:28:37.830117Z","iopub.execute_input":"2022-03-16T09:28:37.830764Z","iopub.status.idle":"2022-03-16T09:28:37.93049Z","shell.execute_reply.started":"2022-03-16T09:28:37.830711Z","shell.execute_reply":"2022-03-16T09:28:37.929858Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## 2) Find Items that were most popular last week\n* We will recommend the 12 most popular items to all the users\n* Extra items will be removed later on while creating submission file so no harm in adding them now\n* Also the problem description says that predicting 12 items for all customers is benificial ","metadata":{}},{"cell_type":"code","source":"print('Latest Date ',df_train_original['t_dat'].max())","metadata":{"execution":{"iopub.status.busy":"2022-03-16T09:28:37.934392Z","iopub.execute_input":"2022-03-16T09:28:37.936283Z","iopub.status.idle":"2022-03-16T09:28:37.945972Z","shell.execute_reply.started":"2022-03-16T09:28:37.936244Z","shell.execute_reply":"2022-03-16T09:28:37.945067Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print('Last Week\\'s Date ',df_train_original['t_dat'].max()-518400000000000) \n# 518400000000000 are nanoseconds in 1 week","metadata":{"execution":{"iopub.status.busy":"2022-03-16T09:28:37.950107Z","iopub.execute_input":"2022-03-16T09:28:37.952174Z","iopub.status.idle":"2022-03-16T09:28:37.958393Z","shell.execute_reply.started":"2022-03-16T09:28:37.952114Z","shell.execute_reply":"2022-03-16T09:28:37.957615Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_train_original = df_train_original.loc[df_train_original.t_dat >= cudf.to_datetime('2020-09-16')]\ntop12 = ' 0' + ' 0'.join(df_train_original.article_id.value_counts().to_pandas().index.astype('str')[:12])\nprint(\"Last week's top 12 popular items:\")\nprint( top12 )","metadata":{"execution":{"iopub.status.busy":"2022-03-16T09:28:37.960118Z","iopub.execute_input":"2022-03-16T09:28:37.960848Z","iopub.status.idle":"2022-03-16T09:28:38.044911Z","shell.execute_reply.started":"2022-03-16T09:28:37.960809Z","shell.execute_reply":"2022-03-16T09:28:38.044067Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## 3) Recommending items by age of customer and other features of article\n* In this we'll be using KNN to predict the nearest/most similar article that the customer will buy\n* Here we are using the concept if a person has bought a product with certain colour, product type(Like tshirt, shorts), material then he is more likely to buy another product with similar characteristic\n* We also take the customer's age into consideration, people of similar age group by similar clothes","metadata":{}},{"cell_type":"code","source":"age_train=cudf.merge(df_train, cust, on='customer_id')\nage_train.head()","metadata":{"execution":{"iopub.status.busy":"2022-03-16T09:28:38.048396Z","iopub.execute_input":"2022-03-16T09:28:38.049178Z","iopub.status.idle":"2022-03-16T09:28:38.104224Z","shell.execute_reply.started":"2022-03-16T09:28:38.049126Z","shell.execute_reply":"2022-03-16T09:28:38.103496Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"age_train=age_train[['index','customer_id','age','article_id']]\nage_train=age_train.fillna({'age':18})\nage_train","metadata":{"execution":{"iopub.status.busy":"2022-03-16T09:28:38.107679Z","iopub.execute_input":"2022-03-16T09:28:38.10841Z","iopub.status.idle":"2022-03-16T09:28:38.195383Z","shell.execute_reply.started":"2022-03-16T09:28:38.10837Z","shell.execute_reply":"2022-03-16T09:28:38.194753Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"arti_age_train = cudf.merge(age_train, articles, on='article_id')\narti_age_train","metadata":{"execution":{"iopub.status.busy":"2022-03-16T09:28:38.198942Z","iopub.execute_input":"2022-03-16T09:28:38.201179Z","iopub.status.idle":"2022-03-16T09:28:38.293259Z","shell.execute_reply.started":"2022-03-16T09:28:38.201139Z","shell.execute_reply":"2022-03-16T09:28:38.292602Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"X=arti_age_train[['age','product_type_no','graphical_appearance_no','colour_group_code']]\nX.head()","metadata":{"execution":{"iopub.status.busy":"2022-03-16T09:28:38.294457Z","iopub.execute_input":"2022-03-16T09:28:38.294859Z","iopub.status.idle":"2022-03-16T09:28:38.314946Z","shell.execute_reply.started":"2022-03-16T09:28:38.294823Z","shell.execute_reply":"2022-03-16T09:28:38.314163Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"Y=arti_age_train[['article_id']]\nY.head()","metadata":{"execution":{"iopub.status.busy":"2022-03-16T09:28:38.316314Z","iopub.execute_input":"2022-03-16T09:28:38.316563Z","iopub.status.idle":"2022-03-16T09:28:38.331848Z","shell.execute_reply.started":"2022-03-16T09:28:38.316529Z","shell.execute_reply":"2022-03-16T09:28:38.331108Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Applying KNN\n* The value of K I have taken is K=10\n* I am using MinkowskiDistance as distance metric","metadata":{}},{"cell_type":"code","source":"from cuml.neighbors import KNeighborsClassifier\n\nknn = KNeighborsClassifier(n_neighbors=10,metric='minkowski')\n\nknn.fit(X[:50000], Y[:50000]) #training only subsample to avoid overfitting and save gpu memory","metadata":{"execution":{"iopub.status.busy":"2022-03-16T09:30:38.819117Z","iopub.execute_input":"2022-03-16T09:30:38.819518Z","iopub.status.idle":"2022-03-16T09:30:38.847401Z","shell.execute_reply.started":"2022-03-16T09:30:38.819482Z","shell.execute_reply":"2022-03-16T09:30:38.846757Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# ans=knn.predict(X) wanted to do this but gpu is going out of memory\nans=knn.predict(X[:100000])","metadata":{"execution":{"iopub.status.busy":"2022-03-16T09:28:38.672824Z","iopub.status.idle":"2022-03-16T09:28:38.673516Z","shell.execute_reply.started":"2022-03-16T09:28:38.673257Z","shell.execute_reply":"2022-03-16T09:28:38.673284Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"ls = ans.to_arrow().to_pylist()","metadata":{"execution":{"iopub.status.busy":"2022-03-16T09:28:38.67463Z","iopub.status.idle":"2022-03-16T09:28:38.675415Z","shell.execute_reply.started":"2022-03-16T09:28:38.675168Z","shell.execute_reply":"2022-03-16T09:28:38.675197Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"arti_age_trai = arti_age_train[['index','customer_id','article_id']]\narti_age_trai","metadata":{"execution":{"iopub.status.busy":"2022-03-16T09:28:38.676618Z","iopub.status.idle":"2022-03-16T09:28:38.677384Z","shell.execute_reply.started":"2022-03-16T09:28:38.677139Z","shell.execute_reply":"2022-03-16T09:28:38.677173Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df = {'index': range(len(arti_age_trai),len(arti_age_trai)+len(ls)), 'customer_id': arti_age_trai['customer_id'][:100000], 'article_id': ls}\narti_age_trai = arti_age_trai.append(df, ignore_index = True)","metadata":{"execution":{"iopub.status.busy":"2022-03-16T09:28:38.67856Z","iopub.status.idle":"2022-03-16T09:28:38.679326Z","shell.execute_reply.started":"2022-03-16T09:28:38.679077Z","shell.execute_reply":"2022-03-16T09:28:38.679104Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"arti_age_trai","metadata":{"execution":{"iopub.status.busy":"2022-03-16T09:28:38.680473Z","iopub.status.idle":"2022-03-16T09:28:38.681257Z","shell.execute_reply.started":"2022-03-16T09:28:38.681002Z","shell.execute_reply":"2022-03-16T09:28:38.681029Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### I am maintaing index because previously I calculated recently most purchased items previous and i will recommend those first and then KNN predictions later","metadata":{}},{"cell_type":"code","source":"arti_age_trai = arti_age_trai[['index','customer_id','article_id']].sort_values('index')\narti_age_trai = arti_age_trai[['customer_id','article_id']]\narti_age_trai=arti_age_trai.reset_index(drop=True)\narti_age_trai","metadata":{"execution":{"iopub.status.busy":"2022-03-16T09:28:38.682395Z","iopub.status.idle":"2022-03-16T09:28:38.683178Z","shell.execute_reply.started":"2022-03-16T09:28:38.682924Z","shell.execute_reply":"2022-03-16T09:28:38.682954Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"arti_age_trai = arti_age_trai.drop_duplicates(['customer_id','article_id'])\narti_age_trai","metadata":{"execution":{"iopub.status.busy":"2022-03-16T09:28:38.684265Z","iopub.status.idle":"2022-03-16T09:28:38.685028Z","shell.execute_reply.started":"2022-03-16T09:28:38.684788Z","shell.execute_reply":"2022-03-16T09:28:38.684815Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_train = arti_age_trai.sort_index()","metadata":{"execution":{"iopub.status.busy":"2022-03-16T09:28:38.68613Z","iopub.status.idle":"2022-03-16T09:28:38.686906Z","shell.execute_reply.started":"2022-03-16T09:28:38.686645Z","shell.execute_reply":"2022-03-16T09:28:38.686673Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"##  Creating Submission File\n* In this file we group all article ids for a customer and then store it as a string as required by submission rules","metadata":{}},{"cell_type":"code","source":"df_train.article_id = ' 0' + df_train.article_id.astype('str')\ndf_train","metadata":{"execution":{"iopub.status.busy":"2022-03-16T09:28:38.688011Z","iopub.status.idle":"2022-03-16T09:28:38.688796Z","shell.execute_reply.started":"2022-03-16T09:28:38.68853Z","shell.execute_reply":"2022-03-16T09:28:38.688557Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"p_df_train = df_train[['customer_id','article_id']].to_pandas() #cudf does not support sum of str in group by and loop is expensive so we convert to pandas\np_df_train","metadata":{"execution":{"iopub.status.busy":"2022-03-16T09:28:38.689952Z","iopub.status.idle":"2022-03-16T09:28:38.690712Z","shell.execute_reply.started":"2022-03-16T09:28:38.690463Z","shell.execute_reply":"2022-03-16T09:28:38.690491Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"temp = p_df_train.groupby('customer_id').sum().reset_index()\ntemp.columns = ['customer_id','prediction']\ndf_train=cudf.DataFrame(temp)","metadata":{"execution":{"iopub.status.busy":"2022-03-16T09:28:38.692018Z","iopub.status.idle":"2022-03-16T09:28:38.692872Z","shell.execute_reply.started":"2022-03-16T09:28:38.69261Z","shell.execute_reply":"2022-03-16T09:28:38.692637Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_train","metadata":{"execution":{"iopub.status.busy":"2022-03-16T09:28:38.694015Z","iopub.status.idle":"2022-03-16T09:28:38.694878Z","shell.execute_reply.started":"2022-03-16T09:28:38.694616Z","shell.execute_reply":"2022-03-16T09:28:38.694643Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_train.rename(columns={'customer_id':'customer_id_edited'},inplace=True)\ndf_train","metadata":{"execution":{"iopub.status.busy":"2022-03-16T09:28:38.696027Z","iopub.status.idle":"2022-03-16T09:28:38.696811Z","shell.execute_reply.started":"2022-03-16T09:28:38.696549Z","shell.execute_reply":"2022-03-16T09:28:38.696576Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"submission = cudf.read_csv('../input/h-and-m-personalized-fashion-recommendations/sample_submission.csv')\nsubmission = submission[['customer_id']]\nsubmission['customer_id_edited'] = submission['customer_id'].str[-16:].str.hex_to_int().astype('int64')\nsubmission = submission.merge(df_train, on='customer_id_edited', how='left').fillna('')\ndel submission['customer_id_edited']\nsubmission","metadata":{"execution":{"iopub.status.busy":"2022-03-16T09:28:38.697975Z","iopub.status.idle":"2022-03-16T09:28:38.69875Z","shell.execute_reply.started":"2022-03-16T09:28:38.698487Z","shell.execute_reply":"2022-03-16T09:28:38.698514Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"submission.prediction = submission.prediction + top12\nsubmission.prediction = submission.prediction.str.strip()\nsubmission.prediction = submission.prediction.str[:131] # 10 * 12 = 120 plus 11 spaces is 131, we do this to only keep 12 predictions for each customer\nsubmission.to_csv('submission.csv',index=False)\nsubmission.head()","metadata":{"execution":{"iopub.status.busy":"2022-03-16T09:28:38.70021Z","iopub.status.idle":"2022-03-16T09:28:38.701067Z","shell.execute_reply.started":"2022-03-16T09:28:38.700788Z","shell.execute_reply":"2022-03-16T09:28:38.700841Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]}]}