{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"## H&M Personalized Fashion Recommendations\n<br>\n\n### Summary : \n\n   <font size=\"3\">In this competition, we are asked to create a product recommendation for 7 products for each customer using the datasets given from us.</font>\n<br>\n\n### Working Method : \n<br>\n\n   <font size=\"3\"> If we group the characteristics of people who buy products sold in a store, we can assume that people in that group will be inclined to buy that product. We will continue this approach in our solution process, take the characteristics of the people who buy each product, classify our products and make an estimate for each customer. We will make these estimations by finding which group the customer belongs to from the products purchased in the past, and by giving the most purchased products in that group. While making these estimates, in order not to destroy the value of the customer's most recent product, we will act as a separate customer who has only purchased the last product for the first 4 products, and will use the other purchased products for the remaining 3 products. If our customer has bought only 1 product, we will act as 1 customer and bring all products from that class. With this distinction, I explained the information in the sub-blocks.</font>\n<br>\n\n### Solution Method : \n<br>\n\n<font size=\"3\">1. A DF will be created including the characteristics of the users who bought each product, and classification will be made with this information.</font><br>\n<font size=\"3\">2. Then, each customer's information will be given to this model and its class will be found.</font><br>\n<font size=\"3\">3. Estimates will be made so that the first four of the products in its class are from the last purchased product, and the remaining 3 are from other purchased products.</font>\n\n<br>\n\n### Datasets : \n<br>\n\n<font size=\"3\"> We have 4 types of data sets. These are in order \"articles, customers, ssample_submission, transactions_train\".\n*     <font size=\"3\">articles : It is our dataset containing products and their properties.</font>\n*     <font size=\"3\">customers : It is our dataset containing customer information.</font>\n*     <font size=\"3\">sample_submission : This is our sample dataset of the requested data.</font>\n*     <font size=\"3\">transactions_train : It is our dataset containing information about customers and their purchases.</font>\n</font>\n<br>\n\n### Process steps :\n<br>\n\n\n<font size=\"3\">1. &nbsp;&nbsp;Step :  <a href=#articlesImport>Importing Data\n</a></font><br>\n<font size=\"3\">&nbsp;&nbsp;&nbsp;&nbsp;   1.1.  <a href=#articlesImport>Articles İmport</a></font><br>\n<font size=\"3\">&nbsp;&nbsp;&nbsp;&nbsp;   1.2.  <a href=#customersImport>Customers İmport</a></font><br>\n<font size=\"3\">&nbsp;&nbsp;&nbsp;&nbsp;   1.3.  <a href=#sampleSubmissionImport>Sample Submission İmport</a></font><br>\n<font size=\"3\">&nbsp;&nbsp;&nbsp;&nbsp;   1.4.  <a href=#transactionsImport>Transactions İmport</a></font><br>\n<font size=\"3\">2. &nbsp;&nbsp;Step : <a href=#transactionsPreprocessing>Transactions DF PreProcessing</a></font><br>\n<font size=\"3\">3. &nbsp;&nbsp;Step : <a href=#transactionsCustomerArticlesDF>Creating  Transactions-Customer-Articles DF</a></font><br>\n<font size=\"3\">4. &nbsp;&nbsp;Step : <a href=#customersPrePro>Customers DF PreProcessing</a></font><br>\n<font size=\"3\">5. &nbsp;&nbsp;Step : <a href=#customersFeatureEngineering>Customers DF Feature Engineering</a></font><br>\n<font size=\"3\">6. &nbsp;&nbsp;Step : <a href=#articlesCustomerDF>Creating Articles-Customers DF</a></font><br>\n<font size=\"3\">7. &nbsp;&nbsp;Step : <a href=#articlesPreProcessing>Articles DF PreProcessing</a></font><br>\n<font size=\"3\">8. &nbsp;&nbsp;Step : <a href=#articlesCustomerValuesDF>Creating ArticlesCustomerValues DF</a></font><br>\n<font size=\"3\">9. &nbsp;&nbsp;Step : <a href=#articlesDFMerge>Combining ArticlesCustomerValues with ArticlesMost DF</a></font><br>\n<font size=\"3\">10. Step : <a href=#articlesDFFeatureEngineering>Articles DF Feature Engineering</a></font><br>\n<font size=\"3\">11. Step : <a href=#classification>Classification Operations ( K-Means, LightGBM )</a></font><br>\n<font size=\"3\">12. Step : <a href=#model>Model Building Stages</a></font><br>\n<font size=\"3\">13. Step : <a href=#customerSubmission>Estimation Operations</a></font><br>\n<br>\n    \n    \n    ","metadata":{}},{"cell_type":"markdown","source":"<a id='articlesImport' style=\"color:black\" /></a>\n# Importing Data","metadata":{}},{"cell_type":"markdown","source":"<a id='articlesImport' style=\"color:black\" /></a>\n# Articles İmport\n<font size=\"3\">We import our Articles data.</font>","metadata":{}},{"cell_type":"code","source":"import pandas as pd\nimport numpy as np\nimport statsmodels.api as sm\nimport matplotlib.pyplot as plt\nimport scipy.stats as stats\nimport seaborn as sns\nsns.set()\nraw_csv_data = pd.read_csv(\"../input/h-and-m-personalized-fashion-recommendations/articles.csv\")\npd.set_option('display.max_columns',30 )\ndf_articles_root = raw_csv_data.copy()\ndf_customers_root_checkpoint = df_articles_root.copy()\ndf_articles_root.head()","metadata":{"scrolled":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id='customersImport' style=\"color:black\" /></a>\n# Customers İmport\n<font size=\"3\">We import the customer dataset</font>","metadata":{}},{"cell_type":"code","source":"raw_csv_data = pd.read_csv(\"../input/h-and-m-personalized-fashion-recommendations/customers.csv\")\ndf_customers_root = raw_csv_data.copy()\n# df_customers_root = df_customers_root.head(500)\ndf_customers_root\ndf_customers_root_checkpoint = df_customers_root.copy()","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id='sampleSubmissionImport' style=\"color:black\" /></a>\n# Sample Submission İmport\n<font size=\"3\">We import the sample submission dataset</font>","metadata":{}},{"cell_type":"code","source":"raw_csv_data = pd.read_csv(\"../input/h-and-m-personalized-fashion-recommendations/sample_submission.csv\")\ndf_sampleSubmission_root = raw_csv_data.copy()\n# display(df)\ndf_sampleSubmission_root.head()","metadata":{"scrolled":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a name='bookmark' />","metadata":{}},{"cell_type":"markdown","source":"<a id='transactionsImport' style=\"color:black\" /></a>\n# Transactions İmport\n<font size=\"3\">We import the Transactions dataset</font>","metadata":{}},{"cell_type":"code","source":"\nraw_csv_data = pd.read_csv(\"../input/h-and-m-personalized-fashion-recommendations/transactions_train.csv\")\ndf_transactionsTrain_root = raw_csv_data.copy()\ndf_transactionsTrain_root = df_transactionsTrain_root.head(10000)\ndf_transactionsTrain_root_checkpoint = df_transactionsTrain_root.copy()\ndf_transactionsTrain_root","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id='transactionsPreprocessing' style=\"color:black\" /></a>\n# Transactions Preprocessing","metadata":{}},{"cell_type":"code","source":"df_transactionsTrain_root_chkpoint = df_transactionsTrain_root.copy()","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_transactionsTrain_root.isnull().sum()","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_transactionsTrain_root['t_dat'].value_counts()","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<font size=\"3\">Since we will buy the last product separately, I will list our transactions. So I'm converting my date variable to date type</font>","metadata":{}},{"cell_type":"markdown","source":"<font size=\"3\">All the dates are the same in the given dataset, I am doing it like this as an example</font>","metadata":{}},{"cell_type":"code","source":"df_transactionsTrain_root.info()","metadata":{"scrolled":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_transactionsTrain_root['t_dat'] = pd.to_datetime(df_transactionsTrain_root['t_dat'])","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_transactionsTrain_root.info()","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_transactionsTrain_root.sort_values(by='t_dat',ascending=False)","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id='transactionsCustomerArticlesDF' style=\"color:black\" /></a>\n# Creating TransactionsCustomerArticles DF","metadata":{}},{"cell_type":"markdown","source":"<font size=\"3\">I create a table with customers and the products they buy.</font>","metadata":{}},{"cell_type":"code","source":"df_transactionsTrain_root_chkpoint = df_transactionsTrain_root.copy()","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<font size=\"3\">Unique number of customers</font>","metadata":{}},{"cell_type":"code","source":"df_transactionsTrain_root['customer_id'].unique().size","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<font size=\"3\">I group transactions by customer id information</font>","metadata":{}},{"cell_type":"code","source":"df_groupedCustomerTransactions = df_transactionsTrain_root.groupby([\"customer_id\"],as_index=False)","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<font size=\"3\">Display of a customer's DF as an example</font>","metadata":{}},{"cell_type":"code","source":"df_groupedCustomerTransactions.get_group(\"000058a12d5b43e67d225668fa1f8d618c13dc232df0cad8ffe7ad4a1091e318\")","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<font size=\"3\">I am creating a table where every customer who buys is once. I will combine it with the products of my customer that I grouped into this table.</font>","metadata":{}},{"cell_type":"code","source":"df_TRS_TR_OneCus = df_transactionsTrain_root.drop_duplicates(subset=['customer_id'])","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_TRS_TR_OneCus","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<font size=\"3\">I place my column where I will put the products purchased by the customer. Sometimes I use the try-except construct as it gives an error when I rerun it to try.</font>","metadata":{}},{"cell_type":"code","source":"try:\n    df_TRS_TR_OneCus.insert(len(df_TRS_TR_OneCus.columns), 'articles', [None] * len(df_TRS_TR_OneCus))\nexcept:\n    print(\"Already created\")\n","metadata":{"scrolled":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_TRS_TR_OneCus","metadata":{"scrolled":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<font size=\"3\">I take the products purchased by each customer from the transaction DF that I have grouped and put them in the articles column in the same table.</font>     \n     ","metadata":{}},{"cell_type":"code","source":"def get_article_indexes(val):\n    groupArticles = df_groupedCustomerTransactions.get_group(val)['article_id'].reset_index()\n    groupArticles = groupArticles.drop(['index'],axis=1)\n    groupArticles = groupArticles.to_numpy()\n    groupArticles = np.reshape(groupArticles, groupArticles.size)\n    \n    df_TRS_TR_OneCus.loc[df_TRS_TR_OneCus['customer_id'] == val,'articles'] = ','.join(map(str, groupArticles))","metadata":{"scrolled":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_TRS_TR_OneCus['customer_id'].apply(get_article_indexes)","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<font size=\"3\">I am removing columns of unnecessary variables that I will no longer use.</font>          ","metadata":{}},{"cell_type":"code","source":"df_TRS_TR_OneCus = df_TRS_TR_OneCus.drop(['sales_channel_id','price','t_dat','article_id'],axis=1,errors = 'ignore')\n# df_transactionsTrain_nonDuplicates\ndf_TRS_TR_OneCus","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<font size=\"3\">I am creating cp for backup purpose. I will usually do this. DF Ready</font>\n","metadata":{}},{"cell_type":"code","source":"df_TRS_TR_OneCus_checkpoint = df_TRS_TR_OneCus.copy()","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_TRS_TR_OneCus","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id='customersPrePro' style=\"color:black\" /></a>\n# Customers PreProcessing","metadata":{}},{"cell_type":"code","source":"# (PreProcessing)\ndf_customers_root_PreProcessing_checkpoint  = df_customers_root.copy()\ndf_customers_root","metadata":{"scrolled":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<font size=\"3\">I'm removing the postal_code variable because I won't be using it</font>","metadata":{}},{"cell_type":"code","source":"df_customers_root = df_customers_root.drop(['postal_code'],axis=1,errors='ignore')","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<font size=\"3\">I see there are null values and I start filling my categorical variables</font>","metadata":{}},{"cell_type":"code","source":"df_customers_root.isnull().sum()","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"for col in df_customers_root:\n    if(col == 'customer_id'):\n        continue\n    print(df_customers_root[col].unique())","metadata":{"scrolled":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_customers_root.info()","metadata":{"scrolled":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_customers_root[\"FN\"].fillna(0,inplace = True)","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_customers_root[\"Active\"].fillna(0,inplace = True)","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_customers_root[\"club_member_status\"].fillna('EMPTY',inplace = True)\ndf_customers_root['club_member_status'].replace('LEFT CLUB', 'LEFT_CLUB', inplace=True)\ndf_customers_root['club_member_status'].replace('PRE-CREATE', 'PRE_CREATE', inplace=True)\n","metadata":{"scrolled":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_customers_root['club_member_status'].value_counts()","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"for col in df_customers_root:\n    if(col == 'customer_id'):\n        continue\n    print(df_customers_root[col].unique())","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_customers_root[\"fashion_news_frequency\"].fillna('EMPTY',inplace = True)\ndf_customers_root['fashion_news_frequency'].replace('NONE', 'EMPTY', inplace=True)\ndf_customers_root['fashion_news_frequency'].replace('None', 'EMPTY', inplace=True)","metadata":{"scrolled":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"for col in df_customers_root:\n    if(col == 'customer_id'):\n        continue\n    print(df_customers_root[col].unique())","metadata":{"scrolled":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<font size=\"3\">Since there are no extreme values in my age value, I can fill it with the average. I fill it with the average.</font>","metadata":{}},{"cell_type":"code","source":"df_customers_root[\"age\"].fillna(df_customers_root[\"age\"].mean(),inplace = True)","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_customers_root.isnull().sum()","metadata":{"scrolled":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"for col in df_customers_root:\n    if(col == 'customer_id'):\n        continue\n    print(df_customers_root[col].unique())","metadata":{"scrolled":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_customers_root","metadata":{"scrolled":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id='customersFeatureEngineering' style=\"color:black\" /></a>\n# Customers Feature Engineering","metadata":{}},{"cell_type":"code","source":"df_customers_root = pd.get_dummies(df_customers_root, columns = [\"club_member_status\"], prefix = [\"CMS\"])","metadata":{"scrolled":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_customers_root =pd.get_dummies(df_customers_root, columns = [\"fashion_news_frequency\"], prefix = [\"FNF\"])","metadata":{"scrolled":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_customers_root","metadata":{"scrolled":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_customers_root.isnull().sum()","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id='articlesCustomerDF' style=\"color:black\" /></a>\n# Creating Articles-Customers DF","metadata":{}},{"cell_type":"code","source":"df_merged_transactionsCustomer = pd.merge(df_transactionsTrain_root,df_customers_root,on='customer_id'\n                                         ,how='left')\n","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_merged_transactionsCustomer","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<font size=\"3\">I'm grouping purchased items for people who buy them.</font>","metadata":{}},{"cell_type":"code","source":"df_groupedBuyingArticles = df_merged_transactionsCustomer.groupby([\"article_id\"],as_index=False)\n","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_merged_transactionsCustomer_checkpoint = df_merged_transactionsCustomer.copy()","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_merged_transactionsCustomer.isnull().sum()","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_groupedBuyingArticles[[\"price\"]].aggregate(\"mean\")","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"   \n<font size=\"3\">The average of the characteristics of the customers who purchased the product. I'm buying. For example, if there is 0.3 cms_active, it means that 3 out of 10 users are active.</font>     ","metadata":{}},{"cell_type":"code","source":"# pd.set_option('display.max_rows', 50)\ndf_groupedBuyingArticles[[\"CMS_ACTIVE\",\"CMS_EMPTY\",\"CMS_LEFT_CLUB\",\"CMS_PRE_CREATE\",\n                    \"FNF_EMPTY\",\"FNF_Monthly\",\"FNF_Regularly\",]].mean()","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_groupedBuyingArticles[[\"age\"]].aggregate(\"mean\")","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_groupedBuyingArticles[[\"FN\"]].aggregate(\"mean\")","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_groupedBuyingArticlesValues = df_groupedBuyingArticles[[\"price\"]].aggregate(\"mean\")\ncacheArr = df_groupedBuyingArticles[[\"customer_id\"]].count().sort_values('customer_id', ascending=False)\ncacheArr2 = df_groupedBuyingArticles[[\"FN\"]].aggregate(\"mean\")\ncacheArr3 = df_groupedBuyingArticles[[\"Active\"]].aggregate(\"mean\")\ncacheArr4 = df_groupedBuyingArticles[[\"age\"]].aggregate(\"mean\")\ncacheArr5 = df_groupedBuyingArticles[[\"CMS_ACTIVE\",\"CMS_EMPTY\",\"CMS_LEFT_CLUB\",\"CMS_PRE_CREATE\",\n                    \"FNF_EMPTY\",\"FNF_Monthly\",\"FNF_Regularly\",]].mean()\ndf_groupedBuyingArticlesValues = pd.merge(df_groupedBuyingArticlesValues,cacheArr, how='outer')\ndf_groupedBuyingArticlesValues = pd.merge(df_groupedBuyingArticlesValues,cacheArr2, how='outer')\ndf_groupedBuyingArticlesValues = pd.merge(df_groupedBuyingArticlesValues,cacheArr3, how='outer')\ndf_groupedBuyingArticlesValues = pd.merge(df_groupedBuyingArticlesValues,cacheArr4, how='outer')\ndf_groupedBuyingArticlesValues = pd.merge(df_groupedBuyingArticlesValues,cacheArr5, how='outer').sort_values('customer_id', ascending=False)\n\ndf_groupedBuyingArticlesValues\n\n\n\n","metadata":{"scrolled":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<font size=\"3\">Two dataframes are ready.</font>","metadata":{}},{"cell_type":"code","source":"display(df_TRS_TR_OneCus)\ndisplay(df_groupedBuyingArticlesValues)","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<font size=\"3\">Now I will create a DF that shows the features of the other products bought by the people who buy these products.</font>","metadata":{}},{"cell_type":"markdown","source":"<a id='articlesPreProcessing' style=\"color:black\" /></a>\n# Articles PreProcessing","metadata":{}},{"cell_type":"code","source":"df_merged_transactionsCustomer","metadata":{"scrolled":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_articles_root.isnull().sum()","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<font size=\"3\">I will not use product descriptions and images. Because the detail in the product table is more than enough for me.</font>","metadata":{}},{"cell_type":"code","source":"df_articles_root = df_articles_root.drop(['detail_desc'],axis=1, errors = 'ignore')\ndf_articles_root","metadata":{"scrolled":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_articles_all = df_articles_root.copy()","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"for col in df_articles_all:\n#     print(col)\n    if(col == 'article_id' ):\n        continue\n    print(col)\n    print(df_articles_all[col].unique())\n    print(df_articles_all[col].unique().size)","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"\n<font size=\"3\">I have two name variables of Promotion/ Other /Offer . I'm fixing this situation</font>","metadata":{}},{"cell_type":"code","source":"df_articles_all.loc[df_articles_all.department_name == 'Promotion/ Other /Offer','department_name'] = \"PromotionOtherOffer\"\ndf_articles_all.loc[df_articles_all.department_name == 'Promotion/Other/Offer','department_name'] = \"PromotionOtherOffer\"\n","metadata":{"scrolled":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"for col in df_articles_all:\n#     print(col)\n    if(col == 'article_id' ):\n        continue\n    print(col)\n    print(df_articles_all[col].unique())\n    print(df_articles_all[col].unique().size)","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"\n<font size=\"3\">I can remove it because values like the product code found here are expressed categorically as _name. so I'm removing it.</font>","metadata":{}},{"cell_type":"code","source":"df_articles_all = df_articles_all.drop(['product_code','product_type_no','graphical_appearance_no',\n                                     'colour_group_code','perceived_colour_value_id',\n                                     'perceived_colour_master_id','department_no',\n                                     'index_group_no','index_code','section_no','garment_group_no'],axis=1,errors='ignore')\n\ndf_articles_all","metadata":{"scrolled":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<font size=\"3\">Since I will create a dummies variable with the variables here, there should be no unacceptable expressions in it. So I'm removing these letters.</font>","metadata":{}},{"cell_type":"code","source":"import re\ncache_df = df_articles_all.copy()\nfor col in cache_df:\n    if(col.endswith(\"_name\")):\n        cache_df[col] = cache_df[col].transform(lambda x:re.sub('[^A-Za-z0-9_]+', '', x))\ndf_articles_all = cache_df\ndf_articles_all","metadata":{"scrolled":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<font size=\"3\">Since processing power is required while shaping the product table, I reduce the table to use only purchased products.\nWe don't have to. I just did it this way because it would take a lot of time. We shouldn't do that in real life.</font>","metadata":{}},{"cell_type":"markdown","source":"<font size=\"3\">df_merged_transactionsCustomer contained only customers and article ids. Here, we said, \"Bring me only the rows that have the article id in this table.\" I have earned about 100000 lines.</font>","metadata":{}},{"cell_type":"code","source":"df_reduced_article = df_articles_all[df_articles_all['article_id'].isin(df_merged_transactionsCustomer['article_id'])]\ndf_reduced_article","metadata":{"scrolled":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id='articlesCustomerValuesDF' style=\"color:black\" /></a>\n# Creating ArticlesCustomerValues DF's","metadata":{}},{"cell_type":"code","source":"df_TransactionsForArticles = df_merged_transactionsCustomer.copy()","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<font size=\"3\">I create my DF where each purchased product is 1 time</font>","metadata":{}},{"cell_type":"code","source":"df_transactionsOnlyOneArt = df_TransactionsForArticles.drop_duplicates(subset=['article_id'])","metadata":{"scrolled":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_transactionsArtGrouped = df_TransactionsForArticles.groupby([\"article_id\"],as_index=False)\n","metadata":{"scrolled":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<font size=\"3\">As you can see, both have the same number of elements.</font>","metadata":{}},{"cell_type":"code","source":"pd.set_option('display.max_colwidth', 50)\ndisplay(df_transactionsOnlyOneArt[['article_id']])\ndisplay(df_transactionsOnlyOneArt)","metadata":{"scrolled":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"\n<font size=\"3\">We have created a DF that brings the other products purchased by the customers who bought each product. base_article_id becomes the id of our main product. Other columns show other products purchased by customers who have purchased that product</font>","metadata":{}},{"cell_type":"code","source":"\nclass Articles:\n    df_articles = None\n    df_articleAndCustomerMost = None\n    \n    # Brings the products bought by the customer who bought that product.\n    \n    def get_customer_most(self,val):\n        \n        articles = df_TRS_TR_OneCus[df_TRS_TR_OneCus.customer_id == val]['articles'].to_numpy()\n        articles = np.fromstring(articles[0], dtype=np.int64, sep=',')\n        articles = np.int64(articles)\n        return df_reduced_article[df_reduced_article['article_id'].isin(articles)]\n\n\n     # It brings the products purchased by the customers who bought the product and combines them with the main product id.\n    \n    def get_article_customers(self,val):\n        self.df_articles = pd.DataFrame()\n        vals = df_transactionsArtGrouped.get_group(int(val))\n        i = 0\n    \n        for customer_id in vals['customer_id']:\n            two = self.get_customer_most(customer_id)\n            \n            if(self.df_articles.empty):\n                self.df_articles = two\n            else:\n                self.df_articles = pd.concat([self.df_articles,two])\n        \n        self.df_articles.insert(0, 'base_article_id', val)\n        if(self.df_articleAndCustomerMost.empty):\n            self.df_articleAndCustomerMost = self.df_articles.copy(deep=True)\n        else:\n            self.df_articleAndCustomerMost = pd.concat([self.df_articleAndCustomerMost,self.df_articles])\n            \n                        \n                \n    def build(self,arr,head):\n        self.df_articleAndCustomerMost = pd.DataFrame()\n        if(head == 0):\n            arr.apply(self.get_article_customers)\n        else:\n            arr.head(head).apply(self.get_article_customers)\n        \n        \n\na = Articles()             \na.build(df_transactionsOnlyOneArt['article_id'],0)\n\ndisplay(a.df_articleAndCustomerMost)\n                                                              ","metadata":{"scrolled":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"\n<font size=\"3\">I take a backup and export our class variable.</font>","metadata":{}},{"cell_type":"code","source":"df_articleAndCustomerMost_checkpoint = a.df_articleAndCustomerMost.copy()\ndf_articleAndCustomerMost = a.df_articleAndCustomerMost\ndf_articleAndCustomerMost","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_articleAndCustomerMost.info()","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<font size=\"3\">I'm grouping the products that have been bought according to the main product id information.</font>","metadata":{}},{"cell_type":"code","source":"df_articleAndCustomerMost_Grouped = df_articleAndCustomerMost.groupby([\"base_article_id\"],as_index=False)","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<font size=\"3\">For example, other products purchased by customers who bought our product with the id \"355569001\".</font>","metadata":{}},{"cell_type":"code","source":"df_articleAndCustomerMost_Grouped.get_group(355569001)","metadata":{"scrolled":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<font size=\"3\">Since I will find a feature for the main product, I create its DF.</font>","metadata":{}},{"cell_type":"code","source":"df_baseArticleOnly = df_articleAndCustomerMost.drop_duplicates(subset=['base_article_id'])\ndf_baseArticleOnly","metadata":{"scrolled":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"\n<font size=\"3\">I take the most recurring features of the products that were taken alongside the main product and combine it with the df I created with df_baseArticleOnly. I add the most_ prefix to the beginning of their names.</font>","metadata":{}},{"cell_type":"markdown","source":"<font size=\"3\">I'm renaming base_article_id to article_id as I'm going to merge this with the customers table</font>","metadata":{}},{"cell_type":"code","source":"class ArticleGroup:\n    \n    \n    df_article_most = None\n    base_arr = None;\n    columns = None;\n        \n   \n    def iteration(self,base_article_id,grouped_arr):\n        values = np.array([])\n        group_arr = grouped_arr.get_group(base_article_id)\n        \n        \n        for col in group_arr:\n            if(col == 'base_article_id' or col == 'article_id'):\n                continue\n            values = np.append(values, group_arr[col].value_counts().index[0])\n        values = np.insert(values,0,base_article_id)\n        \n        cache_df = pd.DataFrame([values], columns=self.columns)\n        \n        if(self.df_article_most.empty):\n            self.df_article_most = cache_df\n        else:          \n            self.df_article_most = pd.concat([self.df_article_most,cache_df])\n            \n    def build(self,arr,grouped_arr):\n        self.df_article_most = pd.DataFrame()\n        self.base_arr = arr\n        self.columns = self.base_arr.columns.tolist()\n        self.columns = ['most_' + sub for sub in self.columns]\n        del self.columns[1]\n        self.columns[0] = 'article_id'\n\n        for base_article_id in arr['base_article_id']:\n         \n            self.iteration(base_article_id,grouped_arr)\n       \n\n        \nb = ArticleGroup()\nb.build(df_baseArticleOnly,df_articleAndCustomerMost_Grouped)\nb.df_article_most\n\n\n","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_oneArticleMost = b.df_article_most","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_oneArticleMost.columns","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<font size=\"3\">For example, the most repetitive features of the products bought by the people who bought the product with the id number 685687003.</font>","metadata":{}},{"cell_type":"code","source":"df_oneArticleMost[df_oneArticleMost['article_id'] == \"685687003\"]","metadata":{"scrolled":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"\n\n<font size=\"3\">For example, people who bought the product with the number 685687004 bought 520 different products.</font>","metadata":{}},{"cell_type":"code","source":"df_articleAndCustomerMost_Grouped['article_id'].count().sort_values('article_id', ascending=False)","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"\n<font size=\"3\">  We have 4 DFs; <br>\n    <br>\n                 1. DF with customers and the products they bought.<br>\n                 2. DF with features of products and customers who have bought it<br>\n                 3. DF with the features of the products they have bought next to the customers who have bought a product<br>\n                 4. The DF we will use to fill in the missing values during the creation of our model values<br>\n</font>     ","metadata":{}},{"cell_type":"code","source":"display(df_TRS_TR_OneCus)\ndisplay(df_groupedBuyingArticlesValues)\ndisplay(df_oneArticleMost)\ndisplay(df_articles_all)","metadata":{"scrolled":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id='articlesDFMerge' style=\"color:black\" /></a>\n# Combining ArticlesCustomerValues with ArticlesMost DF","metadata":{}},{"cell_type":"markdown","source":"<font size=\"3\">Now let's start combining our 2nd and 3rd DF for classification operations.</font>","metadata":{}},{"cell_type":"code","source":"df_oneArticleMost['article_id']=df_oneArticleMost['article_id'].astype('int64')\n\ndf_all_values = pd.merge(df_groupedBuyingArticlesValues,df_oneArticleMost,on='article_id'\n                         ,how='left')\ndf_all_values_CP = df_all_values.copy()\ndf_all_values","metadata":{"scrolled":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id='articlesDFFeatureEngineering' style=\"color:black\" /></a>\n# Articles DF's Feature Engineering","metadata":{}},{"cell_type":"markdown","source":"<font size=\"3\">Now we have the DF of each product with the features of the customers who bought it and the features of the products they bought with it.</font>","metadata":{}},{"cell_type":"markdown","source":"\n<font size=\"3\">I'm putting our model variables into the process of standardizing. If you've come this far, I don't think I need to tell you about them. :)</font>","metadata":{}},{"cell_type":"code","source":"from sklearn import preprocessing \npd.set_option('display.float_format', lambda x: '%.2f' % x)\nnp.set_printoptions(formatter={'float': lambda x: \"{0:0.6f}\".format(x)})","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"\n<font size=\"3\">The reason for keeping the min, max values here is because when I want the class for a customer, these values will be needed when normalizing their age and price.</font>","metadata":{}},{"cell_type":"code","source":"max_age = df_all_values['age'].max()","metadata":{"scrolled":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"min_age = df_all_values['age'].min()","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"max_price = df_all_values['price'].max()","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"min_price = df_all_values['price'].min()","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"cache = np.array(df_all_values.loc[:,'age']).reshape(-1, 1)\nscaler = preprocessing.MinMaxScaler(feature_range = (0,1))\ndf_all_values['age'] = scaler.fit_transform(cache)\ndisplay(scaler.fit_transform(cache))","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"\ncache = np.array(df_all_values.loc[:,'price']).reshape(-1, 1)\nscaler = preprocessing.MinMaxScaler(feature_range = (0,1))\ndf_all_values['price'] = scaler.fit_transform(cache)\ndisplay(scaler.fit_transform(cache))","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<font size=\"3\">Since I will not use the article id value in the model and I want to see which class the product falls into, I index the article id value.</font>","metadata":{}},{"cell_type":"code","source":"df_all_values = df_all_values.set_index('article_id')","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_all_values","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<font size=\"3\">Since the product name value is variable for each product, it will not benefit us. Since we won't be using the product's purchase quantity for our model, I'm removing them.</font>","metadata":{}},{"cell_type":"code","source":"df_all_values_IncCustomerPrirce = df_all_values.copy()","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_all_values_CP2 = df_all_values.copy()\n\ndf_all_values = df_all_values.drop(['customer_id'],axis=1)\ndf_all_values = df_all_values.drop(['most_prod_name'],axis=1)\ndf_all_values\n\n","metadata":{"scrolled":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<font size=\"3\">Names of all our categorical variables</font>","metadata":{}},{"cell_type":"code","source":"for col in df_all_values:\n    if(col == 'price' or col == 'age' or col == 'CMS_ACTIVE' or col == 'CMS_EMPTY'\n       or col == 'CMS_LEFT_CLUB' or col == 'CMS_PRE_CREATE' or col == 'FNF_EMPTY'\n       or col == 'FNF_Monthly' or col == 'FNF_Regularly'):\n        continue\n    print(col)\n    print(df_all_values[col].unique())","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_all_values.columns","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<font size=\"3\">I create Dummies variables</font>","metadata":{}},{"cell_type":"markdown","source":"<font size=\"3\">The reason for not doing 'drop=First' at the time of creating dummies variables is that when the customer is guessing, the variable information here is not preserved during the creation of the dummies variable again. So a variable removed here may be in my table during the estimation phase. That's why I don't remove it.<br>\n</font>","metadata":{}},{"cell_type":"code","source":"df_all_values = pd.get_dummies(df_all_values, columns = [\"most_product_type_name\"], prefix = [\"MPTN\"])\ndf_all_values = pd.get_dummies(df_all_values, columns = [\"most_product_group_name\"], prefix = [\"MPGN\"])\ndf_all_values = pd.get_dummies(df_all_values, columns = [\"most_graphical_appearance_name\"], prefix = [\"GAN\"])\ndf_all_values = pd.get_dummies(df_all_values, columns = [\"most_colour_group_name\"], prefix = [\"CGN\"])\ndf_all_values = pd.get_dummies(df_all_values, columns = [\"most_perceived_colour_value_name\"], prefix = [\"PCVN\"])\ndf_all_values = pd.get_dummies(df_all_values, columns = [\"most_perceived_colour_master_name\"], prefix = [\"PCMN\"])\ndf_all_values = pd.get_dummies(df_all_values, columns = [\"most_department_name\"], prefix = [\"DN\"])\ndf_all_values = pd.get_dummies(df_all_values, columns = [\"most_index_name\"], prefix = [\"IN\"])\ndf_all_values = pd.get_dummies(df_all_values, columns = [\"most_index_group_name\"], prefix = [\"IGN\"])\ndf_all_values = pd.get_dummies(df_all_values, columns = [\"most_section_name\"], prefix = [\"SN\"])\ndf_all_values = pd.get_dummies(df_all_values, columns = [\"most_garment_group_name\"], prefix = [\"GGN\"])\n\ndf_articles_all = pd.get_dummies(df_articles_all, columns = [\"product_type_name\"], prefix = [\"MPTN\"])\ndf_articles_all = pd.get_dummies(df_articles_all, columns = [\"product_group_name\"], prefix = [\"MPGN\"])\ndf_articles_all = pd.get_dummies(df_articles_all, columns = [\"graphical_appearance_name\"], prefix = [\"GAN\"])\ndf_articles_all = pd.get_dummies(df_articles_all, columns = [\"colour_group_name\"], prefix = [\"CGN\"])\ndf_articles_all = pd.get_dummies(df_articles_all, columns = [\"perceived_colour_value_name\"], prefix = [\"PCVN\"])\ndf_articles_all = pd.get_dummies(df_articles_all, columns = [\"perceived_colour_master_name\"], prefix = [\"PCMN\"])\ndf_articles_all = pd.get_dummies(df_articles_all, columns = [\"department_name\"], prefix = [\"DN\"])\ndf_articles_all = pd.get_dummies(df_articles_all, columns = [\"index_name\"], prefix = [\"IN\"])\ndf_articles_all = pd.get_dummies(df_articles_all, columns = [\"index_group_name\"], prefix = [\"IGN\"])\ndf_articles_all = pd.get_dummies(df_articles_all, columns = [\"section_name\"], prefix = [\"SN\"])\ndf_articles_all = pd.get_dummies(df_articles_all, columns = [\"garment_group_name\"], prefix = [\"GGN\"])","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_all_values","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<font size=\"3\">We created the df_articles_all df with all the columns to fill in the missing columns when the customer information is created. For the reason I explained above.</font>","metadata":{}},{"cell_type":"code","source":"df_articles_all = df_articles_all.set_index('article_id')","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_articles_all = df_articles_all.drop(['prod_name'],axis=1)","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_articles_all","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<font size=\"3\">I have created my information that I will use to fill in missing columns and create those that are not. I will use this data in the next block.</font>","metadata":{}},{"cell_type":"code","source":"temp_articles = df_articles_all.reset_index(drop=True).copy()\nexample_modelData = temp_articles.iloc[:1].copy()\nexample_modelData.loc[:] = 0\nexample_modelData","metadata":{"scrolled":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"example_modelData","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"example_data = pd.concat([example_modelData]*df_all_values.shape[0])\ndifferent_cols = example_data.columns.difference(df_all_values.columns)\nexample_data = example_data[different_cols]\nexample_data = example_data.reset_index(drop=True)\ndf_all_values.reset_index(inplace=True)\ndf_all_values = pd.concat([df_all_values,example_data],axis=1)\ndf_all_values = df_all_values.set_index('article_id',drop=True)\ndf_all_values    ","metadata":{"scrolled":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<font size=\"3\">  We have 5 DFs; <br>\n    <br>\n                 1. DF with customers and the products they bought.<br>\n                 2. DF with features of products and customers who have bought it<br>\n                 3. DF with the features of the products they have bought next to the customers who have bought a product<br>\n                 4. The DF we will use to fill in the missing values during the creation of our model values <br>\n                 5. The DF I will use in the modeling phase <br>\n</font>     \n\n","metadata":{}},{"cell_type":"code","source":"display(df_TRS_TR_OneCus)\ndisplay(df_groupedBuyingArticlesValues)\ndisplay(df_oneArticleMost)\ndisplay(df_articles_all)\ndisplay(df_all_values)","metadata":{"scrolled":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"\n<a id='classification' style=\"color:black\" /></a>\n# Classification Operations","metadata":{}},{"cell_type":"markdown","source":"\n<font size=\"3\">We will do clustering, but since we have too many variables, I reduce them. I will find the number of clusters with my reduced variables.</font>","metadata":{}},{"cell_type":"code","source":"from sklearn.decomposition import PCA\nfrom sklearn.preprocessing import scale \nfrom sklearn.metrics import mean_squared_error, r2_score\nimport statsmodels.formula.api as smf\nfrom sklearn.cluster import KMeans\nfrom yellowbrick.cluster import KElbowVisualizer\nfrom sklearn.model_selection import train_test_split, GridSearchCV, cross_val_score\nfrom sklearn.metrics import confusion_matrix, accuracy_score, classification_report\nfrom sklearn.metrics import roc_auc_score,roc_curve\nfrom sklearn import preprocessing ","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"X = df_all_values.copy()","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<font size=\"3\">First, I gave the initial value as 19 variables, let's look at the results.</font>","metadata":{}},{"cell_type":"code","source":"pca = PCA(n_components=19)","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"X_reduced_train = pca.fit_transform(X)\nprint(pca.explained_variance_ratio_)","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<font size=\"3\">Our values are too little I am increasing our value.<font>","metadata":{}},{"cell_type":"code","source":"pca = PCA(n_components=400)","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"X_reduced_train = pca.fit_transform(X)\n# print(pca.explained_variance_ratio_)\nprint(np.cumsum(np.round(pca.explained_variance_ratio_, decimals = 4)*100)[0:135])","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<font size=\"3\">I can explain 95% variance with 135 variables. I will use this value.</font>","metadata":{}},{"cell_type":"code","source":"pca = PCA(n_components=135)\nX_reduced_train = pca.fit_transform(X)","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<font size=\"3\">As an example the post-PCA values of my line 1 variables.</font>","metadata":{}},{"cell_type":"code","source":"X_reduced_train[0:1,:]","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<font size=\"3\">I sort my products into clusters</font>","metadata":{}},{"cell_type":"code","source":"kmeans = KMeans()\nvisualizer = KElbowVisualizer(kmeans, k=(2,100))\nvisualizer.fit(X_reduced_train) \nvisualizer.poof()  ","metadata":{"scrolled":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<font size=\"3\">With Elbow we can find the ideal number of clusters. The method tells us that the value 31 is the ideal number of clusters. I choose the value 31. When I upload the file here and run it, there may be different values. I found 31 when I did. You accept the k value you see here.</font>","metadata":{}},{"cell_type":"code","source":"kmeans = KMeans()","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"cluster = KMeans(n_clusters = 31)\nk_fit = kmeans.fit(X_reduced_train)\ncluster = k_fit.labels_\ncluster","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"X['cluster'] = cluster","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_all_values_clustered_checkpoint = X.copy()\ndf_all_values_clustered = X.copy()","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<font size=\"3\">\n     We say:\n      Person A buys product X with the following features.\n      In the algorithm, it groups people and says:\n      Your characteristics are similar to the person who bought the Y product.\n      You like product</font>","metadata":{}},{"cell_type":"code","source":"df_all_values_clustered","metadata":{"scrolled":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<font size=\"3\">I am creating DF showing product, number of clusters and number of fetches. I will use it in the estimation phase.</font>","metadata":{}},{"cell_type":"code","source":"# df_all_values_clustered['']\ndf_article_cluster = pd.DataFrame({\n    'cluster' : df_all_values_clustered['cluster']\n})\ndf_article_cluster = df_article_cluster.reset_index()\ndf_article_cluster\n\n\ndf_article_cluster = pd.merge(df_article_cluster,df_groupedBuyingArticlesValues.loc[:,['customer_id','article_id']],on='article_id'\n                                         ,how='left')\ncolumns = df_article_cluster.columns.tolist()\ncolumns[2] = 'buying_count'\ndf_article_cluster.columns = columns\ndf_article_cluster\n","metadata":{"scrolled":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<font size=\"3\">Now I create my Train-Test values to build my model</font>","metadata":{}},{"cell_type":"code","source":"X = df_all_values_clustered.drop(['cluster'] , axis=1)","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"y = df_all_values_clustered['cluster']","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"X_train, X_test, y_train, y_test = train_test_split(X, \n                                                    y, \n                                                    test_size=0.25, \n                                                    random_state=42)\n\nprint(\"X_train\", X_train.shape)\n\nprint(\"y_train\",y_train.shape)\n\nprint(\"X_test\",X_test.shape)\n\nprint(\"y_test\",y_test.shape)","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"X_train","metadata":{"scrolled":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id='model' style=\"color:black\" /></a>\n# Model Building Stages","metadata":{}},{"cell_type":"markdown","source":"<font size=\"3\">I'm Creating a Model. I will be using LightGBM. This is the best fit for my computer in terms of performance.</font>","metadata":{}},{"cell_type":"code","source":"from lightgbm import LGBMClassifier","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"lgbm_model = LGBMClassifier().fit(X_train, y_train)","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"y_pred = lgbm_model.predict(X_test)\naccuracy_score(y_test, y_pred)","metadata":{"scrolled":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<font size=\"3\">I was surprised that our value turned out so well. As it is a habit, I will still do a model tuning.</font>","metadata":{}},{"cell_type":"code","source":"# Example\n# lgbm_params = {\n#         'n_estimators': [100, 500, 1000, 2000],\n#         'subsample': [0.6, 0.8, 1.0],\n#         'max_depth': [3, 4, 5,6],\n#         'learning_rate': [0.1,0.01,0.02,0.05],\n#         \"min_child_samples\": [5,10,20]}","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"lgbm_params = {\n        'n_estimators': [100, 500],\n        'subsample': [0.6],\n        'max_depth': [3, 4],\n        'learning_rate': [0.02,0.05],\n        \"min_child_samples\": [10,20]}","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"lgbm = LGBMClassifier()\n\nlgbm_cv_model = GridSearchCV(lgbm, lgbm_params, \n                             cv = 10, \n                             n_jobs = -1, \n                             verbose = 2)\n\n","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"lgbm_cv_model.fit(X_train, y_train)","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"{'learning_rate': 0.02,\n 'max_depth': 4,\n 'min_child_samples': 20,\n 'n_estimators': 500,\n 'subsample': 0.6}","metadata":{}},{"cell_type":"code","source":"lgbm_cv_model.best_params_","metadata":{"scrolled":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<font size=\"3\">We have increased our value by a very small amount.</font>","metadata":{}},{"cell_type":"code","source":"y_pred = lgbm_cv_model.predict(X_test)\naccuracy_score(y_test, y_pred)","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<font size=\"3\">The DFs we have prepared so far.</font>","metadata":{}},{"cell_type":"code","source":"display(df_TRS_TR_OneCus)\ndisplay(df_groupedBuyingArticlesValues)\ndisplay(df_oneArticleMost)\ndisplay(df_all_values)\ndisplay(df_articles_all)\ndisplay(df_article_cluster)\n","metadata":{"scrolled":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id='customerSubmission' style=\"color:black\" /></a>\n# Estimation Operations\n","metadata":{}},{"cell_type":"markdown","source":"<font size=\"3\">We are in the estimation phase.</font>","metadata":{}},{"cell_type":"markdown","source":"<font size=\"3\">I am preparing the DF that I will use to generate the information of a customer in the estimation phase. We could have done this with df_customer_root without provisioning.</font>","metadata":{}},{"cell_type":"code","source":"df_customersWithProperties = pd.merge(df_transactionsTrain_root,df_customers_root,on='customer_id'\n                                         ,how='left')\ndf_customersWithProperties = df_customersWithProperties.sort_values('t_dat', ascending=False)\ndf_customersWithProperties","metadata":{"scrolled":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<font size=\"3\">Receiving products purchased by a customer as an example</font>","metadata":{}},{"cell_type":"markdown","source":"<font size=\"3\">I'll use my customer on line 2 as an example</font>","metadata":{}},{"cell_type":"code","source":"df_TRS_TR_OneCus.iloc[2,:]['customer_id']","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"articles = df_TRS_TR_OneCus.iloc[2,:]['articles']\narticles = articles.split(',')\ndisplay(articles)\ndf_ad = pd.DataFrame(articles,columns = ['articles'])\ndisplay(df_ad)","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_ad['articles']=df_ad['articles'].astype('int64')","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_reduced_article[df_reduced_article['article_id'].isin(df_ad['articles'].to_numpy())]","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<font size=\"3\">I am creating my class that prepares the product information that a customer receives that I will use in the estimation phase.</font>","metadata":{}},{"cell_type":"code","source":"\n\nclass CustomerMostArticles:\n    \n    df_base = None\n    df_customersWithArticlesList = None\n    df_reduced_article = None\n    customer_id = None\n    \n    def __init__(self,customer_id,customersWithArticlesList,article_list):\n        self.customer_id = customer_id;\n        self.customersWithArticlesList = customersWithArticlesList.copy()\n        self.df_reduced_article = article_list.copy()\n        \n        \n        \n    def build(self):\n\n    \n        row = self.customersWithArticlesList[self.customersWithArticlesList['customer_id'] == self.customer_id]\n\n        df_first = None\n        df_others = None\n        df_others_grouped = None\n        df_customersWithArticlesList = None\n        customerArticles = self.customerArticles(row)\n        customerArticlesDetails = self.getCustomerArticlesDetails(customerArticles)\n        customerArticlesDetails = self.dropRedundant(customerArticlesDetails)\n            \n#         print(customerArticlesDetails.shape)\n        if(customerArticlesDetails.shape[0] == 1):\n            \n            df_first  = customerArticlesDetails.iloc[0:1]\n            df_first = self.createFirst(df_first)\n            df_first = df_first.reset_index(drop=True)\n            return df_first\n        else:\n            \n            df_first  = customerArticlesDetails.iloc[0:1]\n            df_others = customerArticlesDetails[1:len(customerArticlesDetails)]\n            df_others_grouped = df_others.groupby([\"article_id\"],as_index=False)\n            df_customersMost = self.getMost(df_others)\n            df_first = self.createFirst(df_first)\n            articles = pd.concat([df_first,df_customersMost]).reset_index(drop=True)\n            return articles\n            \n            \n        \n    def customerArticles(self,row):\n        articles = row['articles'].tolist()\n        articles = articles[0].split(',')\n        for i in range(0, len(articles)):\n            articles[i] = int(articles[i])\n        return articles\n        \n        \n    def getCustomerArticlesDetails(self,articlesList):    \n               \n        return df_reduced_article[df_reduced_article['article_id'].isin(articlesList)]\n        \n    def dropRedundant(self,df):\n        return df.drop(['product_code','product_type_no','graphical_appearance_no',\n                                     'colour_group_code','perceived_colour_value_id',\n                                     'perceived_colour_master_id','department_no','prod_name',\n                                     'index_group_no','index_code','section_no','garment_group_no','detail_desc'],axis=1,errors='ignore')\n    \n    def getMost(self,others_arr):\n        df_article_most = pd.DataFrame()\n        columns = others_arr.columns.tolist()\n        del columns[0]\n        columns = ['most_' + sub for sub in columns]\n        values = np.array([])\n        for col in others_arr:\n            if(col == 'base_article_id' or col == 'article_id'):\n                continue\n            values = np.append(values, others_arr[col].value_counts().index[0])\n        cache_df = pd.DataFrame([values],columns = columns)\n        return cache_df\n    \n    def createFirst(self,first_row):\n       \n        first_row = first_row.drop(['article_id'], axis=1)\n        columns = first_row.columns.tolist()\n        first_row.columns = ['most_' + sub for sub in columns]\n        \n        values = np.array([])\n        for col in first_row:\n            \n            if(col == 'base_article_id' or col == 'article_id'):\n                continue\n            values = np.append(values, first_row[col].value_counts().index[0])\n            \n        \n        cache_df = pd.DataFrame([values],columns = first_row.columns)\n        return cache_df\n    ","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<font size=\"3\">For example, a customer's information</font>","metadata":{}},{"cell_type":"markdown","source":"<font size=\"3\">Get a sample customer id - Use it you want to try it yourself.</font>","metadata":{}},{"cell_type":"markdown","source":"\n<font size=\"3\">Pay attention to whether the customer you sent is within the area I specified with the head.</font>","metadata":{}},{"cell_type":"code","source":"print(df_customersWithProperties.sample().iloc[0]['customer_id'])\ndf_customersWithProperties[df_customersWithProperties.customer_id == df_customersWithProperties.sample().iloc[0]['customer_id']]","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_customerArticleMost = CustomerMostArticles(\"00083cda041544b2fbb0e0d2905ad17da7cf1007526fb4c73235dccbbc132280\",\n                                              df_TRS_TR_OneCus.head(5),df_reduced_article).build()\ndf_customerArticleMost","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<font size=\"3\">I'm creating my class that prepares the customer information that I will use in the prediction phase of a customer.</font>","metadata":{}},{"cell_type":"code","source":"class CustomerInfo:\n    \n    customerWithProperties = None\n    customersWithProperties = None\n    customer_id = None\n    \n    def __init__(self,customer_id,customersWithProperties):\n        self.customersWithProperties = customersWithProperties.copy()\n        self.customerWithProperties = customersWithProperties.drop_duplicates(subset=['customer_id'])\n        self.customer_id = customer_id\n        \n    def build(self):\n        \n        row = self.customerWithProperties[self.customerWithProperties['customer_id'] == self.customer_id]\n        customer_id = self.customer_id\n        \n#       customer_id = \"2648daf1914c9ab1404629a8dfe1f9e58ce5fdd3da5a0d00816c71ea70999825\"\n\n\n        customerData = self.customersWithProperties[self.customersWithProperties.customer_id == customer_id]\n        customerData = customerData.reset_index(drop=True)\n#         display(customerData['age'])\n        if(customerData.shape[0] > 1):\n            customerData['age'] = customerData['age'].mean()\n            customerData['price'] = customerData['price'].mean()\n        customerData = customerData.drop(['t_dat','sales_channel_id','postal_code'],axis=1,errors='ignore')\n        customerData = customerData.iloc[0:1]\n        customerData = pd.concat([customerData]*2).reset_index(drop=True)\n        return customerData\n#         display(customerData)","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<font size=\"3\">For example, a customer's information. I mentioned above why we made 2 lines.</font>","metadata":{}},{"cell_type":"code","source":"df_customer_info = CustomerInfo(\"2648daf1914c9ab1404629a8dfe1f9e58ce5fdd3da5a0d00816c71ea70999825\",df_customersWithProperties).build()\ndf_customer_info","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"pd.set_option('display.float_format', lambda x: '%.2f' % x)\nnp.set_printoptions(formatter={'float': lambda x: \"{0:0.6f}\".format(x)})","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<font size=\"3\">Display of information tables to be combined as an example.</font>","metadata":{}},{"cell_type":"code","source":"df_customer_info","metadata":{"scrolled":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_customerArticleMost","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<font size=\"3\">I am preparing my class where the merge operation will be done.</font>","metadata":{}},{"cell_type":"code","source":"class CustomerDataCreate:\n    \n    df_all_values = None\n    customer_df = None\n    articles_df = None\n    max_price = None\n    min_price = None\n    max_age = None\n    min_age = None\n    model_data = None\n    example_data = None\n    \n    def __init__(self,customer_data,articles_data,max_price,min_price,max_age,min_age,example_data):\n        \n        self.customer_df = customer_data\n        self.articles_df = articles_data\n        self.max_price = max_price\n        self.min_price = min_price\n        self.max_age = max_age\n        self.min_age = min_age\n#         self.build()\n        self.example_data = example_data\n        \n        \n        \n    def build(self):\n\n        self.df_all_values = pd.concat([self.customer_df,self.articles_df],axis=1)  \n\n        self.df_all_values['age'] = self.df_all_values['age'].apply(self.normalizeAge)\n        self.df_all_values['price'] = self.df_all_values['price'].apply(self.normalizePrice)\n        return self.df_all_values\n    \n    def createModelData(self):\n        \n        model_data = self.df_all_values.drop(['customer_id'],axis=1)\n        model_data = pd.get_dummies(model_data, columns = [\"most_product_type_name\"], prefix = [\"MPTN\"])\n        \n        model_data = pd.get_dummies(model_data, columns = [\"most_product_group_name\"], prefix = [\"MPGN\"])\n        model_data = pd.get_dummies(model_data, columns = [\"most_graphical_appearance_name\"], prefix = [\"GAN\"])\n        model_data = pd.get_dummies(model_data, columns = [\"most_colour_group_name\"], prefix = [\"CGN\"])\n        model_data = pd.get_dummies(model_data, columns = [\"most_perceived_colour_value_name\"], prefix = [\"PCVN\"])\n        model_data = pd.get_dummies(model_data, columns = [\"most_perceived_colour_master_name\"], prefix = [\"PCMN\"])\n        model_data = pd.get_dummies(model_data, columns = [\"most_department_name\"], prefix = [\"DN\"])\n        model_data = pd.get_dummies(model_data, columns = [\"most_index_name\"], prefix = [\"IN\"])\n        model_data = pd.get_dummies(model_data, columns = [\"most_index_group_name\"], prefix = [\"IGN\"])\n        model_data = pd.get_dummies(model_data, columns = [\"most_section_name\"], prefix = [\"SN\"])\n        model_data = pd.get_dummies(model_data, columns = [\"most_garment_group_name\"], prefix = [\"GGN\"])\n        \n\n        # Using example_modelData\n        \n        \n        self.example_data = pd.concat([self.example_data]*2)\n\n        different_cols = self.example_data.columns.difference(model_data.columns)\n        self.example_data = self.example_data[different_cols]\n        self.example_data = self.example_data.reset_index(drop=True)\n        \n        model_data = pd.concat([model_data,self.example_data],axis=1)\n        model_data = model_data.set_index('article_id',drop=True)\n        return model_data\n\n        \n    def normalizeAge(self,val):\n        return (val-self.min_age)/(self.max_age-self.min_age)\n        \n    def normalizePrice(self,val):\n        return (val-self.min_price)/(self.max_price-self.min_price)\n        \n        \ncustomerData = CustomerDataCreate(df_customer_info,df_customerArticleMost,max_price,min_price,max_age,min_age,example_modelData)\noneCustomerData = customerData.build()\noneCustomerModelData = customerData.createModelData()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<font size=\"3\">The data representation I created above that I will use to complete the missing information when the dummy variable is created. Thanks to this information, I fill in the missing information for my model after the dummy variables are created.</font>","metadata":{}},{"cell_type":"code","source":"example_modelData","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<font size=\"3\">Display of a customer's information combined and to be inserted into the model</font>","metadata":{}},{"cell_type":"code","source":"display(oneCustomerData)","metadata":{"scrolled":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"display(oneCustomerModelData)","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<font size=\"3\">Estimate of a customer as an example</font>","metadata":{}},{"cell_type":"code","source":"y_pred = lgbm_cv_model.predict(oneCustomerModelData)\n# accuracy_score(y_test, y_pred)","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"y_pred","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_article_cluster[df_article_cluster.cluster == 1]","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<font size=\"3\">I'm writing a main class that will combine all the classes I wrote above and make predictions based on the customer information given.</font>","metadata":{}},{"cell_type":"code","source":"class createCustomerPrediction:\n    \n    \n    instance_CustomerMostArticles = None\n    instance_CustomerInfo = None\n    instance_CustomerDataCreate = None\n    model = None\n    articleWithCluster = None\n    customerWithPredictions = None\n\n    def __init__(self,CustomerMostArticles,CustomerInfo,CustomerDataCreate,model,articleWithCluster):\n        self.instance_CustomerMostArticles = CustomerMostArticles\n        self.instance_CustomerInfo = CustomerInfo\n        self.instance_CustomerDataCreate = CustomerDataCreate\n        self.articleWithCluster = articleWithCluster\n        self.model = model\n        self.customerWithPredictions = pd.DataFrame()\n    \n    def letsPrediction(self,transactionsOneCustomer,articleList,customersWithProperties,\n                      max_price,min_price,max_age,min_age,example_modelData):\n        \n        transactionsOneCustomer = transactionsOneCustomer.reset_index(drop=True)\n        \n        for index in range(len(transactionsOneCustomer.head(5))):\n            row = transactionsOneCustomer.loc[index]\n            \n            df_customerArticleMost = self.instance_CustomerMostArticles(row['customer_id'],transactionsOneCustomer\n                                                                        ,articleList).build()\n            df_customer_info = self.instance_CustomerInfo(row['customer_id'],customersWithProperties).build()\n            \n            customerData = self.instance_CustomerDataCreate(df_customer_info,\n                                              df_customerArticleMost,max_price,min_price,max_age,min_age,example_modelData)        \n            oneCustomerData = customerData.build()\n            oneCustomerModelData = customerData.createModelData()\n            prediction = self.createPrediction(oneCustomerModelData)\n\n            prediction = ','.join(map(str, prediction))\n            cache_df = pd.DataFrame([[row['customer_id'],prediction]], columns=['customer_id','predictions'])\n        \n            if(self.customerWithPredictions.empty):\n                self.customerWithPredictions = cache_df\n            else:          \n                self.customerWithPredictions = pd.concat([self.customerWithPredictions,cache_df])\n                \n                \n    def getPredictions(self):\n        self.customerWithPredictions = self.customerWithPredictions.reset_index(drop=True)\n        return self.customerWithPredictions\n    \n    def createPrediction(self,modelData):\n        prediction = self.model.predict(modelData)\n        return self.getArticles(prediction)\n    \n    def getArticles(self,prediction):\n        self.articleWithCluster\n        first_items = self.articleWithCluster[self.articleWithCluster.cluster == prediction[0]][:4]['article_id'].values\n        if(prediction[0] == prediction[1]):\n            last_items = self.articleWithCluster[self.articleWithCluster.cluster == prediction[1]][4:8]['article_id'].values\n        else:\n            last_items = self.articleWithCluster[self.articleWithCluster.cluster == prediction[1]][:3]['article_id'].values\n\n        return [*first_items,*last_items]\n            \n    \n    \n    \n","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<font size=\"3\">My predictions :)</font>","metadata":{}},{"cell_type":"code","source":"prediction = createCustomerPrediction(CustomerMostArticles,CustomerInfo,CustomerDataCreate,lgbm_cv_model,df_article_cluster)\nprediction.letsPrediction(df_TRS_TR_OneCus,df_reduced_article,df_customersWithProperties,\n                      max_price,min_price,max_age,min_age,example_modelData)\nprediction.getPredictions()","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<font size=\"3\">Our table of products purchased by the customer</font>","metadata":{}},{"cell_type":"code","source":"df_TRS_TR_OneCus","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<font size=\"3\">As an example, let's show the products bought by 1 customer and our estimates.</font>","metadata":{}},{"cell_type":"code","source":"df_TRS_TR_OneCus.loc[0]['customer_id']","metadata":{"scrolled":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import matplotlib.image as mpimg\nPATH = \"../input/h-and-m-personalized-fashion-recommendations/images/\"\n","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<font size=\"3\">In the last section, I used index 0 instead of id. Don't forget to fix it when it works :)</font>\n","metadata":{}},{"cell_type":"code","source":"def showPrediction(customer_id,predictionsDF,transactionsDF, rows=2, columns=7, figsize=(20,10)):\n\n    predictions = predictionsDF.loc[0]['predictions'].split(\",\")\n    bought_items = transactionsDF.loc[0]['articles'].split(\",\")\n    f, ax = plt.subplots(rows, columns, figsize=figsize)\n    for i in range(rows):\n        index = 0\n        for j in range(columns):\n            if i==0:\n                try:\n                    img = mpimg.imread(f'{PATH}0{str(bought_items[index])[:2]}/0{int(bought_items[index])}.jpg')\n                    ax[i,j].imshow(img)\n                    ax[i,j].set_xticks([], [])\n                    ax[i,j].set_yticks([], [])\n                    ax[i,j].grid(False)\n                    ax[i,j].set_title(\"Bought\")\n                    index += 1\n                except IndexError:\n                    continue\n            else:\n                try:\n                    img = mpimg.imread(f'{PATH}0{str(predictions[index])[:2]}/0{int(predictions[index])}.jpg')\n                    ax[i,j].imshow(img)\n                    ax[i,j].set_xticks([], [])\n                    ax[i,j].set_yticks([], [])\n                    ax[i,j].grid(False)\n                    ax[i,j].set_title(\"Prediction\")\n                    index += 1\n                except IndexError:\n                    continue\n                        \n    plt.tight_layout()\n    plt.show()","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"showPrediction(\"000058a12d5b43e67d225668fa1f8d618c13dc232df0cad8ffe7ad4a1091e318\",prediction.getPredictions(),df_TRS_TR_OneCus)","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<font size=\"4\"> Thank you for reading. :)</font>","metadata":{}},{"cell_type":"markdown","source":"<font size=\"4\"> Additional Note: </font>\n<font size=\"4\"> There are many libraries and methods to beautify prediction information. The purpose of this example is to realize a solution with a structure whose algorithm and structure we have established ourselves.</font>","metadata":{}},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]}]}