{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"## The reason why I wrote this notebook\n\nIn this discussions, [Addressing common questions and what the competition is really about](https://www.kaggle.com/c/h-and-m-personalized-fashion-recommendations/discussion/307288) and [Care to share?](https://www.kaggle.com/competitions/h-and-m-personalized-fashion-recommendations/discussion/314458), they mentioned that generating candidates is important to improve score. <br> However, I cannot find some good notebooks for scoring the candidate generation. So I made it!","metadata":{}},{"cell_type":"markdown","source":"1. I evaluate the candidate generation with Local CV. Please refer to [here](https://www.kaggle.com/competitions/h-and-m-personalized-fashion-recommendations/discussion/308919) for making Local CV.\n2. I evaluate the candidate generation with 2 metrics, referenced by [@jacob34's](https://www.kaggle.com/competitions/h-and-m-personalized-fashion-recommendations/discussion/314458), `Recall` and `Multiple Factor`. ","metadata":{}},{"cell_type":"code","source":"import numpy as np\nimport pandas as pd\n\nimport cudf","metadata":{"execution":{"iopub.status.busy":"2022-04-23T02:45:31.379607Z","iopub.execute_input":"2022-04-23T02:45:31.379799Z","iopub.status.idle":"2022-04-23T02:45:34.779354Z","shell.execute_reply.started":"2022-04-23T02:45:31.379775Z","shell.execute_reply":"2022-04-23T02:45:34.778618Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Prepare the Local CV","metadata":{}},{"cell_type":"code","source":"%%time\n\ntransactions = cudf.read_csv(\"../input/h-and-m-personalized-fashion-recommendations/transactions_train.csv\")\ntransactions.t_dat = pd.to_datetime(transactions.t_dat.to_pandas())\ntransactions[\"week\"] = 104 - (transactions.t_dat.max() - transactions.t_dat).dt.days // 7\n\nUSE_WEEKS = 5\nTEST_WEEK = 104\n\nvalid = transactions[transactions['week'] == TEST_WEEK][['customer_id', 'article_id']].to_pandas()\ntransactions = transactions[(transactions.week > TEST_WEEK - USE_WEEKS) & (transactions.week < TEST_WEEK)]  ","metadata":{"execution":{"iopub.status.busy":"2022-04-23T02:45:34.780985Z","iopub.execute_input":"2022-04-23T02:45:34.781232Z","iopub.status.idle":"2022-04-23T02:46:21.752136Z","shell.execute_reply.started":"2022-04-23T02:45:34.781198Z","shell.execute_reply":"2022-04-23T02:46:21.751371Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Generate Candidates","metadata":{}},{"cell_type":"markdown","source":"I made simple candidates by two methods.\n\n1. `previous_week` : Previous week history\n2. `previous_week_top` : (Previous week Popular Top 12 Products) x Customers","metadata":{}},{"cell_type":"code","source":"previous_week = transactions[transactions['week'] == 103][['customer_id', 'article_id']].to_pandas()\ntop_products = pd.DataFrame(data=transactions[transactions['week'] == 103].to_pandas().value_counts('article_id').iloc[:200].index.tolist(),\n                            columns=['article_id'])\nprevious_week_top = transactions[['customer_id']].drop_duplicates().to_pandas().merge(top_products, how='cross')\ncand = pd.concat([previous_week, previous_week_top]).drop_duplicates()","metadata":{"execution":{"iopub.status.busy":"2022-04-23T02:46:21.753374Z","iopub.execute_input":"2022-04-23T02:46:21.753699Z","iopub.status.idle":"2022-04-23T02:46:42.710738Z","shell.execute_reply.started":"2022-04-23T02:46:21.753658Z","shell.execute_reply":"2022-04-23T02:46:42.709990Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Score","metadata":{}},{"cell_type":"code","source":"def score(actual, predict):\n    act_tot = len(actual)\n    pre_tot = len(predict)\n    correct = actual.merge(predict, on=['customer_id', 'article_id'], how='inner').shape[0]\n    print(f\"[+] Recall = {correct/act_tot*100:.1f}% ({correct}/{act_tot})\")\n    print(f\"[+] Multiple Factor = {pre_tot//correct} ({pre_tot}/{correct})\")","metadata":{"execution":{"iopub.status.busy":"2022-04-23T02:46:42.712448Z","iopub.execute_input":"2022-04-23T02:46:42.712697Z","iopub.status.idle":"2022-04-23T02:46:42.719615Z","shell.execute_reply.started":"2022-04-23T02:46:42.712664Z","shell.execute_reply":"2022-04-23T02:46:42.718859Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"score(valid, cand)","metadata":{"execution":{"iopub.status.busy":"2022-04-23T02:46:42.720932Z","iopub.execute_input":"2022-04-23T02:46:42.721420Z","iopub.status.idle":"2022-04-23T02:46:59.834847Z","shell.execute_reply.started":"2022-04-23T02:46:42.721383Z","shell.execute_reply":"2022-04-23T02:46:59.834107Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**If you have some good idea for generating candidates, let's talk together!**\n\n**If this notebook was good for you, Please Upvote!**","metadata":{}}]}