{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"lunana  \nlast update 2022 04 28  \nゆっくりしていってね！  ","metadata":{}},{"cell_type":"markdown","source":"# Trends\n## EDA\nhttps://www.kaggle.com/lunapandachan/h-m-eda-english","metadata":{}},{"cell_type":"markdown","source":"**霊夢:今日はByfone氏のnotebookを高速化するよ。  \n魔理沙:元々は2時間以上かかる。**\n\n**Reimu: Today is Byfone's speed up exam.  \nMarisa: Over two hours to about a hour.**  \n\nhttps://www.kaggle.com/code/byfone/h-m-trending-products-weekly","metadata":{}},{"cell_type":"markdown","source":"**霊夢:データが多いのでtest=Trueの時は1000個のデータを使うよ。**  \n\n**Reimu: Since there is a lot of data, 1000 data will be used when test = True.**","metadata":{}},{"cell_type":"code","source":"test=False","metadata":{"execution":{"iopub.status.busy":"2022-04-29T13:06:10.888348Z","iopub.execute_input":"2022-04-29T13:06:10.888739Z","iopub.status.idle":"2022-04-29T13:06:10.911065Z","shell.execute_reply.started":"2022-04-29T13:06:10.888628Z","shell.execute_reply":"2022-04-29T13:06:10.910416Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import numpy as np\nimport pandas as pd\n\nfrom math import sqrt\nfrom pathlib import Path\nfrom tqdm import tqdm\nimport time\nimport gc\ntqdm.pandas()","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","execution":{"iopub.status.busy":"2022-04-29T13:06:10.915066Z","iopub.execute_input":"2022-04-29T13:06:10.915789Z","iopub.status.idle":"2022-04-29T13:06:10.919932Z","shell.execute_reply.started":"2022-04-29T13:06:10.915755Z","shell.execute_reply":"2022-04-29T13:06:10.919225Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"data_path = Path('../input/h-and-m-personalized-fashion-recommendations/')\nN = 12","metadata":{"execution":{"iopub.status.busy":"2022-04-29T13:06:10.939783Z","iopub.execute_input":"2022-04-29T13:06:10.940343Z","iopub.status.idle":"2022-04-29T13:06:10.946156Z","shell.execute_reply.started":"2022-04-29T13:06:10.940304Z","shell.execute_reply":"2022-04-29T13:06:10.945238Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**魔理沙:N=12は予測アイテムの最大数だぜ。**\n\n**Marisa: N = 12 is the maximum number of predictable items.**","metadata":{}},{"cell_type":"markdown","source":"### Read the train data\ntrainデータの読み込み","metadata":{}},{"cell_type":"code","source":"df = pd.read_csv(data_path / 'transactions_train.csv',\n                 usecols = ['t_dat', 'customer_id', 'article_id'],\n                 dtype={'article_id': str})\n\nif test :\n    df=df[:1000]\n\ndf['t_dat'] = pd.to_datetime(df['t_dat'])\nlast_ts = df['t_dat'].max()","metadata":{"execution":{"iopub.status.busy":"2022-04-29T13:06:10.973378Z","iopub.execute_input":"2022-04-29T13:06:10.974002Z","iopub.status.idle":"2022-04-29T13:07:24.023106Z","shell.execute_reply.started":"2022-04-29T13:06:10.973958Z","shell.execute_reply":"2022-04-29T13:07:24.022274Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**霊夢:customers.csvとarticles.csvは使わないの？？  \n魔理沙:imageも使わないのかよ？？**　　\n\n**Reimu: Don't you use customers.csv and articles.csv??  \nMarisa: Don't you use image too??**","metadata":{}},{"cell_type":"markdown","source":"### Add the last day of billing week\n\n請求週の最終日を追加します","metadata":{}},{"cell_type":"markdown","source":"**霊夢:これは何でやってるのかな？\n魔理沙:必要あるの？**　　\n\n**Reimu: Why are you doing this?\nMarisa: Do you need it?**","metadata":{}},{"cell_type":"code","source":"import datetime\nlast_ts2= df['t_dat'].max()\nfirst_ts= df['t_dat'].min()","metadata":{"execution":{"iopub.status.busy":"2022-04-29T13:07:24.024888Z","iopub.execute_input":"2022-04-29T13:07:24.025189Z","iopub.status.idle":"2022-04-29T13:07:24.231785Z","shell.execute_reply.started":"2022-04-29T13:07:24.025149Z","shell.execute_reply":"2022-04-29T13:07:24.230921Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%time\ndf0=pd.DataFrame(index=[],columns=df.columns)\nwhile last_ts2>first_ts:\n    next_ts=last_ts2-datetime.timedelta(days=7)\n    df2=df[(last_ts2 >= df['t_dat'])&(df['t_dat'] >next_ts) ].copy()\n    df2['ldbw']=last_ts2\n    df0=pd.concat([df0,df2])\n    #print(last_ts)\n    last_ts2=next_ts","metadata":{"execution":{"iopub.status.busy":"2022-04-29T13:07:24.232939Z","iopub.execute_input":"2022-04-29T13:07:24.233187Z","iopub.status.idle":"2022-04-29T13:09:20.891842Z","shell.execute_reply.started":"2022-04-29T13:07:24.233156Z","shell.execute_reply":"2022-04-29T13:09:20.890891Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df0 = df0.sort_values('t_dat').reset_index(drop=True)\ndf0","metadata":{"execution":{"iopub.status.busy":"2022-04-29T13:09:20.894753Z","iopub.execute_input":"2022-04-29T13:09:20.895021Z","iopub.status.idle":"2022-04-29T13:09:28.657405Z","shell.execute_reply.started":"2022-04-29T13:09:20.89499Z","shell.execute_reply":"2022-04-29T13:09:28.65685Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df","metadata":{"execution":{"iopub.status.busy":"2022-04-29T13:09:28.658472Z","iopub.execute_input":"2022-04-29T13:09:28.659064Z","iopub.status.idle":"2022-04-29T13:09:28.672516Z","shell.execute_reply.started":"2022-04-29T13:09:28.659034Z","shell.execute_reply":"2022-04-29T13:09:28.671673Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#df['ldbw'] = df['t_dat'].progress_apply(lambda d: last_ts - (last_ts - d).floor('7D'))\ndf=df0\ngc.collect()","metadata":{"execution":{"iopub.status.busy":"2022-04-29T13:09:28.674186Z","iopub.execute_input":"2022-04-29T13:09:28.674746Z","iopub.status.idle":"2022-04-29T13:09:28.852179Z","shell.execute_reply.started":"2022-04-29T13:09:28.674703Z","shell.execute_reply":"2022-04-29T13:09:28.851315Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import pickle\ndf.to_pickle('df.pkl')","metadata":{"execution":{"iopub.status.busy":"2022-04-29T13:09:28.853265Z","iopub.execute_input":"2022-04-29T13:09:28.853473Z","iopub.status.idle":"2022-04-29T13:09:45.264843Z","shell.execute_reply.started":"2022-04-29T13:09:28.853448Z","shell.execute_reply":"2022-04-29T13:09:45.263812Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Count the number of transactions per week \n\n週あたりのトランザクション数を数える","metadata":{}},{"cell_type":"code","source":"weekly_sales = df.drop('customer_id', axis=1).groupby(['ldbw', 'article_id']).count()\nweekly_sales = weekly_sales.rename(columns={'t_dat': 'count'})\nweekly_sales","metadata":{"execution":{"iopub.status.busy":"2022-04-29T13:09:45.266202Z","iopub.execute_input":"2022-04-29T13:09:45.26649Z","iopub.status.idle":"2022-04-29T13:09:54.637109Z","shell.execute_reply.started":"2022-04-29T13:09:45.266458Z","shell.execute_reply":"2022-04-29T13:09:54.636164Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df = df.join(weekly_sales, on=['ldbw', 'article_id'])","metadata":{"execution":{"iopub.status.busy":"2022-04-29T13:09:54.638574Z","iopub.execute_input":"2022-04-29T13:09:54.639043Z","iopub.status.idle":"2022-04-29T13:10:06.922617Z","shell.execute_reply.started":"2022-04-29T13:09:54.638998Z","shell.execute_reply":"2022-04-29T13:10:06.921798Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df.head()","metadata":{"execution":{"iopub.status.busy":"2022-04-29T13:10:06.92579Z","iopub.execute_input":"2022-04-29T13:10:06.926218Z","iopub.status.idle":"2022-04-29T13:10:06.938949Z","shell.execute_reply.started":"2022-04-29T13:10:06.926161Z","shell.execute_reply":"2022-04-29T13:10:06.93787Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df['count'].unique()","metadata":{"execution":{"iopub.status.busy":"2022-04-29T13:10:06.940505Z","iopub.execute_input":"2022-04-29T13:10:06.940907Z","iopub.status.idle":"2022-04-29T13:10:07.141931Z","shell.execute_reply.started":"2022-04-29T13:10:06.940863Z","shell.execute_reply":"2022-04-29T13:10:07.141098Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Let's assume that in the target week sales will be similar to the last week of the training data\n\n目標週の売上高がトレーニングデータの先週と同様になると仮定しましょう","metadata":{}},{"cell_type":"code","source":"weekly_sales = weekly_sales.reset_index().set_index('article_id')\nlast_day = last_ts.strftime('%Y-%m-%d')\n\ndf = df.join(\n    weekly_sales.loc[weekly_sales['ldbw']==last_day, ['count']],\n    on='article_id', rsuffix=\"_targ\")\n\ndf['count_targ'].fillna(0, inplace=True)\ndel weekly_sales","metadata":{"execution":{"iopub.status.busy":"2022-04-29T13:10:07.143425Z","iopub.execute_input":"2022-04-29T13:10:07.14388Z","iopub.status.idle":"2022-04-29T13:10:17.405579Z","shell.execute_reply.started":"2022-04-29T13:10:07.143836Z","shell.execute_reply":"2022-04-29T13:10:17.404823Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Calculate sales rate adjusted for changes in product popularity \n\n製品の人気の変化に合わせて調整された販売率を計算します","metadata":{}},{"cell_type":"code","source":"df['quotient'] = df['count_targ'] / df['count']","metadata":{"execution":{"iopub.status.busy":"2022-04-29T13:10:17.406773Z","iopub.execute_input":"2022-04-29T13:10:17.407016Z","iopub.status.idle":"2022-04-29T13:10:17.576041Z","shell.execute_reply.started":"2022-04-29T13:10:17.406988Z","shell.execute_reply":"2022-04-29T13:10:17.575426Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df.head()","metadata":{"execution":{"iopub.status.busy":"2022-04-29T13:10:17.577604Z","iopub.execute_input":"2022-04-29T13:10:17.577955Z","iopub.status.idle":"2022-04-29T13:10:17.592214Z","shell.execute_reply.started":"2022-04-29T13:10:17.577914Z","shell.execute_reply":"2022-04-29T13:10:17.591622Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Take supposedly popular products\n\nおそらく人気のある製品を取る","metadata":{}},{"cell_type":"code","source":"target_sales = df.drop('customer_id', axis=1).groupby('article_id')['quotient'].sum()\ngeneral_pred = target_sales.nlargest(N).index.tolist()\ndel target_sales","metadata":{"execution":{"iopub.status.busy":"2022-04-29T13:10:17.593436Z","iopub.execute_input":"2022-04-29T13:10:17.593883Z","iopub.status.idle":"2022-04-29T13:10:26.881297Z","shell.execute_reply.started":"2022-04-29T13:10:17.593845Z","shell.execute_reply":"2022-04-29T13:10:26.880198Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Fill in purchase dictionary\n\n購入辞書に記入する","metadata":{}},{"cell_type":"code","source":"purchase_dict = {}\n\nfor i in tqdm(df.index):\n    cust_id = df.at[i, 'customer_id']\n    art_id = df.at[i, 'article_id']\n    t_dat = df.at[i, 't_dat']\n\n    if cust_id not in purchase_dict:\n        purchase_dict[cust_id] = {}\n\n    if art_id not in purchase_dict[cust_id]:\n        purchase_dict[cust_id][art_id] = 0\n    \n    x = max(1, (last_ts - t_dat).days)\n\n    a, b, c, d = 2.5e4, 1.5e5, 2e-1, 1e3\n    y = a / np.sqrt(x) + b * np.exp(-c*x) - d\n\n    value = df.at[i, 'quotient'] * max(0, y)\n    purchase_dict[cust_id][art_id] += value","metadata":{"execution":{"iopub.status.busy":"2022-04-29T13:10:26.882803Z","iopub.execute_input":"2022-04-29T13:10:26.883029Z","iopub.status.idle":"2022-04-29T13:51:22.886206Z","shell.execute_reply.started":"2022-04-29T13:10:26.883003Z","shell.execute_reply":"2022-04-29T13:51:22.885043Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**魔理沙:個人の趣向に合わせて、購入履歴×トレンドのデータベースを作ったんだな。**\n\n**Marisa: I created a database of purchase history and trends to suit my personal taste.**","metadata":{}},{"cell_type":"markdown","source":"### Make a submission","metadata":{}},{"cell_type":"code","source":"sub = pd.read_csv(data_path / 'sample_submission.csv')\n\npred_list = []\nfor cust_id in tqdm(sub['customer_id']):\n    if cust_id in purchase_dict:\n        series = pd.Series(purchase_dict[cust_id])\n        series = series[series > 0]\n        l = series.nlargest(N).index.tolist()\n        if len(l) < N:\n            l = l + general_pred[:(N-len(l))]\n    else:\n        l = general_pred\n    pred_list.append(' '.join(l))\n\nsub['prediction'] = pred_list\nsub.to_csv('submission.csv', index=None)","metadata":{"execution":{"iopub.status.busy":"2022-04-29T13:51:22.887996Z","iopub.execute_input":"2022-04-29T13:51:22.888354Z","iopub.status.idle":"2022-04-29T14:08:23.575625Z","shell.execute_reply.started":"2022-04-29T13:51:22.88831Z","shell.execute_reply":"2022-04-29T14:08:23.574601Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**霊夢:今まで購入したモノの中で流行にあったものを再購入したリストだぜ。  \n魔理沙:そりゃあ２％しかないわ**\n\n**Reimu:It's a list of repurchased items that were trendy among the items I bought so far.   \nMarisa:That's only 2%**","metadata":{}},{"cell_type":"code","source":"sub.head()","metadata":{"execution":{"iopub.status.busy":"2022-04-29T14:08:23.577021Z","iopub.execute_input":"2022-04-29T14:08:23.577229Z","iopub.status.idle":"2022-04-29T14:08:23.589835Z","shell.execute_reply.started":"2022-04-29T14:08:23.577203Z","shell.execute_reply":"2022-04-29T14:08:23.589203Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**霊夢：残りの９７％は売れ筋の違う柄か、同じ柄の違う型を選んでるんじゃない？\n魔理沙：価格帯もあるよね。3000円の服しか買わない人は3万円の服買わないよね**\n\n**Reimu:The remaining 97% are choosing different patterns that sell well, or different types with the same pattern, aren't they?  \nMarisa:There is also a price range. People who buy only 3,000 yen clothes don't buy 30,000 yen clothes, right?**","metadata":{}}]}