{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"**pixyz  \nlast update 2022 03 12  \nゆっくりしていってね！**","metadata":{}},{"cell_type":"markdown","source":"version8: article_id(購入した商品)、pred_id(購入すべき商品)、confidence(購入数)をまとめたdataframeをpickleで作成しました。SHORT_CUTを追加しました。  \nversion10:客の年齢層毎のitem_of_other_costomersを作成しました。","metadata":{}},{"cell_type":"markdown","source":"<img src=\"https://4.bp.blogspot.com/-6jLig_Zuhyk/UUhH8z560_I/AAAAAAAAO6A/lFCDFT8S1FM/s400/shopping_fasion.png\" width=200>","metadata":{}},{"cell_type":"markdown","source":"**霊夢:H&Mコンペ、客の数も商品の数も多いし、どうやって学習させればいいかわからないなあ**\n\n**魔理沙:「この商品を買った客は、あの商品もよく買ってるぞ！」って言うデータがあればいいのにな。**\n\n**霊夢:無ければ作るしかない！**\n\n**魔理沙:「ある客が買った商品一覧」の辞書と「製品を買った客一覧」の辞書があれば作れそうだな。**\n\n**魔理沙:ということで、「ある客が買った商品一覧」、「製品を買った客一覧」の辞書データと「ある商品を買った人が他に買っている商品ランキング」データをまとめたデータセットを作ったぜ。**\n\n**霊夢:良ければ学習の参考にしてね！**\n\n<br>\n\n**Reimu: H&M competition, there are a lot of customers and products, so I don't know how to learn.**\n\n**Marisa: I wish I had the data that \"customers who bought this product often buy that product too!\"**\n\n**Reimu: If you don't have it, you have to make it!**\n\n**Marisa: It seems that you can make it if you have a dictionary of \"list of products bought by a certain customer\" and a dictionary of \"list of customers who bought products\".**\n\n**Marisa: So, the dictionary data of \"List of products bought by a certain customer\", \"List of customers who bought a product\" and \"Ranking of products bought by a person who bought a certain product\" are summarized. I made a new data set.**\n\n**Reimu: If you like, use it as a reference for learning!**\n","metadata":{}},{"cell_type":"code","source":"SHORT_CUT = True","metadata":{"execution":{"iopub.status.busy":"2022-03-30T10:16:27.571205Z","iopub.execute_input":"2022-03-30T10:16:27.571678Z","iopub.status.idle":"2022-03-30T10:16:27.596798Z","shell.execute_reply.started":"2022-03-30T10:16:27.571558Z","shell.execute_reply":"2022-03-30T10:16:27.595915Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Import","metadata":{}},{"cell_type":"code","source":"import numpy as np\nimport pandas as pd\n\nfrom math import sqrt\nfrom pathlib import Path\nfrom tqdm import tqdm\nimport json\ntqdm.pandas()","metadata":{"execution":{"iopub.status.busy":"2022-03-30T10:16:27.656681Z","iopub.execute_input":"2022-03-30T10:16:27.657028Z","iopub.status.idle":"2022-03-30T10:16:27.662839Z","shell.execute_reply.started":"2022-03-30T10:16:27.656995Z","shell.execute_reply":"2022-03-30T10:16:27.661805Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"class MyEncoder(json.JSONEncoder):\n    def default(self, obj):\n        if isinstance(obj, np.integer):\n            return int(obj)\n","metadata":{"execution":{"iopub.status.busy":"2022-03-30T10:16:27.730027Z","iopub.execute_input":"2022-03-30T10:16:27.730352Z","iopub.status.idle":"2022-03-30T10:16:27.735284Z","shell.execute_reply.started":"2022-03-30T10:16:27.730316Z","shell.execute_reply":"2022-03-30T10:16:27.734577Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"!cp ../input/hm-dictionary/items_of_other_costomers.json ./items_of_other_costomers.json\n!cp ../input/hm-dictionary/items_of_other_costomers.pkl ./items_of_other_costomers.pkl","metadata":{"execution":{"iopub.status.busy":"2022-03-30T10:16:27.807817Z","iopub.execute_input":"2022-03-30T10:16:27.808554Z","iopub.status.idle":"2022-03-30T10:16:32.186179Z","shell.execute_reply.started":"2022-03-30T10:16:27.808508Z","shell.execute_reply":"2022-03-30T10:16:32.185242Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Path Setting","metadata":{}},{"cell_type":"code","source":"data_path = Path('../input/h-and-m-personalized-fashion-recommendations/')\ndict_path = Path('../input/hm-dictionary/')","metadata":{"execution":{"iopub.status.busy":"2022-03-30T10:16:32.189083Z","iopub.execute_input":"2022-03-30T10:16:32.189483Z","iopub.status.idle":"2022-03-30T10:16:32.194592Z","shell.execute_reply.started":"2022-03-30T10:16:32.189433Z","shell.execute_reply":"2022-03-30T10:16:32.193898Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Data Load","metadata":{}},{"cell_type":"code","source":"df = pd.read_csv(data_path / 'transactions_train.csv',\n                 usecols = ['customer_id', 'article_id'],\n                 dtype={'article_id': int})\ndf_articles = pd.read_csv(data_path / \"articles.csv\")","metadata":{"execution":{"iopub.status.busy":"2022-03-30T10:16:32.195691Z","iopub.execute_input":"2022-03-30T10:16:32.196485Z","iopub.status.idle":"2022-03-30T10:17:32.14569Z","shell.execute_reply.started":"2022-03-30T10:16:32.196453Z","shell.execute_reply":"2022-03-30T10:17:32.144714Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"version 10 追記  \n**魔理沙:年齢層ごとに別のランキングを作りたいので、客がどの年齢層に所属しているかを示すデータを作っていくよ。**","metadata":{}},{"cell_type":"code","source":"df_customers = pd.read_csv(data_path / 'customers.csv', usecols=['customer_id', 'age'],index_col = 'customer_id')\ndf_customers.head()","metadata":{"execution":{"iopub.status.busy":"2022-03-30T10:17:32.147749Z","iopub.execute_input":"2022-03-30T10:17:32.148023Z","iopub.status.idle":"2022-03-30T10:17:36.916193Z","shell.execute_reply.started":"2022-03-30T10:17:32.147987Z","shell.execute_reply":"2022-03-30T10:17:36.915267Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"listBin = [-1, 19, 29, 39, 49, 59, 69, 119]\ndf_customers['age_bins'] = pd.cut(df_customers['age'], listBin).astype(str)\ndf_customers.head()","metadata":{"execution":{"iopub.status.busy":"2022-03-30T10:17:36.917757Z","iopub.execute_input":"2022-03-30T10:17:36.918452Z","iopub.status.idle":"2022-03-30T10:17:37.008648Z","shell.execute_reply.started":"2022-03-30T10:17:36.918404Z","shell.execute_reply":"2022-03-30T10:17:37.007481Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"cus_agebins = df_customers['age_bins'].astype(str)","metadata":{"execution":{"iopub.status.busy":"2022-03-30T10:17:37.010158Z","iopub.execute_input":"2022-03-30T10:17:37.0105Z","iopub.status.idle":"2022-03-30T10:17:37.041108Z","shell.execute_reply.started":"2022-03-30T10:17:37.010458Z","shell.execute_reply":"2022-03-30T10:17:37.040277Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"listUniBins = df_customers['age_bins'].unique().tolist()","metadata":{"execution":{"iopub.status.busy":"2022-03-30T10:17:37.042686Z","iopub.execute_input":"2022-03-30T10:17:37.043331Z","iopub.status.idle":"2022-03-30T10:17:37.122031Z","shell.execute_reply.started":"2022-03-30T10:17:37.043262Z","shell.execute_reply":"2022-03-30T10:17:37.121026Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Make dictionary from customer to article","metadata":{}},{"cell_type":"markdown","source":"**魔理沙:まずは、「ある客が買った商品一覧」の辞書を作っていくぜ。**\n\n**Marisa: First, let's make a dictionary of \"a list of products bought by a customer\".**","metadata":{}},{"cell_type":"code","source":"if SHORT_CUT:\n    with open(dict_path / \"dict_c_a.json\", mode=\"r\") as f:\n        ds_dict_c_a = json.load(f)        \nelse:\n    ds_dict_c_a={}\n    for i in tqdm(range(len(df))):\n        customer_id=df.loc[i,\"customer_id\"]\n        article_id=df.loc[i,\"article_id\"]\n        if customer_id in ds_dict_c_a:\n            ds_dict_c_a[customer_id]=ds_dict_c_a[customer_id]+[article_id]        \n        else:\n            ds_dict_c_a[customer_id]=[article_id]    ","metadata":{"execution":{"iopub.status.busy":"2022-03-30T10:17:37.12386Z","iopub.execute_input":"2022-03-30T10:17:37.124129Z","iopub.status.idle":"2022-03-30T10:17:54.003646Z","shell.execute_reply.started":"2022-03-30T10:17:37.124098Z","shell.execute_reply":"2022-03-30T10:17:54.002549Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Make dictionary from article to customer","metadata":{}},{"cell_type":"markdown","source":"**魔理沙:次に、「ある商品を買った客一覧」の辞書を作っていくぜ。**\n\n**Marisa: Next, let's make a dictionary of \"list of customers who bought a certain product\".**","metadata":{}},{"cell_type":"code","source":"# if SHORT_CUT:\n#     with open(dict_path / 'dict_a_c.json', mode=\"r\") as f:\n#         ds_dict_a_c = json.load(f)\n# else:\n#     ds_dict_a_c={}\n#     for i in tqdm(range(len(df))):\n#         article_id=df.loc[i,\"article_id\"]\n#         ds_dict_a_c[int(article_id)]=[]\n#     for i in tqdm(range(len(df))):\n#         customer_id=df.loc[i,\"customer_id\"]\n#         article_id=df.loc[i,\"article_id\"]\n#         ds_dict_a_c[int(article_id)]=ds_dict_a_c[int(article_id)]+[customer_id]","metadata":{"execution":{"iopub.status.busy":"2022-03-30T10:17:54.004926Z","iopub.execute_input":"2022-03-30T10:17:54.00517Z","iopub.status.idle":"2022-03-30T10:17:54.009689Z","shell.execute_reply.started":"2022-03-30T10:17:54.005143Z","shell.execute_reply":"2022-03-30T10:17:54.008734Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"ds_dict_a_c = {listUniBins[i]:{} for i in range(len(listUniBins))}","metadata":{"execution":{"iopub.status.busy":"2022-03-30T10:17:54.01269Z","iopub.execute_input":"2022-03-30T10:17:54.013018Z","iopub.status.idle":"2022-03-30T10:17:54.025923Z","shell.execute_reply.started":"2022-03-30T10:17:54.012969Z","shell.execute_reply":"2022-03-30T10:17:54.024963Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"for i in tqdm(range(len(df))):\n    customer_id=df.loc[i,\"customer_id\"]\n    age_bin = cus_agebins[customer_id] \n    article_id=df.loc[i,\"article_id\"]\n    ds_dict_a_c[age_bin][int(article_id)]=[]\nfor i in tqdm(range(len(df))):\n    customer_id=df.loc[i,\"customer_id\"]\n    age_bin = cus_agebins[customer_id] \n    article_id=df.loc[i,\"article_id\"]\n    ds_dict_a_c[age_bin][int(article_id)]+=[customer_id]","metadata":{"execution":{"iopub.status.busy":"2022-03-30T10:17:54.027259Z","iopub.execute_input":"2022-03-30T10:17:54.027543Z","iopub.status.idle":"2022-03-30T10:47:43.20819Z","shell.execute_reply.started":"2022-03-30T10:17:54.027502Z","shell.execute_reply":"2022-03-30T10:47:43.206506Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Save as json file","metadata":{}},{"cell_type":"markdown","source":"**霊夢:jsonfileにして保存するよ。**\n\n**Reimu: Save it as a jsonfile.**","metadata":{}},{"cell_type":"code","source":"with open(\"dict_c_a.json\", mode=\"w\") as f:\n    ds_dict_c_a = json.dumps(ds_dict_c_a,cls = MyEncoder)\n    f.write(ds_dict_c_a)\nwith open(\"dict_a_c.json\", mode=\"w\") as f:\n    ds_dict_a_c = json.dumps(ds_dict_a_c,cls = MyEncoder)\n    f.write(ds_dict_a_c)","metadata":{"execution":{"iopub.status.busy":"2022-03-30T10:47:43.209732Z","iopub.status.idle":"2022-03-30T10:47:43.210089Z","shell.execute_reply.started":"2022-03-30T10:47:43.209915Z","shell.execute_reply":"2022-03-30T10:47:43.209935Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"ds_dict_a_c ={}\nds_dict_c_a ={}    \nwith open('dict_a_c.json', mode=\"r\") as f:\n    ds_dict_a_c = json.load(f)\nwith open(\"dict_c_a.json\", mode=\"r\") as f:\n    ds_dict_c_a = json.load(f)        ","metadata":{"execution":{"iopub.status.busy":"2022-03-30T10:47:43.211326Z","iopub.status.idle":"2022-03-30T10:47:43.21176Z","shell.execute_reply.started":"2022-03-30T10:47:43.211531Z","shell.execute_reply":"2022-03-30T10:47:43.211555Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(ds_dict_c_a[\"00000dbacae5abe5e23885899a1fa44253a17956c6d1c3d25f88aa139fdfc657\"])","metadata":{"execution":{"iopub.status.busy":"2022-03-30T10:47:43.212715Z","iopub.status.idle":"2022-03-30T10:47:43.213151Z","shell.execute_reply.started":"2022-03-30T10:47:43.212925Z","shell.execute_reply":"2022-03-30T10:47:43.21295Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Make dict article_id to index and index to article_id","metadata":{}},{"cell_type":"markdown","source":"**魔理沙:買った商品集計リストを作成しやすくするために、article_idを空間圧縮していくぜ。**\n\n**Marisa: I'm going to spatially compress the article_id to make it easier to create a list of the items I bought.**","metadata":{}},{"cell_type":"code","source":"df_a_i = {}\ndf_i_a = {}\nfor i in range(len(df_articles)):\n    df_a_i[df_articles.loc[i,\"article_id\"]] = i\n    df_i_a[i] = df_articles.loc[i,\"article_id\"]","metadata":{"execution":{"iopub.status.busy":"2022-03-30T10:47:43.214784Z","iopub.status.idle":"2022-03-30T10:47:43.215231Z","shell.execute_reply.started":"2022-03-30T10:47:43.214996Z","shell.execute_reply":"2022-03-30T10:47:43.21502Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Make dictionary items of other customers top 100","metadata":{}},{"cell_type":"markdown","source":"**魔理沙：そしたらいよいよ、「ある商品を買った人が他に買っている商品ランキング」を作成していくぜ。**\n\n**Marisa: Then, finally, we will create a \"ranking of products that people who bought one product are buying elsewhere\".**","metadata":{}},{"cell_type":"code","source":"# ranking = {}\n# for articl,coslist in tqdm(ds_dict_a_c.items()):\n#     df_count = np.zeros(len(df_articles))\n#     for costomer in coslist:\n#         for x in ds_dict_c_a[costomer]:\n#             df_count[df_a_i[x]] += 1\n#     df_count=np.argsort(df_count)\n#     df_count = df_count[-1::-1]\n#     ranking[int(articl)] = []\n#     for i in range(100):\n#         ranking[int(articl)]=ranking[int(articl)]+[df_i_a[df_count[i]]]\n#     del df_count","metadata":{"execution":{"iopub.status.busy":"2022-03-30T10:47:43.216976Z","iopub.status.idle":"2022-03-30T10:47:43.217259Z","shell.execute_reply.started":"2022-03-30T10:47:43.217115Z","shell.execute_reply":"2022-03-30T10:47:43.21713Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"cols=[\"article_id\",\"pred_id\",\"confidence\"]\ntable = pd.DataFrame(index=[], columns=cols)","metadata":{"execution":{"iopub.status.busy":"2022-03-30T10:47:43.218373Z","iopub.status.idle":"2022-03-30T10:47:43.218801Z","shell.execute_reply.started":"2022-03-30T10:47:43.218569Z","shell.execute_reply":"2022-03-30T10:47:43.218592Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# ##test\n# coslist = ds_dict_a_c['728162001']\n# count = [0]*len(df_articles)\n# for costomer in coslist:\n#     for x in ds_dict_c_a[costomer]:\n#         count[df_a_i[x]] += 1\n            \n# pred_list = sorted(range(len(df_articles)),key = lambda k:count[k],reverse = True)[:100]\n# for i in range(len(pred_list)):\n#     pred_list[i] = df_i_a[pred_list[i]]\n# conf_list = sorted(count,reverse=True)[:100]\n\n# print(pred_list)\n# print(conf_list)","metadata":{"execution":{"iopub.status.busy":"2022-03-30T10:47:43.22028Z","iopub.status.idle":"2022-03-30T10:47:43.220634Z","shell.execute_reply.started":"2022-03-30T10:47:43.220475Z","shell.execute_reply":"2022-03-30T10:47:43.220497Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"for uniBin in listUniBins:\n    if uniBin == 'nan': continue\n    print(uniBin)\n    article_id = []\n    pred_id = []\n    confidence = []\n    for articl,coslist in tqdm(ds_dict_a_c[uniBin].items()):\n        count = [0]*len(df_articles)    \n        for costomer in coslist:\n            for x in ds_dict_c_a[costomer]:\n                count[df_a_i[x]] += 1\n            \n        art_list = [articl]*100\n        pred_list = sorted(range(len(df_articles)),key = lambda k:count[k],reverse = True)[:100]\n        for i in range(len(pred_list)):\n            pred_list[i] = df_i_a[pred_list[i]]\n        conf_list = sorted(count,reverse=True)[:100]\n\n        article_id.extend(art_list)\n        pred_id.extend(pred_list)\n        confidence.extend(conf_list)\n        del art_list\n        del pred_list\n        del conf_list\n    \n    table = pd.DataFrame(list(zip(article_id,pred_id,confidence)),columns = cols)\n    table.to_pickle(f\"items_of_other_costomers_{uniBin}.pkl\")","metadata":{"execution":{"iopub.status.busy":"2022-03-30T10:47:43.221797Z","iopub.status.idle":"2022-03-30T10:47:43.222096Z","shell.execute_reply.started":"2022-03-30T10:47:43.221938Z","shell.execute_reply":"2022-03-30T10:47:43.221955Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# table = pd.DataFrame(list(zip(article_id,pred_id,confidence)),columns = cols)","metadata":{"execution":{"iopub.status.busy":"2022-03-30T10:47:43.223871Z","iopub.status.idle":"2022-03-30T10:47:43.224165Z","shell.execute_reply.started":"2022-03-30T10:47:43.224018Z","shell.execute_reply":"2022-03-30T10:47:43.224034Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"table.head()","metadata":{"execution":{"iopub.status.busy":"2022-03-30T10:47:43.225422Z","iopub.status.idle":"2022-03-30T10:47:43.225707Z","shell.execute_reply.started":"2022-03-30T10:47:43.225563Z","shell.execute_reply":"2022-03-30T10:47:43.225578Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Save as pickle file","metadata":{}},{"cell_type":"markdown","source":"**霊夢:最後にセーブして完了だね。**\n\n**Reimu: Finally save and you're done.**","metadata":{}},{"cell_type":"code","source":"# with open(\"items_of_other_costomers.json\", mode=\"w\") as f:\n#         ranking = json.dumps(ranking,cls = MyEncoder)\n#         f.write(ranking)        ","metadata":{"execution":{"iopub.status.busy":"2022-03-30T10:47:43.226554Z","iopub.status.idle":"2022-03-30T10:47:43.22684Z","shell.execute_reply.started":"2022-03-30T10:47:43.226685Z","shell.execute_reply":"2022-03-30T10:47:43.226701Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# table.to_pickle(\"items_of_other_costomers.pkl\")","metadata":{"execution":{"iopub.status.busy":"2022-03-30T10:47:43.227985Z","iopub.status.idle":"2022-03-30T10:47:43.228292Z","shell.execute_reply.started":"2022-03-30T10:47:43.228123Z","shell.execute_reply":"2022-03-30T10:47:43.228146Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"with open(\"items_of_other_costomers.json\", mode=\"r\") as f:\n        ds_dict = json.load(f)","metadata":{"execution":{"iopub.status.busy":"2022-03-30T10:47:43.229575Z","iopub.status.idle":"2022-03-30T10:47:43.229863Z","shell.execute_reply.started":"2022-03-30T10:47:43.229715Z","shell.execute_reply":"2022-03-30T10:47:43.22973Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**魔理沙:これ使ったらいいスコアでるといいなあ。**\n\n**霊夢:次回このデータを使ってpredictionしてみよう！**\n\n<br>\n\n**Marisa: I hope you get a good score if you use this.**\n\n**Reimu: Let's use this data for prediction next time!**\n\n<br>\n\n**読んでくれてありがとうございました！！よければvoteをよろしくお願いします！！**\n\n**Thank you for reading! !! If you like, please vote!!!**\n\n<img src=\"https://1.bp.blogspot.com/-PNcKwFw1PpM/U1T3oDIr9CI/AAAAAAAAfT4/gEn86X8Ppx0/s400/figure_goodjob.png\" width=200>","metadata":{}}]}