{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"`article_id`s are not random integers. They are strongly tied to the release dates of the articles; articles with smaller ids are older and larger values are newer, and I also estimate that the values are proportional to time.","metadata":{}},{"cell_type":"code","source":"import numpy as np\nimport pandas as pd\nimport matplotlib.pyplot as plt\n\ndi = '/kaggle/input/h-and-m-personalized-fashion-recommendations/'","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","execution":{"iopub.status.busy":"2022-03-19T01:42:04.788157Z","iopub.execute_input":"2022-03-19T01:42:04.788899Z","iopub.status.idle":"2022-03-19T01:42:04.81619Z","shell.execute_reply.started":"2022-03-19T01:42:04.788767Z","shell.execute_reply":"2022-03-19T01:42:04.815199Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"transactions = pd.read_csv(di + 'transactions_train.csv')\ntransactions['t_dat'] = pd.to_datetime(transactions['t_dat'])\n\narticles = pd.read_csv(di + 'articles.csv')","metadata":{"execution":{"iopub.status.busy":"2022-03-19T01:42:04.818589Z","iopub.execute_input":"2022-03-19T01:42:04.819244Z","iopub.status.idle":"2022-03-19T01:43:17.66028Z","shell.execute_reply.started":"2022-03-19T01:42:04.819193Z","shell.execute_reply":"2022-03-19T01:43:17.659349Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Time notation\n\nDays from the first date in the transaction:\n\n* **t = date - 2018-09-20**  in days","metadata":{"execution":{"iopub.status.busy":"2022-03-18T22:28:32.922427Z","iopub.execute_input":"2022-03-18T22:28:32.922785Z","iopub.status.idle":"2022-03-18T22:28:32.929474Z","shell.execute_reply.started":"2022-03-18T22:28:32.922751Z","shell.execute_reply":"2022-03-18T22:28:32.928052Z"}}},{"cell_type":"code","source":"# Define time t by the time from the first transaction in days\nt0 = pd.to_datetime('2018-09-20')\ntransactions['t'] = (transactions['t_dat'] - t0).dt.days","metadata":{"execution":{"iopub.status.busy":"2022-03-19T01:43:17.661975Z","iopub.execute_input":"2022-03-19T01:43:17.662221Z","iopub.status.idle":"2022-03-19T01:43:18.47918Z","shell.execute_reply.started":"2022-03-19T01:43:17.662192Z","shell.execute_reply":"2022-03-19T01:43:18.478288Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Time of first purchase","metadata":{}},{"cell_type":"code","source":"# First t that each article was purchased\n# t_release <= t_first\nt_first = transactions.groupby(['article_id'])[['t']].min()\n\nplt.title('When does each article purchased for the first time?')\nplt.xlabel('article index / 1000')\nplt.ylabel('t_first [day]')\nplt.plot(t_first.values[::1000])\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-03-19T01:43:18.497013Z","iopub.execute_input":"2022-03-19T01:43:18.497203Z","iopub.status.idle":"2022-03-19T01:43:20.443071Z","shell.execute_reply.started":"2022-03-19T01:43:18.497178Z","shell.execute_reply":"2022-03-19T01:43:20.442166Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Clearly, the articles later in the data are first sold later in time. Since unpopular items are not sold immediately after the release, there is randomness in the plot.\n\nArticle index ~40 x 1000 seems to correspond to the beginning of the training data period (2018-09-20).\n\n\n## Relase date estimate\n\nAssuming the article data are in order of release, we can clean the plot by imposing monotonically increasing t_release_estimate.\n\n```\nt_release <= t_release_est <= t_first\nt_release_est[i] <= t_release_est[i + 1]\n```","metadata":{}},{"cell_type":"code","source":"len(articles), len(transactions['article_id'].unique())","metadata":{"execution":{"iopub.status.busy":"2022-03-19T01:43:20.445498Z","iopub.execute_input":"2022-03-19T01:43:20.445722Z","iopub.status.idle":"2022-03-19T01:43:20.814729Z","shell.execute_reply.started":"2022-03-19T01:43:20.445695Z","shell.execute_reply":"2022-03-19T01:43:20.813684Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Include items never purchced, do we have items not relased by the end of train?\n\ndf = articles.merge(t_first, how='left', left_on='article_id', right_index=True)\ndf['t'].fillna(999, inplace=True) # ~1000 items never purchased\ndf.rename(columns={'t': 't_first'}, inplace=True)","metadata":{"execution":{"iopub.status.busy":"2022-03-19T01:43:20.815838Z","iopub.execute_input":"2022-03-19T01:43:20.81604Z","iopub.status.idle":"2022-03-19T01:43:20.898302Z","shell.execute_reply.started":"2022-03-19T01:43:20.816014Z","shell.execute_reply":"2022-03-19T01:43:20.897262Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"t_first = df['t_first'].values\n\nt_min = 999\nt_release_est = np.zeros(len(t_first), dtype=int)\nn = len(t_first)\n\nfor i in range(n - 1, -1, -1):\n    t = t_first[i]\n    t_min = min(t, t_min)  # t_release_est[i] = min(t_first[i], t_release_est[i + 1])\n    t_release_est[i] = t_min","metadata":{"execution":{"iopub.status.busy":"2022-03-19T01:43:20.899221Z","iopub.execute_input":"2022-03-19T01:43:20.899426Z","iopub.status.idle":"2022-03-19T01:43:21.012118Z","shell.execute_reply.started":"2022-03-19T01:43:20.899401Z","shell.execute_reply":"2022-03-19T01:43:21.011238Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plt.xlabel('article index')\nplt.ylabel('release date [day]')\nplt.plot(t_release_est)\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-03-19T01:43:21.013476Z","iopub.execute_input":"2022-03-19T01:43:21.013699Z","iopub.status.idle":"2022-03-19T01:43:21.308132Z","shell.execute_reply.started":"2022-03-19T01:43:21.013672Z","shell.execute_reply":"2022-03-19T01:43:21.307565Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"t_release_est[-8:]","metadata":{"execution":{"iopub.status.busy":"2022-03-19T01:43:21.309097Z","iopub.execute_input":"2022-03-19T01:43:21.309878Z","iopub.status.idle":"2022-03-19T01:43:21.315679Z","shell.execute_reply.started":"2022-03-19T01:43:21.30983Z","shell.execute_reply":"2022-03-19T01:43:21.314847Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Only 2 items have t_release_est = 999; I conclude that articles released after the training period (2020-09-22) are *not* in our article data, and we do not need to recommend such new items.\n\nHowever, notice that new items have short period of time to purchase and therefore have larger possibility that no one has purchased yet.","metadata":{}},{"cell_type":"code","source":"idx = df['t_first'] == 999\nzero_purchased = df.index[idx]\n\nplt.xlabel('article_id index')\nplt.ylabel('number of items never purchased')\nplt.hist(zero_purchased, 101)\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-03-19T01:43:21.316882Z","iopub.execute_input":"2022-03-19T01:43:21.317116Z","iopub.status.idle":"2022-03-19T01:43:21.685072Z","shell.execute_reply.started":"2022-03-19T01:43:21.317087Z","shell.execute_reply":"2022-03-19T01:43:21.684488Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Article index vs article_id","metadata":{}},{"cell_type":"code","source":"plt.figure(figsize=(12, 4))\nplt.subplot(1, 2, 1)\nplt.ylim(0, 733)\nplt.xlabel('article index')\nplt.ylabel('t_release_est [day]')\nplt.plot(t_release_est)\n\nplt.subplot(1, 2, 2)\nplt.xlabel('article_id')\nplt.ylim(0, 633)\nplt.plot(df['article_id'], t_release_est)\n\ntt = np.linspace(0.73e9, 0.95e9)\nplt.plot(tt, 3.1e-6*(tt - tt[0]))\n\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-03-19T01:43:21.685942Z","iopub.execute_input":"2022-03-19T01:43:21.686614Z","iopub.status.idle":"2022-03-19T01:43:22.000328Z","shell.execute_reply.started":"2022-03-19T01:43:21.686578Z","shell.execute_reply":"2022-03-19T01:43:21.999364Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The relation of article_id - t_release_est looks more linear and I speculate article_ids are release dates.","metadata":{}},{"cell_type":"markdown","source":"Not all digits in article_ids are dates; the last 3 digits are not uniform and it must be some index assigned from 1.","metadata":{}},{"cell_type":"code","source":"articles['article_id'].head()","metadata":{"execution":{"iopub.status.busy":"2022-03-19T01:43:22.001599Z","iopub.execute_input":"2022-03-19T01:43:22.001831Z","iopub.status.idle":"2022-03-19T01:43:22.010236Z","shell.execute_reply.started":"2022-03-19T01:43:22.001803Z","shell.execute_reply":"2022-03-19T01:43:22.009313Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"last3 = articles['article_id'].apply(lambda x: int(str(x)[-3:]))\n\nplt.xlabel('Last 3 digits in article_id')\nplt.hist(last3, 41);","metadata":{"execution":{"iopub.status.busy":"2022-03-19T01:43:22.011333Z","iopub.execute_input":"2022-03-19T01:43:22.011769Z","iopub.status.idle":"2022-03-19T01:43:22.425756Z","shell.execute_reply.started":"2022-03-19T01:43:22.011731Z","shell.execute_reply":"2022-03-19T01:43:22.425075Z"},"trusted":true},"execution_count":null,"outputs":[]}]}