{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# NOTE\n\nThis is my attempt at rewriting [this notebook](https://www.kaggle.com/code/cdeotte/recommend-items-purchased-together-0-021) in polars.\n\n[Polars](https://www.pola.rs/) is a DataFrame library written completely in Rust, and is [blazingly fast](https://h2oai.github.io/db-benchmark/).\n\nHere, its performance is comparable to that of `cudf`, without needing a GPU!\n\nIt could probably go even faster, if using the lazy API.","metadata":{}},{"cell_type":"markdown","source":"# Recommend Items Frequently Purchased Together\nThis notebook demonstrates how recommending items that are frequently purchased together is effective. The current best scoring public notebook [here][1] recommends to customers those customers' last purchases and scores public LB 0.020. In this notebook here, we will begin with that idea and add recommending items that are frequently purchased together with a customers' previous purchaes. This notebook improves the LB and scores LB 0.021. This notebook's strategy is as follows:\n* recommend items previously purchased [idea here][1]\n* recommend items that are bought together with previous purchases [idea here][2]\n* recommend popular items [idea here][1]\n\n[1]: https://www.kaggle.com/hengzheng/time-is-our-best-friend-v2\n[2]: https://www.kaggle.com/cdeotte/customers-who-bought-this-frequently-buy-this","metadata":{}},{"cell_type":"markdown","source":"# polars\nWe will use polars for fast dataframe operations","metadata":{}},{"cell_type":"code","source":"!pip install polars","metadata":{"execution":{"iopub.status.busy":"2023-02-11T19:46:29.336637Z","iopub.execute_input":"2023-02-11T19:46:29.337040Z","iopub.status.idle":"2023-02-11T19:46:39.734318Z","shell.execute_reply.started":"2023-02-11T19:46:29.337005Z","shell.execute_reply":"2023-02-11T19:46:39.732908Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import polars as pl\nimport datetime as dt\nimport numpy as np\nprint('polars version', pl.__version__)","metadata":{"execution":{"iopub.status.busy":"2023-02-11T19:46:39.736814Z","iopub.execute_input":"2023-02-11T19:46:39.737174Z","iopub.status.idle":"2023-02-11T19:46:39.743688Z","shell.execute_reply.started":"2023-02-11T19:46:39.737138Z","shell.execute_reply":"2023-02-11T19:46:39.742598Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Load Transactions, Reduce Memory\nDiscussion about reducing memory is [here][1]\n\n[1]: https://www.kaggle.com/c/h-and-m-personalized-fashion-recommendations/discussion/308635","metadata":{}},{"cell_type":"code","source":"%%time\ntrain = (\n    pl.scan_csv(\n        '../input/h-and-m-personalized-fashion-recommendations/transactions_train.csv',\n        parse_dates=True,\n    )\n    .select([\n        pl.col('article_id').cast(pl.Int32),\n        pl.col('t_dat'),\n        pl.col('customer_id'),\n    ])\n)","metadata":{"execution":{"iopub.status.busy":"2023-02-11T19:46:39.745240Z","iopub.execute_input":"2023-02-11T19:46:39.745618Z","iopub.status.idle":"2023-02-11T19:46:39.762017Z","shell.execute_reply.started":"2023-02-11T19:46:39.745568Z","shell.execute_reply":"2023-02-11T19:46:39.761016Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train.head()","metadata":{"execution":{"iopub.status.busy":"2023-02-11T19:46:39.764210Z","iopub.execute_input":"2023-02-11T19:46:39.764600Z","iopub.status.idle":"2023-02-11T19:46:39.800042Z","shell.execute_reply.started":"2023-02-11T19:46:39.764563Z","shell.execute_reply":"2023-02-11T19:46:39.798672Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# (1) Recommend Last Week's Most Popular Items\nAfter recommending previous purchases and items purchased together we will then recommend the 12 most popular items. Therefore if our previous recommendations did not fill up a customer's 12 recommendations, then it will be filled by popular items.","metadata":{}},{"cell_type":"code","source":"%%time\ntop12 = \" 0\" + (\n    \" 0\".join(\n        train.collect()[\"article_id\"]\n        .value_counts()\n        .sort(\"counts\", reverse=True)[:12][\"article_id\"]\n        .cast(pl.Utf8)\n        .to_list()\n    )\n)\n\nprint(\"Last week's top 12 popular items:\")\nprint( top12 )","metadata":{"execution":{"iopub.status.busy":"2023-02-11T19:47:00.108344Z","iopub.execute_input":"2023-02-11T19:47:00.108791Z","iopub.status.idle":"2023-02-11T19:47:18.674858Z","shell.execute_reply.started":"2023-02-11T19:47:00.108754Z","shell.execute_reply":"2023-02-11T19:47:18.673669Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Find Each Customer's Last Week of Purchases\nOur final predictions will have the row order from of our dataframe. Each row of our dataframe will be a prediction. We will create the `predictionstring` later by `train.groupby('customer_id').article_id.sum()`. Since `article_id` is a string, when we groupby sum, it will concatenate all the customer predictions into a single string. It will also create the string in the order of the dataframe. So as we proceed in this notebook, we will order the dataframe how we want our predictions ordered.","metadata":{}},{"cell_type":"code","source":"%%time\ntrain = (\n    train.with_columns([\n        pl.col('t_dat').max().over('customer_id').alias('max_dat')\n    ])\n    .with_columns(\n        (pl.col('max_dat')-pl.col('t_dat')).dt.days().alias('diff_dat'))\n    .filter(pl.col('diff_dat')<=6)\n)","metadata":{"execution":{"iopub.status.busy":"2023-02-11T19:47:39.338289Z","iopub.execute_input":"2023-02-11T19:47:39.338715Z","iopub.status.idle":"2023-02-11T19:47:39.346472Z","shell.execute_reply.started":"2023-02-11T19:47:39.338681Z","shell.execute_reply":"2023-02-11T19:47:39.345331Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# (2) Recommend Most Often Previously Purchased Items\nWe need to sort because we want the most often previously purchased items first. Because this will be the order of our predictons. Since we sort by `ct` and then `t_dat` will will recommend items that have been purchased more frequently first followed by items purchased more recently second.","metadata":{}},{"cell_type":"code","source":"%%time\ntrain = (\n    train\n    .with_columns([\n        pl.col('t_dat').count().over(['customer_id', 'article_id']).alias('ct')\n    ])\n    .sort(['ct','t_dat'], reverse=True)\n    .unique(subset=['customer_id', 'article_id'])\n)","metadata":{"execution":{"iopub.status.busy":"2023-02-11T19:47:43.222318Z","iopub.execute_input":"2023-02-11T19:47:43.222746Z","iopub.status.idle":"2023-02-11T19:47:43.230564Z","shell.execute_reply.started":"2023-02-11T19:47:43.222710Z","shell.execute_reply":"2023-02-11T19:47:43.229290Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train.head()","metadata":{"execution":{"iopub.status.busy":"2023-02-11T19:47:45.756348Z","iopub.execute_input":"2023-02-11T19:47:45.756749Z","iopub.status.idle":"2023-02-11T19:47:45.822170Z","shell.execute_reply.started":"2023-02-11T19:47:45.756715Z","shell.execute_reply":"2023-02-11T19:47:45.820896Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# (3) Recommend Items Purchased Together\nIn my notebook [here][1], we compute a dictionary of items frequently purchased together. We will load and use that dictionary below. Note that we use the command `drop_duplicates` so that we don't recommend an item that the user has already bought and we have already recommended above. \n\nWe concatenate these rows after the rows containing customers' previous purchases. Therefore we will recommend previous items first and then items purchased together second. Note the trick to convert a column of int32 into a prediction string (using groupby agg str sum) is from notebook [here][2]\n\n[1]: https://www.kaggle.com/cdeotte/customers-who-bought-this-frequently-buy-this\n[2]: https://www.kaggle.com/hiroshisakiyama/recommending-items-recently-bought","metadata":{}},{"cell_type":"code","source":"%%time\nimport numpy as np\npairs = np.load('../input/hmitempairs/pairs_cudf.npy',allow_pickle=True).item()\npairs = pl.DataFrame({\n    'article_id': pairs.keys(),\n    'article_id2': pairs.values()\n}\n).with_columns(pl.col('article_id').cast(pl.Int32)).with_columns(pl.col('article_id2').cast(pl.Int32))","metadata":{"execution":{"iopub.status.busy":"2023-02-11T19:47:49.846359Z","iopub.execute_input":"2023-02-11T19:47:49.846802Z","iopub.status.idle":"2023-02-11T19:47:50.043822Z","shell.execute_reply.started":"2023-02-11T19:47:49.846761Z","shell.execute_reply":"2023-02-11T19:47:50.042614Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\ntrain = train.join(pairs.lazy(), on='article_id', how='left')","metadata":{"execution":{"iopub.status.busy":"2023-02-11T19:47:57.562885Z","iopub.execute_input":"2023-02-11T19:47:57.563268Z","iopub.status.idle":"2023-02-11T19:47:57.569494Z","shell.execute_reply.started":"2023-02-11T19:47:57.563230Z","shell.execute_reply":"2023-02-11T19:47:57.568254Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train.head()","metadata":{"execution":{"iopub.status.busy":"2023-02-11T19:48:00.408819Z","iopub.execute_input":"2023-02-11T19:48:00.409204Z","iopub.status.idle":"2023-02-11T19:48:00.471945Z","shell.execute_reply.started":"2023-02-11T19:48:00.409172Z","shell.execute_reply":"2023-02-11T19:48:00.470542Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\ntrain2 = (\n    train.select(['customer_id', 'article_id2'])\n    .filter(~pl.col('article_id2').is_null())\n    .unique(subset=['customer_id', 'article_id2'])\n    .rename({'article_id2': 'article_id'})\n)","metadata":{"execution":{"iopub.status.busy":"2023-02-11T19:48:03.799654Z","iopub.execute_input":"2023-02-11T19:48:03.800054Z","iopub.status.idle":"2023-02-11T19:48:03.807760Z","shell.execute_reply.started":"2023-02-11T19:48:03.800018Z","shell.execute_reply":"2023-02-11T19:48:03.806580Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train2.head()","metadata":{"execution":{"iopub.status.busy":"2023-02-11T19:48:06.216836Z","iopub.execute_input":"2023-02-11T19:48:06.217667Z","iopub.status.idle":"2023-02-11T19:48:06.278854Z","shell.execute_reply.started":"2023-02-11T19:48:06.217626Z","shell.execute_reply":"2023-02-11T19:48:06.277462Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\n# CONCATENATE PAIRED ITEM RECOMMENDATION AFTER PREVIOUS PURCHASED RECOMMENDATIONS\npreds = (\n    pl.concat([\n        train.select(['customer_id', 'article_id']),\n        train2\n    ])\n    .unique(subset=['customer_id','article_id'])\n    .with_columns([\n        (pl.lit(' 0') + pl.col('article_id').cast(pl.Utf8)).alias('article_id')\n    ])\n    .groupby('customer_id')\n    .agg(pl.col('article_id').str.concat(\"\"))\n    .rename({'article_id': 'prediction'})\n)\npreds.head()","metadata":{"execution":{"iopub.status.busy":"2023-02-11T19:48:19.828984Z","iopub.execute_input":"2023-02-11T19:48:19.829459Z","iopub.status.idle":"2023-02-11T19:48:19.896512Z","shell.execute_reply.started":"2023-02-11T19:48:19.829416Z","shell.execute_reply":"2023-02-11T19:48:19.895122Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Write Submission CSV\nWe will merge our predictions onto `sample_submission.csv` and submit to Kaggle.","metadata":{}},{"cell_type":"code","source":"%%time\nsub = (\n    pl.scan_csv('../input/h-and-m-personalized-fashion-recommendations/sample_submission.csv')\n    .select(['customer_id'])\n    .join(preds, on='customer_id', how='left')\n    .fill_null('')\n    .with_columns([\n            (pl.col('prediction') + top12).str.strip().str.slice(0, 131)\n    ])\n)","metadata":{"execution":{"iopub.status.busy":"2023-02-11T19:48:26.011717Z","iopub.execute_input":"2023-02-11T19:48:26.012813Z","iopub.status.idle":"2023-02-11T19:48:26.029439Z","shell.execute_reply.started":"2023-02-11T19:48:26.012770Z","shell.execute_reply":"2023-02-11T19:48:26.028596Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\nsub = sub.collect()\nsub.write_csv('submission.csv')\nsub.head()","metadata":{"execution":{"iopub.status.busy":"2023-02-11T17:20:06.581985Z","iopub.execute_input":"2023-02-11T17:20:06.582737Z","iopub.status.idle":"2023-02-11T17:20:07.287740Z","shell.execute_reply.started":"2023-02-11T17:20:06.582695Z","shell.execute_reply":"2023-02-11T17:20:07.286520Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]}]}