{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# OTTO – Multi-Objective Recommender System","metadata":{}},{"cell_type":"markdown","source":"## Goal of the Competition\nThe goal of this competition is to predict e-commerce clicks, cart additions, and orders. You'll build a multi-objective recommender system based on previous events in a user session.\n\nYour work will help improve the shopping experience for everyone involved. Customers will receive more tailored recommendations while online retailers may increase their sales.\n\n## Context\nOnline shoppers have their pick of millions of products from large retailers. While such variety may be impressive, having so many options to explore can be overwhelming, resulting in shoppers leaving with empty carts. This neither benefits shoppers seeking to make a purchase nor retailers that missed out on sales. This is one reason online retailers rely on recommender systems to guide shoppers to products that best match their interests and motivations. Using data science to enhance retailers' ability to predict which products each customer actually wants to see, add to their cart, and order at any given moment of their visit in real-time could improve your customer experience the next time you shop online with your favorite retailer.\n\nCurrent recommender systems consist of various models with different approaches, ranging from simple matrix factorization to a transformer-type deep neural network. However, no single model exists that can simultaneously optimize multiple objectives. In this competition, you’ll build a single entry to predict click-through, add-to-cart, and conversion rates based on previous same-session events.\n\nWith more than 10 million products from over 19,000 brands, OTTO is the largest German online shop. OTTO is a member of the Hamburg-based, multi-national Otto Group, which also subsidizes Crate & Barrel (USA) and 3 Suisses (France).\n\nYour work will help online retailers select more relevant items from a vast range to recommend to their customers based on their real-time behavior. Improving recommendations will ensure navigating through seemingly endless options is more effortless and engaging for shoppers.","metadata":{}},{"cell_type":"markdown","source":"# Imports 📥","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19"}},{"cell_type":"markdown","source":"**🟦EN** Install [polars](https://www.pola.rs/)\n\n**🟥ES** Instalamos [polars](https://www.pola.rs/)","metadata":{}},{"cell_type":"code","source":"!pip install polars","metadata":{"execution":{"iopub.status.busy":"2022-12-17T08:05:56.317473Z","iopub.execute_input":"2022-12-17T08:05:56.317861Z","iopub.status.idle":"2022-12-17T08:06:13.012584Z","shell.execute_reply.started":"2022-12-17T08:05:56.317830Z","shell.execute_reply":"2022-12-17T08:06:13.011113Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import pandas as pd\nimport seaborn as sns\nimport numpy as np\nimport multiprocessing\nimport polars as pl\nfrom gensim.test.utils import common_texts\nfrom gensim.models import Word2Vec\nimport os\nfrom annoy import AnnoyIndex\nimport collections","metadata":{"execution":{"iopub.status.busy":"2022-12-17T08:06:14.116339Z","iopub.execute_input":"2022-12-17T08:06:14.117096Z","iopub.status.idle":"2022-12-17T08:06:14.123364Z","shell.execute_reply.started":"2022-12-17T08:06:14.117051Z","shell.execute_reply":"2022-12-17T08:06:14.121826Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Constants 📋","metadata":{}},{"cell_type":"code","source":"PATH = \"/kaggle/input/otto-full-optimized-memory-footprint\"\nTEST_NAME = \"test.parquet\"\nTRAIN_NAME = \"train.parquet\"\nTEST_PATH = os.path.join(PATH, TEST_NAME)\nTRAIN_PATH = os.path.join(PATH, TRAIN_NAME)","metadata":{"execution":{"iopub.status.busy":"2022-12-17T08:06:51.027325Z","iopub.execute_input":"2022-12-17T08:06:51.027806Z","iopub.status.idle":"2022-12-17T08:06:51.034347Z","shell.execute_reply.started":"2022-12-17T08:06:51.027770Z","shell.execute_reply":"2022-12-17T08:06:51.033207Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Funcions 📘","metadata":{}},{"cell_type":"code","source":"def m20_mult(x, n_length):\n    x = x[-20:]\n    if n_length > 19:\n        n_length = 19\n    if len(x) >= n_length:\n        y = []\n        y_sum = list(np.zeros(21-len(x)))\n        for i in range(n_length):\n            y.append(list(index.get_nns_by_item(x[-(i+1)], 21 - len(x))[1:]))\n            y_sum += y[-1]\n\n        counter = dict(collections.Counter(y_sum))\n        res = sorted(list(set(y_sum)), key = lambda d: counter[d], reverse=True)\n    else:\n        res = list(index.get_nns_by_item(x[-1], 21 - len(x))[1:])\n        \n    x = list(x) + list(res)[:20-len(x)]\n    \n    return x\n\n","metadata":{"execution":{"iopub.status.busy":"2022-12-17T08:05:45.851409Z","iopub.status.idle":"2022-12-17T08:05:45.851790Z","shell.execute_reply.started":"2022-12-17T08:05:45.851596Z","shell.execute_reply":"2022-12-17T08:05:45.851613Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Data modification 📊","metadata":{}},{"cell_type":"markdown","source":"**🟦EN** Get data using *polar* and the función *read_paquet*\n\n**🟥ES** Cogemos los datos la libreria de *polar* y la función de *read_paquet*","metadata":{}},{"cell_type":"code","source":"train_df = pl.read_parquet(TRAIN_PATH)\ntest_df = pl.read_parquet(TEST_PATH)","metadata":{"execution":{"iopub.status.busy":"2022-12-17T08:06:58.644370Z","iopub.execute_input":"2022-12-17T08:06:58.645125Z","iopub.status.idle":"2022-12-17T08:07:12.845454Z","shell.execute_reply.started":"2022-12-17T08:06:58.645085Z","shell.execute_reply":"2022-12-17T08:07:12.844358Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**🟦EN** Show the raw data\n\n**🟥ES** Mostramos los datos sin modificar","metadata":{}},{"cell_type":"code","source":"train_df.head()","metadata":{"execution":{"iopub.status.busy":"2022-12-17T08:07:37.529312Z","iopub.execute_input":"2022-12-17T08:07:37.530122Z","iopub.status.idle":"2022-12-17T08:07:37.540877Z","shell.execute_reply.started":"2022-12-17T08:07:37.530073Z","shell.execute_reply":"2022-12-17T08:07:37.539632Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test_df.head()","metadata":{"execution":{"iopub.status.busy":"2022-12-17T08:07:37.554418Z","iopub.execute_input":"2022-12-17T08:07:37.554726Z","iopub.status.idle":"2022-12-17T08:07:37.566367Z","shell.execute_reply.started":"2022-12-17T08:07:37.554698Z","shell.execute_reply":"2022-12-17T08:07:37.565174Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sentences_df = pl.concat([train_df, test_df]).groupby(\"session\").agg(pl.col(\"aid\").alias(\"sentence\"))","metadata":{"execution":{"iopub.status.busy":"2022-12-17T08:07:25.476964Z","iopub.execute_input":"2022-12-17T08:07:25.477281Z","iopub.status.idle":"2022-12-17T08:07:37.525925Z","shell.execute_reply.started":"2022-12-17T08:07:25.477251Z","shell.execute_reply":"2022-12-17T08:07:37.524501Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sentences_df.head()","metadata":{"execution":{"iopub.status.busy":"2022-12-17T08:07:53.906876Z","iopub.execute_input":"2022-12-17T08:07:53.907622Z","iopub.status.idle":"2022-12-17T08:07:53.914684Z","shell.execute_reply.started":"2022-12-17T08:07:53.907584Z","shell.execute_reply":"2022-12-17T08:07:53.913524Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**🟦EN** Modify the data to imporve the results of the model.\n\n**🟥ES** Modificamos los datos para mejorar el resultado del entrenamiento","metadata":{}},{"cell_type":"code","source":"test_pred_df = pl.concat([test_df]).groupby(\"session\").agg(pl.col(\"aid\").alias(\"sentence\"))\ntest_pred_df.head()","metadata":{"execution":{"iopub.status.busy":"2022-12-17T08:07:37.658676Z","iopub.status.idle":"2022-12-17T08:07:37.659184Z","shell.execute_reply.started":"2022-12-17T08:07:37.658934Z","shell.execute_reply":"2022-12-17T08:07:37.658959Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sentences_df.head()","metadata":{"execution":{"iopub.status.busy":"2022-12-17T08:07:37.668477Z","iopub.status.idle":"2022-12-17T08:07:37.668952Z","shell.execute_reply.started":"2022-12-17T08:07:37.668714Z","shell.execute_reply":"2022-12-17T08:07:37.668737Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test_pred_df = test_pred_df.to_pandas().rename(columns={'sentence':'labels'})\n\nsentences_df_clicks = pl.concat([test_df]).filter(pl.col('type') == 0)\nsentences_df_carts = pl.concat([test_df]).filter(pl.col('type') == 1)\nsentences_df_orders = pl.concat([test_df]).filter(pl.col('type') == 2)\n\nsentences_df_clicks = sentences_df_clicks.groupby('session').agg(pl.col('aid').alias('sentence'))\nsentences_df_carts = sentences_df_carts.groupby('session').agg(pl.col('aid').alias('sentence'))\nsentences_df_orders = sentences_df_orders.groupby('session').agg(pl.col('aid').alias('sentence'))\n\nsentences_df_clicks = sentences_df_clicks.to_pandas().rename(columns={'sentence':'labels_clicks'})\nsentences_df_carts = sentences_df_carts.to_pandas().rename(columns={'sentence':'labels_carts'})\nsentences_df_orders = sentences_df_orders.to_pandas().rename(columns={'sentence':'labels_orders'})","metadata":{"execution":{"iopub.status.busy":"2022-12-08T14:13:44.592600Z","iopub.execute_input":"2022-12-08T14:13:44.593313Z","iopub.status.idle":"2022-12-08T14:13:48.164442Z","shell.execute_reply.started":"2022-12-08T14:13:44.593267Z","shell.execute_reply":"2022-12-08T14:13:48.161164Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test_pred_df = test_pred_df.merge(sentences_df_clicks, how='left', on='session') \\\n                           .merge(sentences_df_carts, how='left', on='session') \\\n                           .merge(sentences_df_orders, how='left', on='session') \ntest_pred_df.head()","metadata":{"execution":{"iopub.status.busy":"2022-12-08T14:13:48.172014Z","iopub.execute_input":"2022-12-08T14:13:48.173846Z","iopub.status.idle":"2022-12-08T14:13:55.294321Z","shell.execute_reply.started":"2022-12-08T14:13:48.173706Z","shell.execute_reply":"2022-12-08T14:13:55.291822Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sentences_list = sentences_df['sentence'].to_list()","metadata":{"execution":{"iopub.status.busy":"2022-12-08T14:13:55.296788Z","iopub.execute_input":"2022-12-08T14:13:55.297415Z","iopub.status.idle":"2022-12-08T14:14:57.468582Z","shell.execute_reply.started":"2022-12-08T14:13:55.297344Z","shell.execute_reply":"2022-12-08T14:14:57.466969Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"n_cores = multiprocessing.cpu_count() - 1\nw2v = Word2Vec(\n    sentences = sentences_list,\n    vector_size = 100,\n    alpha = 0.02,\n    min_alpha = 0.01,\n    min_count = 1,\n    workers = n_cores)","metadata":{"execution":{"iopub.status.busy":"2022-12-08T14:14:57.470209Z","iopub.execute_input":"2022-12-08T14:14:57.470603Z","iopub.status.idle":"2022-12-08T14:55:55.180657Z","shell.execute_reply.started":"2022-12-08T14:14:57.470569Z","shell.execute_reply":"2022-12-08T14:55:55.179088Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**🟦EN** We load the model we have downloaded to predict the results.\n\n**🟥ES** Cargamos el modelo que hemos descargado para predecir los resultados.","metadata":{}},{"cell_type":"code","source":"w2v.save(\"w2vc.model\")\nmodel = Word2Vec.load(\"w2vc.model\")","metadata":{"execution":{"iopub.status.busy":"2022-12-08T14:55:55.183086Z","iopub.execute_input":"2022-12-08T14:55:55.183520Z","iopub.status.idle":"2022-12-08T14:55:57.791949Z","shell.execute_reply.started":"2022-12-08T14:55:55.183475Z","shell.execute_reply":"2022-12-08T14:55:57.790172Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"aid2idx = {aid: i for i, aid in enumerate(model.wv.index_to_key)}\nindex = AnnoyIndex(100, 'angular')\n\nfor aid, idx in aid2idx.items():\n    index.add_item(aid, model.wv.vectors[idx])\n    \nindex.build(50)","metadata":{"execution":{"iopub.status.busy":"2022-12-08T14:55:57.792940Z","iopub.status.idle":"2022-12-08T14:55:57.793448Z","shell.execute_reply.started":"2022-12-08T14:55:57.793223Z","shell.execute_reply":"2022-12-08T14:55:57.793247Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**🟦EN** Split the labels by carts, order r clicks.\n\n**🟥ES** Separamos las \"lables\" en carts, order, clicks.","metadata":{}},{"cell_type":"code","source":"test_pred_df['labels_carts'] = test_pred_df['labels_carts'].fillna(test_pred_df['labels'])\ntest_pred_df['labels_orders'] = test_pred_df['labels_orders'].fillna(test_pred_df['labels'])\ntest_pred_df['labels_clicks'] = test_pred_df['labels_clicks'].fillna(test_pred_df['labels'])","metadata":{"execution":{"iopub.status.busy":"2022-12-08T14:55:57.795895Z","iopub.status.idle":"2022-12-08T14:55:57.797173Z","shell.execute_reply.started":"2022-12-08T14:55:57.796820Z","shell.execute_reply":"2022-12-08T14:55:57.796856Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test_pred_df['labels'] = test_pred_df.labels.apply(lambda x: list(set(x)))\ntest_pred_df['labels'] = test_pred_df.labels.apply(lambda x: m20_mult(x, len(x)))\ntest_pred_df['labels'] = test_pred_df.labels.apply(lambda x: \" \".join(map(str,x)))\ntest_pred_df = test_pred_df.drop(['labels_clicks', 'labels_carts', 'labels_orders'], axis=1)\nclicks_pred_df = test_pred_df.copy()\nclicks_pred_df.session = clicks_pred_df.session.apply(lambda x: str(x) + '_clicks')\norders_pred_df = test_pred_df.copy()\norders_pred_df.session = orders_pred_df.session.apply(lambda x: str(x) + '_orders')\ncarts_pred_df = test_pred_df.copy()\ncarts_pred_df.session = carts_pred_df.session.apply(lambda x: str(x) + '_carts')","metadata":{"execution":{"iopub.status.busy":"2022-12-08T14:55:57.799122Z","iopub.status.idle":"2022-12-08T14:55:57.800277Z","shell.execute_reply.started":"2022-12-08T14:55:57.799946Z","shell.execute_reply":"2022-12-08T14:55:57.799977Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"pred_df = pd.concat(\n    [clicks_pred_df, orders_pred_df, carts_pred_df]\n)\npred_df.columns = ['session_type', 'labels']","metadata":{"execution":{"iopub.status.busy":"2022-12-08T14:55:57.802058Z","iopub.status.idle":"2022-12-08T14:55:57.803291Z","shell.execute_reply.started":"2022-12-08T14:55:57.802948Z","shell.execute_reply":"2022-12-08T14:55:57.802982Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**🟦EN** Predict the results\n**🟥ES** Predecimos los resultados","metadata":{}},{"cell_type":"code","source":"pred_df = pred_df.sort_values(by='session_type').reset_index()\npred_df = pred_df.drop('index', axis=1)","metadata":{"execution":{"iopub.status.busy":"2022-12-08T14:55:57.805033Z","iopub.status.idle":"2022-12-08T14:55:57.806336Z","shell.execute_reply.started":"2022-12-08T14:55:57.806005Z","shell.execute_reply":"2022-12-08T14:55:57.806039Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Submission 📤","metadata":{}},{"cell_type":"code","source":"pred_df.to_csv(\"submission.csv\", index=False)","metadata":{"execution":{"iopub.status.busy":"2022-12-08T14:55:57.807881Z","iopub.status.idle":"2022-12-08T14:55:57.808523Z","shell.execute_reply.started":"2022-12-08T14:55:57.808300Z","shell.execute_reply":"2022-12-08T14:55:57.808323Z"},"trusted":true},"execution_count":null,"outputs":[]}]}