{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# 🏪 OTTO – Multi-Objective Recommender System\n**Dataset Description**\n\nThe goal of this competition is to predict e-commerce clicks, cart additions, and orders. You'll build a multi-objective recommender system based on previous events in a user session.\n\nThe training data contains full e-commerce session information. For each session in the test data, your task it to predict the aid values for each session type thats occur after the last timestamp ts in the test session. In other words, the test data contains sessions truncated by timestamp, and you are to predict what occurs after the point of truncation.\n\nFor additional background, please see the published OTTO Recommender Systems Dataset GitHub.\n\n**Files**\n\n* **train.jsonl** - the training data, which contains full session data\n    * session - the unique session id\n    * events - the time ordered sequence of events in the session\n    * aid - the article id (product code) of the associated event\n    * ts - the Unix timestamp of the event\n    * type - the event type, i.e., whether a product was clicked, added to the user's cart, or ordered during the session\n* **test.jsonl** - the test data, which contains truncated session data\nyour task is to predict the next aid clicked after the session truncation, as well as the the remaining aids that are added to carts and orders; you may predict up to 20 values for each session type\nsample_submission.csv - a sample submission file in the correct format\n\n\n**Credits**\nI took majority of the ideas from this Notebook, Credits to the author.\nhttps://www.kaggle.com/code/columbia2131/otto-read-a-chunk-of-jsonl-to-manageable-df\n\n...","metadata":{}},{"cell_type":"code","source":"# This Python 3 environment comes with many helpful analytics libraries installed\n# It is defined by the kaggle/python Docker image: https://github.com/kaggle/docker-python\n# For example, here's several helpful packages to load\n\nimport numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\n\n# Input data files are available in the read-only \"../input/\" directory\n# For example, running this (by clicking run or pressing Shift+Enter) will list all files under the input directory\n\nimport os\nfor dirname, _, filenames in os.walk('/kaggle/input'):\n    for filename in filenames:\n        print(os.path.join(dirname, filename))\n\n# You can write up to 20GB to the current directory (/kaggle/working/) that gets preserved as output when you create a version using \"Save & Run All\" \n# You can also write temporary files to /kaggle/temp/, but they won't be saved outside of the current session","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","execution":{"iopub.status.busy":"2022-11-30T01:44:41.668702Z","iopub.execute_input":"2022-11-30T01:44:41.669267Z","iopub.status.idle":"2022-11-30T01:44:41.698993Z","shell.execute_reply.started":"2022-11-30T01:44:41.669121Z","shell.execute_reply":"2022-11-30T01:44:41.697958Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\n# Import all the nesesary libraries...\nfrom pathlib import Path","metadata":{"execution":{"iopub.status.busy":"2022-11-30T01:44:41.700694Z","iopub.execute_input":"2022-11-30T01:44:41.701076Z","iopub.status.idle":"2022-11-30T01:44:41.707750Z","shell.execute_reply.started":"2022-11-30T01:44:41.701047Z","shell.execute_reply":"2022-11-30T01:44:41.706599Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\n# Read the JSON file and creates a pandas df...\ndata_path = Path('/kaggle/input/otto-recommender-system/')\n\n# Create an empty dataframe to store the processed values...\ntrn_data = pd.DataFrame()\n\n\n\n# Read the file in chunks...\nSIZE = 100\nchunks = pd.read_json(data_path / 'train.jsonl', lines = True, chunksize = SIZE)","metadata":{"execution":{"iopub.status.busy":"2022-11-30T01:44:41.709075Z","iopub.execute_input":"2022-11-30T01:44:41.709482Z","iopub.status.idle":"2022-11-30T01:44:41.725062Z","shell.execute_reply.started":"2022-11-30T01:44:41.709452Z","shell.execute_reply":"2022-11-30T01:44:41.723905Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\nlimit = 3 # Maximun number of chunks to load...\n# Read each chunk and add them to the dataframe...\nfor idx, chunk in enumerate(chunks):\n    # Creates date dictionary to store the values...\n    event_dict = {'session': [], 'aid': [], 'timestamp': [], 'even_type': []}\n    \n    # Iterates on each of the chunks and concatenates the data to trn_data...\n    if idx < limit:\n        print(f'Processing chunk: {idx} ...')\n        for session, events in zip(chunk['session'].tolist(), chunk['events'].tolist()):\n            for event in events:\n                event_dict['session'].append(session)\n                event_dict['aid'].append(event['aid'])\n                event_dict['timestamp'].append(event['ts'])\n                event_dict['even_type'].append(event['type'])\n            chunk_session = pd.DataFrame(event_dict)\n            trn_data = pd.concat([trn_data, chunk_session])\n    else: break\n        \ntrn_data = trn_data.reset_index(drop=True)","metadata":{"execution":{"iopub.status.busy":"2022-11-30T01:44:41.728024Z","iopub.execute_input":"2022-11-30T01:44:41.728709Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\n# Review a summary of the datatframe...\ntrn_data.info()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\n# Display the first five rows of data...\ntrn_data.sample(10)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]}]}