{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"code","source":"# This Python 3 environment comes with many helpful analytics libraries installed\n# It is defined by the kaggle/python Docker image: https://github.com/kaggle/docker-python\n# For example, here's several helpful packages to load\n\nimport numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\n\n# Input data files are available in the read-only \"../input/\" directory\n# For example, running this (by clicking run or pressing Shift+Enter) will list all files under the input directory\n\nimport os\nfor dirname, _, filenames in os.walk('/kaggle/input'):\n    for filename in filenames:\n        print(os.path.join(dirname, filename))\n\n# You can write up to 20GB to the current directory (/kaggle/working/) that gets preserved as output when you create a version using \"Save & Run All\" \n# You can also write temporary files to /kaggle/temp/, but they won't be saved outside of the current session","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","execution":{"iopub.status.busy":"2023-01-16T03:13:11.787262Z","iopub.execute_input":"2023-01-16T03:13:11.787659Z","iopub.status.idle":"2023-01-16T03:13:11.798730Z","shell.execute_reply.started":"2023-01-16T03:13:11.787629Z","shell.execute_reply":"2023-01-16T03:13:11.797483Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Goal of the Competition -\n\nThe goal of this competition is to predict e-commerce clicks, cart additions, and orders. You'll build a multi-objective recommender system based on previous events in a user session.","metadata":{}},{"cell_type":"markdown","source":"### Context - \nOnline shoppers have their pick of millions of products from large retailers. While such variety may be impressive, having so many options to explore can be overwhelming, resulting in shoppers leaving with empty carts. This neither benefits shoppers seeking to make a purchase nor retailers that missed out on sales. This is one reason online retailers rely on recommender systems to guide shoppers to products that best match their interests and motivations. Using data science to enhance retailers' ability to predict which products each customer actually wants to see, add to their cart, and order at any given moment of their visit in real-time could improve your customer experience the next time you shop online with your favorite retailer.","metadata":{}},{"cell_type":"markdown","source":"### You’ll build a single entry to predict click-through, add-to-cart, and conversion rates based on previous same-session events.","metadata":{}},{"cell_type":"markdown","source":"Your work will help online retailers select more relevant items from a vast range to recommend to their customers based on their real-time behavior. Improving recommendations will ensure navigating through seemingly endless options is more effortless and engaging for shoppers.","metadata":{}},{"cell_type":"markdown","source":"#Submission File - \n#For each session id and type combination in the test set, you must predict the aid values in the label column, which is space delimited. You can predict up to 20 aid values per row. The file should contain a header and have the following format:\n\n#session_type,labels\n#12906577_clicks,135193 129431 119318 ...\n#12906577_carts,135193 129431 119318 ...\n#12906577_orders,135193 129431 119318 ...\n#12906578_clicks, 135193 129431 119318 ...\n#etc. ","metadata":{"execution":{"iopub.status.busy":"2023-01-14T11:14:19.755825Z","iopub.execute_input":"2023-01-14T11:14:19.756226Z","iopub.status.idle":"2023-01-14T11:14:19.761314Z","shell.execute_reply.started":"2023-01-14T11:14:19.756193Z","shell.execute_reply":"2023-01-14T11:14:19.759931Z"}}},{"cell_type":"markdown","source":"### Files ---\n\ntrain.jsonl - the training data, which contains full session data\n\nsession - the unique session id\n\nevents - the time ordered sequence of events in the session\n\naid - the article id (product code) of the associated event\n\nts - the Unix timestamp of the event\n\ntype - the event type, i.e., whether a product was clicked, added to the user's cart, or ordered during the session\n\n\ntest.jsonl - the test data, which contains truncated session data  .your task is to predict the next aid clicked after the session truncation, as well as the the remaining aids that are added to carts and orders; you may predict up to 20 values for each session type\n\n\nsample_submission.csv - a sample submission file in the correct format","metadata":{}},{"cell_type":"markdown","source":"### Recommendation Systems Terminology - \n\nThe information a system uses to make recommendations. Queries can be a combination of the following:\n\nuser information - the id of the user , items that users previously interacted with\n\nadditional context - time of day , the user's device","metadata":{}},{"cell_type":"markdown","source":"### Embedding - \nA mapping from a discrete set (the set of queries, or the set of items to recommend) to a vector space called the embedding space. Many recommendation systems rely on learning an appropriate embedding representation of the queries and items.","metadata":{}},{"cell_type":"markdown","source":"### Recommendation Systems Overview -  One common architecture for recommendation systems consists of the following components: candidate generation , scoring , re-ranking .\n\nCandidate Generation - In this first stage, the system starts from a potentially huge corpus and generates a much smaller subset of candidates. For example, the candidate generator in YouTube reduces billions of videos down to hundreds or thousands. The model needs to evaluate queries quickly given the enormous size of the corpus. A given model may provide multiple candidate generators, each nominating a different subset of candidates.\n\nScoring - Next, another model scores and ranks the candidates in order to select the set of items (on the order of 10) to display to the user . \n\nRe-ranking - Finally, the system must take into account additional constraints for the final ranking. For example, the system removes items that the user explicitly disliked or boosts the score of fresher content. Re-ranking can also help ensure diversity, freshness, and fairness.","metadata":{}},{"cell_type":"markdown","source":"### Candidate Generation Overview - \n\n\nCandidate generation is the first stage of recommendation. Given a query, the system generates a set of relevant candidates. The following table shows two common candidate generation approaches:\n\ncontent-based filtering - \tUses similarity between items to recommend items similar to what the user likes.\n\ncollaborative filtering - Uses similarities between queries and items simultaneously to provide recommendations.","metadata":{}},{"cell_type":"markdown","source":"### Embedding Space - \n Both content-based and collaborative filtering map each item and each query (or context) to an embedding vector in a common embedding space \n. Typically, the embedding space is low-dimensional (that is, \n is much smaller than the size of the corpus), and captures some latent structure of the item","metadata":{"execution":{"iopub.status.busy":"2023-01-14T11:50:56.005906Z","iopub.execute_input":"2023-01-14T11:50:56.006274Z","iopub.status.idle":"2023-01-14T11:50:56.013808Z","shell.execute_reply.started":"2023-01-14T11:50:56.006245Z","shell.execute_reply":"2023-01-14T11:50:56.012382Z"}}},{"cell_type":"markdown","source":"","metadata":{}},{"cell_type":"markdown","source":"### Similarity Measures - \n\nA similarity measure is a function that takes a pair of embeddings and returns a scalar measuring their similarity. \n\nTo determine the degree of similarity, most recommendation systems rely on one or more of the following: cosine , dot product , Euclidean distance\n\ncosine - This is simply the cosine of the angle between the two vectors, \n\nDot Product - The dot product of two vectors . Thus, if the embeddings are normalized, then dot-product and cosine coincide.","metadata":{"execution":{"iopub.status.busy":"2023-01-14T11:55:16.355548Z","iopub.execute_input":"2023-01-14T11:55:16.355929Z","iopub.status.idle":"2023-01-14T11:55:16.365983Z","shell.execute_reply.started":"2023-01-14T11:55:16.355900Z","shell.execute_reply":"2023-01-14T11:55:16.364536Z"}}},{"cell_type":"markdown","source":"### Which Similarity Measure to Choose - \nCompared to the cosine, the dot product similarity is sensitive to the norm of the embedding. That is, the larger the norm of an embedding, the higher the similarity (for items with an acute angle) and the more likely the item is to be recommended.     This can affect recommendations as follows:\n\nItems that appear very frequently in the training set tend to have embeddings with large norms. If capturing popularity information is desirable, then you should prefer dot product. However, if you're not careful, the popular items may end up dominating the recommendations.\n\nItems that appear very rarely may not be updated frequently during training. Consequently, if they are initialized with a large norm, the system may recommend rare items over more relevant items. To avoid this problem, be careful about embedding initialization, and use appropriate regularization.","metadata":{}},{"cell_type":"markdown","source":"# Content-based Filtering - \nContent-based filtering uses item features to recommend other items similar to what the user likes, based on their previous actions or explicit feedback.\n\n\n# Content-based Filtering Advantages - \nThe model doesn't need any data about other users, since the recommendations are specific to this user. This makes it easier to scale to a large number of users.\n\nThe model can capture the specific interests of a user, and can recommend niche items that very few other users are interested in.\n\n\n# Content-based Filtering Disadvantages - \nSince the feature representation of the items are hand-engineered to some extent, this technique requires a lot of domain knowledge. Therefore, the model can only be as good as the hand-engineered features.\n\nThe model can only make recommendations based on existing interests of the user. In other words, the model has limited ability to expand on the users' existing interests.","metadata":{}},{"cell_type":"markdown","source":"# Collaborative Filtering - \nTo address some of the limitations of content-based filtering, collaborative filtering uses similarities between users and items simultaneously to provide recommendations. This allows for serendipitous recommendations; that is, collaborative filtering models can recommend an item to user A based on the interests of a similar user B. Furthermore, the embeddings can be learned automatically, without relying on hand-engineering of features.","metadata":{}},{"cell_type":"code","source":"#Looking at the distance between the points seems to be a good way to estimate similarity, \n#You can find the distance using the formula for Euclidean distance between two points. \nfrom scipy import spatial\n\na = [1, 2]\nb = [2, 4]\nc = [2.5, 4]\nd = [4.5, 5]\n\n","metadata":{"execution":{"iopub.status.busy":"2023-01-16T03:13:11.836302Z","iopub.execute_input":"2023-01-16T03:13:11.836709Z","iopub.status.idle":"2023-01-16T03:13:11.944536Z","shell.execute_reply.started":"2023-01-16T03:13:11.836674Z","shell.execute_reply":"2023-01-16T03:13:11.942940Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"spatial.distance.euclidean(c, a)\n","metadata":{"execution":{"iopub.status.busy":"2023-01-16T03:13:11.947008Z","iopub.execute_input":"2023-01-16T03:13:11.947691Z","iopub.status.idle":"2023-01-16T03:13:11.955640Z","shell.execute_reply.started":"2023-01-16T03:13:11.947647Z","shell.execute_reply":"2023-01-16T03:13:11.954359Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"spatial.distance.euclidean(c, b)","metadata":{"execution":{"iopub.status.busy":"2023-01-16T03:13:11.957718Z","iopub.execute_input":"2023-01-16T03:13:11.958656Z","iopub.status.idle":"2023-01-16T03:13:11.968235Z","shell.execute_reply.started":"2023-01-16T03:13:11.958613Z","shell.execute_reply":"2023-01-16T03:13:11.966606Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"spatial.distance.euclidean(c, d)","metadata":{"execution":{"iopub.status.busy":"2023-01-16T03:13:11.971277Z","iopub.execute_input":"2023-01-16T03:13:11.972163Z","iopub.status.idle":"2023-01-16T03:13:11.981335Z","shell.execute_reply.started":"2023-01-16T03:13:11.972116Z","shell.execute_reply":"2023-01-16T03:13:11.980193Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"To calculate similarity using angle, you need a function that returns a higher similarity or smaller distance for a lower angle and a lower similarity or larger distance for a higher angle. The cosine of an angle is a function that decreases from 1 to -1 as the angle increases from 0 to 180.\n\n\nUse the cosine of the angle to find the similarity between two users. The higher the angle, the lower will be the cosine and thus, the lower will be the similarity of the users. You can also inverse the value of the cosine of the angle to get the cosine distance between the users by subtracting it from 1.","metadata":{"execution":{"iopub.status.busy":"2023-01-15T05:46:32.122586Z","iopub.execute_input":"2023-01-15T05:46:32.123066Z","iopub.status.idle":"2023-01-15T05:46:32.128509Z","shell.execute_reply.started":"2023-01-15T05:46:32.123032Z","shell.execute_reply":"2023-01-15T05:46:32.127176Z"}}},{"cell_type":"code","source":"from scipy import spatial\na = [1, 2]\nb = [2, 4]\nc = [2.5, 4]\nd = [4.5, 5]","metadata":{"execution":{"iopub.status.busy":"2023-01-16T03:13:11.983314Z","iopub.execute_input":"2023-01-16T03:13:11.984050Z","iopub.status.idle":"2023-01-16T03:13:11.991514Z","shell.execute_reply.started":"2023-01-16T03:13:11.984004Z","shell.execute_reply":"2023-01-16T03:13:11.990270Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"spatial.distance.cosine(c,a)","metadata":{"execution":{"iopub.status.busy":"2023-01-16T03:13:11.993585Z","iopub.execute_input":"2023-01-16T03:13:11.993926Z","iopub.status.idle":"2023-01-16T03:13:12.002256Z","shell.execute_reply.started":"2023-01-16T03:13:11.993895Z","shell.execute_reply":"2023-01-16T03:13:12.001416Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"spatial.distance.cosine(c,b)\n","metadata":{"execution":{"iopub.status.busy":"2023-01-16T03:13:12.003578Z","iopub.execute_input":"2023-01-16T03:13:12.004657Z","iopub.status.idle":"2023-01-16T03:13:12.012746Z","shell.execute_reply.started":"2023-01-16T03:13:12.004613Z","shell.execute_reply":"2023-01-16T03:13:12.011841Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"spatial.distance.cosine(a,b)","metadata":{"execution":{"iopub.status.busy":"2023-01-16T03:13:12.015546Z","iopub.execute_input":"2023-01-16T03:13:12.016495Z","iopub.status.idle":"2023-01-16T03:13:12.024242Z","shell.execute_reply.started":"2023-01-16T03:13:12.016451Z","shell.execute_reply":"2023-01-16T03:13:12.023128Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Surprise is a Python scikit for building and analyzing recommender systems that deal with explicit rating data.\n\n- Dataset module is used to load data from files\n   - To load a dataset, some of the available methods are:\n   \n        - Dataset.load_builtin()\n        - Dataset.load_from_file()\n        - Dataset.load_from_df()\n        \n        \n\n- Reader class is used to parse a file\n\n   - line_format is a string that stores the order of the data with field names separated by a space, as in \"item user rating\".\n- sep is used to specify separator between fields, such as ','.\n- rating_scale is used to specify the rating scale. The default is (1, 5).\n- skip_lines is used to indicate the number of lines to skip at the beginning of the file. The default is 0.","metadata":{}},{"cell_type":"markdown","source":"### K-Nearest Neighbours (k-NN)\nTo find the similarity, you simply have to configure the function by passing a dictionary as an argument to the recommender function. The dictionary should have the required keys, such as the following:\n\n  - name contains the similarity metric to use. Options are cosine, msd, pearson, or pearson_baseline. The default is msd.\n  - user_based is a boolean that tells whether the approach will be user-based or item-based. The default is True, which means the user-based approach will be used.\n  - min_support is the minimum number of common items needed between users to consider them for similarity. For the item-based approach, this corresponds to the minimum number of common users for two items.","metadata":{}},{"cell_type":"code","source":" from surprise import KNNWithMeans\n\n# # To use item-based cosine similarity\n# sim_options = {\n#     \"name\": \"cosine\",\n#     \"user_based\": False,  # Compute  similarities between items\n# }\n# algo = KNNWithMeans(sim_options=sim_options)","metadata":{"execution":{"iopub.status.busy":"2023-01-16T03:13:12.025901Z","iopub.execute_input":"2023-01-16T03:13:12.026596Z","iopub.status.idle":"2023-01-16T03:13:12.033799Z","shell.execute_reply.started":"2023-01-16T03:13:12.026554Z","shell.execute_reply":"2023-01-16T03:13:12.033061Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# from pathlib import Path\n# paths = Path(\"/kaggle/input/otto-recommender-system/train.jsonl\").glob(\"*.json\")\n# df = pd.DataFrame([pd.read_json(p, typ=\"series\") for p in paths])","metadata":{"execution":{"iopub.status.busy":"2023-01-16T03:13:12.036402Z","iopub.execute_input":"2023-01-16T03:13:12.037110Z","iopub.status.idle":"2023-01-16T03:13:12.048078Z","shell.execute_reply.started":"2023-01-16T03:13:12.037067Z","shell.execute_reply":"2023-01-16T03:13:12.047250Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# df","metadata":{"execution":{"iopub.status.busy":"2023-01-16T03:13:12.049682Z","iopub.execute_input":"2023-01-16T03:13:12.050374Z","iopub.status.idle":"2023-01-16T03:13:12.062792Z","shell.execute_reply.started":"2023-01-16T03:13:12.050332Z","shell.execute_reply":"2023-01-16T03:13:12.061537Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from tqdm import tqdm\nimport json\n\n# Open the JSONL file for reading\nwith open('/kaggle/input/otto-recommender-system/train.jsonl', 'r') as f:\n    # Initialize the dictionary to store the counts\n    counts = {}\n    \n    # Read the file line by line\n    for line in tqdm(f):\n        # Parse the JSON object\n        obj = json.loads(line)\n        \n        # Iterate over the events\n        for event in obj['events']:\n            # Get the aid for the event\n            aid = event['aid']\n            \n            # If the aid is not in the dictionary yet, initialize the count to 0\n            if aid not in counts:\n                counts[aid] = 0\n            \n            # Increment the count for the aid\n            counts[aid] += 1\n","metadata":{"execution":{"iopub.status.busy":"2023-01-16T03:13:12.064785Z","iopub.execute_input":"2023-01-16T03:13:12.065793Z","iopub.status.idle":"2023-01-16T03:20:01.666224Z","shell.execute_reply.started":"2023-01-16T03:13:12.065760Z","shell.execute_reply":"2023-01-16T03:20:01.664606Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_data = pd.DataFrame.from_dict(counts, orient='index', columns=['count'])\n\n# Set the aid column as the index\ntrain_data.index.name = 'aid'\ntrain_data = train_data.reset_index()\n\ntrain_data.head()\n","metadata":{"execution":{"iopub.status.busy":"2023-01-16T03:20:01.669389Z","iopub.execute_input":"2023-01-16T03:20:01.669811Z","iopub.status.idle":"2023-01-16T03:20:02.908233Z","shell.execute_reply.started":"2023-01-16T03:20:01.669774Z","shell.execute_reply":"2023-01-16T03:20:02.907115Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# # df_items = train_data.pivot_table(index='aid', columns=['aid'], values='count').max().unstack()\n# # df_items.head(3)\n\n# df_items = train_data.groupby(['count']).max().unstack()\n","metadata":{"execution":{"iopub.status.busy":"2023-01-16T03:20:02.909495Z","iopub.execute_input":"2023-01-16T03:20:02.909959Z","iopub.status.idle":"2023-01-16T03:20:02.915085Z","shell.execute_reply.started":"2023-01-16T03:20:02.909926Z","shell.execute_reply":"2023-01-16T03:20:02.913964Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_data.corr()","metadata":{"execution":{"iopub.status.busy":"2023-01-16T03:23:09.769110Z","iopub.execute_input":"2023-01-16T03:23:09.769588Z","iopub.status.idle":"2023-01-16T03:23:09.856006Z","shell.execute_reply.started":"2023-01-16T03:23:09.769557Z","shell.execute_reply":"2023-01-16T03:23:09.854713Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"\nfrom sklearn.cluster import KMeans\nimport matplotlib.pyplot as plt\n\n\nkmeans = KMeans(n_clusters = 10, init = 'k-means++')\ny_kmeans = kmeans.fit_predict(train_data)\nplt.plot(y_kmeans, \".\")\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-01-16T03:23:20.377407Z","iopub.execute_input":"2023-01-16T03:23:20.378545Z","iopub.status.idle":"2023-01-16T03:23:52.665234Z","shell.execute_reply.started":"2023-01-16T03:23:20.378504Z","shell.execute_reply":"2023-01-16T03:23:52.664057Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"y_kmeans","metadata":{"execution":{"iopub.status.busy":"2023-01-16T03:26:18.088189Z","iopub.execute_input":"2023-01-16T03:26:18.088605Z","iopub.status.idle":"2023-01-16T03:26:18.095925Z","shell.execute_reply.started":"2023-01-16T03:26:18.088573Z","shell.execute_reply":"2023-01-16T03:26:18.094497Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df = pd.DataFrame(y_kmeans).to_csv(\"sample.csv\")\ndf = pd.read_csv(\"sample.csv\")\nprint(df)\n\n\n","metadata":{"execution":{"iopub.status.busy":"2023-01-16T03:31:55.476972Z","iopub.execute_input":"2023-01-16T03:31:55.477860Z","iopub.status.idle":"2023-01-16T03:31:57.859205Z","shell.execute_reply.started":"2023-01-16T03:31:55.477822Z","shell.execute_reply":"2023-01-16T03:31:57.858109Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"submission = pd.read_csv('/kaggle/input/otto-recommender-system/sample_submission.csv')\n\nsubmission.head()","metadata":{"execution":{"iopub.status.busy":"2023-01-16T03:45:30.950348Z","iopub.execute_input":"2023-01-16T03:45:30.950772Z","iopub.status.idle":"2023-01-16T03:45:34.911364Z","shell.execute_reply.started":"2023-01-16T03:45:30.950737Z","shell.execute_reply":"2023-01-16T03:45:34.910547Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"submission['labels'] = pd.Series(y_kmeans)\nsubmission.head()","metadata":{"execution":{"iopub.status.busy":"2023-01-16T03:45:38.820585Z","iopub.execute_input":"2023-01-16T03:45:38.821790Z","iopub.status.idle":"2023-01-16T03:45:39.088088Z","shell.execute_reply.started":"2023-01-16T03:45:38.821748Z","shell.execute_reply":"2023-01-16T03:45:39.086944Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"submission.to_csv('submission.csv', index=False, header=True)","metadata":{"execution":{"iopub.status.busy":"2023-01-16T03:46:17.299686Z","iopub.execute_input":"2023-01-16T03:46:17.300769Z","iopub.status.idle":"2023-01-16T03:46:24.195416Z","shell.execute_reply.started":"2023-01-16T03:46:17.300714Z","shell.execute_reply":"2023-01-16T03:46:24.194268Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]}]}