{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"code","source":"# This Python 3 environment comes with many helpful analytics libraries installed\n# It is defined by the kaggle/python Docker image: https://github.com/kaggle/docker-python\n# For example, here's several helpful packages to load\n\nimport numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\n\n# Input data files are available in the read-only \"../input/\" directory\n# For example, running this (by clicking run or pressing Shift+Enter) will list all files under the input directory\n\nimport os\nfor dirname, _, filenames in os.walk('/kaggle/input'):\n    for filename in filenames:\n        print(os.path.join(dirname, filename))\n\n# You can write up to 20GB to the current directory (/kaggle/working/) that gets preserved as output when you create a version using \"Save & Run All\" \n# You can also write temporary files to /kaggle/temp/, but they won't be saved outside of the current session","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","execution":{"iopub.status.busy":"2023-01-20T00:22:53.864645Z","iopub.execute_input":"2023-01-20T00:22:53.865356Z","iopub.status.idle":"2023-01-20T00:22:53.896934Z","shell.execute_reply.started":"2023-01-20T00:22:53.865261Z","shell.execute_reply":"2023-01-20T00:22:53.896063Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Goal of the Competition -\n\nThe goal of this competition is to predict e-commerce clicks, cart additions, and orders. You'll build a multi-objective recommender system based on previous events in a user session.","metadata":{}},{"cell_type":"markdown","source":"### Context - \nOnline shoppers have their pick of millions of products from large retailers. While such variety may be impressive, having so many options to explore can be overwhelming, resulting in shoppers leaving with empty carts. This neither benefits shoppers seeking to make a purchase nor retailers that missed out on sales. This is one reason online retailers rely on recommender systems to guide shoppers to products that best match their interests and motivations. Using data science to enhance retailers' ability to predict which products each customer actually wants to see, add to their cart, and order at any given moment of their visit in real-time could improve your customer experience the next time you shop online with your favorite retailer.","metadata":{}},{"cell_type":"markdown","source":"### You’ll build a single entry to predict click-through, add-to-cart, and conversion rates based on previous same-session events.","metadata":{}},{"cell_type":"markdown","source":"Your work will help online retailers select more relevant items from a vast range to recommend to their customers based on their real-time behavior. Improving recommendations will ensure navigating through seemingly endless options is more effortless and engaging for shoppers.","metadata":{}},{"cell_type":"markdown","source":"#Submission File - \n#For each session id and type combination in the test set, you must predict the aid values in the label column, which is space delimited. You can predict up to 20 aid values per row. The file should contain a header and have the following format:\n\n#session_type,labels\n#12906577_clicks,135193 129431 119318 ...\n#12906577_carts,135193 129431 119318 ...\n#12906577_orders,135193 129431 119318 ...\n#12906578_clicks, 135193 129431 119318 ...\n#etc. ","metadata":{"execution":{"iopub.status.busy":"2023-01-14T11:14:19.755825Z","iopub.execute_input":"2023-01-14T11:14:19.756226Z","iopub.status.idle":"2023-01-14T11:14:19.761314Z","shell.execute_reply.started":"2023-01-14T11:14:19.756193Z","shell.execute_reply":"2023-01-14T11:14:19.759931Z"}}},{"cell_type":"markdown","source":"### Files ---\n\ntrain.jsonl - the training data, which contains full session data\n\nsession - the unique session id\n\nevents - the time ordered sequence of events in the session\n\naid - the article id (product code) of the associated event\n\nts - the Unix timestamp of the event\n\ntype - the event type, i.e., whether a product was clicked, added to the user's cart, or ordered during the session\n\n\ntest.jsonl - the test data, which contains truncated session data  .your task is to predict the next aid clicked after the session truncation, as well as the the remaining aids that are added to carts and orders; you may predict up to 20 values for each session type\n\n\nsample_submission.csv - a sample submission file in the correct format","metadata":{}},{"cell_type":"markdown","source":"### Recommendation Systems Terminology - \n\nThe information a system uses to make recommendations. Queries can be a combination of the following:\n\nuser information - the id of the user , items that users previously interacted with\n\nadditional context - time of day , the user's device","metadata":{}},{"cell_type":"markdown","source":"### Embedding - \nA mapping from a discrete set (the set of queries, or the set of items to recommend) to a vector space called the embedding space. Many recommendation systems rely on learning an appropriate embedding representation of the queries and items.","metadata":{}},{"cell_type":"markdown","source":"### Recommendation Systems Overview -  One common architecture for recommendation systems consists of the following components: candidate generation , scoring , re-ranking .\n\nCandidate Generation - In this first stage, the system starts from a potentially huge corpus and generates a much smaller subset of candidates. For example, the candidate generator in YouTube reduces billions of videos down to hundreds or thousands. The model needs to evaluate queries quickly given the enormous size of the corpus. A given model may provide multiple candidate generators, each nominating a different subset of candidates.\n\nScoring - Next, another model scores and ranks the candidates in order to select the set of items (on the order of 10) to display to the user . \n\nRe-ranking - Finally, the system must take into account additional constraints for the final ranking. For example, the system removes items that the user explicitly disliked or boosts the score of fresher content. Re-ranking can also help ensure diversity, freshness, and fairness.","metadata":{}},{"cell_type":"markdown","source":"### Candidate Generation Overview - \n\n\nCandidate generation is the first stage of recommendation. Given a query, the system generates a set of relevant candidates. The following table shows two common candidate generation approaches:\n\ncontent-based filtering - \tUses similarity between items to recommend items similar to what the user likes.\n\ncollaborative filtering - Uses similarities between queries and items simultaneously to provide recommendations.","metadata":{}},{"cell_type":"markdown","source":"### Embedding Space - \n Both content-based and collaborative filtering map each item and each query (or context) to an embedding vector in a common embedding space \n. Typically, the embedding space is low-dimensional (that is, \n is much smaller than the size of the corpus), and captures some latent structure of the item","metadata":{"execution":{"iopub.status.busy":"2023-01-14T11:50:56.005906Z","iopub.execute_input":"2023-01-14T11:50:56.006274Z","iopub.status.idle":"2023-01-14T11:50:56.013808Z","shell.execute_reply.started":"2023-01-14T11:50:56.006245Z","shell.execute_reply":"2023-01-14T11:50:56.012382Z"}}},{"cell_type":"markdown","source":"","metadata":{}},{"cell_type":"markdown","source":"### Similarity Measures - \n\nA similarity measure is a function that takes a pair of embeddings and returns a scalar measuring their similarity. \n\nTo determine the degree of similarity, most recommendation systems rely on one or more of the following: cosine , dot product , Euclidean distance\n\ncosine - This is simply the cosine of the angle between the two vectors, \n\nDot Product - The dot product of two vectors . Thus, if the embeddings are normalized, then dot-product and cosine coincide.","metadata":{"execution":{"iopub.status.busy":"2023-01-14T11:55:16.355548Z","iopub.execute_input":"2023-01-14T11:55:16.355929Z","iopub.status.idle":"2023-01-14T11:55:16.365983Z","shell.execute_reply.started":"2023-01-14T11:55:16.355900Z","shell.execute_reply":"2023-01-14T11:55:16.364536Z"}}},{"cell_type":"markdown","source":"### Which Similarity Measure to Choose - \nCompared to the cosine, the dot product similarity is sensitive to the norm of the embedding. That is, the larger the norm of an embedding, the higher the similarity (for items with an acute angle) and the more likely the item is to be recommended.     This can affect recommendations as follows:\n\nItems that appear very frequently in the training set tend to have embeddings with large norms. If capturing popularity information is desirable, then you should prefer dot product. However, if you're not careful, the popular items may end up dominating the recommendations.\n\nItems that appear very rarely may not be updated frequently during training. Consequently, if they are initialized with a large norm, the system may recommend rare items over more relevant items. To avoid this problem, be careful about embedding initialization, and use appropriate regularization.","metadata":{}},{"cell_type":"markdown","source":"# Content-based Filtering - \nContent-based filtering uses item features to recommend other items similar to what the user likes, based on their previous actions or explicit feedback.\n\n\n# Content-based Filtering Advantages - \nThe model doesn't need any data about other users, since the recommendations are specific to this user. This makes it easier to scale to a large number of users.\n\nThe model can capture the specific interests of a user, and can recommend niche items that very few other users are interested in.\n\n\n# Content-based Filtering Disadvantages - \nSince the feature representation of the items are hand-engineered to some extent, this technique requires a lot of domain knowledge. Therefore, the model can only be as good as the hand-engineered features.\n\nThe model can only make recommendations based on existing interests of the user. In other words, the model has limited ability to expand on the users' existing interests.","metadata":{}},{"cell_type":"markdown","source":"# Collaborative Filtering - \nTo address some of the limitations of content-based filtering, collaborative filtering uses similarities between users and items simultaneously to provide recommendations. This allows for serendipitous recommendations; that is, collaborative filtering models can recommend an item to user A based on the interests of a similar user B. Furthermore, the embeddings can be learned automatically, without relying on hand-engineering of features.","metadata":{}},{"cell_type":"markdown","source":"To calculate similarity using angle, you need a function that returns a higher similarity or smaller distance for a lower angle and a lower similarity or larger distance for a higher angle. The cosine of an angle is a function that decreases from 1 to -1 as the angle increases from 0 to 180.\n\n\nUse the cosine of the angle to find the similarity between two users. The higher the angle, the lower will be the cosine and thus, the lower will be the similarity of the users. You can also inverse the value of the cosine of the angle to get the cosine distance between the users by subtracting it from 1.","metadata":{"execution":{"iopub.status.busy":"2023-01-15T05:46:32.122586Z","iopub.execute_input":"2023-01-15T05:46:32.123066Z","iopub.status.idle":"2023-01-15T05:46:32.128509Z","shell.execute_reply.started":"2023-01-15T05:46:32.123032Z","shell.execute_reply":"2023-01-15T05:46:32.127176Z"}}},{"cell_type":"code","source":"from scipy import spatial\na = [1, 2]\nb = [2, 4]\nc = [2.5, 4]\nd = [4.5, 5]","metadata":{"execution":{"iopub.status.busy":"2023-01-20T00:22:54.084310Z","iopub.execute_input":"2023-01-20T00:22:54.084988Z","iopub.status.idle":"2023-01-20T00:22:54.349847Z","shell.execute_reply.started":"2023-01-20T00:22:54.084931Z","shell.execute_reply":"2023-01-20T00:22:54.348521Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"spatial.distance.cosine(c,a)","metadata":{"execution":{"iopub.status.busy":"2023-01-20T00:22:54.351783Z","iopub.execute_input":"2023-01-20T00:22:54.352157Z","iopub.status.idle":"2023-01-20T00:22:54.363798Z","shell.execute_reply.started":"2023-01-20T00:22:54.352124Z","shell.execute_reply":"2023-01-20T00:22:54.362489Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"spatial.distance.cosine(c,b)\n","metadata":{"execution":{"iopub.status.busy":"2023-01-20T00:22:54.365462Z","iopub.execute_input":"2023-01-20T00:22:54.365812Z","iopub.status.idle":"2023-01-20T00:22:54.375646Z","shell.execute_reply.started":"2023-01-20T00:22:54.365782Z","shell.execute_reply":"2023-01-20T00:22:54.374595Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"spatial.distance.cosine(a,b)","metadata":{"execution":{"iopub.status.busy":"2023-01-20T00:22:54.378387Z","iopub.execute_input":"2023-01-20T00:22:54.378796Z","iopub.status.idle":"2023-01-20T00:22:54.386467Z","shell.execute_reply.started":"2023-01-20T00:22:54.378765Z","shell.execute_reply":"2023-01-20T00:22:54.385300Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Surprise is a Python scikit for building and analyzing recommender systems that deal with explicit rating data.\n\n- Dataset module is used to load data from files\n   - To load a dataset, some of the available methods are:\n   \n        - Dataset.load_builtin()\n        - Dataset.load_from_file()\n        - Dataset.load_from_df()\n        \n        \n\n- Reader class is used to parse a file\n\n   - line_format is a string that stores the order of the data with field names separated by a space, as in \"item user rating\".\n- sep is used to specify separator between fields, such as ','.\n- rating_scale is used to specify the rating scale. The default is (1, 5).\n- skip_lines is used to indicate the number of lines to skip at the beginning of the file. The default is 0.","metadata":{}},{"cell_type":"markdown","source":"# Dataset - \n - session - the unique session id\n - events - the time ordered sequence of events in the session\n - aid - the article id (product code) of the associated event\n - ts - the Unix timestamp of the event\n - type - the event type, i.e., whether a product was clicked, added to the user's cart, or ordered during the session","metadata":{}},{"cell_type":"code","source":"# from tqdm import tqdm\n# import json\n\n# # Open the JSONL file for reading\n# with open('/kaggle/input/otto-recommender-system/train.jsonl', 'r') as f:\n#     # Initialize the dictionary to store the counts\n#     counts = {}\n    \n#     # Read the file line by line\n#     for line in tqdm(f):\n#         # Parse the JSON object\n#         obj = json.loads(line)\n        \n#         # Iterate over the events\n#         for event in obj['events']:\n#             # Get the aid for the event\n#             aid = event['aid']\n            \n#             # If the aid is not in the dictionary yet, initialize the count to 0\n#             if aid not in counts:\n#                 counts[aid] = 0\n            \n#             # Increment the count for the aid\n#             counts[aid] += 1\n","metadata":{"execution":{"iopub.status.busy":"2023-01-20T00:22:54.388303Z","iopub.execute_input":"2023-01-20T00:22:54.389127Z","iopub.status.idle":"2023-01-20T00:22:54.396205Z","shell.execute_reply.started":"2023-01-20T00:22:54.389085Z","shell.execute_reply":"2023-01-20T00:22:54.394781Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"!pip install polars\n","metadata":{"execution":{"iopub.status.busy":"2023-01-20T00:22:54.398118Z","iopub.execute_input":"2023-01-20T00:22:54.398937Z","iopub.status.idle":"2023-01-20T00:23:11.990473Z","shell.execute_reply.started":"2023-01-20T00:22:54.398891Z","shell.execute_reply":"2023-01-20T00:23:11.988997Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import polars as pl\n\ntrain = pl.read_parquet('../input/otto-full-optimized-memory-footprint/train.parquet')\ntest = pl.read_parquet('../input/otto-full-optimized-memory-footprint/test.parquet')","metadata":{"execution":{"iopub.status.busy":"2023-01-20T00:23:11.992413Z","iopub.execute_input":"2023-01-20T00:23:11.992814Z","iopub.status.idle":"2023-01-20T00:23:25.204233Z","shell.execute_reply.started":"2023-01-20T00:23:11.992776Z","shell.execute_reply":"2023-01-20T00:23:25.203046Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_data = train.head(50)\n","metadata":{"execution":{"iopub.status.busy":"2023-01-20T00:23:25.205544Z","iopub.execute_input":"2023-01-20T00:23:25.205909Z","iopub.status.idle":"2023-01-20T00:23:25.211264Z","shell.execute_reply.started":"2023-01-20T00:23:25.205875Z","shell.execute_reply":"2023-01-20T00:23:25.210049Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_data.head()","metadata":{"execution":{"iopub.status.busy":"2023-01-20T00:23:25.212788Z","iopub.execute_input":"2023-01-20T00:23:25.214109Z","iopub.status.idle":"2023-01-20T00:23:25.225424Z","shell.execute_reply.started":"2023-01-20T00:23:25.214062Z","shell.execute_reply":"2023-01-20T00:23:25.224176Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Import linear_kernel\nfrom sklearn.metrics.pairwise import linear_kernel\n\n# Compute the cosine similarity matrix\ncosine_sim = linear_kernel(train_data)","metadata":{"execution":{"iopub.status.busy":"2023-01-20T00:23:25.229069Z","iopub.execute_input":"2023-01-20T00:23:25.229483Z","iopub.status.idle":"2023-01-20T00:23:25.699577Z","shell.execute_reply.started":"2023-01-20T00:23:25.229449Z","shell.execute_reply":"2023-01-20T00:23:25.698516Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"cosine_sim.shape\n","metadata":{"execution":{"iopub.status.busy":"2023-01-20T00:23:25.701426Z","iopub.execute_input":"2023-01-20T00:23:25.702125Z","iopub.status.idle":"2023-01-20T00:23:25.709471Z","shell.execute_reply.started":"2023-01-20T00:23:25.702086Z","shell.execute_reply":"2023-01-20T00:23:25.708058Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"cosine_sim[0]\n","metadata":{"execution":{"iopub.status.busy":"2023-01-20T00:23:25.711067Z","iopub.execute_input":"2023-01-20T00:23:25.711527Z","iopub.status.idle":"2023-01-20T00:23:25.722599Z","shell.execute_reply.started":"2023-01-20T00:23:25.711495Z","shell.execute_reply":"2023-01-20T00:23:25.721426Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"- Matrix factorization is to, obviously, factorize a matrix, i.e. to find out two (or more) matrices such that when you multiply them you will get back the original matrix.","metadata":{}},{"cell_type":"markdown","source":"# Matrix factorization follows the following:\n\n - Initialize two random matrices a and b with dimensions m by j and j by n such that when multiplied, their dimension matches the original matrix z (that has dimensions m by n).\n\n - Multiply a by b to achieve an estimate for z.\n\n- Subtract z from y for the known values of z, or some other loss function, to evaluate how far off the estimate is from the real matrix.\n\n- Use gradient descent formulas to adjust each of the values in a and b in the right direction.\n\n- Repeat steps 2 to 4 repeatedly until the error has reached a reasonable value.\n- By multiplying a by b, we now have an estimate for z that not only closely matches the known values of z, but also provides an estimate for the unknown values.","metadata":{}},{"cell_type":"code","source":"#Matrix factorization\nfrom sklearn.decomposition import NMF\nmodel = NMF(n_components=2, init='random', random_state=0)\nW = model.fit_transform(train_data)\nH = model.components_","metadata":{"execution":{"iopub.status.busy":"2023-01-20T00:23:25.723851Z","iopub.execute_input":"2023-01-20T00:23:25.724442Z","iopub.status.idle":"2023-01-20T00:23:25.824763Z","shell.execute_reply.started":"2023-01-20T00:23:25.724409Z","shell.execute_reply":"2023-01-20T00:23:25.823787Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"H","metadata":{"execution":{"iopub.status.busy":"2023-01-20T00:23:25.826144Z","iopub.execute_input":"2023-01-20T00:23:25.827038Z","iopub.status.idle":"2023-01-20T00:23:25.836173Z","shell.execute_reply.started":"2023-01-20T00:23:25.826999Z","shell.execute_reply":"2023-01-20T00:23:25.834769Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import numpy as np\nR_estimated = np.dot(W, H)","metadata":{"execution":{"iopub.status.busy":"2023-01-20T00:23:25.837830Z","iopub.execute_input":"2023-01-20T00:23:25.838256Z","iopub.status.idle":"2023-01-20T00:23:25.845248Z","shell.execute_reply.started":"2023-01-20T00:23:25.838222Z","shell.execute_reply":"2023-01-20T00:23:25.844194Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"R_estimated","metadata":{"execution":{"iopub.status.busy":"2023-01-20T00:53:59.724652Z","iopub.execute_input":"2023-01-20T00:53:59.725270Z","iopub.status.idle":"2023-01-20T00:53:59.739885Z","shell.execute_reply.started":"2023-01-20T00:53:59.725220Z","shell.execute_reply":"2023-01-20T00:53:59.737710Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import pandas as pd\ndf = pd.DataFrame(R_estimated).to_csv(\"sample.csv\")\ndf = pd.read_csv(\"sample.csv\")\nprint(df)\n\n\n","metadata":{"execution":{"iopub.status.busy":"2023-01-20T00:23:25.861182Z","iopub.execute_input":"2023-01-20T00:23:25.862055Z","iopub.status.idle":"2023-01-20T00:23:25.902212Z","shell.execute_reply.started":"2023-01-20T00:23:25.861999Z","shell.execute_reply":"2023-01-20T00:23:25.900732Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import pandas as pd\nsubmission = pd.read_csv('/kaggle/input/otto-recommender-system/sample_submission.csv')\n\nsubmission.head()","metadata":{"execution":{"iopub.status.busy":"2023-01-20T00:23:25.903688Z","iopub.execute_input":"2023-01-20T00:23:25.904069Z","iopub.status.idle":"2023-01-20T00:23:32.599753Z","shell.execute_reply.started":"2023-01-20T00:23:25.904036Z","shell.execute_reply":"2023-01-20T00:23:32.598443Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# To make data 1_Dimensional\n# import itertools\n# R_estimated1 = itertools.chain.from_iterable(R_estimated)","metadata":{"execution":{"iopub.status.busy":"2023-01-20T00:50:35.995469Z","iopub.execute_input":"2023-01-20T00:50:35.995940Z","iopub.status.idle":"2023-01-20T00:50:36.002504Z","shell.execute_reply.started":"2023-01-20T00:50:35.995904Z","shell.execute_reply":"2023-01-20T00:50:36.000815Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Make R_estimated 1-Dimensional\n# R_estimated2 = pd.DataFrame(R_estimated)\n\nR_estimated3 = R_estimated[0] + R_estimated[1]\n","metadata":{"execution":{"iopub.status.busy":"2023-01-20T00:58:07.764148Z","iopub.execute_input":"2023-01-20T00:58:07.764611Z","iopub.status.idle":"2023-01-20T00:58:07.770493Z","shell.execute_reply.started":"2023-01-20T00:58:07.764579Z","shell.execute_reply":"2023-01-20T00:58:07.769150Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"submission['labels'] = pd.Series(R_estimated3)\nsubmission.head()","metadata":{"execution":{"iopub.status.busy":"2023-01-20T00:59:14.785751Z","iopub.execute_input":"2023-01-20T00:59:14.786191Z","iopub.status.idle":"2023-01-20T00:59:14.947659Z","shell.execute_reply.started":"2023-01-20T00:59:14.786152Z","shell.execute_reply":"2023-01-20T00:59:14.946313Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Your submission files contains 5015409 rows ( where each 1/3 corresponds to a specific type clicks/carts/orders).\n\nYour variable R_estimated does not fit the actual shape of your dataframe. Do not forget, that you need to pass a \"labels\" columns where each rows should be a string of candidates ( exemple : 92400 78500 3400 ) using a whitespace between each of them.","metadata":{}},{"cell_type":"code","source":"submission.to_csv('submission.csv', index=False, header=True)","metadata":{"execution":{"iopub.status.busy":"2023-01-20T00:58:30.663140Z","iopub.execute_input":"2023-01-20T00:58:30.663609Z","iopub.status.idle":"2023-01-20T00:58:37.154368Z","shell.execute_reply.started":"2023-01-20T00:58:30.663574Z","shell.execute_reply":"2023-01-20T00:58:37.152650Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]}]}