{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"**What are you trying to do in this notebook?**\n\nTo predict the probability of each team scoring within the next 10 seconds of the game given a snapshot from a Rocket League match. Sounds awesome, right? Well, it's not that simple. The training data is fairly large; trying to read and model it in a single go might pose some challenges. The purpose of this month's competition is for me to explore ways I can take a big dataset and make it manageable within the time and resources I have. For most people, typical brute force approaches aren't going to work well.\n\n**Can I scale down the dataset?**\n\nCan I use, e.g., online learning methods that allow me to train from the data one row at a time? (FTLR is a great place to start if I'm not familiar with online learning! e.g., this notebook)\nCan I figure out a nice set of features to reduce the dataset down to?\nIn addition to that challenge, while my predictions must be made pointwise, the training data is made up of timeseries—maybe I can use that temporal information to improve my model? This competition also has plenty of opportunity for data visualizations. Let's see some pretty graphs!\n\nSo, share your ideas about tackling this beast of a dataset and have a great time!\n\n\n**Why are you trying it?**\n\nThe dataset consists of sequences of snapshots of the state of a Rocket League match, including position and velocity of all players and the ball, as well as extra information. The goal of the competition is to predict -- from a given snapshot in the game -- for each team, the probability that they will score within the next 10 seconds of game time.\n\nThe data was taken from professional Rocket League matches. Each event consists of a chronological series of frames recorded at 10 frames per second. All events begin with a kickoff, and most end in one team scoring a goal, but some are truncated and end with no goal scored due to circumstances which can cause gameplay strategies to shift, for example 1) nearing end of regulation (where the game continues until the ball touches the ground) or 2) becoming non-competitive, eg one team winning by 3+ goals with little time remaining.","metadata":{}},{"cell_type":"code","source":"# This Python 3 environment comes with many helpful analytics libraries installed\n# It is defined by the kaggle/python Docker image: https://github.com/kaggle/docker-python\n# For example, here's several helpful packages to load\n\nimport numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\n\n# Input data files are available in the read-only \"../input/\" directory\n# For example, running this (by clicking run or pressing Shift+Enter) will list all files under the input directory\n\nimport os\nfor dirname, _, filenames in os.walk('/kaggle/input'):\n    for filename in filenames:\n        print(os.path.join(dirname, filename))\n\n# You can write up to 20GB to the current directory (/kaggle/working/) that gets preserved as output when you create a version using \"Save & Run All\" \n# You can also write temporary files to /kaggle/temp/, but they won't be saved outside of the current session","metadata":{"execution":{"iopub.status.busy":"2022-10-19T11:51:58.909087Z","iopub.execute_input":"2022-10-19T11:51:58.909518Z","iopub.status.idle":"2022-10-19T11:51:58.999092Z","shell.execute_reply.started":"2022-10-19T11:51:58.909482Z","shell.execute_reply":"2022-10-19T11:51:58.997713Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import pandas as pd\nimport numpy as np","metadata":{"execution":{"iopub.status.busy":"2022-10-19T11:51:59.001866Z","iopub.execute_input":"2022-10-19T11:51:59.002697Z","iopub.status.idle":"2022-10-19T11:51:59.007727Z","shell.execute_reply.started":"2022-10-19T11:51:59.002650Z","shell.execute_reply":"2022-10-19T11:51:59.006863Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import joblib","metadata":{"execution":{"iopub.status.busy":"2022-10-19T11:51:59.009201Z","iopub.execute_input":"2022-10-19T11:51:59.010050Z","iopub.status.idle":"2022-10-19T11:51:59.047170Z","shell.execute_reply.started":"2022-10-19T11:51:59.010018Z","shell.execute_reply":"2022-10-19T11:51:59.046010Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sub = pd.read_csv('../input/tabular-playground-series-oct-2022/sample_submission.csv')\n\nFEATURES = ['ball_pos_x', 'ball_pos_y', 'ball_pos_z', 'ball_vel_x', 'ball_vel_y',\n       'ball_vel_z', 'p0_pos_x', 'p0_pos_y', 'p0_pos_z', 'p0_vel_x',\n       'p0_vel_y', 'p0_vel_z', 'p0_boost', 'p1_pos_x', 'p1_pos_y', 'p1_pos_z',\n       'p1_vel_x', 'p1_vel_y', 'p1_vel_z', 'p1_boost', 'p2_pos_x', 'p2_pos_y',\n       'p2_pos_z', 'p2_vel_x', 'p2_vel_y', 'p2_vel_z', 'p2_boost', 'p3_pos_x',\n       'p3_pos_y', 'p3_pos_z', 'p3_vel_x', 'p3_vel_y', 'p3_vel_z', 'p3_boost',\n       'p4_pos_x', 'p4_pos_y', 'p4_pos_z', 'p4_vel_x', 'p4_vel_y', 'p4_vel_z',\n       'p4_boost', 'p5_pos_x', 'p5_pos_y', 'p5_pos_z', 'p5_vel_x', 'p5_vel_y',\n       'p5_vel_z', 'p5_boost', 'boost0_timer', 'boost1_timer', 'boost2_timer',\n       'boost3_timer', 'boost4_timer', 'boost5_timer']","metadata":{"execution":{"iopub.status.busy":"2022-10-19T11:51:59.048356Z","iopub.execute_input":"2022-10-19T11:51:59.048701Z","iopub.status.idle":"2022-10-19T11:51:59.460599Z","shell.execute_reply.started":"2022-10-19T11:51:59.048670Z","shell.execute_reply":"2022-10-19T11:51:59.458606Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test = pd.read_csv('../input/tabular-playground-series-oct-2022/test.csv')\ntest = test[FEATURES]","metadata":{"execution":{"iopub.status.busy":"2022-10-19T11:51:59.463795Z","iopub.execute_input":"2022-10-19T11:51:59.464978Z","iopub.status.idle":"2022-10-19T11:52:17.667950Z","shell.execute_reply.started":"2022-10-19T11:51:59.464915Z","shell.execute_reply":"2022-10-19T11:52:17.666649Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"predictions = []\nfor fold in range(10):\n    model = joblib.load('../input/tps-oct-2022-catb-catb/aa/A/model_'+str(fold)+'.pkl')\n    predictions.append([i[1] for i in model.predict_proba(test)])\npredictions = np.average(np.array(predictions),axis=0)\nsub['team_A_scoring_within_10sec'] = predictions\n\npredictions = []\nfor fold in range(10):\n    model = joblib.load('../input/tps-oct-2022-catb-catb/aa/B/model_'+str(fold)+'.pkl')\n    predictions.append([i[1] for i in model.predict_proba(test)])\npredictions = np.average(np.array(predictions),axis=0)\nsub['team_B_scoring_within_10sec'] = predictions","metadata":{"execution":{"iopub.status.busy":"2022-10-19T11:52:17.669273Z","iopub.execute_input":"2022-10-19T11:52:17.669628Z","iopub.status.idle":"2022-10-19T11:55:52.568780Z","shell.execute_reply.started":"2022-10-19T11:52:17.669597Z","shell.execute_reply":"2022-10-19T11:55:52.567464Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sub.to_csv('submission.csv',index = False)\nsub.head()","metadata":{"execution":{"iopub.status.busy":"2022-10-19T11:55:52.570309Z","iopub.execute_input":"2022-10-19T11:55:52.570796Z","iopub.status.idle":"2022-10-19T11:55:55.206971Z","shell.execute_reply.started":"2022-10-19T11:55:52.570751Z","shell.execute_reply":"2022-10-19T11:55:55.205743Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Did it work?**\n\nThe data was taken from professional Rocket League matches. Each event consists of a chronological series of frames recorded at 10 frames per second. All events begin with a kickoff, and most end in one team scoring a goal, but some are truncated and end with no goal scored due to circumstances which can cause gameplay strategies to shift, for example \n1) nearing end of regulation (where the game continues until the ball touches the ground) or \n\n2) becoming non-competitive, eg one team winning by 3+ goals with little time remaining.\n\n\n**What did you not understand about it?**\n\nWell, everything provides in the competition data page. I've no problem while working on it. The dataset consists of sequences of snapshots of the state of a Rocket League match, including position and velocity of all players and the ball, as well as extra information. The goal of the competition is to predict -- from a given snapshot in the game -- for each team, the probability that will score within the next 10 seconds of game time.\n\n**I hope you find this notebook useful , Good Luck!**","metadata":{}}]}