{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"code","source":"# This Python 3 environment comes with many helpful analytics libraries installed\n# It is defined by the kaggle/python Docker image: https://github.com/kaggle/docker-python\n# For example, here's several helpful packages to load\n\nimport numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\nimport seaborn as sns\nimport matplotlib.pyplot as plt\nimport tensorflow as tf\nfrom sklearn.utils import shuffle\nfrom sklearn.model_selection import train_test_split\nimport feather\n\n# Input data files are available in the read-only \"../input/\" directory\n# For example, running this (by clicking run or pressing Shift+Enter) will list all files under the input directory\n\nimport os\nfor dirname, _, filenames in os.walk('/kaggle/input'):\n    for filename in filenames:\n        print(os.path.join(dirname, filename))\n\n# You can write up to 20GB to the current directory (/kaggle/working/) that gets preserved as output when you create a version using \"Save & Run All\" \n# You can also write temporary files to /kaggle/temp/, but they won't be saved outside of the current session","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Import Data\n\nEach training file contains ~2.1 millions of data entries, with values for 61 features and reading the files using pd.read_csv takes ~ 20-30 secs / file. \nThe total number of data entries is therefore ~ 21 millions. \nI have chosen to use a random sample of observations from each training file, because loading all of the files at once crashes the notebook. \nAlso instead of the original csv files I am using the feather files that were created in [this](https://www.kaggle.com/code/gazu468/feather-to-compress-your-data-8x-faster) notebook by Gaju Ahmed. ","metadata":{}},{"cell_type":"code","source":"train = pd.DataFrame()\nfor i in range(10):\n    temp = pd.read_feather('/kaggle/input/tpsoct22-feather-files/train_' + str(i) + '.feather')\n    temp = temp.sample(frac = 0.2, random_state=42).reset_index(drop=True)\n    train = train.append(temp, ignore_index = True)\n    print(f'Reading file train_{i}.feather complete.')\n    \ndel temp\ntest =  pd.read_feather('/kaggle/input/tpsoct22-feather-files/test.feather')","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Explore Data","metadata":{}},{"cell_type":"code","source":"print(f'Shape of training data: {train.shape}')\nprint(f'Shape of testing data: {test.shape}')","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"* Training set contains ~ 4.24 M observations and 61 features\n* Testing set contains ~ 700 k observations and 55 features\nLet's take a look at the features that are present in the two datasets.","metadata":{}},{"cell_type":"code","source":"print(train.columns)\nprint(test.columns)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Features **player_scoring_next, team_scoring_next, team_A_scoring_within_10sec, team_B_scoring_within_10sec** are abscent from the testing set for obvious reasons (the purpose of the challenge is to build a model that predicts how likely is one of the teams to score in the next 10 seconds) \n\nFeatures **game_num, event_time, event_id** are also abscent from the testing set and therefore should be dropped from the training set. \n\nThe **id** feature in the testing set is a simple index and therefore does not offer any information.\n\nFeatures **team_A_scoring_within_10sec** and **team_B_scoring_withing_10sec** are the target values. I will later combine them in a single **label** variable.\n\nLet's inspect the distributions of some of th variables to see how much useful information they carry in relation to the event of a team scoring. For the purspose of taking a look at the distributions I have created the function **plots_distributions** that recieves the data, a string indicating the instance (p0,...,p5, ball, boost1,..., boost5) and a list of string indicating the instance's attribute that we want to observe (positions, velocities, boost or timer) and plots the distribution for 2 cases: (a) one of the teams is close to scoring and (b) no team is close to scoring.","metadata":{}},{"cell_type":"code","source":"positions = ['pos_x', 'pos_y', 'pos_z']\nvelocities = ['vel_x', 'vel_y', 'vel_z']\nboost = ['boost']\ntimer = ['timer']\n\ndef plot_distributions(data, entity, attribute):\n    \n    # data: training data\n    # entity: string that has one of the following values -> 'p0' ... 'p5', 'ball', 'boost0' ... 'boost5'\n    # attribute: list of strings -> ['pos_x', 'pos_y', 'pos_z'], ['vel_x', 'vel_y', 'vel_z'], ['boost'], ['timer']\n    \n    teams_close_to_score = data[data.team_A_scoring_within_10sec.eq(1) | data.team_B_scoring_within_10sec.eq(1)]\n    teams_not_close_to_score = data[data.team_A_scoring_within_10sec.eq(0) & data.team_B_scoring_within_10sec.eq(0)]\n\n    figure, axis = plt.subplots(len(attribute), 2, figsize = (16, 6*(len(attribute))))\n    \n    for i, attr in enumerate(attribute):\n        \n        if (len(attribute) == 1):\n            # in case the attribute list contains one item (['boost'] or ['timer']) \n            axis = np.expand_dims(axis, axis=0)\n        \n        sns.histplot(ax = axis[i, 0], \n                     data = teams_close_to_score, \n                     x = entity + '_' + attr, \n                     hue = 'team_scoring_next',  \n                     bins = 20, \n                     palette = 'pastel')\n        axis[i, 0].set_title('Team close to scoring')\n        sns.histplot(ax = axis[i, 1], \n                     data = teams_not_close_to_score, \n                     x = entity + '_' + attr, \n                     hue = 'team_scoring_next', \n                     bins = 20, \n                     palette = 'pastel')\n        axis[i, 1].set_title('Team not close to scoring')\n    ","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plot_distributions(train, 'ball', positions)\n","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"* From the above histograms, one can see that the feature ball_pos_z has a similar distributies in the two scenarios while ball_pos_y and have different distributions in the two scenarios. For this reason, I have chosen to exclude ball_pos_z from the training features. ","metadata":{}},{"cell_type":"code","source":"plot_distributions(train, 'ball', velocities)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"* Regarding the balls velocities across the three axes, from the above displays of the distributions, it occures that only ball_vel_y differentiates in the two scenarios. That is why I have chosen to exclude ball_vel_x and ball_vel_z from the training features.","metadata":{}},{"cell_type":"code","source":"plot_distributions(train, 'p1', positions)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Positions of different players share similar distributions across the three axis, so for simplicity I have only ploted the distributions for one player. The distributions are especially similar among players of the same team and in a way \"symmetric\" around 0 to the distributions of players from th opposite team. This makes sense since the two goals are on the y axis. Therefore, distributions are different when a team is attacking ther opponent's goal.\n* It seems that the players' positions across the x and z axis have similar distributions in the two scenarios. This means that these features most likely will not offer additional information - when trying to make a prediction - and could be left out. \n* I will only keep the pos_y features for each player. \n","metadata":{}},{"cell_type":"code","source":"plot_distributions(train, 'p1', velocities)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Same as the positions, the velocities of different players share similar distributions across the three axis, so for ease of reading I showcase the distributions of vel_x, vel_y and vel_z for player 1.\n\nFrom these histograms the following observations occur:\n* The distributions of a player's velocity across the x and z axes seem to be similar in both scenarios 1 and 2. This leads us to the cocnlusion that these features most likely will not offer additional information - when trying to make a prediction -  and could be left out.\n* Across the y axis the distributions seem somewhat different in the 2 scenarios, but still not significantly distinguishable.\n\nFeatures of player's velocities across x and z will not be used as input features.\nI will test the model both with and without the velocities across y and check the result.\n(score with the velocities 0.20601, score without: 0.20779)","metadata":{}},{"cell_type":"code","source":"#plot_distributions(train, 'p0', boost)\n#plot_distributions(train, 'p4', boost)\n#plot_distributions(train, 'boost0', timer)\n#plot_distributions(train, 'boost4', timer)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"For the sake of simplicity of reading, I have skipped plotting in this notebook the distributions of players' boosts and boost timers. I turns out that they show similarities in the two scenarios for all players (even ones belonging to opposing teams).","metadata":{}},{"cell_type":"markdown","source":"## Exploring NaN values\n\n* There are a lot of missing values (~ 500k) in the column **team_scoring_next**, but since this column will not be used in the training, the large amount of missing values does not concern us.\n\n* There are also missing values in columns that contain information about the players' position and their velocity.The missing values amount to ~ 5% of the total observations. The rows could be dropped since we have a significant amount of data. However, because the dataset is inbalanced (there are a lot more entries for the teams not being close to score) we should make sure that after droping these rows the dataset does not become even more inbalanced.\n\n* As suggested in many discussion topics, the missing values are probably a result of a player being demolished and therefore the car is out of the court's bounds for a few seconds. In [this](https://www.kaggle.com/code/jcaliz/tps-oct22-animation-eda-keras-baseline) notebook, Jose Cáliz is correctly reasoning that with one player out, the opposing team has an advantage and might be close to score. Therefore, it would be better to keep these observations and find a way to fill in the missing values. In the aformentioned notebook, the author replaces the missing values with zeros. Another interesting idea, in regards to how to fill in the missing data, is presented in [this](https://www.kaggle.com/competitions/tabular-playground-series-oct-2022/discussion/357392) discussion topic by Tarun Yavak. \n\nI have chosen to fill in the missing values with the mean of the corresponding columns. ","metadata":{}},{"cell_type":"code","source":"train_copy = train.drop('team_scoring_next', axis=1)\nteams_close_to_score = train_copy[train_copy.team_A_scoring_within_10sec.eq(1) | train_copy.team_B_scoring_within_10sec.eq(1)]\n\nprint(f'Total number of observations close to score: {teams_close_to_score.shape[0]}')\nprint(f'Total number of observations close to score with demolished players: {teams_close_to_score.isnull().any(axis=1).sum()}')\n","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# train.info(show_counts = True)\n\ntrain_copy = train\ntrain_copy = train_copy.drop('team_scoring_next', axis = 1)\ntrain_copy = train_copy.dropna()\n\nteams_close_to_score = train_copy[train_copy.team_A_scoring_within_10sec.eq(1) | train_copy.team_B_scoring_within_10sec.eq(1)]\nteams_not_close_to_score = train_copy[train_copy.team_A_scoring_within_10sec.eq(0) & train_copy.team_B_scoring_within_10sec.eq(0)]\nprint('Ratio of entities close to score / not close to score if lines containing NaN values are droped:')\nprint(teams_close_to_score.shape[0]/teams_not_close_to_score.shape[0])\n\n\ntrain_copy = train.drop('team_scoring_next', axis = 1)\n\nteams_close_to_score = train_copy[train_copy.team_A_scoring_within_10sec.eq(1) | train_copy.team_B_scoring_within_10sec.eq(1)]\nteams_not_close_to_score = train_copy[train_copy.team_A_scoring_within_10sec.eq(0) & train_copy.team_B_scoring_within_10sec.eq(0)]\nprint('Ratio of entities close to score / not close to score if lines containing NaN values are not droped:')\nprint(teams_close_to_score.shape[0]/teams_not_close_to_score.shape[0])\n\ndel train_copy\ndel teams_close_to_score\ndel teams_not_close_to_score","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The ratio of observations close to score / not close to score does not change significantly, even if we drop the observations that contain NaN values.","metadata":{}},{"cell_type":"markdown","source":"# Train - Validation Split\n\nBecause the data is imbalanced I want to make sure that the training and validation data will have a similar distribution of cases.\nFor the split I will use the train_test_split function from the sklearn.model_selection module. ","metadata":{}},{"cell_type":"code","source":"#create a label variable that is of value:\n#    0 if no team is scoring within 10 seconds, \n#    1 if team A is scoring within 10 seconds and \n#    2 if team B is scoring within 10 seconds\n\ntrain['label'] = pd.DataFrame(train.team_A_scoring_within_10sec + train.team_B_scoring_within_10sec.replace(1,2))\n\ntrain = train.drop(['team_scoring_next', 'team_A_scoring_within_10sec', 'team_B_scoring_within_10sec'], axis=1)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"X_train, X_val, y_train, y_val = train_test_split(train.iloc[:,:-1], train.iloc[:,-1], train_size = 0.9, random_state = 42)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"del train","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"y_train = pd.DataFrame(y_train)\ny_val = pd.DataFrame(y_val)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"y_train.value_counts()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Ratio of score / no score in training set is 0.12633","metadata":{}},{"cell_type":"code","source":"y_val.value_counts()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Ratio of score / no score in training set is 0.1270. \n\nThe 3 cases are distributed in a proportionate way in the training and validation sets.","metadata":{}},{"cell_type":"markdown","source":"## Preprocessing","metadata":{}},{"cell_type":"markdown","source":"Considering the exploration that happened before, I will keep the following features:\n* ball_pos_x/y\n* ball_vel_y\n* pi_pos_y\n* pi_vel_y\n\nMissing values of every column are filled in with the mean value of the corresponding column. ","metadata":{}},{"cell_type":"code","source":"X_train","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"input_features = ['ball_pos_x', 'ball_pos_y', 'ball_vel_y',\n              'p0_pos_y', 'p1_pos_y', 'p2_pos_y', 'p3_pos_y', 'p4_pos_y', 'p5_pos_y',\n              'p0_vel_y', 'p1_vel_y', 'p2_vel_y', 'p3_vel_y', 'p4_vel_y', 'p5_vel_y']\n#target_features = ['team_A_scoring_within_10sec', 'team_B_scoring_within_10sec']\n\nX_train = X_train[input_features].fillna(X_train.mean())\n\nX_val = X_val[input_features].fillna(X_val.mean())\n\nX_test = test[input_features].fillna(X_val.mean())\ndel test\n","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(X_train.shape)\nprint(X_val.shape)\nprint(X_test.shape)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Standardize data\nX_train.iloc[:,:] = X_train.iloc[:,:].apply(lambda x: (x-x.mean())/ x.std(), axis=0)\nX_val.iloc[:,:] = X_val.iloc[:,:].apply(lambda x: (x-x.mean())/ x.std(), axis=0)\nX_test.iloc[:,:] = X_test.iloc[:,:].apply(lambda x: (x-x.mean())/ x.std(), axis=0)\n","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Model - Simple DNN","metadata":{}},{"cell_type":"code","source":"\nmodel = tf.keras.models.Sequential([\n    tf.keras.layers.Dense(64, activation = 'relu'),\n    tf.keras.layers.BatchNormalization(),\n    tf.keras.layers.Dropout(0.3),\n    tf.keras.layers.Dense(32, activation = 'relu'),\n    tf.keras.layers.BatchNormalization(),\n    tf.keras.layers.Dropout(0.3),\n    tf.keras.layers.Dense(16, activation = 'relu'),\n    #tf.keras.layers.Dense(64, activation = 'relu'),\n    #tf.keras.layers.BatchNormalization(),\n    #tf.keras.layers.Dropout(0.3),\n    tf.keras.layers.Dense(3, activation = 'softmax')\n])\n\nmodel.compile(loss = 'sparse_categorical_crossentropy', \n              optimizer = tf.keras.optimizers.Adam(), \n              metrics = ['accuracy']\n)\n\nhistory = model.fit(X_train, \n          y_train, \n          epochs = 20, \n          verbose = 1, \n          batch_size = 256, \n          validation_data = (X_val, y_val)\n          #class_weight = WEIGHTS\n)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"predictions = model.predict(X_test)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#predictions","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"predictions_df = pd.DataFrame(predictions[:,1:])\nindex = pd.DataFrame(predictions_df.index)\nsubmission =pd.concat([index, predictions_df], axis = 1)\nsubmission.columns = ['id', 'team_A_scoring_within_10sec', 'team_B_scoring_within_10sec']\nsubmission.to_csv('submission.csv', index = False)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Future Exploration\n\n* Define new variables that calculate the distances of the ball and the players from the two goals. Check if they are heavily correlated with the pos_y features.\n* Use less of the data and train networks with more nodes (at this moment, when trying to build more complex networks the system runs out of memory and restarts)\n* Maybe try undersampling the observations that are not close to score, so that the ratio of observations score/no score is larger than 0.12 ","metadata":{}},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]}]}