{"cells":[{"metadata":{},"cell_type":"markdown","source":"# NFL Why do players get injured?\n\n![nfl](https://i2.wp.com/www.audiovisual451.com/wp-content/uploads/2020/03/nfl.jpg?w=696&ssl=1)\n\n## Scope of the analysis\n\nThis exercise attempts to provide further insights on how NFL players get injured based on the data provided for this challenge. Most of the reasoning illustrated on this notebook and the subsequent conclusions are based on statistical hypothesis testing that will be also explained. The conclusions of the exercise will be listed at the end of the notebook. Let's get started!","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"import numpy as np\nimport pandas as pd\nimport scipy.stats as stats\nimport seaborn as sns\nimport matplotlib.pyplot as plt\nfrom statsmodels.stats.multicomp import MultiComparison\nfrom statsmodels.stats.weightstats import ttest_ind\nimport scipy.stats as stats\nimport itertools","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## Data:\n\n* InjuryRecord.csv: Contains information on 105 lower-limb injuries that occurred during regular season games over the two seasons.\n* PlayList.csv: Details for the 267,005 player-plays that make up the dataset. Each play is indexed by PlayerKey, GameID, and PlayKey fields.\n* PlayerTrackData.csv: Player level data that describes the location, orientation, speed, and direction of each player during a play recorded at 10 Hz.","execution_count":null},{"metadata":{"_uuid":"d629ff2d2480ee46fbb7e2d37f6b5fab8052498a","_cell_guid":"79c7e3d0-c299-4dcb-8224-4455121ee9b0","trusted":true},"cell_type":"code","source":"df_injuries = pd.read_csv('../input/nfl-playing-surface-analytics/InjuryRecord.csv')\ndf_playlist = pd.read_csv('../input/nfl-playing-surface-analytics/PlayList.csv')\ndf_player_tracks = pd.read_csv('../input/nfl-playing-surface-analytics/PlayerTrackData.csv')","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## Exploring the datasets\n\nWe first read the three datasets provided and get some insights on the kind of data that we have.","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"print('Number of injuries: %d' % len(df_injuries))\nprint('Number of plays: %d' % len(df_playlist))\nprint('Number of player tracks: %d' % len(df_player_tracks))","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"df_injuries.info(verbose=True)\nprint('\\n----------------------------------------------\\n')\ndf_playlist.info(verbose=True)\nprint('\\n----------------------------------------------\\n')\ndf_player_tracks.info(verbose=True)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## Probability of getting injured\n\nInitially, we explore the hypothesis that the probability of getting injured on the synthetic turf is higher. NO other variables are involved in this test. We compare the number of injuries in both surfaces and assess the significance level of the difference between probabilities.","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"# Are there any duplicated IDs?\ndf_playlist['PlayKey'].duplicated().value_counts()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# Are there any null values that need to be dropped?\ndf_playlist[['PlayerKey', 'GameID', 'PlayKey']].isna().sum()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# Join playlist dataset and injury dataset using the keys from the playlist dataset\ndf_playlist_ext = pd.merge(df_playlist, df_injuries, how='left', on=['PlayerKey', 'GameID', 'PlayKey'])\n#df_playlist_ext.info()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# Create a new Boolean column indicating whether the player got injuried in that play.\ndf_playlist_ext['IsInjured'] = (df_playlist_ext['DM_M1']==1) | (df_playlist_ext['DM_M7']==1) | (df_playlist_ext['DM_M28']==1) | (df_playlist_ext['DM_M42']==1)\ndf_playlist_ext['IsInjured'].value_counts()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# compute the injury probability for synthetic and natural turf\np_injury = df_playlist_ext[['FieldType', 'IsInjured']].groupby('FieldType').mean()['IsInjured']\np_injury","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"The null hypothesis **H0** states that there is no difference between the two population proportions, where the two populations or groups are the plays on synthetic and natural turf. More formally, this is stated as the difference between the two proportions being zero:  \n  \n*H0: p1-p2=0*  \n  \nWe employ the **z-score test** for two population proportions. This test is generally used when you want to know whether two populations differ significantly on some characteristic (injuries). The equation is represented in the image below.\n\n![z_2_proportions.png](attachment:z_2_proportions.png)","attachments":{"z_2_proportions.png":{"image/png":"iVBORw0KGgoAAAANSUhEUgAAALsAAABSCAIAAACDuTJ3AAAAGXRFWHRTb2Z0d2FyZQBBZG9iZSBJbWFnZVJlYWR5ccllPAAABpxJREFUeNrsXTt24jAUlTmzFJgihxU4K4A0VLTpTBmadCnTpYESOtpUNMErGK8ghyKwF4+fjYkNtizJ+ll+t8nMJEOeH1fvJ+nixXFMEAhmDNAFCGQMAhmDQMYgkDEIZAwCYR1jwoXnPa7PHfciPMUiRMbo8PP7wyn+9zLsuBcnmzie7fvBGs/UBO+8fhx9zk/dZ4vTj2QPY8C5y/Eh3kyc8mYvOBObwGnlE391KvyLJUa2NQMejASH2GEYqWPOX5+RP38aMr8Zlqyf5v8/fJr7ZLtXXc5ABXiB9tJpYIgwZPzXycg9/Dsm5PvnrJYu0+8sQCchbTvVTBoTjDkdI+I/jNzM8qMHn0THk8JS6X1LgresUhq+vAVEQ0wzzJjzzzcl1GbDmTzsGmhXZZihLsik8bmw3CazQHVMs2Qec5OUkibD288gykbHr+zPcXwIyPZd63RPghlpWtLrPKUx7Q5/LEn/L/+grgz3EdkmaSveDC8BfgvOGPbMjIaMTsYmLTAVYyoDKeQrf7XLpxli9Q7ECQoYtiRamVGXcyWZB/Q1CwMxJmsnahaQP98Ni77nb6kgTry0XcctzaD8eHvzsuU2KfwGrW2EiRhT106E+23B1+HHMvJXr/qHwu3MUNwIpgOfgvMEl1Ub2DHyvc5L/euggRganrYzQ/3Q986+e0+qHYvH1lAm6Un8IMjTdKUfkh9RTqMGMzIS1VlYvRaUkIbiJgcZU7UU6WzInOT7vmrG0M1Ivnv5ZuXP6SFMD/eVoP7brfzSlCNtUOrzPxSMcbyb6xgvUsyYbPLt9oqm5bx+XpLfHstRGDtRlVDgMF6OrtNUqBgt2GtiNeP89UnK9XC4gPMbjh+OMVX53iQBiOPwtbnIhKivNCuxmVFITvrqK1tg7Aye6BR/dHwzfA4rXMDugWNnwTqQlTqKhLPvD6eULuf1ovNH2juYlXjzxQWGUkChrTXT2mJWQnQOtH0lz/PQQX1GZTTBGIPAyheBjEE40SuVuxdET/jQ7kRVcOjvJAuzEi/OP9/OXiFBYB2DMM+Y0zFy9F5jO4QLWwRx1FiCMUYsJa8fK269pRfipls7yMJgSbgQYJQwY3pcxiRvxzPZ3ZT88B7BhTjzHSS7JZPNjjzzXvgU7pUgKc2GvaTLlBziu4NTIFOVft+4hTyWwLk2UGTj6Hn/EAQXX6bb4BC7NFGYbA6BN10wn/gRzUq9TEqgq2DkCpVizrzenLnGyldWgPlYRrkOh1MAUZFo+REqZYxgb53eOrZPkZJJIxbuSgYzN0fcICrCKEMjyBiRpARsSXqM2MKNBSgWoW2gsQYI424iTgUsmCgjISt5VPwuY6OXM5qNTNqG0/xzVBcAU8KUtftqlpItRReXJRwCfsL71lxHbWnXRuDaqi0HZmuvkTTfp74ff5h6KjFL8ntAim7RcjKm7o3In82eI9Y0SzVaqfxilrjGwEBHGVNZAtgzJBUsAV0Eg6KeGGNuOyUGtcG7zgqqTa01MLsk4r3f7ClPjKN95dusNmiBu9klESmyYW5v1bNKPgrtEqRJ6fU6/rFdbVCVkQK3c1Tc3BC7JCRsiayZL01tUJJgqR5JxFTMktnp8j73oPx0o2VC7Snb06n4BAbZjKkY+KZqg08UtcH2KsWZhEwtGEY9jUbSdeWUCi2Xn+6uV1I/yGKtHQZir327NOlqg9CAaFUprm/ZmiQRIcRUbAXo0HW2AQyl2kAePy9LMAmv01/l/MaeVV9Z3GRkTqrKvSPVHzdgHswioVLGd0Kih7qHpM1GNkymNY7w9E/wmEVCpTCGTaLJsKxgk5FN5mnVodLPGOYVMZBQxjSoDRZrO9ggNnPagWokDPRgW51WX1bu7qr6vBbwFcd4s70ZjButLFnpfukJhhhb9xiF4/YlDqRfVpeYYOBT/6SYwZFzCdseaPG3V8kGEts5I8XIStalr5x721zmbWUGz3Ji/bTM39+f/LWXal4170T5n4yF2zZm8PGLMPM3t6DPjLn3bjn+V71TWkjUaEa9OD5vPCLsBl1elvcwlXMonQArvTk3ztemjE83I6aJ4/PzmXCYlL30IehxiLl9T0oV43Whl1yko1FmMKPCGrGii/A4CV6/30mJ3j7lob9EEA2MYTFDWklB+IisJ8Y6xypLPCaloOCY4MF1SxJFEUF0EbLE8blmvnDdkhBUpuoeZIrjmwhsPeqqCjDnN6ni+KgAjTBzahOBjEEgkDEIZAwCGYNAxiCQMQhkDALBhv8CDACKZIre6nm4OwAAAABJRU5ErkJggg=="}},"execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"p_null = df_playlist_ext[['FieldType', 'IsInjured']][df_playlist_ext['FieldType'] == 'Natural'].mean()[0]\n\nn_plays_natural = df_playlist_ext.groupby('FieldType').size()[0]\nn_plays_synthetic = df_playlist_ext.groupby('FieldType').size()[0]\n\ndiv = np.sqrt(p_null * (1-p_null) * (1/n_plays_natural + 1/n_plays_synthetic))\n\n#  compute z-score and p-value\nz = (p_injury[1] - p_injury[0]) / div\n\nprint('The z-score is: {}'.format(z))\nprint('The p-value is: {}'.format(1-stats.norm.cdf(z)))","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"**Conclusion** With a 99% level of confidence we can **reject H0** and therefore we can confidently state that the probability of being injured on synthetic turf is higher than the probability of being injured on natural turf.  \n  \nNow the next question is to discover if any of the given variables affects the probability of getting injured on one or another surface. Additionally, we are also interested on whether they affect the duration of the injury.","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"## Pearson correlation between numerical variables\n\nWe explore the degree of correlation between the numerical variables provided.","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"df_injuries_ext = pd.merge(df_injuries, df_playlist, how='inner', on=['PlayerKey', 'GameID', 'PlayKey'])\n\n# Create a numerical variable from the duration\ndf_injuries_ext['Duration1'] = np.where(df_injuries_ext['DM_M1']>=1, 1, 0)\ndf_injuries_ext['Duration7'] = np.where(df_injuries_ext['DM_M7']>=1, 7, 0)\ndf_injuries_ext['Duration28'] = np.where(df_injuries_ext['DM_M28']>=1, 28, 0)\ndf_injuries_ext['Duration42'] = np.where(df_injuries_ext['DM_M42']>=1, 42, 0)\ndf_injuries_ext['Duration'] = df_injuries_ext[['Duration1', 'Duration7', 'Duration28', 'Duration42']].max(axis=1)\ndf_injuries_ext.drop(columns=['Duration1', 'Duration7', 'Duration28', 'Duration42'], inplace=True)\ndf_injuries_ext.head()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"df_corr = df_injuries_ext[['PlayerDay', 'PlayerGame', 'Temperature', 'PlayerGamePlay', 'Duration']].corr()\n\nfig = plt.figure(figsize=(10,7))\nax = sns.heatmap(df_corr, annot=True, cmap=sns.diverging_palette(240, 10, n=9))\nax.set_title(\"Pearson Correlation between Variables\");\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"**Conclusion**: Most of the numeric variables are uncorrelated with the exception of the day and game, which indicate the day of the season and game and are necessarily correlated. We also observe how the duration of the injury doesn't seem to be related with any of these variables.","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"## Correlation between categorical variables\n\nIn order to assess the relationship between categorical variables in our dataset, with special foculs on the surface of the field, we employ **Cramér's V**. In statistics, Cramér's V is a measure of association between two nominal variables, giving a value between 0 and +1 (inclusive). It is based on Pearson's chi-squared statistic and was published by Harald Cramér in 1946 https://en.wikipedia.org/wiki/Cram%C3%A9r%27s_V.","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"# The following function was originally seen here: https://www.kaggle.com/phaethonprime/eda-and-logistic-regression\n\ndef cramers_corrected_stat(confusion_matrix):\n    \"\"\" calculate Cramers V statistic for categorical-categorical association.\n        uses correction from Bergsma and Wicher, \n        Journal of the Korean Statistical Society 42 (2013): 323-328\n    \"\"\"\n    chi2 = stats.chi2_contingency(confusion_matrix)[0]\n    n = confusion_matrix.sum().sum()\n    phi2 = chi2/n\n    r,k = confusion_matrix.shape\n    phi2corr = max(0, phi2 - ((k-1)*(r-1))/(n-1))    \n    rcorr = r - ((r-1)**2)/(n-1)\n    kcorr = k - ((k-1)**2)/(n-1)\n    return np.sqrt(phi2corr / min( (kcorr-1), (rcorr-1)))","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"cols = ['Surface', 'BodyPart', 'RosterPosition', 'StadiumType', 'FieldType', 'Weather', 'PlayType', 'Position', 'PositionGroup']\ncorr = np.zeros((len(cols),len(cols)))\n\n# Apply previous function to calculate Cramer's V for each pair\nfor col1, col2 in itertools.combinations(cols, 2):\n    idx1, idx2 = cols.index(col1), cols.index(col2)\n    corr[idx1, idx2] = cramers_corrected_stat(pd.crosstab(df_injuries_ext[col1], df_injuries_ext[col2]))\n    corr[idx2, idx1] = corr[idx1, idx2]","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"df_corr = pd.DataFrame(corr, index=cols, columns=cols)\n\n# Plot the correlation matrix\nfig = plt.figure(figsize=(10,7))\nax = sns.heatmap(df_corr, annot=True, cmap=sns.diverging_palette(240, 10, n=9))\nax.set_title(\"Cramer V Correlation between Variables\");\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"**Conclusion**: The most interesting information in the above plot is the relationship between the surface and the weather. The rest of the meaningful correlations are either obvious or not interesting for this analysis.","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"## Statistical hypothesis testing\n\nIn this section, we statistically test the relationship between some of the variables f the dataset. We specifically try to identify conditions in which it is more likely to suffer an injury, or that the injury las a longer duration, based on the number of instances observed. \n\n### About statistical tests with categorical variables\n\nThe Chi-square test of independence measures whether there is a relationship between two categorical variables like this case (surface and duration). However, it comes with a series of assumptions. One of them is that the value of expected cells should be greater than 5 for at least 20% of the cells. This won't be the case for some of the test to be run in this analysis.\n\nThe alternative test when this condition is not fulfilled is the Fisher Exact Test, which is practically applied when the sample sizes are relatively small https://towardsdatascience.com/fishers-exact-test-from-scratch-with-python-2b907f29e593. Still, this test can only be conducted for tables of size 2x2. \n\nSince in our case we have larger tables as we will see below, we stick to the Chi-square test for this analysis. More information regarding how to perform this test in Python here https://pythonfordatascience.org/chi-square-test-of-independence-python/.","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"# Plot the distribution of injuries per surface\nsns.set(style=\"darkgrid\")\nax = sns.countplot(x=\"Surface\", data=df_injuries)\nax.set_title(\"Injury count on both surfaces\");\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"markdown","source":"### Relationship between the surface and duration of the injury\n\nDoes the surface of the turf affect the duration of the injury? Here, the **H0 (Null Hypothesis)** states that there is no relationship between variable one (surface) and variable two (duration), while the **H1 (Alternative Hypothesis)** says that there is a relationship between them.\n\nWe **first** consider the duration as a categorical variable. Thus, when testing the data, the cells should represent counts of cases.","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"sns.set(style=\"darkgrid\")\nax = sns.countplot(x=\"Duration\", data=df_injuries_ext)\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"crosstab = pd.crosstab(df_injuries_ext['Surface'], df_injuries_ext['Duration'])\ncrosstab","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"ret = stats.chi2_contingency(crosstab)\nprint('Pearson Chi-square = %.3f' % ret[0])\nprint('p-value = %.3f > 0.05' % ret[1])\nprint('We cannot reject H0. There is no relationship between Surface and Duration')","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"**Second**, we treat the duration as a numerical variable by making the assumption that the player was injuried the exact number of days indicated in the database. Then, we test if the duration is signifficantly higher on synthetic turf. The plot indicates that indeed it is higher, but we need to assess whether this difference is statistically significant.","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"# Show each observation with a scatterplot\nsns.set(style=\"darkgrid\")\nsns.stripplot(x=\"Duration\", y=\"Surface\", order=['Natural','Synthetic'], data=df_injuries_ext, alpha=.50)\nsns.pointplot(x=\"Duration\", y=\"Surface\", order=['Natural','Synthetic'], data=df_injuries_ext, palette=\"dark\", markers=\"d\")\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# Get each position independently\nd_natural = df_injuries_ext['Duration'][df_injuries_ext['Surface']=='Natural']\nd_synthetic = df_injuries_ext['Duration'][df_injuries_ext['Surface']=='Synthetic']\n\n# Example of how it would be done for only two groups\nret = stats.ttest_ind(d_natural, d_synthetic)\nprint('Independent t-test = %.3f' % ret[0])\nprint('p-value = %.3f > 0.05' % ret[1])\nprint('We cannot reject H0. There is no relationship between Surface and Duration')","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"**Conclusion**: Although there is an interesting effect indicating that the duration may be longer if the injury was suffered on synthetic turf, there are no statistical evidence tu support that affirmation. However, given the little data available, it is hard to say for certain. We need to collect more data!","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"### Relationship between the weather and the surface\n\nFollowing on the correlation observed between weather and surface. Could it be that bad weather in combination with a specific turf surface lead to a higher probability of getting injured? This is the hypothesis that we are following in this section. The **Null Hypothesis (H0)** states that there is no relationship between surface and weather.","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"df_injuries_ext = df_injuries_ext.replace('Mostly sunny', 'Sunny')\ndf_injuries_ext = df_injuries_ext.replace('Mostly Sunny', 'Sunny')\ndf_injuries_ext = df_injuries_ext.replace('Clear Skies', 'Clear')\ndf_injuries_ext = df_injuries_ext.replace('Controlled Climate', 'Clear')\ndf_injuries_ext = df_injuries_ext.replace('Clear and warm', 'Clear')\ndf_injuries_ext = df_injuries_ext.replace('Fair', 'Clear')\ndf_injuries_ext = df_injuries_ext.replace('Clear skies', 'Clear')\ndf_injuries_ext = df_injuries_ext.replace('Cloudy with periods of rain, thunder possible. Winds shifting to WNW, 10-20 mph.', 'Rain')\ndf_injuries_ext = df_injuries_ext.replace('Coudy', 'Cloudy')\ndf_injuries_ext = df_injuries_ext.replace('Mostly cloudy', 'Cloudy')\ndf_injuries_ext = df_injuries_ext.replace('Cloudy and Cool', 'Cloudy')\ndf_injuries_ext = df_injuries_ext.replace('Sun & clouds', 'Partly Cloudy')\ndf_injuries_ext = df_injuries_ext.replace('Light Rain', 'Rain')\ndf_injuries_ext = df_injuries_ext.replace('Rain shower', 'Rain')\ndf_injuries_ext = df_injuries_ext.replace('Cloudy, 50% change of rain', 'Rain')\ndf_injuries_ext = df_injuries_ext.replace('Indoors', 'Indoor')","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"sns.set(style=\"darkgrid\")\nax = sns.countplot(x=\"Weather\", data=df_injuries_ext)\nax.set_xticklabels(ax.get_xticklabels(), rotation=30, ha='right')\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"crosstab = pd.crosstab(df_injuries_ext['Surface'], df_injuries_ext['Weather'])\ncrosstab","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"ret = stats.chi2_contingency(crosstab)\nprint('Pearson Chi-square = %.3f' % ret[0])\nprint('p-value = %.3f < 0.05' % ret[1])\nprint('We can reject H0. There is a relationship between Surface and Weather')","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"In fact we observe an effect on the weather on the surface of the injuries. However, we do not know for which weather in concrete. I have the suspicion though, that this effect is caused by the number of **Indoor** injuries (8 to 0).","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"df_aux = df_injuries_ext[df_injuries_ext['Weather'] != 'Indoor']\ncrosstab = pd.crosstab(df_aux['Surface'], df_aux['Weather'])\nret = stats.chi2_contingency(crosstab)\nprint('Pearson Chi-square = %.3f' % ret[0])\nprint('p-value = %.3f < 0.05' % ret[1])\nprint('We cannot reject H0.')","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"**Conclusion** Bingo! the effect disappears when removing indoor weather from the equation. We don't know beforehand whether all indoor stadiums have a synthetic surface or this is not the case. Therefore, We cannot jump into the conclusion that playing indoors increases the chance of suffering an injury. This is something that has to be studied more in detail. ","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"### Relationship betweein the surface and the body part\n\nDoes the surface of the turf affect the part where the player was injured? Here, the **H0 (Null Hypothesis)** states that there is no relationship between variable one (surface) and variable two (body part), while the **H1 (Alternative Hypothesis)** says that there is a relationship between them.","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"sns.set(style=\"darkgrid\")\nax = sns.countplot(x=\"BodyPart\", data=df_injuries_ext)\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"crosstab = pd.crosstab(df_injuries_ext['Surface'], df_injuries_ext['BodyPart'])\ncrosstab","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"ret = stats.chi2_contingency(crosstab)\nprint('Pearson Chi-square = %.3f' % ret[0])\nprint('p-value = %.3f > 0.05' % ret[1])\nprint('We cannot reject H0. There is no relationship between Surface and BodyPart')","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"crosstab_heel_toe = crosstab[['Ankle', 'Foot']]\nret = stats.fisher_exact(crosstab_heel_toe)\nprint('Pearson Chi-square = %.3f' % ret[0])\nprint('p-value = %.3f > 0.05' % ret[1])\nprint('We cannot reject H0. There is no relationship between Surface and BodyPart for the specific case of heels and toes')","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"**Conclusion**: There is no statistical evidence supporting that the surface has an effect on the number of injuries suffered for any specific part of the body.","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"### Considering the relationship between player position and surface\n\nAlthough we have already observed that these two variables are not correlated, we still want to investigate if players playing in a specific position get injured more often on synthetic turf (our **Null Hypothesis H0**).","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"sns.set(style=\"darkgrid\")\nax = sns.countplot(x=\"PositionGroup\", data=df_injuries_ext)\nax.set_xticklabels(ax.get_xticklabels(), rotation=45, ha='right')\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"crosstab = pd.crosstab(df_injuries_ext['Surface'], df_injuries_ext['PositionGroup'])\ncrosstab","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"ret = stats.chi2_contingency(crosstab)\nprint('Pearson Chi-square = %.3f' % ret[0])\nprint('p-value = %.3f > 0.05' % ret[1])\nprint('We cannot reject H0. There is no relationship between Surface and BodyPart')","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"**Conclusion**: Indeed, there is no evidence regarding the player position having an effect on the injuries per surface.","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"### Interaction between surface and temperature\n\nThe degree of relationship for these two variables is unexplored in the correlation analysis as one is numerical (temperature) and the other one is categorical (surface). Let's start by cleaning the temperature data.","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"# We have some values -999 which are clear NaNs. Replace them\ndf_aux = df_injuries_ext[df_injuries_ext['Temperature'] != -999]","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"sns.set(style=\"darkgrid\")\nsns.stripplot(x=\"Temperature\", y=\"Surface\", data=df_aux, alpha=.70)\nsns.pointplot(x=\"Temperature\", y=\"Surface\", data=df_aux, palette=\"dark\", markers=\"d\")\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"At first sight, we don't observe a clear effect of the temperature on the surface of the injuries. Still, we have to test it statistically.  \n  \nAs the temperature is a numerical variable, we cannot perform the same type of analysis as before. In this case, an **two-sided Student's t-test** for independent samples should be carried out. This statistical test assess whether two independent samples have identical average (expected) values (our **null hypothesis or H0**). https://en.wikipedia.org/wiki/Student%27s_t-test","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"temp_natural = df_aux['Temperature'][df_aux['Surface']=='Natural']\ntemp_synth = df_aux['Temperature'][df_aux['Surface']=='Synthetic']\n\nret = stats.ttest_ind(temp_natural, temp_synth)\nprint('Independent t-test = %.3f' % ret[0])\nprint('p-value = %.3f > 0.05' % ret[1])\nprint('We cannot reject H0. There is no relationship between Surface and Temperature')","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"**Conclusion**: The temperature has no effect on the injuries per surface.","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"## Path analysis for injured players\n\nWe are also provided with the path followed by the player during the entire course of the play where he got injured. In this regard, we will be analyzing the effect of two specific variables on the amount of injuries suffered for any specific turf. One is the **average speed** of the play and the second is the **amount of rotation**.\n\nIn order to get to the bottom of this interactions, we would have to consider with special interest the last part of the player movement. This analysis is outside the scope of this notebook, however, the same methodology would apply.","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"df_tracks_ext = pd.merge(df_player_tracks, df_injuries, on='PlayKey', how='inner')\ndf_tracks_ext.head()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# Scatter plot for one example play (using 39678-2-1, 47813-8-19 and 31070-3-7 as examples)\ndf_track = df_tracks_ext[df_tracks_ext['PlayKey']=='47813-8-19']\nturf_img = plt.imread('../input/customimgs/nfl-turf.jpg')\n\nfig = plt.figure(figsize=(10,6))\nimplot = plt.imshow(turf_img, extent=[-10, 120, -10, 54])\nsns.scatterplot(x='x', y='y', hue='PlayKey', data=df_track)\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# And scatter plot for every play that resulted into an injury\nfig = plt.figure(figsize=(10,6))\nimplot = plt.imshow(turf_img, extent=[-10, 120, -10, 54])\nsns.scatterplot(x='x', y='y', hue='PlayKey', data=df_tracks_ext, legend=False)\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"### Analyzing the average speed\n\nWe look into the average speed of the play and its influence on the injuries suffered for each specific turf. We observe the degree of relationship between the speed and the duration of the injury as well.","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"# Obtain average speed for each play\ndf_speed = df_tracks_ext.groupby('PlayKey')['s'].mean()\ndf_speed_ext = pd.merge(df_speed, df_tracks_ext[['PlayKey', 'Surface']], on='PlayKey', how='inner')\ndf_speed_ext = df_speed_ext.drop_duplicates().reset_index()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# Sort the dataframe by surface\ndf_natural = df_speed_ext.loc[df_speed_ext['Surface'] == 'Natural']\ndf_synthetic = df_speed_ext.loc[df_speed_ext['Surface'] == 'Synthetic']\n\n# Plot distribution of velocities for both groups\nsns.distplot(df_natural['s'], bins=10)\nsns.distplot(df_synthetic['s'], bins=10)\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"sns.set(style=\"darkgrid\")\nsns.stripplot(x=\"s\", y=\"Surface\", data=df_speed_ext, alpha=.70)\nsns.pointplot(x=\"s\", y=\"Surface\", data=df_speed_ext, palette=\"dark\", markers=\"d\")\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"**Conclusion:** In light of the figures above, we do not need to run the statistican tests to observe that the speed for both groups is nearly identical. Our data indicates that **average** speed does not affect the probability of getting injured on one specific surface.  \n  \nWe are also interested on the relationship between the speed and the length of the injury. We observe this through the Pearson correlation between variables.","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"df_aux = pd.concat([df_speed_ext, df_injuries_ext['Duration']], axis=1)\ncorr = df_aux['s'].corr(df_aux['Duration'])\nprint('Pearson correlation: ', corr)\nprint('Degrees of freedom: ', (len(df_speed_ext.index)-2) )","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"**Conclusion:** The correlation is not particularly high, although it has some degree of signifficancy according to this table: https://www.statisticssolutions.com/table-of-critical-values-pearson-correlation/. Therefore there is a mild tendency that higher speed may lead to injuries that last longer.","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"### Analyzing the amount of rotation\n\nIn this exercise, we calculate the amount of rotation as the sum of the differentials between the rotations at each specific time step, for each particular play. In this spirit, we initially consider both the **'o'** and **'dir'** column. We don't know how these orientation values were registered, although it is assumed a certain degree of noise that is neglected for this exercise but should be taken into account for deeper analyses.","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"area = 4 * df_track['s']**2\n\nfig = plt.figure(figsize=(8, 20))\n\n# Plotting direction and orientation in polar coordinates, for one specific play\nax = plt.subplot(1, 2, 1, projection='polar')\nplt.scatter(df_track['dir'], df_track['s'], s=area, alpha=0.50, color='orange')\nax.set_title('Speed and direction.')\nax = plt.subplot(1, 2, 2, projection='polar')\nplt.scatter(df_track['o'], df_track['s'], s=area, alpha=0.50, color='green')\nax.set_title('Speed and orientation.')\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# To get the amount of rotation from orientation we substract the current row with the previous row\ndf_tracks_ext['o_diff'] = df_tracks_ext['o'].diff()\ndf_orientation = df_tracks_ext.groupby('PlayKey')['o_diff'].sum().abs()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# To get the amount of rotation from direction we substract the current row with the previous row\ndf_tracks_ext['dir_diff'] = df_tracks_ext['dir'].diff()\ndf_direction = df_tracks_ext.groupby('PlayKey')['dir_diff'].sum().abs()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# Merge the dataframes\ndf_orientation_ext = pd.merge(df_orientation, df_tracks_ext[['PlayKey', 'Surface']], on='PlayKey', how='inner')\ndf_orientation_ext = df_orientation_ext.drop_duplicates().reset_index()\ndf_direction_ext = pd.merge(df_direction, df_tracks_ext[['PlayKey', 'Surface']], on='PlayKey', how='inner')\ndf_direction_ext = df_direction_ext.drop_duplicates().reset_index()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Somehow surprisingly, these two measures are not highly correlated, as we can see below. The challenge advises not to consider the orientation as a reliable indicator as the measurement varied over seasons. We therefore **won't consider** the orientation for further analysis and analyze the amount of rotation based on the direction.","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"df_direction_ext['dir_diff'].corr(df_orientation_ext['o_diff'])","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# Plot conditional distribution\nsns.set(style=\"darkgrid\")\nax = sns.stripplot(x=\"dir_diff\", y=\"Surface\", data=df_direction_ext, alpha=.70)\nsns.pointplot(x=\"dir_diff\", y=\"Surface\", data=df_direction_ext, palette=\"dark\", markers=\"d\")\nax.set_title('Conditional distribution, direction')\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"temp_natural = df_direction_ext['dir_diff'][df_direction_ext['Surface']=='Natural']\ntemp_synth = df_direction_ext['dir_diff'][df_direction_ext['Surface']=='Synthetic']\n\n# Student's t-test\nret = stats.ttest_ind(temp_natural, temp_synth)\nprint('Independent t-test = %.3f' % ret[0])\nprint('p-value = %.3f > 0.05' % ret[1])\nprint('We cannot reject H0. There is no relationship between Surface and Temperature')","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"df_aux = pd.concat([df_direction_ext, df_injuries_ext['Duration']], axis=1)\ncorr = df_aux['dir_diff'].corr(df_aux['Duration'])\nprint('Pearson correlation: ', corr)\nprint('Degrees of freedom: ', (len(df_speed_ext.index)-2) )","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"**Conclusion:** Both the statistical test and the correlation suggest that the amount of rotation, as considered in this analysis, does not have an effect on the surface nor is related with the duration of the injury","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"## Final remarks\n\nAfter an exhaustive analysis on the data provided we found the following conclusions, supported by statistical evidence:\n* There is a higher probability of getting injured on synthetic turf.\n* The average speed of the play is correlated with the duration of the injury.\n\nWe also found insights on the following effects that should be investigated more carefully:\n* Getting injuried on synthetic turf may lead to being injured for a larger period of time.\n* The weather may affect the probability of getting injured when playing indoors.\n\nThanks for following this notebook till the end :)","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"","execution_count":null,"outputs":[]}],"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat":4,"nbformat_minor":4}