{"cells":[{"metadata":{},"cell_type":"markdown","source":"# NFL 1st and Future - Playing Surface Analytics: The Zoo\n\nContributors:\n* Philipp Singer: https://www.kaggle.com/philippsinger\n* Dmitry Gordeev: https://www.kaggle.com/dott1718"},{"metadata":{},"cell_type":"markdown","source":"![](https://images.unsplash.com/photo-1529663297269-6d349ec39b57?ixlib=rb-1.2.1&ixid=eyJhcHBfaWQiOjEyMDd9&auto=format&fit=crop&w=1050&q=80)\n<font size=\"2\">Source: unsplash.com</font>"},{"metadata":{},"cell_type":"markdown","source":"## Introduction & Summary\n\nArtificial turf plays a crucial role across modern sports like soccer, baseball, tennis, or also American football. While in 2019 already 12 stadiums of the NFL had an artificial surface, this number will rise to **15 in the 2020 season** [1]. Consequently, it is important to understand what impact artificial turf has on aspects of the sport. One of the most important questions to answer is **whether it has an impact on injury rates**. While previous research has been sometimes inconclusive mostly due to lack of historic data, recent studies [2,3] have shown clear evidence of significantly higher injury rates on artificial turf compared to natural grass. \n\nBiomechanical studies have shown that synthetic turf surfaces do not release cleats as readily as natural turf [4] meaning that it is a **harder surface than grass and does not have much \"give\" when forces are placed on it** [5]. Players might thus suffer injuries with a higher rate compared to natural grass as a result of these \"unnatural\" forces. This theory is directly supported by the fact that ankle and foot injuries have particularly high increases of injury rates on artificial turf vs. natural grass [2]. In respective work, the authors also hypothesize that this effect is related to a \"lack of release between a player's shoe and a synthetic turf surface\". \n\nAn area that has been previously mostly unexplored, is to study **whether player movement patterns differ across synthetic turf and natural grass**, and if so, whether these differences in play behavior could explain some of the artefacts found when investigating injury rates on different types of surfaces. In this work, we thus tackle this question by studying full player tracking data of 250 players (100 with injuries) over two regular seasons of the NFL. Next, we summarize our research questions, findings and contributions of this work.\n\n1. Statistical tests on injury rates corroborate previous research showing clear evidence that the **risk of an injury is higher on artificial turf**. Utilizing survival analysis (Cox regression), we also find that the risk of getting an injury accumulates much faster on artificial surface. Additionally, we find that artificial turf elevates the risk of non-knee lower limb injuries significantly more which directly supports the findings of [2]. Repeated games on synthetic turf lower the risk of injury, maybe helping players to adjust posture stability or micro-movement patterns to a harder surface.\n2. To better understand how injury rates and observed effects differ within the data population at hand, we split the data by a player's roster position. While we find that **injuries on artifical turf are more frequent for each player's position**, the prevalence of this effect differs across player groups. Turf type has the most significant impact on injuries of **Cornerbacks who suffer a 600% higher rate of injuries on artificial turf in comparison to natural grass**.\n3. Next, we are interested in studying **whether players' movement patterns differ between synthetic and natural turf** which might be an explanation for observed higher injury rates. To study this and similar questions, we introduce a **flexible neural network model** that allows us to **model movement trajectories** in order to **predict arbitrary aspects of a play** (e.g., type of surface, player role, type of play, etc.). \n4. After demonstrating the **usibility of the model** by successfully predicting the role of a player just by looking at speed and acceleration based features of a movement trajectory, we investigate if **players' movement patterns allow us to predict the surface of the play**. However, our predictive experiments cannot show any strong evidence for such an effect leading to our conclusion that **players do not change their playstyle significantly across surfaces**.\n\nTo summarize, we confirm previous research and find that **playing on artificial turf increases the risk of injury**, this risk seems to differ for different types of player roles.\nHowever, we **cannot assume that different playstyles across surfaces are a leading causal effect for the higher rate of injuries on artificial turf**. Coupled with our findings about types of injuries and the insights from previous research, we rather believe that artefacts of the surface such as the lack of release between a player's shoe and a synthetic turf surface have an effect on the injury rates. However, we cannot observe whether a change in playstyle would be necessary in order to reduce the injury risk. Findings such as the fact that injury risks singificantly increase directly after switching from natural grass to synthetic turf could support this theory. Nonetheless, further research is necessary coupled with the need for more detailed data to study more fine-grained aspect of player movements such as posture or micro-movement in further detail.\n\nThe rest of this kernel includes all supporting experiments of our work and goes into detail about methodology, experiments, and insights.\n"},{"metadata":{"_kg_hide-input":true,"_kg_hide-output":true,"trusted":true},"cell_type":"code","source":"# simply change this to \"cpu\" in order to run on cpu kernels\ndevice = \"cuda\"\n\nimport os\nif True:\n    os.environ[\"OMP_NUM_THREADS\"] = \"1\"\n    os.environ[\"MKL_NUM_THREADS\"] = \"1\"\n    os.environ[\"OPENBLAS_NUM_THREADS\"] = \"1\"\n    os.environ[\"VECLIB_MAXIMUM_THREADS\"] = \"1\"\n    os.environ[\"NUMEXPR_NUM_THREADS\"] = \"1\"\n\n!pip install lifelines\n\nimport numpy as np\nimport pandas as pd\nimport gc\nfrom scipy.stats import norm\nimport warnings\nwarnings.filterwarnings(\"ignore\")\n\nfrom sklearn.preprocessing import OneHotEncoder\n\nfrom sklearn.base import BaseEstimator, TransformerMixin\nfrom sklearn.preprocessing import OneHotEncoder\nfrom scipy.spatial import Voronoi\nimport time\nimport torch\nfrom torch import nn\nfrom tqdm import tqdm_notebook as tqdm\nimport time\nfrom torch.utils.data import DataLoader, Dataset, TensorDataset\nfrom sklearn.preprocessing import StandardScaler\nfrom random import random\nfrom copy import copy\n\nfrom joblib import Parallel, delayed\nimport multiprocessing\nimport torch.nn.functional as F\nfrom sklearn.metrics import roc_auc_score, accuracy_score\nfrom sklearn.model_selection import KFold, GroupKFold, StratifiedKFold\nfrom statsmodels.stats.contingency_tables import mcnemar\n\nfrom lifelines import CoxPHFitter\nimport matplotlib.pyplot as plt\nimport matplotlib.patches as mpatches\nimport seaborn as sns","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## Data load\n\nFortunately, injuries during NFL games are quite rare, however for statistical analysis of the injury reasons it creates additional difficulties. With such low number of injuries it is usually difficult to conduct proper statistical analysis.\nEven though data was collected over two full seasons, even 1-2 extra erroneous or missed injury records can drive significance of statistical tests. Out of **105** injury records (counting a double injury case) available, **28** do not have identifier of the corresponding play. Fortunately, we have the data about the game where each injury accured, so, in order to **preserve all injury information** available, we assign the last plays of these games as those which led to the injury. This allows us to conduct statistical analyses on all 105 injury records at hand."},{"metadata":{"trusted":true,"_kg_hide-input":true,"_kg_hide-output":true},"cell_type":"code","source":"DTYPE = np.float32\nFREQ = 3\n\nInjuryRecord = pd.read_csv('../input/nfl-playing-surface-analytics/InjuryRecord.csv')\n\nPlayList = pd.read_csv('../input/nfl-playing-surface-analytics/PlayList.csv')\nlast_plays = PlayList.groupby('GameID')['PlayKey'].last()\nInjuryRecord.loc[InjuryRecord['PlayKey'].isnull(), 'PlayKey'] = InjuryRecord['GameID'].map(last_plays)[InjuryRecord['PlayKey'].isnull()].values\nInjuryRecord = InjuryRecord.groupby('PlayKey').first().reset_index()\n\nPlayList = PlayList.merge(InjuryRecord[['PlayKey', 'DM_M1']], left_on='PlayKey', right_on='PlayKey', how='left')\nPlayList['DM_M1'] = PlayList['DM_M1'].fillna(0)\n\n\n\nPlayerTrackData = pd.read_csv('../input/nfl-playing-surface-analytics/PlayerTrackData.csv',\\\n                              usecols=[\"PlayKey\", \"time\", \"x\", \"y\", \"dir\"],\\\n                              dtype={\"time\": DTYPE, \"x\": DTYPE, \"y\": DTYPE, \"dir\": DTYPE})\nPlayerTrackData['time'] = (PlayerTrackData['time'] * 10).astype(np.int16)\nPlayerTrackData = PlayerTrackData[PlayerTrackData['time'] % FREQ == 0]\n\ngc.collect()\n\nPlayerTrackData = PlayerTrackData.merge(PlayList[['PlayKey', 'RosterPosition']], left_on='PlayKey', right_on='PlayKey', how='left')\nPlayerTrackData['sx'] = PlayerTrackData['x'].diff().astype(DTYPE) * 10 / FREQ\nPlayerTrackData['sy'] = PlayerTrackData['y'].diff().astype(DTYPE) * 10 / FREQ\nPlayerTrackData['s'] = np.sqrt(PlayerTrackData['sx']**2 + PlayerTrackData['sy']**2).astype(DTYPE)\n\nPlayerTrackData['ax'] = PlayerTrackData['sx'].diff().astype(DTYPE)\nPlayerTrackData['ay'] = PlayerTrackData['sy'].diff().astype(DTYPE)\nPlayerTrackData['a'] = np.sqrt(PlayerTrackData['ax']**2 + PlayerTrackData['ay']**2).astype(DTYPE)\n\nPlayerTrackData['sx_prev'] = PlayerTrackData['sx'].shift()\nPlayerTrackData['sy_prev'] = PlayerTrackData['sy'].shift()\nPlayerTrackData['s_prev'] = PlayerTrackData['s'].shift()\n\nPlayerTrackData = PlayerTrackData[PlayerTrackData['time'] > FREQ].drop(['dir'], axis=1)\n\nPlayerTrackData = PlayerTrackData.merge(PlayList[['PlayKey', 'PlayerKey', 'FieldType']], left_on='PlayKey', right_on='PlayKey')\nPlayerTrackData = PlayerTrackData.merge(InjuryRecord[['PlayKey', 'DM_M1']], left_on='PlayKey', right_on='PlayKey', how='left')\nPlayerTrackData['DM_M1'] = PlayerTrackData['DM_M1'].fillna(0)\n\ngc.collect()\n\nPlayerTrackData['cos_th'] = ((PlayerTrackData['ax']*PlayerTrackData['sx_prev'] + PlayerTrackData['ay']*PlayerTrackData['sy_prev']) / PlayerTrackData['a'] / PlayerTrackData['s_prev']).fillna(1)\nPlayerTrackData['a_fwd'] = PlayerTrackData['a'] * PlayerTrackData['cos_th']\nPlayerTrackData['a_sid'] = np.sqrt(PlayerTrackData['a']**2 - PlayerTrackData['a_fwd']**2)\nPlayerTrackData['a_sid'] = PlayerTrackData['a_sid'].fillna(0) * np.sign(PlayerTrackData['ax']*PlayerTrackData['sy_prev'] - PlayerTrackData['ay']*PlayerTrackData['sx_prev'])\n","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## Correcting the sample bias\n\nThe data provided is a sample of all players in NFL; unfortunately, most of the players who did not suffer any injuries are not available for analysis. However, to properly conduct statistical analytics, we need to account for frequencies of plays on artificial and natural turf. Though it does not have a major impact on statistical tests performed, the resulting injury frequencies are unrealistically high due to us having a rather balanced sample of injured and non-injured players. Therefore, in order to correct the numbers, we calculate a very rough correction term for missing non-injured players and plays. Assuming there are 53 players per team in 32 teams, we estimate there are 1696 players overall, while only 250 players are present in the data. So, for the correction term we assume there are extra 1446 = 1696 - 250 players with same data as non-injured players. We apply it to adjust estimates throughout this kernel if warranted."},{"metadata":{"_kg_hide-input":true,"_kg_hide-output":false,"trusted":true},"cell_type":"code","source":"corr_term = (32 * 53 - 250) / 250\nprint(\"Correction factor:\", corr_term)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## Playing surface impact on injuries\n\nFirst, we study the fundamental question whether injury rates differ statistically between games played on artificial turf and natural grass.\nAn appropriate statistical test to identify if we can claim that players are more prone to injuries when they play on artificial turf is a **chi-squared test** (based on contingency table, other appropriate tests such as Fisher's exact test confirm obtained results). We run the test twice, once to assess how frequent the injuries are **per play**, and second - **per game**. Each test is performed once without and once with the bias correction factor."},{"metadata":{"trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"def csq_test(df, grp=None, target='DM_M1', c_term=1):\n    df0 = df[df[target] >= df.groupby('GameID')[target].cumsum()]\n    \n    if grp is not None:\n        df0 = df0.groupby(grp)['FieldType', target].max().reset_index()\n    \n    cond = df0['FieldType'] == \"Synthetic\"\n\n    e1 = df0[cond][target].sum()\n    n1 = cond.sum() * c_term\n    e2 = df0[~cond][target].sum()\n    n2 = (~cond).sum() * c_term\n    p_hat = (e1 + e2) / (n1 + n2)\n    z = (e1/n1 - e2/n2) / np.sqrt(p_hat * (1-p_hat) * (1/n1 + 1/n2))\n\n    return 1-norm.cdf(z), int(n1+n2), int(e1+e2)\n\nres = csq_test(PlayList, c_term=1)\nprint(f\"p-value by play {res[0]:<6.4f} ({res[2]} injuries from {res[1]} plays)\")\nres = csq_test(PlayList, grp='GameID', c_term=1)\nprint(f\"p-value by game {res[0]:<6.4f} ({res[2]} injuries from {res[1]} games)\")\n\nres = csq_test(PlayList, c_term=corr_term)\nprint(f\"Bias corrected p-value by play {res[0]:<6.4f} ({res[2]} injuries from {res[1]} plays)\")\nres = csq_test(PlayList, grp='GameID', c_term=corr_term)\nprint(f\"Bias corrected p-value by game {res[0]:<6.4f} ({res[2]} injuries from {res[1]} games)\")","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"We can see that in both cases p-values are clearly below 0.01, allowing us to reject the null hypothesis and to claim that **the risk of an injury is significantly higher on artifical turf**. Let us further take a look at how injury risk grows when a player accumulates plays/games on artificial turf compared to natural grass. In order to do so, we run a **survival analysis model** using **Cox regression** with the surface type as the only predictor. We run 2 regressions - one is counting number of plays played on a certain type of surface, and the other is counting number of games."},{"metadata":{"_kg_hide-input":true,"trusted":true},"cell_type":"code","source":"target = 'DM_M1'\n\n# Stop at the first injury\ndf = PlayList[PlayList[target] >= PlayList.groupby('PlayerKey')[target].cumsum()]\ndf = df.groupby(['GameID']).last().reset_index()[['PlayerKey', 'PlayKey', 'PlayerGamePlay', target, 'FieldType']]\ndf['FieldSynthetic'] = (df['FieldType'] == \"Synthetic\").astype(int)\ncph = CoxPHFitter()\ncph.fit(df[['FieldSynthetic', 'PlayerGamePlay', target]], duration_col='PlayerGamePlay', event_col=target)\n\nplay_preds = (1 - cph.predict_survival_function(pd.DataFrame({\"PlayerGamePlay\":[0,0], \"FieldSynthetic\":[0,1]}))[:60]) / corr_term\nplay_preds.columns = ['Natural grass', 'Synthetic turf']\n\n\ndf = PlayList.groupby(['GameID'])[['PlayerKey', 'FieldType', 'PlayerGame', 'DM_M1']].max().reset_index().sort_values(['PlayerKey', 'PlayerGame'])\ndf = df[df[target] >= df.groupby('PlayerKey')[target].cumsum()]\ndf = df.groupby('GameID').last()\ndf = df.groupby(['PlayerKey', 'FieldType']).apply(lambda z: pd.DataFrame(\\\n    {\"ngames\":[len(z)], \"DM_M1\":[z['DM_M1'].max()]}\\\n    )).reset_index().drop(['level_2'], axis=1)\ndf['FieldSynthetic'] = (df['FieldType'] == \"Synthetic\").astype(int)\ncph = CoxPHFitter()\ncph.fit(df[['FieldSynthetic', 'ngames', 'DM_M1']], duration_col='ngames', event_col='DM_M1')\n\ngame_preds = (1 - cph.predict_survival_function(pd.DataFrame({\"PlayerGamePlay\":[0,0], \"FieldSynthetic\":[0,1]}))[:15]) / corr_term\ngame_preds.columns = ['Natural grass', 'Synthetic turf']\n\n\nplt.style.use(['default'])\nfig = plt.figure(figsize = (15, 5))\n\nax = plt.subplot(1, 2, 1)\nplay_preds[['Synthetic turf', 'Natural grass']].plot(colors=['red', 'green'], ax=ax)\nax.yaxis.tick_right()\nplt.title('Injury risk per number of plays')\nplt.xlabel('# plays')\n#plt.ylabel('Injury risk')\nax.yaxis.set_label_position(\"right\")\nax.set_yticklabels(['{:,.1%}'.format(x) for x in ax.get_yticks()])\n\nax = plt.subplot(1, 2, 2)\ngame_preds[['Synthetic turf', 'Natural grass']].plot(colors=['red', 'green'], ax=ax)\nax.yaxis.tick_right()\nplt.title('Injury risk per number of games')\nplt.xlabel('# games')\n#plt.ylabel('Injury risk')\nax.yaxis.set_label_position(\"right\")\nax.set_yticklabels(['{:,.0%}'.format(x) for x in ax.get_yticks()])\nplt.ylim((0, 0.06))\n\nplt.subplots_adjust(right=1.1)\n\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"In the figures above, the x-axis depict the accumulating number of plays (left) and number of games (right), while the y-axis shows the injury risk for synthetic turf (red) and natural grass (green) as calculated by the survival model. \nAs each individual game or play shows higher injury risk on artificial surface, as expected the risk also accumulates much faster over plays/games. **After 60 plays, the risk of an injury on synthetic turf is 0.6% vs. only 0.35% on natural grass.** Similarly, after 15 games, the risk of injury on artificial turf rises to 5.6%, while it is 4% on natural grass."},{"metadata":{"_kg_hide-input":true,"_kg_hide-output":true,"trusted":true},"cell_type":"code","source":"play_preds","execution_count":null,"outputs":[]},{"metadata":{"_kg_hide-input":true,"_kg_hide-output":true,"trusted":true},"cell_type":"code","source":"game_preds","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## Types of injuries\n\nInjuries are mainly of two types: **knee** and **non-knee**, where latter is dominated by ankle injuries. "},{"metadata":{"trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"print(InjuryRecord['BodyPart'].value_counts().to_string())","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Previous research [2] has indicated that ankle and foot injuries have particularly high increases of injury rates on artificial turf vs. natural grass which is what we aim to study next. Limited data available does not let us to confidently draw statistical conclusions by type of the injury. But it is sufficient to gain some interesting insights. "},{"metadata":{"_kg_hide-input":true,"trusted":true},"cell_type":"code","source":"print(\"Knee injuries\")\nPlayList2 = PlayList.merge(InjuryRecord[InjuryRecord['BodyPart'] == 'Knee'][['PlayKey', 'DM_M1']].rename({\"DM_M1\": \"DM_M1_knee\"}, axis=1), \\\n                           left_on='PlayKey', right_on='PlayKey', how='left')\nPlayList2['DM_M1_knee'] = PlayList2['DM_M1_knee'].fillna(0)\n\nprint(f\"\\tOn synthetic turf {100*PlayList2[PlayList2['FieldType'] == 'Synthetic']['DM_M1_knee'].mean()/corr_term:6.4f}% per play, \"\\\n      f\"{100*PlayList2[PlayList2['FieldType'] == 'Synthetic'].groupby('GameID')['DM_M1_knee'].max().mean()/corr_term:4.2f}% per game\")\nprint(f\"\\tOn natural grass  {100*PlayList2[PlayList2['FieldType'] != 'Synthetic']['DM_M1_knee'].mean()/corr_term:6.4f}% per play, \"\\\n      f\"{100*PlayList2[PlayList2['FieldType'] != 'Synthetic'].groupby('GameID')['DM_M1_knee'].max().mean()/corr_term:4.2f}% per game\")\nprint()\n\nres = csq_test(PlayList2, target='DM_M1_knee')\nprint(f\"\\tp-value by play {res[0]:<6.4f}\")\n#print(f\"\\tp-value by play {res[0]:<6.4f} ({res[2]} injuries from {res[1]} plays)\")\nres = csq_test(PlayList2, target='DM_M1_knee', grp='GameID')\nprint(f\"\\tp-value by game {res[0]:<6.4f}\")\n#print(f\"\\tp-value by game {res[0]:<6.4f} ({res[2]} injuries from {res[1]} games)\")\nprint()\n\n\nprint(\"Non-knee injuries\")\nPlayList2 = PlayList.merge(InjuryRecord[InjuryRecord['BodyPart'] != 'Knee'][['PlayKey', 'DM_M1']].rename({\"DM_M1\": \"DM_M1_knee\"}, axis=1), \\\n                           left_on='PlayKey', right_on='PlayKey', how='left')\nPlayList2['DM_M1_knee'] = PlayList2['DM_M1_knee'].fillna(0)\n\nprint(f\"\\tOn synthetic turf {100*PlayList2[PlayList2['FieldType'] == 'Synthetic']['DM_M1_knee'].mean()/corr_term:6.4f}% per play, \"\\\n      f\"{100*PlayList2[PlayList2['FieldType'] == 'Synthetic'].groupby('GameID')['DM_M1_knee'].max().mean()/corr_term:4.2f}% per game\")\nprint(f\"\\tOn natural grass  {100*PlayList2[PlayList2['FieldType'] != 'Synthetic']['DM_M1_knee'].mean()/corr_term:6.4f}% per play, \"\\\n      f\"{100*PlayList2[PlayList2['FieldType'] != 'Synthetic'].groupby('GameID')['DM_M1_knee'].max().mean()/corr_term:4.2f}% per game\")\nprint()\n\nres = csq_test(PlayList2, target='DM_M1_knee')\nprint(f\"\\tp-value by play {res[0]:<6.4f}\")\n#print(f\"\\tp-value by play {res[0]:<6.4f} ({res[2]} injuries from {res[1]} plays)\")\nres = csq_test(PlayList2, target='DM_M1_knee', grp='GameID')\nprint(f\"\\tp-value by game {res[0]:<6.4f}\")\n#print(f\"\\tp-value by game {res[0]:<6.4f} ({res[2]} injuries from {res[1]} games)\")\nprint()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"We see that both groups show increased injury frequency on synthetic turf, allowing us to conclude that the analysis are not dominated by either of the two. From here onward we will not split the injuries by type.\n\nHowever another conclusions is clear from the numbers above: **artificial turf elevates the risk of non-knee (ankle, foot) injuries significantly more**. Even the collected number of records is sufficient to claim the statistical significance on this sub-group of injuries alone. It is aligned with the assumption that synthetic surface is \"harder\" and it puts more \"stress\" on ankles during the game. It is also worth pointing out that a lower increase in frequency of knee injuries might have the same root causes. Increased \"pressure\" on ankles might be the reason of higher knee injury risk.\n\nP-values observed are unstable due to low number of recorded cases. When the analysis will be repeated after 1-2 extra seasons of data, we expect p-values of both types of injuries to drop below 0.01. "},{"metadata":{},"cell_type":"markdown","source":"## Switching the surface\n\nIt is a common observation in sports medicine for players to report more discomfort after the switch from natural surface to an artificial one. As a simplified analysis, a quick check of injury rates split by the surface type of preceding game shows interesting numbers. "},{"metadata":{"_kg_hide-input":true,"trusted":true},"cell_type":"code","source":"df = PlayList.groupby(['GameID'])[['PlayerKey', 'FieldType', 'PlayerGame', 'DM_M1']].max().reset_index()\ndf = df[df['DM_M1'] >= df.groupby('PlayerKey')['DM_M1'].cumsum()]\ndf = df.groupby('GameID').last().reset_index().sort_values(['PlayerKey', 'PlayerGame'])\n\ndf['pred_type'] = df['FieldType'].shift()\ndf = df[df['PlayerGame'] > 1]\n\ndf0 = PlayList[PlayList['FieldType'] == 'Synthetic'].groupby('GameID')[['FieldType', 'DM_M1']].max().reset_index()\nprint(f\"Frequency on synthetic {100*df0['DM_M1'].mean() / corr_term:4.2f}% per game, \"\\\n      f\"{100*PlayList[PlayList['FieldType'] == 'Synthetic']['DM_M1'].mean() / corr_term:6.4f}% per play \"\n      f\"({int(df0['DM_M1'].sum())} injuries in {int(df0['DM_M1'].count() * corr_term)} games)\")\ndf0 = PlayList[PlayList['FieldType'] != 'Synthetic'].groupby('GameID')[['FieldType', 'DM_M1']].max().reset_index()\nprint(f\"Frequency on natural   {100*df0['DM_M1'].mean() / corr_term:4.2f}% per game, \"\\\n      f\"{100*PlayList[PlayList['FieldType'] != 'Synthetic']['DM_M1'].mean() / corr_term:6.4f}% per play \"\n      f\"({int(df0['DM_M1'].sum())} injuries in {int(df0['DM_M1'].count() * corr_term)} games)\")\nprint()\n\nfor f_type in df['FieldType'].unique():\n    print(f\"Previous game on {f_type}\")\n    \n    df1 = df[df['pred_type'] == f_type]\n    df0 = df1[df1['FieldType'] == 'Synthetic']\n    print(f\"\\tSynthetic {100*df0['DM_M1'].mean() / corr_term:4.2f}% ({int(df0['DM_M1'].sum())} injuries in {int(df0['DM_M1'].count() * corr_term)} games)\")\n    df0 = df1[df1['FieldType'] != 'Synthetic']\n    print(f\"\\tNatural   {100*df0['DM_M1'].mean() / corr_term:4.2f}% ({int(df0['DM_M1'].sum())} injuries in {int(df0['DM_M1'].count() * corr_term)} games)\")\n    print()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"We indeed see that **frequency of injuries on synthethic turf are higher when the preceeding game took place on natural grass**. It might be explained that players who had practice on \"harder\" surfaces are able to better adjust the posture to the surface. However, further research is warranted to study this artefact in more detail."},{"metadata":{},"cell_type":"markdown","source":"## Deeper look by player positions"},{"metadata":{},"cell_type":"markdown","source":"Next, we are interested in studying how injury rates differ after splitting by player positions."},{"metadata":{"trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"df = PlayList.groupby(\"GameID\")[['DM_M1', 'FieldType', 'RosterPosition']].max().reset_index()\nstat = df.groupby('RosterPosition', group_keys=True).apply(lambda z: pd.DataFrame(\\\n    {\"# plays\":[len(z)],\n     \"# injuries\":[z['DM_M1'].sum().astype(int)],\n     \"Inj Natural\":[z[z['FieldType'] != \"Synthetic\"]['DM_M1'].mean()/corr_term],\n     \"Inj Synthetic\":[z[z['FieldType'] == \"Synthetic\"]['DM_M1'].mean()/corr_term],\n     \"p-value\":[csq_test(df[df['RosterPosition'] == z['RosterPosition'].values[0]])[0]]\n    }\\\n    )).reset_index().drop(['level_1'], axis=1).\\\n    sort_values(\"# plays\", ascending=False)\n\nstat['p-value'] = stat['p-value'].round(3)\nstat[\"Inj Synthetic\"] = stat[\"Inj Synthetic\"].round(6)\nstat[\"Inj Natural\"] = stat[\"Inj Natural\"].round(6)\nstat[\"Inj diff\"] = ((stat[\"Inj Synthetic\"] / stat[\"Inj Natural\"]).round(2) - 1).fillna(0)\n\n#print(stat.to_string(index=False))\nstat.reset_index(drop=True).style.format({\n    \"Inj Natural\": '{:,.2%}'.format,\n    \"Inj Synthetic\": '{:,.2%}'.format,\n    \"Inj diff\": '+{:,.0%}'.format,\n})\n","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Firstly, **injuries on artifical turf are more frequent for each player's position**. That brings us to an important conclusion - the higher risk of injuries on artificial turf affects players of every role. In other words, it does not matter if a player is sprinting, rushing through the defensive line or trying to tackle or block an opponent - we see an elevated risk of injury.\n\nHowever, **the magnitude of the effect depends on a player's position** as imminent from the varying differences between synthetic and natural injury rates per roster position. Turf type has the most significant impact on injuries of **Cornerbacks**, with the disclaimer of having only 12 records of injuries to judge. However, even with such low numbers we see a statistically significant effect (p-value below 0.01) as almost all the recorded injuries happened on artificial turf. Out of the positions with higher number of observations, we see **Wide Receivers** suffering injuries almost 60% more frequently on artificial turf compared to natural grass, with the second lowest p-value. WR and CB injury risk is also correlated with type of the play, as they get more injuries during passing plays where they are more involved, than during rushing plays. "},{"metadata":{"_kg_hide-input":true,"trusted":true},"cell_type":"code","source":"df = PlayList[['PlayType', 'RosterPosition', 'DM_M1']].copy()\ndf['PlayType'] = df['PlayType'].apply(lambda x: str(x).split()[0])\ndf = df[\\\n    np.logical_and(df['PlayType'].apply(lambda x: str(x)[:7]).isin(['Pass', 'Rush', 'Kickoff']), \\\n                   df['RosterPosition'].isin(['Cornerback', 'Wide Receiver']))]\\\n    .groupby(['RosterPosition', 'PlayType']).apply(lambda z: pd.DataFrame(\\\n    {\"# plays\":[len(z)],\n     \"# injuries\":[z['DM_M1'].sum().astype(int)],\n     \"Injuries freq\":[z['DM_M1'].mean()/corr_term]\n    }\\\n    ))\n\ndf[\"Injuries freq\"] = df[\"Injuries freq\"].round(7)\n\ndf.style.format({\n    \"Injuries freq\": '{:,.4%}'.format,\n})\n","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## Player movement data"},{"metadata":{},"cell_type":"markdown","source":"The provided data gives rich information about players' movements: directions, speed, acceleration. One of the main questions of our work is to study effects of player movement on injury rates with a specific focus on differences between playing surfaces. Based on our initial insights in differences between player positions, we expect the movement profile to also depend on a player's position. For example, we expect higher speed and accelerations from WR and CB compared to other positions. To confirm this, we report a simple initial analysis differentiating speed by player position and conduct most of our remaining experiments split by roster position."},{"metadata":{"_kg_hide-input":true,"_kg_hide-output":true,"trusted":true},"cell_type":"code","source":"df = PlayerTrackData.groupby('PlayKey')[['RosterPosition', 's', 'a', 'a_fwd', 'a_sid', 'DM_M1', 'FieldType']].max().reset_index()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"This is an example of how maximum speed per play is differently distributed for some positions."},{"metadata":{"_kg_hide-input":true,"trusted":true},"cell_type":"code","source":"import matplotlib.pyplot as plt\nimport seaborn as sns\n%matplotlib inline\n\nft = 's'\nnbins = 100\n\nplt.style.use(['default'])\nfig = plt.figure(figsize = (12, 5))\n\n\ndf0 = df[df[ft] <= 12]\n\nsns.distplot(df0[~df0['RosterPosition'].isin(['Cornerback', 'Wide Receiver'])][ft], kde=True, norm_hist=True, bins=nbins, hist_kws={\"histtype\": \"step\"}, color='blue', label='Other')\nsns.distplot(df0[df0['RosterPosition'] == 'Cornerback'][ft], kde=True, norm_hist=True, bins=nbins, hist_kws={\"histtype\": \"step\"}, color='orange', label='Cornerbacks')\nsns.distplot(df0[df0['RosterPosition'] == 'Wide Receiver'][ft], kde=True, norm_hist=True, bins=nbins, hist_kws={\"histtype\": \"step\"}, color='red', label='Wide Receivers')\n\nax = fig.axes[0]\nax.set_yticklabels(['{:,.0%}'.format(x) for x in ax.get_yticks()])\nplt.xlabel(\"Speed yards/s\")\nplt.ylabel(\"\")\nplt.legend()\n\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## Play model\n\nBased on our initial insights, we are interested in now studying whether players' movement patterns have predictive power on certain aspects of a play. To that end, we introduce a **flexible neural network** allowing us to **model movement trajectories** in order to **predict arbitrary aspects of a play**. \n\nThe following figure depicts the general architecture of the network. We represent each individual play as a sequence of features over the course of of the trajectory taken (left-hand part of the figure). \nThe x-axis represents the individual movement frames at consideration; in our experiments we cut the trajectories after 200 frames. Each frame at interest represents one element of the sequence, and each element can have multiple features which are depicted on the y-axis. We only consider **speed** and **acceleration** based properties of a trajectory, but the data can be arbitrarily extended to other types of features in straight-forward fasion (such as orientation, position, etc.). As an example, the red element of the matrix, captures the acceleration of the player after the third frame at consideration in the movement trajectory of respective play.\n\nOn this data representation, the model applies a convolutional layer with a kernel size of 50 and a dimension of 64. This is followed by activation, batch normalization, and average pooling. The final step is then to apply a typical linear layer, and to have the final target classes as outputs. In this visualization, the target consists of three individual classes, but this can be adapted according to the prediction task at hand and we will utilize different targets throughout our experiments shown next."},{"metadata":{},"cell_type":"markdown","source":"![Neural Network Architecture](https://i.imgur.com/kuxIzQL.png)"},{"metadata":{},"cell_type":"markdown","source":"For each of the following experiments, we sample at maximum 10,000 individual plays and run 4-fold player-based cross-validation, meaning that plays of a single player will not overlap between training and validation folds to avoid any leakage. We always report full out-of-fold classification accuracy as well as responsing baseline accuracy always predicting the majority class."},{"metadata":{"trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"nbags = 1\nn_epochs = 50\nnfolds = 4\n\ndef preprocess(df_base, inplace=False):\n    if inplace:\n        df = df_base\n    else:\n        df = df_base.copy()\n    \n    return df\n\ndef pad_sequences(seq, pad=500):\n\n    seq = [i[:pad] for i in seq]\n    seq = np.array([i + [-1.]*(pad-len(i)) for i in seq])\n    return seq\n\nclass FE(BaseEstimator, TransformerMixin):\n    def fit(self, train, *_):\n        \n        return self\n    \n    def transform(self, df_base, *_):\n        \n        D_TYPE = np.float32\n        df = df_base.copy()\n        \n        res = {}\n        \n        pad = 200\n        \n        s = pad_sequences(df[\"s\"].values)\n        a = pad_sequences(df[\"a\"].values)\n        sx = pad_sequences(df[\"sx\"].values)\n        sy = pad_sequences(df[\"sy\"].values)\n        ax = pad_sequences(df[\"ax\"].values)\n        ay = pad_sequences(df[\"ay\"].values)\n        \n        res['sequence'] = np.stack([s, a, sx, sy, ax, ay], axis=2)        \n        \n        res['target'] = df[\"target\"]\n\n        return res\n    \n\n\nclass OwnSampler(torch.utils.data.sampler.Sampler):\n    def __init__(self, idx):\n        self.idx = idx\n    def __iter__(self):\n        return iter(self.idx)\n    def __len__(self):\n        return len(self.idx)\n\nclass NFLDataset(Dataset):\n    def __init__(self, data,  ds=\"train\", mode=\"default\"):\n        fe = FE().fit(data)\n            \n        self.mode=mode\n        \n        self.xkeys = ['sequence']\n        self.data = fe.transform(data)\n\n        self.num_classes = len(data.target[0])\n        self.ds = ds\n        self.num_features = {}\n        for k in self.xkeys:\n            self.num_features[k] = self.data[k].shape[1:]\n        \n    def __len__(self):\n        return len(self.data['target'])\n\n    def __getitem__(self, idx):\n        \n        if self.mode == \"default\":\n            data = self.data\n\n        row = {k:data[k][idx] for k in self.xkeys}\n        \n        if self.ds == \"train\":\n            target = np.array(data[\"target\"][idx])\n        else:\n            target = 0\n        \n        return {**row, 'target': target}\n    \ndim = 64\n\nclass MyModel(nn.Module):\n    def __init__(self, ds_train):\n        super(MyModel, self).__init__()\n        \n        self.cnn = torch.nn.Sequential(\n            torch.nn.Conv1d(ds_train.num_features['sequence'][1], dim, 50),\n            torch.nn.ReLU(),\n            nn.BatchNorm1d(dim, eps=1e-05, momentum=0.1, affine=True),\n        )\n\n        self.classifier = torch.nn.Sequential(\n            \n            torch.nn.Linear(dim, dim),\n            torch.nn.ReLU(),\n            nn.BatchNorm1d(dim, eps=1e-05, momentum=0.1, affine=True),\n\n            torch.nn.Linear(dim, ds_train.num_classes),\n\n            )\n        \n    def forward(self, sequence):\n        \n        sequence = sequence\n        \n        x = self.cnn(sequence.permute(0,2,1)).permute(0,2,1)\n        \n        x = F.avg_pool2d(x,(x.shape[1],1)).squeeze(1)\n        \n        return self.classifier(x)\n\n    \ndef fit_nn(trn_idx, val_idx, ds_train, ds_val_default, y_train, fold, bag, seed):\n        \n    torch.manual_seed(seed)\n    \n    valid_sampler = OwnSampler(val_idx)\n    valid_loader = torch.utils.data.DataLoader(ds_val_default, batch_size=64, num_workers=0, sampler=valid_sampler, pin_memory=True)\n\n    train_sampler = torch.utils.data.sampler.SubsetRandomSampler(trn_idx)\n    train_loader = torch.utils.data.DataLoader(ds_train, batch_size=128, num_workers=0, sampler=train_sampler, pin_memory=True, drop_last=True)\n    \n    model = MyModel(ds_train)\n    \n    model = model.to(device)\n\n    optimizer = torch.optim.Adam(model.parameters(), lr=0.001)\n    criterion = nn.BCEWithLogitsLoss()\n\n    best_score = 0\n    best_preds = None\n    best_epoch = -1\n\n    start_time = time.time()\n    for epoch in range(n_epochs):\n        \n        s = time.time()\n\n        model.train()\n        avg_train_loss = 0\n        optimizer.zero_grad()\n        \n        for idx, data in enumerate(train_loader):\n            \n            \n            \n            X = {k:data[k].to(device).float() for k in ds_train.xkeys}\n            y = data['target'].to(device).float()\n\n            preds = model(**X)\n\n            loss = criterion(preds, y) \n\n            loss.backward()\n\n            optimizer.step()\n            optimizer.zero_grad()\n            avg_train_loss += loss.detach().cpu().numpy() / len(train_loader)\n\n\n        avg_val_loss = 0\n        model.eval()\n        all_preds = []\n        for idx, data in enumerate(valid_loader):\n            X = {k:data[k].to(device).float() for k in ds_train.xkeys}\n            y = data['target'].to(device).float()\n\n            preds = model(**X)\n            loss = criterion(preds, y) \n            avg_val_loss += loss.detach().cpu().numpy() / len(valid_loader) \n            all_preds.append(preds.detach().cpu())\n\n        all_preds = np.vstack(all_preds)\n        \n        score = accuracy_score(np.argmax(y_train[val_idx], axis=1), np.argmax(all_preds, axis=1))\n\n        \n        if score >= best_score:\n            best_score = score\n            best_preds = all_preds\n            best_epoch = epoch\n            #torch.save(model.state_dict(), f\"{fold}_{bag}.pt\")\n\n        #print(f\"{fold:1}:{epoch:2} avg_train_loss {avg_train_loss:<8.4f} avg_val_loss {avg_val_loss:<8.4f} val_accuracy {score:<8.4f}  {time.time()-s:<2.2f}\")\n\n    #print(f\"fold {fold:1} best {best_score:<8.6f} {best_epoch:3}, {(time.time()-start_time) / n_epochs:<2.2f} s/ep\")\n    return best_preds, val_idx","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"df = pd.DataFrame(columns=[\"Roster position\", \"Accuracy\", \"Baseline accuracy\"])\ndef run_model(roster=[\"Wide Receiver\", \"Cornerback\", \"Running Back\"], target=\"RosterPosition\", df=None, verbose=True):\n    \n    np.random.seed(seed=42)\n    train = PlayerTrackData[PlayerTrackData.RosterPosition.isin(roster)]\n    n_targets = train.RosterPosition.nunique()\n\n    #print(train.RosterPosition.unique())\n\n    try:\n        sample = np.random.choice(train.PlayKey.unique(), size=10_000, replace=False)\n    except:\n        sample = train.PlayKey.unique()\n\n    train = train[train.PlayKey.isin(sample)]\n\n    train[\"target\"] = list(OneHotEncoder(sparse=False).fit_transform(train[target].values.reshape(-1,1)))\n\n    train_agg = train.groupby(\"PlayKey\").agg({'s': lambda x: list(x), \n                                              'a': lambda x: list(x),\n                                                'sx': lambda x: list(x), \n                                              'sy': lambda x: list(x), \n                                              'ax': lambda x: list(x), \n                                              'ay': lambda x: list(x), \n                                              'target': lambda x: list(np.mean(x)), \n                                              'PlayerKey': lambda x: max(x)})\n\n    y_train = np.array([np.array(x) for x in train_agg['target'].values])\n    \n    \n\n    train_processed = preprocess(train_agg, inplace=True)\n    ds_train = NFLDataset(train_processed, ds=\"train\", mode=\"default\")\n\n    ds_val_default = copy(ds_train)\n    ds_val_default.mode = \"default\"\n\n    ds_train.num_features\n\n    kf = GroupKFold(n_splits = nfolds)\n\n    start_time = time.time()\n\n    res = Parallel(n_jobs=-1, temp_folder=\"/tmp\", max_nbytes=None, backend=\"multiprocessing\")( \\\n        delayed(fit_nn)(trn_idx, val_idx, ds_train, ds_val_default, y_train, fold, bag, np.random.randint(100_000)) \\\n        for fold, (trn_idx, val_idx) in enumerate(kf.split(np.arange(len(train_agg)), groups=train_agg.PlayerKey.values)) \\\n        for bag in range(nbags))\n\n    from scipy.special import softmax\n\n    oof = np.zeros((len(train_agg), y_train.shape[1]))\n\n    for val, idx in res:\n        oof[idx] += softmax(np.array(val), axis=1) / nbags\n        \n    y_true = np.argmax(y_train,axis=1)\n    y_pred = np.argmax(oof, axis=1)\n    \n    if verbose:\n        print(\"Baseline accuracy:\", y_train.mean(axis=0).max())\n        print(\"Full CV Accuracy: {:<8.4f}\".format(accuracy_score(y_true, y_pred)))\n  \n    \n    if df is not None:\n        df = df.append({\"Roster position\": roster[0], \"Accuracy\":  accuracy_score(y_true, y_pred),\n                        \"Baseline accuracy\": y_train.mean(axis=0).max()}, ignore_index=True)\n    \n    if df is not None:\n        return df","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"To demonstrate the functionality and general predictive power of our model, we first **predict a player's roster position** (limited to *Wide Receiver*, *Cornerback*, and *Running Bac*k) just based on observing speed and acceleration of the corresponding movement trajectory where we already showed earlier that differences in movement patterns seem to exist. "},{"metadata":{"trusted":true,"_kg_hide-input":false},"cell_type":"code","source":"run_model(roster=[\"Wide Receiver\", \"Cornerback\", \"Running Back\"], target=\"RosterPosition\")","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"The results showcase clear predictive power getting close to ~80% accuracy which is significantly higher than a simple baseline accuracy of ~51%. This corroborates our findings from above that player movement patterns differ between separate player positions, even when one just looks at speed and acceleration based features. At the same time, this experiment **demonstrates the usefulness of our proposed neural network model**."},{"metadata":{},"cell_type":"markdown","source":"Coming back to one of the main questions of this analytics report, we want to further understand effects between players' movements and injury rates on different types of surfaces. We have already shown that there is a significantly higher injury rate on plays on artificial turf vs. those on natural turf. While we have also showcased that this effect exists across player roles, the prevalance of it is different across various roles. At the same time, we know that players' movement patterns clearly differ between distinct types of players.\n\nDirectly predicting if a play leads to an injury is not appropriate due to the very low number of events. However, one of our main questions is \n **if players' movements show different patterns across playing surfaces** for which we have a lot of data available. To that end, we utilize our flexible model and attempt to **predict the surface just by looking at players' trajectories**. We start by considering all roster positions and predict a play's surface."},{"metadata":{"trusted":true},"cell_type":"code","source":"run_model(roster=PlayList.RosterPosition.unique(), target=\"FieldType\")","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"We can clearly see that we are unable to outperform a simple baseline model, meaning that most of the trajectory based movement patterns hold no signal in order to predict the surface of a pitch. Based on our insights from above, we now attempt to predict the surface for different roster positions separately."},{"metadata":{"trusted":true},"cell_type":"code","source":"df = run_model(roster=[\"Linebacker\"], target=\"FieldType\", df=df, verbose=False)\ndf = run_model(roster=[\"Offensive Lineman\"], target=\"FieldType\", df=df, verbose=False)\ndf = run_model(roster=[\"Wide Receiver\"], target=\"FieldType\", df=df, verbose=False)\ndf = run_model(roster=[\"Safety\"], target=\"FieldType\", df=df, verbose=False)\ndf = run_model(roster=[\"Defensive Lineman\"], target=\"FieldType\", df=df, verbose=False)\ndf = run_model(roster=[\"Cornerback\"], target=\"FieldType\", df=df, verbose=False)\ndf = run_model(roster=[\"Running Back\"], target=\"FieldType\", df=df, verbose=False)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"plt.style.use(['default'])\nplt.figure(figsize = (12, 8))\n\nsns.swarmplot(x=\"Baseline accuracy\", y=\"Roster position\", data=df, color=\"red\")\nsns.swarmplot(x=\"Accuracy\", y=\"Roster position\", data=df, color=\"blue\")\n\nfor i in range(len(df)):\n    plt.plot([df.iloc[i][\"Baseline accuracy\"], df.iloc[i][\"Accuracy\"]], [i,i], color=\"black\")\n    t = f'Diff: {np.round((df.iloc[i][\"Accuracy\"]-df.iloc[i][\"Baseline accuracy\"]),5)}'\n    plt.annotate(t, xy=(df.iloc[i][\"Baseline accuracy\"]+(df.iloc[i][\"Accuracy\"]-df.iloc[i][\"Baseline accuracy\"])/2, i-0.07), ha='center')\n\np1 = mpatches.Patch(color='red', label='Baseline accuracy')\np2 = mpatches.Patch(color='blue', label='Accuracy')\np = plt.legend(handles=[p1,p2])","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"This figure presents the summarized accuracy results (x-axis) of running the model of predicting the type of surface for each player role separately (y-axis). Again, we contrast the majority baseline accuracy (red) with the actual achieved accuracy (blue); the lines depict the difference in accuracy. While we can see minor improvements over baseline accuracy, we can at maximum see an improvement of ~5% over the baseline accuracy which is not enough to justify any significant predictive effect and we would rather account it to certain noise in the movement patterns. In the experiment above where we predicted the player role, effects where much stronger and supportive of the hypothesis. However, the slight differences between player roles warrant future investigations with more data. The benefit of our proposed model is its flexibility and that it can be also utilized to study different objectives when looking at player trajectories and also incorporate different features.\n\nTo summarize, we cannot find any strong signal in the data for predicting the surface. We thus **cannot find any evidence that players' movement patterns significantly differ between different surfaces** meaning that players do not appear to adapt their behavior based on the type of pitch they play on.\n\nSome observations suggest (but of course, do not claim) that the surface type impacts micro details of player's movement, which might include posture, placement of the feet while running and so on. A single game on synthetic turf might be sufficient for a professional player to adjust such micro details to have a significant impact on the injury risk. In order to test this assumption, more detailed data about players movements is required, such as further sensor data of player movement patterns related to the foot. Also, it would be valuable to access practising data of players in order to study these and similar effects in greater detail."},{"metadata":{},"cell_type":"markdown","source":"### Further remarks\n\n* This kernel currently runs on GPU environment. You can easily change the kernel to run on CPU environment by changing the device variable to ``device = \"cpu\"``. Commenting some of the neural network runs out will also speed up the runtime of the kernel.\n* While simple analyses on weather data did not reveal any interesting facets to report, we found that movement patterns do not differ significantly across weather conditions by applying our machine learning play model to predict weather. \n* Another way to split the data is by type of play (rush, pass, etc.). However, the data correlates with player roles and we could not find any interesting additional effects next to those reported.\n* We believe that these and similar artefacts of plays can be further studied with our introduced neural network model. There is also still quite some room to improve predictions by e.g., tuning (i) the network architecture, (ii) learning procedure, or (iii) utilized features. For example, in some experiments we replaced the CNN layer with an LSTM layer showcasing similar performance, combining these two types of layers might improve predictive power of the model further. Also, for now we have only considered speed and acceleration based features to not introduce any position-based leakages to the model. Carefully investigating these and other properties of movement such as orientation or direction seems worthy to study in future."},{"metadata":{},"cell_type":"markdown","source":"### References\n\n1. https://www.lawnstarter.com/blog/sports-turf/nfl-mlb-teams-artificial-turf-2019/\n2. Mack C, Hershman E, Anderson R, et al. Higher rates of lower extremity injury on synthetic turf compared with natural turf among National Football League athletes: epidemiologic confirmation of a biomechanical hypothesis [published online November 19, 2018]. Am J Sports Med. doi:10.1177/0363546518808499\n3. Loughran, Galvin J., et al. \"Incidence of knee injuries on artificial turf versus natural grass in National Collegiate Athletic Association American football: 2004-2005 through 2013-2014 seasons.\" The American journal of sports medicine 47.6 (2019): 1294-1301.\n4. Kent, Richard, et al. \"The mechanics of American football cleats on natural grass and infill-type artificial playing surfaces with loads relevant to elite athletes.\" Sports biomechanics 14.2 (2015): 246-257.\n5. https://orthoinfo.aaos.org/en/diseases--conditions/turf-toe"}],"metadata":{"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"},"language_info":{"name":"python","version":"3.6.4","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"}},"nbformat":4,"nbformat_minor":1}