{"cells":[{"metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true},"cell_type":"code","source":"# This Python 3 environment comes with many helpful analytics libraries installed\n# It is defined by the kaggle/python docker image: https://github.com/kaggle/docker-python\n# For example, here's several helpful packages to load in \n\nimport numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\nfrom matplotlib import pyplot as plt\nimport seaborn as sns\n\n# Input data files are available in the \"../input/\" directory.\n# For example, running this (by clicking run or pressing Shift+Enter) will list all files under the input directory\n\nimport os\nfor dirname, _, filenames in os.walk('/kaggle/input'):\n    for filename in filenames:\n        print(os.path.join(dirname, filename))\n\n# Any results you write to the current directory are saved as output.","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"d629ff2d2480ee46fbb7e2d37f6b5fab8052498a","_cell_guid":"79c7e3d0-c299-4dcb-8224-4455121ee9b0","trusted":true},"cell_type":"code","source":"playlist = pd.read_csv('/kaggle/input/nfl-playing-surface-analytics/PlayList.csv')\ninjuries = pd.read_csv('/kaggle/input/nfl-playing-surface-analytics/InjuryRecord.csv')\ntrackdata = pd.read_csv('/kaggle/input/nfl-playing-surface-analytics/PlayerTrackData.csv')","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"plt.figure(figsize=(5,3))\ninjuries['Surface'].value_counts().plot(kind='bar')\nplt.title('Injuries By Surface')\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Let us look at testing the hypothesis that the synthetic field has a higher occurance of injuries, in other words\n\n$H_0: \\hat{p}_{syn}=\\hat{p}_{nat}$\n\n$H_A: \\hat{p}_{syn}>\\hat{p}_{nat}$"},{"metadata":{},"cell_type":"markdown","source":"Given that there are a lot of observations, we will just use a pooled Z test statistic, which is defined as\n\n$Z = \\left(\\frac{\\hat{p}_{syn}-\\hat{p}_{nat}}{SE} \\right)$\n\nWhere $SE^2 = p_{pooled}*(1-p_{pooled})*\\left(\\frac{1}{n_{syn}}+\\frac{1}{n_nat}\\right)$\n\n\n$p_{pooled} = n_{injuries}/n_{plays}$"},{"metadata":{"trusted":true},"cell_type":"code","source":"import scipy\nn_injuries = injuries.shape[0]\nn_syn = injuries.Surface.value_counts()[0]\nn_nat = injuries.Surface.value_counts()[1]\nn_plays = playlist.PlayKey.nunique()\nn_plays_syn = playlist.FieldType.value_counts()[1]\nn_plays_nat = playlist.FieldType.value_counts()[0]\n\nprop_syn = n_syn/n_plays_syn\nprint('Proportion of injuries on Synthetic field is ',prop_syn)\nprop_nat = n_nat/n_plays_nat\nprint('Proportion of injuries on Natural field is ',prop_nat)\n\nsigma_syn = np.sqrt(prop_syn*(1-prop_syn)/n_plays_syn)\n#print('SD of Synthetic Field: ',sigma_syn)\n\nsigma_nat = np.sqrt(prop_nat*(1-prop_nat)/n_plays_nat)\n#print('SD of Natural Field: ',sigma_nat)\n\npooled_prop = (n_syn+n_nat)/(n_plays_nat+n_plays_syn)\npooled_SE = np.sqrt(pooled_prop*(1-pooled_prop)*(1/n_plays_syn + 1/n_plays_nat))\n\nZ = (prop_syn - prop_nat)/pooled_SE\nprint('The Z score for the difference in proportions is :',Z)\np =  scipy.stats.norm.cdf(-abs(Z))*2 #twosided\np1 = 1-scipy.stats.norm.cdf(abs(Z))\nprint('Corresponding 2 tailed p value: ',p)\nprint('P Value for testing if p_syn>p_nat: ',p1)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"So a simple two tailed Z test is not significant at a typical 95% level, so we cannot actually say they differ, but there is some evidence at the same 95% level that supports the idea that there is a higher proportion of injuries on Synthetic Fields.  If all of the games were played on natural field, we would expect that there would be almost 23 less injuries in this time frame. \n\nThere are a lot of assumptions going into this, but at a basic level, it is nice to see that Stats 101 is actually useful :)"},{"metadata":{"trusted":true},"cell_type":"code","source":"","execution_count":null,"outputs":[]}],"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat":4,"nbformat_minor":1}