{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.10.13","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":457469,"sourceType":"datasetVersion","datasetId":209295}],"dockerImageVersionId":30673,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"## 1. Introduction\n<font color=#6698FF> Name: Jesse Koert \n    \n<font color=#6698FF> Username in Kaggle: JesseKoert,\n    \n<font color=#6698FF> Score: N.A.\n    \n<font color=#6698FF> Leaderboard rank: N.A.","metadata":{}},{"cell_type":"markdown","source":"#### Importing libaries","metadata":{}},{"cell_type":"code","source":"import pandas as pd # for dataframes\n\nimport matplotlib.pyplot as plt # for plotting\n\n# for calculations\nimport numpy as np\nimport seaborn as sns\n\n# for making an algorithm, cleaning and plotting\nfrom sklearn import ensemble, tree, linear_model\nfrom sklearn.model_selection import train_test_split, cross_val_score\nfrom sklearn.metrics import r2_score, mean_squared_error\nfrom sklearn.utils import shuffle\nfrom sklearn.preprocessing import StandardScaler\nfrom sklearn.linear_model import LogisticRegression\nfrom sklearn import metrics\nfrom sklearn import preprocessing\n\n# neural networking\nfrom keras.models import Sequential\nfrom keras.layers import Dense, Dropout\nimport keras as keras","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","execution":{"iopub.status.busy":"2024-04-21T19:24:07.808093Z","iopub.execute_input":"2024-04-21T19:24:07.808743Z","iopub.status.idle":"2024-04-21T19:24:12.953607Z","shell.execute_reply.started":"2024-04-21T19:24:07.808706Z","shell.execute_reply":"2024-04-21T19:24:12.952280Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## 2. Data\n","metadata":{}},{"cell_type":"markdown","source":"### 2.1 Dataset\n<font color=#6698FF> First, we load the train dataset. Then we plot the first 10 rows to see what the data looks like.","metadata":{}},{"cell_type":"code","source":"# load the dataset in a data frame\nfull_train_data = pd.read_csv('/kaggle/input/parkinsons-disease-speech-signal-features/pd_speech_features.csv')\n\n#print the size of the data\nprint(\"Data size: rows:\", full_train_data.shape[0], \"columns:\", full_train_data.shape[1])\n\n#print the first 10 entries of the dataframe\nfull_train_data.head(10)     ","metadata":{"execution":{"iopub.status.busy":"2024-04-21T19:24:12.955391Z","iopub.execute_input":"2024-04-21T19:24:12.956601Z","iopub.status.idle":"2024-04-21T19:24:13.175623Z","shell.execute_reply.started":"2024-04-21T19:24:12.956540Z","shell.execute_reply":"2024-04-21T19:24:13.174435Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<font color=#6698FF>We can immediatly drop the 'Id' column as it is unnecessary since we imported the data into a dataframe. ","metadata":{}},{"cell_type":"code","source":"# drop the 'id' axis\nfull_train_data = full_train_data.drop('id', axis=1)","metadata":{"execution":{"iopub.status.busy":"2024-04-21T19:24:13.177003Z","iopub.execute_input":"2024-04-21T19:24:13.177320Z","iopub.status.idle":"2024-04-21T19:24:13.186142Z","shell.execute_reply.started":"2024-04-21T19:24:13.177293Z","shell.execute_reply":"2024-04-21T19:24:13.184892Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### 2.2 Data Exploration","metadata":{}},{"cell_type":"markdown","source":"#### 2.2.1 Data cleanup\n<font color=#6698FF> It is important to check if this dataset consists of multiple datatypes, and if there is any missing data. Let's call the .info() function to see the datatypes this file consists of.","metadata":{}},{"cell_type":"code","source":"# get info\nfull_train_data.info()","metadata":{"execution":{"iopub.status.busy":"2024-04-21T19:24:13.189399Z","iopub.execute_input":"2024-04-21T19:24:13.189803Z","iopub.status.idle":"2024-04-21T19:24:13.243767Z","shell.execute_reply.started":"2024-04-21T19:24:13.189770Z","shell.execute_reply":"2024-04-21T19:24:13.242545Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<font color=#6698FF> We see that there are no missing values, so there is no need to encode the data. Let's check for any missing data.\n    \n","metadata":{}},{"cell_type":"code","source":"# printing the amount of Nan and null entries per column\nprint(full_train_data.isna().sum())\nprint(full_train_data.isnull().sum())","metadata":{"execution":{"iopub.status.busy":"2024-04-21T19:29:51.100954Z","iopub.execute_input":"2024-04-21T19:29:51.102269Z","iopub.status.idle":"2024-04-21T19:29:51.119365Z","shell.execute_reply.started":"2024-04-21T19:29:51.102219Z","shell.execute_reply":"2024-04-21T19:29:51.117498Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<font color=#6698FF> Since there are no 'null' and 'NaN' entries in this dataset, we can continue with the data exploration without issues.","metadata":{}},{"cell_type":"markdown","source":"#### 2.2.2 Correlation\n    \n<font color=#6698FF> Let's explore the features of this dataset. To do this, we plot the correlation heatmap of the data features that have the most correlation.","metadata":{}},{"cell_type":"code","source":"# choosing only features that correlate with |value| >0.35\ncormatrix = full_train_data.corr()\nhighest_cor_features = cormatrix.index[abs(cormatrix[\"class\"])>0.35]\n\n# plotting the correlation matrix\nplt.figure(figsize=(10,10))\ng = sns.heatmap(full_train_data[highest_cor_features].corr(),annot=True,cmap=\"RdYlGn\")","metadata":{"execution":{"iopub.status.busy":"2024-04-21T19:31:22.666990Z","iopub.execute_input":"2024-04-21T19:31:22.668547Z","iopub.status.idle":"2024-04-21T19:31:25.304063Z","shell.execute_reply.started":"2024-04-21T19:31:22.668486Z","shell.execute_reply":"2024-04-21T19:31:25.302800Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<font color=#6698FF> In this correlation matrix we can clearly see that some features correlate very highly with each other, and some highly uncorrelated.\n\n<font color=#6698FF> We can also plot these features in scatterplots.    \n    \n","metadata":{}},{"cell_type":"code","source":"sns.set()\nsns.pairplot(full_train_data[highest_cor_features], height = 2.5)\nplt.show();","metadata":{"execution":{"iopub.status.busy":"2024-04-21T19:24:15.675196Z","iopub.execute_input":"2024-04-21T19:24:15.675617Z","iopub.status.idle":"2024-04-21T19:25:41.034899Z","shell.execute_reply.started":"2024-04-21T19:24:15.675581Z","shell.execute_reply":"2024-04-21T19:25:41.032685Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<font color=#6698FF> Here it is again possible to see that these features correlate very nicely, some linear, some exponential.","metadata":{}},{"cell_type":"markdown","source":"#### 2.2.3 Balanced train-test split \n<font color=#6698FF> Now, let's look at the distribution of the 'score' class, eg the 'has Parkinsons Disease' class. Let's do this by making a pieplot.","metadata":{}},{"cell_type":"code","source":"# amount of positive and negative tested entries\nPD_pos_amount = (full_train_data[\"class\"] == 1).sum()\nPD_neg_amount = (full_train_data[\"class\"] == 0).sum()\n\nPychart = [PD_pos_amount, PD_neg_amount]\nPDlabels = [\"Has Parkinsons Disease\", \" Has no Parkinsons Disease\"]\nPDexplode = [0.02, 0]\n\n# want some pie?\nplt.pie(Pychart, labels = PDlabels, explode = PDexplode)\nplt.show() ","metadata":{"execution":{"iopub.status.busy":"2024-04-21T19:25:41.036877Z","iopub.execute_input":"2024-04-21T19:25:41.037307Z","iopub.status.idle":"2024-04-21T19:25:41.202780Z","shell.execute_reply.started":"2024-04-21T19:25:41.037271Z","shell.execute_reply":"2024-04-21T19:25:41.201174Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<font color=#6698FF> It is possible to see that 75% of the subjects has Parkinsons. This presents a problem, because a model that would always give a positive result would always have a 75% accuracy.","metadata":{}},{"cell_type":"markdown","source":"<font color=#6698FF> There are two ways to try to solve this problem: random under- or random oversampling. Undersampling involves randomly removing samples from those with Parkinson's Disease, but there's a risk of losing valuable information in the process. In oversampling we randomly\nduplicate samples without Parkinson's Disease, but there's a potential risk of overfitting the model.\n    \n<font color=#6698FF> In this notebook, I will train my models on both the undersampling, oversampling and original data.\n    ","metadata":{}},{"cell_type":"code","source":"# Splitting the data into a positive and a negative class\nPD_pos = full_train_data[full_train_data['class'] == 1]\nPD_neg = full_train_data[full_train_data['class'] == 0]\n\n# making the over- and undersampled sets\ntrain_os = pd.concat([PD_pos, PD_neg.sample(PD_pos_amount, replace=True)])\ntrain_us = pd.concat([PD_neg, PD_pos.sample(PD_neg_amount, replace=True)])\n\n# making the piechart of the oversampled set\nPD_pos_over = (train_os[\"class\"] == 1).sum()\nPD_neg_over = (train_os[\"class\"] == 0).sum()\nPychart_os = [PD_pos_over, PD_neg_over]\nlabels_os = [\"Has Parkinsons Disease\", \" Has no Parkinsons Disease\"]\nexplode_os = [0.02, 0]\n\nplt.pie(Pychart_os, labels = labels_os, explode = explode_os)\nplt.title(\"Oversampled & balanced dataset\")\nplt.show() \n\n# making the piechart of the undersampled set\nPD_pos_under = (train_us[\"class\"] == 1).sum()\nPD_neg_under = (train_us[\"class\"] == 0).sum()\nPychart_us = [PD_pos_under, PD_neg_under]\nlabels_us = [\"Has Parkinsons Disease\", \" Has no Parkinsons Disease\"]\nexplode_us = [0.02, 0]\n\nplt.pie(Pychart_us, labels = labels_us, explode = explode_us)\nplt.title(\"Undersampled & balanced dataset\")\nplt.show() ","metadata":{"execution":{"iopub.status.busy":"2024-04-21T19:25:41.205082Z","iopub.execute_input":"2024-04-21T19:25:41.206081Z","iopub.status.idle":"2024-04-21T19:25:41.579686Z","shell.execute_reply.started":"2024-04-21T19:25:41.206020Z","shell.execute_reply":"2024-04-21T19:25:41.577747Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<font color=#6698FF> ","metadata":{}},{"cell_type":"markdown","source":"<font color=#6698FF> Now that our datasets have been made, let's split the features and score classes from each other.","metadata":{}},{"cell_type":"code","source":"# full data set\ny = full_train_data.loc[:,'class']\nX = full_train_data.drop('class', axis=1)\n\n# oversampled data set\ny_os = train_os.loc[:,'class']\nX_os = train_os.drop('class', axis=1)\n\n# undersampled data set\ny_us = train_us.loc[:,'class']\nX_us = train_us.drop('class', axis=1)","metadata":{"execution":{"iopub.status.busy":"2024-04-21T19:25:41.588564Z","iopub.execute_input":"2024-04-21T19:25:41.589259Z","iopub.status.idle":"2024-04-21T19:25:41.619598Z","shell.execute_reply.started":"2024-04-21T19:25:41.589203Z","shell.execute_reply":"2024-04-21T19:25:41.617765Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<font color=#6698FF>In the code below, we split the train data into a test and a train set for each set. I have picked a test_size = 0.3, since this delivers a good model accuracy while not overfitting the data.\n\n\n<font color=#6698FF>If the test size is too small, the model gets overfitted to the training data. This is bad, because the model would not give good results when applied to test data.\n\n \n<font color=#6698FF>If the test size is too big, the model would get trained on a too small amount of data. This would also result in a low score since the model does not have enough data to be accurate.","metadata":{}},{"cell_type":"code","source":"# making full data training and testing sets\nX_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.3)\n\n# making oversampled data training and testing sets\nX_train_os, X_test_os, y_train_os, y_test_os = train_test_split(X_os, y_os, test_size=0.3)\n\n# making undersampled data training and testing sets\nX_train_us, X_test_us, y_train_us, y_test_us = train_test_split(X_us, y_us, test_size=0.3)","metadata":{"execution":{"iopub.status.busy":"2024-04-21T19:25:41.621722Z","iopub.execute_input":"2024-04-21T19:25:41.622108Z","iopub.status.idle":"2024-04-21T19:25:41.648945Z","shell.execute_reply.started":"2024-04-21T19:25:41.622078Z","shell.execute_reply":"2024-04-21T19:25:41.647729Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### 2.2.4 Preprocessing\n<font color=#6698FF> The preprocessing for this dataset is very easy, since there are no missing value and everything is in numerical format. We can use the preprocessing function from sklearn for this.\n","metadata":{}},{"cell_type":"code","source":"# scaling the full data set\nscaler = preprocessing.StandardScaler().fit(X_train)\nX_train = scaler.transform(X_train)\nX_test = scaler.transform(X_test)\n\n# scaling the oversampled data set\nscaler_over = preprocessing.StandardScaler().fit(X_train_os)\nX_train_os = scaler_over.transform(X_train_os)\nX_test_os = scaler_over.transform(X_test_os)\n\n# scaling the undersampled data set\nscaler_under = preprocessing.StandardScaler().fit(X_train_us)\nX_train_us = scaler_under.transform(X_train_us)\nX_test_us = scaler_under.transform(X_test_us)","metadata":{"execution":{"iopub.status.busy":"2024-04-21T19:25:41.650270Z","iopub.execute_input":"2024-04-21T19:25:41.650616Z","iopub.status.idle":"2024-04-21T19:25:41.855794Z","shell.execute_reply.started":"2024-04-21T19:25:41.650585Z","shell.execute_reply":"2024-04-21T19:25:41.854662Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"\n## 3. Training and Results\n<font color=#6698FF> In this notebook, I will use linear and neural network models.\n","metadata":{}},{"cell_type":"markdown","source":"### 3.1 Linear Models\n<font color=#6698FF>For the linear models, i use logistic regression to train the model.\n    \n<font color=#6698FF>One advantage of logistic regression is its efficienty, it tends to be computationally efficient and can handle a large number of features without requiring a lot of CPU power.\n    \n<font color=#6698FF>One drawback of logistic regression is that it assumes a linear relationship between the independent variables and the log-odds of the dependent variable. This means it may not perform well when the relationship is non-linear.    ","metadata":{}},{"cell_type":"markdown","source":"#### 3.1.1 Full dataset\n<font color=#6698FF> First, i will make a linear model of the full dataset.","metadata":{}},{"cell_type":"code","source":"model = LogisticRegression()\n\n# fit the model\nmodel.fit(X_train, y_train)\n\n# predict\npredictions=model.predict(X_test)\n\n# accuracy\nscore=model.score(X_test, y_test)\n\n# confusion matrix\ncm = metrics.confusion_matrix(y_test, predictions)\nplt.figure(figsize=(4,4))\nsns.heatmap(cm, annot=True, fmt=\".3f\", linewidths=.5, square = True, cmap = 'Blues_r');\nplt.ylabel('Actual label');\nplt.xlabel('Predicted label');\nall_sample_title = 'Balanced accuracy Score: {0}'.format(round(metrics.balanced_accuracy_score(y_test,predictions),3))\nplt.title(all_sample_title, size = 10);","metadata":{"execution":{"iopub.status.busy":"2024-04-21T19:26:23.944983Z","iopub.execute_input":"2024-04-21T19:26:23.946065Z","iopub.status.idle":"2024-04-21T19:26:24.585839Z","shell.execute_reply.started":"2024-04-21T19:26:23.946015Z","shell.execute_reply":"2024-04-21T19:26:24.584416Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<font color=#6698FF>This model gives a bad score, in comparison with just guessing '1' all the time. It is even .034% worse than guessing. Thus this is a bad way to use the data given to us.","metadata":{}},{"cell_type":"markdown","source":"#### 3.1.2 Undersampled dataset\n<font color=#6698FF>Now, i do the same again for the undersampled dataset.","metadata":{}},{"cell_type":"code","source":"# define the model\nmodel = LogisticRegression()\n# fit the model\nmodel.fit(X_train_us, y_train_us)\n\n# predict the output \npredictions_us=model.predict(X_test_us)\n\n# calculate the accuracy of the model\nscore_us=model.score(X_test_us, y_test_us)\n\n# confusion matrix\ncm = metrics.confusion_matrix(y_test_us, predictions_us)\nplt.figure(figsize=(4,4))\nsns.heatmap(cm, annot=True, fmt=\".3f\", linewidths=.5, square = True, cmap = 'Blues_r');\nplt.ylabel('Actual label');\nplt.xlabel('Predicted label');\nall_sample_title = 'Balanced accuracy Score: {0}'.format(round(metrics.balanced_accuracy_score(y_test_us,predictions_us),3))\nplt.title(all_sample_title, size = 10);","metadata":{"execution":{"iopub.status.busy":"2024-04-21T19:26:24.588348Z","iopub.execute_input":"2024-04-21T19:26:24.588847Z","iopub.status.idle":"2024-04-21T19:26:25.158375Z","shell.execute_reply.started":"2024-04-21T19:26:24.588779Z","shell.execute_reply":"2024-04-21T19:26:25.157232Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<font color=#6698FF> This is already better than training the model using the full data, but not by much. This is still not a great score.","metadata":{}},{"cell_type":"markdown","source":"#### 3.1.3 Oversampled dataset\n<font color=#6698FF>We again do the same but now for the oversampled dataset.","metadata":{}},{"cell_type":"code","source":"# define the model\nmodel = LogisticRegression()\n# fit the model\nmodel.fit(X_train_os, y_train_os)\n\n# predict the output \npredictions_os=model.predict(X_test_os)\n\n# calculate the accuracy of the model\nscore_os=model.score(X_test_os, y_test_os)\n\n# confusion matrix\ncm = metrics.confusion_matrix(y_test_os, predictions_os)\nplt.figure(figsize=(4,4))\nsns.heatmap(cm, annot=True, fmt=\".3f\", linewidths=.5, square = True, cmap = 'Blues_r');\nplt.ylabel('Actual label');\nplt.xlabel('Predicted label');\nall_sample_title = 'Balanced accuracy Score: {0}'.format(round(metrics.balanced_accuracy_score(y_test_os,predictions_os),3))\nplt.title(all_sample_title, size = 10);","metadata":{"execution":{"iopub.status.busy":"2024-04-21T19:26:25.159681Z","iopub.execute_input":"2024-04-21T19:26:25.160018Z","iopub.status.idle":"2024-04-21T19:26:25.724917Z","shell.execute_reply.started":"2024-04-21T19:26:25.159990Z","shell.execute_reply":"2024-04-21T19:26:25.723302Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<font color=#6698FF>This is a clear improvement! Using the oversampled dataset results in a large increase of balanced accuracy, and would thus be the best dataset to use for logistic regression.","metadata":{}},{"cell_type":"markdown","source":"### 3.2 Neural Network\n<font color=#6698FF> I now make neural networks using the datasets as defined before.\n    \n<font color=#6698FF>One advantage of using a neural network is that it can easily detect patterns in data and it makes it highly effective for pattern recognition.\n    \n<font color=#6698FF>One drawback of using a neural network is that it has a 'black-box nature'. It is not clear which features contribute to the model output, and this can be a drawback if interpretability is important.","metadata":{}},{"cell_type":"markdown","source":"#### 3.2.1 Full dataset\n<font color=#6698FF> I start again with the full dataset. First, lets make the neural network.","metadata":{}},{"cell_type":"code","source":"# Create simple Neural Network model\ninput_nodes = X_train.shape[1] # amount of features\nhidden_layer_1_nodes = 40\nhidden_layer_2_nodes = 20\noutput_layer = 1\n\n# initializing a sequential model\nfull_model = Sequential()\n\n# adding layers\nfull_model.add(Dense(hidden_layer_1_nodes,input_dim=input_nodes , activation='relu'))\nfull_model.add(Dropout(0.1))\nfull_model.add(Dense(hidden_layer_2_nodes, activation='relu'))\nfull_model.add(Dropout(0.1))\nfull_model.add(Dense(output_layer, activation='sigmoid'))\n\n# Compiling the ANN\nfull_model.compile(optimizer='adam', loss='binary_crossentropy', metrics=['accuracy'])\n\nhistory = full_model.fit(X_train,y_train,validation_data=(X_test,y_test), epochs=15, batch_size=16, verbose=2)","metadata":{"execution":{"iopub.status.busy":"2024-04-21T19:26:25.728518Z","iopub.execute_input":"2024-04-21T19:26:25.729316Z","iopub.status.idle":"2024-04-21T19:26:29.720877Z","shell.execute_reply.started":"2024-04-21T19:26:25.729262Z","shell.execute_reply":"2024-04-21T19:26:29.719540Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<font color=#6698FF> Now, we print the model accuracy and validation accuracy.","metadata":{}},{"cell_type":"code","source":"# printing\nscore = full_model.evaluate(X_test, y_test, verbose=0)\nprint(f'Test loss: {score[0]} / Test accuracy: {score[1]}')\n\n# plotting\nplt.plot(history.history['accuracy'])\nplt.plot(history.history['val_accuracy'])\nplt.legend(['accuracy', 'val_accuracy'])","metadata":{"execution":{"iopub.status.busy":"2024-04-21T19:26:29.722395Z","iopub.execute_input":"2024-04-21T19:26:29.722782Z","iopub.status.idle":"2024-04-21T19:26:30.275787Z","shell.execute_reply.started":"2024-04-21T19:26:29.722750Z","shell.execute_reply":"2024-04-21T19:26:30.274423Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<font color=#6698FF> From these results it is possible to see that the model is accurate, but not more than the linear oversampled model.","metadata":{}},{"cell_type":"markdown","source":"#### 3.2.2 Undersampled dataset\n<font color=#6698FF> We do the same for the undersampled dataset.","metadata":{}},{"cell_type":"code","source":"# Create simple Neural Network model\ninput_nodes = X_train_us.shape[1] # amount of features\nhidden_layer_1_nodes = 40\nhidden_layer_2_nodes = 20\noutput_layer = 1\n\n# initializing a sequential model\nfull_model = Sequential()\n\n# adding layers\nfull_model.add(Dense(hidden_layer_1_nodes,input_dim=input_nodes , activation='relu'))\nfull_model.add(Dropout(0.1))\nfull_model.add(Dense(hidden_layer_2_nodes, activation='relu'))\nfull_model.add(Dropout(0.1))\nfull_model.add(Dense(output_layer, activation='sigmoid'))\n\n# Compiling the ANN\nfull_model.compile(optimizer='adam', loss='binary_crossentropy', metrics=['accuracy'])\n\nhistory = full_model.fit(X_train_us,y_train_us,validation_data=(X_test_us,y_test_us), epochs=15, batch_size=16, verbose=2)\n\n# printing\nscore = full_model.evaluate(X_test_us, y_test_us, verbose=0)\nprint(f'Test loss: {score[0]} / Test accuracy: {score[1]}')\n\n# plotting\nplt.plot(history.history['accuracy'])\nplt.plot(history.history['val_accuracy'])\nplt.legend(['accuracy', 'val_accuracy'])","metadata":{"execution":{"iopub.status.busy":"2024-04-21T19:26:30.277829Z","iopub.execute_input":"2024-04-21T19:26:30.278244Z","iopub.status.idle":"2024-04-21T19:26:33.904227Z","shell.execute_reply.started":"2024-04-21T19:26:30.278195Z","shell.execute_reply":"2024-04-21T19:26:33.902967Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<font color=#6698FF> Training the neural network on the undersampled dataset results in an worse model result than training it on the full dataset. The accuracy even approaches a value similar to its linear counterpart which was trained on the same data.","metadata":{}},{"cell_type":"markdown","source":"#### 3.2.3 Oversampled dataset\n<font color=#6698FF> And again for the oversampled dataset.","metadata":{}},{"cell_type":"code","source":"# Create simple Neural Network model\ninput_nodes = X_train_os.shape[1] # amount of features\nhidden_layer_1_nodes = 40\nhidden_layer_2_nodes = 20\noutput_layer = 1\n\n# initializing a sequential model\nfull_model = Sequential()\n\n# adding layers\nfull_model.add(Dense(hidden_layer_1_nodes,input_dim=input_nodes , activation='relu'))\nfull_model.add(Dropout(0.1))\nfull_model.add(Dense(hidden_layer_2_nodes, activation='relu'))\nfull_model.add(Dropout(0.1))\nfull_model.add(Dense(output_layer, activation='sigmoid'))\n\n# Compiling the ANN\nfull_model.compile(optimizer='adam', loss='binary_crossentropy', metrics=['accuracy'])\n\nhistory = full_model.fit(X_train_os, y_train_os,validation_data=(X_test_os,y_test_os), epochs=15, batch_size=16, verbose=2)\n\n# printing\nscore = full_model.evaluate(X_test_os, y_test_os, verbose=0)\nprint(f'Test loss: {score[0]} / Test accuracy: {score[1]}')\n\n# plotting\nplt.plot(history.history['accuracy'])\nplt.plot(history.history['val_accuracy'])\nplt.legend(['accuracy', 'val_accuracy'])","metadata":{"execution":{"iopub.status.busy":"2024-04-21T19:26:33.905905Z","iopub.execute_input":"2024-04-21T19:26:33.906237Z","iopub.status.idle":"2024-04-21T19:26:38.903137Z","shell.execute_reply.started":"2024-04-21T19:26:33.906207Z","shell.execute_reply":"2024-04-21T19:26:38.899100Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<font color=#6698FF> Training the network on the oversampled data results in the best model accuracy yet! This is in line with what we would expect, as its linear model counterpart was also the best of its kind.","metadata":{}},{"cell_type":"markdown","source":"## 4. Discussion and Conclusion\n<font color=#6698FF>For the Parkinsons Disease classification dataset, it was not necessary to do a lot of data cleaning, as it already consisted of strictly numerical data without any missing entries. Only the 'id' column needed to be removed.\n\n<font color=#6698FF>I used two models, one linear and one neural network. I tested them using three different datasets, one with the fully unchanged data, one with overbalanced data and one with underbalanced data.\n \n<font color=#6698FF>For this application, a neural network with oversampled data would work best, as it outpreforms the linear models and its neural counterparts every time.\n  \n\n    \n<font color=#6698FF>One of the shortcomings in my approach is that i did not use a feature reducing application. In the correlation matrix, it is possible to see clear grouping of correlations with about the same values. This means I could have possibly been able to reduce the features without losing data, and this would result in a better scalability for large datasets.\n    \n<font color=#6698FF>However, this is more useful in the linear model approach. The neural networks will do feature reduction automatically using a specified amount of nodes. The amount of nodes could possibly be tuned to the amount of these groupings.","metadata":{}},{"cell_type":"markdown","source":"## 5. References\n\n(1) Sakar, C.O., Serbes, G., Gunduz, A., Tunc, H.C., Nizam, H., Sakar, B.E., Tutuncu, M., Aydin, T., Isenkul, M.E. and Apaydin, H., 2018. A comparative analysis of speech signal processing algorithms for Parkinson's disease classification and the use of the tunable Q-factor wavelet transform. Applied Soft Computing, DOI: [Web Link] https://doi.org/10.1016/j.asoc.2018.10.022\n\n(2) https://www.analyticsvidhya.com/blog/2020/07/10-techniques-to-deal-with-class-imbalance-in-machine-learning/\n    \n(3) https://www.kaggle.com/code/pradeep023/parkinson-s-disease-classification","metadata":{}}]}