{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"## 1. Introduction\nName: Alec Daalman\\\nUsername: AlecDaalman\\\nScore: 2.0353\\\nLeaderboard rank: None","metadata":{}},{"cell_type":"markdown","source":"## 2. Data","metadata":{}},{"cell_type":"markdown","source":"### 2.1 Dataset\nIn this section the dataset and libraries are loaded.","metadata":{}},{"cell_type":"code","source":"import os\nimport numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\nimport matplotlib.pyplot as plt # plotting\nfrom sklearn.preprocessing import StandardScaler # normalizing the features\nfrom sklearn.model_selection import train_test_split # splitting the features\nfrom sklearn import linear_model # linear regression model\nfrom sklearn import metrics # accuracy calculation\nfrom keras.models import Sequential # neural network model\nfrom keras.layers import Dense # add NN layers\nfrom keras.layers import Dropout # prevent overfitting\nfrom sklearn.ensemble import RandomForestRegressor # random forests model\n","metadata":{"execution":{"iopub.status.busy":"2023-04-27T10:52:40.304935Z","iopub.execute_input":"2023-04-27T10:52:40.307033Z","iopub.status.idle":"2023-04-27T10:52:51.765359Z","shell.execute_reply.started":"2023-04-27T10:52:40.306942Z","shell.execute_reply":"2023-04-27T10:52:51.763817Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(os.listdir(\"../input\"))\n\ntrain = pd.read_csv('../input/LANL-Earthquake-Prediction/train.csv', dtype={'acoustic_data': np.int16, 'time_to_failure': np.float64})","metadata":{"execution":{"iopub.status.busy":"2023-04-27T10:52:51.767779Z","iopub.execute_input":"2023-04-27T10:52:51.768528Z","iopub.status.idle":"2023-04-27T10:56:42.202871Z","shell.execute_reply.started":"2023-04-27T10:52:51.768488Z","shell.execute_reply":"2023-04-27T10:56:42.201565Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### 2.2 Data Exploration\n\nWhen plotting the data, it can be observed that the deviation of the acoustic data significantly increases (positive and negative) before the the time to failure reaches 0. There are also no outliers or missing data, that need to be removed.\n\nBecause the to be predicted data is time, linear models need to be used. To assess the performance of the models, the Mean Absolute Error is used. This is because it is much easier to interpred the error of the regression model, as it is just the difference between the test ouput and the prediction in milliseconds. (https://www.kaggle.com/competitions/LANL-Earthquake-Prediction/discussion/77525, in the comments)","metadata":{}},{"cell_type":"code","source":"fig, ax = plt.subplots(3,1, figsize=(20,12))\nax[0].scatter(train.index[:10000000],train.acoustic_data[:10000000], c=\"darkred\")\nax[0].set_title(\"Acoustic data of 10 Mio rows\")\nax[0].set_xlabel(\"Index\")\nax[0].set_ylabel(\"Acoustic signal\");\n\nax[1].scatter(train.index[:10000000],train.time_to_failure[:10000000], c=\"darkred\")\nax[1].set_title(\"Quaketime of 10 Mio rows\")\nax[1].set_xlabel(\"Index\")\nax[1].set_ylabel(\"Quaketime in ms\");\n\nax[2].scatter(train.acoustic_data[:10000000],train.time_to_failure[:10000000], c=\"darkred\")\nax[2].set_title(\"Acoustic data vs Quaketime of 10 Mio rows\")\nax[2].set_xlabel(\"Acoustic signal\")\nax[2].set_ylabel(\"Quaketime in ms\");\n\ntrain.isna().sum()\n","metadata":{"execution":{"iopub.status.busy":"2023-04-27T10:56:42.204629Z","iopub.execute_input":"2023-04-27T10:56:42.205034Z","iopub.status.idle":"2023-04-27T10:58:53.851227Z","shell.execute_reply.started":"2023-04-27T10:56:42.204971Z","shell.execute_reply":"2023-04-27T10:58:53.849814Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### 2.3 Data Preparation\n\n<font color=#6698FF>  \nNote that you are only allowed to use training data for preprocessing but you then need to perform similar changes on test data too.\nYou can use [pipelining](https://scikit-learn.org/stable/modules/generated/sklearn.pipeline.Pipeline.html) to help with the preprocessing.</font>","metadata":{}},{"cell_type":"markdown","source":"#### 2.3.1 Feature extraction\nThe features are extracted from time and FFT characteristics of the data. to reduce the amount of data and create better features, the data is segmented into 4194 segments, where every piece has 13 features. Because the data is segmented, the time to failure (TTF) of the corresponding segment is the last TTF in the segment.\n\nThe features are based on: https://www.kaggle.com/code/artgor/earthquakes-fe-more-features-and-samples. The roll mean 10 was chosen specifically because it had the highest importance with their KFold model. The amount of segments are based on: https://www.kaggle.com/code/allunia/shaking-earth.\n","metadata":{}},{"cell_type":"code","source":"rows = 150_000\nsegments = int(np.floor(train.shape[0] / rows))\n\nx = pd.DataFrame(index=range(segments), dtype=np.float64)\ny = pd.DataFrame(index=range(segments), dtype=np.float64,\n                       columns=['time_to_failure'])\n\n#Create features with segmented data\nfor segment in range(segments):\n    samples = train.acoustic_data[segment*rows:(segment+1)*rows]\n    samples_fft = np.fft.fft(samples)\n    real_fft = np.real(samples_fft)\n    imag_fft = np.imag(samples_fft)\n    del samples_fft\n    \n    #Time features\n    x.loc[segment, 'average'] = np.mean(samples)\n    x.loc[segment, 'std'] = np.std(samples)\n    x.loc[segment, 'maximum'] = np.max(samples)\n    x.loc[segment, 'minimum'] = np.min(samples)\n    \n    x_roll_mean = samples.rolling(10).mean().dropna().values\n    x.loc[segment, 'q05_roll_mean_10'] = np.quantile(x_roll_mean, 0.05)\n    \n    #FFT features\n    x.loc[segment, 'average_FFTR'] = real_fft.mean()\n    x.loc[segment, 'std_FFTR'] = real_fft.std()\n    x.loc[segment, 'maximum_FFTR'] = real_fft.max()\n    x.loc[segment, 'minimum_FFTR'] = real_fft.min()\n    \n    x.loc[segment, 'average_FFTI'] = imag_fft.mean()\n    x.loc[segment, 'std_FFTI'] = imag_fft.std()\n    x.loc[segment, 'maximum_FFTI'] = imag_fft.max()\n    x.loc[segment, 'minimum_FFTI'] = imag_fft.min()\n    \n    #Output\n    y.loc[segment, 'time_to_failure'] = train.time_to_failure[(segment+1)*rows]\n    \n    del samples, real_fft, imag_fft # delete segmented data to save RAM\n\nx.to_csv('extracted_features.csv',index=False) #save features in CSV format\ny.to_csv('extracted_outputs.csv',index=False) #save TTF in CSV format\ndel train # delete train to save RAM\n    ","metadata":{"execution":{"iopub.status.busy":"2023-04-27T10:58:53.853882Z","iopub.execute_input":"2023-04-27T10:58:53.855418Z","iopub.status.idle":"2023-04-27T11:00:11.072291Z","shell.execute_reply.started":"2023-04-27T10:58:53.855371Z","shell.execute_reply":"2023-04-27T11:00:11.070808Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### 2.3.2 Normalization\nThe features are normalized to prevent bias and overfitting from certain features.","metadata":{}},{"cell_type":"code","source":"scaler = StandardScaler().fit(x)\nx = scaler.transform(x) # Normalize features","metadata":{"execution":{"iopub.status.busy":"2023-04-27T11:00:11.077275Z","iopub.execute_input":"2023-04-27T11:00:11.077841Z","iopub.status.idle":"2023-04-27T11:00:11.093929Z","shell.execute_reply.started":"2023-04-27T11:00:11.077797Z","shell.execute_reply":"2023-04-27T11:00:11.092053Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### 2.3.4 Train-test split\nThe features and time to failures are shuffeled and spilt into train (80%) and test (20%) data. The random state is set to ensure a reproducable train and test dataset when rerunning.","metadata":{}},{"cell_type":"code","source":"X_train, X_test, y_train, y_test = train_test_split(x, y, test_size=0.2,shuffle=True, random_state=102)","metadata":{"execution":{"iopub.status.busy":"2023-04-27T11:00:11.096219Z","iopub.execute_input":"2023-04-27T11:00:11.096692Z","iopub.status.idle":"2023-04-27T11:00:11.110776Z","shell.execute_reply.started":"2023-04-27T11:00:11.096651Z","shell.execute_reply":"2023-04-27T11:00:11.108969Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"\n## 3. Training and Results\n<font color=#6698FF> Briefly introduce algorithms you choose. \nPresent your balanced accuracies for both test and training data for all classifiers. Analyse the performance on test and training in terms of bias and variance. Give one advantage and one drawback of the method you use.\n<\\font>\n\n","metadata":{}},{"cell_type":"markdown","source":"### 3.1 Linear model\nThe first model is a linear regression model. Linear regression is a simple, linear model that tries to find the best-fit line through the data. It assumes a linear relationship between the input variables and the output variable. Linear regression is easy to interpret and computationally efficient, making it a good choice for simple regression problems with a small number of input variables.\n\nLinear regression tends to have low variance but high bias. This means that it may underfit the training data and generalize poorly to new data. The model may not capture complex nonlinear relationships between input and output variables.","metadata":{}},{"cell_type":"markdown","source":"#### 3.1.1 Model","metadata":{}},{"cell_type":"code","source":"\nmodel_LinReg = linear_model.LinearRegression()\n\nmodel_LinReg.fit(X_train, y_train.time_to_failure)","metadata":{"execution":{"iopub.status.busy":"2023-04-27T11:00:11.112811Z","iopub.execute_input":"2023-04-27T11:00:11.113607Z","iopub.status.idle":"2023-04-27T11:00:11.140174Z","shell.execute_reply.started":"2023-04-27T11:00:11.113563Z","shell.execute_reply":"2023-04-27T11:00:11.138592Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### 3.1.2 Accuracy","metadata":{}},{"cell_type":"code","source":"y_pred_val_LinReg = model_LinReg.predict(X_test)\naccuracy_LinReg = metrics.mean_absolute_error(y_test.time_to_failure, y_pred_val_LinReg,multioutput='raw_values')\n\nprint(accuracy_LinReg)","metadata":{"execution":{"iopub.status.busy":"2023-04-27T11:00:11.141990Z","iopub.execute_input":"2023-04-27T11:00:11.143185Z","iopub.status.idle":"2023-04-27T11:00:11.158822Z","shell.execute_reply.started":"2023-04-27T11:00:11.143095Z","shell.execute_reply":"2023-04-27T11:00:11.155607Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### 3.2 Neural Network Regression\nNeural networks are a type of machine learning algorithm inspired by the structure and function of the human brain. They consist of layers of interconnected nodes that perform complex calculations to transform the input data into the desired output. Neural networks can capture nonlinear relationships between the input and output variables and can handle large and complex datasets.\n\nNeural networks tend to have low bias but high variance. This means that they may overfit the training data and generalize poorly to new data. A NN can be computationally expensive to train and may require a large amount of data to prevent overfitting.","metadata":{}},{"cell_type":"markdown","source":"#### 3.2.1 Model","metadata":{}},{"cell_type":"code","source":"# Create simple Neural Network model\ninput_nodes = 13  # total number of features\nhidden_layer_1_nodes = 10\noutput_layer = 1\n\n# initializing a sequential model\nmodel_NNR = Sequential()\n\n# adding layers\nmodel_NNR.add(Dense(hidden_layer_1_nodes, activation='relu',input_shape=(input_nodes,)))\nmodel_NNR.add(Dropout(0.1))\nmodel_NNR.add(Dense(output_layer, activation='linear'))\n\nmodel_NNR.summary()\n\nmodel_NNR.compile(loss='mean_absolute_error',\n              optimizer='sgd',\n              metrics=['mean_absolute_error'])","metadata":{"execution":{"iopub.status.busy":"2023-04-27T11:00:11.165481Z","iopub.execute_input":"2023-04-27T11:00:11.169306Z","iopub.status.idle":"2023-04-27T11:00:12.007311Z","shell.execute_reply.started":"2023-04-27T11:00:11.169184Z","shell.execute_reply":"2023-04-27T11:00:12.005727Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### 3.2.2 Accuracy","metadata":{}},{"cell_type":"code","source":"history = model_NNR.fit(X_train,y_train.time_to_failure,validation_data=(X_test,y_test.time_to_failure), epochs=100, batch_size=32, verbose=0);\n\nplt.plot(history.history['loss'])\nplt.plot(history.history['val_loss'])\nplt.legend(['loss', 'val_loss'])\n\ny_pred_val_NNR = model_NNR.predict(X_test)\naccuracy_NNR = metrics.mean_absolute_error(y_test.time_to_failure, y_pred_val_NNR,multioutput='raw_values')\nprint(accuracy_NNR)","metadata":{"execution":{"iopub.status.busy":"2023-04-27T11:00:12.008465Z","iopub.execute_input":"2023-04-27T11:00:12.009153Z","iopub.status.idle":"2023-04-27T11:00:38.970708Z","shell.execute_reply.started":"2023-04-27T11:00:12.009089Z","shell.execute_reply":"2023-04-27T11:00:38.968863Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### 3.3 Random Forests\nRandom forests are an ensemble learning method that combines multiple decision trees to make predictions. Each decision tree is trained on a subset of the data and a subset of the input features, and the final prediction is made by aggregating the predictions of all the trees. Random forests can handle both categorical and continuous input variables and can capture nonlinear relationships.\n\nRandom forests also tend to have low bias but high variance. Which means that they may overfit the training data and generalize poorly to new data. Random forests can be computationally expensive and may require careful selection of hyperparameters and feature selection techniques.","metadata":{}},{"cell_type":"markdown","source":"#### 3.3.1 Model","metadata":{}},{"cell_type":"code","source":"# define the model\nmodel_RFC = RandomForestRegressor(max_depth=5, random_state=102)\n\n# fit/train the model on all features\nmodel_RFC.fit(X_train, y_train.time_to_failure)","metadata":{"execution":{"iopub.status.busy":"2023-04-27T11:00:38.972998Z","iopub.execute_input":"2023-04-27T11:00:38.973406Z","iopub.status.idle":"2023-04-27T11:00:40.297272Z","shell.execute_reply.started":"2023-04-27T11:00:38.973367Z","shell.execute_reply":"2023-04-27T11:00:40.295097Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### 3.3.2 Feature importance","metadata":{}},{"cell_type":"code","source":"#get feature importance\nimportance = model_RFC.feature_importances_\n\n#create a dictionary with key=indices, and values=importance\nimportant_features_dict = {}\nfor idx, val in enumerate(importance):\n    important_features_dict[idx] = val\n    \n# sorting \nimportant_features_list = sorted(important_features_dict,\n                                 key=important_features_dict.get,\n                                 reverse=True)\n# print indices of all features\nimportant_features_list[:]","metadata":{"execution":{"iopub.status.busy":"2023-04-27T11:00:40.298807Z","iopub.execute_input":"2023-04-27T11:00:40.299244Z","iopub.status.idle":"2023-04-27T11:00:40.321122Z","shell.execute_reply.started":"2023-04-27T11:00:40.299205Z","shell.execute_reply":"2023-04-27T11:00:40.319750Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### 3.3.3 Accuracy","metadata":{}},{"cell_type":"code","source":"y_pred_val_RFC = model_RFC.predict(X_test)\naccuracy_val_RFC = metrics.mean_absolute_error(y_test.time_to_failure, y_pred_val_RFC,multioutput='raw_values')\n\nprint(accuracy_val_RFC)","metadata":{"execution":{"iopub.status.busy":"2023-04-27T11:00:40.322811Z","iopub.execute_input":"2023-04-27T11:00:40.323177Z","iopub.status.idle":"2023-04-27T11:00:40.351070Z","shell.execute_reply.started":"2023-04-27T11:00:40.323144Z","shell.execute_reply":"2023-04-27T11:00:40.349676Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## 4. Discussion and Conclusion\n\n<font color=#6698FF>Discuss all the choices you made during the process and your final conclusions.  Highlight the strong points of your approach, discuss its shortcomings  and suggest some future approaches that may improve it. Please be self critical here. The assignment is not about achieving a state of the art performance, but about showing what you have learned the concepts during the course.</font>","metadata":{}},{"cell_type":"markdown","source":"### 4.1 Discussion \nIn this section the design choices and predictions are highlighted.\n","metadata":{}},{"cell_type":"markdown","source":"#### 4.1.1 Choices\n\nFeatures: \nFor the features, some generic time and FFT features are generated. To boost the score, one specific feature is generated. The features are not explained using PCA, because after trial and error, this only resulted in a lower accuracy.\n\nLinear regression:\nThe linear regression was chosen an easy and straightforward benchmark. No special design choices were made for this model.\n\nNeural network:\nIn the case of the neural network one hidden layer is chosen, because most neural networks don't increase accuracy much with such little and relatively simple features. As a rule of thumb the amount of neurons in the hidden layer is in between the output size and the input size. When experimenting with more the accuracy did not increase, so the average performed well enough. To prevent overfitting a dropout of 0.1 was set and because we are predicting a continuous value, a linear output activation is used.\n\nRandom forests:\nThe random forests model has a max depth of 5, to get a better accuracy without overfitting the data. ","metadata":{}},{"cell_type":"markdown","source":"#### 4.1.2 Plot of predictions vs test data\nWhen plotting the predicted data of all the models, it can be observed that the linear regression has negative TTF. This means that while the MAE of the model is relatively good compared to the other models, the actual precitions have a large error.","metadata":{}},{"cell_type":"code","source":"y_pred_val_LinReg_plot = pd.DataFrame(y_pred_val_LinReg,columns=['time_to_failure'], index=y_test.index).sort_index(ascending=True)\ny_pred_val_NNR_plot = pd.DataFrame(y_pred_val_NNR,columns=['time_to_failure'], index=y_test.index).sort_index(ascending=True)\ny_pred_val_RFC_plot = pd.DataFrame(y_pred_val_RFC,columns=['time_to_failure'], index=y_test.index).sort_index(ascending=True)\ny_test_plot = y_test.sort_index(ascending=True)\n\nfig, ax = plt.subplots(3,1, figsize=(20,12))\n\nax[0].plot(y_test_plot.index, y_test_plot.time_to_failure, c=\"darkred\")\nax[0].plot(y_pred_val_LinReg_plot.index,y_pred_val_LinReg_plot, c=\"blue\")\nax[0].set_title(\"Linear regression Time to failure compared\")\nax[0].set_xlabel(\"Index\")\nax[0].set_ylabel(\"Quaketime in ms\");\n\nax[1].plot(y_test_plot.index, y_test_plot.time_to_failure, c=\"darkred\")\nax[1].plot(y_pred_val_NNR_plot.index,y_pred_val_NNR_plot, c=\"green\")\nax[1].set_title(\"Neural network Time to failure compared\")\nax[1].set_xlabel(\"Index\")\nax[1].set_ylabel(\"Quaketime in ms\");\n\nax[2].plot(y_test_plot.index, y_test_plot.time_to_failure, c=\"darkred\")\nax[2].plot(y_pred_val_RFC_plot.index,y_pred_val_RFC_plot, c=\"blue\")\nax[2].set_title(\"Random forests Time to failure compared\")\nax[2].set_xlabel(\"Index\")\nax[2].set_ylabel(\"Quaketime in ms\");","metadata":{"execution":{"iopub.status.busy":"2023-04-27T11:00:40.355812Z","iopub.execute_input":"2023-04-27T11:00:40.356226Z","iopub.status.idle":"2023-04-27T11:00:41.109624Z","shell.execute_reply.started":"2023-04-27T11:00:40.356188Z","shell.execute_reply":"2023-04-27T11:00:41.108270Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### 4.2 Conclusion and future approaches\nWhen comparing the models based on predictions and accuracy on the test data, the neural network (while the most complex) results in the best performance when predicting the time to failure of an earthquake. This is however with the limited features and optimization provided. Compared with other (substantially more complicated) methods, the models don't perform really well.  \n\nTo increase the accuracy of the test data, the best way is to create more and better features. Higher quality features can describe patterns in the data better and thus improve the training of  the model to recoginze these patterns.","metadata":{}}]}