{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"code","source":"# This Python 3 environment comes with many helpful analytics libraries installed\n# It is defined by the kaggle/python Docker image: https://github.com/kaggle/docker-python\n# For example, here's several helpful packages to load\n\nimport numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\n\n# Input data files are available in the read-only \"../input/\" directory\n# For example, running this (by clicking run or pressing Shift+Enter) will list all files under the input directory\n\nimport os\nfor dirname, _, filenames in os.walk('/kaggle/input'):\n    for filename in filenames:\n        print(os.path.join(dirname, filename))\n\n# You can write up to 20GB to the current directory (/kaggle/working/) that gets preserved as output when you create a version using \"Save & Run All\" \n# You can also write temporary files to /kaggle/temp/, but they won't be saved outside of the current session\n\n\nimport matplotlib.pyplot as plt\n%matplotlib inline","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","execution":{"iopub.status.busy":"2023-04-28T19:04:04.889455Z","iopub.execute_input":"2023-04-28T19:04:04.890667Z","iopub.status.idle":"2023-04-28T19:04:07.316038Z","shell.execute_reply.started":"2023-04-28T19:04:04.890611Z","shell.execute_reply":"2023-04-28T19:04:07.314680Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<font color=#6698FF> Please write your name, username in Kaggle, score and Leaderboard rank here.\nThijs Geschiere, tjgeschiere, Eliott Schuurman, henk gertman","metadata":{}},{"cell_type":"markdown","source":"### 2.1 Dataset\n<font color=#6698FF>In this section, we load and explore the dataset.","metadata":{}},{"cell_type":"code","source":"#Loadingn all samples from the dataset \ntrain = pd.read_csv(\"../input/LANL-Earthquake-Prediction/train.csv\",\n                    dtype={'acoustic_data': np.int16, 'time_to_failure': np.float64})\n\n#Renaming the dataset to more intuitive names as done in Laura Fink's Notebook: https://www.kaggle.com/code/allunia/shaking-earth?kernelSessionId=10231001\n\ntrain.rename({\"acoustic_data\": \"signal\", \"time_to_failure\": \"quaketime\"}, axis=\"columns\", inplace=True)\nprint(train.quaketime.values.shape)\nprint(train.shape)\ntrain.head(5)\n","metadata":{"execution":{"iopub.status.busy":"2023-04-28T19:04:07.318051Z","iopub.execute_input":"2023-04-28T19:04:07.319099Z","iopub.status.idle":"2023-04-28T19:07:57.409448Z","shell.execute_reply.started":"2023-04-28T19:04:07.319046Z","shell.execute_reply":"2023-04-28T19:07:57.407834Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"\n\n","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#Plotting a part of the Dataset\nfig, ax = plt.subplots(2,1, figsize=(20,12))\nfig.suptitle(\"One Percent of the Given Data Plotted\", fontsize = 20)\nax[0].plot(train.index.values[::100], train.quaketime.values[::100], c=\"darkred\") #stepsize 100 gives 1% of data\nax[0].set_title(\"Quaketime\")\nax[0].set_xlabel(\"Index\")\nax[0].set_ylabel(\"Quaketime in ms\");\n\nax[1].plot(train.index.values[::100], train.signal.values[::100], c=\"mediumseagreen\")\nax[1].set_title(\"Signal\")\nax[1].set_xlabel(\"Index\")\nax[1].set_ylabel(\"Acoustic Signal\");\nplt.tight_layout()\n","metadata":{"execution":{"iopub.status.busy":"2023-04-28T19:09:47.612150Z","iopub.execute_input":"2023-04-28T19:09:47.612743Z","iopub.status.idle":"2023-04-28T19:09:52.744322Z","shell.execute_reply.started":"2023-04-28T19:09:47.612694Z","shell.execute_reply":"2023-04-28T19:09:52.742848Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Research Question: What are the necessary steps to clean and prepare the dataset as we have categorial features and missing data? \n\n\nThe dataset consists of about 600 million datapoints, where each datapoint consists of the seismic signal and the time until the next earthquake. The seismic signal is considered to be the input and the time until the next earthquake, quaketime, the output. The input is a continuous string of acoustic data, which obviously is not categorical. From the acoustic data, features need to be extracted such as mean, standard deviation, minimum and maximum values. The first transformation can be seen in the cell below, as the performance of the model will be evaluated by test segments of 150000 signal datapoints, the train signal shall be divided into 4194 segments which also contain 150000 datapoints. Then from each segment, features shall be extracted.\n","metadata":{}},{"cell_type":"markdown","source":"","metadata":{}},{"cell_type":"code","source":"\n\n\n#The test segments all contain 150000 datapoints, therefore we split the training set into segments of 150000 points as well.\n#Inspired by Gabriel Preda's notebook: https://www.kaggle.com/code/gpreda/lanl-earthquake-eda-and-prediction?kernelSessionId=12610502\nrows = 150000\nsegments = int(np.floor(train.shape[0] / rows)) #Will print 4194, so we need to split the entire train set into 4194 segments.\nprint(segments)\n\nx = np.resize(train.signal,(segments, rows))\ny_full = np.resize(train.quaketime,(segments, rows))\ny = np.array(y_full[:,rows-1])[:,np.newaxis] #We only take the final value of the segment, as only this value needs to be predicted per the competitions rules.\nprint(y.shape)\n\n\n\n","metadata":{"execution":{"iopub.status.busy":"2023-04-28T19:09:52.746587Z","iopub.execute_input":"2023-04-28T19:09:52.746977Z","iopub.status.idle":"2023-04-28T19:09:56.638264Z","shell.execute_reply.started":"2023-04-28T19:09:52.746940Z","shell.execute_reply":"2023-04-28T19:09:56.636674Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### 2.2 Data Exploration\n\n<font color=#6698FF> Explore the features and target variables of the dataset. Think about making some scatter plots, box plots, histograms or printing the data, but feel free to choose any method that suits you.\nWhat do you think is the right performance\nmetric to use for this dataset? Clearly explain which performance metric you\nchoose and why.\nAlgorithmic bias can be a real problem in Machine Learning. Explain what you believe.\n\n</font>","metadata":{}},{"cell_type":"markdown","source":"If algorithmic bias were to manifest itself as producing systematic errors that create unfair outcomes, all must be done to minimize this bias. If, for example some model were to be trained to recognize certain ethnic groups, it is important to not underrepresent any ethnic group in the training data. However, sometimes a lack of data simply exists, in that case the model should not be disregarded, but in some way, a waiver should be included explaing the reason for the bias and and how to interpret the results keeping the potential bias in the back of your mind, in my opinion.","metadata":{}},{"cell_type":"code","source":"#One segment is plotted to get a feel of what we are working with\nfig, ax = plt.subplots(2,1, figsize=(20,12))\nfig.suptitle(\"One Train Segment\", fontsize = 20)\nax[0].plot(range(x.shape[1]), x[0,:], c=\"darkred\") \nax[0].set_title(\"Signal\")\nax[0].set_xlabel(\"Index\")\nax[0].set_ylabel(\"Acoustic Signal\");\n\nax[1].plot(range(x.shape[1]), y_full[0,:], c=\"mediumseagreen\")\nax[1].set_title(\"Quaketime\")\nax[1].set_xlabel(\"Index\")\nax[1].set_ylabel(\"Quaketime in ms\");\nplt.tight_layout()","metadata":{"execution":{"iopub.status.busy":"2023-04-28T19:09:56.641747Z","iopub.execute_input":"2023-04-28T19:09:56.642855Z","iopub.status.idle":"2023-04-28T19:09:57.950813Z","shell.execute_reply.started":"2023-04-28T19:09:56.642799Z","shell.execute_reply":"2023-04-28T19:09:57.949777Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### 2.3 Data Preparation\n\n<font color=#6698FF>  This dataset hasn’t been cleaned yet. Meaning that some attributes (features) are in numerical format and some are in categorial format. Moreover, there are missing values as well. However, all Scikit-learn’s implementations of these algorithms expect numerical features. Check for all features if they are in categorial and use a method to transform them to numerical values. For the numerical data, handle the missing data and normalize the data. \nNote that you are only allowed to use training data for preprocessing but you then need to perform similar changes on test data too.\nYou can use [pipelining](https://scikit-learn.org/stable/modules/generated/sklearn.pipeline.Pipeline.html) to help with the preprocessing.</font>\n\nThe data consists of a signal and an output, both numerical in format. The only preparation deemed necessary was properly segmenting the data as done in segment 2.1, noise could be removed from the source data, however, this is a very time-consuming and intricate process as a lot of knowledge about the generation of the data and the specific lab requirements would be necessary and this data is hard to find.","metadata":{}},{"cell_type":"code","source":"#Calculating the mean of each segment\nlist = []\nfor i in range(segments):\n    mean = x[i,:].mean()\n    list.append(mean)\n    \nX_means = np.array(list)[:,np.newaxis]\nprint(X_means.shape)","metadata":{"execution":{"iopub.status.busy":"2023-04-28T19:09:57.953198Z","iopub.execute_input":"2023-04-28T19:09:57.954562Z","iopub.status.idle":"2023-04-28T19:09:58.469503Z","shell.execute_reply.started":"2023-04-28T19:09:57.954516Z","shell.execute_reply":"2023-04-28T19:09:58.468541Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#Calculating the standard deviation of each segment\nlist1 = []\nfor o in range(segments):\n    std = x[o,:].std()\n    list1.append(std)\n    \nX_stds = np.array(list1)[:,np.newaxis]\nprint(X_stds.shape)\nprint(X_stds[50], X_stds[1050])","metadata":{"execution":{"iopub.status.busy":"2023-04-28T19:09:58.470742Z","iopub.execute_input":"2023-04-28T19:09:58.471556Z","iopub.status.idle":"2023-04-28T19:10:00.051256Z","shell.execute_reply.started":"2023-04-28T19:09:58.471518Z","shell.execute_reply":"2023-04-28T19:10:00.049879Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#Calculating the maximum value of each segment\nlist2 = []\nfor p in range(segments):\n    x_max = np.max(x[p,:])\n    list2.append(x_max)\n    \nX_max = np.array(list2)[:,np.newaxis]\nprint(X_max.shape)\nprint(X_max[50], X_max[1050])","metadata":{"execution":{"iopub.status.busy":"2023-04-28T19:10:00.053068Z","iopub.execute_input":"2023-04-28T19:10:00.053880Z","iopub.status.idle":"2023-04-28T19:10:00.592115Z","shell.execute_reply.started":"2023-04-28T19:10:00.053830Z","shell.execute_reply":"2023-04-28T19:10:00.590627Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#Calculating the average change in each segment, inspired by https://www.kaggle.com/code/artgor/seismic-data-eda-and-baseline?kernelSessionId=12760528\n\nlist3 = []\nfor a in range(segments):\n    x_diff = np.mean(np.diff(x[a,:]))\n    list3.append(x_diff)\n    \nX_diff = np.array(list3)[:,np.newaxis]\nprint(X_diff.shape)\nprint(X_diff[50], X_diff[1050])","metadata":{"execution":{"iopub.status.busy":"2023-04-28T19:10:00.594219Z","iopub.execute_input":"2023-04-28T19:10:00.595080Z","iopub.status.idle":"2023-04-28T19:10:01.264028Z","shell.execute_reply.started":"2023-04-28T19:10:00.595026Z","shell.execute_reply":"2023-04-28T19:10:01.262637Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"list4 = []\nfor b in range(segments):\n    x_first_50000 = x[b,:50000].std()\n    list4.append(x_first_50000)\n    \nX_first_50000 = np.array(list4)[:,np.newaxis]\nprint(X_first_50000.shape)\nprint(X_first_50000[6])","metadata":{"execution":{"iopub.status.busy":"2023-04-28T19:10:01.265536Z","iopub.execute_input":"2023-04-28T19:10:01.266029Z","iopub.status.idle":"2023-04-28T19:10:01.849213Z","shell.execute_reply.started":"2023-04-28T19:10:01.265975Z","shell.execute_reply":"2023-04-28T19:10:01.847639Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"list5 = []\nfor c in range(segments):\n    x_min = np.min(x[c,:])\n    list5.append(x_min)\n    \nX_min = np.array(list5)[:,np.newaxis]\nprint(X_min.shape)\n","metadata":{"execution":{"iopub.status.busy":"2023-04-28T19:10:01.850794Z","iopub.execute_input":"2023-04-28T19:10:01.851186Z","iopub.status.idle":"2023-04-28T19:10:02.391419Z","shell.execute_reply.started":"2023-04-28T19:10:01.851147Z","shell.execute_reply":"2023-04-28T19:10:02.390045Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"list6 = []\nfor d in range(segments):\n    x_q95 = np.quantile(x[d,:], 0.95)\n    list6.append(x_q95)\n    \nX_q95 = np.array(list6)[:,np.newaxis]\nprint(X_q95.shape)","metadata":{"execution":{"iopub.status.busy":"2023-04-28T19:10:02.396072Z","iopub.execute_input":"2023-04-28T19:10:02.396471Z","iopub.status.idle":"2023-04-28T19:10:07.742951Z","shell.execute_reply.started":"2023-04-28T19:10:02.396432Z","shell.execute_reply":"2023-04-28T19:10:07.741468Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"list7 = []\nfor e in range(segments):\n    x_q01 = np.quantile(x[e,:], 0.01)\n    list7.append(x_q01)\n    \nX_q01 = np.array(list7)[:,np.newaxis]\nprint(X_q01.shape)","metadata":{"execution":{"iopub.status.busy":"2023-04-28T19:10:07.744681Z","iopub.execute_input":"2023-04-28T19:10:07.745142Z","iopub.status.idle":"2023-04-28T19:10:12.832597Z","shell.execute_reply.started":"2023-04-28T19:10:07.745094Z","shell.execute_reply":"2023-04-28T19:10:12.831270Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#Here, the Hilbert mean is calculated, inspired by: https://www.kaggle.com/code/artgor/earthquakes-fe-more-features-and-samples?kernelSessionId=13269290\nfrom scipy.signal import hilbert\nfrom scipy.signal import hann\nfrom scipy.signal import convolve\n\nlist8 = []\nfor f in range(segments):\n    x_hil = np.abs(hilbert(x[f,:])).mean()\n    list8.append(x_hil)\n    \nX_hil = np.array(list8)[:,np.newaxis]\nprint(X_hil.shape)","metadata":{"execution":{"iopub.status.busy":"2023-04-28T19:10:12.833806Z","iopub.execute_input":"2023-04-28T19:10:12.834604Z","iopub.status.idle":"2023-04-28T19:10:57.146260Z","shell.execute_reply.started":"2023-04-28T19:10:12.834553Z","shell.execute_reply":"2023-04-28T19:10:57.144669Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#Here, the absolute difference is calculated, inspired by: https://www.kaggle.com/code/artgor/earthquakes-fe-more-features-and-samples?kernelSessionId=13269290\n\nlist9 = []\nfor g in range(segments):\n    x_mm = x[g,:].max() - np.abs(x[g,:].min())\n    list9.append(x_mm)\n    \nX_mm = np.array(list9)[:,np.newaxis]\nprint(X_mm.shape)","metadata":{"execution":{"iopub.status.busy":"2023-04-28T19:10:57.147814Z","iopub.execute_input":"2023-04-28T19:10:57.148265Z","iopub.status.idle":"2023-04-28T19:10:58.179851Z","shell.execute_reply.started":"2023-04-28T19:10:57.148227Z","shell.execute_reply":"2023-04-28T19:10:58.178510Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"\n#Here, Hann window mean is calculated, inspired by: https://www.kaggle.com/code/artgor/earthquakes-fe-more-features-and-samples?kernelSessionId=13269290\n\nlist10 = []\nfor h in range(segments):\n    x_hann = (convolve(x[h,:], hann(150), mode='same') / sum(hann(150))).mean()\n    list10.append(x_hann)\n    \nX_hann = np.array(list10)[:,np.newaxis]\nprint(X_hann.shape)","metadata":{"execution":{"iopub.status.busy":"2023-04-28T19:10:58.181727Z","iopub.execute_input":"2023-04-28T19:10:58.182104Z","iopub.status.idle":"2023-04-28T19:11:34.182371Z","shell.execute_reply.started":"2023-04-28T19:10:58.182065Z","shell.execute_reply":"2023-04-28T19:11:34.181059Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"list11 = []\nfor j in range(segments):\n    x_sum = x[j,:].sum()\n    list11.append(x_sum)\n    \nX_sum = np.array(list11)[:,np.newaxis]\nprint(X_sum.shape)","metadata":{"execution":{"iopub.status.busy":"2023-04-28T19:11:34.184037Z","iopub.execute_input":"2023-04-28T19:11:34.184759Z","iopub.status.idle":"2023-04-28T19:11:34.911500Z","shell.execute_reply.started":"2023-04-28T19:11:34.184707Z","shell.execute_reply":"2023-04-28T19:11:34.910329Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"list12 = []\nfor k in range(segments):\n    x_max_50000 = x[k,:50000].max()\n    list12.append(x_max_50000)\n    \nX_max_50000 = np.array(list12)[:,np.newaxis]\nprint(X_max_50000.shape)\nprint(X_max_50000[6])","metadata":{"execution":{"iopub.status.busy":"2023-04-28T19:11:34.912866Z","iopub.execute_input":"2023-04-28T19:11:34.913213Z","iopub.status.idle":"2023-04-28T19:11:35.098665Z","shell.execute_reply.started":"2023-04-28T19:11:34.913180Z","shell.execute_reply":"2023-04-28T19:11:35.097283Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"list13 = []\nfor l in range(segments):\n    x_std_last = x[l,:-50000].std()\n    list13.append(x_std_last)\n    \nX_std_last = np.array(list13)[:,np.newaxis]\nprint(X_std_last.shape)\nprint(X_std_last[6])","metadata":{"execution":{"iopub.status.busy":"2023-04-28T19:11:35.100258Z","iopub.execute_input":"2023-04-28T19:11:35.100602Z","iopub.status.idle":"2023-04-28T19:11:36.161226Z","shell.execute_reply.started":"2023-04-28T19:11:35.100547Z","shell.execute_reply":"2023-04-28T19:11:36.159963Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"list14 = []\nfor m in range(segments):\n    x_astd = np.abs(x[m,:]).std()\n    list14.append(x_astd)\n    \nX_astd = np.array(list14)[:,np.newaxis]\nprint(X_astd.shape)\nprint(X_astd[6])","metadata":{"execution":{"iopub.status.busy":"2023-04-28T19:11:36.162945Z","iopub.execute_input":"2023-04-28T19:11:36.163315Z","iopub.status.idle":"2023-04-28T19:11:37.905547Z","shell.execute_reply.started":"2023-04-28T19:11:36.163280Z","shell.execute_reply":"2023-04-28T19:11:37.904464Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import statistics\nlist15 = []\nfor n in range(segments):\n    x_median = statistics.median(x[n,:])\n    list15.append(x_median)\n    \nX_median = np.array(list15)[:,np.newaxis]\nprint(X_median.shape)\nprint(X_median[6])","metadata":{"execution":{"iopub.status.busy":"2023-04-28T19:11:37.906816Z","iopub.execute_input":"2023-04-28T19:11:37.907230Z","iopub.status.idle":"2023-04-28T19:14:43.302600Z","shell.execute_reply.started":"2023-04-28T19:11:37.907196Z","shell.execute_reply":"2023-04-28T19:14:43.301388Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#No. of large numbers in a segment len(x[np.abs(x) > 500])\nlist16 = []\nfor q in range(segments):\n    x_big500 = len(x[q,:][np.abs(x[q,:]) > 500])\n    list16.append(x_big500)\n    \nX_big500 = np.array(list16)[:,np.newaxis]\n\nprint(X_big500[X_big500 > 0])","metadata":{"execution":{"iopub.status.busy":"2023-04-28T19:14:43.304531Z","iopub.execute_input":"2023-04-28T19:14:43.304913Z","iopub.status.idle":"2023-04-28T19:14:43.539800Z","shell.execute_reply.started":"2023-04-28T19:14:43.304877Z","shell.execute_reply":"2023-04-28T19:14:43.538606Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"list17 = []\nfor r in range(segments):\n    x_q999 = np.quantile(x[r,:], 0.999)\n    list17.append(x_q999)\n    \nX_q999 = np.array(list17)[:,np.newaxis]\nprint(X_q999.shape)","metadata":{"execution":{"iopub.status.busy":"2023-04-28T19:14:43.541640Z","iopub.execute_input":"2023-04-28T19:14:43.541989Z","iopub.status.idle":"2023-04-28T19:14:48.523127Z","shell.execute_reply.started":"2023-04-28T19:14:43.541955Z","shell.execute_reply":"2023-04-28T19:14:48.521937Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"list18 = []\nfor s in range(segments):\n    x_q001 = np.quantile(x[r,:], 0.001)\n    list18.append(x_q001)\n    \nX_q001 = np.array(list18)[:,np.newaxis]\nprint(X_q001.shape)","metadata":{"execution":{"iopub.status.busy":"2023-04-28T19:14:48.524982Z","iopub.execute_input":"2023-04-28T19:14:48.525531Z","iopub.status.idle":"2023-04-28T19:14:52.711829Z","shell.execute_reply.started":"2023-04-28T19:14:48.525483Z","shell.execute_reply":"2023-04-28T19:14:52.710623Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#inter quartile range \nlist19 = []\nfor t in range(segments):\n    x_iqr = np.subtract(*np.percentile(x[t,:], [75, 25]))\n    list19.append(x_iqr)\n    \nX_iqr = np.array(list19)[:,np.newaxis]\nprint(X_iqr.shape)","metadata":{"execution":{"iopub.status.busy":"2023-04-28T19:14:52.713265Z","iopub.execute_input":"2023-04-28T19:14:52.713614Z","iopub.status.idle":"2023-04-28T19:15:00.846802Z","shell.execute_reply.started":"2023-04-28T19:14:52.713565Z","shell.execute_reply":"2023-04-28T19:15:00.845893Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#Maximum absolute value \nlist20 = []\nfor u in range(segments):\n    x_abs_max = np.abs(x[u,:]).max()\n    list20.append(x_abs_max)\n    \nX_abs_max = np.array(list20)[:,np.newaxis]\nprint(X_abs_max.shape)","metadata":{"execution":{"iopub.status.busy":"2023-04-28T19:15:00.848131Z","iopub.execute_input":"2023-04-28T19:15:00.848678Z","iopub.status.idle":"2023-04-28T19:15:01.481981Z","shell.execute_reply.started":"2023-04-28T19:15:00.848642Z","shell.execute_reply":"2023-04-28T19:15:01.480813Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#Minumum absolute value \nlist21 = []\nfor v in range(segments):\n    x_abs_min = np.abs(x[v,:]).min()\n    list21.append(x_abs_min)\n    \nX_abs_min = np.array(list21)[:,np.newaxis]\nprint(X_abs_min.shape)","metadata":{"execution":{"iopub.status.busy":"2023-04-28T19:15:01.483395Z","iopub.execute_input":"2023-04-28T19:15:01.483821Z","iopub.status.idle":"2023-04-28T19:15:02.110322Z","shell.execute_reply.started":"2023-04-28T19:15:01.483785Z","shell.execute_reply":"2023-04-28T19:15:02.108986Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#Ratio of maximum value to minimum value of a segment \nlist22 = []\nfor w in range(segments):\n    x_max_min = x[w,:].max() / np.abs(x[w,:].min())\n    list22.append(x_max_min)\n    \nX_max_min = np.array(list22)[:,np.newaxis]\nprint(X_max_min.shape)","metadata":{"execution":{"iopub.status.busy":"2023-04-28T19:15:02.111996Z","iopub.execute_input":"2023-04-28T19:15:02.112331Z","iopub.status.idle":"2023-04-28T19:15:03.150114Z","shell.execute_reply.started":"2023-04-28T19:15:02.112298Z","shell.execute_reply":"2023-04-28T19:15:03.148898Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#Max to min difference \nlist23 = []\nfor z in range(segments):\n    x_ratio = x[z,:].max() - np.abs(x[z,:].min())\n    list23.append(x_ratio)\n    \nX_ratio = np.array(list23)[:,np.newaxis]\nprint(X_ratio.shape)","metadata":{"execution":{"iopub.status.busy":"2023-04-28T19:15:03.151715Z","iopub.execute_input":"2023-04-28T19:15:03.152279Z","iopub.status.idle":"2023-04-28T19:15:04.169692Z","shell.execute_reply.started":"2023-04-28T19:15:03.152244Z","shell.execute_reply":"2023-04-28T19:15:04.168564Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#Plotting the average change per segment\nfig, ax = plt.subplots(1,1, figsize=(20,12))\nax.plot(range(X_diff.shape[0]), np.ravel(X_diff), c=\"darkred\") \nax.set_title(\"Average Change per segment\")\nax.set_xlabel(\"Segment\")\nax.set_ylabel(\"Average Change\");\n\n","metadata":{"execution":{"iopub.status.busy":"2023-04-28T19:15:04.176359Z","iopub.execute_input":"2023-04-28T19:15:04.176724Z","iopub.status.idle":"2023-04-28T19:15:04.613641Z","shell.execute_reply.started":"2023-04-28T19:15:04.176690Z","shell.execute_reply":"2023-04-28T19:15:04.612461Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#Combining the calculated features into one array:\ndata = np.vstack((X_means.T, X_stds.T, X_max.T, X_diff.T, X_first_50000.T, X_min.T, X_q95.T, X_q01.T, X_hil.T, X_hann.T, X_mm.T, X_max_50000.T, X_std_last.T, X_astd.T, X_median.T, X_big500.T, X_q999.T, X_q001.T, X_iqr.T, X_abs_max.T, X_abs_min.T, X_max_min.T, X_ratio.T)) #, X_sum.T is left out, because it seems to decrease accuracy.\nprint(data.shape)\n\n#standard normalize the data\nfrom sklearn import preprocessing\nscaler = preprocessing.StandardScaler().fit(data)\ndata = scaler.transform(data)\n","metadata":{"execution":{"iopub.status.busy":"2023-04-28T19:15:04.615149Z","iopub.execute_input":"2023-04-28T19:15:04.616177Z","iopub.status.idle":"2023-04-28T19:15:04.746619Z","shell.execute_reply.started":"2023-04-28T19:15:04.616124Z","shell.execute_reply":"2023-04-28T19:15:04.745295Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from sklearn.decomposition import PCA\nn_components = 5\npca = PCA(n_components = n_components)\npca = pca.fit(data.T)\ndata_pca = pca.transform(data.T).T\nprint(data.shape)\nprint(data_pca.shape)\n\nplt.bar(range(0,n_components), pca.explained_variance_ratio_ * 100, label = \"individual var\");\nplt.step(range(0,n_components), np.cumsum(pca.explained_variance_ratio_) * 100, \"r\", label = \"cumulative var\");\nplt.xlabel(\"principal component index\"); plt.ylabel(\"explained variance ratio %\");\nplt.legend()\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-04-28T19:15:04.748118Z","iopub.execute_input":"2023-04-28T19:15:04.749223Z","iopub.status.idle":"2023-04-28T19:15:05.176146Z","shell.execute_reply.started":"2023-04-28T19:15:04.749183Z","shell.execute_reply":"2023-04-28T19:15:05.174862Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def plot_2features(x_in):\n    #Create a figure\n    plt.figure()\n    colors = ['tab:blue', 'tab:orange', 'tab:green', 'tab:purple','tab:red','tab:pink','tab:grey','tab:brown','tab:olive','tab:cyan']\n\n    plt.scatter(x_in[5,:], x_in[20,:],  marker = 'o', color = colors[0])\n    \n    #plt.axis('scaled')\n    plt.xlabel('feature 1'); plt.ylabel('feature 2');\n    plt.tight_layout()\n    \nplot_2features(data)\n#The plot below shows two features plottet against each other\n","metadata":{"execution":{"iopub.status.busy":"2023-04-28T19:15:05.177703Z","iopub.execute_input":"2023-04-28T19:15:05.178148Z","iopub.status.idle":"2023-04-28T19:15:05.405247Z","shell.execute_reply.started":"2023-04-28T19:15:05.178113Z","shell.execute_reply":"2023-04-28T19:15:05.404100Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"###### 2.3.1 Train-test split\n<font color=#6698FF>In the cell below, we split the train data into a test and a train set. Set a value for the `test_size` yourself. Argue why the test value can not be too small or too large. You can also use k-fold cross validation.\n    \nThe data has been split in a 70% training set and a 30% testing set. Too large a train set would mean that the model would be trained really well for the given training set, but not for another set as it would have overfitted the data, too small a training set and the model would simply not be accurate enough.\n</font>","metadata":{}},{"cell_type":"code","source":"\n\n","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#Splitting the training data into a 70% training part and a 30%testing part\nfrom sklearn.model_selection import train_test_split\nX_train, X_test, Y_train, Y_test = train_test_split(data.T,y,test_size=0.3,shuffle=True, random_state=102)\nprint(X_train.shape)\nprint(Y_train.shape)\n","metadata":{"execution":{"iopub.status.busy":"2023-04-28T19:15:05.406527Z","iopub.execute_input":"2023-04-28T19:15:05.406920Z","iopub.status.idle":"2023-04-28T19:15:05.417029Z","shell.execute_reply.started":"2023-04-28T19:15:05.406885Z","shell.execute_reply":"2023-04-28T19:15:05.415725Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"What model/classifier provides the best result for this application?\n\n\nIf the problem would be categorical and the data linearly seperable, the logistic regression method would be a good fit, however, the data is not categorical,this limits the application to using models like a neural network or linear (or polynomial)regression. A neural network seems to be a good fit as there is quite a lot of processing power availeble and it can compute numerical outputs for the complex dataset.","metadata":{}},{"cell_type":"markdown","source":"What is the best choice for our cost function and performance metrics for this problem?\n\nSince this is not a categorical problem, but rather a regression problem,  the Mean Square Error will be used as a performance metric because the MSE gives an indication of accuracy, which is desired. A performance metric which is less sensitive to ouliers such as Mean Absolute Error is less applicable, as the data contains quite a lot of outliers and they are important for having an accurate model\n\n","metadata":{}},{"cell_type":"markdown","source":"\n## 3. Training and Results\n<font color=#6698FF> Briefly introduce the classification algorithms you choose. \nPresent your accuracies for both test and training data for all classifiers. Analyse the performance on test and training in terms of bias and variance. Give one advantage and one drawback of the method you use.\n    \nA simple linear model is created, with the 22 features a polynomial function will be made to fit the data, the resulting accuracy was about 40% for the test data, one advantage of using the linear model is that it is simple and computationally inexpensive. A drawback of the model is that the dataset must be structured in a way that a line can be drawn through the data with sufficient accuracy, for many datasets this is not the case and another model would be a better suit.","metadata":{}},{"cell_type":"code","source":"#First, a linear model shall be tried, \nfrom sklearn.linear_model import LinearRegression\nfrom sklearn import datasets, svm\n\nmodel = LinearRegression() \n\n# fit the model to data\nmodel.fit(X_train, Y_train)\n\n# predict\npredictions=model.predict(X_test)\n\n# accuracy\nscore=model.score(X_test, Y_test)\nprint(score)\n\n# get importance\nimportance = model.coef_[0]\nprint(np.mean(importance))\n\n#plotting feature importance, normalized to maximum absolute value for proper visualization\nplt.bar([x for x in range(len(importance))], importance/max(abs(importance)))\nplt.show()\n","metadata":{"execution":{"iopub.status.busy":"2023-04-28T19:15:05.418565Z","iopub.execute_input":"2023-04-28T19:15:05.419000Z","iopub.status.idle":"2023-04-28T19:15:05.723622Z","shell.execute_reply.started":"2023-04-28T19:15:05.418965Z","shell.execute_reply":"2023-04-28T19:15:05.722346Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#Here, a svm linear model shall be tried.\nfrom sklearn.svm import NuSVR\n\nsvm = NuSVR()\nsvm.fit(X_train,np.ravel(Y_train))\ny_pred = svm.predict(X_train)\n\n#Accuracy Score:\n# accuracy\nscore=svm.score(X_test, Y_test)\nprint(score)\n\nplt.figure(figsize=(6, 6))\nplt.scatter(np.ravel(Y_train), y_pred)\n#plt.xlim(0, 20)\n#plt.ylim(0, 20)\nplt.xlabel('actual', fontsize=12)\nplt.ylabel('predicted', fontsize=12)\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-04-28T19:15:05.724906Z","iopub.execute_input":"2023-04-28T19:15:05.725484Z","iopub.status.idle":"2023-04-28T19:15:06.785639Z","shell.execute_reply.started":"2023-04-28T19:15:05.725449Z","shell.execute_reply":"2023-04-28T19:15:06.784221Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**NON-LINEAR MODEL**","metadata":{}},{"cell_type":"code","source":"# non linear model\nfrom keras.models import Sequential\nfrom keras.layers import Dense, Dropout\nimport keras as keras\nfrom sklearn import metrics\nfrom keras.datasets import mnist\nfrom tensorflow.keras.utils import to_categorical\nfrom keras.layers import Flatten\n\n# initializing a sequential neural network\nfull_model = Sequential()\n\n# adding layers\nfull_model.add(Dense(60,activation='relu'))\nfull_model.add(Dense(159,activation='relu'))\nfull_model.add(Dropout(.5))\nfull_model.add(Dense(100,activation='relu'))\nfull_model.add(Dense(180,activation='relu'))\nfull_model.add(Dropout(.5))\nfull_model.add(Dense(1))\n\n\n# Compiling the ANN\nfull_model.compile(loss='mean_squared_error', optimizer = 'adam', metrics=['accuracy'])\n\n\n\n\n\nfull_model.fit(X_train,Y_train,validation_data=(X_test,Y_test), epochs=100, batch_size=32)\n\nloss_df = pd.DataFrame(full_model.history.history)\nloss_df.plot(figsize=(12,8))\n\n\npredictions = full_model.predict(X_test).reshape(-1,1)\n\nfig, (ax1) = plt.subplots(1)\n\nax1.plot(Y_test, label = \"Actual\")\nax1.plot(predictions, label = \"predictions\")\nax1.set_ylabel(\"Time to Failure [s]\")\nax1.set_xlabel(\"Index [n]\")\nax1.legend()\n\n","metadata":{"execution":{"iopub.status.busy":"2023-04-28T19:15:06.787340Z","iopub.execute_input":"2023-04-28T19:15:06.788240Z","iopub.status.idle":"2023-04-28T19:16:00.734964Z","shell.execute_reply.started":"2023-04-28T19:15:06.788191Z","shell.execute_reply":"2023-04-28T19:16:00.733666Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(len(predictions))\nfig, (ax2) = plt.subplots(1)\n\nn = 100\nax2.plot(Y_test[:n], label = \"Actual\")\nax2.plot(predictions[:n], label = \"predictions\")\nax2.set_title(\"Subset of Datapoints and Predictions\")\nax2.set_ylabel(\"Time to Failure [s]\")\nax2.set_xlabel(\"Index [n]\")\nax2.legend()","metadata":{"execution":{"iopub.status.busy":"2023-04-28T19:16:00.736679Z","iopub.execute_input":"2023-04-28T19:16:00.737066Z","iopub.status.idle":"2023-04-28T19:16:01.019191Z","shell.execute_reply.started":"2023-04-28T19:16:00.737029Z","shell.execute_reply":"2023-04-28T19:16:01.017872Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fig = plt.figure(figsize=(10,5))\nplt.scatter(Y_test,predictions, label='pred vs actual')\nplt.plot(Y_test,Y_test,'r', label='y=x')\nplt.ylabel('Y_test')\nplt.xlabel('predictions')\nplt.legend()\nprint('RMSE = ', np.sqrt(metrics.mean_squared_error(Y_test, predictions)))","metadata":{"execution":{"iopub.status.busy":"2023-04-28T19:16:01.020503Z","iopub.execute_input":"2023-04-28T19:16:01.020859Z","iopub.status.idle":"2023-04-28T19:16:01.361792Z","shell.execute_reply.started":"2023-04-28T19:16:01.020826Z","shell.execute_reply":"2023-04-28T19:16:01.360528Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#Run after creating the model to create the submission.csv file\n#Submission creation inspired by: https://www.kaggle.com/code/artgor/seismic-data-eda-and-baseline?kernelSessionId=12760528\nfrom tqdm import tqdm\nlist62 = []\nsubmission = pd.read_csv(\"../input/LANL-Earthquake-Prediction/sample_submission.csv\", index_col='seg_id')\nfor i, seg_id in enumerate(tqdm(submission.index)):\n    seg = pd.read_csv('../input/LANL-Earthquake-Prediction/test/' + seg_id + '.csv')\n    x_sub = seg['acoustic_data'].values\n    list62.append(x_sub)\n    \nx = np.array(list62) #This overwrites the data from the training set, with data from the test set used only for generating a score\nprint(x.shape)\n\n#Calculating the mean of each segment\nsegments = x.shape[0]\nlist = []\nfor i in range(segments):\n    mean = x[i,:].mean()\n    list.append(mean)\n    \nX_means = np.array(list)[:,np.newaxis]\nprint(X_means.shape)\n\n#Calculating the standard deviation of each segment\nlist1 = []\nfor o in range(segments):\n    std = x[o,:].std()\n    list1.append(std)\n    \nX_stds = np.array(list1)[:,np.newaxis]\nprint(X_stds.shape)\nprint(X_stds[50], X_stds[1050])\n\n#Calculating the maximum value of each segment\nlist2 = []\nfor p in range(segments):\n    x_max = np.max(x[p,:])\n    list2.append(x_max)\n    \nX_max = np.array(list2)[:,np.newaxis]\nprint(X_max.shape)\nprint(X_max[50], X_max[1050])\n\n#Calculating the average change in each segment, inspired by https://www.kaggle.com/code/artgor/seismic-data-eda-and-baseline?kernelSessionId=12760528\n\nlist3 = []\nfor a in range(segments):\n    x_diff = np.mean(np.diff(x[a,:]))\n    list3.append(x_diff)\n    \nX_diff = np.array(list3)[:,np.newaxis]\nprint(X_diff.shape)\nprint(X_diff[50], X_diff[1050])\n\nlist4 = []\nfor b in range(segments):\n    x_first_50000 = x[b,:50000].std()\n    list4.append(x_first_50000)\n    \nX_first_50000 = np.array(list4)[:,np.newaxis]\nprint(X_first_50000.shape)\nprint(X_first_50000[6])\n\nlist5 = []\nfor c in range(segments):\n    x_min = np.min(x[c,:])\n    list5.append(x_min)\n    \nX_min = np.array(list5)[:,np.newaxis]\nprint(X_min.shape)\n\nlist6 = []\nfor d in range(segments):\n    x_q95 = np.quantile(x[d,:], 0.95)\n    list6.append(x_q95)\n    \nX_q95 = np.array(list6)[:,np.newaxis]\nprint(X_q95.shape)\n\nlist7 = []\nfor e in range(segments):\n    x_q01 = np.quantile(x[e,:], 0.01)\n    list7.append(x_q01)\n    \nX_q01 = np.array(list7)[:,np.newaxis]\nprint(X_q01.shape)\n\n#Here, the Hilbert mean is calculated, inspired by: https://www.kaggle.com/code/artgor/earthquakes-fe-more-features-and-samples?kernelSessionId=13269290\n\n\nlist8 = []\nfor f in range(segments):\n    x_hil = np.abs(hilbert(x[f,:])).mean()\n    list8.append(x_hil)\n    \nX_hil = np.array(list8)[:,np.newaxis]\nprint(X_hil.shape)\n\n#Here, the absolute difference is calculated, inspired by: https://www.kaggle.com/code/artgor/earthquakes-fe-more-features-and-samples?kernelSessionId=13269290\n\nlist9 = []\nfor g in range(segments):\n    x_mm = x[g,:].max() - np.abs(x[g,:].min())\n    list9.append(x_mm)\n    \nX_mm = np.array(list9)[:,np.newaxis]\nprint(X_mm.shape)\n\n\n#Here, Hann window mean is calculated, inspired by: https://www.kaggle.com/code/artgor/earthquakes-fe-more-features-and-samples?kernelSessionId=13269290\n\nlist10 = []\nfor h in range(segments):\n    x_hann = (convolve(x[h,:], hann(150), mode='same') / sum(hann(150))).mean()\n    list10.append(x_hann)\n    \nX_hann = np.array(list10)[:,np.newaxis]\nprint(X_hann.shape)\n\nlist11 = []\nfor j in range(segments):\n    x_sum = x[j,:].sum()\n    list11.append(x_sum)\n    \nX_sum = np.array(list11)[:,np.newaxis]\nprint(X_sum.shape)\n\nlist12 = []\nfor k in range(segments):\n    x_max_50000 = x[k,:50000].max()\n    list12.append(x_max_50000)\n    \nX_max_50000 = np.array(list12)[:,np.newaxis]\nprint(X_max_50000.shape)\nprint(X_max_50000[6])\n\nlist13 = []\nfor l in range(segments):\n    x_std_last = x[l,:-50000].std()\n    list13.append(x_std_last)\n    \nX_std_last = np.array(list13)[:,np.newaxis]\nprint(X_std_last.shape)\nprint(X_std_last[6])\n\nlist14 = []\nfor m in range(segments):\n    x_astd = np.abs(x[m,:]).std()\n    list14.append(x_astd)\n    \nX_astd = np.array(list14)[:,np.newaxis]\nprint(X_astd.shape)\nprint(X_astd[6])\n\nimport statistics\nlist15 = []\nfor n in range(segments):\n    x_median = statistics.median(x[n,:])\n    list15.append(x_median)\n    \nX_median = np.array(list15)[:,np.newaxis]\nprint(X_median.shape)\nprint(X_median[6])\n\n#No. of large numbers in a segment len(x[np.abs(x) > 500])\nlist16 = []\nfor q in range(segments):\n    x_big500 = len(x[q,:][np.abs(x[q,:]) > 500])\n    list16.append(x_big500)\n    \nX_big500 = np.array(list16)[:,np.newaxis]\n\nprint(X_big500[X_big500 > 0])\n\nlist17 = []\nfor r in range(segments):\n    x_q999 = np.quantile(x[r,:], 0.999)\n    list17.append(x_q999)\n    \nX_q999 = np.array(list17)[:,np.newaxis]\nprint(X_q999.shape)\n\nlist18 = []\nfor s in range(segments):\n    x_q001 = np.quantile(x[r,:], 0.001)\n    list18.append(x_q001)\n    \nX_q001 = np.array(list18)[:,np.newaxis]\nprint(X_q001.shape)\n\n#inter quartile range \nlist19 = []\nfor t in range(segments):\n    x_iqr = np.subtract(*np.percentile(x[t,:], [75, 25]))\n    list19.append(x_iqr)\n    \nX_iqr = np.array(list19)[:,np.newaxis]\nprint(X_iqr.shape)\n\n#Maximum absolute value \nlist20 = []\nfor u in range(segments):\n    x_abs_max = np.abs(x[u,:]).max()\n    list20.append(x_abs_max)\n    \nX_abs_max = np.array(list20)[:,np.newaxis]\nprint(X_abs_max.shape)\n\n#Minumum absolute value \nlist21 = []\nfor v in range(segments):\n    x_abs_min = np.abs(x[v,:]).min()\n    list21.append(x_abs_min)\n    \nX_abs_min = np.array(list21)[:,np.newaxis]\nprint(X_abs_min.shape)\n\n#Ratio of maximum value to minimum value of a segment \nlist22 = []\nfor w in range(segments):\n    x_max_min = x[w,:].max() / np.abs(x[w,:].min())\n    list22.append(x_max_min)\n    \nX_max_min = np.array(list22)[:,np.newaxis]\nprint(X_max_min.shape)\n\n#Max to min difference \nlist23 = []\nfor z in range(segments):\n    x_ratio = x[z,:].max() - np.abs(x[z,:].min())\n    list23.append(x_ratio)\n    \nX_ratio = np.array(list23)[:,np.newaxis]\nprint(X_ratio.shape)\n\n#Combining the calculated features into one array:\ndata = np.vstack((X_means.T, X_stds.T, X_max.T, X_diff.T, X_first_50000.T, X_min.T, X_q95.T, X_q01.T, X_hil.T, X_hann.T, X_mm.T, X_max_50000.T, X_std_last.T, X_astd.T, X_median.T, X_big500.T, X_q999.T, X_q001.T, X_iqr.T, X_abs_max.T, X_abs_min.T, X_max_min.T, X_ratio.T)) #, X_sum.T is left out, because it seems to decrease accuracy.\nprint(data.shape)\n\n#standard normalize the data\n\nscaler = preprocessing.StandardScaler().fit(data)\ndata = scaler.transform(data)\n\npredictions = full_model.predict(data.T).reshape(-1,1) #Supposed to be ran after all the other cells are ran, only used for creating the submission file\nprint(predictions.shape)\n\n\nsubmission[\"time_to_failure\"] = predictions #Run after creating the model to create the submission.csv file\n\nsubmission.to_csv('submission.csv')\n\n","metadata":{"execution":{"iopub.status.busy":"2023-04-28T19:21:58.733867Z","iopub.execute_input":"2023-04-28T19:21:58.734523Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## 4. Discussion and Conclusion\n\n\nIn conclusion we can say that given the assignment of generating numerical predictions and the complexity of the dataset, it is better to use a non-linear model, mainly a neural network since it will be able to predict numeric outputs based on the input features and training data and form a good fit with quite a low MSE indicating relatively high accuracy. For a linear model it is also possible to do numeric predictions, however, the linear model simply generates a line to fit the datapoints, which could work well for classification problems where the classes are linearly seperable, howeve, here that is not the case.\n\nIf we look at the linear model, the accuracy was about 40% for the linear model and 24% for the linear SVM model, meaning that it is not very accurate. If we take a look at the non-linear model predictions plotted against the actual values (blue and orange plot), it is not good at predicting very high values, however it does generally follow the trend of the training data. It looks like the model is not good at predicting the time values with a large amplitude, perhaps more features that encorporate different statistal information would result in a greater accuracy.\n\nThe NN model uses 4 hidden layers with variable nodes, the way we decided on this number was simply by trial and error. Too few hidden layers and the output would not make sense, too many hidden layers and nodes and the model would be overfitted. By tweaking those numbers we found a semi-optimal model. Also the neural-network doesn't take many inputs further reducing the accuracy. We found it not to be too worrying since the main goal was to explore the linear and non-linear model on a real-life example where we think we kind of accomplished that goal.\n\nThe features extrated were basic statistical characteristics of the data like, maximum values, averages, standard deviations and rate of change. Including many more and advanced features would probably improve the model accuracy as stated above. To improve this the next time, we would need to dive deeper in what features say the most about when an earthquake is about to happen. Also some optimisation steps were overlooked like Principal Component Analysis, which (in theory) should also simplify the linear model by reducing the inputs and maximising the variance to make it easier for the model to distinguish and find a best fit. \n\nFor this example a non-linear model is better to use for prediction then your linear model. Even though both models are not perfect, it is possible to distinguish a difference in the linear model against the neural-network. The main take-away form these results is that for a more accurate model, both linear and non-linear, it is better to have as many important features as possible while keeping the model managable for the hardware and or software used. The model to be used depends on many factors such as: Is the problem a regression problem or a categorization problem? Is the data linearly seperable? How computationally intensive can the model be? Oftentimes a lot of trial and error is necessary to find the best fit for your dataset, but proper preparation and insight into the problem can guide you in the right direction.","metadata":{}}]}