{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Final Project\n","metadata":{"id":"GF4WsCxg-lL9"}},{"cell_type":"markdown","source":"## 1. Introduction\n","metadata":{"id":"3NnAgB-y0knj"}},{"cell_type":"markdown","source":"<font color=#6698FF> Finn Heijink, Username: Finn Heijink \n\n<font color=#6698FF>Joshua van der Geize, Username: joshuairmf</font>\n\n<font color=#6698FF> Linear model: Private score of 3.33533, Public score of 2.77743\n    \n<font color=#6698FF> Non-linear model: Private score of 2.78641, Public score of 2.21888","metadata":{"id":"DPJV5e8sDVnT"}},{"cell_type":"markdown","source":"## 2. Data\n","metadata":{"id":"PB7FLdQn-dmK"}},{"cell_type":"markdown","source":"### 2.1 Dataset\n<font color=#6698FF>In this section, we load and explore the dataset.","metadata":{"id":"vfGikOxUAxJB"}},{"cell_type":"code","source":"# Import Packages\nimport os\nimport warnings\nimport numpy as np\nimport pandas as pd\nimport seaborn as sns\nfrom scipy import fft\nimport scipy.stats as stats\nimport matplotlib.pyplot as plt\nfrom IPython.display import HTML\nfrom sklearn import preprocessing\nfrom sklearn.decomposition import PCA\nfrom sklearn.model_selection import KFold\nfrom sklearn.metrics import mean_squared_error\nfrom sklearn.linear_model import LinearRegression\nfrom sklearn.ensemble import RandomForestRegressor\nfrom sklearn.model_selection import cross_val_score\nfrom sklearn.model_selection import train_test_split\n\n%matplotlib inline\nsns.set()\n\nwarnings.filterwarnings(\"ignore\", category=DeprecationWarning)\nwarnings.filterwarnings(\"ignore\", category=UserWarning)\nwarnings.filterwarnings(\"ignore\", category=FutureWarning)","metadata":{"id":"elbwcEY42uEv","execution":{"iopub.status.busy":"2023-04-27T15:30:31.808645Z","iopub.execute_input":"2023-04-27T15:30:31.809003Z","iopub.status.idle":"2023-04-27T15:30:31.822309Z","shell.execute_reply.started":"2023-04-27T15:30:31.808953Z","shell.execute_reply":"2023-04-27T15:30:31.821182Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Loading the dataset\nPATH=\"../input/\"\nos.listdir(PATH)","metadata":{"id":"IaLgYknyvOjr","outputId":"a6c7c738-22a6-45a0-d21b-e9732e5e1531","execution":{"iopub.status.busy":"2023-04-27T15:30:31.824368Z","iopub.execute_input":"2023-04-27T15:30:31.824789Z","iopub.status.idle":"2023-04-27T15:30:31.842543Z","shell.execute_reply.started":"2023-04-27T15:30:31.824716Z","shell.execute_reply":"2023-04-27T15:30:31.841562Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Import the training data\n\n%time\ntrain = pd.read_csv(os.path.join(PATH,'train.csv'), low_memory=True, dtype={'acoustic_data': np.int16, 'time_to_failure': np.float32})","metadata":{"id":"GlAGb6lL0Isa","outputId":"eade3f63-dc7f-47ed-dfb2-a1330aa1b57d","execution":{"iopub.status.busy":"2023-04-27T15:31:07.366880Z","iopub.execute_input":"2023-04-27T15:31:07.367257Z","iopub.status.idle":"2023-04-27T15:35:13.500723Z","shell.execute_reply.started":"2023-04-27T15:31:07.367188Z","shell.execute_reply":"2023-04-27T15:35:13.499717Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{"id":"yV7vrCPBlOsX","jupyter":{"source_hidden":true}},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Import the test data \ntest_path = \"../input/test\"\ntest_samples_dir = os.listdir(test_path)\ntest_samples = []\n\n# Convert test data to array of segments\nfor sample in test_samples_dir:\n    test_samples = np.append(test_samples, sample.rsplit('.', 1)[0])\n\ndef get_test_set(seg):\n    return pd.read_csv(os.path.join(test_path,test_samples_dir[seg]), dtype={'acoustic_data': np.int16})","metadata":{"id":"Zlgbdcggn_Gg","execution":{"iopub.status.busy":"2023-04-27T15:30:31.849972Z","iopub.status.idle":"2023-04-27T15:30:31.850462Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### 2.2 Data Exploration\n<font color=#6698FF> We notice that the dataset consists of a very long sequence: we need to generate features based on this sequence. Since the test data provided is given in segments of 150k samples, we split the training data in segments of 150k samples aswell and find features for each segments. We take the average of the time to failure in this segment: there is very little difference in the time to failure over the span of this segment, thus using the average is a viable approximation which allows us to use models with a single output for a single sample.\n    \n<font color=#6698FF> We use the MSE and a plot visualising the data as a metric for the accuracy: the MSE gives a numerical analysis of the model but does not clearly show the shortcomings of the model. There is a pattern in the data itself which we can easily show using plots of the average time to failure and the predicted time to failure. We can do this on the ordered segments for easy visualisation, but note that this is not very realistic: we use this in conjunction with the MSE on the test set (that is to say, 30% of the provided training data which is not used for training) and the training set (the remaining 70% of the provided training data which is used for training) for better accuracy and understanding of the model.","metadata":{"id":"QN5jS2_p_Y9N"}},{"cell_type":"code","source":"# Show the training data for the first 5 samples\ntrain.head(5)","metadata":{"id":"IrIKBNJ03aK1","outputId":"9a7a4111-adbf-47ff-b3da-b1b64f507e77","execution":{"iopub.status.busy":"2023-04-27T15:30:31.854209Z","iopub.status.idle":"2023-04-27T15:30:31.855035Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Pick 1 out of every 100 samples for both the data and the time to failure\ndata_acoustic_preview = train.acoustic_data.values[::100]\ndata_time_to_failure_preview = train.time_to_failure.values[::100]","metadata":{"id":"o2LRKNWFAlkK","execution":{"iopub.status.busy":"2023-04-27T15:30:31.855967Z","iopub.status.idle":"2023-04-27T15:30:31.856558Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Plot the sampled data (1% of the total data)\nfig, ax = plt.subplots(figsize=(10,8))\n\nax.plot(data_acoustic_preview)\nax2 = ax.twinx()\nax2.plot(data_time_to_failure_preview, color='red')\nax.set_title(\"Plot of 1% of the total data\")","metadata":{"id":"FeezfBYg_j2t","outputId":"a618829a-c55b-4157-cab5-d26f9464bff7","execution":{"iopub.status.busy":"2023-04-27T15:30:31.858007Z","iopub.status.idle":"2023-04-27T15:30:31.858590Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Show a closeup of the initial 80k samples of the sampled data\nfig, ax = plt.subplots(figsize=(10,8))\nax.plot(data_acoustic_preview[:80000])\nax2 = ax.twinx()\nax2.plot(data_time_to_failure_preview[:80000], color='red')\nax.set_title(\"Plot of the first 80k samples of 1% of the total data\")","metadata":{"id":"4tuRI7yIGKbD","outputId":"9a24aefe-d35f-4f48-f12d-70fcc23e10dd","execution":{"iopub.status.busy":"2023-04-27T15:30:31.891310Z","iopub.status.idle":"2023-04-27T15:30:31.892540Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Plot a sample segment of the test set\nfig, ax = plt.subplots(figsize=(10,8))\ntest = get_test_set(20)\n\nax.plot(test.acoustic_data.values)\nax.set_title(\"Plot of a segment of the test set\")","metadata":{"id":"psu8FXCTHylR","outputId":"d7022539-dae5-4969-c227-907d27db4928","execution":{"iopub.status.busy":"2023-04-27T15:30:31.893680Z","iopub.status.idle":"2023-04-27T15:30:31.894559Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### 2.3 Data Preparation\n\n<font color=#6698FF> We apply a wide range of features stemming from two notebooks on Kaggle: one focused on simple features regarding the segments itself (https://www.kaggle.com/code/artgor/seismic-data-eda-and-baseline/notebook), and one using the Fourier Transform and more statistical features (https://www.kaggle.com/code/gpreda/lanl-earthquake-eda-and-prediction/notebook). Afterwards, we normalize the features.\n    \n<font color=#6698FF> We split the provided training segments into a 70/30 train/test split. Finally, we apply the same feature selection and normalization to the provided test segments. ","metadata":{"id":"4A5GsVzgATRu"}},{"cell_type":"code","source":"# First, we want to segment our data\n# Segment size is 150.000 to match the length of each segment of the test data\n# Truncate off the end if it doesn't fit\n\ndef make_segments(data):\n    # First we (re)define our data as a local variable\n    segmented = data\n\n    # Truncate it so the size fits\n    segmented = segmented[:int(np.floor(np.size(segmented)/150_000)*150_000)]\n\n    # Reshape into segments, this should result in a 4194x150000 array\n    segmented = np.reshape(segmented, (int(np.size(segmented)/150_000),150_000))\n\n    return segmented\n\nX_segs = make_segments(train.acoustic_data.values)\nY_segs = make_segments(train.time_to_failure.values)\n\n# Checks\nprint(\"Shape of X_train: \" + str(np.shape(X_segs)))\nprint(\"Segment Size: \" + str(np.size(X_segs[0])))\n\nprint(\"Shape of Y_train: \" + str(np.shape(Y_segs)))\nprint(\"Segment Size: \" + str(np.size(Y_segs[0])))","metadata":{"id":"UoJ5E5ZfMPW_","outputId":"347e9392-d47a-405b-cca9-18900b7892fd","execution":{"iopub.status.busy":"2023-04-27T15:30:31.896650Z","iopub.status.idle":"2023-04-27T15:30:31.897591Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<font color=#6698FF> The sizes match, so we can now start using them for training. </font>\n\n<font color=#6698FF> We use two notebooks to find features for the data: one of the feature sets is more oriented towards the statistical features in the time domain, while one of the feature sets is more oriented towards the statistical features in the frequency domain.","metadata":{"id":"_t1YOeZZuAJm"}},{"cell_type":"code","source":"# Take the average of the time to failure over the interval of 150k samples\nttf_average = np.mean(Y_segs, axis=1)\n\n# Define a function that returns features of the dataset in a single feature matrix\ndef feature_extractor(data):\n    features = []\n\n    # Simple Statistical Features (Based on https://www.kaggle.com/code/artgor/seismic-data-eda-and-baseline/notebook)\n    features.append(np.mean(data))\n    features.append(np.std(data))\n    features.append(np.max(data))\n    features.append(np.min(data))\n    features.append(np.mean(np.diff(data)))\n    features.append(np.abs(np.max(data)))\n    features.append(np.abs(np.min(data)))\n    features.append(np.std(data[:50000]))\n    features.append(np.std(data[-50000:]))\n    features.append(np.std(data[:10000]))\n    features.append(np.std(data[-10000:]))\n    features.append(np.mean(data[:50000]))\n    features.append(np.mean(data[-50000:]))\n    features.append(np.mean(data[:10000]))\n    features.append(np.mean(data[-10000:]))\n    features.append(np.min(data[:50000]))\n    features.append(np.min(data[-50000:]))\n    features.append(np.min(data[:10000]))\n    features.append(np.min(data[-10000:]))\n    features.append(np.max(data[:50000]))\n    features.append(np.max(data[-50000:]))\n    features.append(np.max(data[:10000]))\n    features.append(np.max(data[-10000:]))\n\n    # FFT Features (Based on https://www.kaggle.com/code/gpreda/lanl-earthquake-eda-and-prediction/notebook)\n    real_fft = np.real(fft(data))\n    imag_fft = np.imag(fft(data))\n    features.append(np.max(real_fft))\n    features.append(np.min(real_fft))\n    features.append(np.std(real_fft))\n    features.append(np.mean(real_fft))\n    features.append(np.max(imag_fft))\n    features.append(np.min(imag_fft))\n    features.append(np.std(imag_fft))\n    features.append(np.mean(imag_fft))\n    features.append(np.max(real_fft[-5000:]))\n    features.append(np.min(real_fft[-5000:]))\n    features.append(np.std(real_fft[-5000:]))\n    features.append(np.mean(real_fft[-5000:]))\n    features.append(np.max(imag_fft[-5000:]))\n    features.append(np.min(imag_fft[-5000:]))\n    features.append(np.std(imag_fft[-5000:]))\n    features.append(np.mean(imag_fft[-5000:]))\n    features.append(np.max(real_fft[-15000:]))\n    features.append(np.min(real_fft[-15000:]))\n    features.append(np.std(real_fft[-15000:]))\n    features.append(np.mean(real_fft[-15000:]))\n    features.append(np.max(imag_fft[-15000:]))\n    features.append(np.min(imag_fft[-15000:]))\n    features.append(np.std(imag_fft[-15000:]))\n    features.append(np.mean(imag_fft[-15000:]))\n\n    # Quantile Features\n    features.append(np.quantile(np.abs(data), 0.95))\n    features.append(np.quantile(np.abs(data), 0.99))\n    features.append(np.quantile(np.abs(data), 0.05))\n    features.append(np.quantile(np.abs(data), 0.01))\n    features.append(np.quantile(data, 0.95))\n    features.append(np.quantile(data, 0.99))\n    features.append(np.quantile(data, 0.05))\n    features.append(np.quantile(data, 0.01))\n\n    # Others\n    features.append(stats.kurtosis(data))\n    features.append(stats.skew(data))\n    features.append(np.median(data))\n\n    return np.asarray(features)","metadata":{"id":"glAIFw-5tNtZ","outputId":"13cebe5b-2031-4fda-a9e1-5dffd1d3d438","execution":{"iopub.status.busy":"2023-04-27T15:30:31.904629Z","iopub.status.idle":"2023-04-27T15:30:31.905670Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Element-wise feature extraction using the above defined function\nfeatures = []\n\nfor sample in X_segs:\n    features += [feature_extractor(sample)]\n\nprint(np.shape(features)) # This shows the shape of the data: (segments, features)","metadata":{"id":"g0qHnIL_Aqsy","execution":{"iopub.status.busy":"2023-04-27T15:30:31.909991Z","iopub.status.idle":"2023-04-27T15:30:31.911081Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Data Processing\nscaler = preprocessing.StandardScaler().fit(features)\ndata = scaler.transform(features)\n\nfrom sklearn.decomposition import PCA\n\nn = 10 # Amount of PCA components used for the models\n\npca = PCA(n_components=n) # Use PCA with n PCA components\npca = pca.fit(data) # Find the components for the dataset\ndata_pca_skl = pca.fit_transform(data) # Transform the dataset to a dataset of the n PCA components.\n\nfig, ax = plt.subplots(2)\n\n# perform pca on features\nax[0].bar(range(0,n), pca.explained_variance_ratio_, label=\"individual var\");\nax[0].step(range(0,n), np.cumsum(pca.explained_variance_ratio_),'r', label=\"cumulative var\");\nax[0].set_xlabel('Principal component index'); plt.ylabel('explained variance ratio %');\nax[0].legend()\n\nax[1].bar(range(0,3), pca.explained_variance_ratio_[:3], label=\"individual var\");\nax[1].set_xlabel('Principal component index'); plt.ylabel('explained variance ratio %');\nax[1].legend()\n\nprint(data_pca_skl.shape)","metadata":{"id":"USvkN0WtlOFf","execution":{"iopub.status.busy":"2023-04-27T15:30:31.912196Z","iopub.status.idle":"2023-04-27T15:30:31.913013Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<font color=#6698FF> We notice that the first 10 PCA components can already explain almost all of the variance of the dataset. Since we want to reduce the amount of features as much as possible to reduce the complexity of the models, we use the first 10 PCA components","metadata":{}},{"cell_type":"code","source":"# Train-test Split\nX_train, X_test, Y_train, Y_test = train_test_split(data_pca_skl, ttf_average, test_size=0.3, shuffle=True, random_state=102)","metadata":{"id":"bfMfK-UOrLUz","execution":{"iopub.status.busy":"2023-04-27T15:30:31.918908Z","iopub.status.idle":"2023-04-27T15:30:31.919780Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Generate & normalize test data\ntest_data = get_test_set(0).acoustic_data.values\nfeatures = feature_extractor(test_data)\n\n# Extract features of the segments of the test data\nfor i in range(len(test_samples)-1):\n    test_data = get_test_set(i+1).acoustic_data.values\n    feature = feature_extractor(test_data)\n    features = np.vstack((features, feature))\n\n# Normalize test data    \nscaler = preprocessing.StandardScaler().fit(features)\ndata_test = scaler.transform(features)\ndata_test_pca = pca.fit_transform(data_test)","metadata":{"id":"r1fjqtW4Nyxa","execution":{"iopub.status.busy":"2023-04-27T15:30:31.920739Z","iopub.status.idle":"2023-04-27T15:30:31.921454Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"\n## 3. Training and Results\n<font color=#6698FF> Part of the assignment is to use one linear and one non-linear model. For the linear model, Linear Regression was chosen considering the time to failure is not discrete. A linear model is relatively simple, meaning that overfitting is a lot less likely. A problem is that the data is unlikely to be linear, meaning the accuracy will be lower than a non-linear model.\n    \n<font color=#6698FF> For the non-linear model, Random Forest Regression was chosen. The reason this model was chosen is due to good properties regarding noise on the input data: the input data is a physical signal meaning there is noise present which affects the model. One risk of using this model is overfitting: the Random Forest approach can lead to overfitting, albeit less than the individual decision trees inside the Random Forest.","metadata":{"id":"IUcc2ue7yrOv"}},{"cell_type":"code","source":"# Linear Regression\nmodel = LinearRegression()\nmodel.fit(X_train, Y_train)\n\n# Find predictions for the linear model\npredictions_lr = model.predict(data_test_pca)\n\n# Save predictions as results-lr.csv in the correct format for Kaggle\npredictions_df = pd.DataFrame({'seg_id' : test_samples, 'time_to_failure' : predictions_lr})\npredictions_df.reset_index(drop=True, inplace=True)\npredictions_df.to_csv(\"results-lr.csv\", index=False)","metadata":{"execution":{"iopub.status.busy":"2023-04-27T15:30:31.937403Z","iopub.status.idle":"2023-04-27T15:30:31.938367Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Random Forest Approach\nregr = RandomForestRegressor(n_estimators=110,random_state=102)\nregr.fit(X_train, Y_train)\n\n# Find predictions for the non-linear model\npredictions_rf = regr.predict(data_test_pca)\n\n# Save predictions as results-rf.csv in the correct format for Kaggle\npredictions_df = pd.DataFrame({'seg_id' : test_samples, 'time_to_failure' : predictions_rf})\npredictions_df.reset_index(drop=True, inplace=True)\npredictions_df.to_csv(\"results-rf.csv\", index=False)","metadata":{"id":"0MhqiB72PIYy","outputId":"6daca273-f725-4ed3-8e15-4896d7cd4483","execution":{"iopub.status.busy":"2023-04-27T15:30:31.939352Z","iopub.status.idle":"2023-04-27T15:30:31.940089Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Visualize accuracy for both models on training data\npredictions_lr = model.predict(X_test)\npredictions_rf = regr.predict(X_test)\n\nfig, ax = plt.subplots(figsize=(20,5))\n\nax.plot(ttf_average,label='Average TTF')\nax.set_title(\"Plot of the predictions of the model and the total training data\")\n\n# Find the values for the given training data (thus: X_train + X_test ordered in time)\nax.plot(model.predict(data_pca_skl),label='Linear Regression')\nax.plot(regr.predict(data_pca_skl),label='Random Forest')\nax.legend()\n\n# Find the MSE on the training set and on the test set\nmse_lr_test = mean_squared_error(Y_test, predictions_lr)\nmse_rf_test = mean_squared_error(Y_test, predictions_rf)\nmse_lr_total_train = mean_squared_error(Y_train,model.predict(X_train))\nmse_rf_total_train = mean_squared_error(Y_train,regr.predict(X_train))\n\nprint(\"MSE for the Linear Regression model on the training set:\",mse_lr_total_train)\nprint(\"MSE for the Random Forest model on the training set:\",mse_rf_total_train)\nprint(\"MSE for the Linear Regression model on the test set:\",mse_lr_test)\nprint(\"MSE for the Random Forest model on the test set:\",mse_rf_test)","metadata":{"execution":{"iopub.status.busy":"2023-04-27T15:30:31.945033Z","iopub.status.idle":"2023-04-27T15:30:31.945889Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## 4. Discussion and Conclusion\n<font color=#6698FF> The final results of both models when submitted on Kaggle are as follows:\n    \n<font color=#6698FF> Linear regression: Private score of 3.33533, Public score of 2.77743\n    \n<font color=#6698FF> Random forest: Private score of 2.78641, Public score of 2.21888\n    \n<font color=#6698FF> We notice that the results are as expected from a linear and a non-linear model: the non-linear model scores (slightly) better.\n    \n<font color=#6698FF> One very important thing to notice in the review of the accuracy is that of particularly the non-linear model. When we compare the MSE of on the training set and the test set, we notice a very large difference: the MSE for the training set is a lot lower than the MSE on the test set. This is a sign of overfitting, which is likely to happen when using a Random Forest. To obtain better results for the test set, a different model could be used that is less prone to overfitting: an option could be a Neural Network. This will be slightly more sensitive to noise, but would probably result in a lower MSE and thus a higher score if tuned properly. Note that the overfitting does not occur for the linear model as expected.\n    \n<font color=#6698FF> As a final review: we first split the data into segments of 150k samples and took the mean value of the time to failure of this interval, and performed feature extraction on each of these segments. We used PCA to reduce the amount of features since a lot of these contribute very little to the total variance and created two models: linear regression for the linear model, and a random forest for the non-linear model. We used a 70/30 split of the given training set to predict the accuracy of both models. We noticed that the non-linear model performed slightly better on the testing set, but way better on the training set when compared to the linear model. The submissions on Kaggle also show that the non-linear model performs better than the linear model.\n    \n<font color=#6698FF> The model could be improved in the following ways:\n* <font color=#6698FF> More distinct features could be used. We notice that PCA can already explain 90% of the variance using around 8 features, meaning that more distinct features can greatly improve the accuracy of both of the models. A lot of the features that are used differ very little from eachother, which limits the accuracy that can be obtained.\n* <font color=#6698FF> Use a better non-linear model. The Random Forest approach tends to overfit easily, causing the score to be limited. This is very noticeable in the MSE for the test and the training set. We mentioned earlier that using a Neural Network or a different non-linear model could improve the accuracy. The reason a Random Forest was chosen in the first place was since it was felt that this was the best understood non-linear model so far: we did not want to use models which we were not familiar with (e.g. the light gradient-boosting machine used in https://www.kaggle.com/code/gpreda/lanl-earthquake-eda-and-prediction/notebook and https://www.kaggle.com/code/gpreda/lanl-earthquake-eda-and-prediction/notebook).\n* <font color=#6698FF> Tune the hyperparameters of the model better. This was something which was attempted but resulted in multiple issues. If the number of trees in the forest was changed, the MSE of the test data increased: too large and the overfitting would increase, too small and the complexity was not high enough). Since the depth of the forest decides the output range, this also resulted in some issues: if we decreased it the MSE of the test set would also increase, thus it was left on unlimited (to be decided by the SKLearn function). The rest of the hyperparameters were not explored enough: perhaps adapting these could lead to a better accuracy on the test data.   ","metadata":{"id":"yekgnxu7y0gE"}}]}