{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.7.12","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":11000,"databundleVersionId":875412,"sourceType":"competition"}],"dockerImageVersionId":30458,"isInternetEnabled":false,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Final Project\n\nAs in all machine learning problems you should complete the following steps:\n1. Load and explore the data (using plots and histograms )\n2. Clean/preprocess/transform the data if necessary (first performed on training data and next on test data)\n3. Train the machine learning model (in this assignment two models, one linear and one non-linear)\n4. Evaluate and optimise the model\n\nHowever, while performing the steps you should consider our main goals and research questions for doing this project and try to adress them. This can be either in each step or afterwards in discussion and conclusion section (your call!). Below are the research question we are interested in:\n\n* What are the necessary step to clean and prepare the dataset as we have categorial features and missing data.\n* What model/classifier provides the best result for this application.\n* What is the best choice for our cost function and performance metrics for this problem.\n\nPlease also note the following points during the assignment:\n\n* Use functions from open source libraries like sci-kit learn and keras and avoid using your own hand-written functions from previous assignments. You can also use our [cheatsheet](https://colab.research.google.com/drive/12h-QBlsaWXkjGIRoJXfF4yqi1elnX9qn?usp=sharing).\n* Feel free to contact us on Teams if you need more description.\n* This notebook is structured like a scientific paper. The text should provide a high-level overview of your approach. Please don't include any details about your code in the text but add them as comments in the code itself. Your code should be cleane and readable with enough comments.\n* There are some instructions and questions in each section, remove the highlighted text in blue and replace it with your explanations and answers.\n\nYou can delete this section before submission.","metadata":{"id":"GF4WsCxg-lL9"}},{"cell_type":"markdown","source":"## 1. Introduction\n","metadata":{"id":"3NnAgB-y0knj"}},{"cell_type":"markdown","source":"<font color=#6698FF> Tycho Brouwer, score and Leaderboard rank here.\n","metadata":{"id":"DPJV5e8sDVnT"}},{"cell_type":"markdown","source":"## 2. Data\n","metadata":{"id":"PB7FLdQn-dmK"}},{"cell_type":"code","source":"import matplotlib.pyplot as plt\nfrom sklearn import metrics\nfrom sklearn import preprocessing\nfrom sklearn.metrics import balanced_accuracy_score, mean_absolute_error, mean_squared_error\nfrom sklearn.decomposition import PCA\nfrom sklearn.ensemble import RandomForestRegressor \nfrom sklearn.linear_model import LinearRegression, Ridge, Lasso, ElasticNet, BayesianRidge, Perceptron\nfrom sklearn.model_selection import train_test_split\nfrom sklearn.svm import SVR, NuSVR\nfrom sklearn.preprocessing import StandardScaler\nimport seaborn as sns\n%matplotlib inline ","metadata":{"execution":{"iopub.status.busy":"2024-04-22T19:15:22.216080Z","iopub.execute_input":"2024-04-22T19:15:22.217661Z","iopub.status.idle":"2024-04-22T19:15:23.783069Z","shell.execute_reply.started":"2024-04-22T19:15:22.217597Z","shell.execute_reply":"2024-04-22T19:15:23.781539Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### 2.1 Dataset\n<font color=#6698FF>In this section, we load and explore the dataset.","metadata":{"id":"vfGikOxUAxJB"}},{"cell_type":"code","source":"# This Python 3 environment comes with many helpful analytics libraries installed\n# It is defined by the kaggle/python docker image: https://github.com/kaggle/docker-python\n# For example, here's several helpful packages to load in \n\nimport numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\n\n# Input data files are available in the \"../input/\" directory.\n# For example, running this (by clicking run or pressing Shift+Enter) will list the files in the input directory\n\nimport os\nfor dirname, _, filenames in os.walk('/kaggle/input'):\n    for filename in filenames:\n        os.path.join(dirname, filename)\n\n# You can write up to 20GB to the current directory (/kaggle/working/) that gets preserved as output when you create a version using \"Save & Run All\" \n# You can also write temporary files to /kaggle/temp/, but they won't be saved outside of the current session","metadata":{"execution":{"iopub.status.busy":"2024-04-22T19:15:23.785247Z","iopub.execute_input":"2024-04-22T19:15:23.786061Z","iopub.status.idle":"2024-04-22T19:15:24.969048Z","shell.execute_reply.started":"2024-04-22T19:15:23.786019Z","shell.execute_reply":"2024-04-22T19:15:24.967262Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"###### 2.1.1 Train-test split\n<font color=#6698FF>In the below, we split the train data into a test and a train set. Set a value for the `test_size` yourself. Argue why the test value can not be too small or too large. You can also use k-fold cross validation.\n","metadata":{"id":"udqAFg1C3JHe"}},{"cell_type":"code","source":"train = pd.read_csv('../input/LANL-Earthquake-Prediction/train.csv', \n                    dtype={'acoustic_data': np.int16, 'time_to_failure':np.float64})\n\ndata_train, data_test = train_test_split(train, test_size=0.3, shuffle=False)","metadata":{"execution":{"iopub.status.busy":"2024-04-22T19:15:24.974335Z","iopub.execute_input":"2024-04-22T19:15:24.975196Z","iopub.status.idle":"2024-04-22T19:19:27.740620Z","shell.execute_reply.started":"2024-04-22T19:15:24.975134Z","shell.execute_reply":"2024-04-22T19:19:27.738971Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"usually the data is split into 70% for the training and 30% for validation. this is in order to train the model good but also being able to test it and validate that it is indeed a good model. ","metadata":{}},{"cell_type":"code","source":"train.shape","metadata":{"execution":{"iopub.status.busy":"2024-04-22T19:24:38.767040Z","iopub.execute_input":"2024-04-22T19:24:38.767885Z","iopub.status.idle":"2024-04-22T19:24:38.778371Z","shell.execute_reply.started":"2024-04-22T19:24:38.767812Z","shell.execute_reply":"2024-04-22T19:24:38.776837Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### 2.2 Data Exploration\n\n<font color=#6698FF> Explore the features and target variables of the dataset. Think about making some scatter plots, box plots, histograms or printing the data, but feel free to choose any method that suits you.\nWhat do you think is the right performance\nmetric to use for this dataset? Clearly explain which performance metric you\nchoose and why.\nAlgorithmic bias can be a real problem in Machine Learning.  Explain what you believe.</font>","metadata":{"id":"QN5jS2_p_Y9N"}},{"cell_type":"code","source":"train_acoustic_df = train['acoustic_data'].values[::100]\ntrain_acoustic_df.shape\n\ntrain_time_to_failure_df = train['time_to_failure'].values[::100]\n\n\n\ndef plot_acc_ttf_data(train_acoustic_df,train_time_to_failure_df,title='Acoustic data and time to failure: sampled 1%'):\n    fig ,ax1 = plt.subplots(figsize=(12,9))\n    plt.title(title)\n    plt.plot(train_acoustic_df,color='r')\n    ax1.set_ylabel('train_acoustic_df',color='r')\n    plt.legend(['acoustic-data'],loc=(0.01,0.95))\n    ax2 = ax1.twinx()\n    plt.plot(train_time_to_failure_df,color='b')\n    ax2.set_ylabel('time to failure',color='b')\n    plt.legend(['time_to_failure'],loc=(0.01,0.9))\n    plt.grid(True)\n    \nplot_acc_ttf_data(train_acoustic_df,train_time_to_failure_df)","metadata":{"execution":{"iopub.status.busy":"2024-04-22T19:28:01.040051Z","iopub.execute_input":"2024-04-22T19:28:01.040607Z","iopub.status.idle":"2024-04-22T19:28:05.287118Z","shell.execute_reply.started":"2024-04-22T19:28:01.040564Z","shell.execute_reply":"2024-04-22T19:28:05.285776Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"the peak in the acoustic data results in a jump in the time to failure","metadata":{}},{"cell_type":"code","source":"del train_time_to_failure_df\ndel train_acoustic_df","metadata":{"execution":{"iopub.status.busy":"2024-04-22T19:30:53.726334Z","iopub.execute_input":"2024-04-22T19:30:53.727017Z","iopub.status.idle":"2024-04-22T19:30:53.734336Z","shell.execute_reply.started":"2024-04-22T19:30:53.726955Z","shell.execute_reply":"2024-04-22T19:30:53.732791Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_acoustic_df = train['acoustic_data'].values[:6291455]\ntrain_time_to_failure_df = train['time_to_failure'].values[:6291455]","metadata":{"execution":{"iopub.status.busy":"2024-04-22T19:31:02.039465Z","iopub.execute_input":"2024-04-22T19:31:02.039977Z","iopub.status.idle":"2024-04-22T19:31:02.047320Z","shell.execute_reply.started":"2024-04-22T19:31:02.039932Z","shell.execute_reply":"2024-04-22T19:31:02.045819Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plot_acc_ttf_data(train_acoustic_df,train_time_to_failure_df)","metadata":{"execution":{"iopub.status.busy":"2024-04-22T19:33:16.663603Z","iopub.execute_input":"2024-04-22T19:33:16.665293Z","iopub.status.idle":"2024-04-22T19:33:19.838695Z","shell.execute_reply.started":"2024-04-22T19:33:16.665228Z","shell.execute_reply":"2024-04-22T19:33:19.837562Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"the time to failure is just after the acoustic data is measured ","metadata":{}},{"cell_type":"markdown","source":"### 2.3 Data Preparation\n\n<font color=#6698FF>  This dataset hasn’t been cleaned yet. Meaning that some attributes (features) are in numerical format and some are in categorial format. Moreover, there are missing values as well. However, all Scikit-learn’s implementations of these algorithms expect numerical features. Check for all features if they are in categorial and use a method to transform them to numerical values. For the numerical data, handle the missing data and normalize the data. \nNote that you are only allowed to use training data for preprocessing but you then need to perform similar changes on test data too.\nYou can use [pipelining](https://scikit-learn.org/stable/modules/generated/sklearn.pipeline.Pipeline.html) to help with the preprocessing.</font>","metadata":{"id":"4A5GsVzgATRu"}},{"cell_type":"code","source":"rows = 150_000\nsegments = int(np.floor(data_train.shape[0] / rows))","metadata":{"execution":{"iopub.status.busy":"2024-04-22T20:21:59.093010Z","iopub.execute_input":"2024-04-22T20:21:59.093519Z","iopub.status.idle":"2024-04-22T20:21:59.099720Z","shell.execute_reply.started":"2024-04-22T20:21:59.093475Z","shell.execute_reply":"2024-04-22T20:21:59.098342Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def extract_features_from_segment(segment, seg, X):\n    x = pd.Series(seg['acoustic_data'].values)\n\n    X.loc[segment, 'mean'] = x.mean()\n    X.loc[segment, 'std'] = x.std()\n    X.loc[segment, 'max'] = x.max()\n    X.loc[segment, 'min'] = x.min()\n    X.loc[segment, 'kurt'] = x.kurtosis()\n    X.loc[segment, 'skew'] = x.skew()\n    \n#Rolling window of size 100\n    x_roll_std = x.rolling(window = 100).std().dropna().values\n    x_roll_mean = x.rolling(window= 100).mean().dropna().values\n    \n    #Features from rolling mean\n    X.loc[segment, 'rolling_mean'] =x_roll_mean.mean()\n    X.loc[segment, 'rolling_mean'] =x_roll_mean.std()\n    \n    #Features from rolling std\n    X.loc[segment, 'rolling_std'] = x_roll_std.mean()\n    X.loc[segment, 'rolling_std'] = x_roll_std.std()","metadata":{"execution":{"iopub.status.busy":"2024-04-22T20:24:04.971897Z","iopub.execute_input":"2024-04-22T20:24:04.973404Z","iopub.status.idle":"2024-04-22T20:24:04.984223Z","shell.execute_reply.started":"2024-04-22T20:24:04.973323Z","shell.execute_reply":"2024-04-22T20:24:04.982488Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"X_train = pd.DataFrame(index=range(segments), dtype=np.float64)\ny_train = pd.DataFrame(index=range(segments), dtype=np.float64, columns=['time_to_failure'])\nfor segment in range(segments):\n    seg = data_train.iloc[segment*rows:segment*rows+rows]\n    extract_features_from_segment(segment, seg, X_train)\n    y_train.loc[segment, 'time_to_failure'] = seg['time_to_failure'].values[-1]","metadata":{"execution":{"iopub.status.busy":"2024-04-22T20:24:07.708878Z","iopub.execute_input":"2024-04-22T20:24:07.710280Z","iopub.status.idle":"2024-04-22T20:25:04.433707Z","shell.execute_reply.started":"2024-04-22T20:24:07.710206Z","shell.execute_reply":"2024-04-22T20:25:04.431799Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"scaler = StandardScaler()\nscaler.fit(X_train)\nX_train_scaled = scaler.transform(X_train)\n\n\ntest_segments = int(np.floor(data_test.shape[0] / rows))\nX_test = pd.DataFrame(index=range(test_segments), dtype=np.float64)\ny_test = pd.DataFrame(index=range(test_segments), dtype=np.float64, columns=['time_to_failure'])\nfor segment in range(test_segments):\n    seg = data_test.iloc[segment*rows:segment*rows+rows]\n    extract_features_from_segment(segment, seg, X_test)\n    y_test.loc[segment, 'time_to_failure'] = seg['time_to_failure'].values[-1]\nnew_scaler = StandardScaler()\nnew_scaler.fit(X_test)\nX_test_scaled = new_scaler.transform(X_test)","metadata":{"execution":{"iopub.status.busy":"2024-04-22T20:25:29.134122Z","iopub.execute_input":"2024-04-22T20:25:29.134701Z","iopub.status.idle":"2024-04-22T20:25:53.354794Z","shell.execute_reply.started":"2024-04-22T20:25:29.134640Z","shell.execute_reply":"2024-04-22T20:25:53.353253Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"\n## 3. Training and Results\n<font color=#6698FF> Briefly introduce algorithms you choose. \nPresent your balanced accuracies for both test and training data for all classifiers. Analyse the performance on test and training in terms of bias and variance. Give one advantage and one drawback of the method you use.\n","metadata":{"id":"IUcc2ue7yrOv"}},{"cell_type":"code","source":"# define the model\nmodel = LinearRegression()\n# fit the model\nmodel.fit(X_train_scaled, y_train.values.flatten())\n# predict\npred=model.predict(X_train_scaled)\nscore = model.score(X_train_scaled, y_train)\nmae = mean_absolute_error(y_train.values.flatten(), pred)\nmse = mean_squared_error(y_train.values.flatten(), pred)\nprint('score of the linearRegression model: ' + str(score))\nprint('lineaRegression moddel train mse: ' + str(mse))\nprint('lineaRegression moddel train mae: ' + str(mae))\n\n\nplt.scatter(y_train.values.flatten(), pred, label = 'train')\nplt.xlabel('Actual values')\nplt.ylabel('Predicted values')\nplt.legend()\n\n# define the model\nmodel = LinearRegression()\n# fit the model\nmodel.fit(X_train_scaled, y_train.values.flatten())\n# predict\npred=model.predict(X_test_scaled)\nscore = model.score(X_test_scaled, y_test)\nmae = mean_absolute_error(y_test.values.flatten(), pred)\nmse = mean_squared_error(y_test.values.flatten(), pred)\nprint('score of the linearRegression model: ' + str(score))\nprint('lineaRegression moddel test mse: ' + str(mse))\nprint('lineaRegression moddel test mae: ' + str(mae))\n\nplt.scatter(y_test.values.flatten(), pred, label = 'test')\nplt.xlabel('Actual values')\nplt.ylabel('Predicted values')\nplt.legend()","metadata":{"execution":{"iopub.status.busy":"2024-04-22T20:34:53.137836Z","iopub.execute_input":"2024-04-22T20:34:53.138641Z","iopub.status.idle":"2024-04-22T20:34:53.554742Z","shell.execute_reply.started":"2024-04-22T20:34:53.138589Z","shell.execute_reply":"2024-04-22T20:34:53.553128Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"\nmodel = RandomForestRegressor()\nmodel.fit(X_train_scaled, y_train.values.flatten())\npred = model.predict(X_train_scaled)\nMAE = mean_absolute_error(y_train.values.flatten(), pred)\nMSE = mean_squared_error(y_train.values.flatten(), pred)\nscore = model.score(X_train_scaled, y_train)\nprint('regression forest train mae: ', MAE)\nprint('regression forest train mse:', MSE)\nprint(' score regression forest:', score )\nplt.scatter(y_train.values.flatten(), pred, label='train')\nplt.xlabel('Actual values')\nplt.ylabel('Predicted values')\nplt.legend()\n    \n\n\nmodel = RandomForestRegressor()\nmodel.fit(X_test_scaled, y_test.values.flatten())\npred = model.predict(X_test_scaled)\nMAE = mean_absolute_error(y_test.values.flatten(), pred)\nMSE = mean_squared_error(y_test.values.flatten(), pred)\nscore = model.score(X_test_scaled, y_test)\nprint('regression forest test mae: ', MAE)\nprint('regression forest test mse:', MSE)\nprint(' score regression forest:', score )\nplt.scatter(y_test.values.flatten(), pred, label='test')\nplt.xlabel('Actual values')\nplt.ylabel('Predicted values')\nplt.legend()","metadata":{"execution":{"iopub.status.busy":"2024-04-22T20:42:14.460304Z","iopub.execute_input":"2024-04-22T20:42:14.460861Z","iopub.status.idle":"2024-04-22T20:42:17.614620Z","shell.execute_reply.started":"2024-04-22T20:42:14.460815Z","shell.execute_reply":"2024-04-22T20:42:17.613169Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## 4. Discussion and Conclusion\n\n<font color=#6698FF>Discuss all the choices you made during the process and your final conclusions.  Highlight the strong points of your approach, discuss its shortcomings  and suggest some future approaches that may improve it. Please be self critical here. The assignment is not about achieving a state of the art performance, but about showing what you have learned the concepts during the course.</font>","metadata":{"id":"yekgnxu7y0gE"}},{"cell_type":"markdown","source":"there was a lot of data available to use so first I had to devide them into segments inorder to have smaller sets of data. Than the features were extracted for each segment. As last the models were made with a linear one and a non linear, randomforestregressor. The mse of the linear model is suprisingly high in comparison with the non linear one, the mae is hover a bit lower. The last noticeable outcome is the score of the linear model, this is really low. So overall the non linear model is in this case way better. ","metadata":{}},{"cell_type":"markdown","source":"## 5. References","metadata":{"id":"wn7PBEmWESIi"}},{"cell_type":"markdown","source":"notebooks:\nhttps://www.kaggle.com/code/gpreda/lanl-earthquake-eda-and-prediction/notebook\n\nhttps://www.kaggle.com/code/inversion/basic-feature-benchmark\n\nhttps://www.kaggle.com/code/jessteijn/ml23-week8\n\nhttps://www.kaggle.com/code/dineshm2/lanl-earthquake-prediction\n","metadata":{}},{"cell_type":"code","source":"submission = pd.read_csv('/kaggle/input/LANL-Earthquake-Prediction/sample_submission.csv', index_col='seg_id')\nX_test = pd.DataFrame(columns=X_train.columns, dtype=np.float64, index=submission.index)\nfor seg_id in X_test.index:\n    seg = pd.read_csv('/kaggle/input/LANL-Earthquake-Prediction/test/' + seg_id + '.csv')\n    extract_features_from_segment(seg_id, seg, X_test)\n    \nX_test_scaled = scaler.transform(X_test)\nsubmission['time_to_failure'] = model.predict(X_test_scaled)\nsubmission.to_csv('submission.csv')","metadata":{"execution":{"iopub.status.busy":"2024-04-22T20:55:52.570059Z","iopub.execute_input":"2024-04-22T20:55:52.570871Z","iopub.status.idle":"2024-04-22T20:57:45.930483Z","shell.execute_reply.started":"2024-04-22T20:55:52.570812Z","shell.execute_reply":"2024-04-22T20:57:45.929318Z"},"trusted":true},"execution_count":null,"outputs":[]}]}