{"cells":[{"metadata":{"_uuid":"3b29b370d8d261cf6d93330bdc35efe86762e1c1"},"cell_type":"markdown","source":"# Predict time of earthquake with only one feature"},{"metadata":{"_uuid":"cd4dfe550e48a4449c82928294f4bfffc543bee9"},"cell_type":"markdown","source":"This kernel aims at finding a solution that makes sense by using basic models and features.\nI find it useful to understand roughly how the 'acoustic_data' signal is behaving.\nWe are using here the standard deviation of 'acoustic_data' as the only predictive feature."},{"metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"import numpy as np\nimport pandas as pd\nimport os\nimport matplotlib.pyplot as plt\n%matplotlib inline","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"712f85a2381bc3051822e310b4fa40b0050a7154"},"cell_type":"markdown","source":"Let's check the format of the train data."},{"metadata":{"trusted":true,"_uuid":"079164a14d30b98dc233ae99faaad18bb079135c","_kg_hide-input":true},"cell_type":"code","source":"df_train_10M = pd.read_csv('../input/train.csv', nrows=200000000, dtype={'acoustic_data': np.int16, 'time_to_failure': np.float64})\ndf_train_10M.head()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"861088fc79c572b74c0dd52906690f34e4fbfce8"},"cell_type":"markdown","source":"As the submission consists of samples of 150,000 data points, we divide the train set into samples of 150,000 data points as well. \nFor each of these small samples, we calculate the standard deviation of the input data \"acoustic_data\".\nLet's assume that the \"time_to_failure\" can be considered as nearly constant within 150,000 data points, we calculate the mean of \"time_to_failure\" for each of these smaller samples."},{"metadata":{"trusted":true,"_uuid":"0938d2a275af651b31f8c266e20ee5f748d150d8","_kg_hide-input":true},"cell_type":"code","source":"df_train_10M[\"index_obs\"] = (df_train_10M.index.astype(float)/150000).astype(int)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"c097632ecb7728d517198414d560c2818cd5fd8e","_kg_hide-input":true},"cell_type":"code","source":"train_set = df_train_10M.groupby('index_obs').agg({'acoustic_data': 'std', 'time_to_failure': 'mean'})","execution_count":null,"outputs":[]},{"metadata":{"_kg_hide-input":true,"trusted":true,"_uuid":"34e76d858395afed6f270c8e194d48f2f40034b7"},"cell_type":"code","source":"train_set.columns = ['acoustic_data_std', 'time_to_failure_mean']\ntrain_set.head()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"5f142670844a80d206a41d3157c17d678eeeab3d"},"cell_type":"markdown","source":"We delete the original training dataset to free up some memory."},{"metadata":{"trusted":true,"_uuid":"1593c8edd4f60c1393d73cf560ad6abef631c72f","_kg_hide-input":true},"cell_type":"code","source":"del df_train_10M","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"dd7bfeaf50db66001a6dc89d655420969133aec5"},"cell_type":"markdown","source":"We check at which indices the earthquakes happen. These are the indices when the difference between 2 'time_to_failure' is positive."},{"metadata":{"trusted":true,"_uuid":"65a8cc3147823b570961bafabf6119406e906239"},"cell_type":"code","source":"train_set[train_set['time_to_failure_mean'].diff() > 0].head()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"21c87c51ee77c737a4f99f28d5d6063dc36241f3"},"cell_type":"markdown","source":"We smooth out a bit the STD of \"acoustic_data\" using rolling mean."},{"metadata":{"trusted":true,"_uuid":"f05db465e5d37a51ef24d15ae76e86d632e8fff3"},"cell_type":"code","source":"train_set[\"acoustic_data_transform\"] = train_set[\"acoustic_data_std\"].clip(-20, 12).rolling(10, min_periods=1).median()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"55c7f60701df299e5d2195f23e096d7606fcaa74"},"cell_type":"markdown","source":"We plot the STD of 'acoustic_data' over time, the 3 red vertical lines correspond to the 3 first earthquakes.\nAt first glance, you can see that the (smoothed) STD of 'acoustic_data' approximately increases linearly until the time at which the earthquakes happen."},{"metadata":{"trusted":true,"_uuid":"6cb0e5d892a25396a3c44124270d9f5a964552d6"},"cell_type":"code","source":"fig, ax1 = plt.subplots(figsize=(20,12))\nax2 = ax1.twinx()\nax1.plot(train_set[\"acoustic_data_transform\"])\nplt.axvline(x=38, color='r')\nplt.axvline(x=334, color='r')\nplt.axvline(x=697, color='r')\nplt.title('Smoothed standard deviation of acoustic_data vs. time', size=20)\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"71ac59e5ca36ae29b690d570e773ecffb6766c23"},"cell_type":"markdown","source":"Following our first intuition, if \"acoustic_data_std\" is a linear function of  \"time_to_failure_mean\", as we can see on the graph, then \"time_to_failure_mean\" is also a linear function of \"acoustic_data_std\".\n\nLet's use a simple linear regressor to predict the \"time_to_failure_mean\" from the smooth version of \"acoustic_data_std\"."},{"metadata":{"trusted":true,"_uuid":"09b993c7a3834ed4b82e07738ba07e58ab75721e"},"cell_type":"code","source":"from sklearn.linear_model import LinearRegression\nregr = LinearRegression()\nregr.fit(train_set[[\"acoustic_data_transform\"]], train_set[\"time_to_failure_mean\"])\n\nprint('Coefficients: \\n', regr.coef_)\nprint('Intercept: \\n', regr.intercept_)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"ee3ba5c057c220a4300a7cdea67d3068d0ade0d5"},"cell_type":"markdown","source":"Let's predict the results on test set and submit the results."},{"metadata":{"trusted":true,"_uuid":"9dc47f9f09af669c16140138870ee22082660089"},"cell_type":"code","source":"submission_file = pd.read_csv('../input/sample_submission.csv')\nsubmission_file.head()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"b76c33a3a5a47ac2ca2ca06ae0d76f95951eb2e2"},"cell_type":"code","source":"for index, seg_id in enumerate(submission_file['seg_id']):\n    seg = pd.read_csv('../input/test/' + str(seg_id) + '.csv')\n    x = seg['acoustic_data'].values\n    std_x = max(-20, min(12, np.std(x)))\n    submission_file.loc[index, \"time_to_failure\"] = max(0, regr.intercept_ + regr.coef_ * std_x)\n    del seg","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"8288fd76f0ed447ed5b0bbe3af7df9b332994de0"},"cell_type":"code","source":"submission_file.to_csv('submission.csv', index=False)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"a778efdfdf8729095844efc689f309db2481fec1"},"cell_type":"code","source":"","execution_count":null,"outputs":[]}],"metadata":{"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"},"language_info":{"name":"python","version":"3.6.6","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"}},"nbformat":4,"nbformat_minor":1}