{"cells":[{"metadata":{"_uuid":"9cf04d16a9c413fa83cded79243014db9094c8ce"},"cell_type":"markdown","source":"# Background\n\nIn Kaggle competitions visualising and understanding data is often as important as machine learning techniques applied to it. Having deep insights into the data can help to design better features or even spot leaks in the underlying dataset. In this competition we are dealing with a 1-dimensial highly granular signal. Total amout of datapoints is almost a billion, it's easy to miss the forest for the trees.\n\nOne of the most effective techniques in signal processing is DFT (Discrete Fourier Transform). It allows to reduce granularity of the signal across time axis and shift it into frequency dimension. To compact representation further one can apply log filter banks, that combine adjucent frequencies into sub-bands."},{"metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true},"cell_type":"code","source":"import numpy as np\nimport pandas as pd\n\nimport matplotlib.pyplot as plt\nimport seaborn as sns\n\n!pip install python_speech_features\nfrom python_speech_features import logfbank","execution_count":null,"outputs":[]},{"metadata":{"_cell_guid":"79c7e3d0-c299-4dcb-8224-4455121ee9b0","_uuid":"d629ff2d2480ee46fbb7e2d37f6b5fab8052498a","trusted":true,"scrolled":false},"cell_type":"code","source":"train_df = pd.read_csv('../input/train.csv', dtype={'acoustic_data': np.int16, 'time_to_failure': np.float16})\ntrain_df['quake_id'] = (train_df.time_to_failure.diff() > 0.0).cumsum().astype('int16')\ntrain_df.head()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"ae9f345a8b6d14544a7584af7b948fdad499ab3c"},"cell_type":"markdown","source":"# Parsing the signal\nFrom my analysis [here](https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/85350) the sensor works in the following way:\n* It makes measurements in bursts.\n* Each burst has 4096 (in rare cases 4095) measurements (separated by around 1 nanosecond).\n* Measurement bursts happen at a frequency around 1KHz (1000 bursts per second).\n\nBased on the information above it is natural to split signal around the burst boundaries and transform each burst into the frequency domain.\n\n**Note**: I skip the first and last quakes as they are not complete."},{"metadata":{"trusted":true,"_uuid":"5f0025f2c548254b3a5438a49e9288821fd9e1c4"},"cell_type":"code","source":"lfbanks = []\nquake_cnt = 15\nfor quake_id in range(1, quake_cnt + 1):\n    quake_df = train_df.loc[train_df.quake_id == quake_id]\n    burst_id = (quake_df.time_to_failure.diff() < -1e-4).cumsum().astype('int16')\n    lfbank = quake_df.groupby(burst_id).apply(lambda x: logfbank(\n        x.acoustic_data.values, samplerate=4096, winlen=1.0, winstep=1.0, nfilt=32, nfft=4096, highfreq=512\n    ))\n    lfbank = np.vstack(lfbank.values)\n    lfbanks.append(lfbank)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"fea6e636a7539e107a717356dfbe2e90a1690afb"},"cell_type":"code","source":"fig, axes = plt.subplots(quake_cnt, 1, figsize=(20, 4*quake_cnt))\nfig.subplots_adjust(hspace=0.3)\nfor idx in range(1, quake_cnt + 1):\n    sns.heatmap(lfbanks[idx-1].T, ax=axes[idx-1]).set_title('Quake #{}'.format(idx))","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"c3270c2604d3d2748127bef0446f77aac32dd008"},"cell_type":"markdown","source":"# Summary\n\n* The amplitude of the frequencies increases the closer we are to the earthquake event.\n* The amplitude changes are noisy over short time periods, making it hard to predict time to quake based on samples in the test data.\n* Log filter banks can be used as additional input features to your machine learning algorithm. They don't require any tuning and have compact representation.\n\nHappy Kaggling!"},{"metadata":{"trusted":true,"_uuid":"2d98fbda27971a52149841c1bcec155055b5f009"},"cell_type":"code","source":"","execution_count":null,"outputs":[]}],"metadata":{"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"},"language_info":{"name":"python","version":"3.6.6","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"}},"nbformat":4,"nbformat_minor":1}