{"cells": [{"cell_type": "markdown", "metadata": {}, "source": "# LANL Earthquake Prediction \u2014 what the acoustic signal looks like\n\nA compact visual EDA: the training data is one long acoustic stream with a `time_to_failure` label that ramps down to each lab quake. This notebook shows the signal/label relationship and the segment-level statistical features that drive the standard solution. Reads only a slice of the 9 GB file so it runs fast."}, {"cell_type": "code", "execution_count": null, "metadata": {}, "source": "import glob, os\nimport numpy as np, pandas as pd\nimport matplotlib.pyplot as plt\n\nBASE = '/kaggle/input/LANL-Earthquake-Prediction'\nif not os.path.exists(f'{BASE}/train.csv'):\n    BASE = os.path.dirname(glob.glob('/kaggle/input/**/train.csv', recursive=True)[0])\n# read a slice spanning roughly one failure cycle\ntr = pd.read_csv(f'{BASE}/train.csv', nrows=15_000_000,\n                 dtype={'acoustic_data': np.int16, 'time_to_failure': np.float32})\nprint(tr.shape)\ntr.head()", "outputs": []}, {"cell_type": "markdown", "metadata": {}, "source": "## Signal vs time-to-failure\n\n`time_to_failure` decreases monotonically toward each quake, then jumps back up. The acoustic amplitude stays quiet, then bursts shortly **before** failure \u2014 that pre-failure activity is the signal we exploit."}, {"cell_type": "code", "execution_count": null, "metadata": {}, "source": "fig, ax1 = plt.subplots(figsize=(13, 4))\nstep = 50  # downsample for plotting\nax1.plot(tr['acoustic_data'].values[::step], color='steelblue', lw=0.4, label='acoustic_data')\nax1.set_ylabel('acoustic_data', color='steelblue')\nax2 = ax1.twinx()\nax2.plot(tr['time_to_failure'].values[::step], color='crimson', lw=1.2, label='time_to_failure')\nax2.set_ylabel('time_to_failure (s)', color='crimson')\nplt.title('Acoustic signal bursts just before each failure (ttf -> 0)')\nplt.show()", "outputs": []}, {"cell_type": "markdown", "metadata": {}, "source": "## The modelling setup: segment-level features\n\nTest files are independent **150,000-sample** segments; for each you predict the `time_to_failure` at the segment end. So the standard approach turns each 150k chunk into a feature vector (summary statistics of the acoustic amplitude) and regresses it onto the chunk's final ttf. Below we build features over the training slice the same way."}, {"cell_type": "code", "execution_count": null, "metadata": {}, "source": "SEG = 150_000\ndef seg_features(x):\n    x = x.astype(np.float64)\n    ax = np.abs(x)\n    return {\n        'mean': x.mean(), 'std': x.std(), 'min': x.min(), 'max': x.max(),\n        'abs_mean': ax.mean(), 'abs_max': ax.max(),\n        'q95': np.quantile(x, 0.95), 'q05': np.quantile(x, 0.05),\n        'q99': np.quantile(x, 0.99), 'q01': np.quantile(x, 0.01),\n        'kurtosis': pd.Series(x).kurt(), 'skew': pd.Series(x).skew(),\n    }\n\nrows, ys = [], []\nad = tr['acoustic_data'].values\nttf = tr['time_to_failure'].values\nfor i in range(0, len(tr) - SEG, SEG):\n    rows.append(seg_features(ad[i:i+SEG]))\n    ys.append(ttf[i+SEG-1])\nfeat = pd.DataFrame(rows); feat['ttf'] = ys\nprint(feat.shape)\nfeat.head()", "outputs": []}, {"cell_type": "code", "execution_count": null, "metadata": {}, "source": "# which simple features track time-to-failure?\ncorr = feat.corr()['ttf'].drop('ttf').sort_values()\ncorr.plot.barh(figsize=(7, 5), title='Correlation of segment features with time_to_failure')\nplt.tight_layout(); plt.show()\nprint('Signal-amplitude features (std, abs_max, high quantiles) carry most of the signal.')", "outputs": []}, {"cell_type": "markdown", "metadata": {}, "source": "## Takeaways\n\n- Train is one continuous stream; test is independent 150k-sample segments \u2014 featurise each segment and regress onto the end-of-segment `time_to_failure`.\n- Amplitude-dispersion features (`std`, `abs_max`, high quantiles) correlate most with ttf \u2014 the signal grows louder before failure.\n- Natural next steps: richer features (rolling stats, FFT/spectral bands, peak counts), and a GBDT or simple ridge with group-aware CV by earthquake cycle.\n\nFork and extend. \ud83c\udf0e"}], "metadata": {"kernelspec": {"display_name": "Python 3", "language": "python", "name": "python3"}, "language_info": {"name": "python"}}, "nbformat": 4, "nbformat_minor": 5}