{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.10.12","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":39763,"databundleVersionId":11756775,"sourceType":"competition"}],"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Geophysical Waveform Inversion: Simple EDA and Introduction\n\n## Competition Objective & Evaluation Metric\n\nThe **Geophysical Waveform Inversion** competition challenges participants to predict subsurface properties (like a velocity map of the underground) from recorded waveforms of seismic or radar data. In other words, given the **waveform data** (time-series signals measured by sensors on the surface), we need to infer an **image of the subsurface** that produced those signals. This task is essentially a data-driven approach to _full waveform inversion (FWI)_, a geophysical technique that aims to reconstruct subsurface velocity maps from seismic measurements.\n\nParticipants are provided with many examples of waveform recordings (input) and the corresponding subsurface model (output) for training. The goal is to train a model that can learn this mapping and then predict the subsurface model for new waveform data. The competition’s **objective** is to achieve the most accurate inversion of waveforms into subsurface models on the test set.\n\n**Evaluation Metric:** The competition uses a regression error metric to measure how close your predicted model is to the true model. In this case, a common choice (and likely the one used here) is the **Mean Absolute Error (MAE)**. MAE measures the average absolute difference between the predicted values and the true values across all points of the model. In simpler terms, it’s the average of how “far off” each predicted value is from the actual value, without considering whether the error is positive or negative. A lower MAE means your predicted subsurface properties are closer to the true properties. Formally, if $y_{ij}$ is the true value at location $ij$ in the subsurface grid and $\\hat{y}_{ij}$ is your prediction, MAE would be: \n\n$$\n\\text{MAE} = \\frac{1}{N} \\sum_{i,j} \\left| \\hat{y}_{ij} - y_{ij} \\right|\n$$\n\naveraged over all $N$ grid points in all test samples. The **goal** is to minimize this error.\n\n## Dataset Overview and Structure\n\nUnderstanding the dataset structure is key to getting started. This competition’s dataset includes recorded waveforms as inputs and corresponding velocity models (or similar physical property maps) as outputs. The files are organized to separate training data (with known targets) from test data (where targets are hidden and need to be predicted). Here’s an overview of the files provided in the dataset and their roles:\n\n- **Training Inputs:** These are the waveform recordings for each training example. They might be stored as numerical arrays (e.g., NumPy `.npy` files) containing the signals recorded by sensors. Each training sample can include multiple time-series – for example, if there are several sources or shots and multiple receivers, the data could be a collection of waveforms. In this competition, each sample’s seismic data is likely a 3D array (or set of 2D arrays) representing signals over time and space.\n\n- **Training Labels (Target Models):** These are the subsurface models corresponding to each training input. Typically, the subsurface is represented as a grid (2D image or 3D volume) of values (e.g., seismic wave velocity at each point). In the dataset, these are provided for training samples, so your model can learn to predict them. They might also be stored as `.npy` files or perhaps images. Each label is essentially a numerical grid (e.g., a 2D velocity map).\n\n- **Test Inputs:** These are waveform recordings for the test set. They have the same format as the training inputs (same kind of sensors and layout), but the true subsurface models for these are not given. Your task is to use your trained model to predict the subsurface properties for each of these test waveform records.\n\n- **Sample Submission:** A CSV file (`sample_submission.csv`) is provided to show the format in which you must submit your predictions.\n\nLet’s verify the dataset structure using Python:","metadata":{}},{"cell_type":"code","source":"import os\n\nbase_path = \"/kaggle/input/waveform-inversion\"\nprint(os.listdir(base_path))\nif 'train_samples' in os.listdir(base_path):\n    print(\"Train folder contents:\", os.listdir(os.path.join(base_path, 'train_samples')))","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-09T10:01:12.529433Z","iopub.execute_input":"2025-04-09T10:01:12.529759Z","iopub.status.idle":"2025-04-09T10:01:12.537699Z","shell.execute_reply.started":"2025-04-09T10:01:12.529735Z","shell.execute_reply":"2025-04-09T10:01:12.536673Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Exploratory Data Analysis (EDA)\n### Load Sample Data","metadata":{}},{"cell_type":"code","source":"import numpy as np\n\ndata_batch = np.load( \"/kaggle/input/waveform-inversion/train_samples/CurveVel_A/data/data1.npy\" )\nmodel_batch = np.load( \"/kaggle/input/waveform-inversion/train_samples/CurveVel_A/model/model1.npy\" )\nprint(\"Data batch shape:\", data_batch.shape)\nprint(\"Model batch shape:\", model_batch.shape)\n\nsample_seismic = data_batch[0]\nsample_velocity = model_batch[0]\nprint(\"Single sample seismic data shape:\", sample_seismic.shape)\nprint(\"Single sample velocity model shape:\", sample_velocity.shape)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-09T09:57:01.728056Z","iopub.execute_input":"2025-04-09T09:57:01.728407Z","iopub.status.idle":"2025-04-09T09:57:09.304602Z","shell.execute_reply.started":"2025-04-09T09:57:01.728381Z","shell.execute_reply":"2025-04-09T09:57:09.303499Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### Visualize Waveform Data","metadata":{}},{"cell_type":"code","source":"import matplotlib.pyplot as plt\n\ndef viz_waveform_data(sample_seismic):\n    shot0 = sample_seismic[0]\n    \n    plt.figure(figsize=(6,4))\n    plt.imshow(shot0.T, aspect='auto', cmap='seismic', origin='upper')\n    plt.colorbar(label=\"Amplitude\")\n    plt.title(\"Waveform amplitudes (Time vs Receiver)\")\n    plt.xlabel(\"Time index\")\n    plt.ylabel(\"Receiver index\")\n    plt.show()\n    \n    plt.figure(figsize=(6,4))\n    for rec_idx in [0, 35, 69]:\n        trace = shot0[:, rec_idx]\n        plt.plot(trace, label=f\"Receiver {rec_idx}\")\n    plt.title(\"Sample waveforms at selected receivers\")\n    plt.xlabel(\"Time index\")\n    plt.ylabel(\"Amplitude\")\n    plt.legend()\n    plt.show()\n\nviz_waveform_data(sample_seismic)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-09T10:10:09.847429Z","iopub.execute_input":"2025-04-09T10:10:09.847793Z","iopub.status.idle":"2025-04-09T10:10:10.383436Z","shell.execute_reply.started":"2025-04-09T10:10:09.847765Z","shell.execute_reply":"2025-04-09T10:10:10.382168Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### Visualize Velocity Model","metadata":{}},{"cell_type":"code","source":"def viz_velocity_model( sample_velocity ):\n    velocity_grid = sample_velocity.squeeze()\n    \n    plt.figure(figsize=(5,5))\n    plt.imshow(velocity_grid, origin='upper', cmap='viridis')\n    plt.colorbar(label=\"Velocity (m/s)\")\n    plt.title(\"Example Velocity Model\")\n    plt.xlabel(\"Horizontal position\")\n    plt.ylabel(\"Depth\")\n    plt.show()\n\nviz_velocity_model( sample_velocity )","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-09T10:10:38.578714Z","iopub.execute_input":"2025-04-09T10:10:38.579095Z","iopub.status.idle":"2025-04-09T10:10:38.862311Z","shell.execute_reply.started":"2025-04-09T10:10:38.579065Z","shell.execute_reply":"2025-04-09T10:10:38.861297Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### Velocity Value Distribution","metadata":{}},{"cell_type":"code","source":"def viz_velocity_value_distribution(model_batch):\n    all_velocities = model_batch.reshape(-1)\n    print(\"Velocity stats – min:\", all_velocities.min(), \"max:\", all_velocities.max(), \n          \"mean:\", all_velocities.mean())\n    \n    plt.figure(figsize=(6,4))\n    plt.hist(all_velocities, bins=50, color='gray')\n    plt.title(\"Distribution of Velocity Values (Sample Batch)\")\n    plt.xlabel(\"Velocity (m/s)\")\n    plt.ylabel(\"Frequency\")\n    plt.show()\n\nviz_velocity_value_distribution(model_batch)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-09T10:11:14.051756Z","iopub.execute_input":"2025-04-09T10:11:14.052142Z","iopub.status.idle":"2025-04-09T10:11:14.380422Z","shell.execute_reply.started":"2025-04-09T10:11:14.052103Z","shell.execute_reply":"2025-04-09T10:11:14.379421Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"series = [\n    \"CurveVel_A\", \n    \"CurveVel_B\",\n    \"FlatVel_A\",\n    \"FlatVel_B\",\n    \"Style_A\",\n    \"Style_B\",\n]\n\nfor family_name in series:\n    data_file = f\"/kaggle/input/waveform-inversion/train_samples/{family_name}/data/data1.npy\"\n    model_file = f\"/kaggle/input/waveform-inversion/train_samples/{family_name}/model/model1.npy\"\n    print( \"-\" * 30 )\n    print( f\"\\n// {family_name}\" )\n    data_batch = np.load( data_file )\n    model_batch = np.load( model_file )\n    sample_seismic = data_batch[0]\n    sample_velocity = model_batch[0]\n\n    viz_waveform_data(sample_seismic)\n    viz_velocity_model( sample_velocity )\n    viz_velocity_value_distribution(model_batch)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-09T10:16:46.329460Z","iopub.execute_input":"2025-04-09T10:16:46.330997Z","iopub.status.idle":"2025-04-09T10:17:07.772087Z","shell.execute_reply.started":"2025-04-09T10:16:46.330937Z","shell.execute_reply":"2025-04-09T10:17:07.771005Z"}},"outputs":[],"execution_count":null}]}