{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.11.11","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":39763,"databundleVersionId":11756775,"sourceType":"competition"}],"dockerImageVersionId":31012,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# SeismicDataHelper Class & Dataset Exploration Guide\n\nWelcome! Working with OpenFWI seismic–velocity datasets can be frustrating — especially when trying to match and interpret the data correctly. To streamline the process, this notebook introduces a helper utility: `SeismicDataHelper`.\n\nYou can copy the full class implementation (in the hidden cell below) into your project. While not exhaustive, it provides a strong foundation for common tasks.\n\n![Seismic Helper Visualization](https://www.thomasmeli.com/images/kaggle/geo-fwi-2025/seismic_helper.webp)\n\n\n## Purpose of This Notebook\n\nIf you’re new to seismic data workflows, this guide will help you:\n1. Quickly set up your environment and understand the folder structure of the competition.\n2. Enumerate and explore available datasets.\n3. Validate dataset integrity to prevent silent mismatches or errors.\n4. Visualize essential components of seismic and velocity data.\n5. Apply batch processing to analyze or manipulate multiple samples efficiently.\n\n---\n\n## Notebook Layout\nBelow are short examples demonstrating each concept. You can run them in any order. For quick reference:\n\n| Example | Purpose |\n|---------|---------|\n| **1**   | Initialize `SeismicDataHelper` -- Data Family Guide\n| **2**   | Show dataset folders & count samples |\n| **3–4** | Query pairs in specific datasets or across all data |\n| **5**   | Retrieve specific (seis, vel) paths |\n| **6**   | Sample random data pairs |\n| **7**   | Validate file existence, shape, & loadability |\n| **8**   | Visualize a random sample |\n| **9**   | Define & use a custom plotting function |\n| **10**  | Run a batch function across a dataset |\n| **11**  | Plot a spectrogram of a specific trace |\n| **12**  | Plot a single time series for a known sample |\n\n---\n\n**Tip**: After each example, tweak parameters (like dataset folder names or the specific file) to see how it responds. This notebook is yours—use it to freely experiment and adapt these methods for your research or analysis needs.\n\nEnjoy exploring your seismic data!","metadata":{"_uuid":"963fcc74-9611-4bcf-b53d-700d4ec987f4","_cell_guid":"6d21000b-e3b2-482a-8d5f-8345b5d6b1f3","trusted":true,"collapsed":false,"jupyter":{"outputs_hidden":false}}},{"cell_type":"markdown","source":"## Copy the Entire Hidden Cell Below (Including Imports) Below and Use It in Your Projects to Explore ##","metadata":{"_uuid":"3b34f698-adc0-4ccd-9091-313dd9c77da6","_cell_guid":"e1b32247-cc30-4e6d-8efb-cd97c7671e49","trusted":true,"collapsed":false,"jupyter":{"outputs_hidden":false}}},{"cell_type":"code","source":"import os\nimport numpy as np\nimport matplotlib.pyplot as plt\nfrom pathlib import Path\nimport random\nfrom typing import Callable, List, Tuple, Any, Dict, Union, Optional\nimport textwrap \nimport inspect\n\nclass SeismicDataHelper:\n    \"\"\"\n    Manage OpenFWI seismic-velocity pairs under a root folder.\n\n    Datasets scanned (keys):\n      CurveFault_A, CurveFault_B, CurveVel_A, CurveVel_B,\n      FlatFault_A,  FlatFault_B,  FlatVel_A,  FlatVel_B,\n      Style_A,      Style_B\n\n    Quick access:\n      sd = SeismicDataHelper()                      # Default uses kaggle train folder path.\n      OR sd = SeismicDataHelper(\"train_folder_path\")\n\n      sd.explain(folder_name)                -> print a brief, descriptive overview\n      sd.datasets                            -> list of dataset names\n      sd.CurveFault_A                        -> list of (seis_path, vel_path)\n      len(sd.CurveFault_A)                   -> number of pairs in that set\n      len(sd)                                -> total pairs across all sets\n      sd['FlatVel_B']                        -> same as sd.FlatVel_B\n\n    Common methods:\n      folder_name refers to the CurveFault_A or CurveFault_B etc.\n      \n      sd.pairs(folder_name)                 -> all (seis_path, vel_path) pairs\n      sd.sample(folder_name, n)             -> n random pairs (n='all')\n      sd.get(folder_name, seis_file)        -> loads paired velocity array\n      sd.pipeline(folder_name, fn)          -> apply fn(seis, vel) to every pair\n      sd.verify(folder_name, verbose=True)  -> sanity-check all files/arrays\n      sd.plot(folder_name, file=None)       -> plot seismic+velocity (random/file)\n      sd.plot_fn(folder_name, fn, ...)      -> custom fn(X, y, ...) plotting\n      sd.plot_signal_time_series(...)       -> plot seismic signal vs. time\n      sd.sample_pairs(folder_name, n)      -> print random sample file pairs\n    \"\"\"\n\n    def __init__(self, root_dir: str = \"/kaggle/input/waveform-inversion/train_samples\"):\n        \"\"\"\n        Initialize SeismicDataHelper from a root directory of datasets.\n\n        Args:\n            root_dir (str): Directory path containing datasets. Defaults to Kaggle training samples path.\n        \"\"\"\n        self.root_dir = root_dir\n        self._pairs: Dict[str, List[Tuple[str, str]]] = self._scan_pairs()\n\n    def _scan_pairs(self) -> Dict[str, List[Tuple[str, str]]]:\n        pairs: Dict[str, List[Tuple[str, str]]] = {}\n        for folder_name in os.listdir(self.root_dir):\n            ds_dir = os.path.join(self.root_dir, folder_name)\n            if not os.path.isdir(ds_dir):\n                continue\n\n            data_dir = os.path.join(ds_dir, 'data')\n            model_dir = os.path.join(ds_dir, 'model')\n            if os.path.isdir(data_dir) and os.path.isdir(model_dir):\n                lst = []\n                for fname in os.listdir(data_dir):\n                    if fname.startswith('data') and fname.endswith('.npy'):\n                        suf = fname[len('data'):-4]\n                        seis = os.path.join(data_dir, fname)\n                        vel = os.path.join(model_dir, f'model{suf}.npy')\n                        if os.path.exists(vel):\n                            lst.append((seis, vel))\n                if lst:\n                    pairs[folder_name] = sorted(lst)\n                continue\n\n            files = os.listdir(ds_dir)\n            seis_fs = [f for f in files if f.startswith('seis') and f.endswith('.npy')]\n            vel_fs = [f for f in files if f.startswith('vel') and f.endswith('.npy')]\n            if seis_fs and vel_fs:\n                vel_set = set(vel_fs)\n                lst = []\n                for sf in seis_fs:\n                    suf = sf[len('seis'):-4]\n                    vf = f'vel{suf}.npy'\n                    if vf in vel_set:\n                        lst.append((os.path.join(ds_dir, sf), os.path.join(ds_dir, vf)))\n                if lst:\n                    pairs[folder_name] = sorted(lst)\n        return pairs\n\n    @property\n    def datasets(self) -> List[str]:\n        \"\"\"List all available dataset folder names.\"\"\"\n        return list(self._pairs.keys())\n\n    def __repr__(self) -> str:\n        return f\"<SeismicDataHelper datasets={self.datasets}>\"\n\n    def __getattr__(self, key: str) -> Any:\n        if key in self._pairs:\n            return self._pairs[key]\n        raise AttributeError(f\"{self.__class__.__name__!r} has no attribute {key!r}\")\n\n    def __getitem__(self, key: str) -> List[Tuple[str, str]]:\n        return self._pairs[key]\n\n    def __len__(self) -> int:\n        \"\"\"Total number of (seis, vel) pairs across all datasets.\"\"\"\n        return sum(len(v) for v in self._pairs.values())\n\n    def pairs(self, folder_name: str) -> List[Tuple[str, str]]:\n        \"\"\"Return all (seis_path, vel_path) for given dataset folder.\"\"\"\n        return self._pairs[folder_name]\n\n    def sample(self, folder_name: str, n: Union[int, str] = 'all') -> List[Tuple[str, str]]:\n        \"\"\"Return n random (seis, vel) pairs or all if n='all'.\"\"\"\n        allp = self.pairs(folder_name)\n        if isinstance(n, str) and n.lower() == 'all':\n            return allp\n        if not isinstance(n, int) or n < 1:\n            raise ValueError(\"n must be a positive int or 'all'\")\n        return allp if n >= len(allp) else random.sample(allp, n)\n\n    def sample_pairs(self, folder_name: str, n: int = 4) -> None:\n        \"\"\"\n        Print n random (seis, vel) pairs in a simple report format.\n        \"\"\"\n        samples = self.sample(folder_name, n)\n        print(f\"{n} random {folder_name} samples:\")\n        for s, v in samples:\n            print(\"  \", os.path.basename(s), \"↔\", os.path.basename(v))\n\n    def get(self, folder_name: str, seis_file: str) -> np.ndarray:\n        \"\"\"Load the velocity array paired with a given seismic filename.\"\"\"\n        for s, v in self.pairs(folder_name):\n            if os.path.basename(s) == seis_file:\n                return np.load(v)\n        raise KeyError(f\"{seis_file!r} not found in {folder_name!r}\")\n\n    def get_seis(self, folder_name: str, file: str) -> np.ndarray:\n        \"\"\"\n        Load ONLY seismic data from the given folder_name + filename.\n        \"\"\"\n        for s, v in self.pairs(folder_name):\n            if os.path.basename(s) == file:\n                return np.load(s)\n        raise KeyError(f\"{file!r} not found in {folder_name!r}\")\n\n    def get_vel(self, folder_name: str, file: str) -> np.ndarray:\n        \"\"\"\n        Load ONLY velocity data from the given folder_name + filename.\n        \"\"\"\n        for s, v in self.pairs(folder_name):\n            if os.path.basename(s) == file:\n                return np.load(v)\n        raise KeyError(f\"{file!r} not found in {folder_name!r}\")\n        \n    def pipeline(\n        self,\n        folder_name: str,\n        callbacks: Union[Callable[[np.ndarray, np.ndarray], Any],\n                         List[Callable[[np.ndarray, np.ndarray], Any]]]\n    ) -> List[Any]:\n        \"\"\"Apply one or more callbacks to every (seis, vel) pair in dataset.\"\"\"\n        if not isinstance(callbacks, list):\n            callbacks = [callbacks]\n        out = []\n        for s, v in self.pairs(folder_name):\n            seis = np.load(s)\n            vel = np.load(v)\n            for cb in callbacks:\n                out.append(cb(seis, vel))\n        return out\n\n    def verify(self, folder_name: str, verbose: bool = True) -> bool:\n        \"\"\"Ensure all file pairs exist, load properly, and have expected dimensions.\"\"\"\n        for s, v in self.pairs(folder_name):\n            if not os.path.exists(s) or not os.path.exists(v):\n                raise FileNotFoundError(f\"Missing {s} or {v}\")\n            a_s, a_v = np.load(s), np.load(v)\n            if verbose:\n                print(f\"✔ {os.path.basename(s)} <-> {os.path.basename(v)}\")\n            if a_s.ndim < 3 or a_v.ndim < 2:\n                raise ValueError(f\"Bad shapes: {a_s.shape}, {a_v.shape}\")\n        return True\n\n    def plot(self, folder_name: str, file: str = None, index: int = 0):\n        \"\"\"Plot a seismic-velocity pair, selected randomly or by seismic filename.\"\"\"\n        pairs = self.pairs(folder_name)\n        if file:\n            for s, v in pairs:\n                if os.path.basename(s) == file:\n                    seis_p, vel_p = s, v\n                    break\n            else:\n                raise KeyError(f\"{file!r} not in {folder_name!r}\")\n        else:\n            seis_p, vel_p = random.choice(pairs)\n        self._plot_seismic(seis_p, index)\n        self._plot_velocity(vel_p, index)\n\n    def plot_signal_time_series(\n        self,\n        folder_name: str,\n        file: str = None,\n        index: int = 0,\n        src: int = 0,\n        recv: Union[int, None] = None\n    ):\n        \"\"\"Plot a single seismic trace vs. time for a given (src, recv) pair.\"\"\"\n        pairs = self.pairs(folder_name)\n        if file:\n            for s, _ in pairs:\n                if os.path.basename(s) == file:\n                    seis_p = s\n                    break\n            else:\n                raise KeyError(f\"{file!r} not in {folder_name!r}\")\n        else:\n            seis_p, _ = random.choice(pairs)\n        arr = np.load(seis_p)\n        trace_batch = arr[index] if arr.ndim == 4 else arr\n        nrecv = trace_batch.shape[-1]\n        r = recv if recv is not None else nrecv // 2\n        trace = trace_batch[src, :, r]\n        fname = os.path.basename(seis_p)\n        plt.figure()\n        plt.plot(trace)\n        plt.title(f\"{fname} — Trace src={src}, recv={r}\")\n        plt.xlabel(\"Time (sample step)\")\n        plt.ylabel(\"Amplitude\")\n        plt.show()\n\n    def plot_fn(\n        self,\n        folder_name: str,\n        fn: Callable[..., Any],\n        file: str = None,\n        index: int = 0\n    ):\n        \"\"\"\n        Load one sample, then call custom plotting function fn(seis, vel, [optional file]).\n        If fn takes 3 args, we pass (X, y, seismic_path). Otherwise, just (X, y).\n        \"\"\"\n        pairs = self.pairs(folder_name)\n        if file:\n            for s, v in pairs:\n                if os.path.basename(s) == file:\n                    seis_p, vel_p = s, v\n                    break\n            else:\n                raise KeyError(f\"{file!r} not in {folder_name!r}\")\n        else:\n            seis_p, vel_p = random.choice(pairs)\n\n        arr_X = np.load(seis_p)\n        arr_y = np.load(vel_p)\n        X = arr_X[index] if arr_X.ndim == 4 else arr_X\n        if arr_y.ndim == 4:\n            y = arr_y[index, 0]\n        elif arr_y.ndim == 3:\n            y = arr_y[0] if arr_y.shape[0] == 1 else arr_y[index]\n        else:\n            y = arr_y\n\n        # Inspect how many parameters the user function wants:\n        params = inspect.signature(fn).parameters\n        if len(params) == 3:\n            fig = fn(X, y, seis_p)\n        else:\n            fig = fn(X, y)\n\n        if isinstance(fig, plt.Figure):\n            plt.show()\n\n    def _plot_seismic(self, path: str, index: int):\n    \n        arr   = np.load(path)\n        batch = arr[index] if arr.ndim == 4 else arr     # → shape = (nsrc, nt, nrec)\n        nsrc, nt, nrec = batch.shape\n        fname = os.path.basename(path)\n    \n        fig, axs = plt.subplots(1, nsrc, figsize=(3*nsrc, 3))\n        if nsrc == 1:\n            axs = [axs]\n    \n        # extent = [xmin, xmax, ymin, ymax] flips Y so time=0 is at the top\n        extent = [0, nrec, nt, 0]\n    \n        for i, ax in enumerate(axs):\n            ax.imshow(\n                batch[i],          # rows=time, cols=receivers\n                aspect='auto',\n                cmap='gray',\n                extent=extent\n            )\n            ax.set_title(f\"{fname} — Src {i}\")\n            ax.set_xlabel(\"Offset (m)\")\n            ax.set_ylabel(\"Time (sample step)\")\n    \n        plt.tight_layout()\n        plt.show()\n\n\n\n\n    def _plot_velocity(\n        self,\n        path: str | Path,\n        index: int = 0,\n        *,\n        vmin: Optional[float] = None,\n        vmax: Optional[float] = None,\n        cmap: str = \"viridis\",\n    ) -> None:\n        \"\"\"\n        Plot a single 2-D slice of an openWFI P-wave velocity cube.\n        Units of the colors are in meters per second (shown in the title).\n    \n        Parameters\n        ----------\n        path : str | Path\n            Path to the .npy file holding the velocity array.\n        index : int, default 0\n            Which slice (inlines or time steps) to visualise if arr.ndim >= 3.\n        vmin, vmax : float, optional\n            Colour limits in m/s.  Leave None for automatic scaling.\n        cmap : str, default \"viridis\"\n            Any Matplotlib colormap name.\n        \"\"\"\n        arr = np.load(path)\n        fname = Path(path).name\n    \n        # pick out the 2-D slice we actually want to draw\n        if arr.ndim == 4:\n            img = arr[index, 0]\n        elif arr.ndim == 3:\n            img = arr[0] if arr.shape[0] == 1 else arr[index]\n        else:\n            img = arr  # already 2-D\n    \n        fig, ax = plt.subplots(figsize=(6, 5))\n        cax = ax.imshow(img, aspect=\"equal\", vmin=vmin, vmax=vmax, cmap=cmap)\n    \n        # Updated title to include the color units\n        ax.set_title(f\"{fname} — velocity model\\n(color = velocity m/s)\")\n        ax.set_xlabel(\"Horizontal index (sample #)\")\n        ax.set_ylabel(\"Depth index (sample #)\")\n    \n        cbar = fig.colorbar(cax, ax=ax, orientation=\"vertical\", fraction=0.04, pad=0.03)\n        # we can leave the cbar ticks unlabeled since units are in the title\n        cbar.ax.xaxis.set_ticks_position(\"none\")\n    \n        fig.tight_layout(pad=3.0)\n        plt.show()\n\n\n\n    def _wrap_preserving_newlines(self, text, width=65):\n        lines = text.splitlines()\n        wrapped = [\n            textwrap.fill(line, width=width, break_long_words=False, break_on_hyphens=False) if line.strip() else ''\n            for line in lines\n        ]\n        return \"\\n\".join(wrapped)\n\n    \n    def explain(self, folder_name: str):\n        \"\"\"\n        Print a concise, beginner-friendly overview of the specified dataset,\n        highlighting its significance and connection to other folders.\n        \"\"\"\n    \n        # Determine the main dataset type and version (A or B).\n        name_upper = folder_name.upper()\n        main_type = \"Unknown\"\n        version = \"\"\n    \n        if \"FLATVEL\" in name_upper:\n            main_type = \"FlatVel\"\n        elif \"CURVEVEL\" in name_upper:\n            main_type = \"CurveVel\"\n        elif \"FLATFAULT\" in name_upper:\n            main_type = \"FlatFault\"\n        elif \"CURVEFAULT\" in name_upper:\n            main_type = \"CurveFault\"\n        elif \"STYLE\" in name_upper:\n            main_type = \"Style\"\n    \n        if \"_A\" in name_upper:\n            version = \"A (easier version)\"\n        elif \"_B\" in name_upper:\n            version = \"B (harder version)\"\n    \n        dataset_descriptions = {\n            \"FlatVel\": (\n                \"This dataset contains horizontally-layered (flat) subsurface velocity \"\n                \"models without any faults. Seismic signals reflect from layer boundaries \"\n                \"in a more straightforward way, so it's considered simpler than faulted or \"\n                \"curved scenarios.\"\n            ),\n            \"FlatFault\": (\n                \"Like FlatVel, but each velocity model has at least one major fault \"\n                \"offsetting the flat layers. This creates more abrupt jumps in velocity \"\n                \"and thus more complex waveforms for the inversion task.\"\n            ),\n            \"CurveVel\": (\n                \"These velocity models have undulating (curved) layers without faults. \"\n                \"Waveforms reflect and refract in a more complicated manner compared to \"\n                \"flat layers.\"\n            ),\n            \"CurveFault\": (\n                \"Combines curved layering with at least one fault. This is among the most \"\n                \"challenging 2D cases, since faults add discontinuities on top of curved \"\n                \"geometry.\"\n            ),\n            \"Style\": (\n                \"Velocity models generated via style transfer from natural or artistic \"\n                \"images, producing irregular, non-layered patterns. Seismic signals here \"\n                \"can be quite complex and unpredictable compared to layered scenarios.\"\n            ),\n            \"Unknown\": (\n                \"This dataset name doesn't match the typical OpenFWI naming pattern, \"\n                \"so a specialized description isn't available.\"\n            )\n        }\n    \n        desc = dataset_descriptions.get(main_type, dataset_descriptions[\"Unknown\"])\n    \n        explanation_raw = f\"\"\"\n    ┌─────────────────────────────────────────────────────┐\n    │          EXPLANATION FOR:  {folder_name}           │\n    └─────────────────────────────────────────────────────┘\n    \n    This folder is classified as:  {main_type}\n    Version:                       {version or \"N/A\"}\n    \n    {desc}\n    \n    In OpenFWI, {main_type} is connected to other families like:\n      - 'Flat' or 'Curve' indicate the layer geometry (horizontal vs. undulating).\n      - 'Vel' or 'Fault' tell us if there's a fault present or not.\n      - 'Style' stands for more irregular or style-transferred velocity fields.\n      - 'A' versions are often simpler, while 'B' versions add complexity.\n    \n    The /data and /model subfolders hold .npy files that line up one-to-one,\n    where each seismic recording file in /data corresponds to a velocity map\n    in /model with the same file number. Together, they form the input-output\n    pairs used in deep-learning-based Full Waveform Inversion.\n    \n    Enjoy exploring {folder_name}!\n    \"\"\".strip(\"\\n\")\n    \n        # Print the explanation directly (no text wrapping).\n        print(explanation_raw)\n    \n        print(f\"Plotting example of {folder_name}\")\n        self.plot(folder_name)\n","metadata":{"_uuid":"ca7ea24e-6137-4818-937b-0a1b495db3eb","_cell_guid":"0a7665cb-055e-4585-a072-772de950df03","trusted":true,"collapsed":false,"jupyter":{"outputs_hidden":false},"execution":{"iopub.status.busy":"2025-04-25T17:03:06.309681Z","iopub.execute_input":"2025-04-25T17:03:06.309959Z","iopub.status.idle":"2025-04-25T17:03:06.353704Z","shell.execute_reply.started":"2025-04-25T17:03:06.309938Z","shell.execute_reply":"2025-04-25T17:03:06.352722Z"},"_kg_hide-input":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### Example 1: Initialize and inspect the `SeismicDataHelper` object\n\nInitializes the dataset object from the Kaggle waveform inversion path and prints its representation.","metadata":{"_uuid":"aba3c312-3aa8-4f93-9b9f-4273fb00fdac","_cell_guid":"414a787b-4474-47ea-b9cb-b33f522f2868","trusted":true,"collapsed":false,"jupyter":{"outputs_hidden":false}}},{"cell_type":"code","source":"kaggle_train_path = \"/kaggle/input/waveform-inversion/train_samples\"\nsd = SeismicDataHelper(kaggle_train_path)\nprint(sd)","metadata":{"_uuid":"2ec58f07-0f8e-4a21-872c-674bdb61bbed","_cell_guid":"81f21329-e9ba-4fe6-97e9-ddab8a2bddf6","trusted":true,"collapsed":false,"jupyter":{"outputs_hidden":false},"execution":{"iopub.status.busy":"2025-04-25T17:03:09.356312Z","iopub.execute_input":"2025-04-25T17:03:09.357048Z","iopub.status.idle":"2025-04-25T17:03:09.444196Z","shell.execute_reply.started":"2025-04-25T17:03:09.357012Z","shell.execute_reply":"2025-04-25T17:03:09.442893Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### Example 1b: Get an explanation for a family.","metadata":{"_uuid":"0a6ac0fb-8c5a-43ba-822d-9ffb4e9f6b02","_cell_guid":"71a9091f-26e0-4a21-9ea3-3a5feb2fe0c7","trusted":true,"collapsed":false,"jupyter":{"outputs_hidden":false}}},{"cell_type":"code","source":"sd.explain(\"CurveFault_A\")","metadata":{"_uuid":"58fa058c-ceac-4444-be80-c6e63ab3f1ae","_cell_guid":"5794fb31-5afb-4e28-a0af-c9d700a642e8","trusted":true,"collapsed":false,"jupyter":{"outputs_hidden":false},"execution":{"iopub.status.busy":"2025-04-25T01:36:40.886285Z","iopub.execute_input":"2025-04-25T01:36:40.886606Z","iopub.status.idle":"2025-04-25T01:36:46.405400Z","shell.execute_reply.started":"2025-04-25T01:36:40.886578Z","shell.execute_reply":"2025-04-25T01:36:46.404592Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"sd.explain(\"CurveFault_B\")","metadata":{"_uuid":"6d477e55-b7f3-49e2-86f3-719276327cee","_cell_guid":"2772f362-c3a6-45fa-80a4-eff568d8d8dc","trusted":true,"collapsed":false,"jupyter":{"outputs_hidden":false},"execution":{"iopub.status.busy":"2025-04-25T01:36:46.406292Z","iopub.execute_input":"2025-04-25T01:36:46.406513Z","iopub.status.idle":"2025-04-25T01:36:50.735902Z","shell.execute_reply.started":"2025-04-25T01:36:46.406494Z","shell.execute_reply":"2025-04-25T01:36:50.735039Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"sd.explain(\"CurveVel_A\")","metadata":{"_uuid":"9286f12a-0b78-4bf5-b62b-db7dfafb1083","_cell_guid":"fa497a62-1c23-4268-8897-44525958cef1","trusted":true,"collapsed":false,"jupyter":{"outputs_hidden":false},"execution":{"iopub.status.busy":"2025-04-25T01:36:50.737773Z","iopub.execute_input":"2025-04-25T01:36:50.738025Z","iopub.status.idle":"2025-04-25T01:36:55.087246Z","shell.execute_reply.started":"2025-04-25T01:36:50.738005Z","shell.execute_reply":"2025-04-25T01:36:55.086486Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"sd.explain(\"CurveVel_B\")","metadata":{"_uuid":"9f874d6c-a70e-4b98-8ffb-ba6e7203fc78","_cell_guid":"53cdb232-d2f8-43bc-8e86-39fca928c180","trusted":true,"collapsed":false,"jupyter":{"outputs_hidden":false},"execution":{"iopub.status.busy":"2025-04-25T01:36:55.088177Z","iopub.execute_input":"2025-04-25T01:36:55.088475Z","iopub.status.idle":"2025-04-25T01:36:59.445784Z","shell.execute_reply.started":"2025-04-25T01:36:55.088449Z","shell.execute_reply":"2025-04-25T01:36:59.444940Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"sd.explain(\"FlatFault_A\")","metadata":{"_uuid":"46f985bd-d237-49a4-9d5b-bba02bdf5057","_cell_guid":"113f1d9f-ae7f-44f5-8f21-7d8ebafa5b9d","trusted":true,"collapsed":false,"jupyter":{"outputs_hidden":false},"execution":{"iopub.status.busy":"2025-04-25T01:36:59.446804Z","iopub.execute_input":"2025-04-25T01:36:59.447109Z","iopub.status.idle":"2025-04-25T01:37:03.929964Z","shell.execute_reply.started":"2025-04-25T01:36:59.447083Z","shell.execute_reply":"2025-04-25T01:37:03.929058Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"sd.explain(\"FlatFault_B\")","metadata":{"_uuid":"63ce0add-09eb-46ca-b25b-629dc4e0a4a0","_cell_guid":"e1887fcd-687b-4e18-a264-2d744874a39b","trusted":true,"collapsed":false,"jupyter":{"outputs_hidden":false},"execution":{"iopub.status.busy":"2025-04-25T01:37:03.931131Z","iopub.execute_input":"2025-04-25T01:37:03.931431Z","iopub.status.idle":"2025-04-25T01:37:08.546375Z","shell.execute_reply.started":"2025-04-25T01:37:03.931403Z","shell.execute_reply":"2025-04-25T01:37:08.545562Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"sd.explain(\"FlatVel_A\")","metadata":{"_uuid":"7ceb406e-216a-48f8-8b02-b870914b9a79","_cell_guid":"6199df75-341b-45c8-aae0-e8ed9e36cfb6","trusted":true,"collapsed":false,"jupyter":{"outputs_hidden":false},"execution":{"iopub.status.busy":"2025-04-25T01:37:08.547311Z","iopub.execute_input":"2025-04-25T01:37:08.547625Z","iopub.status.idle":"2025-04-25T01:37:13.065537Z","shell.execute_reply.started":"2025-04-25T01:37:08.547598Z","shell.execute_reply":"2025-04-25T01:37:13.064621Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"sd.explain(\"FlatVel_B\")","metadata":{"_uuid":"e7fe0a5d-442c-45bc-81e7-76266bcd43d6","_cell_guid":"a43fbc31-ed5d-4b4e-ac40-940361501eee","trusted":true,"collapsed":false,"jupyter":{"outputs_hidden":false},"execution":{"iopub.status.busy":"2025-04-25T01:37:13.066445Z","iopub.execute_input":"2025-04-25T01:37:13.066697Z","iopub.status.idle":"2025-04-25T01:37:17.569339Z","shell.execute_reply.started":"2025-04-25T01:37:13.066678Z","shell.execute_reply":"2025-04-25T01:37:17.568476Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"sd.explain(\"Style_A\")","metadata":{"_uuid":"bc67813b-50a7-4195-a07e-af99049cf7da","_cell_guid":"1669164d-0d0e-4d01-8c09-2fb623c2063c","trusted":true,"collapsed":false,"jupyter":{"outputs_hidden":false},"execution":{"iopub.status.busy":"2025-04-25T01:37:17.570270Z","iopub.execute_input":"2025-04-25T01:37:17.570508Z","iopub.status.idle":"2025-04-25T01:37:22.112413Z","shell.execute_reply.started":"2025-04-25T01:37:17.570489Z","shell.execute_reply":"2025-04-25T01:37:22.111557Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"sd.explain(\"Style_B\")","metadata":{"_uuid":"ad7219dc-4bf7-4de3-bde3-7c3f9df1ff7c","_cell_guid":"e1917dc4-9497-4dba-8434-b5e67fd57fb3","trusted":true,"collapsed":false,"jupyter":{"outputs_hidden":false},"execution":{"iopub.status.busy":"2025-04-25T01:37:22.113353Z","iopub.execute_input":"2025-04-25T01:37:22.113660Z","iopub.status.idle":"2025-04-25T01:37:27.061881Z","shell.execute_reply.started":"2025-04-25T01:37:22.113634Z","shell.execute_reply":"2025-04-25T01:37:27.061093Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### Example 2: List available datasets\n\nPrints all dataset folder names available in the root directory.","metadata":{"_uuid":"6aa93950-2e34-45e0-8f9e-57c725dabc73","_cell_guid":"50cec6ce-93a6-4624-a8e8-baabe46ce382","trusted":true,"collapsed":false,"jupyter":{"outputs_hidden":false}}},{"cell_type":"code","source":"print(\"Datasets available:\")\nfor name in sd.datasets:\n    print(\" •\", name)","metadata":{"_uuid":"f4703b71-d93c-472f-a4a0-68573cb6b4fa","_cell_guid":"c0608ad3-de72-4e46-bbcf-84ffc5fcfd66","trusted":true,"collapsed":false,"jupyter":{"outputs_hidden":false},"execution":{"iopub.status.busy":"2025-04-25T01:37:27.062912Z","iopub.execute_input":"2025-04-25T01:37:27.063209Z","iopub.status.idle":"2025-04-25T01:37:27.068621Z","shell.execute_reply.started":"2025-04-25T01:37:27.063185Z","shell.execute_reply":"2025-04-25T01:37:27.067600Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### Example 3: Count pairs in a specific dataset\n\nGets the number of seismic/velocity file pairs in the `CurveFault_A` dataset.","metadata":{"_uuid":"15301146-7ff5-46ce-a934-aa221fe28148","_cell_guid":"dea8d8ff-ca58-49e0-9f93-910f85dd75a5","trusted":true,"collapsed":false,"jupyter":{"outputs_hidden":false}}},{"cell_type":"code","source":"count_cf = len(sd.CurveFault_A)\nprint(f\"CurveFault_A has {count_cf} seismic/velocity pairs.\")","metadata":{"_uuid":"d8af98a7-d71c-48b3-aeb4-21da10149b65","_cell_guid":"cfdd576f-9b29-4f50-92b3-776ff7337e4f","trusted":true,"collapsed":false,"jupyter":{"outputs_hidden":false},"execution":{"iopub.status.busy":"2025-04-25T01:37:27.071633Z","iopub.execute_input":"2025-04-25T01:37:27.071875Z","iopub.status.idle":"2025-04-25T01:37:27.084486Z","shell.execute_reply.started":"2025-04-25T01:37:27.071857Z","shell.execute_reply":"2025-04-25T01:37:27.083643Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"sd.get_seis(\"FlatVel_A\", \"data1.npy\")[0]","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-25T17:04:07.034053Z","iopub.execute_input":"2025-04-25T17:04:07.034397Z","iopub.status.idle":"2025-04-25T17:04:07.335883Z","shell.execute_reply.started":"2025-04-25T17:04:07.034371Z","shell.execute_reply":"2025-04-25T17:04:07.335183Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Will work with model soon.","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-25T17:07:22.465839Z","iopub.execute_input":"2025-04-25T17:07:22.466124Z","iopub.status.idle":"2025-04-25T17:07:22.469975Z","shell.execute_reply.started":"2025-04-25T17:07:22.466105Z","shell.execute_reply":"2025-04-25T17:07:22.469060Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### Example 4: Count all pairs across datasets\n\nCalculates the total number of seismic/velocity pairs across all dataset folders.","metadata":{"_uuid":"bb7cc8b0-583c-47ab-b178-6234e00e5b79","_cell_guid":"45b0b6c0-31d9-4101-85dd-49c994c12273","trusted":true,"collapsed":false,"jupyter":{"outputs_hidden":false}}},{"cell_type":"code","source":"total = len(sd)\nprint(f\"Total seismic/velocity pairs across all datasets: {total}\")","metadata":{"_uuid":"65331f08-f0cb-4f94-a4af-cb4c0b7d0a84","_cell_guid":"fb46a9d0-db3b-4dd3-991b-61b7634edec7","trusted":true,"collapsed":false,"jupyter":{"outputs_hidden":false},"execution":{"iopub.status.busy":"2025-04-25T01:37:27.085336Z","iopub.execute_input":"2025-04-25T01:37:27.085620Z","iopub.status.idle":"2025-04-25T01:37:27.103922Z","shell.execute_reply.started":"2025-04-25T01:37:27.085596Z","shell.execute_reply":"2025-04-25T01:37:27.103011Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### Example 5: Get file paths for dataset pairs\n\nPrints the file paths of the first three (seis, vel) pairs from the `FlatVel_B` dataset.","metadata":{"_uuid":"05cb2816-ce22-4212-b97a-d4d9c172fe8e","_cell_guid":"bbb157b5-961a-4a16-9474-884313ab2b25","trusted":true,"collapsed":false,"jupyter":{"outputs_hidden":false}}},{"cell_type":"code","source":"flatvel_b_pairs = sd.pairs(\"FlatVel_B\")\nprint(\"First 3 FlatVel_B pairs:\")\nfor seis_path, vel_path in flatvel_b_pairs[:3]:\n    print(\"  \", seis_path, \"<->\", vel_path)","metadata":{"_uuid":"347b00a6-7c54-447d-b70e-426cb79b631a","_cell_guid":"b06fe39c-2ff6-4fac-8a44-b8edd7044322","trusted":true,"collapsed":false,"jupyter":{"outputs_hidden":false},"execution":{"iopub.status.busy":"2025-04-25T01:37:27.104835Z","iopub.execute_input":"2025-04-25T01:37:27.105179Z","iopub.status.idle":"2025-04-25T01:37:27.120477Z","shell.execute_reply.started":"2025-04-25T01:37:27.105156Z","shell.execute_reply":"2025-04-25T01:37:27.119388Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### Example 6: Random sampling from a dataset\n\nWe now use the class method `sample_pairs` to show a short listing of random pairs.","metadata":{"_uuid":"1a748ebd-08a0-4585-bb37-c779a0deaf73","_cell_guid":"1c82c6d4-6c38-426c-acdf-bc7f9aa107df","trusted":true,"collapsed":false,"jupyter":{"outputs_hidden":false}}},{"cell_type":"code","source":"sd.sample_pairs(\"CurveVel_A\", 4)","metadata":{"_uuid":"484da2cd-2637-40c2-acd1-e57b9e15baa8","_cell_guid":"92c42a3b-1ff3-4044-b97c-31018b9c165f","trusted":true,"collapsed":false,"jupyter":{"outputs_hidden":false},"execution":{"iopub.status.busy":"2025-04-25T01:37:27.121453Z","iopub.execute_input":"2025-04-25T01:37:27.121747Z","iopub.status.idle":"2025-04-25T01:37:27.140629Z","shell.execute_reply.started":"2025-04-25T01:37:27.121722Z","shell.execute_reply":"2025-04-25T01:37:27.139726Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### Example 7: Verify dataset file integrity\n\nChecks that all pairs in `CurveFault_B` exist, load successfully, and have valid shapes. Also prints file names.","metadata":{"_uuid":"314d0a6c-65ff-4cf3-97bf-4e85349961c7","_cell_guid":"07e53497-319e-4c43-a341-3d633f9dbed8","trusted":true,"collapsed":false,"jupyter":{"outputs_hidden":false}}},{"cell_type":"code","source":"try:\n    ok = sd.verify(\"CurveFault_B\", verbose=True)\n    print(\"CurveFault_B verification passed:\", ok)\nexcept Exception as e:\n    print(\"Verification error:\", e)","metadata":{"_uuid":"2762aae9-29d0-4f85-b1ec-f5308012e399","_cell_guid":"bade4b76-ce7a-42cb-9370-a7794afb6ab8","trusted":true,"collapsed":false,"jupyter":{"outputs_hidden":false},"execution":{"iopub.status.busy":"2025-04-25T01:37:27.141512Z","iopub.execute_input":"2025-04-25T01:37:27.141796Z","iopub.status.idle":"2025-04-25T01:37:31.008590Z","shell.execute_reply.started":"2025-04-25T01:37:27.141778Z","shell.execute_reply":"2025-04-25T01:37:31.007710Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### Example 8: Plot a random seismic and velocity pair\n\nVisualizes a randomly selected seismic input and its paired velocity model from the `Style_A` dataset.","metadata":{"_uuid":"97a3f9f3-b064-426f-880b-d4635678fc2f","_cell_guid":"a7c300cb-6f8c-455f-ac39-1acbd418eda5","trusted":true,"collapsed":false,"jupyter":{"outputs_hidden":false}}},{"cell_type":"code","source":"sd.plot(\"Style_A\")","metadata":{"_uuid":"55705765-c5a8-44e8-b4a6-a9fd601fe9d1","_cell_guid":"5730bb66-0e5b-45ca-9ee9-ea4979804bfa","trusted":true,"collapsed":false,"jupyter":{"outputs_hidden":false},"execution":{"iopub.status.busy":"2025-04-25T01:37:31.009481Z","iopub.execute_input":"2025-04-25T01:37:31.009768Z","iopub.status.idle":"2025-04-25T01:37:32.451298Z","shell.execute_reply.started":"2025-04-25T01:37:31.009744Z","shell.execute_reply":"2025-04-25T01:37:32.450318Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### Example 9: Custom plot with a feature function\n\nDemonstrates use of a custom plotting function with `plot_fn`. Here, it compares seismic energy and velocity.","metadata":{"_uuid":"bce1f99b-72c2-40ea-a269-b435b5839b1e","_cell_guid":"92393447-9eec-43c0-924f-cc4552aa2978","trusted":true,"collapsed":false,"jupyter":{"outputs_hidden":false}}},{"cell_type":"code","source":"def line_center(X: np.ndarray, y: np.ndarray):\n    seismic_energy = X.sum(axis=(-1, -2))  # sum over time & receivers\n    center_row = y[y.shape[0]//2, :]       # middle row of velocity\n    fig, ax = plt.subplots()\n    ax.plot(seismic_energy, label=\"Total energy per source\")\n    ax.plot(center_row,   label=\"Velocity at mid-row\")\n    ax.set_title(\"Energy vs. Mid-row Velocity\")\n    ax.set_xlabel(\"Index\")  # dimensionless index\n    ax.set_ylabel(\"Value\")\n    ax.legend()\n    return fig\n\nsd.plot_fn(\"CurveFault_B\", line_center)","metadata":{"_uuid":"db38e69a-c4ad-446b-9bec-e8c860cb7b8d","_cell_guid":"85497951-fd10-4a53-8585-09d57d68685c","trusted":true,"collapsed":false,"jupyter":{"outputs_hidden":false},"execution":{"iopub.status.busy":"2025-04-25T01:37:32.452211Z","iopub.execute_input":"2025-04-25T01:37:32.452979Z","iopub.status.idle":"2025-04-25T01:37:33.007689Z","shell.execute_reply.started":"2025-04-25T01:37:32.452957Z","shell.execute_reply":"2025-04-25T01:37:33.006756Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### Example 10: Apply a batch function across all samples\n\nUses `pipeline` to compute the maximum amplitude from each batch in `FlatFault_A`.","metadata":{"_uuid":"ddef3a13-41b6-4d48-93d5-177b7bb176ea","_cell_guid":"33d70924-de62-4134-bb0b-901db24a1801","trusted":true,"collapsed":false,"jupyter":{"outputs_hidden":false}}},{"cell_type":"code","source":"def max_amplitude(seis: np.ndarray, vel: np.ndarray) -> float:\n    return float(np.abs(seis).max())\n\nmax_vals = sd.pipeline(\"FlatFault_A\", max_amplitude)\nprint(\"Computed max amplitudes for\", len(max_vals), \"batches.\")\nprint(\"First five:\", max_vals[:5])","metadata":{"_uuid":"2797e921-7d92-4dd9-9ea3-4912d13f4319","_cell_guid":"53352ba5-06ba-458a-98bf-97d146aa8a04","trusted":true,"collapsed":false,"jupyter":{"outputs_hidden":false},"execution":{"iopub.status.busy":"2025-04-25T01:37:33.008647Z","iopub.execute_input":"2025-04-25T01:37:33.008936Z","iopub.status.idle":"2025-04-25T01:37:37.976384Z","shell.execute_reply.started":"2025-04-25T01:37:33.008916Z","shell.execute_reply":"2025-04-25T01:37:37.975486Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### Example 11: Use a custom spectrogram plot\n\nDisplays a spectrogram of a mid-receiver trace for a random `CurveFault_B` sample using `plot_fn`.\nNow includes filename in the title and units on the axes.","metadata":{"_uuid":"aa43bba8-d05b-4bcc-abc0-ce88784fd3a0","_cell_guid":"65e4319f-75af-4ec1-86c2-db19e9b32b0d","trusted":true,"collapsed":false,"jupyter":{"outputs_hidden":false}}},{"cell_type":"code","source":"def spectrogram_plot(X: np.ndarray, y: np.ndarray, seis_path: str):\n    trace = X[0, :, X.shape[-1]//2]   # source 0, mid-receiver\n    fname = os.path.basename(seis_path)\n    fig, ax = plt.subplots()\n    ax.specgram(\n        trace,\n        NFFT=256,\n        Fs=1.0,        # sample rate can be treated as 1 if unknown\n        noverlap=128,\n        cmap='magma'\n    )\n    ax.set_title(f\"Spectrogram: {fname} (src=0, mid-receiver)\")\n    ax.set_xlabel(\"Time (sample step)\")\n    ax.set_ylabel(\"Frequency (sample bin)\")\n    return fig\n\nsd.plot_fn(\"CurveFault_B\", spectrogram_plot)","metadata":{"_uuid":"ca7c21a5-be6f-44a1-ba07-f0d5352acdc9","_cell_guid":"2ccec739-d23d-4d4c-aa82-9c4a8db1f668","trusted":true,"collapsed":false,"jupyter":{"outputs_hidden":false},"execution":{"iopub.status.busy":"2025-04-25T01:37:37.977425Z","iopub.execute_input":"2025-04-25T01:37:37.978126Z","iopub.status.idle":"2025-04-25T01:37:38.519722Z","shell.execute_reply.started":"2025-04-25T01:37:37.978096Z","shell.execute_reply":"2025-04-25T01:37:38.516867Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### Example 12: Plot a seismic trace as a time series\n\nVisualizes a specific seismic trace from `CurveFault_A` using `plot_signal_time_series` (now labeled with units).","metadata":{"_uuid":"0e8f900b-3cd5-47e9-a127-471d5e1204ff","_cell_guid":"9ea251f9-9a24-435e-b347-b9bcedf59b16","trusted":true,"collapsed":false,"jupyter":{"outputs_hidden":false}}},{"cell_type":"code","source":"sd.plot_signal_time_series(\"CurveFault_A\", \"seis4_1_0.npy\")","metadata":{"_uuid":"c79b0925-398f-4dff-bea2-16683dbd8aca","_cell_guid":"015bb4e2-6afb-43bf-ba94-bbe19532ddbe","trusted":true,"collapsed":false,"jupyter":{"outputs_hidden":false},"execution":{"iopub.status.busy":"2025-04-25T01:37:38.520767Z","iopub.execute_input":"2025-04-25T01:37:38.521061Z","iopub.status.idle":"2025-04-25T01:37:38.995693Z","shell.execute_reply.started":"2025-04-25T01:37:38.521040Z","shell.execute_reply":"2025-04-25T01:37:38.994812Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Next Steps\n\n### Literature Review ###\n\nIf you want a deeper dive into the background of the problem, AI generated podcasts on this topic, and more, check out the literature review and concept background notebook I've made here\n\n👉 https://www.kaggle.com/code/tpmeli/geo-lit-review-winning-strategies-starter\n\n\n### Exploratory Data Analysis ###\n\nIf you're looking for a comprehensive exploration of geographic and WFI data insights, including detailed visualizations, statistical analyses, and key observations, check out the Exploratory Deep Dive notebook I've prepared here:\n\n👉 https://www.kaggle.com/code/tpmeli/exploratory-deep-dive-geo-wfi-data-insights","metadata":{"_uuid":"0af2a4c6-6e32-4864-906c-8926b922a034","_cell_guid":"18f9ea9d-63ed-4b76-b181-cf5446af574c","trusted":true,"collapsed":false,"jupyter":{"outputs_hidden":false}}}]}