{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.11.13","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":117682,"databundleVersionId":14443416,"sourceType":"competition"}],"dockerImageVersionId":31192,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"## Vesuvius Challenge – Data Exploration & Visualization\n\nIn this notebook I explore the train data for the **Vesuvius Challenge – Surface Detection** competition.  \nMy goal is to understand the 3D volumes, label masks, and scroll structure before training any model.\n\nI work in simple phases:  \n- read the metadata,  \n- load one volume and its mask,  \n- check class balance and intensity,  \n- and visualize some slices and shape types.  \n\nThis EDA helps me see how the scroll surface looks in CT data and what I should be careful about in the training step. Ok Let's begin Phase 0\n\n## **PHASE 0–Imports& Paths**\n\nIn this phase I import the libraries I need for this project.\nI also set the main paths for the Kaggle Vesuvius dataset (csv files, images, labels).","metadata":{}},{"cell_type":"code","source":"import os\nfrom pathlib import Path\nfrom collections import Counter\n\nimport numpy as np\nimport pandas as pd\nimport matplotlib.pyplot as plt\nimport cv2  # <-- we use this instead of skimage\n\n# Kaggle input base path\nBASE_PATH = Path(\"/kaggle/input/vesuvius-challenge-surface-detection\")\n\ntrain_csv_path = BASE_PATH / \"train.csv\"\ntest_csv_path  = BASE_PATH / \"test.csv\"\n\ntrain_images_dir = BASE_PATH / \"train_images\"\ntest_images_dir  = BASE_PATH / \"test_images\"\ntrain_labels_dir = BASE_PATH / \"train_labels\"\n\nprint(\"Train CSV:\", train_csv_path)\nprint(\"Train images dir:\", train_images_dir)\nprint(\"Train labels dir:\", train_labels_dir)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-11-14T13:06:05.049062Z","iopub.execute_input":"2025-11-14T13:06:05.049461Z","iopub.status.idle":"2025-11-14T13:06:05.056474Z","shell.execute_reply.started":"2025-11-14T13:06:05.049428Z","shell.execute_reply":"2025-11-14T13:06:05.055592Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## **PHASE1–TrainCSVExploration**\n\nHere I read the train.csv file and look at the basic information.\nI check how many rows and columns I have, and I see the distribution of scroll_id.","metadata":{}},{"cell_type":"code","source":"# PHASE 1: METADATA EXPLORATION (train.csv)\n\ntrain_df = pd.read_csv(train_csv_path)\ndisplay(train_df.head())\n\nprint(\"\\nShape of train_df:\", train_df.shape)\nprint(\"\\nColumns:\", train_df.columns.tolist())\n\n# scroll_id distribution\nscroll_counts = train_df[\"scroll_id\"].value_counts().sort_index()\nprint(\"\\nScroll ID counts:\")\nprint(scroll_counts)\n\nplt.figure(figsize=(6,4))\nscroll_counts.plot(kind=\"bar\")\nplt.title(\"Number of Volumes per scroll_id\")\nplt.xlabel(\"scroll_id\")\nplt.ylabel(\"count\")\nplt.tight_layout()\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-11-14T13:06:05.058266Z","iopub.execute_input":"2025-11-14T13:06:05.058534Z","iopub.status.idle":"2025-11-14T13:06:05.554362Z","shell.execute_reply.started":"2025-11-14T13:06:05.058513Z","shell.execute_reply":"2025-11-14T13:06:05.553560Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## **PHASE 2 – Load One 3D Volume and Label (with cv2)**\n\nIn this phase I write a helper function to load multi-page TIFF files as 3D arrays.\nThen I choose one example from train_df and load both the volume image and its label mask.","metadata":{}},{"cell_type":"code","source":"def load_tiff_3d(path: Path) -> np.ndarray:\n    \"\"\"\n    Load a multi-page TIFF as a 3D numpy array with shape (z, y, x).\n    Uses cv2.imreadmulti to avoid imagecodecs dependency.\n    \"\"\"\n    ok, frames = cv2.imreadmulti(str(path), flags=cv2.IMREAD_UNCHANGED)\n    if (not ok) or (frames is None) or (len(frames) == 0):\n        raise RuntimeError(f\"cv2.imreadmulti failed for {path}\")\n    \n    # Ensure grayscale (drop channels if any)\n    processed = []\n    for f in frames:\n        if f.ndim == 3:  # (H, W, C)\n            # take first channel; CT images should be single-channel anyway\n            f = f[..., 0]\n        processed.append(f)\n        \n    volume = np.stack(processed, axis=0)  # (z, y, x)\n    return volume\n\n# pick first row as example\nexample_row = train_df.iloc[0]\nexample_id = example_row[\"id\"]\nexample_scroll = example_row[\"scroll_id\"]\n\nprint(\"Example id:\", example_id, \" | scroll_id:\", example_scroll)\n\nimg_path = train_images_dir / f\"{example_id}.tif\"\nlabel_path = train_labels_dir / f\"{example_id}.tif\"\n\nprint(\"Image path:\", img_path)\nprint(\"Label path:\", label_path)\n\nvolume = load_tiff_3d(img_path)\nmask   = load_tiff_3d(label_path)\n\nprint(\"Volume shape (z, y, x):\", volume.shape)\nprint(\"Mask   shape (z, y, x):\", mask.shape)\nprint(\"Volume dtype:\", volume.dtype, \"| Mask dtype:\", mask.dtype)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-11-14T13:06:05.555328Z","iopub.execute_input":"2025-11-14T13:06:05.555640Z","iopub.status.idle":"2025-11-14T13:06:07.263043Z","shell.execute_reply.started":"2025-11-14T13:06:05.555619Z","shell.execute_reply":"2025-11-14T13:06:07.261967Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## **PHASE 3 – Mask Class Distribution (0 / 1 / 2)**\n\nIn this part I check the class distribution inside the label mask.\nI count how many voxels are background (0), surface (1), and unlabeled (2), and I plot them.","metadata":{}},{"cell_type":"code","source":"mask_flat = mask.flatten()\nclass_counts = Counter(mask_flat.tolist())\n\nprint(\"Class counts:\", class_counts)\n\ntotal_voxels = mask_flat.size\nfor c in sorted(class_counts.keys()):\n    print(f\"Class {c}: {class_counts[c]} voxels ({class_counts[c] / total_voxels * 100:.4f}%)\")\n\n# bar plot\nclasses = sorted(class_counts.keys())\ncounts  = [class_counts[c] for c in classes]\n\nplt.figure(figsize=(5,4))\nplt.bar(classes, counts)\nplt.xticks(classes)\nplt.xlabel(\"Class (0=background, 1=surface, 2=unlabeled)\")\nplt.ylabel(\"Voxel count\")\nplt.title(f\"Mask class distribution for id={example_id}\")\nplt.tight_layout()\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-11-14T13:06:07.263966Z","iopub.execute_input":"2025-11-14T13:06:07.264199Z","iopub.status.idle":"2025-11-14T13:06:09.957296Z","shell.execute_reply.started":"2025-11-14T13:06:07.264180Z","shell.execute_reply":"2025-11-14T13:06:09.956379Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## **PHASE 4 – Intensity Statistics and Histogram**\n\nHere I look at the intensity values of the CT volume.\nI print the minimum, maximum and mean values, and I draw a histogram to see the distribution.","metadata":{}},{"cell_type":"code","source":"# PHASE 4: INTENSITY STATS & HISTOGRAM\n\nprint(\"Volume min:\", volume.min(),\n      \"max:\", volume.max(),\n      \"mean:\", float(volume.mean()))\n\n# sample for histogram (avoid huge arrays)\nsample = volume.flatten()\nif sample.size > 1_000_000:\n    sample = np.random.choice(sample, size=1_000_000, replace=False)\n\nplt.figure(figsize=(6,4))\nplt.hist(sample, bins=100)\nplt.title(f\"Intensity Histogram for id={example_id}\")\nplt.xlabel(\"Intensity value\")\nplt.ylabel(\"Frequency\")\nplt.tight_layout()\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-11-14T13:06:09.959253Z","iopub.execute_input":"2025-11-14T13:06:09.959564Z","iopub.status.idle":"2025-11-14T13:06:12.270238Z","shell.execute_reply.started":"2025-11-14T13:06:09.959541Z","shell.execute_reply":"2025-11-14T13:06:12.269277Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## **PHASE 5 – Slice Visualization with Mask Overlay**\n\nIn this phase I visualize some slices from the 3D volume.\nI show the raw CT slice and then I overlay the mask to see where the surface is in red.","metadata":{}},{"cell_type":"code","source":"\n\ndef show_slice(z_idx: int, v: np.ndarray, m: np.ndarray):\n    \"\"\"Show raw slice and mask overlay for given z index.\"\"\"\n    img_slice = v[z_idx]\n    mask_slice = m[z_idx]\n\n    plt.figure(figsize=(10,4))\n\n    # raw image\n    plt.subplot(1, 2, 1)\n    plt.imshow(img_slice, cmap=\"gray\")\n    plt.title(f\"Raw slice z={z_idx}\")\n    plt.axis(\"off\")\n\n    # overlay\n    plt.subplot(1, 2, 2)\n    plt.imshow(img_slice, cmap=\"gray\")\n    \n    surface = mask_slice == 1\n    background = mask_slice == 0\n    unlabeled = mask_slice == 2\n\n    # surface in red\n    plt.imshow(np.ma.masked_where(~surface, surface),\n               cmap=\"Reds\", alpha=0.6)\n    plt.title(f\"Slice z={z_idx} with surface\")\n    plt.axis(\"off\")\n    plt.tight_layout()\n    plt.show()\n\nz_min, z_max = 0, volume.shape[0] - 1\nz_mid = (z_min + z_max) // 2\n\nfor z in [z_min + 5, z_mid, z_max - 5]:\n    z = max(0, min(z, volume.shape[0]-1))\n    show_slice(z, volume, mask)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-11-14T13:07:27.984923Z","iopub.execute_input":"2025-11-14T13:07:27.985420Z","iopub.status.idle":"2025-11-14T13:07:28.821147Z","shell.execute_reply.started":"2025-11-14T13:07:27.985378Z","shell.execute_reply":"2025-11-14T13:07:28.820256Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### **Phase 6 – Shape types of the volumes**\n\nHere I look at the 3D shapes of the first 30 volumes.\n\n- First, I collect the shape for each `scroll_id`.\n- Then I make a small summary table that shows **each shape type** and **how many times** it appears.\n- I also show the shapes per `scroll_id` with counts.\n\nI hope this helps me see that not all volumes have the same shape, so I thought should handle this in the training step (resize, crop, or pad).","metadata":{}},{"cell_type":"code","source":"shapes_info = []\n\n# I check only the first 30 volumes here (you can change the number)\nfor _, row in train_df.head(50).iterrows():\n    vid = row[\"id\"]\n    sid = row[\"scroll_id\"]\n    vpath = train_images_dir / f\"{vid}.tif\"\n    \n    v = load_tiff_3d(vpath)\n    shapes_info.append((sid, v.shape))\n\nshapes_df = pd.DataFrame(shapes_info, columns=[\"scroll_id\", \"shape\"])\n\n\ndisplay(shapes_df.head())\n\n\nprint(\"\\nUnique shapes (all scrolls, first 30 volumes):\")\nshape_counts = shapes_df[\"shape\"].value_counts().reset_index()\nshape_counts.columns = [\"shape\", \"count\"]\ndisplay(shape_counts)\n\n\nprint(\"\\nUnique shapes per scroll_id (with counts):\")\nshape_per_scroll = (\n    shapes_df\n    .groupby([\"scroll_id\", \"shape\"])\n    .size()\n    .reset_index(name=\"count\")\n)\ndisplay(shape_per_scroll)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-11-14T13:08:48.539292Z","iopub.execute_input":"2025-11-14T13:08:48.540270Z","iopub.status.idle":"2025-11-14T13:09:24.610978Z","shell.execute_reply.started":"2025-11-14T13:08:48.540244Z","shell.execute_reply":"2025-11-14T13:09:24.610098Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"","metadata":{"trusted":true},"outputs":[],"execution_count":null}]}