{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Image Subdivision\n\nGiven the limited number of images in our training set, it's highly beneficial to split these images into subareas. This approach helps in:\n\n1. **Training models on a larger fraction of data:** This mitigates the issue of overfitting, which is a common concern when dealing with limited data.\n2. **Expanding the possibilities for ensemble models:** Instead of building ensemble models using the entire images (1, 2, 3), we can create more diverse models using different portions of each image (e.g., 1 part 1, 1 part 2, ...).\n\nIn order to implement this, we will:\n\n1. **Import necessary libraries:** These are the software modules that contain the tools we need to manipulate the images.\n2. **Define image handling functions:** This will include functions to read images, display their subareas, and save these subareas for further use.\n3. **Create new directory with subareas:** We'll establish how we want to split the image fragments. Upon executing the corresponding code, a new folder with all the subdivided image areas will be generated.\n\nWe aim to create `.png` files that split the original `.tif` files into subareas (Horizontal splits here). For the purpose of saving time, we're initially creating subfragments for slices from index 24 to 40. However, if you need a more extensive model, feel free to increase this range. \n\nIn addition to these subfragments, we can also optionally generate `.png` files for the original images. Once all images have been processed, you can compress (zip) the entire directory for ease of upload or download. \n","metadata":{}},{"cell_type":"markdown","source":"## Import Libraries","metadata":{}},{"cell_type":"code","source":"import os\nimport cv2\nfrom tqdm.notebook import tqdm\nimport matplotlib.pyplot as plt","metadata":{"execution":{"iopub.status.busy":"2023-06-16T09:09:12.943608Z","iopub.execute_input":"2023-06-16T09:09:12.943981Z","iopub.status.idle":"2023-06-16T09:09:13.323105Z","shell.execute_reply.started":"2023-06-16T09:09:12.943950Z","shell.execute_reply":"2023-06-16T09:09:13.321775Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Definitions","metadata":{}},{"cell_type":"code","source":"def read_image_mask(fragment_id):\n    images = []\n    idxs = range(from_slice,to_slice)\n    for i in tqdm(idxs):\n        image = cv2.imread(base_path + f\"/train/{fragment_id}/surface_volume/{i:02}.tif\", 0)\n        images.append(image)\n    labels = cv2.imread(base_path + f\"/train/{fragment_id}/inklabels.png\", 0)\n    mask = cv2.imread(base_path + f\"/train/{fragment_id}/mask.png\", 0)\n    return images, labels, mask","metadata":{"execution":{"iopub.status.busy":"2023-06-16T09:09:13.327022Z","iopub.execute_input":"2023-06-16T09:09:13.328050Z","iopub.status.idle":"2023-06-16T09:09:13.335463Z","shell.execute_reply.started":"2023-06-16T09:09:13.328007Z","shell.execute_reply":"2023-06-16T09:09:13.334042Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def show_splits():\n    for j in range(splits): \n        fig, axes = plt.subplots(1, 3, figsize=(12, 4))  # Create a figure with 1 row and 3 columns\n        axes[0].imshow(images[(to_slice-from_slice)//2][j*splitter:(j+1)*splitter])\n        axes[0].set_title('Image')\n        axes[1].imshow(labels[j*splitter:(j+1)*splitter])\n        axes[1].set_title('Labels')\n        axes[2].imshow(mask[j*splitter:(j+1)*splitter])\n        axes[2].set_title('Mask')\n        plt.show()  # Show the figure","metadata":{"execution":{"iopub.status.busy":"2023-06-16T09:09:13.337054Z","iopub.execute_input":"2023-06-16T09:09:13.337571Z","iopub.status.idle":"2023-06-16T09:09:13.352305Z","shell.execute_reply.started":"2023-06-16T09:09:13.337531Z","shell.execute_reply":"2023-06-16T09:09:13.350786Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def save_split(dir_num):\n    # Define the directory and create it if it does not exist\n    dir_path = os.path.join(base_path, save_path, str(dir_num), \"surface_volume\")\n    os.makedirs(dir_path, exist_ok=True)\n    for i in range((to_slice-from_slice)):\n        # Save the image\n        if dir_num in (1,2,3):\n            cv2.imwrite(os.path.join(dir_path, f\"{i:01}.png\"), images[i])\n        else:\n            cv2.imwrite(os.path.join(dir_path, f\"{i:01}.png\"), images[i][s*splitter:(s+1)*splitter])\n    \n    dir_path = os.path.join(base_path, save_path, str(dir_num))\n    if dir_num in (1,2,3):\n        cv2.imwrite(os.path.join(dir_path, \"inklabels.png\"), labels)\n        cv2.imwrite(os.path.join(dir_path, \"mask.png\"), mask)\n    else:\n        cv2.imwrite(os.path.join(dir_path, \"inklabels.png\"), labels[s*splitter:(s+1)*splitter])\n        cv2.imwrite(os.path.join(dir_path, \"mask.png\"), mask[s*splitter:(s+1)*splitter])","metadata":{"execution":{"iopub.status.busy":"2023-06-16T09:09:13.356071Z","iopub.execute_input":"2023-06-16T09:09:13.356553Z","iopub.status.idle":"2023-06-16T09:09:13.369275Z","shell.execute_reply.started":"2023-06-16T09:09:13.356512Z","shell.execute_reply":"2023-06-16T09:09:13.367950Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Create Splits","metadata":{}},{"cell_type":"code","source":"from_slice = 24\nto_slice = 40\n\nbase_path = '/kaggle/input/vesuvius-challenge-ink-detection'\nsave_path = '/kaggle/working/train_sub_png'\nplot = True\nsave_full = True","metadata":{"execution":{"iopub.status.busy":"2023-06-16T09:09:13.370590Z","iopub.execute_input":"2023-06-16T09:09:13.371028Z","iopub.status.idle":"2023-06-16T09:09:13.384221Z","shell.execute_reply.started":"2023-06-16T09:09:13.370987Z","shell.execute_reply":"2023-06-16T09:09:13.383179Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# We split from \nfor fragment_id in range(1,4):\n    images, labels, mask = read_image_mask(fragment_id)\n    if fragment_id == 1:\n        fragment_start = 10\n        splitter = 256 * 16\n    if fragment_id == 2:\n        fragment_start = 20\n        splitter = 256 * 20\n    if fragment_id == 3:\n        fragment_start = 30\n        splitter = 256 * 16\n    splits = int(round(images[0].shape[0]/splitter))\n\n    if plot:\n        show_splits()\n        \n    for s in range(splits): \n        dir_num = fragment_start+s+1\n        save_split(dir_num)\n    if save_full:\n        save_split(fragment_id)","metadata":{"execution":{"iopub.status.busy":"2023-06-16T09:09:13.386079Z","iopub.execute_input":"2023-06-16T09:09:13.386497Z","iopub.status.idle":"2023-06-16T09:14:05.075229Z","shell.execute_reply.started":"2023-06-16T09:09:13.386456Z","shell.execute_reply":"2023-06-16T09:14:05.073973Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]}]}