{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# 🌋[EDA] Vesuvius Challenge - Ink Detection\n---\n**Welcome to my first notebook for the [Ink Detection competition](https://www.kaggle.com/competitions/vesuvius-challenge-ink-detection/overview).** As is customary when beginning a competition, it's important to understand what data are available. By exploring and familiarizing ourselves with the data, we can use them more effectively and gain a better understanding of what we're doing. In this notebook, I will be performing this preliminary work.\n\n**[Here](https://www.youtube.com/watch?v=PpNq2cFotyY) is an interesting introduction video about the project.**\n\n<a href=\"https://zupimages.net/viewer.php?id=23/12/esn8.png\"><img src=\"https://zupimages.net/up/23/12/esn8.png\" alt=\"\" /></a>\n<figcaption><u>Midjourney generation</u></figcaption>\n\n## About the competition\nThe project aims to resurrect an ancient library that was **buried under the ashes of the Vesuvius volcano** nearly 2000 years ago. The library was located in a Roman villa in **Herculaneum**, a town near Pompeii, and contained thousands of scrolls. However, due to the extreme heat of the Vesuvius eruption, **the scrolls were carbonized** and have remained **unreadable** ever since. The competition requires participants to use **3D X-ray scans** to detect ink and read the contents of the scrolls. The discovery of these scrolls several hundred years ago has led to the quest to read them using modern techniques.\n\nThe goal of the competition is to produce a **binary image** of the same size as the scanned fragments, where the ink pixels are represented by 1. This is a segmentation task, and the score will be calculated using the **[Sørensen–Dice coefficient](https://en.wikipedia.org/wiki/S%C3%B8rensen%E2%80%93Dice_coefficient)**, which is essentially the F1-score.\n\nWe hope that this segmentation will help **reveal the knowledge** contained in the ancient texts.\n\n**Let's get started!**","metadata":{}},{"cell_type":"markdown","source":"# Imports","metadata":{}},{"cell_type":"code","source":"!pip install celluloid -q\n\nimport numpy as np\nimport matplotlib.pyplot as plt\nimport pandas as pd\nimport cv2\nfrom IPython.display import HTML, display\n\nfrom celluloid import Camera\nfrom tqdm.notebook import tqdm\nimport gc\nimport seaborn as sns","metadata":{"execution":{"iopub.status.busy":"2023-03-23T10:08:16.832433Z","iopub.execute_input":"2023-03-23T10:08:16.832847Z","iopub.status.idle":"2023-03-23T10:08:17.715142Z","shell.execute_reply.started":"2023-03-23T10:08:16.832803Z","shell.execute_reply":"2023-03-23T10:08:17.713739Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Function to display the scan layers (volume)\nThe following functions create a video to visualize the contents of a given papyrus fragment by combining the scans from the dataset. The dataset contains 5 different scans:\n- **Train: 1 / 2 / 3**\n- **Test: a / b**\n\nThe `display_volume` function uses those values.\n\nThanks to *Leonid Kulyk* for inspiring me this function: [Source](https://www.kaggle.com/code/leonidkulyk/eda-vc-id-volume-layers-animation/notebook)","metadata":{}},{"cell_type":"code","source":"N_SCAN_LAYERS = 65\n\n# Load the volume of the given papyrus fragment\ndef load_volume(set_name, folder_name):\n    path = f'/kaggle/input/vesuvius-challenge-ink-detection/{set_name}/{folder_name}/surface_volume'\n    file_format = \"{:02d}.tif\"\n    volume = []\n    \n    for i in tqdm(range(N_SCAN_LAYERS)):\n        file = f'{path}/{file_format.format(i)}'\n        img = cv2.imread(file, cv2.IMREAD_ANYDEPTH)\n        volume.append(img)\n        \n    return volume\n\n# Display the volume as a video\ndef display_volume(set_name, folder_name):\n    volume = load_volume(set_name, folder_name)\n    fig, ax = plt.subplots()\n    camera = Camera(fig)\n    \n    for i in range(N_SCAN_LAYERS):\n        ax.axis('off')\n        ax.text(0.5, 1.08, f\"Layer {i+1}/{N_SCAN_LAYERS}\", fontweight='bold', fontsize=18,\n                transform=ax.transAxes, horizontalalignment='center')\n        ax.imshow(volume[0], cmap='gray')\n        camera.snap()\n        del volume[0]\n        gc.collect()\n        \n    plt.close(fig)\n    \n    animation = camera.animate()\n    fix_video_adjust = \\\n    '<style> video {margin: 0px; padding: 0px; width:100%; height:auto;} </style>'\n    display(HTML(fix_video_adjust + animation.to_html5_video()))\n    \n    del camera\n    del animation\n    gc.collect()","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2023-03-23T10:08:17.717322Z","iopub.execute_input":"2023-03-23T10:08:17.717768Z","iopub.status.idle":"2023-03-23T10:08:17.730417Z","shell.execute_reply.started":"2023-03-23T10:08:17.717724Z","shell.execute_reply":"2023-03-23T10:08:17.728631Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We will use this function later in the notebook. ","metadata":{}},{"cell_type":"markdown","source":"# Shape of the images\nAs said earlier, the dataset contains **5 fragments** (3 for the training set / 2 for the test set).\n\nAll images present in a papyrus fragment subdirectory have the same size, but the different scans have different sizes. **Let's compare them.**\n\n","metadata":{}},{"cell_type":"code","source":"def get_shape(set_name, folder_name):\n    file = f'/kaggle/input/vesuvius-challenge-ink-detection/{set_name}/{folder_name}/mask.png'\n    img = cv2.imread(file, cv2.IMREAD_ANYDEPTH)\n    return img.shape\n\ndef get_shape_info():\n    samples = []\n    shapes = []\n    px_nbs = []\n    fragments = {'train': ['1', '2', '3'], 'test': ['a', 'b']}\n    for set_name in fragments:\n        for folder_name in fragments[set_name]:\n            samples.append(f'{set_name} {folder_name}')\n            shape = get_shape(set_name, folder_name)\n            shapes.append(shape)\n            px_nbs.append(shape[0] * shape[1])\n    df = pd.DataFrame([shapes, px_nbs], columns=samples, index=['Shape', 'Nb of pixels']).T\n    \n    return df\n\ndf = get_shape_info()\n\nplt.figure(figsize=(10, 4))\nplt.subplot(1, 2, 1)\nsns.barplot(x=df.index, y=df['Nb of pixels'], )\nplt.title(\"Size in pixels for each scan\")\nplt.subplot(1, 2, 2)\nplt.pie(df['Nb of pixels'], labels=df.index, autopct='%0.1f%%')\nplt.title(\"Pixel size as percentage\")\nplt.show()\n\ndf","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2023-03-23T10:09:36.134742Z","iopub.execute_input":"2023-03-23T10:09:36.135153Z","iopub.status.idle":"2023-03-23T10:09:38.273317Z","shell.execute_reply.started":"2023-03-23T10:09:36.135113Z","shell.execute_reply":"2023-03-23T10:09:38.271974Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Insights\n- As we can see, the images have **different sizes** in terms of the number of pixels\n- Train 2 is the largest (almost 50% of the total size), while the scan a in the test set is the smallest.\n- The images in the **training set are landscape oriented** (width is larger than height)\n- The images in the **test set are portray oriented** (height is larger than width)","metadata":{}},{"cell_type":"markdown","source":"# Infrareds, Inklabels and Masks\nThe dataset also includes PNG files, but only the train set contains \"ir\" and \"inklabels\" files since these are information about **what we want to predict.** The masks are binary images that delimit the fragment from the background.\n\n**Let's look at all of these files.**","metadata":{}},{"cell_type":"code","source":"def display_png_files():\n    # Train set\n    files = ['ir.png', 'inklabels.png', 'mask.png']\n    plt.figure(figsize=(12, 18))\n    k = 1\n    for i in range(1, 4):\n        folder = f'/kaggle/input/vesuvius-challenge-ink-detection/train/{i}/'\n        for file in files:\n            filepath = folder + file\n            img = cv2.imread(filepath, cv2.IMREAD_ANYDEPTH)\n            plt.subplot(3, 3, k)\n            plt.imshow(img, cmap='gray')\n            plt.title(f\"Papyrus {i} | {file}\")\n            k += 1\n    plt.show()\n    \n    # Test set\n    plt.figure(figsize=(12, 9))\n    for k, a in enumerate(['a', 'b']):\n        filepath = f'/kaggle/input/vesuvius-challenge-ink-detection/test/{a}/mask.png'\n        img = cv2.imread(filepath, cv2.IMREAD_ANYDEPTH)\n        img.shape\n        plt.subplot(1, 2, k+1)\n        plt.imshow(img, cmap='gray')\n        plt.title(f\"Papyrus {a} | mask.png\")\n    plt.show()\n        \ndisplay_png_files()","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2023-03-23T10:08:20.248842Z","iopub.execute_input":"2023-03-23T10:08:20.249195Z","iopub.status.idle":"2023-03-23T10:08:52.577237Z","shell.execute_reply.started":"2023-03-23T10:08:20.249162Z","shell.execute_reply":"2023-03-23T10:08:52.576097Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Insights\n- The **masks** define the **boundary** between the papyrus fragment and the background.\n- The infrared files reveal the ink on the paper, which is then used to create the label images.\n- The label images are **binary images** where white pixels indicate the location of the ink.","metadata":{}},{"cell_type":"markdown","source":"# How much space takes the ink\nThe task is binary segmentation, aimed at identifying the locations of ink on the paper. Each pixel on the paper can have a value of either 0 or 1.\n\nWe will **calculate the proportion of paper covered by ink** for each papyrus fragment using the masks and ink labels.","metadata":{}},{"cell_type":"code","source":"ratios = []\nfor i in range(1, 4):\n    file_mask = f'/kaggle/input/vesuvius-challenge-ink-detection/train/{i}/mask.png'\n    file_inklabels = f'/kaggle/input/vesuvius-challenge-ink-detection/train/{i}/inklabels.png'\n    mask = cv2.imread(file_mask, cv2.IMREAD_ANYDEPTH)\n    inklabels = cv2.imread(file_inklabels, cv2.IMREAD_ANYDEPTH)\n    total_pixels = np.sum(mask)\n    ink_pixels = np.sum(inklabels)\n    ratios.append(round(ink_pixels / total_pixels * 100, 2))\n\nplt.figure(figsize=(4, 4))\nplots = sns.barplot(x=[\"Train 1\", \"Train 2\", \"Train 3\"], y=ratios)\nfor bar in plots.patches:\n    plots.annotate(format(bar.get_height(), '.2f') + '%',\n                   (bar.get_x() + bar.get_width() / 2,\n                    bar.get_height() * 0.4), ha='center',\n                   size=10, xytext=(0, 8),\n                   textcoords='offset points')\nplt.title(\"Proportion of ink on each papyrus fragment\")\nplt.show()","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2023-03-23T10:08:52.578762Z","iopub.execute_input":"2023-03-23T10:08:52.579558Z","iopub.status.idle":"2023-03-23T10:08:56.256806Z","shell.execute_reply.started":"2023-03-23T10:08:52.579510Z","shell.execute_reply":"2023-03-23T10:08:56.255369Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Insights\n- There is **less than 20%** of ink on each fragment.\n- We can expect having similar values when prediction for the test set.\n- The predictions should only be made on the white space of the masks.","metadata":{}},{"cell_type":"markdown","source":"# Look at the scan volums\nWe will use the function that was created at the beginning of the notebook to display all the **layers** of each fragment. When combined, they create the **volume of the fragment.** \n\n## - Train 1","metadata":{}},{"cell_type":"code","source":"display_volume('train', '1')","metadata":{"execution":{"iopub.status.busy":"2023-03-23T10:08:56.258347Z","iopub.execute_input":"2023-03-23T10:08:56.259224Z","iopub.status.idle":"2023-03-23T10:09:03.427156Z","shell.execute_reply.started":"2023-03-23T10:08:56.259183Z","shell.execute_reply":"2023-03-23T10:09:03.425697Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## - Train 2","metadata":{}},{"cell_type":"code","source":"display_volume('train', '2')","metadata":{"execution":{"iopub.status.busy":"2023-03-23T10:09:03.429101Z","iopub.execute_input":"2023-03-23T10:09:03.429601Z","iopub.status.idle":"2023-03-23T10:09:19.749285Z","shell.execute_reply.started":"2023-03-23T10:09:03.429547Z","shell.execute_reply":"2023-03-23T10:09:19.747811Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## - Train 3","metadata":{}},{"cell_type":"code","source":"display_volume('train', '3')","metadata":{"execution":{"iopub.status.busy":"2023-03-23T10:09:19.754191Z","iopub.execute_input":"2023-03-23T10:09:19.754749Z","iopub.status.idle":"2023-03-23T10:09:25.579473Z","shell.execute_reply.started":"2023-03-23T10:09:19.754700Z","shell.execute_reply":"2023-03-23T10:09:25.578015Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## - Test a","metadata":{}},{"cell_type":"code","source":"display_volume('test', 'a')","metadata":{"execution":{"iopub.status.busy":"2023-03-23T10:09:25.581441Z","iopub.execute_input":"2023-03-23T10:09:25.582974Z","iopub.status.idle":"2023-03-23T10:09:29.327454Z","shell.execute_reply.started":"2023-03-23T10:09:25.582888Z","shell.execute_reply":"2023-03-23T10:09:29.326087Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## - Test b","metadata":{}},{"cell_type":"code","source":"display_volume('test', 'b')","metadata":{"execution":{"iopub.status.busy":"2023-03-23T10:09:29.329224Z","iopub.execute_input":"2023-03-23T10:09:29.329721Z","iopub.status.idle":"2023-03-23T10:09:36.131787Z","shell.execute_reply.started":"2023-03-23T10:09:29.329675Z","shell.execute_reply":"2023-03-23T10:09:36.130422Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Insights:\n- As you can see, it is **impossible to read anything** from these scans.\n- Hopefully a deep learning model will be able to **find patterns** from the ink present on the paper.","metadata":{}},{"cell_type":"markdown","source":"# End\n\n**Thank you for reading!**\n\nIf you enjoyed my work, please upvote it. I would also appreciate any feedback or suggestions for improvement. Thank you!","metadata":{}}]}