{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Loading fragment data with the `scrolldata` package\n\nThe `scrolldata` package offers several utilities for loading and working with Vesuvius Challenge data.\n\n[View on Github](https://github.com/mbr4477/scrolldata)","metadata":{}},{"cell_type":"code","source":"!python -m pip install git+https://github.com/mbr4477/scrolldata.git","metadata":{"execution":{"iopub.status.busy":"2023-04-08T12:37:22.633546Z","iopub.execute_input":"2023-04-08T12:37:22.633962Z","iopub.status.idle":"2023-04-08T12:37:59.674388Z","shell.execute_reply.started":"2023-04-08T12:37:22.633927Z","shell.execute_reply":"2023-04-08T12:37:59.672755Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from scrolldata import Scroll, make_patch_dataset\nimport matplotlib.pyplot as plt","metadata":{"execution":{"iopub.status.busy":"2023-04-08T12:37:59.677131Z","iopub.execute_input":"2023-04-08T12:37:59.677607Z","iopub.status.idle":"2023-04-08T12:37:59.684778Z","shell.execute_reply.started":"2023-04-08T12:37:59.677557Z","shell.execute_reply":"2023-04-08T12:37:59.683356Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Loading Fragment Data\n\nUse `Scroll.from_local` to quickly load a scroll from the Kaggle data.","metadata":{}},{"cell_type":"code","source":"scroll = Scroll.from_local(\n    data_dir=\"/kaggle/input/vesuvius-challenge-ink-detection/train/1/surface_volume\",\n    mask_labels_dir=\"/kaggle/input/vesuvius-challenge-ink-detection/train/1\",\n    downsampling=8\n)","metadata":{"execution":{"iopub.status.busy":"2023-04-08T12:37:59.686868Z","iopub.execute_input":"2023-04-08T12:37:59.687386Z","iopub.status.idle":"2023-04-08T12:38:00.097277Z","shell.execute_reply.started":"2023-04-08T12:37:59.687335Z","shell.execute_reply":"2023-04-08T12:38:00.096054Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We can easily plot a few slices!","metadata":{}},{"cell_type":"code","source":"fig, axs = plt.subplots(ncols=5, figsize=(10,5))\n\nfor i in range(3):\n    axs[i].imshow(\n        scroll.load(start_slice=i * 16, num_slices=1)[0],\n        cmap=\"gray\"\n    )\n    axs[i].axis(\"off\")\n    \naxs[3].imshow(scroll.mask, cmap=\"gray\")\naxs[3].axis(\"off\")\naxs[4].imshow(scroll.ink_labels, cmap=\"gray\")\naxs[4].axis(\"off\")\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-04-08T12:38:55.791592Z","iopub.execute_input":"2023-04-08T12:38:55.792044Z","iopub.status.idle":"2023-04-08T12:38:59.466601Z","shell.execute_reply.started":"2023-04-08T12:38:55.792005Z","shell.execute_reply":"2023-04-08T12:38:59.465344Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Create a Patch Data Set\n\nTo improve memory usage, we can precompute a data set of patches sampled from the fragment. We can then lazily load the samples from disk during training.","metadata":{}},{"cell_type":"code","source":"make_patch_dataset(\n    scroll, \n    patch_size=32, \n    num_patches=2000, \n    holdout_region=(0.4,0.4,0.2,0.2),\n    export=\"/kaggle/working/frag1/patches\",\n    show=True\n)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Inside the export directory, there are now `train`, `val`, and `test` directories. Each contains numbered folders with `inputs.npy` and `targets.npy` files of patch data.\n\nThe `scrolldata` package also includes a PyTorch `PatchDataset` implementation to efficiently load the patches from disk for training.","metadata":{}},{"cell_type":"code","source":"from scrolldata.torch import PatchDataset\ntrainset = PatchDataset(\"/kaggle/working/frag1/patches/train\")","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We can plot some examples.","metadata":{}},{"cell_type":"code","source":"fig, axs = plt.subplots(ncols=8, nrows=2, figsize=(8, 2))\nfor i in range(len(axs[0])):\n    axs[0][i].imshow(trainset[i][\"inputs\"][0], cmap=\"gray\")\n    axs[0][i].axis(\"off\")\n    axs[1][i].imshow(trainset[i][\"targets\"][0], cmap=\"gray\")\n    axs[1][i].axis(\"off\")\nplt.show()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Other Features","metadata":{}},{"cell_type":"markdown","source":"Beyond data loading, the `scrolldata` package also has utilities for computing the F0.5 score and run length encoding.","metadata":{}}]}