{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.11.13","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":117682,"databundleVersionId":14443416,"sourceType":"competition"}],"dockerImageVersionId":31192,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Access to additional unlabeled data\n\n## Dependencies\nLet's install the vesuvius library that will be used to access the available CT scans.\nThe code is available [here](https://github.com/ScrollPrize/villa/tree/main/vesuvius).","metadata":{}},{"cell_type":"code","source":"# install vesuvius & dependencies\n!pip install --quiet vesuvius matplotlib > /dev/null 2>&1\n\n# Accept the Vesuvius Challenge data license (required before first use).\n# This is a non-interactive acceptance; see https://scrollprize.org/data for details.\n!vesuvius.accept_terms --yes","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true,"execution":{"iopub.status.busy":"2025-11-26T22:49:07.447588Z","iopub.execute_input":"2025-11-26T22:49:07.447907Z","iopub.status.idle":"2025-11-26T22:49:30.635793Z","shell.execute_reply.started":"2025-11-26T22:49:07.447873Z","shell.execute_reply":"2025-11-26T22:49:30.634859Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Import and chunk size\nWe are now going to import basic python libraries, like **numpy**, **matplotlib**, and the recently installed **vesuvius** library.\nWe will also define some basic plot properties and the size of an isotropic chunk we want to extract from the CT scan of the scrolls.","metadata":{}},{"cell_type":"code","source":"# imports and plotting config\nimport numpy as np\nimport matplotlib.pyplot as plt\n\nimport vesuvius\nfrom vesuvius import Volume\n\nplt.rcParams[\"figure.figsize\"] = (6, 6)\nplt.rcParams[\"image.cmap\"] = \"gray\"  # grayscale is easier for CT volumes\n\nCHUNK_SIZE = 256","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-11-26T22:49:40.845173Z","iopub.execute_input":"2025-11-26T22:49:40.845552Z","iopub.status.idle":"2025-11-26T22:49:43.231939Z","shell.execute_reply.started":"2025-11-26T22:49:40.845471Z","shell.execute_reply":"2025-11-26T22:49:43.230885Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"We have scanned more than 30 scrolls, but most of them are unreleased.\n\nThe URLs to the CT scans of the released scrolls are available [here](https://scrollprize.org/data_scrolls) and the URLs to fragments (pieces of scrolls mechanically detached) are [here](https://scrollprize.org/data_fragments).\n\nBe careful because not all of them are available as Zarr (the format we need). Some of them could have been normalized/compressed in a non-canonical way.\n\nWe recommend working on these two, which were acquired using a protocol we used for the majority of our unreleased scans:\n- PHerc. 0139 _(scroll)_\n- PHerc. 0009B _(fragment)_\n","metadata":{}},{"cell_type":"code","source":"PHERC_0139_URL_9um = \"https://data.aws.ash2txt.org/samples/PHerc0139/volumes/20250728140407-9.362um-1.2m-113keV-masked.zarr/\"\nPHERC_0009B_URL_9um = \"https://data.aws.ash2txt.org/samples/PHerc0009B/volumes/20250521125136-8.640um-1.2m-116keV-masked.zarr/\"","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-11-26T22:50:12.633176Z","iopub.execute_input":"2025-11-26T22:50:12.633839Z","iopub.status.idle":"2025-11-26T22:50:12.638680Z","shell.execute_reply.started":"2025-11-26T22:50:12.633809Z","shell.execute_reply":"2025-11-26T22:50:12.637658Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"These zarr stores are actually OME-Zarr, and the folder contains several subfolders corresponding to downscaled versions of the same volume.\n\nWe will access to the native scale which is in subfolder **0**.\n\nThe conventional coordinate frame in the volumes is _Z,Y,X_.","metadata":{}},{"cell_type":"code","source":"scroll = Volume(type=\"zarr\", path=PHERC_0139_URL_9um+\"0\")\nprint(f\"Shape: {scroll.shape()}\")\nprint(f\"dtype: {scroll.dtype}\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-11-26T22:50:52.256881Z","iopub.execute_input":"2025-11-26T22:50:52.257232Z","iopub.status.idle":"2025-11-26T22:50:53.706891Z","shell.execute_reply.started":"2025-11-26T22:50:52.257207Z","shell.execute_reply":"2025-11-26T22:50:53.705688Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"Great! We lazy loaded the scan! Let's download a chunk from the full volume by specifying its bounding box!","metadata":{}},{"cell_type":"code","source":"chunk = scroll[10500:10500+CHUNK_SIZE,3000:3000+CHUNK_SIZE,3000:3000+CHUNK_SIZE]\nprint(f\"Chunk shape: {chunk.shape}\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-11-26T22:51:25.017891Z","iopub.execute_input":"2025-11-26T22:51:25.018205Z","iopub.status.idle":"2025-11-26T22:51:26.899312Z","shell.execute_reply.started":"2025-11-26T22:51:25.018181Z","shell.execute_reply":"2025-11-26T22:51:26.898319Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"Let's visualize different slices!","metadata":{}},{"cell_type":"code","source":"fig, axes = plt.subplots(1, 3, figsize=(15, 5))\n\n# Z=100, XY\naxes[0].imshow(chunk[100, :, :])\naxes[0].set_title('XY plane')\naxes[0].axis('off')\n\n# Y=100, XZ\naxes[1].imshow(chunk[:, 100, :])\naxes[1].set_title('XZ plane')\naxes[1].axis('off')\n\n# X=100, YZ\naxes[2].imshow(chunk[:, :, 100])\naxes[2].set_title('YZ plane')\naxes[2].axis('off')\n\nplt.tight_layout()\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-11-26T22:51:56.640572Z","iopub.execute_input":"2025-11-26T22:51:56.641256Z","iopub.status.idle":"2025-11-26T22:51:57.298832Z","shell.execute_reply.started":"2025-11-26T22:51:56.641229Z","shell.execute_reply":"2025-11-26T22:51:57.296884Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"Cool! We have extracted a chunk of _unlabeled_ data that could be used in the competition, either in _unsupervised_ or, after labeling, in _supervised_ approaches!\n\n**Good luck!**","metadata":{}}]}