{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.11.13","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":117682,"databundleVersionId":14443416,"sourceType":"competition"}],"dockerImageVersionId":31192,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Convert Surface 📜 dataset to npz\n\nThis notebook focuses on the initial but vital step of data preparation for the Vesuvius Challenge. It addresses the need to convert large TIFF image and label files into a more manageable and efficient format. \n\n*   Utilize `np.savez_compressed` to save image and label data as separate, compressed `.npz` files, optimizing storage and load times.\n*   Implement distinct processing loops for images and labels, enhancing code clarity and allowing for separate handling if needed.\n\nThis preprocessing step lays the groundwork for subsequent stages of the challenge, such as model training and inference, by providing clean, organized, and efficiently stored data.\n\n---\n\n**Following discussion in https://www.kaggle.com/competitions/vesuvius-challenge-surface-detection/discussion/627529**","metadata":{}},{"cell_type":"code","source":"!pip install -q imagecodecs","metadata":{"trusted":true,"_kg_hide-input":true,"_kg_hide-output":true,"execution":{"iopub.status.busy":"2025-11-20T06:57:44.047357Z","iopub.execute_input":"2025-11-20T06:57:44.047697Z","iopub.status.idle":"2025-11-20T06:57:51.046937Z","shell.execute_reply.started":"2025-11-20T06:57:44.047675Z","shell.execute_reply":"2025-11-20T06:57:51.045027Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"import os\nimport glob\nimport imagecodecs\nimport tifffile\nimport numpy as np\nfrom tqdm.auto import tqdm\n\nroot_dir = \"/kaggle/input/vesuvius-challenge-surface-detection\"\noutput_dir = \"/kaggle/working\"\nfolders = [\"train_images\", \"train_labels\"]\n\nfor folder in folders:\n    folder_path = os.path.join(root_dir, folder)\n    all_files = sorted(glob.glob(os.path.join(folder_path, \"*.tif\")))\n    local_folder = os.path.join(output_dir, folder)\n    os.makedirs(local_folder, exist_ok=True)\n    for file in tqdm(all_files, desc=f\"Processing {folder}\"):\n        base, _ = os.path.splitext(os.path.basename(file))\n        img = tifffile.imread(file)\n        np.savez_compressed(f\"{local_folder}/{base}.npz\", img)","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true,"execution":{"iopub.status.busy":"2025-11-20T06:57:51.048908Z","iopub.execute_input":"2025-11-20T06:57:51.049289Z","execution_failed":"2025-11-20T07:16:53.466Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Upload from Colab\n\n### Why Colab?\nProcessing the Vesuvius Challenge dataset requires significant storage and memory. We use **Google Colab** to bypass local hardware limitations (disk space) and leverage cloud resources.\n\n### Trivial Step-by-Step: Open Kaggle in Colab\n1. **Get API Key**: Go to your [Kaggle Account](https://www.kaggle.com/me/account), scroll to \"API\", and click **Create New API Token**. This downloads a `kaggle.json` file.\n2. **Configure Colab**: \n   - Upload `kaggle.json` to the `/root/.config/kaggle/` folder (create it if missing).\n   - **OR** simply run the `kagglehub.login()` cell which will prompt you to authorize via browser.","metadata":{}},{"cell_type":"code","source":"!echo '{\"username\":\"your-name\",\"key\":\"KEY-HASH-HERE\"}' > /root/.config/kaggle/kaggle.json","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"import json\nimport os\nfrom kaggle.api.kaggle_api_extended import KaggleApi\n\n# Ensure output_dir is defined\n# if 'output_dir' not in locals():\n#     output_dir = \"dataset\"\n\n# Fetch Kaggle username for the metadata ID\napi = KaggleApi()\napi.authenticate()\nusername = api.config_values['username']\n\n# Create dataset-metadata.json\nmetadata = {\n    \"title\": \"Vesuvius Challenge: Surface Detection [NPZ]\",\n    \"id\": f\"{username}/vesuvius-surface-npz\",\n    \"licenses\": [{\"name\": \"CC0-1.0\"}]\n}\n\nmetadata_path = os.path.join(output_dir, \"dataset-metadata.json\")\nwith open(metadata_path, \"w\") as f:\n    json.dump(metadata, f, indent=4)\n\nprint(f\"Metadata saved to {metadata_path}\")\n\n# CLI command to create the dataset\n# --dir-mode zip creates a zip archive of the folder contents before uploading\n!kaggle datasets create -p {output_dir} --dir-mode zip","metadata":{"trusted":true},"outputs":[],"execution_count":null}]}