{"metadata":{"kernelspec":{"name":"python3","display_name":"Python 3","language":"python"},"language_info":{"name":"python","version":"3.10.12","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"gpu","dataSources":[{"sourceId":91196,"databundleVersionId":11377716,"sourceType":"competition"}],"dockerImageVersionId":30919,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":true}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"## Simple baseline with Bioclimatic Cubes — ResNet18 + Binary Cross Entropy\n\nTo demonstrate the potential of single modality data such as Bioclimatic cubes, we provide a straightforward baseline that is baseline on a modified ResNet18 and Binary Cross Entropy but still ranks highly on the leaderboard. The model itself should learn the relationship between the precise climatic history of a given location and its species composition.\n\nConsidering the significant extent for enhancing performance of this baseline, we encourage you to experiment with various techniques, architectures, losses, etc.\n\n#### **Have Fun!**","metadata":{}},{"cell_type":"code","source":"import os\nimport torch\nimport tqdm\nimport numpy as np\nimport pandas as pd\nimport torchvision.models as models\nimport torchvision.transforms as transforms\nimport torch.nn as nn\nfrom torch.utils.data import Dataset, DataLoader\nfrom torch.optim.lr_scheduler import CosineAnnealingLR\nfrom sklearn.metrics import precision_recall_fscore_support","metadata":{"ExecuteTime":{"end_time":"2024-04-30T21:25:07.298310Z","start_time":"2024-04-30T21:25:05.354584Z"},"tags":[],"trusted":true,"execution":{"iopub.status.busy":"2025-03-13T20:09:50.711789Z","iopub.execute_input":"2025-03-13T20:09:50.712144Z","iopub.status.idle":"2025-03-13T20:09:50.717408Z","shell.execute_reply.started":"2025-03-13T20:09:50.712109Z","shell.execute_reply":"2025-03-13T20:09:50.716447Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Data description\n\nThe Bioclimatic Cubes are created from **four** monthly GeoTIFF CHELSA (https://chelsa-climate.org/timeseries/) time series climatic rasters with a resolution of 30 arc seconds, i.e. approximately 1km. The four variables are the precipitation (pr), maximum- (taxmax), minimum- (tasmin), and mean (tax) daily temperatures per month from January 2000 to June 2019. We provide the data in three forms: (i) raw rasters (GeoTiff images), (ii) CSV file with pre-extracted values for each location, i.e., surveyId, and (iii) data cubes as tensor object (.pt).\n\nIn this notebook, we will work with just the cubes. The cubes are structured as follows.\n**Shape**: `(n_year, n_month, n_bio)` where:\n- `n_year` = 19 (ranging from 2000 to 2018)\n- `n_month` = 12 (ranging from January 01 to December 12)\n- `n_bio` = 4 comprising [`pr` (precipitation), `tas` (mean daily air temperature), `tasmin`, `tasmax`]\n\nThe datacubes can simply be loaded as tensors using PyTorch with the following command :\n\n```python\nimport torch\ntorch.load('path_to_file.pt')\n```\n\n**References:**\n- *Karger, D.N., Conrad, O., Böhner, J., Kawohl, T., Kreft, H., Soria-Auza, R.W., Zimmermann, N.E., Linder, P., Kessler, M. (2017): Climatologies at high resolution for the Earth land surface areas. Scientific Data. 4 170122. https://doi.org/10.1038/sdata.2017.122*\n\n- *Karger D.N., Conrad, O., Böhner, J., Kawohl, T., Kreft, H., Soria-Auza, R.W., Zimmermann, N.E, Linder, H.P., Kessler, M. Data from: Climatologies at high resolution for the earth’s land surface areas. Dryad Digital Repository. http://dx.doi.org/doi:10.5061/dryad.kd1d4*","metadata":{"execution":{"iopub.status.busy":"2024-05-01T13:30:07.053659Z","iopub.execute_input":"2024-05-01T13:30:07.054038Z","iopub.status.idle":"2024-05-01T13:30:07.058148Z","shell.execute_reply.started":"2024-05-01T13:30:07.054008Z","shell.execute_reply":"2024-05-01T13:30:07.057269Z"}}},{"cell_type":"markdown","source":"## Prepare custom dataset loader\n\nWe have to sloightly update the Dataset to provide the relevant data in the appropriate format.","metadata":{}},{"cell_type":"code","source":"class TrainDataset(Dataset):\n    def __init__(self, data_dir, metadata, subset, transform=None):\n        self.subset = subset\n        self.transform = transform\n        self.data_dir = data_dir\n        self.metadata = metadata\n        self.metadata = self.metadata.dropna(subset=\"speciesId\").reset_index(drop=True)\n        self.metadata['speciesId'] = self.metadata['speciesId'].astype(int)\n        self.label_dict = self.metadata.groupby('surveyId')['speciesId'].apply(list).to_dict()\n        \n        self.metadata = self.metadata.drop_duplicates(subset=\"surveyId\").reset_index(drop=True)\n\n    def __len__(self):\n        return len(self.metadata)\n\n    def __getitem__(self, idx):\n        \n        survey_id = self.metadata.surveyId[idx]\n        sample = torch.load(os.path.join(self.data_dir, f\"GLC25-PA-{self.subset}-bioclimatic_monthly_{survey_id}_cube.pt\"), weights_only=True)\n        species_ids = self.label_dict.get(survey_id, [])  # Get list of species IDs for the survey ID\n        label = torch.zeros(num_classes)  \n        for species_id in species_ids:\n            label_id = species_id\n            label[label_id] = 1  # Set the corresponding class index to 1 for each species\n\n        # Ensure the sample is in the correct format for the transform\n        if isinstance(sample, torch.Tensor):\n            sample = sample.permute(1, 2, 0)  # Change tensor shape from (C, H, W) to (H, W, C)\n            sample = sample.numpy()  \n\n        if self.transform:\n            sample = self.transform(sample)\n\n        return sample, label, survey_id\n    \nclass TestDataset(TrainDataset):\n    def __init__(self, data_dir, metadata, subset, transform=None):\n        self.subset = subset\n        self.transform = transform\n        self.data_dir = data_dir\n        self.metadata = metadata\n        \n    def __getitem__(self, idx):\n        \n        survey_id = self.metadata.surveyId[idx]\n        sample = torch.load(os.path.join(self.data_dir, f\"GLC25-PA-{self.subset}-bioclimatic_monthly_{survey_id}_cube.pt\"), weights_only=True)\n        \n        if isinstance(sample, torch.Tensor):\n            sample = sample.permute(1, 2, 0)  # Change tensor shape from (C, H, W) to (H, W, C)\n            sample = sample.numpy()\n\n        if self.transform:\n            sample = self.transform(sample)\n\n        return sample, survey_id","metadata":{"ExecuteTime":{"end_time":"2024-04-30T21:25:32.627928Z","start_time":"2024-04-30T21:25:32.612131Z"},"collapsed":false,"jupyter":{"outputs_hidden":false},"tags":[],"execution":{"iopub.status.busy":"2025-03-13T20:10:15.998749Z","iopub.execute_input":"2025-03-13T20:10:15.999141Z","iopub.status.idle":"2025-03-13T20:10:16.010140Z","shell.execute_reply.started":"2025-03-13T20:10:15.999112Z","shell.execute_reply":"2025-03-13T20:10:16.009035Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### Load metadata and prepare data loaders","metadata":{}},{"cell_type":"code","source":"# Dataset and DataLoader\nbatch_size = 64\ntransform = transforms.Compose([\n    transforms.ToTensor()\n])\n\n# Load Training metadata\ntrain_data_path = \"/kaggle/input/geolifeclef-2025/BioclimTimeSeries/cubes/PA-train\"\ntrain_metadata_path = \"/kaggle/input/geolifeclef-2025/GLC25_PA_metadata_train.csv\"\ntrain_metadata = pd.read_csv(train_metadata_path)\ntrain_dataset = TrainDataset(train_data_path, train_metadata, subset=\"train\", transform=transform)\ntrain_loader = DataLoader(train_dataset, batch_size=batch_size, shuffle=True, num_workers=4)\n\n# Load Test metadata\ntest_data_path = \"/kaggle/input/geolifeclef-2025/BioclimTimeSeries/cubes/PA-test\"\ntest_metadata_path = \"/kaggle/input/geolifeclef-2025/GLC25_PA_metadata_test.csv\"\ntest_metadata = pd.read_csv(test_metadata_path)\ntest_dataset = TestDataset(test_data_path, test_metadata, subset=\"test\", transform=transform)\ntest_loader = DataLoader(test_dataset, batch_size=batch_size, shuffle=False, num_workers=4)","metadata":{"ExecuteTime":{"end_time":"2024-04-30T21:25:34.532017Z","start_time":"2024-04-30T21:25:32.615562Z"},"collapsed":false,"jupyter":{"outputs_hidden":false},"tags":[],"execution":{"iopub.status.busy":"2025-03-13T20:10:19.425367Z","iopub.execute_input":"2025-03-13T20:10:19.425671Z","iopub.status.idle":"2025-03-13T20:10:22.786141Z","shell.execute_reply.started":"2025-03-13T20:10:19.425648Z","shell.execute_reply":"2025-03-13T20:10:22.785444Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Define and initialize the ModifiedResNet18 model\n\nTo utilize the bioclimatic cubes, which have a shape of [4,19,12] (RASTER-TYPE, YEAR, and MONTH), some minor adjustments must be made to the vanilla ResNet-18. It's important to note that this is just one method for ensuring compatibility with the unusual tensor shape, and experimentation is encouraged.","metadata":{}},{"cell_type":"code","source":"class ModifiedResNet18(nn.Module):\n    def __init__(self, num_classes):\n        super(ModifiedResNet18, self).__init__()\n\n        self.norm_input = nn.LayerNorm([4,19,12])\n        self.resnet18 = models.resnet18(weights=None)\n        # We have to modify the first convolutional layer to accept 4 channels instead of 3\n        self.resnet18.conv1 = nn.Conv2d(4, 64, kernel_size=3, stride=1, padding=1, bias=False)\n        self.resnet18.maxpool = nn.Identity()\n        self.ln = nn.LayerNorm(1000)\n        self.fc1 = nn.Linear(1000, 2056)\n        self.fc2 = nn.Linear(2056, num_classes)\n\n    def forward(self, x):\n        x = self.norm_input(x)\n        x = self.resnet18(x)\n        x = self.ln(x)\n        x = self.fc1(x)\n        x = self.fc2(x)\n        return x","metadata":{"ExecuteTime":{"end_time":"2024-04-30T21:25:31.014067Z","start_time":"2024-04-30T21:25:31.010060Z"},"collapsed":false,"jupyter":{"outputs_hidden":false},"tags":[],"execution":{"iopub.status.busy":"2025-03-13T19:24:41.227745Z","iopub.execute_input":"2025-03-13T19:24:41.228149Z","iopub.status.idle":"2025-03-13T19:24:41.236236Z","shell.execute_reply.started":"2025-03-13T19:24:41.228125Z","shell.execute_reply":"2025-03-13T19:24:41.235502Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"def set_seed(seed):\n    # Set seed for Python's built-in random number generator\n    torch.manual_seed(seed)\n    # Set seed for numpy\n    np.random.seed(seed)\n    # Set seed for CUDA if available\n    if torch.cuda.is_available():\n        torch.cuda.manual_seed_all(seed)\n        # Set cuDNN's random number generator seed for deterministic behavior\n        torch.backends.cudnn.deterministic = True\n        torch.backends.cudnn.benchmark = False\n\nset_seed(69)","metadata":{"tags":[],"execution":{"iopub.status.busy":"2025-03-13T19:24:41.238441Z","iopub.execute_input":"2025-03-13T19:24:41.238791Z","iopub.status.idle":"2025-03-13T19:24:41.351677Z","shell.execute_reply.started":"2025-03-13T19:24:41.238754Z","shell.execute_reply":"2025-03-13T19:24:41.350752Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Check if cuda is available\ndevice = torch.device(\"cpu\")\n\nif torch.cuda.is_available():\n    device = torch.device(\"cuda\")\n    print(\"DEVICE = CUDA\")\n\nnum_classes = 11255 # Number of all unique classes within the PO and PA data.\nmodel = ModifiedResNet18(num_classes).to(device)","metadata":{"ExecuteTime":{"end_time":"2024-04-30T21:25:31.611823Z","start_time":"2024-04-30T21:25:31.607373Z"},"collapsed":false,"jupyter":{"outputs_hidden":false},"tags":[],"execution":{"iopub.status.busy":"2025-03-13T19:24:41.353022Z","iopub.execute_input":"2025-03-13T19:24:41.353308Z","iopub.status.idle":"2025-03-13T19:24:41.968879Z","shell.execute_reply.started":"2025-03-13T19:24:41.353281Z","shell.execute_reply":"2025-03-13T19:24:41.968221Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Training Loop\n\nNothing special, just a standard Pytorch training loop.","metadata":{}},{"cell_type":"code","source":"# Hyperparameters\nlearning_rate = 0.0002\nnum_epochs = 20\npositive_weigh_factor = 1.0\n\noptimizer = torch.optim.AdamW(model.parameters(), lr=learning_rate)\nscheduler = CosineAnnealingLR(optimizer, T_max=25, verbose=True)","metadata":{"ExecuteTime":{"end_time":"2024-04-30T21:25:32.181927Z","start_time":"2024-04-30T21:25:32.177073Z"},"collapsed":false,"jupyter":{"outputs_hidden":false},"tags":[],"execution":{"iopub.status.busy":"2025-03-13T19:24:41.969714Z","iopub.execute_input":"2025-03-13T19:24:41.970046Z","iopub.status.idle":"2025-03-13T19:24:41.976618Z","shell.execute_reply.started":"2025-03-13T19:24:41.970023Z","shell.execute_reply":"2025-03-13T19:24:41.975930Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"print(f\"Training for {num_epochs} epochs started.\")\n\nfor epoch in range(num_epochs):\n    model.train()\n    for batch_idx, (data, targets, _) in enumerate(train_loader):\n\n        data = data.to(device)\n        targets = targets.to(device)\n\n        optimizer.zero_grad()\n        outputs = model(data)\n\n        pos_weight = targets*positive_weigh_factor  # All positive weights are equal to 10\n        criterion = torch.nn.BCEWithLogitsLoss(pos_weight=pos_weight)\n        loss = criterion(outputs, targets)\n\n        loss.backward()\n        optimizer.step()\n\n        if batch_idx % 278 == 0:\n            print(f\"Epoch {epoch+1}/{num_epochs}, Batch {batch_idx}/{len(train_loader)}, Loss: {loss.item()}\")\n\n    scheduler.step()\n    print(\"Scheduler:\",scheduler.state_dict())\n\n# Save the trained model\nmodel.eval()\ntorch.save(model.state_dict(), \"resnet18-with-bioclimatic-cubes.pth\")","metadata":{"ExecuteTime":{"start_time":"2024-04-30T21:25:34.536634Z"},"collapsed":false,"is_executing":true,"jupyter":{"outputs_hidden":false},"tags":[],"execution":{"iopub.status.busy":"2025-03-13T19:24:41.977401Z","iopub.execute_input":"2025-03-13T19:24:41.977705Z","iopub.status.idle":"2025-03-13T19:42:19.676866Z","shell.execute_reply.started":"2025-03-13T19:24:41.977678Z","shell.execute_reply":"2025-03-13T19:42:19.676006Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Test Loop\n\nAgain, nothing special, just a standard inference.","metadata":{}},{"cell_type":"code","source":"with torch.no_grad():\n    all_predictions = []\n    surveys = []\n    top_k_indices = None\n    for data, surveyID in tqdm.tqdm(test_loader, total=len(test_loader)):\n\n        data = data.to(device)\n        \n        outputs = model(data)\n        predictions = torch.sigmoid(outputs).cpu().numpy()\n\n        # Sellect top-25 values as predictions\n        top_25 = np.argsort(-predictions, axis=1)[:, :25] \n        if top_k_indices is None:\n            top_k_indices = top_25\n        else:\n            top_k_indices = np.concatenate((top_k_indices, top_25), axis=0)\n\n        surveys.extend(surveyID.cpu().numpy())","metadata":{"collapsed":false,"jupyter":{"outputs_hidden":false},"tags":[],"execution":{"iopub.status.busy":"2025-03-13T20:10:24.614911Z","iopub.execute_input":"2025-03-13T20:10:24.615257Z","iopub.status.idle":"2025-03-13T20:10:46.459881Z","shell.execute_reply.started":"2025-03-13T20:10:24.615233Z","shell.execute_reply":"2025-03-13T20:10:46.458822Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Save prediction file! 🎉🥳🙌🤗","metadata":{}},{"cell_type":"code","source":"data_concatenated = [' '.join(map(str, row)) for row in top_k_indices]\n\npd.DataFrame(\n    {'surveyId': surveys,\n     'predictions': data_concatenated,\n    }).to_csv(\"submission.csv\", index = False)","metadata":{"tags":[],"execution":{"iopub.status.busy":"2025-03-13T20:10:46.461461Z","iopub.execute_input":"2025-03-13T20:10:46.461807Z","iopub.status.idle":"2025-03-13T20:10:46.632179Z","shell.execute_reply.started":"2025-03-13T20:10:46.461770Z","shell.execute_reply":"2025-03-13T20:10:46.631302Z"},"trusted":true},"outputs":[],"execution_count":null}]}