{"metadata":{"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"},"language_info":{"name":"python","version":"3.10.12","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"nvidiaTeslaT4","dataSources":[{"sourceId":91196,"databundleVersionId":11377716,"sourceType":"competition"}],"dockerImageVersionId":30919,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":true},"papermill":{"default_parameters":{},"duration":1239.673699,"end_time":"2024-05-05T21:21:35.811865","environment_variables":{},"exception":null,"input_path":"__notebook__.ipynb","output_path":"__notebook__.ipynb","parameters":{},"start_time":"2024-05-05T21:00:56.138166","version":"2.5.0"}},"nbformat_minor":4,"nbformat":4,"cells":[{"id":"e49f73eb","cell_type":"markdown","source":"## Simple baseline with Landsat Cubes — ResNet18 + Binary Cross Entropy [0.26424]\n\nTo demonstrate the potential of other data such as Landsat cubes, we provide a straightforward baseline that is baseline on a modified ResNet18 and Binary Cross Entropy but still ranks highly on the leaderboard. The model itself should learn the relationship between the location [R, G, B, NIR, SWIR1, and SWIR2] value at a given location and its species composition.\n\nConsidering the significant extent for enhancing performance of this baseline, we encourage you to experiment with various techniques, architectures, losses, etc.\n\n#### **Have Fun!**","metadata":{"papermill":{"duration":0.006251,"end_time":"2024-05-05T21:00:58.823056","exception":false,"start_time":"2024-05-05T21:00:58.816805","status":"completed"},"tags":[]}},{"id":"88d8d70b","cell_type":"code","source":"import os\nimport torch\nimport tqdm\nimport numpy as np\nimport pandas as pd\nimport torchvision.models as models\nimport torchvision.transforms as transforms\nimport torch.nn as nn\nfrom torch.utils.data import Dataset, DataLoader\nfrom torch.optim.lr_scheduler import CosineAnnealingLR\nfrom sklearn.metrics import precision_recall_fscore_support","metadata":{"ExecuteTime":{"end_time":"2024-04-30T21:25:07.29831Z","start_time":"2024-04-30T21:25:05.354584Z"},"execution":{"iopub.status.busy":"2025-03-13T20:16:59.136493Z","iopub.execute_input":"2025-03-13T20:16:59.136787Z","iopub.status.idle":"2025-03-13T20:17:06.771058Z","shell.execute_reply.started":"2025-03-13T20:16:59.136757Z","shell.execute_reply":"2025-03-13T20:17:06.770364Z"},"papermill":{"duration":7.910804,"end_time":"2024-05-05T21:01:06.739690","exception":false,"start_time":"2024-05-05T21:00:58.828886","status":"completed"},"tags":[],"trusted":true},"outputs":[],"execution_count":null},{"id":"ad54ede5","cell_type":"markdown","source":"## Data description\n\nSatellite time series data includes over 20 years of Landsat satellite imagery extracted from [Ecodatacube](https://stac.ecodatacube.eu/).\nThe data was acquired through the Landsat satellite program and pre-processed by Ecodatacube to produce raster files scaled to the entire European continent and projected into a unique CRS.\n\nSince the original rasters require a high amount of disk space, we extracted the data points from each spectral band corresponding to all PA and PO locations (i.e., GPS coordinates) and aggregated them in (i) CSV files and (ii) data cubes as tensor objects. Each data point corresponds to the mean value of Landsat's observations at the given location for three months before the given time; e.g., the value of a time series element under column 2012_4 will represent the mean value for that element from October 2012 to December 2012.\n\nIn this notebook, we will work with just the cubes. The cubes are structured as follows.\n**Shape**: `(n_bands, n_quarters, n_years)` where:\n- `n_bands` = 6 comprising [`red`, `green`, `blue`, `nir`, `swir1`, `swir2`]\n- `n_quarters` = 4 \n    - *Quarter 1*: December 2 of previous year until March 20 of current year (winter season proxy),\n    - *Quarter 2*: March 21 until June 24 of current year (spring season proxy),\n    - *Quarter 3*: June 25 until September 12 of current year (summer season proxy),\n    - *Quarter 4*: September 13 until December 1 of current year (fall season proxy).\n- `n_years` = 21 (ranging from 2000 to 2020)\n\nThe datacubes can simply be loaded as tensors using PyTorch with the following command :\n\n```python\nimport torch\ntorch.load('path_to_file.pt')\n```\n\n**References:**\n- *Traceability (lineage): This dataset is a seasonally aggregated and gapfilled version of the Landsat GLAD analysis-ready data product presented by Potapov et al., 2020 ( https://doi.org/10.3390/rs12030426 ).*\n- *Scientific methodology: The Landsat GLAD ARD dataset was aggregated and harmonized using the eumap python package (available at https://eumap.readthedocs.io/en/latest/ ). The full process of gapfilling and harmonization is described in detail in Witjes et al., 2022 (in review, preprint available at https://doi.org/10.21203/rs.3.rs-561383/v3 ).*\n- *Ecodatacube.eu: Analysis-ready open environmental data cube for Europe (https://doi.org/10.21203/rs.3.rs-2277090/v3).*","metadata":{"execution":{"iopub.execute_input":"2024-05-01T13:30:07.054038Z","iopub.status.busy":"2024-05-01T13:30:07.053659Z","iopub.status.idle":"2024-05-01T13:30:07.058148Z","shell.execute_reply":"2024-05-01T13:30:07.057269Z","shell.execute_reply.started":"2024-05-01T13:30:07.054008Z"},"papermill":{"duration":0.005438,"end_time":"2024-05-05T21:01:06.751219","exception":false,"start_time":"2024-05-05T21:01:06.745781","status":"completed"},"tags":[]}},{"id":"f758ba3a","cell_type":"markdown","source":"## Prepare custom dataset loader\n\nWe have to sloightly update the Dataset to provide the relevant data in the appropriate format.","metadata":{"papermill":{"duration":0.005422,"end_time":"2024-05-05T21:01:06.762205","exception":false,"start_time":"2024-05-05T21:01:06.756783","status":"completed"},"tags":[]}},{"id":"6ba67bb2","cell_type":"code","source":"class TrainDataset(Dataset):\n    def __init__(self, data_dir, metadata, subset, transform=None):\n        self.subset = subset\n        self.transform = transform\n        self.data_dir = data_dir\n        self.metadata = metadata\n        self.metadata = self.metadata.dropna(subset=\"speciesId\").reset_index(drop=True)\n        self.metadata['speciesId'] = self.metadata['speciesId'].astype(int)\n        self.label_dict = self.metadata.groupby('surveyId')['speciesId'].apply(list).to_dict()\n        \n        self.metadata = self.metadata.drop_duplicates(subset=\"surveyId\").reset_index(drop=True)\n\n\n    def __len__(self):\n        return len(self.metadata)\n\n    def __getitem__(self, idx):\n        \n        survey_id = self.metadata.surveyId[idx]\n        sample = torch.nan_to_num(torch.load(os.path.join(self.data_dir, f\"GLC25-PA-{self.subset}-landsat-time-series_{survey_id}_cube.pt\"), weights_only=True))\n\n        species_ids = self.label_dict.get(survey_id, [])  # Get list of species IDs for the survey ID\n        label = torch.zeros(num_classes)  # Initialize label tensor\n        for species_id in species_ids:\n            #label_id = self.species_mapping[species_id]  # Get consecutive integer label\n            label_id = species_id\n            label[label_id] = 1  # Set the corresponding class index to 1 for each species\n\n        # Ensure the sample is in the correct format for the transform\n        if isinstance(sample, torch.Tensor):\n            sample = sample.permute(1, 2, 0)  # Change tensor shape from (C, H, W) to (H, W, C)\n            sample = sample.numpy()  # Convert tensor to numpy array\n            #print(sample.shape)\n\n        if self.transform:\n            sample = self.transform(sample)\n\n        return sample, label, survey_id\n    \nclass TestDataset(TrainDataset):\n    def __init__(self, data_dir, metadata, subset, transform=None):\n        self.subset = subset\n        self.transform = transform\n        self.data_dir = data_dir\n        self.metadata = metadata\n        \n    def __getitem__(self, idx):\n        \n        survey_id = self.metadata.surveyId[idx]\n        sample = torch.nan_to_num(torch.load(os.path.join(self.data_dir, f\"GLC25-PA-{self.subset}-landsat_time_series_{survey_id}_cube.pt\"), weights_only=True))\n\n        if isinstance(sample, torch.Tensor):\n            sample = sample.permute(1, 2, 0)  # Change tensor shape from (C, H, W) to (H, W, C)\n            sample = sample.numpy()\n\n        if self.transform:\n            sample = self.transform(sample)\n\n        return sample, survey_id","metadata":{"ExecuteTime":{"end_time":"2024-04-30T21:25:32.627928Z","start_time":"2024-04-30T21:25:32.612131Z"},"collapsed":false,"execution":{"iopub.status.busy":"2025-03-13T20:17:06.771946Z","iopub.execute_input":"2025-03-13T20:17:06.772432Z","iopub.status.idle":"2025-03-13T20:17:06.781139Z","shell.execute_reply.started":"2025-03-13T20:17:06.772403Z","shell.execute_reply":"2025-03-13T20:17:06.780321Z"},"jupyter":{"outputs_hidden":false},"papermill":{"duration":0.022488,"end_time":"2024-05-05T21:01:06.790318","exception":false,"start_time":"2024-05-05T21:01:06.767830","status":"completed"},"tags":[],"trusted":true},"outputs":[],"execution_count":null},{"id":"10cc07d0","cell_type":"markdown","source":"### Load metadata and prepare data loaders","metadata":{"papermill":{"duration":0.00538,"end_time":"2024-05-05T21:01:06.801283","exception":false,"start_time":"2024-05-05T21:01:06.795903","status":"completed"},"tags":[]}},{"id":"f2a1f08d","cell_type":"code","source":"# Dataset and DataLoader\nbatch_size = 64\ntransform = transforms.Compose([\n    transforms.ToTensor()\n])\n\n# Load Training metadata\n\n\"/kaggle/input/geolifeclef-2025/SateliteTimeSeries-Landsat/cubes/PA-train/\"\n\ntrain_data_path = \"/kaggle/input/geolifeclef-2025/SateliteTimeSeries-Landsat/cubes/PA-train/\"\ntrain_metadata_path = \"/kaggle/input/geolifeclef-2025/GLC25_PA_metadata_train.csv\"\ntrain_metadata = pd.read_csv(train_metadata_path)\ntrain_dataset = TrainDataset(train_data_path, train_metadata, subset=\"train\", transform=transform)\ntrain_loader = DataLoader(train_dataset, batch_size=batch_size, shuffle=True, num_workers=4)\n\n# Load Test metadata\ntest_data_path = \"/kaggle/input/geolifeclef-2025/SateliteTimeSeries-Landsat/cubes/PA-test/\"\ntest_metadata_path = \"/kaggle/input/geolifeclef-2025/GLC25_PA_metadata_test.csv\"\ntest_metadata = pd.read_csv(test_metadata_path)\ntest_dataset = TestDataset(test_data_path, test_metadata, subset=\"test\", transform=transform)\ntest_loader = DataLoader(test_dataset, batch_size=batch_size, shuffle=False, num_workers=4)","metadata":{"ExecuteTime":{"end_time":"2024-04-30T21:25:34.532017Z","start_time":"2024-04-30T21:25:32.615562Z"},"collapsed":false,"execution":{"iopub.status.busy":"2025-03-13T20:17:06.781975Z","iopub.execute_input":"2025-03-13T20:17:06.782265Z","iopub.status.idle":"2025-03-13T20:17:11.154722Z","shell.execute_reply.started":"2025-03-13T20:17:06.782238Z","shell.execute_reply":"2025-03-13T20:17:11.154032Z"},"jupyter":{"outputs_hidden":false},"papermill":{"duration":5.726171,"end_time":"2024-05-05T21:01:12.535279","exception":false,"start_time":"2024-05-05T21:01:06.809108","status":"completed"},"tags":[],"trusted":true},"outputs":[],"execution_count":null},{"id":"17dcec3b","cell_type":"markdown","source":"## Define and initialize the ModifiedResNet18 model\n\nTo utilize the landsat cubes, which have a shape of [6,4,21] (BANDs, QUARTERs, and YEARs), some minor adjustments must be made to the vanilla ResNet-18. It's important to note that this is just one method for ensuring compatibility with the unusual tensor shape, and experimentation is encouraged.","metadata":{"papermill":{"duration":0.005747,"end_time":"2024-05-05T21:01:12.547063","exception":false,"start_time":"2024-05-05T21:01:12.541316","status":"completed"},"tags":[]}},{"id":"53d74624","cell_type":"code","source":"class ModifiedResNet18(nn.Module):\n    def __init__(self, num_classes):\n        super(ModifiedResNet18, self).__init__()\n\n        self.norm_input = nn.LayerNorm([6,4,21])\n        self.resnet18 = models.resnet18(weights=None)\n        # We have to modify the first convolutional layer to accept 4 channels instead of 3\n        self.resnet18.conv1 = nn.Conv2d(6, 64, kernel_size=3, stride=1, padding=1, bias=False)\n        self.resnet18.maxpool = nn.Identity()\n        self.ln = nn.LayerNorm(1000)\n        self.fc1 = nn.Linear(1000, 2056)\n        self.fc2 = nn.Linear(2056, num_classes)\n\n    def forward(self, x):\n        x = self.norm_input(x)\n        x = self.resnet18(x)\n        x = self.ln(x)\n        x = self.fc1(x)\n        x = self.fc2(x)\n        return x","metadata":{"ExecuteTime":{"end_time":"2024-04-30T21:25:31.014067Z","start_time":"2024-04-30T21:25:31.01006Z"},"collapsed":false,"execution":{"iopub.status.busy":"2025-03-13T20:17:11.155602Z","iopub.execute_input":"2025-03-13T20:17:11.155839Z","iopub.status.idle":"2025-03-13T20:17:11.163024Z","shell.execute_reply.started":"2025-03-13T20:17:11.155819Z","shell.execute_reply":"2025-03-13T20:17:11.162365Z"},"jupyter":{"outputs_hidden":false},"papermill":{"duration":0.015915,"end_time":"2024-05-05T21:01:12.568601","exception":false,"start_time":"2024-05-05T21:01:12.552686","status":"completed"},"tags":[],"trusted":true},"outputs":[],"execution_count":null},{"id":"7e98ba53","cell_type":"code","source":"def set_seed(seed):\n    # Set seed for Python's built-in random number generator\n    torch.manual_seed(seed)\n    # Set seed for numpy\n    np.random.seed(seed)\n    # Set seed for CUDA if available\n    if torch.cuda.is_available():\n        torch.cuda.manual_seed_all(seed)\n        # Set cuDNN's random number generator seed for deterministic behavior\n        torch.backends.cudnn.deterministic = True\n        torch.backends.cudnn.benchmark = False\n\nset_seed(69)","metadata":{"execution":{"iopub.status.busy":"2025-03-13T20:17:11.164643Z","iopub.execute_input":"2025-03-13T20:17:11.164887Z","iopub.status.idle":"2025-03-13T20:17:11.268616Z","shell.execute_reply.started":"2025-03-13T20:17:11.164866Z","shell.execute_reply":"2025-03-13T20:17:11.267695Z"},"papermill":{"duration":0.058735,"end_time":"2024-05-05T21:01:12.632872","exception":false,"start_time":"2024-05-05T21:01:12.574137","status":"completed"},"tags":[],"trusted":true},"outputs":[],"execution_count":null},{"id":"828330c8","cell_type":"code","source":"# Check if cuda is available\ndevice = torch.device(\"cpu\")\n\nif torch.cuda.is_available():\n    device = torch.device(\"cuda\")\n    print(\"DEVICE = CUDA\")\n\nnum_classes = 11255 # Number of all unique classes within the PO and PA data.\nmodel = ModifiedResNet18(num_classes).to(device)","metadata":{"ExecuteTime":{"end_time":"2024-04-30T21:25:31.611823Z","start_time":"2024-04-30T21:25:31.607373Z"},"collapsed":false,"execution":{"iopub.status.busy":"2025-03-13T20:17:11.269620Z","iopub.execute_input":"2025-03-13T20:17:11.269858Z","iopub.status.idle":"2025-03-13T20:17:11.920107Z","shell.execute_reply.started":"2025-03-13T20:17:11.269838Z","shell.execute_reply":"2025-03-13T20:17:11.919307Z"},"jupyter":{"outputs_hidden":false},"papermill":{"duration":0.676292,"end_time":"2024-05-05T21:01:13.314798","exception":false,"start_time":"2024-05-05T21:01:12.638506","status":"completed"},"tags":[],"trusted":true},"outputs":[],"execution_count":null},{"id":"648576c4","cell_type":"markdown","source":"## Training Loop\n\nNothing special, just a standard Pytorch training loop.","metadata":{"papermill":{"duration":0.005509,"end_time":"2024-05-05T21:01:13.326480","exception":false,"start_time":"2024-05-05T21:01:13.320971","status":"completed"},"tags":[]}},{"id":"e35d34a4","cell_type":"code","source":"# Hyperparameters\nlearning_rate = 0.0002\nnum_epochs = 20\npositive_weigh_factor = 1.0\n\noptimizer = torch.optim.AdamW(model.parameters(), lr=learning_rate)\nscheduler = CosineAnnealingLR(optimizer, T_max=25, verbose=True)","metadata":{"ExecuteTime":{"end_time":"2024-04-30T21:25:32.181927Z","start_time":"2024-04-30T21:25:32.177073Z"},"collapsed":false,"execution":{"iopub.status.busy":"2025-03-13T20:17:11.920875Z","iopub.execute_input":"2025-03-13T20:17:11.921189Z","iopub.status.idle":"2025-03-13T20:17:11.927762Z","shell.execute_reply.started":"2025-03-13T20:17:11.921159Z","shell.execute_reply":"2025-03-13T20:17:11.926976Z"},"jupyter":{"outputs_hidden":false},"papermill":{"duration":0.015096,"end_time":"2024-05-05T21:01:13.347192","exception":false,"start_time":"2024-05-05T21:01:13.332096","status":"completed"},"tags":[],"trusted":true},"outputs":[],"execution_count":null},{"id":"e9d1df4d","cell_type":"code","source":"print(f\"Training for {num_epochs} epochs started.\")\n\nfor epoch in range(num_epochs):\n    model.train()\n    for batch_idx, (data, targets, _) in enumerate(train_loader):\n\n        data = data.to(device)\n        targets = targets.to(device)\n\n        optimizer.zero_grad()\n        outputs = model(data)\n\n        pos_weight = targets*positive_weigh_factor  # All positive weights are equal to 10\n        criterion = torch.nn.BCEWithLogitsLoss(pos_weight=pos_weight)\n        loss = criterion(outputs, targets)\n\n        loss.backward()\n        optimizer.step()\n\n        if batch_idx % 278 == 0:\n            print(f\"Epoch {epoch+1}/{num_epochs}, Batch {batch_idx}/{len(train_loader)}, Loss: {loss.item()}\")\n\n    scheduler.step()\n    print(\"Scheduler:\",scheduler.state_dict())\n\n# Save the trained model\nmodel.eval()\ntorch.save(model.state_dict(), \"resnet18-with-landsat-cubes.pth\")","metadata":{"ExecuteTime":{"start_time":"2024-04-30T21:25:34.536634Z"},"collapsed":false,"execution":{"iopub.status.busy":"2025-03-13T20:17:11.928492Z","iopub.execute_input":"2025-03-13T20:17:11.928704Z","iopub.status.idle":"2025-03-13T20:40:32.881003Z","shell.execute_reply.started":"2025-03-13T20:17:11.928686Z","shell.execute_reply":"2025-03-13T20:40:32.880021Z"},"is_executing":true,"jupyter":{"outputs_hidden":false},"papermill":{"duration":1209.843817,"end_time":"2024-05-05T21:21:23.196800","exception":false,"start_time":"2024-05-05T21:01:13.352983","status":"completed"},"tags":[],"trusted":true},"outputs":[],"execution_count":null},{"id":"ce07f945","cell_type":"markdown","source":"## Test Loop\n\nAgain, nothing special, just a standard inference.","metadata":{"papermill":{"duration":0.014707,"end_time":"2024-05-05T21:21:23.226855","exception":false,"start_time":"2024-05-05T21:21:23.212148","status":"completed"},"tags":[]}},{"id":"5fd35b61","cell_type":"code","source":"with torch.no_grad():\n    all_predictions = []\n    surveys = []\n    top_k_indices = None\n    for data, surveyID in tqdm.tqdm(test_loader, total=len(test_loader)):\n\n        data = data.to(device)\n        \n        outputs = model(data)\n        predictions = torch.sigmoid(outputs).cpu().numpy()\n\n        # Sellect top-25 values as predictions\n        top_25 = np.argsort(-predictions, axis=1)[:, :25] \n        if top_k_indices is None:\n            top_k_indices = top_25\n        else:\n            top_k_indices = np.concatenate((top_k_indices, top_25), axis=0)\n\n        surveys.extend(surveyID.cpu().numpy())","metadata":{"collapsed":false,"execution":{"iopub.status.busy":"2025-03-13T20:40:32.882124Z","iopub.execute_input":"2025-03-13T20:40:32.882443Z","iopub.status.idle":"2025-03-13T20:41:05.981664Z","shell.execute_reply.started":"2025-03-13T20:40:32.882408Z","shell.execute_reply":"2025-03-13T20:41:05.980764Z"},"jupyter":{"outputs_hidden":false},"papermill":{"duration":9.799329,"end_time":"2024-05-05T21:21:33.041149","exception":false,"start_time":"2024-05-05T21:21:23.241820","status":"completed"},"tags":[],"trusted":true},"outputs":[],"execution_count":null},{"id":"5c9d476e","cell_type":"markdown","source":"## Save prediction file! 🎉🥳🙌🤗","metadata":{"papermill":{"duration":0.01651,"end_time":"2024-05-05T21:21:33.076908","exception":false,"start_time":"2024-05-05T21:21:33.060398","status":"completed"},"tags":[]}},{"id":"2e027600","cell_type":"code","source":"data_concatenated = [' '.join(map(str, row)) for row in top_k_indices]\n\npd.DataFrame(\n    {'surveyId': surveys,\n     'predictions': data_concatenated,\n    }).to_csv(\"submission.csv\", index = False)","metadata":{"execution":{"iopub.status.busy":"2025-03-13T20:41:05.982564Z","iopub.execute_input":"2025-03-13T20:41:05.982840Z","iopub.status.idle":"2025-03-13T20:41:06.153832Z","shell.execute_reply.started":"2025-03-13T20:41:05.982816Z","shell.execute_reply":"2025-03-13T20:41:06.153240Z"},"papermill":{"duration":0.124413,"end_time":"2024-05-05T21:21:33.218031","exception":false,"start_time":"2024-05-05T21:21:33.093618","status":"completed"},"tags":[],"trusted":true},"outputs":[],"execution_count":null}]}