{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.10.12","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":91196,"databundleVersionId":11601067,"sourceType":"competition"}],"dockerImageVersionId":30918,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"<h1 style=\"text-align: center;\"><b>EDA Simplified:<span style=\"color:#d5601d;\"> GeoLifeCLEF 2025</span></b></h1>\n<center><img src=\"https://www.kaggle.com/competitions/91196/images/header\" width=700></center>","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19"}},{"cell_type":"markdown","source":"<h2 style=\"background-color: #d5601d; color: white; padding-right: 100vw; background-size:cover; text-align: center; padding: 10px; border-radius: 15px;\">Introduction</h2>\n\nAfter approximately one and a half year of inactivity, EDA Simplified is back on track! Since 2022 in the Great Barrier Reef COTS Competition, we went through competition by competition exploring each data and what they represent, while detailing on what the code does in each cell of our EDA notebooks. And since March 2025 is around the corner, there goes for the action of the FGVC workshops. For our notebook, we're going to go through the data covered in the 2025 GeoLifeCLEF competition.\n\n<h3 style=\"background-color: #d5601d; color: white; padding-right: 100vw; background-size:cover; text-align: center; padding: 7px; border-radius: 15px;\">So, What is GeoLifeCLEF 2025?</h3>\n\nGeoLifeCLEF 2025 is a fine-grained visual competition that focused in predicting the plant presence at given locations throughout Europe. The purpose of this competition is to manage biodiversity and motivate conservation of plants while accelerating the annotation and validation of species observations to produce large, high-quality datasets. Without further ado, let's get to it!","metadata":{}},{"cell_type":"markdown","source":"<h2 style=\"background-color: #d5601d; color: white; padding-right: 100vw; background-size:cover; text-align: center; padding: 10px; border-radius: 15px;\">Loading the Metadata / General Data Overview</h2>\nHow do we start our EDA on GeoLifeCLEF? As always, we import three modules: Pandas for data science, Numpy for possible array usage, and Matplotlib + Seaborn for visualizing what in the GeoLifeCLEF datasets.","metadata":{}},{"cell_type":"code","source":"import pandas as pd\nimport numpy as np\nimport matplotlib.pyplot as plt\nimport seaborn as sns","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-05T04:09:24.132690Z","iopub.execute_input":"2025-04-05T04:09:24.132964Z","iopub.status.idle":"2025-04-05T04:09:25.009667Z","shell.execute_reply.started":"2025-04-05T04:09:24.132942Z","shell.execute_reply":"2025-04-05T04:09:25.009015Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"After we load our necessary modules, let's load the given data from the GeoLifeCLEF dataset! Since there are many, many \"pt\" files and csv files stored in some folders, we only focus on the general metadata by reading the \"PO\" and \"PA\" CSV files and storing it to a dataframe. If you wonder what are those two abbreviations mean, \"PO\" represent Presence only data while \"PA\" represent presence absent data.","metadata":{}},{"cell_type":"code","source":"po_df = pd.read_csv(\"/kaggle/input/geolifeclef-2025/GLC25_P0_metadata_train.csv\")\npa_df = pd.read_csv(\"/kaggle/input/geolifeclef-2025/GLC25_PA_metadata_test.csv\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-05T04:09:25.010517Z","iopub.execute_input":"2025-04-05T04:09:25.010795Z","iopub.status.idle":"2025-04-05T04:09:33.846915Z","shell.execute_reply.started":"2025-04-05T04:09:25.010776Z","shell.execute_reply":"2025-04-05T04:09:33.846130Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"Now that we have two dataframes (PO and PA), let's begin our overview of the metadata! First, let's see how many entities are there in each of the two dataframes.","metadata":{}},{"cell_type":"code","source":"print(\"PO Metadata: \", len(po_df))\nprint(\"PA Metadata: \", len(pa_df))","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-05T04:09:33.848269Z","iopub.execute_input":"2025-04-05T04:09:33.848536Z","iopub.status.idle":"2025-04-05T04:09:33.852960Z","shell.execute_reply.started":"2025-04-05T04:09:33.848515Z","shell.execute_reply":"2025-04-05T04:09:33.852306Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"As we can see, there are 5079787 data entities in the PO dataframe while there are 14784 data entities in the PA dataframe. This highlights that there are many data stored for keeping track of biodiversity in plants.","metadata":{}},{"cell_type":"markdown","source":"Next, let's see how many of the missing values are present in the PO and PA metadata.","metadata":{}},{"cell_type":"code","source":"print(\"PO Missing Values: \", po_df.isna().sum().sum())\nprint(\"PA Missing Values: \", pa_df.isna().sum().sum())","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-05T04:09:33.854001Z","iopub.execute_input":"2025-04-05T04:09:33.854221Z","iopub.status.idle":"2025-04-05T04:09:34.416604Z","shell.execute_reply.started":"2025-04-05T04:09:33.854198Z","shell.execute_reply":"2025-04-05T04:09:34.415862Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"There are 2303 missing values in the PO dataframe and 1484 in the PA dataframe. If we add them together, there are 3787 overall missing values present with a percentage between the ratio of total missing values to all PO and PA values to approximately 7%. In other words, some of the missing data in those two dataframes shows that there were uncertainties or data errors in both of them.","metadata":{}},{"cell_type":"markdown","source":"<h2 style=\"background-color: #d5601d; color: white; padding-right: 100vw; background-size:cover; text-align: center; padding: 10px; border-radius: 15px;\">PO Metadata</h2>\n\nAfter we made our brief overview of the PO and PA dataframes, it's time to start analyzing the PO metadata! To get started, let's display the first five rows of the dataframe containing the PO data.","metadata":{}},{"cell_type":"code","source":"po_df.head()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-05T04:09:34.417230Z","iopub.execute_input":"2025-04-05T04:09:34.417398Z","iopub.status.idle":"2025-04-05T04:09:34.441628Z","shell.execute_reply.started":"2025-04-05T04:09:34.417382Z","shell.execute_reply":"2025-04-05T04:09:34.440855Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"From the PO-only metadata, we counted 12 columns in this dataframe, each containing the data from numerous datasets gathered by the GBIF or Global Biodiversity Information Facility. Regarding the columns we saw, let's take a brief look at it!\n* **publisher**: Who published the metadata of each observations\n* **year**, **month**, & **day**: When was each observation published\n* **lat** & **lon**: The coordinates of each observation data, aka where was it published at\n* **geoUncertaintyInM**: How uncertain each observation is\n* **taxonRank**: What ranking each observation is\n* **date**: Again, when was each observation published\n* **dayOfYear**: What part of day in each year an observation is published\n* **speciesId** & **surveyId**: The identifier for each species or survey in each observation.","metadata":{}},{"cell_type":"markdown","source":"With that brief overview of the PO metadata covered up, let's determine who published most of the observations! To do that, we create a pie chart with the top 100 rows of the PO metadata.","metadata":{}},{"cell_type":"code","source":"publishers = po_df[\"publisher\"].value_counts(normalize=True)\n\nplt.rcParams[\"figure.figsize\"] = [10.00, 10.00]\n\nfig, labels = plt.pie(\n    publishers.values\n)\n\nlabels = ['{0} - {1:1.2f} %'.format(i, j) for i, j in zip(publishers.index, 100.*publishers.values/publishers.values.sum())]\nplt.legend(fig, labels, loc='lower left', bbox_to_anchor=(-0.05, 0.05), fontsize=8)\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-05T04:09:34.442288Z","iopub.execute_input":"2025-04-05T04:09:34.442515Z","iopub.status.idle":"2025-04-05T04:09:34.979841Z","shell.execute_reply.started":"2025-04-05T04:09:34.442497Z","shell.execute_reply":"2025-04-05T04:09:34.979095Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"After displaying this pie-chart, we noticed that there are significant organizations accountable for approximately 45.3% of observations published by \"Pl@ntNet\", while there was roughly 13.6% of the metadata published by the Danish Environmental Protection Agency, 12.3% by iNaturalist.org, 11.8% by the NBIC, and 4.75% by Observation.org. In other words, the 45.3% of the data regarding publishers for the PO dataset was mostly contributed by the \"Pl@ntNet\" organization, particularly because of their large database of plant species in Europe.","metadata":{}},{"cell_type":"markdown","source":"Next, let's determine how many data observations were published in each year, month, and the day from each year with a graph of three histograms.","metadata":{}},{"cell_type":"code","source":"# Note: Data in po_df is truncated by 2000 entities, to reduce lag.\nf, axes = plt.subplots(1, 3, figsize=(30, 15))\n\nsns.histplot(data=po_df.head(2000), x=\"year\", ax=axes[0])\nsns.histplot(data=po_df.head(2000), x=\"month\", ax=axes[1])\nsns.histplot(data=po_df.head(2000), x=\"dayOfYear\", ax=axes[2])\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-05T04:09:34.980456Z","iopub.execute_input":"2025-04-05T04:09:34.980643Z","iopub.status.idle":"2025-04-05T04:09:35.564119Z","shell.execute_reply.started":"2025-04-05T04:09:34.980624Z","shell.execute_reply":"2025-04-05T04:09:35.563398Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"From the first histogram, we noted that the bars was skewed right to the year 2021, as there are approximately 1.7 million entities in that side. Meanwhile in the middle histogram, we saw that it was displaying a uniformal distribution, with around 932K of data recorded in June and 929K in July. Lastly, the last histogram also displayed a nearly uniformal distribution with high amounts of data recorded around days 139 and 241. Regarding the uniformal distributions seen in the month and day of each year, as well as the right-skewed distribution in the year, it hinted us that most observations were recorded in the summer days of 2021.","metadata":{}},{"cell_type":"markdown","source":"Moving on to finding how many species and subspecies were recorded in the PO dataset, we create another pie graph visualizing how many of them were stored in the PO dataframe's `taxonRank` column.","metadata":{}},{"cell_type":"code","source":"taxons = po_df[\"taxonRank\"].value_counts(normalize=True)\n\nfig = plt.pie(\n    taxons.values, \n    labels=taxons.index\n)\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-05T04:09:35.565613Z","iopub.execute_input":"2025-04-05T04:09:35.565821Z","iopub.status.idle":"2025-04-05T04:09:35.895947Z","shell.execute_reply.started":"2025-04-05T04:09:35.565804Z","shell.execute_reply":"2025-04-05T04:09:35.895245Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"As we can see, 98.1% of the taxonomy data were classified as \"SPECIES\", while 1.85% of them were labeled as \"SUBSPECIES\". To summarize that, most of the observations in the PO metadata contained data of plant species to be used for predicting their presence through satellite imagery, environmental rasters, or climate time series.","metadata":{}},{"cell_type":"markdown","source":"Finally, before we get going to plotting coordinates of observation locations with other data fields in the PO metadata, let's visualize the data distribution of the geographic uncertainty in the PO dataframe with a histogram.","metadata":{}},{"cell_type":"code","source":"# Note: Data in po_df is truncated by 2000 entities, to reduce lag.\nsns.histplot(data=po_df.head(2000), x=\"geoUncertaintyInM\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-05T04:09:35.896659Z","iopub.execute_input":"2025-04-05T04:09:35.896830Z","iopub.status.idle":"2025-04-05T04:09:36.079970Z","shell.execute_reply.started":"2025-04-05T04:09:35.896814Z","shell.execute_reply":"2025-04-05T04:09:36.079238Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"After taking a long time for the histogram to be displayed, we observed most of the data were skewed to the left, while there was sparse data on the right of the diagram. Additionally, the data range in this histogram that has the highest counts of data is between 2.95 to 3.04 with 635993 data entities. Moreover, the roughly skewed-left data in the histogram shows us that there were a few uncertainties shown in the `geoUncertaintyInM` section from the PO dataframe.\n\nNow that we have our general visualization of the data from the PO dataframe done, let's move on to visualizing where each observation is recorded with choropleth maps and other corresponding data in the PO metadata!","metadata":{}},{"cell_type":"markdown","source":"<h3 style=\"background-color: #d5601d; color: white; padding-right: 100vw; background-size:cover; text-align: center; padding: 7px; border-radius: 15px;\">Geographical Analysis of the PO Metadata</h3>\n\nAfter visualizing through the PO metadata generally, we understood that there are coordinated stored in this metadata. Without any doubts ahead of us, let's get into visualizing the PO Metadata geographically... within Europe!","metadata":{}},{"cell_type":"markdown","source":"Before we start visualizing the PO Metadata with geographical coordinates, we have to install additional geographic visualization modules, as well as processing the geometry data for the PO dataframe.","metadata":{}},{"cell_type":"code","source":"# Source: mpwolke from https://www.kaggle.com/code/mpwolke/geolifeclef25-maps#Install-Folium-Matplotlib-Mapclassify\n!pip install folium matplotlib mapclassify\nfrom shapely.geometry import Polygon, LineString, Point\nimport geopandas as gpd\nimport tqdm\n\npo_geo_df = po_df.drop_duplicates('surveyId').sample(n=2000, random_state=42)\npo_geo_df.index = range(len(po_geo_df))\n\n# make Point vector \npoint_list = []\nfor i in tqdm.tqdm(range(len(po_geo_df))):\n    x,y = po_df.loc[i, ['lon', 'lat']]\n    poind_i = Point(x,y)\n    point_list.append(poind_i)\n\npo_geo_df.loc[:,'geometry'] = point_list","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-05T04:10:57.420780Z","iopub.execute_input":"2025-04-05T04:10:57.421250Z","iopub.status.idle":"2025-04-05T04:11:01.943911Z","shell.execute_reply.started":"2025-04-05T04:10:57.421225Z","shell.execute_reply":"2025-04-05T04:11:01.943207Z"},"_kg_hide-output":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"After loading the data, let's then convert the geographic-processed PO metadata into a geo-dataframe, then visualize it in Folium.","metadata":{}},{"cell_type":"code","source":"# Source: mpwolke from https://www.kaggle.com/code/mpwolke/geolifeclef25-maps#Install-Folium-Matplotlib-Mapclassify\nvis_geo_po = gpd.GeoDataFrame(po_geo_df, geometry = 'geometry')\nvis_geo_po.crs = ('EPSG:4326')\nvis_geo_po.drop_duplicates(['lon', 'lat']).explore(color = 'green')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-05T04:15:11.109528Z","iopub.execute_input":"2025-04-05T04:15:11.109885Z","iopub.status.idle":"2025-04-05T04:15:15.327553Z","shell.execute_reply.started":"2025-04-05T04:15:11.109860Z","shell.execute_reply":"2025-04-05T04:15:15.326513Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"Whoa, look at that! After running this code segment above, we glimpse a clear view of Europe with the green plots dotted from the Western portion of Europe to the far east side. To add this, the dense cluster of green points suggested that most of the PO metadata were recorded in common areas in France, the UK, Germany, Denmark, Spain, Portugal, Switzerland, Austria, Czech Republic, Slovakia, and Italy.","metadata":{}},{"cell_type":"markdown","source":"Since we vaguely overlooked at the cluster of green plots in the previous plot, let's now create a heatmap to visualize how dense the plots are in the PO metadata.","metadata":{}},{"cell_type":"code","source":"import folium\nfrom folium import plugins\n\nmap = folium.Map(location = [52,9], tiles='Cartodb dark_matter', zoom_start = 3.5)\nheat_data = [[point.xy[1][0], point.xy[0][0]] for point in po_geo_df.geometry ]\nplugins.HeatMap(heat_data).add_to(map)\nmap","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-05T04:40:33.661926Z","iopub.execute_input":"2025-04-05T04:40:33.662234Z","iopub.status.idle":"2025-04-05T04:40:33.761781Z","shell.execute_reply.started":"2025-04-05T04:40:33.662208Z","shell.execute_reply":"2025-04-05T04:40:33.761006Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"As we can see, we observed the red plots combined together throughout Great Britain, the western European region, and areas south of Norway and Sweden, and Finland. In other words, most of the PO metadata were commonly gathered in Western portions of Europe.","metadata":{}},{"cell_type":"markdown","source":"Now that we finished visualizing the PO metadata geographically, let's now pivot into analyzing the contents from the Presence-absent metadata!","metadata":{}},{"cell_type":"markdown","source":"<h2 style=\"background-color: #d5601d; color: white; padding-right: 100vw; background-size:cover; text-align: center; padding: 10px; border-radius: 15px;\">PA Metadata</h2>\n\n**This EDA notebook is under work in progress, stay tuned! Also, it may update weekly or as soon as possible.**","metadata":{}}]}