{"cells":[{"metadata":{},"cell_type":"markdown","source":"This notebook details the data structure and shows how to load the data."},{"metadata":{"ExecuteTime":{"end_time":"2021-03-08T13:34:34.828894Z","start_time":"2021-03-08T13:34:34.414177Z"},"trusted":true},"cell_type":"code","source":"%pylab inline --no-import-all\n\nfrom pathlib import Path","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"We first need to clone our code:"},{"metadata":{"trusted":true},"cell_type":"code","source":"!rm -rf GLC\n!git clone https://github.com/maximiliense/GLC","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Then, we need to define the path to the data:"},{"metadata":{"ExecuteTime":{"end_time":"2021-03-08T13:34:34.832817Z","start_time":"2021-03-08T13:34:34.830546Z"},"trusted":true},"cell_type":"code","source":"# Change this path to adapt to where you downloaded the data\nDATA_PATH = Path(\"../input/geolifeclef-2021/data\")","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"This folder is the path root where the data was downloaded and extracted:"},{"metadata":{"ExecuteTime":{"end_time":"2021-03-08T13:34:34.970444Z","start_time":"2021-03-08T13:34:34.834442Z"},"trusted":true},"cell_type":"code","source":"ls -L $DATA_PATH","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"We can now look into these subfolders and the data they contain."},{"metadata":{},"cell_type":"markdown","source":"# Observations\n\nThe `observations` subfolder contains 4 CSV files:"},{"metadata":{"ExecuteTime":{"end_time":"2021-03-08T13:34:35.08966Z","start_time":"2021-03-08T13:34:34.973157Z"},"trusted":true},"cell_type":"code","source":"ls $DATA_PATH/observations","execution_count":null,"outputs":[]},{"metadata":{"ExecuteTime":{"end_time":"2021-03-05T18:10:04.278128Z","start_time":"2021-03-05T18:10:04.273719Z"}},"cell_type":"markdown","source":"Each of line of those files corresponds to a single observation.\n\nIn the files corresponding to the training data, there are 5 columns:\n- `observation_id`: unique identifier of the observation\n- `latitude`: latitude coordinates of this observation\n- `longitude`: longitude coordinates of this observation\n- `species_id`: identifier of the species observed at that location\n- `subset`: proposed train/val split using the same splitting procedure than for train and test (equal to either \"train\" or \"val\")\n\nIn the files corresponding to the test data, there are only 3 columns:\n- `observation_id`: unique identifier of the observation\n- `latitude`: latitude coordinates of this observation\n- `longitude`: longitude coordinates of this observation\n\nThe goal is then to predict the identifier of the species observed at that location."},{"metadata":{},"cell_type":"markdown","source":"Let's load these CSV files using [pandas](https://pandas.pydata.org/):"},{"metadata":{"ExecuteTime":{"end_time":"2021-03-08T13:34:35.355741Z","start_time":"2021-03-08T13:34:35.091939Z"},"trusted":true},"cell_type":"code","source":"import pandas as pd","execution_count":null,"outputs":[]},{"metadata":{"ExecuteTime":{"end_time":"2021-03-08T13:34:36.424859Z","start_time":"2021-03-08T13:34:35.357542Z"},"trusted":true},"cell_type":"code","source":"df_fr = pd.read_csv(DATA_PATH / \"observations\" / \"observations_fr_train.csv\", sep=\";\", index_col=\"observation_id\")\ndf_us = pd.read_csv(DATA_PATH / \"observations\" / \"observations_us_train.csv\", sep=\";\", index_col=\"observation_id\")\n\ndf = pd.concat((df_fr, df_us))\n\nprint(\"Number of observations for training: {}\".format(len(df)))\n\ndf.head()","execution_count":null,"outputs":[]},{"metadata":{"ExecuteTime":{"end_time":"2021-03-08T13:34:36.466208Z","start_time":"2021-03-08T13:34:36.427435Z"},"trusted":true},"cell_type":"code","source":"df_fr_test = pd.read_csv(DATA_PATH / \"observations\" / \"observations_fr_test.csv\", sep=\";\", index_col=\"observation_id\")\ndf_us_test = pd.read_csv(DATA_PATH / \"observations\" / \"observations_us_test.csv\", sep=\";\", index_col=\"observation_id\")\n\ndf_test = pd.concat((df_fr_test, df_us_test))\n\nprint(\"Number of observations for testing: {}\".format(len(df_test)))\n\ndf_test.head()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"The observations are not uniformly sampled in the two countries as shown the following plots.\nThe training observations are shown in blue while the test ones are shown in red."},{"metadata":{"ExecuteTime":{"end_time":"2021-03-08T13:34:37.790173Z","start_time":"2021-03-08T13:34:36.469043Z"},"trusted":true},"cell_type":"code","source":"def plot_observations_distribution(ax, df, df_test=None, **kwargs):\n    default_kwargs = {\n        \"zorder\": 1,\n        \"alpha\": 0.1,\n        \"s\": 0.5\n    }\n    default_kwargs.update(kwargs)\n    kwargs = default_kwargs\n    \n    ax.scatter(df.longitude, df.latitude, color=\"blue\", **kwargs)\n    \n    if df_test is not None:\n        ax.scatter(df_test.longitude, df_test.latitude, color=\"red\", **kwargs)\n\n    ax.autoscale(enable=True, axis=\"both\", tight=True)\n    ax.axis(\"off\")\n\n\nfig = plt.figure(figsize=(9, 5.5))\nax = fig.gca()\nplot_observations_distribution(ax, df_us, df_us_test)\nax.set_title(\"Observations distribution (US)\")\n\nfig = plt.figure(figsize=(5.5, 5.5))\nax = fig.gca()\nplot_observations_distribution(ax, df_fr, df_fr_test)\nax.set_title(\"Observations distribution (France)\")","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"A close-up view on the region around Montpellier, France, shows the train/test splitting procedure.\n\nNote that there is no geographical overlap between training and test sets."},{"metadata":{"ExecuteTime":{"end_time":"2021-03-08T13:34:38.065912Z","start_time":"2021-03-08T13:34:37.792575Z"},"trusted":true},"cell_type":"code","source":"def select_samples_around_point(df, lon_min, lon_max, lat_min, lat_max):\n    ind = (\n        (lon_min <= df.longitude) & (df.longitude <= lon_max)\n        & (lat_min <= df.latitude) & (df.latitude <= lat_max)\n    )\n    return df[ind]\n\n\nfig, ax = plt.subplots(figsize=(9.5, 7))\n\nkwargs = {\n    \"alpha\": 0.2,\n    \"s\": 5,\n}\ndf_zoom = select_samples_around_point(df_fr, 3, 4.5, 43.25, 44.25)\ndf_zoom_test = select_samples_around_point(df_fr_test, 3, 4.5, 43.25, 44.25)\n\nax = fig.gca()\nplot_observations_distribution(ax, df_zoom, df_zoom_test, **kwargs)\nax.set_title(\"Observations distribution around Montpellier, France\")","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# Metadata\n\nIn the `metadata` folder, some additional data is provided.\nThere are 4 files containing:\n1. GBIF species names associated with the species id provided in the observations in `species_details.csv`\n2. The description of the environmental (bioclimatic and pedological) variables in `environmental_variables.csv`\n3. The labels corresponding to the original land cover codes in `landcover_original_labels.csv`\n4. The suggested alignment of land cover codes between France and US in `landcover_suggested_alignment.csv`"},{"metadata":{"ExecuteTime":{"end_time":"2021-03-08T13:34:38.103773Z","start_time":"2021-03-08T13:34:38.06741Z"},"trusted":true},"cell_type":"code","source":"df_species = pd.read_csv(DATA_PATH / \"metadata\" / \"species_details.csv\", sep=\";\")\n\nprint(\"Total number of species: {}\".format(len(df_species)))\n\ndf_species.head()","execution_count":null,"outputs":[]},{"metadata":{"ExecuteTime":{"end_time":"2021-03-08T13:34:38.121414Z","start_time":"2021-03-08T13:34:38.105224Z"},"trusted":true},"cell_type":"code","source":"df_env_vars = pd.read_csv(DATA_PATH / \"metadata\" / \"environmental_variables.csv\", sep=\";\")\ndf_env_vars.head()","execution_count":null,"outputs":[]},{"metadata":{"ExecuteTime":{"end_time":"2021-03-08T13:34:38.135303Z","start_time":"2021-03-08T13:34:38.123369Z"},"trusted":true},"cell_type":"code","source":"df_landcover_labels = pd.read_csv(DATA_PATH / \"metadata\" / \"landcover_original_labels.csv\", sep=\";\")\ndf_landcover_labels.head()","execution_count":null,"outputs":[]},{"metadata":{"ExecuteTime":{"end_time":"2021-03-08T13:34:38.161394Z","start_time":"2021-03-08T13:34:38.137553Z"},"trusted":true},"cell_type":"code","source":"df_suggested_landcover_alignment = pd.read_csv(DATA_PATH / \"metadata\" / \"landcover_suggested_alignment.csv\", sep=\";\")\ndf_suggested_landcover_alignment.head()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# Patches\n\nThe patches consist of images centered at each observation's location capturing three types of information in the 250m x 250m neighboring square:\n1. remote sensing imagery under the form of RGB-IR images\n2. land cover data\n3. altitude data\n\nThey are located in the `patches` subfolder contains 2 subfolders, one for each country:"},{"metadata":{"ExecuteTime":{"end_time":"2021-03-08T13:34:38.170356Z","start_time":"2021-03-08T13:34:38.16398Z"},"trusted":true},"cell_type":"code","source":"PATCHES_PATH = DATA_PATH / \"patches_sample\"","execution_count":null,"outputs":[]},{"metadata":{"ExecuteTime":{"end_time":"2021-03-08T13:34:38.307414Z","start_time":"2021-03-08T13:34:38.172969Z"},"trusted":true},"cell_type":"code","source":"ls $PATCHES_PATH","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"The first digit of the observation id tells the country it belongs to:\n- `1` for France, thus to be found in subfolder `fr`\n- `2` for US, thus to be found in subfolder `us`\n\nFor instance, `10561949` is an observation made in France whereas `22068175` was observed in the US.\n\nInside those folders, there are two levels of hierarchy, corresponding to the last four digits of the observation id:"},{"metadata":{"ExecuteTime":{"end_time":"2021-03-08T13:34:38.454411Z","start_time":"2021-03-08T13:34:38.309623Z"},"trusted":true},"cell_type":"code","source":"ls $PATCHES_PATH/fr","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"and"},{"metadata":{"ExecuteTime":{"end_time":"2021-03-08T13:34:38.589731Z","start_time":"2021-03-08T13:34:38.457021Z"},"trusted":true},"cell_type":"code","source":"ls $PATCHES_PATH/fr/00","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"To find the files corresponding to an observation:\n1. the first subfolder corresponds to the last two digits,\n2. the second subfolder corresponds to the two digits right before them.\n\nFor instance, the patches corresponding to observation `10561900` can be found in `patches/fr/00/19`, whereas `22068100` can be found in `patches/us/00/81`:"},{"metadata":{"ExecuteTime":{"end_time":"2021-03-08T13:34:38.715235Z","start_time":"2021-03-08T13:34:38.592106Z"},"trusted":true},"cell_type":"code","source":"ls $PATCHES_PATH/fr/00/19/10561900*","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"and"},{"metadata":{"ExecuteTime":{"end_time":"2021-03-08T13:34:38.845241Z","start_time":"2021-03-08T13:34:38.717312Z"},"trusted":true},"cell_type":"code","source":"ls $PATCHES_PATH/us/00/81/22068100*","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"There are 4 files for each observation:\n- a color JPEG image containing an RGB image (`*_rgb.jpg`)\n- a grayscale JPEG image containing a near-infrared image (`*_near_ir.jpg`)\n- a TIFF with Deflate compression containing altitude data (`*_altitude.tif`)\n- a TIFF with Deflate compression containing land cover data (`*_landcover.tif`)\n\nWe provide a loading function which, given an observation id, loads all this data at once using [Pillow](https://pillow.readthedocs.io/en/stable/) for the images and [tiffile](https://github.com/cgohlke/tifffile) for the TIFF files and returns them as a tuple `(rgb, near-ir, altitude, landcover)`:"},{"metadata":{"ExecuteTime":{"end_time":"2021-03-08T13:34:38.882506Z","start_time":"2021-03-08T13:34:38.847185Z"},"trusted":true},"cell_type":"code","source":"from GLC.data_loading.common import load_patch\n\npatch = load_patch(10561900, PATCHES_PATH)\n\nprint(\"Number of data sources: {}\".format(len(patch)))\nprint(\"Arrays shape: {}\".format([p.shape for p in patch]))\nprint(\"Data types: {}\".format([p.dtype for p in patch]))","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"It can also automatically perform the land cover alignment if necessary:"},{"metadata":{"ExecuteTime":{"end_time":"2021-03-08T13:34:38.892165Z","start_time":"2021-03-08T13:34:38.884554Z"},"trusted":true},"cell_type":"code","source":"landcover_mapping = df_suggested_landcover_alignment[\"suggested_landcover_code\"].values\npatch = load_patch(10561900, PATCHES_PATH, landcover_mapping)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"We also provide an visualization function for the patches:"},{"metadata":{"ExecuteTime":{"end_time":"2021-03-08T13:34:39.480455Z","start_time":"2021-03-08T13:34:38.893677Z"},"trusted":true},"cell_type":"code","source":"from GLC.plotting import visualize_observation_patch\n\n# Extracts land cover labels\nlandcover_labels = df_suggested_landcover_alignment[[\"suggested_landcover_code\", \"suggested_landcover_label\"]].drop_duplicates().sort_values(\"suggested_landcover_code\")[\"suggested_landcover_label\"].values\n\nvisualize_observation_patch(patch, landcover_labels)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Similarly, for the observation `22068100`:"},{"metadata":{"ExecuteTime":{"end_time":"2021-03-08T13:34:39.89696Z","start_time":"2021-03-08T13:34:39.484194Z"},"trusted":true},"cell_type":"code","source":"patch = load_patch(22068100, PATCHES_PATH, landcover_mapping)\n\nvisualize_observation_patch(patch, landcover_labels)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# Environmental rasters\n\nThe rasters contain low-resolution environmental data - bioclimatic and pedological data.\n\nThere are two ways to use this data:\n1. directly use the environmental vectors pre-extracted that can be found in the CSV file `pre-extracted/environmental_vectors.csv`\n2. manually extract patches centered at each observation using the rasters located in the `rasters` subfolder"},{"metadata":{},"cell_type":"markdown","source":"## Pre-extracted environmental vectors\n\nThese vectors are ready to be used - see the Random Forest training baseline in the corresponding notebook.\n\nThey are easy to load as they are provided as a CSV file.\n\nEach line of this file correspond to an observation and each column to one of the environmental variable."},{"metadata":{"ExecuteTime":{"end_time":"2021-03-08T13:34:48.989383Z","start_time":"2021-03-08T13:34:39.898763Z"},"trusted":true},"cell_type":"code","source":"df_env = pd.read_csv(DATA_PATH / \"pre-extracted\" / \"environmental_vectors.csv\", sep=\";\", index_col=\"observation_id\")\ndf_env.head()","execution_count":null,"outputs":[]},{"metadata":{"ExecuteTime":{"end_time":"2021-03-05T19:46:04.789976Z","start_time":"2021-03-05T19:46:04.784343Z"}},"cell_type":"markdown","source":"Note that it typically contains NaN values due to absence of data over the seas and oceans for both types of data as well as rivers and others for the pedologic data."},{"metadata":{"ExecuteTime":{"end_time":"2021-03-08T13:34:49.047297Z","start_time":"2021-03-08T13:34:48.991539Z"},"trusted":true},"cell_type":"code","source":"print(\"Variables which can contain NaN values:\")\ndf_env.isna().any()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## Patch extraction from rasters\n\nTo more easily extract patches from the rasters, we provide a `PatchExtractor` class which uses [rasterio](https://github.com/mapbox/rasterio)."},{"metadata":{"ExecuteTime":{"end_time":"2021-03-08T13:34:49.267875Z","start_time":"2021-03-08T13:34:49.049279Z"},"trusted":true},"cell_type":"code","source":"from GLC.data_loading.environmental_raster import PatchExtractor","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"The following code loads the rasters for all the variables and prepares to extract patches of size 256x256.\n\nHere the patches are not of the same resolution as the provided ones as one pixel corresponds to 30arcsec (~1km) for the bioclimatic data and to 250m for the pedologic data.\n\nNote that this uses quite a lot of memory (~18Go) as all the rasters will be loaded in the RAM.\n\nTo avoid this issue, we will only load the bioclimatic rasters here."},{"metadata":{"ExecuteTime":{"end_time":"2021-03-08T13:37:47.42398Z","start_time":"2021-03-08T13:34:49.269255Z"},"trusted":true},"cell_type":"code","source":"extractor = PatchExtractor(DATA_PATH / \"rasters\", size=256)\nextractor.add_all_bioclimatic_rasters()\n\nprint(\"Number of rasters: {}\".format(len(extractor)))","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"To load all the rasters use:\n```python\nextractor.add_all_rasters()\n```\n\nTo load all the pedologic rasters use:\n```python\nextractor.add_all_pedologic_rasters()\n```"},{"metadata":{},"cell_type":"markdown","source":"A patch can then easily to be extracted given the localization using:"},{"metadata":{"ExecuteTime":{"end_time":"2021-03-08T13:37:47.444384Z","start_time":"2021-03-08T13:37:47.428825Z"},"trusted":true},"cell_type":"code","source":"patch = extractor[43.61, 3.88]\n\nprint(\"Patch shape: {}\".format(patch.shape))\nprint(\"Data type: {}\".format(patch.dtype))","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Note that it typically contains NaN values due to absence of data over the seas and oceans for both types of data as well as rivers and others for the pedologic data."},{"metadata":{"ExecuteTime":{"end_time":"2021-03-08T13:37:47.611023Z","start_time":"2021-03-08T13:37:47.445804Z"},"trusted":true},"cell_type":"code","source":"print(\"Contains NaN: {}\".format(np.isnan(patch).any()))","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"A helper function to plot the patches is also provided.\n\nThe following example displays the patches obtained around the region of Montpellier, France."},{"metadata":{"ExecuteTime":{"end_time":"2021-03-08T13:37:52.086992Z","start_time":"2021-03-08T13:37:47.613145Z"},"trusted":true},"cell_type":"code","source":"fig = plt.figure(figsize=(14, 10))\nextractor.plot((43.61, 3.88), fig=fig)","execution_count":null,"outputs":[]}],"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat":4,"nbformat_minor":4}