{"cells":[{"metadata":{},"cell_type":"markdown","source":"This notebook presents some code to compute some basic baselines.\n\nIn particular, it shows how to:\n1. Use the provided validation set\n2. Compute the top-30 metric\n3. Save the predictions on the test in the right format for submission"},{"metadata":{"ExecuteTime":{"end_time":"2021-03-12T13:26:34.897763Z","start_time":"2021-03-12T13:26:34.177058Z"},"trusted":true},"cell_type":"code","source":"%pylab inline --no-import-all\n\nimport os\nfrom pathlib import Path\n\nimport pandas as pd\n\n\n# Change this path to adapt to where you downloaded the data\nBASE_PATH = Path(\"../input/geolifeclef-2021/\")\nDATA_PATH = BASE_PATH / \"data\"\n\n# Create the path to save submission files\nSUBMISSION_PATH = Path(\"submissions\")\nos.makedirs(SUBMISSION_PATH, exist_ok=True)\n\n# Clone the GitHub repository\n!rm -rf GLC\n!git clone https://github.com/maximiliense/GLC","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"We also load the official metric, top-30 error rate, for which we provide efficient implementations:"},{"metadata":{"ExecuteTime":{"end_time":"2021-03-12T13:26:34.912516Z","start_time":"2021-03-12T13:26:34.899648Z"},"trusted":true},"cell_type":"code","source":"from GLC.metrics import top_30_error_rate\nhelp(top_30_error_rate)","execution_count":null,"outputs":[]},{"metadata":{"ExecuteTime":{"end_time":"2021-03-12T13:26:34.917798Z","start_time":"2021-03-12T13:26:34.914363Z"},"trusted":true},"cell_type":"code","source":"from GLC.metrics import top_k_error_rate_from_sets\nhelp(top_k_error_rate_from_sets)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"For submissions, we will also need to predict the top-30 sets for which we also provide an efficient implementation:"},{"metadata":{"ExecuteTime":{"end_time":"2021-03-12T13:26:35.119447Z","start_time":"2021-03-12T13:26:35.115187Z"},"trusted":true},"cell_type":"code","source":"from GLC.metrics import predict_top_30_set\nhelp(predict_top_30_set)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# Observation data loading"},{"metadata":{},"cell_type":"markdown","source":"We first need to load the observation data:"},{"metadata":{"ExecuteTime":{"end_time":"2021-03-12T13:26:37.167338Z","start_time":"2021-03-12T13:26:36.067307Z"},"trusted":true},"cell_type":"code","source":"df_fr = pd.read_csv(DATA_PATH / \"observations\" / \"observations_fr_train.csv\", sep=\";\", index_col=\"observation_id\")\ndf_us = pd.read_csv(DATA_PATH / \"observations\" / \"observations_us_train.csv\", sep=\";\", index_col=\"observation_id\")\ndf = pd.concat((df_fr, df_us))","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Then, we retrieve the train/val split provided:"},{"metadata":{"ExecuteTime":{"end_time":"2021-03-12T13:26:37.747978Z","start_time":"2021-03-12T13:26:37.169035Z"},"trusted":true},"cell_type":"code","source":"obs_id_train = df.index[df[\"subset\"] == \"train\"].values\nobs_id_val = df.index[df[\"subset\"] == \"val\"].values\n\ny_train = df.loc[obs_id_train][\"species_id\"].values\ny_val = df.loc[obs_id_val][\"species_id\"].values\n\nn_val = len(obs_id_val)\nprint(\"Validation set size: {} ({:.1%} of train observations)\".format(n_val, n_val / len(df)))","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"We also load the observation data for the test set:"},{"metadata":{"ExecuteTime":{"end_time":"2021-03-12T13:26:37.79393Z","start_time":"2021-03-12T13:26:37.750408Z"},"trusted":true},"cell_type":"code","source":"df_fr_test = pd.read_csv(DATA_PATH / \"observations\" / \"observations_fr_test.csv\", sep=\";\", index_col=\"observation_id\")\ndf_us_test = pd.read_csv(DATA_PATH / \"observations\" / \"observations_us_test.csv\", sep=\";\", index_col=\"observation_id\")\n\ndf_test = pd.concat((df_fr_test, df_us_test))\n\nobs_id_test = df_test.index\n\nprint(\"Number of observations for testing: {}\".format(len(df_test)))\n\ndf_test.head()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"For submissions, we also need the following mapping to correct a slight misalignment in the test observation ids:"},{"metadata":{"ExecuteTime":{"end_time":"2021-03-12T13:26:37.819873Z","start_time":"2021-03-12T13:26:37.795945Z"},"trusted":true},"cell_type":"code","source":"df_test_obs_id_mapping = pd.read_csv(BASE_PATH / \"test_observation_ids_mapping.csv\", sep=\";\")\ndf_test_obs_id_mapping.head()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# Sample submission file\n\nIn this section, we will demonstrate how to generate the sample submission file provided.\n\nTo do so, we will use this function:"},{"metadata":{"ExecuteTime":{"end_time":"2021-03-12T14:18:06.02525Z","start_time":"2021-03-12T14:18:06.014102Z"},"trusted":true},"cell_type":"code","source":"def generate_submission_file(filename, corrected_observation_ids, s_pred):\n    s_pred = [\n        \" \".join(map(str, pred_set))\n        for pred_set in s_pred\n    ]\n    \n    df = pd.DataFrame({\n        \"Id\": corrected_observation_ids,\n        \"Predicted\": s_pred\n    })\n    df.to_csv(filename, index=False)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"The sample submission consists in always predicting the first 30 species for all the test observations:"},{"metadata":{"ExecuteTime":{"end_time":"2021-03-12T14:18:18.702441Z","start_time":"2021-03-12T14:18:18.592647Z"},"trusted":true},"cell_type":"code","source":"first_30_species = np.arange(30)\ns_pred = np.tile(first_30_species[None], (len(df_test), 1))","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"We can then generate the associated submission file using:\n\n**(NOTE that we need to use the adjusted test ids mapping here)**"},{"metadata":{"ExecuteTime":{"end_time":"2021-03-12T14:18:31.144096Z","start_time":"2021-03-12T14:18:31.127866Z"},"trusted":true},"cell_type":"code","source":"generate_submission_file(SUBMISSION_PATH / \"sample_submission.csv\", df_test_obs_id_mapping[\"Id\"], s_pred)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# Constant baseline: 30 most observed species\n\nThe first baseline consists in predicting the 30 most observed species on the train set which corresponds exactly to the \"Top-30 most present species\":"},{"metadata":{"ExecuteTime":{"end_time":"2021-03-12T13:26:38.031733Z","start_time":"2021-03-12T13:26:37.821697Z"},"trusted":true},"cell_type":"code","source":"species_distribution = df.loc[obs_id_train][\"species_id\"].value_counts(normalize=True)\ntop_30_most_observed = species_distribution.index.values[:30]","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"As expected, it does not perform very well on the validation set:"},{"metadata":{"ExecuteTime":{"end_time":"2021-03-12T13:26:38.426825Z","start_time":"2021-03-12T13:26:38.407391Z"},"trusted":true},"cell_type":"code","source":"s_pred = np.tile(top_30_most_observed[None], (n_val, 1))\nscore = top_k_error_rate_from_sets(y_val, s_pred)\nprint(\"Top-30 error rate: {:.1%}\".format(score))","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"We will however generate the associated submission file on the test using:"},{"metadata":{"ExecuteTime":{"end_time":"2021-03-12T13:26:40.221981Z","start_time":"2021-03-12T13:26:39.222045Z"},"trusted":true},"cell_type":"code","source":"# Compute baseline on the test set\nn_test = len(df_test)\ns_pred = np.tile(top_30_most_observed[None], (n_test, 1))\n\n# Generate the submission file\ngenerate_submission_file(SUBMISSION_PATH / \"constant_top_30_most_present_species_baseline.csv\", df_test_obs_id_mapping[\"Id\"], s_pred)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# Random forest on environmental vectors\n\nA classical approach in ecology is to train Random Forests on environmental vectors.\n\nWe show here how to do so using [scikit-learn](https://scikit-learn.org/).\n\nWe start by loading the environmental vectors:"},{"metadata":{"ExecuteTime":{"end_time":"2021-03-12T13:26:49.451928Z","start_time":"2021-03-12T13:26:40.223618Z"},"trusted":true},"cell_type":"code","source":"df_env = pd.read_csv(DATA_PATH / \"pre-extracted\" / \"environmental_vectors.csv\", sep=\";\", index_col=\"observation_id\")\n\nX_train = df_env.loc[obs_id_train].values\nX_val = df_env.loc[obs_id_val].values\nX_test = df_env.loc[obs_id_test].values","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Then, we need to handle properly the missing values.\n\nFor instance, using `SimpleImputer`:"},{"metadata":{"ExecuteTime":{"end_time":"2021-03-12T13:26:50.817494Z","start_time":"2021-03-12T13:26:49.45442Z"},"trusted":true},"cell_type":"code","source":"from sklearn.impute import SimpleImputer\nimp = SimpleImputer(missing_values=np.nan, strategy=\"mean\")\nimp.fit(X_train)\n\nX_train = imp.transform(X_train)\nX_val = imp.transform(X_val)\nX_test = imp.transform(X_test)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"We can now start training our Random Forest (as there are a lot of observations, over 1.8M, this can take a while):"},{"metadata":{"ExecuteTime":{"end_time":"2021-03-12T13:33:20.000908Z","start_time":"2021-03-12T13:26:50.818924Z"},"trusted":true},"cell_type":"code","source":"from sklearn.ensemble import RandomForestClassifier\nest = RandomForestClassifier(n_estimators=16, max_depth=10, n_jobs=-1)\nest.fit(X_train, y_train)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"As there are a lot of classes (over 30K), we need to be cautious when predicting the scores of the model.\n\nThis can easily take more then 10Go on the validation set.\n\nFor this reason, we will be predict the top-30 sets by batches using the following generic function:"},{"metadata":{"ExecuteTime":{"end_time":"2021-03-12T13:33:20.00567Z","start_time":"2021-03-12T13:33:20.002405Z"},"trusted":true},"cell_type":"code","source":"def batch_predict(predict_func, X, batch_size=1024):\n    res = predict_func(X[:1])\n    n_samples, n_outputs, dtype = X.shape[0], res.shape[1], res.dtype\n    \n    preds = np.empty((n_samples, n_outputs), dtype=dtype)\n    \n    for i in range(0, len(X), batch_size):\n        X_batch = X[i:i+batch_size]\n        preds[i:i+batch_size] = predict_func(X_batch)\n            \n    return preds","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"We can know compute the top-30 error rate on the validation set:"},{"metadata":{"ExecuteTime":{"start_time":"2021-03-12T13:25:29.308Z"},"trusted":true},"cell_type":"code","source":"def predict_func(X):\n    y_score = est.predict_proba(X)\n    s_pred = predict_top_30_set(y_score)\n    return s_pred\n\ns_val = batch_predict(predict_func, X_val, batch_size=1024)\nscore_val = top_k_error_rate_from_sets(y_val, s_val)\nprint(\"Top-30 error rate: {:.1%}\".format(score_val))","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"We now predict the top-30 sets on the test data and save them in a submission file:"},{"metadata":{"ExecuteTime":{"end_time":"2021-03-12T13:04:36.163868Z","start_time":"2021-03-12T13:03:49.421715Z"},"trusted":true},"cell_type":"code","source":"# Compute baseline on the test set\ns_pred = batch_predict(predict_func, X_test, batch_size=1024)\n\n# Generate the submission file\ngenerate_submission_file(SUBMISSION_PATH / \"random_forest_on_environmental_vectors.csv\", df_test_obs_id_mapping[\"Id\"], s_pred)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"","execution_count":null,"outputs":[]}],"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat":4,"nbformat_minor":4}