{
  "id": 325767,
  "title": "Sharing thoughts: Single class labeling is probably misleading models",
  "url": "/competitions/geolifeclef-2022-lifeclef-2022-fgvc9/discussion/325767",
  "author_name": "",
  "post_date": "2022-05-18T08:33:35.576273400Z",
  "votes": 3,
  "comment_count": 6,
  "views": 0,
  "content": "<p>It is sure that for every single sample location, satellite images and environmental info there will be probably multiple species of plants and animals there in that place. But we are training our models by telling them one at a time, one label per each sample data, which makes it change its parameters a little towards to predict that class for that data but, at the same time, erasing also a little the weights of predicting any other class for that observation sample, including other labels that would be there or in a similar environment and the model probably already somehow learnt from previous samples. In other words: when you tell the model what there is in that sample by putting a 1 in the array of ground truth predictions, you are also putting zeros in all of the others, so telling the model what there is not in that sample. </p>\n<p>I know that building datasets is usually far more difficult than building and training models and for this particular case… I cannot imagine the effort of collecting millions of samples in millions of locations for thousand of species, that has been for sure a really impressive work from many people. Even so, from my point of view, a better way of doing it to use this data for AI, could be to focus more on get every single possible species in every location for any season of the year even if that means to have less samples. We could also try to merge labels from close or similar locations but I think that it wouldn't be much accurate as most of the samples are quite far from each other and satellite images wouldn't be the same, so MAYBE this could improve it a little but it would be still quite worse than really having all possible true observations from a single location.</p>\n<p>So, am I right? Anyone else thinks that this is happening as I wrote or maybe what I said is not accurate? To be totally clear, this is just my opinion in order to improve this specific challenge, but also to solving this kind of applications, after being involved in this competition for some weeks and achieving at this current moment the second position in the provisional Leaderboard. I'm totally amazed and grateful for this huge dataset and all the people who has been involved in its creation and I just want to give a constructive and positive thought on how this data should be collected and labeled to be more useful and take more of it with the same effort :). Feel free to correct me if I'm wrong and share your thoughts also about it.</p>",
  "messages": [
    {
      "id": "1793798",
      "postDate": "05/18/2022 08:33:35",
      "content": "<p>It is sure that for every single sample location, satellite images and environmental info there will be probably multiple species of plants and animals there in that place. But we are training our models by telling them one at a time, one label per each sample data, which makes it change its parameters a little towards to predict that class for that data but, at the same time, erasing also a little the weights of predicting any other class for that observation sample, including other labels that would be there or in a similar environment and the model probably already somehow learnt from previous samples. In other words: when you tell the model what there is in that sample by putting a 1 in the array of ground truth predictions, you are also putting zeros in all of the others, so telling the model what there is not in that sample. </p>\n<p>I know that building datasets is usually far more difficult than building and training models and for this particular case… I cannot imagine the effort of collecting millions of samples in millions of locations for thousand of species, that has been for sure a really impressive work from many people. Even so, from my point of view, a better way of doing it to use this data for AI, could be to focus more on get every single possible species in every location for any season of the year even if that means to have less samples. We could also try to merge labels from close or similar locations but I think that it wouldn't be much accurate as most of the samples are quite far from each other and satellite images wouldn't be the same, so MAYBE this could improve it a little but it would be still quite worse than really having all possible true observations from a single location.</p>\n<p>So, am I right? Anyone else thinks that this is happening as I wrote or maybe what I said is not accurate? To be totally clear, this is just my opinion in order to improve this specific challenge, but also to solving this kind of applications, after being involved in this competition for some weeks and achieving at this current moment the second position in the provisional Leaderboard. I'm totally amazed and grateful for this huge dataset and all the people who has been involved in its creation and I just want to give a constructive and positive thought on how this data should be collected and labeled to be more useful and take more of it with the same effort :). Feel free to correct me if I'm wrong and share your thoughts also about it.</p>",
      "rawMarkdown": "It is sure that for every single sample location, satellite images and environmental info there will be probably multiple species of plants and animals there in that place. But we are training our models by telling them one at a time, one label per each sample data, which makes it change its parameters a little towards to predict that class for that data but, at the same time, erasing also a little the weights of predicting any other class for that observation sample, including other labels that would be there or in a similar environment and the model probably already somehow learnt from previous samples. In other words: when you tell the model what there is in that sample by putting a 1 in the array of ground truth predictions, you are also putting zeros in all of the others, so telling the model what there is not in that sample. \n\nI know that building datasets is usually far more difficult than building and training models and for this particular case... I cannot imagine the effort of collecting millions of samples in millions of locations for thousand of species, that has been for sure a really impressive work from many people. Even so, from my point of view, a better way of doing it to use this data for AI, could be to focus more on get every single possible species in every location for any season of the year even if that means to have less samples. We could also try to merge labels from close or similar locations but I think that it wouldn't be much accurate as most of the samples are quite far from each other and satellite images wouldn't be the same, so MAYBE this could improve it a little but it would be still quite worse than really having all possible true observations from a single location.\n\nSo, am I right? Anyone else thinks that this is happening as I wrote or maybe what I said is not accurate? To be totally clear, this is just my opinion in order to improve this specific challenge, but also to solving this kind of applications, after being involved in this competition for some weeks and achieving at this current moment the second position in the provisional Leaderboard. I'm totally amazed and grateful for this huge dataset and all the people who has been involved in its creation and I just want to give a constructive and positive thought on how this data should be collected and labeled to be more useful and take more of it with the same effort :). Feel free to correct me if I'm wrong and share your thoughts also about it.",
      "votes": null
    },
    {
      "id": "1794156",
      "postDate": "05/18/2022 14:42:25",
      "content": "<p>In ecology, what you are telling about is often referred as presence-only species distribution models vs. presence-absence species distribution models (also called site occupancy models). In this Kaggle challenge, we are tackling the presence-only problem which is indeed harder than the presence-absence case where you know all presence and absence of species at each location. Presence-only observations are much cheaper to collect (e.g. through apps like Pl@ntNet or iNaturalist) and have the potential to allow monitoring biodiversity at a higher frequency and higher spatial resolution. Site-occupancy data also exist (e.g. the European Vegetation Archive) but are much more heterogeneous in terms of year of production, spatial uncertainty (up to several kms), spatial extend of the inventory area, etc. Still, we are currently discussing with them to see if this could be part of further challenges. </p>",
      "rawMarkdown": "In ecology, what you are telling about is often referred as presence-only species distribution models vs. presence-absence species distribution models (also called site occupancy models). In this Kaggle challenge, we are tackling the presence-only problem which is indeed harder than the presence-absence case where you know all presence and absence of species at each location. Presence-only observations are much cheaper to collect (e.g. through apps like Pl@ntNet or iNaturalist) and have the potential to allow monitoring biodiversity at a higher frequency and higher spatial resolution. Site-occupancy data also exist (e.g. the European Vegetation Archive) but are much more heterogeneous in terms of year of production, spatial uncertainty (up to several kms), spatial extend of the inventory area, etc. Still, we are currently discussing with them to see if this could be part of further challenges.",
      "votes": null
    },
    {
      "id": "1794179",
      "postDate": "05/18/2022 15:02:14",
      "content": "<p>Thank you very much for the explanation! I'm not very familiar with this sector and now it makes better sense for me the dataset circumstances and trade-offs.</p>",
      "rawMarkdown": "Thank you very much for the explanation! I'm not very familiar with this sector and now it makes better sense for me the dataset circumstances and trade-offs.",
      "votes": null
    },
    {
      "id": "1795109",
      "postDate": "05/19/2022 12:14:54",
      "content": "<p>The local aggregation of presence-only data is indeed an option to try to partly recover presence-absence data while trading off spatial accuracy. We have started to look into this for some other works since a few weeks. If you tried something similar, we would be very interested to see if it works or not on GeoLifeCLEF 2022 😃</p>",
      "rawMarkdown": "The local aggregation of presence-only data is indeed an option to try to partly recover presence-absence data while trading off spatial accuracy. We have started to look into this for some other works since a few weeks. If you tried something similar, we would be very interested to see if it works or not on GeoLifeCLEF 2022 😃",
      "votes": null
    },
    {
      "id": "1795132",
      "postDate": "05/19/2022 13:05:57",
      "content": "<p>I was starting to implement it in a somehow efficient and fast way (of course not going through every pair possibility of samples to see if they are close, that would be around 1e12 checks), let's see if I can finish it and integrate that data in a hole model traning process from scratch, I will tell if I finally make it :)</p>",
      "rawMarkdown": "I was starting to implement it in a somehow efficient and fast way (of course not going through every pair possibility of samples to see if they are close, that would be around 1e12 checks), let's see if I can finish it and integrate that data in a hole model traning process from scratch, I will tell if I finally make it :)",
      "votes": null
    },
    {
      "id": "1799719",
      "postDate": "05/24/2022 08:36:18",
      "content": "<p>Okay so finally I applied it, I aggregated close species id in each observation by clustering them in smaller squared regions in order to reduce the search complexity. I realised that in some cases, there could be up to 400-500 species in less than a squared km aprox. So I decided to just aggregate up to the top-30 closer to the original label. I tried to train a couple of models with the entire inputs that I was already using and testing 3 different loss functions and the validation loss and error has been pretty bad in all cases, the best one achieved a top30 val error of 0.75 or so. In some cases, the train error was decreasing quite good, but validation metrics weren't following it, I tried to simplify the model and apply some common regularization techniques trying to avoid the overfitting and it didn't work. I then tried to put maximum 5 labels instead of 30 and again similar results. I would like to understand better what could be happening there and find better loss functions or strategies, although some are quite restricted due to the large number of classes, features and samples. But anyway, I already spent quite enough time until now on this problem 😅, but also learnt some new very interesting things of course!</p>",
      "rawMarkdown": "Okay so finally I applied it, I aggregated close species id in each observation by clustering them in smaller squared regions in order to reduce the search complexity. I realised that in some cases, there could be up to 400-500 species in less than a squared km aprox. So I decided to just aggregate up to the top-30 closer to the original label. I tried to train a couple of models with the entire inputs that I was already using and testing 3 different loss functions and the validation loss and error has been pretty bad in all cases, the best one achieved a top30 val error of 0.75 or so. In some cases, the train error was decreasing quite good, but validation metrics weren't following it, I tried to simplify the model and apply some common regularization techniques trying to avoid the overfitting and it didn't work. I then tried to put maximum 5 labels instead of 30 and again similar results. I would like to understand better what could be happening there and find better loss functions or strategies, although some are quite restricted due to the large number of classes, features and samples. But anyway, I already spent quite enough time until now on this problem 😅, but also learnt some new very interesting things of course!",
      "votes": null
    },
    {
      "id": "1805912",
      "postDate": "05/30/2022 15:58:29",
      "content": "<p>These ideas of \"learning from presence-only data\" and \"evaluating predictions with presence-only data\" are very interesting to many of us from a research perspective!</p>\n<p>If you're interested in learning more about the presence-only / presence-absence problem from the ecology side:</p>\n<ul>\n<li><a href=\"https://esajournals.onlinelibrary.wiley.com/doi/abs/10.1890/07-2153.1\" target=\"_blank\">Sample selection bias and presence-only distribution models: implications for background and pseudo-absence data</a></li>\n<li><a href=\"https://www.annualreviews.org/doi/abs/10.1146/annurev.ecolsys.110308.120159\" target=\"_blank\">Species Distribution Models: Ecological Explanation and Prediction Across Space and Time\n</a></li>\n</ul>\n<p>There's also some related work from the machine learning / computer vision side from folks involved in GeoLifeCLEF, including:</p>\n<ul>\n<li><a href=\"https://link.springer.com/chapter/10.1007/978-3-319-76445-0_10\" target=\"_blank\">A Deep Learning Approach to Species Distribution Modelling</a></li>\n<li><a href=\"https://journals.plos.org/ploscompbiol/article?id=10.1371/journal.pcbi.1008856\" target=\"_blank\">Convolutional neural networks improve species distribution modelling by capturing the spatial structure of the environment</a></li>\n<li><a href=\"https://arxiv.org/abs/2106.09708\" target=\"_blank\">Multi-Label Learning from Single Positive Labels\n</a></li>\n<li><a href=\"https://arxiv.org/abs/1906.05272\" target=\"_blank\">Presence-Only Geographical Priors for Fine-Grained Image Classification</a></li>\n<li><a href=\"https://arxiv.org/abs/2107.10400\" target=\"_blank\">Species Distribution Modeling for Machine Learning Practitioners: A Review</a></li>\n</ul>",
      "rawMarkdown": "These ideas of \"learning from presence-only data\" and \"evaluating predictions with presence-only data\" are very interesting to many of us from a research perspective!\n\nIf you're interested in learning more about the presence-only / presence-absence problem from the ecology side:\n- [Sample selection bias and presence-only distribution models: implications for background and pseudo-absence data](https://esajournals.onlinelibrary.wiley.com/doi/abs/10.1890/07-2153.1)\n- [Species Distribution Models: Ecological Explanation and Prediction Across Space and Time\n](https://www.annualreviews.org/doi/abs/10.1146/annurev.ecolsys.110308.120159)\n\nThere's also some related work from the machine learning / computer vision side from folks involved in GeoLifeCLEF, including:\n- [A Deep Learning Approach to Species Distribution Modelling](https://link.springer.com/chapter/10.1007/978-3-319-76445-0_10)\n- [Convolutional neural networks improve species distribution modelling by capturing the spatial structure of the environment](https://journals.plos.org/ploscompbiol/article?id=10.1371/journal.pcbi.1008856)\n- [Multi-Label Learning from Single Positive Labels\n](https://arxiv.org/abs/2106.09708)\n- [Presence-Only Geographical Priors for Fine-Grained Image Classification](https://arxiv.org/abs/1906.05272)\n- [Species Distribution Modeling for Machine Learning Practitioners: A Review](https://arxiv.org/abs/2107.10400)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1794156,
      "author_name": "alxjoly",
      "author_url": "",
      "post_date": "05/18/2022 14:42:25",
      "content": "<p>In ecology, what you are telling about is often referred as presence-only species distribution models vs. presence-absence species distribution models (also called site occupancy models). In this Kaggle challenge, we are tackling the presence-only problem which is indeed harder than the presence-absence case where you know all presence and absence of species at each location. Presence-only observations are much cheaper to collect (e.g. through apps like Pl@ntNet or iNaturalist) and have the potential to allow monitoring biodiversity at a higher frequency and higher spatial resolution. Site-occupancy data also exist (e.g. the European Vegetation Archive) but are much more heterogeneous in terms of year of production, spatial uncertainty (up to several kms), spatial extend of the inventory area, etc. Still, we are currently discussing with them to see if this could be part of further challenges. </p>",
      "votes": null,
      "replies": [
        {
          "id": 1794179,
          "author_name": "edomingo",
          "author_url": "",
          "post_date": "05/18/2022 15:02:14",
          "content": "<p>Thank you very much for the explanation! I'm not very familiar with this sector and now it makes better sense for me the dataset circumstances and trade-offs.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1795109,
          "author_name": "tlorieul",
          "author_url": "",
          "post_date": "05/19/2022 12:14:54",
          "content": "<p>The local aggregation of presence-only data is indeed an option to try to partly recover presence-absence data while trading off spatial accuracy. We have started to look into this for some other works since a few weeks. If you tried something similar, we would be very interested to see if it works or not on GeoLifeCLEF 2022 😃</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1795132,
          "author_name": "edomingo",
          "author_url": "",
          "post_date": "05/19/2022 13:05:57",
          "content": "<p>I was starting to implement it in a somehow efficient and fast way (of course not going through every pair possibility of samples to see if they are close, that would be around 1e12 checks), let's see if I can finish it and integrate that data in a hole model traning process from scratch, I will tell if I finally make it :)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1799719,
          "author_name": "edomingo",
          "author_url": "",
          "post_date": "05/24/2022 08:36:18",
          "content": "<p>Okay so finally I applied it, I aggregated close species id in each observation by clustering them in smaller squared regions in order to reduce the search complexity. I realised that in some cases, there could be up to 400-500 species in less than a squared km aprox. So I decided to just aggregate up to the top-30 closer to the original label. I tried to train a couple of models with the entire inputs that I was already using and testing 3 different loss functions and the validation loss and error has been pretty bad in all cases, the best one achieved a top30 val error of 0.75 or so. In some cases, the train error was decreasing quite good, but validation metrics weren't following it, I tried to simplify the model and apply some common regularization techniques trying to avoid the overfitting and it didn't work. I then tried to put maximum 5 labels instead of 30 and again similar results. I would like to understand better what could be happening there and find better loss functions or strategies, although some are quite restricted due to the large number of classes, features and samples. But anyway, I already spent quite enough time until now on this problem 😅, but also learnt some new very interesting things of course!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1805912,
      "author_name": "elijahcole",
      "author_url": "",
      "post_date": "05/30/2022 15:58:29",
      "content": "<p>These ideas of \"learning from presence-only data\" and \"evaluating predictions with presence-only data\" are very interesting to many of us from a research perspective!</p>\n<p>If you're interested in learning more about the presence-only / presence-absence problem from the ecology side:</p>\n<ul>\n<li><a href=\"https://esajournals.onlinelibrary.wiley.com/doi/abs/10.1890/07-2153.1\" target=\"_blank\">Sample selection bias and presence-only distribution models: implications for background and pseudo-absence data</a></li>\n<li><a href=\"https://www.annualreviews.org/doi/abs/10.1146/annurev.ecolsys.110308.120159\" target=\"_blank\">Species Distribution Models: Ecological Explanation and Prediction Across Space and Time\n</a></li>\n</ul>\n<p>There's also some related work from the machine learning / computer vision side from folks involved in GeoLifeCLEF, including:</p>\n<ul>\n<li><a href=\"https://link.springer.com/chapter/10.1007/978-3-319-76445-0_10\" target=\"_blank\">A Deep Learning Approach to Species Distribution Modelling</a></li>\n<li><a href=\"https://journals.plos.org/ploscompbiol/article?id=10.1371/journal.pcbi.1008856\" target=\"_blank\">Convolutional neural networks improve species distribution modelling by capturing the spatial structure of the environment</a></li>\n<li><a href=\"https://arxiv.org/abs/2106.09708\" target=\"_blank\">Multi-Label Learning from Single Positive Labels\n</a></li>\n<li><a href=\"https://arxiv.org/abs/1906.05272\" target=\"_blank\">Presence-Only Geographical Priors for Fine-Grained Image Classification</a></li>\n<li><a href=\"https://arxiv.org/abs/2107.10400\" target=\"_blank\">Species Distribution Modeling for Machine Learning Practitioners: A Review</a></li>\n</ul>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1793798": "It is sure that for every single sample location, satellite images and environmental info there will be probably multiple species of plants and animals there in that place. But we are training our models by telling them one at a time, one label per each sample data, which makes it change its parameters a little towards to predict that class for that data but, at the same time, erasing also a little the weights of predicting any other class for that observation sample, including other labels that would be there or in a similar environment and the model probably already somehow learnt from previous samples. In other words: when you tell the model what there is in that sample by putting a 1 in the array of ground truth predictions, you are also putting zeros in all of the others, so telling the model what there is not in that sample. \n\nI know that building datasets is usually far more difficult than building and training models and for this particular case... I cannot imagine the effort of collecting millions of samples in millions of locations for thousand of species, that has been for sure a really impressive work from many people. Even so, from my point of view, a better way of doing it to use this data for AI, could be to focus more on get every single possible species in every location for any season of the year even if that means to have less samples. We could also try to merge labels from close or similar locations but I think that it wouldn't be much accurate as most of the samples are quite far from each other and satellite images wouldn't be the same, so MAYBE this could improve it a little but it would be still quite worse than really having all possible true observations from a single location.\n\nSo, am I right? Anyone else thinks that this is happening as I wrote or maybe what I said is not accurate? To be totally clear, this is just my opinion in order to improve this specific challenge, but also to solving this kind of applications, after being involved in this competition for some weeks and achieving at this current moment the second position in the provisional Leaderboard. I'm totally amazed and grateful for this huge dataset and all the people who has been involved in its creation and I just want to give a constructive and positive thought on how this data should be collected and labeled to be more useful and take more of it with the same effort :). Feel free to correct me if I'm wrong and share your thoughts also about it.",
    "1794156": "In ecology, what you are telling about is often referred as presence-only species distribution models vs. presence-absence species distribution models (also called site occupancy models). In this Kaggle challenge, we are tackling the presence-only problem which is indeed harder than the presence-absence case where you know all presence and absence of species at each location. Presence-only observations are much cheaper to collect (e.g. through apps like Pl@ntNet or iNaturalist) and have the potential to allow monitoring biodiversity at a higher frequency and higher spatial resolution. Site-occupancy data also exist (e.g. the European Vegetation Archive) but are much more heterogeneous in terms of year of production, spatial uncertainty (up to several kms), spatial extend of the inventory area, etc. Still, we are currently discussing with them to see if this could be part of further challenges.",
    "1794179": "Thank you very much for the explanation! I'm not very familiar with this sector and now it makes better sense for me the dataset circumstances and trade-offs.",
    "1795109": "The local aggregation of presence-only data is indeed an option to try to partly recover presence-absence data while trading off spatial accuracy. We have started to look into this for some other works since a few weeks. If you tried something similar, we would be very interested to see if it works or not on GeoLifeCLEF 2022 😃",
    "1795132": "I was starting to implement it in a somehow efficient and fast way (of course not going through every pair possibility of samples to see if they are close, that would be around 1e12 checks), let's see if I can finish it and integrate that data in a hole model traning process from scratch, I will tell if I finally make it :)",
    "1799719": "Okay so finally I applied it, I aggregated close species id in each observation by clustering them in smaller squared regions in order to reduce the search complexity. I realised that in some cases, there could be up to 400-500 species in less than a squared km aprox. So I decided to just aggregate up to the top-30 closer to the original label. I tried to train a couple of models with the entire inputs that I was already using and testing 3 different loss functions and the validation loss and error has been pretty bad in all cases, the best one achieved a top30 val error of 0.75 or so. In some cases, the train error was decreasing quite good, but validation metrics weren't following it, I tried to simplify the model and apply some common regularization techniques trying to avoid the overfitting and it didn't work. I then tried to put maximum 5 labels instead of 30 and again similar results. I would like to understand better what could be happening there and find better loss functions or strategies, although some are quite restricted due to the large number of classes, features and samples. But anyway, I already spent quite enough time until now on this problem 😅, but also learnt some new very interesting things of course!",
    "1805912": "These ideas of \"learning from presence-only data\" and \"evaluating predictions with presence-only data\" are very interesting to many of us from a research perspective!\n\nIf you're interested in learning more about the presence-only / presence-absence problem from the ecology side:\n- [Sample selection bias and presence-only distribution models: implications for background and pseudo-absence data](https://esajournals.onlinelibrary.wiley.com/doi/abs/10.1890/07-2153.1)\n- [Species Distribution Models: Ecological Explanation and Prediction Across Space and Time\n](https://www.annualreviews.org/doi/abs/10.1146/annurev.ecolsys.110308.120159)\n\nThere's also some related work from the machine learning / computer vision side from folks involved in GeoLifeCLEF, including:\n- [A Deep Learning Approach to Species Distribution Modelling](https://link.springer.com/chapter/10.1007/978-3-319-76445-0_10)\n- [Convolutional neural networks improve species distribution modelling by capturing the spatial structure of the environment](https://journals.plos.org/ploscompbiol/article?id=10.1371/journal.pcbi.1008856)\n- [Multi-Label Learning from Single Positive Labels\n](https://arxiv.org/abs/2106.09708)\n- [Presence-Only Geographical Priors for Fine-Grained Image Classification](https://arxiv.org/abs/1906.05272)\n- [Species Distribution Modeling for Machine Learning Practitioners: A Review](https://arxiv.org/abs/2107.10400)"
  },
  "source": "meta"
}