{
  "id": 574578,
  "title": "Some questions on strange artefacts in the PA data",
  "url": "/competitions/geolifeclef-2025/discussion/574578",
  "author_name": "",
  "post_date": "2025-04-22T14:51:31.238064600Z",
  "votes": 1,
  "comment_count": 3,
  "views": 0,
  "content": "<p>Hi <a href=\"https://www.kaggle.com/picekl\" target=\"_blank\">@picekl</a> </p>\n<p>I would like to ask several questions on some aspects of the training data that I simply do not understand. If you can provide clarifications for those, I will be very grateful, but if you consider that answering these will disclose some internal information, please just tell me that you cannot discuss this matter.</p>\n<ol>\n<li><p>There are some <code>surveyId</code> in the PA data, where there <code>speciesId</code> are duplicated. For example <code>surveyId=128364</code> Do these duplicates have any special meaning?</p></li>\n<li><p>There are some <code>surveyId</code> in the PA data, where <code>areaInM2</code> takes multiple values. For example, <code>surveyId=2756037</code>. Another example is <code>surveyId=789799</code>, where some values of <code>areaInM2</code> are <code>-inf</code>. Could you please clarify whether it is possible to deduce the actual area of the survey from these?</p></li>\n<li><p>What are the safe assumptions to make regarding the consistency of different PA surveys? Especially given that there is so huge variation in area of surveys: from 0.04m2 to 8000m2. Shall we expect that if a species is not listed for some <code>surveyId</code> in the PA data, then it was actually missing in that <code>surveyId</code> because it was not present (or at least it was not detected by survey team)? I have some experience in conducting vegetation surveys in the field, and it is simply very difficult to believe that some 700m2 survey there was only one species found, like in <code>surveyId=168111</code>.</p></li>\n</ol>\n<p>Thank you in advance for your time and hope to hear back from you soon!</p>",
  "messages": [
    {
      "id": "3184862",
      "postDate": "04/22/2025 14:51:31",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/picekl\" target=\"_blank\">@picekl</a> </p>\n<p>I would like to ask several questions on some aspects of the training data that I simply do not understand. If you can provide clarifications for those, I will be very grateful, but if you consider that answering these will disclose some internal information, please just tell me that you cannot discuss this matter.</p>\n<ol>\n<li><p>There are some <code>surveyId</code> in the PA data, where there <code>speciesId</code> are duplicated. For example <code>surveyId=128364</code> Do these duplicates have any special meaning?</p></li>\n<li><p>There are some <code>surveyId</code> in the PA data, where <code>areaInM2</code> takes multiple values. For example, <code>surveyId=2756037</code>. Another example is <code>surveyId=789799</code>, where some values of <code>areaInM2</code> are <code>-inf</code>. Could you please clarify whether it is possible to deduce the actual area of the survey from these?</p></li>\n<li><p>What are the safe assumptions to make regarding the consistency of different PA surveys? Especially given that there is so huge variation in area of surveys: from 0.04m2 to 8000m2. Shall we expect that if a species is not listed for some <code>surveyId</code> in the PA data, then it was actually missing in that <code>surveyId</code> because it was not present (or at least it was not detected by survey team)? I have some experience in conducting vegetation surveys in the field, and it is simply very difficult to believe that some 700m2 survey there was only one species found, like in <code>surveyId=168111</code>.</p></li>\n</ol>\n<p>Thank you in advance for your time and hope to hear back from you soon!</p>",
      "rawMarkdown": "Hi @picekl \n\nI would like to ask several questions on some aspects of the training data that I simply do not understand. If you can provide clarifications for those, I will be very grateful, but if you consider that answering these will disclose some internal information, please just tell me that you cannot discuss this matter.\n\n1. There are some `surveyId` in the PA data, where there `speciesId` are duplicated. For example `surveyId=128364` Do these duplicates have any special meaning?\n\n2. There are some `surveyId` in the PA data, where `areaInM2` takes multiple values. For example, `surveyId=2756037`. Another example is `surveyId=789799`, where some values of `areaInM2` are `-inf`. Could you please clarify whether it is possible to deduce the actual area of the survey from these?\n\n3. What are the safe assumptions to make regarding the consistency of different PA surveys? Especially given that there is so huge variation in area of surveys: from 0.04m2 to 8000m2. Shall we expect that if a species is not listed for some `surveyId` in the PA data, then it was actually missing in that `surveyId` because it was not present (or at least it was not detected by survey team)? I have some experience in conducting vegetation surveys in the field, and it is simply very difficult to believe that some 700m2 survey there was only one species found, like in `surveyId=168111`.\n\nThank you in advance for your time and hope to hear back from you soon!",
      "votes": null
    },
    {
      "id": "3185124",
      "postDate": "04/22/2025 21:39:23",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/gtikho\" target=\"_blank\">@gtikho</a>,</p>\n<p>In general, the collection of the Presence-Absence data is pretty tedious, and from my non-botanist perspective, noisy.<br>\nTherefore, the answers to your questions are pretty simple. Hope it helps.</p>\n<blockquote>\n  <p>There are some surveyId in the PA data, where there speciesId are duplicated. For example surveyId=128364 Do these duplicates have any special meaning?</p>\n</blockquote>\n<p>There is no particular meaning. Probably just a noise. You can remove it.</p>\n<blockquote>\n  <p>There are some surveyId in the PA data, where areaInM2 takes multiple values. For example, surveyId=2756037. Another example is surveyId=789799, where some values of areaInM2 are -inf. Could you please clarify whether it is possible to deduce the actual area of the survey from these?</p>\n</blockquote>\n<p>If there are multiple values and one is a real number, I would suggest using it. Usually, if there is -inf or zero, the data are missing.</p>\n<blockquote>\n  <p>What are the safe assumptions to make regarding the consistency of different PA surveys? Especially given that there is so huge variation in area of surveys: from 0.04m2 to 8000m2. Shall we expect that if a species is not listed for some surveyId in the PA data, then it was actually missing in that surveyId because it was not present (or at least it was not detected by survey team)? I have some experience in conducting vegetation surveys in the field, and it is simply very difficult to believe that some 700m2 survey there was only one species found, like in surveyId=168111.</p>\n</blockquote>\n<p>Usually, if the value is filled, it is more or less accurate. The large areas could include \"water\", so one species is possible.<br>\nDid you check the location?</p>\n<p>Best,<br>\nLukas</p>",
      "rawMarkdown": "Hi @gtikho,\n\nIn general, the collection of the Presence-Absence data is pretty tedious, and from my non-botanist perspective, noisy.\nTherefore, the answers to your questions are pretty simple. Hope it helps.\n\n> There are some surveyId in the PA data, where there speciesId are duplicated. For example surveyId=128364 Do these duplicates have any special meaning?\n\nThere is no particular meaning. Probably just a noise. You can remove it.\n\n\n> There are some surveyId in the PA data, where areaInM2 takes multiple values. For example, surveyId=2756037. Another example is surveyId=789799, where some values of areaInM2 are -inf. Could you please clarify whether it is possible to deduce the actual area of the survey from these?\n\nIf there are multiple values and one is a real number, I would suggest using it. Usually, if there is -inf or zero, the data are missing.\n\n\n> What are the safe assumptions to make regarding the consistency of different PA surveys? Especially given that there is so huge variation in area of surveys: from 0.04m2 to 8000m2. Shall we expect that if a species is not listed for some surveyId in the PA data, then it was actually missing in that surveyId because it was not present (or at least it was not detected by survey team)? I have some experience in conducting vegetation surveys in the field, and it is simply very difficult to believe that some 700m2 survey there was only one species found, like in surveyId=168111.\n\nUsually, if the value is filled, it is more or less accurate. The large areas could include \"water\", so one species is possible.\nDid you check the location?\n\nBest,\nLukas",
      "votes": null
    },
    {
      "id": "3186203",
      "postDate": "04/24/2025 10:50:50",
      "content": "<p>Thanks for the replies!</p>\n<p>Certainly, PA data collection over large area is a time and effort -consuming task. However, I cannot agree with you that it is inherently noisy, but maybe I do not fully understand the meaning that you what you implied but this term. In my opinion, what makes analysis/modeling/prediction of species communities challenging indeed is when there is a large variation in surveying methodologies (including surveyed area), and proficiency levels of the surveying experts across the surveys included to the analysis. Maybe such feat matches your definition of noise, but iI would rather call it inevitable heterogeneity once many data sources are aggregated to a common database like EVA. Generally, the approaches of doing 1m2 survey and 400m2 survey (or even 8000m2) do differ considerably in practice.</p>\n<blockquote>\n  <p>Did you check the location?</p>\n</blockquote>\n<p>Unfortunately, the closest locations of those monospecies surveys are in Denmark, which 1000km away from my place, so it seems like an overkill to go and check in person. On the other hand, Google Street View helped me with getting an overview of a couple locations, here are some screenshots from that virtual trip. Both places do not look very monocultural to me, but clearly this inspection provides quite limited information.</p>\n<p>1) surveyId=3531166 <br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F126148%2F38c5b8a64dde6b72575292afb01536bf%2FScreenshot%202025-04-24%20at%2013.39.08.jpg?generation=1745491344109777&amp;alt=media\" alt=\"\"></p>\n<p>2) surveyId=1962148<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F126148%2F15078cc8cf754787f7f63ec56957e04f%2FScreenshot%202025-04-24%20at%2013.39.35.jpg?generation=1745491357426841&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "Thanks for the replies!\n\nCertainly, PA data collection over large area is a time and effort -consuming task. However, I cannot agree with you that it is inherently noisy, but maybe I do not fully understand the meaning that you what you implied but this term. In my opinion, what makes analysis/modeling/prediction of species communities challenging indeed is when there is a large variation in surveying methodologies (including surveyed area), and proficiency levels of the surveying experts across the surveys included to the analysis. Maybe such feat matches your definition of noise, but iI would rather call it inevitable heterogeneity once many data sources are aggregated to a common database like EVA. Generally, the approaches of doing 1m2 survey and 400m2 survey (or even 8000m2) do differ considerably in practice.\n\n>Did you check the location?\n\nUnfortunately, the closest locations of those monospecies surveys are in Denmark, which 1000km away from my place, so it seems like an overkill to go and check in person. On the other hand, Google Street View helped me with getting an overview of a couple locations, here are some screenshots from that virtual trip. Both places do not look very monocultural to me, but clearly this inspection provides quite limited information.\n\n1) surveyId=3531166 \n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F126148%2F38c5b8a64dde6b72575292afb01536bf%2FScreenshot%202025-04-24%20at%2013.39.08.jpg?generation=1745491344109777&alt=media)\n\n2) surveyId=1962148\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F126148%2F15078cc8cf754787f7f63ec56957e04f%2FScreenshot%202025-04-24%20at%2013.39.35.jpg?generation=1745491357426841&alt=media)",
      "votes": null
    },
    {
      "id": "3187108",
      "postDate": "04/25/2025 15:21:19",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/gtikho\" target=\"_blank\">@gtikho</a> ,</p>\n<p>You are absolutely right that what I call \"noise\" stems from inevitable heterogeneity, whether in methodology, observer proficiency, or survey context. I also agree that 1m2 and 8000m2 surveys are fundamentally different in approach and expectations, and combining such data under a unified structure introduces challenges.</p>\n<p>By the way, the Google Street View check (while they're limited) already shows that there is more than one species, but except for this one, there is no other way to really check if the data is true PA. I discussed PA data \"accuracy\" with an ecologist, and he kind of laughed at me and told me that PA data are never a true PA, not even in 1m2, since it is nearly impossible. Still, even with these issues, the data available through EVA is probably the most precise and large-scale resource that exists. </p>\n<p>Thanks again for sharing your thoughts so openly. Such discussions help everyone better understand the data, which is one of the reasons why we run competitions like this.</p>\n<p>Best regards,<br>\nLukas</p>",
      "rawMarkdown": "Hi @gtikho ,\n\nYou are absolutely right that what I call \"noise\" stems from inevitable heterogeneity, whether in methodology, observer proficiency, or survey context. I also agree that 1m2 and 8000m2 surveys are fundamentally different in approach and expectations, and combining such data under a unified structure introduces challenges.\n\nBy the way, the Google Street View check (while they're limited) already shows that there is more than one species, but except for this one, there is no other way to really check if the data is true PA. I discussed PA data \"accuracy\" with an ecologist, and he kind of laughed at me and told me that PA data are never a true PA, not even in 1m2, since it is nearly impossible. Still, even with these issues, the data available through EVA is probably the most precise and large-scale resource that exists. \n\nThanks again for sharing your thoughts so openly. Such discussions help everyone better understand the data, which is one of the reasons why we run competitions like this.\n\nBest regards,\nLukas",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3185124,
      "author_name": "picekl",
      "author_url": "",
      "post_date": "04/22/2025 21:39:23",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/gtikho\" target=\"_blank\">@gtikho</a>,</p>\n<p>In general, the collection of the Presence-Absence data is pretty tedious, and from my non-botanist perspective, noisy.<br>\nTherefore, the answers to your questions are pretty simple. Hope it helps.</p>\n<blockquote>\n  <p>There are some surveyId in the PA data, where there speciesId are duplicated. For example surveyId=128364 Do these duplicates have any special meaning?</p>\n</blockquote>\n<p>There is no particular meaning. Probably just a noise. You can remove it.</p>\n<blockquote>\n  <p>There are some surveyId in the PA data, where areaInM2 takes multiple values. For example, surveyId=2756037. Another example is surveyId=789799, where some values of areaInM2 are -inf. Could you please clarify whether it is possible to deduce the actual area of the survey from these?</p>\n</blockquote>\n<p>If there are multiple values and one is a real number, I would suggest using it. Usually, if there is -inf or zero, the data are missing.</p>\n<blockquote>\n  <p>What are the safe assumptions to make regarding the consistency of different PA surveys? Especially given that there is so huge variation in area of surveys: from 0.04m2 to 8000m2. Shall we expect that if a species is not listed for some surveyId in the PA data, then it was actually missing in that surveyId because it was not present (or at least it was not detected by survey team)? I have some experience in conducting vegetation surveys in the field, and it is simply very difficult to believe that some 700m2 survey there was only one species found, like in surveyId=168111.</p>\n</blockquote>\n<p>Usually, if the value is filled, it is more or less accurate. The large areas could include \"water\", so one species is possible.<br>\nDid you check the location?</p>\n<p>Best,<br>\nLukas</p>",
      "votes": null,
      "replies": [
        {
          "id": 3186203,
          "author_name": "gtikho",
          "author_url": "",
          "post_date": "04/24/2025 10:50:50",
          "content": "<p>Thanks for the replies!</p>\n<p>Certainly, PA data collection over large area is a time and effort -consuming task. However, I cannot agree with you that it is inherently noisy, but maybe I do not fully understand the meaning that you what you implied but this term. In my opinion, what makes analysis/modeling/prediction of species communities challenging indeed is when there is a large variation in surveying methodologies (including surveyed area), and proficiency levels of the surveying experts across the surveys included to the analysis. Maybe such feat matches your definition of noise, but iI would rather call it inevitable heterogeneity once many data sources are aggregated to a common database like EVA. Generally, the approaches of doing 1m2 survey and 400m2 survey (or even 8000m2) do differ considerably in practice.</p>\n<blockquote>\n  <p>Did you check the location?</p>\n</blockquote>\n<p>Unfortunately, the closest locations of those monospecies surveys are in Denmark, which 1000km away from my place, so it seems like an overkill to go and check in person. On the other hand, Google Street View helped me with getting an overview of a couple locations, here are some screenshots from that virtual trip. Both places do not look very monocultural to me, but clearly this inspection provides quite limited information.</p>\n<p>1) surveyId=3531166 <br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F126148%2F38c5b8a64dde6b72575292afb01536bf%2FScreenshot%202025-04-24%20at%2013.39.08.jpg?generation=1745491344109777&amp;alt=media\" alt=\"\"></p>\n<p>2) surveyId=1962148<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F126148%2F15078cc8cf754787f7f63ec56957e04f%2FScreenshot%202025-04-24%20at%2013.39.35.jpg?generation=1745491357426841&amp;alt=media\" alt=\"\"></p>",
          "votes": null,
          "replies": [
            {
              "id": 3187108,
              "author_name": "picekl",
              "author_url": "",
              "post_date": "04/25/2025 15:21:19",
              "content": "<p>Hi <a href=\"https://www.kaggle.com/gtikho\" target=\"_blank\">@gtikho</a> ,</p>\n<p>You are absolutely right that what I call \"noise\" stems from inevitable heterogeneity, whether in methodology, observer proficiency, or survey context. I also agree that 1m2 and 8000m2 surveys are fundamentally different in approach and expectations, and combining such data under a unified structure introduces challenges.</p>\n<p>By the way, the Google Street View check (while they're limited) already shows that there is more than one species, but except for this one, there is no other way to really check if the data is true PA. I discussed PA data \"accuracy\" with an ecologist, and he kind of laughed at me and told me that PA data are never a true PA, not even in 1m2, since it is nearly impossible. Still, even with these issues, the data available through EVA is probably the most precise and large-scale resource that exists. </p>\n<p>Thanks again for sharing your thoughts so openly. Such discussions help everyone better understand the data, which is one of the reasons why we run competitions like this.</p>\n<p>Best regards,<br>\nLukas</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3184862": "Hi @picekl \n\nI would like to ask several questions on some aspects of the training data that I simply do not understand. If you can provide clarifications for those, I will be very grateful, but if you consider that answering these will disclose some internal information, please just tell me that you cannot discuss this matter.\n\n1. There are some `surveyId` in the PA data, where there `speciesId` are duplicated. For example `surveyId=128364` Do these duplicates have any special meaning?\n\n2. There are some `surveyId` in the PA data, where `areaInM2` takes multiple values. For example, `surveyId=2756037`. Another example is `surveyId=789799`, where some values of `areaInM2` are `-inf`. Could you please clarify whether it is possible to deduce the actual area of the survey from these?\n\n3. What are the safe assumptions to make regarding the consistency of different PA surveys? Especially given that there is so huge variation in area of surveys: from 0.04m2 to 8000m2. Shall we expect that if a species is not listed for some `surveyId` in the PA data, then it was actually missing in that `surveyId` because it was not present (or at least it was not detected by survey team)? I have some experience in conducting vegetation surveys in the field, and it is simply very difficult to believe that some 700m2 survey there was only one species found, like in `surveyId=168111`.\n\nThank you in advance for your time and hope to hear back from you soon!",
    "3185124": "Hi @gtikho,\n\nIn general, the collection of the Presence-Absence data is pretty tedious, and from my non-botanist perspective, noisy.\nTherefore, the answers to your questions are pretty simple. Hope it helps.\n\n> There are some surveyId in the PA data, where there speciesId are duplicated. For example surveyId=128364 Do these duplicates have any special meaning?\n\nThere is no particular meaning. Probably just a noise. You can remove it.\n\n\n> There are some surveyId in the PA data, where areaInM2 takes multiple values. For example, surveyId=2756037. Another example is surveyId=789799, where some values of areaInM2 are -inf. Could you please clarify whether it is possible to deduce the actual area of the survey from these?\n\nIf there are multiple values and one is a real number, I would suggest using it. Usually, if there is -inf or zero, the data are missing.\n\n\n> What are the safe assumptions to make regarding the consistency of different PA surveys? Especially given that there is so huge variation in area of surveys: from 0.04m2 to 8000m2. Shall we expect that if a species is not listed for some surveyId in the PA data, then it was actually missing in that surveyId because it was not present (or at least it was not detected by survey team)? I have some experience in conducting vegetation surveys in the field, and it is simply very difficult to believe that some 700m2 survey there was only one species found, like in surveyId=168111.\n\nUsually, if the value is filled, it is more or less accurate. The large areas could include \"water\", so one species is possible.\nDid you check the location?\n\nBest,\nLukas",
    "3186203": "Thanks for the replies!\n\nCertainly, PA data collection over large area is a time and effort -consuming task. However, I cannot agree with you that it is inherently noisy, but maybe I do not fully understand the meaning that you what you implied but this term. In my opinion, what makes analysis/modeling/prediction of species communities challenging indeed is when there is a large variation in surveying methodologies (including surveyed area), and proficiency levels of the surveying experts across the surveys included to the analysis. Maybe such feat matches your definition of noise, but iI would rather call it inevitable heterogeneity once many data sources are aggregated to a common database like EVA. Generally, the approaches of doing 1m2 survey and 400m2 survey (or even 8000m2) do differ considerably in practice.\n\n>Did you check the location?\n\nUnfortunately, the closest locations of those monospecies surveys are in Denmark, which 1000km away from my place, so it seems like an overkill to go and check in person. On the other hand, Google Street View helped me with getting an overview of a couple locations, here are some screenshots from that virtual trip. Both places do not look very monocultural to me, but clearly this inspection provides quite limited information.\n\n1) surveyId=3531166 \n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F126148%2F38c5b8a64dde6b72575292afb01536bf%2FScreenshot%202025-04-24%20at%2013.39.08.jpg?generation=1745491344109777&alt=media)\n\n2) surveyId=1962148\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F126148%2F15078cc8cf754787f7f63ec56957e04f%2FScreenshot%202025-04-24%20at%2013.39.35.jpg?generation=1745491357426841&alt=media)",
    "3187108": "Hi @gtikho ,\n\nYou are absolutely right that what I call \"noise\" stems from inevitable heterogeneity, whether in methodology, observer proficiency, or survey context. I also agree that 1m2 and 8000m2 surveys are fundamentally different in approach and expectations, and combining such data under a unified structure introduces challenges.\n\nBy the way, the Google Street View check (while they're limited) already shows that there is more than one species, but except for this one, there is no other way to really check if the data is true PA. I discussed PA data \"accuracy\" with an ecologist, and he kind of laughed at me and told me that PA data are never a true PA, not even in 1m2, since it is nearly impossible. Still, even with these issues, the data available through EVA is probably the most precise and large-scale resource that exists. \n\nThanks again for sharing your thoughts so openly. Such discussions help everyone better understand the data, which is one of the reasons why we run competitions like this.\n\nBest regards,\nLukas"
  },
  "source": "meta"
}