{
  "id": 574992,
  "title": "Artefacts in GLC25-PA-test-landcover?",
  "url": "/competitions/geolifeclef-2025/discussion/574992",
  "author_name": "",
  "post_date": "2025-04-25T09:15:39.759978400Z",
  "votes": 2,
  "comment_count": 1,
  "views": 0,
  "content": "<p>Hi <a href=\"https://www.kaggle.com/picekl\" target=\"_blank\">@picekl</a> </p>\n<p>Sorry for keeping pondering you with questions on data quality!</p>\n<p>I have been exploring and trying to adapt the non-satellite modalities that you provide as the competition data. In particular, lately I have been focused on the LandCover data. However, I ran into one annoying issue, which I believe is caused by an artefact in one of the provided datasets.</p>\n<p>Specifically, all column values of <code>GLC25-PA-test-landcover.csv</code> file with <code>surveyId&gt;=5000000</code> are identical. I strongly believe that this is corrupted data and not the correct one. Given that there are more than 10000 of such surveys (68% of test dataset), it makes robust usage of these data challenging.</p>\n<p>Could you provide a comment on this concern of mine, please? If you agree with my claim that these data have been corrupted, shall we expect that there will be an update of this dataset with corrected data?</p>\n<p>In this context, I would also like to clarify what is the organizers' position regarding using publicly available maps/raster data, which is not provided as part of the competition. I have read the relevant paragraph in the Rules, but it sounds a bit vague, and the answer to a previous question about latin names of species was rather implying that it may be possible to use some side data, though it will not be provided by the organizing team. Thus, it would be great to hear your explaination on this matter firsthand.</p>",
  "messages": [
    {
      "id": "3186880",
      "postDate": "04/25/2025 09:15:39",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/picekl\" target=\"_blank\">@picekl</a> </p>\n<p>Sorry for keeping pondering you with questions on data quality!</p>\n<p>I have been exploring and trying to adapt the non-satellite modalities that you provide as the competition data. In particular, lately I have been focused on the LandCover data. However, I ran into one annoying issue, which I believe is caused by an artefact in one of the provided datasets.</p>\n<p>Specifically, all column values of <code>GLC25-PA-test-landcover.csv</code> file with <code>surveyId&gt;=5000000</code> are identical. I strongly believe that this is corrupted data and not the correct one. Given that there are more than 10000 of such surveys (68% of test dataset), it makes robust usage of these data challenging.</p>\n<p>Could you provide a comment on this concern of mine, please? If you agree with my claim that these data have been corrupted, shall we expect that there will be an update of this dataset with corrected data?</p>\n<p>In this context, I would also like to clarify what is the organizers' position regarding using publicly available maps/raster data, which is not provided as part of the competition. I have read the relevant paragraph in the Rules, but it sounds a bit vague, and the answer to a previous question about latin names of species was rather implying that it may be possible to use some side data, though it will not be provided by the organizing team. Thus, it would be great to hear your explaination on this matter firsthand.</p>",
      "rawMarkdown": "Hi @picekl \n\nSorry for keeping pondering you with questions on data quality!\n\nI have been exploring and trying to adapt the non-satellite modalities that you provide as the competition data. In particular, lately I have been focused on the LandCover data. However, I ran into one annoying issue, which I believe is caused by an artefact in one of the provided datasets.\n\nSpecifically, all column values of `GLC25-PA-test-landcover.csv` file with `surveyId>=5000000` are identical. I strongly believe that this is corrupted data and not the correct one. Given that there are more than 10000 of such surveys (68% of test dataset), it makes robust usage of these data challenging.\n\nCould you provide a comment on this concern of mine, please? If you agree with my claim that these data have been corrupted, shall we expect that there will be an update of this dataset with corrected data?\n\nIn this context, I would also like to clarify what is the organizers' position regarding using publicly available maps/raster data, which is not provided as part of the competition. I have read the relevant paragraph in the Rules, but it sounds a bit vague, and the answer to a previous question about latin names of species was rather implying that it may be possible to use some side data, though it will not be provided by the organizing team. Thus, it would be great to hear your explaination on this matter firsthand.",
      "votes": null
    },
    {
      "id": "3187099",
      "postDate": "04/25/2025 15:06:59",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/gtikho\" target=\"_blank\">@gtikho</a>,</p>\n<p>Thank you for raising this. I checked the script, and it was fine, but the data was indeed not. I genuinely do not know what happened.<br>\nAnyway, I re-extracted the data and updated the dataset on Kaggle. It should already be available.</p>\n<p>About the use of available maps/rasters. We already provide many rasters on the <a href=\"https://lab.plantnet.org/seafile/d/979791609d0e4016a68a/\" target=\"_blank\">Seafile</a>. Here is a link to the <a href=\"EnvironmentalValues\" target=\"_blank\">Landcover one</a>. However, you can use any other raster if you find it valuable.</p>\n<p>Best,<br>\nLukas</p>",
      "rawMarkdown": "Hi @gtikho,\n\nThank you for raising this. I checked the script, and it was fine, but the data was indeed not. I genuinely do not know what happened.\nAnyway, I re-extracted the data and updated the dataset on Kaggle. It should already be available.\n\nAbout the use of available maps/rasters. We already provide many rasters on the [Seafile](https://lab.plantnet.org/seafile/d/979791609d0e4016a68a/). Here is a link to the [Landcover one](EnvironmentalValues). However, you can use any other raster if you find it valuable.\n\nBest,\nLukas",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3187099,
      "author_name": "picekl",
      "author_url": "",
      "post_date": "04/25/2025 15:06:59",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/gtikho\" target=\"_blank\">@gtikho</a>,</p>\n<p>Thank you for raising this. I checked the script, and it was fine, but the data was indeed not. I genuinely do not know what happened.<br>\nAnyway, I re-extracted the data and updated the dataset on Kaggle. It should already be available.</p>\n<p>About the use of available maps/rasters. We already provide many rasters on the <a href=\"https://lab.plantnet.org/seafile/d/979791609d0e4016a68a/\" target=\"_blank\">Seafile</a>. Here is a link to the <a href=\"EnvironmentalValues\" target=\"_blank\">Landcover one</a>. However, you can use any other raster if you find it valuable.</p>\n<p>Best,<br>\nLukas</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3186880": "Hi @picekl \n\nSorry for keeping pondering you with questions on data quality!\n\nI have been exploring and trying to adapt the non-satellite modalities that you provide as the competition data. In particular, lately I have been focused on the LandCover data. However, I ran into one annoying issue, which I believe is caused by an artefact in one of the provided datasets.\n\nSpecifically, all column values of `GLC25-PA-test-landcover.csv` file with `surveyId>=5000000` are identical. I strongly believe that this is corrupted data and not the correct one. Given that there are more than 10000 of such surveys (68% of test dataset), it makes robust usage of these data challenging.\n\nCould you provide a comment on this concern of mine, please? If you agree with my claim that these data have been corrupted, shall we expect that there will be an update of this dataset with corrected data?\n\nIn this context, I would also like to clarify what is the organizers' position regarding using publicly available maps/raster data, which is not provided as part of the competition. I have read the relevant paragraph in the Rules, but it sounds a bit vague, and the answer to a previous question about latin names of species was rather implying that it may be possible to use some side data, though it will not be provided by the organizing team. Thus, it would be great to hear your explaination on this matter firsthand.",
    "3187099": "Hi @gtikho,\n\nThank you for raising this. I checked the script, and it was fine, but the data was indeed not. I genuinely do not know what happened.\nAnyway, I re-extracted the data and updated the dataset on Kaggle. It should already be available.\n\nAbout the use of available maps/rasters. We already provide many rasters on the [Seafile](https://lab.plantnet.org/seafile/d/979791609d0e4016a68a/). Here is a link to the [Landcover one](EnvironmentalValues). However, you can use any other raster if you find it valuable.\n\nBest,\nLukas"
  },
  "source": "meta"
}