{
  "id": 276830,
  "title": "Test classes outside of 81k classes from train data",
  "url": "/competitions/landmark-recognition-2021/discussion/276830",
  "author_name": "",
  "post_date": "2021-10-06T14:49:30.708037600Z",
  "votes": 9,
  "comment_count": 3,
  "views": 0,
  "content": "<p>During this competition, I was sure that the test set could contain only classes from the public train set (81k classes). But after reading some solutions, I understood that the private test set included landmarks from the full GLDv2 dataset (203k classes) not presented in the cleaned version (81k classes).</p>\n<p>It was raised in some discussions during the competitions, but I don't see any official confirmation from organizers and unfortunately missed this fact. But I have reread competition rules and descriptions, and I think that the existence of images with classes outside 81k classes from the training dataset violates the competition's official rules.</p>\n<p>So, what I have found in the official competition description:</p>\n<ol>\n<li>\"The training data for this competition comes from a <strong>cleaned version</strong> of the Google Landmarks Dataset v2 (GLDv2)\" – cleaned version of GLDv2 is directly called as a source of train data for this competition.</li>\n<li>\"the private training set contains only a 100k subset of the <strong>total public training set</strong>\" – Private training set is a subset of the training set it means it can't contain any landmark other than 81k from the cleaned GLDv2</li>\n<li>\"This 100k subset contains all of the training set images associated with the landmarks in the private test set\" – it means that there can not be classes in the test set outside of the private train classes → they also should be a subset of 81k classes from a public train set.</li>\n</ol>\n<p>I guess that organizers in #2 by \"total public training set\" wanted to say \"full GLDv2 dataset\", but it is not what is written in the description. </p>\n<p>\"Full GLDv2\" is never called \"train,\" \"full train,\" \"total train,\" or somehow else anywhere in the competition description. Moreover, the full GLDv2 dataset is never mentioned in the description at all, as well as the existence of 203k classes.</p>\n<p>And on the main competition page, it is said that \"there are more than 81K classes in this challenge\" – formally 203k &gt; 81k, but I think that it meant that \"classes\" in this challenge – are classes from the cleaned dataset.</p>\n<p>These are all mentions of datasets in the competition description and rules that I have managed to find. Please point out if I have missed something.</p>\n<p>The official paper about test dataset construction doesn't add any clarity as well. It neither says that test classes come from 81k nor that they contain anything outside these classes.</p>\n<p>So, if we formally read the official description of this competition – the only training dataset mentioned there (and also a public training set) is the cleaned version of GLDv2, which has 81k classes. According to the description, only these classes can be present in the test dataset. And actual test dataset directly violates the competition description. </p>\n<p>Can you please comment on it and share what part of the test set had labels outside 81k to estimate the significance of this issue?</p>\n<p>Overall I just want to point on this problem and stress the importance of clean, straightforward rules that don't allow any misinterpretation (it seems I am not the only one who expected only 81k classes in the test).</p>",
  "messages": [
    {
      "id": "1536247",
      "postDate": "10/06/2021 14:49:30",
      "content": "<p>During this competition, I was sure that the test set could contain only classes from the public train set (81k classes). But after reading some solutions, I understood that the private test set included landmarks from the full GLDv2 dataset (203k classes) not presented in the cleaned version (81k classes).</p>\n<p>It was raised in some discussions during the competitions, but I don't see any official confirmation from organizers and unfortunately missed this fact. But I have reread competition rules and descriptions, and I think that the existence of images with classes outside 81k classes from the training dataset violates the competition's official rules.</p>\n<p>So, what I have found in the official competition description:</p>\n<ol>\n<li>\"The training data for this competition comes from a <strong>cleaned version</strong> of the Google Landmarks Dataset v2 (GLDv2)\" – cleaned version of GLDv2 is directly called as a source of train data for this competition.</li>\n<li>\"the private training set contains only a 100k subset of the <strong>total public training set</strong>\" – Private training set is a subset of the training set it means it can't contain any landmark other than 81k from the cleaned GLDv2</li>\n<li>\"This 100k subset contains all of the training set images associated with the landmarks in the private test set\" – it means that there can not be classes in the test set outside of the private train classes → they also should be a subset of 81k classes from a public train set.</li>\n</ol>\n<p>I guess that organizers in #2 by \"total public training set\" wanted to say \"full GLDv2 dataset\", but it is not what is written in the description. </p>\n<p>\"Full GLDv2\" is never called \"train,\" \"full train,\" \"total train,\" or somehow else anywhere in the competition description. Moreover, the full GLDv2 dataset is never mentioned in the description at all, as well as the existence of 203k classes.</p>\n<p>And on the main competition page, it is said that \"there are more than 81K classes in this challenge\" – formally 203k &gt; 81k, but I think that it meant that \"classes\" in this challenge – are classes from the cleaned dataset.</p>\n<p>These are all mentions of datasets in the competition description and rules that I have managed to find. Please point out if I have missed something.</p>\n<p>The official paper about test dataset construction doesn't add any clarity as well. It neither says that test classes come from 81k nor that they contain anything outside these classes.</p>\n<p>So, if we formally read the official description of this competition – the only training dataset mentioned there (and also a public training set) is the cleaned version of GLDv2, which has 81k classes. According to the description, only these classes can be present in the test dataset. And actual test dataset directly violates the competition description. </p>\n<p>Can you please comment on it and share what part of the test set had labels outside 81k to estimate the significance of this issue?</p>\n<p>Overall I just want to point on this problem and stress the importance of clean, straightforward rules that don't allow any misinterpretation (it seems I am not the only one who expected only 81k classes in the test).</p>",
      "rawMarkdown": "During this competition, I was sure that the test set could contain only classes from the public train set (81k classes). But after reading some solutions, I understood that the private test set included landmarks from the full GLDv2 dataset (203k classes) not presented in the cleaned version (81k classes).\n\nIt was raised in some discussions during the competitions, but I don't see any official confirmation from organizers and unfortunately missed this fact. But I have reread competition rules and descriptions, and I think that the existence of images with classes outside 81k classes from the training dataset violates the competition's official rules.\n\nSo, what I have found in the official competition description:\n\n1. \"The training data for this competition comes from a **cleaned version** of the Google Landmarks Dataset v2 (GLDv2)\" – cleaned version of GLDv2 is directly called as a source of train data for this competition.\n2. \"the private training set contains only a 100k subset of the **total public training set**\" – Private training set is a subset of the training set it means it can't contain any landmark other than 81k from the cleaned GLDv2\n3. \"This 100k subset contains all of the training set images associated with the landmarks in the private test set\" – it means that there can not be classes in the test set outside of the private train classes → they also should be a subset of 81k classes from a public train set.\n\nI guess that organizers in #2 by \"total public training set\" wanted to say \"full GLDv2 dataset\", but it is not what is written in the description. \n\n\"Full GLDv2\" is never called \"train,\" \"full train,\" \"total train,\" or somehow else anywhere in the competition description. Moreover, the full GLDv2 dataset is never mentioned in the description at all, as well as the existence of 203k classes.\n\nAnd on the main competition page, it is said that \"there are more than 81K classes in this challenge\" – formally 203k > 81k, but I think that it meant that \"classes\" in this challenge – are classes from the cleaned dataset.\n\nThese are all mentions of datasets in the competition description and rules that I have managed to find. Please point out if I have missed something.\n\nThe official paper about test dataset construction doesn't add any clarity as well. It neither says that test classes come from 81k nor that they contain anything outside these classes.\n\nSo, if we formally read the official description of this competition – the only training dataset mentioned there (and also a public training set) is the cleaned version of GLDv2, which has 81k classes. According to the description, only these classes can be present in the test dataset. And actual test dataset directly violates the competition description. \n\nCan you please comment on it and share what part of the test set had labels outside 81k to estimate the significance of this issue?\n\nOverall I just want to point on this problem and stress the importance of clean, straightforward rules that don't allow any misinterpretation (it seems I am not the only one who expected only 81k classes in the test).",
      "votes": null
    },
    {
      "id": "1537471",
      "postDate": "10/07/2021 14:01:07",
      "content": "<p>Just a couple of thoughts to add some actionable ideas for the next competitions. </p>\n<p>The most straightforward description is usually the most transparent one. Based on this particular case:</p>\n<ol>\n<li>It is helpful to list requirements for expected submission:<br>\nFor classification – expected number of classes and their list if it is possible<br>\nFor regression – expected range of values (that can be from -inf to +inf)<br>\nFor anything else – competition-specific requirements that clarify the expected answer.</li>\n</ol>\n<p>Sample submission is helpful but does not always answer these questions.<br>\nIt can simplify competitors' lives, especially in kernel competitions (where many people spend the first few (dozens) submissions to understand what is expected by organizers).</p>\n<ol>\n<li>It will be very convenient if kaggle will publish the whole datasets (that is considered the full train dataset) as an official competition dataset. In this case, you could publish the entire GLDv2 as a training dataset + two different train.csv – for all images and the cleaned version only. It would also simplify the process and clarify the rules.</li>\n</ol>\n<p>If the whole dataset is considered too big, publishing the cleaned version as a separate second dataset is possible. If there are some issues related to data protection by competition rule acceptance (and it is the reason not to publish two datasets), it doesn't work anyway. In any competition with extensive datasets, someone creates a public dataset with compressed data that is not protected by rule acceptance.</p>",
      "rawMarkdown": "Just a couple of thoughts to add some actionable ideas for the next competitions. \n\nThe most straightforward description is usually the most transparent one. Based on this particular case:\n1. It is helpful to list requirements for expected submission:\n\tFor classification – expected number of classes and their list if it is possible\n\tFor regression – expected range of values (that can be from -inf to +inf)\n\tFor anything else – competition-specific requirements that clarify the expected answer.\n\nSample submission is helpful but does not always answer these questions.\nIt can simplify competitors' lives, especially in kernel competitions (where many people spend the first few (dozens) submissions to understand what is expected by organizers).\n\n2. It will be very convenient if kaggle will publish the whole datasets (that is considered the full train dataset) as an official competition dataset. In this case, you could publish the entire GLDv2 as a training dataset + two different train.csv – for all images and the cleaned version only. It would also simplify the process and clarify the rules.\n\nIf the whole dataset is considered too big, publishing the cleaned version as a separate second dataset is possible. If there are some issues related to data protection by competition rule acceptance (and it is the reason not to publish two datasets), it doesn't work anyway. In any competition with extensive datasets, someone creates a public dataset with compressed data that is not protected by rule acceptance.",
      "votes": null
    },
    {
      "id": "1537584",
      "postDate": "10/07/2021 15:42:04",
      "content": "<p>From the competition related paper ( <a href=\"https://arxiv.org/abs/2108.08874\" target=\"_blank\">https://arxiv.org/abs/2108.08874</a> )</p>\n<blockquote>\n  <p>Index dataset (recognition challenge): 100,000 images sampled from the GLDv2 training dataset.</p>\n</blockquote>\n<p>thats how I found out. </p>",
      "rawMarkdown": "From the competition related paper ( https://arxiv.org/abs/2108.08874 )\n\n> Index dataset (recognition challenge): 100,000 images sampled from the GLDv2 training dataset.\n\nthats how I found out.",
      "votes": null
    },
    {
      "id": "1538016",
      "postDate": "10/08/2021 01:59:42",
      "content": "<p>Yes. Thanks. But I consider it as an inconclusive statement.<br>\nFormally subsample of cleaned GLDv2 is also a subset of full GLDv2 - so it doesn't contradict \"81k classes theory\".<br>\nSide note: it is called \"index dataset\" in the paper, but in competition, we are talking about \"train dataset\" - so again, it requires some assumptions about what is described here.</p>\n<p>What we have:<br>\nCompetition description that is focused on 81k dataset only + secondary information source that has information that doesn't contradict the idea of 81k classes but can be also be interpreted as 203k classes.</p>\n<p>I don't think it is clean enough for the critical competition rule.</p>\n<p>What could be done better here:</p>\n<ul>\n<li>Say that private test is a subset of full GLDv2 everywhere (so there can be misunderstanding at all)</li>\n<li>Directly say how many (possible) classes are the test and how they are related to 81k and 203k from GLDv2</li>\n</ul>",
      "rawMarkdown": "Yes. Thanks. But I consider it as an inconclusive statement.\nFormally subsample of cleaned GLDv2 is also a subset of full GLDv2 - so it doesn't contradict \"81k classes theory\".\nSide note: it is called \"index dataset\" in the paper, but in competition, we are talking about \"train dataset\" - so again, it requires some assumptions about what is described here.\n\nWhat we have:\nCompetition description that is focused on 81k dataset only + secondary information source that has information that doesn't contradict the idea of 81k classes but can be also be interpreted as 203k classes.\n\nI don't think it is clean enough for the critical competition rule.\n\nWhat could be done better here:\n- Say that private test is a subset of full GLDv2 everywhere (so there can be misunderstanding at all)\n- Directly say how many (possible) classes are the test and how they are related to 81k and 203k from GLDv2",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1537471,
      "author_name": "ilialar",
      "author_url": "",
      "post_date": "10/07/2021 14:01:07",
      "content": "<p>Just a couple of thoughts to add some actionable ideas for the next competitions. </p>\n<p>The most straightforward description is usually the most transparent one. Based on this particular case:</p>\n<ol>\n<li>It is helpful to list requirements for expected submission:<br>\nFor classification – expected number of classes and their list if it is possible<br>\nFor regression – expected range of values (that can be from -inf to +inf)<br>\nFor anything else – competition-specific requirements that clarify the expected answer.</li>\n</ol>\n<p>Sample submission is helpful but does not always answer these questions.<br>\nIt can simplify competitors' lives, especially in kernel competitions (where many people spend the first few (dozens) submissions to understand what is expected by organizers).</p>\n<ol>\n<li>It will be very convenient if kaggle will publish the whole datasets (that is considered the full train dataset) as an official competition dataset. In this case, you could publish the entire GLDv2 as a training dataset + two different train.csv – for all images and the cleaned version only. It would also simplify the process and clarify the rules.</li>\n</ol>\n<p>If the whole dataset is considered too big, publishing the cleaned version as a separate second dataset is possible. If there are some issues related to data protection by competition rule acceptance (and it is the reason not to publish two datasets), it doesn't work anyway. In any competition with extensive datasets, someone creates a public dataset with compressed data that is not protected by rule acceptance.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1537584,
      "author_name": "christofhenkel",
      "author_url": "",
      "post_date": "10/07/2021 15:42:04",
      "content": "<p>From the competition related paper ( <a href=\"https://arxiv.org/abs/2108.08874\" target=\"_blank\">https://arxiv.org/abs/2108.08874</a> )</p>\n<blockquote>\n  <p>Index dataset (recognition challenge): 100,000 images sampled from the GLDv2 training dataset.</p>\n</blockquote>\n<p>thats how I found out. </p>",
      "votes": null,
      "replies": [
        {
          "id": 1538016,
          "author_name": "ilialar",
          "author_url": "",
          "post_date": "10/08/2021 01:59:42",
          "content": "<p>Yes. Thanks. But I consider it as an inconclusive statement.<br>\nFormally subsample of cleaned GLDv2 is also a subset of full GLDv2 - so it doesn't contradict \"81k classes theory\".<br>\nSide note: it is called \"index dataset\" in the paper, but in competition, we are talking about \"train dataset\" - so again, it requires some assumptions about what is described here.</p>\n<p>What we have:<br>\nCompetition description that is focused on 81k dataset only + secondary information source that has information that doesn't contradict the idea of 81k classes but can be also be interpreted as 203k classes.</p>\n<p>I don't think it is clean enough for the critical competition rule.</p>\n<p>What could be done better here:</p>\n<ul>\n<li>Say that private test is a subset of full GLDv2 everywhere (so there can be misunderstanding at all)</li>\n<li>Directly say how many (possible) classes are the test and how they are related to 81k and 203k from GLDv2</li>\n</ul>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1536247": "During this competition, I was sure that the test set could contain only classes from the public train set (81k classes). But after reading some solutions, I understood that the private test set included landmarks from the full GLDv2 dataset (203k classes) not presented in the cleaned version (81k classes).\n\nIt was raised in some discussions during the competitions, but I don't see any official confirmation from organizers and unfortunately missed this fact. But I have reread competition rules and descriptions, and I think that the existence of images with classes outside 81k classes from the training dataset violates the competition's official rules.\n\nSo, what I have found in the official competition description:\n\n1. \"The training data for this competition comes from a **cleaned version** of the Google Landmarks Dataset v2 (GLDv2)\" – cleaned version of GLDv2 is directly called as a source of train data for this competition.\n2. \"the private training set contains only a 100k subset of the **total public training set**\" – Private training set is a subset of the training set it means it can't contain any landmark other than 81k from the cleaned GLDv2\n3. \"This 100k subset contains all of the training set images associated with the landmarks in the private test set\" – it means that there can not be classes in the test set outside of the private train classes → they also should be a subset of 81k classes from a public train set.\n\nI guess that organizers in #2 by \"total public training set\" wanted to say \"full GLDv2 dataset\", but it is not what is written in the description. \n\n\"Full GLDv2\" is never called \"train,\" \"full train,\" \"total train,\" or somehow else anywhere in the competition description. Moreover, the full GLDv2 dataset is never mentioned in the description at all, as well as the existence of 203k classes.\n\nAnd on the main competition page, it is said that \"there are more than 81K classes in this challenge\" – formally 203k > 81k, but I think that it meant that \"classes\" in this challenge – are classes from the cleaned dataset.\n\nThese are all mentions of datasets in the competition description and rules that I have managed to find. Please point out if I have missed something.\n\nThe official paper about test dataset construction doesn't add any clarity as well. It neither says that test classes come from 81k nor that they contain anything outside these classes.\n\nSo, if we formally read the official description of this competition – the only training dataset mentioned there (and also a public training set) is the cleaned version of GLDv2, which has 81k classes. According to the description, only these classes can be present in the test dataset. And actual test dataset directly violates the competition description. \n\nCan you please comment on it and share what part of the test set had labels outside 81k to estimate the significance of this issue?\n\nOverall I just want to point on this problem and stress the importance of clean, straightforward rules that don't allow any misinterpretation (it seems I am not the only one who expected only 81k classes in the test).",
    "1537471": "Just a couple of thoughts to add some actionable ideas for the next competitions. \n\nThe most straightforward description is usually the most transparent one. Based on this particular case:\n1. It is helpful to list requirements for expected submission:\n\tFor classification – expected number of classes and their list if it is possible\n\tFor regression – expected range of values (that can be from -inf to +inf)\n\tFor anything else – competition-specific requirements that clarify the expected answer.\n\nSample submission is helpful but does not always answer these questions.\nIt can simplify competitors' lives, especially in kernel competitions (where many people spend the first few (dozens) submissions to understand what is expected by organizers).\n\n2. It will be very convenient if kaggle will publish the whole datasets (that is considered the full train dataset) as an official competition dataset. In this case, you could publish the entire GLDv2 as a training dataset + two different train.csv – for all images and the cleaned version only. It would also simplify the process and clarify the rules.\n\nIf the whole dataset is considered too big, publishing the cleaned version as a separate second dataset is possible. If there are some issues related to data protection by competition rule acceptance (and it is the reason not to publish two datasets), it doesn't work anyway. In any competition with extensive datasets, someone creates a public dataset with compressed data that is not protected by rule acceptance.",
    "1537584": "From the competition related paper ( https://arxiv.org/abs/2108.08874 )\n\n> Index dataset (recognition challenge): 100,000 images sampled from the GLDv2 training dataset.\n\nthats how I found out.",
    "1538016": "Yes. Thanks. But I consider it as an inconclusive statement.\nFormally subsample of cleaned GLDv2 is also a subset of full GLDv2 - so it doesn't contradict \"81k classes theory\".\nSide note: it is called \"index dataset\" in the paper, but in competition, we are talking about \"train dataset\" - so again, it requires some assumptions about what is described here.\n\nWhat we have:\nCompetition description that is focused on 81k dataset only + secondary information source that has information that doesn't contradict the idea of 81k classes but can be also be interpreted as 203k classes.\n\nI don't think it is clean enough for the critical competition rule.\n\nWhat could be done better here:\n- Say that private test is a subset of full GLDv2 everywhere (so there can be misunderstanding at all)\n- Directly say how many (possible) classes are the test and how they are related to 81k and 203k from GLDv2"
  },
  "source": "meta"
}