{
  "id": 179748,
  "title": "Question to organizers re: 100k data split",
  "url": "/competitions/landmark-recognition-2020/discussion/179748",
  "author_name": "",
  "post_date": "2020-09-02T18:17:53.374504700Z",
  "votes": 2,
  "comment_count": 3,
  "views": 0,
  "content": "<p>I noticed the data description says \"the private training set contains only a 100k subset of the total public training set.\". Does this mean that <em>all</em> of the train/ images during the rerun are present in the 1.5M train/ images I can see now?</p>",
  "messages": [
    {
      "id": "995732",
      "postDate": "09/02/2020 18:17:53",
      "content": "<p>I noticed the data description says \"the private training set contains only a 100k subset of the total public training set.\". Does this mean that <em>all</em> of the train/ images during the rerun are present in the 1.5M train/ images I can see now?</p>",
      "rawMarkdown": "I noticed the data description says \"the private training set contains only a 100k subset of the total public training set.\". Does this mean that *all* of the train/ images during the rerun are present in the 1.5M train/ images I can see now?",
      "votes": null
    },
    {
      "id": "995949",
      "postDate": "09/03/2020 01:24:50",
      "content": "<p>Just wanted to link related discussions here:</p>\n<ul>\n<li><a href=\"https://www.kaggle.com/c/landmark-recognition-2020/discussion/176697\" target=\"_blank\">Can we pre-compute embeddings? Are image ids the same in PB and PV sets?</a></li>\n<li><a href=\"https://www.kaggle.com/c/landmark-recognition-2020/discussion/176319\" target=\"_blank\">How many unique landmarks are in the testing set?</a></li>\n<li><a href=\"https://www.kaggle.com/c/landmark-recognition-2020/discussion/175171\" target=\"_blank\">Are GLRec and GLRet datasets identical?</a></li>\n<li><a href=\"https://www.kaggle.com/c/landmark-recognition-2020/discussion/173111\" target=\"_blank\">What exactly is the 100k private training set?</a></li>\n<li><a href=\"https://www.kaggle.com/c/landmark-recognition-2020/discussion/174919\" target=\"_blank\">Can someone clarify the public and private data set for me?</a></li>\n</ul>",
      "rawMarkdown": "Just wanted to link related discussions here:\n\n* [Can we pre-compute embeddings? Are image ids the same in PB and PV sets?](https://www.kaggle.com/c/landmark-recognition-2020/discussion/176697)\n* [How many unique landmarks are in the testing set?](https://www.kaggle.com/c/landmark-recognition-2020/discussion/176319)\n* [Are GLRec and GLRet datasets identical?](https://www.kaggle.com/c/landmark-recognition-2020/discussion/175171)\n* [What exactly is the 100k private training set?](https://www.kaggle.com/c/landmark-recognition-2020/discussion/173111)\n* [Can someone clarify the public and private data set for me?](https://www.kaggle.com/c/landmark-recognition-2020/discussion/174919)",
      "votes": null
    },
    {
      "id": "996171",
      "postDate": "09/03/2020 06:09:00",
      "content": "<p>Thank you for the references. Seems the organizers have avoided answering this question even in those threads, so perhaps it will remain unknown.</p>",
      "rawMarkdown": "Thank you for the references. Seems the organizers have avoided answering this question even in those threads, so perhaps it will remain unknown.",
      "votes": null
    },
    {
      "id": "996528",
      "postDate": "09/03/2020 11:08:42",
      "content": "<p>I think yes because they have given - 100k \"subset\" of total public training set.</p>",
      "rawMarkdown": "I think yes because they have given - 100k \"subset\" of total public training set.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 995949,
      "author_name": "chankhavu",
      "author_url": "",
      "post_date": "09/03/2020 01:24:50",
      "content": "<p>Just wanted to link related discussions here:</p>\n<ul>\n<li><a href=\"https://www.kaggle.com/c/landmark-recognition-2020/discussion/176697\" target=\"_blank\">Can we pre-compute embeddings? Are image ids the same in PB and PV sets?</a></li>\n<li><a href=\"https://www.kaggle.com/c/landmark-recognition-2020/discussion/176319\" target=\"_blank\">How many unique landmarks are in the testing set?</a></li>\n<li><a href=\"https://www.kaggle.com/c/landmark-recognition-2020/discussion/175171\" target=\"_blank\">Are GLRec and GLRet datasets identical?</a></li>\n<li><a href=\"https://www.kaggle.com/c/landmark-recognition-2020/discussion/173111\" target=\"_blank\">What exactly is the 100k private training set?</a></li>\n<li><a href=\"https://www.kaggle.com/c/landmark-recognition-2020/discussion/174919\" target=\"_blank\">Can someone clarify the public and private data set for me?</a></li>\n</ul>",
      "votes": null,
      "replies": [
        {
          "id": 996171,
          "author_name": "usmannkhan",
          "author_url": "",
          "post_date": "09/03/2020 06:09:00",
          "content": "<p>Thank you for the references. Seems the organizers have avoided answering this question even in those threads, so perhaps it will remain unknown.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 996528,
      "author_name": "josealways123",
      "author_url": "",
      "post_date": "09/03/2020 11:08:42",
      "content": "<p>I think yes because they have given - 100k \"subset\" of total public training set.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "995732": "I noticed the data description says \"the private training set contains only a 100k subset of the total public training set.\". Does this mean that *all* of the train/ images during the rerun are present in the 1.5M train/ images I can see now?",
    "995949": "Just wanted to link related discussions here:\n\n* [Can we pre-compute embeddings? Are image ids the same in PB and PV sets?](https://www.kaggle.com/c/landmark-recognition-2020/discussion/176697)\n* [How many unique landmarks are in the testing set?](https://www.kaggle.com/c/landmark-recognition-2020/discussion/176319)\n* [Are GLRec and GLRet datasets identical?](https://www.kaggle.com/c/landmark-recognition-2020/discussion/175171)\n* [What exactly is the 100k private training set?](https://www.kaggle.com/c/landmark-recognition-2020/discussion/173111)\n* [Can someone clarify the public and private data set for me?](https://www.kaggle.com/c/landmark-recognition-2020/discussion/174919)",
    "996171": "Thank you for the references. Seems the organizers have avoided answering this question even in those threads, so perhaps it will remain unknown.",
    "996528": "I think yes because they have given - 100k \"subset\" of total public training set."
  },
  "source": "meta"
}