{
  "id": 164321,
  "title": "Be ware - Train vs Test data statistics",
  "url": "/competitions/landmark-retrieval-2020/discussion/164321",
  "author_name": "yuval reina",
  "post_date": "2020-07-05T18:07:13.239000",
  "votes": 7,
  "comment_count": 1,
  "views": 0,
  "content": "<p>In this competition the statistics of the train data and the test data are very different</p>\n\n<p>While in the train data the average is <strong>19.4</strong> images per landmark_id, for the test data this number is <strong>67.5</strong>.</p>\n\n<p>This is very significant when calculating MAP@k:\n1. When a leadmark has many images, they are usually diverse and it is harder to detect all of them\n2. When one landmark has a huge number of images, it acts as noise to the other landmarks</p>\n\n<p>This means, it isn't just the average number of images per landmark that is important but also the distribution.</p>\n\n<p>The question is: What would be the statistics of the held out data. </p>",
  "messages": [
    {
      "id": 916516,
      "postDate": "2020-07-05T18:07:13.240Z",
      "content": "<p>In this competition the statistics of the train data and the test data are very different</p>\n\n<p>While in the train data the average is <strong>19.4</strong> images per landmark_id, for the test data this number is <strong>67.5</strong>.</p>\n\n<p>This is very significant when calculating MAP@k:\n1. When a leadmark has many images, they are usually diverse and it is harder to detect all of them\n2. When one landmark has a huge number of images, it acts as noise to the other landmarks</p>\n\n<p>This means, it isn't just the average number of images per landmark that is important but also the distribution.</p>\n\n<p>The question is: What would be the statistics of the held out data. </p>",
      "rawMarkdown": "In this competition the statistics of the train data and the test data are very different\n\nWhile in the train data the average is **19.4** images per landmark_id, for the test data this number is **67.5**.\n\nThis is very significant when calculating MAP@k:\n1. When a leadmark has many images, they are usually diverse and it is harder to detect all of them\n2. When one landmark has a huge number of images, it acts as noise to the other landmarks\n\nThis means, it isn't just the average number of images per landmark that is important but also the distribution.\n\nThe question is: What would be the statistics of the held out data. ",
      "votes": 8
    },
    {
      "id": 916556,
      "postDate": "2020-07-05T18:45:17.010Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 916556,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-07-05T18:45:17.010000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "916516": "In this competition the statistics of the train data and the test data are very different\n\nWhile in the train data the average is **19.4** images per landmark_id, for the test data this number is **67.5**.\n\nThis is very significant when calculating MAP@k:\n1. When a leadmark has many images, they are usually diverse and it is harder to detect all of them\n2. When one landmark has a huge number of images, it acts as noise to the other landmarks\n\nThis means, it isn't just the average number of images per landmark that is important but also the distribution.\n\nThe question is: What would be the statistics of the held out data. ",
    "916556": ""
  }
}