{
  "id": 266416,
  "title": "Host baseline 2020 - LB 0.202",
  "url": "/competitions/landmark-recognition-2021/discussion/266416",
  "author_name": "",
  "post_date": "2021-08-19T04:30:17.567703600Z",
  "votes": 16,
  "comment_count": 5,
  "views": 0,
  "content": "<p>Last year, the hosts provided us with a great baseline solution. It is also working for this year's edition. I reproduced it here:</p>\n<p><a href=\"https://www.kaggle.com/narsil/host-baseline-2020\" target=\"_blank\">https://www.kaggle.com/narsil/host-baseline-2020</a></p>\n<p>Last year this solution scored 0.46 on public LB and 0.43 on private LB. This year it scores 0.202 on public LB, which suggests that 2021 test set poses a greater challenge than 2020 test set, which is great for the competition and will be stimulating our creativity. </p>",
  "messages": [
    {
      "id": "1480543",
      "postDate": "08/19/2021 04:30:17",
      "content": "<p>Last year, the hosts provided us with a great baseline solution. It is also working for this year's edition. I reproduced it here:</p>\n<p><a href=\"https://www.kaggle.com/narsil/host-baseline-2020\" target=\"_blank\">https://www.kaggle.com/narsil/host-baseline-2020</a></p>\n<p>Last year this solution scored 0.46 on public LB and 0.43 on private LB. This year it scores 0.202 on public LB, which suggests that 2021 test set poses a greater challenge than 2020 test set, which is great for the competition and will be stimulating our creativity. </p>",
      "rawMarkdown": "Last year, the hosts provided us with a great baseline solution. It is also working for this year's edition. I reproduced it here:\n\nhttps://www.kaggle.com/narsil/host-baseline-2020\n\nLast year this solution scored 0.46 on public LB and 0.43 on private LB. This year it scores 0.202 on public LB, which suggests that 2021 test set poses a greater challenge than 2020 test set, which is great for the competition and will be stimulating our creativity.",
      "votes": null
    },
    {
      "id": "1480588",
      "postDate": "08/19/2021 04:56:25",
      "content": "<p>Thanks for sharing this. Can you please help me understand the labels in the test set? From the original paper, there are more than 200k landmarks in the whole dataset. But, the data page says that:</p>\n<blockquote>\n  <p>This 100k subset contains all of the training set images associated with the landmarks in the private test set. </p>\n</blockquote>\n<p>Does this mean that we don't have to use images outside this 100k subset since they will not have landmarks present in the test set?</p>",
      "rawMarkdown": "Thanks for sharing this. Can you please help me understand the labels in the test set? From the original paper, there are more than 200k landmarks in the whole dataset. But, the data page says that:\n\n> This 100k subset contains all of the training set images associated with the landmarks in the private test set. \n\nDoes this mean that we don't have to use images outside this 100k subset since they will not have landmarks present in the test set?",
      "votes": null
    },
    {
      "id": "1480610",
      "postDate": "08/19/2021 05:15:33",
      "content": "<p>In this competition, we have something like a private train set (yes, I meant train not test),</p>\n<p>According to the data description page:</p>\n<blockquote>\n  <p>To facilitate recognition-by-retrieval approaches, the private training set contains only a 100k subset of the total public training set. This 100k subset contains all of the training set images associated with the landmarks in the private test set. You may still attach the full training set as an external data set if you wish.</p>\n</blockquote>\n<p>We cannot see this private train set, our notebooks can only access it during the submission rerun. </p>",
      "rawMarkdown": "In this competition, we have something like a private train set (yes, I meant train not test),\n\nAccording to the data description page:\n\n> To facilitate recognition-by-retrieval approaches, the private training set contains only a 100k subset of the total public training set. This 100k subset contains all of the training set images associated with the landmarks in the private test set. You may still attach the full training set as an external data set if you wish.\n\nWe cannot see this private train set, our notebooks can only access it during the submission rerun.",
      "votes": null
    },
    {
      "id": "1481229",
      "postDate": "08/19/2021 11:30:22",
      "content": "<p>Oh, I completely missed that. Thank you.</p>",
      "rawMarkdown": "Oh, I completely missed that. Thank you.",
      "votes": null
    },
    {
      "id": "1508318",
      "postDate": "09/10/2021 06:06:01",
      "content": "<p><a href=\"https://www.kaggle.com/narsil\" target=\"_blank\">@narsil</a> so if I run submission, it will replace train.csv with private one with only 100k images?</p>",
      "rawMarkdown": "narsil so if I run submission, it will replace train.csv with private one with only 100k images?",
      "votes": null
    },
    {
      "id": "1513738",
      "postDate": "09/15/2021 11:29:23",
      "content": "<p><a href=\"https://www.kaggle.com/joven1997\" target=\"_blank\">@joven1997</a> yes</p>",
      "rawMarkdown": "joven1997 yes",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1480588,
      "author_name": "novice03",
      "author_url": "",
      "post_date": "08/19/2021 04:56:25",
      "content": "<p>Thanks for sharing this. Can you please help me understand the labels in the test set? From the original paper, there are more than 200k landmarks in the whole dataset. But, the data page says that:</p>\n<blockquote>\n  <p>This 100k subset contains all of the training set images associated with the landmarks in the private test set. </p>\n</blockquote>\n<p>Does this mean that we don't have to use images outside this 100k subset since they will not have landmarks present in the test set?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1480610,
          "author_name": "narsil",
          "author_url": "",
          "post_date": "08/19/2021 05:15:33",
          "content": "<p>In this competition, we have something like a private train set (yes, I meant train not test),</p>\n<p>According to the data description page:</p>\n<blockquote>\n  <p>To facilitate recognition-by-retrieval approaches, the private training set contains only a 100k subset of the total public training set. This 100k subset contains all of the training set images associated with the landmarks in the private test set. You may still attach the full training set as an external data set if you wish.</p>\n</blockquote>\n<p>We cannot see this private train set, our notebooks can only access it during the submission rerun. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1481229,
          "author_name": "novice03",
          "author_url": "",
          "post_date": "08/19/2021 11:30:22",
          "content": "<p>Oh, I completely missed that. Thank you.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1508318,
          "author_name": "joven1997",
          "author_url": "",
          "post_date": "09/10/2021 06:06:01",
          "content": "<p><a href=\"https://www.kaggle.com/narsil\" target=\"_blank\">@narsil</a> so if I run submission, it will replace train.csv with private one with only 100k images?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1513738,
          "author_name": "narsil",
          "author_url": "",
          "post_date": "09/15/2021 11:29:23",
          "content": "<p><a href=\"https://www.kaggle.com/joven1997\" target=\"_blank\">@joven1997</a> yes</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1480543": "Last year, the hosts provided us with a great baseline solution. It is also working for this year's edition. I reproduced it here:\n\nhttps://www.kaggle.com/narsil/host-baseline-2020\n\nLast year this solution scored 0.46 on public LB and 0.43 on private LB. This year it scores 0.202 on public LB, which suggests that 2021 test set poses a greater challenge than 2020 test set, which is great for the competition and will be stimulating our creativity.",
    "1480588": "Thanks for sharing this. Can you please help me understand the labels in the test set? From the original paper, there are more than 200k landmarks in the whole dataset. But, the data page says that:\n\n> This 100k subset contains all of the training set images associated with the landmarks in the private test set. \n\nDoes this mean that we don't have to use images outside this 100k subset since they will not have landmarks present in the test set?",
    "1480610": "In this competition, we have something like a private train set (yes, I meant train not test),\n\nAccording to the data description page:\n\n> To facilitate recognition-by-retrieval approaches, the private training set contains only a 100k subset of the total public training set. This 100k subset contains all of the training set images associated with the landmarks in the private test set. You may still attach the full training set as an external data set if you wish.\n\nWe cannot see this private train set, our notebooks can only access it during the submission rerun.",
    "1481229": "Oh, I completely missed that. Thank you.",
    "1508318": "narsil so if I run submission, it will replace train.csv with private one with only 100k images?",
    "1513738": "joven1997 yes"
  },
  "source": "meta"
}