{
  "id": 317922,
  "title": "Use of External Data Sets",
  "url": "/competitions/hotel-id-to-combat-human-trafficking-2022-fgvc9/discussion/317922",
  "author_name": "",
  "post_date": "2022-04-09T15:50:24.702445Z",
  "votes": 7,
  "comment_count": 8,
  "views": 0,
  "content": "<p>Let me confirm if I got the rule. Is the use of external datasets prohibited in this contest? I'm interested in using the following datasets:</p>\n<ul>\n<li>2021 Hotel-ID <a href=\"https://arxiv.org/abs/2106.05746\" target=\"_blank\">https://arxiv.org/abs/2106.05746</a> <a href=\"https://www.kaggle.com/c/hotel-id-2021-fgvc8\" target=\"_blank\">https://www.kaggle.com/c/hotel-id-2021-fgvc8</a></li>\n<li>Hotels-50k <a href=\"https://github.com/GWUvision/Hotels-50K\" target=\"_blank\">https://github.com/GWUvision/Hotels-50K</a></li>\n<li>Pre-trained models trained on the dataset above</li>\n</ul>",
  "messages": [
    {
      "id": "1750397",
      "postDate": "04/09/2022 15:50:24",
      "content": "<p>Let me confirm if I got the rule. Is the use of external datasets prohibited in this contest? I'm interested in using the following datasets:</p>\n<ul>\n<li>2021 Hotel-ID <a href=\"https://arxiv.org/abs/2106.05746\" target=\"_blank\">https://arxiv.org/abs/2106.05746</a> <a href=\"https://www.kaggle.com/c/hotel-id-2021-fgvc8\" target=\"_blank\">https://www.kaggle.com/c/hotel-id-2021-fgvc8</a></li>\n<li>Hotels-50k <a href=\"https://github.com/GWUvision/Hotels-50K\" target=\"_blank\">https://github.com/GWUvision/Hotels-50K</a></li>\n<li>Pre-trained models trained on the dataset above</li>\n</ul>",
      "rawMarkdown": "Let me confirm if I got the rule. Is the use of external datasets prohibited in this contest? I'm interested in using the following datasets:\n\n* 2021 Hotel-ID https://arxiv.org/abs/2106.05746 https://www.kaggle.com/c/hotel-id-2021-fgvc8\n* Hotels-50k https://github.com/GWUvision/Hotels-50K\n* Pre-trained models trained on the dataset above",
      "votes": null
    },
    {
      "id": "1750534",
      "postDate": "04/09/2022 18:12:58",
      "content": "<p>Well that's a good question. In the rules section 7. C.</p>\n<blockquote>\n  <p>The general rule is that participants should only use the provided training and validation images for training models to classify the test images.</p>\n</blockquote>\n<p>but also</p>\n<blockquote>\n  <p>Pretrained models may be used to construct the algorithms from publicly available academic datasets (e.g. ImageNet, iNaturalist 2017-2018, Herbarium 2020).</p>\n</blockquote>\n<p>So it's little confusing and should be clarified. I understand it in a way that you can use pretrained models but during training you should use only competition data. But I have no idea if you can use your pretrained models from last year competition or just pretrain model on Hotels-50k, publish it and then use it here for training.</p>\n<p>But maybe that should not be the point. This competition doesn't award points, tiers or prizes so I don't think we should focus on getting better score using huge ensembles or external datasets. We should try to find ways that could help the organizers to improve their current solution and that could be useful in the real world application. This year competition added occlusions in the test dataset in hope that we can find ways how to deal with them. So maybe methods how to train models to handle them or how to calculate image similarity when there is occlusion would be more interesting for the competition host rather than big ensembles trained on external datasets.</p>",
      "rawMarkdown": "Well that's a good question. In the rules section 7. C.\n> The general rule is that participants should only use the provided training and validation images for training models to classify the test images.\n\nbut also\n> Pretrained models may be used to construct the algorithms from publicly available academic datasets (e.g. ImageNet, iNaturalist 2017-2018, Herbarium 2020).\n\nSo it's little confusing and should be clarified. I understand it in a way that you can use pretrained models but during training you should use only competition data. But I have no idea if you can use your pretrained models from last year competition or just pretrain model on Hotels-50k, publish it and then use it here for training.\n\nBut maybe that should not be the point. This competition doesn't award points, tiers or prizes so I don't think we should focus on getting better score using huge ensembles or external datasets. We should try to find ways that could help the organizers to improve their current solution and that could be useful in the real world application. This year competition added occlusions in the test dataset in hope that we can find ways how to deal with them. So maybe methods how to train models to handle them or how to calculate image similarity when there is occlusion would be more interesting for the competition host rather than big ensembles trained on external datasets.",
      "votes": null
    },
    {
      "id": "1750566",
      "postDate": "04/09/2022 19:33:59",
      "content": "<p>I am not interested in rank point nor huge ensemble. I just want to clarify the rules because the advantage of the algorithm I want to try depends on the data size. If the external data is allowed, I would like to try applying DL based image matching techniques.</p>",
      "rawMarkdown": "I am not interested in rank point nor huge ensemble. I just want to clarify the rules because the advantage of the algorithm I want to try depends on the data size. If the external data is allowed, I would like to try applying DL based image matching techniques.",
      "votes": null
    },
    {
      "id": "1756831",
      "postDate": "04/15/2022 23:26:46",
      "content": "<p>The organisers also failed to clarify this point when it was raised last year, doesn't seem like they have any interest in clarifying. In the final talk about this, they cited pre-training on Hotels-50K as being an important element of the top solutions. I'm taking that as precedent that they didn't have an issue with it.</p>\n<p>However, I also can't download the dataset…</p>",
      "rawMarkdown": "The organisers also failed to clarify this point when it was raised last year, doesn't seem like they have any interest in clarifying. In the final talk about this, they cited pre-training on Hotels-50K as being an important element of the top solutions. I'm taking that as precedent that they didn't have an issue with it.\n\nHowever, I also can't download the dataset...",
      "votes": null
    },
    {
      "id": "1766760",
      "postDate": "04/24/2022 19:42:19",
      "content": "<p>Try to replace \"https\" with \"http\" in download_train.py and it should work, as follows:<br>\ndef url_to_image(url):<br>\n    url = url.replace(\"https\", \"http\")<br>\n    resp = opener.open(url)<br>\n    image = np.asarray(bytearray(resp.read()), dtype=\"uint8\")<br>\n    image = cv2.imdecode(image, cv2.IMREAD_UNCHANGED)<br>\n    return image</p>",
      "rawMarkdown": "Try to replace \"https\" with \"http\" in download_train.py and it should work, as follows:\ndef url_to_image(url):\n    url = url.replace(\"https\", \"http\")\n    resp = opener.open(url)\n    image = np.asarray(bytearray(resp.read()), dtype=\"uint8\")\n    image = cv2.imdecode(image, cv2.IMREAD_UNCHANGED)\n    return image",
      "votes": null
    },
    {
      "id": "1768140",
      "postDate": "04/26/2022 02:47:57",
      "content": "<p><a href=\"https://www.kaggle.com/abbystylianou\" target=\"_blank\">@abbystylianou</a> <a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a> Could you clarify the rules? Still waiting for comments.</p>",
      "rawMarkdown": "abbystylianou @sohier Could you clarify the rules? Still waiting for comments.",
      "votes": null
    },
    {
      "id": "1768498",
      "postDate": "04/26/2022 10:44:26",
      "content": "<p>I managed it in the end - used HTTPS but disabled validation. Got just over 900k photos, trained a SwAV model on it, and got no improvement fine tuning that over general pre-trained models… <em>sigh</em></p>",
      "rawMarkdown": "I managed it in the end - used HTTPS but disabled validation. Got just over 900k photos, trained a SwAV model on it, and got no improvement fine tuning that over general pre-trained models... *sigh*",
      "votes": null
    },
    {
      "id": "1771022",
      "postDate": "04/28/2022 20:49:58",
      "content": "<p>We try to carefully navigate this so as to not encourage participants taking a specific approach or using a particular external dataset. However, I will simply say that there are no points awarded for this competition or monetary awards -- the goal of this competition is to help in the development of the best approaches to recognizing hotels to combat human trafficking. If an external dataset fits that goal, then by all means, use it.</p>",
      "rawMarkdown": "We try to carefully navigate this so as to not encourage participants taking a specific approach or using a particular external dataset. However, I will simply say that there are no points awarded for this competition or monetary awards -- the goal of this competition is to help in the development of the best approaches to recognizing hotels to combat human trafficking. If an external dataset fits that goal, then by all means, use it.",
      "votes": null
    },
    {
      "id": "1771024",
      "postDate": "04/28/2022 20:50:22",
      "content": "<p>Copying my same response from below: We try to carefully navigate this so as to not encourage participants taking a specific approach or using a particular external dataset. However, I will simply say that there are no points awarded for this competition or monetary awards -- the goal of this competition is to help in the development of the best approaches to recognizing hotels to combat human trafficking. If an external dataset fits that goal, then by all means, use it.</p>",
      "rawMarkdown": "Copying my same response from below: We try to carefully navigate this so as to not encourage participants taking a specific approach or using a particular external dataset. However, I will simply say that there are no points awarded for this competition or monetary awards -- the goal of this competition is to help in the development of the best approaches to recognizing hotels to combat human trafficking. If an external dataset fits that goal, then by all means, use it.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1750534,
      "author_name": "michaln",
      "author_url": "",
      "post_date": "04/09/2022 18:12:58",
      "content": "<p>Well that's a good question. In the rules section 7. C.</p>\n<blockquote>\n  <p>The general rule is that participants should only use the provided training and validation images for training models to classify the test images.</p>\n</blockquote>\n<p>but also</p>\n<blockquote>\n  <p>Pretrained models may be used to construct the algorithms from publicly available academic datasets (e.g. ImageNet, iNaturalist 2017-2018, Herbarium 2020).</p>\n</blockquote>\n<p>So it's little confusing and should be clarified. I understand it in a way that you can use pretrained models but during training you should use only competition data. But I have no idea if you can use your pretrained models from last year competition or just pretrain model on Hotels-50k, publish it and then use it here for training.</p>\n<p>But maybe that should not be the point. This competition doesn't award points, tiers or prizes so I don't think we should focus on getting better score using huge ensembles or external datasets. We should try to find ways that could help the organizers to improve their current solution and that could be useful in the real world application. This year competition added occlusions in the test dataset in hope that we can find ways how to deal with them. So maybe methods how to train models to handle them or how to calculate image similarity when there is occlusion would be more interesting for the competition host rather than big ensembles trained on external datasets.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1750566,
          "author_name": "confirm",
          "author_url": "",
          "post_date": "04/09/2022 19:33:59",
          "content": "<p>I am not interested in rank point nor huge ensemble. I just want to clarify the rules because the advantage of the algorithm I want to try depends on the data size. If the external data is allowed, I would like to try applying DL based image matching techniques.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1768140,
          "author_name": "confirm",
          "author_url": "",
          "post_date": "04/26/2022 02:47:57",
          "content": "<p><a href=\"https://www.kaggle.com/abbystylianou\" target=\"_blank\">@abbystylianou</a> <a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a> Could you clarify the rules? Still waiting for comments.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1771024,
          "author_name": "abbystylianou",
          "author_url": "",
          "post_date": "04/28/2022 20:50:22",
          "content": "<p>Copying my same response from below: We try to carefully navigate this so as to not encourage participants taking a specific approach or using a particular external dataset. However, I will simply say that there are no points awarded for this competition or monetary awards -- the goal of this competition is to help in the development of the best approaches to recognizing hotels to combat human trafficking. If an external dataset fits that goal, then by all means, use it.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1756831,
      "author_name": "prubyg",
      "author_url": "",
      "post_date": "04/15/2022 23:26:46",
      "content": "<p>The organisers also failed to clarify this point when it was raised last year, doesn't seem like they have any interest in clarifying. In the final talk about this, they cited pre-training on Hotels-50K as being an important element of the top solutions. I'm taking that as precedent that they didn't have an issue with it.</p>\n<p>However, I also can't download the dataset…</p>",
      "votes": null,
      "replies": [
        {
          "id": 1766760,
          "author_name": "alenic",
          "author_url": "",
          "post_date": "04/24/2022 19:42:19",
          "content": "<p>Try to replace \"https\" with \"http\" in download_train.py and it should work, as follows:<br>\ndef url_to_image(url):<br>\n    url = url.replace(\"https\", \"http\")<br>\n    resp = opener.open(url)<br>\n    image = np.asarray(bytearray(resp.read()), dtype=\"uint8\")<br>\n    image = cv2.imdecode(image, cv2.IMREAD_UNCHANGED)<br>\n    return image</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1768498,
          "author_name": "prubyg",
          "author_url": "",
          "post_date": "04/26/2022 10:44:26",
          "content": "<p>I managed it in the end - used HTTPS but disabled validation. Got just over 900k photos, trained a SwAV model on it, and got no improvement fine tuning that over general pre-trained models… <em>sigh</em></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1771022,
          "author_name": "abbystylianou",
          "author_url": "",
          "post_date": "04/28/2022 20:49:58",
          "content": "<p>We try to carefully navigate this so as to not encourage participants taking a specific approach or using a particular external dataset. However, I will simply say that there are no points awarded for this competition or monetary awards -- the goal of this competition is to help in the development of the best approaches to recognizing hotels to combat human trafficking. If an external dataset fits that goal, then by all means, use it.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1750397": "Let me confirm if I got the rule. Is the use of external datasets prohibited in this contest? I'm interested in using the following datasets:\n\n* 2021 Hotel-ID https://arxiv.org/abs/2106.05746 https://www.kaggle.com/c/hotel-id-2021-fgvc8\n* Hotels-50k https://github.com/GWUvision/Hotels-50K\n* Pre-trained models trained on the dataset above",
    "1750534": "Well that's a good question. In the rules section 7. C.\n> The general rule is that participants should only use the provided training and validation images for training models to classify the test images.\n\nbut also\n> Pretrained models may be used to construct the algorithms from publicly available academic datasets (e.g. ImageNet, iNaturalist 2017-2018, Herbarium 2020).\n\nSo it's little confusing and should be clarified. I understand it in a way that you can use pretrained models but during training you should use only competition data. But I have no idea if you can use your pretrained models from last year competition or just pretrain model on Hotels-50k, publish it and then use it here for training.\n\nBut maybe that should not be the point. This competition doesn't award points, tiers or prizes so I don't think we should focus on getting better score using huge ensembles or external datasets. We should try to find ways that could help the organizers to improve their current solution and that could be useful in the real world application. This year competition added occlusions in the test dataset in hope that we can find ways how to deal with them. So maybe methods how to train models to handle them or how to calculate image similarity when there is occlusion would be more interesting for the competition host rather than big ensembles trained on external datasets.",
    "1750566": "I am not interested in rank point nor huge ensemble. I just want to clarify the rules because the advantage of the algorithm I want to try depends on the data size. If the external data is allowed, I would like to try applying DL based image matching techniques.",
    "1756831": "The organisers also failed to clarify this point when it was raised last year, doesn't seem like they have any interest in clarifying. In the final talk about this, they cited pre-training on Hotels-50K as being an important element of the top solutions. I'm taking that as precedent that they didn't have an issue with it.\n\nHowever, I also can't download the dataset...",
    "1766760": "Try to replace \"https\" with \"http\" in download_train.py and it should work, as follows:\ndef url_to_image(url):\n    url = url.replace(\"https\", \"http\")\n    resp = opener.open(url)\n    image = np.asarray(bytearray(resp.read()), dtype=\"uint8\")\n    image = cv2.imdecode(image, cv2.IMREAD_UNCHANGED)\n    return image",
    "1768140": "abbystylianou @sohier Could you clarify the rules? Still waiting for comments.",
    "1768498": "I managed it in the end - used HTTPS but disabled validation. Got just over 900k photos, trained a SwAV model on it, and got no improvement fine tuning that over general pre-trained models... *sigh*",
    "1771022": "We try to carefully navigate this so as to not encourage participants taking a specific approach or using a particular external dataset. However, I will simply say that there are no points awarded for this competition or monetary awards -- the goal of this competition is to help in the development of the best approaches to recognizing hotels to combat human trafficking. If an external dataset fits that goal, then by all means, use it.",
    "1771024": "Copying my same response from below: We try to carefully navigate this so as to not encourage participants taking a specific approach or using a particular external dataset. However, I will simply say that there are no points awarded for this competition or monetary awards -- the goal of this competition is to help in the development of the best approaches to recognizing hotels to combat human trafficking. If an external dataset fits that goal, then by all means, use it."
  },
  "source": "meta"
}