{
  "id": 65069,
  "title": "Welcome from the host team!",
  "url": "/competitions/inclusive-images-challenge/discussion/65069",
  "author_name": "James Atwood",
  "post_date": "2018-09-05T18:37:47.833000",
  "votes": 12,
  "comment_count": 61,
  "views": 0,
  "content": "<p>Hi, we’re the host team, and on behalf of Google AI, we’re excited to welcome you to the Inclusive Images competition! </p>\n\n<p>The main goal of this research competition is to improve how image classifiers can be made to perform well on unseen geographic distributions. The challenge datasets that we provide come from a very different geographic distribution than the Open Images dataset that you will be training on, thus stress-testing your learned classifier’s ability to generalize to new environments and distributions.</p>\n\n<p>In the spirit of the competition, we ask that competitors read the rules carefully and follow them closely. In particular, this competition has strict rules on the allowed data sources, which have been designed to help make sure that the winning solutions are focusing on the algorithmic portion of this research challenge.</p>\n\n<p>The last thing we want to mention is that this is an <em>inclusive</em> competition, and we’ll ask that usage of these forums reflect this spirit.  Please take care to communicate respectfully with your colleagues and co-competitors on this challenge.</p>\n\n<p>We are very excited to host this competition and to see what you are able to come up with!</p>\n\n<p>-- James,\non behalf of Yoni, Pallavi, Tulsee, D., and the rest of the host team</p>",
  "messages": [
    {
      "id": 382123,
      "postDate": "2018-09-05T18:37:47.833Z",
      "content": "<p>Hi, we’re the host team, and on behalf of Google AI, we’re excited to welcome you to the Inclusive Images competition! </p>\n\n<p>The main goal of this research competition is to improve how image classifiers can be made to perform well on unseen geographic distributions. The challenge datasets that we provide come from a very different geographic distribution than the Open Images dataset that you will be training on, thus stress-testing your learned classifier’s ability to generalize to new environments and distributions.</p>\n\n<p>In the spirit of the competition, we ask that competitors read the rules carefully and follow them closely. In particular, this competition has strict rules on the allowed data sources, which have been designed to help make sure that the winning solutions are focusing on the algorithmic portion of this research challenge.</p>\n\n<p>The last thing we want to mention is that this is an <em>inclusive</em> competition, and we’ll ask that usage of these forums reflect this spirit.  Please take care to communicate respectfully with your colleagues and co-competitors on this challenge.</p>\n\n<p>We are very excited to host this competition and to see what you are able to come up with!</p>\n\n<p>-- James,\non behalf of Yoni, Pallavi, Tulsee, D., and the rest of the host team</p>",
      "rawMarkdown": "Hi, we’re the host team, and on behalf of Google AI, we’re excited to welcome you to the Inclusive Images competition! \n\nThe main goal of this research competition is to improve how image classifiers can be made to perform well on unseen geographic distributions. The challenge datasets that we provide come from a very different geographic distribution than the Open Images dataset that you will be training on, thus stress-testing your learned classifier’s ability to generalize to new environments and distributions.\n\nIn the spirit of the competition, we ask that competitors read the rules carefully and follow them closely. In particular, this competition has strict rules on the allowed data sources, which have been designed to help make sure that the winning solutions are focusing on the algorithmic portion of this research challenge.\n\nThe last thing we want to mention is that this is an *inclusive* competition, and we’ll ask that usage of these forums reflect this spirit.  Please take care to communicate respectfully with your colleagues and co-competitors on this challenge.\n\nWe are very excited to host this competition and to see what you are able to come up with!\n\n-- James,\non behalf of Yoni, Pallavi, Tulsee, D., and the rest of the host team\n",
      "votes": 11
    },
    {
      "id": 383776,
      "postDate": "2018-09-09T15:36:10.053Z",
      "content": "<p>Hi,  @James Atwood</p>\n\n<p>We are not allowed to update the model submitted at the end of phase1. We can only use the submitted model and code to generate the final prediction for stage2. Is this a correct interpretation of the rule?</p>",
      "rawMarkdown": "Hi,  @James Atwood\n\nWe are not allowed to update the model submitted at the end of phase1. We can only use the submitted model and code to generate the final prediction for stage2. Is this a correct interpretation of the rule?",
      "votes": 1,
      "replies": [
        {
          "id": 383812,
          "postDate": "2018-09-09T17:21:37.817Z",
          "content": "<p>Hi Ren,</p>\n\n<p>That's correct - the model should be locked in by the end of stage 1, and that's what should be used for stage 2 predictions.</p>\n\n<p>James</p>",
          "rawMarkdown": "Hi Ren,\n\nThat's correct - the model should be locked in by the end of stage 1, and that's what should be used for stage 2 predictions.\n\nJames",
          "votes": 1
        },
        {
          "id": 383816,
          "postDate": "2018-09-09T17:26:24.587Z",
          "content": "<p>Thanks! What about if I use some heuristic to tune the thresholds (most likely a script, no manual work) that turn model outputs into labels. Does it follow the same rule as with model?</p>",
          "rawMarkdown": "Thanks! What about if I use some heuristic to tune the thresholds (most likely a script, no manual work) that turn model outputs into labels. Does it follow the same rule as with model?"
        },
        {
          "id": 385333,
          "postDate": "2018-09-10T19:54:32.363Z",
          "content": "<p>No threshold tuning should be performed after Stage 1 - like the models, thresholds should be locked in at that time.</p>",
          "rawMarkdown": "No threshold tuning should be performed after Stage 1 - like the models, thresholds should be locked in at that time.",
          "votes": 2
        },
        {
          "id": 385393,
          "postDate": "2018-09-10T23:22:13.820Z",
          "content": "<p>Thanks~</p>",
          "rawMarkdown": "Thanks~"
        },
        {
          "id": 386862,
          "postDate": "2018-09-13T19:00:14.910Z",
          "content": "<p>Hi James, I am under the impression that the 12GB validation images from the openimages website are also not allowed to be used for training. Is it correct?</p>",
          "rawMarkdown": "Hi James, I am under the impression that the 12GB validation images from the openimages website are also not allowed to be used for training. Is it correct?"
        },
        {
          "id": 386880,
          "postDate": "2018-09-13T19:50:01.667Z",
          "content": "<p>That is correct.</p>",
          "rawMarkdown": "That is correct.",
          "votes": 1
        },
        {
          "id": 387668,
          "postDate": "2018-09-15T11:45:05.307Z",
          "content": "<p>Thanks~ We are not allowed to use the image meta data for model training. But, can we use the metadata to split datasets to stress test model locally?</p>",
          "rawMarkdown": "Thanks~ We are not allowed to use the image meta data for model training. But, can we use the metadata to split datasets to stress test model locally?",
          "votes": 1
        },
        {
          "id": 388308,
          "postDate": "2018-09-16T18:39:23.573Z",
          "content": "<p>Ren, why would we not be allow image metadata? Is it specified somewhere? It only says about usage during inference: \"Associated metadata such as the image id or the creator's name are not allowed to be used as inputs at inference time.\"</p>",
          "rawMarkdown": "Ren, why would we not be allow image metadata? Is it specified somewhere? It only says about usage during inference: \"Associated metadata such as the image id or the creator's name are not allowed to be used as inputs at inference time.\""
        },
        {
          "id": 388313,
          "postDate": "2018-09-16T18:57:10.440Z",
          "content": "<p>Sorry, I mean to say 'I am under the impression that we are not allowed to use the image metadata for model training.' </p>\n\n<p>Let me rephrase my original question to the host: </p>\n\n<p>'Is it true that, as long as my model don't use image metadata as inputs during inference time, I can do anything with metadata in training?'</p>",
          "rawMarkdown": "Sorry, I mean to say 'I am under the impression that we are not allowed to use the image metadata for model training.' \n\nLet me rephrase my original question to the host: \n\n'Is it true that, as long as my model don't use image metadata as inputs during inference time, I can do anything with metadata in training?'"
        },
        {
          "id": 388314,
          "postDate": "2018-09-16T18:59:13.543Z",
          "content": "<p>a follow-up question for host: must we use the aws as source to get the images or can we download the \"original images\" from the provided flickr URLs? </p>",
          "rawMarkdown": "a follow-up question for host: must we use the aws as source to get the images or can we download the \"original images\" from the provided flickr URLs? "
        },
        {
          "id": 388997,
          "postDate": "2018-09-18T00:46:38.613Z",
          "content": "<ol>\n<li><p>Metadata that you can find in the image files that we have pointed to (e.g., any exif headers, etc) is fine to use for development as long as your model does not rely on it for inference. At inference time, no metadata can be used other than the image itself.</p></li>\n<li><p>Metadata that cannot be found in the image files themselves (e.g., using the attribution files in the challenge dataset or by following links online to find out more about the training images) may not be used at all. That would be considered external data used for development. The attribution files are provided to ensure that the people who donated the images are properly attributed, but they are not intended to be used in the competition at all.</p></li>\n<li><p>If you'd like to download the images a different way (e.g., through the Flickr URLs) that's fine as long as you ensure that you are using the proper subset of images training images [note that the TSV files with urls on the CVDF site contain urls for the  \"Full dataset\" which is larger than the \"Bounding box subset\"].\"</p></li>\n</ol>",
          "rawMarkdown": "1. Metadata that you can find in the image files that we have pointed to (e.g., any exif headers, etc) is fine to use for development as long as your model does not rely on it for inference. At inference time, no metadata can be used other than the image itself.\n\n2. Metadata that cannot be found in the image files themselves (e.g., using the attribution files in the challenge dataset or by following links online to find out more about the training images) may not be used at all. That would be considered external data used for development. The attribution files are provided to ensure that the people who donated the images are properly attributed, but they are not intended to be used in the competition at all.\n\n3. If you'd like to download the images a different way (e.g., through the Flickr URLs) that's fine as long as you ensure that you are using the proper subset of images training images [note that the TSV files with urls on the CVDF site contain urls for the  \"Full dataset\" which is larger than the \"Bounding box subset\"].\"",
          "votes": 2
        },
        {
          "id": 389049,
          "postDate": "2018-09-18T04:08:14.340Z",
          "content": "<p>@Pallavi, thanks for clarifying. I would like to highlight Ren's question just to be sure.</p>\n\n<blockquote>\n  <p>can we use the metadata to split datasets to stress test model locally?</p>\n</blockquote>\n\n<p>For the purpose of this discussion, let's assume that \"metadata\" refers to <strong>any</strong> metadata (present in the image or obtainable in some other way). It looks like the answer is no. Would like to verify that I'm interpreting your answer correctly given that such a rule would be totally unenforceable.</p>",
          "rawMarkdown": "@Pallavi, thanks for clarifying. I would like to highlight Ren's question just to be sure.\n\n&gt;  can we use the metadata to split datasets to stress test model locally?\n\nFor the purpose of this discussion, let's assume that \"metadata\" refers to **any** metadata (present in the image or obtainable in some other way). It looks like the answer is no. Would like to verify that I'm interpreting your answer correctly given that such a rule would be totally unenforceable.",
          "votes": 3
        },
        {
          "id": 389298,
          "postDate": "2018-09-18T13:54:38.157Z",
          "content": "<p>True. There is no way to check whether people used any specially constructed validation set.</p>\n\n<p>From the perspective of encouraging good approaches to the question that this competition is trying to solve, be able to construct experiment settings where training and validation from different geo-locations are scientifically sound. Otherwise, it feels like a lottery to me. </p>",
          "rawMarkdown": "True. There is no way to check whether people used any specially constructed validation set.\n\nFrom the perspective of encouraging good approaches to the question that this competition is trying to solve, be able to construct experiment settings where training and validation from different geo-locations are scientifically sound. Otherwise, it feels like a lottery to me. ",
          "votes": 3
        }
      ]
    },
    {
      "id": 382597,
      "postDate": "2018-09-06T17:37:33.533Z",
      "content": "<p>I am just wondering if there will  be google cloud platform promo credits available .\nIn my opinion, it would be convenient to utilize them within the context of the investigation of the AI application in the discourse of the competition.</p>",
      "rawMarkdown": "I am just wondering if there will  be google cloud platform promo credits available .\nIn my opinion, it would be convenient to utilize them within the context of the investigation of the AI application in the discourse of the competition.",
      "votes": 1,
      "replies": [
        {
          "id": 382628,
          "postDate": "2018-09-06T18:26:45.327Z",
          "content": "<p>Yes, please view <a href=\"https://www.kaggle.com/c/inclusive-images-challenge/discussion/65157\">this post</a>.</p>",
          "rawMarkdown": "Yes, please view [this post][1].\n\n\n[1]: https://www.kaggle.com/c/inclusive-images-challenge/discussion/65157"
        },
        {
          "id": 382983,
          "postDate": "2018-09-07T13:44:34.960Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 383076,
          "postDate": "2018-09-07T16:26:40.067Z",
          "content": "<p>The credits are intended to augment the computational resources available to individuals to work on a problem. It is not necessary to have the credits to make a basic submission. The requirement that you have made a submission is in place to promote that the use of those credits be specific to entrants to the competition for the purposes of this competition. It is not a guarantee that requestors will use the credits towards other means, but it acts as a basic gating method that demonstrates genuine intent of use.</p>",
          "rawMarkdown": "The credits are intended to augment the computational resources available to individuals to work on a problem. It is not necessary to have the credits to make a basic submission. The requirement that you have made a submission is in place to promote that the use of those credits be specific to entrants to the competition for the purposes of this competition. It is not a guarantee that requestors will use the credits towards other means, but it acts as a basic gating method that demonstrates genuine intent of use."
        },
        {
          "id": 383138,
          "postDate": "2018-09-07T19:08:03.887Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 383411,
          "postDate": "2018-09-08T16:29:06.140Z",
          "content": "<p>Just submit the sample submission file and fill out the form, it should be sufficient.</p>",
          "rawMarkdown": "Just submit the sample submission file and fill out the form, it should be sufficient."
        },
        {
          "id": 383414,
          "postDate": "2018-09-08T16:46:28.637Z",
          "rawMarkdown": "",
          "isDeleted": true
        }
      ]
    },
    {
      "id": 409485,
      "postDate": "2018-10-24T11:30:37.327Z",
      "content": "<p>Hi @James and Co.,</p>\n\n<p>I have some questions about data...</p>\n\n<p>Train:\nIn 000002b66c9c498e.jpg, the people are labeled \"hair, mouth, arm, head and so on\"..(train_human_labels.csv)</p>\n\n<p>Test:\nIn 2b2f44594449326f4e52553d.jpg, the people are labeled \"people and shopkeeper\"(tuning_labels.csv)</p>\n\n<p>If my model predicts 2b2f44594449326f4e52553d.jpg, it will give labels about \"people,shopkeeper, head, hair and so on \". Is this wrong?</p>",
      "rawMarkdown": "Hi @James and Co.,\n\nI have some questions about data...\n\nTrain:\nIn 000002b66c9c498e.jpg, the people are labeled \"hair, mouth, arm, head and so on\"..(train_human_labels.csv)\n\nTest:\nIn 2b2f44594449326f4e52553d.jpg, the people are labeled \"people and shopkeeper\"(tuning_labels.csv)\n\nIf my model predicts 2b2f44594449326f4e52553d.jpg, it will give labels about \"people,shopkeeper, head, hair and so on \". Is this wrong?",
      "votes": 2
    },
    {
      "id": 415647,
      "postDate": "2018-11-05T12:46:46.893Z",
      "content": "<p>Hello~</p>\n\n<p>One thing still confuses me is that, is the 1000 tuning set allowed to be used to train our models? The first line on the data page (<a href=\"https://www.kaggle.com/c/inclusive-images-challenge/data\">https://www.kaggle.com/c/inclusive-images-challenge/data</a>) says <code>This competition uses a portion of the Open Images dataset as the training set</code> which to me means the tuning set it is not allowed in training. </p>\n\n<p>Look at the majority of participants in stage 1, they are using something similar to: <a href=\"https://www.kaggle.com/gpreda/last-day-6-lines-baseline-compact-anton-petrov-s\">https://www.kaggle.com/gpreda/last-day-6-lines-baseline-compact-anton-petrov-s</a> , which essentially is a simple model using only the tuning set. I wouldn't consider it legit according to my interpretation of the rule. </p>\n\n<p>In my opinion the tuning set caused more confusion than if it is not provided at all. </p>",
      "rawMarkdown": "Hello~\n\nOne thing still confuses me is that, is the 1000 tuning set allowed to be used to train our models? The first line on the data page (https://www.kaggle.com/c/inclusive-images-challenge/data) says `This competition uses a portion of the Open Images dataset as the training set` which to me means the tuning set it is not allowed in training. \n\nLook at the majority of participants in stage 1, they are using something similar to: https://www.kaggle.com/gpreda/last-day-6-lines-baseline-compact-anton-petrov-s , which essentially is a simple model using only the tuning set. I wouldn't consider it legit according to my interpretation of the rule. \n\nIn my opinion the tuning set caused more confusion than if it is not provided at all. ",
      "replies": [
        {
          "id": 415734,
          "postDate": "2018-11-05T15:19:31.413Z",
          "content": "<p>Hmm, I got a downvote very quickly. Guess my comment hurt someone spent quite some time trying to utilize the tuning set.</p>\n\n<p>Doing some domain adaptation with the help of the tuning set, I hit the .4  range on stage 1 leaderboard fairly early and climbed up to .5 range about a month ago. I think solutions using the tuning set extensively is totally against the spirit of this competition and abandoned it after that.  </p>\n\n<p>How can someone claim a more generalizable solution and the same time take sneak peeks and modify answers? </p>",
          "rawMarkdown": "Hmm, I got a downvote very quickly. Guess my comment hurt someone spent quite some time trying to utilize the tuning set.\n\nDoing some domain adaptation with the help of the tuning set, I hit the .4  range on stage 1 leaderboard fairly early and climbed up to .5 range about a month ago. I think solutions using the tuning set extensively is totally against the spirit of this competition and abandoned it after that.  \n\nHow can someone claim a more generalizable solution and the same time take sneak peeks and modify answers? ",
          "votes": 1
        }
      ]
    },
    {
      "id": 624131,
      "postDate": "2019-09-11T17:50:04.053Z",
      "content": "<p>Hello,\nI have a question regarding the submission of models.\nI have been introduced to the open images challenge lately and I was working to train a model.\nHowever due to lack of time and my late knowledge, I might not be able to submit before the deadline or even submit a good model.\nIs submission still available after the deadline, in case someone just wants to check how good his model performs on the test set ?\nNote: I dont intend to participate conpetitvely ,I just wanted to work on it to see how well my own model will do.</p>",
      "rawMarkdown": "Hello,\nI have a question regarding the submission of models.\nI have been introduced to the open images challenge lately and I was working to train a model.\nHowever due to lack of time and my late knowledge, I might not be able to submit before the deadline or even submit a good model.\nIs submission still available after the deadline, in case someone just wants to check how good his model performs on the test set ?\nNote: I dont intend to participate conpetitvely ,I just wanted to work on it to see how well my own model will do."
    },
    {
      "id": 411794,
      "postDate": "2018-10-29T03:17:44.753Z",
      "content": "<p>Hi James and team,</p>\n\n<p>Are we allowed to use other kind of data like a contextual based not images.\nword hierachy to accomodate a much general probalistic model.</p>\n\n<p>Cheers,\nPrasen</p>",
      "rawMarkdown": "Hi James and team,\n\nAre we allowed to use other kind of data like a contextual based not images.\nword hierachy to accomodate a much general probalistic model.\n\nCheers,\nPrasen\n",
      "replies": [
        {
          "id": 412104,
          "postDate": "2018-10-29T15:15:02.263Z",
          "content": "<p>Hi Prasen,\nAs mentioned in the rules, you are allowed to use the Wikipedia dump linked in the data download instructions. You can use this Wikipedia dump in any way you find suitable, however, you are not allowed to use other external data apart from the Wikipedia dump to derive word hierarchies.</p>",
          "rawMarkdown": "Hi Prasen,\nAs mentioned in the rules, you are allowed to use the Wikipedia dump linked in the data download instructions. You can use this Wikipedia dump in any way you find suitable, however, you are not allowed to use other external data apart from the Wikipedia dump to derive word hierarchies.\n",
          "votes": 1
        },
        {
          "id": 414395,
          "postDate": "2018-11-02T17:11:00.687Z",
          "content": "<p>The wikipedia dump link does not work again.</p>",
          "rawMarkdown": "The wikipedia dump link does not work again."
        },
        {
          "id": 414531,
          "postDate": "2018-11-02T23:28:47.080Z",
          "content": "<p>Please you this link <a href=\"https://dumps.wikimedia.org/enwiki/20180801/enwiki-20180801-pages-articles-multistream.xml.bz2\">https://dumps.wikimedia.org/enwiki/20180801/enwiki-20180801-pages-articles-multistream.xml.bz2</a></p>",
          "rawMarkdown": "Please you this link https://dumps.wikimedia.org/enwiki/20180801/enwiki-20180801-pages-articles-multistream.xml.bz2",
          "votes": 2
        }
      ]
    },
    {
      "id": 407975,
      "postDate": "2018-10-22T05:06:53.400Z",
      "content": "<p>Hi @James, </p>\n\n<p>I have a question.</p>\n\n<p>From the Inclusive Images FAQ</p>\n\n<blockquote>\n  <p>Note that, per the competition rules, any other form of external data\n  or pre-trained models is not permitted. Competitors are strictly\n  prohibited from using other images or data outside of what has been\n  explicitly approved.</p>\n</blockquote>\n\n<p>Does this mean these csv files are considered external data and not permitted to use?</p>\n\n<ul>\n<li><a href=\"https://storage.googleapis.com/openimages/2018_04/class-descriptions-boxable.csv\">https://storage.googleapis.com/openimages/2018_04/class-descriptions-boxable.csv</a></li>\n<li><a href=\"https://storage.googleapis.com/openimages/2018_04/bbox_labels_600_hierarchy.json\">https://storage.googleapis.com/openimages/2018_04/bbox_labels_600_hierarchy.json</a></li>\n</ul>\n\n<p>(from this link <a href=\"https://storage.googleapis.com/openimages/web/download.html\">https://storage.googleapis.com/openimages/web/download.html</a>)</p>",
      "rawMarkdown": "Hi @James, \n\nI have a question.\n\nFrom the Inclusive Images FAQ\n\n&gt; Note that, per the competition rules, any other form of external data\n&gt; or pre-trained models is not permitted. Competitors are strictly\n&gt; prohibited from using other images or data outside of what has been\n&gt; explicitly approved.\n\nDoes this mean these csv files are considered external data and not permitted to use?\n\n- https://storage.googleapis.com/openimages/2018_04/class-descriptions-boxable.csv\n- https://storage.googleapis.com/openimages/2018_04/bbox_labels_600_hierarchy.json\n\n(from this link https://storage.googleapis.com/openimages/web/download.html)",
      "replies": [
        {
          "id": 408227,
          "postDate": "2018-10-22T13:51:34.193Z",
          "content": "<p>Hi Appian,\nPlease do not use the hierarchy.json file (see previous response here: <a href=\"https://www.kaggle.com/c/inclusive-images-challenge/discussion/65230#383157\">https://www.kaggle.com/c/inclusive-images-challenge/discussion/65230#383157</a>).</p>\n\n<p>The class-description-boxable.csv file is a subset of rows of <a href=\"https://storage.googleapis.com/openimages/2018_04/class-descriptions.csv\">https://storage.googleapis.com/openimages/2018_04/class-descriptions.csv</a>.\nThat seems fine to use.</p>",
          "rawMarkdown": "Hi Appian,\nPlease do not use the hierarchy.json file (see previous response here: https://www.kaggle.com/c/inclusive-images-challenge/discussion/65230#383157).\n\nThe class-description-boxable.csv file is a subset of rows of https://storage.googleapis.com/openimages/2018_04/class-descriptions.csv.\nThat seems fine to use.",
          "votes": 1
        },
        {
          "id": 408243,
          "postDate": "2018-10-22T14:29:16.437Z",
          "content": "<p>Oh, I somehow could not find that discussion.\nThank you for answering and pointed out previous response. It really helps.</p>",
          "rawMarkdown": "Oh, I somehow could not find that discussion.\nThank you for answering and pointed out previous response. It really helps."
        }
      ]
    },
    {
      "id": 398213,
      "postDate": "2018-10-03T17:13:13.207Z",
      "content": "<p>Hi, \nI wanted to know if this competition will allow late submissions after the stage 1 deadline. I understand that this will not count towards leaderboard scores or the competition results. But if I wanted to simply evaluate my scores, will that be possible?</p>",
      "rawMarkdown": "Hi, \nI wanted to know if this competition will allow late submissions after the stage 1 deadline. I understand that this will not count towards leaderboard scores or the competition results. But if I wanted to simply evaluate my scores, will that be possible?"
    },
    {
      "id": 388753,
      "postDate": "2018-09-17T14:13:03.317Z",
      "content": "<p>Hi @James, <br>\nWhat number of classes is used to calculate leaderboard?</p>",
      "rawMarkdown": "Hi @James,  \nWhat number of classes is used to calculate leaderboard?",
      "replies": [
        {
          "id": 389008,
          "postDate": "2018-09-18T01:38:27.670Z",
          "content": "<p>Hi Vecxoz,\nThere are 7178 distinct possible labels that can appear in the test set. These are listed in classes-trainable.csv provided on the data page.</p>",
          "rawMarkdown": "Hi Vecxoz,\nThere are 7178 distinct possible labels that can appear in the test set. These are listed in classes-trainable.csv provided on the data page.",
          "votes": 1
        },
        {
          "id": 389139,
          "postDate": "2018-09-18T08:09:02.217Z",
          "content": "<p>Thanks!</p>",
          "rawMarkdown": "Thanks!"
        }
      ]
    },
    {
      "id": 383193,
      "postDate": "2018-09-08T01:28:30.230Z",
      "content": "<p>@julia, a very novice question - I am facing issues in transferring training data from the S3 bucket to Google Cloud Storage Bucket. What Access Key ID and Secret Access Key to use? </p>",
      "rawMarkdown": "@julia, a very novice question - I am facing issues in transferring training data from the S3 bucket to Google Cloud Storage Bucket. What Access Key ID and Secret Access Key to use? ",
      "replies": [
        {
          "id": 383219,
          "postDate": "2018-09-08T03:36:11.883Z",
          "content": "<p>you may approach using your keys, in case you have ones, otherwise the parameter --no-sign-request will waive the need to provide the ones, in my opinion. Moreover, you may approach downloading the dataset using other means, in my opinion. I did not try, but <a href=\"https://datasets.figure-eight.com/figure_eight_datasets/open-images/test_challenge.zip\">https://datasets.figure-eight.com/figure_eight_datasets/open-images/test_challenge.zip</a> will possibly be the same dataset, as it seems to me.</p>\n\n<p>And the <a href=\"https://github.com/cvdfoundation/open-images-dataset#download-images-with-bounding-boxes-annotations\">page</a> will indeed have the reference to the exactly same dataset.\nReference: <a href=\"https://github.com/cvdfoundation/open-images-dataset#download-images-with-bounding-boxes-annotations\">https://github.com/cvdfoundation/open-images-dataset#download-images-with-bounding-boxes-annotations</a></p>\n\n<p>Upd:</p>\n\n<blockquote>\n  <p>--no-sign-request (boolean)</p>\n  \n  <p>Do not sign requests. Credentials will not be loaded if this argument\n  is provided.\n  source: <a href=\"https://docs.aws.amazon.com/cli/latest/reference/\">https://docs.aws.amazon.com/cli/latest/reference/</a></p>\n</blockquote>",
          "rawMarkdown": "you may approach using your keys, in case you have ones, otherwise the parameter --no-sign-request will waive the need to provide the ones, in my opinion. Moreover, you may approach downloading the dataset using other means, in my opinion. I did not try, but https://datasets.figure-eight.com/figure_eight_datasets/open-images/test_challenge.zip will possibly be the same dataset, as it seems to me.\n\nAnd the [page][1] will indeed have the reference to the exactly same dataset.\nReference: https://github.com/cvdfoundation/open-images-dataset#download-images-with-bounding-boxes-annotations\n\n\n  [1]: https://github.com/cvdfoundation/open-images-dataset#download-images-with-bounding-boxes-annotations\n\nUpd:\n\n&gt; --no-sign-request (boolean)\n&gt; \n&gt; Do not sign requests. Credentials will not be loaded if this argument\n&gt; is provided.\nsource: https://docs.aws.amazon.com/cli/latest/reference/"
        },
        {
          "id": 383671,
          "postDate": "2018-09-09T08:36:46.883Z",
          "content": "<p>Does anyone have an idea on how to process the images one by one in a loop using torch, digits, caffe or tensorflow? In a way that it will take from the dataset and will output the compliant format for the submission. \nDid anybody manage to format the dataset in a way it got compatible with AutoML?\nDoes kaggle kernel provide enough space to upload the dataset?\nWill apt-get update succeed on a kernel? ( in my case it halts in a loop at 8%)</p>",
          "rawMarkdown": "Does anyone have an idea on how to process the images one by one in a loop using torch, digits, caffe or tensorflow? In a way that it will take from the dataset and will output the compliant format for the submission. \nDid anybody manage to format the dataset in a way it got compatible with AutoML?\nDoes kaggle kernel provide enough space to upload the dataset?\nWill apt-get update succeed on a kernel? ( in my case it halts in a loop at 8%)"
        },
        {
          "id": 398906,
          "postDate": "2018-10-04T19:16:58.383Z",
          "content": "<p>As I have managed to add the coupon - I will upload the datasets to google storage and will share it for 60 days to public. Regards</p>",
          "rawMarkdown": "As I have managed to add the coupon - I will upload the datasets to google storage and will share it for 60 days to public. Regards"
        },
        {
          "id": 398934,
          "postDate": "2018-10-04T20:27:29.473Z",
          "content": "<p>The bucket has been created and the user: allUsers has been added with Google Storage Viewer permission. Let me know if that permission is redundant for operations with the bucket in a read mode. Otherwise it can be elevated to Storage Admin level that will get the system less reliable though.\n<a href=\"https://console.cloud.google.com/storage/browser/inclusive-images_challenge\">enter link description here</a>\nImages will appear in the bucket as soon as file transfer from aws will finish and then upload to gs will finish\ngs://inclusive-images_challenge</p>",
          "rawMarkdown": "The bucket has been created and the user: allUsers has been added with Google Storage Viewer permission. Let me know if that permission is redundant for operations with the bucket in a read mode. Otherwise it can be elevated to Storage Admin level that will get the system less reliable though.\n[enter link description here][1]\nImages will appear in the bucket as soon as file transfer from aws will finish and then upload to gs will finish\ngs://inclusive-images_challenge\n  [1]: https://console.cloud.google.com/storage/browser/inclusive-images_challenge"
        },
        {
          "id": 399078,
          "postDate": "2018-10-05T06:28:52.133Z",
          "content": "<p>download has been completed and upload has been started</p>",
          "rawMarkdown": "download has been completed and upload has been started"
        },
        {
          "id": 399994,
          "postDate": "2018-10-07T10:16:44.653Z",
          "content": "<p>it seems that the upload got finished, but I started the procedure again just to make sure the whole thing got transferred. In a couple of days the resource could be available for utilization within the context of the challenge and that will waive the instruction step to install awscli and will allow just to start working with gs directly at once, hopefully</p>",
          "rawMarkdown": "it seems that the upload got finished, but I started the procedure again just to make sure the whole thing got transferred. In a couple of days the resource could be available for utilization within the context of the challenge and that will waive the instruction step to install awscli and will allow just to start working with gs directly at once, hopefully"
        },
        {
          "id": 400473,
          "postDate": "2018-10-08T11:36:24.117Z",
          "content": "<p>the upload has been finished</p>",
          "rawMarkdown": "the upload has been finished"
        }
      ]
    },
    {
      "id": 382616,
      "postDate": "2018-09-06T18:11:39.820Z",
      "content": "<p>Thank you for this competition !</p>\n\n<p>PS: The Wikipedia text data <a href=\"https://dumps.wikimedia.org/enwiki/20180720/enwiki-20180720-pages-articles-multistream.xml.bz2&amp;sa=D&amp;source=hangouts&amp;ust=1535834316553000&amp;usg=AFQjCNFkMs_tSMZOTR2OLCDAm9PVJS9LKw\">link</a> in the data tab doesn't work.</p>",
      "rawMarkdown": "Thank you for this competition !\n\nPS: The Wikipedia text data [link][1] in the data tab doesn't work.\n\n\n  [1]: https://dumps.wikimedia.org/enwiki/20180720/enwiki-20180720-pages-articles-multistream.xml.bz2&amp;sa=D&amp;source=hangouts&amp;ust=1535834316553000&amp;usg=AFQjCNFkMs_tSMZOTR2OLCDAm9PVJS9LKw",
      "replies": [
        {
          "id": 382620,
          "postDate": "2018-09-06T18:13:27.150Z",
          "content": "<p>Thanks for flagging that; the link should be working now.</p>",
          "rawMarkdown": "Thanks for flagging that; the link should be working now."
        },
        {
          "id": 383048,
          "postDate": "2018-09-07T15:45:52.703Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 383056,
          "postDate": "2018-09-07T15:55:50.807Z",
          "content": "<p>Hi Andrei, if the link <a href=\"https://www.kaggle.com/c/inclusive-images-challenge#Data-Download-&amp;-Getting-Started\">https://www.kaggle.com/c/inclusive-images-challenge#Data-Download-&amp;-Getting-Started</a> doesn't work for your browser, go to the overview page and hit \"Data Download &amp; Getting Started\" on the left sidebar.  This contains instructions for downloading the correct OpenImages files (also called the correct subset), which are outlined in red on the screenshots on that page.  Hope this helps!</p>",
          "rawMarkdown": "Hi Andrei, if the link https://www.kaggle.com/c/inclusive-images-challenge#Data-Download-&amp;-Getting-Started doesn't work for your browser, go to the overview page and hit \"Data Download &amp; Getting Started\" on the left sidebar.  This contains instructions for downloading the correct OpenImages files (also called the correct subset), which are outlined in red on the screenshots on that page.  Hope this helps!"
        },
        {
          "id": 383061,
          "postDate": "2018-09-07T16:00:00.243Z",
          "content": "<p>got it</p>",
          "rawMarkdown": "got it"
        },
        {
          "id": 383066,
          "postDate": "2018-09-07T16:04:14.307Z",
          "content": "<p>It appears that the open images website is <a href=\"https://storage.googleapis.com/openimages/web/download.html\">https://storage.googleapis.com/openimages/web/download.html</a></p>",
          "rawMarkdown": "It appears that the open images website is https://storage.googleapis.com/openimages/web/download.html"
        },
        {
          "id": 383072,
          "postDate": "2018-09-07T16:20:41.560Z",
          "content": "<p>@Andrei Volodin - Apologies, it looks as though I missed a URL link in those instructions. Yes, that is the correct link. I have updated the instructions accordingly.</p>",
          "rawMarkdown": "@Andrei Volodin - Apologies, it looks as though I missed a URL link in those instructions. Yes, that is the correct link. I have updated the instructions accordingly."
        },
        {
          "id": 383083,
          "postDate": "2018-09-07T16:41:06.663Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 383221,
          "postDate": "2018-09-08T03:49:41.647Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 383224,
          "postDate": "2018-09-08T03:57:22.447Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 400472,
          "postDate": "2018-10-08T11:35:23.130Z",
          "content": "<p>@Julia, the dataset [512gb] has been mirrored to gs://inclusive-images_challenge that is accessible for public. You may like to add the link to the data download section ( though it will be available just for 2 month)</p>",
          "rawMarkdown": "@Julia, the dataset [512gb] has been mirrored to gs://inclusive-images_challenge that is accessible for public. You may like to add the link to the data download section ( though it will be available just for 2 month)"
        },
        {
          "id": 400666,
          "postDate": "2018-10-08T17:40:02.507Z",
          "content": "<p>1743042 is the number of files in the folder</p>",
          "rawMarkdown": "1743042 is the number of files in the folder"
        }
      ]
    },
    {
      "id": 382158,
      "postDate": "2018-09-05T20:09:58.553Z",
      "content": "<p>Seems like a very interesting competition. Thank you for hosting.</p>",
      "rawMarkdown": "Seems like a very interesting competition. Thank you for hosting."
    },
    {
      "id": 382137,
      "postDate": "2018-09-05T19:14:34.667Z",
      "content": "<p>Hi James and Co.,</p>\n\n<p>first of thanks for hosting this competition. I'm looking forward to all the wacky solutions Kagglers will come up with.</p>\n\n<p>My main question is about NIPS presentation - given how highly contested the tickets are. Could you let us know for how many participants you have reserved tickets, and what will be the distribution?</p>\n\n<p>best, \nMiha</p>",
      "rawMarkdown": "Hi James and Co.,\n\nfirst of thanks for hosting this competition. I'm looking forward to all the wacky solutions Kagglers will come up with.\n\nMy main question is about NIPS presentation - given how highly contested the tickets are. Could you let us know for how many participants you have reserved tickets, and what will be the distribution?\n\nbest, \nMiha",
      "replies": [
        {
          "id": 382172,
          "postDate": "2018-09-05T20:38:29.507Z",
          "content": "<p>Hi Miha, thanks for the interest!  As mentioned in the FAQ, we'll be able to provide registration (and some travel expenses) for one representative from each of the top five teams for the workshop component at NIPS.</p>",
          "rawMarkdown": "Hi Miha, thanks for the interest!  As mentioned in the FAQ, we'll be able to provide registration (and some travel expenses) for one representative from each of the top five teams for the workshop component at NIPS.",
          "votes": 1
        },
        {
          "id": 382198,
          "postDate": "2018-09-05T21:11:11.510Z",
          "content": "<p>thanks for the answer Sculley! \nand it's \"only\" for the workshops, isn't it? :/</p>",
          "rawMarkdown": "thanks for the answer Sculley! \nand it's \"only\" for the workshops, isn't it? :/"
        }
      ]
    },
    {
      "id": 424363,
      "postDate": "2018-11-20T01:30:39.350Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 383776,
      "author_name": "Ren",
      "author_url": "",
      "post_date": "2018-09-09T15:36:10.053000",
      "content": "<p>Hi,  @James Atwood</p>\n\n<p>We are not allowed to update the model submitted at the end of phase1. We can only use the submitted model and code to generate the final prediction for stage2. Is this a correct interpretation of the rule?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 383812,
          "author_name": "James Atwood",
          "author_url": "",
          "post_date": "2018-09-09T17:21:37.817000",
          "content": "<p>Hi Ren,</p>\n\n<p>That's correct - the model should be locked in by the end of stage 1, and that's what should be used for stage 2 predictions.</p>\n\n<p>James</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 383816,
          "author_name": "Ren",
          "author_url": "",
          "post_date": "2018-09-09T17:26:24.587000",
          "content": "<p>Thanks! What about if I use some heuristic to tune the thresholds (most likely a script, no manual work) that turn model outputs into labels. Does it follow the same rule as with model?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 385333,
          "author_name": "James Atwood",
          "author_url": "",
          "post_date": "2018-09-10T19:54:32.363000",
          "content": "<p>No threshold tuning should be performed after Stage 1 - like the models, thresholds should be locked in at that time.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 385393,
          "author_name": "Ren",
          "author_url": "",
          "post_date": "2018-09-10T23:22:13.820000",
          "content": "<p>Thanks~</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 386862,
          "author_name": "Ren",
          "author_url": "",
          "post_date": "2018-09-13T19:00:14.910000",
          "content": "<p>Hi James, I am under the impression that the 12GB validation images from the openimages website are also not allowed to be used for training. Is it correct?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 386880,
          "author_name": "Yoni Halpern",
          "author_url": "",
          "post_date": "2018-09-13T19:50:01.667000",
          "content": "<p>That is correct.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 387668,
          "author_name": "Ren",
          "author_url": "",
          "post_date": "2018-09-15T11:45:05.307000",
          "content": "<p>Thanks~ We are not allowed to use the image meta data for model training. But, can we use the metadata to split datasets to stress test model locally?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 388308,
          "author_name": "Miha Skalic",
          "author_url": "",
          "post_date": "2018-09-16T18:39:23.573000",
          "content": "<p>Ren, why would we not be allow image metadata? Is it specified somewhere? It only says about usage during inference: \"Associated metadata such as the image id or the creator's name are not allowed to be used as inputs at inference time.\"</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 388313,
          "author_name": "Ren",
          "author_url": "",
          "post_date": "2018-09-16T18:57:10.440000",
          "content": "<p>Sorry, I mean to say 'I am under the impression that we are not allowed to use the image metadata for model training.' </p>\n\n<p>Let me rephrase my original question to the host: </p>\n\n<p>'Is it true that, as long as my model don't use image metadata as inputs during inference time, I can do anything with metadata in training?'</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 388314,
          "author_name": "Miha Skalic",
          "author_url": "",
          "post_date": "2018-09-16T18:59:13.543000",
          "content": "<p>a follow-up question for host: must we use the aws as source to get the images or can we download the \"original images\" from the provided flickr URLs? </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 388997,
          "author_name": "Pallavi",
          "author_url": "",
          "post_date": "2018-09-18T00:46:38.613000",
          "content": "<ol>\n<li><p>Metadata that you can find in the image files that we have pointed to (e.g., any exif headers, etc) is fine to use for development as long as your model does not rely on it for inference. At inference time, no metadata can be used other than the image itself.</p></li>\n<li><p>Metadata that cannot be found in the image files themselves (e.g., using the attribution files in the challenge dataset or by following links online to find out more about the training images) may not be used at all. That would be considered external data used for development. The attribution files are provided to ensure that the people who donated the images are properly attributed, but they are not intended to be used in the competition at all.</p></li>\n<li><p>If you'd like to download the images a different way (e.g., through the Flickr URLs) that's fine as long as you ensure that you are using the proper subset of images training images [note that the TSV files with urls on the CVDF site contain urls for the  \"Full dataset\" which is larger than the \"Bounding box subset\"].\"</p></li>\n</ol>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 389049,
          "author_name": "Anil Thomas",
          "author_url": "",
          "post_date": "2018-09-18T04:08:14.340000",
          "content": "<p>@Pallavi, thanks for clarifying. I would like to highlight Ren's question just to be sure.</p>\n\n<blockquote>\n  <p>can we use the metadata to split datasets to stress test model locally?</p>\n</blockquote>\n\n<p>For the purpose of this discussion, let's assume that \"metadata\" refers to <strong>any</strong> metadata (present in the image or obtainable in some other way). It looks like the answer is no. Would like to verify that I'm interpreting your answer correctly given that such a rule would be totally unenforceable.</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 389298,
          "author_name": "Ren",
          "author_url": "",
          "post_date": "2018-09-18T13:54:38.157000",
          "content": "<p>True. There is no way to check whether people used any specially constructed validation set.</p>\n\n<p>From the perspective of encouraging good approaches to the question that this competition is trying to solve, be able to construct experiment settings where training and validation from different geo-locations are scientifically sound. Otherwise, it feels like a lottery to me. </p>",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 382597,
      "author_name": "Andrei Volodin",
      "author_url": "",
      "post_date": "2018-09-06T17:37:33.533000",
      "content": "<p>I am just wondering if there will  be google cloud platform promo credits available .\nIn my opinion, it would be convenient to utilize them within the context of the investigation of the AI application in the discourse of the competition.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 382628,
          "author_name": "Julia Elliott",
          "author_url": "",
          "post_date": "2018-09-06T18:26:45.327000",
          "content": "<p>Yes, please view <a href=\"https://www.kaggle.com/c/inclusive-images-challenge/discussion/65157\">this post</a>.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 382983,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-09-07T13:44:34.960000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 383076,
          "author_name": "Julia Elliott",
          "author_url": "",
          "post_date": "2018-09-07T16:26:40.067000",
          "content": "<p>The credits are intended to augment the computational resources available to individuals to work on a problem. It is not necessary to have the credits to make a basic submission. The requirement that you have made a submission is in place to promote that the use of those credits be specific to entrants to the competition for the purposes of this competition. It is not a guarantee that requestors will use the credits towards other means, but it acts as a basic gating method that demonstrates genuine intent of use.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 383138,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-09-07T19:08:03.887000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 383411,
          "author_name": "Ren",
          "author_url": "",
          "post_date": "2018-09-08T16:29:06.140000",
          "content": "<p>Just submit the sample submission file and fill out the form, it should be sufficient.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 383414,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-09-08T16:46:28.637000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 409485,
      "author_name": "Mii",
      "author_url": "",
      "post_date": "2018-10-24T11:30:37.327000",
      "content": "<p>Hi @James and Co.,</p>\n\n<p>I have some questions about data...</p>\n\n<p>Train:\nIn 000002b66c9c498e.jpg, the people are labeled \"hair, mouth, arm, head and so on\"..(train_human_labels.csv)</p>\n\n<p>Test:\nIn 2b2f44594449326f4e52553d.jpg, the people are labeled \"people and shopkeeper\"(tuning_labels.csv)</p>\n\n<p>If my model predicts 2b2f44594449326f4e52553d.jpg, it will give labels about \"people,shopkeeper, head, hair and so on \". Is this wrong?</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 415647,
      "author_name": "Ren",
      "author_url": "",
      "post_date": "2018-11-05T12:46:46.893000",
      "content": "<p>Hello~</p>\n\n<p>One thing still confuses me is that, is the 1000 tuning set allowed to be used to train our models? The first line on the data page (<a href=\"https://www.kaggle.com/c/inclusive-images-challenge/data\">https://www.kaggle.com/c/inclusive-images-challenge/data</a>) says <code>This competition uses a portion of the Open Images dataset as the training set</code> which to me means the tuning set it is not allowed in training. </p>\n\n<p>Look at the majority of participants in stage 1, they are using something similar to: <a href=\"https://www.kaggle.com/gpreda/last-day-6-lines-baseline-compact-anton-petrov-s\">https://www.kaggle.com/gpreda/last-day-6-lines-baseline-compact-anton-petrov-s</a> , which essentially is a simple model using only the tuning set. I wouldn't consider it legit according to my interpretation of the rule. </p>\n\n<p>In my opinion the tuning set caused more confusion than if it is not provided at all. </p>",
      "votes": 0,
      "replies": [
        {
          "id": 415734,
          "author_name": "Ren",
          "author_url": "",
          "post_date": "2018-11-05T15:19:31.413000",
          "content": "<p>Hmm, I got a downvote very quickly. Guess my comment hurt someone spent quite some time trying to utilize the tuning set.</p>\n\n<p>Doing some domain adaptation with the help of the tuning set, I hit the .4  range on stage 1 leaderboard fairly early and climbed up to .5 range about a month ago. I think solutions using the tuning set extensively is totally against the spirit of this competition and abandoned it after that.  </p>\n\n<p>How can someone claim a more generalizable solution and the same time take sneak peeks and modify answers? </p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 624131,
      "author_name": "Hasan N",
      "author_url": "",
      "post_date": "2019-09-11T17:50:04.053000",
      "content": "<p>Hello,\nI have a question regarding the submission of models.\nI have been introduced to the open images challenge lately and I was working to train a model.\nHowever due to lack of time and my late knowledge, I might not be able to submit before the deadline or even submit a good model.\nIs submission still available after the deadline, in case someone just wants to check how good his model performs on the test set ?\nNote: I dont intend to participate conpetitvely ,I just wanted to work on it to see how well my own model will do.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 411794,
      "author_name": "Prasen Jit Singh",
      "author_url": "",
      "post_date": "2018-10-29T03:17:44.753000",
      "content": "<p>Hi James and team,</p>\n\n<p>Are we allowed to use other kind of data like a contextual based not images.\nword hierachy to accomodate a much general probalistic model.</p>\n\n<p>Cheers,\nPrasen</p>",
      "votes": 0,
      "replies": [
        {
          "id": 412104,
          "author_name": "Pallavi",
          "author_url": "",
          "post_date": "2018-10-29T15:15:02.263000",
          "content": "<p>Hi Prasen,\nAs mentioned in the rules, you are allowed to use the Wikipedia dump linked in the data download instructions. You can use this Wikipedia dump in any way you find suitable, however, you are not allowed to use other external data apart from the Wikipedia dump to derive word hierarchies.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 414395,
          "author_name": "Philip Popien",
          "author_url": "",
          "post_date": "2018-11-02T17:11:00.687000",
          "content": "<p>The wikipedia dump link does not work again.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 414531,
          "author_name": "Pallavi",
          "author_url": "",
          "post_date": "2018-11-02T23:28:47.080000",
          "content": "<p>Please you this link <a href=\"https://dumps.wikimedia.org/enwiki/20180801/enwiki-20180801-pages-articles-multistream.xml.bz2\">https://dumps.wikimedia.org/enwiki/20180801/enwiki-20180801-pages-articles-multistream.xml.bz2</a></p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 407975,
      "author_name": "Appian",
      "author_url": "",
      "post_date": "2018-10-22T05:06:53.400000",
      "content": "<p>Hi @James, </p>\n\n<p>I have a question.</p>\n\n<p>From the Inclusive Images FAQ</p>\n\n<blockquote>\n  <p>Note that, per the competition rules, any other form of external data\n  or pre-trained models is not permitted. Competitors are strictly\n  prohibited from using other images or data outside of what has been\n  explicitly approved.</p>\n</blockquote>\n\n<p>Does this mean these csv files are considered external data and not permitted to use?</p>\n\n<ul>\n<li><a href=\"https://storage.googleapis.com/openimages/2018_04/class-descriptions-boxable.csv\">https://storage.googleapis.com/openimages/2018_04/class-descriptions-boxable.csv</a></li>\n<li><a href=\"https://storage.googleapis.com/openimages/2018_04/bbox_labels_600_hierarchy.json\">https://storage.googleapis.com/openimages/2018_04/bbox_labels_600_hierarchy.json</a></li>\n</ul>\n\n<p>(from this link <a href=\"https://storage.googleapis.com/openimages/web/download.html\">https://storage.googleapis.com/openimages/web/download.html</a>)</p>",
      "votes": 0,
      "replies": [
        {
          "id": 408227,
          "author_name": "Yoni Halpern",
          "author_url": "",
          "post_date": "2018-10-22T13:51:34.193000",
          "content": "<p>Hi Appian,\nPlease do not use the hierarchy.json file (see previous response here: <a href=\"https://www.kaggle.com/c/inclusive-images-challenge/discussion/65230#383157\">https://www.kaggle.com/c/inclusive-images-challenge/discussion/65230#383157</a>).</p>\n\n<p>The class-description-boxable.csv file is a subset of rows of <a href=\"https://storage.googleapis.com/openimages/2018_04/class-descriptions.csv\">https://storage.googleapis.com/openimages/2018_04/class-descriptions.csv</a>.\nThat seems fine to use.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 408243,
          "author_name": "Appian",
          "author_url": "",
          "post_date": "2018-10-22T14:29:16.437000",
          "content": "<p>Oh, I somehow could not find that discussion.\nThank you for answering and pointed out previous response. It really helps.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 398213,
      "author_name": "Akshita Bhagia",
      "author_url": "",
      "post_date": "2018-10-03T17:13:13.207000",
      "content": "<p>Hi, \nI wanted to know if this competition will allow late submissions after the stage 1 deadline. I understand that this will not count towards leaderboard scores or the competition results. But if I wanted to simply evaluate my scores, will that be possible?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 388753,
      "author_name": "vecxoz",
      "author_url": "",
      "post_date": "2018-09-17T14:13:03.317000",
      "content": "<p>Hi @James, <br>\nWhat number of classes is used to calculate leaderboard?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 389008,
          "author_name": "Yoni Halpern",
          "author_url": "",
          "post_date": "2018-09-18T01:38:27.670000",
          "content": "<p>Hi Vecxoz,\nThere are 7178 distinct possible labels that can appear in the test set. These are listed in classes-trainable.csv provided on the data page.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 389139,
          "author_name": "vecxoz",
          "author_url": "",
          "post_date": "2018-09-18T08:09:02.217000",
          "content": "<p>Thanks!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 383193,
      "author_name": "av.in",
      "author_url": "",
      "post_date": "2018-09-08T01:28:30.230000",
      "content": "<p>@julia, a very novice question - I am facing issues in transferring training data from the S3 bucket to Google Cloud Storage Bucket. What Access Key ID and Secret Access Key to use? </p>",
      "votes": 0,
      "replies": [
        {
          "id": 383219,
          "author_name": "Andrei Volodin",
          "author_url": "",
          "post_date": "2018-09-08T03:36:11.883000",
          "content": "<p>you may approach using your keys, in case you have ones, otherwise the parameter --no-sign-request will waive the need to provide the ones, in my opinion. Moreover, you may approach downloading the dataset using other means, in my opinion. I did not try, but <a href=\"https://datasets.figure-eight.com/figure_eight_datasets/open-images/test_challenge.zip\">https://datasets.figure-eight.com/figure_eight_datasets/open-images/test_challenge.zip</a> will possibly be the same dataset, as it seems to me.</p>\n\n<p>And the <a href=\"https://github.com/cvdfoundation/open-images-dataset#download-images-with-bounding-boxes-annotations\">page</a> will indeed have the reference to the exactly same dataset.\nReference: <a href=\"https://github.com/cvdfoundation/open-images-dataset#download-images-with-bounding-boxes-annotations\">https://github.com/cvdfoundation/open-images-dataset#download-images-with-bounding-boxes-annotations</a></p>\n\n<p>Upd:</p>\n\n<blockquote>\n  <p>--no-sign-request (boolean)</p>\n  \n  <p>Do not sign requests. Credentials will not be loaded if this argument\n  is provided.\n  source: <a href=\"https://docs.aws.amazon.com/cli/latest/reference/\">https://docs.aws.amazon.com/cli/latest/reference/</a></p>\n</blockquote>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 383671,
          "author_name": "Andrei Volodin",
          "author_url": "",
          "post_date": "2018-09-09T08:36:46.883000",
          "content": "<p>Does anyone have an idea on how to process the images one by one in a loop using torch, digits, caffe or tensorflow? In a way that it will take from the dataset and will output the compliant format for the submission. \nDid anybody manage to format the dataset in a way it got compatible with AutoML?\nDoes kaggle kernel provide enough space to upload the dataset?\nWill apt-get update succeed on a kernel? ( in my case it halts in a loop at 8%)</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 398906,
          "author_name": "Andrei Volodin",
          "author_url": "",
          "post_date": "2018-10-04T19:16:58.383000",
          "content": "<p>As I have managed to add the coupon - I will upload the datasets to google storage and will share it for 60 days to public. Regards</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 398934,
          "author_name": "Andrei Volodin",
          "author_url": "",
          "post_date": "2018-10-04T20:27:29.473000",
          "content": "<p>The bucket has been created and the user: allUsers has been added with Google Storage Viewer permission. Let me know if that permission is redundant for operations with the bucket in a read mode. Otherwise it can be elevated to Storage Admin level that will get the system less reliable though.\n<a href=\"https://console.cloud.google.com/storage/browser/inclusive-images_challenge\">enter link description here</a>\nImages will appear in the bucket as soon as file transfer from aws will finish and then upload to gs will finish\ngs://inclusive-images_challenge</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 399078,
          "author_name": "Andrei Volodin",
          "author_url": "",
          "post_date": "2018-10-05T06:28:52.133000",
          "content": "<p>download has been completed and upload has been started</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 399994,
          "author_name": "Andrei Volodin",
          "author_url": "",
          "post_date": "2018-10-07T10:16:44.653000",
          "content": "<p>it seems that the upload got finished, but I started the procedure again just to make sure the whole thing got transferred. In a couple of days the resource could be available for utilization within the context of the challenge and that will waive the instruction step to install awscli and will allow just to start working with gs directly at once, hopefully</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 400473,
          "author_name": "Andrei Volodin",
          "author_url": "",
          "post_date": "2018-10-08T11:36:24.117000",
          "content": "<p>the upload has been finished</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 382616,
      "author_name": "Fabien Delattre",
      "author_url": "",
      "post_date": "2018-09-06T18:11:39.820000",
      "content": "<p>Thank you for this competition !</p>\n\n<p>PS: The Wikipedia text data <a href=\"https://dumps.wikimedia.org/enwiki/20180720/enwiki-20180720-pages-articles-multistream.xml.bz2&amp;sa=D&amp;source=hangouts&amp;ust=1535834316553000&amp;usg=AFQjCNFkMs_tSMZOTR2OLCDAm9PVJS9LKw\">link</a> in the data tab doesn't work.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 382620,
          "author_name": "Sohier Dane",
          "author_url": "",
          "post_date": "2018-09-06T18:13:27.150000",
          "content": "<p>Thanks for flagging that; the link should be working now.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 383048,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-09-07T15:45:52.703000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 383056,
          "author_name": "D. Sculley",
          "author_url": "",
          "post_date": "2018-09-07T15:55:50.807000",
          "content": "<p>Hi Andrei, if the link <a href=\"https://www.kaggle.com/c/inclusive-images-challenge#Data-Download-&amp;-Getting-Started\">https://www.kaggle.com/c/inclusive-images-challenge#Data-Download-&amp;-Getting-Started</a> doesn't work for your browser, go to the overview page and hit \"Data Download &amp; Getting Started\" on the left sidebar.  This contains instructions for downloading the correct OpenImages files (also called the correct subset), which are outlined in red on the screenshots on that page.  Hope this helps!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 383061,
          "author_name": "Andrei Volodin",
          "author_url": "",
          "post_date": "2018-09-07T16:00:00.243000",
          "content": "<p>got it</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 383066,
          "author_name": "Andrei Volodin",
          "author_url": "",
          "post_date": "2018-09-07T16:04:14.307000",
          "content": "<p>It appears that the open images website is <a href=\"https://storage.googleapis.com/openimages/web/download.html\">https://storage.googleapis.com/openimages/web/download.html</a></p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 383072,
          "author_name": "Julia Elliott",
          "author_url": "",
          "post_date": "2018-09-07T16:20:41.560000",
          "content": "<p>@Andrei Volodin - Apologies, it looks as though I missed a URL link in those instructions. Yes, that is the correct link. I have updated the instructions accordingly.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 383083,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-09-07T16:41:06.663000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 383221,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-09-08T03:49:41.647000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 383224,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-09-08T03:57:22.447000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 400472,
          "author_name": "Andrei Volodin",
          "author_url": "",
          "post_date": "2018-10-08T11:35:23.130000",
          "content": "<p>@Julia, the dataset [512gb] has been mirrored to gs://inclusive-images_challenge that is accessible for public. You may like to add the link to the data download section ( though it will be available just for 2 month)</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 400666,
          "author_name": "Andrei Volodin",
          "author_url": "",
          "post_date": "2018-10-08T17:40:02.507000",
          "content": "<p>1743042 is the number of files in the folder</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 382158,
      "author_name": "Allen",
      "author_url": "",
      "post_date": "2018-09-05T20:09:58.553000",
      "content": "<p>Seems like a very interesting competition. Thank you for hosting.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 382137,
      "author_name": "Miha Skalic",
      "author_url": "",
      "post_date": "2018-09-05T19:14:34.667000",
      "content": "<p>Hi James and Co.,</p>\n\n<p>first of thanks for hosting this competition. I'm looking forward to all the wacky solutions Kagglers will come up with.</p>\n\n<p>My main question is about NIPS presentation - given how highly contested the tickets are. Could you let us know for how many participants you have reserved tickets, and what will be the distribution?</p>\n\n<p>best, \nMiha</p>",
      "votes": 0,
      "replies": [
        {
          "id": 382172,
          "author_name": "D. Sculley",
          "author_url": "",
          "post_date": "2018-09-05T20:38:29.507000",
          "content": "<p>Hi Miha, thanks for the interest!  As mentioned in the FAQ, we'll be able to provide registration (and some travel expenses) for one representative from each of the top five teams for the workshop component at NIPS.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 382198,
          "author_name": "Miha Skalic",
          "author_url": "",
          "post_date": "2018-09-05T21:11:11.510000",
          "content": "<p>thanks for the answer Sculley! \nand it's \"only\" for the workshops, isn't it? :/</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 424363,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-11-20T01:30:39.350000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "382123": "Hi, we’re the host team, and on behalf of Google AI, we’re excited to welcome you to the Inclusive Images competition! \n\nThe main goal of this research competition is to improve how image classifiers can be made to perform well on unseen geographic distributions. The challenge datasets that we provide come from a very different geographic distribution than the Open Images dataset that you will be training on, thus stress-testing your learned classifier’s ability to generalize to new environments and distributions.\n\nIn the spirit of the competition, we ask that competitors read the rules carefully and follow them closely. In particular, this competition has strict rules on the allowed data sources, which have been designed to help make sure that the winning solutions are focusing on the algorithmic portion of this research challenge.\n\nThe last thing we want to mention is that this is an *inclusive* competition, and we’ll ask that usage of these forums reflect this spirit.  Please take care to communicate respectfully with your colleagues and co-competitors on this challenge.\n\nWe are very excited to host this competition and to see what you are able to come up with!\n\n-- James,\non behalf of Yoni, Pallavi, Tulsee, D., and the rest of the host team\n",
    "383776": "Hi,  @James Atwood\n\nWe are not allowed to update the model submitted at the end of phase1. We can only use the submitted model and code to generate the final prediction for stage2. Is this a correct interpretation of the rule?",
    "382597": "I am just wondering if there will  be google cloud platform promo credits available .\nIn my opinion, it would be convenient to utilize them within the context of the investigation of the AI application in the discourse of the competition.",
    "409485": "Hi @James and Co.,\n\nI have some questions about data...\n\nTrain:\nIn 000002b66c9c498e.jpg, the people are labeled \"hair, mouth, arm, head and so on\"..(train_human_labels.csv)\n\nTest:\nIn 2b2f44594449326f4e52553d.jpg, the people are labeled \"people and shopkeeper\"(tuning_labels.csv)\n\nIf my model predicts 2b2f44594449326f4e52553d.jpg, it will give labels about \"people,shopkeeper, head, hair and so on \". Is this wrong?",
    "415647": "Hello~\n\nOne thing still confuses me is that, is the 1000 tuning set allowed to be used to train our models? The first line on the data page (https://www.kaggle.com/c/inclusive-images-challenge/data) says `This competition uses a portion of the Open Images dataset as the training set` which to me means the tuning set it is not allowed in training. \n\nLook at the majority of participants in stage 1, they are using something similar to: https://www.kaggle.com/gpreda/last-day-6-lines-baseline-compact-anton-petrov-s , which essentially is a simple model using only the tuning set. I wouldn't consider it legit according to my interpretation of the rule. \n\nIn my opinion the tuning set caused more confusion than if it is not provided at all. ",
    "624131": "Hello,\nI have a question regarding the submission of models.\nI have been introduced to the open images challenge lately and I was working to train a model.\nHowever due to lack of time and my late knowledge, I might not be able to submit before the deadline or even submit a good model.\nIs submission still available after the deadline, in case someone just wants to check how good his model performs on the test set ?\nNote: I dont intend to participate conpetitvely ,I just wanted to work on it to see how well my own model will do.",
    "411794": "Hi James and team,\n\nAre we allowed to use other kind of data like a contextual based not images.\nword hierachy to accomodate a much general probalistic model.\n\nCheers,\nPrasen\n",
    "407975": "Hi @James, \n\nI have a question.\n\nFrom the Inclusive Images FAQ\n\n&gt; Note that, per the competition rules, any other form of external data\n&gt; or pre-trained models is not permitted. Competitors are strictly\n&gt; prohibited from using other images or data outside of what has been\n&gt; explicitly approved.\n\nDoes this mean these csv files are considered external data and not permitted to use?\n\n- https://storage.googleapis.com/openimages/2018_04/class-descriptions-boxable.csv\n- https://storage.googleapis.com/openimages/2018_04/bbox_labels_600_hierarchy.json\n\n(from this link https://storage.googleapis.com/openimages/web/download.html)",
    "398213": "Hi, \nI wanted to know if this competition will allow late submissions after the stage 1 deadline. I understand that this will not count towards leaderboard scores or the competition results. But if I wanted to simply evaluate my scores, will that be possible?",
    "388753": "Hi @James,  \nWhat number of classes is used to calculate leaderboard?",
    "383193": "@julia, a very novice question - I am facing issues in transferring training data from the S3 bucket to Google Cloud Storage Bucket. What Access Key ID and Secret Access Key to use? ",
    "382616": "Thank you for this competition !\n\nPS: The Wikipedia text data [link][1] in the data tab doesn't work.\n\n\n  [1]: https://dumps.wikimedia.org/enwiki/20180720/enwiki-20180720-pages-articles-multistream.xml.bz2&amp;sa=D&amp;source=hangouts&amp;ust=1535834316553000&amp;usg=AFQjCNFkMs_tSMZOTR2OLCDAm9PVJS9LKw",
    "382158": "Seems like a very interesting competition. Thank you for hosting.",
    "382137": "Hi James and Co.,\n\nfirst of thanks for hosting this competition. I'm looking forward to all the wacky solutions Kagglers will come up with.\n\nMy main question is about NIPS presentation - given how highly contested the tickets are. Could you let us know for how many participants you have reserved tickets, and what will be the distribution?\n\nbest, \nMiha",
    "424363": ""
  }
}