{
  "id": 94811,
  "title": "Notes for Attributes Annotations",
  "url": "/competitions/imaterialist-fashion-2019-FGVC6/discussion/94811",
  "author_name": "",
  "post_date": "2019-06-07T04:22:45.712474100Z",
  "votes": 8,
  "comment_count": 6,
  "views": 0,
  "content": "<p>Thank you guys so much for your interest and participation of this challenge! We have received several questions about annotation procedures of apparel attributes. I will summarize below and hope this is helpful:</p>\n\n<ol>\n<li><p>There are 3 masks that belong to apparel category 27, 28, 33 (apparel parts) that has attributes. This is a human annotation error that we did not catch before releasing this dataset. There are no attributes for the categories larger than 12 for the test set. </p></li>\n<li><p>During the fine-grained attribute annotation process, fashion expert was given a mask and a list of possible options for one super-category of attributes (say \"length\"). Among those possible options for \"length\", we have another 2 options \"not sure\", \"not on the list\". If the annotators chose \"not sure\" or \"not on the list\", we remove this attribute in the released version. That why some masks do not have attributes belong to \"length\". Since it's hard for fashion expert to tell what is the right \"length\" for these masks (one common reason is due to occlusion).</p></li>\n</ol>\n\n<p>Please let me know if you have other questions. You guys' comments are really valuable for us to make this dataset better!</p>",
  "messages": [
    {
      "id": "546966",
      "postDate": "06/07/2019 04:22:45",
      "content": "<p>Thank you guys so much for your interest and participation of this challenge! We have received several questions about annotation procedures of apparel attributes. I will summarize below and hope this is helpful:</p>\n\n<ol>\n<li><p>There are 3 masks that belong to apparel category 27, 28, 33 (apparel parts) that has attributes. This is a human annotation error that we did not catch before releasing this dataset. There are no attributes for the categories larger than 12 for the test set. </p></li>\n<li><p>During the fine-grained attribute annotation process, fashion expert was given a mask and a list of possible options for one super-category of attributes (say \"length\"). Among those possible options for \"length\", we have another 2 options \"not sure\", \"not on the list\". If the annotators chose \"not sure\" or \"not on the list\", we remove this attribute in the released version. That why some masks do not have attributes belong to \"length\". Since it's hard for fashion expert to tell what is the right \"length\" for these masks (one common reason is due to occlusion).</p></li>\n</ol>\n\n<p>Please let me know if you have other questions. You guys' comments are really valuable for us to make this dataset better!</p>",
      "rawMarkdown": "Thank you guys so much for your interest and participation of this challenge! We have received several questions about annotation procedures of apparel attributes. I will summarize below and hope this is helpful:\n\n1. There are 3 masks that belong to apparel category 27, 28, 33 (apparel parts) that has attributes. This is a human annotation error that we did not catch before releasing this dataset. There are no attributes for the categories larger than 12 for the test set. \n\n2.  During the fine-grained attribute annotation process, fashion expert was given a mask and a list of possible options for one super-category of attributes (say \"length\"). Among those possible options for \"length\", we have another 2 options \"not sure\", \"not on the list\". If the annotators chose \"not sure\" or \"not on the list\", we remove this attribute in the released version. That why some masks do not have attributes belong to \"length\". Since it's hard for fashion expert to tell what is the right \"length\" for these masks (one common reason is due to occlusion).\n\nPlease let me know if you have other questions. You guys' comments are really valuable for us to make this dataset better!",
      "votes": null
    },
    {
      "id": "547195",
      "postDate": "06/07/2019 11:44:51",
      "content": "<p>I beleieve the metric must be changed. It must account not the exact match but at least close match for attributes. For example if we have attributes:<code>3_4_7_8_10_11_45</code> if we predict <code>3_7_8_10_11_45</code> now we will get zero score. Which is actualy bad. We should recieve something like <code>5/6</code> of obtained score for this mask. I propose to use something like Jaccard index as coefficeint for attributes.</p>",
      "rawMarkdown": "I beleieve the metric must be changed. It must account not the exact match but at least close match for attributes. For example if we have attributes:` 3_4_7_8_10_11_45` if we predict `3_7_8_10_11_45` now we will get zero score. Which is actualy bad. We should recieve something like `5/6` of obtained score for this mask. I propose to use something like Jaccard index as coefficeint for attributes.",
      "votes": null
    },
    {
      "id": "547498",
      "postDate": "06/07/2019 19:58:04",
      "content": "<p><a href=\"/makeitworkjml\">@makeitworkjml</a> plus one to ZFTurbo. I am not a fashion expert, however, I am an artist and finished an art school, and I just do not agree with some of the human labels. Getting all of them right when the labelling itself is very human-ambiguous makes such metric highly unreasonable and ruins the concept of this part of the competition. \"not sure\"  is highly personal, getting zero score for one \"not sure\" in the test is unreasonable</p>",
      "rawMarkdown": "makeitworkjml plus one to ZFTurbo. I am not a fashion expert, however, I am an artist and finished an art school, and I just do not agree with some of the human labels. Getting all of them right when the labelling itself is very human-ambiguous makes such metric highly unreasonable and ruins the concept of this part of the competition. \"not sure\"  is highly personal, getting zero score for one \"not sure\" in the test is unreasonable",
      "votes": null
    },
    {
      "id": "547555",
      "postDate": "06/07/2019 21:21:30",
      "content": "<p><a href=\"/blondinka\">@blondinka</a> <a href=\"/zfturbo\">@zfturbo</a> In this year's competition, we decided to go with a challenging metric, where to treat mask with category and attributes as one really fine-grained label. Your concerns are noted. And I agree with <a href=\"/zfturbo\">@zfturbo</a> that this metric is not rewarding partially correct predicted attributes. We will certainly think about how to improve this metric for next year. </p>",
      "rawMarkdown": "blondinka @zfturbo In this year's competition, we decided to go with a challenging metric, where to treat mask with category and attributes as one really fine-grained label. Your concerns are noted. And I agree with @zfturbo that this metric is not rewarding partially correct predicted attributes. We will certainly think about how to improve this metric for next year.",
      "votes": null
    },
    {
      "id": "547556",
      "postDate": "06/07/2019 21:21:43",
      "content": "<p><a href=\"/blondinka\">@blondinka</a> , could you give me an example, of human labels that you do not agree with? </p>",
      "rawMarkdown": "blondinka , could you give me an example, of human labels that you do not agree with?",
      "votes": null
    },
    {
      "id": "547909",
      "postDate": "06/08/2019 13:53:07",
      "content": "<p>I will, a bit busy now will publish it later, so you'll have them</p>",
      "rawMarkdown": "I will, a bit busy now will publish it later, so you'll have them",
      "votes": null
    },
    {
      "id": "548137",
      "postDate": "06/08/2019 20:37:03",
      "content": "<p><a href=\"/makeitworkjml\">@makeitworkjml</a> Edit distance may help for proper attributes ranking. I don't know with what this challenge ends up, but looks like we won't see interesting ways how to deal with attributes. We've tried techniques from image captioning, but predictions are usually differs in 1-2 positions with gold, so the current metric just decreased. Better metric might give a rise to more interesting and insightful solutions.</p>",
      "rawMarkdown": "makeitworkjml Edit distance may help for proper attributes ranking. I don't know with what this challenge ends up, but looks like we won't see interesting ways how to deal with attributes. We've tried techniques from image captioning, but predictions are usually differs in 1-2 positions with gold, so the current metric just decreased. Better metric might give a rise to more interesting and insightful solutions.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 547195,
      "author_name": "zfturbo",
      "author_url": "",
      "post_date": "06/07/2019 11:44:51",
      "content": "<p>I beleieve the metric must be changed. It must account not the exact match but at least close match for attributes. For example if we have attributes:<code>3_4_7_8_10_11_45</code> if we predict <code>3_7_8_10_11_45</code> now we will get zero score. Which is actualy bad. We should recieve something like <code>5/6</code> of obtained score for this mask. I propose to use something like Jaccard index as coefficeint for attributes.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 547498,
      "author_name": "blondinka",
      "author_url": "",
      "post_date": "06/07/2019 19:58:04",
      "content": "<p><a href=\"/makeitworkjml\">@makeitworkjml</a> plus one to ZFTurbo. I am not a fashion expert, however, I am an artist and finished an art school, and I just do not agree with some of the human labels. Getting all of them right when the labelling itself is very human-ambiguous makes such metric highly unreasonable and ruins the concept of this part of the competition. \"not sure\"  is highly personal, getting zero score for one \"not sure\" in the test is unreasonable</p>",
      "votes": null,
      "replies": [
        {
          "id": 547556,
          "author_name": "makeitworkjml",
          "author_url": "",
          "post_date": "06/07/2019 21:21:43",
          "content": "<p><a href=\"/blondinka\">@blondinka</a> , could you give me an example, of human labels that you do not agree with? </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 547909,
          "author_name": "blondinka",
          "author_url": "",
          "post_date": "06/08/2019 13:53:07",
          "content": "<p>I will, a bit busy now will publish it later, so you'll have them</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 547555,
      "author_name": "makeitworkjml",
      "author_url": "",
      "post_date": "06/07/2019 21:21:30",
      "content": "<p><a href=\"/blondinka\">@blondinka</a> <a href=\"/zfturbo\">@zfturbo</a> In this year's competition, we decided to go with a challenging metric, where to treat mask with category and attributes as one really fine-grained label. Your concerns are noted. And I agree with <a href=\"/zfturbo\">@zfturbo</a> that this metric is not rewarding partially correct predicted attributes. We will certainly think about how to improve this metric for next year. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 548137,
      "author_name": "harshml",
      "author_url": "",
      "post_date": "06/08/2019 20:37:03",
      "content": "<p><a href=\"/makeitworkjml\">@makeitworkjml</a> Edit distance may help for proper attributes ranking. I don't know with what this challenge ends up, but looks like we won't see interesting ways how to deal with attributes. We've tried techniques from image captioning, but predictions are usually differs in 1-2 positions with gold, so the current metric just decreased. Better metric might give a rise to more interesting and insightful solutions.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "546966": "Thank you guys so much for your interest and participation of this challenge! We have received several questions about annotation procedures of apparel attributes. I will summarize below and hope this is helpful:\n\n1. There are 3 masks that belong to apparel category 27, 28, 33 (apparel parts) that has attributes. This is a human annotation error that we did not catch before releasing this dataset. There are no attributes for the categories larger than 12 for the test set. \n\n2.  During the fine-grained attribute annotation process, fashion expert was given a mask and a list of possible options for one super-category of attributes (say \"length\"). Among those possible options for \"length\", we have another 2 options \"not sure\", \"not on the list\". If the annotators chose \"not sure\" or \"not on the list\", we remove this attribute in the released version. That why some masks do not have attributes belong to \"length\". Since it's hard for fashion expert to tell what is the right \"length\" for these masks (one common reason is due to occlusion).\n\nPlease let me know if you have other questions. You guys' comments are really valuable for us to make this dataset better!",
    "547195": "I beleieve the metric must be changed. It must account not the exact match but at least close match for attributes. For example if we have attributes:` 3_4_7_8_10_11_45` if we predict `3_7_8_10_11_45` now we will get zero score. Which is actualy bad. We should recieve something like `5/6` of obtained score for this mask. I propose to use something like Jaccard index as coefficeint for attributes.",
    "547498": "makeitworkjml plus one to ZFTurbo. I am not a fashion expert, however, I am an artist and finished an art school, and I just do not agree with some of the human labels. Getting all of them right when the labelling itself is very human-ambiguous makes such metric highly unreasonable and ruins the concept of this part of the competition. \"not sure\"  is highly personal, getting zero score for one \"not sure\" in the test is unreasonable",
    "547555": "blondinka @zfturbo In this year's competition, we decided to go with a challenging metric, where to treat mask with category and attributes as one really fine-grained label. Your concerns are noted. And I agree with @zfturbo that this metric is not rewarding partially correct predicted attributes. We will certainly think about how to improve this metric for next year.",
    "547556": "blondinka , could you give me an example, of human labels that you do not agree with?",
    "547909": "I will, a bit busy now will publish it later, so you'll have them",
    "548137": "makeitworkjml Edit distance may help for proper attributes ranking. I don't know with what this challenge ends up, but looks like we won't see interesting ways how to deal with attributes. We've tried techniques from image captioning, but predictions are usually differs in 1-2 positions with gold, so the current metric just decreased. Better metric might give a rise to more interesting and insightful solutions."
  },
  "source": "meta"
}