{
  "id": 95372,
  "title": "Attributes classification solution",
  "url": "/competitions/imaterialist-fashion-2019-FGVC6/writeups/itseezio-nn-attributes-classification-solution",
  "author_name": "",
  "post_date": "2019-06-12T07:14:30.663Z",
  "votes": 3,
  "comment_count": 4,
  "views": 0,
  "content": "<p>Hi all!</p>\n\n<p>We've tried to do attributes classification. The first attempts were with 92-way classifier, but later switched to thechnique from image captioning, which generates next attribute, using image features and knowledge about previous attributes. We've used implementation of the paper <a href=\"https://arxiv.org/pdf/1502.03044.pdf\">\"Show, Attend and Tell: Neural Image Caption Generation with Visual Attention\"</a> from <a href=\"https://github.com/sgrvinod/a-PyTorch-Tutorial-to-Image-Captioning\">this</a> repo. Below is some findings and visualizations.\n<img src=\"http://i68.tinypic.com/rvfrtl.png\" alt=\"t-shirt proper\">\nIt works somehow. Above are the proper attributes for t-shirt. And this with one incorrect attribute, however it looks ambiguous (but I'm not an expert in this issue =)):\n<img src=\"http://i66.tinypic.com/ta2emv.png\" alt=\"t-short ambiguous\">\nSome annotations missed the length of t-shirt, so if model predicts it, this turns to false positive:\n<img src=\"http://i67.tinypic.com/33ngcio.jpg\" alt=\"t-shirt missed length\">\nSame for shorts and other classes. Model can predict correctly the majority of atributes:\n<img src=\"http://i68.tinypic.com/2r6ehc5.png\" alt=\"shorts proper\">\nHowever there are plenty of images with 1 wrongly predicted attribute, so all entry goes to false positive. But they are ambiguous, compare the shorts length on the first two images and style on the last two (frayed in annotations and no/washed in predictions):\n<img src=\"http://i64.tinypic.com/2882ov5.png\" alt=\"mislabeled shorts\"></p>\n\n<p>So with better metric on attributes this challenge can be more interesting.</p>",
  "messages": [
    {
      "id": "550611",
      "postDate": "06/11/2019 20:07:38",
      "content": "<p>Hi all!</p>\n\n<p>We've tried to do attributes classification. The first attempts were with 92-way classifier, but later switched to thechnique from image captioning, which generates next attribute, using image features and knowledge about previous attributes. We've used implementation of the paper <a href=\"https://arxiv.org/pdf/1502.03044.pdf\">\"Show, Attend and Tell: Neural Image Caption Generation with Visual Attention\"</a> from <a href=\"https://github.com/sgrvinod/a-PyTorch-Tutorial-to-Image-Captioning\">this</a> repo. Below is some findings and visualizations.\n<img src=\"http://i68.tinypic.com/rvfrtl.png\" alt=\"t-shirt proper\">\nIt works somehow. Above are the proper attributes for t-shirt. And this with one incorrect attribute, however it looks ambiguous (but I'm not an expert in this issue =)):\n<img src=\"http://i66.tinypic.com/ta2emv.png\" alt=\"t-short ambiguous\">\nSome annotations missed the length of t-shirt, so if model predicts it, this turns to false positive:\n<img src=\"http://i67.tinypic.com/33ngcio.jpg\" alt=\"t-shirt missed length\">\nSame for shorts and other classes. Model can predict correctly the majority of atributes:\n<img src=\"http://i68.tinypic.com/2r6ehc5.png\" alt=\"shorts proper\">\nHowever there are plenty of images with 1 wrongly predicted attribute, so all entry goes to false positive. But they are ambiguous, compare the shorts length on the first two images and style on the last two (frayed in annotations and no/washed in predictions):\n<img src=\"http://i64.tinypic.com/2882ov5.png\" alt=\"mislabeled shorts\"></p>\n\n<p>So with better metric on attributes this challenge can be more interesting.</p>",
      "rawMarkdown": "Hi all!\n\nWe've tried to do attributes classification. The first attempts were with 92-way classifier, but later switched to thechnique from image captioning, which generates next attribute, using image features and knowledge about previous attributes. We've used implementation of the paper [\"Show, Attend and Tell: Neural Image Caption Generation with Visual Attention\"](https://arxiv.org/pdf/1502.03044.pdf) from [this](https://github.com/sgrvinod/a-PyTorch-Tutorial-to-Image-Captioning) repo. Below is some findings and visualizations.\n![t-shirt proper](http://i68.tinypic.com/rvfrtl.png)\nIt works somehow. Above are the proper attributes for t-shirt. And this with one incorrect attribute, however it looks ambiguous (but I'm not an expert in this issue =)):\n![t-short ambiguous](http://i66.tinypic.com/ta2emv.png)\nSome annotations missed the length of t-shirt, so if model predicts it, this turns to false positive:\n![t-shirt missed length](http://i67.tinypic.com/33ngcio.jpg)\nSame for shorts and other classes. Model can predict correctly the majority of atributes:\n![shorts proper](http://i68.tinypic.com/2r6ehc5.png)\nHowever there are plenty of images with 1 wrongly predicted attribute, so all entry goes to false positive. But they are ambiguous, compare the shorts length on the first two images and style on the last two (frayed in annotations and no/washed in predictions):\n![mislabeled shorts](http://i64.tinypic.com/2882ov5.png)\n\nSo with better metric on attributes this challenge can be more interesting.",
      "votes": null
    },
    {
      "id": "551760",
      "postDate": "06/13/2019 04:55:48",
      "content": "<p>Interesting approach! Did you use the pretrained model and its word_map or you trained from scratch. How much time you ran per epoch?</p>",
      "rawMarkdown": "Interesting approach! Did you use the pretrained model and its word_map or you trained from scratch. How much time you ran per epoch?",
      "votes": null
    },
    {
      "id": "551803",
      "postDate": "06/13/2019 06:10:05",
      "content": "<p>It was trained from scratch. The training is really fast, about 1 minute per epoch (the dataset with attributes is not big, about 1k samples per class).</p>",
      "rawMarkdown": "It was trained from scratch. The training is really fast, about 1 minute per epoch (the dataset with attributes is not big, about 1k samples per class).",
      "votes": null
    },
    {
      "id": "551809",
      "postDate": "06/13/2019 06:22:50",
      "content": "<p>That's nice. So your training set did not include images that have the same class ids but without attributes. Did you have to run an additional model before to determine which samples will have attributes?</p>",
      "rawMarkdown": "That's nice. So your training set did not include images that have the same class ids but without attributes. Did you have to run an additional model before to determine which samples will have attributes?",
      "votes": null
    },
    {
      "id": "552200",
      "postDate": "06/13/2019 16:10:34",
      "content": "<p>For the train set if apparel label is just in range [0-12], then attributes for such apparel was not labeled. So I have taken those, wich has attributes, they starts from 0_ ... 12_.</p>",
      "rawMarkdown": "For the train set if apparel label is just in range [0-12], then attributes for such apparel was not labeled. So I have taken those, wich has attributes, they starts from 0_ ... 12_.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 551760,
      "author_name": "ntnguyen1234",
      "author_url": "",
      "post_date": "06/13/2019 04:55:48",
      "content": "<p>Interesting approach! Did you use the pretrained model and its word_map or you trained from scratch. How much time you ran per epoch?</p>",
      "votes": null,
      "replies": [
        {
          "id": 551803,
          "author_name": "harshml",
          "author_url": "",
          "post_date": "06/13/2019 06:10:05",
          "content": "<p>It was trained from scratch. The training is really fast, about 1 minute per epoch (the dataset with attributes is not big, about 1k samples per class).</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 551809,
          "author_name": "ntnguyen1234",
          "author_url": "",
          "post_date": "06/13/2019 06:22:50",
          "content": "<p>That's nice. So your training set did not include images that have the same class ids but without attributes. Did you have to run an additional model before to determine which samples will have attributes?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 552200,
          "author_name": "harshml",
          "author_url": "",
          "post_date": "06/13/2019 16:10:34",
          "content": "<p>For the train set if apparel label is just in range [0-12], then attributes for such apparel was not labeled. So I have taken those, wich has attributes, they starts from 0_ ... 12_.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "550611": "Hi all!\n\nWe've tried to do attributes classification. The first attempts were with 92-way classifier, but later switched to thechnique from image captioning, which generates next attribute, using image features and knowledge about previous attributes. We've used implementation of the paper [\"Show, Attend and Tell: Neural Image Caption Generation with Visual Attention\"](https://arxiv.org/pdf/1502.03044.pdf) from [this](https://github.com/sgrvinod/a-PyTorch-Tutorial-to-Image-Captioning) repo. Below is some findings and visualizations.\n![t-shirt proper](http://i68.tinypic.com/rvfrtl.png)\nIt works somehow. Above are the proper attributes for t-shirt. And this with one incorrect attribute, however it looks ambiguous (but I'm not an expert in this issue =)):\n![t-short ambiguous](http://i66.tinypic.com/ta2emv.png)\nSome annotations missed the length of t-shirt, so if model predicts it, this turns to false positive:\n![t-shirt missed length](http://i67.tinypic.com/33ngcio.jpg)\nSame for shorts and other classes. Model can predict correctly the majority of atributes:\n![shorts proper](http://i68.tinypic.com/2r6ehc5.png)\nHowever there are plenty of images with 1 wrongly predicted attribute, so all entry goes to false positive. But they are ambiguous, compare the shorts length on the first two images and style on the last two (frayed in annotations and no/washed in predictions):\n![mislabeled shorts](http://i64.tinypic.com/2882ov5.png)\n\nSo with better metric on attributes this challenge can be more interesting.",
    "551760": "Interesting approach! Did you use the pretrained model and its word_map or you trained from scratch. How much time you ran per epoch?",
    "551803": "It was trained from scratch. The training is really fast, about 1 minute per epoch (the dataset with attributes is not big, about 1k samples per class).",
    "551809": "That's nice. So your training set did not include images that have the same class ids but without attributes. Did you have to run an additional model before to determine which samples will have attributes?",
    "552200": "For the train set if apparel label is just in range [0-12], then attributes for such apparel was not labeled. So I have taken those, wich has attributes, they starts from 0_ ... 12_."
  },
  "source": "meta"
}