{
  "id": 41123,
  "title": "Classifying products with multiple images",
  "url": "/competitions/cdiscount-image-classification-challenge/discussion/41123",
  "author_name": "",
  "post_date": "2017-10-13T04:22:59.160199300Z",
  "votes": 2,
  "comment_count": 6,
  "views": 0,
  "content": "<p>What is the best way of classifying products that have multiple images?</p>\n\n<p>I am trying to go beyond training a network on single image/product pairs and then combining multiple predictions at the testing stage.</p>\n\n<p>As a start, I am trying to design a trainable network that in takes four images, passes them through the same convolutional layers (or different ones that share the same weights) and then adds the final layers and passes the result to a softmax.</p>\n\n<p>But this is beyond my experience and I can't work out how to do it. I have been googling for ideas, but I am too ignorant to find the right keywords.</p>\n\n<p>Anyone got an bright ideas or papers to read? </p>\n\n<p>This must be a well trodden path.</p>",
  "messages": [
    {
      "id": "230902",
      "postDate": "10/13/2017 04:22:59",
      "content": "<p>What is the best way of classifying products that have multiple images?</p>\n\n<p>I am trying to go beyond training a network on single image/product pairs and then combining multiple predictions at the testing stage.</p>\n\n<p>As a start, I am trying to design a trainable network that in takes four images, passes them through the same convolutional layers (or different ones that share the same weights) and then adds the final layers and passes the result to a softmax.</p>\n\n<p>But this is beyond my experience and I can't work out how to do it. I have been googling for ideas, but I am too ignorant to find the right keywords.</p>\n\n<p>Anyone got an bright ideas or papers to read? </p>\n\n<p>This must be a well trodden path.</p>",
      "rawMarkdown": "What is the best way of classifying products that have multiple images?\n\nI am trying to go beyond training a network on single image/product pairs and then combining multiple predictions at the testing stage.\n\nAs a start, I am trying to design a trainable network that in takes four images, passes them through the same convolutional layers (or different ones that share the same weights) and then adds the final layers and passes the result to a softmax.\n\nBut this is beyond my experience and I can't work out how to do it. I have been googling for ideas, but I am too ignorant to find the right keywords.\n\nAnyone got an bright ideas or papers to read? \n\nThis must be a well trodden path.",
      "votes": null
    },
    {
      "id": "230932",
      "postDate": "10/13/2017 07:05:37",
      "content": "<p>\"i am trying to design a trainable network that in takes four images ...\"</p>\n\n<p>this is more complicated then you think. </p>\n\n<ol>\n<li>The input no. of images can vary from 1 to 4. </li>\n<li>The order (permutation) of the input should not matter. </li>\n</ol>\n\n<p>Your network must take care of that. It can be done by LSTM but i don't think you want to do this unless you are familiar with deep networks (given this huge amount of train data, especially if you are doing end-to-end )</p>\n\n<p>A better idea is some smart post-processing of the results.</p>\n\n<p>And I suggest collect some experimental results of running a network on different images of the same product. Make some statistics to see how these \"scores differ on different images of the same product\" and from that you can think of good post processing methods  (which can be some learning network or simple methods)</p>",
      "rawMarkdown": "\"i am trying to design a trainable network that in takes four images ...\"\n\nthis is more complicated then you think. \n\n1.  The input no. of images can vary from 1 to 4. \n2. The order (permutation) of the input should not matter. \n\nYour network must take care of that. It can be done by LSTM but i don't think you want to do this unless you are familiar with deep networks (given this huge amount of train data, especially if you are doing end-to-end )\n\nA better idea is some smart post-processing of the results.\n\nAnd I suggest collect some experimental results of running a network on different images of the same product. Make some statistics to see how these \"scores differ on different images of the same product\" and from that you can think of good post processing methods  (which can be some learning network or simple methods)",
      "votes": null
    },
    {
      "id": "230934",
      "postDate": "10/13/2017 07:36:56",
      "content": "<p>Google for multi instance classification or image set classification </p>\n\n<p>Abstract. There are classification tasks that take as inputs groups of images rather than single images. In order to address such situations, we introduce a nested multi-instance deep ...</p>",
      "rawMarkdown": "Google for multi instance classification or image set classification \n\nAbstract. There are classification tasks that take as inputs groups of images rather than single images. In order to address such situations, we introduce a nested multi-instance deep ...",
      "votes": null
    },
    {
      "id": "230936",
      "postDate": "10/13/2017 07:44:15",
      "content": "<p>That's very interesting advice. Many thanks. I shall research the data before trying anything too complicated.</p>",
      "rawMarkdown": "That's very interesting advice. Many thanks. I shall research the data before trying anything too complicated.",
      "votes": null
    },
    {
      "id": "230982",
      "postDate": "10/13/2017 10:07:32",
      "content": "<p>Hello James, </p>\n\n<p>you might want to have a look at work of Miech et. al from previous Kaggle competition - <a href=\"https://arxiv.org/pdf/1706.06905.pdf\">https://arxiv.org/pdf/1706.06905.pdf</a></p>\n\n<p>There are several pooling options described and for more details you can follow the references.</p>",
      "rawMarkdown": "Hello James, \n\nyou might want to have a look at work of Miech et. al from previous Kaggle competition - https://arxiv.org/pdf/1706.06905.pdf\n\nThere are several pooling options described and for more details you can follow the references.",
      "votes": null
    },
    {
      "id": "231428",
      "postDate": "10/14/2017 20:10:26",
      "content": "<p>maybe this can be a simple solution:</p>\n\n<ol>\n<li><p>given N images of a product, apply a feature extraction net on each image_n to get feature_n.</p></li>\n<li><p>apply symmetrical function: f = sym_func(feature_1, feature_2 ...).  A symmetrical function is one that does not depend on the order of the input. examples are max pooling or average pooling.</p></li>\n<li><p>apply classifier on f</p></li>\n</ol>\n\n<p>steps 1,2,3 can be train end-to-end as a single network. for reference, refer to:</p>\n\n<p>\"PointNet: Deep Learning on Point Sets for 3D Classification and Segmentation\"</p>\n\n<p>\"Multi-view Convolutional Neural Networks for 3D Shape Recognition\"</p>",
      "rawMarkdown": "maybe this can be a simple solution:\n\n1. given N images of a product, apply a feature extraction net on each image_n to get feature_n.\n\n2. apply symmetrical function: f = sym_func(feature_1, feature_2 ...).  A symmetrical function is one that does not depend on the order of the input. examples are max pooling or average pooling.\n\n3. apply classifier on f\n\nsteps 1,2,3 can be train end-to-end as a single network. for reference, refer to:\n\n\"PointNet: Deep Learning on Point Sets for 3D Classification and Segmentation\"\n\n\"Multi-view Convolutional Neural Networks for 3D Shape Recognition\"",
      "votes": null
    },
    {
      "id": "244705",
      "postDate": "11/16/2017 19:00:52",
      "content": "<p>Wondering why can not it be arranged into multiple rows? If there N images of a product, there will be N rows for the same product and category combination. This may increase processing time, however algorithm and network will be simpler. Is not it?</p>",
      "rawMarkdown": "Wondering why can not it be arranged into multiple rows? If there N images of a product, there will be N rows for the same product and category combination. This may increase processing time, however algorithm and network will be simpler. Is not it?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 230932,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "10/13/2017 07:05:37",
      "content": "<p>\"i am trying to design a trainable network that in takes four images ...\"</p>\n\n<p>this is more complicated then you think. </p>\n\n<ol>\n<li>The input no. of images can vary from 1 to 4. </li>\n<li>The order (permutation) of the input should not matter. </li>\n</ol>\n\n<p>Your network must take care of that. It can be done by LSTM but i don't think you want to do this unless you are familiar with deep networks (given this huge amount of train data, especially if you are doing end-to-end )</p>\n\n<p>A better idea is some smart post-processing of the results.</p>\n\n<p>And I suggest collect some experimental results of running a network on different images of the same product. Make some statistics to see how these \"scores differ on different images of the same product\" and from that you can think of good post processing methods  (which can be some learning network or simple methods)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 230934,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "10/13/2017 07:36:56",
      "content": "<p>Google for multi instance classification or image set classification </p>\n\n<p>Abstract. There are classification tasks that take as inputs groups of images rather than single images. In order to address such situations, we introduce a nested multi-instance deep ...</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 230936,
      "author_name": "jinkos",
      "author_url": "",
      "post_date": "10/13/2017 07:44:15",
      "content": "<p>That's very interesting advice. Many thanks. I shall research the data before trying anything too complicated.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 230982,
      "author_name": "mihaskalic",
      "author_url": "",
      "post_date": "10/13/2017 10:07:32",
      "content": "<p>Hello James, </p>\n\n<p>you might want to have a look at work of Miech et. al from previous Kaggle competition - <a href=\"https://arxiv.org/pdf/1706.06905.pdf\">https://arxiv.org/pdf/1706.06905.pdf</a></p>\n\n<p>There are several pooling options described and for more details you can follow the references.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 231428,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "10/14/2017 20:10:26",
      "content": "<p>maybe this can be a simple solution:</p>\n\n<ol>\n<li><p>given N images of a product, apply a feature extraction net on each image_n to get feature_n.</p></li>\n<li><p>apply symmetrical function: f = sym_func(feature_1, feature_2 ...).  A symmetrical function is one that does not depend on the order of the input. examples are max pooling or average pooling.</p></li>\n<li><p>apply classifier on f</p></li>\n</ol>\n\n<p>steps 1,2,3 can be train end-to-end as a single network. for reference, refer to:</p>\n\n<p>\"PointNet: Deep Learning on Point Sets for 3D Classification and Segmentation\"</p>\n\n<p>\"Multi-view Convolutional Neural Networks for 3D Shape Recognition\"</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 244705,
      "author_name": "ppujari",
      "author_url": "",
      "post_date": "11/16/2017 19:00:52",
      "content": "<p>Wondering why can not it be arranged into multiple rows? If there N images of a product, there will be N rows for the same product and category combination. This may increase processing time, however algorithm and network will be simpler. Is not it?</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "230902": "What is the best way of classifying products that have multiple images?\n\nI am trying to go beyond training a network on single image/product pairs and then combining multiple predictions at the testing stage.\n\nAs a start, I am trying to design a trainable network that in takes four images, passes them through the same convolutional layers (or different ones that share the same weights) and then adds the final layers and passes the result to a softmax.\n\nBut this is beyond my experience and I can't work out how to do it. I have been googling for ideas, but I am too ignorant to find the right keywords.\n\nAnyone got an bright ideas or papers to read? \n\nThis must be a well trodden path.",
    "230932": "\"i am trying to design a trainable network that in takes four images ...\"\n\nthis is more complicated then you think. \n\n1.  The input no. of images can vary from 1 to 4. \n2. The order (permutation) of the input should not matter. \n\nYour network must take care of that. It can be done by LSTM but i don't think you want to do this unless you are familiar with deep networks (given this huge amount of train data, especially if you are doing end-to-end )\n\nA better idea is some smart post-processing of the results.\n\nAnd I suggest collect some experimental results of running a network on different images of the same product. Make some statistics to see how these \"scores differ on different images of the same product\" and from that you can think of good post processing methods  (which can be some learning network or simple methods)",
    "230934": "Google for multi instance classification or image set classification \n\nAbstract. There are classification tasks that take as inputs groups of images rather than single images. In order to address such situations, we introduce a nested multi-instance deep ...",
    "230936": "That's very interesting advice. Many thanks. I shall research the data before trying anything too complicated.",
    "230982": "Hello James, \n\nyou might want to have a look at work of Miech et. al from previous Kaggle competition - https://arxiv.org/pdf/1706.06905.pdf\n\nThere are several pooling options described and for more details you can follow the references.",
    "231428": "maybe this can be a simple solution:\n\n1. given N images of a product, apply a feature extraction net on each image_n to get feature_n.\n\n2. apply symmetrical function: f = sym_func(feature_1, feature_2 ...).  A symmetrical function is one that does not depend on the order of the input. examples are max pooling or average pooling.\n\n3. apply classifier on f\n\nsteps 1,2,3 can be train end-to-end as a single network. for reference, refer to:\n\n\"PointNet: Deep Learning on Point Sets for 3D Classification and Segmentation\"\n\n\"Multi-view Convolutional Neural Networks for 3D Shape Recognition\"",
    "244705": "Wondering why can not it be arranged into multiple rows? If there N images of a product, there will be N rows for the same product and category combination. This may increase processing time, however algorithm and network will be simpler. Is not it?"
  },
  "source": "meta"
}