{
  "id": 19723,
  "title": "Technical questions.",
  "url": "/competitions/yelp-restaurant-photo-classification/discussion/19723",
  "author_name": "",
  "post_date": "2016-03-23T04:38:14.803Z",
  "votes": null,
  "comment_count": 3,
  "views": 788,
  "content": "<p>Hello everyone,</p>\n\n<p>I have several technical questions.</p>\n\n<ol>\n<li><p>All images in the train/test have at most 500 pixels per side. 500x500 is a lot. I guess people make it smaller. How much smaller do you guys do so, that it works? 128x128? 256x256?</p></li>\n<li><p>When I use 128x128 I can fit all dataset into my RAM and my video card does not have problems with it also. But let's say I will want to work with bigger images, so that they will not fit. Can anyone share with me examples of code(I would prefer python) where people are able to train neural net without loading all data into memory at the same time.</p></li>\n<li><p>Let's say I want to use this approach:\nMy target is 9 column array, where each row may have 0 or 1, telling if this  image belongs to a restaurant that have these tags. But because restaurant may belong to several categories I am not treating this problem as classification, but rather as a regression problem with images as input and 9 binary outputs per image as an output. My question is: I do this per image. And I would be able to give an estimate(number from 0 to 1) for each image for each tag. But at the end we are interested assigning categories to restaurants and not images. So, what is good way to merge categories per image to categories per restaurant?</p></li>\n</ol>\n\n<p>Vladimir</p>",
  "messages": [
    {
      "id": "112695",
      "postDate": "03/23/2016 04:38:14",
      "content": "<p>Hello everyone,</p>\n\n<p>I have several technical questions.</p>\n\n<ol>\n<li><p>All images in the train/test have at most 500 pixels per side. 500x500 is a lot. I guess people make it smaller. How much smaller do you guys do so, that it works? 128x128? 256x256?</p></li>\n<li><p>When I use 128x128 I can fit all dataset into my RAM and my video card does not have problems with it also. But let's say I will want to work with bigger images, so that they will not fit. Can anyone share with me examples of code(I would prefer python) where people are able to train neural net without loading all data into memory at the same time.</p></li>\n<li><p>Let's say I want to use this approach:\nMy target is 9 column array, where each row may have 0 or 1, telling if this  image belongs to a restaurant that have these tags. But because restaurant may belong to several categories I am not treating this problem as classification, but rather as a regression problem with images as input and 9 binary outputs per image as an output. My question is: I do this per image. And I would be able to give an estimate(number from 0 to 1) for each image for each tag. But at the end we are interested assigning categories to restaurants and not images. So, what is good way to merge categories per image to categories per restaurant?</p></li>\n</ol>\n\n<p>Vladimir</p>",
      "rawMarkdown": "Hello everyone,\r\n\r\nI have several technical questions.\r\n\r\n1. All images in the train/test have at most 500 pixels per side. 500x500 is a lot. I guess people make it smaller. How much smaller do you guys do so, that it works? 128x128? 256x256?\r\n\r\n2.  When I use 128x128 I can fit all dataset into my RAM and my video card does not have problems with it also. But let's say I will want to work with bigger images, so that they will not fit. Can anyone share with me examples of code(I would prefer python) where people are able to train neural net without loading all data into memory at the same time.\r\n\r\n3. Let's say I want to use this approach:\r\nMy target is 9 column array, where each row may have 0 or 1, telling if this  image belongs to a restaurant that have these tags. But because restaurant may belong to several categories I am not treating this problem as classification, but rather as a regression problem with images as input and 9 binary outputs per image as an output. My question is: I do this per image. And I would be able to give an estimate(number from 0 to 1) for each image for each tag. But at the end we are interested assigning categories to restaurants and not images. So, what is good way to merge categories per image to categories per restaurant?\r\n\r\nVladimir",
      "votes": null
    },
    {
      "id": "112746",
      "postDate": "03/23/2016 14:35:46",
      "content": "<p>Hello, Vladimir</p>\n\n<ol>\n<li>I think most of the people uses pre-trained neural networks(or fine-tunned ones) for this competition, because training a network from scratch is very difficult task, so image size depends on structure of the network. Usually it is about 200x200.</li>\n<li>There is no reason for loading all dataset into RAM. All state-of-the-art neural networks for image classification trained in a &quot;batch&quot; mode, when only small part of data loaded in a memory at the same time, and all deep learning libraries can do it. If you prefer python, check out:  <a href=\"https://github.com/BVLC/caffe/tree/master/examples\">Caffe</a>, <a href=\"https://github.com/fchollet/keras/tree/master/examples\">Keras</a>, <a href=\"https://github.com/Lasagne/Recipes\">Lasagne</a>, <a href=\"https://github.com/dmlc/mxnet/tree/master/example/image-classification\">mxnet</a> </li>\n<li>It's <a href=\"https://en.wikipedia.org/wiki/Multiple-instance_learning\">Multiple-instance</a>, <a href=\"https://en.wikipedia.org/wiki/Multi-label_classification\">Multi-label</a> problem. From wiki: &quot;bag is labeled positive if there is at least one instance in it which is positive&quot;, so if you made a regression per image, i think simple thresholding can gives you a decent result.</li>\n</ol>",
      "rawMarkdown": "Hello, Vladimir\r\n\r\n1. I think most of the people uses pre-trained neural networks(or fine-tunned ones) for this competition, because training a network from scratch is very difficult task, so image size depends on structure of the network. Usually it is about 200x200.\r\n2. There is no reason for loading all dataset into RAM. All state-of-the-art neural networks for image classification trained in a \"batch\" mode, when only small part of data loaded in a memory at the same time, and all deep learning libraries can do it. If you prefer python, check out:  [Caffe][1], [Keras][2], [Lasagne][3], [mxnet][4] \r\n3.  It's [Multiple-instance][5], [Multi-label][6] problem. From wiki: \"bag is labeled positive if there is at least one instance in it which is positive\", so if you made a regression per image, i think simple thresholding can gives you a decent result.\r\n\r\n\r\n  [1]: https://github.com/BVLC/caffe/tree/master/examples\r\n  [2]: https://github.com/fchollet/keras/tree/master/examples\r\n  [3]: https://github.com/Lasagne/Recipes\r\n  [4]: https://github.com/dmlc/mxnet/tree/master/example/image-classification\r\n  [5]: https://en.wikipedia.org/wiki/Multiple-instance_learning\r\n  [6]: https://en.wikipedia.org/wiki/Multi-label_classification",
      "votes": null
    },
    {
      "id": "112809",
      "postDate": "03/24/2016 01:35:11",
      "content": "<p>Thank you. </p>\n\n<p>I have another question if you do not mind. Benchmark based purely on color features gives non trivial result and this may imply that colors are important. Does this make sense try to use all 3 channels for images, or just one channel may be enough to get nontrivial result?</p>",
      "rawMarkdown": "Thank you. \r\n\r\nI have another question if you do not mind. Benchmark based purely on color features gives non trivial result and this may imply that colors are important. Does this make sense try to use all 3 channels for images, or just one channel may be enough to get nontrivial result?",
      "votes": null
    },
    {
      "id": "112823",
      "postDate": "03/24/2016 04:57:50",
      "content": "<p>Benchmark with color features has score 0.64590, but it's pretty low for F1, in my opinion. Only one channel is a loss of information, but convolutional neural network can be less susceptible to this, than hand-crafted color features(histograms or something like that). In both cases, all 3 channels can give better result than just one channel.</p>",
      "rawMarkdown": "Benchmark with color features has score 0.64590, but it's pretty low for F1, in my opinion. Only one channel is a loss of information, but convolutional neural network can be less susceptible to this, than hand-crafted color features(histograms or something like that). In both cases, all 3 channels can give better result than just one channel.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 112746,
      "author_name": "u1234x1234",
      "author_url": "",
      "post_date": "03/23/2016 14:35:46",
      "content": "<p>Hello, Vladimir</p>\n\n<ol>\n<li>I think most of the people uses pre-trained neural networks(or fine-tunned ones) for this competition, because training a network from scratch is very difficult task, so image size depends on structure of the network. Usually it is about 200x200.</li>\n<li>There is no reason for loading all dataset into RAM. All state-of-the-art neural networks for image classification trained in a &quot;batch&quot; mode, when only small part of data loaded in a memory at the same time, and all deep learning libraries can do it. If you prefer python, check out:  <a href=\"https://github.com/BVLC/caffe/tree/master/examples\">Caffe</a>, <a href=\"https://github.com/fchollet/keras/tree/master/examples\">Keras</a>, <a href=\"https://github.com/Lasagne/Recipes\">Lasagne</a>, <a href=\"https://github.com/dmlc/mxnet/tree/master/example/image-classification\">mxnet</a> </li>\n<li>It's <a href=\"https://en.wikipedia.org/wiki/Multiple-instance_learning\">Multiple-instance</a>, <a href=\"https://en.wikipedia.org/wiki/Multi-label_classification\">Multi-label</a> problem. From wiki: &quot;bag is labeled positive if there is at least one instance in it which is positive&quot;, so if you made a regression per image, i think simple thresholding can gives you a decent result.</li>\n</ol>",
      "votes": null,
      "replies": []
    },
    {
      "id": 112809,
      "author_name": "iglovikov",
      "author_url": "",
      "post_date": "03/24/2016 01:35:11",
      "content": "<p>Thank you. </p>\n\n<p>I have another question if you do not mind. Benchmark based purely on color features gives non trivial result and this may imply that colors are important. Does this make sense try to use all 3 channels for images, or just one channel may be enough to get nontrivial result?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 112823,
      "author_name": "u1234x1234",
      "author_url": "",
      "post_date": "03/24/2016 04:57:50",
      "content": "<p>Benchmark with color features has score 0.64590, but it's pretty low for F1, in my opinion. Only one channel is a loss of information, but convolutional neural network can be less susceptible to this, than hand-crafted color features(histograms or something like that). In both cases, all 3 channels can give better result than just one channel.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "112695": "Hello everyone,\r\n\r\nI have several technical questions.\r\n\r\n1. All images in the train/test have at most 500 pixels per side. 500x500 is a lot. I guess people make it smaller. How much smaller do you guys do so, that it works? 128x128? 256x256?\r\n\r\n2.  When I use 128x128 I can fit all dataset into my RAM and my video card does not have problems with it also. But let's say I will want to work with bigger images, so that they will not fit. Can anyone share with me examples of code(I would prefer python) where people are able to train neural net without loading all data into memory at the same time.\r\n\r\n3. Let's say I want to use this approach:\r\nMy target is 9 column array, where each row may have 0 or 1, telling if this  image belongs to a restaurant that have these tags. But because restaurant may belong to several categories I am not treating this problem as classification, but rather as a regression problem with images as input and 9 binary outputs per image as an output. My question is: I do this per image. And I would be able to give an estimate(number from 0 to 1) for each image for each tag. But at the end we are interested assigning categories to restaurants and not images. So, what is good way to merge categories per image to categories per restaurant?\r\n\r\nVladimir",
    "112746": "Hello, Vladimir\r\n\r\n1. I think most of the people uses pre-trained neural networks(or fine-tunned ones) for this competition, because training a network from scratch is very difficult task, so image size depends on structure of the network. Usually it is about 200x200.\r\n2. There is no reason for loading all dataset into RAM. All state-of-the-art neural networks for image classification trained in a \"batch\" mode, when only small part of data loaded in a memory at the same time, and all deep learning libraries can do it. If you prefer python, check out:  [Caffe][1], [Keras][2], [Lasagne][3], [mxnet][4] \r\n3.  It's [Multiple-instance][5], [Multi-label][6] problem. From wiki: \"bag is labeled positive if there is at least one instance in it which is positive\", so if you made a regression per image, i think simple thresholding can gives you a decent result.\r\n\r\n\r\n  [1]: https://github.com/BVLC/caffe/tree/master/examples\r\n  [2]: https://github.com/fchollet/keras/tree/master/examples\r\n  [3]: https://github.com/Lasagne/Recipes\r\n  [4]: https://github.com/dmlc/mxnet/tree/master/example/image-classification\r\n  [5]: https://en.wikipedia.org/wiki/Multiple-instance_learning\r\n  [6]: https://en.wikipedia.org/wiki/Multi-label_classification",
    "112809": "Thank you. \r\n\r\nI have another question if you do not mind. Benchmark based purely on color features gives non trivial result and this may imply that colors are important. Does this make sense try to use all 3 channels for images, or just one channel may be enough to get nontrivial result?",
    "112823": "Benchmark with color features has score 0.64590, but it's pretty low for F1, in my opinion. Only one channel is a loss of information, but convolutional neural network can be less susceptible to this, than hand-crafted color features(histograms or something like that). In both cases, all 3 channels can give better result than just one channel."
  },
  "source": "meta"
}