{
  "id": 19206,
  "title": "Deep learning starter code",
  "url": "/competitions/yelp-restaurant-photo-classification/discussion/19206",
  "author_name": "Nc Chen",
  "post_date": "2016-02-27T11:00:27.947000",
  "votes": 45,
  "comment_count": 46,
  "views": 13433,
  "content": "<p>Hello there,</p>\n\n<p>Since I've benefited from other Kagglers' code (eg. <a href=\"https://www.kaggle.com/c/noaa-right-whale-recognition/forums/t/17555/try-this\">Neon code for the Whale Problem</a> and other shared scripts),  I'd like to share my current solution <a href=\"https://github.com/ncchen55414/Keggle-Yelp/tree/master/CNN_Submission1\">here</a>. (these are ipython notebooks so anyone can view the results without running the code.)</p>\n\n<p>This solution has 3 steps:</p>\n\n<p>Step 1: Use the pre-trained CaffeNet to extract features from images</p>\n\n<p>Step 2: For each business, compute the average of its image features. Use this average as the business feature.</p>\n\n<p>Step3: Train a SVM for multi-label classification and predict.</p>\n\n<p>This gives a LB score of  about 0.76.</p>\n\n<p><strong>Other thoughts:</strong>\nInspired by <a href=\"https://www.kaggle.com/wendykan/yelp-restaurant-photo-classification/expensive-restaurants-look-like-this\">Wendy Kan's script</a>, I'm also trying to get an intuition of  what an expensive or good-for-lunch restaurant looks like. Attached is a t-SNE visualization picture, where we can see small clusters of piazzas/burgers/sandwiches, exactly those food for a quick lunch bite. A food image classifier might be helpful and fun to play with!</p>",
  "messages": [
    {
      "id": 109542,
      "postDate": "2016-02-27T11:00:27.947Z",
      "content": "<p>Hello there,</p>\n\n<p>Since I've benefited from other Kagglers' code (eg. <a href=\"https://www.kaggle.com/c/noaa-right-whale-recognition/forums/t/17555/try-this\">Neon code for the Whale Problem</a> and other shared scripts),  I'd like to share my current solution <a href=\"https://github.com/ncchen55414/Keggle-Yelp/tree/master/CNN_Submission1\">here</a>. (these are ipython notebooks so anyone can view the results without running the code.)</p>\n\n<p>This solution has 3 steps:</p>\n\n<p>Step 1: Use the pre-trained CaffeNet to extract features from images</p>\n\n<p>Step 2: For each business, compute the average of its image features. Use this average as the business feature.</p>\n\n<p>Step3: Train a SVM for multi-label classification and predict.</p>\n\n<p>This gives a LB score of  about 0.76.</p>\n\n<p><strong>Other thoughts:</strong>\nInspired by <a href=\"https://www.kaggle.com/wendykan/yelp-restaurant-photo-classification/expensive-restaurants-look-like-this\">Wendy Kan's script</a>, I'm also trying to get an intuition of  what an expensive or good-for-lunch restaurant looks like. Attached is a t-SNE visualization picture, where we can see small clusters of piazzas/burgers/sandwiches, exactly those food for a quick lunch bite. A food image classifier might be helpful and fun to play with!</p>",
      "rawMarkdown": "Hello there,\r\n\r\nSince I've benefited from other Kagglers' code (eg. [Neon code for the Whale Problem][1] and other shared scripts),  I'd like to share my current solution [here][2]. (these are ipython notebooks so anyone can view the results without running the code.)\r\n\r\nThis solution has 3 steps:\r\n\r\nStep 1: Use the pre-trained CaffeNet to extract features from images\r\n\r\nStep 2: For each business, compute the average of its image features. Use this average as the business feature.\r\n\r\nStep3: Train a SVM for multi-label classification and predict.\r\n\r\nThis gives a LB score of  about 0.76.\r\n\r\n**Other thoughts:**\r\nInspired by [Wendy Kan's script][3], I'm also trying to get an intuition of  what an expensive or good-for-lunch restaurant looks like. Attached is a t-SNE visualization picture, where we can see small clusters of piazzas/burgers/sandwiches, exactly those food for a quick lunch bite. A food image classifier might be helpful and fun to play with!\r\n\r\n\r\n  [1]: https://www.kaggle.com/c/noaa-right-whale-recognition/forums/t/17555/try-this\r\n  [2]: https://github.com/ncchen55414/Keggle-Yelp/tree/master/CNN_Submission1\r\n  [3]: https://www.kaggle.com/wendykan/yelp-restaurant-photo-classification/expensive-restaurants-look-like-this",
      "votes": 45
    },
    {
      "id": 109581,
      "postDate": "2016-02-28T03:16:35.910Z",
      "content": "<p>Thanks for sharing Nina! I especially enjoyed your t-SNE visualization on the good-for-lunch restaurants. My lunch today resembled the burger cluster in the middle-left.</p>\n\n<p>Out of curiosity, what would it take to get your Jupyter notebooks running on Kaggle Scripts (in terms of computational resources and libraries required?). I'd love to see them there if it's computationally feasible.</p>",
      "rawMarkdown": "Thanks for sharing Nina! I especially enjoyed your t-SNE visualization on the good-for-lunch restaurants. My lunch today resembled the burger cluster in the middle-left.\r\n\r\nOut of curiosity, what would it take to get your Jupyter notebooks running on Kaggle Scripts (in terms of computational resources and libraries required?). I'd love to see them there if it's computationally feasible.",
      "votes": 3
    },
    {
      "id": 111725,
      "postDate": "2016-03-16T12:21:01.257Z",
      "content": "<p>Hey Nina,</p>\n\n<p>Thanks a lot for explanation, I think I get it now. So the &quot;mean process&quot; is like a &quot;voting process&quot;, right? </p>",
      "rawMarkdown": "Hey Nina,\r\n\r\nThanks a lot for explanation, I think I get it now. So the \"mean process\" is like a \"voting process\", right? ",
      "votes": 1
    },
    {
      "id": 111685,
      "postDate": "2016-03-16T04:56:26.597Z",
      "content": "<p>Hi Fanghao,</p>\n\n<p>While fc7-features give better leaderboard score, probability-features are more human-understandable. So below I&#8217;ll explain why I take the mean of probability-features of images and use it as a business feature.</p>\n\n<p>The ImageNet includes <a href=\"http://image-net.org/challenges/LSVRC/2014/browse-synsets\">1000 object classes</a>, and the probability-feature is a 1000-dimensional vector, whose i-th component represents the probability of the i-th class. Only a few of the 1000 classes are relevant to restaurants but they can still be informative.</p>\n\n<p>Notice that the mean of probability-vectors is also a probability vector, so the business feature we obtain by taking the mean is a probability vector.  For example, a business feature might look like this:  (70% Plate,  10% Wine,  5% Candle,  15% Others) for a classy-ambience restaurant.</p>\n\n<p>This is my initial motivation before trying other statistics and the more abstract and general FC7 features; we can think each component of the FC7 feature vector as  a score rather than probability.  I hope this makes sense.</p>",
      "rawMarkdown": "Hi Fanghao,\r\n\r\nWhile fc7-features give better leaderboard score, probability-features are more human-understandable. So below I’ll explain why I take the mean of probability-features of images and use it as a business feature.\r\n\r\nThe ImageNet includes [1000 object classes][1], and the probability-feature is a 1000-dimensional vector, whose i-th component represents the probability of the i-th class. Only a few of the 1000 classes are relevant to restaurants but they can still be informative.\r\n\r\nNotice that the mean of probability-vectors is also a probability vector, so the business feature we obtain by taking the mean is a probability vector.  For example, a business feature might look like this:  (70% Plate,  10% Wine,  5% Candle,  15% Others) for a classy-ambience restaurant.\r\n\r\nThis is my initial motivation before trying other statistics and the more abstract and general FC7 features; we can think each component of the FC7 feature vector as  a score rather than probability.  I hope this makes sense.\r\n\r\n\r\n  [1]: http://image-net.org/challenges/LSVRC/2014/browse-synsets",
      "votes": 1
    },
    {
      "id": 111673,
      "postDate": "2016-03-16T02:07:15.320Z",
      "content": "<p>Hey guys, I am quite confusing why we should use the mean of all the features (fc7) for one biz to represent this biz. How can this work well? And I think if we just get the mean of all the images (raw data) for on biz to represent this biz, it won't work, right?</p>",
      "rawMarkdown": "Hey guys, I am quite confusing why we should use the mean of all the features (fc7) for one biz to represent this biz. How can this work well? And I think if we just get the mean of all the images (raw data) for on biz to represent this biz, it won't work, right?",
      "votes": 1
    },
    {
      "id": 111285,
      "postDate": "2016-03-13T10:18:06.130Z",
      "content": "<p>I'm not sure how much we should go into this here since it's kind of tangential to the original topic, but maybe my confusion is shared by some others so I'll give it a shot.</p>\n\n<p>I would agree with the fact that in my case adding another point could change the decision boundary due to the normalization possibly changing. One way around this would be to standardize instead of normalize, so you'd subtract the mean of each feature and divide by it's standard deviation. This way adding a sample is unlikely to change the decision boundary assuming n is large and you would still have the advantage of a similar scale for each feature. </p>\n\n<p>I guess I see the advantage of normalizing your way since you're combining two different spaces. I wonder if doing both directions would yield any kind of better result...</p>\n\n<p>Thanks for the explanations, it's really helpful.</p>",
      "rawMarkdown": "I'm not sure how much we should go into this here since it's kind of tangential to the original topic, but maybe my confusion is shared by some others so I'll give it a shot.\r\n\r\nI would agree with the fact that in my case adding another point could change the decision boundary due to the normalization possibly changing. One way around this would be to standardize instead of normalize, so you'd subtract the mean of each feature and divide by it's standard deviation. This way adding a sample is unlikely to change the decision boundary assuming n is large and you would still have the advantage of a similar scale for each feature. \r\n\r\nI guess I see the advantage of normalizing your way since you're combining two different spaces. I wonder if doing both directions would yield any kind of better result...\r\n\r\nThanks for the explanations, it's really helpful.",
      "votes": 1
    },
    {
      "id": 110422,
      "postDate": "2016-03-05T08:55:48.377Z",
      "content": "<p>Thanks very much Nina, I really enjoyed reading through your code.</p>\n\n<p>I wanted to run it end to end but don't have a GPU, so I edited step 1 to use Histogram of Oriented Gradients features [1] instead of CaffeNet. Feature extraction on CPU takes about 8 hours on a  2013 i7 Macbook Pro. The code and a bit more info are in this pull request. [2]</p>\n\n<p>[1] <a href=\"https://en.wikipedia.org/wiki/Histogram_of_oriented_gradients\">https://en.wikipedia.org/wiki/Histogram_of_oriented_gradients</a></p>\n\n<p>[2] <a href=\"https://github.com/ncchen55414/Kaggle-Yelp/pull/1\">https://github.com/ncchen55414/Kaggle-Yelp/pull/1</a></p>",
      "rawMarkdown": "Thanks very much Nina, I really enjoyed reading through your code.\r\n\r\nI wanted to run it end to end but don't have a GPU, so I edited step 1 to use Histogram of Oriented Gradients features [1] instead of CaffeNet. Feature extraction on CPU takes about 8 hours on a  2013 i7 Macbook Pro. The code and a bit more info are in this pull request. [2]\r\n\r\n[1] https://en.wikipedia.org/wiki/Histogram_of_oriented_gradients\r\n\r\n[2] https://github.com/ncchen55414/Kaggle-Yelp/pull/1",
      "votes": 1
    },
    {
      "id": 110230,
      "postDate": "2016-03-03T23:48:12.153Z",
      "content": "<p>Thanks, Nina and everyone sharing useful tips here. This will be the first time I deal with image recognition using NN, so I'm clueless (now a bit less). Hoping for a top 25% at the end.</p>",
      "rawMarkdown": "Thanks, Nina and everyone sharing useful tips here. This will be the first time I deal with image recognition using NN, so I'm clueless (now a bit less). Hoping for a top 25% at the end.",
      "votes": 1
    },
    {
      "id": 110199,
      "postDate": "2016-03-03T18:43:49.007Z",
      "content": "<p>[quote=Gino;110136]</p>\n\n<p>@u1234x1234 .. Newbie questions...</p>\n\n<p>Can you use inception's pre trained model the same as way Nina uses Caffe's one?   Inception model seems a lot more complex too.  Which layer could you use?  Not even sure if this is practical...or if i am making any sense....  Any info would be appreciated.\nLooking to learn about this this but still trying to install caffe on my mac....and deal with the depencies mess.\nThanks</p>\n\n<p>[/quote]</p>\n\n<p>Yes, you can use plenty of different models the same way, with small changes, mostly in <a href=\"https://github.com/ncchen55414/Kaggle-Yelp/blob/master/CNN_Submission1/Step1_ImageFeatureFc7.ipynb\">Step 1</a>: image preprocessing, layer name, number of extracted features, etc. Check out <a href=\"https://github.com/BVLC/caffe/wiki/Model-Zoo\">Model Zoo</a> for other models, but you should take into account license type. <a href=\"https://www.kaggle.com/c/yelp-restaurant-photo-classification/forums/t/18212/external-data-requests\">Topic about it</a>.</p>\n\n<p>For models trained on the ImageNet dataset: first layers contain a low level information(about edges and blobs), last layers - high-level, semantic information. I think you should start with last but one layer, and try different variants. Last layer(softmax) can gives result &quot;overfitted&quot; on ImageNet classes, but in this task there's no need to distinguish cats from dogs/</p>",
      "rawMarkdown": "[quote=Gino;110136]\r\n\r\n@u1234x1234 .. Newbie questions...\r\n\r\nCan you use inception's pre trained model the same as way Nina uses Caffe's one?   Inception model seems a lot more complex too.  Which layer could you use?  Not even sure if this is practical...or if i am making any sense....  Any info would be appreciated.\r\nLooking to learn about this this but still trying to install caffe on my mac....and deal with the depencies mess.\r\nThanks\r\n\r\n\r\n[/quote]\r\n\r\nYes, you can use plenty of different models the same way, with small changes, mostly in [Step 1][1]: image preprocessing, layer name, number of extracted features, etc. Check out [Model Zoo][2] for other models, but you should take into account license type. [Topic about it][3].\r\n\r\nFor models trained on the ImageNet dataset: first layers contain a low level information(about edges and blobs), last layers - high-level, semantic information. I think you should start with last but one layer, and try different variants. Last layer(softmax) can gives result \"overfitted\" on ImageNet classes, but in this task there's no need to distinguish cats from dogs/\r\n\r\n  [1]: https://github.com/ncchen55414/Kaggle-Yelp/blob/master/CNN_Submission1/Step1_ImageFeatureFc7.ipynb\r\n  [2]: https://github.com/BVLC/caffe/wiki/Model-Zoo\r\n  [3]: https://www.kaggle.com/c/yelp-restaurant-photo-classification/forums/t/18212/external-data-requests",
      "votes": 1
    },
    {
      "id": 109986,
      "postDate": "2016-03-01T22:44:06.100Z",
      "content": "<p>[quote=Johannes Ahlmann;109865]</p>\n\n<p>Thank you very much @Nina Chen for this!</p>\n\n<p>I found some of the labels from pre-trained CaffeNet very unhelpful (i.e. &quot;food&quot;).\nWhile enticing to use &quot;just&quot; the pretrained label output, I think this will be quite limited in application.</p>\n\n<p>To me the pre-trained classes of CaffeNet seem almost orthogonal to the yelp challenge classes.\nThe &quot;best&quot; descriptor of many of the images are words like &quot;food&quot;, &quot;plate&quot;, &quot;menu&quot;, etc. but what we are actually interested in is the &quot;difference that makes the difference&quot; between the target classes, which may not even be easily put into words in the first place.</p>\n\n<p>What we are probably looking for would be words like &quot;posh&quot;, &quot;up-market&quot;, &quot;clean&quot;, &quot;lobster&quot;, etc.\nBut then again, it would probably be much more valuable to train CaffeNet on the images against the target classes, rather than generic &quot;open world&quot; image descriptions.</p>\n\n<p>One idea would be to go through all of CaffeNet's possible/actual classification labels and hand-pick the ones that seem relevant (&quot;lobster&quot; vs. &quot;food&quot;).</p>\n\n<p>What do you think?</p>\n\n<p>[/quote]</p>\n\n<p>You can cut off last fully connected layers because they are responsible for the actual ImageNet labels, and take features from previous layers, which contains just some semantic information.</p>\n\n<p>Regarding picking some subset of activations, there's quite interesting paper [1]:</p>\n\n<p>&quot;First, we find that there is no distinction between individual high level units and\nrandom linear combinations of high level units, according to various methods of\nunit analysis. It suggests that it is the space, rather than the individual units, that\ncontains the semantic information in the high layers of neural networks.&quot;</p>\n\n<p><a href=\"http://arxiv.org/pdf/1312.6199\">[1]</a> Szegedy, C., Zaremba, W., Sutskever, I., Bruna, J., Erhan, D., Goodfellow, I., &amp; Fergus, R. (2013). Intriguing properties of neural networks.</p>",
      "rawMarkdown": "[quote=Johannes Ahlmann;109865]\r\n\r\nThank you very much @Nina Chen for this!\r\n\r\nI found some of the labels from pre-trained CaffeNet very unhelpful (i.e. \"food\").\r\nWhile enticing to use \"just\" the pretrained label output, I think this will be quite limited in application.\r\n\r\nTo me the pre-trained classes of CaffeNet seem almost orthogonal to the yelp challenge classes.\r\nThe \"best\" descriptor of many of the images are words like \"food\", \"plate\", \"menu\", etc. but what we are actually interested in is the \"difference that makes the difference\" between the target classes, which may not even be easily put into words in the first place.\r\n\r\nWhat we are probably looking for would be words like \"posh\", \"up-market\", \"clean\", \"lobster\", etc.\r\nBut then again, it would probably be much more valuable to train CaffeNet on the images against the target classes, rather than generic \"open world\" image descriptions.\r\n\r\nOne idea would be to go through all of CaffeNet's possible/actual classification labels and hand-pick the ones that seem relevant (\"lobster\" vs. \"food\").\r\n\r\nWhat do you think?\r\n\r\n[/quote]\r\n\r\nYou can cut off last fully connected layers because they are responsible for the actual ImageNet labels, and take features from previous layers, which contains just some semantic information.\r\n\r\nRegarding picking some subset of activations, there's quite interesting paper \\[1]:\r\n\r\n\"First, we find that there is no distinction between individual high level units and\r\nrandom linear combinations of high level units, according to various methods of\r\nunit analysis. It suggests that it is the space, rather than the individual units, that\r\ncontains the semantic information in the high layers of neural networks.\"\r\n \r\n[\\[1\\]][1] Szegedy, C., Zaremba, W., Sutskever, I., Bruna, J., Erhan, D., Goodfellow, I., & Fergus, R. (2013). Intriguing properties of neural networks.\r\n\r\n\r\n  [1]: http://arxiv.org/pdf/1312.6199",
      "votes": 1
    },
    {
      "id": 109865,
      "postDate": "2016-03-01T10:04:03.660Z",
      "content": "<p>Thank you very much @Nina Chen for this!</p>\n\n<p>I found some of the labels from pre-trained CaffeNet very unhelpful (i.e. &quot;food&quot;).\nWhile enticing to use &quot;just&quot; the pretrained label output, I think this will be quite limited in application.</p>\n\n<p>To me the pre-trained classes of CaffeNet seem almost orthogonal to the yelp challenge classes.\nThe &quot;best&quot; descriptor of many of the images are words like &quot;food&quot;, &quot;plate&quot;, &quot;menu&quot;, etc. but what we are actually interested in is the &quot;difference that makes the difference&quot; between the target classes, which may not even be easily put into words in the first place.</p>\n\n<p>What we are probably looking for would be words like &quot;posh&quot;, &quot;up-market&quot;, &quot;clean&quot;, &quot;lobster&quot;, etc.\nBut then again, it would probably be much more valuable to train CaffeNet on the images against the target classes, rather than generic &quot;open world&quot; image descriptions.</p>\n\n<p>One idea would be to go through all of CaffeNet's possible/actual classification labels and hand-pick the ones that seem relevant (&quot;lobster&quot; vs. &quot;food&quot;).</p>\n\n<p>What do you think?</p>",
      "rawMarkdown": "Thank you very much @Nina Chen for this!\r\n\r\nI found some of the labels from pre-trained CaffeNet very unhelpful (i.e. \"food\").\r\nWhile enticing to use \"just\" the pretrained label output, I think this will be quite limited in application.\r\n\r\nTo me the pre-trained classes of CaffeNet seem almost orthogonal to the yelp challenge classes.\r\nThe \"best\" descriptor of many of the images are words like \"food\", \"plate\", \"menu\", etc. but what we are actually interested in is the \"difference that makes the difference\" between the target classes, which may not even be easily put into words in the first place.\r\n\r\nWhat we are probably looking for would be words like \"posh\", \"up-market\", \"clean\", \"lobster\", etc.\r\nBut then again, it would probably be much more valuable to train CaffeNet on the images against the target classes, rather than generic \"open world\" image descriptions.\r\n\r\nOne idea would be to go through all of CaffeNet's possible/actual classification labels and hand-pick the ones that seem relevant (\"lobster\" vs. \"food\").\r\n\r\nWhat do you think?",
      "votes": 1
    },
    {
      "id": 109735,
      "postDate": "2016-02-29T22:46:45.110Z",
      "content": "<p>Thanks, @Nina Chen, with fc6 + fc7, I got 0.803 now, tweaking the parameters of svm does not help much. Maybe fine tune the model on this dataset and extract features might help</p>",
      "rawMarkdown": "Thanks, @Nina Chen, with fc6 + fc7, I got 0.803 now, tweaking the parameters of svm does not help much. Maybe fine tune the model on this dataset and extract features might help",
      "votes": 1
    },
    {
      "id": 109728,
      "postDate": "2016-02-29T22:28:47.737Z",
      "content": "<p>[quote=Blue Light;109624]</p>\n\n<p>Hi there,</p>\n\n<p>Thanks for sharing the code.  I've been trying to run code that uses Caffe (as yours does), but it's been incredibly slow. My machine lacks a good GPU, and renting out GPU instances on AWS can get quite expensive.  Is there some other option to get neural net code to run within a reasonable time period without incurring too much cost?</p>\n\n<p>Thanks</p>\n\n<p>[/quote]</p>\n\n<p>Hi @Blue Light,</p>\n\n<p>It might be faster to use Caffe's <a href=\"http://caffe.berkeleyvision.org/gathered/examples/feature_extraction.html\">feature extractor c++ utility</a> than to use the python wrapper. </p>\n\n<p>However, the features extracted are somehow different. (<a href=\"https://github.com/ncchen55414/Keggle-Yelp/blob/master/CNN_Submission1/Extract_feature_python_vs_cpp.ipynb\">Here</a> is my notebook to read the features extracted from c++ utility and compare  them with the python wrapper.) I'm not sure if it matters. </p>\n\n<p>Hopefully someone more knowledgeable can chime in.</p>",
      "rawMarkdown": "[quote=Blue Light;109624]\r\n\r\nHi there,\r\n\r\nThanks for sharing the code.  I've been trying to run code that uses Caffe (as yours does), but it's been incredibly slow. My machine lacks a good GPU, and renting out GPU instances on AWS can get quite expensive.  Is there some other option to get neural net code to run within a reasonable time period without incurring too much cost?\r\n\r\nThanks\r\n\r\n[/quote]\r\n\r\nHi @Blue Light,\r\n\r\nIt might be faster to use Caffe's [feature extractor c++ utility][1] than to use the python wrapper. \r\n\r\nHowever, the features extracted are somehow different. ([Here][2] is my notebook to read the features extracted from c++ utility and compare  them with the python wrapper.) I'm not sure if it matters. \r\n\r\nHopefully someone more knowledgeable can chime in.\r\n\r\n\r\n  [1]: http://caffe.berkeleyvision.org/gathered/examples/feature_extraction.html\r\n  [2]: https://github.com/ncchen55414/Keggle-Yelp/blob/master/CNN_Submission1/Extract_feature_python_vs_cpp.ipynb",
      "votes": 1
    },
    {
      "id": 109617,
      "postDate": "2016-02-28T19:15:14.713Z",
      "content": "<p>[quote=Ben Hamner;109581]</p>\n\n<p>Thanks for sharing Nina! I especially enjoyed your t-SNE visualization on the good-for-lunch restaurants. My lunch today resembled the burger cluster in the middle-left.</p>\n\n<p>Out of curiosity, what would it take to get your Jupyter notebooks running on Kaggle Scripts (in terms of computational resources and libraries required?). I'd love to see them there if it's computationally feasible.</p>\n\n<p>[/quote]</p>\n\n<p>Thanks for your comment Ben. Also thanks Kaggle/Yelp for hosting this competition. The variety of food and restaurants I've tried have significantly increased ever since I started working  on this project and look at food image for hours!</p>\n\n<p>As for the computational resources,</p>\n\n<ul>\n<li>Step1: Image feature extraction is run on a GPU (GTX980 with 4GB RAM) for about 2-3 hours. Caffe is required. The features extracted are 4GB each for training and test images. I'll be happy to share them if they are not so huge; this should help people inaccessible to GPU (/and have slow internet so AWS is not an option).</li>\n<li>GPU or Caffe are not required for other steps. I re-run the code on my Macbook Air 2014 ( i5 with 8GB RAM). Step2 takes about 1 hours, and other steps are done in minutes. Scikit-learn is the most crucial library used.</li>\n</ul>",
      "rawMarkdown": "[quote=Ben Hamner;109581]\r\n\r\nThanks for sharing Nina! I especially enjoyed your t-SNE visualization on the good-for-lunch restaurants. My lunch today resembled the burger cluster in the middle-left.\r\n\r\nOut of curiosity, what would it take to get your Jupyter notebooks running on Kaggle Scripts (in terms of computational resources and libraries required?). I'd love to see them there if it's computationally feasible.\r\n\r\n[/quote]\r\n\r\nThanks for your comment Ben. Also thanks Kaggle/Yelp for hosting this competition. The variety of food and restaurants I've tried have significantly increased ever since I started working  on this project and look at food image for hours!\r\n\r\nAs for the computational resources,\r\n\r\n - Step1: Image feature extraction is run on a GPU (GTX980 with 4GB RAM) for about 2-3 hours. Caffe is required. The features extracted are 4GB each for training and test images. I'll be happy to share them if they are not so huge; this should help people inaccessible to GPU (/and have slow internet so AWS is not an option).\r\n - GPU or Caffe are not required for other steps. I re-run the code on my Macbook Air 2014 ( i5 with 8GB RAM). Step2 takes about 1 hours, and other steps are done in minutes. Scikit-learn is the most crucial library used.\r\n\r\n\r\n  [1]: https://github.com/ncchen55414/Keggle-Yelp/blob/master/CNN_Submission1/my_conda_env.txt",
      "votes": 1
    },
    {
      "id": 111186,
      "postDate": "2016-03-12T16:12:47.510Z",
      "content": "<p>[quote=Blue Light;111185]</p>\n\n<p>[quote=SecondPlan;109735]</p>\n\n<p>Thanks, @Nina Chen, with fc6 + fc7, I got 0.803 now, tweaking the parameters of svm does not help much. Maybe fine tune the model on this dataset and extract features might help</p>\n\n<p>[/quote]\nSimply extending the fc7 feature vector ((of 4096 features) with the fc6 one  (another set of 4096 features) added no improvement on the test set for me.  </p>\n\n<p>Are there any particular guidelines as to how one could get more bang out of adding fc6 features into the mix?</p>\n\n<p>Thanks</p>\n\n<p>[/quote]</p>\n\n<p>Well, you should normalize the vectors before concat them</p>",
      "rawMarkdown": "[quote=Blue Light;111185]\r\n\r\n[quote=SecondPlan;109735]\r\n\r\nThanks, @Nina Chen, with fc6 + fc7, I got 0.803 now, tweaking the parameters of svm does not help much. Maybe fine tune the model on this dataset and extract features might help\r\n\r\n[/quote]\r\nSimply extending the fc7 feature vector ((of 4096 features) with the fc6 one  (another set of 4096 features) added no improvement on the test set for me.  \r\n\r\nAre there any particular guidelines as to how one could get more bang out of adding fc6 features into the mix?\r\n\r\nThanks\r\n\r\n[/quote]\r\n\r\nWell, you should normalize the vectors before concat them",
      "votes": 2
    },
    {
      "id": 111069,
      "postDate": "2016-03-11T07:36:53.137Z",
      "content": "<p>Ayush,\nyou can create custom layers in Python based on caffe.Layer class. There is some example code shipped with caffe, and I just saw that just recently the caffe community added some example code for multi-label classification on the Pascal challenge data:</p>\n\n<p><a href=\"https://github.com/BVLC/caffe/blob/master/examples/pascal-multilabel-with-datalayer.ipynb\">https://github.com/BVLC/caffe/blob/master/examples/pascal-multilabel-with-datalayer.ipynb</a></p>\n\n<p>This should give you a good blueprint for building the layer, and it also demonstrates how to make use of the pretrained alexnet.</p>",
      "rawMarkdown": "Ayush,\r\nyou can create custom layers in Python based on caffe.Layer class. There is some example code shipped with caffe, and I just saw that just recently the caffe community added some example code for multi-label classification on the Pascal challenge data:\r\n\r\nhttps://github.com/BVLC/caffe/blob/master/examples/pascal-multilabel-with-datalayer.ipynb\r\n\r\nThis should give you a good blueprint for building the layer, and it also demonstrates how to make use of the pretrained alexnet.\r\n",
      "votes": 2
    },
    {
      "id": 109982,
      "postDate": "2016-03-01T22:08:01.177Z",
      "content": "<p>[quote=Blue Light;109624]</p>\n\n<p>Hi there,</p>\n\n<p>Thanks for sharing the code.  I've been trying to run code that uses Caffe (as yours does), but it's been incredibly slow. My machine lacks a good GPU, and renting out GPU instances on AWS can get quite expensive.  Is there some other option to get neural net code to run within a reasonable time period without incurring too much cost?</p>\n\n<p>Thanks</p>\n\n<p>[/quote]</p>\n\n<p>There's no need for a good GPU, if you only want to make predictions with pretrained net. For example Inception network [1](it was state-of-the-art 1 year ago) takes 600ms on a single image with a ordinary CPU. Inceptions architectures much more CPU friendly, compared with the other state-of-the-art(VGG-like[2], etc). So, if you have the CPU with 4 threads, you can process 1/0.6 * 60 * 60 * 24 * 4 = 576.000 images per day. There're only 450.000 images in this comptetion. But you should have optimized BLAS library installed.</p>\n\n<p>As Nina Chen noted, there's difference between features extracted with CPU and GPU, but usually it doesnt matter. GPU with different cuDNN versions can gives different results as well.</p>\n\n<p><a href=\"http://arxiv.org/pdf/1502.03167\">[1]</a> Ioffe, S., &amp; Szegedy, C. (2015). Batch normalization: Accelerating deep network training by reducing internal covariate shift.</p>\n\n<p><a href=\"http://arxiv.org/pdf/1409.1556\">[2]</a> Simonyan, K., &amp; Zisserman, A. (2014). Very deep convolutional networks for large-scale image recognition. </p>",
      "rawMarkdown": "[quote=Blue Light;109624]\r\n\r\nHi there,\r\n\r\nThanks for sharing the code.  I've been trying to run code that uses Caffe (as yours does), but it's been incredibly slow. My machine lacks a good GPU, and renting out GPU instances on AWS can get quite expensive.  Is there some other option to get neural net code to run within a reasonable time period without incurring too much cost?\r\n\r\nThanks\r\n\r\n[/quote]\r\n\r\nThere's no need for a good GPU, if you only want to make predictions with pretrained net. For example Inception network \\[1](it was state-of-the-art 1 year ago) takes 600ms on a single image with a ordinary CPU. Inceptions architectures much more CPU friendly, compared with the other state-of-the-art(VGG-like\\[2], etc). So, if you have the CPU with 4 threads, you can process 1/0.6 * 60 * 60 * 24 * 4 = 576.000 images per day. There're only 450.000 images in this comptetion. But you should have optimized BLAS library installed.\r\n\r\nAs Nina Chen noted, there's difference between features extracted with CPU and GPU, but usually it doesnt matter. GPU with different cuDNN versions can gives different results as well.\r\n\r\n[\\[1\\]][1] Ioffe, S., & Szegedy, C. (2015). Batch normalization: Accelerating deep network training by reducing internal covariate shift.\r\n\r\n[\\[2\\]][2] Simonyan, K., & Zisserman, A. (2014). Very deep convolutional networks for large-scale image recognition. \r\n\r\n\r\n  [1]: http://arxiv.org/pdf/1502.03167\r\n  [2]: http://arxiv.org/pdf/1409.1556",
      "votes": 2
    },
    {
      "id": 109658,
      "postDate": "2016-02-29T06:41:39.640Z",
      "content": "<p>@Nina Chen (It seems I cannot @ people in Kaggle)</p>\n\n<p>I was trying something like &quot; training deep networks on single image instances with labels inherited from business&quot; and the results are terrible. I should have read every post carefully.</p>\n\n<p>It is my second year here. Chilly!</p>",
      "rawMarkdown": "@Nina Chen (It seems I cannot @ people in Kaggle)\r\n\r\nI was trying something like \" training deep networks on single image instances with labels inherited from business\" and the results are terrible. I should have read every post carefully.\r\n\r\nIt is my second year here. Chilly!",
      "votes": 2
    },
    {
      "id": 109582,
      "postDate": "2016-02-28T03:43:48.180Z",
      "content": "<p>Up you go!</p>\n\n<p>The notebook introduces</p>\n\n<ol>\n<li>The concept of multiple instance.</li>\n<li>Multiple-label classification SVM.</li>\n</ol>\n\n<p>I have been struggling in dealing with the above two problems in the past two weeks. You are really helping me out.</p>\n\n<p>This is the best thing happened today. The second best thing is that the temp is above 0 in 55414.</p>\n\n<p>Thanks!</p>",
      "rawMarkdown": "Up you go!\r\n\r\nThe notebook introduces\r\n\r\n1. The concept of multiple instance.\r\n2. Multiple-label classification SVM.\r\n\r\nI have been struggling in dealing with the above two problems in the past two weeks. You are really helping me out.\r\n\r\nThis is the best thing happened today. The second best thing is that the temp is above 0 in 55414.\r\n\r\nThanks!",
      "votes": 2
    },
    {
      "id": 1892254,
      "postDate": "2022-08-10T01:47:59.537Z",
      "content": "<p>Thank you for sharing your code. I have a question about Hog_SVM deep learning.<br>\nAt extract_hog_features(train_images), I have error that mismatch with photo_id of train_photo_to_biz_ids and photo name.jpg of train photo file.<br>\nCan you tell me how you solved it?</p>",
      "rawMarkdown": "Thank you for sharing your code. I have a question about Hog_SVM deep learning.\nAt extract_hog_features(train_images), I have error that mismatch with photo_id of train_photo_to_biz_ids and photo name.jpg of train photo file.\nCan you tell me how you solved it?"
    },
    {
      "id": 112792,
      "postDate": "2016-03-23T19:54:42.343Z",
      "content": "<p>[quote=Shouvik Dutta;111280]</p>\n\n<p>Huh ok. Yeah I meant normalize for each dim across all rows of examples, since that would account for different features potentially having different output ranges which could affect the decision boundary found by an svm (unless we concede that these different ranges might be important as determined by the original nn). I'm not sure what sample normalization brings since it divides each partition of your feature vector by a constant which would (I think) simply be reflected in the weights learned by your svm.</p>\n\n<p>[/quote]</p>\n\n<p>Normalization or better say feature scaling <strong>across rows</strong> is standard preprocessing for SVM. &quot;The main advantage of scaling is to avoid attributes in greater numeric ranges dominating those in smaller numeric ranges&quot; See explanation from LIBSVM author:  <a href=\"https://www.csie.ntu.edu.tw/~cjlin/papers/guide/guide.pdf\">2.2 Scaling</a> </p>\n\n<p>Per-instance normalization can work too, e.g. for images. &quot;If your data is stationary (i.e., the statistics for each data dimension follow the same distribution), then you might want to consider subtracting the mean-value for each example (computed per-example).&quot; See Andrew Ng explanation: <a href=\"http://deeplearning.stanford.edu/wiki/index.php/Data_Preprocessing\">Per-example mean subtraction</a></p>\n\n<p>I guess that neural net trained on images inherits this stationary in hidden layers.</p>",
      "rawMarkdown": "[quote=Shouvik Dutta;111280]\r\n\r\nHuh ok. Yeah I meant normalize for each dim across all rows of examples, since that would account for different features potentially having different output ranges which could affect the decision boundary found by an svm (unless we concede that these different ranges might be important as determined by the original nn). I'm not sure what sample normalization brings since it divides each partition of your feature vector by a constant which would (I think) simply be reflected in the weights learned by your svm.\r\n\r\n[/quote]\r\n\r\nNormalization or better say feature scaling **across rows** is standard preprocessing for SVM. \"The main advantage of scaling is to avoid attributes in greater numeric ranges dominating those in smaller numeric ranges\" See explanation from LIBSVM author:  [2.2 Scaling][1] \r\n\r\nPer-instance normalization can work too, e.g. for images. \"If your data is stationary (i.e., the statistics for each data dimension follow the same distribution), then you might want to consider subtracting the mean-value for each example (computed per-example).\" See Andrew Ng explanation: [Per-example mean subtraction][2]\r\n\r\nI guess that neural net trained on images inherits this stationary in hidden layers.\r\n\r\n\r\n  [1]: https://www.csie.ntu.edu.tw/~cjlin/papers/guide/guide.pdf\r\n  [2]: http://deeplearning.stanford.edu/wiki/index.php/Data_Preprocessing"
    },
    {
      "id": 112715,
      "postDate": "2016-03-23T09:14:27.407Z",
      "content": "<p>Hi Guys,</p>\n\n<p>I am still struggling to perform Multi-Instance Classification using Caffe. I am confused about which Loss layer should be used for Multi-Instance Classification (for single label I used SoftmaxLoss layer).</p>",
      "rawMarkdown": "Hi Guys,\r\n\r\nI am still struggling to perform Multi-Instance Classification using Caffe. I am confused about which Loss layer should be used for Multi-Instance Classification (for single label I used SoftmaxLoss layer)."
    },
    {
      "id": 111287,
      "postDate": "2016-03-13T10:34:34.870Z",
      "content": "<p>doing both directions would be a disaster since you completely changed the space and you should test that not just wonder.</p>\n\n<p>Remember <strong>Whatever transformation you do, you will eventually input a n x k matrix to svm, where you have k features, each time you feed a classifier like svm , a 1 x k vector, it changes its weights accordingly, that's how you train a machine learning classifier.</strong></p>\n\n<p>There is no other paradigm to do training </p>",
      "rawMarkdown": "doing both directions would be a disaster since you completely changed the space and you should test that not just wonder.\r\n\r\nRemember **Whatever transformation you do, you will eventually input a n x k matrix to svm, where you have k features, each time you feed a classifier like svm , a 1 x k vector, it changes its weights accordingly, that's how you train a machine learning classifier.**\r\n\r\nThere is no other paradigm to do training \r\n\r\n"
    },
    {
      "id": 111281,
      "postDate": "2016-03-13T09:45:14.200Z",
      "content": "<p>let think this way, does adding 1 extra row of data change the decision boundary completely? In your case, yes, since your have to renormalize for that feature. In my case, no, it just add one feature vector in the space, which might not be a big influencer.</p>\n\n<p>If you do other kaggle competitions, when you have a big matrix n x m, and you want to do svd, makes it n x k, this is feature projection or normalizaiton in a vague broad sense, you can doing the change on the columns instead of rows.</p>",
      "rawMarkdown": "let think this way, does adding 1 extra row of data change the decision boundary completely? In your case, yes, since your have to renormalize for that feature. In my case, no, it just add one feature vector in the space, which might not be a big influencer.\r\n\r\nIf you do other kaggle competitions, when you have a big matrix n x m, and you want to do svd, makes it n x k, this is feature projection or normalizaiton in a vague broad sense, you can doing the change on the columns instead of rows."
    },
    {
      "id": 111280,
      "postDate": "2016-03-13T09:34:34.823Z",
      "content": "<p>Huh ok. Yeah I meant normalize for each dim across all rows of examples, since that would account for different features potentially having different output ranges which could affect the decision boundary found by an svm (unless we concede that these different ranges might be important as determined by the original nn). I'm not sure what sample normalization brings since it divides each partition of your feature vector by a constant which would (I think) simply be reflected in the weights learned by your svm.</p>",
      "rawMarkdown": "Huh ok. Yeah I meant normalize for each dim across all rows of examples, since that would account for different features potentially having different output ranges which could affect the decision boundary found by an svm (unless we concede that these different ranges might be important as determined by the original nn). I'm not sure what sample normalization brings since it divides each partition of your feature vector by a constant which would (I think) simply be reflected in the weights learned by your svm."
    },
    {
      "id": 111278,
      "postDate": "2016-03-13T09:30:08.413Z",
      "content": "<p>sorry, I misunderstand your point, but I did sample-wise normalization and concat. You mean, for each dim of 8092, you normalize across all rows of exmaples? You can try to see whether it works or not. But sample-wise normalization did make sense to me since you are simply projecting  features into new spaces for each sample irrelevant of other data. This way, when you add more and more data, you features in the new space might form some decision boundary for the svm to catch.</p>",
      "rawMarkdown": "sorry, I misunderstand your point, but I did sample-wise normalization and concat. You mean, for each dim of 8092, you normalize across all rows of exmaples? You can try to see whether it works or not. But sample-wise normalization did make sense to me since you are simply projecting  features into new spaces for each sample irrelevant of other data. This way, when you add more and more data, you features in the new space might form some decision boundary for the svm to catch.\r\n"
    },
    {
      "id": 111276,
      "postDate": "2016-03-13T09:17:55.737Z",
      "content": "<p>Ah, when I read your comment I thought you meant to normalize the fc6 vector and the fc7 vector for each sample and then concatenate those, which seemed to make less sense. Thanks!</p>",
      "rawMarkdown": "Ah, when I read your comment I thought you meant to normalize the fc6 vector and the fc7 vector for each sample and then concatenate those, which seemed to make less sense. Thanks!"
    },
    {
      "id": 111275,
      "postDate": "2016-03-13T09:15:38.040Z",
      "content": "<p>you get it now, but that's what I said, somehow I did not make myself clear</p>",
      "rawMarkdown": "you get it now, but that's what I said, somehow I did not make myself clear"
    },
    {
      "id": 111273,
      "postDate": "2016-03-13T09:10:58.807Z",
      "content": "<p>Interesting. With that theory wouldn't it make more sense to normalize each feature separately and then end up with a space of |fc6| + |fc7| features, each of which is normalized? Or am I missing some intuition?</p>",
      "rawMarkdown": "Interesting. With that theory wouldn't it make more sense to normalize each feature separately and then end up with a space of |fc6| + |fc7| features, each of which is normalized? Or am I missing some intuition?"
    },
    {
      "id": 111268,
      "postDate": "2016-03-13T08:47:48.557Z",
      "content": "<p>Well, fc6 and fc7 are different vectors, numbers inside fc6 are meaningful only to fc6, if you concat with fc7 ,  this is like bring two vectors in different space together into a new space,  and normalize them is a simple way to make this new space meaningful.</p>",
      "rawMarkdown": "Well, fc6 and fc7 are different vectors, numbers inside fc6 are meaningful only to fc6, if you concat with fc7 ,  this is like bring two vectors in different space together into a new space,  and normalize them is a simple way to make this new space meaningful."
    },
    {
      "id": 111267,
      "postDate": "2016-03-13T08:24:57.727Z",
      "content": "<p>Why would you want to normalize the vectors? Wouldn't normalizing them have no effect on an SVM since it's just a linear scaling?</p>",
      "rawMarkdown": "Why would you want to normalize the vectors? Wouldn't normalizing them have no effect on an SVM since it's just a linear scaling?"
    },
    {
      "id": 111185,
      "postDate": "2016-03-12T16:11:11.083Z",
      "content": "<p>[quote=SecondPlan;109735]</p>\n\n<p>Thanks, @Nina Chen, with fc6 + fc7, I got 0.803 now, tweaking the parameters of svm does not help much. Maybe fine tune the model on this dataset and extract features might help</p>\n\n<p>[/quote]\nSimply extending the fc7 feature vector ((of 4096 features) with the fc6 one  (another set of 4096 features) added no improvement on the test set for me.  </p>\n\n<p>Are there any particular guidelines as to how one could get more bang out of adding fc6 features into the mix?</p>\n\n<p>Thanks</p>",
      "rawMarkdown": "[quote=SecondPlan;109735]\r\n\r\nThanks, @Nina Chen, with fc6 + fc7, I got 0.803 now, tweaking the parameters of svm does not help much. Maybe fine tune the model on this dataset and extract features might help\r\n\r\n[/quote]\r\nSimply extending the fc7 feature vector ((of 4096 features) with the fc6 one  (another set of 4096 features) added no improvement on the test set for me.  \r\n\r\nAre there any particular guidelines as to how one could get more bang out of adding fc6 features into the mix?\r\n\r\nThanks"
    },
    {
      "id": 111060,
      "postDate": "2016-03-11T06:00:56.637Z",
      "content": "<p>I was exactly looking for this. Could you elaborate on how we can create the custom layers you mentioned. Also what type of data layer you used HDF5/LMDB/IMAGEDATA? I will appreciate if you can help.</p>\n\n<p>Thanks</p>",
      "rawMarkdown": "I was exactly looking for this. Could you elaborate on how we can create the custom layers you mentioned. Also what type of data layer you used HDF5/LMDB/IMAGEDATA? I will appreciate if you can help.\r\n\r\nThanks"
    },
    {
      "id": 111040,
      "postDate": "2016-03-10T22:04:52.060Z",
      "content": "<p>I actually got multi-instance learning working with caffe, but I had to create a custom layer for data loading and augmentation. It load the image ids and labels from a csv file, loads the images from raw jpgs  and feeds image and label vector to the net as seperate matrices.</p>",
      "rawMarkdown": "I actually got multi-instance learning working with caffe, but I had to create a custom layer for data loading and augmentation. It load the image ids and labels from a csv file, loads the images from raw jpgs  and feeds image and label vector to the net as seperate matrices."
    },
    {
      "id": 110891,
      "postDate": "2016-03-09T07:48:30.377Z",
      "content": "<p>Thanks for the response. I have another doubt regarding Multi-Instance Learning. I will appreciate if you can help. In this challenge we are given 9 different business labels. So in total there can be 512 classes (a vector of different labels corresponding to a class) possible based on which label is present for a particular image but out of this (i.e 512 classes) only 173 classes are present in the training set. Is it possible that while testing we end up getting a class out of the remaining 339 classes i.e. a combination of labels which is not present in the training set or we will get the same combination of labels that are present in the training set.</p>",
      "rawMarkdown": "Thanks for the response. I have another doubt regarding Multi-Instance Learning. I will appreciate if you can help. In this challenge we are given 9 different business labels. So in total there can be 512 classes (a vector of different labels corresponding to a class) possible based on which label is present for a particular image but out of this (i.e 512 classes) only 173 classes are present in the training set. Is it possible that while testing we end up getting a class out of the remaining 339 classes i.e. a combination of labels which is not present in the training set or we will get the same combination of labels that are present in the training set."
    },
    {
      "id": 110813,
      "postDate": "2016-03-08T14:01:41.420Z",
      "content": "<p>As far as I understand, caffe currently does not support populating a lmdb library with images and multiple labels. One way you can try is to map multiple label problems to single label problems, for example, '0 0 0 0 0 0 0 1 1 ' can be convert to its label 3 , which is the decimal value of that binary string. If you still want to train a model with multi label and 9 classes, then I would suggest to use lasagne to do that, there are example codes , just search the example code of kaggle facial keypoints recognition, which is treated as multi label regression problem.</p>",
      "rawMarkdown": "As far as I understand, caffe currently does not support populating a lmdb library with images and multiple labels. One way you can try is to map multiple label problems to single label problems, for example, '0 0 0 0 0 0 0 1 1 ' can be convert to its label 3 , which is the decimal value of that binary string. If you still want to train a model with multi label and 9 classes, then I would suggest to use lasagne to do that, there are example codes , just search the example code of kaggle facial keypoints recognition, which is treated as multi label regression problem."
    },
    {
      "id": 110807,
      "postDate": "2016-03-08T11:32:07.817Z",
      "content": "<p>Hi Guys,</p>\n\n<p>I want to train a caffe model (of my own) for multiple labels on images instead of using any pre-trained model for this problem. But most of the information on net is based on training a caffe model for images with a single label.</p>\n\n<p>Please suggest something.</p>\n\n<p>Thanks</p>",
      "rawMarkdown": "Hi Guys,\r\n\r\nI want to train a caffe model (of my own) for multiple labels on images instead of using any pre-trained model for this problem. But most of the information on net is based on training a caffe model for images with a single label.\r\n\r\nPlease suggest something.\r\n\r\nThanks"
    },
    {
      "id": 110263,
      "postDate": "2016-03-04T05:35:55.157Z",
      "content": "<p>Are we allowed to use pre-trained models? isn't that technically using outside information?</p>\n\n<p>edit: never mind, I see that this contest has some special rules. cool stuff!</p>",
      "rawMarkdown": "Are we allowed to use pre-trained models? isn't that technically using outside information?\r\n\r\nedit: never mind, I see that this contest has some special rules. cool stuff!"
    },
    {
      "id": 109624,
      "postDate": "2016-02-28T20:35:26.007Z",
      "content": "<p>Hi there,</p>\n\n<p>Thanks for sharing the code.  I've been trying to run code that uses Caffe (as yours does), but it's been incredibly slow. My machine lacks a good GPU, and renting out GPU instances on AWS can get quite expensive.  Is there some other option to get neural net code to run within a reasonable time period without incurring too much cost?</p>\n\n<p>Thanks</p>",
      "rawMarkdown": "Hi there,\r\n\r\nThanks for sharing the code.  I've been trying to run code that uses Caffe (as yours does), but it's been incredibly slow. My machine lacks a good GPU, and renting out GPU instances on AWS can get quite expensive.  Is there some other option to get neural net code to run within a reasonable time period without incurring too much cost?\r\n\r\nThanks"
    },
    {
      "id": 109619,
      "postDate": "2016-02-28T19:35:51.887Z",
      "content": "<p>[quote=Joseph PENG;109582]</p>\n\n<p>Up you go!</p>\n\n<p>The notebook introduces</p>\n\n<ol>\n<li>The concept of multiple instance.</li>\n<li>Multiple-label classification SVM.</li>\n</ol>\n\n<p>I have been struggling in dealing with the above two problems in the past two weeks. You are really helping me out.</p>\n\n<p>This is the best thing happened today. The second best thing is that the temp is above 0 in 55414.</p>\n\n<p>Thanks!</p>\n\n<p>[/quote]</p>\n\n<p>Hello Joseph,</p>\n\n<p>Glad to help! The keyword &quot;multi-instance&quot; learning is mentioned in <a href=\"https://www.kaggle.com/c/yelp-restaurant-photo-classification/forums/t/18101/welcome/102706#post102706\">ody's message</a>, before that post I was googling something like &quot;neural network with several image inputs&quot; and had no luck.</p>\n\n<p>P.S. I've been in MN for 6 years and now live in the warm California. :-)</p>",
      "rawMarkdown": "[quote=Joseph PENG;109582]\r\n\r\nUp you go!\r\n\r\nThe notebook introduces\r\n\r\n1. The concept of multiple instance.\r\n2. Multiple-label classification SVM.\r\n\r\nI have been struggling in dealing with the above two problems in the past two weeks. You are really helping me out.\r\n\r\nThis is the best thing happened today. The second best thing is that the temp is above 0 in 55414.\r\n\r\nThanks!\r\n\r\n[/quote]\r\n\r\nHello Joseph,\r\n\r\nGlad to help! The keyword \"multi-instance\" learning is mentioned in [ody's message][1], before that post I was googling something like \"neural network with several image inputs\" and had no luck.\r\n\r\nP.S. I've been in MN for 6 years and now live in the warm California. :-)\r\n  [1]: https://www.kaggle.com/c/yelp-restaurant-photo-classification/forums/t/18101/welcome/102706#post102706"
    },
    {
      "id": 109584,
      "postDate": "2016-02-28T03:57:29.550Z",
      "content": "<p>A quick question: how do you find out &quot;multiple instance&quot; and &quot;multiple-label&quot; classification SVM?</p>\n\n<p>I feel it is not from text book like ESL.</p>",
      "rawMarkdown": "A quick question: how do you find out \"multiple instance\" and \"multiple-label\" classification SVM?\r\n\r\nI feel it is not from text book like ESL."
    },
    {
      "id": 109583,
      "postDate": "2016-02-28T03:55:06.240Z",
      "rawMarkdown": ""
    },
    {
      "id": 110136,
      "postDate": "2016-03-03T05:50:54.290Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 114170,
      "postDate": "2016-04-07T22:43:18.220Z",
      "content": "<p>Thanks Nina, Thank you guys for support</p>",
      "rawMarkdown": "Thanks Nina, Thank you guys for support",
      "votes": 1
    },
    {
      "id": 110897,
      "postDate": "2016-03-09T10:52:54.993Z",
      "content": "<p>Thanks a lot for your sharing!</p>",
      "rawMarkdown": "Thanks a lot for your sharing!",
      "votes": 1
    },
    {
      "id": 1033645,
      "postDate": "2020-10-01T07:12:39.973Z",
      "content": "<p>Thanks a lot for your sharing!</p>",
      "rawMarkdown": "Thanks a lot for your sharing!"
    },
    {
      "id": 111070,
      "postDate": "2016-03-11T07:57:25.107Z",
      "content": "<p>Alright thanks Alexander.</p>",
      "rawMarkdown": "Alright thanks Alexander."
    }
  ],
  "comments": [
    {
      "id": 109581,
      "author_name": "Ben Hamner",
      "author_url": "",
      "post_date": "2016-02-28T03:16:35.910000",
      "content": "<p>Thanks for sharing Nina! I especially enjoyed your t-SNE visualization on the good-for-lunch restaurants. My lunch today resembled the burger cluster in the middle-left.</p>\n\n<p>Out of curiosity, what would it take to get your Jupyter notebooks running on Kaggle Scripts (in terms of computational resources and libraries required?). I'd love to see them there if it's computationally feasible.</p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 111725,
      "author_name": "FanghaoLuo",
      "author_url": "",
      "post_date": "2016-03-16T12:21:01.257000",
      "content": "<p>Hey Nina,</p>\n\n<p>Thanks a lot for explanation, I think I get it now. So the &quot;mean process&quot; is like a &quot;voting process&quot;, right? </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 111685,
      "author_name": "Nc Chen",
      "author_url": "",
      "post_date": "2016-03-16T04:56:26.597000",
      "content": "<p>Hi Fanghao,</p>\n\n<p>While fc7-features give better leaderboard score, probability-features are more human-understandable. So below I&#8217;ll explain why I take the mean of probability-features of images and use it as a business feature.</p>\n\n<p>The ImageNet includes <a href=\"http://image-net.org/challenges/LSVRC/2014/browse-synsets\">1000 object classes</a>, and the probability-feature is a 1000-dimensional vector, whose i-th component represents the probability of the i-th class. Only a few of the 1000 classes are relevant to restaurants but they can still be informative.</p>\n\n<p>Notice that the mean of probability-vectors is also a probability vector, so the business feature we obtain by taking the mean is a probability vector.  For example, a business feature might look like this:  (70% Plate,  10% Wine,  5% Candle,  15% Others) for a classy-ambience restaurant.</p>\n\n<p>This is my initial motivation before trying other statistics and the more abstract and general FC7 features; we can think each component of the FC7 feature vector as  a score rather than probability.  I hope this makes sense.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 111673,
      "author_name": "FanghaoLuo",
      "author_url": "",
      "post_date": "2016-03-16T02:07:15.320000",
      "content": "<p>Hey guys, I am quite confusing why we should use the mean of all the features (fc7) for one biz to represent this biz. How can this work well? And I think if we just get the mean of all the images (raw data) for on biz to represent this biz, it won't work, right?</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 111285,
      "author_name": "Shouvik Dutta",
      "author_url": "",
      "post_date": "2016-03-13T10:18:06.130000",
      "content": "<p>I'm not sure how much we should go into this here since it's kind of tangential to the original topic, but maybe my confusion is shared by some others so I'll give it a shot.</p>\n\n<p>I would agree with the fact that in my case adding another point could change the decision boundary due to the normalization possibly changing. One way around this would be to standardize instead of normalize, so you'd subtract the mean of each feature and divide by it's standard deviation. This way adding a sample is unlikely to change the decision boundary assuming n is large and you would still have the advantage of a similar scale for each feature. </p>\n\n<p>I guess I see the advantage of normalizing your way since you're combining two different spaces. I wonder if doing both directions would yield any kind of better result...</p>\n\n<p>Thanks for the explanations, it's really helpful.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 110422,
      "author_name": "thebenedict",
      "author_url": "",
      "post_date": "2016-03-05T08:55:48.377000",
      "content": "<p>Thanks very much Nina, I really enjoyed reading through your code.</p>\n\n<p>I wanted to run it end to end but don't have a GPU, so I edited step 1 to use Histogram of Oriented Gradients features [1] instead of CaffeNet. Feature extraction on CPU takes about 8 hours on a  2013 i7 Macbook Pro. The code and a bit more info are in this pull request. [2]</p>\n\n<p>[1] <a href=\"https://en.wikipedia.org/wiki/Histogram_of_oriented_gradients\">https://en.wikipedia.org/wiki/Histogram_of_oriented_gradients</a></p>\n\n<p>[2] <a href=\"https://github.com/ncchen55414/Kaggle-Yelp/pull/1\">https://github.com/ncchen55414/Kaggle-Yelp/pull/1</a></p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 110230,
      "author_name": "Fernando Inakuko",
      "author_url": "",
      "post_date": "2016-03-03T23:48:12.153000",
      "content": "<p>Thanks, Nina and everyone sharing useful tips here. This will be the first time I deal with image recognition using NN, so I'm clueless (now a bit less). Hoping for a top 25% at the end.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 110199,
      "author_name": "Dmitrii Tsybulevskii",
      "author_url": "",
      "post_date": "2016-03-03T18:43:49.007000",
      "content": "<p>[quote=Gino;110136]</p>\n\n<p>@u1234x1234 .. Newbie questions...</p>\n\n<p>Can you use inception's pre trained model the same as way Nina uses Caffe's one?   Inception model seems a lot more complex too.  Which layer could you use?  Not even sure if this is practical...or if i am making any sense....  Any info would be appreciated.\nLooking to learn about this this but still trying to install caffe on my mac....and deal with the depencies mess.\nThanks</p>\n\n<p>[/quote]</p>\n\n<p>Yes, you can use plenty of different models the same way, with small changes, mostly in <a href=\"https://github.com/ncchen55414/Kaggle-Yelp/blob/master/CNN_Submission1/Step1_ImageFeatureFc7.ipynb\">Step 1</a>: image preprocessing, layer name, number of extracted features, etc. Check out <a href=\"https://github.com/BVLC/caffe/wiki/Model-Zoo\">Model Zoo</a> for other models, but you should take into account license type. <a href=\"https://www.kaggle.com/c/yelp-restaurant-photo-classification/forums/t/18212/external-data-requests\">Topic about it</a>.</p>\n\n<p>For models trained on the ImageNet dataset: first layers contain a low level information(about edges and blobs), last layers - high-level, semantic information. I think you should start with last but one layer, and try different variants. Last layer(softmax) can gives result &quot;overfitted&quot; on ImageNet classes, but in this task there's no need to distinguish cats from dogs/</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 109986,
      "author_name": "Dmitrii Tsybulevskii",
      "author_url": "",
      "post_date": "2016-03-01T22:44:06.100000",
      "content": "<p>[quote=Johannes Ahlmann;109865]</p>\n\n<p>Thank you very much @Nina Chen for this!</p>\n\n<p>I found some of the labels from pre-trained CaffeNet very unhelpful (i.e. &quot;food&quot;).\nWhile enticing to use &quot;just&quot; the pretrained label output, I think this will be quite limited in application.</p>\n\n<p>To me the pre-trained classes of CaffeNet seem almost orthogonal to the yelp challenge classes.\nThe &quot;best&quot; descriptor of many of the images are words like &quot;food&quot;, &quot;plate&quot;, &quot;menu&quot;, etc. but what we are actually interested in is the &quot;difference that makes the difference&quot; between the target classes, which may not even be easily put into words in the first place.</p>\n\n<p>What we are probably looking for would be words like &quot;posh&quot;, &quot;up-market&quot;, &quot;clean&quot;, &quot;lobster&quot;, etc.\nBut then again, it would probably be much more valuable to train CaffeNet on the images against the target classes, rather than generic &quot;open world&quot; image descriptions.</p>\n\n<p>One idea would be to go through all of CaffeNet's possible/actual classification labels and hand-pick the ones that seem relevant (&quot;lobster&quot; vs. &quot;food&quot;).</p>\n\n<p>What do you think?</p>\n\n<p>[/quote]</p>\n\n<p>You can cut off last fully connected layers because they are responsible for the actual ImageNet labels, and take features from previous layers, which contains just some semantic information.</p>\n\n<p>Regarding picking some subset of activations, there's quite interesting paper [1]:</p>\n\n<p>&quot;First, we find that there is no distinction between individual high level units and\nrandom linear combinations of high level units, according to various methods of\nunit analysis. It suggests that it is the space, rather than the individual units, that\ncontains the semantic information in the high layers of neural networks.&quot;</p>\n\n<p><a href=\"http://arxiv.org/pdf/1312.6199\">[1]</a> Szegedy, C., Zaremba, W., Sutskever, I., Bruna, J., Erhan, D., Goodfellow, I., &amp; Fergus, R. (2013). Intriguing properties of neural networks.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 109865,
      "author_name": "Johannes Ahlmann",
      "author_url": "",
      "post_date": "2016-03-01T10:04:03.660000",
      "content": "<p>Thank you very much @Nina Chen for this!</p>\n\n<p>I found some of the labels from pre-trained CaffeNet very unhelpful (i.e. &quot;food&quot;).\nWhile enticing to use &quot;just&quot; the pretrained label output, I think this will be quite limited in application.</p>\n\n<p>To me the pre-trained classes of CaffeNet seem almost orthogonal to the yelp challenge classes.\nThe &quot;best&quot; descriptor of many of the images are words like &quot;food&quot;, &quot;plate&quot;, &quot;menu&quot;, etc. but what we are actually interested in is the &quot;difference that makes the difference&quot; between the target classes, which may not even be easily put into words in the first place.</p>\n\n<p>What we are probably looking for would be words like &quot;posh&quot;, &quot;up-market&quot;, &quot;clean&quot;, &quot;lobster&quot;, etc.\nBut then again, it would probably be much more valuable to train CaffeNet on the images against the target classes, rather than generic &quot;open world&quot; image descriptions.</p>\n\n<p>One idea would be to go through all of CaffeNet's possible/actual classification labels and hand-pick the ones that seem relevant (&quot;lobster&quot; vs. &quot;food&quot;).</p>\n\n<p>What do you think?</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 109735,
      "author_name": "SecondPlan",
      "author_url": "",
      "post_date": "2016-02-29T22:46:45.110000",
      "content": "<p>Thanks, @Nina Chen, with fc6 + fc7, I got 0.803 now, tweaking the parameters of svm does not help much. Maybe fine tune the model on this dataset and extract features might help</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 109728,
      "author_name": "Nc Chen",
      "author_url": "",
      "post_date": "2016-02-29T22:28:47.737000",
      "content": "<p>[quote=Blue Light;109624]</p>\n\n<p>Hi there,</p>\n\n<p>Thanks for sharing the code.  I've been trying to run code that uses Caffe (as yours does), but it's been incredibly slow. My machine lacks a good GPU, and renting out GPU instances on AWS can get quite expensive.  Is there some other option to get neural net code to run within a reasonable time period without incurring too much cost?</p>\n\n<p>Thanks</p>\n\n<p>[/quote]</p>\n\n<p>Hi @Blue Light,</p>\n\n<p>It might be faster to use Caffe's <a href=\"http://caffe.berkeleyvision.org/gathered/examples/feature_extraction.html\">feature extractor c++ utility</a> than to use the python wrapper. </p>\n\n<p>However, the features extracted are somehow different. (<a href=\"https://github.com/ncchen55414/Keggle-Yelp/blob/master/CNN_Submission1/Extract_feature_python_vs_cpp.ipynb\">Here</a> is my notebook to read the features extracted from c++ utility and compare  them with the python wrapper.) I'm not sure if it matters. </p>\n\n<p>Hopefully someone more knowledgeable can chime in.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 109617,
      "author_name": "Nc Chen",
      "author_url": "",
      "post_date": "2016-02-28T19:15:14.713000",
      "content": "<p>[quote=Ben Hamner;109581]</p>\n\n<p>Thanks for sharing Nina! I especially enjoyed your t-SNE visualization on the good-for-lunch restaurants. My lunch today resembled the burger cluster in the middle-left.</p>\n\n<p>Out of curiosity, what would it take to get your Jupyter notebooks running on Kaggle Scripts (in terms of computational resources and libraries required?). I'd love to see them there if it's computationally feasible.</p>\n\n<p>[/quote]</p>\n\n<p>Thanks for your comment Ben. Also thanks Kaggle/Yelp for hosting this competition. The variety of food and restaurants I've tried have significantly increased ever since I started working  on this project and look at food image for hours!</p>\n\n<p>As for the computational resources,</p>\n\n<ul>\n<li>Step1: Image feature extraction is run on a GPU (GTX980 with 4GB RAM) for about 2-3 hours. Caffe is required. The features extracted are 4GB each for training and test images. I'll be happy to share them if they are not so huge; this should help people inaccessible to GPU (/and have slow internet so AWS is not an option).</li>\n<li>GPU or Caffe are not required for other steps. I re-run the code on my Macbook Air 2014 ( i5 with 8GB RAM). Step2 takes about 1 hours, and other steps are done in minutes. Scikit-learn is the most crucial library used.</li>\n</ul>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 111186,
      "author_name": "SecondPlan",
      "author_url": "",
      "post_date": "2016-03-12T16:12:47.510000",
      "content": "<p>[quote=Blue Light;111185]</p>\n\n<p>[quote=SecondPlan;109735]</p>\n\n<p>Thanks, @Nina Chen, with fc6 + fc7, I got 0.803 now, tweaking the parameters of svm does not help much. Maybe fine tune the model on this dataset and extract features might help</p>\n\n<p>[/quote]\nSimply extending the fc7 feature vector ((of 4096 features) with the fc6 one  (another set of 4096 features) added no improvement on the test set for me.  </p>\n\n<p>Are there any particular guidelines as to how one could get more bang out of adding fc6 features into the mix?</p>\n\n<p>Thanks</p>\n\n<p>[/quote]</p>\n\n<p>Well, you should normalize the vectors before concat them</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 111069,
      "author_name": "Alexander Bauer",
      "author_url": "",
      "post_date": "2016-03-11T07:36:53.137000",
      "content": "<p>Ayush,\nyou can create custom layers in Python based on caffe.Layer class. There is some example code shipped with caffe, and I just saw that just recently the caffe community added some example code for multi-label classification on the Pascal challenge data:</p>\n\n<p><a href=\"https://github.com/BVLC/caffe/blob/master/examples/pascal-multilabel-with-datalayer.ipynb\">https://github.com/BVLC/caffe/blob/master/examples/pascal-multilabel-with-datalayer.ipynb</a></p>\n\n<p>This should give you a good blueprint for building the layer, and it also demonstrates how to make use of the pretrained alexnet.</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 109982,
      "author_name": "Dmitrii Tsybulevskii",
      "author_url": "",
      "post_date": "2016-03-01T22:08:01.177000",
      "content": "<p>[quote=Blue Light;109624]</p>\n\n<p>Hi there,</p>\n\n<p>Thanks for sharing the code.  I've been trying to run code that uses Caffe (as yours does), but it's been incredibly slow. My machine lacks a good GPU, and renting out GPU instances on AWS can get quite expensive.  Is there some other option to get neural net code to run within a reasonable time period without incurring too much cost?</p>\n\n<p>Thanks</p>\n\n<p>[/quote]</p>\n\n<p>There's no need for a good GPU, if you only want to make predictions with pretrained net. For example Inception network [1](it was state-of-the-art 1 year ago) takes 600ms on a single image with a ordinary CPU. Inceptions architectures much more CPU friendly, compared with the other state-of-the-art(VGG-like[2], etc). So, if you have the CPU with 4 threads, you can process 1/0.6 * 60 * 60 * 24 * 4 = 576.000 images per day. There're only 450.000 images in this comptetion. But you should have optimized BLAS library installed.</p>\n\n<p>As Nina Chen noted, there's difference between features extracted with CPU and GPU, but usually it doesnt matter. GPU with different cuDNN versions can gives different results as well.</p>\n\n<p><a href=\"http://arxiv.org/pdf/1502.03167\">[1]</a> Ioffe, S., &amp; Szegedy, C. (2015). Batch normalization: Accelerating deep network training by reducing internal covariate shift.</p>\n\n<p><a href=\"http://arxiv.org/pdf/1409.1556\">[2]</a> Simonyan, K., &amp; Zisserman, A. (2014). Very deep convolutional networks for large-scale image recognition. </p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 109658,
      "author_name": "Joseph PENG",
      "author_url": "",
      "post_date": "2016-02-29T06:41:39.640000",
      "content": "<p>@Nina Chen (It seems I cannot @ people in Kaggle)</p>\n\n<p>I was trying something like &quot; training deep networks on single image instances with labels inherited from business&quot; and the results are terrible. I should have read every post carefully.</p>\n\n<p>It is my second year here. Chilly!</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 109582,
      "author_name": "Joseph PENG",
      "author_url": "",
      "post_date": "2016-02-28T03:43:48.180000",
      "content": "<p>Up you go!</p>\n\n<p>The notebook introduces</p>\n\n<ol>\n<li>The concept of multiple instance.</li>\n<li>Multiple-label classification SVM.</li>\n</ol>\n\n<p>I have been struggling in dealing with the above two problems in the past two weeks. You are really helping me out.</p>\n\n<p>This is the best thing happened today. The second best thing is that the temp is above 0 in 55414.</p>\n\n<p>Thanks!</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 1892254,
      "author_name": "kimsuhyeok",
      "author_url": "",
      "post_date": "2022-08-10T01:47:59.537000",
      "content": "<p>Thank you for sharing your code. I have a question about Hog_SVM deep learning.<br>\nAt extract_hog_features(train_images), I have error that mismatch with photo_id of train_photo_to_biz_ids and photo name.jpg of train photo file.<br>\nCan you tell me how you solved it?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 112792,
      "author_name": "rakhlin",
      "author_url": "",
      "post_date": "2016-03-23T19:54:42.343000",
      "content": "<p>[quote=Shouvik Dutta;111280]</p>\n\n<p>Huh ok. Yeah I meant normalize for each dim across all rows of examples, since that would account for different features potentially having different output ranges which could affect the decision boundary found by an svm (unless we concede that these different ranges might be important as determined by the original nn). I'm not sure what sample normalization brings since it divides each partition of your feature vector by a constant which would (I think) simply be reflected in the weights learned by your svm.</p>\n\n<p>[/quote]</p>\n\n<p>Normalization or better say feature scaling <strong>across rows</strong> is standard preprocessing for SVM. &quot;The main advantage of scaling is to avoid attributes in greater numeric ranges dominating those in smaller numeric ranges&quot; See explanation from LIBSVM author:  <a href=\"https://www.csie.ntu.edu.tw/~cjlin/papers/guide/guide.pdf\">2.2 Scaling</a> </p>\n\n<p>Per-instance normalization can work too, e.g. for images. &quot;If your data is stationary (i.e., the statistics for each data dimension follow the same distribution), then you might want to consider subtracting the mean-value for each example (computed per-example).&quot; See Andrew Ng explanation: <a href=\"http://deeplearning.stanford.edu/wiki/index.php/Data_Preprocessing\">Per-example mean subtraction</a></p>\n\n<p>I guess that neural net trained on images inherits this stationary in hidden layers.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 112715,
      "author_name": "Ayush Rai",
      "author_url": "",
      "post_date": "2016-03-23T09:14:27.407000",
      "content": "<p>Hi Guys,</p>\n\n<p>I am still struggling to perform Multi-Instance Classification using Caffe. I am confused about which Loss layer should be used for Multi-Instance Classification (for single label I used SoftmaxLoss layer).</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 111287,
      "author_name": "SecondPlan",
      "author_url": "",
      "post_date": "2016-03-13T10:34:34.870000",
      "content": "<p>doing both directions would be a disaster since you completely changed the space and you should test that not just wonder.</p>\n\n<p>Remember <strong>Whatever transformation you do, you will eventually input a n x k matrix to svm, where you have k features, each time you feed a classifier like svm , a 1 x k vector, it changes its weights accordingly, that's how you train a machine learning classifier.</strong></p>\n\n<p>There is no other paradigm to do training </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 111281,
      "author_name": "SecondPlan",
      "author_url": "",
      "post_date": "2016-03-13T09:45:14.200000",
      "content": "<p>let think this way, does adding 1 extra row of data change the decision boundary completely? In your case, yes, since your have to renormalize for that feature. In my case, no, it just add one feature vector in the space, which might not be a big influencer.</p>\n\n<p>If you do other kaggle competitions, when you have a big matrix n x m, and you want to do svd, makes it n x k, this is feature projection or normalizaiton in a vague broad sense, you can doing the change on the columns instead of rows.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 111280,
      "author_name": "Shouvik Dutta",
      "author_url": "",
      "post_date": "2016-03-13T09:34:34.823000",
      "content": "<p>Huh ok. Yeah I meant normalize for each dim across all rows of examples, since that would account for different features potentially having different output ranges which could affect the decision boundary found by an svm (unless we concede that these different ranges might be important as determined by the original nn). I'm not sure what sample normalization brings since it divides each partition of your feature vector by a constant which would (I think) simply be reflected in the weights learned by your svm.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 111278,
      "author_name": "SecondPlan",
      "author_url": "",
      "post_date": "2016-03-13T09:30:08.413000",
      "content": "<p>sorry, I misunderstand your point, but I did sample-wise normalization and concat. You mean, for each dim of 8092, you normalize across all rows of exmaples? You can try to see whether it works or not. But sample-wise normalization did make sense to me since you are simply projecting  features into new spaces for each sample irrelevant of other data. This way, when you add more and more data, you features in the new space might form some decision boundary for the svm to catch.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 111276,
      "author_name": "Shouvik Dutta",
      "author_url": "",
      "post_date": "2016-03-13T09:17:55.737000",
      "content": "<p>Ah, when I read your comment I thought you meant to normalize the fc6 vector and the fc7 vector for each sample and then concatenate those, which seemed to make less sense. Thanks!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 111275,
      "author_name": "SecondPlan",
      "author_url": "",
      "post_date": "2016-03-13T09:15:38.040000",
      "content": "<p>you get it now, but that's what I said, somehow I did not make myself clear</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 111273,
      "author_name": "Shouvik Dutta",
      "author_url": "",
      "post_date": "2016-03-13T09:10:58.807000",
      "content": "<p>Interesting. With that theory wouldn't it make more sense to normalize each feature separately and then end up with a space of |fc6| + |fc7| features, each of which is normalized? Or am I missing some intuition?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 111268,
      "author_name": "SecondPlan",
      "author_url": "",
      "post_date": "2016-03-13T08:47:48.557000",
      "content": "<p>Well, fc6 and fc7 are different vectors, numbers inside fc6 are meaningful only to fc6, if you concat with fc7 ,  this is like bring two vectors in different space together into a new space,  and normalize them is a simple way to make this new space meaningful.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 111267,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-03-13T08:24:57.727000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 111185,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-03-12T16:11:11.083000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 111060,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-03-11T06:00:56.637000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 111040,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-03-10T22:04:52.060000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 110891,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-03-09T07:48:30.377000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 110813,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-03-08T14:01:41.420000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 110807,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-03-08T11:32:07.817000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 110263,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-03-04T05:35:55.157000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 109624,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-02-28T20:35:26.007000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 109619,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-02-28T19:35:51.887000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 109584,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-02-28T03:57:29.550000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 109583,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-02-28T03:55:06.240000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 110136,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-03-03T05:50:54.290000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 114170,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-04-07T22:43:18.220000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 110897,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-03-09T10:52:54.993000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1033645,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-10-01T07:12:39.973000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 111070,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-03-11T07:57:25.107000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "109542": "Hello there,\r\n\r\nSince I've benefited from other Kagglers' code (eg. [Neon code for the Whale Problem][1] and other shared scripts),  I'd like to share my current solution [here][2]. (these are ipython notebooks so anyone can view the results without running the code.)\r\n\r\nThis solution has 3 steps:\r\n\r\nStep 1: Use the pre-trained CaffeNet to extract features from images\r\n\r\nStep 2: For each business, compute the average of its image features. Use this average as the business feature.\r\n\r\nStep3: Train a SVM for multi-label classification and predict.\r\n\r\nThis gives a LB score of  about 0.76.\r\n\r\n**Other thoughts:**\r\nInspired by [Wendy Kan's script][3], I'm also trying to get an intuition of  what an expensive or good-for-lunch restaurant looks like. Attached is a t-SNE visualization picture, where we can see small clusters of piazzas/burgers/sandwiches, exactly those food for a quick lunch bite. A food image classifier might be helpful and fun to play with!\r\n\r\n\r\n  [1]: https://www.kaggle.com/c/noaa-right-whale-recognition/forums/t/17555/try-this\r\n  [2]: https://github.com/ncchen55414/Keggle-Yelp/tree/master/CNN_Submission1\r\n  [3]: https://www.kaggle.com/wendykan/yelp-restaurant-photo-classification/expensive-restaurants-look-like-this",
    "109581": "Thanks for sharing Nina! I especially enjoyed your t-SNE visualization on the good-for-lunch restaurants. My lunch today resembled the burger cluster in the middle-left.\r\n\r\nOut of curiosity, what would it take to get your Jupyter notebooks running on Kaggle Scripts (in terms of computational resources and libraries required?). I'd love to see them there if it's computationally feasible.",
    "111725": "Hey Nina,\r\n\r\nThanks a lot for explanation, I think I get it now. So the \"mean process\" is like a \"voting process\", right? ",
    "111685": "Hi Fanghao,\r\n\r\nWhile fc7-features give better leaderboard score, probability-features are more human-understandable. So below I’ll explain why I take the mean of probability-features of images and use it as a business feature.\r\n\r\nThe ImageNet includes [1000 object classes][1], and the probability-feature is a 1000-dimensional vector, whose i-th component represents the probability of the i-th class. Only a few of the 1000 classes are relevant to restaurants but they can still be informative.\r\n\r\nNotice that the mean of probability-vectors is also a probability vector, so the business feature we obtain by taking the mean is a probability vector.  For example, a business feature might look like this:  (70% Plate,  10% Wine,  5% Candle,  15% Others) for a classy-ambience restaurant.\r\n\r\nThis is my initial motivation before trying other statistics and the more abstract and general FC7 features; we can think each component of the FC7 feature vector as  a score rather than probability.  I hope this makes sense.\r\n\r\n\r\n  [1]: http://image-net.org/challenges/LSVRC/2014/browse-synsets",
    "111673": "Hey guys, I am quite confusing why we should use the mean of all the features (fc7) for one biz to represent this biz. How can this work well? And I think if we just get the mean of all the images (raw data) for on biz to represent this biz, it won't work, right?",
    "111285": "I'm not sure how much we should go into this here since it's kind of tangential to the original topic, but maybe my confusion is shared by some others so I'll give it a shot.\r\n\r\nI would agree with the fact that in my case adding another point could change the decision boundary due to the normalization possibly changing. One way around this would be to standardize instead of normalize, so you'd subtract the mean of each feature and divide by it's standard deviation. This way adding a sample is unlikely to change the decision boundary assuming n is large and you would still have the advantage of a similar scale for each feature. \r\n\r\nI guess I see the advantage of normalizing your way since you're combining two different spaces. I wonder if doing both directions would yield any kind of better result...\r\n\r\nThanks for the explanations, it's really helpful.",
    "110422": "Thanks very much Nina, I really enjoyed reading through your code.\r\n\r\nI wanted to run it end to end but don't have a GPU, so I edited step 1 to use Histogram of Oriented Gradients features [1] instead of CaffeNet. Feature extraction on CPU takes about 8 hours on a  2013 i7 Macbook Pro. The code and a bit more info are in this pull request. [2]\r\n\r\n[1] https://en.wikipedia.org/wiki/Histogram_of_oriented_gradients\r\n\r\n[2] https://github.com/ncchen55414/Kaggle-Yelp/pull/1",
    "110230": "Thanks, Nina and everyone sharing useful tips here. This will be the first time I deal with image recognition using NN, so I'm clueless (now a bit less). Hoping for a top 25% at the end.",
    "110199": "[quote=Gino;110136]\r\n\r\n@u1234x1234 .. Newbie questions...\r\n\r\nCan you use inception's pre trained model the same as way Nina uses Caffe's one?   Inception model seems a lot more complex too.  Which layer could you use?  Not even sure if this is practical...or if i am making any sense....  Any info would be appreciated.\r\nLooking to learn about this this but still trying to install caffe on my mac....and deal with the depencies mess.\r\nThanks\r\n\r\n\r\n[/quote]\r\n\r\nYes, you can use plenty of different models the same way, with small changes, mostly in [Step 1][1]: image preprocessing, layer name, number of extracted features, etc. Check out [Model Zoo][2] for other models, but you should take into account license type. [Topic about it][3].\r\n\r\nFor models trained on the ImageNet dataset: first layers contain a low level information(about edges and blobs), last layers - high-level, semantic information. I think you should start with last but one layer, and try different variants. Last layer(softmax) can gives result \"overfitted\" on ImageNet classes, but in this task there's no need to distinguish cats from dogs/\r\n\r\n  [1]: https://github.com/ncchen55414/Kaggle-Yelp/blob/master/CNN_Submission1/Step1_ImageFeatureFc7.ipynb\r\n  [2]: https://github.com/BVLC/caffe/wiki/Model-Zoo\r\n  [3]: https://www.kaggle.com/c/yelp-restaurant-photo-classification/forums/t/18212/external-data-requests",
    "109986": "[quote=Johannes Ahlmann;109865]\r\n\r\nThank you very much @Nina Chen for this!\r\n\r\nI found some of the labels from pre-trained CaffeNet very unhelpful (i.e. \"food\").\r\nWhile enticing to use \"just\" the pretrained label output, I think this will be quite limited in application.\r\n\r\nTo me the pre-trained classes of CaffeNet seem almost orthogonal to the yelp challenge classes.\r\nThe \"best\" descriptor of many of the images are words like \"food\", \"plate\", \"menu\", etc. but what we are actually interested in is the \"difference that makes the difference\" between the target classes, which may not even be easily put into words in the first place.\r\n\r\nWhat we are probably looking for would be words like \"posh\", \"up-market\", \"clean\", \"lobster\", etc.\r\nBut then again, it would probably be much more valuable to train CaffeNet on the images against the target classes, rather than generic \"open world\" image descriptions.\r\n\r\nOne idea would be to go through all of CaffeNet's possible/actual classification labels and hand-pick the ones that seem relevant (\"lobster\" vs. \"food\").\r\n\r\nWhat do you think?\r\n\r\n[/quote]\r\n\r\nYou can cut off last fully connected layers because they are responsible for the actual ImageNet labels, and take features from previous layers, which contains just some semantic information.\r\n\r\nRegarding picking some subset of activations, there's quite interesting paper \\[1]:\r\n\r\n\"First, we find that there is no distinction between individual high level units and\r\nrandom linear combinations of high level units, according to various methods of\r\nunit analysis. It suggests that it is the space, rather than the individual units, that\r\ncontains the semantic information in the high layers of neural networks.\"\r\n \r\n[\\[1\\]][1] Szegedy, C., Zaremba, W., Sutskever, I., Bruna, J., Erhan, D., Goodfellow, I., & Fergus, R. (2013). Intriguing properties of neural networks.\r\n\r\n\r\n  [1]: http://arxiv.org/pdf/1312.6199",
    "109865": "Thank you very much @Nina Chen for this!\r\n\r\nI found some of the labels from pre-trained CaffeNet very unhelpful (i.e. \"food\").\r\nWhile enticing to use \"just\" the pretrained label output, I think this will be quite limited in application.\r\n\r\nTo me the pre-trained classes of CaffeNet seem almost orthogonal to the yelp challenge classes.\r\nThe \"best\" descriptor of many of the images are words like \"food\", \"plate\", \"menu\", etc. but what we are actually interested in is the \"difference that makes the difference\" between the target classes, which may not even be easily put into words in the first place.\r\n\r\nWhat we are probably looking for would be words like \"posh\", \"up-market\", \"clean\", \"lobster\", etc.\r\nBut then again, it would probably be much more valuable to train CaffeNet on the images against the target classes, rather than generic \"open world\" image descriptions.\r\n\r\nOne idea would be to go through all of CaffeNet's possible/actual classification labels and hand-pick the ones that seem relevant (\"lobster\" vs. \"food\").\r\n\r\nWhat do you think?",
    "109735": "Thanks, @Nina Chen, with fc6 + fc7, I got 0.803 now, tweaking the parameters of svm does not help much. Maybe fine tune the model on this dataset and extract features might help",
    "109728": "[quote=Blue Light;109624]\r\n\r\nHi there,\r\n\r\nThanks for sharing the code.  I've been trying to run code that uses Caffe (as yours does), but it's been incredibly slow. My machine lacks a good GPU, and renting out GPU instances on AWS can get quite expensive.  Is there some other option to get neural net code to run within a reasonable time period without incurring too much cost?\r\n\r\nThanks\r\n\r\n[/quote]\r\n\r\nHi @Blue Light,\r\n\r\nIt might be faster to use Caffe's [feature extractor c++ utility][1] than to use the python wrapper. \r\n\r\nHowever, the features extracted are somehow different. ([Here][2] is my notebook to read the features extracted from c++ utility and compare  them with the python wrapper.) I'm not sure if it matters. \r\n\r\nHopefully someone more knowledgeable can chime in.\r\n\r\n\r\n  [1]: http://caffe.berkeleyvision.org/gathered/examples/feature_extraction.html\r\n  [2]: https://github.com/ncchen55414/Keggle-Yelp/blob/master/CNN_Submission1/Extract_feature_python_vs_cpp.ipynb",
    "109617": "[quote=Ben Hamner;109581]\r\n\r\nThanks for sharing Nina! I especially enjoyed your t-SNE visualization on the good-for-lunch restaurants. My lunch today resembled the burger cluster in the middle-left.\r\n\r\nOut of curiosity, what would it take to get your Jupyter notebooks running on Kaggle Scripts (in terms of computational resources and libraries required?). I'd love to see them there if it's computationally feasible.\r\n\r\n[/quote]\r\n\r\nThanks for your comment Ben. Also thanks Kaggle/Yelp for hosting this competition. The variety of food and restaurants I've tried have significantly increased ever since I started working  on this project and look at food image for hours!\r\n\r\nAs for the computational resources,\r\n\r\n - Step1: Image feature extraction is run on a GPU (GTX980 with 4GB RAM) for about 2-3 hours. Caffe is required. The features extracted are 4GB each for training and test images. I'll be happy to share them if they are not so huge; this should help people inaccessible to GPU (/and have slow internet so AWS is not an option).\r\n - GPU or Caffe are not required for other steps. I re-run the code on my Macbook Air 2014 ( i5 with 8GB RAM). Step2 takes about 1 hours, and other steps are done in minutes. Scikit-learn is the most crucial library used.\r\n\r\n\r\n  [1]: https://github.com/ncchen55414/Keggle-Yelp/blob/master/CNN_Submission1/my_conda_env.txt",
    "111186": "[quote=Blue Light;111185]\r\n\r\n[quote=SecondPlan;109735]\r\n\r\nThanks, @Nina Chen, with fc6 + fc7, I got 0.803 now, tweaking the parameters of svm does not help much. Maybe fine tune the model on this dataset and extract features might help\r\n\r\n[/quote]\r\nSimply extending the fc7 feature vector ((of 4096 features) with the fc6 one  (another set of 4096 features) added no improvement on the test set for me.  \r\n\r\nAre there any particular guidelines as to how one could get more bang out of adding fc6 features into the mix?\r\n\r\nThanks\r\n\r\n[/quote]\r\n\r\nWell, you should normalize the vectors before concat them",
    "111069": "Ayush,\r\nyou can create custom layers in Python based on caffe.Layer class. There is some example code shipped with caffe, and I just saw that just recently the caffe community added some example code for multi-label classification on the Pascal challenge data:\r\n\r\nhttps://github.com/BVLC/caffe/blob/master/examples/pascal-multilabel-with-datalayer.ipynb\r\n\r\nThis should give you a good blueprint for building the layer, and it also demonstrates how to make use of the pretrained alexnet.\r\n",
    "109982": "[quote=Blue Light;109624]\r\n\r\nHi there,\r\n\r\nThanks for sharing the code.  I've been trying to run code that uses Caffe (as yours does), but it's been incredibly slow. My machine lacks a good GPU, and renting out GPU instances on AWS can get quite expensive.  Is there some other option to get neural net code to run within a reasonable time period without incurring too much cost?\r\n\r\nThanks\r\n\r\n[/quote]\r\n\r\nThere's no need for a good GPU, if you only want to make predictions with pretrained net. For example Inception network \\[1](it was state-of-the-art 1 year ago) takes 600ms on a single image with a ordinary CPU. Inceptions architectures much more CPU friendly, compared with the other state-of-the-art(VGG-like\\[2], etc). So, if you have the CPU with 4 threads, you can process 1/0.6 * 60 * 60 * 24 * 4 = 576.000 images per day. There're only 450.000 images in this comptetion. But you should have optimized BLAS library installed.\r\n\r\nAs Nina Chen noted, there's difference between features extracted with CPU and GPU, but usually it doesnt matter. GPU with different cuDNN versions can gives different results as well.\r\n\r\n[\\[1\\]][1] Ioffe, S., & Szegedy, C. (2015). Batch normalization: Accelerating deep network training by reducing internal covariate shift.\r\n\r\n[\\[2\\]][2] Simonyan, K., & Zisserman, A. (2014). Very deep convolutional networks for large-scale image recognition. \r\n\r\n\r\n  [1]: http://arxiv.org/pdf/1502.03167\r\n  [2]: http://arxiv.org/pdf/1409.1556",
    "109658": "@Nina Chen (It seems I cannot @ people in Kaggle)\r\n\r\nI was trying something like \" training deep networks on single image instances with labels inherited from business\" and the results are terrible. I should have read every post carefully.\r\n\r\nIt is my second year here. Chilly!",
    "109582": "Up you go!\r\n\r\nThe notebook introduces\r\n\r\n1. The concept of multiple instance.\r\n2. Multiple-label classification SVM.\r\n\r\nI have been struggling in dealing with the above two problems in the past two weeks. You are really helping me out.\r\n\r\nThis is the best thing happened today. The second best thing is that the temp is above 0 in 55414.\r\n\r\nThanks!",
    "1892254": "Thank you for sharing your code. I have a question about Hog_SVM deep learning.\nAt extract_hog_features(train_images), I have error that mismatch with photo_id of train_photo_to_biz_ids and photo name.jpg of train photo file.\nCan you tell me how you solved it?",
    "112792": "[quote=Shouvik Dutta;111280]\r\n\r\nHuh ok. Yeah I meant normalize for each dim across all rows of examples, since that would account for different features potentially having different output ranges which could affect the decision boundary found by an svm (unless we concede that these different ranges might be important as determined by the original nn). I'm not sure what sample normalization brings since it divides each partition of your feature vector by a constant which would (I think) simply be reflected in the weights learned by your svm.\r\n\r\n[/quote]\r\n\r\nNormalization or better say feature scaling **across rows** is standard preprocessing for SVM. \"The main advantage of scaling is to avoid attributes in greater numeric ranges dominating those in smaller numeric ranges\" See explanation from LIBSVM author:  [2.2 Scaling][1] \r\n\r\nPer-instance normalization can work too, e.g. for images. \"If your data is stationary (i.e., the statistics for each data dimension follow the same distribution), then you might want to consider subtracting the mean-value for each example (computed per-example).\" See Andrew Ng explanation: [Per-example mean subtraction][2]\r\n\r\nI guess that neural net trained on images inherits this stationary in hidden layers.\r\n\r\n\r\n  [1]: https://www.csie.ntu.edu.tw/~cjlin/papers/guide/guide.pdf\r\n  [2]: http://deeplearning.stanford.edu/wiki/index.php/Data_Preprocessing",
    "112715": "Hi Guys,\r\n\r\nI am still struggling to perform Multi-Instance Classification using Caffe. I am confused about which Loss layer should be used for Multi-Instance Classification (for single label I used SoftmaxLoss layer).",
    "111287": "doing both directions would be a disaster since you completely changed the space and you should test that not just wonder.\r\n\r\nRemember **Whatever transformation you do, you will eventually input a n x k matrix to svm, where you have k features, each time you feed a classifier like svm , a 1 x k vector, it changes its weights accordingly, that's how you train a machine learning classifier.**\r\n\r\nThere is no other paradigm to do training \r\n\r\n",
    "111281": "let think this way, does adding 1 extra row of data change the decision boundary completely? In your case, yes, since your have to renormalize for that feature. In my case, no, it just add one feature vector in the space, which might not be a big influencer.\r\n\r\nIf you do other kaggle competitions, when you have a big matrix n x m, and you want to do svd, makes it n x k, this is feature projection or normalizaiton in a vague broad sense, you can doing the change on the columns instead of rows.",
    "111280": "Huh ok. Yeah I meant normalize for each dim across all rows of examples, since that would account for different features potentially having different output ranges which could affect the decision boundary found by an svm (unless we concede that these different ranges might be important as determined by the original nn). I'm not sure what sample normalization brings since it divides each partition of your feature vector by a constant which would (I think) simply be reflected in the weights learned by your svm.",
    "111278": "sorry, I misunderstand your point, but I did sample-wise normalization and concat. You mean, for each dim of 8092, you normalize across all rows of exmaples? You can try to see whether it works or not. But sample-wise normalization did make sense to me since you are simply projecting  features into new spaces for each sample irrelevant of other data. This way, when you add more and more data, you features in the new space might form some decision boundary for the svm to catch.\r\n",
    "111276": "Ah, when I read your comment I thought you meant to normalize the fc6 vector and the fc7 vector for each sample and then concatenate those, which seemed to make less sense. Thanks!",
    "111275": "you get it now, but that's what I said, somehow I did not make myself clear",
    "111273": "Interesting. With that theory wouldn't it make more sense to normalize each feature separately and then end up with a space of |fc6| + |fc7| features, each of which is normalized? Or am I missing some intuition?",
    "111268": "Well, fc6 and fc7 are different vectors, numbers inside fc6 are meaningful only to fc6, if you concat with fc7 ,  this is like bring two vectors in different space together into a new space,  and normalize them is a simple way to make this new space meaningful.",
    "111267": "Why would you want to normalize the vectors? Wouldn't normalizing them have no effect on an SVM since it's just a linear scaling?",
    "111185": "[quote=SecondPlan;109735]\r\n\r\nThanks, @Nina Chen, with fc6 + fc7, I got 0.803 now, tweaking the parameters of svm does not help much. Maybe fine tune the model on this dataset and extract features might help\r\n\r\n[/quote]\r\nSimply extending the fc7 feature vector ((of 4096 features) with the fc6 one  (another set of 4096 features) added no improvement on the test set for me.  \r\n\r\nAre there any particular guidelines as to how one could get more bang out of adding fc6 features into the mix?\r\n\r\nThanks",
    "111060": "I was exactly looking for this. Could you elaborate on how we can create the custom layers you mentioned. Also what type of data layer you used HDF5/LMDB/IMAGEDATA? I will appreciate if you can help.\r\n\r\nThanks",
    "111040": "I actually got multi-instance learning working with caffe, but I had to create a custom layer for data loading and augmentation. It load the image ids and labels from a csv file, loads the images from raw jpgs  and feeds image and label vector to the net as seperate matrices.",
    "110891": "Thanks for the response. I have another doubt regarding Multi-Instance Learning. I will appreciate if you can help. In this challenge we are given 9 different business labels. So in total there can be 512 classes (a vector of different labels corresponding to a class) possible based on which label is present for a particular image but out of this (i.e 512 classes) only 173 classes are present in the training set. Is it possible that while testing we end up getting a class out of the remaining 339 classes i.e. a combination of labels which is not present in the training set or we will get the same combination of labels that are present in the training set.",
    "110813": "As far as I understand, caffe currently does not support populating a lmdb library with images and multiple labels. One way you can try is to map multiple label problems to single label problems, for example, '0 0 0 0 0 0 0 1 1 ' can be convert to its label 3 , which is the decimal value of that binary string. If you still want to train a model with multi label and 9 classes, then I would suggest to use lasagne to do that, there are example codes , just search the example code of kaggle facial keypoints recognition, which is treated as multi label regression problem.",
    "110807": "Hi Guys,\r\n\r\nI want to train a caffe model (of my own) for multiple labels on images instead of using any pre-trained model for this problem. But most of the information on net is based on training a caffe model for images with a single label.\r\n\r\nPlease suggest something.\r\n\r\nThanks",
    "110263": "Are we allowed to use pre-trained models? isn't that technically using outside information?\r\n\r\nedit: never mind, I see that this contest has some special rules. cool stuff!",
    "109624": "Hi there,\r\n\r\nThanks for sharing the code.  I've been trying to run code that uses Caffe (as yours does), but it's been incredibly slow. My machine lacks a good GPU, and renting out GPU instances on AWS can get quite expensive.  Is there some other option to get neural net code to run within a reasonable time period without incurring too much cost?\r\n\r\nThanks",
    "109619": "[quote=Joseph PENG;109582]\r\n\r\nUp you go!\r\n\r\nThe notebook introduces\r\n\r\n1. The concept of multiple instance.\r\n2. Multiple-label classification SVM.\r\n\r\nI have been struggling in dealing with the above two problems in the past two weeks. You are really helping me out.\r\n\r\nThis is the best thing happened today. The second best thing is that the temp is above 0 in 55414.\r\n\r\nThanks!\r\n\r\n[/quote]\r\n\r\nHello Joseph,\r\n\r\nGlad to help! The keyword \"multi-instance\" learning is mentioned in [ody's message][1], before that post I was googling something like \"neural network with several image inputs\" and had no luck.\r\n\r\nP.S. I've been in MN for 6 years and now live in the warm California. :-)\r\n  [1]: https://www.kaggle.com/c/yelp-restaurant-photo-classification/forums/t/18101/welcome/102706#post102706",
    "109584": "A quick question: how do you find out \"multiple instance\" and \"multiple-label\" classification SVM?\r\n\r\nI feel it is not from text book like ESL.",
    "109583": "",
    "110136": "",
    "114170": "Thanks Nina, Thank you guys for support",
    "110897": "Thanks a lot for your sharing!",
    "1033645": "Thanks a lot for your sharing!",
    "111070": "Alright thanks Alexander."
  }
}