{
  "id": 20340,
  "title": "Solution Sharing",
  "url": "/competitions/yelp-restaurant-photo-classification/writeups/prithvi-solution-sharing",
  "author_name": "",
  "post_date": "2016-04-22T18:49:28.840Z",
  "votes": 6,
  "comment_count": 4,
  "views": 2112,
  "content": "<p>I have shared my code here: <a href=\"https://github.com/prith189/Yelp_Restaurant_Photo_Classification\">https://github.com/prith189/Yelp_Restaurant_Photo_Classification</a></p>\n\n<p>Thanks to Nina Chen for the starter <a href=\"https://www.kaggle.com/c/yelp-restaurant-photo-classification/forums/t/19206/deep-learning-starter-code\">code</a> and MXNET team for the pretrained Inception net</p>\n\n<p>My final submission was an enhanced version of this.</p>\n\n<p>Approach:</p>\n\n<p>Method1: Build the business feature set by taking the mean of all the image features. Build 9 binary classification models</p>\n\n<p>Method2: Build the business feature set by first clustering similar images and then taking the mean of similar image sets in each business. Build 9 binary classification models</p>\n\n<p>Ensemble: Use a neural net and use the predictions from method1 and 2. Here the final layer uses 9 nodes and the hope is that we take into account the dependancy between labels.</p>\n\n<p>I have used the image features from the last fully connected layer in the Inception net in the code I've shared. Adding features from the softmax layer improves the score a little bit.</p>\n\n<p>Thanks,\nPrithvi</p>",
  "messages": [
    {
      "id": "116266",
      "postDate": "04/22/2016 18:49:28",
      "content": "<p>I have shared my code here: <a href=\"https://github.com/prith189/Yelp_Restaurant_Photo_Classification\">https://github.com/prith189/Yelp_Restaurant_Photo_Classification</a></p>\n\n<p>Thanks to Nina Chen for the starter <a href=\"https://www.kaggle.com/c/yelp-restaurant-photo-classification/forums/t/19206/deep-learning-starter-code\">code</a> and MXNET team for the pretrained Inception net</p>\n\n<p>My final submission was an enhanced version of this.</p>\n\n<p>Approach:</p>\n\n<p>Method1: Build the business feature set by taking the mean of all the image features. Build 9 binary classification models</p>\n\n<p>Method2: Build the business feature set by first clustering similar images and then taking the mean of similar image sets in each business. Build 9 binary classification models</p>\n\n<p>Ensemble: Use a neural net and use the predictions from method1 and 2. Here the final layer uses 9 nodes and the hope is that we take into account the dependancy between labels.</p>\n\n<p>I have used the image features from the last fully connected layer in the Inception net in the code I've shared. Adding features from the softmax layer improves the score a little bit.</p>\n\n<p>Thanks,\nPrithvi</p>",
      "rawMarkdown": "I have shared my code here: https://github.com/prith189/Yelp_Restaurant_Photo_Classification\r\n\r\nThanks to Nina Chen for the starter [code][1] and MXNET team for the pretrained Inception net\r\n\r\nMy final submission was an enhanced version of this.\r\n\r\nApproach:\r\n\r\nMethod1: Build the business feature set by taking the mean of all the image features. Build 9 binary classification models\r\n\r\nMethod2: Build the business feature set by first clustering similar images and then taking the mean of similar image sets in each business. Build 9 binary classification models\r\n\r\nEnsemble: Use a neural net and use the predictions from method1 and 2. Here the final layer uses 9 nodes and the hope is that we take into account the dependancy between labels.\r\n\r\nI have used the image features from the last fully connected layer in the Inception net in the code I've shared. Adding features from the softmax layer improves the score a little bit.\r\n\r\nThanks,\r\nPrithvi\r\n\r\n\r\n  [1]: https://www.kaggle.com/c/yelp-restaurant-photo-classification/forums/t/19206/deep-learning-starter-code",
      "votes": null
    },
    {
      "id": "116273",
      "postDate": "04/22/2016 20:20:01",
      "content": "<p>I've shared my solution here: <a href=\"https://github.com/bzshang/yelp-photo-classification\">https://github.com/bzshang/yelp-photo-classification</a></p>\n\n<p>I fine-tuned the mxnet Inception network (and modified the output layers to fit the 9 class multi-label problem) and fed the network several epochs of the restaurant image + augmentation (crop, translate, scaling, minor rotations). Used AWS GPU Spot Instance for this. Thankfully this competition was ending soon as the prices started getting very expensive/unpredictable during the last few weeks!</p>\n\n<p>The output features (=1024 from next to last layer) from the net were then used as inputs into support vector, logistic regression, and random forest one-vs-rest classifiers (used scikit-learn here), and final model was ensemble of these.</p>\n\n<p>Fine-tuning the network and then feeding into classical ML models was good enough for top 10 at the time, without much focus on the multi-instance problem or engineering better business features (I just used averages of images within same business). Using random forest in the ML mix and a slightly lower-biased label threshold (~0.46 from CV) were helpful too.</p>\n\n<p>For CNN, I must share this link I found: \n<a href=\"http://lamda.nju.edu.cn/weixs/project/CNNTricks/CNNTricks.html\">http://lamda.nju.edu.cn/weixs/project/CNNTricks/CNNTricks.html</a> It really summarizes the techniques in a succinct way.</p>\n\n<p>Also, big recommendation for mxnet. ImageRecordIter has very good performance.</p>\n\n<p>It is great to see many GPU-free solutions working well, and I'm looking forward to learning more about them!</p>",
      "rawMarkdown": "I've shared my solution here: https://github.com/bzshang/yelp-photo-classification\r\n\r\nI fine-tuned the mxnet Inception network (and modified the output layers to fit the 9 class multi-label problem) and fed the network several epochs of the restaurant image + augmentation (crop, translate, scaling, minor rotations). Used AWS GPU Spot Instance for this. Thankfully this competition was ending soon as the prices started getting very expensive/unpredictable during the last few weeks!\r\n\r\nThe output features (=1024 from next to last layer) from the net were then used as inputs into support vector, logistic regression, and random forest one-vs-rest classifiers (used scikit-learn here), and final model was ensemble of these.\r\n\r\nFine-tuning the network and then feeding into classical ML models was good enough for top 10 at the time, without much focus on the multi-instance problem or engineering better business features (I just used averages of images within same business). Using random forest in the ML mix and a slightly lower-biased label threshold (~0.46 from CV) were helpful too.\r\n\r\nFor CNN, I must share this link I found: \r\nhttp://lamda.nju.edu.cn/weixs/project/CNNTricks/CNNTricks.html It really summarizes the techniques in a succinct way.\r\n\r\nAlso, big recommendation for mxnet. ImageRecordIter has very good performance.\r\n\r\nIt is great to see many GPU-free solutions working well, and I'm looking forward to learning more about them!",
      "votes": null
    },
    {
      "id": "116445",
      "postDate": "04/24/2016 09:47:41",
      "content": "<p>Congratulations, B Shang.\nDo you train SVM, RF, LR on individual image features, or feed to the classifiers business averages? I'd be surprised to know you were able to train SVM on 230,000 instances of 1000 dimensionality.</p>\n\n<p>I'd also like to recommend great articles by Xiu-Shen Wei et al (the author of the link you shared) &quot;Deep Spatial Pyramid Ensemble for Cultural Event Recognition&quot; and &quot;Deep Spatial Pyramid: The Devil is Once Again in the Details&quot;</p>",
      "rawMarkdown": "Congratulations, B Shang.\r\nDo you train SVM, RF, LR on individual image features, or feed to the classifiers business averages? I'd be surprised to know you were able to train SVM on 230,000 instances of 1000 dimensionality.\r\n\r\nI'd also like to recommend great articles by Xiu-Shen Wei et al (the author of the link you shared) \"Deep Spatial Pyramid Ensemble for Cultural Event Recognition\" and \"Deep Spatial Pyramid: The Devil is Once Again in the Details\"",
      "votes": null
    },
    {
      "id": "116524",
      "postDate": "04/25/2016 00:04:54",
      "content": "<p>[quote=rakhlin;116445]</p>\n\n<p>Congratulations, B Shang.\nDo you train SVM, RF, LR on individual image features, or feed to the classifiers business averages? I'd be surprised to know you were able to train SVM on 230,000 instances of 1000 dimensionality.</p>\n\n<p>I'd also like to recommend great articles by Xiu-Shen Wei et al (the author of the link you shared) &quot;Deep Spatial Pyramid Ensemble for Cultural Event Recognition&quot; and &quot;Deep Spatial Pyramid: The Devil is Once Again in the Details&quot;</p>\n\n<p>[/quote]</p>\n\n<p>I fed to the classifiers the business averages. Thanks for the links!</p>",
      "rawMarkdown": "[quote=rakhlin;116445]\r\n\r\nCongratulations, B Shang.\r\nDo you train SVM, RF, LR on individual image features, or feed to the classifiers business averages? I'd be surprised to know you were able to train SVM on 230,000 instances of 1000 dimensionality.\r\n\r\nI'd also like to recommend great articles by Xiu-Shen Wei et al (the author of the link you shared) \"Deep Spatial Pyramid Ensemble for Cultural Event Recognition\" and \"Deep Spatial Pyramid: The Devil is Once Again in the Details\"\r\n\r\n[/quote]\r\n\r\nI fed to the classifiers the business averages. Thanks for the links!",
      "votes": null
    },
    {
      "id": "119373",
      "postDate": "05/09/2016 16:07:36",
      "content": "<p>Code for the #7 solution can be found here: <a href=\"https://github.com/bluelight773/Kaggle_Yelp_Photo_Top10_Solution\">https://github.com/bluelight773/Kaggle_Yelp_Photo_Top10_Solution</a></p>\n\n<p>I believe this is a fairly elegant solution as ultimately the predictions are based on only 1 GBC classifier that relies on the averages of the image features extracted from the following 4 layers:</p>\n\n<ol>\n<li>FC6 of VGGPLACES (MIT Places) </li>\n<li>FC7 of VGGPLACES (MIT Places) </li>\n<li>P3 of Inception V3</li>\n<li>SM of Inception V3</li>\n</ol>\n\n<p>There were certain directions to elaborate on the above with some promise, but for better or worse, hardware and time limitations made it difficult to pursue them.</p>",
      "rawMarkdown": "Code for the #7 solution can be found here: https://github.com/bluelight773/Kaggle_Yelp_Photo_Top10_Solution\r\n\r\nI believe this is a fairly elegant solution as ultimately the predictions are based on only 1 GBC classifier that relies on the averages of the image features extracted from the following 4 layers:\r\n\r\n 1. FC6 of VGGPLACES (MIT Places) \r\n 2. FC7 of VGGPLACES (MIT Places) \r\n 3. P3 of Inception V3\r\n 4. SM of Inception V3\r\n\r\nThere were certain directions to elaborate on the above with some promise, but for better or worse, hardware and time limitations made it difficult to pursue them.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 116273,
      "author_name": "bshang",
      "author_url": "",
      "post_date": "04/22/2016 20:20:01",
      "content": "<p>I've shared my solution here: <a href=\"https://github.com/bzshang/yelp-photo-classification\">https://github.com/bzshang/yelp-photo-classification</a></p>\n\n<p>I fine-tuned the mxnet Inception network (and modified the output layers to fit the 9 class multi-label problem) and fed the network several epochs of the restaurant image + augmentation (crop, translate, scaling, minor rotations). Used AWS GPU Spot Instance for this. Thankfully this competition was ending soon as the prices started getting very expensive/unpredictable during the last few weeks!</p>\n\n<p>The output features (=1024 from next to last layer) from the net were then used as inputs into support vector, logistic regression, and random forest one-vs-rest classifiers (used scikit-learn here), and final model was ensemble of these.</p>\n\n<p>Fine-tuning the network and then feeding into classical ML models was good enough for top 10 at the time, without much focus on the multi-instance problem or engineering better business features (I just used averages of images within same business). Using random forest in the ML mix and a slightly lower-biased label threshold (~0.46 from CV) were helpful too.</p>\n\n<p>For CNN, I must share this link I found: \n<a href=\"http://lamda.nju.edu.cn/weixs/project/CNNTricks/CNNTricks.html\">http://lamda.nju.edu.cn/weixs/project/CNNTricks/CNNTricks.html</a> It really summarizes the techniques in a succinct way.</p>\n\n<p>Also, big recommendation for mxnet. ImageRecordIter has very good performance.</p>\n\n<p>It is great to see many GPU-free solutions working well, and I'm looking forward to learning more about them!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 116445,
      "author_name": "rakhlin",
      "author_url": "",
      "post_date": "04/24/2016 09:47:41",
      "content": "<p>Congratulations, B Shang.\nDo you train SVM, RF, LR on individual image features, or feed to the classifiers business averages? I'd be surprised to know you were able to train SVM on 230,000 instances of 1000 dimensionality.</p>\n\n<p>I'd also like to recommend great articles by Xiu-Shen Wei et al (the author of the link you shared) &quot;Deep Spatial Pyramid Ensemble for Cultural Event Recognition&quot; and &quot;Deep Spatial Pyramid: The Devil is Once Again in the Details&quot;</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 116524,
      "author_name": "bshang",
      "author_url": "",
      "post_date": "04/25/2016 00:04:54",
      "content": "<p>[quote=rakhlin;116445]</p>\n\n<p>Congratulations, B Shang.\nDo you train SVM, RF, LR on individual image features, or feed to the classifiers business averages? I'd be surprised to know you were able to train SVM on 230,000 instances of 1000 dimensionality.</p>\n\n<p>I'd also like to recommend great articles by Xiu-Shen Wei et al (the author of the link you shared) &quot;Deep Spatial Pyramid Ensemble for Cultural Event Recognition&quot; and &quot;Deep Spatial Pyramid: The Devil is Once Again in the Details&quot;</p>\n\n<p>[/quote]</p>\n\n<p>I fed to the classifiers the business averages. Thanks for the links!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 119373,
      "author_name": "bluelight773",
      "author_url": "",
      "post_date": "05/09/2016 16:07:36",
      "content": "<p>Code for the #7 solution can be found here: <a href=\"https://github.com/bluelight773/Kaggle_Yelp_Photo_Top10_Solution\">https://github.com/bluelight773/Kaggle_Yelp_Photo_Top10_Solution</a></p>\n\n<p>I believe this is a fairly elegant solution as ultimately the predictions are based on only 1 GBC classifier that relies on the averages of the image features extracted from the following 4 layers:</p>\n\n<ol>\n<li>FC6 of VGGPLACES (MIT Places) </li>\n<li>FC7 of VGGPLACES (MIT Places) </li>\n<li>P3 of Inception V3</li>\n<li>SM of Inception V3</li>\n</ol>\n\n<p>There were certain directions to elaborate on the above with some promise, but for better or worse, hardware and time limitations made it difficult to pursue them.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "116266": "I have shared my code here: https://github.com/prith189/Yelp_Restaurant_Photo_Classification\r\n\r\nThanks to Nina Chen for the starter [code][1] and MXNET team for the pretrained Inception net\r\n\r\nMy final submission was an enhanced version of this.\r\n\r\nApproach:\r\n\r\nMethod1: Build the business feature set by taking the mean of all the image features. Build 9 binary classification models\r\n\r\nMethod2: Build the business feature set by first clustering similar images and then taking the mean of similar image sets in each business. Build 9 binary classification models\r\n\r\nEnsemble: Use a neural net and use the predictions from method1 and 2. Here the final layer uses 9 nodes and the hope is that we take into account the dependancy between labels.\r\n\r\nI have used the image features from the last fully connected layer in the Inception net in the code I've shared. Adding features from the softmax layer improves the score a little bit.\r\n\r\nThanks,\r\nPrithvi\r\n\r\n\r\n  [1]: https://www.kaggle.com/c/yelp-restaurant-photo-classification/forums/t/19206/deep-learning-starter-code",
    "116273": "I've shared my solution here: https://github.com/bzshang/yelp-photo-classification\r\n\r\nI fine-tuned the mxnet Inception network (and modified the output layers to fit the 9 class multi-label problem) and fed the network several epochs of the restaurant image + augmentation (crop, translate, scaling, minor rotations). Used AWS GPU Spot Instance for this. Thankfully this competition was ending soon as the prices started getting very expensive/unpredictable during the last few weeks!\r\n\r\nThe output features (=1024 from next to last layer) from the net were then used as inputs into support vector, logistic regression, and random forest one-vs-rest classifiers (used scikit-learn here), and final model was ensemble of these.\r\n\r\nFine-tuning the network and then feeding into classical ML models was good enough for top 10 at the time, without much focus on the multi-instance problem or engineering better business features (I just used averages of images within same business). Using random forest in the ML mix and a slightly lower-biased label threshold (~0.46 from CV) were helpful too.\r\n\r\nFor CNN, I must share this link I found: \r\nhttp://lamda.nju.edu.cn/weixs/project/CNNTricks/CNNTricks.html It really summarizes the techniques in a succinct way.\r\n\r\nAlso, big recommendation for mxnet. ImageRecordIter has very good performance.\r\n\r\nIt is great to see many GPU-free solutions working well, and I'm looking forward to learning more about them!",
    "116445": "Congratulations, B Shang.\r\nDo you train SVM, RF, LR on individual image features, or feed to the classifiers business averages? I'd be surprised to know you were able to train SVM on 230,000 instances of 1000 dimensionality.\r\n\r\nI'd also like to recommend great articles by Xiu-Shen Wei et al (the author of the link you shared) \"Deep Spatial Pyramid Ensemble for Cultural Event Recognition\" and \"Deep Spatial Pyramid: The Devil is Once Again in the Details\"",
    "116524": "[quote=rakhlin;116445]\r\n\r\nCongratulations, B Shang.\r\nDo you train SVM, RF, LR on individual image features, or feed to the classifiers business averages? I'd be surprised to know you were able to train SVM on 230,000 instances of 1000 dimensionality.\r\n\r\nI'd also like to recommend great articles by Xiu-Shen Wei et al (the author of the link you shared) \"Deep Spatial Pyramid Ensemble for Cultural Event Recognition\" and \"Deep Spatial Pyramid: The Devil is Once Again in the Details\"\r\n\r\n[/quote]\r\n\r\nI fed to the classifiers the business averages. Thanks for the links!",
    "119373": "Code for the #7 solution can be found here: https://github.com/bluelight773/Kaggle_Yelp_Photo_Top10_Solution\r\n\r\nI believe this is a fairly elegant solution as ultimately the predictions are based on only 1 GBC classifier that relies on the averages of the image features extracted from the following 4 layers:\r\n\r\n 1. FC6 of VGGPLACES (MIT Places) \r\n 2. FC7 of VGGPLACES (MIT Places) \r\n 3. P3 of Inception V3\r\n 4. SM of Inception V3\r\n\r\nThere were certain directions to elaborate on the above with some promise, but for better or worse, hardware and time limitations made it difficult to pursue them."
  },
  "source": "meta"
}