{
  "id": 20329,
  "title": "What methods do you use?",
  "url": "/competitions/yelp-restaurant-photo-classification/discussion/20329",
  "author_name": "",
  "post_date": "2016-04-22T08:00:09.667Z",
  "votes": 1,
  "comment_count": 3,
  "views": 794,
  "content": "<p><a href=\"http://158.109.8.37/files/Amo2013.pdf\">Amores&#8217; survey paper</a> categorizes methods for solving multi-instance classification problems into three paradigms:</p>\n\n<p>(1) Instance Space paradigm, where each instance is associated with a label (positive/negative) or a score. Example: MI-SVM mentioned in <a href=\"https://www.kaggle.com/c/yelp-restaurant-photo-classification/forums/t/20112/multi-instance-svm-misvm\">this thread</a>.</p>\n\n<p>(2) Embedded Space paradigm, where each bag is represented by a feature vector. Example: SimpleMI.</p>\n\n<p>(3) Bag Space paradigm, where we discriminate two bags by defining a distance function between two bags. Discussed below.</p>\n\n<p>The paper compares the performance of more than 12 methods on 10 data sets and concludes that, while no method is absolutely better than others, all methods in the Instance Space paradigm perform worse than the EarthMover&#8217;sDistance+SVM method. </p>\n\n<p><img src=\"http://i.imgur.com/eEnhLra.png//\" alt=\"Crop Paper\" title></p>\n\n<p>I've tried <a href=\"https://github.com/ncchen55414/Kaggle-Yelp/tree/master/BagDistance_Submission1\">Chamfer distance (much simpler than EMD)+SVM</a>. The resulted private public scores ranged from 0.818 to 0.822. I&#8217;m wondering if anyone has got better scores with a better metric to measure similarities between restaurants?</p>\n\n<p>I&#8217;m also very curious about single-model performance of other methods, whether successful or unsuccessful. Please share your methods! Thanks.</p>",
  "messages": [
    {
      "id": "116141",
      "postDate": "04/22/2016 08:00:09",
      "content": "<p><a href=\"http://158.109.8.37/files/Amo2013.pdf\">Amores&#8217; survey paper</a> categorizes methods for solving multi-instance classification problems into three paradigms:</p>\n\n<p>(1) Instance Space paradigm, where each instance is associated with a label (positive/negative) or a score. Example: MI-SVM mentioned in <a href=\"https://www.kaggle.com/c/yelp-restaurant-photo-classification/forums/t/20112/multi-instance-svm-misvm\">this thread</a>.</p>\n\n<p>(2) Embedded Space paradigm, where each bag is represented by a feature vector. Example: SimpleMI.</p>\n\n<p>(3) Bag Space paradigm, where we discriminate two bags by defining a distance function between two bags. Discussed below.</p>\n\n<p>The paper compares the performance of more than 12 methods on 10 data sets and concludes that, while no method is absolutely better than others, all methods in the Instance Space paradigm perform worse than the EarthMover&#8217;sDistance+SVM method. </p>\n\n<p><img src=\"http://i.imgur.com/eEnhLra.png//\" alt=\"Crop Paper\" title></p>\n\n<p>I've tried <a href=\"https://github.com/ncchen55414/Kaggle-Yelp/tree/master/BagDistance_Submission1\">Chamfer distance (much simpler than EMD)+SVM</a>. The resulted private public scores ranged from 0.818 to 0.822. I&#8217;m wondering if anyone has got better scores with a better metric to measure similarities between restaurants?</p>\n\n<p>I&#8217;m also very curious about single-model performance of other methods, whether successful or unsuccessful. Please share your methods! Thanks.</p>",
      "rawMarkdown": "[Amores’ survey paper][1] categorizes methods for solving multi-instance classification problems into three paradigms:\r\n\r\n(1) Instance Space paradigm, where each instance is associated with a label (positive/negative) or a score. Example: MI-SVM mentioned in [this thread][2].\r\n\r\n(2) Embedded Space paradigm, where each bag is represented by a feature vector. Example: SimpleMI.\r\n\r\n(3) Bag Space paradigm, where we discriminate two bags by defining a distance function between two bags. Discussed below.\r\n\r\nThe paper compares the performance of more than 12 methods on 10 data sets and concludes that, while no method is absolutely better than others, all methods in the Instance Space paradigm perform worse than the EarthMover’sDistance+SVM method. \r\n\r\n![Crop Paper][3]\r\n\r\nI've tried [Chamfer distance (much simpler than EMD)+SVM][4]. The resulted private public scores ranged from 0.818 to 0.822. I’m wondering if anyone has got better scores with a better metric to measure similarities between restaurants?\r\n\r\nI’m also very curious about single-model performance of other methods, whether successful or unsuccessful. Please share your methods! Thanks.\r\n\r\n\r\n  [1]: http://158.109.8.37/files/Amo2013.pdf\r\n  [2]: https://www.kaggle.com/c/yelp-restaurant-photo-classification/forums/t/20112/multi-instance-svm-misvm\r\n  [3]: http://i.imgur.com/eEnhLra.png//\r\n  [4]: https://github.com/ncchen55414/Kaggle-Yelp/tree/master/BagDistance_Submission1",
      "votes": null
    },
    {
      "id": "116247",
      "postDate": "04/22/2016 17:37:25",
      "content": "<p>I am not quite sure this problem falls into multi-instance classification category. Has anyone tried tracing classifier output (e.g. &quot;good for kids&quot;) to the input images that prompted the category most? Do these images present any tractable relevance to the category?</p>",
      "rawMarkdown": "I am not quite sure this problem falls into multi-instance classification category. Has anyone tried tracing classifier output (e.g. \"good for kids\") to the input images that prompted the category most? Do these images present any tractable relevance to the category?",
      "votes": null
    },
    {
      "id": "116264",
      "postDate": "04/22/2016 18:47:10",
      "content": "<p>I treat this problem as a multi-instance learning problem, in the same spirit as <a href=\"http://cseweb.ucsd.edu/~jfoulds/FouldsAndFrankMIreview.pdf#page=5\">A Review of Multi-instance Learning Assumptions</a>:\n<em>We contend that the term &quot;multi-instance learning&quot; should contrast directly with &quot;single instance\nlearning&quot;, and connotes any type of learning where several instances can be included\nwithin a single learning example, regardless of the assumptions used.</em></p>\n\n<p>This paper summarizes different assumptions used for multi-instance learning problems. Some of them (Instance Space paradigm) assume that each instance has a hidden class label, and some of them don't. </p>\n\n<p>I agree with you that most single photos cannot be classified into business labels, so I didn't try any methods in the Instance Space paradigm. Though <a href=\"https://www.kaggle.com/c/yelp-restaurant-photo-classification/forums/t/18762/help-some-ideas-on-getting-started/108748#post108748\">others</a> seem to get decent results with that. </p>",
      "rawMarkdown": "I treat this problem as a multi-instance learning problem, in the same spirit as [A Review of Multi-instance Learning Assumptions][1]:\r\n*We contend that the term \"multi-instance learning\" should contrast directly with \"single instance\r\nlearning\", and connotes any type of learning where several instances can be included\r\nwithin a single learning example, regardless of the assumptions used.*\r\n\r\nThis paper summarizes different assumptions used for multi-instance learning problems. Some of them (Instance Space paradigm) assume that each instance has a hidden class label, and some of them don't. \r\n\r\nI agree with you that most single photos cannot be classified into business labels, so I didn't try any methods in the Instance Space paradigm. Though [others][2] seem to get decent results with that. \r\n\r\n\r\n  [1]: http://cseweb.ucsd.edu/~jfoulds/FouldsAndFrankMIreview.pdf#page=5\r\n  [2]: https://www.kaggle.com/c/yelp-restaurant-photo-classification/forums/t/18762/help-some-ideas-on-getting-started/108748#post108748",
      "votes": null
    },
    {
      "id": "116276",
      "postDate": "04/22/2016 20:37:36",
      "content": "<p>I've shared by solution here: <a href=\"https://www.kaggle.com/c/yelp-restaurant-photo-classification/forums/t/20340/solution-sharing/116273#post116273\">https://www.kaggle.com/c/yelp-restaurant-photo-classification/forums/t/20340/solution-sharing/116273#post116273</a></p>\n\n<p>Multi-instance problem is interesting. Looking forward to learning more about other methods.</p>",
      "rawMarkdown": "I've shared by solution here: https://www.kaggle.com/c/yelp-restaurant-photo-classification/forums/t/20340/solution-sharing/116273#post116273\r\n\r\nMulti-instance problem is interesting. Looking forward to learning more about other methods.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 116247,
      "author_name": "rakhlin",
      "author_url": "",
      "post_date": "04/22/2016 17:37:25",
      "content": "<p>I am not quite sure this problem falls into multi-instance classification category. Has anyone tried tracing classifier output (e.g. &quot;good for kids&quot;) to the input images that prompted the category most? Do these images present any tractable relevance to the category?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 116264,
      "author_name": "ncchen",
      "author_url": "",
      "post_date": "04/22/2016 18:47:10",
      "content": "<p>I treat this problem as a multi-instance learning problem, in the same spirit as <a href=\"http://cseweb.ucsd.edu/~jfoulds/FouldsAndFrankMIreview.pdf#page=5\">A Review of Multi-instance Learning Assumptions</a>:\n<em>We contend that the term &quot;multi-instance learning&quot; should contrast directly with &quot;single instance\nlearning&quot;, and connotes any type of learning where several instances can be included\nwithin a single learning example, regardless of the assumptions used.</em></p>\n\n<p>This paper summarizes different assumptions used for multi-instance learning problems. Some of them (Instance Space paradigm) assume that each instance has a hidden class label, and some of them don't. </p>\n\n<p>I agree with you that most single photos cannot be classified into business labels, so I didn't try any methods in the Instance Space paradigm. Though <a href=\"https://www.kaggle.com/c/yelp-restaurant-photo-classification/forums/t/18762/help-some-ideas-on-getting-started/108748#post108748\">others</a> seem to get decent results with that. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 116276,
      "author_name": "bshang",
      "author_url": "",
      "post_date": "04/22/2016 20:37:36",
      "content": "<p>I've shared by solution here: <a href=\"https://www.kaggle.com/c/yelp-restaurant-photo-classification/forums/t/20340/solution-sharing/116273#post116273\">https://www.kaggle.com/c/yelp-restaurant-photo-classification/forums/t/20340/solution-sharing/116273#post116273</a></p>\n\n<p>Multi-instance problem is interesting. Looking forward to learning more about other methods.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "116141": "[Amores’ survey paper][1] categorizes methods for solving multi-instance classification problems into three paradigms:\r\n\r\n(1) Instance Space paradigm, where each instance is associated with a label (positive/negative) or a score. Example: MI-SVM mentioned in [this thread][2].\r\n\r\n(2) Embedded Space paradigm, where each bag is represented by a feature vector. Example: SimpleMI.\r\n\r\n(3) Bag Space paradigm, where we discriminate two bags by defining a distance function between two bags. Discussed below.\r\n\r\nThe paper compares the performance of more than 12 methods on 10 data sets and concludes that, while no method is absolutely better than others, all methods in the Instance Space paradigm perform worse than the EarthMover’sDistance+SVM method. \r\n\r\n![Crop Paper][3]\r\n\r\nI've tried [Chamfer distance (much simpler than EMD)+SVM][4]. The resulted private public scores ranged from 0.818 to 0.822. I’m wondering if anyone has got better scores with a better metric to measure similarities between restaurants?\r\n\r\nI’m also very curious about single-model performance of other methods, whether successful or unsuccessful. Please share your methods! Thanks.\r\n\r\n\r\n  [1]: http://158.109.8.37/files/Amo2013.pdf\r\n  [2]: https://www.kaggle.com/c/yelp-restaurant-photo-classification/forums/t/20112/multi-instance-svm-misvm\r\n  [3]: http://i.imgur.com/eEnhLra.png//\r\n  [4]: https://github.com/ncchen55414/Kaggle-Yelp/tree/master/BagDistance_Submission1",
    "116247": "I am not quite sure this problem falls into multi-instance classification category. Has anyone tried tracing classifier output (e.g. \"good for kids\") to the input images that prompted the category most? Do these images present any tractable relevance to the category?",
    "116264": "I treat this problem as a multi-instance learning problem, in the same spirit as [A Review of Multi-instance Learning Assumptions][1]:\r\n*We contend that the term \"multi-instance learning\" should contrast directly with \"single instance\r\nlearning\", and connotes any type of learning where several instances can be included\r\nwithin a single learning example, regardless of the assumptions used.*\r\n\r\nThis paper summarizes different assumptions used for multi-instance learning problems. Some of them (Instance Space paradigm) assume that each instance has a hidden class label, and some of them don't. \r\n\r\nI agree with you that most single photos cannot be classified into business labels, so I didn't try any methods in the Instance Space paradigm. Though [others][2] seem to get decent results with that. \r\n\r\n\r\n  [1]: http://cseweb.ucsd.edu/~jfoulds/FouldsAndFrankMIreview.pdf#page=5\r\n  [2]: https://www.kaggle.com/c/yelp-restaurant-photo-classification/forums/t/18762/help-some-ideas-on-getting-started/108748#post108748",
    "116276": "I've shared by solution here: https://www.kaggle.com/c/yelp-restaurant-photo-classification/forums/t/20340/solution-sharing/116273#post116273\r\n\r\nMulti-instance problem is interesting. Looking forward to learning more about other methods."
  },
  "source": "meta"
}