{
  "id": 20112,
  "title": "Multi-Instance SVM (misvm)",
  "url": "/competitions/yelp-restaurant-photo-classification/discussion/20112",
  "author_name": "",
  "post_date": "2016-04-13T16:33:38.987Z",
  "votes": null,
  "comment_count": 4,
  "views": 881,
  "content": "<p>I came across Gary Doran's presumably awesome &quot;misvm&quot; (<a href=\"https://github.com/garydoranjr/misvm\">https://github.com/garydoranjr/misvm</a>) which I believe could be very useful in this competition.</p>\n\n<p>Unfortunately I ran out of time and couldn't try the same yet for the competition, but would assume that it would give a strong improvement over more simple approaches like averaging the vectors within a bag/business.</p>\n\n<p>Did anyone use Multi-Instance SVM, or another Multi-Instance library?</p>\n\n<p>The papers associated with the above library suggest that instance-based and hybrid solutions are very slow with a large number of instances. So my plan would be to start evaluating &quot;SIL&quot;, &quot;NST&quot;, &quot;STK&quot;.</p>\n\n<p>Depending on the outcome I would then evaluate the performance/feasibility for some of the other algorithms.</p>\n\n<p>Note: The algorithms do not implement the classifier base class from sklearn, but they follow a similar convention and from my understanding the &quot;predict&quot; method actually returns &quot;real-valued predictions [-1, 1]&quot; which I believe we could scale to [0, 1] and interpret as probabilities.</p>\n\n<p>If we created a thin abstraction layer to return the above as a new method &quot;predict_proba&quot;, I wonder if we could leverage the same in other sklearn functionality like OneVsRestClassifier, CrossValidation, etc.</p>",
  "messages": [
    {
      "id": "114783",
      "postDate": "04/13/2016 16:33:38",
      "content": "<p>I came across Gary Doran's presumably awesome &quot;misvm&quot; (<a href=\"https://github.com/garydoranjr/misvm\">https://github.com/garydoranjr/misvm</a>) which I believe could be very useful in this competition.</p>\n\n<p>Unfortunately I ran out of time and couldn't try the same yet for the competition, but would assume that it would give a strong improvement over more simple approaches like averaging the vectors within a bag/business.</p>\n\n<p>Did anyone use Multi-Instance SVM, or another Multi-Instance library?</p>\n\n<p>The papers associated with the above library suggest that instance-based and hybrid solutions are very slow with a large number of instances. So my plan would be to start evaluating &quot;SIL&quot;, &quot;NST&quot;, &quot;STK&quot;.</p>\n\n<p>Depending on the outcome I would then evaluate the performance/feasibility for some of the other algorithms.</p>\n\n<p>Note: The algorithms do not implement the classifier base class from sklearn, but they follow a similar convention and from my understanding the &quot;predict&quot; method actually returns &quot;real-valued predictions [-1, 1]&quot; which I believe we could scale to [0, 1] and interpret as probabilities.</p>\n\n<p>If we created a thin abstraction layer to return the above as a new method &quot;predict_proba&quot;, I wonder if we could leverage the same in other sklearn functionality like OneVsRestClassifier, CrossValidation, etc.</p>",
      "rawMarkdown": "I came across Gary Doran's presumably awesome \"misvm\" (https://github.com/garydoranjr/misvm) which I believe could be very useful in this competition.\r\n\r\nUnfortunately I ran out of time and couldn't try the same yet for the competition, but would assume that it would give a strong improvement over more simple approaches like averaging the vectors within a bag/business.\r\n\r\nDid anyone use Multi-Instance SVM, or another Multi-Instance library?\r\n\r\nThe papers associated with the above library suggest that instance-based and hybrid solutions are very slow with a large number of instances. So my plan would be to start evaluating \"SIL\", \"NST\", \"STK\".\r\n\r\nDepending on the outcome I would then evaluate the performance/feasibility for some of the other algorithms.\r\n\r\nNote: The algorithms do not implement the classifier base class from sklearn, but they follow a similar convention and from my understanding the \"predict\" method actually returns \"real-valued predictions [-1, 1]\" which I believe we could scale to [0, 1] and interpret as probabilities.\r\n\r\nIf we created a thin abstraction layer to return the above as a new method \"predict_proba\", I wonder if we could leverage the same in other sklearn functionality like OneVsRestClassifier, CrossValidation, etc.",
      "votes": null
    },
    {
      "id": "114792",
      "postDate": "04/13/2016 18:03:25",
      "content": "<p>I have tried this library, but couldn't get any decent result, mostly due to the bad scalability: it's hard to fit whole dataset in the memory, and training is verrry slow,  i.e. typical problems of kernel methods.</p>",
      "rawMarkdown": "I have tried this library, but couldn't get any decent result, mostly due to the bad scalability: it's hard to fit whole dataset in the memory, and training is verrry slow,  i.e. typical problems of kernel methods.",
      "votes": null
    },
    {
      "id": "114829",
      "postDate": "04/14/2016 02:31:26",
      "content": "<p>Hello u1234x1234, <br>\nThanks for sharing your knowledge and congratulations for winning the competition! </p>\n\n<p>Looking forward to hearing  and learning from your solutions. </p>",
      "rawMarkdown": "Hello u1234x1234,    \r\nThanks for sharing your knowledge and congratulations for winning the competition! \r\n\r\nLooking forward to hearing  and learning from your solutions.",
      "votes": null
    },
    {
      "id": "114850",
      "postDate": "04/14/2016 09:22:45",
      "content": "<p>Are you guys sure this problem formally fits Multi-Instance Learning? No single photo can be classified in terms of business labeling (maybe with exception of 5 category &quot;has_alcohol&quot;)</p>",
      "rawMarkdown": "Are you guys sure this problem formally fits Multi-Instance Learning? No single photo can be classified in terms of business labeling (maybe with exception of 5 category \"has_alcohol\")",
      "votes": null
    },
    {
      "id": "114851",
      "postDate": "04/14/2016 10:16:45",
      "content": "<p>I tried MISVM , Multiple-Instance Support Vector Machines algorithm by Gary Doran. For Bag level prediction ( for each business_id ) ,  Normalised set kernel (NSK) turned to perform well with feature space normalized kernel setting (linear_fs). To use MISVM , problem statement need to be treated little differently. MISVM predicts each label seperately , we need to combine the positives for business_id label prediction for final prediction. Though F1 score (for each label ) was around 0.8 , but when combined to predict for business_id result came down drastically on submission. One of my submission scored 0.56 on the same. </p>\n\n<p>u1234x1234 is right , it would not fit in the memory with setting for instance level prediction methods. In this competition , bag level prediction was required. So , NSK would have done the job with no Memory Error and much less training time.</p>\n\n<p><a href=\"https://etd.ohiolink.edu/!etd.send_file?accession=case1417736923&disposition=inline\">https://etd.ohiolink.edu/!etd.send_file?accession=case1417736923&amp;disposition=inline</a></p>\n\n<p>P.S.  Congratulation u1234x1234   :)</p>",
      "rawMarkdown": "I tried MISVM , Multiple-Instance Support Vector Machines algorithm by Gary Doran. For Bag level prediction ( for each business_id ) ,  Normalised set kernel (NSK) turned to perform well with feature space normalized kernel setting (linear_fs). To use MISVM , problem statement need to be treated little differently. MISVM predicts each label seperately , we need to combine the positives for business_id label prediction for final prediction. Though F1 score (for each label ) was around 0.8 , but when combined to predict for business_id result came down drastically on submission. One of my submission scored 0.56 on the same. \r\n\r\nu1234x1234 is right , it would not fit in the memory with setting for instance level prediction methods. In this competition , bag level prediction was required. So , NSK would have done the job with no Memory Error and much less training time.\r\n\r\nhttps://etd.ohiolink.edu/!etd.send_file?accession=case1417736923&disposition=inline\r\n\r\nP.S.  Congratulation u1234x1234   :)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 114792,
      "author_name": "u1234x1234",
      "author_url": "",
      "post_date": "04/13/2016 18:03:25",
      "content": "<p>I have tried this library, but couldn't get any decent result, mostly due to the bad scalability: it's hard to fit whole dataset in the memory, and training is verrry slow,  i.e. typical problems of kernel methods.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 114829,
      "author_name": "ncchen",
      "author_url": "",
      "post_date": "04/14/2016 02:31:26",
      "content": "<p>Hello u1234x1234, <br>\nThanks for sharing your knowledge and congratulations for winning the competition! </p>\n\n<p>Looking forward to hearing  and learning from your solutions. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 114850,
      "author_name": "rakhlin",
      "author_url": "",
      "post_date": "04/14/2016 09:22:45",
      "content": "<p>Are you guys sure this problem formally fits Multi-Instance Learning? No single photo can be classified in terms of business labeling (maybe with exception of 5 category &quot;has_alcohol&quot;)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 114851,
      "author_name": "yardstick17",
      "author_url": "",
      "post_date": "04/14/2016 10:16:45",
      "content": "<p>I tried MISVM , Multiple-Instance Support Vector Machines algorithm by Gary Doran. For Bag level prediction ( for each business_id ) ,  Normalised set kernel (NSK) turned to perform well with feature space normalized kernel setting (linear_fs). To use MISVM , problem statement need to be treated little differently. MISVM predicts each label seperately , we need to combine the positives for business_id label prediction for final prediction. Though F1 score (for each label ) was around 0.8 , but when combined to predict for business_id result came down drastically on submission. One of my submission scored 0.56 on the same. </p>\n\n<p>u1234x1234 is right , it would not fit in the memory with setting for instance level prediction methods. In this competition , bag level prediction was required. So , NSK would have done the job with no Memory Error and much less training time.</p>\n\n<p><a href=\"https://etd.ohiolink.edu/!etd.send_file?accession=case1417736923&disposition=inline\">https://etd.ohiolink.edu/!etd.send_file?accession=case1417736923&amp;disposition=inline</a></p>\n\n<p>P.S.  Congratulation u1234x1234   :)</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "114783": "I came across Gary Doran's presumably awesome \"misvm\" (https://github.com/garydoranjr/misvm) which I believe could be very useful in this competition.\r\n\r\nUnfortunately I ran out of time and couldn't try the same yet for the competition, but would assume that it would give a strong improvement over more simple approaches like averaging the vectors within a bag/business.\r\n\r\nDid anyone use Multi-Instance SVM, or another Multi-Instance library?\r\n\r\nThe papers associated with the above library suggest that instance-based and hybrid solutions are very slow with a large number of instances. So my plan would be to start evaluating \"SIL\", \"NST\", \"STK\".\r\n\r\nDepending on the outcome I would then evaluate the performance/feasibility for some of the other algorithms.\r\n\r\nNote: The algorithms do not implement the classifier base class from sklearn, but they follow a similar convention and from my understanding the \"predict\" method actually returns \"real-valued predictions [-1, 1]\" which I believe we could scale to [0, 1] and interpret as probabilities.\r\n\r\nIf we created a thin abstraction layer to return the above as a new method \"predict_proba\", I wonder if we could leverage the same in other sklearn functionality like OneVsRestClassifier, CrossValidation, etc.",
    "114792": "I have tried this library, but couldn't get any decent result, mostly due to the bad scalability: it's hard to fit whole dataset in the memory, and training is verrry slow,  i.e. typical problems of kernel methods.",
    "114829": "Hello u1234x1234,    \r\nThanks for sharing your knowledge and congratulations for winning the competition! \r\n\r\nLooking forward to hearing  and learning from your solutions.",
    "114850": "Are you guys sure this problem formally fits Multi-Instance Learning? No single photo can be classified in terms of business labeling (maybe with exception of 5 category \"has_alcohol\")",
    "114851": "I tried MISVM , Multiple-Instance Support Vector Machines algorithm by Gary Doran. For Bag level prediction ( for each business_id ) ,  Normalised set kernel (NSK) turned to perform well with feature space normalized kernel setting (linear_fs). To use MISVM , problem statement need to be treated little differently. MISVM predicts each label seperately , we need to combine the positives for business_id label prediction for final prediction. Though F1 score (for each label ) was around 0.8 , but when combined to predict for business_id result came down drastically on submission. One of my submission scored 0.56 on the same. \r\n\r\nu1234x1234 is right , it would not fit in the memory with setting for instance level prediction methods. In this competition , bag level prediction was required. So , NSK would have done the job with no Memory Error and much less training time.\r\n\r\nhttps://etd.ohiolink.edu/!etd.send_file?accession=case1417736923&disposition=inline\r\n\r\nP.S.  Congratulation u1234x1234   :)"
  },
  "source": "meta"
}