{
  "id": 20114,
  "title": "On multiple-label problem",
  "url": "/competitions/yelp-restaurant-photo-classification/discussion/20114",
  "author_name": "",
  "post_date": "2016-04-13T17:01:57.270Z",
  "votes": 2,
  "comment_count": 7,
  "views": 1267,
  "content": "<p>The most simple approach for multiple-label problem is to treat them as n independent classes. This method scales linearly. However, it suffers from several problems. Obviously, the independent assumption, at the very least, does not work well under complicated cases. </p>\n\n<p>I have not come up with any good ideas to deal with the disadvantages. But yesterday, I came across an interesting paper. At the first glance, their approach is novel and promising.</p>\n\n<p>Check it out!</p>\n\n<p><a href=\"http://link.springer.com/article/10.1007/s10994-011-5276-1\">Compressed labeling on distilled labelsets for multi-label learning</a></p>",
  "messages": [
    {
      "id": "114787",
      "postDate": "04/13/2016 17:01:57",
      "content": "<p>The most simple approach for multiple-label problem is to treat them as n independent classes. This method scales linearly. However, it suffers from several problems. Obviously, the independent assumption, at the very least, does not work well under complicated cases. </p>\n\n<p>I have not come up with any good ideas to deal with the disadvantages. But yesterday, I came across an interesting paper. At the first glance, their approach is novel and promising.</p>\n\n<p>Check it out!</p>\n\n<p><a href=\"http://link.springer.com/article/10.1007/s10994-011-5276-1\">Compressed labeling on distilled labelsets for multi-label learning</a></p>",
      "rawMarkdown": "The most simple approach for multiple-label problem is to treat them as n independent classes. This method scales linearly. However, it suffers from several problems. Obviously, the independent assumption, at the very least, does not work well under complicated cases. \r\n\r\nI have not come up with any good ideas to deal with the disadvantages. But yesterday, I came across an interesting paper. At the first glance, their approach is novel and promising.\r\n\r\n Check it out!\r\n\r\n [Compressed labeling on distilled labelsets for multi-label learning][1]\r\n\r\n\r\n  [1]: http://link.springer.com/article/10.1007/s10994-011-5276-1",
      "votes": null
    },
    {
      "id": "114848",
      "postDate": "04/14/2016 08:35:02",
      "content": "<p>There are bayesian approaches to address label interdependency, like ensembled probabilistic classifier chain, see <a href=\"http://icml2010.haifa.il.ibm.com/papers/589.pdf\">Bayes Optimal Multilabel Classification via Probabilistic Classifier Chains</a>. I approached this issue in several ways. EPCC applied to raw probabilities had a small advantage over independent labels. But making label probabilities meta-features (via CV), stacking them and training another non-linear classifier was more advantageous. What works better depends on tasks, of course. I believe EPCC still has its value, and the paper confirms that. Also, keep in mind that this competition is difficult to draw conclusions because it is subject to data bias (2000 businesses is too few).</p>",
      "rawMarkdown": "There are bayesian approaches to address label interdependency, like ensembled probabilistic classifier chain, see [Bayes Optimal Multilabel Classification via Probabilistic Classifier Chains][1]. I approached this issue in several ways. EPCC applied to raw probabilities had a small advantage over independent labels. But making label probabilities meta-features (via CV), stacking them and training another non-linear classifier was more advantageous. What works better depends on tasks, of course. I believe EPCC still has its value, and the paper confirms that. Also, keep in mind that this competition is difficult to draw conclusions because it is subject to data bias (2000 businesses is too few).\r\n\r\n\r\n  [1]: http://icml2010.haifa.il.ibm.com/papers/589.pdf",
      "votes": null
    },
    {
      "id": "115010",
      "postDate": "04/15/2016 16:04:01",
      "content": "<blockquote>\n  <p>But making label probabilities meta-features (via CV), stacking them\n  and training another non-linear classifier was more advantageous.</p>\n</blockquote>\n\n<p>Any paper recommendation on this idea?</p>\n\n<p>[quote=rakhlin;114848]</p>\n\n<p>There are bayesian approaches to address label interdependency, like ensembled probabilistic classifier chain, see <a href=\"http://icml2010.haifa.il.ibm.com/papers/589.pdf\">Bayes Optimal Multilabel Classification via Probabilistic Classifier Chains</a>. I approached this issue in several ways. EPCC applied to raw probabilities had a small advantage over independent labels. But making label probabilities meta-features (via CV), stacking them and training another non-linear classifier was more advantageous. What works better depends on tasks, of course. I believe EPCC still has its value, and the paper confirms that. Also, keep in mind that this competition is difficult to draw conclusions because it is subject to data bias (2000 businesses is too few).</p>\n\n<p>[/quote]</p>",
      "rawMarkdown": "> But making label probabilities meta-features (via CV), stacking them\r\n> and training another non-linear classifier was more advantageous.\r\n\r\nAny paper recommendation on this idea?\r\n\r\n[quote=rakhlin;114848]\r\n\r\nThere are bayesian approaches to address label interdependency, like ensembled probabilistic classifier chain, see [Bayes Optimal Multilabel Classification via Probabilistic Classifier Chains][1]. I approached this issue in several ways. EPCC applied to raw probabilities had a small advantage over independent labels. But making label probabilities meta-features (via CV), stacking them and training another non-linear classifier was more advantageous. What works better depends on tasks, of course. I believe EPCC still has its value, and the paper confirms that. Also, keep in mind that this competition is difficult to draw conclusions because it is subject to data bias (2000 businesses is too few).\r\n\r\n\r\n  [1]: http://icml2010.haifa.il.ibm.com/papers/589.pdf\r\n\r\n[/quote]",
      "votes": null
    },
    {
      "id": "115013",
      "postDate": "04/15/2016 16:33:05",
      "content": "<p>This is a general technique. Here is code for blending, written by Emanuele Olivetti <a href=\"https://github.com/emanuele/kaggle_pbr/blob/master/blend.py\">https://github.com/emanuele/kaggle_pbr/blob/master/blend.py</a></p>",
      "rawMarkdown": "This is a general technique. Here is code for blending, written by Emanuele Olivetti https://github.com/emanuele/kaggle_pbr/blob/master/blend.py",
      "votes": null
    },
    {
      "id": "115760",
      "postDate": "04/19/2016 22:59:09",
      "content": "<p>I used the following approach:</p>\n\n<p>Stage 1: Train each label independent and obtain predicted labels</p>\n\n<p>Stage2: Add predicted labels to the feature set and train each label for the second time</p>",
      "rawMarkdown": "I used the following approach:\r\n\r\nStage 1: Train each label independent and obtain predicted labels\r\n\r\nStage2: Add predicted labels to the feature set and train each label for the second time",
      "votes": null
    },
    {
      "id": "115761",
      "postDate": "04/19/2016 23:19:01",
      "content": "<p>Does it improve the result?</p>\n\n<p>We could add all but one label into feature set, then predict on the one that left. Repeat.</p>",
      "rawMarkdown": "Does it improve the result?\r\n\r\nWe could add all but one label into feature set, then predict on the one that left. Repeat.",
      "votes": null
    },
    {
      "id": "115764",
      "postDate": "04/19/2016 23:31:15",
      "content": "<p>[quote=ZFTurbo;115760]</p>\n\n<p>I used the following approach:</p>\n\n<p>Stage 1: Train each label independent and obtain predicted labels</p>\n\n<p>Stage2: Add predicted labels to the feature set and train each label for the second time</p>\n\n<p>[/quote]\nThis is similar to the link I posted above - Probabilistic Classifier Chains, - but they add labels one by one and train classifiers successively. In my experience, adding labels to ordinary features has little effect in our problem, probably because of already huge dimensionality. But making labels new meta features via CV and training a new classifier(s) exclusively on them has a positive effect.</p>",
      "rawMarkdown": "[quote=ZFTurbo;115760]\r\n\r\nI used the following approach:\r\n\r\nStage 1: Train each label independent and obtain predicted labels\r\n\r\nStage2: Add predicted labels to the feature set and train each label for the second time\r\n\r\n[/quote]\r\nThis is similar to the link I posted above - Probabilistic Classifier Chains, - but they add labels one by one and train classifiers successively. In my experience, adding labels to ordinary features has little effect in our problem, probably because of already huge dimensionality. But making labels new meta features via CV and training a new classifier(s) exclusively on them has a positive effect.",
      "votes": null
    },
    {
      "id": "115797",
      "postDate": "04/20/2016 07:02:03",
      "content": "<p>[quote=Joseph PENG;115761]</p>\n\n<p>Does it improve the result?</p>\n\n<p>We could add all but one label into feature set, then predict on the one that left. Repeat.</p>\n\n<p>[/quote]</p>\n\n<p>There was a little improvement so I add it to my dataflow. But I didn't test enough to tell for sure.</p>",
      "rawMarkdown": "[quote=Joseph PENG;115761]\r\n\r\nDoes it improve the result?\r\n\r\nWe could add all but one label into feature set, then predict on the one that left. Repeat.\r\n\r\n[/quote]\r\n\r\nThere was a little improvement so I add it to my dataflow. But I didn't test enough to tell for sure.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 114848,
      "author_name": "rakhlin",
      "author_url": "",
      "post_date": "04/14/2016 08:35:02",
      "content": "<p>There are bayesian approaches to address label interdependency, like ensembled probabilistic classifier chain, see <a href=\"http://icml2010.haifa.il.ibm.com/papers/589.pdf\">Bayes Optimal Multilabel Classification via Probabilistic Classifier Chains</a>. I approached this issue in several ways. EPCC applied to raw probabilities had a small advantage over independent labels. But making label probabilities meta-features (via CV), stacking them and training another non-linear classifier was more advantageous. What works better depends on tasks, of course. I believe EPCC still has its value, and the paper confirms that. Also, keep in mind that this competition is difficult to draw conclusions because it is subject to data bias (2000 businesses is too few).</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 115010,
      "author_name": "josephpeng",
      "author_url": "",
      "post_date": "04/15/2016 16:04:01",
      "content": "<blockquote>\n  <p>But making label probabilities meta-features (via CV), stacking them\n  and training another non-linear classifier was more advantageous.</p>\n</blockquote>\n\n<p>Any paper recommendation on this idea?</p>\n\n<p>[quote=rakhlin;114848]</p>\n\n<p>There are bayesian approaches to address label interdependency, like ensembled probabilistic classifier chain, see <a href=\"http://icml2010.haifa.il.ibm.com/papers/589.pdf\">Bayes Optimal Multilabel Classification via Probabilistic Classifier Chains</a>. I approached this issue in several ways. EPCC applied to raw probabilities had a small advantage over independent labels. But making label probabilities meta-features (via CV), stacking them and training another non-linear classifier was more advantageous. What works better depends on tasks, of course. I believe EPCC still has its value, and the paper confirms that. Also, keep in mind that this competition is difficult to draw conclusions because it is subject to data bias (2000 businesses is too few).</p>\n\n<p>[/quote]</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 115013,
      "author_name": "rakhlin",
      "author_url": "",
      "post_date": "04/15/2016 16:33:05",
      "content": "<p>This is a general technique. Here is code for blending, written by Emanuele Olivetti <a href=\"https://github.com/emanuele/kaggle_pbr/blob/master/blend.py\">https://github.com/emanuele/kaggle_pbr/blob/master/blend.py</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 115760,
      "author_name": "zfturbo",
      "author_url": "",
      "post_date": "04/19/2016 22:59:09",
      "content": "<p>I used the following approach:</p>\n\n<p>Stage 1: Train each label independent and obtain predicted labels</p>\n\n<p>Stage2: Add predicted labels to the feature set and train each label for the second time</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 115761,
      "author_name": "josephpeng",
      "author_url": "",
      "post_date": "04/19/2016 23:19:01",
      "content": "<p>Does it improve the result?</p>\n\n<p>We could add all but one label into feature set, then predict on the one that left. Repeat.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 115764,
      "author_name": "rakhlin",
      "author_url": "",
      "post_date": "04/19/2016 23:31:15",
      "content": "<p>[quote=ZFTurbo;115760]</p>\n\n<p>I used the following approach:</p>\n\n<p>Stage 1: Train each label independent and obtain predicted labels</p>\n\n<p>Stage2: Add predicted labels to the feature set and train each label for the second time</p>\n\n<p>[/quote]\nThis is similar to the link I posted above - Probabilistic Classifier Chains, - but they add labels one by one and train classifiers successively. In my experience, adding labels to ordinary features has little effect in our problem, probably because of already huge dimensionality. But making labels new meta features via CV and training a new classifier(s) exclusively on them has a positive effect.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 115797,
      "author_name": "zfturbo",
      "author_url": "",
      "post_date": "04/20/2016 07:02:03",
      "content": "<p>[quote=Joseph PENG;115761]</p>\n\n<p>Does it improve the result?</p>\n\n<p>We could add all but one label into feature set, then predict on the one that left. Repeat.</p>\n\n<p>[/quote]</p>\n\n<p>There was a little improvement so I add it to my dataflow. But I didn't test enough to tell for sure.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "114787": "The most simple approach for multiple-label problem is to treat them as n independent classes. This method scales linearly. However, it suffers from several problems. Obviously, the independent assumption, at the very least, does not work well under complicated cases. \r\n\r\nI have not come up with any good ideas to deal with the disadvantages. But yesterday, I came across an interesting paper. At the first glance, their approach is novel and promising.\r\n\r\n Check it out!\r\n\r\n [Compressed labeling on distilled labelsets for multi-label learning][1]\r\n\r\n\r\n  [1]: http://link.springer.com/article/10.1007/s10994-011-5276-1",
    "114848": "There are bayesian approaches to address label interdependency, like ensembled probabilistic classifier chain, see [Bayes Optimal Multilabel Classification via Probabilistic Classifier Chains][1]. I approached this issue in several ways. EPCC applied to raw probabilities had a small advantage over independent labels. But making label probabilities meta-features (via CV), stacking them and training another non-linear classifier was more advantageous. What works better depends on tasks, of course. I believe EPCC still has its value, and the paper confirms that. Also, keep in mind that this competition is difficult to draw conclusions because it is subject to data bias (2000 businesses is too few).\r\n\r\n\r\n  [1]: http://icml2010.haifa.il.ibm.com/papers/589.pdf",
    "115010": "> But making label probabilities meta-features (via CV), stacking them\r\n> and training another non-linear classifier was more advantageous.\r\n\r\nAny paper recommendation on this idea?\r\n\r\n[quote=rakhlin;114848]\r\n\r\nThere are bayesian approaches to address label interdependency, like ensembled probabilistic classifier chain, see [Bayes Optimal Multilabel Classification via Probabilistic Classifier Chains][1]. I approached this issue in several ways. EPCC applied to raw probabilities had a small advantage over independent labels. But making label probabilities meta-features (via CV), stacking them and training another non-linear classifier was more advantageous. What works better depends on tasks, of course. I believe EPCC still has its value, and the paper confirms that. Also, keep in mind that this competition is difficult to draw conclusions because it is subject to data bias (2000 businesses is too few).\r\n\r\n\r\n  [1]: http://icml2010.haifa.il.ibm.com/papers/589.pdf\r\n\r\n[/quote]",
    "115013": "This is a general technique. Here is code for blending, written by Emanuele Olivetti https://github.com/emanuele/kaggle_pbr/blob/master/blend.py",
    "115760": "I used the following approach:\r\n\r\nStage 1: Train each label independent and obtain predicted labels\r\n\r\nStage2: Add predicted labels to the feature set and train each label for the second time",
    "115761": "Does it improve the result?\r\n\r\nWe could add all but one label into feature set, then predict on the one that left. Repeat.",
    "115764": "[quote=ZFTurbo;115760]\r\n\r\nI used the following approach:\r\n\r\nStage 1: Train each label independent and obtain predicted labels\r\n\r\nStage2: Add predicted labels to the feature set and train each label for the second time\r\n\r\n[/quote]\r\nThis is similar to the link I posted above - Probabilistic Classifier Chains, - but they add labels one by one and train classifiers successively. In my experience, adding labels to ordinary features has little effect in our problem, probably because of already huge dimensionality. But making labels new meta features via CV and training a new classifier(s) exclusively on them has a positive effect.",
    "115797": "[quote=Joseph PENG;115761]\r\n\r\nDoes it improve the result?\r\n\r\nWe could add all but one label into feature set, then predict on the one that left. Repeat.\r\n\r\n[/quote]\r\n\r\nThere was a little improvement so I add it to my dataflow. But I didn't test enough to tell for sure."
  },
  "source": "meta"
}