{
  "id": 20127,
  "title": "method sharing (3rd place solution)",
  "url": "/competitions/yelp-restaurant-photo-classification/writeups/y-method-sharing-3rd-place-solution",
  "author_name": "",
  "post_date": "2016-04-14T14:47:59.393Z",
  "votes": null,
  "comment_count": 4,
  "views": 1827,
  "content": "",
  "messages": [
    {
      "id": "114874",
      "postDate": "04/14/2016 14:47:59",
      "content": "",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "114876",
      "postDate": "04/14/2016 14:56:06",
      "content": "<p>Congratulations to all top finishers!</p>\n\n<p>Here is a brief description of my single model prediction method:</p>\n\n<p><strong>Photo-level feature extraction</strong> I use MXNet pretrained models that can be downloaded here(<a href=\"http://data.dmlc.ml/\">http://data.dmlc.ml/</a>)</p>\n\n<p><strong>Business-level feature generation</strong>\nFor each business_id, I calculate 60th, 80th, 90th, 95th, 100th percentiles, mean and std deviation of the photo-level features.</p>\n\n<p><strong>Classification</strong>\nI use xgboost to train 9 binary classification models. To capture the dependecies between the labels, I use Classifier Chains (<a href=\"https://en.wikipedia.org/wiki/Classifier_chains\">https://en.wikipedia.org/wiki/Classifier_chains</a>). It improves the F1 score by about 0.003-0.005 compared to Binary Relevance model.</p>\n\n<p><strong>Prediction</strong>\nIf the class probability from the binary classifier is greater than a certain cutoff value, I included the label to the final prediction. For the cutoff value, I used 0.4 (It was determined by experiments).</p>\n\n<p>With this, I got 0.82786 in the private LB. Ensemble of a few models (same method, different pretrained nets and different parameters) scored 0.83115.</p>\n\n<p>All the calculation was done on my laptop without GPU ;)</p>\n\n<p>I'm attaching the code here.</p>",
      "rawMarkdown": "Congratulations to all top finishers!\r\n\r\nHere is a brief description of my single model prediction method:\r\n\r\n**Photo-level feature extraction** I use MXNet pretrained models that can be downloaded here(http://data.dmlc.ml/)\r\n\r\n**Business-level feature generation**\r\nFor each business_id, I calculate 60th, 80th, 90th, 95th, 100th percentiles, mean and std deviation of the photo-level features.\r\n\r\n**Classification**\r\nI use xgboost to train 9 binary classification models. To capture the dependecies between the labels, I use Classifier Chains (https://en.wikipedia.org/wiki/Classifier_chains). It improves the F1 score by about 0.003-0.005 compared to Binary Relevance model.\r\n\r\n**Prediction**\r\nIf the class probability from the binary classifier is greater than a certain cutoff value, I included the label to the final prediction. For the cutoff value, I used 0.4 (It was determined by experiments).\r\n\r\nWith this, I got 0.82786 in the private LB. Ensemble of a few models (same method, different pretrained nets and different parameters) scored 0.83115.\r\n\r\n\r\nAll the calculation was done on my laptop without GPU ;)\r\n\r\nI'm attaching the code here.",
      "votes": null
    },
    {
      "id": "114915",
      "postDate": "04/14/2016 18:23:18",
      "content": "<p>Congratulations and thanks for the code. I can't figure out form R code, do you stack these 5 percentiles, mean and std into single feature set? Lets's say, you extract from MXNet model 1000 image features which translates into 7000 business features?</p>\n\n<p>I have similar approach, but didn't stack them into single data set. Instead I created 2 business sets for every model (using only mean and max) and fused them later on.</p>\n\n<p>P.S. also tried classifier chains and obtained similar gain</p>",
      "rawMarkdown": "Congratulations and thanks for the code. I can't figure out form R code, do you stack these 5 percentiles, mean and std into single feature set? Lets's say, you extract from MXNet model 1000 image features which translates into 7000 business features?\r\n\r\nI have similar approach, but didn't stack them into single data set. Instead I created 2 business sets for every model (using only mean and max) and fused them later on.\r\n\r\nP.S. also tried classifier chains and obtained similar gain",
      "votes": null
    },
    {
      "id": "114942",
      "postDate": "04/14/2016 22:10:37",
      "content": "<p>@rakhlin, I created a single business feature set containing 7n features from n class predictions extracted from the MXNet model.</p>",
      "rawMarkdown": "rakhlin, I created a single business feature set containing 7n features from n class predictions extracted from the MXNet model.",
      "votes": null
    },
    {
      "id": "134789",
      "postDate": "09/08/2016 13:13:44",
      "content": "<p>Amazing, thanks !!</p>",
      "rawMarkdown": "Amazing, thanks !!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 114876,
      "author_name": "tyi2000",
      "author_url": "",
      "post_date": "04/14/2016 14:56:06",
      "content": "<p>Congratulations to all top finishers!</p>\n\n<p>Here is a brief description of my single model prediction method:</p>\n\n<p><strong>Photo-level feature extraction</strong> I use MXNet pretrained models that can be downloaded here(<a href=\"http://data.dmlc.ml/\">http://data.dmlc.ml/</a>)</p>\n\n<p><strong>Business-level feature generation</strong>\nFor each business_id, I calculate 60th, 80th, 90th, 95th, 100th percentiles, mean and std deviation of the photo-level features.</p>\n\n<p><strong>Classification</strong>\nI use xgboost to train 9 binary classification models. To capture the dependecies between the labels, I use Classifier Chains (<a href=\"https://en.wikipedia.org/wiki/Classifier_chains\">https://en.wikipedia.org/wiki/Classifier_chains</a>). It improves the F1 score by about 0.003-0.005 compared to Binary Relevance model.</p>\n\n<p><strong>Prediction</strong>\nIf the class probability from the binary classifier is greater than a certain cutoff value, I included the label to the final prediction. For the cutoff value, I used 0.4 (It was determined by experiments).</p>\n\n<p>With this, I got 0.82786 in the private LB. Ensemble of a few models (same method, different pretrained nets and different parameters) scored 0.83115.</p>\n\n<p>All the calculation was done on my laptop without GPU ;)</p>\n\n<p>I'm attaching the code here.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 114915,
      "author_name": "rakhlin",
      "author_url": "",
      "post_date": "04/14/2016 18:23:18",
      "content": "<p>Congratulations and thanks for the code. I can't figure out form R code, do you stack these 5 percentiles, mean and std into single feature set? Lets's say, you extract from MXNet model 1000 image features which translates into 7000 business features?</p>\n\n<p>I have similar approach, but didn't stack them into single data set. Instead I created 2 business sets for every model (using only mean and max) and fused them later on.</p>\n\n<p>P.S. also tried classifier chains and obtained similar gain</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 114942,
      "author_name": "tyi2000",
      "author_url": "",
      "post_date": "04/14/2016 22:10:37",
      "content": "<p>@rakhlin, I created a single business feature set containing 7n features from n class predictions extracted from the MXNet model.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 134789,
      "author_name": "superds",
      "author_url": "",
      "post_date": "09/08/2016 13:13:44",
      "content": "<p>Amazing, thanks !!</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "114874": "",
    "114876": "Congratulations to all top finishers!\r\n\r\nHere is a brief description of my single model prediction method:\r\n\r\n**Photo-level feature extraction** I use MXNet pretrained models that can be downloaded here(http://data.dmlc.ml/)\r\n\r\n**Business-level feature generation**\r\nFor each business_id, I calculate 60th, 80th, 90th, 95th, 100th percentiles, mean and std deviation of the photo-level features.\r\n\r\n**Classification**\r\nI use xgboost to train 9 binary classification models. To capture the dependecies between the labels, I use Classifier Chains (https://en.wikipedia.org/wiki/Classifier_chains). It improves the F1 score by about 0.003-0.005 compared to Binary Relevance model.\r\n\r\n**Prediction**\r\nIf the class probability from the binary classifier is greater than a certain cutoff value, I included the label to the final prediction. For the cutoff value, I used 0.4 (It was determined by experiments).\r\n\r\nWith this, I got 0.82786 in the private LB. Ensemble of a few models (same method, different pretrained nets and different parameters) scored 0.83115.\r\n\r\n\r\nAll the calculation was done on my laptop without GPU ;)\r\n\r\nI'm attaching the code here.",
    "114915": "Congratulations and thanks for the code. I can't figure out form R code, do you stack these 5 percentiles, mean and std into single feature set? Lets's say, you extract from MXNet model 1000 image features which translates into 7000 business features?\r\n\r\nI have similar approach, but didn't stack them into single data set. Instead I created 2 business sets for every model (using only mean and max) and fused them later on.\r\n\r\nP.S. also tried classifier chains and obtained similar gain",
    "114942": "rakhlin, I created a single business feature set containing 7n features from n class predictions extracted from the MXNet model.",
    "134789": "Amazing, thanks !!"
  },
  "source": "meta"
}