{
  "id": 19004,
  "title": "Surprisingly good baseline",
  "url": "/competitions/yelp-restaurant-photo-classification/discussion/19004",
  "author_name": "Johannes Ahlmann",
  "post_date": "2016-02-16T15:57:43.263000",
  "votes": 0,
  "comment_count": 0,
  "views": 620,
  "content": "<p>Hi,</p>\n\n<p>I output the frequency of outcomes in the training data to create a submission with all entries predicting the same expected outcome.</p>\n\n<p>Training category frequency:</p>\n\n<pre><code>[('6', 1360),\n ('5', 1249),\n ('8', 1238),\n ('2', 1026),\n ('3', 1003),\n ('1', 993),\n ('0', 671),\n ('7', 572),\n ('4', 547)]\n</code></pre>\n\n<p>I should have run all possible combinations against the evaluation function, but instead I just tried picking categories from the top one by one. Adding the '7' reduces the score a little, but I believe a submission with always predicting all categories would still be very strong ;)</p>\n\n<p>The strongest submission on the public leaderboard was the attached with a surprisingly strong leaderboard score of 0.63791. This is a just a dumb baseline submission, so don't be surprised if it overfits the private leaderboard ;)</p>",
  "messages": [
    {
      "id": 108292,
      "postDate": "2016-02-16T15:57:43.263Z",
      "content": "<p>Hi,</p>\n\n<p>I output the frequency of outcomes in the training data to create a submission with all entries predicting the same expected outcome.</p>\n\n<p>Training category frequency:</p>\n\n<pre><code>[('6', 1360),\n ('5', 1249),\n ('8', 1238),\n ('2', 1026),\n ('3', 1003),\n ('1', 993),\n ('0', 671),\n ('7', 572),\n ('4', 547)]\n</code></pre>\n\n<p>I should have run all possible combinations against the evaluation function, but instead I just tried picking categories from the top one by one. Adding the '7' reduces the score a little, but I believe a submission with always predicting all categories would still be very strong ;)</p>\n\n<p>The strongest submission on the public leaderboard was the attached with a surprisingly strong leaderboard score of 0.63791. This is a just a dumb baseline submission, so don't be surprised if it overfits the private leaderboard ;)</p>",
      "rawMarkdown": "Hi,\r\n\r\nI output the frequency of outcomes in the training data to create a submission with all entries predicting the same expected outcome.\r\n\r\nTraining category frequency:\r\n\r\n    [('6', 1360),\r\n     ('5', 1249),\r\n     ('8', 1238),\r\n     ('2', 1026),\r\n     ('3', 1003),\r\n     ('1', 993),\r\n     ('0', 671),\r\n     ('7', 572),\r\n     ('4', 547)]\r\n\r\nI should have run all possible combinations against the evaluation function, but instead I just tried picking categories from the top one by one. Adding the '7' reduces the score a little, but I believe a submission with always predicting all categories would still be very strong ;)\r\n\r\nThe strongest submission on the public leaderboard was the attached with a surprisingly strong leaderboard score of 0.63791. This is a just a dumb baseline submission, so don't be surprised if it overfits the private leaderboard ;)"
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "108292": "Hi,\r\n\r\nI output the frequency of outcomes in the training data to create a submission with all entries predicting the same expected outcome.\r\n\r\nTraining category frequency:\r\n\r\n    [('6', 1360),\r\n     ('5', 1249),\r\n     ('8', 1238),\r\n     ('2', 1026),\r\n     ('3', 1003),\r\n     ('1', 993),\r\n     ('0', 671),\r\n     ('7', 572),\r\n     ('4', 547)]\r\n\r\nI should have run all possible combinations against the evaluation function, but instead I just tried picking categories from the top one by one. Adding the '7' reduces the score a little, but I believe a submission with always predicting all categories would still be very strong ;)\r\n\r\nThe strongest submission on the public leaderboard was the attached with a surprisingly strong leaderboard score of 0.63791. This is a just a dumb baseline submission, so don't be surprised if it overfits the private leaderboard ;)"
  }
}