{
  "id": 16341,
  "title": "A Better Benchmark",
  "url": "/competitions/noaa-right-whale-recognition/discussion/16341",
  "author_name": "Paul H",
  "post_date": "2015-09-06T13:19:14.020000",
  "votes": 4,
  "comment_count": 0,
  "views": 680,
  "content": "<p>No, this is not a beat the benchmark script...  Abhishek?</p>\n\n<p>I would like a better benchmark to see whether my modeling attempts do better than picking a random whale from the population, using the training set whale probabilities as representative of the population.</p>\n\n<hr>\n\n<pre><code># Create a better benchmark, using whale probabilities from train.csv\n# should score 5.97927\n\nimport pandas as pd\n\n# read data files\nsample = pd.read_csv('data/sample_submission.csv', index_col='Image')\ntrain = pd.read_csv('data/train.csv')\n\n# calculate probabilities\nwhale_count = train.groupby('whaleID')['whaleID'].count()\nwhale_probs = whale_count/whale_count.sum()\ncols = ','.join(whale_probs.index.values.tolist())\nvals = ','.join(['%.8f' % num for num in whale_probs.values.tolist()])\n\n# write submission file\nf = open('benchmark.csv', 'w')\nf.write('Image,' + cols + '\\n')\nfor fname in sample.index.tolist():\n  f.write(fname + ',' + vals + '\\n')\nf.close()\n</code></pre>",
  "messages": [
    {
      "id": 91695,
      "postDate": "2015-09-06T13:19:14.020Z",
      "content": "<p>No, this is not a beat the benchmark script...  Abhishek?</p>\n\n<p>I would like a better benchmark to see whether my modeling attempts do better than picking a random whale from the population, using the training set whale probabilities as representative of the population.</p>\n\n<hr>\n\n<pre><code># Create a better benchmark, using whale probabilities from train.csv\n# should score 5.97927\n\nimport pandas as pd\n\n# read data files\nsample = pd.read_csv('data/sample_submission.csv', index_col='Image')\ntrain = pd.read_csv('data/train.csv')\n\n# calculate probabilities\nwhale_count = train.groupby('whaleID')['whaleID'].count()\nwhale_probs = whale_count/whale_count.sum()\ncols = ','.join(whale_probs.index.values.tolist())\nvals = ','.join(['%.8f' % num for num in whale_probs.values.tolist()])\n\n# write submission file\nf = open('benchmark.csv', 'w')\nf.write('Image,' + cols + '\\n')\nfor fname in sample.index.tolist():\n  f.write(fname + ',' + vals + '\\n')\nf.close()\n</code></pre>",
      "rawMarkdown": "No, this is not a beat the benchmark script...  Abhishek?\r\n\r\nI would like a better benchmark to see whether my modeling attempts do better than picking a random whale from the population, using the training set whale probabilities as representative of the population.\r\n\r\n---\r\n\r\n    # Create a better benchmark, using whale probabilities from train.csv\r\n    # should score 5.97927\r\n    \r\n    import pandas as pd\r\n    \r\n    # read data files\r\n    sample = pd.read_csv('data/sample_submission.csv', index_col='Image')\r\n    train = pd.read_csv('data/train.csv')\r\n    \r\n    # calculate probabilities\r\n    whale_count = train.groupby('whaleID')['whaleID'].count()\r\n    whale_probs = whale_count/whale_count.sum()\r\n    cols = ','.join(whale_probs.index.values.tolist())\r\n    vals = ','.join(['%.8f' % num for num in whale_probs.values.tolist()])\r\n    \r\n    # write submission file\r\n    f = open('benchmark.csv', 'w')\r\n    f.write('Image,' + cols + '\\n')\r\n    for fname in sample.index.tolist():\r\n      f.write(fname + ',' + vals + '\\n')\r\n    f.close()\r\n\r\n",
      "votes": 4
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "91695": "No, this is not a beat the benchmark script...  Abhishek?\r\n\r\nI would like a better benchmark to see whether my modeling attempts do better than picking a random whale from the population, using the training set whale probabilities as representative of the population.\r\n\r\n---\r\n\r\n    # Create a better benchmark, using whale probabilities from train.csv\r\n    # should score 5.97927\r\n    \r\n    import pandas as pd\r\n    \r\n    # read data files\r\n    sample = pd.read_csv('data/sample_submission.csv', index_col='Image')\r\n    train = pd.read_csv('data/train.csv')\r\n    \r\n    # calculate probabilities\r\n    whale_count = train.groupby('whaleID')['whaleID'].count()\r\n    whale_probs = whale_count/whale_count.sum()\r\n    cols = ','.join(whale_probs.index.values.tolist())\r\n    vals = ','.join(['%.8f' % num for num in whale_probs.values.tolist()])\r\n    \r\n    # write submission file\r\n    f = open('benchmark.csv', 'w')\r\n    f.write('Image,' + cols + '\\n')\r\n    for fname in sample.index.tolist():\r\n      f.write(fname + ',' + vals + '\\n')\r\n    f.close()\r\n\r\n"
  }
}