{
  "id": 2736,
  "title": "Running Basic benchmark",
  "url": "/competitions/predict-closed-questions-on-stack-overflow/discussion/2736",
  "author_name": "",
  "post_date": "2012-09-22T05:28:28.023Z",
  "votes": null,
  "comment_count": 2,
  "views": 1556,
  "content": "<p>In the readme on git it is mentioned that we need only train-sample.csv and public_leaderboard.csv to run the benchmark code. But when I ran the code it says train.csv not available? On what data set is training of basic benchmark is done, and why do we\r\n need botth files train.csv and train-sample.csv</p>",
  "messages": [
    {
      "id": "14693",
      "postDate": "09/22/2012 05:28:28",
      "content": "<p>In the readme on git it is mentioned that we need only train-sample.csv and public_leaderboard.csv to run the benchmark code. But when I ran the code it says train.csv not available? On what data set is training of basic benchmark is done, and why do we\r\n need botth files train.csv and train-sample.csv</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "14706",
      "postDate": "09/22/2012 14:17:06",
      "content": "<p>The majority benchmarks require the larger train.csv unless they're modified.</p>\r\n<p>The models created using the stratified train-sample.csv data are heavily biased away from the &quot;Open&quot; class as there's an equal number of &quot;Open&quot; and not open training examples whilst the real world data is heavily skewed towards &quot;Open&quot;. Hence, to improve\r\n the scores, the probabilities returned from the trained models are scaled according to the distribution found in train.csv.</p>\r\n<p>If you want to avoid grabbing train.csv just for five values, you can use this precomputed value by replacing the call to get the priors where appropriate.</p>\r\n<p>new_priors = [0.00913477057600471, 0.004645859639795308, 0.005200965546050945, 0.9791913907850639, 0.0018270134530850952]</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "14734",
      "postDate": "09/24/2012 02:10:25",
      "content": "<p>Thanks, I can use the pre-calculated values as a workaround for my problem.</p>\r\n<p>I downloaded the full train.csv, but when running basic_benchmark.py I get</p>\r\n<p>File &quot;mydir/StackOverflowChallenge/kaggle/basic<em>benchmark.py&quot;, line 43, in <br>\r\nmain()<br>\r\nFile &quot;mydir/StackOverflowChallenge/kaggle/basic</em>benchmark.py&quot;, line 35, in main<br>\r\nnew<em>priors = cu.get</em>priors(full<em>train</em>file)<br>\r\nFile &quot;mydir/StackOverflowChallenge/kaggle/competition<em>utilities.py&quot;, line 56, in get</em>priors<br>\r\nclosed<em>reasons = [r[14] for r in get</em>reader(file_name)]</p>\r\n<p>_csv.Error: newline inside string</p>\r\n<p>I read (on stack overflow) that this can be caused by having commas or quotes inside a comma-separated string field. Did anyone else have this problem? Is the train.csv file well-formed?\r\n</p>\r\n<p>Strange as i would have expected any parsing error related to , &quot; to be common enough to trigger an error in train-sample.csv as well, if it triggerred a problem in train.csv</p>",
      "rawMarkdown": "",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 14706,
      "author_name": "smerity",
      "author_url": "",
      "post_date": "09/22/2012 14:17:06",
      "content": "<p>The majority benchmarks require the larger train.csv unless they're modified.</p>\r\n<p>The models created using the stratified train-sample.csv data are heavily biased away from the &quot;Open&quot; class as there's an equal number of &quot;Open&quot; and not open training examples whilst the real world data is heavily skewed towards &quot;Open&quot;. Hence, to improve\r\n the scores, the probabilities returned from the trained models are scaled according to the distribution found in train.csv.</p>\r\n<p>If you want to avoid grabbing train.csv just for five values, you can use this precomputed value by replacing the call to get the priors where appropriate.</p>\r\n<p>new_priors = [0.00913477057600471, 0.004645859639795308, 0.005200965546050945, 0.9791913907850639, 0.0018270134530850952]</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 14734,
      "author_name": "darkoram",
      "author_url": "",
      "post_date": "09/24/2012 02:10:25",
      "content": "<p>Thanks, I can use the pre-calculated values as a workaround for my problem.</p>\r\n<p>I downloaded the full train.csv, but when running basic_benchmark.py I get</p>\r\n<p>File &quot;mydir/StackOverflowChallenge/kaggle/basic<em>benchmark.py&quot;, line 43, in <br>\r\nmain()<br>\r\nFile &quot;mydir/StackOverflowChallenge/kaggle/basic</em>benchmark.py&quot;, line 35, in main<br>\r\nnew<em>priors = cu.get</em>priors(full<em>train</em>file)<br>\r\nFile &quot;mydir/StackOverflowChallenge/kaggle/competition<em>utilities.py&quot;, line 56, in get</em>priors<br>\r\nclosed<em>reasons = [r[14] for r in get</em>reader(file_name)]</p>\r\n<p>_csv.Error: newline inside string</p>\r\n<p>I read (on stack overflow) that this can be caused by having commas or quotes inside a comma-separated string field. Did anyone else have this problem? Is the train.csv file well-formed?\r\n</p>\r\n<p>Strange as i would have expected any parsing error related to , &quot; to be common enough to trigger an error in train-sample.csv as well, if it triggerred a problem in train.csv</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "14693": "",
    "14706": "",
    "14734": ""
  },
  "source": "meta"
}