{
  "id": 30998,
  "title": "Small Amount of Train Data",
  "url": "/competitions/noaa-fisheries-steller-sea-lion-population-count/discussion/30998",
  "author_name": "RichardHerbert",
  "post_date": "2017-04-01T22:16:12.282000",
  "votes": 0,
  "comment_count": 2,
  "views": 0,
  "content": "<p>Does the training data seem small to anyone else? The Train folder has less than 1,000 images, while the test has almost over 18,000. Am I missing something? That seems like an awfully small dataset to train to high accuracy, while providing nearly 20,000 test examples. </p>",
  "messages": [
    {
      "id": 172131,
      "postDate": "2017-04-02T05:10:33.910Z",
      "content": "<p>Hi RichardHerbert,</p>\n\n<p>Remember that test dataset include dummy images. Some part of 20,000 test images are used for evaluation.\nIn Data description page, it is written as follows, \n<code>As an anti-cheating measure, Kaggle has added some extra images to the test set which will not be counted towards your score.</code></p>",
      "rawMarkdown": "Hi RichardHerbert,\n\nRemember that test dataset include dummy images. Some part of 20,000 test images are used for evaluation.\nIn Data description page, it is written as follows, \n`As an anti-cheating measure, Kaggle has added some extra images to the test set which will not be counted towards your score.`"
    },
    {
      "id": 172096,
      "postDate": "2017-04-01T22:16:12.283Z",
      "content": "<p>Does the training data seem small to anyone else? The Train folder has less than 1,000 images, while the test has almost over 18,000. Am I missing something? That seems like an awfully small dataset to train to high accuracy, while providing nearly 20,000 test examples. </p>",
      "rawMarkdown": "Does the training data seem small to anyone else? The Train folder has less than 1,000 images, while the test has almost over 18,000. Am I missing something? That seems like an awfully small dataset to train to high accuracy, while providing nearly 20,000 test examples. "
    },
    {
      "id": 173950,
      "postDate": "2017-04-09T15:20:35.283Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 172131,
      "author_name": "toshi_k",
      "author_url": "",
      "post_date": "2017-04-02T05:10:33.910000",
      "content": "<p>Hi RichardHerbert,</p>\n\n<p>Remember that test dataset include dummy images. Some part of 20,000 test images are used for evaluation.\nIn Data description page, it is written as follows, \n<code>As an anti-cheating measure, Kaggle has added some extra images to the test set which will not be counted towards your score.</code></p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 173950,
      "author_name": "",
      "author_url": "",
      "post_date": "2017-04-09T15:20:35.283000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "172131": "Hi RichardHerbert,\n\nRemember that test dataset include dummy images. Some part of 20,000 test images are used for evaluation.\nIn Data description page, it is written as follows, \n`As an anti-cheating measure, Kaggle has added some extra images to the test set which will not be counted towards your score.`",
    "172096": "Does the training data seem small to anyone else? The Train folder has less than 1,000 images, while the test has almost over 18,000. Am I missing something? That seems like an awfully small dataset to train to high accuracy, while providing nearly 20,000 test examples. ",
    "173950": ""
  }
}