{
  "id": 40428,
  "title": "Create Train,Validate and Test Set",
  "url": "/competitions/cdiscount-image-classification-challenge/discussion/40428",
  "author_name": "",
  "post_date": "2017-10-02T21:03:19.220678200Z",
  "votes": 2,
  "comment_count": 3,
  "views": 0,
  "content": "<p>Hi I have a question regarding split data for validation. Obviously we cannot load all the data into memory and split on run time. Can someone share how do you create the train validate and test dataset from the orginal training set and do validation? Thanks.</p>",
  "messages": [
    {
      "id": "226624",
      "postDate": "10/02/2017 21:03:19",
      "content": "<p>Hi I have a question regarding split data for validation. Obviously we cannot load all the data into memory and split on run time. Can someone share how do you create the train validate and test dataset from the orginal training set and do validation? Thanks.</p>",
      "rawMarkdown": "Hi I have a question regarding split data for validation. Obviously we cannot load all the data into memory and split on run time. Can someone share how do you create the train validate and test dataset from the orginal training set and do validation? Thanks.",
      "votes": null
    },
    {
      "id": "226763",
      "postDate": "10/03/2017 02:23:13",
      "content": "<ol>\n<li><p>ZFTurbo's kernel <a href=\"https://www.kaggle.com/zfturbo/simple-squeeze-baseline\">https://www.kaggle.com/zfturbo/simple-squeeze-baseline</a> .But it seems that the classes of train data may not be 5270.</p></li>\n<li><p>Analog's kernel <a href=\"https://www.kaggle.com/humananalog/keras-generator-for-reading-directly-from-bson\">https://www.kaggle.com/humananalog/keras-generator-for-reading-directly-from-bson</a> . It confirms the train data and the val data have 5270 classes, but if you don't use keras(like pytorch), the process  of reading data will be slower.</p></li>\n<li><p>Transform <code>bson</code> to images. Of course, it needs more space of your computer. But in  my way, it performs better.</p></li>\n</ol>\n\n<p>By the way , it seems that not so many people use pytorch.</p>\n\n<p>Maybe those can help you, and I am also learning from that.</p>\n\n<p>Hoping for better ways!</p>\n\n<p>Good Luck!</p>",
      "rawMarkdown": "1. ZFTurbo's kernel https://www.kaggle.com/zfturbo/simple-squeeze-baseline .But it seems that the classes of train data may not be 5270.\n\n2. Analog's kernel https://www.kaggle.com/humananalog/keras-generator-for-reading-directly-from-bson . It confirms the train data and the val data have 5270 classes, but if you don't use keras(like pytorch), the process  of reading data will be slower.\n\n3. Transform ```bson``` to images. Of course, it needs more space of your computer. But in  my way, it performs better.\n\nBy the way , it seems that not so many people use pytorch.\n\nMaybe those can help you, and I am also learning from that.\n\nHoping for better ways!\n\nGood Luck!",
      "votes": null
    },
    {
      "id": "227023",
      "postDate": "10/03/2017 15:17:19",
      "content": "<p>Thanks for sharing. For the 3rd method, do you split training data into train and validation on your desk when you read bson?</p>",
      "rawMarkdown": "Thanks for sharing. For the 3rd method, do you split training data into train and validation on your desk when you read bson?",
      "votes": null
    },
    {
      "id": "227402",
      "postDate": "10/04/2017 08:54:37",
      "content": "<p>Yes, I did it. And I confirms both training data and validation data have 5270 classes.</p>",
      "rawMarkdown": "Yes, I did it. And I confirms both training data and validation data have 5270 classes.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 226763,
      "author_name": "zhangsongwei",
      "author_url": "",
      "post_date": "10/03/2017 02:23:13",
      "content": "<ol>\n<li><p>ZFTurbo's kernel <a href=\"https://www.kaggle.com/zfturbo/simple-squeeze-baseline\">https://www.kaggle.com/zfturbo/simple-squeeze-baseline</a> .But it seems that the classes of train data may not be 5270.</p></li>\n<li><p>Analog's kernel <a href=\"https://www.kaggle.com/humananalog/keras-generator-for-reading-directly-from-bson\">https://www.kaggle.com/humananalog/keras-generator-for-reading-directly-from-bson</a> . It confirms the train data and the val data have 5270 classes, but if you don't use keras(like pytorch), the process  of reading data will be slower.</p></li>\n<li><p>Transform <code>bson</code> to images. Of course, it needs more space of your computer. But in  my way, it performs better.</p></li>\n</ol>\n\n<p>By the way , it seems that not so many people use pytorch.</p>\n\n<p>Maybe those can help you, and I am also learning from that.</p>\n\n<p>Hoping for better ways!</p>\n\n<p>Good Luck!</p>",
      "votes": null,
      "replies": [
        {
          "id": 227023,
          "author_name": "wenbo5565",
          "author_url": "",
          "post_date": "10/03/2017 15:17:19",
          "content": "<p>Thanks for sharing. For the 3rd method, do you split training data into train and validation on your desk when you read bson?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 227402,
          "author_name": "zhangsongwei",
          "author_url": "",
          "post_date": "10/04/2017 08:54:37",
          "content": "<p>Yes, I did it. And I confirms both training data and validation data have 5270 classes.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "226624": "Hi I have a question regarding split data for validation. Obviously we cannot load all the data into memory and split on run time. Can someone share how do you create the train validate and test dataset from the orginal training set and do validation? Thanks.",
    "226763": "1. ZFTurbo's kernel https://www.kaggle.com/zfturbo/simple-squeeze-baseline .But it seems that the classes of train data may not be 5270.\n\n2. Analog's kernel https://www.kaggle.com/humananalog/keras-generator-for-reading-directly-from-bson . It confirms the train data and the val data have 5270 classes, but if you don't use keras(like pytorch), the process  of reading data will be slower.\n\n3. Transform ```bson``` to images. Of course, it needs more space of your computer. But in  my way, it performs better.\n\nBy the way , it seems that not so many people use pytorch.\n\nMaybe those can help you, and I am also learning from that.\n\nHoping for better ways!\n\nGood Luck!",
    "227023": "Thanks for sharing. For the 3rd method, do you split training data into train and validation on your desk when you read bson?",
    "227402": "Yes, I did it. And I confirms both training data and validation data have 5270 classes."
  },
  "source": "meta"
}