{
  "id": 41027,
  "title": "Retraining a Network with Random Access?",
  "url": "/competitions/cdiscount-image-classification-challenge/discussion/41027",
  "author_name": "",
  "post_date": "2017-10-11T18:13:38.315447900Z",
  "votes": null,
  "comment_count": 2,
  "views": 0,
  "content": "<p>What's the best way to train a network if we can't fit the entire training set into memory?</p>\n\n<p>Split the training set into chunks, and retrain the network on each several times? Or train on each chunk for several epochs then move onto another one? Or is there a better way?</p>",
  "messages": [
    {
      "id": "230325",
      "postDate": "10/11/2017 18:13:38",
      "content": "<p>What's the best way to train a network if we can't fit the entire training set into memory?</p>\n\n<p>Split the training set into chunks, and retrain the network on each several times? Or train on each chunk for several epochs then move onto another one? Or is there a better way?</p>",
      "rawMarkdown": "What's the best way to train a network if we can't fit the entire training set into memory?\n\nSplit the training set into chunks, and retrain the network on each several times? Or train on each chunk for several epochs then move onto another one? Or is there a better way?",
      "votes": null
    },
    {
      "id": "230328",
      "postDate": "10/11/2017 18:22:22",
      "content": "<p>I assume you're using a GPU for the actual training. You can never fit the entire training set on GPU memory anyway, only small batches of between 32 and 256 images. So there isn't really a need to fit the entire training set on the CPU either. </p>\n\n<p>You can load images from disk to fill up a batch and send this batch to the GPU. As long as creating the batches is faster than the GPU pass, this is fine. (Note that loading images from a hard disk may actually be too slow, which is why it's recommended to use an SSD for storing your training data.)</p>",
      "rawMarkdown": "I assume you're using a GPU for the actual training. You can never fit the entire training set on GPU memory anyway, only small batches of between 32 and 256 images. So there isn't really a need to fit the entire training set on the CPU either. \n\nYou can load images from disk to fill up a batch and send this batch to the GPU. As long as creating the batches is faster than the GPU pass, this is fine. (Note that loading images from a hard disk may actually be too slow, which is why it's recommended to use an SSD for storing your training data.)",
      "votes": null
    },
    {
      "id": "232370",
      "postDate": "10/17/2017 12:09:33",
      "content": "<p>Take a look at <a href=\"https://www.kaggle.com/humananalog/keras-generator-for-reading-directly-from-bson\">this</a> and <a href=\"https://www.kaggle.com/aloisiodn/fast-thread-safe-keras-generator-from-bin-files\">this</a>. The first one indexes the image file and feed the images in batches with random access directly from original BSON files. The second one, uses the same indexes of the first but create separate binary files for train and validation. This approach eliminates the need of random access, thus improving access time, with the drawback of using more disk space. </p>",
      "rawMarkdown": "Take a look at [this][1] and [this][2]. The first one indexes the image file and feed the images in batches with random access directly from original BSON files. The second one, uses the same indexes of the first but create separate binary files for train and validation. This approach eliminates the need of random access, thus improving access time, with the drawback of using more disk space. \n\n\n  [1]: https://www.kaggle.com/humananalog/keras-generator-for-reading-directly-from-bson\n  [2]: https://www.kaggle.com/aloisiodn/fast-thread-safe-keras-generator-from-bin-files",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 230328,
      "author_name": "humananalog",
      "author_url": "",
      "post_date": "10/11/2017 18:22:22",
      "content": "<p>I assume you're using a GPU for the actual training. You can never fit the entire training set on GPU memory anyway, only small batches of between 32 and 256 images. So there isn't really a need to fit the entire training set on the CPU either. </p>\n\n<p>You can load images from disk to fill up a batch and send this batch to the GPU. As long as creating the batches is faster than the GPU pass, this is fine. (Note that loading images from a hard disk may actually be too slow, which is why it's recommended to use an SSD for storing your training data.)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 232370,
      "author_name": "aloisiodn",
      "author_url": "",
      "post_date": "10/17/2017 12:09:33",
      "content": "<p>Take a look at <a href=\"https://www.kaggle.com/humananalog/keras-generator-for-reading-directly-from-bson\">this</a> and <a href=\"https://www.kaggle.com/aloisiodn/fast-thread-safe-keras-generator-from-bin-files\">this</a>. The first one indexes the image file and feed the images in batches with random access directly from original BSON files. The second one, uses the same indexes of the first but create separate binary files for train and validation. This approach eliminates the need of random access, thus improving access time, with the drawback of using more disk space. </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "230325": "What's the best way to train a network if we can't fit the entire training set into memory?\n\nSplit the training set into chunks, and retrain the network on each several times? Or train on each chunk for several epochs then move onto another one? Or is there a better way?",
    "230328": "I assume you're using a GPU for the actual training. You can never fit the entire training set on GPU memory anyway, only small batches of between 32 and 256 images. So there isn't really a need to fit the entire training set on the CPU either. \n\nYou can load images from disk to fill up a batch and send this batch to the GPU. As long as creating the batches is faster than the GPU pass, this is fine. (Note that loading images from a hard disk may actually be too slow, which is why it's recommended to use an SSD for storing your training data.)",
    "232370": "Take a look at [this][1] and [this][2]. The first one indexes the image file and feed the images in batches with random access directly from original BSON files. The second one, uses the same indexes of the first but create separate binary files for train and validation. This approach eliminates the need of random access, thus improving access time, with the drawback of using more disk space. \n\n\n  [1]: https://www.kaggle.com/humananalog/keras-generator-for-reading-directly-from-bson\n  [2]: https://www.kaggle.com/aloisiodn/fast-thread-safe-keras-generator-from-bin-files"
  },
  "source": "meta"
}