{
  "id": 128986,
  "title": "How to use all of the available data in one epoch?",
  "url": "/competitions/bengaliai-cv19/discussion/128986",
  "author_name": "",
  "post_date": "2020-02-04T18:22:29.234543400Z",
  "votes": 1,
  "comment_count": 2,
  "views": 0,
  "content": "<p>I've seen implementations where people train multiple epochs on one of the parquet files, then they continue training on the next file and so on. I think this could really spoil the training as what you're effectively doing is training a model on a smaller amount of data than available, and then doing further training with a different dataset and so on.</p>\n\n<p>I was looking into ways of loading rows from a csv on the fly (in Python using <code>yield</code> to create a generator), and I have two questions, the first being most important\n1) I'm guessing this will be the bottleneck in the speed. Although one epoch takes a while for training, probably it takes longer to load all the data.\n2) If I'm wrong with 1, then is there a way to load parquet data on the fly?</p>\n\n<p>Bonus) Stepping outside the box which I've constricted this question into, what are some approaches you've tried to mitigate the original problem I proposed?</p>",
  "messages": [
    {
      "id": "736947",
      "postDate": "02/04/2020 18:22:29",
      "content": "<p>I've seen implementations where people train multiple epochs on one of the parquet files, then they continue training on the next file and so on. I think this could really spoil the training as what you're effectively doing is training a model on a smaller amount of data than available, and then doing further training with a different dataset and so on.</p>\n\n<p>I was looking into ways of loading rows from a csv on the fly (in Python using <code>yield</code> to create a generator), and I have two questions, the first being most important\n1) I'm guessing this will be the bottleneck in the speed. Although one epoch takes a while for training, probably it takes longer to load all the data.\n2) If I'm wrong with 1, then is there a way to load parquet data on the fly?</p>\n\n<p>Bonus) Stepping outside the box which I've constricted this question into, what are some approaches you've tried to mitigate the original problem I proposed?</p>",
      "rawMarkdown": "I've seen implementations where people train multiple epochs on one of the parquet files, then they continue training on the next file and so on. I think this could really spoil the training as what you're effectively doing is training a model on a smaller amount of data than available, and then doing further training with a different dataset and so on.\n\nI was looking into ways of loading rows from a csv on the fly (in Python using `yield` to create a generator), and I have two questions, the first being most important\n1) I'm guessing this will be the bottleneck in the speed. Although one epoch takes a while for training, probably it takes longer to load all the data.\n2) If I'm wrong with 1, then is there a way to load parquet data on the fly?\n\nBonus) Stepping outside the box which I've constricted this question into, what are some approaches you've tried to mitigate the original problem I proposed?",
      "votes": null
    },
    {
      "id": "737077",
      "postDate": "02/04/2020 22:12:48",
      "content": "<p>A lot of people have preprocessed the data and saved in a format where it can all be loaded into memory or split it into individual image files for a smoother flow of data. There are several different approaches if you search in the current notebooks or datasets.</p>",
      "rawMarkdown": "A lot of people have preprocessed the data and saved in a format where it can all be loaded into memory or split it into individual image files for a smoother flow of data. There are several different approaches if you search in the current notebooks or datasets.",
      "votes": null
    },
    {
      "id": "737395",
      "postDate": "02/05/2020 09:46:22",
      "content": "<p>Valid query. It will give better results to load all the data at once and train the model. But this will pose memory issues maybe. Could TensorFlow datasets prefetch feature be used here to optimize this. Not sure.</p>",
      "rawMarkdown": "Valid query. It will give better results to load all the data at once and train the model. But this will pose memory issues maybe. Could TensorFlow datasets prefetch feature be used here to optimize this. Not sure.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 737077,
      "author_name": "nathanpw",
      "author_url": "",
      "post_date": "02/04/2020 22:12:48",
      "content": "<p>A lot of people have preprocessed the data and saved in a format where it can all be loaded into memory or split it into individual image files for a smoother flow of data. There are several different approaches if you search in the current notebooks or datasets.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 737395,
      "author_name": "anirbank",
      "author_url": "",
      "post_date": "02/05/2020 09:46:22",
      "content": "<p>Valid query. It will give better results to load all the data at once and train the model. But this will pose memory issues maybe. Could TensorFlow datasets prefetch feature be used here to optimize this. Not sure.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "736947": "I've seen implementations where people train multiple epochs on one of the parquet files, then they continue training on the next file and so on. I think this could really spoil the training as what you're effectively doing is training a model on a smaller amount of data than available, and then doing further training with a different dataset and so on.\n\nI was looking into ways of loading rows from a csv on the fly (in Python using `yield` to create a generator), and I have two questions, the first being most important\n1) I'm guessing this will be the bottleneck in the speed. Although one epoch takes a while for training, probably it takes longer to load all the data.\n2) If I'm wrong with 1, then is there a way to load parquet data on the fly?\n\nBonus) Stepping outside the box which I've constricted this question into, what are some approaches you've tried to mitigate the original problem I proposed?",
    "737077": "A lot of people have preprocessed the data and saved in a format where it can all be loaded into memory or split it into individual image files for a smoother flow of data. There are several different approaches if you search in the current notebooks or datasets.",
    "737395": "Valid query. It will give better results to load all the data at once and train the model. But this will pose memory issues maybe. Could TensorFlow datasets prefetch feature be used here to optimize this. Not sure."
  },
  "source": "meta"
}