{
  "id": 42683,
  "title": "Prediction by a generator (Keras)",
  "url": "/competitions/cdiscount-image-classification-challenge/discussion/42683",
  "author_name": "",
  "post_date": "2017-11-03T00:51:01.484980Z",
  "votes": null,
  "comment_count": 3,
  "views": 0,
  "content": "<p>Is it possible to make predictions for test data by using a generator in Keras?\nThe dataset we're given is sorted by product_id, and for each product, there are 1-4 pictures. So we need to take an average for a product with multiple images. </p>\n\n<p>Because of this and machine memory issue, we need to run a loop for each product. In each loop, we get predictions from our trained model, take an average, and obtain a class label for that product. But, is there any smarter way to utilize a data generator and efficiently make predictions? Now my GPUs are starved during prediction...</p>",
  "messages": [
    {
      "id": "239235",
      "postDate": "11/03/2017 00:51:01",
      "content": "<p>Is it possible to make predictions for test data by using a generator in Keras?\nThe dataset we're given is sorted by product_id, and for each product, there are 1-4 pictures. So we need to take an average for a product with multiple images. </p>\n\n<p>Because of this and machine memory issue, we need to run a loop for each product. In each loop, we get predictions from our trained model, take an average, and obtain a class label for that product. But, is there any smarter way to utilize a data generator and efficiently make predictions? Now my GPUs are starved during prediction...</p>",
      "rawMarkdown": "Is it possible to make predictions for test data by using a generator in Keras?\nThe dataset we're given is sorted by product_id, and for each product, there are 1-4 pictures. So we need to take an average for a product with multiple images. \n\nBecause of this and machine memory issue, we need to run a loop for each product. In each loop, we get predictions from our trained model, take an average, and obtain a class label for that product. But, is there any smarter way to utilize a data generator and efficiently make predictions? Now my GPUs are starved during prediction...",
      "votes": null
    },
    {
      "id": "239252",
      "postDate": "11/03/2017 02:56:40",
      "content": "<p>Did you look at this kernel?</p>\n\n<p><a href=\"https://www.kaggle.com/humananalog/keras-generator-for-reading-directly-from-bson\">https://www.kaggle.com/humananalog/keras-generator-for-reading-directly-from-bson</a></p>",
      "rawMarkdown": "Did you look at this kernel?\n\nhttps://www.kaggle.com/humananalog/keras-generator-for-reading-directly-from-bson",
      "votes": null
    },
    {
      "id": "239262",
      "postDate": "11/03/2017 04:40:03",
      "content": "<p>Yes, but it only is for training. For prediction, he uses a modified generator to generate only 1-4 images for each product. The speed of that approach is bounded by speed of loop and utilization of GPUs goes low. I wanted to know how to use the original generator in his kernel (or whatever generator proposed by anybody) to make predictions.</p>",
      "rawMarkdown": "Yes, but it only is for training. For prediction, he uses a modified generator to generate only 1-4 images for each product. The speed of that approach is bounded by speed of loop and utilization of GPUs goes low. I wanted to know how to use the original generator in his kernel (or whatever generator proposed by anybody) to make predictions.",
      "votes": null
    },
    {
      "id": "239409",
      "postDate": "11/03/2017 11:39:37",
      "content": "<p>If you use the original generator to make predictions, you're going to run out of memory. Version 1 of the kernel shows how to do this. First you make a Pandas DataFrame that contains one row for each test image, then you create a generator, and finally you call <code>predict_generator()</code>. However, this creates an output array that is about 60GB in size, too much to fit into memory. So if you really want to do this, you should make predictions on smaller chunks of the test set.</p>",
      "rawMarkdown": "If you use the original generator to make predictions, you're going to run out of memory. Version 1 of the kernel shows how to do this. First you make a Pandas DataFrame that contains one row for each test image, then you create a generator, and finally you call `predict_generator()`. However, this creates an output array that is about 60GB in size, too much to fit into memory. So if you really want to do this, you should make predictions on smaller chunks of the test set.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 239252,
      "author_name": "sgalib",
      "author_url": "",
      "post_date": "11/03/2017 02:56:40",
      "content": "<p>Did you look at this kernel?</p>\n\n<p><a href=\"https://www.kaggle.com/humananalog/keras-generator-for-reading-directly-from-bson\">https://www.kaggle.com/humananalog/keras-generator-for-reading-directly-from-bson</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 239262,
          "author_name": "llain21",
          "author_url": "",
          "post_date": "11/03/2017 04:40:03",
          "content": "<p>Yes, but it only is for training. For prediction, he uses a modified generator to generate only 1-4 images for each product. The speed of that approach is bounded by speed of loop and utilization of GPUs goes low. I wanted to know how to use the original generator in his kernel (or whatever generator proposed by anybody) to make predictions.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 239409,
          "author_name": "humananalog",
          "author_url": "",
          "post_date": "11/03/2017 11:39:37",
          "content": "<p>If you use the original generator to make predictions, you're going to run out of memory. Version 1 of the kernel shows how to do this. First you make a Pandas DataFrame that contains one row for each test image, then you create a generator, and finally you call <code>predict_generator()</code>. However, this creates an output array that is about 60GB in size, too much to fit into memory. So if you really want to do this, you should make predictions on smaller chunks of the test set.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "239235": "Is it possible to make predictions for test data by using a generator in Keras?\nThe dataset we're given is sorted by product_id, and for each product, there are 1-4 pictures. So we need to take an average for a product with multiple images. \n\nBecause of this and machine memory issue, we need to run a loop for each product. In each loop, we get predictions from our trained model, take an average, and obtain a class label for that product. But, is there any smarter way to utilize a data generator and efficiently make predictions? Now my GPUs are starved during prediction...",
    "239252": "Did you look at this kernel?\n\nhttps://www.kaggle.com/humananalog/keras-generator-for-reading-directly-from-bson",
    "239262": "Yes, but it only is for training. For prediction, he uses a modified generator to generate only 1-4 images for each product. The speed of that approach is bounded by speed of loop and utilization of GPUs goes low. I wanted to know how to use the original generator in his kernel (or whatever generator proposed by anybody) to make predictions.",
    "239409": "If you use the original generator to make predictions, you're going to run out of memory. Version 1 of the kernel shows how to do this. First you make a Pandas DataFrame that contains one row for each test image, then you create a generator, and finally you call `predict_generator()`. However, this creates an output array that is about 60GB in size, too much to fit into memory. So if you really want to do this, you should make predictions on smaller chunks of the test set."
  },
  "source": "meta"
}