{
  "id": 20664,
  "title": "Data can't fit in memory",
  "url": "/competitions/state-farm-distracted-driver-detection/discussion/20664",
  "author_name": "",
  "post_date": "2016-05-03T11:34:49.737Z",
  "votes": null,
  "comment_count": 12,
  "views": 3954,
  "content": "<p>Hi,</p>\n\n<p>I have been using keras for this competition.my submission scored 1.13 on leader board. Right now, am trying to increase the dimension of the image(from 32*32 to 224*224) .but i can't fit this to memory.Please suggest some good idea to have this loaded in RAM (at a stretch or as a generator).</p>\n\n<p>My Laptop RAM - 8GB</p>\n\n<p>NJ/-</p>",
  "messages": [
    {
      "id": "118349",
      "postDate": "05/03/2016 11:34:49",
      "content": "<p>Hi,</p>\n\n<p>I have been using keras for this competition.my submission scored 1.13 on leader board. Right now, am trying to increase the dimension of the image(from 32*32 to 224*224) .but i can't fit this to memory.Please suggest some good idea to have this loaded in RAM (at a stretch or as a generator).</p>\n\n<p>My Laptop RAM - 8GB</p>\n\n<p>NJ/-</p>",
      "rawMarkdown": "Hi,\r\n\r\nI have been using keras for this competition.my submission scored 1.13 on leader board. Right now, am trying to increase the dimension of the image(from 32*32 to 224*224) .but i can't fit this to memory.Please suggest some good idea to have this loaded in RAM (at a stretch or as a generator).\r\n\r\n\r\n\r\nMy Laptop RAM - 8GB\r\n\r\n\r\nNJ/-",
      "votes": null
    },
    {
      "id": "118692",
      "postDate": "05/04/2016 19:16:53",
      "content": "<p>You can use AWS but you have to pay to use the service.</p>",
      "rawMarkdown": "You can use AWS but you have to pay to use the service.",
      "votes": null
    },
    {
      "id": "118731",
      "postDate": "05/05/2016 01:16:29",
      "content": "<p>Maybe you can run a manual train.\nAnd every time you train a batch, then you load the batch data to ram.</p>",
      "rawMarkdown": "Maybe you can run a manual train.\r\nAnd every time you train a batch, then you load the batch data to ram.",
      "votes": null
    },
    {
      "id": "118754",
      "postDate": "05/05/2016 06:22:25",
      "content": "<p>Check what data type your array is using.  Maybe you can more efficiently store the information without loss of integrity.</p>",
      "rawMarkdown": "Check what data type your array is using.  Maybe you can more efficiently store the information without loss of integrity.",
      "votes": null
    },
    {
      "id": "118783",
      "postDate": "05/05/2016 09:29:28",
      "content": "<p>You don't have to load all the dataset to memory, you could load only the batch you are going to feed through the network. </p>",
      "rawMarkdown": "You don't have to load all the dataset to memory, you could load only the batch you are going to feed through the network.",
      "votes": null
    },
    {
      "id": "119096",
      "postDate": "05/07/2016 06:55:34",
      "content": "<p>[quote=NJ/-;118349]</p>\n\n<p>Hi,</p>\n\n<p>I have been using keras for this competition.my submission scored 1.13 on leader board. Right now, am trying to increase the dimension of the image(from 32*32 to 224*224) .but i can't fit this to memory.Please suggest some good idea to have this loaded in RAM (at a stretch or as a generator).</p>\n\n<p>My Laptop RAM - 8GB</p>\n\n<p>NJ/-</p>\n\n<p>[/quote]\nAny success with this? I am facing the same problem. My starter script is that shared by ZFTurbo. What about you?</p>",
      "rawMarkdown": "[quote=NJ/-;118349]\r\n\r\nHi,\r\n\r\nI have been using keras for this competition.my submission scored 1.13 on leader board. Right now, am trying to increase the dimension of the image(from 32*32 to 224*224) .but i can't fit this to memory.Please suggest some good idea to have this loaded in RAM (at a stretch or as a generator).\r\n\r\n\r\n\r\nMy Laptop RAM - 8GB\r\n\r\n\r\nNJ/-\r\n\r\n[/quote]\r\nAny success with this? I am facing the same problem. My starter script is that shared by ZFTurbo. What about you?",
      "votes": null
    },
    {
      "id": "119137",
      "postDate": "05/07/2016 14:43:26",
      "content": "<p>@Abhijay</p>\n\n<p>I'm Trying with generator now. </p>",
      "rawMarkdown": "Abhijay\r\n\r\nI'm Trying with generator now.",
      "votes": null
    },
    {
      "id": "119147",
      "postDate": "05/07/2016 16:37:01",
      "content": "<p>You can use memory mapped files to load and work the training data. It's a bit like a swap file; memory mapped files create a window and a virtual address space larger than RAM. When data is accessed that's not loaded into the window, the kernel swaps in the required pages from the file so you can work on that piece of data. So with this method, random access can be a bit slow, sequential access should be slightly better. </p>\n\n<p>With this method, I can access a 48gb file (test data) using only 24gb of RAM without running out of memory, because the kernel unloads and reloads data for me automatically.</p>\n\n<p>Here's how to do this in python:</p>\n\n<p>As I'm using numpy to represent my arrays/matrix, I create a file of the correct size first. Notice how it's opened write/append:</p>\n\n<pre><code>train_data = np.memmap(ROOT_DIR+&quot;train.driver&quot;, dtype='float32', mode='w+', shape=(TRAIN_ROWS, IMG_SIZE))\n</code></pre>\n\n<p>Then you can assign items row by row or any way you want. Eventually the numpy array is flushed and written to disk:</p>\n\n<pre><code>train_data[tidx,:] = &lt;my_numpy_array_structure_also_of_float32&gt;\n</code></pre>\n\n<p>Flush and close the file at the end:</p>\n\n<pre><code>train_data.flush()\ntrain_data.close()\n</code></pre>\n\n<p>Open it when you're trying to access data, again through a memmap, but this time in read mode:</p>\n\n<pre><code>train = np.memmap(trainfilename, dtype='float32', mode='r', shape=(TRAIN_ROWS, IMG_SIZE))\n</code></pre>\n\n<p>Create an index to access the array and shuffle the indices to get random data:</p>\n\n<pre><code>train_idxs = [i for i in range(train.shape[0])]\nshuffle(train_idxs)\n</code></pre>\n\n<p>It works with ZFTurbo's code, but you can write your own method to access train_data in a batched way with tensorflow:</p>\n\n<pre><code>def next_batch(start,train,labels,batch_size=250):\n    newstart = start+batch_size\n    if newstart &gt; train.shape[0]:\n        newstart = 0\n    idxs = train_idxs[start:start+batch_size]\n    return train[idxs,:], labels[idxs,:], newstart\n</code></pre>\n\n<p>Obviously with random access and a large matrix there will be a delay accessing memory because the kernel needs to move the window across your data. The larger you specify your window, the quicker this can be done.</p>",
      "rawMarkdown": "You can use memory mapped files to load and work the training data. It's a bit like a swap file; memory mapped files create a window and a virtual address space larger than RAM. When data is accessed that's not loaded into the window, the kernel swaps in the required pages from the file so you can work on that piece of data. So with this method, random access can be a bit slow, sequential access should be slightly better. \r\n\r\nWith this method, I can access a 48gb file (test data) using only 24gb of RAM without running out of memory, because the kernel unloads and reloads data for me automatically.\r\n\r\nHere's how to do this in python:\r\n\r\nAs I'm using numpy to represent my arrays/matrix, I create a file of the correct size first. Notice how it's opened write/append:\r\n\r\n    train_data = np.memmap(ROOT_DIR+\"train.driver\", dtype='float32', mode='w+', shape=(TRAIN_ROWS, IMG_SIZE))\r\n\r\nThen you can assign items row by row or any way you want. Eventually the numpy array is flushed and written to disk:\r\n\r\n    train_data[tidx,:] = <my_numpy_array_structure_also_of_float32>\r\n\r\nFlush and close the file at the end:\r\n\r\n    train_data.flush()\r\n    train_data.close()\r\n\r\nOpen it when you're trying to access data, again through a memmap, but this time in read mode:\r\n\r\n    train = np.memmap(trainfilename, dtype='float32', mode='r', shape=(TRAIN_ROWS, IMG_SIZE))\r\n\r\nCreate an index to access the array and shuffle the indices to get random data:\r\n\r\n    train_idxs = [i for i in range(train.shape[0])]\r\n    shuffle(train_idxs)\r\n\r\nIt works with ZFTurbo's code, but you can write your own method to access train_data in a batched way with tensorflow:\r\n\r\n    def next_batch(start,train,labels,batch_size=250):\r\n        newstart = start+batch_size\r\n        if newstart > train.shape[0]:\r\n            newstart = 0\r\n        idxs = train_idxs[start:start+batch_size]\r\n        return train[idxs,:], labels[idxs,:], newstart\r\n\r\nObviously with random access and a large matrix there will be a delay accessing memory because the kernel needs to move the window across your data. The larger you specify your window, the quicker this can be done.",
      "votes": null
    },
    {
      "id": "119195",
      "postDate": "05/07/2016 23:32:50",
      "content": "<p>I've been preparing a keras tutorial for the MNIST dataset which is aimed at folks attempting the <a href=\"https://www.kaggle.com/c/painter-by-numbers\">Painter By Numbers</a> competition - but perhaps those of you working on the Distracted Driver competition might also find it helpful. The network architecture is a little different from this competition because Painter By Numbers needs to examine two images at the same time using a siamese network. However, the process of using a generator to load chunks of the training set is the same.</p>\n\n<p>Code is at <a>https://github.com/small-yellow-duck/kaggle_art/blob/master/mnist_siamese_generator2.py</a>: </p>",
      "rawMarkdown": "I've been preparing a keras tutorial for the MNIST dataset which is aimed at folks attempting the [Painter By Numbers][1] competition - but perhaps those of you working on the Distracted Driver competition might also find it helpful. The network architecture is a little different from this competition because Painter By Numbers needs to examine two images at the same time using a siamese network. However, the process of using a generator to load chunks of the training set is the same.\r\n\r\nCode is at [https://github.com/small-yellow-duck/kaggle_art/blob/master/mnist_siamese_generator2.py][2]: \r\n\r\n\r\n  [1]: https://www.kaggle.com/c/painter-by-numbers\r\n  [2]: http://%20https://github.com/small-yellow-duck/kaggle_art/blob/master/mnist_siamese_generator2.py",
      "votes": null
    },
    {
      "id": "121895",
      "postDate": "05/30/2016 19:21:02",
      "content": "<p>@Gerard Toonstra  would you be willing to post more of your code?  I can't get it running.</p>",
      "rawMarkdown": "Gerard Toonstra  would you be willing to post more of your code?  I can't get it running.",
      "votes": null
    },
    {
      "id": "121900",
      "postDate": "05/30/2016 19:43:10",
      "content": "<p>Evan,</p>\n\n<p>Start with a toy example. Create an array of increasing numbers, then assign that to the memory mapped file. Open the memory mapped file with a hex editor and see if it's all there. Then attempt to read lines from it when you open it from a different script. </p>\n\n<p>Once you have that going, apply it to the problem at hand. </p>\n\n<pre><code>import numpy as np\n\nTOY_ROWS=5\nTOY_SIZE=5\ntrain_data = np.memmap(&quot;toy.example&quot;, dtype='float32', mode='w+', shape=(TOY_ROWS, TOY_SIZE))\nfor x in xrange(TOY_ROWS):\n    train_data[x,:] = np.arange(x*TOY_SIZE, (x+1)*TOY_SIZE)\ntrain_data.flush()\n</code></pre>\n\n<p>I didn't actually run that code, so there may be minor bugs, but it should produce an example file locally. Then read this out as:</p>\n\n<pre><code>import numpy as np\n\ntrain = np.memmap(&quot;toy.example&quot;, dtype='float32', mode='r', shape=(TOY_ROWS, TOY_SIZE))\nprint(train.shape)\nfor x in xrange(TOY_ROWS):\n    print(train[x,:])\n</code></pre>",
      "rawMarkdown": "Evan,\r\n\r\nStart with a toy example. Create an array of increasing numbers, then assign that to the memory mapped file. Open the memory mapped file with a hex editor and see if it's all there. Then attempt to read lines from it when you open it from a different script. \r\n\r\nOnce you have that going, apply it to the problem at hand. \r\n\r\n    import numpy as np\r\n\r\n    TOY_ROWS=5\r\n    TOY_SIZE=5\r\n    train_data = np.memmap(\"toy.example\", dtype='float32', mode='w+', shape=(TOY_ROWS, TOY_SIZE))\r\n    for x in xrange(TOY_ROWS):\r\n        train_data[x,:] = np.arange(x*TOY_SIZE, (x+1)*TOY_SIZE)\r\n    train_data.flush()\r\n\r\nI didn't actually run that code, so there may be minor bugs, but it should produce an example file locally. Then read this out as:\r\n\r\n    import numpy as np\r\n    \r\n    train = np.memmap(\"toy.example\", dtype='float32', mode='r', shape=(TOY_ROWS, TOY_SIZE))\r\n    print(train.shape)\r\n    for x in xrange(TOY_ROWS):\r\n        print(train[x,:])",
      "votes": null
    },
    {
      "id": "121929",
      "postDate": "05/31/2016 02:53:32",
      "content": "<p>thanks Gerard.  I'll see if I can get it to work.</p>",
      "rawMarkdown": "thanks Gerard.  I'll see if I can get it to work.",
      "votes": null
    },
    {
      "id": "121941",
      "postDate": "05/31/2016 05:40:02",
      "content": "<p>Another solution which I find much more convenient than memmaps is hdf5 with h5py.</p>\n\n<p>Keras can run directly on h5py files, eliminating the need to manually specify the training on each batch.</p>\n\n<pre><code>f = h5py.File('cache.hdf5', 'w')\n\ndataset = f.create_dataset('cacheName', (nImgs, nChannels, imgHeight, imgWidth), dtype='float32')\n\nfor i, img in enumerate(images):\n    # fill the thing\n    dataset[i] = img\n\n#read it back\ndataset = h5py.File('cache.hdf5', 'r')\ndataset = dataset.get('cacheName')\n</code></pre>",
      "rawMarkdown": "Another solution which I find much more convenient than memmaps is hdf5 with h5py.\r\n\r\nKeras can run directly on h5py files, eliminating the need to manually specify the training on each batch.\r\n\r\n    f = h5py.File('cache.hdf5', 'w')\r\n    \r\n    dataset = f.create_dataset('cacheName', (nImgs, nChannels, imgHeight, imgWidth), dtype='float32')\r\n    \r\n    for i, img in enumerate(images):\r\n        # fill the thing\r\n        dataset[i] = img\r\n    \r\n    #read it back\r\n    dataset = h5py.File('cache.hdf5', 'r')\r\n    dataset = dataset.get('cacheName')",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 118692,
      "author_name": "notaapple",
      "author_url": "",
      "post_date": "05/04/2016 19:16:53",
      "content": "<p>You can use AWS but you have to pay to use the service.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 118731,
      "author_name": "yelrose",
      "author_url": "",
      "post_date": "05/05/2016 01:16:29",
      "content": "<p>Maybe you can run a manual train.\nAnd every time you train a batch, then you load the batch data to ram.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 118754,
      "author_name": "austin6",
      "author_url": "",
      "post_date": "05/05/2016 06:22:25",
      "content": "<p>Check what data type your array is using.  Maybe you can more efficiently store the information without loss of integrity.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 118783,
      "author_name": "galloguille",
      "author_url": "",
      "post_date": "05/05/2016 09:29:28",
      "content": "<p>You don't have to load all the dataset to memory, you could load only the batch you are going to feed through the network. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 119096,
      "author_name": "abhijayvuyyuru",
      "author_url": "",
      "post_date": "05/07/2016 06:55:34",
      "content": "<p>[quote=NJ/-;118349]</p>\n\n<p>Hi,</p>\n\n<p>I have been using keras for this competition.my submission scored 1.13 on leader board. Right now, am trying to increase the dimension of the image(from 32*32 to 224*224) .but i can't fit this to memory.Please suggest some good idea to have this loaded in RAM (at a stretch or as a generator).</p>\n\n<p>My Laptop RAM - 8GB</p>\n\n<p>NJ/-</p>\n\n<p>[/quote]\nAny success with this? I am facing the same problem. My starter script is that shared by ZFTurbo. What about you?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 119137,
      "author_name": "nithinjames",
      "author_url": "",
      "post_date": "05/07/2016 14:43:26",
      "content": "<p>@Abhijay</p>\n\n<p>I'm Trying with generator now. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 119147,
      "author_name": "remap1",
      "author_url": "",
      "post_date": "05/07/2016 16:37:01",
      "content": "<p>You can use memory mapped files to load and work the training data. It's a bit like a swap file; memory mapped files create a window and a virtual address space larger than RAM. When data is accessed that's not loaded into the window, the kernel swaps in the required pages from the file so you can work on that piece of data. So with this method, random access can be a bit slow, sequential access should be slightly better. </p>\n\n<p>With this method, I can access a 48gb file (test data) using only 24gb of RAM without running out of memory, because the kernel unloads and reloads data for me automatically.</p>\n\n<p>Here's how to do this in python:</p>\n\n<p>As I'm using numpy to represent my arrays/matrix, I create a file of the correct size first. Notice how it's opened write/append:</p>\n\n<pre><code>train_data = np.memmap(ROOT_DIR+&quot;train.driver&quot;, dtype='float32', mode='w+', shape=(TRAIN_ROWS, IMG_SIZE))\n</code></pre>\n\n<p>Then you can assign items row by row or any way you want. Eventually the numpy array is flushed and written to disk:</p>\n\n<pre><code>train_data[tidx,:] = &lt;my_numpy_array_structure_also_of_float32&gt;\n</code></pre>\n\n<p>Flush and close the file at the end:</p>\n\n<pre><code>train_data.flush()\ntrain_data.close()\n</code></pre>\n\n<p>Open it when you're trying to access data, again through a memmap, but this time in read mode:</p>\n\n<pre><code>train = np.memmap(trainfilename, dtype='float32', mode='r', shape=(TRAIN_ROWS, IMG_SIZE))\n</code></pre>\n\n<p>Create an index to access the array and shuffle the indices to get random data:</p>\n\n<pre><code>train_idxs = [i for i in range(train.shape[0])]\nshuffle(train_idxs)\n</code></pre>\n\n<p>It works with ZFTurbo's code, but you can write your own method to access train_data in a batched way with tensorflow:</p>\n\n<pre><code>def next_batch(start,train,labels,batch_size=250):\n    newstart = start+batch_size\n    if newstart &gt; train.shape[0]:\n        newstart = 0\n    idxs = train_idxs[start:start+batch_size]\n    return train[idxs,:], labels[idxs,:], newstart\n</code></pre>\n\n<p>Obviously with random access and a large matrix there will be a delay accessing memory because the kernel needs to move the window across your data. The larger you specify your window, the quicker this can be done.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 119195,
      "author_name": "smallyellowduck",
      "author_url": "",
      "post_date": "05/07/2016 23:32:50",
      "content": "<p>I've been preparing a keras tutorial for the MNIST dataset which is aimed at folks attempting the <a href=\"https://www.kaggle.com/c/painter-by-numbers\">Painter By Numbers</a> competition - but perhaps those of you working on the Distracted Driver competition might also find it helpful. The network architecture is a little different from this competition because Painter By Numbers needs to examine two images at the same time using a siamese network. However, the process of using a generator to load chunks of the training set is the same.</p>\n\n<p>Code is at <a>https://github.com/small-yellow-duck/kaggle_art/blob/master/mnist_siamese_generator2.py</a>: </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 121895,
      "author_name": "evanvanness",
      "author_url": "",
      "post_date": "05/30/2016 19:21:02",
      "content": "<p>@Gerard Toonstra  would you be willing to post more of your code?  I can't get it running.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 121900,
      "author_name": "remap1",
      "author_url": "",
      "post_date": "05/30/2016 19:43:10",
      "content": "<p>Evan,</p>\n\n<p>Start with a toy example. Create an array of increasing numbers, then assign that to the memory mapped file. Open the memory mapped file with a hex editor and see if it's all there. Then attempt to read lines from it when you open it from a different script. </p>\n\n<p>Once you have that going, apply it to the problem at hand. </p>\n\n<pre><code>import numpy as np\n\nTOY_ROWS=5\nTOY_SIZE=5\ntrain_data = np.memmap(&quot;toy.example&quot;, dtype='float32', mode='w+', shape=(TOY_ROWS, TOY_SIZE))\nfor x in xrange(TOY_ROWS):\n    train_data[x,:] = np.arange(x*TOY_SIZE, (x+1)*TOY_SIZE)\ntrain_data.flush()\n</code></pre>\n\n<p>I didn't actually run that code, so there may be minor bugs, but it should produce an example file locally. Then read this out as:</p>\n\n<pre><code>import numpy as np\n\ntrain = np.memmap(&quot;toy.example&quot;, dtype='float32', mode='r', shape=(TOY_ROWS, TOY_SIZE))\nprint(train.shape)\nfor x in xrange(TOY_ROWS):\n    print(train[x,:])\n</code></pre>",
      "votes": null,
      "replies": []
    },
    {
      "id": 121929,
      "author_name": "evanvanness",
      "author_url": "",
      "post_date": "05/31/2016 02:53:32",
      "content": "<p>thanks Gerard.  I'll see if I can get it to work.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 121941,
      "author_name": "gpistre",
      "author_url": "",
      "post_date": "05/31/2016 05:40:02",
      "content": "<p>Another solution which I find much more convenient than memmaps is hdf5 with h5py.</p>\n\n<p>Keras can run directly on h5py files, eliminating the need to manually specify the training on each batch.</p>\n\n<pre><code>f = h5py.File('cache.hdf5', 'w')\n\ndataset = f.create_dataset('cacheName', (nImgs, nChannels, imgHeight, imgWidth), dtype='float32')\n\nfor i, img in enumerate(images):\n    # fill the thing\n    dataset[i] = img\n\n#read it back\ndataset = h5py.File('cache.hdf5', 'r')\ndataset = dataset.get('cacheName')\n</code></pre>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "118349": "Hi,\r\n\r\nI have been using keras for this competition.my submission scored 1.13 on leader board. Right now, am trying to increase the dimension of the image(from 32*32 to 224*224) .but i can't fit this to memory.Please suggest some good idea to have this loaded in RAM (at a stretch or as a generator).\r\n\r\n\r\n\r\nMy Laptop RAM - 8GB\r\n\r\n\r\nNJ/-",
    "118692": "You can use AWS but you have to pay to use the service.",
    "118731": "Maybe you can run a manual train.\r\nAnd every time you train a batch, then you load the batch data to ram.",
    "118754": "Check what data type your array is using.  Maybe you can more efficiently store the information without loss of integrity.",
    "118783": "You don't have to load all the dataset to memory, you could load only the batch you are going to feed through the network.",
    "119096": "[quote=NJ/-;118349]\r\n\r\nHi,\r\n\r\nI have been using keras for this competition.my submission scored 1.13 on leader board. Right now, am trying to increase the dimension of the image(from 32*32 to 224*224) .but i can't fit this to memory.Please suggest some good idea to have this loaded in RAM (at a stretch or as a generator).\r\n\r\n\r\n\r\nMy Laptop RAM - 8GB\r\n\r\n\r\nNJ/-\r\n\r\n[/quote]\r\nAny success with this? I am facing the same problem. My starter script is that shared by ZFTurbo. What about you?",
    "119137": "Abhijay\r\n\r\nI'm Trying with generator now.",
    "119147": "You can use memory mapped files to load and work the training data. It's a bit like a swap file; memory mapped files create a window and a virtual address space larger than RAM. When data is accessed that's not loaded into the window, the kernel swaps in the required pages from the file so you can work on that piece of data. So with this method, random access can be a bit slow, sequential access should be slightly better. \r\n\r\nWith this method, I can access a 48gb file (test data) using only 24gb of RAM without running out of memory, because the kernel unloads and reloads data for me automatically.\r\n\r\nHere's how to do this in python:\r\n\r\nAs I'm using numpy to represent my arrays/matrix, I create a file of the correct size first. Notice how it's opened write/append:\r\n\r\n    train_data = np.memmap(ROOT_DIR+\"train.driver\", dtype='float32', mode='w+', shape=(TRAIN_ROWS, IMG_SIZE))\r\n\r\nThen you can assign items row by row or any way you want. Eventually the numpy array is flushed and written to disk:\r\n\r\n    train_data[tidx,:] = <my_numpy_array_structure_also_of_float32>\r\n\r\nFlush and close the file at the end:\r\n\r\n    train_data.flush()\r\n    train_data.close()\r\n\r\nOpen it when you're trying to access data, again through a memmap, but this time in read mode:\r\n\r\n    train = np.memmap(trainfilename, dtype='float32', mode='r', shape=(TRAIN_ROWS, IMG_SIZE))\r\n\r\nCreate an index to access the array and shuffle the indices to get random data:\r\n\r\n    train_idxs = [i for i in range(train.shape[0])]\r\n    shuffle(train_idxs)\r\n\r\nIt works with ZFTurbo's code, but you can write your own method to access train_data in a batched way with tensorflow:\r\n\r\n    def next_batch(start,train,labels,batch_size=250):\r\n        newstart = start+batch_size\r\n        if newstart > train.shape[0]:\r\n            newstart = 0\r\n        idxs = train_idxs[start:start+batch_size]\r\n        return train[idxs,:], labels[idxs,:], newstart\r\n\r\nObviously with random access and a large matrix there will be a delay accessing memory because the kernel needs to move the window across your data. The larger you specify your window, the quicker this can be done.",
    "119195": "I've been preparing a keras tutorial for the MNIST dataset which is aimed at folks attempting the [Painter By Numbers][1] competition - but perhaps those of you working on the Distracted Driver competition might also find it helpful. The network architecture is a little different from this competition because Painter By Numbers needs to examine two images at the same time using a siamese network. However, the process of using a generator to load chunks of the training set is the same.\r\n\r\nCode is at [https://github.com/small-yellow-duck/kaggle_art/blob/master/mnist_siamese_generator2.py][2]: \r\n\r\n\r\n  [1]: https://www.kaggle.com/c/painter-by-numbers\r\n  [2]: http://%20https://github.com/small-yellow-duck/kaggle_art/blob/master/mnist_siamese_generator2.py",
    "121895": "Gerard Toonstra  would you be willing to post more of your code?  I can't get it running.",
    "121900": "Evan,\r\n\r\nStart with a toy example. Create an array of increasing numbers, then assign that to the memory mapped file. Open the memory mapped file with a hex editor and see if it's all there. Then attempt to read lines from it when you open it from a different script. \r\n\r\nOnce you have that going, apply it to the problem at hand. \r\n\r\n    import numpy as np\r\n\r\n    TOY_ROWS=5\r\n    TOY_SIZE=5\r\n    train_data = np.memmap(\"toy.example\", dtype='float32', mode='w+', shape=(TOY_ROWS, TOY_SIZE))\r\n    for x in xrange(TOY_ROWS):\r\n        train_data[x,:] = np.arange(x*TOY_SIZE, (x+1)*TOY_SIZE)\r\n    train_data.flush()\r\n\r\nI didn't actually run that code, so there may be minor bugs, but it should produce an example file locally. Then read this out as:\r\n\r\n    import numpy as np\r\n    \r\n    train = np.memmap(\"toy.example\", dtype='float32', mode='r', shape=(TOY_ROWS, TOY_SIZE))\r\n    print(train.shape)\r\n    for x in xrange(TOY_ROWS):\r\n        print(train[x,:])",
    "121929": "thanks Gerard.  I'll see if I can get it to work.",
    "121941": "Another solution which I find much more convenient than memmaps is hdf5 with h5py.\r\n\r\nKeras can run directly on h5py files, eliminating the need to manually specify the training on each batch.\r\n\r\n    f = h5py.File('cache.hdf5', 'w')\r\n    \r\n    dataset = f.create_dataset('cacheName', (nImgs, nChannels, imgHeight, imgWidth), dtype='float32')\r\n    \r\n    for i, img in enumerate(images):\r\n        # fill the thing\r\n        dataset[i] = img\r\n    \r\n    #read it back\r\n    dataset = h5py.File('cache.hdf5', 'r')\r\n    dataset = dataset.get('cacheName')"
  },
  "source": "meta"
}