{
  "id": 34066,
  "title": "Tile database",
  "url": "/competitions/noaa-fisheries-steller-sea-lion-population-count/discussion/34066",
  "author_name": "",
  "post_date": "2017-06-02T23:24:39.840939100Z",
  "votes": null,
  "comment_count": 4,
  "views": 0,
  "content": "<p>Does anyone who train model on tile level use some sort of database (leveldb,lmdb,hdf5)?\nLooks like lots of small files on HDD is not very good solution.</p>",
  "messages": [
    {
      "id": "188465",
      "postDate": "06/02/2017 23:24:39",
      "content": "<p>Does anyone who train model on tile level use some sort of database (leveldb,lmdb,hdf5)?\nLooks like lots of small files on HDD is not very good solution.</p>",
      "rawMarkdown": "Does anyone who train model on tile level use some sort of database (leveldb,lmdb,hdf5)?\nLooks like lots of small files on HDD is not very good solution.",
      "votes": null
    },
    {
      "id": "188469",
      "postDate": "06/02/2017 23:54:00",
      "content": "<p>Just a quick test:</p>\n\n<p>Initial image size 10.3 Mb.</p>\n\n<blockquote>\n<pre><code>du -sh DB/*\n39M   DB/db_type_1_image_id_0\n221M  DB/db_type_2_image_id_0\n177M  DB/db_type_3_image_id_0\n</code></pre>\n</blockquote>\n\n<pre><code>import cv2\nimport numpy as np\nimport os\n\nimage_id= 0\ntile_size= 256\ntile_stride= tile_size/2\n\ntrain_image_dir='Train'\ndatabase_dir='DB'\n\nif not os.path.exists(database_dir):\n    os.makedirs(database_dir)\n\n#Just create folder, cut image to tiles and store as jpg images in dir\ndef create_database_1(image_id):\n    img= cv2.imread(train_image_dir+'/'+str(image_id)+'.jpg')\n\n    if not os.path.exists(database_dir+'/db_1_'+str(image_id)):\n        os.makedirs(database_dir+'/db_type_1_image_id_'+str(image_id))\n\n    w= img.shape[1]\n    h= img.shape[0]\n\n    counter=0\n    for y in range(0, h, tile_stride):\n        for x in range(0, w, tile_stride):\n            if(y+tile_size &amp;lt; h and x+tile_size &amp;lt; w):\n                tile_img= img[y:y+tile_size,x:x+tile_size]\n\n                cv2.imwrite(database_dir+'/db_type_1_image_id_'+str(image_id)+'/'+str(counter)+'.jpg',tile_img)\n\n                counter=counter+1\n\n    print 'Numer of tiles:',counter\n\n#Create .npy file with all tiles via numpy\ndef create_database_2(image_id):\n    img= cv2.imread(train_image_dir+'/'+str(image_id)+'.jpg')\n\n    if not os.path.exists(database_dir+'/db_type_2_image_id_'+str(image_id)):\n        os.makedirs(database_dir+'/db_type_2_image_id_'+str(image_id))\n\n    w= img.shape[1]\n    h= img.shape[0]\n\n    arr_list=[]\n    counter=0\n    for y in range(0, h, tile_stride):\n        for x in range(0, w, tile_stride):\n            if(y+tile_size &amp;lt; h and x+tile_size &amp;lt; w):\n                tile_img= img[y:y+tile_size,x:x+tile_size]\n                arr_list.append(tile_img)\n                counter=counter+1\n\n    arr= np.asarray(arr_list)\n    print 'arr.shape', arr.shape\n\n    np.save(database_dir+'/db_type_2_image_id_'+str(image_id)+'/tiles.npy', arr)\n\n    print 'Numer of tiles:',counter\n\n#Create .npz file with all tiles via numpy\ndef create_database_3(image_id):\n    img= cv2.imread(train_image_dir+'/'+str(image_id)+'.jpg')\n\n    if not os.path.exists(database_dir+'/db_type_3_image_id_'+str(image_id)):\n        os.makedirs(database_dir+'/db_type_3_image_id_'+str(image_id))\n\n    w= img.shape[1]\n    h= img.shape[0]\n\n    arr_list=[]\n    counter=0\n    for y in range(0, h, tile_stride):\n        for x in range(0, w, tile_stride):\n            if(y+tile_size &amp;lt; h and x+tile_size &amp;lt; w):\n                tile_img= img[y:y+tile_size,x:x+tile_size]\n                arr_list.append(tile_img)\n                counter=counter+1\n\n    arr= np.asarray(arr_list)\n    print 'arr.shape', arr.shape\n\n    np.savez_compressed(database_dir+'/db_type_3_image_id_'+str(image_id)+'/tiles.npz', arr)\n\n    print 'Numer of tiles:',counter\n\ncreate_database_1(image_id)\ncreate_database_2(image_id)\ncreate_database_3(image_id)\n</code></pre>",
      "rawMarkdown": "Just a quick test:\n\nInitial image size 10.3 Mb.\n\n&gt;     du -sh DB/*\n&gt;     39M\tDB/db_type_1_image_id_0\n&gt;     221M\tDB/db_type_2_image_id_0\n&gt;     177M\tDB/db_type_3_image_id_0\n\n\n    import cv2\n    import numpy as np\n    import os\n    \n    image_id= 0\n    tile_size= 256\n    tile_stride= tile_size/2\n    \n    train_image_dir='Train'\n    database_dir='DB'\n    \n    if not os.path.exists(database_dir):\n        os.makedirs(database_dir)\n        \n    #Just create folder, cut image to tiles and store as jpg images in dir\n    def create_database_1(image_id):\n    \timg= cv2.imread(train_image_dir+'/'+str(image_id)+'.jpg')\n    \t\n    \tif not os.path.exists(database_dir+'/db_1_'+str(image_id)):\n    \t\tos.makedirs(database_dir+'/db_type_1_image_id_'+str(image_id))\n    \t\n    \tw= img.shape[1]\n    \th= img.shape[0]\n    \n    \tcounter=0\n    \tfor y in range(0, h, tile_stride):\n    \t\tfor x in range(0, w, tile_stride):\n    \t\t\tif(y+tile_size &lt; h and x+tile_size &lt; w):\n    \t\t\t\ttile_img= img[y:y+tile_size,x:x+tile_size]\n    \t\t\t\t\n    \t\t\t\tcv2.imwrite(database_dir+'/db_type_1_image_id_'+str(image_id)+'/'+str(counter)+'.jpg',tile_img)\n    \t\t\t\t\n    \t\t\t\tcounter=counter+1\n    \t\n    \tprint 'Numer of tiles:',counter\n    \n    #Create .npy file with all tiles via numpy\n    def create_database_2(image_id):\n    \timg= cv2.imread(train_image_dir+'/'+str(image_id)+'.jpg')\n    \t\n    \tif not os.path.exists(database_dir+'/db_type_2_image_id_'+str(image_id)):\n    \t\tos.makedirs(database_dir+'/db_type_2_image_id_'+str(image_id))\n    \t\n    \tw= img.shape[1]\n    \th= img.shape[0]\n    \t\n    \tarr_list=[]\n    \tcounter=0\n    \tfor y in range(0, h, tile_stride):\n    \t\tfor x in range(0, w, tile_stride):\n    \t\t\tif(y+tile_size &lt; h and x+tile_size &lt; w):\n    \t\t\t\ttile_img= img[y:y+tile_size,x:x+tile_size]\n    \t\t\t\tarr_list.append(tile_img)\n    \t\t\t\tcounter=counter+1\n    \t\n    \tarr= np.asarray(arr_list)\n    \tprint 'arr.shape', arr.shape\n    \t\n    \tnp.save(database_dir+'/db_type_2_image_id_'+str(image_id)+'/tiles.npy', arr)\n    \t\n    \tprint 'Numer of tiles:',counter\n    \n    #Create .npz file with all tiles via numpy\n    def create_database_3(image_id):\n    \timg= cv2.imread(train_image_dir+'/'+str(image_id)+'.jpg')\n    \t\n    \tif not os.path.exists(database_dir+'/db_type_3_image_id_'+str(image_id)):\n    \t\tos.makedirs(database_dir+'/db_type_3_image_id_'+str(image_id))\n    \t\n    \tw= img.shape[1]\n    \th= img.shape[0]\n    \t\n    \tarr_list=[]\n    \tcounter=0\n    \tfor y in range(0, h, tile_stride):\n    \t\tfor x in range(0, w, tile_stride):\n    \t\t\tif(y+tile_size &lt; h and x+tile_size &lt; w):\n    \t\t\t\ttile_img= img[y:y+tile_size,x:x+tile_size]\n    \t\t\t\tarr_list.append(tile_img)\n    \t\t\t\tcounter=counter+1\n    \t\n    \tarr= np.asarray(arr_list)\n    \tprint 'arr.shape', arr.shape\n    \t\n    \tnp.savez_compressed(database_dir+'/db_type_3_image_id_'+str(image_id)+'/tiles.npz', arr)\n    \t\n    \tprint 'Numer of tiles:',counter\n    \t\n    create_database_1(image_id)\n    create_database_2(image_id)\n    create_database_3(image_id)",
      "votes": null
    },
    {
      "id": "188542",
      "postDate": "06/03/2017 06:38:19",
      "content": "<p>I'm loading most images in memory as uint8 numpy arrays (they are stored in this form on disk), and then generating patches on the fly during training. But that barely fits in 64GB RAM :(</p>",
      "rawMarkdown": "I'm loading most images in memory as uint8 numpy arrays (they are stored in this form on disk), and then generating patches on the fly during training. But that barely fits in 64GB RAM :(",
      "votes": null
    },
    {
      "id": "188565",
      "postDate": "06/03/2017 08:13:01",
      "content": "<p>Why do you store images as numpy arrays? for faster loading?\nAlso if you store one .npy file per image then you can just get random batch from this image and not from entire dataset which sounds suboptimal, my idea is to save shuffled tiles from all images to some db and then sequentially read from it (like in Caffe).</p>",
      "rawMarkdown": "Why do you store images as numpy arrays? for faster loading?\nAlso if you store one .npy file per image then you can just get random batch from this image and not from entire dataset which sounds suboptimal, my idea is to save shuffled tiles from all images to some db and then sequentially read from it (like in Caffe).",
      "votes": null
    },
    {
      "id": "188573",
      "postDate": "06/03/2017 08:35:27",
      "content": "<p>Yes, I store them in numpy arrays for faster loading: for example if do some experiment on 100 images, and they are already in disk cache, then loading takes just a couple of seconds. I don't have an SSD, just an HDD, so I optimize for that.</p>\n\n<p>I'm sampling random batches from the entire dataset: all images are in memory, so first I sample a random image, then random location, and form the whole batch this way, so it's completely random.</p>",
      "rawMarkdown": "Yes, I store them in numpy arrays for faster loading: for example if do some experiment on 100 images, and they are already in disk cache, then loading takes just a couple of seconds. I don't have an SSD, just an HDD, so I optimize for that.\n\nI'm sampling random batches from the entire dataset: all images are in memory, so first I sample a random image, then random location, and form the whole batch this way, so it's completely random.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 188469,
      "author_name": "mrgloom",
      "author_url": "",
      "post_date": "06/02/2017 23:54:00",
      "content": "<p>Just a quick test:</p>\n\n<p>Initial image size 10.3 Mb.</p>\n\n<blockquote>\n<pre><code>du -sh DB/*\n39M   DB/db_type_1_image_id_0\n221M  DB/db_type_2_image_id_0\n177M  DB/db_type_3_image_id_0\n</code></pre>\n</blockquote>\n\n<pre><code>import cv2\nimport numpy as np\nimport os\n\nimage_id= 0\ntile_size= 256\ntile_stride= tile_size/2\n\ntrain_image_dir='Train'\ndatabase_dir='DB'\n\nif not os.path.exists(database_dir):\n    os.makedirs(database_dir)\n\n#Just create folder, cut image to tiles and store as jpg images in dir\ndef create_database_1(image_id):\n    img= cv2.imread(train_image_dir+'/'+str(image_id)+'.jpg')\n\n    if not os.path.exists(database_dir+'/db_1_'+str(image_id)):\n        os.makedirs(database_dir+'/db_type_1_image_id_'+str(image_id))\n\n    w= img.shape[1]\n    h= img.shape[0]\n\n    counter=0\n    for y in range(0, h, tile_stride):\n        for x in range(0, w, tile_stride):\n            if(y+tile_size &amp;lt; h and x+tile_size &amp;lt; w):\n                tile_img= img[y:y+tile_size,x:x+tile_size]\n\n                cv2.imwrite(database_dir+'/db_type_1_image_id_'+str(image_id)+'/'+str(counter)+'.jpg',tile_img)\n\n                counter=counter+1\n\n    print 'Numer of tiles:',counter\n\n#Create .npy file with all tiles via numpy\ndef create_database_2(image_id):\n    img= cv2.imread(train_image_dir+'/'+str(image_id)+'.jpg')\n\n    if not os.path.exists(database_dir+'/db_type_2_image_id_'+str(image_id)):\n        os.makedirs(database_dir+'/db_type_2_image_id_'+str(image_id))\n\n    w= img.shape[1]\n    h= img.shape[0]\n\n    arr_list=[]\n    counter=0\n    for y in range(0, h, tile_stride):\n        for x in range(0, w, tile_stride):\n            if(y+tile_size &amp;lt; h and x+tile_size &amp;lt; w):\n                tile_img= img[y:y+tile_size,x:x+tile_size]\n                arr_list.append(tile_img)\n                counter=counter+1\n\n    arr= np.asarray(arr_list)\n    print 'arr.shape', arr.shape\n\n    np.save(database_dir+'/db_type_2_image_id_'+str(image_id)+'/tiles.npy', arr)\n\n    print 'Numer of tiles:',counter\n\n#Create .npz file with all tiles via numpy\ndef create_database_3(image_id):\n    img= cv2.imread(train_image_dir+'/'+str(image_id)+'.jpg')\n\n    if not os.path.exists(database_dir+'/db_type_3_image_id_'+str(image_id)):\n        os.makedirs(database_dir+'/db_type_3_image_id_'+str(image_id))\n\n    w= img.shape[1]\n    h= img.shape[0]\n\n    arr_list=[]\n    counter=0\n    for y in range(0, h, tile_stride):\n        for x in range(0, w, tile_stride):\n            if(y+tile_size &amp;lt; h and x+tile_size &amp;lt; w):\n                tile_img= img[y:y+tile_size,x:x+tile_size]\n                arr_list.append(tile_img)\n                counter=counter+1\n\n    arr= np.asarray(arr_list)\n    print 'arr.shape', arr.shape\n\n    np.savez_compressed(database_dir+'/db_type_3_image_id_'+str(image_id)+'/tiles.npz', arr)\n\n    print 'Numer of tiles:',counter\n\ncreate_database_1(image_id)\ncreate_database_2(image_id)\ncreate_database_3(image_id)\n</code></pre>",
      "votes": null,
      "replies": []
    },
    {
      "id": 188542,
      "author_name": "lopuhin",
      "author_url": "",
      "post_date": "06/03/2017 06:38:19",
      "content": "<p>I'm loading most images in memory as uint8 numpy arrays (they are stored in this form on disk), and then generating patches on the fly during training. But that barely fits in 64GB RAM :(</p>",
      "votes": null,
      "replies": [
        {
          "id": 188565,
          "author_name": "mrgloom",
          "author_url": "",
          "post_date": "06/03/2017 08:13:01",
          "content": "<p>Why do you store images as numpy arrays? for faster loading?\nAlso if you store one .npy file per image then you can just get random batch from this image and not from entire dataset which sounds suboptimal, my idea is to save shuffled tiles from all images to some db and then sequentially read from it (like in Caffe).</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 188573,
          "author_name": "lopuhin",
          "author_url": "",
          "post_date": "06/03/2017 08:35:27",
          "content": "<p>Yes, I store them in numpy arrays for faster loading: for example if do some experiment on 100 images, and they are already in disk cache, then loading takes just a couple of seconds. I don't have an SSD, just an HDD, so I optimize for that.</p>\n\n<p>I'm sampling random batches from the entire dataset: all images are in memory, so first I sample a random image, then random location, and form the whole batch this way, so it's completely random.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "188465": "Does anyone who train model on tile level use some sort of database (leveldb,lmdb,hdf5)?\nLooks like lots of small files on HDD is not very good solution.",
    "188469": "Just a quick test:\n\nInitial image size 10.3 Mb.\n\n&gt;     du -sh DB/*\n&gt;     39M\tDB/db_type_1_image_id_0\n&gt;     221M\tDB/db_type_2_image_id_0\n&gt;     177M\tDB/db_type_3_image_id_0\n\n\n    import cv2\n    import numpy as np\n    import os\n    \n    image_id= 0\n    tile_size= 256\n    tile_stride= tile_size/2\n    \n    train_image_dir='Train'\n    database_dir='DB'\n    \n    if not os.path.exists(database_dir):\n        os.makedirs(database_dir)\n        \n    #Just create folder, cut image to tiles and store as jpg images in dir\n    def create_database_1(image_id):\n    \timg= cv2.imread(train_image_dir+'/'+str(image_id)+'.jpg')\n    \t\n    \tif not os.path.exists(database_dir+'/db_1_'+str(image_id)):\n    \t\tos.makedirs(database_dir+'/db_type_1_image_id_'+str(image_id))\n    \t\n    \tw= img.shape[1]\n    \th= img.shape[0]\n    \n    \tcounter=0\n    \tfor y in range(0, h, tile_stride):\n    \t\tfor x in range(0, w, tile_stride):\n    \t\t\tif(y+tile_size &lt; h and x+tile_size &lt; w):\n    \t\t\t\ttile_img= img[y:y+tile_size,x:x+tile_size]\n    \t\t\t\t\n    \t\t\t\tcv2.imwrite(database_dir+'/db_type_1_image_id_'+str(image_id)+'/'+str(counter)+'.jpg',tile_img)\n    \t\t\t\t\n    \t\t\t\tcounter=counter+1\n    \t\n    \tprint 'Numer of tiles:',counter\n    \n    #Create .npy file with all tiles via numpy\n    def create_database_2(image_id):\n    \timg= cv2.imread(train_image_dir+'/'+str(image_id)+'.jpg')\n    \t\n    \tif not os.path.exists(database_dir+'/db_type_2_image_id_'+str(image_id)):\n    \t\tos.makedirs(database_dir+'/db_type_2_image_id_'+str(image_id))\n    \t\n    \tw= img.shape[1]\n    \th= img.shape[0]\n    \t\n    \tarr_list=[]\n    \tcounter=0\n    \tfor y in range(0, h, tile_stride):\n    \t\tfor x in range(0, w, tile_stride):\n    \t\t\tif(y+tile_size &lt; h and x+tile_size &lt; w):\n    \t\t\t\ttile_img= img[y:y+tile_size,x:x+tile_size]\n    \t\t\t\tarr_list.append(tile_img)\n    \t\t\t\tcounter=counter+1\n    \t\n    \tarr= np.asarray(arr_list)\n    \tprint 'arr.shape', arr.shape\n    \t\n    \tnp.save(database_dir+'/db_type_2_image_id_'+str(image_id)+'/tiles.npy', arr)\n    \t\n    \tprint 'Numer of tiles:',counter\n    \n    #Create .npz file with all tiles via numpy\n    def create_database_3(image_id):\n    \timg= cv2.imread(train_image_dir+'/'+str(image_id)+'.jpg')\n    \t\n    \tif not os.path.exists(database_dir+'/db_type_3_image_id_'+str(image_id)):\n    \t\tos.makedirs(database_dir+'/db_type_3_image_id_'+str(image_id))\n    \t\n    \tw= img.shape[1]\n    \th= img.shape[0]\n    \t\n    \tarr_list=[]\n    \tcounter=0\n    \tfor y in range(0, h, tile_stride):\n    \t\tfor x in range(0, w, tile_stride):\n    \t\t\tif(y+tile_size &lt; h and x+tile_size &lt; w):\n    \t\t\t\ttile_img= img[y:y+tile_size,x:x+tile_size]\n    \t\t\t\tarr_list.append(tile_img)\n    \t\t\t\tcounter=counter+1\n    \t\n    \tarr= np.asarray(arr_list)\n    \tprint 'arr.shape', arr.shape\n    \t\n    \tnp.savez_compressed(database_dir+'/db_type_3_image_id_'+str(image_id)+'/tiles.npz', arr)\n    \t\n    \tprint 'Numer of tiles:',counter\n    \t\n    create_database_1(image_id)\n    create_database_2(image_id)\n    create_database_3(image_id)",
    "188542": "I'm loading most images in memory as uint8 numpy arrays (they are stored in this form on disk), and then generating patches on the fly during training. But that barely fits in 64GB RAM :(",
    "188565": "Why do you store images as numpy arrays? for faster loading?\nAlso if you store one .npy file per image then you can just get random batch from this image and not from entire dataset which sounds suboptimal, my idea is to save shuffled tiles from all images to some db and then sequentially read from it (like in Caffe).",
    "188573": "Yes, I store them in numpy arrays for faster loading: for example if do some experiment on 100 images, and they are already in disk cache, then loading takes just a couple of seconds. I don't have an SSD, just an HDD, so I optimize for that.\n\nI'm sampling random batches from the entire dataset: all images are in memory, so first I sample a random image, then random location, and form the whole batch this way, so it's completely random."
  },
  "source": "meta"
}