{
  "id": 70092,
  "title": "Is 32 GB RAM sufficient for this competition?",
  "url": "/competitions/human-protein-atlas-image-classification/discussion/70092",
  "author_name": "",
  "post_date": "2018-10-30T18:30:04.757504200Z",
  "votes": 2,
  "comment_count": 43,
  "views": 0,
  "content": "<p>Hi everyone,</p>\n\n<p>I consider starting this competition with my local computer. Data amount for this competition is huge. However, I am not sure about whether my laptop is gonna be enough for this competition. Is having a i7 processor(7th generation) and GTX 1050 sufficient? Thanks.</p>",
  "messages": [
    {
      "id": "412762",
      "postDate": "10/30/2018 18:30:04",
      "content": "<p>Hi everyone,</p>\n\n<p>I consider starting this competition with my local computer. Data amount for this competition is huge. However, I am not sure about whether my laptop is gonna be enough for this competition. Is having a i7 processor(7th generation) and GTX 1050 sufficient? Thanks.</p>",
      "rawMarkdown": "Hi everyone,\n\nI consider starting this competition with my local computer. Data amount for this competition is huge. However, I am not sure about whether my laptop is gonna be enough for this competition. Is having a i7 processor(7th generation) and GTX 1050 sufficient? Thanks.",
      "votes": null
    },
    {
      "id": "412765",
      "postDate": "10/30/2018 18:35:39",
      "content": "<p>In general it wont be great but it will work. Using 512x512 you will have to use generators and small batch sizes. My dev system is 64GB, I7-5930k, GTX Titan X, GTX 970. Most of the time I am using Keras, splitting batches between the two cards, limited by the 3.5GB usable memory on the 970. </p>",
      "rawMarkdown": "In general it wont be great but it will work. Using 512x512 you will have to use generators and small batch sizes. My dev system is 64GB, I7-5930k, GTX Titan X, GTX 970. Most of the time I am using Keras, splitting batches between the two cards, limited by the 3.5GB usable memory on the 970.",
      "votes": null
    },
    {
      "id": "412784",
      "postDate": "10/30/2018 19:11:28",
      "content": "<p>IMHO with 32GB you can only do serious 256x256 experiments on at least 1070. If you are aiming for high silver and above, you'll need at least 64GB ram and 1080Ti.</p>",
      "rawMarkdown": "IMHO with 32GB you can only do serious 256x256 experiments on at least 1070. If you are aiming for high silver and above, you'll need at least 64GB ram and 1080Ti.",
      "votes": null
    },
    {
      "id": "412814",
      "postDate": "10/30/2018 20:31:02",
      "content": "<p>What would you say were the requirements for the TGS Salt Prediction Challenge?</p>",
      "rawMarkdown": "What would you say were the requirements for the TGS Salt Prediction Challenge?",
      "votes": null
    },
    {
      "id": "412817",
      "postDate": "10/30/2018 20:34:03",
      "content": "<p>You can always work in float16</p>",
      "rawMarkdown": "You can always work in float16",
      "votes": null
    },
    {
      "id": "412820",
      "postDate": "10/30/2018 20:55:04",
      "content": "<p>This is definitely high requirement than the TGS challenge. I started with a 24GB machine then and ran into issues. I don't think I would be able to run this challenge in 24GB.</p>",
      "rawMarkdown": "This is definitely high requirement than the TGS challenge. I started with a 24GB machine then and ran into issues. I don't think I would be able to run this challenge in 24GB.",
      "votes": null
    },
    {
      "id": "412883",
      "postDate": "10/31/2018 00:47:43",
      "content": "<p>Whoa... feels like I should give up...</p>",
      "rawMarkdown": "Whoa... feels like I should give up...",
      "votes": null
    },
    {
      "id": "412887",
      "postDate": "10/31/2018 00:55:52",
      "content": "<p>FYI: <a href=\"https://stackoverflow.com/questions/46613748/float16-vs-float32-for-convolutional-neural-networks\">https://stackoverflow.com/questions/46613748/float16-vs-float32-for-convolutional-neural-networks</a></p>",
      "rawMarkdown": "FYI: https://stackoverflow.com/questions/46613748/float16-vs-float32-for-convolutional-neural-networks",
      "votes": null
    },
    {
      "id": "412890",
      "postDate": "10/31/2018 01:11:53",
      "content": "<p>Last time I tried that in Keras I got nan for the loss. Can't explain why. Maybe newer versions could go with that... ?</p>",
      "rawMarkdown": "Last time I tried that in Keras I got nan for the loss. Can't explain why. Maybe newer versions could go with that... ?",
      "votes": null
    },
    {
      "id": "413059",
      "postDate": "10/31/2018 08:11:38",
      "content": "<p>with fastai is pretty straightforward</p>",
      "rawMarkdown": "with fastai is pretty straightforward",
      "votes": null
    },
    {
      "id": "413061",
      "postDate": "10/31/2018 08:12:23",
      "content": "<p>Try my smaller custom datasets, 64, 128 and 256. </p>",
      "rawMarkdown": "Try my smaller custom datasets, 64, 128 and 256.",
      "votes": null
    },
    {
      "id": "413142",
      "postDate": "10/31/2018 11:06:37",
      "content": "<p>I believe 512 brings better results, though. (Still testing). </p>",
      "rawMarkdown": "I believe 512 brings better results, though. (Still testing).",
      "votes": null
    },
    {
      "id": "413261",
      "postDate": "10/31/2018 15:37:07",
      "content": "<p>Ultimately, I agree. My best working models are 512x512. I still think with 32GB ram you can do it. With 64GB sometimes I have 2 Jupyter notebooks open running different experiments on each video card and it works out ok. I'm working on moving up to the full size images.</p>",
      "rawMarkdown": "Ultimately, I agree. My best working models are 512x512. I still think with 32GB ram you can do it. With 64GB sometimes I have 2 Jupyter notebooks open running different experiments on each video card and it works out ok. I'm working on moving up to the full size images.",
      "votes": null
    },
    {
      "id": "413288",
      "postDate": "10/31/2018 16:23:59",
      "content": "<p>A good example right now is my model 19. It's currently training on both video cards, batch sized limited by the 3.5GB GTX970. I also have 7 zip running, unpacking the huge tiff files, chrome with a few tabs open. Total memory used is only 22GB</p>",
      "rawMarkdown": "A good example right now is my model 19. It's currently training on both video cards, batch sized limited by the 3.5GB GTX970. I also have 7 zip running, unpacking the huge tiff files, chrome with a few tabs open. Total memory used is only 22GB",
      "votes": null
    },
    {
      "id": "413354",
      "postDate": "10/31/2018 19:18:02",
      "content": "<p>You may have forgotten to do clipping in focal loss.</p>",
      "rawMarkdown": "You may have forgotten to do clipping in focal loss.",
      "votes": null
    },
    {
      "id": "413356",
      "postDate": "10/31/2018 19:19:35",
      "content": "<p>I tend to do more augmentations and would like all my images loaded in memory before training. So even with 512x512 it's a pain for me. Now with batch size=48 I can barely do 3 batches per second, compared to over 10 for 256x256.</p>",
      "rawMarkdown": "I tend to do more augmentations and would like all my images loaded in memory before training. So even with 512x512 it's a pain for me. Now with batch size=48 I can barely do 3 batches per second, compared to over 10 for 256x256.",
      "votes": null
    },
    {
      "id": "413361",
      "postDate": "10/31/2018 19:31:05",
      "content": "<p>My model and my GPU becomes the bottleneck at the higher resolutions. The conv layers take a lot of time. If I train only the dense half I am limited by CPU for augmentations. Without augmentation the limit is the SSD speed, at 240MB/sec.  I do run into memory issues if I try to load them all, gave up on that approach. Especially with the tif files...</p>",
      "rawMarkdown": "My model and my GPU becomes the bottleneck at the higher resolutions. The conv layers take a lot of time. If I train only the dense half I am limited by CPU for augmentations. Without augmentation the limit is the SSD speed, at 240MB/sec.  I do run into memory issues if I try to load them all, gave up on that approach. Especially with the tif files...",
      "votes": null
    },
    {
      "id": "413363",
      "postDate": "10/31/2018 19:38:59",
      "content": "<p>Ok... found the problem(s)</p>\n\n<p>I need to set <code>epsilon</code> (the small constant to avoid dividing by zero) to a much higher value, such as <code>1e-4</code> or higher for training to succeed. </p>\n\n<p>Also, due to the Tensorflow implementation of BatchNormalization, it's necessary to rewrite this Keras layer in a way that supports the rest of the model in float16.</p>\n\n<p>Now, a newbie question.</p>\n\n<p>I managed to train in float16, but shouldn't this occupy less GPU? I mean, shouldn't I be able to use bigger batches? <br>\nAnd if batches are the same size, shouldn't I be able to train a lot faster?   </p>\n\n<p>Tests are showing absolutely no difference in speed or batch size for float16 and float32 training. </p>",
      "rawMarkdown": "Ok... found the problem(s)\n\nI need to set `epsilon` (the small constant to avoid dividing by zero) to a much higher value, such as `1e-4` or higher for training to succeed. \n\nAlso, due to the Tensorflow implementation of BatchNormalization, it's necessary to rewrite this Keras layer in a way that supports the rest of the model in float16.\n\nNow, a newbie question.\n\nI managed to train in float16, but shouldn't this occupy less GPU? I mean, shouldn't I be able to use bigger batches?    \nAnd if batches are the same size, shouldn't I be able to train a lot faster?   \n\nTests are showing absolutely no difference in speed or batch size for float16 and float32 training.",
      "votes": null
    },
    {
      "id": "414036",
      "postDate": "11/02/2018 01:59:39",
      "content": "<p>I have two 1080 TI cards but just 32GB ram.. I probably have to upgrade ram ..</p>",
      "rawMarkdown": "I have two 1080 TI cards but just 32GB ram.. I probably have to upgrade ram ..",
      "votes": null
    },
    {
      "id": "414049",
      "postDate": "11/02/2018 02:26:03",
      "content": "<p>I've started working with the images at 1024x1024. 64gb is needed at this size </p>",
      "rawMarkdown": "I've started working with the images at 1024x1024. 64gb is needed at this size",
      "votes": null
    },
    {
      "id": "414110",
      "postDate": "11/02/2018 06:44:12",
      "content": "<p>I don't see why you would need this much RAM. Save the images in an uncompressed format on an SSD and load them on the fly with background workers -&gt; all RAM issues solved :-)</p>",
      "rawMarkdown": "I don't see why you would need this much RAM. Save the images in an uncompressed format on an SSD and load them on the fly with background workers -&gt; all RAM issues solved :-)",
      "votes": null
    },
    {
      "id": "414133",
      "postDate": "11/02/2018 07:33:06",
      "content": "<p>I moved up to 1536x1536 and had to do this. Both my training and validation load from SSD now. While training python is only using 10GB ram.</p>",
      "rawMarkdown": "I moved up to 1536x1536 and had to do this. Both my training and validation load from SSD now. While training python is only using 10GB ram.",
      "votes": null
    },
    {
      "id": "414680",
      "postDate": "11/03/2018 11:23:12",
      "content": "<p>Daniel, did you succeed in implementing a BN layer in Keras with float16 support?</p>",
      "rawMarkdown": "Daniel, did you succeed in implementing a BN layer in Keras with float16 support?",
      "votes": null
    },
    {
      "id": "414966",
      "postDate": "11/04/2018 01:44:51",
      "content": "<p>I made casts to float32 (because tensorflow backend demands that for fused batch norm) in a custom BN layer. (This is still faster and occupies less memory than using a standard batchnormalization in float16 - Tested).    </p>\n\n<p>But this layer with float32 will reflect in the optimizer, and the optimizer will need to be customized for creating moments and learning rates with the proper format depending on which weight matrix it's calculating. </p>\n\n<p>I made everything working, both cases:</p>\n\n<ul>\n<li>Convs in float16, BN in float32, mixed optimizer    </li>\n<li>Convs and BN in float16 (not using fused batch norm)    </li>\n</ul>\n\n<p>The first occupies roughly the same GPU memory and has about the same speed as the original model entirely in float32 (I don't understand why)    </p>\n\n<p>The second is slower and occupies more memory. (I understand even less... fused batch norm must be a real good thing)    </p>",
      "rawMarkdown": "I made casts to float32 (because tensorflow backend demands that for fused batch norm) in a custom BN layer. (This is still faster and occupies less memory than using a standard batchnormalization in float16 - Tested).    \n\nBut this layer with float32 will reflect in the optimizer, and the optimizer will need to be customized for creating moments and learning rates with the proper format depending on which weight matrix it's calculating. \n\nI made everything working, both cases:\n\n - Convs in float16, BN in float32, mixed optimizer    \n - Convs and BN in float16 (not using fused batch norm)    \n\nThe first occupies roughly the same GPU memory and has about the same speed as the original model entirely in float32 (I don't understand why)    \n\nThe second is slower and occupies more memory. (I understand even less... fused batch norm must be a real good thing)",
      "votes": null
    },
    {
      "id": "414973",
      "postDate": "11/04/2018 02:03:29",
      "content": "<p>There's a github ticket open in keras for the batch norm issue. I ran into it trying float16</p>",
      "rawMarkdown": "There's a github ticket open in keras for the batch norm issue. I ran into it trying float16",
      "votes": null
    },
    {
      "id": "415004",
      "postDate": "11/04/2018 05:21:19",
      "content": "<p>Can we load the data by chunks and then train them?</p>",
      "rawMarkdown": "Can we load the data by chunks and then train them?",
      "votes": null
    },
    {
      "id": "415191",
      "postDate": "11/04/2018 16:07:25",
      "content": "<p>Check out these two kernels for the solutions:</p>\n\n<ul>\n<li><a href=\"https://www.kaggle.com/danmoller/keras-training-with-float16-test-kernel-1\">https://www.kaggle.com/danmoller/keras-training-with-float16-test-kernel-1</a> (BN is float 32)</li>\n<li><a href=\"https://www.kaggle.com/danmoller/keras-training-with-float16-test-kernel-2\">https://www.kaggle.com/danmoller/keras-training-with-float16-test-kernel-2</a> (BN is float 16)</li>\n</ul>\n\n<p>They don't seem a good thing to do, though, unless there is a bug or I'm missing something.    </p>",
      "rawMarkdown": "Check out these two kernels for the solutions:\n\n- https://www.kaggle.com/danmoller/keras-training-with-float16-test-kernel-1 (BN is float 32)\n- https://www.kaggle.com/danmoller/keras-training-with-float16-test-kernel-2 (BN is float 16)\n\nThey don't seem a good thing to do, though, unless there is a bug or I'm missing something.",
      "votes": null
    },
    {
      "id": "415299",
      "postDate": "11/04/2018 20:50:16",
      "content": "<p>I want to update my comments on this. As I tested at 1024 and 2048, I couldn't even keep the validation set in memory. In Keras I've switched to using data generators from directories for everything. It is a bit slower but works well with low memory usage. At 1024 I'm using less that 16GB system memory.</p>",
      "rawMarkdown": "I want to update my comments on this. As I tested at 1024 and 2048, I couldn't even keep the validation set in memory. In Keras I've switched to using data generators from directories for everything. It is a bit slower but works well with low memory usage. At 1024 I'm using less that 16GB system memory.",
      "votes": null
    },
    {
      "id": "415502",
      "postDate": "11/05/2018 08:31:52",
      "content": "<p>1 1080 Ti and 32 GB ram. No real problem using 512x512. I’ll try the big ones too but it may be impractical.</p>",
      "rawMarkdown": "1 1080 Ti and 32 GB ram. No real problem using 512x512. I’ll try the big ones too but it may be impractical.",
      "votes": null
    },
    {
      "id": "415508",
      "postDate": "11/05/2018 08:42:10",
      "content": "<p>I  Would suggest loading it in batches. Use generators and load the data. I think 32GB is enough.</p>",
      "rawMarkdown": "I  Would suggest loading it in batches. Use generators and load the data. I think 32GB is enough.",
      "votes": null
    },
    {
      "id": "415512",
      "postDate": "11/05/2018 08:46:58",
      "content": "<p>I am working on both 512x512 and 1024x1024 with 2 GPUs (two trainings in parallel) and just 32GB of RAM. Loading everything from SSD is what makes it possible, but you will need background workers for that to prevent IO bottlenecks</p>",
      "rawMarkdown": "I am working on both 512x512 and 1024x1024 with 2 GPUs (two trainings in parallel) and just 32GB of RAM. Loading everything from SSD is what makes it possible, but you will need background workers for that to prevent IO bottlenecks",
      "votes": null
    },
    {
      "id": "415543",
      "postDate": "11/05/2018 09:47:42",
      "content": "<p>When using generators, 16 is quite ok. </p>\n\n<p>The problem might be the loading and augmenting speed. </p>",
      "rawMarkdown": "When using generators, 16 is quite ok. \n\nThe problem might be the loading and augmenting speed.",
      "votes": null
    },
    {
      "id": "415549",
      "postDate": "11/05/2018 09:54:50",
      "content": "<p>Augmentations are indeed a problem. My 8C/16T CPU cannot keep up with feeding my two GPUs :-c</p>",
      "rawMarkdown": "Augmentations are indeed a problem. My 8C/16T CPU cannot keep up with feeding my two GPUs :-c",
      "votes": null
    },
    {
      "id": "415562",
      "postDate": "11/05/2018 10:18:52",
      "content": "<p>if you are using keras, then try setting multi processing to true or num_workers to -1 or (maximum cores -1)</p>",
      "rawMarkdown": "if you are using keras, then try setting multi processing to true or num_workers to -1 or (maximum cores -1)",
      "votes": null
    },
    {
      "id": "415575",
      "postDate": "11/05/2018 10:48:52",
      "content": "<p>Not using Keras and I would not point out 8C/16T if I was not using them already ;-) Any augmentation that needs resampling into a new image grid (elastic deformation, rotation, scaling, ..) is expensive</p>",
      "rawMarkdown": "Not using Keras and I would not point out 8C/16T if I was not using them already ;-) Any augmentation that needs resampling into a new image grid (elastic deformation, rotation, scaling, ..) is expensive",
      "votes": null
    },
    {
      "id": "415589",
      "postDate": "11/05/2018 11:13:49",
      "content": "<p>sorry, my bad. i actually wanted to reply to Daniel Moller's post. I just saw yours. And yeah.. its an expensive process for sure. But, don't you  think..we can just do it once...and then use this generated data for all of the further model tuning?</p>",
      "rawMarkdown": "sorry, my bad. i actually wanted to reply to Daniel Moller's post. I just saw yours. And yeah.. its an expensive process for sure. But, don't you  think..we can just do it once...and then use this generated data for all of the further model tuning?",
      "votes": null
    },
    {
      "id": "415603",
      "postDate": "11/05/2018 11:33:55",
      "content": "<p>It might be good depending on how much augmentation is necessary.</p>\n\n<p>If the problem needs a lot, I don't think it's very good, unless you've got tons of disk. Because strong augmentations rely a lot on random generations and saving augmented images would not bring the same variability. </p>",
      "rawMarkdown": "It might be good depending on how much augmentation is necessary.\n\nIf the problem needs a lot, I don't think it's very good, unless you've got tons of disk. Because strong augmentations rely a lot on random generations and saving augmented images would not bring the same variability.",
      "votes": null
    },
    {
      "id": "415605",
      "postDate": "11/05/2018 11:36:32",
      "content": "<p>Excatly. Doing the augmentations on the fly will allow new augmentation parameters each time an example is chosen thus creating a lot more variability. How much this actually matters I don't know</p>",
      "rawMarkdown": "Excatly. Doing the augmentations on the fly will allow new augmentation parameters each time an example is chosen thus creating a lot more variability. How much this actually matters I don't know",
      "votes": null
    },
    {
      "id": "415626",
      "postDate": "11/05/2018 12:12:36",
      "content": "<p>I have a feeling this problem will not require heavy augmentation (but I'm just starting).   </p>\n\n<p>But we can always get 8 flips (faster than distortions/rotations) from any data combining:</p>\n\n<pre><code>flipMode = random.randint(0,7)\nif flipMode in [4,5,6,7]:\n    x = np.flip(x,axis1)\nif flipMode in [1,3,5,7]:\n    x = np.flip(x,axis2)\nif flipMode in [2,3,6,7]:\n    x = np.swapaxes(x,axis1,axis2)\n</code></pre>",
      "rawMarkdown": "I have a feeling this problem will not require heavy augmentation (but I'm just starting).   \n\nBut we can always get 8 flips (faster than distortions/rotations) from any data combining:\n\n    flipMode = random.randint(0,7)\n    if flipMode in [4,5,6,7]:\n        x = np.flip(x,axis1)\n    if flipMode in [1,3,5,7]:\n        x = np.flip(x,axis2)\n    if flipMode in [2,3,6,7]:\n        x = np.swapaxes(x,axis1,axis2)",
      "votes": null
    },
    {
      "id": "415630",
      "postDate": "11/05/2018 12:18:42",
      "content": "<p>You may be right. I did not investigate it at all. Currently, I am doing rotations, deformations, scaling, gamma, gaussian noise, axes transpose, mirror and contrast (each of those with some probability for all samples)</p>",
      "rawMarkdown": "You may be right. I did not investigate it at all. Currently, I am doing rotations, deformations, scaling, gamma, gaussian noise, axes transpose, mirror and contrast (each of those with some probability for all samples)",
      "votes": null
    },
    {
      "id": "416745",
      "postDate": "11/07/2018 08:03:03",
      "content": "<p>well first of all you should worry more about your GPU RAM as this limits you batch size and all, (as larger batch size is usually better for models with BatchNorm).\nif you worry about running out RAM, you can always use data generators and online image augmentation. with a decent CPU(which you have) and SSD you won't have too much speed loss imo. don't forget to use multiprocessing for your data generator though.</p>",
      "rawMarkdown": "well first of all you should worry more about your GPU RAM as this limits you batch size and all, (as larger batch size is usually better for models with BatchNorm).\nif you worry about running out RAM, you can always use data generators and online image augmentation. with a decent CPU(which you have) and SSD you won't have too much speed loss imo. don't forget to use multiprocessing for your data generator though.",
      "votes": null
    },
    {
      "id": "416783",
      "postDate": "11/07/2018 09:19:44",
      "content": "<p>Yes, I agree. The GPU RAM seems to be the bottleneck here</p>",
      "rawMarkdown": "Yes, I agree. The GPU RAM seems to be the bottleneck here",
      "votes": null
    },
    {
      "id": "417057",
      "postDate": "11/07/2018 17:06:42",
      "content": "<p>I'm quite lacking in RAM and am able to train ok by loading from SSD. I'm using 512x512 images with a 1070 and 8GB of system RAM.</p>",
      "rawMarkdown": "I'm quite lacking in RAM and am able to train ok by loading from SSD. I'm using 512x512 images with a 1070 and 8GB of system RAM.",
      "votes": null
    },
    {
      "id": "417232",
      "postDate": "11/08/2018 01:03:01",
      "content": "<p>So its kinda like a dropout?</p>",
      "rawMarkdown": "So its kinda like a dropout?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 412765,
      "author_name": "ldm314",
      "author_url": "",
      "post_date": "10/30/2018 18:35:39",
      "content": "<p>In general it wont be great but it will work. Using 512x512 you will have to use generators and small batch sizes. My dev system is 64GB, I7-5930k, GTX Titan X, GTX 970. Most of the time I am using Keras, splitting batches between the two cards, limited by the 3.5GB usable memory on the 970. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 412784,
      "author_name": "alexanderliao",
      "author_url": "",
      "post_date": "10/30/2018 19:11:28",
      "content": "<p>IMHO with 32GB you can only do serious 256x256 experiments on at least 1070. If you are aiming for high silver and above, you'll need at least 64GB ram and 1080Ti.</p>",
      "votes": null,
      "replies": [
        {
          "id": 412814,
          "author_name": "danmoller",
          "author_url": "",
          "post_date": "10/30/2018 20:31:02",
          "content": "<p>What would you say were the requirements for the TGS Salt Prediction Challenge?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 412820,
          "author_name": "ldm314",
          "author_url": "",
          "post_date": "10/30/2018 20:55:04",
          "content": "<p>This is definitely high requirement than the TGS challenge. I started with a 24GB machine then and ran into issues. I don't think I would be able to run this challenge in 24GB.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 412883,
          "author_name": "danmoller",
          "author_url": "",
          "post_date": "10/31/2018 00:47:43",
          "content": "<p>Whoa... feels like I should give up...</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 413061,
          "author_name": "tcapelle",
          "author_url": "",
          "post_date": "10/31/2018 08:12:23",
          "content": "<p>Try my smaller custom datasets, 64, 128 and 256. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 413142,
          "author_name": "danmoller",
          "author_url": "",
          "post_date": "10/31/2018 11:06:37",
          "content": "<p>I believe 512 brings better results, though. (Still testing). </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 413261,
          "author_name": "ldm314",
          "author_url": "",
          "post_date": "10/31/2018 15:37:07",
          "content": "<p>Ultimately, I agree. My best working models are 512x512. I still think with 32GB ram you can do it. With 64GB sometimes I have 2 Jupyter notebooks open running different experiments on each video card and it works out ok. I'm working on moving up to the full size images.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 413288,
          "author_name": "ldm314",
          "author_url": "",
          "post_date": "10/31/2018 16:23:59",
          "content": "<p>A good example right now is my model 19. It's currently training on both video cards, batch sized limited by the 3.5GB GTX970. I also have 7 zip running, unpacking the huge tiff files, chrome with a few tabs open. Total memory used is only 22GB</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 413356,
          "author_name": "alexanderliao",
          "author_url": "",
          "post_date": "10/31/2018 19:19:35",
          "content": "<p>I tend to do more augmentations and would like all my images loaded in memory before training. So even with 512x512 it's a pain for me. Now with batch size=48 I can barely do 3 batches per second, compared to over 10 for 256x256.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 413361,
          "author_name": "ldm314",
          "author_url": "",
          "post_date": "10/31/2018 19:31:05",
          "content": "<p>My model and my GPU becomes the bottleneck at the higher resolutions. The conv layers take a lot of time. If I train only the dense half I am limited by CPU for augmentations. Without augmentation the limit is the SSD speed, at 240MB/sec.  I do run into memory issues if I try to load them all, gave up on that approach. Especially with the tif files...</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 412817,
      "author_name": "tcapelle",
      "author_url": "",
      "post_date": "10/30/2018 20:34:03",
      "content": "<p>You can always work in float16</p>",
      "votes": null,
      "replies": [
        {
          "id": 412887,
          "author_name": "alexanderliao",
          "author_url": "",
          "post_date": "10/31/2018 00:55:52",
          "content": "<p>FYI: <a href=\"https://stackoverflow.com/questions/46613748/float16-vs-float32-for-convolutional-neural-networks\">https://stackoverflow.com/questions/46613748/float16-vs-float32-for-convolutional-neural-networks</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 412890,
          "author_name": "danmoller",
          "author_url": "",
          "post_date": "10/31/2018 01:11:53",
          "content": "<p>Last time I tried that in Keras I got nan for the loss. Can't explain why. Maybe newer versions could go with that... ?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 413059,
          "author_name": "tcapelle",
          "author_url": "",
          "post_date": "10/31/2018 08:11:38",
          "content": "<p>with fastai is pretty straightforward</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 413354,
          "author_name": "alexanderliao",
          "author_url": "",
          "post_date": "10/31/2018 19:18:02",
          "content": "<p>You may have forgotten to do clipping in focal loss.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 413363,
          "author_name": "danmoller",
          "author_url": "",
          "post_date": "10/31/2018 19:38:59",
          "content": "<p>Ok... found the problem(s)</p>\n\n<p>I need to set <code>epsilon</code> (the small constant to avoid dividing by zero) to a much higher value, such as <code>1e-4</code> or higher for training to succeed. </p>\n\n<p>Also, due to the Tensorflow implementation of BatchNormalization, it's necessary to rewrite this Keras layer in a way that supports the rest of the model in float16.</p>\n\n<p>Now, a newbie question.</p>\n\n<p>I managed to train in float16, but shouldn't this occupy less GPU? I mean, shouldn't I be able to use bigger batches? <br>\nAnd if batches are the same size, shouldn't I be able to train a lot faster?   </p>\n\n<p>Tests are showing absolutely no difference in speed or batch size for float16 and float32 training. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 414680,
          "author_name": "stecasasso",
          "author_url": "",
          "post_date": "11/03/2018 11:23:12",
          "content": "<p>Daniel, did you succeed in implementing a BN layer in Keras with float16 support?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 414966,
          "author_name": "danmoller",
          "author_url": "",
          "post_date": "11/04/2018 01:44:51",
          "content": "<p>I made casts to float32 (because tensorflow backend demands that for fused batch norm) in a custom BN layer. (This is still faster and occupies less memory than using a standard batchnormalization in float16 - Tested).    </p>\n\n<p>But this layer with float32 will reflect in the optimizer, and the optimizer will need to be customized for creating moments and learning rates with the proper format depending on which weight matrix it's calculating. </p>\n\n<p>I made everything working, both cases:</p>\n\n<ul>\n<li>Convs in float16, BN in float32, mixed optimizer    </li>\n<li>Convs and BN in float16 (not using fused batch norm)    </li>\n</ul>\n\n<p>The first occupies roughly the same GPU memory and has about the same speed as the original model entirely in float32 (I don't understand why)    </p>\n\n<p>The second is slower and occupies more memory. (I understand even less... fused batch norm must be a real good thing)    </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 414973,
          "author_name": "ldm314",
          "author_url": "",
          "post_date": "11/04/2018 02:03:29",
          "content": "<p>There's a github ticket open in keras for the batch norm issue. I ran into it trying float16</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 415191,
          "author_name": "danmoller",
          "author_url": "",
          "post_date": "11/04/2018 16:07:25",
          "content": "<p>Check out these two kernels for the solutions:</p>\n\n<ul>\n<li><a href=\"https://www.kaggle.com/danmoller/keras-training-with-float16-test-kernel-1\">https://www.kaggle.com/danmoller/keras-training-with-float16-test-kernel-1</a> (BN is float 32)</li>\n<li><a href=\"https://www.kaggle.com/danmoller/keras-training-with-float16-test-kernel-2\">https://www.kaggle.com/danmoller/keras-training-with-float16-test-kernel-2</a> (BN is float 16)</li>\n</ul>\n\n<p>They don't seem a good thing to do, though, unless there is a bug or I'm missing something.    </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 417232,
          "author_name": "kokecacao",
          "author_url": "",
          "post_date": "11/08/2018 01:03:01",
          "content": "<p>So its kinda like a dropout?</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 414036,
      "author_name": "vivek8943",
      "author_url": "",
      "post_date": "11/02/2018 01:59:39",
      "content": "<p>I have two 1080 TI cards but just 32GB ram.. I probably have to upgrade ram ..</p>",
      "votes": null,
      "replies": [
        {
          "id": 414049,
          "author_name": "ldm314",
          "author_url": "",
          "post_date": "11/02/2018 02:26:03",
          "content": "<p>I've started working with the images at 1024x1024. 64gb is needed at this size </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 415502,
          "author_name": "petewills",
          "author_url": "",
          "post_date": "11/05/2018 08:31:52",
          "content": "<p>1 1080 Ti and 32 GB ram. No real problem using 512x512. I’ll try the big ones too but it may be impractical.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 415512,
          "author_name": "fabianisensee",
          "author_url": "",
          "post_date": "11/05/2018 08:46:58",
          "content": "<p>I am working on both 512x512 and 1024x1024 with 2 GPUs (two trainings in parallel) and just 32GB of RAM. Loading everything from SSD is what makes it possible, but you will need background workers for that to prevent IO bottlenecks</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 417057,
          "author_name": "pjbutcher",
          "author_url": "",
          "post_date": "11/07/2018 17:06:42",
          "content": "<p>I'm quite lacking in RAM and am able to train ok by loading from SSD. I'm using 512x512 images with a 1070 and 8GB of system RAM.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 414110,
      "author_name": "fabianisensee",
      "author_url": "",
      "post_date": "11/02/2018 06:44:12",
      "content": "<p>I don't see why you would need this much RAM. Save the images in an uncompressed format on an SSD and load them on the fly with background workers -&gt; all RAM issues solved :-)</p>",
      "votes": null,
      "replies": [
        {
          "id": 414133,
          "author_name": "ldm314",
          "author_url": "",
          "post_date": "11/02/2018 07:33:06",
          "content": "<p>I moved up to 1536x1536 and had to do this. Both my training and validation load from SSD now. While training python is only using 10GB ram.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 415004,
      "author_name": "konohayui",
      "author_url": "",
      "post_date": "11/04/2018 05:21:19",
      "content": "<p>Can we load the data by chunks and then train them?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 415299,
      "author_name": "ldm314",
      "author_url": "",
      "post_date": "11/04/2018 20:50:16",
      "content": "<p>I want to update my comments on this. As I tested at 1024 and 2048, I couldn't even keep the validation set in memory. In Keras I've switched to using data generators from directories for everything. It is a bit slower but works well with low memory usage. At 1024 I'm using less that 16GB system memory.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 415508,
      "author_name": "kalyanrdy",
      "author_url": "",
      "post_date": "11/05/2018 08:42:10",
      "content": "<p>I  Would suggest loading it in batches. Use generators and load the data. I think 32GB is enough.</p>",
      "votes": null,
      "replies": [
        {
          "id": 415543,
          "author_name": "danmoller",
          "author_url": "",
          "post_date": "11/05/2018 09:47:42",
          "content": "<p>When using generators, 16 is quite ok. </p>\n\n<p>The problem might be the loading and augmenting speed. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 415549,
          "author_name": "fabianisensee",
          "author_url": "",
          "post_date": "11/05/2018 09:54:50",
          "content": "<p>Augmentations are indeed a problem. My 8C/16T CPU cannot keep up with feeding my two GPUs :-c</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 415562,
          "author_name": "kalyanrdy",
          "author_url": "",
          "post_date": "11/05/2018 10:18:52",
          "content": "<p>if you are using keras, then try setting multi processing to true or num_workers to -1 or (maximum cores -1)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 415575,
          "author_name": "fabianisensee",
          "author_url": "",
          "post_date": "11/05/2018 10:48:52",
          "content": "<p>Not using Keras and I would not point out 8C/16T if I was not using them already ;-) Any augmentation that needs resampling into a new image grid (elastic deformation, rotation, scaling, ..) is expensive</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 415589,
          "author_name": "kalyanrdy",
          "author_url": "",
          "post_date": "11/05/2018 11:13:49",
          "content": "<p>sorry, my bad. i actually wanted to reply to Daniel Moller's post. I just saw yours. And yeah.. its an expensive process for sure. But, don't you  think..we can just do it once...and then use this generated data for all of the further model tuning?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 415603,
          "author_name": "danmoller",
          "author_url": "",
          "post_date": "11/05/2018 11:33:55",
          "content": "<p>It might be good depending on how much augmentation is necessary.</p>\n\n<p>If the problem needs a lot, I don't think it's very good, unless you've got tons of disk. Because strong augmentations rely a lot on random generations and saving augmented images would not bring the same variability. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 415605,
          "author_name": "fabianisensee",
          "author_url": "",
          "post_date": "11/05/2018 11:36:32",
          "content": "<p>Excatly. Doing the augmentations on the fly will allow new augmentation parameters each time an example is chosen thus creating a lot more variability. How much this actually matters I don't know</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 415626,
          "author_name": "danmoller",
          "author_url": "",
          "post_date": "11/05/2018 12:12:36",
          "content": "<p>I have a feeling this problem will not require heavy augmentation (but I'm just starting).   </p>\n\n<p>But we can always get 8 flips (faster than distortions/rotations) from any data combining:</p>\n\n<pre><code>flipMode = random.randint(0,7)\nif flipMode in [4,5,6,7]:\n    x = np.flip(x,axis1)\nif flipMode in [1,3,5,7]:\n    x = np.flip(x,axis2)\nif flipMode in [2,3,6,7]:\n    x = np.swapaxes(x,axis1,axis2)\n</code></pre>",
          "votes": null,
          "replies": []
        },
        {
          "id": 415630,
          "author_name": "fabianisensee",
          "author_url": "",
          "post_date": "11/05/2018 12:18:42",
          "content": "<p>You may be right. I did not investigate it at all. Currently, I am doing rotations, deformations, scaling, gamma, gaussian noise, axes transpose, mirror and contrast (each of those with some probability for all samples)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 416745,
      "author_name": "marsyeti",
      "author_url": "",
      "post_date": "11/07/2018 08:03:03",
      "content": "<p>well first of all you should worry more about your GPU RAM as this limits you batch size and all, (as larger batch size is usually better for models with BatchNorm).\nif you worry about running out RAM, you can always use data generators and online image augmentation. with a decent CPU(which you have) and SSD you won't have too much speed loss imo. don't forget to use multiprocessing for your data generator though.</p>",
      "votes": null,
      "replies": [
        {
          "id": 416783,
          "author_name": "stecasasso",
          "author_url": "",
          "post_date": "11/07/2018 09:19:44",
          "content": "<p>Yes, I agree. The GPU RAM seems to be the bottleneck here</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "412762": "Hi everyone,\n\nI consider starting this competition with my local computer. Data amount for this competition is huge. However, I am not sure about whether my laptop is gonna be enough for this competition. Is having a i7 processor(7th generation) and GTX 1050 sufficient? Thanks.",
    "412765": "In general it wont be great but it will work. Using 512x512 you will have to use generators and small batch sizes. My dev system is 64GB, I7-5930k, GTX Titan X, GTX 970. Most of the time I am using Keras, splitting batches between the two cards, limited by the 3.5GB usable memory on the 970.",
    "412784": "IMHO with 32GB you can only do serious 256x256 experiments on at least 1070. If you are aiming for high silver and above, you'll need at least 64GB ram and 1080Ti.",
    "412814": "What would you say were the requirements for the TGS Salt Prediction Challenge?",
    "412817": "You can always work in float16",
    "412820": "This is definitely high requirement than the TGS challenge. I started with a 24GB machine then and ran into issues. I don't think I would be able to run this challenge in 24GB.",
    "412883": "Whoa... feels like I should give up...",
    "412887": "FYI: https://stackoverflow.com/questions/46613748/float16-vs-float32-for-convolutional-neural-networks",
    "412890": "Last time I tried that in Keras I got nan for the loss. Can't explain why. Maybe newer versions could go with that... ?",
    "413059": "with fastai is pretty straightforward",
    "413061": "Try my smaller custom datasets, 64, 128 and 256.",
    "413142": "I believe 512 brings better results, though. (Still testing).",
    "413261": "Ultimately, I agree. My best working models are 512x512. I still think with 32GB ram you can do it. With 64GB sometimes I have 2 Jupyter notebooks open running different experiments on each video card and it works out ok. I'm working on moving up to the full size images.",
    "413288": "A good example right now is my model 19. It's currently training on both video cards, batch sized limited by the 3.5GB GTX970. I also have 7 zip running, unpacking the huge tiff files, chrome with a few tabs open. Total memory used is only 22GB",
    "413354": "You may have forgotten to do clipping in focal loss.",
    "413356": "I tend to do more augmentations and would like all my images loaded in memory before training. So even with 512x512 it's a pain for me. Now with batch size=48 I can barely do 3 batches per second, compared to over 10 for 256x256.",
    "413361": "My model and my GPU becomes the bottleneck at the higher resolutions. The conv layers take a lot of time. If I train only the dense half I am limited by CPU for augmentations. Without augmentation the limit is the SSD speed, at 240MB/sec.  I do run into memory issues if I try to load them all, gave up on that approach. Especially with the tif files...",
    "413363": "Ok... found the problem(s)\n\nI need to set `epsilon` (the small constant to avoid dividing by zero) to a much higher value, such as `1e-4` or higher for training to succeed. \n\nAlso, due to the Tensorflow implementation of BatchNormalization, it's necessary to rewrite this Keras layer in a way that supports the rest of the model in float16.\n\nNow, a newbie question.\n\nI managed to train in float16, but shouldn't this occupy less GPU? I mean, shouldn't I be able to use bigger batches?    \nAnd if batches are the same size, shouldn't I be able to train a lot faster?   \n\nTests are showing absolutely no difference in speed or batch size for float16 and float32 training.",
    "414036": "I have two 1080 TI cards but just 32GB ram.. I probably have to upgrade ram ..",
    "414049": "I've started working with the images at 1024x1024. 64gb is needed at this size",
    "414110": "I don't see why you would need this much RAM. Save the images in an uncompressed format on an SSD and load them on the fly with background workers -&gt; all RAM issues solved :-)",
    "414133": "I moved up to 1536x1536 and had to do this. Both my training and validation load from SSD now. While training python is only using 10GB ram.",
    "414680": "Daniel, did you succeed in implementing a BN layer in Keras with float16 support?",
    "414966": "I made casts to float32 (because tensorflow backend demands that for fused batch norm) in a custom BN layer. (This is still faster and occupies less memory than using a standard batchnormalization in float16 - Tested).    \n\nBut this layer with float32 will reflect in the optimizer, and the optimizer will need to be customized for creating moments and learning rates with the proper format depending on which weight matrix it's calculating. \n\nI made everything working, both cases:\n\n - Convs in float16, BN in float32, mixed optimizer    \n - Convs and BN in float16 (not using fused batch norm)    \n\nThe first occupies roughly the same GPU memory and has about the same speed as the original model entirely in float32 (I don't understand why)    \n\nThe second is slower and occupies more memory. (I understand even less... fused batch norm must be a real good thing)",
    "414973": "There's a github ticket open in keras for the batch norm issue. I ran into it trying float16",
    "415004": "Can we load the data by chunks and then train them?",
    "415191": "Check out these two kernels for the solutions:\n\n- https://www.kaggle.com/danmoller/keras-training-with-float16-test-kernel-1 (BN is float 32)\n- https://www.kaggle.com/danmoller/keras-training-with-float16-test-kernel-2 (BN is float 16)\n\nThey don't seem a good thing to do, though, unless there is a bug or I'm missing something.",
    "415299": "I want to update my comments on this. As I tested at 1024 and 2048, I couldn't even keep the validation set in memory. In Keras I've switched to using data generators from directories for everything. It is a bit slower but works well with low memory usage. At 1024 I'm using less that 16GB system memory.",
    "415502": "1 1080 Ti and 32 GB ram. No real problem using 512x512. I’ll try the big ones too but it may be impractical.",
    "415508": "I  Would suggest loading it in batches. Use generators and load the data. I think 32GB is enough.",
    "415512": "I am working on both 512x512 and 1024x1024 with 2 GPUs (two trainings in parallel) and just 32GB of RAM. Loading everything from SSD is what makes it possible, but you will need background workers for that to prevent IO bottlenecks",
    "415543": "When using generators, 16 is quite ok. \n\nThe problem might be the loading and augmenting speed.",
    "415549": "Augmentations are indeed a problem. My 8C/16T CPU cannot keep up with feeding my two GPUs :-c",
    "415562": "if you are using keras, then try setting multi processing to true or num_workers to -1 or (maximum cores -1)",
    "415575": "Not using Keras and I would not point out 8C/16T if I was not using them already ;-) Any augmentation that needs resampling into a new image grid (elastic deformation, rotation, scaling, ..) is expensive",
    "415589": "sorry, my bad. i actually wanted to reply to Daniel Moller's post. I just saw yours. And yeah.. its an expensive process for sure. But, don't you  think..we can just do it once...and then use this generated data for all of the further model tuning?",
    "415603": "It might be good depending on how much augmentation is necessary.\n\nIf the problem needs a lot, I don't think it's very good, unless you've got tons of disk. Because strong augmentations rely a lot on random generations and saving augmented images would not bring the same variability.",
    "415605": "Excatly. Doing the augmentations on the fly will allow new augmentation parameters each time an example is chosen thus creating a lot more variability. How much this actually matters I don't know",
    "415626": "I have a feeling this problem will not require heavy augmentation (but I'm just starting).   \n\nBut we can always get 8 flips (faster than distortions/rotations) from any data combining:\n\n    flipMode = random.randint(0,7)\n    if flipMode in [4,5,6,7]:\n        x = np.flip(x,axis1)\n    if flipMode in [1,3,5,7]:\n        x = np.flip(x,axis2)\n    if flipMode in [2,3,6,7]:\n        x = np.swapaxes(x,axis1,axis2)",
    "415630": "You may be right. I did not investigate it at all. Currently, I am doing rotations, deformations, scaling, gamma, gaussian noise, axes transpose, mirror and contrast (each of those with some probability for all samples)",
    "416745": "well first of all you should worry more about your GPU RAM as this limits you batch size and all, (as larger batch size is usually better for models with BatchNorm).\nif you worry about running out RAM, you can always use data generators and online image augmentation. with a decent CPU(which you have) and SSD you won't have too much speed loss imo. don't forget to use multiprocessing for your data generator though.",
    "416783": "Yes, I agree. The GPU RAM seems to be the bottleneck here",
    "417057": "I'm quite lacking in RAM and am able to train ok by loading from SSD. I'm using 512x512 images with a 1070 and 8GB of system RAM.",
    "417232": "So its kinda like a dropout?"
  },
  "source": "meta"
}