{
  "id": 77302,
  "title": "64th place solution",
  "url": "/competitions/human-protein-atlas-image-classification/writeups/robin-smits-64th-place-solution",
  "author_name": "",
  "post_date": "2019-01-11T21:23:03.287Z",
  "votes": 7,
  "comment_count": 10,
  "views": 0,
  "content": "<p>I joined this competition already in the early stages with a personal goal to apply all my current knowledge and to learn a lot of new things. And hopefully achieve a nice score ;-)</p>\n\n<p>Here some things I tried in the first weeks that worked...but not good enough. Self designed CNN's and training from scratch...too slow and quickly overfitting. Fully Convolutional Networks and train from scratch. Didn't seem to overfit but to max out at certain level of Loss and F1. </p>\n\n<p>I then started out trying transfer-learning with Resnet50, VGG16 setup with Keras and Tensorflow backend. This looked really good but with not being able to use 4 seperate channels as input on Keras pretrained models (anybody found a way to do that?) I decided to look beyond Keras.</p>\n\n<p>Seeing that there were various kernels that used Pytorch where it was only a few lines of python code to modify the number of input channels I decided to give Pytorch a try. Also the list of pretrained models for Pytorch is really impressive.</p>\n\n<p>My final solution that I used:</p>\n\n<ul>\n<li>Pretrained Models: BN-Inception and NASNET Large (only used for 1 model..good results but way to heavy for my 1070 Ti)</li>\n<li>6-folds CV</li>\n<li>Multiple runs with variations in batch size, seed, learning rate and pretrained model.</li>\n<li>Epochs 20-25</li>\n<li>Adam optimizer with learning rate either 0.001 or 0.0005. I tried some stepping schedules but didn't notice any significant difference.</li>\n<li>Batch sizes varying between 24 - 36.</li>\n<li>Image size: mostly 512 pixels but also some with 448 pixels.</li>\n<li>Image augmentations: Rotation, Flip and Shear.</li>\n<li>Binary Cross Entropy loss</li>\n<li>Oversampling of the minority classes.</li>\n<li>No Test-Time Augmentation</li>\n<li>Optimal Threshold search on each epoch.</li>\n<li>I generated a full probs file and optimal threshold file after each epoch.</li>\n</ul>\n\n<p>For my final submission I selected multiple good folds from the various runs. From each selected fold I then used between 3 to 6 files with the probabilities and between 8 to 12 files with the optimal thresholds.  I ended up using 53 probability files and 127 optimal threshold files and use simple averaging to generate the final values. Being unsure if I should use a fixed threshold for all classes or use the average for each class instead I did multiple submission for both.</p>\n\n<p>The submission with a fixed threshold of 0.2 was my personal best. However the other submissions with an average threshold for each class are on the private leaderboard almost just as good and multiple ones are even better compared to the fixed threshold used for that same submission. </p>",
  "messages": [
    {
      "id": "454219",
      "postDate": "01/11/2019 09:09:14",
      "content": "<p>I joined this competition already in the early stages with a personal goal to apply all my current knowledge and to learn a lot of new things. And hopefully achieve a nice score ;-)</p>\n\n<p>Here some things I tried in the first weeks that worked...but not good enough. Self designed CNN's and training from scratch...too slow and quickly overfitting. Fully Convolutional Networks and train from scratch. Didn't seem to overfit but to max out at certain level of Loss and F1. </p>\n\n<p>I then started out trying transfer-learning with Resnet50, VGG16 setup with Keras and Tensorflow backend. This looked really good but with not being able to use 4 seperate channels as input on Keras pretrained models (anybody found a way to do that?) I decided to look beyond Keras.</p>\n\n<p>Seeing that there were various kernels that used Pytorch where it was only a few lines of python code to modify the number of input channels I decided to give Pytorch a try. Also the list of pretrained models for Pytorch is really impressive.</p>\n\n<p>My final solution that I used:</p>\n\n<ul>\n<li>Pretrained Models: BN-Inception and NASNET Large (only used for 1 model..good results but way to heavy for my 1070 Ti)</li>\n<li>6-folds CV</li>\n<li>Multiple runs with variations in batch size, seed, learning rate and pretrained model.</li>\n<li>Epochs 20-25</li>\n<li>Adam optimizer with learning rate either 0.001 or 0.0005. I tried some stepping schedules but didn't notice any significant difference.</li>\n<li>Batch sizes varying between 24 - 36.</li>\n<li>Image size: mostly 512 pixels but also some with 448 pixels.</li>\n<li>Image augmentations: Rotation, Flip and Shear.</li>\n<li>Binary Cross Entropy loss</li>\n<li>Oversampling of the minority classes.</li>\n<li>No Test-Time Augmentation</li>\n<li>Optimal Threshold search on each epoch.</li>\n<li>I generated a full probs file and optimal threshold file after each epoch.</li>\n</ul>\n\n<p>For my final submission I selected multiple good folds from the various runs. From each selected fold I then used between 3 to 6 files with the probabilities and between 8 to 12 files with the optimal thresholds.  I ended up using 53 probability files and 127 optimal threshold files and use simple averaging to generate the final values. Being unsure if I should use a fixed threshold for all classes or use the average for each class instead I did multiple submission for both.</p>\n\n<p>The submission with a fixed threshold of 0.2 was my personal best. However the other submissions with an average threshold for each class are on the private leaderboard almost just as good and multiple ones are even better compared to the fixed threshold used for that same submission. </p>",
      "rawMarkdown": "I joined this competition already in the early stages with a personal goal to apply all my current knowledge and to learn a lot of new things. And hopefully achieve a nice score ;-)\n\nHere some things I tried in the first weeks that worked...but not good enough. Self designed CNN's and training from scratch...too slow and quickly overfitting. Fully Convolutional Networks and train from scratch. Didn't seem to overfit but to max out at certain level of Loss and F1. \n\nI then started out trying transfer-learning with Resnet50, VGG16 setup with Keras and Tensorflow backend. This looked really good but with not being able to use 4 seperate channels as input on Keras pretrained models (anybody found a way to do that?) I decided to look beyond Keras.\n\nSeeing that there were various kernels that used Pytorch where it was only a few lines of python code to modify the number of input channels I decided to give Pytorch a try. Also the list of pretrained models for Pytorch is really impressive.\n\nMy final solution that I used:\n\n - Pretrained Models: BN-Inception and NASNET Large (only used for 1 model..good results but way to heavy for my 1070 Ti)\n - 6-folds CV\n - Multiple runs with variations in batch size, seed, learning rate and pretrained model.\n - Epochs 20-25\n - Adam optimizer with learning rate either 0.001 or 0.0005. I tried some stepping schedules but didn't notice any significant difference.\n - Batch sizes varying between 24 - 36.\n - Image size: mostly 512 pixels but also some with 448 pixels.\n - Image augmentations: Rotation, Flip and Shear.\n - Binary Cross Entropy loss\n - Oversampling of the minority classes.\n - No Test-Time Augmentation\n - Optimal Threshold search on each epoch.\n - I generated a full probs file and optimal threshold file after each epoch.\n\nFor my final submission I selected multiple good folds from the various runs. From each selected fold I then used between 3 to 6 files with the probabilities and between 8 to 12 files with the optimal thresholds.  I ended up using 53 probability files and 127 optimal threshold files and use simple averaging to generate the final values. Being unsure if I should use a fixed threshold for all classes or use the average for each class instead I did multiple submission for both.\n\nThe submission with a fixed threshold of 0.2 was my personal best. However the other submissions with an average threshold for each class are on the private leaderboard almost just as good and multiple ones are even better compared to the fixed threshold used for that same submission.",
      "votes": null
    },
    {
      "id": "454269",
      "postDate": "01/11/2019 10:33:49",
      "content": "<p>this code worked ok, but improvement was practically zero:</p>\n\n<pre><code>conv_4to3 = Conv2D(3, 1,\n            kernel_regularizer=l2(0.0005),\n            padding=\"same\",\n            bias_initializer='zeros')(inp_mask)\n\npretrain_model_mask = ResNet50(input_shape = (512,512,3),#SWITCH\n</code></pre>\n\n<p>it takes the 4th layer and \"merge it\" into the other 3. It needs to learn how to do it best and I probably didn't freeze the right layers, etc so It could in theory work.... probably it created a big havoc till it stabilized, explaining the long training</p>",
      "rawMarkdown": "this code worked ok, but improvement was practically zero:\n\n    conv_4to3 = Conv2D(3, 1,\n                kernel_regularizer=l2(0.0005),\n                padding=\"same\",\n                bias_initializer='zeros')(inp_mask)\n\n    pretrain_model_mask = ResNet50(input_shape = (512,512,3),#SWITCH\n\nit takes the 4th layer and \"merge it\" into the other 3. It needs to learn how to do it best and I probably didn't freeze the right layers, etc so It could in theory work.... probably it created a big havoc till it stabilized, explaining the long training",
      "votes": null
    },
    {
      "id": "454621",
      "postDate": "01/11/2019 22:13:38",
      "content": "<p>Interesting..seems like the merge of a 4th channel into the other 3 channels loses any benefit if you don't do it the right way.\nMay'be I will give it a try...thanks for the info anyway.</p>",
      "rawMarkdown": "Interesting..seems like the merge of a 4th channel into the other 3 channels loses any benefit if you don't do it the right way.\nMay'be I will give it a try...thanks for the info anyway.",
      "votes": null
    },
    {
      "id": "454666",
      "postDate": "01/12/2019 00:16:39",
      "content": "<p>Note that most writeups mention that the yellow channel was useless for them, so it might have worked as much as possible. I am interested in your findings, so if you remember, please keep me posted... </p>",
      "rawMarkdown": "Note that most writeups mention that the yellow channel was useless for them, so it might have worked as much as possible. I am interested in your findings, so if you remember, please keep me posted...",
      "votes": null
    },
    {
      "id": "454845",
      "postDate": "01/12/2019 11:01:31",
      "content": "<p>Yes I have red those....interesting because I noticed in my setup a very slight improvement all other things beging equal. It wasn't much but enough to keep me using all channels. I guess that also depends on the complete model setup.\nI think I'll have some time coming week to try 2 runs with my Resnet50 setup... will let you know here when I have my results.</p>",
      "rawMarkdown": "Yes I have red those....interesting because I noticed in my setup a very slight improvement all other things beging equal. It wasn't much but enough to keep me using all channels. I guess that also depends on the complete model setup.\nI think I'll have some time coming week to try 2 runs with my Resnet50 setup... will let you know here when I have my results.",
      "votes": null
    },
    {
      "id": "455549",
      "postDate": "01/14/2019 06:31:59",
      "content": "<p>Thanks. Would be very interesting. I am also running now all sorts of tests with rn50 based on the higher places solutions. One thing that bothers me... you say that batch size is 32 with 512x512... I couldn't fit more then 12 on my titan X! is inception that much smaller then rn50?</p>",
      "rawMarkdown": "Thanks. Would be very interesting. I am also running now all sorts of tests with rn50 based on the higher places solutions. One thing that bothers me... you say that batch size is 32 with 512x512... I couldn't fit more then 12 on my titan X! is inception that much smaller then rn50?",
      "votes": null
    },
    {
      "id": "455782",
      "postDate": "01/14/2019 15:32:20",
      "content": "<p>On my personal Datascience PC I have an NVidia GTX 1070 Ti card. With a Pytorch BN Inception pretrained model I could load a batch size of 26 images with size 512 * 512. A batch size of 28 or higher would give me a Cuda out of memory message consistently. With 448 * 448 pixs I could do batch sizes of 32. I also run the model on an Azure VM with a K80. I could use batch sizes of 44 with the memory available. However a full fold with 20 epochs needed about 32 hours to run...so very impressive batch size there but the speed was horrible.</p>",
      "rawMarkdown": "On my personal Datascience PC I have an NVidia GTX 1070 Ti card. With a Pytorch BN Inception pretrained model I could load a batch size of 26 images with size 512 * 512. A batch size of 28 or higher would give me a Cuda out of memory message consistently. With 448 * 448 pixs I could do batch sizes of 32. I also run the model on an Azure VM with a K80. I could use batch sizes of 44 with the memory available. However a full fold with 20 epochs needed about 32 hours to run...so very impressive batch size there but the speed was horrible.",
      "votes": null
    },
    {
      "id": "462220",
      "postDate": "01/27/2019 22:30:36",
      "content": "<p>Hey Moshel, So I did some experiments where I lined up a Resnet50 model to be as much as possible the same as the BN-Inception approach. Training 2 epochs head only, then 20 epochs full model. The outcomes vary a little bit each run..but I get about 0.050 - 0.070 difference between Resnet50 and BN Inception. I'am not sure if that all comes down to the differences in the 2 networks but it is quite a difference. I presume that further optimizing and tuning parameters for the Resnet50 would make the difference much smaller.</p>\n\n<p>On my 1070 Ti I could fit a max batch size of 8 for 512x512....which makes training of a Resnet50 also consume more time.</p>",
      "rawMarkdown": "Hey Moshel, So I did some experiments where I lined up a Resnet50 model to be as much as possible the same as the BN-Inception approach. Training 2 epochs head only, then 20 epochs full model. The outcomes vary a little bit each run..but I get about 0.050 - 0.070 difference between Resnet50 and BN Inception. I'am not sure if that all comes down to the differences in the 2 networks but it is quite a difference. I presume that further optimizing and tuning parameters for the Resnet50 would make the difference much smaller.\n\nOn my 1070 Ti I could fit a max batch size of 8 for 512x512....which makes training of a Resnet50 also consume more time.",
      "votes": null
    },
    {
      "id": "462242",
      "postDate": "01/27/2019 23:58:23",
      "content": "<p>I also found that rn50 performed better. As for the batch size, my experiments showed very little difference in time (up to 10% iirc) in training time per epoch with larger batches, which surprised me. Might be some other bottleneck. However, with large batches i got smoother convergence so needed less epochs. These are all observations and ymmv. If you want larger batches for the gradient, you can look into additive gradient or something which delays the backprop till x batches where processed. It's next on my \"to do\" list, as i had real problems with training rn50 on 1024*1024 even with 12gb</p>",
      "rawMarkdown": "I also found that rn50 performed better. As for the batch size, my experiments showed very little difference in time (up to 10% iirc) in training time per epoch with larger batches, which surprised me. Might be some other bottleneck. However, with large batches i got smoother convergence so needed less epochs. These are all observations and ymmv. If you want larger batches for the gradient, you can look into additive gradient or something which delays the backprop till x batches where processed. It's next on my \"to do\" list, as i had real problems with training rn50 on 1024*1024 even with 12gb",
      "votes": null
    },
    {
      "id": "462244",
      "postDate": "01/28/2019 00:01:30",
      "content": "<p>Oh wait! Re reading your reply i realized bn inception worked better for you.\nDeep learning is a dark art. For me rn50 always worked better (3ch)</p>",
      "rawMarkdown": "Oh wait! Re reading your reply i realized bn inception worked better for you.\nDeep learning is a dark art. For me rn50 always worked better (3ch)",
      "votes": null
    },
    {
      "id": "462377",
      "postDate": "01/28/2019 07:57:47",
      "content": "<p>Yes could be that with more epochs and different parameters rn50 would be better...however time needed to train is also an important factor.\nWith the time available in a day I could sometimes try 3 different things on a day...with more epochs that would be a lot less.\nAnd yes deep learning is a dark art.. I like that one ;-)</p>",
      "rawMarkdown": "Yes could be that with more epochs and different parameters rn50 would be better...however time needed to train is also an important factor.\nWith the time available in a day I could sometimes try 3 different things on a day...with more epochs that would be a lot less.\nAnd yes deep learning is a dark art.. I like that one ;-)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 454269,
      "author_name": "moshel",
      "author_url": "",
      "post_date": "01/11/2019 10:33:49",
      "content": "<p>this code worked ok, but improvement was practically zero:</p>\n\n<pre><code>conv_4to3 = Conv2D(3, 1,\n            kernel_regularizer=l2(0.0005),\n            padding=\"same\",\n            bias_initializer='zeros')(inp_mask)\n\npretrain_model_mask = ResNet50(input_shape = (512,512,3),#SWITCH\n</code></pre>\n\n<p>it takes the 4th layer and \"merge it\" into the other 3. It needs to learn how to do it best and I probably didn't freeze the right layers, etc so It could in theory work.... probably it created a big havoc till it stabilized, explaining the long training</p>",
      "votes": null,
      "replies": [
        {
          "id": 454621,
          "author_name": "rsmits",
          "author_url": "",
          "post_date": "01/11/2019 22:13:38",
          "content": "<p>Interesting..seems like the merge of a 4th channel into the other 3 channels loses any benefit if you don't do it the right way.\nMay'be I will give it a try...thanks for the info anyway.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 454666,
          "author_name": "moshel",
          "author_url": "",
          "post_date": "01/12/2019 00:16:39",
          "content": "<p>Note that most writeups mention that the yellow channel was useless for them, so it might have worked as much as possible. I am interested in your findings, so if you remember, please keep me posted... </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 454845,
          "author_name": "rsmits",
          "author_url": "",
          "post_date": "01/12/2019 11:01:31",
          "content": "<p>Yes I have red those....interesting because I noticed in my setup a very slight improvement all other things beging equal. It wasn't much but enough to keep me using all channels. I guess that also depends on the complete model setup.\nI think I'll have some time coming week to try 2 runs with my Resnet50 setup... will let you know here when I have my results.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 455549,
          "author_name": "moshel",
          "author_url": "",
          "post_date": "01/14/2019 06:31:59",
          "content": "<p>Thanks. Would be very interesting. I am also running now all sorts of tests with rn50 based on the higher places solutions. One thing that bothers me... you say that batch size is 32 with 512x512... I couldn't fit more then 12 on my titan X! is inception that much smaller then rn50?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 455782,
          "author_name": "rsmits",
          "author_url": "",
          "post_date": "01/14/2019 15:32:20",
          "content": "<p>On my personal Datascience PC I have an NVidia GTX 1070 Ti card. With a Pytorch BN Inception pretrained model I could load a batch size of 26 images with size 512 * 512. A batch size of 28 or higher would give me a Cuda out of memory message consistently. With 448 * 448 pixs I could do batch sizes of 32. I also run the model on an Azure VM with a K80. I could use batch sizes of 44 with the memory available. However a full fold with 20 epochs needed about 32 hours to run...so very impressive batch size there but the speed was horrible.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 462220,
          "author_name": "rsmits",
          "author_url": "",
          "post_date": "01/27/2019 22:30:36",
          "content": "<p>Hey Moshel, So I did some experiments where I lined up a Resnet50 model to be as much as possible the same as the BN-Inception approach. Training 2 epochs head only, then 20 epochs full model. The outcomes vary a little bit each run..but I get about 0.050 - 0.070 difference between Resnet50 and BN Inception. I'am not sure if that all comes down to the differences in the 2 networks but it is quite a difference. I presume that further optimizing and tuning parameters for the Resnet50 would make the difference much smaller.</p>\n\n<p>On my 1070 Ti I could fit a max batch size of 8 for 512x512....which makes training of a Resnet50 also consume more time.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 462242,
          "author_name": "moshel",
          "author_url": "",
          "post_date": "01/27/2019 23:58:23",
          "content": "<p>I also found that rn50 performed better. As for the batch size, my experiments showed very little difference in time (up to 10% iirc) in training time per epoch with larger batches, which surprised me. Might be some other bottleneck. However, with large batches i got smoother convergence so needed less epochs. These are all observations and ymmv. If you want larger batches for the gradient, you can look into additive gradient or something which delays the backprop till x batches where processed. It's next on my \"to do\" list, as i had real problems with training rn50 on 1024*1024 even with 12gb</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 462244,
          "author_name": "moshel",
          "author_url": "",
          "post_date": "01/28/2019 00:01:30",
          "content": "<p>Oh wait! Re reading your reply i realized bn inception worked better for you.\nDeep learning is a dark art. For me rn50 always worked better (3ch)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 462377,
          "author_name": "rsmits",
          "author_url": "",
          "post_date": "01/28/2019 07:57:47",
          "content": "<p>Yes could be that with more epochs and different parameters rn50 would be better...however time needed to train is also an important factor.\nWith the time available in a day I could sometimes try 3 different things on a day...with more epochs that would be a lot less.\nAnd yes deep learning is a dark art.. I like that one ;-)</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "454219": "I joined this competition already in the early stages with a personal goal to apply all my current knowledge and to learn a lot of new things. And hopefully achieve a nice score ;-)\n\nHere some things I tried in the first weeks that worked...but not good enough. Self designed CNN's and training from scratch...too slow and quickly overfitting. Fully Convolutional Networks and train from scratch. Didn't seem to overfit but to max out at certain level of Loss and F1. \n\nI then started out trying transfer-learning with Resnet50, VGG16 setup with Keras and Tensorflow backend. This looked really good but with not being able to use 4 seperate channels as input on Keras pretrained models (anybody found a way to do that?) I decided to look beyond Keras.\n\nSeeing that there were various kernels that used Pytorch where it was only a few lines of python code to modify the number of input channels I decided to give Pytorch a try. Also the list of pretrained models for Pytorch is really impressive.\n\nMy final solution that I used:\n\n - Pretrained Models: BN-Inception and NASNET Large (only used for 1 model..good results but way to heavy for my 1070 Ti)\n - 6-folds CV\n - Multiple runs with variations in batch size, seed, learning rate and pretrained model.\n - Epochs 20-25\n - Adam optimizer with learning rate either 0.001 or 0.0005. I tried some stepping schedules but didn't notice any significant difference.\n - Batch sizes varying between 24 - 36.\n - Image size: mostly 512 pixels but also some with 448 pixels.\n - Image augmentations: Rotation, Flip and Shear.\n - Binary Cross Entropy loss\n - Oversampling of the minority classes.\n - No Test-Time Augmentation\n - Optimal Threshold search on each epoch.\n - I generated a full probs file and optimal threshold file after each epoch.\n\nFor my final submission I selected multiple good folds from the various runs. From each selected fold I then used between 3 to 6 files with the probabilities and between 8 to 12 files with the optimal thresholds.  I ended up using 53 probability files and 127 optimal threshold files and use simple averaging to generate the final values. Being unsure if I should use a fixed threshold for all classes or use the average for each class instead I did multiple submission for both.\n\nThe submission with a fixed threshold of 0.2 was my personal best. However the other submissions with an average threshold for each class are on the private leaderboard almost just as good and multiple ones are even better compared to the fixed threshold used for that same submission.",
    "454269": "this code worked ok, but improvement was practically zero:\n\n    conv_4to3 = Conv2D(3, 1,\n                kernel_regularizer=l2(0.0005),\n                padding=\"same\",\n                bias_initializer='zeros')(inp_mask)\n\n    pretrain_model_mask = ResNet50(input_shape = (512,512,3),#SWITCH\n\nit takes the 4th layer and \"merge it\" into the other 3. It needs to learn how to do it best and I probably didn't freeze the right layers, etc so It could in theory work.... probably it created a big havoc till it stabilized, explaining the long training",
    "454621": "Interesting..seems like the merge of a 4th channel into the other 3 channels loses any benefit if you don't do it the right way.\nMay'be I will give it a try...thanks for the info anyway.",
    "454666": "Note that most writeups mention that the yellow channel was useless for them, so it might have worked as much as possible. I am interested in your findings, so if you remember, please keep me posted...",
    "454845": "Yes I have red those....interesting because I noticed in my setup a very slight improvement all other things beging equal. It wasn't much but enough to keep me using all channels. I guess that also depends on the complete model setup.\nI think I'll have some time coming week to try 2 runs with my Resnet50 setup... will let you know here when I have my results.",
    "455549": "Thanks. Would be very interesting. I am also running now all sorts of tests with rn50 based on the higher places solutions. One thing that bothers me... you say that batch size is 32 with 512x512... I couldn't fit more then 12 on my titan X! is inception that much smaller then rn50?",
    "455782": "On my personal Datascience PC I have an NVidia GTX 1070 Ti card. With a Pytorch BN Inception pretrained model I could load a batch size of 26 images with size 512 * 512. A batch size of 28 or higher would give me a Cuda out of memory message consistently. With 448 * 448 pixs I could do batch sizes of 32. I also run the model on an Azure VM with a K80. I could use batch sizes of 44 with the memory available. However a full fold with 20 epochs needed about 32 hours to run...so very impressive batch size there but the speed was horrible.",
    "462220": "Hey Moshel, So I did some experiments where I lined up a Resnet50 model to be as much as possible the same as the BN-Inception approach. Training 2 epochs head only, then 20 epochs full model. The outcomes vary a little bit each run..but I get about 0.050 - 0.070 difference between Resnet50 and BN Inception. I'am not sure if that all comes down to the differences in the 2 networks but it is quite a difference. I presume that further optimizing and tuning parameters for the Resnet50 would make the difference much smaller.\n\nOn my 1070 Ti I could fit a max batch size of 8 for 512x512....which makes training of a Resnet50 also consume more time.",
    "462242": "I also found that rn50 performed better. As for the batch size, my experiments showed very little difference in time (up to 10% iirc) in training time per epoch with larger batches, which surprised me. Might be some other bottleneck. However, with large batches i got smoother convergence so needed less epochs. These are all observations and ymmv. If you want larger batches for the gradient, you can look into additive gradient or something which delays the backprop till x batches where processed. It's next on my \"to do\" list, as i had real problems with training rn50 on 1024*1024 even with 12gb",
    "462244": "Oh wait! Re reading your reply i realized bn inception worked better for you.\nDeep learning is a dark art. For me rn50 always worked better (3ch)",
    "462377": "Yes could be that with more epochs and different parameters rn50 would be better...however time needed to train is also an important factor.\nWith the time available in a day I could sometimes try 3 different things on a day...with more epochs that would be a lot less.\nAnd yes deep learning is a dark art.. I like that one ;-)"
  },
  "source": "meta"
}