{
  "id": 38125,
  "title": "how to move to 0.997",
  "url": "/competitions/carvana-image-masking-challenge/discussion/38125",
  "author_name": "",
  "post_date": "2017-08-15T09:41:28.975770700Z",
  "votes": 81,
  "comment_count": 61,
  "views": 0,
  "content": "<p>UPDATED! the software and model for producing this results can be found at:\n<a href=\"https://www.kaggle.com/c/carvana-image-masking-challenge/discussion/37208\">https://www.kaggle.com/c/carvana-image-masking-challenge/discussion/37208</a></p>\n\n<p>.</p>\n\n<p>Here are some tips on how to move to 0.997.</p>\n\n<p>.</p>\n\n<p><strong>(1) Which deep learning framework to use?</strong></p>\n\n<p>Both unet imlementation in Keras/tensorflow and pytorch are good starters:</p>\n\n<p><a href=\"https://www.kaggle.com/c/carvana-image-masking-challenge/discussion/37523\">https://www.kaggle.com/c/carvana-image-masking-challenge/discussion/37523</a></p>\n\n<p><a href=\"https://www.kaggle.com/c/carvana-image-masking-challenge/discussion/37208\">https://www.kaggle.com/c/carvana-image-masking-challenge/discussion/37208</a></p>\n\n<p>.</p>\n\n<p><strong>(2) Which CNN model to use?</strong></p>\n\n<p>Unet is sufficient. However, do experiment with different depths, how many convolution filters to use per scale, etc? If you do not get the parameters right, you do not get the best performance.</p>\n\n<p>.</p>\n\n<p><strong>(3) Which image resolution to use?</strong></p>\n\n<p>I would suggest 1024x1024. (this is what i use for my 0.997 solution)</p>\n\n<p>.</p>\n\n<p><strong>(4) Generalization and validation/training split</strong></p>\n\n<p>Any good split is ok, because the train and test dataset are very similar. But be careful about the validation error during training. There are some ground truth error, please see:</p>\n\n<p><a href=\"https://www.kaggle.com/c/carvana-image-masking-challenge/discussion/37229\">https://www.kaggle.com/c/carvana-image-masking-challenge/discussion/37229</a></p>\n\n<p>Do examine the validation error (e.g. by visual inspection) and make sure high error are not due to ground truth error. If you can get training/validation error of near 0.997, you should get 0.997 at the LB.</p>\n\n<p>Note that the LB page may not show the correct ranking. A better way to judge if your new submission is better then previous ones is to use the 'sort by public score' function on your submission page, please see:</p>\n\n<p><a href=\"https://www.kaggle.com/c/carvana-image-masking-challenge/discussion/37137\">https://www.kaggle.com/c/carvana-image-masking-challenge/discussion/37137</a></p>\n\n<p>.</p>\n\n<p><strong>(5) Training hyper parameters</strong></p>\n\n<p>Try different learning rates. I use about 45 epoch. Try different batch size. I use 3 images per batch, and accumulate 5 batches before updating the gradient. i.e. my effective batch size is 3x5 = 15. I just use SDG with momentum with manual rate scheduling.</p>\n\n<p>.</p>\n\n<p><strong>(6) Ensemble and test-time augmentation</strong></p>\n\n<p>They may improvement results. But I did not use them in my 0.997 solution. </p>\n\n<p>.</p>\n\n<p><strong>(7) Pre or post processing?</strong></p>\n\n<p>They may improvement results. But I did not them. Just a simple CNN prediction can give 0.997. I use cv2.INTER_LINEAR for downsize and upsize.</p>\n\n<p>.</p>\n\n<p><strong>(6) Train sample augmentation</strong></p>\n\n<p>In my experiments, some train augmentation is required. However, do be careful. Too much augmentation actually reduces accuracy.</p>\n\n<p>.</p>\n\n<p><strong>(7) Loss function for back propagation?</strong></p>\n\n<p>Use binary cross entropy and dice loss. Weigh pixels near boundary. </p>\n\n<p>.</p>\n\n<p><strong>(8) Use pretrained model?</strong></p>\n\n<p>I do not use. I train from sratch.</p>\n\n<p>.</p>\n\n<p><strong>(9) Any other tricks?</strong></p>\n\n<p>No. Just tune your learning carefully. Do proper experiments and record your results. Make your work process efficient. Due to large image image size and huge number of test images (100064), it does take some time to generate results.  </p>\n\n<p>my machine: pascal titan-X gpu. Train time per epoch = 15 min (11.5 hours in total training).  Time to make submission csv file = 1.5 hr to make prediction,  35 min to encode rle and save csv.</p>",
  "messages": [
    {
      "id": "213810",
      "postDate": "08/15/2017 09:41:28",
      "content": "<p>UPDATED! the software and model for producing this results can be found at:\n<a href=\"https://www.kaggle.com/c/carvana-image-masking-challenge/discussion/37208\">https://www.kaggle.com/c/carvana-image-masking-challenge/discussion/37208</a></p>\n\n<p>.</p>\n\n<p>Here are some tips on how to move to 0.997.</p>\n\n<p>.</p>\n\n<p><strong>(1) Which deep learning framework to use?</strong></p>\n\n<p>Both unet imlementation in Keras/tensorflow and pytorch are good starters:</p>\n\n<p><a href=\"https://www.kaggle.com/c/carvana-image-masking-challenge/discussion/37523\">https://www.kaggle.com/c/carvana-image-masking-challenge/discussion/37523</a></p>\n\n<p><a href=\"https://www.kaggle.com/c/carvana-image-masking-challenge/discussion/37208\">https://www.kaggle.com/c/carvana-image-masking-challenge/discussion/37208</a></p>\n\n<p>.</p>\n\n<p><strong>(2) Which CNN model to use?</strong></p>\n\n<p>Unet is sufficient. However, do experiment with different depths, how many convolution filters to use per scale, etc? If you do not get the parameters right, you do not get the best performance.</p>\n\n<p>.</p>\n\n<p><strong>(3) Which image resolution to use?</strong></p>\n\n<p>I would suggest 1024x1024. (this is what i use for my 0.997 solution)</p>\n\n<p>.</p>\n\n<p><strong>(4) Generalization and validation/training split</strong></p>\n\n<p>Any good split is ok, because the train and test dataset are very similar. But be careful about the validation error during training. There are some ground truth error, please see:</p>\n\n<p><a href=\"https://www.kaggle.com/c/carvana-image-masking-challenge/discussion/37229\">https://www.kaggle.com/c/carvana-image-masking-challenge/discussion/37229</a></p>\n\n<p>Do examine the validation error (e.g. by visual inspection) and make sure high error are not due to ground truth error. If you can get training/validation error of near 0.997, you should get 0.997 at the LB.</p>\n\n<p>Note that the LB page may not show the correct ranking. A better way to judge if your new submission is better then previous ones is to use the 'sort by public score' function on your submission page, please see:</p>\n\n<p><a href=\"https://www.kaggle.com/c/carvana-image-masking-challenge/discussion/37137\">https://www.kaggle.com/c/carvana-image-masking-challenge/discussion/37137</a></p>\n\n<p>.</p>\n\n<p><strong>(5) Training hyper parameters</strong></p>\n\n<p>Try different learning rates. I use about 45 epoch. Try different batch size. I use 3 images per batch, and accumulate 5 batches before updating the gradient. i.e. my effective batch size is 3x5 = 15. I just use SDG with momentum with manual rate scheduling.</p>\n\n<p>.</p>\n\n<p><strong>(6) Ensemble and test-time augmentation</strong></p>\n\n<p>They may improvement results. But I did not use them in my 0.997 solution. </p>\n\n<p>.</p>\n\n<p><strong>(7) Pre or post processing?</strong></p>\n\n<p>They may improvement results. But I did not them. Just a simple CNN prediction can give 0.997. I use cv2.INTER_LINEAR for downsize and upsize.</p>\n\n<p>.</p>\n\n<p><strong>(6) Train sample augmentation</strong></p>\n\n<p>In my experiments, some train augmentation is required. However, do be careful. Too much augmentation actually reduces accuracy.</p>\n\n<p>.</p>\n\n<p><strong>(7) Loss function for back propagation?</strong></p>\n\n<p>Use binary cross entropy and dice loss. Weigh pixels near boundary. </p>\n\n<p>.</p>\n\n<p><strong>(8) Use pretrained model?</strong></p>\n\n<p>I do not use. I train from sratch.</p>\n\n<p>.</p>\n\n<p><strong>(9) Any other tricks?</strong></p>\n\n<p>No. Just tune your learning carefully. Do proper experiments and record your results. Make your work process efficient. Due to large image image size and huge number of test images (100064), it does take some time to generate results.  </p>\n\n<p>my machine: pascal titan-X gpu. Train time per epoch = 15 min (11.5 hours in total training).  Time to make submission csv file = 1.5 hr to make prediction,  35 min to encode rle and save csv.</p>",
      "rawMarkdown": "UPDATED! the software and model for producing this results can be found at:\nhttps://www.kaggle.com/c/carvana-image-masking-challenge/discussion/37208\n\n.\n\nHere are some tips on how to move to 0.997.\n\n.\n\n\n**(1) Which deep learning framework to use?**\n\nBoth unet imlementation in Keras/tensorflow and pytorch are good starters:\n\nhttps://www.kaggle.com/c/carvana-image-masking-challenge/discussion/37523\n\nhttps://www.kaggle.com/c/carvana-image-masking-challenge/discussion/37208\n\n.\n\n**(2) Which CNN model to use?**\n\nUnet is sufficient. However, do experiment with different depths, how many convolution filters to use per scale, etc? If you do not get the parameters right, you do not get the best performance.\n\n.\n\n **(3) Which image resolution to use?**\n\nI would suggest 1024x1024. (this is what i use for my 0.997 solution)\n\n.\n\n **(4) Generalization and validation/training split**\n\nAny good split is ok, because the train and test dataset are very similar. But be careful about the validation error during training. There are some ground truth error, please see:\n\nhttps://www.kaggle.com/c/carvana-image-masking-challenge/discussion/37229\n\nDo examine the validation error (e.g. by visual inspection) and make sure high error are not due to ground truth error. If you can get training/validation error of near 0.997, you should get 0.997 at the LB.\n\nNote that the LB page may not show the correct ranking. A better way to judge if your new submission is better then previous ones is to use the 'sort by public score' function on your submission page, please see:\n\nhttps://www.kaggle.com/c/carvana-image-masking-challenge/discussion/37137\n\n.\n\n\n\n\n**(5) Training hyper parameters**\n\nTry different learning rates. I use about 45 epoch. Try different batch size. I use 3 images per batch, and accumulate 5 batches before updating the gradient. i.e. my effective batch size is 3x5 = 15. I just use SDG with momentum with manual rate scheduling.\n\n.\n\n**(6) Ensemble and test-time augmentation**\n\nThey may improvement results. But I did not use them in my 0.997 solution. \n\n.\n\n**(7) Pre or post processing?**\n\nThey may improvement results. But I did not them. Just a simple CNN prediction can give 0.997. I use cv2.INTER_LINEAR for downsize and upsize.\n\n.\n\n**(6) Train sample augmentation**\n\nIn my experiments, some train augmentation is required. However, do be careful. Too much augmentation actually reduces accuracy.\n\n.\n\n **(7) Loss function for back propagation?**\n\nUse binary cross entropy and dice loss. Weigh pixels near boundary. \n\n.\n\n **(8) Use pretrained model?**\n\nI do not use. I train from sratch.\n\n.\n\n **(9) Any other tricks?**\n\nNo. Just tune your learning carefully. Do proper experiments and record your results. Make your work process efficient. Due to large image image size and huge number of test images (100064), it does take some time to generate results.  \n\nmy machine: pascal titan-X gpu. Train time per epoch = 15 min (11.5 hours in total training).  Time to make submission csv file = 1.5 hr to make prediction,  35 min to encode rle and save csv.",
      "votes": null
    },
    {
      "id": "213813",
      "postDate": "08/15/2017 10:01:13",
      "content": "<p>Good summary.\nOne question about 3): 1024x1024 rescaled to 1280x1918 or sliced into 1024x1024 pieces?</p>",
      "rawMarkdown": "Good summary.\nOne question about 3): 1024x1024 rescaled to 1280x1918 or sliced into 1024x1024 pieces?",
      "votes": null
    },
    {
      "id": "213814",
      "postDate": "08/15/2017 10:02:31",
      "content": "<p>1024x1024 rescaled to 1280x1918. Only one-step resize, no cropping.</p>",
      "rawMarkdown": "1024x1024 rescaled to 1280x1918. Only one-step resize, no cropping.",
      "votes": null
    },
    {
      "id": "213815",
      "postDate": "08/15/2017 10:10:28",
      "content": "<p>May also be worth testing variations of the Unet architecture, like replacing the VGG-like architecture with residual or inception blocks.  Not necessary for 0.997 but could be for 0.998 :)</p>",
      "rawMarkdown": "May also be worth testing variations of the Unet architecture, like replacing the VGG-like architecture with residual or inception blocks.  Not necessary for 0.997 but could be for 0.998 :)",
      "votes": null
    },
    {
      "id": "213842",
      "postDate": "08/15/2017 11:30:46",
      "content": "<p>0.996 ers are looking at you</p>",
      "rawMarkdown": "0.996 ers are looking at you",
      "votes": null
    },
    {
      "id": "213857",
      "postDate": "08/15/2017 12:05:27",
      "content": "<p>What lr-scheduler do you use for optimization?</p>",
      "rawMarkdown": "What lr-scheduler do you use for optimization?",
      "votes": null
    },
    {
      "id": "213860",
      "postDate": "08/15/2017 12:08:50",
      "content": "<p>note that this has to depend on your network, data augment, data batch size, loss,  etc. What works for me may not work for you.</p>\n\n<p>For me, i use conv-bn-relu. for epoch 0 to 40: 0.01,  40 to 45: 0.005, 45 to 47: 0.001. My training loss is as attached:</p>\n\n<pre><code>   LR=Step Learning Rates\n   rates=[' 0.0100', ' 0.0050', ' 0.0010', '-1.0000', '-1.0000']\n   steps=['      0', '     35', '     40', '     42', '     44']\n\n epoch    iter      rate   | valid_loss/acc | train_loss/acc | batch_loss/acc ...             \n ---------------------------------------------------------------------------------------------\n   1.0    1440    0.0100   | 0.0670  0.9900 | 0.0807  0.9883 | 0.0897  0.9885  |  14.8 min    \n   2.0    1440    0.0100   | 0.0470  0.9937 | 0.0566  0.9922 | 0.0572  0.9918  |  14.7 min    \n   3.0    1440    0.0100   | 0.0380  0.9948 | 0.0496  0.9929 | 0.0369  0.9949  |  14.7 min    \n   4.0    1440    0.0100   | 0.0400  0.9944 | 0.0414  0.9944 | 0.0504  0.9934  |  14.7 min    \n   5.0    1440    0.0100   | 0.0317  0.9956 | 0.0400  0.9946 | 0.0821  0.9896  |  14.7 min    \n   6.0    1440    0.0100   | 0.0310  0.9956 | 0.0401  0.9947 | 0.0391  0.9950  |  14.7 min    \n   7.0    1440    0.0100   | 0.0314  0.9955 | 0.0397  0.9946 | 0.0364  0.9954  |  14.7 min    \n   8.0    1440    0.0100   | 0.0286  0.9959 | 0.0344  0.9952 | 0.0267  0.9959  |  14.7 min    \n   9.0    1440    0.0100   | 0.0286  0.9959 | 0.0375  0.9951 | 0.0296  0.9953  |  14.7 min    \n  10.0    1440    0.0100   | 0.0284  0.9959 | 0.0321  0.9956 | 0.0435  0.9948  |  14.9 min    \n  11.0    1440    0.0100   | 0.0280  0.9960 | 0.0325  0.9954 | 0.0253  0.9967  |  15.1 min    \n  12.0    1440    0.0100   | 0.0270  0.9961 | 0.0349  0.9950 | 0.0352  0.9945  |  14.9 min    \n  13.0    1440    0.0100   | 0.0266  0.9962 | 0.0335  0.9953 | 0.0296  0.9952  |  14.9 min    \n  14.0    1440    0.0100   | 0.0263  0.9962 | 0.0324  0.9955 | 0.0287  0.9962  |  15.1 min    \n  15.0    1440    0.0100   | 0.0272  0.9961 | 0.0292  0.9959 | 0.0332  0.9950  |  14.9 min    \n  16.0    1440    0.0100   | 0.0269  0.9961 | 0.0288  0.9960 | 0.0228  0.9966  |  15.0 min    \n  17.0    1440    0.0100   | 0.0263  0.9962 | 0.0332  0.9954 | 0.0256  0.9964  |  14.8 min    \n  18.0    1440    0.0100   | 0.0254  0.9963 | 0.0283  0.9962 | 0.0498  0.9916  |  14.9 min    \n  19.0    1440    0.0100   | 0.0249  0.9964 | 0.0285  0.9960 | 0.0248  0.9963  |  14.8 min    \n  20.0    1440    0.0100   | 0.0247  0.9964 | 0.0277  0.9961 | 0.0314  0.9952  |  14.9 min    \n  21.0    1440    0.0100   | 0.0248  0.9964 | 0.0274  0.9961 | 0.0240  0.9961  |  15.0 min    \n  22.0    1440    0.0100   | 0.0251  0.9963 | 0.0252  0.9965 | 0.0214  0.9969  |  14.9 min    \n  23.0    1440    0.0100   | 0.0252  0.9963 | 0.0284  0.9961 | 0.0343  0.9952  |  14.9 min    \n  24.0    1440    0.0100   | 0.0243  0.9965 | 0.0261  0.9963 | 0.0246  0.9966  |  14.8 min    \n  25.0    1440    0.0100   | 0.0248  0.9964 | 0.0279  0.9959 | 0.0197  0.9967  |  14.6 min    \n  26.0    1440    0.0100   | 0.0244  0.9964 | 0.0269  0.9961 | 0.0366  0.9952  |  14.6 min    \n  27.0    1440    0.0100   | 0.0240  0.9965 | 0.0275  0.9962 | 0.0273  0.9960  |  14.5 min    \n  28.0    1440    0.0100   | 0.0245  0.9964 | 0.0251  0.9964 | 0.0340  0.9953  |  14.5 min    \n  29.0    1440    0.0100   | 0.0244  0.9965 | 0.0281  0.9961 | 0.0220  0.9967  |  14.5 min    \n  30.0    1440    0.0100   | 0.0241  0.9965 | 0.0298  0.9958 | 0.0364  0.9954  |  14.5 min    \n  31.0    1440    0.0100   | 0.0235  0.9966 | 0.0240  0.9965 | 0.0273  0.9962  |  14.5 min    \n  32.0    1440    0.0100   | 0.0237  0.9965 | 0.0261  0.9963 | 0.0257  0.9962  |  14.5 min    \n  33.0    1440    0.0100   | 0.0235  0.9966 | 0.0240  0.9965 | 0.0191  0.9972  |  14.5 min    \n  34.0    1440    0.0100   | 0.0234  0.9966 | 0.0237  0.9965 | 0.0332  0.9955  |  14.5 min    \n\n  make some change to dataset : reduce augmentation\n  ... stop and resume training ....\n\n  LR=Step Learning Rates\n rates=[' 0.0100', ' 0.0050', ' 0.0010', '-1.0000', '-1.0000']\n steps=['      0', '     40', '     45', '     47', '     44']\n\n  epoch    iter      rate   | valid_loss/acc | train_loss/acc | batch_loss/acc ... \n  --------------------------------------------------------------------------------------------------\n   34.0    1440    0.0100   | 0.0231  0.9966 | 0.0230  0.9969 | 0.0262  0.9966  |  14.6 min \n   35.0    1440    0.0100   | 0.0234  0.9966 | 0.0216  0.9969 | 0.0247  0.9963  |  14.8 min \n   36.0    1440    0.0100   | 0.0231  0.9966 | 0.0214  0.9970 | 0.0220  0.9967  |  14.8 min \n   37.0    1440    0.0100   | 0.0238  0.9965 | 0.0216  0.9971 | 0.0281  0.9963  |  14.8 min \n   38.0    1440    0.0100   | 0.0228  0.9967 | 0.0237  0.9967 | 0.0247  0.9965  |  14.8 min \n   39.0    1440    0.0100   | 0.0226  0.9967 | 0.0230  0.9969 | 0.0188  0.9973  |  15.0 min \n   40.0    1440    0.0100   | 0.0230  0.9966 | 0.0224  0.9969 | 0.0231  0.9971  |  15.4 min \n   41.0    1440    0.0050   | 0.0224  0.9967 | 0.0224  0.9970 | 0.0180  0.9975  |  15.4 min \n   42.0    1440    0.0050   | 0.0225  0.9967 | 0.0218  0.9970 | 0.0216  0.9968  |  15.4 min \n   43.0    1440    0.0050   | 0.0224  0.9967 | 0.0202  0.9973 | 0.0214  0.9972  |  15.4 min \n   44.0    1440    0.0050   | 0.0224  0.9967 | 0.0196  0.9972 | 0.0172  0.9977  |  14.8 min \n</code></pre>",
      "rawMarkdown": "note that this has to depend on your network, data augment, data batch size, loss,  etc. What works for me may not work for you.\n\nFor me, i use conv-bn-relu. for epoch 0 to 40: 0.01,  40 to 45: 0.005, 45 to 47: 0.001. My training loss is as attached:\n\n  \n       LR=Step Learning Rates\n       rates=[' 0.0100', ' 0.0050', ' 0.0010', '-1.0000', '-1.0000']\n       steps=['      0', '     35', '     40', '     42', '     44']\n\n     epoch    iter      rate   | valid_loss/acc | train_loss/acc | batch_loss/acc ...             \n     ---------------------------------------------------------------------------------------------\n       1.0    1440    0.0100   | 0.0670  0.9900 | 0.0807  0.9883 | 0.0897  0.9885  |  14.8 min    \n       2.0    1440    0.0100   | 0.0470  0.9937 | 0.0566  0.9922 | 0.0572  0.9918  |  14.7 min    \n       3.0    1440    0.0100   | 0.0380  0.9948 | 0.0496  0.9929 | 0.0369  0.9949  |  14.7 min    \n       4.0    1440    0.0100   | 0.0400  0.9944 | 0.0414  0.9944 | 0.0504  0.9934  |  14.7 min    \n       5.0    1440    0.0100   | 0.0317  0.9956 | 0.0400  0.9946 | 0.0821  0.9896  |  14.7 min    \n       6.0    1440    0.0100   | 0.0310  0.9956 | 0.0401  0.9947 | 0.0391  0.9950  |  14.7 min    \n       7.0    1440    0.0100   | 0.0314  0.9955 | 0.0397  0.9946 | 0.0364  0.9954  |  14.7 min    \n       8.0    1440    0.0100   | 0.0286  0.9959 | 0.0344  0.9952 | 0.0267  0.9959  |  14.7 min    \n       9.0    1440    0.0100   | 0.0286  0.9959 | 0.0375  0.9951 | 0.0296  0.9953  |  14.7 min    \n      10.0    1440    0.0100   | 0.0284  0.9959 | 0.0321  0.9956 | 0.0435  0.9948  |  14.9 min    \n      11.0    1440    0.0100   | 0.0280  0.9960 | 0.0325  0.9954 | 0.0253  0.9967  |  15.1 min    \n      12.0    1440    0.0100   | 0.0270  0.9961 | 0.0349  0.9950 | 0.0352  0.9945  |  14.9 min    \n      13.0    1440    0.0100   | 0.0266  0.9962 | 0.0335  0.9953 | 0.0296  0.9952  |  14.9 min    \n      14.0    1440    0.0100   | 0.0263  0.9962 | 0.0324  0.9955 | 0.0287  0.9962  |  15.1 min    \n      15.0    1440    0.0100   | 0.0272  0.9961 | 0.0292  0.9959 | 0.0332  0.9950  |  14.9 min    \n      16.0    1440    0.0100   | 0.0269  0.9961 | 0.0288  0.9960 | 0.0228  0.9966  |  15.0 min    \n      17.0    1440    0.0100   | 0.0263  0.9962 | 0.0332  0.9954 | 0.0256  0.9964  |  14.8 min    \n      18.0    1440    0.0100   | 0.0254  0.9963 | 0.0283  0.9962 | 0.0498  0.9916  |  14.9 min    \n      19.0    1440    0.0100   | 0.0249  0.9964 | 0.0285  0.9960 | 0.0248  0.9963  |  14.8 min    \n      20.0    1440    0.0100   | 0.0247  0.9964 | 0.0277  0.9961 | 0.0314  0.9952  |  14.9 min    \n      21.0    1440    0.0100   | 0.0248  0.9964 | 0.0274  0.9961 | 0.0240  0.9961  |  15.0 min    \n      22.0    1440    0.0100   | 0.0251  0.9963 | 0.0252  0.9965 | 0.0214  0.9969  |  14.9 min    \n      23.0    1440    0.0100   | 0.0252  0.9963 | 0.0284  0.9961 | 0.0343  0.9952  |  14.9 min    \n      24.0    1440    0.0100   | 0.0243  0.9965 | 0.0261  0.9963 | 0.0246  0.9966  |  14.8 min    \n      25.0    1440    0.0100   | 0.0248  0.9964 | 0.0279  0.9959 | 0.0197  0.9967  |  14.6 min    \n      26.0    1440    0.0100   | 0.0244  0.9964 | 0.0269  0.9961 | 0.0366  0.9952  |  14.6 min    \n      27.0    1440    0.0100   | 0.0240  0.9965 | 0.0275  0.9962 | 0.0273  0.9960  |  14.5 min    \n      28.0    1440    0.0100   | 0.0245  0.9964 | 0.0251  0.9964 | 0.0340  0.9953  |  14.5 min    \n      29.0    1440    0.0100   | 0.0244  0.9965 | 0.0281  0.9961 | 0.0220  0.9967  |  14.5 min    \n      30.0    1440    0.0100   | 0.0241  0.9965 | 0.0298  0.9958 | 0.0364  0.9954  |  14.5 min    \n      31.0    1440    0.0100   | 0.0235  0.9966 | 0.0240  0.9965 | 0.0273  0.9962  |  14.5 min    \n      32.0    1440    0.0100   | 0.0237  0.9965 | 0.0261  0.9963 | 0.0257  0.9962  |  14.5 min    \n      33.0    1440    0.0100   | 0.0235  0.9966 | 0.0240  0.9965 | 0.0191  0.9972  |  14.5 min    \n      34.0    1440    0.0100   | 0.0234  0.9966 | 0.0237  0.9965 | 0.0332  0.9955  |  14.5 min    \n\n      make some change to dataset : reduce augmentation\n      ... stop and resume training ....\n\n      LR=Step Learning Rates\n     rates=[' 0.0100', ' 0.0050', ' 0.0010', '-1.0000', '-1.0000']\n     steps=['      0', '     40', '     45', '     47', '     44']\n\n      epoch    iter      rate   | valid_loss/acc | train_loss/acc | batch_loss/acc ... \n      --------------------------------------------------------------------------------------------------\n       34.0    1440    0.0100   | 0.0231  0.9966 | 0.0230  0.9969 | 0.0262  0.9966  |  14.6 min \n       35.0    1440    0.0100   | 0.0234  0.9966 | 0.0216  0.9969 | 0.0247  0.9963  |  14.8 min \n       36.0    1440    0.0100   | 0.0231  0.9966 | 0.0214  0.9970 | 0.0220  0.9967  |  14.8 min \n       37.0    1440    0.0100   | 0.0238  0.9965 | 0.0216  0.9971 | 0.0281  0.9963  |  14.8 min \n       38.0    1440    0.0100   | 0.0228  0.9967 | 0.0237  0.9967 | 0.0247  0.9965  |  14.8 min \n       39.0    1440    0.0100   | 0.0226  0.9967 | 0.0230  0.9969 | 0.0188  0.9973  |  15.0 min \n       40.0    1440    0.0100   | 0.0230  0.9966 | 0.0224  0.9969 | 0.0231  0.9971  |  15.4 min \n       41.0    1440    0.0050   | 0.0224  0.9967 | 0.0224  0.9970 | 0.0180  0.9975  |  15.4 min \n       42.0    1440    0.0050   | 0.0225  0.9967 | 0.0218  0.9970 | 0.0216  0.9968  |  15.4 min \n       43.0    1440    0.0050   | 0.0224  0.9967 | 0.0202  0.9973 | 0.0214  0.9972  |  15.4 min \n       44.0    1440    0.0050   | 0.0224  0.9967 | 0.0196  0.9972 | 0.0172  0.9977  |  14.8 min",
      "votes": null
    },
    {
      "id": "213920",
      "postDate": "08/15/2017 14:34:00",
      "content": "<p>You said: \"Accumulate batches before updating\". How do you do it in keras?</p>\n\n<p>(5) Training hyper parameters</p>\n\n<p>Try different learning rates. I use about 45 epoch. Try different batch size. I use 3 images per batch, and <strong>accumulate 5 batches</strong> before updating the gradient. i.e. my effective batch size is 3x5 = 15. I just use SDG with momentum with manual rate scheduling.</p>",
      "rawMarkdown": "You said: \"Accumulate batches before updating\". How do you do it in keras?\n\n(5) Training hyper parameters\n\nTry different learning rates. I use about 45 epoch. Try different batch size. I use 3 images per batch, and **accumulate 5 batches** before updating the gradient. i.e. my effective batch size is 3x5 = 15. I just use SDG with momentum with manual rate scheduling.",
      "votes": null
    },
    {
      "id": "213921",
      "postDate": "08/15/2017 14:34:01",
      "content": "<p>@Heng CherKeng\nYou said \"accumulate batches before updating\":\n&gt;(5) Training hyper parameters\nTry different learning rates. I use about 45 epoch. Try different batch size. I use 3 images per batch, and <strong>accumulate 5 batches</strong> before updating the gradient. i.e. my effective batch size is 3x5 = 15. I just use SDG with momentum with manual rate scheduling.</p>\n\n<p>How do you do it in Keras?</p>",
      "rawMarkdown": "Heng CherKeng\nYou said \"accumulate batches before updating\":\n&gt;(5) Training hyper parameters\nTry different learning rates. I use about 45 epoch. Try different batch size. I use 3 images per batch, and **accumulate 5 batches** before updating the gradient. i.e. my effective batch size is 3x5 = 15. I just use SDG with momentum with manual rate scheduling.\n\nHow do you do it in Keras?",
      "votes": null
    },
    {
      "id": "213997",
      "postDate": "08/15/2017 17:16:08",
      "content": "<p>There is a way to do it here: <a href=\"https://github.com/fchollet/keras/issues/3556\">https://github.com/fchollet/keras/issues/3556</a></p>",
      "rawMarkdown": "There is a way to do it here: https://github.com/fchollet/keras/issues/3556",
      "votes": null
    },
    {
      "id": "214027",
      "postDate": "08/15/2017 19:12:47",
      "content": "<p>You da best. One question though. What did you mean by \"Use binary cross entropy and dice loss. Weigh pixels near boundary.\"?</p>",
      "rawMarkdown": "You da best. One question though. What did you mean by \"Use binary cross entropy and dice loss. Weigh pixels near boundary.\"?",
      "votes": null
    },
    {
      "id": "214062",
      "postDate": "08/15/2017 21:19:19",
      "content": "<p>check the dicsussion at: <a href=\"https://www.kaggle.com/c/carvana-image-masking-challenge/discussion/37208\">https://www.kaggle.com/c/carvana-image-masking-challenge/discussion/37208</a></p>",
      "rawMarkdown": "check the dicsussion at: https://www.kaggle.com/c/carvana-image-masking-challenge/discussion/37208",
      "votes": null
    },
    {
      "id": "214159",
      "postDate": "08/16/2017 04:52:26",
      "content": "<p>Thanks a lot for sharing. Curious what caused the unstable validation scores in <a href=\"https://www.kaggle.com/c/carvana-image-masking-challenge/discussion/37208\">https://www.kaggle.com/c/carvana-image-masking-challenge/discussion/37208</a>, and what did you do to fix it in this training?</p>",
      "rawMarkdown": "Thanks a lot for sharing. Curious what caused the unstable validation scores in https://www.kaggle.com/c/carvana-image-masking-challenge/discussion/37208, and what did you do to fix it in this training?",
      "votes": null
    },
    {
      "id": "214166",
      "postDate": "08/16/2017 05:26:08",
      "content": "<p>@Heng CherKeng here we meet again. </p>\n\n<p>Hi everyone, any advise for the resource challenged, I meant CPU users not GPU for this competition?</p>",
      "rawMarkdown": "Heng CherKeng here we meet again. \n\nHi everyone, any advise for the resource challenged, I meant CPU users not GPU for this competition?",
      "votes": null
    },
    {
      "id": "214215",
      "postDate": "08/16/2017 08:45:29",
      "content": "<p>Thanks for this wonderful post. Learn't many topics related to segmentation from your posts. I have one question regarding \"Train sample augmentation\". What sort of augmentation is normally required, I understand it varies from subject to subject, in general are there any list of things that we can try? My knowledge about image processing is limited. If you have any reference to read / suggestion it will be helpful. If I am asking too much information, that might create problems for the competition then there is no need to divulge the information now. I will wait for the competition to end and for more details to come out. It been a pleasure and great learning reading your posts on this forum and also of a lot of other Kagglers who have greatly contributed to my learning..\nThank you Kaggle and Kagglers..</p>",
      "rawMarkdown": "Thanks for this wonderful post. Learn't many topics related to segmentation from your posts. I have one question regarding \"Train sample augmentation\". What sort of augmentation is normally required, I understand it varies from subject to subject, in general are there any list of things that we can try? My knowledge about image processing is limited. If you have any reference to read / suggestion it will be helpful. If I am asking too much information, that might create problems for the competition then there is no need to divulge the information now. I will wait for the competition to end and for more details to come out. It been a pleasure and great learning reading your posts on this forum and also of a lot of other Kagglers who have greatly contributed to my learning..\nThank you Kaggle and Kagglers..",
      "votes": null
    },
    {
      "id": "214237",
      "postDate": "08/16/2017 10:52:03",
      "content": "<p>Follow Peter's post and wx405557858's suggestion, you can replace Adam in Keras with this :)</p>\n\n<pre><code>class Adam(Optimizer):\n\"\"\"Adam optimizer.\n\nDefault parameters follow those provided in the original paper.\n\n# Arguments\n    lr: float &gt;= 0. Learning rate.\n    beta_1: float, 0 &lt; beta &lt; 1. Generally close to 1.\n    beta_2: float, 0 &lt; beta &lt; 1. Generally close to 1.\n    epsilon: float &gt;= 0. Fuzz factor.\n    decay: float &gt;= 0. Learning rate decay over each update.\n\n# References\n    - [Adam - A Method for Stochastic Optimization](http://arxiv.org/abs/1412.6980v8)\n\"\"\"\n\ndef __init__(self, lr=0.001, beta_1=0.9, beta_2=0.999,\n             epsilon=1e-8, decay=0., accumulator=5., **kwargs):\n    super(Adam, self).__init__(**kwargs)\n    self.iterations = K.variable(0, name='iterations')\n    self.lr = K.variable(lr, name='lr')\n    self.beta_1 = K.variable(beta_1, name='beta_1')\n    self.beta_2 = K.variable(beta_2, name='beta_2')\n    self.epsilon = epsilon\n    self.decay = K.variable(decay, name='decay')\n    self.initial_decay = decay\n    self.accumulator = K.variable(accumulator, name='accumulator')\n\ndef get_updates(self, params, constraints, loss):\n    grads = self.get_gradients(loss, params)\n    self.updates = [K.update_add(self.iterations, 1)]\n\n    lr = self.lr\n    if self.initial_decay &gt; 0:\n        lr *= (1. / (1. + self.decay * self.iterations))\n\n    t = self.iterations + 1\n    lr_t = lr * (K.sqrt(1. - K.pow(self.beta_2, t)) /\n                 (1. - K.pow(self.beta_1, t)))\n\n    ms = [K.zeros(K.get_variable_shape(p), dtype=K.dtype(p)) for p in params]\n    vs = [K.zeros(K.get_variable_shape(p), dtype=K.dtype(p)) for p in params]\n    gs = [K.zeros(K.get_variable_shape(p), dtype=K.dtype(p)) for p in params]\n\n    self.weights = [self.iterations] + ms + vs\n\n    for p, g, m, v, ga in zip(params, grads, ms, vs, gs):\n\n\n        flag = K.equal(self.iterations % self.accumulator, 0)\n        flag = K.cast(flag, dtype='float32')\n\n        ga_t = (1 - flag) * (ga + g)\n\n        m_t = (self.beta_1 * m) + (1. - self.beta_1) * (ga + flag * g) / self.accumulator\n        v_t = (self.beta_2 * v) + (1. - self.beta_2) * K.square((ga + flag * g) / self.accumulator)\n        p_t = p - lr_t * m_t / (K.sqrt(v_t) + self.epsilon)\n\n\n        self.updates.append(K.update(m, flag * m_t + (1 - flag) * m))\n        self.updates.append(K.update(v, flag * v_t + (1 - flag) * v))\n        self.updates.append(K.update(ga, ga_t))\n\n        new_p = p_t\n        # apply constraints\n        if p in constraints:\n            c = constraints[p]\n            new_p = c(new_p)\n        self.updates.append(K.update(p, new_p))\n    return self.updates\n\ndef get_config(self):\n    config = {'lr': float(K.get_value(self.lr)),\n              'beta_1': float(K.get_value(self.beta_1)),\n              'beta_2': float(K.get_value(self.beta_2)),\n              'decay': float(K.get_value(self.decay)),\n              'accumulator': float(K.get_value(self.accumulator)),\n              'epsilon': self.epsilon}\n    base_config = super(Adam, self).get_config()\n    return dict(list(base_config.items()) + list(config.items()))\n</code></pre>",
      "rawMarkdown": "Follow Peter's post and wx405557858's suggestion, you can replace Adam in Keras with this :)\n\n    class Adam(Optimizer):\n    \"\"\"Adam optimizer.\n\n    Default parameters follow those provided in the original paper.\n\n    # Arguments\n        lr: float &gt;= 0. Learning rate.\n        beta_1: float, 0 &lt; beta &lt; 1. Generally close to 1.\n        beta_2: float, 0 &lt; beta &lt; 1. Generally close to 1.\n        epsilon: float &gt;= 0. Fuzz factor.\n        decay: float &gt;= 0. Learning rate decay over each update.\n\n    # References\n        - [Adam - A Method for Stochastic Optimization](http://arxiv.org/abs/1412.6980v8)\n    \"\"\"\n\n    def __init__(self, lr=0.001, beta_1=0.9, beta_2=0.999,\n                 epsilon=1e-8, decay=0., accumulator=5., **kwargs):\n        super(Adam, self).__init__(**kwargs)\n        self.iterations = K.variable(0, name='iterations')\n        self.lr = K.variable(lr, name='lr')\n        self.beta_1 = K.variable(beta_1, name='beta_1')\n        self.beta_2 = K.variable(beta_2, name='beta_2')\n        self.epsilon = epsilon\n        self.decay = K.variable(decay, name='decay')\n        self.initial_decay = decay\n        self.accumulator = K.variable(accumulator, name='accumulator')\n\n    def get_updates(self, params, constraints, loss):\n        grads = self.get_gradients(loss, params)\n        self.updates = [K.update_add(self.iterations, 1)]\n\n        lr = self.lr\n        if self.initial_decay &gt; 0:\n            lr *= (1. / (1. + self.decay * self.iterations))\n\n        t = self.iterations + 1\n        lr_t = lr * (K.sqrt(1. - K.pow(self.beta_2, t)) /\n                     (1. - K.pow(self.beta_1, t)))\n\n        ms = [K.zeros(K.get_variable_shape(p), dtype=K.dtype(p)) for p in params]\n        vs = [K.zeros(K.get_variable_shape(p), dtype=K.dtype(p)) for p in params]\n        gs = [K.zeros(K.get_variable_shape(p), dtype=K.dtype(p)) for p in params]\n\n        self.weights = [self.iterations] + ms + vs\n\n        for p, g, m, v, ga in zip(params, grads, ms, vs, gs):\n\n\n            flag = K.equal(self.iterations % self.accumulator, 0)\n            flag = K.cast(flag, dtype='float32')\n\n            ga_t = (1 - flag) * (ga + g)\n\n            m_t = (self.beta_1 * m) + (1. - self.beta_1) * (ga + flag * g) / self.accumulator\n            v_t = (self.beta_2 * v) + (1. - self.beta_2) * K.square((ga + flag * g) / self.accumulator)\n            p_t = p - lr_t * m_t / (K.sqrt(v_t) + self.epsilon)\n\n\n            self.updates.append(K.update(m, flag * m_t + (1 - flag) * m))\n            self.updates.append(K.update(v, flag * v_t + (1 - flag) * v))\n            self.updates.append(K.update(ga, ga_t))\n\n            new_p = p_t\n            # apply constraints\n            if p in constraints:\n                c = constraints[p]\n                new_p = c(new_p)\n            self.updates.append(K.update(p, new_p))\n        return self.updates\n\n    def get_config(self):\n        config = {'lr': float(K.get_value(self.lr)),\n                  'beta_1': float(K.get_value(self.beta_1)),\n                  'beta_2': float(K.get_value(self.beta_2)),\n                  'decay': float(K.get_value(self.decay)),\n                  'accumulator': float(K.get_value(self.accumulator)),\n                  'epsilon': self.epsilon}\n        base_config = super(Adam, self).get_config()\n        return dict(list(base_config.items()) + list(config.items()))",
      "votes": null
    },
    {
      "id": "214239",
      "postDate": "08/16/2017 10:53:59",
      "content": "<p>I don't think you can get good results efficiently with cpu. I think gpu is a must.</p>",
      "rawMarkdown": "I don't think you can get good results efficiently with cpu. I think gpu is a must.",
      "votes": null
    },
    {
      "id": "214241",
      "postDate": "08/16/2017 10:55:42",
      "content": "<p>I think it is due to small batch size. But i cannot confirm this. It can also be due to wrong label. The fluctuating is solved by accumulating gradient for larger effective batch size </p>",
      "rawMarkdown": "I think it is due to small batch size. But i cannot confirm this. It can also be due to wrong label. The fluctuating is solved by accumulating gradient for larger effective batch size",
      "votes": null
    },
    {
      "id": "214271",
      "postDate": "08/16/2017 13:13:51",
      "content": "<p>Hi Heng!</p>\n\n<p>Can you please clarify how many \"poolings\" are you using in your Unet for 1024x1024?\nI now use 4 pooling layers for 480x320, and it seems to be too small. So to save structure of network I should be use 6 pooling network for 1918x1280</p>",
      "rawMarkdown": "Hi Heng!\n\nCan you please clarify how many \"poolings\" are you using in your Unet for 1024x1024?\nI now use 4 pooling layers for 480x320, and it seems to be too small. So to save structure of network I should be use 6 pooling network for 1918x1280",
      "votes": null
    },
    {
      "id": "214324",
      "postDate": "08/16/2017 14:38:33",
      "content": "<p>shifting, horizontally flipping, random crops, scaling, rotating, color jittering.  A good summary on CNN tricks (including data augmentation): <a href=\"http://lamda.nju.edu.cn/weixs/project/CNNTricks/CNNTricks.html\">http://lamda.nju.edu.cn/weixs/project/CNNTricks/CNNTricks.html</a></p>",
      "rawMarkdown": "shifting, horizontally flipping, random crops, scaling, rotating, color jittering.  A good summary on CNN tricks (including data augmentation): http://lamda.nju.edu.cn/weixs/project/CNNTricks/CNNTricks.html",
      "votes": null
    },
    {
      "id": "214332",
      "postDate": "08/16/2017 15:13:28",
      "content": "<pre><code>class SGDAccum(Optimizer):\n\"\"\"Stochastic gradient descent optimizer.\nIncludes support for momentum,\nlearning rate decay, and Nesterov momentum.\n# Arguments\n    lr: float &gt;= 0. Learning rate.\n    momentum: float &gt;= 0. Parameter updates momentum.\n    decay: float &gt;= 0. Learning rate decay over each update.\n    nesterov: boolean. Whether to apply Nesterov momentum.\n\"\"\"\n\ndef __init__(self, lr=0.01, momentum=0., decay=0.,\n             nesterov=False, accum_iters=5, **kwargs):\n    super(SGDAccum, self).__init__(**kwargs)\n    with K.name_scope(self.__class__.__name__):\n        self.iterations = K.variable(0., name='iterations')\n        self.lr = K.variable(lr, name='lr')\n        self.momentum = K.variable(momentum, name='momentum')\n        self.decay = K.variable(decay, name='decay')\n        self.accum_iters = K.variable(accum_iters, name='accum_iters')\n    self.initial_decay = decay\n    self.nesterov = nesterov\n\n@interfaces.legacy_get_updates_support\ndef get_updates(self, loss, params):\n    grads = self.get_gradients(loss, params)\n    self.updates = []\n\n    lr = self.lr\n    if self.initial_decay &gt; 0:\n        lr *= (1. / (1. + self.decay * self.iterations))\n        self.updates.append(K.update_add(self.iterations, 1))\n\n    # momentum\n    shapes = [K.get_variable_shape(p) for p in params]\n    moments = [K.zeros(shape) for shape in shapes]\n    gradients = [K.zeros(shape) for shape in shapes]\n\n    self.weights = [self.iterations] + moments\n\n    for p, g, m, gg in zip(params, grads, moments, gradients):\n\n        flag = K.equal(self.iterations % self.accum_iters, 0)\n\n        if K.eval(flag):\n            flag = 1\n\n        gg_t = (1 - flag) * (gg + g)\n\n        v = self.momentum * m - lr * \\\n            K.square((gg + flag * g) / self.accum_iters)  # velocity\n\n        self.updates.append(K.update(m, flag * v + (1 - flag) * m))\n        self.updates.append((gg, gg_t))\n\n        if self.nesterov:\n            new_p = p + self.momentum * v - lr * g\n        else:\n            new_p = p + v\n\n        # Apply constraints.\n        if getattr(p, 'constraint', None) is not None:\n            new_p = p.constraint(new_p)\n\n        self.updates.append(K.update(p, new_p))\n\n    return self.updates\n\ndef get_config(self):\n    config = {'lr': float(K.get_value(self.lr)),\n              'momentum': float(K.get_value(self.momentum)),\n              'decay': float(K.get_value(self.decay)),\n              'nesterov': self.nesterov}\n    base_config = super(SGDAccum, self).get_config()\n    return dict(list(base_config.items()) + list(config.items()))\n</code></pre>",
      "rawMarkdown": "class SGDAccum(Optimizer):\n    \"\"\"Stochastic gradient descent optimizer.\n    Includes support for momentum,\n    learning rate decay, and Nesterov momentum.\n    # Arguments\n        lr: float &gt;= 0. Learning rate.\n        momentum: float &gt;= 0. Parameter updates momentum.\n        decay: float &gt;= 0. Learning rate decay over each update.\n        nesterov: boolean. Whether to apply Nesterov momentum.\n    \"\"\"\n\n    def __init__(self, lr=0.01, momentum=0., decay=0.,\n                 nesterov=False, accum_iters=5, **kwargs):\n        super(SGDAccum, self).__init__(**kwargs)\n        with K.name_scope(self.__class__.__name__):\n            self.iterations = K.variable(0., name='iterations')\n            self.lr = K.variable(lr, name='lr')\n            self.momentum = K.variable(momentum, name='momentum')\n            self.decay = K.variable(decay, name='decay')\n            self.accum_iters = K.variable(accum_iters, name='accum_iters')\n        self.initial_decay = decay\n        self.nesterov = nesterov\n\n    @interfaces.legacy_get_updates_support\n    def get_updates(self, loss, params):\n        grads = self.get_gradients(loss, params)\n        self.updates = []\n\n        lr = self.lr\n        if self.initial_decay &gt; 0:\n            lr *= (1. / (1. + self.decay * self.iterations))\n            self.updates.append(K.update_add(self.iterations, 1))\n\n        # momentum\n        shapes = [K.get_variable_shape(p) for p in params]\n        moments = [K.zeros(shape) for shape in shapes]\n        gradients = [K.zeros(shape) for shape in shapes]\n\n        self.weights = [self.iterations] + moments\n\n        for p, g, m, gg in zip(params, grads, moments, gradients):\n\n            flag = K.equal(self.iterations % self.accum_iters, 0)\n\n            if K.eval(flag):\n                flag = 1\n\n            gg_t = (1 - flag) * (gg + g)\n\n            v = self.momentum * m - lr * \\\n                K.square((gg + flag * g) / self.accum_iters)  # velocity\n\n            self.updates.append(K.update(m, flag * v + (1 - flag) * m))\n            self.updates.append((gg, gg_t))\n\n            if self.nesterov:\n                new_p = p + self.momentum * v - lr * g\n            else:\n                new_p = p + v\n\n            # Apply constraints.\n            if getattr(p, 'constraint', None) is not None:\n                new_p = p.constraint(new_p)\n\n            self.updates.append(K.update(p, new_p))\n\n        return self.updates\n\n    def get_config(self):\n        config = {'lr': float(K.get_value(self.lr)),\n                  'momentum': float(K.get_value(self.momentum)),\n                  'decay': float(K.get_value(self.decay)),\n                  'nesterov': self.nesterov}\n        base_config = super(SGDAccum, self).get_config()\n        return dict(list(base_config.items()) + list(config.items()))",
      "votes": null
    },
    {
      "id": "214333",
      "postDate": "08/16/2017 15:14:01",
      "content": "<p>A small hack for SGD Accumulator for keras, lemme know if it works.</p>",
      "rawMarkdown": "A small hack for SGD Accumulator for keras, lemme know if it works.",
      "votes": null
    },
    {
      "id": "214363",
      "postDate": "08/16/2017 16:34:04",
      "content": "<p>What changes you made in the original SGD?</p>",
      "rawMarkdown": "What changes you made in the original SGD?",
      "votes": null
    },
    {
      "id": "214377",
      "postDate": "08/16/2017 17:40:06",
      "content": "<p>This will accumulate the gradients and update the weights on accum_iters'th  iteration.</p>",
      "rawMarkdown": "This will accumulate the gradients and update the weights on accum_iters'th  iteration.",
      "votes": null
    },
    {
      "id": "214500",
      "postDate": "08/17/2017 06:54:27",
      "content": "<p>Thanks..</p>",
      "rawMarkdown": "Thanks..",
      "votes": null
    },
    {
      "id": "214509",
      "postDate": "08/17/2017 07:52:25",
      "content": "<p>Hi, Kaggoo, thanks for your sharing. Could you please the mean of <code>accumulator</code>?  Are the parameters  updated after 32 samples if we set <code>accumulator=5</code>? Could you please show how to calculate it? The batch_size was 4. Thank you.</p>",
      "rawMarkdown": "Hi, Kaggoo, thanks for your sharing. Could you please the mean of ```accumulator```?  Are the parameters  updated after 32 samples if we set ```accumulator=5```? Could you please show how to calculate it? The batch_size was 4. Thank you.",
      "votes": null
    },
    {
      "id": "214521",
      "postDate": "08/17/2017 08:19:15",
      "content": "<p>Thanks for your sharing. The parameter <code>accum_iters</code> is 5. If the previous batch_size=4, the final batch_size=32 when set the parameter? Could you show how do you calculate that? \nThanks!</p>",
      "rawMarkdown": "Thanks for your sharing. The parameter ```accum_iters``` is 5. If the previous batch_size=4, the final batch_size=32 when set the parameter? Could you show how do you calculate that? \nThanks!",
      "votes": null
    },
    {
      "id": "214596",
      "postDate": "08/17/2017 14:33:27",
      "content": "<p>Thanks @Heng,</p>\n\n<p>I am thinking of trying <a href=\"https://www.floydhub.com/\">Floyd</a>. Has anybody here used it? Any advise on the best way to set it up for this contest? I briefly checked their website but will register in a few days when I get a chance.</p>",
      "rawMarkdown": "Thanks @Heng,\n\nI am thinking of trying [Floyd](https://www.floydhub.com/). Has anybody here used it? Any advise on the best way to set it up for this contest? I briefly checked their website but will register in a few days when I get a chance.",
      "votes": null
    },
    {
      "id": "214661",
      "postDate": "08/17/2017 18:47:18",
      "content": "<p>I've used Floyd (but not for Kaggle). It's super intuitive, and easy to use, but make sure you test your code on a local machine before bringing it onto the cloud to save time and money. Also, they don't handle generators as well as they could- their I/O is slower than desired. As much as you can, limit I/O and do things in memory. </p>\n\n<p>GCP also offers a $300 dollar credit if you haven't used it for a project before- might be worth checking out. </p>",
      "rawMarkdown": "I've used Floyd (but not for Kaggle). It's super intuitive, and easy to use, but make sure you test your code on a local machine before bringing it onto the cloud to save time and money. Also, they don't handle generators as well as they could- their I/O is slower than desired. As much as you can, limit I/O and do things in memory. \n\n GCP also offers a $300 dollar credit if you haven't used it for a project before- might be worth checking out.",
      "votes": null
    },
    {
      "id": "214673",
      "postDate": "08/17/2017 20:17:18",
      "content": "<p>Hi Rohit. Thanks for sharing the SGD hacker. In the code you used a decorator, \"interfaces.legacy_get_updates_support\". Where is it defined?</p>",
      "rawMarkdown": "Hi Rohit. Thanks for sharing the SGD hacker. In the code you used a decorator, \"interfaces.legacy_get_updates_support\". Where is it defined?",
      "votes": null
    },
    {
      "id": "214681",
      "postDate": "08/17/2017 20:38:59",
      "content": "<p>I use a GPU VM on Google cloud platform that gives 300$ free trial to any new user. Once on the cloud you have two ways to do machine learning: </p>\n\n<ul>\n<li><p>Using a virtual machine (a google cloud computing engine instance). Simple as using the local machine. The cost is based on the machine you've chosed and the time for which it was used.</p></li>\n<li><p>Serverless, via ML engine, the google service provided with python Tensorflow API. In this case you pay only according to the \"job\" that you are required. It's the most innovative and appropriate way to do ML in cloud, but it's a bit more difficult and still I don't manage how to build my package for its API.</p></li>\n</ul>\n\n<p>I suggest you to start with the first option, create a VM, install what you want, use Google Storage browser API to load your python file and the kaggle-cli package ( <a href=\"https://github.com/floydwch/kaggle-cli\">https://github.com/floydwch/kaggle-cli</a> ) to load the data-set. </p>\n\n<p>Always remember to stop or delete the VM after you've used it, try to do the bulk of the debugging of your code locally to avoid wasting money, and the 300$ free should be enough to do many things.</p>",
      "rawMarkdown": "I use a GPU VM on Google cloud platform that gives 300$ free trial to any new user. Once on the cloud you have two ways to do machine learning: \n\n- Using a virtual machine (a google cloud computing engine instance). Simple as using the local machine. The cost is based on the machine you've chosed and the time for which it was used.\n\n- Serverless, via ML engine, the google service provided with python Tensorflow API. In this case you pay only according to the \"job\" that you are required. It's the most innovative and appropriate way to do ML in cloud, but it's a bit more difficult and still I don't manage how to build my package for its API.\n\nI suggest you to start with the first option, create a VM, install what you want, use Google Storage browser API to load your python file and the kaggle-cli package ( https://github.com/floydwch/kaggle-cli ) to load the data-set. \n\nAlways remember to stop or delete the VM after you've used it, try to do the bulk of the debugging of your code locally to avoid wasting money, and the 300$ free should be enough to do many things.",
      "votes": null
    },
    {
      "id": "214709",
      "postDate": "08/18/2017 00:47:33",
      "content": "<p>Hi Heng. Thanks for sharing your hyperparameter setting. When you trained your networking using different learning rate. Did you stop the model.fit(), change the learning rate and start again? Or is there a way to define multiple learning rates by hooking the learning rate to the number of epochs? </p>",
      "rawMarkdown": "Hi Heng. Thanks for sharing your hyperparameter setting. When you trained your networking using different learning rate. Did you stop the model.fit(), change the learning rate and start again? Or is there a way to define multiple learning rates by hooking the learning rate to the number of epochs?",
      "votes": null
    },
    {
      "id": "214879",
      "postDate": "08/18/2017 16:48:41",
      "content": "<p>Final batch size will be accum_iters*batch_size (you defined it to be 4 so, final=20)</p>",
      "rawMarkdown": "Final batch size will be accum_iters*batch_size (you defined it to be 4 so, final=20)",
      "votes": null
    },
    {
      "id": "214880",
      "postDate": "08/18/2017 16:50:05",
      "content": "<p>\"interfaces.legacy_get_updates_support\" in interfaces.py script of keras, this code needs to be pasted in optimizers.py in keras.</p>",
      "rawMarkdown": "\"interfaces.legacy_get_updates_support\" in interfaces.py script of keras, this code needs to be pasted in optimizers.py in keras.",
      "votes": null
    },
    {
      "id": "214969",
      "postDate": "08/19/2017 00:51:12",
      "content": "<p>@To Train Them, thanks for your useful suggestions, I will try Floyd as I have used the GCP credits on another project.</p>\n\n<p>@Brasnold, thanks for your kind advise and suggestions.</p>\n\n<p>For the people who down voted my asking this question. I do not understand what in my asking a question would rub anyone the wrong way to prompt such strong response. Can someone tell me what is wrong in asking a question of a kind and generous person like @Heng who I have communicated with in the past via other competitions? I believe we all have learnt from a lot of generous Kagglers during these contests by having discussions.</p>",
      "rawMarkdown": "To Train Them, thanks for your useful suggestions, I will try Floyd as I have used the GCP credits on another project.\n\n@Brasnold, thanks for your kind advise and suggestions.\n\nFor the people who down voted my asking this question. I do not understand what in my asking a question would rub anyone the wrong way to prompt such strong response. Can someone tell me what is wrong in asking a question of a kind and generous person like @Heng who I have communicated with in the past via other competitions? I believe we all have learnt from a lot of generous Kagglers during these contests by having discussions.",
      "votes": null
    },
    {
      "id": "214973",
      "postDate": "08/19/2017 01:33:44",
      "content": "<p>here is a very detailed google cloud platform tutorial from cs231n (<a href=\"http://cs231n.github.io/gce-tutorial/\">http://cs231n.github.io/gce-tutorial/</a>).  I followed the steps but got stuck on requiring gpu resources, good luck.</p>",
      "rawMarkdown": "here is a very detailed google cloud platform tutorial from cs231n (http://cs231n.github.io/gce-tutorial/).  I followed the steps but got stuck on requiring gpu resources, good luck.",
      "votes": null
    },
    {
      "id": "215147",
      "postDate": "08/20/2017 04:18:32",
      "content": "<p>Thanks Heng.  What does it mean when learning rate becomes negative?</p>\n\n<blockquote>\n  <blockquote>\n    <p>rates=[' 0.0100', ' 0.0050', ' 0.0010', '-1.0000', '-1.0000']</p>\n  </blockquote>\n</blockquote>",
      "rawMarkdown": "Thanks Heng.  What does it mean when learning rate becomes negative?\n\n&gt;&gt; rates=[' 0.0100', ' 0.0050', ' 0.0010', '-1.0000', '-1.0000']",
      "votes": null
    },
    {
      "id": "215152",
      "postDate": "08/20/2017 05:23:37",
      "content": "<p>It gets terminated</p>",
      "rawMarkdown": "It gets terminated",
      "votes": null
    },
    {
      "id": "215212",
      "postDate": "08/20/2017 14:08:09",
      "content": "<p>Thank you for clarity</p>",
      "rawMarkdown": "Thank you for clarity",
      "votes": null
    },
    {
      "id": "215263",
      "postDate": "08/20/2017 19:46:56",
      "content": "<p>@YaGana, why there are down votes: This is not the right place for your question. This thread isn't about resources and cloud services. First, you are polluting this thread with your question and second, you should have opened a new thread or post it into a respective one. There might be other users with the same question as well and they won't find it here. For the sake of sharing and finding knowledge post thread-unrelated questions into their own threads.</p>",
      "rawMarkdown": "YaGana, why there are down votes: This is not the right place for your question. This thread isn't about resources and cloud services. First, you are polluting this thread with your question and second, you should have opened a new thread or post it into a respective one. There might be other users with the same question as well and they won't find it here. For the sake of sharing and finding knowledge post thread-unrelated questions into their own threads.",
      "votes": null
    },
    {
      "id": "215356",
      "postDate": "08/21/2017 09:35:36",
      "content": "<p>@firolino thank you. My question is more directed to Heng but I get your point. As a novice, I cannot ask or message people directly yet. I am still fairly new to Kaggle. I wish the 1st person that down voted instead pointed that out. </p>\n\n<p>I will post it as an independent thread for others to contribute or learn from. I apologize everyone for hijacking the thread. </p>\n\n<p>If you want to respond to my question with suggestions please do it here (<a href=\"https://www.kaggle.com/c/carvana-image-masking-challenge/discussion/38375\">https://www.kaggle.com/c/carvana-image-masking-challenge/discussion/38375</a>)</p>\n\n<p>@Max thanks.</p>",
      "rawMarkdown": "firolino thank you. My question is more directed to Heng but I get your point. As a novice, I cannot ask or message people directly yet. I am still fairly new to Kaggle. I wish the 1st person that down voted instead pointed that out. \n\nI will post it as an independent thread for others to contribute or learn from. I apologize everyone for hijacking the thread. \n\nIf you want to respond to my question with suggestions please do it here (https://www.kaggle.com/c/carvana-image-masking-challenge/discussion/38375)\n\n@Max thanks.",
      "votes": null
    },
    {
      "id": "215406",
      "postDate": "08/21/2017 13:54:36",
      "content": "<p>Just want to know how can I accumulate the gradient before update loss?</p>\n\n<p>So that means only every 5 steps I should call optimizer.step() and optimizer.zero_grad()?</p>\n\n<p>Is there any other thing I need to pay attention?</p>",
      "rawMarkdown": "Just want to know how can I accumulate the gradient before update loss?\n\nSo that means only every 5 steps I should call optimizer.step() and optimizer.zero_grad()?\n\nIs there any other thing I need to pay attention?",
      "votes": null
    },
    {
      "id": "215407",
      "postDate": "08/21/2017 14:00:33",
      "content": "<p>for example:</p>\n\n<pre><code> optimizer.zero_grad() \n\n x,y_hat =  .... get one batch of data\n y = net(x)\n loss.backward(y, y_hat)  #accumulate gradient  for 1 times \n\n x,y_hat =  .... get one batch of data\n y = net(x)\n loss.backward(y, y_hat)  #accumulate gradient  for 2 times \n\n x,y_hat=  .... get one batch of data\n y = net(x)\n loss.backward(y, y_hat)  #accumulate gradient  for 3 times \n\n optimizer.step()  #update accumulated gradients\n\n print(loss.data[0]) #print only loss for lass batch , not accumulated loss \n</code></pre>",
      "rawMarkdown": "for example:\n\n \n     optimizer.zero_grad() \n\n     x,y_hat =  .... get one batch of data\n     y = net(x)\n     loss.backward(y, y_hat)  #accumulate gradient  for 1 times \n\n     x,y_hat =  .... get one batch of data\n     y = net(x)\n     loss.backward(y, y_hat)  #accumulate gradient  for 2 times \n\n     x,y_hat=  .... get one batch of data\n     y = net(x)\n     loss.backward(y, y_hat)  #accumulate gradient  for 3 times \n\n     optimizer.step()  #update accumulated gradients\n\n     print(loss.data[0]) #print only loss for lass batch , not accumulated loss",
      "votes": null
    },
    {
      "id": "215417",
      "postDate": "08/21/2017 15:08:23",
      "content": "<p>Thanks!</p>",
      "rawMarkdown": "Thanks!",
      "votes": null
    },
    {
      "id": "215617",
      "postDate": "08/22/2017 10:47:09",
      "content": "<p>Sorry, it seems that it doesn't work in my train.</p>",
      "rawMarkdown": "Sorry, it seems that it doesn't work in my train.",
      "votes": null
    },
    {
      "id": "216098",
      "postDate": "08/24/2017 09:23:28",
      "content": "<p>if self.initial_decay is 0, then self.iterations doesn't make change?</p>",
      "rawMarkdown": "if self.initial_decay is 0, then self.iterations doesn't make change?",
      "votes": null
    },
    {
      "id": "216495",
      "postDate": "08/26/2017 01:46:01",
      "content": "<p>Note to everyone- accumulator has to be a float or it throws an error from internal calculations. Spent 2 hours looking for an implementation bug that didn't exist....</p>",
      "rawMarkdown": "Note to everyone- accumulator has to be a float or it throws an error from internal calculations. Spent 2 hours looking for an implementation bug that didn't exist....",
      "votes": null
    },
    {
      "id": "216624",
      "postDate": "08/27/2017 02:25:20",
      "content": "<p>Could you tell me is the accumulate gradient's method is simply add the gradients together or average evergytime gradient ? If it just add togather I think maybe it is no benefit.  </p>",
      "rawMarkdown": "Could you tell me is the accumulate gradient's method is simply add the gradients together or average evergytime gradient ? If it just add togather I think maybe it is no benefit.",
      "votes": null
    },
    {
      "id": "216628",
      "postDate": "08/27/2017 03:03:22",
      "content": "<p>If you add together the gradients, it effectively gives you a larger batch size, which does help when everyone is optimizing for the 4-5th digit. </p>",
      "rawMarkdown": "If you add together the gradients, it effectively gives you a larger batch size, which does help when everyone is optimizing for the 4-5th digit.",
      "votes": null
    },
    {
      "id": "216674",
      "postDate": "08/27/2017 12:03:23",
      "content": "<p>I used flag = K.cast(flag, dtype='float32') to make sure the flag is float32 and thus the remaining formula should work. Tested on Keras with TF backend.</p>",
      "rawMarkdown": "I used flag = K.cast(flag, dtype='float32') to make sure the flag is float32 and thus the remaining formula should work. Tested on Keras with TF backend.",
      "votes": null
    },
    {
      "id": "216770",
      "postDate": "08/28/2017 03:34:41",
      "content": "<p>I meant about the accumulator variable (5. instead of 5). But the flag is important also</p>",
      "rawMarkdown": "I meant about the accumulator variable (5. instead of 5). But the flag is important also",
      "votes": null
    },
    {
      "id": "217777",
      "postDate": "08/31/2017 21:58:10",
      "content": "<p>So how exactly do you use this, just create a file with it, and import it where the model is made? does it only depend on keras.backend and keras.optimizers.Optimizer? sorry, I just have never played around with custom optimizers in keras</p>",
      "rawMarkdown": "So how exactly do you use this, just create a file with it, and import it where the model is made? does it only depend on keras.backend and keras.optimizers.Optimizer? sorry, I just have never played around with custom optimizers in keras",
      "votes": null
    },
    {
      "id": "217779",
      "postDate": "08/31/2017 22:05:42",
      "content": "<p>You have to find where keras is installed on your computer and add this to the file optimizers.py</p>",
      "rawMarkdown": "You have to find where keras is installed on your computer and add this to the file optimizers.py",
      "votes": null
    },
    {
      "id": "218035",
      "postDate": "09/01/2017 22:55:41",
      "content": "<p>Do you get a speed up when increasing the accumulation here? I placed it in optimizers.py and it worked but no matter how much I increase the accumulation, it is still longer than normal adam. For example, normal adam with batch_size=2, takes for me ~55 minutes,  and with batch_size=2 and accumulation=100 on Adam_accumulate I get ~1 hour. </p>",
      "rawMarkdown": "Do you get a speed up when increasing the accumulation here? I placed it in optimizers.py and it worked but no matter how much I increase the accumulation, it is still longer than normal adam. For example, normal adam with batch_size=2, takes for me ~55 minutes,  and with batch_size=2 and accumulation=100 on Adam_accumulate I get ~1 hour.",
      "votes": null
    },
    {
      "id": "218079",
      "postDate": "09/02/2017 05:44:03",
      "content": "<p>Do you keep the accumulated gradients in the RAM? What if the effective size of all GPUs combined is equal to the size of RAM?</p>",
      "rawMarkdown": "Do you keep the accumulated gradients in the RAM? What if the effective size of all GPUs combined is equal to the size of RAM?",
      "votes": null
    },
    {
      "id": "218080",
      "postDate": "09/02/2017 05:58:11",
      "content": "<p>No. If anything it slows it down. The goal is to increase batch size since there are limitations on VRAM.</p>",
      "rawMarkdown": "No. If anything it slows it down. The goal is to increase batch size since there are limitations on VRAM.",
      "votes": null
    },
    {
      "id": "218084",
      "postDate": "09/02/2017 06:22:21",
      "content": "<p>I know they're stored in RAM for tensorflow, but this is a workaround if you only have a single GPU. I'm sure if you have enough vram to have a batch size of 15 for 1024x1024 images the results would be slightly better.</p>",
      "rawMarkdown": "I know they're stored in RAM for tensorflow, but this is a workaround if you only have a single GPU. I'm sure if you have enough vram to have a batch size of 15 for 1024x1024 images the results would be slightly better.",
      "votes": null
    },
    {
      "id": "218097",
      "postDate": "09/02/2017 08:14:33",
      "content": "<p>@Steven, so if my training takes 12hrs, using this Adam, will result in may be twice the training time i.e. 20hrs but I can go from batch_size=10 to 12 without memory problem? These numbers are just examples as it all depends on one's setup. Am I right?</p>",
      "rawMarkdown": "Steven, so if my training takes 12hrs, using this Adam, will result in may be twice the training time i.e. 20hrs but I can go from batch_size=10 to 12 without memory problem? These numbers are just examples as it all depends on one's setup. Am I right?",
      "votes": null
    },
    {
      "id": "218098",
      "postDate": "09/02/2017 08:40:43",
      "content": "<p>Using the accumulator shouldn't double the training time. In case of @mshliselberg, the training duration increased from 55 mins to 60 mins (per epoch I guess). \nAlso, you don't incease your batch size, but accumulate the batches. So with a batch size of 10, you perform updates every 10*x samples (where x is an integer).</p>",
      "rawMarkdown": "Using the accumulator shouldn't double the training time. In case of @mshliselberg, the training duration increased from 55 mins to 60 mins (per epoch I guess). \nAlso, you don't incease your batch size, but accumulate the batches. So with a batch size of 10, you perform updates every 10*x samples (where x is an integer).",
      "votes": null
    },
    {
      "id": "218241",
      "postDate": "09/03/2017 03:54:53",
      "content": "<p>So I am also using pascal titan-X and when using batch size of 2 with pixel budget &lt;1024^2 It takes me still ~1 hour to train per epoch, how did you manage 15 minutes. I understand your using accumulation but i thought that doesnt really help speedup as mentioned in the comments here (thanks to @Steven Nguyen). Is there anything else you have done for speedup?</p>",
      "rawMarkdown": "So I am also using pascal titan-X and when using batch size of 2 with pixel budget &lt;1024^2 It takes me still ~1 hour to train per epoch, how did you manage 15 minutes. I understand your using accumulation but i thought that doesnt really help speedup as mentioned in the comments here (thanks to @Steven Nguyen). Is there anything else you have done for speedup?",
      "votes": null
    },
    {
      "id": "218243",
      "postDate": "09/03/2017 03:58:23",
      "content": "<p>maybe your network structure is different from mine.</p>",
      "rawMarkdown": "maybe your network structure is different from mine.",
      "votes": null
    },
    {
      "id": "218246",
      "postDate": "09/03/2017 04:14:31",
      "content": "<p>Thanks  @Markus. I missread  @mshliselberg's comment.</p>",
      "rawMarkdown": "Thanks  @Markus. I missread  @mshliselberg's comment.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 213813,
      "author_name": "timjoseph",
      "author_url": "",
      "post_date": "08/15/2017 10:01:13",
      "content": "<p>Good summary.\nOne question about 3): 1024x1024 rescaled to 1280x1918 or sliced into 1024x1024 pieces?</p>",
      "votes": null,
      "replies": [
        {
          "id": 213814,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "08/15/2017 10:02:31",
          "content": "<p>1024x1024 rescaled to 1280x1918. Only one-step resize, no cropping.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 213815,
      "author_name": "petrosgk",
      "author_url": "",
      "post_date": "08/15/2017 10:10:28",
      "content": "<p>May also be worth testing variations of the Unet architecture, like replacing the VGG-like architecture with residual or inception blocks.  Not necessary for 0.997 but could be for 0.998 :)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 213842,
      "author_name": "zhihang",
      "author_url": "",
      "post_date": "08/15/2017 11:30:46",
      "content": "<p>0.996 ers are looking at you</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 213857,
      "author_name": "authman",
      "author_url": "",
      "post_date": "08/15/2017 12:05:27",
      "content": "<p>What lr-scheduler do you use for optimization?</p>",
      "votes": null,
      "replies": [
        {
          "id": 213860,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "08/15/2017 12:08:50",
          "content": "<p>note that this has to depend on your network, data augment, data batch size, loss,  etc. What works for me may not work for you.</p>\n\n<p>For me, i use conv-bn-relu. for epoch 0 to 40: 0.01,  40 to 45: 0.005, 45 to 47: 0.001. My training loss is as attached:</p>\n\n<pre><code>   LR=Step Learning Rates\n   rates=[' 0.0100', ' 0.0050', ' 0.0010', '-1.0000', '-1.0000']\n   steps=['      0', '     35', '     40', '     42', '     44']\n\n epoch    iter      rate   | valid_loss/acc | train_loss/acc | batch_loss/acc ...             \n ---------------------------------------------------------------------------------------------\n   1.0    1440    0.0100   | 0.0670  0.9900 | 0.0807  0.9883 | 0.0897  0.9885  |  14.8 min    \n   2.0    1440    0.0100   | 0.0470  0.9937 | 0.0566  0.9922 | 0.0572  0.9918  |  14.7 min    \n   3.0    1440    0.0100   | 0.0380  0.9948 | 0.0496  0.9929 | 0.0369  0.9949  |  14.7 min    \n   4.0    1440    0.0100   | 0.0400  0.9944 | 0.0414  0.9944 | 0.0504  0.9934  |  14.7 min    \n   5.0    1440    0.0100   | 0.0317  0.9956 | 0.0400  0.9946 | 0.0821  0.9896  |  14.7 min    \n   6.0    1440    0.0100   | 0.0310  0.9956 | 0.0401  0.9947 | 0.0391  0.9950  |  14.7 min    \n   7.0    1440    0.0100   | 0.0314  0.9955 | 0.0397  0.9946 | 0.0364  0.9954  |  14.7 min    \n   8.0    1440    0.0100   | 0.0286  0.9959 | 0.0344  0.9952 | 0.0267  0.9959  |  14.7 min    \n   9.0    1440    0.0100   | 0.0286  0.9959 | 0.0375  0.9951 | 0.0296  0.9953  |  14.7 min    \n  10.0    1440    0.0100   | 0.0284  0.9959 | 0.0321  0.9956 | 0.0435  0.9948  |  14.9 min    \n  11.0    1440    0.0100   | 0.0280  0.9960 | 0.0325  0.9954 | 0.0253  0.9967  |  15.1 min    \n  12.0    1440    0.0100   | 0.0270  0.9961 | 0.0349  0.9950 | 0.0352  0.9945  |  14.9 min    \n  13.0    1440    0.0100   | 0.0266  0.9962 | 0.0335  0.9953 | 0.0296  0.9952  |  14.9 min    \n  14.0    1440    0.0100   | 0.0263  0.9962 | 0.0324  0.9955 | 0.0287  0.9962  |  15.1 min    \n  15.0    1440    0.0100   | 0.0272  0.9961 | 0.0292  0.9959 | 0.0332  0.9950  |  14.9 min    \n  16.0    1440    0.0100   | 0.0269  0.9961 | 0.0288  0.9960 | 0.0228  0.9966  |  15.0 min    \n  17.0    1440    0.0100   | 0.0263  0.9962 | 0.0332  0.9954 | 0.0256  0.9964  |  14.8 min    \n  18.0    1440    0.0100   | 0.0254  0.9963 | 0.0283  0.9962 | 0.0498  0.9916  |  14.9 min    \n  19.0    1440    0.0100   | 0.0249  0.9964 | 0.0285  0.9960 | 0.0248  0.9963  |  14.8 min    \n  20.0    1440    0.0100   | 0.0247  0.9964 | 0.0277  0.9961 | 0.0314  0.9952  |  14.9 min    \n  21.0    1440    0.0100   | 0.0248  0.9964 | 0.0274  0.9961 | 0.0240  0.9961  |  15.0 min    \n  22.0    1440    0.0100   | 0.0251  0.9963 | 0.0252  0.9965 | 0.0214  0.9969  |  14.9 min    \n  23.0    1440    0.0100   | 0.0252  0.9963 | 0.0284  0.9961 | 0.0343  0.9952  |  14.9 min    \n  24.0    1440    0.0100   | 0.0243  0.9965 | 0.0261  0.9963 | 0.0246  0.9966  |  14.8 min    \n  25.0    1440    0.0100   | 0.0248  0.9964 | 0.0279  0.9959 | 0.0197  0.9967  |  14.6 min    \n  26.0    1440    0.0100   | 0.0244  0.9964 | 0.0269  0.9961 | 0.0366  0.9952  |  14.6 min    \n  27.0    1440    0.0100   | 0.0240  0.9965 | 0.0275  0.9962 | 0.0273  0.9960  |  14.5 min    \n  28.0    1440    0.0100   | 0.0245  0.9964 | 0.0251  0.9964 | 0.0340  0.9953  |  14.5 min    \n  29.0    1440    0.0100   | 0.0244  0.9965 | 0.0281  0.9961 | 0.0220  0.9967  |  14.5 min    \n  30.0    1440    0.0100   | 0.0241  0.9965 | 0.0298  0.9958 | 0.0364  0.9954  |  14.5 min    \n  31.0    1440    0.0100   | 0.0235  0.9966 | 0.0240  0.9965 | 0.0273  0.9962  |  14.5 min    \n  32.0    1440    0.0100   | 0.0237  0.9965 | 0.0261  0.9963 | 0.0257  0.9962  |  14.5 min    \n  33.0    1440    0.0100   | 0.0235  0.9966 | 0.0240  0.9965 | 0.0191  0.9972  |  14.5 min    \n  34.0    1440    0.0100   | 0.0234  0.9966 | 0.0237  0.9965 | 0.0332  0.9955  |  14.5 min    \n\n  make some change to dataset : reduce augmentation\n  ... stop and resume training ....\n\n  LR=Step Learning Rates\n rates=[' 0.0100', ' 0.0050', ' 0.0010', '-1.0000', '-1.0000']\n steps=['      0', '     40', '     45', '     47', '     44']\n\n  epoch    iter      rate   | valid_loss/acc | train_loss/acc | batch_loss/acc ... \n  --------------------------------------------------------------------------------------------------\n   34.0    1440    0.0100   | 0.0231  0.9966 | 0.0230  0.9969 | 0.0262  0.9966  |  14.6 min \n   35.0    1440    0.0100   | 0.0234  0.9966 | 0.0216  0.9969 | 0.0247  0.9963  |  14.8 min \n   36.0    1440    0.0100   | 0.0231  0.9966 | 0.0214  0.9970 | 0.0220  0.9967  |  14.8 min \n   37.0    1440    0.0100   | 0.0238  0.9965 | 0.0216  0.9971 | 0.0281  0.9963  |  14.8 min \n   38.0    1440    0.0100   | 0.0228  0.9967 | 0.0237  0.9967 | 0.0247  0.9965  |  14.8 min \n   39.0    1440    0.0100   | 0.0226  0.9967 | 0.0230  0.9969 | 0.0188  0.9973  |  15.0 min \n   40.0    1440    0.0100   | 0.0230  0.9966 | 0.0224  0.9969 | 0.0231  0.9971  |  15.4 min \n   41.0    1440    0.0050   | 0.0224  0.9967 | 0.0224  0.9970 | 0.0180  0.9975  |  15.4 min \n   42.0    1440    0.0050   | 0.0225  0.9967 | 0.0218  0.9970 | 0.0216  0.9968  |  15.4 min \n   43.0    1440    0.0050   | 0.0224  0.9967 | 0.0202  0.9973 | 0.0214  0.9972  |  15.4 min \n   44.0    1440    0.0050   | 0.0224  0.9967 | 0.0196  0.9972 | 0.0172  0.9977  |  14.8 min \n</code></pre>",
          "votes": null,
          "replies": []
        },
        {
          "id": 214159,
          "author_name": "luckyguy",
          "author_url": "",
          "post_date": "08/16/2017 04:52:26",
          "content": "<p>Thanks a lot for sharing. Curious what caused the unstable validation scores in <a href=\"https://www.kaggle.com/c/carvana-image-masking-challenge/discussion/37208\">https://www.kaggle.com/c/carvana-image-masking-challenge/discussion/37208</a>, and what did you do to fix it in this training?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 214241,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "08/16/2017 10:55:42",
          "content": "<p>I think it is due to small batch size. But i cannot confirm this. It can also be due to wrong label. The fluctuating is solved by accumulating gradient for larger effective batch size </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 214709,
          "author_name": "xiaokangwang",
          "author_url": "",
          "post_date": "08/18/2017 00:47:33",
          "content": "<p>Hi Heng. Thanks for sharing your hyperparameter setting. When you trained your networking using different learning rate. Did you stop the model.fit(), change the learning rate and start again? Or is there a way to define multiple learning rates by hooking the learning rate to the number of epochs? </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 215147,
          "author_name": "jackkwok",
          "author_url": "",
          "post_date": "08/20/2017 04:18:32",
          "content": "<p>Thanks Heng.  What does it mean when learning rate becomes negative?</p>\n\n<blockquote>\n  <blockquote>\n    <p>rates=[' 0.0100', ' 0.0050', ' 0.0010', '-1.0000', '-1.0000']</p>\n  </blockquote>\n</blockquote>",
          "votes": null,
          "replies": []
        },
        {
          "id": 215152,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "08/20/2017 05:23:37",
          "content": "<p>It gets terminated</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 213920,
      "author_name": "lnicalo",
      "author_url": "",
      "post_date": "08/15/2017 14:34:00",
      "content": "<p>You said: \"Accumulate batches before updating\". How do you do it in keras?</p>\n\n<p>(5) Training hyper parameters</p>\n\n<p>Try different learning rates. I use about 45 epoch. Try different batch size. I use 3 images per batch, and <strong>accumulate 5 batches</strong> before updating the gradient. i.e. my effective batch size is 3x5 = 15. I just use SDG with momentum with manual rate scheduling.</p>",
      "votes": null,
      "replies": [
        {
          "id": 213997,
          "author_name": "petrosgk",
          "author_url": "",
          "post_date": "08/15/2017 17:16:08",
          "content": "<p>There is a way to do it here: <a href=\"https://github.com/fchollet/keras/issues/3556\">https://github.com/fchollet/keras/issues/3556</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 214237,
          "author_name": "lpachuong",
          "author_url": "",
          "post_date": "08/16/2017 10:52:03",
          "content": "<p>Follow Peter's post and wx405557858's suggestion, you can replace Adam in Keras with this :)</p>\n\n<pre><code>class Adam(Optimizer):\n\"\"\"Adam optimizer.\n\nDefault parameters follow those provided in the original paper.\n\n# Arguments\n    lr: float &gt;= 0. Learning rate.\n    beta_1: float, 0 &lt; beta &lt; 1. Generally close to 1.\n    beta_2: float, 0 &lt; beta &lt; 1. Generally close to 1.\n    epsilon: float &gt;= 0. Fuzz factor.\n    decay: float &gt;= 0. Learning rate decay over each update.\n\n# References\n    - [Adam - A Method for Stochastic Optimization](http://arxiv.org/abs/1412.6980v8)\n\"\"\"\n\ndef __init__(self, lr=0.001, beta_1=0.9, beta_2=0.999,\n             epsilon=1e-8, decay=0., accumulator=5., **kwargs):\n    super(Adam, self).__init__(**kwargs)\n    self.iterations = K.variable(0, name='iterations')\n    self.lr = K.variable(lr, name='lr')\n    self.beta_1 = K.variable(beta_1, name='beta_1')\n    self.beta_2 = K.variable(beta_2, name='beta_2')\n    self.epsilon = epsilon\n    self.decay = K.variable(decay, name='decay')\n    self.initial_decay = decay\n    self.accumulator = K.variable(accumulator, name='accumulator')\n\ndef get_updates(self, params, constraints, loss):\n    grads = self.get_gradients(loss, params)\n    self.updates = [K.update_add(self.iterations, 1)]\n\n    lr = self.lr\n    if self.initial_decay &gt; 0:\n        lr *= (1. / (1. + self.decay * self.iterations))\n\n    t = self.iterations + 1\n    lr_t = lr * (K.sqrt(1. - K.pow(self.beta_2, t)) /\n                 (1. - K.pow(self.beta_1, t)))\n\n    ms = [K.zeros(K.get_variable_shape(p), dtype=K.dtype(p)) for p in params]\n    vs = [K.zeros(K.get_variable_shape(p), dtype=K.dtype(p)) for p in params]\n    gs = [K.zeros(K.get_variable_shape(p), dtype=K.dtype(p)) for p in params]\n\n    self.weights = [self.iterations] + ms + vs\n\n    for p, g, m, v, ga in zip(params, grads, ms, vs, gs):\n\n\n        flag = K.equal(self.iterations % self.accumulator, 0)\n        flag = K.cast(flag, dtype='float32')\n\n        ga_t = (1 - flag) * (ga + g)\n\n        m_t = (self.beta_1 * m) + (1. - self.beta_1) * (ga + flag * g) / self.accumulator\n        v_t = (self.beta_2 * v) + (1. - self.beta_2) * K.square((ga + flag * g) / self.accumulator)\n        p_t = p - lr_t * m_t / (K.sqrt(v_t) + self.epsilon)\n\n\n        self.updates.append(K.update(m, flag * m_t + (1 - flag) * m))\n        self.updates.append(K.update(v, flag * v_t + (1 - flag) * v))\n        self.updates.append(K.update(ga, ga_t))\n\n        new_p = p_t\n        # apply constraints\n        if p in constraints:\n            c = constraints[p]\n            new_p = c(new_p)\n        self.updates.append(K.update(p, new_p))\n    return self.updates\n\ndef get_config(self):\n    config = {'lr': float(K.get_value(self.lr)),\n              'beta_1': float(K.get_value(self.beta_1)),\n              'beta_2': float(K.get_value(self.beta_2)),\n              'decay': float(K.get_value(self.decay)),\n              'accumulator': float(K.get_value(self.accumulator)),\n              'epsilon': self.epsilon}\n    base_config = super(Adam, self).get_config()\n    return dict(list(base_config.items()) + list(config.items()))\n</code></pre>",
          "votes": null,
          "replies": []
        },
        {
          "id": 214509,
          "author_name": "zhangsongwei",
          "author_url": "",
          "post_date": "08/17/2017 07:52:25",
          "content": "<p>Hi, Kaggoo, thanks for your sharing. Could you please the mean of <code>accumulator</code>?  Are the parameters  updated after 32 samples if we set <code>accumulator=5</code>? Could you please show how to calculate it? The batch_size was 4. Thank you.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 216495,
          "author_name": "rhgrossm",
          "author_url": "",
          "post_date": "08/26/2017 01:46:01",
          "content": "<p>Note to everyone- accumulator has to be a float or it throws an error from internal calculations. Spent 2 hours looking for an implementation bug that didn't exist....</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 216674,
          "author_name": "lpachuong",
          "author_url": "",
          "post_date": "08/27/2017 12:03:23",
          "content": "<p>I used flag = K.cast(flag, dtype='float32') to make sure the flag is float32 and thus the remaining formula should work. Tested on Keras with TF backend.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 216770,
          "author_name": "rhgrossm",
          "author_url": "",
          "post_date": "08/28/2017 03:34:41",
          "content": "<p>I meant about the accumulator variable (5. instead of 5). But the flag is important also</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 217777,
          "author_name": "mshliselberg",
          "author_url": "",
          "post_date": "08/31/2017 21:58:10",
          "content": "<p>So how exactly do you use this, just create a file with it, and import it where the model is made? does it only depend on keras.backend and keras.optimizers.Optimizer? sorry, I just have never played around with custom optimizers in keras</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 217779,
          "author_name": "stevenknguyen",
          "author_url": "",
          "post_date": "08/31/2017 22:05:42",
          "content": "<p>You have to find where keras is installed on your computer and add this to the file optimizers.py</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 218035,
          "author_name": "mshliselberg",
          "author_url": "",
          "post_date": "09/01/2017 22:55:41",
          "content": "<p>Do you get a speed up when increasing the accumulation here? I placed it in optimizers.py and it worked but no matter how much I increase the accumulation, it is still longer than normal adam. For example, normal adam with batch_size=2, takes for me ~55 minutes,  and with batch_size=2 and accumulation=100 on Adam_accumulate I get ~1 hour. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 218080,
          "author_name": "stevenknguyen",
          "author_url": "",
          "post_date": "09/02/2017 05:58:11",
          "content": "<p>No. If anything it slows it down. The goal is to increase batch size since there are limitations on VRAM.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 218097,
          "author_name": "sheriytm",
          "author_url": "",
          "post_date": "09/02/2017 08:14:33",
          "content": "<p>@Steven, so if my training takes 12hrs, using this Adam, will result in may be twice the training time i.e. 20hrs but I can go from batch_size=10 to 12 without memory problem? These numbers are just examples as it all depends on one's setup. Am I right?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 218098,
          "author_name": "depthfirstsearch",
          "author_url": "",
          "post_date": "09/02/2017 08:40:43",
          "content": "<p>Using the accumulator shouldn't double the training time. In case of @mshliselberg, the training duration increased from 55 mins to 60 mins (per epoch I guess). \nAlso, you don't incease your batch size, but accumulate the batches. So with a batch size of 10, you perform updates every 10*x samples (where x is an integer).</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 218246,
          "author_name": "sheriytm",
          "author_url": "",
          "post_date": "09/03/2017 04:14:31",
          "content": "<p>Thanks  @Markus. I missread  @mshliselberg's comment.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 213921,
      "author_name": "lnicalo",
      "author_url": "",
      "post_date": "08/15/2017 14:34:01",
      "content": "<p>@Heng CherKeng\nYou said \"accumulate batches before updating\":\n&gt;(5) Training hyper parameters\nTry different learning rates. I use about 45 epoch. Try different batch size. I use 3 images per batch, and <strong>accumulate 5 batches</strong> before updating the gradient. i.e. my effective batch size is 3x5 = 15. I just use SDG with momentum with manual rate scheduling.</p>\n\n<p>How do you do it in Keras?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 214027,
      "author_name": "harungunaydin",
      "author_url": "",
      "post_date": "08/15/2017 19:12:47",
      "content": "<p>You da best. One question though. What did you mean by \"Use binary cross entropy and dice loss. Weigh pixels near boundary.\"?</p>",
      "votes": null,
      "replies": [
        {
          "id": 214062,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "08/15/2017 21:19:19",
          "content": "<p>check the dicsussion at: <a href=\"https://www.kaggle.com/c/carvana-image-masking-challenge/discussion/37208\">https://www.kaggle.com/c/carvana-image-masking-challenge/discussion/37208</a></p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 214166,
      "author_name": "sheriytm",
      "author_url": "",
      "post_date": "08/16/2017 05:26:08",
      "content": "<p>@Heng CherKeng here we meet again. </p>\n\n<p>Hi everyone, any advise for the resource challenged, I meant CPU users not GPU for this competition?</p>",
      "votes": null,
      "replies": [
        {
          "id": 214239,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "08/16/2017 10:53:59",
          "content": "<p>I don't think you can get good results efficiently with cpu. I think gpu is a must.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 214596,
          "author_name": "sheriytm",
          "author_url": "",
          "post_date": "08/17/2017 14:33:27",
          "content": "<p>Thanks @Heng,</p>\n\n<p>I am thinking of trying <a href=\"https://www.floydhub.com/\">Floyd</a>. Has anybody here used it? Any advise on the best way to set it up for this contest? I briefly checked their website but will register in a few days when I get a chance.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 214661,
          "author_name": "rhgrossm",
          "author_url": "",
          "post_date": "08/17/2017 18:47:18",
          "content": "<p>I've used Floyd (but not for Kaggle). It's super intuitive, and easy to use, but make sure you test your code on a local machine before bringing it onto the cloud to save time and money. Also, they don't handle generators as well as they could- their I/O is slower than desired. As much as you can, limit I/O and do things in memory. </p>\n\n<p>GCP also offers a $300 dollar credit if you haven't used it for a project before- might be worth checking out. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 214681,
          "author_name": "brasnold",
          "author_url": "",
          "post_date": "08/17/2017 20:38:59",
          "content": "<p>I use a GPU VM on Google cloud platform that gives 300$ free trial to any new user. Once on the cloud you have two ways to do machine learning: </p>\n\n<ul>\n<li><p>Using a virtual machine (a google cloud computing engine instance). Simple as using the local machine. The cost is based on the machine you've chosed and the time for which it was used.</p></li>\n<li><p>Serverless, via ML engine, the google service provided with python Tensorflow API. In this case you pay only according to the \"job\" that you are required. It's the most innovative and appropriate way to do ML in cloud, but it's a bit more difficult and still I don't manage how to build my package for its API.</p></li>\n</ul>\n\n<p>I suggest you to start with the first option, create a VM, install what you want, use Google Storage browser API to load your python file and the kaggle-cli package ( <a href=\"https://github.com/floydwch/kaggle-cli\">https://github.com/floydwch/kaggle-cli</a> ) to load the data-set. </p>\n\n<p>Always remember to stop or delete the VM after you've used it, try to do the bulk of the debugging of your code locally to avoid wasting money, and the 300$ free should be enough to do many things.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 214969,
          "author_name": "sheriytm",
          "author_url": "",
          "post_date": "08/19/2017 00:51:12",
          "content": "<p>@To Train Them, thanks for your useful suggestions, I will try Floyd as I have used the GCP credits on another project.</p>\n\n<p>@Brasnold, thanks for your kind advise and suggestions.</p>\n\n<p>For the people who down voted my asking this question. I do not understand what in my asking a question would rub anyone the wrong way to prompt such strong response. Can someone tell me what is wrong in asking a question of a kind and generous person like @Heng who I have communicated with in the past via other competitions? I believe we all have learnt from a lot of generous Kagglers during these contests by having discussions.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 214973,
          "author_name": "wangshuo0225",
          "author_url": "",
          "post_date": "08/19/2017 01:33:44",
          "content": "<p>here is a very detailed google cloud platform tutorial from cs231n (<a href=\"http://cs231n.github.io/gce-tutorial/\">http://cs231n.github.io/gce-tutorial/</a>).  I followed the steps but got stuck on requiring gpu resources, good luck.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 215263,
          "author_name": "firolino",
          "author_url": "",
          "post_date": "08/20/2017 19:46:56",
          "content": "<p>@YaGana, why there are down votes: This is not the right place for your question. This thread isn't about resources and cloud services. First, you are polluting this thread with your question and second, you should have opened a new thread or post it into a respective one. There might be other users with the same question as well and they won't find it here. For the sake of sharing and finding knowledge post thread-unrelated questions into their own threads.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 215356,
          "author_name": "sheriytm",
          "author_url": "",
          "post_date": "08/21/2017 09:35:36",
          "content": "<p>@firolino thank you. My question is more directed to Heng but I get your point. As a novice, I cannot ask or message people directly yet. I am still fairly new to Kaggle. I wish the 1st person that down voted instead pointed that out. </p>\n\n<p>I will post it as an independent thread for others to contribute or learn from. I apologize everyone for hijacking the thread. </p>\n\n<p>If you want to respond to my question with suggestions please do it here (<a href=\"https://www.kaggle.com/c/carvana-image-masking-challenge/discussion/38375\">https://www.kaggle.com/c/carvana-image-masking-challenge/discussion/38375</a>)</p>\n\n<p>@Max thanks.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 214215,
      "author_name": "sayantan",
      "author_url": "",
      "post_date": "08/16/2017 08:45:29",
      "content": "<p>Thanks for this wonderful post. Learn't many topics related to segmentation from your posts. I have one question regarding \"Train sample augmentation\". What sort of augmentation is normally required, I understand it varies from subject to subject, in general are there any list of things that we can try? My knowledge about image processing is limited. If you have any reference to read / suggestion it will be helpful. If I am asking too much information, that might create problems for the competition then there is no need to divulge the information now. I will wait for the competition to end and for more details to come out. It been a pleasure and great learning reading your posts on this forum and also of a lot of other Kagglers who have greatly contributed to my learning..\nThank you Kaggle and Kagglers..</p>",
      "votes": null,
      "replies": [
        {
          "id": 214324,
          "author_name": "heyt0ny",
          "author_url": "",
          "post_date": "08/16/2017 14:38:33",
          "content": "<p>shifting, horizontally flipping, random crops, scaling, rotating, color jittering.  A good summary on CNN tricks (including data augmentation): <a href=\"http://lamda.nju.edu.cn/weixs/project/CNNTricks/CNNTricks.html\">http://lamda.nju.edu.cn/weixs/project/CNNTricks/CNNTricks.html</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 214500,
          "author_name": "sayantan",
          "author_url": "",
          "post_date": "08/17/2017 06:54:27",
          "content": "<p>Thanks..</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 214271,
      "author_name": "gadgysaidoff",
      "author_url": "",
      "post_date": "08/16/2017 13:13:51",
      "content": "<p>Hi Heng!</p>\n\n<p>Can you please clarify how many \"poolings\" are you using in your Unet for 1024x1024?\nI now use 4 pooling layers for 480x320, and it seems to be too small. So to save structure of network I should be use 6 pooling network for 1918x1280</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 214332,
      "author_name": "rrqqmm",
      "author_url": "",
      "post_date": "08/16/2017 15:13:28",
      "content": "<pre><code>class SGDAccum(Optimizer):\n\"\"\"Stochastic gradient descent optimizer.\nIncludes support for momentum,\nlearning rate decay, and Nesterov momentum.\n# Arguments\n    lr: float &gt;= 0. Learning rate.\n    momentum: float &gt;= 0. Parameter updates momentum.\n    decay: float &gt;= 0. Learning rate decay over each update.\n    nesterov: boolean. Whether to apply Nesterov momentum.\n\"\"\"\n\ndef __init__(self, lr=0.01, momentum=0., decay=0.,\n             nesterov=False, accum_iters=5, **kwargs):\n    super(SGDAccum, self).__init__(**kwargs)\n    with K.name_scope(self.__class__.__name__):\n        self.iterations = K.variable(0., name='iterations')\n        self.lr = K.variable(lr, name='lr')\n        self.momentum = K.variable(momentum, name='momentum')\n        self.decay = K.variable(decay, name='decay')\n        self.accum_iters = K.variable(accum_iters, name='accum_iters')\n    self.initial_decay = decay\n    self.nesterov = nesterov\n\n@interfaces.legacy_get_updates_support\ndef get_updates(self, loss, params):\n    grads = self.get_gradients(loss, params)\n    self.updates = []\n\n    lr = self.lr\n    if self.initial_decay &gt; 0:\n        lr *= (1. / (1. + self.decay * self.iterations))\n        self.updates.append(K.update_add(self.iterations, 1))\n\n    # momentum\n    shapes = [K.get_variable_shape(p) for p in params]\n    moments = [K.zeros(shape) for shape in shapes]\n    gradients = [K.zeros(shape) for shape in shapes]\n\n    self.weights = [self.iterations] + moments\n\n    for p, g, m, gg in zip(params, grads, moments, gradients):\n\n        flag = K.equal(self.iterations % self.accum_iters, 0)\n\n        if K.eval(flag):\n            flag = 1\n\n        gg_t = (1 - flag) * (gg + g)\n\n        v = self.momentum * m - lr * \\\n            K.square((gg + flag * g) / self.accum_iters)  # velocity\n\n        self.updates.append(K.update(m, flag * v + (1 - flag) * m))\n        self.updates.append((gg, gg_t))\n\n        if self.nesterov:\n            new_p = p + self.momentum * v - lr * g\n        else:\n            new_p = p + v\n\n        # Apply constraints.\n        if getattr(p, 'constraint', None) is not None:\n            new_p = p.constraint(new_p)\n\n        self.updates.append(K.update(p, new_p))\n\n    return self.updates\n\ndef get_config(self):\n    config = {'lr': float(K.get_value(self.lr)),\n              'momentum': float(K.get_value(self.momentum)),\n              'decay': float(K.get_value(self.decay)),\n              'nesterov': self.nesterov}\n    base_config = super(SGDAccum, self).get_config()\n    return dict(list(base_config.items()) + list(config.items()))\n</code></pre>",
      "votes": null,
      "replies": [
        {
          "id": 214333,
          "author_name": "rrqqmm",
          "author_url": "",
          "post_date": "08/16/2017 15:14:01",
          "content": "<p>A small hack for SGD Accumulator for keras, lemme know if it works.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 214363,
          "author_name": "svesal",
          "author_url": "",
          "post_date": "08/16/2017 16:34:04",
          "content": "<p>What changes you made in the original SGD?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 214377,
          "author_name": "rrqqmm",
          "author_url": "",
          "post_date": "08/16/2017 17:40:06",
          "content": "<p>This will accumulate the gradients and update the weights on accum_iters'th  iteration.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 214521,
          "author_name": "zhangsongwei",
          "author_url": "",
          "post_date": "08/17/2017 08:19:15",
          "content": "<p>Thanks for your sharing. The parameter <code>accum_iters</code> is 5. If the previous batch_size=4, the final batch_size=32 when set the parameter? Could you show how do you calculate that? \nThanks!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 214673,
          "author_name": "xiaokangwang",
          "author_url": "",
          "post_date": "08/17/2017 20:17:18",
          "content": "<p>Hi Rohit. Thanks for sharing the SGD hacker. In the code you used a decorator, \"interfaces.legacy_get_updates_support\". Where is it defined?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 214879,
          "author_name": "rrqqmm",
          "author_url": "",
          "post_date": "08/18/2017 16:48:41",
          "content": "<p>Final batch size will be accum_iters*batch_size (you defined it to be 4 so, final=20)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 214880,
          "author_name": "rrqqmm",
          "author_url": "",
          "post_date": "08/18/2017 16:50:05",
          "content": "<p>\"interfaces.legacy_get_updates_support\" in interfaces.py script of keras, this code needs to be pasted in optimizers.py in keras.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 215617,
          "author_name": "zhangsongwei",
          "author_url": "",
          "post_date": "08/22/2017 10:47:09",
          "content": "<p>Sorry, it seems that it doesn't work in my train.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 216098,
          "author_name": "hw199212291",
          "author_url": "",
          "post_date": "08/24/2017 09:23:28",
          "content": "<p>if self.initial_decay is 0, then self.iterations doesn't make change?</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 215212,
      "author_name": "cbaldie",
      "author_url": "",
      "post_date": "08/20/2017 14:08:09",
      "content": "<p>Thank you for clarity</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 215406,
      "author_name": "strideradu",
      "author_url": "",
      "post_date": "08/21/2017 13:54:36",
      "content": "<p>Just want to know how can I accumulate the gradient before update loss?</p>\n\n<p>So that means only every 5 steps I should call optimizer.step() and optimizer.zero_grad()?</p>\n\n<p>Is there any other thing I need to pay attention?</p>",
      "votes": null,
      "replies": [
        {
          "id": 215407,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "08/21/2017 14:00:33",
          "content": "<p>for example:</p>\n\n<pre><code> optimizer.zero_grad() \n\n x,y_hat =  .... get one batch of data\n y = net(x)\n loss.backward(y, y_hat)  #accumulate gradient  for 1 times \n\n x,y_hat =  .... get one batch of data\n y = net(x)\n loss.backward(y, y_hat)  #accumulate gradient  for 2 times \n\n x,y_hat=  .... get one batch of data\n y = net(x)\n loss.backward(y, y_hat)  #accumulate gradient  for 3 times \n\n optimizer.step()  #update accumulated gradients\n\n print(loss.data[0]) #print only loss for lass batch , not accumulated loss \n</code></pre>",
          "votes": null,
          "replies": []
        },
        {
          "id": 215417,
          "author_name": "strideradu",
          "author_url": "",
          "post_date": "08/21/2017 15:08:23",
          "content": "<p>Thanks!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 216624,
          "author_name": "",
          "author_url": "",
          "post_date": "08/27/2017 02:25:20",
          "content": "<p>Could you tell me is the accumulate gradient's method is simply add the gradients together or average evergytime gradient ? If it just add togather I think maybe it is no benefit.  </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 216628,
          "author_name": "stevenknguyen",
          "author_url": "",
          "post_date": "08/27/2017 03:03:22",
          "content": "<p>If you add together the gradients, it effectively gives you a larger batch size, which does help when everyone is optimizing for the 4-5th digit. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 218079,
          "author_name": "harungunaydin",
          "author_url": "",
          "post_date": "09/02/2017 05:44:03",
          "content": "<p>Do you keep the accumulated gradients in the RAM? What if the effective size of all GPUs combined is equal to the size of RAM?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 218084,
          "author_name": "stevenknguyen",
          "author_url": "",
          "post_date": "09/02/2017 06:22:21",
          "content": "<p>I know they're stored in RAM for tensorflow, but this is a workaround if you only have a single GPU. I'm sure if you have enough vram to have a batch size of 15 for 1024x1024 images the results would be slightly better.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 218241,
      "author_name": "mshliselberg",
      "author_url": "",
      "post_date": "09/03/2017 03:54:53",
      "content": "<p>So I am also using pascal titan-X and when using batch size of 2 with pixel budget &lt;1024^2 It takes me still ~1 hour to train per epoch, how did you manage 15 minutes. I understand your using accumulation but i thought that doesnt really help speedup as mentioned in the comments here (thanks to @Steven Nguyen). Is there anything else you have done for speedup?</p>",
      "votes": null,
      "replies": [
        {
          "id": 218243,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "09/03/2017 03:58:23",
          "content": "<p>maybe your network structure is different from mine.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "213810": "UPDATED! the software and model for producing this results can be found at:\nhttps://www.kaggle.com/c/carvana-image-masking-challenge/discussion/37208\n\n.\n\nHere are some tips on how to move to 0.997.\n\n.\n\n\n**(1) Which deep learning framework to use?**\n\nBoth unet imlementation in Keras/tensorflow and pytorch are good starters:\n\nhttps://www.kaggle.com/c/carvana-image-masking-challenge/discussion/37523\n\nhttps://www.kaggle.com/c/carvana-image-masking-challenge/discussion/37208\n\n.\n\n**(2) Which CNN model to use?**\n\nUnet is sufficient. However, do experiment with different depths, how many convolution filters to use per scale, etc? If you do not get the parameters right, you do not get the best performance.\n\n.\n\n **(3) Which image resolution to use?**\n\nI would suggest 1024x1024. (this is what i use for my 0.997 solution)\n\n.\n\n **(4) Generalization and validation/training split**\n\nAny good split is ok, because the train and test dataset are very similar. But be careful about the validation error during training. There are some ground truth error, please see:\n\nhttps://www.kaggle.com/c/carvana-image-masking-challenge/discussion/37229\n\nDo examine the validation error (e.g. by visual inspection) and make sure high error are not due to ground truth error. If you can get training/validation error of near 0.997, you should get 0.997 at the LB.\n\nNote that the LB page may not show the correct ranking. A better way to judge if your new submission is better then previous ones is to use the 'sort by public score' function on your submission page, please see:\n\nhttps://www.kaggle.com/c/carvana-image-masking-challenge/discussion/37137\n\n.\n\n\n\n\n**(5) Training hyper parameters**\n\nTry different learning rates. I use about 45 epoch. Try different batch size. I use 3 images per batch, and accumulate 5 batches before updating the gradient. i.e. my effective batch size is 3x5 = 15. I just use SDG with momentum with manual rate scheduling.\n\n.\n\n**(6) Ensemble and test-time augmentation**\n\nThey may improvement results. But I did not use them in my 0.997 solution. \n\n.\n\n**(7) Pre or post processing?**\n\nThey may improvement results. But I did not them. Just a simple CNN prediction can give 0.997. I use cv2.INTER_LINEAR for downsize and upsize.\n\n.\n\n**(6) Train sample augmentation**\n\nIn my experiments, some train augmentation is required. However, do be careful. Too much augmentation actually reduces accuracy.\n\n.\n\n **(7) Loss function for back propagation?**\n\nUse binary cross entropy and dice loss. Weigh pixels near boundary. \n\n.\n\n **(8) Use pretrained model?**\n\nI do not use. I train from sratch.\n\n.\n\n **(9) Any other tricks?**\n\nNo. Just tune your learning carefully. Do proper experiments and record your results. Make your work process efficient. Due to large image image size and huge number of test images (100064), it does take some time to generate results.  \n\nmy machine: pascal titan-X gpu. Train time per epoch = 15 min (11.5 hours in total training).  Time to make submission csv file = 1.5 hr to make prediction,  35 min to encode rle and save csv.",
    "213813": "Good summary.\nOne question about 3): 1024x1024 rescaled to 1280x1918 or sliced into 1024x1024 pieces?",
    "213814": "1024x1024 rescaled to 1280x1918. Only one-step resize, no cropping.",
    "213815": "May also be worth testing variations of the Unet architecture, like replacing the VGG-like architecture with residual or inception blocks.  Not necessary for 0.997 but could be for 0.998 :)",
    "213842": "0.996 ers are looking at you",
    "213857": "What lr-scheduler do you use for optimization?",
    "213860": "note that this has to depend on your network, data augment, data batch size, loss,  etc. What works for me may not work for you.\n\nFor me, i use conv-bn-relu. for epoch 0 to 40: 0.01,  40 to 45: 0.005, 45 to 47: 0.001. My training loss is as attached:\n\n  \n       LR=Step Learning Rates\n       rates=[' 0.0100', ' 0.0050', ' 0.0010', '-1.0000', '-1.0000']\n       steps=['      0', '     35', '     40', '     42', '     44']\n\n     epoch    iter      rate   | valid_loss/acc | train_loss/acc | batch_loss/acc ...             \n     ---------------------------------------------------------------------------------------------\n       1.0    1440    0.0100   | 0.0670  0.9900 | 0.0807  0.9883 | 0.0897  0.9885  |  14.8 min    \n       2.0    1440    0.0100   | 0.0470  0.9937 | 0.0566  0.9922 | 0.0572  0.9918  |  14.7 min    \n       3.0    1440    0.0100   | 0.0380  0.9948 | 0.0496  0.9929 | 0.0369  0.9949  |  14.7 min    \n       4.0    1440    0.0100   | 0.0400  0.9944 | 0.0414  0.9944 | 0.0504  0.9934  |  14.7 min    \n       5.0    1440    0.0100   | 0.0317  0.9956 | 0.0400  0.9946 | 0.0821  0.9896  |  14.7 min    \n       6.0    1440    0.0100   | 0.0310  0.9956 | 0.0401  0.9947 | 0.0391  0.9950  |  14.7 min    \n       7.0    1440    0.0100   | 0.0314  0.9955 | 0.0397  0.9946 | 0.0364  0.9954  |  14.7 min    \n       8.0    1440    0.0100   | 0.0286  0.9959 | 0.0344  0.9952 | 0.0267  0.9959  |  14.7 min    \n       9.0    1440    0.0100   | 0.0286  0.9959 | 0.0375  0.9951 | 0.0296  0.9953  |  14.7 min    \n      10.0    1440    0.0100   | 0.0284  0.9959 | 0.0321  0.9956 | 0.0435  0.9948  |  14.9 min    \n      11.0    1440    0.0100   | 0.0280  0.9960 | 0.0325  0.9954 | 0.0253  0.9967  |  15.1 min    \n      12.0    1440    0.0100   | 0.0270  0.9961 | 0.0349  0.9950 | 0.0352  0.9945  |  14.9 min    \n      13.0    1440    0.0100   | 0.0266  0.9962 | 0.0335  0.9953 | 0.0296  0.9952  |  14.9 min    \n      14.0    1440    0.0100   | 0.0263  0.9962 | 0.0324  0.9955 | 0.0287  0.9962  |  15.1 min    \n      15.0    1440    0.0100   | 0.0272  0.9961 | 0.0292  0.9959 | 0.0332  0.9950  |  14.9 min    \n      16.0    1440    0.0100   | 0.0269  0.9961 | 0.0288  0.9960 | 0.0228  0.9966  |  15.0 min    \n      17.0    1440    0.0100   | 0.0263  0.9962 | 0.0332  0.9954 | 0.0256  0.9964  |  14.8 min    \n      18.0    1440    0.0100   | 0.0254  0.9963 | 0.0283  0.9962 | 0.0498  0.9916  |  14.9 min    \n      19.0    1440    0.0100   | 0.0249  0.9964 | 0.0285  0.9960 | 0.0248  0.9963  |  14.8 min    \n      20.0    1440    0.0100   | 0.0247  0.9964 | 0.0277  0.9961 | 0.0314  0.9952  |  14.9 min    \n      21.0    1440    0.0100   | 0.0248  0.9964 | 0.0274  0.9961 | 0.0240  0.9961  |  15.0 min    \n      22.0    1440    0.0100   | 0.0251  0.9963 | 0.0252  0.9965 | 0.0214  0.9969  |  14.9 min    \n      23.0    1440    0.0100   | 0.0252  0.9963 | 0.0284  0.9961 | 0.0343  0.9952  |  14.9 min    \n      24.0    1440    0.0100   | 0.0243  0.9965 | 0.0261  0.9963 | 0.0246  0.9966  |  14.8 min    \n      25.0    1440    0.0100   | 0.0248  0.9964 | 0.0279  0.9959 | 0.0197  0.9967  |  14.6 min    \n      26.0    1440    0.0100   | 0.0244  0.9964 | 0.0269  0.9961 | 0.0366  0.9952  |  14.6 min    \n      27.0    1440    0.0100   | 0.0240  0.9965 | 0.0275  0.9962 | 0.0273  0.9960  |  14.5 min    \n      28.0    1440    0.0100   | 0.0245  0.9964 | 0.0251  0.9964 | 0.0340  0.9953  |  14.5 min    \n      29.0    1440    0.0100   | 0.0244  0.9965 | 0.0281  0.9961 | 0.0220  0.9967  |  14.5 min    \n      30.0    1440    0.0100   | 0.0241  0.9965 | 0.0298  0.9958 | 0.0364  0.9954  |  14.5 min    \n      31.0    1440    0.0100   | 0.0235  0.9966 | 0.0240  0.9965 | 0.0273  0.9962  |  14.5 min    \n      32.0    1440    0.0100   | 0.0237  0.9965 | 0.0261  0.9963 | 0.0257  0.9962  |  14.5 min    \n      33.0    1440    0.0100   | 0.0235  0.9966 | 0.0240  0.9965 | 0.0191  0.9972  |  14.5 min    \n      34.0    1440    0.0100   | 0.0234  0.9966 | 0.0237  0.9965 | 0.0332  0.9955  |  14.5 min    \n\n      make some change to dataset : reduce augmentation\n      ... stop and resume training ....\n\n      LR=Step Learning Rates\n     rates=[' 0.0100', ' 0.0050', ' 0.0010', '-1.0000', '-1.0000']\n     steps=['      0', '     40', '     45', '     47', '     44']\n\n      epoch    iter      rate   | valid_loss/acc | train_loss/acc | batch_loss/acc ... \n      --------------------------------------------------------------------------------------------------\n       34.0    1440    0.0100   | 0.0231  0.9966 | 0.0230  0.9969 | 0.0262  0.9966  |  14.6 min \n       35.0    1440    0.0100   | 0.0234  0.9966 | 0.0216  0.9969 | 0.0247  0.9963  |  14.8 min \n       36.0    1440    0.0100   | 0.0231  0.9966 | 0.0214  0.9970 | 0.0220  0.9967  |  14.8 min \n       37.0    1440    0.0100   | 0.0238  0.9965 | 0.0216  0.9971 | 0.0281  0.9963  |  14.8 min \n       38.0    1440    0.0100   | 0.0228  0.9967 | 0.0237  0.9967 | 0.0247  0.9965  |  14.8 min \n       39.0    1440    0.0100   | 0.0226  0.9967 | 0.0230  0.9969 | 0.0188  0.9973  |  15.0 min \n       40.0    1440    0.0100   | 0.0230  0.9966 | 0.0224  0.9969 | 0.0231  0.9971  |  15.4 min \n       41.0    1440    0.0050   | 0.0224  0.9967 | 0.0224  0.9970 | 0.0180  0.9975  |  15.4 min \n       42.0    1440    0.0050   | 0.0225  0.9967 | 0.0218  0.9970 | 0.0216  0.9968  |  15.4 min \n       43.0    1440    0.0050   | 0.0224  0.9967 | 0.0202  0.9973 | 0.0214  0.9972  |  15.4 min \n       44.0    1440    0.0050   | 0.0224  0.9967 | 0.0196  0.9972 | 0.0172  0.9977  |  14.8 min",
    "213920": "You said: \"Accumulate batches before updating\". How do you do it in keras?\n\n(5) Training hyper parameters\n\nTry different learning rates. I use about 45 epoch. Try different batch size. I use 3 images per batch, and **accumulate 5 batches** before updating the gradient. i.e. my effective batch size is 3x5 = 15. I just use SDG with momentum with manual rate scheduling.",
    "213921": "Heng CherKeng\nYou said \"accumulate batches before updating\":\n&gt;(5) Training hyper parameters\nTry different learning rates. I use about 45 epoch. Try different batch size. I use 3 images per batch, and **accumulate 5 batches** before updating the gradient. i.e. my effective batch size is 3x5 = 15. I just use SDG with momentum with manual rate scheduling.\n\nHow do you do it in Keras?",
    "213997": "There is a way to do it here: https://github.com/fchollet/keras/issues/3556",
    "214027": "You da best. One question though. What did you mean by \"Use binary cross entropy and dice loss. Weigh pixels near boundary.\"?",
    "214062": "check the dicsussion at: https://www.kaggle.com/c/carvana-image-masking-challenge/discussion/37208",
    "214159": "Thanks a lot for sharing. Curious what caused the unstable validation scores in https://www.kaggle.com/c/carvana-image-masking-challenge/discussion/37208, and what did you do to fix it in this training?",
    "214166": "Heng CherKeng here we meet again. \n\nHi everyone, any advise for the resource challenged, I meant CPU users not GPU for this competition?",
    "214215": "Thanks for this wonderful post. Learn't many topics related to segmentation from your posts. I have one question regarding \"Train sample augmentation\". What sort of augmentation is normally required, I understand it varies from subject to subject, in general are there any list of things that we can try? My knowledge about image processing is limited. If you have any reference to read / suggestion it will be helpful. If I am asking too much information, that might create problems for the competition then there is no need to divulge the information now. I will wait for the competition to end and for more details to come out. It been a pleasure and great learning reading your posts on this forum and also of a lot of other Kagglers who have greatly contributed to my learning..\nThank you Kaggle and Kagglers..",
    "214237": "Follow Peter's post and wx405557858's suggestion, you can replace Adam in Keras with this :)\n\n    class Adam(Optimizer):\n    \"\"\"Adam optimizer.\n\n    Default parameters follow those provided in the original paper.\n\n    # Arguments\n        lr: float &gt;= 0. Learning rate.\n        beta_1: float, 0 &lt; beta &lt; 1. Generally close to 1.\n        beta_2: float, 0 &lt; beta &lt; 1. Generally close to 1.\n        epsilon: float &gt;= 0. Fuzz factor.\n        decay: float &gt;= 0. Learning rate decay over each update.\n\n    # References\n        - [Adam - A Method for Stochastic Optimization](http://arxiv.org/abs/1412.6980v8)\n    \"\"\"\n\n    def __init__(self, lr=0.001, beta_1=0.9, beta_2=0.999,\n                 epsilon=1e-8, decay=0., accumulator=5., **kwargs):\n        super(Adam, self).__init__(**kwargs)\n        self.iterations = K.variable(0, name='iterations')\n        self.lr = K.variable(lr, name='lr')\n        self.beta_1 = K.variable(beta_1, name='beta_1')\n        self.beta_2 = K.variable(beta_2, name='beta_2')\n        self.epsilon = epsilon\n        self.decay = K.variable(decay, name='decay')\n        self.initial_decay = decay\n        self.accumulator = K.variable(accumulator, name='accumulator')\n\n    def get_updates(self, params, constraints, loss):\n        grads = self.get_gradients(loss, params)\n        self.updates = [K.update_add(self.iterations, 1)]\n\n        lr = self.lr\n        if self.initial_decay &gt; 0:\n            lr *= (1. / (1. + self.decay * self.iterations))\n\n        t = self.iterations + 1\n        lr_t = lr * (K.sqrt(1. - K.pow(self.beta_2, t)) /\n                     (1. - K.pow(self.beta_1, t)))\n\n        ms = [K.zeros(K.get_variable_shape(p), dtype=K.dtype(p)) for p in params]\n        vs = [K.zeros(K.get_variable_shape(p), dtype=K.dtype(p)) for p in params]\n        gs = [K.zeros(K.get_variable_shape(p), dtype=K.dtype(p)) for p in params]\n\n        self.weights = [self.iterations] + ms + vs\n\n        for p, g, m, v, ga in zip(params, grads, ms, vs, gs):\n\n\n            flag = K.equal(self.iterations % self.accumulator, 0)\n            flag = K.cast(flag, dtype='float32')\n\n            ga_t = (1 - flag) * (ga + g)\n\n            m_t = (self.beta_1 * m) + (1. - self.beta_1) * (ga + flag * g) / self.accumulator\n            v_t = (self.beta_2 * v) + (1. - self.beta_2) * K.square((ga + flag * g) / self.accumulator)\n            p_t = p - lr_t * m_t / (K.sqrt(v_t) + self.epsilon)\n\n\n            self.updates.append(K.update(m, flag * m_t + (1 - flag) * m))\n            self.updates.append(K.update(v, flag * v_t + (1 - flag) * v))\n            self.updates.append(K.update(ga, ga_t))\n\n            new_p = p_t\n            # apply constraints\n            if p in constraints:\n                c = constraints[p]\n                new_p = c(new_p)\n            self.updates.append(K.update(p, new_p))\n        return self.updates\n\n    def get_config(self):\n        config = {'lr': float(K.get_value(self.lr)),\n                  'beta_1': float(K.get_value(self.beta_1)),\n                  'beta_2': float(K.get_value(self.beta_2)),\n                  'decay': float(K.get_value(self.decay)),\n                  'accumulator': float(K.get_value(self.accumulator)),\n                  'epsilon': self.epsilon}\n        base_config = super(Adam, self).get_config()\n        return dict(list(base_config.items()) + list(config.items()))",
    "214239": "I don't think you can get good results efficiently with cpu. I think gpu is a must.",
    "214241": "I think it is due to small batch size. But i cannot confirm this. It can also be due to wrong label. The fluctuating is solved by accumulating gradient for larger effective batch size",
    "214271": "Hi Heng!\n\nCan you please clarify how many \"poolings\" are you using in your Unet for 1024x1024?\nI now use 4 pooling layers for 480x320, and it seems to be too small. So to save structure of network I should be use 6 pooling network for 1918x1280",
    "214324": "shifting, horizontally flipping, random crops, scaling, rotating, color jittering.  A good summary on CNN tricks (including data augmentation): http://lamda.nju.edu.cn/weixs/project/CNNTricks/CNNTricks.html",
    "214332": "class SGDAccum(Optimizer):\n    \"\"\"Stochastic gradient descent optimizer.\n    Includes support for momentum,\n    learning rate decay, and Nesterov momentum.\n    # Arguments\n        lr: float &gt;= 0. Learning rate.\n        momentum: float &gt;= 0. Parameter updates momentum.\n        decay: float &gt;= 0. Learning rate decay over each update.\n        nesterov: boolean. Whether to apply Nesterov momentum.\n    \"\"\"\n\n    def __init__(self, lr=0.01, momentum=0., decay=0.,\n                 nesterov=False, accum_iters=5, **kwargs):\n        super(SGDAccum, self).__init__(**kwargs)\n        with K.name_scope(self.__class__.__name__):\n            self.iterations = K.variable(0., name='iterations')\n            self.lr = K.variable(lr, name='lr')\n            self.momentum = K.variable(momentum, name='momentum')\n            self.decay = K.variable(decay, name='decay')\n            self.accum_iters = K.variable(accum_iters, name='accum_iters')\n        self.initial_decay = decay\n        self.nesterov = nesterov\n\n    @interfaces.legacy_get_updates_support\n    def get_updates(self, loss, params):\n        grads = self.get_gradients(loss, params)\n        self.updates = []\n\n        lr = self.lr\n        if self.initial_decay &gt; 0:\n            lr *= (1. / (1. + self.decay * self.iterations))\n            self.updates.append(K.update_add(self.iterations, 1))\n\n        # momentum\n        shapes = [K.get_variable_shape(p) for p in params]\n        moments = [K.zeros(shape) for shape in shapes]\n        gradients = [K.zeros(shape) for shape in shapes]\n\n        self.weights = [self.iterations] + moments\n\n        for p, g, m, gg in zip(params, grads, moments, gradients):\n\n            flag = K.equal(self.iterations % self.accum_iters, 0)\n\n            if K.eval(flag):\n                flag = 1\n\n            gg_t = (1 - flag) * (gg + g)\n\n            v = self.momentum * m - lr * \\\n                K.square((gg + flag * g) / self.accum_iters)  # velocity\n\n            self.updates.append(K.update(m, flag * v + (1 - flag) * m))\n            self.updates.append((gg, gg_t))\n\n            if self.nesterov:\n                new_p = p + self.momentum * v - lr * g\n            else:\n                new_p = p + v\n\n            # Apply constraints.\n            if getattr(p, 'constraint', None) is not None:\n                new_p = p.constraint(new_p)\n\n            self.updates.append(K.update(p, new_p))\n\n        return self.updates\n\n    def get_config(self):\n        config = {'lr': float(K.get_value(self.lr)),\n                  'momentum': float(K.get_value(self.momentum)),\n                  'decay': float(K.get_value(self.decay)),\n                  'nesterov': self.nesterov}\n        base_config = super(SGDAccum, self).get_config()\n        return dict(list(base_config.items()) + list(config.items()))",
    "214333": "A small hack for SGD Accumulator for keras, lemme know if it works.",
    "214363": "What changes you made in the original SGD?",
    "214377": "This will accumulate the gradients and update the weights on accum_iters'th  iteration.",
    "214500": "Thanks..",
    "214509": "Hi, Kaggoo, thanks for your sharing. Could you please the mean of ```accumulator```?  Are the parameters  updated after 32 samples if we set ```accumulator=5```? Could you please show how to calculate it? The batch_size was 4. Thank you.",
    "214521": "Thanks for your sharing. The parameter ```accum_iters``` is 5. If the previous batch_size=4, the final batch_size=32 when set the parameter? Could you show how do you calculate that? \nThanks!",
    "214596": "Thanks @Heng,\n\nI am thinking of trying [Floyd](https://www.floydhub.com/). Has anybody here used it? Any advise on the best way to set it up for this contest? I briefly checked their website but will register in a few days when I get a chance.",
    "214661": "I've used Floyd (but not for Kaggle). It's super intuitive, and easy to use, but make sure you test your code on a local machine before bringing it onto the cloud to save time and money. Also, they don't handle generators as well as they could- their I/O is slower than desired. As much as you can, limit I/O and do things in memory. \n\n GCP also offers a $300 dollar credit if you haven't used it for a project before- might be worth checking out.",
    "214673": "Hi Rohit. Thanks for sharing the SGD hacker. In the code you used a decorator, \"interfaces.legacy_get_updates_support\". Where is it defined?",
    "214681": "I use a GPU VM on Google cloud platform that gives 300$ free trial to any new user. Once on the cloud you have two ways to do machine learning: \n\n- Using a virtual machine (a google cloud computing engine instance). Simple as using the local machine. The cost is based on the machine you've chosed and the time for which it was used.\n\n- Serverless, via ML engine, the google service provided with python Tensorflow API. In this case you pay only according to the \"job\" that you are required. It's the most innovative and appropriate way to do ML in cloud, but it's a bit more difficult and still I don't manage how to build my package for its API.\n\nI suggest you to start with the first option, create a VM, install what you want, use Google Storage browser API to load your python file and the kaggle-cli package ( https://github.com/floydwch/kaggle-cli ) to load the data-set. \n\nAlways remember to stop or delete the VM after you've used it, try to do the bulk of the debugging of your code locally to avoid wasting money, and the 300$ free should be enough to do many things.",
    "214709": "Hi Heng. Thanks for sharing your hyperparameter setting. When you trained your networking using different learning rate. Did you stop the model.fit(), change the learning rate and start again? Or is there a way to define multiple learning rates by hooking the learning rate to the number of epochs?",
    "214879": "Final batch size will be accum_iters*batch_size (you defined it to be 4 so, final=20)",
    "214880": "\"interfaces.legacy_get_updates_support\" in interfaces.py script of keras, this code needs to be pasted in optimizers.py in keras.",
    "214969": "To Train Them, thanks for your useful suggestions, I will try Floyd as I have used the GCP credits on another project.\n\n@Brasnold, thanks for your kind advise and suggestions.\n\nFor the people who down voted my asking this question. I do not understand what in my asking a question would rub anyone the wrong way to prompt such strong response. Can someone tell me what is wrong in asking a question of a kind and generous person like @Heng who I have communicated with in the past via other competitions? I believe we all have learnt from a lot of generous Kagglers during these contests by having discussions.",
    "214973": "here is a very detailed google cloud platform tutorial from cs231n (http://cs231n.github.io/gce-tutorial/).  I followed the steps but got stuck on requiring gpu resources, good luck.",
    "215147": "Thanks Heng.  What does it mean when learning rate becomes negative?\n\n&gt;&gt; rates=[' 0.0100', ' 0.0050', ' 0.0010', '-1.0000', '-1.0000']",
    "215152": "It gets terminated",
    "215212": "Thank you for clarity",
    "215263": "YaGana, why there are down votes: This is not the right place for your question. This thread isn't about resources and cloud services. First, you are polluting this thread with your question and second, you should have opened a new thread or post it into a respective one. There might be other users with the same question as well and they won't find it here. For the sake of sharing and finding knowledge post thread-unrelated questions into their own threads.",
    "215356": "firolino thank you. My question is more directed to Heng but I get your point. As a novice, I cannot ask or message people directly yet. I am still fairly new to Kaggle. I wish the 1st person that down voted instead pointed that out. \n\nI will post it as an independent thread for others to contribute or learn from. I apologize everyone for hijacking the thread. \n\nIf you want to respond to my question with suggestions please do it here (https://www.kaggle.com/c/carvana-image-masking-challenge/discussion/38375)\n\n@Max thanks.",
    "215406": "Just want to know how can I accumulate the gradient before update loss?\n\nSo that means only every 5 steps I should call optimizer.step() and optimizer.zero_grad()?\n\nIs there any other thing I need to pay attention?",
    "215407": "for example:\n\n \n     optimizer.zero_grad() \n\n     x,y_hat =  .... get one batch of data\n     y = net(x)\n     loss.backward(y, y_hat)  #accumulate gradient  for 1 times \n\n     x,y_hat =  .... get one batch of data\n     y = net(x)\n     loss.backward(y, y_hat)  #accumulate gradient  for 2 times \n\n     x,y_hat=  .... get one batch of data\n     y = net(x)\n     loss.backward(y, y_hat)  #accumulate gradient  for 3 times \n\n     optimizer.step()  #update accumulated gradients\n\n     print(loss.data[0]) #print only loss for lass batch , not accumulated loss",
    "215417": "Thanks!",
    "215617": "Sorry, it seems that it doesn't work in my train.",
    "216098": "if self.initial_decay is 0, then self.iterations doesn't make change?",
    "216495": "Note to everyone- accumulator has to be a float or it throws an error from internal calculations. Spent 2 hours looking for an implementation bug that didn't exist....",
    "216624": "Could you tell me is the accumulate gradient's method is simply add the gradients together or average evergytime gradient ? If it just add togather I think maybe it is no benefit.",
    "216628": "If you add together the gradients, it effectively gives you a larger batch size, which does help when everyone is optimizing for the 4-5th digit.",
    "216674": "I used flag = K.cast(flag, dtype='float32') to make sure the flag is float32 and thus the remaining formula should work. Tested on Keras with TF backend.",
    "216770": "I meant about the accumulator variable (5. instead of 5). But the flag is important also",
    "217777": "So how exactly do you use this, just create a file with it, and import it where the model is made? does it only depend on keras.backend and keras.optimizers.Optimizer? sorry, I just have never played around with custom optimizers in keras",
    "217779": "You have to find where keras is installed on your computer and add this to the file optimizers.py",
    "218035": "Do you get a speed up when increasing the accumulation here? I placed it in optimizers.py and it worked but no matter how much I increase the accumulation, it is still longer than normal adam. For example, normal adam with batch_size=2, takes for me ~55 minutes,  and with batch_size=2 and accumulation=100 on Adam_accumulate I get ~1 hour.",
    "218079": "Do you keep the accumulated gradients in the RAM? What if the effective size of all GPUs combined is equal to the size of RAM?",
    "218080": "No. If anything it slows it down. The goal is to increase batch size since there are limitations on VRAM.",
    "218084": "I know they're stored in RAM for tensorflow, but this is a workaround if you only have a single GPU. I'm sure if you have enough vram to have a batch size of 15 for 1024x1024 images the results would be slightly better.",
    "218097": "Steven, so if my training takes 12hrs, using this Adam, will result in may be twice the training time i.e. 20hrs but I can go from batch_size=10 to 12 without memory problem? These numbers are just examples as it all depends on one's setup. Am I right?",
    "218098": "Using the accumulator shouldn't double the training time. In case of @mshliselberg, the training duration increased from 55 mins to 60 mins (per epoch I guess). \nAlso, you don't incease your batch size, but accumulate the batches. So with a batch size of 10, you perform updates every 10*x samples (where x is an integer).",
    "218241": "So I am also using pascal titan-X and when using batch size of 2 with pixel budget &lt;1024^2 It takes me still ~1 hour to train per epoch, how did you manage 15 minutes. I understand your using accumulation but i thought that doesnt really help speedup as mentioned in the comments here (thanks to @Steven Nguyen). Is there anything else you have done for speedup?",
    "218243": "maybe your network structure is different from mine.",
    "218246": "Thanks  @Markus. I missread  @mshliselberg's comment."
  },
  "source": "meta"
}