{
  "id": 48293,
  "title": "From 0.959 LB to...",
  "url": "/competitions/sp-society-camera-model-identification/discussion/48293",
  "author_name": "",
  "post_date": "2018-01-25T21:04:49.125506600Z",
  "votes": 56,
  "comment_count": 136,
  "views": 0,
  "content": "<p>Got 0.959 LB with a single Densenet201 model trained in less than 1 day (my code is available on github, see other thread).</p>\n\n<p>So far implementation has no real novel approaches: just no hidden bugs in the code (I hope), use pre-trained model and Gleb's dataset; along with basic train, validation and test augmentation.</p>\n\n<p>For me the challenge is how to get to 0.97+ with a single model. Some ideas Im going to try:</p>\n\n<ul>\n<li>Loss/accuracy function that mimics the organization (right now I'm using global accuracy) </li>\n<li>Class-aware sampling for training </li>\n<li>Mixup</li>\n<li>Balanced validation set (assuming validation is equally balanced as\ntraining)</li>\n<li>More data</li>\n<li>New NN architecture/tuning</li>\n<li>Build confusion matrix to diagnose issues</li>\n</ul>\n\n<p>If you have any other ideas you're willing to share, feel free to comment.</p>",
  "messages": [
    {
      "id": "274067",
      "postDate": "01/25/2018 21:04:49",
      "content": "<p>Got 0.959 LB with a single Densenet201 model trained in less than 1 day (my code is available on github, see other thread).</p>\n\n<p>So far implementation has no real novel approaches: just no hidden bugs in the code (I hope), use pre-trained model and Gleb's dataset; along with basic train, validation and test augmentation.</p>\n\n<p>For me the challenge is how to get to 0.97+ with a single model. Some ideas Im going to try:</p>\n\n<ul>\n<li>Loss/accuracy function that mimics the organization (right now I'm using global accuracy) </li>\n<li>Class-aware sampling for training </li>\n<li>Mixup</li>\n<li>Balanced validation set (assuming validation is equally balanced as\ntraining)</li>\n<li>More data</li>\n<li>New NN architecture/tuning</li>\n<li>Build confusion matrix to diagnose issues</li>\n</ul>\n\n<p>If you have any other ideas you're willing to share, feel free to comment.</p>",
      "rawMarkdown": "Got 0.959 LB with a single Densenet201 model trained in less than 1 day (my code is available on github, see other thread).\n\nSo far implementation has no real novel approaches: just no hidden bugs in the code (I hope), use pre-trained model and Gleb's dataset; along with basic train, validation and test augmentation.\n\nFor me the challenge is how to get to 0.97+ with a single model. Some ideas Im going to try:\n\n - Loss/accuracy function that mimics the organization (right now I'm using global accuracy) \n - Class-aware sampling for training \n - Mixup\n - Balanced validation set (assuming validation is equally balanced as\n   training)\n - More data\n - New NN architecture/tuning\n - Build confusion matrix to diagnose issues\n\nIf you have any other ideas you're willing to share, feel free to comment.",
      "votes": null
    },
    {
      "id": "274085",
      "postDate": "01/25/2018 21:34:20",
      "content": "<p>You're just killing it, man!</p>",
      "rawMarkdown": "You're just killing it, man!",
      "votes": null
    },
    {
      "id": "274169",
      "postDate": "01/26/2018 01:15:41",
      "content": "<p>Very nice Andres, thanks for sharing your approach! </p>\n\n<p>Couple (minor) optimization suggestions - </p>\n\n<ol>\n<li><p>Multiprocessing - Have a look at the consumer-producer(s) design pattern. Invoking pool-map on every batch yield isn't particularly efficient. I suspect that your GPUs are the bottleneck which is masking this issue. As an aside, GIL doesn't mean you should always default to processes over threads. Most functions in C-native Python libraries such as numpy release the GIL. The book \"Fluent Python\" is a great read to learn more about this and other lower level language structures.</p></li>\n<li><p>Increasing your GPU throughput - You've mentioned gradient check-pointing but also have a look at replacing the TF backend with MXNet. Note that this would require using a forked version of Keras.</p></li>\n<li><p>Downclocking your GPUs - this might be controversial but consider downclocking your GPUs if you have a hobbyist home setup (use nvidia-smi to set persistence mode with the PM flag, and then using the PL flag to set the target TDP). If you're running your GPUs 24/7 like most of us are, it reduces thermal wear in the long-term for little-to-no performance decrease. I personally run my 1080 Tis at 200W (down from the default 270W) and have measured a &lt; 3% decrease. </p></li>\n</ol>",
      "rawMarkdown": "Very nice Andres, thanks for sharing your approach! \n\nCouple (minor) optimization suggestions - \n\n1. Multiprocessing - Have a look at the consumer-producer(s) design pattern. Invoking pool-map on every batch yield isn't particularly efficient. I suspect that your GPUs are the bottleneck which is masking this issue. As an aside, GIL doesn't mean you should always default to processes over threads. Most functions in C-native Python libraries such as numpy release the GIL. The book \"Fluent Python\" is a great read to learn more about this and other lower level language structures.\n\n2. Increasing your GPU throughput - You've mentioned gradient check-pointing but also have a look at replacing the TF backend with MXNet. Note that this would require using a forked version of Keras.\n\n3. Downclocking your GPUs - this might be controversial but consider downclocking your GPUs if you have a hobbyist home setup (use nvidia-smi to set persistence mode with the PM flag, and then using the PL flag to set the target TDP). If you're running your GPUs 24/7 like most of us are, it reduces thermal wear in the long-term for little-to-no performance decrease. I personally run my 1080 Tis at 200W (down from the default 270W) and have measured a &lt; 3% decrease.",
      "votes": null
    },
    {
      "id": "274203",
      "postDate": "01/26/2018 03:23:44",
      "content": "<p>That's awesome</p>",
      "rawMarkdown": "That's awesome",
      "votes": null
    },
    {
      "id": "274287",
      "postDate": "01/26/2018 08:49:03",
      "content": "<p>Thanks for the suggestions. Just bought the \"Fluent Python\" book. :-) Will also try your other suggestions, Im going to check CNTK too. </p>\n\n<p>I believe using a MXNet would preclude me from using imagenet pretrained weights for latest networks (NasNet, DenseNet, etc.). Do you know if it is the case?</p>",
      "rawMarkdown": "Thanks for the suggestions. Just bought the \"Fluent Python\" book. :-) Will also try your other suggestions, Im going to check CNTK too. \n\nI believe using a MXNet would preclude me from using imagenet pretrained weights for latest networks (NasNet, DenseNet, etc.). Do you know if it is the case?",
      "votes": null
    },
    {
      "id": "274306",
      "postDate": "01/26/2018 09:42:36",
      "content": "<p>That's my understanding but I'm happy to be corrected. My team-mate and I have stuck to vanilla Keras because of that. So there's a trade-off between flexibility and performance. </p>\n\n<p>In any case, we're not using any of the cutting edge models (e.g., DenseNet, NASNet) for this competition. </p>",
      "rawMarkdown": "That's my understanding but I'm happy to be corrected. My team-mate and I have stuck to vanilla Keras because of that. So there's a trade-off between flexibility and performance. \n\nIn any case, we're not using any of the cutting edge models (e.g., DenseNet, NASNet) for this competition.",
      "votes": null
    },
    {
      "id": "274308",
      "postDate": "01/26/2018 09:50:41",
      "content": "<p>Im curious: are you using an ensemble of models or a single model for your score?</p>",
      "rawMarkdown": "Im curious: are you using an ensemble of models or a single model for your score?",
      "votes": null
    },
    {
      "id": "274321",
      "postDate": "01/26/2018 10:34:13",
      "content": "<p>We're ensembling but by default rather than having benchmarked it against single models.  I doubt this is a major factor. </p>\n\n<p>Having observed the trajectory of the top teams' scores for the past month, I suspect we are sniffing around the same local optima and the spread in performance is largely due to hyper-parameters/randomness.</p>\n\n<p>Hopefully someone will come in with a novel solution and blow the LB up!</p>",
      "rawMarkdown": "We're ensembling but by default rather than having benchmarked it against single models.  I doubt this is a major factor. \n\nHaving observed the trajectory of the top teams' scores for the past month, I suspect we are sniffing around the same local optima and the spread in performance is largely due to hyper-parameters/randomness.\n\nHopefully someone will come in with a novel solution and blow the LB up!",
      "votes": null
    },
    {
      "id": "274329",
      "postDate": "01/26/2018 11:06:54",
      "content": "<p>Thanks for the power limit tip.</p>\n\n<blockquote>\n  <p>set persistence mode with the PM flag</p>\n</blockquote>\n\n<p><a href=\"http://docs.nvidia.com/deploy/driver-persistence/index.html#persistence-daemon\">NVIDIA said</a> they'll eventually stop supporting persistence mode, and are focusing on the persistence daemon. <a href=\"http://docs.nvidia.com/deploy/driver-persistence/index.html#installation\">Getting the daemon to start at startup</a> is a little weird, though. On Ubuntu, I had to unpack /usr/share/doc/NVIDIA_GLX-1.0/sample/nvidia-persistenced-init.tar.bz2 and then run the install script. <a href=\"https://devtalk.nvidia.com/default/topic/995248/cuda-setup-and-installation/setting-up-nvidia-persistenced/post/5090647/#5090647\">One person says</a> to add \"--persistence-mode\" to the service start line in the template file, but that seems unneeded.</p>",
      "rawMarkdown": "Thanks for the power limit tip.\n\n&gt; set persistence mode with the PM flag\n\n[NVIDIA said][1] they'll eventually stop supporting persistence mode, and are focusing on the persistence daemon. [Getting the daemon to start at startup][2] is a little weird, though. On Ubuntu, I had to unpack /usr/share/doc/NVIDIA_GLX-1.0/sample/nvidia-persistenced-init.tar.bz2 and then run the install script. [One person says][3] to add \"--persistence-mode\" to the service start line in the template file, but that seems unneeded.\n\n  [1]: http://docs.nvidia.com/deploy/driver-persistence/index.html#persistence-daemon\n  [2]: http://docs.nvidia.com/deploy/driver-persistence/index.html#installation\n  [3]: https://devtalk.nvidia.com/default/topic/995248/cuda-setup-and-installation/setting-up-nvidia-persistenced/post/5090647/#5090647",
      "votes": null
    },
    {
      "id": "274386",
      "postDate": "01/26/2018 14:19:27",
      "content": "<p>Hey everyone for me the code is not working. It's stuck like this for almost a day now:</p>\n\n<pre><code>       HTC-1-M7:  1023 (13.0%)\n       iPhone-6:   823 (10.5%)\n       Motorola-Droid-Maxx:   825 (10.5%)\n       Motorola-X:   275 (03.5%)\n       Samsung-Galaxy-S4:  1412 (18.0%)\n       iPhone-4s:   774 (09.9%)\n       LG-Nexus-5x:   680 (08.7%)\n       Motorola-Nexus-6:   926 (11.8%)\n       Samsung-Galaxy-Note3:   548 (07.0%)\n       Sony-NEX-7:   557 (07.1%)\n       validation steps = 0\n       Epoch 1/200\n</code></pre>\n\n<p>Any suggestions?</p>",
      "rawMarkdown": "Hey everyone for me the code is not working. It's stuck like this for almost a day now:\n\n           HTC-1-M7:  1023 (13.0%)\n           iPhone-6:   823 (10.5%)\n           Motorola-Droid-Maxx:   825 (10.5%)\n           Motorola-X:   275 (03.5%)\n           Samsung-Galaxy-S4:  1412 (18.0%)\n           iPhone-4s:   774 (09.9%)\n           LG-Nexus-5x:   680 (08.7%)\n           Motorola-Nexus-6:   926 (11.8%)\n           Samsung-Galaxy-Note3:   548 (07.0%)\n           Sony-NEX-7:   557 (07.1%)\n           validation steps = 0\n           Epoch 1/200\nAny suggestions?",
      "votes": null
    },
    {
      "id": "274388",
      "postDate": "01/26/2018 14:25:17",
      "content": "<p>I assume you're running it with <code>-x</code>, i.e. using Gleb's dataset. \nDid you sync the <code>val_images</code> directory @ <a href=\"https://github.com/antorsae/sp-society-camera-model-identification/tree/master/val_images\">https://github.com/antorsae/sp-society-camera-model-identification/tree/master/val_images</a> and downloaded all images?</p>\n\n<p>When using <code>-x</code> I wanted to have the validation as distinct as possible to the train set, so:</p>\n\n<pre><code>    ids_train = ids\n    ids_val   = [ ]\n\n    extra_train_ids = [os.path.join(EXTRA_TRAIN_FOLDER,line.rstrip('\\n')) for line in open(os.path.join(EXTRA_TRAIN_FOLDER, 'good_jpgs'))]\n    extra_train_ids.sort()\n    ids_train.extend(extra_train_ids)\n\n    extra_val_ids = glob.glob(join(EXTRA_VAL_FOLDER,'*/*.jpg'))\n    extra_val_ids.sort()\n    ids_val.extend(extra_val_ids)\n</code></pre>\n\n<p>do a <code>print(ids_val)</code> after that line and LMK what you get.</p>",
      "rawMarkdown": "I assume you're running it with `-x`, i.e. using Gleb's dataset. \nDid you sync the `val_images` directory @ https://github.com/antorsae/sp-society-camera-model-identification/tree/master/val_images and downloaded all images?\n\nWhen using `-x` I wanted to have the validation as distinct as possible to the train set, so:\n\n        ids_train = ids\n        ids_val   = [ ]\n\n        extra_train_ids = [os.path.join(EXTRA_TRAIN_FOLDER,line.rstrip('\\n')) for line in open(os.path.join(EXTRA_TRAIN_FOLDER, 'good_jpgs'))]\n        extra_train_ids.sort()\n        ids_train.extend(extra_train_ids)\n\n        extra_val_ids = glob.glob(join(EXTRA_VAL_FOLDER,'*/*.jpg'))\n        extra_val_ids.sort()\n        ids_val.extend(extra_val_ids)\n\ndo a `print(ids_val)` after that line and LMK what you get.",
      "votes": null
    },
    {
      "id": "274392",
      "postDate": "01/26/2018 14:33:44",
      "content": "<p>Hi Andres, thanks for the quick response. I haven't downloaded all the images, I that was happening automatically. Should I'll be also creating any additional directories for the images? On another note I am getting the value of <code>validation_steps=0</code> in the fit.generator so as a consequence I had to hardcode the value to 6 in order for the code not  to throw an error. Let me try download all the images and run again give you back the output of <code>print(ids_val)</code>.</p>",
      "rawMarkdown": "Hi Andres, thanks for the quick response. I haven't downloaded all the images, I that was happening automatically. Should I'll be also creating any additional directories for the images? On another note I am getting the value of `validation_steps=0` in the fit.generator so as a consequence I had to hardcode the value to 6 in order for the code not  to throw an error. Let me try download all the images and run again give you back the output of `print(ids_val)`.",
      "votes": null
    },
    {
      "id": "274434",
      "postDate": "01/26/2018 15:41:34",
      "content": "<p>Download Gleb's training set by running <code>find . -name \"urls_*\" -execdir wget -nc --tries=10 -i {} \\;</code> in the <code>flickr_images</code> directory and adjust accordingly in the <code>val_images</code> one.</p>",
      "rawMarkdown": "Download Gleb's training set by running `find . -name \"urls_*\" -execdir wget -nc --tries=10 -i {} \\;` in the `flickr_images` directory and adjust accordingly in the `val_images` one.",
      "votes": null
    },
    {
      "id": "274437",
      "postDate": "01/26/2018 15:48:31",
      "content": "<p>I got 0.946 on the public LB using Andres code with \n<code>python train.py -g 1 -b 16 -cs 512 -cm ResNet50 -x -l 1e-4 -uiw</code>\ntraining for 91 epochs on a single V100. Training took about 10 hours. Thanks Andres, I have learned a lot from looking at your code.</p>",
      "rawMarkdown": "I got 0.946 on the public LB using Andres code with \n`python train.py -g 1 -b 16 -cs 512 -cm ResNet50 -x -l 1e-4 -uiw`\ntraining for 91 epochs on a single V100. Training took about 10 hours. Thanks Andres, I have learned a lot from looking at your code.",
      "votes": null
    },
    {
      "id": "274449",
      "postDate": "01/26/2018 16:12:43",
      "content": "<p>Also don't forget to rename any .JPG to .jpg as mentioned on the other thread.</p>",
      "rawMarkdown": "Also don't forget to rename any .JPG to .jpg as mentioned on the other thread.",
      "votes": null
    },
    {
      "id": "274491",
      "postDate": "01/26/2018 17:52:03",
      "content": "<p>Thank you Andres &amp; Alberto, we are back in business fellas :).</p>",
      "rawMarkdown": "Thank you Andres &amp; Alberto, we are back in business fellas :).",
      "votes": null
    },
    {
      "id": "274548",
      "postDate": "01/26/2018 19:26:37",
      "content": "<p>Good to hear. There's a bug in the code re: augmentation. LMK if you can spot it :-) I will submit fixed version with other improvements this weekend.</p>",
      "rawMarkdown": "Good to hear. There's a bug in the code re: augmentation. LMK if you can spot it :-) I will submit fixed version with other improvements this weekend.",
      "votes": null
    },
    {
      "id": "274558",
      "postDate": "01/26/2018 19:41:27",
      "content": "<p>is bug in the train part?</p>",
      "rawMarkdown": "is bug in the train part?",
      "votes": null
    },
    {
      "id": "274561",
      "postDate": "01/26/2018 19:45:38",
      "content": "<p>Yes, may or may not decrease accuracy (I guess it does), but I want to retrain and see tomorrow. Will submit then.</p>",
      "rawMarkdown": "Yes, may or may not decrease accuracy (I guess it does), but I want to retrain and see tomorrow. Will submit then.",
      "votes": null
    },
    {
      "id": "274578",
      "postDate": "01/26/2018 20:15:51",
      "content": "<p>I noticed something in the gamma augmentation, the formula would be pow(x, 1/gamma) but it's written as pow(x, gamma); I think it doesn't matter much as 1/0.8==1.25 and 1/1.2==0.83, so in the end gamma08 generates gamma 1.2 and vice-versa :-)</p>",
      "rawMarkdown": "I noticed something in the gamma augmentation, the formula would be pow(x, 1/gamma) but it's written as pow(x, gamma); I think it doesn't matter much as 1/0.8==1.25 and 1/1.2==0.83, so in the end gamma08 generates gamma 1.2 and vice-versa :-)",
      "votes": null
    },
    {
      "id": "274581",
      "postDate": "01/26/2018 20:21:29",
      "content": "<p>The bug I was referring to is not related to gamma. Re: gamma by looking at <a href=\"http://scikit-image.org/docs/dev/api/skimage.exposure.html#skimage.exposure.adjust_gamma\">http://scikit-image.org/docs/dev/api/skimage.exposure.html#skimage.exposure.adjust_gamma</a> the formula is O = I**gamma so I think the current formula is OK; why should it be 1/gamma?</p>",
      "rawMarkdown": "The bug I was referring to is not related to gamma. Re: gamma by looking at http://scikit-image.org/docs/dev/api/skimage.exposure.html#skimage.exposure.adjust_gamma the formula is O = I**gamma so I think the current formula is OK; why should it be 1/gamma?",
      "votes": null
    },
    {
      "id": "274585",
      "postDate": "01/26/2018 20:34:31",
      "content": "<p>There is some confusion about which formula is actually \"gamma correction\":\n<a href=\"https://stackoverflow.com/questions/16521003/gamma-correction-formula-gamma-or-1-gamma\">https://stackoverflow.com/questions/16521003/gamma-correction-formula-gamma-or-1-gamma</a></p>",
      "rawMarkdown": "There is some confusion about which formula is actually \"gamma correction\":\nhttps://stackoverflow.com/questions/16521003/gamma-correction-formula-gamma-or-1-gamma",
      "votes": null
    },
    {
      "id": "274616",
      "postDate": "01/26/2018 22:05:08",
      "content": "<p>I am getting this error :</p>\n\n<p>$ python train.py -g 1 -b 16 -cs 512 -cm ResNet50 -l 1e-4 -uiw</p>\n\n<p>2018-01-26 23:01:46.708921: E tensorflow/stream_executor/cuda/cuda_dnn.cc:385] could not create cudnn handle: CUDNN_STATUS_INTERNAL_ERROR\n2018-01-26 23:01:46.708954: E tensorflow/stream_executor/cuda/cuda_dnn.cc:352] could not destroy cudnn handle: CUDNN_STATUS_BAD_PARAM\n2018-01-26 23:01:46.708961: F tensorflow/core/kernels/conv_ops.cc:667] Check failed: stream-&gt;parent()-&gt;GetConvolveAlgorithms( conv_parameters.ShouldIncludeWinogradNonfusedAlgo</p>\n\n<p>Everything works fine If I run Resnet50 from keras examples.</p>\n\n<p>Anyone has an idea what might be the reason ? </p>",
      "rawMarkdown": "I am getting this error :\n\n$ python train.py -g 1 -b 16 -cs 512 -cm ResNet50 -l 1e-4 -uiw\n\n2018-01-26 23:01:46.708921: E tensorflow/stream_executor/cuda/cuda_dnn.cc:385] could not create cudnn handle: CUDNN_STATUS_INTERNAL_ERROR\n2018-01-26 23:01:46.708954: E tensorflow/stream_executor/cuda/cuda_dnn.cc:352] could not destroy cudnn handle: CUDNN_STATUS_BAD_PARAM\n2018-01-26 23:01:46.708961: F tensorflow/core/kernels/conv_ops.cc:667] Check failed: stream-&gt;parent()-&gt;GetConvolveAlgorithms( conv_parameters.ShouldIncludeWinogradNonfusedAlgo\n\nEverything works fine If I run Resnet50 from keras examples.\n\nAnyone has an idea what might be the reason ?",
      "votes": null
    },
    {
      "id": "274626",
      "postDate": "01/26/2018 23:00:21",
      "content": "<p>Maybe it is a versioning issue. What about downgrading keras?</p>",
      "rawMarkdown": "Maybe it is a versioning issue. What about downgrading keras?",
      "votes": null
    },
    {
      "id": "274628",
      "postDate": "01/26/2018 23:04:24",
      "content": "<p>I use the same keras version and tf as OP \nWhat version are you using ?</p>",
      "rawMarkdown": "I use the same keras version and tf as OP \nWhat version are you using ?",
      "votes": null
    },
    {
      "id": "274631",
      "postDate": "01/26/2018 23:14:59",
      "content": "<p>keras 2.1.2, tensorflow 1.4.0 and it works well. I also tried to upgrade keras to 2.1.3 and got the error you mentioned above. </p>",
      "rawMarkdown": "keras 2.1.2, tensorflow 1.4.0 and it works well. I also tried to upgrade keras to 2.1.3 and got the error you mentioned above.",
      "votes": null
    },
    {
      "id": "274635",
      "postDate": "01/26/2018 23:30:38",
      "content": "<p>I use Keras 2.1.3 and TF 1.4.1</p>",
      "rawMarkdown": "I use Keras 2.1.3 and TF 1.4.1",
      "votes": null
    },
    {
      "id": "274638",
      "postDate": "01/26/2018 23:40:23",
      "content": "<p>Check your CUDA &amp; cudnn versions. Default Tensorflow from pip compiled with cuda 9.0 and cudnn 7, so you need exactly them.</p>",
      "rawMarkdown": "Check your CUDA &amp; cudnn versions. Default Tensorflow from pip compiled with cuda 9.0 and cudnn 7, so you need exactly them.",
      "votes": null
    },
    {
      "id": "274649",
      "postDate": "01/27/2018 01:09:42",
      "content": "<p>Another possible reason: Anaconda. It messes with cublas libraries and tensorflow going crazy. Solution: probably, simplest way -- just use system python for this code.</p>",
      "rawMarkdown": "Another possible reason: Anaconda. It messes with cublas libraries and tensorflow going crazy. Solution: probably, simplest way -- just use system python for this code.",
      "votes": null
    },
    {
      "id": "274709",
      "postDate": "01/27/2018 03:53:03",
      "content": "<p>First of all, really appreciate sharing the code. It really, really helps.</p>\n\n<p>I managed to run it on one of my machines, but now when I am trying to run the latest version (downloaded today) on another machine, I am getting really weird number of images in different classes. One of them (Moto X) even has 0 images. However, when I check the folders, I see that all of the images are there. Anyone else have the same issue? Any idea how to fix it?</p>\n\n<p><code>HTC-1-M7:   748 (13.9%) \n              iPhone-6:   548 (10.2%)\n   Motorola-Droid-Maxx:   550 (10.2%)\n            Motorola-X:     0 (00.0%)\n     Samsung-Galaxy-S4:  1137 (21.2%)\n             iPhone-4s:   499 (09.3%)\n           LG-Nexus-5x:   405 (07.5%)\n      Motorola-Nexus-6:   651 (12.1%)\n  Samsung-Galaxy-Note3:   273 (05.1%)\n            Sony-NEX-7:   557 (10.4%)\n</code></p>",
      "rawMarkdown": "First of all, really appreciate sharing the code. It really, really helps.\n\nI managed to run it on one of my machines, but now when I am trying to run the latest version (downloaded today) on another machine, I am getting really weird number of images in different classes. One of them (Moto X) even has 0 images. However, when I check the folders, I see that all of the images are there. Anyone else have the same issue? Any idea how to fix it?\n\n```              HTC-1-M7:   748 (13.9%) \n              iPhone-6:   548 (10.2%)\n   Motorola-Droid-Maxx:   550 (10.2%)\n            Motorola-X:     0 (00.0%)\n     Samsung-Galaxy-S4:  1137 (21.2%)\n             iPhone-4s:   499 (09.3%)\n           LG-Nexus-5x:   405 (07.5%)\n      Motorola-Nexus-6:   651 (12.1%)\n  Samsung-Galaxy-Note3:   273 (05.1%)\n            Sony-NEX-7:   557 (10.4%)\n```",
      "votes": null
    },
    {
      "id": "274761",
      "postDate": "01/27/2018 08:15:15",
      "content": "<p>Hi, <br>\nI am using Keras, with CUDA_VISIBLE_DEVICES to mask to use 2 GTX1080 GPU,\nbut when I run DenseNet201 model, will another two fc layers,\nIt give me a lot of errors : </p>\n\n<pre><code>ran out of memory trying to allocate XXXX MiB\n</code></pre>\n\n<p>I tried to use 4 GTX1080 GPU,  but still get same errors.\nIt seems that it GPU memory doesn't increase, it performance is same as single GTX1080\nCan you give me some suggestions? What should I do to implement this model to multi-gpu?</p>",
      "rawMarkdown": "Hi,  \nI am using Keras, with CUDA_VISIBLE_DEVICES to mask to use 2 GTX1080 GPU,\nbut when I run DenseNet201 model, will another two fc layers,\nIt give me a lot of errors : \n\n    ran out of memory trying to allocate XXXX MiB\n\nI tried to use 4 GTX1080 GPU,  but still get same errors.\nIt seems that it GPU memory doesn't increase, it performance is same as single GTX1080\nCan you give me some suggestions? What should I do to implement this model to multi-gpu?",
      "votes": null
    },
    {
      "id": "274773",
      "postDate": "01/27/2018 09:02:47",
      "content": "<p>OK, getting 96.2 LB now single model.\nCode is updated in the repo:</p>\n\n<ul>\n<li>Added GPL3 license: basically if you modify it, share the code.</li>\n<li>Fixed orientation flip augmentation bug (need to confirm current one is OK.</li>\n<li>Control TTA and prints class distribution after <code>-t</code>.</li>\n</ul>",
      "rawMarkdown": "OK, getting 96.2 LB now single model.\nCode is updated in the repo:\n\n - Added GPL3 license: basically if you modify it, share the code.\n - Fixed orientation flip augmentation bug (need to confirm current one is OK.\n - Control TTA and prints class distribution after `-t`.",
      "votes": null
    },
    {
      "id": "274782",
      "postDate": "01/27/2018 09:48:17",
      "content": "<p>I created a fresh env of py3.5and it worked, before I was using py3.6.</p>\n\n<p>Thanks guys !</p>",
      "rawMarkdown": "I created a fresh env of py3.5and it worked, before I was using py3.6.\n\nThanks guys !",
      "votes": null
    },
    {
      "id": "274821",
      "postDate": "01/27/2018 12:45:17",
      "content": "<p>do you do <code>-g 2</code> for 2 GPUs?</p>",
      "rawMarkdown": "do you do `-g 2` for 2 GPUs?",
      "votes": null
    },
    {
      "id": "274860",
      "postDate": "01/27/2018 14:58:24",
      "content": "<p>Can you share how you created the environment. I still haven't been able to run Andres code anymore. I'm actually creating a new code base and deconstructing it part by part.</p>",
      "rawMarkdown": "Can you share how you created the environment. I still haven't been able to run Andres code anymore. I'm actually creating a new code base and deconstructing it part by part.",
      "votes": null
    },
    {
      "id": "274864",
      "postDate": "01/27/2018 15:09:43",
      "content": "<p>i use anaconda : \n<a href=\"http://uoa-eresearch.github.io/eresearch-cookbook/recipe/2014/11/20/conda/\">http://uoa-eresearch.github.io/eresearch-cookbook/recipe/2014/11/20/conda/</a></p>",
      "rawMarkdown": "i use anaconda : \nhttp://uoa-eresearch.github.io/eresearch-cookbook/recipe/2014/11/20/conda/",
      "votes": null
    },
    {
      "id": "274868",
      "postDate": "01/27/2018 15:18:32",
      "content": "<p>On linux I use <a href=\"https://github.com/pypa/pipenv\">pipenv</a></p>",
      "rawMarkdown": "On linux I use [pipenv][1]\n\n\n  [1]: https://github.com/pypa/pipenv",
      "votes": null
    },
    {
      "id": "274871",
      "postDate": "01/27/2018 15:23:48",
      "content": "<p>Yes, but that doesn't give me versions/packages.</p>\n\n<p>I have several virtenvs but something happened trying to upgrade tensorflow and I haven't been able to run Andres code since.</p>",
      "rawMarkdown": "Yes, but that doesn't give me versions/packages.\n\nI have several virtenvs but something happened trying to upgrade tensorflow and I haven't been able to run Andres code since.",
      "votes": null
    },
    {
      "id": "274872",
      "postDate": "01/27/2018 15:23:56",
      "content": "<p>Alberto, this may be an overkill but here's my base environment to use with anaconda (I use miniconda, but regular anaconda would work too). There's many packages that are not needed for this project so feel free to edit the file by hand and remove the ones you think are not needed.</p>\n\n<p><a href=\"https://conda.io/docs/user-guide/tasks/manage-environments.html\">https://conda.io/docs/user-guide/tasks/manage-environments.html</a></p>",
      "rawMarkdown": "Alberto, this may be an overkill but here's my base environment to use with anaconda (I use miniconda, but regular anaconda would work too). There's many packages that are not needed for this project so feel free to edit the file by hand and remove the ones you think are not needed.\n\nhttps://conda.io/docs/user-guide/tasks/manage-environments.html",
      "votes": null
    },
    {
      "id": "274879",
      "postDate": "01/27/2018 15:39:00",
      "content": "<p>I'm using Docker for Andres's code:</p>\n\n<pre><code>alias andres=\"docker run --runtime=nvidia --init -it --rm \\\n              --ipc=host \\\n              -v $CODE:/code -v $DATA:/data \\\n              -w=/code/sp-society-camera-model-identification \\\n              mwksmith/cam:andres\"\n</code></pre>\n\n<p>When you get into the container enter \"sa\" to activate the conda environment, and then create sym links to your data folders according to the globals set in train.py. Then run train.py.</p>\n\n<p>I'm pushing the Docker Image to Docker Hub now. You can also build it yourself with <a href=\"https://github.com/antorsae/sp-society-camera-model-identification/blob/master/Dockerfile\">the Dockerfile</a>.</p>\n\n<p>Edit: Removed <code>--shm-size</code>. It appears to be redundant with respect to <code>--ipc=host</code>.</p>",
      "rawMarkdown": "I'm using Docker for Andres's code:\n\n    alias andres=\"docker run --runtime=nvidia --init -it --rm \\\n                  --ipc=host \\\n                  -v $CODE:/code -v $DATA:/data \\\n                  -w=/code/sp-society-camera-model-identification \\\n                  mwksmith/cam:andres\"\n\nWhen you get into the container enter \"sa\" to activate the conda environment, and then create sym links to your data folders according to the globals set in train.py. Then run train.py.\n\nI'm pushing the Docker Image to Docker Hub now. You can also build it yourself with [the Dockerfile][1].\n\n\n  [1]: https://github.com/antorsae/sp-society-camera-model-identification/blob/master/Dockerfile\n\nEdit: Removed `--shm-size`. It appears to be redundant with respect to `--ipc=host`.",
      "votes": null
    },
    {
      "id": "274913",
      "postDate": "01/27/2018 18:06:40",
      "content": "<blockquote>\n  <p><strong>Andres Torrubia wrote</strong></p>\n  \n  <blockquote>\n    <p>Download Gleb's training set by running <code>find . -name \"urls_*\" -execdir wget -nc --tries=10 -i {} \\;</code> in the <code>flickr_images</code> directory and adjust accordingly in the <code>val_images</code> one.</p>\n  </blockquote>\n</blockquote>\n\n<p>Here's the breakdown of this command.</p>\n\n<ul>\n<li><code>find</code>: returns file paths that match a given pattern</li>\n<li><code>.</code>: search working directory recursively (search its subdirectories, their subdirectories, and so on)</li>\n<li><code>-name \"urls_*\"</code>: the given file path pattern</li>\n<li><code>-execdir</code>: execute the given command on each file path found, <strong>in the directory of the file.</strong></li>\n<li><code>wget</code>: download what a given URL points to</li>\n<li><code>-nc</code>: don't download files that already exist</li>\n<li><code>-i</code>: read URLs from a given file</li>\n<li><code>{}</code>: placeholder for each file that <code>find</code> found</li>\n<li><code>\\;</code>: end command</li>\n</ul>",
      "rawMarkdown": "&gt; **Andres Torrubia wrote**\n&gt; \n&gt; &gt; Download Gleb's training set by running `find . -name \"urls_*\" -execdir wget -nc --tries=10 -i {} \\;` in the `flickr_images` directory and adjust accordingly in the `val_images` one.\n\nHere's the breakdown of this command.\n\n- `find`: returns file paths that match a given pattern\n- `.`: search working directory recursively (search its subdirectories, their subdirectories, and so on)\n-  `-name \"urls_*\"`: the given file path pattern\n- `-execdir`: execute the given command on each file path found, **in the directory of the file.**\n- `wget`: download what a given URL points to\n- `-nc`: don't download files that already exist\n- `-i`: read URLs from a given file\n- `{}`: placeholder for each file that `find` found\n- `\\;`: end command",
      "votes": null
    },
    {
      "id": "274916",
      "postDate": "01/27/2018 18:15:59",
      "content": "<p>One caveat when doing this on <code>val_images</code>. In the <code>val_images/moto_maxx</code> directory a bunch of images will get downloaed with *.1 *.2 *.3 ... extensions. Rename them to _1.jpg _2.jpg _3.jpg otherwise the code will not use them and if using <code>-x</code> validation set for that class will be heavily underrepresented. </p>",
      "rawMarkdown": "One caveat when doing this on `val_images`. In the `val_images/moto_maxx` directory a bunch of images will get downloaed with *.1 *.2 *.3 ... extensions. Rename them to _1.jpg _2.jpg _3.jpg otherwise the code will not use them and if using `-x` validation set for that class will be heavily underrepresented.",
      "votes": null
    },
    {
      "id": "274917",
      "postDate": "01/27/2018 18:18:47",
      "content": "<p>Thanks, it may end up being what I need.</p>\n\n<p>I finally finished going through every line and did some re-structuring of the code from a couple of days ago. I made it a little more robust in handling errors. Unfortunately I also got rid of the paralleled implementation of the generator so it is now slower. So far it started to run, so I may finally be in the right track.</p>\n\n<p>For anyone interested my fork is at: <a href=\"https://github.com/albertoa/sp-society-camera-model-identification\">https://github.com/albertoa/sp-society-camera-model-identification</a>\nI would only recommend it at this stage for the added comments. I will improve the generator and look at your latest changes to see what else would be useful. For now I'm just happy I'm back training networks.</p>",
      "rawMarkdown": "Thanks, it may end up being what I need.\n\nI finally finished going through every line and did some re-structuring of the code from a couple of days ago. I made it a little more robust in handling errors. Unfortunately I also got rid of the paralleled implementation of the generator so it is now slower. So far it started to run, so I may finally be in the right track.\n\nFor anyone interested my fork is at: https://github.com/albertoa/sp-society-camera-model-identification\nI would only recommend it at this stage for the added comments. I will improve the generator and look at your latest changes to see what else would be useful. For now I'm just happy I'm back training networks.",
      "votes": null
    },
    {
      "id": "274918",
      "postDate": "01/27/2018 18:22:30",
      "content": "<p>Thanks, this may be the fastest approach for me to get his latest release working. I'll try it as soon as the current network training is done.</p>",
      "rawMarkdown": "Thanks, this may be the fastest approach for me to get his latest release working. I'll try it as soon as the current network training is done.",
      "votes": null
    },
    {
      "id": "274940",
      "postDate": "01/27/2018 21:02:18",
      "content": "<p>Got 0.969 LB w/ single model. \nI think there's a chance to get to 0.98 with a single model.</p>",
      "rawMarkdown": "Got 0.969 LB w/ single model. \nI think there's a chance to get to 0.98 with a single model.",
      "votes": null
    },
    {
      "id": "274945",
      "postDate": "01/27/2018 21:19:15",
      "content": "<p>With the same code? </p>",
      "rawMarkdown": "With the same code?",
      "votes": null
    },
    {
      "id": "274952",
      "postDate": "01/27/2018 21:46:44",
      "content": "<p>Almost. Diff is 20 lines of code. Will submit after more changes I have planned.</p>",
      "rawMarkdown": "Almost. Diff is 20 lines of code. Will submit after more changes I have planned.",
      "votes": null
    },
    {
      "id": "274965",
      "postDate": "01/27/2018 22:57:10",
      "content": "<p>Chun, thanks for the note on power. I'm curious though, why set the power limit instead of the temperature thresholds: slow_threshold and max_threshold?</p>",
      "rawMarkdown": "Chun, thanks for the note on power. I'm curious though, why set the power limit instead of the temperature thresholds: slow_threshold and max_threshold?",
      "votes": null
    },
    {
      "id": "274970",
      "postDate": "01/27/2018 23:15:07",
      "content": "<p>you better keep some of that code for yourself if you hope to be in the money :)</p>",
      "rawMarkdown": "you better keep some of that code for yourself if you hope to be in the money :)",
      "votes": null
    },
    {
      "id": "274971",
      "postDate": "01/27/2018 23:18:04",
      "content": "<p>I know. This all adds to the excitement ;-)</p>",
      "rawMarkdown": "I know. This all adds to the excitement ;-)",
      "votes": null
    },
    {
      "id": "274982",
      "postDate": "01/28/2018 00:25:06",
      "content": "<p>I bet top teams are already very busy training <em>a lot</em> of diverse models (many folds probably) and will do it up to the last minute, so we could see a lot of movements near the end :)</p>",
      "rawMarkdown": "I bet top teams are already very busy training *a lot* of diverse models (many folds probably) and will do it up to the last minute, so we could see a lot of movements near the end :)",
      "votes": null
    },
    {
      "id": "274998",
      "postDate": "01/28/2018 01:27:57",
      "content": "<blockquote>\n  <p><strong>Sergey Mushinskiy wrote</strong></p>\n  \n  <blockquote>\n    <p>I bet top teams are already very busy training <em>a lot</em> of diverse models (many folds probably)</p>\n  </blockquote>\n</blockquote>\n\n<p>Hi Sergey, what do you mean by \"fold\"?</p>",
      "rawMarkdown": "&gt; **Sergey Mushinskiy wrote**\n&gt; \n&gt; &gt; I bet top teams are already very busy training *a lot* of diverse models (many folds probably)\n\nHi Sergey, what do you mean by \"fold\"?",
      "votes": null
    },
    {
      "id": "275023",
      "postDate": "01/28/2018 04:16:23",
      "content": "<p>Thanks Andres for sharing your solution. I was running your code. The code run for 44 epochs and stopped spitting the following error. Do you have any ideas how to debug this error?</p>\n\n<p>read Thread-4:\nTraceback (most recent call last):\n  File \"/home/wxk/anaconda2/envs/py3/lib/python3.6/threading.py\", line 916, in _bootstrap_inner\n    self.run()\n  File \"/home/wxk/anaconda2/envs/py3/lib/python3.6/threading.py\", line 864, in run\n    self._target(*self._args, **self._kwargs)\n  File \"/home/wxk/anaconda2/envs/py3/lib/python3.6/multiprocessing/pool.py\", line 429, in _handle_results\n    task = get()\n  File \"/home/wxk/anaconda2/envs/py3/lib/python3.6/multiprocessing/connection.py\", line 251, in recv\n    return _ForkingPickler.loads(buf.getbuffer())\nTypeError: <strong>init</strong>() missing 1 required positional argument: 'code'</p>",
      "rawMarkdown": "Thanks Andres for sharing your solution. I was running your code. The code run for 44 epochs and stopped spitting the following error. Do you have any ideas how to debug this error?\n\nread Thread-4:\nTraceback (most recent call last):\n  File \"/home/wxk/anaconda2/envs/py3/lib/python3.6/threading.py\", line 916, in _bootstrap_inner\n    self.run()\n  File \"/home/wxk/anaconda2/envs/py3/lib/python3.6/threading.py\", line 864, in run\n    self._target(*self._args, **self._kwargs)\n  File \"/home/wxk/anaconda2/envs/py3/lib/python3.6/multiprocessing/pool.py\", line 429, in _handle_results\n    task = get()\n  File \"/home/wxk/anaconda2/envs/py3/lib/python3.6/multiprocessing/connection.py\", line 251, in recv\n    return _ForkingPickler.loads(buf.getbuffer())\nTypeError: __init__() missing 1 required positional argument: 'code'",
      "votes": null
    },
    {
      "id": "275024",
      "postDate": "01/28/2018 04:22:56",
      "content": "<p>I have seen that before. To me it was happening on epoch 1. There seems to be a problem with the multiprocess code that triggers it, but other than going into a single CPU for the generator I couldn't solve it.</p>",
      "rawMarkdown": "I have seen that before. To me it was happening on epoch 1. There seems to be a problem with the multiprocess code that triggers it, but other than going into a single CPU for the generator I couldn't solve it.",
      "votes": null
    },
    {
      "id": "275028",
      "postDate": "01/28/2018 04:43:56",
      "content": "<p>Thanks Alberto for your reply. Then how to set the code to one CPU? </p>\n\n<p>BTW, do you know how the follow code works? I was trying to  figure out how the models (Resnet50, Densenet201 and so on) are created.\ngetattr(globals()[classifier_module_name], 'preprocess_input')</p>",
      "rawMarkdown": "Thanks Alberto for your reply. Then how to set the code to one CPU? \n\nBTW, do you know how the follow code works? I was trying to  figure out how the models (Resnet50, Densenet201 and so on) are created.\ngetattr(globals()[classifier_module_name], 'preprocess_input')",
      "votes": null
    },
    {
      "id": "275033",
      "postDate": "01/28/2018 05:02:20",
      "content": "<p>I was afraid you were going to ask. I have a fork of the github code, but I found a major issue with it that I will fix tomorrow, gives wrong results ;-(.\nAdres is working on a new release, hopefully his has the fix.\nThe lines in question should be:\n    p = Pool(cpu_count()-2)\n            batch_results = p.map(process_item_func, item_batch)\nand how the batches are put together. But that's where I messed up my fork, so I rather not recommend anything until I know I can in fact run it properly.</p>",
      "rawMarkdown": "I was afraid you were going to ask. I have a fork of the github code, but I found a major issue with it that I will fix tomorrow, gives wrong results ;-(.\nAdres is working on a new release, hopefully his has the fix.\nThe lines in question should be:\n    p = Pool(cpu_count()-2)\n            batch_results = p.map(process_item_func, item_batch)\nand how the batches are put together. But that's where I messed up my fork, so I rather not recommend anything until I know I can in fact run it properly.",
      "votes": null
    },
    {
      "id": "275043",
      "postDate": "01/28/2018 05:12:57",
      "content": "<p>Thanks Alberto for your kind reply. </p>",
      "rawMarkdown": "Thanks Alberto for your kind reply.",
      "votes": null
    },
    {
      "id": "275048",
      "postDate": "01/28/2018 05:26:36",
      "content": "<p>By the way I am having the same issue with the docker image from Matt Kleinsmith. </p>\n\n<p>[...]\n      File \"/opt/conda/envs/tf/lib/python3.5/multiprocessing/connection.py\", line 251, in recv\n        return ForkingPickler.loads(buf.getbuffer())\n    TypeError: <strong>init</strong>() missing 1 required positional argument: 'code'</p>\n\n<p>I can not tell you why you and I get the error and yet so many others don't. We are the wrong kind of lucky I guess. </p>",
      "rawMarkdown": "By the way I am having the same issue with the docker image from Matt Kleinsmith. \n\n[...]\n      File \"/opt/conda/envs/tf/lib/python3.5/multiprocessing/connection.py\", line 251, in recv\n        return ForkingPickler.loads(buf.getbuffer())\n    TypeError: __init__() missing 1 required positional argument: 'code'\n\nI can not tell you why you and I get the error and yet so many others don't. We are the wrong kind of lucky I guess.",
      "votes": null
    },
    {
      "id": "275049",
      "postDate": "01/28/2018 05:30:46",
      "content": "<p>Thanks Alberto. </p>\n\n<p>Do you know how the networks become global variables in this line: getattr(globals()[classifier_module_name], 'preprocess_input')?</p>",
      "rawMarkdown": "Thanks Alberto. \n\nDo you know how the networks become global variables in this line: getattr(globals()[classifier_module_name], 'preprocess_input')?",
      "votes": null
    },
    {
      "id": "275051",
      "postDate": "01/28/2018 05:33:58",
      "content": "<p>Thanks a lot for sharing your code! Really helps students like me learn a lot! :)</p>",
      "rawMarkdown": "Thanks a lot for sharing your code! Really helps students like me learn a lot! :)",
      "votes": null
    },
    {
      "id": "275054",
      "postDate": "01/28/2018 05:39:39",
      "content": "<p>The key to that is:</p>\n\n<p>from keras.applications import *</p>\n\n<p>He is using the classifier_to_module dictionary so that the correct preprocess_input function is called for whichever model is being used.</p>",
      "rawMarkdown": "The key to that is:\n\nfrom keras.applications import *\n\nHe is using the classifier_to_module dictionary so that the correct preprocess_input function is called for whichever model is being used.",
      "votes": null
    },
    {
      "id": "275058",
      "postDate": "01/28/2018 05:47:51",
      "content": "<p>I see. Thanks Alberto. I was stilling creating the models myself. Keras is becoming more handy.</p>",
      "rawMarkdown": "I see. Thanks Alberto. I was stilling creating the models myself. Keras is becoming more handy.",
      "votes": null
    },
    {
      "id": "275124",
      "postDate": "01/28/2018 11:30:30",
      "content": "<p>An exception happened in the mutiprocess code and Python does not tell you the exception (very likely a file was not read correctly) so it tanks with an obscure error. Making the code run in 1 thread only is painfully slow unless you resort to some sort of caching.</p>",
      "rawMarkdown": "An exception happened in the mutiprocess code and Python does not tell you the exception (very likely a file was not read correctly) so it tanks with an obscure error. Making the code run in 1 thread only is painfully slow unless you resort to some sort of caching.",
      "votes": null
    },
    {
      "id": "275125",
      "postDate": "01/28/2018 11:32:42",
      "content": "<p>I mean folds like in k-fold cross-validation. Split dataset strategically into, say, 5 parts and train 5 model on 4 of them (different each time) and use 1 as validation. It is very common strategy in general data science but obviously somewhat rare in deep learning (given computational resources needed). However, it allows to build honest second layer model on top of networks predictions.</p>",
      "rawMarkdown": "I mean folds like in k-fold cross-validation. Split dataset strategically into, say, 5 parts and train 5 model on 4 of them (different each time) and use 1 as validation. It is very common strategy in general data science but obviously somewhat rare in deep learning (given computational resources needed). However, it allows to build honest second layer model on top of networks predictions.",
      "votes": null
    },
    {
      "id": "275246",
      "postDate": "01/28/2018 15:13:46",
      "content": "<p>Hi everyone I just wanted to update on some things that I've noticed while using @Andres code which he kindly provided to us. @Andres once again thanks for sharing the code with us. In my system I haven't noticed a big difference in regards to gaining speed training from the multiprocessing generator. Incidentally I am using python3 and running the code without the multiprocessing, I see that multiple cores are still utilized with the difference in regards to the multiprocessor generator, that they are not  &gt;90% occupied all the time, which in some cases is not ideal to have all the cores fully loaded, especially in a shared system. No matter what you set the number of processes in the code still all cores are fully occupied.  That being said, there are some alternatives which I haven't tested yet but I thought I should share them here in case they are useful or maybe someone has already tried them. The first alternative is to use sth like <a href=\"http://tensorpack.readthedocs.io/en/latest/\">tensorpack</a>. The other one is to use sth like <code>producer-consumer</code> pattern that @Chun Ming Lee already mentioned. With all the deadlines I didn't have the time to incorporate that into @Andres code but there's a simple example in the <code>consumer-producer.py</code> file which can be easily extended into @Andres code if anyone wants to.  What does though really make a difference regarding training speed is doing multi-gpu training, almost reduces the training time in half plus you might be able to use larger batch_size. That's all for now folks, I'll update later on different models. Cheers!</p>",
      "rawMarkdown": "Hi everyone I just wanted to update on some things that I've noticed while using @Andres code which he kindly provided to us. @Andres once again thanks for sharing the code with us. In my system I haven't noticed a big difference in regards to gaining speed training from the multiprocessing generator. Incidentally I am using python3 and running the code without the multiprocessing, I see that multiple cores are still utilized with the difference in regards to the multiprocessor generator, that they are not  &gt;90% occupied all the time, which in some cases is not ideal to have all the cores fully loaded, especially in a shared system. No matter what you set the number of processes in the code still all cores are fully occupied.  That being said, there are some alternatives which I haven't tested yet but I thought I should share them here in case they are useful or maybe someone has already tried them. The first alternative is to use sth like [tensorpack][1]. The other one is to use sth like `producer-consumer` pattern that @Chun Ming Lee already mentioned. With all the deadlines I didn't have the time to incorporate that into @Andres code but there's a simple example in the `consumer-producer.py` file which can be easily extended into @Andres code if anyone wants to.  What does though really make a difference regarding training speed is doing multi-gpu training, almost reduces the training time in half plus you might be able to use larger batch_size. That's all for now folks, I'll update later on different models. Cheers!\n\n\n  [1]: http://tensorpack.readthedocs.io/en/latest/",
      "votes": null
    },
    {
      "id": "275247",
      "postDate": "01/28/2018 15:21:19",
      "content": "<p>I have a question @Chun Ming Lee regarding point 2. Does is make any difference which backend you use? Ultimately now all of them rely on the same lower layer which is cuda and cudnn? Which IMHO it won't make any big difference but then again I might be wrong. Thanks!</p>",
      "rawMarkdown": "I have a question @Chun Ming Lee regarding point 2. Does is make any difference which backend you use? Ultimately now all of them rely on the same lower layer which is cuda and cudnn? Which IMHO it won't make any big difference but then again I might be wrong. Thanks!",
      "votes": null
    },
    {
      "id": "275268",
      "postDate": "01/28/2018 16:08:48",
      "content": "<p>I got errors like this, any suggestion?</p>\n\n<p>Training set distribution:\n              HTC-1-M7:  1014 (13.3%)\n              iPhone-6:   793 (10.4%)\n   Motorola-Droid-Maxx:   767 (10.0%)\n            Motorola-X:   599 (07.8%)\n     Samsung-Galaxy-S4:  1380 (18.1%)\n             iPhone-4s:   740 (09.7%)\n           LG-Nexus-5x:   656 (08.6%)\n      Motorola-Nexus-6:   801 (10.5%)\n  Samsung-Galaxy-Note3:   365 (04.8%)\n            Sony-NEX-7:   519 (06.8%)\nValidation set distribution:\n              HTC-1-M7:    48 (10.0%)\n              iPhone-6:    48 (10.0%)\n   Motorola-Droid-Maxx:    48 (10.0%)\n            Motorola-X:    48 (10.0%)\n     Samsung-Galaxy-S4:    48 (10.0%)\n             iPhone-4s:    48 (10.0%)\n           LG-Nexus-5x:    48 (10.0%)\n      Motorola-Nexus-6:    48 (10.0%)\n  Samsung-Galaxy-Note3:    48 (10.0%)\n            Sony-NEX-7:    48 (10.0%)\nEpoch 1/200\n165/954 [====&gt;.........................] - ETA: 14:01 - loss: 2.2015 - acc: 0.2167Exception in thread Thread-8:\nTraceback (most recent call last):\n  File \"/usr/lib/python2.7/threading.py\", line 801, in <strong>bootstrap_inner\n    self.run()\n  File \"/usr/lib/python2.7/threading.py\", line 754, in run\n    self.__target(*self.__args, **self.__kwargs)\n  File \"/usr/lib/python2.7/multiprocessing/pool.py\", line 389, in _handle_results\n    task = get()\nTypeError: ('__init</strong>() takes exactly 3 arguments (2 given)', , (u'tjDecompressHeader2() failed with error -1 and error string Not a JPEG file: starts with 0x89 0x50',))</p>\n\n<p>175/954 [====&gt;.........................] - ETA: 13:49 - loss: 2.1799 - acc: 0.2279</p>",
      "rawMarkdown": "I got errors like this, any suggestion?\n\n\nTraining set distribution:\n              HTC-1-M7:  1014 (13.3%)\n              iPhone-6:   793 (10.4%)\n   Motorola-Droid-Maxx:   767 (10.0%)\n            Motorola-X:   599 (07.8%)\n     Samsung-Galaxy-S4:  1380 (18.1%)\n             iPhone-4s:   740 (09.7%)\n           LG-Nexus-5x:   656 (08.6%)\n      Motorola-Nexus-6:   801 (10.5%)\n  Samsung-Galaxy-Note3:   365 (04.8%)\n            Sony-NEX-7:   519 (06.8%)\nValidation set distribution:\n              HTC-1-M7:    48 (10.0%)\n              iPhone-6:    48 (10.0%)\n   Motorola-Droid-Maxx:    48 (10.0%)\n            Motorola-X:    48 (10.0%)\n     Samsung-Galaxy-S4:    48 (10.0%)\n             iPhone-4s:    48 (10.0%)\n           LG-Nexus-5x:    48 (10.0%)\n      Motorola-Nexus-6:    48 (10.0%)\n  Samsung-Galaxy-Note3:    48 (10.0%)\n            Sony-NEX-7:    48 (10.0%)\nEpoch 1/200\n165/954 [====&gt;.........................] - ETA: 14:01 - loss: 2.2015 - acc: 0.2167Exception in thread Thread-8:\nTraceback (most recent call last):\n  File \"/usr/lib/python2.7/threading.py\", line 801, in __bootstrap_inner\n    self.run()\n  File \"/usr/lib/python2.7/threading.py\", line 754, in run\n    self.__target(*self.__args, **self.__kwargs)\n  File \"/usr/lib/python2.7/multiprocessing/pool.py\", line 389, in _handle_results\n    task = get()\nTypeError: ('__init__() takes exactly 3 arguments (2 given)',",
      "votes": null
    },
    {
      "id": "275275",
      "postDate": "01/28/2018 16:18:53",
      "content": "<p>You have broken jpegs in your train. I use code like this to check and remove any offending files:</p>\n\n<pre><code>def check_remove_broken(img_path):\ntry:\n    x = jpeg.JPEG(img_path).decode()\nexcept Exception:\n    print('Decoding error:', img_path)\n    os.remove(img_path)\n\np = Pool(cpu_count() - 2)\np.map(check_remove_broken, tqdm(ids_train))\n</code></pre>",
      "rawMarkdown": "You have broken jpegs in your train. I use code like this to check and remove any offending files:\n\n    def check_remove_broken(img_path):\n    try:\n        x = jpeg.JPEG(img_path).decode()\n    except Exception:\n        print('Decoding error:', img_path)\n        os.remove(img_path)\n\n    p = Pool(cpu_count() - 2)\n    p.map(check_remove_broken, tqdm(ids_train))",
      "votes": null
    },
    {
      "id": "275277",
      "postDate": "01/28/2018 16:25:00",
      "content": "<p>Thanks a lot for your timely and helpful response, I will try it.</p>",
      "rawMarkdown": "Thanks a lot for your timely and helpful response, I will try it.",
      "votes": null
    },
    {
      "id": "275285",
      "postDate": "01/28/2018 16:43:19",
      "content": "<p>Make sure to reload ids_train after cleaning or your will get \"File not found\" errors instead :)</p>",
      "rawMarkdown": "Make sure to reload ids_train after cleaning or your will get \"File not found\" errors instead :)",
      "votes": null
    },
    {
      "id": "275293",
      "postDate": "01/28/2018 17:33:55",
      "content": "<p>I removed 2 bad images, but got error again, It seems still broken image problem</p>\n\n<p>249/954 [======&gt;.......................] - ETA: 12:20 - loss: 2.1828 - acc: 0.2269['flickr_images/./sony_nex7/37130939252_0b932c54bb_o.jpg', 'flickr_images/./moto_maxx/38349362074_91f99e4ed8_o.jpg', 'flickr_images/./iphone_4s/38769820101_87c2d4fb0c_o.jpg', 'flickr_images/./samsung_s4/37180498790_5f224f9468_o.jpg', '../train/Motorola-X/(MotoX)106.jpg', 'flickr_images/./moto_maxx/38351113384_7295402f17_o.jpg', 'flickr_images/./moto_maxx/38348060164_76a98301fe_o.jpg', 'flickr_images/./nexus_6/36259418946_76c889cca2_o.jpg']\n250/954 [======&gt;.......................] - ETA: 12:18 - loss: 2.1825 - acc: 0.2280['flickr_images/./htc_m7/35835183885_0ee8244504_o.jpg', '../train/Motorola-X/(MotoX)8.jpg', 'flickr_images/./samsung_s4/37983843806_3c0336c1fe_o.jpg', 'flickr_images/./moto_x/23931672383_701bb2cb6b_o_d.jpg', 'flickr_images/./moto_x/31257251621_cfeb71556d_o_d.jpg', '../train/Motorola-Nexus-6/(MotoNex6)103.jpg', 'flickr_images/./samsung_s4/36962967424_6851c6e1d8_o.jpg', 'flickr_images/./iphone_6/27638535039_94c75816b2_o.jpg']\n251/954 [======&gt;.......................] - ETA: 12:17 - loss: 2.1818 - acc: 0.2286['flickr_images/./samsung_s4/26260305859_75196e6006_o.jpg', 'flickr_images/./iphone_6/38718281414_c16867e84d_o.jpg', 'flickr_images/./sony_nex7/23485515358_af58b1be1e_o.jpg', 'flickr_images/./iphone_6/39398306292_ab49153dc1_o.jpg', '../train/Samsung-Galaxy-Note3/(GalaxyN3)114.jpg', '../train/Motorola-X/(MotoX)167.jpg', 'flickr_images/./nexus_6/37396245501_50d30cc8bf_o.jpg', 'flickr_images/./nexus_6/37823866942_4b5e78914f_o.jpg']\n252/954 [======&gt;.......................] - ETA: 12:16 - loss: 2.1814 - acc: 0.2287Traceback (most recent call last):\n  File \"train.py\", line 605, in \n    class_weight=class_weight)\n  File \"/usr/local/lib/python2.7/dist-packages/keras/legacy/interfaces.py\", line 91, in wrapper\n    return func(*args, **kwargs)\n  File \"/usr/local/lib/python2.7/dist-packages/keras/engine/training.py\", line 2145, in fit_generator\n    generator_output = next(output_generator)\n  File \"/usr/local/lib/python2.7/dist-packages/keras/utils/data_utils.py\", line 770, in get\n    six.reraise(value.<strong>class</strong>, value, value.<strong>traceback</strong>)\n  File \"/usr/local/lib/python2.7/dist-packages/keras/utils/data_utils.py\", line 635, in _data_generator_task\n    generator_output = next(self._generator)\n  File \"train.py\", line 371, in gen\n    X[batch_idx], O[batch_idx], y[batch_idx] = batch_result\nValueError: could not broadcast input array from shape (0,512,3) into shape (512,512,3)\nroot@yz114:/data/work/camera_model/sp-society-camera-model-identification# </p>",
      "rawMarkdown": "I removed 2 bad images, but got error again, It seems still broken image problem\n\n249/954 [======&gt;.......................] - ETA: 12:20 - loss: 2.1828 - acc: 0.2269['flickr_images/./sony_nex7/37130939252_0b932c54bb_o.jpg', 'flickr_images/./moto_maxx/38349362074_91f99e4ed8_o.jpg', 'flickr_images/./iphone_4s/38769820101_87c2d4fb0c_o.jpg', 'flickr_images/./samsung_s4/37180498790_5f224f9468_o.jpg', '../train/Motorola-X/(MotoX)106.jpg', 'flickr_images/./moto_maxx/38351113384_7295402f17_o.jpg', 'flickr_images/./moto_maxx/38348060164_76a98301fe_o.jpg', 'flickr_images/./nexus_6/36259418946_76c889cca2_o.jpg']\n250/954 [======&gt;.......................] - ETA: 12:18 - loss: 2.1825 - acc: 0.2280['flickr_images/./htc_m7/35835183885_0ee8244504_o.jpg', '../train/Motorola-X/(MotoX)8.jpg', 'flickr_images/./samsung_s4/37983843806_3c0336c1fe_o.jpg', 'flickr_images/./moto_x/23931672383_701bb2cb6b_o_d.jpg', 'flickr_images/./moto_x/31257251621_cfeb71556d_o_d.jpg', '../train/Motorola-Nexus-6/(MotoNex6)103.jpg', 'flickr_images/./samsung_s4/36962967424_6851c6e1d8_o.jpg', 'flickr_images/./iphone_6/27638535039_94c75816b2_o.jpg']\n251/954 [======&gt;.......................] - ETA: 12:17 - loss: 2.1818 - acc: 0.2286['flickr_images/./samsung_s4/26260305859_75196e6006_o.jpg', 'flickr_images/./iphone_6/38718281414_c16867e84d_o.jpg', 'flickr_images/./sony_nex7/23485515358_af58b1be1e_o.jpg', 'flickr_images/./iphone_6/39398306292_ab49153dc1_o.jpg', '../train/Samsung-Galaxy-Note3/(GalaxyN3)114.jpg', '../train/Motorola-X/(MotoX)167.jpg', 'flickr_images/./nexus_6/37396245501_50d30cc8bf_o.jpg', 'flickr_images/./nexus_6/37823866942_4b5e78914f_o.jpg']\n252/954 [======&gt;.......................] - ETA: 12:16 - loss: 2.1814 - acc: 0.2287Traceback (most recent call last):\n  File \"train.py\", line 605, in",
      "votes": null
    },
    {
      "id": "275325",
      "postDate": "01/28/2018 20:25:45",
      "content": "<p>Nice work... Thats impressive!</p>",
      "rawMarkdown": "Nice work... Thats impressive!",
      "votes": null
    },
    {
      "id": "275329",
      "postDate": "01/28/2018 20:46:37",
      "content": "<p>The recv return _ForkingPickler.loads(buf.getbuffer()) TypeError: init() missing 1...  error seems to be related to either non jpg files or files missing/corrupted.  I replaced:</p>\n\n<pre><code>img = load_img_fast_jpg(item)\n</code></pre>\n\n<p>with\n    try:\n        img = jpeg.JPEG(item).decode()\n    except:\n        img = np.array(Image.open(item))</p>\n\n<p>(I will add another try block in my code for the Image.open as well, I just didn't get around to doing it)</p>\n\n<p>Furthermore I made sure the directories AND files in Flicker do match Andres files. I particularly had to look at his: flickr_images/good_jpgs and flickr_images/low-quality files.</p>\n\n<p>After downloading some of the images I didn't have, removing some entries I was able to get his code to run without seeing the particular error. The run was using Matt Kleinsmith's docker container, since I still didn't trust my environment to have the right versions of anything anymore.</p>\n\n<p>To further improve the code I think the proper solution would be to change the return None within process_item and change it so that it can return from the \"child\" processes properly without causing issues in Python's pool management and then at the \"parent\" handle error conditions prior to: for batch_result in batch_results: </p>\n\n<p>XiaokangWang if you get a chance make the changes and let us know if that solves it for you as well.</p>\n\n<p>Anyway, I hope this helps anybody that is having reliability issues. I can't believe it took me so long to track the thing down.</p>\n\n<p>Andres, si algun dia estoy en Valencia para la fallas hazme un favor, pasate y me empujas a la hoguera ;-)</p>",
      "rawMarkdown": "The recv return _ForkingPickler.loads(buf.getbuffer()) TypeError: init() missing 1...  error seems to be related to either non jpg files or files missing/corrupted.  I replaced:\n\n    img = load_img_fast_jpg(item)\n\nwith\n    try:\n        img = jpeg.JPEG(item).decode()\n    except:\n        img = np.array(Image.open(item))\n\n(I will add another try block in my code for the Image.open as well, I just didn't get around to doing it)\n\nFurthermore I made sure the directories AND files in Flicker do match Andres files. I particularly had to look at his: flickr_images/good_jpgs and flickr_images/low-quality files.\n\nAfter downloading some of the images I didn't have, removing some entries I was able to get his code to run without seeing the particular error. The run was using Matt Kleinsmith's docker container, since I still didn't trust my environment to have the right versions of anything anymore.\n\nTo further improve the code I think the proper solution would be to change the return None within process_item and change it so that it can return from the \"child\" processes properly without causing issues in Python's pool management and then at the \"parent\" handle error conditions prior to: for batch_result in batch_results: \n\nXiaokangWang if you get a chance make the changes and let us know if that solves it for you as well.\n\nAnyway, I hope this helps anybody that is having reliability issues. I can't believe it took me so long to track the thing down.\n\nAndres, si algun dia estoy en Valencia para la fallas hazme un favor, pasate y me empujas a la hoguera ;-)",
      "votes": null
    },
    {
      "id": "275339",
      "postDate": "01/28/2018 22:00:10",
      "content": "<p>If I have the time I will try the consumer/producer model. I'm curious to see whether is more efficient than <code>Pool</code>. I started reading the book Fluent Python as suggested by Chun Ming Lee.</p>\n\n<p>Re: file mismatch errors hitting you in the face (and you didn't know where the punch came from) yes - the code/error handling is a bit messy, but hey, once you fix all files it works (you need to triple check the files).</p>\n\n<p>Alberto, soy de Alicante... pero igualmente te puedo empujar a la hoguera, de hecho aquí las llamamos Las Hogueras...  :-)</p>",
      "rawMarkdown": "If I have the time I will try the consumer/producer model. I'm curious to see whether is more efficient than `Pool`. I started reading the book Fluent Python as suggested by Chun Ming Lee.\n\nRe: file mismatch errors hitting you in the face (and you didn't know where the punch came from) yes - the code/error handling is a bit messy, but hey, once you fix all files it works (you need to triple check the files).\n\nAlberto, soy de Alicante... pero igualmente te puedo empujar a la hoguera, de hecho aquí las llamamos Las Hogueras...  :-)",
      "votes": null
    },
    {
      "id": "275347",
      "postDate": "01/28/2018 23:09:31",
      "content": "<p>Thanks Andres for your reply. There are some image that can not be read. </p>",
      "rawMarkdown": "Thanks Andres for your reply. There are some image that can not be read.",
      "votes": null
    },
    {
      "id": "275352",
      "postDate": "01/28/2018 23:41:43",
      "content": "<p>Hi Alberto, I will give it a try. I was thinking why we just use Image.open rather than both jpeg and Image.open</p>",
      "rawMarkdown": "Hi Alberto, I will give it a try. I was thinking why we just use Image.open rather than both jpeg and Image.open",
      "votes": null
    },
    {
      "id": "275374",
      "postDate": "01/29/2018 01:03:36",
      "content": "<p>That should also work, as it is a more generalized library. There is a performance hit, although I haven't timed it to see its significance.</p>",
      "rawMarkdown": "That should also work, as it is a more generalized library. There is a performance hit, although I haven't timed it to see its significance.",
      "votes": null
    },
    {
      "id": "275386",
      "postDate": "01/29/2018 02:12:52",
      "content": "<p>@Andres, I wouldn't spend too much time on optimizing your MP code. Parallel code is extremely bug-prone and the only reason I have a decent working implementation is that I spent half of a previous competition (CDiscount) working outs bugs. </p>\n\n<p>And with &lt;=2 GPUs, your CPUs won't be the bottleneck. </p>",
      "rawMarkdown": "Andres, I wouldn't spend too much time on optimizing your MP code. Parallel code is extremely bug-prone and the only reason I have a decent working implementation is that I spent half of a previous competition (CDiscount) working outs bugs. \n\nAnd with &lt;=2 GPUs, your CPUs won't be the bottleneck.",
      "votes": null
    },
    {
      "id": "275387",
      "postDate": "01/29/2018 02:16:01",
      "content": "<p>At the margins we're dealing with, they do matter. </p>\n\n<p>To give you an example, until a couple months back, Keras' implementation of RNNs (LSTM &amp; GRU) were substantially slower than bare-metal TF implementations. </p>\n\n<p>And this extends to stuff like pre-trained model weights leading to materially different results depending on whether you use TF + TF-weights or Theano + Theano weights. </p>",
      "rawMarkdown": "At the margins we're dealing with, they do matter. \n\nTo give you an example, until a couple months back, Keras' implementation of RNNs (LSTM &amp; GRU) were substantially slower than bare-metal TF implementations. \n\nAnd this extends to stuff like pre-trained model weights leading to materially different results depending on whether you use TF + TF-weights or Theano + Theano weights.",
      "votes": null
    },
    {
      "id": "275404",
      "postDate": "01/29/2018 03:36:34",
      "content": "<p>It seems not data problem, I tried without -x option, still got\nbatch_result ValueError: could not broadcast input array from shape (0,512,3) into shape (512,512,3) </p>",
      "rawMarkdown": "It seems not data problem, I tried without -x option, still got\nbatch_result ValueError: could not broadcast input array from shape (0,512,3) into shape (512,512,3)",
      "votes": null
    },
    {
      "id": "275407",
      "postDate": "01/29/2018 03:47:18",
      "content": "<p>How in the world did you get a 0? That looks like a bad image for sure.</p>",
      "rawMarkdown": "How in the world did you get a 0? That looks like a bad image for sure.",
      "votes": null
    },
    {
      "id": "275449",
      "postDate": "01/29/2018 06:49:27",
      "content": "<p>It seems not a bad image.\nI printed  the image name, but got the same error from on different images. \nsome times, it was shape (0,512,3) into shape (512,512,3)\nsome times, shape (512,0,3) into shape (512,512,3)</p>",
      "rawMarkdown": "It seems not a bad image.\nI printed  the image name, but got the same error from on different images. \nsome times, it was shape (0,512,3) into shape (512,512,3)\nsome times, shape (512,0,3) into shape (512,512,3)",
      "votes": null
    },
    {
      "id": "275453",
      "postDate": "01/29/2018 06:55:27",
      "content": "<p>Run it with -v to see the shapes of the image. I had that bug a few days ago but was fixed. It was due to wrong offsets in random crops. </p>",
      "rawMarkdown": "Run it with -v to see the shapes of the image. I had that bug a few days ago but was fixed. It was due to wrong offsets in random crops.",
      "votes": null
    },
    {
      "id": "275454",
      "postDate": "01/29/2018 07:03:49",
      "content": "<p>How to fix it</p>",
      "rawMarkdown": "How to fix it",
      "votes": null
    },
    {
      "id": "275455",
      "postDate": "01/29/2018 07:12:45",
      "content": "<p>Are you using the latest code in the repo? It's fixed there.</p>",
      "rawMarkdown": "Are you using the latest code in the repo? It's fixed there.",
      "votes": null
    },
    {
      "id": "275466",
      "postDate": "01/29/2018 07:52:32",
      "content": "<p>I checked again, it is the latest code.</p>",
      "rawMarkdown": "I checked again, it is the latest code.",
      "votes": null
    },
    {
      "id": "275492",
      "postDate": "01/29/2018 09:35:00",
      "content": "<p>I changed the random crops part as follows, then I can finish a whole epoch now</p>\n\n<pre><code>if random_crop:\n    freedom_x, freedom_y = img.shape[1] - crop_size, img.shape[0] - crop_size\n    if freedom_x &gt; 0:\n        center_x += np.random.randint(math.ceil(-freedom_x/2)+1, freedom_x - math.ceil(freedom_x/2)-1 )\n    if freedom_y &gt; 0:\n        center_y += np.random.randint(math.ceil(-freedom_y/2)+1, freedom_y - math.ceil(freedom_y/2)-1 )\n</code></pre>",
      "rawMarkdown": "I changed the random crops part as follows, then I can finish a whole epoch now\n\n    if random_crop:\n        freedom_x, freedom_y = img.shape[1] - crop_size, img.shape[0] - crop_size\n        if freedom_x &gt; 0:\n            center_x += np.random.randint(math.ceil(-freedom_x/2)+1, freedom_x - math.ceil(freedom_x/2)-1 )\n        if freedom_y &gt; 0:\n            center_y += np.random.randint(math.ceil(-freedom_y/2)+1, freedom_y - math.ceil(freedom_y/2)-1 )",
      "votes": null
    },
    {
      "id": "275509",
      "postDate": "01/29/2018 11:07:43",
      "content": "<p>Yeah, better safe than sorry and 2 pixels is no biggie. I wonder why Im not experiencing it, b/c I had the same issue and applied <code>math.ceil</code> and <code>math.floor</code> very carefully to avoid edge conditions.</p>",
      "rawMarkdown": "Yeah, better safe than sorry and 2 pixels is no biggie. I wonder why Im not experiencing it, b/c I had the same issue and applied `math.ceil` and `math.floor` very carefully to avoid edge conditions.",
      "votes": null
    },
    {
      "id": "275521",
      "postDate": "01/29/2018 12:07:43",
      "content": "<p>@Chun Ming Lee, thanks for your reply!</p>",
      "rawMarkdown": "Chun Ming Lee, thanks for your reply!",
      "votes": null
    },
    {
      "id": "275524",
      "postDate": "01/29/2018 12:18:06",
      "content": "<p>Some updates folks. My first observation is that vggish type of networks are not a good option for the task at hand. Slow training times and low accuracy. Batch size has quite an impact on accuracy, that goes for all type of networks. </p>",
      "rawMarkdown": "Some updates folks. My first observation is that vggish type of networks are not a good option for the task at hand. Slow training times and low accuracy. Batch size has quite an impact on accuracy, that goes for all type of networks.",
      "votes": null
    },
    {
      "id": "275858",
      "postDate": "01/30/2018 05:46:05",
      "content": "<p>Hey Andres Torrubia, Is the high pass filter that you used from a slide working? I am a Student,trying several approaches with severe failures. Any suggestions to improve my score would be appreciated</p>",
      "rawMarkdown": "Hey Andres Torrubia, Is the high pass filter that you used from a slide working? I am a Student,trying several approaches with severe failures. Any suggestions to improve my score would be appreciated",
      "votes": null
    },
    {
      "id": "275947",
      "postDate": "01/30/2018 09:48:11",
      "content": "<p>Thanks for your valuable inputs. Could you help me with why batch size is playing a significant role, I would assume it has to do with how much data we can fit in the ram for training and I would expect with a larger batch size the convergence to happen faster compared to lower batch size assuming other parameters to be same. It would be helpful if you can lead me to why there is a decrease/increase in accuracy based on batch size.</p>\n\n<p>What batch size did you explore for this particular data set ? and what worked best for you?</p>",
      "rawMarkdown": "Thanks for your valuable inputs. Could you help me with why batch size is playing a significant role, I would assume it has to do with how much data we can fit in the ram for training and I would expect with a larger batch size the convergence to happen faster compared to lower batch size assuming other parameters to be same. It would be helpful if you can lead me to why there is a decrease/increase in accuracy based on batch size.\n\nWhat batch size did you explore for this particular data set ? and what worked best for you?",
      "votes": null
    },
    {
      "id": "275949",
      "postDate": "01/30/2018 09:56:00",
      "content": "<p>Batch size interacts with learning rate, as a rule of thumb, if you increase batch size performs a similar role as decreasing learning rate, and learning rate is probably the single most important hyper-parameter to fine tune to find fast convergence. See also <a href=\"https://arxiv.org/abs/1711.00489\">Don't Decay the Learning Rate, Increase the Batch Size</a></p>\n\n<p>Also, classifiers/feature extractors use Batch Normalization which is very depending on batch size. See <a href=\"https://arxiv.org/abs/1502.03167\">Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift</a>. As a fun aside, what is covariate shift?</p>\n\n<p>I'm using Keras so I have to find a good LR by hand. Fast.ai just released their Pytorch-based framework that supports the Learning Rate Finder algorithm <a href=\"http://www.fast.ai/2018/01/26/v2-launch/\">fast.ai v2</a></p>",
      "rawMarkdown": "Batch size interacts with learning rate, as a rule of thumb, if you increase batch size performs a similar role as decreasing learning rate, and learning rate is probably the single most important hyper-parameter to fine tune to find fast convergence. See also [Don't Decay the Learning Rate, Increase the Batch Size][1]\n\nAlso, classifiers/feature extractors use Batch Normalization which is very depending on batch size. See [Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift][2]. As a fun aside, what is covariate shift?\n\nI'm using Keras so I have to find a good LR by hand. Fast.ai just released their Pytorch-based framework that supports the Learning Rate Finder algorithm [fast.ai v2][3]\n\n\n  [1]: https://arxiv.org/abs/1711.00489 \"Don't Decay the Learning Rate, Increase the Batch Size\"\n  [2]: https://arxiv.org/abs/1502.03167 \"Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift\"\n  [3]: http://www.fast.ai/2018/01/26/v2-launch/",
      "votes": null
    },
    {
      "id": "275964",
      "postDate": "01/30/2018 10:32:31",
      "content": "<p>The way I picture the batch size trade-off is -</p>\n\n<ol>\n<li>Smaller batches: noisier gradient updates but potentially gets you out of sharp local minima</li>\n<li>Larger batches: cleaner gradient updates potentially converging faster, but at the risk of getting stuck in a local minima.</li>\n</ol>\n\n<p>As Andres mentioned, there's a fair bit of ongoing research about this with two general schools of thought - the \"small batch size is better\" camp, and the Facebook etc. camp which has produced research claiming they can scale up to BS of thousands by playing around with LR and other settings. </p>\n\n<p>What I've seen in the few Kaggle competitions I've participated in is - it depends. You're going to have test out various combinations on the problem you're working on.</p>\n\n<p>The choice of optimizer is pretty important as well. As a newbie, I defaulted to Adam but there's literature suggesting that adaptive optimizers (e.g., Adam, RMSProp etc.) generally perform worse or at best equal to vanilla SGD algorithms. (<a href=\"https://arxiv.org/pdf/1705.08292.pdf\">\"The Marginal Value of Adaptive Gradient Methods in Machine Learning\"</a>)</p>",
      "rawMarkdown": "The way I picture the batch size trade-off is -\n\n 1. Smaller batches: noisier gradient updates but potentially gets you out of sharp local minima\n 2. Larger batches: cleaner gradient updates potentially converging faster, but at the risk of getting stuck in a local minima.\n\nAs Andres mentioned, there's a fair bit of ongoing research about this with two general schools of thought - the \"small batch size is better\" camp, and the Facebook etc. camp which has produced research claiming they can scale up to BS of thousands by playing around with LR and other settings. \n\nWhat I've seen in the few Kaggle competitions I've participated in is - it depends. You're going to have test out various combinations on the problem you're working on.\n\nThe choice of optimizer is pretty important as well. As a newbie, I defaulted to Adam but there's literature suggesting that adaptive optimizers (e.g., Adam, RMSProp etc.) generally perform worse or at best equal to vanilla SGD algorithms. ([\"The Marginal Value of Adaptive Gradient Methods in Machine Learning\"][1])\n\n  [1]: https://arxiv.org/pdf/1705.08292.pdf",
      "votes": null
    },
    {
      "id": "275973",
      "postDate": "01/30/2018 10:59:59",
      "content": "<p>@Andres Torrubia Thanks for your contribution so far, has been a great learning experience for me. I read the article you suggested \"Don't Decay the Learning Rate, Increase the Batch Size\", it was an interesting read, however, it was mostly around how It reaches equivalent test accuracies after the same number of training epochs leading to greater parallelism and shorter training times, whereas I was inferring from @kirk's post that it was affecting the accuracies on validation(assuming time is not a constraint here). </p>\n\n<p>@Chung Ming Lee Thanks for your explanation, based on your points one can infer that small batches can lead to better results at the cost of  greater time complexity as they have lesser likelihood of getting stuck in local minima, I am not sure should I take this inference as conclusion because you mentioned there are researchers in this area working to better validate it.\nI'll read further on the literature you suggested to get a better hold of it.</p>",
      "rawMarkdown": "Andres Torrubia Thanks for your contribution so far, has been a great learning experience for me. I read the article you suggested \"Don't Decay the Learning Rate, Increase the Batch Size\", it was an interesting read, however, it was mostly around how It reaches equivalent test accuracies after the same number of training epochs leading to greater parallelism and shorter training times, whereas I was inferring from @kirk's post that it was affecting the accuracies on validation(assuming time is not a constraint here). \n\n@Chung Ming Lee Thanks for your explanation, based on your points one can infer that small batches can lead to better results at the cost of  greater time complexity as they have lesser likelihood of getting stuck in local minima, I am not sure should I take this inference as conclusion because you mentioned there are researchers in this area working to better validate it.\nI'll read further on the literature you suggested to get a better hold of it.",
      "votes": null
    },
    {
      "id": "276112",
      "postDate": "01/30/2018 18:15:37",
      "content": "<p>I am getting below error:\nTraceback (most recent call last):\n  File \"train.py\", line 591, in \n    class_weight=class_weight)\n  File \"/home/rbhat/anaconda2/envs/mob/lib/python3.5/site-packages/keras/legacy/interfaces.py\", line 91, in wrapper\n    return func(*args, **kwargs)\n  File \"/home/rbhat/anaconda2/envs/mob/lib/python3.5/site-packages/keras/engine/training.py\", line 2145, in fit_generator\n    generator_output = next(output_generator)\n  File \"/home/rbhat/anaconda2/envs/mob/lib/python3.5/site-packages/keras/utils/data_utils.py\", line 770, in get\n    six.reraise(value.<strong>class</strong>, value, value.<strong>traceback</strong>)\n  File \"/home/rbhat/.local/lib/python3.5/site-packages/six.py\", line 693, in reraise\n    raise value\n  File \"/home/rbhat/anaconda2/envs/mob/lib/python3.5/site-packages/keras/utils/data_utils.py\", line 635, in _data_generator_task\n    generator_output = next(self._generator)\n  File \"train.py\", line 356, in gen\n    batch_results = p.map(process_item_func, item_batch)\n  File \"/home/rbhat/anaconda2/envs/mob/lib/python3.5/multiprocessing/pool.py\", line 266, in map\n    return self._map_async(func, iterable, mapstar, chunksize).get()\n  File \"/home/rbhat/anaconda2/envs/mob/lib/python3.5/multiprocessing/pool.py\", line 644, in get\n    raise self._value\nOSError: Could not load libjpeg-turbo library</p>\n\n<p>Does anyone faced the same issue?</p>",
      "rawMarkdown": "I am getting below error:\nTraceback (most recent call last):\n  File \"train.py\", line 591, in",
      "votes": null
    },
    {
      "id": "276128",
      "postDate": "01/30/2018 19:00:07",
      "content": "<p>You need to install libjpeg-turbo </p>\n\n<pre><code>sudo apt install libturbojpeg\n</code></pre>",
      "rawMarkdown": "You need to install libjpeg-turbo \n\n\n\n    sudo apt install libturbojpeg",
      "votes": null
    },
    {
      "id": "276203",
      "postDate": "01/31/2018 00:29:25",
      "content": "<p>@Andres Check <code>clr_callback.py</code> a.k.a rate finder. Disclaimer, for me it didn't work. What I mean is that you'll have to wait 2-8 times the number of iterations per epoch for a phase to change in rate finding. I'd rather kill myself than have to wait that much time. I'm pretty sure that anyone can do much better at finding suitable lr by plug and play than waiting for any automatic algorithm to find the best optimum.</p>",
      "rawMarkdown": "Andres Check `clr_callback.py` a.k.a rate finder. Disclaimer, for me it didn't work. What I mean is that you'll have to wait 2-8 times the number of iterations per epoch for a phase to change in rate finding. I'd rather kill myself than have to wait that much time. I'm pretty sure that anyone can do much better at finding suitable lr by plug and play than waiting for any automatic algorithm to find the best optimum.",
      "votes": null
    },
    {
      "id": "276241",
      "postDate": "01/31/2018 02:48:55",
      "content": "<p>add the lib64 directory to LD_LIBRARY_PATH</p>",
      "rawMarkdown": "add the lib64 directory to LD_LIBRARY_PATH",
      "votes": null
    },
    {
      "id": "276292",
      "postDate": "01/31/2018 05:33:41",
      "content": "<p>@kirk Im going to try it. This looks like a LR sweeper rather than a LR finder. I will let you know...</p>",
      "rawMarkdown": "kirk Im going to try it. This looks like a LR sweeper rather than a LR finder. I will let you know...",
      "votes": null
    },
    {
      "id": "276451",
      "postDate": "01/31/2018 14:22:33",
      "content": "<p>Hey @Andres, I just have a question. I've noticed that whenever I resume training after checkpointing a model at some good <code>x</code> accuracy I always pay a loss of <code>x-20%</code> and no matter how long I train I can never hit again the same <code>x</code> accuracy. Have you experienced the same issue? And there is a discussion on github saying that there is an error with <code>save.model()</code> not saving the state of the optimizer.</p>",
      "rawMarkdown": "Hey @Andres, I just have a question. I've noticed that whenever I resume training after checkpointing a model at some good `x` accuracy I always pay a loss of `x-20%` and no matter how long I train I can never hit again the same `x` accuracy. Have you experienced the same issue? And there is a discussion on github saying that there is an error with `save.model()` not saving the state of the optimizer.",
      "votes": null
    },
    {
      "id": "276458",
      "postDate": "01/31/2018 14:49:21",
      "content": "<p>What Keras version are you using?</p>",
      "rawMarkdown": "What Keras version are you using?",
      "votes": null
    },
    {
      "id": "276477",
      "postDate": "01/31/2018 15:33:04",
      "content": "<pre><code>In [3]: keras.__version__\nOut[3]: '2.0.8'\n</code></pre>\n\n<p>When you resume training do you also use the flags <code>-l</code>, <code>-uiw</code>?</p>",
      "rawMarkdown": "In [3]: keras.__version__\n    Out[3]: '2.0.8'\n\nWhen you resume training do you also use the flags `-l`, `-uiw`?",
      "votes": null
    },
    {
      "id": "276480",
      "postDate": "01/31/2018 15:38:58",
      "content": "<p>I use 2.1.3</p>\n\n<p>Biggest reason for penalty is learning rate is not saved, so if you don't specify it it will default to the initial so, check the last learning rate and put it with -l. -uiw not needed (has no effect ) when used with -m or -w</p>",
      "rawMarkdown": "I use 2.1.3\n\nBiggest reason for penalty is learning rate is not saved, so if you don't specify it it will default to the initial so, check the last learning rate and put it with -l. -uiw not needed (has no effect ) when used with -m or -w",
      "votes": null
    },
    {
      "id": "276488",
      "postDate": "01/31/2018 15:46:53",
      "content": "<p>Sorry my mistake I was checking keras version in terminal without being ssh in the actual server. I have the same keras version= 2.1.3. About your second comment here is the difficult part, during training the checkpoints are saving the model at best val_accuray but we don't have information on the actual learning rate at that stage. In the end we have a model saved in <code>.hdf5</code> format. How can we know the learning rate used when that particular checkpoint was saved?</p>",
      "rawMarkdown": "Sorry my mistake I was checking keras version in terminal without being ssh in the actual server. I have the same keras version= 2.1.3. About your second comment here is the difficult part, during training the checkpoints are saving the model at best val_accuray but we don't have information on the actual learning rate at that stage. In the end we have a model saved in `.hdf5` format. How can we know the learning rate used when that particular checkpoint was saved?",
      "votes": null
    },
    {
      "id": "276512",
      "postDate": "01/31/2018 16:48:44",
      "content": "<p>You may just guess it (or calculate). There is schedule in the script, halving LR after 5 epochs after last improvement. Just take a look at string of saved models you can decode between which of them there was a halving. And initial rate is given </p>",
      "rawMarkdown": "You may just guess it (or calculate). There is schedule in the script, halving LR after 5 epochs after last improvement. Just take a look at string of saved models you can decode between which of them there was a halving. And initial rate is given",
      "votes": null
    },
    {
      "id": "276520",
      "postDate": "01/31/2018 16:59:12",
      "content": "<p>@Sergey thanks. A minimal example would help. Particularly I am interested in this bit <code>Just take a look at string of saved models</code>. I am assuming that you imply that I should have all the model checkpoints in place. What if I have deleted all of them apart the one with highest accuracy. I am also very interested to know if there is an option saving the learning rate along with the model when you use <code>ModelCheckpoint</code> callback?</p>",
      "rawMarkdown": "Sergey thanks. A minimal example would help. Particularly I am interested in this bit `Just take a look at string of saved models`. I am assuming that you imply that I should have all the model checkpoints in place. What if I have deleted all of them apart the one with highest accuracy. I am also very interested to know if there is an option saving the learning rate along with the model when you use `ModelCheckpoint` callback?",
      "votes": null
    },
    {
      "id": "276541",
      "postDate": "01/31/2018 17:44:58",
      "content": "<p>Keras saves each model with highest metric, so if you have several of them for example: \nepoch1  0.3\nepoch2  0.35\nepoch3  0.41\nepoch12 0.51\nepoch15 0.55\nepoch25 0.6\nyou can see where halving happen: between model where there were more then 5 epochs between models</p>",
      "rawMarkdown": "Keras saves each model with highest metric, so if you have several of them for example: \nepoch1\t0.3\nepoch2\t0.35\nepoch3\t0.41\nepoch12\t0.51\nepoch15\t0.55\nepoch25\t0.6\nyou can see where halving happen: between model where there were more then 5 epochs between models",
      "votes": null
    },
    {
      "id": "276570",
      "postDate": "01/31/2018 19:32:00",
      "content": "<p>Check documentation of <code>ModelCheckpoint</code> and see whether learning save in the <code>log</code> key which were the keywords like <code>val_acc</code> <code>epoch</code> etc are made available and it's how I construct the filename for the model, then if LR is available there place it accordingly in the filename and extract it with the <code>re.match</code> ... I use to extract <code>epoch</code> from the filename upon loading. LMK if it works.</p>",
      "rawMarkdown": "Check documentation of `ModelCheckpoint` and see whether learning save in the `log` key which were the keywords like `val_acc` `epoch` etc are made available and it's how I construct the filename for the model, then if LR is available there place it accordingly in the filename and extract it with the `re.match` ... I use to extract `epoch` from the filename upon loading. LMK if it works.",
      "votes": null
    },
    {
      "id": "276624",
      "postDate": "01/31/2018 23:58:09",
      "content": "<p>@Andres thanks for the reply. Let's make a concrete example. Here is what is saved <code>VGG19_do0.3_doc0.0_avg-epoch086-val_acc0.329167.hdf5</code>. From this I don't really understand how you can extract the learning rate? How do you resume training in your case? Where do you get the learning rate from?</p>",
      "rawMarkdown": "Andres thanks for the reply. Let's make a concrete example. Here is what is saved `VGG19_do0.3_doc0.0_avg-epoch086-val_acc0.329167.hdf5`. From this I don't really understand how you can extract the learning rate? How do you resume training in your case? Where do you get the learning rate from?",
      "votes": null
    },
    {
      "id": "276733",
      "postDate": "02/01/2018 10:00:27",
      "content": "<p>What I do is look at the output of training, e.g:</p>\n\n<pre><code>poch 19/200\n449/450 [============================&gt;.] - ETA: 0s - loss: 0.3424 - acc: 0.8899\nEpoch 00019: ReduceLROnPlateau reducing learning rate to 4.999999873689376e-05.\n450/450 [==============================] - 186s 413ms/step - loss: 0.3420 - acc: 0.8900 - val_loss: 0.6567 - val_acc: 0.8356\nEpoch 20/200\n450/450 [==============================] - 186s 414ms/step - loss: 0.2574 - acc: 0.9219 - val_loss: 0.4105 - val_acc: 0.9008\n</code></pre>\n\n<p>check the latest <code>ReduceLROnPlateau</code> output, i.e. <code>Epoch 00019: ReduceLROnPlateau reducing learning rate to 4.999999873689376e-05</code> and when I restart training manually add <code>-l 5e-5</code>.</p>",
      "rawMarkdown": "What I do is look at the output of training, e.g:\n\n    poch 19/200\n    449/450 [============================&gt;.] - ETA: 0s - loss: 0.3424 - acc: 0.8899\n    Epoch 00019: ReduceLROnPlateau reducing learning rate to 4.999999873689376e-05.\n    450/450 [==============================] - 186s 413ms/step - loss: 0.3420 - acc: 0.8900 - val_loss: 0.6567 - val_acc: 0.8356\n    Epoch 20/200\n    450/450 [==============================] - 186s 414ms/step - loss: 0.2574 - acc: 0.9219 - val_loss: 0.4105 - val_acc: 0.9008\n\ncheck the latest `ReduceLROnPlateau` output, i.e. `Epoch 00019: ReduceLROnPlateau reducing learning rate to 4.999999873689376e-05` and when I restart training manually add `-l 5e-5`.",
      "votes": null
    },
    {
      "id": "277053",
      "postDate": "02/02/2018 03:44:07",
      "content": "<p>Great approach! Way to go.</p>",
      "rawMarkdown": "Great approach! Way to go.",
      "votes": null
    },
    {
      "id": "277308",
      "postDate": "02/02/2018 18:56:02",
      "content": "<p>Hi @Andres, apologies I couldn't respond earlier, I was traveling. I am assuming that you get the LR from the training output which probably you have saved somewhere. I don't have that. Imagine that you trained days ago in tmux or session. And now you've released that session and you don't have that info anymore, you haven't saved it anywhere, is there any other alternative? I tried <code>model = load_model()</code>, <code>model.fit()</code> but couldn't get any info like the one you gave: <code>449/450 [============================&gt;.] - ETA: 0s - loss: 0.3424 - acc: 0.8899\nEpoch 00019: ReduceLROnPlateau reducing learning rate to 4.999999873689376e-05.</code></p>",
      "rawMarkdown": "Hi @Andres, apologies I couldn't respond earlier, I was traveling. I am assuming that you get the LR from the training output which probably you have saved somewhere. I don't have that. Imagine that you trained days ago in tmux or session. And now you've released that session and you don't have that info anymore, you haven't saved it anywhere, is there any other alternative? I tried `model = load_model()`, `model.fit()` but couldn't get any info like the one you gave: `449/450 [============================&gt;.] - ETA: 0s - loss: 0.3424 - acc: 0.8899\nEpoch 00019: ReduceLROnPlateau reducing learning rate to 4.999999873689376e-05.`",
      "votes": null
    },
    {
      "id": "277315",
      "postDate": "02/02/2018 19:14:07",
      "content": "<p>Yes, I get the training for the output, using <code>tmux</code>.</p>\n\n<p>I have digged into the saved model and it does indeed contain the last LR:</p>\n\n<pre><code>&gt; from keras.models import load_model, Model\n&gt; from keras import backend as K\n&gt; model = load_model('models/your_model_____.hdf5')\n&gt; model.optimizer.__dict__.keys()\ndict_keys(['updates', 'weights', 'iterations', 'lr', 'beta_1', 'beta_2', 'decay', 'epsilon', 'initial_decay', 'amsgrad'])\n&gt; model.optimizer.__dict__['lr']\n&lt;tf.Variable 'Adam/lr:0' shape=() dtype=float32_ref&gt;\n&gt; K.eval(model.optimizer.__dict__['lr'])\n5e-05\n</code></pre>\n\n<p>Now up to you...</p>",
      "rawMarkdown": "Yes, I get the training for the output, using `tmux`.\n\nI have digged into the saved model and it does indeed contain the last LR:\n\n    &gt; from keras.models import load_model, Model\n    &gt; from keras import backend as K\n    &gt; model = load_model('models/your_model_____.hdf5')\n    &gt; model.optimizer.__dict__.keys()\n    dict_keys(['updates', 'weights', 'iterations', 'lr', 'beta_1', 'beta_2', 'decay', 'epsilon', 'initial_decay', 'amsgrad'])\n    &gt; model.optimizer.__dict__['lr']",
      "votes": null
    },
    {
      "id": "277355",
      "postDate": "02/02/2018 21:46:46",
      "content": "<p>@Andres thanks a lot but, here's what I get:\n<img src=\"http://i63.tinypic.com/2r3cql1.png\" alt=\"error\"></p>",
      "rawMarkdown": "Andres thanks a lot but, here's what I get:\n![error][1]\n\n\n  [1]: http://i63.tinypic.com/2r3cql1.png",
      "votes": null
    },
    {
      "id": "277357",
      "postDate": "02/02/2018 21:55:28",
      "content": "<p>I use nohup to save my process and help me retrospect better for future runs.</p>",
      "rawMarkdown": "I use nohup to save my process and help me retrospect better for future runs.",
      "votes": null
    },
    {
      "id": "277359",
      "postDate": "02/02/2018 22:04:49",
      "content": "<p>You need to be on Keras 2.1.3, TF 1.4.1 and Python 3.6</p>\n\n<p>If you saved that model w/ older version of Keras you're probably out of luck with that one.</p>",
      "rawMarkdown": "You need to be on Keras 2.1.3, TF 1.4.1 and Python 3.6\n\nIf you saved that model w/ older version of Keras you're probably out of luck with that one.",
      "votes": null
    },
    {
      "id": "277369",
      "postDate": "02/02/2018 22:57:52",
      "content": "<p>@andres and @kirk can you help me with using Densenet201, I tried passing denset as parameter but I get this error,</p>\n\n<pre><code>python train.py -g 1 -b 8 -cs 512 -cm Densenet210 -x -l 1e-4 -uiw \n</code></pre>\n\n<p>Error:</p>\n\n<pre><code>Traceback (most recent call last):\n File \"train.py\", line 466, in &lt;module&gt;\n classifier = globals()[args.classifier]\n KeyError: 'Densenet210'\n</code></pre>\n\n<p>any help is appreciated. Resnet50 works fine for me though</p>",
      "rawMarkdown": "andres and @kirk can you help me with using Densenet201, I tried passing denset as parameter but I get this error,\n\n    python train.py -g 1 -b 8 -cs 512 -cm Densenet210 -x -l 1e-4 -uiw \n\nError:\n\n    Traceback (most recent call last):\n     File \"train.py\", line 466, in",
      "votes": null
    },
    {
      "id": "277378",
      "postDate": "02/02/2018 23:29:57",
      "content": "<p>Change \"Densenet210\" to \"DenseNet201\".</p>",
      "rawMarkdown": "Change \"Densenet210\" to \"DenseNet201\".",
      "votes": null
    },
    {
      "id": "277380",
      "postDate": "02/02/2018 23:48:04",
      "content": "<p>Apologies for such a typo, but I did try DenseNet201, However issue was with keras version.\nUpdating keras fixed it.</p>",
      "rawMarkdown": "Apologies for such a typo, but I did try DenseNet201, However issue was with keras version.\nUpdating keras fixed it.",
      "votes": null
    },
    {
      "id": "277389",
      "postDate": "02/03/2018 00:37:37",
      "content": "<p>@Andres it is keras 2.1.3 and python 3.6 and tf 1.4.1.</p>",
      "rawMarkdown": "Andres it is keras 2.1.3 and python 3.6 and tf 1.4.1.",
      "votes": null
    },
    {
      "id": "277463",
      "postDate": "02/03/2018 08:18:14",
      "content": "<p>You need to have the model compiled to access <code>.optimizer</code>, in my latest code:</p>\n\n<p><code>model = load_model(args.model, compile=False if args.test or (args.learning_rate is not None) else True)</code></p>",
      "rawMarkdown": "You need to have the model compiled to access `.optimizer`, in my latest code:\n\n`model = load_model(args.model, compile=False if args.test or (args.learning_rate is not None) else True)`",
      "votes": null
    },
    {
      "id": "277543",
      "postDate": "02/03/2018 14:20:14",
      "content": "<p>Where can I find the lib.inputgenerator  and  lib.classhelper?</p>",
      "rawMarkdown": "Where can I find the lib.inputgenerator  and  lib.classhelper?",
      "votes": null
    },
    {
      "id": "277548",
      "postDate": "02/03/2018 14:36:23",
      "content": "<p>If you are trying to use my fork please don't. I moved away from it because it was generating the wrong results. There was a bug causing wrong labeling during training and I never fixed it. Go ahead and use the official one from Andres: <a href=\"https://github.com/antorsae/sp-society-camera-model-identification\">https://github.com/antorsae/sp-society-camera-model-identification</a>\nSorry for any confusion.</p>",
      "rawMarkdown": "If you are trying to use my fork please don't. I moved away from it because it was generating the wrong results. There was a bug causing wrong labeling during training and I never fixed it. Go ahead and use the official one from Andres: https://github.com/antorsae/sp-society-camera-model-identification\nSorry for any confusion.",
      "votes": null
    },
    {
      "id": "277627",
      "postDate": "02/03/2018 19:00:08",
      "content": "<p>@Andres awesome thanks a bunch ;)</p>",
      "rawMarkdown": "Andres awesome thanks a bunch ;)",
      "votes": null
    },
    {
      "id": "277901",
      "postDate": "02/04/2018 19:16:12",
      "content": "<p>I give up with this fucking keras. No more, I've had enough. Even after setting the learning rate as @Andres suggested here is what I see.\n<img src=\"http://i66.tinypic.com/34phpgm.png\" alt=\"Training\"></p>",
      "rawMarkdown": "I give up with this fucking keras. No more, I've had enough. Even after setting the learning rate as @Andres suggested here is what I see.\n![Training][1]\n\n\n  [1]: http://i66.tinypic.com/34phpgm.png",
      "votes": null
    },
    {
      "id": "277999",
      "postDate": "02/05/2018 03:40:05",
      "content": "<p>python train.py -g 1 -b 8 -cs 229 -cm Densenet201 -l 5e-5 -uiw</p>\n\n<p>Traceback (most recent call last):\n  File \"train.py\", line 467, in \n    classifier = globals()[args.classifier]\nKeyError: 'Densenet201'</p>\n\n<p>Resnet50 works fine, How do you fix this?</p>",
      "rawMarkdown": "python train.py -g 1 -b 8 -cs 229 -cm Densenet201 -l 5e-5 -uiw\n\nTraceback (most recent call last):\n  File \"train.py\", line 467, in",
      "votes": null
    },
    {
      "id": "278002",
      "postDate": "02/05/2018 03:52:57",
      "content": "<p>Your keras is old version, updating it will work</p>",
      "rawMarkdown": "Your keras is old version, updating it will work",
      "votes": null
    },
    {
      "id": "278007",
      "postDate": "02/05/2018 04:48:30",
      "content": "<p>@YangLu update your keras to keras 2.1.3, previous keras doesn't have DenseNet pre-trained models.</p>\n\n<p><a href=\"https://keras.io/applications/\">Keras Documentation</a></p>",
      "rawMarkdown": "YangLu update your keras to keras 2.1.3, previous keras doesn't have DenseNet pre-trained models.\n\n[Keras Documentation][1]\n\n\n  [1]: https://keras.io/applications/",
      "votes": null
    },
    {
      "id": "278039",
      "postDate": "02/05/2018 07:18:32",
      "content": "<p>My keras is 2.1.3, I use serveral conda envs, maybe this disturb.</p>",
      "rawMarkdown": "My keras is 2.1.3, I use serveral conda envs, maybe this disturb.",
      "votes": null
    },
    {
      "id": "278042",
      "postDate": "02/05/2018 07:26:41",
      "content": "<p>check if this gets executed in your python console from your present environment:</p>\n\n<pre><code>from keras.applications.densenet import DenseNet201\n</code></pre>\n\n<p>If this works out well, then you have keras 2.1.3 and  are good to go.</p>",
      "rawMarkdown": "check if this gets executed in your python console from your present environment:\n\n    from keras.applications.densenet import DenseNet201\n\nIf this works out well, then you have keras 2.1.3 and  are good to go.",
      "votes": null
    },
    {
      "id": "278044",
      "postDate": "02/05/2018 07:33:34",
      "content": "<p>YangLu, how about to try DenseNet201, not Densenet201?</p>",
      "rawMarkdown": "YangLu, how about to try DenseNet201, not Densenet201?",
      "votes": null
    },
    {
      "id": "278104",
      "postDate": "02/05/2018 11:37:01",
      "content": "<p>Thanks @RK @shivrajp @Yaozj, It's my low level spelling mistakes. By the way, -cs 512 is too big for me , my one 1070Ti  can  work on \"-cs 229\".</p>",
      "rawMarkdown": "Thanks @RK @shivrajp @Yaozj, It's my low level spelling mistakes. By the way, -cs 512 is too big for me , my one 1070Ti  can  work on \"-cs 229\".",
      "votes": null
    },
    {
      "id": "278321",
      "postDate": "02/05/2018 22:31:21",
      "content": "<p>Yes， I have got the same problem， DenseNet needs huge memory.\nMoreover, the loss decrease very slowly, Maybe the learning rate is not set properly</p>",
      "rawMarkdown": "Yes， I have got the same problem， DenseNet needs huge memory.\nMoreover, the loss decrease very slowly, Maybe the learning rate is not set properly",
      "votes": null
    },
    {
      "id": "278347",
      "postDate": "02/06/2018 00:48:11",
      "content": "<p>Awesome work! Thanks!</p>",
      "rawMarkdown": "Awesome work! Thanks!",
      "votes": null
    },
    {
      "id": "278512",
      "postDate": "02/06/2018 11:18:52",
      "content": "<p>I have another problem wiht extra dataset, Have you fixed it?\n541/953 [================&gt;.............] - ETA: 2:36 - loss: 2.0163 - acc: 0.3073\nException in thread Thread-8:\nFile \"/home/yl/miniconda3/envs/gluon/lib/python3.6/multiprocessing/connection.py\", line 251, in recv\n    return _ForkingPickler.loads(buf.getbuffer())\nTypeError: <strong>init</strong>() missing 1 required positional argument: 'code'</p>\n\n<h1>Add script as follows , fix it.</h1>\n\n<p>load_img  = lambda img_path: np.array(Image.open(img_path))</p>\n\n<p>def load_img_fast_jpg(img_path):\n    try:\n        x = jpeg.JPEG(img_path).decode()\n        return x\n    except:\n        return load_img(img_path)</p>",
      "rawMarkdown": "I have another problem wiht extra dataset, Have you fixed it?\n541/953 [================&gt;.............] - ETA: 2:36 - loss: 2.0163 - acc: 0.3073\nException in thread Thread-8:\nFile \"/home/yl/miniconda3/envs/gluon/lib/python3.6/multiprocessing/connection.py\", line 251, in recv\n    return _ForkingPickler.loads(buf.getbuffer())\nTypeError: __init__() missing 1 required positional argument: 'code'\n\n# Add script as follows , fix it.\n\nload_img  = lambda img_path: np.array(Image.open(img_path))\n\ndef load_img_fast_jpg(img_path):\n    try:\n        x = jpeg.JPEG(img_path).decode()\n        return x\n    except:\n        return load_img(img_path)",
      "votes": null
    },
    {
      "id": "278883",
      "postDate": "02/07/2018 02:15:26",
      "content": "<p>Hi guys,</p>\n\n<p>When I use the preprocessing kernel filters # <a href=\"http://www.lirmm.fr/~chaumont/publications/WIFS-2016_TUAMA_COMBY_CHAUMONT_Camera_Model_Identification_With_CNN_slides.pdf\">http://www.lirmm.fr/~chaumont/publications/WIFS-2016_TUAMA_COMBY_CHAUMONT_Camera_Model_Identification_With_CNN_slides.pdf</a></p>\n\n<pre><code> kernel_filter = 1/12. * np.array([\\\n        [-1,  2,  -2,  2, -1],  \\\n        [ 2, -6,   8, -6,  2],  \\\n        [-2,  8, -12,  8, -2],  \\\n        [ 2, -6,   8, -6,  2],  \\\n        [-1,  2,  -2,  2, -1]]) \n</code></pre>\n\n<p>The images after the filter are basically completely black with speckles of white. Is this supposed to be the case? </p>\n\n<p>Thanks for the help!</p>",
      "rawMarkdown": "Hi guys,\n\nWhen I use the preprocessing kernel filters # http://www.lirmm.fr/~chaumont/publications/WIFS-2016_TUAMA_COMBY_CHAUMONT_Camera_Model_Identification_With_CNN_slides.pdf\n       \n     kernel_filter = 1/12. * np.array([\\\n            [-1,  2,  -2,  2, -1],  \\\n            [ 2, -6,   8, -6,  2],  \\\n            [-2,  8, -12,  8, -2],  \\\n            [ 2, -6,   8, -6,  2],  \\\n            [-1,  2,  -2,  2, -1]]) \n\nThe images after the filter are basically completely black with speckles of white. Is this supposed to be the case? \n\nThanks for the help!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 274085,
      "author_name": "ceperaang",
      "author_url": "",
      "post_date": "01/25/2018 21:34:20",
      "content": "<p>You're just killing it, man!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 274169,
      "author_name": "leecming",
      "author_url": "",
      "post_date": "01/26/2018 01:15:41",
      "content": "<p>Very nice Andres, thanks for sharing your approach! </p>\n\n<p>Couple (minor) optimization suggestions - </p>\n\n<ol>\n<li><p>Multiprocessing - Have a look at the consumer-producer(s) design pattern. Invoking pool-map on every batch yield isn't particularly efficient. I suspect that your GPUs are the bottleneck which is masking this issue. As an aside, GIL doesn't mean you should always default to processes over threads. Most functions in C-native Python libraries such as numpy release the GIL. The book \"Fluent Python\" is a great read to learn more about this and other lower level language structures.</p></li>\n<li><p>Increasing your GPU throughput - You've mentioned gradient check-pointing but also have a look at replacing the TF backend with MXNet. Note that this would require using a forked version of Keras.</p></li>\n<li><p>Downclocking your GPUs - this might be controversial but consider downclocking your GPUs if you have a hobbyist home setup (use nvidia-smi to set persistence mode with the PM flag, and then using the PL flag to set the target TDP). If you're running your GPUs 24/7 like most of us are, it reduces thermal wear in the long-term for little-to-no performance decrease. I personally run my 1080 Tis at 200W (down from the default 270W) and have measured a &lt; 3% decrease. </p></li>\n</ol>",
      "votes": null,
      "replies": [
        {
          "id": 274287,
          "author_name": "antorsae",
          "author_url": "",
          "post_date": "01/26/2018 08:49:03",
          "content": "<p>Thanks for the suggestions. Just bought the \"Fluent Python\" book. :-) Will also try your other suggestions, Im going to check CNTK too. </p>\n\n<p>I believe using a MXNet would preclude me from using imagenet pretrained weights for latest networks (NasNet, DenseNet, etc.). Do you know if it is the case?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 274306,
          "author_name": "leecming",
          "author_url": "",
          "post_date": "01/26/2018 09:42:36",
          "content": "<p>That's my understanding but I'm happy to be corrected. My team-mate and I have stuck to vanilla Keras because of that. So there's a trade-off between flexibility and performance. </p>\n\n<p>In any case, we're not using any of the cutting edge models (e.g., DenseNet, NASNet) for this competition. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 274308,
          "author_name": "antorsae",
          "author_url": "",
          "post_date": "01/26/2018 09:50:41",
          "content": "<p>Im curious: are you using an ensemble of models or a single model for your score?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 274321,
          "author_name": "leecming",
          "author_url": "",
          "post_date": "01/26/2018 10:34:13",
          "content": "<p>We're ensembling but by default rather than having benchmarked it against single models.  I doubt this is a major factor. </p>\n\n<p>Having observed the trajectory of the top teams' scores for the past month, I suspect we are sniffing around the same local optima and the spread in performance is largely due to hyper-parameters/randomness.</p>\n\n<p>Hopefully someone will come in with a novel solution and blow the LB up!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 274329,
          "author_name": "kleinsmith",
          "author_url": "",
          "post_date": "01/26/2018 11:06:54",
          "content": "<p>Thanks for the power limit tip.</p>\n\n<blockquote>\n  <p>set persistence mode with the PM flag</p>\n</blockquote>\n\n<p><a href=\"http://docs.nvidia.com/deploy/driver-persistence/index.html#persistence-daemon\">NVIDIA said</a> they'll eventually stop supporting persistence mode, and are focusing on the persistence daemon. <a href=\"http://docs.nvidia.com/deploy/driver-persistence/index.html#installation\">Getting the daemon to start at startup</a> is a little weird, though. On Ubuntu, I had to unpack /usr/share/doc/NVIDIA_GLX-1.0/sample/nvidia-persistenced-init.tar.bz2 and then run the install script. <a href=\"https://devtalk.nvidia.com/default/topic/995248/cuda-setup-and-installation/setting-up-nvidia-persistenced/post/5090647/#5090647\">One person says</a> to add \"--persistence-mode\" to the service start line in the template file, but that seems unneeded.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 274965,
          "author_name": "albertoa",
          "author_url": "",
          "post_date": "01/27/2018 22:57:10",
          "content": "<p>Chun, thanks for the note on power. I'm curious though, why set the power limit instead of the temperature thresholds: slow_threshold and max_threshold?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 275247,
          "author_name": "kirk86",
          "author_url": "",
          "post_date": "01/28/2018 15:21:19",
          "content": "<p>I have a question @Chun Ming Lee regarding point 2. Does is make any difference which backend you use? Ultimately now all of them rely on the same lower layer which is cuda and cudnn? Which IMHO it won't make any big difference but then again I might be wrong. Thanks!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 275387,
          "author_name": "leecming",
          "author_url": "",
          "post_date": "01/29/2018 02:16:01",
          "content": "<p>At the margins we're dealing with, they do matter. </p>\n\n<p>To give you an example, until a couple months back, Keras' implementation of RNNs (LSTM &amp; GRU) were substantially slower than bare-metal TF implementations. </p>\n\n<p>And this extends to stuff like pre-trained model weights leading to materially different results depending on whether you use TF + TF-weights or Theano + Theano weights. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 275521,
          "author_name": "kirk86",
          "author_url": "",
          "post_date": "01/29/2018 12:07:43",
          "content": "<p>@Chun Ming Lee, thanks for your reply!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 274203,
      "author_name": "albertoa",
      "author_url": "",
      "post_date": "01/26/2018 03:23:44",
      "content": "<p>That's awesome</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 274386,
      "author_name": "kirk86",
      "author_url": "",
      "post_date": "01/26/2018 14:19:27",
      "content": "<p>Hey everyone for me the code is not working. It's stuck like this for almost a day now:</p>\n\n<pre><code>       HTC-1-M7:  1023 (13.0%)\n       iPhone-6:   823 (10.5%)\n       Motorola-Droid-Maxx:   825 (10.5%)\n       Motorola-X:   275 (03.5%)\n       Samsung-Galaxy-S4:  1412 (18.0%)\n       iPhone-4s:   774 (09.9%)\n       LG-Nexus-5x:   680 (08.7%)\n       Motorola-Nexus-6:   926 (11.8%)\n       Samsung-Galaxy-Note3:   548 (07.0%)\n       Sony-NEX-7:   557 (07.1%)\n       validation steps = 0\n       Epoch 1/200\n</code></pre>\n\n<p>Any suggestions?</p>",
      "votes": null,
      "replies": [
        {
          "id": 274388,
          "author_name": "antorsae",
          "author_url": "",
          "post_date": "01/26/2018 14:25:17",
          "content": "<p>I assume you're running it with <code>-x</code>, i.e. using Gleb's dataset. \nDid you sync the <code>val_images</code> directory @ <a href=\"https://github.com/antorsae/sp-society-camera-model-identification/tree/master/val_images\">https://github.com/antorsae/sp-society-camera-model-identification/tree/master/val_images</a> and downloaded all images?</p>\n\n<p>When using <code>-x</code> I wanted to have the validation as distinct as possible to the train set, so:</p>\n\n<pre><code>    ids_train = ids\n    ids_val   = [ ]\n\n    extra_train_ids = [os.path.join(EXTRA_TRAIN_FOLDER,line.rstrip('\\n')) for line in open(os.path.join(EXTRA_TRAIN_FOLDER, 'good_jpgs'))]\n    extra_train_ids.sort()\n    ids_train.extend(extra_train_ids)\n\n    extra_val_ids = glob.glob(join(EXTRA_VAL_FOLDER,'*/*.jpg'))\n    extra_val_ids.sort()\n    ids_val.extend(extra_val_ids)\n</code></pre>\n\n<p>do a <code>print(ids_val)</code> after that line and LMK what you get.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 274392,
          "author_name": "kirk86",
          "author_url": "",
          "post_date": "01/26/2018 14:33:44",
          "content": "<p>Hi Andres, thanks for the quick response. I haven't downloaded all the images, I that was happening automatically. Should I'll be also creating any additional directories for the images? On another note I am getting the value of <code>validation_steps=0</code> in the fit.generator so as a consequence I had to hardcode the value to 6 in order for the code not  to throw an error. Let me try download all the images and run again give you back the output of <code>print(ids_val)</code>.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 274434,
          "author_name": "antorsae",
          "author_url": "",
          "post_date": "01/26/2018 15:41:34",
          "content": "<p>Download Gleb's training set by running <code>find . -name \"urls_*\" -execdir wget -nc --tries=10 -i {} \\;</code> in the <code>flickr_images</code> directory and adjust accordingly in the <code>val_images</code> one.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 274449,
          "author_name": "albertoa",
          "author_url": "",
          "post_date": "01/26/2018 16:12:43",
          "content": "<p>Also don't forget to rename any .JPG to .jpg as mentioned on the other thread.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 274491,
          "author_name": "kirk86",
          "author_url": "",
          "post_date": "01/26/2018 17:52:03",
          "content": "<p>Thank you Andres &amp; Alberto, we are back in business fellas :).</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 274913,
          "author_name": "kleinsmith",
          "author_url": "",
          "post_date": "01/27/2018 18:06:40",
          "content": "<blockquote>\n  <p><strong>Andres Torrubia wrote</strong></p>\n  \n  <blockquote>\n    <p>Download Gleb's training set by running <code>find . -name \"urls_*\" -execdir wget -nc --tries=10 -i {} \\;</code> in the <code>flickr_images</code> directory and adjust accordingly in the <code>val_images</code> one.</p>\n  </blockquote>\n</blockquote>\n\n<p>Here's the breakdown of this command.</p>\n\n<ul>\n<li><code>find</code>: returns file paths that match a given pattern</li>\n<li><code>.</code>: search working directory recursively (search its subdirectories, their subdirectories, and so on)</li>\n<li><code>-name \"urls_*\"</code>: the given file path pattern</li>\n<li><code>-execdir</code>: execute the given command on each file path found, <strong>in the directory of the file.</strong></li>\n<li><code>wget</code>: download what a given URL points to</li>\n<li><code>-nc</code>: don't download files that already exist</li>\n<li><code>-i</code>: read URLs from a given file</li>\n<li><code>{}</code>: placeholder for each file that <code>find</code> found</li>\n<li><code>\\;</code>: end command</li>\n</ul>",
          "votes": null,
          "replies": []
        },
        {
          "id": 274916,
          "author_name": "antorsae",
          "author_url": "",
          "post_date": "01/27/2018 18:15:59",
          "content": "<p>One caveat when doing this on <code>val_images</code>. In the <code>val_images/moto_maxx</code> directory a bunch of images will get downloaed with *.1 *.2 *.3 ... extensions. Rename them to _1.jpg _2.jpg _3.jpg otherwise the code will not use them and if using <code>-x</code> validation set for that class will be heavily underrepresented. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 274437,
      "author_name": "jfkingiii",
      "author_url": "",
      "post_date": "01/26/2018 15:48:31",
      "content": "<p>I got 0.946 on the public LB using Andres code with \n<code>python train.py -g 1 -b 16 -cs 512 -cm ResNet50 -x -l 1e-4 -uiw</code>\ntraining for 91 epochs on a single V100. Training took about 10 hours. Thanks Andres, I have learned a lot from looking at your code.</p>",
      "votes": null,
      "replies": [
        {
          "id": 274548,
          "author_name": "antorsae",
          "author_url": "",
          "post_date": "01/26/2018 19:26:37",
          "content": "<p>Good to hear. There's a bug in the code re: augmentation. LMK if you can spot it :-) I will submit fixed version with other improvements this weekend.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 274558,
          "author_name": "",
          "author_url": "",
          "post_date": "01/26/2018 19:41:27",
          "content": "<p>is bug in the train part?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 274561,
          "author_name": "antorsae",
          "author_url": "",
          "post_date": "01/26/2018 19:45:38",
          "content": "<p>Yes, may or may not decrease accuracy (I guess it does), but I want to retrain and see tomorrow. Will submit then.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 274578,
          "author_name": "aamaia",
          "author_url": "",
          "post_date": "01/26/2018 20:15:51",
          "content": "<p>I noticed something in the gamma augmentation, the formula would be pow(x, 1/gamma) but it's written as pow(x, gamma); I think it doesn't matter much as 1/0.8==1.25 and 1/1.2==0.83, so in the end gamma08 generates gamma 1.2 and vice-versa :-)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 274581,
          "author_name": "antorsae",
          "author_url": "",
          "post_date": "01/26/2018 20:21:29",
          "content": "<p>The bug I was referring to is not related to gamma. Re: gamma by looking at <a href=\"http://scikit-image.org/docs/dev/api/skimage.exposure.html#skimage.exposure.adjust_gamma\">http://scikit-image.org/docs/dev/api/skimage.exposure.html#skimage.exposure.adjust_gamma</a> the formula is O = I**gamma so I think the current formula is OK; why should it be 1/gamma?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 274585,
          "author_name": "jfkingiii",
          "author_url": "",
          "post_date": "01/26/2018 20:34:31",
          "content": "<p>There is some confusion about which formula is actually \"gamma correction\":\n<a href=\"https://stackoverflow.com/questions/16521003/gamma-correction-formula-gamma-or-1-gamma\">https://stackoverflow.com/questions/16521003/gamma-correction-formula-gamma-or-1-gamma</a></p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 274616,
      "author_name": "CVxTz",
      "author_url": "",
      "post_date": "01/26/2018 22:05:08",
      "content": "<p>I am getting this error :</p>\n\n<p>$ python train.py -g 1 -b 16 -cs 512 -cm ResNet50 -l 1e-4 -uiw</p>\n\n<p>2018-01-26 23:01:46.708921: E tensorflow/stream_executor/cuda/cuda_dnn.cc:385] could not create cudnn handle: CUDNN_STATUS_INTERNAL_ERROR\n2018-01-26 23:01:46.708954: E tensorflow/stream_executor/cuda/cuda_dnn.cc:352] could not destroy cudnn handle: CUDNN_STATUS_BAD_PARAM\n2018-01-26 23:01:46.708961: F tensorflow/core/kernels/conv_ops.cc:667] Check failed: stream-&gt;parent()-&gt;GetConvolveAlgorithms( conv_parameters.ShouldIncludeWinogradNonfusedAlgo</p>\n\n<p>Everything works fine If I run Resnet50 from keras examples.</p>\n\n<p>Anyone has an idea what might be the reason ? </p>",
      "votes": null,
      "replies": [
        {
          "id": 274626,
          "author_name": "neongen",
          "author_url": "",
          "post_date": "01/26/2018 23:00:21",
          "content": "<p>Maybe it is a versioning issue. What about downgrading keras?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 274628,
          "author_name": "CVxTz",
          "author_url": "",
          "post_date": "01/26/2018 23:04:24",
          "content": "<p>I use the same keras version and tf as OP \nWhat version are you using ?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 274631,
          "author_name": "neongen",
          "author_url": "",
          "post_date": "01/26/2018 23:14:59",
          "content": "<p>keras 2.1.2, tensorflow 1.4.0 and it works well. I also tried to upgrade keras to 2.1.3 and got the error you mentioned above. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 274635,
          "author_name": "antorsae",
          "author_url": "",
          "post_date": "01/26/2018 23:30:38",
          "content": "<p>I use Keras 2.1.3 and TF 1.4.1</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 274638,
          "author_name": "ceperaang",
          "author_url": "",
          "post_date": "01/26/2018 23:40:23",
          "content": "<p>Check your CUDA &amp; cudnn versions. Default Tensorflow from pip compiled with cuda 9.0 and cudnn 7, so you need exactly them.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 274649,
          "author_name": "ceperaang",
          "author_url": "",
          "post_date": "01/27/2018 01:09:42",
          "content": "<p>Another possible reason: Anaconda. It messes with cublas libraries and tensorflow going crazy. Solution: probably, simplest way -- just use system python for this code.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 274782,
          "author_name": "CVxTz",
          "author_url": "",
          "post_date": "01/27/2018 09:48:17",
          "content": "<p>I created a fresh env of py3.5and it worked, before I was using py3.6.</p>\n\n<p>Thanks guys !</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 274860,
          "author_name": "albertoa",
          "author_url": "",
          "post_date": "01/27/2018 14:58:24",
          "content": "<p>Can you share how you created the environment. I still haven't been able to run Andres code anymore. I'm actually creating a new code base and deconstructing it part by part.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 274864,
          "author_name": "CVxTz",
          "author_url": "",
          "post_date": "01/27/2018 15:09:43",
          "content": "<p>i use anaconda : \n<a href=\"http://uoa-eresearch.github.io/eresearch-cookbook/recipe/2014/11/20/conda/\">http://uoa-eresearch.github.io/eresearch-cookbook/recipe/2014/11/20/conda/</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 274868,
          "author_name": "neongen",
          "author_url": "",
          "post_date": "01/27/2018 15:18:32",
          "content": "<p>On linux I use <a href=\"https://github.com/pypa/pipenv\">pipenv</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 274871,
          "author_name": "albertoa",
          "author_url": "",
          "post_date": "01/27/2018 15:23:48",
          "content": "<p>Yes, but that doesn't give me versions/packages.</p>\n\n<p>I have several virtenvs but something happened trying to upgrade tensorflow and I haven't been able to run Andres code since.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 274879,
          "author_name": "kleinsmith",
          "author_url": "",
          "post_date": "01/27/2018 15:39:00",
          "content": "<p>I'm using Docker for Andres's code:</p>\n\n<pre><code>alias andres=\"docker run --runtime=nvidia --init -it --rm \\\n              --ipc=host \\\n              -v $CODE:/code -v $DATA:/data \\\n              -w=/code/sp-society-camera-model-identification \\\n              mwksmith/cam:andres\"\n</code></pre>\n\n<p>When you get into the container enter \"sa\" to activate the conda environment, and then create sym links to your data folders according to the globals set in train.py. Then run train.py.</p>\n\n<p>I'm pushing the Docker Image to Docker Hub now. You can also build it yourself with <a href=\"https://github.com/antorsae/sp-society-camera-model-identification/blob/master/Dockerfile\">the Dockerfile</a>.</p>\n\n<p>Edit: Removed <code>--shm-size</code>. It appears to be redundant with respect to <code>--ipc=host</code>.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 274918,
          "author_name": "albertoa",
          "author_url": "",
          "post_date": "01/27/2018 18:22:30",
          "content": "<p>Thanks, this may be the fastest approach for me to get his latest release working. I'll try it as soon as the current network training is done.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 274709,
      "author_name": "tunguz",
      "author_url": "",
      "post_date": "01/27/2018 03:53:03",
      "content": "<p>First of all, really appreciate sharing the code. It really, really helps.</p>\n\n<p>I managed to run it on one of my machines, but now when I am trying to run the latest version (downloaded today) on another machine, I am getting really weird number of images in different classes. One of them (Moto X) even has 0 images. However, when I check the folders, I see that all of the images are there. Anyone else have the same issue? Any idea how to fix it?</p>\n\n<p><code>HTC-1-M7:   748 (13.9%) \n              iPhone-6:   548 (10.2%)\n   Motorola-Droid-Maxx:   550 (10.2%)\n            Motorola-X:     0 (00.0%)\n     Samsung-Galaxy-S4:  1137 (21.2%)\n             iPhone-4s:   499 (09.3%)\n           LG-Nexus-5x:   405 (07.5%)\n      Motorola-Nexus-6:   651 (12.1%)\n  Samsung-Galaxy-Note3:   273 (05.1%)\n            Sony-NEX-7:   557 (10.4%)\n</code></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 274761,
      "author_name": "youngnotnaive",
      "author_url": "",
      "post_date": "01/27/2018 08:15:15",
      "content": "<p>Hi, <br>\nI am using Keras, with CUDA_VISIBLE_DEVICES to mask to use 2 GTX1080 GPU,\nbut when I run DenseNet201 model, will another two fc layers,\nIt give me a lot of errors : </p>\n\n<pre><code>ran out of memory trying to allocate XXXX MiB\n</code></pre>\n\n<p>I tried to use 4 GTX1080 GPU,  but still get same errors.\nIt seems that it GPU memory doesn't increase, it performance is same as single GTX1080\nCan you give me some suggestions? What should I do to implement this model to multi-gpu?</p>",
      "votes": null,
      "replies": [
        {
          "id": 274821,
          "author_name": "antorsae",
          "author_url": "",
          "post_date": "01/27/2018 12:45:17",
          "content": "<p>do you do <code>-g 2</code> for 2 GPUs?</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 274773,
      "author_name": "antorsae",
      "author_url": "",
      "post_date": "01/27/2018 09:02:47",
      "content": "<p>OK, getting 96.2 LB now single model.\nCode is updated in the repo:</p>\n\n<ul>\n<li>Added GPL3 license: basically if you modify it, share the code.</li>\n<li>Fixed orientation flip augmentation bug (need to confirm current one is OK.</li>\n<li>Control TTA and prints class distribution after <code>-t</code>.</li>\n</ul>",
      "votes": null,
      "replies": []
    },
    {
      "id": 274872,
      "author_name": "antorsae",
      "author_url": "",
      "post_date": "01/27/2018 15:23:56",
      "content": "<p>Alberto, this may be an overkill but here's my base environment to use with anaconda (I use miniconda, but regular anaconda would work too). There's many packages that are not needed for this project so feel free to edit the file by hand and remove the ones you think are not needed.</p>\n\n<p><a href=\"https://conda.io/docs/user-guide/tasks/manage-environments.html\">https://conda.io/docs/user-guide/tasks/manage-environments.html</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 274917,
          "author_name": "albertoa",
          "author_url": "",
          "post_date": "01/27/2018 18:18:47",
          "content": "<p>Thanks, it may end up being what I need.</p>\n\n<p>I finally finished going through every line and did some re-structuring of the code from a couple of days ago. I made it a little more robust in handling errors. Unfortunately I also got rid of the paralleled implementation of the generator so it is now slower. So far it started to run, so I may finally be in the right track.</p>\n\n<p>For anyone interested my fork is at: <a href=\"https://github.com/albertoa/sp-society-camera-model-identification\">https://github.com/albertoa/sp-society-camera-model-identification</a>\nI would only recommend it at this stage for the added comments. I will improve the generator and look at your latest changes to see what else would be useful. For now I'm just happy I'm back training networks.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 277543,
          "author_name": "yyll008",
          "author_url": "",
          "post_date": "02/03/2018 14:20:14",
          "content": "<p>Where can I find the lib.inputgenerator  and  lib.classhelper?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 277548,
          "author_name": "albertoa",
          "author_url": "",
          "post_date": "02/03/2018 14:36:23",
          "content": "<p>If you are trying to use my fork please don't. I moved away from it because it was generating the wrong results. There was a bug causing wrong labeling during training and I never fixed it. Go ahead and use the official one from Andres: <a href=\"https://github.com/antorsae/sp-society-camera-model-identification\">https://github.com/antorsae/sp-society-camera-model-identification</a>\nSorry for any confusion.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 274940,
      "author_name": "antorsae",
      "author_url": "",
      "post_date": "01/27/2018 21:02:18",
      "content": "<p>Got 0.969 LB w/ single model. \nI think there's a chance to get to 0.98 with a single model.</p>",
      "votes": null,
      "replies": [
        {
          "id": 274945,
          "author_name": "tunguz",
          "author_url": "",
          "post_date": "01/27/2018 21:19:15",
          "content": "<p>With the same code? </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 274952,
          "author_name": "antorsae",
          "author_url": "",
          "post_date": "01/27/2018 21:46:44",
          "content": "<p>Almost. Diff is 20 lines of code. Will submit after more changes I have planned.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 274970,
          "author_name": "godaibo",
          "author_url": "",
          "post_date": "01/27/2018 23:15:07",
          "content": "<p>you better keep some of that code for yourself if you hope to be in the money :)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 274971,
          "author_name": "antorsae",
          "author_url": "",
          "post_date": "01/27/2018 23:18:04",
          "content": "<p>I know. This all adds to the excitement ;-)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 274982,
          "author_name": "ceperaang",
          "author_url": "",
          "post_date": "01/28/2018 00:25:06",
          "content": "<p>I bet top teams are already very busy training <em>a lot</em> of diverse models (many folds probably) and will do it up to the last minute, so we could see a lot of movements near the end :)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 274998,
          "author_name": "kleinsmith",
          "author_url": "",
          "post_date": "01/28/2018 01:27:57",
          "content": "<blockquote>\n  <p><strong>Sergey Mushinskiy wrote</strong></p>\n  \n  <blockquote>\n    <p>I bet top teams are already very busy training <em>a lot</em> of diverse models (many folds probably)</p>\n  </blockquote>\n</blockquote>\n\n<p>Hi Sergey, what do you mean by \"fold\"?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 275125,
          "author_name": "ceperaang",
          "author_url": "",
          "post_date": "01/28/2018 11:32:42",
          "content": "<p>I mean folds like in k-fold cross-validation. Split dataset strategically into, say, 5 parts and train 5 model on 4 of them (different each time) and use 1 as validation. It is very common strategy in general data science but obviously somewhat rare in deep learning (given computational resources needed). However, it allows to build honest second layer model on top of networks predictions.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 275023,
      "author_name": "xiaokangwang",
      "author_url": "",
      "post_date": "01/28/2018 04:16:23",
      "content": "<p>Thanks Andres for sharing your solution. I was running your code. The code run for 44 epochs and stopped spitting the following error. Do you have any ideas how to debug this error?</p>\n\n<p>read Thread-4:\nTraceback (most recent call last):\n  File \"/home/wxk/anaconda2/envs/py3/lib/python3.6/threading.py\", line 916, in _bootstrap_inner\n    self.run()\n  File \"/home/wxk/anaconda2/envs/py3/lib/python3.6/threading.py\", line 864, in run\n    self._target(*self._args, **self._kwargs)\n  File \"/home/wxk/anaconda2/envs/py3/lib/python3.6/multiprocessing/pool.py\", line 429, in _handle_results\n    task = get()\n  File \"/home/wxk/anaconda2/envs/py3/lib/python3.6/multiprocessing/connection.py\", line 251, in recv\n    return _ForkingPickler.loads(buf.getbuffer())\nTypeError: <strong>init</strong>() missing 1 required positional argument: 'code'</p>",
      "votes": null,
      "replies": [
        {
          "id": 275024,
          "author_name": "albertoa",
          "author_url": "",
          "post_date": "01/28/2018 04:22:56",
          "content": "<p>I have seen that before. To me it was happening on epoch 1. There seems to be a problem with the multiprocess code that triggers it, but other than going into a single CPU for the generator I couldn't solve it.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 275028,
          "author_name": "xiaokangwang",
          "author_url": "",
          "post_date": "01/28/2018 04:43:56",
          "content": "<p>Thanks Alberto for your reply. Then how to set the code to one CPU? </p>\n\n<p>BTW, do you know how the follow code works? I was trying to  figure out how the models (Resnet50, Densenet201 and so on) are created.\ngetattr(globals()[classifier_module_name], 'preprocess_input')</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 275033,
          "author_name": "albertoa",
          "author_url": "",
          "post_date": "01/28/2018 05:02:20",
          "content": "<p>I was afraid you were going to ask. I have a fork of the github code, but I found a major issue with it that I will fix tomorrow, gives wrong results ;-(.\nAdres is working on a new release, hopefully his has the fix.\nThe lines in question should be:\n    p = Pool(cpu_count()-2)\n            batch_results = p.map(process_item_func, item_batch)\nand how the batches are put together. But that's where I messed up my fork, so I rather not recommend anything until I know I can in fact run it properly.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 275043,
          "author_name": "xiaokangwang",
          "author_url": "",
          "post_date": "01/28/2018 05:12:57",
          "content": "<p>Thanks Alberto for your kind reply. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 275048,
          "author_name": "albertoa",
          "author_url": "",
          "post_date": "01/28/2018 05:26:36",
          "content": "<p>By the way I am having the same issue with the docker image from Matt Kleinsmith. </p>\n\n<p>[...]\n      File \"/opt/conda/envs/tf/lib/python3.5/multiprocessing/connection.py\", line 251, in recv\n        return ForkingPickler.loads(buf.getbuffer())\n    TypeError: <strong>init</strong>() missing 1 required positional argument: 'code'</p>\n\n<p>I can not tell you why you and I get the error and yet so many others don't. We are the wrong kind of lucky I guess. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 275049,
          "author_name": "xiaokangwang",
          "author_url": "",
          "post_date": "01/28/2018 05:30:46",
          "content": "<p>Thanks Alberto. </p>\n\n<p>Do you know how the networks become global variables in this line: getattr(globals()[classifier_module_name], 'preprocess_input')?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 275054,
          "author_name": "albertoa",
          "author_url": "",
          "post_date": "01/28/2018 05:39:39",
          "content": "<p>The key to that is:</p>\n\n<p>from keras.applications import *</p>\n\n<p>He is using the classifier_to_module dictionary so that the correct preprocess_input function is called for whichever model is being used.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 275058,
          "author_name": "xiaokangwang",
          "author_url": "",
          "post_date": "01/28/2018 05:47:51",
          "content": "<p>I see. Thanks Alberto. I was stilling creating the models myself. Keras is becoming more handy.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 275124,
          "author_name": "antorsae",
          "author_url": "",
          "post_date": "01/28/2018 11:30:30",
          "content": "<p>An exception happened in the mutiprocess code and Python does not tell you the exception (very likely a file was not read correctly) so it tanks with an obscure error. Making the code run in 1 thread only is painfully slow unless you resort to some sort of caching.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 275329,
          "author_name": "albertoa",
          "author_url": "",
          "post_date": "01/28/2018 20:46:37",
          "content": "<p>The recv return _ForkingPickler.loads(buf.getbuffer()) TypeError: init() missing 1...  error seems to be related to either non jpg files or files missing/corrupted.  I replaced:</p>\n\n<pre><code>img = load_img_fast_jpg(item)\n</code></pre>\n\n<p>with\n    try:\n        img = jpeg.JPEG(item).decode()\n    except:\n        img = np.array(Image.open(item))</p>\n\n<p>(I will add another try block in my code for the Image.open as well, I just didn't get around to doing it)</p>\n\n<p>Furthermore I made sure the directories AND files in Flicker do match Andres files. I particularly had to look at his: flickr_images/good_jpgs and flickr_images/low-quality files.</p>\n\n<p>After downloading some of the images I didn't have, removing some entries I was able to get his code to run without seeing the particular error. The run was using Matt Kleinsmith's docker container, since I still didn't trust my environment to have the right versions of anything anymore.</p>\n\n<p>To further improve the code I think the proper solution would be to change the return None within process_item and change it so that it can return from the \"child\" processes properly without causing issues in Python's pool management and then at the \"parent\" handle error conditions prior to: for batch_result in batch_results: </p>\n\n<p>XiaokangWang if you get a chance make the changes and let us know if that solves it for you as well.</p>\n\n<p>Anyway, I hope this helps anybody that is having reliability issues. I can't believe it took me so long to track the thing down.</p>\n\n<p>Andres, si algun dia estoy en Valencia para la fallas hazme un favor, pasate y me empujas a la hoguera ;-)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 275339,
          "author_name": "antorsae",
          "author_url": "",
          "post_date": "01/28/2018 22:00:10",
          "content": "<p>If I have the time I will try the consumer/producer model. I'm curious to see whether is more efficient than <code>Pool</code>. I started reading the book Fluent Python as suggested by Chun Ming Lee.</p>\n\n<p>Re: file mismatch errors hitting you in the face (and you didn't know where the punch came from) yes - the code/error handling is a bit messy, but hey, once you fix all files it works (you need to triple check the files).</p>\n\n<p>Alberto, soy de Alicante... pero igualmente te puedo empujar a la hoguera, de hecho aquí las llamamos Las Hogueras...  :-)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 275347,
          "author_name": "xiaokangwang",
          "author_url": "",
          "post_date": "01/28/2018 23:09:31",
          "content": "<p>Thanks Andres for your reply. There are some image that can not be read. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 275352,
          "author_name": "xiaokangwang",
          "author_url": "",
          "post_date": "01/28/2018 23:41:43",
          "content": "<p>Hi Alberto, I will give it a try. I was thinking why we just use Image.open rather than both jpeg and Image.open</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 275374,
          "author_name": "albertoa",
          "author_url": "",
          "post_date": "01/29/2018 01:03:36",
          "content": "<p>That should also work, as it is a more generalized library. There is a performance hit, although I haven't timed it to see its significance.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 275386,
          "author_name": "leecming",
          "author_url": "",
          "post_date": "01/29/2018 02:12:52",
          "content": "<p>@Andres, I wouldn't spend too much time on optimizing your MP code. Parallel code is extremely bug-prone and the only reason I have a decent working implementation is that I spent half of a previous competition (CDiscount) working outs bugs. </p>\n\n<p>And with &lt;=2 GPUs, your CPUs won't be the bottleneck. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 275051,
      "author_name": "grassknoted",
      "author_url": "",
      "post_date": "01/28/2018 05:33:58",
      "content": "<p>Thanks a lot for sharing your code! Really helps students like me learn a lot! :)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 275246,
      "author_name": "kirk86",
      "author_url": "",
      "post_date": "01/28/2018 15:13:46",
      "content": "<p>Hi everyone I just wanted to update on some things that I've noticed while using @Andres code which he kindly provided to us. @Andres once again thanks for sharing the code with us. In my system I haven't noticed a big difference in regards to gaining speed training from the multiprocessing generator. Incidentally I am using python3 and running the code without the multiprocessing, I see that multiple cores are still utilized with the difference in regards to the multiprocessor generator, that they are not  &gt;90% occupied all the time, which in some cases is not ideal to have all the cores fully loaded, especially in a shared system. No matter what you set the number of processes in the code still all cores are fully occupied.  That being said, there are some alternatives which I haven't tested yet but I thought I should share them here in case they are useful or maybe someone has already tried them. The first alternative is to use sth like <a href=\"http://tensorpack.readthedocs.io/en/latest/\">tensorpack</a>. The other one is to use sth like <code>producer-consumer</code> pattern that @Chun Ming Lee already mentioned. With all the deadlines I didn't have the time to incorporate that into @Andres code but there's a simple example in the <code>consumer-producer.py</code> file which can be easily extended into @Andres code if anyone wants to.  What does though really make a difference regarding training speed is doing multi-gpu training, almost reduces the training time in half plus you might be able to use larger batch_size. That's all for now folks, I'll update later on different models. Cheers!</p>",
      "votes": null,
      "replies": [
        {
          "id": 275524,
          "author_name": "kirk86",
          "author_url": "",
          "post_date": "01/29/2018 12:18:06",
          "content": "<p>Some updates folks. My first observation is that vggish type of networks are not a good option for the task at hand. Slow training times and low accuracy. Batch size has quite an impact on accuracy, that goes for all type of networks. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 275947,
          "author_name": "shivrajp",
          "author_url": "",
          "post_date": "01/30/2018 09:48:11",
          "content": "<p>Thanks for your valuable inputs. Could you help me with why batch size is playing a significant role, I would assume it has to do with how much data we can fit in the ram for training and I would expect with a larger batch size the convergence to happen faster compared to lower batch size assuming other parameters to be same. It would be helpful if you can lead me to why there is a decrease/increase in accuracy based on batch size.</p>\n\n<p>What batch size did you explore for this particular data set ? and what worked best for you?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 275949,
          "author_name": "antorsae",
          "author_url": "",
          "post_date": "01/30/2018 09:56:00",
          "content": "<p>Batch size interacts with learning rate, as a rule of thumb, if you increase batch size performs a similar role as decreasing learning rate, and learning rate is probably the single most important hyper-parameter to fine tune to find fast convergence. See also <a href=\"https://arxiv.org/abs/1711.00489\">Don't Decay the Learning Rate, Increase the Batch Size</a></p>\n\n<p>Also, classifiers/feature extractors use Batch Normalization which is very depending on batch size. See <a href=\"https://arxiv.org/abs/1502.03167\">Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift</a>. As a fun aside, what is covariate shift?</p>\n\n<p>I'm using Keras so I have to find a good LR by hand. Fast.ai just released their Pytorch-based framework that supports the Learning Rate Finder algorithm <a href=\"http://www.fast.ai/2018/01/26/v2-launch/\">fast.ai v2</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 275964,
          "author_name": "leecming",
          "author_url": "",
          "post_date": "01/30/2018 10:32:31",
          "content": "<p>The way I picture the batch size trade-off is -</p>\n\n<ol>\n<li>Smaller batches: noisier gradient updates but potentially gets you out of sharp local minima</li>\n<li>Larger batches: cleaner gradient updates potentially converging faster, but at the risk of getting stuck in a local minima.</li>\n</ol>\n\n<p>As Andres mentioned, there's a fair bit of ongoing research about this with two general schools of thought - the \"small batch size is better\" camp, and the Facebook etc. camp which has produced research claiming they can scale up to BS of thousands by playing around with LR and other settings. </p>\n\n<p>What I've seen in the few Kaggle competitions I've participated in is - it depends. You're going to have test out various combinations on the problem you're working on.</p>\n\n<p>The choice of optimizer is pretty important as well. As a newbie, I defaulted to Adam but there's literature suggesting that adaptive optimizers (e.g., Adam, RMSProp etc.) generally perform worse or at best equal to vanilla SGD algorithms. (<a href=\"https://arxiv.org/pdf/1705.08292.pdf\">\"The Marginal Value of Adaptive Gradient Methods in Machine Learning\"</a>)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 275973,
          "author_name": "shivrajp",
          "author_url": "",
          "post_date": "01/30/2018 10:59:59",
          "content": "<p>@Andres Torrubia Thanks for your contribution so far, has been a great learning experience for me. I read the article you suggested \"Don't Decay the Learning Rate, Increase the Batch Size\", it was an interesting read, however, it was mostly around how It reaches equivalent test accuracies after the same number of training epochs leading to greater parallelism and shorter training times, whereas I was inferring from @kirk's post that it was affecting the accuracies on validation(assuming time is not a constraint here). </p>\n\n<p>@Chung Ming Lee Thanks for your explanation, based on your points one can infer that small batches can lead to better results at the cost of  greater time complexity as they have lesser likelihood of getting stuck in local minima, I am not sure should I take this inference as conclusion because you mentioned there are researchers in this area working to better validate it.\nI'll read further on the literature you suggested to get a better hold of it.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 276203,
          "author_name": "kirk86",
          "author_url": "",
          "post_date": "01/31/2018 00:29:25",
          "content": "<p>@Andres Check <code>clr_callback.py</code> a.k.a rate finder. Disclaimer, for me it didn't work. What I mean is that you'll have to wait 2-8 times the number of iterations per epoch for a phase to change in rate finding. I'd rather kill myself than have to wait that much time. I'm pretty sure that anyone can do much better at finding suitable lr by plug and play than waiting for any automatic algorithm to find the best optimum.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 276292,
          "author_name": "antorsae",
          "author_url": "",
          "post_date": "01/31/2018 05:33:41",
          "content": "<p>@kirk Im going to try it. This looks like a LR sweeper rather than a LR finder. I will let you know...</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 276451,
          "author_name": "kirk86",
          "author_url": "",
          "post_date": "01/31/2018 14:22:33",
          "content": "<p>Hey @Andres, I just have a question. I've noticed that whenever I resume training after checkpointing a model at some good <code>x</code> accuracy I always pay a loss of <code>x-20%</code> and no matter how long I train I can never hit again the same <code>x</code> accuracy. Have you experienced the same issue? And there is a discussion on github saying that there is an error with <code>save.model()</code> not saving the state of the optimizer.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 276458,
          "author_name": "antorsae",
          "author_url": "",
          "post_date": "01/31/2018 14:49:21",
          "content": "<p>What Keras version are you using?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 276477,
          "author_name": "kirk86",
          "author_url": "",
          "post_date": "01/31/2018 15:33:04",
          "content": "<pre><code>In [3]: keras.__version__\nOut[3]: '2.0.8'\n</code></pre>\n\n<p>When you resume training do you also use the flags <code>-l</code>, <code>-uiw</code>?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 276480,
          "author_name": "antorsae",
          "author_url": "",
          "post_date": "01/31/2018 15:38:58",
          "content": "<p>I use 2.1.3</p>\n\n<p>Biggest reason for penalty is learning rate is not saved, so if you don't specify it it will default to the initial so, check the last learning rate and put it with -l. -uiw not needed (has no effect ) when used with -m or -w</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 276488,
          "author_name": "kirk86",
          "author_url": "",
          "post_date": "01/31/2018 15:46:53",
          "content": "<p>Sorry my mistake I was checking keras version in terminal without being ssh in the actual server. I have the same keras version= 2.1.3. About your second comment here is the difficult part, during training the checkpoints are saving the model at best val_accuray but we don't have information on the actual learning rate at that stage. In the end we have a model saved in <code>.hdf5</code> format. How can we know the learning rate used when that particular checkpoint was saved?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 276512,
          "author_name": "ceperaang",
          "author_url": "",
          "post_date": "01/31/2018 16:48:44",
          "content": "<p>You may just guess it (or calculate). There is schedule in the script, halving LR after 5 epochs after last improvement. Just take a look at string of saved models you can decode between which of them there was a halving. And initial rate is given </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 276520,
          "author_name": "kirk86",
          "author_url": "",
          "post_date": "01/31/2018 16:59:12",
          "content": "<p>@Sergey thanks. A minimal example would help. Particularly I am interested in this bit <code>Just take a look at string of saved models</code>. I am assuming that you imply that I should have all the model checkpoints in place. What if I have deleted all of them apart the one with highest accuracy. I am also very interested to know if there is an option saving the learning rate along with the model when you use <code>ModelCheckpoint</code> callback?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 276541,
          "author_name": "ceperaang",
          "author_url": "",
          "post_date": "01/31/2018 17:44:58",
          "content": "<p>Keras saves each model with highest metric, so if you have several of them for example: \nepoch1  0.3\nepoch2  0.35\nepoch3  0.41\nepoch12 0.51\nepoch15 0.55\nepoch25 0.6\nyou can see where halving happen: between model where there were more then 5 epochs between models</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 276570,
          "author_name": "antorsae",
          "author_url": "",
          "post_date": "01/31/2018 19:32:00",
          "content": "<p>Check documentation of <code>ModelCheckpoint</code> and see whether learning save in the <code>log</code> key which were the keywords like <code>val_acc</code> <code>epoch</code> etc are made available and it's how I construct the filename for the model, then if LR is available there place it accordingly in the filename and extract it with the <code>re.match</code> ... I use to extract <code>epoch</code> from the filename upon loading. LMK if it works.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 276624,
          "author_name": "kirk86",
          "author_url": "",
          "post_date": "01/31/2018 23:58:09",
          "content": "<p>@Andres thanks for the reply. Let's make a concrete example. Here is what is saved <code>VGG19_do0.3_doc0.0_avg-epoch086-val_acc0.329167.hdf5</code>. From this I don't really understand how you can extract the learning rate? How do you resume training in your case? Where do you get the learning rate from?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 276733,
          "author_name": "antorsae",
          "author_url": "",
          "post_date": "02/01/2018 10:00:27",
          "content": "<p>What I do is look at the output of training, e.g:</p>\n\n<pre><code>poch 19/200\n449/450 [============================&gt;.] - ETA: 0s - loss: 0.3424 - acc: 0.8899\nEpoch 00019: ReduceLROnPlateau reducing learning rate to 4.999999873689376e-05.\n450/450 [==============================] - 186s 413ms/step - loss: 0.3420 - acc: 0.8900 - val_loss: 0.6567 - val_acc: 0.8356\nEpoch 20/200\n450/450 [==============================] - 186s 414ms/step - loss: 0.2574 - acc: 0.9219 - val_loss: 0.4105 - val_acc: 0.9008\n</code></pre>\n\n<p>check the latest <code>ReduceLROnPlateau</code> output, i.e. <code>Epoch 00019: ReduceLROnPlateau reducing learning rate to 4.999999873689376e-05</code> and when I restart training manually add <code>-l 5e-5</code>.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 277308,
          "author_name": "kirk86",
          "author_url": "",
          "post_date": "02/02/2018 18:56:02",
          "content": "<p>Hi @Andres, apologies I couldn't respond earlier, I was traveling. I am assuming that you get the LR from the training output which probably you have saved somewhere. I don't have that. Imagine that you trained days ago in tmux or session. And now you've released that session and you don't have that info anymore, you haven't saved it anywhere, is there any other alternative? I tried <code>model = load_model()</code>, <code>model.fit()</code> but couldn't get any info like the one you gave: <code>449/450 [============================&gt;.] - ETA: 0s - loss: 0.3424 - acc: 0.8899\nEpoch 00019: ReduceLROnPlateau reducing learning rate to 4.999999873689376e-05.</code></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 277315,
          "author_name": "antorsae",
          "author_url": "",
          "post_date": "02/02/2018 19:14:07",
          "content": "<p>Yes, I get the training for the output, using <code>tmux</code>.</p>\n\n<p>I have digged into the saved model and it does indeed contain the last LR:</p>\n\n<pre><code>&gt; from keras.models import load_model, Model\n&gt; from keras import backend as K\n&gt; model = load_model('models/your_model_____.hdf5')\n&gt; model.optimizer.__dict__.keys()\ndict_keys(['updates', 'weights', 'iterations', 'lr', 'beta_1', 'beta_2', 'decay', 'epsilon', 'initial_decay', 'amsgrad'])\n&gt; model.optimizer.__dict__['lr']\n&lt;tf.Variable 'Adam/lr:0' shape=() dtype=float32_ref&gt;\n&gt; K.eval(model.optimizer.__dict__['lr'])\n5e-05\n</code></pre>\n\n<p>Now up to you...</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 277355,
          "author_name": "kirk86",
          "author_url": "",
          "post_date": "02/02/2018 21:46:46",
          "content": "<p>@Andres thanks a lot but, here's what I get:\n<img src=\"http://i63.tinypic.com/2r3cql1.png\" alt=\"error\"></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 277357,
          "author_name": "shivrajp",
          "author_url": "",
          "post_date": "02/02/2018 21:55:28",
          "content": "<p>I use nohup to save my process and help me retrospect better for future runs.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 277359,
          "author_name": "antorsae",
          "author_url": "",
          "post_date": "02/02/2018 22:04:49",
          "content": "<p>You need to be on Keras 2.1.3, TF 1.4.1 and Python 3.6</p>\n\n<p>If you saved that model w/ older version of Keras you're probably out of luck with that one.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 277369,
          "author_name": "shivrajp",
          "author_url": "",
          "post_date": "02/02/2018 22:57:52",
          "content": "<p>@andres and @kirk can you help me with using Densenet201, I tried passing denset as parameter but I get this error,</p>\n\n<pre><code>python train.py -g 1 -b 8 -cs 512 -cm Densenet210 -x -l 1e-4 -uiw \n</code></pre>\n\n<p>Error:</p>\n\n<pre><code>Traceback (most recent call last):\n File \"train.py\", line 466, in &lt;module&gt;\n classifier = globals()[args.classifier]\n KeyError: 'Densenet210'\n</code></pre>\n\n<p>any help is appreciated. Resnet50 works fine for me though</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 277378,
          "author_name": "kleinsmith",
          "author_url": "",
          "post_date": "02/02/2018 23:29:57",
          "content": "<p>Change \"Densenet210\" to \"DenseNet201\".</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 277380,
          "author_name": "shivrajp",
          "author_url": "",
          "post_date": "02/02/2018 23:48:04",
          "content": "<p>Apologies for such a typo, but I did try DenseNet201, However issue was with keras version.\nUpdating keras fixed it.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 277389,
          "author_name": "kirk86",
          "author_url": "",
          "post_date": "02/03/2018 00:37:37",
          "content": "<p>@Andres it is keras 2.1.3 and python 3.6 and tf 1.4.1.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 277463,
          "author_name": "antorsae",
          "author_url": "",
          "post_date": "02/03/2018 08:18:14",
          "content": "<p>You need to have the model compiled to access <code>.optimizer</code>, in my latest code:</p>\n\n<p><code>model = load_model(args.model, compile=False if args.test or (args.learning_rate is not None) else True)</code></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 277627,
          "author_name": "kirk86",
          "author_url": "",
          "post_date": "02/03/2018 19:00:08",
          "content": "<p>@Andres awesome thanks a bunch ;)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 277901,
          "author_name": "kirk86",
          "author_url": "",
          "post_date": "02/04/2018 19:16:12",
          "content": "<p>I give up with this fucking keras. No more, I've had enough. Even after setting the learning rate as @Andres suggested here is what I see.\n<img src=\"http://i66.tinypic.com/34phpgm.png\" alt=\"Training\"></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 277999,
          "author_name": "yyll008",
          "author_url": "",
          "post_date": "02/05/2018 03:40:05",
          "content": "<p>python train.py -g 1 -b 8 -cs 229 -cm Densenet201 -l 5e-5 -uiw</p>\n\n<p>Traceback (most recent call last):\n  File \"train.py\", line 467, in \n    classifier = globals()[args.classifier]\nKeyError: 'Densenet201'</p>\n\n<p>Resnet50 works fine, How do you fix this?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 278002,
          "author_name": "yaozhenjie",
          "author_url": "",
          "post_date": "02/05/2018 03:52:57",
          "content": "<p>Your keras is old version, updating it will work</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 278007,
          "author_name": "shivrajp",
          "author_url": "",
          "post_date": "02/05/2018 04:48:30",
          "content": "<p>@YangLu update your keras to keras 2.1.3, previous keras doesn't have DenseNet pre-trained models.</p>\n\n<p><a href=\"https://keras.io/applications/\">Keras Documentation</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 278039,
          "author_name": "yyll008",
          "author_url": "",
          "post_date": "02/05/2018 07:18:32",
          "content": "<p>My keras is 2.1.3, I use serveral conda envs, maybe this disturb.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 278042,
          "author_name": "shivrajp",
          "author_url": "",
          "post_date": "02/05/2018 07:26:41",
          "content": "<p>check if this gets executed in your python console from your present environment:</p>\n\n<pre><code>from keras.applications.densenet import DenseNet201\n</code></pre>\n\n<p>If this works out well, then you have keras 2.1.3 and  are good to go.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 278044,
          "author_name": "ryuichi0704",
          "author_url": "",
          "post_date": "02/05/2018 07:33:34",
          "content": "<p>YangLu, how about to try DenseNet201, not Densenet201?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 278104,
          "author_name": "yyll008",
          "author_url": "",
          "post_date": "02/05/2018 11:37:01",
          "content": "<p>Thanks @RK @shivrajp @Yaozj, It's my low level spelling mistakes. By the way, -cs 512 is too big for me , my one 1070Ti  can  work on \"-cs 229\".</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 278321,
          "author_name": "yaozhenjie",
          "author_url": "",
          "post_date": "02/05/2018 22:31:21",
          "content": "<p>Yes， I have got the same problem， DenseNet needs huge memory.\nMoreover, the loss decrease very slowly, Maybe the learning rate is not set properly</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 278512,
          "author_name": "yyll008",
          "author_url": "",
          "post_date": "02/06/2018 11:18:52",
          "content": "<p>I have another problem wiht extra dataset, Have you fixed it?\n541/953 [================&gt;.............] - ETA: 2:36 - loss: 2.0163 - acc: 0.3073\nException in thread Thread-8:\nFile \"/home/yl/miniconda3/envs/gluon/lib/python3.6/multiprocessing/connection.py\", line 251, in recv\n    return _ForkingPickler.loads(buf.getbuffer())\nTypeError: <strong>init</strong>() missing 1 required positional argument: 'code'</p>\n\n<h1>Add script as follows , fix it.</h1>\n\n<p>load_img  = lambda img_path: np.array(Image.open(img_path))</p>\n\n<p>def load_img_fast_jpg(img_path):\n    try:\n        x = jpeg.JPEG(img_path).decode()\n        return x\n    except:\n        return load_img(img_path)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 275268,
      "author_name": "yaozhenjie",
      "author_url": "",
      "post_date": "01/28/2018 16:08:48",
      "content": "<p>I got errors like this, any suggestion?</p>\n\n<p>Training set distribution:\n              HTC-1-M7:  1014 (13.3%)\n              iPhone-6:   793 (10.4%)\n   Motorola-Droid-Maxx:   767 (10.0%)\n            Motorola-X:   599 (07.8%)\n     Samsung-Galaxy-S4:  1380 (18.1%)\n             iPhone-4s:   740 (09.7%)\n           LG-Nexus-5x:   656 (08.6%)\n      Motorola-Nexus-6:   801 (10.5%)\n  Samsung-Galaxy-Note3:   365 (04.8%)\n            Sony-NEX-7:   519 (06.8%)\nValidation set distribution:\n              HTC-1-M7:    48 (10.0%)\n              iPhone-6:    48 (10.0%)\n   Motorola-Droid-Maxx:    48 (10.0%)\n            Motorola-X:    48 (10.0%)\n     Samsung-Galaxy-S4:    48 (10.0%)\n             iPhone-4s:    48 (10.0%)\n           LG-Nexus-5x:    48 (10.0%)\n      Motorola-Nexus-6:    48 (10.0%)\n  Samsung-Galaxy-Note3:    48 (10.0%)\n            Sony-NEX-7:    48 (10.0%)\nEpoch 1/200\n165/954 [====&gt;.........................] - ETA: 14:01 - loss: 2.2015 - acc: 0.2167Exception in thread Thread-8:\nTraceback (most recent call last):\n  File \"/usr/lib/python2.7/threading.py\", line 801, in <strong>bootstrap_inner\n    self.run()\n  File \"/usr/lib/python2.7/threading.py\", line 754, in run\n    self.__target(*self.__args, **self.__kwargs)\n  File \"/usr/lib/python2.7/multiprocessing/pool.py\", line 389, in _handle_results\n    task = get()\nTypeError: ('__init</strong>() takes exactly 3 arguments (2 given)', , (u'tjDecompressHeader2() failed with error -1 and error string Not a JPEG file: starts with 0x89 0x50',))</p>\n\n<p>175/954 [====&gt;.........................] - ETA: 13:49 - loss: 2.1799 - acc: 0.2279</p>",
      "votes": null,
      "replies": [
        {
          "id": 275275,
          "author_name": "ceperaang",
          "author_url": "",
          "post_date": "01/28/2018 16:18:53",
          "content": "<p>You have broken jpegs in your train. I use code like this to check and remove any offending files:</p>\n\n<pre><code>def check_remove_broken(img_path):\ntry:\n    x = jpeg.JPEG(img_path).decode()\nexcept Exception:\n    print('Decoding error:', img_path)\n    os.remove(img_path)\n\np = Pool(cpu_count() - 2)\np.map(check_remove_broken, tqdm(ids_train))\n</code></pre>",
          "votes": null,
          "replies": []
        },
        {
          "id": 275277,
          "author_name": "yaozhenjie",
          "author_url": "",
          "post_date": "01/28/2018 16:25:00",
          "content": "<p>Thanks a lot for your timely and helpful response, I will try it.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 275285,
          "author_name": "ceperaang",
          "author_url": "",
          "post_date": "01/28/2018 16:43:19",
          "content": "<p>Make sure to reload ids_train after cleaning or your will get \"File not found\" errors instead :)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 275293,
          "author_name": "yaozhenjie",
          "author_url": "",
          "post_date": "01/28/2018 17:33:55",
          "content": "<p>I removed 2 bad images, but got error again, It seems still broken image problem</p>\n\n<p>249/954 [======&gt;.......................] - ETA: 12:20 - loss: 2.1828 - acc: 0.2269['flickr_images/./sony_nex7/37130939252_0b932c54bb_o.jpg', 'flickr_images/./moto_maxx/38349362074_91f99e4ed8_o.jpg', 'flickr_images/./iphone_4s/38769820101_87c2d4fb0c_o.jpg', 'flickr_images/./samsung_s4/37180498790_5f224f9468_o.jpg', '../train/Motorola-X/(MotoX)106.jpg', 'flickr_images/./moto_maxx/38351113384_7295402f17_o.jpg', 'flickr_images/./moto_maxx/38348060164_76a98301fe_o.jpg', 'flickr_images/./nexus_6/36259418946_76c889cca2_o.jpg']\n250/954 [======&gt;.......................] - ETA: 12:18 - loss: 2.1825 - acc: 0.2280['flickr_images/./htc_m7/35835183885_0ee8244504_o.jpg', '../train/Motorola-X/(MotoX)8.jpg', 'flickr_images/./samsung_s4/37983843806_3c0336c1fe_o.jpg', 'flickr_images/./moto_x/23931672383_701bb2cb6b_o_d.jpg', 'flickr_images/./moto_x/31257251621_cfeb71556d_o_d.jpg', '../train/Motorola-Nexus-6/(MotoNex6)103.jpg', 'flickr_images/./samsung_s4/36962967424_6851c6e1d8_o.jpg', 'flickr_images/./iphone_6/27638535039_94c75816b2_o.jpg']\n251/954 [======&gt;.......................] - ETA: 12:17 - loss: 2.1818 - acc: 0.2286['flickr_images/./samsung_s4/26260305859_75196e6006_o.jpg', 'flickr_images/./iphone_6/38718281414_c16867e84d_o.jpg', 'flickr_images/./sony_nex7/23485515358_af58b1be1e_o.jpg', 'flickr_images/./iphone_6/39398306292_ab49153dc1_o.jpg', '../train/Samsung-Galaxy-Note3/(GalaxyN3)114.jpg', '../train/Motorola-X/(MotoX)167.jpg', 'flickr_images/./nexus_6/37396245501_50d30cc8bf_o.jpg', 'flickr_images/./nexus_6/37823866942_4b5e78914f_o.jpg']\n252/954 [======&gt;.......................] - ETA: 12:16 - loss: 2.1814 - acc: 0.2287Traceback (most recent call last):\n  File \"train.py\", line 605, in \n    class_weight=class_weight)\n  File \"/usr/local/lib/python2.7/dist-packages/keras/legacy/interfaces.py\", line 91, in wrapper\n    return func(*args, **kwargs)\n  File \"/usr/local/lib/python2.7/dist-packages/keras/engine/training.py\", line 2145, in fit_generator\n    generator_output = next(output_generator)\n  File \"/usr/local/lib/python2.7/dist-packages/keras/utils/data_utils.py\", line 770, in get\n    six.reraise(value.<strong>class</strong>, value, value.<strong>traceback</strong>)\n  File \"/usr/local/lib/python2.7/dist-packages/keras/utils/data_utils.py\", line 635, in _data_generator_task\n    generator_output = next(self._generator)\n  File \"train.py\", line 371, in gen\n    X[batch_idx], O[batch_idx], y[batch_idx] = batch_result\nValueError: could not broadcast input array from shape (0,512,3) into shape (512,512,3)\nroot@yz114:/data/work/camera_model/sp-society-camera-model-identification# </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 275404,
          "author_name": "yaozhenjie",
          "author_url": "",
          "post_date": "01/29/2018 03:36:34",
          "content": "<p>It seems not data problem, I tried without -x option, still got\nbatch_result ValueError: could not broadcast input array from shape (0,512,3) into shape (512,512,3) </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 275407,
          "author_name": "albertoa",
          "author_url": "",
          "post_date": "01/29/2018 03:47:18",
          "content": "<p>How in the world did you get a 0? That looks like a bad image for sure.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 275449,
          "author_name": "yaozhenjie",
          "author_url": "",
          "post_date": "01/29/2018 06:49:27",
          "content": "<p>It seems not a bad image.\nI printed  the image name, but got the same error from on different images. \nsome times, it was shape (0,512,3) into shape (512,512,3)\nsome times, shape (512,0,3) into shape (512,512,3)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 275453,
          "author_name": "antorsae",
          "author_url": "",
          "post_date": "01/29/2018 06:55:27",
          "content": "<p>Run it with -v to see the shapes of the image. I had that bug a few days ago but was fixed. It was due to wrong offsets in random crops. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 275454,
          "author_name": "yaozhenjie",
          "author_url": "",
          "post_date": "01/29/2018 07:03:49",
          "content": "<p>How to fix it</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 275455,
          "author_name": "antorsae",
          "author_url": "",
          "post_date": "01/29/2018 07:12:45",
          "content": "<p>Are you using the latest code in the repo? It's fixed there.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 275466,
          "author_name": "yaozhenjie",
          "author_url": "",
          "post_date": "01/29/2018 07:52:32",
          "content": "<p>I checked again, it is the latest code.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 275492,
          "author_name": "yaozhenjie",
          "author_url": "",
          "post_date": "01/29/2018 09:35:00",
          "content": "<p>I changed the random crops part as follows, then I can finish a whole epoch now</p>\n\n<pre><code>if random_crop:\n    freedom_x, freedom_y = img.shape[1] - crop_size, img.shape[0] - crop_size\n    if freedom_x &gt; 0:\n        center_x += np.random.randint(math.ceil(-freedom_x/2)+1, freedom_x - math.ceil(freedom_x/2)-1 )\n    if freedom_y &gt; 0:\n        center_y += np.random.randint(math.ceil(-freedom_y/2)+1, freedom_y - math.ceil(freedom_y/2)-1 )\n</code></pre>",
          "votes": null,
          "replies": []
        },
        {
          "id": 275509,
          "author_name": "antorsae",
          "author_url": "",
          "post_date": "01/29/2018 11:07:43",
          "content": "<p>Yeah, better safe than sorry and 2 pixels is no biggie. I wonder why Im not experiencing it, b/c I had the same issue and applied <code>math.ceil</code> and <code>math.floor</code> very carefully to avoid edge conditions.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 275325,
      "author_name": "omsuchak",
      "author_url": "",
      "post_date": "01/28/2018 20:25:45",
      "content": "<p>Nice work... Thats impressive!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 275858,
      "author_name": "sabbiracoustic1006",
      "author_url": "",
      "post_date": "01/30/2018 05:46:05",
      "content": "<p>Hey Andres Torrubia, Is the high pass filter that you used from a slide working? I am a Student,trying several approaches with severe failures. Any suggestions to improve my score would be appreciated</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 276112,
      "author_name": "rajeshbhat",
      "author_url": "",
      "post_date": "01/30/2018 18:15:37",
      "content": "<p>I am getting below error:\nTraceback (most recent call last):\n  File \"train.py\", line 591, in \n    class_weight=class_weight)\n  File \"/home/rbhat/anaconda2/envs/mob/lib/python3.5/site-packages/keras/legacy/interfaces.py\", line 91, in wrapper\n    return func(*args, **kwargs)\n  File \"/home/rbhat/anaconda2/envs/mob/lib/python3.5/site-packages/keras/engine/training.py\", line 2145, in fit_generator\n    generator_output = next(output_generator)\n  File \"/home/rbhat/anaconda2/envs/mob/lib/python3.5/site-packages/keras/utils/data_utils.py\", line 770, in get\n    six.reraise(value.<strong>class</strong>, value, value.<strong>traceback</strong>)\n  File \"/home/rbhat/.local/lib/python3.5/site-packages/six.py\", line 693, in reraise\n    raise value\n  File \"/home/rbhat/anaconda2/envs/mob/lib/python3.5/site-packages/keras/utils/data_utils.py\", line 635, in _data_generator_task\n    generator_output = next(self._generator)\n  File \"train.py\", line 356, in gen\n    batch_results = p.map(process_item_func, item_batch)\n  File \"/home/rbhat/anaconda2/envs/mob/lib/python3.5/multiprocessing/pool.py\", line 266, in map\n    return self._map_async(func, iterable, mapstar, chunksize).get()\n  File \"/home/rbhat/anaconda2/envs/mob/lib/python3.5/multiprocessing/pool.py\", line 644, in get\n    raise self._value\nOSError: Could not load libjpeg-turbo library</p>\n\n<p>Does anyone faced the same issue?</p>",
      "votes": null,
      "replies": [
        {
          "id": 276128,
          "author_name": "igorkrashenyi",
          "author_url": "",
          "post_date": "01/30/2018 19:00:07",
          "content": "<p>You need to install libjpeg-turbo </p>\n\n<pre><code>sudo apt install libturbojpeg\n</code></pre>",
          "votes": null,
          "replies": []
        },
        {
          "id": 276241,
          "author_name": "yaozhenjie",
          "author_url": "",
          "post_date": "01/31/2018 02:48:55",
          "content": "<p>add the lib64 directory to LD_LIBRARY_PATH</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 277053,
      "author_name": "humbledumble",
      "author_url": "",
      "post_date": "02/02/2018 03:44:07",
      "content": "<p>Great approach! Way to go.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 278347,
      "author_name": "yyll008",
      "author_url": "",
      "post_date": "02/06/2018 00:48:11",
      "content": "<p>Awesome work! Thanks!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 278883,
      "author_name": "bopengiowa",
      "author_url": "",
      "post_date": "02/07/2018 02:15:26",
      "content": "<p>Hi guys,</p>\n\n<p>When I use the preprocessing kernel filters # <a href=\"http://www.lirmm.fr/~chaumont/publications/WIFS-2016_TUAMA_COMBY_CHAUMONT_Camera_Model_Identification_With_CNN_slides.pdf\">http://www.lirmm.fr/~chaumont/publications/WIFS-2016_TUAMA_COMBY_CHAUMONT_Camera_Model_Identification_With_CNN_slides.pdf</a></p>\n\n<pre><code> kernel_filter = 1/12. * np.array([\\\n        [-1,  2,  -2,  2, -1],  \\\n        [ 2, -6,   8, -6,  2],  \\\n        [-2,  8, -12,  8, -2],  \\\n        [ 2, -6,   8, -6,  2],  \\\n        [-1,  2,  -2,  2, -1]]) \n</code></pre>\n\n<p>The images after the filter are basically completely black with speckles of white. Is this supposed to be the case? </p>\n\n<p>Thanks for the help!</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "274067": "Got 0.959 LB with a single Densenet201 model trained in less than 1 day (my code is available on github, see other thread).\n\nSo far implementation has no real novel approaches: just no hidden bugs in the code (I hope), use pre-trained model and Gleb's dataset; along with basic train, validation and test augmentation.\n\nFor me the challenge is how to get to 0.97+ with a single model. Some ideas Im going to try:\n\n - Loss/accuracy function that mimics the organization (right now I'm using global accuracy) \n - Class-aware sampling for training \n - Mixup\n - Balanced validation set (assuming validation is equally balanced as\n   training)\n - More data\n - New NN architecture/tuning\n - Build confusion matrix to diagnose issues\n\nIf you have any other ideas you're willing to share, feel free to comment.",
    "274085": "You're just killing it, man!",
    "274169": "Very nice Andres, thanks for sharing your approach! \n\nCouple (minor) optimization suggestions - \n\n1. Multiprocessing - Have a look at the consumer-producer(s) design pattern. Invoking pool-map on every batch yield isn't particularly efficient. I suspect that your GPUs are the bottleneck which is masking this issue. As an aside, GIL doesn't mean you should always default to processes over threads. Most functions in C-native Python libraries such as numpy release the GIL. The book \"Fluent Python\" is a great read to learn more about this and other lower level language structures.\n\n2. Increasing your GPU throughput - You've mentioned gradient check-pointing but also have a look at replacing the TF backend with MXNet. Note that this would require using a forked version of Keras.\n\n3. Downclocking your GPUs - this might be controversial but consider downclocking your GPUs if you have a hobbyist home setup (use nvidia-smi to set persistence mode with the PM flag, and then using the PL flag to set the target TDP). If you're running your GPUs 24/7 like most of us are, it reduces thermal wear in the long-term for little-to-no performance decrease. I personally run my 1080 Tis at 200W (down from the default 270W) and have measured a &lt; 3% decrease.",
    "274203": "That's awesome",
    "274287": "Thanks for the suggestions. Just bought the \"Fluent Python\" book. :-) Will also try your other suggestions, Im going to check CNTK too. \n\nI believe using a MXNet would preclude me from using imagenet pretrained weights for latest networks (NasNet, DenseNet, etc.). Do you know if it is the case?",
    "274306": "That's my understanding but I'm happy to be corrected. My team-mate and I have stuck to vanilla Keras because of that. So there's a trade-off between flexibility and performance. \n\nIn any case, we're not using any of the cutting edge models (e.g., DenseNet, NASNet) for this competition.",
    "274308": "Im curious: are you using an ensemble of models or a single model for your score?",
    "274321": "We're ensembling but by default rather than having benchmarked it against single models.  I doubt this is a major factor. \n\nHaving observed the trajectory of the top teams' scores for the past month, I suspect we are sniffing around the same local optima and the spread in performance is largely due to hyper-parameters/randomness.\n\nHopefully someone will come in with a novel solution and blow the LB up!",
    "274329": "Thanks for the power limit tip.\n\n&gt; set persistence mode with the PM flag\n\n[NVIDIA said][1] they'll eventually stop supporting persistence mode, and are focusing on the persistence daemon. [Getting the daemon to start at startup][2] is a little weird, though. On Ubuntu, I had to unpack /usr/share/doc/NVIDIA_GLX-1.0/sample/nvidia-persistenced-init.tar.bz2 and then run the install script. [One person says][3] to add \"--persistence-mode\" to the service start line in the template file, but that seems unneeded.\n\n  [1]: http://docs.nvidia.com/deploy/driver-persistence/index.html#persistence-daemon\n  [2]: http://docs.nvidia.com/deploy/driver-persistence/index.html#installation\n  [3]: https://devtalk.nvidia.com/default/topic/995248/cuda-setup-and-installation/setting-up-nvidia-persistenced/post/5090647/#5090647",
    "274386": "Hey everyone for me the code is not working. It's stuck like this for almost a day now:\n\n           HTC-1-M7:  1023 (13.0%)\n           iPhone-6:   823 (10.5%)\n           Motorola-Droid-Maxx:   825 (10.5%)\n           Motorola-X:   275 (03.5%)\n           Samsung-Galaxy-S4:  1412 (18.0%)\n           iPhone-4s:   774 (09.9%)\n           LG-Nexus-5x:   680 (08.7%)\n           Motorola-Nexus-6:   926 (11.8%)\n           Samsung-Galaxy-Note3:   548 (07.0%)\n           Sony-NEX-7:   557 (07.1%)\n           validation steps = 0\n           Epoch 1/200\nAny suggestions?",
    "274388": "I assume you're running it with `-x`, i.e. using Gleb's dataset. \nDid you sync the `val_images` directory @ https://github.com/antorsae/sp-society-camera-model-identification/tree/master/val_images and downloaded all images?\n\nWhen using `-x` I wanted to have the validation as distinct as possible to the train set, so:\n\n        ids_train = ids\n        ids_val   = [ ]\n\n        extra_train_ids = [os.path.join(EXTRA_TRAIN_FOLDER,line.rstrip('\\n')) for line in open(os.path.join(EXTRA_TRAIN_FOLDER, 'good_jpgs'))]\n        extra_train_ids.sort()\n        ids_train.extend(extra_train_ids)\n\n        extra_val_ids = glob.glob(join(EXTRA_VAL_FOLDER,'*/*.jpg'))\n        extra_val_ids.sort()\n        ids_val.extend(extra_val_ids)\n\ndo a `print(ids_val)` after that line and LMK what you get.",
    "274392": "Hi Andres, thanks for the quick response. I haven't downloaded all the images, I that was happening automatically. Should I'll be also creating any additional directories for the images? On another note I am getting the value of `validation_steps=0` in the fit.generator so as a consequence I had to hardcode the value to 6 in order for the code not  to throw an error. Let me try download all the images and run again give you back the output of `print(ids_val)`.",
    "274434": "Download Gleb's training set by running `find . -name \"urls_*\" -execdir wget -nc --tries=10 -i {} \\;` in the `flickr_images` directory and adjust accordingly in the `val_images` one.",
    "274437": "I got 0.946 on the public LB using Andres code with \n`python train.py -g 1 -b 16 -cs 512 -cm ResNet50 -x -l 1e-4 -uiw`\ntraining for 91 epochs on a single V100. Training took about 10 hours. Thanks Andres, I have learned a lot from looking at your code.",
    "274449": "Also don't forget to rename any .JPG to .jpg as mentioned on the other thread.",
    "274491": "Thank you Andres &amp; Alberto, we are back in business fellas :).",
    "274548": "Good to hear. There's a bug in the code re: augmentation. LMK if you can spot it :-) I will submit fixed version with other improvements this weekend.",
    "274558": "is bug in the train part?",
    "274561": "Yes, may or may not decrease accuracy (I guess it does), but I want to retrain and see tomorrow. Will submit then.",
    "274578": "I noticed something in the gamma augmentation, the formula would be pow(x, 1/gamma) but it's written as pow(x, gamma); I think it doesn't matter much as 1/0.8==1.25 and 1/1.2==0.83, so in the end gamma08 generates gamma 1.2 and vice-versa :-)",
    "274581": "The bug I was referring to is not related to gamma. Re: gamma by looking at http://scikit-image.org/docs/dev/api/skimage.exposure.html#skimage.exposure.adjust_gamma the formula is O = I**gamma so I think the current formula is OK; why should it be 1/gamma?",
    "274585": "There is some confusion about which formula is actually \"gamma correction\":\nhttps://stackoverflow.com/questions/16521003/gamma-correction-formula-gamma-or-1-gamma",
    "274616": "I am getting this error :\n\n$ python train.py -g 1 -b 16 -cs 512 -cm ResNet50 -l 1e-4 -uiw\n\n2018-01-26 23:01:46.708921: E tensorflow/stream_executor/cuda/cuda_dnn.cc:385] could not create cudnn handle: CUDNN_STATUS_INTERNAL_ERROR\n2018-01-26 23:01:46.708954: E tensorflow/stream_executor/cuda/cuda_dnn.cc:352] could not destroy cudnn handle: CUDNN_STATUS_BAD_PARAM\n2018-01-26 23:01:46.708961: F tensorflow/core/kernels/conv_ops.cc:667] Check failed: stream-&gt;parent()-&gt;GetConvolveAlgorithms( conv_parameters.ShouldIncludeWinogradNonfusedAlgo\n\nEverything works fine If I run Resnet50 from keras examples.\n\nAnyone has an idea what might be the reason ?",
    "274626": "Maybe it is a versioning issue. What about downgrading keras?",
    "274628": "I use the same keras version and tf as OP \nWhat version are you using ?",
    "274631": "keras 2.1.2, tensorflow 1.4.0 and it works well. I also tried to upgrade keras to 2.1.3 and got the error you mentioned above.",
    "274635": "I use Keras 2.1.3 and TF 1.4.1",
    "274638": "Check your CUDA &amp; cudnn versions. Default Tensorflow from pip compiled with cuda 9.0 and cudnn 7, so you need exactly them.",
    "274649": "Another possible reason: Anaconda. It messes with cublas libraries and tensorflow going crazy. Solution: probably, simplest way -- just use system python for this code.",
    "274709": "First of all, really appreciate sharing the code. It really, really helps.\n\nI managed to run it on one of my machines, but now when I am trying to run the latest version (downloaded today) on another machine, I am getting really weird number of images in different classes. One of them (Moto X) even has 0 images. However, when I check the folders, I see that all of the images are there. Anyone else have the same issue? Any idea how to fix it?\n\n```              HTC-1-M7:   748 (13.9%) \n              iPhone-6:   548 (10.2%)\n   Motorola-Droid-Maxx:   550 (10.2%)\n            Motorola-X:     0 (00.0%)\n     Samsung-Galaxy-S4:  1137 (21.2%)\n             iPhone-4s:   499 (09.3%)\n           LG-Nexus-5x:   405 (07.5%)\n      Motorola-Nexus-6:   651 (12.1%)\n  Samsung-Galaxy-Note3:   273 (05.1%)\n            Sony-NEX-7:   557 (10.4%)\n```",
    "274761": "Hi,  \nI am using Keras, with CUDA_VISIBLE_DEVICES to mask to use 2 GTX1080 GPU,\nbut when I run DenseNet201 model, will another two fc layers,\nIt give me a lot of errors : \n\n    ran out of memory trying to allocate XXXX MiB\n\nI tried to use 4 GTX1080 GPU,  but still get same errors.\nIt seems that it GPU memory doesn't increase, it performance is same as single GTX1080\nCan you give me some suggestions? What should I do to implement this model to multi-gpu?",
    "274773": "OK, getting 96.2 LB now single model.\nCode is updated in the repo:\n\n - Added GPL3 license: basically if you modify it, share the code.\n - Fixed orientation flip augmentation bug (need to confirm current one is OK.\n - Control TTA and prints class distribution after `-t`.",
    "274782": "I created a fresh env of py3.5and it worked, before I was using py3.6.\n\nThanks guys !",
    "274821": "do you do `-g 2` for 2 GPUs?",
    "274860": "Can you share how you created the environment. I still haven't been able to run Andres code anymore. I'm actually creating a new code base and deconstructing it part by part.",
    "274864": "i use anaconda : \nhttp://uoa-eresearch.github.io/eresearch-cookbook/recipe/2014/11/20/conda/",
    "274868": "On linux I use [pipenv][1]\n\n\n  [1]: https://github.com/pypa/pipenv",
    "274871": "Yes, but that doesn't give me versions/packages.\n\nI have several virtenvs but something happened trying to upgrade tensorflow and I haven't been able to run Andres code since.",
    "274872": "Alberto, this may be an overkill but here's my base environment to use with anaconda (I use miniconda, but regular anaconda would work too). There's many packages that are not needed for this project so feel free to edit the file by hand and remove the ones you think are not needed.\n\nhttps://conda.io/docs/user-guide/tasks/manage-environments.html",
    "274879": "I'm using Docker for Andres's code:\n\n    alias andres=\"docker run --runtime=nvidia --init -it --rm \\\n                  --ipc=host \\\n                  -v $CODE:/code -v $DATA:/data \\\n                  -w=/code/sp-society-camera-model-identification \\\n                  mwksmith/cam:andres\"\n\nWhen you get into the container enter \"sa\" to activate the conda environment, and then create sym links to your data folders according to the globals set in train.py. Then run train.py.\n\nI'm pushing the Docker Image to Docker Hub now. You can also build it yourself with [the Dockerfile][1].\n\n\n  [1]: https://github.com/antorsae/sp-society-camera-model-identification/blob/master/Dockerfile\n\nEdit: Removed `--shm-size`. It appears to be redundant with respect to `--ipc=host`.",
    "274913": "&gt; **Andres Torrubia wrote**\n&gt; \n&gt; &gt; Download Gleb's training set by running `find . -name \"urls_*\" -execdir wget -nc --tries=10 -i {} \\;` in the `flickr_images` directory and adjust accordingly in the `val_images` one.\n\nHere's the breakdown of this command.\n\n- `find`: returns file paths that match a given pattern\n- `.`: search working directory recursively (search its subdirectories, their subdirectories, and so on)\n-  `-name \"urls_*\"`: the given file path pattern\n- `-execdir`: execute the given command on each file path found, **in the directory of the file.**\n- `wget`: download what a given URL points to\n- `-nc`: don't download files that already exist\n- `-i`: read URLs from a given file\n- `{}`: placeholder for each file that `find` found\n- `\\;`: end command",
    "274916": "One caveat when doing this on `val_images`. In the `val_images/moto_maxx` directory a bunch of images will get downloaed with *.1 *.2 *.3 ... extensions. Rename them to _1.jpg _2.jpg _3.jpg otherwise the code will not use them and if using `-x` validation set for that class will be heavily underrepresented.",
    "274917": "Thanks, it may end up being what I need.\n\nI finally finished going through every line and did some re-structuring of the code from a couple of days ago. I made it a little more robust in handling errors. Unfortunately I also got rid of the paralleled implementation of the generator so it is now slower. So far it started to run, so I may finally be in the right track.\n\nFor anyone interested my fork is at: https://github.com/albertoa/sp-society-camera-model-identification\nI would only recommend it at this stage for the added comments. I will improve the generator and look at your latest changes to see what else would be useful. For now I'm just happy I'm back training networks.",
    "274918": "Thanks, this may be the fastest approach for me to get his latest release working. I'll try it as soon as the current network training is done.",
    "274940": "Got 0.969 LB w/ single model. \nI think there's a chance to get to 0.98 with a single model.",
    "274945": "With the same code?",
    "274952": "Almost. Diff is 20 lines of code. Will submit after more changes I have planned.",
    "274965": "Chun, thanks for the note on power. I'm curious though, why set the power limit instead of the temperature thresholds: slow_threshold and max_threshold?",
    "274970": "you better keep some of that code for yourself if you hope to be in the money :)",
    "274971": "I know. This all adds to the excitement ;-)",
    "274982": "I bet top teams are already very busy training *a lot* of diverse models (many folds probably) and will do it up to the last minute, so we could see a lot of movements near the end :)",
    "274998": "&gt; **Sergey Mushinskiy wrote**\n&gt; \n&gt; &gt; I bet top teams are already very busy training *a lot* of diverse models (many folds probably)\n\nHi Sergey, what do you mean by \"fold\"?",
    "275023": "Thanks Andres for sharing your solution. I was running your code. The code run for 44 epochs and stopped spitting the following error. Do you have any ideas how to debug this error?\n\nread Thread-4:\nTraceback (most recent call last):\n  File \"/home/wxk/anaconda2/envs/py3/lib/python3.6/threading.py\", line 916, in _bootstrap_inner\n    self.run()\n  File \"/home/wxk/anaconda2/envs/py3/lib/python3.6/threading.py\", line 864, in run\n    self._target(*self._args, **self._kwargs)\n  File \"/home/wxk/anaconda2/envs/py3/lib/python3.6/multiprocessing/pool.py\", line 429, in _handle_results\n    task = get()\n  File \"/home/wxk/anaconda2/envs/py3/lib/python3.6/multiprocessing/connection.py\", line 251, in recv\n    return _ForkingPickler.loads(buf.getbuffer())\nTypeError: __init__() missing 1 required positional argument: 'code'",
    "275024": "I have seen that before. To me it was happening on epoch 1. There seems to be a problem with the multiprocess code that triggers it, but other than going into a single CPU for the generator I couldn't solve it.",
    "275028": "Thanks Alberto for your reply. Then how to set the code to one CPU? \n\nBTW, do you know how the follow code works? I was trying to  figure out how the models (Resnet50, Densenet201 and so on) are created.\ngetattr(globals()[classifier_module_name], 'preprocess_input')",
    "275033": "I was afraid you were going to ask. I have a fork of the github code, but I found a major issue with it that I will fix tomorrow, gives wrong results ;-(.\nAdres is working on a new release, hopefully his has the fix.\nThe lines in question should be:\n    p = Pool(cpu_count()-2)\n            batch_results = p.map(process_item_func, item_batch)\nand how the batches are put together. But that's where I messed up my fork, so I rather not recommend anything until I know I can in fact run it properly.",
    "275043": "Thanks Alberto for your kind reply.",
    "275048": "By the way I am having the same issue with the docker image from Matt Kleinsmith. \n\n[...]\n      File \"/opt/conda/envs/tf/lib/python3.5/multiprocessing/connection.py\", line 251, in recv\n        return ForkingPickler.loads(buf.getbuffer())\n    TypeError: __init__() missing 1 required positional argument: 'code'\n\nI can not tell you why you and I get the error and yet so many others don't. We are the wrong kind of lucky I guess.",
    "275049": "Thanks Alberto. \n\nDo you know how the networks become global variables in this line: getattr(globals()[classifier_module_name], 'preprocess_input')?",
    "275051": "Thanks a lot for sharing your code! Really helps students like me learn a lot! :)",
    "275054": "The key to that is:\n\nfrom keras.applications import *\n\nHe is using the classifier_to_module dictionary so that the correct preprocess_input function is called for whichever model is being used.",
    "275058": "I see. Thanks Alberto. I was stilling creating the models myself. Keras is becoming more handy.",
    "275124": "An exception happened in the mutiprocess code and Python does not tell you the exception (very likely a file was not read correctly) so it tanks with an obscure error. Making the code run in 1 thread only is painfully slow unless you resort to some sort of caching.",
    "275125": "I mean folds like in k-fold cross-validation. Split dataset strategically into, say, 5 parts and train 5 model on 4 of them (different each time) and use 1 as validation. It is very common strategy in general data science but obviously somewhat rare in deep learning (given computational resources needed). However, it allows to build honest second layer model on top of networks predictions.",
    "275246": "Hi everyone I just wanted to update on some things that I've noticed while using @Andres code which he kindly provided to us. @Andres once again thanks for sharing the code with us. In my system I haven't noticed a big difference in regards to gaining speed training from the multiprocessing generator. Incidentally I am using python3 and running the code without the multiprocessing, I see that multiple cores are still utilized with the difference in regards to the multiprocessor generator, that they are not  &gt;90% occupied all the time, which in some cases is not ideal to have all the cores fully loaded, especially in a shared system. No matter what you set the number of processes in the code still all cores are fully occupied.  That being said, there are some alternatives which I haven't tested yet but I thought I should share them here in case they are useful or maybe someone has already tried them. The first alternative is to use sth like [tensorpack][1]. The other one is to use sth like `producer-consumer` pattern that @Chun Ming Lee already mentioned. With all the deadlines I didn't have the time to incorporate that into @Andres code but there's a simple example in the `consumer-producer.py` file which can be easily extended into @Andres code if anyone wants to.  What does though really make a difference regarding training speed is doing multi-gpu training, almost reduces the training time in half plus you might be able to use larger batch_size. That's all for now folks, I'll update later on different models. Cheers!\n\n\n  [1]: http://tensorpack.readthedocs.io/en/latest/",
    "275247": "I have a question @Chun Ming Lee regarding point 2. Does is make any difference which backend you use? Ultimately now all of them rely on the same lower layer which is cuda and cudnn? Which IMHO it won't make any big difference but then again I might be wrong. Thanks!",
    "275268": "I got errors like this, any suggestion?\n\n\nTraining set distribution:\n              HTC-1-M7:  1014 (13.3%)\n              iPhone-6:   793 (10.4%)\n   Motorola-Droid-Maxx:   767 (10.0%)\n            Motorola-X:   599 (07.8%)\n     Samsung-Galaxy-S4:  1380 (18.1%)\n             iPhone-4s:   740 (09.7%)\n           LG-Nexus-5x:   656 (08.6%)\n      Motorola-Nexus-6:   801 (10.5%)\n  Samsung-Galaxy-Note3:   365 (04.8%)\n            Sony-NEX-7:   519 (06.8%)\nValidation set distribution:\n              HTC-1-M7:    48 (10.0%)\n              iPhone-6:    48 (10.0%)\n   Motorola-Droid-Maxx:    48 (10.0%)\n            Motorola-X:    48 (10.0%)\n     Samsung-Galaxy-S4:    48 (10.0%)\n             iPhone-4s:    48 (10.0%)\n           LG-Nexus-5x:    48 (10.0%)\n      Motorola-Nexus-6:    48 (10.0%)\n  Samsung-Galaxy-Note3:    48 (10.0%)\n            Sony-NEX-7:    48 (10.0%)\nEpoch 1/200\n165/954 [====&gt;.........................] - ETA: 14:01 - loss: 2.2015 - acc: 0.2167Exception in thread Thread-8:\nTraceback (most recent call last):\n  File \"/usr/lib/python2.7/threading.py\", line 801, in __bootstrap_inner\n    self.run()\n  File \"/usr/lib/python2.7/threading.py\", line 754, in run\n    self.__target(*self.__args, **self.__kwargs)\n  File \"/usr/lib/python2.7/multiprocessing/pool.py\", line 389, in _handle_results\n    task = get()\nTypeError: ('__init__() takes exactly 3 arguments (2 given)',",
    "275275": "You have broken jpegs in your train. I use code like this to check and remove any offending files:\n\n    def check_remove_broken(img_path):\n    try:\n        x = jpeg.JPEG(img_path).decode()\n    except Exception:\n        print('Decoding error:', img_path)\n        os.remove(img_path)\n\n    p = Pool(cpu_count() - 2)\n    p.map(check_remove_broken, tqdm(ids_train))",
    "275277": "Thanks a lot for your timely and helpful response, I will try it.",
    "275285": "Make sure to reload ids_train after cleaning or your will get \"File not found\" errors instead :)",
    "275293": "I removed 2 bad images, but got error again, It seems still broken image problem\n\n249/954 [======&gt;.......................] - ETA: 12:20 - loss: 2.1828 - acc: 0.2269['flickr_images/./sony_nex7/37130939252_0b932c54bb_o.jpg', 'flickr_images/./moto_maxx/38349362074_91f99e4ed8_o.jpg', 'flickr_images/./iphone_4s/38769820101_87c2d4fb0c_o.jpg', 'flickr_images/./samsung_s4/37180498790_5f224f9468_o.jpg', '../train/Motorola-X/(MotoX)106.jpg', 'flickr_images/./moto_maxx/38351113384_7295402f17_o.jpg', 'flickr_images/./moto_maxx/38348060164_76a98301fe_o.jpg', 'flickr_images/./nexus_6/36259418946_76c889cca2_o.jpg']\n250/954 [======&gt;.......................] - ETA: 12:18 - loss: 2.1825 - acc: 0.2280['flickr_images/./htc_m7/35835183885_0ee8244504_o.jpg', '../train/Motorola-X/(MotoX)8.jpg', 'flickr_images/./samsung_s4/37983843806_3c0336c1fe_o.jpg', 'flickr_images/./moto_x/23931672383_701bb2cb6b_o_d.jpg', 'flickr_images/./moto_x/31257251621_cfeb71556d_o_d.jpg', '../train/Motorola-Nexus-6/(MotoNex6)103.jpg', 'flickr_images/./samsung_s4/36962967424_6851c6e1d8_o.jpg', 'flickr_images/./iphone_6/27638535039_94c75816b2_o.jpg']\n251/954 [======&gt;.......................] - ETA: 12:17 - loss: 2.1818 - acc: 0.2286['flickr_images/./samsung_s4/26260305859_75196e6006_o.jpg', 'flickr_images/./iphone_6/38718281414_c16867e84d_o.jpg', 'flickr_images/./sony_nex7/23485515358_af58b1be1e_o.jpg', 'flickr_images/./iphone_6/39398306292_ab49153dc1_o.jpg', '../train/Samsung-Galaxy-Note3/(GalaxyN3)114.jpg', '../train/Motorola-X/(MotoX)167.jpg', 'flickr_images/./nexus_6/37396245501_50d30cc8bf_o.jpg', 'flickr_images/./nexus_6/37823866942_4b5e78914f_o.jpg']\n252/954 [======&gt;.......................] - ETA: 12:16 - loss: 2.1814 - acc: 0.2287Traceback (most recent call last):\n  File \"train.py\", line 605, in",
    "275325": "Nice work... Thats impressive!",
    "275329": "The recv return _ForkingPickler.loads(buf.getbuffer()) TypeError: init() missing 1...  error seems to be related to either non jpg files or files missing/corrupted.  I replaced:\n\n    img = load_img_fast_jpg(item)\n\nwith\n    try:\n        img = jpeg.JPEG(item).decode()\n    except:\n        img = np.array(Image.open(item))\n\n(I will add another try block in my code for the Image.open as well, I just didn't get around to doing it)\n\nFurthermore I made sure the directories AND files in Flicker do match Andres files. I particularly had to look at his: flickr_images/good_jpgs and flickr_images/low-quality files.\n\nAfter downloading some of the images I didn't have, removing some entries I was able to get his code to run without seeing the particular error. The run was using Matt Kleinsmith's docker container, since I still didn't trust my environment to have the right versions of anything anymore.\n\nTo further improve the code I think the proper solution would be to change the return None within process_item and change it so that it can return from the \"child\" processes properly without causing issues in Python's pool management and then at the \"parent\" handle error conditions prior to: for batch_result in batch_results: \n\nXiaokangWang if you get a chance make the changes and let us know if that solves it for you as well.\n\nAnyway, I hope this helps anybody that is having reliability issues. I can't believe it took me so long to track the thing down.\n\nAndres, si algun dia estoy en Valencia para la fallas hazme un favor, pasate y me empujas a la hoguera ;-)",
    "275339": "If I have the time I will try the consumer/producer model. I'm curious to see whether is more efficient than `Pool`. I started reading the book Fluent Python as suggested by Chun Ming Lee.\n\nRe: file mismatch errors hitting you in the face (and you didn't know where the punch came from) yes - the code/error handling is a bit messy, but hey, once you fix all files it works (you need to triple check the files).\n\nAlberto, soy de Alicante... pero igualmente te puedo empujar a la hoguera, de hecho aquí las llamamos Las Hogueras...  :-)",
    "275347": "Thanks Andres for your reply. There are some image that can not be read.",
    "275352": "Hi Alberto, I will give it a try. I was thinking why we just use Image.open rather than both jpeg and Image.open",
    "275374": "That should also work, as it is a more generalized library. There is a performance hit, although I haven't timed it to see its significance.",
    "275386": "Andres, I wouldn't spend too much time on optimizing your MP code. Parallel code is extremely bug-prone and the only reason I have a decent working implementation is that I spent half of a previous competition (CDiscount) working outs bugs. \n\nAnd with &lt;=2 GPUs, your CPUs won't be the bottleneck.",
    "275387": "At the margins we're dealing with, they do matter. \n\nTo give you an example, until a couple months back, Keras' implementation of RNNs (LSTM &amp; GRU) were substantially slower than bare-metal TF implementations. \n\nAnd this extends to stuff like pre-trained model weights leading to materially different results depending on whether you use TF + TF-weights or Theano + Theano weights.",
    "275404": "It seems not data problem, I tried without -x option, still got\nbatch_result ValueError: could not broadcast input array from shape (0,512,3) into shape (512,512,3)",
    "275407": "How in the world did you get a 0? That looks like a bad image for sure.",
    "275449": "It seems not a bad image.\nI printed  the image name, but got the same error from on different images. \nsome times, it was shape (0,512,3) into shape (512,512,3)\nsome times, shape (512,0,3) into shape (512,512,3)",
    "275453": "Run it with -v to see the shapes of the image. I had that bug a few days ago but was fixed. It was due to wrong offsets in random crops.",
    "275454": "How to fix it",
    "275455": "Are you using the latest code in the repo? It's fixed there.",
    "275466": "I checked again, it is the latest code.",
    "275492": "I changed the random crops part as follows, then I can finish a whole epoch now\n\n    if random_crop:\n        freedom_x, freedom_y = img.shape[1] - crop_size, img.shape[0] - crop_size\n        if freedom_x &gt; 0:\n            center_x += np.random.randint(math.ceil(-freedom_x/2)+1, freedom_x - math.ceil(freedom_x/2)-1 )\n        if freedom_y &gt; 0:\n            center_y += np.random.randint(math.ceil(-freedom_y/2)+1, freedom_y - math.ceil(freedom_y/2)-1 )",
    "275509": "Yeah, better safe than sorry and 2 pixels is no biggie. I wonder why Im not experiencing it, b/c I had the same issue and applied `math.ceil` and `math.floor` very carefully to avoid edge conditions.",
    "275521": "Chun Ming Lee, thanks for your reply!",
    "275524": "Some updates folks. My first observation is that vggish type of networks are not a good option for the task at hand. Slow training times and low accuracy. Batch size has quite an impact on accuracy, that goes for all type of networks.",
    "275858": "Hey Andres Torrubia, Is the high pass filter that you used from a slide working? I am a Student,trying several approaches with severe failures. Any suggestions to improve my score would be appreciated",
    "275947": "Thanks for your valuable inputs. Could you help me with why batch size is playing a significant role, I would assume it has to do with how much data we can fit in the ram for training and I would expect with a larger batch size the convergence to happen faster compared to lower batch size assuming other parameters to be same. It would be helpful if you can lead me to why there is a decrease/increase in accuracy based on batch size.\n\nWhat batch size did you explore for this particular data set ? and what worked best for you?",
    "275949": "Batch size interacts with learning rate, as a rule of thumb, if you increase batch size performs a similar role as decreasing learning rate, and learning rate is probably the single most important hyper-parameter to fine tune to find fast convergence. See also [Don't Decay the Learning Rate, Increase the Batch Size][1]\n\nAlso, classifiers/feature extractors use Batch Normalization which is very depending on batch size. See [Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift][2]. As a fun aside, what is covariate shift?\n\nI'm using Keras so I have to find a good LR by hand. Fast.ai just released their Pytorch-based framework that supports the Learning Rate Finder algorithm [fast.ai v2][3]\n\n\n  [1]: https://arxiv.org/abs/1711.00489 \"Don't Decay the Learning Rate, Increase the Batch Size\"\n  [2]: https://arxiv.org/abs/1502.03167 \"Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift\"\n  [3]: http://www.fast.ai/2018/01/26/v2-launch/",
    "275964": "The way I picture the batch size trade-off is -\n\n 1. Smaller batches: noisier gradient updates but potentially gets you out of sharp local minima\n 2. Larger batches: cleaner gradient updates potentially converging faster, but at the risk of getting stuck in a local minima.\n\nAs Andres mentioned, there's a fair bit of ongoing research about this with two general schools of thought - the \"small batch size is better\" camp, and the Facebook etc. camp which has produced research claiming they can scale up to BS of thousands by playing around with LR and other settings. \n\nWhat I've seen in the few Kaggle competitions I've participated in is - it depends. You're going to have test out various combinations on the problem you're working on.\n\nThe choice of optimizer is pretty important as well. As a newbie, I defaulted to Adam but there's literature suggesting that adaptive optimizers (e.g., Adam, RMSProp etc.) generally perform worse or at best equal to vanilla SGD algorithms. ([\"The Marginal Value of Adaptive Gradient Methods in Machine Learning\"][1])\n\n  [1]: https://arxiv.org/pdf/1705.08292.pdf",
    "275973": "Andres Torrubia Thanks for your contribution so far, has been a great learning experience for me. I read the article you suggested \"Don't Decay the Learning Rate, Increase the Batch Size\", it was an interesting read, however, it was mostly around how It reaches equivalent test accuracies after the same number of training epochs leading to greater parallelism and shorter training times, whereas I was inferring from @kirk's post that it was affecting the accuracies on validation(assuming time is not a constraint here). \n\n@Chung Ming Lee Thanks for your explanation, based on your points one can infer that small batches can lead to better results at the cost of  greater time complexity as they have lesser likelihood of getting stuck in local minima, I am not sure should I take this inference as conclusion because you mentioned there are researchers in this area working to better validate it.\nI'll read further on the literature you suggested to get a better hold of it.",
    "276112": "I am getting below error:\nTraceback (most recent call last):\n  File \"train.py\", line 591, in",
    "276128": "You need to install libjpeg-turbo \n\n\n\n    sudo apt install libturbojpeg",
    "276203": "Andres Check `clr_callback.py` a.k.a rate finder. Disclaimer, for me it didn't work. What I mean is that you'll have to wait 2-8 times the number of iterations per epoch for a phase to change in rate finding. I'd rather kill myself than have to wait that much time. I'm pretty sure that anyone can do much better at finding suitable lr by plug and play than waiting for any automatic algorithm to find the best optimum.",
    "276241": "add the lib64 directory to LD_LIBRARY_PATH",
    "276292": "kirk Im going to try it. This looks like a LR sweeper rather than a LR finder. I will let you know...",
    "276451": "Hey @Andres, I just have a question. I've noticed that whenever I resume training after checkpointing a model at some good `x` accuracy I always pay a loss of `x-20%` and no matter how long I train I can never hit again the same `x` accuracy. Have you experienced the same issue? And there is a discussion on github saying that there is an error with `save.model()` not saving the state of the optimizer.",
    "276458": "What Keras version are you using?",
    "276477": "In [3]: keras.__version__\n    Out[3]: '2.0.8'\n\nWhen you resume training do you also use the flags `-l`, `-uiw`?",
    "276480": "I use 2.1.3\n\nBiggest reason for penalty is learning rate is not saved, so if you don't specify it it will default to the initial so, check the last learning rate and put it with -l. -uiw not needed (has no effect ) when used with -m or -w",
    "276488": "Sorry my mistake I was checking keras version in terminal without being ssh in the actual server. I have the same keras version= 2.1.3. About your second comment here is the difficult part, during training the checkpoints are saving the model at best val_accuray but we don't have information on the actual learning rate at that stage. In the end we have a model saved in `.hdf5` format. How can we know the learning rate used when that particular checkpoint was saved?",
    "276512": "You may just guess it (or calculate). There is schedule in the script, halving LR after 5 epochs after last improvement. Just take a look at string of saved models you can decode between which of them there was a halving. And initial rate is given",
    "276520": "Sergey thanks. A minimal example would help. Particularly I am interested in this bit `Just take a look at string of saved models`. I am assuming that you imply that I should have all the model checkpoints in place. What if I have deleted all of them apart the one with highest accuracy. I am also very interested to know if there is an option saving the learning rate along with the model when you use `ModelCheckpoint` callback?",
    "276541": "Keras saves each model with highest metric, so if you have several of them for example: \nepoch1\t0.3\nepoch2\t0.35\nepoch3\t0.41\nepoch12\t0.51\nepoch15\t0.55\nepoch25\t0.6\nyou can see where halving happen: between model where there were more then 5 epochs between models",
    "276570": "Check documentation of `ModelCheckpoint` and see whether learning save in the `log` key which were the keywords like `val_acc` `epoch` etc are made available and it's how I construct the filename for the model, then if LR is available there place it accordingly in the filename and extract it with the `re.match` ... I use to extract `epoch` from the filename upon loading. LMK if it works.",
    "276624": "Andres thanks for the reply. Let's make a concrete example. Here is what is saved `VGG19_do0.3_doc0.0_avg-epoch086-val_acc0.329167.hdf5`. From this I don't really understand how you can extract the learning rate? How do you resume training in your case? Where do you get the learning rate from?",
    "276733": "What I do is look at the output of training, e.g:\n\n    poch 19/200\n    449/450 [============================&gt;.] - ETA: 0s - loss: 0.3424 - acc: 0.8899\n    Epoch 00019: ReduceLROnPlateau reducing learning rate to 4.999999873689376e-05.\n    450/450 [==============================] - 186s 413ms/step - loss: 0.3420 - acc: 0.8900 - val_loss: 0.6567 - val_acc: 0.8356\n    Epoch 20/200\n    450/450 [==============================] - 186s 414ms/step - loss: 0.2574 - acc: 0.9219 - val_loss: 0.4105 - val_acc: 0.9008\n\ncheck the latest `ReduceLROnPlateau` output, i.e. `Epoch 00019: ReduceLROnPlateau reducing learning rate to 4.999999873689376e-05` and when I restart training manually add `-l 5e-5`.",
    "277053": "Great approach! Way to go.",
    "277308": "Hi @Andres, apologies I couldn't respond earlier, I was traveling. I am assuming that you get the LR from the training output which probably you have saved somewhere. I don't have that. Imagine that you trained days ago in tmux or session. And now you've released that session and you don't have that info anymore, you haven't saved it anywhere, is there any other alternative? I tried `model = load_model()`, `model.fit()` but couldn't get any info like the one you gave: `449/450 [============================&gt;.] - ETA: 0s - loss: 0.3424 - acc: 0.8899\nEpoch 00019: ReduceLROnPlateau reducing learning rate to 4.999999873689376e-05.`",
    "277315": "Yes, I get the training for the output, using `tmux`.\n\nI have digged into the saved model and it does indeed contain the last LR:\n\n    &gt; from keras.models import load_model, Model\n    &gt; from keras import backend as K\n    &gt; model = load_model('models/your_model_____.hdf5')\n    &gt; model.optimizer.__dict__.keys()\n    dict_keys(['updates', 'weights', 'iterations', 'lr', 'beta_1', 'beta_2', 'decay', 'epsilon', 'initial_decay', 'amsgrad'])\n    &gt; model.optimizer.__dict__['lr']",
    "277355": "Andres thanks a lot but, here's what I get:\n![error][1]\n\n\n  [1]: http://i63.tinypic.com/2r3cql1.png",
    "277357": "I use nohup to save my process and help me retrospect better for future runs.",
    "277359": "You need to be on Keras 2.1.3, TF 1.4.1 and Python 3.6\n\nIf you saved that model w/ older version of Keras you're probably out of luck with that one.",
    "277369": "andres and @kirk can you help me with using Densenet201, I tried passing denset as parameter but I get this error,\n\n    python train.py -g 1 -b 8 -cs 512 -cm Densenet210 -x -l 1e-4 -uiw \n\nError:\n\n    Traceback (most recent call last):\n     File \"train.py\", line 466, in",
    "277378": "Change \"Densenet210\" to \"DenseNet201\".",
    "277380": "Apologies for such a typo, but I did try DenseNet201, However issue was with keras version.\nUpdating keras fixed it.",
    "277389": "Andres it is keras 2.1.3 and python 3.6 and tf 1.4.1.",
    "277463": "You need to have the model compiled to access `.optimizer`, in my latest code:\n\n`model = load_model(args.model, compile=False if args.test or (args.learning_rate is not None) else True)`",
    "277543": "Where can I find the lib.inputgenerator  and  lib.classhelper?",
    "277548": "If you are trying to use my fork please don't. I moved away from it because it was generating the wrong results. There was a bug causing wrong labeling during training and I never fixed it. Go ahead and use the official one from Andres: https://github.com/antorsae/sp-society-camera-model-identification\nSorry for any confusion.",
    "277627": "Andres awesome thanks a bunch ;)",
    "277901": "I give up with this fucking keras. No more, I've had enough. Even after setting the learning rate as @Andres suggested here is what I see.\n![Training][1]\n\n\n  [1]: http://i66.tinypic.com/34phpgm.png",
    "277999": "python train.py -g 1 -b 8 -cs 229 -cm Densenet201 -l 5e-5 -uiw\n\nTraceback (most recent call last):\n  File \"train.py\", line 467, in",
    "278002": "Your keras is old version, updating it will work",
    "278007": "YangLu update your keras to keras 2.1.3, previous keras doesn't have DenseNet pre-trained models.\n\n[Keras Documentation][1]\n\n\n  [1]: https://keras.io/applications/",
    "278039": "My keras is 2.1.3, I use serveral conda envs, maybe this disturb.",
    "278042": "check if this gets executed in your python console from your present environment:\n\n    from keras.applications.densenet import DenseNet201\n\nIf this works out well, then you have keras 2.1.3 and  are good to go.",
    "278044": "YangLu, how about to try DenseNet201, not Densenet201?",
    "278104": "Thanks @RK @shivrajp @Yaozj, It's my low level spelling mistakes. By the way, -cs 512 is too big for me , my one 1070Ti  can  work on \"-cs 229\".",
    "278321": "Yes， I have got the same problem， DenseNet needs huge memory.\nMoreover, the loss decrease very slowly, Maybe the learning rate is not set properly",
    "278347": "Awesome work! Thanks!",
    "278512": "I have another problem wiht extra dataset, Have you fixed it?\n541/953 [================&gt;.............] - ETA: 2:36 - loss: 2.0163 - acc: 0.3073\nException in thread Thread-8:\nFile \"/home/yl/miniconda3/envs/gluon/lib/python3.6/multiprocessing/connection.py\", line 251, in recv\n    return _ForkingPickler.loads(buf.getbuffer())\nTypeError: __init__() missing 1 required positional argument: 'code'\n\n# Add script as follows , fix it.\n\nload_img  = lambda img_path: np.array(Image.open(img_path))\n\ndef load_img_fast_jpg(img_path):\n    try:\n        x = jpeg.JPEG(img_path).decode()\n        return x\n    except:\n        return load_img(img_path)",
    "278883": "Hi guys,\n\nWhen I use the preprocessing kernel filters # http://www.lirmm.fr/~chaumont/publications/WIFS-2016_TUAMA_COMBY_CHAUMONT_Camera_Model_Identification_With_CNN_slides.pdf\n       \n     kernel_filter = 1/12. * np.array([\\\n            [-1,  2,  -2,  2, -1],  \\\n            [ 2, -6,   8, -6,  2],  \\\n            [-2,  8, -12,  8, -2],  \\\n            [ 2, -6,   8, -6,  2],  \\\n            [-1,  2,  -2,  2, -1]]) \n\nThe images after the filter are basically completely black with speckles of white. Is this supposed to be the case? \n\nThanks for the help!"
  },
  "source": "meta"
}