{
  "id": 47896,
  "title": "Sharing my code, looking for suggestions (0.934 LB)",
  "url": "/competitions/sp-society-camera-model-identification/discussion/47896",
  "author_name": "",
  "post_date": "2018-01-20T08:48:21.519063900Z",
  "votes": 60,
  "comment_count": 92,
  "views": 0,
  "content": "<p>Hi guys,</p>\n\n<p>I just started yesterday and I'd like to share my progress:</p>\n\n<p><a href=\"https://github.com/antorsae/sp-society-camera-model-identification\">https://github.com/antorsae/sp-society-camera-model-identification</a></p>\n\n<p>Current approach:</p>\n\n<ul>\n<li>Resizes train/val images to 512x512 (same as target test images)</li>\n<li>Takes random crops located at the edges of the images (total 4 possible different crops per image) of size <code>-cs</code> (defaults to 299)</li>\n<li>Since I believe the features learned by the classifier are location dependent (i.e. features on the top-left crop will be different than bottom-right crop) I concatenate the relative location of the crop to the features prior to the FC layer.</li>\n<li>Uses any of the Keras applications invoked by command line (e.g. <code>-cm ResNet50</code> uses ResNet50)</li>\n<li>Uses pooling as specified by <code>--pooling</code> i.e. avg|max|none</li>\n<li>Optionally applies a kernel filter as per this Slide 13 @ <a href=\"http://www.lirmm.fr/~chaumont/publications/WIFS-2016_TUAMA_COMBY_CHAUMONT_Camera_Model_Identification_With_CNN_slides.pdf\">http://www.lirmm.fr/~chaumont/publications/WIFS-2016_TUAMA_COMBY_CHAUMONT_Camera_Model_Identification_With_CNN_slides.pdf</a> with <code>-kf</code></li>\n<li>Loading full-size JPGs takes a lot of time so on first iteration of the Keras generator it builds a dictionary in memory with the resize (and preprocessed) images so after first 4 epochs (4 is b/c <code>sub_batch_size</code>)  it goes significantly faster (on my 2 GPUs subsequent epoch take ~20 secs).</li>\n</ul>\n\n<p>So far I only get ~0.87% train acc, and ~0.70% val acc. I have not submitted LB yet.</p>\n\n<p>Any comments / suggestions welcome!</p>",
  "messages": [
    {
      "id": "271358",
      "postDate": "01/20/2018 08:48:21",
      "content": "<p>Hi guys,</p>\n\n<p>I just started yesterday and I'd like to share my progress:</p>\n\n<p><a href=\"https://github.com/antorsae/sp-society-camera-model-identification\">https://github.com/antorsae/sp-society-camera-model-identification</a></p>\n\n<p>Current approach:</p>\n\n<ul>\n<li>Resizes train/val images to 512x512 (same as target test images)</li>\n<li>Takes random crops located at the edges of the images (total 4 possible different crops per image) of size <code>-cs</code> (defaults to 299)</li>\n<li>Since I believe the features learned by the classifier are location dependent (i.e. features on the top-left crop will be different than bottom-right crop) I concatenate the relative location of the crop to the features prior to the FC layer.</li>\n<li>Uses any of the Keras applications invoked by command line (e.g. <code>-cm ResNet50</code> uses ResNet50)</li>\n<li>Uses pooling as specified by <code>--pooling</code> i.e. avg|max|none</li>\n<li>Optionally applies a kernel filter as per this Slide 13 @ <a href=\"http://www.lirmm.fr/~chaumont/publications/WIFS-2016_TUAMA_COMBY_CHAUMONT_Camera_Model_Identification_With_CNN_slides.pdf\">http://www.lirmm.fr/~chaumont/publications/WIFS-2016_TUAMA_COMBY_CHAUMONT_Camera_Model_Identification_With_CNN_slides.pdf</a> with <code>-kf</code></li>\n<li>Loading full-size JPGs takes a lot of time so on first iteration of the Keras generator it builds a dictionary in memory with the resize (and preprocessed) images so after first 4 epochs (4 is b/c <code>sub_batch_size</code>)  it goes significantly faster (on my 2 GPUs subsequent epoch take ~20 secs).</li>\n</ul>\n\n<p>So far I only get ~0.87% train acc, and ~0.70% val acc. I have not submitted LB yet.</p>\n\n<p>Any comments / suggestions welcome!</p>",
      "rawMarkdown": "Hi guys,\n\nI just started yesterday and I'd like to share my progress:\n\nhttps://github.com/antorsae/sp-society-camera-model-identification\n\nCurrent approach:\n\n- Resizes train/val images to 512x512 (same as target test images)\n- Takes random crops located at the edges of the images (total 4 possible different crops per image) of size `-cs` (defaults to 299)\n- Since I believe the features learned by the classifier are location dependent (i.e. features on the top-left crop will be different than bottom-right crop) I concatenate the relative location of the crop to the features prior to the FC layer.\n- Uses any of the Keras applications invoked by command line (e.g. `-cm ResNet50` uses ResNet50)\n- Uses pooling as specified by `--pooling` i.e. avg|max|none\n- Optionally applies a kernel filter as per this Slide 13 @ http://www.lirmm.fr/~chaumont/publications/WIFS-2016_TUAMA_COMBY_CHAUMONT_Camera_Model_Identification_With_CNN_slides.pdf with `-kf`\n- Loading full-size JPGs takes a lot of time so on first iteration of the Keras generator it builds a dictionary in memory with the resize (and preprocessed) images so after first 4 epochs (4 is b/c `sub_batch_size`)  it goes significantly faster (on my 2 GPUs subsequent epoch take ~20 secs).\n\nSo far I only get ~0.87% train acc, and ~0.70% val acc. I have not submitted LB yet.\n\nAny comments / suggestions welcome!",
      "votes": null
    },
    {
      "id": "271445",
      "postDate": "01/20/2018 13:08:03",
      "content": "<p>Hey Andres!</p>\n\n<p>Nice to see you ;)</p>\n\n<p>Why do you resize the train images? The test images are not resized, they are center cropped:</p>\n\n<blockquote>\n  <p>While the train data includes full images, the test data contains only single 512 x 512 pixel blocks cropped from the center of a single image taken with the device.</p>\n</blockquote>",
      "rawMarkdown": "Hey Andres!\n\nNice to see you ;)\n\nWhy do you resize the train images? The test images are not resized, they are center cropped:\n\n&gt; While the train data includes full images, the test data contains only single 512 x 512 pixel blocks cropped from the center of a single image taken with the device.",
      "votes": null
    },
    {
      "id": "271448",
      "postDate": "01/20/2018 13:21:16",
      "content": "<p>Hi Ali! yes, I just noticed... just pushed new code, much better now:</p>\n\n<p><code>Epoch 29/200\n139/139 [==============================] - 53s 384ms/step - loss: 0.2487 - acc: 0.9101 - val_loss: 0.4792 - val_acc: 0.8750</code></p>",
      "rawMarkdown": "Hi Ali! yes, I just noticed... just pushed new code, much better now:\n\n`Epoch 29/200\n139/139 [==============================] - 53s 384ms/step - loss: 0.2487 - acc: 0.9101 - val_loss: 0.4792 - val_acc: 0.8750`",
      "votes": null
    },
    {
      "id": "272157",
      "postDate": "01/22/2018 11:29:43",
      "content": "<p>I think it is not a good idea to resize the training images</p>",
      "rawMarkdown": "I think it is not a good idea to resize the training images",
      "votes": null
    },
    {
      "id": "272178",
      "postDate": "01/22/2018 12:54:19",
      "content": "<p>Agreed, I've updated the code and it doesn't resize anymore. </p>",
      "rawMarkdown": "Agreed, I've updated the code and it doesn't resize anymore.",
      "votes": null
    },
    {
      "id": "272181",
      "postDate": "01/22/2018 12:55:57",
      "content": "<p>Update:\n- No resizing, just takes 512x512 crops from the center</p>\n\n<p>The good news is that it achieves ~90% val acc, however LB is ~30-40% only.  Currently investigating why.</p>",
      "rawMarkdown": "Update:\n- No resizing, just takes 512x512 crops from the center\n\nThe good news is that it achieves ~90% val acc, however LB is ~30-40% only.  Currently investigating why.",
      "votes": null
    },
    {
      "id": "272213",
      "postDate": "01/22/2018 14:32:19",
      "content": "<p>I have similar issue. I'm trying to resolve it with no success.  If I will resolve it somehow, I will let you know.\nMy model is ResNet50 finetuned with random crop from full size images. </p>",
      "rawMarkdown": "I have similar issue. I'm trying to resolve it with no success.  If I will resolve it somehow, I will let you know.\nMy model is ResNet50 finetuned with random crop from full size images.",
      "votes": null
    },
    {
      "id": "272216",
      "postDate": "01/22/2018 14:54:07",
      "content": "<p>What size of your crop? I have the same issue.</p>",
      "rawMarkdown": "What size of your crop? I have the same issue.",
      "votes": null
    },
    {
      "id": "272217",
      "postDate": "01/22/2018 14:56:03",
      "content": "<p>512x512, same as the test size.</p>",
      "rawMarkdown": "512x512, same as the test size.",
      "votes": null
    },
    {
      "id": "272235",
      "postDate": "01/22/2018 15:42:37",
      "content": "<p>It seems the validation accuracy on its own does not mean much. I had a model with both training and validation accuracy above 98%, but the LB was only 64%. On the other hand, I had a model with train accuracy equal to 91 and validation accuracy equal to 95, and the LB accuracy for that model was 85.2. I am totally confused!</p>",
      "rawMarkdown": "It seems the validation accuracy on its own does not mean much. I had a model with both training and validation accuracy above 98%, but the LB was only 64%. On the other hand, I had a model with train accuracy equal to 91 and validation accuracy equal to 95, and the LB accuracy for that model was 85.2. I am totally confused!",
      "votes": null
    },
    {
      "id": "272236",
      "postDate": "01/22/2018 15:44:02",
      "content": "<p>Maybe I should reconsider to create a whole new validation set. I think this is the key point. Maybe the high number of submissions of top LBs is a confirmation! I think they have tried many models until they have found a good one.</p>",
      "rawMarkdown": "Maybe I should reconsider to create a whole new validation set. I think this is the key point. Maybe the high number of submissions of top LBs is a confirmation! I think they have tried many models until they have found a good one.",
      "votes": null
    },
    {
      "id": "272242",
      "postDate": "01/22/2018 15:58:09",
      "content": "<p>I find when I use the validation data from Gleb's post, my validation accuracy is close to my LB score. I trained a model to 0.79 on the Gleb validation set and got 0.804 on the public leaderboard.</p>",
      "rawMarkdown": "I find when I use the validation data from Gleb's post, my validation accuracy is close to my LB score. I trained a model to 0.79 on the Gleb validation set and got 0.804 on the public leaderboard.",
      "votes": null
    },
    {
      "id": "272300",
      "postDate": "01/22/2018 19:20:08",
      "content": "<p>From my experience, I think you could get a LB score &gt; 0.90 just cropping to 112x112. That way you can train your models faster. Make some experiments until you get good parameters (LR, dense layers size, drop, augmentation). When you have a good one, scale up the cropping size to get better results.\nRegarding the validation/LB gap, just for reference, I got 0.957/0.922 (val/LB). Using additional data should help to reduce the gap.</p>",
      "rawMarkdown": "From my experience, I think you could get a LB score &gt; 0.90 just cropping to 112x112. That way you can train your models faster. Make some experiments until you get good parameters (LR, dense layers size, drop, augmentation). When you have a good one, scale up the cropping size to get better results.\nRegarding the validation/LB gap, just for reference, I got 0.957/0.922 (val/LB). Using additional data should help to reduce the gap.",
      "votes": null
    },
    {
      "id": "272308",
      "postDate": "01/22/2018 19:39:37",
      "content": "<p>Hi Daniel, You are fine tuning some pre trained model or using your own model?</p>",
      "rawMarkdown": "Hi Daniel, You are fine tuning some pre trained model or using your own model?",
      "votes": null
    },
    {
      "id": "272500",
      "postDate": "01/23/2018 06:20:52",
      "content": "<p>Hi Daniel, did you use Gleb's dataset for your 0.922 LB or just the organization-provided dataset?</p>",
      "rawMarkdown": "Hi Daniel, did you use Gleb's dataset for your 0.922 LB or just the organization-provided dataset?",
      "votes": null
    },
    {
      "id": "272526",
      "postDate": "01/23/2018 07:27:46",
      "content": "<p>Update:\nUsing Gleb's dataset I get 0.76 LB with a ResNet50 and pre-trained weights and a fixed 224 crop size.</p>\n\n<p>Next steps:</p>\n\n<ul>\n<li>Balance classes (DONE)</li>\n<li>Fine-tune FC, dropout, whether to use high-pass kernel filtering (<code>-kf</code>) and determine best architecture</li>\n<li>Account for different weights (0.3 vs. 0.7) in loss function</li>\n<li>Get more data (I think this is critical)</li>\n</ul>",
      "rawMarkdown": "Update:\nUsing Gleb's dataset I get 0.76 LB with a ResNet50 and pre-trained weights and a fixed 224 crop size.\n\nNext steps:\n\n - Balance classes (DONE)\n - Fine-tune FC, dropout, whether to use high-pass kernel filtering (`-kf`) and determine best architecture\n - Account for different weights (0.3 vs. 0.7) in loss function\n - Get more data (I think this is critical)",
      "votes": null
    },
    {
      "id": "272776",
      "postDate": "01/23/2018 17:21:30",
      "content": "<p>Just the organization-provided dataset. Now using Gleb's dataset too.</p>",
      "rawMarkdown": "Just the organization-provided dataset. Now using Gleb's dataset too.",
      "votes": null
    },
    {
      "id": "272780",
      "postDate": "01/23/2018 17:42:40",
      "content": "<p>So the .76 is just with the built in Keras ResNet50? Since you mention fine-tuning the FC does that mean you didn't even replace the FC?</p>",
      "rawMarkdown": "So the .76 is just with the built in Keras ResNet50? Since you mention fine-tuning the FC does that mean you didn't even replace the FC?",
      "votes": null
    },
    {
      "id": "272798",
      "postDate": "01/23/2018 18:28:38",
      "content": "<p>I've already changed the code a lot but IIRC it was something like <code>python train.py -l 1e-3 -b 64 -g 2 -cm ResNet50 -cs 224 -x  -p avg -do 0.3  -kf</code></p>\n\n<p>I added random crops to the current code so score may be different now. I've also added crop ensembling to <code>-test</code> so it may help.\nSummary of paramaters:\n - <code>-b 64</code> batch size for each GPU\n - <code>-g 2</code> use 2 GPUs\n - <code>-cm ResNet50</code> use ResNet50 as feature extractor\n - <code>-cs 224</code> crop size\n - <code>-x</code> use Gleb's dataset\n - <code>-p avg</code> use average pooling on the classifier (ResNet50)\n - <code>-do 0.3</code> use 0.3 dropout rate on my FC layers \n - <code>-kf</code> use a high pass pre-processing filter.</p>\n\n<p>You may want to take a look at the code to try it out. Once you have a model with decent <code>val_acc</code> do <code>-m path_to_saved_model.hdf5 -t</code> to generate <code>submission.csv</code> </p>",
      "rawMarkdown": "I've already changed the code a lot but IIRC it was something like `python train.py -l 1e-3 -b 64 -g 2 -cm ResNet50 -cs 224 -x  -p avg -do 0.3  -kf`\n\nI added random crops to the current code so score may be different now. I've also added crop ensembling to `-test` so it may help.\nSummary of paramaters:\n - `-b 64` batch size for each GPU\n - `-g 2` use 2 GPUs\n - `-cm ResNet50` use ResNet50 as feature extractor\n - `-cs 224` crop size\n - `-x` use Gleb's dataset\n - `-p avg` use average pooling on the classifier (ResNet50)\n - `-do 0.3` use 0.3 dropout rate on my FC layers \n - `-kf` use a high pass pre-processing filter.\n\nYou may want to take a look at the code to try it out. Once you have a model with decent `val_acc` do `-m path_to_saved_model.hdf5 -t` to generate `submission.csv`",
      "votes": null
    },
    {
      "id": "272959",
      "postDate": "01/24/2018 01:25:22",
      "content": "<p>I went ahead and got your code from github. </p>\n\n<p>I edited the post as the errors where caused by an unknown issue.</p>\n\n<p>I have successfully run your code and got results in the  0.503 LB using the flags above (except for the flickr data, which I still have to get properly)</p>\n\n<p>For my submissions (0.517 LB) I was using a resnet like architecture with 3 stages training from scratch. It was taking 3-6 hours for training and your model is only taking me 1 hour. So now I have something better to look at.</p>\n\n<p>Looking at the CaCNN code and the paper you reference I guess I'll be running that next.</p>\n\n<p>Thanks for sharing the code.</p>",
      "rawMarkdown": "I went ahead and got your code from github. \n\nI edited the post as the errors where caused by an unknown issue.\n\nI have successfully run your code and got results in the  0.503 LB using the flags above (except for the flickr data, which I still have to get properly)\n\nFor my submissions (0.517 LB) I was using a resnet like architecture with 3 stages training from scratch. It was taking 3-6 hours for training and your model is only taking me 1 hour. So now I have something better to look at.\n\nLooking at the CaCNN code and the paper you reference I guess I'll be running that next.\n\nThanks for sharing the code.",
      "votes": null
    },
    {
      "id": "273376",
      "postDate": "01/24/2018 14:30:24",
      "content": "<p>Great. I am finishing a multiprocess generator which does not need to cache (memory or disk) and has the benefit of being able to select random crops from all across the image. Will be ready in a few hours. Stay tuned.</p>",
      "rawMarkdown": "Great. I am finishing a multiprocess generator which does not need to cache (memory or disk) and has the benefit of being able to select random crops from all across the image. Will be ready in a few hours. Stay tuned.",
      "votes": null
    },
    {
      "id": "273388",
      "postDate": "01/24/2018 14:55:25",
      "content": "<p>Great, I look forward to looking at it.</p>\n\n<p>By the way let me know if you would like to merge as a team. </p>\n\n<p>I've been using hdf5 files with unprocessed and pre-processed patches, speed wise they work well. I'm at the stage of testing multiple networks and improving results.</p>",
      "rawMarkdown": "Great, I look forward to looking at it.\n\nBy the way let me know if you would like to merge as a team. \n\nI've been using hdf5 files with unprocessed and pre-processed patches, speed wise they work well. I'm at the stage of testing multiple networks and improving results.",
      "votes": null
    },
    {
      "id": "273414",
      "postDate": "01/24/2018 15:25:45",
      "content": "<p>Code is updated, now with multiprocess generator which is more flexible, especially wrt selecting random crops (previously only a smaller version of the original image was saved in the cache so cropping was more limited). </p>\n\n<p>I think you need a powerful CPU to do all augmentations, etc. and saturate your GPUs. </p>",
      "rawMarkdown": "Code is updated, now with multiprocess generator which is more flexible, especially wrt selecting random crops (previously only a smaller version of the original image was saved in the cache so cropping was more limited). \n\nI think you need a powerful CPU to do all augmentations, etc. and saturate your GPUs.",
      "votes": null
    },
    {
      "id": "273523",
      "postDate": "01/24/2018 18:07:19",
      "content": "<p>The new code achieves ~0.73 just after 8 epochs:\n<code>$ python train.py -g 2 -b 16 -cs 512 -cm ResNet50 -x -l 1e-4 -uiw</code></p>\n\n<p>Epoch 1/100\n507/507 [==============================] - 227s 447ms/step - loss: 1.6933 - acc: 0.4243 - val_loss: 1.6242 - val_acc: 0.5395\nEpoch 2/100\n507/507 [==============================] - 205s 405ms/step - loss: 1.0983 - acc: 0.6495 - val_loss: 1.1173 - val_acc: 0.6711\nEpoch 3/100\n507/507 [==============================] - 207s 408ms/step - loss: 0.8839 - acc: 0.7187 - val_loss: 1.6343 - val_acc: 0.5230\nEpoch 4/100\n507/507 [==============================] - 206s 406ms/step - loss: 0.7618 - acc: 0.7625 - val_loss: 0.9526 - val_acc: 0.7105\nEpoch 5/100\n507/507 [==============================] - 204s 403ms/step - loss: 0.6861 - acc: 0.7855 - val_loss: 1.0580 - val_acc: 0.6941\nEpoch 6/100\n507/507 [==============================] - 203s 399ms/step - loss: 0.6055 - acc: 0.8104 - val_loss: 0.7093 - val_acc: 0.7599\nEpoch 7/100\n507/507 [==============================] - 197s 389ms/step - loss: 0.5509 - acc: 0.8274 - val_loss: 1.1975 - val_acc: 0.6875\nEpoch 8/100\n507/507 [==============================] - 351s 693ms/step - loss: 0.5159 - acc: 0.8410 - val_loss: 0.8306 - val_acc: 0.8125</p>\n\n<p>(it then crashed due to a memory leak... need to investigate)</p>\n\n<p>then:\n<code>$ python train.py -g 1 -b 16 -cs 512 -cm ResNet50 -x -l 1e-4 -uiw -m models/ResNet50_do0.3_avg-epoch008-val_acc0.812500.hdf5 -t\n</code></p>\n\n<p>then:\n<code>$ kg submit submission.csv\n0.733\n</code></p>\n\n<p>I've noticed that in Gleb's dataset some images did not match any resolution, so I discard them. Also Sony-NEX-7 directory had files ending in JPG, not jpg... so my code was not using them. Renamed then.</p>\n\n<p>Finally I noted some images are portrait whereas others are landscape, so now I'm augmenting via orientation change during training.</p>",
      "rawMarkdown": "The new code achieves ~0.73 just after 8 epochs:\n`$ python train.py -g 2 -b 16 -cs 512 -cm ResNet50 -x -l 1e-4 -uiw`\n\nEpoch 1/100\n507/507 [==============================] - 227s 447ms/step - loss: 1.6933 - acc: 0.4243 - val_loss: 1.6242 - val_acc: 0.5395\nEpoch 2/100\n507/507 [==============================] - 205s 405ms/step - loss: 1.0983 - acc: 0.6495 - val_loss: 1.1173 - val_acc: 0.6711\nEpoch 3/100\n507/507 [==============================] - 207s 408ms/step - loss: 0.8839 - acc: 0.7187 - val_loss: 1.6343 - val_acc: 0.5230\nEpoch 4/100\n507/507 [==============================] - 206s 406ms/step - loss: 0.7618 - acc: 0.7625 - val_loss: 0.9526 - val_acc: 0.7105\nEpoch 5/100\n507/507 [==============================] - 204s 403ms/step - loss: 0.6861 - acc: 0.7855 - val_loss: 1.0580 - val_acc: 0.6941\nEpoch 6/100\n507/507 [==============================] - 203s 399ms/step - loss: 0.6055 - acc: 0.8104 - val_loss: 0.7093 - val_acc: 0.7599\nEpoch 7/100\n507/507 [==============================] - 197s 389ms/step - loss: 0.5509 - acc: 0.8274 - val_loss: 1.1975 - val_acc: 0.6875\nEpoch 8/100\n507/507 [==============================] - 351s 693ms/step - loss: 0.5159 - acc: 0.8410 - val_loss: 0.8306 - val_acc: 0.8125\n\n(it then crashed due to a memory leak... need to investigate)\n\nthen:\n```$ python train.py -g 1 -b 16 -cs 512 -cm ResNet50 -x -l 1e-4 -uiw -m models/ResNet50_do0.3_avg-epoch008-val_acc0.812500.hdf5 -t\n```\n\nthen:\n```$ kg submit submission.csv\n0.733\n```\n\nI've noticed that in Gleb's dataset some images did not match any resolution, so I discard them. Also Sony-NEX-7 directory had files ending in JPG, not jpg... so my code was not using them. Renamed then.\n\nFinally I noted some images are portrait whereas others are landscape, so now I'm augmenting via orientation change during training.",
      "votes": null
    },
    {
      "id": "273550",
      "postDate": "01/24/2018 18:55:50",
      "content": "<p>I'm still screwing up the execution somehow. Once again I got a run with no learning:</p>\n\n<p>Epoch 100/100\n245/245 [==============================] - 81s - loss: 2.2325 - acc: 0.1793 - val_loss: 2.3201 - val_acc: 0.1562</p>\n\n<p>I'm wiping out the whole models directory and restarted the training with:</p>\n\n<p>-l 1e-3 -b 32 -g 1 -cm ResNet50 -cs 224 -p avg -do 0.3 -x -kf -uiw</p>\n\n<p>We'll see if this run it learns (I wasn't using the uiw before but I can't see how that would explain the model not learning). </p>\n\n<p>Not sure why you got JPG, all my links in flickr_images/sony_nex7/urls_final were lower case.</p>",
      "rawMarkdown": "I'm still screwing up the execution somehow. Once again I got a run with no learning:\n\nEpoch 100/100\n245/245 [==============================] - 81s - loss: 2.2325 - acc: 0.1793 - val_loss: 2.3201 - val_acc: 0.1562\n\nI'm wiping out the whole models directory and restarted the training with:\n\n-l 1e-3 -b 32 -g 1 -cm ResNet50 -cs 224 -p avg -do 0.3 -x -kf -uiw\n\nWe'll see if this run it learns (I wasn't using the uiw before but I can't see how that would explain the model not learning). \n\nNot sure why you got JPG, all my links in flickr_images/sony_nex7/urls_final were lower case.",
      "votes": null
    },
    {
      "id": "273553",
      "postDate": "01/24/2018 19:07:53",
      "content": "<p>Try changing (lowering) the learning rate. Also, intuitively I don't think <code>-uiw</code> would work OK with <code>-kf</code> the reason being is <code>-kf</code> applies a high-pass filter on the images and they cease to look anything like an image, and also the normalization is different than in imagenet (mean is similar but std and range are not).</p>\n\n<p>The JPG vs. jpg was on the train directory (organization dataset).</p>",
      "rawMarkdown": "Try changing (lowering) the learning rate. Also, intuitively I don't think `-uiw` would work OK with `-kf` the reason being is `-kf` applies a high-pass filter on the images and they cease to look anything like an image, and also the normalization is different than in imagenet (mean is similar but std and range are not).\n\nThe JPG vs. jpg was on the train directory (organization dataset).",
      "votes": null
    },
    {
      "id": "273591",
      "postDate": "01/24/2018 21:19:39",
      "content": "<p>Update:  0.868 LB now with latest code changes and <code>python train.py -g 2 -b 16 -cs 512 -cm ResNet50 -x -l 1e-4 -uiw</code></p>",
      "rawMarkdown": "Update:  0.868 LB now with latest code changes and `python train.py -g 2 -b 16 -cs 512 -cm ResNet50 -x -l 1e-4 -uiw`",
      "votes": null
    },
    {
      "id": "273628",
      "postDate": "01/24/2018 22:02:15",
      "content": "<p>Awesome! What was the change compare to your last model?</p>",
      "rawMarkdown": "Awesome! What was the change compare to your last model?",
      "votes": null
    },
    {
      "id": "273637",
      "postDate": "01/24/2018 22:16:39",
      "content": "<p>I rewrote the generator to make it multiprocessor friendly. Architecturally is the same but I think there may have been a bug hidden in the old code. Also the new generator has more liberty in selecting random crops and now does orientation augmentation too. Will leave it overnight to see how far it converges.</p>",
      "rawMarkdown": "I rewrote the generator to make it multiprocessor friendly. Architecturally is the same but I think there may have been a bug hidden in the old code. Also the new generator has more liberty in selecting random crops and now does orientation augmentation too. Will leave it overnight to see how far it converges.",
      "votes": null
    },
    {
      "id": "273652",
      "postDate": "01/24/2018 23:21:27",
      "content": "<p>I can't seem to get it to run. When I saw this before my directories didn't match the script expectations, but this time I seem to have the right numbers on the image data:</p>\n\n<pre><code>Image set counts\n</code></pre>\n\n<p>extra_train_ids:  5368\n  extra_val_ids:   315</p>\n\n<pre><code>  ids_train:  8118 steps=  507\n    ids_val:   315 steps=   19\n</code></pre>\n\n<hr>\n\n<pre><code>          HTC-1-M7:  1023 (12.6%)\n          iPhone-6:   823 (10.1%)\n</code></pre>\n\n<p>Motorola-Droid-Maxx:   825 (10.2%)\n            Motorola-X:   275 (03.4%)\n     Samsung-Galaxy-S4:  1412 (17.4%)\n             iPhone-4s:   774 (09.5%)\n           LG-Nexus-5x:   680 (08.4%)\n      Motorola-Nexus-6:   926 (11.4%)\n  Samsung-Galaxy-Note3:   548 (06.8%)\n            Sony-NEX-7:   832 (10.2%)</p>\n\n<pre><code>Exception in thread Thread-1:\n</code></pre>\n\n<p>multiprocessing.pool.RemoteTraceback: \n\"\"\"\nTraceback (most recent call last):\n  File \"/usr/lib/python3.5/multiprocessing/pool.py\", line 119, in worker\n    result = (True, func(*args, **kwds))\n  File \"/usr/lib/python3.5/multiprocessing/pool.py\", line 44, in mapstar\n    return list(map(*args))\n  File \"train.py\", line 265, in process_item\n    img = preprocess_image(img)\n  File \"train.py\", line 197, in preprocess_image\n    return preprocess_input_function(img.astype(np.float32))\n  File \"/home/alonsoa/projects/virtenv/py3cv3/lib/python3.5/site-packages/keras/applications/imagenet_utils.py\", line 33, in preprocess_input\n    x = x[:, :, :, ::-1]\nIndexError: too many indices for array\n\"\"\"</p>\n\n<p>Anything jumps at you why   return preprocess_input_function(img.astype(np.float32)) would fail even though the image is shaped (512, 512, 3)\"</p>",
      "rawMarkdown": "I can't seem to get it to run. When I saw this before my directories didn't match the script expectations, but this time I seem to have the right numbers on the image data:\n\n    Image set counts\n\nextra_train_ids:  5368\n  extra_val_ids:   315\n\n      ids_train:  8118 steps=  507\n        ids_val:   315 steps=   19\n____________________________________________________________________________________________________\n\n              HTC-1-M7:  1023 (12.6%)\n              iPhone-6:   823 (10.1%)\n   Motorola-Droid-Maxx:   825 (10.2%)\n            Motorola-X:   275 (03.4%)\n     Samsung-Galaxy-S4:  1412 (17.4%)\n             iPhone-4s:   774 (09.5%)\n           LG-Nexus-5x:   680 (08.4%)\n      Motorola-Nexus-6:   926 (11.4%)\n  Samsung-Galaxy-Note3:   548 (06.8%)\n            Sony-NEX-7:   832 (10.2%)\n\n\n    Exception in thread Thread-1:\nmultiprocessing.pool.RemoteTraceback: \n\"\"\"\nTraceback (most recent call last):\n  File \"/usr/lib/python3.5/multiprocessing/pool.py\", line 119, in worker\n    result = (True, func(*args, **kwds))\n  File \"/usr/lib/python3.5/multiprocessing/pool.py\", line 44, in mapstar\n    return list(map(*args))\n  File \"train.py\", line 265, in process_item\n    img = preprocess_image(img)\n  File \"train.py\", line 197, in preprocess_image\n    return preprocess_input_function(img.astype(np.float32))\n  File \"/home/alonsoa/projects/virtenv/py3cv3/lib/python3.5/site-packages/keras/applications/imagenet_utils.py\", line 33, in preprocess_input\n    x = x[:, :, :, ::-1]\nIndexError: too many indices for array\n\"\"\"\n\nAnything jumps at you why   return preprocess_input_function(img.astype(np.float32)) would fail even though the image is shaped (512, 512, 3)\"",
      "votes": null
    },
    {
      "id": "273655",
      "postDate": "01/24/2018 23:36:00",
      "content": "<p>I had the same error. Sometimes the images don't download correctly and they read in as an array with less than three dimensions. I identified some of these images and downloaded them again, but had trouble getting them all. So I added logic that would check that len(img.shape) == 3 and if not move on to the next image.</p>",
      "rawMarkdown": "I had the same error. Sometimes the images don't download correctly and they read in as an array with less than three dimensions. I identified some of these images and downloaded them again, but had trouble getting them all. So I added logic that would check that len(img.shape) == 3 and if not move on to the next image.",
      "votes": null
    },
    {
      "id": "273657",
      "postDate": "01/24/2018 23:43:03",
      "content": "<p>Thanks, unfortunately there already is:</p>\n\n<pre><code>    if img.ndim != 3:\n    return None\n</code></pre>",
      "rawMarkdown": "Thanks, unfortunately there already is:\n\n        if img.ndim != 3:\n        return None",
      "votes": null
    },
    {
      "id": "273704",
      "postDate": "01/25/2018 02:35:14",
      "content": "<p>Well done, Andres. I have some question. Did you crop image to 512*512? Did you crop image from center or just randomly? how did you set your validation set? Thanks in advance~</p>",
      "rawMarkdown": "Well done, Andres. I have some question. Did you crop image to 512*512? Did you crop image from center or just randomly? how did you set your validation set? Thanks in advance~",
      "votes": null
    },
    {
      "id": "273711",
      "postDate": "01/25/2018 02:53:59",
      "content": "<p>He crops around the center at 2 * patch_size. Then a random manipulation (or not) and a final crop.</p>\n\n<p>Validation for the above results seems to be based on gleb's additional data (see the -x flag)</p>",
      "rawMarkdown": "He crops around the center at 2 * patch_size. Then a random manipulation (or not) and a final crop.\n\nValidation for the above results seems to be based on gleb's additional data (see the -x flag)",
      "votes": null
    },
    {
      "id": "273713",
      "postDate": "01/25/2018 02:55:51",
      "content": "<p>I ended up having to change the line to:</p>\n\n<pre><code>    return preprocess_input_function(np.expand_dims(img.astype(np.float32), axis=0))\n</code></pre>\n\n<p>as Keras' preprocess_input expects an array of images.</p>",
      "rawMarkdown": "I ended up having to change the line to:\n\n        return preprocess_input_function(np.expand_dims(img.astype(np.float32), axis=0))\n\nas Keras' preprocess_input expects an array of images.",
      "votes": null
    },
    {
      "id": "273718",
      "postDate": "01/25/2018 02:59:33",
      "content": "<p>The new code successfully kills (I mean gives a full workout ;-) ) all my CPU cores.</p>\n\n<p>After 1.5 hours I'm still at step 180 out of 980 of epoch 1. I'll give it 10-12 hours and see where it gets.</p>",
      "rawMarkdown": "The new code successfully kills (I mean gives a full workout ;-) ) all my CPU cores.\n\nAfter 1.5 hours I'm still at step 180 out of 980 of epoch 1. I'll give it 10-12 hours and see where it gets.",
      "votes": null
    },
    {
      "id": "273813",
      "postDate": "01/25/2018 08:05:33",
      "content": "<p>What version of Keras are you running? Mine:</p>\n\n<pre><code>$ python -c 'import keras;print(keras.__version__)'\n/home/antor/miniconda3/lib/python3.6/site-packages/h5py/__init__.py:36: FutureWarning: Conversion of the second argument of issubdtype from `float` to `np.floating` is deprecated. In future, it will be treated as `np.float64 == np.dtype(float).type`.\n  from ._conv import register_converters as _register_converters\nUsing TensorFlow backend.\n2.1.3\n</code></pre>\n\n<p>There's slightly different preprocessing functions on keras.applications and the code selects the appropriate one for each of the networks in keras.applications.* and defaults to xception (which is just normalizing between -1 and 1 iirc) for all others).</p>",
      "rawMarkdown": "What version of Keras are you running? Mine:\n\n    $ python -c 'import keras;print(keras.__version__)'\n    /home/antor/miniconda3/lib/python3.6/site-packages/h5py/__init__.py:36: FutureWarning: Conversion of the second argument of issubdtype from `float` to `np.floating` is deprecated. In future, it will be treated as `np.float64 == np.dtype(float).type`.\n      from ._conv import register_converters as _register_converters\n    Using TensorFlow backend.\n    2.1.3\n\nThere's slightly different preprocessing functions on keras.applications and the code selects the appropriate one for each of the networks in keras.applications.* and defaults to xception (which is just normalizing between -1 and 1 iirc) for all others).",
      "votes": null
    },
    {
      "id": "273814",
      "postDate": "01/25/2018 08:10:02",
      "content": "<p>Here's the output for me:</p>\n\n<pre><code>              HTC-1-M7:  1023 (12.6%)\n              iPhone-6:   823 (10.1%)\n   Motorola-Droid-Maxx:   825 (10.2%)\n            Motorola-X:   275 (03.4%)\n     Samsung-Galaxy-S4:  1412 (17.4%)\n             iPhone-4s:   774 (09.5%)\n           LG-Nexus-5x:   680 (08.4%)\n      Motorola-Nexus-6:   926 (11.4%)\n  Samsung-Galaxy-Note3:   548 (06.8%)\n            Sony-NEX-7:   832 (10.2%)\nEpoch 59/200\n507/507 [==============================] - 231s 456ms/step - loss: 0.2178 - acc: 0.9339 - val_loss: 0.2700 - val_acc: 0.9243\n</code></pre>\n\n<p>Each epoch takes ~230 secs or less on my machine (2 x 1080 Tis + AMD Threadripper 1950X...). Re: samples, there's ~507*16 (batch size)= ~8112 samples per epoch with <code>-x</code>. Some of them are now discarded b/c of resolution.</p>",
      "rawMarkdown": "Here's the output for me:\n\n                  HTC-1-M7:  1023 (12.6%)\n                  iPhone-6:   823 (10.1%)\n       Motorola-Droid-Maxx:   825 (10.2%)\n                Motorola-X:   275 (03.4%)\n         Samsung-Galaxy-S4:  1412 (17.4%)\n                 iPhone-4s:   774 (09.5%)\n               LG-Nexus-5x:   680 (08.4%)\n          Motorola-Nexus-6:   926 (11.4%)\n      Samsung-Galaxy-Note3:   548 (06.8%)\n                Sony-NEX-7:   832 (10.2%)\n    Epoch 59/200\n    507/507 [==============================] - 231s 456ms/step - loss: 0.2178 - acc: 0.9339 - val_loss: 0.2700 - val_acc: 0.9243\n\nEach epoch takes ~230 secs or less on my machine (2 x 1080 Tis + AMD Threadripper 1950X...). Re: samples, there's ~507*16 (batch size)= ~8112 samples per epoch with `-x`. Some of them are now discarded b/c of resolution.",
      "votes": null
    },
    {
      "id": "273865",
      "postDate": "01/25/2018 11:00:07",
      "content": "<p>Update, I left it overnight and it now achieves 0.934 LB with a single model. At this point I'm going to try a few new ideas. Kaggle competitions are usually won by huge ensembles and while \"everything goes\" I'm going to try to stick to simple stuff. Some ideas</p>\n\n<ul>\n<li>Class-aware sampling vs. class weighting</li>\n<li>Mixup</li>\n<li>New dataset</li>\n<li>Pretraining or faux-labeling</li>\n<li>Better TTA</li>\n</ul>",
      "rawMarkdown": "Update, I left it overnight and it now achieves 0.934 LB with a single model. At this point I'm going to try a few new ideas. Kaggle competitions are usually won by huge ensembles and while \"everything goes\" I'm going to try to stick to simple stuff. Some ideas\n\n - Class-aware sampling vs. class weighting\n - Mixup\n - New dataset\n - Pretraining or faux-labeling\n - Better TTA",
      "votes": null
    },
    {
      "id": "273916",
      "postDate": "01/25/2018 13:51:55",
      "content": "<p>I am running Keras (2.0.6) and I was running tensorflow (1.3.0)</p>\n\n<p>Turns out that last night I was getting CUDNN_STATUS_INTERNAL_ERROR which made me upgrade tensorflow to (1.4.1). After doing that I was no longer using the GPU and I didn't notice. That explains my previous comment about the CPU high usage, it was running the model. After upgrading tensorflow-gpu as well to (1.4.1) I'm back on the GPU (a GTX 1080Ti )</p>\n\n<p>As for samples the only difference between us is your Sony-NEX-7:   832 (10.2%) vs. mine Sony-NEX-7:   557 (07.1%) so I'll double check which samples I'm missing.</p>\n\n<p>I'm still having some performance and learning issues though. I kept running out of memory unless I brought the batch size to 1</p>\n\n<pre><code> -g 1 -b 1 -cs 512 -cm ResNet50 -x -l 1e-4 -uiw --max-epoch 350\n</code></pre>\n\n<p>Epoch 19/350\n7843/7843 [==============================] - 1159s - loss: 1.3477 - acc: 0.5560 - val_loss: 3.1935 - val_acc: 0.1108</p>\n\n<p>I'll upgrade keras and rerun</p>\n\n<p>Thanks for all the help. I hope my info helps anybody else trying to reproduce the results.</p>",
      "rawMarkdown": "I am running Keras (2.0.6) and I was running tensorflow (1.3.0)\n\nTurns out that last night I was getting CUDNN_STATUS_INTERNAL_ERROR which made me upgrade tensorflow to (1.4.1). After doing that I was no longer using the GPU and I didn't notice. That explains my previous comment about the CPU high usage, it was running the model. After upgrading tensorflow-gpu as well to (1.4.1) I'm back on the GPU (a GTX 1080Ti )\n\nAs for samples the only difference between us is your Sony-NEX-7:   832 (10.2%) vs. mine Sony-NEX-7:   557 (07.1%) so I'll double check which samples I'm missing.\n\nI'm still having some performance and learning issues though. I kept running out of memory unless I brought the batch size to 1\n\n     -g 1 -b 1 -cs 512 -cm ResNet50 -x -l 1e-4 -uiw --max-epoch 350\n\nEpoch 19/350\n7843/7843 [==============================] - 1159s - loss: 1.3477 - acc: 0.5560 - val_loss: 3.1935 - val_acc: 0.1108\n\nI'll upgrade keras and rerun\n\nThanks for all the help. I hope my info helps anybody else trying to reproduce the results.",
      "votes": null
    },
    {
      "id": "273921",
      "postDate": "01/25/2018 13:59:26",
      "content": "<p>The discrepancy in Sony-NEX-7 is b/c in that folder the original files are .JPG and should be renamed to .jpg.</p>\n\n<p>Re: versions </p>\n\n<pre><code>$ python -c 'import tensorflow as tf;import keras;print(tf.__version__, keras.__version__)'\n/home/antor/miniconda3/lib/python3.6/importlib/_bootstrap.py:219: RuntimeWarning: compiletime version 3.5 of module 'tensorflow.python.framework.fast_tensor_util' does not match runtime version 3.6\n  return f(*args, **kwds)\nUsing TensorFlow backend.\n1.4.1 2.1.3\n</code></pre>\n\n<p>Re: batch size you should be able to fit more in a 1080 Ti. I am testing the code on a different machine (AMD Ryzen 1600X + 1070 Ti) and this works w/o issue <code>$ python train.py -g 1 -b 8 -cs 512 -cm ResNet50 -x -l 1e-4</code>. Check with <code>nvidia-smi</code> that no memory is being used by the GPU before launching the script.</p>",
      "rawMarkdown": "The discrepancy in Sony-NEX-7 is b/c in that folder the original files are .JPG and should be renamed to .jpg.\n\nRe: versions \n\n    $ python -c 'import tensorflow as tf;import keras;print(tf.__version__, keras.__version__)'\n    /home/antor/miniconda3/lib/python3.6/importlib/_bootstrap.py:219: RuntimeWarning: compiletime version 3.5 of module 'tensorflow.python.framework.fast_tensor_util' does not match runtime version 3.6\n      return f(*args, **kwds)\n    Using TensorFlow backend.\n    1.4.1 2.1.3\n\nRe: batch size you should be able to fit more in a 1080 Ti. I am testing the code on a different machine (AMD Ryzen 1600X + 1070 Ti) and this works w/o issue `$ python train.py -g 1 -b 8 -cs 512 -cm ResNet50 -x -l 1e-4`. Check with `nvidia-smi` that no memory is being used by the GPU before launching the script.",
      "votes": null
    },
    {
      "id": "273957",
      "postDate": "01/25/2018 15:22:48",
      "content": "<p>As per your prior suggestion I already renamed the files. I think my downloads are missing some though:</p>\n\n<pre><code>vdir flickr_images/sony_nex7/|wc\n    558    6129   74089\n</code></pre>\n\n<p>Now that I have the same versions: 1.4.1 2.1.3 I'll try again.\nI have used your script before with -b 8, I just haven't been able since yesterday. I am using the card for my display, so about .5GB are used, but not enough to keep me from using -b 8, there has to be something on my end I have messed up.</p>",
      "rawMarkdown": "As per your prior suggestion I already renamed the files. I think my downloads are missing some though:\n\n    vdir flickr_images/sony_nex7/|wc\n        558    6129   74089\n\nNow that I have the same versions: 1.4.1 2.1.3 I'll try again.\nI have used your script before with -b 8, I just haven't been able since yesterday. I am using the card for my display, so about .5GB are used, but not enough to keep me from using -b 8, there has to be something on my end I have messed up.",
      "votes": null
    },
    {
      "id": "273980",
      "postDate": "01/25/2018 16:05:32",
      "content": "<pre><code>$ vdir flickr_images/sony_nex7/|wc\n   1302   11711   98848\n</code></pre>\n\n<p>I did <code>find . -name \"urls_*\" -execdir wget -nc --tries=10 -i  {} \\;</code> on the <code>flick_images</code> directory b/c sometimes wget had to retry.</p>",
      "rawMarkdown": "$ vdir flickr_images/sony_nex7/|wc\n       1302   11711   98848\n\nI did `find . -name \"urls_*\" -execdir wget -nc --tries=10 -i  {} \\;` on the `flick_images` directory b/c sometimes wget had to retry.",
      "votes": null
    },
    {
      "id": "274082",
      "postDate": "01/25/2018 21:29:58",
      "content": "<p>Update: LB 0.959 using single DenseNet201. Further discussion moved to: <a href=\"https://www.kaggle.com/c/sp-society-camera-model-identification/discussion/48293\">https://www.kaggle.com/c/sp-society-camera-model-identification/discussion/48293</a></p>",
      "rawMarkdown": "Update: LB 0.959 using single DenseNet201. Further discussion moved to: https://www.kaggle.com/c/sp-society-camera-model-identification/discussion/48293",
      "votes": null
    },
    {
      "id": "274108",
      "postDate": "01/25/2018 22:08:32",
      "content": "<p>Very Good progress. Many thanks for sharing with us. I have one question: Is it very important to use Glep data? Because I did not notice much difference. </p>",
      "rawMarkdown": "Very Good progress. Many thanks for sharing with us. I have one question: Is it very important to use Glep data? Because I did not notice much difference.",
      "votes": null
    },
    {
      "id": "274111",
      "postDate": "01/25/2018 22:24:24",
      "content": "<p>Good question. I will test without it.</p>",
      "rawMarkdown": "Good question. I will test without it.",
      "votes": null
    },
    {
      "id": "274117",
      "postDate": "01/25/2018 22:37:27",
      "content": "<pre><code>Traceback (most recent call last):\n  File \"/home/ppleskov/anaconda3/lib/python3.6/threading.py\", line 916, in _bootstrap_inner\n    self.run()\n  File \"/home/ppleskov/anaconda3/lib/python3.6/threading.py\", line 864, in run\n    self._target(*self._args, **self._kwargs)\n  File \"/home/ppleskov/anaconda3/lib/python3.6/multiprocessing/pool.py\", line 429, in _handle_results\n    task = get()\n  File \"/home/ppleskov/anaconda3/lib/python3.6/multiprocessing/connection.py\", line 251, in recv\n    return _ForkingPickler.loads(buf.getbuffer())\nTypeError: __init__() missing 1 required positional argument: 'code'\n</code></pre>\n\n<p>anyone got this issue?</p>",
      "rawMarkdown": "Traceback (most recent call last):\n      File \"/home/ppleskov/anaconda3/lib/python3.6/threading.py\", line 916, in _bootstrap_inner\n        self.run()\n      File \"/home/ppleskov/anaconda3/lib/python3.6/threading.py\", line 864, in run\n        self._target(*self._args, **self._kwargs)\n      File \"/home/ppleskov/anaconda3/lib/python3.6/multiprocessing/pool.py\", line 429, in _handle_results\n        task = get()\n      File \"/home/ppleskov/anaconda3/lib/python3.6/multiprocessing/connection.py\", line 251, in recv\n        return _ForkingPickler.loads(buf.getbuffer())\n    TypeError: __init__() missing 1 required positional argument: 'code'\n\nanyone got this issue?",
      "votes": null
    },
    {
      "id": "274124",
      "postDate": "01/25/2018 23:10:50",
      "content": "<p>I think there's an exception in the <code>process_item</code> function and it's not easy to see the exception from the main thread. Use <code>--verbose</code> and put some prints around <code>img = load_img_fast_jpg(item)</code>. I bet some images fail to load.</p>",
      "rawMarkdown": "I think there's an exception in the `process_item` function and it's not easy to see the exception from the main thread. Use `--verbose` and put some prints around `img = load_img_fast_jpg(item)`. I bet some images fail to load.",
      "votes": null
    },
    {
      "id": "274127",
      "postDate": "01/25/2018 23:11:49",
      "content": "<p>I'd like know of an easier way to catch/see exceptions when using <code>Pool</code></p>",
      "rawMarkdown": "I'd like know of an easier way to catch/see exceptions when using `Pool`",
      "votes": null
    },
    {
      "id": "274145",
      "postDate": "01/25/2018 23:42:56",
      "content": "<p>Yes, it's like Andres said above, the jpeg4py library can only load jpegs but in Gleb's dataset there are PNGs with the jpg extension.</p>\n\n<p>You could remove those files or use something like this:</p>\n\n<pre><code>def load_img_fast_jpg(img_path):\ntry:\n    x = jpeg.JPEG(img_path).decode()\n    return x\nexcept:\n    return load_img(img_path)\n</code></pre>",
      "rawMarkdown": "Yes, it's like Andres said above, the jpeg4py library can only load jpegs but in Gleb's dataset there are PNGs with the jpg extension.\n\nYou could remove those files or use something like this:\n\n    def load_img_fast_jpg(img_path):\n    try:\n        x = jpeg.JPEG(img_path).decode()\n        return x\n    except:\n        return load_img(img_path)",
      "votes": null
    },
    {
      "id": "274146",
      "postDate": "01/25/2018 23:46:37",
      "content": "<p>thanks a lot, that helped</p>",
      "rawMarkdown": "thanks a lot, that helped",
      "votes": null
    },
    {
      "id": "274267",
      "postDate": "01/26/2018 07:42:54",
      "content": "<p>For me its stuck at this stage without any explicit information:</p>\n\n<pre><code>              HTC-1-M7:  1023 (13.0%)\n              iPhone-6:   823 (10.5%)\n              Motorola-Droid-Maxx:   825 (10.5%)\n              Motorola-X:   275 (03.5%)\n             Samsung-Galaxy-S4:  1412 (18.0%)\n             iPhone-4s:   774 (09.9%)\n             LG-Nexus-5x:   680 (08.7%)\n             Motorola-Nexus-6:   926 (11.8%)\n             Samsung-Galaxy-Note3:   548 (07.0%)\n             Sony-NEX-7:   557 (07.1%)\n             validation steps = 0\n             Epoch 1/200\n</code></pre>\n\n<p>No particular output, nothing, even with the <code>-v</code> flag.  Notice the <code>validation_steps=0</code> is the value that the <code>validation_steps</code> in the generator was getting without modifying the code, I had to hardcode that into sth like 6 otherwise I was getting errors. Any suggestions?</p>",
      "rawMarkdown": "For me its stuck at this stage without any explicit information:\n\n                  HTC-1-M7:  1023 (13.0%)\n                  iPhone-6:   823 (10.5%)\n                  Motorola-Droid-Maxx:   825 (10.5%)\n                  Motorola-X:   275 (03.5%)\n                 Samsung-Galaxy-S4:  1412 (18.0%)\n                 iPhone-4s:   774 (09.9%)\n                 LG-Nexus-5x:   680 (08.7%)\n                 Motorola-Nexus-6:   926 (11.8%)\n                 Samsung-Galaxy-Note3:   548 (07.0%)\n                 Sony-NEX-7:   557 (07.1%)\n                 validation steps = 0\n                 Epoch 1/200\nNo particular output, nothing, even with the `-v` flag.  Notice the `validation_steps=0` is the value that the `validation_steps` in the generator was getting without modifying the code, I had to hardcode that into sth like 6 otherwise I was getting errors. Any suggestions?",
      "votes": null
    },
    {
      "id": "274396",
      "postDate": "01/26/2018 14:37:11",
      "content": "<p>ValueError: <code>validation_steps=None</code> is only valid for a generator based on the <code>keras.utils.Sequence</code> class. Please specify <code>validation_steps</code> or use the <code>keras.utils.Sequence</code> class.</p>\n\n<p>I also have some problem, when i look back to the code, i found it already used <code>keras.utils.Sequence</code> class, is there any one know how to reslove this? </p>\n\n<p>Thanks</p>",
      "rawMarkdown": "ValueError: `validation_steps=None` is only valid for a generator based on the `keras.utils.Sequence` class. Please specify `validation_steps` or use the `keras.utils.Sequence` class.\n\nI also have some problem, when i look back to the code, i found it already used `keras.utils.Sequence` class, is there any one know how to reslove this? \n\nThanks",
      "votes": null
    },
    {
      "id": "274420",
      "postDate": "01/26/2018 15:25:43",
      "content": "<p>I just hardcoded a value. For instance put a value of 6 and see if it works.</p>",
      "rawMarkdown": "I just hardcoded a value. For instance put a value of 6 and see if it works.",
      "votes": null
    },
    {
      "id": "274446",
      "postDate": "01/26/2018 16:06:50",
      "content": "<p>validation steps should be: number of validations images / batch size</p>\n\n<p>I added the following code to help me debug the image lists:</p>\n\n<pre><code>    print(\"\\nImage set counts\\n\")\nif args.extra_dataset:\n    print('{:&gt;15}: {:5d}'.format(\"extra_train_ids\",len(extra_train_ids)))\n    print('{:&gt;15}: {:5d}'.format(\"extra_val_ids\",len(extra_val_ids)))\n    print(\"\\n\")\nprint('{:&gt;15}: {:5d} steps={:5d}'.format(\n    \"ids_train\",len(ids_train),int(math.ceil(len(ids_train)  // args.batch_size))))\nprint('{:&gt;15}: {:5d} steps={:5d}'.format(\n    \"ids_val\",len(ids_val),int(math.ceil(len(ids_val) // args.batch_size))))\nprint('_' * 100)\nprint(\"\\n\")\n</code></pre>\n\n<p>(pasting code doesn't format well, pay attention to indentation)</p>",
      "rawMarkdown": "validation steps should be: number of validations images / batch size\n\nI added the following code to help me debug the image lists:\n\n        print(\"\\nImage set counts\\n\")\n    if args.extra_dataset:\n        print('{:&gt;15}: {:5d}'.format(\"extra_train_ids\",len(extra_train_ids)))\n        print('{:&gt;15}: {:5d}'.format(\"extra_val_ids\",len(extra_val_ids)))\n        print(\"\\n\")\n    print('{:&gt;15}: {:5d} steps={:5d}'.format(\n        \"ids_train\",len(ids_train),int(math.ceil(len(ids_train)  // args.batch_size))))\n    print('{:&gt;15}: {:5d} steps={:5d}'.format(\n        \"ids_val\",len(ids_val),int(math.ceil(len(ids_val) // args.batch_size))))\n    print('_' * 100)\n    print(\"\\n\")\n\n(pasting code doesn't format well, pay attention to indentation)",
      "votes": null
    },
    {
      "id": "274452",
      "postDate": "01/26/2018 16:19:07",
      "content": "<p>Andres, trying to upgrade tensorflow got my system unstable. I know you version of libraries, since we both use 1080ti can I ask you the version of the OS and the nvidia drivers?</p>\n\n<p>Now trying to run your script I only get:</p>\n\n<pre><code>E tensorflow/stream_executor/cuda/cuda_dnn.cc:385] could not create cudnn handle: CUDNN_STATUS_INTERNAL_ERROR\n</code></pre>\n\n<p>I may end up having to rebuild the system.</p>\n\n<p>Thanks in advance</p>",
      "rawMarkdown": "Andres, trying to upgrade tensorflow got my system unstable. I know you version of libraries, since we both use 1080ti can I ask you the version of the OS and the nvidia drivers?\n\nNow trying to run your script I only get:\n\n    E tensorflow/stream_executor/cuda/cuda_dnn.cc:385] could not create cudnn handle: CUDNN_STATUS_INTERNAL_ERROR\n\nI may end up having to rebuild the system.\n\nThanks in advance",
      "votes": null
    },
    {
      "id": "274460",
      "postDate": "01/26/2018 16:35:01",
      "content": "<p>May I ask if you have downloaded all the images? For me is the case I believe because I haven't downloaded all the images?</p>",
      "rawMarkdown": "May I ask if you have downloaded all the images? For me is the case I believe because I haven't downloaded all the images?",
      "votes": null
    },
    {
      "id": "274479",
      "postDate": "01/26/2018 17:16:39",
      "content": "<p>Yes, I downloaded all the images. It took me a couple of tries and had to make sure my paths were setup the way that the script expects them. You can change things around by changing the lines:</p>\n\n<pre><code>    extra_train_ids = [os.path.join(EXTRA_TRAIN_FOLDER,line.rstrip('\\n')) for line in open(os.path.join(EXTRA_TRAIN_FOLDER, 'good_jpgs'))]\n    extra_train_ids.sort()\n    ids_train.extend(extra_train_ids)\n\n    extra_val_ids = glob.glob(join(EXTRA_VAL_FOLDER,'*/*.jpg'))\n    extra_val_ids.sort()\n    ids_val.extend(extra_val_ids)\n</code></pre>\n\n<p>Then you can setup the extra images anywhere you want.</p>\n\n<p>These are my counts:</p>\n\n<p>Image set counts</p>\n\n<p>extra_train_ids:  6119\n  extra_val_ids:   316</p>\n\n<pre><code>  ids_train:  8594 steps= 8594\n    ids_val:   316 steps=  316\n</code></pre>\n\n<hr>\n\n<pre><code>          HTC-1-M7:  1024 (11.9%)\n          iPhone-6:   825 (09.6%)\n</code></pre>\n\n<p>Motorola-Droid-Maxx:   825 (09.6%)\n            Motorola-X:   275 (03.2%)\n     Samsung-Galaxy-S4:  1412 (16.4%)\n             iPhone-4s:   779 (09.1%)\n           LG-Nexus-5x:   680 (07.9%)\n      Motorola-Nexus-6:   926 (10.8%)\n  Samsung-Galaxy-Note3:   549 (06.4%)\n            Sony-NEX-7:  1299 (15.1%)</p>",
      "rawMarkdown": "Yes, I downloaded all the images. It took me a couple of tries and had to make sure my paths were setup the way that the script expects them. You can change things around by changing the lines:\n\n\n        extra_train_ids = [os.path.join(EXTRA_TRAIN_FOLDER,line.rstrip('\\n')) for line in open(os.path.join(EXTRA_TRAIN_FOLDER, 'good_jpgs'))]\n        extra_train_ids.sort()\n        ids_train.extend(extra_train_ids)\n\n        extra_val_ids = glob.glob(join(EXTRA_VAL_FOLDER,'*/*.jpg'))\n        extra_val_ids.sort()\n        ids_val.extend(extra_val_ids)\n\nThen you can setup the extra images anywhere you want.\n\nThese are my counts:\n\nImage set counts\n\nextra_train_ids:  6119\n  extra_val_ids:   316\n\n      ids_train:  8594 steps= 8594\n        ids_val:   316 steps=  316\n____________________________________________________________________________________________________\n\n              HTC-1-M7:  1024 (11.9%)\n              iPhone-6:   825 (09.6%)\n   Motorola-Droid-Maxx:   825 (09.6%)\n            Motorola-X:   275 (03.2%)\n     Samsung-Galaxy-S4:  1412 (16.4%)\n             iPhone-4s:   779 (09.1%)\n           LG-Nexus-5x:   680 (07.9%)\n      Motorola-Nexus-6:   926 (10.8%)\n  Samsung-Galaxy-Note3:   549 (06.4%)\n            Sony-NEX-7:  1299 (15.1%)",
      "votes": null
    },
    {
      "id": "274534",
      "postDate": "01/26/2018 18:54:25",
      "content": "<p>Here you go:</p>\n\n<pre><code>$ nvidia-smi\nFri Jan 26 19:53:31 2018\n+-----------------------------------------------------------------------------+\n| NVIDIA-SMI 390.12                 Driver Version: 390.12                    |\n|-------------------------------+----------------------+----------------------+\n| GPU  Name        Persistence-M| Bus-Id        Disp.A | Volatile Uncorr. ECC |\n| Fan  Temp  Perf  Pwr:Usage/Cap|         Memory-Usage | GPU-Util  Compute M. |\n|===============================+======================+======================|\n|   0  GeForce GTX 108...  Off  | 00000000:09:00.0 Off |                  N/A |\n| 46%   67C    P2   252W / 280W |  10801MiB / 11178MiB |     96%      Default |\n+-------------------------------+----------------------+----------------------+\n|   1  GeForce GTX 108...  Off  | 00000000:41:00.0  On |                  N/A |\n| 54%   69C    P2   197W / 280W |  10821MiB / 11170MiB |     48%      Default |\n+-------------------------------+----------------------+----------------------+\n\n+-----------------------------------------------------------------------------+\n| Processes:                                                       GPU Memory |\n|  GPU       PID   Type   Process name                             Usage      |\n|=============================================================================|\n|    0     51088      C   python                                     10789MiB |\n|    1      1830      G   /usr/lib/xorg/Xorg                           336MiB |\n|    1      2283      G   compiz                                       249MiB |\n|    1     51088      C   python                                     10223MiB |\n+-----------------------------------------------------------------------------+\n$ uname -a\nLinux cineubuntu 4.13.0-31-generic #34~16.04.1-Ubuntu SMP Fri Jan 19 17:11:01 UTC 2018 x86_64 x86_64 x86_64 GNU/Linux\n</code></pre>\n\n<p>$ sudo apt list --installed | egrep \"(nvidia|cuda)\"</p>\n\n<p>WARNING: apt does not have a stable CLI interface. Use with caution in scripts.</p>\n\n<p>cuda/unknown,now 9.1.85-1 amd64 [installed]\ncuda-8-0/unknown,unknown,now 8.0.61-1 amd64 [installed,auto-removable]\ncuda-9-1/unknown,now 9.1.85-1 amd64 [installed,automatic]\ncuda-command-line-tools-8-0/unknown,unknown,now 8.0.61-1 amd64 [installed,auto-removable]\ncuda-command-line-tools-9-1/unknown,now 9.1.85-1 amd64 [installed,automatic]\ncuda-compiler-9-1/unknown,now 9.1.85-1 amd64 [installed,automatic]\ncuda-core-8-0/unknown,unknown,now 8.0.61-1 amd64 [installed,auto-removable]\ncuda-cublas-8-0/unknown,now 8.0.61.2-1 amd64 [installed,auto-removable]\ncuda-cublas-9-1/unknown,now 9.1.85-1 amd64 [installed,upgradable to: 9.1.85.1-1]\ncuda-cublas-dev-8-0/unknown,now 8.0.61.2-1 amd64 [installed,auto-removable]\ncuda-cublas-dev-9-1/unknown,now 9.1.85-1 amd64 [installed,upgradable to: 9.1.85.1-1]\ncuda-cudart-8-0/unknown,unknown,now 8.0.61-1 amd64 [installed,auto-removable]\ncuda-cudart-9-1/unknown,now 9.1.85-1 amd64 [installed,automatic]\ncuda-cudart-dev-8-0/unknown,unknown,now 8.0.61-1 amd64 [installed,auto-removable]\ncuda-cudart-dev-9-1/unknown,now 9.1.85-1 amd64 [installed,automatic]\ncuda-cufft-8-0/unknown,unknown,now 8.0.61-1 amd64 [installed,auto-removable]\ncuda-cufft-9-1/unknown,now 9.1.85-1 amd64 [installed,automatic]\ncuda-cufft-dev-8-0/unknown,unknown,now 8.0.61-1 amd64 [installed,auto-removable]\ncuda-cufft-dev-9-1/unknown,now 9.1.85-1 amd64 [installed,automatic]\ncuda-cuobjdump-9-1/unknown,now 9.1.85-1 amd64 [installed,automatic]\ncuda-cupti-9-1/unknown,now 9.1.85-1 amd64 [installed,automatic]\ncuda-curand-8-0/unknown,unknown,now 8.0.61-1 amd64 [installed,auto-removable]\ncuda-curand-9-1/unknown,now 9.1.85-1 amd64 [installed,automatic]\ncuda-curand-dev-8-0/unknown,unknown,now 8.0.61-1 amd64 [installed,auto-removable]\ncuda-curand-dev-9-1/unknown,now 9.1.85-1 amd64 [installed,automatic]\ncuda-cusolver-8-0/unknown,unknown,now 8.0.61-1 amd64 [installed,auto-removable]\ncuda-cusolver-9-1/unknown,now 9.1.85-1 amd64 [installed,automatic]\ncuda-cusolver-dev-8-0/unknown,unknown,now 8.0.61-1 amd64 [installed,auto-removable]\ncuda-cusolver-dev-9-1/unknown,now 9.1.85-1 amd64 [installed,automatic]\ncuda-cusparse-8-0/unknown,unknown,now 8.0.61-1 amd64 [installed,auto-removable]\ncuda-cusparse-9-1/unknown,now 9.1.85-1 amd64 [installed,automatic]\ncuda-cusparse-dev-8-0/unknown,unknown,now 8.0.61-1 amd64 [installed,auto-removable]\ncuda-cusparse-dev-9-1/unknown,now 9.1.85-1 amd64 [installed,automatic]\ncuda-demo-suite-8-0/unknown,unknown,now 8.0.61-1 amd64 [installed,auto-removable]\ncuda-demo-suite-9-1/unknown,now 9.1.85-1 amd64 [installed,automatic]\ncuda-documentation-8-0/unknown,unknown,now 8.0.61-1 amd64 [installed,auto-removable]\ncuda-documentation-9-1/unknown,now 9.1.85-1 amd64 [installed,automatic]\ncuda-driver-dev-8-0/unknown,unknown,now 8.0.61-1 amd64 [installed,auto-removable]\ncuda-driver-dev-9-1/unknown,now 9.1.85-1 amd64 [installed,automatic]\ncuda-drivers/unknown,now 390.12-1 amd64 [installed,automatic]\ncuda-gdb-9-1/unknown,now 9.1.85-1 amd64 [installed,automatic]\ncuda-gpu-library-advisor-9-1/unknown,now 9.1.85-1 amd64 [installed,automatic]\ncuda-libraries-9-1/unknown,now 9.1.85-1 amd64 [installed,automatic]\ncuda-libraries-dev-9-1/unknown,now 9.1.85-1 amd64 [installed,automatic]\ncuda-license-8-0/unknown,unknown,now 8.0.61-1 amd64 [installed,auto-removable]\ncuda-license-9-1/unknown,now 9.1.85-1 amd64 [installed,automatic]\ncuda-memcheck-9-1/unknown,now 9.1.85-1 amd64 [installed,automatic]\ncuda-misc-headers-8-0/unknown,unknown,now 8.0.61-1 amd64 [installed,auto-removable]\ncuda-misc-headers-9-1/unknown,now 9.1.85-1 amd64 [installed,automatic]\ncuda-npp-8-0/unknown,unknown,now 8.0.61-1 amd64 [installed,auto-removable]\ncuda-npp-9-1/unknown,now 9.1.85-1 amd64 [installed,automatic]\ncuda-npp-dev-8-0/unknown,unknown,now 8.0.61-1 amd64 [installed,auto-removable]\ncuda-npp-dev-9-1/unknown,now 9.1.85-1 amd64 [installed,automatic]\ncuda-nsight-9-1/unknown,now 9.1.85-1 amd64 [installed,automatic]\ncuda-nvcc-9-1/unknown,now 9.1.85-1 amd64 [installed,automatic]\ncuda-nvdisasm-9-1/unknown,now 9.1.85-1 amd64 [installed,automatic]\ncuda-nvgraph-8-0/unknown,unknown,now 8.0.61-1 amd64 [installed,auto-removable]\ncuda-nvgraph-9-1/unknown,now 9.1.85-1 amd64 [installed,automatic]\ncuda-nvgraph-dev-8-0/unknown,unknown,now 8.0.61-1 amd64 [installed,auto-removable]\ncuda-nvgraph-dev-9-1/unknown,now 9.1.85-1 amd64 [installed,automatic]\ncuda-nvml-dev-8-0/unknown,unknown,now 8.0.61-1 amd64 [installed,auto-removable]\ncuda-nvml-dev-9-1/unknown,now 9.1.85-1 amd64 [installed,automatic]\ncuda-nvprof-9-1/unknown,now 9.1.85-1 amd64 [installed,automatic]\ncuda-nvprune-9-1/unknown,now 9.1.85-1 amd64 [installed,automatic]\ncuda-nvrtc-8-0/unknown,unknown,now 8.0.61-1 amd64 [installed,auto-removable]\ncuda-nvrtc-9-1/unknown,now 9.1.85-1 amd64 [installed,automatic]\ncuda-nvrtc-dev-8-0/unknown,unknown,now 8.0.61-1 amd64 [installed,auto-removable]\ncuda-nvrtc-dev-9-1/unknown,now 9.1.85-1 amd64 [installed,automatic]\ncuda-nvtx-9-1/unknown,now 9.1.85-1 amd64 [installed,automatic]\ncuda-nvvp-9-1/unknown,now 9.1.85-1 amd64 [installed,automatic]\ncuda-repo-ubuntu1604/unknown,now 9.1.85-1 amd64 [installed]\ncuda-repo-ubuntu1604-8-0-local-ga2/now 8.0.61-1 amd64 [installed,local]\ncuda-runtime-8-0/unknown,unknown,now 8.0.61-1 amd64 [installed,auto-removable]\ncuda-runtime-9-1/unknown,now 9.1.85-1 amd64 [installed,automatic]\ncuda-samples-8-0/unknown,unknown,now 8.0.61-1 amd64 [installed,auto-removable]\ncuda-samples-9-1/unknown,now 9.1.85-1 amd64 [installed,automatic]\ncuda-toolkit-8-0/unknown,unknown,now 8.0.61-1 amd64 [installed,auto-removable]\ncuda-toolkit-9-1/unknown,now 9.1.85-1 amd64 [installed,automatic]\ncuda-tools-9-1/unknown,now 9.1.85-1 amd64 [installed,automatic]\ncuda-visual-tools-8-0/unknown,unknown,now 8.0.61-1 amd64 [installed,auto-removable]\ncuda-visual-tools-9-1/unknown,now 9.1.85-1 amd64 [installed,automatic]\nlibcuda1-375/unknown,now 390.12-0ubuntu1 amd64 [installed]\nlibcuda1-390/unknown,now 390.12-0ubuntu1 amd64 [installed,automatic]\nlibcudart7.5/xenial,now 7.5.18-0ubuntu1 amd64 [installed,automatic]\nnvidia-375/unknown,now 390.12-0ubuntu1 amd64 [installed,auto-removable]\nnvidia-375-dev/unknown,now 390.12-0ubuntu1 amd64 [installed]\nnvidia-390/unknown,now 390.12-0ubuntu1 amd64 [installed,automatic]\nnvidia-390-dev/unknown,now 390.12-0ubuntu1 amd64 [installed,automatic]\nnvidia-cuda-dev/xenial,now 7.5.18-0ubuntu1 amd64 [installed,automatic]\nnvidia-cuda-doc/xenial,xenial,now 7.5.18-0ubuntu1 all [installed,automatic]\nnvidia-cuda-gdb/xenial,now 7.5.18-0ubuntu1 amd64 [installed,automatic]\nnvidia-cuda-toolkit/xenial,now 7.5.18-0ubuntu1 amd64 [installed]\nnvidia-modprobe/unknown,now 390.12-0ubuntu1 amd64 [installed,automatic]\nnvidia-opencl-dev/xenial,now 7.5.18-0ubuntu1 amd64 [installed,automatic]\nnvidia-opencl-icd-375/unknown,now 390.12-0ubuntu1 amd64 [installed,auto-removable]\nnvidia-opencl-icd-390/unknown,now 390.12-0ubuntu1 amd64 [installed,automatic]\nnvidia-prime/xenial,now 0.8.2 amd64 [installed]\nnvidia-profiler/xenial,now 7.5.18-0ubuntu1 amd64 [installed,automatic]\nnvidia-settings/unknown,now 390.12-0ubuntu1 amd64 [installed,automatic]\nnvidia-visual-profiler/xenial,now 7.5.18-0ubuntu1 amd64 [installed,automatic]</p>",
      "rawMarkdown": "Here you go:\n\n    $ nvidia-smi\n    Fri Jan 26 19:53:31 2018\n    +-----------------------------------------------------------------------------+\n    | NVIDIA-SMI 390.12                 Driver Version: 390.12                    |\n    |-------------------------------+----------------------+----------------------+\n    | GPU  Name        Persistence-M| Bus-Id        Disp.A | Volatile Uncorr. ECC |\n    | Fan  Temp  Perf  Pwr:Usage/Cap|         Memory-Usage | GPU-Util  Compute M. |\n    |===============================+======================+======================|\n    |   0  GeForce GTX 108...  Off  | 00000000:09:00.0 Off |                  N/A |\n    | 46%   67C    P2   252W / 280W |  10801MiB / 11178MiB |     96%      Default |\n    +-------------------------------+----------------------+----------------------+\n    |   1  GeForce GTX 108...  Off  | 00000000:41:00.0  On |                  N/A |\n    | 54%   69C    P2   197W / 280W |  10821MiB / 11170MiB |     48%      Default |\n    +-------------------------------+----------------------+----------------------+\n    \n    +-----------------------------------------------------------------------------+\n    | Processes:                                                       GPU Memory |\n    |  GPU       PID   Type   Process name                             Usage      |\n    |=============================================================================|\n    |    0     51088      C   python                                     10789MiB |\n    |    1      1830      G   /usr/lib/xorg/Xorg                           336MiB |\n    |    1      2283      G   compiz                                       249MiB |\n    |    1     51088      C   python                                     10223MiB |\n    +-----------------------------------------------------------------------------+\n    $ uname -a\n    Linux cineubuntu 4.13.0-31-generic #34~16.04.1-Ubuntu SMP Fri Jan 19 17:11:01 UTC 2018 x86_64 x86_64 x86_64 GNU/Linux\n$ sudo apt list --installed | egrep \"(nvidia|cuda)\"\n\nWARNING: apt does not have a stable CLI interface. Use with caution in scripts.\n\ncuda/unknown,now 9.1.85-1 amd64 [installed]\ncuda-8-0/unknown,unknown,now 8.0.61-1 amd64 [installed,auto-removable]\ncuda-9-1/unknown,now 9.1.85-1 amd64 [installed,automatic]\ncuda-command-line-tools-8-0/unknown,unknown,now 8.0.61-1 amd64 [installed,auto-removable]\ncuda-command-line-tools-9-1/unknown,now 9.1.85-1 amd64 [installed,automatic]\ncuda-compiler-9-1/unknown,now 9.1.85-1 amd64 [installed,automatic]\ncuda-core-8-0/unknown,unknown,now 8.0.61-1 amd64 [installed,auto-removable]\ncuda-cublas-8-0/unknown,now 8.0.61.2-1 amd64 [installed,auto-removable]\ncuda-cublas-9-1/unknown,now 9.1.85-1 amd64 [installed,upgradable to: 9.1.85.1-1]\ncuda-cublas-dev-8-0/unknown,now 8.0.61.2-1 amd64 [installed,auto-removable]\ncuda-cublas-dev-9-1/unknown,now 9.1.85-1 amd64 [installed,upgradable to: 9.1.85.1-1]\ncuda-cudart-8-0/unknown,unknown,now 8.0.61-1 amd64 [installed,auto-removable]\ncuda-cudart-9-1/unknown,now 9.1.85-1 amd64 [installed,automatic]\ncuda-cudart-dev-8-0/unknown,unknown,now 8.0.61-1 amd64 [installed,auto-removable]\ncuda-cudart-dev-9-1/unknown,now 9.1.85-1 amd64 [installed,automatic]\ncuda-cufft-8-0/unknown,unknown,now 8.0.61-1 amd64 [installed,auto-removable]\ncuda-cufft-9-1/unknown,now 9.1.85-1 amd64 [installed,automatic]\ncuda-cufft-dev-8-0/unknown,unknown,now 8.0.61-1 amd64 [installed,auto-removable]\ncuda-cufft-dev-9-1/unknown,now 9.1.85-1 amd64 [installed,automatic]\ncuda-cuobjdump-9-1/unknown,now 9.1.85-1 amd64 [installed,automatic]\ncuda-cupti-9-1/unknown,now 9.1.85-1 amd64 [installed,automatic]\ncuda-curand-8-0/unknown,unknown,now 8.0.61-1 amd64 [installed,auto-removable]\ncuda-curand-9-1/unknown,now 9.1.85-1 amd64 [installed,automatic]\ncuda-curand-dev-8-0/unknown,unknown,now 8.0.61-1 amd64 [installed,auto-removable]\ncuda-curand-dev-9-1/unknown,now 9.1.85-1 amd64 [installed,automatic]\ncuda-cusolver-8-0/unknown,unknown,now 8.0.61-1 amd64 [installed,auto-removable]\ncuda-cusolver-9-1/unknown,now 9.1.85-1 amd64 [installed,automatic]\ncuda-cusolver-dev-8-0/unknown,unknown,now 8.0.61-1 amd64 [installed,auto-removable]\ncuda-cusolver-dev-9-1/unknown,now 9.1.85-1 amd64 [installed,automatic]\ncuda-cusparse-8-0/unknown,unknown,now 8.0.61-1 amd64 [installed,auto-removable]\ncuda-cusparse-9-1/unknown,now 9.1.85-1 amd64 [installed,automatic]\ncuda-cusparse-dev-8-0/unknown,unknown,now 8.0.61-1 amd64 [installed,auto-removable]\ncuda-cusparse-dev-9-1/unknown,now 9.1.85-1 amd64 [installed,automatic]\ncuda-demo-suite-8-0/unknown,unknown,now 8.0.61-1 amd64 [installed,auto-removable]\ncuda-demo-suite-9-1/unknown,now 9.1.85-1 amd64 [installed,automatic]\ncuda-documentation-8-0/unknown,unknown,now 8.0.61-1 amd64 [installed,auto-removable]\ncuda-documentation-9-1/unknown,now 9.1.85-1 amd64 [installed,automatic]\ncuda-driver-dev-8-0/unknown,unknown,now 8.0.61-1 amd64 [installed,auto-removable]\ncuda-driver-dev-9-1/unknown,now 9.1.85-1 amd64 [installed,automatic]\ncuda-drivers/unknown,now 390.12-1 amd64 [installed,automatic]\ncuda-gdb-9-1/unknown,now 9.1.85-1 amd64 [installed,automatic]\ncuda-gpu-library-advisor-9-1/unknown,now 9.1.85-1 amd64 [installed,automatic]\ncuda-libraries-9-1/unknown,now 9.1.85-1 amd64 [installed,automatic]\ncuda-libraries-dev-9-1/unknown,now 9.1.85-1 amd64 [installed,automatic]\ncuda-license-8-0/unknown,unknown,now 8.0.61-1 amd64 [installed,auto-removable]\ncuda-license-9-1/unknown,now 9.1.85-1 amd64 [installed,automatic]\ncuda-memcheck-9-1/unknown,now 9.1.85-1 amd64 [installed,automatic]\ncuda-misc-headers-8-0/unknown,unknown,now 8.0.61-1 amd64 [installed,auto-removable]\ncuda-misc-headers-9-1/unknown,now 9.1.85-1 amd64 [installed,automatic]\ncuda-npp-8-0/unknown,unknown,now 8.0.61-1 amd64 [installed,auto-removable]\ncuda-npp-9-1/unknown,now 9.1.85-1 amd64 [installed,automatic]\ncuda-npp-dev-8-0/unknown,unknown,now 8.0.61-1 amd64 [installed,auto-removable]\ncuda-npp-dev-9-1/unknown,now 9.1.85-1 amd64 [installed,automatic]\ncuda-nsight-9-1/unknown,now 9.1.85-1 amd64 [installed,automatic]\ncuda-nvcc-9-1/unknown,now 9.1.85-1 amd64 [installed,automatic]\ncuda-nvdisasm-9-1/unknown,now 9.1.85-1 amd64 [installed,automatic]\ncuda-nvgraph-8-0/unknown,unknown,now 8.0.61-1 amd64 [installed,auto-removable]\ncuda-nvgraph-9-1/unknown,now 9.1.85-1 amd64 [installed,automatic]\ncuda-nvgraph-dev-8-0/unknown,unknown,now 8.0.61-1 amd64 [installed,auto-removable]\ncuda-nvgraph-dev-9-1/unknown,now 9.1.85-1 amd64 [installed,automatic]\ncuda-nvml-dev-8-0/unknown,unknown,now 8.0.61-1 amd64 [installed,auto-removable]\ncuda-nvml-dev-9-1/unknown,now 9.1.85-1 amd64 [installed,automatic]\ncuda-nvprof-9-1/unknown,now 9.1.85-1 amd64 [installed,automatic]\ncuda-nvprune-9-1/unknown,now 9.1.85-1 amd64 [installed,automatic]\ncuda-nvrtc-8-0/unknown,unknown,now 8.0.61-1 amd64 [installed,auto-removable]\ncuda-nvrtc-9-1/unknown,now 9.1.85-1 amd64 [installed,automatic]\ncuda-nvrtc-dev-8-0/unknown,unknown,now 8.0.61-1 amd64 [installed,auto-removable]\ncuda-nvrtc-dev-9-1/unknown,now 9.1.85-1 amd64 [installed,automatic]\ncuda-nvtx-9-1/unknown,now 9.1.85-1 amd64 [installed,automatic]\ncuda-nvvp-9-1/unknown,now 9.1.85-1 amd64 [installed,automatic]\ncuda-repo-ubuntu1604/unknown,now 9.1.85-1 amd64 [installed]\ncuda-repo-ubuntu1604-8-0-local-ga2/now 8.0.61-1 amd64 [installed,local]\ncuda-runtime-8-0/unknown,unknown,now 8.0.61-1 amd64 [installed,auto-removable]\ncuda-runtime-9-1/unknown,now 9.1.85-1 amd64 [installed,automatic]\ncuda-samples-8-0/unknown,unknown,now 8.0.61-1 amd64 [installed,auto-removable]\ncuda-samples-9-1/unknown,now 9.1.85-1 amd64 [installed,automatic]\ncuda-toolkit-8-0/unknown,unknown,now 8.0.61-1 amd64 [installed,auto-removable]\ncuda-toolkit-9-1/unknown,now 9.1.85-1 amd64 [installed,automatic]\ncuda-tools-9-1/unknown,now 9.1.85-1 amd64 [installed,automatic]\ncuda-visual-tools-8-0/unknown,unknown,now 8.0.61-1 amd64 [installed,auto-removable]\ncuda-visual-tools-9-1/unknown,now 9.1.85-1 amd64 [installed,automatic]\nlibcuda1-375/unknown,now 390.12-0ubuntu1 amd64 [installed]\nlibcuda1-390/unknown,now 390.12-0ubuntu1 amd64 [installed,automatic]\nlibcudart7.5/xenial,now 7.5.18-0ubuntu1 amd64 [installed,automatic]\nnvidia-375/unknown,now 390.12-0ubuntu1 amd64 [installed,auto-removable]\nnvidia-375-dev/unknown,now 390.12-0ubuntu1 amd64 [installed]\nnvidia-390/unknown,now 390.12-0ubuntu1 amd64 [installed,automatic]\nnvidia-390-dev/unknown,now 390.12-0ubuntu1 amd64 [installed,automatic]\nnvidia-cuda-dev/xenial,now 7.5.18-0ubuntu1 amd64 [installed,automatic]\nnvidia-cuda-doc/xenial,xenial,now 7.5.18-0ubuntu1 all [installed,automatic]\nnvidia-cuda-gdb/xenial,now 7.5.18-0ubuntu1 amd64 [installed,automatic]\nnvidia-cuda-toolkit/xenial,now 7.5.18-0ubuntu1 amd64 [installed]\nnvidia-modprobe/unknown,now 390.12-0ubuntu1 amd64 [installed,automatic]\nnvidia-opencl-dev/xenial,now 7.5.18-0ubuntu1 amd64 [installed,automatic]\nnvidia-opencl-icd-375/unknown,now 390.12-0ubuntu1 amd64 [installed,auto-removable]\nnvidia-opencl-icd-390/unknown,now 390.12-0ubuntu1 amd64 [installed,automatic]\nnvidia-prime/xenial,now 0.8.2 amd64 [installed]\nnvidia-profiler/xenial,now 7.5.18-0ubuntu1 amd64 [installed,automatic]\nnvidia-settings/unknown,now 390.12-0ubuntu1 amd64 [installed,automatic]\nnvidia-visual-profiler/xenial,now 7.5.18-0ubuntu1 amd64 [installed,automatic]",
      "votes": null
    },
    {
      "id": "274537",
      "postDate": "01/26/2018 19:04:18",
      "content": "<p>You can also consider TF optimized wheels:\n<a href=\"https://github.com/mind/wheels\">https://github.com/mind/wheels</a></p>",
      "rawMarkdown": "You can also consider TF optimized wheels:\nhttps://github.com/mind/wheels",
      "votes": null
    },
    {
      "id": "274557",
      "postDate": "01/26/2018 19:38:46",
      "content": "<p>Thanks I may have to try wheels. Going to a newer tensorflow then cuda9 then the nvidia-390 drivers is what got my system unstable. I haven't been able to run your scripts since then. Since you are running that I now know it is possible.</p>\n\n<p>Are you using one of the 1080s for your display? Using the 390 drivers I did notice a huge lag on X and that's when I started to downgrade back.</p>",
      "rawMarkdown": "Thanks I may have to try wheels. Going to a newer tensorflow then cuda9 then the nvidia-390 drivers is what got my system unstable. I haven't been able to run your scripts since then. Since you are running that I now know it is possible.\n\nAre you using one of the 1080s for your display? Using the 390 drivers I did notice a huge lag on X and that's when I started to downgrade back.",
      "votes": null
    },
    {
      "id": "274560",
      "postDate": "01/26/2018 19:44:44",
      "content": "<p>Although I have cuda9 installed I use cuda8 too, which I what Im using. I have that computer plugged to a monitor which pretends it's on, but I always work via ssh / tmux with that machine, so I didn't really use the video cards in Linux as display much.</p>",
      "rawMarkdown": "Although I have cuda9 installed I use cuda8 too, which I what Im using. I have that computer plugged to a monitor which pretends it's on, but I always work via ssh / tmux with that machine, so I didn't really use the video cards in Linux as display much.",
      "votes": null
    },
    {
      "id": "275128",
      "postDate": "01/28/2018 11:39:58",
      "content": "<p>Hey Andres, \nthanks for the updates to your code!</p>\n\n<p>I am now getting the following error message:</p>\n\n<pre><code>  File \"train.py\", line 284, in process_item\nimg = np.rot90(_img, 1, (0,1))\nTypeError: rot90() takes from 1 to 2 positional arguments but 3 were given\n</code></pre>\n\n<p>Should that be np.rot90(1,(0,1)?\nThanks</p>",
      "rawMarkdown": "Hey Andres, \nthanks for the updates to your code!\n\nI am now getting the following error message:\n\n      File \"train.py\", line 284, in process_item\n    img = np.rot90(_img, 1, (0,1))\n    TypeError: rot90() takes from 1 to 2 positional arguments but 3 were given\n\nShould that be np.rot90(1,(0,1)?\nThanks",
      "votes": null
    },
    {
      "id": "275135",
      "postDate": "01/28/2018 11:54:33",
      "content": "<p>What version of np are you running?\n<a href=\"https://docs.scipy.org/doc/numpy-1.13.0/reference/generated/numpy.rot90.html\">https://docs.scipy.org/doc/numpy-1.13.0/reference/generated/numpy.rot90.html</a></p>",
      "rawMarkdown": "What version of np are you running?\nhttps://docs.scipy.org/doc/numpy-1.13.0/reference/generated/numpy.rot90.html",
      "votes": null
    },
    {
      "id": "275147",
      "postDate": "01/28/2018 12:03:47",
      "content": "<p>Thanks for the quick reply. \nI was on 1.11, have upgraded to 1.14 and now it's working!</p>\n\n<p>Btw. I read you were using resnet50 quite a lot. Any success looking into other models?\nI guess even if other models don't quite reach resnet50's performance, they might be useful for ensembling.</p>",
      "rawMarkdown": "Thanks for the quick reply. \nI was on 1.11, have upgraded to 1.14 and now it's working!\n\nBtw. I read you were using resnet50 quite a lot. Any success looking into other models?\nI guess even if other models don't quite reach resnet50's performance, they might be useful for ensembling.",
      "votes": null
    },
    {
      "id": "275634",
      "postDate": "01/29/2018 17:14:10",
      "content": "<blockquote>\n  <p>Exception in thread Thread-1:\n  multiprocessing.pool.RemoteTraceback:\n  \"\"\"\n  Traceback (most recent call last):\n    File \"/usr/lib/python3.4/multiprocessing/pool.py\", line 119, in worker\n      result = (True, func(*args, **kwds))\n    File \"/usr/lib/python3.4/multiprocessing/pool.py\", line 44, in mapstar\n      return list(map(*args))\n    File \"train.py\", line 252, in process_item\n      img = load_img_fast_jpg(item)\n    File \"train.py\", line 148, in \n      load_img_fast_jpg  = lambda img_path: jpeg.JPEG(img_path).decode()\n    File \"/home/xiaodong_he/py3env/lib/python3.4/site-packages/jpeg4py/<em>py.py\", line 128,                                                                                                       in <strong>init</strong>\n      super(JPEG, self).<strong>init</strong>(lib</em>)\n    File \"/home/xiaodong_he/py3env/lib/python3.4/site-packages/jpeg4py/_py.py\", line 64,                                                                                                       in <strong>init</strong>\n      jpeg.initialize()\n    File \"/home/xiaodong_he/py3env/lib/python3.4/site-packages/jpeg4py/_cffi.py\", line 21                                                                                                      2, in initialize\n      _initialize(backends)\n    File \"/home/xiaodong_he/py3env/lib/python3.4/site-packages/jpeg4py/_cffi.py\", line 19                                                                                                      3, in _initialize\n      raise OSError(\"Could not load libjpeg-turbo library\")\n  OSError: Could not load libjpeg-turbo library\n  \"\"\"</p>\n</blockquote>\n\n<p>anyone got some idea about this error?thanks </p>",
      "rawMarkdown": "&gt; Exception in thread Thread-1:\nmultiprocessing.pool.RemoteTraceback:\n\"\"\"\nTraceback (most recent call last):\n  File \"/usr/lib/python3.4/multiprocessing/pool.py\", line 119, in worker\n    result = (True, func(*args, **kwds))\n  File \"/usr/lib/python3.4/multiprocessing/pool.py\", line 44, in mapstar\n    return list(map(*args))\n  File \"train.py\", line 252, in process_item\n    img = load_img_fast_jpg(item)\n  File \"train.py\", line 148, in",
      "votes": null
    },
    {
      "id": "275636",
      "postDate": "01/29/2018 17:17:46",
      "content": "<p>You'll need to use <code>apt-get install libturbojpeg</code> or equivalent.</p>",
      "rawMarkdown": "You'll need to use `apt-get install libturbojpeg` or equivalent.",
      "votes": null
    },
    {
      "id": "275642",
      "postDate": "01/29/2018 17:26:39",
      "content": "<p>it does work ,many thanks</p>",
      "rawMarkdown": "it does work ,many thanks",
      "votes": null
    },
    {
      "id": "275872",
      "postDate": "01/30/2018 06:11:11",
      "content": "<p>The same question with HamYad. By the way, my validation accuracy is 80%+ but the LB score is only about 0.4, could you tell me how to improve?</p>",
      "rawMarkdown": "The same question with HamYad. By the way, my validation accuracy is 80%+ but the LB score is only about 0.4, could you tell me how to improve?",
      "votes": null
    },
    {
      "id": "275878",
      "postDate": "01/30/2018 06:26:10",
      "content": "<p>Excuse me. Do you mean that you trained your model by all the training data and validated the model by the validation data from Gleb's post?</p>",
      "rawMarkdown": "Excuse me. Do you mean that you trained your model by all the training data and validated the model by the validation data from Gleb's post?",
      "votes": null
    },
    {
      "id": "275927",
      "postDate": "01/30/2018 08:56:27",
      "content": "<p>Hi Daniel,  when you cropping train set to 112x112 and use it to train model, what do you use for test set, which size is 512x512? You just resize it from 512 to 112 or may be you get crop from test? Thanks!</p>",
      "rawMarkdown": "Hi Daniel,  when you cropping train set to 112x112 and use it to train model, what do you use for test set, which size is 512x512? You just resize it from 512 to 112 or may be you get crop from test? Thanks!",
      "votes": null
    },
    {
      "id": "275946",
      "postDate": "01/30/2018 09:45:18",
      "content": "<p>I find out, that resizing is bad practice.... Getting crop of test size is much better. May be it help somebody :)</p>",
      "rawMarkdown": "I find out, that resizing is bad practice.... Getting crop of test size is much better. May be it help somebody :)",
      "votes": null
    },
    {
      "id": "275954",
      "postDate": "01/30/2018 10:15:47",
      "content": "<p>Hi Alex. You're looking for image noise patterns. Particularly, it looks like camera model create some kind of grid fingerprint. Resizing (scaling) from 512x512 to 112x112 doesn't lost much information about image content, but it does about noise.  From my side, I get worst predictions for jpg70 and scale0.5 altered images.</p>",
      "rawMarkdown": "Hi Alex. You're looking for image noise patterns. Particularly, it looks like camera model create some kind of grid fingerprint. Resizing (scaling) from 512x512 to 112x112 doesn't lost much information about image content, but it does about noise.  From my side, I get worst predictions for jpg70 and scale0.5 altered images.",
      "votes": null
    },
    {
      "id": "275959",
      "postDate": "01/30/2018 10:25:49",
      "content": "<p>Daniel, how do you quantify worst predictions: looking at the soft-probabilities (softmax) below a threshold, e.g. 0.7?</p>\n\n<p>In my case I'm doing that I have worst predictions for Nexus-5. </p>",
      "rawMarkdown": "Daniel, how do you quantify worst predictions: looking at the soft-probabilities (softmax) below a threshold, e.g. 0.7?\n\nIn my case I'm doing that I have worst predictions for Nexus-5.",
      "votes": null
    },
    {
      "id": "275963",
      "postDate": "01/30/2018 10:31:07",
      "content": "<p>Cross validation. Same here, LG Nexus 5x and Motorola Nexus 6.</p>",
      "rawMarkdown": "Cross validation. Same here, LG Nexus 5x and Motorola Nexus 6.",
      "votes": null
    },
    {
      "id": "276139",
      "postDate": "01/30/2018 19:48:59",
      "content": "<p>I'm having issues with those two as well, but also the iphone-4s which for me is the worst.</p>",
      "rawMarkdown": "I'm having issues with those two as well, but also the iphone-4s which for me is the worst.",
      "votes": null
    },
    {
      "id": "276142",
      "postDate": "01/30/2018 20:01:06",
      "content": "<p>Look at the data you should, young jedi. :-)</p>",
      "rawMarkdown": "Look at the data you should, young jedi. :-)",
      "votes": null
    },
    {
      "id": "276227",
      "postDate": "01/31/2018 02:12:44",
      "content": "<p>Hi Andres! My model also achieves ~90% val acc, but LB is ~30-40% only. But I can't find out a solution. Could you tell me how to deal with it?</p>",
      "rawMarkdown": "Hi Andres! My model also achieves ~90% val acc, but LB is ~30-40% only. But I can't find out a solution. Could you tell me how to deal with it?",
      "votes": null
    },
    {
      "id": "276279",
      "postDate": "01/31/2018 04:50:55",
      "content": "<p>Maybe you should use proper preprocessing(e.g. jpeg compression ) to training set and validation set.</p>",
      "rawMarkdown": "Maybe you should use proper preprocessing(e.g. jpeg compression ) to training set and validation set.",
      "votes": null
    },
    {
      "id": "276296",
      "postDate": "01/31/2018 05:48:38",
      "content": "<p>Test data is different than train in 3 big ways (then there's smaller issues);</p>\n\n<ol>\n<li>50% of it is manipulated as per the instructions</li>\n<li>Taken from a different device (BIG issue)</li>\n<li>Center-cropped, fixed res.</li>\n</ol>\n\n<p>You need to address all 3.</p>",
      "rawMarkdown": "Test data is different than train in 3 big ways (then there's smaller issues);\n\n 1. 50% of it is manipulated as per the instructions\n 2. Taken from a different device (BIG issue)\n 3. Center-cropped, fixed res.\n\nYou need to address all 3.",
      "votes": null
    },
    {
      "id": "276670",
      "postDate": "02/01/2018 05:01:10",
      "content": "<p>Thanks Andres. I've looked into my preprocession. Now it seems I'm going on a right way. Thanks very much!</p>",
      "rawMarkdown": "Thanks Andres. I've looked into my preprocession. Now it seems I'm going on a right way. Thanks very much!",
      "votes": null
    },
    {
      "id": "277221",
      "postDate": "02/02/2018 15:12:41",
      "content": "<p>Awesome work! Thanks for your sharing!</p>\n\n<p>When run this code find this error:\nTraceback (most recent call last):\n  File \"train.py\", line 540, in \n    max_classes_val_count = max(classes_val_count)\nValueError: max() arg is an empty sequence</p>\n\n<p>I use ne GTX1070Ti, TF 1.4.1; keras 2.1.3; numpy 1.13.1</p>",
      "rawMarkdown": "Awesome work! Thanks for your sharing!\n\n\nWhen run this code find this error:\nTraceback (most recent call last):\n  File \"train.py\", line 540, in",
      "votes": null
    },
    {
      "id": "277233",
      "postDate": "02/02/2018 15:55:42",
      "content": "<p>Looks your validation set is empty. Check whether you have downloaded val set accordingly.</p>",
      "rawMarkdown": "Looks your validation set is empty. Check whether you have downloaded val set accordingly.",
      "votes": null
    },
    {
      "id": "277245",
      "postDate": "02/02/2018 16:28:18",
      "content": "<p>Wow! You are simply a grand master! Thanks for sharing.</p>",
      "rawMarkdown": "Wow! You are simply a grand master! Thanks for sharing.",
      "votes": null
    },
    {
      "id": "277381",
      "postDate": "02/02/2018 23:50:44",
      "content": "<p>I check it ,val_images folder is ok. Thanks for your help!</p>\n\n<p>Last time, I run:python train.py -g 1 -b 8 -cs 512 -cm ResNet50 -x -l 1e-4</p>\n\n<p>This time ,I  run:python train.py -g 1 -b 8 -cs 512 -cm ResNet50 -l 1e-4,\nerror:</p>\n\n<p>/keras/engine/training.py\", line 2053, in fit_generator\n    raise ValueError('<code>validation_steps=None</code> is only valid for a'\nValueError: <code>validation_steps=None</code> is only valid for a generator based on the <code>keras.utils.Sequence</code> class. Please specify <code>validation_steps</code> or use the <code>keras.utils.Sequence</code> class.</p>",
      "rawMarkdown": "I check it ,val_images folder is ok. Thanks for your help!\n\nLast time, I run:python train.py -g 1 -b 8 -cs 512 -cm ResNet50 -x -l 1e-4\n\nThis time ,I  run:python train.py -g 1 -b 8 -cs 512 -cm ResNet50 -l 1e-4,\nerror:\n\n/keras/engine/training.py\", line 2053, in fit_generator\n    raise ValueError('`validation_steps=None` is only valid for a'\nValueError: `validation_steps=None` is only valid for a generator based on the `keras.utils.Sequence` class. Please specify `validation_steps` or use the `keras.utils.Sequence` class.",
      "votes": null
    },
    {
      "id": "277549",
      "postDate": "02/03/2018 14:38:49",
      "content": "<p>Brilliant!\nnot use -x , it works,  without extra dataset. I have put the floders(<a href=\"https://www.kaggle.com/c/sp-society-camera-model-identification/discussion/47235\">https://www.kaggle.com/c/sp-society-camera-model-identification/discussion/47235</a> )  in it , while not work.</p>",
      "rawMarkdown": "Brilliant!\nnot use -x , it works,  without extra dataset. I have put the floders(https://www.kaggle.com/c/sp-society-camera-model-identification/discussion/47235 )  in it , while not work.",
      "votes": null
    },
    {
      "id": "277551",
      "postDate": "02/03/2018 14:44:31",
      "content": "<p>You have to download the extra data in order to use it</p>",
      "rawMarkdown": "You have to download the extra data in order to use it",
      "votes": null
    },
    {
      "id": "277706",
      "postDate": "02/04/2018 01:04:02",
      "content": "<p>extra dataset seem importance, without extra dataset (python train.py -g 1 -b 8 -cs 512 -cm ResNet50 -l 1e-4), epoch054-val_acc0.303427.</p>",
      "rawMarkdown": "extra dataset seem importance, without extra dataset (python train.py -g 1 -b 8 -cs 512 -cm ResNet50 -l 1e-4), epoch054-val_acc0.303427.",
      "votes": null
    },
    {
      "id": "278353",
      "postDate": "02/06/2018 01:05:20",
      "content": "<p>I run: python train.py -g 1 -b 8 -cs 229 -cm DenseNet201 -x -l 1e-4 -uiw</p>\n\n<p>The same error, I have try to fix it by add this:</p>\n\n<p>def load_img_fast_jpg(img_path):\ntry:\n    x = jpeg.JPEG(img_path).decode()\n    return x\nexcept:\n    return load_img(img_path)</p>\n\n<p>But not solve it.</p>\n\n<p>Exception in thread Thread-8:\nTraceback (most recent call last):\n  File \"/home/yl/miniconda3/envs/tensorflow/lib/python3.6/threading.py\", line 916, in _bootstrap_inner\n    self.run()\n  File \"/home/yl/miniconda3/envs/tensorflow/lib/python3.6/threading.py\", line 864, in run\n    self._target(*self._args, **self._kwargs)\n  File \"/home/yl/miniconda3/envs/tensorflow/lib/python3.6/multiprocessing/pool.py\", line 463, in _handle_results\n    task = get()\n  File \"/home/miniconda3/envs/tensorflow/lib/python3.6/multiprocessing/connection.py\", line 251, in recv\n    return _ForkingPickler.loads(buf.getbuffer())\nTypeError: <strong>init</strong>() missing 1 required positional argument: 'code'</p>",
      "rawMarkdown": "I run: python train.py -g 1 -b 8 -cs 229 -cm DenseNet201 -x -l 1e-4 -uiw\n\nThe same error, I have try to fix it by add this:\n\ndef load_img_fast_jpg(img_path):\ntry:\n    x = jpeg.JPEG(img_path).decode()\n    return x\nexcept:\n    return load_img(img_path)\n\nBut not solve it.\n\nException in thread Thread-8:\nTraceback (most recent call last):\n  File \"/home/yl/miniconda3/envs/tensorflow/lib/python3.6/threading.py\", line 916, in _bootstrap_inner\n    self.run()\n  File \"/home/yl/miniconda3/envs/tensorflow/lib/python3.6/threading.py\", line 864, in run\n    self._target(*self._args, **self._kwargs)\n  File \"/home/yl/miniconda3/envs/tensorflow/lib/python3.6/multiprocessing/pool.py\", line 463, in _handle_results\n    task = get()\n  File \"/home/miniconda3/envs/tensorflow/lib/python3.6/multiprocessing/connection.py\", line 251, in recv\n    return _ForkingPickler.loads(buf.getbuffer())\nTypeError: __init__() missing 1 required positional argument: 'code'",
      "votes": null
    },
    {
      "id": "278376",
      "postDate": "02/06/2018 02:40:51",
      "content": "<p>Really good job on the script. Did you design the CaCNN network just from reading the paper alone? </p>\n\n<p>Awesome work!</p>",
      "rawMarkdown": "Really good job on the script. Did you design the CaCNN network just from reading the paper alone? \n\nAwesome work!",
      "votes": null
    },
    {
      "id": "278394",
      "postDate": "02/06/2018 03:56:05",
      "content": "<p>I add this code as follows, but not wok. How to fix?\nThanks.</p>\n\n<p>def load_img_fast_jpg(img_path):\n    try:\n        x = jpeg.JPEG(img_path).decode()\n        return x\n    except:\n        return load_img(img_path)</p>\n\n<p>load_img_fast_jpg  = lambda img_path: jpeg.JPEG(img_path).decode()\nload_img  = lambda img_path: np.array(Image.open(img_path))</p>",
      "rawMarkdown": "I add this code as follows, but not wok. How to fix?\nThanks.\n\n\ndef load_img_fast_jpg(img_path):\n    try:\n        x = jpeg.JPEG(img_path).decode()\n        return x\n    except:\n        return load_img(img_path)\n\nload_img_fast_jpg  = lambda img_path: jpeg.JPEG(img_path).decode()\nload_img  = lambda img_path: np.array(Image.open(img_path))",
      "votes": null
    },
    {
      "id": "379945",
      "postDate": "09/01/2018 09:19:34",
      "content": "<p>Sir, Thanks for sharing your code.I am currently working on Camera Model Identification for my undergrad project.  Can i train my own dataset in your given model ? if i can , what do i really have to do with the code. TIA</p>",
      "rawMarkdown": "Sir, Thanks for sharing your code.I am currently working on Camera Model Identification for my undergrad project.  Can i train my own dataset in your given model ? if i can , what do i really have to do with the code. TIA",
      "votes": null
    },
    {
      "id": "380143",
      "postDate": "09/01/2018 19:53:49",
      "content": "<p>Sure, you can use. Code should be self-explanatory though.</p>",
      "rawMarkdown": "Sure, you can use. Code should be self-explanatory though.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 271445,
      "author_name": "kaliev",
      "author_url": "",
      "post_date": "01/20/2018 13:08:03",
      "content": "<p>Hey Andres!</p>\n\n<p>Nice to see you ;)</p>\n\n<p>Why do you resize the train images? The test images are not resized, they are center cropped:</p>\n\n<blockquote>\n  <p>While the train data includes full images, the test data contains only single 512 x 512 pixel blocks cropped from the center of a single image taken with the device.</p>\n</blockquote>",
      "votes": null,
      "replies": [
        {
          "id": 271448,
          "author_name": "antorsae",
          "author_url": "",
          "post_date": "01/20/2018 13:21:16",
          "content": "<p>Hi Ali! yes, I just noticed... just pushed new code, much better now:</p>\n\n<p><code>Epoch 29/200\n139/139 [==============================] - 53s 384ms/step - loss: 0.2487 - acc: 0.9101 - val_loss: 0.4792 - val_acc: 0.8750</code></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 272235,
          "author_name": "hamyadlab",
          "author_url": "",
          "post_date": "01/22/2018 15:42:37",
          "content": "<p>It seems the validation accuracy on its own does not mean much. I had a model with both training and validation accuracy above 98%, but the LB was only 64%. On the other hand, I had a model with train accuracy equal to 91 and validation accuracy equal to 95, and the LB accuracy for that model was 85.2. I am totally confused!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 272236,
          "author_name": "hamyadlab",
          "author_url": "",
          "post_date": "01/22/2018 15:44:02",
          "content": "<p>Maybe I should reconsider to create a whole new validation set. I think this is the key point. Maybe the high number of submissions of top LBs is a confirmation! I think they have tried many models until they have found a good one.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 272242,
          "author_name": "jfkingiii",
          "author_url": "",
          "post_date": "01/22/2018 15:58:09",
          "content": "<p>I find when I use the validation data from Gleb's post, my validation accuracy is close to my LB score. I trained a model to 0.79 on the Gleb validation set and got 0.804 on the public leaderboard.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 275878,
          "author_name": "hzywish",
          "author_url": "",
          "post_date": "01/30/2018 06:26:10",
          "content": "<p>Excuse me. Do you mean that you trained your model by all the training data and validated the model by the validation data from Gleb's post?</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 272157,
      "author_name": "hamyadlab",
      "author_url": "",
      "post_date": "01/22/2018 11:29:43",
      "content": "<p>I think it is not a good idea to resize the training images</p>",
      "votes": null,
      "replies": [
        {
          "id": 272178,
          "author_name": "antorsae",
          "author_url": "",
          "post_date": "01/22/2018 12:54:19",
          "content": "<p>Agreed, I've updated the code and it doesn't resize anymore. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 272181,
      "author_name": "antorsae",
      "author_url": "",
      "post_date": "01/22/2018 12:55:57",
      "content": "<p>Update:\n- No resizing, just takes 512x512 crops from the center</p>\n\n<p>The good news is that it achieves ~90% val acc, however LB is ~30-40% only.  Currently investigating why.</p>",
      "votes": null,
      "replies": [
        {
          "id": 272213,
          "author_name": "melgor",
          "author_url": "",
          "post_date": "01/22/2018 14:32:19",
          "content": "<p>I have similar issue. I'm trying to resolve it with no success.  If I will resolve it somehow, I will let you know.\nMy model is ResNet50 finetuned with random crop from full size images. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 272216,
          "author_name": "igormunizims",
          "author_url": "",
          "post_date": "01/22/2018 14:54:07",
          "content": "<p>What size of your crop? I have the same issue.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 272217,
          "author_name": "antorsae",
          "author_url": "",
          "post_date": "01/22/2018 14:56:03",
          "content": "<p>512x512, same as the test size.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 272300,
          "author_name": "danielfg",
          "author_url": "",
          "post_date": "01/22/2018 19:20:08",
          "content": "<p>From my experience, I think you could get a LB score &gt; 0.90 just cropping to 112x112. That way you can train your models faster. Make some experiments until you get good parameters (LR, dense layers size, drop, augmentation). When you have a good one, scale up the cropping size to get better results.\nRegarding the validation/LB gap, just for reference, I got 0.957/0.922 (val/LB). Using additional data should help to reduce the gap.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 272308,
          "author_name": "igormunizims",
          "author_url": "",
          "post_date": "01/22/2018 19:39:37",
          "content": "<p>Hi Daniel, You are fine tuning some pre trained model or using your own model?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 272500,
          "author_name": "antorsae",
          "author_url": "",
          "post_date": "01/23/2018 06:20:52",
          "content": "<p>Hi Daniel, did you use Gleb's dataset for your 0.922 LB or just the organization-provided dataset?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 272776,
          "author_name": "danielfg",
          "author_url": "",
          "post_date": "01/23/2018 17:21:30",
          "content": "<p>Just the organization-provided dataset. Now using Gleb's dataset too.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 275927,
          "author_name": "famazon",
          "author_url": "",
          "post_date": "01/30/2018 08:56:27",
          "content": "<p>Hi Daniel,  when you cropping train set to 112x112 and use it to train model, what do you use for test set, which size is 512x512? You just resize it from 512 to 112 or may be you get crop from test? Thanks!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 275946,
          "author_name": "famazon",
          "author_url": "",
          "post_date": "01/30/2018 09:45:18",
          "content": "<p>I find out, that resizing is bad practice.... Getting crop of test size is much better. May be it help somebody :)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 275954,
          "author_name": "danielfg",
          "author_url": "",
          "post_date": "01/30/2018 10:15:47",
          "content": "<p>Hi Alex. You're looking for image noise patterns. Particularly, it looks like camera model create some kind of grid fingerprint. Resizing (scaling) from 512x512 to 112x112 doesn't lost much information about image content, but it does about noise.  From my side, I get worst predictions for jpg70 and scale0.5 altered images.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 275959,
          "author_name": "antorsae",
          "author_url": "",
          "post_date": "01/30/2018 10:25:49",
          "content": "<p>Daniel, how do you quantify worst predictions: looking at the soft-probabilities (softmax) below a threshold, e.g. 0.7?</p>\n\n<p>In my case I'm doing that I have worst predictions for Nexus-5. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 275963,
          "author_name": "danielfg",
          "author_url": "",
          "post_date": "01/30/2018 10:31:07",
          "content": "<p>Cross validation. Same here, LG Nexus 5x and Motorola Nexus 6.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 276139,
          "author_name": "albertoa",
          "author_url": "",
          "post_date": "01/30/2018 19:48:59",
          "content": "<p>I'm having issues with those two as well, but also the iphone-4s which for me is the worst.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 276142,
          "author_name": "antorsae",
          "author_url": "",
          "post_date": "01/30/2018 20:01:06",
          "content": "<p>Look at the data you should, young jedi. :-)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 276227,
          "author_name": "hzywish",
          "author_url": "",
          "post_date": "01/31/2018 02:12:44",
          "content": "<p>Hi Andres! My model also achieves ~90% val acc, but LB is ~30-40% only. But I can't find out a solution. Could you tell me how to deal with it?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 276279,
          "author_name": "wuzuping",
          "author_url": "",
          "post_date": "01/31/2018 04:50:55",
          "content": "<p>Maybe you should use proper preprocessing(e.g. jpeg compression ) to training set and validation set.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 276296,
          "author_name": "antorsae",
          "author_url": "",
          "post_date": "01/31/2018 05:48:38",
          "content": "<p>Test data is different than train in 3 big ways (then there's smaller issues);</p>\n\n<ol>\n<li>50% of it is manipulated as per the instructions</li>\n<li>Taken from a different device (BIG issue)</li>\n<li>Center-cropped, fixed res.</li>\n</ol>\n\n<p>You need to address all 3.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 276670,
          "author_name": "hzywish",
          "author_url": "",
          "post_date": "02/01/2018 05:01:10",
          "content": "<p>Thanks Andres. I've looked into my preprocession. Now it seems I'm going on a right way. Thanks very much!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 272526,
      "author_name": "antorsae",
      "author_url": "",
      "post_date": "01/23/2018 07:27:46",
      "content": "<p>Update:\nUsing Gleb's dataset I get 0.76 LB with a ResNet50 and pre-trained weights and a fixed 224 crop size.</p>\n\n<p>Next steps:</p>\n\n<ul>\n<li>Balance classes (DONE)</li>\n<li>Fine-tune FC, dropout, whether to use high-pass kernel filtering (<code>-kf</code>) and determine best architecture</li>\n<li>Account for different weights (0.3 vs. 0.7) in loss function</li>\n<li>Get more data (I think this is critical)</li>\n</ul>",
      "votes": null,
      "replies": [
        {
          "id": 272780,
          "author_name": "albertoa",
          "author_url": "",
          "post_date": "01/23/2018 17:42:40",
          "content": "<p>So the .76 is just with the built in Keras ResNet50? Since you mention fine-tuning the FC does that mean you didn't even replace the FC?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 272798,
          "author_name": "antorsae",
          "author_url": "",
          "post_date": "01/23/2018 18:28:38",
          "content": "<p>I've already changed the code a lot but IIRC it was something like <code>python train.py -l 1e-3 -b 64 -g 2 -cm ResNet50 -cs 224 -x  -p avg -do 0.3  -kf</code></p>\n\n<p>I added random crops to the current code so score may be different now. I've also added crop ensembling to <code>-test</code> so it may help.\nSummary of paramaters:\n - <code>-b 64</code> batch size for each GPU\n - <code>-g 2</code> use 2 GPUs\n - <code>-cm ResNet50</code> use ResNet50 as feature extractor\n - <code>-cs 224</code> crop size\n - <code>-x</code> use Gleb's dataset\n - <code>-p avg</code> use average pooling on the classifier (ResNet50)\n - <code>-do 0.3</code> use 0.3 dropout rate on my FC layers \n - <code>-kf</code> use a high pass pre-processing filter.</p>\n\n<p>You may want to take a look at the code to try it out. Once you have a model with decent <code>val_acc</code> do <code>-m path_to_saved_model.hdf5 -t</code> to generate <code>submission.csv</code> </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 272959,
          "author_name": "albertoa",
          "author_url": "",
          "post_date": "01/24/2018 01:25:22",
          "content": "<p>I went ahead and got your code from github. </p>\n\n<p>I edited the post as the errors where caused by an unknown issue.</p>\n\n<p>I have successfully run your code and got results in the  0.503 LB using the flags above (except for the flickr data, which I still have to get properly)</p>\n\n<p>For my submissions (0.517 LB) I was using a resnet like architecture with 3 stages training from scratch. It was taking 3-6 hours for training and your model is only taking me 1 hour. So now I have something better to look at.</p>\n\n<p>Looking at the CaCNN code and the paper you reference I guess I'll be running that next.</p>\n\n<p>Thanks for sharing the code.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 273376,
          "author_name": "antorsae",
          "author_url": "",
          "post_date": "01/24/2018 14:30:24",
          "content": "<p>Great. I am finishing a multiprocess generator which does not need to cache (memory or disk) and has the benefit of being able to select random crops from all across the image. Will be ready in a few hours. Stay tuned.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 273388,
          "author_name": "albertoa",
          "author_url": "",
          "post_date": "01/24/2018 14:55:25",
          "content": "<p>Great, I look forward to looking at it.</p>\n\n<p>By the way let me know if you would like to merge as a team. </p>\n\n<p>I've been using hdf5 files with unprocessed and pre-processed patches, speed wise they work well. I'm at the stage of testing multiple networks and improving results.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 273414,
          "author_name": "antorsae",
          "author_url": "",
          "post_date": "01/24/2018 15:25:45",
          "content": "<p>Code is updated, now with multiprocess generator which is more flexible, especially wrt selecting random crops (previously only a smaller version of the original image was saved in the cache so cropping was more limited). </p>\n\n<p>I think you need a powerful CPU to do all augmentations, etc. and saturate your GPUs. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 273523,
          "author_name": "antorsae",
          "author_url": "",
          "post_date": "01/24/2018 18:07:19",
          "content": "<p>The new code achieves ~0.73 just after 8 epochs:\n<code>$ python train.py -g 2 -b 16 -cs 512 -cm ResNet50 -x -l 1e-4 -uiw</code></p>\n\n<p>Epoch 1/100\n507/507 [==============================] - 227s 447ms/step - loss: 1.6933 - acc: 0.4243 - val_loss: 1.6242 - val_acc: 0.5395\nEpoch 2/100\n507/507 [==============================] - 205s 405ms/step - loss: 1.0983 - acc: 0.6495 - val_loss: 1.1173 - val_acc: 0.6711\nEpoch 3/100\n507/507 [==============================] - 207s 408ms/step - loss: 0.8839 - acc: 0.7187 - val_loss: 1.6343 - val_acc: 0.5230\nEpoch 4/100\n507/507 [==============================] - 206s 406ms/step - loss: 0.7618 - acc: 0.7625 - val_loss: 0.9526 - val_acc: 0.7105\nEpoch 5/100\n507/507 [==============================] - 204s 403ms/step - loss: 0.6861 - acc: 0.7855 - val_loss: 1.0580 - val_acc: 0.6941\nEpoch 6/100\n507/507 [==============================] - 203s 399ms/step - loss: 0.6055 - acc: 0.8104 - val_loss: 0.7093 - val_acc: 0.7599\nEpoch 7/100\n507/507 [==============================] - 197s 389ms/step - loss: 0.5509 - acc: 0.8274 - val_loss: 1.1975 - val_acc: 0.6875\nEpoch 8/100\n507/507 [==============================] - 351s 693ms/step - loss: 0.5159 - acc: 0.8410 - val_loss: 0.8306 - val_acc: 0.8125</p>\n\n<p>(it then crashed due to a memory leak... need to investigate)</p>\n\n<p>then:\n<code>$ python train.py -g 1 -b 16 -cs 512 -cm ResNet50 -x -l 1e-4 -uiw -m models/ResNet50_do0.3_avg-epoch008-val_acc0.812500.hdf5 -t\n</code></p>\n\n<p>then:\n<code>$ kg submit submission.csv\n0.733\n</code></p>\n\n<p>I've noticed that in Gleb's dataset some images did not match any resolution, so I discard them. Also Sony-NEX-7 directory had files ending in JPG, not jpg... so my code was not using them. Renamed then.</p>\n\n<p>Finally I noted some images are portrait whereas others are landscape, so now I'm augmenting via orientation change during training.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 273550,
          "author_name": "albertoa",
          "author_url": "",
          "post_date": "01/24/2018 18:55:50",
          "content": "<p>I'm still screwing up the execution somehow. Once again I got a run with no learning:</p>\n\n<p>Epoch 100/100\n245/245 [==============================] - 81s - loss: 2.2325 - acc: 0.1793 - val_loss: 2.3201 - val_acc: 0.1562</p>\n\n<p>I'm wiping out the whole models directory and restarted the training with:</p>\n\n<p>-l 1e-3 -b 32 -g 1 -cm ResNet50 -cs 224 -p avg -do 0.3 -x -kf -uiw</p>\n\n<p>We'll see if this run it learns (I wasn't using the uiw before but I can't see how that would explain the model not learning). </p>\n\n<p>Not sure why you got JPG, all my links in flickr_images/sony_nex7/urls_final were lower case.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 273553,
          "author_name": "antorsae",
          "author_url": "",
          "post_date": "01/24/2018 19:07:53",
          "content": "<p>Try changing (lowering) the learning rate. Also, intuitively I don't think <code>-uiw</code> would work OK with <code>-kf</code> the reason being is <code>-kf</code> applies a high-pass filter on the images and they cease to look anything like an image, and also the normalization is different than in imagenet (mean is similar but std and range are not).</p>\n\n<p>The JPG vs. jpg was on the train directory (organization dataset).</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 273652,
          "author_name": "albertoa",
          "author_url": "",
          "post_date": "01/24/2018 23:21:27",
          "content": "<p>I can't seem to get it to run. When I saw this before my directories didn't match the script expectations, but this time I seem to have the right numbers on the image data:</p>\n\n<pre><code>Image set counts\n</code></pre>\n\n<p>extra_train_ids:  5368\n  extra_val_ids:   315</p>\n\n<pre><code>  ids_train:  8118 steps=  507\n    ids_val:   315 steps=   19\n</code></pre>\n\n<hr>\n\n<pre><code>          HTC-1-M7:  1023 (12.6%)\n          iPhone-6:   823 (10.1%)\n</code></pre>\n\n<p>Motorola-Droid-Maxx:   825 (10.2%)\n            Motorola-X:   275 (03.4%)\n     Samsung-Galaxy-S4:  1412 (17.4%)\n             iPhone-4s:   774 (09.5%)\n           LG-Nexus-5x:   680 (08.4%)\n      Motorola-Nexus-6:   926 (11.4%)\n  Samsung-Galaxy-Note3:   548 (06.8%)\n            Sony-NEX-7:   832 (10.2%)</p>\n\n<pre><code>Exception in thread Thread-1:\n</code></pre>\n\n<p>multiprocessing.pool.RemoteTraceback: \n\"\"\"\nTraceback (most recent call last):\n  File \"/usr/lib/python3.5/multiprocessing/pool.py\", line 119, in worker\n    result = (True, func(*args, **kwds))\n  File \"/usr/lib/python3.5/multiprocessing/pool.py\", line 44, in mapstar\n    return list(map(*args))\n  File \"train.py\", line 265, in process_item\n    img = preprocess_image(img)\n  File \"train.py\", line 197, in preprocess_image\n    return preprocess_input_function(img.astype(np.float32))\n  File \"/home/alonsoa/projects/virtenv/py3cv3/lib/python3.5/site-packages/keras/applications/imagenet_utils.py\", line 33, in preprocess_input\n    x = x[:, :, :, ::-1]\nIndexError: too many indices for array\n\"\"\"</p>\n\n<p>Anything jumps at you why   return preprocess_input_function(img.astype(np.float32)) would fail even though the image is shaped (512, 512, 3)\"</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 273655,
          "author_name": "jfkingiii",
          "author_url": "",
          "post_date": "01/24/2018 23:36:00",
          "content": "<p>I had the same error. Sometimes the images don't download correctly and they read in as an array with less than three dimensions. I identified some of these images and downloaded them again, but had trouble getting them all. So I added logic that would check that len(img.shape) == 3 and if not move on to the next image.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 273657,
          "author_name": "albertoa",
          "author_url": "",
          "post_date": "01/24/2018 23:43:03",
          "content": "<p>Thanks, unfortunately there already is:</p>\n\n<pre><code>    if img.ndim != 3:\n    return None\n</code></pre>",
          "votes": null,
          "replies": []
        },
        {
          "id": 273713,
          "author_name": "albertoa",
          "author_url": "",
          "post_date": "01/25/2018 02:55:51",
          "content": "<p>I ended up having to change the line to:</p>\n\n<pre><code>    return preprocess_input_function(np.expand_dims(img.astype(np.float32), axis=0))\n</code></pre>\n\n<p>as Keras' preprocess_input expects an array of images.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 273813,
          "author_name": "antorsae",
          "author_url": "",
          "post_date": "01/25/2018 08:05:33",
          "content": "<p>What version of Keras are you running? Mine:</p>\n\n<pre><code>$ python -c 'import keras;print(keras.__version__)'\n/home/antor/miniconda3/lib/python3.6/site-packages/h5py/__init__.py:36: FutureWarning: Conversion of the second argument of issubdtype from `float` to `np.floating` is deprecated. In future, it will be treated as `np.float64 == np.dtype(float).type`.\n  from ._conv import register_converters as _register_converters\nUsing TensorFlow backend.\n2.1.3\n</code></pre>\n\n<p>There's slightly different preprocessing functions on keras.applications and the code selects the appropriate one for each of the networks in keras.applications.* and defaults to xception (which is just normalizing between -1 and 1 iirc) for all others).</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 273916,
          "author_name": "albertoa",
          "author_url": "",
          "post_date": "01/25/2018 13:51:55",
          "content": "<p>I am running Keras (2.0.6) and I was running tensorflow (1.3.0)</p>\n\n<p>Turns out that last night I was getting CUDNN_STATUS_INTERNAL_ERROR which made me upgrade tensorflow to (1.4.1). After doing that I was no longer using the GPU and I didn't notice. That explains my previous comment about the CPU high usage, it was running the model. After upgrading tensorflow-gpu as well to (1.4.1) I'm back on the GPU (a GTX 1080Ti )</p>\n\n<p>As for samples the only difference between us is your Sony-NEX-7:   832 (10.2%) vs. mine Sony-NEX-7:   557 (07.1%) so I'll double check which samples I'm missing.</p>\n\n<p>I'm still having some performance and learning issues though. I kept running out of memory unless I brought the batch size to 1</p>\n\n<pre><code> -g 1 -b 1 -cs 512 -cm ResNet50 -x -l 1e-4 -uiw --max-epoch 350\n</code></pre>\n\n<p>Epoch 19/350\n7843/7843 [==============================] - 1159s - loss: 1.3477 - acc: 0.5560 - val_loss: 3.1935 - val_acc: 0.1108</p>\n\n<p>I'll upgrade keras and rerun</p>\n\n<p>Thanks for all the help. I hope my info helps anybody else trying to reproduce the results.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 273921,
          "author_name": "antorsae",
          "author_url": "",
          "post_date": "01/25/2018 13:59:26",
          "content": "<p>The discrepancy in Sony-NEX-7 is b/c in that folder the original files are .JPG and should be renamed to .jpg.</p>\n\n<p>Re: versions </p>\n\n<pre><code>$ python -c 'import tensorflow as tf;import keras;print(tf.__version__, keras.__version__)'\n/home/antor/miniconda3/lib/python3.6/importlib/_bootstrap.py:219: RuntimeWarning: compiletime version 3.5 of module 'tensorflow.python.framework.fast_tensor_util' does not match runtime version 3.6\n  return f(*args, **kwds)\nUsing TensorFlow backend.\n1.4.1 2.1.3\n</code></pre>\n\n<p>Re: batch size you should be able to fit more in a 1080 Ti. I am testing the code on a different machine (AMD Ryzen 1600X + 1070 Ti) and this works w/o issue <code>$ python train.py -g 1 -b 8 -cs 512 -cm ResNet50 -x -l 1e-4</code>. Check with <code>nvidia-smi</code> that no memory is being used by the GPU before launching the script.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 273957,
          "author_name": "albertoa",
          "author_url": "",
          "post_date": "01/25/2018 15:22:48",
          "content": "<p>As per your prior suggestion I already renamed the files. I think my downloads are missing some though:</p>\n\n<pre><code>vdir flickr_images/sony_nex7/|wc\n    558    6129   74089\n</code></pre>\n\n<p>Now that I have the same versions: 1.4.1 2.1.3 I'll try again.\nI have used your script before with -b 8, I just haven't been able since yesterday. I am using the card for my display, so about .5GB are used, but not enough to keep me from using -b 8, there has to be something on my end I have messed up.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 273980,
          "author_name": "antorsae",
          "author_url": "",
          "post_date": "01/25/2018 16:05:32",
          "content": "<pre><code>$ vdir flickr_images/sony_nex7/|wc\n   1302   11711   98848\n</code></pre>\n\n<p>I did <code>find . -name \"urls_*\" -execdir wget -nc --tries=10 -i  {} \\;</code> on the <code>flick_images</code> directory b/c sometimes wget had to retry.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 273591,
      "author_name": "antorsae",
      "author_url": "",
      "post_date": "01/24/2018 21:19:39",
      "content": "<p>Update:  0.868 LB now with latest code changes and <code>python train.py -g 2 -b 16 -cs 512 -cm ResNet50 -x -l 1e-4 -uiw</code></p>",
      "votes": null,
      "replies": [
        {
          "id": 273628,
          "author_name": "shunjiading",
          "author_url": "",
          "post_date": "01/24/2018 22:02:15",
          "content": "<p>Awesome! What was the change compare to your last model?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 273637,
          "author_name": "antorsae",
          "author_url": "",
          "post_date": "01/24/2018 22:16:39",
          "content": "<p>I rewrote the generator to make it multiprocessor friendly. Architecturally is the same but I think there may have been a bug hidden in the old code. Also the new generator has more liberty in selecting random crops and now does orientation augmentation too. Will leave it overnight to see how far it converges.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 273704,
          "author_name": "wuzuping",
          "author_url": "",
          "post_date": "01/25/2018 02:35:14",
          "content": "<p>Well done, Andres. I have some question. Did you crop image to 512*512? Did you crop image from center or just randomly? how did you set your validation set? Thanks in advance~</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 273711,
          "author_name": "albertoa",
          "author_url": "",
          "post_date": "01/25/2018 02:53:59",
          "content": "<p>He crops around the center at 2 * patch_size. Then a random manipulation (or not) and a final crop.</p>\n\n<p>Validation for the above results seems to be based on gleb's additional data (see the -x flag)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 273718,
          "author_name": "albertoa",
          "author_url": "",
          "post_date": "01/25/2018 02:59:33",
          "content": "<p>The new code successfully kills (I mean gives a full workout ;-) ) all my CPU cores.</p>\n\n<p>After 1.5 hours I'm still at step 180 out of 980 of epoch 1. I'll give it 10-12 hours and see where it gets.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 273814,
          "author_name": "antorsae",
          "author_url": "",
          "post_date": "01/25/2018 08:10:02",
          "content": "<p>Here's the output for me:</p>\n\n<pre><code>              HTC-1-M7:  1023 (12.6%)\n              iPhone-6:   823 (10.1%)\n   Motorola-Droid-Maxx:   825 (10.2%)\n            Motorola-X:   275 (03.4%)\n     Samsung-Galaxy-S4:  1412 (17.4%)\n             iPhone-4s:   774 (09.5%)\n           LG-Nexus-5x:   680 (08.4%)\n      Motorola-Nexus-6:   926 (11.4%)\n  Samsung-Galaxy-Note3:   548 (06.8%)\n            Sony-NEX-7:   832 (10.2%)\nEpoch 59/200\n507/507 [==============================] - 231s 456ms/step - loss: 0.2178 - acc: 0.9339 - val_loss: 0.2700 - val_acc: 0.9243\n</code></pre>\n\n<p>Each epoch takes ~230 secs or less on my machine (2 x 1080 Tis + AMD Threadripper 1950X...). Re: samples, there's ~507*16 (batch size)= ~8112 samples per epoch with <code>-x</code>. Some of them are now discarded b/c of resolution.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 273865,
      "author_name": "antorsae",
      "author_url": "",
      "post_date": "01/25/2018 11:00:07",
      "content": "<p>Update, I left it overnight and it now achieves 0.934 LB with a single model. At this point I'm going to try a few new ideas. Kaggle competitions are usually won by huge ensembles and while \"everything goes\" I'm going to try to stick to simple stuff. Some ideas</p>\n\n<ul>\n<li>Class-aware sampling vs. class weighting</li>\n<li>Mixup</li>\n<li>New dataset</li>\n<li>Pretraining or faux-labeling</li>\n<li>Better TTA</li>\n</ul>",
      "votes": null,
      "replies": []
    },
    {
      "id": 274082,
      "author_name": "antorsae",
      "author_url": "",
      "post_date": "01/25/2018 21:29:58",
      "content": "<p>Update: LB 0.959 using single DenseNet201. Further discussion moved to: <a href=\"https://www.kaggle.com/c/sp-society-camera-model-identification/discussion/48293\">https://www.kaggle.com/c/sp-society-camera-model-identification/discussion/48293</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 274108,
          "author_name": "hamyadlab",
          "author_url": "",
          "post_date": "01/25/2018 22:08:32",
          "content": "<p>Very Good progress. Many thanks for sharing with us. I have one question: Is it very important to use Glep data? Because I did not notice much difference. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 274111,
          "author_name": "antorsae",
          "author_url": "",
          "post_date": "01/25/2018 22:24:24",
          "content": "<p>Good question. I will test without it.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 275872,
          "author_name": "hzywish",
          "author_url": "",
          "post_date": "01/30/2018 06:11:11",
          "content": "<p>The same question with HamYad. By the way, my validation accuracy is 80%+ but the LB score is only about 0.4, could you tell me how to improve?</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 274117,
      "author_name": "",
      "author_url": "",
      "post_date": "01/25/2018 22:37:27",
      "content": "<pre><code>Traceback (most recent call last):\n  File \"/home/ppleskov/anaconda3/lib/python3.6/threading.py\", line 916, in _bootstrap_inner\n    self.run()\n  File \"/home/ppleskov/anaconda3/lib/python3.6/threading.py\", line 864, in run\n    self._target(*self._args, **self._kwargs)\n  File \"/home/ppleskov/anaconda3/lib/python3.6/multiprocessing/pool.py\", line 429, in _handle_results\n    task = get()\n  File \"/home/ppleskov/anaconda3/lib/python3.6/multiprocessing/connection.py\", line 251, in recv\n    return _ForkingPickler.loads(buf.getbuffer())\nTypeError: __init__() missing 1 required positional argument: 'code'\n</code></pre>\n\n<p>anyone got this issue?</p>",
      "votes": null,
      "replies": [
        {
          "id": 274124,
          "author_name": "antorsae",
          "author_url": "",
          "post_date": "01/25/2018 23:10:50",
          "content": "<p>I think there's an exception in the <code>process_item</code> function and it's not easy to see the exception from the main thread. Use <code>--verbose</code> and put some prints around <code>img = load_img_fast_jpg(item)</code>. I bet some images fail to load.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 274127,
          "author_name": "antorsae",
          "author_url": "",
          "post_date": "01/25/2018 23:11:49",
          "content": "<p>I'd like know of an easier way to catch/see exceptions when using <code>Pool</code></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 274145,
          "author_name": "aamaia",
          "author_url": "",
          "post_date": "01/25/2018 23:42:56",
          "content": "<p>Yes, it's like Andres said above, the jpeg4py library can only load jpegs but in Gleb's dataset there are PNGs with the jpg extension.</p>\n\n<p>You could remove those files or use something like this:</p>\n\n<pre><code>def load_img_fast_jpg(img_path):\ntry:\n    x = jpeg.JPEG(img_path).decode()\n    return x\nexcept:\n    return load_img(img_path)\n</code></pre>",
          "votes": null,
          "replies": []
        },
        {
          "id": 274146,
          "author_name": "",
          "author_url": "",
          "post_date": "01/25/2018 23:46:37",
          "content": "<p>thanks a lot, that helped</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 278353,
          "author_name": "yyll008",
          "author_url": "",
          "post_date": "02/06/2018 01:05:20",
          "content": "<p>I run: python train.py -g 1 -b 8 -cs 229 -cm DenseNet201 -x -l 1e-4 -uiw</p>\n\n<p>The same error, I have try to fix it by add this:</p>\n\n<p>def load_img_fast_jpg(img_path):\ntry:\n    x = jpeg.JPEG(img_path).decode()\n    return x\nexcept:\n    return load_img(img_path)</p>\n\n<p>But not solve it.</p>\n\n<p>Exception in thread Thread-8:\nTraceback (most recent call last):\n  File \"/home/yl/miniconda3/envs/tensorflow/lib/python3.6/threading.py\", line 916, in _bootstrap_inner\n    self.run()\n  File \"/home/yl/miniconda3/envs/tensorflow/lib/python3.6/threading.py\", line 864, in run\n    self._target(*self._args, **self._kwargs)\n  File \"/home/yl/miniconda3/envs/tensorflow/lib/python3.6/multiprocessing/pool.py\", line 463, in _handle_results\n    task = get()\n  File \"/home/miniconda3/envs/tensorflow/lib/python3.6/multiprocessing/connection.py\", line 251, in recv\n    return _ForkingPickler.loads(buf.getbuffer())\nTypeError: <strong>init</strong>() missing 1 required positional argument: 'code'</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 278394,
          "author_name": "yyll008",
          "author_url": "",
          "post_date": "02/06/2018 03:56:05",
          "content": "<p>I add this code as follows, but not wok. How to fix?\nThanks.</p>\n\n<p>def load_img_fast_jpg(img_path):\n    try:\n        x = jpeg.JPEG(img_path).decode()\n        return x\n    except:\n        return load_img(img_path)</p>\n\n<p>load_img_fast_jpg  = lambda img_path: jpeg.JPEG(img_path).decode()\nload_img  = lambda img_path: np.array(Image.open(img_path))</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 274267,
      "author_name": "kirk86",
      "author_url": "",
      "post_date": "01/26/2018 07:42:54",
      "content": "<p>For me its stuck at this stage without any explicit information:</p>\n\n<pre><code>              HTC-1-M7:  1023 (13.0%)\n              iPhone-6:   823 (10.5%)\n              Motorola-Droid-Maxx:   825 (10.5%)\n              Motorola-X:   275 (03.5%)\n             Samsung-Galaxy-S4:  1412 (18.0%)\n             iPhone-4s:   774 (09.9%)\n             LG-Nexus-5x:   680 (08.7%)\n             Motorola-Nexus-6:   926 (11.8%)\n             Samsung-Galaxy-Note3:   548 (07.0%)\n             Sony-NEX-7:   557 (07.1%)\n             validation steps = 0\n             Epoch 1/200\n</code></pre>\n\n<p>No particular output, nothing, even with the <code>-v</code> flag.  Notice the <code>validation_steps=0</code> is the value that the <code>validation_steps</code> in the generator was getting without modifying the code, I had to hardcode that into sth like 6 otherwise I was getting errors. Any suggestions?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 274396,
      "author_name": "allenpeng0209",
      "author_url": "",
      "post_date": "01/26/2018 14:37:11",
      "content": "<p>ValueError: <code>validation_steps=None</code> is only valid for a generator based on the <code>keras.utils.Sequence</code> class. Please specify <code>validation_steps</code> or use the <code>keras.utils.Sequence</code> class.</p>\n\n<p>I also have some problem, when i look back to the code, i found it already used <code>keras.utils.Sequence</code> class, is there any one know how to reslove this? </p>\n\n<p>Thanks</p>",
      "votes": null,
      "replies": [
        {
          "id": 274420,
          "author_name": "kirk86",
          "author_url": "",
          "post_date": "01/26/2018 15:25:43",
          "content": "<p>I just hardcoded a value. For instance put a value of 6 and see if it works.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 274446,
          "author_name": "albertoa",
          "author_url": "",
          "post_date": "01/26/2018 16:06:50",
          "content": "<p>validation steps should be: number of validations images / batch size</p>\n\n<p>I added the following code to help me debug the image lists:</p>\n\n<pre><code>    print(\"\\nImage set counts\\n\")\nif args.extra_dataset:\n    print('{:&gt;15}: {:5d}'.format(\"extra_train_ids\",len(extra_train_ids)))\n    print('{:&gt;15}: {:5d}'.format(\"extra_val_ids\",len(extra_val_ids)))\n    print(\"\\n\")\nprint('{:&gt;15}: {:5d} steps={:5d}'.format(\n    \"ids_train\",len(ids_train),int(math.ceil(len(ids_train)  // args.batch_size))))\nprint('{:&gt;15}: {:5d} steps={:5d}'.format(\n    \"ids_val\",len(ids_val),int(math.ceil(len(ids_val) // args.batch_size))))\nprint('_' * 100)\nprint(\"\\n\")\n</code></pre>\n\n<p>(pasting code doesn't format well, pay attention to indentation)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 274460,
          "author_name": "kirk86",
          "author_url": "",
          "post_date": "01/26/2018 16:35:01",
          "content": "<p>May I ask if you have downloaded all the images? For me is the case I believe because I haven't downloaded all the images?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 274479,
          "author_name": "albertoa",
          "author_url": "",
          "post_date": "01/26/2018 17:16:39",
          "content": "<p>Yes, I downloaded all the images. It took me a couple of tries and had to make sure my paths were setup the way that the script expects them. You can change things around by changing the lines:</p>\n\n<pre><code>    extra_train_ids = [os.path.join(EXTRA_TRAIN_FOLDER,line.rstrip('\\n')) for line in open(os.path.join(EXTRA_TRAIN_FOLDER, 'good_jpgs'))]\n    extra_train_ids.sort()\n    ids_train.extend(extra_train_ids)\n\n    extra_val_ids = glob.glob(join(EXTRA_VAL_FOLDER,'*/*.jpg'))\n    extra_val_ids.sort()\n    ids_val.extend(extra_val_ids)\n</code></pre>\n\n<p>Then you can setup the extra images anywhere you want.</p>\n\n<p>These are my counts:</p>\n\n<p>Image set counts</p>\n\n<p>extra_train_ids:  6119\n  extra_val_ids:   316</p>\n\n<pre><code>  ids_train:  8594 steps= 8594\n    ids_val:   316 steps=  316\n</code></pre>\n\n<hr>\n\n<pre><code>          HTC-1-M7:  1024 (11.9%)\n          iPhone-6:   825 (09.6%)\n</code></pre>\n\n<p>Motorola-Droid-Maxx:   825 (09.6%)\n            Motorola-X:   275 (03.2%)\n     Samsung-Galaxy-S4:  1412 (16.4%)\n             iPhone-4s:   779 (09.1%)\n           LG-Nexus-5x:   680 (07.9%)\n      Motorola-Nexus-6:   926 (10.8%)\n  Samsung-Galaxy-Note3:   549 (06.4%)\n            Sony-NEX-7:  1299 (15.1%)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 274452,
      "author_name": "albertoa",
      "author_url": "",
      "post_date": "01/26/2018 16:19:07",
      "content": "<p>Andres, trying to upgrade tensorflow got my system unstable. I know you version of libraries, since we both use 1080ti can I ask you the version of the OS and the nvidia drivers?</p>\n\n<p>Now trying to run your script I only get:</p>\n\n<pre><code>E tensorflow/stream_executor/cuda/cuda_dnn.cc:385] could not create cudnn handle: CUDNN_STATUS_INTERNAL_ERROR\n</code></pre>\n\n<p>I may end up having to rebuild the system.</p>\n\n<p>Thanks in advance</p>",
      "votes": null,
      "replies": [
        {
          "id": 274534,
          "author_name": "antorsae",
          "author_url": "",
          "post_date": "01/26/2018 18:54:25",
          "content": "<p>Here you go:</p>\n\n<pre><code>$ nvidia-smi\nFri Jan 26 19:53:31 2018\n+-----------------------------------------------------------------------------+\n| NVIDIA-SMI 390.12                 Driver Version: 390.12                    |\n|-------------------------------+----------------------+----------------------+\n| GPU  Name        Persistence-M| Bus-Id        Disp.A | Volatile Uncorr. ECC |\n| Fan  Temp  Perf  Pwr:Usage/Cap|         Memory-Usage | GPU-Util  Compute M. |\n|===============================+======================+======================|\n|   0  GeForce GTX 108...  Off  | 00000000:09:00.0 Off |                  N/A |\n| 46%   67C    P2   252W / 280W |  10801MiB / 11178MiB |     96%      Default |\n+-------------------------------+----------------------+----------------------+\n|   1  GeForce GTX 108...  Off  | 00000000:41:00.0  On |                  N/A |\n| 54%   69C    P2   197W / 280W |  10821MiB / 11170MiB |     48%      Default |\n+-------------------------------+----------------------+----------------------+\n\n+-----------------------------------------------------------------------------+\n| Processes:                                                       GPU Memory |\n|  GPU       PID   Type   Process name                             Usage      |\n|=============================================================================|\n|    0     51088      C   python                                     10789MiB |\n|    1      1830      G   /usr/lib/xorg/Xorg                           336MiB |\n|    1      2283      G   compiz                                       249MiB |\n|    1     51088      C   python                                     10223MiB |\n+-----------------------------------------------------------------------------+\n$ uname -a\nLinux cineubuntu 4.13.0-31-generic #34~16.04.1-Ubuntu SMP Fri Jan 19 17:11:01 UTC 2018 x86_64 x86_64 x86_64 GNU/Linux\n</code></pre>\n\n<p>$ sudo apt list --installed | egrep \"(nvidia|cuda)\"</p>\n\n<p>WARNING: apt does not have a stable CLI interface. Use with caution in scripts.</p>\n\n<p>cuda/unknown,now 9.1.85-1 amd64 [installed]\ncuda-8-0/unknown,unknown,now 8.0.61-1 amd64 [installed,auto-removable]\ncuda-9-1/unknown,now 9.1.85-1 amd64 [installed,automatic]\ncuda-command-line-tools-8-0/unknown,unknown,now 8.0.61-1 amd64 [installed,auto-removable]\ncuda-command-line-tools-9-1/unknown,now 9.1.85-1 amd64 [installed,automatic]\ncuda-compiler-9-1/unknown,now 9.1.85-1 amd64 [installed,automatic]\ncuda-core-8-0/unknown,unknown,now 8.0.61-1 amd64 [installed,auto-removable]\ncuda-cublas-8-0/unknown,now 8.0.61.2-1 amd64 [installed,auto-removable]\ncuda-cublas-9-1/unknown,now 9.1.85-1 amd64 [installed,upgradable to: 9.1.85.1-1]\ncuda-cublas-dev-8-0/unknown,now 8.0.61.2-1 amd64 [installed,auto-removable]\ncuda-cublas-dev-9-1/unknown,now 9.1.85-1 amd64 [installed,upgradable to: 9.1.85.1-1]\ncuda-cudart-8-0/unknown,unknown,now 8.0.61-1 amd64 [installed,auto-removable]\ncuda-cudart-9-1/unknown,now 9.1.85-1 amd64 [installed,automatic]\ncuda-cudart-dev-8-0/unknown,unknown,now 8.0.61-1 amd64 [installed,auto-removable]\ncuda-cudart-dev-9-1/unknown,now 9.1.85-1 amd64 [installed,automatic]\ncuda-cufft-8-0/unknown,unknown,now 8.0.61-1 amd64 [installed,auto-removable]\ncuda-cufft-9-1/unknown,now 9.1.85-1 amd64 [installed,automatic]\ncuda-cufft-dev-8-0/unknown,unknown,now 8.0.61-1 amd64 [installed,auto-removable]\ncuda-cufft-dev-9-1/unknown,now 9.1.85-1 amd64 [installed,automatic]\ncuda-cuobjdump-9-1/unknown,now 9.1.85-1 amd64 [installed,automatic]\ncuda-cupti-9-1/unknown,now 9.1.85-1 amd64 [installed,automatic]\ncuda-curand-8-0/unknown,unknown,now 8.0.61-1 amd64 [installed,auto-removable]\ncuda-curand-9-1/unknown,now 9.1.85-1 amd64 [installed,automatic]\ncuda-curand-dev-8-0/unknown,unknown,now 8.0.61-1 amd64 [installed,auto-removable]\ncuda-curand-dev-9-1/unknown,now 9.1.85-1 amd64 [installed,automatic]\ncuda-cusolver-8-0/unknown,unknown,now 8.0.61-1 amd64 [installed,auto-removable]\ncuda-cusolver-9-1/unknown,now 9.1.85-1 amd64 [installed,automatic]\ncuda-cusolver-dev-8-0/unknown,unknown,now 8.0.61-1 amd64 [installed,auto-removable]\ncuda-cusolver-dev-9-1/unknown,now 9.1.85-1 amd64 [installed,automatic]\ncuda-cusparse-8-0/unknown,unknown,now 8.0.61-1 amd64 [installed,auto-removable]\ncuda-cusparse-9-1/unknown,now 9.1.85-1 amd64 [installed,automatic]\ncuda-cusparse-dev-8-0/unknown,unknown,now 8.0.61-1 amd64 [installed,auto-removable]\ncuda-cusparse-dev-9-1/unknown,now 9.1.85-1 amd64 [installed,automatic]\ncuda-demo-suite-8-0/unknown,unknown,now 8.0.61-1 amd64 [installed,auto-removable]\ncuda-demo-suite-9-1/unknown,now 9.1.85-1 amd64 [installed,automatic]\ncuda-documentation-8-0/unknown,unknown,now 8.0.61-1 amd64 [installed,auto-removable]\ncuda-documentation-9-1/unknown,now 9.1.85-1 amd64 [installed,automatic]\ncuda-driver-dev-8-0/unknown,unknown,now 8.0.61-1 amd64 [installed,auto-removable]\ncuda-driver-dev-9-1/unknown,now 9.1.85-1 amd64 [installed,automatic]\ncuda-drivers/unknown,now 390.12-1 amd64 [installed,automatic]\ncuda-gdb-9-1/unknown,now 9.1.85-1 amd64 [installed,automatic]\ncuda-gpu-library-advisor-9-1/unknown,now 9.1.85-1 amd64 [installed,automatic]\ncuda-libraries-9-1/unknown,now 9.1.85-1 amd64 [installed,automatic]\ncuda-libraries-dev-9-1/unknown,now 9.1.85-1 amd64 [installed,automatic]\ncuda-license-8-0/unknown,unknown,now 8.0.61-1 amd64 [installed,auto-removable]\ncuda-license-9-1/unknown,now 9.1.85-1 amd64 [installed,automatic]\ncuda-memcheck-9-1/unknown,now 9.1.85-1 amd64 [installed,automatic]\ncuda-misc-headers-8-0/unknown,unknown,now 8.0.61-1 amd64 [installed,auto-removable]\ncuda-misc-headers-9-1/unknown,now 9.1.85-1 amd64 [installed,automatic]\ncuda-npp-8-0/unknown,unknown,now 8.0.61-1 amd64 [installed,auto-removable]\ncuda-npp-9-1/unknown,now 9.1.85-1 amd64 [installed,automatic]\ncuda-npp-dev-8-0/unknown,unknown,now 8.0.61-1 amd64 [installed,auto-removable]\ncuda-npp-dev-9-1/unknown,now 9.1.85-1 amd64 [installed,automatic]\ncuda-nsight-9-1/unknown,now 9.1.85-1 amd64 [installed,automatic]\ncuda-nvcc-9-1/unknown,now 9.1.85-1 amd64 [installed,automatic]\ncuda-nvdisasm-9-1/unknown,now 9.1.85-1 amd64 [installed,automatic]\ncuda-nvgraph-8-0/unknown,unknown,now 8.0.61-1 amd64 [installed,auto-removable]\ncuda-nvgraph-9-1/unknown,now 9.1.85-1 amd64 [installed,automatic]\ncuda-nvgraph-dev-8-0/unknown,unknown,now 8.0.61-1 amd64 [installed,auto-removable]\ncuda-nvgraph-dev-9-1/unknown,now 9.1.85-1 amd64 [installed,automatic]\ncuda-nvml-dev-8-0/unknown,unknown,now 8.0.61-1 amd64 [installed,auto-removable]\ncuda-nvml-dev-9-1/unknown,now 9.1.85-1 amd64 [installed,automatic]\ncuda-nvprof-9-1/unknown,now 9.1.85-1 amd64 [installed,automatic]\ncuda-nvprune-9-1/unknown,now 9.1.85-1 amd64 [installed,automatic]\ncuda-nvrtc-8-0/unknown,unknown,now 8.0.61-1 amd64 [installed,auto-removable]\ncuda-nvrtc-9-1/unknown,now 9.1.85-1 amd64 [installed,automatic]\ncuda-nvrtc-dev-8-0/unknown,unknown,now 8.0.61-1 amd64 [installed,auto-removable]\ncuda-nvrtc-dev-9-1/unknown,now 9.1.85-1 amd64 [installed,automatic]\ncuda-nvtx-9-1/unknown,now 9.1.85-1 amd64 [installed,automatic]\ncuda-nvvp-9-1/unknown,now 9.1.85-1 amd64 [installed,automatic]\ncuda-repo-ubuntu1604/unknown,now 9.1.85-1 amd64 [installed]\ncuda-repo-ubuntu1604-8-0-local-ga2/now 8.0.61-1 amd64 [installed,local]\ncuda-runtime-8-0/unknown,unknown,now 8.0.61-1 amd64 [installed,auto-removable]\ncuda-runtime-9-1/unknown,now 9.1.85-1 amd64 [installed,automatic]\ncuda-samples-8-0/unknown,unknown,now 8.0.61-1 amd64 [installed,auto-removable]\ncuda-samples-9-1/unknown,now 9.1.85-1 amd64 [installed,automatic]\ncuda-toolkit-8-0/unknown,unknown,now 8.0.61-1 amd64 [installed,auto-removable]\ncuda-toolkit-9-1/unknown,now 9.1.85-1 amd64 [installed,automatic]\ncuda-tools-9-1/unknown,now 9.1.85-1 amd64 [installed,automatic]\ncuda-visual-tools-8-0/unknown,unknown,now 8.0.61-1 amd64 [installed,auto-removable]\ncuda-visual-tools-9-1/unknown,now 9.1.85-1 amd64 [installed,automatic]\nlibcuda1-375/unknown,now 390.12-0ubuntu1 amd64 [installed]\nlibcuda1-390/unknown,now 390.12-0ubuntu1 amd64 [installed,automatic]\nlibcudart7.5/xenial,now 7.5.18-0ubuntu1 amd64 [installed,automatic]\nnvidia-375/unknown,now 390.12-0ubuntu1 amd64 [installed,auto-removable]\nnvidia-375-dev/unknown,now 390.12-0ubuntu1 amd64 [installed]\nnvidia-390/unknown,now 390.12-0ubuntu1 amd64 [installed,automatic]\nnvidia-390-dev/unknown,now 390.12-0ubuntu1 amd64 [installed,automatic]\nnvidia-cuda-dev/xenial,now 7.5.18-0ubuntu1 amd64 [installed,automatic]\nnvidia-cuda-doc/xenial,xenial,now 7.5.18-0ubuntu1 all [installed,automatic]\nnvidia-cuda-gdb/xenial,now 7.5.18-0ubuntu1 amd64 [installed,automatic]\nnvidia-cuda-toolkit/xenial,now 7.5.18-0ubuntu1 amd64 [installed]\nnvidia-modprobe/unknown,now 390.12-0ubuntu1 amd64 [installed,automatic]\nnvidia-opencl-dev/xenial,now 7.5.18-0ubuntu1 amd64 [installed,automatic]\nnvidia-opencl-icd-375/unknown,now 390.12-0ubuntu1 amd64 [installed,auto-removable]\nnvidia-opencl-icd-390/unknown,now 390.12-0ubuntu1 amd64 [installed,automatic]\nnvidia-prime/xenial,now 0.8.2 amd64 [installed]\nnvidia-profiler/xenial,now 7.5.18-0ubuntu1 amd64 [installed,automatic]\nnvidia-settings/unknown,now 390.12-0ubuntu1 amd64 [installed,automatic]\nnvidia-visual-profiler/xenial,now 7.5.18-0ubuntu1 amd64 [installed,automatic]</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 274537,
          "author_name": "antorsae",
          "author_url": "",
          "post_date": "01/26/2018 19:04:18",
          "content": "<p>You can also consider TF optimized wheels:\n<a href=\"https://github.com/mind/wheels\">https://github.com/mind/wheels</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 274557,
          "author_name": "albertoa",
          "author_url": "",
          "post_date": "01/26/2018 19:38:46",
          "content": "<p>Thanks I may have to try wheels. Going to a newer tensorflow then cuda9 then the nvidia-390 drivers is what got my system unstable. I haven't been able to run your scripts since then. Since you are running that I now know it is possible.</p>\n\n<p>Are you using one of the 1080s for your display? Using the 390 drivers I did notice a huge lag on X and that's when I started to downgrade back.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 274560,
          "author_name": "antorsae",
          "author_url": "",
          "post_date": "01/26/2018 19:44:44",
          "content": "<p>Although I have cuda9 installed I use cuda8 too, which I what Im using. I have that computer plugged to a monitor which pretends it's on, but I always work via ssh / tmux with that machine, so I didn't really use the video cards in Linux as display much.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 275128,
      "author_name": "fabsta",
      "author_url": "",
      "post_date": "01/28/2018 11:39:58",
      "content": "<p>Hey Andres, \nthanks for the updates to your code!</p>\n\n<p>I am now getting the following error message:</p>\n\n<pre><code>  File \"train.py\", line 284, in process_item\nimg = np.rot90(_img, 1, (0,1))\nTypeError: rot90() takes from 1 to 2 positional arguments but 3 were given\n</code></pre>\n\n<p>Should that be np.rot90(1,(0,1)?\nThanks</p>",
      "votes": null,
      "replies": [
        {
          "id": 275135,
          "author_name": "antorsae",
          "author_url": "",
          "post_date": "01/28/2018 11:54:33",
          "content": "<p>What version of np are you running?\n<a href=\"https://docs.scipy.org/doc/numpy-1.13.0/reference/generated/numpy.rot90.html\">https://docs.scipy.org/doc/numpy-1.13.0/reference/generated/numpy.rot90.html</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 275147,
          "author_name": "fabsta",
          "author_url": "",
          "post_date": "01/28/2018 12:03:47",
          "content": "<p>Thanks for the quick reply. \nI was on 1.11, have upgraded to 1.14 and now it's working!</p>\n\n<p>Btw. I read you were using resnet50 quite a lot. Any success looking into other models?\nI guess even if other models don't quite reach resnet50's performance, they might be useful for ensembling.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 275634,
      "author_name": "trytrytrykaggle",
      "author_url": "",
      "post_date": "01/29/2018 17:14:10",
      "content": "<blockquote>\n  <p>Exception in thread Thread-1:\n  multiprocessing.pool.RemoteTraceback:\n  \"\"\"\n  Traceback (most recent call last):\n    File \"/usr/lib/python3.4/multiprocessing/pool.py\", line 119, in worker\n      result = (True, func(*args, **kwds))\n    File \"/usr/lib/python3.4/multiprocessing/pool.py\", line 44, in mapstar\n      return list(map(*args))\n    File \"train.py\", line 252, in process_item\n      img = load_img_fast_jpg(item)\n    File \"train.py\", line 148, in \n      load_img_fast_jpg  = lambda img_path: jpeg.JPEG(img_path).decode()\n    File \"/home/xiaodong_he/py3env/lib/python3.4/site-packages/jpeg4py/<em>py.py\", line 128,                                                                                                       in <strong>init</strong>\n      super(JPEG, self).<strong>init</strong>(lib</em>)\n    File \"/home/xiaodong_he/py3env/lib/python3.4/site-packages/jpeg4py/_py.py\", line 64,                                                                                                       in <strong>init</strong>\n      jpeg.initialize()\n    File \"/home/xiaodong_he/py3env/lib/python3.4/site-packages/jpeg4py/_cffi.py\", line 21                                                                                                      2, in initialize\n      _initialize(backends)\n    File \"/home/xiaodong_he/py3env/lib/python3.4/site-packages/jpeg4py/_cffi.py\", line 19                                                                                                      3, in _initialize\n      raise OSError(\"Could not load libjpeg-turbo library\")\n  OSError: Could not load libjpeg-turbo library\n  \"\"\"</p>\n</blockquote>\n\n<p>anyone got some idea about this error?thanks </p>",
      "votes": null,
      "replies": [
        {
          "id": 275636,
          "author_name": "kleinsmith",
          "author_url": "",
          "post_date": "01/29/2018 17:17:46",
          "content": "<p>You'll need to use <code>apt-get install libturbojpeg</code> or equivalent.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 275642,
          "author_name": "trytrytrykaggle",
          "author_url": "",
          "post_date": "01/29/2018 17:26:39",
          "content": "<p>it does work ,many thanks</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 277221,
      "author_name": "yyll008",
      "author_url": "",
      "post_date": "02/02/2018 15:12:41",
      "content": "<p>Awesome work! Thanks for your sharing!</p>\n\n<p>When run this code find this error:\nTraceback (most recent call last):\n  File \"train.py\", line 540, in \n    max_classes_val_count = max(classes_val_count)\nValueError: max() arg is an empty sequence</p>\n\n<p>I use ne GTX1070Ti, TF 1.4.1; keras 2.1.3; numpy 1.13.1</p>",
      "votes": null,
      "replies": [
        {
          "id": 277233,
          "author_name": "antorsae",
          "author_url": "",
          "post_date": "02/02/2018 15:55:42",
          "content": "<p>Looks your validation set is empty. Check whether you have downloaded val set accordingly.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 277381,
          "author_name": "yyll008",
          "author_url": "",
          "post_date": "02/02/2018 23:50:44",
          "content": "<p>I check it ,val_images folder is ok. Thanks for your help!</p>\n\n<p>Last time, I run:python train.py -g 1 -b 8 -cs 512 -cm ResNet50 -x -l 1e-4</p>\n\n<p>This time ,I  run:python train.py -g 1 -b 8 -cs 512 -cm ResNet50 -l 1e-4,\nerror:</p>\n\n<p>/keras/engine/training.py\", line 2053, in fit_generator\n    raise ValueError('<code>validation_steps=None</code> is only valid for a'\nValueError: <code>validation_steps=None</code> is only valid for a generator based on the <code>keras.utils.Sequence</code> class. Please specify <code>validation_steps</code> or use the <code>keras.utils.Sequence</code> class.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 277549,
          "author_name": "yyll008",
          "author_url": "",
          "post_date": "02/03/2018 14:38:49",
          "content": "<p>Brilliant!\nnot use -x , it works,  without extra dataset. I have put the floders(<a href=\"https://www.kaggle.com/c/sp-society-camera-model-identification/discussion/47235\">https://www.kaggle.com/c/sp-society-camera-model-identification/discussion/47235</a> )  in it , while not work.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 277551,
          "author_name": "igormunizims",
          "author_url": "",
          "post_date": "02/03/2018 14:44:31",
          "content": "<p>You have to download the extra data in order to use it</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 277706,
          "author_name": "yyll008",
          "author_url": "",
          "post_date": "02/04/2018 01:04:02",
          "content": "<p>extra dataset seem importance, without extra dataset (python train.py -g 1 -b 8 -cs 512 -cm ResNet50 -l 1e-4), epoch054-val_acc0.303427.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 277245,
      "author_name": "vivalavida1989",
      "author_url": "",
      "post_date": "02/02/2018 16:28:18",
      "content": "<p>Wow! You are simply a grand master! Thanks for sharing.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 278376,
      "author_name": "bopengiowa",
      "author_url": "",
      "post_date": "02/06/2018 02:40:51",
      "content": "<p>Really good job on the script. Did you design the CaCNN network just from reading the paper alone? </p>\n\n<p>Awesome work!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 379945,
      "author_name": "fahad924",
      "author_url": "",
      "post_date": "09/01/2018 09:19:34",
      "content": "<p>Sir, Thanks for sharing your code.I am currently working on Camera Model Identification for my undergrad project.  Can i train my own dataset in your given model ? if i can , what do i really have to do with the code. TIA</p>",
      "votes": null,
      "replies": [
        {
          "id": 380143,
          "author_name": "antorsae",
          "author_url": "",
          "post_date": "09/01/2018 19:53:49",
          "content": "<p>Sure, you can use. Code should be self-explanatory though.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "271358": "Hi guys,\n\nI just started yesterday and I'd like to share my progress:\n\nhttps://github.com/antorsae/sp-society-camera-model-identification\n\nCurrent approach:\n\n- Resizes train/val images to 512x512 (same as target test images)\n- Takes random crops located at the edges of the images (total 4 possible different crops per image) of size `-cs` (defaults to 299)\n- Since I believe the features learned by the classifier are location dependent (i.e. features on the top-left crop will be different than bottom-right crop) I concatenate the relative location of the crop to the features prior to the FC layer.\n- Uses any of the Keras applications invoked by command line (e.g. `-cm ResNet50` uses ResNet50)\n- Uses pooling as specified by `--pooling` i.e. avg|max|none\n- Optionally applies a kernel filter as per this Slide 13 @ http://www.lirmm.fr/~chaumont/publications/WIFS-2016_TUAMA_COMBY_CHAUMONT_Camera_Model_Identification_With_CNN_slides.pdf with `-kf`\n- Loading full-size JPGs takes a lot of time so on first iteration of the Keras generator it builds a dictionary in memory with the resize (and preprocessed) images so after first 4 epochs (4 is b/c `sub_batch_size`)  it goes significantly faster (on my 2 GPUs subsequent epoch take ~20 secs).\n\nSo far I only get ~0.87% train acc, and ~0.70% val acc. I have not submitted LB yet.\n\nAny comments / suggestions welcome!",
    "271445": "Hey Andres!\n\nNice to see you ;)\n\nWhy do you resize the train images? The test images are not resized, they are center cropped:\n\n&gt; While the train data includes full images, the test data contains only single 512 x 512 pixel blocks cropped from the center of a single image taken with the device.",
    "271448": "Hi Ali! yes, I just noticed... just pushed new code, much better now:\n\n`Epoch 29/200\n139/139 [==============================] - 53s 384ms/step - loss: 0.2487 - acc: 0.9101 - val_loss: 0.4792 - val_acc: 0.8750`",
    "272157": "I think it is not a good idea to resize the training images",
    "272178": "Agreed, I've updated the code and it doesn't resize anymore.",
    "272181": "Update:\n- No resizing, just takes 512x512 crops from the center\n\nThe good news is that it achieves ~90% val acc, however LB is ~30-40% only.  Currently investigating why.",
    "272213": "I have similar issue. I'm trying to resolve it with no success.  If I will resolve it somehow, I will let you know.\nMy model is ResNet50 finetuned with random crop from full size images.",
    "272216": "What size of your crop? I have the same issue.",
    "272217": "512x512, same as the test size.",
    "272235": "It seems the validation accuracy on its own does not mean much. I had a model with both training and validation accuracy above 98%, but the LB was only 64%. On the other hand, I had a model with train accuracy equal to 91 and validation accuracy equal to 95, and the LB accuracy for that model was 85.2. I am totally confused!",
    "272236": "Maybe I should reconsider to create a whole new validation set. I think this is the key point. Maybe the high number of submissions of top LBs is a confirmation! I think they have tried many models until they have found a good one.",
    "272242": "I find when I use the validation data from Gleb's post, my validation accuracy is close to my LB score. I trained a model to 0.79 on the Gleb validation set and got 0.804 on the public leaderboard.",
    "272300": "From my experience, I think you could get a LB score &gt; 0.90 just cropping to 112x112. That way you can train your models faster. Make some experiments until you get good parameters (LR, dense layers size, drop, augmentation). When you have a good one, scale up the cropping size to get better results.\nRegarding the validation/LB gap, just for reference, I got 0.957/0.922 (val/LB). Using additional data should help to reduce the gap.",
    "272308": "Hi Daniel, You are fine tuning some pre trained model or using your own model?",
    "272500": "Hi Daniel, did you use Gleb's dataset for your 0.922 LB or just the organization-provided dataset?",
    "272526": "Update:\nUsing Gleb's dataset I get 0.76 LB with a ResNet50 and pre-trained weights and a fixed 224 crop size.\n\nNext steps:\n\n - Balance classes (DONE)\n - Fine-tune FC, dropout, whether to use high-pass kernel filtering (`-kf`) and determine best architecture\n - Account for different weights (0.3 vs. 0.7) in loss function\n - Get more data (I think this is critical)",
    "272776": "Just the organization-provided dataset. Now using Gleb's dataset too.",
    "272780": "So the .76 is just with the built in Keras ResNet50? Since you mention fine-tuning the FC does that mean you didn't even replace the FC?",
    "272798": "I've already changed the code a lot but IIRC it was something like `python train.py -l 1e-3 -b 64 -g 2 -cm ResNet50 -cs 224 -x  -p avg -do 0.3  -kf`\n\nI added random crops to the current code so score may be different now. I've also added crop ensembling to `-test` so it may help.\nSummary of paramaters:\n - `-b 64` batch size for each GPU\n - `-g 2` use 2 GPUs\n - `-cm ResNet50` use ResNet50 as feature extractor\n - `-cs 224` crop size\n - `-x` use Gleb's dataset\n - `-p avg` use average pooling on the classifier (ResNet50)\n - `-do 0.3` use 0.3 dropout rate on my FC layers \n - `-kf` use a high pass pre-processing filter.\n\nYou may want to take a look at the code to try it out. Once you have a model with decent `val_acc` do `-m path_to_saved_model.hdf5 -t` to generate `submission.csv`",
    "272959": "I went ahead and got your code from github. \n\nI edited the post as the errors where caused by an unknown issue.\n\nI have successfully run your code and got results in the  0.503 LB using the flags above (except for the flickr data, which I still have to get properly)\n\nFor my submissions (0.517 LB) I was using a resnet like architecture with 3 stages training from scratch. It was taking 3-6 hours for training and your model is only taking me 1 hour. So now I have something better to look at.\n\nLooking at the CaCNN code and the paper you reference I guess I'll be running that next.\n\nThanks for sharing the code.",
    "273376": "Great. I am finishing a multiprocess generator which does not need to cache (memory or disk) and has the benefit of being able to select random crops from all across the image. Will be ready in a few hours. Stay tuned.",
    "273388": "Great, I look forward to looking at it.\n\nBy the way let me know if you would like to merge as a team. \n\nI've been using hdf5 files with unprocessed and pre-processed patches, speed wise they work well. I'm at the stage of testing multiple networks and improving results.",
    "273414": "Code is updated, now with multiprocess generator which is more flexible, especially wrt selecting random crops (previously only a smaller version of the original image was saved in the cache so cropping was more limited). \n\nI think you need a powerful CPU to do all augmentations, etc. and saturate your GPUs.",
    "273523": "The new code achieves ~0.73 just after 8 epochs:\n`$ python train.py -g 2 -b 16 -cs 512 -cm ResNet50 -x -l 1e-4 -uiw`\n\nEpoch 1/100\n507/507 [==============================] - 227s 447ms/step - loss: 1.6933 - acc: 0.4243 - val_loss: 1.6242 - val_acc: 0.5395\nEpoch 2/100\n507/507 [==============================] - 205s 405ms/step - loss: 1.0983 - acc: 0.6495 - val_loss: 1.1173 - val_acc: 0.6711\nEpoch 3/100\n507/507 [==============================] - 207s 408ms/step - loss: 0.8839 - acc: 0.7187 - val_loss: 1.6343 - val_acc: 0.5230\nEpoch 4/100\n507/507 [==============================] - 206s 406ms/step - loss: 0.7618 - acc: 0.7625 - val_loss: 0.9526 - val_acc: 0.7105\nEpoch 5/100\n507/507 [==============================] - 204s 403ms/step - loss: 0.6861 - acc: 0.7855 - val_loss: 1.0580 - val_acc: 0.6941\nEpoch 6/100\n507/507 [==============================] - 203s 399ms/step - loss: 0.6055 - acc: 0.8104 - val_loss: 0.7093 - val_acc: 0.7599\nEpoch 7/100\n507/507 [==============================] - 197s 389ms/step - loss: 0.5509 - acc: 0.8274 - val_loss: 1.1975 - val_acc: 0.6875\nEpoch 8/100\n507/507 [==============================] - 351s 693ms/step - loss: 0.5159 - acc: 0.8410 - val_loss: 0.8306 - val_acc: 0.8125\n\n(it then crashed due to a memory leak... need to investigate)\n\nthen:\n```$ python train.py -g 1 -b 16 -cs 512 -cm ResNet50 -x -l 1e-4 -uiw -m models/ResNet50_do0.3_avg-epoch008-val_acc0.812500.hdf5 -t\n```\n\nthen:\n```$ kg submit submission.csv\n0.733\n```\n\nI've noticed that in Gleb's dataset some images did not match any resolution, so I discard them. Also Sony-NEX-7 directory had files ending in JPG, not jpg... so my code was not using them. Renamed then.\n\nFinally I noted some images are portrait whereas others are landscape, so now I'm augmenting via orientation change during training.",
    "273550": "I'm still screwing up the execution somehow. Once again I got a run with no learning:\n\nEpoch 100/100\n245/245 [==============================] - 81s - loss: 2.2325 - acc: 0.1793 - val_loss: 2.3201 - val_acc: 0.1562\n\nI'm wiping out the whole models directory and restarted the training with:\n\n-l 1e-3 -b 32 -g 1 -cm ResNet50 -cs 224 -p avg -do 0.3 -x -kf -uiw\n\nWe'll see if this run it learns (I wasn't using the uiw before but I can't see how that would explain the model not learning). \n\nNot sure why you got JPG, all my links in flickr_images/sony_nex7/urls_final were lower case.",
    "273553": "Try changing (lowering) the learning rate. Also, intuitively I don't think `-uiw` would work OK with `-kf` the reason being is `-kf` applies a high-pass filter on the images and they cease to look anything like an image, and also the normalization is different than in imagenet (mean is similar but std and range are not).\n\nThe JPG vs. jpg was on the train directory (organization dataset).",
    "273591": "Update:  0.868 LB now with latest code changes and `python train.py -g 2 -b 16 -cs 512 -cm ResNet50 -x -l 1e-4 -uiw`",
    "273628": "Awesome! What was the change compare to your last model?",
    "273637": "I rewrote the generator to make it multiprocessor friendly. Architecturally is the same but I think there may have been a bug hidden in the old code. Also the new generator has more liberty in selecting random crops and now does orientation augmentation too. Will leave it overnight to see how far it converges.",
    "273652": "I can't seem to get it to run. When I saw this before my directories didn't match the script expectations, but this time I seem to have the right numbers on the image data:\n\n    Image set counts\n\nextra_train_ids:  5368\n  extra_val_ids:   315\n\n      ids_train:  8118 steps=  507\n        ids_val:   315 steps=   19\n____________________________________________________________________________________________________\n\n              HTC-1-M7:  1023 (12.6%)\n              iPhone-6:   823 (10.1%)\n   Motorola-Droid-Maxx:   825 (10.2%)\n            Motorola-X:   275 (03.4%)\n     Samsung-Galaxy-S4:  1412 (17.4%)\n             iPhone-4s:   774 (09.5%)\n           LG-Nexus-5x:   680 (08.4%)\n      Motorola-Nexus-6:   926 (11.4%)\n  Samsung-Galaxy-Note3:   548 (06.8%)\n            Sony-NEX-7:   832 (10.2%)\n\n\n    Exception in thread Thread-1:\nmultiprocessing.pool.RemoteTraceback: \n\"\"\"\nTraceback (most recent call last):\n  File \"/usr/lib/python3.5/multiprocessing/pool.py\", line 119, in worker\n    result = (True, func(*args, **kwds))\n  File \"/usr/lib/python3.5/multiprocessing/pool.py\", line 44, in mapstar\n    return list(map(*args))\n  File \"train.py\", line 265, in process_item\n    img = preprocess_image(img)\n  File \"train.py\", line 197, in preprocess_image\n    return preprocess_input_function(img.astype(np.float32))\n  File \"/home/alonsoa/projects/virtenv/py3cv3/lib/python3.5/site-packages/keras/applications/imagenet_utils.py\", line 33, in preprocess_input\n    x = x[:, :, :, ::-1]\nIndexError: too many indices for array\n\"\"\"\n\nAnything jumps at you why   return preprocess_input_function(img.astype(np.float32)) would fail even though the image is shaped (512, 512, 3)\"",
    "273655": "I had the same error. Sometimes the images don't download correctly and they read in as an array with less than three dimensions. I identified some of these images and downloaded them again, but had trouble getting them all. So I added logic that would check that len(img.shape) == 3 and if not move on to the next image.",
    "273657": "Thanks, unfortunately there already is:\n\n        if img.ndim != 3:\n        return None",
    "273704": "Well done, Andres. I have some question. Did you crop image to 512*512? Did you crop image from center or just randomly? how did you set your validation set? Thanks in advance~",
    "273711": "He crops around the center at 2 * patch_size. Then a random manipulation (or not) and a final crop.\n\nValidation for the above results seems to be based on gleb's additional data (see the -x flag)",
    "273713": "I ended up having to change the line to:\n\n        return preprocess_input_function(np.expand_dims(img.astype(np.float32), axis=0))\n\nas Keras' preprocess_input expects an array of images.",
    "273718": "The new code successfully kills (I mean gives a full workout ;-) ) all my CPU cores.\n\nAfter 1.5 hours I'm still at step 180 out of 980 of epoch 1. I'll give it 10-12 hours and see where it gets.",
    "273813": "What version of Keras are you running? Mine:\n\n    $ python -c 'import keras;print(keras.__version__)'\n    /home/antor/miniconda3/lib/python3.6/site-packages/h5py/__init__.py:36: FutureWarning: Conversion of the second argument of issubdtype from `float` to `np.floating` is deprecated. In future, it will be treated as `np.float64 == np.dtype(float).type`.\n      from ._conv import register_converters as _register_converters\n    Using TensorFlow backend.\n    2.1.3\n\nThere's slightly different preprocessing functions on keras.applications and the code selects the appropriate one for each of the networks in keras.applications.* and defaults to xception (which is just normalizing between -1 and 1 iirc) for all others).",
    "273814": "Here's the output for me:\n\n                  HTC-1-M7:  1023 (12.6%)\n                  iPhone-6:   823 (10.1%)\n       Motorola-Droid-Maxx:   825 (10.2%)\n                Motorola-X:   275 (03.4%)\n         Samsung-Galaxy-S4:  1412 (17.4%)\n                 iPhone-4s:   774 (09.5%)\n               LG-Nexus-5x:   680 (08.4%)\n          Motorola-Nexus-6:   926 (11.4%)\n      Samsung-Galaxy-Note3:   548 (06.8%)\n                Sony-NEX-7:   832 (10.2%)\n    Epoch 59/200\n    507/507 [==============================] - 231s 456ms/step - loss: 0.2178 - acc: 0.9339 - val_loss: 0.2700 - val_acc: 0.9243\n\nEach epoch takes ~230 secs or less on my machine (2 x 1080 Tis + AMD Threadripper 1950X...). Re: samples, there's ~507*16 (batch size)= ~8112 samples per epoch with `-x`. Some of them are now discarded b/c of resolution.",
    "273865": "Update, I left it overnight and it now achieves 0.934 LB with a single model. At this point I'm going to try a few new ideas. Kaggle competitions are usually won by huge ensembles and while \"everything goes\" I'm going to try to stick to simple stuff. Some ideas\n\n - Class-aware sampling vs. class weighting\n - Mixup\n - New dataset\n - Pretraining or faux-labeling\n - Better TTA",
    "273916": "I am running Keras (2.0.6) and I was running tensorflow (1.3.0)\n\nTurns out that last night I was getting CUDNN_STATUS_INTERNAL_ERROR which made me upgrade tensorflow to (1.4.1). After doing that I was no longer using the GPU and I didn't notice. That explains my previous comment about the CPU high usage, it was running the model. After upgrading tensorflow-gpu as well to (1.4.1) I'm back on the GPU (a GTX 1080Ti )\n\nAs for samples the only difference between us is your Sony-NEX-7:   832 (10.2%) vs. mine Sony-NEX-7:   557 (07.1%) so I'll double check which samples I'm missing.\n\nI'm still having some performance and learning issues though. I kept running out of memory unless I brought the batch size to 1\n\n     -g 1 -b 1 -cs 512 -cm ResNet50 -x -l 1e-4 -uiw --max-epoch 350\n\nEpoch 19/350\n7843/7843 [==============================] - 1159s - loss: 1.3477 - acc: 0.5560 - val_loss: 3.1935 - val_acc: 0.1108\n\nI'll upgrade keras and rerun\n\nThanks for all the help. I hope my info helps anybody else trying to reproduce the results.",
    "273921": "The discrepancy in Sony-NEX-7 is b/c in that folder the original files are .JPG and should be renamed to .jpg.\n\nRe: versions \n\n    $ python -c 'import tensorflow as tf;import keras;print(tf.__version__, keras.__version__)'\n    /home/antor/miniconda3/lib/python3.6/importlib/_bootstrap.py:219: RuntimeWarning: compiletime version 3.5 of module 'tensorflow.python.framework.fast_tensor_util' does not match runtime version 3.6\n      return f(*args, **kwds)\n    Using TensorFlow backend.\n    1.4.1 2.1.3\n\nRe: batch size you should be able to fit more in a 1080 Ti. I am testing the code on a different machine (AMD Ryzen 1600X + 1070 Ti) and this works w/o issue `$ python train.py -g 1 -b 8 -cs 512 -cm ResNet50 -x -l 1e-4`. Check with `nvidia-smi` that no memory is being used by the GPU before launching the script.",
    "273957": "As per your prior suggestion I already renamed the files. I think my downloads are missing some though:\n\n    vdir flickr_images/sony_nex7/|wc\n        558    6129   74089\n\nNow that I have the same versions: 1.4.1 2.1.3 I'll try again.\nI have used your script before with -b 8, I just haven't been able since yesterday. I am using the card for my display, so about .5GB are used, but not enough to keep me from using -b 8, there has to be something on my end I have messed up.",
    "273980": "$ vdir flickr_images/sony_nex7/|wc\n       1302   11711   98848\n\nI did `find . -name \"urls_*\" -execdir wget -nc --tries=10 -i  {} \\;` on the `flick_images` directory b/c sometimes wget had to retry.",
    "274082": "Update: LB 0.959 using single DenseNet201. Further discussion moved to: https://www.kaggle.com/c/sp-society-camera-model-identification/discussion/48293",
    "274108": "Very Good progress. Many thanks for sharing with us. I have one question: Is it very important to use Glep data? Because I did not notice much difference.",
    "274111": "Good question. I will test without it.",
    "274117": "Traceback (most recent call last):\n      File \"/home/ppleskov/anaconda3/lib/python3.6/threading.py\", line 916, in _bootstrap_inner\n        self.run()\n      File \"/home/ppleskov/anaconda3/lib/python3.6/threading.py\", line 864, in run\n        self._target(*self._args, **self._kwargs)\n      File \"/home/ppleskov/anaconda3/lib/python3.6/multiprocessing/pool.py\", line 429, in _handle_results\n        task = get()\n      File \"/home/ppleskov/anaconda3/lib/python3.6/multiprocessing/connection.py\", line 251, in recv\n        return _ForkingPickler.loads(buf.getbuffer())\n    TypeError: __init__() missing 1 required positional argument: 'code'\n\nanyone got this issue?",
    "274124": "I think there's an exception in the `process_item` function and it's not easy to see the exception from the main thread. Use `--verbose` and put some prints around `img = load_img_fast_jpg(item)`. I bet some images fail to load.",
    "274127": "I'd like know of an easier way to catch/see exceptions when using `Pool`",
    "274145": "Yes, it's like Andres said above, the jpeg4py library can only load jpegs but in Gleb's dataset there are PNGs with the jpg extension.\n\nYou could remove those files or use something like this:\n\n    def load_img_fast_jpg(img_path):\n    try:\n        x = jpeg.JPEG(img_path).decode()\n        return x\n    except:\n        return load_img(img_path)",
    "274146": "thanks a lot, that helped",
    "274267": "For me its stuck at this stage without any explicit information:\n\n                  HTC-1-M7:  1023 (13.0%)\n                  iPhone-6:   823 (10.5%)\n                  Motorola-Droid-Maxx:   825 (10.5%)\n                  Motorola-X:   275 (03.5%)\n                 Samsung-Galaxy-S4:  1412 (18.0%)\n                 iPhone-4s:   774 (09.9%)\n                 LG-Nexus-5x:   680 (08.7%)\n                 Motorola-Nexus-6:   926 (11.8%)\n                 Samsung-Galaxy-Note3:   548 (07.0%)\n                 Sony-NEX-7:   557 (07.1%)\n                 validation steps = 0\n                 Epoch 1/200\nNo particular output, nothing, even with the `-v` flag.  Notice the `validation_steps=0` is the value that the `validation_steps` in the generator was getting without modifying the code, I had to hardcode that into sth like 6 otherwise I was getting errors. Any suggestions?",
    "274396": "ValueError: `validation_steps=None` is only valid for a generator based on the `keras.utils.Sequence` class. Please specify `validation_steps` or use the `keras.utils.Sequence` class.\n\nI also have some problem, when i look back to the code, i found it already used `keras.utils.Sequence` class, is there any one know how to reslove this? \n\nThanks",
    "274420": "I just hardcoded a value. For instance put a value of 6 and see if it works.",
    "274446": "validation steps should be: number of validations images / batch size\n\nI added the following code to help me debug the image lists:\n\n        print(\"\\nImage set counts\\n\")\n    if args.extra_dataset:\n        print('{:&gt;15}: {:5d}'.format(\"extra_train_ids\",len(extra_train_ids)))\n        print('{:&gt;15}: {:5d}'.format(\"extra_val_ids\",len(extra_val_ids)))\n        print(\"\\n\")\n    print('{:&gt;15}: {:5d} steps={:5d}'.format(\n        \"ids_train\",len(ids_train),int(math.ceil(len(ids_train)  // args.batch_size))))\n    print('{:&gt;15}: {:5d} steps={:5d}'.format(\n        \"ids_val\",len(ids_val),int(math.ceil(len(ids_val) // args.batch_size))))\n    print('_' * 100)\n    print(\"\\n\")\n\n(pasting code doesn't format well, pay attention to indentation)",
    "274452": "Andres, trying to upgrade tensorflow got my system unstable. I know you version of libraries, since we both use 1080ti can I ask you the version of the OS and the nvidia drivers?\n\nNow trying to run your script I only get:\n\n    E tensorflow/stream_executor/cuda/cuda_dnn.cc:385] could not create cudnn handle: CUDNN_STATUS_INTERNAL_ERROR\n\nI may end up having to rebuild the system.\n\nThanks in advance",
    "274460": "May I ask if you have downloaded all the images? For me is the case I believe because I haven't downloaded all the images?",
    "274479": "Yes, I downloaded all the images. It took me a couple of tries and had to make sure my paths were setup the way that the script expects them. You can change things around by changing the lines:\n\n\n        extra_train_ids = [os.path.join(EXTRA_TRAIN_FOLDER,line.rstrip('\\n')) for line in open(os.path.join(EXTRA_TRAIN_FOLDER, 'good_jpgs'))]\n        extra_train_ids.sort()\n        ids_train.extend(extra_train_ids)\n\n        extra_val_ids = glob.glob(join(EXTRA_VAL_FOLDER,'*/*.jpg'))\n        extra_val_ids.sort()\n        ids_val.extend(extra_val_ids)\n\nThen you can setup the extra images anywhere you want.\n\nThese are my counts:\n\nImage set counts\n\nextra_train_ids:  6119\n  extra_val_ids:   316\n\n      ids_train:  8594 steps= 8594\n        ids_val:   316 steps=  316\n____________________________________________________________________________________________________\n\n              HTC-1-M7:  1024 (11.9%)\n              iPhone-6:   825 (09.6%)\n   Motorola-Droid-Maxx:   825 (09.6%)\n            Motorola-X:   275 (03.2%)\n     Samsung-Galaxy-S4:  1412 (16.4%)\n             iPhone-4s:   779 (09.1%)\n           LG-Nexus-5x:   680 (07.9%)\n      Motorola-Nexus-6:   926 (10.8%)\n  Samsung-Galaxy-Note3:   549 (06.4%)\n            Sony-NEX-7:  1299 (15.1%)",
    "274534": "Here you go:\n\n    $ nvidia-smi\n    Fri Jan 26 19:53:31 2018\n    +-----------------------------------------------------------------------------+\n    | NVIDIA-SMI 390.12                 Driver Version: 390.12                    |\n    |-------------------------------+----------------------+----------------------+\n    | GPU  Name        Persistence-M| Bus-Id        Disp.A | Volatile Uncorr. ECC |\n    | Fan  Temp  Perf  Pwr:Usage/Cap|         Memory-Usage | GPU-Util  Compute M. |\n    |===============================+======================+======================|\n    |   0  GeForce GTX 108...  Off  | 00000000:09:00.0 Off |                  N/A |\n    | 46%   67C    P2   252W / 280W |  10801MiB / 11178MiB |     96%      Default |\n    +-------------------------------+----------------------+----------------------+\n    |   1  GeForce GTX 108...  Off  | 00000000:41:00.0  On |                  N/A |\n    | 54%   69C    P2   197W / 280W |  10821MiB / 11170MiB |     48%      Default |\n    +-------------------------------+----------------------+----------------------+\n    \n    +-----------------------------------------------------------------------------+\n    | Processes:                                                       GPU Memory |\n    |  GPU       PID   Type   Process name                             Usage      |\n    |=============================================================================|\n    |    0     51088      C   python                                     10789MiB |\n    |    1      1830      G   /usr/lib/xorg/Xorg                           336MiB |\n    |    1      2283      G   compiz                                       249MiB |\n    |    1     51088      C   python                                     10223MiB |\n    +-----------------------------------------------------------------------------+\n    $ uname -a\n    Linux cineubuntu 4.13.0-31-generic #34~16.04.1-Ubuntu SMP Fri Jan 19 17:11:01 UTC 2018 x86_64 x86_64 x86_64 GNU/Linux\n$ sudo apt list --installed | egrep \"(nvidia|cuda)\"\n\nWARNING: apt does not have a stable CLI interface. Use with caution in scripts.\n\ncuda/unknown,now 9.1.85-1 amd64 [installed]\ncuda-8-0/unknown,unknown,now 8.0.61-1 amd64 [installed,auto-removable]\ncuda-9-1/unknown,now 9.1.85-1 amd64 [installed,automatic]\ncuda-command-line-tools-8-0/unknown,unknown,now 8.0.61-1 amd64 [installed,auto-removable]\ncuda-command-line-tools-9-1/unknown,now 9.1.85-1 amd64 [installed,automatic]\ncuda-compiler-9-1/unknown,now 9.1.85-1 amd64 [installed,automatic]\ncuda-core-8-0/unknown,unknown,now 8.0.61-1 amd64 [installed,auto-removable]\ncuda-cublas-8-0/unknown,now 8.0.61.2-1 amd64 [installed,auto-removable]\ncuda-cublas-9-1/unknown,now 9.1.85-1 amd64 [installed,upgradable to: 9.1.85.1-1]\ncuda-cublas-dev-8-0/unknown,now 8.0.61.2-1 amd64 [installed,auto-removable]\ncuda-cublas-dev-9-1/unknown,now 9.1.85-1 amd64 [installed,upgradable to: 9.1.85.1-1]\ncuda-cudart-8-0/unknown,unknown,now 8.0.61-1 amd64 [installed,auto-removable]\ncuda-cudart-9-1/unknown,now 9.1.85-1 amd64 [installed,automatic]\ncuda-cudart-dev-8-0/unknown,unknown,now 8.0.61-1 amd64 [installed,auto-removable]\ncuda-cudart-dev-9-1/unknown,now 9.1.85-1 amd64 [installed,automatic]\ncuda-cufft-8-0/unknown,unknown,now 8.0.61-1 amd64 [installed,auto-removable]\ncuda-cufft-9-1/unknown,now 9.1.85-1 amd64 [installed,automatic]\ncuda-cufft-dev-8-0/unknown,unknown,now 8.0.61-1 amd64 [installed,auto-removable]\ncuda-cufft-dev-9-1/unknown,now 9.1.85-1 amd64 [installed,automatic]\ncuda-cuobjdump-9-1/unknown,now 9.1.85-1 amd64 [installed,automatic]\ncuda-cupti-9-1/unknown,now 9.1.85-1 amd64 [installed,automatic]\ncuda-curand-8-0/unknown,unknown,now 8.0.61-1 amd64 [installed,auto-removable]\ncuda-curand-9-1/unknown,now 9.1.85-1 amd64 [installed,automatic]\ncuda-curand-dev-8-0/unknown,unknown,now 8.0.61-1 amd64 [installed,auto-removable]\ncuda-curand-dev-9-1/unknown,now 9.1.85-1 amd64 [installed,automatic]\ncuda-cusolver-8-0/unknown,unknown,now 8.0.61-1 amd64 [installed,auto-removable]\ncuda-cusolver-9-1/unknown,now 9.1.85-1 amd64 [installed,automatic]\ncuda-cusolver-dev-8-0/unknown,unknown,now 8.0.61-1 amd64 [installed,auto-removable]\ncuda-cusolver-dev-9-1/unknown,now 9.1.85-1 amd64 [installed,automatic]\ncuda-cusparse-8-0/unknown,unknown,now 8.0.61-1 amd64 [installed,auto-removable]\ncuda-cusparse-9-1/unknown,now 9.1.85-1 amd64 [installed,automatic]\ncuda-cusparse-dev-8-0/unknown,unknown,now 8.0.61-1 amd64 [installed,auto-removable]\ncuda-cusparse-dev-9-1/unknown,now 9.1.85-1 amd64 [installed,automatic]\ncuda-demo-suite-8-0/unknown,unknown,now 8.0.61-1 amd64 [installed,auto-removable]\ncuda-demo-suite-9-1/unknown,now 9.1.85-1 amd64 [installed,automatic]\ncuda-documentation-8-0/unknown,unknown,now 8.0.61-1 amd64 [installed,auto-removable]\ncuda-documentation-9-1/unknown,now 9.1.85-1 amd64 [installed,automatic]\ncuda-driver-dev-8-0/unknown,unknown,now 8.0.61-1 amd64 [installed,auto-removable]\ncuda-driver-dev-9-1/unknown,now 9.1.85-1 amd64 [installed,automatic]\ncuda-drivers/unknown,now 390.12-1 amd64 [installed,automatic]\ncuda-gdb-9-1/unknown,now 9.1.85-1 amd64 [installed,automatic]\ncuda-gpu-library-advisor-9-1/unknown,now 9.1.85-1 amd64 [installed,automatic]\ncuda-libraries-9-1/unknown,now 9.1.85-1 amd64 [installed,automatic]\ncuda-libraries-dev-9-1/unknown,now 9.1.85-1 amd64 [installed,automatic]\ncuda-license-8-0/unknown,unknown,now 8.0.61-1 amd64 [installed,auto-removable]\ncuda-license-9-1/unknown,now 9.1.85-1 amd64 [installed,automatic]\ncuda-memcheck-9-1/unknown,now 9.1.85-1 amd64 [installed,automatic]\ncuda-misc-headers-8-0/unknown,unknown,now 8.0.61-1 amd64 [installed,auto-removable]\ncuda-misc-headers-9-1/unknown,now 9.1.85-1 amd64 [installed,automatic]\ncuda-npp-8-0/unknown,unknown,now 8.0.61-1 amd64 [installed,auto-removable]\ncuda-npp-9-1/unknown,now 9.1.85-1 amd64 [installed,automatic]\ncuda-npp-dev-8-0/unknown,unknown,now 8.0.61-1 amd64 [installed,auto-removable]\ncuda-npp-dev-9-1/unknown,now 9.1.85-1 amd64 [installed,automatic]\ncuda-nsight-9-1/unknown,now 9.1.85-1 amd64 [installed,automatic]\ncuda-nvcc-9-1/unknown,now 9.1.85-1 amd64 [installed,automatic]\ncuda-nvdisasm-9-1/unknown,now 9.1.85-1 amd64 [installed,automatic]\ncuda-nvgraph-8-0/unknown,unknown,now 8.0.61-1 amd64 [installed,auto-removable]\ncuda-nvgraph-9-1/unknown,now 9.1.85-1 amd64 [installed,automatic]\ncuda-nvgraph-dev-8-0/unknown,unknown,now 8.0.61-1 amd64 [installed,auto-removable]\ncuda-nvgraph-dev-9-1/unknown,now 9.1.85-1 amd64 [installed,automatic]\ncuda-nvml-dev-8-0/unknown,unknown,now 8.0.61-1 amd64 [installed,auto-removable]\ncuda-nvml-dev-9-1/unknown,now 9.1.85-1 amd64 [installed,automatic]\ncuda-nvprof-9-1/unknown,now 9.1.85-1 amd64 [installed,automatic]\ncuda-nvprune-9-1/unknown,now 9.1.85-1 amd64 [installed,automatic]\ncuda-nvrtc-8-0/unknown,unknown,now 8.0.61-1 amd64 [installed,auto-removable]\ncuda-nvrtc-9-1/unknown,now 9.1.85-1 amd64 [installed,automatic]\ncuda-nvrtc-dev-8-0/unknown,unknown,now 8.0.61-1 amd64 [installed,auto-removable]\ncuda-nvrtc-dev-9-1/unknown,now 9.1.85-1 amd64 [installed,automatic]\ncuda-nvtx-9-1/unknown,now 9.1.85-1 amd64 [installed,automatic]\ncuda-nvvp-9-1/unknown,now 9.1.85-1 amd64 [installed,automatic]\ncuda-repo-ubuntu1604/unknown,now 9.1.85-1 amd64 [installed]\ncuda-repo-ubuntu1604-8-0-local-ga2/now 8.0.61-1 amd64 [installed,local]\ncuda-runtime-8-0/unknown,unknown,now 8.0.61-1 amd64 [installed,auto-removable]\ncuda-runtime-9-1/unknown,now 9.1.85-1 amd64 [installed,automatic]\ncuda-samples-8-0/unknown,unknown,now 8.0.61-1 amd64 [installed,auto-removable]\ncuda-samples-9-1/unknown,now 9.1.85-1 amd64 [installed,automatic]\ncuda-toolkit-8-0/unknown,unknown,now 8.0.61-1 amd64 [installed,auto-removable]\ncuda-toolkit-9-1/unknown,now 9.1.85-1 amd64 [installed,automatic]\ncuda-tools-9-1/unknown,now 9.1.85-1 amd64 [installed,automatic]\ncuda-visual-tools-8-0/unknown,unknown,now 8.0.61-1 amd64 [installed,auto-removable]\ncuda-visual-tools-9-1/unknown,now 9.1.85-1 amd64 [installed,automatic]\nlibcuda1-375/unknown,now 390.12-0ubuntu1 amd64 [installed]\nlibcuda1-390/unknown,now 390.12-0ubuntu1 amd64 [installed,automatic]\nlibcudart7.5/xenial,now 7.5.18-0ubuntu1 amd64 [installed,automatic]\nnvidia-375/unknown,now 390.12-0ubuntu1 amd64 [installed,auto-removable]\nnvidia-375-dev/unknown,now 390.12-0ubuntu1 amd64 [installed]\nnvidia-390/unknown,now 390.12-0ubuntu1 amd64 [installed,automatic]\nnvidia-390-dev/unknown,now 390.12-0ubuntu1 amd64 [installed,automatic]\nnvidia-cuda-dev/xenial,now 7.5.18-0ubuntu1 amd64 [installed,automatic]\nnvidia-cuda-doc/xenial,xenial,now 7.5.18-0ubuntu1 all [installed,automatic]\nnvidia-cuda-gdb/xenial,now 7.5.18-0ubuntu1 amd64 [installed,automatic]\nnvidia-cuda-toolkit/xenial,now 7.5.18-0ubuntu1 amd64 [installed]\nnvidia-modprobe/unknown,now 390.12-0ubuntu1 amd64 [installed,automatic]\nnvidia-opencl-dev/xenial,now 7.5.18-0ubuntu1 amd64 [installed,automatic]\nnvidia-opencl-icd-375/unknown,now 390.12-0ubuntu1 amd64 [installed,auto-removable]\nnvidia-opencl-icd-390/unknown,now 390.12-0ubuntu1 amd64 [installed,automatic]\nnvidia-prime/xenial,now 0.8.2 amd64 [installed]\nnvidia-profiler/xenial,now 7.5.18-0ubuntu1 amd64 [installed,automatic]\nnvidia-settings/unknown,now 390.12-0ubuntu1 amd64 [installed,automatic]\nnvidia-visual-profiler/xenial,now 7.5.18-0ubuntu1 amd64 [installed,automatic]",
    "274537": "You can also consider TF optimized wheels:\nhttps://github.com/mind/wheels",
    "274557": "Thanks I may have to try wheels. Going to a newer tensorflow then cuda9 then the nvidia-390 drivers is what got my system unstable. I haven't been able to run your scripts since then. Since you are running that I now know it is possible.\n\nAre you using one of the 1080s for your display? Using the 390 drivers I did notice a huge lag on X and that's when I started to downgrade back.",
    "274560": "Although I have cuda9 installed I use cuda8 too, which I what Im using. I have that computer plugged to a monitor which pretends it's on, but I always work via ssh / tmux with that machine, so I didn't really use the video cards in Linux as display much.",
    "275128": "Hey Andres, \nthanks for the updates to your code!\n\nI am now getting the following error message:\n\n      File \"train.py\", line 284, in process_item\n    img = np.rot90(_img, 1, (0,1))\n    TypeError: rot90() takes from 1 to 2 positional arguments but 3 were given\n\nShould that be np.rot90(1,(0,1)?\nThanks",
    "275135": "What version of np are you running?\nhttps://docs.scipy.org/doc/numpy-1.13.0/reference/generated/numpy.rot90.html",
    "275147": "Thanks for the quick reply. \nI was on 1.11, have upgraded to 1.14 and now it's working!\n\nBtw. I read you were using resnet50 quite a lot. Any success looking into other models?\nI guess even if other models don't quite reach resnet50's performance, they might be useful for ensembling.",
    "275634": "&gt; Exception in thread Thread-1:\nmultiprocessing.pool.RemoteTraceback:\n\"\"\"\nTraceback (most recent call last):\n  File \"/usr/lib/python3.4/multiprocessing/pool.py\", line 119, in worker\n    result = (True, func(*args, **kwds))\n  File \"/usr/lib/python3.4/multiprocessing/pool.py\", line 44, in mapstar\n    return list(map(*args))\n  File \"train.py\", line 252, in process_item\n    img = load_img_fast_jpg(item)\n  File \"train.py\", line 148, in",
    "275636": "You'll need to use `apt-get install libturbojpeg` or equivalent.",
    "275642": "it does work ,many thanks",
    "275872": "The same question with HamYad. By the way, my validation accuracy is 80%+ but the LB score is only about 0.4, could you tell me how to improve?",
    "275878": "Excuse me. Do you mean that you trained your model by all the training data and validated the model by the validation data from Gleb's post?",
    "275927": "Hi Daniel,  when you cropping train set to 112x112 and use it to train model, what do you use for test set, which size is 512x512? You just resize it from 512 to 112 or may be you get crop from test? Thanks!",
    "275946": "I find out, that resizing is bad practice.... Getting crop of test size is much better. May be it help somebody :)",
    "275954": "Hi Alex. You're looking for image noise patterns. Particularly, it looks like camera model create some kind of grid fingerprint. Resizing (scaling) from 512x512 to 112x112 doesn't lost much information about image content, but it does about noise.  From my side, I get worst predictions for jpg70 and scale0.5 altered images.",
    "275959": "Daniel, how do you quantify worst predictions: looking at the soft-probabilities (softmax) below a threshold, e.g. 0.7?\n\nIn my case I'm doing that I have worst predictions for Nexus-5.",
    "275963": "Cross validation. Same here, LG Nexus 5x and Motorola Nexus 6.",
    "276139": "I'm having issues with those two as well, but also the iphone-4s which for me is the worst.",
    "276142": "Look at the data you should, young jedi. :-)",
    "276227": "Hi Andres! My model also achieves ~90% val acc, but LB is ~30-40% only. But I can't find out a solution. Could you tell me how to deal with it?",
    "276279": "Maybe you should use proper preprocessing(e.g. jpeg compression ) to training set and validation set.",
    "276296": "Test data is different than train in 3 big ways (then there's smaller issues);\n\n 1. 50% of it is manipulated as per the instructions\n 2. Taken from a different device (BIG issue)\n 3. Center-cropped, fixed res.\n\nYou need to address all 3.",
    "276670": "Thanks Andres. I've looked into my preprocession. Now it seems I'm going on a right way. Thanks very much!",
    "277221": "Awesome work! Thanks for your sharing!\n\n\nWhen run this code find this error:\nTraceback (most recent call last):\n  File \"train.py\", line 540, in",
    "277233": "Looks your validation set is empty. Check whether you have downloaded val set accordingly.",
    "277245": "Wow! You are simply a grand master! Thanks for sharing.",
    "277381": "I check it ,val_images folder is ok. Thanks for your help!\n\nLast time, I run:python train.py -g 1 -b 8 -cs 512 -cm ResNet50 -x -l 1e-4\n\nThis time ,I  run:python train.py -g 1 -b 8 -cs 512 -cm ResNet50 -l 1e-4,\nerror:\n\n/keras/engine/training.py\", line 2053, in fit_generator\n    raise ValueError('`validation_steps=None` is only valid for a'\nValueError: `validation_steps=None` is only valid for a generator based on the `keras.utils.Sequence` class. Please specify `validation_steps` or use the `keras.utils.Sequence` class.",
    "277549": "Brilliant!\nnot use -x , it works,  without extra dataset. I have put the floders(https://www.kaggle.com/c/sp-society-camera-model-identification/discussion/47235 )  in it , while not work.",
    "277551": "You have to download the extra data in order to use it",
    "277706": "extra dataset seem importance, without extra dataset (python train.py -g 1 -b 8 -cs 512 -cm ResNet50 -l 1e-4), epoch054-val_acc0.303427.",
    "278353": "I run: python train.py -g 1 -b 8 -cs 229 -cm DenseNet201 -x -l 1e-4 -uiw\n\nThe same error, I have try to fix it by add this:\n\ndef load_img_fast_jpg(img_path):\ntry:\n    x = jpeg.JPEG(img_path).decode()\n    return x\nexcept:\n    return load_img(img_path)\n\nBut not solve it.\n\nException in thread Thread-8:\nTraceback (most recent call last):\n  File \"/home/yl/miniconda3/envs/tensorflow/lib/python3.6/threading.py\", line 916, in _bootstrap_inner\n    self.run()\n  File \"/home/yl/miniconda3/envs/tensorflow/lib/python3.6/threading.py\", line 864, in run\n    self._target(*self._args, **self._kwargs)\n  File \"/home/yl/miniconda3/envs/tensorflow/lib/python3.6/multiprocessing/pool.py\", line 463, in _handle_results\n    task = get()\n  File \"/home/miniconda3/envs/tensorflow/lib/python3.6/multiprocessing/connection.py\", line 251, in recv\n    return _ForkingPickler.loads(buf.getbuffer())\nTypeError: __init__() missing 1 required positional argument: 'code'",
    "278376": "Really good job on the script. Did you design the CaCNN network just from reading the paper alone? \n\nAwesome work!",
    "278394": "I add this code as follows, but not wok. How to fix?\nThanks.\n\n\ndef load_img_fast_jpg(img_path):\n    try:\n        x = jpeg.JPEG(img_path).decode()\n        return x\n    except:\n        return load_img(img_path)\n\nload_img_fast_jpg  = lambda img_path: jpeg.JPEG(img_path).decode()\nload_img  = lambda img_path: np.array(Image.open(img_path))",
    "379945": "Sir, Thanks for sharing your code.I am currently working on Camera Model Identification for my undergrad project.  Can i train my own dataset in your given model ? if i can , what do i really have to do with the code. TIA",
    "380143": "Sure, you can use. Code should be self-explanatory though."
  },
  "source": "meta"
}